diegomrsantos opened a new issue, #4232:
URL: https://github.com/apache/iggy/issues/4232

   I propose tracking the remaining sink delivery work under one agreed 
contract, with separate implementation issues and recovery tests as acceptance 
criteria.
   
   The runtime can save a consumer offset before the sink delivers the 
corresponding records. If delivery fails, restarting with the same consumer 
group can skip those records. They may remain in Iggy's retained log, but 
normal recovery does not deliver them.
   
   A local HTTP regression test reproduces this sequence: the endpoint rejects 
a record with HTTP 503, the runtime restarts with the same consumer group after 
the endpoint recovers, and a later record arrives while the rejected record 
does not. This test covers rejection followed by restart; abrupt termination 
needs additional coverage.
   
   [Issue #2928](https://github.com/apache/iggy/issues/2928) and discussions 
[#3203](https://github.com/apache/iggy/discussions/3203) and 
[#3518](https://github.com/apache/iggy/discussions/3518) already describe the 
underlying problem and proposed approaches. [PR 
#3954](https://github.com/apache/iggy/pull/3954), by kriti-sc, adds control 
over commit timing and provides implementation work to build on. [PR 
#4152](https://github.com/apache/iggy/pull/4152) has since improved failure 
reporting and accounting, but explicitly does not introduce batch replay. The 
status of [#2927](https://github.com/apache/iggy/issues/2927) should be 
reconciled with that merged work.
   
   The proposed contract is:
   
   > For each partition, the runtime advances the recovery checkpoint only 
after every record it passes has satisfied the configured delivery contract, 
including all required destination writes. Failed or uncertain delivery remains 
recoverable. Recovery may produce duplicates.
   
   Acceptance into a plugin's memory buffer must be distinct from durable 
delivery. Unexpected decode, transform, and serialization failures must remain 
unresolved until recovered or handled by an explicit policy. A configured 
filter is a separate outcome. If a dead-letter policy is supported, its write 
must succeed before the checkpoint advances, and the resulting guarantee must 
name that alternative destination.
   
   Each supported sink configuration must define what its delivery 
acknowledgment proves. The contract also depends on source durability, 
sufficient retention, and eventual recovery. Missing replay data must produce a 
visible recovery failure.
   
   Proposed work:
   
   - [ ] A. Define the sink delivery and recovery contract.
   - [ ] B. Implement explicit sink delivery acknowledgments and safe 
checkpoint advancement.
   - [ ] C. Preserve failed records throughout the sink pipeline.
   - [ ] D. Migrate sink plugins and document supported configurations.
   - [ ] E. Add a shared recovery conformance suite.
   - [ ] F. Introduce safe defaults, migration guidance, and operational 
visibility.
   
   Existing issues and PRs should be linked to these work items where their 
scope matches. The test suite should begin alongside the contract work, with 
regression coverage included in each implementation change.
   
   This tracker is complete when the agreed delivery contract is enforced for 
every configuration advertised as supporting it, incompatible configurations 
fail explicitly, required recovery tests pass in CI, and the migration and 
operational documentation are available.
   
   Maintainer feedback is requested on the target contract, the initial set of 
supported sinks, the interface migration, and how to continue the work in #3954.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to