diegomrsantos opened a new issue, #4232: URL: https://github.com/apache/iggy/issues/4232
I propose tracking the remaining sink delivery work under one agreed contract, with separate implementation issues and recovery tests as acceptance criteria. The runtime can save a consumer offset before the sink delivers the corresponding records. If delivery fails, restarting with the same consumer group can skip those records. They may remain in Iggy's retained log, but normal recovery does not deliver them. A local HTTP regression test reproduces this sequence: the endpoint rejects a record with HTTP 503, the runtime restarts with the same consumer group after the endpoint recovers, and a later record arrives while the rejected record does not. This test covers rejection followed by restart; abrupt termination needs additional coverage. [Issue #2928](https://github.com/apache/iggy/issues/2928) and discussions [#3203](https://github.com/apache/iggy/discussions/3203) and [#3518](https://github.com/apache/iggy/discussions/3518) already describe the underlying problem and proposed approaches. [PR #3954](https://github.com/apache/iggy/pull/3954), by kriti-sc, adds control over commit timing and provides implementation work to build on. [PR #4152](https://github.com/apache/iggy/pull/4152) has since improved failure reporting and accounting, but explicitly does not introduce batch replay. The status of [#2927](https://github.com/apache/iggy/issues/2927) should be reconciled with that merged work. The proposed contract is: > For each partition, the runtime advances the recovery checkpoint only after every record it passes has satisfied the configured delivery contract, including all required destination writes. Failed or uncertain delivery remains recoverable. Recovery may produce duplicates. Acceptance into a plugin's memory buffer must be distinct from durable delivery. Unexpected decode, transform, and serialization failures must remain unresolved until recovered or handled by an explicit policy. A configured filter is a separate outcome. If a dead-letter policy is supported, its write must succeed before the checkpoint advances, and the resulting guarantee must name that alternative destination. Each supported sink configuration must define what its delivery acknowledgment proves. The contract also depends on source durability, sufficient retention, and eventual recovery. Missing replay data must produce a visible recovery failure. Proposed work: - [ ] A. Define the sink delivery and recovery contract. - [ ] B. Implement explicit sink delivery acknowledgments and safe checkpoint advancement. - [ ] C. Preserve failed records throughout the sink pipeline. - [ ] D. Migrate sink plugins and document supported configurations. - [ ] E. Add a shared recovery conformance suite. - [ ] F. Introduce safe defaults, migration guidance, and operational visibility. Existing issues and PRs should be linked to these work items where their scope matches. The test suite should begin alongside the contract work, with regression coverage included in each implementation change. This tracker is complete when the agreed delivery contract is enforced for every configuration advertised as supporting it, incompatible configurations fail explicitly, required recovery tests pass in CI, and the migration and operational documentation are available. Maintainer feedback is requested on the target contract, the initial set of supported sinks, the interface migration, and how to continue the work in #3954. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
