jphjsoares commented on issue #3793: URL: https://github.com/apache/iggy/issues/3793#issuecomment-5627582770
Hey @krishvishal, I was able to reproduce it, so I don't think it's fixed yet (even though I saw some promising PRs getting merged the last month). Repro: `cargo nextest run -p integration --retries 0 -E "test(/should_purge_topic_and_clear_consumer_offsets.*restart_on/)"` failed 1 in 79 runs for me on `d693eb93d`. Two setup notes: `--features vsr` no longer exists (I ran the bare command), and the test is currently pinned to `cluster_nodes = 1`, which can't hit the repair path — I ran it with a local override back to 3 nodes. The failure looks like what's reported here: the surviving peer served the full pre-purge window (`from_op=1`, served through 27, no eviction notice), the restarted node completed the repair, and the partition came back with all segments plus both consumer offsets at their old values. My understanding is still shallow, but it looks like the case left open by the TODO in https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/server/src/partition_reconciler.rs#L1424 — the restarted node had already purged (so its own gate was open) with a zeroed floor (it isn't persisted), and the serving peer hadn't applied the purge commit yet, so nothing "fenced" the replay. Before I attempt a fix, I'd appreciate some guidance: would persisting the purge floor plus some generation check on the repair path be a reasonable direction, or is the checkpoint barrier the way this should be solved? It definitely sounds like the most "correct" solution. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
