jphjsoares commented on issue #3793:
URL: https://github.com/apache/iggy/issues/3793#issuecomment-5627582770

   Hey @krishvishal, I was able to reproduce it, so I don't think it's fixed 
yet (even though I saw some promising PRs getting merged the last month).
   
   Repro: `cargo nextest run -p integration --retries 0 -E 
"test(/should_purge_topic_and_clear_consumer_offsets.*restart_on/)"` failed 1 
in 79 runs for me on `d693eb93d`. Two setup notes: `--features vsr` no longer 
exists (I ran the bare command), and the test is currently pinned to 
`cluster_nodes = 1`, which can't hit the repair path — I ran it with a local 
override back to 3 nodes.
   
   The failure looks like what's reported here: the surviving peer served the 
full pre-purge window (`from_op=1`, served through 27, no eviction notice), the 
restarted node completed the repair, and the partition came back with all 
segments plus both consumer offsets at their old values.
   
   My understanding is still shallow, but it looks like the case left open by 
the TODO in 
https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/server/src/partition_reconciler.rs#L1424
 — the restarted node had already purged (so its own gate was open) with a 
zeroed floor (it isn't persisted), and the serving peer hadn't applied the 
purge commit yet, so nothing "fenced" the replay.
   
   Before I attempt a fix, I'd appreciate some guidance: would persisting the 
purge floor plus some generation check on the repair path be a reasonable 
direction, or is the checkpoint barrier the way this should be solved? It 
definitely sounds like the most "correct" solution.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to