krishvishal commented on issue #3793:
URL: https://github.com/apache/iggy/issues/3793#issuecomment-5630051088

   Thanks for reproducing it, and for the `--features vsr` / `cluster_nodes = 
1` notes. I'll fix the repro in the issue.
   
   Yeah, that's pretty much what's going on. One thing to add: the [serve-side 
check](https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/shard/src/lib.rs#L4884)
 only kicks in when the serving peer's metadata already has the purge 
committed. In your run it didn't have it yet. That's why the peer had floor 0 
and served from op 1. The [receive-side 
check](https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/shard/src/lib.rs#L5413)
 on the restarted node just compares its own committed vs applied generation, 
which were equal, and its floor was back to 0 after the restart. Neither node 
has any idea what the other one applied.
   
   I don't think persisting the floor + a generation check is the way to go:
   
   1. The floor is whatever [the local op 
counter](https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/partitions/src/iggy_partition.rs#L6636)
 happens to be when that node's reconciler runs. Two replicas on the same 
generation can have different floors, restart or not. Saving it to disk doesn't 
help with that.
   2. For `replicated` topics the partition journal lives only in memory. If 
the whole cluster goes down, op numbers start over when it comes back, and a 
floor saved earlier ends up pointing at unrelated ops.
   3. The generation check needs new fields in `RequestPrepares` / 
`RepairRangeReply` (wire format change). It also still leaves (1).
   4. #4092 already persists a purge marker for `persisted` topics, we'd be 
stepping on that.
   
   IMO the checkpoint barrier is the real fix. Right now purge changes 
partition data outside the partition log and nothing makes replicas agree on 
when it happens. Truncation gets away with that because it works off offsets. 
With purge, the timing decides what gets wiped. If purge becomes an op in the 
partition's own consensus group (proposed by the partition primary, kind of 
like auto-commit offsets), every replica applies it at the same op. That op is 
the floor, and repair just replays it like anything else.
   
   Could you write up a short design in this issue before starting on code? 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to