krishvishal commented on issue #3793: URL: https://github.com/apache/iggy/issues/3793#issuecomment-5630051088
Thanks for reproducing it, and for the `--features vsr` / `cluster_nodes = 1` notes. I'll fix the repro in the issue. Yeah, that's pretty much what's going on. One thing to add: the [serve-side check](https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/shard/src/lib.rs#L4884) only kicks in when the serving peer's metadata already has the purge committed. In your run it didn't have it yet. That's why the peer had floor 0 and served from op 1. The [receive-side check](https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/shard/src/lib.rs#L5413) on the restarted node just compares its own committed vs applied generation, which were equal, and its floor was back to 0 after the restart. Neither node has any idea what the other one applied. I don't think persisting the floor + a generation check is the way to go: 1. The floor is whatever [the local op counter](https://github.com/apache/iggy/blob/d693eb93d878a7759221fba05dcea20e52dd1249/core/partitions/src/iggy_partition.rs#L6636) happens to be when that node's reconciler runs. Two replicas on the same generation can have different floors, restart or not. Saving it to disk doesn't help with that. 2. For `replicated` topics the partition journal lives only in memory. If the whole cluster goes down, op numbers start over when it comes back, and a floor saved earlier ends up pointing at unrelated ops. 3. The generation check needs new fields in `RequestPrepares` / `RepairRangeReply` (wire format change). It also still leaves (1). 4. #4092 already persists a purge marker for `persisted` topics, we'd be stepping on that. IMO the checkpoint barrier is the real fix. Right now purge changes partition data outside the partition log and nothing makes replicas agree on when it happens. Truncation gets away with that because it works off offsets. With purge, the timing decides what gets wiped. If purge becomes an op in the partition's own consensus group (proposed by the partition primary, kind of like auto-commit offsets), every replica applies it at the same op. That op is the floor, and repair just replays it like anything else. Could you write up a short design in this issue before starting on code? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
