This is an automated email from the ASF dual-hosted git repository. krishvishal pushed a commit to branch sim-workload-faults in repository https://gitbox.apache.org/repos/asf/iggy.git
commit 4e9d69cc5994769e63b6f74279fa79f56a5c76b9 Author: Krishna Vishal <[email protected]> AuthorDate: Sat Aug 15 12:24:46 2026 +0530 fix(simulator): restore the partition consensus frontier on restart A restarted replica adopted its retained partition log but rebuilt the group's consensus at op 0, so it advertised an empty frontier in its `DoViewChange` while holding a log full of ops. A quorum of such replicas merged to a log shorter than what peers had already committed -- one run merged `op_head=2` against a replica whose `commit_min` was 34 -- and the new primary then replayed from the start into a replica far ahead of it, tripping the sequential-advance assert in `advance_commit_min`. The metadata plane already restores this in `restore_metadata_consensus`; the partition plane had the log carried across but nothing that read it back into consensus. The commit watermark is derived the same way, as the highest `commit` any journaled prepare stamped, which is a lower bound that re-commits on rejoin. Under `--crash-primary` this takes the sweep from 17 of 20 to 19 of 20, and it also removes the `transmute_header` validity panic entirely, which turned out to share the cause: a group whose frontier disagreed with its log built a prepare its own peers could not revalidate. Lives in `init_partition`, which is simulator-gated, so no production path changes. --- core/shard/src/lib.rs | 30 ++++++++++++++++++++++++++++++ 1 file changed, 30 insertions(+) diff --git a/core/shard/src/lib.rs b/core/shard/src/lib.rs index 607713758..2e477816b 100644 --- a/core/shard/src/lib.rs +++ b/core/shard/src/lib.rs @@ -3483,6 +3483,36 @@ where // recovered offset space) meaningful. if let Some((log, durable_offset, write_offset)) = retained { partition.adopt_retained_log(log, durable_offset, write_offset); + // Restore the consensus frontier from the log we just adopted, the + // partition-plane counterpart of what `restore_metadata_consensus` + // does for shard 0. Without it a restarted replica rebuilds its + // consensus at op 0 while holding a log full of ops, and then + // ADVERTISES that empty frontier in its `DoViewChange`. A quorum of + // such replicas merges to a log shorter than what peers have already + // committed, and the new primary replays from the start into a replica + // whose `commit_min` is far ahead, tripping the sequential-advance + // assert in `advance_commit_min`. + // + // The commit watermark is the highest `commit` any journaled prepare + // stamped: a lower bound, since a prepare records the primary's commit + // point at send time, so the true point may be one higher and + // re-commits on rejoin. Same rule `SimJournal::recovery_commit_watermark` + // applies on the metadata side. + let journal = &partition.log.journal().inner; + if let Some(head) = journal.last_op() { + let mut watermark = 0; + for op in 1..=head { + if let Some(header) = journal.header_by_op(op) { + watermark = watermark.max(header.commit); + } + } + let consensus = partition.consensus(); + consensus.sequencer().set_sequence(head); + consensus.restore_commit_state(watermark, watermark); + if let Some(header) = journal.header_by_op(head) { + consensus.set_last_prepare_checksum(header.checksum); + } + } } // The SAME call the boot paths make, not a copy of it: this restore is // a max against what the segments already proved, and a harness running
