This is an automated email from the ASF dual-hosted git repository.

krishvishal pushed a commit to branch sim-workload-faults
in repository https://gitbox.apache.org/repos/asf/iggy.git

commit 4e9d69cc5994769e63b6f74279fa79f56a5c76b9
Author: Krishna Vishal <[email protected]>
AuthorDate: Sat Aug 15 12:24:46 2026 +0530

    fix(simulator): restore the partition consensus frontier on restart
    
    A restarted replica adopted its retained partition log but rebuilt the
    group's consensus at op 0, so it advertised an empty frontier in its
    `DoViewChange` while holding a log full of ops. A quorum of such replicas
    merged to a log shorter than what peers had already committed -- one run
    merged `op_head=2` against a replica whose `commit_min` was 34 -- and the
    new primary then replayed from the start into a replica far ahead of it,
    tripping the sequential-advance assert in `advance_commit_min`.
    
    The metadata plane already restores this in `restore_metadata_consensus`;
    the partition plane had the log carried across but nothing that read it
    back into consensus. The commit watermark is derived the same way, as the
    highest `commit` any journaled prepare stamped, which is a lower bound
    that re-commits on rejoin.
    
    Under `--crash-primary` this takes the sweep from 17 of 20 to 19 of 20,
    and it also removes the `transmute_header` validity panic entirely, which
    turned out to share the cause: a group whose frontier disagreed with its
    log built a prepare its own peers could not revalidate.
    
    Lives in `init_partition`, which is simulator-gated, so no production path
    changes.
---
 core/shard/src/lib.rs | 30 ++++++++++++++++++++++++++++++
 1 file changed, 30 insertions(+)

diff --git a/core/shard/src/lib.rs b/core/shard/src/lib.rs
index 607713758..2e477816b 100644
--- a/core/shard/src/lib.rs
+++ b/core/shard/src/lib.rs
@@ -3483,6 +3483,36 @@ where
         // recovered offset space) meaningful.
         if let Some((log, durable_offset, write_offset)) = retained {
             partition.adopt_retained_log(log, durable_offset, write_offset);
+            // Restore the consensus frontier from the log we just adopted, the
+            // partition-plane counterpart of what `restore_metadata_consensus`
+            // does for shard 0. Without it a restarted replica rebuilds its
+            // consensus at op 0 while holding a log full of ops, and then
+            // ADVERTISES that empty frontier in its `DoViewChange`. A quorum 
of
+            // such replicas merges to a log shorter than what peers have 
already
+            // committed, and the new primary replays from the start into a 
replica
+            // whose `commit_min` is far ahead, tripping the sequential-advance
+            // assert in `advance_commit_min`.
+            //
+            // The commit watermark is the highest `commit` any journaled 
prepare
+            // stamped: a lower bound, since a prepare records the primary's 
commit
+            // point at send time, so the true point may be one higher and
+            // re-commits on rejoin. Same rule 
`SimJournal::recovery_commit_watermark`
+            // applies on the metadata side.
+            let journal = &partition.log.journal().inner;
+            if let Some(head) = journal.last_op() {
+                let mut watermark = 0;
+                for op in 1..=head {
+                    if let Some(header) = journal.header_by_op(op) {
+                        watermark = watermark.max(header.commit);
+                    }
+                }
+                let consensus = partition.consensus();
+                consensus.sequencer().set_sequence(head);
+                consensus.restore_commit_state(watermark, watermark);
+                if let Some(header) = journal.header_by_op(head) {
+                    consensus.set_last_prepare_checksum(header.checksum);
+                }
+            }
         }
         // The SAME call the boot paths make, not a copy of it: this restore is
         // a max against what the segments already proved, and a harness 
running

Reply via email to