krishvishal opened a new pull request, #3975:
URL: https://github.com/apache/iggy/pull/3975

   ## The defect
   
   A solo node ACKs sends from its in-memory journal. The client receives a 
concrete base offset before the threshold-gated flush writes it to a segment. A 
SIGKILL in this window loses the only record that the offset was issued.
   
   ```text
   LIFE 1  threshold = 4 messages
           offsets 0 1 2 3 | 4 5 6 7 8 9    all ACKed
           disk:   [0..3]  | RAM only         lost on SIGKILL
   
   BOOT    counter := highest segment offset + 1 = 4
   
   LIFE 2  next send ACKed at offset 4         already issued for another 
message
   ```
   
   Two messages now share offset 4. Consumers positioned above the reissued 
range also miss the new messages.
   
   `offset_frontier` already existed in the superblock, but stable-view traffic 
never persisted it because its write gate only ran on view changes.
   
   ## The fix
   
   Add `offset_reserved`, a monotonic ceiling on offsets that may have been 
issued. Before an offset can escape, the append fence reserves its block in the 
superblock. On boot, minting starts at this ceiling.
   
   ```text
   send -> mint -> reserve_offsets_through -> journal -+-> commit -> ACK
                      superblock write                 +-> poll tier
                                                       +-> prepare to peers
   ```
   
   The reservation covers both primary mints and backup re-stamps. It requires 
one write per block, not per batch. With the default 64 Ki lease at 100,000 
messages per second, this is about three fsyncs per second. A crash wastes at 
most one block of the `u64` offset space.
   
   The lease is configurable through `[partition] offset_reservation_lease`. A 
failed reservation rejects the append.
   
   ### Why a second field?
   
   ```text
   segments  [0 ..... 3]
   journal            [4 ..... 9]   ACKed, RAM only
                                 ^10                              ^65537
                      offset_frontier                     offset_reserved
   
                      offsets known to exist              offsets that may
                                                          have been issued
   ```
   
   Combining the fields would make valid transfer offers inside the reserved 
block look like rewinds. The rewind guard must instead compare against stored 
data: sized segments plus the resident journal. The append counter may be one 
lease block ahead after recovery.
   
   ### Segment re-anchoring
   
   A gap inside a segment is not recoverable. `recover_segment_bounds` expects 
contiguous offsets and truncates everything after a gap, allowing another crash 
to reissue confirmed offsets.
   
   ```text
   BEFORE                              AFTER BOOT RE-ANCHORING
   
   00000000000000000000.log [0..3]     00000000000000000000.log [0..3] SEALED
     append 65537 inside it             00000000000000065537.log [empty]
     next boot truncates suffix         gap lies on the segment boundary
   ```
   
   An empty tail that falsely claims a range is removed. A sized tail is 
sealed, and a new segment is created at the frontier. Graceful shutdown 
collapses the reservation to the frontier, so only crashes spend offsets.
   
   This applies only to solo groups. A backup rejects prepares whose 
`base_offset` does not continue its counter.
   
   The superblock grows from 66 to 74 bytes. Older records are rejected, so 
data directories from earlier builds are wiped.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to