krishvishal opened a new pull request, #3975:
URL: https://github.com/apache/iggy/pull/3975
## The defect
A solo node ACKs sends from its in-memory journal. The client receives a
concrete base offset before the threshold-gated flush writes it to a segment. A
SIGKILL in this window loses the only record that the offset was issued.
```text
LIFE 1 threshold = 4 messages
offsets 0 1 2 3 | 4 5 6 7 8 9 all ACKed
disk: [0..3] | RAM only lost on SIGKILL
BOOT counter := highest segment offset + 1 = 4
LIFE 2 next send ACKed at offset 4 already issued for another
message
```
Two messages now share offset 4. Consumers positioned above the reissued
range also miss the new messages.
`offset_frontier` already existed in the superblock, but stable-view traffic
never persisted it because its write gate only ran on view changes.
## The fix
Add `offset_reserved`, a monotonic ceiling on offsets that may have been
issued. Before an offset can escape, the append fence reserves its block in the
superblock. On boot, minting starts at this ceiling.
```text
send -> mint -> reserve_offsets_through -> journal -+-> commit -> ACK
superblock write +-> poll tier
+-> prepare to peers
```
The reservation covers both primary mints and backup re-stamps. It requires
one write per block, not per batch. With the default 64 Ki lease at 100,000
messages per second, this is about three fsyncs per second. A crash wastes at
most one block of the `u64` offset space.
The lease is configurable through `[partition] offset_reservation_lease`. A
failed reservation rejects the append.
### Why a second field?
```text
segments [0 ..... 3]
journal [4 ..... 9] ACKed, RAM only
^10 ^65537
offset_frontier offset_reserved
offsets known to exist offsets that may
have been issued
```
Combining the fields would make valid transfer offers inside the reserved
block look like rewinds. The rewind guard must instead compare against stored
data: sized segments plus the resident journal. The append counter may be one
lease block ahead after recovery.
### Segment re-anchoring
A gap inside a segment is not recoverable. `recover_segment_bounds` expects
contiguous offsets and truncates everything after a gap, allowing another crash
to reissue confirmed offsets.
```text
BEFORE AFTER BOOT RE-ANCHORING
00000000000000000000.log [0..3] 00000000000000000000.log [0..3] SEALED
append 65537 inside it 00000000000000065537.log [empty]
next boot truncates suffix gap lies on the segment boundary
```
An empty tail that falsely claims a range is removed. A sized tail is
sealed, and a new segment is created at the frontier. Graceful shutdown
collapses the reservation to the frontier, so only crashes spend offsets.
This applies only to solo groups. A backup rejects prepares whose
`base_offset` does not continue its counter.
The superblock grows from 66 to 74 bytes. Older records are rejected, so
data directories from earlier builds are wiped.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]