diegomrsantos opened a new issue, #4239:
URL: https://github.com/apache/iggy/issues/4239

   Following [hubcio's suggestion in 
#4128](https://github.com/apache/iggy/issues/4128#issuecomment-5632296400), 
make production purge cleanup, completion and consumer offset recovery usable 
with `SimStorage` through `DurableStorage`. This would let us test what 
survives power loss before relying on a fix for #4128.
   
   `SimStorage` already models directory sync failures and `Crash::PowerLoss`, 
including restoring file entries whose deletion was not synced. The missing 
part is connecting production code to that storage. The [existing storage purge 
test](https://github.com/apache/iggy/blob/ad8951f57/core/simulator/src/storage/tests.rs#L2484)
 implements its own cleanup sequence and propagates sync errors, while 
production logs those errors and continues. The [cluster restart 
path](https://github.com/apache/iggy/blob/ad8951f57/core/simulator/src/lib.rs#L1364)
 retains consumer offsets in memory. Neither path currently exercises the 
recovery sequence needed for this bug.
   
   The implementation should
   
   - Share the production cleanup and completion control flow with a focused 
storage harness, including the decisions made after errors. Sharing only the 
filesystem calls would still let the test behave differently from production.
   - Route the relevant directory scans, offset file deletion, directory syncs 
and persistence of `purge.gen` through `DurableStorage`.
   - Preserve separate directories for individual consumer offsets, consumer 
group offsets and the generation marker. Syncing the marker's directory must 
not make deletions in the offset directories durable.
   - Reload `purge.gen` and both kinds of consumer offsets from the simulated 
filesystem through shared production recovery code. Checks of `Next` must use 
that recovered state.
   - Add passing controls for successful purge and recovery under both 
`replicated` and `persisted` consumer offset durability. Check individual 
consumers and consumer groups, and verify that `Next` returns the complete 
fresh history. Make that history durable before any simulated power loss so 
these checks isolate offset recovery from message durability.
   
   Keep this as a preparatory refactor that preserves production behavior. A 
focused harness using the existing storage simulator is enough to start. 
Converting the whole cluster simulator is not required.
   
   The regression for #4128 and #4130 can then fail each offset directory sync 
separately after successful deletion, allow subsequent operations and simulate 
power loss. That regression should check that the completion marker cannot 
suppress required cleanup, stale offsets disappear after cleanup, and retries 
preserve acknowledged fresh writes. #4130 already reproduces the production 
sync failure but does not cover that full recovery sequence.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to