diegomrsantos opened a new issue, #4239: URL: https://github.com/apache/iggy/issues/4239
Following [hubcio's suggestion in #4128](https://github.com/apache/iggy/issues/4128#issuecomment-5632296400), make production purge cleanup, completion and consumer offset recovery usable with `SimStorage` through `DurableStorage`. This would let us test what survives power loss before relying on a fix for #4128. `SimStorage` already models directory sync failures and `Crash::PowerLoss`, including restoring file entries whose deletion was not synced. The missing part is connecting production code to that storage. The [existing storage purge test](https://github.com/apache/iggy/blob/ad8951f57/core/simulator/src/storage/tests.rs#L2484) implements its own cleanup sequence and propagates sync errors, while production logs those errors and continues. The [cluster restart path](https://github.com/apache/iggy/blob/ad8951f57/core/simulator/src/lib.rs#L1364) retains consumer offsets in memory. Neither path currently exercises the recovery sequence needed for this bug. The implementation should - Share the production cleanup and completion control flow with a focused storage harness, including the decisions made after errors. Sharing only the filesystem calls would still let the test behave differently from production. - Route the relevant directory scans, offset file deletion, directory syncs and persistence of `purge.gen` through `DurableStorage`. - Preserve separate directories for individual consumer offsets, consumer group offsets and the generation marker. Syncing the marker's directory must not make deletions in the offset directories durable. - Reload `purge.gen` and both kinds of consumer offsets from the simulated filesystem through shared production recovery code. Checks of `Next` must use that recovered state. - Add passing controls for successful purge and recovery under both `replicated` and `persisted` consumer offset durability. Check individual consumers and consumer groups, and verify that `Next` returns the complete fresh history. Make that history durable before any simulated power loss so these checks isolate offset recovery from message durability. Keep this as a preparatory refactor that preserves production behavior. A focused harness using the existing storage simulator is enough to start. Converting the whole cluster simulator is not required. The regression for #4128 and #4130 can then fail each offset directory sync separately after successful deletion, allow subsequent operations and simulate power loss. That regression should check that the completion marker cannot suppress required cleanup, stale offsets disappear after cleanup, and retries preserve acknowledged fresh writes. #4130 already reproduces the production sync failure but does not cover that full recovery sequence. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
