jphjsoares commented on issue #3793: URL: https://github.com/apache/iggy/issues/3793#issuecomment-5652850384
My design proposal is to represent the purge as **a partition-plane consensus operation**, tentatively named as, `PurgePartition`. `G` = metadata purge generation `P` = the consensus operation number of `PurgePartition` - **Metadata and proposal:** Metadata continues to record purge intent, and the public purge request remains acknowledged at the metadata level while partition cleanup converges asynchronously. The reconciler should stop invoking local partition cleanup directly. The owning primary proposes `P`; followers wait for the replicated operation. Dropped requests and primary changes must be retryable, with at most one purge proposal outstanding per partition. - **Staging:** For the initial implementation, the primary can stop admitting new sends and consumer-offset writes while `P` is outstanding. It must not reserve or assign speculative post-purge offsets. Any already-replicated post-`P` operations must remain staged and must not be materialized or flushed past `P` until the purge completes. The shard pump must continue processing consensus traffic while requests are held. - **Commit:** `P` is journaled during prepare ingestion and applied in partition-log order. The commit path should flush only operations before `P`, execute the existing cleanup using `P` as the agreed boundary, then process operations after `P`. A cleanup failure must prevent the local applied/commit frontier from advancing past `P`, with retry or fencing as appropriate. A duplicate `(incarnation, G)` proposal may advance the log but must be a no-op: it must not wipe post-purge data or move the established boundary. - **Repair:** `P` must travel through repair as a partition operation, not as a metadata operation. Repaired operations must be committed in log order. A completed durable marker for `(G, P, incarnation)` may establish the purge boundary, but repaired operations at or before `P` must not be materialized again. If no unambiguous matching boundary exists in the repair history or durable checkpoint state, the replica must use state transfer before skipping or materializing the missing prefix. `P` identifies the purge boundary, but it is not a snapshot of the state that survives the purge, especially the table of already-committed client requests. Most likely, recovery must preserve or transfer that state when the `pre-P` log is unavailable. - **Persistence:** Use a crash-safe marker containing `(G, P, incarnation, state)`. Record the purge intent before destructive cleanup, and mark it complete only after the cleanup and required durability steps finish. On restart, an incomplete marker resumes the purge; a complete marker makes replay of the same `P` idempotent and must not delete post-`P` data. Legacy generation-only or locally chosen-floor records cannot establish a consensus boundary and must be handled conservatively. Let me know your opinions about this. Please point out any incorrect assumptions, missing constraints, or alternative approaches. Once the behavior and compatibility decisions are agreed on, I can look into the implementation. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
