jphjsoares commented on issue #3793:
URL: https://github.com/apache/iggy/issues/3793#issuecomment-5652850384

   My design proposal is to represent the purge as **a partition-plane 
consensus operation**, tentatively named as, `PurgePartition`.
   
   `G` = metadata purge generation  
   `P` = the consensus operation number of `PurgePartition`
   
   - **Metadata and proposal:** Metadata continues to record purge intent, and 
the public purge request remains acknowledged at the metadata level while 
partition cleanup converges asynchronously. The reconciler should stop invoking 
local partition cleanup directly. The owning primary proposes `P`; followers 
wait for the replicated operation. Dropped requests and primary changes must be 
retryable, with at most one purge proposal outstanding per partition.
   - **Staging:** For the initial implementation, the primary can stop 
admitting new sends and consumer-offset writes while `P` is outstanding. It 
must not reserve or assign speculative post-purge offsets. Any 
already-replicated post-`P` operations must remain staged and must not be 
materialized or flushed past `P` until the purge completes. The shard pump must 
continue processing consensus traffic while requests are held.
   - **Commit:** `P` is journaled during prepare ingestion and applied in 
partition-log order. The commit path should flush only operations before `P`, 
execute the existing cleanup using `P` as the agreed boundary, then process 
operations after `P`. A cleanup failure must prevent the local applied/commit 
frontier from advancing past `P`, with retry or fencing as appropriate. A 
duplicate `(incarnation, G)` proposal may advance the log but must be a no-op: 
it must not wipe post-purge data or move the established boundary.
   - **Repair:** `P` must travel through repair as a partition operation, not 
as a metadata operation. Repaired operations must be committed in log order. A 
completed durable marker for `(G, P, incarnation)` may establish the purge 
boundary, but repaired operations at or before `P` must not be materialized 
again. If no unambiguous matching boundary exists in the repair history or 
durable checkpoint state, the replica must use state transfer before skipping 
or materializing the missing prefix. `P` identifies the purge boundary, but it 
is not a snapshot of the state that survives the purge, especially the table of 
already-committed client requests. Most likely, recovery must preserve or 
transfer that state when the `pre-P` log is unavailable.
   - **Persistence:** Use a crash-safe marker containing `(G, P, incarnation, 
state)`. Record the purge intent before destructive cleanup, and mark it 
complete only after the cleanup and required durability steps finish. On 
restart, an incomplete marker resumes the purge; a complete marker makes replay 
of the same `P` idempotent and must not delete post-`P` data. Legacy 
generation-only or locally chosen-floor records cannot establish a consensus 
boundary and must be handled conservatively.
   
   Let me know your opinions about this. Please point out any incorrect 
assumptions, missing constraints, or alternative approaches. Once the behavior 
and compatibility decisions are agreed on, I can look into the implementation.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to