DanielLeens opened a new pull request, #12096:
URL: https://github.com/apache/seatunnel/pull/12096

   ### Purpose of this pull request
   
   Fixes #12094.
   
   `CheckpointCoordinator` previously rebuilt `completedCheckpointIds` as an 
empty queue after an active-master failover. Checkpoints written by earlier 
coordinator instances were therefore invisible to retention cleanup and could 
accumulate beyond `checkpoint.storage.max-retained`.
   
   This patch:
   
   - rebuilds the durable checkpoint retention queue for the current job and 
pipeline when a running coordinator is recreated;
   - sorts storage enumeration results by checkpoint ID before selecting the 
oldest files;
   - enforces the configured maximum after recovery and after every subsequent 
durable checkpoint;
   - keeps explicit checkpoint restores to a different destination job isolated 
from the source job's retention state;
   - makes LocalFile and HDFS batch deletion failures observable while keeping 
retention cleanup best-effort at the coordinator layer, so cleanup failures do 
not fail checkpoint completion or active-master recovery and remain eligible 
for retry.
   
   ### Does this PR introduce _any_ user-facing change?
   
   Yes.
   
   Previously, a running Zeta job could retain more checkpoint files than 
`checkpoint.storage.max-retained` after one or more active-master failovers. 
After this change, the coordinator restores the existing checkpoint IDs and 
removes the oldest files until the configured bound is satisfied. A transient 
cleanup failure is logged and retried after the next completed checkpoint 
without interrupting the job.
   
   ### How was this patch tested?
   
   Added regression tests covering:
   
   - unordered checkpoint enumeration and exact cleanup during active-master 
coordinator recreation;
   - exact retention on the next completed checkpoint;
   - isolation when restoring from a different source job;
   - cleanup failure not interrupting coordinator recreation or checkpoint 
completion, with retry on the next checkpoint;
   - LocalFile delete exceptions and HDFS `FileSystem.delete` returning `false`.
   
   Formatting verification completed successfully:
   
   ```shell
   ./mvnw -nsu -pl 
seatunnel-engine/seatunnel-engine-server,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-local-file,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-hdfs
 -am spotless:apply
   ./mvnw -nsu -pl 
seatunnel-engine/seatunnel-engine-server,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-local-file,seatunnel-engine/seatunnel-engine-storage/checkpoint-storage-plugins/checkpoint-storage-hdfs
 -am spotless:check
   ```
   
   Functional tests are delegated to GitHub CI and were not executed locally.
   
   ### Check list
   
   * [x] No new Jar binary package is added.
   * [x] Documentation changes are not required because no configuration or 
public API changes.
   * [x] `incompatible-changes.md` changes are not required because the patch 
restores the documented retention behavior.
   * [x] This PR does not change connector code.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to