davidzollo commented on PR #12031: URL: https://github.com/apache/seatunnel/pull/12031#issuecomment-5549650459
### CI update for head `08b795b` — this PR's own test now passes; the one remaining failure is unrelated The redesign is in (`f34f4f1`) and it worked. On run [33942789441](https://github.com/davidzollo/seatunnel/actions/runs/33942789441): ``` Tests run: 3, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 82.729 s - in org.apache.seatunnel.engine.e2e.CheckpointCoordinatorFailoverIT ``` green on **both** `engine-v2-it (8)` and `engine-v2-it (11)`, so `testBatchJobCompletesAfterMasterFailoverDuringCloseHandshake` is fixed on both JDKs. #### What the fix actually changed Per my previous analysis, the assertion was written against a total that cannot occur. Correcting that, rather than loosening the bound: - `CLOSE_HANDSHAKE_STARTING_SUBTASKS` `4` -> `2`. `env.parallelism = 2` scales each source's *reader* tasks; `PhysicalPlanGenerator#getEnumeratorTask` allocates exactly one split-enumerator coordinator subtask per source action regardless of reader parallelism, and the enumerator coordinator is what the ready-to-close handshake tracks. So the real total is one per pipeline, two overall, and the partial state the test waits for is `1` — reachable, where the old `(0, 4)` window was not. - Poll interval `20ms` -> `1ms`. The window is genuinely transient (once `table_fast`'s lone starting subtask reports ready, its `COMPLETED_POINT_TYPE` checkpoint fires immediately and `CheckpointCoordinator#shutdown` removes the `readyToCloseImapKey` entry), so this buys more chances to land inside it. The 30s `atMost` bound and the strict `0 < n < 2` assertion are unchanged — nothing was weakened. #### The remaining red job is not this PR `engine-v2-it (8)` failed in a test this PR does not touch: ``` ClusterFaultToleranceTwoPipelineIT.testTwoPipelineStreamJobRestoreIn2NodeMasterDown <<< ERROR! java.lang.RuntimeException: java.util.concurrent.ExecutionException: SeaTunnelEngineRetryableException: Can not get coordinator service from an active master node. at ClusterFaultToleranceTwoPipelineIT.java:755 ``` This PR changes exactly two files — `CheckpointCoordinatorFailoverIT.java` and `batch_fake_to_localfile_close_handshake_failover_template.conf` — and neither is on that test's path. The failure is a master-election timing race inside that test's own wait (it gave up after 17s), and the same test passed on `engine-v2-it (11)` in this run. I have re-run **only** that one job (attempt 2); the passing JDK 11 result is preserved. Nothing further is pending from my side on this PR. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
