SEZ9 commented on issue #12353: URL: https://github.com/apache/seatunnel/issues/12353#issuecomment-5723586754
Correction to my earlier comment. The "suite composition" observation I offered there is **wrong**, and I would rather retract it than leave it in the thread as something for someone to chase. I wrote that this test fails 8/8 on #12299 but passes 3/3 on #12298, and that the only correlate I could see was suite size (204 tests vs 203). The 3/3 was an artifact of my own measurement: I only looked at the newest workflow run on that branch, and GitHub carries unchanged jobs forward into each `rerun --failed` attempt with an identical `started_at`, so a single pass appeared as eight. Reading the two earlier runs on the same branch gives the real record for `engine-v2-it (8, ubuntu-latest)` on #12298: | Run | Attempt | `SplitClusterFaultToleranceIT` | Suite | |---|---|---|---| | [34798396872](https://github.com/SEZ9/seatunnel/actions/runs/34798396872) | 1 | **fail** | 204 | | [34926657845](https://github.com/SEZ9/seatunnel/actions/runs/34926657845) | 1 | **fail** | 204 | | 34926657845 | 2 | **fail** | 204 | | 34926657845 | 3 | **fail** | 204 | | 34926657845 | 4 | pass | 204 | | [34993525332](https://github.com/SEZ9/seatunnel/actions/runs/34993525332) | 1–8 (one execution) | pass | 204 | | 34993525332 | 9 | pass | 204 | | 34993525332 | 10 | pass | 204 | So #12298 fails it **4 of 8** real executions, every one of them at 204 tests. The suite-size correlate is directly falsified — the same suite size both passes and fails. Corrected totals for JDK 8 / Linux across everything I have measured: - #12298 — 4 / 8 - #12299 — 8 / 8 - `dev` @ `75b60fa14` + a comment-only change ([run 35166337277](https://github.com/SEZ9/seatunnel/actions/runs/35166337277)) — 1 / 1 **13 failures in 17 real executions (~76%)**, spread over three code bases with no branch correlation. That is a plainer story than the one I told: a base-level race with a high hit rate, not something specific to a branch or a suite. One detail that may be worth something, since @DanielLeens named the mechanism precisely. Attempts 9 and 10 above were single-job reruns I triggered on the newest head; in both, `SplitClusterFaultToleranceIT` passed and `BackpressureSlowSinkIT.testCheckpointsKeepCompletingUnderSustainedBackpressure` failed instead. Two flaky-under-load tests trading places on the same runner fits the described control path: reaching `failedTaskNum > 0` priority in `getPipelineEndState()` requires the crashed worker's `FAILED` callback to land inside the cancel-ack window, so runner speed decides whether the race is won or lost. It is not evidence for a fix, just consistent with the diagnosis rather than with an ordering bug that would fail every time. @tomatotomata — agreed that #12311 looks like the implementation for this, and I have no stake in where the extra cases land. My data above is neutral between the two options you offered. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
