SEZ9 commented on issue #12353: URL: https://github.com/apache/seatunnel/issues/12353#issuecomment-5724550938
Two facts about this test's provenance and its behaviour on `dev` that I don't think have been stated in the thread, and that argue against anyone treating it as flaky. **The test is nine days old, and it was added as coverage for a known defect.** `testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck` came in on 2026-09-09 with `20900262e`, @davidzollo's #12030 — *"[Test][E2E] Add regression coverage for task stuck in CANCELING when worker crashes mid-cancel"*. So it is not a long-standing test that recently became unstable; it is a regression test for a bug that was never fixed, doing exactly what it was written to do. That reading also fits @DanielLeens's analysis above: the `failedTaskNum > 0` priority in `getPipelineEndState()` has presumably always behaved this way, and #12030 simply made it visible. **On `dev` it alternates between the JDK legs, which is what a race looks like from the outside.** Reading the completed `Build` runs on `apache/seatunnel@dev` (most are `cancelled` by the next merge, so only these finished): | `dev` run | Date | `engine-v2-it (8)` | `engine-v2-it (11)` | |---|---|---|---| | [34801745423](https://github.com/apache/seatunnel/actions/runs/34801745423) | 09-14 03:11 | **fail** | pass | | [34821874806](https://github.com/apache/seatunnel/actions/runs/34821874806) | 09-14 08:16 | pass | **fail** | | [34935118103](https://github.com/apache/seatunnel/actions/runs/34935118103) | 09-15 06:01 | **fail** | **fail** | | [34995029902](https://github.com/apache/seatunnel/actions/runs/34995029902) | 09-15 16:26 | **fail** | pass | | [35063682504](https://github.com/apache/seatunnel/actions/runs/35063682504) | 09-16 06:26 | pass | **fail** | | [35219017572](https://github.com/apache/seatunnel/actions/runs/35219017572) | 09-17 12:03 | **fail** | pass | Every completed `dev` run since this test landed has a red `engine-v2-it` leg, and which JDK loses the race varies run to run. One caveat on that table: these are job-level conclusions, and `engine-v2-it` also carries `BackpressureSlowSinkIT` (#12313 is the fix for that one), so not every red cell above is necessarily this test — I confirmed the class-level cause on the 09-17 run and on my own branches, not on all six. Worth noting the 09-07 run's `engine-v2-it (8)` was already red two days *before* this test existed, so there is at least one other pre-existing problem in that job. I did not chase it down. None of this changes the diagnosis, and I have nothing to add to the fix — #12311 looks right to me. I'm posting it because "flaky, will clear on rerun" is the natural assumption for a red e2e job, I made that assumption myself earlier in this thread and was wrong, and the provenance makes it clear the assumption is unavailable here. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
