SEZ9 commented on issue #12353:
URL: https://github.com/apache/seatunnel/issues/12353#issuecomment-5724550938

   Two facts about this test's provenance and its behaviour on `dev` that I 
don't think have been stated in the thread, and that argue against anyone 
treating it as flaky.
   
   **The test is nine days old, and it was added as coverage for a known 
defect.** `testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck` came in 
on 2026-09-09 with `20900262e`, @davidzollo's #12030 — *"[Test][E2E] Add 
regression coverage for task stuck in CANCELING when worker crashes 
mid-cancel"*. So it is not a long-standing test that recently became unstable; 
it is a regression test for a bug that was never fixed, doing exactly what it 
was written to do. That reading also fits @DanielLeens's analysis above: the 
`failedTaskNum > 0` priority in `getPipelineEndState()` has presumably always 
behaved this way, and #12030 simply made it visible.
   
   **On `dev` it alternates between the JDK legs, which is what a race looks 
like from the outside.** Reading the completed `Build` runs on 
`apache/seatunnel@dev` (most are `cancelled` by the next merge, so only these 
finished):
   
   | `dev` run | Date | `engine-v2-it (8)` | `engine-v2-it (11)` |
   |---|---|---|---|
   | 
[34801745423](https://github.com/apache/seatunnel/actions/runs/34801745423) | 
09-14 03:11 | **fail** | pass |
   | 
[34821874806](https://github.com/apache/seatunnel/actions/runs/34821874806) | 
09-14 08:16 | pass | **fail** |
   | 
[34935118103](https://github.com/apache/seatunnel/actions/runs/34935118103) | 
09-15 06:01 | **fail** | **fail** |
   | 
[34995029902](https://github.com/apache/seatunnel/actions/runs/34995029902) | 
09-15 16:26 | **fail** | pass |
   | 
[35063682504](https://github.com/apache/seatunnel/actions/runs/35063682504) | 
09-16 06:26 | pass | **fail** |
   | 
[35219017572](https://github.com/apache/seatunnel/actions/runs/35219017572) | 
09-17 12:03 | **fail** | pass |
   
   Every completed `dev` run since this test landed has a red `engine-v2-it` 
leg, and which JDK loses the race varies run to run. One caveat on that table: 
these are job-level conclusions, and `engine-v2-it` also carries 
`BackpressureSlowSinkIT` (#12313 is the fix for that one), so not every red 
cell above is necessarily this test — I confirmed the class-level cause on the 
09-17 run and on my own branches, not on all six.
   
   Worth noting the 09-07 run's `engine-v2-it (8)` was already red two days 
*before* this test existed, so there is at least one other pre-existing problem 
in that job. I did not chase it down.
   
   None of this changes the diagnosis, and I have nothing to add to the fix — 
#12311 looks right to me. I'm posting it because "flaky, will clear on rerun" 
is the natural assumption for a red e2e job, I made that assumption myself 
earlier in this thread and was wrong, and the provenance makes it clear the 
assumption is unavailable here.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to