SEZ9 commented on issue #12353:
URL: https://github.com/apache/seatunnel/issues/12353#issuecomment-5707445866

   Corroborating data point on a different platform and JDK, in case it helps 
narrow this down.
   
   I hit the same failure on **JDK 8 / Linux** (`ubuntu-latest`), where you saw 
it on JDK 17 / Windows, so it is not JDK- or OS-specific:
   
   ```
   [ERROR] 
SplitClusterFaultToleranceIT.testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck:448
           ->assertEventuallyCanceled:556 ยป ConditionTimeout
   org.awaitility.core.ConditionTimeoutException: ... expected: <CANCELED> but 
was: <FAILED> within 1 minutes.
   Caused by: org.opentest4j.AssertionFailedError: expected: <CANCELED> but 
was: <FAILED>
   [ERROR] Tests run: 203, Failures: 0, Errors: 1, Skipped: 6
   ```
   
   To be sure it was not something in my own branches, I ran it on a branch 
that is `dev` @ `75b60fa14` plus a three-line comment in 
`seatunnel-engine/README.md` and nothing else โ€” the comment exists only so 
change detection schedules `engine-v2-it`. It fails there: [run 
35166337277](https://github.com/SEZ9/seatunnel/actions/runs/35166337277).
   
   Two further observations from repeated runs that might be useful:
   
   - **The JDK 11 leg of that same control run passed.** So on Linux it is 
JDK-8-reproducible but not JDK-11-reproducible in a single sample โ€” consistent 
with a race rather than a deterministic ordering bug, even though it looks 
deterministic on any one leg.
   - **It failed 8 out of 8 real executions** of `engine-v2-it (8)` on one of 
my pull requests (#12299, which does not touch the engine's cancel path), while 
**passing 3 out of 3** on another (#12298, same base). The only difference I 
can point at is that #12298 adds one test to that e2e module, so it runs 204 
tests where the base and #12299 run 203. Given the test deliberately races a 
cancel request against a worker shutdown, suite composition shifting the timing 
is a plausible mechanism โ€” but I have not tested that and offer it only as 
something to rule out, not as a finding.
   
   Your pointer at `SubPlan.getPipelineEndState()` / 
`addPhysicalVertexCallBack()` matches what the assertion implies: the job 
reaches a terminal state, just `FAILED` instead of `CANCELED`, so the 
terminal-state decision is being made from the crashed worker's vertex outcome 
rather than from the already-recorded cancellation intent. I have not read that 
path closely enough to say more.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to