SEZ9 commented on issue #12353:
URL: https://github.com/apache/seatunnel/issues/12353#issuecomment-5707445866
Corroborating data point on a different platform and JDK, in case it helps
narrow this down.
I hit the same failure on **JDK 8 / Linux** (`ubuntu-latest`), where you saw
it on JDK 17 / Windows, so it is not JDK- or OS-specific:
```
[ERROR]
SplitClusterFaultToleranceIT.testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck:448
->assertEventuallyCanceled:556 ยป ConditionTimeout
org.awaitility.core.ConditionTimeoutException: ... expected: <CANCELED> but
was: <FAILED> within 1 minutes.
Caused by: org.opentest4j.AssertionFailedError: expected: <CANCELED> but
was: <FAILED>
[ERROR] Tests run: 203, Failures: 0, Errors: 1, Skipped: 6
```
To be sure it was not something in my own branches, I ran it on a branch
that is `dev` @ `75b60fa14` plus a three-line comment in
`seatunnel-engine/README.md` and nothing else โ the comment exists only so
change detection schedules `engine-v2-it`. It fails there: [run
35166337277](https://github.com/SEZ9/seatunnel/actions/runs/35166337277).
Two further observations from repeated runs that might be useful:
- **The JDK 11 leg of that same control run passed.** So on Linux it is
JDK-8-reproducible but not JDK-11-reproducible in a single sample โ consistent
with a race rather than a deterministic ordering bug, even though it looks
deterministic on any one leg.
- **It failed 8 out of 8 real executions** of `engine-v2-it (8)` on one of
my pull requests (#12299, which does not touch the engine's cancel path), while
**passing 3 out of 3** on another (#12298, same base). The only difference I
can point at is that #12298 adds one test to that e2e module, so it runs 204
tests where the base and #12299 run 203. Given the test deliberately races a
cancel request against a worker shutdown, suite composition shifting the timing
is a plausible mechanism โ but I have not tested that and offer it only as
something to rule out, not as a finding.
Your pointer at `SubPlan.getPipelineEndState()` /
`addPhysicalVertexCallBack()` matches what the assertion implies: the job
reaches a terminal state, just `FAILED` instead of `CANCELED`, so the
terminal-state decision is being made from the crashed worker's vertex outcome
rather than from the already-recorded cancellation intent. I have not read that
path closely enough to say more.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]