DanielLeens commented on PR #11727:
URL: https://github.com/apache/seatunnel/pull/11727#issuecomment-5706506926
Thanks for the ping, @abdessalems — good question, and easy to check
precisely since I can pull the actual job logs rather than guess from the red X.
I pulled both `engine-v2-it` job logs from the rerun on `2dc21468105`:
- **JDK 8** (`Run / engine-v2-it (8, ubuntu-latest)`):
`SplitClusterFaultToleranceIT.testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck`
failed with `expected: <CANCELED> but was: <FAILED>` — this is exactly the
same failure I traced on `ce142f830` and reported back then. It's a
pre-existing race on `dev` itself: when a worker is lost mid-cancel, the vertex
sometimes resolves to `FAILED` instead of `CANCELED`. Tracking fix:
apache/seatunnel#12311 ("Resolve a CANCELING vertex to CANCELED, not FAILED,
when its worker is lost") — **still open, not yet merged into `dev`**.
- **JDK 11**: `SplitClusterFaultToleranceIT` actually passed this time (0
failures, 1 skipped). Instead,
`BackpressureSlowSinkIT.testCheckpointsKeepCompletingUnderSustainedBackpressure`
failed: `expected at least 3 additional checkpoints to complete during the 90s
sustained backpressure window, only observed 0`. This is a different, also
pre-existing `dev` flake — a source checkpoint-lock starvation issue where the
reader thread can beat `triggerBarrier` on an unfair monitor. Tracking fixes:
apache/seatunnel#12316 (engine-side fix) and #12313 (test determinism
follow-up) — **both still open, not yet merged into `dev`**.
So to directly answer your question: it's not a new failure and not
something introduced by this PR's diff (which only touches
`BlockingWorker.run()`'s class loader resolution and has no relationship to
checkpoint-barrier injection or cancel/failover state resolution). Both are
known, already-diagnosed `dev`-level flakes with fixes already up for review.
One important nuance though: since #12311/#12316/#12313 haven't landed on
`dev` yet, **syncing this branch with the latest `dev` right now would not make
these two pass** — you'd likely just trade one flaky leg for the other on a
rerun, as you've already been seeing across the last few reruns. I'd treat both
as known/tracked flakes that don't block this PR, and I don't think there's
anything actionable for you to change here. If you want, a plain rerun of just
those two failed legs is reasonable; otherwise this is fine to leave as-is from
a source-review standpoint — my source-level review on this head stands.
The other reds you mentioned (`all-connectors-it-2/6/7`,
`paimon-connector-it`) I'd bucket the same way unless a rerun shows a stack
trace that actually touches something in your diff — happy to take a look at
those logs too if they keep failing.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]