DanielLeens commented on PR #11727:
URL: https://github.com/apache/seatunnel/pull/11727#issuecomment-5706506926

   Thanks for the ping, @abdessalems — good question, and easy to check 
precisely since I can pull the actual job logs rather than guess from the red X.
   
   I pulled both `engine-v2-it` job logs from the rerun on `2dc21468105`:
   
   - **JDK 8** (`Run / engine-v2-it (8, ubuntu-latest)`): 
`SplitClusterFaultToleranceIT.testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck`
 failed with `expected: <CANCELED> but was: <FAILED>` — this is exactly the 
same failure I traced on `ce142f830` and reported back then. It's a 
pre-existing race on `dev` itself: when a worker is lost mid-cancel, the vertex 
sometimes resolves to `FAILED` instead of `CANCELED`. Tracking fix: 
apache/seatunnel#12311 ("Resolve a CANCELING vertex to CANCELED, not FAILED, 
when its worker is lost") — **still open, not yet merged into `dev`**.
   - **JDK 11**: `SplitClusterFaultToleranceIT` actually passed this time (0 
failures, 1 skipped). Instead, 
`BackpressureSlowSinkIT.testCheckpointsKeepCompletingUnderSustainedBackpressure`
 failed: `expected at least 3 additional checkpoints to complete during the 90s 
sustained backpressure window, only observed 0`. This is a different, also 
pre-existing `dev` flake — a source checkpoint-lock starvation issue where the 
reader thread can beat `triggerBarrier` on an unfair monitor. Tracking fixes: 
apache/seatunnel#12316 (engine-side fix) and #12313 (test determinism 
follow-up) — **both still open, not yet merged into `dev`**.
   
   So to directly answer your question: it's not a new failure and not 
something introduced by this PR's diff (which only touches 
`BlockingWorker.run()`'s class loader resolution and has no relationship to 
checkpoint-barrier injection or cancel/failover state resolution). Both are 
known, already-diagnosed `dev`-level flakes with fixes already up for review.
   
   One important nuance though: since #12311/#12316/#12313 haven't landed on 
`dev` yet, **syncing this branch with the latest `dev` right now would not make 
these two pass** — you'd likely just trade one flaky leg for the other on a 
rerun, as you've already been seeing across the last few reruns. I'd treat both 
as known/tracked flakes that don't block this PR, and I don't think there's 
anything actionable for you to change here. If you want, a plain rerun of just 
those two failed legs is reasonable; otherwise this is fine to leave as-is from 
a source-review standpoint — my source-level review on this head stands.
   
   The other reds you mentioned (`all-connectors-it-2/6/7`, 
`paimon-connector-it`) I'd bucket the same way unless a rerun shows a stack 
trace that actually touches something in your diff — happy to take a look at 
those logs too if they keep failing.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to