SEZ9 commented on issue #12353:
URL: https://github.com/apache/seatunnel/issues/12353#issuecomment-5723586754

   Correction to my earlier comment. The "suite composition" observation I 
offered there is **wrong**, and I would rather retract it than leave it in the 
thread as something for someone to chase.
   
   I wrote that this test fails 8/8 on #12299 but passes 3/3 on #12298, and 
that the only correlate I could see was suite size (204 tests vs 203). The 3/3 
was an artifact of my own measurement: I only looked at the newest workflow run 
on that branch, and GitHub carries unchanged jobs forward into each `rerun 
--failed` attempt with an identical `started_at`, so a single pass appeared as 
eight. Reading the two earlier runs on the same branch gives the real record 
for `engine-v2-it (8, ubuntu-latest)` on #12298:
   
   | Run | Attempt | `SplitClusterFaultToleranceIT` | Suite |
   |---|---|---|---|
   | [34798396872](https://github.com/SEZ9/seatunnel/actions/runs/34798396872) 
| 1 | **fail** | 204 |
   | [34926657845](https://github.com/SEZ9/seatunnel/actions/runs/34926657845) 
| 1 | **fail** | 204 |
   | 34926657845 | 2 | **fail** | 204 |
   | 34926657845 | 3 | **fail** | 204 |
   | 34926657845 | 4 | pass | 204 |
   | [34993525332](https://github.com/SEZ9/seatunnel/actions/runs/34993525332) 
| 1–8 (one execution) | pass | 204 |
   | 34993525332 | 9 | pass | 204 |
   | 34993525332 | 10 | pass | 204 |
   
   So #12298 fails it **4 of 8** real executions, every one of them at 204 
tests. The suite-size correlate is directly falsified — the same suite size 
both passes and fails. Corrected totals for JDK 8 / Linux across everything I 
have measured:
   
   - #12298 — 4 / 8
   - #12299 — 8 / 8
   - `dev` @ `75b60fa14` + a comment-only change ([run 
35166337277](https://github.com/SEZ9/seatunnel/actions/runs/35166337277)) — 1 / 
1
   
   **13 failures in 17 real executions (~76%)**, spread over three code bases 
with no branch correlation. That is a plainer story than the one I told: a 
base-level race with a high hit rate, not something specific to a branch or a 
suite.
   
   One detail that may be worth something, since @DanielLeens named the 
mechanism precisely. Attempts 9 and 10 above were single-job reruns I triggered 
on the newest head; in both, `SplitClusterFaultToleranceIT` passed and 
`BackpressureSlowSinkIT.testCheckpointsKeepCompletingUnderSustainedBackpressure`
 failed instead. Two flaky-under-load tests trading places on the same runner 
fits the described control path: reaching `failedTaskNum > 0` priority in 
`getPipelineEndState()` requires the crashed worker's `FAILED` callback to land 
inside the cancel-ack window, so runner speed decides whether the race is won 
or lost. It is not evidence for a fix, just consistent with the diagnosis 
rather than with an ordering bug that would fail every time.
   
   @tomatotomata — agreed that #12311 looks like the implementation for this, 
and I have no stake in where the extra cases land. My data above is neutral 
between the two options you offered.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to