DanielLeens commented on PR #12134:
URL: https://github.com/apache/seatunnel/pull/12134#issuecomment-6033681815

   Thanks @SEZ9. The reason is in the commit message of `872701ec8`, and I have 
now also added it to the PR description (section "Note on the widened 
assertion"), so it is not only in history.
   
   Short version: `testStreamJobFailsAfterCheckpointTriggerDispatchFailure` is 
a separate copy of the test added in #12288 and flaked the same way here. 
`CheckpointManager#sendOperationToMemberNode` is used by both the barrier 
dispatch in `triggerCheckpoint` and the post-completion notify in 
`notifyCompleted`, so the injected fault can surface through either one 
depending on timing. Both fail the job through `handleCoordinatorError`; only 
the `CheckpointCloseReason` label differs, so asserting one value was flaky on 
timing alone. The deterministic variant from #12288 needs a production-code 
addition, which is out of scope for a PR about the savepoint redrive.
   
   CI on the current head `38c06bbd501`: fork run 37388052981 concluded 
`success`. That is one run for this head, not a flake-rate measurement.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to