SEZ9 commented on PR #10808:
URL: https://github.com/apache/seatunnel/pull/10808#issuecomment-5385641573

   Thanks @DanielLeens for the careful follow-up — withdrawing the approval 
after the `Build` check finished with `FAILURE` was the right call, and no 
apology needed.
   
   Your triage is convincing: the timeout happens during test startup while 
waiting for the cluster to report 2 members, before any job is submitted, so 
the checkpoint-barrier `.join()` path from last round cannot be the cause. And 
since the head is still `e92eea319952d44365cec9678e7d1d97eafa96cd`, this 
failure (run `32327424045`, job `96384616750`) is on the exact diff under 
review.
   
   One problem: your comment appears to have been cut off right at the 
root-cause diff. Could you repost the code excerpt and the rest of your 
explanation? I don't want to guess at the fix from a truncated analysis.
   
   Once the fix is pushed, please also confirm the earlier round's points on 
the new head: the CLI fallback coordinator selection possibly diverging from 
the server's actual active coordinator, the incompatible-changes doc coverage 
of the member-list semantics and the standby/failover window, the telemetry 
doc's "bounded timeout" value and configurability, some indication in 
member-list output when the active coordinator cannot be resolved, and the 
`java.util.Collections` import cleanup in the test.
   
   I'll hold off on merging until the test failure is understood and fixed. 
Thanks again for the thorough follow-through.
   
   <!-- streview-comment:488 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to