goutamadwant commented on PR #12311: URL: https://github.com/apache/seatunnel/pull/12311#issuecomment-5841261302
I ran `SplitClusterFaultToleranceIT#testStreamJobCancelResolvesWhenWorkerCrashesBeforeCancelAck` in a loop locally to see whether this fixes the `engine-v2-it` flake. Setup: macOS, 10 cores, with other builds running on the same machine. Each run was a fresh JVM. Base was `dev` at `deb16a3c3`. Head was this PR's diff applied to the same commit. It applies cleanly, and none of the files it touches have changed on `dev` since this branch's base. | | base | this PR | |---|---|---| | JDK 8 (1.8.0_172), host load 26-63 | 12 / 20 failed | 0 / 20 failed | | JDK 11, `-XX:ActiveProcessorCount=2`, quiet host | 4 / 10 failed | 0 / 10 failed | Every base failure was the one CI shows: `assertEventuallyCanceled` fails with `expected: <CANCELED> but was: <FAILED>`. `CoordinatorServiceLostMemberResolutionTest` also passes. An earlier JDK 11 batch also hit `Node failed to start!` on both base and head. I was running other in-process cluster tests at the same time, so this was port and join interference from my setup. The quiet-host rerun above had none of it. On the change itself: only a vertex that is already CANCELING moves to CANCELED when its worker is lost. DEPLOYING and RUNNING still become FAILED. I also checked the paths where CANCELING does not come from a user cancel: - A pipeline that is cancelling because a sibling task failed still ends FAILED, through `failedTaskNum` and the FAILING check in `SubPlan#getPipelineEndState`. - A pipeline cancelled by `handleCheckpointError` still ends FAILED through the checkpoint coordinator status check in the same method. So I don't see restore or failure reporting being masked. LGTM from the testing side. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
