CryoThrust commented on issue #12139: URL: https://github.com/apache/seatunnel/issues/12139#issuecomment-5556476068
I traced the reported call chain in dev: NotifyTaskRestoreOperation catches the restoreState failure and sends CheckpointErrorReportOperation; its runInternal calls CheckpointManager.reportCheckpointErrorFromTask synchronously on the Hazelcast operation thread. CheckpointCoordinator.handleCoordinatorError then calls JobMaster.handleCheckpointError, which synchronously transitions SubPlan to CANCELING and can enter the restore/state-process path. The important invariant seems to be that a remote operation must report the error and return promptly; pipeline cancellation, resource release, and restore/retry should be serialized on the JobMaster or checkpoint coordinator executor. A likely shape is to keep the existing single-terminal guard in CheckpointCoordinator, then enqueue the state transition exactly once on its executor (or a JobMaster-owned executor), rather than invoking SubPlan state changes inline from CheckpointErrorReportOperation. Simply making the operation fire-and-forget without preserving that guard could race duplicate error reports. A focused regression should make a source enumerator throw during restoreState and assert both: (1) the Hazelcast operation completes without waiting for restore/retry work, and (2) after the configured retry budget the pipeline reaches FAILED and the job future completes with the original privilege error. This would distinguish the operation-thread deadlock from the connector exception itself. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
