DanielLeens opened a new issue, #12117: URL: https://github.com/apache/seatunnel/issues/12117
## Description This is a focused design and implementation task for the Zeta coordinator executor. It is not a regression report; no runtime failure was reproduced for this issue. Verified at `dev` commit `97d461bc0773399d632fd078735736ecd44f5f0b` (2026-09-05): - `CoordinatorServiceConfig` exposes `core-thread-num` (default `10`) and `max-thread-num` (default `Integer.MAX_VALUE`): `ServerConfigOptions.java:456-466` (`470-480` at `04fabece50`). - `CoordinatorService.createCoordinatorExecutor()` (`CoordinatorService.java:287-297`) builds `new ThreadPoolExecutor(core, max, 60s, new SynchronousQueue<>(), factory, RejectionCountingHandler)`. The same factory is used at construction (`:257`) and when `checkNewActiveMaster()` rebuilds the pool after failover. With a `SynchronousQueue` and an unbounded maximum, every handoff that finds no idle worker creates a new thread. - The executor is not admission-only. `CoordinatorService.java:474` submits the job body to it and `jobMaster.run()` (`:491`) runs there; `JobMaster.run()` blocks in `jobMasterCompleteFuture.join()` (`JobMaster.java:692`) for the whole job lifetime; pipeline-end callbacks are dispatched asynchronously (`PhysicalPlan.addPipelineEndCallback`, `PhysicalPlan.java:148-150`) on the executor the plan was built with (`JobMaster.java:298`). One live thread per running job is therefore parked in this pool. Why the obvious fix is not safe: bounding `max-thread-num` and adding a work queue would let job-lifetime waiters fill the pool while the completion callbacks those jobs need queue behind them, which is a progress deadlock. This is a source-derived dependency hazard, not a reproduced failure. ## Expected outcome 1. Separate admission and RPC work from job-lifetime waits: a dedicated lifecycle executor for `JobMaster.run()` and pipeline-end callbacks, or removal of the blocking wait with equivalent completion ownership. 2. Only after that separation, choose a bounded default for the admission pool using the existing configuration keys, and verify the identical configuration is in effect after a `checkNewActiveMaster()` rebuild. 3. Acceptance: job completion, cancellation, checkpointing, and restore all succeed while admission capacity is saturated; the master thread count stays bounded under a submission storm and under a failover that restores many jobs at once. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
