DanielLeens opened a new issue, #12117:
URL: https://github.com/apache/seatunnel/issues/12117

   ## Description
   
   This is a focused design and implementation task for the Zeta coordinator 
executor. It is not a regression report; no runtime failure was reproduced for 
this issue.
   
   Verified at `dev` commit `97d461bc0773399d632fd078735736ecd44f5f0b` 
(2026-09-05):
   
   - `CoordinatorServiceConfig` exposes `core-thread-num` (default `10`) and 
`max-thread-num` (default `Integer.MAX_VALUE`): 
`ServerConfigOptions.java:456-466` (`470-480` at `04fabece50`).
   - `CoordinatorService.createCoordinatorExecutor()` 
(`CoordinatorService.java:287-297`) builds `new ThreadPoolExecutor(core, max, 
60s, new SynchronousQueue<>(), factory, RejectionCountingHandler)`. The same 
factory is used at construction (`:257`) and when `checkNewActiveMaster()` 
rebuilds the pool after failover. With a `SynchronousQueue` and an unbounded 
maximum, every handoff that finds no idle worker creates a new thread.
   - The executor is not admission-only. `CoordinatorService.java:474` submits 
the job body to it and `jobMaster.run()` (`:491`) runs there; `JobMaster.run()` 
blocks in `jobMasterCompleteFuture.join()` (`JobMaster.java:692`) for the whole 
job lifetime; pipeline-end callbacks are dispatched asynchronously 
(`PhysicalPlan.addPipelineEndCallback`, `PhysicalPlan.java:148-150`) on the 
executor the plan was built with (`JobMaster.java:298`). One live thread per 
running job is therefore parked in this pool.
   
   Why the obvious fix is not safe: bounding `max-thread-num` and adding a work 
queue would let job-lifetime waiters fill the pool while the completion 
callbacks those jobs need queue behind them, which is a progress deadlock. This 
is a source-derived dependency hazard, not a reproduced failure.
   
   ## Expected outcome
   
   1. Separate admission and RPC work from job-lifetime waits: a dedicated 
lifecycle executor for `JobMaster.run()` and pipeline-end callbacks, or removal 
of the blocking wait with equivalent completion ownership.
   2. Only after that separation, choose a bounded default for the admission 
pool using the existing configuration keys, and verify the identical 
configuration is in effect after a `checkNewActiveMaster()` rebuild.
   3. Acceptance: job completion, cancellation, checkpointing, and restore all 
succeed while admission capacity is saturated; the master thread count stays 
bounded under a submission storm and under a failover that restores many jobs 
at once.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to