CryoThrust commented on issue #12121:
URL: https://github.com/apache/seatunnel/issues/12121#issuecomment-5595002395

   I reviewed the current `TaskExecutionService` path and did not find an 
active implementation PR for this issue. A useful first contract may separate 
three budgets rather than applying one global thread cap:
   
   - a global ceiling for promoted cooperative workers;
   - a per-job (or per task-group) ceiling so one noisy job cannot consume the 
pool;
   - reserved capacity for readiness/coordination work, so an exhausted 
promotion budget cannot prevent source/sink startup or coordinator callbacks.
   
   The admission decision should be explicit: when a timeout would promote work 
but the budget is exhausted, keep the tracker queued with a bounded 
retry/backoff and expose a reason (`BUDGET_EXHAUSTED`) rather than silently 
dropping it. The budget should be released when the promoted worker returns to 
the queue or the task group terminates, not merely when a timer fires.
   
   For deterministic tests, I would inject the clock and promotion policy, then 
cover (1) one job exhausting its own budget while another job still progresses, 
(2) global budget exhaustion with reserved readiness capacity, (3) 
release/reuse after a worker yields, (4) cancellation while a promotion is 
pending, and (5) shutdown without stranded queue entries. Metrics can be added 
after the policy contract is accepted.
   
   Would maintainers prefer the first PR to contain only a policy/budget object 
and unit tests, or should it also wire the decision into `TaskCallTimer` and 
`RunBusWorkSupplier`? I can prepare the smallest slice that matches the 
expected review boundary.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to