yrenat opened a new issue, #8618:
URL: https://github.com/apache/texera/issues/8618

   ### Feature Summary
   
   Follow-up to #6046.
   
   ### Problem
   
   The idle Kubernetes computing unit sweep treats a computing unit as busy 
whenever any of its
   `workflow_executions` rows carries a non-terminal status code (`status NOT 
IN (3, 4, 5)`). That
   test has no time bound, so an execution row left stuck in a non-terminal 
state keeps its computing unit off the sweep indefinitely, even though the unit 
is doing no work.
   
   But the problem is, nothing else reclaims it either: 
`ComputingUnitHelpers.reconcileVanishedKubernetesUnits` only
   runs when someone calls a listing endpoint, and it only checks whether the 
pod is already gone,
   not whether the execution row is stuck somewhere. A live pod with a stuck 
row is missed on both paths.
   
   ### Proposed Solution or Design
   
   ### Possible directions
   
   - Ignore a non-terminal execution row whose `last_update_time` is older than 
its own timeout.
   - Check execution status codes against actual pod state on a schedule, 
rather than only on a
     listing request.
   
   Either is a larger change than #6046 should carry, hence this follow-up.
   
   ### Affected Area
   
   _No response_


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to