yrenat opened a new issue, #8618:
URL: https://github.com/apache/texera/issues/8618
### Feature Summary
Follow-up to #6046.
### Problem
The idle Kubernetes computing unit sweep treats a computing unit as busy
whenever any of its
`workflow_executions` rows carries a non-terminal status code (`status NOT
IN (3, 4, 5)`). That
test has no time bound, so an execution row left stuck in a non-terminal
state keeps its computing unit off the sweep indefinitely, even though the unit
is doing no work.
But the problem is, nothing else reclaims it either:
`ComputingUnitHelpers.reconcileVanishedKubernetesUnits` only
runs when someone calls a listing endpoint, and it only checks whether the
pod is already gone,
not whether the execution row is stuck somewhere. A live pod with a stuck
row is missed on both paths.
### Proposed Solution or Design
### Possible directions
- Ignore a non-terminal execution row whose `last_update_time` is older than
its own timeout.
- Check execution status codes against actual pod state on a schedule,
rather than only on a
listing request.
Either is a larger change than #6046 should carry, hence this follow-up.
### Affected Area
_No response_
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]