Dale Richardson created YUNIKORN-3375:
-----------------------------------------
Summary: Node delete/re-add silently corrupts queue accounting:
adopted pods are never re-reported to the core
Key: YUNIKORN-3375
URL: https://issues.apache.org/jira/browse/YUNIKORN-3375
Project: Apache YuniKorn
Issue Type: Bug
Components: shim - kubernetes
Reporter: Dale Richardson
When a Node object is deleted and re-created while its pods keep running, the
shim and the core permanently disagree about the node's allocations, and queue
accounting is silently corrupted.
*Mechanism* (all verified in code):
# Node delete: the shim orphans the node's bound pods
({{{}Context.deleteNodeInternal{}}} -> {{{}SchedulerCache.RemoveNode{}}}); the
core removes the node, removes the allocations from the applications and
*decrements queue allocated resources*
({{{}partition.removeNodeAllocations{}}}), then notifies the shim with
{{{}TerminationType STOPPED_BY_RM{}}}.
# The shim callback intentionally only forgets the pod in the cache and does
not fail the task ({{{}scheduler_callback.go{}}} ~117) – this filter is correct
and must stay: dispatching the release would delete running pods.
# Node re-add: {{registerNodes}} sends a bare {{si.NodeInfo}} (it cannot carry
allocations – {{existingAllocations}} is a reserved/removed SI field), and the
orphan-adoption path re-attaches the pods to the shim cache via
{{updateYuniKornPod}} -> {{SetTaskPod}} on the still-Bound task – {*}nothing
re-reports the allocations to the core{*}.
# No reconciliation exists. The divergence lasts until each pod terminates, or
until a scheduler restart (restart recovery does re-report).
{*}The asymmetry that shows this is an oversight{*}: the same adoption path
already re-reports *foreign* pods to the core ({{{}updateForeignPod{}}} ->
{{CreateAllocationForForeignPod}} -> {{{}UpdateAllocation{}}}). Only the
YuniKorn-managed branch is silent, because the existing Bound task suppresses
task creation.
*Impact* :
- Node-level double allocation does NOT occur: the shim predicate callback
runs {{NodeResourcesFit}} against the shim cache, which contains the adopted
pods, so a genuinely full node rejects new placements (kubelet admission is a
second backstop).
- The real damage is {*}accounting integrity{*}: the queue's allocated
resource was decremented while the pods keep running, and is never
re-incremented. A queue with max=X can end up running more than X (survivors
plus a full X of new admissions); user quotas, fair-share ordering, preemption
inputs, autoscaler outstanding-request calculations, metrics and the web UI are
all silently wrong. The core also wastes scheduling cycles proposing a node it
believes is empty while the shim rejects every attempt.
{*}Reachability{*}: {{kubectl delete node}} on a live node (or automation doing
the equivalent) followed by kubelet self-re-registration recreates the Node
object within seconds; the pod-GC controller quarantines orphaned pods for ~40s
before deleting them, so pods surviving a brief flap is the normal outcome, not
a race. Cluster-autoscaler is not a trigger (it drains first). Rare but
ordinary operations; when it fires it affects every pod on the node, for an
unbounded duration on long-running workloads.
{*}Reproduction{*}: deterministic in a fault-injection test built on the
MockScheduler harness (bind churn + node remove/re-add through the informer
path): after settle, pods placed in the shim cache have no matching core
allocation (e.g. one flapped node: shim 40 pods, core 0 allocations). 3/3
flapped runs reproduce; 0/2 static controls. Test to be contributed.
{*}Suggested fix{*}: on adoption of a YuniKorn-managed pod whose task is Bound,
re-report the allocation the same way the foreign-pod branch and restart
recovery already do ({{{}UpdateAllocation{}}} with NodeID set -> the core's
{{MarkPreviouslyAllocated}} path). Alternatively (or additionally) reconsider
overloading {{STOPPED_BY_RM}} for core-initiated node-removal releases, so the
shim can distinguish them.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]