Dale Richardson created YUNIKORN-3355:
-----------------------------------------

             Summary: Failed bind leaves a stale node assignment in the shim 
cache
                 Key: YUNIKORN-3355
                 URL: https://issues.apache.org/jira/browse/YUNIKORN-3355
             Project: Apache YuniKorn
          Issue Type: Bug
          Components: core - scheduler, shim - kubernetes
            Reporter: Dale Richardson


If Pod bind calls fail while allocations are in flight (API server outage, 
restart, or sustained errors under load), YuniKorn can permanently strand the 
affected pods: they remain {{Pending}} in Kubernetes forever, while the core's 
queue books them as {*}allocated{*}. The scheduler logs "No outstanding apps 
found" and never retries. Pod metadata updates do not resurrect them; only 
deleting/recreating the pods or restarting the scheduler recovers. No 
WARN/ERROR is logged for the inconsistency.

This is very likely the root cause of YUNIKORN-3128 ("Yunikorn ignores pending 
pods after apiserver errors"): we reproduced the identical end state 
deterministically and traced the mechanism through the code.
h2. Reproduction (deterministic when timed right; reproduced on 1.9.0 and 
master a82e4f92)
 # kind cluster (K8s v1.36.1), 500 KWOK fake nodes, YuniKorn standard mode 
(Helm defaults, admission controller disabled).
 # Create 3,000 pods with {{schedulerName: yunikorn}} at high rate (~2,000 
pods/s, client-go at concurrency 16).
 # ~3s in, while Binding POSTs are in flight, {{kill -9}} the kube-apiserver; 
kubelet restarts it ~10s later.
 # Wait for recovery and observe.

The kill must land while binds are in flight — killing before binding starts 
recovers cleanly. Any failure mode producing a burst of failed binds should 
trigger it (YUNIKORN-3128 saw it from apiserver errors under load, with no full 
outage).
h2. Observed end state (master build, 242 of 3,000 pods affected)
 * 242 pods {{Pending}} with empty {{{}spec.nodeName{}}}, permanently (observed 
>12 min, no recovery).
 * Queue REST API: {{root.default}} shows {{allocatedResource pods: 3000}} — 
the queue books all 3,000 {*}including the 242 never bound{*}; the applications 
list is empty (app completed and was removed).
 * Scheduler log: exactly 242 {{added existing allocation}} lines (1:1 with 
stuck pods), each with a {{{}targetNode{}}}, ~30-45s after the outage; later 
{{{}No outstanding apps found for a while{}}}.
 * New pods (a different app) schedule normally — the damage is silent quota 
leakage plus the stranded pods.

h2. Root cause (verified against master a82e4f92)

Four-step interaction between the shim cache's assumed-pod handling and the 
core's existing-allocation semantics:
 # *Assume stamps NodeName on the cached copy.* {{Context.AssumePod}} 
deep-copies the pod, sets {{{}assumedPod.Spec.NodeName = node{}}}, and stores 
it in the scheduler cache ({{{}pkg/cache/context.go:869-871{}}}). The cache 
records the assignment in {{{}assignedPods{}}}.
 # *ForgetPod does not undo the assignment.* On release after the failed bind, 
{{Context.ForgetPod}} fetches the pod from the cache — the assumed copy with 
NodeName set — and passes it to {{SchedulerCache.forgetPod}} 
({{{}pkg/cache/context.go:879-884{}}}). {{forgetPod}} calls {{updatePod(pod)}} 
with that copy and deletes only the {{assumedPods}} marker 
({{{}pkg/cache/external/scheduler_cache.go:494-506{}}}). Inside {{updatePod}} 
the copy still satisfies {{{}IsAssignedPod{}}}, so it is re-added to the node 
and re-recorded in {{{}assignedPods{}}}. Net effect: after a failed bind the 
cache permanently believes the pod is assigned.
 # *The stale assignment is copied onto the real pod object.* On the next 
informer event for the pod (watch reconnect / relist after the outage), 
{{updatePod}} sees the incoming pod has {{Spec.NodeName == ""}} while 
{{assignedPods}} has an entry, and executes {{pod.Spec.NodeName = nodeName}} 
("use existing assignment", 
{{{}pkg/cache/external/scheduler_cache.go:~355-358{}}}). Note this also 
*mutates the shared informer-cache object* (pods are stored by reference via 
{{{}utils.Convert2Pod{}}}), poisoning every other consumer of that object.
 # *Task re-creation turns the stale assignment into a phantom placed 
allocation.* The bind-failure storm failed the tasks and completed/removed the 
application, so the informer event re-creates the app and task 
({{{}ensureAppAndTaskCreated{}}}); {{Task.updateAllocation}} builds the 
allocation with {{NodeID: task.pod.Spec.NodeName}} 
({{{}pkg/cache/task.go:319-322{}}}, {{{}pkg/common/si_helper.go:134{}}}) — now 
non-empty due to step 3. The core's {{PartitionContext.UpdateAllocation}} 
treats any allocation with a NodeID as already placed ("handling existing 
allocation" / "added existing allocation", 
{{{}pkg/scheduler/partition.go:~1240-1259{}}}) and books it against the queue 
without ever scheduling or binding it.

The pod is now invisible to the scheduler (no pending ask), unbound in 
Kubernetes, and counted against queue quota. When the re-created app later 
completes, the phantom allocations remain in queue accounting with no owning 
application.
h2. Impact
 * Pods stranded Pending indefinitely after any burst of bind failures; silent 
— scheduler appears healthy.
 * Queue quota silently leaked by phantom allocations; in quota-tight clusters 
this can starve a queue completely.
 * Shared informer objects mutated (step 3) — undefined behaviour for all other 
informer consumers.
 * Amplified by scheduler throughput: the faster the core allocates, the more 
binds are in flight during a blip, the larger the phantom set (relevant to the 
YUNIKORN-3350 throughput work).

h2. Suggested fixes (layered; any of the first three breaks the chain)
 # {{forgetPod}} must actually revert the assignment: clear {{Spec.NodeName}} 
on (a copy of) the pod before re-inserting, or restore the informer version, 
and remove the {{assignedPods}} entry.
 # {{updatePod}} must not copy a stale assignment onto an unassigned incoming 
pod when that assignment came from an assumed (never confirmed bound) pod — and 
must never mutate the incoming informer-owned object (copy-on-write instead).
 # The shim should only report {{NodeID}} to the core from the *informer's* 
view of {{spec.nodeName}} (ground truth of what is actually bound), never from 
cache-internal assumed state.
 # Defence in depth: bounded retry of transient bind failures (connection 
refused / timeout / 429) before failing the task; an assumed-pod expiry 
analogous to kube-scheduler's; a core-side invariant warning when queue 
allocated resources exist with no owning application.

h2. Environment

kind v0.32 / K8s v1.36.1 single control plane, 500 KWOK fake nodes; reproduced 
on YuniKorn 1.9.0 (Helm) and a master build (k8shim a82e4f92). Load generator 
and repro script available on request.

 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to