Dale Richardson created YUNIKORN-3355:
-----------------------------------------
Summary: Failed bind leaves a stale node assignment in the shim
cache
Key: YUNIKORN-3355
URL: https://issues.apache.org/jira/browse/YUNIKORN-3355
Project: Apache YuniKorn
Issue Type: Bug
Components: core - scheduler, shim - kubernetes
Reporter: Dale Richardson
If Pod bind calls fail while allocations are in flight (API server outage,
restart, or sustained errors under load), YuniKorn can permanently strand the
affected pods: they remain {{Pending}} in Kubernetes forever, while the core's
queue books them as {*}allocated{*}. The scheduler logs "No outstanding apps
found" and never retries. Pod metadata updates do not resurrect them; only
deleting/recreating the pods or restarting the scheduler recovers. No
WARN/ERROR is logged for the inconsistency.
This is very likely the root cause of YUNIKORN-3128 ("Yunikorn ignores pending
pods after apiserver errors"): we reproduced the identical end state
deterministically and traced the mechanism through the code.
h2. Reproduction (deterministic when timed right; reproduced on 1.9.0 and
master a82e4f92)
# kind cluster (K8s v1.36.1), 500 KWOK fake nodes, YuniKorn standard mode
(Helm defaults, admission controller disabled).
# Create 3,000 pods with {{schedulerName: yunikorn}} at high rate (~2,000
pods/s, client-go at concurrency 16).
# ~3s in, while Binding POSTs are in flight, {{kill -9}} the kube-apiserver;
kubelet restarts it ~10s later.
# Wait for recovery and observe.
The kill must land while binds are in flight — killing before binding starts
recovers cleanly. Any failure mode producing a burst of failed binds should
trigger it (YUNIKORN-3128 saw it from apiserver errors under load, with no full
outage).
h2. Observed end state (master build, 242 of 3,000 pods affected)
* 242 pods {{Pending}} with empty {{{}spec.nodeName{}}}, permanently (observed
>12 min, no recovery).
* Queue REST API: {{root.default}} shows {{allocatedResource pods: 3000}} —
the queue books all 3,000 {*}including the 242 never bound{*}; the applications
list is empty (app completed and was removed).
* Scheduler log: exactly 242 {{added existing allocation}} lines (1:1 with
stuck pods), each with a {{{}targetNode{}}}, ~30-45s after the outage; later
{{{}No outstanding apps found for a while{}}}.
* New pods (a different app) schedule normally — the damage is silent quota
leakage plus the stranded pods.
h2. Root cause (verified against master a82e4f92)
Four-step interaction between the shim cache's assumed-pod handling and the
core's existing-allocation semantics:
# *Assume stamps NodeName on the cached copy.* {{Context.AssumePod}}
deep-copies the pod, sets {{{}assumedPod.Spec.NodeName = node{}}}, and stores
it in the scheduler cache ({{{}pkg/cache/context.go:869-871{}}}). The cache
records the assignment in {{{}assignedPods{}}}.
# *ForgetPod does not undo the assignment.* On release after the failed bind,
{{Context.ForgetPod}} fetches the pod from the cache — the assumed copy with
NodeName set — and passes it to {{SchedulerCache.forgetPod}}
({{{}pkg/cache/context.go:879-884{}}}). {{forgetPod}} calls {{updatePod(pod)}}
with that copy and deletes only the {{assumedPods}} marker
({{{}pkg/cache/external/scheduler_cache.go:494-506{}}}). Inside {{updatePod}}
the copy still satisfies {{{}IsAssignedPod{}}}, so it is re-added to the node
and re-recorded in {{{}assignedPods{}}}. Net effect: after a failed bind the
cache permanently believes the pod is assigned.
# *The stale assignment is copied onto the real pod object.* On the next
informer event for the pod (watch reconnect / relist after the outage),
{{updatePod}} sees the incoming pod has {{Spec.NodeName == ""}} while
{{assignedPods}} has an entry, and executes {{pod.Spec.NodeName = nodeName}}
("use existing assignment",
{{{}pkg/cache/external/scheduler_cache.go:~355-358{}}}). Note this also
*mutates the shared informer-cache object* (pods are stored by reference via
{{{}utils.Convert2Pod{}}}), poisoning every other consumer of that object.
# *Task re-creation turns the stale assignment into a phantom placed
allocation.* The bind-failure storm failed the tasks and completed/removed the
application, so the informer event re-creates the app and task
({{{}ensureAppAndTaskCreated{}}}); {{Task.updateAllocation}} builds the
allocation with {{NodeID: task.pod.Spec.NodeName}}
({{{}pkg/cache/task.go:319-322{}}}, {{{}pkg/common/si_helper.go:134{}}}) — now
non-empty due to step 3. The core's {{PartitionContext.UpdateAllocation}}
treats any allocation with a NodeID as already placed ("handling existing
allocation" / "added existing allocation",
{{{}pkg/scheduler/partition.go:~1240-1259{}}}) and books it against the queue
without ever scheduling or binding it.
The pod is now invisible to the scheduler (no pending ask), unbound in
Kubernetes, and counted against queue quota. When the re-created app later
completes, the phantom allocations remain in queue accounting with no owning
application.
h2. Impact
* Pods stranded Pending indefinitely after any burst of bind failures; silent
— scheduler appears healthy.
* Queue quota silently leaked by phantom allocations; in quota-tight clusters
this can starve a queue completely.
* Shared informer objects mutated (step 3) — undefined behaviour for all other
informer consumers.
* Amplified by scheduler throughput: the faster the core allocates, the more
binds are in flight during a blip, the larger the phantom set (relevant to the
YUNIKORN-3350 throughput work).
h2. Suggested fixes (layered; any of the first three breaks the chain)
# {{forgetPod}} must actually revert the assignment: clear {{Spec.NodeName}}
on (a copy of) the pod before re-inserting, or restore the informer version,
and remove the {{assignedPods}} entry.
# {{updatePod}} must not copy a stale assignment onto an unassigned incoming
pod when that assignment came from an assumed (never confirmed bound) pod — and
must never mutate the incoming informer-owned object (copy-on-write instead).
# The shim should only report {{NodeID}} to the core from the *informer's*
view of {{spec.nodeName}} (ground truth of what is actually bound), never from
cache-internal assumed state.
# Defence in depth: bounded retry of transient bind failures (connection
refused / timeout / 429) before failing the task; an assumed-pod expiry
analogous to kube-scheduler's; a core-side invariant warning when queue
allocated resources exist with no owning application.
h2. Environment
kind v0.32 / K8s v1.36.1 single control plane, 500 KWOK fake nodes; reproduced
on YuniKorn 1.9.0 (Helm) and a master build (k8shim a82e4f92). Load generator
and repro script available on request.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]