[
https://issues.apache.org/jira/browse/YUNIKORN-3355?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated YUNIKORN-3355:
-------------------------------------
Labels: pull-request-available (was: )
> Failed bind leaves a stale node assignment in the shim cache
> ------------------------------------------------------------
>
> Key: YUNIKORN-3355
> URL: https://issues.apache.org/jira/browse/YUNIKORN-3355
> Project: Apache YuniKorn
> Issue Type: Bug
> Components: core - scheduler, shim - kubernetes
> Reporter: Dale Richardson
> Assignee: Aditya Maheshwari
> Priority: Major
> Labels: pull-request-available
>
> If Pod bind calls fail while allocations are in flight (API server outage,
> restart, or sustained errors under load), YuniKorn can permanently strand the
> affected pods: they remain {{Pending}} in Kubernetes forever, while the
> core's queue books them as {*}allocated{*}. The scheduler logs "No
> outstanding apps found" and never retries. Pod metadata updates do not
> resurrect them; only deleting/recreating the pods or restarting the scheduler
> recovers. No WARN/ERROR is logged for the inconsistency.
> This is very likely the root cause of YUNIKORN-3128 ("Yunikorn ignores
> pending pods after apiserver errors"): we reproduced the identical end state
> deterministically and traced the mechanism through the code.
> h2. Reproduction (deterministic when timed right; reproduced on 1.9.0 and
> master a82e4f92)
> # kind cluster (K8s v1.36.1), 500 KWOK fake nodes, YuniKorn standard mode
> (Helm defaults, admission controller disabled).
> # Create 3,000 pods with {{schedulerName: yunikorn}} at high rate (~2,000
> pods/s, client-go at concurrency 16).
> # ~3s in, while Binding POSTs are in flight, {{kill -9}} the kube-apiserver;
> kubelet restarts it ~10s later.
> # Wait for recovery and observe.
> The kill must land while binds are in flight — killing before binding starts
> recovers cleanly. Any failure mode producing a burst of failed binds should
> trigger it (YUNIKORN-3128 saw it from apiserver errors under load, with no
> full outage).
> h2. Observed end state (master build, 242 of 3,000 pods affected)
> * 242 pods {{Pending}} with empty {{{}spec.nodeName{}}}, permanently
> (observed >12 min, no recovery).
> * Queue REST API: {{root.default}} shows {{allocatedResource pods: 3000}} —
> the queue books all 3,000 {*}including the 242 never bound{*}; the
> applications list is empty (app completed and was removed).
> * Scheduler log: exactly 242 {{added existing allocation}} lines (1:1 with
> stuck pods), each with a {{{}targetNode{}}}, ~30-45s after the outage; later
> {{{}No outstanding apps found for a while{}}}.
> * New pods (a different app) schedule normally — the damage is silent quota
> leakage plus the stranded pods.
> h2. Root cause (verified against master a82e4f92)
> Four-step interaction between the shim cache's assumed-pod handling and the
> core's existing-allocation semantics:
> # *Assume stamps NodeName on the cached copy.* {{Context.AssumePod}}
> deep-copies the pod, sets {{{}assumedPod.Spec.NodeName = node{}}}, and stores
> it in the scheduler cache ({{{}pkg/cache/context.go:869-871{}}}). The cache
> records the assignment in {{{}assignedPods{}}}.
> # *The stale node assignment survives task teardown — by either of two
> code-verified paths.* Our repro logs cannot distinguish which ran (container
> log rotation consumed the failure window), but both preserve the assignment:
> *# The release callback ran.* {{Context.ForgetPod}} fetches the pod from the
> scheduler cache — which returns the *assumed copy* with {{Spec.NodeName}} set
> — and passes it to {{SchedulerCache.forgetPod}}
> ({{pkg/cache/context.go:879-884}}). {{forgetPod}} deletes only the
> {{assumedPods}} marker and calls {{updatePod(pod)}} with that copy
> ({{pkg/cache/external/scheduler_cache.go:494-506}}); inside {{updatePod}} the
> copy still satisfies {{IsAssignedPod}}, so it is re-added to the node and
> re-recorded in {{assignedPods}}. The "forget" is a no-op for the assignment.
> *# The release callback never ran.* Under a mass bind-failure storm the async
> event pipeline can drop or indefinitely delay the release delivery (the
> drop-on-full behaviour of {{rmproxy.enqueueAndCheckFull}} /
> {{dispatcher.asyncDispatch}}), in which case {{ForgetPod}} is never invoked
> and the assumed state trivially survives — still marked as an active assume.
> # Either way, the invariant that matters is: *no path in bind-failure
> teardown removes the node assignment (or clears the stamped {{NodeName}})
> from the scheduler cache.* The subsequent steps are identical for both
> branches.
> ## The stale assignment is copied onto the real pod object.* On the next
> informer event for the pod (watch reconnect / relist after the outage),
> {{updatePod}} sees the incoming pod has {{Spec.NodeName == ""}} while
> {{assignedPods}} has an entry, and executes {{pod.Spec.NodeName = nodeName}}
> ("use existing assignment",
> {{{}pkg/cache/external/scheduler_cache.go:~355-358{}}}). Note this also
> *mutates the shared informer-cache object* (pods are stored by reference via
> {{{}utils.Convert2Pod{}}}), poisoning every other consumer of that object.
> ## Task re-creation turns the stale assignment into a phantom placed
> allocation.* The bind-failure storm failed the tasks and completed/removed
> the application, so the informer event re-creates the app and task
> ({{{}ensureAppAndTaskCreated{}}}); {{Task.updateAllocation}} builds the
> allocation with {{NodeID: task.pod.Spec.NodeName}}
> ({{{}pkg/cache/task.go:319-322{}}}, {{{}pkg/common/si_helper.go:134{}}}) —
> now non-empty due to step 3. The core's {{PartitionContext.UpdateAllocation}}
> treats any allocation with a NodeID as already placed ("handling existing
> allocation" / "added existing allocation",
> {{{}pkg/scheduler/partition.go:~1240-1259{}}}) and books it against the queue
> without ever scheduling or binding it.
> The pod is now invisible to the scheduler (no pending ask), unbound in
> Kubernetes, and counted against queue quota. When the re-created app later
> completes, the phantom allocations remain in queue accounting with no owning
> application.
> h2. Impact
> * Pods stranded Pending indefinitely after any burst of bind failures;
> silent — scheduler appears healthy.
> * Queue quota silently leaked by phantom allocations; in quota-tight
> clusters this can starve a queue completely.
> * Shared informer objects mutated (step 3) — undefined behaviour for all
> other informer consumers.
> * Amplified by scheduler throughput: the faster the core allocates, the more
> binds are in flight during a blip, the larger the phantom set (relevant to
> the YUNIKORN-3350 throughput work).
> h2. Suggested fixes (revised)
> # *Primary:* derive {{NodeID}} for core registration only from the
> *informer's* view of {{spec.nodeName}} (ground truth of what is bound), never
> from cache-internal assumed state ({{si_helper.go:134}} call sites). This
> breaks the chain in both branches above.
> # *Backstop for branch (b):* expire assumed pods that receive no bind
> confirmation (kube-scheduler prior art, ~30s TTL) so lost release callbacks
> cannot leave permanent assumed state.
> # *Hygiene for branch (a):* {{forgetPod}} must actually revert the assignment
> (clear {{NodeName}} on the copy / restore the informer version, remove the
> {{assignedPods}} entry); and {{updatePod}} should stamp an existing
> assignment only while the pod is still actively assumed ({{assumedPods}}
> entry present) — and must never mutate the informer-owned object.
> # *Root cause of branch (b), separate scope:* make the release/allocation
> event pipeline bounded-blocking instead of drop-on-full so teardown signals
> cannot be silently lost.
> # Defence in depth unchanged: bounded transient-error bind retry
> (YUNIKORN-2804), and a core-side invariant warning when queue allocations
> exist with no owning application.
> h2. Environment
> kind v0.32 / K8s v1.36.1 single control plane, 500 KWOK fake nodes;
> reproduced on YuniKorn 1.9.0 (Helm) and a master build (k8shim a82e4f92).
> Load generator and repro script available on request.
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]