[
https://issues.apache.org/jira/browse/YUNIKORN-3373?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated YUNIKORN-3373:
-------------------------------------
Labels: pull-request-available (was: )
> Add an invariant property fuzzer for the shim scheduler cache
> -------------------------------------------------------------
>
> Key: YUNIKORN-3373
> URL: https://issues.apache.org/jira/browse/YUNIKORN-3373
> Project: Apache YuniKorn
> Issue Type: Bug
> Components: shim - kubernetes
> Reporter: Dale Richardson
> Priority: Major
> Labels: pull-request-available
>
> The scheduler cache in the shim ({{pkg/cache/external}}) tracks pods across
> informer events, the assume/forget cycle of scheduling, and node lifecycle.
> Defects in this bookkeeping surface as phantom allocations and wedged pods,
> are only reachable through unusual event orderings (typically after a
> failure), and have historically required cluster-level reproduction to find:
> YUNIKORN-3355 was diagnosed from a 10 second apiserver outage during a 10,000
> pod bind burst on a test rig.
> This adds a randomized property test that drives the operations the cache
> sees in production -- informer add/update/delete with and without an
> assignment, assume, forget, node add and remove -- against a small reference
> model, and asserts six invariants after every single operation:
> # an assumed pod is visible on the node it is assumed on
> # a pod that never bound is not assigned to a node
> # the assignment map and the node infos hold the same view
> # objects handed to the cache are never mutated (they can be owned by the
> informer cache)
> # scheduling state only covers cached pods
> # the cache matches the reference model
> Sequences are generated from fixed seeds and are fully deterministic; a
> failure reports the seed, the step number and the tail of the operation
> history, so it replays exactly. An extra seed can be supplied through
> {{YUNIKORN_CACHE_FUZZ_SEED}} for exploratory runs.
> *The fuzzer currently FAILS on master.* On every seed, within a few dozen
> operations, it finds two live defects:
> # YUNIKORN-3355: a forgotten pod keeps its node assignment and informer
> updates are re-stamped with it. The minimal sequence found is three
> operations: update (pending pod), assume, forget. The fuzzer also
> independently catches that path mutating the pod object owned by the informer
> cache.
> # A pod assumed on a node that is removed is orphaned still carrying the node
> name the shim stamped on it at assume time; when the node comes back the
> orphan adoption path re-assigns the never-bound pod to it and task recovery
> reports an existing allocation to the core for a bind that never happened.
> (Separate bug report to follow; link to be added here.)
> The PR is therefore a draft until both fixes land; after that it passes and
> becomes a permanent regression guard for both defects and for the cache
> invariants generally.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]