[
https://issues.apache.org/jira/browse/YUNIKORN-3423?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Huangpinhan reassigned YUNIKORN-3423:
-------------------------------------
Assignee: (was: Huangpinhan)
> postAppAccepted reads taskGroups and taskMap without the application lock
> -------------------------------------------------------------------------
>
> Key: YUNIKORN-3423
> URL: https://issues.apache.org/jira/browse/YUNIKORN-3423
> Project: Apache YuniKorn
> Issue Type: Sub-task
> Components: shim - kubernetes
> Reporter: Dale Richardson
> Priority: Major
>
> {{Application.postAppAccepted}} is called from {{Schedule()}} on the
> scheduling ticker with no lock, outside any state-machine transition, unlike
> the other state handlers which are FSM callbacks under {{app.handle}}'s lock.
> It reads {{taskGroups}} and, through {{skipReservationStage}}, ranges over
> {{taskMap}} while the pod informer path keeps adding tasks to that map under
> the lock; pods arrive for an application after it is Accepted as a matter of
> course.
> Confirmed with the race detector: the test on the fork fix branch reproduces
> it on master every time (three reports on this chain). Because the read is a
> {{range}}, the runtime consequence is {{fatal error: concurrent map iteration
> and map write}}, which kills the scheduler.
> Fix: snapshot the decision inputs under one {{RLock}} and release it before
> logging and dispatching. Wrapping the whole body self-deadlocks because
> {{GetAllocatedTasks()}} takes the read lock internally; use the lock-free
> {{getTasks}} twin inside the locked region. A fix exists on the tigerquoll
> fork (PR #12 there) and will be filed once this has a number.
> Marker: {{postAppAccepted}} in {{application.go}} carries this JIRA; the fix
> removes it.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]