[ 
https://issues.apache.org/jira/browse/YUNIKORN-3430?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

jimmycmlin reassigned YUNIKORN-3430:
------------------------------------

    Assignee: jimmycmlin

> registerNodesInternal releases and retakes the context lock around the wait 
> group
> ---------------------------------------------------------------------------------
>
>                 Key: YUNIKORN-3430
>                 URL: https://issues.apache.org/jira/browse/YUNIKORN-3430
>             Project: Apache YuniKorn
>          Issue Type: Sub-task
>          Components: shim - kubernetes
>            Reporter: Dale Richardson
>            Assignee: jimmycmlin
>            Priority: Minor
>
> {{Context.registerNodesInternal}} is documented as requiring the context 
> write lock, yet it unlocks and re-locks {{ctx.lock}} around its wait for the 
> core's node-accepted events, splitting the caller's critical section in two. 
> {{updateNodeInternal}} continues after the call with adopted-pod registration 
> and {{enableNode}} on state it established before the gap, and the wait has 
> no timeout (YUNIKORN-3239 saw it last 53 minutes).
> No corruption has been shown; the hazard is that nothing re-validates the 
> invariants after the wait, and the analyser has to exclude the whole function 
> from checking.
> Fix: build the request under the lock, release, send and wait, then re-take 
> the lock to apply the results, instead of unlocking inside the callee.
> Marker: {{registerNodesInternal}} in {{context.go}} carries a function-level 
> {{+checklocksignore}} for this JIRA; the fix removes it.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to