wilfred-s commented on code in PR #1062:
URL: https://github.com/apache/yunikorn-k8shim/pull/1062#discussion_r3773024357
##########
pkg/cache/task.go:
##########
@@ -669,6 +709,22 @@ func (task *Task)
rollbackOnAssumePodFailure(allocationKey, nodeID string) {
zap.String("allocationKey", allocationKey))
}
+// rescheduleOnBindFailure is called when volume or pod binding fails after
all retries.
+// Move the task back to Scheduling before rolling back the allocation, so a
+// subsequent TaskAllocated event from the core (which requires Scheduling) is
accepted.
+// Must be called without holding the task lock.
+func (task *Task) rescheduleOnBindFailure(allocationKey, nodeID, eventReason,
eventMsg string) {
+ // Move the task back to Scheduling before releasing to the core, so
the re-delivered
+ // allocation (valid only from the Scheduling state) is accepted by the
state machine.
+ if err := task.handle(NewRescheduleTaskEvent(task.applicationID,
task.taskID)); err != nil {
Review Comment:
You must hold the task lock to do the (volume)binding. The calling go
routine has the write lock. The handle will block as it needs a read lock.
The pattern used in the k8shim is not good for this. You either need to
dispatch an event or simply change the state back via
`task.sm.SetState(states.Scheduling)`. Preference is a direct state setting as
it has less overhead than an event dispatch.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]