adityadtu5 commented on code in PR #1062:
URL: https://github.com/apache/yunikorn-k8shim/pull/1062#discussion_r3774435759
##########
pkg/cache/task.go:
##########
@@ -669,6 +709,22 @@ func (task *Task)
rollbackOnAssumePodFailure(allocationKey, nodeID string) {
zap.String("allocationKey", allocationKey))
}
+// rescheduleOnBindFailure is called when volume or pod binding fails after
all retries.
+// Move the task back to Scheduling before rolling back the allocation, so a
+// subsequent TaskAllocated event from the core (which requires Scheduling) is
accepted.
+// Must be called without holding the task lock.
+func (task *Task) rescheduleOnBindFailure(allocationKey, nodeID, eventReason,
eventMsg string) {
+ // Move the task back to Scheduling before releasing to the core, so
the re-delivered
+ // allocation (valid only from the Scheduling state) is accepted by the
state machine.
+ if err := task.handle(NewRescheduleTaskEvent(task.applicationID,
task.taskID)); err != nil {
Review Comment:
I think it is better to directly update the state as this all is happening
on shim side only. Making this change.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]