adityadtu5 commented on PR #1052: URL: https://github.com/apache/yunikorn-k8shim/pull/1052#issuecomment-5178109204
> 1. Fix the flaky test > 2. **Follow-up:** don't fetch PodVolumes with FindPodVolumes before reverting. This object should be stored separately in Context or in the scheduler cache. It's not a problem right now, but if we want to change the rollback behavior to be in-line with the default scheduler, then we have to perform a revert in case of a bind failure (we don't do this atm). So if we do it to achieve K8s parity (which is desirable), then the revert operation can fail inside the volumeBinder's cache. > 3. **Follow-up:** we return nil from AssumePod if either the node/pod is not found. I just spotted this. This is a pre-existing problem, not related to this change. Investigate this in a separate JIRA (low prio). > > More details on 1) > > Revert scenarios > > Scenario Assume happened? Re-fetch works? > All retries fail at AssumePodVolumes No (self-reverted) N/A > Assume succeeds, bind proceeds Yes N/A (no revert needed) > Assume succeeds, then core release -> ForgetPod (mostly a non-issue for us) Yes Risky, especially dynamic PVCs > Assume succeeds, bind fails, cleanup needed Yes Risky > So the gap we have is "Assume succeeds, bind fails, cleanup needed". This is worth implementing as a follow-up, but it has the PodVolumes impact. 1) I tried running locally the test and have not faced any failures. The application state assertion is not very important for this flow, so I will remove it. 2) We can not save the pod volumes into scheduler cache as revert needs to work even if yunikorn lose its state in between. 3) Yes there are jiras in place to handle volume and pod bind failures. See epic - YUNIKORN-2804 for more details. They will be done once this PR is merged as this PR is base for them. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
