elihschiff opened a new pull request, #558:
URL: https://github.com/apache/yunikorn-k8shim/pull/558

   ### What is this PR for?
   In my cluster I am seeing occasional events like this on some of my pods 
which causes them to get stuck.
   
   `54m     Warning  OutOfpods           
pod/tg-spark-executor-640b4349263cc74570ae3a1e-0     Node didn't have enough 
resource: pods, requested: 1, used: 12, capacity: 12`
   
   From what I can tell, nodes all have a `pods` resource value on them. 
Because this value already exists on nodes, my PR adds a pods=1 value to add 
pods inside the scheduler. Then yunikorn uses the same resource scheduling 
logic used for memory, cpu, ect. to limit the number of pods running on a node.
   
   ### What type of PR is it?
   * [x] - Bug Fix
   * [ ] - Improvement
   * [ ] - Feature
   * [ ] - Documentation
   * [ ] - Hot Fix
   * [ ] - Refactoring
   
   ### Todos
   * [ ] - This change only works once the core repo is updated 
https://github.com/apache/yunikorn-core/pull/521. Therefore that change will 
need to be merged first.
   
   ### What is the Jira issue?
   https://issues.apache.org/jira/browse/YUNIKORN-1559
   
   ### How should this be tested?
   
   I am not 100% sure how best to test this. I fixed the tests including the 
e2e tests. I am testing this on my kubernetes clusters and it seems to have 
fixed the issue but I am not 100% sure if there is anything I should be worried 
about.
   
   ### Screenshots (if appropriate)
   
   ### Questions:
   * [ ] - The licenses files need update.
   * [ ] - There is breaking changes for older versions.
   * [ ] - It needs documentation.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to