Michael Semb Wever created CASSANDRA-21576:
----------------------------------------------

             Summary: CI: Prevent/kill Jenkins k8s agent orphans and 
unschedulable agent pods
                 Key: CASSANDRA-21576
                 URL: https://issues.apache.org/jira/browse/CASSANDRA-21576
             Project: Apache Cassandra
          Issue Type: Improvement
          Components: CI
            Reporter: Michael Semb Wever


node pools can get stuck at max desired size, wasting a lot of resources and $$$

Either the instanceCap goes beyond the node group sizes, and it thrashes, or 
(for example) the controller fails its liveness probe, gets SIGKILLed, then 
orphans agent pods forever.  Manually trying to scale down the node groups 
won't help the as the jenkins keeps trying to scale up again…


 patch: 
https://github.com/apache/cassandra/compare/cassandra-5.0...thelastpickle:cassandra:mck/fix_ci_gha/5.0



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to