Gyula Fora created FLINK-40954:
----------------------------------
Summary: Force delete pods stuck Terminating on cluster resource
cleanup timeout
Key: FLINK-40954
URL: https://issues.apache.org/jira/browse/FLINK-40954
Project: Flink
Issue Type: Improvement
Components: Kubernetes Operator
Reporter: Gyula Fora
Assignee: Gyula Fora
deleteClusterInternal blocks on the JobManager/TaskManager Deployment
disappearing, which never happens if one of its pods is stuck Terminating
(e.g. its node went NotReady and the kubelet can never confirm the
deletion). Today this times out and the identical wait is retried forever,
both on the suspend path (leaves the deployment at zero JM/TM) and on the
CR cleanup path (pins the finalizer and blocks all future redeploys under
the same name).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)