[ 
https://issues.apache.org/jira/browse/FLINK-40226?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18098564#comment-18098564
 ] 

Aniruddh J commented on FLINK-40226:
------------------------------------

If this looks like a valid case, I would like to fix it. 

> Standalone session cluster JM recovery enters infinite MISSING loop and fails 
> with HTTP 409 on TM Deployment when TaskManager pods still exist
> ----------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40226
>                 URL: https://issues.apache.org/jira/browse/FLINK-40226
>             Project: Flink
>          Issue Type: Bug
>          Components: Kubernetes Operator
>    Affects Versions: 1.11.0, 1.14.0
>         Environment: Operator version  1.11.0, 1.14.0 (the ones I have tested 
> it on)
> Flink version1.20.x (standalone mode)
> Cluster typeStandalone Kubernetes session cluster ({{{}FlinkDeployment{}}} 
> with no {{{}spec.job{}}})
> Kubernetes reproduced on OpenShift and vanilla Kubernetes
> HA Enabled 
>            Reporter: Aniruddh J
>            Priority: Major
>
> When the *JobManager (JM) Deployment* of a standalone Kubernetes session 
> cluster is deleted externally (e.g. by an admin) while the {*}TaskManager 
> (TM) Deployment and its pods are still running or terminating{*}, the 
> operator enters a state where it can never successfully recover the session 
> cluster. Two distinct but related defects combine to produce this outcome:
>  # *Bug 1 — Infinite MISSING status loop on submit failure.* 
> {{recoverSession()}} in {{SessionReconciler}} does not set the 
> {{JobManagerDeploymentStatus}} to {{DEPLOYING}} before calling 
> {{{}submitSessionCluster(){}}}. If the submit call throws, the controller's 
> error-handling path resets the status back to {{{}MISSING{}}}, causing 
> {{shouldRecoverDeployment()}} to fire on every subsequent reconcile cycle — 
> producing an unbound retry loop with no back-off.
>  # *Bug 2 — HTTP 409 AlreadyExists on TM Deployment re-creation.* 
> {{deployClusterInternal()}} in {{KubernetesStandaloneClusterDescriptor}} uses 
> a bare Kubernetes {{CREATE}} (HTTP POST) for the TaskManager Deployment via 
> {{{}Fabric8FlinkStandaloneKubeClient.createTaskManagerDeployment(){}}}. If 
> the TM Deployment already exists — which is the normal state when only the JM 
> was deleted — the API server returns {{{}HTTP 409 Conflict 
> (AlreadyExists){}}}, causing every recovery attempt to fail. This failure 
> feeds back into Bug 1.
>  
> Together, the operator logs a continuous stream of reconciliation errors and 
> the session cluster remains permanently unrecoverable until manual operator 
> intervention.
>  
>  
> h2. Steps to Reproduce
>  # Deploy a standalone Kubernetes session cluster (HA mode enabled) using 
> {{{}FlinkDeployment{}}}. Wait for {{{}jobManagerDeploymentStatus: READY{}}}.
>  # Manually delete only the *JM Deployment* from the namespace:
> {{kubectl delete deployment <cluster-id> -n <namespace>}}
>  # Do _not_ delete the TM Deployment or its pods — allow them to continue 
> running.
>  # Observe the operator logs. The operator detects {{MISSING}} and calls 
> {{{}recoverSession(){}}}.
>  # {{deployClusterInternal()}} attempts to re-create the TM Deployment via a 
> bare HTTP POST. Because the TM Deployment still exists, the API server 
> returns {{{}HTTP 409 AlreadyExists{}}}.
>  # The 409 exception propagates through {{submitSessionCluster()}} → 
> {{recoverSession()}} → controller error handler.
>  # The controller error handler resets the status to {{MISSING}} from the 
> in-memory cache (which was never updated to {{{}DEPLOYING{}}}).
>  # On the next reconcile cycle (15 s default), {{shouldRecoverDeployment()}} 
> evaluates to {{true}} again, and the loop repeats indefinitely.
>  
> h2. Expected Behaviour
>  # When a JM-recovery submit fails for any transient reason (network error, 
> 409, throttling), the {{jobManagerDeploymentStatus}} should be advanced to 
> {{DEPLOYING}} _before_ the submit attempt, so that the error path does not 
> restore {{MISSING}} and re-trigger an immediate retry.
>  # A re-submit when only the JM Deployment is absent and the TM Deployment 
> still exists should succeed without a 409. The desired TM spec is unchanged, 
> so an idempotent create-or-replace is semantically correct.
>  # A failed recovery attempt should not proactively delete or disturb 
> still-running TM pods as part of its error rollback.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to