robertpi opened a new issue, #863:
URL: https://github.com/apache/spark-kubernetes-operator/issues/863

   **Component:** spark-kubernetes-operator (reconciler / driver observers)     
                                                                                
                                                   
   **Version:** 0.8.0                                                           
                                                                                
          
   **API:** `spark.apache.org/v1`                                               
                                                                                
                                                   
   **Reproduce difficulty:** Medium — any pre-container-start pod death on a 
scheduler with fast pod cleanup reproduces it.                                  
                                                      
                                                                                
                                                                                
                                                   
   ## Problem                                                                   
                                                                                
                                                   
                                                                                
                                                                                
                                                   
   A `SparkApplication` can remain stuck in `DriverRequested` forever when the 
driver Pod reaches a terminal `Failed` phase **before any container starts**. 
Kubernetes then reports the pod as `Failed` with an **empty 
`containerStatuses`** (no container ever initialized), which the operator's 
`observeDriverTermination()` misreads as "pod may be in pending state" and 
ignores. The CR is never reconciled to `Failed`, and — with `restartPolicy: 
Never` / `maxRestartAttempts: 0` — nothing retries it. The job stays wedged 
indefinitely (7h+ in our productions) while owning upstream job/lock state.
                                                                                
                                                                   
   ## Reproduction scenario (two independent triggers)
                                                                                
                                        
   1. Submit a `SparkApplication`. Operator creates the driver Pod, state → 
`DriverRequested`.
   2. The driver Pod is scheduled onto a node but the container **never 
starts**.
   3. The node then either:
      - **A)** is deleted/disappears — Kubernetes PodGC marks the orphaned pod 
`Failed`, condition `DisruptionTarget`, reason `DeletionByPodGC`, message 
`"PodGC: node no longer exists"`; or
      - **B)** hits node **DiskPressure** — the pod is evicted with `reason: 
Evicted`.
   4. Resulting driver pod status: `phase: Failed`, `containerStatuses: []`, no 
`startTime`, no pod IP.
   5. Operator keeps the `SparkApplication` in `DriverRequested`; no state 
transition ever occurs.
   
   ## Root cause
   
   In 
`spark-operator/src/main/java/org/apache/spark/k8s/operator/reconciler/observers/BaseAppDriverObserver.java`
 (tag `0.8.0`), `observeDriverTermination()`:
   
   ```java
   protected Optional<ApplicationState> observeDriverTermination(
       final Pod driverPod, final boolean driverReady, final ApplicationSpec 
spec) {
     PodStatus status = driverPod.getStatus();
     if (status == null
         || status.getContainerStatuses() == null
         || status.getContainerStatuses().isEmpty()) {
       log.debug("Cannot determine driver pod status, the pod may in pending 
state.");
       return Optional.empty();            // <-- returns before checking pod 
PHASE
     }
   
     if (PodPhase.FAILED == PodPhase.getPhase(driverPod)) { ... }   // 
unreachable when containerStatuses is empty
     if (PodPhase.SUCCEEDED == PodPhase.getPhase(driverPod)) { ... }
   ```
   
   The guard treats "empty `containerStatuses`" as *pending* and returns before 
ever inspecting the pod **phase**. A terminal `Failed` pod with zero containers 
(pod GC'd before first start, or evicted before first start) is therefore never 
detected as terminated. `AppDriverReadyObserver`/`AppDriverStartObserver` then 
can't advance the application, and the `DriverRequested → DriverStartTimedOut` 
transition never happens within a reasonable window.
   
   ## Expected behavior
   
   If the driver pod's **phase** is `Failed` (resp. `Succeeded`), the operator 
should transition the application to `Failed`/`Succeeded` **regardless of 
`containerStatuses`**, since no container may have ever started. 
`containerStatuses` should only be consulted to refine the reason (e.g. 
`DriverEvicted`) or to classify per-container failures when containers did run.
   
   ## Proposed fix
                                                                                
                                                                                
                        
   Reorder the phase checks ahead of the empty-`containerStatuses` guard:
                                                                                
    
   ```java                                                                      
                                                                                
                                                   
   if (status == null) {                                                        
                                                                                
                                                   
     log.debug("Cannot determine driver pod status, status is null.");          
                                                                                
                                                   
     return Optional.empty();                                                   
                                                                                
                                                   
   }                                                                            
                                                                                
                                                   
   
   if (PodPhase.FAILED == PodPhase.getPhase(driverPod)) {
     // ... Failed / DriverEvicted as today
   }
   if (PodPhase.SUCCEEDED == PodPhase.getPhase(driverPod)) {
     // ... Succeeded as today
   }
   
   if (status.getContainerStatuses() == null || 
status.getContainerStatuses().isEmpty()) {
     log.debug("Cannot determine driver pod status, the pod may be in pending 
state.");  return Optional.empty();
   }
   // existing container-status-based logic...
   ```
   
   Consider also treating the `DisruptionTarget`/`DeletionByPodGC` condition as 
evidence of a terminal driver pod (node no longer exists), independent of 
container status.
   
   ## Impact
   
   - Jobs hold upstream orchestration state (job locks / status) for 
hours-to-days because the CR never transitions.
   - Any pod death before container start reproduces it: node cleanup (PodGC), 
disk/memory eviction, spot/preemptible reclaim, node NotReady+taint eviction.
   - The same class of defect affects executors ("executors never started") if 
`instanceConfig` ever gets enabled, since 
`ExecutorStartTimeoutMillis`/`ExecutorsStartTimedOut` depend on the operator 
observing real pod state.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to