MJ Deng created SPARK-59972:
-------------------------------

             Summary: Dynamic allocation can stall after an unfinished task is 
killed
                 Key: SPARK-59972
                 URL: https://issues.apache.org/jira/browse/SPARK-59972
             Project: Spark
          Issue Type: Bug
          Components: Spark Core
    Affects Versions: 4.3.0
            Reporter: MJ Deng


ExecutorAllocationManager does not count an unfinished regular task as pending 
again when its running attempt is explicitly killed.

A possible event sequence is:
 # A task starts on an available executor while the executor target is zero.
 # Starting the task drains the scheduler queue and clears the scheduler 
backlog timer.
 # SparkContext.killTaskAttempt kills the task before it succeeds.
 # The scheduler requeues the unfinished task.
 # ExecutorAllocationListener.onTaskEnd ignores TaskKilled, so it does not mark 
the task as pending or restart the backlog timer.
 # If the available executor is subsequently lost before the task relaunches, 
dynamic allocation does not request a replacement executor and the stage can 
stall.

TaskKilled cannot always be treated as pending because a regular attempt may be 
killed after a speculative copy has already succeeded, as covered by 
SPARK-30511.

The listener should therefore count a killed regular task as pending and 
restart the backlog timer only when:
 * the stage attempt is still active;
 * the killed attempt is non-speculative;
 * the task index has not already succeeded; and
 * the task index was known to be running.

Killed speculative attempts and task-end events from completed stages should 
remainĀ 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to