MJ Deng created SPARK-59972:
-------------------------------
Summary: Dynamic allocation can stall after an unfinished task is
killed
Key: SPARK-59972
URL: https://issues.apache.org/jira/browse/SPARK-59972
Project: Spark
Issue Type: Bug
Components: Spark Core
Affects Versions: 4.3.0
Reporter: MJ Deng
ExecutorAllocationManager does not count an unfinished regular task as pending
again when its running attempt is explicitly killed.
A possible event sequence is:
# A task starts on an available executor while the executor target is zero.
# Starting the task drains the scheduler queue and clears the scheduler
backlog timer.
# SparkContext.killTaskAttempt kills the task before it succeeds.
# The scheduler requeues the unfinished task.
# ExecutorAllocationListener.onTaskEnd ignores TaskKilled, so it does not mark
the task as pending or restart the backlog timer.
# If the available executor is subsequently lost before the task relaunches,
dynamic allocation does not request a replacement executor and the stage can
stall.
TaskKilled cannot always be treated as pending because a regular attempt may be
killed after a speculative copy has already succeeded, as covered by
SPARK-30511.
The listener should therefore count a killed regular task as pending and
restart the backlog timer only when:
* the stage attempt is still active;
* the killed attempt is non-speculative;
* the task index has not already succeeded; and
* the task index was known to be running.
Killed speculative attempts and task-end events from completed stages should
remainĀ
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]