[
https://issues.apache.org/jira/browse/SPARK-58924?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Dongjoon Hyun updated SPARK-58924:
----------------------------------
Description:
This issue aims to allow a running application to temporarily give back its
resources without being killed. Executors are drained gracefully with no loss
of in-progress task work or shuffle output, and the application can be resumed
later without recomputing the shuffle output already written. Cached blocks are
not preserved and are recomputed after resuming.
This takes effect only when *spark.decommission.enabled* is true, the shuffle
data is kept outside the executors (through either
*spark.shuffle.service.enabled* or a
ShuffleDataIO with reliable storage), and the cluster manager can hold
executors: Standalone, YARN, and Kubernetes (with the default "direct"
allocator).
!Diagram.png!
was:
This issue aims to allow a running application to temporarily give back its
resources without being killed. Executors are drained gracefully with no loss
of in-progress task work or shuffle output, and the application can be resumed
later without recomputing the shuffle output already written. Cached blocks are
not preserved and are recomputed after resuming.
!Diagram.png!
> Graceful Application Suspend and Resume Framework
> -------------------------------------------------
>
> Key: SPARK-58924
> URL: https://issues.apache.org/jira/browse/SPARK-58924
> Project: Spark
> Issue Type: Umbrella
> Components: Spark Core
> Affects Versions: 4.4.0
> Reporter: Dongjoon Hyun
> Assignee: Dongjoon Hyun
> Priority: Critical
> Fix For: 4.4.0
>
> Attachments: Diagram.png
>
>
> This issue aims to allow a running application to temporarily give back its
> resources without being killed. Executors are drained gracefully with no loss
> of in-progress task work or shuffle output, and the application can be
> resumed later without recomputing the shuffle output already written. Cached
> blocks are not preserved and are recomputed after resuming.
> This takes effect only when *spark.decommission.enabled* is true, the shuffle
> data is kept outside the executors (through either
> *spark.shuffle.service.enabled* or a
> ShuffleDataIO with reliable storage), and the cluster manager can hold
> executors: Standalone, YARN, and Kubernetes (with the default "direct"
> allocator).
> !Diagram.png!
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]