yanbin.zhang created SPARK-59813:
------------------------------------

             Summary: Dump thread stacks when SparkContext initialization times 
out in ApplicationMaster
                 Key: SPARK-59813
                 URL: https://issues.apache.org/jira/browse/SPARK-59813
             Project: Spark
          Issue Type: Improvement
          Components: YARN
    Affects Versions: 3.3.2
            Reporter: yanbin.zhang


When SparkContext initialization times out in YARN cluster mode 
(spark.yarn.am.waitTime),
the ApplicationMaster fails the application with exit code 13 and only advises 
to
"check earlier log output for errors". However, in real production incidents 
the driver
thread is often blocked silently (e.g., a slow/contended node, or a blocking 
call that
produces no logging) and there is no earlier log output at all, leaving 
operators with
no way to tell where the driver thread was stuck.

Production example: a Kyuubi Spark engine failed with
"java.util.concurrent.TimeoutException: Futures timed out after [100000 
milliseconds]"
(exit code 13). The driver stdout showed ~80 seconds of complete silence right 
after
SparkEnv initialization finished. Cross-checking HDFS NameNode and Hive 
Metastore logs
proved the driver issued zero external I/O during the hang, but without a 
thread dump
there was no way to pinpoint the blocking frame.

Proposal: when the timeout fires, log (1) the stack trace of the user 
application
thread (driver), and (2) a full thread dump, before failing the application. 
The dump
is best-effort so it never masks the original timeout error. This follows the 
same
pattern as the existing spark.task.reaper.threadDump on the executor side.

PR: https://github.com/apache/spark/pull/58991



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to