[ https://issues.apache.org/jira/browse/SPARK-5721?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14317895#comment-14317895 ]
Lianhui Wang edited comment on SPARK-5721 at 2/12/15 9:39 AM: -------------------------------------------------------------- yes, i think it is same as SPARK-5759. now in https://github.com/apache/spark/pull/4554 i log the container ID and host name for exception. was (Author: lianhuiwang): yes, i think it is same as SPARK-5759. now in https://github.com/apache/spark/pull/4554 i report the container ID and host name for exception. > Propagate missing external shuffle service errors to client > ----------------------------------------------------------- > > Key: SPARK-5721 > URL: https://issues.apache.org/jira/browse/SPARK-5721 > Project: Spark > Issue Type: Bug > Components: Spark Core, YARN > Reporter: Kostas Sakellis > > When spark.shuffle.service.enabled=true, the yarn AM expects to find an aux > service running in the namenode. If it cannot find one an exception like this > is present in the app master logs. > {noformat} > Exception in thread "ContainerLauncher #0" Exception in thread > "ContainerLauncher #1" java.lang.Error: > org.apache.hadoop.yarn.exceptions.InvalidAuxServiceException: The > auxService:spark_shuffle does not exist > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1151) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: org.apache.hadoop.yarn.exceptions.InvalidAuxServiceException: The > auxService:spark_shuffle does not exist > at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method) > at > sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57) > at > sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45) > at java.lang.reflect.Constructor.newInstance(Constructor.java:526) > at > org.apache.hadoop.yarn.api.records.impl.pb.SerializedExceptionPBImpl.instantiateException(SerializedExceptionPBImpl.java:168) > at > org.apache.hadoop.yarn.api.records.impl.pb.SerializedExceptionPBImpl.deSerialize(SerializedExceptionPBImpl.java:106) > at > org.apache.hadoop.yarn.client.api.impl.NMClientImpl.startContainer(NMClientImpl.java:206) > at > org.apache.spark.deploy.yarn.ExecutorRunnable.startContainer(ExecutorRunnable.scala:110) > at > org.apache.spark.deploy.yarn.ExecutorRunnable.run(ExecutorRunnable.scala:65) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > ... 2 more > java.lang.Error: > org.apache.hadoop.yarn.exceptions.InvalidAuxServiceException: The > auxService:spark_shuffle does not exist > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1151) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: org.apache.hadoop.yarn.exceptions.InvalidAuxServiceException: The > auxService:spark_shuffle does not exist > at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method) > at > sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57) > at > sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45) > at java.lang.reflect.Constructor.newInstance(Constructor.java:526) > at > org.apache.hadoop.yarn.api.records.impl.pb.SerializedExceptionPBImpl.instantiateException(SerializedExceptionPBImpl.java:168) > at > org.apache.hadoop.yarn.api.records.impl.pb.SerializedExceptionPBImpl.deSerialize(SerializedExceptionPBImpl.java:106) > at > org.apache.hadoop.yarn.client.api.impl.NMClientImpl.startContainer(NMClientImpl.java:206) > at > org.apache.spark.deploy.yarn.ExecutorRunnable.startContainer(ExecutorRunnable.scala:110) > at > org.apache.spark.deploy.yarn.ExecutorRunnable.run(ExecutorRunnable.scala:65) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > ... 2 more > {noformat} > We should propagate this error to the driver (in yarn-client mode) because it > is otherwise unclear why the number of executors expected are not starting up. -- This message was sent by Atlassian JIRA (v6.3.4#6332) --------------------------------------------------------------------- To unsubscribe, e-mail: issues-unsubscr...@spark.apache.org For additional commands, e-mail: issues-h...@spark.apache.org