[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14212378#comment-14212378 ] Hudson commented on YARN-2846: -- SUCCESS: Integrated in Hadoop-Mapreduce-trunk-Java8 #5 (See [https://builds.apache.org/job/Hadoop-Mapreduce-trunk-Java8/5/]) YARN-2846. Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart. Contributed by Junping Du (jlowe: rev 33ea5ae92b9dd3abace104903d9a94d17dd75af5) * hadoop-yarn-project/CHANGES.txt * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/containermanager/launcher/RecoveredContainerLaunch.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/LinuxContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/ContainerExecutor.java > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Fix For: 2.6.0 > > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.y
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14212353#comment-14212353 ] Hudson commented on YARN-2846: -- FAILURE: Integrated in Hadoop-Mapreduce-trunk #1957 (See [https://builds.apache.org/job/Hadoop-Mapreduce-trunk/1957/]) YARN-2846. Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart. Contributed by Junping Du (jlowe: rev 33ea5ae92b9dd3abace104903d9a94d17dd75af5) * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/ContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/containermanager/launcher/RecoveredContainerLaunch.java * hadoop-yarn-project/CHANGES.txt * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/LinuxContainerExecutor.java > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Fix For: 2.6.0 > > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.se
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14212290#comment-14212290 ] Hudson commented on YARN-2846: -- SUCCESS: Integrated in Hadoop-Hdfs-trunk-Java8 #5 (See [https://builds.apache.org/job/Hadoop-Hdfs-trunk-Java8/5/]) YARN-2846. Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart. Contributed by Junping Du (jlowe: rev 33ea5ae92b9dd3abace104903d9a94d17dd75af5) * hadoop-yarn-project/CHANGES.txt * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/ContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/containermanager/launcher/RecoveredContainerLaunch.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/LinuxContainerExecutor.java > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Fix For: 2.6.0 > > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14212277#comment-14212277 ] Hudson commented on YARN-2846: -- SUCCESS: Integrated in Hadoop-Hdfs-trunk #1933 (See [https://builds.apache.org/job/Hadoop-Hdfs-trunk/1933/]) YARN-2846. Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart. Contributed by Junping Du (jlowe: rev 33ea5ae92b9dd3abace104903d9a94d17dd75af5) * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/LinuxContainerExecutor.java * hadoop-yarn-project/CHANGES.txt * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/ContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/containermanager/launcher/RecoveredContainerLaunch.java > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Fix For: 2.6.0 > > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server.nodem
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14212185#comment-14212185 ] Hudson commented on YARN-2846: -- SUCCESS: Integrated in Hadoop-Yarn-trunk #743 (See [https://builds.apache.org/job/Hadoop-Yarn-trunk/743/]) YARN-2846. Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart. Contributed by Junping Du (jlowe: rev 33ea5ae92b9dd3abace104903d9a94d17dd75af5) * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/ContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/LinuxContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/containermanager/launcher/RecoveredContainerLaunch.java * hadoop-yarn-project/CHANGES.txt > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Fix For: 2.6.0 > > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server.nodeman
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14212156#comment-14212156 ] Hudson commented on YARN-2846: -- FAILURE: Integrated in Hadoop-Yarn-trunk-Java8 #5 (See [https://builds.apache.org/job/Hadoop-Yarn-trunk-Java8/5/]) YARN-2846. Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart. Contributed by Junping Du (jlowe: rev 33ea5ae92b9dd3abace104903d9a94d17dd75af5) * hadoop-yarn-project/CHANGES.txt * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/containermanager/launcher/RecoveredContainerLaunch.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/ContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/LinuxContainerExecutor.java > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Fix For: 2.6.0 > > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14209974#comment-14209974 ] Hudson commented on YARN-2846: -- SUCCESS: Integrated in Hadoop-trunk-Commit #6534 (See [https://builds.apache.org/job/Hadoop-trunk-Commit/6534/]) YARN-2846. Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart. Contributed by Junping Du (jlowe: rev 33ea5ae92b9dd3abace104903d9a94d17dd75af5) * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/containermanager/launcher/RecoveredContainerLaunch.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/ContainerExecutor.java * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager/src/main/java/org/apache/hadoop/yarn/server/nodemanager/LinuxContainerExecutor.java * hadoop-yarn-project/CHANGES.txt > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Fix For: 2.6.0 > > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server.n
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14209948#comment-14209948 ] Jason Lowe commented on YARN-2846: -- Agreed. Committing this. > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:177) > ... 6 more > {code} > In reacquireContainer() of ContainerExecutor.java, the while loop of checking > container process (AM container) will be interrupted by NM stop. The > IOException get thrown and failed to generate an ExitCodeFile for the running > container. Later, the IOException will be caught in upper call > (RecoveredContainerLaunch.call()) and the ExitCode (by default to be LOST > without any setting) get persistent in NMStateStore. > After NM restart again, this container is recovered as COMPLETE state but > exit code is LOST (154) - cause this (AM) container get killed later. > We should get rid of recording the exit code of running containers if > detecting process is interrupted. -- This message was sent by Atlassian JIRA (v6.3.4#6332)
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14209920#comment-14209920 ] Junping Du commented on YARN-2846: -- bq. If we're going to kill normal containers on shutdown then why wouldn't we also kill containers we are recovering as well? For the NM restart scenario we're not supposed to be killing any containers. My bad. Sorry for my expression above which isn't right. Yes. We shouldn't kill containers for normal containers (fresh) and recovered containers (survival from NM restart before). bq. it's essentially a question of why doesn't interrupting the ContainerLaunch thread manifest as a container completing as it did for a recovered container. Agree this is important. For ContainerLaunch (take DefaultContainerExecutor as an example), I think thread are blocking in launchContainer() {code} if (isContainerActive(containerId)) { shExec.execute(); } {code} The shExec.execute() will call Shell.runCommand() with building a new process for the command (with an error monitoring thread). The thread will be waiting at : {code} // wait for the process to finish and check the exit code exitCode = process.waitFor(); {code} It is also possible for InterruptedException get thrown there but the trigger event is not the kill of NM but kill of shell process (so not affected by NM kill). That may be the root cause for the different behavior now for fresh container and recovered container. This is not my final conclusion, but I would prefer to fix the existing significant bug (block container recovery for recovered containers) here and we can do more investigation later. [~jlowe], what do you think? > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLau
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14208087#comment-14208087 ] Jason Lowe commented on YARN-2846: -- Thanks, Junping, patch looks better. I'm +1 pending investigation of the ContainerLaunch path and why we don't have to deal with thread interruption there. bq. But if regular ContainerLaunch get interrupted, we may not care if running container exit code as these running container should be killed soon If we're going to kill normal containers on shutdown then why wouldn't we also kill containers we are recovering as well? For the NM restart scenario we're not supposed to be killing any containers, so it's essentially a question of why doesn't interrupting the ContainerLaunch thread manifest as a container completing as it did for a recovered container. If we know why that's not possible then we can put in the patch as-is, otherwise I'm wondering if there's another hole we need to plug. > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:177) > ... 6 more > {
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14207659#comment-14207659 ] Hadoop QA commented on YARN-2846: - {color:red}-1 overall{color}. Here are the results of testing the latest attachment http://issues.apache.org/jira/secure/attachment/12680993/YARN-2846.patch against trunk revision 46f6f9d. {color:green}+1 @author{color}. The patch does not contain any @author tags. {color:red}-1 tests included{color}. The patch doesn't appear to include any new or modified tests. Please justify why no new tests are needed for this patch. Also please list what manual steps were performed to verify this patch. {color:green}+1 javac{color}. The applied patch does not increase the total number of javac compiler warnings. {color:green}+1 javadoc{color}. There were no new javadoc warning messages. {color:green}+1 eclipse:eclipse{color}. The patch built with eclipse:eclipse. {color:green}+1 findbugs{color}. The patch does not introduce any new Findbugs (version 2.0.3) warnings. {color:green}+1 release audit{color}. The applied patch does not increase the total number of release audit warnings. {color:green}+1 core tests{color}. The patch passed unit tests in hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager. {color:green}+1 contrib tests{color}. The patch passed contrib unit tests. Test results: https://builds.apache.org/job/PreCommit-YARN-Build/5822//testReport/ Console output: https://builds.apache.org/job/PreCommit-YARN-Build/5822//console This message is automatically generated. > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82)
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14207639#comment-14207639 ] Junping Du commented on YARN-2846: -- Thanks [~jlowe] for review and comments. The latest patch addressed your comments. bq. I'm curious why we're not seeing a similar issue with regular ContainerLaunch threads, as they should be interrupted as well. Are those threads silently swallowing the interrupt? Because otherwise I would expect us to log a container completion just like we were doing with a recovered container. I am not sure on this. But if regular ContainerLaunch get interrupted, we may not care if running container exit code as these running container should be killed soon (because NM daemon stop). Am I missing anything here? > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Attachments: YARN-2846-demo.patch, YARN-2846.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:177) > ... 6 more > {code} > In reacquireContainer() of ContainerExecutor.java, the while loop of checking > container process (AM container) will be interrupted by NM stop. The > IOException get thrown and failed to
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14206616#comment-14206616 ] Jason Lowe commented on YARN-2846: -- Thanks for the report and patch, Junping! Nit: If reacquireContainer is going to allow InterruptedException to be thrown then I'd rather remove the try/catch around the Thread.sleep call and just let the exception be thrown directly from there. We can let the code catching the exception deal with any logging/etc as appropriate for that caller. In this case we can move the log message to RecoveredContainerLaunch when it fields the InterruptedException and chooses not to propagate it upwards. I'm curious why we're not seeing a similar issue with regular ContainerLaunch threads, as they should be interrupted as well. Are those threads silently swallowing the interrupt? Because otherwise I would expect us to log a container completion just like we were doing with a recovered container. > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Attachments: YARN-2846-demo.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:46) > at java.util.concurrent.FutureTask.run(FutureTask.java:262) > at > java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) > at > java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) > at java.lang.Thread.run(Thread.java:745) > Caused by: java.lang.InterruptedException: sleep interrupted > at java.lang.Thread.sleep(Native Method) > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:177) > ... 6 more > {code} > In reacquireCo
[jira] [Commented] (YARN-2846) Incorrect persist exit code for running containers in reacquireContainer() that interrupted by NodeManager restart.
[ https://issues.apache.org/jira/browse/YARN-2846?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14206601#comment-14206601 ] Hadoop QA commented on YARN-2846: - {color:red}-1 overall{color}. Here are the results of testing the latest attachment http://issues.apache.org/jira/secure/attachment/12680801/YARN-2846-demo.patch against trunk revision 58e9bf4. {color:green}+1 @author{color}. The patch does not contain any @author tags. {color:red}-1 tests included{color}. The patch doesn't appear to include any new or modified tests. Please justify why no new tests are needed for this patch. Also please list what manual steps were performed to verify this patch. {color:green}+1 javac{color}. The applied patch does not increase the total number of javac compiler warnings. {color:green}+1 javadoc{color}. There were no new javadoc warning messages. {color:green}+1 eclipse:eclipse{color}. The patch built with eclipse:eclipse. {color:green}+1 findbugs{color}. The patch does not introduce any new Findbugs (version 2.0.3) warnings. {color:green}+1 release audit{color}. The applied patch does not increase the total number of release audit warnings. {color:green}+1 core tests{color}. The patch passed unit tests in hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-nodemanager. {color:green}+1 contrib tests{color}. The patch passed contrib unit tests. Test results: https://builds.apache.org/job/PreCommit-YARN-Build/5814//testReport/ Console output: https://builds.apache.org/job/PreCommit-YARN-Build/5814//console This message is automatically generated. > Incorrect persist exit code for running containers in reacquireContainer() > that interrupted by NodeManager restart. > --- > > Key: YARN-2846 > URL: https://issues.apache.org/jira/browse/YARN-2846 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Junping Du >Assignee: Junping Du >Priority: Blocker > Attachments: YARN-2846-demo.patch > > > The NM restart work preserving feature could make running AM container get > LOST and killed during stop NM daemon. The exception is like below: > {code} > 2014-11-11 00:48:35,214 INFO monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(408)) - Memory usage of ProcessTree 22140 for > container-id container_1415666714233_0001_01_84: 53.8 MB of 512 MB > physical memory used; 931.3 MB of 1.0 GB virtual memory used > 2014-11-11 00:48:35,223 ERROR nodemanager.NodeManager > (SignalLogger.java:handle(60)) - RECEIVED SIGNAL 15: SIGTERM > 2014-11-11 00:48:35,299 INFO mortbay.log (Slf4jLog.java:info(67)) - Stopped > HttpServer2$SelectChannelConnectorWithSafeStartup@0.0.0.0:50060 > 2014-11-11 00:48:35,337 INFO containermanager.ContainerManagerImpl > (ContainerManagerImpl.java:cleanUpApplicationsOnNMShutDown(512)) - > Applications still running : [application_1415666714233_0001] > 2014-11-11 00:48:35,338 INFO ipc.Server (Server.java:stop(2437)) - Stopping > server on 45454 > 2014-11-11 00:48:35,344 INFO ipc.Server (Server.java:run(706)) - Stopping > IPC Server listener on 45454 > 2014-11-11 00:48:35,346 INFO logaggregation.LogAggregationService > (LogAggregationService.java:serviceStop(141)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService > waiting for pending aggregation during exit > 2014-11-11 00:48:35,347 INFO ipc.Server (Server.java:run(832)) - Stopping > IPC Server Responder > 2014-11-11 00:48:35,347 INFO logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:abortLogAggregation(502)) - Aborting log > aggregation for application_1415666714233_0001 > 2014-11-11 00:48:35,348 WARN logaggregation.AppLogAggregatorImpl > (AppLogAggregatorImpl.java:run(382)) - Aggregation did not complete for > application application_1415666714233_0001 > 2014-11-11 00:48:35,358 WARN monitor.ContainersMonitorImpl > (ContainersMonitorImpl.java:run(476)) - > org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl > is interrupted. Exiting. > 2014-11-11 00:48:35,406 ERROR launcher.RecoveredContainerLaunch > (RecoveredContainerLaunch.java:call(87)) - Unable to recover container > container_1415666714233_0001_01_01 > java.io.IOException: Interrupted while waiting for process 20001 to exit > at > org.apache.hadoop.yarn.server.nodemanager.ContainerExecutor.reacquireContainer(ContainerExecutor.java:180) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.launcher.RecoveredContainerLaunch.call(RecoveredContainerLaunch.java:82) > at