[jira] [Updated] (YARN-4152) NM crash with NPE when LogAggregationService#stopContainer called for absent container
[ https://issues.apache.org/jira/browse/YARN-4152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Rohith Sharma K S updated YARN-4152: Component/s: nodemanager > NM crash with NPE when LogAggregationService#stopContainer called for absent > container > -- > > Key: YARN-4152 > URL: https://issues.apache.org/jira/browse/YARN-4152 > Project: Hadoop YARN > Issue Type: Bug > Components: nodemanager >Reporter: Bibin A Chundatt >Assignee: Bibin A Chundatt >Priority: Critical > Fix For: 2.8.0 > > Attachments: 0001-YARN-4152.patch, 0002-YARN-4152.patch, > 0003-YARN-4152.patch > > > NM crash during of log aggregation. > Ran Pi job with 500 container and killed application in between > *Logs* > {code} > 2015-09-12 18:44:25,597 WARN > org.apache.hadoop.yarn.server.nodemanager.DefaultContainerExecutor: Exit code > from container container_e51_1442063466801_0001_01_99 is : 143 > 2015-09-12 18:44:25,670 WARN > org.apache.hadoop.yarn.server.nodemanager.containermanager.ContainerManagerImpl: > Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101 > 2015-09-12 18:44:25,670 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.application.ApplicationImpl: > Removing container_e51_1442063466801_0001_01_000101 from application > application_1442063466801_0001 > 2015-09-12 18:44:25,670 FATAL org.apache.hadoop.yarn.event.AsyncDispatcher: > Error in dispatcher thread > java.lang.NullPointerException > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.stopContainer(LogAggregationService.java:422) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:456) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:68) > at > org.apache.hadoop.yarn.event.AsyncDispatcher.dispatch(AsyncDispatcher.java:183) > at > org.apache.hadoop.yarn.event.AsyncDispatcher$1.run(AsyncDispatcher.java:109) > at java.lang.Thread.run(Thread.java:745) > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.AuxServices: Got > event CONTAINER_STOP for appId application_1442063466801_0001 > 2015-09-12 18:44:25,692 INFO org.apache.hadoop.yarn.event.AsyncDispatcher: > Exiting, bbye.. > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.NMAuditLogger: USER=dsperf > OPERATION=Container Finished - SucceededTARGET=ContainerImpl > RESULT=SUCCESS APPID=application_1442063466801_0001 > CONTAINERID=container_e51_1442063466801_0001_01_000100 > {code} > *Analysis* > Looks like for absent container also {{stopContainer}} is called > {code} > case CONTAINER_FINISHED: > LogHandlerContainerFinishedEvent containerFinishEvent = > (LogHandlerContainerFinishedEvent) event; > stopContainer(containerFinishEvent.getContainerId(), > containerFinishEvent.getExitCode()); > break; > {code} > *Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101* > Should skip when {{null==context.getContainers().get(containerId)}} -- This message was sent by Atlassian JIRA (v6.3.4#6332)
[jira] [Updated] (YARN-4152) NM crash with NPE when LogAggregationService#stopContainer called for absent container
[ https://issues.apache.org/jira/browse/YARN-4152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Rohith Sharma K S updated YARN-4152: Component/s: log-aggregation > NM crash with NPE when LogAggregationService#stopContainer called for absent > container > -- > > Key: YARN-4152 > URL: https://issues.apache.org/jira/browse/YARN-4152 > Project: Hadoop YARN > Issue Type: Bug > Components: log-aggregation, nodemanager >Reporter: Bibin A Chundatt >Assignee: Bibin A Chundatt >Priority: Critical > Fix For: 2.8.0 > > Attachments: 0001-YARN-4152.patch, 0002-YARN-4152.patch, > 0003-YARN-4152.patch > > > NM crash during of log aggregation. > Ran Pi job with 500 container and killed application in between > *Logs* > {code} > 2015-09-12 18:44:25,597 WARN > org.apache.hadoop.yarn.server.nodemanager.DefaultContainerExecutor: Exit code > from container container_e51_1442063466801_0001_01_99 is : 143 > 2015-09-12 18:44:25,670 WARN > org.apache.hadoop.yarn.server.nodemanager.containermanager.ContainerManagerImpl: > Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101 > 2015-09-12 18:44:25,670 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.application.ApplicationImpl: > Removing container_e51_1442063466801_0001_01_000101 from application > application_1442063466801_0001 > 2015-09-12 18:44:25,670 FATAL org.apache.hadoop.yarn.event.AsyncDispatcher: > Error in dispatcher thread > java.lang.NullPointerException > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.stopContainer(LogAggregationService.java:422) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:456) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:68) > at > org.apache.hadoop.yarn.event.AsyncDispatcher.dispatch(AsyncDispatcher.java:183) > at > org.apache.hadoop.yarn.event.AsyncDispatcher$1.run(AsyncDispatcher.java:109) > at java.lang.Thread.run(Thread.java:745) > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.AuxServices: Got > event CONTAINER_STOP for appId application_1442063466801_0001 > 2015-09-12 18:44:25,692 INFO org.apache.hadoop.yarn.event.AsyncDispatcher: > Exiting, bbye.. > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.NMAuditLogger: USER=dsperf > OPERATION=Container Finished - SucceededTARGET=ContainerImpl > RESULT=SUCCESS APPID=application_1442063466801_0001 > CONTAINERID=container_e51_1442063466801_0001_01_000100 > {code} > *Analysis* > Looks like for absent container also {{stopContainer}} is called > {code} > case CONTAINER_FINISHED: > LogHandlerContainerFinishedEvent containerFinishEvent = > (LogHandlerContainerFinishedEvent) event; > stopContainer(containerFinishEvent.getContainerId(), > containerFinishEvent.getExitCode()); > break; > {code} > *Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101* > Should skip when {{null==context.getContainers().get(containerId)}} -- This message was sent by Atlassian JIRA (v6.3.4#6332)
[jira] [Updated] (YARN-4152) NM crash with NPE when LogAggregationService#stopContainer called for absent container
[ https://issues.apache.org/jira/browse/YARN-4152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Bibin A Chundatt updated YARN-4152: --- Attachment: 0003-YARN-4152.patch Hi [~sunilg] Thnks for comments.Updated patch as per comments The operation performed was kill app and context container entry is not available. > NM crash with NPE when LogAggregationService#stopContainer called for absent > container > -- > > Key: YARN-4152 > URL: https://issues.apache.org/jira/browse/YARN-4152 > Project: Hadoop YARN > Issue Type: Bug >Reporter: Bibin A Chundatt >Assignee: Bibin A Chundatt >Priority: Critical > Attachments: 0001-YARN-4152.patch, 0002-YARN-4152.patch, > 0003-YARN-4152.patch > > > NM crash during of log aggregation. > Ran Pi job with 500 container and killed application in between > *Logs* > {code} > 2015-09-12 18:44:25,597 WARN > org.apache.hadoop.yarn.server.nodemanager.DefaultContainerExecutor: Exit code > from container container_e51_1442063466801_0001_01_99 is : 143 > 2015-09-12 18:44:25,670 WARN > org.apache.hadoop.yarn.server.nodemanager.containermanager.ContainerManagerImpl: > Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101 > 2015-09-12 18:44:25,670 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.application.ApplicationImpl: > Removing container_e51_1442063466801_0001_01_000101 from application > application_1442063466801_0001 > 2015-09-12 18:44:25,670 FATAL org.apache.hadoop.yarn.event.AsyncDispatcher: > Error in dispatcher thread > java.lang.NullPointerException > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.stopContainer(LogAggregationService.java:422) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:456) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:68) > at > org.apache.hadoop.yarn.event.AsyncDispatcher.dispatch(AsyncDispatcher.java:183) > at > org.apache.hadoop.yarn.event.AsyncDispatcher$1.run(AsyncDispatcher.java:109) > at java.lang.Thread.run(Thread.java:745) > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.AuxServices: Got > event CONTAINER_STOP for appId application_1442063466801_0001 > 2015-09-12 18:44:25,692 INFO org.apache.hadoop.yarn.event.AsyncDispatcher: > Exiting, bbye.. > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.NMAuditLogger: USER=dsperf > OPERATION=Container Finished - SucceededTARGET=ContainerImpl > RESULT=SUCCESS APPID=application_1442063466801_0001 > CONTAINERID=container_e51_1442063466801_0001_01_000100 > {code} > *Analysis* > Looks like for absent container also {{stopContainer}} is called > {code} > case CONTAINER_FINISHED: > LogHandlerContainerFinishedEvent containerFinishEvent = > (LogHandlerContainerFinishedEvent) event; > stopContainer(containerFinishEvent.getContainerId(), > containerFinishEvent.getExitCode()); > break; > {code} > *Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101* > Should skip when {{null==context.getContainers().get(containerId)}} -- This message was sent by Atlassian JIRA (v6.3.4#6332)
[jira] [Updated] (YARN-4152) NM crash with NPE when LogAggregationService#stopContainer called for absent container
[ https://issues.apache.org/jira/browse/YARN-4152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Bibin A Chundatt updated YARN-4152: --- Attachment: 0002-YARN-4152.patch Hi [~rohithsharma] Thnks for looking into the issue. Updated patch as per comments > NM crash with NPE when LogAggregationService#stopContainer called for absent > container > -- > > Key: YARN-4152 > URL: https://issues.apache.org/jira/browse/YARN-4152 > Project: Hadoop YARN > Issue Type: Bug >Reporter: Bibin A Chundatt >Assignee: Bibin A Chundatt >Priority: Critical > Attachments: 0001-YARN-4152.patch, 0002-YARN-4152.patch > > > NM crash during of log aggregation. > Ran Pi job with 500 container and killed application in between > *Logs* > {code} > 2015-09-12 18:44:25,597 WARN > org.apache.hadoop.yarn.server.nodemanager.DefaultContainerExecutor: Exit code > from container container_e51_1442063466801_0001_01_99 is : 143 > 2015-09-12 18:44:25,670 WARN > org.apache.hadoop.yarn.server.nodemanager.containermanager.ContainerManagerImpl: > Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101 > 2015-09-12 18:44:25,670 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.application.ApplicationImpl: > Removing container_e51_1442063466801_0001_01_000101 from application > application_1442063466801_0001 > 2015-09-12 18:44:25,670 FATAL org.apache.hadoop.yarn.event.AsyncDispatcher: > Error in dispatcher thread > java.lang.NullPointerException > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.stopContainer(LogAggregationService.java:422) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:456) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:68) > at > org.apache.hadoop.yarn.event.AsyncDispatcher.dispatch(AsyncDispatcher.java:183) > at > org.apache.hadoop.yarn.event.AsyncDispatcher$1.run(AsyncDispatcher.java:109) > at java.lang.Thread.run(Thread.java:745) > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.AuxServices: Got > event CONTAINER_STOP for appId application_1442063466801_0001 > 2015-09-12 18:44:25,692 INFO org.apache.hadoop.yarn.event.AsyncDispatcher: > Exiting, bbye.. > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.NMAuditLogger: USER=dsperf > OPERATION=Container Finished - SucceededTARGET=ContainerImpl > RESULT=SUCCESS APPID=application_1442063466801_0001 > CONTAINERID=container_e51_1442063466801_0001_01_000100 > {code} > *Analysis* > Looks like for absent container also {{stopContainer}} is called > {code} > case CONTAINER_FINISHED: > LogHandlerContainerFinishedEvent containerFinishEvent = > (LogHandlerContainerFinishedEvent) event; > stopContainer(containerFinishEvent.getContainerId(), > containerFinishEvent.getExitCode()); > break; > {code} > *Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101* > Should skip when {{null==context.getContainers().get(containerId)}} -- This message was sent by Atlassian JIRA (v6.3.4#6332)
[jira] [Updated] (YARN-4152) NM crash with NPE when LogAggregationService#stopContainer called for absent container
[ https://issues.apache.org/jira/browse/YARN-4152?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ] Bibin A Chundatt updated YARN-4152: --- Summary: NM crash with NPE when LogAggregationService#stopContainer called for absent container (was: NM crash when LogAggregationService#stopContainer called for absent container) > NM crash with NPE when LogAggregationService#stopContainer called for absent > container > -- > > Key: YARN-4152 > URL: https://issues.apache.org/jira/browse/YARN-4152 > Project: Hadoop YARN > Issue Type: Bug >Reporter: Bibin A Chundatt >Assignee: Bibin A Chundatt >Priority: Critical > Attachments: 0001-YARN-4152.patch > > > NM crash during of log aggregation. > Ran Pi job with 500 container and killed application in between > *Logs* > {code} > 2015-09-12 18:44:25,597 WARN > org.apache.hadoop.yarn.server.nodemanager.DefaultContainerExecutor: Exit code > from container container_e51_1442063466801_0001_01_99 is : 143 > 2015-09-12 18:44:25,670 WARN > org.apache.hadoop.yarn.server.nodemanager.containermanager.ContainerManagerImpl: > Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101 > 2015-09-12 18:44:25,670 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.application.ApplicationImpl: > Removing container_e51_1442063466801_0001_01_000101 from application > application_1442063466801_0001 > 2015-09-12 18:44:25,670 FATAL org.apache.hadoop.yarn.event.AsyncDispatcher: > Error in dispatcher thread > java.lang.NullPointerException > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.stopContainer(LogAggregationService.java:422) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:456) > at > org.apache.hadoop.yarn.server.nodemanager.containermanager.logaggregation.LogAggregationService.handle(LogAggregationService.java:68) > at > org.apache.hadoop.yarn.event.AsyncDispatcher.dispatch(AsyncDispatcher.java:183) > at > org.apache.hadoop.yarn.event.AsyncDispatcher$1.run(AsyncDispatcher.java:109) > at java.lang.Thread.run(Thread.java:745) > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.containermanager.AuxServices: Got > event CONTAINER_STOP for appId application_1442063466801_0001 > 2015-09-12 18:44:25,692 INFO org.apache.hadoop.yarn.event.AsyncDispatcher: > Exiting, bbye.. > 2015-09-12 18:44:25,692 INFO > org.apache.hadoop.yarn.server.nodemanager.NMAuditLogger: USER=dsperf > OPERATION=Container Finished - SucceededTARGET=ContainerImpl > RESULT=SUCCESS APPID=application_1442063466801_0001 > CONTAINERID=container_e51_1442063466801_0001_01_000100 > {code} > *Analysis* > Looks like for absent container also {{stopContainer}} is called > {code} > case CONTAINER_FINISHED: > LogHandlerContainerFinishedEvent containerFinishEvent = > (LogHandlerContainerFinishedEvent) event; > stopContainer(containerFinishEvent.getContainerId(), > containerFinishEvent.getExitCode()); > break; > {code} > *Event EventType: KILL_CONTAINER sent to absent container > container_e51_1442063466801_0001_01_000101* > Should skip when {{null==context.getContainers().get(containerId)}} -- This message was sent by Atlassian JIRA (v6.3.4#6332)