[
https://issues.apache.org/jira/browse/YARN-8193?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17096807#comment-17096807
]
Jonathan Hung edited comment on YARN-8193 at 4/30/20, 5:32 PM:
---------------------------------------------------------------
Hit this issue on 2.10.0 cluster. Reuploading patch to trigger jenkins
was (Author: jhung):
Hit this issue on 2.10.0 cluster. Reuploading patch
> YARN RM hangs abruptly (stops allocating resources) when running successive
> applications.
> -----------------------------------------------------------------------------------------
>
> Key: YARN-8193
> URL: https://issues.apache.org/jira/browse/YARN-8193
> Project: Hadoop YARN
> Issue Type: Bug
> Components: yarn
> Reporter: Zian Chen
> Assignee: Zian Chen
> Priority: Blocker
> Fix For: 3.2.0, 3.1.1
>
> Attachments: YARN-8193-branch-2-001.patch,
> YARN-8193-branch-2.10-001.patch, YARN-8193-branch-2.9.0-001.patch,
> YARN-8193.001.patch, YARN-8193.002.patch
>
>
> When running massive queries successively, at some point RM just hangs and
> stops allocating resources. At the point RM get hangs, YARN throw
> NullPointerException at RegularContainerAllocator.getLocalityWaitFactor.
> There's sufficient space given to yarn.nodemanager.local-dirs (not a node
> health issue, RM didn't report any node being unhealthy). There is no fixed
> trigger for this (query or operation).
> This problem goes away on restarting ResourceManager. No NM restart is
> required.
>
>
--
This message was sent by Atlassian Jira
(v8.3.4#803005)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]