[ https://issues.apache.org/jira/browse/YARN-2456?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14132704#comment-14132704 ]
Hudson commented on YARN-2456: ------------------------------ FAILURE: Integrated in Hadoop-Mapreduce-trunk #1895 (See [https://builds.apache.org/job/Hadoop-Mapreduce-trunk/1895/]) YARN-2456. Possible livelock in CapacityScheduler when RM is recovering (xgong: rev e65ae575a059a426c4c38fdabe22a31eabbb349e) * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-resourcemanager/src/main/java/org/apache/hadoop/yarn/server/resourcemanager/recovery/RMStateStore.java * hadoop-yarn-project/CHANGES.txt * hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-resourcemanager/src/test/java/org/apache/hadoop/yarn/server/resourcemanager/TestRMRestart.java > Possible livelock in CapacityScheduler when RM is recovering apps > ----------------------------------------------------------------- > > Key: YARN-2456 > URL: https://issues.apache.org/jira/browse/YARN-2456 > Project: Hadoop YARN > Issue Type: Sub-task > Components: resourcemanager > Reporter: Jian He > Assignee: Jian He > Fix For: 2.6.0 > > Attachments: YARN-2456.1.patch, YARN-2456.2.patch > > > Consider this scenario: > 1. RM is configured with a single queue and only one application can be > active at a time. > 2. Submit App1 which uses up the queue's whole capacity > 3. Submit App2 which remains pending. > 4. Restart RM. > 5. App2 is recovered before App1, so App2 is added to the activeApplications > list. Now App1 remains pending (because of max-active-app limit) > 6. All containers of App1 are now recovered when NM registers, and use up the > whole queue capacity again. > 7. Since the queue is full, App2 cannot proceed to allocate AM container. > 8. In the meanwhile, App1 cannot proceed to become active because of the > max-active-app limit -- This message was sent by Atlassian JIRA (v6.3.4#6332)