Gary Yao created FLINK-9190: ------------------------------- Summary: YarnResourceManager sometimes does not request new Containers Key: FLINK-9190 URL: https://issues.apache.org/jira/browse/FLINK-9190 Project: Flink Issue Type: Bug Components: Distributed Coordination, YARN Affects Versions: 1.5.0 Environment: Hadoop 2.8.3 ZooKeeper 3.4.5 Flink 71c3cd2781d36e0a03d022a38cc4503d343f7ff8 Reporter: Gary Yao Attachments: yarn-logs
*Description* The {{YarnResourceManager}} does not request new containers if {{TaskManagers}} are killed rapidly in succession. After 5 minutes the job is restarted due to {{NoResourceAvailableException}}, and the job runs normally afterwards. I suspect that {{TaskManager}} failures are not registered if the failure occurs before the {{TaskManager}} registers with the master. Logs are attached; I added additional log statements to {{YarnResourceManager.onContainersCompleted}} and YarnResourceManager.onContainersAllocated}}. *Expected Behavior* The {{YarnResourceManager}} should recognize that the container is completed and keep requesting new containers. The job should run as soon as resources are available. -- This message was sent by Atlassian JIRA (v7.6.3#76005)