Gary Yao created FLINK-9190:
-------------------------------

             Summary: YarnResourceManager sometimes does not request new 
Containers
                 Key: FLINK-9190
                 URL: https://issues.apache.org/jira/browse/FLINK-9190
             Project: Flink
          Issue Type: Bug
          Components: Distributed Coordination, YARN
    Affects Versions: 1.5.0
         Environment: Hadoop 2.8.3
ZooKeeper 3.4.5
Flink 71c3cd2781d36e0a03d022a38cc4503d343f7ff8
            Reporter: Gary Yao
         Attachments: yarn-logs

*Description*
The {{YarnResourceManager}} does not request new containers if {{TaskManagers}} 
are killed rapidly in succession. After 5 minutes the job is restarted due to 
{{NoResourceAvailableException}}, and the job runs normally afterwards. I 
suspect that {{TaskManager}} failures are not registered if the failure occurs 
before the {{TaskManager}} registers with the master. Logs are attached; I 
added additional log statements to 
{{YarnResourceManager.onContainersCompleted}} and 
YarnResourceManager.onContainersAllocated}}.

*Expected Behavior*
The {{YarnResourceManager}} should recognize that the container is completed 
and keep requesting new containers. The job should run as soon as resources are 
available. 






--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Reply via email to