[jira] [Commented] (FLINK-9190) YarnResourceManager sometimes does not request new Containers

ASF GitHub Bot (JIRA) Mon, 30 Apr 2018 12:04:43 -0700

    [ 
https://issues.apache.org/jira/browse/FLINK-9190?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16458891#comment-16458891
 ]


ASF GitHub Bot commented on FLINK-9190:
---------------------------------------

Github user StephanEwen commented on the issue:

    https://github.com/apache/flink/pull/5931
  
     @GJL Briefly digging through the log, there are a few strange things 
happening:
    
      -  `YarnResourceManager` still has 8 pending requests even when 11 
containers are running:
    ```Received new container: container_1524853016208_0001_01_000184 - 
Remaining pending container requests: 8```
    
      - Some slots are requested and then the requests are cancelled again
      - In the end, one request is not fulfilled: 
`aeec2a9f010a187e04e31e6efd6f0f88`
    
    Might be an inconsistency in either in the `SlotManager` or `SlotPool`.


> YarnResourceManager sometimes does not request new Containers
> -------------------------------------------------------------
>
>                 Key: FLINK-9190
>                 URL: https://issues.apache.org/jira/browse/FLINK-9190
>             Project: Flink
>          Issue Type: Bug
>          Components: Distributed Coordination, YARN
>    Affects Versions: 1.5.0
>         Environment: Hadoop 2.8.3
> ZooKeeper 3.4.5
> Flink 71c3cd2781d36e0a03d022a38cc4503d343f7ff8
>            Reporter: Gary Yao
>            Assignee: Gary Yao
>            Priority: Blocker
>              Labels: flip-6
>             Fix For: 1.5.0
>
>         Attachments: yarn-logs
>
>
> *Description*
> The {{YarnResourceManager}} does not request new containers if 
> {{TaskManagers}} are killed rapidly in succession. After 5 minutes the job is 
> restarted due to {{NoResourceAvailableException}}, and the job runs normally 
> afterwards. I suspect that {{TaskManager}} failures are not registered if the 
> failure occurs before the {{TaskManager}} registers with the master. Logs are 
> attached; I added additional log statements to 
> {{YarnResourceManager.onContainersCompleted}} and 
> {{YarnResourceManager.onContainersAllocated}}.
> *Expected Behavior*
> The {{YarnResourceManager}} should recognize that the container is completed 
> and keep requesting new containers. The job should run as soon as resources 
> are available. 



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

[jira] [Commented] (FLINK-9190) YarnResourceManager sometimes does not request new Containers

Reply via email to