[ 
https://issues.apache.org/jira/browse/FLINK-10868?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16920778#comment-16920778
 ] 

huanyang commented on FLINK-10868:
----------------------------------

Hi Peter&Till:


We have introduced the FLINK-10868 patch (mainly batch tasks) online, we found 
that JM occasionally lost contact during use, and the use of [multi-threaded 
start Container|https://issues.apache.org/jira/browse/FLINK-13184]  is 
mitigated.


There are therefore two suggestions: 
1. Parameter control time interval. At present, the default time interval of 1 
min is used, which is too short for batch tasks; 
2. Parameter Control When the failed Container number reaches 
MAXIMUM_WORKERS_FAILURE_RATE and JM disconnects whether to perform OnFatalError 
so that the batch tasks can exit as soon as possible.

> Flink's JobCluster ResourceManager doesn't use maximum-failed-containers as 
> limit of resource acquirement
> ---------------------------------------------------------------------------------------------------------
>
>                 Key: FLINK-10868
>                 URL: https://issues.apache.org/jira/browse/FLINK-10868
>             Project: Flink
>          Issue Type: Bug
>          Components: Deployment / Mesos, Deployment / YARN
>    Affects Versions: 1.6.2, 1.7.0
>            Reporter: Zhenqiu Huang
>            Assignee: Zhenqiu Huang
>            Priority: Major
>              Labels: pull-request-available
>          Time Spent: 0.5h
>  Remaining Estimate: 0h
>
> Currently, YarnResourceManager does use yarn.maximum-failed-containers as 
> limit of resource acquirement. In worse case, when new start containers 
> consistently fail, YarnResourceManager will goes into an infinite resource 
> acquirement process without failing the job. Together with the 
> https://issues.apache.org/jira/browse/FLINK-10848, It will quick occupy all 
> resources of yarn queue.



--
This message was sent by Atlassian Jira
(v8.3.2#803003)

Reply via email to