Karthik Kambatla updated YARN-4180:
    Attachment: YARN-4180.002.patch

Re-uploading the same patch to see if Jenkins kicks in. 

By the way, I ran the test locally and it passes. +1, even if Jenkins doesn't 
kick in. 

> AMLauncher does not retry on failures when talking to NM 
> ---------------------------------------------------------
>                 Key: YARN-4180
>                 URL: https://issues.apache.org/jira/browse/YARN-4180
>             Project: Hadoop YARN
>          Issue Type: Bug
>          Components: resourcemanager
>    Affects Versions: 2.7.1
>            Reporter: Anubhav Dhoot
>            Assignee: Anubhav Dhoot
>            Priority: Critical
>         Attachments: YARN-4180.001.patch, YARN-4180.002.patch, 
> YARN-4180.002.patch, YARN-4180.002.patch
> We see issues with RM trying to launch a container while a NM is restarting 
> and we get exceptions like NMNotReadyException. While YARN-3842 added retry 
> for other clients of NM (AMs mainly) its not used by AMLauncher in RM causing 
> there intermittent errors to cause job failures. This can manifest during 
> rolling restart of NMs. 

This message was sent by Atlassian JIRA

Reply via email to