[
https://issues.apache.org/jira/browse/YARN-1618?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13883659#comment-13883659
]
Karthik Kambatla commented on YARN-1618:
----------------------------------------
bq. I am not sure this would happen in real life since only a START event would
trigger going to the scheduler and introduce the possibility of a REJECTED
event.
Actually, looking at all the places an APP_REJECTED event is called, found that
YARN-674 triggers APP_REJECTED on NEW. We could either change this to KILL, or
update our comments in RMAppEventType to reflect that APP_REJECTED could come
from places other than the scheduler. [~bikassaha] - thoughts?
> Applications transition from NEW to FINAL_SAVING, and try to update
> non-existing entries in the state-store
> -----------------------------------------------------------------------------------------------------------
>
> Key: YARN-1618
> URL: https://issues.apache.org/jira/browse/YARN-1618
> Project: Hadoop YARN
> Issue Type: Sub-task
> Components: resourcemanager
> Affects Versions: 2.2.0
> Reporter: Karthik Kambatla
> Assignee: Karthik Kambatla
> Priority: Blocker
> Attachments: yarn-1618-1.patch, yarn-1618-2.patch
>
>
> YARN-891 augments the RMStateStore to store information on completed
> applications. In the process, it adds transitions from NEW to FINAL_SAVING.
> This leads to the RM trying to update entries in the state-store that do not
> exist. On ZKRMStateStore, this leads to the RM crashing.
> Previous description:
> ZKRMStateStore fails to handle updates to znodes that don't exist. For
> instance, this can happen when an app transitions from NEW to FINAL_SAVING.
> In these cases, the store should create the missing znode and handle the
> update.
--
This message was sent by Atlassian JIRA
(v6.1.5#6160)