[
https://issues.apache.org/jira/browse/OOZIE-1735?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Purshotam Shah updated OOZIE-1735:
----------------------------------
Description:
We should support resuming of failed coordinator job. Job are set to failed if
there are runtime error( like SQL timeout).
In current scenario there is no way to recover beside running SQL.
Resuming of failed coordinator job should also set pending to 1 ,reset
doneMaterialization and last modified to current time. So that materialization
continues.
We should also provide an option of resuming failed action. The behavior will
be same as killed option.
was:
We should support rerunning of failed job. Job are set to failed if there are
runtime error( like SQL timeout).
In current scenario there is no way to recover beside running SQL.
Rerun should set coord status to running and also set pending to 1 ,reset
doneMaterialization and last modified to current time. So that materialization
continues.
We should also provide an option of resuming failed action. The behavior will
be same as killed option.
> Support resuming of failed coordinator job and rerun of a failed coordinator
> action
> -----------------------------------------------------------------------------------
>
> Key: OOZIE-1735
> URL: https://issues.apache.org/jira/browse/OOZIE-1735
> Project: Oozie
> Issue Type: Bug
> Reporter: Purshotam Shah
> Assignee: Purshotam Shah
>
> We should support resuming of failed coordinator job. Job are set to failed
> if there are runtime error( like SQL timeout).
> In current scenario there is no way to recover beside running SQL.
> Resuming of failed coordinator job should also set pending to 1 ,reset
> doneMaterialization and last modified to current time. So that
> materialization continues.
> We should also provide an option of resuming failed action. The behavior will
> be same as killed option.
--
This message was sent by Atlassian JIRA
(v6.2#6252)