amoghrajesh commented on code in PR #72100:
URL: https://github.com/apache/airflow/pull/72100#discussion_r4122610287


##########
airflow-core/docs/core-concepts/task-state-store.rst:
##########
@@ -285,7 +285,7 @@ If the worker process crashes, the task instance is 
retried. Task store data wri
 Deferrable tasks
 ~~~~~~~~~~~~~~~~
 
-Once a task defers, the Triggerer handles continuity across poke cycles. Use 
task state store in deferrable tasks only when you need to survive an 
operator-initiated clear, not for normal poke continuity.
+Once a task defers, the Triggerer handles continuity across poke cycles. 
Clearing a deferred task does not synchronously cancel its trigger: the 
Triggerer only notices the trigger is orphaned on its next iteration, then 
cancels it via the trigger's ``on_kill``, bounded by ``[triggerer] 
on_kill_timeout``. A new attempt can therefore start before that cancellation 
finishes. Most triggers implement ``on_kill`` to cancel the external job there, 
so the next attempt usually finds nothing left to reconnect to, but this is not 
guaranteed. The state store still matters for the small set of triggers that 
don't implement ``on_kill`` (for example ``GlueJobCompleteTrigger`` and 
``LivyTrigger``): for those, keep task state (``keep_task_state``) when 
clearing so the next attempt reconnects to the job still running instead of 
submitting a duplicate.

Review Comment:
   Fixed. Removed the wrong `keep_task_state` advice for 
`GlueJobCompleteTrigger`/`LivyTrigger` in both files, replaced with the actual 
fix (durable=True for Glue, cancel the batch manually for Livy).



##########
airflow-core/newsfragments/72100.significant.rst:
##########
@@ -0,0 +1,59 @@
+Clearing a task now discards its task state store entries by default
+
+Clearing a task instance discards its ``task_state_store`` entries, so the 
next attempt starts from
+the beginning instead of resuming from a checkpoint or reconnecting to an 
external job recorded by
+the attempt that was cleared.
+
+Retries are unaffected. They keep task state exactly as before, which is what 
crash recovery relies
+on. Only a deliberate clear discards.
+
+This only applies to clearing individual task instances (the task-instance 
clear endpoint / dialog,
+and ``airflowctl dags clear``, which clears every task instance in the matched 
Dag run(s) through
+the same endpoint). Clearing an entire Dag run through the "Clear Run" 
dialog/API, and marking a
+task as failed or success (which clears downstream tasks as a side effect), 
still keep task state
+unconditionally today; extending discard-by-default to those paths is tracked 
in
+`#72929 <https://github.com/apache/airflow/issues/72929>`_.
+
+**Why**
+
+Clearing means "run this again". A checkpoint records how far a task got, not 
what it got there
+with, so resuming after the code or the upstream data changed left work done 
before the fix in place
+and silently mixed it with the corrected work. Clearing a task whose external 
job had already
+succeeded was worse: the operator read the stored result back and returned in 
seconds having run
+nothing.
+
+**Keeping the old behaviour**
+
+Pass ``keep_task_state=True`` to the clear task instances endpoint, or tick 
"keep task state" in the
+clear dialog. Use it when nothing about the inputs or the code changed and the 
task should carry on
+where it stopped, or when an external job is still running and you want the 
next attempt to
+reconnect rather than submit a duplicate. To make "keep task state" the 
default so you don't have to
+tick it on every clear, turn it on under Settings > Clearing > "Keep task 
state on clear"; that
+default is per-browser and does not affect anyone else on the same deployment.
+
+Operators with durable execution are worth particular attention. Clearing a 
*failed* task never runs
+``on_kill``, so an external job that outlived its worker is still running, and 
discarding the stored
+id means submitting a second one. The same applies to operators configured to 
leave their job alive
+on kill, such as ``KubernetesPodOperator`` with ``on_kill_action="keep_pod"``.

Review Comment:
   Fixed. Dropped the KubernetesPodOperator/keep_pod example (it doesn't hold 
with the default reattach_on_restart=True), kept the GlueJobOperator one in 
both places.
   
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to