shivaam commented on code in PR #68961:
URL: https://github.com/apache/airflow/pull/68961#discussion_r4174316278
##########
airflow-core/src/airflow/api_fastapi/core_api/openapi/_private_ui.yaml:
##########
@@ -3643,6 +3643,9 @@ components:
description: Interval in seconds between the reference time and the
deadline.
Null for a dynamic interval (e.g. a VariableInterval) whose value
is only
resolved at scheduler evaluation time.
+ fire_on_failure:
Review Comment:
I am adding a third approach below :)
My understanding is that failure notifications and deadline notifications
serve different purposes. on_failure_callback tells users that a Dag has
failed; a deadline alert tells them that the expected outcome wasn’t achieved
by the deadline. I’d prefer keeping that distinction explicit so users know
they would need both.
Under that approach, deadline alerts would retain their deadline-based
timing. Users who already receive failure notifications could optionally
suppress the pending deadline alert after failure, perhaps through a flag such
as suppress_deadline_on_failure. Users wanting immediate failure notifications
would use on_failure_callback.
For example, a Dag could fail at 7 a.m. with a 9 a.m. deadline. The on-call
receives the failure notification, clears the failed tasks, and attempts
recovery. If the pending deadline alert was cancelled at 7 a.m., they would
lose the notification that recovery still hadn’t succeeded by 9 a.m.
That leaves a question: should clearing the failed tasks restore deadline
monitoring? From an on-call perspective, I would want that coverage restored
while attempting recovery.
Would this separation make sense, rather than making deadline callbacks fire
immediately on failure?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]