Hi all,

A few of us have been working on this proposal and would like community 
feedback on:

AIP-97 Task Failure Classification

https://cwiki.apache.org/confluence/x/0ITMFw

Airflow’s executor runs separately from the task worker, so it can still read 
backend failure details after the worker stops. AIP-97 passes confirmed causes, 
such as Kubernetes preemption, into failure handling for listeners, logs and 
metrics. When the cause is unclear, it stays unclassified.

This helps integrations distinguish infrastructure interruptions from 
application errors and route alerts appropriately. Ordinary retries, clears and 
DAG callback interfaces stay unchanged.



We’ve separated infrastructure retry decisions into:

AIP-122 Infrastructure Retries

https://cwiki.apache.org/confluence/x/-pXwGg

When enabled, AIP-122 uses confirmed infrastructure failures to preserve 
ordinary task retries, within an administrator-configured limit.

For example, consider a Spark task with retries=1. Its first attempt is 
interrupted by a confirmed Kubernetes preemption. The retry starts, but the 
Spark job fails during that second attempt. There is now no retry left: the 
user-defined retry was consumed by an infrastructure interruption rather than 
an application failure. With AIP-122 enabled, the first infrastructure failure 
can receive an infrastructure retry, leaving the ordinary retry available for 
the later Spark failure.

AIP-97 can deliver value independently while we work with the community on 
AIP-122.

Both AIPs include implementation POCs, runnable examples and end-to-end 
evidence.



Thank you,

Stefan, on behalf of the AIP-97 and AIP-122 co-authors



Reference: Previous dev-list discussion 
<https://lists.apache.org/thread/g4jhd0vz71x6jm4z8p6hjv9z7rmtb3jk>, bcc’d folks 
from there FYI.

Reply via email to