kaxil opened a new pull request, #73450:
URL: https://github.com/apache/airflow/pull/73450

   `LLMRetryPolicy` asked the model four things: a `category` typed as a bare 
`str`, `should_retry`, `suggested_delay_seconds`, and a prose `reasoning`. 
Three of those were the model reciting back a table the prompt had just given 
it, and one was a free-text field. A classifier model such as TypeSafe's Jev 
refuses free text and unbounded numbers before the request leaves the process. 
The retry policy runs on every task failure, which is where a 200 ms call at 
$0.042 per million input tokens pays off most, and it was the one Common AI 
surface such a model could not run on. On a text model, `category="auth", 
should_retry=True` was reachable and trusted.
   
   The model now answers one question: which of the policy's `categories` this 
failure is. Everything else comes from the author's table in the worker:
   
   ```python
   LLMRetryPolicy(
       llm_conn_id="jev_default",
       min_confidence=0.6,
       categories={
           "queued": ErrorCategory("Statement queued or a concurrency limit 
reached.", delay=timedelta(seconds=120)),
           "warehouse_suspended": ErrorCategory("The warehouse is suspended and 
will auto-resume.", delay=timedelta(seconds=30)),
           "schema_drift": ErrorCategory("A referenced column or table does not 
exist.", retry=False),
           "permanent": ErrorCategory("Will fail identically on every 
attempt.", retry=False, min_confidence=0.85),
       },
       fallback_rules=[...],
   )
   ```
   
   Each `ErrorCategory` carries what the model reads (`description`, sent in 
the output schema next to the name), what the policy does (`retry`, `delay`), 
and how sure the model has to be (`min_confidence`). Under the bar, or when a 
bar is set and the model reported no confidence, the answer is discarded and 
the policy takes the path it already took when the model call failed: 
`fallback_rules`, then the task's own retry behaviour. The seven previous 
categories remain as `DEFAULT_CATEGORIES`, an example taxonomy with the same 
retry/fail split and delays as before.
   
   The `retry_reason` a person reads is now generated from the decision, 
`category=network confidence=0.91 threshold=0.60 action=retry delay=10s`, 
instead of whatever the model wrote.
   
   ## Design rationale
   
   **The category set is a constructor parameter, not a module-level 
`Literal`.** A classifier needs the set closed per request, and pinning seven 
names would close it for everyone and break every custom taxonomy already 
written against `instructions`. With the names and descriptions in the schema, 
`instructions` goes back to being what it is good for: teaching the model your 
stack's error strings.
   
   **Per-category bars inherit the policy bar, and need one to inherit from.** 
Same rule as `BranchOption.min_confidence` in #73368: a category bar with no 
policy bar is rejected at construction, so a category added later without a bar 
of its own still inherits one.
   
   **An unsure answer falls back; it does not open a review.** A retry policy 
runs at failure time on the worker, with no task to pause, so under the bar it 
takes the `fallback_rules` path it already had for a failed model call, and 
`RetryDecision.default()` after that keeps the task's own backoff and remaining 
attempts. The bar is a flat `min_confidence` rather than the operators' 
`DecisionPolicy`, because the only other field on that object is 
`on_uncertain`, which picks review or fail, and neither applies at failure time 
on a worker. The per-option override is spelled the same on both surfaces 
(`BranchOption.min_confidence`, `ErrorCategory.min_confidence`).
   
   **No `reasoning` field, even optional.** The classifier adapter refuses a 
text field before the request is sent, so keeping one, even optional, would 
need a special case per model type. The generated `retry_reason` line lists the 
category, confidence, bar and action.
   
   **The `Choices` builder moves into the decision utils.** The branch operator 
and the retry policy present described options to the model the same way, on 
pydantic-ai 2.46+ as `Choices` and before that as an `Enum` with the same 
schema. The branch operator's behaviour is unchanged; only where the helper 
lives moved.
   
   Common AI is 0.x, so the removed names (`ErrorClassification`, 
`should_retry`, `suggested_delay_seconds`, `reasoning`) go without a 
deprecation path. An existing `LLMRetryPolicy(llm_conn_id=...)` keeps working 
with the same defaults. Two things do change for 0.9.0 users, and the changelog 
note says so: an import of `ErrorClassification` fails at parse time, and 
custom `instructions` that told the model which categories to retry or what 
delay to use still parse but no longer steer anything, because those decisions 
moved into `categories`. The 0.9.0 docs recommended writing delays into 
`instructions` (the Snowflake example asked for 120 seconds on `rate_limit`), 
and a Dag that followed them now gets the default 60 seconds with no error.
   
   ## Gotchas
   
   - `min_confidence` with a text model discards every answer, by design: a 
model swap must not silently switch off a control the author set. Remove the 
bar to run a text model.
   - Calibrate the bar from your own failures. Run with no bar first and read 
the logged `confidence=` values per category; the confidence measures how 
concentrated the probability distribution is on one category, it is not a 
probability of being correct, and it moves between model versions, so pin 
`typesafe:jev-1.13.0` rather than `jev-latest`.
   
   Part of the classifier-model work started in #73363 and #73368.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to