kaxil opened a new pull request, #74266:
URL: https://github.com/apache/airflow/pull/74266

   pydantic-ai 2.53.0 added `SystemOneModel` (pydantic/pydantic-ai#8942), which 
runs any decision model served over the same `POST /v1/systemone` API as 
TypeSafe's Jev: decision models in Ollama 0.35+, AWS's [Strands 
Decider](https://github.com/strands-labs/strands-decider), 
[Kev](https://github.com/jaredpalmer/kev), CLM, Laya. The Common AI docs still 
said Jev was the only decision model pydantic-ai supports.
   
   This needs no code change. `PydanticAIHook` resolves an unknown prefix 
through its generic path, `infer_provider_class(prefix)(api_key=conn.password, 
base_url=conn.host)`, and `SystemOneProvider` takes exactly those two 
arguments. A `pydanticai` connection with the server URL in **Host** and 
`system-one:<model>` as its **Model** is the whole integration. The end-to-end 
run below confirms it.
   
   So this PR is docs and examples:
   
   - `classifier_models.rst` becomes `decision_models.rst` (redirect added). 
"Decision model" is the term pydantic-ai, Strands and Cloudflare all use now, 
and the operators already call the gate `decision_policy`. Setup covers both 
backends, `typesafe:` and `system-one:`.
   - `example_classifier_model.py` becomes `example_decision_model.py`, and its 
Dags become `example_decision_model_branch` / 
`example_decision_model_confidence`. These Dags, 
`example_llm_branch_decision_policy` and the decision-model retry policies now 
read a `decision_default` connection and take the model from it rather than 
hard-coding `model_id="typesafe:jev-1.13.0"`, so the same Dag runs on either 
backend.
   - Docstrings and the retry-policy log hint stop saying "(TypeSafe's)".
   
   ## Design rationale
   
   **The examples now describe every option.** The first end-to-end run failed 
with a 422: for an option with no description, pydantic-ai sends `criteria: 
{option: null}`, and Strands Decider 0.1.0 accepts only string criteria. It 
also requires question text, which an operator with the default empty 
`system_prompt` does not send. The branch example now passes `branches=`, and 
the classify example uses an `Enum` with `UseEnumMemberDocstrings` instead of a 
bare `Literal`. The decision models page documents both requirements, since 
they are the first thing a Strands Decider user would hit, and the `branches` 
docstring on `LLMBranchOperator` says the same.
   
   **A `system-one:` model's option cap is the server's.** pydantic-ai refuses 
a question over Jev's cap before sending it. A model reached through a 
connection carries no `DecisionModelProfile`, so a question over a System One 
server's cap (Ollama's is 26) comes back as an HTTP error instead. The page 
says so.
   
   **Cloudflare's Clef is not named.** Its changelog says it "follows the 
System One API", but Workers AI serves it at `/ai/run`. I have not confirmed 
that `SystemOneModel` can reach it, so the docs don't claim it.
   
   ## Gotchas
   
   - `example_decision_model.py` imports `UseEnumMemberDocstrings`, which 
shipped in pydantic-ai-slim 2.46. The provider floor stays at 2.33. The 
example-Dag import test already skips under lowest dependencies, and the page 
and the example state the versions needed (2.46 for the examples, 2.53 for 
`system-one:`). On an older pydantic-ai, a `system-one:` connection fails with 
the hook's existing "is not a provider pydantic-ai recognizes" error.
   - Example Dag ids and the examples' connection id change (`jev_default` to 
`decision_default`). The retry-policy example keeps its note that its 
confidence bars were calibrated on `jev-1.13.0` and need measuring again for 
another model.
   - These examples were not re-run against Jev for this PR. For Jev, the 
change to them is the connection id and the added option descriptions.
   
   ## End-to-end run
   
   Airflow `main` under breeze, sqlite, pydantic-ai-slim 2.53.0. Strands 
Decider 2B (`StrandsAgents/strands-decider-2B-hobson-v19`) served locally by 
`strands-decider serve`. The three Dags ran unmodified from the provider's 
example-Dag bundle. The classify step was re-run through `PydanticAIHook` after 
the example's `Enum` got its class docstring: still `resource`, confidence 0.61.
   
   | Dag | Result |
   |---|---|
   | `example_llm_branch_decision_policy` | Picked `rerun` at confidence 0.74, 
above the 0.6 bar. `page_oncall` and `ignore` skipped |
   | `example_decision_model_branch` | Picked `grant_bucket_write`. 
`restore_deleted_bucket` and `wait_and_retry` skipped |
   | `example_decision_model_confidence` | Classified `resource` at 0.63, and 
`act` filed it for review (0.5 to 0.8 band) |
   
   The connection, with the server URL in Host and the `system-one:` model:
   
   ![Connection list](./e2e_conn_list.png)
   ![Connection model field](./e2e_conn_edit2.png)
   
   The `decision` XCom from `example_llm_branch_decision_policy`, with the 
model, confidence and probabilities Strands Decider returned:
   
   ![Decision record](./e2e_decision_xcom.png)
   
   `example_decision_model_confidence`. The red runs in the grid are the 
earlier attempts: the previous example's undescribed options (the 422 above), 
plus a crash of the local model server.
   
   ![Classify result](./e2e_classify_xcom.png)
   
   ## Follow-ups
   
   `ToolCallJudge` (pydantic-ai-harness 0.53) is a natural fit for unattended 
agents, but it builds its judge model when the Dag file is parsed, so today its 
credentials can only come from environment variables, not an Airflow 
connection. That gets its own PR.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to