kaxil opened a new pull request, #73962: URL: https://github.com/apache/airflow/pull/73962
Agents whose loop is the Anthropic SDK's `client.beta.messages.tool_runner` can now use common.ai's connection-backed toolsets (`SQLToolset`, `HookToolset`, `ObjectStorageToolset`, sandboxes) without being rebuilt as an `AgentOperator` or a Pydantic AI agent: ```python tools = AirflowTools(SQLToolset(db_conn_id="warehouse")) runner = client.beta.messages.tool_runner(model=..., max_tokens=16000, tools=tools.tools, messages=[...]) message = tools.run(runner) ``` `AirflowTools` (for `Anthropic`) and `AsyncAirflowTools` (for `AsyncAnthropic`) sit next to the Strands and ADK adapters from #73898 and are built on the same framework-neutral `AirflowTool` interface. They make the same promises: - Every result goes through the secret masker. - A failure the model can correct reaches it as a `tool_result` with `is_error: true`, counted against the tool's retry limit once per model turn. - Any other failure fails the task. - Calls are counted under `common_ai.tool_calls` with `framework=anthropic`. The adapters are experimental. There's a guide under Frameworks and a row on the stability page. Checked end to end against Claude through a gateway connection with four runs: - the example Dag; - a registered connection password inside a query result, which reached the model as `***`; - a query on a table outside `allowed_tables`, which came back as `is_error` and which the model recovered from; - a hook that raises, which failed the task after exactly one model request. ## Design rationale **Why `tools.run(runner)` and not the runner's own `until_done()`.** The runner catches every exception a tool raises and sends it to the model as an error result. On its own, a hook raising or a connection's credentials being rejected becomes one more model turn, and the task can succeed on an answer written around the failure. The SDK has no hook to stop that. `run()` drives the loop and calls the runner's public `generate_tool_call_response()` inside a per-turn scope; the runner caches that response and sends it with the next request. That way a failure is raised before the next request, and the retry limit knows which calls share a turn. Driving the runner without `run()` still works, but a failure then goes to the model, and the adapter logs a warning when that happens. **Tools run only on `tool_use` turns.** The runner itself runs a turn's tool calls only when the turn ended in `tool_use`. A turn cut off at `max_tokens` can hold a tool call with incomplete arguments, and `pause_turn` and `compaction` turns are resumed without running any tools. `run()` follows the same rule. The SDK's runner gained that routing in 1.1.0 (1.0.0 ran tools on every stop reason except `refusal`), so the floor of common.ai's `anthropic` extra moves from 1.0.0 to 1.1.0. **After one call in a turn fails, the turn's other calls are skipped.** The runner calls them one at a time, and their results would never be sent, so a tool that writes would run for nothing. **Why common.ai and not the Anthropic provider.** The adapter adapts common.ai's tool interface, which is experimental and 0.x, and it uses common.ai internals: the per-turn retry scope and the metric's framework label. common.ai's `anthropic` extra already carries the SDK for `@task.llm_batch`. The adapter accepts any `Anthropic` or `AsyncAnthropic` client, so it does not depend on the Anthropic provider. The guide points to `AnthropicHook(...).get_conn()` for Bedrock, Vertex AI and Foundry. **`run_coroutine_sync` now copies the caller's context into its worker thread.** When a loop is already running (an `async def` task, a notebook), the shared helper runs the coroutine on a worker thread. Until now that thread did not see the caller's context variables, so the sync adapter lost track of the run and a fatal failure went to the model. Copying the context fixes this for every caller of the helper. The helper also now calls `asyncio.run` outside its `except RuntimeError` block, so a tool's exception no longer arrives in the task log chained to a "no running event loop" error. ## Gotchas - `tool_runner` is a beta helper of the Anthropic SDK, which is one reason the adapter is experimental. - The streaming runner (`tool_runner(..., stream=True)`) is not supported. --- * Read the **[Pull Request Guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#pull-request-guidelines)** for more information. Note: commit author/co-author name and email in commits become permanently public when merged. * For fundamental code changes, an Airflow Improvement Proposal ([AIP](https://cwiki.apache.org/confluence/display/AIRFLOW/Airflow+Improvement+Proposals)) is needed. * When adding dependency, check compliance with the [ASF 3rd Party License Policy](https://www.apache.org/legal/resolved.html#category-x). * For significant user-facing changes create newsfragment: `{pr_number}.significant.rst`, in [airflow-core/newsfragments](https://github.com/apache/airflow/tree/main/airflow-core/newsfragments). You can add this file in a follow-up commit after the PR is created so you know the PR number. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
