Lee-W commented on code in PR #74026:
URL: https://github.com/apache/airflow/pull/74026#discussion_r4183839425
##########
providers/common/ai/docs/frameworks/index.rst:
##########
@@ -79,10 +79,17 @@ is Airflow's and which part stays yours.
:doc:`../rag_pipelines`.
- No agent
- ``[llamaindex]`` extra
-
-The Strands and ADK integrations, the framework-neutral tool interface under
them, and
-the tracing helper are experimental: they can change or be removed in a minor
release of
-this provider. See :ref:`howto/stability`.
+ * - Claude Agent SDK, through ``HarnessOperator``
Review Comment:
After playing with this a bit, I think the more useful way to frame it is by
where the agent loop runs and what it can touch:
- `AnthropicAgentSessionOperator` / `OpenAIAgentSessionOperator`: the vendor
runs the agent loop on its own servers, in an environment and with resources
defined on the vendor side. Airflow starts the session and waits for it.
- `AgentOperator` (+ `SandboxToolset`): Pydantic AI runs a generic agent
loop on the worker, and the model gets Airflow toolsets, including shell and
file access inside a sandbox.
- What I'm after here: the vendor's own coding agent (Claude Code, Codex),
with its built-in tools and agent loop, running inside a sandbox that the
Airflow deployment provisions and controls.
Compared to the vendor session operators, the work happens in infrastructure
chosen by the deployment (Docker, Modal, OpenSandbox, ...), against a repo
checked out into that sandbox, with egress controlled by the sandbox backend
and results exported to object storage, using credentials from Airflow
connections.
Compared to `AgentOperator`, the sandbox can be the same, but the loop is
the vendor's coding agent—with its own editing tools, context management, and
so on—rather than a generic Pydantic AI loop with generic tools.
Think scheduled "bump dependencies / fix lint / work on this issue" tasks.
This is what I'm currently thinking of:
```python id="f2r1kq"
# common.ai
class BaseAgentOperator(BaseOperator):
# common prompt / output / Airflow integration
...
class AgentOperator(BaseAgentOperator):
# Pydantic AI implementation
...
# anthropic
class ClaudeCodeOperator(BaseAgentOperator):
# Claude Agent SDK implementation
...
# openai
class CodexOperator(BaseAgentOperator):
# Codex implementation
...
```
A few notes on that shape:
- On the OpenAI side, I'm targeting Codex rather than the OpenAI Agents SDK.
The Agents SDK is a framework in roughly the same space as Pydantic AI, so an
operator on top of it would mostly just swap out the framework underneath.
Codex is the actual counterpart to Claude Code.
- Naming the operators after the products makes it clear that they're
different from what we already have.
- The sandbox side (provisioning, repo checkout, exports, egress,
credentials) is vendor-neutral and already lives in `common.ai`, so it would
stay there, while the vendor operators build on top of it.
- I'm not sure yet how much `BaseAgentOperator` can actually share. Today
the overlap is mostly prompt / connection / model / system prompt / XCom;
durable execution, HITL, usage limits, and tool approval are all tied to
Pydantic AI types. I'd rather extract the common base once the first vendor
operator exists and we can see the real duplication, rather than guess at the
abstraction now.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]