Lee-W commented on code in PR #74026:
URL: https://github.com/apache/airflow/pull/74026#discussion_r4183839425


##########
providers/common/ai/docs/frameworks/index.rst:
##########
@@ -79,10 +79,17 @@ is Airflow's and which part stays yours.
        :doc:`../rag_pipelines`.
      - No agent
      - ``[llamaindex]`` extra
-
-The Strands and ADK integrations, the framework-neutral tool interface under 
them, and
-the tracing helper are experimental: they can change or be removed in a minor 
release of
-this provider. See :ref:`howto/stability`.
+   * - Claude Agent SDK, through ``HarnessOperator``

Review Comment:
   After playing with this a bit, I think the more useful way to frame it is by 
where the agent loop runs and what it can touch:
   
   - `AnthropicAgentSessionOperator` / `OpenAIAgentSessionOperator`: the vendor 
runs the agent loop on its own servers, in an environment and with resources 
defined on the vendor side. Airflow starts the session and waits for it.
   - `AgentOperator` (+ `SandboxToolset`): Pydantic AI runs a generic agent 
loop on the worker, and the model gets Airflow toolsets, including shell and 
file access inside a sandbox.
   - What I'm after here: the vendor's own coding agent (Claude Code, Codex), 
with its built-in tools and agent loop, running inside a sandbox that the 
Airflow deployment provisions and controls.
   
   Compared to the vendor session operators, the work happens in infrastructure 
chosen by the deployment (Docker, Modal, OpenSandbox, ...), against a repo 
checked out into that sandbox, with egress controlled by the sandbox backend 
and results exported to object storage, using credentials from Airflow 
connections.
   
   Compared to `AgentOperator`, the sandbox can be the same, but the loop is 
the vendor's coding agent—with its own editing tools, context management, and 
so on—rather than a generic Pydantic AI loop with generic tools.
   
   Think scheduled "bump dependencies / fix lint / work on this issue" tasks.
   
   This is what I'm currently thinking of:
   
   ```python id="f2r1kq"
   # common.ai
   class BaseAgentOperator(BaseOperator):
       # common prompt / output / Airflow integration
       ...
   
   
   class AgentOperator(BaseAgentOperator):
       # Pydantic AI implementation
       ...
   
   
   # anthropic
   class ClaudeCodeOperator(BaseAgentOperator):
       # Claude Agent SDK implementation
       ...
   
   
   # openai
   class CodexOperator(BaseAgentOperator):
       # Codex implementation
       ...
   ```
   
   A few notes on that shape:
   
   - On the OpenAI side, I'm targeting Codex rather than the OpenAI Agents SDK. 
The Agents SDK is a framework in roughly the same space as Pydantic AI, so an 
operator on top of it would mostly just swap out the framework underneath. 
Codex is the actual counterpart to Claude Code.
   - Naming the operators after the products makes it clear that they're 
different from what we already have.
   - The sandbox side (provisioning, repo checkout, exports, egress, 
credentials) is vendor-neutral and already lives in `common.ai`, so it would 
stay there, while the vendor operators build on top of it.
   - I'm not sure yet how much `BaseAgentOperator` can actually share. Today 
the overlap is mostly prompt / connection / model / system prompt / XCom; 
durable execution, HITL, usage limits, and tool approval are all tied to 
Pydantic AI types. I'd rather extract the common base once the first vendor 
operator exists and we can see the real duplication, rather than guess at the 
abstraction now.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to