uranusjr commented on code in PR #70657:
URL: https://github.com/apache/airflow/pull/70657#discussion_r3793754355


##########
airflow-core/docs/core-concepts/overview.rst:
##########
@@ -192,90 +192,30 @@ code is never executed in the context of the *scheduler*.
     bundle version when dispatching each task. If needed, the cadence of sync 
and scan
     of the *Dag bundle* can be configured.
 
-Task execution architecture
----------------------------
-
-The diagrams above show how Airflow's components are *deployed*. The diagrams 
below instead show what happens
-*inside a worker when a task actually runs* — how the Task SDK, the Supervisor 
and Coordinator processes, and
-the language runtimes work together: which processes are involved, and the 
classes and protocols they use to
-communicate.
-
-.. _overview-task-sdk-execution-architecture:
-
-Python Task SDK execution
-.........................
-
-When a *worker* actually runs a task, it does not run the user's code 
directly. Instead it starts a
-lightweight **Supervisor** that runs in its own **native operating-system 
process** and
-*forks* a second native process in which the **Task SDK** runtime 
(``task_runner``) executes the user code.
-The two processes talk over a socket, and the Supervisor is the only side that 
ever holds the short-lived
-task JWT or talks to the *Execution API* — the user's code never sees the 
token and never touches the
-database.
-
-The same runtime can also run *in-process* (a single Python process, no fork, 
no sockets, no HTTP) for
-``dag.test()`` and local runs. The diagram below contrasts the two paths and 
marks where each Python process
-lives:
-
-.. image:: ../img/diagram_task_sdk_execution_architecture.png
-
-The message flow of a supervised run — startup, running the user code, proxied 
Connection/Variable/XCom
-lookups, heartbeats, and reporting the final state — is shown below as a 
sequence diagram, with each process
-on its own lifeline. The **Supervisor** sits in the middle, so the Task ↔ 
Supervisor request/response
-round-trip (the task asks for a Connection/Variable/XCom and gets the answer 
back) reads as arrows going back
-and forth between neighboring lifelines. Each arrow is numbered, colored by 
its sender, and labeled with the
-message class or protocol used:
-
-.. image:: ../img/diagram_task_sdk_execution_sequence.png
+.. _overview-task-execution-architecture:
 
-.. _overview-non-python-language-sdks:
-
-Non-Python language SDKs (Go and Java)
-......................................
-
-The Task Execution Interface (TEI) introduced in AIP-72 is language-agnostic, 
so a task can also be written in
-a **compiled, non-Python language**. A Python Dag still declares the task with 
``@task.stub(queue=...)`` (so
-Python and non-Python tasks can be mixed in one Dag), but the actual work is 
delegated to the matching runtime.
-There are currently **two different integration styles** — the Go SDK runs a 
standalone worker, while the Java
-SDK plugs into the existing Python Supervisor.
-
-.. _overview-go-sdk-architecture:
-
-**Go SDK — standalone edge worker.** The `Go Task SDK
-<https://github.com/apache/airflow/blob/main/go-sdk/README.md>`_ has **no 
Python Supervisor and no msgpack
-stdin socket**. A long-running, compiled **edge worker** 
(``airflow-go-edge-worker``) *pulls* work from the
-**Edge Executor API**, launches the user's compiled Dag bundle as a 
**go-plugin (gRPC) subprocess**, and
-invokes the task over gRPC. The task then uses the **native TEI client** to 
reach the **Execution API**
-directly over HTTPS — so, unlike the Python task, it holds the task JWT itself:
-
-.. image:: ../img/diagram_native_language_sdk_architecture.png
+Task execution architecture
+----------------------------
 
-.. _overview-java-sdk-architecture:
+The diagrams above show how Airflow's components are *deployed*.
+This section covers what you should expect to be running on a *worker* while 
tasks execute, so you can size workers for the memory and start-up cost that 
each concurrent task instance adds.
 
-**Java (JVM) SDK — Coordinator plugged into the Supervisor.** The `Java Task 
SDK
-<https://github.com/apache/airflow/blob/main/java-sdk/README.md>`_ takes the 
opposite approach: it *reuses* the
-existing Python Supervisor through a new **Coordinator** layer. 
``CoordinatorManager`` resolves the task's
-``queue`` to a ``BaseCoordinator`` — ``JavaCoordinator`` for the ``java`` 
queue, or the built-in
-``_PythonCoordinator`` otherwise. ``JavaCoordinator`` opens two loopback-TCP 
servers, spawns a **JVM bundle
-process** with ``subprocess.Popen``, and drives it with 
``_JavaActivitySubprocess`` (a subclass of the shared
-``ActivitySubprocess``). The JVM connects *back* over TCP and speaks the 
**same msgpack protocol** as a Python
-task, so the Python side heartbeats, manages state, and **proxies every 
Execution-API call** — meaning the JVM
-task, like a Python task, never holds the task JWT itself:
+A *worker* never runs task code in its own process.
+For every task instance it starts a **new subprocess**, supervises it, and 
tears it down once the task instance finishes.
+That is true whatever language the task is written in — only the kind of 
subprocess differs:
 
-.. image:: ../img/diagram_java_sdk_execution_architecture.png
+* **Python** — a forked Python interpreter.
+* **Java** — a brand-new JVM instance.
+* **Go** — a brand-new process of the compiled task binary.
 
-The end-to-end workflow of a Java task — from ``@task.stub`` through the 
coordinator, the JVM subprocess, the
-proxied Connection/Variable/XCom lookups, and reporting the final state — is 
shown below as a sequence diagram.
-As above, the **Supervisor** is the central lifeline, so the JVM ↔ Supervisor 
round-trip over loopback TCP is
-drawn as arrows going back and forth to its neighbours:
+.. image:: ../img/diagram_task_execution_architecture.png
 
-.. image:: ../img/diagram_java_sdk_execution_sequence.png
+Nothing is pooled or reused between task instances, so *N* concurrent task 
instances cost *N* subprocesses plus anything the task code itself spawns.
+Sizing a worker therefore means budgeting for the peak number of task 
instances it runs at once, not for the worker process alone.
 
-.. note::
+The worker process is also the only side that holds the task's short-lived 
credentials and talks to the *Execution API*, so the subprocess running the 
task code never reaches the metadata database directly, in any language.

Review Comment:
   Not sure if we want to add a note here saying the Edge Worker is another 
story. (It does not follow the same rules even in Python.)



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to