GitHub user TongyiDai added a comment to the discussion: airflow.sdk.api.client.ServerResponseError
The important part is the API detail, not the final `ServerResponseError`: the worker tried to mark this task instance as running, but the API says its `previous_state` was already `running`. That means the same TI launch was delivered/executed more than once, or a stale Kubernetes worker pod started after another launch had already claimed it. Your tuning makes that race much easier to trigger: `worker_pods_creation_batch_size: 1000`, `parallelism: 500`, very large scheduling batches, and `job_heartbeat_sec: 300` are unusually aggressive. A five-minute scheduler heartbeat also delays failure/adoption decisions. I would diagnose it in this order: 1. Query the TI by its UUID `01a0a88a-...` and map index 87; identify every worker pod created for try 2 and compare pod creation times/UIDs. 2. Inspect scheduler and KubernetesExecutor logs around correlation ID `01a0a89d-dcac-70f7-ad43-cb175020ceff` for duplicate queue/adoption events. 3. Confirm scheduler replicas are running the same Airflow image/config and that all components use the same JWT secret. 4. Reduce `worker_pods_creation_batch_size` to a conservative value and restore a normal heartbeat while reproducing. Scaling it far above the API/scheduler's processing capacity does not increase useful throughput. 5. Upgrade from 3.0.2 to a current supported Airflow 3 patch before deeper investigation; many executor/API consistency fixes landed after the initial 3.0 releases. Do not retry this as an application exception—the worker is being SIGKILLed because its launch is stale. If it persists on a current version with conservative settings, report it with the TI UUID, both pod UIDs, scheduler logs, correlation ID, and exact executor-event timeline. GitHub link: https://github.com/apache/airflow/discussions/73230#discussioncomment-18459916 ---- This is an automatically sent email for [email protected]. To unsubscribe, please send an email to: [email protected]
