I like the idea Amogh. I think **just changing** pause behaviour might not be good idea for compatibility (someone might rely on it's current behaviour).
But maybe we could have different behaviour of two different buttons/behaviour: "Pause" (current behaviour - stop current tasks and it's downstream), "Pause and drain" - set the Dag to "paused" but let the running dag_runs to continue until completion. It might however require some scheduler changes - I think currently is_paused is used to select run-eligible tasks ? J. On Tue, Sep 8, 2026 at 7:41 AM Amogh Desai <[email protected]> wrote: > Thanks for picking a 4-year-old issue. > The ask itself makes sense. > > One design question that might eventually come up in this thread / PR is: > why do we need a new persisted state > instead of just changing what is_paused does? i.e: keep gating new run > creation on is_paused, but stop freezing > downstream tasks in created runs. That would get you graceful draining > without a schema change or migration. > > From my reading, doing that would silently change the semantics of > *is_paused*. Today, freezes downstream tasks in-flight. > Some operators pause specifically to stop a run's progress, not just to > block new ones. Folding drain behavior into *is_paused* > might change that for everyone already relying on the current freeze + > pause behavior, with no option to opt out, yes? > Worth clarifying. > > Thanks & Regards, > Amogh Desai > > > On Sun, Sep 6, 2026 at 3:15 AM Dheeraj Turaga <[email protected]> > wrote: > > > Hey Everyone, > > > > I wanted to start this discussion to get feedback on PR #72407, which > > proposes adding a "draining" scheduling state to Airflow. > > > > The Problem > > > > Currently, pausing a DAG stops it mid-flight. While running tasks are > > allowed to finish, queued and downstream tasks remain stranded until the > > DAG is unpaused. There is currently no native way to stop starting new > runs > > while allowing those already in flight to complete. > > > > This is a significant pain point during upgrades and maintenance. The > > current workaround as described in the four-year-old Issue #22006 > requires > > users to manually rewrite every DAG schedule to None, wait for runs to > > drain, and then restore the schedules afterward. This process is invasive > > and prone to error. > > > > Proposed Solution: The draining State > > > > The PR introduces a draining state that sits between active and paused. > Key > > behaviors include: > > > > - No New Scheduled Runs: The scheduler creates no new runs for a > draining > > DAG (including scheduled, asset-triggered, and partitioned/rollup paths). > > - Completion of In-Flight Runs: is_paused remains false during the > drain, > > meaning task instances in existing runs are still scheduled and finish > > normally. > > - Automatic Convergence: Once no unfinished runs remain, the scheduler > > automatically moves the DAG to the paused state and writes a > > drain_completed audit log entry. > > > > The core property of this feature is that draining is transient, not a > > third resting state; it always converges to paused. > > > > Implementation Details > > > > Explicit run creation via manual triggers, TriggerDagRunOperator, asset > > materialization, or backfills remains allowed during draining, mirroring > > the behavior of a paused DAG. An earlier revision that blocked these > > actions was reverted to ensure draining is not stricter than the state it > > converges into. > > > > Links: > > > > - PR: https://github.com/apache/airflow/pull/72407 > > - Issue: https://github.com/apache/airflow/issues/22006 > > > > Thanks, > > Dheeraj > > >
