Hi Constance,

I agree with you that we need a bulk pause/drain for ease of use.

I already have the PR ready https://github.com/apache/airflow/pull/73055

Will ensure everything needed for drain is in before Airflow 3.4

Thanks,
Dheeraj

On Mon, Sep 14, 2026 at 4:02 PM Constance Martineau <[email protected]>
wrote:

> Sending this again, as I just realized my response did not go through!
>
> Hi Dheeraj,
>
> I think draining is a useful feature! I don’t have a strong opinion on
> whether or not draining should be its own state (having it be a transient
> state seems reasonable, and it seems like you have consensus). What I am
> questioning is the need to have this feature be a per dag option. The
> Airflow UI is already really busy for individual dags, and having a modal
> pop up to ask whether you really mean drain or pause will be really
> annoying, especially for operators who will need to do this for every
> active dag at scale. If there are 200 dags, that’s 400 clicks. Since
> upgrade prep is the primary use-case, why not start by making this a bulk
> feature?
>
> This only fell on my radar because of the dev list conversation, but I’m
> happy to take this to the PR or start a new email chain if you prefer.
>
> Constance
>
> On Fri, Sep 11, 2026 at 1:06 PM Dheeraj Turaga <[email protected]>
> wrote:
>
>> Hi Constance,
>>
>> As someone who operates both as a Dag Author and Airflow Instance manager,
>> I see having the ability to drain useful.
>>
>> As a dag author, there are many cases where I want existing long running
>> scheduled dag runs
>> to finish before I disable the pipeline before new runs kick in.
>> I also in some cases may want to pause or "halt" the dag runs
>> immediately.
>> This current pr addresses this particular usecase
>>
>> As a Airflow Infra manager, I would like to drain the instance before a
>> planned maintenance
>> and want to give every dag the opportunity to finish before I do so.
>> The current pr does not address this bulk action/cli action (YET).
>> I plan to create follow up prs to address the scenario
>>
>> In any case, both scenarios would require a general agreement on the need
>> for a
>> transient "drain" state which Is what im hoping to get consensus on with
>> this thread.
>>
>> Thanks,
>> Dheeraj
>>
>> On Thu, Sep 10, 2026 at 7:28 PM Constance Martineau via dev <
>> [email protected]> wrote:
>>
>>> Hi all,
>>>
>>> Jumping in with a related question. Skimming the original issue, the
>>> motivating case is explicitly instance-wide: "we don't want to fail/kill
>>> running DAGs... Once none of the DAGs are in a running state, we perform
>>> the upgrades... Once upgrades are done, we update the DAG schedules back
>>> to
>>> their original value". That reads as an operator preparing for a
>>> maintenance window, not a Dag owner reaching for a per-Dag control.
>>>
>>> I have two UX concerns, from each side of that use case. The design
>>> reserves the drain-or-pause choice for the moment a Dag has runs
>>> in-flight,
>>> based on the logic that this is when the choice actually matters.
>>> However,
>>> for a Dag owner, having runs in flight isn't the exception, it's
>>> precisely
>>> the state a Dag is in when I need to pause it, because something running
>>> is
>>> going wrong. That's exactly the moment I want pause to still mean "stop
>>> it," not a fork to resolve under pressure. And as an operator preparing
>>> for
>>> an upgrade, having to toggle and confirm this per Dag, one at a time,
>>> across however many active Dags there are, is close to the same manual,
>>> error-prone process this is meant to replace.
>>>
>>> Would a bulk action cover the use-case better? Like, drain and pause
>>> everything currently active in one click, and resume everything that was
>>> active in one click afterward? That directly matches the upgrade scenario
>>> without adding a new decision to the single-Dag pause flow. A per-Dag
>>> version only justifies its complexity if there's a real need to
>>> gracefully
>>> drain one specific Dag independent of an upgrade. I haven't seen that use
>>> case motivated anywhere in this thread, every example so far points to
>>> draining everything before touching the cluster.
>>>
>>> Curious what others think, happy to be wrong here.
>>>
>>> Thanks, Constance
>>>
>>> On Thu, Sep 10, 2026 at 12:20 PM Shubham Raj <[email protected]>
>>> wrote:
>>>
>>> > Hey,
>>> >
>>> > I agree with the direction of the PR. Given that is_paused serves
>>> several
>>> > critical functions, overloading it to include draining behavior would
>>> > likely cause unintended side effects across the system. By keeping
>>> > "draining" as a distinct, transient state, we can avoid those
>>> complications
>>> > while effectively solving the need for graceful task completion. This
>>> will
>>> > also significant improvement for managing Dags during upgrades.
>>> >
>>> >
>>> > On Wed, 9 Sep 2026 at 20:15, Elad Kalif <[email protected]> wrote:
>>> >
>>> > > > The drain here is a transient which eventually converges to pause.
>>> > >
>>> > > I agree. drain is intermediate state. When draining is completed the
>>> > status
>>> > > should be changed to paused.
>>> > >
>>> > > On Wed, Sep 9, 2026 at 1:49 PM Jarek Potiuk <[email protected]>
>>> wrote:
>>> > >
>>> > > > > The drain here is a transient which eventually converges to
>>> pause.
>>> > > >
>>> > > > Yep. That was my thinking. You summarized it all in one sentence :)
>>> > > >
>>> > > > On Wed, Sep 9, 2026 at 7:03 AM Dheeraj Turaga <
>>> [email protected]
>>> > >
>>> > > > wrote:
>>> > > >
>>> > > > > Hi Amogh,
>>> > > > >
>>> > > > > On Jarek's thought, I agree with him that just changing the pause
>>> > > > behavior
>>> > > > > is probably not a good idea.
>>> > > > > I can see instances where a user may want to instantly pause a
>>> dag in
>>> > > one
>>> > > > > case
>>> > > > > but want to drain a dag in another case. I myself have such a
>>> need
>>> > for
>>> > > > > different dags.
>>> > > > >
>>> > > > > For example, during airflow upgrades, I would like to gracefully
>>> > drain
>>> > > my
>>> > > > > whole instance before upgrade.
>>> > > > > While during regular operation, I may want to prevent my dag from
>>> > > > launching
>>> > > > > further tasks.
>>> > > > >
>>> > > > > The drain here is a transient which eventually converges to
>>> pause.
>>> > > > >
>>> > > > > Draining keeps task scheduling enabled, and only automated dagrun
>>> > > > creation
>>> > > > > paths are gated.
>>> > > > > *is_paused* is set to true only when the dagrun are completed.
>>> Manual
>>> > > > > triggers are still allowed
>>> > > > > simialr to the behavior when *is_paused* set to false.
>>> > > > >
>>> > > > > It would be great to get your thoughts on the PR once you get a
>>> > chance
>>> > > to
>>> > > > > play with it.
>>> > > > >
>>> > > > > Dheeraj
>>> > > > >
>>> > > > > On Tue, Sep 8, 2026 at 11:19 PM Amogh Desai <
>>> [email protected]>
>>> > > > wrote:
>>> > > > >
>>> > > > > > Yep, *is_paused* does more than block new runs. The scheduler
>>> > checks
>>> > > it
>>> > > > > > directly when deciding which
>>> > > > > > tasks to run next. That is what freezes downstream tasks
>>> today, and
>>> > > it
>>> > > > is
>>> > > > > > also why I would push back on
>>> > > > > > "same is_paused but two buttons."
>>> > > > > >
>>> > > > > > If you are suggesting that "Pause and drain" sets
>>> *is_paused=True*
>>> > > > right
>>> > > > > > away, those two checks still need a way
>>> > > > > > to know "this one is a soft pause, let it keep going." That
>>> means
>>> > > > adding
>>> > > > > a
>>> > > > > > second flag anyway, just hidden behind
>>> > > > > > the same name. And *is_paused* shows up in other places like
>>> > backfill
>>> > > > > pause
>>> > > > > > logic and asset triggered runs. Each of
>>> > > > > > those would need checking to make sure a "soft paused" Dag
>>> still
>>> > acts
>>> > > > > like
>>> > > > > > an active one everywhere except the one
>>> > > > > > spot it should not. From the PR, it avoids all that by keeping
>>> > > > > > *is_paused=False* during drain, so every existing check
>>> > > > > > keeps working as it should. It only adds new checks in a few
>>> places
>>> > > > that
>>> > > > > > decide whether to start a new run.
>>> > > > > >
>>> > > > > > Thanks & Regards,
>>> > > > > > Amogh Desai
>>> > > > > >
>>> > > > > >
>>> > > > > > On Tue, Sep 8, 2026 at 4:29 PM Jarek Potiuk <[email protected]>
>>> > > wrote:
>>> > > > > >
>>> > > > > > > I like the idea Amogh. I think **just changing** pause
>>> behaviour
>>> > > > might
>>> > > > > > not
>>> > > > > > > be good idea for compatibility (someone might rely on it's
>>> > current
>>> > > > > > > behaviour).
>>> > > > > > >
>>> > > > > > > But maybe we could have different behaviour of two
>>> > > > > > > different buttons/behaviour: "Pause" (current behaviour -
>>> stop
>>> > > > current
>>> > > > > > > tasks and it's downstream), "Pause and drain" - set the Dag
>>> to
>>> > > > "paused"
>>> > > > > > but
>>> > > > > > > let the running dag_runs to continue until completion.
>>> > > > > > >
>>> > > > > > > It might however require some scheduler changes - I think
>>> > currently
>>> > > > > > > is_paused is used to select run-eligible tasks ?
>>> > > > > > >
>>> > > > > > > J.
>>> > > > > > >
>>> > > > > > >
>>> > > > > > > On Tue, Sep 8, 2026 at 7:41 AM Amogh Desai <
>>> > [email protected]>
>>> > > > > > wrote:
>>> > > > > > >
>>> > > > > > > > Thanks for picking a 4-year-old issue.
>>> > > > > > > > The ask itself makes sense.
>>> > > > > > > >
>>> > > > > > > > One design question that might eventually come up in this
>>> > thread
>>> > > /
>>> > > > PR
>>> > > > > > is:
>>> > > > > > > > why do we need a new persisted state
>>> > > > > > > > instead of just changing what is_paused does? i.e: keep
>>> gating
>>> > > new
>>> > > > > run
>>> > > > > > > > creation on is_paused, but stop freezing
>>> > > > > > > > downstream tasks in created runs. That would get you
>>> graceful
>>> > > > > draining
>>> > > > > > > > without a schema change or migration.
>>> > > > > > > >
>>> > > > > > > > From my reading, doing that would silently change the
>>> semantics
>>> > > of
>>> > > > > > > > *is_paused*. Today, freezes downstream tasks in-flight.
>>> > > > > > > > Some operators pause specifically to stop a run's
>>> progress, not
>>> > > > just
>>> > > > > to
>>> > > > > > > > block new ones. Folding drain behavior into *is_paused*
>>> > > > > > > > might change that for everyone already relying on the
>>> current
>>> > > > freeze
>>> > > > > +
>>> > > > > > > > pause behavior, with no option to opt out, yes?
>>> > > > > > > > Worth clarifying.
>>> > > > > > > >
>>> > > > > > > > Thanks & Regards,
>>> > > > > > > > Amogh Desai
>>> > > > > > > >
>>> > > > > > > >
>>> > > > > > > > On Sun, Sep 6, 2026 at 3:15 AM Dheeraj Turaga <
>>> > > > > [email protected]
>>> > > > > > >
>>> > > > > > > > wrote:
>>> > > > > > > >
>>> > > > > > > > > Hey Everyone,
>>> > > > > > > > >
>>> > > > > > > > > I wanted to start this discussion to get feedback on PR
>>> > #72407,
>>> > > > > which
>>> > > > > > > > > proposes adding a "draining" scheduling state to Airflow.
>>> > > > > > > > >
>>> > > > > > > > > The Problem
>>> > > > > > > > >
>>> > > > > > > > > Currently, pausing a DAG stops it mid-flight. While
>>> running
>>> > > tasks
>>> > > > > are
>>> > > > > > > > > allowed to finish, queued and downstream tasks remain
>>> > stranded
>>> > > > > until
>>> > > > > > > the
>>> > > > > > > > > DAG is unpaused. There is currently no native way to stop
>>> > > > starting
>>> > > > > > new
>>> > > > > > > > runs
>>> > > > > > > > > while allowing those already in flight to complete.
>>> > > > > > > > >
>>> > > > > > > > > This is a significant pain point during upgrades and
>>> > > maintenance.
>>> > > > > The
>>> > > > > > > > > current workaround as described in the four-year-old
>>> Issue
>>> > > #22006
>>> > > > > > > > requires
>>> > > > > > > > > users to manually rewrite every DAG schedule to None,
>>> wait
>>> > for
>>> > > > runs
>>> > > > > > to
>>> > > > > > > > > drain, and then restore the schedules afterward. This
>>> process
>>> > > is
>>> > > > > > > invasive
>>> > > > > > > > > and prone to error.
>>> > > > > > > > >
>>> > > > > > > > > Proposed Solution: The draining State
>>> > > > > > > > >
>>> > > > > > > > > The PR introduces a draining state that sits between
>>> active
>>> > and
>>> > > > > > paused.
>>> > > > > > > > Key
>>> > > > > > > > > behaviors include:
>>> > > > > > > > >
>>> > > > > > > > >   - No New Scheduled Runs: The scheduler creates no new
>>> runs
>>> > > for
>>> > > > a
>>> > > > > > > > draining
>>> > > > > > > > > DAG (including scheduled, asset-triggered, and
>>> > > partitioned/rollup
>>> > > > > > > paths).
>>> > > > > > > > >   - Completion of In-Flight Runs: is_paused remains false
>>> > > during
>>> > > > > the
>>> > > > > > > > drain,
>>> > > > > > > > > meaning task instances in existing runs are still
>>> scheduled
>>> > and
>>> > > > > > finish
>>> > > > > > > > > normally.
>>> > > > > > > > >   - Automatic Convergence: Once no unfinished runs
>>> remain,
>>> > the
>>> > > > > > > scheduler
>>> > > > > > > > > automatically moves the DAG to the paused state and
>>> writes a
>>> > > > > > > > > drain_completed audit log entry.
>>> > > > > > > > >
>>> > > > > > > > > The core property of this feature is that draining is
>>> > > transient,
>>> > > > > not
>>> > > > > > a
>>> > > > > > > > > third resting state; it always converges to paused.
>>> > > > > > > > >
>>> > > > > > > > > Implementation Details
>>> > > > > > > > >
>>> > > > > > > > > Explicit run creation via manual triggers,
>>> > > TriggerDagRunOperator,
>>> > > > > > asset
>>> > > > > > > > > materialization, or backfills remains allowed during
>>> > draining,
>>> > > > > > > mirroring
>>> > > > > > > > > the behavior of a paused DAG. An earlier revision that
>>> > blocked
>>> > > > > these
>>> > > > > > > > > actions was reverted to ensure draining is not stricter
>>> than
>>> > > the
>>> > > > > > state
>>> > > > > > > it
>>> > > > > > > > > converges into.
>>> > > > > > > > >
>>> > > > > > > > > Links:
>>> > > > > > > > >
>>> > > > > > > > >   - PR: https://github.com/apache/airflow/pull/72407
>>> > > > > > > > >   - Issue:
>>> https://github.com/apache/airflow/issues/22006
>>> > > > > > > > >
>>> > > > > > > > > Thanks,
>>> > > > > > > > > Dheeraj
>>> > > > > > > > >
>>> > > > > > > >
>>> > > > > > >
>>> > > > > >
>>> > > > >
>>> > > >
>>> > >
>>> >
>>>
>>

Reply via email to