Sending this again, as I just realized my response did not go through! Hi Dheeraj,
I think draining is a useful feature! I don’t have a strong opinion on whether or not draining should be its own state (having it be a transient state seems reasonable, and it seems like you have consensus). What I am questioning is the need to have this feature be a per dag option. The Airflow UI is already really busy for individual dags, and having a modal pop up to ask whether you really mean drain or pause will be really annoying, especially for operators who will need to do this for every active dag at scale. If there are 200 dags, that’s 400 clicks. Since upgrade prep is the primary use-case, why not start by making this a bulk feature? This only fell on my radar because of the dev list conversation, but I’m happy to take this to the PR or start a new email chain if you prefer. Constance On Fri, Sep 11, 2026 at 1:06 PM Dheeraj Turaga <[email protected]> wrote: > Hi Constance, > > As someone who operates both as a Dag Author and Airflow Instance manager, > I see having the ability to drain useful. > > As a dag author, there are many cases where I want existing long running > scheduled dag runs > to finish before I disable the pipeline before new runs kick in. > I also in some cases may want to pause or "halt" the dag runs immediately. > This current pr addresses this particular usecase > > As a Airflow Infra manager, I would like to drain the instance before a > planned maintenance > and want to give every dag the opportunity to finish before I do so. > The current pr does not address this bulk action/cli action (YET). > I plan to create follow up prs to address the scenario > > In any case, both scenarios would require a general agreement on the need > for a > transient "drain" state which Is what im hoping to get consensus on with > this thread. > > Thanks, > Dheeraj > > On Thu, Sep 10, 2026 at 7:28 PM Constance Martineau via dev < > [email protected]> wrote: > >> Hi all, >> >> Jumping in with a related question. Skimming the original issue, the >> motivating case is explicitly instance-wide: "we don't want to fail/kill >> running DAGs... Once none of the DAGs are in a running state, we perform >> the upgrades... Once upgrades are done, we update the DAG schedules back >> to >> their original value". That reads as an operator preparing for a >> maintenance window, not a Dag owner reaching for a per-Dag control. >> >> I have two UX concerns, from each side of that use case. The design >> reserves the drain-or-pause choice for the moment a Dag has runs >> in-flight, >> based on the logic that this is when the choice actually matters. However, >> for a Dag owner, having runs in flight isn't the exception, it's precisely >> the state a Dag is in when I need to pause it, because something running >> is >> going wrong. That's exactly the moment I want pause to still mean "stop >> it," not a fork to resolve under pressure. And as an operator preparing >> for >> an upgrade, having to toggle and confirm this per Dag, one at a time, >> across however many active Dags there are, is close to the same manual, >> error-prone process this is meant to replace. >> >> Would a bulk action cover the use-case better? Like, drain and pause >> everything currently active in one click, and resume everything that was >> active in one click afterward? That directly matches the upgrade scenario >> without adding a new decision to the single-Dag pause flow. A per-Dag >> version only justifies its complexity if there's a real need to gracefully >> drain one specific Dag independent of an upgrade. I haven't seen that use >> case motivated anywhere in this thread, every example so far points to >> draining everything before touching the cluster. >> >> Curious what others think, happy to be wrong here. >> >> Thanks, Constance >> >> On Thu, Sep 10, 2026 at 12:20 PM Shubham Raj <[email protected]> >> wrote: >> >> > Hey, >> > >> > I agree with the direction of the PR. Given that is_paused serves >> several >> > critical functions, overloading it to include draining behavior would >> > likely cause unintended side effects across the system. By keeping >> > "draining" as a distinct, transient state, we can avoid those >> complications >> > while effectively solving the need for graceful task completion. This >> will >> > also significant improvement for managing Dags during upgrades. >> > >> > >> > On Wed, 9 Sep 2026 at 20:15, Elad Kalif <[email protected]> wrote: >> > >> > > > The drain here is a transient which eventually converges to pause. >> > > >> > > I agree. drain is intermediate state. When draining is completed the >> > status >> > > should be changed to paused. >> > > >> > > On Wed, Sep 9, 2026 at 1:49 PM Jarek Potiuk <[email protected]> wrote: >> > > >> > > > > The drain here is a transient which eventually converges to pause. >> > > > >> > > > Yep. That was my thinking. You summarized it all in one sentence :) >> > > > >> > > > On Wed, Sep 9, 2026 at 7:03 AM Dheeraj Turaga < >> [email protected] >> > > >> > > > wrote: >> > > > >> > > > > Hi Amogh, >> > > > > >> > > > > On Jarek's thought, I agree with him that just changing the pause >> > > > behavior >> > > > > is probably not a good idea. >> > > > > I can see instances where a user may want to instantly pause a >> dag in >> > > one >> > > > > case >> > > > > but want to drain a dag in another case. I myself have such a need >> > for >> > > > > different dags. >> > > > > >> > > > > For example, during airflow upgrades, I would like to gracefully >> > drain >> > > my >> > > > > whole instance before upgrade. >> > > > > While during regular operation, I may want to prevent my dag from >> > > > launching >> > > > > further tasks. >> > > > > >> > > > > The drain here is a transient which eventually converges to pause. >> > > > > >> > > > > Draining keeps task scheduling enabled, and only automated dagrun >> > > > creation >> > > > > paths are gated. >> > > > > *is_paused* is set to true only when the dagrun are completed. >> Manual >> > > > > triggers are still allowed >> > > > > simialr to the behavior when *is_paused* set to false. >> > > > > >> > > > > It would be great to get your thoughts on the PR once you get a >> > chance >> > > to >> > > > > play with it. >> > > > > >> > > > > Dheeraj >> > > > > >> > > > > On Tue, Sep 8, 2026 at 11:19 PM Amogh Desai < >> [email protected]> >> > > > wrote: >> > > > > >> > > > > > Yep, *is_paused* does more than block new runs. The scheduler >> > checks >> > > it >> > > > > > directly when deciding which >> > > > > > tasks to run next. That is what freezes downstream tasks today, >> and >> > > it >> > > > is >> > > > > > also why I would push back on >> > > > > > "same is_paused but two buttons." >> > > > > > >> > > > > > If you are suggesting that "Pause and drain" sets >> *is_paused=True* >> > > > right >> > > > > > away, those two checks still need a way >> > > > > > to know "this one is a soft pause, let it keep going." That >> means >> > > > adding >> > > > > a >> > > > > > second flag anyway, just hidden behind >> > > > > > the same name. And *is_paused* shows up in other places like >> > backfill >> > > > > pause >> > > > > > logic and asset triggered runs. Each of >> > > > > > those would need checking to make sure a "soft paused" Dag still >> > acts >> > > > > like >> > > > > > an active one everywhere except the one >> > > > > > spot it should not. From the PR, it avoids all that by keeping >> > > > > > *is_paused=False* during drain, so every existing check >> > > > > > keeps working as it should. It only adds new checks in a few >> places >> > > > that >> > > > > > decide whether to start a new run. >> > > > > > >> > > > > > Thanks & Regards, >> > > > > > Amogh Desai >> > > > > > >> > > > > > >> > > > > > On Tue, Sep 8, 2026 at 4:29 PM Jarek Potiuk <[email protected]> >> > > wrote: >> > > > > > >> > > > > > > I like the idea Amogh. I think **just changing** pause >> behaviour >> > > > might >> > > > > > not >> > > > > > > be good idea for compatibility (someone might rely on it's >> > current >> > > > > > > behaviour). >> > > > > > > >> > > > > > > But maybe we could have different behaviour of two >> > > > > > > different buttons/behaviour: "Pause" (current behaviour - stop >> > > > current >> > > > > > > tasks and it's downstream), "Pause and drain" - set the Dag to >> > > > "paused" >> > > > > > but >> > > > > > > let the running dag_runs to continue until completion. >> > > > > > > >> > > > > > > It might however require some scheduler changes - I think >> > currently >> > > > > > > is_paused is used to select run-eligible tasks ? >> > > > > > > >> > > > > > > J. >> > > > > > > >> > > > > > > >> > > > > > > On Tue, Sep 8, 2026 at 7:41 AM Amogh Desai < >> > [email protected]> >> > > > > > wrote: >> > > > > > > >> > > > > > > > Thanks for picking a 4-year-old issue. >> > > > > > > > The ask itself makes sense. >> > > > > > > > >> > > > > > > > One design question that might eventually come up in this >> > thread >> > > / >> > > > PR >> > > > > > is: >> > > > > > > > why do we need a new persisted state >> > > > > > > > instead of just changing what is_paused does? i.e: keep >> gating >> > > new >> > > > > run >> > > > > > > > creation on is_paused, but stop freezing >> > > > > > > > downstream tasks in created runs. That would get you >> graceful >> > > > > draining >> > > > > > > > without a schema change or migration. >> > > > > > > > >> > > > > > > > From my reading, doing that would silently change the >> semantics >> > > of >> > > > > > > > *is_paused*. Today, freezes downstream tasks in-flight. >> > > > > > > > Some operators pause specifically to stop a run's progress, >> not >> > > > just >> > > > > to >> > > > > > > > block new ones. Folding drain behavior into *is_paused* >> > > > > > > > might change that for everyone already relying on the >> current >> > > > freeze >> > > > > + >> > > > > > > > pause behavior, with no option to opt out, yes? >> > > > > > > > Worth clarifying. >> > > > > > > > >> > > > > > > > Thanks & Regards, >> > > > > > > > Amogh Desai >> > > > > > > > >> > > > > > > > >> > > > > > > > On Sun, Sep 6, 2026 at 3:15 AM Dheeraj Turaga < >> > > > > [email protected] >> > > > > > > >> > > > > > > > wrote: >> > > > > > > > >> > > > > > > > > Hey Everyone, >> > > > > > > > > >> > > > > > > > > I wanted to start this discussion to get feedback on PR >> > #72407, >> > > > > which >> > > > > > > > > proposes adding a "draining" scheduling state to Airflow. >> > > > > > > > > >> > > > > > > > > The Problem >> > > > > > > > > >> > > > > > > > > Currently, pausing a DAG stops it mid-flight. While >> running >> > > tasks >> > > > > are >> > > > > > > > > allowed to finish, queued and downstream tasks remain >> > stranded >> > > > > until >> > > > > > > the >> > > > > > > > > DAG is unpaused. There is currently no native way to stop >> > > > starting >> > > > > > new >> > > > > > > > runs >> > > > > > > > > while allowing those already in flight to complete. >> > > > > > > > > >> > > > > > > > > This is a significant pain point during upgrades and >> > > maintenance. >> > > > > The >> > > > > > > > > current workaround as described in the four-year-old Issue >> > > #22006 >> > > > > > > > requires >> > > > > > > > > users to manually rewrite every DAG schedule to None, wait >> > for >> > > > runs >> > > > > > to >> > > > > > > > > drain, and then restore the schedules afterward. This >> process >> > > is >> > > > > > > invasive >> > > > > > > > > and prone to error. >> > > > > > > > > >> > > > > > > > > Proposed Solution: The draining State >> > > > > > > > > >> > > > > > > > > The PR introduces a draining state that sits between >> active >> > and >> > > > > > paused. >> > > > > > > > Key >> > > > > > > > > behaviors include: >> > > > > > > > > >> > > > > > > > > - No New Scheduled Runs: The scheduler creates no new >> runs >> > > for >> > > > a >> > > > > > > > draining >> > > > > > > > > DAG (including scheduled, asset-triggered, and >> > > partitioned/rollup >> > > > > > > paths). >> > > > > > > > > - Completion of In-Flight Runs: is_paused remains false >> > > during >> > > > > the >> > > > > > > > drain, >> > > > > > > > > meaning task instances in existing runs are still >> scheduled >> > and >> > > > > > finish >> > > > > > > > > normally. >> > > > > > > > > - Automatic Convergence: Once no unfinished runs remain, >> > the >> > > > > > > scheduler >> > > > > > > > > automatically moves the DAG to the paused state and >> writes a >> > > > > > > > > drain_completed audit log entry. >> > > > > > > > > >> > > > > > > > > The core property of this feature is that draining is >> > > transient, >> > > > > not >> > > > > > a >> > > > > > > > > third resting state; it always converges to paused. >> > > > > > > > > >> > > > > > > > > Implementation Details >> > > > > > > > > >> > > > > > > > > Explicit run creation via manual triggers, >> > > TriggerDagRunOperator, >> > > > > > asset >> > > > > > > > > materialization, or backfills remains allowed during >> > draining, >> > > > > > > mirroring >> > > > > > > > > the behavior of a paused DAG. An earlier revision that >> > blocked >> > > > > these >> > > > > > > > > actions was reverted to ensure draining is not stricter >> than >> > > the >> > > > > > state >> > > > > > > it >> > > > > > > > > converges into. >> > > > > > > > > >> > > > > > > > > Links: >> > > > > > > > > >> > > > > > > > > - PR: https://github.com/apache/airflow/pull/72407 >> > > > > > > > > - Issue: https://github.com/apache/airflow/issues/22006 >> > > > > > > > > >> > > > > > > > > Thanks, >> > > > > > > > > Dheeraj >> > > > > > > > > >> > > > > > > > >> > > > > > > >> > > > > > >> > > > > >> > > > >> > > >> > >> >
