ashb opened a new pull request, #74222:
URL: https://github.com/apache/airflow/pull/74222

   ## Note to reviewers
   
   This PR desc is still in progress of being drafted, so it might not fully 
make sense yet. Bear with me please.
   
   Additionally, the commit count looks scar, yes, but _most_ of it is fairly 
mechanical adding of the workingset condition.
   
   ## Description
   
   This is a data-model change precursor for the Task Loop implementation.
   
   The subject sounds simple, but unfortunately this change is far more 
complicated
   than I'd like, but for a couple of good reasons.
   
   First though, why do we want this: to support Task Loops (AIP-111) which lets
   the number of loops of a task group be decided, and created just-in-time at
   runtime. All  good so far, and no changes needed far.
   
   However things change when you get to clearing TIs. Task Loops are most 
useful
   when the output of one loop feeds in to the next. So the only sensible 
default
   is that when you clear mid-way through a loop, say you clear iterations 2, 
3, 4,
   and 5, we restart and run 2' onwards. Since it is a runtime decision is is no
   guarantee how many it will run, and most crucially, we don't want to make the
   previous TIs invisible.
   
   Visually, this looks like
   
   <img width="1672" height="484" alt="image" 
src="https://github.com/user-attachments/assets/cbbf3ea6-499b-42fe-8cd8-955c3f5310ff";
 />
   
   (where each box there is a set of tasks from a looped task group for AIP-111)
   
   We could choose to just keep the retired/superseded TI in the existing
   task_instance_history and delete/loose the XCom and other data. But that 
struck
   me as very sub-optimal, and I'm using this to also fix a long term pain 
point:
   that having TI and TIHistory tables separate causes some other odd quirks.
   
   So why the `is_workingset` column? This is to keep the scheduler performant. 
The
   current main reason to have TI and TIHistory separate is so that the 
scheduler
   can 
   
   Now to the gnarliest point of this change: xcom_v2 et al. Why do this? Simply
   put, it's because of the time to migrate the data. With the new FK on TI, we
   would need to update _every single tuple_ on XCom table. I ran some tests on 
a
   copy of a production database with 309 million XCom rows on how long it would
   have taken to migrate XCom to store the creating task_instance_id on it: I 
gave
   up and cancelled it after 14.5 hours. 309M is larger that most people will 
have,
   but more than 14.5 hours is simply not worth it.
   
   
   So instead I've got for the "generational table" approach, where we write to 
a
   new table, and read from both. It's slightly more complicate than that due 
to a
   few shadowing cases etc, but that's the gist of the approach.
   
   <img width="1674" height="700" alt="image (1)" 
src="https://github.com/user-attachments/assets/018d0611-8bef-48d0-8467-8568d7832622";
 />
   
   
   Now, I hear you asking: "Doesn't this mean we have to carry this legacy code
   path until Airflow 4 at an unknown time in the future?" Nope, I've got this
   covered, just not included in this PR. In a future PR (likely landing for 
3.4.x
   or 3.5.0) I plan on having the scheduler "slowly" (but not too slowly) 
migrate
   chunks of the data from the v1 table to v2, and then in 3.6 we add a 
migration
   that does the final "move all data form xcom_v1 to xcom_v2" and document "You
   need to upgrade to 3.4.x/3.5.0, and wait for X time, before upgrading to 3.6,
   else expect a very slow migration" (some status in the admin screens to be 
added
   to make this visible.)
   
   [ed: Still to add:]
   
   - Why "NOT VALID" (faster, still enforced for new writes.)
   - why "legacy-owner" tableĀ (to maintain FK cascaded)
   
   <img width="1668" height="882" alt="image (2)" 
src="https://github.com/user-attachments/assets/164861a5-cf6b-44f4-a850-1eadcbc06203";
 />
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to