Hi Shahar, Very good points. All of them are valid and worth deeper thought and discussion.
Resource termination and cost information are indeed crucial. We should have strong gatekeeping there, and being able to run tests "as needed" by a maintainer or steward is likely best, as that allows for enforced teardowns of resources. "Required" is an interesting concept here. While it might seem like a "pay to play" model, relying solely on unit tests when contributing to features that interact with live services reflects a lack of care. Most cloud providers, including Google, offer free tiers for testing. When someone claims they cannot test with real services, it is rarely a money issue; it is usually a refusal to invest the time to set up a free account and learn how system tests work. Time and attention have always been our currency. We have never judged contributions by lines of code or the number of PRs, but by the effort someone invests in the community. This gets to the root of our current issues with AI-generated low-effort contributions. It is frustrating when contributors expect maintainers to spend time reviewing work that they themselves spent no time testing. Claiming an inability to test with real services often just means a lack of desire to put in the necessary effort. In those cases, we should firmly close the PR without hesitation. We should be both helpful and firm: - Helpful: Guide contributors on how to run system tests, utilize free tiers, and potentially access project-provided credits (e.g., from Google or Amazon). We should provide clear documentation, warnings about resource teardowns, and helpful tooling. - Firm: Clearly communicate our expectations. If contributors are unwilling to follow instructions or set up free accounts to test with real services, they should contribute elsewhere. is it 'pay to play' - yes. With time and attention. not money. This is the essence of volunteer conteibutions and Open Source movement. It is risky to trust untested code for external services. AI agents make confident mistakes regarding API parameters and flows. Without real service testing, there is no proof the code works, or if the underlying issue was even real to begin with. This standard is not new—it has always been our implicit expectation. We would not accept production code written purely from API docs without validation, nor would we write code that way ourselves. With modern agentic workflows, the manual effort for proper engineering practices—such as reproducing issues, running local test environments like Breeze, executing live DAGs, and capturing logs or screenshots—can now be automated. We can encode these standards into our project guidelines (e.g., AGENTS.md) so that automated tools perform and verify these steps. Ultimately, we must hold contributors to this standard. We expect the quality of incoming contributions to match or exceed what maintainers produce using their own workflows. We should not accept lower quality simply because it is easier to submit. For contributors who prefer not to use live services, there are plenty of areas in Airflow that can be tested sufficiently with unit and integration tests. Best regards, Jarek On Wed, Oct 7, 2026, 11:57 Shahar Epstein <[email protected]> wrote: > Thanks, Ulada, for your feedback, and Jarek for the suggested solution. > > To put things in context, the trigger for this discussion is this PR #74305 > <https://github.com/apache/airflow/pull/74305>. I merged it after > reviewing > it using Magpie, but without running the system tests. The existing system > test did not catch the issue because its queries completed before the > trigger’s first poll. I’ve raised PR #74356 > <https://github.com/apache/airflow/pull/74356> to cover the long-running > case in the system tests, which I tested on my personal GCP project. > > To the point: > > Considering that the system tests require a GCP account and can incur > charges, I don’t think requiring contributors to run them on their own for > every functional change is optimal for contributors or maintainers: > - Contributors outside the Google provider team may be gatekept because > they don’t have the resources to run these tests (please refer to the “pay > to play > <https://lists.apache.org/thread.html/b9wmh7nmpy0dv73mgtl8n3x87n7wv7v0>” > discussion we had about running CI on forks). > - For maintainers, this creates a bottleneck while we wait for the system > tests to be run by contributors or the Google team. > > If the Google team wants this enforced for every functional or auth-related > PR, I’m asking for the team's cooperation in integrating the tests into CI, > so maintainers can trigger the relevant system tests on the GCP > infrastructure from contributors' branches. > > Moreover, Jarek, if we add AGENTS.md instructions requiring system tests > for any provider whose tests use paid services: > 1. The agent must first estimate the cost and ask the user for explicit > approval before running tests if the estimated cost is more than $0. > 2. The agent must ensure that cloud resources created by the system tests > are terminated. > > > Shahar > > On Wed, Oct 7, 2026 at 11:10 AM Jarek Potiuk <[email protected]> wrote: > > > Hello Ulada, > > > > First of all - there is something wrong with delivery of this message. > > While > > I (and few other people) see it in the mailing list interface [1] it > > apparently has not been delivered to at least some of us. > > > > I will follow-up with ASF Infrastructure on that. > > > > But let me respond to the merit in-line :) > > > > [1] https://lists.apache.org/thread/7z1q39d8fq6ny9jj6omxkrk9o1mogmk7 > > > > On 2026-10-05 13:46 UTC Ulada Zakharava (xWF) wrote: > > > Hi everyone, > > > > > > We’ve noticed an increasing number of PRs submitted to the Google > > Provider > > > where contributors—often using AI tooling—don't have access to a GCP > > > project to test their changes. > > > > Yeah. This is something we see as well. The good thing about it is that > > those same AI tools actually can run system tests when they are asked to. > > They will need to have a GCP account for it - but I personally closed > > a few PRs earlier where the users responded that they will not test it, > > because they do not have an account. I think we should automate that > > (and we can - and we will even be able to do more of that when you > > have our Airflow instance with mainteance Dags that Shubham works on > > prototype of - where a lot of those maintenance rules will be easier > > to manage, run and tap into credited / ASF tokens to do some of the > > analysis - on top of the "deterministic" checks we will do. > > > > > While we love community contributions, many of these PRs touch critical > > > logic like authentication, base hooks, or core operator behavior with > > only > > > mocked unit tests. Maintainers are regularly asked to pull these > branches > > > and run manual or system tests to see if they actually work against > live > > > APIs. > > > > > > Simply put, this doesn't scale. Pulling branches to run ad-hoc system > > tests > > > takes significant time, and foundational changes often require running > > > multiple full suites. Merging without verification isn't an option > > either, > > > as catching regressions during release candidate testing blocks > releases > > > and creates painful reverts. > > > > Absolutely - and we feeel it in other places as well - and we need to > > definitely > > add even more harness to do those kind of things and automate it > > relentlessly - > > part of the problem is also scale of the number of PRs we receive - and > > the overall > > growth of those "filtering" that we have to do. I personally think (and I > > am > > step-by-step working on it, and welcome any incremental contributions > > here) - where we should **relentlessly** automate and add more harness > for > > things like that. > > > > There are several things we can do about it: > > > > * Rules that we agree on - those are easiest to do, but also least > > dependable - > > and this is the same "scale" problem in reverse - we simply have so many > > rules > > that we would have to follow that it would be insane to expect all > > maintainers > > to remember and follow them - especially when there are people who are > > merging > > and reviewing things "manually". > > > > * AGENTS.md describing the rules and expectations. This is IMHO a very > > efficient > > way to target big percentage of the issues (and I am going to add a few > > PRs as > > a follow-up to this one. We can add expectations for the agents to > > **always** > > do system test and to **refuse** to submit PR when the system tests > cannot > > be > > run - and direct people to contribute in other areas. Not 100% harness, > > because > > human (and other agentic instructions) can override it - but it should > > work in > > vast majority of agentic cases. Also if the maintainers (which they > > increasingly > > do) use agents to review and comments on the PRs - those same AGENTS.md > > instructions > > will be read and adversarial reviews on the PR will catch that there are > > no proofs > > (logs/screenshots) of running the system tests. This will give almost > 100% > > catch > > rate for those who use good agents for review (I do it now for **ALL** > PRs > > I merge) > > > > * deterministic/Agentic checks which will automate triage for all our PRs > > - while > > GitHub CI already handles LOTS of issues and they are fixed by the agents > > automatically, we cannot (for now) afford running agentic / AI checks > > for every single PR separately - because this means that usage of tokens > > can be > > used/abused by other agents submitting the PR. This is where the idea of > > continuously > > running Airflow that will run scheduled triage and will take appropriate > > actions > > (with HITL) that Shubham is working on, is going to be really useful - > > this is where > > we can bulk-check 100s of PRs at a time - much more efficiently, with > more > > control > > and where we will be able to iterate and experiment with various rules - > > same as > > I manually did with Magpie triage. > > > > This is something we should work as part of the "improve our processes in > > the world > > of AI" I will be workign on in the coming months - together with others - > > Shahar, > > Shubham, and other people who are now very active in the CI work would > > surely > > join the efforts and we will work out some good ways of managing it - I > > hope so that > > we can free maintainer time to not have to remember and even think about > > those > > rules, but to enforce them automatically (but with proper overlook). > > > > > Proposed Rule > > > > > > We want to agree on a clear baseline for PRs touching functional or > auth > > > logic in the Google Provider: > > > > > > 1. > > > > > > *Proof of testing required:* Authors must test against a real GCP > > > project and include proof (logs, CLI output, or run results) in the > PR > > > description. > > > > Yep. This should be hard requirement. Let's encode it in the AGENTS.md > now > > -\ > > I will open PRs later today. And those rules will be applied not only to > > Google > > - but for **all** providers that have system tests. With the growing > number > > of providers we have - it will not scale if we do not have the same rule > > for all provider "stewards". And since we already have common framework > for > > System Tests this could be a common rule in "providers/AGENTS.md" for all > > providers that have system tests. > > > > > 2. > > > > > > *Untested PRs will be put on hold:* Changes affecting runtime > behavior > > > without evidence of testing shouldn't be reviewed or merged until > > validated. > > > > Yeah. For now, they will not pass agentic review - which should be **good > > enough*, > > > > > 3. > > > > > > *Keep system tests up to date:* New operators or behavior changes > > should > > > include or update automated system tests, or at minimum show manual > > > end-to-end run results. > > > > This we can actually even deterministically implement as a prek hook - > > again, same > > for all providers that have system tests. I am all for it - and maybe > > someone > > who is also taking care about CI could implement it - we have several > > similar prek > > hooks where we expect unit tests to be written and we keep ever-shrinking > > list of > > exceptions, this could be done in a very similar way. > > > > > > > > > > What we're doing on our end > > > > > > We understand that not everyone has an active GCP environment. On the > > > Google/maintainer side, we’re looking into ways to make it easier to > > > trigger specific existing system tests directly on PR branches via CI > for > > > trusted contributors. Until that workflow is ready, the responsibility > of > > > proving a change works has to sit with the author. > > > > I can help with that I think. Recently I've been trying an approach in > > Magpie > > where maintainer triggers additional things when PR > > of certain kind is created - by commenting on the PR. For example what it > > helps > > in Magpie is to create preview of docs build - so that you can see them > > hosted > > `prNNNN.magpie.staged.apache.org` (I have PR opened to > > infrastructure-actions > > to make this action "ASF-wide" > > https://github.com/apache/infrastructure-actions/pull/1332 > > It has this nice feature that you can even select part of the website > > and it will take the screenshot and open PR in the place from where the > > part of website came so you can paste-in the comment (see the PR > examples). > > I planned to add it to Airflow, and we can do very similar thing for > > system testing. All we will need to do is have a google account that we > > can use > > and then maintainer will be able to comment "/run-system-tests" on the PR > > - and > > the system tests will be run there. We can fully automate that - and when > > it is > > maintainer controlled, it will mitigate abuse capabilities. > > > > Such setting can be "sticky" - this is what I've done in the website > > action in > > Magpie - once maintainer commments on the PR, every new fix /rebase/push > > will > > trigger such system tests. > > > > This is ultimately that we will have to implement with all stakeholders > of > > ours - I will lead that effort in the coming months. > > > > > > > > What do you all think? We'd love to hear feedback before updating the > > > provider contribution docs. > > > > We hear and understand you :) .. And we will work together to make it > > better. > > > > > > > > Best regards, > > > > > > Ulada > > > > > > > --------------------------------------------------------------------- > > To unsubscribe, e-mail: [email protected] > > For additional commands, e-mail: [email protected] > > > > >
