Hello Ulada,

First of all - there is something wrong with delivery of this message. While
I (and few other people) see it in the mailing list interface [1] it
apparently has not been delivered to at least some of us.

I will follow-up with ASF Infrastructure on that.

But let me respond to the merit in-line :) 

[1] https://lists.apache.org/thread/7z1q39d8fq6ny9jj6omxkrk9o1mogmk7

On 2026-10-05 13:46 UTC Ulada Zakharava (xWF) wrote:
> Hi everyone,
> 
> We’ve noticed an increasing number of PRs submitted to the Google Provider
> where contributors—often using AI tooling—don't have access to a GCP
> project to test their changes.

Yeah. This is something we see as well. The good thing about it is that
those same AI tools actually can run system tests when they are asked to.
They will need to have a GCP account for it - but I personally closed
a few PRs earlier where the users responded that they will not test it,
because they do not have an account. I think we should automate that
(and we can - and we will even be able to do more of that when you
have our Airflow instance with mainteance Dags that Shubham works on
prototype of - where a lot of those maintenance rules will be easier
to manage, run and tap into credited / ASF tokens to do some of the
analysis - on top of the "deterministic" checks we will do.

> While we love community contributions, many of these PRs touch critical
> logic like authentication, base hooks, or core operator behavior with only
> mocked unit tests. Maintainers are regularly asked to pull these branches
> and run manual or system tests to see if they actually work against live
> APIs.
> 
> Simply put, this doesn't scale. Pulling branches to run ad-hoc system tests
> takes significant time, and foundational changes often require running
> multiple full suites. Merging without verification isn't an option either,
> as catching regressions during release candidate testing blocks releases
> and creates painful reverts.

Absolutely - and we feeel it in other places as well - and we need to definitely
add even more harness to do those kind of things and automate it relentlessly - 
part of the problem is also scale of the number of PRs we receive - and the 
overall
growth of those "filtering" that we have to do. I personally think (and I am
step-by-step working on it, and welcome any incremental contributions
here) - where we should **relentlessly** automate and add more harness for
things like that. 

There are several things we can do about it:

* Rules that we agree on - those are easiest to do, but also least dependable -
and this is the same "scale" problem in reverse - we simply have so many rules
that we would have to follow that it would be insane to expect all maintainers
to remember and follow them - especially when there are people who are merging
and reviewing things "manually".

* AGENTS.md describing the rules and expectations. This is IMHO a very efficient
way to target big percentage of the issues (and I am going to add a few PRs as
a follow-up to this one. We can add expectations for the agents to **always**
do system test and to **refuse** to submit PR when the system tests cannot be
run - and direct people to contribute in other areas. Not 100% harness, because
human (and other agentic instructions) can override it - but it should work in
vast majority of agentic cases. Also if the maintainers (which they increasingly
do) use agents to review and comments on the PRs - those same AGENTS.md 
instructions
will be read and adversarial reviews on the PR will catch that there are no 
proofs
(logs/screenshots) of running the system tests. This will give almost 100% catch
rate for those who use good agents for review (I do it now for **ALL** PRs I 
merge)

* deterministic/Agentic checks which will automate triage for all our PRs - 
while
GitHub CI already handles LOTS of issues and they are fixed by the agents
automatically, we cannot (for now) afford running agentic / AI checks
for every single PR separately - because this means that usage of tokens can be
used/abused by other agents submitting the PR. This is where the idea of 
continuously
running Airflow that will run scheduled triage and will take appropriate actions
(with HITL) that Shubham is working on, is going to be really useful - this is 
where
we can bulk-check 100s of PRs at a time - much more efficiently, with more 
control
and where we will be able to iterate and experiment with various rules - same as
I manually did with Magpie triage.

This is something we should work as part of the "improve our processes in the 
world
of AI" I will be workign on in the coming months - together with others - 
Shahar,
Shubham, and other people who are now very active in the CI work would surely
join the efforts and we will work out some good ways of managing it - I hope so 
that
we can free maintainer time to not have to remember and even think about those
rules, but to enforce them automatically (but with proper overlook).

> Proposed Rule
> 
> We want to agree on a clear baseline for PRs touching functional or auth
> logic in the Google Provider:
> 
>    1.
> 
>    *Proof of testing required:* Authors must test against a real GCP
>    project and include proof (logs, CLI output, or run results) in the PR
>    description.

Yep. This should be hard requirement. Let's encode it in the AGENTS.md now -\
I will open PRs later today. And those rules will be applied not only to Google
- but for **all** providers that have system tests. With the growing number
of providers we have - it will not scale if we do not have the same rule
for all provider "stewards". And since we already have common framework for
System Tests this could be a common rule in "providers/AGENTS.md" for all
providers that have system tests.

>    2.
> 
>    *Untested PRs will be put on hold:* Changes affecting runtime behavior
>    without evidence of testing shouldn't be reviewed or merged until 
> validated.

Yeah. For now, they will not pass agentic review - which should be **good 
enough*,

>    3.
> 
>    *Keep system tests up to date:* New operators or behavior changes should
>    include or update automated system tests, or at minimum show manual
>    end-to-end run results.

This we can actually even deterministically implement as a prek hook - again, 
same
for all providers that have system tests. I am all for it - and maybe someone
who is also taking care about CI could implement it - we have several similar 
prek
hooks where we expect unit tests to be written and we keep ever-shrinking list 
of
exceptions, this could be done in a very similar way.


> 
> What we're doing on our end
> 
> We understand that not everyone has an active GCP environment. On the
> Google/maintainer side, we’re looking into ways to make it easier to
> trigger specific existing system tests directly on PR branches via CI for
> trusted contributors. Until that workflow is ready, the responsibility of
> proving a change works has to sit with the author.

I can help with that I think. Recently I've been trying an approach in Magpie
where maintainer triggers additional things when PR
of certain kind is created - by commenting on the PR. For example what it helps
in Magpie is to create preview of docs build - so that you can see them hosted
`prNNNN.magpie.staged.apache.org` (I have PR opened to infrastructure-actions
to make this action "ASF-wide" 
https://github.com/apache/infrastructure-actions/pull/1332
It has this nice feature that you can even select part of the website
and it will take the screenshot and open PR in the place from where the
part of website came so you can paste-in the comment (see the PR examples).
I planned to add it to Airflow, and we can do very similar thing for
system testing. All we will need to do is have a google account that we can use
and then maintainer will be able to comment "/run-system-tests" on the PR - and
the system tests will be run there. We can fully automate that - and when it is
maintainer controlled, it will mitigate abuse capabilities.

Such setting can be "sticky" - this is what I've done in the website action in
Magpie - once maintainer commments on the PR, every new fix /rebase/push will
trigger such system tests. 

This is ultimately that we will have to implement with all stakeholders of
ours - I will lead that effort in the coming months.

> 
> What do you all think? We'd love to hear feedback before updating the
> provider contribution docs.

We hear and understand you :) .. And we will work together to make it better.

> 
> Best regards,
> 
> Ulada
> 

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to