zozo123 opened a new pull request, #74055:
URL: https://github.com/apache/airflow/pull/74055

   On pull requests, a change to any Python file under `devel-common/`, or to 
anything in
   `airflow-core/tests/unit/utils/`, forces the full test matrix 
(`TESTS_UTILS_FILES`). On top of that,
   `find_provider_affected()` treats every `devel-common/` file as common 
provider code and selects all
   providers. Most of those files do not need that:
   
   - **`devel-common/src/tests_common/` helpers** are mostly imported by a 
handful of test files. Of
     the 62 helpers, 18 are loaded by every test run (the pytest plugin, 
anything it imports,
     conftest and package `__init__` modules). The other 44 are not. 
`mock_context.py`, for example,
     reaches 9 test files in 3 providers (through `run_deferrable.py`).
   - **`airflow-core/tests/unit/utils/`** holds only `test_*.py` files. Its few 
cross-imports are from
     other core tests, never from providers.
   - **`devel-common/src/sphinx_exts/`** only affects the docs build.
   
   With this change, a changed `tests_common` helper selects the test files 
that import it, directly or
   through other helpers, as if those files had changed. Only helpers that 
every test run loads still
   force the full matrix. The `airflow-core` unit-test directory is mapped like 
any other core test
   files, and Sphinx extensions trigger the docs build (`DOC_FILES`).
   
   ### What it saves, measured
   
   I replayed `SelectiveChecks` on the exact inputs (changed files, labels, 
commit) recorded in the
   *Build info* logs of real `ci-amd.yml` PR runs, once with `main` and once 
with this branch. For each
   run whose outputs changed, I estimated the new cost from that run's actual 
runner time, scaled by
   the median runner time of green PR runs with the new versus the old 
selective-check outcome.
   
   | Sample of PR runs | Runs whose matrix shrinks | From PRs | Runner time 
saved | Runs that got bigger |
   |---|---|---|---|---|
   | 28 Sep – 1 Oct (819 runs, 3.6 days) | 86 (28 green, 58 failed/cancelled) | 
16 | 269 runner-hours/day | 0 |
   | 14 – 18 Sep (694 runs over 60 jobs, 4.3 days) | 37 (10 green, 27 
failed/cancelled) | 4 | 106 runner-hours/day | 0 |
   
   Every changed run goes from `full-tests-needed=true` to a selective run, 
typically "all core test
   types plus the affected providers and their direct dependents". Those cost a 
median 12.5 h on green
   runs, versus 24.1 h for a full run. PRs that edit test helpers tend to be 
pushed many times, which is
   why failed and cancelled runs carry most of the saving.
   
   ### Checks run
   
   - `dev/breeze/tests/test_selective_checks.py`: 237 tests pass. The three 
cases whose expected
     outputs change fail on `main`. The two that keep the full matrix (the 
pytest plugin, and a helper it
     imports) guard against narrowing too much.
   - The `_imports_module` cases cover `from module import`, `from package 
import name` (including
     multi-line), `import`, dotted strings such as `pytest_plugins` entries, 
and near-misses.
   - Every `tests_common` module classifies without errors. The `git grep` scan 
only runs when a
     helper changed, and took about 0.8 s for `mock_context.py`.
   - prek pre-commit and manual stages pass.
   
   The mypy hooks are unaffected: the helper itself stays in the changed-file 
list, so
   `mypy-devel-common` still runs whenever a helper changes.
   
   ---
   
   ##### Was generative AI tooling used to co-author this PR?
   
   - [X] Yes — Claude Code (Opus 5.5)
   
   Generated-by: Claude Code (Opus 5.5) following [the 
guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to