zozo123 opened a new pull request, #74055:
URL: https://github.com/apache/airflow/pull/74055
On pull requests, a change to any Python file under `devel-common/`, or to
anything in
`airflow-core/tests/unit/utils/`, forces the full test matrix
(`TESTS_UTILS_FILES`). On top of that,
`find_provider_affected()` treats every `devel-common/` file as common
provider code and selects all
providers. Most of those files do not need that:
- **`devel-common/src/tests_common/` helpers** are mostly imported by a
handful of test files. Of
the 62 helpers, 18 are loaded by every test run (the pytest plugin,
anything it imports,
conftest and package `__init__` modules). The other 44 are not.
`mock_context.py`, for example,
reaches 9 test files in 3 providers (through `run_deferrable.py`).
- **`airflow-core/tests/unit/utils/`** holds only `test_*.py` files. Its few
cross-imports are from
other core tests, never from providers.
- **`devel-common/src/sphinx_exts/`** only affects the docs build.
With this change, a changed `tests_common` helper selects the test files
that import it, directly or
through other helpers, as if those files had changed. Only helpers that
every test run loads still
force the full matrix. The `airflow-core` unit-test directory is mapped like
any other core test
files, and Sphinx extensions trigger the docs build (`DOC_FILES`).
### What it saves, measured
I replayed `SelectiveChecks` on the exact inputs (changed files, labels,
commit) recorded in the
*Build info* logs of real `ci-amd.yml` PR runs, once with `main` and once
with this branch. For each
run whose outputs changed, I estimated the new cost from that run's actual
runner time, scaled by
the median runner time of green PR runs with the new versus the old
selective-check outcome.
| Sample of PR runs | Runs whose matrix shrinks | From PRs | Runner time
saved | Runs that got bigger |
|---|---|---|---|---|
| 28 Sep – 1 Oct (819 runs, 3.6 days) | 86 (28 green, 58 failed/cancelled) |
16 | 269 runner-hours/day | 0 |
| 14 – 18 Sep (694 runs over 60 jobs, 4.3 days) | 37 (10 green, 27
failed/cancelled) | 4 | 106 runner-hours/day | 0 |
Every changed run goes from `full-tests-needed=true` to a selective run,
typically "all core test
types plus the affected providers and their direct dependents". Those cost a
median 12.5 h on green
runs, versus 24.1 h for a full run. PRs that edit test helpers tend to be
pushed many times, which is
why failed and cancelled runs carry most of the saving.
### Checks run
- `dev/breeze/tests/test_selective_checks.py`: 237 tests pass. The three
cases whose expected
outputs change fail on `main`. The two that keep the full matrix (the
pytest plugin, and a helper it
imports) guard against narrowing too much.
- The `_imports_module` cases cover `from module import`, `from package
import name` (including
multi-line), `import`, dotted strings such as `pytest_plugins` entries,
and near-misses.
- Every `tests_common` module classifies without errors. The `git grep` scan
only runs when a
helper changed, and took about 0.8 s for `mock_context.py`.
- prek pre-commit and manual stages pass.
The mypy hooks are unaffected: the helper itself stays in the changed-file
list, so
`mypy-devel-common` still runs whenever a helper changes.
---
##### Was generative AI tooling used to co-author this PR?
- [X] Yes — Claude Code (Opus 5.5)
Generated-by: Claude Code (Opus 5.5) following [the
guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]