GitHub user zozo123 closed a discussion: CI: make environment materialization scale with unique environments, not matrix jobs
This is an **ideation / measurement discussion**, rather than a proposal for one specific implementation. Airflow has done a lot of work to avoid unnecessary CI work through selective checks, image build caching, and matrix reduction. There appears to be another large optimization dimension: > **Many short-lived CI jobs independently materialize the same multi-GB > execution environment before doing useful work.** The question I would like to explore here is: > **Can the cost of preparing CI environments become proportional to the number > of unique environments we need, rather than proportional to the number of > matrix jobs consuming them?** This discussion can stay open as a place to collect measurements, experiments and smaller PRs. --- ## Current observation I measured the completed September 28 AMD canary, [run 36432564874](https://github.com/apache/airflow/actions/runs/36432564874). Approximately: | Metric | AMD canary | |---|---:| | Jobs | 283 | | Sum of runner time | 59.68 h | | Jobs preparing CI/PROD images | 231 | | Image-preparation runner time | 16.32 h | | Share of runner compute | **27.3%** | | Nominal repeated image-artifact payload | **~529 GB** | The CI-image artifacts were approximately 2.8 GB compressed, while the loaded image was approximately 8 GB. A representative Postgres shard looked roughly like: ```text setup Breeze ~0m14s restore image artifact ~7m24s docker load ~2m00s actual tests ~19m48s ``` The effect is much larger for short jobs. In that same run: ```text 53 jobs >50% of lifetime preparing the environment 102 jobs >=30% ``` Some examples: ```text Helm statsd total ~7.6m environment ~6.9m useful work ~0.7m Integration mongo total ~7.9m environment ~6.9m Integration redis total ~7.2m environment ~6.2m ``` So this is not only a bandwidth issue. For some matrix entries the execution environment costs considerably more than the test itself. A systematic sample of completed AMD PR workflows also showed image preparation at roughly one third of aggregate runner time, so the pattern does not appear limited to scheduled canaries. --- ## Current shape The current CI path intentionally builds the image once and distributes it to downstream jobs as an artifact/stash: ```text build image once | docker save | stash artifact | +------------+-------------+ | | | runner 1 runner 2 runner N | | | restore restore restore | | | docker load docker load docker load | | | tests tests tests ``` `.github/actions/prepare_breeze_and_image/action.yml` currently restores the image stash into `/mnt` and then invokes `breeze ... image load`. The main AMD workflow also deliberately uses: ```yaml push-image: "false" upload-image-artifact: "true" ``` This architecture made sense when artifact sharing replaced the older `pull_request_target` / registry approach in [#43268](https://github.com/apache/airflow/issues/43268). I do **not** suggest simply reverting that security decision. However, image size and CI fan-out have grown enough that the trade-off seems worth measuring again. There is also useful precedent in [#70618](https://github.com/apache/airflow/pull/70618): another Airflow CI path was changed to build a CI image once and share it between consumers, and cross-run reuse through the registry was identified there as a possible direction. --- ## Desired property Conceptually, today the environment cost is approximately: ```text number_of_jobs × materialize(environment) ``` Could we move closer to: ```text number_of_unique_environments × materialize(environment) + number_of_jobs × cheap_attach(environment) ``` while keeping each job isolated and disposable? For example: ```text immutable environment E | content-addressed storage | +---------------+---------------+ | | | fresh job A fresh job B fresh job C | | | private writable private writable private writable overlay overlay overlay | | | tests tests tests ``` The workspace, secrets, processes and writable state can remain ephemeral. Only immutable environment content needs reuse. --- ## Possible directions to benchmark These are deliberately alternatives / combinations rather than a proposed design: 1. **OCI registry transport by immutable digest** Instead of `docker save -> artifact -> restore -> docker load`, publish trusted images/layers once and pull the exact digest. This removes the tar-style transport but still requires cold runners to transfer layers. 2. **Runner-local or shared content-addressed OCI storage** Ephemeral jobs could start on clean machines/VMs while immutable image layers are already available locally or through a nearby content-addressed store. A cache hit should mean exact content identity, not a mutable image tag. 3. **Read-only image snapshots + private CoW overlays** If a runner implementation supports cheap filesystem/block snapshots, many isolated jobs could attach the same immutable parent image and receive independent writable overlays. The runner remains disposable; only trusted immutable content persists. 4. **Cache-affinity scheduling** If several jobs need the same Python/platform/environment, prefer hosts where those exact layers already exist rather than repeatedly warming unrelated machines. 5. **Pre-warmed ephemeral runners** Frequently used CI/service images could be present in the runner image/snapshot itself. This may be especially useful for stable dependencies and service images such as PostgreSQL/MySQL/Redis/Kubernetes helpers. 6. **Separate environment inputs from source inputs** A potentially deeper direction is: ```text Dockerfile / OS deps / Python / constraints | v environment digest E | +------+------+------+ | | | | source source source source SHA1 SHA2 SHA3 SHA4 ``` Many Airflow changes alter Python/source files without changing the OS/tool/dependency environment. Where safe, Breeze's source-mount capabilities might allow those commits to reuse the same dependency/tooling environment rather than making the source SHA part of a multi-GB image that must then be redistributed. 7. **Smaller purpose-built environments** A job doing tens of seconds of Helm or Task SDK work may not need the same large environment as a full Airflow Python integration shard. Environment size could potentially follow job requirements. 8. **Amortize setup for very small jobs** Where GitHub job isolation is not itself important to the test, several very small logical shards might run on one already-prepared environment while retaining separate test reporting. 9. **Alternative/larger runner experiments** Different runner implementations or larger machines are worth measuring as a control, but raw CPU alone does not solve repeated image transport/materialization. An interesting runner experiment would combine fresh job isolation with persistent immutable environment content. These directions are not mutually exclusive. --- ## Security / correctness constraints Any optimization here should preserve the current trust model. I think the useful invariants are approximately: ```text selected jobs before == selected jobs after test commands before == test commands after requested environment digest == executed environment digest requested platform == executed platform requested Python == executed Python untrusted job: may read approved immutable cache must not publish trusted cache state job completion: destroys workspace destroys secrets destroys private writable overlay cache miss: falls back to known-correct materialization path ``` In particular: ### Do not use mutable environment identity Prefer an OCI/image digest or another content-derived identity rather than: ```text airflow-ci:latest ``` ### Do not let fork PRs poison trusted caches One possible model is: ```text trusted global immutable parent | read-only for PR jobs | +-------+-------+ | | workflow A workflow B private state private state ``` Trusted canaries/main jobs could produce reusable parents. Untrusted PR jobs could consume approved state but not mutate what subsequent trusted jobs see. The existing artifact mechanism should remain a valid fallback where those guarantees cannot be provided. --- ## Platform scope I do not think an experiment needs to solve every GitHub Actions platform before it is useful. The expensive Airflow path being discussed here is already primarily Linux. A first experiment could target only `linux/amd64` trusted canaries while leaving ARM and other jobs untouched. If a runner/backend supports Linux and Windows but not macOS, that should not by itself prevent testing this idea. Likewise, lack of ARM support should not block an AMD experiment. --- ## Measurement Before changing PR behavior, I suggest testing this on trusted scheduled canaries. For each experiment collect: ```text runner allocation -> first useful test prepare_breeze_and_image wall time artifact / registry bytes transferred image unpack/load time test wall time total runner-minutes workflow wall time cache hit/miss CPU utilization where available test/result parity ``` And compare distributions across multiple runs rather than one best case: ```text median p90/p95 cold cache warm cache failure/fallback behavior ``` A useful experiment matrix might be: ```text A. current GitHub-hosted path B. alternate runner, current image path -> isolates hardware/runner effects C. alternate transport/cache, same runner class -> isolates environment-materialization effects D. warm content-addressed/snapshot environment -> tests the intended steady state ``` --- ## Expected magnitude / Amdahl's law This should not be sold as “10x Airflow CI”. In the measured full AMD canary, environment preparation was about 27% of total runner compute. Even reducing that to zero would therefore cap the whole-workflow compute improvement from this optimization alone at roughly: ```text 1 / (1 - 0.27) ~= 1.37x ``` So a full 2x CI improvement would require another independent reduction in actual executed work, for example further matrix/test-impact optimization. However, individual setup-dominated jobs can improve dramatically. For a job like: ```text 6.9m environment 0.7m useful work ``` reducing environment attachment to seconds approaches an order-of-magnitude improvement for that job without skipping any tests. That is still valuable because it reduces runner consumption, network traffic and latency while preserving coverage. --- ## Why keep this as an ideation discussion? I don't think we know yet whether the best answer is: ```text GHCR runner-local OCI cache filesystem/block snapshots source/environment separation smaller CI images pre-warmed ephemeral runners job packing cache-affinity scheduling some combination of these ``` I would prefer to use this discussion to: ```text measure -> form hypotheses -> run small trusted-canary experiments -> compare -> land independent improvements ``` rather than select the implementation first. The architectural question is the interesting part: > **Can Airflow keep jobs ephemeral and isolated while making expensive > immutable CI environment state reusable?** Or equivalently: > **Can environment provisioning become proportional to what changed / how many > unique environments exist, rather than proportional to the number of CI > shards?** GitHub link: https://github.com/apache/airflow/discussions/73856 ---- This is an automatically sent email for [email protected]. To unsubscribe, please send an email to: [email protected]
