GitHub user zozo123 closed a discussion: CI: make environment materialization 
scale with unique environments, not matrix jobs

This is an **ideation / measurement discussion**, rather than a proposal for 
one specific implementation.

Airflow has done a lot of work to avoid unnecessary CI work through selective 
checks, image build caching, and matrix reduction.

There appears to be another large optimization dimension:

> **Many short-lived CI jobs independently materialize the same multi-GB 
> execution environment before doing useful work.**

The question I would like to explore here is:

> **Can the cost of preparing CI environments become proportional to the number 
> of unique environments we need, rather than proportional to the number of 
> matrix jobs consuming them?**

This discussion can stay open as a place to collect measurements, experiments 
and smaller PRs.

---

## Current observation

I measured the completed September 28 AMD canary, [run 
36432564874](https://github.com/apache/airflow/actions/runs/36432564874).

Approximately:

| Metric | AMD canary |
|---|---:|
| Jobs | 283 |
| Sum of runner time | 59.68 h |
| Jobs preparing CI/PROD images | 231 |
| Image-preparation runner time | 16.32 h |
| Share of runner compute | **27.3%** |
| Nominal repeated image-artifact payload | **~529 GB** |

The CI-image artifacts were approximately 2.8 GB compressed, while the loaded 
image was approximately 8 GB.

A representative Postgres shard looked roughly like:

```text
setup Breeze             ~0m14s
restore image artifact   ~7m24s
docker load              ~2m00s
actual tests             ~19m48s
```

The effect is much larger for short jobs.

In that same run:

```text
53 jobs  >50% of lifetime preparing the environment
102 jobs >=30%
```

Some examples:

```text
Helm statsd
    total        ~7.6m
    environment  ~6.9m
    useful work  ~0.7m

Integration mongo
    total        ~7.9m
    environment  ~6.9m

Integration redis
    total        ~7.2m
    environment  ~6.2m
```

So this is not only a bandwidth issue. For some matrix entries the execution 
environment costs considerably more than the test itself.

A systematic sample of completed AMD PR workflows also showed image preparation 
at roughly one third of aggregate runner time, so the pattern does not appear 
limited to scheduled canaries.

---

## Current shape

The current CI path intentionally builds the image once and distributes it to 
downstream jobs as an artifact/stash:

```text
                   build image once
                          |
                     docker save
                          |
                    stash artifact
                          |
             +------------+-------------+
             |            |             |
          runner 1     runner 2      runner N
             |            |             |
          restore       restore        restore
             |            |             |
        docker load   docker load    docker load
             |            |             |
           tests          tests          tests
```

`.github/actions/prepare_breeze_and_image/action.yml` currently restores the 
image stash into `/mnt` and then invokes `breeze ... image load`.

The main AMD workflow also deliberately uses:

```yaml
push-image: "false"
upload-image-artifact: "true"
```

This architecture made sense when artifact sharing replaced the older 
`pull_request_target` / registry approach in 
[#43268](https://github.com/apache/airflow/issues/43268).

I do **not** suggest simply reverting that security decision.

However, image size and CI fan-out have grown enough that the trade-off seems 
worth measuring again.

There is also useful precedent in 
[#70618](https://github.com/apache/airflow/pull/70618): another Airflow CI path 
was changed to build a CI image once and share it between consumers, and 
cross-run reuse through the registry was identified there as a possible 
direction.

---

## Desired property

Conceptually, today the environment cost is approximately:

```text
number_of_jobs × materialize(environment)
```

Could we move closer to:

```text
number_of_unique_environments × materialize(environment)
+
number_of_jobs × cheap_attach(environment)
```

while keeping each job isolated and disposable?

For example:

```text
                    immutable environment E
                              |
                   content-addressed storage
                              |
              +---------------+---------------+
              |               |               |
         fresh job A     fresh job B     fresh job C
              |               |               |
        private writable private writable private writable
            overlay         overlay         overlay
              |               |               |
             tests           tests           tests
```

The workspace, secrets, processes and writable state can remain ephemeral.

Only immutable environment content needs reuse.

---

## Possible directions to benchmark

These are deliberately alternatives / combinations rather than a proposed 
design:

1. **OCI registry transport by immutable digest**

   Instead of `docker save -> artifact -> restore -> docker load`, publish 
trusted images/layers once and pull the exact digest.

   This removes the tar-style transport but still requires cold runners to 
transfer layers.

2. **Runner-local or shared content-addressed OCI storage**

   Ephemeral jobs could start on clean machines/VMs while immutable image 
layers are already available locally or through a nearby content-addressed 
store.

   A cache hit should mean exact content identity, not a mutable image tag.

3. **Read-only image snapshots + private CoW overlays**

   If a runner implementation supports cheap filesystem/block snapshots, many 
isolated jobs could attach the same immutable parent image and receive 
independent writable overlays.

   The runner remains disposable; only trusted immutable content persists.

4. **Cache-affinity scheduling**

   If several jobs need the same Python/platform/environment, prefer hosts 
where those exact layers already exist rather than repeatedly warming unrelated 
machines.

5. **Pre-warmed ephemeral runners**

   Frequently used CI/service images could be present in the runner 
image/snapshot itself.

   This may be especially useful for stable dependencies and service images 
such as PostgreSQL/MySQL/Redis/Kubernetes helpers.

6. **Separate environment inputs from source inputs**

   A potentially deeper direction is:

   ```text
   Dockerfile / OS deps / Python / constraints
                       |
                       v
               environment digest E
                       |
                +------+------+------+
                |      |      |      |
              source source source source
               SHA1   SHA2   SHA3   SHA4
   ```

   Many Airflow changes alter Python/source files without changing the 
OS/tool/dependency environment.

   Where safe, Breeze's source-mount capabilities might allow those commits to 
reuse the same dependency/tooling environment rather than making the source SHA 
part of a multi-GB image that must then be redistributed.

7. **Smaller purpose-built environments**

   A job doing tens of seconds of Helm or Task SDK work may not need the same 
large environment as a full Airflow Python integration shard.

   Environment size could potentially follow job requirements.

8. **Amortize setup for very small jobs**

   Where GitHub job isolation is not itself important to the test, several very 
small logical shards might run on one already-prepared environment while 
retaining separate test reporting.

9. **Alternative/larger runner experiments**

   Different runner implementations or larger machines are worth measuring as a 
control, but raw CPU alone does not solve repeated image 
transport/materialization.

   An interesting runner experiment would combine fresh job isolation with 
persistent immutable environment content.

These directions are not mutually exclusive.

---

## Security / correctness constraints

Any optimization here should preserve the current trust model.

I think the useful invariants are approximately:

```text
selected jobs before == selected jobs after
test commands before == test commands after

requested environment digest == executed environment digest
requested platform           == executed platform
requested Python             == executed Python

untrusted job:
    may read approved immutable cache
    must not publish trusted cache state

job completion:
    destroys workspace
    destroys secrets
    destroys private writable overlay

cache miss:
    falls back to known-correct materialization path
```

In particular:

### Do not use mutable environment identity

Prefer an OCI/image digest or another content-derived identity rather than:

```text
airflow-ci:latest
```

### Do not let fork PRs poison trusted caches

One possible model is:

```text
trusted global immutable parent
                |
        read-only for PR jobs
                |
        +-------+-------+
        |               |
 workflow A          workflow B
 private state      private state
```

Trusted canaries/main jobs could produce reusable parents.

Untrusted PR jobs could consume approved state but not mutate what subsequent 
trusted jobs see.

The existing artifact mechanism should remain a valid fallback where those 
guarantees cannot be provided.

---

## Platform scope

I do not think an experiment needs to solve every GitHub Actions platform 
before it is useful.

The expensive Airflow path being discussed here is already primarily Linux.

A first experiment could target only `linux/amd64` trusted canaries while 
leaving ARM and other jobs untouched.

If a runner/backend supports Linux and Windows but not macOS, that should not 
by itself prevent testing this idea.

Likewise, lack of ARM support should not block an AMD experiment.

---

## Measurement

Before changing PR behavior, I suggest testing this on trusted scheduled 
canaries.

For each experiment collect:

```text
runner allocation -> first useful test
prepare_breeze_and_image wall time
artifact / registry bytes transferred
image unpack/load time
test wall time
total runner-minutes
workflow wall time
cache hit/miss
CPU utilization where available
test/result parity
```

And compare distributions across multiple runs rather than one best case:

```text
median
p90/p95
cold cache
warm cache
failure/fallback behavior
```

A useful experiment matrix might be:

```text
A. current GitHub-hosted path

B. alternate runner, current image path
   -> isolates hardware/runner effects

C. alternate transport/cache, same runner class
   -> isolates environment-materialization effects

D. warm content-addressed/snapshot environment
   -> tests the intended steady state
```

---

## Expected magnitude / Amdahl's law

This should not be sold as “10x Airflow CI”.

In the measured full AMD canary, environment preparation was about 27% of total 
runner compute.

Even reducing that to zero would therefore cap the whole-workflow compute 
improvement from this optimization alone at roughly:

```text
1 / (1 - 0.27) ~= 1.37x
```

So a full 2x CI improvement would require another independent reduction in 
actual executed work, for example further matrix/test-impact optimization.

However, individual setup-dominated jobs can improve dramatically.

For a job like:

```text
6.9m environment
0.7m useful work
```

reducing environment attachment to seconds approaches an order-of-magnitude 
improvement for that job without skipping any tests.

That is still valuable because it reduces runner consumption, network traffic 
and latency while preserving coverage.

---

## Why keep this as an ideation discussion?

I don't think we know yet whether the best answer is:

```text
GHCR
runner-local OCI cache
filesystem/block snapshots
source/environment separation
smaller CI images
pre-warmed ephemeral runners
job packing
cache-affinity scheduling
some combination of these
```

I would prefer to use this discussion to:

```text
measure
    ->
form hypotheses
    ->
run small trusted-canary experiments
    ->
compare
    ->
land independent improvements
```

rather than select the implementation first.

The architectural question is the interesting part:

> **Can Airflow keep jobs ephemeral and isolated while making expensive 
> immutable CI environment state reusable?**

Or equivalently:

> **Can environment provisioning become proportional to what changed / how many 
> unique environments exist, rather than proportional to the number of CI 
> shards?**

GitHub link: https://github.com/apache/airflow/discussions/73856

----
This is an automatically sent email for [email protected].
To unsubscribe, please send an email to: [email protected]

Reply via email to