This is an automated email from the ASF dual-hosted git repository.

github-merge-queue[bot] pushed a commit to branch 
gh-readonly-queue/main/pr-7172-5de9bf94766d441e7d1655bf326a9208dc6769a5
in repository https://gitbox.apache.org/repos/asf/texera.git

commit 092a02c46eae42305f18cd1025ae34a33b41c1d7
Author: Yicong Huang <[email protected]>
AuthorDate: Wed Aug 26 07:31:28 2026 +0000

    ci: route merge-queue builds to self-hosted runner tiers with hosted 
fallback (#7172)
    
    ### What changes were proposed in this PR?
    
    Merge-queue builds default to four self-hosted runner tiers; PR runs opt
    in via a `ci:self-hosted` label, and push/scheduled runs stay
    GitHub-hosted.
    
    - `arc-heavy-linux`: amber, amber-integration, frontend
    - `arc-medium-linux`: platform, pyamber
    - `arc-light-linux`: service jobs and the precheck/aggregator control
    plane
    - `arc-bench-linux`: dedicated bench box, so bench numbers stay
    comparable
    
    A `runner-heartbeat` workflow proves every 15 minutes that the fleet is
    usable, not just reachable: disk and inode headroom, a live docker
    daemon, the tools builds assume. A stale or failed beat routes precheck
    back to `ubuntu-latest` and flips `ARC_AVAILABLE`, consulted by
    first-hop `runs-on` expressions and usable as a manual kill switch.
    `ARC_STATUS_TOKEN` optionally adds per-tier congestion probing.
    
    These runners are not ephemeral, so every self-hosted job ends with an
    `if: always()` step dropping what it produced: the workspace, its own
    `/tmp` scratch, and the containers it started with `docker run -d`,
    which `services:` never removes and which would block the next job's
    containers. Entry points clear their own leftovers first, since a killed
    job never reaches cleanup.
    
    Step conditions become `runner.os == 'Linux'`, so the matrices are
    label-agnostic. 25 runners are registered; no new credentials needed.
    
    ### Any related issues, documentation, discussions?
    
    Closes #7171
    Discussion: #6973
    
    ### How was this PR tested?
    
    - Repeated full CI rounds on a fork against a live 25-runner fleet
    ([Yicong-Huang#25](https://github.com/Yicong-Huang/texera/pull/25)),
    every job landing on its intended tier.
    - Fallback verified per tier: with a scale set offline, precheck routed
    that tier back to `ubuntu-latest`.
    - This PR carries `ci:self-hosted`, so its own rounds exercise the
    cleanup and self-healing paths on the real fleet.
    - Label-free rounds confirmed hosted behavior is unchanged.
    
    ### Was this PR authored or co-authored using generative AI tooling?
    
    Generated-by: Claude Code (Claude Fable 5, Claude Opus 5)
    
    ---------
    
    Co-authored-by: Claude Fable 5 <[email protected]>
---
 .github/workflows/backport-checks.yml  |  23 +++-
 .github/workflows/benchmarks.yml       |  85 ++++++++++++-
 .github/workflows/build.yml            | 201 ++++++++++++++++++++++++++----
 .github/workflows/check-header.yml     |  20 ++-
 .github/workflows/precheck.yml         | 216 ++++++++++++++++++++++++++++++++-
 .github/workflows/required-checks.yml  |  19 ++-
 .github/workflows/runner-heartbeat.yml | 199 ++++++++++++++++++++++++++++++
 7 files changed, 734 insertions(+), 29 deletions(-)

diff --git a/.github/workflows/backport-checks.yml 
b/.github/workflows/backport-checks.yml
index 7288928ef3..9bfcc75d7f 100644
--- a/.github/workflows/backport-checks.yml
+++ b/.github/workflows/backport-checks.yml
@@ -41,6 +41,9 @@ on:
       - unlabeled
 
 permissions:
+  # actions: read is required by the called precheck.yml (backlog probe);
+  # a reusable workflow's permissions must be a subset of its caller's.
+  actions: read
   checks: write
   contents: read
   pull-requests: read
@@ -62,7 +65,9 @@ jobs:
   apply-check:
     needs: precheck
     if: ${{ needs.precheck.outputs.backport_targets != '[]' }}
-    runs-on: ubuntu-latest
+    # Control-plane resilience: follow precheck's probed routing (label
+    # present AND the scale set has online runners); ubuntu-latest otherwise.
+    runs-on: ${{ needs.precheck.outputs.light_runner == 'arc-light-linux' && 
'arc-light-linux' || 'ubuntu-latest' }}
     outputs:
       buildable: ${{ steps.check.outputs.buildable }}
     steps:
@@ -110,6 +115,22 @@ jobs:
           name: backport-preflight
           path: preflight/
           retention-days: 3
+      - name: Clean up job artifacts
+        # Classic self-hosted runners hand this workspace to the next job
+        # as-is. The dry-run cherry-picks leave it parked on a release branch,
+        # so also abort any half-applied pick before cleaning. Runs after the
+        # artifact upload above, so the preflight results are already safe.
+        # always() so a failed or cancelled job cleans up too; the
+        # GitHub-hosted fallback is a throwaway VM and skips it. .git stays so
+        # the next checkout is incremental; the fallback covers a job whose
+        # checkout never completed.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          git cherry-pick --abort 2>/dev/null || true
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   # Build the backported tree for targets that applied cleanly, so a clean
   # cherry-pick that nonetheless fails to compile on the release branch is
diff --git a/.github/workflows/benchmarks.yml b/.github/workflows/benchmarks.yml
index a427a18646..87dc28b426 100644
--- a/.github/workflows/benchmarks.yml
+++ b/.github/workflows/benchmarks.yml
@@ -88,6 +88,8 @@ on:
   workflow_dispatch:
 
 permissions:
+  # actions: read backs the queued-job backlog probe for runner routing.
+  actions: read
   contents: write
 
 concurrency:
@@ -101,9 +103,15 @@ jobs:
     # run). Lifted from required-checks.yml's precheck so the trigger
     # surface matches amber-integration exactly.
     name: Precheck
-    runs-on: ubuntu-latest
+    # Control-plane resilience: with the ci:self-hosted label this job
+    # runs on arc-light-linux so a jammed GitHub-hosted pool cannot stall
+    # the pipeline. Availability cannot be probed before this job runs, so
+    # the repo variable ARC_AVAILABLE is the kill switch: set it to
+    # 'false' when the cluster is down to route back to ubuntu-latest.
+    runs-on: ${{ (github.event_name == 'pull_request' && 
contains(github.event.pull_request.labels.*.name, 'ci:self-hosted') && 
vars.ARC_AVAILABLE != 'false') && 'arc-light-linux' || 'ubuntu-latest' }}
     outputs:
       run_bench: ${{ steps.decide.outputs.run_bench }}
+      linux_runner: ${{ steps.decide.outputs.linux_runner }}
     steps:
       - name: Wait for Pull Request Labeler
         if: github.event_name == 'pull_request'
@@ -132,6 +140,12 @@ jobs:
       - name: Decide whether to run bench
         id: decide
         uses: actions/github-script@v9
+        env:
+          # Optional PAT (administration:read) for the runner availability
+          # probe; see precheck.yml.
+          ARC_STATUS_TOKEN: ${{ secrets.ARC_STATUS_TOKEN }}
+          # Kill switch mirrored by runner-heartbeat.yml's watchdog.
+          ARC_AVAILABLE: ${{ vars.ARC_AVAILABLE }}
         with:
           script: |
             const eventName = context.eventName;
@@ -139,6 +153,7 @@ jobs:
               // push to main / workflow_dispatch always run.
               core.info(`event=${eventName} — running unconditionally`);
               core.setOutput("run_bench", "true");
+              core.setOutput("linux_runner", "ubuntu-latest");
               return;
             }
             // Re-fetch labels: the labeler may have just added some.
@@ -168,12 +183,59 @@ jobs:
                 : "No trigger label present; skipping bench."
             );
             core.setOutput("run_bench", shouldRun ? "true" : "false");
+            // ci:self-hosted routes the bench onto the dedicated
+            // arc-bench-linux runner. Availability-only fallback: benchmark
+            // numbers must come from consistent hardware, so a busy bench
+            // runner means WAIT (queue), never drift to GitHub-hosted —
+            // only a set with no online runner at all falls back.
+            let linuxRunner = labels.includes("ci:self-hosted") ? 
"arc-bench-linux" : "ubuntu-latest";
+            if (linuxRunner === "arc-bench-linux" && process.env.ARC_AVAILABLE 
=== "false") {
+              linuxRunner = "ubuntu-latest";
+              core.warning("ARC_AVAILABLE=false; bench falls back to 
ubuntu-latest.");
+            }
+            if (linuxRunner === "arc-bench-linux" && 
process.env.ARC_STATUS_TOKEN) {
+              try {
+                // Raw request rather than the SDK because this call needs a
+                // second token (administration:read); see the same probe in
+                // precheck.yml for why one Octokit client cannot carry it.
+                const res = await fetch(
+                  
`https://api.github.com/repos/${context.repo.owner}/${context.repo.repo}/actions/runners?per_page=100`,
+                  {
+                    headers: {
+                      authorization: `Bearer ${process.env.ARC_STATUS_TOKEN}`,
+                      accept: "application/vnd.github+json",
+                    },
+                  }
+                );
+                if (!res.ok) throw new Error(`runner list HTTP ${res.status}`);
+                const online = (await res.json()).runners?.some(
+                  (r) => r.status === "online" && r.labels.some((l) => l.name 
=== "arc-bench-linux")
+                );
+                if (!online) {
+                  linuxRunner = "ubuntu-latest";
+                  core.warning("arc-bench-linux has no online runner; falling 
back to ubuntu-latest.");
+                }
+              } catch (e) {
+                core.warning(`Runner availability probe failed (${e.message}); 
trusting the label.`);
+              }
+            }
+            core.setOutput("linux_runner", linuxRunner);
+            core.info(`Linux runner: ${linuxRunner}`);
+
+      - name: Clean up job artifacts
+        # No checkout here, but a classic self-hosted runner still keeps the
+        # workspace directory it created for this job; leave it empty for the
+        # next one. always() so a failed or cancelled job cleans up too.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        run: |
+          find "${GITHUB_WORKSPACE:?}" -mindepth 1 -maxdepth 1 -exec rm -rf -- 
{} + 2>/dev/null ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   bench:
     name: Bench
     needs: precheck
     if: ${{ needs.precheck.outputs.run_bench == 'true' }}
-    runs-on: ubuntu-latest
+    runs-on: ${{ needs.precheck.outputs.linux_runner }}
     env:
       JAVA_OPTS: -Xms2048M -Xmx2048M -Xss6M -XX:ReservedCodeCacheSize=256M 
-Dfile.encoding=UTF-8
       JVM_OPTS: -Xms2048M -Xmx2048M -Xss6M -XX:ReservedCodeCacheSize=256M 
-Dfile.encoding=UTF-8
@@ -416,3 +478,22 @@ jobs:
           alert-threshold: "150%"
           comment-on-alert: false
           summary-always: true
+
+      - name: Clean up job artifacts
+        # arc-bench-linux is a dedicated machine every bench run lands on, so
+        # its leftovers (sbt target/, bench-results/, the main-baseline
+        # checkout) accumulate faster than anywhere else — and stale state on
+        # the bench box is a comparability risk, not just a disk one. Runs
+        # after the artifact upload above, so no results are lost. always() so
+        # a failed or cancelled run cleans up too; the GitHub-hosted fallback
+        # is a throwaway VM and skips it. `git clean -x` covers gitignored
+        # build output, `-ff` nested git dirs; .git stays so the next
+        # checkout is incremental. The fallback covers a run whose checkout
+        # never completed.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          rm -rf /tmp/protoc.zip
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml
index 6a2a6cff2c..d0864968d9 100644
--- a/.github/workflows/build.yml
+++ b/.github/workflows/build.yml
@@ -36,6 +36,24 @@ on:
         required: false
         type: string
         default: ""
+      # Runner labels for Linux jobs, three tiers (precheck decides via the
+      # ci:self-hosted PR label; GitHub-hosted ubuntu-latest otherwise).
+      # heavy: amber / amber-integration / frontend (8C/16G class; the
+      # Angular prod build OOMs under the medium tier's 8Gi limit).
+      heavy_runner:
+        required: false
+        type: string
+        default: "ubuntu-latest"
+      # medium: platform / platform-integration / pyamber (4C/8G class).
+      medium_runner:
+        required: false
+        type: string
+        default: "ubuntu-latest"
+      # light: infra / agent-service / pyright (2C/4G class).
+      light_runner:
+        required: false
+        type: string
+        default: "ubuntu-latest"
       run_frontend:
         required: false
         type: boolean
@@ -88,11 +106,11 @@ jobs:
     runs-on: ${{ matrix.os }}
     strategy:
       matrix:
-        os: [ubuntu-latest, windows-latest, macos-latest]
+        os: ["${{ inputs.heavy_runner }}", windows-latest, macos-latest]
         include:
           - os: macos-latest
             arch: arm64
-          - os: ubuntu-latest
+          - os: "${{ inputs.heavy_runner }}"
             arch: x64
           - os: windows-latest
             arch: x64
@@ -133,12 +151,12 @@ jobs:
       - name: Prod build
         run: yarn --cwd frontend run build:ci
       - name: Check bundled npm packages against per-module LICENSE-binary 
files
-        if: matrix.os == 'ubuntu-latest'
+        if: runner.os == 'Linux'
         run: ./bin/licensing/check_binary_deps.py ${{ inputs.mode == 'PR' && 
'--ignore-transitive-version' || '' }} npm frontend/dist/3rdpartylicenses.json
       - name: Run frontend unit tests
         run: yarn --cwd frontend run test:ci
       - name: Upload frontend coverage to Codecov
-        if: matrix.os == 'ubuntu-latest' && always()
+        if: runner.os == 'Linux' && always()
         uses: codecov/codecov-action@fb8b3582c8e4def4969c97caa2f19720cb33a72f 
# v7.0.0
         with:
           token: ${{ secrets.CODECOV_TOKEN }}
@@ -149,7 +167,7 @@ jobs:
         # vitest.config.ts adds a `junit` reporter that writes to junit.xml
         # in the working dir; the @angular/build:unit-test runner forwards
         # vitest config through unchanged.
-        if: matrix.os == 'ubuntu-latest' && !cancelled()
+        if: runner.os == 'Linux' && !cancelled()
         uses: codecov/codecov-action@fb8b3582c8e4def4969c97caa2f19720cb33a72f 
# v7.0.0
         with:
           token: ${{ secrets.CODECOV_TOKEN }}
@@ -159,14 +177,14 @@ jobs:
           disable_search: true
           fail_ci_if_error: false
       - name: Install Playwright Chromium
-        run: yarn --cwd frontend playwright install ${{ matrix.os == 
'ubuntu-latest' && '--with-deps' || '' }} chromium
+        run: yarn --cwd frontend playwright install ${{ runner.os == 'Linux' 
&& '--with-deps' || '' }} chromium
       - name: Run frontend browser-mode tests
         run: yarn --cwd frontend ng run gui:test-browser
       - name: Upload frontend browser-mode test results to Codecov
         # vitest.browser.config.ts emits junit-browser.xml (distinct from
         # the unit-test report). Same `frontend` flag — Codecov merges
         # multi-file uploads under one flag.
-        if: matrix.os == 'ubuntu-latest' && !cancelled()
+        if: runner.os == 'Linux' && !cancelled()
         uses: codecov/codecov-action@fb8b3582c8e4def4969c97caa2f19720cb33a72f 
# v7.0.0
         with:
           token: ${{ secrets.CODECOV_TOKEN }}
@@ -175,6 +193,31 @@ jobs:
           report_type: test_results
           disable_search: true
           fail_ci_if_error: false
+      - name: Clean up job artifacts
+        # The arc-* runners are classic (non-ephemeral): the same workspace,
+        # /tmp and docker daemon are handed to the next job that lands on the
+        # machine, so this job's output (here: node_modules/, .yarn/cache,
+        # dist/) would pile up until the runner's root disk fills. Every
+        # self-hosted job below ends with the same kind of step.
+        #
+        # always() so a failed or cancelled job cleans up too. GitHub-hosted
+        # legs (windows / macos / ubuntu-latest fallback) are throwaway VMs
+        # and skip it. `git clean -x` covers gitignored build output, `-ff`
+        # nested git dirs; .git itself stays so the next checkout is
+        # incremental (fixed size, unlike the build output). The fallback
+        # covers a job whose checkout never completed, and a cleanup that
+        # still cannot finish degrades to a warning rather than turning an
+        # otherwise-green (and merge-gating) job red.
+        #
+        # Deliberately kept inline rather than in a shared script: backport
+        # runs replace the tree with a release branch that predates any new
+        # script, and cleanup must not depend on the tree it is cleaning.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   amber:
     # The amber job runs the cross-cutting Scala lints (scalafmtCheckAll,
@@ -185,7 +228,7 @@ jobs:
     if: ${{ inputs.run_amber }}
     strategy:
       matrix:
-        os: [ubuntu-latest]
+        os: ["${{ inputs.heavy_runner }}"]
         java-version: [17]
     runs-on: ${{ matrix.os }}
     env:
@@ -254,6 +297,14 @@ jobs:
         #       ./bin/licensing/generate_notice_binary.py amber/NOTICE-binary 
/tmp/dists/amber-*/lib --extras amber/NOTICE-binary-python
         run: |
           set -euo pipefail
+          # Self-heal before unpacking: on a classic runner /tmp survives the
+          # job, and a previous run's unpacked dist makes unzip stop on an
+          # interactive overwrite prompt (no stdin -> exit 1). The cleanup step
+          # at the end of every job removes this, but a job killed outright
+          # (OOM, runner restart) never gets there, so the entry point cannot
+          # assume a clean slate. Every `docker run --name` below is guarded
+          # the same way.
+          rm -rf /tmp/dists/amber-*
           mkdir -p /tmp/dists
           unzip -q amber/target/universal/amber-*.zip -d /tmp/dists/
 
@@ -352,6 +403,18 @@ jobs:
           flags: amber
           report_type: test_results
           fail_ci_if_error: false
+      - name: Clean up job artifacts
+        # Drops the sbt target/ trees, the dist zip, and the dist this job
+        # unpacked under /tmp (which also survives the job on a classic
+        # runner). The `services:` postgres is removed by the runner itself.
+        # See the frontend job above for the full rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          rm -rf /tmp/dists/amber-* /tmp/notice-amber.txt
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   amber-integration:
     # Runs Scala tests tagged @org.apache.texera.amber.tags.IntegrationTest —
@@ -374,7 +437,7 @@ jobs:
       # the docker image, macOS uses brew + the upstream
       # aarch64-apple-darwin lakekeeper tarball.
       matrix:
-        os: [ubuntu-latest, macos-latest]
+        os: ["${{ inputs.heavy_runner }}", macos-latest]
         java-version: [17]
     runs-on: ${{ matrix.os }}
     env:
@@ -452,6 +515,9 @@ jobs:
         # Linux docker image here.
         run: |
           if [ "$RUNNER_OS" = "Linux" ]; then
+            # Self-heal: a killed job can leave this container behind on a
+            # classic runner, and the name would then be taken.
+            docker rm -f -v postgres >/dev/null 2>&1 || true
             docker run -d --name postgres \
               -p 5432:5432 \
               -e POSTGRES_PASSWORD=postgres \
@@ -498,6 +564,7 @@ jobs:
         # S3-protocol surface that has been stable across releases.
         run: |
           if [ "$RUNNER_OS" = "Linux" ]; then
+            docker rm -f -v minio >/dev/null 2>&1 || true
             docker run -d --name minio --network host \
               -e MINIO_ROOT_USER=texera_minio \
               -e MINIO_ROOT_PASSWORD=password \
@@ -534,6 +601,7 @@ jobs:
               -e LAKEKEEPER__PG_DATABASE_URL_WRITE \
               -e LAKEKEEPER__PG_ENCRYPTION_KEY \
               vakamo/lakekeeper:${LAKEKEEPER_VERSION} migrate
+            docker rm -f -v lakekeeper >/dev/null 2>&1 || true
             docker run -d --name lakekeeper --network host \
               -e LAKEKEEPER__PG_DATABASE_URL_READ \
               -e LAKEKEEPER__PG_DATABASE_URL_WRITE \
@@ -678,11 +746,13 @@ jobs:
         # `amber` and runs in parallel (as platform-integration does vs
         # platform); the sbt compile is already warm from the integration-test
         # run above. ubuntu-only: the boot is pure-JVM and OS-independent.
-        if: matrix.os == 'ubuntu-latest'
+        if: runner.os == 'Linux'
         run: sbt "WorkflowExecutionService/dist"
       - name: Unzip amber dist
-        if: matrix.os == 'ubuntu-latest'
+        if: runner.os == 'Linux'
         run: |
+          # Self-heal: see the amber job's unzip step.
+          rm -rf /tmp/dists/amber-*
           mkdir -p /tmp/dists
           unzip -q amber/target/universal/amber-*.zip -d /tmp/dists/
       - name: Smoke-test texera-web boots
@@ -692,7 +762,7 @@ jobs:
         # working directory for a directory named `amber`; run from the 
checkout
         # root (the default) that resolves to ./amber and reads
         # amber/src/main/resources/web-config.yml (port 8080).
-        if: matrix.os == 'ubuntu-latest'
+        if: runner.os == 'Linux'
         env:
           # Quiet boot logs, same wiring as the test steps above. Safe here:
           # smoke-boot's verdict is LISTEN-based, never log-scraping (#6332).
@@ -709,7 +779,7 @@ jobs:
         # so no iceberg / S3 access. Config resolves via Utils.amberHomePath to
         # amber/src/main/resources/computing-unit-master-config.yml (port 
8085).
         # Reuses the amber dist built + unzipped above for the texera-web boot.
-        if: matrix.os == 'ubuntu-latest'
+        if: runner.os == 'Linux'
         env:
           # Quiet boot logs, same wiring as the test steps above. Safe here:
           # smoke-boot's verdict is LISTEN-based, never log-scraping (#6332).
@@ -737,6 +807,21 @@ jobs:
           flags: amber-integration
           report_type: test_results
           fail_ci_if_error: false
+      - name: Clean up job artifacts
+        # This job starts its infra with `docker run -d` rather than
+        # `services:`, so nothing removes those containers for it — and on a
+        # classic runner they keep running, holding ports 5432 / 9000 / 8181
+        # against the next job (whose own `docker run` would then fail on the
+        # name and the port). -v also drops the anonymous volumes postgres
+        # declares. See the frontend job above for the full rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          docker rm -f -v postgres minio lakekeeper >/dev/null 2>&1 || true
+          rm -rf /tmp/dists/amber-* /tmp/protoc.zip
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   platform:
     # Per-service build, test, and license check for the non-amber Scala
@@ -745,7 +830,7 @@ jobs:
     # cover every module in the amber job above, so this matrix skips them.
     if: ${{ inputs.run_platform }}
     name: ${{ format('platform{0} ({1})', inputs.job_name_suffix, 
matrix.service) }}
-    runs-on: ubuntu-latest
+    runs-on: ${{ inputs.medium_runner }}
     strategy:
       fail-fast: false
       matrix:
@@ -843,6 +928,8 @@ jobs:
         if: ${{ steps.module_check.outputs.present == 'true' }}
         run: |
           set -euo pipefail
+          # Self-heal: see the amber job's unzip step.
+          rm -rf /tmp/dists/${{ matrix.service }}-*
           mkdir -p /tmp/dists
           unzip -q ${{ matrix.service }}/target/universal/${{ matrix.service 
}}-*.zip -d /tmp/dists/
 
@@ -879,6 +966,17 @@ jobs:
           report_type: test_results
           disable_search: true
           fail_ci_if_error: false
+      - name: Clean up job artifacts
+        # Six matrix entries share one runner pool, each leaving its own sbt
+        # target/ tree and unpacked dist behind. See the frontend job above
+        # for the full rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          rm -rf /tmp/dists/${{ matrix.service }}-* "/tmp/notice-${{ 
matrix.service }}.txt"
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   platform-integration:
     # Boot smoke test for the platform services (mirrors amber-integration: an
@@ -892,7 +990,7 @@ jobs:
     # alone. Unit tests + coverage stay in `platform`. See #6273.
     if: ${{ inputs.run_platform_integration }}
     name: ${{ format('platform-integration{0} ({1})', inputs.job_name_suffix, 
matrix.service) }}
-    runs-on: ubuntu-latest
+    runs-on: ${{ inputs.medium_runner }}
     strategy:
       fail-fast: false
       matrix:
@@ -984,6 +1082,8 @@ jobs:
       - name: Unzip ${{ matrix.service }} dist
         if: ${{ steps.module_check.outputs.present == 'true' }}
         run: |
+          # Self-heal: see the amber job's unzip step.
+          rm -rf /tmp/dists/${{ matrix.service }}-*
           mkdir -p /tmp/dists
           unzip -q ${{ matrix.service }}/target/universal/${{ matrix.service 
}}-*.zip -d /tmp/dists/
       - name: Start MinIO
@@ -991,6 +1091,7 @@ jobs:
         # that service needs an object store, so the rest of the matrix skips 
this.
         if: ${{ matrix.object_store && steps.module_check.outputs.present == 
'true' }}
         run: |
+          docker rm -f -v minio >/dev/null 2>&1 || true
           docker run -d --name minio --network host \
             -e MINIO_ROOT_USER=texera_minio \
             -e MINIO_ROOT_PASSWORD=password \
@@ -1008,6 +1109,7 @@ jobs:
         # adapted to CI creds (postgres/postgres @ localhost).
         if: ${{ matrix.object_store && steps.module_check.outputs.present == 
'true' }}
         run: |
+          docker rm -f -v lakefs >/dev/null 2>&1 || true
           docker run -d --name lakefs --network host \
             -e LAKEFS_DATABASE_TYPE=postgres \
             -e 
"LAKEFS_DATABASE_POSTGRES_CONNECTION_STRING=postgres://postgres:postgres@localhost:5432/texera_lakefs?sslmode=disable"
 \
@@ -1035,12 +1137,25 @@ jobs:
           # smoke-boot's verdict is LISTEN-based, never log-scraping (#6332).
           TEXERA_SERVICE_LOG_LEVEL: ${{ runner.debug == '1' && 'DEBUG' || 
'WARN' }}
         run: .github/scripts/smoke-boot.sh "/tmp/dists/${{ matrix.service 
}}-*/bin/${{ matrix.service }}" "${{ matrix.port }}"
+      - name: Clean up job artifacts
+        # minio / lakefs are started with `docker run -d` for the file-service
+        # entry only, but removing them unconditionally is harmless (a missing
+        # container is ignored). See the frontend job above for the full
+        # rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          docker rm -f -v minio lakefs >/dev/null 2>&1 || true
+          rm -rf /tmp/dists/${{ matrix.service }}-*
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   pyamber:
     if: ${{ inputs.run_pyamber }}
     strategy:
       matrix:
-        os: [ubuntu-latest]
+        os: ["${{ inputs.medium_runner }}"]
         python-version: ["3.11", "3.12", "3.13"]
     runs-on: ${{ matrix.os }}
     services:
@@ -1152,6 +1267,17 @@ jobs:
           report_type: test_results
           disable_search: true
           fail_ci_if_error: false
+      - name: Clean up job artifacts
+        # Three python legs land on the same pool; each leaves generated proto
+        # bindings, .pytest_cache and coverage output in the workspace. See
+        # the frontend job above for the full rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          rm -rf /tmp/protoc.zip /tmp/pip-licenses.csv
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   agent-service:
     if: ${{ inputs.run_agent_service }}
@@ -1160,7 +1286,7 @@ jobs:
     strategy:
       fail-fast: false
       matrix:
-        os: [ubuntu-latest, macos-latest]
+        os: ["${{ inputs.light_runner }}", macos-latest]
         bun-version: ["1.3.3"]
     defaults:
       run:
@@ -1182,12 +1308,12 @@ jobs:
       - name: Install production dependencies
         run: bun install --production --frozen-lockfile
       - name: Generate agent-service license manifest
-        if: matrix.os == 'ubuntu-latest'
+        if: runner.os == 'Linux'
         run: |
           mkdir -p dist
           bun run bin/collect-licenses.ts > dist/3rdpartylicenses.json
       - name: Check bundled agent-service packages against per-module 
LICENSE-binary files
-        if: matrix.os == 'ubuntu-latest'
+        if: runner.os == 'Linux'
         run: ../bin/licensing/check_binary_deps.py ${{ inputs.mode == 'PR' && 
'--ignore-transitive-version' || '' }} agent-npm dist/3rdpartylicenses.json
       - name: Install development dependencies
         run: bun install --frozen-lockfile
@@ -1207,7 +1333,7 @@ jobs:
           TEXERA_SERVICE_LOG_LEVEL: ${{ runner.debug == '1' && 'DEBUG' || 
'WARN' }}
         run: bun test --coverage --coverage-reporter=lcov --reporter=junit 
--reporter-outfile=junit.xml
       - name: Upload agent-service coverage to Codecov
-        if: matrix.os == 'ubuntu-latest' && always()
+        if: runner.os == 'Linux' && always()
         uses: codecov/codecov-action@fb8b3582c8e4def4969c97caa2f19720cb33a72f 
# v7.0.0
         with:
           token: ${{ secrets.CODECOV_TOKEN }}
@@ -1217,7 +1343,7 @@ jobs:
       - name: Upload agent-service test results to Codecov
         # Test Analytics ingestion. Runs on the same ubuntu leg that
         # uploads coverage. `!cancelled()` so test failures still upload.
-        if: matrix.os == 'ubuntu-latest' && !cancelled()
+        if: runner.os == 'Linux' && !cancelled()
         uses: codecov/codecov-action@fb8b3582c8e4def4969c97caa2f19720cb33a72f 
# v7.0.0
         with:
           token: ${{ secrets.CODECOV_TOKEN }}
@@ -1226,6 +1352,16 @@ jobs:
           report_type: test_results
           disable_search: true
           fail_ci_if_error: false
+      - name: Clean up job artifacts
+        # working-directory overrides the job default (agent-service) so the
+        # clean covers the whole workspace, not just that subdirectory. See
+        # the frontend job above for the full rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   infra:
     # Generic project-tooling job — runs the lightweight (no docker, no
@@ -1242,7 +1378,7 @@ jobs:
     strategy:
       fail-fast: false
       matrix:
-        os: [ubuntu-latest, macos-latest]
+        os: ["${{ inputs.light_runner }}", macos-latest]
         python-version: ["3.12"]
     steps:
       - name: Checkout Texera
@@ -1288,6 +1424,15 @@ jobs:
         # pass, and any future tooling test suite under bin/ comes
         # along without this workflow being edited.
         run: python -m pytest bin/ -v --tb=short
+      - name: Clean up job artifacts
+        # Leftovers here are the checkout plus the pytest caches. See the
+        # frontend job above for the full rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
 
   pyright-language-service:
     # Typecheck-only job for the standalone Pyright LSP bridge (the language
@@ -1301,7 +1446,7 @@ jobs:
     # stays a fast, docker-free job.
     if: ${{ inputs.run_pyright_language_service }}
     name: ${{ format('pyright-language-service{0}', inputs.job_name_suffix) }}
-    runs-on: ubuntu-latest
+    runs-on: ${{ inputs.light_runner }}
     defaults:
       run:
         working-directory: pyright-language-service
@@ -1330,3 +1475,13 @@ jobs:
       - name: Typecheck
         # allowImportingTsExtensions in tsconfig.json requires --noEmit.
         run: yarn tsc --noEmit
+      - name: Clean up job artifacts
+        # working-directory overrides the job default 
(pyright-language-service)
+        # so the clean covers the whole workspace, node_modules/ included. See
+        # the frontend job above for the full rationale.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
diff --git a/.github/workflows/check-header.yml 
b/.github/workflows/check-header.yml
index 003747c6e1..55ad399d29 100644
--- a/.github/workflows/check-header.yml
+++ b/.github/workflows/check-header.yml
@@ -27,7 +27,25 @@ on:
 jobs:
   test:
     name: Check License Headers
-    runs-on: ubuntu-latest
+    # Control-plane resilience: merge-queue entries (and labeled PRs) run
+    # this job on arc-light-linux so a jammed GitHub-hosted pool cannot
+    # stall the pipeline. Availability cannot be probed before this job runs, 
so
+    # the repo variable ARC_AVAILABLE is the kill switch: set it to
+    # 'false' when the cluster is down to route back to ubuntu-latest.
+    runs-on: ${{ ((github.event_name == 'merge_group' || (github.event_name == 
'pull_request' && contains(github.event.pull_request.labels.*.name, 
'ci:self-hosted'))) && vars.ARC_AVAILABLE != 'false') && 'arc-light-linux' || 
'ubuntu-latest' }}
     steps:
       - uses: actions/checkout@v7
       - uses: apache/skywalking-eyes@61275cc80d0798a405cb070f7d3a8aaf7cf2c2c1 
# v0.8.0
+      - name: Clean up job artifacts
+        # Classic self-hosted runners hand this workspace to the next job
+        # as-is; drop whatever license-eye wrote into it. always() so a failed
+        # or cancelled job cleans up too; the GitHub-hosted fallback is a
+        # throwaway VM and skips it. .git stays so the next checkout is
+        # incremental; the fallback covers a job whose checkout never
+        # completed.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        working-directory: ${{ github.workspace }}
+        run: |
+          git clean -xdff ||
+            find . -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
diff --git a/.github/workflows/precheck.yml b/.github/workflows/precheck.yml
index 4ce1d24092..c9274af81b 100644
--- a/.github/workflows/precheck.yml
+++ b/.github/workflows/precheck.yml
@@ -50,8 +50,19 @@ on:
         value: ${{ jobs.decide.outputs.run_pyright_language_service }}
       backport_targets:
         value: ${{ jobs.decide.outputs.backport_targets }}
+      # Runner labels for Linux build jobs, three tiers: the self-hosted
+      # scale sets when the PR carries the ci:self-hosted label,
+      # GitHub-hosted ubuntu-latest otherwise.
+      heavy_runner:
+        value: ${{ jobs.decide.outputs.heavy_runner }}
+      medium_runner:
+        value: ${{ jobs.decide.outputs.medium_runner }}
+      light_runner:
+        value: ${{ jobs.decide.outputs.light_runner }}
 
 permissions:
+  # actions: read backs the queued-job backlog probe for runner routing.
+  actions: read
   checks: read
   # contents: read backs the git-data commit lookups (tree SHAs) in the
   # merge-queue fast path.
@@ -61,7 +72,12 @@ permissions:
 jobs:
   decide:
     name: Precheck
-    runs-on: ubuntu-latest
+    # Control-plane resilience: merge-queue entries (and labeled PRs) run
+    # this job on arc-light-linux so a jammed GitHub-hosted pool cannot
+    # stall the pipeline. Availability cannot be probed before this job runs, 
so
+    # the repo variable ARC_AVAILABLE is the kill switch: set it to
+    # 'false' when the cluster is down to route back to ubuntu-latest.
+    runs-on: ${{ ((github.event_name == 'merge_group' || (github.event_name == 
'pull_request' && contains(github.event.pull_request.labels.*.name, 
'ci:self-hosted'))) && vars.ARC_AVAILABLE != 'false') && 'arc-light-linux' || 
'ubuntu-latest' }}
     outputs:
       run_frontend: ${{ steps.decide.outputs.run_frontend }}
       run_amber: ${{ steps.decide.outputs.run_amber }}
@@ -73,6 +89,9 @@ jobs:
       run_infra: ${{ steps.decide.outputs.run_infra }}
       run_pyright_language_service: ${{ 
steps.decide.outputs.run_pyright_language_service }}
       backport_targets: ${{ steps.decide.outputs.backport_targets }}
+      heavy_runner: ${{ steps.decide.outputs.heavy_runner }}
+      medium_runner: ${{ steps.decide.outputs.medium_runner }}
+      light_runner: ${{ steps.decide.outputs.light_runner }}
     steps:
       - name: Wait for Pull Request Labeler
         if: github.event_name == 'pull_request'
@@ -101,6 +120,14 @@ jobs:
       - name: Decide which jobs to run
         id: decide
         uses: actions/github-script@v9
+        env:
+          # Fine-grained PAT with administration:read on this repo; used
+          # only to probe self-hosted runner availability. Optional — when
+          # absent the ci:self-hosted label is trusted without probing.
+          ARC_STATUS_TOKEN: ${{ secrets.ARC_STATUS_TOKEN }}
+          # Manual/auto kill switch mirrored by runner-heartbeat.yml's
+          # watchdog; 'false' forces every Linux tier back to GitHub-hosted.
+          ARC_AVAILABLE: ${{ vars.ARC_AVAILABLE }}
         with:
           script: |
             const eventName = context.eventName;
@@ -314,6 +341,184 @@ jobs:
               runPyrightLanguageService = false;
             }
 
+            // Runner tiers: merge queue entries default to the
+            // self-hosted scale sets (the queue is the latency-critical
+            // serialized path most hurt by hosted-pool congestion);
+            // PR events opt in via the ci:self-hosted label. Everything
+            // else (push / dispatch) stays GitHub-hosted. Tiers:
+            // arc-heavy-linux for the amber/sbt stacks, arc-medium-linux
+            // for platform and pyamber, arc-light-linux for the small
+            // service jobs.
+            //
+            // Fallback probing, two levels, evaluated per tier:
+            //
+            //  Level 1 (availability): a set with no online runner (cluster
+            //  down, scale set removed) falls back to GitHub-hosted so a
+            //  labeled PR never queues forever. Relies on each set keeping
+            //  warm runners (minRunners > 0) — an idle scale-set with
+            //  minRunners: 0 registers no runners and would look
+            //  unavailable.
+            //
+            //  Level 2 (congestion): predict queueing before it happens by
+            //  comparing this round's demand (the stack decisions above
+            //  tell us how many jobs each tier is about to receive)
+            //  against idle supply (runner busy flags) plus the backlog
+            //  of already-queued jobs waiting for the same label. Blind
+            //  spot: k8s-side capacity (Pending pods, scale-up headroom)
+            //  is invisible to the GitHub API.
+            //
+            // Runner listing needs ARC_STATUS_TOKEN (administration:read);
+            // without it the label is trusted as-is. Any probe error
+            // degrades to trusting the label.
+            const selfHosted = eventName === "merge_group" || 
labels.includes("ci:self-hosted");
+            const tiers = {
+              heavy_runner: selfHosted ? "arc-heavy-linux" : "ubuntu-latest",
+              medium_runner: selfHosted ? "arc-medium-linux" : "ubuntu-latest",
+              light_runner: selfHosted ? "arc-light-linux" : "ubuntu-latest",
+            };
+            // Per-tier job counts this round. PLATFORM_SERVICES mirrors the
+            // service matrix size in build.yml — keep in sync.
+            const PLATFORM_SERVICES = 6;
+            const demand = {
+              "arc-heavy-linux":
+                (runAmber ? 1 : 0) + (runAmberIntegration ? 1 : 0) + 
(runFrontend ? 1 : 0),
+              "arc-medium-linux":
+                (runPlatform ? PLATFORM_SERVICES : 0) +
+                (runPlatformIntegration ? PLATFORM_SERVICES : 0) +
+                (runPyamber ? 3 : 0),
+              "arc-light-linux":
+                (runInfra ? 1 : 0) +
+                (runAgentService ? 1 : 0) +
+                (runPyrightLanguageService ? 1 : 0),
+            };
+            // Kill switch: ARC_AVAILABLE=false (manual, or flipped by the
+            // runner-heartbeat watchdog) forces every tier to GitHub-hosted.
+            if (selfHosted && process.env.ARC_AVAILABLE === "false") {
+              for (const k of Object.keys(tiers)) tiers[k] = "ubuntu-latest";
+              core.warning("ARC_AVAILABLE=false; routing all Linux tiers to 
GitHub-hosted runners.");
+            }
+            // Heartbeat fallback: runner-heartbeat.yml proves fleet liveness
+            // every 15 minutes. Stale last-success (with history present)
+            // means the fleet is down — fall back without needing any extra
+            // credential (plain GITHUB_TOKEN, actions: read).
+            if (selfHosted && tiers.light_runner !== "ubuntu-latest") {
+              try {
+                const STALE_MS = 35 * 60 * 1000;
+                const ok = await github.rest.actions.listWorkflowRuns({
+                  owner: context.repo.owner,
+                  repo: context.repo.repo,
+                  workflow_id: "runner-heartbeat.yml",
+                  status: "success",
+                  per_page: 1,
+                });
+                const any = await github.rest.actions.listWorkflowRuns({
+                  owner: context.repo.owner,
+                  repo: context.repo.repo,
+                  workflow_id: "runner-heartbeat.yml",
+                  per_page: 1,
+                });
+                const hasHistory = (any.data.workflow_runs ?? []).length > 0;
+                const last = ok.data.workflow_runs?.[0];
+                const lastOk = last ? new Date(last.updated_at).getTime() : 0;
+                if (hasHistory && Date.now() - lastOk > STALE_MS) {
+                  for (const k of Object.keys(tiers)) tiers[k] = 
"ubuntu-latest";
+                  core.warning("Runner heartbeat is stale; routing all Linux 
tiers to GitHub-hosted runners.");
+                }
+              } catch (e) {
+                core.info(`Heartbeat check skipped (${e.message}).`);
+              }
+            }
+            if (selfHosted && process.env.ARC_STATUS_TOKEN) {
+              try {
+                // Raw request rather than
+                // github.rest.actions.listSelfHostedRunnersForRepo: the
+                // `github` client is bound to the workflow GITHUB_TOKEN, and
+                // Octokit's auth hook overwrites the authorization header on
+                // every request, so one client cannot carry a second token.
+                // Listing runners needs administration:read, which only
+                // ARC_STATUS_TOKEN has. The backlog probe below stays on the
+                // SDK because actions:read is within the workflow token's
+                // reach — the split follows token scope, not style.
+                //
+                // Giving this call the SDK back means moving it into its own
+                // github-script step with `github-token: ARC_STATUS_TOKEN`,
+                // which also buys `retries` and github.paginate (this request
+                // caps at 100 runners today). Worth doing once the fleet
+                // approaches that size; see the discussion on #7172.
+                const res = await fetch(
+                  
`https://api.github.com/repos/${context.repo.owner}/${context.repo.repo}/actions/runners?per_page=100`,
+                  {
+                    headers: {
+                      authorization: `Bearer ${process.env.ARC_STATUS_TOKEN}`,
+                      accept: "application/vnd.github+json",
+                    },
+                  }
+                );
+                if (!res.ok) throw new Error(`runner list HTTP ${res.status}`);
+                const online = {};
+                const idle = {};
+                for (const r of (await res.json()).runners ?? []) {
+                  if (r.status !== "online") continue;
+                  for (const l of r.labels) {
+                    online[l.name] = (online[l.name] ?? 0) + 1;
+                    if (!r.busy) idle[l.name] = (idle[l.name] ?? 0) + 1;
+                  }
+                }
+                // Backlog: queued jobs per label across currently queued
+                // runs (capped sample; uses the workflow GITHUB_TOKEN).
+                const backlog = {};
+                try {
+                  const { data: queued } = await 
github.rest.actions.listWorkflowRunsForRepo({
+                    owner: context.repo.owner,
+                    repo: context.repo.repo,
+                    status: "queued",
+                    per_page: 10,
+                  });
+                  for (const run of queued.workflow_runs ?? []) {
+                    const { data: jobsData } = await 
github.rest.actions.listJobsForWorkflowRun({
+                      owner: context.repo.owner,
+                      repo: context.repo.repo,
+                      run_id: run.id,
+                      per_page: 100,
+                    });
+                    for (const job of jobsData.jobs ?? []) {
+                      if (job.status !== "queued") continue;
+                      for (const l of job.labels) backlog[l] = (backlog[l] ?? 
0) + 1;
+                    }
+                  }
+                } catch (e) {
+                  core.warning(`Backlog probe failed (${e.message}); assuming 
no backlog.`);
+                }
+                for (const [output, scaleSet] of Object.entries({ ...tiers })) 
{
+                  if (scaleSet === "ubuntu-latest") continue;
+                  const nOnline = online[scaleSet] ?? 0;
+                  const nIdle = idle[scaleSet] ?? 0;
+                  const nBacklog = backlog[scaleSet] ?? 0;
+                  const nDemand = demand[scaleSet] ?? 0;
+                  let reason = null;
+                  if (nOnline === 0) {
+                    reason = "no online runner";
+                  } else if (nIdle === 0 && nBacklog >= 2) {
+                    reason = `all ${nOnline} runners busy with ${nBacklog} 
jobs already queued`;
+                  } else if (nDemand > nIdle + 2) {
+                    reason = `demand ${nDemand} exceeds idle ${nIdle} + slack`;
+                  }
+                  if (reason) {
+                    tiers[output] = "ubuntu-latest";
+                    core.warning(`${scaleSet}: ${reason}; falling back to 
ubuntu-latest.`);
+                  } else {
+                    core.info(`${scaleSet}: online=${nOnline} idle=${nIdle} 
backlog=${nBacklog} demand=${nDemand} — ok`);
+                  }
+                }
+              } catch (e) {
+                core.warning(`Runner availability probe failed (${e.message}); 
trusting the label.`);
+              }
+            } else if (selfHosted) {
+              core.info("ARC_STATUS_TOKEN not set; trusting the ci:self-hosted 
label without probing.");
+            }
+            for (const [output, runner] of Object.entries(tiers)) 
core.setOutput(output, runner);
+            core.info(`Linux runners: heavy=${tiers.heavy_runner} 
medium=${tiers.medium_runner} light=${tiers.light_runner}`);
+
             core.setOutput("run_frontend", runFrontend ? "true" : "false");
             core.setOutput("run_amber", runAmber ? "true" : "false");
             core.setOutput("run_amber_integration", runAmberIntegration ? 
"true" : "false");
@@ -344,3 +549,12 @@ jobs:
               core.info(`Backport targets: ${targets.join(", ")}`);
             }
             core.setOutput("backport_targets", JSON.stringify(targets));
+
+      - name: Clean up job artifacts
+        # No checkout here, but a classic self-hosted runner still keeps the
+        # workspace directory it created for this job; leave it empty for the
+        # next one. always() so a failed or cancelled job cleans up too.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        run: |
+          find "${GITHUB_WORKSPACE:?}" -mindepth 1 -maxdepth 1 -exec rm -rf -- 
{} + 2>/dev/null ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
diff --git a/.github/workflows/required-checks.yml 
b/.github/workflows/required-checks.yml
index 1db8a52668..b3a6895046 100644
--- a/.github/workflows/required-checks.yml
+++ b/.github/workflows/required-checks.yml
@@ -33,6 +33,9 @@ on:
   workflow_dispatch:
 
 permissions:
+  # actions: read is required by the called precheck.yml (backlog probe);
+  # a reusable workflow's permissions must be a subset of its caller's.
+  actions: read
   checks: write
   contents: read
   pull-requests: read
@@ -56,6 +59,9 @@ jobs:
     needs: precheck
     uses: ./.github/workflows/build.yml
     with:
+      heavy_runner: ${{ needs.precheck.outputs.heavy_runner }}
+      medium_runner: ${{ needs.precheck.outputs.medium_runner }}
+      light_runner: ${{ needs.precheck.outputs.light_runner }}
       run_frontend: ${{ needs.precheck.outputs.run_frontend == 'true' }}
       run_amber: ${{ needs.precheck.outputs.run_amber == 'true' }}
       run_amber_integration: ${{ needs.precheck.outputs.run_amber_integration 
== 'true' }}
@@ -72,7 +78,9 @@ jobs:
     name: Required Checks
     needs: [precheck, build]
     if: always()
-    runs-on: ubuntu-latest
+    # Control-plane resilience: follow precheck's probed routing (label
+    # present AND the scale set has online runners); ubuntu-latest otherwise.
+    runs-on: ${{ needs.precheck.outputs.light_runner == 'arc-light-linux' && 
'arc-light-linux' || 'ubuntu-latest' }}
     steps:
       - name: Verify all required checks succeeded or were skipped
         run: |
@@ -93,3 +101,12 @@ jobs:
             exit 1
           fi
           echo "All required checks succeeded or were skipped."
+      - name: Clean up job artifacts
+        # No checkout here, but a classic self-hosted runner still keeps the
+        # workspace directory it created for this job; leave it empty for the
+        # next one. always() so the aggregator cleans up even when it reports
+        # a failure above.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        run: |
+          find "${GITHUB_WORKSPACE:?}" -mindepth 1 -maxdepth 1 -exec rm -rf -- 
{} + 2>/dev/null ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
diff --git a/.github/workflows/runner-heartbeat.yml 
b/.github/workflows/runner-heartbeat.yml
new file mode 100644
index 0000000000..9eb13dd446
--- /dev/null
+++ b/.github/workflows/runner-heartbeat.yml
@@ -0,0 +1,199 @@
+# Licensed to the Apache Software Foundation (ASF) under one
+# or more contributor license agreements.  See the NOTICE file
+# distributed with this work for additional information
+# regarding copyright ownership.  The ASF licenses this file
+# to you under the Apache License, Version 2.0 (the
+# "License"); you may not use this file except in compliance
+# with the License.  You may obtain a copy of the License at
+#
+#   http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing,
+# software distributed under the License is distributed on an
+# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+# KIND, either express or implied.  See the License for the
+# specific language governing permissions and limitations
+# under the License.
+
+# Liveness and health probe for the self-hosted runner fleet. `beat` runs on
+# the fleet every 15 minutes and, beyond proving the node is reachable,
+# checks the things that make a classic (non-ephemeral) runner unusable
+# without making it unreachable: disk and inode headroom, a responsive docker
+# daemon, and the tools every build job assumes. A stale last-success means
+# the fleet is down OR unhealthy. Two consumers act on that signal:
+#   - precheck.yml queries the last successful heartbeat directly (plain
+#     GITHUB_TOKEN) and falls back to GitHub-hosted runners for the build
+#     tiers when it is stale.
+#   - `watchdog` mirrors the verdict into the ARC_AVAILABLE repository
+#     variable, which the first-hop runs-on expressions (unable to call
+#     APIs) consult. The workflow GITHUB_TOKEN cannot write variables
+#     ("Resource not accessible by integration"), so watchdog uses
+#     AUTO_MERGE_TOKEN (collaborator suffices for the Variables API) and
+#     degrades to log-only when the secret is absent.
+name: Runner Heartbeat
+
+on:
+  schedule:
+    - cron: "*/15 * * * *"
+  workflow_dispatch:
+
+permissions:
+  actions: read
+
+concurrency:
+  group: runner-heartbeat
+  # A beat stuck queueing on a dead fleet is superseded by the next tick
+  # instead of piling up.
+  cancel-in-progress: true
+
+jobs:
+  beat:
+    runs-on: arc-light-linux
+    timeout-minutes: 5
+    steps:
+      - name: Prove the fleet is alive
+        run: echo "self-hosted fleet alive at $(date -u +%FT%TZ) on 
$RUNNER_NAME"
+      - name: Check node health
+        # Reachable is not the same as healthy. On a classic (non-ephemeral)
+        # runner the node keeps its disk and its docker daemon across every
+        # job it has ever run, so it can answer this heartbeat while being
+        # unable to complete any real build — a full disk, a wedged docker
+        # daemon or a tool that went missing all surface as job failures that
+        # read like code problems. Failing here is the point: the watchdog
+        # below flips ARC_AVAILABLE to false and CI routes back to
+        # GitHub-hosted runners until the node is fixed.
+        #
+        # Scope: this beat runs on one arc-light-linux runner, so what it
+        # proves about disk is that node's disk, not the heavy/medium/bench
+        # tiers'. Turning `beat` into a matrix over all four labels is the
+        # natural extension if the tiers start diverging.
+        run: |
+          set -uo pipefail
+          status=0
+          err()  { echo "::error::$1"; status=1; }
+          warn() { echo "::warning::$1"; }
+
+          # Disk is the realistic killer: build output accumulates until the
+          # root volume fills and the node starts evicting. / and the work
+          # directory can be separate volumes, so check both.
+          for mount in / "${GITHUB_WORKSPACE:-/}"; do
+            free_gib=$(($(df -Pk "$mount" | awk 'NR==2 {print $4}') / 1024 / 
1024))
+            used_pct=$(df -Pk "$mount" | awk 'NR==2 {gsub(/%/,"",$5); print 
$5}')
+            echo "disk $mount: ${free_gib}GiB free, ${used_pct}% used"
+            if [ "$free_gib" -lt 5 ]; then
+              err "disk $mount is nearly full: ${free_gib}GiB free 
(${used_pct}% used)"
+            elif [ "$free_gib" -lt 15 ]; then
+              warn "disk $mount is filling up: ${free_gib}GiB free 
(${used_pct}% used)"
+            fi
+            # Inodes can run out with disk to spare (millions of small files
+            # from node_modules / coursier). Some filesystems report '-'.
+            inode_pct=$(df -Pi "$mount" 2>/dev/null | awk 'NR==2 
{gsub(/%/,"",$5); print $5}')
+            case "$inode_pct" in
+              ''|*[!0-9]*) ;;
+              *) if [ "$inode_pct" -ge 90 ]; then warn "inode usage on $mount 
is ${inode_pct}%"; fi ;;
+            esac
+          done
+
+          # Tools every build job assumes. docker is absent by design on some
+          # images (the light tier runs no containers), so it only warns.
+          for tool in git curl unzip tar; do
+            command -v "$tool" >/dev/null 2>&1 || err "required tool missing: 
$tool"
+          done
+          if command -v docker >/dev/null 2>&1; then
+            if docker info >/dev/null 2>&1; then
+              running=$(docker ps -q | wc -l | tr -d ' ')
+              total=$(docker ps -aq | wc -l | tr -d ' ')
+              echo "docker: daemon ok, ${running} running / ${total} total 
container(s)"
+              # Leftovers here mean jobs are dying before their cleanup step;
+              # the named ones (postgres, minio, lakekeeper, lakefs) would then
+              # collide with the next job that wants the same name and port.
+              if [ "$total" -gt 5 ]; then
+                warn "${total} containers left on this node; job cleanup is 
not keeping up"
+              fi
+            else
+              err "docker CLI is present but the daemon is not responding"
+            fi
+          else
+            warn "docker is not available on this runner"
+          fi
+
+          if [ "$status" -ne 0 ]; then
+            echo "node is unhealthy — failing the heartbeat so CI falls back 
to GitHub-hosted runners"
+            exit 1
+          fi
+          echo "node healthy"
+      - name: Clean up job artifacts
+        # The beat checks nothing out, but it runs every 15 minutes on a
+        # classic runner, so even an empty workspace directory is worth
+        # leaving clean. always() so a failed or cancelled beat cleans up too.
+        if: ${{ always() && runner.environment == 'self-hosted' }}
+        run: |
+          find "${GITHUB_WORKSPACE:?}" -mindepth 1 -maxdepth 1 -exec rm -rf -- 
{} + 2>/dev/null ||
+            echo "::warning::workspace cleanup incomplete on 
${RUNNER_NAME:-this runner}"
+
+  watchdog:
+    runs-on: ubuntu-slim
+    timeout-minutes: 5
+    steps:
+      - name: Mirror fleet health into ARC_AVAILABLE
+        uses: actions/github-script@v9
+        env:
+          VAR_TOKEN: ${{ secrets.AUTO_MERGE_TOKEN }}
+        with:
+          script: |
+            const STALE_MS = 35 * 60 * 1000;
+            const { data } = await github.rest.actions.listWorkflowRuns({
+              owner: context.repo.owner,
+              repo: context.repo.repo,
+              workflow_id: "runner-heartbeat.yml",
+              status: "success",
+              per_page: 1,
+            });
+            const last = data.workflow_runs?.[0];
+            const lastOk = last ? new Date(last.updated_at).getTime() : 0;
+            const healthy = Date.now() - lastOk <= STALE_MS;
+            core.info(
+              `fleet ${healthy ? "healthy" : "STALE"}; last successful 
heartbeat: ` +
+              (lastOk ? new Date(lastOk).toISOString() : "never")
+            );
+            if (!process.env.VAR_TOKEN) {
+              core.warning("AUTO_MERGE_TOKEN not set; cannot update 
ARC_AVAILABLE (log-only).");
+              return;
+            }
+            const desired = healthy ? "true" : "false";
+            // Raw requests rather than github.rest.actions.{get,create,update}
+            // RepoVariable: the `github` client above is bound to the workflow
+            // GITHUB_TOKEN, which cannot write variables, and Octokit's auth
+            // hook overwrites the authorization header on every request, so it
+            // cannot carry AUTO_MERGE_TOKEN. Unlike the runner probe in
+            // precheck.yml, every call in this step could use that token, so
+            // moving the whole step onto `github-token: AUTO_MERGE_TOKEN`
+            // would put all three back on the SDK.
+            const headers = {
+              authorization: `Bearer ${process.env.VAR_TOKEN}`,
+              accept: "application/vnd.github+json",
+            };
+            const base = 
`https://api.github.com/repos/${context.repo.owner}/${context.repo.repo}/actions/variables`;
+            const cur = await fetch(`${base}/ARC_AVAILABLE`, { headers });
+            if (cur.status === 404) {
+              const res = await fetch(base, {
+                method: "POST",
+                headers,
+                body: JSON.stringify({ name: "ARC_AVAILABLE", value: desired 
}),
+              });
+              if (!res.ok) throw new Error(`create ARC_AVAILABLE: HTTP 
${res.status}`);
+              core.notice(`ARC_AVAILABLE created as ${desired}`);
+            } else if (cur.ok) {
+              const { value } = await cur.json();
+              if (value !== desired) {
+                const res = await fetch(`${base}/ARC_AVAILABLE`, {
+                  method: "PATCH",
+                  headers,
+                  body: JSON.stringify({ name: "ARC_AVAILABLE", value: desired 
}),
+                });
+                if (!res.ok) throw new Error(`update ARC_AVAILABLE: HTTP 
${res.status}`);
+                core.notice(`ARC_AVAILABLE flipped to ${desired}`);
+              }
+            } else {
+              throw new Error(`read ARC_AVAILABLE: HTTP ${cur.status}`);
+            }

Reply via email to