Hi Ayush, thanks for looking into this. We could also move some of our CI workflows to ubuntu-slim. It has following limits per PR https://github.com/apache/airflow/pull/72852:
ubuntu-slim's constraints, verified against the runner image rather than > assumed: > > 1. No Docker daemon — client only, unprivileged, no DinD. Everything > touching Breeze, image builds or Docker-container actions stays put. > 2. python3 is Ubuntu's system interpreter — PEP 668 externally > managed, so a bare pip install fails. > 3. Minimal toolset — git, gh, jq, node, curl present; no Java, no Go. > actions/setup-* still works (tool-cache downloads). > 4. One core — jobs doing real work need timeout headroom. > > Yufei On Wed, Sep 9, 2026 at 5:57 PM Ayush Saxena <[email protected]> wrote: > Hi Yufei, > I tried to grab some information around it, the best we have around per > project is this report [1] from Apace Infra, It shows the usage to people, > if they are part of that project. > > Regarding the queue, it is not per project it is at ASF level, the response > ``` > there is no per-project budgeting. One giant pool, and everybody dirties > it. the infra-reports is the best available, given limited introspection > available from GitHub.com > the short answer is that all the runner queues are full, despite our large > queue slot allocation > ``` > > I even while lurking around found some projects are trying to move to > ubuntu-slim for jobs which are small and takes less time <15mins to avoid > queuing ubuntu-latest like here: > https://github.com/apache/magpie/pull/1190 > https://github.com/apache/airflow/pull/72852 > > So, maybe we can also move a few of our jobs to ubuntu-slim, the reasoning > mentioned to move to ubuntu-slim was that no job can run on it more than 15 > mins and chances of runners blocked by some huge job isn't possible because > it hard aborts at 15 minutes. > > There is an Infra script added here [2] for watching out who is using how > much in case of such outages, but nothing around to solve it yet. > > That is the best I could find :-) > > -Ayush > > [1] https://infra-reports.apache.org/#ghactions > [2] https://github.com/apache/infrastructure-actions/pull/1259 > > On 2026/09/09 21:11:00 Yufei Gu wrote: > > Hi all, > > > > Polaris GitHub Actions workflows are experiencing long queue times, > > affecting multiple PRs and main-branch CI today (Sept. 9). Eight PR CI > runs > > were queued and two were still in progress. > > > > One main-branch CI run waited roughly four hours before its initial jobs > > started. It now has 10 successful jobs, but another 10 have been waiting > > for runners for about 27 minutes: > > https://github.com/apache/polaris/actions/runs/34377371899 > > > > Similar queueing was also visible in other Apache projects earlier today. > > ASF shared runner capacity may be a factor, but I haven’t confirmed > whether > > the organization’s concurrency limit is currently reached. > > > > Is anyone aware of an ongoing Infra issue, or able to help confirm the > > current shared runner capacity? > > > > Thanks, > > Yufei > > >
