Thanks for bringing that up! You have my +1 :) I've already implemented some of the guardrails as part of the EKS solution of AIP-118 : ephemeral runners, runner replica ceiling, and node capacity ceiling. Depending on the actual implementation we choose for AIP-118 (EKS/RunsOn/another alternative), we'll adjust the appropriate monitoring and alerting tooling accordingly - but overall it seems that suggested tools will become handy. Regarding your question - I support limiting in the first stage to branches in the upstream repository only. Later on, after ensuring that we're good in terms of security and approval procedure - we could expand to fork branches.
Shahar On Fri, Sep 18, 2026 at 4:22 PM Jarek Potiuk <[email protected]> wrote: > Hi all, > > We run a dedicated AWS account for the project. Today it hosts the > documentation site, and we will soon be running self-hosted CI runners > on EKS there and possibly an Airflow instance :). > > Right now that account has no cost guardrails, no anomaly monitoring > and no automated spend controls. I would like to fix that before the > runners go live rather than after, and I would like some feedback on > the approach. > > The concern > > Self-hosted runners attached to a public repository are a well-known > abuse target. A workflow submitted from a fork can run attacker > controlled code on project-funded infrastructure, and mining is the > usual payload. We get a lot of fork PRs, which is normally a strength, > but combined with self-hosted runners and no guardrails it becomes an > exposure. > > The approach > > Two things shaped the design more than anything technical. > > First, we have no on-call rotation and we are not going to have one. > Alerts land in a channel and get read when somebody is around. That is > just the reality of a volunteer project. The consequence is that > prevention beats detection: an alert saying a miner is eating the fleet > is not worth much at 03:00 if it gets read at 09:00. A hard job timeout, > a runner replica ceiling and a budget action all act without a human. > > Second, and slightly against the title of this thread: anomaly > detection is the wrong primary tool for abuse. CloudWatch trains its > model on a trailing window, so sustained abuse gets learned as the new > normal within about two weeks, and the alarm goes quiet exactly when the > problem has become entrenched. Anomaly bands are good at catching > changes and bad at catching persistent conditions. So the proposal puts > hard caps first and uses anomaly detection as a secondary wide net. > > So the shape is: autonomous guardrails first, static thresholds for > things that should be zero, anomaly bands only where there is real > seasonality to learn (CI load and docs traffic both have a weekly > rhythm; error rates do not). > > Phasing > > Phase 0 needs no EKS and could be done this week: AWS Cost Anomaly > Detection, a budget with an action attached, alerts to a channel. Cost > Anomaly Detection is free, and since the account is Airflow-only, > anything anomalous in the bill is ours by definition. > > Phase 1 covers the docs site: a few CloudFront alarms plus a synthetic > canary. The canary matters more than it looks, because error-rate > alarms cannot see "site is up and serving a broken docs build" - every > request returns 200 and every dashboard stays green. > > Phase 2 is the runner fleet, and the guardrails there are a prerequisite > for enabling runners, not a follow-up. > > The question I most want input on > > Should fork PRs run on self-hosted runners at all? > > The safest topology is fork PRs on GitHub-hosted runners, with > self-hosted reserved for post-merge and trusted contributors. That is > also the least convenient one and it has a cost implication given our > PR volume. If we do want fork PRs on self-hosted runners, then at > minimum we need the approval requirement for outside and first-time > contributors turned on. > > There will be another proposal on how to deal with PR overload > related to it. However, at this stage, we just need to agree on > monitoring and alerting > before we set everything up is a good idea. > > I have a fuller draft written up covering the metric-by-metric > decisions, alternatives considered and cost implications. > > AIP proposal here: > > https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451974695/AIP-119+Anomaly+monitoring+and+guardrails+for+the+Airflow+AWS+account > > A few things I could not settle on my own: our baseline monthly spend, > whether the docs site is S3 behind CloudFront or plain S3 hosting, and > whether we are going with Actions Runner Controller for the runners. > > But that will come after we implement AIP-118 possibly. > > J. > > --------------------------------------------------------------------- > To unsubscribe, e-mail: [email protected] > For additional commands, e-mail: [email protected] > >
