Thanks for bringing that up! You have my +1 :)

I've already implemented some of the guardrails as part of the EKS solution
of AIP-118 :
ephemeral runners, runner replica ceiling, and node capacity ceiling.
Depending on the actual implementation we choose for AIP-118
(EKS/RunsOn/another alternative), we'll adjust the appropriate monitoring
and alerting tooling accordingly - but overall it seems that suggested
tools will become handy.
Regarding your question - I support limiting in the first stage to branches
in the upstream repository only. Later on, after ensuring that we're good
in terms of security and approval procedure - we could expand to fork
branches.


Shahar


On Fri, Sep 18, 2026 at 4:22 PM Jarek Potiuk <[email protected]> wrote:

> Hi all,
>
> We run a dedicated AWS account for the project. Today it hosts the
> documentation site, and we will soon be running self-hosted CI runners
> on EKS there and possibly an Airflow instance :).
>
> Right now that account has no cost guardrails, no anomaly monitoring
> and no automated spend controls. I would like to fix that before the
> runners go live rather than after, and I would like some feedback on
> the approach.
>
> The concern
>
> Self-hosted runners attached to a public repository are a well-known
> abuse target. A workflow submitted from a fork can run attacker
> controlled code on project-funded infrastructure, and mining is the
> usual payload. We get a lot of fork PRs, which is normally a strength,
> but combined with self-hosted runners and no guardrails it becomes an
> exposure.
>
> The approach
>
> Two things shaped the design more than anything technical.
>
> First, we have no on-call rotation and we are not going to have one.
> Alerts land in a channel and get read when somebody is around. That is
> just the reality of a volunteer project. The consequence is that
> prevention beats detection: an alert saying a miner is eating the fleet
> is not worth much at 03:00 if it gets read at 09:00. A hard job timeout,
> a runner replica ceiling and a budget action all act without a human.
>
> Second, and slightly against the title of this thread: anomaly
> detection is the wrong primary tool for abuse. CloudWatch trains its
> model on a trailing window, so sustained abuse gets learned as the new
> normal within about two weeks, and the alarm goes quiet exactly when the
> problem has become entrenched. Anomaly bands are good at catching
> changes and bad at catching persistent conditions. So the proposal puts
> hard caps first and uses anomaly detection as a secondary wide net.
>
> So the shape is: autonomous guardrails first, static thresholds for
> things that should be zero, anomaly bands only where there is real
> seasonality to learn (CI load and docs traffic both have a weekly
> rhythm; error rates do not).
>
> Phasing
>
> Phase 0 needs no EKS and could be done this week: AWS Cost Anomaly
> Detection, a budget with an action attached, alerts to a channel. Cost
> Anomaly Detection is free, and since the account is Airflow-only,
> anything anomalous in the bill is ours by definition.
>
> Phase 1 covers the docs site: a few CloudFront alarms plus a synthetic
> canary. The canary matters more than it looks, because error-rate
> alarms cannot see "site is up and serving a broken docs build" - every
> request returns 200 and every dashboard stays green.
>
> Phase 2 is the runner fleet, and the guardrails there are a prerequisite
> for enabling runners, not a follow-up.
>
> The question I most want input on
>
> Should fork PRs run on self-hosted runners at all?
>
> The safest topology is fork PRs on GitHub-hosted runners, with
> self-hosted reserved for post-merge and trusted contributors. That is
> also the least convenient one and it has a cost implication given our
> PR volume. If we do want fork PRs on self-hosted runners, then at
> minimum we need the approval requirement for outside and first-time
> contributors turned on.
>
> There will be another proposal on how to deal with PR overload
> related to it. However, at this stage, we just need to agree on
> monitoring and alerting
> before we set everything up is a good idea.
>
> I have a fuller draft written up covering the metric-by-metric
> decisions, alternatives considered and cost implications.
>
> AIP proposal here:
>
> https://cwiki.apache.org/confluence/spaces/AIRFLOW/pages/451974695/AIP-119+Anomaly+monitoring+and+guardrails+for+the+Airflow+AWS+account
>
> A few things I could not settle on my own: our baseline monthly spend,
> whether the docs site is S3 behind CloudFront or plain S3 hosting, and
> whether we are going with Actions Runner Controller for the runners.
>
> But that will come after we implement AIP-118 possibly.
>
> J.
>
> ---------------------------------------------------------------------
> To unsubscribe, e-mail: [email protected]
> For additional commands, e-mail: [email protected]
>
>

Reply via email to