On Thu, Sep 24, 2026 at 11:47:04AM -0700, Shakeel Butt wrote:

> Options we looked at
> ====================
> 
> A. Keep the shared kworker, and only charge the cgroup.
> 
>   A1. Measure how long the reclaim ran, then charge that time to the
>       memcg's cgroup.
> 
>   A2. Let the kworker say "charge this cgroup" while it works, the
>       same way set_active_memcg() works for memory.
> 
> B. Also make the cgroup's CPU limits apply.
> 
>   B1. Run the kworker in the cgroup's scheduling group while it works,
>       so cpu.weight and cpu.max apply to it.
> 
>   B2. Run the reclaim at full speed, then take its CPU time out of the
>       cgroup's cpu.max quota afterwards (back-charging).
> 
> C. Do the reclaim inside the cgroup.
> 
>   C1. Record the overage as a debt on the memcg, and let the memcg's
>       own tasks pay it the next time they charge memory or return to
>       user space.
> 
>   C2. Give each memcg its own reclaim thread that lives in the cgroup.
> 
>   C3. Use a cgroup-aware workqueue, as in [1], or per-cgroup worker
>       pools.

> What we chose and why
> =====================
> 
> We chose A2, together with B2 for cpu.max.
> 
> We do not throttle the reclaim itself. Adding limits to it is more
> complicated and most probably unneeded, as we envision that we will
> need concurrent background reclaimers instead of throttling in a
> real-world environment. We are working on developing a system to
> balance the rate of allocations/charges with the rate of reclaim, to
> keep the system always running effectively.
> 
> In addition, reclaim takes sleeping locks, like i_mmap_rwsem and the
> anon_vma lock in rmap walks, and fs locks in shrinkers. With a low
> cgroup weight, the kworker can be preempted while it holds them, and
> then tasks in other cgroups wait on it. It also hurts the workqueue. A
> kworker that is runnable but not running still counts as running, and
> the workqueue only spots CPU hogs by the CPU time they use. So other
> work queued on that CPU's system_wq waits too.

With proxy execution, this should be alleviated. Then all tasks waiting
on a lock acquisition will contribute to the runnability of the lock
holder.

> Still, the reclaim should not be free CPU time on top of the cgroup's
> limit. With A2 alone, the reclaim shows up in cpu.stat, but the
> cgroup's own tasks still get their full cpu.max quota. B2 fixes that
> without slowing the reclaim down. The kworker runs at full speed, and
> afterwards its time is taken out of the cgroup's cpu.max quota, so the
> cgroup's own tasks get less CPU time instead. This is the
> back-charging Tejun described [2]. It only covers cpu.max.
> Back-charging cpu.weight is future work.

Fiddling with weight sounds like horrible garbage. If you find you need
that, I would really rather you went C[23].

> We did not take C2 or C3. A high_work run asks for only 64 pages,
> which is far too little to pay for moving a thread into a cgroup
> through the global cgroup locks. A thread per memcg means thousands of
> threads. cgroup v2 does not allow tasks in a non-leaf domain cgroup.
> And a kernel thread in a cgroup shows up in cgroup.procs and blocks
> rmdir.

You don't move the threads, you spawn them on group creation and leave
them there. I'm sure you can fudge the rmdir thing with less ugly than
you're proposing here and in the other series.



Reply via email to