On Mon, Sep 14, 2026 at 08:17:09PM +0200, Kumar Kartikeya Dwivedi wrote: > On Fri Sep 11, 2026 at 4:20 AM CEST, Hui Zhu wrote: > > From: Hui Zhu <[email protected]> > > > > BPF programs can observe memory pressure on a cgroup (e.g. refault > > stats via bpf_mem_cgroup_page_state()), but cannot act on it: > > triggering reclaim on a chosen cgroup requires writing to > > memory.reclaim, which BPF cannot do. Add bpf_proactive_reclaim(), > > a sleepable kfunc which performs one proactive reclaim pass on a > > given memory cgroup, similar to a write to memory.reclaim but > > without retrying until the target is reached, so that when and how > > hard to reclaim is BPF policy rather than hard-coded thresholds. > > > > Since some bpf program types may be invoked while holding fs locks, > > limit the kfunc to BPF_PROG_TYPE_SYSCALL only, to avoid deadlocking > > in filesystem shrinkers on the reclaim path. A SYSCALL program can > > invoke the kfunc directly, or asynchronously from its bpf_wq or > > task_work callbacks, which run in process context and keep the > > SYSCALL program type. > > > > The reclaim target of a single call is capped at MEMCG_CHARGE_BATCH, > > following the precedent of high_work_func(), the memory.high > > workqueue fallback. Note that only the reclaim target is capped: the > > actual scanning work and its duration are not bounded. Reclaiming more > > than one batch is left to the BPF program rather than enforced by the > > kfunc: with one call per bpf_wq callback and the same work item > > requeued for the next batch, the program can also stop submitting > > batches in between, e.g. once the target cgroup is dying. > > > > Signed-off-by: Hui Zhu <[email protected]> > > --- > > mm/bpf_memcontrol.c | 88 ++++++++++++++++++++++++++++++++++++++++++++- > > mm/internal.h | 10 +++--- > > 2 files changed, 93 insertions(+), 5 deletions(-) > > > > diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c > > index 716df49d7647..d8827bc388ef 100644 > > --- a/mm/bpf_memcontrol.c > > +++ b/mm/bpf_memcontrol.c > > @@ -8,6 +8,8 @@ > > #include <linux/memcontrol.h> > > #include <linux/bpf.h> > > > > +#include "internal.h" > > + > > __bpf_kfunc_start_defs(); > > > > /** > > @@ -159,6 +161,74 @@ __bpf_kfunc void bpf_mem_cgroup_flush_stats(struct > > mem_cgroup *memcg) > > mem_cgroup_flush_stats(memcg); > > } > > > > +/** > > + * bpf_proactive_reclaim - proactively reclaim memory from a memory > > + * cgroup > > + * @memcg: the target memory cgroup to reclaim from > > + * @size: the amount of memory to reclaim, in bytes, clamped to > > + * MEMCG_CHARGE_BATCH (64 pages) > > + * @swappiness: the reclaim swappiness, in the range > > + * [MIN_SWAPPINESS, SWAPPINESS_ANON_ONLY], where > > + * SWAPPINESS_ANON_ONLY means anon-only reclaim, or -1 to use > > + * the memcg's own swappiness > > + * > > + * Trigger one proactive reclaim pass on @memcg, similar to a write to > > + * memory.reclaim, but without retrying until @size is reached. > > + * > > + * Only the reclaim target is capped: @size is clamped to > > + * MEMCG_CHARGE_BATCH, following the precedent of high_work_func(), > > + * the memory.high workqueue fallback, which bounds each reclaim > > + * request the same way. The actual scanning work and its duration > > + * are not bounded. To reclaim more, call this kfunc repeatedly > > + * instead of passing a larger @size. > > + * > > + * The kfunc can be called directly from a BPF_PROG_TYPE_SYSCALL > > + * program, synchronously in the context of the thread running the > > + * program, or from the bpf_wq and task_work callbacks of a SYSCALL > > + * program, which run in process context and keep the SYSCALL program > > + * type. It is registered for BPF_PROG_TYPE_SYSCALL only, because > > + * generic sleepable programs may run with filesystem locks held or > > + * in NOFS/NOIO contexts, where the reclaim path could deadlock on > > + * those locks via filesystem shrinkers. > > + * > > + * For asynchronous reclaim of more than one batch, driving the > > + * reclaim from a bpf_wq callback is recommended: call this kfunc > > + * once per callback and requeue the same work item for the next > > + * batch, instead of looping inside the callback and monopolizing a > > + * workqueue worker, and give each target memcg its own work item, > > + * as high_work_func() does with one work item per memcg. Whether > > + * to submit the next batch is up to the BPF program, which can stop > > + * at any point, e.g. once the target cgroup is dying. > > + * > > + * Return: The amount of memory reclaimed, in bytes, or 0 if @size is > > + * smaller than a page, or (unsigned long)-1 if @swappiness is out of > > + * range. > > + */ > > +__bpf_kfunc unsigned long bpf_proactive_reclaim(struct mem_cgroup *memcg, > > + unsigned long size, > > + int swappiness) > > I'm going to have to request one final change, sorry. > > I think long is more meaningful as return value than unsigned long. It's > already > restricted to MEMCG_CHARGE_BATCH. We return error for various arguments, we > should probably change to -EINVAL. > > Apart from that it looks ok to me, but please, also wait for Shakeel to review > the set before you respin v11.
My only feedback is: please don't write essays in the comments. Just couple of sentences should be sufficient. (Please ask your AI to be very very concise for the comments and commit messages.) Other than that please follow Kumar's suggestion and send the next version.

