On Sat, Sep 19, 2026 at 4:15 AM Shakeel Butt <[email protected]> wrote: > > On Fri, Sep 18, 2026 at 12:45:01PM -0700, JP Kobryn wrote: > > On 9/17/26 11:58 PM, Hui Zhu wrote: > > > From: Hui Zhu <[email protected]> > > > > > > BPF programs can observe memory pressure on a cgroup, e.g. refault > > > stats via bpf_mem_cgroup_page_state(), but cannot act on it: > > > triggering reclaim requires writing to memory.reclaim, which BPF > > > cannot do. > > > > > > Add bpf_proactive_reclaim(), a sleepable kfunc performing one > > > proactive reclaim pass on a memcg, like a write to memory.reclaim > > > but without retrying until the target is reached, so that when and > > > how hard to reclaim is BPF policy rather than hard-coded thresholds. > > > The reclaim target of a single call is capped at MEMCG_CHARGE_BATCH, > > > as high_work_func() does for memory.high; reclaiming more is left to > > > the program, which can call the kfunc once per bpf_wq callback and > > > stop at any point. It is limited to BPF_PROG_TYPE_SYSCALL, because > > > other sleepable programs may run with filesystem locks held, on > > > which the reclaim path could deadlock via filesystem shrinkers. > > > > > > Convert MIN_SWAPPINESS, MAX_SWAPPINESS and SWAPPINESS_ANON_ONLY from > > > macros to an enum so that they are emitted into BTF and usable from > > > BPF programs via vmlinux.h. > > > > > > Signed-off-by: Hui Zhu <[email protected]> > > > Acked-by: Shakeel Butt <[email protected]> > > > --- > > > mm/bpf_memcontrol.c | 62 ++++++++++++++++++++++++++++++++++++++++++++- > > > mm/internal.h | 10 +++++--- > > > 2 files changed, 67 insertions(+), 5 deletions(-) > > > > > > diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c > > > index 716df49d7647..c5d7f29ade85 100644 > > > --- a/mm/bpf_memcontrol.c > > > +++ b/mm/bpf_memcontrol.c > > > @@ -8,6 +8,8 @@ > > > #include <linux/memcontrol.h> > > > #include <linux/bpf.h> > > > +#include "internal.h" > > > + > > > __bpf_kfunc_start_defs(); > > > /** > > > @@ -159,6 +161,48 @@ __bpf_kfunc void bpf_mem_cgroup_flush_stats(struct > > > mem_cgroup *memcg) > > > mem_cgroup_flush_stats(memcg); > > > } > > > +/** > > > + * bpf_proactive_reclaim - proactively reclaim memory from a memory > > > cgroup > > > + * @memcg: the target memory cgroup to reclaim from. > > > + * @size: the amount of memory to reclaim, in bytes, clamped to > > > + * MEMCG_CHARGE_BATCH. > > > + * @swappiness: the reclaim swappiness, in the range [MIN_SWAPPINESS, > > > + * SWAPPINESS_ANON_ONLY], or -1 to use the memcg's own > > > + * swappiness. > > > > The ANON_ONLY enumerator shouldn't be included as part of the range. I > > would change this to [MIN_SWAPPINESS, MAX_SWAPPINESS] and then specify > > that ANON_ONLY is a special mode like -1 is. > > > > > + * > > > + * Performs one proactive reclaim pass on @memcg, like a write to > > > + * memory.reclaim but without retrying until @size is reached. Call it > > > + * repeatedly to reclaim more than one batch. > > > + * > > > + * Only available to BPF_PROG_TYPE_SYSCALL, because other sleepable > > > programs > > > + * may run with filesystem locks held, which the reclaim path can > > > deadlock > > > + * on via filesystem shrinkers. > > > + * > > > + * Return: The amount of memory reclaimed, in bytes, or a negative error. > > > + */ > > > +__bpf_kfunc long bpf_proactive_reclaim(struct mem_cgroup *memcg, > > > + unsigned long size, > > > + int swappiness) > > > +{ > > > + unsigned long nr_reclaimed; > > > + unsigned long nr_pages; > > > + > > > + if (swappiness < -1 || swappiness > SWAPPINESS_ANON_ONLY) > > > + return -EINVAL; > > > > Related to the previous comment, you treat the special values as part of > > the range. It works currently, but creates a layout dependency on the > > enum. I think it would be more future-proof if you did: > > > > if (swappiness != -1 && swappiness != SWAPPINESS_ANON_ONLY) { > > if (swappiness < MIN_SWAPPINESS || swappiness > MAX_SWAPPINESS) > > return -EINVAL; > > } > > > > The previous comments I brought up are now resolved, so assuming you'll > > make the changes above you can include: > > > > Reviewed-by: JP Kobryn <[email protected]> > > Hui, please make these changes and just send this patch in next version.
Hi JP and Shakeel, This feels a bit weird to me, since 0–201 are all valid swappiness values. 201 (`SWAPPINESS_ANON_ONLY`) is just as valid a swappiness value as 200. In contrast, -1 is a different story, as it indicates completely different semantics. If 201 is treated as a special value, then 0 (`MIN_SWAPPINESS`) should be treated the same way for MGLRU, since 0 means file-only reclaim there. This isn't the case for CLRU, where swappiness 0 can still reclaim anon folios. If possible, I'd vote for keeping Hui's current code :-) Best Regards Barry

