On Fri, Sep 11, 2026 at 11:40:59AM +0800, Matthias Goergens wrote:
> The bcachefs ktest allocation-leak check writes rcutree.do_rcu_barrier
> before reading /proc/allocinfo. While testing bcachefs performance
> changes, small objects released with kfree_rcu() remained visible after
> repeated writes to the hook and 20 seconds of waiting, causing otherwise
> clean tests to fail their leak check.
>
> The test assumes a stronger contract than the hook currently documents:
> rcu_barrier() waits for ordinary callbacks, but does not flush objects
> still held in kfree_rcu() batching or per-CPU SLUB sheaves. The retained
> population eventually fell as a sheaf filled; there is no evidence here
> of unbounded growth or OOM.
>
> Changing the hook to drain kvfree_rcu() work let the same unmodified
> bcachefs workload pass its allocation check. All eight checkpoints in
> one VM, after 50 through 400 option changes, reported zero retained
> reconcile_scan objects. This motivated the separate private-cache test
> used to isolate the incomplete drain from bcachefs.
>
> Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed
> when kvfree_rcu_barrier() was added in 2024, to restore a clean baseline
> between userspace benchmark runs. The discussion concluded that keeping
> the existing hook name, adding the second operation and documenting both
> was the safest compatibility choice, but the follow-up was not added.
>
> Add that drain and document the stronger test interface. Always retain
> the existing start-rate limit and perform the kvfree_rcu() drain: an
> unrelated ordinary barrier does not establish that this work completed.
>
> Retain the entry ordinary-barrier sequence snapshot. After draining,
> skip the final ordinary barrier only if that snapshot is complete,
> preserving the memory barrier on the completion path. Otherwise, invoke
> rcu_barrier() explicitly. This keeps the ordinary-callback guarantee
> independent of whether kvfree_rcu_barrier() embeds an ordinary barrier.
>
> Clarify that the documented completion guarantee covers work queued
> before the request, without preventing new work from being queued.
>
> Earlier validation of the unconditional-drain version used four fresh
> VM pairs with a private-cache fixture: controls retained the queued
> object (60 to 60 active objects), and treatments drained it (60 to 59).
> An ordinary-callback test passed on both kernels. Those runs predated
> the guarded skip and do not validate that change. No elapsed-time
> improvement is claimed.
>
> Link: https://lore.kernel.org/all/[email protected]/
> Signed-off-by: Matthias Goergens <[email protected]>
Queued for further review and testing, thank you!
Thanx, Paul
> ---
> .../admin-guide/kernel-parameters.txt | 9 ++++--
> kernel/rcu/tree.c | 30 ++++++++++++-------
> 2 files changed, 26 insertions(+), 13 deletions(-)
>
> diff --git a/Documentation/admin-guide/kernel-parameters.txt
> b/Documentation/admin-guide/kernel-parameters.txt
> index 68647ff4bdd2..914b65ae9413 100644
> --- a/Documentation/admin-guide/kernel-parameters.txt
> +++ b/Documentation/admin-guide/kernel-parameters.txt
> @@ -5699,9 +5699,12 @@ Kernel parameters
> there is an ongoing too-long CSD-lock wait.
>
> rcutree.do_rcu_barrier= [KNL]
> - Request a call to rcu_barrier(). This is
> - throttled so that userspace tests can safely
> - hammer on the sysfs variable if they so choose.
> + Wait for deferred kfree_rcu() frees and ordinary
> + call_rcu() callbacks queued before this request to
> + complete. This does not prevent new work from being
> + queued concurrently. Requests are throttled so that
> + userspace tests can safely hammer on the sysfs
> + variable if they so choose.
> If triggered before the RCU grace-period machinery
> is fully active, this will error out with EAGAIN.
>
> diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> index 96848fc1f02b..93b71682306c 100644
> --- a/kernel/rcu/tree.c
> +++ b/kernel/rcu/tree.c
> @@ -3989,12 +3989,12 @@ EXPORT_SYMBOL_GPL(rcu_barrier);
> static unsigned long rcu_barrier_last_throttle;
>
> /**
> - * rcu_barrier_throttled - Do rcu_barrier(), but limit to one per second
> + * rcu_barrier_throttled - Drain deferred RCU frees, but rate-limit starts
> *
> - * This can be thought of as guard rails around rcu_barrier() that
> - * permits unrestricted userspace use, at least assuming the hardware's
> - * try_cmpxchg() is robust. There will be at most one call per second to
> - * rcu_barrier() system-wide from use of this function, which means that
> + * This can be thought of as guard rails around the deferred-free barriers
> + * that permit unrestricted userspace use, at least assuming the hardware's
> + * try_cmpxchg() is robust. There will be at most one drain operation
> started
> + * per sixteenth of a second from use of this function, which means that
> * callers might needlessly wait a second or three.
> *
> * This is intended for use by test suites to avoid OOM by flushing RCU
> @@ -4016,14 +4016,24 @@ static void rcu_barrier_throttled(void)
> while (time_in_range(j, old, old + HZ / 16) ||
> !try_cmpxchg(&rcu_barrier_last_throttle, &old, j)) {
> schedule_timeout_idle(HZ / 16);
> - if (rcu_seq_done(&rcu_state.barrier_sequence, s)) {
> - smp_mb(); /* caller's subsequent code after above
> check. */
> - return;
> - }
> j = jiffies;
> old = READ_ONCE(rcu_barrier_last_throttle);
> }
> - rcu_barrier();
> + /*
> + * kfree_rcu() can retain objects outside the ordinary callback lists in
> + * per-CPU SLUB sheaves and kvfree_rcu batches. Always drain those
> queues:
> + * an ordinary barrier does not establish that this work was drained.
> + */
> + kvfree_rcu_barrier();
> + /*
> + * A completed barrier can still cover ordinary callbacks queued before
> + * our entry snapshot. Otherwise, retain an explicit ordinary barrier
> + * without depending on the implementation of kvfree_rcu_barrier().
> + */
> + if (rcu_seq_done(&rcu_state.barrier_sequence, s))
> + smp_mb(); /* caller's subsequent code after above check. */
> + else
> + rcu_barrier();
> }
>
> /*
> --
> 2.55.0
>