On Fri, Sep 11, 2026 at 01:00:40AM +0800, Matthias Goergens wrote:
> The bcachefs ktest allocation-leak check writes rcutree.do_rcu_barrier
> before reading /proc/allocinfo. While testing bcachefs performance
> changes, small objects released with kfree_rcu() remained visible after
> repeated writes to the hook and 20 seconds of waiting, causing otherwise
> clean tests to fail their leak check.
> 
> The test assumes a stronger contract than the hook currently documents:
> rcu_barrier() waits for ordinary callbacks, but does not flush objects
> still held in kfree_rcu() batching or per-CPU SLUB sheaves. The retained
> population eventually fell as a sheaf filled; there is no evidence here
> of unbounded growth or OOM.
> 
> Changing the hook to drain kvfree_rcu() work let the same unmodified
> bcachefs workload pass its allocation check. All eight checkpoints in
> one VM, after 50 through 400 option changes, reported zero retained
> reconcile_scan objects. This motivated the separate private-cache
> reproducer used to isolate the incomplete drain from bcachefs.
> 
> Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed
> when the former API was added in 2024, to restore a clean baseline
> between userspace benchmark runs. The discussion concluded that keeping
> the existing hook name, adding the second operation and documenting both
> was the safest compatibility choice, but the follow-up was not added.
> 
> Add that drain and document the stronger test interface. Keep the
> explicit rcu_barrier() so the hook's ordinary-callback contract does not
> depend on kvfree_rcu_barrier() reaching an ordinary barrier internally.
> 
> Do not reuse the ordinary rcu_barrier() sequence as an early-completion
> check while throttling: an unrelated ordinary barrier does not establish
> that kvfree_rcu() work was drained. Retain the existing start-rate limit.
> 
> Four fresh VM pairs with the full private-cache fixture retained the
> queued object without the patch (60 to 60 active objects) and drained it
> with the patch (60 to 59). A separate ordinary-callback regression test
> passed on both kernels.
> 
> Link: https://lore.kernel.org/all/[email protected]/
> Signed-off-by: Matthias Goergens <[email protected]>
> ---
>  .../admin-guide/kernel-parameters.txt         |  7 ++---
>  kernel/rcu/tree.c                             | 27 ++++++++++++-------
>  2 files changed, 21 insertions(+), 13 deletions(-)
> 
> diff --git a/Documentation/admin-guide/kernel-parameters.txt 
> b/Documentation/admin-guide/kernel-parameters.txt
> index 68647ff4bdd2..244a53166249 100644
> --- a/Documentation/admin-guide/kernel-parameters.txt
> +++ b/Documentation/admin-guide/kernel-parameters.txt
> @@ -5699,9 +5699,10 @@ Kernel parameters
>                       there is an ongoing too-long CSD-lock wait.
>  
>       rcutree.do_rcu_barrier= [KNL]
> -                     Request a call to rcu_barrier().  This is
> -                     throttled so that userspace tests can safely
> -                     hammer on the sysfs variable if they so choose.
> +                     Request that deferred kfree_rcu() objects and
> +                     ordinary call_rcu() callbacks be drained.  This is
> +                     throttled so that userspace tests can safely hammer
> +                     on the sysfs variable if they so choose.
>                       If triggered before the RCU grace-period machinery
>                       is fully active, this will error out with EAGAIN.
>  
> diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> index 96848fc1f02b..014e28ec3bd3 100644
> --- a/kernel/rcu/tree.c
> +++ b/kernel/rcu/tree.c
> @@ -3989,12 +3989,12 @@ EXPORT_SYMBOL_GPL(rcu_barrier);
>  static unsigned long rcu_barrier_last_throttle;
>  
>  /**
> - * rcu_barrier_throttled - Do rcu_barrier(), but limit to one per second
> + * rcu_barrier_throttled - Drain deferred RCU frees, but rate-limit starts
>   *
> - * This can be thought of as guard rails around rcu_barrier() that
> - * permits unrestricted userspace use, at least assuming the hardware's
> - * try_cmpxchg() is robust.  There will be at most one call per second to
> - * rcu_barrier() system-wide from use of this function, which means that
> + * This can be thought of as guard rails around the deferred-free barriers
> + * that permit unrestricted userspace use, at least assuming the hardware's
> + * try_cmpxchg() is robust.  There will be at most one drain operation 
> started
> + * per sixteenth of a second from use of this function, which means that
>   * callers might needlessly wait a second or three.
>   *
>   * This is intended for use by test suites to avoid OOM by flushing RCU
> @@ -4011,18 +4011,25 @@ static void rcu_barrier_throttled(void)
>  {
>       unsigned long j = jiffies;
>       unsigned long old = READ_ONCE(rcu_barrier_last_throttle);
> -     unsigned long s = rcu_seq_snap(&rcu_state.barrier_sequence);
>  
>       while (time_in_range(j, old, old + HZ / 16) ||
>              !try_cmpxchg(&rcu_barrier_last_throttle, &old, j)) {
>               schedule_timeout_idle(HZ / 16);
> -             if (rcu_seq_done(&rcu_state.barrier_sequence, s)) {
> -                     smp_mb(); /* caller's subsequent code after above 
> check. */
> -                     return;

Don't we still want to skip the rcu_barrier() in this case?  Or am I missing
something subtle here?

                                                                Thanx, Paul

> -             }
>               j = jiffies;
>               old = READ_ONCE(rcu_barrier_last_throttle);
>       }
> +     /*
> +      * kfree_rcu() can retain objects outside the ordinary callback lists in
> +      * per-CPU SLUB sheaves and kvfree_rcu batches.  Test suites use this 
> hook
> +      * to prevent deferred frees from spilling into the following test, so
> +      * drain those queues as well as ordinary call_rcu() callbacks.
> +      *
> +      * kvfree_rcu_barrier() currently includes an ordinary barrier, but that
> +      * is not part of its documented API.  Keep the explicit rcu_barrier() 
> so
> +      * this hook's original contract does not depend on slab implementation
> +      * details.
> +      */
> +     kvfree_rcu_barrier();
>       rcu_barrier();
>  }
>  
> -- 
> 2.55.0
> 

Reply via email to