v1 tried to make the sweep loop cheaper: raise the batch from 32 to
1024, and add a cond_resched() so the walk could not hold a CPU. Peter
pointed out that the loop does not need to be batched at all. Allocating
inline with GFP_NOWAIT is legal under rcu_read_lock() because it cannot
sleep, so a single sweep can serve every task and the pre-allocated
array becomes an out-of-memory fallback.

That removes the quadratic behaviour rather than dividing it by a
constant, so both v1 patches are dropped in favour of this one.

Measured at 400000 threads on a 60-core Sapphire Rapids machine,
PREEMPT_LAZY:

        v1 base (batch 32)      227.8 s     12503 sweeps
        v1 patch (batch 1024)     7.2 s       391 sweeps
        v2 (this patch)           0.092 s       1 sweep

With one sweep there is no retry loop left, so the v1 cond_resched()
patch has nothing to attach to and is dropped too.

v1: https://lore.kernel.org/all/[email protected]/

Changes since v1:
 - replace both patches with Peter's inline GFP_NOWAIT approach
 - comment why returning from inside scoped_guard() skips the free
   loop safely

Vineet Gupta (1):
  tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT

 kernel/trace/fgraph.c | 38 ++++++++++++++++++++++++--------------
 1 file changed, 24 insertions(+), 14 deletions(-)

-- 
2.53.0-Meta


Reply via email to