v1 tried to make the sweep loop cheaper: raise the batch from 32 to
1024, and add a cond_resched() so the walk could not hold a CPU. Peter
pointed out that the loop does not need to be batched at all. Allocating
inline with GFP_NOWAIT is legal under rcu_read_lock() because it cannot
sleep, so a single sweep can serve every task and the pre-allocated
array becomes an out-of-memory fallback.
That removes the quadratic behaviour rather than dividing it by a
constant, so both v1 patches are dropped in favour of this one.
Measured at 400000 threads on a 60-core Sapphire Rapids machine,
PREEMPT_LAZY:
v1 base (batch 32) 227.8 s 12503 sweeps
v1 patch (batch 1024) 7.2 s 391 sweeps
v2 (this patch) 0.092 s 1 sweep
With one sweep there is no retry loop left, so the v1 cond_resched()
patch has nothing to attach to and is dropped too.
v1: https://lore.kernel.org/all/[email protected]/
Changes since v1:
- replace both patches with Peter's inline GFP_NOWAIT approach
- comment why returning from inside scoped_guard() skips the free
loop safely
Vineet Gupta (1):
tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
kernel/trace/fgraph.c | 38 ++++++++++++++++++++++++--------------
1 file changed, 24 insertions(+), 14 deletions(-)
--
2.53.0-Meta