When ftrace graphing is turned on, all tasks in the system missing
return stack page are assigned one. This is done in a simplistic
multi-sweep loop of FTRACE_RETSTACK_ALLOC_SIZE (currently 32) tasks
at a time as follows:

   start_graph_tracing()
        do {
                alloc_retstack_tasklist
        } while (-EAGAIN);

   alloc_retstack_tasklist()
        alloc x32           # GFP_KERNEL, may sleep
        rcu_read_lock()     # preempt off
                for_each_process_thread    walk N_total, no cond_resched
                     t->ret_stack = new_page
        rcu_read_unlock()   # preempt enable but no explicit yield

Each successive iteration of loop invokes for_each_process_thread()
which doesn't support cursor based resume and always restarts from the
init_task. Thus each successive loop needs to skip the tasks assigned
ret_stack in prior sweeps and thus take longer and longer to find the
candidate 32 tasks. And while the loop end calls rcu unlock,
and re-enables preemption briefly, there is no explicit yield.

The total number of iterations turn out to be

        N_total * N_null / (2 * FTRACE_RETSTACK_ALLOC_SIZE)

where N_total is the number of threads on the system and N_null the number
of those whose ret_stack is still NULL. After boot with no fgraph user that
is every thread.

This is not a tracing-only path. Since commit 4346ba160409 ("fprobe:
Rewrite fprobe on function-graph tracer") fprobe is built on fgraph,
so an ordinary kprobe_multi BPF attach that flips ftrace_graph_active
from 0 to 1 pays the whole cost inside one bpf() syscall. Commit
2c67dc457bc6 ("tracing: fprobe: optimization for entry only case")
narrows that to return and session probes but does not remove it.

A host running hundreds of thousands of threads, with the current batch
size of 32, turns the ftrace_graph_active 0 -> 1 transition into a
multi-second, and eventually multi-minute, operation which showed up as
RCU stall and softlockup_panic on heavily loaded Meta Fleet machines
with preemption disabled.

    rcu: INFO: rcu_sched self-detected stall on CPU
    rcu: 0-....: (20718 ticks this GP) (t=21000 jiffies g=913 q=8 ncpus=8)
    CPU: 0 UID: 0 PID: 160141 Comm: bpftrace
     ___slab_alloc+0x549/0xa20
     kmem_cache_alloc_noprof+0x16c/0x340
     register_ftrace_graph+0x3d7/0x6c0
     register_fprobe_ips+0x2d5/0x310
     bpf_kprobe_multi_link_attach+0x218/0x7e0
     __sys_bpf+0x267a/0x27a0

Fix this by allocating inline with GFP_NOWAIT instead of pre-allocating
a fixed batch. GFP_NOWAIT does not sleep, so it is legal inside the RCU
read-side critical section.

A single sweep now serves every task, so the walk becomes O(N) rather
than O(N_total * N_null), and the retry loop is no longer the normal
path.

scoped_guard(rcu) replaces the open-coded rcu_read_lock() and the
goto unlock, which is what makes returning directly from inside the
critical section legal.

The pre-allocated array remains as a fallback, although the testing
below never needed it.

smp_store_release() is a drop-in replacement for the smp_wmb() and
plain store it replaces.

Measured on a 60-core Sapphire Rapids machine running a PREEMPT_LAZY
kernel (CONFIG_PREEMPT_DYNAMIC=y, mode lazy), timing the write that
drives the ftrace_graph_active 0 -> 1 transition. Threads are spawned
while fgraph is inactive so every one of them has ret_stack == NULL,
and the loop was instrumented to count sweeps:

        threads   before              after
                  sweeps   time       sweeps   time
        100000     3126     15.3 s      1       0.019 s
        200000     6253     59.1 s      1       0.047 s
        400000    12503    227.8 s      1       0.092 s

The sweep count is 1 throughout. Time now scales linearly with the
thread count at roughly 230 ns per task, which is the allocation and
initialisation of one shadow stack.

Fixes: f201ae2356c7 ("tracing/function-return-tracer: store return stack into 
task_struct and allocate it dynamically")
Suggested-by: Peter Zijlstra <[email protected]>
Signed-off-by: Vineet Gupta <[email protected]>
Tested-by: Vineet Gupta <[email protected]>
Cc: [email protected]
---
 kernel/trace/fgraph.c | 38 ++++++++++++++++++++++++--------------
 1 file changed, 24 insertions(+), 14 deletions(-)

diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c
index ed455b53513b..336f40dc6618 100644
--- a/kernel/trace/fgraph.c
+++ b/kernel/trace/fgraph.c
@@ -1036,10 +1036,9 @@ trace_func_graph_ent_t ftrace_graph_entry = 
ftrace_graph_entry_stub;
 /* Try to assign a return stack array on FTRACE_RETSTACK_ALLOC_SIZE tasks. */
 static int alloc_retstack_tasklist(unsigned long **ret_stack_list)
 {
-       int i;
-       int ret = 0;
        int start = 0, end = FTRACE_RETSTACK_ALLOC_SIZE;
        struct task_struct *g, *t;
+       int i, ret = 0;
 
        if (WARN_ON_ONCE(!fgraph_stack_cachep))
                return -ENOMEM;
@@ -1054,26 +1053,37 @@ static int alloc_retstack_tasklist(unsigned long 
**ret_stack_list)
                }
        }
 
-       rcu_read_lock();
-       for_each_process_thread(g, t) {
-               if (start == end) {
-                       ret = -EAGAIN;
-                       goto unlock;
-               }
+       scoped_guard (rcu) {
+               for_each_process_thread(g, t) {
+                       unsigned long *rs;
+
+                       if (t->ret_stack)
+                               continue;
+
+                       rs = kmem_cache_alloc(fgraph_stack_cachep, GFP_NOWAIT);
+                       if (!rs) {
+                               /*
+                                * Returning from inside scoped_guard() drops
+                                * the RCU read lock, but skips the free loop
+                                * below. That is only safe because start ==
+                                * end here, which leaves that loop nothing to
+                                * free. Keep the two in step if this
+                                * exhaustion check ever changes.
+                                */
+                               if (start == end)
+                                       return -EAGAIN;
+                               rs = ret_stack_list[start++];
+                       }
 
-               if (t->ret_stack == NULL) {
                        atomic_set(&t->trace_overrun, 0);
-                       ret_stack_init_task_vars(ret_stack_list[start]);
+                       ret_stack_init_task_vars(rs);
                        t->curr_ret_stack = 0;
                        t->curr_ret_depth = -1;
                        /* Make sure the tasks see the 0 first: */
-                       smp_wmb();
-                       t->ret_stack = ret_stack_list[start++];
+                       smp_store_release(&t->ret_stack, rs);
                }
        }
 
-unlock:
-       rcu_read_unlock();
 free:
        for (i = start; i < end; i++)
                kmem_cache_free(fgraph_stack_cachep, ret_stack_list[i]);
-- 
2.53.0-Meta


Reply via email to