Re: [PATCH v3 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
From: Andrii Nakryiko
Date: Thu Oct 01 2026 - 09:38:05 EST
On Thu, Oct 1, 2026 at 11:29 AM Vineet Gupta <vineet.gupta@xxxxxxxxx> wrote:
>
> When ftrace graphing is turned on, all tasks in the system missing
> return stack page are assigned one. This is done in a simplistic
> multi-sweep loop of FTRACE_RETSTACK_ALLOC_SIZE (currently 32) tasks
> at a time as follows:
>
> start_graph_tracing()
> do {
> alloc_retstack_tasklist
> } while (-EAGAIN);
>
> alloc_retstack_tasklist()
> alloc x32 # GFP_KERNEL, may sleep
> rcu_read_lock() # preempt off
> for_each_process_thread walk N_total, no cond_resched
> t->ret_stack = new_page
> rcu_read_unlock() # preempt enable but no explicit yield
>
> Each successive iteration of loop invokes for_each_process_thread()
> which doesn't support cursor based resume and always restarts from the
> init_task. Thus each successive loop needs to skip the tasks assigned
> ret_stack in prior sweeps and thus take longer and longer to find the
> candidate 32 tasks. And while the loop end calls rcu unlock,
> and re-enables preemption briefly, there is no explicit yield.
>
> The total number of iterations turn out to be
>
> N_total * N_null / (2 * FTRACE_RETSTACK_ALLOC_SIZE)
>
> where N_total is the number of threads on the system and N_null the number
> of those whose ret_stack is still NULL. After boot with no fgraph user that
> is every thread.
>
> This is not a tracing-only path. Since commit 4346ba160409 ("fprobe:
> Rewrite fprobe on function-graph tracer") fprobe is built on fgraph,
> so an ordinary kprobe_multi BPF attach that flips ftrace_graph_active
> from 0 to 1 pays the whole cost inside one bpf() syscall. Commit
> 2c67dc457bc6 ("tracing: fprobe: optimization for entry only case")
> narrows that to return and session probes but does not remove it.
>
> A host running hundreds of thousands of threads, with the current batch
> size of 32, turns the ftrace_graph_active 0 -> 1 transition into a
> multi-second, and eventually multi-minute, operation which showed up as
> RCU stall and softlockup_panic on heavily loaded Meta Fleet machines
> with preemption disabled.
>
> rcu: INFO: rcu_sched self-detected stall on CPU
> rcu: 0-....: (20718 ticks this GP) (t=21000 jiffies g=913 q=8 ncpus=8)
> CPU: 0 UID: 0 PID: 160141 Comm: bpftrace
> ___slab_alloc+0x549/0xa20
> kmem_cache_alloc_noprof+0x16c/0x340
> register_ftrace_graph+0x3d7/0x6c0
> register_fprobe_ips+0x2d5/0x310
> bpf_kprobe_multi_link_attach+0x218/0x7e0
> __sys_bpf+0x267a/0x27a0
>
> Fix this by allocating inline with GFP_NOWAIT instead of pre-allocating
> a fixed batch. GFP_NOWAIT does not sleep, so it is legal inside the RCU
> read-side critical section.
>
> A single sweep now serves every task, so the walk becomes O(N) rather
> than O(N_total * N_null), and the retry loop is no longer the normal
> path.
>
> scoped_guard(rcu) replaces the open-coded rcu_read_lock() and the
> goto unlock, which is what makes returning directly from inside the
> critical section legal.
>
> The pre-allocated array remains as a fallback, although the testing
> below never needed it.
>
> smp_store_release() is a drop-in replacement for the smp_wmb() and
> plain store it replaces.
>
> Measured on a 60-core Sapphire Rapids machine running a PREEMPT_LAZY
> kernel (CONFIG_PREEMPT_DYNAMIC=y, mode lazy), timing the write that
> drives the ftrace_graph_active 0 -> 1 transition. Threads are spawned
> while fgraph is inactive so every one of them has ret_stack == NULL,
> and the loop was instrumented to count sweeps:
>
> threads before after
> sweeps time sweeps time
> 100000 3126 15.3 s 1 0.019 s
> 200000 6253 59.1 s 1 0.047 s
> 400000 12503 227.8 s 1 0.092 s
>
> The sweep count is 1 throughout. Time now scales linearly with the
> thread count at roughly 230 ns per task, which is the allocation and
> initialisation of one shadow stack.
>
> Fixes: f201ae2356c7 ("tracing/function-return-tracer: store return stack into task_struct and allocate it dynamically")
> Suggested-by: Peter Zijlstra <peterz@xxxxxxxxxxxxx>
> Signed-off-by: Vineet Gupta <vineet.gupta@xxxxxxxxx>
> Tested-by: Vineet Gupta <vineet.gupta@xxxxxxxxx>
> Cc: stable@xxxxxxxxxxxxxxx
> ---
> kernel/trace/fgraph.c | 44 ++++++++++++++++++++++++++++---------------
> 1 file changed, 29 insertions(+), 15 deletions(-)
>
Acked-by: Andrii Nakryiko <andrii@xxxxxxxxxx>
[...]