* [PATCH v2] tracing: fgraph: fix the quadratic shadow stack walk @ 2026-09-29 0:54 Vineet Gupta 2026-09-29 0:54 ` [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT Vineet Gupta 0 siblings, 1 reply; 5+ messages in thread From: Vineet Gupta @ 2026-09-29 0:54 UTC (permalink / raw) To: rostedt, mhiramat Cc: mark.rutland, mathieu.desnoyers, andrii, peterz, linux-trace-kernel, linux-kernel, bpf, kernel-team, Vineet Gupta v1 tried to make the sweep loop cheaper: raise the batch from 32 to 1024, and add a cond_resched() so the walk could not hold a CPU. Peter pointed out that the loop does not need to be batched at all. Allocating inline with GFP_NOWAIT is legal under rcu_read_lock() because it cannot sleep, so a single sweep can serve every task and the pre-allocated array becomes an out-of-memory fallback. That removes the quadratic behaviour rather than dividing it by a constant, so both v1 patches are dropped in favour of this one. Measured at 400000 threads on a 60-core Sapphire Rapids machine, PREEMPT_LAZY: v1 base (batch 32) 227.8 s 12503 sweeps v1 patch (batch 1024) 7.2 s 391 sweeps v2 (this patch) 0.092 s 1 sweep With one sweep there is no retry loop left, so the v1 cond_resched() patch has nothing to attach to and is dropped too. v1: https://lore.kernel.org/all/20260922225526.1554758-1-vineet.gupta@linux.dev/ Changes since v1: - replace both patches with Peter's inline GFP_NOWAIT approach - comment why returning from inside scoped_guard() skips the free loop safely Vineet Gupta (1): tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT kernel/trace/fgraph.c | 38 ++++++++++++++++++++++++-------------- 1 file changed, 24 insertions(+), 14 deletions(-) -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT 2026-09-29 0:54 [PATCH v2] tracing: fgraph: fix the quadratic shadow stack walk Vineet Gupta @ 2026-09-29 0:54 ` Vineet Gupta [not found] ` <20260929010327.024281F000FF@smtp.kernel.org> 0 siblings, 1 reply; 5+ messages in thread From: Vineet Gupta @ 2026-09-29 0:54 UTC (permalink / raw) To: rostedt, mhiramat Cc: mark.rutland, mathieu.desnoyers, andrii, peterz, linux-trace-kernel, linux-kernel, bpf, kernel-team, Vineet Gupta, stable When ftrace graphing is turned on, all tasks in the system missing return stack page are assigned one. This is done in a simplistic multi-sweep loop of FTRACE_RETSTACK_ALLOC_SIZE (currently 32) tasks at a time as follows: start_graph_tracing() do { alloc_retstack_tasklist } while (-EAGAIN); alloc_retstack_tasklist() alloc x32 # GFP_KERNEL, may sleep rcu_read_lock() # preempt off for_each_process_thread walk N_total, no cond_resched t->ret_stack = new_page rcu_read_unlock() # preempt enable but no explicit yield Each successive iteration of loop invokes for_each_process_thread() which doesn't support cursor based resume and always restarts from the init_task. Thus each successive loop needs to skip the tasks assigned ret_stack in prior sweeps and thus take longer and longer to find the candidate 32 tasks. And while the loop end calls rcu unlock, and re-enables preemption briefly, there is no explicit yield. The total number of iterations turn out to be N_total * N_null / (2 * FTRACE_RETSTACK_ALLOC_SIZE) where N_total is the number of threads on the system and N_null the number of those whose ret_stack is still NULL. After boot with no fgraph user that is every thread. This is not a tracing-only path. Since commit 4346ba160409 ("fprobe: Rewrite fprobe on function-graph tracer") fprobe is built on fgraph, so an ordinary kprobe_multi BPF attach that flips ftrace_graph_active from 0 to 1 pays the whole cost inside one bpf() syscall. Commit 2c67dc457bc6 ("tracing: fprobe: optimization for entry only case") narrows that to return and session probes but does not remove it. A host running hundreds of thousands of threads, with the current batch size of 32, turns the ftrace_graph_active 0 -> 1 transition into a multi-second, and eventually multi-minute, operation which showed up as RCU stall and softlockup_panic on heavily loaded Meta Fleet machines with preemption disabled. rcu: INFO: rcu_sched self-detected stall on CPU rcu: 0-....: (20718 ticks this GP) (t=21000 jiffies g=913 q=8 ncpus=8) CPU: 0 UID: 0 PID: 160141 Comm: bpftrace ___slab_alloc+0x549/0xa20 kmem_cache_alloc_noprof+0x16c/0x340 register_ftrace_graph+0x3d7/0x6c0 register_fprobe_ips+0x2d5/0x310 bpf_kprobe_multi_link_attach+0x218/0x7e0 __sys_bpf+0x267a/0x27a0 Fix this by allocating inline with GFP_NOWAIT instead of pre-allocating a fixed batch. GFP_NOWAIT does not sleep, so it is legal inside the RCU read-side critical section. A single sweep now serves every task, so the walk becomes O(N) rather than O(N_total * N_null), and the retry loop is no longer the normal path. scoped_guard(rcu) replaces the open-coded rcu_read_lock() and the goto unlock, which is what makes returning directly from inside the critical section legal. The pre-allocated array remains as a fallback, although the testing below never needed it. smp_store_release() is a drop-in replacement for the smp_wmb() and plain store it replaces. Measured on a 60-core Sapphire Rapids machine running a PREEMPT_LAZY kernel (CONFIG_PREEMPT_DYNAMIC=y, mode lazy), timing the write that drives the ftrace_graph_active 0 -> 1 transition. Threads are spawned while fgraph is inactive so every one of them has ret_stack == NULL, and the loop was instrumented to count sweeps: threads before after sweeps time sweeps time 100000 3126 15.3 s 1 0.019 s 200000 6253 59.1 s 1 0.047 s 400000 12503 227.8 s 1 0.092 s The sweep count is 1 throughout. Time now scales linearly with the thread count at roughly 230 ns per task, which is the allocation and initialisation of one shadow stack. Fixes: f201ae2356c7 ("tracing/function-return-tracer: store return stack into task_struct and allocate it dynamically") Suggested-by: Peter Zijlstra <peterz@infradead.org> Signed-off-by: Vineet Gupta <vineet.gupta@linux.dev> Tested-by: Vineet Gupta <vineet.gupta@linux.dev> Cc: stable@vger.kernel.org --- kernel/trace/fgraph.c | 38 ++++++++++++++++++++++++-------------- 1 file changed, 24 insertions(+), 14 deletions(-) diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c index ed455b53513b..336f40dc6618 100644 --- a/kernel/trace/fgraph.c +++ b/kernel/trace/fgraph.c @@ -1036,10 +1036,9 @@ trace_func_graph_ent_t ftrace_graph_entry = ftrace_graph_entry_stub; /* Try to assign a return stack array on FTRACE_RETSTACK_ALLOC_SIZE tasks. */ static int alloc_retstack_tasklist(unsigned long **ret_stack_list) { - int i; - int ret = 0; int start = 0, end = FTRACE_RETSTACK_ALLOC_SIZE; struct task_struct *g, *t; + int i, ret = 0; if (WARN_ON_ONCE(!fgraph_stack_cachep)) return -ENOMEM; @@ -1054,26 +1053,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list) } } - rcu_read_lock(); - for_each_process_thread(g, t) { - if (start == end) { - ret = -EAGAIN; - goto unlock; - } + scoped_guard (rcu) { + for_each_process_thread(g, t) { + unsigned long *rs; + + if (t->ret_stack) + continue; + + rs = kmem_cache_alloc(fgraph_stack_cachep, GFP_NOWAIT); + if (!rs) { + /* + * Returning from inside scoped_guard() drops + * the RCU read lock, but skips the free loop + * below. That is only safe because start == + * end here, which leaves that loop nothing to + * free. Keep the two in step if this + * exhaustion check ever changes. + */ + if (start == end) + return -EAGAIN; + rs = ret_stack_list[start++]; + } - if (t->ret_stack == NULL) { atomic_set(&t->trace_overrun, 0); - ret_stack_init_task_vars(ret_stack_list[start]); + ret_stack_init_task_vars(rs); t->curr_ret_stack = 0; t->curr_ret_depth = -1; /* Make sure the tasks see the 0 first: */ - smp_wmb(); - t->ret_stack = ret_stack_list[start++]; + smp_store_release(&t->ret_stack, rs); } } -unlock: - rcu_read_unlock(); free: for (i = start; i < end; i++) kmem_cache_free(fgraph_stack_cachep, ret_stack_list[i]); -- 2.53.0-Meta ^ permalink raw reply [flat|nested] 5+ messages in thread
[parent not found: <20260929010327.024281F000FF@smtp.kernel.org>]
* Re: [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT [not found] ` <20260929010327.024281F000FF@smtp.kernel.org> @ 2026-09-29 1:22 ` Vineet Gupta 2026-09-29 2:25 ` Andrii Nakryiko 0 siblings, 1 reply; 5+ messages in thread From: Vineet Gupta @ 2026-09-29 1:22 UTC (permalink / raw) To: sashiko-reviews, Peter Zijlstra, rostedt, Andrii Nakryiko Cc: bpf, linux-trace-kernel, LKML, kernel-team On 9/28/26 6:03 PM, sashiko-bot@kernel.org wrote: > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: > - [Low] cleanup: mixing goto-based error handling and scope-based cleanup > -- > > commit a6de528fef00ef67da15965a7a8b3944fa92b053 > Author: Vineet Gupta <vineet.gupta@linux.dev> > Subject: tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT > > This commit aims to resolve performance regressions with ftrace graph tracing > when there are a large number of tasks. It replaces the pre-allocated batch > approach with inline allocation using GFP_NOWAIT, avoiding repeated sweeps of > the task list. Additionally, it transitions the RCU read-side section to use > scoped_guard(rcu) instead of manual rcu_read_lock and rcu_read_unlock calls. > >> diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c >> index 40d373d65f9b9..1c5ccc3f80d53 100644 >> --- a/kernel/trace/fgraph.c >> +++ b/kernel/trace/fgraph.c > [ ... ] >> @@ -1057,26 +1056,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list) >> } >> } >> >> - rcu_read_lock(); >> - for_each_process_thread(g, t) { >> - if (start == end) { >> - ret = -EAGAIN; >> - goto unlock; >> - } >> + scoped_guard (rcu) { > [Severity: Low] > Does this mix goto-based error handling and scope-based cleanup within the same > function? > > Looking at alloc_retstack_tasklist(), the function retains the legacy goto free > path in the initial pre-allocation loop (just before this block), while > introducing scoped_guard(rcu) here. > > The cleanup subsystem guidelines recommend that for a given routine, we should > convert all resources that need a goto cleanup to scope-based cleanup, or > convert none of them. Mixing these paradigms creates confusing ownership > semantics and is highly prone to double-free or resource leak regressions upon > future modifications. > > Could we either convert the ret_stack_list array cleanup to use scope-based > helpers, or stick to the manual rcu_read_lock() and rcu_read_unlock() here to > keep the error handling style consistent? And a third option is to drop the fallback allocation completely and return -EAGAIN and come back. Granted I don't have much experience of typical handling of GFP_NOWAIT fails and retrying immediately: would that recover at all or does that take us back to where we started ? Peter Thx, -Vineet ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT 2026-09-29 1:22 ` Vineet Gupta @ 2026-09-29 2:25 ` Andrii Nakryiko 2026-09-29 8:03 ` Peter Zijlstra 0 siblings, 1 reply; 5+ messages in thread From: Andrii Nakryiko @ 2026-09-29 2:25 UTC (permalink / raw) To: Vineet Gupta Cc: sashiko-reviews, Peter Zijlstra, rostedt, Andrii Nakryiko, bpf, linux-trace-kernel, LKML, kernel-team On Mon, Sep 28, 2026 at 6:22 PM Vineet Gupta <vineet.gupta@linux.dev> wrote: > > On 9/28/26 6:03 PM, sashiko-bot@kernel.org wrote: > > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: > > - [Low] cleanup: mixing goto-based error handling and scope-based cleanup > > -- > > > > commit a6de528fef00ef67da15965a7a8b3944fa92b053 > > Author: Vineet Gupta <vineet.gupta@linux.dev> > > Subject: tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT > > > > This commit aims to resolve performance regressions with ftrace graph tracing > > when there are a large number of tasks. It replaces the pre-allocated batch > > approach with inline allocation using GFP_NOWAIT, avoiding repeated sweeps of > > the task list. Additionally, it transitions the RCU read-side section to use > > scoped_guard(rcu) instead of manual rcu_read_lock and rcu_read_unlock calls. > > > >> diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c > >> index 40d373d65f9b9..1c5ccc3f80d53 100644 > >> --- a/kernel/trace/fgraph.c > >> +++ b/kernel/trace/fgraph.c > > [ ... ] > >> @@ -1057,26 +1056,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list) > >> } > >> } > >> > >> - rcu_read_lock(); > >> - for_each_process_thread(g, t) { > >> - if (start == end) { > >> - ret = -EAGAIN; > >> - goto unlock; > >> - } > >> + scoped_guard (rcu) { > > [Severity: Low] > > Does this mix goto-based error handling and scope-based cleanup within the same > > function? > > > > Looking at alloc_retstack_tasklist(), the function retains the legacy goto free > > path in the initial pre-allocation loop (just before this block), while > > introducing scoped_guard(rcu) here. > > > > The cleanup subsystem guidelines recommend that for a given routine, we should > > convert all resources that need a goto cleanup to scope-based cleanup, or > > convert none of them. Mixing these paradigms creates confusing ownership > > semantics and is highly prone to double-free or resource leak regressions upon > > future modifications. > > > > Could we either convert the ret_stack_list array cleanup to use scope-based > > helpers, or stick to the manual rcu_read_lock() and rcu_read_unlock() here to > > keep the error handling style consistent? > > And a third option is to drop the fallback allocation completely and > return -EAGAIN and come back. Granted I don't have much experience of > typical handling of GFP_NOWAIT fails and retrying immediately: would > that recover at all or does that take us back to where we started ? Peter > Not Peter, but I don't see a problem with scoped rcu and goto-based free/cleanup. Unless Peter objects, let's keep it as is? > Thx, > -Vineet ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT 2026-09-29 2:25 ` Andrii Nakryiko @ 2026-09-29 8:03 ` Peter Zijlstra 0 siblings, 0 replies; 5+ messages in thread From: Peter Zijlstra @ 2026-09-29 8:03 UTC (permalink / raw) To: Andrii Nakryiko Cc: Vineet Gupta, sashiko-reviews, rostedt, Andrii Nakryiko, bpf, linux-trace-kernel, LKML, kernel-team On Mon, Sep 28, 2026 at 07:25:00PM -0700, Andrii Nakryiko wrote: > On Mon, Sep 28, 2026 at 6:22 PM Vineet Gupta <vineet.gupta@linux.dev> wrote: > > > > On 9/28/26 6:03 PM, sashiko-bot@kernel.org wrote: > > > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider: > > > - [Low] cleanup: mixing goto-based error handling and scope-based cleanup > > > -- > > > > > > commit a6de528fef00ef67da15965a7a8b3944fa92b053 > > > Author: Vineet Gupta <vineet.gupta@linux.dev> > > > Subject: tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT > > > > > > This commit aims to resolve performance regressions with ftrace graph tracing > > > when there are a large number of tasks. It replaces the pre-allocated batch > > > approach with inline allocation using GFP_NOWAIT, avoiding repeated sweeps of > > > the task list. Additionally, it transitions the RCU read-side section to use > > > scoped_guard(rcu) instead of manual rcu_read_lock and rcu_read_unlock calls. > > > > > >> diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c > > >> index 40d373d65f9b9..1c5ccc3f80d53 100644 > > >> --- a/kernel/trace/fgraph.c > > >> +++ b/kernel/trace/fgraph.c > > > [ ... ] > > >> @@ -1057,26 +1056,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list) > > >> } > > >> } > > >> > > >> - rcu_read_lock(); > > >> - for_each_process_thread(g, t) { > > >> - if (start == end) { > > >> - ret = -EAGAIN; > > >> - goto unlock; > > >> - } > > >> + scoped_guard (rcu) { > > > [Severity: Low] > > > Does this mix goto-based error handling and scope-based cleanup within the same > > > function? > > > > > > Looking at alloc_retstack_tasklist(), the function retains the legacy goto free > > > path in the initial pre-allocation loop (just before this block), while > > > introducing scoped_guard(rcu) here. > > > > > > The cleanup subsystem guidelines recommend that for a given routine, we should > > > convert all resources that need a goto cleanup to scope-based cleanup, or > > > convert none of them. Mixing these paradigms creates confusing ownership > > > semantics and is highly prone to double-free or resource leak regressions upon > > > future modifications. > > > > > > Could we either convert the ret_stack_list array cleanup to use scope-based > > > helpers, or stick to the manual rcu_read_lock() and rcu_read_unlock() here to > > > keep the error handling style consistent? > > > > And a third option is to drop the fallback allocation completely and > > return -EAGAIN and come back. Granted I don't have much experience of > > typical handling of GFP_NOWAIT fails and retrying immediately: would > > that recover at all or does that take us back to where we started ? Peter > > > > Not Peter, but I don't see a problem with scoped rcu and goto-based > free/cleanup. Unless Peter objects, let's keep it as is? Yeah, the silly robot is being silly. Code is fine as is. The guideline is just that, a guide. Its not a hard requirement. It is possible to create a terrible mess of things when doing a partial conversion of a large multi-stage goto unwind fest. But in general, goto over the scope (as here) is fine. Also goto out of a scope is also fine. ^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-29 8:03 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 0:54 [PATCH v2] tracing: fgraph: fix the quadratic shadow stack walk Vineet Gupta
2026-09-29 0:54 ` [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT Vineet Gupta
[not found] ` <20260929010327.024281F000FF@smtp.kernel.org>
2026-09-29 1:22 ` Vineet Gupta
2026-09-29 2:25 ` Andrii Nakryiko
2026-09-29 8:03 ` Peter Zijlstra
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®