mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2] tracing: fgraph: fix the quadratic shadow stack walk
@ 2026-09-29  0:54 Vineet Gupta
  2026-09-29  0:54 ` [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT Vineet Gupta
  0 siblings, 1 reply; 5+ messages in thread
From: Vineet Gupta @ 2026-09-29  0:54 UTC (permalink / raw)
  To: rostedt, mhiramat
  Cc: mark.rutland, mathieu.desnoyers, andrii, peterz,
	linux-trace-kernel, linux-kernel, bpf, kernel-team, Vineet Gupta

v1 tried to make the sweep loop cheaper: raise the batch from 32 to
1024, and add a cond_resched() so the walk could not hold a CPU. Peter
pointed out that the loop does not need to be batched at all. Allocating
inline with GFP_NOWAIT is legal under rcu_read_lock() because it cannot
sleep, so a single sweep can serve every task and the pre-allocated
array becomes an out-of-memory fallback.

That removes the quadratic behaviour rather than dividing it by a
constant, so both v1 patches are dropped in favour of this one.

Measured at 400000 threads on a 60-core Sapphire Rapids machine,
PREEMPT_LAZY:

	v1 base (batch 32)      227.8 s     12503 sweeps
	v1 patch (batch 1024)     7.2 s       391 sweeps
	v2 (this patch)           0.092 s       1 sweep

With one sweep there is no retry loop left, so the v1 cond_resched()
patch has nothing to attach to and is dropped too.

v1: https://lore.kernel.org/all/20260922225526.1554758-1-vineet.gupta@linux.dev/

Changes since v1:
 - replace both patches with Peter's inline GFP_NOWAIT approach
 - comment why returning from inside scoped_guard() skips the free
   loop safely

Vineet Gupta (1):
  tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT

 kernel/trace/fgraph.c | 38 ++++++++++++++++++++++++--------------
 1 file changed, 24 insertions(+), 14 deletions(-)

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
  2026-09-29  0:54 [PATCH v2] tracing: fgraph: fix the quadratic shadow stack walk Vineet Gupta
@ 2026-09-29  0:54 ` Vineet Gupta
       [not found]   ` <20260929010327.024281F000FF@smtp.kernel.org>
  0 siblings, 1 reply; 5+ messages in thread
From: Vineet Gupta @ 2026-09-29  0:54 UTC (permalink / raw)
  To: rostedt, mhiramat
  Cc: mark.rutland, mathieu.desnoyers, andrii, peterz,
	linux-trace-kernel, linux-kernel, bpf, kernel-team, Vineet Gupta,
	stable

When ftrace graphing is turned on, all tasks in the system missing
return stack page are assigned one. This is done in a simplistic
multi-sweep loop of FTRACE_RETSTACK_ALLOC_SIZE (currently 32) tasks
at a time as follows:

   start_graph_tracing()
	do {
		alloc_retstack_tasklist
	} while (-EAGAIN);

   alloc_retstack_tasklist()
	alloc x32           # GFP_KERNEL, may sleep
	rcu_read_lock()     # preempt off
		for_each_process_thread    walk N_total, no cond_resched
                     t->ret_stack = new_page
	rcu_read_unlock()   # preempt enable but no explicit yield

Each successive iteration of loop invokes for_each_process_thread()
which doesn't support cursor based resume and always restarts from the
init_task. Thus each successive loop needs to skip the tasks assigned
ret_stack in prior sweeps and thus take longer and longer to find the
candidate 32 tasks. And while the loop end calls rcu unlock,
and re-enables preemption briefly, there is no explicit yield.

The total number of iterations turn out to be

	N_total * N_null / (2 * FTRACE_RETSTACK_ALLOC_SIZE)

where N_total is the number of threads on the system and N_null the number
of those whose ret_stack is still NULL. After boot with no fgraph user that
is every thread.

This is not a tracing-only path. Since commit 4346ba160409 ("fprobe:
Rewrite fprobe on function-graph tracer") fprobe is built on fgraph,
so an ordinary kprobe_multi BPF attach that flips ftrace_graph_active
from 0 to 1 pays the whole cost inside one bpf() syscall. Commit
2c67dc457bc6 ("tracing: fprobe: optimization for entry only case")
narrows that to return and session probes but does not remove it.

A host running hundreds of thousands of threads, with the current batch
size of 32, turns the ftrace_graph_active 0 -> 1 transition into a
multi-second, and eventually multi-minute, operation which showed up as
RCU stall and softlockup_panic on heavily loaded Meta Fleet machines
with preemption disabled.

    rcu: INFO: rcu_sched self-detected stall on CPU
    rcu: 0-....: (20718 ticks this GP) (t=21000 jiffies g=913 q=8 ncpus=8)
    CPU: 0 UID: 0 PID: 160141 Comm: bpftrace
     ___slab_alloc+0x549/0xa20
     kmem_cache_alloc_noprof+0x16c/0x340
     register_ftrace_graph+0x3d7/0x6c0
     register_fprobe_ips+0x2d5/0x310
     bpf_kprobe_multi_link_attach+0x218/0x7e0
     __sys_bpf+0x267a/0x27a0

Fix this by allocating inline with GFP_NOWAIT instead of pre-allocating
a fixed batch. GFP_NOWAIT does not sleep, so it is legal inside the RCU
read-side critical section.

A single sweep now serves every task, so the walk becomes O(N) rather
than O(N_total * N_null), and the retry loop is no longer the normal
path.

scoped_guard(rcu) replaces the open-coded rcu_read_lock() and the
goto unlock, which is what makes returning directly from inside the
critical section legal.

The pre-allocated array remains as a fallback, although the testing
below never needed it.

smp_store_release() is a drop-in replacement for the smp_wmb() and
plain store it replaces.

Measured on a 60-core Sapphire Rapids machine running a PREEMPT_LAZY
kernel (CONFIG_PREEMPT_DYNAMIC=y, mode lazy), timing the write that
drives the ftrace_graph_active 0 -> 1 transition. Threads are spawned
while fgraph is inactive so every one of them has ret_stack == NULL,
and the loop was instrumented to count sweeps:

	threads   before              after
	          sweeps   time       sweeps   time
	100000     3126     15.3 s      1       0.019 s
	200000     6253     59.1 s      1       0.047 s
	400000    12503    227.8 s      1       0.092 s

The sweep count is 1 throughout. Time now scales linearly with the
thread count at roughly 230 ns per task, which is the allocation and
initialisation of one shadow stack.

Fixes: f201ae2356c7 ("tracing/function-return-tracer: store return stack into task_struct and allocate it dynamically")
Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Vineet Gupta <vineet.gupta@linux.dev>
Tested-by: Vineet Gupta <vineet.gupta@linux.dev>
Cc: stable@vger.kernel.org
---
 kernel/trace/fgraph.c | 38 ++++++++++++++++++++++++--------------
 1 file changed, 24 insertions(+), 14 deletions(-)

diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c
index ed455b53513b..336f40dc6618 100644
--- a/kernel/trace/fgraph.c
+++ b/kernel/trace/fgraph.c
@@ -1036,10 +1036,9 @@ trace_func_graph_ent_t ftrace_graph_entry = ftrace_graph_entry_stub;
 /* Try to assign a return stack array on FTRACE_RETSTACK_ALLOC_SIZE tasks. */
 static int alloc_retstack_tasklist(unsigned long **ret_stack_list)
 {
-	int i;
-	int ret = 0;
 	int start = 0, end = FTRACE_RETSTACK_ALLOC_SIZE;
 	struct task_struct *g, *t;
+	int i, ret = 0;
 
 	if (WARN_ON_ONCE(!fgraph_stack_cachep))
 		return -ENOMEM;
@@ -1054,26 +1053,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list)
 		}
 	}
 
-	rcu_read_lock();
-	for_each_process_thread(g, t) {
-		if (start == end) {
-			ret = -EAGAIN;
-			goto unlock;
-		}
+	scoped_guard (rcu) {
+		for_each_process_thread(g, t) {
+			unsigned long *rs;
+
+			if (t->ret_stack)
+				continue;
+
+			rs = kmem_cache_alloc(fgraph_stack_cachep, GFP_NOWAIT);
+			if (!rs) {
+				/*
+				 * Returning from inside scoped_guard() drops
+				 * the RCU read lock, but skips the free loop
+				 * below. That is only safe because start ==
+				 * end here, which leaves that loop nothing to
+				 * free. Keep the two in step if this
+				 * exhaustion check ever changes.
+				 */
+				if (start == end)
+					return -EAGAIN;
+				rs = ret_stack_list[start++];
+			}
 
-		if (t->ret_stack == NULL) {
 			atomic_set(&t->trace_overrun, 0);
-			ret_stack_init_task_vars(ret_stack_list[start]);
+			ret_stack_init_task_vars(rs);
 			t->curr_ret_stack = 0;
 			t->curr_ret_depth = -1;
 			/* Make sure the tasks see the 0 first: */
-			smp_wmb();
-			t->ret_stack = ret_stack_list[start++];
+			smp_store_release(&t->ret_stack, rs);
 		}
 	}
 
-unlock:
-	rcu_read_unlock();
 free:
 	for (i = start; i < end; i++)
 		kmem_cache_free(fgraph_stack_cachep, ret_stack_list[i]);
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
       [not found]   ` <20260929010327.024281F000FF@smtp.kernel.org>
@ 2026-09-29  1:22     ` Vineet Gupta
  2026-09-29  2:25       ` Andrii Nakryiko
  0 siblings, 1 reply; 5+ messages in thread
From: Vineet Gupta @ 2026-09-29  1:22 UTC (permalink / raw)
  To: sashiko-reviews, Peter Zijlstra, rostedt, Andrii Nakryiko
  Cc: bpf, linux-trace-kernel, LKML, kernel-team

On 9/28/26 6:03 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Low] cleanup: mixing goto-based error handling and scope-based cleanup
> --
>
> commit a6de528fef00ef67da15965a7a8b3944fa92b053
> Author: Vineet Gupta <vineet.gupta@linux.dev>
> Subject: tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
>
> This commit aims to resolve performance regressions with ftrace graph tracing
> when there are a large number of tasks. It replaces the pre-allocated batch
> approach with inline allocation using GFP_NOWAIT, avoiding repeated sweeps of
> the task list. Additionally, it transitions the RCU read-side section to use
> scoped_guard(rcu) instead of manual rcu_read_lock and rcu_read_unlock calls.
>
>> diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c
>> index 40d373d65f9b9..1c5ccc3f80d53 100644
>> --- a/kernel/trace/fgraph.c
>> +++ b/kernel/trace/fgraph.c
> [ ... ]
>> @@ -1057,26 +1056,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list)
>>   		}
>>   	}
>>   
>> -	rcu_read_lock();
>> -	for_each_process_thread(g, t) {
>> -		if (start == end) {
>> -			ret = -EAGAIN;
>> -			goto unlock;
>> -		}
>> +	scoped_guard (rcu) {
> [Severity: Low]
> Does this mix goto-based error handling and scope-based cleanup within the same
> function?
>
> Looking at alloc_retstack_tasklist(), the function retains the legacy goto free
> path in the initial pre-allocation loop (just before this block), while
> introducing scoped_guard(rcu) here.
>
> The cleanup subsystem guidelines recommend that for a given routine, we should
> convert all resources that need a goto cleanup to scope-based cleanup, or
> convert none of them. Mixing these paradigms creates confusing ownership
> semantics and is highly prone to double-free or resource leak regressions upon
> future modifications.
>
> Could we either convert the ret_stack_list array cleanup to use scope-based
> helpers, or stick to the manual rcu_read_lock() and rcu_read_unlock() here to
> keep the error handling style consistent?

And a third option is to drop the fallback allocation completely and 
return -EAGAIN and come back. Granted I don't have much experience of 
typical handling of GFP_NOWAIT fails and retrying immediately: would 
that recover at all or does that take us back to where we started ? Peter

Thx,
-Vineet

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
  2026-09-29  1:22     ` Vineet Gupta
@ 2026-09-29  2:25       ` Andrii Nakryiko
  2026-09-29  8:03         ` Peter Zijlstra
  0 siblings, 1 reply; 5+ messages in thread
From: Andrii Nakryiko @ 2026-09-29  2:25 UTC (permalink / raw)
  To: Vineet Gupta
  Cc: sashiko-reviews, Peter Zijlstra, rostedt, Andrii Nakryiko, bpf,
	linux-trace-kernel, LKML, kernel-team

On Mon, Sep 28, 2026 at 6:22 PM Vineet Gupta <vineet.gupta@linux.dev> wrote:
>
> On 9/28/26 6:03 PM, sashiko-bot@kernel.org wrote:
> > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> > - [Low] cleanup: mixing goto-based error handling and scope-based cleanup
> > --
> >
> > commit a6de528fef00ef67da15965a7a8b3944fa92b053
> > Author: Vineet Gupta <vineet.gupta@linux.dev>
> > Subject: tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
> >
> > This commit aims to resolve performance regressions with ftrace graph tracing
> > when there are a large number of tasks. It replaces the pre-allocated batch
> > approach with inline allocation using GFP_NOWAIT, avoiding repeated sweeps of
> > the task list. Additionally, it transitions the RCU read-side section to use
> > scoped_guard(rcu) instead of manual rcu_read_lock and rcu_read_unlock calls.
> >
> >> diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c
> >> index 40d373d65f9b9..1c5ccc3f80d53 100644
> >> --- a/kernel/trace/fgraph.c
> >> +++ b/kernel/trace/fgraph.c
> > [ ... ]
> >> @@ -1057,26 +1056,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list)
> >>              }
> >>      }
> >>
> >> -    rcu_read_lock();
> >> -    for_each_process_thread(g, t) {
> >> -            if (start == end) {
> >> -                    ret = -EAGAIN;
> >> -                    goto unlock;
> >> -            }
> >> +    scoped_guard (rcu) {
> > [Severity: Low]
> > Does this mix goto-based error handling and scope-based cleanup within the same
> > function?
> >
> > Looking at alloc_retstack_tasklist(), the function retains the legacy goto free
> > path in the initial pre-allocation loop (just before this block), while
> > introducing scoped_guard(rcu) here.
> >
> > The cleanup subsystem guidelines recommend that for a given routine, we should
> > convert all resources that need a goto cleanup to scope-based cleanup, or
> > convert none of them. Mixing these paradigms creates confusing ownership
> > semantics and is highly prone to double-free or resource leak regressions upon
> > future modifications.
> >
> > Could we either convert the ret_stack_list array cleanup to use scope-based
> > helpers, or stick to the manual rcu_read_lock() and rcu_read_unlock() here to
> > keep the error handling style consistent?
>
> And a third option is to drop the fallback allocation completely and
> return -EAGAIN and come back. Granted I don't have much experience of
> typical handling of GFP_NOWAIT fails and retrying immediately: would
> that recover at all or does that take us back to where we started ? Peter
>

Not Peter, but I don't see a problem with scoped rcu and goto-based
free/cleanup. Unless Peter objects, let's keep it as is?

> Thx,
> -Vineet

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
  2026-09-29  2:25       ` Andrii Nakryiko
@ 2026-09-29  8:03         ` Peter Zijlstra
  0 siblings, 0 replies; 5+ messages in thread
From: Peter Zijlstra @ 2026-09-29  8:03 UTC (permalink / raw)
  To: Andrii Nakryiko
  Cc: Vineet Gupta, sashiko-reviews, rostedt, Andrii Nakryiko, bpf,
	linux-trace-kernel, LKML, kernel-team

On Mon, Sep 28, 2026 at 07:25:00PM -0700, Andrii Nakryiko wrote:
> On Mon, Sep 28, 2026 at 6:22 PM Vineet Gupta <vineet.gupta@linux.dev> wrote:
> >
> > On 9/28/26 6:03 PM, sashiko-bot@kernel.org wrote:
> > > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> > > - [Low] cleanup: mixing goto-based error handling and scope-based cleanup
> > > --
> > >
> > > commit a6de528fef00ef67da15965a7a8b3944fa92b053
> > > Author: Vineet Gupta <vineet.gupta@linux.dev>
> > > Subject: tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT
> > >
> > > This commit aims to resolve performance regressions with ftrace graph tracing
> > > when there are a large number of tasks. It replaces the pre-allocated batch
> > > approach with inline allocation using GFP_NOWAIT, avoiding repeated sweeps of
> > > the task list. Additionally, it transitions the RCU read-side section to use
> > > scoped_guard(rcu) instead of manual rcu_read_lock and rcu_read_unlock calls.
> > >
> > >> diff --git a/kernel/trace/fgraph.c b/kernel/trace/fgraph.c
> > >> index 40d373d65f9b9..1c5ccc3f80d53 100644
> > >> --- a/kernel/trace/fgraph.c
> > >> +++ b/kernel/trace/fgraph.c
> > > [ ... ]
> > >> @@ -1057,26 +1056,37 @@ static int alloc_retstack_tasklist(unsigned long **ret_stack_list)
> > >>              }
> > >>      }
> > >>
> > >> -    rcu_read_lock();
> > >> -    for_each_process_thread(g, t) {
> > >> -            if (start == end) {
> > >> -                    ret = -EAGAIN;
> > >> -                    goto unlock;
> > >> -            }
> > >> +    scoped_guard (rcu) {
> > > [Severity: Low]
> > > Does this mix goto-based error handling and scope-based cleanup within the same
> > > function?
> > >
> > > Looking at alloc_retstack_tasklist(), the function retains the legacy goto free
> > > path in the initial pre-allocation loop (just before this block), while
> > > introducing scoped_guard(rcu) here.
> > >
> > > The cleanup subsystem guidelines recommend that for a given routine, we should
> > > convert all resources that need a goto cleanup to scope-based cleanup, or
> > > convert none of them. Mixing these paradigms creates confusing ownership
> > > semantics and is highly prone to double-free or resource leak regressions upon
> > > future modifications.
> > >
> > > Could we either convert the ret_stack_list array cleanup to use scope-based
> > > helpers, or stick to the manual rcu_read_lock() and rcu_read_unlock() here to
> > > keep the error handling style consistent?
> >
> > And a third option is to drop the fallback allocation completely and
> > return -EAGAIN and come back. Granted I don't have much experience of
> > typical handling of GFP_NOWAIT fails and retrying immediately: would
> > that recover at all or does that take us back to where we started ? Peter
> >
> 
> Not Peter, but I don't see a problem with scoped rcu and goto-based
> free/cleanup. Unless Peter objects, let's keep it as is?

Yeah, the silly robot is being silly. Code is fine as is.

The guideline is just that, a guide. Its not a hard requirement. It is
possible to create a terrible mess of things when doing a partial
conversion of a large multi-stage goto unwind fest. But in general, goto
over the scope (as here) is fine. Also goto out of a scope is also fine.


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-29  8:03 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29  0:54 [PATCH v2] tracing: fgraph: fix the quadratic shadow stack walk Vineet Gupta
2026-09-29  0:54 ` [PATCH v2 1/1] tracing: fgraph: allocate shadow stacks inline with GFP_NOWAIT Vineet Gupta
     [not found]   ` <20260929010327.024281F000FF@smtp.kernel.org>
2026-09-29  1:22     ` Vineet Gupta
2026-09-29  2:25       ` Andrii Nakryiko
2026-09-29  8:03         ` Peter Zijlstra

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®