* [RFC PATCH 1/3] rcu/tasks: Export call_rcu_tasks_rude()
2026-09-28 14:29 [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
@ 2026-09-28 14:29 ` Masami Hiramatsu (Google)
2026-09-28 15:17 ` bot+bpf-ci
2026-09-28 14:29 ` [RFC PATCH 2/3] rhashtable: Add use_tasks_rude parameter to defer bucket table free Masami Hiramatsu (Google)
` (2 subsequent siblings)
3 siblings, 1 reply; 6+ messages in thread
From: Masami Hiramatsu (Google) @ 2026-09-28 14:29 UTC (permalink / raw)
To: Paul E . McKenney, Frederic Weisbecker, Neeraj Upadhyay,
Joel Fernandes, Josh Triplett, Boqun Feng, Uladzislau Rezki,
Thomas Graf, Herbert Xu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Steven Rostedt, Masami Hiramatsu, Andrew Morton
Cc: Mathieu Desnoyers, Lai Jiangshan, Zqiang, John Fastabend,
Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, rcu, linux-kernel, linux-crypto,
bpf, linux-trace-kernel
From: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Currently call_rcu_tasks_rude() is static and private to
kernel/rcu/tasks.h, while synchronize_rcu_tasks_rude() is exported.
Subsystems that execute handlers under preempt_disable() (such as
fprobe and BPF kprobe-multi) cannot safely wait on
synchronize_rcu_tasks_rude() in atomic or asynchronous contexts
without an asynchronous call_rcu_tasks_rude() variant.
Make call_rcu_tasks_rude() non-static, declare it in
include/linux/rcupdate.h, and export it via EXPORT_SYMBOL_GPL().
Also provide fallbacks to call_rcu when Tasks Rude RCU is not
configured.
Assisted-by: LLM
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
---
include/linux/rcupdate.h | 6 ++++++
kernel/rcu/tasks.h | 8 +++-----
2 files changed, 9 insertions(+), 5 deletions(-)
diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h
index 44c07a66edff..dca61b52dc8c 100644
--- a/include/linux/rcupdate.h
+++ b/include/linux/rcupdate.h
@@ -198,7 +198,11 @@ void rcu_tasks_torture_stats_print(char *tt, char *tf);
# ifdef CONFIG_TASKS_RUDE_RCU
void synchronize_rcu_tasks_rude(void);
+void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func);
void rcu_tasks_rude_torture_stats_print(char *tt, char *tf);
+# else
+# define call_rcu_tasks_rude call_rcu
+# define synchronize_rcu_tasks_rude synchronize_rcu
# endif
#define rcu_note_voluntary_context_switch(t) rcu_tasks_qs(t, false)
@@ -210,6 +214,8 @@ void exit_tasks_rcu_finish(void);
#define rcu_note_voluntary_context_switch(t) do { } while (0)
#define call_rcu_tasks call_rcu
#define synchronize_rcu_tasks synchronize_rcu
+#define call_rcu_tasks_rude call_rcu
+#define synchronize_rcu_tasks_rude synchronize_rcu
static inline void exit_tasks_rcu_start(void) { }
static inline void exit_tasks_rcu_finish(void) { }
#endif /* #else #ifdef CONFIG_TASKS_RCU_GENERIC */
diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h
index 627295396cd9..f0b524d5f3f5 100644
--- a/kernel/rcu/tasks.h
+++ b/kernel/rcu/tasks.h
@@ -1355,7 +1355,7 @@ static void rcu_tasks_rude_wait_gp(struct rcu_tasks *rtp)
schedule_on_each_cpu(rcu_tasks_be_rude);
}
-static void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func);
+void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func);
DEFINE_RCU_TASKS(rcu_tasks_rude, rcu_tasks_rude_wait_gp, call_rcu_tasks_rude,
"RCU Tasks Rude");
@@ -1376,14 +1376,12 @@ DEFINE_RCU_TASKS(rcu_tasks_rude, rcu_tasks_rude_wait_gp, call_rcu_tasks_rude,
*
* See the description of call_rcu() for more detailed information on
* memory ordering guarantees.
- *
- * This is no longer exported, and is instead reserved for use by
- * synchronize_rcu_tasks_rude().
*/
-static void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func)
+void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func)
{
call_rcu_tasks_generic(rhp, func, &rcu_tasks_rude);
}
+EXPORT_SYMBOL_GPL(call_rcu_tasks_rude);
/**
* synchronize_rcu_tasks_rude - wait for a rude rcu-tasks grace period
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [RFC PATCH 1/3] rcu/tasks: Export call_rcu_tasks_rude()
2026-09-28 14:29 ` [RFC PATCH 1/3] rcu/tasks: Export call_rcu_tasks_rude() Masami Hiramatsu (Google)
@ 2026-09-28 15:17 ` bot+bpf-ci
0 siblings, 0 replies; 6+ messages in thread
From: bot+bpf-ci @ 2026-09-28 15:17 UTC (permalink / raw)
To: mhiramat, paulmck, frederic, neeraj.upadhyay, joelagnelf, josh,
boqun, urezki, tgraf, herbert, ast, daniel, andrii, eddyz87,
memxor, rostedt, mhiramat, akpm
Cc: mathieu.desnoyers, jiangshanlai, qiang.zhang, john.fastabend,
martin.lau, song, yonghong.song, jolsa, emil, ihor.solodrai, rcu,
linux-kernel, linux-crypto, bpf, linux-trace-kernel, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 5217 bytes --]
> diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h
> index 44c07a66edff..dca61b52dc8c 100644
> --- a/include/linux/rcupdate.h
> +++ b/include/linux/rcupdate.h
> @@ -198,7 +198,11 @@ void rcu_tasks_torture_stats_print(char *tt, char *tf);
>
> # ifdef CONFIG_TASKS_RUDE_RCU
> void synchronize_rcu_tasks_rude(void);
> +void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func);
> void rcu_tasks_rude_torture_stats_print(char *tt, char *tf);
> +# else
> +# define call_rcu_tasks_rude call_rcu
> +# define synchronize_rcu_tasks_rude synchronize_rcu
> # endif
>
> #define rcu_note_voluntary_context_switch(t) rcu_tasks_qs(t, false)
[ ... ]
> diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h
> index 627295396cd9..f0b524d5f3f5 100644
> --- a/kernel/rcu/tasks.h
> +++ b/kernel/rcu/tasks.h
[ ... ]
> @@ -1376,14 +1376,12 @@ DEFINE_RCU_TASKS(rcu_tasks_rude, rcu_tasks_rude_wait_gp, call_rcu_tasks_rude,
> *
> * See the description of call_rcu() for more detailed information on
> * memory ordering guarantees.
> - *
> - * This is no longer exported, and is instead reserved for use by
> - * synchronize_rcu_tasks_rude().
> */
> -static void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func)
> +void call_rcu_tasks_rude(struct rcu_head *rhp, rcu_callback_t func)
> {
> call_rcu_tasks_generic(rhp, func, &rcu_tasks_rude);
> }
> +EXPORT_SYMBOL_GPL(call_rcu_tasks_rude);
Do call_rcu_tasks_rude() and synchronize_rcu_tasks_rude() actually provide
the same grace period guarantee?
On architectures that select ARCH_WANTS_NO_INSTR (x86, arm64, riscv, s390,
loongarch), synchronize_rcu_tasks_rude() returns without waiting at all
unless CONFIG_FORCE_TASKS_RUDE_RCU is set:
kernel/rcu/tasks.h:synchronize_rcu_tasks_rude() {
if (!IS_ENABLED(CONFIG_ARCH_WANTS_NO_INSTR) ||
IS_ENABLED(CONFIG_FORCE_TASKS_RUDE_RCU))
synchronize_rcu_tasks_generic(&rcu_tasks_rude);
}
That option depends on RCU_EXPERT and defaults to n.
But call_rcu_tasks_rude() never takes that shortcut. Every callback goes
through this path:
call_rcu_tasks_rude()
-> call_rcu_tasks_generic()
-> rcu_tasks_one_gp()
-> rcu_tasks_rude_wait_gp()
-> schedule_on_each_cpu()
which always waits.
The rest of the series relies on these two functions being equivalent.
Commit f0b3c7fc5f82 ("fprobe: Switch fprobe and BPF kprobe-multi to
Tasks-Rude RCU") replaces synchronize_rcu() with
synchronize_rcu_tasks_rude() in unregister_fprobe() and in the error path
of register_fprobe_ips().
On a default x86 or arm64 build (FPROBE selects FUNCTION_GRAPH_TRACER,
which depends on FUNCTION_TRACER, which selects TASKS_RUDE_RCU),
unregister_fprobe() then returns without waiting for handlers that are
still running. Example path:
trace_fprobe_release()
-> unregister_trace_fprobe()
-> __unregister_trace_fprobe()
-> unregister_fprobe(&tf->fp)
-> unregister_fprobe_async() + synchronize_rcu_tasks_rude() [no-op]
-> free_trace_fprobe(tf)
-> kfree(tf)
Meanwhile another CPU can be inside fprobe_ftrace_entry() or
fprobe_return(), holding fp from READ_ONCE(node->fp) or from the shadow
stack, and calling fp->entry_handler / fp->exit_handler, which uses
container_of() to reach the freed trace_fprobe. The same thing happens when
the samples/fprobe module unloads after unregister_fprobe().
Before the series, synchronize_rcu() waited for these preempt-disabled
sections, because consolidated RCU waits for preempt-disabled regions where
RCU is watching.
Other points:
(1) The commit message says callers "cannot safely wait on
synchronize_rcu_tasks_rude()" and need an async variant. On these
architectures that wait is a no-op, so the premise is wrong, and the
synchronize_rcu_tasks_rude() kerneldoc ("Control will return to the
caller some time after a rude rcu-tasks grace period has elapsed") does
not match its implementation.
(2) Every grace period needed by a call_rcu_tasks_rude() callback now
queues work on and context-switches every online CPU, including idle
and nohz_full CPUs. For the synchronous API that cost was deliberately
removed on these architectures. Once the series moves fprobe/kprobe-multi
teardown and rhashtable resizing onto this API, the cost is paid on
every teardown and resize.
(3) The removed comment ("This is no longer exported, and is instead
reserved for use by synchronize_rcu_tasks_rude()") recorded a
deliberate decision to retire the async Rude API. The changelog does
not explain why that decision is being reversed.
Options include making call_rcu_tasks_rude() match
synchronize_rcu_tasks_rude(), for example falling back to call_rcu() under
the same ARCH_WANTS_NO_INSTR condition. Alternatively the series could keep
using call_rcu()/synchronize_rcu(): these already cover preempt-disabled
readers on noinstr architectures, and rcu_read_lock_sched_notrace() can
replace rcu_read_lock() in the handlers.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36438713281
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC PATCH 2/3] rhashtable: Add use_tasks_rude parameter to defer bucket table free
2026-09-28 14:29 [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
2026-09-28 14:29 ` [RFC PATCH 1/3] rcu/tasks: Export call_rcu_tasks_rude() Masami Hiramatsu (Google)
@ 2026-09-28 14:29 ` Masami Hiramatsu (Google)
2026-09-28 14:29 ` [RFC PATCH 3/3] fprobe: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
2026-09-28 16:25 ` [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: " Paul E. McKenney
3 siblings, 0 replies; 6+ messages in thread
From: Masami Hiramatsu (Google) @ 2026-09-28 14:29 UTC (permalink / raw)
To: Paul E . McKenney, Frederic Weisbecker, Neeraj Upadhyay,
Joel Fernandes, Josh Triplett, Boqun Feng, Uladzislau Rezki,
Thomas Graf, Herbert Xu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Steven Rostedt, Masami Hiramatsu, Andrew Morton
Cc: Mathieu Desnoyers, Lai Jiangshan, Zqiang, John Fastabend,
Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, rcu, linux-kernel, linux-crypto,
bpf, linux-trace-kernel
From: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Currently rhashtable always uses call_rcu() to free old bucket tables
upon table resizing. However, some callers (such as fprobe) operate
under preempt_disable() without holding rcu_read_lock(), relying on
Tasks Rude RCU grace periods instead of standard RCU.
Add a use_tasks_rude boolean flag to struct rhashtable_params. When
enabled, rhashtable_rehash_table() frees old bucket tables using
call_rcu_tasks_rude() instead of call_rcu(). Existing callers continue
to default to call_rcu().
Assisted-by: LLM
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
---
include/linux/rhashtable-types.h | 2 ++
lib/rhashtable.c | 5 ++++-
2 files changed, 6 insertions(+), 1 deletion(-)
diff --git a/include/linux/rhashtable-types.h b/include/linux/rhashtable-types.h
index 57c11ec9dc64..24282ae68409 100644
--- a/include/linux/rhashtable-types.h
+++ b/include/linux/rhashtable-types.h
@@ -52,6 +52,7 @@ typedef int (*rht_obj_cmpfn_t)(struct rhashtable_compare_arg *arg,
* @min_size: Minimum size while shrinking
* @insecure_elasticity: Set to true to disable chain length checks
* @automatic_shrinking: Enable automatic shrinking of tables
+ * @use_tasks_rude: Use call_rcu_tasks_rude() to free bucket tables
* @hashfn: Hash function (default: jhash2 if !(key_len % 4), or jhash)
* @obj_hashfn: Function to hash object
* @obj_cmpfn: Function to compare key with object
@@ -65,6 +66,7 @@ struct rhashtable_params {
u16 min_size;
bool insecure_elasticity;
bool automatic_shrinking;
+ bool use_tasks_rude;
rht_hashfn_t hashfn;
rht_obj_hashfn_t obj_hashfn;
rht_obj_cmpfn_t obj_cmpfn;
diff --git a/lib/rhashtable.c b/lib/rhashtable.c
index 6362896e4f09..b183fb112a70 100644
--- a/lib/rhashtable.c
+++ b/lib/rhashtable.c
@@ -359,7 +359,10 @@ static int rhashtable_rehash_table(struct rhashtable *ht)
* rhashtable_walk_stop() can use rcu_head_after_call_rcu()
* to check if it should not re-link the table.
*/
- call_rcu(&old_tbl->rcu, bucket_table_free_rcu);
+ if (ht->p.use_tasks_rude)
+ call_rcu_tasks_rude(&old_tbl->rcu, bucket_table_free_rcu);
+ else
+ call_rcu(&old_tbl->rcu, bucket_table_free_rcu);
spin_unlock(&ht->lock);
return rht_dereference(new_tbl->future_tbl, ht) ? -EAGAIN : 0;
^ permalink raw reply [flat|nested] 6+ messages in thread* [RFC PATCH 3/3] fprobe: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU
2026-09-28 14:29 [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
2026-09-28 14:29 ` [RFC PATCH 1/3] rcu/tasks: Export call_rcu_tasks_rude() Masami Hiramatsu (Google)
2026-09-28 14:29 ` [RFC PATCH 2/3] rhashtable: Add use_tasks_rude parameter to defer bucket table free Masami Hiramatsu (Google)
@ 2026-09-28 14:29 ` Masami Hiramatsu (Google)
2026-09-28 16:25 ` [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: " Paul E. McKenney
3 siblings, 0 replies; 6+ messages in thread
From: Masami Hiramatsu (Google) @ 2026-09-28 14:29 UTC (permalink / raw)
To: Paul E . McKenney, Frederic Weisbecker, Neeraj Upadhyay,
Joel Fernandes, Josh Triplett, Boqun Feng, Uladzislau Rezki,
Thomas Graf, Herbert Xu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Steven Rostedt, Masami Hiramatsu, Andrew Morton
Cc: Mathieu Desnoyers, Lai Jiangshan, Zqiang, John Fastabend,
Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, rcu, linux-kernel, linux-crypto,
bpf, linux-trace-kernel
From: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Currently, fprobe entry and exit callbacks are called from the tracing
path where preemption is disabled. However, because rhashtable deferred
bucket table reclamation using standard RCU and unregister_fprobe()
waited for standard RCU grace periods, guard(rcu)() and rcu_read_lock()
were used around hash lookups. Calling rcu_read_lock() in the trace path
introduces unnecessary overhead and potential recursion risks.
Furthermore, BPF_LINK_TYPE_KPROBE_MULTI attaches to fprobe via
register_fprobe_ips() and unregisters it asynchronously via
unregister_fprobe_async(), relying on bpf_link_free() to wait for
an RCU grace period before freeing the link structure.
If fprobe switches to Tasks-Rude RCU without simultaneously updating
BPF, a Use-After-Free race window opens during bpf_link_free() because
standard RCU grace periods do not wait for pure preempt-disabled
execution contexts to complete.
Atomically switch both fprobe and BPF kprobe-multi to Tasks-Rude RCU.
Assisted-by: LLM
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
---
kernel/bpf/syscall.c | 2 +
kernel/trace/fprobe.c | 66 +++++++++++++++++++++++++++----------------------
2 files changed, 39 insertions(+), 29 deletions(-)
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index c7bc9ba9b331..f9356262ff76 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -3350,6 +3350,8 @@ static void bpf_link_free(struct bpf_link *link)
/* We need to do a SRCU grace period wait for non-faultable tracepoint BPF links. */
else if (bpf_link_is_tracepoint(link))
call_tracepoint_unregister_atomic(&link->rcu, bpf_link_defer_dealloc_rcu_gp);
+ else if (link->type == BPF_LINK_TYPE_KPROBE_MULTI)
+ call_rcu_tasks_rude(&link->rcu, bpf_link_defer_dealloc_rcu_gp);
else
call_rcu(&link->rcu, bpf_link_defer_dealloc_rcu_gp);
} else if (ops->dealloc) {
diff --git a/kernel/trace/fprobe.c b/kernel/trace/fprobe.c
index 9f2d98181779..b7850df17575 100644
--- a/kernel/trace/fprobe.c
+++ b/kernel/trace/fprobe.c
@@ -36,11 +36,11 @@
*
* When unregistering the fprobe, fprobe_hlist::fp and fprobe_hlist::array[*].fp
* are set NULL and delete those from both hash tables (by hlist_del_rcu).
- * After an RCU grace period, the fprobe_hlist itself will be released.
+ * After a Tasks-Rude RCU grace period, the fprobe_hlist itself will be released.
*
* fprobe_table and fprobe_ip_table can be accessed from either
* - Normal hlist traversal and RCU add/del under 'fprobe_mutex' is held.
- * - RCU hlist traversal under disabling preempt
+ * - Tasks-Rude RCU / preempt-disabled hlist traversal
*/
static struct hlist_head fprobe_table[FPROBE_TABLE_SIZE];
static struct rhltable fprobe_ip_table;
@@ -76,8 +76,14 @@ static const struct rhashtable_params fprobe_rht_params = {
.obj_hashfn = fprobe_node_obj_hashfn,
.obj_cmpfn = fprobe_node_cmp,
.automatic_shrinking = true,
+ .use_tasks_rude = true,
};
+DEFINE_LOCK_GUARD_0(rcu_sched_notrace, rcu_read_lock_sched_notrace(),
+ rcu_read_unlock_sched_notrace())
+DECLARE_LOCK_GUARD_0_ATTRS(rcu_sched_notrace, __acquires_shared(RCU),
+ __releases_shared(RCU))
+
/* Node insertion and deletion requires the fprobe_mutex */
static int __insert_fprobe_node(struct fprobe_hlist_node *node, struct fprobe *fp)
{
@@ -333,27 +339,22 @@ static void fprobe_ftrace_entry(unsigned long ip, unsigned long parent_ip,
if (bit < 0)
return;
- /*
- * ftrace_test_recursion_trylock() disables preemption, but
- * rhltable_lookup() checks whether rcu_read_lcok is held.
- * So we take rcu_read_lock() here.
- */
- rcu_read_lock();
- head = rhltable_lookup(&fprobe_ip_table, &ip, fprobe_rht_params);
-
- rhl_for_each_entry_rcu(node, pos, head, hlist) {
- if (node->addr != ip)
- break;
- fp = READ_ONCE(node->fp);
- if (unlikely(!fp || fprobe_disabled(fp) || fp->exit_handler))
- continue;
+ scoped_guard(rcu_sched_notrace) {
+ head = rhltable_lookup(&fprobe_ip_table, &ip, fprobe_rht_params);
- if (fprobe_shared_with_kprobes(fp))
- __fprobe_kprobe_handler(ip, parent_ip, fp, fregs, NULL);
- else
- __fprobe_handler(ip, parent_ip, fp, fregs, NULL);
+ rhl_for_each_entry_rcu(node, pos, head, hlist) {
+ if (node->addr != ip)
+ break;
+ fp = READ_ONCE(node->fp);
+ if (unlikely(!fp || fprobe_disabled(fp) || fp->exit_handler))
+ continue;
+
+ if (fprobe_shared_with_kprobes(fp))
+ __fprobe_kprobe_handler(ip, parent_ip, fp, fregs, NULL);
+ else
+ __fprobe_handler(ip, parent_ip, fp, fregs, NULL);
+ }
}
- rcu_read_unlock();
ftrace_test_recursion_unlock(bit);
}
NOKPROBE_SYMBOL(fprobe_ftrace_entry);
@@ -452,7 +453,7 @@ static bool fprobe_exists_on_hash(unsigned long ip, bool ftrace)
struct fprobe_hlist_node *node;
struct fprobe *fp;
- guard(rcu)();
+ guard(rcu_sched_notrace)();
head = rhltable_lookup(&fprobe_ip_table, &ip,
fprobe_rht_params);
if (!head)
@@ -526,7 +527,7 @@ static bool fprobe_exists_on_hash(unsigned long ip, bool ftrace __maybe_unused)
struct fprobe_hlist_node *node;
struct fprobe *fp;
- guard(rcu)();
+ guard(rcu_sched_notrace)();
head = rhltable_lookup(&fprobe_ip_table, &ip,
fprobe_rht_params);
if (!head)
@@ -570,7 +571,7 @@ static int fprobe_fgraph_entry(struct ftrace_graph_ent *trace, struct fgraph_ops
if (WARN_ON_ONCE(!fregs))
return 0;
- guard(rcu)();
+ guard(rcu_sched_notrace)();
head = rhltable_lookup(&fprobe_ip_table, &func, fprobe_rht_params);
reserved_words = 0;
rhl_for_each_entry_rcu(node, pos, head, hlist) {
@@ -671,7 +672,7 @@ static void fprobe_return(struct ftrace_graph_ret *trace,
size_words = SIZE_IN_LONG(size);
ret_ip = ftrace_regs_get_instruction_pointer(fregs);
- preempt_disable_notrace();
+ guard(rcu_sched_notrace)();
curr = 0;
while (size_words > curr) {
@@ -687,7 +688,6 @@ static void fprobe_return(struct ftrace_graph_ret *trace,
}
curr += size;
}
- preempt_enable_notrace();
}
NOKPROBE_SYMBOL(fprobe_return);
@@ -1025,7 +1025,7 @@ int register_fprobe_ips(struct fprobe *fp, unsigned long *addrs, int num)
if (ret) {
unregister_fprobe_nolock(fp);
/* In error case, wait for clean up safely. */
- synchronize_rcu();
+ synchronize_rcu_tasks_rude();
}
return ret;
@@ -1070,6 +1070,14 @@ bool fprobe_is_registered(struct fprobe *fp)
return true;
}
+static void free_fprobe_hlist_array(struct rcu_head *head)
+{
+ struct fprobe_hlist *hlist_array;
+
+ hlist_array = container_of(head, struct fprobe_hlist, rcu);
+ kfree(hlist_array);
+}
+
static int unregister_fprobe_nolock(struct fprobe *fp)
{
struct fprobe_hlist *hlist_array = fp->hlist_array;
@@ -1101,7 +1109,7 @@ static int unregister_fprobe_nolock(struct fprobe *fp)
else
fprobe_graph_remove_ips(addrs, count);
- kfree_rcu(hlist_array, rcu);
+ call_rcu_tasks_rude(&hlist_array->rcu, free_fprobe_hlist_array);
fp->hlist_array = NULL;
kfree(addrs);
@@ -1140,7 +1148,7 @@ int unregister_fprobe(struct fprobe *fp)
int ret = unregister_fprobe_async(fp);
if (!ret)
- synchronize_rcu();
+ synchronize_rcu_tasks_rude();
return ret;
}
EXPORT_SYMBOL_GPL(unregister_fprobe);
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU
2026-09-28 14:29 [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
` (2 preceding siblings ...)
2026-09-28 14:29 ` [RFC PATCH 3/3] fprobe: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
@ 2026-09-28 16:25 ` Paul E. McKenney
3 siblings, 0 replies; 6+ messages in thread
From: Paul E. McKenney @ 2026-09-28 16:25 UTC (permalink / raw)
To: Masami Hiramatsu (Google)
Cc: Frederic Weisbecker, Neeraj Upadhyay, Joel Fernandes,
Josh Triplett, Boqun Feng, Uladzislau Rezki, Thomas Graf,
Herbert Xu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi, Steven Rostedt,
Andrew Morton, Mathieu Desnoyers, Lai Jiangshan, Zqiang,
John Fastabend, Martin KaFai Lau, Song Liu, Yonghong Song,
Jiri Olsa, Emil Tsalapatis, Ihor Solodrai, rcu, linux-kernel,
linux-crypto, bpf, linux-trace-kernel, Roman Gushchin,
Chris Mason
On Mon, Sep 28, 2026 at 11:29:13PM +0900, Masami Hiramatsu (Google) wrote:
> Hi,
>
> Here is an RFC patch series which removes standard RCU read lock
> (guard(rcu) and rcu_read_lock()) from fprobe callback paths
> by switching fprobe and BPF multi-kprobe to Tasks-Rude RCU.
>
> Motivation & Problem
> ====================
>
> Currently, fprobe entry and exit callbacks (fprobe_fgraph_entry,
> fprobe_return, and fprobe_ftrace_entry) execute in the tracing hot
> path where preemption is disabled.
>
> However, fprobe was forced to wrap hash lookups with guard(rcu)() and
> rcu_read_lock() because:
>
> - rhashtable defers bucket table deallocation using standard RCU
> - unregister_fprobe() and BPF waited for a standard RCU grace period
>
> Taking standard RCU read locks in the tracing fast path introduces
> several drawbacks:
>
> 1. Unnecessary Runtime Overhead:
> Every probe hit manipulates current->rcu_read_lock_nesting with
> memory barriers, and rcu_read_unlock() adds conditional branches
> to check for special quiescent processing.
> On debug kernels with CONFIG_PROVE_RCU=y or CONFIG_LOCKDEP=y,
> this additionally acquires and releases lockdep maps on every hit,
> introducing severe lockdep hashing overhead and tracer recursion risks.
>
> 2. Fragile Dependency on rcu_is_watching():
> Standard RCU treats idle CPUs (and user-space on nohz_full CPUs) as
> Extended Quiescent States (EQS). If a function is traced while
> rcu_is_watching() is false, standard RCU is blind to the read-side
> critical section. In such contexts, synchronize_rcu() does not wait
> for the reader (risking Use-After-Free), and lockdep emits an
> "RCU-illegal: rcu_read_lock() used while not watching!" warning.
> To avoid this, ftrace callbacks normally require FTRACE_OPS_FL_RCU,
> adding extra trampoline check overhead.
>
> 3. Asymmetric Synchronization with fprobe_return():
> fprobe_return() executes under preempt_disable_notrace() without
> holding rcu_read_lock(). Prior to this series, unregister_fprobe()
> only waited on synchronize_rcu(), which does NOT wait for pure
> preempt-disabled sections, leaving a potential Use-After-Free window
> during probe unregistration.
>
> I've tried to fix the last UAF with simply introducing guard(rcu)()[1]
> but Sashiko found the 2nd problem [2]. So I decided to implement this
> series.
Unless I am missing something subtle, Sashiko needs to be taught a
little bit more about RCU. Preemption-disabled regions of code really
are valid RCU readers. If you have a reproducer showing that this is
not the case in some situation, that would be a bug in RCU.
Adding Roman Gushchin and Chris Mason on CC for their thoughts.
Thanx, Paul
> [1] https://lore.kernel.org/all/179055575009.241711.6358052647499787191.stgit@devnote2/
> [2] https://lore.kernel.org/all/20260928005114.9C9FC1F000FF@smtp.kernel.org/
>
>
> Solution: Tasks-Rude RCU
> ========================
>
> Because fprobe callbacks already run strictly within preempt-disabled
> contexts, we can transition fprobe and its deferred table reclamation
> to Tasks-Rude RCU:
>
> - Tasks-Rude RCU detects grace periods via schedule_on_each_cpu(),
> forcing a schedule on every online CPU. This guarantees that all
> preempt-disabled sections that began prior to the grace period have
> completed before memory is reclaimed.
> - Unlike standard RCU, Tasks-Rude RCU does not rely on dyntick-idle /
> EQS tracking. It does NOT require rcu_is_watching() to be true and
> does not trigger lockdep warnings in pre-RCU/idle execution paths.
> - Within fprobe, guard(rcu)() is replaced with guard(rcu_sched_notrace)(),
> reducing the lookup lock to pure, non-tracing preempt counter
> increments without lockdep or RCU state manipulation.
>
> Feedback and suggestions from RCU, BPF, and tracing maintainers are welcome!
>
> Thank you,
>
> ---
> base-commit: 5bfa9f1a9dcb6ecb607adbc1c0226605c972935b
>
> Masami Hiramatsu (Google) (3):
> rcu/tasks: Export call_rcu_tasks_rude()
> rhashtable: Add use_tasks_rude parameter to defer bucket table free
> fprobe: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU
>
>
> include/linux/rcupdate.h | 6 +++
> include/linux/rhashtable-types.h | 2 +
> kernel/bpf/syscall.c | 2 +
> kernel/rcu/tasks.h | 8 ++---
> kernel/trace/fprobe.c | 66 +++++++++++++++++++++-----------------
> lib/rhashtable.c | 5 ++-
> 6 files changed, 54 insertions(+), 35 deletions(-)
>
> --
> Masami Hiramatsu (Google) <mhiramat@kernel.org>
^ permalink raw reply [flat|nested] 6+ messages in thread