mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU
@ 2026-09-28 14:29 Masami Hiramatsu (Google)
  2026-09-28 14:29 ` [RFC PATCH 1/3] rcu/tasks: Export call_rcu_tasks_rude() Masami Hiramatsu (Google)
                   ` (3 more replies)
  0 siblings, 4 replies; 6+ messages in thread
From: Masami Hiramatsu (Google) @ 2026-09-28 14:29 UTC (permalink / raw)
  To: Paul E . McKenney, Frederic Weisbecker, Neeraj Upadhyay,
	Joel Fernandes, Josh Triplett, Boqun Feng, Uladzislau Rezki,
	Thomas Graf, Herbert Xu, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Steven Rostedt, Masami Hiramatsu, Andrew Morton
  Cc: Mathieu Desnoyers, Lai Jiangshan, Zqiang, John Fastabend,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, rcu, linux-kernel, linux-crypto,
	bpf, linux-trace-kernel

Hi,

Here is an RFC patch series which removes standard RCU read lock
(guard(rcu) and rcu_read_lock()) from fprobe callback paths
by switching fprobe and BPF multi-kprobe to Tasks-Rude RCU.

Motivation & Problem
====================

Currently, fprobe entry and exit callbacks (fprobe_fgraph_entry,
fprobe_return, and fprobe_ftrace_entry) execute in the tracing hot
path where preemption is disabled.

However, fprobe was forced to wrap hash lookups with guard(rcu)() and
rcu_read_lock() because:

- rhashtable defers bucket table deallocation using standard RCU
- unregister_fprobe() and BPF waited for a standard RCU grace period

Taking standard RCU read locks in the tracing fast path introduces
several drawbacks:

1. Unnecessary Runtime Overhead:
   Every probe hit manipulates current->rcu_read_lock_nesting with
   memory barriers, and rcu_read_unlock() adds conditional branches
   to check for special quiescent processing.
   On debug kernels with CONFIG_PROVE_RCU=y or CONFIG_LOCKDEP=y,
   this additionally acquires and releases lockdep maps on every hit,
   introducing severe lockdep hashing overhead and tracer recursion risks.

2. Fragile Dependency on rcu_is_watching():
   Standard RCU treats idle CPUs (and user-space on nohz_full CPUs) as
   Extended Quiescent States (EQS). If a function is traced while
   rcu_is_watching() is false, standard RCU is blind to the read-side
   critical section. In such contexts, synchronize_rcu() does not wait
   for the reader (risking Use-After-Free), and lockdep emits an
   "RCU-illegal: rcu_read_lock() used while not watching!" warning.
   To avoid this, ftrace callbacks normally require FTRACE_OPS_FL_RCU,
   adding extra trampoline check overhead.

3. Asymmetric Synchronization with fprobe_return():
   fprobe_return() executes under preempt_disable_notrace() without
   holding rcu_read_lock(). Prior to this series, unregister_fprobe()
   only waited on synchronize_rcu(), which does NOT wait for pure
   preempt-disabled sections, leaving a potential Use-After-Free window
   during probe unregistration.

I've tried to fix the last UAF with simply introducing guard(rcu)()[1]
but Sashiko found the 2nd problem [2]. So I decided to implement this
series.

[1] https://lore.kernel.org/all/179055575009.241711.6358052647499787191.stgit@devnote2/
[2] https://lore.kernel.org/all/20260928005114.9C9FC1F000FF@smtp.kernel.org/


Solution: Tasks-Rude RCU
========================

Because fprobe callbacks already run strictly within preempt-disabled
contexts, we can transition fprobe and its deferred table reclamation
to Tasks-Rude RCU:

- Tasks-Rude RCU detects grace periods via schedule_on_each_cpu(),
  forcing a schedule on every online CPU. This guarantees that all
  preempt-disabled sections that began prior to the grace period have
  completed before memory is reclaimed.
- Unlike standard RCU, Tasks-Rude RCU does not rely on dyntick-idle /
  EQS tracking. It does NOT require rcu_is_watching() to be true and
  does not trigger lockdep warnings in pre-RCU/idle execution paths.
- Within fprobe, guard(rcu)() is replaced with guard(rcu_sched_notrace)(),
  reducing the lookup lock to pure, non-tracing preempt counter
  increments without lockdep or RCU state manipulation.

Feedback and suggestions from RCU, BPF, and tracing maintainers are welcome!

Thank you,

---
base-commit: 5bfa9f1a9dcb6ecb607adbc1c0226605c972935b

Masami Hiramatsu (Google) (3):
      rcu/tasks: Export call_rcu_tasks_rude()
      rhashtable: Add use_tasks_rude parameter to defer bucket table free
      fprobe: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU


 include/linux/rcupdate.h         |    6 +++
 include/linux/rhashtable-types.h |    2 +
 kernel/bpf/syscall.c             |    2 +
 kernel/rcu/tasks.h               |    8 ++---
 kernel/trace/fprobe.c            |   66 +++++++++++++++++++++-----------------
 lib/rhashtable.c                 |    5 ++-
 6 files changed, 54 insertions(+), 35 deletions(-)

-- 
Masami Hiramatsu (Google) <mhiramat@kernel.org>

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-28 16:25 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 14:29 [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
2026-09-28 14:29 ` [RFC PATCH 1/3] rcu/tasks: Export call_rcu_tasks_rude() Masami Hiramatsu (Google)
2026-09-28 15:17   ` bot+bpf-ci
2026-09-28 14:29 ` [RFC PATCH 2/3] rhashtable: Add use_tasks_rude parameter to defer bucket table free Masami Hiramatsu (Google)
2026-09-28 14:29 ` [RFC PATCH 3/3] fprobe: Switch fprobe and BPF kprobe-multi to Tasks-Rude RCU Masami Hiramatsu (Google)
2026-09-28 16:25 ` [RFC PATCH 0/3] fprobe, rcu/tasks, rhashtable: " Paul E. McKenney

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®