* [PATCH hazptr 0/4] Hazard pointer updates
@ 2026-09-27 15:51 Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
` (5 more replies)
0 siblings, 6 replies; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 15:51 UTC (permalink / raw)
To: Paul E . McKenney
Cc: linux-kernel, Mathieu Desnoyers, Boqun Feng, Bradley Morgan,
Gary Guo, rcu, lkmm
Hi Paul,
This series applies on top of "hazptr: handle NULL address in
hazptr_detach" you have in your rcu dev tree.
This first patch addresses a race identified by Boqun Feng in the
two-phase wildcard scheme.
Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
Those were discussed at length in a prior version of hazard pointer
patches.
Patch 4 introduces a "try acquire" helper to allow the fast path
to not rely on wildcards, while keeping the wildcard forward
progress guarantees in the acquire slow path, used on fast path
failure.
Thanks,
Mathieu
Mathieu Desnoyers (4):
hazptr: Fix two-phase hazptr_synchronize race with detach
compiler.h: Introduce ptr_eq() to preserve address dependency
Documentation: RCU: Refer to ptr_eq()
hazptr: Introduce "try acquire" fast path, fallback to overflow list
Cc: Paul E. McKenney <paulmck@kernel.org>
Cc: Boqun Feng <boqun@kernel.org>
Cc: Bradley Morgan <brads@mainlining.org>
Cc: Gary Guo <gary@garyguo.net>
Cc: <rcu@vger.kernel.org>
Cc: <lkmm@lists.linux.dev>
Documentation/RCU/rcu_dereference.rst | 38 +++++++-
include/linux/compiler.h | 63 ++++++++++++
include/linux/hazptr.h | 47 +++++----
kernel/hazptr.c | 135 +++++++++++++++-----------
4 files changed, 203 insertions(+), 80 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 22+ messages in thread
* [PATCH hazptr 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach
2026-09-27 15:51 [PATCH hazptr 0/4] Hazard pointer updates Mathieu Desnoyers
@ 2026-09-27 15:51 ` Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 2/4] compiler.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
` (4 subsequent siblings)
5 siblings, 0 replies; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 15:51 UTC (permalink / raw)
To: Paul E . McKenney
Cc: linux-kernel, Mathieu Desnoyers, Boqun Feng, Bradley Morgan,
Gary Guo, rcu, lkmm
Boqun Feng pointed out that a detach happening concurrently with
hazptr_synchronize can miss a slot. Indeed, scanning the per-CPU
slots needs to be done *before* scanning the overflow lists. Fix the
implementation accordingly.
Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Reported-by: Boqun Feng <boqun@kernel.org>
Cc: Paul E. McKenney <paulmck@kernel.org>
Cc: Boqun Feng <boqun@kernel.org>
Cc: Bradley Morgan <brads@mainlining.org>
Cc: Gary Guo <gary@garyguo.net>
Cc: <rcu@vger.kernel.org>
Cc: <lkmm@lists.linux.dev>
---
Changes since v0:
- Load the new overflow list phase word in "hazptr_chain_backup_slot".
- Rename the lock to hazptr_phase_lock to make it clear that it
protects both the wildcard and the overflow list phases.
---
kernel/hazptr.c | 50 +++++++++++++++++++++++++++++++++++++++----------
1 file changed, 40 insertions(+), 10 deletions(-)
diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index d3d1050d92cf..13faa5ba7677 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -13,6 +13,8 @@
#include <linux/list.h>
#include <linux/export.h>
+static DEFINE_MUTEX(hazptr_phase_lock); /* Protect the wildcard and list phase flip. */
+
/*
* The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
* hazptr_synchronize forward progress even with a steady stream of readers.
@@ -20,10 +22,12 @@
* This also affects the overflow list selection: the current list used by
* readers is array[(unsigned long) hazptr_wildcard - 1].
*/
-static DEFINE_MUTEX(hazptr_wildcard_lock); /* Protect the wildcard flip. */
void *hazptr_wildcard = (void *) 1UL;
EXPORT_SYMBOL_GPL(hazptr_wildcard);
+/* The current overflow list phase. */
+static unsigned int hazptr_overflow_list_phase;
+
struct hazptr_overflow_list {
raw_spinlock_t lock; /* Lock protecting overflow list and list generation. */
struct hlist_head head; /* Overflow list head. */
@@ -53,6 +57,12 @@ void *flip_wildcard(void *wildcard)
return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
}
+static
+unsigned int flip_list_phase(unsigned int phase)
+{
+ return 1 - phase;
+}
+
static
bool is_wildcard(void *addr)
{
@@ -178,15 +188,12 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
}
static
-void hazptr_scan_period(void *addr, void *scan_wildcard)
+void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
{
- unsigned int scan_idx = (unsigned long) scan_wildcard - 1;
int cpu;
/* Scan all CPUs slots. */
for_each_possible_cpu(cpu) {
- struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
-
/*
* Scan CPU slots.
* Forward progress against recurring wildcards is guaranteed
@@ -199,6 +206,17 @@ void hazptr_scan_period(void *addr, void *scan_wildcard)
* to acquire that same hazard pointer value.
*/
hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
+ }
+}
+
+static
+void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
+{
+ int cpu;
+
+ /* Scan all CPUs overflow lists. */
+ for_each_possible_cpu(cpu) {
+ struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
/*
* Scan backup slots in percpu overflow lists.
@@ -218,6 +236,7 @@ void hazptr_scan_period(void *addr, void *scan_wildcard)
*/
void hazptr_synchronize(void *addr)
{
+ unsigned int scan_list_phase;
void *scan_wildcard;
/*
@@ -235,18 +254,29 @@ void hazptr_synchronize(void *addr)
/* Memory ordering: Store A before Load B. */
smp_mb();
- guard(mutex)(&hazptr_wildcard_lock);
+ guard(mutex)(&hazptr_phase_lock);
+
+ /* Scan per-CPU slots. */
scan_wildcard = flip_wildcard(hazptr_wildcard);
- hazptr_scan_period(addr, scan_wildcard);
- WRITE_ONCE(hazptr_wildcard, scan_wildcard); /* Flip the current wildcard. */
- hazptr_scan_period(addr, flip_wildcard(scan_wildcard));
+ hazptr_scan_cpu_slots_period(addr, scan_wildcard);
+ WRITE_ONCE(hazptr_wildcard, scan_wildcard); /* Flip the current wildcard. */
+ hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
+
+ /*
+ * Scan overflow lists *after* scanning per-CPU slots. See
+ * hazptr_promote_to_backup_slot() for scan ordering requirement.
+ */
+ scan_list_phase = flip_list_phase(hazptr_overflow_list_phase);
+ hazptr_scan_overflow_list_period(addr, scan_list_phase);
+ WRITE_ONCE(hazptr_overflow_list_phase, scan_list_phase); /* Flip the current list phase. */
+ hazptr_scan_overflow_list_period(addr, flip_list_phase(scan_list_phase));
}
EXPORT_SYMBOL_GPL(hazptr_synchronize);
struct hazptr_slot *hazptr_chain_backup_slot(struct hazptr_ctx *ctx)
{
struct hazptr_overflow_list_flip *overflow_list_flip = this_cpu_ptr(&percpu_overflow_list_flip);
- unsigned int list_idx = (unsigned long) READ_ONCE(hazptr_wildcard) - 1;
+ unsigned int list_idx = READ_ONCE(hazptr_overflow_list_phase);
struct hazptr_overflow_list *overflow_list = &overflow_list_flip->array[list_idx];
struct hazptr_slot *slot = &ctx->backup_slot.slot;
--
2.43.0
^ permalink raw reply [flat|nested] 22+ messages in thread
* [PATCH hazptr 2/4] compiler.h: Introduce ptr_eq() to preserve address dependency
2026-09-27 15:51 [PATCH hazptr 0/4] Hazard pointer updates Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
@ 2026-09-27 15:51 ` Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 3/4] Documentation: RCU: Refer to ptr_eq() Mathieu Desnoyers
` (3 subsequent siblings)
5 siblings, 0 replies; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 15:51 UTC (permalink / raw)
To: Paul E . McKenney
Cc: linux-kernel, Mathieu Desnoyers, Linus Torvalds, Boqun Feng,
Joel Fernandes, Alan Stern, Greg Kroah-Hartman,
Sebastian Andrzej Siewior, Will Deacon, Peter Zijlstra,
John Stultz, Neeraj Upadhyay, Frederic Weisbecker, Josh Triplett,
Uladzislau Rezki, Steven Rostedt, Lai Jiangshan, Zqiang,
Ingo Molnar, Waiman Long, Mark Rutland, Thomas Gleixner,
Vlastimil Babka, maged.michael, Mateusz Guzik, Gary Guo,
Jonas Oberhauser, rcu, linux-mm, lkmm, Nikita Popov, llvm
Compiler CSE and SSA GVN optimizations can cause the address dependency
of addresses returned by rcu_dereference to be lost when comparing those
pointers with either constants or previously loaded pointers.
Introduce ptr_eq() to compare two addresses while preserving the address
dependencies for later use of the address. It should be used when
comparing an address returned by rcu_dereference().
This is needed to prevent the compiler CSE and SSA GVN optimizations
from using @a (or @b) in places where the source refers to @b (or @a)
based on the fact that after the comparison, the two are known to be
equal, which does not preserve address dependencies and allows the
following misordering speculations:
- If @b is a constant, the compiler can issue the loads which depend
on @a before loading @a.
- If @b is a register populated by a prior load, weakly-ordered
CPUs can speculate loads which depend on @a before loading @a.
The same logic applies with @a and @b swapped.
Suggested-by: Linus Torvalds <torvalds@linux-foundation.org>
Suggested-by: Boqun Feng <boqun.feng@gmail.com>
Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Reviewed-by: Boqun Feng <boqun.feng@gmail.com>
Reviewed-by: Joel Fernandes (Google) <joel@joelfernandes.org>
Tested-by: Joel Fernandes (Google) <joel@joelfernandes.org>
Acked-by: "Paul E. McKenney" <paulmck@kernel.org>
Acked-by: Alan Stern <stern@rowland.harvard.edu>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Cc: "Paul E. McKenney" <paulmck@kernel.org>
Cc: Will Deacon <will@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Alan Stern <stern@rowland.harvard.edu>
Cc: John Stultz <jstultz@google.com>
Cc: Neeraj Upadhyay <Neeraj.Upadhyay@amd.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Frederic Weisbecker <frederic@kernel.org>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Josh Triplett <josh@joshtriplett.org>
Cc: Uladzislau Rezki <urezki@gmail.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Lai Jiangshan <jiangshanlai@gmail.com>
Cc: Zqiang <qiang.zhang1211@gmail.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Waiman Long <longman@redhat.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Vlastimil Babka <vbabka@suse.cz>
Cc: maged.michael@gmail.com
Cc: Mateusz Guzik <mjguzik@gmail.com>
Cc: Gary Guo <gary@garyguo.net>
Cc: Jonas Oberhauser <jonas.oberhauser@huaweicloud.com>
Cc: rcu@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: lkmm@lists.linux.dev
Cc: Nikita Popov <github@npopov.com>
Cc: llvm@lists.linux.dev
---
Changes since v0:
- Include feedback from Alan Stern.
---
include/linux/compiler.h | 63 ++++++++++++++++++++++++++++++++++++++++
1 file changed, 63 insertions(+)
diff --git a/include/linux/compiler.h b/include/linux/compiler.h
index 70cb31d60538..0a8c8e53cdad 100644
--- a/include/linux/compiler.h
+++ b/include/linux/compiler.h
@@ -166,6 +166,69 @@ void ftrace_likely_update(struct ftrace_likely_data *f, int val,
__PASTE(name, \
__PASTE(_, __COUNTER__)))
+/*
+ * Compare two addresses while preserving the address dependencies for
+ * later use of the address. It should be used when comparing an address
+ * returned by rcu_dereference().
+ *
+ * This is needed to prevent the compiler CSE and SSA GVN optimizations
+ * from using @a (or @b) in places where the source refers to @b (or @a)
+ * based on the fact that after the comparison, the two are known to be
+ * equal, which does not preserve address dependencies and allows the
+ * following misordering speculations:
+ *
+ * - If @b is a constant, the compiler can issue the loads which depend
+ * on @a before loading @a.
+ * - If @b is a register populated by a prior load, weakly-ordered
+ * CPUs can speculate loads which depend on @a before loading @a.
+ *
+ * The same logic applies with @a and @b swapped.
+ *
+ * Return value: true if pointers are equal, false otherwise.
+ *
+ * The compiler barrier() is ineffective at fixing this issue. It does
+ * not prevent the compiler CSE from losing the address dependency:
+ *
+ * int fct_2_volatile_barriers(void)
+ * {
+ * int *a, *b;
+ *
+ * do {
+ * a = READ_ONCE(p);
+ * asm volatile ("" : : : "memory");
+ * b = READ_ONCE(p);
+ * } while (a != b);
+ * asm volatile ("" : : : "memory"); <-- barrier()
+ * return *b;
+ * }
+ *
+ * With gcc 14.2 (arm64):
+ *
+ * fct_2_volatile_barriers:
+ * adrp x0, .LANCHOR0
+ * add x0, x0, :lo12:.LANCHOR0
+ * .L2:
+ * ldr x1, [x0] <-- x1 populated by first load.
+ * ldr x2, [x0]
+ * cmp x1, x2
+ * bne .L2
+ * ldr w0, [x1] <-- x1 is used for access which should depend on b.
+ * ret
+ *
+ * On weakly-ordered architectures, this lets CPU speculation use the
+ * result from the first load to speculate "ldr w0, [x1]" before
+ * "ldr x2, [x0]".
+ * Based on the RCU documentation, the control dependency does not
+ * prevent the CPU from speculating loads.
+ */
+static __always_inline
+int ptr_eq(const volatile void *a, const volatile void *b)
+{
+ OPTIMIZER_HIDE_VAR(a);
+ OPTIMIZER_HIDE_VAR(b);
+ return a == b;
+}
+
/**
* data_race - mark an expression as containing intentional data races
*
--
2.43.0
^ permalink raw reply [flat|nested] 22+ messages in thread
* [PATCH hazptr 3/4] Documentation: RCU: Refer to ptr_eq()
2026-09-27 15:51 [PATCH hazptr 0/4] Hazard pointer updates Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 2/4] compiler.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
@ 2026-09-27 15:51 ` Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
` (2 subsequent siblings)
5 siblings, 0 replies; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 15:51 UTC (permalink / raw)
To: Paul E . McKenney
Cc: linux-kernel, Mathieu Desnoyers, Alan Stern, Joel Fernandes,
Greg Kroah-Hartman, Sebastian Andrzej Siewior, Will Deacon,
Peter Zijlstra, Boqun Feng, John Stultz, Neeraj Upadhyay,
Linus Torvalds, Frederic Weisbecker, Josh Triplett,
Uladzislau Rezki, Steven Rostedt, Lai Jiangshan, Zqiang,
Ingo Molnar, Waiman Long, Mark Rutland, Thomas Gleixner,
Vlastimil Babka, maged.michael, Mateusz Guzik, Gary Guo,
Jonas Oberhauser, rcu, linux-mm, lkmm, Nikita Popov, llvm
Refer to ptr_eq() in the rcu_dereference() documentation.
ptr_eq() is a mechanism that preserves address dependencies when
comparing pointers, and should be favored when comparing a pointer
obtained from rcu_dereference() against another pointer.
Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Acked-by: Alan Stern <stern@rowland.harvard.edu>
Acked-by: Paul E. McKenney <paulmck@kernel.org>
Reviewed-by: Joel Fernandes (Google) <joel@joelfernandes.org>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Cc: "Paul E. McKenney" <paulmck@kernel.org>
Cc: Will Deacon <will@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Alan Stern <stern@rowland.harvard.edu>
Cc: John Stultz <jstultz@google.com>
Cc: Neeraj Upadhyay <Neeraj.Upadhyay@amd.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Frederic Weisbecker <frederic@kernel.org>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Josh Triplett <josh@joshtriplett.org>
Cc: Uladzislau Rezki <urezki@gmail.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Lai Jiangshan <jiangshanlai@gmail.com>
Cc: Zqiang <qiang.zhang1211@gmail.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Waiman Long <longman@redhat.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Vlastimil Babka <vbabka@suse.cz>
Cc: maged.michael@gmail.com
Cc: Mateusz Guzik <mjguzik@gmail.com>
Cc: Gary Guo <gary@garyguo.net>
Cc: Jonas Oberhauser <jonas.oberhauser@huaweicloud.com>
Cc: rcu@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: lkmm@lists.linux.dev
Cc: Nikita Popov <github@npopov.com>
Cc: llvm@lists.linux.dev
---
Changes since v1:
- Include feedback from Paul E. McKenney.
Changes since v0:
- Include feedback from Alan Stern.
---
Documentation/RCU/rcu_dereference.rst | 38 +++++++++++++++++++++++----
1 file changed, 33 insertions(+), 5 deletions(-)
diff --git a/Documentation/RCU/rcu_dereference.rst b/Documentation/RCU/rcu_dereference.rst
index 5bc3785ebfc2..85b5f79ebf14 100644
--- a/Documentation/RCU/rcu_dereference.rst
+++ b/Documentation/RCU/rcu_dereference.rst
@@ -104,11 +104,12 @@ readers working properly:
after such branches, but can speculate loads, which can again
result in misordering bugs.
-- Be very careful about comparing pointers obtained from
- rcu_dereference() against non-NULL values. As Linus Torvalds
- explained, if the two pointers are equal, the compiler could
- substitute the pointer you are comparing against for the pointer
- obtained from rcu_dereference(). For example::
+- Use operations that preserve address dependencies (such as
+ "ptr_eq()") to compare pointers obtained from rcu_dereference()
+ against non-NULL pointers. As Linus Torvalds explained, if the
+ two pointers are equal, the compiler could substitute the
+ pointer you are comparing against for the pointer obtained from
+ rcu_dereference(). For example::
p = rcu_dereference(gp);
if (p == &default_struct)
@@ -125,6 +126,29 @@ readers working properly:
On ARM and Power hardware, the load from "default_struct.a"
can now be speculated, such that it might happen before the
rcu_dereference(). This could result in bugs due to misordering.
+ Performing the comparison with "ptr_eq()" ensures the compiler
+ does not perform such transformation.
+
+ If the comparison is against another pointer, the compiler is
+ allowed to use either pointer for the following accesses, which
+ loses the address dependency and allows weakly-ordered
+ architectures such as ARM and PowerPC to speculate the
+ address-dependent load before rcu_dereference(). For example::
+
+ p1 = READ_ONCE(gp);
+ p2 = rcu_dereference(gp);
+ if (p1 == p2) /* BUGGY!!! */
+ do_default(p2->a);
+
+ The compiler can use p1->a rather than p2->a, destroying the
+ address dependency. Performing the comparison with "ptr_eq()"
+ ensures the compiler preserves the address dependencies.
+ Corrected code::
+
+ p1 = READ_ONCE(gp);
+ p2 = rcu_dereference(gp);
+ if (ptr_eq(p1, p2))
+ do_default(p2->a);
However, comparisons are OK in the following cases:
@@ -204,6 +228,10 @@ readers working properly:
comparison will provide exactly the information that the
compiler needs to deduce the value of the pointer.
+ When in doubt, use operations that preserve address dependencies
+ (such as "ptr_eq()") to compare pointers obtained from
+ rcu_dereference() against non-NULL pointers.
+
- Disable any value-speculation optimizations that your compiler
might provide, especially if you are making use of feedback-based
optimizations that take data collected from prior runs. Such
--
2.43.0
^ permalink raw reply [flat|nested] 22+ messages in thread
* [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 15:51 [PATCH hazptr 0/4] Hazard pointer updates Mathieu Desnoyers
` (2 preceding siblings ...)
2026-09-27 15:51 ` [PATCH hazptr 3/4] Documentation: RCU: Refer to ptr_eq() Mathieu Desnoyers
@ 2026-09-27 15:51 ` Mathieu Desnoyers
2026-09-27 16:40 ` Boqun Feng
2026-09-27 16:07 ` [PATCH hazptr 0/4] Hazard pointer updates Bradley Morgan
2026-09-27 16:15 ` Boqun Feng
5 siblings, 1 reply; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 15:51 UTC (permalink / raw)
To: Paul E . McKenney
Cc: linux-kernel, Mathieu Desnoyers, Boqun Feng, Bradley Morgan,
Gary Guo, rcu, lkmm
Introduce a "try acquire" hazard pointer fast path, which performs an
early load of the address to store it into the hazard pointer slot, and
then re-loads that address after a barrier to check whether it has
changed meanwhile.
On comparison failure, rather than re-try, guarantee forward progress by
falling back to the __hazptr_acquire slow path on failure.
The acquire slow path attempts a try-acquire for any available per-CPU
slot. If that fails, it chains the backup slot into the overflow list,
therefore guaranteeing forward progress for both hazard pointer
read-side and synchronize:
- Readers set the wildcard, and then proceed to set the more
specific address to replace the wildcard.
- One synchronize alternates between two overflow list periods,
scanning each one while readers are added to the other period,
thus preventing a steady flow of readers from preventing
synchronize forward progress.
With this change, the scan on per-CPU slots don't need to expect a
wildcard anymore, because none can be produced by readers. Wildcards are
only expected within overflow lists.
Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Paul E. McKenney <paulmck@kernel.org>
Cc: Boqun Feng <boqun@kernel.org>
Cc: Bradley Morgan <brads@mainlining.org>
Cc: Gary Guo <gary@garyguo.net>
Cc: <rcu@vger.kernel.org>
Cc: <lkmm@lists.linux.dev>
---
include/linux/hazptr.h | 47 +++++++++++--------
kernel/hazptr.c | 103 ++++++++++++++++++++---------------------
2 files changed, 76 insertions(+), 74 deletions(-)
diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
index d1670121947a..fcf017ca2255 100644
--- a/include/linux/hazptr.h
+++ b/include/linux/hazptr.h
@@ -29,9 +29,6 @@
/* 4 slots (each sizeof(hazptr_slot_item)) fit in a single 64-byte cache line. */
#define NR_HAZPTR_PERCPU_SLOTS 4
-/* The current hazard pointer wildcard. */
-extern void *hazptr_wildcard;
-
/*
* Hazard pointer slot.
*/
@@ -190,6 +187,31 @@ void hazptr_note_context_switch(void)
}
}
+/* Try hazard pointer protection. */
+static inline
+void *__hazptr_try_acquire(struct hazptr_ctx *ctx, void * const *addr_p, struct hazptr_slot *slot)
+{
+ void *early_addr, *addr;
+
+ if (unlikely(slot->addr))
+ return NULL;
+ early_addr = READ_ONCE(*addr_p); /* Early load. */
+ WRITE_ONCE(slot->addr, early_addr); /* Store B */
+ /* Memory ordering: Store B before Load A. */
+ smp_mb();
+ addr = READ_ONCE(*addr_p); /* Load A */
+ /*
+ * Validate that address did not change between Early load and Load A.
+ * Use ptr_eq() to make sure that result from Load A is returned to the
+ * caller to preserve address dependency.
+ */
+ if (unlikely(!ptr_eq(addr, early_addr))) {
+ WRITE_ONCE(slot->addr, NULL);
+ return NULL;
+ }
+ return addr;
+}
+
/**
* hazptr_acquire - Load pointer at address and protect with hazard pointer.
*
@@ -245,24 +267,9 @@ void *hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
ctx->acquire_cpu = smp_processor_id();
ctx->acquire_caller = _THIS_IP_;
#endif
- if (unlikely(slot->addr))
+ addr = __hazptr_try_acquire(ctx, addr_p, slot);
+ if (unlikely(!addr))
return __hazptr_acquire(ctx, addr_p);
- WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard)); /* Store B */
-
- /* Memory ordering: Store B before Load A. */
- smp_mb();
-
- /*
- * Load @addr_p after storing wildcard to the hazard pointer slot.
- */
- addr = READ_ONCE(*addr_p); /* Load A */
-
- /*
- * We don't care about ordering of Store C. It will simply
- * replace the wildcard by a more specific address. If addr is
- * NULL, we simply store NULL into the slot.
- */
- WRITE_ONCE(slot->addr, addr); /* Store C */
slot_item->ctx.ctx = ctx;
ctx->slot = slot;
return addr;
diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index 13faa5ba7677..3ca73b56a5c2 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -13,17 +13,9 @@
#include <linux/list.h>
#include <linux/export.h>
-static DEFINE_MUTEX(hazptr_phase_lock); /* Protect the wildcard and list phase flip. */
+#define HAZPTR_WILDCARD ((void *) 1UL)
-/*
- * The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
- * hazptr_synchronize forward progress even with a steady stream of readers.
- * This wildcard value is used by acquire to temporarily tag the per-CPU slots.
- * This also affects the overflow list selection: the current list used by
- * readers is array[(unsigned long) hazptr_wildcard - 1].
- */
-void *hazptr_wildcard = (void *) 1UL;
-EXPORT_SYMBOL_GPL(hazptr_wildcard);
+static DEFINE_MUTEX(hazptr_phase_lock); /* Protect the list phase flip. */
/* The current overflow list phase. */
static unsigned int hazptr_overflow_list_phase;
@@ -41,6 +33,10 @@ struct hazptr_overflow_list {
* successively iterates on both lists. Therefore, only list removals
* can cause the iteration to retry, and the number of removals is
* limited to the number of list elements.
+ *
+ * Due to the overflow list raw spin lock, the hazard pointer readers are
+ * blocking, starvation-free with bounded waiting, assuming bounded critical
+ * sections and no NMI or virtualization-induced holder preemption.
*/
struct hazptr_overflow_list_flip {
struct hazptr_overflow_list array[2];
@@ -51,26 +47,12 @@ static DEFINE_PER_CPU(struct hazptr_overflow_list_flip, percpu_overflow_list_fli
DEFINE_PER_CPU(struct hazptr_percpu_slots, hazptr_percpu_slots);
EXPORT_PER_CPU_SYMBOL_GPL(hazptr_percpu_slots);
-static
-void *flip_wildcard(void *wildcard)
-{
- return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
-}
-
static
unsigned int flip_list_phase(unsigned int phase)
{
return 1 - phase;
}
-static
-bool is_wildcard(void *addr)
-{
- if ((unsigned long) addr == 1UL || (unsigned long) addr == 2UL)
- return true;
- return false;
-}
-
static
struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
{
@@ -96,16 +78,36 @@ struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
*/
void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
{
- struct hazptr_slot *slot = hazptr_get_free_percpu_slot(ctx);
+ struct hazptr_slot *slot;
void *addr;
/*
- * If all the per-CPU slots are already in use, fallback
- * to the backup slot.
+ * In case we are called due to nested use of hazard pointers,
+ * try a slot protection with per-CPU slots.
*/
- if (unlikely(!slot))
- slot = hazptr_chain_backup_slot(ctx);
- WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard)); /* Store B */
+ slot = hazptr_get_free_percpu_slot(ctx);
+ if (likely(slot)) {
+ addr = __hazptr_try_acquire(ctx, addr_p, slot);
+ if (addr) {
+ ctx->slot = slot;
+ return addr;
+ }
+ }
+
+ /*
+ * The backup slot overflow list guarantees forward progress of both
+ * hazard pointer readers and synchronize:
+ *
+ * - Readers set the wildcard, and then proceed to set the more
+ * specific address to replace the wildcard.
+ *
+ * - One synchronize alternates between two overflow list periods,
+ * scanning each one while readers are added to the other period,
+ * thus preventing a steady flow of readers from preventing
+ * synchronize forward progress.
+ */
+ slot = hazptr_chain_backup_slot(ctx);
+ WRITE_ONCE(slot->addr, HAZPTR_WILDCARD); /* Store B */
/* Memory ordering: Store B before Load A. */
smp_mb();
@@ -121,20 +123,21 @@ void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
* NULL, we simply store NULL into the slot.
*/
WRITE_ONCE(slot->addr, addr); /* Store C */
+
ctx->slot = slot;
- if (!addr && hazptr_slot_is_backup(ctx, slot))
+ if (!addr)
hazptr_unchain_backup_slot(ctx);
return addr;
}
EXPORT_SYMBOL_GPL(__hazptr_acquire);
/*
- * Perform piecewise iteration on overflow list waiting until "addr" is
- * not present. Raw spinlock is released and taken between each list
- * item and busy loop iteration. The overflow list generation is checked
- * each time the lock is taken to validate that the list has not changed
- * before resuming iteration or busy wait. If the generation has
- * changed, retry the entire list traversal.
+ * Perform piecewise iteration on overflow list waiting until "addr" and
+ * wildcard are not present. Raw spinlock is released and taken between each
+ * list item and busy loop iteration. The overflow list generation is checked
+ * each time the lock is taken to validate that the list has not changed before
+ * resuming iteration or busy wait. If the generation has changed, retry the
+ * entire list traversal.
*/
static
void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list, void *addr)
@@ -147,13 +150,11 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
retry:
snapshot_gen = overflow_list->gen;
hlist_for_each_entry(backup_slot, &overflow_list->head, overflow_node) {
- /* Busy-wait if node is found. */
+ /* Busy-wait if addr or wildcard are found. */
for (;;) {
void *load_addr = smp_load_acquire(&backup_slot->slot.addr); /* Load B */
- /* We don't expect wildcards in overflow list. */
- WARN_ON_ONCE(is_wildcard(load_addr));
- if (load_addr != addr)
+ if (load_addr != addr && load_addr != HAZPTR_WILDCARD)
break;
raw_spin_unlock_irqrestore(&overflow_list->lock, flags);
cpu_relax();
@@ -174,7 +175,7 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
}
static
-void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
+void hazptr_synchronize_cpu_slots(int cpu, void *addr)
{
struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
unsigned int idx;
@@ -182,13 +183,13 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
struct hazptr_slot_item *item = &percpu_slots->items[idx];
- /* Busy-wait if node is found. */
- smp_cond_load_acquire(&item->slot.addr, VAL != addr && VAL != scan_wildcard); /* Load B */
+ /* Busy-wait if addr is found. */
+ smp_cond_load_acquire(&item->slot.addr, VAL != addr); /* Load B */
}
}
static
-void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
+void hazptr_scan_cpu_slots(void *addr)
{
int cpu;
@@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
for_each_possible_cpu(cpu) {
/*
* Scan CPU slots.
- * Forward progress against recurring wildcards is guaranteed
- * by scanning for one wildcard while new elements use the
- * other wildcard value (1UL vs 2UL).
* Forward progress against recurring single hazard pointer
* values is guaranteed by the fact that a hazard pointer
* is not reclaimed nor reused until the scan for that hazard
* pointer completes, which prevents a steady flow of readers
* to acquire that same hazard pointer value.
*/
- hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
+ hazptr_synchronize_cpu_slots(cpu, addr);
}
}
@@ -237,7 +235,6 @@ void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
void hazptr_synchronize(void *addr)
{
unsigned int scan_list_phase;
- void *scan_wildcard;
/*
* Busy-wait should only be done from preemptible context.
@@ -251,16 +248,14 @@ void hazptr_synchronize(void *addr)
*/
if (!addr)
return;
+
/* Memory ordering: Store A before Load B. */
smp_mb();
guard(mutex)(&hazptr_phase_lock);
/* Scan per-CPU slots. */
- scan_wildcard = flip_wildcard(hazptr_wildcard);
- hazptr_scan_cpu_slots_period(addr, scan_wildcard);
- WRITE_ONCE(hazptr_wildcard, scan_wildcard); /* Flip the current wildcard. */
- hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
+ hazptr_scan_cpu_slots(addr);
/*
* Scan overflow lists *after* scanning per-CPU slots. See
--
2.43.0
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 0/4] Hazard pointer updates
2026-09-27 15:51 [PATCH hazptr 0/4] Hazard pointer updates Mathieu Desnoyers
` (3 preceding siblings ...)
2026-09-27 15:51 ` [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
@ 2026-09-27 16:07 ` Bradley Morgan
2026-09-27 16:27 ` Mathieu Desnoyers
2026-09-27 16:15 ` Boqun Feng
5 siblings, 1 reply; 22+ messages in thread
From: Bradley Morgan @ 2026-09-27 16:07 UTC (permalink / raw)
To: Mathieu Desnoyers, Paul E . McKenney
Cc: linux-kernel, Boqun Feng, Gary Guo, rcu, lkmm
On 27 September 2026 16:51:27 BST, Mathieu Desnoyers
<mathieu.desnoyers@efficios.com> wrote:
>Hi Paul,
>
>This series applies on top of "hazptr: handle NULL address in
>hazptr_detach" you have in your rcu dev tree.
>
>This first patch addresses a race identified by Boqun Feng in the
>two-phase wildcard scheme.
>
>Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
>Those were discussed at length in a prior version of hazard pointer
>patches.
>
>Patch 4 introduces a "try acquire" helper to allow the fast path
>to not rely on wildcards, while keeping the wildcard forward
>progress guarantees in the acquire slow path, used on fast path
>failure.
Hi, here is a hazptr perf test on powerpc
REAL kill_fasync(), ns per call, best of 3, 100k calls:
(stock = rwlock walk, conv = hazptr walk, same v3 tree ± the conversion)
shape stock conv delta
1 node, 1 walker 59 59 +0.0% (singleton: identical)
16 nodes, 1 walker 539 539 +0.0% (uncontended: identical)
16 nodes, 4 walkers 509 134 -73.7% ← rwlock readers contend
16 nodes, 8 walkers 313 113 -63.9% ← same list, 8 cpus
64 nodes, 1 walker 1979 2039 +3.0% (pure walk: hazptr tax)
64 nodes, 4 walkers 1914 509 -73.4%
64 nodes, 8 walkers 1015 382 -62.4%
Its SLOWER than rcu, but beats rwlock
SIGIO delivery, plain mode, 8 ptys 1 listener each (identical harness):
BASELINE (rwlock) 472/s
CONVERTED v4 (hazptr) 452/s ← singleton fast path: gap 12% → 4%
With a few changes, it was 12% slower than rcu before.
Do you want those changes?
My idea is, we find something that would put use to hazptr, here is what I
tried
diff --git a/fs/fcntl.c b/fs/fcntl.c
index c158f082f1da..bb04076ff6d9 100644
--- a/fs/fcntl.c
+++ b/fs/fcntl.c
@@ -17,6 +17,7 @@
#include <linux/slab.h>
#include <linux/module.h>
#include <linux/pipe_fs_i.h>
+#include <linux/hazptr.h>
#include <linux/security.h>
#include <linux/ptrace.h>
#include <linux/signal.h>
@@ -1009,15 +1010,20 @@ int fasync_remove_entry(struct file *filp, struct fasync_struct **fapp)
if (fa->fa_file != filp)
continue;
- write_lock_irq(&fa->fa_lock);
+ /*
+ * Make the file invisible to the walk before unlinking,
+ * then wait for any in-flight send_sigio() to be done with
+ * the node before freeing it. The walk holds a hazard
+ * pointer to this node, so it cannot already be freed.
+ */
fa->fa_file = NULL;
- write_unlock_irq(&fa->fa_lock);
-
*fp = fa->fa_next;
- kfree_rcu(fa, fa_rcu);
+ spin_unlock(&fasync_lock);
+ spin_unlock(&filp->f_lock);
+ hazptr_synchronize(fa);
+ fasync_free(fa);
filp->f_flags &= ~FASYNC;
- result = 1;
- break;
+ return 1;
}
spin_unlock(&fasync_lock);
spin_unlock(&filp->f_lock);
@@ -1056,13 +1062,10 @@ struct fasync_struct *fasync_insert_entry(int fd, struct file *filp, struct fasy
if (fa->fa_file != filp)
continue;
- write_lock_irq(&fa->fa_lock);
- fa->fa_fd = fd;
- write_unlock_irq(&fa->fa_lock);
+ WRITE_ONCE(fa->fa_fd, fd);
goto out;
}
- rwlock_init(&new->fa_lock);
new->magic = FASYNC_MAGIC;
new->fa_file = filp;
new->fa_fd = fd;
@@ -1121,44 +1124,51 @@ EXPORT_SYMBOL(fasync_helper);
/*
* rcu_read_lock() is held
*/
-static void kill_fasync_rcu(struct fasync_struct *fa, int sig, int band)
+void kill_fasync(struct fasync_struct **fp, int sig, int band)
{
+ struct hazptr_ctx cur, nxt;
+ struct fasync_struct *fa;
+
+ /* First a quick test without locking: usually
+ * the list is empty.
+ */
+ fa = READ_ONCE(*fp);
+ if (!fa)
+ return;
+
+ /*
+ * Hand-over-hand with two ping-ponged contexts: the next node
+ * must be acquired before the current one is released, but a
+ * hazptr_ctx may only front one live slot at a time.
+ */
+ cur = (struct hazptr_ctx){ };
+ nxt = (struct hazptr_ctx){ };
+ fa = hazptr_acquire(&cur, (void * const *)fp);
while (fa) {
- struct fown_struct *fown;
- unsigned long flags;
+ struct fasync_struct *next;
if (fa->magic != FASYNC_MAGIC) {
printk(KERN_ERR "kill_fasync: bad magic number in "
"fasync_struct!\n");
- return;
+ break;
}
- read_lock_irqsave(&fa->fa_lock, flags);
+
if (fa->fa_file) {
- fown = file_f_owner(fa->fa_file);
- if (!fown)
- goto next;
- /* Don't send SIGURG to processes which have not set a
- queued signum: SIGURG has its own default signalling
- mechanism. */
- if (!(sig == SIGURG && fown->signum == 0))
+ struct fown_struct *fown = file_f_owner(fa->fa_file);
+
+ if (fown &&
+ /* Don't send SIGURG to processes which have not set a
+ queued signum: SIGURG has its own default signalling
+ mechanism. */
+ !(sig == SIGURG && fown->signum == 0))
send_sigio(fown, fa->fa_fd, band);
}
-next:
- read_unlock_irqrestore(&fa->fa_lock, flags);
- fa = rcu_dereference(fa->fa_next);
- }
-}
-
-void kill_fasync(struct fasync_struct **fp, int sig, int band)
-{
- /* First a quick test without locking: usually
- * the list is empty.
- */
- if (*fp) {
- rcu_read_lock();
- kill_fasync_rcu(rcu_dereference(*fp), sig, band);
- rcu_read_unlock();
+ next = hazptr_acquire(&nxt, (void * const *)&fa->fa_next);
+ hazptr_release(&cur, fa);
+ swap(cur, nxt);
+ fa = next;
}
+ hazptr_release(&cur, fa);
}
EXPORT_SYMBOL(kill_fasync);
diff --git a/include/linux/fs.h b/include/linux/fs.h
index f9d1e05e8ae6..852f1b8c00af 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -1365,7 +1365,6 @@ static inline struct dentry *file_dentry(const struct file *file)
}
struct fasync_struct {
- rwlock_t fa_lock;
int magic;
int fa_fd;
struct fasync_struct *fa_next; /* singly linked list */
Anything I did wrong? No?
>
>Thanks,
>
>Mathieu
>
>Mathieu Desnoyers (4):
> hazptr: Fix two-phase hazptr_synchronize race with detach
> compiler.h: Introduce ptr_eq() to preserve address dependency
> Documentation: RCU: Refer to ptr_eq()
> hazptr: Introduce "try acquire" fast path, fallback to overflow list
>
>Cc: Paul E. McKenney <paulmck@kernel.org>
>Cc: Boqun Feng <boqun@kernel.org>
>Cc: Bradley Morgan <brads@mainlining.org>
>Cc: Gary Guo <gary@garyguo.net>
>Cc: <rcu@vger.kernel.org>
>Cc: <lkmm@lists.linux.dev>
>
> Documentation/RCU/rcu_dereference.rst | 38 +++++++-
> include/linux/compiler.h | 63 ++++++++++++
> include/linux/hazptr.h | 47 +++++----
> kernel/hazptr.c | 135 +++++++++++++++-----------
> 4 files changed, 203 insertions(+), 80 deletions(-)
>
>
--- Thanks!
"I'm not a very positive person" - Linus torvalds
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 0/4] Hazard pointer updates
2026-09-27 15:51 [PATCH hazptr 0/4] Hazard pointer updates Mathieu Desnoyers
` (4 preceding siblings ...)
2026-09-27 16:07 ` [PATCH hazptr 0/4] Hazard pointer updates Bradley Morgan
@ 2026-09-27 16:15 ` Boqun Feng
2026-09-27 16:20 ` Mathieu Desnoyers
5 siblings, 1 reply; 22+ messages in thread
From: Boqun Feng @ 2026-09-27 16:15 UTC (permalink / raw)
To: Mathieu Desnoyers
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu,
lkmm, Lian Wang, Kunwu Chan
On Sun, Sep 27, 2026 at 11:51:27AM -0400, Mathieu Desnoyers wrote:
> Hi Paul,
>
> This series applies on top of "hazptr: handle NULL address in
> hazptr_detach" you have in your rcu dev tree.
>
> This first patch addresses a race identified by Boqun Feng in the
> two-phase wildcard scheme.
>
Thank you! That looks good from a quick look. I will give a deep look
later on (I was traveling).
For the following versions/patches, could you also Cc Kuwu Chan and
Liang Wang (Cced) ? They are helping on the scan thread and lockdep
integration, so it'll be good to keep them in the loop. Thank you!
> Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
> Those were discussed at length in a prior version of hazard pointer
> patches.
>
> Patch 4 introduces a "try acquire" helper to allow the fast path
> to not rely on wildcards, while keeping the wildcard forward
> progress guarantees in the acquire slow path, used on fast path
> failure.
>
Nice! I want to point out that later on if the reader can handle the
race with the updater (i.e. the readers don't need the progress
guarantees from hazptr, then we can expose a hazptr_try_acquire() for
exactly that). Although I do have some design trade-off questions for
patch 4.
Regards,
Boqun
> Thanks,
>
> Mathieu
>
> Mathieu Desnoyers (4):
> hazptr: Fix two-phase hazptr_synchronize race with detach
> compiler.h: Introduce ptr_eq() to preserve address dependency
> Documentation: RCU: Refer to ptr_eq()
> hazptr: Introduce "try acquire" fast path, fallback to overflow list
>
> Cc: Paul E. McKenney <paulmck@kernel.org>
> Cc: Boqun Feng <boqun@kernel.org>
> Cc: Bradley Morgan <brads@mainlining.org>
> Cc: Gary Guo <gary@garyguo.net>
> Cc: <rcu@vger.kernel.org>
> Cc: <lkmm@lists.linux.dev>
>
> Documentation/RCU/rcu_dereference.rst | 38 +++++++-
> include/linux/compiler.h | 63 ++++++++++++
> include/linux/hazptr.h | 47 +++++----
> kernel/hazptr.c | 135 +++++++++++++++-----------
> 4 files changed, 203 insertions(+), 80 deletions(-)
>
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 0/4] Hazard pointer updates
2026-09-27 16:15 ` Boqun Feng
@ 2026-09-27 16:20 ` Mathieu Desnoyers
2026-09-27 16:22 ` Bradley Morgan
0 siblings, 1 reply; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 16:20 UTC (permalink / raw)
To: Boqun Feng
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu,
lkmm, Lian Wang, Kunwu Chan
On 2026-09-27 12:15, Boqun Feng wrote:
> On Sun, Sep 27, 2026 at 11:51:27AM -0400, Mathieu Desnoyers wrote:
>> Hi Paul,
>>
>> This series applies on top of "hazptr: handle NULL address in
>> hazptr_detach" you have in your rcu dev tree.
>>
>> This first patch addresses a race identified by Boqun Feng in the
>> two-phase wildcard scheme.
>>
>
> Thank you! That looks good from a quick look. I will give a deep look
> later on (I was traveling).
>
> For the following versions/patches, could you also Cc Kuwu Chan and
> Liang Wang (Cced) ? They are helping on the scan thread and lockdep
> integration, so it'll be good to keep them in the loop. Thank you!
Noted, I've added them to my patch commit messages, so they should
be there for the next round.
>
>> Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
>> Those were discussed at length in a prior version of hazard pointer
>> patches.
>>
>> Patch 4 introduces a "try acquire" helper to allow the fast path
>> to not rely on wildcards, while keeping the wildcard forward
>> progress guarantees in the acquire slow path, used on fast path
>> failure.
>>
>
> Nice! I want to point out that later on if the reader can handle the
> race with the updater (i.e. the readers don't need the progress
> guarantees from hazptr, then we can expose a hazptr_try_acquire() for
> exactly that). Although I do have some design trade-off questions for
> patch 4.
Sure, if you have a use-case in mind I'm interested to hear about it!
Thanks,
Mathieu
>
> Regards,
> Boqun
>
>> Thanks,
>>
>> Mathieu
>>
>> Mathieu Desnoyers (4):
>> hazptr: Fix two-phase hazptr_synchronize race with detach
>> compiler.h: Introduce ptr_eq() to preserve address dependency
>> Documentation: RCU: Refer to ptr_eq()
>> hazptr: Introduce "try acquire" fast path, fallback to overflow list
>>
>> Cc: Paul E. McKenney <paulmck@kernel.org>
>> Cc: Boqun Feng <boqun@kernel.org>
>> Cc: Bradley Morgan <brads@mainlining.org>
>> Cc: Gary Guo <gary@garyguo.net>
>> Cc: <rcu@vger.kernel.org>
>> Cc: <lkmm@lists.linux.dev>
>>
>> Documentation/RCU/rcu_dereference.rst | 38 +++++++-
>> include/linux/compiler.h | 63 ++++++++++++
>> include/linux/hazptr.h | 47 +++++----
>> kernel/hazptr.c | 135 +++++++++++++++-----------
>> 4 files changed, 203 insertions(+), 80 deletions(-)
>>
>> --
>> 2.43.0
>>
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 0/4] Hazard pointer updates
2026-09-27 16:20 ` Mathieu Desnoyers
@ 2026-09-27 16:22 ` Bradley Morgan
0 siblings, 0 replies; 22+ messages in thread
From: Bradley Morgan @ 2026-09-27 16:22 UTC (permalink / raw)
To: Mathieu Desnoyers, Boqun Feng
Cc: Paul E . McKenney, linux-kernel, Gary Guo, rcu, lkmm, Lian Wang,
Kunwu Chan
On 27 September 2026 17:20:36 BST, Mathieu Desnoyers
<mathieu.desnoyers@efficios.com> wrote:
>On 2026-09-27 12:15, Boqun Feng wrote:
>> On Sun, Sep 27, 2026 at 11:51:27AM -0400, Mathieu Desnoyers wrote:
>>> Hi Paul,
>>>
>>> This series applies on top of "hazptr: handle NULL address in
>>> hazptr_detach" you have in your rcu dev tree.
>>>
>>> This first patch addresses a race identified by Boqun Feng in the
>>> two-phase wildcard scheme.
>>>
>>
>> Thank you! That looks good from a quick look. I will give a deep look
>> later on (I was traveling).
>>
>> For the following versions/patches, could you also Cc Kuwu Chan and
>> Liang Wang (Cced) ? They are helping on the scan thread and lockdep
>> integration, so it'll be good to keep them in the loop. Thank you!
>
>Noted, I've added them to my patch commit messages, so they should
>be there for the next round.
>
Maybe my tests could help you? Sorry to impede but the things I sent may
need a but if discussi9n around it
>>
>>> Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
>>> Those were discussed at length in a prior version of hazard pointer
>>> patches.
>>>
>>> Patch 4 introduces a "try acquire" helper to allow the fast path
>>> to not rely on wildcards, while keeping the wildcard forward
>>> progress guarantees in the acquire slow path, used on fast path
>>> failure.
>>>
>>
>> Nice! I want to point out that later on if the reader can handle the
>> race with the updater (i.e. the readers don't need the progress
>> guarantees from hazptr, then we can expose a hazptr_try_acquire() for
>> exactly that). Although I do have some design trade-off questions for
>> patch 4.
>
>Sure, if you have a use-case in mind I'm interested to hear about it!
>
>Thanks,
>
>Mathieu
>
>>
>> Regards,
>> Boqun
>>
>>> Thanks,
>>>
>>> Mathieu
>>>
>>> Mathieu Desnoyers (4):
>>> hazptr: Fix two-phase hazptr_synchronize race with detach
>>> compiler.h: Introduce ptr_eq() to preserve address dependency
>>> Documentation: RCU: Refer to ptr_eq()
>>> hazptr: Introduce "try acquire" fast path, fallback to overflow list
>>>
>>> Cc: Paul E. McKenney <paulmck@kernel.org>
>>> Cc: Boqun Feng <boqun@kernel.org>
>>> Cc: Bradley Morgan <brads@mainlining.org>
>>> Cc: Gary Guo <gary@garyguo.net>
>>> Cc: <rcu@vger.kernel.org>
>>> Cc: <lkmm@lists.linux.dev>
>>>
>>> Documentation/RCU/rcu_dereference.rst | 38 +++++++-
>>> include/linux/compiler.h | 63 ++++++++++++
>>> include/linux/hazptr.h | 47 +++++----
>>> kernel/hazptr.c | 135
>+++++++++++++++-----------
>>> 4 files changed, 203 insertions(+), 80 deletions(-)
>>>
>>> --
>>> 2.43.0
>>>
>
>
>
--- Thanks!
https://lore.kernel.org/all/EE579805-42F2-4C58-B752-F28779EEB717@grrlz.net/
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 0/4] Hazard pointer updates
2026-09-27 16:07 ` [PATCH hazptr 0/4] Hazard pointer updates Bradley Morgan
@ 2026-09-27 16:27 ` Mathieu Desnoyers
2026-09-27 16:33 ` Bradley Morgan
0 siblings, 1 reply; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 16:27 UTC (permalink / raw)
To: Bradley Morgan, Paul E . McKenney
Cc: linux-kernel, Boqun Feng, Gary Guo, rcu, lkmm
On 2026-09-27 12:07, Bradley Morgan wrote:
> On 27 September 2026 16:51:27 BST, Mathieu Desnoyers
> <mathieu.desnoyers@efficios.com> wrote:
>> Hi Paul,
>>
>> This series applies on top of "hazptr: handle NULL address in
>> hazptr_detach" you have in your rcu dev tree.
>>
>> This first patch addresses a race identified by Boqun Feng in the
>> two-phase wildcard scheme.
>>
>> Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
>> Those were discussed at length in a prior version of hazard pointer
>> patches.
>>
>> Patch 4 introduces a "try acquire" helper to allow the fast path
>> to not rely on wildcards, while keeping the wildcard forward
>> progress guarantees in the acquire slow path, used on fast path
>> failure.
>
> Hi, here is a hazptr perf test on powerpc
>
> REAL kill_fasync(), ns per call, best of 3, 100k calls:
> (stock = rwlock walk, conv = hazptr walk, same v3 tree ± the conversion)
>
> shape stock conv delta
> 1 node, 1 walker 59 59 +0.0% (singleton: identical)
> 16 nodes, 1 walker 539 539 +0.0% (uncontended: identical)
> 16 nodes, 4 walkers 509 134 -73.7% ← rwlock readers contend
> 16 nodes, 8 walkers 313 113 -63.9% ← same list, 8 cpus
> 64 nodes, 1 walker 1979 2039 +3.0% (pure walk: hazptr tax)
> 64 nodes, 4 walkers 1914 509 -73.4%
> 64 nodes, 8 walkers 1015 382 -62.4%
>
> Its SLOWER than rcu, but beats rwlock
Two feedback points:
1) The comparison I think Boqun cares mostly about is with expedited
RCU grace periods, this is where we suspect there is a significant
benefit to using hazptr rather than RCU to eliminate those IPIs
on synchronize.
It's good to know that it performs better than rwlock (albeit it's
not surprising).
2) I'm concerned about what looks like a use of hazptr to protect linked
lists elements in your benchmark (did I miss anything ?).
RCU read-side critical sections protect all elements of a linked list
naturally, but hazptr requires more care. See this comment above
hazptr_acquire:
* This protection is unconditional, and has limitations similar to
* that of unconditional reference-counter acquisition. In particular,
* although holding a hazard pointer prevents a hazard-pointer-protected
* object from being freed, it does not prevent that object from being
* removed from a linked data structure, and does not prevent other
* hazard-pointer-protected objects referenced by this object from being
* both removed and freed. At which point, invoking hazptr_acquire()
* on these dangling pointers would be a bug. On the other hand, use of
* hazptr_acquire() is safe for immortal pointers to objects that do not
* themselves contain pointers to hazard-pointer-protected objects.
* Other (more complex) use cases are also possible.
Does the pointer you protect qualify as an "immortal" pointer, or it's
a linked list "next" pointer ?
Thanks,
Mathieu
>
> SIGIO delivery, plain mode, 8 ptys 1 listener each (identical harness):
>
> BASELINE (rwlock) 472/s
> CONVERTED v4 (hazptr) 452/s ← singleton fast path: gap 12% → 4%
>
> With a few changes, it was 12% slower than rcu before.
>
> Do you want those changes?
>
> My idea is, we find something that would put use to hazptr, here is what I
> tried
>
>
> diff --git a/fs/fcntl.c b/fs/fcntl.c
> index c158f082f1da..bb04076ff6d9 100644
> --- a/fs/fcntl.c
> +++ b/fs/fcntl.c
> @@ -17,6 +17,7 @@
> #include <linux/slab.h>
> #include <linux/module.h>
> #include <linux/pipe_fs_i.h>
> +#include <linux/hazptr.h>
> #include <linux/security.h>
> #include <linux/ptrace.h>
> #include <linux/signal.h>
> @@ -1009,15 +1010,20 @@ int fasync_remove_entry(struct file *filp, struct fasync_struct **fapp)
> if (fa->fa_file != filp)
> continue;
>
> - write_lock_irq(&fa->fa_lock);
> + /*
> + * Make the file invisible to the walk before unlinking,
> + * then wait for any in-flight send_sigio() to be done with
> + * the node before freeing it. The walk holds a hazard
> + * pointer to this node, so it cannot already be freed.
> + */
> fa->fa_file = NULL;
> - write_unlock_irq(&fa->fa_lock);
> -
> *fp = fa->fa_next;
> - kfree_rcu(fa, fa_rcu);
> + spin_unlock(&fasync_lock);
> + spin_unlock(&filp->f_lock);
> + hazptr_synchronize(fa);
> + fasync_free(fa);
> filp->f_flags &= ~FASYNC;
> - result = 1;
> - break;
> + return 1;
> }
> spin_unlock(&fasync_lock);
> spin_unlock(&filp->f_lock);
> @@ -1056,13 +1062,10 @@ struct fasync_struct *fasync_insert_entry(int fd, struct file *filp, struct fasy
> if (fa->fa_file != filp)
> continue;
>
> - write_lock_irq(&fa->fa_lock);
> - fa->fa_fd = fd;
> - write_unlock_irq(&fa->fa_lock);
> + WRITE_ONCE(fa->fa_fd, fd);
> goto out;
> }
>
> - rwlock_init(&new->fa_lock);
> new->magic = FASYNC_MAGIC;
> new->fa_file = filp;
> new->fa_fd = fd;
> @@ -1121,44 +1124,51 @@ EXPORT_SYMBOL(fasync_helper);
> /*
> * rcu_read_lock() is held
> */
> -static void kill_fasync_rcu(struct fasync_struct *fa, int sig, int band)
> +void kill_fasync(struct fasync_struct **fp, int sig, int band)
> {
> + struct hazptr_ctx cur, nxt;
> + struct fasync_struct *fa;
> +
> + /* First a quick test without locking: usually
> + * the list is empty.
> + */
> + fa = READ_ONCE(*fp);
> + if (!fa)
> + return;
> +
> + /*
> + * Hand-over-hand with two ping-ponged contexts: the next node
> + * must be acquired before the current one is released, but a
> + * hazptr_ctx may only front one live slot at a time.
> + */
> + cur = (struct hazptr_ctx){ };
> + nxt = (struct hazptr_ctx){ };
> + fa = hazptr_acquire(&cur, (void * const *)fp);
> while (fa) {
> - struct fown_struct *fown;
> - unsigned long flags;
> + struct fasync_struct *next;
>
> if (fa->magic != FASYNC_MAGIC) {
> printk(KERN_ERR "kill_fasync: bad magic number in "
> "fasync_struct!\n");
> - return;
> + break;
> }
> - read_lock_irqsave(&fa->fa_lock, flags);
> +
> if (fa->fa_file) {
> - fown = file_f_owner(fa->fa_file);
> - if (!fown)
> - goto next;
> - /* Don't send SIGURG to processes which have not set a
> - queued signum: SIGURG has its own default signalling
> - mechanism. */
> - if (!(sig == SIGURG && fown->signum == 0))
> + struct fown_struct *fown = file_f_owner(fa->fa_file);
> +
> + if (fown &&
> + /* Don't send SIGURG to processes which have not set a
> + queued signum: SIGURG has its own default signalling
> + mechanism. */
> + !(sig == SIGURG && fown->signum == 0))
> send_sigio(fown, fa->fa_fd, band);
> }
> -next:
> - read_unlock_irqrestore(&fa->fa_lock, flags);
> - fa = rcu_dereference(fa->fa_next);
> - }
> -}
> -
> -void kill_fasync(struct fasync_struct **fp, int sig, int band)
> -{
> - /* First a quick test without locking: usually
> - * the list is empty.
> - */
> - if (*fp) {
> - rcu_read_lock();
> - kill_fasync_rcu(rcu_dereference(*fp), sig, band);
> - rcu_read_unlock();
> + next = hazptr_acquire(&nxt, (void * const *)&fa->fa_next);
> + hazptr_release(&cur, fa);
> + swap(cur, nxt);
> + fa = next;
> }
> + hazptr_release(&cur, fa);
> }
> EXPORT_SYMBOL(kill_fasync);
>
> diff --git a/include/linux/fs.h b/include/linux/fs.h
> index f9d1e05e8ae6..852f1b8c00af 100644
> --- a/include/linux/fs.h
> +++ b/include/linux/fs.h
> @@ -1365,7 +1365,6 @@ static inline struct dentry *file_dentry(const struct file *file)
> }
>
> struct fasync_struct {
> - rwlock_t fa_lock;
> int magic;
> int fa_fd;
> struct fasync_struct *fa_next; /* singly linked list */
>
>
> Anything I did wrong? No?
>
>
>>
>> Thanks,
>>
>> Mathieu
>>
>> Mathieu Desnoyers (4):
>> hazptr: Fix two-phase hazptr_synchronize race with detach
>> compiler.h: Introduce ptr_eq() to preserve address dependency
>> Documentation: RCU: Refer to ptr_eq()
>> hazptr: Introduce "try acquire" fast path, fallback to overflow list
>>
>> Cc: Paul E. McKenney <paulmck@kernel.org>
>> Cc: Boqun Feng <boqun@kernel.org>
>> Cc: Bradley Morgan <brads@mainlining.org>
>> Cc: Gary Guo <gary@garyguo.net>
>> Cc: <rcu@vger.kernel.org>
>> Cc: <lkmm@lists.linux.dev>
>>
>> Documentation/RCU/rcu_dereference.rst | 38 +++++++-
>> include/linux/compiler.h | 63 ++++++++++++
>> include/linux/hazptr.h | 47 +++++----
>> kernel/hazptr.c | 135 +++++++++++++++-----------
>> 4 files changed, 203 insertions(+), 80 deletions(-)
>>
>>
>
> --- Thanks!
> "I'm not a very positive person" - Linus torvalds
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 0/4] Hazard pointer updates
2026-09-27 16:27 ` Mathieu Desnoyers
@ 2026-09-27 16:33 ` Bradley Morgan
2026-09-27 16:45 ` Mathieu Desnoyers
0 siblings, 1 reply; 22+ messages in thread
From: Bradley Morgan @ 2026-09-27 16:33 UTC (permalink / raw)
To: Mathieu Desnoyers, Paul E . McKenney
Cc: linux-kernel, Boqun Feng, Gary Guo, rcu, lkmm
On 27 September 2026 17:27:33 BST, Mathieu Desnoyers
<mathieu.desnoyers@efficios.com> wrote:
>On 2026-09-27 12:07, Bradley Morgan wrote:
>> On 27 September 2026 16:51:27 BST, Mathieu Desnoyers
>> <mathieu.desnoyers@efficios.com> wrote:
>>> Hi Paul,
>>>
>>> This series applies on top of "hazptr: handle NULL address in
>>> hazptr_detach" you have in your rcu dev tree.
>>>
>>> This first patch addresses a race identified by Boqun Feng in the
>>> two-phase wildcard scheme.
>>>
>>> Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
>>> Those were discussed at length in a prior version of hazard pointer
>>> patches.
>>>
>>> Patch 4 introduces a "try acquire" helper to allow the fast path
>>> to not rely on wildcards, while keeping the wildcard forward
>>> progress guarantees in the acquire slow path, used on fast path
>>> failure.
>>
>> Hi, here is a hazptr perf test on powerpc
>>
>> REAL kill_fasync(), ns per call, best of 3, 100k calls:
>> (stock = rwlock walk, conv = hazptr walk, same v3 tree ± the conversion)
>>
>> shape stock conv delta
>> 1 node, 1 walker 59 59 +0.0% (singleton: identical)
>> 16 nodes, 1 walker 539 539 +0.0% (uncontended: identical)
>> 16 nodes, 4 walkers 509 134 -73.7% ← rwlock readers contend
>> 16 nodes, 8 walkers 313 113 -63.9% ← same list, 8 cpus
>> 64 nodes, 1 walker 1979 2039 +3.0% (pure walk: hazptr tax)
>> 64 nodes, 4 walkers 1914 509 -73.4%
>> 64 nodes, 8 walkers 1015 382 -62.4%
>>
>> Its SLOWER than rcu, but beats rwlock
>
>Two feedback points:
>
>1) The comparison I think Boqun cares mostly about is with expedited
> RCU grace periods, this is where we suspect there is a significant
> benefit to using hazptr rather than RCU to eliminate those IPIs
> on synchronize.
>
> It's good to know that it performs better than rwlock (albeit it's
> not surprising).
>
>2) I'm concerned about what looks like a use of hazptr to protect linked
> lists elements in your benchmark (did I miss anything ?).
>
> RCU read-side critical sections protect all elements of a linked list
> naturally, but hazptr requires more care. See this comment above
> hazptr_acquire:
>
> * This protection is unconditional, and has limitations similar to
> * that of unconditional reference-counter acquisition. In particular,
> * although holding a hazard pointer prevents a hazard-pointer-protected
> * object from being freed, it does not prevent that object from being
> * removed from a linked data structure, and does not prevent other
> * hazard-pointer-protected objects referenced by this object from being
> * both removed and freed. At which point, invoking hazptr_acquire()
> * on these dangling pointers would be a bug. On the other hand, use of
> * hazptr_acquire() is safe for immortal pointers to objects that do not
> * themselves contain pointers to hazard-pointer-protected objects.
> * Other (more complex) use cases are also possible.
>
>Does the pointer you protect qualify as an "immortal" pointer, or it's
>a linked list "next" pointer ?
>
Hmm. Do you have a idea on what you could metaphorically convert, with a
core subsystem?
I'll give anything you want me to do a try
>Thanks,
>
>Mathieu
>
>>
>> SIGIO delivery, plain mode, 8 ptys 1 listener each (identical harness):
>>
>> BASELINE (rwlock) 472/s
>> CONVERTED v4 (hazptr) 452/s ← singleton fast path: gap 12% → 4%
>>
>> With a few changes, it was 12% slower than rcu before.
>>
>> Do you want those changes?
>>
>> My idea is, we find something that would put use to hazptr, here is what
>I
>> tried
>>
>>
>> diff --git a/fs/fcntl.c b/fs/fcntl.c
>> index c158f082f1da..bb04076ff6d9 100644
>> --- a/fs/fcntl.c
>> +++ b/fs/fcntl.c
>> @@ -17,6 +17,7 @@
>> #include <linux/slab.h>
>> #include <linux/module.h>
>> #include <linux/pipe_fs_i.h>
>> +#include <linux/hazptr.h>
>> #include <linux/security.h>
>> #include <linux/ptrace.h>
>> #include <linux/signal.h>
>> @@ -1009,15 +1010,20 @@ int fasync_remove_entry(struct file *filp,
>struct fasync_struct **fapp)
>> if (fa->fa_file != filp)
>> continue;
>> - write_lock_irq(&fa->fa_lock);
>> + /*
>> + * Make the file invisible to the walk before unlinking,
>> + * then wait for any in-flight send_sigio() to be done with
>> + * the node before freeing it. The walk holds a hazard
>> + * pointer to this node, so it cannot already be freed.
>> + */
>> fa->fa_file = NULL;
>> - write_unlock_irq(&fa->fa_lock);
>> -
>> *fp = fa->fa_next;
>> - kfree_rcu(fa, fa_rcu);
>> + spin_unlock(&fasync_lock);
>> + spin_unlock(&filp->f_lock);
>> + hazptr_synchronize(fa);
>> + fasync_free(fa);
>> filp->f_flags &= ~FASYNC;
>> - result = 1;
>> - break;
>> + return 1;
>> }
>> spin_unlock(&fasync_lock);
>> spin_unlock(&filp->f_lock);
>> @@ -1056,13 +1062,10 @@ struct fasync_struct *fasync_insert_entry(int
>fd, struct file *filp, struct fasy
>> if (fa->fa_file != filp)
>> continue;
>> - write_lock_irq(&fa->fa_lock);
>> - fa->fa_fd = fd;
>> - write_unlock_irq(&fa->fa_lock);
>> + WRITE_ONCE(fa->fa_fd, fd);
>> goto out;
>> }
>> - rwlock_init(&new->fa_lock);
>> new->magic = FASYNC_MAGIC;
>> new->fa_file = filp;
>> new->fa_fd = fd;
>> @@ -1121,44 +1124,51 @@ EXPORT_SYMBOL(fasync_helper);
>> /*
>> * rcu_read_lock() is held
>> */
>> -static void kill_fasync_rcu(struct fasync_struct *fa, int sig, int
>band)
>> +void kill_fasync(struct fasync_struct **fp, int sig, int band)
>> {
>> + struct hazptr_ctx cur, nxt;
>> + struct fasync_struct *fa;
>> +
>> + /* First a quick test without locking: usually
>> + * the list is empty.
>> + */
>> + fa = READ_ONCE(*fp);
>> + if (!fa)
>> + return;
>> +
>> + /*
>> + * Hand-over-hand with two ping-ponged contexts: the next node
>> + * must be acquired before the current one is released, but a
>> + * hazptr_ctx may only front one live slot at a time.
>> + */
>> + cur = (struct hazptr_ctx){ };
>> + nxt = (struct hazptr_ctx){ };
>> + fa = hazptr_acquire(&cur, (void * const *)fp);
>> while (fa) {
>> - struct fown_struct *fown;
>> - unsigned long flags;
>> + struct fasync_struct *next;
>> if (fa->magic != FASYNC_MAGIC) {
>> printk(KERN_ERR "kill_fasync: bad magic number in "
>> "fasync_struct!\n");
>> - return;
>> + break;
>> }
>> - read_lock_irqsave(&fa->fa_lock, flags);
>> +
>> if (fa->fa_file) {
>> - fown = file_f_owner(fa->fa_file);
>> - if (!fown)
>> - goto next;
>> - /* Don't send SIGURG to processes which have not set a
>> - queued signum: SIGURG has its own default signalling
>> - mechanism. */
>> - if (!(sig == SIGURG && fown->signum == 0))
>> + struct fown_struct *fown = file_f_owner(fa->fa_file);
>> +
>> + if (fown &&
>> + /* Don't send SIGURG to processes which have not set a
>> + queued signum: SIGURG has its own default signalling
>> + mechanism. */
>> + !(sig == SIGURG && fown->signum == 0))
>> send_sigio(fown, fa->fa_fd, band);
>> }
>> -next:
>> - read_unlock_irqrestore(&fa->fa_lock, flags);
>> - fa = rcu_dereference(fa->fa_next);
>> - }
>> -}
>> -
>> -void kill_fasync(struct fasync_struct **fp, int sig, int band)
>> -{
>> - /* First a quick test without locking: usually
>> - * the list is empty.
>> - */
>> - if (*fp) {
>> - rcu_read_lock();
>> - kill_fasync_rcu(rcu_dereference(*fp), sig, band);
>> - rcu_read_unlock();
>> + next = hazptr_acquire(&nxt, (void * const *)&fa->fa_next);
>> + hazptr_release(&cur, fa);
>> + swap(cur, nxt);
>> + fa = next;
>> }
>> + hazptr_release(&cur, fa);
>> }
>> EXPORT_SYMBOL(kill_fasync);
>> diff --git a/include/linux/fs.h b/include/linux/fs.h
>> index f9d1e05e8ae6..852f1b8c00af 100644
>> --- a/include/linux/fs.h
>> +++ b/include/linux/fs.h
>> @@ -1365,7 +1365,6 @@ static inline struct dentry *file_dentry(const
>struct file *file)
>> }
>> struct fasync_struct {
>> - rwlock_t fa_lock;
>> int magic;
>> int fa_fd;
>> struct fasync_struct *fa_next; /* singly linked list */
>>
>>
>> Anything I did wrong? No?
>>
>>
>>>
>>> Thanks,
>>>
>>> Mathieu
>>>
>>> Mathieu Desnoyers (4):
>>> hazptr: Fix two-phase hazptr_synchronize race with detach
>>> compiler.h: Introduce ptr_eq() to preserve address dependency
>>> Documentation: RCU: Refer to ptr_eq()
>>> hazptr: Introduce "try acquire" fast path, fallback to overflow list
>>>
>>> Cc: Paul E. McKenney <paulmck@kernel.org>
>>> Cc: Boqun Feng <boqun@kernel.org>
>>> Cc: Bradley Morgan <brads@mainlining.org>
>>> Cc: Gary Guo <gary@garyguo.net>
>>> Cc: <rcu@vger.kernel.org>
>>> Cc: <lkmm@lists.linux.dev>
>>>
>>> Documentation/RCU/rcu_dereference.rst | 38 +++++++-
>>> include/linux/compiler.h | 63 ++++++++++++
>>> include/linux/hazptr.h | 47 +++++----
>>> kernel/hazptr.c | 135 +++++++++++++++-----------
>>> 4 files changed, 203 insertions(+), 80 deletions(-)
>>>
>>>
>>
>> --- Thanks!
>> "I'm not a very positive person" - Linus torvalds
>
>
>
--- Thanks!
https://lore.kernel.org/all/EE579805-42F2-4C58-B752-F28779EEB717@grrlz.net/
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 15:51 ` [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
@ 2026-09-27 16:40 ` Boqun Feng
2026-09-27 17:15 ` Mathieu Desnoyers
2026-09-28 9:27 ` Kunwu Chan
0 siblings, 2 replies; 22+ messages in thread
From: Boqun Feng @ 2026-09-27 16:40 UTC (permalink / raw)
To: Mathieu Desnoyers
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu, lkmm
On Sun, Sep 27, 2026 at 11:51:31AM -0400, Mathieu Desnoyers wrote:
> Introduce a "try acquire" hazard pointer fast path, which performs an
> early load of the address to store it into the hazard pointer slot, and
> then re-loads that address after a barrier to check whether it has
> changed meanwhile.
>
> On comparison failure, rather than re-try, guarantee forward progress by
> falling back to the __hazptr_acquire slow path on failure.
>
> The acquire slow path attempts a try-acquire for any available per-CPU
> slot. If that fails, it chains the backup slot into the overflow list,
> therefore guaranteeing forward progress for both hazard pointer
> read-side and synchronize:
>
> - Readers set the wildcard, and then proceed to set the more
> specific address to replace the wildcard.
>
> - One synchronize alternates between two overflow list periods,
> scanning each one while readers are added to the other period,
> thus preventing a steady flow of readers from preventing
> synchronize forward progress.
>
> With this change, the scan on per-CPU slots don't need to expect a
> wildcard anymore, because none can be produced by readers. Wildcards are
> only expected within overflow lists.
>
Ok, I was missing something, but I think it's better to call it out.
Wildcards can only exist in the overflow lists when the context is not
preemptible. In other words, there won't be a preempted readers blocking
the synchronize_hazptr() with a wilcard in the overflow list.
So no more design trade-off question from me :)
Regards,
Boqun
> Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
> Cc: Paul E. McKenney <paulmck@kernel.org>
> Cc: Boqun Feng <boqun@kernel.org>
> Cc: Bradley Morgan <brads@mainlining.org>
> Cc: Gary Guo <gary@garyguo.net>
> Cc: <rcu@vger.kernel.org>
> Cc: <lkmm@lists.linux.dev>
> ---
> include/linux/hazptr.h | 47 +++++++++++--------
> kernel/hazptr.c | 103 ++++++++++++++++++++---------------------
> 2 files changed, 76 insertions(+), 74 deletions(-)
>
> diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
> index d1670121947a..fcf017ca2255 100644
> --- a/include/linux/hazptr.h
> +++ b/include/linux/hazptr.h
> @@ -29,9 +29,6 @@
> /* 4 slots (each sizeof(hazptr_slot_item)) fit in a single 64-byte cache line. */
> #define NR_HAZPTR_PERCPU_SLOTS 4
>
> -/* The current hazard pointer wildcard. */
> -extern void *hazptr_wildcard;
> -
> /*
> * Hazard pointer slot.
> */
> @@ -190,6 +187,31 @@ void hazptr_note_context_switch(void)
> }
> }
>
> +/* Try hazard pointer protection. */
> +static inline
> +void *__hazptr_try_acquire(struct hazptr_ctx *ctx, void * const *addr_p, struct hazptr_slot *slot)
> +{
> + void *early_addr, *addr;
> +
> + if (unlikely(slot->addr))
> + return NULL;
> + early_addr = READ_ONCE(*addr_p); /* Early load. */
> + WRITE_ONCE(slot->addr, early_addr); /* Store B */
> + /* Memory ordering: Store B before Load A. */
> + smp_mb();
> + addr = READ_ONCE(*addr_p); /* Load A */
> + /*
> + * Validate that address did not change between Early load and Load A.
> + * Use ptr_eq() to make sure that result from Load A is returned to the
> + * caller to preserve address dependency.
> + */
> + if (unlikely(!ptr_eq(addr, early_addr))) {
> + WRITE_ONCE(slot->addr, NULL);
> + return NULL;
> + }
> + return addr;
> +}
> +
> /**
> * hazptr_acquire - Load pointer at address and protect with hazard pointer.
> *
> @@ -245,24 +267,9 @@ void *hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> ctx->acquire_cpu = smp_processor_id();
> ctx->acquire_caller = _THIS_IP_;
> #endif
> - if (unlikely(slot->addr))
> + addr = __hazptr_try_acquire(ctx, addr_p, slot);
> + if (unlikely(!addr))
> return __hazptr_acquire(ctx, addr_p);
> - WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard)); /* Store B */
> -
> - /* Memory ordering: Store B before Load A. */
> - smp_mb();
> -
> - /*
> - * Load @addr_p after storing wildcard to the hazard pointer slot.
> - */
> - addr = READ_ONCE(*addr_p); /* Load A */
> -
> - /*
> - * We don't care about ordering of Store C. It will simply
> - * replace the wildcard by a more specific address. If addr is
> - * NULL, we simply store NULL into the slot.
> - */
> - WRITE_ONCE(slot->addr, addr); /* Store C */
> slot_item->ctx.ctx = ctx;
> ctx->slot = slot;
> return addr;
> diff --git a/kernel/hazptr.c b/kernel/hazptr.c
> index 13faa5ba7677..3ca73b56a5c2 100644
> --- a/kernel/hazptr.c
> +++ b/kernel/hazptr.c
> @@ -13,17 +13,9 @@
> #include <linux/list.h>
> #include <linux/export.h>
>
> -static DEFINE_MUTEX(hazptr_phase_lock); /* Protect the wildcard and list phase flip. */
> +#define HAZPTR_WILDCARD ((void *) 1UL)
>
> -/*
> - * The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
> - * hazptr_synchronize forward progress even with a steady stream of readers.
> - * This wildcard value is used by acquire to temporarily tag the per-CPU slots.
> - * This also affects the overflow list selection: the current list used by
> - * readers is array[(unsigned long) hazptr_wildcard - 1].
> - */
> -void *hazptr_wildcard = (void *) 1UL;
> -EXPORT_SYMBOL_GPL(hazptr_wildcard);
> +static DEFINE_MUTEX(hazptr_phase_lock); /* Protect the list phase flip. */
>
> /* The current overflow list phase. */
> static unsigned int hazptr_overflow_list_phase;
> @@ -41,6 +33,10 @@ struct hazptr_overflow_list {
> * successively iterates on both lists. Therefore, only list removals
> * can cause the iteration to retry, and the number of removals is
> * limited to the number of list elements.
> + *
> + * Due to the overflow list raw spin lock, the hazard pointer readers are
> + * blocking, starvation-free with bounded waiting, assuming bounded critical
> + * sections and no NMI or virtualization-induced holder preemption.
> */
> struct hazptr_overflow_list_flip {
> struct hazptr_overflow_list array[2];
> @@ -51,26 +47,12 @@ static DEFINE_PER_CPU(struct hazptr_overflow_list_flip, percpu_overflow_list_fli
> DEFINE_PER_CPU(struct hazptr_percpu_slots, hazptr_percpu_slots);
> EXPORT_PER_CPU_SYMBOL_GPL(hazptr_percpu_slots);
>
> -static
> -void *flip_wildcard(void *wildcard)
> -{
> - return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
> -}
> -
> static
> unsigned int flip_list_phase(unsigned int phase)
> {
> return 1 - phase;
> }
>
> -static
> -bool is_wildcard(void *addr)
> -{
> - if ((unsigned long) addr == 1UL || (unsigned long) addr == 2UL)
> - return true;
> - return false;
> -}
> -
> static
> struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
> {
> @@ -96,16 +78,36 @@ struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
> */
> void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> {
> - struct hazptr_slot *slot = hazptr_get_free_percpu_slot(ctx);
> + struct hazptr_slot *slot;
> void *addr;
>
> /*
> - * If all the per-CPU slots are already in use, fallback
> - * to the backup slot.
> + * In case we are called due to nested use of hazard pointers,
> + * try a slot protection with per-CPU slots.
> */
> - if (unlikely(!slot))
> - slot = hazptr_chain_backup_slot(ctx);
> - WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard)); /* Store B */
> + slot = hazptr_get_free_percpu_slot(ctx);
> + if (likely(slot)) {
> + addr = __hazptr_try_acquire(ctx, addr_p, slot);
> + if (addr) {
> + ctx->slot = slot;
> + return addr;
> + }
> + }
> +
> + /*
> + * The backup slot overflow list guarantees forward progress of both
> + * hazard pointer readers and synchronize:
> + *
> + * - Readers set the wildcard, and then proceed to set the more
> + * specific address to replace the wildcard.
> + *
> + * - One synchronize alternates between two overflow list periods,
> + * scanning each one while readers are added to the other period,
> + * thus preventing a steady flow of readers from preventing
> + * synchronize forward progress.
> + */
> + slot = hazptr_chain_backup_slot(ctx);
> + WRITE_ONCE(slot->addr, HAZPTR_WILDCARD); /* Store B */
>
> /* Memory ordering: Store B before Load A. */
> smp_mb();
> @@ -121,20 +123,21 @@ void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> * NULL, we simply store NULL into the slot.
> */
> WRITE_ONCE(slot->addr, addr); /* Store C */
> +
> ctx->slot = slot;
> - if (!addr && hazptr_slot_is_backup(ctx, slot))
> + if (!addr)
> hazptr_unchain_backup_slot(ctx);
> return addr;
> }
> EXPORT_SYMBOL_GPL(__hazptr_acquire);
>
> /*
> - * Perform piecewise iteration on overflow list waiting until "addr" is
> - * not present. Raw spinlock is released and taken between each list
> - * item and busy loop iteration. The overflow list generation is checked
> - * each time the lock is taken to validate that the list has not changed
> - * before resuming iteration or busy wait. If the generation has
> - * changed, retry the entire list traversal.
> + * Perform piecewise iteration on overflow list waiting until "addr" and
> + * wildcard are not present. Raw spinlock is released and taken between each
> + * list item and busy loop iteration. The overflow list generation is checked
> + * each time the lock is taken to validate that the list has not changed before
> + * resuming iteration or busy wait. If the generation has changed, retry the
> + * entire list traversal.
> */
> static
> void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list, void *addr)
> @@ -147,13 +150,11 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
> retry:
> snapshot_gen = overflow_list->gen;
> hlist_for_each_entry(backup_slot, &overflow_list->head, overflow_node) {
> - /* Busy-wait if node is found. */
> + /* Busy-wait if addr or wildcard are found. */
> for (;;) {
> void *load_addr = smp_load_acquire(&backup_slot->slot.addr); /* Load B */
>
> - /* We don't expect wildcards in overflow list. */
> - WARN_ON_ONCE(is_wildcard(load_addr));
> - if (load_addr != addr)
> + if (load_addr != addr && load_addr != HAZPTR_WILDCARD)
> break;
> raw_spin_unlock_irqrestore(&overflow_list->lock, flags);
> cpu_relax();
> @@ -174,7 +175,7 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
> }
>
> static
> -void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
> +void hazptr_synchronize_cpu_slots(int cpu, void *addr)
> {
> struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
> unsigned int idx;
> @@ -182,13 +183,13 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
> for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
> struct hazptr_slot_item *item = &percpu_slots->items[idx];
>
> - /* Busy-wait if node is found. */
> - smp_cond_load_acquire(&item->slot.addr, VAL != addr && VAL != scan_wildcard); /* Load B */
> + /* Busy-wait if addr is found. */
> + smp_cond_load_acquire(&item->slot.addr, VAL != addr); /* Load B */
> }
> }
>
> static
> -void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
> +void hazptr_scan_cpu_slots(void *addr)
> {
> int cpu;
>
> @@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
> for_each_possible_cpu(cpu) {
> /*
> * Scan CPU slots.
> - * Forward progress against recurring wildcards is guaranteed
> - * by scanning for one wildcard while new elements use the
> - * other wildcard value (1UL vs 2UL).
> * Forward progress against recurring single hazard pointer
> * values is guaranteed by the fact that a hazard pointer
> * is not reclaimed nor reused until the scan for that hazard
> * pointer completes, which prevents a steady flow of readers
> * to acquire that same hazard pointer value.
(Not a comment to this patch, but I think it's worth bringing up)
I want to point out this is not true for the lockdep use case, because
the we need to protect a hash list deletion there, and we use the
address of the hash bucket there. It's proven fine in practice because
the readers are rare (we only call the reader is_dynamic_key() in
register_lock_class(), that is every time you have a new lock class to
register).
Maybe what we want to say here is that "if the users guarantee no steady
flow of the same hazard pointer value, we guarantee forward progress".
Thoughts?
The rest looks good to me.
Regards,
Boqun
> */
> - hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
> + hazptr_synchronize_cpu_slots(cpu, addr);
> }
> }
>
> @@ -237,7 +235,6 @@ void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
> void hazptr_synchronize(void *addr)
> {
> unsigned int scan_list_phase;
> - void *scan_wildcard;
>
> /*
> * Busy-wait should only be done from preemptible context.
> @@ -251,16 +248,14 @@ void hazptr_synchronize(void *addr)
> */
> if (!addr)
> return;
> +
> /* Memory ordering: Store A before Load B. */
> smp_mb();
>
> guard(mutex)(&hazptr_phase_lock);
>
> /* Scan per-CPU slots. */
> - scan_wildcard = flip_wildcard(hazptr_wildcard);
> - hazptr_scan_cpu_slots_period(addr, scan_wildcard);
> - WRITE_ONCE(hazptr_wildcard, scan_wildcard); /* Flip the current wildcard. */
> - hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
> + hazptr_scan_cpu_slots(addr);
>
> /*
> * Scan overflow lists *after* scanning per-CPU slots. See
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 0/4] Hazard pointer updates
2026-09-27 16:33 ` Bradley Morgan
@ 2026-09-27 16:45 ` Mathieu Desnoyers
0 siblings, 0 replies; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 16:45 UTC (permalink / raw)
To: Bradley Morgan, Paul E . McKenney
Cc: linux-kernel, Boqun Feng, Gary Guo, rcu, lkmm
On 2026-09-27 12:33, Bradley Morgan wrote:
> On 27 September 2026 17:27:33 BST, Mathieu Desnoyers
> <mathieu.desnoyers@efficios.com> wrote:
>> On 2026-09-27 12:07, Bradley Morgan wrote:
>>> On 27 September 2026 16:51:27 BST, Mathieu Desnoyers
>>> <mathieu.desnoyers@efficios.com> wrote:
>>>> Hi Paul,
>>>>
>>>> This series applies on top of "hazptr: handle NULL address in
>>>> hazptr_detach" you have in your rcu dev tree.
>>>>
>>>> This first patch addresses a race identified by Boqun Feng in the
>>>> two-phase wildcard scheme.
>>>>
>>>> Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
>>>> Those were discussed at length in a prior version of hazard pointer
>>>> patches.
>>>>
>>>> Patch 4 introduces a "try acquire" helper to allow the fast path
>>>> to not rely on wildcards, while keeping the wildcard forward
>>>> progress guarantees in the acquire slow path, used on fast path
>>>> failure.
>>>
>>> Hi, here is a hazptr perf test on powerpc
>>>
>>> REAL kill_fasync(), ns per call, best of 3, 100k calls:
>>> (stock = rwlock walk, conv = hazptr walk, same v3 tree ± the conversion)
>>>
>>> shape stock conv delta
>>> 1 node, 1 walker 59 59 +0.0% (singleton: identical)
>>> 16 nodes, 1 walker 539 539 +0.0% (uncontended: identical)
>>> 16 nodes, 4 walkers 509 134 -73.7% ← rwlock readers contend
>>> 16 nodes, 8 walkers 313 113 -63.9% ← same list, 8 cpus
>>> 64 nodes, 1 walker 1979 2039 +3.0% (pure walk: hazptr tax)
>>> 64 nodes, 4 walkers 1914 509 -73.4%
>>> 64 nodes, 8 walkers 1015 382 -62.4%
>>>
>>> Its SLOWER than rcu, but beats rwlock
>>
>> Two feedback points:
>>
>> 1) The comparison I think Boqun cares mostly about is with expedited
>> RCU grace periods, this is where we suspect there is a significant
>> benefit to using hazptr rather than RCU to eliminate those IPIs
>> on synchronize.
>>
>> It's good to know that it performs better than rwlock (albeit it's
>> not surprising).
>>
>> 2) I'm concerned about what looks like a use of hazptr to protect linked
>> lists elements in your benchmark (did I miss anything ?).
>>
>> RCU read-side critical sections protect all elements of a linked list
>> naturally, but hazptr requires more care. See this comment above
>> hazptr_acquire:
>>
>> * This protection is unconditional, and has limitations similar to
>> * that of unconditional reference-counter acquisition. In particular,
>> * although holding a hazard pointer prevents a hazard-pointer-protected
>> * object from being freed, it does not prevent that object from being
>> * removed from a linked data structure, and does not prevent other
>> * hazard-pointer-protected objects referenced by this object from being
>> * both removed and freed. At which point, invoking hazptr_acquire()
>> * on these dangling pointers would be a bug. On the other hand, use of
>> * hazptr_acquire() is safe for immortal pointers to objects that do not
>> * themselves contain pointers to hazard-pointer-protected objects.
>> * Other (more complex) use cases are also possible.
>>
>> Does the pointer you protect qualify as an "immortal" pointer, or it's
>> a linked list "next" pointer ?
>>
>
> Hmm. Do you have a idea on what you could metaphorically convert, with a
> core subsystem?
>
> I'll give anything you want me to do a try
At a high level, I suspect it would be good to start by digging into
users of synchronize_rcu_expedited(), to see if a few of those may be
good candidates.
I would also favor scenarios where RCU (or locking) are used to protect
the existence of an object reachable from a global pointer, and use
hazptr to protect that object. Note that initially this precludes
a hazptr-protected list, because the next pointers would sit in
prior objects which are themselves hazptr-protected, which is not
sufficient to guarantee existence against a hand in hand traversal.
However, if you have a pointer to an object "side-car" structure
(e.g. optional extra metadata), and the object containing the pointer
is guaranteed to exist by another mechanism (e.g. RCU, locking), then
that pointer-to-side-car-object field would be a good candidate for
hazptr (AFAIU).
I'm not saying the linked lists could not be done, but it would require
more care.
Thanks,
Mathieu
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 16:40 ` Boqun Feng
@ 2026-09-27 17:15 ` Mathieu Desnoyers
2026-09-27 17:24 ` Boqun Feng
` (2 more replies)
2026-09-28 9:27 ` Kunwu Chan
1 sibling, 3 replies; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 17:15 UTC (permalink / raw)
To: Boqun Feng
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu, lkmm
On 2026-09-27 12:40, Boqun Feng wrote:
> On Sun, Sep 27, 2026 at 11:51:31AM -0400, Mathieu Desnoyers wrote:
>> Introduce a "try acquire" hazard pointer fast path, which performs an
>> early load of the address to store it into the hazard pointer slot, and
>> then re-loads that address after a barrier to check whether it has
>> changed meanwhile.
>>
>> On comparison failure, rather than re-try, guarantee forward progress by
>> falling back to the __hazptr_acquire slow path on failure.
>>
>> The acquire slow path attempts a try-acquire for any available per-CPU
>> slot. If that fails, it chains the backup slot into the overflow list,
>> therefore guaranteeing forward progress for both hazard pointer
>> read-side and synchronize:
>>
>> - Readers set the wildcard, and then proceed to set the more
>> specific address to replace the wildcard.
>>
>> - One synchronize alternates between two overflow list periods,
>> scanning each one while readers are added to the other period,
>> thus preventing a steady flow of readers from preventing
>> synchronize forward progress.
>>
>> With this change, the scan on per-CPU slots don't need to expect a
>> wildcard anymore, because none can be produced by readers. Wildcards are
>> only expected within overflow lists.
>>
>
> Ok, I was missing something, but I think it's better to call it out.
> Wildcards can only exist in the overflow lists when the context is not
> preemptible. In other words, there won't be a preempted readers blocking
> the synchronize_hazptr() with a wilcard in the overflow list.
Exactly ! Wildcard slots only exist during the short time-frame of the
preempt-off read-side code region (few instructions). And with this
patch, this does not even happen very often, because the fast path don't
rely on the wildcards.
>
> So no more design trade-off question from me :)
>
Are you sure ? Scrolling down....
>> @@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
>> for_each_possible_cpu(cpu) {
>> /*
>> * Scan CPU slots.
>> - * Forward progress against recurring wildcards is guaranteed
>> - * by scanning for one wildcard while new elements use the
>> - * other wildcard value (1UL vs 2UL).
>> * Forward progress against recurring single hazard pointer
>> * values is guaranteed by the fact that a hazard pointer
>> * is not reclaimed nor reused until the scan for that hazard
>> * pointer completes, which prevents a steady flow of readers
>> * to acquire that same hazard pointer value.
>
> (Not a comment to this patch, but I think it's worth bringing up)
>
> I want to point out this is not true for the lockdep use case, because
> the we need to protect a hash list deletion there, and we use the
> address of the hash bucket there. It's proven fine in practice because
> the readers are rare (we only call the reader is_dynamic_key() in
> register_lock_class(), that is every time you have a new lock class to
> register).
>
> Maybe what we want to say here is that "if the users guarantee no steady
> flow of the same hazard pointer value, we guarantee forward progress".
> Thoughts?
AFAIU, your approach to protect lockdep linked lists is to use the
address of the hash bucket to protect the traversal. As this address is
invariant (global array item address), that address should be fine
to fulfill hazptr requirements, but it has downsides: rather than
protecting the specific nodes being retired, the whole hash chain is
protected. This means that, as you point out, many readers retiring
nodes from a given bucket (except the first node) could end up holding a
continuous stream of hazptr for a given hazptr value, preventing
progress of hazptr synchronize.
It's also coarser: per-bucket rather than per-node.
Am I missing something here ?
One honest question: is this pattern something we expect to
see often ? If so, then we may want to introduce a notion of
hazptr protection "period" flip (similar to some RCU implementations),
where we tag the low bit of the slot pointer (0 vs 1), and alternate
between the two periods in synchronize. This would prevent a steady-flow
of same-value readers from preventing synchronize forward progress.
Thoughts ?
Thanks,
Mathieu
>
> The rest looks good to me.
>
> Regards,
> Boqun
>
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 17:15 ` Mathieu Desnoyers
@ 2026-09-27 17:24 ` Boqun Feng
2026-09-27 17:36 ` Mathieu Desnoyers
2026-09-27 17:26 ` Boqun Feng
2026-09-27 22:39 ` Gary Guo
2 siblings, 1 reply; 22+ messages in thread
From: Boqun Feng @ 2026-09-27 17:24 UTC (permalink / raw)
To: Mathieu Desnoyers
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu, lkmm
On Sun, Sep 27, 2026 at 01:15:39PM -0400, Mathieu Desnoyers wrote:
[...]
> > > @@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
> > > for_each_possible_cpu(cpu) {
> > > /*
> > > * Scan CPU slots.
> > > - * Forward progress against recurring wildcards is guaranteed
> > > - * by scanning for one wildcard while new elements use the
> > > - * other wildcard value (1UL vs 2UL).
> > > * Forward progress against recurring single hazard pointer
> > > * values is guaranteed by the fact that a hazard pointer
> > > * is not reclaimed nor reused until the scan for that hazard
> > > * pointer completes, which prevents a steady flow of readers
> > > * to acquire that same hazard pointer value.
> >
> > (Not a comment to this patch, but I think it's worth bringing up)
> >
> > I want to point out this is not true for the lockdep use case, because
> > the we need to protect a hash list deletion there, and we use the
> > address of the hash bucket there. It's proven fine in practice because
> > the readers are rare (we only call the reader is_dynamic_key() in
> > register_lock_class(), that is every time you have a new lock class to
> > register).
> >
> > Maybe what we want to say here is that "if the users guarantee no steady
> > flow of the same hazard pointer value, we guarantee forward progress".
> > Thoughts?
>
> AFAIU, your approach to protect lockdep linked lists is to use the
> address of the hash bucket to protect the traversal. As this address is
> invariant (global array item address), that address should be fine
> to fulfill hazptr requirements, but it has downsides: rather than
> protecting the specific nodes being retired, the whole hash chain is
> protected. This means that, as you point out, many readers retiring
> nodes from a given bucket (except the first node) could end up holding a
> continuous stream of hazptr for a given hazptr value, preventing
> progress of hazptr synchronize.
>
> It's also coarser: per-bucket rather than per-node.
>
> Am I missing something here ?
>
No, you got it right, but as I said, we can use it in lockdep since the
readers are relatively rare, so not an issue here.
> One honest question: is this pattern something we expect to
> see often ? If so, then we may want to introduce a notion of
I honestly don't know. But in my opinion, we'd better focus on finding
more typical usage of hazptr (i.e. protecting actual object). So ...
> hazptr protection "period" flip (similar to some RCU implementations),
> where we tag the low bit of the slot pointer (0 vs 1), and alternate
> between the two periods in synchronize. This would prevent a steady-flow
> of same-value readers from preventing synchronize forward progress.
>
> Thoughts ?
>
... I will say let's add it only if we have more users of this pattern.
Regards,
Boqun
> Thanks,
>
> Mathieu
>
> >
> > The rest looks good to me.
> >
> > Regards,
> > Boqun
> >
>
>
> --
> Mathieu Desnoyers
> EfficiOS Inc.
> https://www.efficios.com
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 17:15 ` Mathieu Desnoyers
2026-09-27 17:24 ` Boqun Feng
@ 2026-09-27 17:26 ` Boqun Feng
2026-09-27 22:39 ` Gary Guo
2 siblings, 0 replies; 22+ messages in thread
From: Boqun Feng @ 2026-09-27 17:26 UTC (permalink / raw)
To: Mathieu Desnoyers
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu, lkmm
On Sun, Sep 27, 2026 at 01:15:39PM -0400, Mathieu Desnoyers wrote:
[...]
> > > With this change, the scan on per-CPU slots don't need to expect a
> > > wildcard anymore, because none can be produced by readers. Wildcards are
> > > only expected within overflow lists.
> > >
> >
> > Ok, I was missing something, but I think it's better to call it out.
> > Wildcards can only exist in the overflow lists when the context is not
> > preemptible. In other words, there won't be a preempted readers blocking
> > the synchronize_hazptr() with a wilcard in the overflow list.
>
> Exactly ! Wildcard slots only exist during the short time-frame of the
> preempt-off read-side code region (few instructions). And with this
> patch, this does not even happen very often, because the fast path don't
> rely on the wildcards.
>
> >
> > So no more design trade-off question from me :)
> >
>
> Are you sure ? Scrolling down....
>
I do feel with scan thread support and call_hazptr() support, we will
need to tweak the design and implementation some more, but right now it
does look good to me :)
Regards,
Boqun
[...]
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 17:24 ` Boqun Feng
@ 2026-09-27 17:36 ` Mathieu Desnoyers
2026-09-27 18:16 ` Boqun Feng
0 siblings, 1 reply; 22+ messages in thread
From: Mathieu Desnoyers @ 2026-09-27 17:36 UTC (permalink / raw)
To: Boqun Feng
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu,
lkmm, Lian Wang, Kunwu Chan
On 2026-09-27 13:24, Boqun Feng wrote:
> On Sun, Sep 27, 2026 at 01:15:39PM -0400, Mathieu Desnoyers wrote:
> [...]
>>>> @@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
>>>> for_each_possible_cpu(cpu) {
>>>> /*
>>>> * Scan CPU slots.
>>>> - * Forward progress against recurring wildcards is guaranteed
>>>> - * by scanning for one wildcard while new elements use the
>>>> - * other wildcard value (1UL vs 2UL).
>>>> * Forward progress against recurring single hazard pointer
>>>> * values is guaranteed by the fact that a hazard pointer
>>>> * is not reclaimed nor reused until the scan for that hazard
>>>> * pointer completes, which prevents a steady flow of readers
>>>> * to acquire that same hazard pointer value.
>>>
>>> (Not a comment to this patch, but I think it's worth bringing up)
>>>
>>> I want to point out this is not true for the lockdep use case, because
>>> the we need to protect a hash list deletion there, and we use the
>>> address of the hash bucket there. It's proven fine in practice because
>>> the readers are rare (we only call the reader is_dynamic_key() in
>>> register_lock_class(), that is every time you have a new lock class to
>>> register).
>>>
>>> Maybe what we want to say here is that "if the users guarantee no steady
>>> flow of the same hazard pointer value, we guarantee forward progress".
>>> Thoughts?
>>
>> AFAIU, your approach to protect lockdep linked lists is to use the
>> address of the hash bucket to protect the traversal. As this address is
>> invariant (global array item address), that address should be fine
>> to fulfill hazptr requirements, but it has downsides: rather than
>> protecting the specific nodes being retired, the whole hash chain is
>> protected. This means that, as you point out, many readers retiring
>> nodes from a given bucket (except the first node) could end up holding a
>> continuous stream of hazptr for a given hazptr value, preventing
>> progress of hazptr synchronize.
>>
>> It's also coarser: per-bucket rather than per-node.
>>
>> Am I missing something here ?
>>
>
> No, you got it right, but as I said, we can use it in lockdep since the
> readers are relatively rare, so not an issue here.
>
>> One honest question: is this pattern something we expect to
>> see often ? If so, then we may want to introduce a notion of
>
> I honestly don't know. But in my opinion, we'd better focus on finding
> more typical usage of hazptr (i.e. protecting actual object). So ...
>
>> hazptr protection "period" flip (similar to some RCU implementations),
>> where we tag the low bit of the slot pointer (0 vs 1), and alternate
>> between the two periods in synchronize. This would prevent a steady-flow
>> of same-value readers from preventing synchronize forward progress.
>>
>> Thoughts ?
>>
>
> ... I will say let's add it only if we have more users of this pattern.
>
I am concerned about this because many uses of RCU in the Linux
kernel protects linked list traversals. Turning a RCU-protected list
traversal into a hazptr protected traversal is not as simple as
acquiring each hazptr hand in hand.
Your own use-case for lockdep is indeed a linked list traversal,
and you need to use a work-around: protect the address of the
hash bucket head.
I'm just wondering if this work-around will end up being the
"blessed" way for protecting linked list traversals with hazptr,
or whether we should consider alternatives ?
This ties into finding additional usage for hazptr, because linked
lists are so prevalent in the kernel.
Thanks,
Mathieu
> Regards,
> Boqun
>
>> Thanks,
>>
>> Mathieu
>>
>>>
>>> The rest looks good to me.
>>>
>>> Regards,
>>> Boqun
>>>
>>
>>
>> --
>> Mathieu Desnoyers
>> EfficiOS Inc.
>> https://www.efficios.com
--
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 17:36 ` Mathieu Desnoyers
@ 2026-09-27 18:16 ` Boqun Feng
0 siblings, 0 replies; 22+ messages in thread
From: Boqun Feng @ 2026-09-27 18:16 UTC (permalink / raw)
To: Mathieu Desnoyers
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu,
lkmm, Lian Wang, Kunwu Chan
On Sun, Sep 27, 2026 at 01:36:54PM -0400, Mathieu Desnoyers wrote:
> On 2026-09-27 13:24, Boqun Feng wrote:
> > On Sun, Sep 27, 2026 at 01:15:39PM -0400, Mathieu Desnoyers wrote:
> > [...]
> > > > > @@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
> > > > > for_each_possible_cpu(cpu) {
> > > > > /*
> > > > > * Scan CPU slots.
> > > > > - * Forward progress against recurring wildcards is guaranteed
> > > > > - * by scanning for one wildcard while new elements use the
> > > > > - * other wildcard value (1UL vs 2UL).
> > > > > * Forward progress against recurring single hazard pointer
> > > > > * values is guaranteed by the fact that a hazard pointer
> > > > > * is not reclaimed nor reused until the scan for that hazard
> > > > > * pointer completes, which prevents a steady flow of readers
> > > > > * to acquire that same hazard pointer value.
> > > >
> > > > (Not a comment to this patch, but I think it's worth bringing up)
> > > >
> > > > I want to point out this is not true for the lockdep use case, because
> > > > the we need to protect a hash list deletion there, and we use the
> > > > address of the hash bucket there. It's proven fine in practice because
> > > > the readers are rare (we only call the reader is_dynamic_key() in
> > > > register_lock_class(), that is every time you have a new lock class to
> > > > register).
> > > >
> > > > Maybe what we want to say here is that "if the users guarantee no steady
> > > > flow of the same hazard pointer value, we guarantee forward progress".
> > > > Thoughts?
> > >
> > > AFAIU, your approach to protect lockdep linked lists is to use the
> > > address of the hash bucket to protect the traversal. As this address is
> > > invariant (global array item address), that address should be fine
> > > to fulfill hazptr requirements, but it has downsides: rather than
> > > protecting the specific nodes being retired, the whole hash chain is
> > > protected. This means that, as you point out, many readers retiring
> > > nodes from a given bucket (except the first node) could end up holding a
> > > continuous stream of hazptr for a given hazptr value, preventing
> > > progress of hazptr synchronize.
> > >
> > > It's also coarser: per-bucket rather than per-node.
> > >
> > > Am I missing something here ?
> > >
> >
> > No, you got it right, but as I said, we can use it in lockdep since the
> > readers are relatively rare, so not an issue here.
> >
> > > One honest question: is this pattern something we expect to
> > > see often ? If so, then we may want to introduce a notion of
> >
> > I honestly don't know. But in my opinion, we'd better focus on finding
> > more typical usage of hazptr (i.e. protecting actual object). So ...
> >
> > > hazptr protection "period" flip (similar to some RCU implementations),
> > > where we tag the low bit of the slot pointer (0 vs 1), and alternate
> > > between the two periods in synchronize. This would prevent a steady-flow
> > > of same-value readers from preventing synchronize forward progress.
> > >
> > > Thoughts ?
> > >
> >
> > ... I will say let's add it only if we have more users of this pattern.
> >
>
> I am concerned about this because many uses of RCU in the Linux
> kernel protects linked list traversals. Turning a RCU-protected list
> traversal into a hazptr protected traversal is not as simple as
> acquiring each hazptr hand in hand.
>
I actually don't think converting RCU usage into hazptr would be a good
starting point for finding good hazptr usage. RCU readers are faster,
and a lot of existing RCU users do want the reader side to be as fast as
possible. So even though this is a problem, but I don't think that's an
urgent problem to resolve. Of course, if one can find a faster hazptr
reader implemenation, then it may be a different story.
> Your own use-case for lockdep is indeed a linked list traversal,
> and you need to use a work-around: protect the address of the
> hash bucket head.
>
So the lockdep usage is IMO a very special case, but it does show the
advantages of hazptr.
> I'm just wondering if this work-around will end up being the
> "blessed" way for protecting linked list traversals with hazptr,
> or whether we should consider alternatives ?
>
It depends on how many uses of the linked list traversals really care
about the reader side performance. I'm not sure we can easily find
another one like lockdep. Hence I think we should only add the
optimization when we find more of such a usage. Make sense?
Personally, I also want to see a refcount-like usage for hazptr, that'll
be very exciting for me :D
Regards,
Boqun
> This ties into finding additional usage for hazptr, because linked
> lists are so prevalent in the kernel.
>
> Thanks,
>
> Mathieu
>
> > Regards,
> > Boqun
> >
> > > Thanks,
> > >
> > > Mathieu
> > >
> > > >
> > > > The rest looks good to me.
> > > >
> > > > Regards,
> > > > Boqun
> > > >
> > >
> > >
> > > --
> > > Mathieu Desnoyers
> > > EfficiOS Inc.
> > > https://www.efficios.com
>
>
> --
> Mathieu Desnoyers
> EfficiOS Inc.
> https://www.efficios.com
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 17:15 ` Mathieu Desnoyers
2026-09-27 17:24 ` Boqun Feng
2026-09-27 17:26 ` Boqun Feng
@ 2026-09-27 22:39 ` Gary Guo
2026-09-28 9:12 ` Boqun Feng
2 siblings, 1 reply; 22+ messages in thread
From: Gary Guo @ 2026-09-27 22:39 UTC (permalink / raw)
To: Mathieu Desnoyers, Boqun Feng
Cc: Paul E . McKenney, linux-kernel, Bradley Morgan, Gary Guo, rcu, lkmm
On Sun Sep 27, 2026 at 6:15 PM BST, Mathieu Desnoyers wrote:
> On 2026-09-27 12:40, Boqun Feng wrote:
>
>> I want to point out this is not true for the lockdep use case, because
>> the we need to protect a hash list deletion there, and we use the
>> address of the hash bucket there. It's proven fine in practice because
>> the readers are rare (we only call the reader is_dynamic_key() in
>> register_lock_class(), that is every time you have a new lock class to
>> register).
>>
>> Maybe what we want to say here is that "if the users guarantee no steady
>> flow of the same hazard pointer value, we guarantee forward progress".
>> Thoughts?
>
> AFAIU, your approach to protect lockdep linked lists is to use the
> address of the hash bucket to protect the traversal. As this address is
> invariant (global array item address), that address should be fine
> to fulfill hazptr requirements, but it has downsides: rather than
> protecting the specific nodes being retired, the whole hash chain is
> protected. This means that, as you point out, many readers retiring
> nodes from a given bucket (except the first node) could end up holding a
> continuous stream of hazptr for a given hazptr value, preventing
> progress of hazptr synchronize.
>
> It's also coarser: per-bucket rather than per-node.
>
> Am I missing something here ?
>
> One honest question: is this pattern something we expect to
> see often ? If so, then we may want to introduce a notion of
> hazptr protection "period" flip (similar to some RCU implementations),
> where we tag the low bit of the slot pointer (0 vs 1), and alternate
> between the two periods in synchronize. This would prevent a steady-flow
> of same-value readers from preventing synchronize forward progress.
Slightly off topic, but I have a use-case in mind (in case you're not already
aware) where the address is fixed like the lockdep class, but it does not suffer
the forward progress guarantee.
I have been wanting to use hazptr for revocable for quite a while (I think I
chatted with Boqun about this last LPC). For the revocable use case, the
protected pointer is fixed, however there is an additional boolean flag to
determine if the resource been revoked or not.
Something like this:
void *revocable_try_access(struct revocable *rev) {
struct hazptr_ctx ctx = {};
if (READ_ONCE(rev->revoked))
return NULL;
// note the & cancels out with the * in acquire, so the address is fixed.
hazptr_acquire(&ctx, &rev);
if (READ_ONCE(rev->revoked))
return NULL;
return rev->res;
}
void revocable_revoke(struct revocable *rev) {
WRITE_ONCE(rev->revoked, true);
hazptr_synchronize(&rev);
}
So while the address is fixed, we have a different field to do the unpublishing
part.
For some context, the current Rust revocable implementation uses RCU, but this
is limiting the case where it can be used. The current C revocable series in
https://lore.kernel.org/all/20260912123529.7951-1-tzungbi@kernel.org/ uses SRCU.
I think this use case is a good one for hazptr, in fact, I have encouraged Alvin
Sun to try it out and you can see an implementation (Rust) in here:
https://lore.kernel.org/rust-for-linux/20260326-b4-tyr-debugfs-v1-6-074badd18716@linux.dev/
Although, over the course of the year, we have been reducing the amount of
revocable usage and shifting to represent things with lifetime..
Best,
Gary
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 22:39 ` Gary Guo
@ 2026-09-28 9:12 ` Boqun Feng
2026-09-28 11:32 ` Gary Guo
0 siblings, 1 reply; 22+ messages in thread
From: Boqun Feng @ 2026-09-28 9:12 UTC (permalink / raw)
To: Gary Guo
Cc: Mathieu Desnoyers, Paul E . McKenney, linux-kernel,
Bradley Morgan, rcu, lkmm
On Sun, Sep 27, 2026 at 11:39:00PM +0100, Gary Guo wrote:
> On Sun Sep 27, 2026 at 6:15 PM BST, Mathieu Desnoyers wrote:
> > On 2026-09-27 12:40, Boqun Feng wrote:
> >
> >> I want to point out this is not true for the lockdep use case, because
> >> the we need to protect a hash list deletion there, and we use the
> >> address of the hash bucket there. It's proven fine in practice because
> >> the readers are rare (we only call the reader is_dynamic_key() in
> >> register_lock_class(), that is every time you have a new lock class to
> >> register).
> >>
> >> Maybe what we want to say here is that "if the users guarantee no steady
> >> flow of the same hazard pointer value, we guarantee forward progress".
> >> Thoughts?
> >
> > AFAIU, your approach to protect lockdep linked lists is to use the
> > address of the hash bucket to protect the traversal. As this address is
> > invariant (global array item address), that address should be fine
> > to fulfill hazptr requirements, but it has downsides: rather than
> > protecting the specific nodes being retired, the whole hash chain is
> > protected. This means that, as you point out, many readers retiring
> > nodes from a given bucket (except the first node) could end up holding a
> > continuous stream of hazptr for a given hazptr value, preventing
> > progress of hazptr synchronize.
> >
> > It's also coarser: per-bucket rather than per-node.
> >
> > Am I missing something here ?
> >
> > One honest question: is this pattern something we expect to
> > see often ? If so, then we may want to introduce a notion of
> > hazptr protection "period" flip (similar to some RCU implementations),
> > where we tag the low bit of the slot pointer (0 vs 1), and alternate
> > between the two periods in synchronize. This would prevent a steady-flow
> > of same-value readers from preventing synchronize forward progress.
>
> Slightly off topic, but I have a use-case in mind (in case you're not already
> aware) where the address is fixed like the lockdep class, but it does not suffer
> the forward progress guarantee.
>
> I have been wanting to use hazptr for revocable for quite a while (I think I
> chatted with Boqun about this last LPC). For the revocable use case, the
> protected pointer is fixed, however there is an additional boolean flag to
> determine if the resource been revoked or not.
>
> Something like this:
>
> void *revocable_try_access(struct revocable *rev) {
> struct hazptr_ctx ctx = {};
> if (READ_ONCE(rev->revoked))
> return NULL;
> // note the & cancels out with the * in acquire, so the address is fixed.
> hazptr_acquire(&ctx, &rev);
> if (READ_ONCE(rev->revoked))
> return NULL;
> return rev->res;
> }
>
> void revocable_revoke(struct revocable *rev) {
> WRITE_ONCE(rev->revoked, true);
> hazptr_synchronize(&rev);
> }
>
> So while the address is fixed, we have a different field to do the unpublishing
> part.
>
> For some context, the current Rust revocable implementation uses RCU, but this
> is limiting the case where it can be used. The current C revocable series in
> https://lore.kernel.org/all/20260912123529.7951-1-tzungbi@kernel.org/ uses SRCU.
>
> I think this use case is a good one for hazptr, in fact, I have encouraged Alvin
> Sun to try it out and you can see an implementation (Rust) in here:
> https://lore.kernel.org/rust-for-linux/20260326-b4-tyr-debugfs-v1-6-074badd18716@linux.dev/
> Although, over the course of the year, we have been reducing the amount of
> revocable usage and shifting to represent things with lifetime..
>
Paul asked me for the Rust usage a few days ago, and since we moved to
a lifetime based approach [1], which is better IMO, I don't think
switching to hazptr will should observable difference here, especially
for a real workload improvement. It's still worth trying to see how
hazptr in Rust would work for this, but it's less a sufficient condition
to merge hazptr from my understanding. Of course I could be wrong.
[1]: https://lore.kernel.org/rust-for-linux/20260517000149.3226762-1-dakr@kernel.org/
Regards,
Boqun
Regards,
Boqun
> Best,
> Gary
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-27 16:40 ` Boqun Feng
2026-09-27 17:15 ` Mathieu Desnoyers
@ 2026-09-28 9:27 ` Kunwu Chan
1 sibling, 0 replies; 22+ messages in thread
From: Kunwu Chan @ 2026-09-28 9:27 UTC (permalink / raw)
To: Boqun Feng
Cc: Kunwu Chan, Mathieu Desnoyers, Paul E . McKenney, linux-kernel,
Bradley Morgan, Gary Guo, rcu, lkmm
On Sun, 27 Sep 2026 18:40:53 +0200 Boqun Feng <boqun@kernel.org> wrote:
> On Sun, Sep 27, 2026 at 11:51:31AM -0400, Mathieu Desnoyers wrote:
> > Introduce a "try acquire" hazard pointer fast path, which performs an
> > early load of the address to store it into the hazard pointer slot, and
> > then re-loads that address after a barrier to check whether it has
> > changed meanwhile.
> >
> > On comparison failure, rather than re-try, guarantee forward progress by
> > falling back to the __hazptr_acquire slow path on failure.
> >
> > The acquire slow path attempts a try-acquire for any available per-CPU
> > slot. If that fails, it chains the backup slot into the overflow list,
> > therefore guaranteeing forward progress for both hazard pointer
> > read-side and synchronize:
> >
> > - Readers set the wildcard, and then proceed to set the more
> > specific address to replace the wildcard.
> >
> > - One synchronize alternates between two overflow list periods,
> > scanning each one while readers are added to the other period,
> > thus preventing a steady flow of readers from preventing
> > synchronize forward progress.
> >
> > With this change, the scan on per-CPU slots don't need to expect a
> > wildcard anymore, because none can be produced by readers. Wildcards are
> > only expected within overflow lists.
> >
>
> Ok, I was missing something, but I think it's better to call it out.
> Wildcards can only exist in the overflow lists when the context is not
> preemptible. In other words, there won't be a preempted readers blocking
> the synchronize_hazptr() with a wilcard in the overflow list.
>
> So no more design trade-off question from me :)
>
> Regards,
> Boqun
>
> > Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
> > Cc: Paul E. McKenney <paulmck@kernel.org>
> > Cc: Boqun Feng <boqun@kernel.org>
> > Cc: Bradley Morgan <brads@mainlining.org>
> > Cc: Gary Guo <gary@garyguo.net>
> > Cc: <rcu@vger.kernel.org>
> > Cc: <lkmm@lists.linux.dev>
> > ---
> > include/linux/hazptr.h | 47 +++++++++++--------
> > kernel/hazptr.c | 103 ++++++++++++++++++++---------------------
> > 2 files changed, 76 insertions(+), 74 deletions(-)
> >
> > diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
> > index d1670121947a..fcf017ca2255 100644
> > --- a/include/linux/hazptr.h
> > +++ b/include/linux/hazptr.h
> > @@ -29,9 +29,6 @@
> > /* 4 slots (each sizeof(hazptr_slot_item)) fit in a single 64-byte cache line. */
> > #define NR_HAZPTR_PERCPU_SLOTS 4
> >
> > -/* The current hazard pointer wildcard. */
> > -extern void *hazptr_wildcard;
> > -
> > /*
> > * Hazard pointer slot.
> > */
> > @@ -190,6 +187,31 @@ void hazptr_note_context_switch(void)
> > }
> > }
> >
> > +/* Try hazard pointer protection. */
> > +static inline
> > +void *__hazptr_try_acquire(struct hazptr_ctx *ctx, void * const *addr_p, struct hazptr_slot *slot)
> > +{
> > + void *early_addr, *addr;
> > +
> > + if (unlikely(slot->addr))
> > + return NULL;
> > + early_addr = READ_ONCE(*addr_p); /* Early load. */
> > + WRITE_ONCE(slot->addr, early_addr); /* Store B */
> > + /* Memory ordering: Store B before Load A. */
> > + smp_mb();
> > + addr = READ_ONCE(*addr_p); /* Load A */
> > + /*
> > + * Validate that address did not change between Early load and Load A.
> > + * Use ptr_eq() to make sure that result from Load A is returned to the
> > + * caller to preserve address dependency.
> > + */
> > + if (unlikely(!ptr_eq(addr, early_addr))) {
> > + WRITE_ONCE(slot->addr, NULL);
> > + return NULL;
> > + }
> > + return addr;
> > +}
> > +
> > /**
> > * hazptr_acquire - Load pointer at address and protect with hazard pointer.
> > *
> > @@ -245,24 +267,9 @@ void *hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> > ctx->acquire_cpu = smp_processor_id();
> > ctx->acquire_caller = _THIS_IP_;
> > #endif
> > - if (unlikely(slot->addr))
> > + addr = __hazptr_try_acquire(ctx, addr_p, slot);
> > + if (unlikely(!addr))
> > return __hazptr_acquire(ctx, addr_p);
> > - WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard)); /* Store B */
> > -
> > - /* Memory ordering: Store B before Load A. */
> > - smp_mb();
> > -
> > - /*
> > - * Load @addr_p after storing wildcard to the hazard pointer slot.
> > - */
> > - addr = READ_ONCE(*addr_p); /* Load A */
> > -
> > - /*
> > - * We don't care about ordering of Store C. It will simply
> > - * replace the wildcard by a more specific address. If addr is
> > - * NULL, we simply store NULL into the slot.
> > - */
> > - WRITE_ONCE(slot->addr, addr); /* Store C */
> > slot_item->ctx.ctx = ctx;
> > ctx->slot = slot;
> > return addr;
> > diff --git a/kernel/hazptr.c b/kernel/hazptr.c
> > index 13faa5ba7677..3ca73b56a5c2 100644
> > --- a/kernel/hazptr.c
> > +++ b/kernel/hazptr.c
> > @@ -13,17 +13,9 @@
> > #include <linux/list.h>
> > #include <linux/export.h>
> >
> > -static DEFINE_MUTEX(hazptr_phase_lock); /* Protect the wildcard and list phase flip. */
> > +#define HAZPTR_WILDCARD ((void *) 1UL)
> >
> > -/*
> > - * The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
> > - * hazptr_synchronize forward progress even with a steady stream of readers.
> > - * This wildcard value is used by acquire to temporarily tag the per-CPU slots.
> > - * This also affects the overflow list selection: the current list used by
> > - * readers is array[(unsigned long) hazptr_wildcard - 1].
> > - */
> > -void *hazptr_wildcard = (void *) 1UL;
> > -EXPORT_SYMBOL_GPL(hazptr_wildcard);
> > +static DEFINE_MUTEX(hazptr_phase_lock); /* Protect the list phase flip. */
> >
> > /* The current overflow list phase. */
> > static unsigned int hazptr_overflow_list_phase;
> > @@ -41,6 +33,10 @@ struct hazptr_overflow_list {
> > * successively iterates on both lists. Therefore, only list removals
> > * can cause the iteration to retry, and the number of removals is
> > * limited to the number of list elements.
> > + *
> > + * Due to the overflow list raw spin lock, the hazard pointer readers are
> > + * blocking, starvation-free with bounded waiting, assuming bounded critical
> > + * sections and no NMI or virtualization-induced holder preemption.
> > */
> > struct hazptr_overflow_list_flip {
> > struct hazptr_overflow_list array[2];
> > @@ -51,26 +47,12 @@ static DEFINE_PER_CPU(struct hazptr_overflow_list_flip, percpu_overflow_list_fli
> > DEFINE_PER_CPU(struct hazptr_percpu_slots, hazptr_percpu_slots);
> > EXPORT_PER_CPU_SYMBOL_GPL(hazptr_percpu_slots);
> >
> > -static
> > -void *flip_wildcard(void *wildcard)
> > -{
> > - return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
> > -}
> > -
> > static
> > unsigned int flip_list_phase(unsigned int phase)
> > {
> > return 1 - phase;
> > }
> >
> > -static
> > -bool is_wildcard(void *addr)
> > -{
> > - if ((unsigned long) addr == 1UL || (unsigned long) addr == 2UL)
> > - return true;
> > - return false;
> > -}
> > -
> > static
> > struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
> > {
> > @@ -96,16 +78,36 @@ struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
> > */
> > void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> > {
> > - struct hazptr_slot *slot = hazptr_get_free_percpu_slot(ctx);
> > + struct hazptr_slot *slot;
> > void *addr;
> >
> > /*
> > - * If all the per-CPU slots are already in use, fallback
> > - * to the backup slot.
> > + * In case we are called due to nested use of hazard pointers,
> > + * try a slot protection with per-CPU slots.
> > */
> > - if (unlikely(!slot))
> > - slot = hazptr_chain_backup_slot(ctx);
> > - WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard)); /* Store B */
> > + slot = hazptr_get_free_percpu_slot(ctx);
> > + if (likely(slot)) {
> > + addr = __hazptr_try_acquire(ctx, addr_p, slot);
> > + if (addr) {
> > + ctx->slot = slot;
> > + return addr;
> > + }
> > + }
> > +
> > + /*
> > + * The backup slot overflow list guarantees forward progress of both
> > + * hazard pointer readers and synchronize:
> > + *
> > + * - Readers set the wildcard, and then proceed to set the more
> > + * specific address to replace the wildcard.
> > + *
> > + * - One synchronize alternates between two overflow list periods,
> > + * scanning each one while readers are added to the other period,
> > + * thus preventing a steady flow of readers from preventing
> > + * synchronize forward progress.
> > + */
> > + slot = hazptr_chain_backup_slot(ctx);
> > + WRITE_ONCE(slot->addr, HAZPTR_WILDCARD); /* Store B */
> >
> > /* Memory ordering: Store B before Load A. */
> > smp_mb();
> > @@ -121,20 +123,21 @@ void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> > * NULL, we simply store NULL into the slot.
> > */
> > WRITE_ONCE(slot->addr, addr); /* Store C */
> > +
> > ctx->slot = slot;
> > - if (!addr && hazptr_slot_is_backup(ctx, slot))
> > + if (!addr)
> > hazptr_unchain_backup_slot(ctx);
> > return addr;
> > }
> > EXPORT_SYMBOL_GPL(__hazptr_acquire);
> >
> > /*
> > - * Perform piecewise iteration on overflow list waiting until "addr" is
> > - * not present. Raw spinlock is released and taken between each list
> > - * item and busy loop iteration. The overflow list generation is checked
> > - * each time the lock is taken to validate that the list has not changed
> > - * before resuming iteration or busy wait. If the generation has
> > - * changed, retry the entire list traversal.
> > + * Perform piecewise iteration on overflow list waiting until "addr" and
> > + * wildcard are not present. Raw spinlock is released and taken between each
> > + * list item and busy loop iteration. The overflow list generation is checked
> > + * each time the lock is taken to validate that the list has not changed before
> > + * resuming iteration or busy wait. If the generation has changed, retry the
> > + * entire list traversal.
> > */
> > static
> > void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list, void *addr)
> > @@ -147,13 +150,11 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
> > retry:
> > snapshot_gen = overflow_list->gen;
> > hlist_for_each_entry(backup_slot, &overflow_list->head, overflow_node) {
> > - /* Busy-wait if node is found. */
> > + /* Busy-wait if addr or wildcard are found. */
> > for (;;) {
> > void *load_addr = smp_load_acquire(&backup_slot->slot.addr); /* Load B */
> >
> > - /* We don't expect wildcards in overflow list. */
> > - WARN_ON_ONCE(is_wildcard(load_addr));
> > - if (load_addr != addr)
> > + if (load_addr != addr && load_addr != HAZPTR_WILDCARD)
> > break;
> > raw_spin_unlock_irqrestore(&overflow_list->lock, flags);
> > cpu_relax();
> > @@ -174,7 +175,7 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
> > }
> >
> > static
> > -void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
> > +void hazptr_synchronize_cpu_slots(int cpu, void *addr)
> > {
> > struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
> > unsigned int idx;
> > @@ -182,13 +183,13 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
> > for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
> > struct hazptr_slot_item *item = &percpu_slots->items[idx];
> >
> > - /* Busy-wait if node is found. */
> > - smp_cond_load_acquire(&item->slot.addr, VAL != addr && VAL != scan_wildcard); /* Load B */
> > + /* Busy-wait if addr is found. */
> > + smp_cond_load_acquire(&item->slot.addr, VAL != addr); /* Load B */
> > }
> > }
> >
> > static
> > -void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
> > +void hazptr_scan_cpu_slots(void *addr)
> > {
> > int cpu;
> >
> > @@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
> > for_each_possible_cpu(cpu) {
> > /*
> > * Scan CPU slots.
> > - * Forward progress against recurring wildcards is guaranteed
> > - * by scanning for one wildcard while new elements use the
> > - * other wildcard value (1UL vs 2UL).
> > * Forward progress against recurring single hazard pointer
> > * values is guaranteed by the fact that a hazard pointer
> > * is not reclaimed nor reused until the scan for that hazard
> > * pointer completes, which prevents a steady flow of readers
> > * to acquire that same hazard pointer value.
>
> (Not a comment to this patch, but I think it's worth bringing up)
>
> I want to point out this is not true for the lockdep use case, because
> the we need to protect a hash list deletion there, and we use the
> address of the hash bucket there. It's proven fine in practice because
> the readers are rare (we only call the reader is_dynamic_key() in
> register_lock_class(), that is every time you have a new lock class to
> register).
>
> Maybe what we want to say here is that "if the users guarantee no steady
> flow of the same hazard pointer value, we guarantee forward progress".
> Thoughts?
Thanks Mathieu, and thanks Boqun for the review.
I hit the same ordering requirement in the shared-scan
implementation I'm working on. hazptr_promote_to_backup_slot()
links the backup slot into the overflow list before clearing the
per-CPU slot, so the scan has to cover all per-CPU slots before
the overflow lists. My series carries the same ordering fix in
the scan-kthread path.
For the lockdep case, I've tested the conversion with the
wq_churn/LOCKDEP torture scenario and PROVE_LOCKING enabled.
The hazard pointer is the hash bucket address there, so I agree
that the forward-progress guarantee should be stated conditional
on there being no steady flow of the same hazard pointer value.
I'm also working on the scan-thread side, where concurrent
hazptr_synchronize() callers share one scan cycle, based on
the scan-kthread approach from your shazptr series [1].
[1] https://lore.kernel.org/lkml/20250625031101.12555-1-boqun.feng@gmail.com/
The v2 series is testing, I'll include the rcuscale results
in the cover letter.
Thanks,
Kunwu
>
> The rest looks good to me.
>
> Regards,
> Boqun
>
> > */
> > - hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
> > + hazptr_synchronize_cpu_slots(cpu, addr);
> > }
> > }
> >
> > @@ -237,7 +235,6 @@ void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
> > void hazptr_synchronize(void *addr)
> > {
> > unsigned int scan_list_phase;
> > - void *scan_wildcard;
> >
> > /*
> > * Busy-wait should only be done from preemptible context.
> > @@ -251,16 +248,14 @@ void hazptr_synchronize(void *addr)
> > */
> > if (!addr)
> > return;
> > +
> > /* Memory ordering: Store A before Load B. */
> > smp_mb();
> >
> > guard(mutex)(&hazptr_phase_lock);
> >
> > /* Scan per-CPU slots. */
> > - scan_wildcard = flip_wildcard(hazptr_wildcard);
> > - hazptr_scan_cpu_slots_period(addr, scan_wildcard);
> > - WRITE_ONCE(hazptr_wildcard, scan_wildcard); /* Flip the current wildcard. */
> > - hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
> > + hazptr_scan_cpu_slots(addr);
> >
> > /*
> > * Scan overflow lists *after* scanning per-CPU slots. See
> > --
> > 2.43.0
> >
>
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
2026-09-28 9:12 ` Boqun Feng
@ 2026-09-28 11:32 ` Gary Guo
0 siblings, 0 replies; 22+ messages in thread
From: Gary Guo @ 2026-09-28 11:32 UTC (permalink / raw)
To: Boqun Feng, Gary Guo
Cc: Mathieu Desnoyers, Paul E . McKenney, linux-kernel,
Bradley Morgan, rcu, lkmm
On Mon Sep 28, 2026 at 10:12 AM BST, Boqun Feng wrote:
> On Sun, Sep 27, 2026 at 11:39:00PM +0100, Gary Guo wrote:
>> On Sun Sep 27, 2026 at 6:15 PM BST, Mathieu Desnoyers wrote:
>> > On 2026-09-27 12:40, Boqun Feng wrote:
>> >
>> >> I want to point out this is not true for the lockdep use case, because
>> >> the we need to protect a hash list deletion there, and we use the
>> >> address of the hash bucket there. It's proven fine in practice because
>> >> the readers are rare (we only call the reader is_dynamic_key() in
>> >> register_lock_class(), that is every time you have a new lock class to
>> >> register).
>> >>
>> >> Maybe what we want to say here is that "if the users guarantee no steady
>> >> flow of the same hazard pointer value, we guarantee forward progress".
>> >> Thoughts?
>> >
>> > AFAIU, your approach to protect lockdep linked lists is to use the
>> > address of the hash bucket to protect the traversal. As this address is
>> > invariant (global array item address), that address should be fine
>> > to fulfill hazptr requirements, but it has downsides: rather than
>> > protecting the specific nodes being retired, the whole hash chain is
>> > protected. This means that, as you point out, many readers retiring
>> > nodes from a given bucket (except the first node) could end up holding a
>> > continuous stream of hazptr for a given hazptr value, preventing
>> > progress of hazptr synchronize.
>> >
>> > It's also coarser: per-bucket rather than per-node.
>> >
>> > Am I missing something here ?
>> >
>> > One honest question: is this pattern something we expect to
>> > see often ? If so, then we may want to introduce a notion of
>> > hazptr protection "period" flip (similar to some RCU implementations),
>> > where we tag the low bit of the slot pointer (0 vs 1), and alternate
>> > between the two periods in synchronize. This would prevent a steady-flow
>> > of same-value readers from preventing synchronize forward progress.
>>
>> Slightly off topic, but I have a use-case in mind (in case you're not already
>> aware) where the address is fixed like the lockdep class, but it does not suffer
>> the forward progress guarantee.
>>
>> I have been wanting to use hazptr for revocable for quite a while (I think I
>> chatted with Boqun about this last LPC). For the revocable use case, the
>> protected pointer is fixed, however there is an additional boolean flag to
>> determine if the resource been revoked or not.
>>
>> Something like this:
>>
>> void *revocable_try_access(struct revocable *rev) {
>> struct hazptr_ctx ctx = {};
>> if (READ_ONCE(rev->revoked))
>> return NULL;
>> // note the & cancels out with the * in acquire, so the address is fixed.
>> hazptr_acquire(&ctx, &rev);
>> if (READ_ONCE(rev->revoked))
>> return NULL;
>> return rev->res;
>> }
>>
>> void revocable_revoke(struct revocable *rev) {
>> WRITE_ONCE(rev->revoked, true);
>> hazptr_synchronize(&rev);
>> }
>>
>> So while the address is fixed, we have a different field to do the unpublishing
>> part.
>>
>> For some context, the current Rust revocable implementation uses RCU, but this
>> is limiting the case where it can be used. The current C revocable series in
>> https://lore.kernel.org/all/20260912123529.7951-1-tzungbi@kernel.org/ uses SRCU.
>>
>> I think this use case is a good one for hazptr, in fact, I have encouraged Alvin
>> Sun to try it out and you can see an implementation (Rust) in here:
>> https://lore.kernel.org/rust-for-linux/20260326-b4-tyr-debugfs-v1-6-074badd18716@linux.dev/
>> Although, over the course of the year, we have been reducing the amount of
>> revocable usage and shifting to represent things with lifetime..
>>
>
> Paul asked me for the Rust usage a few days ago, and since we moved to
> a lifetime based approach [1], which is better IMO, I don't think
> switching to hazptr will should observable difference here, especially
> for a real workload improvement. It's still worth trying to see how
> hazptr in Rust would work for this, but it's less a sufficient condition
> to merge hazptr from my understanding. Of course I could be wrong.
>
> [1]: https://lore.kernel.org/rust-for-linux/20260517000149.3226762-1-dakr@kernel.org/
What we get rid of is to use `Revocable` (or `DevRes`) to protect resources
where we _know_ that the resource is alive (for example, in device private data,
because it's teared down by driver-core on unbind).
However, there are cases where we can still have the `Revocable` pattern, for
example when a class device is exposed to userspace and the bus device is
unbinding. This is a pretty recurring pattern too, and cannot be handled via
lifetime alone. This is what motivates the C revocable series.
Best,
Gary
^ permalink raw reply [flat|nested] 22+ messages in thread
end of thread, other threads:[~2026-09-28 11:32 UTC | newest]
Thread overview: 22+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 15:51 [PATCH hazptr 0/4] Hazard pointer updates Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 2/4] compiler.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 3/4] Documentation: RCU: Refer to ptr_eq() Mathieu Desnoyers
2026-09-27 15:51 ` [PATCH hazptr 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
2026-09-27 16:40 ` Boqun Feng
2026-09-27 17:15 ` Mathieu Desnoyers
2026-09-27 17:24 ` Boqun Feng
2026-09-27 17:36 ` Mathieu Desnoyers
2026-09-27 18:16 ` Boqun Feng
2026-09-27 17:26 ` Boqun Feng
2026-09-27 22:39 ` Gary Guo
2026-09-28 9:12 ` Boqun Feng
2026-09-28 11:32 ` Gary Guo
2026-09-28 9:27 ` Kunwu Chan
2026-09-27 16:07 ` [PATCH hazptr 0/4] Hazard pointer updates Bradley Morgan
2026-09-27 16:27 ` Mathieu Desnoyers
2026-09-27 16:33 ` Bradley Morgan
2026-09-27 16:45 ` Mathieu Desnoyers
2026-09-27 16:15 ` Boqun Feng
2026-09-27 16:20 ` Mathieu Desnoyers
2026-09-27 16:22 ` Bradley Morgan
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®