mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH hazptr v2 0/4] Hazard pointer updates
@ 2026-10-09 18:21 Mathieu Desnoyers
  2026-10-09 18:21 ` [PATCH hazptr v2 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
                   ` (3 more replies)
  0 siblings, 4 replies; 9+ messages in thread
From: Mathieu Desnoyers @ 2026-10-09 18:21 UTC (permalink / raw)
  To: Paul E . McKenney
  Cc: linux-kernel, Mathieu Desnoyers, Boqun Feng, Bradley Morgan,
	Gary Guo, rcu, lkmm, Lian Wang, Kunwu Chan

Hi Paul,

This series applies on top of "hazptr: handle NULL address in
hazptr_detach" you have in your rcu dev tree.

This first patch addresses a race identified by Boqun Feng in the
two-phase wildcard scheme. (no change since the last series)

Patches 2-3 are prerequisites for using ptr_eq() in the 4th patch.
Those were discussed at length in a prior version of hazard pointer
patches.

Since the last series, I have modified patch 2 to take into account
feedback from Linus: There is now a x86-specific implementation of
ptr_eq, an asm-generic fallback, and both sit in newly introduced
"ptreq.h" headers.  The documentation is moved to linux/ptreq.h, which
includes asm/ptreq.h. This reuses all the usual mechanisms for
arch-specific implementation with asm-generic fallback.

Patch 4 introduces a "try acquire" helper to allow the fast path
to not rely on wildcards, while keeping the wildcard forward
progress guarantees in the acquire slow path, used on fast path
failure.

Thanks,

Mathieu

Mathieu Desnoyers (4):
  hazptr: Fix two-phase hazptr_synchronize race with detach
  ptreq.h: Introduce ptr_eq() to preserve address dependency
  Documentation: RCU: Refer to ptr_eq()
  hazptr: Introduce "try acquire" fast path, fallback to overflow list

Cc: Paul E. McKenney <paulmck@kernel.org>
Cc: Boqun Feng <boqun@kernel.org>
Cc: Bradley Morgan <brads@mainlining.org>
Cc: Gary Guo <gary@garyguo.net>
Cc: <rcu@vger.kernel.org>
Cc: <lkmm@lists.linux.dev>
Cc: Lian Wang <lianux.mm@gmail.com>
Cc: Kunwu Chan <kunwu.chan@gmail.com>

 Documentation/RCU/rcu_dereference.rst |  38 +++++++-
 arch/x86/include/asm/ptreq.h          |  16 +++
 include/asm-generic/Kbuild            |   1 +
 include/asm-generic/ptreq.h           |  16 +++
 include/linux/hazptr.h                |  48 +++++----
 include/linux/ptreq.h                 |  63 ++++++++++++
 kernel/hazptr.c                       | 135 +++++++++++++++-----------
 7 files changed, 237 insertions(+), 80 deletions(-)
 create mode 100644 arch/x86/include/asm/ptreq.h
 create mode 100644 include/asm-generic/ptreq.h
 create mode 100644 include/linux/ptreq.h

-- 
2.43.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH hazptr v2 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach
  2026-10-09 18:21 [PATCH hazptr v2 0/4] Hazard pointer updates Mathieu Desnoyers
@ 2026-10-09 18:21 ` Mathieu Desnoyers
  2026-10-09 18:21 ` [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 9+ messages in thread
From: Mathieu Desnoyers @ 2026-10-09 18:21 UTC (permalink / raw)
  To: Paul E . McKenney
  Cc: linux-kernel, Mathieu Desnoyers, Boqun Feng, Bradley Morgan,
	Gary Guo, rcu, lkmm, Lian Wang, Kunwu Chan

Boqun Feng pointed out that a detach happening concurrently with
hazptr_synchronize can miss a slot. Indeed, scanning the per-CPU
slots needs to be done *before* scanning the overflow lists. Fix the
implementation accordingly.

Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Reported-by: Boqun Feng <boqun@kernel.org>
Cc: Paul E. McKenney <paulmck@kernel.org>
Cc: Boqun Feng <boqun@kernel.org>
Cc: Bradley Morgan <brads@mainlining.org>
Cc: Gary Guo <gary@garyguo.net>
Cc: <rcu@vger.kernel.org>
Cc: <lkmm@lists.linux.dev>
Cc: Lian Wang <lianux.mm@gmail.com>
Cc: Kunwu Chan <kunwu.chan@gmail.com>
---
Changes since v0:
- Load the new overflow list phase word in "hazptr_chain_backup_slot".
- Rename the lock to hazptr_phase_lock to make it clear that it
  protects both the wildcard and the overflow list phases.
---
 kernel/hazptr.c | 50 +++++++++++++++++++++++++++++++++++++++----------
 1 file changed, 40 insertions(+), 10 deletions(-)

diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index d3d1050d92cf..13faa5ba7677 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -13,6 +13,8 @@
 #include <linux/list.h>
 #include <linux/export.h>
 
+static DEFINE_MUTEX(hazptr_phase_lock);	/* Protect the wildcard and list phase flip. */
+
 /*
  * The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
  * hazptr_synchronize forward progress even with a steady stream of readers.
@@ -20,10 +22,12 @@
  * This also affects the overflow list selection: the current list used by
  * readers is array[(unsigned long) hazptr_wildcard - 1].
  */
-static DEFINE_MUTEX(hazptr_wildcard_lock);	/* Protect the wildcard flip. */
 void *hazptr_wildcard = (void *) 1UL;
 EXPORT_SYMBOL_GPL(hazptr_wildcard);
 
+/* The current overflow list phase. */
+static unsigned int hazptr_overflow_list_phase;
+
 struct hazptr_overflow_list {
 	raw_spinlock_t lock;		/* Lock protecting overflow list and list generation. */
 	struct hlist_head head;		/* Overflow list head. */
@@ -53,6 +57,12 @@ void *flip_wildcard(void *wildcard)
 	return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
 }
 
+static
+unsigned int flip_list_phase(unsigned int phase)
+{
+	return 1 - phase;
+}
+
 static
 bool is_wildcard(void *addr)
 {
@@ -178,15 +188,12 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
 }
 
 static
-void hazptr_scan_period(void *addr, void *scan_wildcard)
+void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
 {
-	unsigned int scan_idx = (unsigned long) scan_wildcard - 1;
 	int cpu;
 
 	/* Scan all CPUs slots. */
 	for_each_possible_cpu(cpu) {
-		struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
-
 		/*
 		 * Scan CPU slots.
 		 * Forward progress against recurring wildcards is guaranteed
@@ -199,6 +206,17 @@ void hazptr_scan_period(void *addr, void *scan_wildcard)
 		 * to acquire that same hazard pointer value.
 		 */
 		hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
+	}
+}
+
+static
+void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
+{
+	int cpu;
+
+	/* Scan all CPUs overflow lists. */
+	for_each_possible_cpu(cpu) {
+		struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
 
 		/*
 		 * Scan backup slots in percpu overflow lists.
@@ -218,6 +236,7 @@ void hazptr_scan_period(void *addr, void *scan_wildcard)
  */
 void hazptr_synchronize(void *addr)
 {
+	unsigned int scan_list_phase;
 	void *scan_wildcard;
 
 	/*
@@ -235,18 +254,29 @@ void hazptr_synchronize(void *addr)
 	/* Memory ordering: Store A before Load B. */
 	smp_mb();
 
-	guard(mutex)(&hazptr_wildcard_lock);
+	guard(mutex)(&hazptr_phase_lock);
+
+	/* Scan per-CPU slots. */
 	scan_wildcard = flip_wildcard(hazptr_wildcard);
-	hazptr_scan_period(addr, scan_wildcard);
-	WRITE_ONCE(hazptr_wildcard, scan_wildcard);	/* Flip the current wildcard. */
-	hazptr_scan_period(addr, flip_wildcard(scan_wildcard));
+	hazptr_scan_cpu_slots_period(addr, scan_wildcard);
+	WRITE_ONCE(hazptr_wildcard, scan_wildcard);			/* Flip the current wildcard. */
+	hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
+
+	/*
+	 * Scan overflow lists *after* scanning per-CPU slots. See
+	 * hazptr_promote_to_backup_slot() for scan ordering requirement.
+	 */
+	scan_list_phase = flip_list_phase(hazptr_overflow_list_phase);
+	hazptr_scan_overflow_list_period(addr, scan_list_phase);
+	WRITE_ONCE(hazptr_overflow_list_phase, scan_list_phase);	/* Flip the current list phase. */
+	hazptr_scan_overflow_list_period(addr, flip_list_phase(scan_list_phase));
 }
 EXPORT_SYMBOL_GPL(hazptr_synchronize);
 
 struct hazptr_slot *hazptr_chain_backup_slot(struct hazptr_ctx *ctx)
 {
 	struct hazptr_overflow_list_flip *overflow_list_flip = this_cpu_ptr(&percpu_overflow_list_flip);
-	unsigned int list_idx = (unsigned long) READ_ONCE(hazptr_wildcard) - 1;
+	unsigned int list_idx = READ_ONCE(hazptr_overflow_list_phase);
 	struct hazptr_overflow_list *overflow_list = &overflow_list_flip->array[list_idx];
 	struct hazptr_slot *slot = &ctx->backup_slot.slot;
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency
  2026-10-09 18:21 [PATCH hazptr v2 0/4] Hazard pointer updates Mathieu Desnoyers
  2026-10-09 18:21 ` [PATCH hazptr v2 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
@ 2026-10-09 18:21 ` Mathieu Desnoyers
  2026-10-09 19:03   ` Boqun Feng
  2026-10-09 21:55   ` Gary Guo
  2026-10-09 18:21 ` [PATCH hazptr v2 3/4] Documentation: RCU: Refer to ptr_eq() Mathieu Desnoyers
  2026-10-09 18:21 ` [PATCH hazptr v2 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
  3 siblings, 2 replies; 9+ messages in thread
From: Mathieu Desnoyers @ 2026-10-09 18:21 UTC (permalink / raw)
  To: Paul E . McKenney
  Cc: linux-kernel, Mathieu Desnoyers, Linus Torvalds, Boqun Feng,
	Greg Kroah-Hartman, Sebastian Andrzej Siewior, Will Deacon,
	Peter Zijlstra, Alan Stern, John Stultz, Frederic Weisbecker,
	Joel Fernandes, Josh Triplett, Uladzislau Rezki, Steven Rostedt,
	Lai Jiangshan, Zqiang, Ingo Molnar, Waiman Long, Mark Rutland,
	Thomas Gleixner, Vlastimil Babka, maged.michael, Mateusz Guzik,
	Gary Guo, rcu, linux-mm, lkmm, Nikita Popov, llvm, Lian Wang,
	Kunwu Chan

Compiler CSE and SSA GVN optimizations can cause the address dependency
of addresses returned by rcu_dereference to be lost when comparing those
pointers with either constants or previously loaded pointers.

Introduce ptr_eq() to compare two addresses while preserving the address
dependencies for later use of the address. It should be used when
comparing an address returned by rcu_dereference().

This is needed to prevent the compiler CSE and SSA GVN optimizations
from using @a (or @b) in places where the source refers to @b (or @a)
based on the fact that after the comparison, the two are known to be
equal, which does not preserve address dependencies and allows the
following misordering speculations:

- If @b is a constant, the compiler can issue the loads which depend
  on @a before loading @a.
- If @b is a register populated by a prior load, weakly-ordered
  CPUs can speculate loads which depend on @a before loading @a.

The same logic applies with @a and @b swapped.

The header "linux/ptreq.h" is the expected include target. It
contains the documentation of the ptr_eq() API.

An architecture specific implementation of the pointer comparison
can be implemented by each architecture as asm/ptreq.h. The x86
implementation is provided initially. If no implementation header is
present for the architecture, an arch-agnostic fallback based on
OPTIMIZER_HIDE_VAR is included from asm-generic.

Suggested-by: Linus Torvalds <torvalds@linux-foundation.org>
Suggested-by: Boqun Feng <boqun.feng@gmail.com>
Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Cc: "Paul E. McKenney" <paulmck@kernel.org>
Cc: Will Deacon <will@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Alan Stern <stern@rowland.harvard.edu>
Cc: John Stultz <jstultz@google.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Frederic Weisbecker <frederic@kernel.org>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Josh Triplett <josh@joshtriplett.org>
Cc: Uladzislau Rezki <urezki@gmail.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Lai Jiangshan <jiangshanlai@gmail.com>
Cc: Zqiang <qiang.zhang1211@gmail.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Waiman Long <longman@redhat.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Vlastimil Babka <vbabka@suse.cz>
Cc: maged.michael@gmail.com
Cc: Mateusz Guzik <mjguzik@gmail.com>
Cc: Gary Guo <gary@garyguo.net>
Cc: rcu@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: lkmm@lists.linux.dev
Cc: Nikita Popov <github@npopov.com>
Cc: llvm@lists.linux.dev
Cc: Lian Wang <lianux.mm@gmail.com>
Cc: Kunwu Chan <kunwu.chan@gmail.com>
---
Changes since v1:
- Move to linux/ptreq.h, implement x86-specific comparison
  and asm-generic header (fallback).

Changes since v0:
- Include feedback from Alan Stern.
---
 arch/x86/include/asm/ptreq.h | 16 +++++++++
 include/asm-generic/Kbuild   |  1 +
 include/asm-generic/ptreq.h  | 16 +++++++++
 include/linux/ptreq.h        | 63 ++++++++++++++++++++++++++++++++++++
 4 files changed, 96 insertions(+)
 create mode 100644 arch/x86/include/asm/ptreq.h
 create mode 100644 include/asm-generic/ptreq.h
 create mode 100644 include/linux/ptreq.h

diff --git a/arch/x86/include/asm/ptreq.h b/arch/x86/include/asm/ptreq.h
new file mode 100644
index 000000000000..593b45f80451
--- /dev/null
+++ b/arch/x86/include/asm/ptreq.h
@@ -0,0 +1,16 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _ASM_X86_PTREQ_H
+#define _ASM_X86_PTREQ_H
+
+#include <linux/types.h>
+#include <asm/asm.h>
+
+static __always_inline
+bool ptr_eq(const volatile void *a, const volatile void *b)
+{
+	bool ret;
+	asm(__ASM_SIZE(cmp) " %1,%2" : "=@ccz" (ret) : "r" (a), "r" (b));
+	return ret;
+}
+
+#endif	/* _ASM_X86_PTREQ_H */
diff --git a/include/asm-generic/Kbuild b/include/asm-generic/Kbuild
index 2bc00c67dc54..e1d95dea09d5 100644
--- a/include/asm-generic/Kbuild
+++ b/include/asm-generic/Kbuild
@@ -47,6 +47,7 @@ mandatory-y += percpu.h
 mandatory-y += percpu_types.h
 mandatory-y += pgalloc.h
 mandatory-y += preempt.h
+mandatory-y += ptreq.h
 mandatory-y += rqspinlock.h
 mandatory-y += runtime-const.h
 mandatory-y += rwonce.h
diff --git a/include/asm-generic/ptreq.h b/include/asm-generic/ptreq.h
new file mode 100644
index 000000000000..db5a88628b1e
--- /dev/null
+++ b/include/asm-generic/ptreq.h
@@ -0,0 +1,16 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef __ASM_GENERIC_PTREQ_H
+#define __ASM_GENERIC_PTREQ_H
+
+#include <linux/types.h>
+#include <linux/compiler.h>
+
+static __always_inline
+bool ptr_eq(const volatile void *a, const volatile void *b)
+{
+	OPTIMIZER_HIDE_VAR(a);
+	OPTIMIZER_HIDE_VAR(b);
+	return a == b;
+}
+
+#endif	/* __ASM_GENERIC_PTREQ_H */
diff --git a/include/linux/ptreq.h b/include/linux/ptreq.h
new file mode 100644
index 000000000000..5499f65da30a
--- /dev/null
+++ b/include/linux/ptreq.h
@@ -0,0 +1,63 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_PTREQ_H
+#define _LINUX_PTREQ_H
+
+/*
+ * Compare two addresses while preserving the address dependencies for
+ * later use of the address. It should be used when comparing an address
+ * returned by rcu_dereference().
+ *
+ * This is needed to prevent the compiler CSE and SSA GVN optimizations
+ * from using @a (or @b) in places where the source refers to @b (or @a)
+ * based on the fact that after the comparison, the two are known to be
+ * equal, which does not preserve address dependencies and allows the
+ * following misordering speculations:
+ *
+ * - If @b is a constant, the compiler can issue the loads which depend
+ *   on @a before loading @a.
+ * - If @b is a register populated by a prior load, weakly-ordered
+ *   CPUs can speculate loads which depend on @a before loading @a.
+ *
+ * The same logic applies with @a and @b swapped.
+ *
+ * Return value: true if pointers are equal, false otherwise.
+ *
+ * The compiler barrier() is ineffective at fixing this issue. It does
+ * not prevent the compiler CSE from losing the address dependency:
+ *
+ * int fct_2_volatile_barriers(void)
+ * {
+ *     int *a, *b;
+ *
+ *     do {
+ *         a = READ_ONCE(p);
+ *         asm volatile ("" : : : "memory");
+ *         b = READ_ONCE(p);
+ *     } while (a != b);
+ *     asm volatile ("" : : : "memory");  <-- barrier()
+ *     return *b;
+ * }
+ *
+ * With gcc 14.2 (arm64):
+ *
+ * fct_2_volatile_barriers:
+ *         adrp    x0, .LANCHOR0
+ *         add     x0, x0, :lo12:.LANCHOR0
+ * .L2:
+ *         ldr     x1, [x0]  <-- x1 populated by first load.
+ *         ldr     x2, [x0]
+ *         cmp     x1, x2
+ *         bne     .L2
+ *         ldr     w0, [x1]  <-- x1 is used for access which should depend on b.
+ *         ret
+ *
+ * On weakly-ordered architectures, this lets CPU speculation use the
+ * result from the first load to speculate "ldr w0, [x1]" before
+ * "ldr x2, [x0]".
+ * Based on the RCU documentation, the control dependency does not
+ * prevent the CPU from speculating loads.
+ */
+
+#include <asm/ptreq.h>
+
+#endif	/* _LINUX_PTREQ_H */
-- 
2.43.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH hazptr v2 3/4] Documentation: RCU: Refer to ptr_eq()
  2026-10-09 18:21 [PATCH hazptr v2 0/4] Hazard pointer updates Mathieu Desnoyers
  2026-10-09 18:21 ` [PATCH hazptr v2 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
  2026-10-09 18:21 ` [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
@ 2026-10-09 18:21 ` Mathieu Desnoyers
  2026-10-09 18:21 ` [PATCH hazptr v2 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
  3 siblings, 0 replies; 9+ messages in thread
From: Mathieu Desnoyers @ 2026-10-09 18:21 UTC (permalink / raw)
  To: Paul E . McKenney
  Cc: linux-kernel, Mathieu Desnoyers, Alan Stern, Joel Fernandes,
	Greg Kroah-Hartman, Sebastian Andrzej Siewior, Will Deacon,
	Peter Zijlstra, Boqun Feng, John Stultz, Linus Torvalds,
	Frederic Weisbecker, Josh Triplett, Uladzislau Rezki,
	Steven Rostedt, Lai Jiangshan, Zqiang, Ingo Molnar, Waiman Long,
	Mark Rutland, Thomas Gleixner, Vlastimil Babka, maged.michael,
	Mateusz Guzik, Gary Guo, rcu, linux-mm, lkmm, Nikita Popov, llvm,
	Lian Wang, Kunwu Chan

Refer to ptr_eq() in the rcu_dereference() documentation.

ptr_eq() is a mechanism that preserves address dependencies when
comparing pointers, and should be favored when comparing a pointer
obtained from rcu_dereference() against another pointer.

Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Acked-by: Alan Stern <stern@rowland.harvard.edu>
Acked-by: Paul E. McKenney <paulmck@kernel.org>
Reviewed-by: Joel Fernandes (Google) <joel@joelfernandes.org>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
Cc: "Paul E. McKenney" <paulmck@kernel.org>
Cc: Will Deacon <will@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Alan Stern <stern@rowland.harvard.edu>
Cc: John Stultz <jstultz@google.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Boqun Feng <boqun.feng@gmail.com>
Cc: Frederic Weisbecker <frederic@kernel.org>
Cc: Joel Fernandes <joel@joelfernandes.org>
Cc: Josh Triplett <josh@joshtriplett.org>
Cc: Uladzislau Rezki <urezki@gmail.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Lai Jiangshan <jiangshanlai@gmail.com>
Cc: Zqiang <qiang.zhang1211@gmail.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Waiman Long <longman@redhat.com>
Cc: Mark Rutland <mark.rutland@arm.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Vlastimil Babka <vbabka@suse.cz>
Cc: maged.michael@gmail.com
Cc: Mateusz Guzik <mjguzik@gmail.com>
Cc: Gary Guo <gary@garyguo.net>
Cc: rcu@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: lkmm@lists.linux.dev
Cc: Nikita Popov <github@npopov.com>
Cc: llvm@lists.linux.dev
Cc: Lian Wang <lianux.mm@gmail.com>
Cc: Kunwu Chan <kunwu.chan@gmail.com>
---
Changes since v1:
- Include feedback from Paul E. McKenney.

Changes since v0:
- Include feedback from Alan Stern.
---
 Documentation/RCU/rcu_dereference.rst | 38 +++++++++++++++++++++++----
 1 file changed, 33 insertions(+), 5 deletions(-)

diff --git a/Documentation/RCU/rcu_dereference.rst b/Documentation/RCU/rcu_dereference.rst
index 5bc3785ebfc2..85b5f79ebf14 100644
--- a/Documentation/RCU/rcu_dereference.rst
+++ b/Documentation/RCU/rcu_dereference.rst
@@ -104,11 +104,12 @@ readers working properly:
 	after such branches, but can speculate loads, which can again
 	result in misordering bugs.
 
--	Be very careful about comparing pointers obtained from
-	rcu_dereference() against non-NULL values.  As Linus Torvalds
-	explained, if the two pointers are equal, the compiler could
-	substitute the pointer you are comparing against for the pointer
-	obtained from rcu_dereference().  For example::
+-	Use operations that preserve address dependencies (such as
+	"ptr_eq()") to compare pointers obtained from rcu_dereference()
+	against non-NULL pointers. As Linus Torvalds explained, if the
+	two pointers are equal, the compiler could substitute the
+	pointer you are comparing against for the pointer obtained from
+	rcu_dereference().  For example::
 
 		p = rcu_dereference(gp);
 		if (p == &default_struct)
@@ -125,6 +126,29 @@ readers working properly:
 	On ARM and Power hardware, the load from "default_struct.a"
 	can now be speculated, such that it might happen before the
 	rcu_dereference().  This could result in bugs due to misordering.
+	Performing the comparison with "ptr_eq()" ensures the compiler
+	does not perform such transformation.
+
+	If the comparison is against another pointer, the compiler is
+	allowed to use either pointer for the following accesses, which
+	loses the address dependency and allows weakly-ordered
+	architectures such as ARM and PowerPC to speculate the
+	address-dependent load before rcu_dereference().  For example::
+
+		p1 = READ_ONCE(gp);
+		p2 = rcu_dereference(gp);
+		if (p1 == p2)  /* BUGGY!!! */
+			do_default(p2->a);
+
+	The compiler can use p1->a rather than p2->a, destroying the
+	address dependency.  Performing the comparison with "ptr_eq()"
+	ensures the compiler preserves the address dependencies.
+	Corrected code::
+
+		p1 = READ_ONCE(gp);
+		p2 = rcu_dereference(gp);
+		if (ptr_eq(p1, p2))
+			do_default(p2->a);
 
 	However, comparisons are OK in the following cases:
 
@@ -204,6 +228,10 @@ readers working properly:
 		comparison will provide exactly the information that the
 		compiler needs to deduce the value of the pointer.
 
+	When in doubt, use operations that preserve address dependencies
+	(such as "ptr_eq()") to compare pointers obtained from
+	rcu_dereference() against non-NULL pointers.
+
 -	Disable any value-speculation optimizations that your compiler
 	might provide, especially if you are making use of feedback-based
 	optimizations that take data collected from prior runs.  Such
-- 
2.43.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH hazptr v2 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
  2026-10-09 18:21 [PATCH hazptr v2 0/4] Hazard pointer updates Mathieu Desnoyers
                   ` (2 preceding siblings ...)
  2026-10-09 18:21 ` [PATCH hazptr v2 3/4] Documentation: RCU: Refer to ptr_eq() Mathieu Desnoyers
@ 2026-10-09 18:21 ` Mathieu Desnoyers
  2026-10-10 19:39   ` Bradley Morgan
  3 siblings, 1 reply; 9+ messages in thread
From: Mathieu Desnoyers @ 2026-10-09 18:21 UTC (permalink / raw)
  To: Paul E . McKenney
  Cc: linux-kernel, Mathieu Desnoyers, Gary Guo, Boqun Feng,
	Bradley Morgan, rcu, lkmm, Lian Wang, Kunwu Chan

Introduce a "try acquire" hazard pointer fast path, which performs an
early load of the address to store it into the hazard pointer slot, and
then re-loads that address after a barrier to check whether it has
changed meanwhile.

On comparison failure, rather than re-try, guarantee forward progress by
falling back to the __hazptr_acquire slow path on failure.

The acquire slow path attempts a try-acquire for any available per-CPU
slot. If that fails, it chains the backup slot into the overflow list,
therefore guaranteeing forward progress for both hazard pointer
read-side and synchronize:

- Readers set the wildcard, and then proceed to set the more
  specific address to replace the wildcard.

- One synchronize alternates between two overflow list periods,
  scanning each one while readers are added to the other period,
  thus preventing a steady flow of readers from preventing
  synchronize forward progress.

With this change, the scan on per-CPU slots don't need to expect a
wildcard anymore, because none can be produced by readers. Wildcards are
only expected within overflow lists.

Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Suggested-by: Gary Guo <gary@garyguo.net>
Link: https://lore.kernel.org/lkmm/DLK9MW1N86KD.2GRWUAD2NX4FU@garyguo.net/
Reviewed-by: Gary Guo <gary@garyguo.net>
Cc: Paul E. McKenney <paulmck@kernel.org>
Cc: Boqun Feng <boqun@kernel.org>
Cc: Bradley Morgan <brads@mainlining.org>
Cc: Gary Guo <gary@garyguo.net>
Cc: <rcu@vger.kernel.org>
Cc: <lkmm@lists.linux.dev>
Cc: Lian Wang <lianux.mm@gmail.com>
Cc: Kunwu Chan <kunwu.chan@gmail.com>
---
 include/linux/hazptr.h |  48 +++++++++++--------
 kernel/hazptr.c        | 103 ++++++++++++++++++++---------------------
 2 files changed, 77 insertions(+), 74 deletions(-)

diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
index d1670121947a..df2aca9aa26b 100644
--- a/include/linux/hazptr.h
+++ b/include/linux/hazptr.h
@@ -25,13 +25,11 @@
 #include <linux/types.h>
 #include <linux/cleanup.h>
 #include <linux/sched.h>
+#include <linux/ptreq.h>
 
 /* 4 slots (each sizeof(hazptr_slot_item)) fit in a single 64-byte cache line. */
 #define NR_HAZPTR_PERCPU_SLOTS	4
 
-/* The current hazard pointer wildcard. */
-extern void *hazptr_wildcard;
-
 /*
  * Hazard pointer slot.
  */
@@ -190,6 +188,31 @@ void hazptr_note_context_switch(void)
 	}
 }
 
+/* Try hazard pointer protection. */
+static inline
+void *__hazptr_try_acquire(struct hazptr_ctx *ctx, void * const *addr_p, struct hazptr_slot *slot)
+{
+	void *early_addr, *addr;
+
+	if (unlikely(slot->addr))
+		return NULL;
+	early_addr = READ_ONCE(*addr_p);	/* Early load. */
+	WRITE_ONCE(slot->addr, early_addr);	/* Store B */
+	/* Memory ordering: Store B before Load A. */
+	smp_mb();
+	addr = READ_ONCE(*addr_p);		/* Load A */
+	/*
+	 * Validate that address did not change between Early load and Load A.
+	 * Use ptr_eq() to make sure that result from Load A is returned to the
+	 * caller to preserve address dependency.
+	 */
+	if (unlikely(!ptr_eq(addr, early_addr))) {
+		WRITE_ONCE(slot->addr, NULL);
+		return NULL;
+	}
+	return addr;
+}
+
 /**
  * hazptr_acquire - Load pointer at address and protect with hazard pointer.
  *
@@ -245,24 +268,9 @@ void *hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
 	ctx->acquire_cpu = smp_processor_id();
 	ctx->acquire_caller = _THIS_IP_;
 #endif
-	if (unlikely(slot->addr))
+	addr = __hazptr_try_acquire(ctx, addr_p, slot);
+	if (unlikely(!addr))
 		return __hazptr_acquire(ctx, addr_p);
-	WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard));	/* Store B */
-
-	/* Memory ordering: Store B before Load A. */
-	smp_mb();
-
-	/*
-	 * Load @addr_p after storing wildcard to the hazard pointer slot.
-	 */
-	addr = READ_ONCE(*addr_p);	/* Load A */
-
-	/*
-	 * We don't care about ordering of Store C. It will simply
-	 * replace the wildcard by a more specific address. If addr is
-	 * NULL, we simply store NULL into the slot.
-	 */
-	WRITE_ONCE(slot->addr, addr);	/* Store C */
 	slot_item->ctx.ctx = ctx;
 	ctx->slot = slot;
 	return addr;
diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index 13faa5ba7677..3ca73b56a5c2 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -13,17 +13,9 @@
 #include <linux/list.h>
 #include <linux/export.h>
 
-static DEFINE_MUTEX(hazptr_phase_lock);	/* Protect the wildcard and list phase flip. */
+#define HAZPTR_WILDCARD	((void *) 1UL)
 
-/*
- * The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
- * hazptr_synchronize forward progress even with a steady stream of readers.
- * This wildcard value is used by acquire to temporarily tag the per-CPU slots.
- * This also affects the overflow list selection: the current list used by
- * readers is array[(unsigned long) hazptr_wildcard - 1].
- */
-void *hazptr_wildcard = (void *) 1UL;
-EXPORT_SYMBOL_GPL(hazptr_wildcard);
+static DEFINE_MUTEX(hazptr_phase_lock);	/* Protect the list phase flip. */
 
 /* The current overflow list phase. */
 static unsigned int hazptr_overflow_list_phase;
@@ -41,6 +33,10 @@ struct hazptr_overflow_list {
  * successively iterates on both lists. Therefore, only list removals
  * can cause the iteration to retry, and the number of removals is
  * limited to the number of list elements.
+ *
+ * Due to the overflow list raw spin lock, the hazard pointer readers are
+ * blocking, starvation-free with bounded waiting, assuming bounded critical
+ * sections and no NMI or virtualization-induced holder preemption.
  */
 struct hazptr_overflow_list_flip {
 	struct hazptr_overflow_list array[2];
@@ -51,26 +47,12 @@ static DEFINE_PER_CPU(struct hazptr_overflow_list_flip, percpu_overflow_list_fli
 DEFINE_PER_CPU(struct hazptr_percpu_slots, hazptr_percpu_slots);
 EXPORT_PER_CPU_SYMBOL_GPL(hazptr_percpu_slots);
 
-static
-void *flip_wildcard(void *wildcard)
-{
-	return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
-}
-
 static
 unsigned int flip_list_phase(unsigned int phase)
 {
 	return 1 - phase;
 }
 
-static
-bool is_wildcard(void *addr)
-{
-	if ((unsigned long) addr == 1UL || (unsigned long) addr == 2UL)
-		return true;
-	return false;
-}
-
 static
 struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
 {
@@ -96,16 +78,36 @@ struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
  */
 void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
 {
-	struct hazptr_slot *slot = hazptr_get_free_percpu_slot(ctx);
+	struct hazptr_slot *slot;
 	void *addr;
 
 	/*
-	 * If all the per-CPU slots are already in use, fallback
-	 * to the backup slot.
+	 * In case we are called due to nested use of hazard pointers,
+	 * try a slot protection with per-CPU slots.
 	 */
-	if (unlikely(!slot))
-		slot = hazptr_chain_backup_slot(ctx);
-	WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard));	/* Store B */
+	slot = hazptr_get_free_percpu_slot(ctx);
+	if (likely(slot)) {
+		addr = __hazptr_try_acquire(ctx, addr_p, slot);
+		if (addr) {
+			ctx->slot = slot;
+			return addr;
+		}
+	}
+
+	/*
+	 * The backup slot overflow list guarantees forward progress of both
+	 * hazard pointer readers and synchronize:
+	 *
+	 * - Readers set the wildcard, and then proceed to set the more
+	 *   specific address to replace the wildcard.
+	 *
+	 * - One synchronize alternates between two overflow list periods,
+	 *   scanning each one while readers are added to the other period,
+	 *   thus preventing a steady flow of readers from preventing
+	 *   synchronize forward progress.
+	 */
+	slot = hazptr_chain_backup_slot(ctx);
+	WRITE_ONCE(slot->addr, HAZPTR_WILDCARD);	/* Store B */
 
 	/* Memory ordering: Store B before Load A. */
 	smp_mb();
@@ -121,20 +123,21 @@ void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
 	 * NULL, we simply store NULL into the slot.
 	 */
 	WRITE_ONCE(slot->addr, addr);	/* Store C */
+
 	ctx->slot = slot;
-	if (!addr && hazptr_slot_is_backup(ctx, slot))
+	if (!addr)
 		hazptr_unchain_backup_slot(ctx);
 	return addr;
 }
 EXPORT_SYMBOL_GPL(__hazptr_acquire);
 
 /*
- * Perform piecewise iteration on overflow list waiting until "addr" is
- * not present. Raw spinlock is released and taken between each list
- * item and busy loop iteration. The overflow list generation is checked
- * each time the lock is taken to validate that the list has not changed
- * before resuming iteration or busy wait. If the generation has
- * changed, retry the entire list traversal.
+ * Perform piecewise iteration on overflow list waiting until "addr" and
+ * wildcard are not present. Raw spinlock is released and taken between each
+ * list item and busy loop iteration. The overflow list generation is checked
+ * each time the lock is taken to validate that the list has not changed before
+ * resuming iteration or busy wait. If the generation has changed, retry the
+ * entire list traversal.
  */
 static
 void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list, void *addr)
@@ -147,13 +150,11 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
 retry:
 	snapshot_gen = overflow_list->gen;
 	hlist_for_each_entry(backup_slot, &overflow_list->head, overflow_node) {
-		/* Busy-wait if node is found. */
+		/* Busy-wait if addr or wildcard are found. */
 		for (;;) {
 			void *load_addr = smp_load_acquire(&backup_slot->slot.addr);	/* Load B */
 
-			/* We don't expect wildcards in overflow list. */
-			WARN_ON_ONCE(is_wildcard(load_addr));
-			if (load_addr != addr)
+			if (load_addr != addr && load_addr != HAZPTR_WILDCARD)
 				break;
 			raw_spin_unlock_irqrestore(&overflow_list->lock, flags);
 			cpu_relax();
@@ -174,7 +175,7 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
 }
 
 static
-void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
+void hazptr_synchronize_cpu_slots(int cpu, void *addr)
 {
 	struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
 	unsigned int idx;
@@ -182,13 +183,13 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
 	for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
 		struct hazptr_slot_item *item = &percpu_slots->items[idx];
 
-		/* Busy-wait if node is found. */
-		smp_cond_load_acquire(&item->slot.addr, VAL != addr && VAL != scan_wildcard); /* Load B */
+		/* Busy-wait if addr is found. */
+		smp_cond_load_acquire(&item->slot.addr, VAL != addr); /* Load B */
 	}
 }
 
 static
-void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
+void hazptr_scan_cpu_slots(void *addr)
 {
 	int cpu;
 
@@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
 	for_each_possible_cpu(cpu) {
 		/*
 		 * Scan CPU slots.
-		 * Forward progress against recurring wildcards is guaranteed
-		 * by scanning for one wildcard while new elements use the
-		 * other wildcard value (1UL vs 2UL).
 		 * Forward progress against recurring single hazard pointer
 		 * values is guaranteed by the fact that a hazard pointer
 		 * is not reclaimed nor reused until the scan for that hazard
 		 * pointer completes, which prevents a steady flow of readers
 		 * to acquire that same hazard pointer value.
 		 */
-		hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
+		hazptr_synchronize_cpu_slots(cpu, addr);
 	}
 }
 
@@ -237,7 +235,6 @@ void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
 void hazptr_synchronize(void *addr)
 {
 	unsigned int scan_list_phase;
-	void *scan_wildcard;
 
 	/*
 	 * Busy-wait should only be done from preemptible context.
@@ -251,16 +248,14 @@ void hazptr_synchronize(void *addr)
 	 */
 	if (!addr)
 		return;
+
 	/* Memory ordering: Store A before Load B. */
 	smp_mb();
 
 	guard(mutex)(&hazptr_phase_lock);
 
 	/* Scan per-CPU slots. */
-	scan_wildcard = flip_wildcard(hazptr_wildcard);
-	hazptr_scan_cpu_slots_period(addr, scan_wildcard);
-	WRITE_ONCE(hazptr_wildcard, scan_wildcard);			/* Flip the current wildcard. */
-	hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
+	hazptr_scan_cpu_slots(addr);
 
 	/*
 	 * Scan overflow lists *after* scanning per-CPU slots. See
-- 
2.43.0


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency
  2026-10-09 18:21 ` [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
@ 2026-10-09 19:03   ` Boqun Feng
  2026-10-09 19:07     ` Mathieu Desnoyers
  2026-10-09 21:55   ` Gary Guo
  1 sibling, 1 reply; 9+ messages in thread
From: Boqun Feng @ 2026-10-09 19:03 UTC (permalink / raw)
  To: Mathieu Desnoyers
  Cc: Paul E . McKenney, linux-kernel, Linus Torvalds, Boqun Feng,
	Greg Kroah-Hartman, Sebastian Andrzej Siewior, Will Deacon,
	Peter Zijlstra, Alan Stern, John Stultz, Frederic Weisbecker,
	Joel Fernandes, Josh Triplett, Uladzislau Rezki, Steven Rostedt,
	Lai Jiangshan, Zqiang, Ingo Molnar, Waiman Long, Mark Rutland,
	Thomas Gleixner, Vlastimil Babka, maged.michael, Mateusz Guzik,
	Gary Guo, rcu, linux-mm, lkmm, Nikita Popov, llvm, Lian Wang,
	Kunwu Chan

On Fri, Oct 09, 2026 at 02:21:33PM -0400, Mathieu Desnoyers wrote:
> Compiler CSE and SSA GVN optimizations can cause the address dependency
> of addresses returned by rcu_dereference to be lost when comparing those
> pointers with either constants or previously loaded pointers.
> 
> Introduce ptr_eq() to compare two addresses while preserving the address
> dependencies for later use of the address. It should be used when
> comparing an address returned by rcu_dereference().
> 
> This is needed to prevent the compiler CSE and SSA GVN optimizations
> from using @a (or @b) in places where the source refers to @b (or @a)
> based on the fact that after the comparison, the two are known to be
> equal, which does not preserve address dependencies and allows the
> following misordering speculations:
> 
> - If @b is a constant, the compiler can issue the loads which depend
>   on @a before loading @a.
> - If @b is a register populated by a prior load, weakly-ordered
>   CPUs can speculate loads which depend on @a before loading @a.
> 
> The same logic applies with @a and @b swapped.
> 
> The header "linux/ptreq.h" is the expected include target. It
> contains the documentation of the ptr_eq() API.
> 
> An architecture specific implementation of the pointer comparison
> can be implemented by each architecture as asm/ptreq.h. The x86
> implementation is provided initially. If no implementation header is
> present for the architecture, an arch-agnostic fallback based on
> OPTIMIZER_HIDE_VAR is included from asm-generic.
> 
> Suggested-by: Linus Torvalds <torvalds@linux-foundation.org>
> Suggested-by: Boqun Feng <boqun.feng@gmail.com>
> Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
> Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
> Cc: "Paul E. McKenney" <paulmck@kernel.org>
> Cc: Will Deacon <will@kernel.org>
> Cc: Peter Zijlstra <peterz@infradead.org>
> Cc: Boqun Feng <boqun.feng@gmail.com>
> Cc: Alan Stern <stern@rowland.harvard.edu>
> Cc: John Stultz <jstultz@google.com>
> Cc: Linus Torvalds <torvalds@linux-foundation.org>
> Cc: Boqun Feng <boqun.feng@gmail.com>
> Cc: Frederic Weisbecker <frederic@kernel.org>
> Cc: Joel Fernandes <joel@joelfernandes.org>
> Cc: Josh Triplett <josh@joshtriplett.org>
> Cc: Uladzislau Rezki <urezki@gmail.com>
> Cc: Steven Rostedt <rostedt@goodmis.org>
> Cc: Lai Jiangshan <jiangshanlai@gmail.com>
> Cc: Zqiang <qiang.zhang1211@gmail.com>
> Cc: Ingo Molnar <mingo@redhat.com>
> Cc: Waiman Long <longman@redhat.com>
> Cc: Mark Rutland <mark.rutland@arm.com>
> Cc: Thomas Gleixner <tglx@linutronix.de>
> Cc: Vlastimil Babka <vbabka@suse.cz>
> Cc: maged.michael@gmail.com
> Cc: Mateusz Guzik <mjguzik@gmail.com>
> Cc: Gary Guo <gary@garyguo.net>
> Cc: rcu@vger.kernel.org
> Cc: linux-mm@kvack.org
> Cc: lkmm@lists.linux.dev
> Cc: Nikita Popov <github@npopov.com>
> Cc: llvm@lists.linux.dev
> Cc: Lian Wang <lianux.mm@gmail.com>
> Cc: Kunwu Chan <kunwu.chan@gmail.com>
> ---
> Changes since v1:
> - Move to linux/ptreq.h, implement x86-specific comparison

I prefer the name ptr_eq.h over ptreq.h :)

With that name fix or there is a strong reason of ptreq.h:

Reviewed-by: Boqun Feng <boqun@kernel.org>

Regards,
Boqun

>   and asm-generic header (fallback).
> 
> Changes since v0:
> - Include feedback from Alan Stern.
> ---
>  arch/x86/include/asm/ptreq.h | 16 +++++++++
>  include/asm-generic/Kbuild   |  1 +
>  include/asm-generic/ptreq.h  | 16 +++++++++
>  include/linux/ptreq.h        | 63 ++++++++++++++++++++++++++++++++++++
>  4 files changed, 96 insertions(+)
>  create mode 100644 arch/x86/include/asm/ptreq.h
>  create mode 100644 include/asm-generic/ptreq.h
>  create mode 100644 include/linux/ptreq.h
> 
> diff --git a/arch/x86/include/asm/ptreq.h b/arch/x86/include/asm/ptreq.h
> new file mode 100644
> index 000000000000..593b45f80451
> --- /dev/null
> +++ b/arch/x86/include/asm/ptreq.h
> @@ -0,0 +1,16 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _ASM_X86_PTREQ_H
> +#define _ASM_X86_PTREQ_H
> +
> +#include <linux/types.h>
> +#include <asm/asm.h>
> +
> +static __always_inline
> +bool ptr_eq(const volatile void *a, const volatile void *b)
> +{
> +	bool ret;
> +	asm(__ASM_SIZE(cmp) " %1,%2" : "=@ccz" (ret) : "r" (a), "r" (b));
> +	return ret;
> +}
> +
> +#endif	/* _ASM_X86_PTREQ_H */
> diff --git a/include/asm-generic/Kbuild b/include/asm-generic/Kbuild
> index 2bc00c67dc54..e1d95dea09d5 100644
> --- a/include/asm-generic/Kbuild
> +++ b/include/asm-generic/Kbuild
> @@ -47,6 +47,7 @@ mandatory-y += percpu.h
>  mandatory-y += percpu_types.h
>  mandatory-y += pgalloc.h
>  mandatory-y += preempt.h
> +mandatory-y += ptreq.h
>  mandatory-y += rqspinlock.h
>  mandatory-y += runtime-const.h
>  mandatory-y += rwonce.h
> diff --git a/include/asm-generic/ptreq.h b/include/asm-generic/ptreq.h
> new file mode 100644
> index 000000000000..db5a88628b1e
> --- /dev/null
> +++ b/include/asm-generic/ptreq.h
> @@ -0,0 +1,16 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef __ASM_GENERIC_PTREQ_H
> +#define __ASM_GENERIC_PTREQ_H
> +
> +#include <linux/types.h>
> +#include <linux/compiler.h>
> +
> +static __always_inline
> +bool ptr_eq(const volatile void *a, const volatile void *b)
> +{
> +	OPTIMIZER_HIDE_VAR(a);
> +	OPTIMIZER_HIDE_VAR(b);
> +	return a == b;
> +}
> +
> +#endif	/* __ASM_GENERIC_PTREQ_H */
> diff --git a/include/linux/ptreq.h b/include/linux/ptreq.h
> new file mode 100644
> index 000000000000..5499f65da30a
> --- /dev/null
> +++ b/include/linux/ptreq.h
> @@ -0,0 +1,63 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _LINUX_PTREQ_H
> +#define _LINUX_PTREQ_H
> +
> +/*
> + * Compare two addresses while preserving the address dependencies for
> + * later use of the address. It should be used when comparing an address
> + * returned by rcu_dereference().
> + *
> + * This is needed to prevent the compiler CSE and SSA GVN optimizations
> + * from using @a (or @b) in places where the source refers to @b (or @a)
> + * based on the fact that after the comparison, the two are known to be
> + * equal, which does not preserve address dependencies and allows the
> + * following misordering speculations:
> + *
> + * - If @b is a constant, the compiler can issue the loads which depend
> + *   on @a before loading @a.
> + * - If @b is a register populated by a prior load, weakly-ordered
> + *   CPUs can speculate loads which depend on @a before loading @a.
> + *
> + * The same logic applies with @a and @b swapped.
> + *
> + * Return value: true if pointers are equal, false otherwise.
> + *
> + * The compiler barrier() is ineffective at fixing this issue. It does
> + * not prevent the compiler CSE from losing the address dependency:
> + *
> + * int fct_2_volatile_barriers(void)
> + * {
> + *     int *a, *b;
> + *
> + *     do {
> + *         a = READ_ONCE(p);
> + *         asm volatile ("" : : : "memory");
> + *         b = READ_ONCE(p);
> + *     } while (a != b);
> + *     asm volatile ("" : : : "memory");  <-- barrier()
> + *     return *b;
> + * }
> + *
> + * With gcc 14.2 (arm64):
> + *
> + * fct_2_volatile_barriers:
> + *         adrp    x0, .LANCHOR0
> + *         add     x0, x0, :lo12:.LANCHOR0
> + * .L2:
> + *         ldr     x1, [x0]  <-- x1 populated by first load.
> + *         ldr     x2, [x0]
> + *         cmp     x1, x2
> + *         bne     .L2
> + *         ldr     w0, [x1]  <-- x1 is used for access which should depend on b.
> + *         ret
> + *
> + * On weakly-ordered architectures, this lets CPU speculation use the
> + * result from the first load to speculate "ldr w0, [x1]" before
> + * "ldr x2, [x0]".
> + * Based on the RCU documentation, the control dependency does not
> + * prevent the CPU from speculating loads.
> + */
> +
> +#include <asm/ptreq.h>
> +
> +#endif	/* _LINUX_PTREQ_H */
> -- 
> 2.43.0
> 

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency
  2026-10-09 19:03   ` Boqun Feng
@ 2026-10-09 19:07     ` Mathieu Desnoyers
  0 siblings, 0 replies; 9+ messages in thread
From: Mathieu Desnoyers @ 2026-10-09 19:07 UTC (permalink / raw)
  To: Boqun Feng
  Cc: Paul E . McKenney, linux-kernel, Linus Torvalds, Boqun Feng,
	Greg Kroah-Hartman, Sebastian Andrzej Siewior, Will Deacon,
	Peter Zijlstra, Alan Stern, John Stultz, Frederic Weisbecker,
	Joel Fernandes, Josh Triplett, Uladzislau Rezki, Steven Rostedt,
	Lai Jiangshan, Zqiang, Ingo Molnar, Waiman Long, Mark Rutland,
	Thomas Gleixner, Vlastimil Babka, maged.michael, Mateusz Guzik,
	Gary Guo, rcu, linux-mm, lkmm, Nikita Popov, llvm, Lian Wang,
	Kunwu Chan

On 2026-10-09 15:03, Boqun Feng wrote:
> On Fri, Oct 09, 2026 at 02:21:33PM -0400, Mathieu Desnoyers wrote:
[...]
>>
>> An architecture specific implementation of the pointer comparison
>> can be implemented by each architecture as asm/ptreq.h. The x86
>> implementation is provided initially. If no implementation header is
>> present for the architecture, an arch-agnostic fallback based on
>> OPTIMIZER_HIDE_VAR is included from asm-generic.
>>
[...]
> 
> I prefer the name ptr_eq.h over ptreq.h :)
> 
> With that name fix or there is a strong reason of ptreq.h:
> 
> Reviewed-by: Boqun Feng <boqun@kernel.org>
I took rwonce.h (implementing READ_ONCE()/WRITE_ONCE()) as an
inspiration for the file name. I'm open to change it to whatever
people think best. I have no strong preference for either ptreq.h
or ptr_eq.h. I'll let people chime in before respinning an updated
series if a file name change is needed.

Thanks!

Mathieu

-- 
Mathieu Desnoyers
EfficiOS Inc.
https://www.efficios.com

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency
  2026-10-09 18:21 ` [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
  2026-10-09 19:03   ` Boqun Feng
@ 2026-10-09 21:55   ` Gary Guo
  1 sibling, 0 replies; 9+ messages in thread
From: Gary Guo @ 2026-10-09 21:55 UTC (permalink / raw)
  To: Mathieu Desnoyers, Paul E . McKenney
  Cc: linux-kernel, Linus Torvalds, Boqun Feng, Greg Kroah-Hartman,
	Sebastian Andrzej Siewior, Will Deacon, Peter Zijlstra,
	Alan Stern, John Stultz, Frederic Weisbecker, Joel Fernandes,
	Josh Triplett, Uladzislau Rezki, Steven Rostedt, Lai Jiangshan,
	Zqiang, Ingo Molnar, Waiman Long, Mark Rutland, Thomas Gleixner,
	Vlastimil Babka, maged.michael, Mateusz Guzik, Gary Guo, rcu,
	linux-mm, lkmm, Nikita Popov, llvm, Lian Wang, Kunwu Chan

On Fri Oct 9, 2026 at 7:21 PM BST, Mathieu Desnoyers wrote:
> Compiler CSE and SSA GVN optimizations can cause the address dependency
> of addresses returned by rcu_dereference to be lost when comparing those
> pointers with either constants or previously loaded pointers.
>
> Introduce ptr_eq() to compare two addresses while preserving the address
> dependencies for later use of the address. It should be used when
> comparing an address returned by rcu_dereference().
>
> This is needed to prevent the compiler CSE and SSA GVN optimizations
> from using @a (or @b) in places where the source refers to @b (or @a)
> based on the fact that after the comparison, the two are known to be
> equal, which does not preserve address dependencies and allows the
> following misordering speculations:
>
> - If @b is a constant, the compiler can issue the loads which depend
>   on @a before loading @a.
> - If @b is a register populated by a prior load, weakly-ordered
>   CPUs can speculate loads which depend on @a before loading @a.
>
> The same logic applies with @a and @b swapped.
>
> The header "linux/ptreq.h" is the expected include target. It
> contains the documentation of the ptr_eq() API.
>
> An architecture specific implementation of the pointer comparison
> can be implemented by each architecture as asm/ptreq.h. The x86
> implementation is provided initially. If no implementation header is
> present for the architecture, an arch-agnostic fallback based on
> OPTIMIZER_HIDE_VAR is included from asm-generic.
>
> Suggested-by: Linus Torvalds <torvalds@linux-foundation.org>
> Suggested-by: Boqun Feng <boqun.feng@gmail.com>
> Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
> Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
> Cc: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
> Cc: "Paul E. McKenney" <paulmck@kernel.org>
> Cc: Will Deacon <will@kernel.org>
> Cc: Peter Zijlstra <peterz@infradead.org>
> Cc: Boqun Feng <boqun.feng@gmail.com>
> Cc: Alan Stern <stern@rowland.harvard.edu>
> Cc: John Stultz <jstultz@google.com>
> Cc: Linus Torvalds <torvalds@linux-foundation.org>
> Cc: Boqun Feng <boqun.feng@gmail.com>
> Cc: Frederic Weisbecker <frederic@kernel.org>
> Cc: Joel Fernandes <joel@joelfernandes.org>
> Cc: Josh Triplett <josh@joshtriplett.org>
> Cc: Uladzislau Rezki <urezki@gmail.com>
> Cc: Steven Rostedt <rostedt@goodmis.org>
> Cc: Lai Jiangshan <jiangshanlai@gmail.com>
> Cc: Zqiang <qiang.zhang1211@gmail.com>
> Cc: Ingo Molnar <mingo@redhat.com>
> Cc: Waiman Long <longman@redhat.com>
> Cc: Mark Rutland <mark.rutland@arm.com>
> Cc: Thomas Gleixner <tglx@linutronix.de>
> Cc: Vlastimil Babka <vbabka@suse.cz>
> Cc: maged.michael@gmail.com
> Cc: Mateusz Guzik <mjguzik@gmail.com>
> Cc: Gary Guo <gary@garyguo.net>
> Cc: rcu@vger.kernel.org
> Cc: linux-mm@kvack.org
> Cc: lkmm@lists.linux.dev
> Cc: Nikita Popov <github@npopov.com>
> Cc: llvm@lists.linux.dev
> Cc: Lian Wang <lianux.mm@gmail.com>
> Cc: Kunwu Chan <kunwu.chan@gmail.com>
> ---
> Changes since v1:
> - Move to linux/ptreq.h, implement x86-specific comparison
>   and asm-generic header (fallback).
>
> Changes since v0:
> - Include feedback from Alan Stern.
> ---
>  arch/x86/include/asm/ptreq.h | 16 +++++++++
>  include/asm-generic/Kbuild   |  1 +
>  include/asm-generic/ptreq.h  | 16 +++++++++

I agree with Boqun that this should be ptr_eq. There is no need to save that
character :)

>  include/linux/ptreq.h        | 63 ++++++++++++++++++++++++++++++++++++
>  4 files changed, 96 insertions(+)
>  create mode 100644 arch/x86/include/asm/ptreq.h
>  create mode 100644 include/asm-generic/ptreq.h
>  create mode 100644 include/linux/ptreq.h
>
> diff --git a/arch/x86/include/asm/ptreq.h b/arch/x86/include/asm/ptreq.h
> new file mode 100644
> index 000000000000..593b45f80451
> --- /dev/null
> +++ b/arch/x86/include/asm/ptreq.h
> @@ -0,0 +1,16 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef _ASM_X86_PTREQ_H
> +#define _ASM_X86_PTREQ_H
> +
> +#include <linux/types.h>
> +#include <asm/asm.h>
> +
> +static __always_inline
> +bool ptr_eq(const volatile void *a, const volatile void *b)
> +{
> +	bool ret;
> +	asm(__ASM_SIZE(cmp) " %1,%2" : "=@ccz" (ret) : "r" (a), "r" (b));
> +	return ret;
> +}
> +
> +#endif	/* _ASM_X86_PTREQ_H */
> diff --git a/include/asm-generic/Kbuild b/include/asm-generic/Kbuild
> index 2bc00c67dc54..e1d95dea09d5 100644
> --- a/include/asm-generic/Kbuild
> +++ b/include/asm-generic/Kbuild
> @@ -47,6 +47,7 @@ mandatory-y += percpu.h
>  mandatory-y += percpu_types.h
>  mandatory-y += pgalloc.h
>  mandatory-y += preempt.h
> +mandatory-y += ptreq.h
>  mandatory-y += rqspinlock.h
>  mandatory-y += runtime-const.h
>  mandatory-y += rwonce.h
> diff --git a/include/asm-generic/ptreq.h b/include/asm-generic/ptreq.h
> new file mode 100644
> index 000000000000..db5a88628b1e
> --- /dev/null
> +++ b/include/asm-generic/ptreq.h
> @@ -0,0 +1,16 @@
> +/* SPDX-License-Identifier: GPL-2.0 */
> +#ifndef __ASM_GENERIC_PTREQ_H
> +#define __ASM_GENERIC_PTREQ_H
> +
> +#include <linux/types.h>
> +#include <linux/compiler.h>
> +
> +static __always_inline
> +bool ptr_eq(const volatile void *a, const volatile void *b)

I suppose this is `const volatile` to take advantage of implicit conversion that
adds cv-qualifiers to be added implicitly. Might worth a mention in somewhere
(commit message maybe) because I was confused for a second on why this needs
to be volatile.

> +{
> +	OPTIMIZER_HIDE_VAR(a);
> +	OPTIMIZER_HIDE_VAR(b);
> +	return a == b;
> +}

One thought on inline assembly. For RISC-V, branching with comparison is a
single instruction, but there's no instruction for obtaining the boolean for
a == b.

One option is to do

    static __always_inline
    bool ptr_eq_likely(const volatile void *a, const volatile void *b) {
        asm goto ("bne %0, %1, %l2" : : "r" (a), "r" (b) :: not_eq);
        return true;
    not_eq:
        return likely(false);
    }

and then the optimizer can turn a ptr_eq call + a branch to a single
instruction. This, however, would require likely + unlikely variants.

Alternatively, we can just not have all the complexity, and just emit one extra
instruction:

    static __always_inline
    bool ptr_eq(const volatile void *a, const volatile void *b) {
        unsigned long diff;
        asm("xor %0, %1, %2" : "=r" (diff) : "r" (a), "r" (b));
        return diff == 0;
    }

Thoughts?

Best,
Gary


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH hazptr v2 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list
  2026-10-09 18:21 ` [PATCH hazptr v2 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
@ 2026-10-10 19:39   ` Bradley Morgan
  0 siblings, 0 replies; 9+ messages in thread
From: Bradley Morgan @ 2026-10-10 19:39 UTC (permalink / raw)
  To: Mathieu Desnoyers, Paul E . McKenney
  Cc: linux-kernel, Gary Guo, Boqun Feng, rcu, lkmm, Lian Wang, Kunwu Chan

On 9 October 2026 19:21:35 BST, Mathieu Desnoyers
<mathieu.desnoyers@efficios.com> wrote:
>Introduce a "try acquire" hazard pointer fast path, which performs an
>early load of the address to store it into the hazard pointer slot, and
>then re-loads that address after a barrier to check whether it has
>changed meanwhile.
>
>On comparison failure, rather than re-try, guarantee forward progress by
>falling back to the __hazptr_acquire slow path on failure.
>
>The acquire slow path attempts a try-acquire for any available per-CPU
>slot. If that fails, it chains the backup slot into the overflow list,
>therefore guaranteeing forward progress for both hazard pointer
>read-side and synchronize:
>
>- Readers set the wildcard, and then proceed to set the more
>  specific address to replace the wildcard.
>
>- One synchronize alternates between two overflow list periods,
>  scanning each one while readers are added to the other period,
>  thus preventing a steady flow of readers from preventing
>  synchronize forward progress.
>
>With this change, the scan on per-CPU slots don't need to expect a
>wildcard anymore, because none can be produced by readers. Wildcards are
>only expected within overflow lists.
>
>Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
>Suggested-by: Gary Guo <gary@garyguo.net>
>Link: https://lore.kernel.org/lkmm/DLK9MW1N86KD.2GRWUAD2NX4FU@garyguo.net/
>Reviewed-by: Gary Guo <gary@garyguo.net>

Hey, looks fine to me! 

Reviewed-by: Bradley Morgan <brads@mainlining.org>

(Well needed in my opinion)

>Cc: Paul E. McKenney <paulmck@kernel.org>
>Cc: Boqun Feng <boqun@kernel.org>
>Cc: Bradley Morgan <brads@mainlining.org>
>Cc: Gary Guo <gary@garyguo.net>
>Cc: <rcu@vger.kernel.org>
>Cc: <lkmm@lists.linux.dev>
>Cc: Lian Wang <lianux.mm@gmail.com>
>Cc: Kunwu Chan <kunwu.chan@gmail.com>
>---
> include/linux/hazptr.h |  48 +++++++++++--------
> kernel/hazptr.c        | 103 ++++++++++++++++++++---------------------
> 2 files changed, 77 insertions(+), 74 deletions(-)
>
>diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
>index d1670121947a..df2aca9aa26b 100644
>--- a/include/linux/hazptr.h
>+++ b/include/linux/hazptr.h
>@@ -25,13 +25,11 @@
> #include <linux/types.h>
> #include <linux/cleanup.h>
> #include <linux/sched.h>
>+#include <linux/ptreq.h>
> 
> /* 4 slots (each sizeof(hazptr_slot_item)) fit in a single 64-byte cache
> line. */
> #define NR_HAZPTR_PERCPU_SLOTS	4
> 
>-/* The current hazard pointer wildcard. */
>-extern void *hazptr_wildcard;
>-
> /*
>  * Hazard pointer slot.
>  */
>@@ -190,6 +188,31 @@ void hazptr_note_context_switch(void)
> 	}
> }
> 
>+/* Try hazard pointer protection. */
>+static inline
>+void *__hazptr_try_acquire(struct hazptr_ctx *ctx, void * const *addr_p, struct hazptr_slot *slot)
>+{
>+	void *early_addr, *addr;
>+
>+	if (unlikely(slot->addr))
>+		return NULL;
>+	early_addr = READ_ONCE(*addr_p);	/* Early load. */
>+	WRITE_ONCE(slot->addr, early_addr);	/* Store B */
>+	/* Memory ordering: Store B before Load A. */
>+	smp_mb();
>+	addr = READ_ONCE(*addr_p);		/* Load A */
>+	/*
>+	 * Validate that address did not change between Early load and Load A.
>+	 * Use ptr_eq() to make sure that result from Load A is returned to the
>+	 * caller to preserve address dependency.
>+	 */
>+	if (unlikely(!ptr_eq(addr, early_addr))) {
>+		WRITE_ONCE(slot->addr, NULL);
>+		return NULL;
>+	}
>+	return addr;
>+}
>+
> /**
>  * hazptr_acquire - Load pointer at address and protect with hazard pointer.
>  *
>@@ -245,24 +268,9 @@ void *hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> 	ctx->acquire_cpu = smp_processor_id();
> 	ctx->acquire_caller = _THIS_IP_;
> #endif
>-	if (unlikely(slot->addr))
>+	addr = __hazptr_try_acquire(ctx, addr_p, slot);
>+	if (unlikely(!addr))
> 		return __hazptr_acquire(ctx, addr_p);
>-	WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard));	/* Store B */
>-
>-	/* Memory ordering: Store B before Load A. */
>-	smp_mb();
>-
>-	/*
>-	 * Load @addr_p after storing wildcard to the hazard pointer slot.
>-	 */
>-	addr = READ_ONCE(*addr_p);	/* Load A */
>-
>-	/*
>-	 * We don't care about ordering of Store C. It will simply
>-	 * replace the wildcard by a more specific address. If addr is
>-	 * NULL, we simply store NULL into the slot.
>-	 */
>-	WRITE_ONCE(slot->addr, addr);	/* Store C */
> 	slot_item->ctx.ctx = ctx;
> 	ctx->slot = slot;
> 	return addr;
>diff --git a/kernel/hazptr.c b/kernel/hazptr.c
>index 13faa5ba7677..3ca73b56a5c2 100644
>--- a/kernel/hazptr.c
>+++ b/kernel/hazptr.c
>@@ -13,17 +13,9 @@
> #include <linux/list.h>
> #include <linux/export.h>
> 
>-static DEFINE_MUTEX(hazptr_phase_lock);	/* Protect the wildcard and list phase flip. */
>+#define HAZPTR_WILDCARD	((void *) 1UL)
> 
>-/*
>- * The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
>- * hazptr_synchronize forward progress even with a steady stream of readers.
>- * This wildcard value is used by acquire to temporarily tag the per-CPU slots.
>- * This also affects the overflow list selection: the current list used by
>- * readers is array[(unsigned long) hazptr_wildcard - 1].
>- */
>-void *hazptr_wildcard = (void *) 1UL;
>-EXPORT_SYMBOL_GPL(hazptr_wildcard);
>+static DEFINE_MUTEX(hazptr_phase_lock);	/* Protect the list phase flip. */
> 
> /* The current overflow list phase. */
> static unsigned int hazptr_overflow_list_phase;
>@@ -41,6 +33,10 @@ struct hazptr_overflow_list {
>  * successively iterates on both lists. Therefore, only list removals
>  * can cause the iteration to retry, and the number of removals is
>  * limited to the number of list elements.
>+ *
>+ * Due to the overflow list raw spin lock, the hazard pointer readers are
>+ * blocking, starvation-free with bounded waiting, assuming bounded critical
>+ * sections and no NMI or virtualization-induced holder preemption.
>  */
> struct hazptr_overflow_list_flip {
> 	struct hazptr_overflow_list array[2];
>@@ -51,26 +47,12 @@ static DEFINE_PER_CPU(struct hazptr_overflow_list_flip, percpu_overflow_list_fli
> DEFINE_PER_CPU(struct hazptr_percpu_slots, hazptr_percpu_slots);
> EXPORT_PER_CPU_SYMBOL_GPL(hazptr_percpu_slots);
> 
>-static
>-void *flip_wildcard(void *wildcard)
>-{
>-	return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
>-}
>-
> static
> unsigned int flip_list_phase(unsigned int phase)
> {
> 	return 1 - phase;
> }
> 
>-static
>-bool is_wildcard(void *addr)
>-{
>-	if ((unsigned long) addr == 1UL || (unsigned long) addr == 2UL)
>-		return true;
>-	return false;
>-}
>-
> static
> struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
> {
>@@ -96,16 +78,36 @@ struct hazptr_slot *hazptr_get_free_percpu_slot(struct hazptr_ctx *ctx)
>  */
> void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> {
>-	struct hazptr_slot *slot = hazptr_get_free_percpu_slot(ctx);
>+	struct hazptr_slot *slot;
> 	void *addr;
> 
> 	/*
>-	 * If all the per-CPU slots are already in use, fallback
>-	 * to the backup slot.
>+	 * In case we are called due to nested use of hazard pointers,
>+	 * try a slot protection with per-CPU slots.
> 	 */
>-	if (unlikely(!slot))
>-		slot = hazptr_chain_backup_slot(ctx);
>-	WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard));	/* Store B */
>+	slot = hazptr_get_free_percpu_slot(ctx);
>+	if (likely(slot)) {
>+		addr = __hazptr_try_acquire(ctx, addr_p, slot);
>+		if (addr) {
>+			ctx->slot = slot;
>+			return addr;
>+		}
>+	}
>+
>+	/*
>+	 * The backup slot overflow list guarantees forward progress of both
>+	 * hazard pointer readers and synchronize:
>+	 *
>+	 * - Readers set the wildcard, and then proceed to set the more
>+	 *   specific address to replace the wildcard.
>+	 *
>+	 * - One synchronize alternates between two overflow list periods,
>+	 *   scanning each one while readers are added to the other period,
>+	 *   thus preventing a steady flow of readers from preventing
>+	 *   synchronize forward progress.
>+	 */
>+	slot = hazptr_chain_backup_slot(ctx);
>+	WRITE_ONCE(slot->addr, HAZPTR_WILDCARD);	/* Store B */
> 
> 	/* Memory ordering: Store B before Load A. */
> 	smp_mb();
>@@ -121,20 +123,21 @@ void *__hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
> 	 * NULL, we simply store NULL into the slot.
> 	 */
> 	WRITE_ONCE(slot->addr, addr);	/* Store C */
>+
> 	ctx->slot = slot;
>-	if (!addr && hazptr_slot_is_backup(ctx, slot))
>+	if (!addr)
> 		hazptr_unchain_backup_slot(ctx);
> 	return addr;
> }
> EXPORT_SYMBOL_GPL(__hazptr_acquire);
> 
> /*
>- * Perform piecewise iteration on overflow list waiting until "addr" is
>- * not present. Raw spinlock is released and taken between each list
>- * item and busy loop iteration. The overflow list generation is checked
>- * each time the lock is taken to validate that the list has not changed
>- * before resuming iteration or busy wait. If the generation has
>- * changed, retry the entire list traversal.
>+ * Perform piecewise iteration on overflow list waiting until "addr" and
>+ * wildcard are not present. Raw spinlock is released and taken between each
>+ * list item and busy loop iteration. The overflow list generation is checked
>+ * each time the lock is taken to validate that the list has not changed before
>+ * resuming iteration or busy wait. If the generation has changed, retry the
>+ * entire list traversal.
>  */
> static
> void hazptr_synchronize_overflow_list(struct hazptr_overflow_list
> *overflow_list, void *addr)
>@@ -147,13 +150,11 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
> retry:
> 	snapshot_gen = overflow_list->gen;
> 	hlist_for_each_entry(backup_slot, &overflow_list->head, overflow_node) {
>-		/* Busy-wait if node is found. */
>+		/* Busy-wait if addr or wildcard are found. */
> 		for (;;) {
> 			void *load_addr = smp_load_acquire(&backup_slot->slot.addr);	/* Load B */
> 
>-			/* We don't expect wildcards in overflow list. */
>-			WARN_ON_ONCE(is_wildcard(load_addr));
>-			if (load_addr != addr)
>+			if (load_addr != addr && load_addr != HAZPTR_WILDCARD)
> 				break;
> 			raw_spin_unlock_irqrestore(&overflow_list->lock, flags);
> 			cpu_relax();
>@@ -174,7 +175,7 @@ void hazptr_synchronize_overflow_list(struct hazptr_overflow_list *overflow_list
> }
> 
> static
>-void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
>+void hazptr_synchronize_cpu_slots(int cpu, void *addr)
> {
> 	struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
> 	unsigned int idx;
>@@ -182,13 +183,13 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
> 	for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
> 		struct hazptr_slot_item *item = &percpu_slots->items[idx];
> 
>-		/* Busy-wait if node is found. */
>-		smp_cond_load_acquire(&item->slot.addr, VAL != addr && VAL != scan_wildcard); /* Load B */
>+		/* Busy-wait if addr is found. */
>+		smp_cond_load_acquire(&item->slot.addr, VAL != addr); /* Load B */
> 	}
> }
> 
> static
>-void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
>+void hazptr_scan_cpu_slots(void *addr)
> {
> 	int cpu;
> 
>@@ -196,16 +197,13 @@ void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
> 	for_each_possible_cpu(cpu) {
> 		/*
> 		 * Scan CPU slots.
>-		 * Forward progress against recurring wildcards is guaranteed
>-		 * by scanning for one wildcard while new elements use the
>-		 * other wildcard value (1UL vs 2UL).
> 		 * Forward progress against recurring single hazard pointer
> 		 * values is guaranteed by the fact that a hazard pointer
> 		 * is not reclaimed nor reused until the scan for that hazard
> 		 * pointer completes, which prevents a steady flow of readers
> 		 * to acquire that same hazard pointer value.
> 		 */
>-		hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
>+		hazptr_synchronize_cpu_slots(cpu, addr);
> 	}
> }
> 
>@@ -237,7 +235,6 @@ void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
> void hazptr_synchronize(void *addr)
> {
> 	unsigned int scan_list_phase;
>-	void *scan_wildcard;
> 
> 	/*
> 	 * Busy-wait should only be done from preemptible context.
>@@ -251,16 +248,14 @@ void hazptr_synchronize(void *addr)
> 	 */
> 	if (!addr)
> 		return;
>+
> 	/* Memory ordering: Store A before Load B. */
> 	smp_mb();
> 
> 	guard(mutex)(&hazptr_phase_lock);
> 
> 	/* Scan per-CPU slots. */
>-	scan_wildcard = flip_wildcard(hazptr_wildcard);
>-	hazptr_scan_cpu_slots_period(addr, scan_wildcard);
>-	WRITE_ONCE(hazptr_wildcard, scan_wildcard);			/* Flip the current wildcard. */
>-	hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
>+	hazptr_scan_cpu_slots(addr);
> 
> 	/*
> 	 * Scan overflow lists *after* scanning per-CPU slots. See
>

--- Thanks!
"I'm not a very positive person" - Linus torvalds

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-10-10 19:39 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 18:21 [PATCH hazptr v2 0/4] Hazard pointer updates Mathieu Desnoyers
2026-10-09 18:21 ` [PATCH hazptr v2 1/4] hazptr: Fix two-phase hazptr_synchronize race with detach Mathieu Desnoyers
2026-10-09 18:21 ` [PATCH hazptr v2 2/4] ptreq.h: Introduce ptr_eq() to preserve address dependency Mathieu Desnoyers
2026-10-09 19:03   ` Boqun Feng
2026-10-09 19:07     ` Mathieu Desnoyers
2026-10-09 21:55   ` Gary Guo
2026-10-09 18:21 ` [PATCH hazptr v2 3/4] Documentation: RCU: Refer to ptr_eq() Mathieu Desnoyers
2026-10-09 18:21 ` [PATCH hazptr v2 4/4] hazptr: Introduce "try acquire" fast path, fallback to overflow list Mathieu Desnoyers
2026-10-10 19:39   ` Bradley Morgan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®