* [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan
@ 2026-10-02 17:08 Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 01/15] hazptr: add shared scan kthread Kunwu Chan
` (15 more replies)
0 siblings, 16 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
This series extends Paul McKenney's v3 hazptr implementation [1]
with a shared scan path for concurrent hazptr_synchronize() callers,
and adapts the lockdep dynamic-key hashlist use case from Boqun
Feng's 2025 shazptr series [2] to the current hazptr API.
The series also adds rcuscale support and torture coverage for
the hazptr implementation.
The lockdep conversion replaces the expedited RCU wait in
lockdep_unregister_key() with hazptr_synchronize(). This limits
the wait to hazard pointers protecting the target hash bucket
instead of waiting for a system-wide expedited RCU grace period.
[1] https://lore.kernel.org/all/20260919000056.3132131-26-paulmck@kernel.org/
[2] https://lore.kernel.org/lkml/20250625031101.12555-1-boqun.feng@gmail.com/
Performance data
================
All measurements are on an ARM64 KVM guest with
HAZPTR_SCALE_NR_OBJS=8 and round-robin updater selection; unless
noted otherwise, nreaders=0. Scale tests run with
PROVE_LOCKING=n.
rcuscale synchronize latency (96 CPUs, nw=1):
nreaders=0 nreaders=1 nreaders=4 nreaders=96
scale avg avg avg avg
--------------------------------------------------------------------
hazptr 46 us 47 us 102 us 106 ms
RCU 8.3 ms 8.5 ms 16.6 ms 54.4 ms
SRCU 8.2 ms 8.0 ms 11.7 ms 11.8 ms
Hazptr single-writer latency remains below 110 us with up to
four readers on this 96-CPU guest, but rises to 106 ms when
all 96 CPUs hold hazard pointers.
Under 16 concurrent synchronize callers (nw=16), per-writer
latency at four CPU counts:
hazptr nw=16 RCU nw=16 SRCU nw=16
CPUs avg avg avg
-----------------------------------------------------------
24 8.0 ms 13.6 ms 9.2 ms
96 8.0 ms 13.3 ms 8.6 ms
128 7.9 ms 13.0 ms 11.7 ms
256 8.0 ms 20.8 ms 9.6 ms
Hazptr remains around 8.0 ms across these CPU counts, with the
scan-kthread retry interval contributing to the latency.
RCU rises to ~21 ms at 256 CPUs while SRCU stays around 9-12 ms.
Reader-side overhead (refscale, 96 readers on the 96-CPU guest,
3 runs each):
hazptr 25.7 ns/op
RCU 96.7 ns/op
SRCU 134.6 ns/op
lockdep workload -- tc qdisc mq x100, ARM64 KVM, PROVE_LOCKING=y,
96 background hazptr readers, function-call IPIs per 100 ops:
CPUs hazptr exp RCU reduction
------------------------------------------------------------
24 196 207 5%
96 189 265 29%
128 193 266 27%
256 222 384 42%
Hazptr issues fewer function-call IPIs at each CPU count, with
the difference reaching 42% at 256 CPUs.
The hazptr.sh test suite passes, including the lockdep scenarios
and the 8-to-256 CPU sweep, with no lockdep warnings, deadlocks,
or crashes.
Additional x86 server testing with Lian Wang is planned(maybe after
LPC).
Changes since RFC/WIP
=====================
- RFC/WIP: https://lore.kernel.org/all/20260922070950.4173245-1-kunwu.chan@gmail.com/
- Split the original 4-patch RFC/WIP into smaller commits covering
shared scanning, correctness, API support, lockdep, scaling, and
torture testing.
- Incorporated Boqun Feng's review feedback: use a Bloom filter to
avoid per-waiter allocation, add scoped_guard() support, and add
a debug option to force the hazptr acquire slow path.
- Fixed scan ordering around backup-slot promotion by scanning all
per-CPU slots before the overflow lists, with a separate
overflow-list phase.
- Simplified the scan cycle to flip first and drain only the old
wildcard generation, with herd7-verified LKMM tests for both the
in-flight and resolved publication cases.
- Extended rcuscale and hazptrtorture coverage, added a selftest
script for the torture configurations, and fixed the
hazptr_release() kernel-doc.
Kunwu Chan (15):
hazptr: add shared scan kthread
hazptr: use Bloom filter for shared scan waiters
hazptr: scan all per-CPU slots before overflow lists
hazptr: add scoped_guard() support
hazptr: add debug option to force the acquire slow path
hazptr: elide redundant first drain pass
Documentation/litmus-tests: add hazptr wildcard-flip escape test
locking/lockdep: use hazptr to wait for dynamic key lookups
rcuscale: add hazptr scale type
hazptr: fix kernel-doc of hazptr_release()
Documentation/litmus-tests: add hazptr acquire-before-scan test
hazptrtorture: add slowpath and lockdep scenarios
hazptrtorture: add READERS4 and READERS0 torture configs
hazptrtorture: add 128- and 256-CPU configs
selftests/rcutorture: add hazptr torture test script
Documentation/litmus-tests/README | 13 +
.../hazptr/hazptr-acquire-before-scan.litmus | 45 +++
.../hazptr/hazptr-wildcard-flip-escape.litmus | 46 +++
include/linux/hazptr.h | 56 ++-
kernel/hazptr.c | 336 +++++++++++++++++-
kernel/locking/lockdep.c | 25 +-
kernel/rcu/Kconfig.debug | 10 +
kernel/rcu/hazptrtorture.c | 57 ++-
kernel/rcu/rcuscale.c | 70 +++-
.../selftests/rcutorture/bin/hazptr.sh | 146 ++++++++
.../rcutorture/configs/hazptr/CFLIST | 6 +
.../rcutorture/configs/hazptr/CPU128 | 16 +
.../rcutorture/configs/hazptr/CPU128.boot | 1 +
.../rcutorture/configs/hazptr/CPU256 | 16 +
.../rcutorture/configs/hazptr/CPU256.boot | 1 +
.../rcutorture/configs/hazptr/LOCKDEP | 17 +
.../rcutorture/configs/hazptr/LOCKDEP.boot | 1 +
.../rcutorture/configs/hazptr/READERS0 | 16 +
.../rcutorture/configs/hazptr/READERS0.boot | 2 +
.../rcutorture/configs/hazptr/READERS4 | 16 +
.../rcutorture/configs/hazptr/READERS4.boot | 2 +
.../rcutorture/configs/hazptr/SLOWPATH | 16 +
22 files changed, 881 insertions(+), 33 deletions(-)
create mode 100644 Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus
create mode 100644 Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus
create mode 100755 tools/testing/selftests/rcutorture/bin/hazptr.sh
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 01/15] hazptr: add shared scan kthread
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 02/15] hazptr: use Bloom filter for shared scan waiters Kunwu Chan
` (14 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Batch concurrent hazptr_synchronize() callers into a shared scan
cycle, avoiding redundant wildcard flips and slot scans.
Queue waiters to a kthread and perform a two-phase wildcard scan
once for all queued waiters. After both phases, complete waiters
whose address is no longer held by any slot; waiters that remain
blocked are retried in a later cycle.
Waiters are embedded in the caller's stack frame, so no dynamic
allocation is needed in the synchronize path.
Fall back to the direct scan if the scan kthread is unavailable.
Adapted from the scan-kthread approach in Boqun Feng's shazptr
implementation.
Link: https://lore.kernel.org/lkml/20250625031101.12555-1-boqun.feng@gmail.com/
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
kernel/hazptr.c | 201 ++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 201 insertions(+)
diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index d3d1050d92cf..9e274a691af5 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -12,6 +12,10 @@
#include <linux/mutex.h>
#include <linux/list.h>
#include <linux/export.h>
+#include <linux/completion.h>
+#include <linux/kthread.h>
+#include <linux/swait.h>
+#include <linux/sched.h>
/*
* The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
@@ -209,12 +213,179 @@ void hazptr_scan_period(void *addr, void *scan_wildcard)
}
}
+struct hazptr_waiter {
+ struct list_head node;
+ void *addr;
+ struct completion done;
+};
+
+struct hazptr_scan_state {
+ struct task_struct *kthread;
+ struct swait_queue_head wq;
+ bool wakeup;
+ struct mutex lock;
+ struct list_head pending;
+ struct list_head scanning; /* kthread only */
+};
+static struct hazptr_scan_state hazptr_scan;
+
+/*
+ * Check per-CPU slots before overflow-list slots to match the
+ * acquisition ordering of promoted slots.
+ */
+static bool hazptr_value_present(void *val)
+{
+ int cpu;
+
+ for_each_possible_cpu(cpu) {
+ struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
+ struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
+ unsigned int idx;
+
+ for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
+ struct hazptr_slot_item *item = &percpu_slots->items[idx];
+
+ /* Pairs with smp_store_release in hazptr_release(). */
+ if (smp_load_acquire(&item->slot.addr) == val)
+ return true;
+ }
+ for (int i = 0; i < 2; i++) {
+ struct hazptr_overflow_list *list = &overflow_list_flip->array[i];
+ struct hazptr_backup_slot *b;
+ unsigned long flags;
+
+ raw_spin_lock_irqsave(&list->lock, flags);
+ hlist_for_each_entry(b, &list->head, overflow_node) {
+ /* Pairs with smp_store_release in hazptr_release(). */
+ if (smp_load_acquire(&b->slot.addr) == val) {
+ raw_spin_unlock_irqrestore(&list->lock, flags);
+ return true;
+ }
+ }
+ raw_spin_unlock_irqrestore(&list->lock, flags);
+ }
+ }
+ return false;
+}
+
+/*
+ * Wait until no per-CPU slot or overflow-list slot holds @wc.
+ * Callers must ensure that the wildcard value in use by new acquires
+ * differs from @wc, so that the set of slots holding @wc only
+ * shrinks, which guarantees forward progress.
+ */
+static void hazptr_drain_wildcard(void *wc)
+{
+ while (hazptr_value_present(wc))
+ cond_resched();
+}
+
+/*
+ * Move pending waiters to ->scanning and perform a two-phase
+ * wildcard scan shared by all waiters.
+ */
+static void hazptr_scan_do_cycle(void)
+{
+ void *scan_wildcard, *old_wildcard;
+ struct hazptr_waiter *w, *n;
+ LIST_HEAD(done);
+
+ mutex_lock(&hazptr_wildcard_lock);
+
+ mutex_lock(&hazptr_scan.lock);
+ list_splice_tail_init(&hazptr_scan.pending, &hazptr_scan.scanning);
+ mutex_unlock(&hazptr_scan.lock);
+
+ if (list_empty(&hazptr_scan.scanning)) {
+ mutex_unlock(&hazptr_wildcard_lock);
+ return;
+ }
+
+ /* Pass 1: drain the unpublished wildcard. */
+ scan_wildcard = flip_wildcard(READ_ONCE(hazptr_wildcard));
+ hazptr_drain_wildcard(scan_wildcard);
+
+ /* Flip so new acquires use the new generation. */
+ WRITE_ONCE(hazptr_wildcard, scan_wildcard);
+ old_wildcard = flip_wildcard(scan_wildcard);
+
+ /* Pass 2: drain the old wildcard. */
+ hazptr_drain_wildcard(old_wildcard);
+
+ /* Complete waiters whose address is no longer held by any slot. */
+ list_for_each_entry_safe(w, n, &hazptr_scan.scanning, node) {
+ if (!hazptr_value_present(w->addr))
+ list_move(&w->node, &done);
+ }
+
+ mutex_unlock(&hazptr_wildcard_lock);
+
+ list_for_each_entry_safe(w, n, &done, node) {
+ list_del_init(&w->node);
+ complete(&w->done);
+ }
+}
+
+static int hazptr_scan_kthread(void *unused)
+{
+ for (;;) {
+ bool idle;
+
+ swait_event_idle_exclusive(hazptr_scan.wq,
+ READ_ONCE(hazptr_scan.wakeup));
+
+ hazptr_scan_do_cycle();
+
+ mutex_lock(&hazptr_scan.lock);
+ idle = list_empty(&hazptr_scan.pending) &&
+ list_empty(&hazptr_scan.scanning);
+ if (idle)
+ WRITE_ONCE(hazptr_scan.wakeup, false);
+ mutex_unlock(&hazptr_scan.lock);
+
+ if (idle)
+ continue;
+ /* Waiters still blocked: retry after a short delay. */
+ schedule_timeout_idle(1);
+ }
+ return 0;
+}
+
+/*
+ * Queue @addr for the shared scan. The waiter lives on the
+ * caller's stack, so no allocation is needed.
+ */
+static void hazptr_synchronize_queued(void *addr)
+{
+ struct hazptr_waiter waiter = {
+ .addr = addr,
+ };
+
+ init_completion(&waiter.done);
+ INIT_LIST_HEAD(&waiter.node);
+
+ /* Enqueue and wake the scan kthread. */
+ mutex_lock(&hazptr_scan.lock);
+ list_add_tail(&waiter.node, &hazptr_scan.pending);
+ if (!READ_ONCE(hazptr_scan.wakeup)) {
+ WRITE_ONCE(hazptr_scan.wakeup, true);
+ swake_up_one(&hazptr_scan.wq);
+ }
+ mutex_unlock(&hazptr_scan.lock);
+
+ /* Sleep until the scan kthread completes this waiter. */
+ wait_for_completion(&waiter.done);
+}
+
/*
* hazptr_synchronize: Wait until @addr is released from all slots.
*
* Wait to observe that each slot contains a value that differs from
* @addr before returning.
* Should be called from preemptible context.
+ *
+ * If available, queue the caller for a shared scan; otherwise use
+ * the direct scan path.
*/
void hazptr_synchronize(void *addr)
{
@@ -235,6 +406,13 @@ void hazptr_synchronize(void *addr)
/* Memory ordering: Store A before Load B. */
smp_mb();
+ /* Pairs with smp_store_release in hazptr_scan_init(). */
+ if (smp_load_acquire(&hazptr_scan.kthread)) {
+ hazptr_synchronize_queued(addr);
+ return;
+ }
+
+ /* Fallback: use the direct scan path. */
guard(mutex)(&hazptr_wildcard_lock);
scan_wildcard = flip_wildcard(hazptr_wildcard);
hazptr_scan_period(addr, scan_wildcard);
@@ -282,3 +460,26 @@ void __init hazptr_init(void)
}
}
}
+
+/*
+ * Initialize the scan kthread. Failure falls back to the
+ * direct scan.
+ */
+static int __init hazptr_scan_init(void)
+{
+ struct task_struct *t;
+
+ init_swait_queue_head(&hazptr_scan.wq);
+ mutex_init(&hazptr_scan.lock);
+ INIT_LIST_HEAD(&hazptr_scan.pending);
+ INIT_LIST_HEAD(&hazptr_scan.scanning);
+
+ t = kthread_run(hazptr_scan_kthread, NULL, "hazptr_scan");
+ if (!IS_ERR(t))
+ /* Pairs with smp_load_acquire in hazptr_synchronize(). */
+ smp_store_release(&hazptr_scan.kthread, t);
+ else
+ pr_warn("hazptr: scan kthread failed, using direct scan\n");
+ return 0;
+}
+core_initcall(hazptr_scan_init);
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 02/15] hazptr: use Bloom filter for shared scan waiters
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 01/15] hazptr: add shared scan kthread Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 03/15] hazptr: scan all per-CPU slots before overflow lists Kunwu Chan
` (13 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Replace the per-waiter exact-match scan with a shared Bloom filter
built during the second drain pass: each observed non-wildcard
address sets three bits, so a waiter whose address hashes to an
unset bit can be completed without another slot walk.
A Bloom filter has no false negatives, so an address absent from
the final drain scan cannot be held by an observed slot; false
positives only delay completion to a later scan cycle.
The filter uses a one-page bitmap with three multiply-shift hash
functions. It lives in the hazptr_scan_state and is only touched
by the scan kthread, so waiter state remains on the caller's stack
and no dynamic allocation is needed.
Suggested-by: Boqun Feng <boqun@kernel.org>
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
kernel/hazptr.c | 93 +++++++++++++++++++++++++++++++++++++++++++------
1 file changed, 83 insertions(+), 10 deletions(-)
diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index 9e274a691af5..900ba35de2cb 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -7,6 +7,7 @@
*/
#include <linux/hazptr.h>
+#include <linux/bitops.h>
#include <linux/percpu.h>
#include <linux/spinlock.h>
#include <linux/mutex.h>
@@ -16,6 +17,7 @@
#include <linux/kthread.h>
#include <linux/swait.h>
#include <linux/sched.h>
+#include <linux/bitmap.h>
/*
* The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
@@ -213,6 +215,55 @@ void hazptr_scan_period(void *addr, void *scan_wildcard)
}
}
+/* Number of hash functions for the Bloom filter. */
+#define HAZPTR_BLOOM_HASHES 3
+
+/* Number of bits in the Bloom filter bitmap: one page. */
+#define HAZPTR_BLOOM_NBITS (PAGE_SIZE * 8)
+
+struct hazptr_bloom {
+ unsigned long map[PAGE_SIZE / sizeof(unsigned long)];
+};
+
+/*
+ * Multiply @addr by a per-hash odd constant and use the high bits
+ * as the Bloom filter index.
+ */
+static unsigned long hazptr_bloom_hash(void *addr, unsigned int i)
+{
+ static const u64 mult[HAZPTR_BLOOM_HASHES] = {
+ 0x9E3779B97F4A7C15ULL,
+ 0xC2B2AE3D27D4EB4FULL,
+ 0x165667B19E3779F9ULL,
+ };
+ u64 hash = (u64)(unsigned long)addr * mult[i];
+
+ return hash >> (64 - ilog2(HAZPTR_BLOOM_NBITS));
+}
+
+static void hazptr_bloom_reset(struct hazptr_bloom *bloom)
+{
+ bitmap_zero(bloom->map, HAZPTR_BLOOM_NBITS);
+}
+
+static void hazptr_bloom_add(struct hazptr_bloom *bloom, void *addr)
+{
+ unsigned int i;
+
+ for (i = 0; i < HAZPTR_BLOOM_HASHES; i++)
+ __set_bit(hazptr_bloom_hash(addr, i), bloom->map);
+}
+
+static bool hazptr_bloom_contains(const struct hazptr_bloom *bloom, void *addr)
+{
+ unsigned int i;
+
+ for (i = 0; i < HAZPTR_BLOOM_HASHES; i++)
+ if (!test_bit(hazptr_bloom_hash(addr, i), bloom->map))
+ return false;
+ return true;
+}
+
struct hazptr_waiter {
struct list_head node;
void *addr;
@@ -226,17 +277,26 @@ struct hazptr_scan_state {
struct mutex lock;
struct list_head pending;
struct list_head scanning; /* kthread only */
+ struct hazptr_bloom bloom; /* kthread only */
};
static struct hazptr_scan_state hazptr_scan;
/*
- * Check per-CPU slots before overflow-list slots to match the
- * acquisition ordering of promoted slots.
+ * Walk all slots and return true if @watch is present. If @bloom
+ * is non-NULL, record observed non-wildcard addresses in it.
+ *
+ * Per-CPU slots are examined before overflow-list slots on each CPU
+ * to preserve the acquisition ordering required by the promote path:
+ * synchronize must observe the per-CPU slot release before the
+ * overflow-list entry can be missed.
*/
-static bool hazptr_value_present(void *val)
+static bool hazptr_scan_walk(void *watch, struct hazptr_bloom *bloom)
{
int cpu;
+ if (bloom)
+ hazptr_bloom_reset(bloom);
+
for_each_possible_cpu(cpu) {
struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
@@ -244,10 +304,14 @@ static bool hazptr_value_present(void *val)
for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
struct hazptr_slot_item *item = &percpu_slots->items[idx];
+ void *v;
/* Pairs with smp_store_release in hazptr_release(). */
- if (smp_load_acquire(&item->slot.addr) == val)
+ v = smp_load_acquire(&item->slot.addr);
+ if (v == watch)
return true;
+ if (bloom && v && !is_wildcard(v))
+ hazptr_bloom_add(bloom, v);
}
for (int i = 0; i < 2; i++) {
struct hazptr_overflow_list *list = &overflow_list_flip->array[i];
@@ -256,11 +320,16 @@ static bool hazptr_value_present(void *val)
raw_spin_lock_irqsave(&list->lock, flags);
hlist_for_each_entry(b, &list->head, overflow_node) {
+ void *v;
+
/* Pairs with smp_store_release in hazptr_release(). */
- if (smp_load_acquire(&b->slot.addr) == val) {
+ v = smp_load_acquire(&b->slot.addr);
+ if (v == watch) {
raw_spin_unlock_irqrestore(&list->lock, flags);
return true;
}
+ if (bloom && v && !is_wildcard(v))
+ hazptr_bloom_add(bloom, v);
}
raw_spin_unlock_irqrestore(&list->lock, flags);
}
@@ -276,7 +345,7 @@ static bool hazptr_value_present(void *val)
*/
static void hazptr_drain_wildcard(void *wc)
{
- while (hazptr_value_present(wc))
+ while (hazptr_scan_walk(wc, NULL))
cond_resched();
}
@@ -309,12 +378,16 @@ static void hazptr_scan_do_cycle(void)
WRITE_ONCE(hazptr_wildcard, scan_wildcard);
old_wildcard = flip_wildcard(scan_wildcard);
- /* Pass 2: drain the old wildcard. */
- hazptr_drain_wildcard(old_wildcard);
+ /*
+ * Pass 2: drain the old wildcard while collecting observed
+ * addresses into the Bloom filter.
+ */
+ while (hazptr_scan_walk(old_wildcard, &hazptr_scan.bloom))
+ cond_resched();
- /* Complete waiters whose address is no longer held by any slot. */
+ /* Complete waiters whose address is not in the filter. */
list_for_each_entry_safe(w, n, &hazptr_scan.scanning, node) {
- if (!hazptr_value_present(w->addr))
+ if (!hazptr_bloom_contains(&hazptr_scan.bloom, w->addr))
list_move(&w->node, &done);
}
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 03/15] hazptr: scan all per-CPU slots before overflow lists
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 01/15] hazptr: add shared scan kthread Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 02/15] hazptr: use Bloom filter for shared scan waiters Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 04/15] hazptr: add scoped_guard() support Kunwu Chan
` (12 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Boqun Feng pointed out that a promote racing a scan can cause the
scan to miss a hazard pointer: hazptr_promote_to_backup_slot()
chains the backup slot into an overflow list before clearing the
per-CPU slot, so the scan must observe the old location before the
new one. The scan currently visits each CPU's overflow lists right
after that CPU's per-CPU slots, so a promote that moves a hazard
pointer from a per-CPU slot of a later-scanned CPU into an overflow
list of an earlier-scanned CPU can be missed.
The direct two-phase scan has a related race because the overflow
list selection follows the wildcard phase: a promote racing the
wildcard flip can move a backup slot to the list not covered by the
second phase.
Fix this by scanning all per-CPU slots before any overflow-list
slot, in both the shared-scan walk and the direct fallback scan.
Give the overflow lists a flip phase of their own, independent of
the wildcard, so each list is scanned while it is the non-live one:
new backup slots are chained to the other list, preserving forward
progress.
The same race was also addressed by Mathieu Desnoyers.
Fixes: 0b8114b25f17 ("hazptr: Implement two-phase wildcard scan")
Reported-by: Boqun Feng <boqun@kernel.org>
Link: https://lore.kernel.org/all/20260925195958.4766-1-mathieu.desnoyers@efficios.com/
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
kernel/hazptr.c | 88 +++++++++++++++++++++++++++++++++++--------------
1 file changed, 64 insertions(+), 24 deletions(-)
diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index 900ba35de2cb..799784f698ed 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -23,13 +23,20 @@
* The current hazard pointer wildcard. Flips between 1UL and 2UL to guarantee
* hazptr_synchronize forward progress even with a steady stream of readers.
* This wildcard value is used by acquire to temporarily tag the per-CPU slots.
- * This also affects the overflow list selection: the current list used by
- * readers is array[(unsigned long) hazptr_wildcard - 1].
*/
-static DEFINE_MUTEX(hazptr_wildcard_lock); /* Protect the wildcard flip. */
+static DEFINE_MUTEX(hazptr_phase_lock);
+/* Protect the wildcard and overflow-list phase flips. */
void *hazptr_wildcard = (void *) 1UL;
EXPORT_SYMBOL_GPL(hazptr_wildcard);
+/*
+ * The current overflow list phase. Independent from the wildcard so that the
+ * overflow list scan can be placed after the per-CPU slot scan while still
+ * scanning the non-live list: the current list used by readers is
+ * array[hazptr_overflow_list_phase].
+ */
+static unsigned int hazptr_overflow_list_phase;
+
struct hazptr_overflow_list {
raw_spinlock_t lock; /* Lock protecting overflow list and list generation. */
struct hlist_head head; /* Overflow list head. */
@@ -59,6 +66,12 @@ void *flip_wildcard(void *wildcard)
return ((unsigned long) wildcard == 1UL) ? (void *) 2UL : (void *) 1UL;
}
+static
+unsigned int flip_list_phase(unsigned int phase)
+{
+ return 1 - phase;
+}
+
static
bool is_wildcard(void *addr)
{
@@ -184,15 +197,12 @@ void hazptr_synchronize_cpu_slots(int cpu, void *addr, void *scan_wildcard)
}
static
-void hazptr_scan_period(void *addr, void *scan_wildcard)
+void hazptr_scan_cpu_slots_period(void *addr, void *scan_wildcard)
{
- unsigned int scan_idx = (unsigned long) scan_wildcard - 1;
int cpu;
/* Scan all CPUs slots. */
for_each_possible_cpu(cpu) {
- struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
-
/*
* Scan CPU slots.
* Forward progress against recurring wildcards is guaranteed
@@ -205,12 +215,22 @@ void hazptr_scan_period(void *addr, void *scan_wildcard)
* to acquire that same hazard pointer value.
*/
hazptr_synchronize_cpu_slots(cpu, addr, scan_wildcard);
+ }
+}
+
+static
+void hazptr_scan_overflow_list_period(void *addr, unsigned int scan_idx)
+{
+ int cpu;
+
+ /*
+ * Scan backup slots in percpu overflow lists.
+ * Forward progress is guaranteed by scanning one list
+ * while new elements are added into the other list.
+ */
+ for_each_possible_cpu(cpu) {
+ struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
- /*
- * Scan backup slots in percpu overflow lists.
- * Forward progress is guaranteed by scanning one list
- * while new elements are added into the other list.
- */
hazptr_synchronize_overflow_list(&overflow_list_flip->array[scan_idx], addr);
}
}
@@ -285,10 +305,13 @@ static struct hazptr_scan_state hazptr_scan;
* Walk all slots and return true if @watch is present. If @bloom
* is non-NULL, record observed non-wildcard addresses in it.
*
- * Per-CPU slots are examined before overflow-list slots on each CPU
- * to preserve the acquisition ordering required by the promote path:
- * synchronize must observe the per-CPU slot release before the
- * overflow-list entry can be missed.
+ * All per-CPU slots are examined before any overflow-list slot: a
+ * promote (hazptr_detach() or a context switch) moves a hazard
+ * pointer from a per-CPU slot to an overflow list by chaining the
+ * backup slot before clearing the per-CPU slot, see
+ * hazptr_promote_to_backup_slot(). Therefore the scan must observe
+ * the old location before the new location for every slot/list
+ * pair, including pairs on different CPUs.
*/
static bool hazptr_scan_walk(void *watch, struct hazptr_bloom *bloom)
{
@@ -297,9 +320,9 @@ static bool hazptr_scan_walk(void *watch, struct hazptr_bloom *bloom)
if (bloom)
hazptr_bloom_reset(bloom);
+ /* Scan all per-CPU slots before any overflow-list slot. */
for_each_possible_cpu(cpu) {
struct hazptr_percpu_slots *percpu_slots = per_cpu_ptr(&hazptr_percpu_slots, cpu);
- struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
unsigned int idx;
for (idx = 0; idx < NR_HAZPTR_PERCPU_SLOTS; idx++) {
@@ -313,6 +336,10 @@ static bool hazptr_scan_walk(void *watch, struct hazptr_bloom *bloom)
if (bloom && v && !is_wildcard(v))
hazptr_bloom_add(bloom, v);
}
+ }
+ for_each_possible_cpu(cpu) {
+ struct hazptr_overflow_list_flip *overflow_list_flip = per_cpu_ptr(&percpu_overflow_list_flip, cpu);
+
for (int i = 0; i < 2; i++) {
struct hazptr_overflow_list *list = &overflow_list_flip->array[i];
struct hazptr_backup_slot *b;
@@ -359,14 +386,14 @@ static void hazptr_scan_do_cycle(void)
struct hazptr_waiter *w, *n;
LIST_HEAD(done);
- mutex_lock(&hazptr_wildcard_lock);
+ mutex_lock(&hazptr_phase_lock);
mutex_lock(&hazptr_scan.lock);
list_splice_tail_init(&hazptr_scan.pending, &hazptr_scan.scanning);
mutex_unlock(&hazptr_scan.lock);
if (list_empty(&hazptr_scan.scanning)) {
- mutex_unlock(&hazptr_wildcard_lock);
+ mutex_unlock(&hazptr_phase_lock);
return;
}
@@ -391,7 +418,7 @@ static void hazptr_scan_do_cycle(void)
list_move(&w->node, &done);
}
- mutex_unlock(&hazptr_wildcard_lock);
+ mutex_unlock(&hazptr_phase_lock);
list_for_each_entry_safe(w, n, &done, node) {
list_del_init(&w->node);
@@ -463,6 +490,7 @@ static void hazptr_synchronize_queued(void *addr)
void hazptr_synchronize(void *addr)
{
void *scan_wildcard;
+ unsigned int scan_list_phase;
/*
* Busy-wait should only be done from preemptible context.
@@ -486,18 +514,30 @@ void hazptr_synchronize(void *addr)
}
/* Fallback: use the direct scan path. */
- guard(mutex)(&hazptr_wildcard_lock);
+ guard(mutex)(&hazptr_phase_lock);
scan_wildcard = flip_wildcard(hazptr_wildcard);
- hazptr_scan_period(addr, scan_wildcard);
+ /* Scan per-CPU slots. */
+ hazptr_scan_cpu_slots_period(addr, scan_wildcard);
WRITE_ONCE(hazptr_wildcard, scan_wildcard); /* Flip the current wildcard. */
- hazptr_scan_period(addr, flip_wildcard(scan_wildcard));
+ hazptr_scan_cpu_slots_period(addr, flip_wildcard(scan_wildcard));
+
+ /*
+ * Scan overflow lists *after* scanning all per-CPU slots, see
+ * hazptr_promote_to_backup_slot(). Flip the overflow list phase
+ * between the two lists so that each list is scanned while it is
+ * the non-live one.
+ */
+ scan_list_phase = flip_list_phase(hazptr_overflow_list_phase);
+ hazptr_scan_overflow_list_period(addr, scan_list_phase);
+ WRITE_ONCE(hazptr_overflow_list_phase, scan_list_phase);
+ hazptr_scan_overflow_list_period(addr, flip_list_phase(scan_list_phase));
}
EXPORT_SYMBOL_GPL(hazptr_synchronize);
struct hazptr_slot *hazptr_chain_backup_slot(struct hazptr_ctx *ctx)
{
struct hazptr_overflow_list_flip *overflow_list_flip = this_cpu_ptr(&percpu_overflow_list_flip);
- unsigned int list_idx = (unsigned long) READ_ONCE(hazptr_wildcard) - 1;
+ unsigned int list_idx = READ_ONCE(hazptr_overflow_list_phase);
struct hazptr_overflow_list *overflow_list = &overflow_list_flip->array[list_idx];
struct hazptr_slot *slot = &ctx->backup_slot.slot;
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 04/15] hazptr: add scoped_guard() support
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (2 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 03/15] hazptr: scan all per-CPU slots before overflow lists Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 05/15] hazptr: add debug option to force the acquire slow path Kunwu Chan
` (11 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add a scoped_guard() class for hazptr so that callers can pair
hazptr_acquire() with hazptr_release() on scope exit.
The guard holds a pointer to a caller-provided hazptr_ctx rather
than embedding it. The ctx must remain at a stable address while
the hazard pointer is held because the slow path can chain its
backup slot into the overflow list by address.
Suggested-by: Boqun Feng <boqun@kernel.org>
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
include/linux/hazptr.h | 51 ++++++++++++++++++++++++++++++++++++++++++
1 file changed, 51 insertions(+)
diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
index d1670121947a..275cbe54b075 100644
--- a/include/linux/hazptr.h
+++ b/include/linux/hazptr.h
@@ -320,6 +320,57 @@ void hazptr_release(struct hazptr_ctx *ctx, void *addr)
hazptr_unchain_backup_slot(ctx);
}
+/**
+ * struct hazptr_guard - cleanup.h guard for a single hazard pointer
+ *
+ * @ctx: Pointer to the caller-provided hazard-pointer context.
+ * @addr: The address returned by the matching hazptr_acquire().
+ */
+struct hazptr_guard {
+ struct hazptr_ctx *ctx;
+ void *addr;
+};
+
+/**
+ * hazptr_guard_acquire - Acquire a hazard pointer for scoped_guard() use
+ *
+ * @ctx: The caller-provided hazard-pointer context.
+ * @addr_p: Pointer to the pointer that is to be hazard-pointer protected.
+ *
+ * Acquire a hazard pointer using @ctx and return a guard that
+ * cleanup.h releases via hazptr_release() on scope exit. The guard
+ * only carries the ctx pointer and the address returned by
+ * hazptr_acquire().
+ */
+static inline struct hazptr_guard
+hazptr_guard_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
+{
+ struct hazptr_guard g;
+
+ g.ctx = ctx;
+ g.addr = hazptr_acquire(ctx, addr_p);
+ return g;
+}
+
+DEFINE_CLASS(hazptr, struct hazptr_guard,
+ hazptr_release(_T.ctx, _T.addr),
+ hazptr_guard_acquire(ctx, addr_p),
+ struct hazptr_ctx *ctx, void * const *addr_p)
+DEFINE_CLASS_IS_UNCONDITIONAL(hazptr)
+
+/*
+ * The ctx must remain at a stable address while the hazard pointer
+ * is held because the slow path can chain its backup slot
+ * into the overflow list by address.
+ *
+ * struct hazptr_ctx ctx;
+ * void *obj = ptr;
+ *
+ * scoped_guard(hazptr, &ctx, &obj) {
+ * ...
+ * }
+ */
+
void hazptr_init(void);
#endif /* _LINUX_HAZPTR_H */
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 05/15] hazptr: add debug option to force the acquire slow path
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (3 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 04/15] hazptr: add scoped_guard() support Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 06/15] hazptr: elide redundant first drain pass Kunwu Chan
` (10 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add a debug option that makes hazptr_acquire() always enter
__hazptr_acquire(), even when a per-CPU slot is available.
This allows the slow path, including its lock acquisition, to be
exercised by lockdep-heavy workloads. It also lets hazptrtorture
exercise the slow path without requiring the per-CPU slots to be
exhausted.
Suggested-by: Boqun Feng <boqun@kernel.org>
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
include/linux/hazptr.h | 3 ++-
kernel/rcu/Kconfig.debug | 10 ++++++++++
2 files changed, 12 insertions(+), 1 deletion(-)
diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
index 275cbe54b075..4ee02d23abc9 100644
--- a/include/linux/hazptr.h
+++ b/include/linux/hazptr.h
@@ -245,7 +245,8 @@ void *hazptr_acquire(struct hazptr_ctx *ctx, void * const *addr_p)
ctx->acquire_cpu = smp_processor_id();
ctx->acquire_caller = _THIS_IP_;
#endif
- if (unlikely(slot->addr))
+ if (IS_ENABLED(CONFIG_HAZPTR_ACQUIRE_FORCE_SLOWPATH) ||
+ unlikely(slot->addr))
return __hazptr_acquire(ctx, addr_p);
WRITE_ONCE(slot->addr, READ_ONCE(hazptr_wildcard)); /* Store B */
diff --git a/kernel/rcu/Kconfig.debug b/kernel/rcu/Kconfig.debug
index 7c3c6017b266..3a2e9d7f58f8 100644
--- a/kernel/rcu/Kconfig.debug
+++ b/kernel/rcu/Kconfig.debug
@@ -265,4 +265,14 @@ config HAZPTR_DEBUG
Say Y here if you want to enable those assert, N otherwise.
+config HAZPTR_ACQUIRE_FORCE_SLOWPATH
+ bool "Force hazard-pointer acquire into the slow path"
+ depends on DEBUG_KERNEL
+ default n
+ help
+ Make hazptr_acquire() always enter __hazptr_acquire().
+ This is useful for testing slow-path locking, especially
+ with lockdep. It also lets hazptrtorture exercise the slow
+ path without filling the per-CPU slots.
+
endmenu # "RCU Debugging"
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 06/15] hazptr: elide redundant first drain pass
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (4 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 05/15] hazptr: add debug option to force the acquire slow path Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 07/15] Documentation/litmus-tests: add hazptr wildcard-flip escape test Kunwu Chan
` (9 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Each scan cycle drained both wildcard generations: the
unpublished "other" generation (pass 1) and the pre-flip current
generation (pass 2). Pass 1 is not needed for correctness, so
flip the wildcard first and drain only the old generation while
building the Bloom filter.
An acquire that read the old wildcard before the flip may publish
it into its slot after the drain has passed that slot. Such a
straggling acquire cannot have loaded the pre-unpublish pointer;
see the following LKMM test for the ordering argument.
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
kernel/hazptr.c | 36 +++++++++++++-----------------------
1 file changed, 13 insertions(+), 23 deletions(-)
diff --git a/kernel/hazptr.c b/kernel/hazptr.c
index 799784f698ed..86336f229f57 100644
--- a/kernel/hazptr.c
+++ b/kernel/hazptr.c
@@ -365,24 +365,19 @@ static bool hazptr_scan_walk(void *watch, struct hazptr_bloom *bloom)
}
/*
- * Wait until no per-CPU slot or overflow-list slot holds @wc.
- * Callers must ensure that the wildcard value in use by new acquires
- * differs from @wc, so that the set of slots holding @wc only
- * shrinks, which guarantees forward progress.
- */
-static void hazptr_drain_wildcard(void *wc)
-{
- while (hazptr_scan_walk(wc, NULL))
- cond_resched();
-}
-
-/*
- * Move pending waiters to ->scanning and perform a two-phase
- * wildcard scan shared by all waiters.
+ * Move pending waiters to ->scanning, flip the wildcard, then
+ * drain the old generation while collecting observed addresses
+ * into the Bloom filter. New acquires use the new generation,
+ * so old-generation slots normally only drain.
+ *
+ * An acquire that read the old wildcard before the flip may
+ * publish it after the scanner has passed its slot, but cannot
+ * have loaded the pre-unpublish pointer. See
+ * Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus.
*/
static void hazptr_scan_do_cycle(void)
{
- void *scan_wildcard, *old_wildcard;
+ void *old_wildcard;
struct hazptr_waiter *w, *n;
LIST_HEAD(done);
@@ -397,16 +392,11 @@ static void hazptr_scan_do_cycle(void)
return;
}
- /* Pass 1: drain the unpublished wildcard. */
- scan_wildcard = flip_wildcard(READ_ONCE(hazptr_wildcard));
- hazptr_drain_wildcard(scan_wildcard);
-
- /* Flip so new acquires use the new generation. */
- WRITE_ONCE(hazptr_wildcard, scan_wildcard);
- old_wildcard = flip_wildcard(scan_wildcard);
+ old_wildcard = READ_ONCE(hazptr_wildcard);
+ WRITE_ONCE(hazptr_wildcard, flip_wildcard(old_wildcard));
/*
- * Pass 2: drain the old wildcard while collecting observed
+ * Drain the old wildcard while collecting observed
* addresses into the Bloom filter.
*/
while (hazptr_scan_walk(old_wildcard, &hazptr_scan.bloom))
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 07/15] Documentation/litmus-tests: add hazptr wildcard-flip escape test
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (5 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 06/15] hazptr: elide redundant first drain pass Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 08/15] locking/lockdep: use hazptr to wait for dynamic key lookups Kunwu Chan
` (8 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add a three-process litmus test for the in-flight wildcard window
of the single-pass wildcard-flip scan.
P0 models hazptr_synchronize(): it unpublishes the object and
enqueues the waiter after its memory barrier.
P1 models the scan kthread: it picks up the waiter, flips the
wildcard, and performs the final hazard-pointer walk.
P2 models a straggling hazptr_acquire(): it reads and publishes
the old wildcard, then loads the object after its publication
barrier.
The forbidden outcome is that the final walk misses the
straggler's slot while the straggler still loads the unpublished
object.
Verified with herd7 (linux-kernel.cfg): Never 0 13.
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
.../hazptr/hazptr-wildcard-flip-escape.litmus | 46 +++++++++++++++++++
1 file changed, 46 insertions(+)
create mode 100644 Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus
diff --git a/Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus b/Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus
new file mode 100644
index 000000000000..de467bb82d20
--- /dev/null
+++ b/Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus
@@ -0,0 +1,46 @@
+C hazptr-wildcard-flip-escape
+
+(*
+ * Result: Never
+ *
+ * Check that a straggling acquire which published the old
+ * wildcard after the final walk cannot still load the
+ * unpublished object.
+ *)
+
+{
+ int ptr = 3;
+ int hp = 0;
+ int wildcard = 1;
+ int enq = 0;
+}
+
+P0(int *ptr, int *enq)
+{
+ WRITE_ONCE(*ptr, 0);
+ smp_mb();
+ smp_store_release(enq, 1);
+}
+
+P1(int *enq, int *wildcard, int *hp)
+{
+ int r1;
+ int r2;
+
+ r1 = smp_load_acquire(enq);
+ WRITE_ONCE(*wildcard, 2);
+ r2 = READ_ONCE(*hp);
+}
+
+P2(int *wildcard, int *hp, int *ptr)
+{
+ int r0;
+ int r3;
+
+ r0 = READ_ONCE(*wildcard);
+ WRITE_ONCE(*hp, r0);
+ smp_mb();
+ r3 = READ_ONCE(*ptr);
+}
+
+exists (1:r1=1 /\ 1:r2=0 /\ 2:r3=3)
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 08/15] locking/lockdep: use hazptr to wait for dynamic key lookups
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (6 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 07/15] Documentation/litmus-tests: add hazptr wildcard-flip escape test Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 09/15] rcuscale: add hazptr scale type Kunwu Chan
` (7 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
lockdep_unregister_key() waits for is_dynamic_key() callers with
synchronize_rcu_expedited(), which requires system-wide coordination.
Make is_dynamic_key() protect the hash bucket with a hazard pointer
and use hazptr_synchronize() to wait for traversals of that bucket.
The hash bucket returned by keyhashentry() has a stable address, so
it can be used as the hazptr synchronization target. Keep the key
hashlist lifetime RCU-based with hlist_del_rcu() and call_rcu().
Protect the traversal with the hazptr scoped_guard() added earlier
in this series.
This adapts the lockdep use case from Boqun Feng's shazptr series
to the current hazptr API.
Link: https://lore.kernel.org/lkml/20250625031101.12555-1-boqun.feng@gmail.com/
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
kernel/locking/lockdep.c | 25 +++++++++++++++----------
1 file changed, 15 insertions(+), 10 deletions(-)
diff --git a/kernel/locking/lockdep.c b/kernel/locking/lockdep.c
index c56a7f91d72e..d1ce5ace7fbf 100644
--- a/kernel/locking/lockdep.c
+++ b/kernel/locking/lockdep.c
@@ -58,6 +58,7 @@
#include <linux/context_tracking.h>
#include <linux/console.h>
#include <linux/kasan.h>
+#include <linux/hazptr.h>
#include <asm/sections.h>
@@ -1280,14 +1281,19 @@ static bool is_dynamic_key(const struct lock_class_key *key)
hash_head = keyhashentry(key);
- rcu_read_lock();
- hlist_for_each_entry_rcu(k, hash_head, hash_entry) {
- if (k == key) {
- found = true;
- break;
+ {
+ struct hazptr_ctx ctx;
+ void *bucket = hash_head;
+
+ scoped_guard(hazptr, &ctx, &bucket) {
+ hlist_for_each_entry_rcu(k, hash_head, hash_entry, 1) {
+ if (k == key) {
+ found = true;
+ break;
+ }
+ }
}
}
- rcu_read_unlock();
return found;
}
@@ -6683,11 +6689,10 @@ void lockdep_unregister_key(struct lock_class_key *key)
*
* Some operations like __qdisc_destroy() will call this in a debug
* kernel, and the network traffic is disabled while waiting, hence
- * the delay of the wait matters in debugging cases. Currently use a
- * synchronize_rcu_expedited() to speed up the wait at the cost of
- * system IPIs. TODO: Replace RCU with hazptr for this.
+ * the delay of the wait matters in debugging cases. Replace the
+ * expedited RCU wait with hazptr_synchronize().
*/
- synchronize_rcu_expedited();
+ hazptr_synchronize(keyhashentry(key));
}
EXPORT_SYMBOL_GPL(lockdep_unregister_key);
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 09/15] rcuscale: add hazptr scale type
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (7 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 08/15] locking/lockdep: use hazptr to wait for dynamic key lookups Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 10/15] hazptr: fix kernel-doc of hazptr_release() Kunwu Chan
` (6 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add hazptr reader and synchronize operations so that synchronize
latency can be measured alongside RCU and SRCU.
Use a small set of objects so readers do not always protect the
same target. Each reader protects the object selected by its CPU,
while the updater round-robins through the objects.
Use hazptr_synchronize() for both normal and expedited scale
operations, since hazptr has no separate expedited operation.
Link: https://lore.kernel.org/lkml/20250625031101.12555-1-boqun.feng@gmail.com/
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
kernel/rcu/rcuscale.c | 70 ++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 69 insertions(+), 1 deletion(-)
diff --git a/kernel/rcu/rcuscale.c b/kernel/rcu/rcuscale.c
index 1097ec15879c..5aca31c3c01d 100644
--- a/kernel/rcu/rcuscale.c
+++ b/kernel/rcu/rcuscale.c
@@ -39,6 +39,7 @@
#include <linux/torture.h>
#include <linux/vmalloc.h>
#include <linux/rcupdate_trace.h>
+#include <linux/hazptr.h>
#include <linux/sched/debug.h>
#include "rcu.h"
@@ -418,6 +419,71 @@ static struct rcu_scale_ops tasks_tracing_ops = {
#endif // #else // #ifdef CONFIG_TASKS_TRACE_RCU
+#if IS_ENABLED(CONFIG_HAZPTR_TORTURE_TEST)
+
+#define HAZPTR_SCALE_NR_OBJS 8
+static int hazptr_scale_obj[HAZPTR_SCALE_NR_OBJS];
+static atomic_t hazptr_scale_sync_idx = ATOMIC_INIT(0);
+
+struct hazptr_scale_state {
+ struct hazptr_ctx ctx;
+ void *addr;
+};
+static DEFINE_PER_CPU(struct hazptr_scale_state, hazptr_scale_state);
+
+static int hazptr_scale_read_lock(void)
+{
+ struct hazptr_scale_state *state = this_cpu_ptr(&hazptr_scale_state);
+ int idx = smp_processor_id() % HAZPTR_SCALE_NR_OBJS;
+ void *ptr = &hazptr_scale_obj[idx];
+
+ state->addr = hazptr_acquire(&state->ctx, &ptr);
+ return 0;
+}
+
+static void hazptr_scale_read_unlock(int idx)
+{
+ struct hazptr_scale_state *state = this_cpu_ptr(&hazptr_scale_state);
+
+ hazptr_release(&state->ctx, state->addr);
+}
+
+static unsigned long hazptr_scale_completed(void)
+{
+ return 0;
+}
+
+static void hazptr_scale_sync(void)
+{
+ int idx = atomic_fetch_inc(&hazptr_scale_sync_idx) % HAZPTR_SCALE_NR_OBJS;
+
+ hazptr_synchronize(&hazptr_scale_obj[idx]);
+}
+
+static void hazptr_scale_sync_exp(void)
+{
+ int idx = atomic_fetch_inc(&hazptr_scale_sync_idx) % HAZPTR_SCALE_NR_OBJS;
+
+ hazptr_synchronize(&hazptr_scale_obj[idx]);
+}
+
+static struct rcu_scale_ops hazptr_scale_ops = {
+ .ptype = 0,
+ .readlock = hazptr_scale_read_lock,
+ .readunlock = hazptr_scale_read_unlock,
+ .get_gp_seq = hazptr_scale_completed,
+ .gp_diff = NULL,
+ .exp_completed = hazptr_scale_completed,
+ .sync = hazptr_scale_sync,
+ .exp_sync = hazptr_scale_sync_exp,
+ .name = "hazptr",
+};
+
+#define HAZPTR_SCALE_OPS &hazptr_scale_ops,
+#else
+#define HAZPTR_SCALE_OPS
+#endif
+
static unsigned long rcuscale_seq_diff(unsigned long new, unsigned long old)
{
if (!cur_ops->gp_diff)
@@ -1110,7 +1176,9 @@ rcu_scale_init(void)
long i;
long j;
static struct rcu_scale_ops *scale_ops[] = {
- &rcu_ops, &srcu_ops, &srcud_ops, TASKS_OPS TASKS_RUDE_OPS TASKS_TRACING_OPS
+ &rcu_ops, &srcu_ops, &srcud_ops,
+ TASKS_OPS TASKS_RUDE_OPS TASKS_TRACING_OPS
+ HAZPTR_SCALE_OPS
};
if (!torture_init_begin(scale_type, verbose))
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 10/15] hazptr: fix kernel-doc of hazptr_release()
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (8 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 09/15] rcuscale: add hazptr scale type Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 11/15] Documentation/litmus-tests: add hazptr acquire-before-scan test Kunwu Chan
` (5 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
The second parameter of hazptr_release() is named @addr, not
@addr_p. It is the address passed to the matching
hazptr_acquire().
Fix the kernel-doc header to match the API.
Fixes: 17969a0e7f7e ("hazptr: Upgrade kernel-doc headers")
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
include/linux/hazptr.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/include/linux/hazptr.h b/include/linux/hazptr.h
index 4ee02d23abc9..07f5d9ef5a6d 100644
--- a/include/linux/hazptr.h
+++ b/include/linux/hazptr.h
@@ -293,7 +293,7 @@ static inline void hazptr_release_debug(struct hazptr_ctx *ctx, void *addr) { }
* hazptr_release - Release the specified hazard pointer
*
* @ctx: The hazard-pointer context that was passed to hazptr_acquire().
- * @addr_p: The pointer that is to be hazard-pointer unprotected.
+ * @addr: The address passed to the matching hazptr_acquire().
*
* Release the protected hazard pointer recorded in @ctx.
*
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 11/15] Documentation/litmus-tests: add hazptr acquire-before-scan test
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (9 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 10/15] hazptr: fix kernel-doc of hazptr_release() Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 12/15] hazptrtorture: add slowpath and lockdep scenarios Kunwu Chan
` (4 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add an LKMM test for the hazard-pointer publication protocol used
by the lockdep conversion.
The test checks that a scan cannot miss a hazard-pointer
publication while the reader still observes the pre-unpublish
pointer. The smp_mb() pair provides the required ordering.
This covers the resolved-publication case. The in-flight wildcard
window of the wildcard-flip scan is covered by
hazptr-wildcard-flip-escape.litmus.
Verified with herd7: Never 0 3.
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
Documentation/litmus-tests/README | 13 ++++++
.../hazptr/hazptr-acquire-before-scan.litmus | 45 +++++++++++++++++++
2 files changed, 58 insertions(+)
create mode 100644 Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus
diff --git a/Documentation/litmus-tests/README b/Documentation/litmus-tests/README
index 4d4ec9c6f2cc..c2bf20fa4092 100644
--- a/Documentation/litmus-tests/README
+++ b/Documentation/litmus-tests/README
@@ -97,3 +97,16 @@ SRCU-fastpath-scan-before-anchor.litmus
period, violating the SRCU grace-period guarantee. See
SRCU-fastpath-anchor-before-scan.litmus for the opposite
ordering.
+
+
+hazptr (/hazptr directory)
+--------------------------
+
+hazptr-acquire-before-scan.litmus
+ Test the resolved-publication ordering between hazard-pointer
+ publication and the scan.
+
+hazptr-wildcard-flip-escape.litmus
+ Test the in-flight wildcard window of the single-pass wildcard-flip
+ scan: a straggling acquire that published the old wildcard after
+ the final walk cannot still load the unpublished object.
diff --git a/Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus b/Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus
new file mode 100644
index 000000000000..6dc13830c07c
--- /dev/null
+++ b/Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus
@@ -0,0 +1,45 @@
+C hazptr-acquire-before-scan
+
+(*
+ * Result: Never
+ *
+ * The reclaimer unpublishes the pointer, executes smp_mb(), then
+ * scans the hazard-pointer slot. The reader publishes the
+ * protected address, executes smp_mb(), then loads the pointer.
+ *
+ * The smp_mb() pair forbids the reclaimer from missing the
+ * publication while the reader still observes the pre-unpublish
+ * pointer.
+ *
+ * This is the publication protocol used by the lockdep
+ * is_dynamic_key()/lockdep_unregister_key() conversion.
+ *
+ * This models the resolved-publication case. The in-flight
+ * wildcard window of the wildcard-flip scan is covered by
+ * hazptr-wildcard-flip-escape.litmus.
+ *)
+
+{
+int ptr = 1;
+int hp = 0;
+}
+
+P0(int *ptr, int *hp)
+{
+ int r0;
+
+ WRITE_ONCE(*ptr, 0); /* Unpublish. */
+ smp_mb();
+ r0 = READ_ONCE(*hp); /* Scan. */
+}
+
+P1(int *ptr, int *hp)
+{
+ int r0;
+
+ WRITE_ONCE(*hp, 1); /* Publish. */
+ smp_mb();
+ r0 = READ_ONCE(*ptr); /* Load pointer. */
+}
+
+exists (0:r0=0 /\ 1:r0=1)
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 12/15] hazptrtorture: add slowpath and lockdep scenarios
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (10 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 11/15] Documentation/litmus-tests: add hazptr acquire-before-scan test Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 13/15] hazptrtorture: add READERS4 and READERS0 torture configs Kunwu Chan
` (3 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add a wq_churn kthread and module parameter to exercise lockdep
dynamic-key handling by repeatedly creating and destroying
workqueues. Each workqueue creation registers a dynamic lock
class key and each destruction unregisters it, exercising both
is_dynamic_key() and lockdep_unregister_key().
Add SLOWPATH and LOCKDEP torture scenarios to exercise the
hazptr acquire slow path, including its use by lockdep under
CONFIG_PROVE_LOCKING.
Add both scenarios to CFLIST so that torture.sh --do-hazptr
includes them.
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
kernel/rcu/hazptrtorture.c | 57 +++++++++++++++++--
.../rcutorture/configs/hazptr/CFLIST | 2 +
.../rcutorture/configs/hazptr/LOCKDEP | 17 ++++++
.../rcutorture/configs/hazptr/LOCKDEP.boot | 1 +
.../rcutorture/configs/hazptr/SLOWPATH | 16 ++++++
5 files changed, 89 insertions(+), 4 deletions(-)
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH
diff --git a/kernel/rcu/hazptrtorture.c b/kernel/rcu/hazptrtorture.c
index 7c8b5899fb01..7c144b9bca6a 100644
--- a/kernel/rcu/hazptrtorture.c
+++ b/kernel/rcu/hazptrtorture.c
@@ -23,6 +23,7 @@
#include <linux/torture.h>
#include <linux/hazptr.h>
#include <linux/rcupdate.h>
+#include <linux/workqueue.h>
#include "rcu.h"
@@ -49,6 +50,8 @@ torture_param(int, shutdown_secs, 0, "Shutdown time (s), <= zero to disable.");
torture_param(int, stat_interval, 60, "Number of seconds between stats printk()s");
torture_param(int, stutter, 5, "Number of seconds to run/halt test");
torture_param(int, verbose, 1, "Enable verbose debugging printk()s");
+torture_param(bool, wq_churn, false,
+ "Churn workqueues to exercise lockdep dynamic keys");
static char *torture_type = "hazptr";
module_param(torture_type, charp, 0444);
@@ -60,6 +63,7 @@ static struct task_struct *preempt_task;
static struct task_struct **reader_tasks;
static struct task_struct *do_pending_task;
static struct task_struct *stats_task;
+static struct task_struct *wq_churn_task;
#define HAZPTR_TORTURE_PIPE_LEN 10
@@ -87,6 +91,7 @@ static DEFINE_PER_CPU(atomic_long_t, hazptr_torture_acquires_irq);
static DEFINE_PER_CPU(atomic_long_t, hazptr_torture_releases_irq);
static DEFINE_PER_CPU(atomic_long_t, hazptr_torture_releases_defer);
static DEFINE_PER_CPU(atomic_long_t, hazptr_torture_releases_undefer);
+static atomic_t n_hazptr_torture_wq_churn;
static struct list_head hazptr_torture_removed;
// State for a deferred (AKA pending) hazard pointer
@@ -643,13 +648,14 @@ hazptr_torture_stats_print(void)
atomic_read(&n_hazptr_torture_alloc_fail),
atomic_read(&n_hazptr_torture_free));
torture_onoff_stats();
- pr_cont("acq: %lld rel: %lld acqirq: %lld relirq: %lld reldefer: %lld relundefer %lld\n",
+ pr_cont("acq: %lld rel: %lld acqirq: %lld relirq: %lld reldefer: %lld relundefer %lld wqchurn: %d\n",
torture_sum_pcpu_atomic_long(&hazptr_torture_acquires),
torture_sum_pcpu_atomic_long(&hazptr_torture_releases),
torture_sum_pcpu_atomic_long(&hazptr_torture_acquires_irq),
torture_sum_pcpu_atomic_long(&hazptr_torture_releases_irq),
torture_sum_pcpu_atomic_long(&hazptr_torture_releases_defer),
- torture_sum_pcpu_atomic_long(&hazptr_torture_releases_undefer));
+ torture_sum_pcpu_atomic_long(&hazptr_torture_releases_undefer),
+ atomic_read(&n_hazptr_torture_wq_churn));
pr_alert("%s%s ", torture_type, TORTURE_FLAG);
if (i > 1) {
@@ -720,14 +726,14 @@ hazptr_torture_print_module_parms(struct hazptr_torture_ops *cur_ops, const char
"defer_modulus=%d irq_acquire=%d irq_release=%d kthread_do_pending_ms=%d "
"onoff_interval=%d onoff_holdoff=%d "
"preempt_duration=%d preempt_interval=%d "
- "reader_sleep_us=%d "
+ "reader_sleep_us=%d wq_churn=%d "
"shuffle_interval=%d shutdown_secs=%d stat_interval=%d stutter=%d "
"verbose=%d\n",
torture_type, tag, nrealreaders, nwriters,
defer_modulus, irq_acquire, irq_release, kthread_do_pending_ms,
onoff_interval, onoff_holdoff,
preempt_duration, preempt_interval,
- reader_sleep_us,
+ reader_sleep_us, wq_churn,
shuffle_interval, shutdown_secs, stat_interval, stutter,
verbose);
}
@@ -761,6 +767,42 @@ static int hazptr_torture_preempt(void *unused)
return 0;
}
+static void hazptr_torture_wq_churn_fn(struct work_struct *work)
+{
+}
+
+/*
+ * Repeatedly create and destroy workqueues. Each workqueue creation
+ * registers a lockdep dynamic key and each destruction unregisters it,
+ * which exercises both is_dynamic_key() and lockdep_unregister_key(),
+ * the two lockdep paths that use hazptr. Most useful in kernels with
+ * CONFIG_PROVE_LOCKING=y and CONFIG_HAZPTR_ACQUIRE_FORCE_SLOWPATH=y.
+ */
+static int hazptr_torture_wq_churn(void *unused)
+{
+ DEFINE_TORTURE_RANDOM(rand);
+
+ VERBOSE_TOROUT_STRING("hazptr_torture_wq_churn task started");
+ do {
+ struct workqueue_struct *wq;
+ struct work_struct churn_work;
+
+ wq = alloc_workqueue("hazptr-torture-churn", WQ_PERCPU, 1);
+ if (!wq) {
+ torture_hrtimeout_ms(10, 10, &rand);
+ continue;
+ }
+ INIT_WORK(&churn_work, hazptr_torture_wq_churn_fn);
+ queue_work(wq, &churn_work);
+ flush_work(&churn_work);
+ destroy_workqueue(wq);
+ atomic_inc(&n_hazptr_torture_wq_churn);
+ stutter_wait("hazptr_torture_wq_churn");
+ } while (!torture_must_stop());
+ torture_kthread_stopping("hazptr_torture_wq_churn");
+ return 0;
+}
+
static void
hazptr_torture_cleanup(void)
{
@@ -775,6 +817,7 @@ hazptr_torture_cleanup(void)
torture_stop_kthread(hazptr_torture_do_pending, do_pending_task);
torture_stop_kthread(hazptr_torture_preempt, preempt_task);
+ torture_stop_kthread(hazptr_torture_wq_churn, wq_churn_task);
torture_stop_kthread(hazptr_torture_writer, writer_task);
if (reader_tasks) {
@@ -853,6 +896,7 @@ static int __init hazptr_torture_init(void)
atomic_set(&n_hazptr_torture_alloc_fail, 0);
atomic_set(&n_hazptr_torture_free, 0);
atomic_set(&n_hazptr_torture_error, 0);
+ atomic_set(&n_hazptr_torture_wq_churn, 0);
for (i = 0; i < HAZPTR_TORTURE_PIPE_LEN + 1; i++)
atomic_set(&hazptr_torture_wcount[i], 0);
for_each_possible_cpu(cpu) {
@@ -928,6 +972,11 @@ static int __init hazptr_torture_init(void)
if (torture_init_error(firsterr))
goto unwind;
}
+ if (wq_churn) {
+ firsterr = torture_create_kthread(hazptr_torture_wq_churn, NULL, wq_churn_task);
+ if (torture_init_error(firsterr))
+ goto unwind;
+ }
torture_init_end();
return 0;
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST b/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
index 4d62eb4a39f9..380b8402ccc9 100644
--- a/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
@@ -1,2 +1,4 @@
NOPREEMPT
PREEMPT
+SLOWPATH
+LOCKDEP
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP b/tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP
new file mode 100644
index 000000000000..129769e2e32e
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP
@@ -0,0 +1,17 @@
+CONFIG_SMP=y
+CONFIG_NR_CPUS=16
+CONFIG_PREEMPT_NONE=n
+CONFIG_PREEMPT_VOLUNTARY=n
+CONFIG_PREEMPT=y
+CONFIG_HZ_PERIODIC=n
+CONFIG_NO_HZ_IDLE=y
+CONFIG_NO_HZ_FULL=n
+CONFIG_HOTPLUG_CPU=y
+CONFIG_SUSPEND=n
+CONFIG_HIBERNATION=n
+CONFIG_DEBUG_LOCK_ALLOC=y
+CONFIG_PROVE_LOCKING=y
+CONFIG_DEBUG_OBJECTS_RCU_HEAD=n
+CONFIG_DEBUG_KERNEL=y
+CONFIG_HAZPTR_DEBUG=y
+CONFIG_HAZPTR_ACQUIRE_FORCE_SLOWPATH=y
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot b/tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot
new file mode 100644
index 000000000000..7b16720214bd
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot
@@ -0,0 +1 @@
+hazptrtorture.wq_churn=1
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH b/tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH
new file mode 100644
index 000000000000..4d13028ca84c
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH
@@ -0,0 +1,16 @@
+CONFIG_SMP=y
+CONFIG_NR_CPUS=16
+CONFIG_PREEMPT_NONE=n
+CONFIG_PREEMPT_VOLUNTARY=n
+CONFIG_PREEMPT=y
+CONFIG_HZ_PERIODIC=n
+CONFIG_NO_HZ_IDLE=y
+CONFIG_NO_HZ_FULL=n
+CONFIG_HOTPLUG_CPU=y
+CONFIG_SUSPEND=n
+CONFIG_HIBERNATION=n
+CONFIG_PROVE_LOCKING=n
+CONFIG_DEBUG_OBJECTS_RCU_HEAD=n
+CONFIG_DEBUG_KERNEL=y
+CONFIG_HAZPTR_DEBUG=y
+CONFIG_HAZPTR_ACQUIRE_FORCE_SLOWPATH=y
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 13/15] hazptrtorture: add READERS4 and READERS0 torture configs
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (11 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 12/15] hazptrtorture: add slowpath and lockdep scenarios Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 14/15] hazptrtorture: add 128- and 256-CPU configs Kunwu Chan
` (2 subsequent siblings)
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add READERS4 and READERS0 configs to cover reader-active and
reader-free synchronize paths.
Both use the PREEMPT baseline and hazptr-stack torture type,
with four readers for READERS4 and no readers for READERS0.
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
.../selftests/rcutorture/configs/hazptr/CFLIST | 2 ++
.../selftests/rcutorture/configs/hazptr/READERS0 | 16 ++++++++++++++++
.../rcutorture/configs/hazptr/READERS0.boot | 2 ++
.../selftests/rcutorture/configs/hazptr/READERS4 | 16 ++++++++++++++++
.../rcutorture/configs/hazptr/READERS4.boot | 2 ++
5 files changed, 38 insertions(+)
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST b/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
index 380b8402ccc9..1826be30c066 100644
--- a/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
@@ -1,4 +1,6 @@
NOPREEMPT
PREEMPT
+READERS0
+READERS4
SLOWPATH
LOCKDEP
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/READERS0 b/tools/testing/selftests/rcutorture/configs/hazptr/READERS0
new file mode 100644
index 000000000000..c5b12f3bef74
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/READERS0
@@ -0,0 +1,16 @@
+CONFIG_SMP=y
+CONFIG_NR_CPUS=16
+CONFIG_PREEMPT_NONE=n
+CONFIG_PREEMPT_VOLUNTARY=n
+CONFIG_PREEMPT=y
+CONFIG_HZ_PERIODIC=n
+CONFIG_NO_HZ_IDLE=y
+CONFIG_NO_HZ_FULL=n
+CONFIG_HOTPLUG_CPU=y
+CONFIG_SUSPEND=n
+CONFIG_HIBERNATION=n
+CONFIG_DEBUG_LOCK_ALLOC=n
+CONFIG_PROVE_LOCKING=n
+CONFIG_DEBUG_OBJECTS_RCU_HEAD=n
+CONFIG_DEBUG_KERNEL=y
+CONFIG_HAZPTR_DEBUG=y
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot b/tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot
new file mode 100644
index 000000000000..f5564c218902
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot
@@ -0,0 +1,2 @@
+hazptrtorture.nreaders=0
+hazptrtorture.torture_type=hazptr-stack
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/READERS4 b/tools/testing/selftests/rcutorture/configs/hazptr/READERS4
new file mode 100644
index 000000000000..c5b12f3bef74
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/READERS4
@@ -0,0 +1,16 @@
+CONFIG_SMP=y
+CONFIG_NR_CPUS=16
+CONFIG_PREEMPT_NONE=n
+CONFIG_PREEMPT_VOLUNTARY=n
+CONFIG_PREEMPT=y
+CONFIG_HZ_PERIODIC=n
+CONFIG_NO_HZ_IDLE=y
+CONFIG_NO_HZ_FULL=n
+CONFIG_HOTPLUG_CPU=y
+CONFIG_SUSPEND=n
+CONFIG_HIBERNATION=n
+CONFIG_DEBUG_LOCK_ALLOC=n
+CONFIG_PROVE_LOCKING=n
+CONFIG_DEBUG_OBJECTS_RCU_HEAD=n
+CONFIG_DEBUG_KERNEL=y
+CONFIG_HAZPTR_DEBUG=y
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot b/tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot
new file mode 100644
index 000000000000..0d0df3f7dca4
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot
@@ -0,0 +1,2 @@
+hazptrtorture.nreaders=4
+hazptrtorture.torture_type=hazptr-stack
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 14/15] hazptrtorture: add 128- and 256-CPU configs
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (12 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 13/15] hazptrtorture: add READERS4 and READERS0 torture configs Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 15/15] selftests/rcutorture: add hazptr torture test script Kunwu Chan
2026-10-02 17:12 ` [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Bradley Morgan
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add CPU128 and CPU256 configs for hazptr-stack torture testing.
These configs exercise the shared-scan path at larger CPU counts.
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
.../selftests/rcutorture/configs/hazptr/CFLIST | 2 ++
.../selftests/rcutorture/configs/hazptr/CPU128 | 16 ++++++++++++++++
.../rcutorture/configs/hazptr/CPU128.boot | 1 +
.../selftests/rcutorture/configs/hazptr/CPU256 | 16 ++++++++++++++++
.../rcutorture/configs/hazptr/CPU256.boot | 1 +
5 files changed, 36 insertions(+)
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256
create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST b/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
index 1826be30c066..90ef09b0e654 100644
--- a/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/CFLIST
@@ -4,3 +4,5 @@ READERS0
READERS4
SLOWPATH
LOCKDEP
+CPU128
+CPU256
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/CPU128 b/tools/testing/selftests/rcutorture/configs/hazptr/CPU128
new file mode 100644
index 000000000000..079b463607e8
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/CPU128
@@ -0,0 +1,16 @@
+CONFIG_SMP=y
+CONFIG_NR_CPUS=128
+CONFIG_PREEMPT_NONE=n
+CONFIG_PREEMPT_VOLUNTARY=n
+CONFIG_PREEMPT=y
+CONFIG_HZ_PERIODIC=n
+CONFIG_NO_HZ_IDLE=y
+CONFIG_NO_HZ_FULL=n
+CONFIG_HOTPLUG_CPU=y
+CONFIG_SUSPEND=n
+CONFIG_HIBERNATION=n
+CONFIG_DEBUG_LOCK_ALLOC=n
+CONFIG_PROVE_LOCKING=n
+CONFIG_DEBUG_OBJECTS_RCU_HEAD=n
+CONFIG_DEBUG_KERNEL=y
+CONFIG_HAZPTR_DEBUG=y
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot b/tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot
new file mode 100644
index 000000000000..1d09a6446080
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot
@@ -0,0 +1 @@
+hazptrtorture.torture_type=hazptr-stack
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/CPU256 b/tools/testing/selftests/rcutorture/configs/hazptr/CPU256
new file mode 100644
index 000000000000..b00a984977f3
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/CPU256
@@ -0,0 +1,16 @@
+CONFIG_SMP=y
+CONFIG_NR_CPUS=256
+CONFIG_PREEMPT_NONE=n
+CONFIG_PREEMPT_VOLUNTARY=n
+CONFIG_PREEMPT=y
+CONFIG_HZ_PERIODIC=n
+CONFIG_NO_HZ_IDLE=y
+CONFIG_NO_HZ_FULL=n
+CONFIG_HOTPLUG_CPU=y
+CONFIG_SUSPEND=n
+CONFIG_HIBERNATION=n
+CONFIG_DEBUG_LOCK_ALLOC=n
+CONFIG_PROVE_LOCKING=n
+CONFIG_DEBUG_OBJECTS_RCU_HEAD=n
+CONFIG_DEBUG_KERNEL=y
+CONFIG_HAZPTR_DEBUG=y
diff --git a/tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot b/tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot
new file mode 100644
index 000000000000..1d09a6446080
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot
@@ -0,0 +1 @@
+hazptrtorture.torture_type=hazptr-stack
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 15/15] selftests/rcutorture: add hazptr torture test script
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (13 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 14/15] hazptrtorture: add 128- and 256-CPU configs Kunwu Chan
@ 2026-10-02 17:08 ` Kunwu Chan
2026-10-02 17:12 ` [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Bradley Morgan
15 siblings, 0 replies; 17+ messages in thread
From: Kunwu Chan @ 2026-10-02 17:08 UTC (permalink / raw)
To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
Add hazptr.sh as a one-command regression test for the hazptr
torture configs.
Run NOPREEMPT, PREEMPT, READERS0, and READERS4, followed by
LOCKDEP with wq_churn and SLOWPATH with CONFIG_PROVE_LOCKING=y.
The latter two also check the console for lockdep warnings.
With --allcpus, repeat the PREEMPT test at CPU counts 8, 16, 32,
and so on up to 256.
Follow the pattern of the existing srcu_lockdep.sh script.
Signed-off-by: Kunwu Chan <kunwu.chan@gmail.com>
---
.../selftests/rcutorture/bin/hazptr.sh | 146 ++++++++++++++++++
1 file changed, 146 insertions(+)
create mode 100755 tools/testing/selftests/rcutorture/bin/hazptr.sh
diff --git a/tools/testing/selftests/rcutorture/bin/hazptr.sh b/tools/testing/selftests/rcutorture/bin/hazptr.sh
new file mode 100755
index 000000000000..ef636e519a55
--- /dev/null
+++ b/tools/testing/selftests/rcutorture/bin/hazptr.sh
@@ -0,0 +1,146 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0+
+#
+# Run hazptr torture tests and report any that fail.
+#
+# Usage: hazptr.sh [ --cpus N ] [ --allcpus ] [ --datestamp string ]
+#
+# Copyright (C) 2026 Kunwu Chan
+#
+# Authors: Kunwu Chan <kunwu.chan@gmail.com>
+
+usage () {
+ echo "Usage: $scriptname optional arguments:"
+ echo " --cpus N override CPU count (default: auto-detect)"
+ echo " --allcpus test at CPU counts 8,16,32,...,256"
+ echo " --datestamp string"
+ exit 1
+}
+
+ds=`date +%Y.%m.%d-%H.%M.%S`-hazptr
+scriptname="$0"
+
+RCUTORTURE="`pwd`/tools/testing/selftests/rcutorture"; export RCUTORTURE
+PATH=${RCUTORTURE}/bin:$PATH; export PATH
+. functions.sh
+
+cpus="" # empty = let kvm.sh auto-detect
+allcpus=0
+
+while test $# -gt 0
+do
+ case "$1" in
+ --datestamp)
+ checkarg --datestamp "(relative pathname)" "$#" "$2" '^[a-zA-Z0-9._/-]*$' '^--'
+ ds=$2
+ shift
+ ;;
+ --cpus)
+ checkarg --cpus "(number)" "$#" "$2" '^[0-9]\+$' '^--'
+ cpus="$2"
+ shift
+ ;;
+ --allcpus)
+ allcpus=1
+ ;;
+ *)
+ echo Unknown argument $1
+ usage
+ ;;
+ esac
+ shift
+done
+
+nerrs=0
+
+run_kvm () {
+ local ds_suffix="$1"
+ local configs="$2"
+ local cpu_arg="$3"
+ local kconfig="$4"
+ local resdir="$RCUTORTURE/res/$ds/$ds_suffix"
+
+ mkdir -p "$resdir"
+
+ tools/testing/selftests/rcutorture/bin/kvm.sh \
+ --torture hazptr \
+ --duration 5s \
+ --configs "$configs" \
+ ${cpu_arg:+--cpus "$cpu_arg"} \
+ ${kconfig:+--kconfig "$kconfig"} \
+ --trust-make \
+ --datestamp "$ds/$ds_suffix" > "$resdir/kvm.sh.out" 2>&1
+ return $?
+}
+
+check_lockdep () {
+ local ds_suffix="$1" config_dir="$2" label="$3"
+ local resdir="$RCUTORTURE/res/$ds/$ds_suffix"
+ local console="$resdir/$config_dir/console.log"
+
+ if ! test -f "$console"
+ then
+ echo "Missing console.log for $label ($ds_suffix)" > "$resdir/kvm.sh.err"
+ return 1
+ fi
+ if grep -qE "WARNING: possible (recursive locking|circular locking dependency)" "$console"
+ then
+ echo "Lockdep warning in $label" > "$resdir/kvm.sh.err"
+ return 1
+ fi
+ return 0
+}
+
+# Basic torture: all hazptr configs, CPU count auto-detected
+
+echo "--- hazptr basic torture${cpus:+ ($cpus CPUs)} ---"
+run_kvm basic "NOPREEMPT PREEMPT READERS0 READERS4" "$cpus" || nerrs=$((nerrs+1))
+
+# LOCKDEP: verify no lockdep splats with PROVE_LOCKING + wq_churn
+
+echo "--- hazptr lockdep + wq_churn ---"
+if run_kvm lockdep-wq LOCKDEP "$cpus"
+then
+ check_lockdep lockdep-wq LOCKDEP "LOCKDEP wq_churn" ||
+ nerrs=$((nerrs+1))
+else
+ nerrs=$((nerrs+1))
+fi
+
+# SLOWPATH + PROVE_LOCKING: verify forced-slowpath does not deadlock
+
+echo "--- hazptr SLOWPATH + lockdep ---"
+if run_kvm slowpath-lockdep SLOWPATH "$cpus" \
+ "CONFIG_PROVE_LOCKING=y"
+then
+ check_lockdep slowpath-lockdep SLOWPATH \
+ "SLOWPATH+PROVE_LOCKING" ||
+ nerrs=$((nerrs+1))
+else
+ nerrs=$((nerrs+1))
+fi
+
+# CPU scaling: optionally test at multiple CPU counts
+
+if test "$allcpus" -ne 0
+then
+ max_cpus=`grep '^processor' /proc/cpuinfo 2>/dev/null | wc -l`
+ for n in 8 16 32 64 128 256
+ do
+ if test "$n" -le "$max_cpus"
+ then
+ echo "--- hazptr CPU=$n ---"
+ run_kvm "cpu$n" "PREEMPT" "$n" "CONFIG_NR_CPUS=$n" || nerrs=$((nerrs+1))
+ fi
+ done
+fi
+
+# Report
+
+if test "$nerrs" -ne 0
+then
+ echo "hazptr.sh: $nerrs test(s) failed"
+ exit 1
+fi
+echo "hazptr.sh: all tests passed"
+exit 0
--
2.43.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
` (14 preceding siblings ...)
2026-10-02 17:08 ` [PATCH RFC v2 15/15] selftests/rcutorture: add hazptr torture test script Kunwu Chan
@ 2026-10-02 17:12 ` Bradley Morgan
15 siblings, 0 replies; 17+ messages in thread
From: Bradley Morgan @ 2026-10-02 17:12 UTC (permalink / raw)
To: Kunwu Chan, paulmck, corbet, mingo, frederic, neeraj.upadhyay,
josh, urezki, dave, lianux.mm
Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
qiang.zhang, kunwu.chan, linux-kernel, linux-arch, lkmm,
linux-doc, rcu, linux-kselftest
On 2 October 2026 18:08:32 BST, Kunwu Chan <kunwu.chan@gmail.com> wrote:
>This series extends Paul McKenney's v3 hazptr implementation [1]
>with a shared scan path for concurrent hazptr_synchronize() callers,
>and adapts the lockdep dynamic-key hashlist use case from Boqun
>Feng's 2025 shazptr series [2] to the current hazptr API.
>
>The series also adds rcuscale support and torture coverage for
>the hazptr implementation.
>
>The lockdep conversion replaces the expedited RCU wait in
>lockdep_unregister_key() with hazptr_synchronize(). This limits
>the wait to hazard pointers protecting the target hash bucket
>instead of waiting for a system-wide expedited RCU grace period.
>
>[1]
>https://lore.kernel.org/all/20260919000056.3132131-26-paulmck@kernel.org/
>[2]
>https://lore.kernel.org/lkml/20250625031101.12555-1-boqun.feng@gmail.com/
>
This is well needed. Thanks for the patch, I'll maybe review maybe test if
I have time, cough cough, school!
>Performance data
>================
>
>All measurements are on an ARM64 KVM guest with
>HAZPTR_SCALE_NR_OBJS=8 and round-robin updater selection; unless
>noted otherwise, nreaders=0. Scale tests run with
>PROVE_LOCKING=n.
>
>rcuscale synchronize latency (96 CPUs, nw=1):
>
> nreaders=0 nreaders=1 nreaders=4 nreaders=96
> scale avg avg avg avg
> --------------------------------------------------------------------
> hazptr 46 us 47 us 102 us 106 ms
> RCU 8.3 ms 8.5 ms 16.6 ms 54.4 ms
> SRCU 8.2 ms 8.0 ms 11.7 ms 11.8 ms
>
>Hazptr single-writer latency remains below 110 us with up to
>four readers on this 96-CPU guest, but rises to 106 ms when
>all 96 CPUs hold hazard pointers.
>
>Under 16 concurrent synchronize callers (nw=16), per-writer
>latency at four CPU counts:
>
> hazptr nw=16 RCU nw=16 SRCU nw=16
> CPUs avg avg avg
> -----------------------------------------------------------
> 24 8.0 ms 13.6 ms 9.2 ms
> 96 8.0 ms 13.3 ms 8.6 ms
> 128 7.9 ms 13.0 ms 11.7 ms
> 256 8.0 ms 20.8 ms 9.6 ms
>
>Hazptr remains around 8.0 ms across these CPU counts, with the
>scan-kthread retry interval contributing to the latency.
>RCU rises to ~21 ms at 256 CPUs while SRCU stays around 9-12 ms.
>
>Reader-side overhead (refscale, 96 readers on the 96-CPU guest,
>3 runs each):
>
> hazptr 25.7 ns/op
> RCU 96.7 ns/op
> SRCU 134.6 ns/op
>
>lockdep workload -- tc qdisc mq x100, ARM64 KVM, PROVE_LOCKING=y,
>96 background hazptr readers, function-call IPIs per 100 ops:
>
> CPUs hazptr exp RCU reduction
> ------------------------------------------------------------
> 24 196 207 5%
> 96 189 265 29%
> 128 193 266 27%
> 256 222 384 42%
>
>Hazptr issues fewer function-call IPIs at each CPU count, with
>the difference reaching 42% at 256 CPUs.
>
>The hazptr.sh test suite passes, including the lockdep scenarios
>and the 8-to-256 CPU sweep, with no lockdep warnings, deadlocks,
>or crashes.
>Additional x86 server testing with Lian Wang is planned(maybe after
>LPC).
>
>Changes since RFC/WIP
>=====================
>
>- RFC/WIP: https://lore.kernel.org/all/20260922070950.4173245-1-kunwu.chan@gmail.com/
>
>- Split the original 4-patch RFC/WIP into smaller commits covering
> shared scanning, correctness, API support, lockdep, scaling, and
> torture testing.
>
>- Incorporated Boqun Feng's review feedback: use a Bloom filter to
> avoid per-waiter allocation, add scoped_guard() support, and add
> a debug option to force the hazptr acquire slow path.
>
>- Fixed scan ordering around backup-slot promotion by scanning all
> per-CPU slots before the overflow lists, with a separate
> overflow-list phase.
>
>- Simplified the scan cycle to flip first and drain only the old
> wildcard generation, with herd7-verified LKMM tests for both the
> in-flight and resolved publication cases.
>
>- Extended rcuscale and hazptrtorture coverage, added a selftest
> script for the torture configurations, and fixed the
> hazptr_release() kernel-doc.
>
>Kunwu Chan (15):
> hazptr: add shared scan kthread
> hazptr: use Bloom filter for shared scan waiters
> hazptr: scan all per-CPU slots before overflow lists
> hazptr: add scoped_guard() support
> hazptr: add debug option to force the acquire slow path
> hazptr: elide redundant first drain pass
> Documentation/litmus-tests: add hazptr wildcard-flip escape test
> locking/lockdep: use hazptr to wait for dynamic key lookups
> rcuscale: add hazptr scale type
> hazptr: fix kernel-doc of hazptr_release()
> Documentation/litmus-tests: add hazptr acquire-before-scan test
> hazptrtorture: add slowpath and lockdep scenarios
> hazptrtorture: add READERS4 and READERS0 torture configs
> hazptrtorture: add 128- and 256-CPU configs
> selftests/rcutorture: add hazptr torture test script
>
> Documentation/litmus-tests/README | 13 +
> .../hazptr/hazptr-acquire-before-scan.litmus | 45 +++
> .../hazptr/hazptr-wildcard-flip-escape.litmus | 46 +++
> include/linux/hazptr.h | 56 ++-
> kernel/hazptr.c | 336 +++++++++++++++++-
> kernel/locking/lockdep.c | 25 +-
> kernel/rcu/Kconfig.debug | 10 +
> kernel/rcu/hazptrtorture.c | 57 ++-
> kernel/rcu/rcuscale.c | 70 +++-
> .../selftests/rcutorture/bin/hazptr.sh | 146 ++++++++
> .../rcutorture/configs/hazptr/CFLIST | 6 +
> .../rcutorture/configs/hazptr/CPU128 | 16 +
> .../rcutorture/configs/hazptr/CPU128.boot | 1 +
> .../rcutorture/configs/hazptr/CPU256 | 16 +
> .../rcutorture/configs/hazptr/CPU256.boot | 1 +
> .../rcutorture/configs/hazptr/LOCKDEP | 17 +
> .../rcutorture/configs/hazptr/LOCKDEP.boot | 1 +
> .../rcutorture/configs/hazptr/READERS0 | 16 +
> .../rcutorture/configs/hazptr/READERS0.boot | 2 +
> .../rcutorture/configs/hazptr/READERS4 | 16 +
> .../rcutorture/configs/hazptr/READERS4.boot | 2 +
> .../rcutorture/configs/hazptr/SLOWPATH | 16 +
> 22 files changed, 881 insertions(+), 33 deletions(-)
> create mode 100644
> Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus
> create mode 100644
> Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus
> create mode 100755 tools/testing/selftests/rcutorture/bin/hazptr.sh
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/CPU128
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/CPU256
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/READERS0
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/READERS4
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot
> create mode 100644
> tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH
>
>
--- Thanks!
"I'm not a very positive person" - Linus torvalds
^ permalink raw reply [flat|nested] 17+ messages in thread
end of thread, other threads:[~2026-10-02 17:13 UTC | newest]
Thread overview: 17+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 17:08 [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 01/15] hazptr: add shared scan kthread Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 02/15] hazptr: use Bloom filter for shared scan waiters Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 03/15] hazptr: scan all per-CPU slots before overflow lists Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 04/15] hazptr: add scoped_guard() support Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 05/15] hazptr: add debug option to force the acquire slow path Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 06/15] hazptr: elide redundant first drain pass Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 07/15] Documentation/litmus-tests: add hazptr wildcard-flip escape test Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 08/15] locking/lockdep: use hazptr to wait for dynamic key lookups Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 09/15] rcuscale: add hazptr scale type Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 10/15] hazptr: fix kernel-doc of hazptr_release() Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 11/15] Documentation/litmus-tests: add hazptr acquire-before-scan test Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 12/15] hazptrtorture: add slowpath and lockdep scenarios Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 13/15] hazptrtorture: add READERS4 and READERS0 torture configs Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 14/15] hazptrtorture: add 128- and 256-CPU configs Kunwu Chan
2026-10-02 17:08 ` [PATCH RFC v2 15/15] selftests/rcutorture: add hazptr torture test script Kunwu Chan
2026-10-02 17:12 ` [PATCH RFC v2 00/15] hazptr: batch synchronize operations through a shared scan Bradley Morgan
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®