mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH RFC v3 00/15] hazptr: batch synchronize operations through a shared scan
@ 2026-10-05 17:15 Kunwu Chan
  2026-10-05 17:15 ` [PATCH RFC v3 01/15] hazptr: add shared scan kthread Kunwu Chan
                   ` (14 more replies)
  0 siblings, 15 replies; 18+ messages in thread
From: Kunwu Chan @ 2026-10-05 17:15 UTC (permalink / raw)
  To: paulmck, corbet, mingo, frederic, neeraj.upadhyay, josh, urezki,
	dave, lianux.mm
  Cc: stern, parri.andrea, will, peterz, boqun, npiggin, dhowells,
	j.alglave, luc.maranget, akiyks, dlustig, joelagnelf, skhan,
	rdunlap, longman, rostedt, mathieu.desnoyers, jiangshanlai,
	qiang.zhang, kunwu.chan, brads, linux-kernel, linux-arch, lkmm,
	linux-doc, rcu, linux-kselftest

This series extends Mathieu Desnoyers's hazard pointer implementation [1]
with a shared scan path for concurrent hazptr_synchronize() callers,
and adapts the lockdep dynamic-key hashlist use case from Boqun
Feng's 2025 shazptr series [2] to the current hazptr API.

The series also adds rcuscale support, torture coverage, and litmus
tests for the hazptr implementation.

The lockdep conversion replaces the expedited RCU wait in
lockdep_unregister_key() with hazptr_synchronize().  This limits
the wait to hazard pointers protecting the target hash bucket
instead of waiting for a system-wide expedited RCU grace period.

[1] https://lore.kernel.org/all/20260919000056.3132131-26-paulmck@kernel.org/
[2] https://lore.kernel.org/lkml/20250625031101.12555-1-boqun.feng@gmail.com/

Performance data
================

The measurements below were collected on an ARM64 KVM guest running
on an ARM64 server (96 vCPUs, 8 GB RAM), unless noted otherwise.
The kernel is based on Paul McKenney's -rcu tree "dev" branch at
commit d21906b0aa1e ("doc: Document additional RCU task-stall dump
information") with the full hazptr patch series applied.

CONFIG_PREEMPT=y, CONFIG_PREEMPT_RCU=y, CONFIG_NR_CPUS=256.
Configuration is saved alongside each result set.

Following Paul's suggestion, the reader-side refscale numbers are
reported separately for CONFIG_PROVE_LOCKING=n and
CONFIG_PROVE_LOCKING=y, since lockdep instrumentation significantly
affects the measured RCU and SRCU reader-side costs but has little
effect on the hazptr reader path in this setup.

Lockdep workload -- tc qdisc mq x100, CONFIG_PROVE_LOCKING=y,
96 background hazptr readers:

  function-call IPIs (100 tc qdisc operations):  183

For reference, insmod/rmmod x10 in the same setup generated 4930
function-call IPIs.

The lockdep_unregister_key() path is the motivating use case for
this series: it replaces a system-wide synchronize_rcu_expedited()
call with a hazptr_synchronize() scoped to the per-key hash bucket.

Writer-side effective per-GP time (rcuscale, nwriters=16, nreaders=0,
one run per CPU count; computed as total test duration divided by
the total number of grace periods across all writers):

                         hazptr      RCU         Tree SRCU
  CPUs                   per-GP      per-GP      per-GP
  -----------------------------------------------------------
   24                    498 us      863 us      632 us
   96                    493 us      865 us      599 us
  128                    493 us      873 us      789 us
  256                    500 us      1162 us     476 us

At 96 CPUs the measurements are from the comparison blocks in the
same guest boot (order: hazptr nw=16, hazptr nw=1, hazptr nw=16,
rcu nw=1, rcu nw=16, srcu nw=1, srcu nw=16).  The hazptr scalability
row uses the first hazptr nw=16 block, which is consistent with the
single per-CPU runs at 24, 128, and 256 CPUs.  RCU and Tree SRCU
use their standard implementations.

Hazptr effective per-GP time stays nearly constant from 24 to 256
CPUs, remaining within 493-500 us, consistent with the shared-scan
kthread amortising the scan across concurrent synchronize callers.
At 96 CPUs, the effective per-GP times are approximately 865 us for
RCU, 599 us for SRCU, and 493 us for hazptr with 16 concurrent
callers.

For comparison, the nwriters=1, nreaders=0 case measured about 19 us
per grace period in this run at 96 CPUs.

Reader-side overhead (refscale, nreaders=-1 for 75% of 96 online
CPUs, nruns=5, median of 5 runs per scale type):

  The lockdep=n and lockdep=y refscale numbers were collected in
  separate QEMU boots from kernels built from the same source tree
  with PROVE_LOCKING toggled.  The first hazptr block in each boot
  is used.  All 96 vCPUs were visible to the guest (CONFIG_NR_CPUS=256).

                   PROVE_LOCKING=n     PROVE_LOCKING=y
  hazptr           26.1 ns             25.6 ns
  RCU               5.0 ns            165.5 ns
  SRCU             38.8 ns            192.3 ns

Hazptr reader-side overhead is nearly unchanged with and without
CONFIG_PROVE_LOCKING in this setup, while RCU and SRCU show
substantially higher costs when lockdep is enabled.  This is why
the refscale results are reported separately for the two
configurations.

Hazptr torture regression (CONFIG_PROVE_LOCKING=y):
basic, lockdep+wq_churn, and SLOWPATH PASS at 96 CPUs;
CPU sweep (8/16/32/64/128/256) also PASS with no lockdep warnings.

Additional x86 server testing with Lian Wang is planned(maybe after
LPC).

Test reproducibility
---------------------

Across five refscale runs, the standard deviation was 0.16 ns for
hazptr and 0.12 ns for RCU with PROVE_LOCKING=n.  Rcuscale data
points are based on a single test run per CPU count, but each run
covers thousands of individual grace periods and the hazptr 
measurements remain within 493-500us across the tested CPU counts.
The lockdep workload measurement uses 100 tc qdisc operations in 
a single QEMU boot; repeated runs agree to within ~10 IPIs on 
this server.

Changes since v2
================

- v2: https://lore.kernel.org/all/20261002170847.3653663-1-kunwu.chan@gmail.com/

Only cover letter changes; no code changes since v2.

- Fixed attribution: the base hazptr implementation is Mathieu
  Desnoyers's work.

- Split the refscale reader-overhead table into separate
  PROVE_LOCKING=n and PROVE_LOCKING=y columns, following Paul's
  suggestion, so that the lockdep impact on RCU and SRCU fast
  paths is visible rather than folded into a single number.

- Replaced the rcuscale "per-writer latency" column with
  per-grace-period values.  The new metric (total test duration
  divided by total grace periods) is more directly interpretable
  when comparing 1-vs-16 concurrent synchronize callers.  The
  nw=16 comparison now uses the same rcuscale test block for
  hazptr, RCU, and SRCU at each CPU count.  The nw=1 baseline
  is provided for reference so that the reader can see the single-
  writer cost and the per-GP cost under 16 concurrent callers side
  by side.

- Added test-condition and reproducibility notes (kernel commit,
  PREEMPT model, visible CPU count, boot ordering, standard
  deviation, sample sizes).

- Removed the unreviewed srcua scale-type patch from this series;
  it will be posted separately.

Changes since RFC/WIP
=====================

- RFC/WIP: https://lore.kernel.org/all/20260922070950.4173245-1-kunwu.chan@gmail.com/

- Split the original 4-patch RFC/WIP into smaller commits covering
  shared scanning, correctness, API support, lockdep, scaling, and
  torture testing.

- Incorporated Boqun Feng's review feedback: use a Bloom filter to
  avoid per-waiter allocation, add scoped_guard() support, and add
  a debug option to force the hazptr acquire slow path.

- Fixed scan ordering around backup-slot promotion by scanning all
  per-CPU slots before the overflow lists, with a separate
  overflow-list phase.

- Simplified the scan cycle to flip first and drain only the old
  wildcard generation, with herd7-verified LKMM tests for both the
  in-flight and resolved publication cases.

- Extended rcuscale and hazptrtorture coverage, added a selftest
  script for the torture configurations, and fixed the
  hazptr_release() kernel-doc.


Kunwu Chan (15):
  hazptr: add shared scan kthread
  hazptr: use Bloom filter for shared scan waiters
  hazptr: scan all per-CPU slots before overflow lists
  hazptr: add scoped_guard() support
  hazptr: add debug option to force the acquire slow path
  hazptr: elide redundant first drain pass
  Documentation/litmus-tests: add hazptr wildcard-flip escape test
  locking/lockdep: use hazptr to wait for dynamic key lookups
  rcuscale: add hazptr scale type
  hazptr: fix kernel-doc of hazptr_release()
  Documentation/litmus-tests: add hazptr acquire-before-scan test
  hazptrtorture: add slowpath and lockdep scenarios
  hazptrtorture: add READERS4 and READERS0 torture configs
  hazptrtorture: add 128- and 256-CPU configs
  selftests/rcutorture: add hazptr torture test script

 Documentation/litmus-tests/README             |  13 +
 .../hazptr/hazptr-acquire-before-scan.litmus  |  45 +++
 .../hazptr/hazptr-wildcard-flip-escape.litmus |  46 +++
 include/linux/hazptr.h                        |  56 ++-
 kernel/hazptr.c                               | 336 +++++++++++++++++-
 kernel/locking/lockdep.c                      |  25 +-
 kernel/rcu/Kconfig.debug                      |  10 +
 kernel/rcu/hazptrtorture.c                    |  57 ++-
 kernel/rcu/rcuscale.c                         |  70 +++-
 .../selftests/rcutorture/bin/hazptr.sh        | 146 ++++++++
 .../rcutorture/configs/hazptr/CFLIST          |   6 +
 .../rcutorture/configs/hazptr/CPU128          |  16 +
 .../rcutorture/configs/hazptr/CPU128.boot     |   1 +
 .../rcutorture/configs/hazptr/CPU256          |  16 +
 .../rcutorture/configs/hazptr/CPU256.boot     |   1 +
 .../rcutorture/configs/hazptr/LOCKDEP         |  17 +
 .../rcutorture/configs/hazptr/LOCKDEP.boot    |   1 +
 .../rcutorture/configs/hazptr/READERS0        |  16 +
 .../rcutorture/configs/hazptr/READERS0.boot   |   2 +
 .../rcutorture/configs/hazptr/READERS4        |  16 +
 .../rcutorture/configs/hazptr/READERS4.boot   |   2 +
 .../rcutorture/configs/hazptr/SLOWPATH        |  16 +
 22 files changed, 881 insertions(+), 33 deletions(-)
 create mode 100644 Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus
 create mode 100644 Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus
 create mode 100755 tools/testing/selftests/rcutorture/bin/hazptr.sh
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot
 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH

base-commit: d21906b0aa1e9573cdb5e7acaca44966b9d1dcd2
-- 
2.43.0


^ permalink raw reply	[flat|nested] 18+ messages in thread

end of thread, other threads:[~2026-10-06  6:13 UTC | newest]

Thread overview: 18+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-05 17:15 [PATCH RFC v3 00/15] hazptr: batch synchronize operations through a shared scan Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 01/15] hazptr: add shared scan kthread Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 02/15] hazptr: use Bloom filter for shared scan waiters Kunwu Chan
2026-10-06  5:51   ` Mathieu Desnoyers
2026-10-06  6:13     ` KunWu Chan
2026-10-05 17:15 ` [PATCH RFC v3 03/15] hazptr: scan all per-CPU slots before overflow lists Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 04/15] hazptr: add scoped_guard() support Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 05/15] hazptr: add debug option to force the acquire slow path Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 06/15] hazptr: elide redundant first drain pass Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 07/15] Documentation/litmus-tests: add hazptr wildcard-flip escape test Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 08/15] locking/lockdep: use hazptr to wait for dynamic key lookups Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 09/15] rcuscale: add hazptr scale type Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 10/15] hazptr: fix kernel-doc of hazptr_release() Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 11/15] Documentation/litmus-tests: add hazptr acquire-before-scan test Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 12/15] hazptrtorture: add slowpath and lockdep scenarios Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 13/15] hazptrtorture: add READERS4 and READERS0 torture configs Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 14/15] hazptrtorture: add 128- and 256-CPU configs Kunwu Chan
2026-10-05 17:15 ` [PATCH RFC v3 15/15] selftests/rcutorture: add hazptr torture test script Kunwu Chan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®