* Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA [not found] ` <aqvq8-6IJBDer90O@thinkstation> @ 2026-09-17 14:10 ` Harry Yoo 2026-09-17 15:29 ` Peter Zijlstra ` (2 more replies) 0 siblings, 3 replies; 6+ messages in thread From: Harry Yoo @ 2026-09-17 14:10 UTC (permalink / raw) To: Tim Menninger Cc: linux-mm, Chuck Lever, linux-nfs, Jon Curley, Eric Badger, Vlastimil Babka, Andrew Morton, Hao Li, Christoph Lameter, David Rientjes, Roman Gushchin, Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers, Adrian Hunter, James Clark, linux-perf-users, linux-kernel ... now I realize we need to add PERFORMANCE EVENTS SUBSYSTEM folks as well ;-) Hmm, sounds like perf lock is somehow triggering slab allocations and interfering the workload. It could be because SLUB is merging nfs_page cache with some other cache that perf uses. Could you please check if it reproduces with slab_nomerge kernel parameter? On Thu, Sep 17, 2026 at 02:46:24PM +0100, Harry Yoo wrote: > Hi Tim and Chuck, thanks for reporting this to linux-mm. > Will take a look at this but let me Cc SLAB ALLOCATOR folks here first. > > -- > Cheers, > Harry / Hyeonggon > > On Wed, Sep 16, 2026 at 11:22:27PM +0000, Tim Menninger wrote: > > Chuck Lever suggested I bring this to linux-mm after we found some SLUB > > measurements we could not account for while investigating an NFS/RDMA > > throughput regression. > > > > The original NFS discussion is here for context: > > > > https://lore.kernel.org/linux-nfs/0e2f9688-097c-4bb9-a7ec-82b41bdb3653@slotpi15m67/ > > > > The NFS regression itself has been separated from this issue. What > > remains interesting here is the behavior of the nfs_page slab cache on > > this machine, particularly the measurements with and without perf lock. > > > > The system is: > > > > Intel Xeon Silver 4516Y+ > > 2 sockets > > 24 cores/socket > > 2 threads/core > > 96 logical CPUs > > > > NUMA node0 CPUs: 0-23,48-71 > > NUMA node1 CPUs: 24-47,72-95 > > > > The workload is a high-throughput NFS/RDMA direct-read workload using > > 1 MiB I/O, 160 threads, and iodepth 64. > > > > The relevant debug options are all disabled: > > > > # CONFIG_KASAN is not set > > # CONFIG_PROVE_LOCKING is not set > > # CONFIG_LOCK_STAT is not set > > # CONFIG_DEBUG_SPINLOCK is not set > > # CONFIG_DEBUG_LIST is not set > > > > All measurements below were collected on the unpatched base kernel: > > > > $ git rev-parse HEAD > > 940de590b839f71d6dc846160534bf202401b8b7 > > > > $ uname -r > > 7.3.0-rc1-mainline+ > > > > The initial observation was a high apparent contention rate on the > > nfs_page slab's list_lock. In 10-second perf-lock captures I was seeing > > roughly 4.5M contended acquisitions and about 170-185 us of average > > reported wait. > > > > Chuck reproduced a similar acquisition rate on a single-node EPYC > > system, but saw only about 7 ns average wait and fewer than 100 > > cmpxchg_double_fail events over a corresponding interval. He suggested > > checking cmpxchg_double_fail because __slab_free() drops list_lock and > > retries when the freelist cmpxchg fails. > > > > I repeated the measurements in three placement configurations: > > > > A. workload unpinned, CQs all on node0 > > B. workload pinned to node0, CQs all on node0 > > C. workload unpinned, CQs balanced across the nodes > > > > Without perf lock, throughput is similar in all three: > > > > A. unpinned / CQs node0: ~45.5 GB/s > > B. node0 pinned / CQs node0: ~45.7 GB/s > > C. unpinned / balanced CQs: ~45.5 GB/s > > > > During the perf-lock captures, throughput is approximately 25 GB/s. > > > > For each instrumented 10-second window I ran: > > > > sudo perf lock record -a -g -o "$D/perf-locks.data" -- sleep 10 > > mpstat -P ALL 1 10 > > > > Before and after the same window I sampled the counters under: > > > > /sys/kernel/slab/nfs_page/ > > > > I also collected separate 10-second counter and mpstat windows under > > the same workload configurations without perf lock. > > > > The resulting slab counter deltas were: > > > > A B C > > unpinned/node0 node0/node0 unpinned/balanced > > > > free_fastpath > > instrumented 655,917,745 457,936,140 328,670,030 > > uninstrumented 34,906,233 119,984,226 59,469,931 > > > > free_slowpath > > instrumented 236,735,724 7,720,917 270,716,506 > > uninstrumented 85,039,753 2,814 60,145,870 > > > > sheaf_flush > > instrumented 39,842,700 42,706,800 8,341,440 > > uninstrumented 1,286,400 13,487,700 887,700 > > > > barn_put_fail > > instrumented 663,994 711,745 139,037 > > uninstrumented 21,464 224,762 14,765 > > > > barn_get_fail > > instrumented 4,609,925 840,344 4,651,225 > > uninstrumented 1,438,643 224,787 1,017,006 > > > > cmpxchg_double_fail > > instrumented 70,524 8,077 28,062 > > uninstrumented 7,471 713 1,503 > > > > alloc_slowpath > > all cases 0 0 0 > > > > The SLUB counter profile changes substantially with perf lock despite > > the lower NFS throughput, and the exact mix depends strongly on > > placement. > > > > The uninstrumented node0/node0 run also reproduces the barn/sheaf > > relationship Chuck pointed out earlier. The nfs_page sheaf capacity is > > 60: > > > > sheaf_flush / 60 = 13,487,700 / 60 = 224,795 > > barn_put_fail = 224,762 > > > > The placement dependence of free_slowpath is also large. It falls from > > about 85M events per 10 seconds in the unpinned/node0 case to 2,814 in > > the node0/node0 case. > > > > The reported perf-lock result, however, is similar across all three > > placements: > > > > contentions total wait average wait > > unpinned/node0 4,782,839 14.39 min 180.47 us > > node0/node0 4,582,012 14.13 min 185.05 us > > unpinned/balanced 4,572,196 12.81 min 168.14 us > > > > I also revisited an inconsistency Chuck noticed in my earlier > > measurements. Previously I had compared aggregate perf-lock wait from > > one 10-second capture with CPU utilization measured during a different > > window. > > > > I now have paired 10-second mpstat samples for each placement, with and > > without perf lock. The node values below are averages of the per-CPU > > %idle values for the CPUs in each NUMA node: > > > > system-wide node0 node1 > > %idle %idle %idle > > > > A. unpinned / CQs node0 > > uninstrumented 31.46 8.0 54.6 > > instrumented 3.60 0.06 7.1 > > > > B. node0 pinned / CQs node0 > > uninstrumented 84.26 69.2 99.4 > > instrumented 12.07 10.4 13.8 > > > > C. unpinned / balanced CQs > > uninstrumented 66.86 66.7 66.9 > > instrumented 14.83 19.0 10.6 > > > > This resolves the accounting inconsistency in my earlier measurements. > > The large aggregate perf-lock wait and high idle percentage had come > > from different windows. In the aligned samples, the system is much > > busier during the perf-lock capture than in the corresponding > > uninstrumented run. > > > > I am still unsure how representative the reported ~170-185 us average > > wait is of the uninstrumented workload. > > > > The remaining number I am less sure how to interpret is > > cmpxchg_double_fail. In the uninstrumented windows I see: > > > > unpinned / CQs node0: 7,471 / 10 sec > > node0 pinned / CQs node0: 713 / 10 sec > > unpinned / balanced CQs: 1,503 / 10 sec > > > > compared with fewer than 100 in Chuck's test. > > > > I understand that cmpxchg_double_fail counts failed slab freelist > > updates rather than failed logical frees, so I am not sure what the > > appropriate denominator is here. In particular, the node0/node0 case > > still has 713 failures while sustaining full throughput and only 2,814 > > free_slowpath events over the interval. > > > > My questions are: > > > > 1. Do the uninstrumented cmpxchg_double_fail rates above look abnormal > > for this workload/topology, or are they within the range one would > > expect from this degree of concurrency and NUMA placement? > > > > 2. Is there a less invasive way you would recommend measuring the > > nfs_page list_lock/freelist contention? I would like to distinguish > > the steady-state behavior from what is observed during the perf-lock > > capture. > > > > Thanks, > > Tim -- Cheers, Harry / Hyeonggon ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA 2026-09-17 14:10 ` [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA Harry Yoo @ 2026-09-17 15:29 ` Peter Zijlstra 2026-09-18 7:04 ` Namhyung Kim 2026-09-17 16:25 ` Tim Menninger 2026-09-18 7:14 ` Namhyung Kim 2 siblings, 1 reply; 6+ messages in thread From: Peter Zijlstra @ 2026-09-17 15:29 UTC (permalink / raw) To: Harry Yoo Cc: Tim Menninger, linux-mm, Chuck Lever, linux-nfs, Jon Curley, Eric Badger, Vlastimil Babka, Andrew Morton, Hao Li, Christoph Lameter, David Rientjes, Roman Gushchin, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers, Adrian Hunter, James Clark, linux-perf-users, linux-kernel On Thu, Sep 17, 2026 at 03:10:47PM +0100, Harry Yoo wrote: > ... now I realize we need to add PERFORMANCE EVENTS SUBSYSTEM folks > as well ;-) > > Hmm, sounds like perf lock is somehow triggering slab allocations > and interfering the workload. perf should not have runtime memory allocation, its build to run from NMI context. It can allocate some memory on event creation and the like, but once its on, there should be no more allocations happening. ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA 2026-09-17 15:29 ` Peter Zijlstra @ 2026-09-18 7:04 ` Namhyung Kim 0 siblings, 0 replies; 6+ messages in thread From: Namhyung Kim @ 2026-09-18 7:04 UTC (permalink / raw) To: Peter Zijlstra Cc: Harry Yoo, Tim Menninger, linux-mm, Chuck Lever, linux-nfs, Jon Curley, Eric Badger, Vlastimil Babka, Andrew Morton, Hao Li, Christoph Lameter, David Rientjes, Roman Gushchin, Ingo Molnar, Arnaldo Carvalho de Melo, Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers, Adrian Hunter, James Clark, linux-perf-users, linux-kernel Hello, On Thu, Sep 17, 2026 at 05:29:59PM +0200, Peter Zijlstra wrote: > On Thu, Sep 17, 2026 at 03:10:47PM +0100, Harry Yoo wrote: > > ... now I realize we need to add PERFORMANCE EVENTS SUBSYSTEM folks > > as well ;-) > > > > Hmm, sounds like perf lock is somehow triggering slab allocations > > and interfering the workload. > > perf should not have runtime memory allocation, its build to run from > NMI context. It can allocate some memory on event creation and the like, > but once its on, there should be no more allocations happening. Right, perf lock record command will save the event info to a ring buffer without allocation. It may lose some events when the buffer is full though. Thanks, Namhyung ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA 2026-09-17 14:10 ` [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA Harry Yoo 2026-09-17 15:29 ` Peter Zijlstra @ 2026-09-17 16:25 ` Tim Menninger 2026-09-18 7:08 ` Vlastimil Babka (SUSE) 2026-09-18 7:14 ` Namhyung Kim 2 siblings, 1 reply; 6+ messages in thread From: Tim Menninger @ 2026-09-17 16:25 UTC (permalink / raw) To: Harry Yoo Cc: linux-mm, Chuck Lever, linux-nfs, Jon Curley, Eric Badger, Vlastimil Babka, Andrew Morton, Hao Li, Christoph Lameter, David Rientjes, Roman Gushchin, Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers, Adrian Hunter, James Clark, linux-perf-users, linux-kernel It reproduces with slab_nomerge. I booted the same kernel with: $ cat /proc/cmdline ... slab_nomerge ... Uninstrumented throughput remains in the high-45 to low-46 GB/s range, as before. During perf lock it was around 30 GB/s in these runs, versus roughly 25 GB/s previously. The dominant lock measurements are essentially unchanged: default slab_nomerge unpinned/node0 contentions 4,782,839 4,858,759 total wait 14.39 min 14.26 min average wait 180.47 us 176.13 us node0/node0 contentions 4,582,012 4,686,729 total wait 14.13 min 14.34 min average wait 185.05 us 183.53 us unpinned/balanced contentions 4,572,196 4,544,677 total wait 12.81 min 12.56 min average wait 168.14 us 165.84 us The uninstrumented cmpxchg_double_fail deltas were: default slab_nomerge unpinned/node0 7,471 8,892 node0/node0 713 457 unpinned/balanced 1,503 1,504 The CPU-utilization change while perf lock is active also remains. System-wide %idle for uninstrumented/instrumented windows was: default slab_nomerge uninstr. instr. uninstr. instr. unpinned/node0 31.46% 3.60% 23.02% 2.15% node0/node0 84.26% 12.07% 84.62% 11.85% unpinned/balanced 66.86% 14.83% 66.58% 14.78% Thanks, Tim ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA 2026-09-17 16:25 ` Tim Menninger @ 2026-09-18 7:08 ` Vlastimil Babka (SUSE) 0 siblings, 0 replies; 6+ messages in thread From: Vlastimil Babka (SUSE) @ 2026-09-18 7:08 UTC (permalink / raw) To: Tim Menninger, Harry Yoo Cc: linux-mm, Chuck Lever, linux-nfs, Jon Curley, Eric Badger, Andrew Morton, Hao Li, Christoph Lameter, David Rientjes, Roman Gushchin, Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers, Adrian Hunter, James Clark, linux-perf-users, linux-kernel On 9/17/26 18:25, Tim Menninger wrote: > It reproduces with slab_nomerge. > > I booted the same kernel with: > > $ cat /proc/cmdline > ... slab_nomerge ... > > Uninstrumented throughput remains in the high-45 to low-46 GB/s > range, as before. During perf lock it was around 30 GB/s in these > runs, versus roughly 25 GB/s previously. > > The dominant lock measurements are essentially unchanged: > > default slab_nomerge > unpinned/node0 > contentions 4,782,839 4,858,759 > total wait 14.39 min 14.26 min > average wait 180.47 us 176.13 us > > node0/node0 > contentions 4,582,012 4,686,729 > total wait 14.13 min 14.34 min > average wait 185.05 us 183.53 us > > unpinned/balanced > contentions 4,572,196 4,544,677 > total wait 12.81 min 12.56 min > average wait 168.14 us 165.84 us > > The uninstrumented cmpxchg_double_fail deltas were: > > default slab_nomerge > unpinned/node0 7,471 8,892 > node0/node0 713 457 > unpinned/balanced 1,503 1,504 > > The CPU-utilization change while perf lock is active also remains. > System-wide %idle for uninstrumented/instrumented windows was: > > default slab_nomerge > uninstr. instr. uninstr. instr. > unpinned/node0 31.46% 3.60% 23.02% 2.15% > node0/node0 84.26% 12.07% 84.62% 11.85% > unpinned/balanced 66.86% 14.83% 66.58% 14.78% I believe the weirdest part in the original report was how slub stat counters such as free_fastpath etc got order of magnitude higher with instrumented vs non-instrumented. Did any of that change with slab_nomerge? > Thanks, > Tim ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA 2026-09-17 14:10 ` [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA Harry Yoo 2026-09-17 15:29 ` Peter Zijlstra 2026-09-17 16:25 ` Tim Menninger @ 2026-09-18 7:14 ` Namhyung Kim 2 siblings, 0 replies; 6+ messages in thread From: Namhyung Kim @ 2026-09-18 7:14 UTC (permalink / raw) To: Harry Yoo Cc: Tim Menninger, linux-mm, Chuck Lever, linux-nfs, Jon Curley, Eric Badger, Vlastimil Babka, Andrew Morton, Hao Li, Christoph Lameter, David Rientjes, Roman Gushchin, Peter Zijlstra, Ingo Molnar, Arnaldo Carvalho de Melo, Mark Rutland, Alexander Shishkin, Jiri Olsa, Ian Rogers, Adrian Hunter, James Clark, linux-perf-users, linux-kernel On Thu, Sep 17, 2026 at 03:10:47PM +0100, Harry Yoo wrote: > ... now I realize we need to add PERFORMANCE EVENTS SUBSYSTEM folks > as well ;-) > > Hmm, sounds like perf lock is somehow triggering slab allocations > and interfering the workload. > > It could be because SLUB is merging nfs_page cache with some other > cache that perf uses. > > Could you please check if it reproduces with slab_nomerge kernel > parameter? > > On Thu, Sep 17, 2026 at 02:46:24PM +0100, Harry Yoo wrote: > > Hi Tim and Chuck, thanks for reporting this to linux-mm. > > Will take a look at this but let me Cc SLAB ALLOCATOR folks here first. > > > > -- > > Cheers, > > Harry / Hyeonggon > > [SNIP] > > On Wed, Sep 16, 2026 at 11:22:27PM +0000, Tim Menninger wrote: > > > My questions are: > > > > > > 1. Do the uninstrumented cmpxchg_double_fail rates above look abnormal > > > for this workload/topology, or are they within the range one would > > > expect from this degree of concurrency and NUMA placement? > > > > > > 2. Is there a less invasive way you would recommend measuring the > > > nfs_page list_lock/freelist contention? I would like to distinguish > > > the steady-state behavior from what is observed during the perf-lock > > > capture. You can use the BPF version of perf lock to check lock contention like below. (it only work with 'contention' subcommand.) It won't create a data file and just shows the result. $ sudo perf lock con -ab sleep 10 or $ sudo perf lock con -ab -E 5 sleep 10 to see top-5 locks only. Thanks, Namhyung ^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-18 7:14 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
[not found] <20260916232227.4098143-1-tmenninger@everpuredata.com>
[not found] ` <aqvq8-6IJBDer90O@thinkstation>
2026-09-17 14:10 ` [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA Harry Yoo
2026-09-17 15:29 ` Peter Zijlstra
2026-09-18 7:04 ` Namhyung Kim
2026-09-17 16:25 ` Tim Menninger
2026-09-18 7:08 ` Vlastimil Babka (SUSE)
2026-09-18 7:14 ` Namhyung Kim
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®