From: Harry Yoo <harry@kernel.org>
To: Tim Menninger <tmenninger@everpuredata.com>
Cc: linux-mm@kvack.org, Chuck Lever <cel@kernel.org>,
linux-nfs@vger.kernel.org, Jon Curley <jcurley@everpuredata.com>,
Eric Badger <ebadger@everpuredata.com>,
Vlastimil Babka <vbabka@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
Hao Li <hao.li@linux.dev>, Christoph Lameter <cl@gentwo.org>,
David Rientjes <rientjes@google.com>,
Roman Gushchin <roman.gushchin@linux.dev>,
Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
Arnaldo Carvalho de Melo <acme@kernel.org>,
Namhyung Kim <namhyung@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Alexander Shishkin <alexander.shishkin@linux.intel.com>,
Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
James Clark <james.clark@linaro.org>,
linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA
Date: Thu, 17 Sep 2026 15:10:47 +0100 [thread overview]
Message-ID: <aqvy8NqRRRmUlxjR@thinkstation> (raw)
In-Reply-To: <aqvq8-6IJBDer90O@thinkstation>
... now I realize we need to add PERFORMANCE EVENTS SUBSYSTEM folks
as well ;-)
Hmm, sounds like perf lock is somehow triggering slab allocations
and interfering the workload.
It could be because SLUB is merging nfs_page cache with some other
cache that perf uses.
Could you please check if it reproduces with slab_nomerge kernel
parameter?
On Thu, Sep 17, 2026 at 02:46:24PM +0100, Harry Yoo wrote:
> Hi Tim and Chuck, thanks for reporting this to linux-mm.
> Will take a look at this but let me Cc SLAB ALLOCATOR folks here first.
>
> --
> Cheers,
> Harry / Hyeonggon
>
> On Wed, Sep 16, 2026 at 11:22:27PM +0000, Tim Menninger wrote:
> > Chuck Lever suggested I bring this to linux-mm after we found some SLUB
> > measurements we could not account for while investigating an NFS/RDMA
> > throughput regression.
> >
> > The original NFS discussion is here for context:
> >
> > https://lore.kernel.org/linux-nfs/0e2f9688-097c-4bb9-a7ec-82b41bdb3653@slotpi15m67/
> >
> > The NFS regression itself has been separated from this issue. What
> > remains interesting here is the behavior of the nfs_page slab cache on
> > this machine, particularly the measurements with and without perf lock.
> >
> > The system is:
> >
> > Intel Xeon Silver 4516Y+
> > 2 sockets
> > 24 cores/socket
> > 2 threads/core
> > 96 logical CPUs
> >
> > NUMA node0 CPUs: 0-23,48-71
> > NUMA node1 CPUs: 24-47,72-95
> >
> > The workload is a high-throughput NFS/RDMA direct-read workload using
> > 1 MiB I/O, 160 threads, and iodepth 64.
> >
> > The relevant debug options are all disabled:
> >
> > # CONFIG_KASAN is not set
> > # CONFIG_PROVE_LOCKING is not set
> > # CONFIG_LOCK_STAT is not set
> > # CONFIG_DEBUG_SPINLOCK is not set
> > # CONFIG_DEBUG_LIST is not set
> >
> > All measurements below were collected on the unpatched base kernel:
> >
> > $ git rev-parse HEAD
> > 940de590b839f71d6dc846160534bf202401b8b7
> >
> > $ uname -r
> > 7.3.0-rc1-mainline+
> >
> > The initial observation was a high apparent contention rate on the
> > nfs_page slab's list_lock. In 10-second perf-lock captures I was seeing
> > roughly 4.5M contended acquisitions and about 170-185 us of average
> > reported wait.
> >
> > Chuck reproduced a similar acquisition rate on a single-node EPYC
> > system, but saw only about 7 ns average wait and fewer than 100
> > cmpxchg_double_fail events over a corresponding interval. He suggested
> > checking cmpxchg_double_fail because __slab_free() drops list_lock and
> > retries when the freelist cmpxchg fails.
> >
> > I repeated the measurements in three placement configurations:
> >
> > A. workload unpinned, CQs all on node0
> > B. workload pinned to node0, CQs all on node0
> > C. workload unpinned, CQs balanced across the nodes
> >
> > Without perf lock, throughput is similar in all three:
> >
> > A. unpinned / CQs node0: ~45.5 GB/s
> > B. node0 pinned / CQs node0: ~45.7 GB/s
> > C. unpinned / balanced CQs: ~45.5 GB/s
> >
> > During the perf-lock captures, throughput is approximately 25 GB/s.
> >
> > For each instrumented 10-second window I ran:
> >
> > sudo perf lock record -a -g -o "$D/perf-locks.data" -- sleep 10
> > mpstat -P ALL 1 10
> >
> > Before and after the same window I sampled the counters under:
> >
> > /sys/kernel/slab/nfs_page/
> >
> > I also collected separate 10-second counter and mpstat windows under
> > the same workload configurations without perf lock.
> >
> > The resulting slab counter deltas were:
> >
> > A B C
> > unpinned/node0 node0/node0 unpinned/balanced
> >
> > free_fastpath
> > instrumented 655,917,745 457,936,140 328,670,030
> > uninstrumented 34,906,233 119,984,226 59,469,931
> >
> > free_slowpath
> > instrumented 236,735,724 7,720,917 270,716,506
> > uninstrumented 85,039,753 2,814 60,145,870
> >
> > sheaf_flush
> > instrumented 39,842,700 42,706,800 8,341,440
> > uninstrumented 1,286,400 13,487,700 887,700
> >
> > barn_put_fail
> > instrumented 663,994 711,745 139,037
> > uninstrumented 21,464 224,762 14,765
> >
> > barn_get_fail
> > instrumented 4,609,925 840,344 4,651,225
> > uninstrumented 1,438,643 224,787 1,017,006
> >
> > cmpxchg_double_fail
> > instrumented 70,524 8,077 28,062
> > uninstrumented 7,471 713 1,503
> >
> > alloc_slowpath
> > all cases 0 0 0
> >
> > The SLUB counter profile changes substantially with perf lock despite
> > the lower NFS throughput, and the exact mix depends strongly on
> > placement.
> >
> > The uninstrumented node0/node0 run also reproduces the barn/sheaf
> > relationship Chuck pointed out earlier. The nfs_page sheaf capacity is
> > 60:
> >
> > sheaf_flush / 60 = 13,487,700 / 60 = 224,795
> > barn_put_fail = 224,762
> >
> > The placement dependence of free_slowpath is also large. It falls from
> > about 85M events per 10 seconds in the unpinned/node0 case to 2,814 in
> > the node0/node0 case.
> >
> > The reported perf-lock result, however, is similar across all three
> > placements:
> >
> > contentions total wait average wait
> > unpinned/node0 4,782,839 14.39 min 180.47 us
> > node0/node0 4,582,012 14.13 min 185.05 us
> > unpinned/balanced 4,572,196 12.81 min 168.14 us
> >
> > I also revisited an inconsistency Chuck noticed in my earlier
> > measurements. Previously I had compared aggregate perf-lock wait from
> > one 10-second capture with CPU utilization measured during a different
> > window.
> >
> > I now have paired 10-second mpstat samples for each placement, with and
> > without perf lock. The node values below are averages of the per-CPU
> > %idle values for the CPUs in each NUMA node:
> >
> > system-wide node0 node1
> > %idle %idle %idle
> >
> > A. unpinned / CQs node0
> > uninstrumented 31.46 8.0 54.6
> > instrumented 3.60 0.06 7.1
> >
> > B. node0 pinned / CQs node0
> > uninstrumented 84.26 69.2 99.4
> > instrumented 12.07 10.4 13.8
> >
> > C. unpinned / balanced CQs
> > uninstrumented 66.86 66.7 66.9
> > instrumented 14.83 19.0 10.6
> >
> > This resolves the accounting inconsistency in my earlier measurements.
> > The large aggregate perf-lock wait and high idle percentage had come
> > from different windows. In the aligned samples, the system is much
> > busier during the perf-lock capture than in the corresponding
> > uninstrumented run.
> >
> > I am still unsure how representative the reported ~170-185 us average
> > wait is of the uninstrumented workload.
> >
> > The remaining number I am less sure how to interpret is
> > cmpxchg_double_fail. In the uninstrumented windows I see:
> >
> > unpinned / CQs node0: 7,471 / 10 sec
> > node0 pinned / CQs node0: 713 / 10 sec
> > unpinned / balanced CQs: 1,503 / 10 sec
> >
> > compared with fewer than 100 in Chuck's test.
> >
> > I understand that cmpxchg_double_fail counts failed slab freelist
> > updates rather than failed logical frees, so I am not sure what the
> > appropriate denominator is here. In particular, the node0/node0 case
> > still has 713 failures while sustaining full throughput and only 2,814
> > free_slowpath events over the interval.
> >
> > My questions are:
> >
> > 1. Do the uninstrumented cmpxchg_double_fail rates above look abnormal
> > for this workload/topology, or are they within the range one would
> > expect from this degree of concurrency and NUMA placement?
> >
> > 2. Is there a less invasive way you would recommend measuring the
> > nfs_page list_lock/freelist contention? I would like to distinguish
> > the steady-state behavior from what is observed during the perf-lock
> > capture.
> >
> > Thanks,
> > Tim
--
Cheers,
Harry / Hyeonggon
next parent reply other threads:[~2026-09-17 14:10 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <20260916232227.4098143-1-tmenninger@everpuredata.com>
[not found] ` <aqvq8-6IJBDer90O@thinkstation>
2026-09-17 14:10 ` Harry Yoo [this message]
2026-09-17 15:29 ` Peter Zijlstra
2026-09-17 16:25 ` Tim Menninger
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqvy8NqRRRmUlxjR@thinkstation \
--to=harry@kernel.org \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=akpm@linux-foundation.org \
--cc=alexander.shishkin@linux.intel.com \
--cc=cel@kernel.org \
--cc=cl@gentwo.org \
--cc=ebadger@everpuredata.com \
--cc=hao.li@linux.dev \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jcurley@everpuredata.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-nfs@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mingo@redhat.com \
--cc=namhyung@kernel.org \
--cc=peterz@infradead.org \
--cc=rientjes@google.com \
--cc=roman.gushchin@linux.dev \
--cc=tmenninger@everpuredata.com \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®