From: Ilya Gladyshev <ilya.gladyshev@linux.dev>
To: ilya.gladyshev@linux.dev
Cc: akpm@linux-foundation.org, andrew+netdev@lunn.ch,
apopple@nvidia.com, artem.kuzin@huawei.com,
baolin.wang@linux.alibaba.com, david@kernel.org,
Liam.Howlett@oracle.com, edumazet@google.com,
harry.yoo@oracle.com, hramamurthy@google.com, ivgorbunov@me.com,
joshwash@google.com, kirill@shutemov.name,
linux-kernel@vger.kernel.org, linux-mm@kvack.org,
lorenzo.stoakes@oracle.com, mhocko@suse.com,
muchun.song@linux.dev, pfalcato@suse.de, rppt@kernel.org,
surenb@google.com, torvalds@linuxfoundation.org, vbabka@suse.cz,
willy@infradead.org, yuzhao@google.com, ziy@nvidia.com
Subject: [PATCH v6 0/3] mm: improve folio refcount scalability
Date: Sat, 12 Sep 2026 22:50:07 +0300 [thread overview]
Message-ID: <cover.1789239015.git.ilya.gladyshev@linux.dev> (raw)
From: Gladyshev Ilya <ilya.gladyshev@linux.dev>
Recap
-----
This patchset addresses a scalability issue of a folio's add_unless() operation,
noticeable during contended IO reads from the same page [folio_try_get()]. The
main idea is to replace CAS loop with optimistic increment (and deal with
failure later). This requires splitting refcount into counter and separate
"dead/frozen" bit.
To allow for such modification, this patchset also slightly refactors page_ref
API, consolidating all implementation logic inside mm headers. For more
information, check individual commit messages. The original performance issue
and previous attempts by other people can be found in [1][2].
Performance
-----------
To my regret, I don't have any access to high-core CPUs that I can
benchmark on, and a 12 vcpu laptop isn't really a scalability test. So,
here I can only paste my previous measurements on Linux 6.15. To be fair,
none of the related code paths really changed, so I don't expect any changes in the
numbers here.
Performance was measured using a simple custom benchmark based on
will-it-scale[3]. This benchmark spawns N pinned threads/processes that
execute the following loop:
``
char buf[]
fd = open(/* same file in tmpfs */);
while (true) {
pread(fd, buf, /* read size = */ 64, /* offset = 0 */)
}
``
While this is a synthetic load, it does highlight existing issue and
doesn't differ much from the benchmarking in patch [2].
This benchmark measures operations per second in the inner loop and the
results across all workers. Performance was tested on top of v6.15 kernel
on two platforms. Since threads and processes showed similar performance on
both systems, only the thread results are provided below. The performance
improvement scales linearly between the CPU counts shown.
Platform 1: 2 x E5-2690 v3, 12C/12T each [disabled SMT]
#threads | vanilla | patched | boost (%)
1 | 1343381 | 1344401 | +0.1
2 | 2186160 | 2455837 | +12.3
5 | 5277092 | 6108030 | +15.7
10 | 5858123 | 7506328 | +28.1
12 | 6484445 | 8137706 | +25.5
/* Cross socket NUMA */
14 | 3145860 | 4247391 | +35.0
16 | 2350840 | 4262707 | +81.3
18 | 2378825 | 4121415 | +73.2
20 | 2438475 | 4683548 | +92.1
24 | 2325998 | 4529737 | +94.7
Platform 2: 2 x AMD EPYC 9654, 96C/192T each [enabled SMT]
#threads | vanilla | patched | boost (%)
1 | 1077276 | 1081653 | +0.4
5 | 4286838 | 4682513 | +9.2
10 | 1698095 | 1902753 | +12.1
20 | 1662266 | 1921603 | +15.6
49 | 1486745 | 1828926 | +23.0
97 | 1617365 | 2052635 | +26.9
/* Cross socket NUMA */
105 | 1368319 | 1798862 | +31.5
136 | 1008071 | 1393055 | +38.2
168 | 879332 | 1245210 | +41.6
/* SMT */
193 | 905432 | 1294833 | +43.0
289 | 851988 | 1313110 | +54.1
353 | 771288 | 1347165 | +74.7
Changes since v4 (last significant checkpoint)
---
- Lower Google GVE pagecnt_bias
- VM_BUG_ON -> VM_WARN_ON_ONCE
- Refactor missing API calls in mm/memory-failure.c
- rebase
- Make __page_is_frozen() public and introduce folio_is_frozen()
counterpart [3].
- Make commit messages more informative
Link to v4: https://lore.kernel.org/linux-mm/df26082871b4c65b2bd38d409026237c08572836@linux.dev/
[1]: https://lore.kernel.org/linux-mm/CAHk-=wj00-nGmXEkxY=-=Z_qP6kiGUziSFvxHJ9N-cLWry5zpA@mail.gmail.com/
[2]: https://lore.kernel.org/linux-mm/20251017141536.577466-1-kirill@shutemov.name/
[3]: https://lore.kernel.org/all/aqAIFV4nOGPbWiDS@thinkstation/
---
Ilya Gladyshev (3):
gve: reduce pagecnt_bias to USHRT_MAX
mm: drop page refcount zero state semantics
mm: implement page refcount locking via dedicated bit
.../ethernet/google/gve/gve_buffer_mgmt_dqo.c | 4 +-
drivers/net/ethernet/google/gve/gve_rx.c | 8 +--
drivers/net/ethernet/google/gve/gve_utils.c | 6 +-
drivers/net/ethernet/google/gve/gve_utils.h | 2 +-
drivers/pci/p2pdma.c | 4 +-
drivers/virtio/virtio_mem.c | 2 +-
include/linux/mm.h | 2 +-
include/linux/page-flags.h | 13 ++++
include/linux/page_ref.h | 69 ++++++++++++++++---
kernel/liveupdate/kexec_handover.c | 6 +-
lib/test_hmm.c | 4 +-
mm/hugetlb.c | 2 +-
mm/internal.h | 2 +-
mm/memory-failure.c | 8 +--
mm/memremap.c | 4 +-
mm/mm_init.c | 6 +-
mm/page_alloc.c | 10 +--
mm/page_frag_cache.c | 2 +-
18 files changed, 108 insertions(+), 46 deletions(-)
--
2.55.0
next reply other threads:[~2026-09-12 19:50 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-12 19:50 Ilya Gladyshev [this message]
2026-09-12 19:50 ` [PATCH v6 1/3] gve: reduce pagecnt_bias to USHRT_MAX Ilya Gladyshev
2026-09-12 19:50 ` [PATCH v6 2/3] mm: drop page refcount zero state semantics Ilya Gladyshev
2026-09-12 19:50 ` [PATCH v6 3/3] mm: implement page refcount locking via dedicated bit Ilya Gladyshev
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cover.1789239015.git.ilya.gladyshev@linux.dev \
--to=ilya.gladyshev@linux.dev \
--cc=Liam.Howlett@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=andrew+netdev@lunn.ch \
--cc=apopple@nvidia.com \
--cc=artem.kuzin@huawei.com \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=edumazet@google.com \
--cc=harry.yoo@oracle.com \
--cc=hramamurthy@google.com \
--cc=ivgorbunov@me.com \
--cc=joshwash@google.com \
--cc=kirill@shutemov.name \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=lorenzo.stoakes@oracle.com \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=pfalcato@suse.de \
--cc=rppt@kernel.org \
--cc=surenb@google.com \
--cc=torvalds@linuxfoundation.org \
--cc=vbabka@suse.cz \
--cc=willy@infradead.org \
--cc=yuzhao@google.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®