From: Ehab Ababneh <ehab.ababneh@intel.com>
To: ziy@nvidia.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org
Cc: buddy.lumpkin@oracle.com, akpm@linux-foundation.org,
kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev,
baohua@kernel.org, axelrasmussen@google.com, yuanchu@google.com,
weixugc@google.com, david@kernel.org, ljs@kernel.org,
liam@infradead.org, vbabka@kernel.org, rppt@kernel.org,
surenb@google.com, mhocko@suse.com, jackmanb@google.com,
hannes@cmpxchg.org, baoquan.he@linux.dev,
baolin.wang@linux.alibaba.com,
Ehab Ababneh <ehab.ababneh@intel.com>
Subject: [RFC PATCH v2 0/5] mm/vmscan: adaptive multi-threaded kswapd for NUMA-aware reclaim
Date: Fri, 9 Oct 2026 13:32:59 -0700 [thread overview]
Message-ID: <cover.1791577417.git.ehab.ababneh@intel.com> (raw)
In-Reply-To: <SJ1PR11MB6105E073723CFEEE8C8EDBF48C832@SJ1PR11MB6105.namprd11.prod.outlook.com>
This is a follow-up to the RFC for adaptive multi-threaded kswapd
on NUMA. It handles concurrent MGLRU aging, adjusts wakeups to node
CPU load, and scales the hopeless threshold for multiple workers.
Under memory pressure, an lruvec can end up with only two MGLRU
generations. That leaves an older and a younger cohort, a bit like
inactive and active LRU lists, though MGLRU tracks age differently.
Regarding the concerns raised in review:
- In the updated Cassandra comparison, four kswapd threads reduced p99
latency from 6.5 ms to 3.35 ms (about 49%) and increased throughput by
10.3%. We believe the latency gain comes from the reclaim design,
not from changes to LRU locking.
- This newer result compares four kswapd threads with one. The earlier
7.3% throughput result was measured with eight threads.
We also saw periods with zero file scans. The MIN_NR_GENS guard may
contribute: when only two generations remain, it makes scan_folios()
return without scanning the oldest one. This patch tests whether letting
kswapd scan that generation helps reclaim progress. The current traces
do not tie the zero-scan events to the same lruvec, so this remains a
hypothesis; more per-lruvec tracing is needed to compare it with lock
contention.
Changes since v1:
- Add a diagnostic kswapd scan path at MIN_NR_GENS.
- Scale the hopeless-node retry threshold by the active worker count.
Testing:
- Cassandra workload with one and four kswapd threads.
Buddy Lumpkin (1):
vmscan: Support multiple kswapd threads per node
Ehab Ababneh (4):
mm/vmscan: silence spurious max_seq race warning
mm/vmscan: make kswapd wakeups NUMA load-aware
mm/vmscan: let kswapd scan minimum-generation lruvecs
mm/vmscan: scale hopeless threshold for kswapd workers
include/linux/mmzone.h | 5 +-
include/trace/events/vmscan.h | 32 +++
mm/compaction.c | 8 +-
mm/internal.h | 3 +
mm/page_alloc.c | 26 +++
mm/vmscan.c | 426 ++++++++++++++++++++++++++++++++--
6 files changed, 473 insertions(+), 27 deletions(-)
--
2.43.0
next prev parent reply other threads:[~2026-10-09 20:30 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-08 22:10 [RFC PATCH 0/3] " Ehab Ababneh
2026-09-08 22:10 ` [RFC PATCH 1/3] vmscan: Support multiple kswapd threads per node Ehab Ababneh
2026-09-08 22:10 ` [RFC PATCH 2/3] mm/vmscan: handle racing max_seq advancement Ehab Ababneh
2026-09-08 22:10 ` [RFC PATCH 3/3] mm/vmscan: make kswapd wakeups NUMA load-aware Ehab Ababneh
2026-09-09 2:16 ` [RFC PATCH 0/3] mm/vmscan: adaptive multi-threaded kswapd for NUMA-aware reclaim Zi Yan
2026-09-22 21:23 ` Ababneh, Ehab
2026-10-09 20:32 ` Ehab Ababneh [this message]
2026-10-09 20:33 ` [RFC PATCH v2 1/5] vmscan: Support multiple kswapd threads per node Ehab Ababneh
2026-10-10 8:51 ` Barry Song
2026-10-09 20:33 ` [RFC PATCH v2 2/5] mm/vmscan: silence spurious max_seq race warning Ehab Ababneh
2026-10-09 20:33 ` [RFC PATCH v2 3/5] mm/vmscan: make kswapd wakeups NUMA load-aware Ehab Ababneh
2026-10-09 20:33 ` [RFC PATCH v2 4/5] mm/vmscan: let kswapd scan minimum-generation lruvecs Ehab Ababneh
2026-10-09 20:33 ` [RFC PATCH v2 5/5] mm/vmscan: scale hopeless threshold for kswapd workers Ehab Ababneh
2026-10-09 20:40 ` [RFC PATCH v2 0/5] mm/vmscan: adaptive multi-threaded kswapd for NUMA-aware reclaim Lorenzo Stoakes (ARM)
2026-10-10 8:19 ` Barry Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cover.1791577417.git.ehab.ababneh@intel.com \
--to=ehab.ababneh@intel.com \
--cc=akpm@linux-foundation.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=baoquan.he@linux.dev \
--cc=buddy.lumpkin@oracle.com \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=jackmanb@google.com \
--cc=kasong@tencent.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=qi.zheng@linux.dev \
--cc=rppt@kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=yuanchu@google.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®