From: Davidlohr Bueso <dave@stgolabs.net>
To: Bharata B Rao <bharata@amd.com>
Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org,
jic23@kernel.org, dave.hansen@intel.com, gourry@gourry.net,
mgorman@techsingularity.net, mingo@redhat.com,
peterz@infradead.org, raghavendra.kt@amd.com, riel@surriel.com,
rientjes@google.com, sj@kernel.org, weixugc@google.com,
willy@infradead.org, ying.huang@linux.alibaba.com,
ziy@nvidia.com, nifan.cxl@gmail.com, xuezhengchu@huawei.com,
yiannis@zptcorp.com, akpm@linux-foundation.org, david@kernel.org,
byungchul@sk.com, kinseyho@google.com, joshua.hahnjy@gmail.com,
yuanchu@google.com, balbirs@nvidia.com, alok.rathore@samsung.com,
shivankg@amd.com, donettom@linux.ibm.com
Subject: Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure
Date: Tue, 29 Sep 2026 12:56:46 -0700 [thread overview]
Message-ID: <20260929195646.4ttbl22clnrrktg7@offworld> (raw)
In-Reply-To: <31442591-d020-47b5-8f1c-b87fb632c226@amd.com>
On Sun, 27 Sep 2026, Bharata B Rao wrote:
>Hotness promotion engine
>------------------------
>This is not something that was written for pghot from scratch, but instead it is
>the same hot page promotion engine that is part of NUMAB2 which is now
>generalized and moved to pghot. So this is not the complexity that pghot
>introduces afresh.
>
>So considering all these, I see pghot as a light-weight and low-overhead
>mechanism to track per-PFN hotness and do async batch migration. Initial
>versions had fancy double data structures; a large hash a small binary tree of
>promotion-ready records and associated synchronization mechanism, but that is
>all past now.
I think we are all in agreement that the async batch migration is wanted.
>
>pghot interface for sampling
>============================
>pghot_record_access() interface was designed keeping the existing NUMAB2 source
>in mind. It fits that and it fits other sources like IBS Memory Profiler. So
>sampling sources report an access and the shared promotion engine acts upon it.
>
>But for sources like CHMU, from what you describe, I gather that a bulk
>reporting interface plus an indication to bypass the engine to treat the PFNs as
>migrate-ready, is what is required. Should those migrate-ready PFNs go through
>the regular pghot tracking (getting into section hotmaps to be picked up by
>kmigrated) or even that should be bypassed?
>
>However, in the context of PTE A bit based source, I have often thought about
>extending the interface for
>
>- bulk reporting where more than one PFN gets reported.
>- indicating the bypass options (frequency check bypass, recency check bypass etc)
>
>NUMAB2 has to perform better in pghot
>=====================================
>pghot is about a sub-system that makes it possible to have multiple sources to
>coexist with reuse of common hot page promotion engine.
>
>NUMAB2 source resides within the scheduler and the promotion engine is also part
>of the scheduler. Through pghot, I am separating the source (NUMA hint faults)
>from the engine and moving that existing engine into pghot, to a common place
>where it gets reused for other sources as well.
>
>It is the same NUMA hint faults and more or less the same engine and hence my
>main objective is to ensure that there is no regression during this move.
>Additional performance optimizations can be done to the engine itself separately
>but that shouldn't be the baseline expectation from pghot.
>
>Is moving hot page promotion out of scheduler into a dedicated system, a good
>thing in general? I believe so as scheduler isn't the right place for it to
>reside. However I would like to hear from scheduler folks on this.
So two of the autonuma balancing og authors are scheduler experts - and iirc
*the* reason back then was locality. And Peter has already nacked the IBS
stuff in the past. But indeed the batch async part would be good to get nack/ack;
albeit the cgroup charging situation.
>
>Do we even need a centralized hot page promotion engine?
>========================================================
>NUMAB2 is good and serves as a good baseline for any new source that comes up.
>But with different kinds of sources becoming available, do we want all of them
>to duplicate the hot page heuristics and promote hot pages on their own? I
>thought that may not be preferable and hence started this effort.
What are these sources that will become available?
>CXL HMU may not need the promotion engine, but IBS Memory Profiler needs. It
>needs a promotion engine with full recency and frequency considerations before
>promoting. I don't think an arch driver like IBS Memory Profiler should be doing
>hotness heuristics within itself but instead be using the existing engine. In
>fact in my early posts, the driver based on primary IBS instance was feeding the
>"access sample" as "NUMA hint fault" to NUMAB1/B2 so that rest of NUMA
>Balancing/Hot page promotion engine just worked. But I think pghot is a better
>approach than that.
I agree that the IBS driver should not be doing hotness heuristics, it's the hw
that should.
>
>Why full pghot? Isn't async batch migration enough?
>===================================================
>Some of the above reasons apply but during the course of iterations, I have had
>implementations of just the migrator (kmigrated [1]).
>
>If every sub-system/source has intelligence of its own and just wants to
>handover a list of pages to async migrator thread, that's not much of an effort
>as this implementation showed.
>
>But then if some source needs rate-limiting and another source needs only
>hottest pages to be promoted, then again we overlap with the existing NUMAB2 engine.
>
>Then if we want to be slightly generic and want two sources to complement each
>other or the promoter to differentiate between lukewarm vs hot pages/regions
>then we may have to maintain hotness records and may soon end up with something
>similar to pghot's hotness tracking and reporting mechanism.
>
>Hardware sources have to out-perform NUMAB2
>===========================================
>Different sources will have different characteristics and capabilities and will
>help different workloads differently. So it is the choice that one could
>provide. Sometimes sources can complement each other as well.
>
>What IBS Memory Profiler has shown is that it can match and/or exceed (for
>Graph500, ptr-chase, llama) NUMAB2 with no hint faults overhead [2]. Also please
>check the initial XSBench numbers in my inline reply.
It's all about the numbers, and the ones you have just don't really sell - which is
one of the reasons this has been going on for years. If numa balancing didn't
exist, then maybe adding all this would make sense. For the XSBench I don't think
making decisions based on the non-overcommitted case is worthwhile.
>
>Cost of an unused source
>========================
>Not all the sources are required for every situation. Sources can be disabled at
>compile time or not enabled at run time with no cost or effect on other sources.
>But hotmap allocations would remain as a static cost even when no source is
>enabled (built out at compile time, the map is gone; sources off at runtime, the
>map remains)
>
>Tracking granularity: per-PFN vs region
>=======================================
>Often times this question comes up when pghot is compared with DAMON.
>
>Firstly, the baseline that we have (NUMAB2), tracks hotness at page granularity.
>To support this, pghot started with per-PFN granularity. Naturally two concerns
>come up:
>
>1. Memory overhead: I have shown the numbers above. It is lower-tier only and
>not much IMHO.
>
>2. Scan/CPU overhead: pghot started with scanning all PFNs, that was expensive.
>Then it introduced hotness bit per memory section so that only those sections
>which are marked hot are scanned. Now I have added (yet to be posted) a
>sub-section level hotness tracking where a hotness bit is maintained for each
>fixed 2M region within a section. This has considerably reduced the CPU overhead
>for kmigrated thread. While the ptr-chase numbers that I shared with the
>separated out IBS RFC v0 post of IBS Memory Profiler [2] does show the kmigrated
>utilization numbers with sub-section tracking, I plan to have some more numbers
>ready for LPC.
I look forward to seeing any new numbers you have. Region granularity is certainly
more aligned with willy as well as chmu.
>I am beginning to feel that this may be a good middle ground between real
>region-level tracking (where all pages of region are promoted irrespective of
>their real hotness) vs scanning at sub-section granularity but performing
>per-PFN precise promotion.
[...]
>I would be interested to understand more about how you ran XSBench, the
>parameters used, the promotion stats, if demotion stats etc. Do share when you
>get time.
The actual command is 'XSBench -g 90424' so only set the gridpoints, everything
else is default, so total ~44Gb footprint. I don't have the vmstats currently
(I do not run these benchmarks) but will share them once I get them - I can
affirm that demotion is in fact enabled, so full TPP up and down.
For the chmu: 32GB device (1:1 dram and cxl), this is with a 4k unit size, 1s
epoch, reporting mode is always on, threshold value is 1024.
The numa balancing mode was set to 3, the rest used the default values:
pghot_freq_threshold=2, pghot_promote_freq_window_ms=3000,
pghot_promote_rage_limit_MBps=65536, kmigrated_sleep_ms=100, kmigrated_batch_nr=512.
I also have numbers for two more benchmarks with a real chmu, with basically
the same numab parameters:
(i) TaoBench almost 2x throughput, going from ~250 qps to ~480 qps, vs NUMAB3.
(dram:cxl is 16:32Gb with a 32Gb memsize, num_clients=2, clients_per_thread=75)
(ii) Graph500 only shows a smaller ~20% improvement vs NUMAB3 (bfs mean time drops
from 5.0 to 4.15 secs). This was for a memory ratio dram:cxl as 64:32Gb. The
algorithm is BFS, edge factor 512, 16 mpi processes.
(mpiexec.openmpi -n 16 ./graph500_reference_bfs 22 512)
>
>I did a round of testing to compare base-NUMAB2, pghot-hintfaults(NUMAB2) and
>pghot-hwhints (IBS Memory Profiler). I am still experimenting with options,
>placement etc but here are my initial numbers:
>
>Non-Overcommitted case: XSBench working set fits fully within toptier but starts
>on lower tier before the measurement phase
Those are nice numbers, but I don't think this is the methodology to use...
It is more representative for the workload's working set to be > total dram and
therefore spill into slower tier(s), instead of artificially starting in the slow
memory and moving up. Do you have data for the over committed case?
Thanks,
Davidlohr
>(XSBench -s XL -g 238847 -t 64 -p 60000000 -l 34, footprint ~116.56 GiB)
>
> Runtime (s) lookups/s Promotions (pages)
>base-NUMAB0 498.6 4.09M 0
>base-NUMAB2 422.4 4.83M 30.5M
>pghot-hintfaults 329.6 6.20M 30.5M
>pghot-hwhints 109.2 18.69M 2.7M
>
>Same amount of pages promoted by both base-NUMAB2 and pghot-hintfaults but
>runtime is better with pghot. This could be the benefit of async batched
>migration showing and no adverse effect of losing cache locality.
>
>pghot-hwhints shows good results. As I said this is just a first peek to the
>experimental numbers, I should have more concrete numbers and conclusion in LPC.
>
>[1] Kmigrated -
>https://lore.kernel.org/linux-mm/20250616133931.206626-1-bharata@amd.com/#t
>[2] IBS Memory Profiler RFC v0 -
>https://lore.kernel.org/linux-mm/20260924062206.319314-1-bharata@amd.com/
>
>Regards,
>Bharata.
prev parent reply other threads:[~2026-09-29 20:13 UTC|newest]
Thread overview: 56+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-28 5:43 Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 1/8] mm: migrate: Allow misplaced migration without VMA Bharata B Rao
2026-09-29 22:06 ` Davidlohr Bueso
2026-07-28 5:43 ` [PATCH v8 2/8] mm: migrate: Add promote_misplaced_memcg_folios() Bharata B Rao
2026-07-30 6:34 ` Bharata B Rao
2026-09-29 22:08 ` Davidlohr Bueso
2026-07-28 5:43 ` [PATCH v8 3/8] mm: Hot page tracking and promotion - pghot Bharata B Rao
2026-07-31 16:14 ` Bharata B Rao
2026-09-27 23:25 ` Davidlohr Bueso
2026-09-28 4:14 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 4/8] mm: pghot: Precision mode for pghot Bharata B Rao
2026-07-31 16:27 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 5/8] mm: sched: move NUMA balancing tiering promotion to pghot Bharata B Rao
2026-08-03 8:23 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 6/8] x86/ibs: Move IBS caps definitions into its own header Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 7/8] x86/mm/ibs: In-kernel driver for AMD IBS Memory Profiler Bharata B Rao
2026-08-04 5:00 ` Bharata B Rao
2026-07-28 5:43 ` [PATCH v8 8/8] x86/mm/ibs: Add runtime controls for IBS memprofiler Bharata B Rao
2026-08-04 5:20 ` Bharata B Rao
2026-07-28 5:55 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - microbenchmark numbers Bharata B Rao
2026-07-28 5:59 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - NAS BT Bharata B Rao
2026-07-28 6:02 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - Graph500 Bharata B Rao
2026-07-28 6:05 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - redis-memtier Bharata B Rao
2026-07-28 6:17 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - llama-bench Bharata B Rao
2026-07-28 18:14 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure Andrew Morton
2026-07-28 18:24 ` Matthew Wilcox
2026-07-28 18:57 ` Gregory Price
2026-07-28 19:20 ` David Hildenbrand (Arm)
2026-07-28 19:59 ` Gregory Price
2026-07-29 11:45 ` Bharata B Rao
2026-08-10 3:38 ` Yongting Lin
2026-08-10 4:16 ` Matthew Wilcox
2026-08-10 5:35 ` Bharata B Rao
2026-08-11 7:15 ` Yongting Lin
2026-08-13 2:21 ` Gregory Price
2026-08-10 14:37 ` SJ Park
2026-08-11 6:37 ` Yongting Lin
2026-07-29 9:35 ` Bharata B Rao
2026-07-29 13:54 ` SJ Park
2026-08-04 1:23 ` SJ Park
2026-08-06 5:49 ` Bharata B Rao
2026-08-06 13:44 ` SJ Park
2026-08-10 4:46 ` Bharata B Rao
2026-08-10 14:25 ` SJ Park
2026-09-11 21:08 ` Joshua Hahn
2026-09-16 3:08 ` Bharata B Rao
2026-09-16 20:52 ` Joshua Hahn
2026-09-17 5:23 ` Bharata B Rao
2026-09-27 23:21 ` Davidlohr Bueso
2026-09-28 4:13 ` Bharata B Rao
2026-09-25 2:12 ` Davidlohr Bueso
2026-09-25 9:50 ` SJ Park
2026-09-25 16:30 ` Davidlohr Bueso
2026-09-25 16:57 ` Gregory Price
2026-09-27 15:51 ` Bharata B Rao
2026-09-29 19:56 ` Davidlohr Bueso [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929195646.4ttbl22clnrrktg7@offworld \
--to=dave@stgolabs.net \
--cc=akpm@linux-foundation.org \
--cc=alok.rathore@samsung.com \
--cc=balbirs@nvidia.com \
--cc=bharata@amd.com \
--cc=byungchul@sk.com \
--cc=dave.hansen@intel.com \
--cc=david@kernel.org \
--cc=donettom@linux.ibm.com \
--cc=gourry@gourry.net \
--cc=jic23@kernel.org \
--cc=joshua.hahnjy@gmail.com \
--cc=kinseyho@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mgorman@techsingularity.net \
--cc=mingo@redhat.com \
--cc=nifan.cxl@gmail.com \
--cc=peterz@infradead.org \
--cc=raghavendra.kt@amd.com \
--cc=riel@surriel.com \
--cc=rientjes@google.com \
--cc=shivankg@amd.com \
--cc=sj@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=xuezhengchu@huawei.com \
--cc=yiannis@zptcorp.com \
--cc=ying.huang@linux.alibaba.com \
--cc=yuanchu@google.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®