mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Davidlohr Bueso <dave@stgolabs.net>
To: Bharata B Rao <bharata@amd.com>
Cc: linux-kernel@vger.kernel.org, linux-mm@kvack.org,
	jic23@kernel.org, dave.hansen@intel.com, gourry@gourry.net,
	mgorman@techsingularity.net, mingo@redhat.com,
	peterz@infradead.org, raghavendra.kt@amd.com, riel@surriel.com,
	rientjes@google.com, sj@kernel.org, weixugc@google.com,
	willy@infradead.org, ying.huang@linux.alibaba.com,
	ziy@nvidia.com, nifan.cxl@gmail.com, xuezhengchu@huawei.com,
	yiannis@zptcorp.com, akpm@linux-foundation.org, david@kernel.org,
	byungchul@sk.com, kinseyho@google.com, joshua.hahnjy@gmail.com,
	yuanchu@google.com, balbirs@nvidia.com, alok.rathore@samsung.com,
	shivankg@amd.com, donettom@linux.ibm.com
Subject: Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure
Date: Tue, 29 Sep 2026 12:56:46 -0700	[thread overview]
Message-ID: <20260929195646.4ttbl22clnrrktg7@offworld> (raw)
In-Reply-To: <31442591-d020-47b5-8f1c-b87fb632c226@amd.com>

On Sun, 27 Sep 2026, Bharata B Rao wrote:

>Hotness promotion engine
>------------------------
>This is not something that was written for pghot from scratch, but instead it is
>the same hot page promotion engine that is part of NUMAB2 which is now
>generalized and moved to pghot. So this is not the complexity that pghot
>introduces afresh.
>
>So considering all these, I see pghot as a light-weight and low-overhead
>mechanism to track per-PFN hotness and do async batch migration. Initial
>versions had fancy double data structures; a large hash a small binary tree of
>promotion-ready records and associated synchronization mechanism, but that is
>all past now.

I think we are all in agreement that the async batch migration is wanted.

>
>pghot interface for sampling
>============================
>pghot_record_access() interface was designed keeping the existing NUMAB2 source
>in mind. It fits that and it fits other sources like IBS Memory Profiler. So
>sampling sources report an access and the shared promotion engine acts upon it.
>
>But for sources like CHMU, from what you describe, I gather that a bulk
>reporting interface plus an indication to bypass the engine to treat the PFNs as
>migrate-ready, is what is required. Should those migrate-ready PFNs go through
>the regular pghot tracking (getting into section hotmaps to be picked up by
>kmigrated) or even that should be bypassed?
>
>However, in the context of PTE A bit based source, I have often thought about
>extending the interface for
>
>- bulk reporting where more than one PFN gets reported.
>- indicating the bypass options (frequency check bypass, recency check bypass etc)
>	
>NUMAB2 has to perform better in pghot
>=====================================
>pghot is about a sub-system that makes it possible to have multiple sources to
>coexist with reuse of common hot page promotion engine.
>
>NUMAB2 source resides within the scheduler and the promotion engine is also part
>of the scheduler. Through pghot, I am separating the source (NUMA hint faults)
>from the engine and moving that existing engine into pghot, to a common place
>where it gets reused for other sources as well.
>
>It is the same NUMA hint faults and more or less the same engine and hence my
>main objective is to ensure that there is no regression during this move.
>Additional performance optimizations can be done to the engine itself separately
>but that shouldn't be the baseline expectation from pghot.
>
>Is moving hot page promotion out of scheduler into a dedicated system, a good
>thing in general? I believe so as scheduler isn't the right place for it to
>reside. However I would like to hear from scheduler folks on this.

So two of the autonuma balancing og authors are scheduler experts - and iirc
*the* reason back then was locality. And Peter has already nacked the IBS
stuff in the past. But indeed the batch async part would be good to get nack/ack;
albeit the cgroup charging situation.

>
>Do we even need a centralized hot page promotion engine?
>========================================================
>NUMAB2 is good and serves as a good baseline for any new source that comes up.
>But with different kinds of sources becoming available, do we want all of them
>to  duplicate the hot page heuristics and promote hot pages on their own? I
>thought that may not be preferable and hence started this effort.

What are these sources that will become available?

>CXL HMU may not need the promotion engine,  but IBS Memory Profiler needs. It
>needs a promotion engine with full recency and frequency considerations before
>promoting. I don't think an arch driver like IBS Memory Profiler should be doing
>hotness heuristics within itself but instead be using the existing engine. In
>fact in my early posts, the driver based on primary IBS instance was feeding the
>"access sample" as "NUMA hint fault" to NUMAB1/B2 so that rest of NUMA
>Balancing/Hot page promotion engine just worked. But I think pghot is a better
>approach than that.

I agree that the IBS driver should not be doing hotness heuristics, it's the hw
that should.

>
>Why full pghot? Isn't async batch migration enough?
>===================================================
>Some of the above reasons apply but during the course of iterations, I have had
>implementations of just the migrator (kmigrated [1]).
>
>If every sub-system/source has intelligence of its own and just wants to
>handover a list of pages to async migrator thread, that's not much of an effort
>as this implementation showed.
>
>But then if some source needs rate-limiting and another source needs only
>hottest pages to be promoted, then again we overlap with the existing NUMAB2 engine.
>
>Then if we want to be slightly generic and want two sources to complement each
>other or the promoter to differentiate between lukewarm vs hot pages/regions
>then we may have to maintain hotness records and may soon end up with something
>similar to pghot's hotness tracking and reporting mechanism.
>
>Hardware sources have to out-perform NUMAB2
>===========================================
>Different sources will have different characteristics and capabilities and will
>help different workloads differently. So it is the choice that one could
>provide. Sometimes sources can complement each other as well.
>
>What IBS Memory Profiler has shown is that it can match and/or exceed (for
>Graph500, ptr-chase, llama) NUMAB2 with no hint faults overhead [2]. Also please
>check the initial XSBench numbers in my inline reply.

It's all about the numbers, and the ones you have just don't really sell - which is
one of the reasons this has been going on for years. If numa balancing didn't
exist, then maybe adding all this would make sense. For the XSBench I don't think
making decisions based on the non-overcommitted case is worthwhile.

>
>Cost of an unused source
>========================
>Not all the sources are required for every situation. Sources can be disabled at
>compile time or not enabled at run time with no cost or effect on other sources.
>But hotmap allocations would remain as a static cost even when no source is
>enabled (built out at compile time, the map is gone; sources off at runtime, the
>map remains)
>
>Tracking granularity: per-PFN vs region
>=======================================
>Often times this question comes up when pghot is compared with DAMON.
>
>Firstly, the baseline that we have (NUMAB2), tracks hotness at page granularity.
>To support this, pghot started with per-PFN granularity. Naturally two concerns
>come up:
>
>1. Memory overhead: I have shown the numbers above. It is lower-tier only and
>not much IMHO.
>
>2. Scan/CPU overhead: pghot started with scanning all PFNs, that was expensive.
>Then it introduced hotness bit per memory section so that only those sections
>which are marked hot are scanned. Now I have added (yet to be posted) a
>sub-section level hotness tracking where a hotness bit is maintained for each
>fixed 2M region within a section. This has considerably reduced the CPU overhead
>for kmigrated thread. While the ptr-chase numbers that I shared with the
>separated out IBS RFC v0 post of IBS Memory Profiler [2] does show the kmigrated
>utilization numbers with sub-section tracking, I plan to have some more numbers
>ready for LPC.

I look forward to seeing any new numbers you have. Region granularity is certainly
more aligned with willy as well as chmu.

>I am beginning to feel that this may be a good middle ground between real
>region-level tracking (where all pages of region are promoted irrespective of
>their real hotness) vs scanning at sub-section granularity but performing
>per-PFN precise promotion.

[...]

>I would be interested to understand more about how you ran XSBench, the
>parameters used, the promotion stats, if demotion stats etc. Do share when you
>get time.

The actual command is 'XSBench -g 90424' so only set the gridpoints, everything
else is default, so total ~44Gb footprint. I don't have the vmstats currently
(I do not run these benchmarks) but will share them once I get them - I can
affirm that demotion is in fact enabled, so full TPP up and down.

For the chmu: 32GB device (1:1 dram and cxl), this is with a 4k unit size, 1s
epoch, reporting mode is always on, threshold value is 1024.

The numa balancing mode was set to 3, the rest used the default values:
pghot_freq_threshold=2, pghot_promote_freq_window_ms=3000,
pghot_promote_rage_limit_MBps=65536, kmigrated_sleep_ms=100, kmigrated_batch_nr=512.

I also have numbers for two more benchmarks with a real chmu, with basically
the same numab parameters:

(i) TaoBench almost 2x throughput, going from ~250 qps to ~480 qps, vs NUMAB3.
     (dram:cxl is 16:32Gb with a 32Gb memsize, num_clients=2, clients_per_thread=75)

(ii) Graph500 only shows a smaller ~20% improvement vs NUMAB3 (bfs mean time drops
      from 5.0 to 4.15 secs). This was for a memory ratio dram:cxl as 64:32Gb. The
      algorithm is BFS, edge factor 512, 16 mpi processes.
      (mpiexec.openmpi -n 16 ./graph500_reference_bfs 22 512)

>
>I did a round of testing to compare base-NUMAB2, pghot-hintfaults(NUMAB2) and
>pghot-hwhints (IBS Memory Profiler). I am still experimenting with options,
>placement etc but here are my initial numbers:
>
>Non-Overcommitted case: XSBench working set fits fully within toptier but starts
>on lower tier before the measurement phase

Those are nice numbers, but I don't think this is the methodology to use...
It is more representative for the workload's working set to be > total dram and
therefore spill into slower tier(s), instead of artificially starting in the slow
memory and moving up. Do you have data for the over committed case?

Thanks,
Davidlohr

>(XSBench -s XL -g 238847 -t 64 -p 60000000 -l 34, footprint ~116.56 GiB)
>
>			Runtime (s)	lookups/s	Promotions (pages)
>base-NUMAB0		498.6		4.09M		0
>base-NUMAB2		422.4		4.83M		30.5M
>pghot-hintfaults	329.6		6.20M		30.5M
>pghot-hwhints		109.2		18.69M		2.7M
>
>Same amount of pages promoted by both base-NUMAB2 and pghot-hintfaults but
>runtime is better with pghot. This could be the benefit of async batched
>migration showing and no adverse effect of losing cache locality.
>
>pghot-hwhints shows good results. As I said this is just a first peek to the
>experimental numbers, I should have more concrete numbers and conclusion in LPC.
>
>[1] Kmigrated -
>https://lore.kernel.org/linux-mm/20250616133931.206626-1-bharata@amd.com/#t
>[2] IBS Memory Profiler RFC v0 -
>https://lore.kernel.org/linux-mm/20260924062206.319314-1-bharata@amd.com/
>
>Regards,
>Bharata.

      reply	other threads:[~2026-09-29 20:13 UTC|newest]

Thread overview: 56+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-28  5:43 Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 1/8] mm: migrate: Allow misplaced migration without VMA Bharata B Rao
2026-09-29 22:06   ` Davidlohr Bueso
2026-07-28  5:43 ` [PATCH v8 2/8] mm: migrate: Add promote_misplaced_memcg_folios() Bharata B Rao
2026-07-30  6:34   ` Bharata B Rao
2026-09-29 22:08   ` Davidlohr Bueso
2026-07-28  5:43 ` [PATCH v8 3/8] mm: Hot page tracking and promotion - pghot Bharata B Rao
2026-07-31 16:14   ` Bharata B Rao
2026-09-27 23:25   ` Davidlohr Bueso
2026-09-28  4:14     ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 4/8] mm: pghot: Precision mode for pghot Bharata B Rao
2026-07-31 16:27   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 5/8] mm: sched: move NUMA balancing tiering promotion to pghot Bharata B Rao
2026-08-03  8:23   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 6/8] x86/ibs: Move IBS caps definitions into its own header Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 7/8] x86/mm/ibs: In-kernel driver for AMD IBS Memory Profiler Bharata B Rao
2026-08-04  5:00   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 8/8] x86/mm/ibs: Add runtime controls for IBS memprofiler Bharata B Rao
2026-08-04  5:20   ` Bharata B Rao
2026-07-28  5:55 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - microbenchmark numbers Bharata B Rao
2026-07-28  5:59 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - NAS BT Bharata B Rao
2026-07-28  6:02 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - Graph500 Bharata B Rao
2026-07-28  6:05 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - redis-memtier Bharata B Rao
2026-07-28  6:17 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - llama-bench Bharata B Rao
2026-07-28 18:14 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure Andrew Morton
2026-07-28 18:24   ` Matthew Wilcox
2026-07-28 18:57     ` Gregory Price
2026-07-28 19:20       ` David Hildenbrand (Arm)
2026-07-28 19:59         ` Gregory Price
2026-07-29 11:45         ` Bharata B Rao
2026-08-10  3:38     ` Yongting Lin
2026-08-10  4:16       ` Matthew Wilcox
2026-08-10  5:35         ` Bharata B Rao
2026-08-11  7:15         ` Yongting Lin
2026-08-13  2:21         ` Gregory Price
2026-08-10 14:37       ` SJ Park
2026-08-11  6:37         ` Yongting Lin
2026-07-29  9:35   ` Bharata B Rao
2026-07-29 13:54     ` SJ Park
2026-08-04  1:23       ` SJ Park
2026-08-06  5:49   ` Bharata B Rao
2026-08-06 13:44     ` SJ Park
2026-08-10  4:46       ` Bharata B Rao
2026-08-10 14:25         ` SJ Park
2026-09-11 21:08 ` Joshua Hahn
2026-09-16  3:08   ` Bharata B Rao
2026-09-16 20:52     ` Joshua Hahn
2026-09-17  5:23       ` Bharata B Rao
2026-09-27 23:21     ` Davidlohr Bueso
2026-09-28  4:13       ` Bharata B Rao
2026-09-25  2:12 ` Davidlohr Bueso
2026-09-25  9:50   ` SJ Park
2026-09-25 16:30     ` Davidlohr Bueso
2026-09-25 16:57       ` Gregory Price
2026-09-27 15:51   ` Bharata B Rao
2026-09-29 19:56     ` Davidlohr Bueso [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929195646.4ttbl22clnrrktg7@offworld \
    --to=dave@stgolabs.net \
    --cc=akpm@linux-foundation.org \
    --cc=alok.rathore@samsung.com \
    --cc=balbirs@nvidia.com \
    --cc=bharata@amd.com \
    --cc=byungchul@sk.com \
    --cc=dave.hansen@intel.com \
    --cc=david@kernel.org \
    --cc=donettom@linux.ibm.com \
    --cc=gourry@gourry.net \
    --cc=jic23@kernel.org \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kinseyho@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mgorman@techsingularity.net \
    --cc=mingo@redhat.com \
    --cc=nifan.cxl@gmail.com \
    --cc=peterz@infradead.org \
    --cc=raghavendra.kt@amd.com \
    --cc=riel@surriel.com \
    --cc=rientjes@google.com \
    --cc=shivankg@amd.com \
    --cc=sj@kernel.org \
    --cc=weixugc@google.com \
    --cc=willy@infradead.org \
    --cc=xuezhengchu@huawei.com \
    --cc=yiannis@zptcorp.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=yuanchu@google.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®