mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Bharata B Rao <bharata@amd.com>
To: Matthew Wilcox <willy@infradead.org>,
	Yongting Lin <linyongting@bytedance.com>
Cc: <Jonathan.Cameron@huawei.com>, <akpm@linux-foundation.org>,
	<alok.rathore@samsung.com>, <balbirs@nvidia.com>,
	<byungchul@sk.com>, <dave.hansen@intel.com>, <dave@stgolabs.net>,
	<david@kernel.org>, <donettom@linux.ibm.com>, <gourry@gourry.net>,
	<joshua.hahnjy@gmail.com>, <kinseyho@google.com>,
	<linux-kernel@vger.kernel.org>, <linux-mm@kvack.org>,
	<mgorman@techsingularity.net>, <mingo@redhat.com>,
	<nifan.cxl@gmail.com>, <peterz@infradead.org>,
	<raghavendra.kt@amd.com>, <riel@surriel.com>,
	<rientjes@google.com>, <shivankg@amd.com>, <sj@kernel.org>,
	<weixugc@google.com>, <xuezhengchu@huawei.com>,
	<yiannis@zptcorp.com>, <ying.huang@linux.alibaba.com>,
	<yuanchu@google.com>, <ziy@nvidia.com>
Subject: Re: [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure
Date: Mon, 10 Aug 2026 11:05:37 +0530	[thread overview]
Message-ID: <bf4d7be1-5fb3-4d26-b6ed-4a835f77d424@amd.com> (raw)
In-Reply-To: <anlQp9LQw471t3YU@casper.infradead.org>

On 10-Aug-26 9:46 AM, Matthew Wilcox wrote:
> On Mon, Aug 10, 2026 at 11:38:14AM +0800, Yongting Lin wrote:
>> On Tue, Jul 28, 2026 at 07:24:46PM +0100, Matthew Wilcox wrote:
>>> I don't think we should do this.  I don't think it's a use-case that
>>> customers actually want; I think it's something CPU vendors think their
>>> customers want.  I've said this repeatedly over the months this patchset
>>> has been in development.  I say no.
>>
>> From our perspective, this is not merely a hypothetical use case. We
>> are actively developing and evaluating CXL tiered-memory systems, and
>> expect promoting CXL-resident pages that become hot back to DRAM to be
>> a practical requirement.
> 
> But what *is* your use case?  Truly hot data ends up in L1/L2/L3 no
> matter whether it came across the ridiculously high latency CXL link or
> from regular DRAM.  So we're talking about "warm" data that's presumably
> measured in gigabytes (since current server CPUs have about half a
> gigabyte of LLC)
> 
> I see the basis for saying "this task is low priority, it gets its memory
> allocated on CXL" and "this task is high priority, it gets its memory
> allocated on DRAM".  I don't see the use case where we're trying to figure
> out that region A of this task is sufficiently warmer than region B, and so
> we want to demote region B back to CXL and promote region A back to DRAM.
> 
> I particularly don't see a case for trying to do it on single-page
> granularity.  There's just too much damn information to track.  Maybe at
> a 2MB or 1GB boundary, but not per page.

Replying to the observation about per-page tracking information...

Since pghot has seen several versions and implementations over time, things have
changed multiple times and it is also easy to miss the information about what
gets tracked per-page and how it gets used from my long cover note. Hence
pasting some info here if it helps.

How pghot works
===============
- Tracks frequency and last access time.
- Additionally, the accessing NUMA node ID (NID) for each recorded
  access is tracked in precision mode.
- These hotness parameters are maintained in a per-PFN hotness record
  within a per-section map hanging off the existing mem_section data
  structure. The map is RCU-protected and also carries a section-level
  "hot" flag used to gate scanning by kmigrated.
  - In default mode, one byte (u8) is used for the hotness record. 5 bits
    store time using a bucketing scheme representing up to 4s with
    HZ=1000. The default toptier NID (0) is used as the promotion target
    and can be changed via the vm.pghot_target_nid sysctl.
  - In precision mode, 4 bytes (u32) are used for each hotness record.
    14 bits store time, representing around 16s with HZ=1000, and the
    accessing NID is stored per-PFN as the promotion target.
- Classifies pages as hot based on configurable thresholds.
- Pages classified as hot are marked ready for migration using the ready
  bit (MSB of the hotness record in both modes).
- Per-lower-tier-node kmigrated threads periodically scan the PFNs of
  lower-tier nodes, checking the migration-ready bit to perform batched
  migrations. The scan interval and batch size are configurable via
  debugfs tunables.

Memory overhead
===============
Default mode: 1 byte per lower-tier PFN. For 1TB of lower-tier memory
this amounts to 256MB overhead (assuming 4K pages).

Precision mode: 4 bytes per lower-tier PFN. For 1TB of lower-tier memory
this amounts to 1GB overhead.

Regards,
Bharata.

  reply	other threads:[~2026-08-10  5:35 UTC|newest]

Thread overview: 36+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-28  5:43 Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 1/8] mm: migrate: Allow misplaced migration without VMA Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 2/8] mm: migrate: Add promote_misplaced_memcg_folios() Bharata B Rao
2026-07-30  6:34   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 3/8] mm: Hot page tracking and promotion - pghot Bharata B Rao
2026-07-31 16:14   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 4/8] mm: pghot: Precision mode for pghot Bharata B Rao
2026-07-31 16:27   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 5/8] mm: sched: move NUMA balancing tiering promotion to pghot Bharata B Rao
2026-08-03  8:23   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 6/8] x86/ibs: Move IBS caps definitions into its own header Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 7/8] x86/mm/ibs: In-kernel driver for AMD IBS Memory Profiler Bharata B Rao
2026-08-04  5:00   ` Bharata B Rao
2026-07-28  5:43 ` [PATCH v8 8/8] x86/mm/ibs: Add runtime controls for IBS memprofiler Bharata B Rao
2026-08-04  5:20   ` Bharata B Rao
2026-07-28  5:55 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - microbenchmark numbers Bharata B Rao
2026-07-28  5:59 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - NAS BT Bharata B Rao
2026-07-28  6:02 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - Graph500 Bharata B Rao
2026-07-28  6:05 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - redis-memtier Bharata B Rao
2026-07-28  6:17 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure - llama-bench Bharata B Rao
2026-07-28 18:14 ` [PATCH v8 0/8] mm: Hot page tracking and promotion infrastructure Andrew Morton
2026-07-28 18:24   ` Matthew Wilcox
2026-07-28 18:57     ` Gregory Price
2026-07-28 19:20       ` David Hildenbrand (Arm)
2026-07-28 19:59         ` Gregory Price
2026-07-29 11:45         ` Bharata B Rao
2026-08-10  3:38     ` Yongting Lin
2026-08-10  4:16       ` Matthew Wilcox
2026-08-10  5:35         ` Bharata B Rao [this message]
2026-07-29  9:35   ` Bharata B Rao
2026-07-29 13:54     ` SJ Park
2026-08-04  1:23       ` SJ Park
2026-08-06  5:49   ` Bharata B Rao
2026-08-06 13:44     ` SJ Park
2026-08-10  4:46       ` Bharata B Rao
2026-08-10 14:25         ` SJ Park

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=bf4d7be1-5fb3-4d26-b6ed-4a835f77d424@amd.com \
    --to=bharata@amd.com \
    --cc=Jonathan.Cameron@huawei.com \
    --cc=akpm@linux-foundation.org \
    --cc=alok.rathore@samsung.com \
    --cc=balbirs@nvidia.com \
    --cc=byungchul@sk.com \
    --cc=dave.hansen@intel.com \
    --cc=dave@stgolabs.net \
    --cc=david@kernel.org \
    --cc=donettom@linux.ibm.com \
    --cc=gourry@gourry.net \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kinseyho@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linyongting@bytedance.com \
    --cc=mgorman@techsingularity.net \
    --cc=mingo@redhat.com \
    --cc=nifan.cxl@gmail.com \
    --cc=peterz@infradead.org \
    --cc=raghavendra.kt@amd.com \
    --cc=riel@surriel.com \
    --cc=rientjes@google.com \
    --cc=shivankg@amd.com \
    --cc=sj@kernel.org \
    --cc=weixugc@google.com \
    --cc=willy@infradead.org \
    --cc=xuezhengchu@huawei.com \
    --cc=yiannis@zptcorp.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=yuanchu@google.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome