mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Li Zhe" <lizhe.67@bytedance.com>
To: "David Hildenbrand (Arm)" <david@kernel.org>,
	 <akpm@linux-foundation.org>, <ljs@kernel.org>,
	<liam@infradead.org>,  <rppt@kernel.org>, <mhocko@suse.com>,
	<corbet@lwn.net>,  <skhan@linuxfoundation.org>, <ziy@nvidia.com>,
	<joshua.hahnjy@gmail.com>,  <gourry@gourry.net>,
	<ying.huang@linux.alibaba.com>,  <apopple@nvidia.com>
Cc: <linux-mm@kvack.org>, <linux-doc@vger.kernel.org>,
	 <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
Date: Wed, 30 Sep 2026 19:52:05 +0800	[thread overview]
Message-ID: <808292b4-841c-45bd-aacd-b487be1858ca@bytedance.com> (raw)
In-Reply-To: <cfdcaaa3-a9e7-4352-ab01-4922ebb8e0c9@kernel.org>

On 9/30/26 7:26 PM, David Hildenbrand (Arm) wrote:
> On 9/30/26 09:26, Li Zhe wrote:
>> MPOL_WEIGHTED_INTERLEAVE is useful on tiered-memory systems because it
>> can seed a workload's new allocations across fast memory and slower
>> capacity memory according to a configured ratio.
>>
>> That initial placement is useful for workload managers and orchestration
>> systems.  They can take the amount of fast memory and slower capacity
>> memory on a machine into account before starting a workload, and choose a
>> weighted policy that seeds the workload across the tiers at allocation
>> time.  This avoids starting from an all-fast or all-slow placement and
>> then relying on promotion or demotion to reshape a large working set.
>>
>> After those pages have been placed, however, the policy cannot currently
>> opt in to migrate-on-fault placement.  set_mempolicy() and mbind()
>> reject MPOL_WEIGHTED_INTERLEAVE when MPOL_F_NUMA_BALANCING is specified,
>> so memory tiering cannot promote hot pages that were initially placed on
>> the slower nodes by the weighted policy.
>>
>> Initial placement is only a starting point.  Pages initially allocated on
>> fast memory are not necessarily the long-term hot pages, and pages
>> initially allocated on slower memory may become hot as the workload's hot
>> set changes.  The policy therefore needs to be able to combine weighted
>> initial placement with memory tiering's NUMA fault based hot-page
>> promotion.
> Ok, so memory allocation will respect the weights but balancing will ignore
> them? That really sounds rather odd to me.
>
> And I assume that was the reason why we might have disallowed the combination:
> it turns a weighted mechanism into an unweighted mechanism.
>
> So are we really sure these semantics that you would essentially set in stone
> here are the semantics we want? (ignoring weights)
Yes, that is a fair concern. It does look odd if the weights are
interpreted as a hard resident placement ratio.

My understanding of the existing MPOL_WEIGHTED_INTERLEAVE ABI is that
the weights are allocation weights, not a long-term resident ratio. The
sysfs ABI documentation says that these weights only affect new
allocations, and that changing them at runtime will not migrate already
allocated pages.  The implementation also has normal allocation fallback
if the selected weighted target cannot satisfy the allocation.

The use case here follows that interpretation.  The weights are used to
seed the initial placement across memory tiers.  After that, with an
explicit MPOL_F_NUMA_BALANCING opt-in, memory tiering would promote hot
pages based on access patterns, not based on the original allocation
ratio.  Users that want weighted allocation without that behavior would
keep using MPOL_WEIGHTED_INTERLEAVE without MPOL_F_NUMA_BALANCING.

That said, I agree that allowing this flag combination would set the
semantics for it.  If reusing MPOL_F_NUMA_BALANCING for this is too
ambiguous, do you think we should model this as a separate opt-in ABI
for "weighted initial placement plus access-based balancing" instead?
For example, a separate flag or policy mode would make it clearer that
the weights are not intended to constrain migrate-on-fault placement.

If the allocation-only interpretation of the weights is acceptable, I
can make it explicit in the commit message and documentation in v2.
Otherwise I would appreciate your suggestion on the preferred interface.

Thanks,
Zhe

  reply	other threads:[~2026-09-30 11:52 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  7:26 Li Zhe
2026-09-30  8:43 ` Gregory Price
2026-09-30  8:59   ` Li Zhe
2026-09-30  9:11     ` Gregory Price
2026-09-30 10:59       ` Joshua Hahn
2026-09-30 11:22       ` Li Zhe
2026-09-30 12:25         ` Gregory Price
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52   ` Li Zhe [this message]
2026-09-30 12:11     ` Joshua Hahn
2026-09-30 14:02       ` Li Zhe
2026-09-30 14:50         ` Gregory Price
2026-09-30 15:10           ` Zi Yan
2026-09-30 12:03   ` Gregory Price
2026-09-30 12:32     ` David Hildenbrand (Arm)
2026-09-30 13:40       ` Gregory Price

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=808292b4-841c-45bd-aacd-b487be1858ca@bytedance.com \
    --to=lizhe.67@bytedance.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=joshua.hahnjy@gmail.com \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=rppt@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®