From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Zi Yan <ziy@nvidia.com>, Gregory Price <gourry@gourry.net>,
Li Zhe <lizhe.67@bytedance.com>,
Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
rppt@kernel.org, mhocko@suse.com, corbet@lwn.net,
skhan@linuxfoundation.org, ying.huang@linux.alibaba.com,
apopple@nvidia.com, linux-mm@kvack.org,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
Date: Thu, 1 Oct 2026 12:54:22 +0200 [thread overview]
Message-ID: <280965b2-c6eb-4f01-a0ac-c4fb7e518f97@kernel.org> (raw)
In-Reply-To: <A005B05D-A85C-4153-A65D-0BDF9D8500DF@nvidia.com>
On 9/30/26 17:10, Zi Yan wrote:
> On 30 Sep 2026, at 10:50, Gregory Price wrote:
>
>> On Wed, Sep 30, 2026 at 10:02:23PM +0800, Li Zhe wrote:
>>> The use case is closer to the second one, but the main motivation is not
>>> only startup-time performance. The more important point is that each
>>> workload has its own DDR/CXL budget assigned by the workload manager.
>>>
>>
>> mempolicy is the wrong interface to do budget policy, you'd be better
>> off looking at Joshua's memcg tiered node solutions for that.
>>
>>> MPOL_WEIGHTED_INTERLEAVE is useful because it lets new allocations
>>> follow that per-workload budget from the beginning, instead of placing
>>> everything on one tier first and correcting the placement later. This is
>>> important for workload orchestration because the workload starts from a
>>> placement close to its assigned DDR/CXL ratio.
>>>
>>
>> This however is reasonable to me - you'd prefer to spread out the cost
>> of initial faulting placement explicitly, rather than simply take
>> fallbacks when the top-tier budget becomes pressured.
>>
>> i.e. w/o interleave:
>>
>> [node 0 ]
>> ^^^^^^^^^^^^^^ alloc until full
>> vvvvvvvvv fallback
>> [node 1 ]
>>
>> in this scenario you end up with considerable hot memory
>> on the remote node consolidated in time-space (everything
>> allocated after node0 becomes full skews heavily toward
>> node1)
>>
>> Tiering then likely takes many faults to rebalance after
>> you've already reached node0 limits.
>>
>> w/ interleave
>>
>> [node 0 ]
>> ^^^vv^^^vv^^^vv^^^vv^^^vv^^^vv....
>> [node 1 ]
>>
>> In this scenario you do an initial fill distributed by weight
>> and then let tiering figure it out without consolidating all
>> of the pressure to the point where node0 has no space left.
>>
>> That said - this seems like mostly an initial-fill problem, after
>> you initially fill your memory, a new allocation largely implies
>> the memory is hot - and you probably prefer that to be local.
>
> If this is a initial-fill problem, can userspace set weighted interleave
> initially? The program or harness can observe memory usage of the program
> or related NUMA nodes and switch the policy to numa balancing via
> set_mempolicy() or mbind() without MPOL_MF_MOVE after certain threshold
> is met?
You mean: use the weighted policy initially and then switch to a NUMA-balancing
one which doesn't involve the weights anymore?
That makes more sense to me. Although I struggle to see why an effectively
"let's put random memory on slow and others at hot" is a good starting point to
later let if be fixed up by actual balancing/tiering.
It all sounds a bit hackish. :)
--
Cheers,
David
next prev parent reply other threads:[~2026-10-01 10:54 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 7:26 Li Zhe
2026-09-30 8:43 ` Gregory Price
2026-09-30 8:59 ` Li Zhe
2026-09-30 9:11 ` Gregory Price
2026-09-30 10:59 ` Joshua Hahn
2026-09-30 11:22 ` Li Zhe
2026-09-30 12:25 ` Gregory Price
2026-10-02 7:25 ` Li Zhe
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52 ` Li Zhe
2026-09-30 12:11 ` Joshua Hahn
2026-09-30 14:02 ` Li Zhe
2026-09-30 14:50 ` Gregory Price
2026-09-30 15:10 ` Zi Yan
2026-10-01 10:54 ` David Hildenbrand (Arm) [this message]
2026-10-01 11:18 ` Zi Yan
2026-10-01 13:26 ` Gregory Price
2026-10-02 7:56 ` David Hildenbrand (Arm)
2026-10-02 8:15 ` Li Zhe
2026-10-02 12:05 ` Gregory Price
2026-09-30 12:03 ` Gregory Price
2026-09-30 12:32 ` David Hildenbrand (Arm)
2026-09-30 13:40 ` Gregory Price
2026-10-01 10:55 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=280965b2-c6eb-4f01-a0ac-c4fb7e518f97@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=corbet@lwn.net \
--cc=gourry@gourry.net \
--cc=joshua.hahnjy@gmail.com \
--cc=liam@infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=lizhe.67@bytedance.com \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=rppt@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®