From: Zi Yan <ziy@nvidia.com>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Gregory Price <gourry@gourry.net>,
Li Zhe <lizhe.67@bytedance.com>,
Joshua Hahn <joshua.hahnjy@gmail.com>,
akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
rppt@kernel.org, mhocko@suse.com, corbet@lwn.net,
skhan@linuxfoundation.org, ying.huang@linux.alibaba.com,
apopple@nvidia.com, linux-mm@kvack.org,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
Date: Thu, 01 Oct 2026 07:18:11 -0400 [thread overview]
Message-ID: <4CDEE1DB-1847-42EB-9801-6B3968590C33@nvidia.com> (raw)
In-Reply-To: <280965b2-c6eb-4f01-a0ac-c4fb7e518f97@kernel.org>
On 1 Oct 2026, at 6:54, David Hildenbrand (Arm) wrote:
> On 9/30/26 17:10, Zi Yan wrote:
>> On 30 Sep 2026, at 10:50, Gregory Price wrote:
>>
>>> On Wed, Sep 30, 2026 at 10:02:23PM +0800, Li Zhe wrote:
>>>> The use case is closer to the second one, but the main motivation is not
>>>> only startup-time performance. The more important point is that each
>>>> workload has its own DDR/CXL budget assigned by the workload manager.
>>>>
>>>
>>> mempolicy is the wrong interface to do budget policy, you'd be better
>>> off looking at Joshua's memcg tiered node solutions for that.
>>>
>>>> MPOL_WEIGHTED_INTERLEAVE is useful because it lets new allocations
>>>> follow that per-workload budget from the beginning, instead of placing
>>>> everything on one tier first and correcting the placement later. This is
>>>> important for workload orchestration because the workload starts from a
>>>> placement close to its assigned DDR/CXL ratio.
>>>>
>>>
>>> This however is reasonable to me - you'd prefer to spread out the cost
>>> of initial faulting placement explicitly, rather than simply take
>>> fallbacks when the top-tier budget becomes pressured.
>>>
>>> i.e. w/o interleave:
>>>
>>> [node 0 ]
>>> ^^^^^^^^^^^^^^ alloc until full
>>> vvvvvvvvv fallback
>>> [node 1 ]
>>>
>>> in this scenario you end up with considerable hot memory
>>> on the remote node consolidated in time-space (everything
>>> allocated after node0 becomes full skews heavily toward
>>> node1)
>>>
>>> Tiering then likely takes many faults to rebalance after
>>> you've already reached node0 limits.
>>>
>>> w/ interleave
>>>
>>> [node 0 ]
>>> ^^^vv^^^vv^^^vv^^^vv^^^vv^^^vv....
>>> [node 1 ]
>>>
>>> In this scenario you do an initial fill distributed by weight
>>> and then let tiering figure it out without consolidating all
>>> of the pressure to the point where node0 has no space left.
>>>
>>> That said - this seems like mostly an initial-fill problem, after
>>> you initially fill your memory, a new allocation largely implies
>>> the memory is hot - and you probably prefer that to be local.
>>
>> If this is a initial-fill problem, can userspace set weighted interleave
>> initially? The program or harness can observe memory usage of the program
>> or related NUMA nodes and switch the policy to numa balancing via
>> set_mempolicy() or mbind() without MPOL_MF_MOVE after certain threshold
>> is met?
>
> You mean: use the weighted policy initially and then switch to a NUMA-balancing
> one which doesn't involve the weights anymore?
>
> That makes more sense to me. Although I struggle to see why an effectively
> "let's put random memory on slow and others at hot" is a good starting point to
> later let if be fixed up by actual balancing/tiering.
Statistically speaking, unless the hot/cold page distribution follows a power-law
distribution, random placement usually provides an average placement result.
So it is not a bad start point.
>
> It all sounds a bit hackish. :)
Best Regards,
Yan, Zi
next prev parent reply other threads:[~2026-10-01 11:18 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 7:26 Li Zhe
2026-09-30 8:43 ` Gregory Price
2026-09-30 8:59 ` Li Zhe
2026-09-30 9:11 ` Gregory Price
2026-09-30 10:59 ` Joshua Hahn
2026-09-30 11:22 ` Li Zhe
2026-09-30 12:25 ` Gregory Price
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52 ` Li Zhe
2026-09-30 12:11 ` Joshua Hahn
2026-09-30 14:02 ` Li Zhe
2026-09-30 14:50 ` Gregory Price
2026-09-30 15:10 ` Zi Yan
2026-10-01 10:54 ` David Hildenbrand (Arm)
2026-10-01 11:18 ` Zi Yan [this message]
2026-10-01 13:26 ` Gregory Price
2026-09-30 12:03 ` Gregory Price
2026-09-30 12:32 ` David Hildenbrand (Arm)
2026-09-30 13:40 ` Gregory Price
2026-10-01 10:55 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4CDEE1DB-1847-42EB-9801-6B3968590C33@nvidia.com \
--to=ziy@nvidia.com \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=corbet@lwn.net \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=joshua.hahnjy@gmail.com \
--cc=liam@infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=lizhe.67@bytedance.com \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=rppt@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=ying.huang@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®