mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Zi Yan <ziy@nvidia.com>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Gregory Price <gourry@gourry.net>,
	Li Zhe <lizhe.67@bytedance.com>,
	Joshua Hahn <joshua.hahnjy@gmail.com>,
	akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
	rppt@kernel.org, mhocko@suse.com, corbet@lwn.net,
	skhan@linuxfoundation.org, ying.huang@linux.alibaba.com,
	apopple@nvidia.com, linux-mm@kvack.org,
	linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
Date: Thu, 01 Oct 2026 07:18:11 -0400	[thread overview]
Message-ID: <4CDEE1DB-1847-42EB-9801-6B3968590C33@nvidia.com> (raw)
In-Reply-To: <280965b2-c6eb-4f01-a0ac-c4fb7e518f97@kernel.org>

On 1 Oct 2026, at 6:54, David Hildenbrand (Arm) wrote:

> On 9/30/26 17:10, Zi Yan wrote:
>> On 30 Sep 2026, at 10:50, Gregory Price wrote:
>>
>>> On Wed, Sep 30, 2026 at 10:02:23PM +0800, Li Zhe wrote:
>>>> The use case is closer to the second one, but the main motivation is not
>>>> only startup-time performance.  The more important point is that each
>>>> workload has its own DDR/CXL budget assigned by the workload manager.
>>>>
>>>
>>> mempolicy is the wrong interface to do budget policy, you'd be better
>>> off looking at Joshua's memcg tiered node solutions for that.
>>>
>>>> MPOL_WEIGHTED_INTERLEAVE is useful because it lets new allocations
>>>> follow that per-workload budget from the beginning, instead of placing
>>>> everything on one tier first and correcting the placement later. This is
>>>> important for workload orchestration because the workload starts from a
>>>> placement close to its assigned DDR/CXL ratio.
>>>>
>>>
>>> This however is reasonable to me - you'd prefer to spread out the cost
>>> of initial faulting placement explicitly, rather than simply take
>>> fallbacks when the top-tier budget becomes pressured.
>>>
>>> i.e. w/o interleave:
>>>
>>> [node 0       ]
>>> ^^^^^^^^^^^^^^ alloc until full
>>>                 vvvvvvvvv fallback
>>>                 [node 1          ]
>>>
>>> in this scenario you end up with considerable hot memory
>>> on the remote node consolidated in time-space (everything
>>> allocated after node0 becomes full skews heavily toward
>>> node1)
>>>
>>> Tiering then likely takes many faults to rebalance after
>>> you've already reached node0 limits.
>>>
>>> w/ interleave
>>>
>>> [node 0                                      ]
>>>  ^^^vv^^^vv^^^vv^^^vv^^^vv^^^vv....
>>> [node 1                                      ]
>>>
>>> In this scenario you do an initial fill distributed by weight
>>> and then let tiering figure it out without consolidating all
>>> of the pressure to the point where node0 has no space left.
>>>
>>> That said - this seems like mostly an initial-fill problem, after
>>> you initially fill your memory, a new allocation largely implies
>>> the memory is hot - and you probably prefer that to be local.
>>
>> If this is a initial-fill problem, can userspace set weighted interleave
>> initially? The program or harness can observe memory usage of the program
>> or related NUMA nodes and switch the policy to numa balancing via
>> set_mempolicy() or mbind() without MPOL_MF_MOVE after certain threshold
>> is met?
>
> You mean: use the weighted policy initially and then switch to a NUMA-balancing
> one which doesn't involve the weights anymore?
>
> That makes more sense to me. Although I struggle to see why an effectively
> "let's put random memory on slow and others at hot" is a good starting point to
> later let if be fixed up by actual balancing/tiering.

Statistically speaking, unless the hot/cold page distribution follows a power-law
distribution, random placement usually provides an average placement result.
So it is not a bad start point.

>
> It all sounds a bit hackish. :)


Best Regards,
Yan, Zi

  reply	other threads:[~2026-10-01 11:18 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  7:26 Li Zhe
2026-09-30  8:43 ` Gregory Price
2026-09-30  8:59   ` Li Zhe
2026-09-30  9:11     ` Gregory Price
2026-09-30 10:59       ` Joshua Hahn
2026-09-30 11:22       ` Li Zhe
2026-09-30 12:25         ` Gregory Price
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52   ` Li Zhe
2026-09-30 12:11     ` Joshua Hahn
2026-09-30 14:02       ` Li Zhe
2026-09-30 14:50         ` Gregory Price
2026-09-30 15:10           ` Zi Yan
2026-10-01 10:54             ` David Hildenbrand (Arm)
2026-10-01 11:18               ` Zi Yan [this message]
2026-10-01 13:26               ` Gregory Price
2026-09-30 12:03   ` Gregory Price
2026-09-30 12:32     ` David Hildenbrand (Arm)
2026-09-30 13:40       ` Gregory Price
2026-10-01 10:55         ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=4CDEE1DB-1847-42EB-9801-6B3968590C33@nvidia.com \
    --to=ziy@nvidia.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=joshua.hahnjy@gmail.com \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=lizhe.67@bytedance.com \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=rppt@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=ying.huang@linux.alibaba.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®