mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Li Zhe" <lizhe.67@bytedance.com>
To: "David Hildenbrand (Arm)" <david@kernel.org>,
	 "Gregory Price" <gourry@gourry.net>
Cc: "Zi Yan" <ziy@nvidia.com>,
	"Joshua Hahn" <joshua.hahnjy@gmail.com>,
	 <akpm@linux-foundation.org>, <ljs@kernel.org>,
	<liam@infradead.org>,  <rppt@kernel.org>, <mhocko@suse.com>,
	<corbet@lwn.net>,  <skhan@linuxfoundation.org>,
	<ying.huang@linux.alibaba.com>,  <apopple@nvidia.com>,
	<linux-mm@kvack.org>, <linux-doc@vger.kernel.org>,
	 <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
Date: Fri, 2 Oct 2026 16:15:45 +0800	[thread overview]
Message-ID: <cac17315-7157-4f97-8ad4-98ef268f1b8a@bytedance.com> (raw)
In-Reply-To: <44d15437-6aae-4d7c-bfd4-296276583826@kernel.org>

On 10/2/26 3:56 PM, David Hildenbrand (Arm) wrote:
> On 10/1/26 15:26, Gregory Price wrote:
>> On Thu, Oct 01, 2026 at 12:54:22PM +0200, David Hildenbrand (Arm) wrote:
>>>> If this is a initial-fill problem, can userspace set weighted interleave
>>>> initially? The program or harness can observe memory usage of the program
>>>> or related NUMA nodes and switch the policy to numa balancing via
>>>> set_mempolicy() or mbind() without MPOL_MF_MOVE after certain threshold
>>>> is met?
>>> You mean: use the weighted policy initially and then switch to a NUMA-balancing
>>> one which doesn't involve the weights anymore?
>>>
>>> That makes more sense to me. Although I struggle to see why an effectively
>>> "let's put random memory on slow and others at hot" is a good starting point to
>>> later let if be fixed up by actual balancing/tiering.
>>>
>>> It all sounds a bit hackish. :)
>>>
>> It is a bit of a non-combo (Nonbo).  You're using weighted interleave
>> with the intent of spreading out the bandwidth utilization (and maybe to
>> offset some reclaim behavior? *shrug*) but then undo all the placement
>> with tiering.
> I can understand the "random initial placement will help if you cross your
> fingers" argument from Zi.
>
> But then the question really is whether the app should then change the policy
> after the initial placement was done and the weighted stuff no longer makes a
> lot of sense.
Yes, that model makes sense if the application or runtime can
cooperate with the policy switch.

One limitation is that this is not fully transparent for existing
workloads. set_mempolicy() updates the calling task's policy, and
mbind() updates VMAs in the calling mm. move_pages() and
migrate_pages() can move pages of another process, but they do not
change that process's future allocation policy.

So I agree this staged approach is worth considering, but it also has
some deployment cost for workloads that cannot participate in the
policy switch.

Thanks,
Zhe

>
>> But, in defense of the hackery - I have seen strategies like this work
>> to optimize startup times and then let tiering optimize runtime.  Phased
>> execution gets funky like that.
> No doubt about that.

  reply	other threads:[~2026-10-02  8:16 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  7:26 Li Zhe
2026-09-30  8:43 ` Gregory Price
2026-09-30  8:59   ` Li Zhe
2026-09-30  9:11     ` Gregory Price
2026-09-30 10:59       ` Joshua Hahn
2026-09-30 11:22       ` Li Zhe
2026-09-30 12:25         ` Gregory Price
2026-10-02  7:25           ` Li Zhe
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52   ` Li Zhe
2026-09-30 12:11     ` Joshua Hahn
2026-09-30 14:02       ` Li Zhe
2026-09-30 14:50         ` Gregory Price
2026-09-30 15:10           ` Zi Yan
2026-10-01 10:54             ` David Hildenbrand (Arm)
2026-10-01 11:18               ` Zi Yan
2026-10-01 13:26               ` Gregory Price
2026-10-02  7:56                 ` David Hildenbrand (Arm)
2026-10-02  8:15                   ` Li Zhe [this message]
2026-10-02 12:05                     ` Gregory Price
2026-09-30 12:03   ` Gregory Price
2026-09-30 12:32     ` David Hildenbrand (Arm)
2026-09-30 13:40       ` Gregory Price
2026-10-01 10:55         ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=cac17315-7157-4f97-8ad4-98ef268f1b8a@bytedance.com \
    --to=lizhe.67@bytedance.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=joshua.hahnjy@gmail.com \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=rppt@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®