mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Gregory Price <gourry@gourry.net>
To: Li Zhe <lizhe.67@bytedance.com>
Cc: "David Hildenbrand (Arm)" <david@kernel.org>,
	Zi Yan <ziy@nvidia.com>,  Joshua Hahn <joshua.hahnjy@gmail.com>,
	akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
	 rppt@kernel.org, mhocko@suse.com, corbet@lwn.net,
	skhan@linuxfoundation.org,  ying.huang@linux.alibaba.com,
	apopple@nvidia.com, linux-mm@kvack.org,
	 linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
Date: Fri, 2 Oct 2026 08:05:51 -0400	[thread overview]
Message-ID: <ar-b2nxDbQjcDbLY@gourry-fedora-PF4VCD3F> (raw)
In-Reply-To: <cac17315-7157-4f97-8ad4-98ef268f1b8a@bytedance.com>

On Fri, Oct 02, 2026 at 04:15:45PM +0800, Li Zhe wrote:
> On 10/2/26 3:56 PM, David Hildenbrand (Arm) wrote:
> > I can understand the "random initial placement will help if you cross your
> > fingers" argument from Zi.
> >
> > But then the question really is whether the app should then change the policy
> > after the initial placement was done and the weighted stuff no longer makes a
> > lot of sense.
> Yes, that model makes sense if the application or runtime can
> cooperate with the policy switch.
> 
> One limitation is that this is not fully transparent for existing
> workloads. set_mempolicy() updates the calling task's policy, and
> mbind() updates VMAs in the calling mm. move_pages() and
> migrate_pages() can move pages of another process, but they do not
> change that process's future allocation policy.
> 
> So I agree this staged approach is worth considering, but it also has
> some deployment cost for workloads that cannot participate in the
> policy switch.
>

if the use-case is vma (mbind), such a switch can make sense.

if the use-case is task policy (set_mempolicy), such a switch is
not a realistic solution.

In userspace we tend to use task policy by way of numactl:

  numactl --interleave=all ./my_program

This calls set_mempolicy for the numactl task and then exec's into
my_program with the inherited mempolicy.  That mempolicy is dup'd on
fork / clone.

Changing the mempolicy from that point requires every task in the
workload to call set_mempolicy() again.

There is no way to externally change another task's mempolicy.

I attempted this during the initial weighted-interleave exploration:

https://lore.kernel.org/all/20231122211200.31620-1-gregory.price@memverge.com/

but we did not see the need for it once we landed on sysfs controls for
weights.  In addition - there are MANY `current` assumptions hard coded
into the mempolicy and cgroup stack - getting such things dug out would
be (will be?) very painful.

~Gregory


  reply	other threads:[~2026-10-02 12:05 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  7:26 Li Zhe
2026-09-30  8:43 ` Gregory Price
2026-09-30  8:59   ` Li Zhe
2026-09-30  9:11     ` Gregory Price
2026-09-30 10:59       ` Joshua Hahn
2026-09-30 11:22       ` Li Zhe
2026-09-30 12:25         ` Gregory Price
2026-10-02  7:25           ` Li Zhe
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52   ` Li Zhe
2026-09-30 12:11     ` Joshua Hahn
2026-09-30 14:02       ` Li Zhe
2026-09-30 14:50         ` Gregory Price
2026-09-30 15:10           ` Zi Yan
2026-10-01 10:54             ` David Hildenbrand (Arm)
2026-10-01 11:18               ` Zi Yan
2026-10-01 13:26               ` Gregory Price
2026-10-02  7:56                 ` David Hildenbrand (Arm)
2026-10-02  8:15                   ` Li Zhe
2026-10-02 12:05                     ` Gregory Price [this message]
2026-09-30 12:03   ` Gregory Price
2026-09-30 12:32     ` David Hildenbrand (Arm)
2026-09-30 13:40       ` Gregory Price
2026-10-01 10:55         ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar-b2nxDbQjcDbLY@gourry-fedora-PF4VCD3F \
    --to=gourry@gourry.net \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=joshua.hahnjy@gmail.com \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=lizhe.67@bytedance.com \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=rppt@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®