mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
@ 2026-09-30  7:26 Li Zhe
  2026-09-30  8:43 ` Gregory Price
  2026-09-30 11:26 ` David Hildenbrand (Arm)
  0 siblings, 2 replies; 14+ messages in thread
From: Li Zhe @ 2026-09-30  7:26 UTC (permalink / raw)
  To: akpm, david, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, gourry, ying.huang, apopple
  Cc: linux-mm, linux-doc, linux-kernel, lizhe.67

MPOL_WEIGHTED_INTERLEAVE is useful on tiered-memory systems because it
can seed a workload's new allocations across fast memory and slower
capacity memory according to a configured ratio.

That initial placement is useful for workload managers and orchestration
systems.  They can take the amount of fast memory and slower capacity
memory on a machine into account before starting a workload, and choose a
weighted policy that seeds the workload across the tiers at allocation
time.  This avoids starting from an all-fast or all-slow placement and
then relying on promotion or demotion to reshape a large working set.

After those pages have been placed, however, the policy cannot currently
opt in to migrate-on-fault placement.  set_mempolicy() and mbind()
reject MPOL_WEIGHTED_INTERLEAVE when MPOL_F_NUMA_BALANCING is specified,
so memory tiering cannot promote hot pages that were initially placed on
the slower nodes by the weighted policy.

Initial placement is only a starting point.  Pages initially allocated on
fast memory are not necessarily the long-term hot pages, and pages
initially allocated on slower memory may become hot as the workload's hot
set changes.  The policy therefore needs to be able to combine weighted
initial placement with memory tiering's NUMA fault based hot-page
promotion.

Allow MPOL_F_NUMA_BALANCING for MPOL_WEIGHTED_INTERLEAVE.  As with
MPOL_BIND and MPOL_PREFERRED_MANY, keep migration constrained by the
policy nodemask: if the CPU's node is outside the nodemask, do not
migrate the folio there.

This is an opt-in behavior.  MPOL_WEIGHTED_INTERLEAVE without
MPOL_F_NUMA_BALANCING keeps its existing behavior and does not
participate in NUMA balancing.

The weighted interleave weights continue to control new allocations only.
They are not a target residency ratio after migrate-on-fault placement.

Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
 Documentation/admin-guide/mm/numa_memory_policy.rst | 10 ++++++++++
 mm/mempolicy.c                                      |  8 +++++++-
 2 files changed, 17 insertions(+), 1 deletion(-)

diff --git a/Documentation/admin-guide/mm/numa_memory_policy.rst b/Documentation/admin-guide/mm/numa_memory_policy.rst
index 90ab26e805a9a..f10a8297fb178 100644
--- a/Documentation/admin-guide/mm/numa_memory_policy.rst
+++ b/Documentation/admin-guide/mm/numa_memory_policy.rst
@@ -259,6 +259,16 @@ MPOL_WEIGHTED_INTERLEAVE
 	weight.  For example if nodes [0,1] are weighted [5,2], 5 pages
 	will be allocated on node0 for every 2 pages allocated on node1.
 
+	When MPOL_F_NUMA_BALANCING is specified, migrate-on-fault
+	placement is allowed within the policy nodemask.  This is an
+	opt-in behavior; without MPOL_F_NUMA_BALANCING,
+	MPOL_WEIGHTED_INTERLEAVE keeps its existing behavior and does
+	not participate in NUMA balancing.  The flag can be used by
+	memory tiering to promote hot pages that were initially
+	allocated on slower nodes.  The weights still affect new
+	allocations only and do not define a target resident ratio after
+	page migration.
+
 NUMA memory policy supports the following optional mode flags:
 
 MPOL_F_STATIC_NODES
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index 79053ece02cd4..f71cdb488bbdc 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -1731,7 +1731,8 @@ static inline int sanitize_mpol_flags(int *mode, unsigned short *flags)
 	if ((*flags & MPOL_F_STATIC_NODES) && (*flags & MPOL_F_RELATIVE_NODES))
 		return -EINVAL;
 	if (*flags & MPOL_F_NUMA_BALANCING) {
-		if (*mode == MPOL_BIND || *mode == MPOL_PREFERRED_MANY)
+		if (*mode == MPOL_BIND || *mode == MPOL_PREFERRED_MANY ||
+		    *mode == MPOL_WEIGHTED_INTERLEAVE)
 			*flags |= (MPOL_F_MOF | MPOL_F_MORON);
 		else
 			return -EINVAL;
@@ -3005,6 +3006,11 @@ int mpol_misplaced(struct folio *folio, struct vm_fault *vmf,
 		break;
 
 	case MPOL_WEIGHTED_INTERLEAVE:
+		if (pol->flags & MPOL_F_MORON) {
+			if (node_isset(thisnid, pol->nodes))
+				break;
+			goto out;
+		}
 		polnid = weighted_interleave_nid(pol, ilx);
 		break;
 
-- 
2.20.1

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30  7:26 [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy Li Zhe
@ 2026-09-30  8:43 ` Gregory Price
  2026-09-30  8:59   ` Li Zhe
  2026-09-30 11:26 ` David Hildenbrand (Arm)
  1 sibling, 1 reply; 14+ messages in thread
From: Gregory Price @ 2026-09-30  8:43 UTC (permalink / raw)
  To: Li Zhe
  Cc: akpm, david, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On Wed, Sep 30, 2026 at 03:26:17PM +0800, Li Zhe wrote:
> Allow MPOL_F_NUMA_BALANCING for MPOL_WEIGHTED_INTERLEAVE.  As with
> MPOL_BIND and MPOL_PREFERRED_MANY, keep migration constrained by the
> policy nodemask: if the CPU's node is outside the nodemask, do not
> migrate the folio there.
>

Why for WEIGHTED_INTERLEAVE and not INTERLEAVE as well?

Otherwise this seems reasonable.

~Gregory

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30  8:43 ` Gregory Price
@ 2026-09-30  8:59   ` Li Zhe
  2026-09-30  9:11     ` Gregory Price
  0 siblings, 1 reply; 14+ messages in thread
From: Li Zhe @ 2026-09-30  8:59 UTC (permalink / raw)
  To: Gregory Price
  Cc: akpm, david, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On 9/30/26 4:43 PM, Gregory Price wrote:
> On Wed, Sep 30, 2026 at 03:26:17PM +0800, Li Zhe wrote:
>> Allow MPOL_F_NUMA_BALANCING for MPOL_WEIGHTED_INTERLEAVE.  As with
>> MPOL_BIND and MPOL_PREFERRED_MANY, keep migration constrained by the
>> policy nodemask: if the CPU's node is outside the nodemask, do not
>> migrate the folio there.
>>
> Why for WEIGHTED_INTERLEAVE and not INTERLEAVE as well?


Thanks for pointing this out.

I focused on MPOL_WEIGHTED_INTERLEAVE in v1 because the motivating use
case is to seed memory across tiers with a configurable ratio at
allocation time, and then let NUMA balancing/memory tiering adjust the
placement based on access patterns. With equal weights,
MPOL_WEIGHTED_INTERLEAVE can also cover the regular interleave
allocation pattern.

That said, I agree that MPOL_INTERLEAVE can be handled consistently as
well. Since this is still an explicit opt-in via MPOL_F_NUMA_BALANCING,
unless others see a reason to keep MPOL_INTERLEAVE out, I will extend
this in v2 to cover both MPOL_INTERLEAVE and MPOL_WEIGHTED_INTERLEAVE
with the same nodemask constraint.

Thanks,
Zhe

>
> Otherwise this seems reasonable.
>
> ~Gregory

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30  8:59   ` Li Zhe
@ 2026-09-30  9:11     ` Gregory Price
  2026-09-30 10:59       ` Joshua Hahn
  2026-09-30 11:22       ` Li Zhe
  0 siblings, 2 replies; 14+ messages in thread
From: Gregory Price @ 2026-09-30  9:11 UTC (permalink / raw)
  To: Li Zhe
  Cc: akpm, david, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On Wed, Sep 30, 2026 at 04:59:50PM +0800, Li Zhe wrote:
> On 9/30/26 4:43 PM, Gregory Price wrote:
> > On Wed, Sep 30, 2026 at 03:26:17PM +0800, Li Zhe wrote:
> >> Allow MPOL_F_NUMA_BALANCING for MPOL_WEIGHTED_INTERLEAVE.  As with
> >> MPOL_BIND and MPOL_PREFERRED_MANY, keep migration constrained by the
> >> policy nodemask: if the CPU's node is outside the nodemask, do not
> >> migrate the folio there.
> >>
> > Why for WEIGHTED_INTERLEAVE and not INTERLEAVE as well?
> 
> 
> Thanks for pointing this out.
> 
> I focused on MPOL_WEIGHTED_INTERLEAVE in v1 because the motivating use
> case is to seed memory across tiers with a configurable ratio at
> allocation time, and then let NUMA balancing/memory tiering adjust the
> placement based on access patterns. With equal weights,
> MPOL_WEIGHTED_INTERLEAVE can also cover the regular interleave
> allocation pattern.
> 
> That said, I agree that MPOL_INTERLEAVE can be handled consistently as
> well. Since this is still an explicit opt-in via MPOL_F_NUMA_BALANCING,
> unless others see a reason to keep MPOL_INTERLEAVE out, I will extend
> this in v2 to cover both MPOL_INTERLEAVE and MPOL_WEIGHTED_INTERLEAVE
> with the same nodemask constraint.
> 

Looking at it, is there an actual reason to limit F_MORON at all? Or
should we just lift

                if (pol->flags & MPOL_F_MORON) {
                        /*
                         * Optimize placement among multiple nodes
                         * via NUMA balancing
                         */
                        if (node_isset(thisnid, pol->nodes))
                                break;
                        goto out;
                }

Out ahead of everything?

~Gregory

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30  9:11     ` Gregory Price
@ 2026-09-30 10:59       ` Joshua Hahn
  2026-09-30 11:22       ` Li Zhe
  1 sibling, 0 replies; 14+ messages in thread
From: Joshua Hahn @ 2026-09-30 10:59 UTC (permalink / raw)
  To: Gregory Price
  Cc: Li Zhe, akpm, david, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On Wed, 30 Sep 2026 05:11:28 -0400 Gregory Price <gourry@gourry.net> wrote:

> On Wed, Sep 30, 2026 at 04:59:50PM +0800, Li Zhe wrote:
> > On 9/30/26 4:43 PM, Gregory Price wrote:
> > > On Wed, Sep 30, 2026 at 03:26:17PM +0800, Li Zhe wrote:
> > >> Allow MPOL_F_NUMA_BALANCING for MPOL_WEIGHTED_INTERLEAVE.  As with
> > >> MPOL_BIND and MPOL_PREFERRED_MANY, keep migration constrained by the
> > >> policy nodemask: if the CPU's node is outside the nodemask, do not
> > >> migrate the folio there.
> > >>
> > > Why for WEIGHTED_INTERLEAVE and not INTERLEAVE as well?
> > 
> > 
> > Thanks for pointing this out.
> > 
> > I focused on MPOL_WEIGHTED_INTERLEAVE in v1 because the motivating use
> > case is to seed memory across tiers with a configurable ratio at
> > allocation time, and then let NUMA balancing/memory tiering adjust the
> > placement based on access patterns. With equal weights,
> > MPOL_WEIGHTED_INTERLEAVE can also cover the regular interleave
> > allocation pattern.
> > 
> > That said, I agree that MPOL_INTERLEAVE can be handled consistently as
> > well. Since this is still an explicit opt-in via MPOL_F_NUMA_BALANCING,
> > unless others see a reason to keep MPOL_INTERLEAVE out, I will extend
> > this in v2 to cover both MPOL_INTERLEAVE and MPOL_WEIGHTED_INTERLEAVE
> > with the same nodemask constraint.
> > 
> 
> Looking at it, is there an actual reason to limit F_MORON at all? Or
> should we just lift
> 
>                 if (pol->flags & MPOL_F_MORON) {
>                         /*
>                          * Optimize placement among multiple nodes
>                          * via NUMA balancing
>                          */
>                         if (node_isset(thisnid, pol->nodes))
>                                 break;
>                         goto out;
>                 }
> 
> Out ahead of everything?

+1, I like this idea, especialy since the switch-case break case already
leads to an MPOL_F_MORON check at the bottom. It would be nice to just
consolidate this into one generic path at the top and simplify the
exit path as well.

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30  9:11     ` Gregory Price
  2026-09-30 10:59       ` Joshua Hahn
@ 2026-09-30 11:22       ` Li Zhe
  2026-09-30 12:25         ` Gregory Price
  1 sibling, 1 reply; 14+ messages in thread
From: Li Zhe @ 2026-09-30 11:22 UTC (permalink / raw)
  To: Gregory Price
  Cc: akpm, david, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On 9/30/26 5:11 PM, Gregory Price wrote:
> On Wed, Sep 30, 2026 at 04:59:50PM +0800, Li Zhe wrote:
>> On 9/30/26 4:43 PM, Gregory Price wrote:
>>> On Wed, Sep 30, 2026 at 03:26:17PM +0800, Li Zhe wrote:
>>>> Allow MPOL_F_NUMA_BALANCING for MPOL_WEIGHTED_INTERLEAVE.  As with
>>>> MPOL_BIND and MPOL_PREFERRED_MANY, keep migration constrained by the
>>>> policy nodemask: if the CPU's node is outside the nodemask, do not
>>>> migrate the folio there.
>>>>
>>> Why for WEIGHTED_INTERLEAVE and not INTERLEAVE as well?
>>
>> Thanks for pointing this out.
>>
>> I focused on MPOL_WEIGHTED_INTERLEAVE in v1 because the motivating use
>> case is to seed memory across tiers with a configurable ratio at
>> allocation time, and then let NUMA balancing/memory tiering adjust the
>> placement based on access patterns. With equal weights,
>> MPOL_WEIGHTED_INTERLEAVE can also cover the regular interleave
>> allocation pattern.
>>
>> That said, I agree that MPOL_INTERLEAVE can be handled consistently as
>> well. Since this is still an explicit opt-in via MPOL_F_NUMA_BALANCING,
>> unless others see a reason to keep MPOL_INTERLEAVE out, I will extend
>> this in v2 to cover both MPOL_INTERLEAVE and MPOL_WEIGHTED_INTERLEAVE
>> with the same nodemask constraint.
>>
> Looking at it, is there an actual reason to limit F_MORON at all? Or
> should we just lift
>
>                  if (pol->flags & MPOL_F_MORON) {
>                          /*
>                           * Optimize placement among multiple nodes
>                           * via NUMA balancing
>                           */
>                          if (node_isset(thisnid, pol->nodes))
>                                  break;
>                          goto out;
>                  }
>
> Out ahead of everything?
Yes, that makes sense to me.

I do not see a reason to keep the MPOL_F_MORON handling limited to
specific policy cases. The flag already means that migrate-on-fault
placement should target the accessing CPU's node, and the nodemask check
is the common constraint we want for all policies that opt in to this
behavior.

I will rework v2 to handle MPOL_F_MORON before the policy-specific
switch. Then the switch can remain responsible for the non-MPOL_F_MORON
misplaced logic, while MPOL_BIND, MPOL_PREFERRED_MANY, MPOL_INTERLEAVE
and MPOL_WEIGHTED_INTERLEAVE can share the same migrate-on-fault path.

Thanks,
Zhe

>
> ~Gregory

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30  7:26 [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy Li Zhe
  2026-09-30  8:43 ` Gregory Price
@ 2026-09-30 11:26 ` David Hildenbrand (Arm)
  2026-09-30 11:52   ` Li Zhe
  2026-09-30 12:03   ` Gregory Price
  1 sibling, 2 replies; 14+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-30 11:26 UTC (permalink / raw)
  To: Li Zhe, akpm, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, gourry, ying.huang, apopple
  Cc: linux-mm, linux-doc, linux-kernel

On 9/30/26 09:26, Li Zhe wrote:
> MPOL_WEIGHTED_INTERLEAVE is useful on tiered-memory systems because it
> can seed a workload's new allocations across fast memory and slower
> capacity memory according to a configured ratio.
> 
> That initial placement is useful for workload managers and orchestration
> systems.  They can take the amount of fast memory and slower capacity
> memory on a machine into account before starting a workload, and choose a
> weighted policy that seeds the workload across the tiers at allocation
> time.  This avoids starting from an all-fast or all-slow placement and
> then relying on promotion or demotion to reshape a large working set.
> 
> After those pages have been placed, however, the policy cannot currently
> opt in to migrate-on-fault placement.  set_mempolicy() and mbind()
> reject MPOL_WEIGHTED_INTERLEAVE when MPOL_F_NUMA_BALANCING is specified,
> so memory tiering cannot promote hot pages that were initially placed on
> the slower nodes by the weighted policy.
> 
> Initial placement is only a starting point.  Pages initially allocated on
> fast memory are not necessarily the long-term hot pages, and pages
> initially allocated on slower memory may become hot as the workload's hot
> set changes.  The policy therefore needs to be able to combine weighted
> initial placement with memory tiering's NUMA fault based hot-page
> promotion.

Ok, so memory allocation will respect the weights but balancing will ignore
them? That really sounds rather odd to me.

And I assume that was the reason why we might have disallowed the combination:
it turns a weighted mechanism into an unweighted mechanism.

So are we really sure these semantics that you would essentially set in stone
here are the semantics we want? (ignoring weights)

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30 11:26 ` David Hildenbrand (Arm)
@ 2026-09-30 11:52   ` Li Zhe
  2026-09-30 12:11     ` Joshua Hahn
  2026-09-30 12:03   ` Gregory Price
  1 sibling, 1 reply; 14+ messages in thread
From: Li Zhe @ 2026-09-30 11:52 UTC (permalink / raw)
  To: David Hildenbrand (Arm),
	akpm, ljs, liam, rppt, mhocko, corbet, skhan, ziy, joshua.hahnjy,
	gourry, ying.huang, apopple
  Cc: linux-mm, linux-doc, linux-kernel

On 9/30/26 7:26 PM, David Hildenbrand (Arm) wrote:
> On 9/30/26 09:26, Li Zhe wrote:
>> MPOL_WEIGHTED_INTERLEAVE is useful on tiered-memory systems because it
>> can seed a workload's new allocations across fast memory and slower
>> capacity memory according to a configured ratio.
>>
>> That initial placement is useful for workload managers and orchestration
>> systems.  They can take the amount of fast memory and slower capacity
>> memory on a machine into account before starting a workload, and choose a
>> weighted policy that seeds the workload across the tiers at allocation
>> time.  This avoids starting from an all-fast or all-slow placement and
>> then relying on promotion or demotion to reshape a large working set.
>>
>> After those pages have been placed, however, the policy cannot currently
>> opt in to migrate-on-fault placement.  set_mempolicy() and mbind()
>> reject MPOL_WEIGHTED_INTERLEAVE when MPOL_F_NUMA_BALANCING is specified,
>> so memory tiering cannot promote hot pages that were initially placed on
>> the slower nodes by the weighted policy.
>>
>> Initial placement is only a starting point.  Pages initially allocated on
>> fast memory are not necessarily the long-term hot pages, and pages
>> initially allocated on slower memory may become hot as the workload's hot
>> set changes.  The policy therefore needs to be able to combine weighted
>> initial placement with memory tiering's NUMA fault based hot-page
>> promotion.
> Ok, so memory allocation will respect the weights but balancing will ignore
> them? That really sounds rather odd to me.
>
> And I assume that was the reason why we might have disallowed the combination:
> it turns a weighted mechanism into an unweighted mechanism.
>
> So are we really sure these semantics that you would essentially set in stone
> here are the semantics we want? (ignoring weights)
Yes, that is a fair concern. It does look odd if the weights are
interpreted as a hard resident placement ratio.

My understanding of the existing MPOL_WEIGHTED_INTERLEAVE ABI is that
the weights are allocation weights, not a long-term resident ratio. The
sysfs ABI documentation says that these weights only affect new
allocations, and that changing them at runtime will not migrate already
allocated pages.  The implementation also has normal allocation fallback
if the selected weighted target cannot satisfy the allocation.

The use case here follows that interpretation.  The weights are used to
seed the initial placement across memory tiers.  After that, with an
explicit MPOL_F_NUMA_BALANCING opt-in, memory tiering would promote hot
pages based on access patterns, not based on the original allocation
ratio.  Users that want weighted allocation without that behavior would
keep using MPOL_WEIGHTED_INTERLEAVE without MPOL_F_NUMA_BALANCING.

That said, I agree that allowing this flag combination would set the
semantics for it.  If reusing MPOL_F_NUMA_BALANCING for this is too
ambiguous, do you think we should model this as a separate opt-in ABI
for "weighted initial placement plus access-based balancing" instead?
For example, a separate flag or policy mode would make it clearer that
the weights are not intended to constrain migrate-on-fault placement.

If the allocation-only interpretation of the weights is acceptable, I
can make it explicit in the commit message and documentation in v2.
Otherwise I would appreciate your suggestion on the preferred interface.

Thanks,
Zhe

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30 11:26 ` David Hildenbrand (Arm)
  2026-09-30 11:52   ` Li Zhe
@ 2026-09-30 12:03   ` Gregory Price
  2026-09-30 12:32     ` David Hildenbrand (Arm)
  1 sibling, 1 reply; 14+ messages in thread
From: Gregory Price @ 2026-09-30 12:03 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Li Zhe, akpm, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On Wed, Sep 30, 2026 at 01:26:17PM +0200, David Hildenbrand (Arm) wrote:
> > 
> > Initial placement is only a starting point.  Pages initially allocated on
> > fast memory are not necessarily the long-term hot pages, and pages
> > initially allocated on slower memory may become hot as the workload's hot
> > set changes.  The policy therefore needs to be able to combine weighted
> > initial placement with memory tiering's NUMA fault based hot-page
> > promotion.
> 
> Ok, so memory allocation will respect the weights but balancing will ignore
> them? That really sounds rather odd to me.
> 
> And I assume that was the reason why we might have disallowed the combination:
> it turns a weighted mechanism into an unweighted mechanism.
> 
> So are we really sure these semantics that you would essentially set in stone
> here are the semantics we want? (ignoring weights)
>

Right, but you're requesting this configuration.  If you didn't set
F_NUMA_BALANCING then you get the existing semantics.

In the existing semantics, if we're INTERLEAVE and WEIGHTED_INTERLEAVE
we just always say "no tiering for you".

But as an opt-in option?  I don't quite see the argument for saying
interleaved regions (or tasks) to be opted-out if the user asks for it -
that just seems like an arbitrary limitation.

~Gregory

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30 11:52   ` Li Zhe
@ 2026-09-30 12:11     ` Joshua Hahn
  2026-09-30 14:02       ` Li Zhe
  0 siblings, 1 reply; 14+ messages in thread
From: Joshua Hahn @ 2026-09-30 12:11 UTC (permalink / raw)
  To: Li Zhe
  Cc: David Hildenbrand (Arm),
	akpm, ljs, liam, rppt, mhocko, corbet, skhan, ziy, joshua.hahnjy,
	gourry, ying.huang, apopple, linux-mm, linux-doc, linux-kernel

On Wed, 30 Sep 2026 19:52:05 +0800 "Li Zhe" <lizhe.67@bytedance.com> wrote:

> On 9/30/26 7:26 PM, David Hildenbrand (Arm) wrote:
> > On 9/30/26 09:26, Li Zhe wrote:
> >> MPOL_WEIGHTED_INTERLEAVE is useful on tiered-memory systems because it
> >> can seed a workload's new allocations across fast memory and slower
> >> capacity memory according to a configured ratio.
> >>
> >> That initial placement is useful for workload managers and orchestration
> >> systems.  They can take the amount of fast memory and slower capacity
> >> memory on a machine into account before starting a workload, and choose a
> >> weighted policy that seeds the workload across the tiers at allocation
> >> time.  This avoids starting from an all-fast or all-slow placement and
> >> then relying on promotion or demotion to reshape a large working set.
> >>
> >> After those pages have been placed, however, the policy cannot currently
> >> opt in to migrate-on-fault placement.  set_mempolicy() and mbind()
> >> reject MPOL_WEIGHTED_INTERLEAVE when MPOL_F_NUMA_BALANCING is specified,
> >> so memory tiering cannot promote hot pages that were initially placed on
> >> the slower nodes by the weighted policy.
> >>
> >> Initial placement is only a starting point.  Pages initially allocated on
> >> fast memory are not necessarily the long-term hot pages, and pages
> >> initially allocated on slower memory may become hot as the workload's hot
> >> set changes.  The policy therefore needs to be able to combine weighted
> >> initial placement with memory tiering's NUMA fault based hot-page
> >> promotion.
> > Ok, so memory allocation will respect the weights but balancing will ignore
> > them? That really sounds rather odd to me.
> >
> > And I assume that was the reason why we might have disallowed the combination:
> > it turns a weighted mechanism into an unweighted mechanism.
> >
> > So are we really sure these semantics that you would essentially set in stone
> > here are the semantics we want? (ignoring weights)
> Yes, that is a fair concern. It does look odd if the weights are
> interpreted as a hard resident placement ratio.
> 
> My understanding of the existing MPOL_WEIGHTED_INTERLEAVE ABI is that
> the weights are allocation weights, not a long-term resident ratio. The
> sysfs ABI documentation says that these weights only affect new
> allocations, and that changing them at runtime will not migrate already
> allocated pages.  The implementation also has normal allocation fallback
> if the selected weighted target cannot satisfy the allocation.
> 
> The use case here follows that interpretation.  The weights are used to
> seed the initial placement across memory tiers.  After that, with an
> explicit MPOL_F_NUMA_BALANCING opt-in, memory tiering would promote hot
> pages based on access patterns, not based on the original allocation
> ratio.  Users that want weighted allocation without that behavior would
> keep using MPOL_WEIGHTED_INTERLEAVE without MPOL_F_NUMA_BALANCING.
> 
> That said, I agree that allowing this flag combination would set the
> semantics for it.  If reusing MPOL_F_NUMA_BALANCING for this is too
> ambiguous, do you think we should model this as a separate opt-in ABI
> for "weighted initial placement plus access-based balancing" instead?
> For example, a separate flag or policy mode would make it clearer that
> the weights are not intended to constrain migrate-on-fault placement.
> 
> If the allocation-only interpretation of the weights is acceptable, I
> can make it explicit in the commit message and documentation in v2.
> Otherwise I would appreciate your suggestion on the preferred interface.

I see two usecases for weighted interleave. Let's say you first allocate
all the cold memory, and then all the hot memory using an interleave
policy. This leaves the same hotness in both tiers since they are
interleaved:

In one scenario you might genuinely want to keep some hot memory in
both tiers to maximize bandwidth utilization.

But in the other case when you are not limited by bandwidth, you might
indeed want to start out with a weighted interleave allocation to
reduce variance on where hot memory lands, and then make promotion /
demotion decisions based on access.

Zhe, do you have a specific use-case in mind? For the second use
case I would be curious to see if you see any meaningful performance
differences in startup time when comapring a workload using other
mempolicies for the allocation and then tiering vs. using weighted
interleave and then tiering. 

But yeah, I think David is right that we should go over the semantics
and really make sure this is what we want to commit to.

Thanks again Zhe! Have a great day,
Joshua

> Thanks,
> Zhe

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30 11:22       ` Li Zhe
@ 2026-09-30 12:25         ` Gregory Price
  0 siblings, 0 replies; 14+ messages in thread
From: Gregory Price @ 2026-09-30 12:25 UTC (permalink / raw)
  To: Li Zhe
  Cc: akpm, david, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On Wed, Sep 30, 2026 at 07:22:10PM +0800, Li Zhe wrote:
> Yes, that makes sense to me.
> 
> I do not see a reason to keep the MPOL_F_MORON handling limited to
> specific policy cases. The flag already means that migrate-on-fault
> placement should target the accessing CPU's node, and the nodemask check
> is the common constraint we want for all policies that opt in to this
> behavior.
> 
> I will rework v2 to handle MPOL_F_MORON before the policy-specific
> switch. Then the switch can remain responsible for the non-MPOL_F_MORON
> misplaced logic, while MPOL_BIND, MPOL_PREFERRED_MANY, MPOL_INTERLEAVE
> and MPOL_WEIGHTED_INTERLEAVE can share the same migrate-on-fault path.
>

I do think it's worth breaking this up a bit, we can decide whether we
want to apply it to interleave separate of the larger change.

there is also slightly different behavior for task-mempolicy and
vma-mempolicy, because vma-mempolicy applies interleave via index while
task mempolicy does it based on a rolling counter, so you'll want to
think about what happens on repeated faults

i.e.

1) fault in a VMA interleaved based on idx
2) a bunch of tiering and pageout/swap happens
3) we fault a page back in based on idx - that means this page is
   forever-faulted onto that location rather than taking that as an
   indication that maybe it should be local

Maybe we're ok with that, but we should probably think about it a bit.

We may also want to limit this based on NUMA balancing being in tiering
mode vs normal mode.  This only makes sense in tiering mode, in my
opinion.  In normal mode i'm not sure it makes as much sense - and
that's probably where this all came from.

tl;dr: If we want to change this behavior for interleave, we should
probably give more thought for how it should apply more generally
instead of just hacking on support to one or two modes.

~Gregory

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30 12:03   ` Gregory Price
@ 2026-09-30 12:32     ` David Hildenbrand (Arm)
  2026-09-30 13:40       ` Gregory Price
  0 siblings, 1 reply; 14+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-30 12:32 UTC (permalink / raw)
  To: Gregory Price
  Cc: Li Zhe, akpm, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On 9/30/26 14:03, Gregory Price wrote:
> On Wed, Sep 30, 2026 at 01:26:17PM +0200, David Hildenbrand (Arm) wrote:
>>>
>>> Initial placement is only a starting point.  Pages initially allocated on
>>> fast memory are not necessarily the long-term hot pages, and pages
>>> initially allocated on slower memory may become hot as the workload's hot
>>> set changes.  The policy therefore needs to be able to combine weighted
>>> initial placement with memory tiering's NUMA fault based hot-page
>>> promotion.
>>
>> Ok, so memory allocation will respect the weights but balancing will ignore
>> them? That really sounds rather odd to me.
>>
>> And I assume that was the reason why we might have disallowed the combination:
>> it turns a weighted mechanism into an unweighted mechanism.
>>
>> So are we really sure these semantics that you would essentially set in stone
>> here are the semantics we want? (ignoring weights)
>>
> 
> Right, but you're requesting this configuration.  If you didn't set
> F_NUMA_BALANCING then you get the existing semantics.
> 
> In the existing semantics, if we're INTERLEAVE and WEIGHTED_INTERLEAVE
> we just always say "no tiering for you".
> 
> But as an opt-in option?  I don't quite see the argument for saying
> interleaved regions (or tasks) to be opted-out if the user asks for it -
> that just seems like an arbitrary limitation.

Just to be clear, what I am saying is: the proposed semantics are inconsistent
(weighted when mechanism A honors them and mechanism B doesn't honor them), but
once we set these semantics in stone like that, we cannot easily change them
later because some user might depend on that behavior.

So you'll need yet another flag to say "NUMA balancing really also honors the
weights". And that's where it all gets ugly.

-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30 12:32     ` David Hildenbrand (Arm)
@ 2026-09-30 13:40       ` Gregory Price
  0 siblings, 0 replies; 14+ messages in thread
From: Gregory Price @ 2026-09-30 13:40 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Li Zhe, akpm, ljs, liam, rppt, mhocko, corbet, skhan, ziy,
	joshua.hahnjy, ying.huang, apopple, linux-mm, linux-doc,
	linux-kernel

On Wed, Sep 30, 2026 at 02:32:16PM +0200, David Hildenbrand (Arm) wrote:
> > 
> > Right, but you're requesting this configuration.  If you didn't set
> > F_NUMA_BALANCING then you get the existing semantics.
> > 
> > In the existing semantics, if we're INTERLEAVE and WEIGHTED_INTERLEAVE
> > we just always say "no tiering for you".
> > 
> > But as an opt-in option?  I don't quite see the argument for saying
> > interleaved regions (or tasks) to be opted-out if the user asks for it -
> > that just seems like an arbitrary limitation.
> 
> Just to be clear, what I am saying is: the proposed semantics are inconsistent
> (weighted when mechanism A honors them and mechanism B doesn't honor them), but
> once we set these semantics in stone like that, we cannot easily change them
> later because some user might depend on that behavior.
> 
> So you'll need yet another flag to say "NUMA balancing really also honors the
> weights". And that's where it all gets ugly.
> 

I would agree with you if demotion didn't entirely ignoring mempolicy.

On a tiered system, mempolicy *only* applies to initial (or fault-in)
placement, and otherwise is more or less entirely ignored.  This is
a case where it's not ignored - which is actually inconsistent with
the rest of the tiering tools (demotion, damon).

This is why I said this distinction may only make sense in TIERING
mode, while in NORMAL mode you likely just want to fault it back to
its original interleave location.

Also file interleave (indexed by offset) vs task (counter-incremented)
policies are affected by these changes very differently, which I asked
for some thought on.

As tiering develops, the less I think task-mempolicy as a whole makes
much sense (because it's at-best advisory).

~Gregory

^ permalink raw reply	[flat|nested] 14+ messages in thread

* Re: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
  2026-09-30 12:11     ` Joshua Hahn
@ 2026-09-30 14:02       ` Li Zhe
  0 siblings, 0 replies; 14+ messages in thread
From: Li Zhe @ 2026-09-30 14:02 UTC (permalink / raw)
  To: Joshua Hahn
  Cc: David Hildenbrand (Arm),
	akpm, ljs, liam, rppt, mhocko, corbet, skhan, ziy, gourry,
	ying.huang, apopple, linux-mm, linux-doc, linux-kernel

On 9/30/26 8:11 PM, Joshua Hahn wrote:
> On Wed, 30 Sep 2026 19:52:05 +0800 "Li Zhe" <lizhe.67@bytedance.com> wrote:
>
>> On 9/30/26 7:26 PM, David Hildenbrand (Arm) wrote:
>>> On 9/30/26 09:26, Li Zhe wrote:
>>>> MPOL_WEIGHTED_INTERLEAVE is useful on tiered-memory systems because it
>>>> can seed a workload's new allocations across fast memory and slower
>>>> capacity memory according to a configured ratio.
>>>>
>>>> That initial placement is useful for workload managers and orchestration
>>>> systems.  They can take the amount of fast memory and slower capacity
>>>> memory on a machine into account before starting a workload, and choose a
>>>> weighted policy that seeds the workload across the tiers at allocation
>>>> time.  This avoids starting from an all-fast or all-slow placement and
>>>> then relying on promotion or demotion to reshape a large working set.
>>>>
>>>> After those pages have been placed, however, the policy cannot currently
>>>> opt in to migrate-on-fault placement.  set_mempolicy() and mbind()
>>>> reject MPOL_WEIGHTED_INTERLEAVE when MPOL_F_NUMA_BALANCING is specified,
>>>> so memory tiering cannot promote hot pages that were initially placed on
>>>> the slower nodes by the weighted policy.
>>>>
>>>> Initial placement is only a starting point.  Pages initially allocated on
>>>> fast memory are not necessarily the long-term hot pages, and pages
>>>> initially allocated on slower memory may become hot as the workload's hot
>>>> set changes.  The policy therefore needs to be able to combine weighted
>>>> initial placement with memory tiering's NUMA fault based hot-page
>>>> promotion.
>>> Ok, so memory allocation will respect the weights but balancing will ignore
>>> them? That really sounds rather odd to me.
>>>
>>> And I assume that was the reason why we might have disallowed the combination:
>>> it turns a weighted mechanism into an unweighted mechanism.
>>>
>>> So are we really sure these semantics that you would essentially set in stone
>>> here are the semantics we want? (ignoring weights)
>> Yes, that is a fair concern. It does look odd if the weights are
>> interpreted as a hard resident placement ratio.
>>
>> My understanding of the existing MPOL_WEIGHTED_INTERLEAVE ABI is that
>> the weights are allocation weights, not a long-term resident ratio. The
>> sysfs ABI documentation says that these weights only affect new
>> allocations, and that changing them at runtime will not migrate already
>> allocated pages.  The implementation also has normal allocation fallback
>> if the selected weighted target cannot satisfy the allocation.
>>
>> The use case here follows that interpretation.  The weights are used to
>> seed the initial placement across memory tiers.  After that, with an
>> explicit MPOL_F_NUMA_BALANCING opt-in, memory tiering would promote hot
>> pages based on access patterns, not based on the original allocation
>> ratio.  Users that want weighted allocation without that behavior would
>> keep using MPOL_WEIGHTED_INTERLEAVE without MPOL_F_NUMA_BALANCING.
>>
>> That said, I agree that allowing this flag combination would set the
>> semantics for it.  If reusing MPOL_F_NUMA_BALANCING for this is too
>> ambiguous, do you think we should model this as a separate opt-in ABI
>> for "weighted initial placement plus access-based balancing" instead?
>> For example, a separate flag or policy mode would make it clearer that
>> the weights are not intended to constrain migrate-on-fault placement.
>>
>> If the allocation-only interpretation of the weights is acceptable, I
>> can make it explicit in the commit message and documentation in v2.
>> Otherwise I would appreciate your suggestion on the preferred interface.
> I see two usecases for weighted interleave. Let's say you first allocate
> all the cold memory, and then all the hot memory using an interleave
> policy. This leaves the same hotness in both tiers since they are
> interleaved:
>
> In one scenario you might genuinely want to keep some hot memory in
> both tiers to maximize bandwidth utilization.
>
> But in the other case when you are not limited by bandwidth, you might
> indeed want to start out with a weighted interleave allocation to
> reduce variance on where hot memory lands, and then make promotion /
> demotion decisions based on access.
>
> Zhe, do you have a specific use-case in mind? For the second use
> case I would be curious to see if you see any meaningful performance
> differences in startup time when comapring a workload using other
> mempolicies for the allocation and then tiering vs. using weighted
> interleave and then tiering.
>
> But yeah, I think David is right that we should go over the semantics
> and really make sure this is what we want to commit to.
>
> Thanks again Zhe! Have a great day,
> Joshua

The use case is closer to the second one, but the main motivation is not
only startup-time performance.  The more important point is that each
workload has its own DDR/CXL budget assigned by the workload manager.

MPOL_WEIGHTED_INTERLEAVE is useful because it lets new allocations
follow that per-workload budget from the beginning, instead of placing
everything on one tier first and correcting the placement later. This is
important for workload orchestration because the workload starts from a
placement close to its assigned DDR/CXL ratio.

After that, we still need memory tiering because the initial weighted
placement does not know which pages will be hot. Some pages initially
placed on CXL may become hot, and some pages initially placed on DDR may
be cold. The intended model is therefore: weighted interleave provides
the per-workload allocation ratio, while promotion/demotion and reclaim
exchange hot and cold pages and try to keep the workload within its
assigned tier budget.

So I do not expect MPOL_WEIGHTED_INTERLEAVE alone to be a hard resident
ratio after migrations. The ratio is part of the workload placement
policy managed by the control plane. The kernel side still needs to let
hot pages from the slower tier be promoted when the workload explicitly
opts in to NUMA balancing.

Thanks a lot for taking a look at this.

Thanks,
Zhe

>
>> Thanks,
>> Zhe

^ permalink raw reply	[flat|nested] 14+ messages in thread

end of thread, other threads:[~2026-09-30 14:02 UTC | newest]

Thread overview: 14+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30  7:26 [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy Li Zhe
2026-09-30  8:43 ` Gregory Price
2026-09-30  8:59   ` Li Zhe
2026-09-30  9:11     ` Gregory Price
2026-09-30 10:59       ` Joshua Hahn
2026-09-30 11:22       ` Li Zhe
2026-09-30 12:25         ` Gregory Price
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52   ` Li Zhe
2026-09-30 12:11     ` Joshua Hahn
2026-09-30 14:02       ` Li Zhe
2026-09-30 12:03   ` Gregory Price
2026-09-30 12:32     ` David Hildenbrand (Arm)
2026-09-30 13:40       ` Gregory Price

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®