From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Gregory Price <gourry@gourry.net>, linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
mhocko@suse.com, mingo@redhat.com, peterz@infradead.org,
juri.lelli@redhat.com, vincent.guittot@linaro.org,
dietmar.eggemann@arm.com, rostedt@goodmis.org,
bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
kprateek.nayak@amd.com, ziy@nvidia.com,
baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
matthew.brost@intel.com, joshua.hahnjy@gmail.com,
rakie.kim@sk.com, byungchul@sk.com, ying.huang@linux.alibaba.com,
apopple@nvidia.com, jannh@google.com, pfalcato@suse.de,
hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com,
stable@vger.kernel.org
Subject: Re: [PATCH v3 2/7] mm: allow shared folios to be promoted to a fast tier
Date: Thu, 24 Sep 2026 13:59:15 +0200 [thread overview]
Message-ID: <f0cd5c54-deca-42e7-bc85-1838dfa917f3@kernel.org> (raw)
In-Reply-To: <20260922182928.2199090-3-gourry@gourry.net>
On 9/22/26 20:29, Gregory Price wrote:
> From: "Gregory Price (Meta)" <gourry@gourry.net>
>
> NUMA balancing rejects shared copy-on-write folios and executable
> file folios mapped by multiple processes to avoid east-west
> migration bouncing. These checks break promotion from slow memory.
Makes sense. So in general we're happy to promote if there is no chance of
bouncing (as would happen with ordinary numa migration).
>
> Allow such folios to participate when the folio being migrated is
> a low-tier folio and the destination is top-tier (south->north).
> This allows promotion, but prevents east-west or north->south
> migrations (north->south is handled by reclaim demotion).
>
> Keep the existing restrictions for ordinary placement and for
> migrations that are not slow-to-top-tier promotions.
>
> Rename folio_use_access_time() to folio_in_lowtier() so the source-tier
> condition is more obvious (timing is the mechanism, not the condition).
> The helper retains its existing behavior.
>
> Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory tiering system")
> Cc: stable@vger.kernel.org
> Suggested-by: Zi Yan <ziy@nvidia.com>
> Link: https://lore.kernel.org/r/DLHS4KFPQ86I.1J4LN3352RI71@nvidia.com
> Assisted-by: LLM
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
[...]
> --- a/mm/memory-tiers.c
> +++ b/mm/memory-tiers.c
> @@ -53,16 +53,19 @@ static const struct bus_type memory_tier_subsys = {
>
> #ifdef CONFIG_NUMA_BALANCING
> /**
> - * folio_use_access_time - check if a folio reuses cpupid for page access time
> + * folio_in_lowtier - check if a folio is in a tiering-managed lower tier
> * @folio: folio to check
> *
> * folio's _last_cpupid field is repurposed by memory tiering. In memory
> * tiering mode, cpupid of slow memory folio (not toptier memory) is used to
> * record page access time.
> *
> - * Return: the folio _last_cpupid is used to record page access time
> + * If memory tiering is disabled, then lowtier has no appreciable meaning,
> + * so we return false (the folio should not be migrated on this distinction).
> + *
> + * Return: true if memory tiering can promote the folio.
I wonder if a better name would just make the comment above (about memory
tiering) implicit and make more sense of the sysctl_numa_balancing_mode &
NUMA_BALANCING_MEMORY_TIERING) check.
folio_promotable()
folio_is_promotable()
folio_numa_promotable()
etc.
because that seems to be what we really test with both things combined.
> +++ b/mm/mempolicy.c
> @@ -864,8 +864,13 @@ bool folio_can_map_prot_numa(struct folio *folio, struct vm_area_struct *vma,
> if (!folio || folio_is_zone_device(folio) || folio_test_ksm(folio))
> return false;
>
> - /* Also skip shared copy-on-write folios */
> - if (vma_is_cow_mapping(vma) && folio_maybe_mapped_shared(folio))
> + /*
> + * Shared copy-on-write folios are poor east-west placement candidates.
> + * When tiering is enabled, folio_in_lowtier() identifies a promotable
> + * folio on a low tier, which needs a hint fault for promotion.
> + */
> + if (vma_is_cow_mapping(vma) && folio_maybe_mapped_shared(folio) &&
> + !folio_in_lowtier(folio))
> return false;
>
> /* Folios are pinned and can't be migrated */
> @@ -891,7 +896,8 @@ bool folio_can_map_prot_numa(struct folio *folio, struct vm_area_struct *vma,
> */
> if (vma_is_single_threaded_private(vma) && nid == numa_node_id())
> return false;
> - if (folio_use_access_time(folio))
> +
> + if (folio_in_lowtier(folio))
> folio_xchg_access_time(folio, jiffies_to_msecs(jiffies));
>
> return true;
> diff --git a/mm/migrate.c b/mm/migrate.c
> index 7bdcdb57652f8..6a08690cfe219 100644
> --- a/mm/migrate.c
> +++ b/mm/migrate.c
> @@ -2694,14 +2694,20 @@ int migrate_misplaced_folio_prepare(struct folio *folio,
>
> if (folio_is_file_lru(folio)) {
> /*
> - * Do not migrate file folios that are mapped in multiple
> - * processes with execute permissions as they are probably
> - * shared libraries.
> + * Limit east-east migration of file folios mapped in
> + * multiple processes with execute permissions as they
> + * are probably shared libraries (limits bouncing).
> + *
> + * If this is a low-tier folio, only migrate if the target
> + * node is toptier (this allows south->north migration while
> + * disallowing east-west migration between slow tiers).
> *
> * See folio_maybe_mapped_shared() on possible imprecision
> * when we cannot easily detect if a folio is shared.
> */
> - if ((vma->vm_flags & VM_EXEC) && folio_maybe_mapped_shared(folio))
> + if ((vma->vm_flags & VM_EXEC) &&
> + folio_maybe_mapped_shared(folio) &&
> + (!folio_in_lowtier(folio) || !node_is_toptier(node)))
> return -EACCES;
>
> /*
Sashiko has some comment about anonymous shared COW folios. I think it has a
point, but didn't look too closely.
--
Cheers,
David
next prev parent reply other threads:[~2026-09-24 11:59 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 18:29 [PATCH v3 0/7] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-22 18:29 ` [PATCH v3 1/7] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-24 11:45 ` David Hildenbrand (Arm)
2026-09-24 14:08 ` Gregory Price
2026-09-22 18:29 ` [PATCH v3 2/7] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-24 11:59 ` David Hildenbrand (Arm) [this message]
2026-09-24 14:07 ` Gregory Price
2026-09-24 15:32 ` David Hildenbrand (Arm)
2026-09-24 15:37 ` Gregory Price
2026-09-24 15:42 ` Zi Yan
2026-09-22 18:29 ` [PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-09-22 18:29 ` [PATCH v3 4/7] sched/numa: separate VMA placement from scan continuation Gregory Price
2026-09-22 18:29 ` [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion Gregory Price
2026-09-22 18:29 ` [PATCH v3 6/7] mm: use BIT() for change_protection() flags Gregory Price
2026-09-22 18:29 ` [PATCH v3 7/7] mm: use VMA flag helpers in NUMA balancing Gregory Price
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=f0cd5c54-deca-42e7-bc85-1838dfa917f3@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bsegall@google.com \
--cc=byungchul@sk.com \
--cc=dev.jain@arm.com \
--cc=dietmar.eggemann@arm.com \
--cc=gourry@gourry.net \
--cc=hannes@cmpxchg.org \
--cc=jannh@google.com \
--cc=joshua.hahnjy@gmail.com \
--cc=juri.lelli@redhat.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kprateek.nayak@amd.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=matthew.brost@intel.com \
--cc=mgorman@suse.de \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=nico.pache@linux.dev \
--cc=peterz@infradead.org \
--cc=pfalcato@suse.de \
--cc=raghavendra.kt@amd.com \
--cc=rakie.kim@sk.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shy828301@gmail.com \
--cc=stable@vger.kernel.org \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®