mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Gregory Price <gourry@gourry.net>
To: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
	akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
	liam@infradead.org, vbabka@kernel.org, rppt@kernel.org,
	surenb@google.com, mhocko@suse.com, mingo@redhat.com,
	peterz@infradead.org, juri.lelli@redhat.com,
	vincent.guittot@linaro.org, dietmar.eggemann@arm.com,
	rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de,
	vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com,
	baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
	ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
	lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
	matthew.brost@intel.com, joshua.hahnjy@gmail.com,
	rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net,
	ying.huang@linux.alibaba.com, apopple@nvidia.com,
	jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org,
	shy828301@gmail.com, raghavendra.kt@amd.com,
	stable@vger.kernel.org
Subject: [PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode
Date: Tue, 22 Sep 2026 14:29:24 -0400	[thread overview]
Message-ID: <20260922182928.2199090-4-gourry@gourry.net> (raw)
In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net>

From: "Gregory Price (Meta)" <gourry@gourry.net>

Commit 4591ce4f2d22 ("sched/numa: Do not trap hinting faults for
shared libraries") excludes file-backed read-only VMAs from NUMA
hint faulting to prevent east-west placement bouncing.

This filter hides hot file folios on slow memory from promotion.

Scan those VMAs when tiering is enabled, but make their scans promotion
only to retains the existing restriction. Keep the historical VMA
predicate unchanged for backport-ability.

Read the balancing mode once per task_numa_work() invocation and build
the protection flags for each VMA from that snapshot. This keeps the PTE
and PMD paths on the same policy for an entire protection walk (which
may be split across multiple scanning periods).

On a host with 768 GB of DRAM and 256 GB of CXL memory running two
database services using ~430GB each, 169 MB of their shared 185 MB
main binary accumulated on CXL before permanently stuck there.

With the change, the binary tier residency tracks its runtime hotness.

Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory tiering system")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
---
 kernel/sched/fair.c | 31 +++++++++++++++++++++----------
 1 file changed, 21 insertions(+), 10 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index dc78d24ed8bc0..8a4687f67d82f 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4119,21 +4119,21 @@ static bool vma_is_accessed(struct mm_struct *mm, struct vm_area_struct *vma)
  */
 static void task_numa_work(struct callback_head *work)
 {
+	const unsigned int numab_mode = READ_ONCE(sysctl_numa_balancing_mode);
+	const bool tiering = numab_mode & NUMA_BALANCING_MEMORY_TIERING;
 	unsigned long migrate, next_scan, now = jiffies;
 	struct task_struct *p = current;
 	struct mm_struct *mm = p->mm;
 	u64 runtime = p->se.sum_exec_runtime;
 	struct vm_area_struct *vma;
-	unsigned long cp_flags = MM_CP_PROT_NUMA;
+	unsigned long cp_flags;
 	unsigned long start, end;
 	unsigned long nr_pte_updates = 0;
 	long pages, virtpages;
 	struct vma_iterator vmi;
 	bool vma_pids_skipped;
 	bool vma_pids_forced = false;
-
-	if (!(READ_ONCE(sysctl_numa_balancing_mode) & NUMA_BALANCING_NORMAL))
-		cp_flags |= MM_CP_PROT_NUMA_PROMO_ONLY;
+	bool placement_scan;
 
 	WARN_ON_ONCE(p != container_of(work, struct task_struct, numa_work));
 
@@ -4221,13 +4221,19 @@ static void task_numa_work(struct callback_head *work)
 		}
 
 		/*
-		 * Shared library pages mapped by multiple processes are not
-		 * migrated as it is expected they are cache replicated. Avoid
-		 * hinting faults in read-only file-backed mappings or the vDSO
-		 * as migrating the pages will be of marginal benefit.
+		 * Shared library pages mapped by multiple processes are limited
+		 * to south->north migrations as it is expected they are cache
+		 * replicated. The benefit of east-west migration in this case
+		 * is at best marginal and may be harmful due to TLB/cache
+		 * invalidation.
+		 *
+		 * Allow promotion as a cold page incurring many cache-misses
+		 * under cache pressure can drive considerable bandwidth.
 		 */
-		if (!vma->vm_mm ||
-		    (vma->vm_file && (vma->vm_flags & (VM_READ|VM_WRITE)) == (VM_READ))) {
+		placement_scan = !(vma->vm_file &&
+			(vma->vm_flags & (VM_READ | VM_WRITE)) == VM_READ);
+
+		if (!vma->vm_mm || (!placement_scan && !tiering)) {
 			trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO);
 			continue;
 		}
@@ -4307,6 +4313,11 @@ static void task_numa_work(struct callback_head *work)
 			continue;
 		}
 
+		placement_scan &= numab_mode & NUMA_BALANCING_NORMAL;
+		cp_flags = MM_CP_PROT_NUMA;
+		if (!placement_scan)
+			cp_flags |= MM_CP_PROT_NUMA_PROMO_ONLY;
+
 		do {
 			start = max(start, vma->vm_start);
 			end = ALIGN(start + (pages << PAGE_SHIFT), HPAGE_SIZE);
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-09-22 18:29 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 18:29 [PATCH v3 0/7] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-22 18:29 ` [PATCH v3 1/7] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-24 11:45   ` David Hildenbrand (Arm)
2026-09-24 14:08     ` Gregory Price
2026-09-22 18:29 ` [PATCH v3 2/7] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-24 11:59   ` David Hildenbrand (Arm)
2026-09-24 14:07     ` Gregory Price
2026-09-24 15:32       ` David Hildenbrand (Arm)
2026-09-24 15:37         ` Gregory Price
2026-09-24 15:42           ` Zi Yan
2026-09-22 18:29 ` Gregory Price [this message]
2026-09-24 20:53   ` [PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode David Hildenbrand (Arm)
2026-09-25  0:35     ` Gregory Price
2026-09-25 10:04       ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 4/7] sched/numa: separate VMA placement from scan continuation Gregory Price
2026-09-25 10:53   ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion Gregory Price
2026-09-25 10:55   ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 6/7] mm: use BIT() for change_protection() flags Gregory Price
2026-09-24 20:38   ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 7/7] mm: use VMA flag helpers in NUMA balancing Gregory Price
2026-09-24 20:39   ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922182928.2199090-4-gourry@gourry.net \
    --to=gourry@gourry.net \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=bsegall@google.com \
    --cc=byungchul@sk.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=hannes@cmpxchg.org \
    --cc=jannh@google.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=juri.lelli@redhat.com \
    --cc=kas@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=kprateek.nayak@amd.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=matthew.brost@intel.com \
    --cc=mgorman@suse.de \
    --cc=mhocko@suse.com \
    --cc=mingo@redhat.com \
    --cc=nico.pache@linux.dev \
    --cc=peterz@infradead.org \
    --cc=pfalcato@suse.de \
    --cc=raghavendra.kt@amd.com \
    --cc=rakie.kim@sk.com \
    --cc=rostedt@goodmis.org \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shy828301@gmail.com \
    --cc=stable@vger.kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®