From: Gregory Price <gourry@gourry.net>
To: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
liam@infradead.org, vbabka@kernel.org, rppt@kernel.org,
surenb@google.com, mhocko@suse.com, mingo@redhat.com,
peterz@infradead.org, juri.lelli@redhat.com,
vincent.guittot@linaro.org, dietmar.eggemann@arm.com,
rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de,
vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com,
baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
gourry@gourry.net, joshua.hahnjy@gmail.com, rakie.kim@sk.com,
ying.huang@linux.alibaba.com, matthew.brost@intel.com,
byungchul@sk.com, apopple@nvidia.com, jannh@google.com,
pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com,
osalvador@suse.de, raghavendra.kt@amd.com,
stable@vger.kernel.org
Subject: [PATCH v4 5/7] sched/numa: scan PID-inactive VMAs for promotion
Date: Wed, 30 Sep 2026 07:22:04 -0400 [thread overview]
Message-ID: <20260930112206.205083-6-gourry@gourry.net> (raw)
In-Reply-To: <20260930112206.205083-1-gourry@gourry.net>
From: "Gregory Price (Meta)" <gourry@gourry.net>
Commit fc137c0ddab2 ("sched/numa: enhance vma scanning logic") skips
VMAs without recent PID activity. Since only NUMA hint faults record
that activity, the filter can suppress the fault needed to promote hot
slow-tier memory.
Let memory tiering bypass the PID scan filter. In combined mode, use the
placement decision to restrict top-tier sampling to VMAs that need it,
while inactive VMAs still receive promotion-only scans.
Make this decision only when a VMA scan starts. placement_due is
evaluated for the current task, and a partially scanned VMA is usually
resumed by a different thread, so re-evaluating it for each chunk would
demote in-progress placement scans and starve large VMAs of placement
sampling in multi-threaded processes.
Track the last completed placement scan separately from scans of any
kind. Promotion-only scans still update prev_scan_seq, but do not
advance the placement-starvation horizon.
On a host with 768 GB of DRAM and 256 GB of CXL memory, one large shmem
VMA consumed most scanning activity, while 2,537 other VMAs covering
84 GB were skipped as inactive. One stand-out result: a hot 20 GB hash
table ended up trapped entirely on CXL and drove CXL bandwidth
utilization beyond sustainable levels - resulting in a large regression.
With this series, the hot hash table ends up split evenly between DRAM
and CXL, tier residency tracks runtime load, and CXL bandwidth
utilization drops from 45GB/s (maxed) to 5-10GB/s, while DRAM bandwidth
utilization increases from ~200GB/s to 250GB/s+, resulting in major
throughput improvements for the database workload.
Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
---
include/linux/mm_types.h | 7 +++++++
kernel/sched/fair.c | 31 +++++++++++++++++++++++++------
2 files changed, 32 insertions(+), 6 deletions(-)
diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
index d0ac21d02c47..86cd12a61755 100644
--- a/include/linux/mm_types.h
+++ b/include/linux/mm_types.h
@@ -804,6 +804,13 @@ struct vma_numab_state {
*/
int prev_scan_seq;
+ /*
+ * MM scan sequence ID when the VMA was last scanned for placement.
+ * The starvation horizon in vma_needs_placement_scan() counts against
+ * this, so promotion-only scans cannot postpone placement indefinitely.
+ */
+ int prev_placement_scan_seq;
+
/*
* The in-progress scan of this VMA is promotion-only.
* Resumed scans finish with the policy they started with.
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 3b30786ded3b..6002882b9d30 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4105,7 +4105,7 @@ static bool vma_needs_placement_scan(struct mm_struct *mm,
* threads can help scan this vma, force a vma scan.
*/
if (READ_ONCE(mm->numa_scan_seq) >
- (vma->numab_state->prev_scan_seq + get_nr_threads(current)))
+ (vma->numab_state->prev_placement_scan_seq + get_nr_threads(current)))
return true;
return false;
@@ -4279,7 +4279,8 @@ static void task_numa_work(struct callback_head *work)
* to prevent VMAs being skipped prematurely on the
* first scan:
*/
- vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
+ vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
+ vma->numab_state->prev_placement_scan_seq = mm->numa_scan_seq - 1;
}
/*
@@ -4312,10 +4313,13 @@ static void task_numa_work(struct callback_head *work)
* Do not scan the VMA if a task has not accessed it, unless no other
* VMA candidate exists. If a scan is already in-progress, finish it,
* but track continuation separately from starting a new one.
+ *
+ * The PID filter must not gate promotion. Allow PID-inactive VMAs
+ * to proceed when memory tiering is enabled.
*/
placement_due = vma_needs_placement_scan(mm, vma);
scan_started = mm->numa_scan_offset > vma->vm_start;
- pid_scan_allowed = vma_pids_forced || placement_due;
+ pid_scan_allowed = tiering || vma_pids_forced || placement_due;
if (!pid_scan_allowed) {
if (scan_started) {
@@ -4327,9 +4331,18 @@ static void task_numa_work(struct callback_head *work)
}
}
- /* Keep scan policy stable while processing a VMA in chunks. */
- if (scan_started && vma->numab_state->promo_only)
+ /*
+ * Keep scan policy stable while processing a VMA in chunks.
+ * placement_due is evaluated for the current task, and a
+ * different thread will usually resume the scan. Finish an
+ * in-progress scan with the policy it started with.
+ */
+ if (scan_started) {
+ if (vma->numab_state->promo_only)
+ placement_scan = false;
+ } else if (tiering && !placement_due) {
placement_scan = false;
+ }
vma->numab_state->promo_only = !placement_scan;
cp_flags = MM_CP_PROT_NUMA;
@@ -4362,8 +4375,14 @@ static void task_numa_work(struct callback_head *work)
cond_resched();
} while (end != vma->vm_end);
- /* VMA scan is complete, do not scan until next sequence. */
+ /*
+ * VMA scan is complete, do not scan until next sequence.
+ * A promotion-only scan did not cover top-tier folios, so it
+ * does not count towards the placement-scan starvation check.
+ */
vma->numab_state->prev_scan_seq = mm->numa_scan_seq;
+ if (placement_scan)
+ vma->numab_state->prev_placement_scan_seq = mm->numa_scan_seq;
vma->numab_state->promo_only = false;
/*
--
2.55.0
next prev parent reply other threads:[~2026-09-30 11:22 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 11:21 [PATCH v4 0/7] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-30 11:22 ` [PATCH v4 1/7] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-30 11:22 ` [PATCH v4 2/7] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-10-01 10:34 ` David Hildenbrand (Arm)
2026-09-30 11:22 ` [PATCH v4 3/7] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-10-01 10:34 ` David Hildenbrand (Arm)
2026-09-30 11:22 ` [PATCH v4 4/7] sched/numa: separate VMA placement from scan continuation Gregory Price
2026-10-01 10:44 ` David Hildenbrand (Arm)
2026-10-01 13:29 ` Gregory Price
2026-09-30 11:22 ` Gregory Price [this message]
2026-10-01 10:45 ` [PATCH v4 5/7] sched/numa: scan PID-inactive VMAs for promotion David Hildenbrand (Arm)
2026-09-30 11:22 ` [PATCH v4 6/7] mm: use BIT() for change_protection() flags Gregory Price
2026-09-30 11:22 ` [PATCH v4 7/7] mm: use VMA flag helpers in NUMA balancing Gregory Price
2026-10-01 10:45 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260930112206.205083-6-gourry@gourry.net \
--to=gourry@gourry.net \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bsegall@google.com \
--cc=byungchul@sk.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=dietmar.eggemann@arm.com \
--cc=hannes@cmpxchg.org \
--cc=jannh@google.com \
--cc=joshua.hahnjy@gmail.com \
--cc=juri.lelli@redhat.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kprateek.nayak@amd.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=matthew.brost@intel.com \
--cc=mgorman@suse.de \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=nico.pache@linux.dev \
--cc=osalvador@suse.de \
--cc=peterz@infradead.org \
--cc=pfalcato@suse.de \
--cc=raghavendra.kt@amd.com \
--cc=rakie.kim@sk.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shy828301@gmail.com \
--cc=stable@vger.kernel.org \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®