From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Gregory Price <gourry@gourry.net>, linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
mhocko@suse.com, mingo@redhat.com, peterz@infradead.org,
juri.lelli@redhat.com, vincent.guittot@linaro.org,
dietmar.eggemann@arm.com, rostedt@goodmis.org,
bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
kprateek.nayak@amd.com, ziy@nvidia.com,
baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
matthew.brost@intel.com, joshua.hahnjy@gmail.com,
rakie.kim@sk.com, byungchul@sk.com, ying.huang@linux.alibaba.com,
apopple@nvidia.com, jannh@google.com, pfalcato@suse.de,
hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com,
stable@vger.kernel.org
Subject: Re: [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion
Date: Fri, 25 Sep 2026 12:55:09 +0200 [thread overview]
Message-ID: <ad99014e-3714-45a5-ad23-aa7fe72745cb@kernel.org> (raw)
In-Reply-To: <20260922182928.2199090-6-gourry@gourry.net>
On 9/22/26 20:29, Gregory Price wrote:
> From: "Gregory Price (Meta)" <gourry@gourry.net>
>
> Commit fc137c0ddab2 ("sched/numa: enhance vma scanning logic") skips
> VMAs without recent PID activity. Since only NUMA hint faults record
> that activity, the filter can suppress the fault needed to promote hot
> slow-tier memory.
>
> Let memory tiering bypass the PID scan filter. In combined mode, use the
> placement decision to restrict top-tier sampling to VMAs that need it,
> while inactive VMAs still receive promotion-only scans.
>
> Track the last completed placement scan separately from scans of any
> kind. Promotion-only scans still update prev_scan_seq, but do not
> advance the placement-starvation horizon.
>
> On a host with 768 GB of DRAM and 256 GB of CXL memory, one large shmem
> VMA consumed most scanning activity, while 2,537 other VMAs covering
> 84 GB were skipped as inactive. One stand-out result: a hot 20 GB hash
> table ended up trapped entirely on CXL and drove CXL bandwidth
> utilization beyond sustainable levels - resulting in a large regression.
>
> With this series, the hot hash table ends up split evenly between DRAM
> and CXL, tier residency tracked runtime load, and CXL bandwidth
> utilization drops from 45GB/s (maxed) to 5-10GB/s, while DRAM bandwidth
> utilization increases from ~200GB/s to 250GB/s+, resulting in major
> throughput improvements for the database workload.
>
> Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic")
> Cc: stable@vger.kernel.org
> Assisted-by: OpenAI:gpt-5
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
> include/linux/mm_types.h | 7 +++++++
> kernel/sched/fair.c | 26 +++++++++++++++++++++-----
> 2 files changed, 28 insertions(+), 5 deletions(-)
>
> diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
> index fd35db969bc94..dcca3ead9db59 100644
> --- a/include/linux/mm_types.h
> +++ b/include/linux/mm_types.h
> @@ -804,6 +804,13 @@ struct vma_numab_state {
> */
> int prev_scan_seq;
>
> + /*
> + * MM scan sequence ID when the VMA was last scanned for placement.
> + * The starvation horizon in vma_needs_placement_scan() counts against
> + * this, so promotion-only scans cannot postpone placement indefinitely.
> + */
> + int prev_placement_scan_seq;
> +
> /* Preserve placement-scan eligibility during an in-progress scan. */
> bool placement_scan;
> };
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index a2849e72c4e26..412c72084a63d 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4097,7 +4097,7 @@ static bool vma_needs_placement_scan(struct mm_struct *mm,
> * threads can help scan this vma, force a vma scan.
> */
> if (READ_ONCE(mm->numa_scan_seq) >
> - (vma->numab_state->prev_scan_seq + get_nr_threads(current)))
> + (vma->numab_state->prev_placement_scan_seq + get_nr_threads(current)))
> return true;
>
> return false;
> @@ -4267,7 +4267,8 @@ static void task_numa_work(struct callback_head *work)
> * to prevent VMAs being skipped prematurely on the
> * first scan:
> */
> - vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
> + vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
> + vma->numab_state->prev_placement_scan_seq = mm->numa_scan_seq - 1;
> }
>
> /*
> @@ -4300,10 +4301,13 @@ static void task_numa_work(struct callback_head *work)
> * Do not scan the VMA if a task has not accessed it, unless no other
> * VMA candidate exists. If a scan is already in-progress, finish it,
> * but track continuation separately from starting a new one.
> + *
> + * The PID filter must not gate promotion. Allow PID-inactive VMAs
> + * to proceed when memory tiering is enabled.
> */
> placement_due = vma_needs_placement_scan(mm, vma);
> scan_started = mm->numa_scan_offset > vma->vm_start;
> - pid_scan_allowed = vma_pids_forced || placement_due;
> + pid_scan_allowed = tiering || vma_pids_forced || placement_due;
>
> if (!pid_scan_allowed) {
> if (scan_started) {
> @@ -4315,10 +4319,16 @@ static void task_numa_work(struct callback_head *work)
> }
> }
>
> - /* Keep scan policy stable while processing a VMA in chunks.*/
> + /*
> + * Keep scan policy stable while processing a VMA in chunks.
> + * A fault in one chunk can make a VMA placement-eligible. Keep a
> + * promotion-only decision sticky for the rest of a partial scan.
> + */
> placement_scan &= numab_mode & NUMA_BALANCING_NORMAL;
> if (scan_started)
> placement_scan &= vma->numab_state->placement_scan;
> + else if (tiering)
> + placement_scan &= placement_due;
Same comment. Apart from that LGTM.
--
Cheers,
David
next prev parent reply other threads:[~2026-09-25 10:55 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 18:29 [PATCH v3 0/7] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-22 18:29 ` [PATCH v3 1/7] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-24 11:45 ` David Hildenbrand (Arm)
2026-09-24 14:08 ` Gregory Price
2026-09-22 18:29 ` [PATCH v3 2/7] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-24 11:59 ` David Hildenbrand (Arm)
2026-09-24 14:07 ` Gregory Price
2026-09-24 15:32 ` David Hildenbrand (Arm)
2026-09-24 15:37 ` Gregory Price
2026-09-24 15:42 ` Zi Yan
2026-09-22 18:29 ` [PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-09-24 20:53 ` David Hildenbrand (Arm)
2026-09-25 0:35 ` Gregory Price
2026-09-25 10:04 ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 4/7] sched/numa: separate VMA placement from scan continuation Gregory Price
2026-09-25 10:53 ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion Gregory Price
2026-09-25 10:55 ` David Hildenbrand (Arm) [this message]
2026-09-22 18:29 ` [PATCH v3 6/7] mm: use BIT() for change_protection() flags Gregory Price
2026-09-24 20:38 ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 7/7] mm: use VMA flag helpers in NUMA balancing Gregory Price
2026-09-24 20:39 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ad99014e-3714-45a5-ad23-aa7fe72745cb@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bsegall@google.com \
--cc=byungchul@sk.com \
--cc=dev.jain@arm.com \
--cc=dietmar.eggemann@arm.com \
--cc=gourry@gourry.net \
--cc=hannes@cmpxchg.org \
--cc=jannh@google.com \
--cc=joshua.hahnjy@gmail.com \
--cc=juri.lelli@redhat.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kprateek.nayak@amd.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=matthew.brost@intel.com \
--cc=mgorman@suse.de \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=nico.pache@linux.dev \
--cc=peterz@infradead.org \
--cc=pfalcato@suse.de \
--cc=raghavendra.kt@amd.com \
--cc=rakie.kim@sk.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shy828301@gmail.com \
--cc=stable@vger.kernel.org \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®