mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Gregory Price <gourry@gourry.net>, linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
	akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
	vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
	mhocko@suse.com, mingo@redhat.com, peterz@infradead.org,
	juri.lelli@redhat.com, vincent.guittot@linaro.org,
	dietmar.eggemann@arm.com, rostedt@goodmis.org,
	bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
	kprateek.nayak@amd.com, ziy@nvidia.com,
	baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
	ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
	lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
	matthew.brost@intel.com, joshua.hahnjy@gmail.com,
	rakie.kim@sk.com, byungchul@sk.com, ying.huang@linux.alibaba.com,
	apopple@nvidia.com, jannh@google.com, pfalcato@suse.de,
	hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com,
	stable@vger.kernel.org
Subject: Re: [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion
Date: Fri, 25 Sep 2026 12:55:09 +0200	[thread overview]
Message-ID: <ad99014e-3714-45a5-ad23-aa7fe72745cb@kernel.org> (raw)
In-Reply-To: <20260922182928.2199090-6-gourry@gourry.net>

On 9/22/26 20:29, Gregory Price wrote:
> From: "Gregory Price (Meta)" <gourry@gourry.net>
> 
> Commit fc137c0ddab2 ("sched/numa: enhance vma scanning logic") skips
> VMAs without recent PID activity. Since only NUMA hint faults record
> that activity, the filter can suppress the fault needed to promote hot
> slow-tier memory.
> 
> Let memory tiering bypass the PID scan filter. In combined mode, use the
> placement decision to restrict top-tier sampling to VMAs that need it,
> while inactive VMAs still receive promotion-only scans.
> 
> Track the last completed placement scan separately from scans of any
> kind. Promotion-only scans still update prev_scan_seq, but do not
> advance the placement-starvation horizon.
> 
> On a host with 768 GB of DRAM and 256 GB of CXL memory, one large shmem
> VMA consumed most scanning activity, while 2,537 other VMAs covering
> 84 GB were skipped as inactive. One stand-out result: a hot 20 GB hash
> table ended up trapped entirely on CXL and drove CXL bandwidth
> utilization beyond sustainable levels - resulting in a large regression.
> 
> With this series, the hot hash table ends up split evenly between DRAM
> and CXL, tier residency tracked runtime load, and CXL bandwidth
> utilization drops from 45GB/s (maxed) to 5-10GB/s, while DRAM bandwidth
> utilization increases from ~200GB/s to 250GB/s+, resulting in major
> throughput improvements for the database workload.
> 
> Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic")
> Cc: stable@vger.kernel.org
> Assisted-by: OpenAI:gpt-5
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
>  include/linux/mm_types.h |  7 +++++++
>  kernel/sched/fair.c      | 26 +++++++++++++++++++++-----
>  2 files changed, 28 insertions(+), 5 deletions(-)
> 
> diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
> index fd35db969bc94..dcca3ead9db59 100644
> --- a/include/linux/mm_types.h
> +++ b/include/linux/mm_types.h
> @@ -804,6 +804,13 @@ struct vma_numab_state {
>  	 */
>  	int prev_scan_seq;
>  
> +	/*
> +	 * MM scan sequence ID when the VMA was last scanned for placement.
> +	 * The starvation horizon in vma_needs_placement_scan() counts against
> +	 * this, so promotion-only scans cannot postpone placement indefinitely.
> +	 */
> +	int prev_placement_scan_seq;
> +
>  	/* Preserve placement-scan eligibility during an in-progress scan. */
>  	bool placement_scan;
>  };
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index a2849e72c4e26..412c72084a63d 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4097,7 +4097,7 @@ static bool vma_needs_placement_scan(struct mm_struct *mm,
>  	 * threads can help scan this vma, force a vma scan.
>  	 */
>  	if (READ_ONCE(mm->numa_scan_seq) >
> -	   (vma->numab_state->prev_scan_seq + get_nr_threads(current)))
> +	   (vma->numab_state->prev_placement_scan_seq + get_nr_threads(current)))
>  		return true;
>  
>  	return false;
> @@ -4267,7 +4267,8 @@ static void task_numa_work(struct callback_head *work)
>  			 * to prevent VMAs being skipped prematurely on the
>  			 * first scan:
>  			 */
> -			 vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
> +			vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
> +			vma->numab_state->prev_placement_scan_seq = mm->numa_scan_seq - 1;
>  		}
>  
>  		/*
> @@ -4300,10 +4301,13 @@ static void task_numa_work(struct callback_head *work)
>  		 * Do not scan the VMA if a task has not accessed it, unless no other
>  		 * VMA candidate exists. If a scan is already in-progress, finish it,
>  		 * but track continuation separately from starting a new one.
> +		 *
> +		 * The PID filter must not gate promotion. Allow PID-inactive VMAs
> +		 * to proceed when memory tiering is enabled.
>  		 */
>  		placement_due = vma_needs_placement_scan(mm, vma);
>  		scan_started = mm->numa_scan_offset > vma->vm_start;
> -		pid_scan_allowed = vma_pids_forced || placement_due;
> +		pid_scan_allowed = tiering || vma_pids_forced || placement_due;
>  
>  		if (!pid_scan_allowed) {
>  			if (scan_started) {
> @@ -4315,10 +4319,16 @@ static void task_numa_work(struct callback_head *work)
>  			}
>  		}
>  
> -		/* Keep scan policy stable while processing a VMA in chunks.*/
> +		/*
> +		 * Keep scan policy stable while processing a VMA in chunks.
> +		 * A fault in one chunk can make a VMA placement-eligible. Keep a
> +		 * promotion-only decision sticky for the rest of a partial scan.
> +		 */
>  		placement_scan &= numab_mode & NUMA_BALANCING_NORMAL;
>  		if (scan_started)
>  			placement_scan &= vma->numab_state->placement_scan;
> +		else if (tiering)
> +			placement_scan &= placement_due;

Same comment. Apart from that LGTM.

-- 
Cheers,

David

  reply	other threads:[~2026-09-25 10:55 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 18:29 [PATCH v3 0/7] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-22 18:29 ` [PATCH v3 1/7] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-24 11:45   ` David Hildenbrand (Arm)
2026-09-24 14:08     ` Gregory Price
2026-09-22 18:29 ` [PATCH v3 2/7] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-24 11:59   ` David Hildenbrand (Arm)
2026-09-24 14:07     ` Gregory Price
2026-09-24 15:32       ` David Hildenbrand (Arm)
2026-09-24 15:37         ` Gregory Price
2026-09-24 15:42           ` Zi Yan
2026-09-22 18:29 ` [PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-09-24 20:53   ` David Hildenbrand (Arm)
2026-09-25  0:35     ` Gregory Price
2026-09-25 10:04       ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 4/7] sched/numa: separate VMA placement from scan continuation Gregory Price
2026-09-25 10:53   ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion Gregory Price
2026-09-25 10:55   ` David Hildenbrand (Arm) [this message]
2026-09-22 18:29 ` [PATCH v3 6/7] mm: use BIT() for change_protection() flags Gregory Price
2026-09-24 20:38   ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 7/7] mm: use VMA flag helpers in NUMA balancing Gregory Price
2026-09-24 20:39   ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ad99014e-3714-45a5-ad23-aa7fe72745cb@kernel.org \
    --to=david@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=bsegall@google.com \
    --cc=byungchul@sk.com \
    --cc=dev.jain@arm.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=jannh@google.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=juri.lelli@redhat.com \
    --cc=kas@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=kprateek.nayak@amd.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=matthew.brost@intel.com \
    --cc=mgorman@suse.de \
    --cc=mhocko@suse.com \
    --cc=mingo@redhat.com \
    --cc=nico.pache@linux.dev \
    --cc=peterz@infradead.org \
    --cc=pfalcato@suse.de \
    --cc=raghavendra.kt@amd.com \
    --cc=rakie.kim@sk.com \
    --cc=rostedt@goodmis.org \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shy828301@gmail.com \
    --cc=stable@vger.kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®