mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Andrea Righi <arighi@nvidia.com>
To: Vladimir Vdovin <deliran@verdict.gg>
Cc: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
	Changwoo Min <changwoo@igalia.com>,
	Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>, Hui Su <sh_def@163.com>,
	sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks
Date: Fri, 2 Oct 2026 16:23:56 +0200	[thread overview]
Message-ID: <ar--fEmqMyX8odNn@gpd4> (raw)
In-Reply-To: <20261002124559.10367-1-deliran@verdict.gg>

Hi Vladimir,

On Fri, Oct 02, 2026 at 03:45:31PM +0300, Vladimir Vdovin wrote:
> Automatic NUMA balancing is effectively off for every task in the ext
> class, even when kernel.numa_balancing is enabled.
> 
> The periodic scan is queued from the fair tick only:
> 
>   task_tick_fair()
>     if (static_branch_unlikely(&sched_numa_balancing))
>             task_tick_numa(rq, curr);
> 
> task_tick_scx() has no counterpart, so task_numa_work() never runs for
> an SCX task and no PTEs are made PROT_NONE. The fault side does not
> depend on the scheduling class: task_numa_fault() still accounts faults,
> runs task_numa_placement() and migrates misplaced folios. It just never
> gets anything to work on.
> 
> Numbers from a 2-node, 160-CPU KVM hypervisor on 6.18.5 with a BPF
> scheduler attached, detached for five minutes, then attached again:
> 
>                          SCX      fair (5 min)   SCX again
>   numa_pte_updates/s     0        2.4M - 3.9M    0
>   numa_hint_faults/s     ~0       10K - 34K      200 - 300, decaying
> 
> Those five minutes under fair changed the placement data considerably.
> Measured from the BPF scheduler, as a share of task runtime:
> 
>                                          before    after
>   task has no numa_preferred_nid         58%       10%
>   task runs on its numa_preferred_nid    37%       83%
> 
> So under SCX p->numa_preferred_nid stays at whatever the last fair
> period left behind, tasks created while a BPF scheduler is loaded never
> get one, and memory is never migrated towards its users. A BPF scheduler
> that wants to keep tasks close to their memory has nothing to read and
> no way to start the scan itself.
> 
> Call task_tick_numa() from task_tick_scx(). It only needs
> p->se.sum_exec_runtime, which SCX maintains through update_curr_common().
> 
> Keep task placement out of it: after a fault numa_migrate_preferred()
> calls task_numa_migrate(), which picks a CPU from fair load statistics
> and moves the task. Those statistics say nothing about SCX tasks and
> CPU selection belongs to the BPF scheduler, so skip that step for them.
> Scanning, fault accounting, numa_preferred_nid and folio migration keep
> working.
> 
> Signed-off-by: Vladimir Vdovin <deliran@verdict.gg>

Thanks for looking at this! I actually have a patch series in my backlog to
introduce NUMA-balancing support in sched_ext. It makes NUMA hinting scans
opt-in for BPF schedulers and exposes a task's preferred NUMA node to BPF.

It also lets a BPF scheduler set a per-task memory target, which NUMA balancing
can use when migrating pages after hinting faults. I think that could give
schedulers a useful way to coordinate CPU and memory placement, includeing when
device locality constratins CPU choice.

We're actually going to discuss this topic at LPC next week in Prague during the
sched_ext MC session.

I can probably send the patch series at this point, it's not completely
well-tested, but it might be useful to have it on the list, so that we can start
discussing and improving it. I'll keep you in the loop.

Thanks,
-Andrea

> ---
> This is an RFC: the patch applies to v7.3-rc5 but has not been built or
> run yet. The numbers above are from an unpatched 6.18.5 and only show
> the problem. I would like to agree on the direction before testing it
> properly.
> 
> Questions:
> 
> 1. Is the missing scan intentional, or has nobody needed it so far?
> 
> 2. Is calling a fair.c function from the ext tick acceptable? It keeps
>    the logic in the class tick, in line with the per-class approach from
>    the "sched: Fix execution-context tick handling under proxy
>    execution" discussion, but it does make task_tick_numa() non-static.
>    I can rebase on top of that series once it settles.
> 
> 3. Is skipping task_numa_migrate() for SCX tasks the right split, or
>    should the BPF scheduler get a say there through an ops callback?
> 
> 4. Should this be opt-in through an ops flag, so that BPF schedulers
>    that do not want the scan overhead are unaffected?
> 
>  kernel/sched/ext/ext.c | 3 +++
>  kernel/sched/fair.c    | 14 +++++++++-----
>  kernel/sched/sched.h   | 5 +++++
>  3 files changed, 17 insertions(+), 5 deletions(-)
> 
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index e56c3c9..82c7d6a 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -3840,6 +3840,9 @@ static void task_tick_scx(struct rq *rq, struct task_struct *curr, int queued)
>  
>  	if (!curr->scx.slice)
>  		resched_curr(rq);
> +
> +	if (!queued && static_branch_unlikely(&sched_numa_balancing))
> +		task_tick_numa(rq, curr);
>  }
>  
>  #ifdef CONFIG_EXT_GROUP_SCHED
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 57360f5..669c8ab 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -3642,6 +3642,14 @@ static void numa_migrate_preferred(struct task_struct *p)
>  	if (task_node(p) == p->numa_preferred_nid)
>  		return;
>  
> +	/*
> +	 * CPU selection for SCX tasks belongs to the BPF scheduler, which can
> +	 * act on p->numa_preferred_nid itself. The statistics used below only
> +	 * describe fair tasks anyway.
> +	 */
> +	if (task_on_scx(p))
> +		return;
> +
>  	/* Otherwise, try migrate to a CPU on the preferred node */
>  	task_numa_migrate(p);
>  }
> @@ -4630,7 +4638,7 @@ void init_numa_balancing(u64 clone_flags, struct task_struct *p)
>  /*
>   * Drive the periodic memory faults..
>   */
> -static void task_tick_numa(struct rq *rq, struct task_struct *curr)
> +void task_tick_numa(struct rq *rq, struct task_struct *curr)
>  {
>  	struct callback_head *work = &curr->numa_work;
>  	u64 period, now;
> @@ -4696,10 +4704,6 @@ static void update_scan_period(struct task_struct *p, int new_cpu)
>  
>  #else /* !CONFIG_NUMA_BALANCING: */
>  
> -static void task_tick_numa(struct rq *rq, struct task_struct *curr)
> -{
> -}
> -
>  static inline void account_numa_enqueue(struct rq *rq, struct task_struct *p)
>  {
>  }
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index e656c70..f6774bc 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -2133,6 +2133,7 @@ extern int migrate_task_to(struct task_struct *p, int cpu);
>  extern int migrate_swap(struct task_struct *p, struct task_struct *t,
>  			int cpu, int scpu);
>  extern void init_numa_balancing(u64 clone_flags, struct task_struct *p);
> +extern void task_tick_numa(struct rq *rq, struct task_struct *curr);
>  
>  #else /* !CONFIG_NUMA_BALANCING: */
>  
> @@ -2141,6 +2142,10 @@ init_numa_balancing(u64 clone_flags, struct task_struct *p)
>  {
>  }
>  
> +static inline void task_tick_numa(struct rq *rq, struct task_struct *curr)
> +{
> +}
> +
>  #endif /* !CONFIG_NUMA_BALANCING */
>  
>  int task_llc(const struct task_struct *p);
> 
> base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
> -- 
> 2.47.0
> 

  reply	other threads:[~2026-10-02 14:24 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-02 12:45 Vladimir Vdovin
2026-10-02 14:23 ` Andrea Righi [this message]
2026-10-02 15:08   ` Vladimir Vdovin

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ar--fEmqMyX8odNn@gpd4 \
    --to=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=changwoo@igalia.com \
    --cc=deliran@verdict.gg \
    --cc=dietmar.eggemann@arm.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=sh_def@163.com \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=void@manifault.com \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®