* [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks
@ 2026-10-02 12:45 Vladimir Vdovin
2026-10-02 14:23 ` Andrea Righi
0 siblings, 1 reply; 3+ messages in thread
From: Vladimir Vdovin @ 2026-10-02 12:45 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min, Ingo Molnar,
Peter Zijlstra, Juri Lelli, Vincent Guittot
Cc: Vladimir Vdovin, Dietmar Eggemann, Steven Rostedt, Ben Segall,
Mel Gorman, Valentin Schneider, K Prateek Nayak, Hui Su,
sched-ext, linux-kernel
Automatic NUMA balancing is effectively off for every task in the ext
class, even when kernel.numa_balancing is enabled.
The periodic scan is queued from the fair tick only:
task_tick_fair()
if (static_branch_unlikely(&sched_numa_balancing))
task_tick_numa(rq, curr);
task_tick_scx() has no counterpart, so task_numa_work() never runs for
an SCX task and no PTEs are made PROT_NONE. The fault side does not
depend on the scheduling class: task_numa_fault() still accounts faults,
runs task_numa_placement() and migrates misplaced folios. It just never
gets anything to work on.
Numbers from a 2-node, 160-CPU KVM hypervisor on 6.18.5 with a BPF
scheduler attached, detached for five minutes, then attached again:
SCX fair (5 min) SCX again
numa_pte_updates/s 0 2.4M - 3.9M 0
numa_hint_faults/s ~0 10K - 34K 200 - 300, decaying
Those five minutes under fair changed the placement data considerably.
Measured from the BPF scheduler, as a share of task runtime:
before after
task has no numa_preferred_nid 58% 10%
task runs on its numa_preferred_nid 37% 83%
So under SCX p->numa_preferred_nid stays at whatever the last fair
period left behind, tasks created while a BPF scheduler is loaded never
get one, and memory is never migrated towards its users. A BPF scheduler
that wants to keep tasks close to their memory has nothing to read and
no way to start the scan itself.
Call task_tick_numa() from task_tick_scx(). It only needs
p->se.sum_exec_runtime, which SCX maintains through update_curr_common().
Keep task placement out of it: after a fault numa_migrate_preferred()
calls task_numa_migrate(), which picks a CPU from fair load statistics
and moves the task. Those statistics say nothing about SCX tasks and
CPU selection belongs to the BPF scheduler, so skip that step for them.
Scanning, fault accounting, numa_preferred_nid and folio migration keep
working.
Signed-off-by: Vladimir Vdovin <deliran@verdict.gg>
---
This is an RFC: the patch applies to v7.3-rc5 but has not been built or
run yet. The numbers above are from an unpatched 6.18.5 and only show
the problem. I would like to agree on the direction before testing it
properly.
Questions:
1. Is the missing scan intentional, or has nobody needed it so far?
2. Is calling a fair.c function from the ext tick acceptable? It keeps
the logic in the class tick, in line with the per-class approach from
the "sched: Fix execution-context tick handling under proxy
execution" discussion, but it does make task_tick_numa() non-static.
I can rebase on top of that series once it settles.
3. Is skipping task_numa_migrate() for SCX tasks the right split, or
should the BPF scheduler get a say there through an ops callback?
4. Should this be opt-in through an ops flag, so that BPF schedulers
that do not want the scan overhead are unaffected?
kernel/sched/ext/ext.c | 3 +++
kernel/sched/fair.c | 14 +++++++++-----
kernel/sched/sched.h | 5 +++++
3 files changed, 17 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index e56c3c9..82c7d6a 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -3840,6 +3840,9 @@ static void task_tick_scx(struct rq *rq, struct task_struct *curr, int queued)
if (!curr->scx.slice)
resched_curr(rq);
+
+ if (!queued && static_branch_unlikely(&sched_numa_balancing))
+ task_tick_numa(rq, curr);
}
#ifdef CONFIG_EXT_GROUP_SCHED
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 57360f5..669c8ab 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3642,6 +3642,14 @@ static void numa_migrate_preferred(struct task_struct *p)
if (task_node(p) == p->numa_preferred_nid)
return;
+ /*
+ * CPU selection for SCX tasks belongs to the BPF scheduler, which can
+ * act on p->numa_preferred_nid itself. The statistics used below only
+ * describe fair tasks anyway.
+ */
+ if (task_on_scx(p))
+ return;
+
/* Otherwise, try migrate to a CPU on the preferred node */
task_numa_migrate(p);
}
@@ -4630,7 +4638,7 @@ void init_numa_balancing(u64 clone_flags, struct task_struct *p)
/*
* Drive the periodic memory faults..
*/
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
+void task_tick_numa(struct rq *rq, struct task_struct *curr)
{
struct callback_head *work = &curr->numa_work;
u64 period, now;
@@ -4696,10 +4704,6 @@ static void update_scan_period(struct task_struct *p, int new_cpu)
#else /* !CONFIG_NUMA_BALANCING: */
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
-{
-}
-
static inline void account_numa_enqueue(struct rq *rq, struct task_struct *p)
{
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c70..f6774bc 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2133,6 +2133,7 @@ extern int migrate_task_to(struct task_struct *p, int cpu);
extern int migrate_swap(struct task_struct *p, struct task_struct *t,
int cpu, int scpu);
extern void init_numa_balancing(u64 clone_flags, struct task_struct *p);
+extern void task_tick_numa(struct rq *rq, struct task_struct *curr);
#else /* !CONFIG_NUMA_BALANCING: */
@@ -2141,6 +2142,10 @@ init_numa_balancing(u64 clone_flags, struct task_struct *p)
{
}
+static inline void task_tick_numa(struct rq *rq, struct task_struct *curr)
+{
+}
+
#endif /* !CONFIG_NUMA_BALANCING */
int task_llc(const struct task_struct *p);
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.47.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks
2026-10-02 12:45 [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks Vladimir Vdovin
@ 2026-10-02 14:23 ` Andrea Righi
2026-10-02 15:08 ` Vladimir Vdovin
0 siblings, 1 reply; 3+ messages in thread
From: Andrea Righi @ 2026-10-02 14:23 UTC (permalink / raw)
To: Vladimir Vdovin
Cc: Tejun Heo, David Vernet, Changwoo Min, Ingo Molnar,
Peter Zijlstra, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Hui Su, sched-ext, linux-kernel
Hi Vladimir,
On Fri, Oct 02, 2026 at 03:45:31PM +0300, Vladimir Vdovin wrote:
> Automatic NUMA balancing is effectively off for every task in the ext
> class, even when kernel.numa_balancing is enabled.
>
> The periodic scan is queued from the fair tick only:
>
> task_tick_fair()
> if (static_branch_unlikely(&sched_numa_balancing))
> task_tick_numa(rq, curr);
>
> task_tick_scx() has no counterpart, so task_numa_work() never runs for
> an SCX task and no PTEs are made PROT_NONE. The fault side does not
> depend on the scheduling class: task_numa_fault() still accounts faults,
> runs task_numa_placement() and migrates misplaced folios. It just never
> gets anything to work on.
>
> Numbers from a 2-node, 160-CPU KVM hypervisor on 6.18.5 with a BPF
> scheduler attached, detached for five minutes, then attached again:
>
> SCX fair (5 min) SCX again
> numa_pte_updates/s 0 2.4M - 3.9M 0
> numa_hint_faults/s ~0 10K - 34K 200 - 300, decaying
>
> Those five minutes under fair changed the placement data considerably.
> Measured from the BPF scheduler, as a share of task runtime:
>
> before after
> task has no numa_preferred_nid 58% 10%
> task runs on its numa_preferred_nid 37% 83%
>
> So under SCX p->numa_preferred_nid stays at whatever the last fair
> period left behind, tasks created while a BPF scheduler is loaded never
> get one, and memory is never migrated towards its users. A BPF scheduler
> that wants to keep tasks close to their memory has nothing to read and
> no way to start the scan itself.
>
> Call task_tick_numa() from task_tick_scx(). It only needs
> p->se.sum_exec_runtime, which SCX maintains through update_curr_common().
>
> Keep task placement out of it: after a fault numa_migrate_preferred()
> calls task_numa_migrate(), which picks a CPU from fair load statistics
> and moves the task. Those statistics say nothing about SCX tasks and
> CPU selection belongs to the BPF scheduler, so skip that step for them.
> Scanning, fault accounting, numa_preferred_nid and folio migration keep
> working.
>
> Signed-off-by: Vladimir Vdovin <deliran@verdict.gg>
Thanks for looking at this! I actually have a patch series in my backlog to
introduce NUMA-balancing support in sched_ext. It makes NUMA hinting scans
opt-in for BPF schedulers and exposes a task's preferred NUMA node to BPF.
It also lets a BPF scheduler set a per-task memory target, which NUMA balancing
can use when migrating pages after hinting faults. I think that could give
schedulers a useful way to coordinate CPU and memory placement, includeing when
device locality constratins CPU choice.
We're actually going to discuss this topic at LPC next week in Prague during the
sched_ext MC session.
I can probably send the patch series at this point, it's not completely
well-tested, but it might be useful to have it on the list, so that we can start
discussing and improving it. I'll keep you in the loop.
Thanks,
-Andrea
> ---
> This is an RFC: the patch applies to v7.3-rc5 but has not been built or
> run yet. The numbers above are from an unpatched 6.18.5 and only show
> the problem. I would like to agree on the direction before testing it
> properly.
>
> Questions:
>
> 1. Is the missing scan intentional, or has nobody needed it so far?
>
> 2. Is calling a fair.c function from the ext tick acceptable? It keeps
> the logic in the class tick, in line with the per-class approach from
> the "sched: Fix execution-context tick handling under proxy
> execution" discussion, but it does make task_tick_numa() non-static.
> I can rebase on top of that series once it settles.
>
> 3. Is skipping task_numa_migrate() for SCX tasks the right split, or
> should the BPF scheduler get a say there through an ops callback?
>
> 4. Should this be opt-in through an ops flag, so that BPF schedulers
> that do not want the scan overhead are unaffected?
>
> kernel/sched/ext/ext.c | 3 +++
> kernel/sched/fair.c | 14 +++++++++-----
> kernel/sched/sched.h | 5 +++++
> 3 files changed, 17 insertions(+), 5 deletions(-)
>
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index e56c3c9..82c7d6a 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -3840,6 +3840,9 @@ static void task_tick_scx(struct rq *rq, struct task_struct *curr, int queued)
>
> if (!curr->scx.slice)
> resched_curr(rq);
> +
> + if (!queued && static_branch_unlikely(&sched_numa_balancing))
> + task_tick_numa(rq, curr);
> }
>
> #ifdef CONFIG_EXT_GROUP_SCHED
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 57360f5..669c8ab 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -3642,6 +3642,14 @@ static void numa_migrate_preferred(struct task_struct *p)
> if (task_node(p) == p->numa_preferred_nid)
> return;
>
> + /*
> + * CPU selection for SCX tasks belongs to the BPF scheduler, which can
> + * act on p->numa_preferred_nid itself. The statistics used below only
> + * describe fair tasks anyway.
> + */
> + if (task_on_scx(p))
> + return;
> +
> /* Otherwise, try migrate to a CPU on the preferred node */
> task_numa_migrate(p);
> }
> @@ -4630,7 +4638,7 @@ void init_numa_balancing(u64 clone_flags, struct task_struct *p)
> /*
> * Drive the periodic memory faults..
> */
> -static void task_tick_numa(struct rq *rq, struct task_struct *curr)
> +void task_tick_numa(struct rq *rq, struct task_struct *curr)
> {
> struct callback_head *work = &curr->numa_work;
> u64 period, now;
> @@ -4696,10 +4704,6 @@ static void update_scan_period(struct task_struct *p, int new_cpu)
>
> #else /* !CONFIG_NUMA_BALANCING: */
>
> -static void task_tick_numa(struct rq *rq, struct task_struct *curr)
> -{
> -}
> -
> static inline void account_numa_enqueue(struct rq *rq, struct task_struct *p)
> {
> }
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index e656c70..f6774bc 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -2133,6 +2133,7 @@ extern int migrate_task_to(struct task_struct *p, int cpu);
> extern int migrate_swap(struct task_struct *p, struct task_struct *t,
> int cpu, int scpu);
> extern void init_numa_balancing(u64 clone_flags, struct task_struct *p);
> +extern void task_tick_numa(struct rq *rq, struct task_struct *curr);
>
> #else /* !CONFIG_NUMA_BALANCING: */
>
> @@ -2141,6 +2142,10 @@ init_numa_balancing(u64 clone_flags, struct task_struct *p)
> {
> }
>
> +static inline void task_tick_numa(struct rq *rq, struct task_struct *curr)
> +{
> +}
> +
> #endif /* !CONFIG_NUMA_BALANCING */
>
> int task_llc(const struct task_struct *p);
>
> base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
> --
> 2.47.0
>
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks
2026-10-02 14:23 ` Andrea Righi
@ 2026-10-02 15:08 ` Vladimir Vdovin
0 siblings, 0 replies; 3+ messages in thread
From: Vladimir Vdovin @ 2026-10-02 15:08 UTC (permalink / raw)
To: Andrea Righi, Vladimir Vdovin
Cc: Tejun Heo, David Vernet, Changwoo Min, Ingo Molnar,
Peter Zijlstra, Juri Lelli, Vincent Guittot, Dietmar Eggemann,
Steven Rostedt, Ben Segall, Mel Gorman, Valentin Schneider,
K Prateek Nayak, Hui Su, sched-ext, linux-kernel
Hi Andrea,
On Fri, Oct 02, 2026, Andrea Righi wrote:
> Thanks for looking at this! I actually have a patch series in my backlog to
> introduce NUMA-balancing support in sched_ext. It makes NUMA hinting scans
> opt-in for BPF schedulers and exposes a task's preferred NUMA node to BPF.
That is great news, thanks. My patch was only a small sketch to show the
problem, so I am happy to leave it at that and follow your series
instead.
Opt-in scanning plus the preferred node exposed to BPF sounds like what
I was looking for. My scheduler runs on KVM hypervisors and reads
p->numa_preferred_nid to choose a home node for each vCPU thread,
falling back to the thread group's majority when a task has none. Under
SCX that value stops being updated, which is how I noticed.
> It also lets a BPF scheduler set a per-task memory target, which NUMA balancing
> can use when migrating pages after hinting faults.
This part sounds interesting too. Guests that span two nodes are the
hard case for me: today I can only follow the memory, I cannot ask for
it to follow the vCPUs.
> I can probably send the patch series at this point, it's not completely
> well-tested, but it might be useful to have it on the list, so that we can start
> discussing and improving it. I'll keep you in the loop.
Thanks, I will be watching for it. If it is of any use, I can try it on
a few hypervisors (2- and 4-node hosts) and share numbers for scan rate,
hint faults and time spent on the preferred node.
Thanks,
Vladimir
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-02 15:08 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 12:45 [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks Vladimir Vdovin
2026-10-02 14:23 ` Andrea Righi
2026-10-02 15:08 ` Vladimir Vdovin
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®