mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks
@ 2026-10-02 12:45 Vladimir Vdovin
  2026-10-02 14:23 ` Andrea Righi
  0 siblings, 1 reply; 3+ messages in thread
From: Vladimir Vdovin @ 2026-10-02 12:45 UTC (permalink / raw)
  To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min, Ingo Molnar,
	Peter Zijlstra, Juri Lelli, Vincent Guittot
  Cc: Vladimir Vdovin, Dietmar Eggemann, Steven Rostedt, Ben Segall,
	Mel Gorman, Valentin Schneider, K Prateek Nayak, Hui Su,
	sched-ext, linux-kernel

Automatic NUMA balancing is effectively off for every task in the ext
class, even when kernel.numa_balancing is enabled.

The periodic scan is queued from the fair tick only:

  task_tick_fair()
    if (static_branch_unlikely(&sched_numa_balancing))
            task_tick_numa(rq, curr);

task_tick_scx() has no counterpart, so task_numa_work() never runs for
an SCX task and no PTEs are made PROT_NONE. The fault side does not
depend on the scheduling class: task_numa_fault() still accounts faults,
runs task_numa_placement() and migrates misplaced folios. It just never
gets anything to work on.

Numbers from a 2-node, 160-CPU KVM hypervisor on 6.18.5 with a BPF
scheduler attached, detached for five minutes, then attached again:

                         SCX      fair (5 min)   SCX again
  numa_pte_updates/s     0        2.4M - 3.9M    0
  numa_hint_faults/s     ~0       10K - 34K      200 - 300, decaying

Those five minutes under fair changed the placement data considerably.
Measured from the BPF scheduler, as a share of task runtime:

                                         before    after
  task has no numa_preferred_nid         58%       10%
  task runs on its numa_preferred_nid    37%       83%

So under SCX p->numa_preferred_nid stays at whatever the last fair
period left behind, tasks created while a BPF scheduler is loaded never
get one, and memory is never migrated towards its users. A BPF scheduler
that wants to keep tasks close to their memory has nothing to read and
no way to start the scan itself.

Call task_tick_numa() from task_tick_scx(). It only needs
p->se.sum_exec_runtime, which SCX maintains through update_curr_common().

Keep task placement out of it: after a fault numa_migrate_preferred()
calls task_numa_migrate(), which picks a CPU from fair load statistics
and moves the task. Those statistics say nothing about SCX tasks and
CPU selection belongs to the BPF scheduler, so skip that step for them.
Scanning, fault accounting, numa_preferred_nid and folio migration keep
working.

Signed-off-by: Vladimir Vdovin <deliran@verdict.gg>
---
This is an RFC: the patch applies to v7.3-rc5 but has not been built or
run yet. The numbers above are from an unpatched 6.18.5 and only show
the problem. I would like to agree on the direction before testing it
properly.

Questions:

1. Is the missing scan intentional, or has nobody needed it so far?

2. Is calling a fair.c function from the ext tick acceptable? It keeps
   the logic in the class tick, in line with the per-class approach from
   the "sched: Fix execution-context tick handling under proxy
   execution" discussion, but it does make task_tick_numa() non-static.
   I can rebase on top of that series once it settles.

3. Is skipping task_numa_migrate() for SCX tasks the right split, or
   should the BPF scheduler get a say there through an ops callback?

4. Should this be opt-in through an ops flag, so that BPF schedulers
   that do not want the scan overhead are unaffected?

 kernel/sched/ext/ext.c | 3 +++
 kernel/sched/fair.c    | 14 +++++++++-----
 kernel/sched/sched.h   | 5 +++++
 3 files changed, 17 insertions(+), 5 deletions(-)

diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index e56c3c9..82c7d6a 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -3840,6 +3840,9 @@ static void task_tick_scx(struct rq *rq, struct task_struct *curr, int queued)
 
 	if (!curr->scx.slice)
 		resched_curr(rq);
+
+	if (!queued && static_branch_unlikely(&sched_numa_balancing))
+		task_tick_numa(rq, curr);
 }
 
 #ifdef CONFIG_EXT_GROUP_SCHED
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 57360f5..669c8ab 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3642,6 +3642,14 @@ static void numa_migrate_preferred(struct task_struct *p)
 	if (task_node(p) == p->numa_preferred_nid)
 		return;
 
+	/*
+	 * CPU selection for SCX tasks belongs to the BPF scheduler, which can
+	 * act on p->numa_preferred_nid itself. The statistics used below only
+	 * describe fair tasks anyway.
+	 */
+	if (task_on_scx(p))
+		return;
+
 	/* Otherwise, try migrate to a CPU on the preferred node */
 	task_numa_migrate(p);
 }
@@ -4630,7 +4638,7 @@ void init_numa_balancing(u64 clone_flags, struct task_struct *p)
 /*
  * Drive the periodic memory faults..
  */
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
+void task_tick_numa(struct rq *rq, struct task_struct *curr)
 {
 	struct callback_head *work = &curr->numa_work;
 	u64 period, now;
@@ -4696,10 +4704,6 @@ static void update_scan_period(struct task_struct *p, int new_cpu)
 
 #else /* !CONFIG_NUMA_BALANCING: */
 
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
-{
-}
-
 static inline void account_numa_enqueue(struct rq *rq, struct task_struct *p)
 {
 }
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c70..f6774bc 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2133,6 +2133,7 @@ extern int migrate_task_to(struct task_struct *p, int cpu);
 extern int migrate_swap(struct task_struct *p, struct task_struct *t,
 			int cpu, int scpu);
 extern void init_numa_balancing(u64 clone_flags, struct task_struct *p);
+extern void task_tick_numa(struct rq *rq, struct task_struct *curr);
 
 #else /* !CONFIG_NUMA_BALANCING: */
 
@@ -2141,6 +2142,10 @@ init_numa_balancing(u64 clone_flags, struct task_struct *p)
 {
 }
 
+static inline void task_tick_numa(struct rq *rq, struct task_struct *curr)
+{
+}
+
 #endif /* !CONFIG_NUMA_BALANCING */
 
 int task_llc(const struct task_struct *p);

base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
-- 
2.47.0

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-02 15:08 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 12:45 [RFC PATCH] sched_ext: Drive the NUMA balancing scan for SCX tasks Vladimir Vdovin
2026-10-02 14:23 ` Andrea Righi
2026-10-02 15:08   ` Vladimir Vdovin

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®