* [PATCH 1/5] sched/numa: Let other scheduling classes drive NUMA scanning
2026-10-04 7:27 [PATCHSET sched_ext/for-7.4] sched_ext: Add NUMA balancing support Andrea Righi
@ 2026-10-04 7:27 ` Andrea Righi
2026-10-04 7:27 ` [PATCH 2/5] sched/numa: Leave the placement of a BPF-scheduled task to its scheduler Andrea Righi
` (3 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Andrea Righi @ 2026-10-04 7:27 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Vladimir Vdovin,
Emil Tsalapatis, Christian Loehle, Balbir Singh, Lee Trager,
sched-ext, linux-kernel
The periodic scan that installs the NUMA hinting faults is driven by
task_tick_numa(), which is reachable only from task_tick_fair(). What
the scan feeds is class independent, but a task that is not in the fair
class is never scanned, records no faults and never gets a preferred
node.
Make task_tick_numa() available to the other scheduling classes (e.g.,
sched_ext). Declare it in sched.h and move the !CONFIG_NUMA_BALANCING
stub there, so a class that has no interest in NUMA balancing keeps
costing nothing.
No functional change: task_tick_fair() remains the only caller.
Reported-by: Vladimir Vdovin <deliran@verdict.gg>
Link: https://lore.kernel.org/r/20261002124559.10367-1-deliran@verdict.gg
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/sched/fair.c | 6 +-----
kernel/sched/sched.h | 5 +++++
2 files changed, 6 insertions(+), 5 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 03206e15e6fe4..e2d52faacdb7a 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4431,7 +4431,7 @@ void init_numa_balancing(u64 clone_flags, struct task_struct *p)
/*
* Drive the periodic memory faults..
*/
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
+void task_tick_numa(struct rq *rq, struct task_struct *curr)
{
struct callback_head *work = &curr->numa_work;
u64 period, now;
@@ -4497,10 +4497,6 @@ static void update_scan_period(struct task_struct *p, int new_cpu)
#else /* !CONFIG_NUMA_BALANCING: */
-static void task_tick_numa(struct rq *rq, struct task_struct *curr)
-{
-}
-
static inline void account_numa_enqueue(struct rq *rq, struct task_struct *p)
{
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 0a34f0da1ec36..30232341db9de 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -2137,6 +2137,7 @@ extern int migrate_task_to(struct task_struct *p, int cpu);
extern int migrate_swap(struct task_struct *p, struct task_struct *t,
int cpu, int scpu);
extern void init_numa_balancing(u64 clone_flags, struct task_struct *p);
+extern void task_tick_numa(struct rq *rq, struct task_struct *curr);
#else /* !CONFIG_NUMA_BALANCING: */
@@ -2145,6 +2146,10 @@ init_numa_balancing(u64 clone_flags, struct task_struct *p)
{
}
+static inline void task_tick_numa(struct rq *rq, struct task_struct *curr)
+{
+}
+
#endif /* !CONFIG_NUMA_BALANCING */
int task_llc(const struct task_struct *p);
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH 2/5] sched/numa: Leave the placement of a BPF-scheduled task to its scheduler
2026-10-04 7:27 [PATCHSET sched_ext/for-7.4] sched_ext: Add NUMA balancing support Andrea Righi
2026-10-04 7:27 ` [PATCH 1/5] sched/numa: Let other scheduling classes drive NUMA scanning Andrea Righi
@ 2026-10-04 7:27 ` Andrea Righi
2026-10-04 7:27 ` [PATCH 3/5] sched_ext: Scan NUMA hinting faults for opted-in BPF schedulers Andrea Righi
` (2 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Andrea Righi @ 2026-10-04 7:27 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Vladimir Vdovin,
Emil Tsalapatis, Christian Loehle, Balbir Singh, Lee Trager,
sched-ext, linux-kernel
A task under a BPF scheduler can take NUMA hinting faults via
task_numa_fault() and then reach numa_migrate_preferred(), which may
migrate the task to its preferred node through task_numa_migrate().
The migration of such a task should always be driven by its BPF
scheduler, which owns the placement of its tasks.
Keep collecting the statistics, which is what makes
p->numa_preferred_nid worth reading, but skip the migration for a task
owned by a BPF scheduler. The scheduler can read the preferred node and
decide for itself whether moving the task is worth the cost.
Reported-by: Vladimir Vdovin <deliran@verdict.gg>
Link: https://lore.kernel.org/r/20261002124559.10367-1-deliran@verdict.gg
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/sched/fair.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index e2d52faacdb7a..019c1331545bf 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3446,6 +3446,14 @@ static void numa_migrate_preferred(struct task_struct *p)
if (task_node(p) == p->numa_preferred_nid)
return;
+ /*
+ * A task under a BPF scheduler is placed by that scheduler. Keep the
+ * statistics coming, which is what makes p->numa_preferred_nid worth
+ * reading, but leave the placement alone.
+ */
+ if (task_on_scx(p))
+ return;
+
/* Otherwise, try migrate to a CPU on the preferred node */
task_numa_migrate(p);
}
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH 3/5] sched_ext: Scan NUMA hinting faults for opted-in BPF schedulers
2026-10-04 7:27 [PATCHSET sched_ext/for-7.4] sched_ext: Add NUMA balancing support Andrea Righi
2026-10-04 7:27 ` [PATCH 1/5] sched/numa: Let other scheduling classes drive NUMA scanning Andrea Righi
2026-10-04 7:27 ` [PATCH 2/5] sched/numa: Leave the placement of a BPF-scheduled task to its scheduler Andrea Righi
@ 2026-10-04 7:27 ` Andrea Righi
2026-10-04 7:27 ` [PATCH 4/5] sched_ext: Add scx_bpf_task_numa_nid() Andrea Righi
2026-10-04 7:27 ` [PATCH 5/5] selftests/sched_ext: Add a test for scx_bpf_task_numa_nid() Andrea Righi
4 siblings, 0 replies; 6+ messages in thread
From: Andrea Righi @ 2026-10-04 7:27 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Vladimir Vdovin,
Emil Tsalapatis, Christian Loehle, Balbir Singh, Lee Trager,
sched-ext, linux-kernel
A task running under a BPF scheduler never has its address space scanned
for NUMA hinting faults. As a result, mm/memory.c never calls
task_numa_fault() and the task's NUMA fault statistics and preferred
node remain empty.
Scanning is not free and a scheduler which does not use the result
should not pay for it. Add SCX_OPS_NUMA_BALANCING to let a BPF scheduler
request the NUMA hinting-fault scan for its tasks. Drive the scan from
the sched_ext tick for tasks owned by an opted-in scheduler, the way
task_tick_fair() does, under the existing sched_numa_balancing static
key. The hrtick, which only fires to end the running task's slice, does
not drive the scan, as in task_tick_fair().
The hinting faults then work for these tasks as they do for fair tasks:
they maintain the NUMA statistics, the preferred node and potentially
migrate the memory a task accesses toward the node the task runs on. The
tasks themselves are not migrated, the BPF scheduler remains responsible
for placing them.
No new accounting is needed to pace the scan, task_tick_numa() spaces
scans by the task's CPU time, p->se.sum_exec_runtime and sched_ext
already maintains that field through update_curr_common().
Reported-by: Vladimir Vdovin <deliran@verdict.gg>
Link: https://lore.kernel.org/r/20261002124559.10367-1-deliran@verdict.gg
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/sched/ext/ext.c | 13 +++++++++++++
kernel/sched/ext/internal.h | 16 +++++++++++++++-
tools/sched_ext/include/scx/compat.h | 1 +
tools/sched_ext/include/scx/enum_defs.autogen.h | 1 +
tools/sched_ext/include/scx/enums_abi.autogen.h | 3 ++-
5 files changed, 32 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 96c904b396019..b22db5ccb91ab 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -4227,6 +4227,19 @@ static void task_tick_scx(struct rq *rq, struct task_struct *donor, int queued)
else if (SCX_HAS_OP(sch, tick))
SCX_CALL_OP_TASK(sch, tick, rq, donor);
+ /*
+ * If requested by the scheduler, drive the NUMA hinting-fault scan as
+ * task_tick_fair() does. Neither the scan nor the statistics it feeds
+ * depend on the scheduling class. Where the task then runs is left to
+ * that scheduler: numa_migrate_preferred() does not move a task it owns.
+ *
+ * The tick that only refreshes an already queued task does not scan,
+ * matching the @queued check in task_tick_fair().
+ */
+ if (!queued && (sch->ops.flags & SCX_OPS_NUMA_BALANCING) &&
+ static_branch_unlikely(&sched_numa_balancing))
+ task_tick_numa(rq, donor);
+
if (!donor->scx.slice) {
/* the slice can't be trusted while bypassing */
if (READ_ONCE(donor->scx.lazy_resched) &&
diff --git a/kernel/sched/ext/internal.h b/kernel/sched/ext/internal.h
index d65fec631bdf7..10cbbee48e094 100644
--- a/kernel/sched/ext/internal.h
+++ b/kernel/sched/ext/internal.h
@@ -243,6 +243,19 @@ enum scx_ops_flags {
*/
SCX_OPS_ENQ_BLOCKED = 1LLU << 10,
+ /*
+ * If set, drive automatic NUMA hinting-fault scans for tasks owned by
+ * this scheduler. The faults maintain the tasks' NUMA statistics and
+ * preferred node, which can be queried with scx_bpf_task_numa_nid(),
+ * and migrate the memory a task accesses toward the node it runs on, as
+ * NUMA balancing does for fair tasks. Tasks themselves are never
+ * migrated: the BPF scheduler remains responsible for placing them.
+ *
+ * If clear, sched_ext does not initiate NUMA hinting-fault scans for the
+ * scheduler's tasks.
+ */
+ SCX_OPS_NUMA_BALANCING = 1LLU << 11,
+
SCX_OPS_ALL_FLAGS = SCX_OPS_KEEP_BUILTIN_IDLE |
SCX_OPS_ENQ_LAST |
SCX_OPS_ENQ_EXITING |
@@ -253,7 +266,8 @@ enum scx_ops_flags {
SCX_OPS_ALWAYS_ENQ_IMMED |
SCX_OPS_TID_TO_TASK |
SCX_OPS_LAZY_RESCHED |
- SCX_OPS_ENQ_BLOCKED,
+ SCX_OPS_ENQ_BLOCKED |
+ SCX_OPS_NUMA_BALANCING,
/* high 8 bits are internal, don't include in SCX_OPS_ALL_FLAGS */
__SCX_OPS_INTERNAL_MASK = 0xffLLU << 56,
diff --git a/tools/sched_ext/include/scx/compat.h b/tools/sched_ext/include/scx/compat.h
index 07fb55e63c57c..f4c35f8750730 100644
--- a/tools/sched_ext/include/scx/compat.h
+++ b/tools/sched_ext/include/scx/compat.h
@@ -215,6 +215,7 @@ static inline bool __COMPAT_struct_has_field(const char *type, const char *field
#define SCX_OPS_BUILTIN_IDLE_PER_NODE SCX_OPS_FLAG(SCX_OPS_BUILTIN_IDLE_PER_NODE)
#define SCX_OPS_ALWAYS_ENQ_IMMED SCX_OPS_FLAG(SCX_OPS_ALWAYS_ENQ_IMMED)
#define SCX_OPS_ENQ_BLOCKED SCX_OPS_FLAG(SCX_OPS_ENQ_BLOCKED)
+#define SCX_OPS_NUMA_BALANCING SCX_OPS_FLAG(SCX_OPS_NUMA_BALANCING)
#define SCX_PICK_IDLE_FLAG(name) __COMPAT_ENUM_OR_ZERO("scx_pick_idle_cpu_flags", #name)
diff --git a/tools/sched_ext/include/scx/enum_defs.autogen.h b/tools/sched_ext/include/scx/enum_defs.autogen.h
index 4fa82033b99c9..1bed9c9a96f29 100644
--- a/tools/sched_ext/include/scx/enum_defs.autogen.h
+++ b/tools/sched_ext/include/scx/enum_defs.autogen.h
@@ -172,6 +172,7 @@
#define HAVE_SCX_OPS_TID_TO_TASK
#define HAVE_SCX_OPS_LAZY_RESCHED
#define HAVE_SCX_OPS_ENQ_BLOCKED
+#define HAVE_SCX_OPS_NUMA_BALANCING
#define HAVE_SCX_OPS_ALL_FLAGS
#define HAVE___SCX_OPS_INTERNAL_MASK
#define HAVE_SCX_OPS_HAS_CPU_PREEMPT
diff --git a/tools/sched_ext/include/scx/enums_abi.autogen.h b/tools/sched_ext/include/scx/enums_abi.autogen.h
index 1186daa9cb141..1cc1aad845023 100644
--- a/tools/sched_ext/include/scx/enums_abi.autogen.h
+++ b/tools/sched_ext/include/scx/enums_abi.autogen.h
@@ -184,7 +184,8 @@ static const struct __scx_enum_abi_val __scx_enum_abi_vals[]
{ "scx_ops_flags", "SCX_OPS_TID_TO_TASK", 0x100LLU },
{ "scx_ops_flags", "SCX_OPS_LAZY_RESCHED", 0x200LLU },
{ "scx_ops_flags", "SCX_OPS_ENQ_BLOCKED", 0x400LLU },
- { "scx_ops_flags", "SCX_OPS_ALL_FLAGS", 0x7ffLLU },
+ { "scx_ops_flags", "SCX_OPS_NUMA_BALANCING", 0x800LLU },
+ { "scx_ops_flags", "SCX_OPS_ALL_FLAGS", 0xfffLLU },
{ "scx_ops_flags", "__SCX_OPS_INTERNAL_MASK", 0xff00000000000000LLU },
{ "scx_ops_flags", "SCX_OPS_HAS_CPU_PREEMPT", 0x100000000000000LLU },
{ "scx_ops_state", "SCX_OPSS_NONE", 0x0LLU },
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH 4/5] sched_ext: Add scx_bpf_task_numa_nid()
2026-10-04 7:27 [PATCHSET sched_ext/for-7.4] sched_ext: Add NUMA balancing support Andrea Righi
` (2 preceding siblings ...)
2026-10-04 7:27 ` [PATCH 3/5] sched_ext: Scan NUMA hinting faults for opted-in BPF schedulers Andrea Righi
@ 2026-10-04 7:27 ` Andrea Righi
2026-10-04 7:27 ` [PATCH 5/5] selftests/sched_ext: Add a test for scx_bpf_task_numa_nid() Andrea Righi
4 siblings, 0 replies; 6+ messages in thread
From: Andrea Righi @ 2026-10-04 7:27 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Vladimir Vdovin,
Emil Tsalapatis, Christian Loehle, Balbir Singh, Lee Trager,
sched-ext, linux-kernel
With the hinting-fault scan running for tasks under a BPF scheduler,
p->numa_preferred_nid is maintained for them. Give schedulers a way to
read it.
scx_bpf_task_numa_nid() returns the node a task's memory accesses are
concentrated on, or NUMA_NO_NODE when the task has no preference. A BPF
scheduler can use it to keep a task near its memory when placing or
balancing it.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/sched/ext/ext.c | 26 ++++++++++++++++++++++++
tools/sched_ext/include/scx/compat.bpf.h | 10 +++++++++
2 files changed, 36 insertions(+)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index b22db5ccb91ab..21caafd2cfe0c 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -11140,6 +11140,31 @@ __bpf_kfunc s32 scx_bpf_task_cid(const struct task_struct *p)
return tbl[task_cpu(p)];
}
+/**
+ * scx_bpf_task_numa_nid - Node NUMA balancing would prefer a task to run on
+ * @p: task of interest
+ *
+ * Return p->numa_preferred_nid: the node NUMA balancing has derived from @p's
+ * hinting faults as the best one to run @p on, which is the node with the most
+ * faults, adjusted for CPU-less nodes and for the NUMA group @p may be part of.
+ * Return NUMA_NO_NODE when @p has no preference yet. A task has none until its
+ * address space has been scanned a few times and every task has none with
+ * kernel.numa_balancing disabled.
+ *
+ * The value is a hint, not an instruction: nothing in the kernel moves a task
+ * under a BPF scheduler to this node and the scheduler is free to weigh it
+ * against its own placement rules. The value is updated under @p's rq lock,
+ * which this kfunc doesn't take, so it is a snapshot.
+ */
+__bpf_kfunc s32 scx_bpf_task_numa_nid(const struct task_struct *p)
+{
+#ifdef CONFIG_NUMA_BALANCING
+ return READ_ONCE(p->numa_preferred_nid);
+#else
+ return NUMA_NO_NODE;
+#endif
+}
+
/**
* scx_bpf_locked_rq - Return the rq currently locked by SCX
* @aux: implicit BPF argument to access bpf_prog_aux hidden from BPF progs
@@ -11446,6 +11471,7 @@ BTF_ID_FLAGS(func, scx_bpf_put_cpumask, KF_RELEASE)
BTF_ID_FLAGS(func, scx_bpf_task_running, KF_RCU)
BTF_ID_FLAGS(func, scx_bpf_task_cpu, KF_RCU)
BTF_ID_FLAGS(func, scx_bpf_task_cid, KF_RCU)
+BTF_ID_FLAGS(func, scx_bpf_task_numa_nid, KF_RCU)
BTF_ID_FLAGS(func, scx_bpf_locked_rq, KF_IMPLICIT_ARGS | KF_RET_NULL)
BTF_ID_FLAGS(func, scx_bpf_cpu_curr, KF_IMPLICIT_ARGS | KF_RET_NULL | KF_RCU_PROTECTED)
BTF_ID_FLAGS(func, scx_bpf_cid_curr, KF_IMPLICIT_ARGS | KF_RET_NULL | KF_RCU_PROTECTED)
diff --git a/tools/sched_ext/include/scx/compat.bpf.h b/tools/sched_ext/include/scx/compat.bpf.h
index c26024a6e2b1b..64c2855d37875 100644
--- a/tools/sched_ext/include/scx/compat.bpf.h
+++ b/tools/sched_ext/include/scx/compat.bpf.h
@@ -478,6 +478,16 @@ static inline void scx_bpf_task_set_dsq_vtime(struct task_struct *p, u64 vtime)
p->scx.dsq_vtime = vtime;
}
+/* v7.4: Add scx_bpf_task_numa_nid(). */
+s32 scx_bpf_task_numa_nid___new(const struct task_struct *p) __ksym __weak;
+
+static inline s32 scx_bpf_task_numa_nid(const struct task_struct *p)
+{
+ if (bpf_ksym_exists(scx_bpf_task_numa_nid___new))
+ return scx_bpf_task_numa_nid___new(p);
+ return NUMA_NO_NODE;
+}
+
/* v7.4: Add scx_bpf_task_set_lazy_resched(). */
bool scx_bpf_task_set_lazy_resched___new(struct task_struct *p, bool lazy) __ksym __weak;
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH 5/5] selftests/sched_ext: Add a test for scx_bpf_task_numa_nid()
2026-10-04 7:27 [PATCHSET sched_ext/for-7.4] sched_ext: Add NUMA balancing support Andrea Righi
` (3 preceding siblings ...)
2026-10-04 7:27 ` [PATCH 4/5] sched_ext: Add scx_bpf_task_numa_nid() Andrea Righi
@ 2026-10-04 7:27 ` Andrea Righi
4 siblings, 0 replies; 6+ messages in thread
From: Andrea Righi @ 2026-10-04 7:27 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
Valentin Schneider, K Prateek Nayak, Vladimir Vdovin,
Emil Tsalapatis, Christian Loehle, Balbir Singh, Lee Trager,
sched-ext, linux-kernel
Add a scheduler which opts into NUMA hinting-fault scanning and reads
every running task's preferred node. Check that each result is either
NUMA_NO_NODE or a node reported by the kernel.
The test runs two threads touching memory for a few seconds, which on a
NUMA machine with numa_balancing enabled is enough for its tasks to be
scanned and acquire a preferred node. Two threads are needed because
the scan skips the pages of a single-threaded process that are already
on the node it runs on. Whether a node is acquired depends on the
machine, so the test only reports the counts it observed and does not
require one.
Assisted-by: LLM
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
tools/testing/selftests/sched_ext/Makefile | 1 +
.../selftests/sched_ext/numa_nid.bpf.c | 52 ++++++++
tools/testing/selftests/sched_ext/numa_nid.c | 111 ++++++++++++++++++
3 files changed, 164 insertions(+)
create mode 100644 tools/testing/selftests/sched_ext/numa_nid.bpf.c
create mode 100644 tools/testing/selftests/sched_ext/numa_nid.c
diff --git a/tools/testing/selftests/sched_ext/Makefile b/tools/testing/selftests/sched_ext/Makefile
index c5d3a2eaea7da..707ab511acb51 100644
--- a/tools/testing/selftests/sched_ext/Makefile
+++ b/tools/testing/selftests/sched_ext/Makefile
@@ -186,6 +186,7 @@ auto-test-targets := \
non_scx_kfunc_deny \
nohz_tick \
numa \
+ numa_nid \
allowed_cpus \
peek_dsq \
prog_run \
diff --git a/tools/testing/selftests/sched_ext/numa_nid.bpf.c b/tools/testing/selftests/sched_ext/numa_nid.bpf.c
new file mode 100644
index 0000000000000..54f7bdbc0a3d0
--- /dev/null
+++ b/tools/testing/selftests/sched_ext/numa_nid.bpf.c
@@ -0,0 +1,52 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A scheduler that reads the NUMA node a task's hinting faults point at.
+ *
+ * Every task that runs is asked for its preferred node. The answer is either
+ * NUMA_NO_NODE, which a task keeps until its address space has been scanned a
+ * few times, or a node id the kernel knows about. Anything else means the
+ * value reaching BPF is not the one the placement code maintains.
+ *
+ * Copyright (c) 2026 NVIDIA Corporation.
+ */
+
+#include <scx/common.bpf.h>
+
+char _license[] SEC("license") = "GPL";
+
+UEI_DEFINE(uei);
+
+/* Tasks seen with and without a preferred node. */
+u64 nr_with_nid;
+u64 nr_without_nid;
+
+void BPF_STRUCT_OPS(numa_nid_running, struct task_struct *p)
+{
+ s32 nid = scx_bpf_task_numa_nid(p);
+
+ if (nid == NUMA_NO_NODE) {
+ __sync_fetch_and_add(&nr_without_nid, 1);
+ return;
+ }
+
+ if (nid < 0 || nid >= scx_bpf_nr_node_ids()) {
+ scx_bpf_error("task %d reported node %d, kernel has %d node ids",
+ p->pid, nid, scx_bpf_nr_node_ids());
+ return;
+ }
+
+ __sync_fetch_and_add(&nr_with_nid, 1);
+}
+
+void BPF_STRUCT_OPS(numa_nid_exit, struct scx_exit_info *ei)
+{
+ UEI_RECORD(uei, ei);
+}
+
+SEC(".struct_ops.link")
+struct sched_ext_ops numa_nid_ops = {
+ .running = (void *)numa_nid_running,
+ .exit = (void *)numa_nid_exit,
+ .flags = SCX_OPS_NUMA_BALANCING,
+ .name = "numa_nid",
+};
diff --git a/tools/testing/selftests/sched_ext/numa_nid.c b/tools/testing/selftests/sched_ext/numa_nid.c
new file mode 100644
index 0000000000000..3aa1630ff6844
--- /dev/null
+++ b/tools/testing/selftests/sched_ext/numa_nid.c
@@ -0,0 +1,111 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA Corporation.
+ */
+#define _GNU_SOURCE
+#include <bpf/bpf.h>
+#include <pthread.h>
+#include <scx/common.h>
+#include <stdlib.h>
+#include <time.h>
+#include <unistd.h>
+#include "numa_nid.bpf.skel.h"
+#include "scx_test.h"
+
+#define WORKLOAD_SIZE (64 << 20)
+#define WORKLOAD_THREADS 2
+#define WORKLOAD_SECONDS 3
+
+/*
+ * Touch a private buffer for a while. The hinting-fault scan skips the pages
+ * of a single-threaded process that are already on the node it runs on, so
+ * one such thread alone would never take a hinting fault: run two of them.
+ */
+static void *touch_memory(void *arg)
+{
+ struct timespec start, now;
+ volatile char *buffer;
+ size_t offset;
+
+ buffer = malloc(WORKLOAD_SIZE);
+ if (!buffer)
+ return NULL;
+
+ clock_gettime(CLOCK_MONOTONIC, &start);
+ do {
+ for (offset = 0; offset < WORKLOAD_SIZE; offset += 4096)
+ buffer[offset]++;
+ clock_gettime(CLOCK_MONOTONIC, &now);
+ } while (now.tv_sec - start.tv_sec < WORKLOAD_SECONDS);
+
+ free((void *)buffer);
+ return NULL;
+}
+
+static enum scx_test_status setup(void **ctx)
+{
+ struct numa_nid *skel;
+
+ skel = numa_nid__open();
+ SCX_FAIL_IF(!skel, "Failed to open");
+ SCX_ENUM_INIT(skel);
+ SCX_FAIL_IF(numa_nid__load(skel), "Failed to load skel");
+
+ *ctx = skel;
+
+ return SCX_TEST_PASS;
+}
+
+static enum scx_test_status run(void *ctx)
+{
+ struct numa_nid *skel = ctx;
+ pthread_t threads[WORKLOAD_THREADS - 1];
+ struct bpf_link *link;
+ int i;
+
+ link = bpf_map__attach_struct_ops(skel->maps.numa_nid_ops);
+ SCX_FAIL_IF(!link, "Failed to attach scheduler");
+
+ /*
+ * Give the scan something to work on, so that the tasks of this test
+ * can acquire a preferred node on a NUMA machine with numa_balancing
+ * enabled. The test does not require one, as it depends on the
+ * machine: the scheduler reports an error through UEI if it ever sees
+ * a node id that is not valid and the counts below show what it saw.
+ */
+ for (i = 0; i < WORKLOAD_THREADS - 1; i++) {
+ if (pthread_create(&threads[i], NULL, touch_memory, NULL)) {
+ SCX_ERR("Failed to create a worker thread");
+ bpf_link__destroy(link);
+ return SCX_TEST_FAIL;
+ }
+ }
+ touch_memory(NULL);
+ for (i = 0; i < WORKLOAD_THREADS - 1; i++)
+ pthread_join(threads[i], NULL);
+
+ fprintf(stderr, "tasks with a preferred node: %lu, without: %lu\n",
+ (unsigned long)skel->bss->nr_with_nid,
+ (unsigned long)skel->bss->nr_without_nid);
+
+ bpf_link__destroy(link);
+ SCX_EQ(skel->data->uei.kind, EXIT_KIND(SCX_EXIT_UNREG));
+
+ return SCX_TEST_PASS;
+}
+
+static void cleanup(void *ctx)
+{
+ struct numa_nid *skel = ctx;
+
+ numa_nid__destroy(skel);
+}
+
+struct scx_test numa_nid = {
+ .name = "numa_nid",
+ .description = "Read a task's preferred NUMA node from BPF",
+ .setup = setup,
+ .run = run,
+ .cleanup = cleanup,
+};
+REGISTER_SCX_TEST(&numa_nid)
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread