* [PATCHSET sched_ext/for-7.4] sched_ext: Add scx_bpf_cgroup_nr_cpus()
@ 2026-09-29 8:37 Andrea Righi
2026-09-29 8:37 ` [PATCH 1/3] cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus() Andrea Righi
` (2 more replies)
0 siblings, 3 replies; 5+ messages in thread
From: Andrea Righi @ 2026-09-29 8:37 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Waiman Long, Ridong Chen, Johannes Weiner, Michal Koutny,
sched-ext, cgroups, linux-kernel
Hierarchical BPF schedulers need to know how many CPUs a cgroup can run on when
distributing group-wide weights. For example, fair's group share calculation
scales a group's weight by the smaller of its estimated runnable task count and
the number of CPUs in its effective cpuset (tg_cpus() -> cpuset_num_cpus()).
Without that count, a BPF scheduler can only use the machine-wide CPU count and
may give a group confined to a few CPUs excessive shares.
This patchset adds scx_bpf_cgroup_nr_cpus(), a kfunc that returns the same count
fair uses. This allows scx_eevdf to implement proper group scheduling, computing
per-cgroup shares the same way fair does, including for groups confined to a
subset of CPUs through cpuset.
Andrea Righi (3):
cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus()
sched_ext: Introduce scx_bpf_cgroup_nr_cpus()
selftests/sched_ext: Test scx_bpf_cgroup_nr_cpus()
kernel/cgroup/cpuset.c | 6 +-
kernel/sched/ext/ext.c | 22 ++
tools/sched_ext/include/scx/common.bpf.h | 1 +
tools/sched_ext/include/scx/compat.bpf.h | 7 +
tools/testing/selftests/sched_ext/Makefile | 6 +-
.../selftests/sched_ext/cgroup_nr_cpus.bpf.c | 60 +++
tools/testing/selftests/sched_ext/cgroup_nr_cpus.c | 419 +++++++++++++++++++++
tools/testing/selftests/sched_ext/config | 1 +
8 files changed, 519 insertions(+), 3 deletions(-)
create mode 100644 tools/testing/selftests/sched_ext/cgroup_nr_cpus.bpf.c
create mode 100644 tools/testing/selftests/sched_ext/cgroup_nr_cpus.c
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 1/3] cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus()
2026-09-29 8:37 [PATCHSET sched_ext/for-7.4] sched_ext: Add scx_bpf_cgroup_nr_cpus() Andrea Righi
@ 2026-09-29 8:37 ` Andrea Righi
2026-09-29 13:47 ` Michal Koutný
2026-09-29 8:37 ` [PATCH 2/3] sched_ext: Introduce scx_bpf_cgroup_nr_cpus() Andrea Righi
2026-09-29 8:37 ` [PATCH 3/3] selftests/sched_ext: Test scx_bpf_cgroup_nr_cpus() Andrea Righi
2 siblings, 1 reply; 5+ messages in thread
From: Andrea Righi @ 2026-09-29 8:37 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Waiman Long, Ridong Chen, Johannes Weiner, Michal Koutny,
sched-ext, cgroups, linux-kernel
cpuset_num_cpus() enters its RCU read-side section only after checking
is_in_v2_mode(). When cpuset is bound to a v1 hierarchy, is_in_v2_mode()
dereferences cpuset_cgrp_subsys.root, which is freed via kfree_rcu()
once that hierarchy is destroyed and cpuset is rebound to the default
hierarchy. A preemptible caller outside RCU can therefore read the flags
of a freed root.
The only current caller, fair's group share calculation, runs under the
rq lock with preemption disabled, so it can't hit this. However, the
helper already means to protect itself with RCU, and upcoming sched_ext
support exposes it to sleepable BPF programs.
Take the RCU read lock before is_in_v2_mode() so that the whole lookup
is protected regardless of the caller's context.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/cgroup/cpuset.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index d100634fa12b7..03fa9472c0893 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -4292,8 +4292,12 @@ int cpuset_num_cpus(struct cgroup *cgrp)
int nr = num_online_cpus();
struct cpuset *cs;
+ /*
+ * is_in_v2_mode() dereferences cpuset's hierarchy root, which can be a
+ * v1 root freed via kfree_rcu() on unmount.
+ */
+ guard(rcu)();
if (is_in_v2_mode()) {
- guard(rcu)();
cs = css_cs(cgroup_e_css(cgrp, &cpuset_cgrp_subsys));
if (cs)
nr = cpumask_weight(cs->effective_cpus);
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 2/3] sched_ext: Introduce scx_bpf_cgroup_nr_cpus()
2026-09-29 8:37 [PATCHSET sched_ext/for-7.4] sched_ext: Add scx_bpf_cgroup_nr_cpus() Andrea Righi
2026-09-29 8:37 ` [PATCH 1/3] cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus() Andrea Righi
@ 2026-09-29 8:37 ` Andrea Righi
2026-09-29 8:37 ` [PATCH 3/3] selftests/sched_ext: Test scx_bpf_cgroup_nr_cpus() Andrea Righi
2 siblings, 0 replies; 5+ messages in thread
From: Andrea Righi @ 2026-09-29 8:37 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Waiman Long, Ridong Chen, Johannes Weiner, Michal Koutny,
sched-ext, cgroups, linux-kernel
Hierarchical BPF schedulers may need to know how many CPUs a cgroup can
run on when distributing group-wide weights.
For example, fair's group share calculation scales a group's weight by
the smaller of its estimated runnable task count and the number of CPUs
in its effective cpuset. Without that count, a BPF scheduler can only
use the machine-wide CPU count and may give a group confined to a few
CPUs excessive shares.
A task's effective affinity can include per-task restrictions, so it
cannot reliably represent the group's cpuset. Reading the cpuset
directly from BPF would require CO-RE access to private cgroup and
cpuset structures, including inherited cpuset resolution, making the
approach fragile across kernel versions.
Add scx_bpf_cgroup_nr_cpus(), which returns the same count fair uses
through cpuset_num_cpus(). The count is independent of CPU or CID
numbering, so the kfunc is available to both cpu-form and cid-form
schedulers from any context. Also add the a compat wrapper which falls
back to nr_cpu_ids on older kernels.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/sched/ext/ext.c | 22 ++++++++++++++++++++++
tools/sched_ext/include/scx/common.bpf.h | 1 +
tools/sched_ext/include/scx/compat.bpf.h | 7 +++++++
3 files changed, 30 insertions(+)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index e48aac45afeaf..bb668a1749f55 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -11350,6 +11350,27 @@ __bpf_kfunc struct cgroup *scx_bpf_task_cgroup(struct task_struct *p,
cgroup_get(cgrp);
return cgrp;
}
+
+/**
+ * scx_bpf_cgroup_nr_cpus - Return the number of CPUs in a cgroup's cpuset
+ * @cgrp: cgroup of interest
+ *
+ * Return the number of CPUs in @cgrp's effective cpuset, which is inherited
+ * from the nearest ancestor with the cpuset controller enabled. This matches
+ * the count used by fair's group share calculation and lets hierarchical BPF
+ * schedulers bound a group's weight by the CPUs it can actually run on.
+ *
+ * On cgroup v1 without the cpuset_v2_mode mount option, or when cpuset is not
+ * configured, the number of online CPUs is returned. The result may be 0 for a
+ * partition root which has no effective CPUs left.
+ *
+ * The value is a snapshot and may change at any time through cpuset updates or
+ * CPU hotplug. Schedulers should re-read it when recomputing group state.
+ */
+__bpf_kfunc u32 scx_bpf_cgroup_nr_cpus(struct cgroup *cgrp)
+{
+ return cpuset_num_cpus(cgrp);
+}
#endif /* CONFIG_CGROUP_SCHED */
__bpf_kfunc_end_defs();
@@ -11397,6 +11418,7 @@ BTF_ID_FLAGS(func, scx_bpf_now)
BTF_ID_FLAGS(func, scx_bpf_events, KF_IMPLICIT_ARGS)
#ifdef CONFIG_CGROUP_SCHED
BTF_ID_FLAGS(func, scx_bpf_task_cgroup, KF_IMPLICIT_ARGS | KF_RCU | KF_ACQUIRE)
+BTF_ID_FLAGS(func, scx_bpf_cgroup_nr_cpus, KF_RCU)
#endif
BTF_ID_FLAGS(func, scx_bpf_sub_grant, KF_IMPLICIT_ARGS)
BTF_ID_FLAGS(func, scx_bpf_sub_revoke, KF_IMPLICIT_ARGS)
diff --git a/tools/sched_ext/include/scx/common.bpf.h b/tools/sched_ext/include/scx/common.bpf.h
index 5003e33f51423..28337ae8d9e66 100644
--- a/tools/sched_ext/include/scx/common.bpf.h
+++ b/tools/sched_ext/include/scx/common.bpf.h
@@ -104,6 +104,7 @@ struct task_struct *scx_bpf_cpu_curr(s32 cpu) __ksym __weak;
struct task_struct *scx_bpf_tid_to_task(u64 tid) __ksym __weak;
u64 scx_bpf_now(void) __ksym __weak;
void scx_bpf_events(struct scx_event_stats *events, size_t events__sz) __ksym __weak;
+u32 scx_bpf_cgroup_nr_cpus(struct cgroup *cgrp) __ksym __weak;
s32 scx_bpf_cpu_to_cid(s32 cpu) __ksym __weak;
s32 scx_bpf_cid_to_cpu(s32 cid) __ksym __weak;
s32 scx_bpf_cid_node(s32 cid) __ksym __weak;
diff --git a/tools/sched_ext/include/scx/compat.bpf.h b/tools/sched_ext/include/scx/compat.bpf.h
index c5c90d1d81e04..a793c5e8814b9 100644
--- a/tools/sched_ext/include/scx/compat.bpf.h
+++ b/tools/sched_ext/include/scx/compat.bpf.h
@@ -248,6 +248,13 @@ static inline bool __COMPAT_is_enq_cpu_selected(u64 enq_flags)
(bpf_ksym_exists(scx_bpf_cid_node) ? \
scx_bpf_cid_node(cid) : NUMA_NO_NODE)
+/*
+ * v7.4: Add scx_bpf_cgroup_nr_cpus().
+ */
+#define __COMPAT_scx_bpf_cgroup_nr_cpus(cgrp) \
+ (bpf_ksym_exists(scx_bpf_cgroup_nr_cpus) ? \
+ scx_bpf_cgroup_nr_cpus(cgrp) : scx_bpf_nr_cpu_ids())
+
/*
* v6.18: Add a helper to retrieve the current task running on a CPU.
*
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 3/3] selftests/sched_ext: Test scx_bpf_cgroup_nr_cpus()
2026-09-29 8:37 [PATCHSET sched_ext/for-7.4] sched_ext: Add scx_bpf_cgroup_nr_cpus() Andrea Righi
2026-09-29 8:37 ` [PATCH 1/3] cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus() Andrea Righi
2026-09-29 8:37 ` [PATCH 2/3] sched_ext: Introduce scx_bpf_cgroup_nr_cpus() Andrea Righi
@ 2026-09-29 8:37 ` Andrea Righi
2 siblings, 0 replies; 5+ messages in thread
From: Andrea Righi @ 2026-09-29 8:37 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: Waiman Long, Ridong Chen, Johannes Weiner, Michal Koutny,
sched-ext, cgroups, linux-kernel
Verify that scx_bpf_cgroup_nr_cpus() reports the number of CPUs in a
cgroup's effective cpuset.
Create a cgroup-v2 subtree where a child owns a cpuset and a leaf
inherits it. Sample the kfunc from a BPF_PROG_TYPE_SYSCALL program and
compare it with cpuset.cpus.effective for the root, the cpuset owner and
the inheriting leaf, without a scheduler attached. Attach a scheduler
and verify the value observed from ops.cgroup_init(). Then restrict the
child's cpuset to a non-contiguous and a single-CPU set, and finally
disable cpuset below the parent so that both cgroups inherit the
parent's cpuset.
A disabled cpuset css stays attached to its cgroup until it is
asynchronously offlined, so poll briefly for the last step. Skip the
test when the required controllers or two CPUs are unavailable.
Assisted-by: LLM
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
tools/testing/selftests/sched_ext/Makefile | 6 +-
.../selftests/sched_ext/cgroup_nr_cpus.bpf.c | 60 +++
.../selftests/sched_ext/cgroup_nr_cpus.c | 419 ++++++++++++++++++
tools/testing/selftests/sched_ext/config | 1 +
4 files changed, 484 insertions(+), 2 deletions(-)
create mode 100644 tools/testing/selftests/sched_ext/cgroup_nr_cpus.bpf.c
create mode 100644 tools/testing/selftests/sched_ext/cgroup_nr_cpus.c
diff --git a/tools/testing/selftests/sched_ext/Makefile b/tools/testing/selftests/sched_ext/Makefile
index 286b510fd76e1..ff59c95b9b386 100644
--- a/tools/testing/selftests/sched_ext/Makefile
+++ b/tools/testing/selftests/sched_ext/Makefile
@@ -10,6 +10,7 @@ TEST_GEN_MODS_DIR := test_modules
# override lib.mk's default rules
OVERRIDE_TARGETS := 1
include ../lib.mk
+include ../cgroup/lib/libcgroup.mk
CURDIR := $(abspath .)
REPOROOT := $(abspath ../../../..)
@@ -154,7 +155,7 @@ $(INCLUDE_DIR)/%.bpf.skel.h: $(SCXOBJ_DIR)/%.bpf.o $(INCLUDE_DIR)/vmlinux.h $(BP
override define CLEAN
rm -rf $(OUTPUT_DIR)
- rm -f $(TEST_GEN_PROGS)
+ rm -f $(TEST_GEN_PROGS) $(EXTRA_CLEAN)
endef
# Every testcase takes all of the BPF progs are dependencies by default. This
@@ -163,6 +164,7 @@ endef
all_test_bpfprogs := $(foreach prog,$(wildcard *.bpf.c),$(INCLUDE_DIR)/$(patsubst %.c,%.skel.h,$(prog)))
auto-test-targets := \
+ cgroup_nr_cpus \
create_dsq \
dequeue \
dequeue_iter \
@@ -216,7 +218,7 @@ $(testcase-targets): $(SCXOBJ_DIR)/%.o: %.c $(SCXOBJ_DIR)/runner.o $(all_test_bp
$(SCXOBJ_DIR)/util.o: util.c | $(SCXOBJ_DIR)
$(CC) $(CFLAGS) -c $< -o $@
-$(OUTPUT)/runner: $(SCXOBJ_DIR)/runner.o $(SCXOBJ_DIR)/util.o $(BPFOBJ) $(testcase-targets)
+$(OUTPUT)/runner: $(SCXOBJ_DIR)/runner.o $(SCXOBJ_DIR)/util.o $(BPFOBJ) $(LIBCGROUP_O) $(testcase-targets)
@echo "$(testcase-targets)"
$(CC) $(CFLAGS) -o $@ $^ $(LDFLAGS)
diff --git a/tools/testing/selftests/sched_ext/cgroup_nr_cpus.bpf.c b/tools/testing/selftests/sched_ext/cgroup_nr_cpus.bpf.c
new file mode 100644
index 0000000000000..841c83abb41bb
--- /dev/null
+++ b/tools/testing/selftests/sched_ext/cgroup_nr_cpus.bpf.c
@@ -0,0 +1,60 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Validate scx_bpf_cgroup_nr_cpus() from both BPF_PROG_TYPE_SYSCALL and
+ * struct_ops contexts.
+ *
+ * Copyright (c) 2026 NVIDIA Corporation.
+ */
+
+#include <scx/common.bpf.h>
+
+char _license[] SEC("license") = "GPL";
+
+UEI_DEFINE(uei);
+
+/* input to cgroup_nr_cpus_read() */
+u64 query_cgid;
+/* output of cgroup_nr_cpus_read(), -1 if @query_cgid couldn't be resolved */
+s64 query_nr_cpus = -1;
+
+/* recorded by ops.cgroup_init() for @init_cgid */
+u64 init_cgid;
+s64 init_nr_cpus = -1;
+
+SEC("syscall")
+int cgroup_nr_cpus_read(void *ctx)
+{
+ struct cgroup *cgrp;
+
+ query_nr_cpus = -1;
+
+ cgrp = bpf_cgroup_from_id(query_cgid);
+ if (!cgrp)
+ return -ENOENT;
+
+ query_nr_cpus = scx_bpf_cgroup_nr_cpus(cgrp);
+ bpf_cgroup_release(cgrp);
+
+ return 0;
+}
+
+s32 BPF_STRUCT_OPS(cgroup_nr_cpus_cgroup_init, struct cgroup *cgrp,
+ struct scx_cgroup_init_args *args)
+{
+ if (cgrp->kn->id == init_cgid)
+ init_nr_cpus = scx_bpf_cgroup_nr_cpus(cgrp);
+
+ return 0;
+}
+
+void BPF_STRUCT_OPS(cgroup_nr_cpus_exit, struct scx_exit_info *ei)
+{
+ UEI_RECORD(uei, ei);
+}
+
+SEC(".struct_ops.link")
+struct sched_ext_ops cgroup_nr_cpus_ops = {
+ .cgroup_init = (void *)cgroup_nr_cpus_cgroup_init,
+ .exit = (void *)cgroup_nr_cpus_exit,
+ .name = "cgroup_nr_cpus",
+};
diff --git a/tools/testing/selftests/sched_ext/cgroup_nr_cpus.c b/tools/testing/selftests/sched_ext/cgroup_nr_cpus.c
new file mode 100644
index 0000000000000..4973f70803532
--- /dev/null
+++ b/tools/testing/selftests/sched_ext/cgroup_nr_cpus.c
@@ -0,0 +1,419 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Verify that scx_bpf_cgroup_nr_cpus() reports the number of CPUs in a
+ * cgroup's effective cpuset, including inherited and updated cpusets.
+ *
+ * Copyright (c) 2026 NVIDIA Corporation.
+ */
+
+#define _GNU_SOURCE
+#include <bpf/bpf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <limits.h>
+#include <linux/limits.h>
+#include <scx/common.h>
+#include <stdbool.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+
+#include "cgroup_nr_cpus.bpf.skel.h"
+#include "cgroup_util.h"
+#include "scx_test.h"
+
+/*
+ * Hierarchy under the cgroup2 root, all with the cpu controller enabled so
+ * that ops.cgroup_init() runs for each of them:
+ *
+ * parent cpuset enabled by the root, enables cpuset for its children
+ * parent/child owns a cpuset
+ * parent/child/leaf no cpuset of its own, inherits child's
+ */
+struct cgroup_nr_cpus_ctx {
+ struct cgroup_nr_cpus *skel;
+ struct bpf_link *link;
+ char root[PATH_MAX];
+ char parent[PATH_MAX];
+ char child[PATH_MAX];
+ char leaf[PATH_MAX];
+ bool parent_created;
+ bool child_created;
+ bool leaf_created;
+};
+
+static int join_path(char *dst, size_t dst_size, const char *parent, const char *name)
+{
+ int ret;
+
+ ret = snprintf(dst, dst_size, "%s/%s", parent, name);
+ if (ret < 0 || (size_t)ret >= dst_size)
+ return -ENAMETOOLONG;
+ return 0;
+}
+
+static u64 cgroup_id(const char *path)
+{
+ union {
+ u64 id;
+ unsigned char bytes[8];
+ } id = {};
+ struct file_handle *handle;
+ int mount_id, ret;
+
+ handle = calloc(1, sizeof(*handle) + sizeof(id));
+ if (!handle)
+ return 0;
+ handle->handle_bytes = sizeof(id);
+ ret = name_to_handle_at(AT_FDCWD, path, handle, &mount_id, 0);
+ if (!ret && handle->handle_bytes == sizeof(id))
+ memcpy(id.bytes, handle->f_handle, sizeof(id));
+ free(handle);
+
+ return ret ? 0 : id.id;
+}
+
+/*
+ * Parse a cpulist such as "0-3,8,10-11". Return the number of CPUs and the
+ * lowest and highest CPU in @first and @last, or -errno on failure.
+ */
+static int parse_cpulist(const char *cpulist, u32 *first, u32 *last)
+{
+ const char *p = cpulist;
+ u32 lowest = UINT_MAX, highest = 0;
+ int count = 0;
+
+ while (*p && *p != '\n') {
+ unsigned long start, end_cpu;
+ char *end;
+
+ errno = 0;
+ start = strtoul(p, &end, 10);
+ if (errno || end == p || start > INT_MAX)
+ return -EINVAL;
+ end_cpu = start;
+ p = end;
+ if (*p == '-') {
+ end_cpu = strtoul(p + 1, &end, 10);
+ if (errno || end == p + 1 || end_cpu > INT_MAX || end_cpu < start)
+ return -EINVAL;
+ p = end;
+ }
+ if (end_cpu - start + 1 > (unsigned long)(INT_MAX - count))
+ return -EOVERFLOW;
+ if (start < lowest)
+ lowest = start;
+ if (end_cpu > highest)
+ highest = end_cpu;
+ count += end_cpu - start + 1;
+ if (*p == ',')
+ p++;
+ else if (*p && *p != '\n')
+ return -EINVAL;
+ }
+
+ if (first)
+ *first = lowest;
+ if (last)
+ *last = highest;
+ return count;
+}
+
+/*
+ * Number of CPUs in @cgroup's cpuset.cpus.effective, or -errno.
+ *
+ * A sparse cpulist can exceed a page on large systems. cg_read() does a single
+ * bounded read, so size the buffer for the worst case of NR_CPUS=8192 and
+ * reject a read that fills it rather than parsing a truncated list.
+ */
+static int effective_nr_cpus(const char *cgroup, u32 *first, u32 *last)
+{
+ static char buf[65536];
+
+ if (cg_read(cgroup, "cpuset.cpus.effective", buf, sizeof(buf)))
+ return -EIO;
+ if (strlen(buf) >= sizeof(buf) - 1)
+ return -EOVERFLOW;
+ return parse_cpulist(buf, first, last);
+}
+
+/* Run the SYSCALL program to sample scx_bpf_cgroup_nr_cpus() for @path. */
+static int kfunc_nr_cpus(struct cgroup_nr_cpus_ctx *ctx, const char *path, s64 *nr_cpus)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, topts);
+ u64 cgid;
+ int err;
+
+ cgid = cgroup_id(path);
+ if (!cgid) {
+ SCX_ERR("Failed to read cgroup ID of %s", path);
+ return -ENOENT;
+ }
+
+ ctx->skel->bss->query_cgid = cgid;
+ err = bpf_prog_test_run_opts(bpf_program__fd(ctx->skel->progs.cgroup_nr_cpus_read),
+ &topts);
+ if (err || topts.retval) {
+ SCX_ERR("BPF_PROG_RUN failed for %s (err=%d retval=%d)",
+ path, err, (int)topts.retval);
+ return err ?: -EIO;
+ }
+
+ *nr_cpus = ctx->skel->data->query_nr_cpus;
+ return 0;
+}
+
+static bool check_nr_cpus(struct cgroup_nr_cpus_ctx *ctx, const char *path, int expected,
+ const char *what)
+{
+ s64 nr_cpus;
+
+ if (kfunc_nr_cpus(ctx, path, &nr_cpus))
+ return false;
+ if (nr_cpus != expected) {
+ SCX_ERR("%s: expected %d CPUs, got %lld", what, expected,
+ (long long)nr_cpus);
+ return false;
+ }
+ return true;
+}
+
+/*
+ * Like check_nr_cpus() but tolerate a transient mismatch. A cpuset css being
+ * disabled stays attached to its cgroup until it's asynchronously offlined, and
+ * cpuset_num_cpus() keeps reporting its stale mask until then.
+ */
+static bool wait_nr_cpus(struct cgroup_nr_cpus_ctx *ctx, const char *path, int expected,
+ const char *what)
+{
+ s64 nr_cpus = -1;
+ int i;
+
+ for (i = 0; i < 1000; i++) {
+ if (kfunc_nr_cpus(ctx, path, &nr_cpus))
+ return false;
+ if (nr_cpus == expected)
+ return true;
+ usleep(1000);
+ }
+ SCX_ERR("%s: expected %d CPUs, got %lld", what, expected, (long long)nr_cpus);
+ return false;
+}
+
+static bool controller_enabled(const char *cgroup, const char *file, const char *controller)
+{
+ char buf[4096], *saveptr, *token;
+
+ if (cg_read(cgroup, file, buf, sizeof(buf)))
+ return false;
+ for (token = strtok_r(buf, "\n ", &saveptr); token;
+ token = strtok_r(NULL, "\n ", &saveptr))
+ if (!strcmp(token, controller))
+ return true;
+ return false;
+}
+
+static void cleanup_ctx(struct cgroup_nr_cpus_ctx *ctx)
+{
+ bpf_link__destroy(ctx->link);
+ cgroup_nr_cpus__destroy(ctx->skel);
+ if (ctx->leaf_created)
+ cg_destroy(ctx->leaf);
+ if (ctx->child_created)
+ cg_destroy(ctx->child);
+ if (ctx->parent_created)
+ cg_destroy(ctx->parent);
+}
+
+/*
+ * Enable @controller in the root's subtree_control if needed. Like the cgroup
+ * selftests, leave it enabled afterwards: the root is shared, and another
+ * manager may start relying on the controller while the test runs.
+ */
+static enum scx_test_status enable_controller(const char *root, const char *controller)
+{
+ char value[32];
+
+ if (controller_enabled(root, "cgroup.subtree_control", controller))
+ return SCX_TEST_PASS;
+ if (!controller_enabled(root, "cgroup.controllers", controller))
+ return SCX_TEST_SKIP;
+
+ snprintf(value, sizeof(value), "+%s", controller);
+ if (cg_write(root, "cgroup.subtree_control", value))
+ return SCX_TEST_SKIP;
+ return SCX_TEST_PASS;
+}
+
+static enum scx_test_status setup_cgroups(struct cgroup_nr_cpus_ctx *ctx)
+{
+ enum scx_test_status status;
+ char name[64];
+
+ if (cg_find_unified_root(ctx->root, sizeof(ctx->root), NULL))
+ return SCX_TEST_SKIP;
+
+ status = enable_controller(ctx->root, "cpu");
+ if (status != SCX_TEST_PASS)
+ return status;
+ status = enable_controller(ctx->root, "cpuset");
+ if (status != SCX_TEST_PASS)
+ return status;
+
+ snprintf(name, sizeof(name), "scx_nr_cpus_%d", getpid());
+ if (join_path(ctx->parent, sizeof(ctx->parent), ctx->root, name) ||
+ join_path(ctx->child, sizeof(ctx->child), ctx->parent, "child") ||
+ join_path(ctx->leaf, sizeof(ctx->leaf), ctx->child, "leaf")) {
+ SCX_ERR("Cgroup path is too long");
+ return SCX_TEST_FAIL;
+ }
+
+ if (cg_create(ctx->parent)) {
+ SCX_ERR("Failed to create cgroup %s", ctx->parent);
+ return SCX_TEST_FAIL;
+ }
+ ctx->parent_created = true;
+ if (cg_write(ctx->parent, "cgroup.subtree_control", "+cpu +cpuset")) {
+ SCX_ERR("Failed to enable controllers in %s", ctx->parent);
+ return SCX_TEST_FAIL;
+ }
+ if (cg_create(ctx->child)) {
+ SCX_ERR("Failed to create cgroup %s", ctx->child);
+ return SCX_TEST_FAIL;
+ }
+ ctx->child_created = true;
+ if (cg_write(ctx->child, "cgroup.subtree_control", "+cpu")) {
+ SCX_ERR("Failed to enable cpu in %s", ctx->child);
+ return SCX_TEST_FAIL;
+ }
+ if (cg_create(ctx->leaf)) {
+ SCX_ERR("Failed to create cgroup %s", ctx->leaf);
+ return SCX_TEST_FAIL;
+ }
+ ctx->leaf_created = true;
+
+ return SCX_TEST_PASS;
+}
+
+static enum scx_test_status run(void *arg)
+{
+ struct cgroup_nr_cpus_ctx ctx = {};
+ enum scx_test_status status;
+ char value[32];
+ u32 first, last;
+ int nr_root, nr_child;
+
+ (void)arg;
+
+ /*
+ * SCX_ENUM_INIT() exits the process if vmlinux BTF can't be loaded, so
+ * run it before creating any cgroups that would then be left behind.
+ */
+ ctx.skel = cgroup_nr_cpus__open();
+ if (!ctx.skel) {
+ SCX_ERR("Failed to open skel");
+ return SCX_TEST_FAIL;
+ }
+ SCX_ENUM_INIT(ctx.skel);
+
+ status = setup_cgroups(&ctx);
+ if (status != SCX_TEST_PASS)
+ goto out;
+ status = SCX_TEST_FAIL;
+
+ nr_root = effective_nr_cpus(ctx.root, NULL, NULL);
+ nr_child = effective_nr_cpus(ctx.child, &first, &last);
+ if (nr_root < 0 || nr_child < 0) {
+ SCX_ERR("Failed to read effective cpusets");
+ goto out;
+ }
+ /* The effective cpuset can be empty, e.g. under a partition root. */
+ if (nr_child < 2) {
+ status = SCX_TEST_SKIP;
+ goto out;
+ }
+
+ ctx.skel->bss->init_cgid = cgroup_id(ctx.leaf);
+ if (!ctx.skel->bss->init_cgid) {
+ SCX_ERR("Failed to read cgroup ID of %s", ctx.leaf);
+ goto out;
+ }
+ if (cgroup_nr_cpus__load(ctx.skel)) {
+ SCX_ERR("Failed to load skel");
+ goto out;
+ }
+
+ /* The kfunc must be callable without a scheduler attached. */
+ if (!check_nr_cpus(&ctx, ctx.root, nr_root, "root") ||
+ !check_nr_cpus(&ctx, ctx.child, nr_child, "child") ||
+ !check_nr_cpus(&ctx, ctx.leaf, nr_child, "inherited leaf"))
+ goto out;
+
+ /* ops.cgroup_init() runs for existing cgroups when attaching. */
+ ctx.link = bpf_map__attach_struct_ops(ctx.skel->maps.cgroup_nr_cpus_ops);
+ if (!ctx.link) {
+ SCX_ERR("Failed to attach scheduler");
+ goto out;
+ }
+ if (ctx.skel->data->init_nr_cpus != nr_child) {
+ SCX_ERR("ops.cgroup_init(): expected %d CPUs, got %lld", nr_child,
+ (long long)ctx.skel->data->init_nr_cpus);
+ goto out;
+ }
+
+ /* Non-contiguous cpuset, observed by the owner and by the inheritor. */
+ if (nr_child > 2) {
+ snprintf(value, sizeof(value), "%u,%u", first, last);
+ if (cg_write(ctx.child, "cpuset.cpus", value)) {
+ SCX_ERR("Failed to set cpuset.cpus=%s for %s", value, ctx.child);
+ goto out;
+ }
+ if (!check_nr_cpus(&ctx, ctx.child, 2, "sparse child") ||
+ !check_nr_cpus(&ctx, ctx.leaf, 2, "sparse inherited leaf"))
+ goto out;
+ }
+
+ snprintf(value, sizeof(value), "%u", first);
+ if (cg_write(ctx.child, "cpuset.cpus", value)) {
+ SCX_ERR("Failed to set cpuset.cpus=%s for %s", value, ctx.child);
+ goto out;
+ }
+ if (!check_nr_cpus(&ctx, ctx.child, 1, "single-CPU child") ||
+ !check_nr_cpus(&ctx, ctx.leaf, 1, "single-CPU inherited leaf"))
+ goto out;
+
+ /*
+ * Disabling cpuset below @parent makes @child and @leaf inherit
+ * @parent's effective cpuset, which spans all of the root's CPUs.
+ */
+ if (cg_write(ctx.parent, "cgroup.subtree_control", "-cpuset")) {
+ SCX_ERR("Failed to disable cpuset in %s", ctx.parent);
+ goto out;
+ }
+ nr_child = effective_nr_cpus(ctx.parent, NULL, NULL);
+ if (nr_child < 0) {
+ SCX_ERR("Failed to read effective cpuset of %s", ctx.parent);
+ goto out;
+ }
+ if (!wait_nr_cpus(&ctx, ctx.child, nr_child, "child after cpuset disable") ||
+ !wait_nr_cpus(&ctx, ctx.leaf, nr_child, "leaf after cpuset disable"))
+ goto out;
+
+ if (ctx.skel->data->uei.kind != EXIT_KIND(SCX_EXIT_NONE)) {
+ SCX_ERR("Scheduler exited unexpectedly");
+ goto out;
+ }
+
+ status = SCX_TEST_PASS;
+out:
+ cleanup_ctx(&ctx);
+ return status;
+}
+
+struct scx_test cgroup_nr_cpus = {
+ .name = "cgroup_nr_cpus",
+ .description = "Verify scx_bpf_cgroup_nr_cpus() reports effective cpuset CPU counts",
+ .run = run,
+};
+REGISTER_SCX_TEST(&cgroup_nr_cpus)
diff --git a/tools/testing/selftests/sched_ext/config b/tools/testing/selftests/sched_ext/config
index affa3cf33470a..8173f9170ffe8 100644
--- a/tools/testing/selftests/sched_ext/config
+++ b/tools/testing/selftests/sched_ext/config
@@ -1,6 +1,7 @@
CONFIG_SCHED_CLASS_EXT=y
CONFIG_CGROUPS=y
CONFIG_CGROUP_SCHED=y
+CONFIG_CPUSETS=y
CONFIG_EXT_GROUP_SCHED=y
CONFIG_BPF=y
CONFIG_BPF_SYSCALL=y
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH 1/3] cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus()
2026-09-29 8:37 ` [PATCH 1/3] cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus() Andrea Righi
@ 2026-09-29 13:47 ` Michal Koutný
0 siblings, 0 replies; 5+ messages in thread
From: Michal Koutný @ 2026-09-29 13:47 UTC (permalink / raw)
To: Andrea Righi
Cc: Tejun Heo, David Vernet, Changwoo Min, Waiman Long, Ridong Chen,
Johannes Weiner, sched-ext, cgroups, linux-kernel,
Peter Zijlstra
[-- Attachment #1: Type: text/plain, Size: 3061 bytes --]
Hi.
On Tue, Sep 29, 2026 at 10:37:38AM +0200, Andrea Righi <arighi@nvidia.com> wrote:
> cpuset_num_cpus() enters its RCU read-side section only after checking
> is_in_v2_mode(). When cpuset is bound to a v1 hierarchy, is_in_v2_mode()
> dereferences cpuset_cgrp_subsys.root, which is freed via kfree_rcu()
> once that hierarchy is destroyed and cpuset is rebound to the default
> hierarchy. A preemptible caller outside RCU can therefore read the flags
> of a freed root.
>
> The only current caller, fair's group share calculation, runs under the
> rq lock with preemption disabled, so it can't hit this. However, the
> helper already means to protect itself with RCU, and upcoming sched_ext
> support exposes it to sleepable BPF programs.
>
> Take the RCU read lock before is_in_v2_mode() so that the whole lookup
> is protected regardless of the caller's context.
This feels like mere querying of the mode shouldn't require such
constraints (despite it's needed anyway later down). But it could truly
happen with the novel usage (CONFIG_CPUSET_V1 && unmounting cpuset
hierarchy for some reason, I wonder how you noticed :)).
Then I'd welcome more structured approach with at least:
diff --git a/include/linux/cgroup-defs.h b/include/linux/cgroup-defs.h
index 3754d697854b3..7f4d346cfb119 100644
--- a/include/linux/cgroup-defs.h
+++ b/include/linux/cgroup-defs.h
@@ -841,7 +841,7 @@ struct cgroup_subsys {
const char *legacy_name;
/* link to parent, protected by cgroup_lock() */
- struct cgroup_root *root;
+ struct cgroup_root __rcu *root;
/* idr for css->id */
struct idr css_idr;
diff --git a/kernel/cgroup/cgroup.c b/kernel/cgroup/cgroup.c
index 227d09704ca59..a718b5f521fb2 100644
--- a/kernel/cgroup/cgroup.c
+++ b/kernel/cgroup/cgroup.c
@@ -1909,7 +1909,7 @@ int rebind_subsystems(struct cgroup_root *dst_root, u32 ss_mask)
/* rebind */
RCU_INIT_POINTER(scgrp->subsys[ssid], NULL);
rcu_assign_pointer(dcgrp->subsys[ssid], css);
- ss->root = dst_root;
+ rcu_assign_pointer(ss->root, dst_root);
spin_lock_irq(&css_set_lock);
css->cgroup = dcgrp;
However, if I zoom out, I see that the intention of reading cpuset's
nr_cpus from the scheduler is meant for setups where cpuset tree ~ cpu
tree:
| * This only really works for cgroup-v2 where all the controllers are mounted
| * in the same hierarchy. If not cgroup-v2 or no cpuset controller is
| * configured it reverts to num_online_cpus().
Hence it may be just OK to do:
int nr = num_online_cpus();
struct cpuset *cs;
- if (is_in_v2_mode()) {
+ if (cpuset_v2()) {
guard(rcu)();
cs = css_cs(cgroup_e_css(cgrp, &cpuset_cgrp_subsys));
if (cs)
I hope Waiman seconds this -- if a feature depends on shared tree,
there's only so much that 'cpuset_v2_mode' can guarantee.
0.02€,
Michal
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 265 bytes --]
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-29 13:48 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 8:37 [PATCHSET sched_ext/for-7.4] sched_ext: Add scx_bpf_cgroup_nr_cpus() Andrea Righi
2026-09-29 8:37 ` [PATCH 1/3] cgroup/cpuset: Protect is_in_v2_mode() in cpuset_num_cpus() Andrea Righi
2026-09-29 13:47 ` Michal Koutný
2026-09-29 8:37 ` [PATCH 2/3] sched_ext: Introduce scx_bpf_cgroup_nr_cpus() Andrea Righi
2026-09-29 8:37 ` [PATCH 3/3] selftests/sched_ext: Test scx_bpf_cgroup_nr_cpus() Andrea Righi
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®