From: Tejun Heo <tj@kernel.org>
To: sched-ext@lists.linux.dev
Cc: David Vernet <void@manifault.com>,
Andrea Righi <arighi@nvidia.com>,
Changwoo Min <changwoo@igalia.com>,
Emil Tsalapatis <emil@etsalapatis.com>,
David Dai <david.dai@linux.dev>,
linux-kernel@vger.kernel.org, Tejun Heo <tj@kernel.org>
Subject: [PATCH 2/2] sched_ext: Report a sub-scheduler's attach-time ecaps before lifting its enable bypass
Date: Thu, 8 Oct 2026 23:43:56 -1000 [thread overview]
Message-ID: <20261009094356.2852794-3-tj@kernel.org> (raw)
In-Reply-To: <20261009094356.2852794-1-tj@kernel.org>
A sub-scheduler learns which cids it holds through ops.sub_ecaps_updated().
Today, the grants its parent makes when it attaches are reported only after
its enable bypass is lifted, which is also when its tasks are handed to it.
So the sub-scheduler gets its first ops.enqueue() calls before it has been
notified of a single cid and has to hold the tasks somewhere until the
grants arrive.
Report the attach-time grants before lifting bypass instead. The
sub-scheduler then has its complete cid view when the first task arrives and
can place it right away. Nothing else reaches its ops in the meantime, as it
is still bypassing.
Signed-off-by: Tejun Heo <tj@kernel.org>
---
kernel/sched/ext/sub.c | 73 +++++++++++++++++++++++++++++++++++++-----
1 file changed, 65 insertions(+), 8 deletions(-)
diff --git a/kernel/sched/ext/sub.c b/kernel/sched/ext/sub.c
index 313e69841bc9..889fe322198e 100644
--- a/kernel/sched/ext/sub.c
+++ b/kernel/sched/ext/sub.c
@@ -1069,6 +1069,9 @@ static void discard_queued_syncs(struct rq *rq)
* @prev: @rq's previous task from the in-progress dispatch
* @report: whether a change is reported to the sched
*
+ * @prev may be NULL when seeding a bypassing sched. Only a nested
+ * scx_bpf_sub_dispatch() reads it, and a sched being enabled has no children.
+ *
* pshard->caps[] is the target configuration. pcpu->ecaps is the effective
* transposed copy owned by the cid's cpu and written only here under @rq's
* lock.
@@ -1107,13 +1110,21 @@ static u64 sync_pcpu_ecaps(struct rq *rq, struct scx_sched_pcpu *pcpu, s32 cid,
if (ecaps != pcpu->reported_ecaps &&
SCX_HAS_OP(pcpu->sch, sub_ecaps_updated) && report) {
struct scx_dsp_ctx *dspc = &this_cpu_ptr(pcpu->sch->pcpu)->dsp_ctx;
+ struct task_struct *prev_stash;
dspc->rq = rq;
- /* stash @prev so nested dispatches can access it */
+ /*
+ * Stash @prev for nested dispatches. The enable path calls this
+ * while the cpu's own dispatch may be inside an op with the rq
+ * lock dropped. That dispatch needs its stash back when it
+ * resumes, so save and restore it. A bypassing sched's op never
+ * drops the lock, so nothing else writes the stash in between.
+ */
+ prev_stash = rq->scx.sub_dispatch_prev;
rq->scx.sub_dispatch_prev = prev;
SCX_CALL_OP(pcpu->sch, sub_ecaps_updated, rq, scx_cpu_arg(cpu_of(rq)),
pcpu->reported_ecaps, ecaps);
- rq->scx.sub_dispatch_prev = NULL;
+ rq->scx.sub_dispatch_prev = prev_stash;
scx_flush_dispatch_buf(pcpu->sch, rq);
pcpu->reported_ecaps = ecaps;
}
@@ -1184,6 +1195,48 @@ void __scx_process_sync_ecaps(struct rq *rq, struct task_struct *prev)
scx_schedule_reenq_local(rq, SCX_REENQ_CAP_REVOKE);
}
+/**
+ * scx_sub_seed_ecaps - Report @sch's attach-time ecaps before lifting bypass
+ * @sch: sub-scheduler being enabled, still bypassing
+ *
+ * The grants made during ops.sub_attach() were applied without calling
+ * ops.sub_ecaps_updated(), as @sch had no ops registered yet or was already
+ * bypassing. Report them before lifting bypass, so that @sch is notified of its
+ * cids before any task reaches its ops. A sync still queued on a cpu finds
+ * nothing new afterwards. Kicks and inserts from the op are dropped while
+ * bypassing. Lifting bypass reschedules all cpus.
+ */
+static void scx_sub_seed_ecaps(struct scx_sched *sch)
+{
+ s32 cpu;
+
+ for_each_possible_cpu(cpu) {
+ struct rq *rq = cpu_rq(cpu);
+ s32 cid, shard;
+
+ /* unpinned, the lock state the op is called in from dispatch */
+ guard(raw_spin_rq_lock_irqsave)(rq);
+
+ /*
+ * When a cpu goes offline, its ecaps are cleared and must stay
+ * zero until it comes back online. Don't write non-zero ecaps
+ * on an inactive cpu.
+ */
+ if (!cpu_active(cpu))
+ continue;
+ cid = __scx_cpu_to_cid(cpu);
+ shard = rcu_dereference_all(scx_cid_to_shard)[cid];
+
+ /*
+ * An aborting sched gets no op calls.
+ * __scx_process_sync_ecaps() gets that from its bypass test,
+ * but @sch is bypassing here, so test aborting directly.
+ */
+ sync_pcpu_ecaps(rq, per_cpu_ptr(sch->pcpu, cpu), cid, shard, NULL,
+ !READ_ONCE(sch->aborting));
+ }
+}
+
/**
* scx_unbypass_replay_ecaps - Replay a bypass-suppressed ecaps notification
* @rq: rq of the cpu leaving bypass
@@ -1192,9 +1245,8 @@ void __scx_process_sync_ecaps(struct rq *rq, struct task_struct *prev)
* scx_process_sync_ecaps() consumes syncs while bypassing without delivering
* ops.sub_ecaps_updated(), leaving reported_ecaps stale. Nothing re-queues a
* sync when bypass lifts, so without a replay a cid that never changes again
- * would never be notified. The attach-time initial grants are the acute case
- * as they are consumed during the enable bypass window. Re-queue a sync for
- * any undelivered delta so the next dispatch delivers it.
+ * would never be notified. Re-queue a sync for any undelivered delta so the
+ * next dispatch delivers it.
*/
void scx_unbypass_replay_ecaps(struct rq *rq, struct scx_sched *sch)
{
@@ -2145,10 +2197,15 @@ void scx_sub_enable_workfn(struct kthread_work *work)
scx_cgroup_unlock();
percpu_up_write(&scx_fork_rwsem);
- scx_bypass(sch, false);
-
- /* @sch is enabled; deliver any caps owed since its sub_attach() */
+ /*
+ * @sch is enabled but still bypassing. Report the caps and ecaps
+ * granted during ops.sub_attach() before lifting bypass and letting its
+ * tasks reach ops.enqueue().
+ */
scx_sub_seed_caps(sch);
+ scx_sub_seed_ecaps(sch);
+
+ scx_bypass(sch, false);
pr_info("sched_ext: BPF sub-scheduler \"%s\" enabled\n", sch->ops.name);
kobject_uevent(&sch->kobj, KOBJ_ADD);
--
2.55.0
prev parent reply other threads:[~2026-10-09 9:43 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 9:43 [PATCHSET sched_ext/for-7.4] " Tejun Heo
2026-10-09 9:43 ` [PATCH 1/2] sched_ext: Factor out sync_pcpu_ecaps() Tejun Heo
2026-10-09 9:43 ` Tejun Heo [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261009094356.2852794-3-tj@kernel.org \
--to=tj@kernel.org \
--cc=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=david.dai@linux.dev \
--cc=emil@etsalapatis.com \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®