From: Tao Cui <cui.tao@linux.dev>
To: Tejun Heo <tj@kernel.org>
Cc: cui.tao@linux.dev, David Vernet <void@manifault.com>,
sched-ext@lists.linux.dev, Andrea Righi <arighi@nvidia.com>,
linux-kernel@vger.kernel.org, Changwoo Min <changwoo@igalia.com>
Subject: [REPORT] sched_ext: sub-scheduler tasks are severely under-scheduled vs root-owned tasks
Date: Thu, 8 Oct 2026 16:36:49 +0800 [thread overview]
Message-ID: <163e2397-0e3e-4cc0-842b-0d36598d13d7@linux.dev> (raw)
Hello Tejun,
During recent testing of the sub-scheduler attach flow, I hit what
looks like a dispatch starvation for claimed tasks and want to check
whether it's a known limitation of the sub-sched dispatch model before
digging further.
Setup:
A kvm guest (4 vCPUs, HZ=1000) running a sched_ext/for-7.4 based tree
with the caps-clear patch applied. A pair of minimal cid-form test
schedulers: a parent that grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED over
a cgroup to a child, and the child writing cpuperf target 1 from
ops.tick(). A busy-loop task (`taskset 1 sh -c 'while :; do :; done'`)
is placed in the child's cgroup (pinned to cpu0), and
rq->scx.cpuperf_target plus per-task CPU time are sampled every second.
Two attach orders are tested:
(a) the busy-loop task is moved into the cgroup before the child
attaches;
(b) the task is moved in after the child attaches.
Observation:
Order (a), 8 rounds: the child claims the task in every round
(verified: the enable walk's handover fires, p->scx.sched == child),
but the task then runs at ~0.3% duty cycle - it accumulates 0 or
~20-27 ticks per 10s window (one SCX_SLICE_DFL slice) while an
identical task owned by the parent runs at ~100% (1000 ticks/s).
Order (b), 8 rounds: the child claims the task via the migration path
and the task runs normally (no starvation observed in 16 rounds across
two parent-load conditions).
Based on the current test data, the disparity is ~300x: root-owned
tasks run at full duty while child-owned tasks get one slice and then
starve. The starvation is silent: attach succeeds, the grant succeeds,
the task's cgroup membership is correct - there is no error, warning,
or counter that indicates anything is wrong.
Probes and exclusions:
- ops.dispatch() for the child's cid is never invoked: a counter
bumped in ops.dispatch() when cid == 0 stays 0 across entire runs,
including rounds where the task is starved.
- ops.tick() for the child fires only when the task actually runs
(the tick count matches the starved duty cycle).
- The parent's ops.dispatch() delegation via
scx_bpf_sub_dispatch(cgroup_id) does not change the behavior.
- An explicit scx_bpf_kick_cid() on the task's cid in ops.enqueue()
does not change the behavior.
- The enable walk's claim is verified: pass 1 (init) and pass 2
(handover) both process the task, and a post-handover ownership
check (scx_task_on_sched()) passes.
- Adding a root-owned busy task on another CPU raises the failure rate
from ~25% to ~100%, which suggests the claimed task only gets to run
when some unrelated event runs balance on its CPU.
This is independent of the caps-clear patch (it reproduces with and
without it). The root cause is likely in the sub-sched attach->dispatch
path itself, so any sub-scheduler deployment would be affected.
Questions:
1. Is this a known gap in the sub-sched dispatch model - claimed tasks
on idle CPUs waiting for a delegation that only happens when the
parent's dispatch runs for that CPU?
2. Is the intended pattern for parents to kick CPUs on the child's
behalf (scx_bpf_kick_cid() after grant/claim), or should the core
route the balance to the child?
The repro is deterministic (8/8 rounds with the starvation in order
(a) under root-side load).
Thanks.
--
Tao
next reply other threads:[~2026-10-08 8:37 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 8:36 Tao Cui [this message]
2026-10-08 9:10 ` Tejun Heo
2026-10-09 1:51 ` Tao Cui
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=163e2397-0e3e-4cc0-842b-0d36598d13d7@linux.dev \
--to=cui.tao@linux.dev \
--cc=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®