From: Tejun Heo <tj@kernel.org>
To: David Vernet <void@manifault.com>,
Andrea Righi <arighi@nvidia.com>,
Changwoo Min <changwoo@igalia.com>
Cc: Emil Tsalapatis <emil@etsalapatis.com>,
David Dai <david.dai@linux.dev>,
Dan Schatzberg <schatzberg.dan@gmail.com>,
Xiangyu Bu <xbu@meta.com>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH sched_ext/for-7.3-fixes] sched_ext: Fix stall when a task enqueued with SCX_ENQ_LAST lands back on its CPU's local DSQ
Date: Wed, 07 Oct 2026 06:26:47 -1000 [thread overview]
Message-ID: <2b405553751fe811b4cd2c24f307d961@kernel.org> (raw)
With SCX_OPS_ENQ_LAST, the last runnable task on a CPU is passed to
ops.enqueue() instead of being kept running. Exiting tasks without
SCX_OPS_ENQ_EXITING, migration disabled tasks without
SCX_OPS_ENQ_MIGRATION_DISABLED and tasks on an offline rq skip ops.enqueue()
and go straight to the local DSQ. This happens in put_prev_task_scx() after
the pick has settled on idle, so nothing reschedules the CPU, which idles
with the task queued until an unrelated wakeup lands on it or the watchdog
fires.
scx_mitosis, which sets SCX_OPS_ENQ_LAST without SCX_OPS_ENQ_EXITING, hit
this during mass cgroup teardown: exiting tasks stuck in exit_mmap() for
over 40 seconds.
A task put on the local DSQ runs, as anywhere else. If the SCX_ENQ_LAST
enqueue left the task on this CPU's local DSQ, set need_resched on the task
being switched in, as proxy_resched_idle() does. resched_curr() would flag
the task being put instead, and __schedule() clears that right after the
pick.
Fixes: f0e1a0643a59 ("sched_ext: Implement BPF extensible scheduler class")
Cc: stable@vger.kernel.org # v6.12+
Reported-by: Xiangyu Bu <xbu@meta.com>
Link: https://github.com/sched-ext/scx/pull/3870
Signed-off-by: Tejun Heo <tj@kernel.org>
Cc: Dan Schatzberg <schatzberg.dan@gmail.com>
---
kernel/sched/ext/ext.c | 13 +++++++++++--
kernel/sched/ext/internal.h | 5 +++--
2 files changed, 14 insertions(+), 4 deletions(-)
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -3213,8 +3213,7 @@ static void put_prev_task_scx(struct rq
/*
* If @p is runnable but we're about to enter a lower
* sched_class, %SCX_OPS_ENQ_LAST must be set. Tell
- * ops.enqueue() that @p is the only one available for this cpu,
- * which should trigger an explicit follow-up scheduling event.
+ * ops.enqueue() that @p is the only one available for this cpu.
* This doesn't apply if the baseline access on the CPU is lost.
*
* Under core scheduling, a pick dispatches only when nothing is
@@ -3226,6 +3225,16 @@ static void put_prev_task_scx(struct rq
WARN_ON_ONCE(!sched_core_enabled(rq) &&
!(sch->ops.flags & SCX_OPS_ENQ_LAST));
scx_do_enqueue_task(rq, p, SCX_ENQ_LAST, -1);
+
+ /*
+ * A task put on the local DSQ runs, as anywhere else.
+ * Here the insert can't reschedule on its own: the pick
+ * has settled on @next and @p is still curr, so
+ * resched_curr() would flag @p and __schedule() clears
+ * that right after the pick. Flag @next instead.
+ */
+ if (p->scx.dsq == &rq->scx.local_dsq)
+ set_tsk_need_resched(next);
} else {
scx_do_enqueue_task(rq, p, 0, -1);
}
--- a/kernel/sched/ext/internal.h
+++ b/kernel/sched/ext/internal.h
@@ -1774,8 +1774,9 @@ enum scx_enq_flags {
* %SCX_OPS_ENQ_LAST is specified, they're ops.enqueue()'d with the
* %SCX_ENQ_LAST flag set.
*
- * The BPF scheduler is responsible for triggering a follow-up
- * scheduling event. Otherwise, Execution may stall.
+ * If the task is queued on the local DSQ of the CPU it was running on,
+ * it continues to run. Otherwise, the CPU goes idle. A scheduler that
+ * wants a full dispatch cycle on the CPU should kick it.
*/
SCX_ENQ_LAST = 1LLU << 41,
reply other threads:[~2026-10-07 16:26 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2b405553751fe811b4cd2c24f307d961@kernel.org \
--to=tj@kernel.org \
--cc=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=david.dai@linux.dev \
--cc=emil@etsalapatis.com \
--cc=linux-kernel@vger.kernel.org \
--cc=schatzberg.dan@gmail.com \
--cc=sched-ext@lists.linux.dev \
--cc=void@manifault.com \
--cc=xbu@meta.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®