* [PATCH sched_ext/for-7.4] sched_ext: Keep proxy donors with slice left on the local DSQ
@ 2026-10-01 19:12 Andrea Righi
0 siblings, 0 replies; only message in thread
From: Andrea Righi @ 2026-10-01 19:12 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Changwoo Min
Cc: John Stultz, sched-ext, linux-kernel
Commit ee172227d0dc ("sched_ext: Delegate proxy donor admission to BPF
schedulers") makes put_prev_task_scx() pass a retained proxy donor to
ops.enqueue() with SCX_ENQ_BLOCKED.
Some of these puts are only proxy bookkeeping; proxy_resched_idle()
drops the rq's donor reference before switching to idle when:
- find_proxy_task() finds a remote mutex owner before the donor
switches out,
- proxy_migrate_task() detaches the donor before migration,
- proxy_deactivate() blocks a donor whose owner cannot run.
In the first case, BPF has already selected a donor with slice left,
returning it to ops.enqueue() forces BPF to dispatch it again before
proxy resolution can continue. In the other two cases, the caller
deactivates the donor immediately after ops.enqueue(), undoing any
placement BPF makes.
Keep a donor with slice left at the head of the local DSQ instead, so
the next pick can resolve its owner, or deactivation can remove it
without an unnecessary BPF handoff.
A donor placed with SCX_ENQ_IMMED retains SCX_TASK_IMMED and needs
special care. Normally, putting such a task with slice left reports an
IMMED preemption and returns it to BPF. Proxy resolution may put the
already selected donor to idle only to drop rq references. Treating that
as preemption would repeat the same BPF placement unnecessarily. Mark a
blocked pick as awaiting proxy resolution (SCX_RQ_PROXY_PICK_PENDING),
so this put retains the donor locally. A real preemption still returns
it to BPF. Moreover, carry SCX_ENQ_IMMED on the local insertion so
sub-scheduler capability checks use SCX_CAP_ENQ_IMMED rather than
SCX_CAP_ENQ.
For donors with slice left, this leaves BPF to handle meaningful
placement decisions rather than transient proxy-bookkeeping puts.
Fixes: ee172227d0dc ("sched_ext: Delegate proxy donor admission to BPF schedulers")
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
kernel/sched/ext/ext.c | 62 +++++++++++++++++++++++++++----------
kernel/sched/ext/internal.h | 7 +++++
kernel/sched/sched.h | 1 +
3 files changed, 53 insertions(+), 17 deletions(-)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 96c904b396019..2c5e9eccf170a 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -3344,6 +3344,20 @@ static void set_next_task_scx(struct rq *rq, struct task_struct *p, enum snt_e t
bool first = type == SNT_PICK;
bool can_stop_tick;
+ /*
+ * A blocked donor picked by sched_ext still needs proxy resolution;
+ * find_proxy_task() may put it to idle while looking for its mutex
+ * owner. That put is bookkeeping, not an IMMED preemption, so mark the
+ * pick so put_prev_task_scx() can distinguish the two.
+ *
+ * With proxy execution, a blocked mutex waiter can stay on the runqueue
+ * as a donor and be picked again (SNT_REPICK). Core scheduling's
+ * forced-idle pick selects idle instead, so it must not mark a proxy put.
+ */
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
+ if (sched_proxy_exec() && p->is_blocked && (type == SNT_PICK || type == SNT_REPICK))
+ rq->scx.flags |= SCX_RQ_PROXY_PICK_PENDING;
+
if (type == SNT_REPICK)
return;
@@ -3423,6 +3437,7 @@ void scx_proxy_donor_start(struct rq *rq)
struct task_struct *donor = rq->donor;
lockdep_assert_rq_held(rq);
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
if (donor->sched_class == &ext_sched_class && (donor->scx.flags & SCX_TASK_QUEUED))
scx_start_task_running(rq, donor);
@@ -3483,8 +3498,12 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
struct task_struct *next)
{
struct scx_sched *sch = scx_task_sched(p);
+ bool proxy_put = p->is_blocked && next == rq->idle &&
+ (rq->scx.flags & SCX_RQ_PROXY_PICK_PENDING);
bool rescue_keep = false;
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
+
/* see kick_sync_wait_bal_cb() */
smp_store_release(&rq->scx.kick_sync, rq->scx.kick_sync + 1);
@@ -3516,23 +3535,13 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
if (p->scx.flags & SCX_TASK_QUEUED) {
set_task_runnable(rq, p);
- /* Delegate retained donor admission to its owning BPF scheduler. */
- if (p->is_blocked) {
- /*
- * If the donor is the same and only the mutex owner
- * changes, avoid triggering another ops.enqueue(): the
- * BPF scheduler has already admitted the donor, so it
- * can continue running.
- */
- if (next == p)
- goto switch_class;
-
- if (WARN_ON_ONCE(!sch))
- goto switch_class;
- WARN_ON_ONCE(!(sch->ops.flags & SCX_OPS_ENQ_BLOCKED));
- scx_do_enqueue_task(rq, p, 0, -1);
+ /*
+ * If the donor is the same and only the mutex owner changes,
+ * avoid triggering another ops.enqueue(): the BPF scheduler has
+ * already admitted the donor, so it can continue running.
+ */
+ if (p->is_blocked && next == p)
goto switch_class;
- }
/*
* If @p has slice left and is being put, @p is getting
@@ -3541,12 +3550,17 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
* DSQ unless it was an IMMED task. IMMED tasks should not
* linger on a busy CPU, reenqueue them to the BPF scheduler.
*
+ * proxy_resched_idle() also puts a blocked donor while resolving its
+ * mutex owner. That put is bookkeeping, not a preemption, so retain
+ * the donor locally even if it is IMMED. The deferred local check
+ * returns it to BPF if this CPU becomes unavailable.
+ *
* An open rescue must keep @p on the local DSQ even if the
* scheduler zeroed the slice in ops.stopping() above.
*/
if ((p->scx.slice || unlikely(p == scx_rescuee(rq))) &&
!scx_bypassing(sch, cpu_of(rq))) {
- if (p->scx.flags & SCX_TASK_IMMED) {
+ if ((p->scx.flags & SCX_TASK_IMMED) && !proxy_put) {
p->scx.flags |= SCX_TASK_REENQ_PREEMPTED;
scx_do_enqueue_task(rq, p, SCX_ENQ_REENQ, -1);
} else {
@@ -3563,6 +3577,8 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
enq_flags |= SCX_ENQ_HEAD;
} else {
enq_flags |= SCX_ENQ_HEAD;
+ if (proxy_put && (p->scx.flags & SCX_TASK_IMMED))
+ enq_flags |= SCX_ENQ_IMMED;
}
scx_dispatch_enqueue(sch, rq, &rq->scx.local_dsq, p, 0, 0,
@@ -3571,6 +3587,15 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
goto switch_class;
}
+ /* Delegate retained donor admission to its owning BPF scheduler. */
+ if (p->is_blocked) {
+ if (WARN_ON_ONCE(!sch))
+ goto switch_class;
+ WARN_ON_ONCE(!(sch->ops.flags & SCX_OPS_ENQ_BLOCKED));
+ scx_do_enqueue_task(rq, p, 0, -1);
+ goto switch_class;
+ }
+
/*
* If @p is runnable but we're about to enter a lower
* sched_class, %SCX_OPS_ENQ_LAST must be set. Tell
@@ -3752,6 +3777,9 @@ do_pick_task_scx(struct rq *rq, struct rq_flags *rf, bool force_scx)
enum scx_dsp_verdict verdict;
struct task_struct *p;
+ /* A retry can abandon a provisional blocked-donor pick. */
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
+
/* see kick_sync_wait_bal_cb() */
smp_store_release(&rq->scx.kick_sync, rq->scx.kick_sync + 1);
diff --git a/kernel/sched/ext/internal.h b/kernel/sched/ext/internal.h
index d65fec631bdf7..7aa9b567bd73f 100644
--- a/kernel/sched/ext/internal.h
+++ b/kernel/sched/ext/internal.h
@@ -1829,6 +1829,13 @@ enum scx_enq_flags {
/*
* The task is blocked on a mutex and is being kept runnable as a proxy
* donor. Only passed to ops.enqueue() when %SCX_OPS_ENQ_BLOCKED is set.
+ *
+ * Blocking on the mutex does not enqueue the task by itself. A donor put
+ * with slice left stays at the head of the local DSQ, except when an
+ * IMMED donor is preempted or cannot run on the CPU. It is passed to
+ * ops.enqueue() when its slice runs out, another SCX task preempts it,
+ * an IMMED placement cannot be kept, or proxy execution moves it to the
+ * CPU of the mutex owner.
*/
SCX_ENQ_BLOCKED = 1LLU << 42,
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 0a34f0da1ec36..c7f0fc8fecec2 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -793,6 +793,7 @@ enum scx_rq_flags {
SCX_RQ_ROOT_IDLE_RENOTIFY = 1 << 8, /* the root is owed update_idle() */
SCX_RQ_PROXY_RETRY = 1 << 9, /* proxy-rejected tasks need retry */
SCX_RQ_PROXY_TICK = 1 << 10, /* proxy execution requires the tick */
+ SCX_RQ_PROXY_PICK_PENDING = 1 << 11, /* blocked donor awaits proxy resolution */
SCX_RQ_IN_WAKEUP = 1 << 16,
SCX_RQ_IN_DISPATCH = 1 << 17,
--
2.55.0
^ permalink raw reply [flat|nested] only message in thread
only message in thread, other threads:[~2026-10-01 19:12 UTC | newest]
Thread overview: (only message) (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 19:12 [PATCH sched_ext/for-7.4] sched_ext: Keep proxy donors with slice left on the local DSQ Andrea Righi
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®