mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v4 sched_ext/for-7.4] sched_ext: Keep proxy donors with slice left on the local DSQ
@ 2026-10-10  9:06 Andrea Righi
  0 siblings, 0 replies; only message in thread
From: Andrea Righi @ 2026-10-10  9:06 UTC (permalink / raw)
  To: Tejun Heo, David Vernet, Changwoo Min
  Cc: John Stultz, sched-ext, linux-kernel

Commit ee172227d0dc ("sched_ext: Delegate proxy donor admission to BPF
schedulers") makes put_prev_task_scx() pass a retained proxy donor to
ops.enqueue() with SCX_ENQ_BLOCKED.

Some of these puts are only proxy bookkeeping; proxy_resched_idle()
drops the rq's donor reference before switching to idle when:
 - find_proxy_task() cannot run the mutex owner yet and retries through
   idle (e.g., the owner is on a remote CPU),
 - proxy_migrate_task() detaches the donor before migration,
 - proxy_deactivate() blocks a task in the chain whose owner cannot run.

In the first case, BPF has already selected a donor with slice left.
Returning it to ops.enqueue() forces BPF to dispatch it again before
proxy resolution can continue. In the other two cases, the donor is
typically deactivated right after ops.enqueue(), undoing any placement
BPF makes.

As also discussed at the sched_ext microconference at Linux Plumbers
2026, these ops.enqueue/dequeue() invocations do not represent any
meaningful scheduling events, so we should avoid triggering them.

A more reasonable semantic is to keep a donor with slice left at the
head of the local DSQ, so the next pick can resolve its owner, or
deactivation can remove it without an unnecessary BPF handoff.

An IMMED donor can stay there only for a bookkeeping put during proxy
resolution. Mark a blocked pick as awaiting resolution with
SCX_RQ_PROXY_PENDING to distinguish that put from a real preemption and
re-insert the donor with SCX_ENQ_IMMED, so that it only needs
SCX_CAP_ENQ_IMMED rather than SCX_CAP_ENQ. A preempted IMMED donor still
returns to BPF. If a higher priority class takes the CPU once proxy
resolution completes, schedule a local reenqueue, so that a parked IMMED
donor returns to BPF instead of lingering on the local DSQ. Update the
SCX_ENQ_IMMED and SCX_ENQ_BLOCKED documentation accordingly.

Moreover, a donor that has run out of slice now follows the regular put
path, so it can receive SCX_ENQ_LAST when the CPU is about to go idle
and the scheduler needs to arrange a follow-up scheduling event. This
does not apply to bookkeeping puts, where the switch to idle is only
temporary. The blocked-donor specific sanity checks go away together
with the special case.

When a retained donor is removed ahead of a scheduler ownership or
policy change, mark the proxy reset with SCX_RQ_PROXY_BLOCKING and skip
reenqueueing the active donor: sched_proxy_block_task() will dequeue it
next. Report it as not runnable to ops.stopping(), as a regular dequeue
of the current donor does. This avoids a transient ops.enqueue/dequeue()
pair and a false ENQ_LAST warning.

With this, the sched_ext core absorbs the transient proxy-bookkeeping
puts and the BPF scheduler only sees a blocked donor when there is an
actual placement decision to make.

Fixes: ee172227d0dc ("sched_ext: Delegate proxy donor admission to BPF schedulers")
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
Changes in v4:
 - Schedule a local reenqueue when a higher priority class takes the CPU after
   the transient proxy idle, so that an IMMED donor parked on the local DSQ does
   not linger behind it (Tejun Heo)
 - Skip the SCX_ENQ_LAST path for proxy bookkeeping puts, fixing the false
   ENQ_LAST warning when the donor is not kept on the local DSQ (Tejun Heo)
 - Rename SCX_RQ_PROXY_PICK_PENDING to SCX_RQ_PROXY_PENDING
 - Link to v3: https://lore.kernel.org/r/20261009200827.4026499-1-arighi@nvidia.com

 kernel/sched/ext/ext.c      | 84+++++++++++++++++++++++++++----------
 kernel/sched/ext/internal.h | 10 ++++-
 kernel/sched/sched.h        |  2 +
 3 files changed, 72 insertions(+), 24 deletions(-)

diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 248b39d09ce3c..49de031204d19 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -24,6 +24,17 @@
 
 DEFINE_RAW_SPINLOCK(scx_sched_lock);
 
+/*
+ * Block a retained proxy donor, telling put_prev_task_scx() that the donor is
+ * about to be dequeued and must not be reenqueued.
+ */
+static void scx_block_proxy_donor(struct rq *rq, struct task_struct *p)
+{
+	rq->scx.flags |= SCX_RQ_PROXY_BLOCKING;
+	sched_proxy_block_task(rq, p);
+	rq->scx.flags &= ~SCX_RQ_PROXY_BLOCKING;
+}
+
 bool __scx_allow_proxy_exec(const struct task_struct *p)
 {
 	struct scx_sched *sch;
@@ -47,7 +58,7 @@ void scx_prepare_task_sched_change(struct task_struct *p)
 	lockdep_assert_rq_held(task_rq(p));
 
 	update_rq_clock(task_rq(p));
-	sched_proxy_block_task(task_rq(p), p);
+	scx_block_proxy_donor(task_rq(p), p);
 }
 
 /*
@@ -1189,6 +1200,14 @@ void scx_proxy_reenqueue_retry(struct rq *rq, struct task_struct *next)
 		scx_proxy_update_tick(rq, next);
 #endif
 
+	/*
+	 * A higher priority class may wake up while @rq is transiently idle
+	 * for proxy resolution. wakeup_preempt_scx() isn't invoked in that
+	 * case, so reenqueue IMMED donors parked on the local DSQ here.
+	 */
+	if (rq->scx.nr_immed && sched_class_above(rq->next_class, &ext_sched_class))
+		scx_schedule_reenq_local(rq, 0);
+
 	if (rq->scx.flags & SCX_RQ_PROXY_RETRY) {
 		rq->scx.flags &= ~SCX_RQ_PROXY_RETRY;
 		schedule_deferred_locked(rq);
@@ -3346,6 +3365,16 @@ static void set_next_task_scx(struct rq *rq, struct task_struct *p, enum snt_e t
 	bool first = type == SNT_PICK;
 	bool can_stop_tick;
 
+	/*
+	 * A blocked pick is provisional until proxy resolution completes, see
+	 * put_prev_task_scx(). A retry can repick the same donor, so
+	 * SNT_REPICK must set it again.
+	 */
+	rq->scx.flags &= ~SCX_RQ_PROXY_PENDING;
+	if (sched_proxy_exec() && p->is_blocked &&
+	    (type == SNT_PICK || type == SNT_REPICK))
+		rq->scx.flags |= SCX_RQ_PROXY_PENDING;
+
 	if (type == SNT_REPICK)
 		return;
 
@@ -3425,6 +3454,7 @@ void scx_proxy_donor_start(struct rq *rq)
 	struct task_struct *donor = rq->donor;
 
 	lockdep_assert_rq_held(rq);
+	rq->scx.flags &= ~SCX_RQ_PROXY_PENDING;
 
 	if (donor->sched_class == &ext_sched_class && (donor->scx.flags & SCX_TASK_QUEUED))
 		scx_start_task_running(rq, donor);
@@ -3485,8 +3515,13 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
 			      struct task_struct *next)
 {
 	struct scx_sched *sch = scx_task_sched(p);
+	bool proxy_put = p->is_blocked && next == rq->idle &&
+			 (rq->scx.flags & SCX_RQ_PROXY_PENDING);
+	bool proxy_block = rq->scx.flags & SCX_RQ_PROXY_BLOCKING;
 	bool rescue_keep = false;
 
+	rq->scx.flags &= ~SCX_RQ_PROXY_PENDING;
+
 	/* see kick_sync_wait_bal_cb() */
 	smp_store_release(&rq->scx.kick_sync, rq->scx.kick_sync + 1);
 
@@ -3510,31 +3545,24 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
 	if (next != p && (p->scx.flags & SCX_TASK_QUEUED) &&
 	    (p->scx.flags & SCX_TASK_RUN_TRACKED)) {
 		if (SCX_HAS_OP(sch, stopping))
-			SCX_CALL_OP_TASK(sch, stopping, rq, p, true);
+			SCX_CALL_OP_TASK(sch, stopping, rq, p, !proxy_block);
 
 		p->scx.flags &= ~SCX_TASK_RUN_TRACKED;
 	}
 
 	if (p->scx.flags & SCX_TASK_QUEUED) {
-		set_task_runnable(rq, p);
+		if (proxy_block)
+			goto switch_class;
 
-		/* Delegate retained donor admission to its owning BPF scheduler. */
-		if (p->is_blocked) {
-			/*
-			 * If the donor is the same and only the mutex owner
-			 * changes, avoid triggering another ops.enqueue(): the
-			 * BPF scheduler has already admitted the donor, so it
-			 * can continue running.
-			 */
-			if (next == p)
-				goto switch_class;
+		set_task_runnable(rq, p);
 
-			if (WARN_ON_ONCE(!sch))
-				goto switch_class;
-			WARN_ON_ONCE(!(sch->ops.flags & SCX_OPS_ENQ_BLOCKED));
-			scx_do_enqueue_task(rq, p, 0, -1);
+		/*
+		 * If the donor is the same and only the mutex owner changes,
+		 * avoid triggering another ops.enqueue(): the BPF scheduler has
+		 * already admitted the donor, so it can continue running.
+		 */
+		if (p->is_blocked && next == p)
 			goto switch_class;
-		}
 
 		/*
 		 * If @p has slice left and is being put, @p is getting
@@ -3543,12 +3571,16 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
 		 * DSQ unless it was an IMMED task. IMMED tasks should not
 		 * linger on a busy CPU, reenqueue them to the BPF scheduler.
 		 *
+		 * The exception is an IMMED donor put by proxy_resched_idle():
+		 * the CPU isn't busy, proxy resolution is only retrying through
+		 * idle, so keep the donor where it can be picked again.
+		 *
 		 * An open rescue must keep @p on the local DSQ even if the
 		 * scheduler zeroed the slice in ops.stopping() above.
 		 */
 		if ((p->scx.slice || unlikely(p == scx_rescuee(rq))) &&
 		    !scx_bypassing(sch, cpu_of(rq))) {
-			if (p->scx.flags & SCX_TASK_IMMED) {
+			if ((p->scx.flags & SCX_TASK_IMMED) && !proxy_put) {
 				p->scx.flags |= SCX_TASK_REENQ_PREEMPTED;
 				scx_do_enqueue_task(rq, p, SCX_ENQ_REENQ, -1);
 			} else {
@@ -3565,6 +3597,9 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
 						enq_flags |= SCX_ENQ_HEAD;
 				} else {
 					enq_flags |= SCX_ENQ_HEAD;
+					/* only require SCX_CAP_ENQ_IMMED, see scx_caps_for_enq() */
+					if (proxy_put && (p->scx.flags & SCX_TASK_IMMED))
+						enq_flags |= SCX_ENQ_IMMED;
 				}
 
 				scx_dispatch_enqueue(sch, rq, &rq->scx.local_dsq, p, 0, 0,
@@ -3578,13 +3613,15 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
 		 * sched_class, %SCX_OPS_ENQ_LAST must be set. Tell
 		 * ops.enqueue() that @p is the only one available for this cpu,
 		 * which should trigger an explicit follow-up scheduling event.
-		 * This doesn't apply if the baseline access on the CPU is lost.
+		 * This doesn't apply if the baseline access on the CPU is lost or
+		 * proxy resolution temporarily switches to idle.
 		 *
 		 * Under core scheduling, a pick dispatches only when nothing is
 		 * locally runnable and can legitimately go idle with @p still
 		 * runnable (see do_pick_task_scx()).
 		 */
-		if (next && sched_class_above(&ext_sched_class, next->sched_class) &&
+		if (!proxy_put && next &&
+		    sched_class_above(&ext_sched_class, next->sched_class) &&
 		    scx_task_can_stay_on_cpu(rq, p)) {
 			WARN_ON_ONCE(!sched_core_enabled(rq) &&
 				     !(sch->ops.flags & SCX_OPS_ENQ_LAST));
@@ -3756,6 +3793,9 @@ do_pick_task_scx(struct rq *rq, struct rq_flags *rf, bool force_scx)
 	enum scx_dsp_verdict verdict;
 	struct task_struct *p;
 
+	/* A retry can abandon a provisional blocked-donor pick. */
+	rq->scx.flags &= ~SCX_RQ_PROXY_PENDING;
+
 	/* see kick_sync_wait_bal_cb() */
 	smp_store_release(&rq->scx.kick_sync, rq->scx.kick_sync + 1);
 
@@ -4776,7 +4816,7 @@ void scx_prepare_setscheduler(struct task_struct *p, int policy)
 	lockdep_assert_rq_held(task_rq(p));
 
 	if (scx_enabled() && p->policy != policy && policy == SCHED_EXT)
-		sched_proxy_block_task(task_rq(p), p);
+		scx_block_proxy_donor(task_rq(p), p);
 }
 
 static void process_ddsp_deferred_locals(struct rq *rq)
diff --git a/kernel/sched/ext/internal.h b/kernel/sched/ext/internal.h
index 8dcab02a38ab1..8821eb7d143e3 100644
--- a/kernel/sched/ext/internal.h
+++ b/kernel/sched/ext/internal.h
@@ -1818,8 +1818,8 @@ enum scx_enq_flags {
 	/*
 	 * Only allowed on local DSQs. Guarantees that the task either gets
 	 * on the CPU immediately and stays on it, or gets reenqueued back
-	 * to the BPF scheduler. It will never linger on a local DSQ or be
-	 * silently put back after preemption.
+	 * to the BPF scheduler. A blocked proxy donor may be put back on the
+	 * local DSQ temporarily while proxy execution resolves its mutex owner.
 	 *
 	 * The protection persists until the next fresh enqueue - it
 	 * survives SAVE/RESTORE cycles, slice extensions and preemption.
@@ -1865,6 +1865,12 @@ enum scx_enq_flags {
 	/*
 	 * The task is blocked on a mutex and is being kept runnable as a proxy
 	 * donor. Only passed to ops.enqueue() when %SCX_OPS_ENQ_BLOCKED is set.
+	 *
+	 * Blocking on the mutex does not enqueue the task by itself. A donor put
+	 * with slice left stays at the head of the local DSQ, except when an
+	 * IMMED donor is preempted or cannot stay local. It is passed to
+	 * ops.enqueue() when its slice runs out, the IMMED placement cannot be
+	 * kept, or proxy execution moves it to the owner's CPU.
 	 */
 	SCX_ENQ_BLOCKED		= 1LLU << 42,
 
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index fc65296f5c745..3118502d64a9e 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -793,6 +793,8 @@ enum scx_rq_flags {
 	SCX_RQ_ROOT_IDLE_RENOTIFY	= 1 << 8, /* the root is owed update_idle() */
 	SCX_RQ_PROXY_RETRY	= 1 << 9, /* proxy-rejected tasks need retry */
 	SCX_RQ_PROXY_TICK	= 1 << 10, /* proxy execution requires the tick */
+	SCX_RQ_PROXY_PENDING	= 1 << 11, /* blocked donor awaits proxy resolution */
+	SCX_RQ_PROXY_BLOCKING	= 1 << 12, /* donor is being removed for sched change */
 
 	SCX_RQ_IN_WAKEUP	= 1 << 16,
 	SCX_RQ_IN_DISPATCH	= 1 << 17,
-- 
2.56.0


^ permalink raw reply	[flat|nested] only message in thread

only message in thread, other threads:[~2026-10-10  9:06 UTC | newest]

Thread overview: (only message) (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-10  9:06 [PATCH v4 sched_ext/for-7.4] sched_ext: Keep proxy donors with slice left on the local DSQ Andrea Righi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®