mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Andrea Righi <arighi@nvidia.com>
To: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
	Changwoo Min <changwoo@igalia.com>,
	John Stultz <jstultz@google.com>
Cc: Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	Christian Loehle <christian.loehle@arm.com>,
	David Dai <david.dai@linux.dev>, Koba Ko <kobak@nvidia.com>,
	Aiqun Yu <aiqun.yu@oss.qualcomm.com>,
	Shuah Khan <shuah@kernel.org>,
	sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH 3/9] sched_ext: Fix TOCTOU race in consume_remote_task()
Date: Mon,  6 Jul 2026 08:50:58 +0200	[thread overview]
Message-ID: <20260706070410.282826-4-arighi@nvidia.com> (raw)
In-Reply-To: <20260706070410.282826-1-arighi@nvidia.com>

When pulling a task from a non-local DSQ, scx_consume_dispatch_q()
checks whether the task can run on the destination rq via
task_can_run_on_remote_rq(). However, it then drops the destination rq
lock and locks the source rq in consume_remote_task() ->
unlink_dsq_and_lock_src_rq(). During this window, the task might become
migration disabled, making it invalid to migrate to the destination rq.

With proxy execution enabled, a mutex-intensive workload such as
stress-ng --pipeherd 0 can trigger this race when a donor becomes active
on its source rq during the unlocked window. Migrating the active
execution context can make it resume with IRQ and preemption state from
the wrong scheduling path, triggering sleeping-while-atomic warnings and
subsequent lockdep corruption.

Re-evaluate task_can_run_on_remote_rq() after locking the source rq. Use
a non-enforcing check because eligibility may have changed during the
lock handoff independently of BPF policy; reporting that transient race
through scx_error() would unnecessarily abort the scheduler. If the task
can no longer migrate, clear its DSQ association, reset the holding CPU,
and enqueue it to the global DSQ instead. Normal consumption filtering
then skips it on ineligible CPUs while allowing an eligible CPU to pick
it without forcing it onto the source local DSQ.

Also reject tasks which are currently executing. Although an on-CPU task
should normally not be available for remote DSQ consumption, checking it
explicitly prevents a stale or racy candidate from reaching
deactivate_task() and set_task_cpu().

While the destination rq lock is dropped, clear the tracked rq state and
restore it after reacquiring the lock. Otherwise, a nested ops.dequeue()
callback can attempt to restore an rq which is no longer locked.

Closing the race and preserving rq tracking across the lock handoff are
prerequisites for correctly supporting proxy execution with sched_ext.

Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
 kernel/sched/ext/ext.c | 62 ++++++++++++++++++++++++++++++++++++------
 1 file changed, 53 insertions(+), 9 deletions(-)

diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 242f6ffbba350..4c6cd694c86db 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -2131,8 +2131,10 @@ static void move_remote_task_to_local_dsq(struct task_struct *p, u64 enq_flags,
  * - The BPF scheduler is bypassed while the rq is offline and we can always say
  *   no to the BPF scheduler initiated migrations while offline.
  *
- * The caller must ensure that @p and @rq are on different CPUs.
- * If enforce == true, caller must hold @p's rq lock.
+ * The caller must ensure that @p and @rq are on different CPUs. If @enforce is
+ * true, report violations attributable to BPF-directed migrations. The caller
+ * must hold @p's rq lock to avoid reporting a transient race as a scheduler
+ * error.
  */
 static bool task_can_run_on_remote_rq(struct scx_sched *sch,
 				      struct task_struct *p, struct rq *rq,
@@ -2150,6 +2152,10 @@ static bool task_can_run_on_remote_rq(struct scx_sched *sch,
 
 	WARN_ON_ONCE(task_cpu(p) == cpu);
 
+	/* Don't migrate a task which is running on a CPU. */
+	if (task_on_cpu(task_rq(p), p))
+		return false;
+
 	/*
 	 * If @p has migration disabled, @p->cpus_ptr is updated to contain only
 	 * the pinned CPU in migrate_disable_switch() while @p is being switched
@@ -2240,20 +2246,58 @@ static bool unlink_dsq_and_lock_src_rq(struct task_struct *p,
 		!WARN_ON_ONCE(src_rq != task_rq(p));
 }
 
-static bool consume_remote_task(struct rq *this_rq,
+static bool consume_remote_task(struct scx_sched *sch, struct rq *this_rq,
 				struct task_struct *p, u64 enq_flags,
 				struct scx_dispatch_q *dsq, struct rq *src_rq)
 {
+	struct rq *tracked_rq = scx_locked_rq();
+	bool consumed = false;
+
+	/*
+	 * consume_remote_task() may be called from an SCX op with @this_rq
+	 * recorded as the currently locked rq. Clear the tracking while the rq
+	 * lock is dropped so nested callbacks don't save and later try to restore
+	 * an rq which isn't locked anymore.
+	 */
+	if (tracked_rq) {
+		WARN_ON_ONCE(tracked_rq != this_rq);
+		update_locked_rq(NULL);
+	}
 	raw_spin_rq_unlock(this_rq);
 
 	if (unlink_dsq_and_lock_src_rq(p, dsq, src_rq)) {
+		/*
+		 * Eligibility may have changed while switching rq locks. This is a
+		 * kernel-side race, not an invalid BPF placement request, so don't
+		 * abort the scheduler on failure. Fall back to the global DSQ, where
+		 * normal consumption filters can select an eligible CPU without
+		 * forcing the task onto the source local DSQ.
+		 */
+		if (unlikely(!task_can_run_on_remote_rq(sch, p, this_rq, false))) {
+			p->scx.dsq = NULL;
+			p->scx.holding_cpu = -1;
+			scx_dispatch_enqueue(sch, src_rq,
+					     find_global_dsq(sch, task_cpu(p)), p,
+					     enq_flags | SCX_ENQ_CLEAR_OPSS |
+					     SCX_ENQ_GDSQ_FALLBACK);
+			if (sched_class_above(p->sched_class, src_rq->donor->sched_class))
+				resched_curr(src_rq);
+			raw_spin_rq_unlock(src_rq);
+			goto relock;
+		}
 		move_remote_task_to_local_dsq(p, enq_flags, src_rq, this_rq);
-		return true;
-	} else {
-		raw_spin_rq_unlock(src_rq);
-		raw_spin_rq_lock(this_rq);
-		return false;
+		consumed = true;
+		goto restore;
 	}
+	raw_spin_rq_unlock(src_rq);
+
+relock:
+	raw_spin_rq_lock(this_rq);
+restore:
+	if (tracked_rq)
+		update_locked_rq(tracked_rq);
+
+	return consumed;
 }
 
 /**
@@ -2363,7 +2407,7 @@ bool scx_consume_dispatch_q(struct scx_sched *sch, struct rq *rq,
 		}
 
 		if (task_can_run_on_remote_rq(sch, p, rq, false)) {
-			if (likely(consume_remote_task(rq, p, enq_flags, dsq, task_rq)))
+			if (likely(consume_remote_task(sch, rq, p, enq_flags, dsq, task_rq)))
 				return true;
 			goto retry;
 		}
-- 
2.55.0


  parent reply	other threads:[~2026-07-06  7:04 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-06  6:50 [PATCHSET v3 sched_ext/for-7.3] sched: Make proxy execution compatible with sched_ext Andrea Righi
2026-07-06  6:50 ` [PATCH 1/9] sched/core: Drop mutex locks before proxy rescheduling Andrea Righi
2026-07-06  8:45   ` K Prateek Nayak
2026-07-06  6:50 ` [PATCH 2/9] sched_ext: Split curr|donor references properly Andrea Righi
2026-07-06  6:50 ` Andrea Righi [this message]
2026-07-06  6:50 ` [PATCH 4/9] sched_ext: Handle blocked donor migration with proxy execution Andrea Righi
2026-07-06  6:51 ` [PATCH 5/9] sched_ext: Fix ops.running/stopping() pairing for proxy-exec donors Andrea Righi
2026-07-06  6:51 ` [PATCH 6/9] sched_ext: Delegate proxy donor admission to BPF schedulers Andrea Righi
2026-07-06  7:49   ` K Prateek Nayak
2026-07-06  6:51 ` [PATCH 7/9] sched_ext: Add selftest for blocked donor admission Andrea Righi
2026-07-06  6:51 ` [PATCH 8/9] sched_ext: scx_qmap: Add proxy execution support Andrea Righi
2026-07-06  6:51 ` [PATCH 9/9] sched: Allow enabling proxy exec with sched_ext Andrea Righi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260706070410.282826-4-arighi@nvidia.com \
    --to=arighi@nvidia.com \
    --cc=aiqun.yu@oss.qualcomm.com \
    --cc=bsegall@google.com \
    --cc=changwoo@igalia.com \
    --cc=christian.loehle@arm.com \
    --cc=david.dai@linux.dev \
    --cc=dietmar.eggemann@arm.com \
    --cc=jstultz@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kobak@nvidia.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=shuah@kernel.org \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=void@manifault.com \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome