mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH sched_ext/for-7.3-fixes] sched_ext: Unlink pending local reenqueues before freeing scheduler
@ 2026-09-16 14:58 Andrea Righi
  2026-09-16 16:33 ` Cheng-Yang Chou
  2026-09-16 22:00 ` Tejun Heo
  0 siblings, 2 replies; 3+ messages in thread
From: Andrea Righi @ 2026-09-16 14:58 UTC (permalink / raw)
  To: Tejun Heo, David Vernet, Changwoo Min
  Cc: Cheng-Yang Chou, sched-ext, linux-kernel

A deferred local DSQ reenqueue embeds its list node in struct
scx_sched_pcpu and links it into rq->scx.deferred_reenq_locals.

A scheduler teardown enters bypass before the RCU grace period, which
prevents new requests. However, this does not guarantee that an
already-linked request has been consumed. An RCU grace period waits for
active readers and it does not flush a pending rq deferred request.

scx_sched_free_rcu_work() assumes that each node has been removed. It only
warns when one remains linked and then frees sch->pcpu. The rq list
therefore retains a pointer into freed per-CPU storage. A later
run_deferred() can walk the stale node, derive sch_pcpu and sch from it,
and dereference freed memory.

This was reproduced on a 352-CPU arm64 machine by repeatedly attaching
an SCX scheduler, running hackbench, and detaching it:

  for i in $(seq 1 100); do
          sudo timeout --signal=INT 15s \
                  scx_cidland --stats 1 &
          sleep 1
          hackbench -l 2000 -g 100
          wait
  done

The scheduler was enabled and disabled twice in close succession. On the
second detach, scx_sched_free_rcu_work() reported the pending node:

  WARNING: kernel/sched/ext/ext.c:5351 at
           scx_sched_free_rcu_work+0x460/0x5a0
  Workqueue: events_unbound scx_sched_free_rcu_work

After scx_cidland was attached again, a hackbench worker hit the stale
entry two seconds later:

  Unable to handle kernel paging request at virtual address
  000000000010b8bf
  CPU: 89 PID: 590569 Comm: hackbench
  pc : run_deferred+0x148/0x5a8
  lr : run_deferred+0x194/0x5a8
  x1 : 000000000010b8bf
  Call trace:
    run_deferred+0x148/0x5a8
    task_woken_scx+0x1c/0x40
    wake_up_new_task+0x1d4/0x448
    kernel_clone+0x1b8/0x5e8

The oops left interrupts disabled and was followed by persistent RCU
stalls, making the system unusable.

Cancel any pending local reenqueue before freeing sch->pcpu. Take the rq
lock first to wait for any in-flight run_deferred() and prevent another
one from starting, then unlink the request under deferred_reenq_lock
(the request is obsolete once its scheduler is being torn down).

Fixes: 0d8c551dd5de ("sched_ext: Make scx_bpf_reenqueue_local() sub-sched aware")
Cc: stable@vger.kernel.org # v7.1+
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
 kernel/sched/ext/ext.c | 17 ++++++++++++-----
 1 file changed, 12 insertions(+), 5 deletions(-)

diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 70b711c4de6e1..94ee33ec88de6 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -5380,13 +5380,20 @@ static void scx_sched_free_rcu_work(struct work_struct *work)
 
 	for_each_possible_cpu(cpu) {
 		struct scx_sched_pcpu *pcpu = per_cpu_ptr(sch->pcpu, cpu);
+		struct rq *rq = cpu_rq(cpu);
 
 		/*
-		 * $sch would have entered bypass mode before the RCU grace
-		 * period. As that blocks new deferrals, all
-		 * deferred_reenq_local_node's must be off-list by now.
+		 * Bypass blocks new deferrals, but a request queued before bypass
+		 * may still be pending. As run_deferred() runs under the rq lock,
+		 * take it to wait for any in-flight processing before unlinking the
+		 * now-obsolete request.
 		 */
-		WARN_ON_ONCE(!list_empty(&pcpu->deferred_reenq_local.node));
+		scoped_guard (rq_lock_irqsave, rq) {
+			guard(raw_spinlock)(&rq->scx.deferred_reenq_lock);
+
+			if (!list_empty(&pcpu->deferred_reenq_local.node))
+				list_del_init(&pcpu->deferred_reenq_local.node);
+		}
 
 		/* remove the queued ecaps sync so the pcpu can be freed */
 		scx_discard_ecaps_to_sync(cpu, pcpu);
@@ -5395,7 +5402,7 @@ static void scx_sched_free_rcu_work(struct work_struct *work)
 		 * Bypass blocks new kicks. Flush the kick irq_work so this
 		 * pcpu's to_kick_node is off the list before it is freed.
 		 */
-		irq_work_sync(&cpu_rq(cpu)->scx.kick_cpus_irq_work);
+		irq_work_sync(&rq->scx.kick_cpus_irq_work);
 		WARN_ON_ONCE(!list_empty(&pcpu->to_kick_node));
 		free_cpumask_var(pcpu->cpus_to_kick);
 		free_cpumask_var(pcpu->cpus_to_kick_if_idle);
-- 
2.55.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH sched_ext/for-7.3-fixes] sched_ext: Unlink pending local reenqueues before freeing scheduler
  2026-09-16 14:58 [PATCH sched_ext/for-7.3-fixes] sched_ext: Unlink pending local reenqueues before freeing scheduler Andrea Righi
@ 2026-09-16 16:33 ` Cheng-Yang Chou
  2026-09-16 22:00 ` Tejun Heo
  1 sibling, 0 replies; 3+ messages in thread
From: Cheng-Yang Chou @ 2026-09-16 16:33 UTC (permalink / raw)
  To: Andrea Righi
  Cc: Tejun Heo, David Vernet, Changwoo Min, sched-ext, linux-kernel,
	Ching-Chun Huang, Chia-Ping Tsai, chengyang.chou

Hi Andrea,

On Wed, Sep 16, 2026 at 04:58:07PM +0200, Andrea Righi wrote:
[...]

> @@ -5380,13 +5380,20 @@ static void scx_sched_free_rcu_work(struct work_struct *work)
>  
>  	for_each_possible_cpu(cpu) {
>  		struct scx_sched_pcpu *pcpu = per_cpu_ptr(sch->pcpu, cpu);
> +		struct rq *rq = cpu_rq(cpu);
>  
>  		/*
> -		 * $sch would have entered bypass mode before the RCU grace
> -		 * period. As that blocks new deferrals, all
> -		 * deferred_reenq_local_node's must be off-list by now.
> +		 * Bypass blocks new deferrals, but a request queued before bypass
> +		 * may still be pending. As run_deferred() runs under the rq lock,
> +		 * take it to wait for any in-flight processing before unlinking the
> +		 * now-obsolete request.
>  		 */
> -		WARN_ON_ONCE(!list_empty(&pcpu->deferred_reenq_local.node));
> +		scoped_guard (rq_lock_irqsave, rq) {
> +			guard(raw_spinlock)(&rq->scx.deferred_reenq_lock);
> +
> +			if (!list_empty(&pcpu->deferred_reenq_local.node))
> +				list_del_init(&pcpu->deferred_reenq_local.node);
> +		}
>  
>  		/* remove the queued ecaps sync so the pcpu can be freed */
>  		scx_discard_ecaps_to_sync(cpu, pcpu);
> @@ -5395,7 +5402,7 @@ static void scx_sched_free_rcu_work(struct work_struct *work)
>  		 * Bypass blocks new kicks. Flush the kick irq_work so this
>  		 * pcpu's to_kick_node is off the list before it is freed.
>  		 */
> -		irq_work_sync(&cpu_rq(cpu)->scx.kick_cpus_irq_work);
> +		irq_work_sync(&rq->scx.kick_cpus_irq_work);
>  		WARN_ON_ONCE(!list_empty(&pcpu->to_kick_node));
>  		free_cpumask_var(pcpu->cpus_to_kick);
>  		free_cpumask_var(pcpu->cpus_to_kick_if_idle);
> -- 
> 2.55.0
> 

Though I couldn't reproduce the issue on my end, the fix lgtm. Thanks!

Acked-by: Cheng-Yang Chou <yphbchou0911@gmail.com>

-- 
Cheers,
Cheng-Yang

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH sched_ext/for-7.3-fixes] sched_ext: Unlink pending local reenqueues before freeing scheduler
  2026-09-16 14:58 [PATCH sched_ext/for-7.3-fixes] sched_ext: Unlink pending local reenqueues before freeing scheduler Andrea Righi
  2026-09-16 16:33 ` Cheng-Yang Chou
@ 2026-09-16 22:00 ` Tejun Heo
  1 sibling, 0 replies; 3+ messages in thread
From: Tejun Heo @ 2026-09-16 22:00 UTC (permalink / raw)
  To: Andrea Righi
  Cc: Tejun Heo, David Vernet, Changwoo Min, Cheng-Yang Chou,
	Emil Tsalapatis, Ching-Chun Huang, Chia-Ping Tsai,
	chengyang.chou, sched-ext, linux-kernel

Hello,

On Wed, Sep 16, 2026 at 04:58:07PM +0200, Andrea Righi wrote:
> A scheduler teardown enters bypass before the RCU grace period, which
> prevents new requests. However, this does not guarantee that an
> already-linked request has been consumed. An RCU grace period waits for
> active readers and it does not flush a pending rq deferred request.

Every link is followed by a scheduled run_deferred(), so a request that is
still linked when the scheduler is freed means that schedule was lost. The
WARN is doing its job there. Unlinking would hide the next one.

The lost schedule comes from move_remote_task_to_local_dsq(). It stashes the
mover's enq_flags for the destination enqueue and, since 57ccf5ccdc56
("sched_ext: Fix enqueue_task_scx() truncation of upper enqueue flags"),
those flags decide SCX_RQ_IN_WAKEUP. An IMMED insert into a busy remote CPU
requests a local reenqueue while the flag is set, schedule_deferred_locked()
leaves it to task_woken_scx(), and nothing calls that after activate_task().
Fix posted here:

  https://lore.kernel.org/all/20260916215713.2701551-1-tj@kernel.org/

Thanks.

-- 
tejun

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-16 22:00 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-16 14:58 [PATCH sched_ext/for-7.3-fixes] sched_ext: Unlink pending local reenqueues before freeing scheduler Andrea Righi
2026-09-16 16:33 ` Cheng-Yang Chou
2026-09-16 22:00 ` Tejun Heo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®