mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched()
@ 2026-06-30  8:47 Sechang Lim
  2026-07-03 15:33 ` Gabriele Monaco
  2026-09-28 17:19 ` Andrea Righi
  0 siblings, 2 replies; 5+ messages in thread
From: Sechang Lim @ 2026-06-30  8:47 UTC (permalink / raw)
  To: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Gabriele Monaco,
	linux-kernel

set_tsk_need_resched() tests TIF_NEED_RESCHED, calls
__trace_set_need_resched() if the flag is clear, then sets it via
set_tsk_thread_flag().  A BPF raw_tp program attached to
sched_set_need_resched executes synchronously inside __bpf_trace_run().
On return, __bpf_trace_run() drops the RCU lock with
rcu_read_unlock_migrate(), which on the preempt-or-BH-disabled path
calls set_need_resched_current() -> set_tsk_need_resched() again.

set_tsk_thread_flag() follows the tracepoint call, so every re-entrant
frame sees TIF_NEED_RESCHED clear and calls __trace_set_need_resched()
again:

  BUG: TASK stack guard page was hit at ffffc9001224ff98
  Oops: stack guard page: 0000 [#1] SMP KASAN PTI
  RIP: 0010:__bpf_trace_sched_set_need_resched_tp+0x1c/0x190
  Call Trace:
   trace_sched_set_need_resched_tp+0x110/0x130
   set_tsk_need_resched include/linux/sched.h:2076
   set_need_resched_current include/linux/sched.h:2094
   rcu_read_unlock_special+0x43a/0x440
   __rcu_read_unlock+0x9e/0x120
   rcu_read_unlock_migrate+0xa9/0x240
   __bpf_trace_run+0x131/0x180
   bpf_trace_run3+0x333/0x430
   __bpf_trace_sched_set_need_resched_tp+0x13a/0x190
   trace_sched_set_need_resched_tp+0x110/0x130
   set_tsk_need_resched include/linux/sched.h:2076
   ...

__resched_curr() has the same ordering, firing the tracepoint before
setting the flag via set_ti_thread_flag() or set_nr_and_not_polling().
Fix it for consistency.

Replace the separate test_tsk_thread_flag() + set_tsk_thread_flag() pair
in set_tsk_need_resched() with test_and_set_tsk_thread_flag().  In
__resched_curr(), move the tracepoint call after the flag is set in
each path.

Fixes: adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
Signed-off-by: Sechang Lim <rhkrqnwk98@gmail.com>
---
v3:
 - reorder need_ipi variable. (K Prateek Nayak)

v2:
 - https://lore.kernel.org/all/20260627081657.499781-1-rhkrqnwk98@gmail.com/

v1:
 - https://lore.kernel.org/all/20260625065656.392182-1-rhkrqnwk98@gmail.com/

 include/linux/sched.h | 5 ++---
 kernel/sched/core.c   | 7 +++++--
 2 files changed, 7 insertions(+), 5 deletions(-)

diff --git a/include/linux/sched.h b/include/linux/sched.h
index ee06cba5c6f5..c9efd08dae92 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -2071,10 +2071,9 @@ static inline int test_tsk_thread_flag(struct task_struct *tsk, int flag)
 
 static inline void set_tsk_need_resched(struct task_struct *tsk)
 {
-	if (tracepoint_enabled(sched_set_need_resched_tp) &&
-	    !test_tsk_thread_flag(tsk, TIF_NEED_RESCHED))
+	if (!test_and_set_tsk_thread_flag(tsk, TIF_NEED_RESCHED) &&
+	    tracepoint_enabled(sched_set_need_resched_tp))
 		__trace_set_need_resched(tsk, TIF_NEED_RESCHED);
-	set_tsk_thread_flag(tsk,TIF_NEED_RESCHED);
 }
 
 static inline void clear_tsk_need_resched(struct task_struct *tsk)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index b8871449d3c6..19de28f0d85a 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -1171,6 +1171,7 @@ static void __resched_curr(struct rq *rq, int tif)
 {
 	struct task_struct *curr = rq->curr;
 	struct thread_info *cti = task_thread_info(curr);
+	bool need_ipi;
 	int cpu;
 
 	lockdep_assert_rq_held(rq);
@@ -1187,15 +1188,17 @@ static void __resched_curr(struct rq *rq, int tif)
 
 	cpu = cpu_of(rq);
 
-	trace_sched_set_need_resched_tp(curr, cpu, tif);
 	if (cpu == smp_processor_id()) {
 		set_ti_thread_flag(cti, tif);
 		if (tif == TIF_NEED_RESCHED)
 			set_preempt_need_resched();
+		trace_sched_set_need_resched_tp(curr, cpu, tif);
 		return;
 	}
 
-	if (set_nr_and_not_polling(cti, tif)) {
+	need_ipi = set_nr_and_not_polling(cti, tif);
+	trace_sched_set_need_resched_tp(curr, cpu, tif);
+	if (need_ipi) {
 		if (tif == TIF_NEED_RESCHED)
 			smp_send_reschedule(cpu);
 	} else {
-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched()
  2026-06-30  8:47 [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched() Sechang Lim
@ 2026-07-03 15:33 ` Gabriele Monaco
  2026-09-28 17:19 ` Andrea Righi
  1 sibling, 0 replies; 5+ messages in thread
From: Gabriele Monaco @ 2026-07-03 15:33 UTC (permalink / raw)
  To: Sechang Lim, Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot
  Cc: Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, linux-kernel

On Tue, 2026-06-30 at 08:47 +0000, Sechang Lim wrote:
> set_tsk_need_resched() tests TIF_NEED_RESCHED, calls
> __trace_set_need_resched() if the flag is clear, then sets it via
> set_tsk_thread_flag().  A BPF raw_tp program attached to
> sched_set_need_resched executes synchronously inside __bpf_trace_run().
> On return, __bpf_trace_run() drops the RCU lock with
> rcu_read_unlock_migrate(), which on the preempt-or-BH-disabled path
> calls set_need_resched_current() -> set_tsk_need_resched() again.
> 
> set_tsk_thread_flag() follows the tracepoint call, so every re-entrant
> frame sees TIF_NEED_RESCHED clear and calls __trace_set_need_resched()
> again:
> 
>   BUG: TASK stack guard page was hit at ffffc9001224ff98
>   Oops: stack guard page: 0000 [#1] SMP KASAN PTI
>   RIP: 0010:__bpf_trace_sched_set_need_resched_tp+0x1c/0x190
>   Call Trace:
>    trace_sched_set_need_resched_tp+0x110/0x130
>    set_tsk_need_resched include/linux/sched.h:2076
>    set_need_resched_current include/linux/sched.h:2094
>    rcu_read_unlock_special+0x43a/0x440
>    __rcu_read_unlock+0x9e/0x120
>    rcu_read_unlock_migrate+0xa9/0x240
>    __bpf_trace_run+0x131/0x180
>    bpf_trace_run3+0x333/0x430
>    __bpf_trace_sched_set_need_resched_tp+0x13a/0x190
>    trace_sched_set_need_resched_tp+0x110/0x130
>    set_tsk_need_resched include/linux/sched.h:2076
>    ...
> 
> __resched_curr() has the same ordering, firing the tracepoint before
> setting the flag via set_ti_thread_flag() or set_nr_and_not_polling().
> Fix it for consistency.
> 
> Replace the separate test_tsk_thread_flag() + set_tsk_thread_flag() pair
> in set_tsk_need_resched() with test_and_set_tsk_thread_flag().  In
> __resched_curr(), move the tracepoint call after the flag is set in
> each path.
> 
> Fixes: adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
> Signed-off-by: Sechang Lim <rhkrqnwk98@gmail.com>
> ---

Works fine with the existing rv monitors, thanks!

Acked-by: Gabriele Monaco <gmonaco@redhat.com>

> v3:
>  - reorder need_ipi variable. (K Prateek Nayak)
> 
> v2:
>  - https://lore.kernel.org/all/20260627081657.499781-1-rhkrqnwk98@gmail.com/
> 
> v1:
>  - https://lore.kernel.org/all/20260625065656.392182-1-rhkrqnwk98@gmail.com/
> 
>  include/linux/sched.h | 5 ++---
>  kernel/sched/core.c   | 7 +++++--
>  2 files changed, 7 insertions(+), 5 deletions(-)
> 
> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index ee06cba5c6f5..c9efd08dae92 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -2071,10 +2071,9 @@ static inline int test_tsk_thread_flag(struct
> task_struct *tsk, int flag)
>  
>  static inline void set_tsk_need_resched(struct task_struct *tsk)
>  {
> -	if (tracepoint_enabled(sched_set_need_resched_tp) &&
> -	    !test_tsk_thread_flag(tsk, TIF_NEED_RESCHED))
> +	if (!test_and_set_tsk_thread_flag(tsk, TIF_NEED_RESCHED) &&
> +	    tracepoint_enabled(sched_set_need_resched_tp))
>  		__trace_set_need_resched(tsk, TIF_NEED_RESCHED);
> -	set_tsk_thread_flag(tsk,TIF_NEED_RESCHED);
>  }
>  
>  static inline void clear_tsk_need_resched(struct task_struct *tsk)
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index b8871449d3c6..19de28f0d85a 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -1171,6 +1171,7 @@ static void __resched_curr(struct rq *rq, int tif)
>  {
>  	struct task_struct *curr = rq->curr;
>  	struct thread_info *cti = task_thread_info(curr);
> +	bool need_ipi;
>  	int cpu;
>  
>  	lockdep_assert_rq_held(rq);
> @@ -1187,15 +1188,17 @@ static void __resched_curr(struct rq *rq, int tif)
>  
>  	cpu = cpu_of(rq);
>  
> -	trace_sched_set_need_resched_tp(curr, cpu, tif);
>  	if (cpu == smp_processor_id()) {
>  		set_ti_thread_flag(cti, tif);
>  		if (tif == TIF_NEED_RESCHED)
>  			set_preempt_need_resched();
> +		trace_sched_set_need_resched_tp(curr, cpu, tif);
>  		return;
>  	}
>  
> -	if (set_nr_and_not_polling(cti, tif)) {
> +	need_ipi = set_nr_and_not_polling(cti, tif);
> +	trace_sched_set_need_resched_tp(curr, cpu, tif);
> +	if (need_ipi) {
>  		if (tif == TIF_NEED_RESCHED)
>  			smp_send_reschedule(cpu);
>  	} else {


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched()
  2026-06-30  8:47 [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched() Sechang Lim
  2026-07-03 15:33 ` Gabriele Monaco
@ 2026-09-28 17:19 ` Andrea Righi
  2026-09-29  8:50   ` Gabriele Monaco
  1 sibling, 1 reply; 5+ messages in thread
From: Andrea Righi @ 2026-09-28 17:19 UTC (permalink / raw)
  To: Sechang Lim
  Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, Gabriele Monaco,
	linux-kernel

Hi Sechang,

On Tue, Jun 30, 2026 at 08:47:37AM +0000, Sechang Lim wrote:
> set_tsk_need_resched() tests TIF_NEED_RESCHED, calls
> __trace_set_need_resched() if the flag is clear, then sets it via
> set_tsk_thread_flag().  A BPF raw_tp program attached to
> sched_set_need_resched executes synchronously inside __bpf_trace_run().
> On return, __bpf_trace_run() drops the RCU lock with
> rcu_read_unlock_migrate(), which on the preempt-or-BH-disabled path
> calls set_need_resched_current() -> set_tsk_need_resched() again.
> 
> set_tsk_thread_flag() follows the tracepoint call, so every re-entrant
> frame sees TIF_NEED_RESCHED clear and calls __trace_set_need_resched()
> again:
> 
>   BUG: TASK stack guard page was hit at ffffc9001224ff98
>   Oops: stack guard page: 0000 [#1] SMP KASAN PTI
>   RIP: 0010:__bpf_trace_sched_set_need_resched_tp+0x1c/0x190
>   Call Trace:
>    trace_sched_set_need_resched_tp+0x110/0x130
>    set_tsk_need_resched include/linux/sched.h:2076
>    set_need_resched_current include/linux/sched.h:2094
>    rcu_read_unlock_special+0x43a/0x440
>    __rcu_read_unlock+0x9e/0x120
>    rcu_read_unlock_migrate+0xa9/0x240
>    __bpf_trace_run+0x131/0x180
>    bpf_trace_run3+0x333/0x430
>    __bpf_trace_sched_set_need_resched_tp+0x13a/0x190
>    trace_sched_set_need_resched_tp+0x110/0x130
>    set_tsk_need_resched include/linux/sched.h:2076
>    ...
> 
> __resched_curr() has the same ordering, firing the tracepoint before
> setting the flag via set_ti_thread_flag() or set_nr_and_not_polling().
> Fix it for consistency.
> 
> Replace the separate test_tsk_thread_flag() + set_tsk_thread_flag() pair
> in set_tsk_need_resched() with test_and_set_tsk_thread_flag().  In
> __resched_curr(), move the tracepoint call after the flag is set in
> each path.

This leaves the race Prateek described in the v2 discussion: a remote CPU sets
TIF_NEED_RESCHED under rq->lock, but the target CPU can emit sched_entry_tp()
before acquiring that lock, ahead of the need-resched tracepoint.

I reproduced it with this applied to tip/master. With the RV nrp monitor
enabled, 300 runs of "perf bench sched messaging -g 20 -l 100" produced:

  rv: monitor nrp does not allow event schedule_entry_preempt on state any_thread_running

Applying the following on top fixed the ordering, the same 300-run test then
completed without an RV violation. If it makes sense, could you fold this change
into v4?

Thanks,
-Andrea

diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 639b5df7cf130..91a15a4234280 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -7143,9 +7143,6 @@ static void __sched notrace __schedule(int sched_mode)
 	struct rq *rq;
 	int cpu;
 
-	/* Trace preemptions consistently with task switches */
-	trace_sched_entry_tp(sched_mode == SM_PREEMPT);
-
 	cpu = smp_processor_id();
 	rq = cpu_rq(cpu);
 	prev = rq->curr;
@@ -7178,6 +7175,9 @@ static void __sched notrace __schedule(int sched_mode)
 	rq_lock(rq, &rf);
 	smp_mb__after_spinlock();
 
+	/* Trace preemptions consistently with task switches */
+	trace_sched_entry_tp(sched_mode == SM_PREEMPT);
+
 	hrtick_schedule_enter(rq);
 
 	/* Promote REQ to ACT */


> 
> Fixes: adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
> Signed-off-by: Sechang Lim <rhkrqnwk98@gmail.com>
> ---
> v3:
>  - reorder need_ipi variable. (K Prateek Nayak)
> 
> v2:
>  - https://lore.kernel.org/all/20260627081657.499781-1-rhkrqnwk98@gmail.com/
> 
> v1:
>  - https://lore.kernel.org/all/20260625065656.392182-1-rhkrqnwk98@gmail.com/
> 
>  include/linux/sched.h | 5 ++---
>  kernel/sched/core.c   | 7 +++++--
>  2 files changed, 7 insertions(+), 5 deletions(-)
> 
> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index ee06cba5c6f5..c9efd08dae92 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -2071,10 +2071,9 @@ static inline int test_tsk_thread_flag(struct task_struct *tsk, int flag)
>  
>  static inline void set_tsk_need_resched(struct task_struct *tsk)
>  {
> -	if (tracepoint_enabled(sched_set_need_resched_tp) &&
> -	    !test_tsk_thread_flag(tsk, TIF_NEED_RESCHED))
> +	if (!test_and_set_tsk_thread_flag(tsk, TIF_NEED_RESCHED) &&
> +	    tracepoint_enabled(sched_set_need_resched_tp))
>  		__trace_set_need_resched(tsk, TIF_NEED_RESCHED);
> -	set_tsk_thread_flag(tsk,TIF_NEED_RESCHED);
>  }
>  
>  static inline void clear_tsk_need_resched(struct task_struct *tsk)
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index b8871449d3c6..19de28f0d85a 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -1171,6 +1171,7 @@ static void __resched_curr(struct rq *rq, int tif)
>  {
>  	struct task_struct *curr = rq->curr;
>  	struct thread_info *cti = task_thread_info(curr);
> +	bool need_ipi;
>  	int cpu;
>  
>  	lockdep_assert_rq_held(rq);
> @@ -1187,15 +1188,17 @@ static void __resched_curr(struct rq *rq, int tif)
>  
>  	cpu = cpu_of(rq);
>  
> -	trace_sched_set_need_resched_tp(curr, cpu, tif);
>  	if (cpu == smp_processor_id()) {
>  		set_ti_thread_flag(cti, tif);
>  		if (tif == TIF_NEED_RESCHED)
>  			set_preempt_need_resched();
> +		trace_sched_set_need_resched_tp(curr, cpu, tif);
>  		return;
>  	}
>  
> -	if (set_nr_and_not_polling(cti, tif)) {
> +	need_ipi = set_nr_and_not_polling(cti, tif);
> +	trace_sched_set_need_resched_tp(curr, cpu, tif);
> +	if (need_ipi) {
>  		if (tif == TIF_NEED_RESCHED)
>  			smp_send_reschedule(cpu);
>  	} else {
> -- 
> 2.43.0
> 

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched()
  2026-09-28 17:19 ` Andrea Righi
@ 2026-09-29  8:50   ` Gabriele Monaco
  2026-09-29  9:30     ` Andrea Righi
  0 siblings, 1 reply; 5+ messages in thread
From: Gabriele Monaco @ 2026-09-29  8:50 UTC (permalink / raw)
  To: Andrea Righi, Sechang Lim
  Cc: Ingo Molnar, Peter Zijlstra, Juri Lelli, Vincent Guittot,
	Dietmar Eggemann, Steven Rostedt, Ben Segall, Mel Gorman,
	Valentin Schneider, K Prateek Nayak, linux-kernel

Hi Andrea,

On Mon, 2026-09-28 at 19:19 +0200, Andrea Righi wrote:
> This leaves the race Prateek described in the v2 discussion: a remote CPU sets
> TIF_NEED_RESCHED under rq->lock, but the target CPU can emit sched_entry_tp()
> before acquiring that lock, ahead of the need-resched tracepoint.
> 
> I reproduced it with this applied to tip/master. With the RV nrp monitor
> enabled, 300 runs of "perf bench sched messaging -g 20 -l 100" produced:
> 
>   rv: monitor nrp does not allow event schedule_entry_preempt on state
> any_thread_running
> 
> Applying the following on top fixed the ordering, the same 300-run test then
> completed without an RV violation. If it makes sense, could you fold this
> change into v4?
> 
> Thanks,
> -Andrea

thanks for looking into this. Is this change required for anything else besides
fixing the nrp monitor after changing the order with need_resched?

I believe this would break the sts monitor which expects sched_entry before
disabling interrupts.

Both monitors could be adapted to either case, but if your change is just for
the sake of nrp, I think it's easier to just allow this race, since nrp is
already allowing the race with interrupts.
I haven't tested yet, but something like allowing a sched_entry_preempt without
need_resched set but provided it's going to be set before sched_exit would
probably do (it'd allow also independent need_resched there, but those will then
need their own preemption too).

I'd say if you don't need to move sched_entry for other reasons we can wait to
apply this change until I see what's better for the models.

Thanks,
Gabriele

> 
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 639b5df7cf130..91a15a4234280 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -7143,9 +7143,6 @@ static void __sched notrace __schedule(int sched_mode)
>  	struct rq *rq;
>  	int cpu;
>  
> -	/* Trace preemptions consistently with task switches */
> -	trace_sched_entry_tp(sched_mode == SM_PREEMPT);
> -
>  	cpu = smp_processor_id();
>  	rq = cpu_rq(cpu);
>  	prev = rq->curr;
> @@ -7178,6 +7175,9 @@ static void __sched notrace __schedule(int sched_mode)
>  	rq_lock(rq, &rf);
>  	smp_mb__after_spinlock();
>  
> +	/* Trace preemptions consistently with task switches */
> +	trace_sched_entry_tp(sched_mode == SM_PREEMPT);
> +
>  	hrtick_schedule_enter(rq);
>  
>  	/* Promote REQ to ACT */
> 
> 
> > 
> > Fixes: adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
> > Signed-off-by: Sechang Lim <rhkrqnwk98@gmail.com>
> > ---
> > v3:
> >  - reorder need_ipi variable. (K Prateek Nayak)
> > 
> > v2:
> >  - https://lore.kernel.org/all/20260627081657.499781-1-rhkrqnwk98@gmail.com/
> > 
> > v1:
> >  - https://lore.kernel.org/all/20260625065656.392182-1-rhkrqnwk98@gmail.com/
> > 
> >  include/linux/sched.h | 5 ++---
> >  kernel/sched/core.c   | 7 +++++--
> >  2 files changed, 7 insertions(+), 5 deletions(-)
> > 
> > diff --git a/include/linux/sched.h b/include/linux/sched.h
> > index ee06cba5c6f5..c9efd08dae92 100644
> > --- a/include/linux/sched.h
> > +++ b/include/linux/sched.h
> > @@ -2071,10 +2071,9 @@ static inline int test_tsk_thread_flag(struct
> > task_struct *tsk, int flag)
> >  
> >  static inline void set_tsk_need_resched(struct task_struct *tsk)
> >  {
> > -	if (tracepoint_enabled(sched_set_need_resched_tp) &&
> > -	    !test_tsk_thread_flag(tsk, TIF_NEED_RESCHED))
> > +	if (!test_and_set_tsk_thread_flag(tsk, TIF_NEED_RESCHED) &&
> > +	    tracepoint_enabled(sched_set_need_resched_tp))
> >  		__trace_set_need_resched(tsk, TIF_NEED_RESCHED);
> > -	set_tsk_thread_flag(tsk,TIF_NEED_RESCHED);
> >  }
> >  
> >  static inline void clear_tsk_need_resched(struct task_struct *tsk)
> > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > index b8871449d3c6..19de28f0d85a 100644
> > --- a/kernel/sched/core.c
> > +++ b/kernel/sched/core.c
> > @@ -1171,6 +1171,7 @@ static void __resched_curr(struct rq *rq, int tif)
> >  {
> >  	struct task_struct *curr = rq->curr;
> >  	struct thread_info *cti = task_thread_info(curr);
> > +	bool need_ipi;
> >  	int cpu;
> >  
> >  	lockdep_assert_rq_held(rq);
> > @@ -1187,15 +1188,17 @@ static void __resched_curr(struct rq *rq, int tif)
> >  
> >  	cpu = cpu_of(rq);
> >  
> > -	trace_sched_set_need_resched_tp(curr, cpu, tif);
> >  	if (cpu == smp_processor_id()) {
> >  		set_ti_thread_flag(cti, tif);
> >  		if (tif == TIF_NEED_RESCHED)
> >  			set_preempt_need_resched();
> > +		trace_sched_set_need_resched_tp(curr, cpu, tif);
> >  		return;
> >  	}
> >  
> > -	if (set_nr_and_not_polling(cti, tif)) {
> > +	need_ipi = set_nr_and_not_polling(cti, tif);
> > +	trace_sched_set_need_resched_tp(curr, cpu, tif);
> > +	if (need_ipi) {
> >  		if (tif == TIF_NEED_RESCHED)
> >  			smp_send_reschedule(cpu);
> >  	} else {
> > -- 
> > 2.43.0
> > 


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched()
  2026-09-29  8:50   ` Gabriele Monaco
@ 2026-09-29  9:30     ` Andrea Righi
  0 siblings, 0 replies; 5+ messages in thread
From: Andrea Righi @ 2026-09-29  9:30 UTC (permalink / raw)
  To: Gabriele Monaco
  Cc: Sechang Lim, Ingo Molnar, Peter Zijlstra, Juri Lelli,
	Vincent Guittot, Dietmar Eggemann, Steven Rostedt, Ben Segall,
	Mel Gorman, Valentin Schneider, K Prateek Nayak, linux-kernel

Hi Gabriele,

On Tue, Sep 29, 2026 at 10:50:48AM +0200, Gabriele Monaco wrote:
> Hi Andrea,
> 
> On Mon, 2026-09-28 at 19:19 +0200, Andrea Righi wrote:
> > This leaves the race Prateek described in the v2 discussion: a remote CPU sets
> > TIF_NEED_RESCHED under rq->lock, but the target CPU can emit sched_entry_tp()
> > before acquiring that lock, ahead of the need-resched tracepoint.
> > 
> > I reproduced it with this applied to tip/master. With the RV nrp monitor
> > enabled, 300 runs of "perf bench sched messaging -g 20 -l 100" produced:
> > 
> >   rv: monitor nrp does not allow event schedule_entry_preempt on state
> > any_thread_running
> > 
> > Applying the following on top fixed the ordering, the same 300-run test then
> > completed without an RV violation. If it makes sense, could you fold this
> > change into v4?
> > 
> > Thanks,
> > -Andrea
> 
> thanks for looking into this. Is this change required for anything else besides
> fixing the nrp monitor after changing the order with need_resched?

No, I was mainly trying to help get this fix upstream: a sched_ext kselftest I'm
working on hits that issue and I'm holding off on submitting the test until the
fix lands. Moving sched_entry was only my attempt to address the nrp warning I
saw while testing the fix.

> 
> I believe this would break the sts monitor which expects sched_entry before
> disabling interrupts.
> 
> Both monitors could be adapted to either case, but if your change is just for
> the sake of nrp, I think it's easier to just allow this race, since nrp is
> already allowing the race with interrupts.
> I haven't tested yet, but something like allowing a sched_entry_preempt without
> need_resched set but provided it's going to be set before sched_exit would
> probably do (it'd allow also independent need_resched there, but those will then
> need their own preemption too).
> 
> I'd say if you don't need to move sched_entry for other reasons we can wait to
> apply this change until I see what's better for the models.

That makes sense, we can ignore the sched_entry change for now. With that,
Sechang's fix looks good to me.

Reviewed-by: Andrea Righi <arighi@nvidia.com>

Thanks,
-Andrea

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-29  9:30 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-06-30  8:47 [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched() Sechang Lim
2026-07-03 15:33 ` Gabriele Monaco
2026-09-28 17:19 ` Andrea Righi
2026-09-29  8:50   ` Gabriele Monaco
2026-09-29  9:30     ` Andrea Righi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®