mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Gabriele Monaco <gmonaco@redhat.com>
To: Andrea Righi <arighi@nvidia.com>, Sechang Lim <rhkrqnwk98@gmail.com>
Cc: Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	 Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann	 <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall	 <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider	 <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
		linux-kernel@vger.kernel.org
Subject: Re: [PATCH v3] sched: set TIF_NEED_RESCHED before calling __trace_set_need_resched()
Date: Tue, 29 Sep 2026 10:50:48 +0200	[thread overview]
Message-ID: <96b3b5d5ec0eae499f1a98845720776b6fab84d4.camel@redhat.com> (raw)
In-Reply-To: <arqhnt6lW8qiCo19@gpd4>

Hi Andrea,

On Mon, 2026-09-28 at 19:19 +0200, Andrea Righi wrote:
> This leaves the race Prateek described in the v2 discussion: a remote CPU sets
> TIF_NEED_RESCHED under rq->lock, but the target CPU can emit sched_entry_tp()
> before acquiring that lock, ahead of the need-resched tracepoint.
> 
> I reproduced it with this applied to tip/master. With the RV nrp monitor
> enabled, 300 runs of "perf bench sched messaging -g 20 -l 100" produced:
> 
>   rv: monitor nrp does not allow event schedule_entry_preempt on state
> any_thread_running
> 
> Applying the following on top fixed the ordering, the same 300-run test then
> completed without an RV violation. If it makes sense, could you fold this
> change into v4?
> 
> Thanks,
> -Andrea

thanks for looking into this. Is this change required for anything else besides
fixing the nrp monitor after changing the order with need_resched?

I believe this would break the sts monitor which expects sched_entry before
disabling interrupts.

Both monitors could be adapted to either case, but if your change is just for
the sake of nrp, I think it's easier to just allow this race, since nrp is
already allowing the race with interrupts.
I haven't tested yet, but something like allowing a sched_entry_preempt without
need_resched set but provided it's going to be set before sched_exit would
probably do (it'd allow also independent need_resched there, but those will then
need their own preemption too).

I'd say if you don't need to move sched_entry for other reasons we can wait to
apply this change until I see what's better for the models.

Thanks,
Gabriele

> 
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 639b5df7cf130..91a15a4234280 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -7143,9 +7143,6 @@ static void __sched notrace __schedule(int sched_mode)
>  	struct rq *rq;
>  	int cpu;
>  
> -	/* Trace preemptions consistently with task switches */
> -	trace_sched_entry_tp(sched_mode == SM_PREEMPT);
> -
>  	cpu = smp_processor_id();
>  	rq = cpu_rq(cpu);
>  	prev = rq->curr;
> @@ -7178,6 +7175,9 @@ static void __sched notrace __schedule(int sched_mode)
>  	rq_lock(rq, &rf);
>  	smp_mb__after_spinlock();
>  
> +	/* Trace preemptions consistently with task switches */
> +	trace_sched_entry_tp(sched_mode == SM_PREEMPT);
> +
>  	hrtick_schedule_enter(rq);
>  
>  	/* Promote REQ to ACT */
> 
> 
> > 
> > Fixes: adcc3bfa8806 ("sched: Adapt sched tracepoints for RV task model")
> > Signed-off-by: Sechang Lim <rhkrqnwk98@gmail.com>
> > ---
> > v3:
> >  - reorder need_ipi variable. (K Prateek Nayak)
> > 
> > v2:
> >  - https://lore.kernel.org/all/20260627081657.499781-1-rhkrqnwk98@gmail.com/
> > 
> > v1:
> >  - https://lore.kernel.org/all/20260625065656.392182-1-rhkrqnwk98@gmail.com/
> > 
> >  include/linux/sched.h | 5 ++---
> >  kernel/sched/core.c   | 7 +++++--
> >  2 files changed, 7 insertions(+), 5 deletions(-)
> > 
> > diff --git a/include/linux/sched.h b/include/linux/sched.h
> > index ee06cba5c6f5..c9efd08dae92 100644
> > --- a/include/linux/sched.h
> > +++ b/include/linux/sched.h
> > @@ -2071,10 +2071,9 @@ static inline int test_tsk_thread_flag(struct
> > task_struct *tsk, int flag)
> >  
> >  static inline void set_tsk_need_resched(struct task_struct *tsk)
> >  {
> > -	if (tracepoint_enabled(sched_set_need_resched_tp) &&
> > -	    !test_tsk_thread_flag(tsk, TIF_NEED_RESCHED))
> > +	if (!test_and_set_tsk_thread_flag(tsk, TIF_NEED_RESCHED) &&
> > +	    tracepoint_enabled(sched_set_need_resched_tp))
> >  		__trace_set_need_resched(tsk, TIF_NEED_RESCHED);
> > -	set_tsk_thread_flag(tsk,TIF_NEED_RESCHED);
> >  }
> >  
> >  static inline void clear_tsk_need_resched(struct task_struct *tsk)
> > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > index b8871449d3c6..19de28f0d85a 100644
> > --- a/kernel/sched/core.c
> > +++ b/kernel/sched/core.c
> > @@ -1171,6 +1171,7 @@ static void __resched_curr(struct rq *rq, int tif)
> >  {
> >  	struct task_struct *curr = rq->curr;
> >  	struct thread_info *cti = task_thread_info(curr);
> > +	bool need_ipi;
> >  	int cpu;
> >  
> >  	lockdep_assert_rq_held(rq);
> > @@ -1187,15 +1188,17 @@ static void __resched_curr(struct rq *rq, int tif)
> >  
> >  	cpu = cpu_of(rq);
> >  
> > -	trace_sched_set_need_resched_tp(curr, cpu, tif);
> >  	if (cpu == smp_processor_id()) {
> >  		set_ti_thread_flag(cti, tif);
> >  		if (tif == TIF_NEED_RESCHED)
> >  			set_preempt_need_resched();
> > +		trace_sched_set_need_resched_tp(curr, cpu, tif);
> >  		return;
> >  	}
> >  
> > -	if (set_nr_and_not_polling(cti, tif)) {
> > +	need_ipi = set_nr_and_not_polling(cti, tif);
> > +	trace_sched_set_need_resched_tp(curr, cpu, tif);
> > +	if (need_ipi) {
> >  		if (tif == TIF_NEED_RESCHED)
> >  			smp_send_reschedule(cpu);
> >  	} else {
> > -- 
> > 2.43.0
> > 


  reply	other threads:[~2026-09-29  8:50 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-30  8:47 Sechang Lim
2026-07-03 15:33 ` Gabriele Monaco
2026-09-28 17:19 ` Andrea Righi
2026-09-29  8:50   ` Gabriele Monaco [this message]
2026-09-29  9:30     ` Andrea Righi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=96b3b5d5ec0eae499f1a98845720776b6fab84d4.camel@redhat.com \
    --to=gmonaco@redhat.com \
    --cc=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rhkrqnwk98@gmail.com \
    --cc=rostedt@goodmis.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®