mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Shrikanth Hegde <sshegde@linux.ibm.com>
To: Yury Norov <ynorov@nvidia.com>, Yury Norov <yury.norov@gmail.com>
Cc: "Bradley Morgan" <include@grrlz.net>,
	linux-kernel@vger.kernel.org, "Oleg Nesterov" <oleg@redhat.com>,
	"Paul E . McKenney" <paulmck@kernel.org>,
	"Peter Zijlstra" <peterz@infradead.org>,
	"Phil Auld" <pauld@redhat.com>,
	"Sebastian Andrzej Siewior" <bigeasy@linutronix.de>,
	"Tejun Heo" <tj@kernel.org>,
	"Thomas Weißschuh" <thomas.weissschuh@linutronix.de>,
	"Valentin Schneider" <vschneid@redhat.com>
Subject: Re: [PATCH v2] stop_machine: Make stop_one_cpu_nowait() return void
Date: Wed, 29 Jul 2026 21:00:03 +0530	[thread overview]
Message-ID: <b81fad02-ab22-45d8-927a-2beabbbe8a32@linux.ibm.com> (raw)
In-Reply-To: <20260729022355.325058-1-ynorov@nvidia.com>

Hi Yury.

On 7/29/26 7:53 AM, Yury Norov wrote:
> No caller checks the return value from stop_one_cpu_nowait(). All
> callers require the callback to run and arrange for the target CPU's
> stopper to remain enabled while queuing the work. In particular, commit
> f0498d2a54e7 ("sched: Fix stop_one_cpu_nowait() vs hotplug") added
> preemption protection to the scheduler callers so that queuing must
> succeed once the target CPU has been observed online.
> 
> Therefore, a failure is an unrecoverable violation rather than a condition
> individual callers can recover from. Diagnose it with WARN_ON_ONCE() in
> stop_one_cpu_nowait(). A check in the common helper covers current and
> future callers consistently, while individual checks would duplicate
> the same non-recoverable handling at every call site.
> 
> Make the function return void because there is no longer a meaningful
> result for callers to consume.
> 
> On UP, warn if the supplied CPU is not the current CPU because the work
> cannot be scheduled in that case.
> 
> CC: Bradley Morgan <include@grrlz.net>
> Signed-off-by: Yury Norov <ynorov@nvidia.com>
> ---
> v1: https://lore.kernel.org/all/20260724232325.594212-1-ynorov@nvidia.com/
> v2:
>   - carefully include linux/bug.h (Bradley);
>   - update the CONTEXT section (Bradley).
> 
>   include/linux/stop_machine.h | 21 ++++++++++-----------
>   kernel/stop_machine.c        | 14 ++++++--------
>   2 files changed, 16 insertions(+), 19 deletions(-)
> 
> diff --git a/include/linux/stop_machine.h b/include/linux/stop_machine.h
> index 01011113d226..84e7fb627ba4 100644
> --- a/include/linux/stop_machine.h
> +++ b/include/linux/stop_machine.h
> @@ -2,6 +2,7 @@
>   #ifndef _LINUX_STOP_MACHINE
>   #define _LINUX_STOP_MACHINE
>   
> +#include <linux/bug.h>
>   #include <linux/cpu.h>
>   #include <linux/cpumask_types.h>
>   #include <linux/smp.h>
> @@ -31,7 +32,7 @@ struct cpu_stop_work {
>   
>   int stop_one_cpu(unsigned int cpu, cpu_stop_fn_t fn, void *arg);
>   int stop_two_cpus(unsigned int cpu1, unsigned int cpu2, cpu_stop_fn_t fn, void *arg);
> -bool stop_one_cpu_nowait(unsigned int cpu, cpu_stop_fn_t fn, void *arg,
> +void stop_one_cpu_nowait(unsigned int cpu, cpu_stop_fn_t fn, void *arg,
>   			 struct cpu_stop_work *work_buf);
>   void stop_machine_park(int cpu);
>   void stop_machine_unpark(int cpu);
> @@ -68,19 +69,17 @@ static void stop_one_cpu_nowait_workfn(struct work_struct *work)
>   	preempt_enable();
>   }
>   
> -static inline bool stop_one_cpu_nowait(unsigned int cpu,
> +static inline void stop_one_cpu_nowait(unsigned int cpu,
>   				       cpu_stop_fn_t fn, void *arg,
>   				       struct cpu_stop_work *work_buf)
>   {
> -	if (cpu == smp_processor_id()) {
> -		INIT_WORK(&work_buf->work, stop_one_cpu_nowait_workfn);
> -		work_buf->fn = fn;
> -		work_buf->arg = arg;
> -		schedule_work(&work_buf->work);
> -		return true;
> -	}
> -
> -	return false;
> +	if (WARN_ON_ONCE(cpu != smp_processor_id()))
> +		return;
> +
> +	INIT_WORK(&work_buf->work, stop_one_cpu_nowait_workfn);
> +	work_buf->fn = fn;
> +	work_buf->arg = arg;
> +	schedule_work(&work_buf->work);
>   }
>   
>   static inline void print_stop_info(const char *log_lvl, struct task_struct *task) { }
> diff --git a/kernel/stop_machine.c b/kernel/stop_machine.c
> index 773d8e9ae30c..d085ba1f4b44 100644
> --- a/kernel/stop_machine.c
> +++ b/kernel/stop_machine.c
> @@ -7,6 +7,7 @@
>    * Copyright (C) 2010		SUSE Linux Products GmbH
>    * Copyright (C) 2010		Tejun Heo <tj@kernel.org>
>    */
> +#include <linux/bug.h>
>   #include <linux/compiler.h>
>   #include <linux/completion.h>
>   #include <linux/cpu.h>
> @@ -376,17 +377,14 @@ int stop_two_cpus(unsigned int cpu1, unsigned int cpu2, cpu_stop_fn_t fn, void *
>    * and will remain untouched until stopper starts executing @fn.
>    *
>    * CONTEXT:
> - * Don't care.
> - *
> - * RETURNS:
> - * true if cpu_stop_work was queued successfully and @fn will be called,
> - * false otherwise.
> + * Don't care, but the caller must ensure @cpu's stopper stays enabled
> + * until the work is queued, e.g. by preempt_disable().

nit:
* until the work is queued, i.e. ensuring preemption/irq disabled. ?

Because, if irq are disabled it is equally good enough.
#define preemptible()   (preempt_count() == 0 && !irqs_disabled())


>    */
> -bool stop_one_cpu_nowait(unsigned int cpu, cpu_stop_fn_t fn, void *arg,
> -			struct cpu_stop_work *work_buf)
> +void stop_one_cpu_nowait(unsigned int cpu, cpu_stop_fn_t fn, void *arg,
> +			 struct cpu_stop_work *work_buf)
>   {
>   	*work_buf = (struct cpu_stop_work){ .fn = fn, .arg = arg, .caller = _RET_IP_, };
> -	return cpu_stop_queue_work(cpu, work_buf);
> +	WARN_ON_ONCE(!cpu_stop_queue_work(cpu, work_buf));
>   }
>   
>   static bool queue_stop_cpus_work(const struct cpumask *cpumask,

Other than that,
Reviewed-by: Shrikanth Hegde <sshegde@linux.ibm.com>

  parent reply	other threads:[~2026-07-29 15:30 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-29  2:23 Yury Norov
2026-07-29 12:45 ` Bradley Morgan
2026-07-29 12:51 ` Peter Zijlstra
2026-07-29 13:40   ` Yury Norov
2026-07-29 15:30 ` Shrikanth Hegde [this message]
2026-08-03  7:48 ` [tip: sched/core] " tip-bot2 for Yury Norov

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=b81fad02-ab22-45d8-927a-2beabbbe8a32@linux.ibm.com \
    --to=sshegde@linux.ibm.com \
    --cc=bigeasy@linutronix.de \
    --cc=include@grrlz.net \
    --cc=linux-kernel@vger.kernel.org \
    --cc=oleg@redhat.com \
    --cc=pauld@redhat.com \
    --cc=paulmck@kernel.org \
    --cc=peterz@infradead.org \
    --cc=thomas.weissschuh@linutronix.de \
    --cc=tj@kernel.org \
    --cc=vschneid@redhat.com \
    --cc=ynorov@nvidia.com \
    --cc=yury.norov@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome