mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu()
@ 2026-09-01 15:22 Michal Blaszczyk
  2026-09-01 16:20 ` Andrea Righi
  0 siblings, 1 reply; 3+ messages in thread
From: Michal Blaszczyk @ 2026-09-01 15:22 UTC (permalink / raw)
  To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min
  Cc: Michal Blaszczyk, Kuba Piecuch, sched-ext, linux-kernel

In scx_idle_test_and_clear_cpu(), the shared idle_smts mask is modified
locklessly by concurrent CPUs. Currently, the code uses
__cpumask_clear_cpu() to clear a CPU from the mask. Because this is
a non-atomic read-modify-write operation, concurrent modifications to
different bits within the same memory word can lead to data races and
lost updates.

Fix this by replacing __cpumask_clear_cpu() with the atomic
cpumask_clear_cpu().

Fixes: 48849271e661 ("sched_ext: idle: Per-node idle cpumasks")
Signed-off-by: Michal Blaszczyk <michalblk@google.com>
---
 kernel/sched/ext/idle.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
index d2973fb3af6d..8985b48c83a5 100644
--- a/kernel/sched/ext/idle.c
+++ b/kernel/sched/ext/idle.c
@@ -104,7 +104,7 @@ static bool scx_idle_test_and_clear_cpu(int cpu)
 		if (cpumask_intersects(smt, idle_smts))
 			cpumask_andnot(idle_smts, idle_smts, smt);
 		else if (cpumask_test_cpu(cpu, idle_smts))
-			__cpumask_clear_cpu(cpu, idle_smts);
+			cpumask_clear_cpu(cpu, idle_smts);
 	}
 
 	return cpumask_test_and_clear_cpu(cpu, idle_cpus);
-- 
2.55.0.897.gb25b4bd76c-goog


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu()
  2026-09-01 15:22 [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu() Michal Blaszczyk
@ 2026-09-01 16:20 ` Andrea Righi
  2026-09-02  8:07   ` Michał Błaszczyk
  0 siblings, 1 reply; 3+ messages in thread
From: Andrea Righi @ 2026-09-01 16:20 UTC (permalink / raw)
  To: Michal Blaszczyk
  Cc: Tejun Heo, David Vernet, Changwoo Min, Kuba Piecuch, sched-ext,
	linux-kernel

Hi Michal,

On Tue, Sep 01, 2026 at 03:22:12PM +0000, Michal Blaszczyk wrote:
> In scx_idle_test_and_clear_cpu(), the shared idle_smts mask is modified
> locklessly by concurrent CPUs. Currently, the code uses
> __cpumask_clear_cpu() to clear a CPU from the mask. Because this is
> a non-atomic read-modify-write operation, concurrent modifications to
> different bits within the same memory word can lead to data races and
> lost updates.
> 
> Fix this by replacing __cpumask_clear_cpu() with the atomic
> cpumask_clear_cpu().

I think the atomic clear makes sense here, but, as sashiko also pointed out, it
does not fully address the race, because idle_smts is also modified by the
non-atomic cpumask_andnot() below and cpumask_or() in update_builtin_idle().

> 
> Fixes: 48849271e661 ("sched_ext: idle: Per-node idle cpumasks")

And the race existed way before this commit, the idle SMT tracking has been
always documented as racy and self-correcting.

This change may still be a best-effort improvement, but the commit message
should describe it in this way. Did you notice any improvements/benefits with
some workloads with this patch applied?

Thanks,
-Andrea

> Signed-off-by: Michal Blaszczyk <michalblk@google.com>
> ---
>  kernel/sched/ext/idle.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
> 
> diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
> index d2973fb3af6d..8985b48c83a5 100644
> --- a/kernel/sched/ext/idle.c
> +++ b/kernel/sched/ext/idle.c
> @@ -104,7 +104,7 @@ static bool scx_idle_test_and_clear_cpu(int cpu)
>  		if (cpumask_intersects(smt, idle_smts))
>  			cpumask_andnot(idle_smts, idle_smts, smt);
>  		else if (cpumask_test_cpu(cpu, idle_smts))
> -			__cpumask_clear_cpu(cpu, idle_smts);
> +			cpumask_clear_cpu(cpu, idle_smts);
>  	}
>  
>  	return cpumask_test_and_clear_cpu(cpu, idle_cpus);
> -- 
> 2.55.0.897.gb25b4bd76c-goog
> 

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu()
  2026-09-01 16:20 ` Andrea Righi
@ 2026-09-02  8:07   ` Michał Błaszczyk
  0 siblings, 0 replies; 3+ messages in thread
From: Michał Błaszczyk @ 2026-09-02  8:07 UTC (permalink / raw)
  To: Andrea Righi
  Cc: Tejun Heo, David Vernet, Changwoo Min, Kuba Piecuch, sched-ext,
	linux-kernel

Hi Andrea,

On Tue, Sep 1, 2026 at 6:20 PM Andrea Righi <arighi@nvidia.com> wrote:
> I think the atomic clear makes sense here, but, as sashiko also pointed out, it
> does not fully address the race, because idle_smts is also modified by the
> non-atomic cpumask_andnot() below and cpumask_or() in update_builtin_idle().
>
> > Fixes: 48849271e661 ("sched_ext: idle: Per-node idle cpumasks")
>
> And the race existed way before this commit, the idle SMT tracking has been
> always documented as racy and self-correcting.
>
> This change may still be a best-effort improvement, but the commit message
> should describe it in this way. Did you notice any improvements/benefits with
> some workloads with this patch applied?

I was investigating an automated static analysis report from Sashiko
which flagged the use of __cpumask_clear_cpu() as a concurrency bug
that could lead to lost updates. I mistakenly assumed this specific
call was an isolated oversight and that the other surrounding cpumask_*
operations were properly atomic.

I hadn't initially realized that cpumask_andnot() and cpumask_or() are
also non-atomic bulk operations. After reading your explanation and
looking closer at the code, it's clear to me now that this non-atomic
handling was intentional to avoid lock overhead on the fast path.

Since my original intent was to fix what I incorrectly thought was a
strict logical bug, and not a profiled optimization, I think it makes
the most sense to just drop this patch. Making just this one operation
atomic while the rest of the mask is manipulated non-atomically would
just be inconsistent and add unnecessary overhead.

Thank you for taking the time to look at this and explain the
context to me.

Best,
Michal

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-02  8:07 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-01 15:22 [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu() Michal Blaszczyk
2026-09-01 16:20 ` Andrea Righi
2026-09-02  8:07   ` Michał Błaszczyk

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®