* [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu()
@ 2026-09-01 15:22 Michal Blaszczyk
2026-09-01 16:20 ` Andrea Righi
0 siblings, 1 reply; 3+ messages in thread
From: Michal Blaszczyk @ 2026-09-01 15:22 UTC (permalink / raw)
To: Tejun Heo, David Vernet, Andrea Righi, Changwoo Min
Cc: Michal Blaszczyk, Kuba Piecuch, sched-ext, linux-kernel
In scx_idle_test_and_clear_cpu(), the shared idle_smts mask is modified
locklessly by concurrent CPUs. Currently, the code uses
__cpumask_clear_cpu() to clear a CPU from the mask. Because this is
a non-atomic read-modify-write operation, concurrent modifications to
different bits within the same memory word can lead to data races and
lost updates.
Fix this by replacing __cpumask_clear_cpu() with the atomic
cpumask_clear_cpu().
Fixes: 48849271e661 ("sched_ext: idle: Per-node idle cpumasks")
Signed-off-by: Michal Blaszczyk <michalblk@google.com>
---
kernel/sched/ext/idle.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
index d2973fb3af6d..8985b48c83a5 100644
--- a/kernel/sched/ext/idle.c
+++ b/kernel/sched/ext/idle.c
@@ -104,7 +104,7 @@ static bool scx_idle_test_and_clear_cpu(int cpu)
if (cpumask_intersects(smt, idle_smts))
cpumask_andnot(idle_smts, idle_smts, smt);
else if (cpumask_test_cpu(cpu, idle_smts))
- __cpumask_clear_cpu(cpu, idle_smts);
+ cpumask_clear_cpu(cpu, idle_smts);
}
return cpumask_test_and_clear_cpu(cpu, idle_cpus);
--
2.55.0.897.gb25b4bd76c-goog
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu()
2026-09-01 15:22 [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu() Michal Blaszczyk
@ 2026-09-01 16:20 ` Andrea Righi
2026-09-02 8:07 ` Michał Błaszczyk
0 siblings, 1 reply; 3+ messages in thread
From: Andrea Righi @ 2026-09-01 16:20 UTC (permalink / raw)
To: Michal Blaszczyk
Cc: Tejun Heo, David Vernet, Changwoo Min, Kuba Piecuch, sched-ext,
linux-kernel
Hi Michal,
On Tue, Sep 01, 2026 at 03:22:12PM +0000, Michal Blaszczyk wrote:
> In scx_idle_test_and_clear_cpu(), the shared idle_smts mask is modified
> locklessly by concurrent CPUs. Currently, the code uses
> __cpumask_clear_cpu() to clear a CPU from the mask. Because this is
> a non-atomic read-modify-write operation, concurrent modifications to
> different bits within the same memory word can lead to data races and
> lost updates.
>
> Fix this by replacing __cpumask_clear_cpu() with the atomic
> cpumask_clear_cpu().
I think the atomic clear makes sense here, but, as sashiko also pointed out, it
does not fully address the race, because idle_smts is also modified by the
non-atomic cpumask_andnot() below and cpumask_or() in update_builtin_idle().
>
> Fixes: 48849271e661 ("sched_ext: idle: Per-node idle cpumasks")
And the race existed way before this commit, the idle SMT tracking has been
always documented as racy and self-correcting.
This change may still be a best-effort improvement, but the commit message
should describe it in this way. Did you notice any improvements/benefits with
some workloads with this patch applied?
Thanks,
-Andrea
> Signed-off-by: Michal Blaszczyk <michalblk@google.com>
> ---
> kernel/sched/ext/idle.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
> index d2973fb3af6d..8985b48c83a5 100644
> --- a/kernel/sched/ext/idle.c
> +++ b/kernel/sched/ext/idle.c
> @@ -104,7 +104,7 @@ static bool scx_idle_test_and_clear_cpu(int cpu)
> if (cpumask_intersects(smt, idle_smts))
> cpumask_andnot(idle_smts, idle_smts, smt);
> else if (cpumask_test_cpu(cpu, idle_smts))
> - __cpumask_clear_cpu(cpu, idle_smts);
> + cpumask_clear_cpu(cpu, idle_smts);
> }
>
> return cpumask_test_and_clear_cpu(cpu, idle_cpus);
> --
> 2.55.0.897.gb25b4bd76c-goog
>
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu()
2026-09-01 16:20 ` Andrea Righi
@ 2026-09-02 8:07 ` Michał Błaszczyk
0 siblings, 0 replies; 3+ messages in thread
From: Michał Błaszczyk @ 2026-09-02 8:07 UTC (permalink / raw)
To: Andrea Righi
Cc: Tejun Heo, David Vernet, Changwoo Min, Kuba Piecuch, sched-ext,
linux-kernel
Hi Andrea,
On Tue, Sep 1, 2026 at 6:20 PM Andrea Righi <arighi@nvidia.com> wrote:
> I think the atomic clear makes sense here, but, as sashiko also pointed out, it
> does not fully address the race, because idle_smts is also modified by the
> non-atomic cpumask_andnot() below and cpumask_or() in update_builtin_idle().
>
> > Fixes: 48849271e661 ("sched_ext: idle: Per-node idle cpumasks")
>
> And the race existed way before this commit, the idle SMT tracking has been
> always documented as racy and self-correcting.
>
> This change may still be a best-effort improvement, but the commit message
> should describe it in this way. Did you notice any improvements/benefits with
> some workloads with this patch applied?
I was investigating an automated static analysis report from Sashiko
which flagged the use of __cpumask_clear_cpu() as a concurrency bug
that could lead to lost updates. I mistakenly assumed this specific
call was an isolated oversight and that the other surrounding cpumask_*
operations were properly atomic.
I hadn't initially realized that cpumask_andnot() and cpumask_or() are
also non-atomic bulk operations. After reading your explanation and
looking closer at the code, it's clear to me now that this non-atomic
handling was intentional to avoid lock overhead on the fast path.
Since my original intent was to fix what I incorrectly thought was a
strict logical bug, and not a profiled optimization, I think it makes
the most sense to just drop this patch. Making just this one operation
atomic while the rest of the mask is manipulated non-atomically would
just be inconsistent and add unnecessary overhead.
Thank you for taking the time to look at this and explain the
context to me.
Best,
Michal
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-02 8:07 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-01 15:22 [PATCH] sched_ext: Use atomic cpumask_clear_cpu in scx_idle_test_and_clear_cpu() Michal Blaszczyk
2026-09-01 16:20 ` Andrea Righi
2026-09-02 8:07 ` Michał Błaszczyk
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®