mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/1] Fix an HWCR RMW race
@ 2026-06-09 21:15 Jim Mattson
  2026-06-09 21:15 ` [PATCH 1/1] x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting() Jim Mattson
  0 siblings, 1 reply; 5+ messages in thread
From: Jim Mattson @ 2026-06-09 21:15 UTC (permalink / raw)
  To: Borislav Petkov, Thomas Gleixner, x86, linux-kernel, yosry; +Cc: Jim Mattson

I was backporting commit 65f55a301766 ("x86/CPU/AMD: Add CPUID faulting
support") to a local branch based on Linux v6.12, when our internal Sashiko
asked:

> Can this corrupt MSR_K7_HWCR? disable_cpuid()->set_cpuid_faulting() is
> called with preemption disabled, but interrupts are still enabled. Since
> msr_set_bit() performs a read-modify-write without disabling interrupts,
> if an IPI arrives between the read and write and modifies MSR_K7_HWCR
> (e.g. acpi-cpufreq toggling Core Performance Boost), the IPI's update
> will be lost.

To confirm that this wasn't just AI slop, I set up an empirical test on a
Turin system. First, I replaced the amd-pstate cpufreq driver with
acpi-cpufreq. Then I ran a test program, where one thread repeatedly reads
CPU0's HWCR, toggles /sys/devices/system/cpu/cpufreq/boost, reads CPU0's
HWCR again, and then verifies that the CPB_DIS bit has flipped. A second
thread, pinned to CPU0, repeatedly calls arch_prctl(ARCH_SET_CPUID, <val>),
where <val> alternates between 0 and 1. With the second thread running, the
first thread soon fails the verification step, indicating that the CPB_DIS
bit change is, in fact, lost.

Jim Mattson (1):
  x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting()

 arch/x86/kernel/process.c | 4 ++++
 1 file changed, 4 insertions(+)


base-commit: 2d3090a8aeb596a26935db0955d46c9a5db5c6ce
-- 
2.54.0.1099.g489fc7bff1-goog


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH 1/1] x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting()
  2026-06-09 21:15 [PATCH 0/1] Fix an HWCR RMW race Jim Mattson
@ 2026-06-09 21:15 ` Jim Mattson
  2026-06-09 21:41   ` Borislav Petkov
  0 siblings, 1 reply; 5+ messages in thread
From: Jim Mattson @ 2026-06-09 21:15 UTC (permalink / raw)
  To: Borislav Petkov, Thomas Gleixner, x86, linux-kernel, yosry; +Cc: Jim Mattson

Since msr_set_bit() and msr_clear_bit() perform a non-atomic update to an
MSR, they can race with a write to the same MSR from interrupt context.  On
AMD CPUs, set_cpuid_faulting() uses these functions to modify MSR_K7_HWCR
from process context, and boost_set_msr() modifies MSR_K7_HWCR from
interrupt context.

To prevent the race, disable interrupts on the AMD path through
set_cpuid_faulting(). Note that when set_cpuid_faulting() is called from
__switch_to_xtra(), interrupts are already disabled.

Reported-by: Sashiko (gemini/gemini-3.1-pro-preview)
Fixes: 65f55a301766 ("x86/CPU/AMD: Add CPUID faulting support")
Signed-off-by: Jim Mattson <jmattson@google.com>
---
 arch/x86/kernel/process.c | 4 ++++
 1 file changed, 4 insertions(+)

diff --git a/arch/x86/kernel/process.c b/arch/x86/kernel/process.c
index 4c718f8adc59..92492a63108f 100644
--- a/arch/x86/kernel/process.c
+++ b/arch/x86/kernel/process.c
@@ -354,10 +354,14 @@ static void set_cpuid_faulting(bool on)
 		this_cpu_write(msr_misc_features_shadow, msrval);
 		wrmsrq(MSR_MISC_FEATURES_ENABLES, msrval);
 	} else if (boot_cpu_data.x86_vendor == X86_VENDOR_AMD) {
+		unsigned long flags;
+
+		local_irq_save(flags);
 		if (on)
 			msr_set_bit(MSR_K7_HWCR, MSR_K7_HWCR_CPUID_USER_DIS_BIT);
 		else
 			msr_clear_bit(MSR_K7_HWCR, MSR_K7_HWCR_CPUID_USER_DIS_BIT);
+		local_irq_restore(flags);
 	}
 }
 
-- 
2.54.0.1099.g489fc7bff1-goog


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH 1/1] x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting()
  2026-06-09 21:15 ` [PATCH 1/1] x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting() Jim Mattson
@ 2026-06-09 21:41   ` Borislav Petkov
  2026-06-09 22:08     ` Jim Mattson
  0 siblings, 1 reply; 5+ messages in thread
From: Borislav Petkov @ 2026-06-09 21:41 UTC (permalink / raw)
  To: Jim Mattson; +Cc: Thomas Gleixner, x86, linux-kernel, yosry

On Tue, Jun 09, 2026 at 02:15:47PM -0700, Jim Mattson wrote:
> +		unsigned long flags;
> +
> +		local_irq_save(flags);
>  		if (on)
>  			msr_set_bit(MSR_K7_HWCR, MSR_K7_HWCR_CPUID_USER_DIS_BIT);
>  		else
>  			msr_clear_bit(MSR_K7_HWCR, MSR_K7_HWCR_CPUID_USER_DIS_BIT);
> +		local_irq_restore(flags);
>  	}

Can we carve out this logic that modifies bits in HWCR into a separate
function and call it from everywhere it needs so that it is very obvious that
synchronization needs to happen?

Thx.

-- 
Regards/Gruss,
    Boris.

https://people.kernel.org/tglx/notes-about-netiquette

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH 1/1] x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting()
  2026-06-09 21:41   ` Borislav Petkov
@ 2026-06-09 22:08     ` Jim Mattson
  2026-06-09 23:27       ` Borislav Petkov
  0 siblings, 1 reply; 5+ messages in thread
From: Jim Mattson @ 2026-06-09 22:08 UTC (permalink / raw)
  To: Borislav Petkov; +Cc: Thomas Gleixner, x86, linux-kernel, yosry

()

On Tue, Jun 9, 2026 at 2:41 PM Borislav Petkov <bp@alien8.de> wrote:
>
> On Tue, Jun 09, 2026 at 02:15:47PM -0700, Jim Mattson wrote:
> > +             unsigned long flags;
> > +
> > +             local_irq_save(flags);
> >               if (on)
> >                       msr_set_bit(MSR_K7_HWCR, MSR_K7_HWCR_CPUID_USER_DIS_BIT);
> >               else
> >                       msr_clear_bit(MSR_K7_HWCR, MSR_K7_HWCR_CPUID_USER_DIS_BIT);
> > +             local_irq_restore(flags);
> >       }
>
> Can we carve out this logic that modifies bits in HWCR into a separate
> function and call it from everywhere it needs so that it is very obvious that
> synchronization needs to happen?

Yes, I'd be happy to.

Most HWCR updates are in init code, so they should be fine. Should I
call the helper from init code anyway?

The one potential HWCR update race I am quite uncertain about is:

> static int cpufreq_boost_down_prep(unsigned int cpu)
> {
>         /*
>          * Clear the boost-disable bit on the CPU_DOWN path so that
>          * this cpu cannot block the remaining ones from boosting.
>          */
>         return boost_set_msr(1);
> }

If this can race with boost_set_msr_each(), then the bigger problem is
that boost_set_msr_each() might set the boost-disable bit between the
call to cpufreq_boost_down_prep() and the CPU going offline.

I couldn't convince myself that cpufreq_boost_down_prep() was always
called on the CPU going offline. The call to boost_set_msr() on the
current CPU suggests it is, but the 'cpu' argument suggests it isn't.
If policies are shared, my reading of __cpufreq_offline() is that
cpufreq_driver->exit() [which calls cpufreq_boost_down_prep()] is only
called when the last CPU sharing the policy goes offline, potentially
leaving CPB_DIS set on the rest.

Anyway, since boost_set_msr() is called from both process context and
interrupt context, perhaps I should use the helper in boost_set_msr()
as well, just for consistency.

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH 1/1] x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting()
  2026-06-09 22:08     ` Jim Mattson
@ 2026-06-09 23:27       ` Borislav Petkov
  0 siblings, 0 replies; 5+ messages in thread
From: Borislav Petkov @ 2026-06-09 23:27 UTC (permalink / raw)
  To: Jim Mattson; +Cc: Thomas Gleixner, x86, linux-kernel, yosry

On Tue, Jun 09, 2026 at 03:08:22PM -0700, Jim Mattson wrote:
> Yes, I'd be happy to.

Thanks!

I haven't heard that in a while - it is usually grumbling :-P

> Most HWCR updates are in init code, so they should be fine. Should I
> call the helper from init code anyway?

Pfff, looking around the tree, we've accumulated a lot of HWCR pokers... I'm
thinking perhaps we should leave the obviously non-problematic ones alone for
now and convert only the problematic ones at first and then slowly, convert
the rest. Otherwise, that'll be a bunch of patch churn at once...

> If this can race with boost_set_msr_each(),

You mean if someone echoes into sysfs to toggle boost and yanks the module
at the same time? Pff, I fail to see the valid use case here but I won't be
surprised.

> then the bigger problem is that boost_set_msr_each() might set the
> boost-disable bit between the call to cpufreq_boost_down_prep() and the CPU
> going offline.

Offline? I think that's the module exit callback acpi_cpufreq_cpu_exit() which
calls it.

/me greps some more...

Uff, that's hit down the cpuhp_cpufreq_offline() path which is the hotplug
callback.

> I couldn't convince myself that cpufreq_boost_down_prep() was always
> called on the CPU going offline. The call to boost_set_msr() on the
> current CPU suggests it is, but the 'cpu' argument suggests it isn't.
> If policies are shared, my reading of __cpufreq_offline() is that
> cpufreq_driver->exit() [which calls cpufreq_boost_down_prep()] is only
> called when the last CPU sharing the policy goes offline, potentially
> leaving CPB_DIS set on the rest.

Yah, I see it the same. Unless we both are missing something, this looks
broken and prolly no one caught it because we're using amd-state now. And that
thing doesn't even touch CPB_DIS but does this pm_qos stuff in
amd_pstate_cpu_boost_update().

Dunno, do we care about a driver we don't use anymore... maybe, but low
prio...

> Anyway, since boost_set_msr() is called from both process context and
> interrupt context, perhaps I should use the helper in boost_set_msr()
> as well, just for consistency.

Yah, stick the call in the case ...HYGON/...AMD: branch and should be good.

I can test it then on some machines here.

Thx.

-- 
Regards/Gruss,
    Boris.

https://people.kernel.org/tglx/notes-about-netiquette

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-06-09 23:28 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-06-09 21:15 [PATCH 0/1] Fix an HWCR RMW race Jim Mattson
2026-06-09 21:15 ` [PATCH 1/1] x86/CPU/AMD: Avoid racy updates to MSR_K7_HWCR in set_cpuid_faulting() Jim Mattson
2026-06-09 21:41   ` Borislav Petkov
2026-06-09 22:08     ` Jim Mattson
2026-06-09 23:27       ` Borislav Petkov

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®