mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check()
@ 2026-10-01 16:12 Hitesh Murali via B4 Relay
  2026-10-01 16:24 ` Dave Hansen
  2026-10-01 16:34 ` Edgecombe, Rick P
  0 siblings, 2 replies; 4+ messages in thread
From: Hitesh Murali via B4 Relay @ 2026-10-01 16:12 UTC (permalink / raw)
  To: Dave Hansen, Andy Lutomirski, Peter Zijlstra, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, x86, H. Peter Anvin,
	Rick Edgecombe, Yu-cheng Yu, Mike Rapoport (IBM)
  Cc: linux-kernel, Hitesh Murali

From: Hitesh Murali <hitesh.murali@imperva.com>

spurious_kernel_fault() only considers faults whose error code is exactly
X86_PF_WRITE | X86_PF_PROT or X86_PF_INSTR | X86_PF_PROT. A write fault
that reaches spurious_kernel_fault_check() is therefore a normal store,
never a shadow stack access, and with CR0.WP set, which the kernel pins,
the hardware permits a normal store only if _PAGE_RW is set. The check is
applied to 4K, 2M and 1G leaves and to the PMD table entry, and _PAGE_RW
means the same at each of these levels.

Commit bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
made pte_write() return true for Write=0,Dirty=1 entries, the encoding of
shadow stack memory, so that core mm treats shadow stacks as writable.
pte_shstk() decides this from X86_FEATURE_SHSTK alone, so it applies on
every shadow stack capable CPU, also with CONFIG_X86_USER_SHADOW_STACK=n.
For a kernel mapping the hardware does not agree: the kernel never sets
IA32_S_CET.SH_STK_EN, so a supervisor Write=0,Dirty=1 entry is an
ordinary read-only entry.

Kernel read-only mappings are created without Dirty since commit
f788b71768ff ("x86/mm: Remove _PAGE_DIRTY from kernel RO pages"), but a
store made while _PAGE_RW is temporarily set leaves Dirty behind, and the
kernel does not clear it again. A later normal store to such an entry
raises a protection fault, which is then classified as spurious. The
store is restarted, faults again, and the CPU makes no progress.

When the soft lockup watchdog reports it, it reports a stuck CPU at the
store, which reads as a long running loop rather than as a write to a
read-only page. This was found with an out-of-tree module that writes to
.rodata; on a distribution kernel carrying bb3aadf7d446, on Sapphire
Rapids, one CPU looped on the store for days.

Test _PAGE_RW directly. The fault then takes the regular
bad_area_nosemaphore() path: an exception table fixup where the access
has one, as for any other read-only page, and otherwise an oops that
reports the faulting address and the page table entry.

Fixes: bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
Assisted-by: LLM
Signed-off-by: Hitesh Murali <hitesh.murali@imperva.com>
Cc: stable@vger.kernel.org
---
Observed on a distribution kernel carrying bb3aadf7d446 (RHEL 9.8,
5.14.0-687.10.1.el9_8) on Sapphire Rapids, with a proprietary out-of-tree
module storing into sys_call_table:

  watchdog: BUG: soft lockup - CPU#53 stuck for 23s! [<task>:5016]
  RIP: 0010:<module store function>+0x1d/0x30 [<out-of-tree module>]
    faulting insn: mov %r8,(%rsi,%rdx,8)       (plain store, not WRSS)
  CR0: 0000000080050033   (WP=1)
  CR2: ffffffff96c02250   (= sys_call_table + 0x2a * 8)
  CR4: 0000000000f73ef0   (CET=1)
  Kernel panic - not syncing: softlockup: hung tasks
                                  (softlockup_panic=1 on this host)

  vmcore page table walk of CR2:
  PMD: 8000004af5e001e1 -> 2MB page, PRESENT|ACCESSED|DIRTY|PSE|GLOBAL|NX

Reproduced on mainline v7.3-rc5-37-g551c722f4080 without this patch, on
an Amazon EC2 m7i.metal-24xl (Xeon Platinum 8488C, user shadow stack
enabled, CR4.CET=1). Only two small GPL test modules were loaded
(taint O, E). They are available on request.

1) A test module sets _PAGE_RW on the live leaf that maps sys_call_table,
   stores an entry's own value back, and clears _PAGE_RW again. As read
   back through lookup_address():

     before: level=2M flags=0x80000000000001a1 RW=0 DIRTY=0 pte_write()=0
     after:  level=2M flags=0x80000000000001e1 RW=0 DIRTY=1 pte_write()=1

2) A second module stores sys_call_table[39]'s own value back (a plain
   MOV; sys_call_table is no longer used for x86-64 dispatch, so the
   store would be harmless if it completed). It never completes. Nine
   minutes later the task was still running, with CPU time equal to wall
   time, and the leaf was unchanged. An NMI backtrace (sysrq-l), trimmed:

     RIP: 0010:sct_store_init+0xaa/0xff0 [sct_store]
     Code: ... <48> 89 83 38 01 00 00     mov %rax,0x138(%rbx)
     RBX: ffffffffa2a029c0 (sys_call_table)
     CR0: 0000000080050033 (WP=1)  CR2: ffffffffa2a02af8  CR4: 0000000000f73ef0 (CET=1)
     Call Trace:
      do_one_initcall
      do_init_module
      init_module_from_file
      __x64_sys_finit_module

   CR2 is sys_call_table + 39 * 8, the slot being written: the same
   leaf state, the same kind of store, and the same fault address as on
   the distribution kernel.

Testing:
 - W=1 build of arch/x86/mm/fault.o on v7.3-rc5-37-g551c722f4080, no
   warnings. The patched kernel was built from the same tree and config.
 - The patched kernel has not been boot or runtime tested yet. The
   expected result is an oops at the store ("supervisor write access in
   kernel mode", "permissions violation") instead of the loop.
---
 arch/x86/mm/fault.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
index aa88370ce739..84298cb38fae 100644
--- a/arch/x86/mm/fault.c
+++ b/arch/x86/mm/fault.c
@@ -955,7 +955,13 @@ do_sigbus(struct pt_regs *regs, unsigned long error_code, unsigned long address,
 
 static int spurious_kernel_fault_check(unsigned long error_code, pte_t *pte)
 {
-	if ((error_code & X86_PF_WRITE) && !pte_write(*pte))
+	/*
+	 * Only normal stores get here: the caller's exact error code match
+	 * filters out shadow stack accesses, and a normal store requires
+	 * _PAGE_RW. Do not use pte_write(), which also reports
+	 * Write=0,Dirty=1 as writable.
+	 */
+	if ((error_code & X86_PF_WRITE) && !(pte_flags(*pte) & _PAGE_RW))
 		return 0;
 
 	if ((error_code & X86_PF_INSTR) && !pte_exec(*pte))

---
base-commit: 551c722f40809618230001baccf219193e22fc5a
change-id: 20260930-x86-fault-spurious-rw-89333ae30644

Best regards,
--  
Hitesh Murali <hitesh.murali@imperva.com>



^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check()
  2026-10-01 16:12 [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check() Hitesh Murali via B4 Relay
@ 2026-10-01 16:24 ` Dave Hansen
       [not found]   ` <MR0P264MB69837D7B9DF99BBA6333DB60E68A2@MR0P264MB6983.FRAP264.PROD.OUTLOOK.COM>
  2026-10-01 16:34 ` Edgecombe, Rick P
  1 sibling, 1 reply; 4+ messages in thread
From: Dave Hansen @ 2026-10-01 16:24 UTC (permalink / raw)
  To: hitesh.murali, Dave Hansen, Andy Lutomirski, Peter Zijlstra,
	Thomas Gleixner, Ingo Molnar, Borislav Petkov, x86,
	H. Peter Anvin, Rick Edgecombe, Yu-cheng Yu, Mike Rapoport (IBM)
  Cc: linux-kernel

On 10/1/26 09:12, Hitesh Murali via B4 Relay wrote:
> Kernel read-only mappings are created without Dirty since commit
> f788b71768ff ("x86/mm: Remove _PAGE_DIRTY from kernel RO pages"), but a
> store made while _PAGE_RW is temporarily set leaves Dirty behind, and the
> kernel does not clear it again. A later normal store to such an entry
> raises a protection fault, which is then classified as spurious. The
> store is restarted, faults again, and the CPU makes no progress.

I'm having a really hard time parsing through this changelog. Was this
all from the LLM?

Can this issue actually be triggered in a normal kernel? The root of
this problem is that PTE-modifying kernel code is leaving _PAGE_DIRTY
behind during a RW=>RO transition. We have code to prevent that from
happening.

So are the out-of-tree modules just being naughty and directly
manipulating page tables? Care to post your "small GPL test modules"?

Yeah, there still might be a buglet in spurious fault detection that
would surface in the face of _other_ kernel bugs, but it'd be extra low
priority.

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check()
  2026-10-01 16:12 [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check() Hitesh Murali via B4 Relay
  2026-10-01 16:24 ` Dave Hansen
@ 2026-10-01 16:34 ` Edgecombe, Rick P
  1 sibling, 0 replies; 4+ messages in thread
From: Edgecombe, Rick P @ 2026-10-01 16:34 UTC (permalink / raw)
  To: Yu, Yu-cheng, luto, hitesh.murali, bp, peterz, x86, mingo, rppt,
	dave.hansen, tglx, hpa
  Cc: linux-kernel

On Thu, 2026-10-01 at 19:12 +0300, Hitesh Murali via B4 Relay wrote:
> From: Hitesh Murali <hitesh.murali@imperva.com>
> 
> spurious_kernel_fault() only considers faults whose error code is exactly
> X86_PF_WRITE | X86_PF_PROT or X86_PF_INSTR | X86_PF_PROT. A write fault
> that reaches spurious_kernel_fault_check() is therefore a normal store,
> never a shadow stack access, and with CR0.WP set, which the kernel pins,
> the hardware permits a normal store only if _PAGE_RW is set. The check is
> applied to 4K, 2M and 1G leaves and to the PMD table entry, and _PAGE_RW
> means the same at each of these levels.
> 
> Commit bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> made pte_write() return true for Write=0,Dirty=1 entries, the encoding of
> shadow stack memory, so that core mm treats shadow stacks as writable.
> pte_shstk() decides this from X86_FEATURE_SHSTK alone, so it applies on
> every shadow stack capable CPU, also with CONFIG_X86_USER_SHADOW_STACK=n.
> For a kernel mapping the hardware does not agree: the kernel never sets
> IA32_S_CET.SH_STK_EN, so a supervisor Write=0,Dirty=1 entry is an
> ordinary read-only entry.

I've always wondered if we really need spurious_kernel_fault(). Does anyone know
what functionality depends on it? Also, always seemed strange that it works from
user accesses to the kernel half of the address space.

> 
> Kernel read-only mappings are created without Dirty since commit
> f788b71768ff ("x86/mm: Remove _PAGE_DIRTY from kernel RO pages"), but a
> store made while _PAGE_RW is temporarily set leaves Dirty behind, and the
> kernel does not clear it again.

The intention was to prevent RO,Dirty mappings all together unless they were
really intended to be shadow stack. So maybe that is the right fix. Do you know
what flow left the PTE in this state? It was this test module? Or something in
the upstream kernel?

>  A later normal store to such an entry
> raises a protection fault, which is then classified as spurious. The
> store is restarted, faults again, and the CPU makes no progress.
> 
> When the soft lockup watchdog reports it, it reports a stuck CPU at the
> store, which reads as a long running loop rather than as a write to a
> read-only page. This was found with an out-of-tree module that writes to
> .rodata; on a distribution kernel carrying bb3aadf7d446, on Sapphire
> Rapids, one CPU looped on the store for days.
> 
> Test _PAGE_RW directly. The fault then takes the regular
> bad_area_nosemaphore() path: an exception table fixup where the access
> has one, as for any other read-only page, and otherwise an oops that
> reports the faulting address and the page table entry.
> 
> Fixes: bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> Assisted-by: LLM
> Signed-off-by: Hitesh Murali <hitesh.murali@imperva.com>
> Cc: stable@vger.kernel.org



^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check()
       [not found]   ` <MR0P264MB69837D7B9DF99BBA6333DB60E68A2@MR0P264MB6983.FRAP264.PROD.OUTLOOK.COM>
@ 2026-10-01 20:37     ` Dave Hansen
  0 siblings, 0 replies; 4+ messages in thread
From: Dave Hansen @ 2026-10-01 20:37 UTC (permalink / raw)
  To: K MURALI Hitesh, hitesh.murali, Dave Hansen, Andy Lutomirski,
	Peter Zijlstra, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
	x86, H. Peter Anvin, Rick Edgecombe, Yu-cheng Yu,
	Mike Rapoport (IBM)
  Cc: linux-kernel

On 10/1/26 13:29, K MURALI Hitesh wrote:
> THALES GROUP LIMITED DISTRIBUTION to email recipients

This mail is pretty badly malformed and won't make it to the mailing
list archives anyway.

I'll take your patch as a very good report of a theoretical issue and
squirrel it away to look at on some future rainy day that I have some
spare time. I'll credit you as the bug reporter if it ever gets fixed.

I'm not going to be able to take a patch from you to fix it today,
though. It's not a practical bug that needs fixing.

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-01 20:37 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 16:12 [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check() Hitesh Murali via B4 Relay
2026-10-01 16:24 ` Dave Hansen
     [not found]   ` <MR0P264MB69837D7B9DF99BBA6333DB60E68A2@MR0P264MB6983.FRAP264.PROD.OUTLOOK.COM>
2026-10-01 20:37     ` Dave Hansen
2026-10-01 16:34 ` Edgecombe, Rick P

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®