mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "Edgecombe, Rick P" <rick.p.edgecombe@intel.com>
To: "Yu, Yu-cheng" <yu-cheng.yu@intel.com>,
	"luto@kernel.org" <luto@kernel.org>,
	"hitesh.murali@imperva.com" <hitesh.murali@imperva.com>,
	"bp@alien8.de" <bp@alien8.de>,
	"peterz@infradead.org" <peterz@infradead.org>,
	"x86@kernel.org" <x86@kernel.org>,
	"mingo@redhat.com" <mingo@redhat.com>,
	"rppt@kernel.org" <rppt@kernel.org>,
	"dave.hansen@linux.intel.com" <dave.hansen@linux.intel.com>,
	"tglx@kernel.org" <tglx@kernel.org>,
	"hpa@zytor.com" <hpa@zytor.com>
Cc: "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] x86/mm/fault: Test _PAGE_RW, not pte_write(), in spurious_kernel_fault_check()
Date: Thu, 1 Oct 2026 16:34:17 +0000	[thread overview]
Message-ID: <10955849251def44fdc006926b78b1dc39301032.camel@intel.com> (raw)
In-Reply-To: <20261001-x86-fault-spurious-rw-v1-1-7fe1189b5efd@imperva.com>

On Thu, 2026-10-01 at 19:12 +0300, Hitesh Murali via B4 Relay wrote:
> From: Hitesh Murali <hitesh.murali@imperva.com>
> 
> spurious_kernel_fault() only considers faults whose error code is exactly
> X86_PF_WRITE | X86_PF_PROT or X86_PF_INSTR | X86_PF_PROT. A write fault
> that reaches spurious_kernel_fault_check() is therefore a normal store,
> never a shadow stack access, and with CR0.WP set, which the kernel pins,
> the hardware permits a normal store only if _PAGE_RW is set. The check is
> applied to 4K, 2M and 1G leaves and to the PMD table entry, and _PAGE_RW
> means the same at each of these levels.
> 
> Commit bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> made pte_write() return true for Write=0,Dirty=1 entries, the encoding of
> shadow stack memory, so that core mm treats shadow stacks as writable.
> pte_shstk() decides this from X86_FEATURE_SHSTK alone, so it applies on
> every shadow stack capable CPU, also with CONFIG_X86_USER_SHADOW_STACK=n.
> For a kernel mapping the hardware does not agree: the kernel never sets
> IA32_S_CET.SH_STK_EN, so a supervisor Write=0,Dirty=1 entry is an
> ordinary read-only entry.

I've always wondered if we really need spurious_kernel_fault(). Does anyone know
what functionality depends on it? Also, always seemed strange that it works from
user accesses to the kernel half of the address space.

> 
> Kernel read-only mappings are created without Dirty since commit
> f788b71768ff ("x86/mm: Remove _PAGE_DIRTY from kernel RO pages"), but a
> store made while _PAGE_RW is temporarily set leaves Dirty behind, and the
> kernel does not clear it again.

The intention was to prevent RO,Dirty mappings all together unless they were
really intended to be shadow stack. So maybe that is the right fix. Do you know
what flow left the PTE in this state? It was this test module? Or something in
the upstream kernel?

>  A later normal store to such an entry
> raises a protection fault, which is then classified as spurious. The
> store is restarted, faults again, and the CPU makes no progress.
> 
> When the soft lockup watchdog reports it, it reports a stuck CPU at the
> store, which reads as a long running loop rather than as a write to a
> read-only page. This was found with an out-of-tree module that writes to
> .rodata; on a distribution kernel carrying bb3aadf7d446, on Sapphire
> Rapids, one CPU looped on the store for days.
> 
> Test _PAGE_RW directly. The fault then takes the regular
> bad_area_nosemaphore() path: an exception table fixup where the access
> has one, as for any other read-only page, and otherwise an oops that
> reports the faulting address and the page table entry.
> 
> Fixes: bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> Assisted-by: LLM
> Signed-off-by: Hitesh Murali <hitesh.murali@imperva.com>
> Cc: stable@vger.kernel.org



      parent reply	other threads:[~2026-10-01 16:34 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 16:12 Hitesh Murali via B4 Relay
2026-10-01 16:24 ` Dave Hansen
     [not found]   ` <MR0P264MB69837D7B9DF99BBA6333DB60E68A2@MR0P264MB6983.FRAP264.PROD.OUTLOOK.COM>
2026-10-01 20:37     ` Dave Hansen
2026-10-01 16:34 ` Edgecombe, Rick P [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=10955849251def44fdc006926b78b1dc39301032.camel@intel.com \
    --to=rick.p.edgecombe@intel.com \
    --cc=bp@alien8.de \
    --cc=dave.hansen@linux.intel.com \
    --cc=hitesh.murali@imperva.com \
    --cc=hpa@zytor.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=luto@kernel.org \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rppt@kernel.org \
    --cc=tglx@kernel.org \
    --cc=x86@kernel.org \
    --cc=yu-cheng.yu@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®