mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] KVM: x86/mmu: Cope with the SS bit being set in #NPF exit codes
@ 2026-10-09 13:04 Paolo Bonzini
  2026-10-10  5:34 ` kernel test robot
  0 siblings, 1 reply; 2+ messages in thread
From: Paolo Bonzini @ 2026-10-09 13:04 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: SimonP, stable

When running Hyper-V from Windows 11, KVM is seeing the SS bit set
in nested page fault error codes produced for shadowed NPT.  The AMD
manual is unclear about when SS is set, including whether it can be set
when SupervisorShadowStackEn (bit 6) is not set in the MISC_CTL field
of the VMCB; in fact, bisection shows that while SHSTK was available
to guests since Linux 6.18, this started happening with commit
687ee95c1b6d ("KVM: nSVM: enable GMET for guests"), which seems unrelated.

Nevertheless, the unexpected SS bit causes out of bounds accesses in
permission_fault(), seen as either WARNs or UBSAN reports:

  kernel: UBSAN: array-index-out-of-bounds in arch/x86/kvm/mmu.h:225:27
  kernel: index 34 is out of range for type 'u16 [16]'

(in this case, the page fault error code is 34*2=0x44, i.e. SS|U; the
report shows SS|U|W happening as well).  Fix the handling of unknown
error codes by dropping SS and (with a WARN) any bits not among those
that permission_fault() expects to receive.

At the same time, given the uncertainty about when hardware sets SS,
follow the processor's behavior when constructing the fault's error
code, and preserve the SS bit as passed to FNAME(walk_addr_generic).
While the behavior is surprising, there's a possibility that Hyper-V
expects it (Hyper-V does not even enable CET unless it finds MBEC/GMET!),
so retain the information when reflecting the fault to L1 instead of
unconditionally discarding it.

Note that, even though KVM always tries to run L1 with GMET enabled, it
copies the GMET-enable bit of the VMCB12 to the VMCB02 when L1 requests
usage of NPT.  Therefore, if the theory suggested by the bisection result
is correct and GMET enables setting the SS bit as well, this would not
affect hypervisors that enable NPT but not GMET.

Reported-by: SimonP <simonp.git@mailbox.org>
Fixes: 687ee95c1b6d ("KVM: nSVM: enable GMET for guests")
Cc: stable@vger.kernel.org # 7.2+
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 arch/x86/kvm/mmu.h             | 16 ++++++++++++----
 arch/x86/kvm/mmu/paging_tmpl.h |  2 +-
 2 files changed, 13 insertions(+), 5 deletions(-)

diff --git a/arch/x86/kvm/mmu.h b/arch/x86/kvm/mmu.h
index 40dcbcbae173..36a3a42e0b18 100644
--- a/arch/x86/kvm/mmu.h
+++ b/arch/x86/kvm/mmu.h
@@ -275,8 +275,18 @@ static inline u8 permission_fault(struct kvm_vcpu *vcpu, struct kvm_pagewalk *w,
 				  unsigned pte_access, unsigned pte_pkey,
 				  u64 access)
 {
-	/* strip nested paging fault error codes */
-	unsigned int pfec = access;
+	/*
+	 * Of the error codes that do not contribute to the index in
+	 * fmt->permissions, PK is not included in EPT violation bits and
+	 * not supported by shadow paging, and RSVD faults are handled
+	 * elsewhere.  SS however is included in the NPT exit code.
+	 */
+	unsigned int pfec = access & ~PFERR_SS_MASK;
+	if (WARN_ON_ONCE(pfec & ~(PFERR_PRESENT_MASK | PFERR_WRITE_MASK |
+				  PFERR_USER_MASK | PFERR_FETCH_MASK)))
+		pfec &= (PFERR_PRESENT_MASK | PFERR_WRITE_MASK |
+			 PFERR_USER_MASK | PFERR_FETCH_MASK);
+
 	unsigned long rflags = kvm_x86_call(get_rflags)(vcpu);
 
 	/*
@@ -301,8 +311,6 @@ static inline u8 permission_fault(struct kvm_vcpu *vcpu, struct kvm_pagewalk *w,
 	kvm_mmu_refresh_passthrough_bits(vcpu, w);
 
 	fault = (fmt->permissions[index] >> pte_access) & 1;
-
-	WARN_ON_ONCE(pfec & (PFERR_PK_MASK | PFERR_SS_MASK | PFERR_RSVD_MASK));
 	if (unlikely(fmt->pkru_mask)) {
 		u32 pkru_bits, offset;
 
diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h
index dc0155ed1cbf..a1afe2ae28ca 100644
--- a/arch/x86/kvm/mmu/paging_tmpl.h
+++ b/arch/x86/kvm/mmu/paging_tmpl.h
@@ -504,7 +504,7 @@ static int FNAME(walk_addr_generic)(struct guest_walker *walker,
 	return 1;
 
 error:
-	errcode |= write_fault | user_fault;
+	errcode |= access & (PFERR_WRITE_MASK | PFERR_USER_MASK | PFERR_SS_MASK);
 	if (fetch_fault && has_pferr_fetch(w))
 		errcode |= PFERR_FETCH_MASK;
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [PATCH] KVM: x86/mmu: Cope with the SS bit being set in #NPF exit codes
  2026-10-09 13:04 [PATCH] KVM: x86/mmu: Cope with the SS bit being set in #NPF exit codes Paolo Bonzini
@ 2026-10-10  5:34 ` kernel test robot
  0 siblings, 0 replies; 2+ messages in thread
From: kernel test robot @ 2026-10-10  5:34 UTC (permalink / raw)
  To: Paolo Bonzini, linux-kernel, kvm; +Cc: oe-kbuild-all, SimonP, stable

Hi Paolo,

kernel test robot noticed the following build warnings:

[auto build test WARNING on kvm/queue]
[also build test WARNING on kvm/next mst-vhost/linux-next linus/master v7.3-rc6 next-20261008]
[cannot apply to kvm/linux-next]
[If your patch is applied to the wrong git tree, kindly drop us a note.
And when submitting patch, we suggest to use '--base' as documented in
https://git-scm.com/docs/git-format-patch#_base_tree_information]

url:    https://github.com/intel-lab-lkp/linux/commits/Paolo-Bonzini/KVM-x86-mmu-Cope-with-the-SS-bit-being-set-in-NPF-exit-codes/20261009-150450
base:   https://git.kernel.org/pub/scm/virt/kvm/kvm.git queue
patch link:    https://lore.kernel.org/r/20261009130450.361690-1-pbonzini%40redhat.com
patch subject: [PATCH] KVM: x86/mmu: Cope with the SS bit being set in #NPF exit codes
config: x86_64-randconfig-1016-20261010 (https://download.01.org/0day-ci/archive/20261010/202610101512.fYgKYoD8-lkp@intel.com/config)
compiler: gcc-14 (Debian 14.2.0-19) 14.2.0
reproduce (this is a W=1 build): (https://download.01.org/0day-ci/archive/20261010/202610101512.fYgKYoD8-lkp@intel.com/reproduce)

If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <lkp@intel.com>
| Closes: https://lore.kernel.org/oe-kbuild-all/202610101512.fYgKYoD8-lkp@intel.com/

All warnings (new ones prefixed by >>):

   In file included from arch/x86/kvm/mmu/mmu.c:5413:
   arch/x86/kvm/mmu/paging_tmpl.h: In function 'ept_walk_addr_generic':
>> arch/x86/kvm/mmu/paging_tmpl.h:331:19: warning: unused variable 'user_fault' [-Wunused-variable]
     331 |         const int user_fault  = access & PFERR_USER_MASK;
         |                   ^~~~~~~~~~
   In file included from arch/x86/kvm/mmu/mmu.c:5417:
   arch/x86/kvm/mmu/paging_tmpl.h: In function 'paging64_walk_addr_generic':
>> arch/x86/kvm/mmu/paging_tmpl.h:331:19: warning: unused variable 'user_fault' [-Wunused-variable]
     331 |         const int user_fault  = access & PFERR_USER_MASK;
         |                   ^~~~~~~~~~
   In file included from arch/x86/kvm/mmu/mmu.c:5421:
   arch/x86/kvm/mmu/paging_tmpl.h: In function 'paging32_walk_addr_generic':
>> arch/x86/kvm/mmu/paging_tmpl.h:331:19: warning: unused variable 'user_fault' [-Wunused-variable]
     331 |         const int user_fault  = access & PFERR_USER_MASK;
         |                   ^~~~~~~~~~


vim +/user_fault +331 arch/x86/kvm/mmu/paging_tmpl.h

be94f6b71067df4 arch/x86/kvm/paging_tmpl.h     Huaitong Han        2016-03-22  282  
ac889bdd23df1ee arch/x86/kvm/mmu/paging_tmpl.h Paolo Bonzini       2026-05-10  283  static inline bool FNAME(is_last_gpte)(struct kvm_pagewalk *w,
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  284  				       unsigned int level, unsigned int gpte)
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  285  {
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  286  	/*
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  287  	 * For EPT and PAE paging (both variants), bit 7 is either reserved at
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  288  	 * all level or indicates a huge page (ignoring CR3/EPTP).  In either
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  289  	 * case, bit 7 being set terminates the walk.
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  290  	 */
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  291  #if PTTYPE == 32
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  292  	/*
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  293  	 * 32-bit paging requires special handling because bit 7 is ignored if
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  294  	 * CR4.PSE=0, not reserved.  Clear bit 7 in the gpte if the level is
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  295  	 * greater than the last level for which bit 7 is the PAGE_SIZE bit.
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  296  	 *
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  297  	 * The RHS has bit 7 set iff level < (2 + PSE).  If it is clear, bit 7
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  298  	 * is not reserved and does not indicate a large page at this level,
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  299  	 * so clear PT_PAGE_SIZE_MASK in gpte if that is the case.
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  300  	 */
d869ff104efb1f2 arch/x86/kvm/mmu/paging_tmpl.h Paolo Bonzini       2026-05-11  301  	gpte &= level - (PT32_ROOT_LEVEL + w->cpu_role.ext.cr4_pse);
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  302  #endif
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  303  	/*
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  304  	 * PG_LEVEL_4K always terminates.  The RHS has bit 7 set
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  305  	 * iff level <= PG_LEVEL_4K, which for our purpose means
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  306  	 * level == PG_LEVEL_4K; set PT_PAGE_SIZE_MASK in gpte then.
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  307  	 */
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  308  	gpte |= level - PG_LEVEL_4K - 1;
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  309  
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  310  	return gpte & PT_PAGE_SIZE_MASK;
7cd138db5cae0da arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2021-06-22  311  }
ac79c978f173586 drivers/kvm/paging_tmpl.h      Avi Kivity          2007-01-05  312  /*
736c291c9f36b07 arch/x86/kvm/mmu/paging_tmpl.h Sean Christopherson 2019-12-06  313   * Fetch a guest pte for a guest virtual address, or for an L2's GPA.
ac79c978f173586 drivers/kvm/paging_tmpl.h      Avi Kivity          2007-01-05  314   */
1e301feb079e8ee arch/x86/kvm/paging_tmpl.h     Joerg Roedel        2010-09-10  315  static int FNAME(walk_addr_generic)(struct guest_walker *walker,
ac889bdd23df1ee arch/x86/kvm/mmu/paging_tmpl.h Paolo Bonzini       2026-05-10  316  				    struct kvm_vcpu *vcpu, struct kvm_pagewalk *w,
5b22bbe717d9b96 arch/x86/kvm/mmu/paging_tmpl.h Lai Jiangshan       2022-03-11  317  				    gpa_t addr, u64 access)
6aa8b732ca01c3d drivers/kvm/paging_tmpl.h      Avi Kivity          2006-12-10  318  {
8cbc70696f149e4 arch/x86/kvm/paging_tmpl.h     Avi Kivity          2012-09-16  319  	int ret;
42bf3f0a1f5a25b drivers/kvm/paging_tmpl.h      Avi Kivity          2007-10-17  320  	pt_element_t pte;
3f649ab728cda80 arch/x86/kvm/mmu/paging_tmpl.h Kees Cook           2020-06-03  321  	pt_element_t __user *ptep_user;
cea0f0e7ea54753 drivers/kvm/paging_tmpl.h      Avi Kivity          2007-01-05  322  	gfn_t table_gfn;
0780516a18f87e8 arch/x86/kvm/paging_tmpl.h     Paolo Bonzini       2017-05-11  323  	u64 pt_access, pte_access;
0780516a18f87e8 arch/x86/kvm/paging_tmpl.h     Paolo Bonzini       2017-05-11  324  	unsigned index, accessed_dirty, pte_pkey;
5b22bbe717d9b96 arch/x86/kvm/mmu/paging_tmpl.h Lai Jiangshan       2022-03-11  325  	u64 nested_access;
42bf3f0a1f5a25b drivers/kvm/paging_tmpl.h      Avi Kivity          2007-10-17  326  	gpa_t pte_gpa;
86407bcb5c8320a arch/x86/kvm/paging_tmpl.h     Paolo Bonzini       2017-03-30  327  	bool have_ad;
134291bf3cb434a arch/x86/kvm/paging_tmpl.h     Takuya Yoshikawa    2011-07-01  328  	int offset;
0780516a18f87e8 arch/x86/kvm/paging_tmpl.h     Paolo Bonzini       2017-05-11  329  	u64 walk_nx_mask = 0;
134291bf3cb434a arch/x86/kvm/paging_tmpl.h     Takuya Yoshikawa    2011-07-01  330  	const int write_fault = access & PFERR_WRITE_MASK;
134291bf3cb434a arch/x86/kvm/paging_tmpl.h     Takuya Yoshikawa    2011-07-01 @331  	const int user_fault  = access & PFERR_USER_MASK;
134291bf3cb434a arch/x86/kvm/paging_tmpl.h     Takuya Yoshikawa    2011-07-01  332  	const int fetch_fault = access & PFERR_FETCH_MASK;
bb24edbb673f6f3 arch/x86/kvm/mmu/paging_tmpl.h Kevin Cheng         2026-05-22  333  	/*
bb24edbb673f6f3 arch/x86/kvm/mmu/paging_tmpl.h Kevin Cheng         2026-05-22  334  	 * Note! Track the error_code that's common to legacy shadow paging
bb24edbb673f6f3 arch/x86/kvm/mmu/paging_tmpl.h Kevin Cheng         2026-05-22  335  	 * and NPT shadow paging as a u16 to guard against unintentionally
bb24edbb673f6f3 arch/x86/kvm/mmu/paging_tmpl.h Kevin Cheng         2026-05-22  336  	 * setting any of bits 63:16.  Architecturally, the #PF error code is
bb24edbb673f6f3 arch/x86/kvm/mmu/paging_tmpl.h Kevin Cheng         2026-05-22  337  	 * 32 bits, and Intel CPUs don't support settings bits 31:16.
bb24edbb673f6f3 arch/x86/kvm/mmu/paging_tmpl.h Kevin Cheng         2026-05-22  338  	 */
134291bf3cb434a arch/x86/kvm/paging_tmpl.h     Takuya Yoshikawa    2011-07-01  339  	u16 errcode = 0;
13d22b6aebb000a arch/x86/kvm/paging_tmpl.h     Avi Kivity          2012-09-12  340  	gpa_t real_gpa;
13d22b6aebb000a arch/x86/kvm/paging_tmpl.h     Avi Kivity          2012-09-12  341  	gfn_t gfn;
6aa8b732ca01c3d drivers/kvm/paging_tmpl.h      Avi Kivity          2006-12-10  342  
6fbc277053836a4 arch/x86/kvm/paging_tmpl.h     Xiao Guangrong      2012-06-20  343  	trace_kvm_mmu_pagetable_walk(addr, access);
92c1c1e85bdd72b arch/x86/kvm/paging_tmpl.h     Takuya Yoshikawa    2011-07-01  344  retry_walk:
d869ff104efb1f2 arch/x86/kvm/mmu/paging_tmpl.h Paolo Bonzini       2026-05-11  345  	walker->level = w->cpu_role.base.level;
d9af9341707f67c arch/x86/kvm/mmu/paging_tmpl.h Paolo Bonzini       2026-04-05  346  	pte           = kvm_mmu_get_guest_pgd(vcpu, w);
d869ff104efb1f2 arch/x86/kvm/mmu/paging_tmpl.h Paolo Bonzini       2026-05-11  347  	have_ad       = PT_HAVE_ACCESSED_DIRTY(w);
1e301feb079e8ee arch/x86/kvm/paging_tmpl.h     Joerg Roedel        2010-09-10  348  

--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-10  5:35 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 13:04 [PATCH] KVM: x86/mmu: Cope with the SS bit being set in #NPF exit codes Paolo Bonzini
2026-10-10  5:34 ` kernel test robot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®