mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2] KVM: x86: Clear CR3[63:32] on SMM entry when running on Hyper-V
@ 2026-09-29  1:33 Qiliang Yuan
  2026-09-29  1:49 ` Qiliang Yuan
  0 siblings, 1 reply; 2+ messages in thread
From: Qiliang Yuan @ 2026-09-29  1:33 UTC (permalink / raw)
  To: Sean Christopherson, Paolo Bonzini, Thomas Gleixner, Ingo Molnar,
	Borislav Petkov, Dave Hansen, x86, H. Peter Anvin
  Cc: kvm, linux-kernel, Vitaly Kuznetsov, K. Y. Srinivasan,
	Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li, linux-hyperv,
	Qiliang Yuan

enter_smm() clears CR0.PG and EFER.LMA/LME but leaves CR3 as-is, so a
64-bit guest whose page tables live above 4GiB enters SMM with
CR3[63:32] != 0 while outside of long mode.  Bare-metal SVM accepts that
state, but Hyper-V's emulation of VMRUN for a nested hypervisor rejects
it as an invalid VMCB, and the vCPU dies on the very first instruction
of the SMI handler:

  KVM: entry failed, hardware error 0xffffffff
  EIP=00008000 EFL=00000002 [-------] CPL=0 II=0 A20=1 SMM=1 HLT=0
  CS =f900 7bff9000 ffffffff 00809300
  CR0=00050032 CR2=2a91044a CR3=77681000 CR4=00000000
  EFER=0000000000000000

QEMU prints only CR3[31:0] here; the CR3 saved in SMRAM for this vCPU
was 0x277681000.  All vCPUs that failed had a CR3 above 4GiB, while the
one vCPU whose CR3 was below 4GiB entered SMM without issue.

This reproduces reliably when booting a Windows 11 guest (8GiB of RAM)
with Secure Boot OVMF, i.e. with SMM enabled, in KVM on WSL2 on an AMD
host, with Windows 11 25H2 build 26200.9457, the latest released build.

With paging disabled outside of long mode, only CR3[31:0] is
reachable: a MOV to CR3 can only write 32 bits, and the SMI handler
must load its own CR3 before enabling paging.  RSM restores the full
CR3 from the SMRAM state-save area, which was written before this point.
Clear the upper 32 bits when KVM runs on Hyper-V, so that the SMM entry
state passes Hyper-V's VMRUN consistency checks.  Leave the behavior on
other hosts unchanged, as the problem belongs to Hyper-V and should be
fixed there.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
V1 -> V2:
- Only clear CR3[63:32] when KVM runs on Hyper-V, leave other hosts
  unchanged (Sean)
- Note in the changelog that the latest released Windows 11 build
  (25H2, 26200.9457) is still affected
- Cc Hyper-V maintainers and linux-hyperv

v1: https://lore.kernel.org/r/20260928-kvm-smm-cr3-upper-bits-v1-1-138f52dd531e@gmail.com
---
 arch/x86/kvm/smm.c | 15 +++++++++++++++
 1 file changed, 15 insertions(+)

diff --git a/arch/x86/kvm/smm.c b/arch/x86/kvm/smm.c
index f623c5986119..34c16b8c2c0b 100644
--- a/arch/x86/kvm/smm.c
+++ b/arch/x86/kvm/smm.c
@@ -2,6 +2,7 @@
 #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
 
 #include <linux/kvm_host.h>
+#include <asm/hypervisor.h>
 #include "x86.h"
 #include "kvm_cache_regs.h"
 #include "kvm_emulate.h"
@@ -361,6 +362,20 @@ void enter_smm(struct kvm_vcpu *vcpu)
 	if (guest_cpu_cap_has(vcpu, X86_FEATURE_LM))
 		if (kvm_x86_call(set_efer)(vcpu, 0))
 			goto error;
+
+	/*
+	 * Work around Hyper-V rejecting VMRUN for a nested hypervisor when
+	 * CR3[63:32] != 0 with EFER.LMA=0, which is exactly the state left
+	 * behind by entering SMM with page tables above 4GiB.  CR3 is
+	 * unmodified on SMM entry, but with paging disabled and outside of
+	 * long mode bits 63:32 are unreachable, and RSM restores the full
+	 * value from SMRAM, so clearing them is invisible to the guest.
+	 */
+	if (hypervisor_is_type(X86_HYPER_MS_HYPERV) &&
+	    kvm_read_cr3(vcpu) >> 32) {
+		vcpu->arch.cr3 = (u32)vcpu->arch.cr3;
+		kvm_register_mark_dirty(vcpu, VCPU_EXREG_CR3);
+	}
 #endif
 
 	vcpu->arch.cpuid_dynamic_bits_dirty = true;

---
base-commit: eb3f4b7426cfd2b79d65b7d37155480b32259a11
change-id: 20260928-kvm-smm-cr3-upper-bits-9819a1a2f337

Best regards,
-- 
Qiliang Yuan <odys.yuan@gmail.com>


^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [PATCH v2] KVM: x86: Clear CR3[63:32] on SMM entry when running on Hyper-V
  2026-09-29  1:33 [PATCH v2] KVM: x86: Clear CR3[63:32] on SMM entry when running on Hyper-V Qiliang Yuan
@ 2026-09-29  1:49 ` Qiliang Yuan
  0 siblings, 0 replies; 2+ messages in thread
From: Qiliang Yuan @ 2026-09-29  1:49 UTC (permalink / raw)
  To: Sean Christopherson, Paolo Bonzini, Thomas Gleixner, Ingo Molnar,
	Borislav Petkov, Dave Hansen, x86, H. Peter Anvin
  Cc: Qiliang Yuan, kvm, linux-kernel, Vitaly Kuznetsov,
	K. Y. Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
	linux-hyperv

Some more data on this, in case it helps either the KVM or the Hyper-V
side.  It is also tracked, with the WSL logs, at
https://github.com/microsoft/WSL/issues/41709


1. Setup
--------

  Host:   Windows 11 Pro 25H2, OS build 26200.9457 (fully up to date
          through Windows Update), WSL 2.9.3.0, AMD Ryzen 9 7940HX
  L1:     WSL2 kernel 6.18.35.2-microsoft-standard-WSL2, kvm_amd nested=1
  L2:     Windows 11 25H2 guest, q35, 4 vCPUs, 8GiB RAM,
          OVMF_CODE_4M.ms.fd (Secure Boot, SMM), swtpm,
          QEMU 10.2.1, libvirt 12.0.0

Wei mentioned that this may already be fixed in Hyper-V.  Since it
still reproduces on the latest released build above, the fix does not
seem to be in released builds yet.  Which build or channel carries it?
I am happy to test it.


2. Not a recent regression
--------------------------

The libvirt log of this guest on the same machine shows the same
"KVM: entry failed, hardware error 0xffffffff" going back a year:

  WSL kernel                    Period                    Boots  Failures
  5.15.167.4-microsoft-...      2025-09-01 .. 2025-10-01    127       133
  6.18.35.2-microsoft-...       2026-09-24 .. 2026-09-28     15        21
  6.18.35.2 + this patch (v2)   2026-09-29                    1         0


3. Testing of v2
----------------

I applied v2 to the WSL2 6.18.35.2 kernel and traced enter_smm() with
bpftrace, recording vcpu->arch.cr3 on entry and on return, while the
guest booted from firmware to the Windows desktop with SMM and Secure
Boot enabled:

  SMM entries in total:                        7532
  entries with CR3 above 4GiB before entry:    1185
  of those, CR3[63:32] cleared on return:      1185
  VMRUN failures:                                 0

For example:

  vcpu=1 cr3 before=0x13aa12000 after=0x3aa12000

Every one of those 1185 entries would have hit the failure without the
patch, and the guest returned from SMM normally in all cases.


4. The rejected VMCB
--------------------

To capture the failing state without restarting WSL, I built an
unpatched kvm.ko/kvm-amd.ko for the same running kernel, loaded them
with kvm_amd.dump_invalid_vmcb=1 and booted the guest.  It failed at
SMM entry again, with CR3 above 4GiB.

The relevant fields are exit_code ffffffff, rip 8000 (first instruction
of the SMI handler), cr0 00050032 (PE=0, PG=0), efer 00001000 (SVME
only, LMA=0), event_inj 0 (no pending event), and cr3 0000000262906000,
i.e. CR3[63:32] = 0x2 outside of long mode.  The full dump:

  SVM vCPU1 VMCB 000000004a03e896, last attempted VMRUN on CPU 2
  VMCB Control Area:
  cr_read:            0010
  cr_write:           0010
  dr_read:            00ff
  dr_write:           00ff
  exceptions:         00060042
  intercepts:         bddc8027 00006e7f
  pause filter count: 3000
  pause filter threshold:128
  iopm_base_pa:       000000034604c000
  msrpm_base_pa:      00000003d50da000
  tsc_offset:         ffffbe16edc138b1
  asid:               8
  tlb_ctl:            0
  int_ctl:            01000000
  int_vector:         00000000
  int_state:          00000000
  exit_code:          ffffffff
  exit_info1:         0000000000000000
  exit_info2:         0000000000000000
  exit_int_info:      00000000
  exit_int_info_err:  00000000
  nested_ctl:         1
  nested_cr3:         0000000260eac000
  avic_vapic_bar:     0000000000000000
  ghcb:               0000000000000000
  event_inj:          00000000
  event_inj_err:      00000000
  virt_ext:           0
  next_rip:           0000000000000000
  avic_backing_page:  0000000000000000
  avic_logical_id:    0000000000000000
  avic_physical_id:   0000000000000000
  vmsa_pa:            0000000000000000
  allowed_sev_features:0000000000000000
  guest_sev_features: 0000000000000000
  VMCB State Save Area:
  es:   s: 0000 a: 0893 l: ffffffff b: 0000000000000000
  cs:   s: fb00 a: 0893 l: ffffffff b: 000000007bffb000
  ss:   s: 0000 a: 0893 l: ffffffff b: 0000000000000000
  ds:   s: 0000 a: 0893 l: ffffffff b: 0000000000000000
  fs:   s: 0000 a: 0893 l: ffffffff b: 0000000000000000
  gs:   s: 0000 a: 0893 l: ffffffff b: 0000000000000000
  gdtr: s: 0000 a: 0000 l: 00000057 b: ffff9a0133f5ffb0
  ldtr: s: 0000 a: 0000 l: 00000000 b: 0000000000000000
  idtr: s: 0000 a: 0000 l: 00000000 b: 0000000000000000
  tr:   s: 0040 a: 008b l: 00000067 b: ffff9a0133f5e000
  vmpl: 0   cpl:  0               efer:          0000000000001000
  cr0:            0000000000050032 cr2:          ffffe7888c4ec690
  cr3:            0000000262906000 cr4:          0000000000000040
  dr6:            00000000ffff0ff0 dr7:          0000000000000400
  rip:            0000000000008000 rflags:       0000000000000002
  rsp:            ffffbf0c30b864d8 rax:          0000000000000000
  s_cet:          0000000000000000 ssp:          0000000000000000
  isst_addr:      0000000000000000
  star:           0023001000000000 lstar:        fffff801b9ac1840
  cstar:          fffff801b9ac1300 sfmask:       0000000000004700
  kernel_gs_base: 000000caa8494000 sysenter_cs:  0000000000000000
  sysenter_esp:   0000000000000000 sysenter_eip: 0000000000000000
  gpat:           0007010600070106 dbgctl:       0000000000000000
  br_from:        0000000000000000 br_to:        0000000000000000
  excp_from:      0000000000000000 excp_to:      0000000000000000
  rax:            0000000000000000 rbx:          fffff8014d990018
  rcx:            00000000000000b2 rdx:          00000000000000b2
  rsi:            0000000000000200 rdi:          0000000000000218
  rbp:            ffffbf0c30b86500 rsp:          ffffbf0c30b864d8
  r8:             0000000000000000 r9:           0000000000000000
  r10:            0000000000000000 r11:          ffff89fdab000000
  r12:            ffffbf0c30b86720 r13:          fffff8014d990060
  r14:            fffff8014d990078 r15:          0000000000000002

And QEMU's view of the same vCPU (it only prints CR3[31:0] here):

  KVM: entry failed, hardware error 0xffffffff
  EAX=00000000 EBX=4d990018 ECX=000000b2 EDX=000000b2
  ESI=00000200 EDI=00000218 EBP=30b86500 ESP=30b864d8
  EIP=00008000 EFL=00000002 [-------] CPL=0 II=0 A20=1 SMM=1 HLT=0
  ES =0000 00000000 ffffffff 00809300
  CS =fb00 7bffb000 ffffffff 00809300
  SS =0000 00000000 ffffffff 00809300
  DS =0000 00000000 ffffffff 00809300
  FS =0000 00000000 ffffffff 00809300
  GS =0000 00000000 ffffffff 00809300
  LDT=0000 00000000 00000000 00000000
  TR =0040 33f5e000 00000067 00008b00
  GDT=     33f5ffb0 00000057
  IDT=     00000000 00000000
  CR0=00050032 CR2=8c4ec690 CR3=62906000 CR4=00000000
  DR0=0000000000000000 DR1=0000000000000000 DR2=0000000000000000 DR3=0000000000000000
  DR6=00000000ffff0ff0 DR7=0000000000000400
  EFER=0000000000000000
  Code=00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 <bb> 4d 80 2e a1 38 fb 48 2e 89 07 2e 66 a1 30 fb 2e 66 89 47 02 2e 66 0f 01 17 b8 08 00 2e


Thanks,
Qiliang

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-29  1:49 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29  1:33 [PATCH v2] KVM: x86: Clear CR3[63:32] on SMM entry when running on Hyper-V Qiliang Yuan
2026-09-29  1:49 ` Qiliang Yuan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®