mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM
@ 2026-10-01 20:22 Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm Sean Christopherson
                   ` (10 more replies)
  0 siblings, 11 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Fix a class of bugs (limited to nested VMX, as far as we know) where KVM can
corrupt an unrelated process' memory if KVM (attempts to) write to guest memory
during VM destruction.  Because current->mm usually isn't kvm->mm when a VM is
dying, e.g. because the associated kvm->mm process has already exited, writing
to what KVM thinks is guest memory will corrupt the current address space if
the associated userspace address happens to be writable in the victim.

Patch 1 is a blanket fix for the uaccess paths.  AFAIK, x86's nVMX is the only
path in KVM that screws up, but auditing kvm_arch_destroy_vm() proved to be
infeasible for a human (I didn't throw AI at it, yet...), and I can't think of
any downsides to going straight to a broader fix.

Patch 2 fixes what I assume is a blatant PPC bug.  KVM PPC completely ignores
kvm_arch_flush_shadow_all(), i.e. AFAICT, doesn't tear down its page tables
when the owning process exits.  I don't have much confidence in the "fix", in
part because it seems impossible that such a blatant bug could have gone
unnoticed, but also because I went with a very naive approach of invoking
kvm_arch_flush_shadow_memslot() for each memslot.  The changelog is pretty
sparse, because I didn't know how to describe the issue byeond "this is
completely broken".

Patches 3-5 implement more agressive hardening to nuke the memslots before
calling kvm_arch_destroy_vm(), e.g. to guard against writing to guest memory
during kvm_arch_destroy_vm() without going through uaccess.  Setting dummy
memslots feels a little hacky, but kvm_arch_flush_shadow_all() should have
purged everything that effectively caches memslots, so it seems like the right
approach?

The remaining patches fudge around the nVMX bugs (KVM abuses its nested VM-Exit
flow to forcefully take a vCPU out of L2, which has been an endless source of
pain, but is also equally difficult to fix properly), and add more hardening to
detect KVM bugs (though the uaccess+memslot changes earlier in the series should
render any bugs benign).

v1: https://lore.kernel.org/all/20260908132838.2116068-1-jmattson@google.com

Jim Mattson (1):
  KVM: nVMX: Don't flush shadow VMCS12 to guest memory during vCPU
    teardown

Sean Christopherson (9):
  KVM: Reject user accesses to guest memory if current->mm != kvm->mm
  KVM: PPC: Flush/zap all memslots on kvm_arch_flush_shadow_all()
  KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is
    freed
  KVM: Disallow setting memslots when the VM is being destroyed
  KVM: Destroy memslots immediately after mmu_notifiers are unregistered
  KVM: WARN if KVM attempts to do guest-related uaccess with "wrong"
    process
  KVM: WARN and reject guest-based uaccess if VM is dying
  KVM: nVMX: Don't try to load eVMCS12 page when the VM is dying
  KVM: Pre-check uaccesses in KVM's APIs to read/write guest memory

 arch/powerpc/include/asm/kvm_host.h |  1 -
 arch/powerpc/kvm/powerpc.c          | 11 ++++
 arch/x86/kvm/vmx/nested.c           | 14 ++++-
 arch/x86/kvm/vmx/sgx.c              |  2 +-
 arch/x86/kvm/vmx/vmx.c              |  8 +--
 arch/x86/kvm/x86.c                  | 36 +++++-------
 include/linux/kvm_host.h            | 30 +++++++++-
 virt/kvm/kvm_main.c                 | 87 +++++++++++++++++++----------
 8 files changed, 126 insertions(+), 63 deletions(-)


base-commit: d4b7fb647204f0c81dfeae2d1a708e4d858e0c94
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 21:08   ` James Houghton
  2026-10-01 20:22 ` [PATCH v2 02/10] KVM: PPC: Flush/zap all memslots on kvm_arch_flush_shadow_all() Sean Christopherson
                   ` (9 subsequent siblings)
  10 siblings, 1 reply; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Reject user accesses to guest memory, which are supposed to be done only
in the context of KVM_RUN or similar operations, if the current address
space is not the VM's (host userspace) address space.  If KVM writes to
guest memory after the owning host process has exited, or if the VM is
being destroyed in the context of a different process, then writing using
the wrong address space will corrupt a different process' memory.

Reject the access but don't WARN() or KVM_BUG_ON() event though attempting
to access guest memory with a mismatched address space is a blatant KVM
bug, because unfortunately KVM is buggy.  On KVM VMX, when a vCPU is
destroyed while L2 is active, KVM synthesizes a nested VM-Exit to force the
vCPU out of L2 in order to free the nested VMX assets, and a side effect of
a nested VM-Exit is that it flushes the cached shadow VMCS12 back to guest
memory:

  vmx_vcpu_free()
  |-> nested_vmx_free_vcpu()
      |-> vmx_leave_nested()
          |-> nested_vmx_vmexit(vcpu, -1, 0, 0)
              |-> nested_flush_cached_shadow_vmcs12()
                  |-> kvm_write_guest_cached()
                      |-> __copy_to_user(ghc->hva, ...)

Fix the bug broadly even though the "real" bug is that KVM abuses the
nested VM-Exit flow for non-architectural purposes, as there may be other
such violations lurking.  For now, punt on fixing individual bugs and
hardening the common flows, e.g. with WARNs.

Opportunistically provide wrappers in anticipation of adding more checks
and hardening, i.e. growing the logic beyond checking current->mm.

Fixes: 61ada7488ffd ("KVM: nVMX: Cache shadow vmcs12 on VMEntry and flush to memory on VMExit")
Cc: stable@vger.kernel.org
Reported-by: Jim Mattson <jmattson@google.com>
Closes: https://lore.kernel.org/all/20260908132838.2116068-1-jmattson@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 arch/x86/kvm/vmx/sgx.c   |  2 +-
 arch/x86/kvm/vmx/vmx.c   |  8 ++++----
 include/linux/kvm_host.h | 25 +++++++++++++++++++++++--
 virt/kvm/kvm_main.c      | 24 ++++++++++++------------
 4 files changed, 40 insertions(+), 19 deletions(-)

diff --git a/arch/x86/kvm/vmx/sgx.c b/arch/x86/kvm/vmx/sgx.c
index 771c75a58343..52a6d0f8bb9e 100644
--- a/arch/x86/kvm/vmx/sgx.c
+++ b/arch/x86/kvm/vmx/sgx.c
@@ -64,7 +64,7 @@ static void sgx_handle_emulation_failure(struct kvm_vcpu *vcpu, u64 addr,
 static int sgx_read_hva(struct kvm_vcpu *vcpu, unsigned long hva, void *data,
 			unsigned int size)
 {
-	if (__copy_from_user(data, (void __user *)hva, size)) {
+	if (kvm_copy_from_user(vcpu->kvm, data, (void __user *)hva, size)) {
 		sgx_handle_emulation_failure(vcpu, hva, size);
 		return -EFAULT;
 	}
diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
index 612ab07d4100..b2ffa6002944 100644
--- a/arch/x86/kvm/vmx/vmx.c
+++ b/arch/x86/kvm/vmx/vmx.c
@@ -4011,16 +4011,16 @@ static int init_rmode_tss(struct kvm *kvm, void __user *ua)
 	int i;
 
 	for (i = 0; i < 3; i++) {
-		if (__copy_to_user(ua + PAGE_SIZE * i, zero_page, PAGE_SIZE))
+		if (kvm_copy_to_user(kvm, ua + PAGE_SIZE * i, zero_page, PAGE_SIZE))
 			return -EFAULT;
 	}
 
 	data = TSS_BASE_SIZE + TSS_REDIRECTION_SIZE;
-	if (__copy_to_user(ua + TSS_IOPB_BASE_OFFSET, &data, sizeof(u16)))
+	if (kvm_copy_to_user(kvm, ua + TSS_IOPB_BASE_OFFSET, &data, sizeof(u16)))
 		return -EFAULT;
 
 	data = ~0;
-	if (__copy_to_user(ua + RMODE_TSS_SIZE - 1, &data, sizeof(u8)))
+	if (kvm_copy_to_user(kvm, ua + RMODE_TSS_SIZE - 1, &data, sizeof(u8)))
 		return -EFAULT;
 
 	return 0;
@@ -4055,7 +4055,7 @@ static int init_rmode_identity_map(struct kvm *kvm)
 	for (i = 0; i < (PAGE_SIZE / sizeof(tmp)); i++) {
 		tmp = (i << 22) + (_PAGE_PRESENT | _PAGE_RW | _PAGE_USER |
 			_PAGE_ACCESSED | _PAGE_DIRTY | _PAGE_PSE);
-		if (__copy_to_user(uaddr + i * sizeof(tmp), &tmp, sizeof(tmp))) {
+		if (kvm_copy_to_user(kvm, uaddr + i * sizeof(tmp), &tmp, sizeof(tmp))) {
 			r = -EFAULT;
 			goto out;
 		}
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 3dd04605f2e5..b37cf275ef79 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -1350,13 +1350,34 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 			      gpa_t gpa, unsigned long len);
 
+static __always_inline __must_check bool kvm_can_do_uaccess(struct kvm *kvm)
+{
+	return current->mm == kvm->mm;
+}
+
+#define BUILD_KVM_COPY_USER_WRAPPER(fn, to_user, from_user)				\
+static __always_inline __must_check unsigned long kvm_##fn(struct kvm *kvm,		\
+							   void to_user *to,		\
+							   const void from_user *from,	\
+							   unsigned long n)		\
+{											\
+	if (!kvm_can_do_uaccess(kvm))							\
+		return n;								\
+											\
+	return __##fn(to, from, n);							\
+}
+BUILD_KVM_COPY_USER_WRAPPER(copy_from_user, , __user)
+BUILD_KVM_COPY_USER_WRAPPER(copy_from_user_inatomic, , __user)
+BUILD_KVM_COPY_USER_WRAPPER(copy_to_user, __user, )
+BUILD_KVM_COPY_USER_WRAPPER(copy_to_user_inatomic, __user, )
+
 #define __kvm_get_guest(kvm, gfn, offset, v)				\
 ({									\
 	unsigned long __addr = gfn_to_hva(kvm, gfn);			\
 	typeof(v) __user *__uaddr = (typeof(__uaddr))(__addr + offset);	\
 	int __ret = -EFAULT;						\
 									\
-	if (!kvm_is_error_hva(__addr))					\
+	if (!kvm_is_error_hva(__addr) && kvm_can_do_uaccess(kvm))	\
 		__ret = get_user(v, __uaddr);				\
 	__ret;								\
 })
@@ -1376,7 +1397,7 @@ int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 	typeof(v) __user *__uaddr = (typeof(__uaddr))(__addr + offset);	\
 	int __ret = -EFAULT;						\
 									\
-	if (!kvm_is_error_hva(__addr))					\
+	if (!kvm_is_error_hva(__addr) && kvm_can_do_uaccess(kvm))	\
 		__ret = put_user(v, __uaddr);				\
 	if (!__ret)							\
 		mark_page_dirty(kvm, gfn);				\
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 108d42c5c1d6..9a24c3064896 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -3192,8 +3192,8 @@ static int next_segment(unsigned long len, int offset)
 }
 
 /* Copy @len bytes from guest memory at '(@gfn * PAGE_SIZE) + @offset' to @data */
-static int __kvm_read_guest_page(struct kvm_memory_slot *slot, gfn_t gfn,
-				 void *data, int offset, int len)
+static int __kvm_read_guest_page(struct kvm *kvm, struct kvm_memory_slot *slot,
+				 gfn_t gfn, void *data, int offset, int len)
 {
 	int r;
 	unsigned long addr;
@@ -3204,7 +3204,7 @@ static int __kvm_read_guest_page(struct kvm_memory_slot *slot, gfn_t gfn,
 	addr = gfn_to_hva_memslot_prot(slot, gfn, NULL);
 	if (kvm_is_error_hva(addr))
 		return -EFAULT;
-	r = __copy_from_user(data, (void __user *)addr + offset, len);
+	r = kvm_copy_from_user(kvm, data, (void __user *)addr + offset, len);
 	if (r)
 		return -EFAULT;
 	return 0;
@@ -3215,7 +3215,7 @@ int kvm_read_guest_page(struct kvm *kvm, gfn_t gfn, void *data, int offset,
 {
 	struct kvm_memory_slot *slot = gfn_to_memslot(kvm, gfn);
 
-	return __kvm_read_guest_page(slot, gfn, data, offset, len);
+	return __kvm_read_guest_page(kvm, slot, gfn, data, offset, len);
 }
 EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_read_guest_page);
 
@@ -3224,7 +3224,7 @@ int kvm_vcpu_read_guest_page(struct kvm_vcpu *vcpu, gfn_t gfn, void *data,
 {
 	struct kvm_memory_slot *slot = kvm_vcpu_gfn_to_memslot(vcpu, gfn);
 
-	return __kvm_read_guest_page(slot, gfn, data, offset, len);
+	return __kvm_read_guest_page(vcpu->kvm, slot, gfn, data, offset, len);
 }
 EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_read_guest_page);
 
@@ -3268,8 +3268,8 @@ int kvm_vcpu_read_guest(struct kvm_vcpu *vcpu, gpa_t gpa, void *data, unsigned l
 }
 EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_read_guest);
 
-static int __kvm_read_guest_atomic(struct kvm_memory_slot *slot, gfn_t gfn,
-			           void *data, int offset, unsigned long len)
+static int __kvm_read_guest_atomic(struct kvm *kvm, struct kvm_memory_slot *slot,
+				   gfn_t gfn, void *data, int offset, unsigned long len)
 {
 	int r;
 	unsigned long addr;
@@ -3281,7 +3281,7 @@ static int __kvm_read_guest_atomic(struct kvm_memory_slot *slot, gfn_t gfn,
 	if (kvm_is_error_hva(addr))
 		return -EFAULT;
 	pagefault_disable();
-	r = __copy_from_user_inatomic(data, (void __user *)addr + offset, len);
+	r = kvm_copy_from_user_inatomic(kvm, data, (void __user *)addr + offset, len);
 	pagefault_enable();
 	if (r)
 		return -EFAULT;
@@ -3295,7 +3295,7 @@ int kvm_vcpu_read_guest_atomic(struct kvm_vcpu *vcpu, gpa_t gpa,
 	struct kvm_memory_slot *slot = kvm_vcpu_gfn_to_memslot(vcpu, gfn);
 	int offset = offset_in_page(gpa);
 
-	return __kvm_read_guest_atomic(slot, gfn, data, offset, len);
+	return __kvm_read_guest_atomic(vcpu->kvm, slot, gfn, data, offset, len);
 }
 EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_vcpu_read_guest_atomic);
 
@@ -3313,7 +3313,7 @@ static int __kvm_write_guest_page(struct kvm *kvm,
 	addr = gfn_to_hva_memslot(memslot, gfn);
 	if (kvm_is_error_hva(addr))
 		return -EFAULT;
-	r = __copy_to_user((void __user *)addr + offset, data, len);
+	r = kvm_copy_to_user(kvm, (void __user *)addr + offset, data, len);
 	if (r)
 		return -EFAULT;
 	mark_page_dirty_in_slot(kvm, memslot, gfn);
@@ -3451,7 +3451,7 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 	if (unlikely(!ghc->memslot))
 		return kvm_write_guest(kvm, gpa, data, len);
 
-	r = __copy_to_user((void __user *)ghc->hva + offset, data, len);
+	r = kvm_copy_to_user(kvm, (void __user *)ghc->hva + offset, data, len);
 	if (r)
 		return -EFAULT;
 	mark_page_dirty_in_slot(kvm, ghc->memslot, gpa >> PAGE_SHIFT);
@@ -3489,7 +3489,7 @@ int kvm_read_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 	if (unlikely(!ghc->memslot))
 		return kvm_read_guest(kvm, gpa, data, len);
 
-	r = __copy_from_user(data, (void __user *)ghc->hva + offset, len);
+	r = kvm_copy_from_user(kvm, data, (void __user *)ghc->hva + offset, len);
 	if (r)
 		return -EFAULT;
 
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 02/10] KVM: PPC: Flush/zap all memslots on kvm_arch_flush_shadow_all()
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 03/10] KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is freed Sean Christopherson
                   ` (8 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Flush/zap every memslot in kvm_arch_flush_shadow_all(), i.e. when the
owning userspace process is exiting, as KVM must ensure all associated
memory is unmapped from the guest.

Cc: stable@vger.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 arch/powerpc/include/asm/kvm_host.h |  1 -
 arch/powerpc/kvm/powerpc.c          | 11 +++++++++++
 2 files changed, 11 insertions(+), 1 deletion(-)

diff --git a/arch/powerpc/include/asm/kvm_host.h b/arch/powerpc/include/asm/kvm_host.h
index 2d139c807577..1c8d9d7360e9 100644
--- a/arch/powerpc/include/asm/kvm_host.h
+++ b/arch/powerpc/include/asm/kvm_host.h
@@ -903,7 +903,6 @@ struct kvm_vcpu_arch {
 #define __KVM_HAVE_CREATE_DEVICE
 
 static inline void kvm_arch_memslots_updated(struct kvm *kvm, u64 gen) {}
-static inline void kvm_arch_flush_shadow_all(struct kvm *kvm) {}
 static inline void kvm_arch_vcpu_blocking(struct kvm_vcpu *vcpu) {}
 static inline void kvm_arch_vcpu_unblocking(struct kvm_vcpu *vcpu) {}
 
diff --git a/arch/powerpc/kvm/powerpc.c b/arch/powerpc/kvm/powerpc.c
index 9194cf492d1c..9e9d0865c2f6 100644
--- a/arch/powerpc/kvm/powerpc.c
+++ b/arch/powerpc/kvm/powerpc.c
@@ -764,6 +764,17 @@ void kvm_arch_commit_memory_region(struct kvm *kvm,
 	kvmppc_core_commit_memory_region(kvm, old, new, change);
 }
 
+void kvm_arch_flush_shadow_all(struct kvm *kvm)
+{
+	struct kvm_memory_slot *slot;
+	struct kvm_memslots *slots;
+	int bkt;
+
+	slots = kvm_memslots(kvm);
+	kvm_for_each_memslot(slot, bkt, slots)
+		kvmppc_core_flush_memslot(kvm, slot);
+}
+
 void kvm_arch_flush_shadow_memslot(struct kvm *kvm,
 				   struct kvm_memory_slot *slot)
 {
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 03/10] KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is freed
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 02/10] KVM: PPC: Flush/zap all memslots on kvm_arch_flush_shadow_all() Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 21:26   ` James Houghton
  2026-10-01 20:22 ` [PATCH v2 04/10] KVM: Disallow setting memslots when the VM is being destroyed Sean Christopherson
                   ` (7 subsequent siblings)
  10 siblings, 1 reply; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Unmap KVM-created VMAs for KVM-internal memslots when the associated
memslot is freed instead of unmapping via x86's internal API.  This will
allow deleting/freeing all memslots before kvm_arch_destroy_vm() without
leaking VMAs, and in general is a net positive as it makes it far less
likely that KVM will leak VMAs, e.g. in the unlikely scenario some other
flow deletes KVM-internal memslots without going through
__x86_set_memory_region().

Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 arch/x86/kvm/x86.c | 36 ++++++++++++++----------------------
 1 file changed, 14 insertions(+), 22 deletions(-)

diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index 79468ddfe473..e91bd79dd3cc 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -9925,7 +9925,7 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type)
  * @size > 0 to install a new slot, while @size == 0 to uninstall a
  * slot.  The return code can be one of the following:
  *
- *   HVA:           on success (uninstall will return a bogus HVA)
+ *   HVA:           on success (uninstall will return a NULL HVA)
  *   -errno:        on error
  *
  * The caller should always use IS_ERR() to check the return value
@@ -9938,10 +9938,10 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type)
 void __user * __x86_set_memory_region(struct kvm *kvm, int id, gpa_t gpa,
 				      u32 size)
 {
-	int i, r;
-	unsigned long hva, old_npages;
 	struct kvm_memslots *slots = kvm_memslots(kvm);
 	struct kvm_memory_slot *slot;
+	unsigned long hva;
+	int i, r;
 
 	lockdep_assert_held(&kvm->slots_lock);
 
@@ -9965,8 +9965,7 @@ void __user * __x86_set_memory_region(struct kvm *kvm, int id, gpa_t gpa,
 		if (!slot || !slot->npages)
 			return NULL;
 
-		old_npages = slot->npages;
-		hva = slot->userspace_addr;
+		hva = 0;
 	}
 
 	for (i = 0; i < kvm_arch_nr_memslot_as_ids(kvm); i++) {
@@ -9982,9 +9981,6 @@ void __user * __x86_set_memory_region(struct kvm *kvm, int id, gpa_t gpa,
 			return ERR_PTR_USR(r);
 	}
 
-	if (!size)
-		vm_munmap(hva, old_npages * PAGE_SIZE);
-
 	return (void __user *)hva;
 }
 EXPORT_SYMBOL_FOR_KVM_INTERNAL(__x86_set_memory_region);
@@ -10013,20 +10009,6 @@ void kvm_arch_pre_destroy_vm(struct kvm *kvm)
 
 void kvm_arch_destroy_vm(struct kvm *kvm)
 {
-	if (current->mm == kvm->mm) {
-		/*
-		 * Free memory regions allocated on behalf of userspace,
-		 * unless the memory map has changed due to process exit
-		 * or fd copying.
-		 */
-		mutex_lock(&kvm->slots_lock);
-		__x86_set_memory_region(kvm, APIC_ACCESS_PAGE_PRIVATE_MEMSLOT,
-					0, 0);
-		__x86_set_memory_region(kvm, IDENTITY_PAGETABLE_PRIVATE_MEMSLOT,
-					0, 0);
-		__x86_set_memory_region(kvm, TSS_PRIVATE_MEMSLOT, 0, 0);
-		mutex_unlock(&kvm->slots_lock);
-	}
 	if (kvm->arch.created_mediated_pmu)
 		perf_release_mediated_pmu();
 	kvm_destroy_vcpus(kvm);
@@ -10066,6 +10048,16 @@ void kvm_arch_free_memslot(struct kvm *kvm, struct kvm_memory_slot *slot)
 	}
 
 	kvm_page_track_free_memslot(slot);
+
+	/*
+	 * Free memory regions allocated on behalf of userspace, unless the
+	 * memory map has changed due to process exit or fd copying.  Leak the
+	 * mapping on failure, e.g. if the task is killed, worst case scenario,
+	 * the page(s) will be reclaimed when the process exits.
+	 */
+	if (current->mm == kvm->mm && slot->id >= KVM_USER_MEM_SLOTS &&
+	    !WARN_ON_ONCE(!slot->npages))
+		vm_munmap(slot->userspace_addr, slot->npages * PAGE_SIZE);
 }
 
 int memslot_rmap_alloc(struct kvm_memory_slot *slot, unsigned long npages)
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 04/10] KVM: Disallow setting memslots when the VM is being destroyed
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (2 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 03/10] KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is freed Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 05/10] KVM: Destroy memslots immediately after mmu_notifiers are unregistered Sean Christopherson
                   ` (6 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Now that KVM doesn't delete KVM-internal memslots as an unnecessary side
effect during VM destruction, WARN and reject any attempt to set memslots
after the VM's refcount has hit 0.  There's obviously no need to CREATE,
MOVE, or do a FLAGS_ONLY update when a VM is being destroyed, and there
should be no reason for arch code to manually DELETE a memslot: once KVM
KVM unregisters its mmu_notifier and does the final kvm_flush_shadow_all(),
there absolutely must not be any outstanding references to memslots.  I.e.
if arch code "needs" to manually DELETE a memslot, then it's already buggy.

Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 virt/kvm/kvm_main.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 9a24c3064896..f368240aa1cd 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -2015,6 +2015,9 @@ static int kvm_set_memory_region(struct kvm *kvm,
 
 	lockdep_assert_held(&kvm->slots_lock);
 
+	if (WARN_ON_ONCE(!refcount_read(&kvm->users_count)))
+		return -EIO;
+
 	r = check_memory_region_flags(kvm, mem);
 	if (r)
 		return r;
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 05/10] KVM: Destroy memslots immediately after mmu_notifiers are unregistered
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (3 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 04/10] KVM: Disallow setting memslots when the VM is being destroyed Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 06/10] KVM: WARN if KVM attempts to do guest-related uaccess with "wrong" process Sean Christopherson
                   ` (5 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Instally dummy, empty memslots immediately after unregistering KVM's
mmu_notifier during VM destruction to harden against accessing memory from
the wrong address space when tearing down a VM.  Because kvm_destroy_vm()
often runs when the associated VM's file is being released, current->mm is
often no longer kvm->mm, i.e. using the memslots to access userspace memory
is inherently broken/dangerous.

While KVM's APIs to read/write guest memory explicitly reject accesses if
current->mm != kvm->mm, taking away the memslots adds another layer of
defense and helps guard against rogue accesses that don't go through KVM's
standard API, or that do GUP+kmap().

Note, simply hoisting memslot destruction above kvm_arch_destroy_vm()
without installing dummy slots is not a viable alternative.  Doing so would
require auditing all the paths of kvm_arch_destroy_vm(), which is nearly
infeasible, and missing even one case would result in a NULL pointer
dereference and/or use-after-free, neither of which is a substantially
better outcome than the status quo.

Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 virt/kvm/kvm_main.c | 45 +++++++++++++++++++++++++++++++--------------
 1 file changed, 31 insertions(+), 14 deletions(-)

diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index f368240aa1cd..770a2c3bd445 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -955,23 +955,42 @@ static void kvm_free_memslot(struct kvm *kvm, struct kvm_memory_slot *slot)
 	kfree(slot);
 }
 
-static void kvm_free_memslots(struct kvm *kvm, struct kvm_memslots *slots)
+static const struct kvm_memslots kvm_empty_memslots = {
+	.generation = -1ull,
+	.hva_tree = RB_ROOT_CACHED,
+	.gfn_tree = RB_ROOT,
+	.id_hash[0 ... (ARRAY_SIZE(kvm_empty_memslots.id_hash) - 1)] = HLIST_HEAD_INIT,
+	.node_idx = 0,
+};
+
+static void kvm_destroy_memslots(struct kvm *kvm)
 {
 	struct hlist_node *idnode;
 	struct kvm_memory_slot *memslot;
-	int bkt;
+	int bkt, i;
+
+	/*
+	 * Install empty memslots to guard against memslot lookups while the VM
+	 * is being destroyed.  Consuming memslots at this stage is a KVM bug,
+	 * but "gracefully do nothing" is a much better outcome than "crash the
+	 * host" or "corrupt random memory" when there inevitably is a bug.
+	 */
+	mutex_lock(&kvm->slots_lock);
+	for (i = 0; i < kvm_arch_nr_memslot_as_ids(kvm); i++)
+		rcu_assign_pointer(kvm->memslots[i], &kvm_empty_memslots);
+
+	synchronize_srcu_expedited(&kvm->srcu);
+	mutex_unlock(&kvm->slots_lock);
 
 	/*
 	 * The same memslot objects live in both active and inactive sets,
-	 * arbitrarily free using index '1' so the second invocation of this
-	 * function isn't operating over a structure with dangling pointers
-	 * (even though this function isn't actually touching them).
+	 * arbitrarily free using index '1'.
 	 */
-	if (!slots->node_idx)
-		return;
-
-	hash_for_each_safe(slots->id_hash, bkt, idnode, memslot, id_node[1])
-		kvm_free_memslot(kvm, memslot);
+	for (i = 0; i < kvm_arch_nr_memslot_as_ids(kvm); i++) {
+		hash_for_each_safe(kvm->__memslots[i][1].id_hash, bkt, idnode,
+				   memslot, id_node[1])
+			kvm_free_memslot(kvm, memslot);
+	}
 }
 
 static umode_t kvm_stats_debugfs_mode(const struct kvm_stats_desc *desc)
@@ -1302,12 +1321,10 @@ static void kvm_destroy_vm(struct kvm *kvm)
 		kvm->mn_active_invalidate_count = 0;
 	else
 		WARN_ON(kvm->mmu_invalidate_in_progress);
+	kvm_destroy_memslots(kvm);
+
 	kvm_arch_destroy_vm(kvm);
 	kvm_destroy_devices(kvm);
-	for (i = 0; i < kvm_arch_nr_memslot_as_ids(kvm); i++) {
-		kvm_free_memslots(kvm, &kvm->__memslots[i][0]);
-		kvm_free_memslots(kvm, &kvm->__memslots[i][1]);
-	}
 	cleanup_srcu_struct(&kvm->irq_srcu);
 	srcu_barrier(&kvm->srcu);
 	cleanup_srcu_struct(&kvm->srcu);
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 06/10] KVM: WARN if KVM attempts to do guest-related uaccess with "wrong" process
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (4 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 05/10] KVM: Destroy memslots immediately after mmu_notifiers are unregistered Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 07/10] KVM: WARN and reject guest-based uaccess if VM is dying Sean Christopherson
                   ` (4 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Now that memslots are torn down before kvm_arch_destroy_vm(), i.e. now that
the massive, unaudited path that is known to have at least one rogue access
to guest memory should be short-circuited before attempting to access user
memory, WARN if KVM attempts to do a user access to guest memory with the
wrong address space.

Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 include/linux/kvm_host.h | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index b37cf275ef79..2a88e4ce145e 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -1352,7 +1352,7 @@ int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 
 static __always_inline __must_check bool kvm_can_do_uaccess(struct kvm *kvm)
 {
-	return current->mm == kvm->mm;
+	return !WARN_ON_ONCE(current->mm != kvm->mm);
 }
 
 #define BUILD_KVM_COPY_USER_WRAPPER(fn, to_user, from_user)				\
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 07/10] KVM: WARN and reject guest-based uaccess if VM is dying
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (5 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 06/10] KVM: WARN if KVM attempts to do guest-related uaccess with "wrong" process Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 08/10] KVM: nVMX: Don't flush shadow VMCS12 to guest memory during vCPU teardown Sean Christopherson
                   ` (3 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

WARN and reject user accesses to guest memory if the associated VM is dying
even if the current address space happens to be the correct address space.
Accessing guest memory after the last reference to the VM has been put may
be "fine" from a safety perspective, but it's still a KVM bug.

Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 include/linux/kvm_host.h | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 2a88e4ce145e..0ee81754d730 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -1352,7 +1352,8 @@ int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 
 static __always_inline __must_check bool kvm_can_do_uaccess(struct kvm *kvm)
 {
-	return !WARN_ON_ONCE(current->mm != kvm->mm);
+	return !WARN_ON_ONCE(current->mm != kvm->mm ||
+			     !refcount_read(&kvm->users_count));
 }
 
 #define BUILD_KVM_COPY_USER_WRAPPER(fn, to_user, from_user)				\
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 08/10] KVM: nVMX: Don't flush shadow VMCS12 to guest memory during vCPU teardown
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (6 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 07/10] KVM: WARN and reject guest-based uaccess if VM is dying Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 09/10] KVM: nVMX: Don't try to load eVMCS12 page when the VM is dying Sean Christopherson
                   ` (2 subsequent siblings)
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

From: Jim Mattson <jmattson@google.com>

When a vCPU is destroyed while L2 is active, KVM synthesizes a nested
VM-Exit, which flushes the cached shadow VMCS12 back to guest memory:

  vmx_vcpu_free()
  |-> nested_vmx_free_vcpu()
      |-> vmx_leave_nested()
          |-> nested_vmx_vmexit(vcpu, -1, 0, 0)
              |-> nested_flush_cached_shadow_vmcs12()
                  |-> kvm_write_guest_cached()
                      |-> __copy_to_user(ghc->hva, ...)

Accessing user memory via a memslot during VM destruction is broken, as
there are no guarantees that current->mm == kvm->mm when the VM is dying,
because the last reference to the VM can be put from a different process
than the original creating processes.

And even if the original process does put the final reference, during
process exit, do_exit() calls exit_mm() before closing file descriptors, so
vCPU destruction runs with current->mm == NULL on a borrowed lazy TLB
active_mm.  If the borrowed address space has a writable mapping at the
to-be-written userspace address, KVM will corrupt an unrelated task's
memory since uaccess APIs, including __copy_to_user(), don't sanity check
current->mm (and *can't* sanity perform KVM's current->mm == kvm->mm check
since that is firmly a KVM-only concept).

Hack-a-fix the nVMX flow even though KVM now protects against bad uaccess
reads/writes in the core APIs, as doing so will allow adding even more
sanity checks in KVM's APIs to help detect other buggy code.  Add a TODO to
call out that checking if KVM can do a uaccess for the VM is a hack; nVMX
really needs to stop abusing __nested_vmx_vmexit() when destroying a vCPU.

Fixes: 61ada7488ffd ("KVM: nVMX: Cache shadow vmcs12 on VMEntry and flush to memory on VMExit")
Signed-off-by: Jim Mattson <jmattson@google.com>
[sean: key off __kvm_can_do_uaccess(), add TODO]
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 arch/x86/kvm/vmx/nested.c | 7 ++++++-
 include/linux/kvm_host.h  | 8 ++++++--
 2 files changed, 12 insertions(+), 3 deletions(-)

diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
index 151873407abd..8e31eba4d9fa 100644
--- a/arch/x86/kvm/vmx/nested.c
+++ b/arch/x86/kvm/vmx/nested.c
@@ -5134,8 +5134,13 @@ void __nested_vmx_vmexit(struct kvm_vcpu *vcpu, u32 vm_exit_reason,
 		 * Otherwise, this flush will dirty guest memory at a
 		 * point it is already assumed by user-space to be
 		 * immutable.
+		 *
+		 * TODO: Drop the explicit check on being able to access guest
+		 *       memory once KVM no longer abuses the nested VM-Exit
+		 *       flow when destroying a vCPU.
 		 */
-		nested_flush_cached_shadow_vmcs12(vcpu, vmcs12);
+		if (__kvm_can_do_uaccess(vcpu->kvm))
+			nested_flush_cached_shadow_vmcs12(vcpu, vmcs12);
 	} else {
 		/*
 		 * The only expected VM-instruction error is "VM entry with
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 0ee81754d730..0c58a4945595 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -1350,10 +1350,14 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 			      gpa_t gpa, unsigned long len);
 
+static __always_inline __must_check bool __kvm_can_do_uaccess(struct kvm *kvm)
+{
+	return current->mm == kvm->mm && refcount_read(&kvm->users_count);
+}
+
 static __always_inline __must_check bool kvm_can_do_uaccess(struct kvm *kvm)
 {
-	return !WARN_ON_ONCE(current->mm != kvm->mm ||
-			     !refcount_read(&kvm->users_count));
+	return !WARN_ON_ONCE(!__kvm_can_do_uaccess(kvm));
 }
 
 #define BUILD_KVM_COPY_USER_WRAPPER(fn, to_user, from_user)				\
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 09/10] KVM: nVMX: Don't try to load eVMCS12 page when the VM is dying
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (7 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 08/10] KVM: nVMX: Don't flush shadow VMCS12 to guest memory during vCPU teardown Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-01 20:22 ` [PATCH v2 10/10] KVM: Pre-check uaccesses in KVM's APIs to read/write guest memory Sean Christopherson
  2026-10-02 20:30 ` [syzbot ci] Re: KVM: Fix+harden against bad uaccess using dying VM syzbot ci
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Don't try to get/load the eVMCS12 page when KVM can't do uaccesses to guest
memory, i.e. when the VM dying, as reading/writing guest memory via the
associated userspace address space is obviously broken if the address space
is inactive and/or has already been torn down.

Hack-a-fix the eVMCS code even though KVM now protects against bad uaccess
reads/writes in the core APIs, as doing so will allow adding WARNs in said
APIs to help detect other buggy code.  Add a TODO to call out that checking
if KVM can do a uaccess for the VM is a hack; nVMX really needs to stop
abusing __nested_vmx_vmexit() when destroying a vCPU.

E.g. without the "fix", adding the sanity checks will trigger:

  ------------[ cut here ]------------
  WARNING: ./include/linux/kvm_host.h:1359 at kvm_read_guest_offset_cached+0x1b0/0x200 [kvm], CPU#14: hyperv_evmcs/27428
  Call Trace:
   <TASK>
   nested_get_evmptr+0x4f/0x80 [kvm_intel]
   nested_vmx_handle_enlightened_vmptrld+0x2a/0x130 [kvm_intel]
   __nested_vmx_vmexit+0xbe/0x8e0 [kvm_intel]
   nested_vmx_free_vcpu+0x36/0x50 [kvm_intel]
   vmx_vcpu_free+0x73/0x120 [kvm_intel]
   kvm_arch_vcpu_destroy+0x81/0x1a0 [kvm]
   kvm_destroy_vcpus+0x9d/0x160 [kvm]
   kvm_arch_destroy_vm+0x25/0x90 [kvm]
   kvm_put_kvm+0x342/0x460 [kvm]
   kvm_vm_stats_release+0x12/0x20 [kvm]
   __fput+0xfe/0x280
   __se_sys_close+0x76/0xe0
   do_syscall_64+0xfe/0x440
   entry_SYSCALL_64_after_hwframe+0x4b/0x53
  RIP: 0033:0x4a3822
   </TASK>
  ---[ end trace 0000000000000000 ]---

Fixes: f5c7e8425f18 ("KVM: nVMX: Always make an attempt to map eVMCS after migration")
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 arch/x86/kvm/vmx/nested.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
index 8e31eba4d9fa..07ec0dd014b5 100644
--- a/arch/x86/kvm/vmx/nested.c
+++ b/arch/x86/kvm/vmx/nested.c
@@ -5091,8 +5091,13 @@ void __nested_vmx_vmexit(struct kvm_vcpu *vcpu, u32 vm_exit_reason,
 		 * Enlightened VMCS after migration and we still need to
 		 * do that when something is forcing L2->L1 exit prior to
 		 * the first L2 run.
+		 *
+		 * TODO: Drop the explicit check on being able to access guest
+		 *       memory once KVM no longer abuses the nested VM-Exit
+		 *       flow when destroying a vCPU.
 		 */
-		(void)nested_get_evmcs_page(vcpu);
+		if (__kvm_can_do_uaccess(vcpu->kvm))
+			(void)nested_get_evmcs_page(vcpu);
 #endif
 	}
 
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v2 10/10] KVM: Pre-check uaccesses in KVM's APIs to read/write guest memory
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (8 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 09/10] KVM: nVMX: Don't try to load eVMCS12 page when the VM is dying Sean Christopherson
@ 2026-10-01 20:22 ` Sean Christopherson
  2026-10-02 20:30 ` [syzbot ci] Re: KVM: Fix+harden against bad uaccess using dying VM syzbot ci
  10 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 20:22 UTC (permalink / raw)
  To: Madhavan Srinivasan, Sean Christopherson, Paolo Bonzini
  Cc: Nicholas Piggin, linuxppc-dev, kvm, linux-kernel, Jim Mattson

Add a check on whether or not doing a user access is ok in KVM's APIs for
accessing guest memory, *before* doing memslot lookup/validation, so that
buggy KVM accesses to guest memory more likely to fail noisily.  Because
KVM installs dummy/empty memslots early in VM destruction, bugs in the VM
teardown path will fail on the memslot checks without getting to the WARNs
buried in kvm_copy_{to,from}_user() and friends.

Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 virt/kvm/kvm_main.c | 15 ++++++++++-----
 1 file changed, 10 insertions(+), 5 deletions(-)

diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 770a2c3bd445..80b19d47840b 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -3211,6 +3211,11 @@ static int next_segment(unsigned long len, int offset)
 		return len;
 }
 
+static bool kvm_is_guest_access_ok(struct kvm *kvm, int offset, int len, int size)
+{
+	return kvm_can_do_uaccess(kvm) && !WARN_ON_ONCE(offset + len > PAGE_SIZE);
+}
+
 /* Copy @len bytes from guest memory at '(@gfn * PAGE_SIZE) + @offset' to @data */
 static int __kvm_read_guest_page(struct kvm *kvm, struct kvm_memory_slot *slot,
 				 gfn_t gfn, void *data, int offset, int len)
@@ -3218,7 +3223,7 @@ static int __kvm_read_guest_page(struct kvm *kvm, struct kvm_memory_slot *slot,
 	int r;
 	unsigned long addr;
 
-	if (WARN_ON_ONCE(offset + len > PAGE_SIZE))
+	if (!kvm_is_guest_access_ok(kvm, offset, len, PAGE_SIZE))
 		return -EFAULT;
 
 	addr = gfn_to_hva_memslot_prot(slot, gfn, NULL);
@@ -3294,7 +3299,7 @@ static int __kvm_read_guest_atomic(struct kvm *kvm, struct kvm_memory_slot *slot
 	int r;
 	unsigned long addr;
 
-	if (WARN_ON_ONCE(offset + len > PAGE_SIZE))
+	if (!kvm_is_guest_access_ok(kvm, offset, len, PAGE_SIZE))
 		return -EFAULT;
 
 	addr = gfn_to_hva_memslot_prot(slot, gfn, NULL);
@@ -3327,7 +3332,7 @@ static int __kvm_write_guest_page(struct kvm *kvm,
 	int r;
 	unsigned long addr;
 
-	if (WARN_ON_ONCE(offset + len > PAGE_SIZE))
+	if (!kvm_is_guest_access_ok(kvm, offset, len, PAGE_SIZE))
 		return -EFAULT;
 
 	addr = gfn_to_hva_memslot(memslot, gfn);
@@ -3457,7 +3462,7 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 	int r;
 	gpa_t gpa = ghc->gpa + offset;
 
-	if (WARN_ON_ONCE(len + offset > ghc->len))
+	if (!kvm_is_guest_access_ok(kvm, offset, len, ghc->len))
 		return -EINVAL;
 
 	if (slots->generation != ghc->generation) {
@@ -3495,7 +3500,7 @@ int kvm_read_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc,
 	int r;
 	gpa_t gpa = ghc->gpa + offset;
 
-	if (WARN_ON_ONCE(len + offset > ghc->len))
+	if (!kvm_is_guest_access_ok(kvm, offset, len, ghc->len))
 		return -EINVAL;
 
 	if (slots->generation != ghc->generation) {
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm
  2026-10-01 20:22 ` [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm Sean Christopherson
@ 2026-10-01 21:08   ` James Houghton
  2026-10-01 21:18     ` Sean Christopherson
  0 siblings, 1 reply; 17+ messages in thread
From: James Houghton @ 2026-10-01 21:08 UTC (permalink / raw)
  To: Sean Christopherson
  Cc: Madhavan Srinivasan, Paolo Bonzini, Nicholas Piggin,
	linuxppc-dev, kvm, linux-kernel, Jim Mattson

On Thu, Oct 1, 2026 at 1:24 PM Sean Christopherson <seanjc@google.com> wrote:
>
> Reject user accesses to guest memory, which are supposed to be done only
> in the context of KVM_RUN or similar operations, if the current address
> space is not the VM's (host userspace) address space.  If KVM writes to
> guest memory after the owning host process has exited, or if the VM is
> being destroyed in the context of a different process, then writing using
> the wrong address space will corrupt a different process' memory.
>
> Reject the access but don't WARN() or KVM_BUG_ON() event though attempting
> to access guest memory with a mismatched address space is a blatant KVM
> bug, because unfortunately KVM is buggy.  On KVM VMX, when a vCPU is
> destroyed while L2 is active, KVM synthesizes a nested VM-Exit to force the
> vCPU out of L2 in order to free the nested VMX assets, and a side effect of
> a nested VM-Exit is that it flushes the cached shadow VMCS12 back to guest
> memory:
>
>   vmx_vcpu_free()
>   |-> nested_vmx_free_vcpu()
>       |-> vmx_leave_nested()
>           |-> nested_vmx_vmexit(vcpu, -1, 0, 0)
>               |-> nested_flush_cached_shadow_vmcs12()
>                   |-> kvm_write_guest_cached()
>                       |-> __copy_to_user(ghc->hva, ...)
>
> Fix the bug broadly even though the "real" bug is that KVM abuses the
> nested VM-Exit flow for non-architectural purposes, as there may be other
> such violations lurking.  For now, punt on fixing individual bugs and
> hardening the common flows, e.g. with WARNs.
>
> Opportunistically provide wrappers in anticipation of adding more checks
> and hardening, i.e. growing the logic beyond checking current->mm.
>
> Fixes: 61ada7488ffd ("KVM: nVMX: Cache shadow vmcs12 on VMEntry and flush to memory on VMExit")
> Cc: stable@vger.kernel.org
> Reported-by: Jim Mattson <jmattson@google.com>
> Closes: https://lore.kernel.org/all/20260908132838.2116068-1-jmattson@google.com
> Signed-off-by: Sean Christopherson <seanjc@google.com>

Thanks, Sean. Feel free to add:

Reviewed-by: James Houghton <jthoughton@google.com>

I wonder if it makes sense to add similar hardening to
kvm_faultin_pfn(). What do you think?

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm
  2026-10-01 21:08   ` James Houghton
@ 2026-10-01 21:18     ` Sean Christopherson
  2026-10-01 21:28       ` James Houghton
  0 siblings, 1 reply; 17+ messages in thread
From: Sean Christopherson @ 2026-10-01 21:18 UTC (permalink / raw)
  To: James Houghton
  Cc: Madhavan Srinivasan, Paolo Bonzini, Nicholas Piggin,
	linuxppc-dev, kvm, linux-kernel, Jim Mattson

On Thu, Oct 01, 2026, James Houghton wrote:
> On Thu, Oct 1, 2026 at 1:24 PM Sean Christopherson <seanjc@google.com> wrote:
> >
> > Reject user accesses to guest memory, which are supposed to be done only
> > in the context of KVM_RUN or similar operations, if the current address
> > space is not the VM's (host userspace) address space.  If KVM writes to
> > guest memory after the owning host process has exited, or if the VM is
> > being destroyed in the context of a different process, then writing using
> > the wrong address space will corrupt a different process' memory.
> >
> > Reject the access but don't WARN() or KVM_BUG_ON() event though attempting
> > to access guest memory with a mismatched address space is a blatant KVM
> > bug, because unfortunately KVM is buggy.  On KVM VMX, when a vCPU is
> > destroyed while L2 is active, KVM synthesizes a nested VM-Exit to force the
> > vCPU out of L2 in order to free the nested VMX assets, and a side effect of
> > a nested VM-Exit is that it flushes the cached shadow VMCS12 back to guest
> > memory:
> >
> >   vmx_vcpu_free()
> >   |-> nested_vmx_free_vcpu()
> >       |-> vmx_leave_nested()
> >           |-> nested_vmx_vmexit(vcpu, -1, 0, 0)
> >               |-> nested_flush_cached_shadow_vmcs12()
> >                   |-> kvm_write_guest_cached()
> >                       |-> __copy_to_user(ghc->hva, ...)
> >
> > Fix the bug broadly even though the "real" bug is that KVM abuses the
> > nested VM-Exit flow for non-architectural purposes, as there may be other
> > such violations lurking.  For now, punt on fixing individual bugs and
> > hardening the common flows, e.g. with WARNs.
> >
> > Opportunistically provide wrappers in anticipation of adding more checks
> > and hardening, i.e. growing the logic beyond checking current->mm.
> >
> > Fixes: 61ada7488ffd ("KVM: nVMX: Cache shadow vmcs12 on VMEntry and flush to memory on VMExit")
> > Cc: stable@vger.kernel.org
> > Reported-by: Jim Mattson <jmattson@google.com>
> > Closes: https://lore.kernel.org/all/20260908132838.2116068-1-jmattson@google.com
> > Signed-off-by: Sean Christopherson <seanjc@google.com>
> 
> Thanks, Sean. Feel free to add:
> 
> Reviewed-by: James Houghton <jthoughton@google.com>
> 
> I wonder if it makes sense to add similar hardening to kvm_faultin_pfn().
> What do you think?

I'm not opposed to explicitly hardening kvm_faultin_pfn(), but I don't think it
would add much value in practice.  Far more arch code uses __kvm_faultin_pfn()
directly, and that doesn't have a @vcpu or @vm pointer to do the check.  We could
obviously "fix" that, but nuking the memslots (patches 3-5) will prevent all but
the most ridiculous bugs.  Getting anywhere near __kvm_faultin_pfn() with the
wrong mm would either mean KVM is doing something amazingly stupid during VM
teardown, or I guess maybe the scheduler or preempt notifiers went off the rails?

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v2 03/10] KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is freed
  2026-10-01 20:22 ` [PATCH v2 03/10] KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is freed Sean Christopherson
@ 2026-10-01 21:26   ` James Houghton
  0 siblings, 0 replies; 17+ messages in thread
From: James Houghton @ 2026-10-01 21:26 UTC (permalink / raw)
  To: Sean Christopherson
  Cc: Madhavan Srinivasan, Paolo Bonzini, Nicholas Piggin,
	linuxppc-dev, kvm, linux-kernel, Jim Mattson

On Thu, Oct 1, 2026 at 1:26 PM Sean Christopherson <seanjc@google.com> wrote:
>
> Unmap KVM-created VMAs for KVM-internal memslots when the associated
> memslot is freed instead of unmapping via x86's internal API.  This will
> allow deleting/freeing all memslots before kvm_arch_destroy_vm() without
> leaking VMAs, and in general is a net positive as it makes it far less
> likely that KVM will leak VMAs, e.g. in the unlikely scenario some other
> flow deletes KVM-internal memslots without going through
> __x86_set_memory_region().
>
> Signed-off-by: Sean Christopherson <seanjc@google.com>

LGTM; this code is so much better. Feel free to add:

Reviewed-by: James Houghton <jthoughton@google.com>

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm
  2026-10-01 21:18     ` Sean Christopherson
@ 2026-10-01 21:28       ` James Houghton
  0 siblings, 0 replies; 17+ messages in thread
From: James Houghton @ 2026-10-01 21:28 UTC (permalink / raw)
  To: Sean Christopherson
  Cc: Madhavan Srinivasan, Paolo Bonzini, Nicholas Piggin,
	linuxppc-dev, kvm, linux-kernel, Jim Mattson

On Thu, Oct 1, 2026 at 2:18 PM Sean Christopherson <seanjc@google.com> wrote:
>
> On Thu, Oct 01, 2026, James Houghton wrote:
> > I wonder if it makes sense to add similar hardening to kvm_faultin_pfn().
> > What do you think?
>
> I'm not opposed to explicitly hardening kvm_faultin_pfn(), but I don't think it
> would add much value in practice.  Far more arch code uses __kvm_faultin_pfn()
> directly, and that doesn't have a @vcpu or @vm pointer to do the check.  We could
> obviously "fix" that, but nuking the memslots (patches 3-5) will prevent all but
> the most ridiculous bugs.  Getting anywhere near __kvm_faultin_pfn() with the
> wrong mm would either mean KVM is doing something amazingly stupid during VM
> teardown, or I guess maybe the scheduler or preempt notifiers went off the rails?

Yeah, fair enough. It would be of pretty limited benefit. Thanks.

^ permalink raw reply	[flat|nested] 17+ messages in thread

* [syzbot ci] Re: KVM: Fix+harden against bad uaccess using dying VM
  2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
                   ` (9 preceding siblings ...)
  2026-10-01 20:22 ` [PATCH v2 10/10] KVM: Pre-check uaccesses in KVM's APIs to read/write guest memory Sean Christopherson
@ 2026-10-02 20:30 ` syzbot ci
  2026-10-02 20:39   ` Sean Christopherson
  10 siblings, 1 reply; 17+ messages in thread
From: syzbot ci @ 2026-10-02 20:30 UTC (permalink / raw)
  To: jmattson, kvm, linux-kernel, linuxppc-dev, maddy, npiggin,
	pbonzini, seanjc
  Cc: syzbot, syzkaller-bugs

syzbot ci has tested the following series

[v2] KVM: Fix+harden against bad uaccess using dying VM
https://lore.kernel.org/all/20261001202234.3794060-1-seanjc@google.com
* [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm
* [PATCH v2 02/10] KVM: PPC: Flush/zap all memslots on kvm_arch_flush_shadow_all()
* [PATCH v2 03/10] KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is freed
* [PATCH v2 04/10] KVM: Disallow setting memslots when the VM is being destroyed
* [PATCH v2 05/10] KVM: Destroy memslots immediately after mmu_notifiers are unregistered
* [PATCH v2 06/10] KVM: WARN if KVM attempts to do guest-related uaccess with "wrong" process
* [PATCH v2 07/10] KVM: WARN and reject guest-based uaccess if VM is dying
* [PATCH v2 08/10] KVM: nVMX: Don't flush shadow VMCS12 to guest memory during vCPU teardown
* [PATCH v2 09/10] KVM: nVMX: Don't try to load eVMCS12 page when the VM is dying
* [PATCH v2 10/10] KVM: Pre-check uaccesses in KVM's APIs to read/write guest memory

and found the following issue:
WARNING in __kvm_read_guest_page

Full report is available here:
https://ci.syzbot.org/series/5f8da53a-c584-4176-abcb-7a29eb02e4c4

***

WARNING in __kvm_read_guest_page

tree:      kvm-next
URL:       https://kernel.googlesource.com/pub/scm/virt/kvm/kvm/
base:      d4b7fb647204f0c81dfeae2d1a708e4d858e0c94
arch:      amd64
compiler:  Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
config:    https://ci.syzbot.org/builds/4d4c10d0-f56c-448f-8782-ad4dab8f46d0/config
syz repro: https://ci.syzbot.org/findings/dc5f6c42-c3ab-41b0-9384-1a15798e4820/syz_repro

------------[ cut here ]------------
!__kvm_can_do_uaccess(kvm)
WARNING: ./include/linux/kvm_host.h:1360 at kvm_can_do_uaccess include/linux/kvm_host.h:1360 [inline], CPU#1: syz.1.18/5858
WARNING: ./include/linux/kvm_host.h:1360 at kvm_is_guest_access_ok virt/kvm/kvm_main.c:3216 [inline], CPU#1: syz.1.18/5858
WARNING: ./include/linux/kvm_host.h:1360 at __kvm_read_guest_page+0x38e/0x440 virt/kvm/kvm_main.c:3226, CPU#1: syz.1.18/5858
Modules linked in:
CPU: 1 UID: 0 PID: 5858 Comm: syz.1.18 Not tainted syzkaller #0 PREEMPT(full) 
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014
RIP: 0010:kvm_can_do_uaccess include/linux/kvm_host.h:1360 [inline]
RIP: 0010:kvm_is_guest_access_ok virt/kvm/kvm_main.c:3216 [inline]
RIP: 0010:__kvm_read_guest_page+0x38e/0x440 virt/kvm/kvm_main.c:3226
Code: f2 ff ff ff 0f 44 d8 31 ff e8 5e a4 89 00 89 d8 48 83 c4 30 5b 41 5c 41 5d 41 5e 41 5f 5d e9 09 ac a8 0a cc e8 83 9e 89 00 90 <0f> 0b 90 bb f2 ff ff ff eb da e8 73 9e 89 00 90 0f 0b 90 bb f2 ff
RSP: 0018:ffffc90003157778 EFLAGS: 00010293
RAX: ffffffff813e2d22 RBX: 0000000000000000 RCX: ffff888174382580
RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000000
RBP: 1ffff1102217d82a R08: ffff888110bee183 R09: 1ffff1102217dc30
R10: dffffc0000000000 R11: ffffed102217dc31 R12: ffff888110bee180
R13: ffff888174382b40 R14: 1ffff1102217dc30 R15: 0000000000000000
FS:  000055557c0db500(0000) GS:ffff8882a8cda000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007f458dfeb840 CR3: 000000016c9d0000 CR4: 0000000000352ef0
Call Trace:
 <TASK>
 kvm_vcpu_read_guest+0x64/0x140 virt/kvm/kvm_main.c:3284
 nested_vmx_load_msr+0x133/0x4d0 arch/x86/kvm/vmx/nested.c:1107
 load_vmcs12_host_state+0x10ca/0x1e20 arch/x86/kvm/vmx/nested.c:4930
 __nested_vmx_vmexit+0x1800/0x2d10 arch/x86/kvm/vmx/nested.c:5206
 nested_vmx_vmexit arch/x86/kvm/vmx/nested.h:46 [inline]
 vmx_leave_nested arch/x86/kvm/vmx/nested.c:6891 [inline]
 nested_vmx_free_vcpu+0x8e/0xd0 arch/x86/kvm/vmx/nested.c:389
 vmx_vcpu_free+0x109/0x3e0 arch/x86/kvm/vmx/vmx.c:7653
 kvm_arch_vcpu_destroy+0x154/0x380 arch/x86/kvm/x86.c:9463
 kvm_vcpu_destroy virt/kvm/kvm_main.c:469 [inline]
 kvm_destroy_vcpus+0x123/0x380 virt/kvm/kvm_main.c:489
 kvm_arch_destroy_vm+0x55/0x120 arch/x86/kvm/x86.c:10014
 kvm_destroy_vm virt/kvm/kvm_main.c:1326 [inline]
 kvm_put_kvm+0x93a/0xc80 virt/kvm/kvm_main.c:1359
 kvm_vm_release+0x43/0x50 virt/kvm/kvm_main.c:1382
 __fput+0x418/0xa50 fs/file_table.c:512
 task_work_run+0x1d9/0x270 kernel/task_work.c:233
 resume_user_mode_work include/linux/resume_user_mode.h:50 [inline]
 __exit_to_user_mode_loop kernel/entry/common.c:70 [inline]
 exit_to_user_mode_loop+0x204/0x770 kernel/entry/common.c:101
 __exit_to_user_mode_prepare include/linux/irq-entry-common.h:207 [inline]
 syscall_exit_to_user_mode_prepare include/linux/irq-entry-common.h:230 [inline]
 syscall_exit_to_user_mode include/linux/entry-common.h:336 [inline]
 do_syscall_64+0x328/0x520 arch/x86/entry/syscall_64.c:89
 entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f7fec79e159
Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 e8 ff ff ff f7 d8 64 89 01 48
RSP: 002b:00007fff0d2850c8 EFLAGS: 00000246 ORIG_RAX: 00000000000001b4
RAX: 0000000000000000 RBX: 00007f7feca27da0 RCX: 00007f7fec79e159
RDX: 0000000000000000 RSI: 000000000000001e RDI: 0000000000000003
RBP: 00007f7feca27da0 R08: 00007f7feca26038 R09: 0000000000000000
R10: 000000000003fdb8 R11: 0000000000000246 R12: 000000000000ec04
R13: 00007f7feca25fac R14: 000000000000e938 R15: 00007fff0d2851d0
 </TASK>


***

If these findings have caused you to resend the series or submit a
separate fix, please add the following tag to your commit message:
  Tested-by: syzbot@syzkaller.appspotmail.com

---
This report is generated by a bot. It may contain errors.
syzbot ci engineers can be reached at syzkaller@googlegroups.com.

To test a fix for this bug, please reply with `#syz test`
(on a separate line) and attach the patch to the email.

Notes:
- The patch will be applied on top of the tested series (as an
  incremental fix).
- To test a new version of the whole series, please send it directly
  to syzbot@lists.linux.dev.
- Arguments like custom git repos and branches are not supported.

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [syzbot ci] Re: KVM: Fix+harden against bad uaccess using dying VM
  2026-10-02 20:30 ` [syzbot ci] Re: KVM: Fix+harden against bad uaccess using dying VM syzbot ci
@ 2026-10-02 20:39   ` Sean Christopherson
  0 siblings, 0 replies; 17+ messages in thread
From: Sean Christopherson @ 2026-10-02 20:39 UTC (permalink / raw)
  To: syzbot ci
  Cc: jmattson, kvm, linux-kernel, linuxppc-dev, maddy, npiggin,
	pbonzini, syzbot, syzkaller-bugs

On Fri, Oct 02, 2026, syzbot ci wrote:
> ------------[ cut here ]------------
> !__kvm_can_do_uaccess(kvm)
> WARNING: ./include/linux/kvm_host.h:1360 at kvm_can_do_uaccess include/linux/kvm_host.h:1360 [inline], CPU#1: syz.1.18/5858
> WARNING: ./include/linux/kvm_host.h:1360 at kvm_is_guest_access_ok virt/kvm/kvm_main.c:3216 [inline], CPU#1: syz.1.18/5858
> WARNING: ./include/linux/kvm_host.h:1360 at __kvm_read_guest_page+0x38e/0x440 virt/kvm/kvm_main.c:3226, CPU#1: syz.1.18/5858
> Modules linked in:
> CPU: 1 UID: 0 PID: 5858 Comm: syz.1.18 Not tainted syzkaller #0 PREEMPT(full) 
> Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014
> RIP: 0010:kvm_can_do_uaccess include/linux/kvm_host.h:1360 [inline]
> RIP: 0010:kvm_is_guest_access_ok virt/kvm/kvm_main.c:3216 [inline]
> RIP: 0010:__kvm_read_guest_page+0x38e/0x440 virt/kvm/kvm_main.c:3226
> Code: f2 ff ff ff 0f 44 d8 31 ff e8 5e a4 89 00 89 d8 48 83 c4 30 5b 41 5c 41 5d 41 5e 41 5f 5d e9 09 ac a8 0a cc e8 83 9e 89 00 90 <0f> 0b 90 bb f2 ff ff ff eb da e8 73 9e 89 00 90 0f 0b 90 bb f2 ff
> RSP: 0018:ffffc90003157778 EFLAGS: 00010293
> RAX: ffffffff813e2d22 RBX: 0000000000000000 RCX: ffff888174382580
> RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000000
> RBP: 1ffff1102217d82a R08: ffff888110bee183 R09: 1ffff1102217dc30
> R10: dffffc0000000000 R11: ffffed102217dc31 R12: ffff888110bee180
> R13: ffff888174382b40 R14: 1ffff1102217dc30 R15: 0000000000000000
> FS:  000055557c0db500(0000) GS:ffff8882a8cda000(0000) knlGS:0000000000000000
> CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> CR2: 00007f458dfeb840 CR3: 000000016c9d0000 CR4: 0000000000352ef0
> Call Trace:
>  <TASK>
>  kvm_vcpu_read_guest+0x64/0x140 virt/kvm/kvm_main.c:3284
>  nested_vmx_load_msr+0x133/0x4d0 arch/x86/kvm/vmx/nested.c:1107

Oh man.  vmx_leave_nested() is so broken.  If loading MSRs on nested VM-Exit is
broken (and it obviously is), then storing MSRs on nested VM-Exit is also broken,
i.e. there's at least a second case where nVMX can write to random process memory
on vCPU teardown (shadow vmcs12 being the other one).

It probably makes sense to go straight to open coding punting the vCPU out of L2
in vmx_leave_nested() instead of hack-a-fixing a bunch of flows.  I.e. replace
patch 8 and 9 with a proper fix.

^ permalink raw reply	[flat|nested] 17+ messages in thread

end of thread, other threads:[~2026-10-02 20:39 UTC | newest]

Thread overview: 17+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 20:22 [PATCH v2 00/10] KVM: Fix+harden against bad uaccess using dying VM Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 01/10] KVM: Reject user accesses to guest memory if current->mm != kvm->mm Sean Christopherson
2026-10-01 21:08   ` James Houghton
2026-10-01 21:18     ` Sean Christopherson
2026-10-01 21:28       ` James Houghton
2026-10-01 20:22 ` [PATCH v2 02/10] KVM: PPC: Flush/zap all memslots on kvm_arch_flush_shadow_all() Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 03/10] KVM: x86: Unmap VMAs for KVM-internal memslots when the memslot is freed Sean Christopherson
2026-10-01 21:26   ` James Houghton
2026-10-01 20:22 ` [PATCH v2 04/10] KVM: Disallow setting memslots when the VM is being destroyed Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 05/10] KVM: Destroy memslots immediately after mmu_notifiers are unregistered Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 06/10] KVM: WARN if KVM attempts to do guest-related uaccess with "wrong" process Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 07/10] KVM: WARN and reject guest-based uaccess if VM is dying Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 08/10] KVM: nVMX: Don't flush shadow VMCS12 to guest memory during vCPU teardown Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 09/10] KVM: nVMX: Don't try to load eVMCS12 page when the VM is dying Sean Christopherson
2026-10-01 20:22 ` [PATCH v2 10/10] KVM: Pre-check uaccesses in KVM's APIs to read/write guest memory Sean Christopherson
2026-10-02 20:30 ` [syzbot ci] Re: KVM: Fix+harden against bad uaccess using dying VM syzbot ci
2026-10-02 20:39   ` Sean Christopherson

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®