* [PATCH v2 1/2] KVM: x86/mmu: Report a memory fault exit when the fault handler EFAULTs
2026-09-20 7:24 [PATCH v2 0/2] KVM: x86/mmu: Fill memory_fault for unresolvable guest page faults mike.malyshev
@ 2026-09-20 7:24 ` mike.malyshev
2026-09-20 7:24 ` [PATCH v2 2/2] KVM: x86/mmu: Convert can't-happen fault path EFAULTs to KVM_BUG_ON() + -EIO mike.malyshev
1 sibling, 0 replies; 3+ messages in thread
From: mike.malyshev @ 2026-09-20 7:24 UTC (permalink / raw)
To: seanjc, pbonzini, kvm
Cc: amoorthy, tglx, mingo, bp, dave.hansen, hpa, x86, chao.p.peng,
xiaoyao.li, yu.c.zhang, linux-kernel, Mikhail Malyshev
From: Anish Moorthy <amoorthy@google.com>
KVM_CAP_MEMORY_FAULT_INFO documents that KVM_RUN will fill
kvm_run.memory_fault when KVM cannot resolve a guest page fault VM-Exit,
"e.g. if there is a valid memslot but no backing VMA for the
corresponding host virtual address". kvm_handle_error_pfn() does not
honor that guarantee. It returns a bare -EFAULT, leaving userspace with
no indication of which guest physical address faulted, or even that the
exit was a memory fault at all. Userspace cannot distinguish a
transient, resolvable condition from a fatal one, so in practice the VMM
terminates the guest.
Fill kvm_run.memory_fault before returning -EFAULT.
A concrete user is an Intel integrated GPU assigned to a guest via
vfio-pci. The guest driver clears PCI_COMMAND.MEM on one vCPU while
another vCPU is mid-MMIO to a BAR of the same device. Clearing
PCI_COMMAND.MEM makes vfio-pci zap the BAR's mmap, so the second vCPU's
fault finds a valid memslot whose VMA can no longer supply a PFN, and
KVM_RUN fails with a bare -EFAULT. The VM dies, even though the guest
did nothing architecturally invalid and the condition clears as soon as
the driver re-enables memory decoding.
Reproduce by pairing a vCPU that spins on accesses to the assigned
device's BAR0 with a vCPU that toggles PCI_COMMAND.MEM; the race is hit
within minutes. The same crash has been observed in the field on
production edge hardware.
Reporting the fault does not by itself define the access semantics.
Userspace still has to decide what a read or write to a BAR with memory
decoding disabled returns. But it is the information userspace needs in
order to make that decision instead of killing the guest.
Suggested-by: Sean Christopherson <seanjc@google.com>
Fixes: 16f95f3b95ca ("KVM: Add KVM_EXIT_MEMORY_FAULT exit to report faults to userspace")
Link: https://lore.kernel.org/all/20240809205158.1340255-1-amoorthy@google.com/
Link: https://lore.kernel.org/all/Zr-8M9rYplgN6IS3@google.com/
Signed-off-by: Anish Moorthy <amoorthy@google.com>
Co-developed-by: Mikhail Malyshev <mike.malyshev@gmail.com>
Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
---
arch/x86/kvm/mmu/mmu.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 9788ff1803740..244575f576071 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -3616,6 +3616,7 @@ static int kvm_handle_error_pfn(struct kvm_vcpu *vcpu, struct kvm_page_fault *fa
return RET_PF_RETRY;
}
+ kvm_mmu_prepare_memory_fault_exit(vcpu, fault);
return -EFAULT;
}
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread* [PATCH v2 2/2] KVM: x86/mmu: Convert can't-happen fault path EFAULTs to KVM_BUG_ON() + -EIO
2026-09-20 7:24 [PATCH v2 0/2] KVM: x86/mmu: Fill memory_fault for unresolvable guest page faults mike.malyshev
2026-09-20 7:24 ` [PATCH v2 1/2] KVM: x86/mmu: Report a memory fault exit when the fault handler EFAULTs mike.malyshev
@ 2026-09-20 7:24 ` mike.malyshev
1 sibling, 0 replies; 3+ messages in thread
From: mike.malyshev @ 2026-09-20 7:24 UTC (permalink / raw)
To: seanjc, pbonzini, kvm
Cc: amoorthy, tglx, mingo, bp, dave.hansen, hpa, x86, chao.p.peng,
xiaoyao.li, yu.c.zhang, linux-kernel, Mikhail Malyshev
From: Mikhail Malyshev <mike.malyshev@gmail.com>
Four -EFAULT returns in x86's page fault path are guarded by
WARN_ON_ONCE() because they describe KVM bugs, not conditions the guest
or userspace can reach:
- direct_map() and FNAME(fetch)() completing the walk at a level other
than the goal level, i.e. KVM's shadow page table walk disagreeing
with the mapping level KVM itself just computed.
- a CR2 with bits 63:32 set on 32-bit KVM, which the hardware cannot
produce.
- PFERR_PRIVATE_ACCESS set on a reserved-bit fault. KVM sets that
synthetic bit itself, and only when PFERR_RSVD_MASK is clear.
Returning -EFAULT for a KVM bug conflicts with KVM_CAP_MEMORY_FAULT_INFO,
which promises that an -EFAULT out of KVM_RUN on a guest page fault
VM-Exit is accompanied by kvm_run.memory_fault. Describing a broken
KVM invariant as a memory fault would be actively misleading, as there
is no guest access for userspace to resolve and retry.
Convert the four sites to KVM_BUG_ON() + -EIO, the established way to
report that KVM is hosed, so that -EFAULT out of the fault path always
means "guest memory access KVM could not resolve".
Suggested-by: Sean Christopherson <seanjc@google.com>
Link: https://lore.kernel.org/all/Zr-8M9rYplgN6IS3@google.com/
Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
---
arch/x86/kvm/mmu/mmu.c | 12 ++++++------
arch/x86/kvm/mmu/paging_tmpl.h | 4 ++--
2 files changed, 8 insertions(+), 8 deletions(-)
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 244575f576071..6529ff9e98358 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -3577,8 +3577,8 @@ static int direct_map(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault)
fault->req_level >= it.level);
}
- if (WARN_ON_ONCE(it.level != fault->goal_level))
- return -EFAULT;
+ if (KVM_BUG_ON(it.level != fault->goal_level, vcpu->kvm))
+ return -EIO;
ret = mmu_set_spte(vcpu, fault->slot, it.sptep, access,
base_gfn, fault->pfn, fault);
@@ -4942,8 +4942,8 @@ int kvm_handle_page_fault(struct kvm_vcpu *vcpu, u64 error_code,
#ifndef CONFIG_X86_64
/* A 64-bit CR2 should be impossible on 32-bit KVM. */
- if (WARN_ON_ONCE(fault_address >> 32))
- return -EFAULT;
+ if (KVM_BUG_ON(fault_address >> 32, vcpu->kvm))
+ return -EIO;
#endif
/*
* Legacy #PF exception only have a 32-bit error code. Simply drop the
@@ -6658,8 +6658,8 @@ int noinline kvm_mmu_page_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, u64 err
r = RET_PF_INVALID;
if (unlikely(error_code & PFERR_RSVD_MASK)) {
- if (WARN_ON_ONCE(error_code & PFERR_PRIVATE_ACCESS))
- return -EFAULT;
+ if (KVM_BUG_ON(error_code & PFERR_PRIVATE_ACCESS, vcpu->kvm))
+ return -EIO;
r = handle_mmio_page_fault(vcpu, cr2_or_gpa, direct);
if (r == RET_PF_EMULATE)
diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h
index c8ec47b09264b..8e350095508c5 100644
--- a/arch/x86/kvm/mmu/paging_tmpl.h
+++ b/arch/x86/kvm/mmu/paging_tmpl.h
@@ -778,8 +778,8 @@ static int FNAME(fetch)(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault,
fault->req_level >= it.level);
}
- if (WARN_ON_ONCE(it.level != fault->goal_level))
- return -EFAULT;
+ if (KVM_BUG_ON(it.level != fault->goal_level, vcpu->kvm))
+ return -EIO;
ret = mmu_set_spte(vcpu, fault->slot, it.sptep, gw->pte_access,
base_gfn, fault->pfn, fault);
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread