* [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled
@ 2026-07-23 9:44 Paolo Bonzini
2026-07-24 22:36 ` Sean Christopherson
` (2 more replies)
0 siblings, 3 replies; 17+ messages in thread
From: Paolo Bonzini @ 2026-07-23 9:44 UTC (permalink / raw)
To: linux-kernel, kvm; +Cc: Vitaly Kuznetsov, Alexander Lougovski
Red Hat is seeing multiple reports of Windows memory corruptions
(and consequent BSODs) with hv-tlbflush=on, on AMD processors only.
The crashes, while extremely rare, happen even with a stock configuration,
but with Driver Verifier enabled they can be detected after approximately
200 VM hours. In particular, Alexander Lougovski measured the following:
- on AMD Turin, 15 crashes in 3300 VM hours
- on AMD Milan, 2 crashes in 500 VM hours (there are fewer hours
here due to the host being smaller)
- on Intel Sapphire Rapids, 0 crashes in 8000 VM hours
- on AMD Turin with full TLB flush (not exactly this patch but
similar), no crashes in ~2 weeks of run time which should also
be ~7000 VM hours
For Turin, the microcode version was 0x0b002162, which (assuming
this is the same issue) should not be affected by the problem listed in
https://knowledge.broadcom.com/external/article/419026/bsod-on-virtual-machines-running-on-amd.html;
on the other hand that problem should not apply to earlier processors.
AMD has not provided any information or analysis yet, and when we asked
we didn't know yet that it reproduced on Milan as well.
As to the workload, Alexander threw more or less everything at the same
time at the VM:
- a full Windows Defender scan every 30 minutes
- a disk I/O job
- a loop doing repeated mmap of system files (mostly to hope that
it triggers some consistency check in the Windows memory manager)
- SQL Express 2022 + StressDB (1.6M rows), with the host doing queries
(75% write/25% read) via sqlcmd
Driver Verifier is able to detect BSODs more or less at the same time as
the pages are freed. They mostly happen in the Windows Defender filter
driver, but occasionally also in the networking stack (e.g., afd.sys)
or elsewhere in the filesystem stack (e.g., fltmgr.sys).
The flush is issued from kvm_hv_vcpu_flush_tlb(), which receives the
cross-CPU requests from the Hyper-V TLB flush hypercalls via a kfifo
and is invoked by the KVM_REQ_HV_TLB_FLUSH request. The mechanism is
the same for both Intel and AMD, and the handler for both vendors is
a simple INVVPID(ADDR)/INVLPGA instruction.
Because the request is handled on the destination CPU, there is a question
of what happens if the VM is migrated across physical CPUs. In that case,
the INVLPGA instruction would use a stale svm->vmcb->control.asid; but
if anything that might do an *unnecessary* flush (on an asid that's being
used for another VM) and then pre_svm_run() would force a full TLB rebuild.
So, for lack of better ideas, this patch forces a full ASID bump in
svm_flush_tlb_gva(). To avoid paying the price on Intel and also to
avoid unnecessary loops on AMD, the flush_tlb_gva op now returns whether
it did a full flush or not; kvm_hv_vcpu_flush_tlb() takes note and exits
its loops immediately. While there is an obvious performance impact,
about half of the benefit from Hyper-V tlbflush is preserved (10% vs. 20%
on the SQL Server workload).
kvm_mmu_invalidate_addr() is the only other caller of the flush_tlb_gva op.
The change would have a performance impact on every intercepted INVLPG and,
for nested SVM, on every L1 INVLPGA. For INVLPGA specifically, this covers
the same suspected issue but for nested hypervisors, so it is correct to
apply the workaround; for INVLPG on shadow paging, instead, the impact
would be stronger and, due to lack of data, for now the use of INVLPGA is
left in place in svm_flush_tlb_gva().
Analyzed-by: Vitaly Kuznetsov <vkuznets@redhat.com>
Analyzed-by: Alexander Lougovski <alougovsk@redhat.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
arch/x86/include/asm/kvm_host.h | 2 +-
arch/x86/kvm/hyperv.c | 7 ++++---
arch/x86/kvm/mmu/mmu.c | 2 +-
arch/x86/kvm/svm/svm.c | 27 ++++++++++++++++++++-------
arch/x86/kvm/vmx/main.c | 4 ++--
arch/x86/kvm/vmx/vmx.c | 2 +-
arch/x86/kvm/vmx/x86_ops.h | 2 +-
7 files changed, 30 insertions(+), 16 deletions(-)
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index b517257a6315..eca04d4b974e 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -1751,7 +1751,7 @@ struct kvm_x86_ops {
* Can potentially get non-canonical addresses through INVLPGs, which
* the implementation may choose to ignore if appropriate.
*/
- void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr);
+ void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr, bool *full);
/*
* Flush any TLB entries created by the guest. Like tlb_flush_gva(),
diff --git a/arch/x86/kvm/hyperv.c b/arch/x86/kvm/hyperv.c
index 1ee0d23f8949..5ec1bf28a195 100644
--- a/arch/x86/kvm/hyperv.c
+++ b/arch/x86/kvm/hyperv.c
@@ -1974,6 +1974,7 @@ int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu)
u64 entries[KVM_HV_TLB_FLUSH_FIFO_SIZE];
int i, j, count;
gva_t gva;
+ bool full = false;
if (!tdp_enabled || !hv_vcpu)
return -EINVAL;
@@ -1982,7 +1983,7 @@ int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu)
count = kfifo_out(&tlb_flush_fifo->entries, entries, KVM_HV_TLB_FLUSH_FIFO_SIZE);
- for (i = 0; i < count; i++) {
+ for (i = 0; i < count && !full; i++) {
if (entries[i] == KVM_HV_TLB_FLUSHALL_ENTRY)
goto out_flush_all;
@@ -1991,11 +1992,11 @@ int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu)
* pages to flush.
*/
gva = entries[i] & PAGE_MASK;
- for (j = 0; j < (entries[i] & ~PAGE_MASK) + 1; j++) {
+ for (j = 0; j < (entries[i] & ~PAGE_MASK) + 1 && !full; j++) {
if (is_noncanonical_invlpg_address(gva + j * PAGE_SIZE, vcpu))
continue;
- kvm_x86_call(flush_tlb_gva)(vcpu, gva + j * PAGE_SIZE);
+ kvm_x86_call(flush_tlb_gva)(vcpu, gva + j * PAGE_SIZE, &full);
}
++vcpu->stat.tlb_flush;
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 6c13da942bfc..bea5499fc9f1 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -6672,7 +6672,7 @@ void kvm_mmu_invalidate_addr(struct kvm_vcpu *vcpu, struct kvm_pagewalk *w,
if (is_noncanonical_invlpg_address(addr, vcpu))
return;
- kvm_x86_call(flush_tlb_gva)(vcpu, addr);
+ kvm_x86_call(flush_tlb_gva)(vcpu, addr, NULL);
if (tdp_enabled)
return;
diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index ef69a51ab27f..0bd63970305b 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -4222,13 +4222,6 @@ static void svm_flush_tlb_all(struct kvm_vcpu *vcpu)
svm_flush_tlb_asid(vcpu);
}
-static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva)
-{
- struct vcpu_svm *svm = to_svm(vcpu);
-
- invlpga(gva, svm->vmcb->control.asid);
-}
-
static void svm_flush_tlb_guest(struct kvm_vcpu *vcpu)
{
kvm_register_mark_dirty(vcpu, VCPU_REG_ERAPS);
@@ -4236,6 +4229,26 @@ static void svm_flush_tlb_guest(struct kvm_vcpu *vcpu)
svm_flush_tlb_asid(vcpu);
}
+static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva, bool *full)
+{
+ struct vcpu_svm *svm = to_svm(vcpu);
+
+ /*
+ * INVLPGA has had errata on Genoa and Turin, and even on older
+ * generations there were reports of Windows BSODs if INVLPGA
+ * was used for Hyper-V tlbflush. Use it only for shadow paging
+ * where it seems to be okay.
+ */
+ if (!npt_enabled) {
+ invlpga(gva, svm->vmcb->control.asid);
+ return;
+ }
+
+ svm_flush_tlb_guest(vcpu);
+ if (full)
+ *full = true;
+}
+
static inline void sync_cr8_to_lapic(struct kvm_vcpu *vcpu)
{
struct vcpu_svm *svm = to_svm(vcpu);
diff --git a/arch/x86/kvm/vmx/main.c b/arch/x86/kvm/vmx/main.c
index 83d9921277ea..f204a0fc0a57 100644
--- a/arch/x86/kvm/vmx/main.c
+++ b/arch/x86/kvm/vmx/main.c
@@ -535,12 +535,12 @@ static void vt_flush_tlb_current(struct kvm_vcpu *vcpu)
vmx_flush_tlb_current(vcpu);
}
-static void vt_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr)
+static void vt_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr, bool *full)
{
if (is_td_vcpu(vcpu))
return;
- vmx_flush_tlb_gva(vcpu, addr);
+ vmx_flush_tlb_gva(vcpu, addr, full);
}
static void vt_flush_tlb_guest(struct kvm_vcpu *vcpu)
diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
index 3681d565f177..e27084d30a5e 100644
--- a/arch/x86/kvm/vmx/vmx.c
+++ b/arch/x86/kvm/vmx/vmx.c
@@ -3374,7 +3374,7 @@ void vmx_flush_tlb_current(struct kvm_vcpu *vcpu)
vpid_sync_context(vmx_get_current_vpid(vcpu));
}
-void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr)
+void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr, bool *full)
{
/*
* vpid_sync_vcpu_addr() is a nop if vpid==0, see the comment in
diff --git a/arch/x86/kvm/vmx/x86_ops.h b/arch/x86/kvm/vmx/x86_ops.h
index 409858074246..17595d52985c 100644
--- a/arch/x86/kvm/vmx/x86_ops.h
+++ b/arch/x86/kvm/vmx/x86_ops.h
@@ -82,7 +82,7 @@ void vmx_set_rflags(struct kvm_vcpu *vcpu, unsigned long rflags);
bool vmx_get_if_flag(struct kvm_vcpu *vcpu);
void vmx_flush_tlb_all(struct kvm_vcpu *vcpu);
void vmx_flush_tlb_current(struct kvm_vcpu *vcpu);
-void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr);
+void vmx_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t addr, bool *full);
void vmx_flush_tlb_guest(struct kvm_vcpu *vcpu);
void vmx_set_interrupt_shadow(struct kvm_vcpu *vcpu, int mask);
u32 vmx_get_interrupt_shadow(struct kvm_vcpu *vcpu);
--
2.55.0
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-23 9:44 [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled Paolo Bonzini @ 2026-07-24 22:36 ` Sean Christopherson 2026-07-25 13:37 ` Paolo Bonzini 2026-07-24 23:42 ` Yosry Ahmed 2026-07-27 19:23 ` Tycho Andersen 2 siblings, 1 reply; 17+ messages in thread From: Sean Christopherson @ 2026-07-24 22:36 UTC (permalink / raw) To: Paolo Bonzini Cc: linux-kernel, kvm, Vitaly Kuznetsov, Alexander Lougovski, Yosry Ahmed +Yosry, who has been digging deep on SVM TLB crud. On Thu, Jul 23, 2026, Paolo Bonzini wrote: > The flush is issued from kvm_hv_vcpu_flush_tlb(), which receives the > cross-CPU requests from the Hyper-V TLB flush hypercalls via a kfifo > and is invoked by the KVM_REQ_HV_TLB_FLUSH request. The mechanism is > the same for both Intel and AMD, and the handler for both vendors is > a simple INVVPID(ADDR)/INVLPGA instruction. > > Because the request is handled on the destination CPU, there is a question > of what happens if the VM is migrated across physical CPUs. In that case, > the INVLPGA instruction would use a stale svm->vmcb->control.asid; but > if anything that might do an *unnecessary* flush (on an asid that's being > used for another VM) and then pre_svm_run() would force a full TLB rebuild. > > So, for lack of better ideas, this patch forces a full ASID bump in > svm_flush_tlb_gva(). To avoid paying the price on Intel and also to > avoid unnecessary loops on AMD, the flush_tlb_gva op now returns whether > it did a full flush or not; kvm_hv_vcpu_flush_tlb() takes note and exits > its loops immediately. While there is an obvious performance impact, > about half of the benefit from Hyper-V tlbflush is preserved (10% vs. 20% > on the SQL Server workload). > > kvm_mmu_invalidate_addr() is the only other caller of the flush_tlb_gva op. > The change would have a performance impact on every intercepted INVLPG and, > for nested SVM, on every L1 INVLPGA. For INVLPGA specifically, this covers > the same suspected issue but for nested hypervisors, so it is correct to > apply the workaround; for INVLPG on shadow paging, instead, the impact > would be stronger and, due to lack of data, for now the use of INVLPGA is > left in place in svm_flush_tlb_gva(). > > Analyzed-by: Vitaly Kuznetsov <vkuznets@redhat.com> > Analyzed-by: Alexander Lougovski <alougovsk@redhat.com> > Signed-off-by: Paolo Bonzini <pbonzini@redhat.com> > --- > arch/x86/include/asm/kvm_host.h | 2 +- > arch/x86/kvm/hyperv.c | 7 ++++--- > arch/x86/kvm/mmu/mmu.c | 2 +- > arch/x86/kvm/svm/svm.c | 27 ++++++++++++++++++++------- > arch/x86/kvm/vmx/main.c | 4 ++-- > arch/x86/kvm/vmx/vmx.c | 2 +- > arch/x86/kvm/vmx/x86_ops.h | 2 +- > 7 files changed, 30 insertions(+), 16 deletions(-) > > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h > index b517257a6315..eca04d4b974e 100644 > --- a/arch/x86/include/asm/kvm_host.h > +++ b/arch/x86/include/asm/kvm_host.h > @@ -1751,7 +1751,7 @@ struct kvm_x86_ops { > * Can potentially get non-canonical addresses through INVLPGs, which > * the implementation may choose to ignore if appropriate. > */ > - void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr); > + void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr, bool *full); LOL, why on earth are you using an out-param? If we do this at runtime, just return a bool, at least that way we don't have to churn every call-site. But I would much rather handle this by nuking .flush_tlb_gva at setup, e.g. diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h index 3776cf5382a2..4235331b71a0 100644 --- a/arch/x86/include/asm/kvm-x86-ops.h +++ b/arch/x86/include/asm/kvm-x86-ops.h @@ -61,7 +61,7 @@ KVM_X86_OP(flush_tlb_current) KVM_X86_OP_OPTIONAL(flush_remote_tlbs) KVM_X86_OP_OPTIONAL(flush_remote_tlbs_range) #endif -KVM_X86_OP(flush_tlb_gva) +KVM_X86_OP_OPTIONAL(flush_tlb_gva) KVM_X86_OP(flush_tlb_guest) KVM_X86_OP(vcpu_pre_run) KVM_X86_OP(vcpu_run) diff --git a/arch/x86/kvm/hyperv.c b/arch/x86/kvm/hyperv.c index 4438ecac9a89..7b1c6391f878 100644 --- a/arch/x86/kvm/hyperv.c +++ b/arch/x86/kvm/hyperv.c @@ -1978,7 +1978,8 @@ int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu) count = kfifo_out(&tlb_flush_fifo->entries, entries, KVM_HV_TLB_FLUSH_FIFO_SIZE); for (i = 0; i < count; i++) { - if (entries[i] == KVM_HV_TLB_FLUSHALL_ENTRY) + if (entries[i] == KVM_HV_TLB_FLUSHALL_ENTRY || + !kvm_x86_ops.flush_tlb_gva) goto out_flush_all; /* diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index f0144ae8d891..bf16119640dc 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -6555,7 +6555,10 @@ void kvm_mmu_invalidate_addr(struct kvm_vcpu *vcpu, struct kvm_mmu *mmu, if (is_noncanonical_invlpg_address(addr, vcpu)) return; - kvm_x86_call(flush_tlb_gva)(vcpu, addr); + if (kvm_x86_ops.flush_tlb_gva) + kvm_x86_call(flush_tlb_gva)(vcpu, addr); + else + kvm_make_request(KVM_REQ_TLB_FLUSH_GUEST, vcpu); } if (!mmu->sync_spte) diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c index c46a34aeb3df..9e04f016d939 100644 --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ -5688,6 +5688,15 @@ static __init int svm_hardware_setup(void) if (!enable_pmu) pr_info("PMU virtualization is disabled\n"); + /* + * INVLPGA has had errata on Genoa and Turin, and even on older + * generations there were reports of Windows BSODs if INVLPGA + * was used for Hyper-V tlbflush. Use it only for shadow paging + * where it seems to be okay. + */ + if (npt_enabled) + svm_x86_ops.flush_tlb_gva = NULL; + svm_set_cpu_caps(); kvm_caps.inapplicable_quirks &= ~KVM_X86_QUIRK_CD_NW_CLEARED; ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-24 22:36 ` Sean Christopherson @ 2026-07-25 13:37 ` Paolo Bonzini 0 siblings, 0 replies; 17+ messages in thread From: Paolo Bonzini @ 2026-07-25 13:37 UTC (permalink / raw) To: Sean Christopherson Cc: Kernel Mailing List, Linux, kvm, Vitaly Kuznetsov, Alexander Lougovski, Yosry Ahmed Il sab 25 lug 2026, 00:36 Sean Christopherson <seanjc@google.com> ha scritto: > > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h > > index b517257a6315..eca04d4b974e 100644 > > --- a/arch/x86/include/asm/kvm_host.h > > +++ b/arch/x86/include/asm/kvm_host.h > > @@ -1751,7 +1751,7 @@ struct kvm_x86_ops { > > * Can potentially get non-canonical addresses through INVLPGs, which > > * the implementation may choose to ignore if appropriate. > > */ > > - void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr); > > + void (*flush_tlb_gva)(struct kvm_vcpu *vcpu, gva_t addr, bool *full); > > LOL, why on earth are you using an out-param? If we do this at runtime, just > return a bool, at least that way we don't have to churn every call-site. To be precise the churn is only one call site out of two, and in fact it's the one which is mangled even more by your proposal below. :) More seriously, I used an out parameter because returning 0/1 or false/true is impenetrable for the Hyper-V call site, while 0/-EOPNOTSUPP (i.e. do not flush at all on NPT) would force changes in kvm_mmu_invalidate_addr(). So, instead of making things good for one call site at the expense of the other, the optional out param is more self documenting in svm.c and is either good or bearable for the callers: Hyper-V doesn't care either way (it doesn't e.g. need the return value within an "if" or "while"), and kvm_mmu_invalidate_addr() just gets an extra NULL argument. Overall it's a matter of taste, I understand if you're not convinced but I can say it wasn't out of a whim. In your defense maybe I should have put this argument in the commit message somewhere? AI also complained, but I shelved it because "Sean surely has better taste than the AI"... :) > But I would much rather handle this by nuking .flush_tlb_gva at setup, e.g. > > diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c > index f0144ae8d891..bf16119640dc 100644 > --- a/arch/x86/kvm/mmu/mmu.c > +++ b/arch/x86/kvm/mmu/mmu.c > @@ -6555,7 +6555,10 @@ void kvm_mmu_invalidate_addr(struct kvm_vcpu *vcpu, struct kvm_mmu *mmu, > if (is_noncanonical_invlpg_address(addr, vcpu)) > return; > > - kvm_x86_call(flush_tlb_gva)(vcpu, addr); > + if (kvm_x86_ops.flush_tlb_gva) > + kvm_x86_call(flush_tlb_gva)(vcpu, addr); > + else > + kvm_make_request(KVM_REQ_TLB_FLUSH_GUEST, vcpu); > } > > if (!mmu->sync_spte) This is a variation on the -EOPNOTSUPP convention, for which I didn't like having this "else" in the caller that doesn't already have a full-flush fallback. If you insist, I guess I would be fine with adding a kvm_flush_tlb_gva() wrapper, use it in mmu.c, and check for NULL in hyperv.c... but in the end is it really better than the out param version? In fact I am not even sure it is better than returning -EOPNOTSUPP. From the svm.c point of view it certainly is very clean, but for everyone else it's super easy to forget about checking the callback; and you don't even have __must_check to save you. Paolo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-23 9:44 [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled Paolo Bonzini 2026-07-24 22:36 ` Sean Christopherson @ 2026-07-24 23:42 ` Yosry Ahmed 2026-07-24 23:50 ` Yosry Ahmed 2026-07-25 13:32 ` Paolo Bonzini 2026-07-27 19:23 ` Tycho Andersen 2 siblings, 2 replies; 17+ messages in thread From: Yosry Ahmed @ 2026-07-24 23:42 UTC (permalink / raw) To: Paolo Bonzini; +Cc: linux-kernel, kvm, Vitaly Kuznetsov, Alexander Lougovski On Thu, Jul 23, 2026 at 2:44 AM Paolo Bonzini <pbonzini@redhat.com> wrote: > > Red Hat is seeing multiple reports of Windows memory corruptions > (and consequent BSODs) with hv-tlbflush=on, on AMD processors only. > The crashes, while extremely rare, happen even with a stock configuration, > but with Driver Verifier enabled they can be detected after approximately > 200 VM hours. In particular, Alexander Lougovski measured the following: > > - on AMD Turin, 15 crashes in 3300 VM hours > > - on AMD Milan, 2 crashes in 500 VM hours (there are fewer hours > here due to the host being smaller) > > - on Intel Sapphire Rapids, 0 crashes in 8000 VM hours > > - on AMD Turin with full TLB flush (not exactly this patch but > similar), no crashes in ~2 weeks of run time which should also > be ~7000 VM hours > > For Turin, the microcode version was 0x0b002162, which (assuming > this is the same issue) should not be affected by the problem listed in > https://knowledge.broadcom.com/external/article/419026/bsod-on-virtual-machines-running-on-amd.html; > on the other hand that problem should not apply to earlier processors. > AMD has not provided any information or analysis yet, and when we asked > we didn't know yet that it reproduced on Milan as well. > > As to the workload, Alexander threw more or less everything at the same > time at the VM: > > - a full Windows Defender scan every 30 minutes > > - a disk I/O job > > - a loop doing repeated mmap of system files (mostly to hope that > it triggers some consistency check in the Windows memory manager) > > - SQL Express 2022 + StressDB (1.6M rows), with the host doing queries > (75% write/25% read) via sqlcmd > > Driver Verifier is able to detect BSODs more or less at the same time as > the pages are freed. They mostly happen in the Windows Defender filter > driver, but occasionally also in the networking stack (e.g., afd.sys) > or elsewhere in the filesystem stack (e.g., fltmgr.sys). > > The flush is issued from kvm_hv_vcpu_flush_tlb(), which receives the > cross-CPU requests from the Hyper-V TLB flush hypercalls via a kfifo > and is invoked by the KVM_REQ_HV_TLB_FLUSH request. The mechanism is > the same for both Intel and AMD, and the handler for both vendors is > a simple INVVPID(ADDR)/INVLPGA instruction. > > Because the request is handled on the destination CPU, there is a question > of what happens if the VM is migrated across physical CPUs. In that case, > the INVLPGA instruction would use a stale svm->vmcb->control.asid; but > if anything that might do an *unnecessary* flush (on an asid that's being > used for another VM) and then pre_svm_run() would force a full TLB rebuild. > > So, for lack of better ideas, this patch forces a full ASID bump in > svm_flush_tlb_gva(). To avoid paying the price on Intel and also to > avoid unnecessary loops on AMD, the flush_tlb_gva op now returns whether > it did a full flush or not; kvm_hv_vcpu_flush_tlb() takes note and exits > its loops immediately. While there is an obvious performance impact, > about half of the benefit from Hyper-V tlbflush is preserved (10% vs. 20% > on the SQL Server workload). > > kvm_mmu_invalidate_addr() is the only other caller of the flush_tlb_gva op. > The change would have a performance impact on every intercepted INVLPG and, > for nested SVM, on every L1 INVLPGA. For INVLPGA specifically, this covers > the same suspected issue but for nested hypervisors, so it is correct to > apply the workaround; for INVLPG on shadow paging, instead, the impact > would be stronger and, due to lack of data, for now the use of INVLPGA is > left in place in svm_flush_tlb_gva(). > > Analyzed-by: Vitaly Kuznetsov <vkuznets@redhat.com> > Analyzed-by: Alexander Lougovski <alougovsk@redhat.com> > Signed-off-by: Paolo Bonzini <pbonzini@redhat.com> [..] > +static void svm_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva, bool *full) > +{ > + struct vcpu_svm *svm = to_svm(vcpu); > + > + /* > + * INVLPGA has had errata on Genoa and Turin, and even on older > + * generations there were reports of Windows BSODs if INVLPGA > + * was used for Hyper-V tlbflush. Use it only for shadow paging > + * where it seems to be okay. Is this an actual errata documented by AMD, or is this just an empirical observation? I ask because the APM says: --- The input address is always interpreted as a guest virtual address, so INVLPGA is typically meaningful only when used with shadow page tables; it does not provide a means to invalidate a nested translation by guest physical address --- While this is terrible wording, it seems like KVM should *not* be using INVLPGA when TDP is enabled. Looks like kvm_mmu_invalidate_addr() might be doing the right thing, but it seems like kvm_hv_vcpu_flush_tlb() shouldn't be calling flush_tlb_gva() with TDP enabled to begin with, at least on AMD? I don't have enough context about what kvm_hv_vcpu_flush_tlb() is doing to know if flush_tlb_gva() makes sense on Intel. But at least on AMD, looks like it should always just do a full ASID flush (since it falls back to a full flush with TDP disabled anyway)? So maybe something like: int kvm_hv_vcpu_flush_tlb(struct kvm_vcpu *vcpu) { ... if (AMD CPU) goto out_flush_all; ... } I also love Sean's idea, I think it's good to harden against this by nullifying flush_tlb_gva, and maybe add a helper that does the fallback: static void kvm_vcpu_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva) { if (kvm_x86_ops.flush_tlb_gva) kvm_x86_call(flush_tlb_gva)(vcpu, addr); else kvm_make_request(KVM_REQ_TLB_FLUSH_GUEST, vcpu); } Hmm actually we check KVM_REQ_TLB_FLUSH_GUEST before KVM_REQ_HV_TLB_FLUSH, so maybe just call kvm_vcpu_flush_tlb_guest() directly for the fallback: static void kvm_vcpu_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva) { if (kvm_x86_ops.flush_tlb_gva) kvm_x86_call(flush_tlb_gva)(vcpu, addr); else kvm_vcpu_flush_tlb_guest(vcpu); } And if we go this route, I think we can key off the presence of flush_tlb_gva in kvm_hv_vcpu_flush_tlb() instead of checking for an AMD CPU. We can probably break it down into a stable-friendly fix that just jumps to out_flush_all on AMD CPUs, then the hardening on top. ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-24 23:42 ` Yosry Ahmed @ 2026-07-24 23:50 ` Yosry Ahmed 2026-07-25 13:32 ` Paolo Bonzini 1 sibling, 0 replies; 17+ messages in thread From: Yosry Ahmed @ 2026-07-24 23:50 UTC (permalink / raw) To: Paolo Bonzini; +Cc: linux-kernel, kvm, Vitaly Kuznetsov, Alexander Lougovski > I also love Sean's idea, I think it's good to harden against this by > nullifying flush_tlb_gva, and maybe add a helper that does the > fallback: > > static void kvm_vcpu_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva) > { > if (kvm_x86_ops.flush_tlb_gva) > kvm_x86_call(flush_tlb_gva)(vcpu, addr); > else > kvm_make_request(KVM_REQ_TLB_FLUSH_GUEST, vcpu); > } > > Hmm actually we check KVM_REQ_TLB_FLUSH_GUEST before > KVM_REQ_HV_TLB_FLUSH, so maybe just call kvm_vcpu_flush_tlb_guest() > directly for the fallback: > > static void kvm_vcpu_flush_tlb_gva(struct kvm_vcpu *vcpu, gva_t gva) > { > if (kvm_x86_ops.flush_tlb_gva) > kvm_x86_call(flush_tlb_gva)(vcpu, addr); > else > kvm_vcpu_flush_tlb_guest(vcpu); > } > > And if we go this route, I think we can key off the presence of > flush_tlb_gva in kvm_hv_vcpu_flush_tlb() instead of checking for an > AMD CPU. We can probably break it down into a stable-friendly fix that > just jumps to out_flush_all on AMD CPUs, then the hardening on top. Although it could be less crud if we just made svm_flush_tlb_gva() do the fallback. kvm_hv_vcpu_flush_tlb() would need to key-off AMD CPU though instead of the presence of flush_tlb_gva. Pick your poison, I guess. ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-24 23:42 ` Yosry Ahmed 2026-07-24 23:50 ` Yosry Ahmed @ 2026-07-25 13:32 ` Paolo Bonzini 2026-07-25 21:33 ` Yosry Ahmed 1 sibling, 1 reply; 17+ messages in thread From: Paolo Bonzini @ 2026-07-25 13:32 UTC (permalink / raw) To: Yosry Ahmed Cc: Kernel Mailing List, Linux, kvm, Vitaly Kuznetsov, Alexander Lougovski Il sab 25 lug 2026, 01:44 Yosry Ahmed <yosry@kernel.org> ha scritto: > > + /* > > + * INVLPGA has had errata on Genoa and Turin, and even on older > > + * generations there were reports of Windows BSODs if INVLPGA > > + * was used for Hyper-V tlbflush. Use it only for shadow paging > > + * where it seems to be okay. > > Is this an actual errata documented by AMD, or is this just an > empirical observation? There is the VMware knowledge base in the commit message that requires new microcode: > > For Turin, the microcode version was 0x0b002162, which (assuming > > this is the same issue) should not be affected by the problem listed in > > https://knowledge.broadcom.com/external/article/419026/bsod-on-virtual-machines-running-on-amd.html; > > on the other hand that problem should not apply to earlier processors. > > AMD has not provided any information or analysis yet, and when we asked > > we didn't know yet that it reproduced on Milan as well. and I interpreted that as an erratum. But we reproduced it also on Milan and with supposedly fixed microcode. > I ask because the APM says: > --- > The input address is always interpreted as a guest virtual address, so > INVLPGA is typically meaningful only when used with shadow page > tables; it does not provide a means to invalidate a nested translation > by guest physical address > --- > > While this is terrible wording, it seems like KVM should *not* be > using INVLPGA when TDP is enabled While it is certainly an odd case, here the guest has requested to do an invalidation by GVA on its behalf, so INVLPGA should have worked. "Not typically meaningful" is a friendly hint to read the manual twice, but reality seems to be more like "doesn't actually flush the right entries" when NPT is in use. Especially since Intel has INVVPID for the exact same operation and it works just fine. Paolo ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-25 13:32 ` Paolo Bonzini @ 2026-07-25 21:33 ` Yosry Ahmed 2026-09-03 21:16 ` Joris de Vries 0 siblings, 1 reply; 17+ messages in thread From: Yosry Ahmed @ 2026-07-25 21:33 UTC (permalink / raw) To: Paolo Bonzini Cc: Kernel Mailing List, Linux, kvm, Vitaly Kuznetsov, Alexander Lougovski On Sat, Jul 25, 2026 at 6:32 AM Paolo Bonzini <pbonzini@redhat.com> wrote: > > Il sab 25 lug 2026, 01:44 Yosry Ahmed <yosry@kernel.org> ha scritto: > > > + /* > > > + * INVLPGA has had errata on Genoa and Turin, and even on older > > > + * generations there were reports of Windows BSODs if INVLPGA > > > + * was used for Hyper-V tlbflush. Use it only for shadow paging > > > + * where it seems to be okay. > > > > Is this an actual errata documented by AMD, or is this just an > > empirical observation? > > There is the VMware knowledge base in the commit message that requires > new microcode: > > > > For Turin, the microcode version was 0x0b002162, which (assuming > > > this is the same issue) should not be affected by the problem listed in > > > https://knowledge.broadcom.com/external/article/419026/bsod-on-virtual-machines-running-on-amd.html; > > > on the other hand that problem should not apply to earlier processors. > > > AMD has not provided any information or analysis yet, and when we asked > > > we didn't know yet that it reproduced on Milan as well. > > and I interpreted that as an erratum. But we reproduced it also on > Milan and with supposedly fixed microcode. > > > I ask because the APM says: > > --- > > The input address is always interpreted as a guest virtual address, so > > INVLPGA is typically meaningful only when used with shadow page > > tables; it does not provide a means to invalidate a nested translation > > by guest physical address > > --- > > > > While this is terrible wording, it seems like KVM should *not* be > > using INVLPGA when TDP is enabled > > While it is certainly an odd case, here the guest has requested to do > an invalidation by GVA on its behalf, so INVLPGA should have worked. > > "Not typically meaningful" is a friendly hint to read the manual > twice, but reality seems to be more like "doesn't actually flush the > right entries" when NPT is in use. Especially since Intel has INVVPID > for the exact same operation and it works just fine. Yeah I agree that it makes sense that INVLPGA should work if we are just flushing a GVA on behalf of the guest (e.g. guest making a hypercall instead of INVLPG). I think we probably need clarification from AMD about what the intention is. Is INVLPGA expected to be broken when NPT is enabled (in which case the wording in the APM needs fixing), or is INVLPGA expected to work but is actually broken. IIUC the erratum you referred to should already be fixed with the new microcode, so maybe there's another problem with INVLPGA? The third possibility is that INVLPGA works as intended but the Hyper-V code in KVM is broken in some other way. I am not sure if we can ask, or if people at Microsoft can answer, but it would definitely help to know how Hyper-V handles these cases. Does it actually use INVLPGA to flush GVAs on behalf of the guest with NPT enabeld? If yes, there's a good chance it's KVM that's broken here? ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-25 21:33 ` Yosry Ahmed @ 2026-09-03 21:16 ` Joris de Vries 0 siblings, 0 replies; 17+ messages in thread From: Joris de Vries @ 2026-09-03 21:16 UTC (permalink / raw) To: Yosry Ahmed Cc: Paolo Bonzini, Kernel Mailing List, Linux, kvm, Vitaly Kuznetsov, Alexander Lougovski > On 25 Jul 2026, at 23:33, Yosry Ahmed <yosry@kernel.org> wrote: > > On Sat, Jul 25, 2026 at 6:32 AM Paolo Bonzini <pbonzini@redhat.com> wrote: >> >> Il sab 25 lug 2026, 01:44 Yosry Ahmed <yosry@kernel.org> ha scritto: >>>> + /* >>>> + * INVLPGA has had errata on Genoa and Turin, and even on older >>>> + * generations there were reports of Windows BSODs if INVLPGA >>>> + * was used for Hyper-V tlbflush. Use it only for shadow paging >>>> + * where it seems to be okay. >>> >>> Is this an actual errata documented by AMD, or is this just an >>> empirical observation? >> >> There is the VMware knowledge base in the commit message that requires >> new microcode: >> >>>> For Turin, the microcode version was 0x0b002162, which (assuming >>>> this is the same issue) should not be affected by the problem listed in >>>> https://knowledge.broadcom.com/external/article/419026/bsod-on-virtual-machines-running-on-amd.html; >>>> on the other hand that problem should not apply to earlier processors. >>>> AMD has not provided any information or analysis yet, and when we asked >>>> we didn't know yet that it reproduced on Milan as well. >> >> and I interpreted that as an erratum. But we reproduced it also on >> Milan and with supposedly fixed microcode. >> >>> I ask because the APM says: >>> --- >>> The input address is always interpreted as a guest virtual address, so >>> INVLPGA is typically meaningful only when used with shadow page >>> tables; it does not provide a means to invalidate a nested translation >>> by guest physical address >>> --- >>> >>> While this is terrible wording, it seems like KVM should *not* be >>> using INVLPGA when TDP is enabled >> >> While it is certainly an odd case, here the guest has requested to do >> an invalidation by GVA on its behalf, so INVLPGA should have worked. >> >> "Not typically meaningful" is a friendly hint to read the manual >> twice, but reality seems to be more like "doesn't actually flush the >> right entries" when NPT is in use. Especially since Intel has INVVPID >> for the exact same operation and it works just fine. > > Yeah I agree that it makes sense that INVLPGA should work if we are > just flushing a GVA on behalf of the guest (e.g. guest making a > hypercall instead of INVLPG). > > I think we probably need clarification from AMD about what the > intention is. Is INVLPGA expected to be broken when NPT is enabled (in > which case the wording in the APM needs fixing), or is INVLPGA > expected to work but is actually broken. IIUC the erratum you referred > to should already be fixed with the new microcode, so maybe there's > another problem with INVLPGA? > > The third possibility is that INVLPGA works as intended but the > Hyper-V code in KVM is broken in some other way. I am not sure if we > can ask, or if people at Microsoft can answer, but it would definitely > help to know how Hyper-V handles these cases. Does it actually use > INVLPGA to flush GVAs on behalf of the guest with NPT enabeld? If yes, > there's a good chance it's KVM that's broken here? Hello all I have encountered a bug that to my unknowing eye looks similar and/or related to this discussion. A session with Claude to tweak a local LLM configuration on a kvm guest with GPU passthrough on an AMD host turned into a multi-day debugging session to chase seemingly random guest crashes that entirely subsided when setting npt=0: During debugging this, I searched the kvm list for npt & amd and came across this thread. If my message is correctly placed here, please let me know. Please allow me to forward Claude’s message and analysis. I have tried as much as I can to have it ground its missive in factual findings, but do realize my knowledge is not sufficient to fully contain mistakes and would be grateful for any assistance. Joris Summary ------- With AMD nested paging enabled, a KVM guest reliably kernel-panics with a reserved-bit page fault whenever it tears down a 2 MiB-aligned mapping whose lower page-table page is then freed and reused for other data. - THP collapse/split does this fast: a small stressor panics a 1-vCPU guest in 2-3 minutes, deterministically. - Ordinary glibc thread-stack recycling does it slowly: an idle Ubuntu guest panics in ~35 minutes. The NPT in host memory is structurally correct and is never modified by KVM in the minutes before a panic; the guest page tables are correct; the host is unaffected and logs nothing. The stale value the hardware page-walker uses exists in no memory - it is inside the CPU's own page-walk cache. Every invalidation KVM can issue - guest INVLPG/INVPCID (native or intercepted and re-issued), VMCB TLB_CONTROL at its strongest value on every VMRUN, a fresh ASID after each guest invalidation or on every VMRUN - fails to evict it. Only npt=0 avoids it. This is the same class of bug as the kvm-list thread above, reached locally (single vCPU, no hypercalls, native invalidation) rather than via the Hyper-V TLB-flush hypercall path. Environment ----------- Host CPU AMD Ryzen 9 5900X "Vermeer", family 19h model 21h stepping 0, microcode 0x0a201030, Gigabyte BIOS F41c (2026-08-18) Host memory 4 x 16 GiB DDR4-2400 JEDEC (no XMP), no ECC, no EDAC Host kernel Ubuntu 7.0.0-30-generic; also reproduced on clean mainline v7.3-rc1 Hypervisor QEMU/libvirt, q35, <vmcoreinfo/> + pvpanic, no device assignment (GPU passthrough removable, not required) Guest stock Ubuntu 26.04, kernel 6.12 or 7.0 (does not matter), 1 vCPU, 6 GiB, NO hv-* enlightenments Trigger kvm_amd npt=1 (module default). npt=0 -> stable. Symptom ------- A guest process takes a page fault at a virtual address that is a 2 MiB boundary minus 0x48 or 0x50 (a glibc thread-stack TCB slot), the fault code carries the reserved-bit flag, and the kernel panics on the death of init: systemd[1]: segfault at 741bf3ffffb8 ip 00007fd69b0a0675 sp 00007ffc6813ea88 error 44 in libc.so.6[a0675,...] systemd[1]: segfault at 741bf3ffffb0 ip 0000583d040c17a8 sp 00007ffc6813dda0 error 46 in systemd[157a8,...] Kernel panic - not syncing: Attempted to kill init! exitcode=0x0000008b error 44 is 0x2c - user, read, reserved-bit set; writes fault with 0x2e. A reserved-bit #PF means the hardware page-walker traversed a paging structure marked "present" and found reserved bits in it. In a full cascade a dozen unrelated processes fault at the identical address, all attributed to the same logical CPU. The host stays up and silent throughout: no MCE, no EDAC, no IOMMU fault, no KVM warning. The determining factor is a 2 MiB-aligned region whose lower page-table page is freed and reused. One page-table page maps exactly one 2 MiB-aligned region and is referenced by exactly one PMD entry; when the whole region is torn down, that page-table page is freed and the PMD entry cleared. A THP split/collapse churns this fast; glibc's thread-stack cache releasing a thread's 2 MiB-aligned stack and handing the address to a new thread churns it slowly but continuously on any thread-heavy guest. Relationship to the INVLPGA / hv-tlbflush thread ----------------------------------------------- Red Hat is chasing the same class of bug from the Hyper-V direction: Windows guests with hv-tlbflush=on BSOD on AMD only (Turin 15 crashes / 3300 VM-hours; Milan 2 / 500; Intel Sapphire Rapids 0 / 8000; Turin with a full ASID flush 0 / ~7000). There the guest offloads its cross-vCPU TLB shootdowns to KVM via hypercall, KVM issues INVLPGA on the target vCPU (kvm_hv_vcpu_flush_tlb -> svm_flush_tlb_gva), and that fails to flush the nested part. Broadcom KB 419026 attributes a related Turin case to microcode and lists a fix (-> 0x0B002151); Red Hat still reproduces on Milan past that. This report is the same failure reached without any of that: a stock Linux guest, one vCPU, no enlightenments, no hypercalls, issuing its own native INVLPG/INVPCID (and the CR3 reload the kernel uses for a >2 MiB range flush), none of which NPT intercepts. The full-ASID-flush fix works for the hypercall path because the hypercall is a VMEXIT at exactly the invalidation point; the native local path has no exit there, and forcing a full flush - or a fresh ASID - on every VMRUN does not help (modules A, F above). So if it is the same silicon defect, the KVM fix is correct but covers only guests that offload TLB shootdowns via hypercall; native invalidation - every Linux guest, and Windows without hv-tlbflush - is not covered and cannot be, from the hypervisor. Evidence it is the CPU, not KVM ------------------------------ 1. KVM never touches the NPT before a panic. A bpftrace script filtered to one guest's struct kvm, covering every NPT-mutation and TLB-flush path, over a complete 53-minute run that ended in a panic: present-leaf SPTE zaps / repoints (handle_changed_spte) .... 0 kvm_flush_remote_tlbs / _range ............................ 0 kvm_unmap_gfn_range (mmu-notifier) ....................... 0 svm_flush_tlb_asid / _current / _all ..................... 0 The guest's RAM is fully NPT-mapped after boot and nothing perturbs it - not KSM, not page migration, not an mmu-notifier. There is no KVM flush to race with or to get wrong. svm_flush_tlb_gva() is likewise never called on this path (0 hits), so commit 26505e1b5b has nothing to act on here. 2. The NPT and the guest page tables in memory are correct. A drgn walk of the frozen guest's nested page tables straight out of host-physical memory: huge (2 MiB / 1 GiB) NPT leaves ............... 0 PD entries -> 4 KiB page tables .............. 3082 qemu /proc/<pid>/smaps AnonHugePages ......... 0 kB (host THP = never) The NPT is 4 KiB-only, so the stale entry is not a huge NPT mapping - it is an upper-level step of the 2-D walk (a guest PMD fused with its nested translation). For the faulting task's CR3 the NPT resolves cleanly through all four levels, no reserved bits, to a host page that is byte-identical - full-page MD5, not spot checks - to the same page in the guest core dump. The faulting VA's PML4 slot reads 0 in both: the guest's own page tables say the address is unmapped and call for a plain not-present #PF (error 4/6). The CPU raised a reserved-bit one. The value the walker used is in no memory anywhere. Repeated on four separate cores including a 6.12-kernel guest. 3. page_poison rules out the DRAM / use-after-free path. The guest runs page_poison=1 init_on_free=1 slub_debug=FZ. Freed pages are overwritten; a use-after-free would surface as poison, not as a coherent stale translation. The page tables in the cores are self-consistent. The discrepancy is translation state, not bytes. What does not fix it ------------------- All rows still panic with the identical signature unless noted. Guest uptime at panic; npt=1; hugestress2 unless the row predates it. npt=0 (software shadow paging) ..................... clean (only fix) single-VM 25 min clean rejects "same as npt=1" at p<0.01; bare-metal hugestress2 45 min clean commit 26505e1b5b backported to the 7.0 host ....... panic ~30 min mainline v7.3-rc1 host ............................. panic 31 min BIOS F31->F41c, ucode 0x0a20102e->0x0a201030 ....... panic 12 min HWCR[TlbCacheDis]=1 (AMD flush-filter disabled) .... panic ~24 min x2 module A: TLB_CONTROL_FLUSH_ALL_ASID every VMRUN ... panic 14 min module B: guest INVLPG intercepted + hv-mediated flush (~112k exits/s) ......................... panic 33 min module Q: force-intercept guest INVPCID + INVLPG under NPT, fresh ASID after each ............. panic 2 min module F: new_asid() unconditionally in pre_svm_run (fresh ASID every VMRUN) ......... panic 29 min KSM off / 1 vCPU / tdp_mmu=0 / pku,ospke masked ... panic GPU passthrough removed / CPU pinning removed ..... panic 30 min transparent_hugepage=never in the guest .......... panic 45 min (removes the fast THP path; the slow glibc-thread-stack path still panics - the fault landed on a systemd thread-stack address, not the workload's regions) The corrupting collapse -> teardown -> re-access sequence completes inside a single VMRUN, in guest code the hypervisor never traps, so no between-VMRUN action - which is all KVM has - can land in the window. The only untried lever is an nCR3 (NPT root) reload per guest invalidation, which is a full MMU reset per INVPCID and not viable. AMD erratum search ------------------ Checked AMD publication 56683 (Revision Guide for Family 19h Models 00h-0Fh, i.e. Milan - the same Zen 3 core; AMD publishes no client revision guide, so Vermeer / model 21h has none) and the RemembERR errata database. No published erratum matches. The closest siblings, all marked "no fix planned" and affecting both Zen 3 steppings: 1193 Page Remapping Without Invalidation May Cause Missed Detection of Self-Modifying Code. If a PTE with the Accessed bit set has its physical page base changed without first making the translation a permission violation and then invalidating it, the CPU may execute stale instructions. Same failure class - a stale translation surviving a change of a page's backing - but scoped to instruction fetch. This report is the data-side / nested-walk analog. Note the workaround text: plain "clear the PTE + INVLPG" is stated to be insufficient. 1277 IOMMU May Mishandle Fault on Skipped Page Directory Entry Levels. When guest and nested page tables are enabled, a nested walk that skips a PDE level is mishandled. Confirms nested-walk defects exist in this silicon. 1455 PCID-Based INVLPGB May Fail to Flush Global Translations under specific conditions. Confirms "an invalidation instruction does not flush" precedent on Zen 3. Workarounds ----------- kvm_amd npt=0 The only reliable option. KVM's shadow MMU does the guest page walk in software and never routes the hardware 2-D walker through NPT to a guest page-table page. Cost is a higher VM-exit rate on guest page-table edits; near zero for steady GPU / compute workloads. guest transparent_hugepage=never Removes the fast (THP) path only. A guest with no huge-page activity at all still panicked at 45 minutes on an ambient systemd thread-stack address. Rate reduction, not a fix. AMD microcode The actual fix for the silicon. A related Turin case has one (Broadcom KB 419026); no Zen 3 client microcode fix is known. Diagnostic module diffs ----------------------- Built against linux-source-7.0.0 (== 7.0.0-30.30), vermagic-matched, disassembly-verified, loaded as a drop-in kvm-amd.ko. --- 8< --- module A: flush all ASIDs on every VMRUN --- --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ svm_vcpu_enter_exit() amd_clear_divider(); + /* force a full flush of ALL ASIDs on every VMRUN, unconditionally */ + svm->vmcb->control.tlb_ctl = TLB_CONTROL_FLUSH_ALL_ASID; + if (sev_es_guest(vcpu->kvm)) __svm_sev_es_vcpu_run(svm, ...); else __svm_vcpu_run(svm, spec_ctrl_intercepted); --- 8< --- --- 8< --- module B: route guest INVLPG through the hypervisor --- --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ init_vmcb(), if (npt_enabled) control->nested_ctl |= SVM_NESTED_CTL_NP_ENABLE; - svm_clr_intercept(svm, INTERCEPT_INVLPG); + /* keep INVLPG intercepted under NPT */ clr_exception_intercept(svm, PF_VECTOR); @@ svm_flush_tlb_gva() struct vcpu_svm *svm = to_svm(vcpu); + /* no INVLPGA under NPT (unreliable, cf. 26505e1b5b); hypervisor- + * mediated full guest-ASID flush instead */ + if (npt_enabled) { + svm_flush_tlb_asid(vcpu); + return; + } + invlpga(gva, svm->vmcb->control.asid); --- 8< --- --- 8< --- module F: brand-new ASID on every VMRUN --- --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ pre_svm_run() if (sev_guest(vcpu->kvm)) return pre_sev_run(svm, vcpu->cpu); - /* FIXME: handle wraparound of asid_generation */ - if (svm->current_vmcb->asid_generation != sd->asid_generation) - new_asid(svm, sd); + /* assign a brand-new ASID on EVERY VMRUN. Unlike module A (which + * re-flushes the SAME asid) this changes the ASID *tag* so the CPU + * cannot consult a stale NPT-derived walk-cache entry left under + * the old tag. */ + new_asid(svm, sd); return 0; --- 8< --- --- 8< --- module Q: intercept guest INVPCID+INVLPG, fresh ASID each --- --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ svm_recalc_instruction_intercepts() if (kvm_cpu_cap_has(X86_FEATURE_INVPCID)) { - if (!npt_enabled || - !guest_cpu_cap_has(&svm->vcpu, X86_FEATURE_INVPCID)) - svm_set_intercept(svm, INTERCEPT_INVPCID); - else - svm_clr_intercept(svm, INTERCEPT_INVPCID); + /* always intercept INVPCID, even under NPT */ + svm_set_intercept(svm, INTERCEPT_INVPCID); } @@ init_vmcb(), if (npt_enabled) - svm_clr_intercept(svm, INTERCEPT_INVLPG); + /* keep INVLPG intercepted under NPT */ clr_exception_intercept(svm, PF_VECTOR); @@ invlpg_interception() kvm_mmu_invlpg(vcpu, to_svm(vcpu)->vmcb->control.exit_info_1); + to_svm(vcpu)->current_vmcb->asid_generation--; /* -> new_asid() */ return kvm_skip_emulated_instruction(vcpu); @@ invpcid_interception() - return kvm_handle_invpcid(vcpu, type, gva); + { + int _r = kvm_handle_invpcid(vcpu, type, gva); + svm->current_vmcb->asid_generation--; /* -> new_asid() */ + return _r; + } --- 8< --- Reproducer ---------- Build: gcc -O2 -static -pthread -o hugestress2 hugestress2.c Run in the guest as root (needs /proc/self/pagemap and /dev/kmsg). Panic in 2-3 min on a 1-vCPU npt=1 guest; clean under npt=0. --- 8< --- hugestress2.c --- /* hugestress2 - 2 MiB THP collapse/teardown churn with a canary. * On canary corruption: resolve the bad page's guest-physical address * via /proc/self/pagemap, dump context to /dev/kmsg + console, stop the * workers and FREEZE (keep the VM alive and the page mapped for * virsh dump). * build: gcc -O2 -static -pthread -o hugestress2 hugestress2.c */ #define _GNU_SOURCE #include <stdio.h> #include <stdlib.h> #include <string.h> #include <stdint.h> #include <unistd.h> #include <fcntl.h> #include <time.h> #include <sys/mman.h> #include <pthread.h> #include <signal.h> #ifndef MADV_COLLAPSE #define MADV_COLLAPSE 25 #endif #define HP (2UL*1024*1024) #define NREG 48 #define NTHR 3 static int kmsg = -1; static volatile int freeze = 0; static void say(const char *s){ dprintf(1,"%s\n",s); dprintf(2,"%s\n",s); if(kmsg>=0){ char b[300]; int n=snprintf(b,sizeof b,"hugestress2: %s\n",s); if(write(kmsg,b,n)){} } } static uint64_t rnd(uint64_t *s){ *s=*s*6364136223846793005ULL+1442695040888963407ULL; return *s; } static unsigned char canpat(size_t i){ return (unsigned char)((i*2654435761u)>>24); } /* virtual addr -> guest physical addr via /proc/self/pagemap (needs root) */ static uint64_t v2p(void *va){ static int pm = -2; if(pm==-2) pm = open("/proc/self/pagemap", O_RDONLY); if(pm<0) return 0; uint64_t off = ((uint64_t)va/4096)*8, ent=0; if(pread(pm,&ent,8,off)!=8) return 0; if(!(ent & (1ULL<<63))) return 0; /* not present */ return ((ent & ((1ULL<<55)-1))*4096) | ((uint64_t)va & 4095); } static void *worker(void *arg){ uint64_t s=(uint64_t)(long)arg*0x9e3779b97f4a7c15ULL ^ (uint64_t)time(0); unsigned char *reg[NREG]; memset(reg,0,sizeof reg); while(!freeze){ int i=rnd(&s)%NREG; if(!reg[i]){ unsigned char *m=mmap(0,HP,PROT_READ|PROT_WRITE,MAP_PRIVATE|MAP_ANONYMOUS,-1,0); if(m==MAP_FAILED) continue; madvise(m,HP,MADV_HUGEPAGE); for(size_t o=0;o<HP;o+=4096) m[o]=1; madvise(m,HP,MADV_COLLAPSE); m[HP-0x48]=0x5a; reg[i]=m; } else { unsigned char *m=reg[i]; switch(rnd(&s)&7){ case 0: mprotect(m,HP,PROT_READ); mprotect(m,HP,PROT_READ|PROT_WRITE); break; case 1: madvise(m+HP/2,HP/2,MADV_DONTNEED); break; case 2: mprotect(m+HP/2,0x1000,PROT_NONE); mprotect(m+HP/2,0x1000,PROT_READ|PROT_WRITE); break; case 3: { unsigned char *n=mremap(m,HP,HP,MREMAP_MAYMOVE); if(n!=MAP_FAILED){reg[i]=n; n[HP-0x48]=0x5a;} break; } case 4: madvise(m,HP,MADV_DONTNEED); m[0]=1; m[HP-0x48]=0x5a; break; case 5: madvise(m,HP,MADV_COLLAPSE); break; case 6: { volatile unsigned char *p=m+HP-0x48; unsigned char v=*p; *p=v+1; break; } default: munmap(m,HP); reg[i]=0; break; } } } return 0; } int main(void){ setvbuf(stdout,0,_IONBF,0); kmsg=open("/dev/kmsg",O_WRONLY|O_CLOEXEC); say("hugestress2 start"); size_t CN=64*HP; /* 128 MiB canary */ unsigned char *can=mmap(0,CN,PROT_READ|PROT_WRITE,MAP_PRIVATE|MAP_ANONYMOUS,-1,0); if(can==MAP_FAILED){ say("canary mmap failed"); return 1; } madvise(can,CN,MADV_HUGEPAGE); for(size_t i=0;i<CN;i++) can[i]=canpat(i); madvise(can,CN,MADV_COLLAPSE); { char b[160]; snprintf(b,sizeof b,"canary va=%p..%p pa[0]=%#lx pa[mid]=%#lx", (void*)can,(void*)(can+CN), v2p(can), v2p(can+CN/2)); say(b); } pthread_t t[NTHR]; for(int i=0;i<NTHR;i++) pthread_create(&t[i],0,worker,(void*)(long)(i+1)); unsigned long pass=0; time_t t0=time(0); for(;;){ for(size_t i=0;i<CN;i+=53){ unsigned char got=can[i], want=canpat(i); if(got!=want){ freeze=1; /* stop workers, keep state */ char b[256]; void *va=&can[i]; uint64_t pa=v2p(va); snprintf(b,sizeof b,"CANARY CORRUPT off=%zu va=%p pa=%#lx got=%02x want=%02x pass=%lu t=%lds", i,va,pa,got,want,pass,(long)(time(0)-t0)); say(b); /* is it a stale-TLB flicker or stable memory corruption? */ for(int k=0;k<8;k++){ snprintf(b,sizeof b," reread[%d]=%02x pa=%#lx",k,can[i],v2p(va)); say(b); usleep(200000); } /* hexdump 128 bytes around, aligned */ size_t base=i & ~63UL; char line[160]; int p=0; p+=snprintf(line,sizeof line," dump %p:",(void*)&can[base]); for(size_t j=0;j<128 && base+j<CN;j++){ p+=snprintf(line+p,sizeof line-p," %02x",can[base+j]); if((j&15)==15){ say(line); p=snprintf(line,sizeof line," +%zu:",j+1); } } say("FROZEN - virsh dump now. (workers stopped, page still mapped)"); for(;;) pause(); } } if((++pass%5000)==0){ char b[80]; snprintf(b,sizeof b,"canary ok pass=%lu t=%lds",pass,(long)(time(0)-t0)); say(b); } } } --- 8< --- hugestress2 also runs clean for 45 minutes on the bare-metal host, and Debian-minimal userspace runs it clean for hours on both a Debian 6.1 and an Ubuntu 7.0 guest kernel - the guest kernel version is not the variable, the volume of 2 MiB-aligned teardown+reuse is. ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-23 9:44 [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled Paolo Bonzini 2026-07-24 22:36 ` Sean Christopherson 2026-07-24 23:42 ` Yosry Ahmed @ 2026-07-27 19:23 ` Tycho Andersen 2026-07-27 19:58 ` Yosry Ahmed 2026-07-28 17:13 ` Alexander Lougovski 2 siblings, 2 replies; 17+ messages in thread From: Tycho Andersen @ 2026-07-27 19:23 UTC (permalink / raw) To: Paolo Bonzini Cc: linux-kernel, kvm, Vitaly Kuznetsov, Alexander Lougovski, Tom Lendacky, Borislav Petkov Hi all, On Thu, Jul 23, 2026 at 11:44:19AM +0200, Paolo Bonzini wrote: > As to the workload, Alexander threw more or less everything at the same > time at the VM: > > - a full Windows Defender scan every 30 minutes > > - a disk I/O job > > - a loop doing repeated mmap of system files (mostly to hope that > it triggers some consistency check in the Windows memory manager) > > - SQL Express 2022 + StressDB (1.6M rows), with the host doing queries > (75% write/25% read) via sqlcmd Can you share the code for the mmap + disk io + DB content generator? We're looking at this and would like to reproduce. Thanks, Tycho ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-27 19:23 ` Tycho Andersen @ 2026-07-27 19:58 ` Yosry Ahmed 2026-07-29 23:41 ` Paolo Bonzini 2026-07-28 17:13 ` Alexander Lougovski 1 sibling, 1 reply; 17+ messages in thread From: Yosry Ahmed @ 2026-07-27 19:58 UTC (permalink / raw) To: Tycho Andersen, Paolo Bonzini Cc: linux-kernel, kvm, Vitaly Kuznetsov, Alexander Lougovski, Tom Lendacky, Borislav Petkov, Sean Christopherson On Mon, Jul 27, 2026 at 12:23 PM Tycho Andersen <tycho@kernel.org> wrote: > > Hi all, > > On Thu, Jul 23, 2026 at 11:44:19AM +0200, Paolo Bonzini wrote: > > As to the workload, Alexander threw more or less everything at the same > > time at the VM: > > > > - a full Windows Defender scan every 30 minutes > > > > - a disk I/O job > > > > - a loop doing repeated mmap of system files (mostly to hope that > > it triggers some consistency check in the Windows memory manager) > > > > - SQL Express 2022 + StressDB (1.6M rows), with the host doing queries > > (75% write/25% read) via sqlcmd > > Can you share the code for the mmap + disk io + DB content generator? > > We're looking at this and would like to reproduce. Not sure if this is relevant, but I see that hyperv_tlb_flush is flaky with npt=0 on Turin at the tip of kvm-x86/next (commit 567329869b9c7). It's not likely because the problem fixed by this patch happens specifically with npt=1 and not npt=0 AFAICT, but it is a weird coincidence :) Maybe something is wrong in the Hyper-V code that manifests differently with npt=0/1? ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-27 19:58 ` Yosry Ahmed @ 2026-07-29 23:41 ` Paolo Bonzini 2026-07-30 0:24 ` Yosry Ahmed 0 siblings, 1 reply; 17+ messages in thread From: Paolo Bonzini @ 2026-07-29 23:41 UTC (permalink / raw) To: Yosry Ahmed, Tycho Andersen Cc: linux-kernel, kvm, Vitaly Kuznetsov, Tom Lendacky, Borislav Petkov, Sean Christopherson, Alexander Lougovski On 7/27/26 21:58, Yosry Ahmed wrote: > On Mon, Jul 27, 2026 at 12:23 PM Tycho Andersen <tycho@kernel.org> wrote: >> >> Hi all, >> >> On Thu, Jul 23, 2026 at 11:44:19AM +0200, Paolo Bonzini wrote: >>> As to the workload, Alexander threw more or less everything at the same >>> time at the VM: >>> >>> - a full Windows Defender scan every 30 minutes >>> >>> - a disk I/O job >>> >>> - a loop doing repeated mmap of system files (mostly to hope that >>> it triggers some consistency check in the Windows memory manager) >>> >>> - SQL Express 2022 + StressDB (1.6M rows), with the host doing queries >>> (75% write/25% read) via sqlcmd >> >> Can you share the code for the mmap + disk io + DB content generator? >> >> We're looking at this and would like to reproduce. > > Not sure if this is relevant, but I see that hyperv_tlb_flush is flaky > with npt=0 on Turin at the tip of kvm-x86/next (commit 567329869b9c7). Same on Milan, but ept=0 works. > It's not likely because the problem fixed by this patch happens > specifically with npt=1 and not npt=0 AFAICT, but it is a weird > coincidence :) And also because this patch uses, in the end, the same flush mechanism that npt=0 and ept=0 already use. So that bug is probably in npt=0. That said, I'll do a quick experiment of running hyperv_tlb_flush in a loop with this patch. Paolo > Maybe something is wrong in the Hyper-V code that manifests > differently with npt=0/1? ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-29 23:41 ` Paolo Bonzini @ 2026-07-30 0:24 ` Yosry Ahmed 0 siblings, 0 replies; 17+ messages in thread From: Yosry Ahmed @ 2026-07-30 0:24 UTC (permalink / raw) To: Paolo Bonzini Cc: Tycho Andersen, linux-kernel, kvm, Vitaly Kuznetsov, Tom Lendacky, Borislav Petkov, Sean Christopherson, Alexander Lougovski On Wed, Jul 29, 2026 at 4:41 PM Paolo Bonzini <pbonzini@redhat.com> wrote: > > On 7/27/26 21:58, Yosry Ahmed wrote: > > On Mon, Jul 27, 2026 at 12:23 PM Tycho Andersen <tycho@kernel.org> wrote: > >> > >> Hi all, > >> > >> On Thu, Jul 23, 2026 at 11:44:19AM +0200, Paolo Bonzini wrote: > >>> As to the workload, Alexander threw more or less everything at the same > >>> time at the VM: > >>> > >>> - a full Windows Defender scan every 30 minutes > >>> > >>> - a disk I/O job > >>> > >>> - a loop doing repeated mmap of system files (mostly to hope that > >>> it triggers some consistency check in the Windows memory manager) > >>> > >>> - SQL Express 2022 + StressDB (1.6M rows), with the host doing queries > >>> (75% write/25% read) via sqlcmd > >> > >> Can you share the code for the mmap + disk io + DB content generator? > >> > >> We're looking at this and would like to reproduce. > > > > Not sure if this is relevant, but I see that hyperv_tlb_flush is flaky > > with npt=0 on Turin at the tip of kvm-x86/next (commit 567329869b9c7). > > Same on Milan, but ept=0 works. > > > It's not likely because the problem fixed by this patch happens > > specifically with npt=1 and not npt=0 AFAICT, but it is a weird > > coincidence :) > > And also because this patch uses, in the end, the same flush mechanism > that npt=0 and ept=0 already use. So that bug is probably in npt=0. > > That said, I'll do a quick experiment of running hyperv_tlb_flush in a > loop with this patch. I don't think this patch changes the behavior with npt=0 tho, right? I would be surprised if the flakiness went away. I mainly mentioned this in case the (preexisting) cause of the flakiness is somehow related to the problem this patch is trying to fix. ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-27 19:23 ` Tycho Andersen 2026-07-27 19:58 ` Yosry Ahmed @ 2026-07-28 17:13 ` Alexander Lougovski 2026-07-28 21:34 ` Tycho Andersen 1 sibling, 1 reply; 17+ messages in thread From: Alexander Lougovski @ 2026-07-28 17:13 UTC (permalink / raw) To: tycho Cc: Alexander Lougovski, bp, kvm, linux-kernel, pbonzini, thomas.lendacky, vkuznets On Mon, Jul 27, 2026 at 01:23:03PM -0600, Tycho Andersen wrote: > Can you share the code for the mmap + disk io + DB content generator? > > We're looking at this and would like to reproduce. Hi Tycho, Sure, here's everything below. The setup is a Windows Server 2022 guest on QEMU/KVM with Driver Verifier enabled (/standard /all, which includes Special Pool). All three workloads run concurrently inside the guest. 1) Memory-mapped file workload (PowerShell, runs as a scheduled task) Two concurrent background jobs, each in an infinite loop: pick a random system DLL, read it into a byte array, create a memory-mapped file from it, touch pages through the view accessor, then dispose. Every 50 iterations, GC.Collect() forces the CLR to finalize all the disposed mappings at once, which triggers a burst of PTE teardowns in the Windows memory manager - in theory that should put some pressure on the TLB. --- 8< --- workload-1582.ps1 --- $ErrorActionPreference = 'Stop' $logFile = "C:\drivers\workload-1582-errors.log" $files = @( "$env:SystemRoot\System32\ntdll.dll", "$env:SystemRoot\System32\kernel32.dll", "$env:SystemRoot\System32\user32.dll", "$env:SystemRoot\System32\advapi32.dll", "$env:SystemRoot\System32\ole32.dll", "$env:SystemRoot\System32\shell32.dll", "$env:SystemRoot\System32\comctl32.dll", "$env:SystemRoot\System32\msvcrt.dll" ) $block = { param($files, $logFile) $i = 0 while ($true) { $f = $files[(Get-Random -Maximum $files.Count)] try { $bytes = [System.IO.File]::ReadAllBytes($f) $ms = [System.IO.MemoryStream]::new($bytes) $mmf = [System.IO.MemoryMappedFiles.MemoryMappedFile]::CreateFromFile( $f, [System.IO.FileMode]::Open, $null, 0, [System.IO.MemoryMappedFiles.MemoryMappedFileAccess]::Read) $view = $mmf.CreateViewAccessor(0, 0, [System.IO.MemoryMappedFiles.MemoryMappedFileAccess]::Read) for ($p = 0; $p -lt [Math]::Min($view.Capacity, 65536); $p += 4096) { $null = $view.ReadByte($p) } $view.Dispose() $mmf.Dispose() $ms.Dispose() } catch { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') iter=$i file=$f error=$($_.Exception.Message)" | Out-File -Append $logFile } $i++ if ($i % 50 -eq 0) { try { [GC]::Collect() [GC]::WaitForPendingFinalizers() } catch { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') iter=$i GC error=$($_.Exception.Message)" | Out-File -Append $logFile } } } } "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') Workload1582 starting" | Out-File -Append $logFile 1..2 | ForEach-Object { Start-Job -ScriptBlock $block -ArgumentList (,$files), $logFile } while ($true) { $jobs = Get-Job $failed = $jobs | Where-Object { $_.State -eq 'Failed' } if ($failed) { foreach ($j in $failed) { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') Job $($j.Id) FAILED: $($j.ChildJobs[0].JobStateInfo.Reason)" | Out-File -Append $logFile Remove-Job $j -Force Start-Job -ScriptBlock $block -ArgumentList (,$files), $logFile } } Start-Sleep -Seconds 60 } --- 8< --- Register as a startup task: schtasks /create /tn Workload1582 /sc onstart /ru SYSTEM /f ^ /tr "powershell.exe -ExecutionPolicy Bypass -File C:\workload-1582.ps1" 2) Disk I/O (DiskSpd, runs as a scheduled task) Single-threaded, 32 outstanding I/Os, 64K block size, 50/50 read/write, 128MB file, loop forever. --- 8< --- diskspd-loop.cmd --- :loop C:\drivers\diskspd.exe -d604800 -t1 -o32 -b64K -w50 -Sh -c128M C:\diskspd-io.dat goto loop --- 8< --- DiskSpd binary: https://github.com/microsoft/diskspd/releases Register: schtasks /create /tn DiskSpdIO /sc onstart /ru SYSTEM /f ^ /tr "C:\drivers\diskspd-loop.cmd" 3) SQL stress (host-side script, drives queries into the guest) SQL Server Express 2022 runs inside the guest. The database (StressDB, SIMPLE recovery model) has a single table with ~1.6M rows: RandomData ( ID INT IDENTITY PRIMARY KEY, Col1 UNIQUEIDENTIFIER, Col2 NVARCHAR(100), Col3 NVARCHAR(100), Col4 NVARCHAR(100), Col5 DATETIME2, Col6 FLOAT, Col7 BIGINT, Col8 VARBINARY(200) ) In our setup the database is baked into the golden VM image, so we don't have a standalone creation script. To reproduce, seed random data and double via INSERT...SELECT until you reach ~1.6M rows. The exact count doesn't matter much — it just needs to be larger than the SQL Express buffer pool (1.4 GB) so that queries cause constant cache misses and disk I/O. The host runs 3 sqlcmd streams per VM, 75% writes / 25% reads: --- 8< --- sql-stress-75w.sh --- #!/usr/bin/env bash SQLCMD=/opt/mssql-tools18/bin/sqlcmd query_stream() { local TARGET=$1 while true; do case $((RANDOM % 4)) in 0) Q="SET NOCOUNT ON; DECLARE @s INT = ABS(CHECKSUM(NEWID())) % 2800000 + 1; UPDATE StressDB.dbo.RandomData SET Col2 = REPLICATE(N'X', 90 + ABS(CHECKSUM(NEWID())) % 10), Col3 = REPLICATE(N'Y', 90 + ABS(CHECKSUM(NEWID())) % 10), Col5 = SYSDATETIME(), Col7 = ABS(CHECKSUM(NEWID())) WHERE ID BETWEEN @s AND @s + 50000" ;; 1) Q="SET NOCOUNT ON; DECLARE @s INT = ABS(CHECKSUM(NEWID())) % 2800000 + 1; UPDATE StressDB.dbo.RandomData SET Col4 = REPLICATE(N'Z', 90 + ABS(CHECKSUM(NEWID())) % 10), Col6 = RAND(CHECKSUM(NEWID())) * 1000, Col8 = CAST(REPLICATE(0x42, 100 + ABS(CHECKSUM(NEWID())) % 100) AS VARBINARY(200)) WHERE ID BETWEEN @s AND @s + 50000" ;; 2) Q="SET NOCOUNT ON; DECLARE @s INT = ABS(CHECKSUM(NEWID())) % 2800000 + 1; UPDATE StressDB.dbo.RandomData SET Col2 = REPLICATE(N'W', 90 + ABS(CHECKSUM(NEWID())) % 10), Col5 = SYSDATETIME(), Col7 = ABS(CHECKSUM(NEWID())), Col8 = CAST(REPLICATE(0x43, 150) AS VARBINARY(200)) WHERE ID BETWEEN @s AND @s + 50000; CHECKPOINT" ;; 3) Q="SET NOCOUNT ON; DBCC DROPCLEANBUFFERS WITH NO_INFOMSGS; CHECKPOINT; SELECT TOP 5000 ID, Col2, Col3, Col4, Col8 FROM StressDB.dbo.RandomData WITH (NOLOCK) WHERE Col7 > ABS(CHECKSUM(NEWID())) % 2000000000 ORDER BY Col5 DESC" ;; esac $SQLCMD -S "$TARGET" -U sa -P 'TestPass123!' -C -Q "$Q" -t 120 > /dev/null 2>&1 done } echo "Starting 3 streams per VM (42 VMs = 126 streams)..." echo "75% writes (cases 0,1,2): UPDATE 50K rows with wide columns + CHECKPOINT on case 2" echo "25% reads (case 3): DROPCLEANBUFFERS + CHECKPOINT + table scan" for N in $(seq 1 42); do IP="192.168.100.$((10 + N))" query_stream "$IP,1433" & query_stream "$IP,1433" & query_stream "$IP,1433" & done wait --- 8< --- On top of all that, Windows Defender runs a full scan of all drives every 30 minutes via a scheduled task. Al ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-28 17:13 ` Alexander Lougovski @ 2026-07-28 21:34 ` Tycho Andersen 2026-07-30 11:02 ` Alexander Lougovski 0 siblings, 1 reply; 17+ messages in thread From: Tycho Andersen @ 2026-07-28 21:34 UTC (permalink / raw) To: Alexander Lougovski Cc: bp, kvm, linux-kernel, pbonzini, thomas.lendacky, vkuznets Hi Alexander, On Tue, Jul 28, 2026 at 07:13:56PM +0200, Alexander Lougovski wrote: > On Mon, Jul 27, 2026 at 01:23:03PM -0600, Tycho Andersen wrote: > > Can you share the code for the mmap + disk io + DB content generator? > > > > We're looking at this and would like to reproduce. > > Hi Tycho, > > Sure, here's everything below. Thanks for this, I think I've got everything set up the way you've described on top of 7.2-rc5. One question, > query_stream() { > local TARGET=$1 > while true; do > case $((RANDOM % 4)) in > 0) Q="SET NOCOUNT ON; DECLARE @s INT = ABS(CHECKSUM(NEWID())) % 2800000 + 1; UPDATE StressDB.dbo.RandomData SET Col2 = REPLICATE(N'X', 90 + ABS(CHECKSUM(NEWID())) % 10), Col3 = REPLICATE(N'Y', 90 + ABS(CHECKSUM(NEWID())) % 10), Col5 = SYSDATETIME(), Col7 = ABS(CHECKSUM(NEWID())) WHERE ID BETWEEN @s AND @s + 50000" ;; my LLM complained about the '% 2800000', looks like it's selecting from 1..2.8M in a 1.6M table, so the updates above 1.6M don't affect any rows. I'll leave these grinding for now, I'm trying to scare up more hardware to run more copies since it's so rare. Tycho ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-28 21:34 ` Tycho Andersen @ 2026-07-30 11:02 ` Alexander Lougovski 2026-07-30 16:26 ` Tycho Andersen 0 siblings, 1 reply; 17+ messages in thread From: Alexander Lougovski @ 2026-07-30 11:02 UTC (permalink / raw) To: tycho Cc: Alexander Lougovski, bp, kvm, linux-kernel, pbonzini, thomas.lendacky, vkuznets On Tue, Jul 28, 2026 at 03:34:36PM -0600, Tycho Andersen wrote: > my LLM complained about the '% 2800000', looks like it's selecting > from 1..2.8M in a 1.6M table, so the updates above 1.6M don't affect > any rows. Hi Tycho, ah, true. Your LLM is correct. I've been experimenting with the different sizes of DB to ensure more memory churn (avoid sql reading from the cache) but then settled back to 1.6M. Another thing to mention, it looks like 1582 workload version I've shared doesn't work correctly. So I'm attaching an improved version which mitigates problems - the version I shared before, was still generating a noticeable amount of TLB flushes but still wasn't exactly working as it's supposed to. So here is the improved version: ---------8<------------- $ErrorActionPreference = 'Stop' $logFile = "C:\drivers\workload-1582-errors.log" $block = { param($logFile) $files = @( "$env:SystemRoot\System32\ntdll.dll", "$env:SystemRoot\System32\kernel32.dll", "$env:SystemRoot\System32\user32.dll", "$env:SystemRoot\System32\advapi32.dll", "$env:SystemRoot\System32\ole32.dll", "$env:SystemRoot\System32\shell32.dll", "$env:SystemRoot\System32\comctl32.dll", "$env:SystemRoot\System32\msvcrt.dll" ) $i = 0 while ($true) { $f = $files[(Get-Random -Maximum $files.Count)] try { $fs = [System.IO.FileStream]::new($f, [System.IO.FileMode]::Open, [System.IO.FileAccess]::Read, [System.IO.FileShare]::Read) $mmf = [System.IO.MemoryMappedFiles.MemoryMappedFile]::CreateFromFile( $fs, [NullString]::Value, 0, [System.IO.MemoryMappedFiles.MemoryMappedFileAccess]::Read, [System.IO.HandleInheritability]::None, $false) $view = $mmf.CreateViewAccessor(0, 0, [System.IO.MemoryMappedFiles.MemoryMappedFileAccess]::Read) for ($p = 0; $p -lt [Math]::Min($view.Capacity, 65536); $p += 4096) { $null = $view.ReadByte($p) } $view.Dispose() $mmf.Dispose() } catch { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') iter=$i file=$f error=$($_.Exception.Message)" | Out-File -Append $logFile } $i++ if ($i % 50 -eq 0) { try { [GC]::Collect() [GC]::WaitForPendingFinalizers() } catch { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') iter=$i GC error=$($_.Exception.Message)" | Out-File -Append $logFile } } } } "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') Workload1582 starting (PID=$PID)" | Out-File -Append $logFile try { 1..2 | ForEach-Object { Start-Job -ScriptBlock $block -ArgumentList $logFile } while ($true) { $jobs = Get-Job $failed = $jobs | Where-Object { $_.State -eq 'Failed' } if ($failed) { foreach ($j in $failed) { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') Job $($j.Id) FAILED: $($j.ChildJobs[0].JobStateInfo.Reason)" | Out-File -Append $logFile Remove-Job $j -Force Start-Job -ScriptBlock $block -ArgumentList $logFile } } $running = @(Get-Job | Where-Object { $_.State -eq 'Running' }) if ($running.Count -eq 0) { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') ALL JOBS DEAD, restarting" | Out-File -Append $logFile 1..2 | ForEach-Object { Start-Job -ScriptBlock $block -ArgumentList $logFile } } Start-Sleep -Seconds 60 } } catch { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') PARENT CAUGHT EXCEPTION: $($_.Exception.GetType().FullName): $($_.Exception.Message)" | Out-File -Append $logFile "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') STACK: $($_.ScriptStackTrace)" | Out-File -Append $logFile } finally { "$(Get-Date -Format 'yyyy-MM-dd HH:mm:ss') PARENT EXITING (PID=$PID) LastExitCode=$LASTEXITCODE" | Out-File -Append $logFile } ---------8<------------- Also here is my qemu cmd: /usr/libexec/qemu-kvm -name vm-1,debug-threads=on -machine pc-q35-rhel9.6.0,usb=off,dump-guest-core=off,hpet=off,acpi=on -accel kvm -cpu host,hv-time=on,hv-relaxed=on,hv-vapic=on,hv-spinlocks=0x1fff,hv-vpindex=on,hv-runtime=on,hv-synic=on,hv-reset=on,hv-frequencies=on,hv-tlbflush=on,hv-ipi=on -m 16384 -object memory-backend-ram,id=ram-node0,size=16G,host-nodes=0,policy=bind -numa node,nodeid=0,cpus=0-3,memdev=ram-node0 -smp 4,sockets=1,dies=1,cores=4,threads=1 -rtc base=localtime,driftfix=slew -global kvm-pit.lost_tick_policy=delay -device virtio-blk-pci,drive=drive0,write-cache=on -drive file=/mnt/nfs-fleet/vm-1.raw,if=none,id=drive0,format=raw,aio=native,cache.direct=on,discard=unmap -device virtio-net-pci,netdev=net0,mac=XX:XX:XX:XX:XX:XX,mq=on,vectors=18,host_mtu=9000 -netdev tap,id=net0,ifname=tap1,script=no,downscript=no -device virtio-serial-pci -device virtio-balloon-pci -device virtio-rng-pci,rng=objrng0 -object rng-random,id=objrng0,filename=/dev/urandom -device vmcoreinfo -vga virtio -serial unix:/tmp/serial-1.sock,server,nowait -qmp unix:/tmp/qmp-1.sock,server,nowait -display none -daemonize which follows the issue reported in the virtio-win project: https://github.com/virtio-win/kvm-guest-drivers-windows/issues/1582 Thanks, Al ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-30 11:02 ` Alexander Lougovski @ 2026-07-30 16:26 ` Tycho Andersen 2026-08-01 14:53 ` Alexander Lougovski 0 siblings, 1 reply; 17+ messages in thread From: Tycho Andersen @ 2026-07-30 16:26 UTC (permalink / raw) To: Alexander Lougovski Cc: bp, kvm, linux-kernel, pbonzini, thomas.lendacky, vkuznets On Thu, Jul 30, 2026 at 01:02:46PM +0200, Alexander Lougovski wrote: > On Tue, Jul 28, 2026 at 03:34:36PM -0600, Tycho Andersen wrote: > > my LLM complained about the '% 2800000', looks like it's selecting > > from 1..2.8M in a 1.6M table, so the updates above 1.6M don't affect > > any rows. > > Hi Tycho, > ah, true. Your LLM is correct. I've been experimenting with the different sizes of DB to ensure more memory churn (avoid sql reading from the cache) but then settled back to 1.6M. Ok, makes sense. I found enough hardware yesterday to run ~35 VMs with a bit of memory pressure, so smaller is better for me. When you say "200 VM hours", I guess that's an average? Or do you find they need to run that long to see the fault? Also, how are you detecting the BSOD? I'm just waiting for ssh to stop responding and screen capping the VNC, but maybe there's a better way. > > Another thing to mention, it looks like 1582 workload version I've shared doesn't work correctly. So I'm attaching an improved version which mitigates problems - the version I shared before, was still generating a noticeable amount of TLB flushes but still wasn't exactly working as it's supposed to. So here is the improved version: Thanks for this, I'll update. Though based on: #!/usr/bin/env bpftrace kprobe:svm_flush_tlb_gva { $vcpu = (struct kvm_vcpu *)arg0; printf("%-14llu %-7d %-16s %-4d 0x%lx\n", nsecs / 1000000, pid, comm, $vcpu->vcpu_id, arg1); } I was already getting quite a few of these flushes, so your scripts are working somewhat :) > -cpu host,hv-time=on,hv-relaxed=on,hv-vapic=on,hv-spinlocks=0x1fff,hv-vpindex=on,hv-runtime=on,hv-synic=on,hv-reset=on,hv-frequencies=on,hv-tlbflush=on,hv-ipi=on Thanks for this as well, I only had hv-tlbflush=on, I'll go ahead and adapt my command line to yours. Cheers, Tycho ^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled 2026-07-30 16:26 ` Tycho Andersen @ 2026-08-01 14:53 ` Alexander Lougovski 0 siblings, 0 replies; 17+ messages in thread From: Alexander Lougovski @ 2026-08-01 14:53 UTC (permalink / raw) To: tycho Cc: Alexander Lougovski, bp, kvm, linux-kernel, pbonzini, thomas.lendacky, vkuznets On Thu, Jul 30, 2026 at 10:26:46AM -0600, Tycho Andersen wrote: Hi Tycho, > Ok, makes sense. I found enough hardware yesterday to run ~35 VMs with > a bit of memory pressure, so smaller is better for me. I run it on a 2 socket single host with 128c/256t and 1.5TB RAM - not really pushing memory limits. So cannot comment if actually memory pressure is required. > When you say "200 VM hours", I guess that's an average? Or do you find > they need to run that long to see the fault? Most of the times I would say I see the first BSOD on 42VM fleet within first 10-11 hours after experiment's start -> equovalent of 400-500 VM-hours. However if running over a course of several days, average time drops to more like 200-300 VM-hours per crash. Also sometimes they batch up. E.g. a few within 1-2 hours and then nothing for a day. > Also, how are you detecting the BSOD? I'm just waiting for ssh to stop > responding and screen capping the VNC, but maybe there's a better way. Yeah that's a bit if pain. SSH is not always reliable, since I'm somewhat thrashing VMs (also storage under the sql pressure get's somewhat slow with latency spikes up to several seconds in VM) - timeout aren't that uncommon . So what appeared to work better is using SAC over serial. Here is a script I use. As a bonus it also keeps a track of VMs' uptimes. ---------8<---------------- #!/usr/bin/env bash LOG=/tmp/serial-stress.log UPTIME_LOG=/tmp/serial-uptimes.log CYCLE=0 while true; do CYCLE=$((CYCLE + 1)) TS=$(date '+%Y-%m-%d %H:%M:%S') OK=0 FAIL="" UPTIMES="" for N in $(seq 1 42); do UP=$( (sleep 0.2; printf "\r\nid\r\n"; sleep 1.2) | timeout 3 socat UNIX-CONNECT:/tmp/serial-${N}.sock STDIO 2>/dev/null | grep -o "Time since last reboot:.*" | sed 's/Time since last reboot: //' | tr -d '\r') if [ -z "$UP" ]; then FAIL="$FAIL vm-$N" else OK=$((OK+1)) UPTIMES="$UPTIMES vm-$N=$UP" fi done if [ -n "$FAIL" ]; then echo "[$TS] cycle=$CYCLE ${OK}/42 FAIL:$FAIL" >> $LOG else echo "[$TS] cycle=$CYCLE 42/42 OK" >> $LOG fi if [ $((CYCLE % 10)) -eq 0 ]; then echo "[$TS] cycle=$CYCLE$UPTIMES" >> $UPTIME_LOG fi sleep 5 done --------8<--------------- most of the time you if vm is listed in the output (I personally redirect it to stress-serial.log file and watch it) as FAILED for 2-3 consequative rounds - it is a BSOD, then your agent can take a screenshot and confirm. > Thanks for this as well, I only had hv-tlbflush=on, I'll go ahead and > adapt my command line to yours. Not 100% if others do matter, but that's how I run it. Another thing to notice, I'm somehow more successful reproducing BSODs when storage path is stressed. When I stress networking path I get BSOD waaaaay more seldome, and cannot ATM state that these BSODs have the same root cause even thought they also indicate signs of memory corruption/stale cache. Hope it helps, Let me know if you were able to hit a BSOD. With 35 VMs I would expect to see it within first 15 hours or so. Thanks, Al ^ permalink raw reply [flat|nested] 17+ messages in thread
end of thread, other threads:[~2026-09-03 21:16 UTC | newest] Thread overview: 17+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-07-23 9:44 [PATCH] KVM: SVM: make svm_flush_tlb_gva do a full asid flush if NPT enabled Paolo Bonzini 2026-07-24 22:36 ` Sean Christopherson 2026-07-25 13:37 ` Paolo Bonzini 2026-07-24 23:42 ` Yosry Ahmed 2026-07-24 23:50 ` Yosry Ahmed 2026-07-25 13:32 ` Paolo Bonzini 2026-07-25 21:33 ` Yosry Ahmed 2026-09-03 21:16 ` Joris de Vries 2026-07-27 19:23 ` Tycho Andersen 2026-07-27 19:58 ` Yosry Ahmed 2026-07-29 23:41 ` Paolo Bonzini 2026-07-30 0:24 ` Yosry Ahmed 2026-07-28 17:13 ` Alexander Lougovski 2026-07-28 21:34 ` Tycho Andersen 2026-07-30 11:02 ` Alexander Lougovski 2026-07-30 16:26 ` Tycho Andersen 2026-08-01 14:53 ` Alexander Lougovski
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®