mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 00/11] KVM: fix issues with stale control fields
@ 2026-09-26  5:32 Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 01/11] KVM: SVM: Preserve TLB control (i.e. pending TLB flush) on failed VMRUN Paolo Bonzini
                   ` (12 more replies)
  0 siblings, 13 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm

Fix two bugs where the guest could do stupid things on purpose to
cause problems in the host.

Patches 1-5 cover cases where actions done through VMCB control fields
have to be redone if VMRUN fails.  In particular, failed VMRUNs can
cause pending TLB flushes to be dropped.

Patch 6 fixes a case where eVMCS execution controls can cause the
host to use a stale MSR permission bitmap.  Patches 7-11 are tests
for nested x2APIC; don't run them on an unpatched kernel.

Paolo

Sean Christopherson (11):
  KVM: SVM: Preserve TLB control (i.e. pending TLB flush) on failed
    VMRUN
  KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN
    succeeded
  KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl
  KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on
    successful VMRUN
  KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is
    intercepted
  KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are
    modified
  KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested
  KVM: selftests: Run the nested x2APIC with and without APICv being
    inhibited in L2
  KVM: selftests: Verify that L0's TPR doesn't get clobbered
  KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC
    virt
  KVM: selftests: Extend nested x2APIC test to validate using eVMCS for
    vmcs12

 arch/x86/kvm/svm/sev.c                        |   1 -
 arch/x86/kvm/svm/svm.c                        |  43 ++--
 arch/x86/kvm/svm/svm.h                        |   1 +
 arch/x86/kvm/vmx/nested.c                     |   6 +
 tools/testing/selftests/kvm/Makefile.kvm      |   1 +
 .../selftests/kvm/x86/nested_x2apic_test.c    | 235 ++++++++++++++++++
 6 files changed, 263 insertions(+), 24 deletions(-)
 create mode 100644 tools/testing/selftests/kvm/x86/nested_x2apic_test.c

-- 
2.52.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 01/11] KVM: SVM: Preserve TLB control (i.e. pending TLB flush) on failed VMRUN
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 02/11] KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded Paolo Bonzini
                   ` (11 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm
  Cc: Sean Christopherson, stable, Stefan Teodorescu, Yosry Ahmed,
	Tom Lendacky, Jim Mattson

From: Sean Christopherson <seanjc@google.com>

Don't reset the VMCB's TLB control back to "do nothing" on a failed VMRUN,
as empirical testing shows that the CPU performs the requested TLB flush if
and only if VMRUN is successful, i.e. clearing TLB control on a failed
VMRUN effectively drops a TLB flush.

Explicitly track the need to flush all ASIDs on a per-CPU basis, as the
ASID reuse condition is tied to the pCPU, not to the vCPU.  As a bonus,
this also obviates the need to avoid clobbering FLUSH_ALL_ASID with
TLB_CONTROL_FLUSH_ASID, e.g. in svm_flush_tlb_asid().

Deliberately don't bother saving/restoring the "old" tlb_ctl on failure,
in quotes because it's not exactly the old tlb_ctl, it's the tlb_ctl from
after pre_svm_run(), but before updating tlb_ctl for flush_all_asids.  If
VMRUN fails and TLB_CONTROL_FLUSH_ALL_ASID is forced, then the next
successful run of the VMCB *may* unnecessarily flush all ASIDs, which
strictly speaking could result in noisy neighbor issues.  However, the
fact that new_asid() is already guest-triggerable, because of KVM's flawed
behavior of clearing the ASID on emulated INIT, means that a guest can
already trigger a flush of all ASIDs at roughly the same rate.  And once
KVM stops clobbering the ASID on emulated INIT, *or* assigns a static ASID
to each vCPU, this flaw goes away.

Fixes: 38e5e92fe8c0 ("KVM: SVM: Implement Flush-By-Asid feature")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Cc: Tom Lendacky <thomas.lendacky@amd.com>
Cc: Jim Mattson <jmattson@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-2-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 arch/x86/kvm/svm/svm.c | 13 ++++++++++---
 arch/x86/kvm/svm/svm.h |  1 +
 2 files changed, 11 insertions(+), 3 deletions(-)

diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index 7d59d301e1e5..afbaaaab84ed 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -1902,8 +1902,7 @@ static void new_asid(struct vcpu_svm *svm, struct svm_cpu_data *sd)
 	if (sd->next_asid > sd->max_asid) {
 		++sd->asid_generation;
 		sd->next_asid = sd->min_asid;
-		svm->vmcb->control.tlb_ctl = TLB_CONTROL_FLUSH_ALL_ASID;
-		vmcb_mark_dirty(svm->vmcb, VMCB_ASID);
+		sd->flush_all_asids = true;
 	}
 
 	svm->current_vmcb->asid_generation = sd->asid_generation;
@@ -4528,6 +4527,11 @@ static __no_kcsan fastpath_t svm_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags)
 		svm->vmcb->control.asid = svm->asid;
 		vmcb_mark_dirty(svm->vmcb, VMCB_ASID);
 	}
+	if (this_cpu_ptr(&svm_data)->flush_all_asids) {
+		svm->vmcb->control.tlb_ctl = TLB_CONTROL_FLUSH_ALL_ASID;
+		vmcb_mark_dirty(svm->vmcb, VMCB_ASID);
+	}
+
 	svm->vmcb->save.cr2 = vcpu->arch.cr2;
 
 	if (guest_cpu_cap_has(vcpu, X86_FEATURE_ERAPS) &&
@@ -4618,7 +4622,10 @@ static __no_kcsan fastpath_t svm_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags)
 		vcpu->arch.nested_run_pending = 0;
 	}
 
-	svm->vmcb->control.tlb_ctl = TLB_CONTROL_DO_NOTHING;
+	if (!svm_is_vmrun_failure(svm->vmcb->control.exit_code)) {
+		this_cpu_ptr(&svm_data)->flush_all_asids = false;
+		svm->vmcb->control.tlb_ctl = TLB_CONTROL_DO_NOTHING;
+	}
 
 	/*
 	 * Unconditionally mask off the CLEAR_RAP bit, the AND is just as cheap
diff --git a/arch/x86/kvm/svm/svm.h b/arch/x86/kvm/svm/svm.h
index e958943b8162..84f19026d3e8 100644
--- a/arch/x86/kvm/svm/svm.h
+++ b/arch/x86/kvm/svm/svm.h
@@ -376,6 +376,7 @@ struct svm_cpu_data {
 	u32 next_asid;
 	u32 min_asid;
 
+	bool flush_all_asids;
 	bool bp_spec_reduce_set;
 
 	struct vmcb *save_area;
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 02/11] KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 01/11] KVM: SVM: Preserve TLB control (i.e. pending TLB flush) on failed VMRUN Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 03/11] KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl Paolo Bonzini
                   ` (10 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson, stable

From: Sean Christopherson <seanjc@google.com>

Leave control.erap_ctl and control.clean as-is in the VMCS if VMRUN fails,
because as per AMD:

  there's no explicit architectural guarantee about the behavior in the
  presence of VMRUN failures. So the best thing to do would be to assume
  that if VMRUN fails, the actions requested in the control fields may not
  have been performed.

Cc: stable@vger.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 arch/x86/kvm/svm/svm.c | 18 +++++++++---------
 1 file changed, 9 insertions(+), 9 deletions(-)

diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index afbaaaab84ed..b63e7c69aa1c 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -4625,17 +4625,17 @@ static __no_kcsan fastpath_t svm_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags)
 	if (!svm_is_vmrun_failure(svm->vmcb->control.exit_code)) {
 		this_cpu_ptr(&svm_data)->flush_all_asids = false;
 		svm->vmcb->control.tlb_ctl = TLB_CONTROL_DO_NOTHING;
+
+		/*
+		 * Unconditionally mask off the CLEAR_RAP bit, the AND is just
+		 * as cheap as the TEST+Jcc to avoid it.
+		 */
+		if (cpu_feature_enabled(X86_FEATURE_ERAPS))
+			svm->vmcb->control.erap_ctl &= ~ERAP_CONTROL_CLEAR_RAP;
+
+		vmcb_mark_all_clean(svm->vmcb);
 	}
 
-	/*
-	 * Unconditionally mask off the CLEAR_RAP bit, the AND is just as cheap
-	 * as the TEST+Jcc to avoid it.
-	 */
-	if (cpu_feature_enabled(X86_FEATURE_ERAPS))
-		svm->vmcb->control.erap_ctl &= ~ERAP_CONTROL_CLEAR_RAP;
-
-	vmcb_mark_all_clean(svm->vmcb);
-
 	/* if exit due to PF check for async PF */
 	if (svm->vmcb->control.exit_code == SVM_EXIT_EXCP_BASE + PF_VECTOR)
 		vcpu->arch.apf.host_apf_flags =
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 03/11] KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 01/11] KVM: SVM: Preserve TLB control (i.e. pending TLB flush) on failed VMRUN Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 02/11] KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 04/11] KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN Paolo Bonzini
                   ` (9 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson, Yosry Ahmed

From: Sean Christopherson <seanjc@google.com>

Don't mark the ASID as dirty in the VMCB when requesting a TLB flush via
control.tlb_ctl.  Per "15.15.3 VMCB Clean Field" of the July 2026, Revision
3.45 version of the APM:

  The following are explicitly not cached and not represented by Clean bits:

    * TLB_Control

Fixes: 7e8e6eed75e2 ("KVM: SVM: Move asid to vcpu_svm")
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-4-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 arch/x86/kvm/svm/sev.c | 1 -
 arch/x86/kvm/svm/svm.c | 4 +---
 2 files changed, 1 insertion(+), 4 deletions(-)

diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c
index 6aa86a78711e..3448d56520c6 100644
--- a/arch/x86/kvm/svm/sev.c
+++ b/arch/x86/kvm/svm/sev.c
@@ -3620,7 +3620,6 @@ int pre_sev_run(struct vcpu_svm *svm, int cpu)
 
 	sd->sev_vmcbs[asid] = svm->vmcb;
 	svm->vmcb->control.tlb_ctl = TLB_CONTROL_FLUSH_ASID;
-	vmcb_mark_dirty(svm->vmcb, VMCB_ASID);
 	return 0;
 }
 
diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index b63e7c69aa1c..830ace75e986 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -4527,10 +4527,8 @@ static __no_kcsan fastpath_t svm_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags)
 		svm->vmcb->control.asid = svm->asid;
 		vmcb_mark_dirty(svm->vmcb, VMCB_ASID);
 	}
-	if (this_cpu_ptr(&svm_data)->flush_all_asids) {
+	if (this_cpu_ptr(&svm_data)->flush_all_asids)
 		svm->vmcb->control.tlb_ctl = TLB_CONTROL_FLUSH_ALL_ASID;
-		vmcb_mark_dirty(svm->vmcb, VMCB_ASID);
-	}
 
 	svm->vmcb->save.cr2 = vcpu->arch.cr2;
 
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 04/11] KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (2 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 03/11] KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 05/11] KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted Paolo Bonzini
                   ` (8 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson

From: Sean Christopherson <seanjc@google.com>

Don't (re)read PERF_CNTR_GLOBAL_CTL from hardware on a failed VMRUN, as the
purpose of the read is to synchronize KVM's cache with any writes done by
the guest, and the guest can't possibly have modified the MSR if it never
got a chance to run.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260904170642.3291466-5-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 arch/x86/kvm/svm/svm.c | 7 ++++---
 1 file changed, 4 insertions(+), 3 deletions(-)

diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index 830ace75e986..f4f488328ea4 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -4632,6 +4632,10 @@ static __no_kcsan fastpath_t svm_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags)
 			svm->vmcb->control.erap_ctl &= ~ERAP_CONTROL_CLEAR_RAP;
 
 		vmcb_mark_all_clean(svm->vmcb);
+
+		if (!msr_write_intercepted(svm, MSR_AMD64_PERF_CNTR_GLOBAL_CTL))
+			rdmsrq(MSR_AMD64_PERF_CNTR_GLOBAL_CTL,
+			       vcpu_to_pmu(vcpu)->global_ctrl);
 	}
 
 	/* if exit due to PF check for async PF */
@@ -4641,9 +4645,6 @@ static __no_kcsan fastpath_t svm_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags)
 
 	kvm_clear_available_registers(vcpu, SVM_REGS_LAZY_LOAD_SET);
 
-	if (!msr_write_intercepted(svm, MSR_AMD64_PERF_CNTR_GLOBAL_CTL))
-		rdmsrq(MSR_AMD64_PERF_CNTR_GLOBAL_CTL, vcpu_to_pmu(vcpu)->global_ctrl);
-
 	trace_kvm_exit(vcpu, KVM_ISA_SVM);
 
 	svm_complete_interrupts(vcpu);
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 05/11] KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (3 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 04/11] KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 06/11] KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified Paolo Bonzini
                   ` (7 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson, stable, Stefan Teodorescu

From: Sean Christopherson <seanjc@google.com>

Use the MSR permission bitmap of the active VMCB instead of assuming that
KVM is always using vmcb02's bitmap when L2 is active, as KVM uses msrpm02
if and only if L1 wants to intercept MSR accesses, i.e. if and only if KVM
needs to merge msprm01 with msrpm12.

Don't bother tracking the virtual address of the bitmap that's being used,
as __va() is cheap on x86, and caching the virtual address would introduce
yet another source of potentially stale information.

Fixes: b2ac58f90540 ("KVM/SVM: Allow direct access to MSR_IA32_SPEC_CTRL")
Cc: stable@vger.kernel.org
Reported-by: Stefan Teodorescu <fane@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260826195833.844526-1-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 arch/x86/kvm/svm/svm.c | 11 +----------
 1 file changed, 1 insertion(+), 10 deletions(-)

diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index f4f488328ea4..1f279cd91ecf 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -673,16 +673,7 @@ static void clr_dr_intercepts(struct vcpu_svm *svm)
 
 static bool msr_write_intercepted(struct vcpu_svm *svm, u32 msr)
 {
-	/*
-	 * For non-nested case:
-	 * If the L01 MSR bitmap does not intercept the MSR, then we need to
-	 * save it.
-	 *
-	 * For nested case:
-	 * If the L02 MSR bitmap does not intercept the MSR, then we need to
-	 * save it.
-	 */
-	void *msrpm = is_guest_mode(&svm->vcpu) ? svm->nested.msrpm : svm->msrpm;
+	void *msrpm = __va(svm->vmcb->control.msrpm_base_pa);
 
 	return svm_test_msr_bitmap_write(msrpm, msr);
 }
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 06/11] KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (4 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 05/11] KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 07/11] KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested Paolo Bonzini
                   ` (6 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson, stable, Vitaly Kuznetsov

From: Sean Christopherson <seanjc@google.com>

Force a refresh of the vmcs02 MSR bitmap during nested VM-Enter if the
runtime eVMCS controls (pin, primary, secondary, etc.) are being updated.
If L1 isn't intercepting TPR writes, runs L2 with TPR virtualization, and
then runs the same L2 with TPR virtualization disabled, KVM will fail to
refresh msr_bitmap02 and leave TPR in passthrough mode even though TPR
virtualization is disabled.  I.e. failure to refresh the bitmap lets L2 (or
L1 by proxy) read and write L0's TPR.

Fixes: 502d2bf5f2fd ("KVM: nVMX: Implement Enlightened MSR Bitmap feature")
Cc: stable@vger.kernel.org
Reviewed-by: Vitaly Kuznetsov <vkuznets@redhat.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 arch/x86/kvm/vmx/nested.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
index 40c1a5f6fa8a..b25216862740 100644
--- a/arch/x86/kvm/vmx/nested.c
+++ b/arch/x86/kvm/vmx/nested.c
@@ -1754,6 +1754,9 @@ static void copy_vmcs12_to_shadow(struct vcpu_vmx *vmx)
 static void copy_enlightened_to_vmcs12(struct vcpu_vmx *vmx, u32 hv_clean_fields)
 {
 #ifdef CONFIG_KVM_HYPERV
+	const u64 runtime_controls = HV_VMX_ENLIGHTENED_CLEAN_FIELD_CONTROL_GRP1 |
+				     HV_VMX_ENLIGHTENED_CLEAN_FIELD_CONTROL_GRP2 |
+				     HV_VMX_ENLIGHTENED_CLEAN_FIELD_CONTROL_PROC;
 	struct vmcs12 *vmcs12 = vmx->nested.cached_vmcs12;
 	struct hv_enlightened_vmcs *evmcs = nested_vmx_evmcs(vmx);
 	struct kvm_vcpu_hv *hv_vcpu = to_hv_vcpu(&vmx->vcpu);
@@ -1762,6 +1765,9 @@ static void copy_enlightened_to_vmcs12(struct vcpu_vmx *vmx, u32 hv_clean_fields
 	vmcs12->tpr_threshold = evmcs->tpr_threshold;
 	vmcs12->guest_rip = evmcs->guest_rip;
 
+	if ((hv_clean_fields & runtime_controls) != runtime_controls)
+		vmx->nested.force_msr_bitmap_recalc = true;
+
 	if (unlikely(!(hv_clean_fields &
 		       HV_VMX_ENLIGHTENED_CLEAN_FIELD_ENLIGHTENMENTSCONTROL))) {
 		hv_vcpu->nested.pa_page_gpa = evmcs->partition_assist_page;
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 07/11] KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (5 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 06/11] KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 08/11] KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2 Paolo Bonzini
                   ` (5 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson

From: Sean Christopherson <seanjc@google.com>

Add a selftest to verify that KVM intercepts x2APIC MSR accesses for L1
after APICv is inhibited while L2 is active.  This is a regression test for
an AVIC bug where KVM would skip updating x2APIC MSR intercepts while L2
is active, thus giving L1 access to a wide swath of L0's x2APIC surface.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260710162052.2188574-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 tools/testing/selftests/kvm/Makefile.kvm      |   1 +
 .../selftests/kvm/x86/nested_x2apic_test.c    | 117 ++++++++++++++++++
 2 files changed, 118 insertions(+)
 create mode 100644 tools/testing/selftests/kvm/x86/nested_x2apic_test.c

diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm
index 6a1482e3a286..752f81b070fa 100644
--- a/tools/testing/selftests/kvm/Makefile.kvm
+++ b/tools/testing/selftests/kvm/Makefile.kvm
@@ -103,6 +103,7 @@ TEST_GEN_PROGS_x86 += x86/nested_tdp_fault_test
 TEST_GEN_PROGS_x86 += x86/nested_tsc_adjust_test
 TEST_GEN_PROGS_x86 += x86/nested_tsc_scaling_test
 TEST_GEN_PROGS_x86 += x86/nested_vmsave_vmload_test
+TEST_GEN_PROGS_x86 += x86/nested_x2apic_test
 TEST_GEN_PROGS_x86 += x86/platform_info_test
 TEST_GEN_PROGS_x86 += x86/pmu_counters_test
 TEST_GEN_PROGS_x86 += x86/pmu_event_filter_test
diff --git a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
new file mode 100644
index 000000000000..e77b4c347272
--- /dev/null
+++ b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
@@ -0,0 +1,117 @@
+// SPDX-License-Identifier: GPL-2.0-only
+#include "test_util.h"
+#include "kvm_util.h"
+#include "processor.h"
+#include "vmx.h"
+#include "svm_util.h"
+
+/*
+ * Use the kernel's posted interrupt vectors to minimize the risk of crashing
+ * the host if KVM is buggy.  Note, the vectors aren't set in stone, ideally
+ * these will be kept up-to-date if the kernel vectors change, but it's "fine"
+ * if they are stale.
+ */
+#define POSTED_INTR_VECTOR		0xf2
+#define POSTED_INTR_WAKEUP_VECTOR	0xf1
+#define POSTED_INTR_NESTED_VECTOR	0xf0
+
+static volatile unsigned int nr_irqs;
+
+static void guest_irq_handler(struct ex_regs *regs)
+{
+	nr_irqs++;
+	x2apic_write_reg(APIC_EOI, 0);
+}
+
+static void l2_guest_code(void)
+{
+	wrmsr(MSR_IA32_APICBASE, rdmsr(MSR_IA32_APICBASE) & GENMASK_ULL(11, 0));
+	asm volatile("cpuid" ::: "eax", "ebx", "ecx", "edx");
+}
+
+static void l1_svm_code(struct svm_test_data *svm)
+{
+	struct vmcb_control_area *ctrl = &svm->vmcb->control;
+
+	generic_svm_setup(svm, l2_guest_code);
+	ctrl->intercept |= BIT_ULL(INTERCEPT_CPUID) | BIT_ULL(INTERCEPT_MSR_PROT);
+
+	run_guest(svm->vmcb, svm->vmcb_gpa);
+	GUEST_ASSERT_EQ(ctrl->exit_code, SVM_EXIT_CPUID);
+
+	stgi();
+}
+
+static void l1_vmx_code(struct vmx_pages *vmx)
+{
+	u64 control;
+
+	GUEST_ASSERT_EQ(prepare_for_vmx_operation(vmx), true);
+	GUEST_ASSERT_EQ(load_vmcs(vmx), true);
+
+	prepare_vmcs(vmx, NULL);
+	GUEST_ASSERT_EQ(vmwrite(GUEST_RIP, (unsigned long)l2_guest_code), 0);
+
+	control = vmreadz(CPU_BASED_VM_EXEC_CONTROL);
+	control |= CPU_BASED_USE_MSR_BITMAPS;
+	GUEST_ASSERT_EQ(vmwrite(CPU_BASED_VM_EXEC_CONTROL, control), 0);
+
+	GUEST_ASSERT(!vmlaunch());
+	GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), EXIT_REASON_CPUID);
+}
+
+static void l1_guest_code(void *test_data)
+{
+	x2apic_enable();
+
+	if (this_cpu_has(X86_FEATURE_SVM))
+		l1_svm_code(test_data);
+	else
+		l1_vmx_code(test_data);
+
+	sti_nop();
+
+	x2apic_write_reg(APIC_ICR, APIC_DEST_SELF | APIC_INT_ASSERT | POSTED_INTR_VECTOR);
+	x2apic_write_reg(APIC_ICR, APIC_DEST_SELF | APIC_INT_ASSERT | POSTED_INTR_WAKEUP_VECTOR);
+	x2apic_write_reg(APIC_ICR, APIC_DEST_SELF | APIC_INT_ASSERT | POSTED_INTR_NESTED_VECTOR);
+	GUEST_ASSERT_EQ(nr_irqs, 3);
+	GUEST_DONE();
+}
+
+int main(int argc, char *argv[])
+{
+	gva_t nested_test_data_gva;
+	struct kvm_vcpu *vcpu;
+	struct kvm_vm *vm;
+	struct ucall uc;
+
+	TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM) || kvm_cpu_has(X86_FEATURE_VMX));
+
+	vm = vm_create_with_one_vcpu(&vcpu, l1_guest_code);
+	vm_install_exception_handler(vm, POSTED_INTR_VECTOR, guest_irq_handler);
+	vm_install_exception_handler(vm, POSTED_INTR_WAKEUP_VECTOR, guest_irq_handler);
+	vm_install_exception_handler(vm, POSTED_INTR_NESTED_VECTOR, guest_irq_handler);
+
+	if (kvm_cpu_has(X86_FEATURE_SVM))
+		vcpu_alloc_svm(vm, &nested_test_data_gva);
+	else
+		vcpu_alloc_vmx(vm, &nested_test_data_gva);
+
+	vcpu_args_set(vcpu, 1, nested_test_data_gva);
+
+	vcpu_run(vcpu);
+
+	TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO);
+
+	switch (get_ucall(vcpu, &uc)) {
+	case UCALL_DONE:
+		break;
+	case UCALL_ABORT:
+		REPORT_GUEST_ASSERT(uc);
+		break;
+	default:
+		TEST_FAIL("Expected DONE, got unexpected ucall %lu", uc.cmd);
+	}
+
+	kvm_vm_free(vm);
+}
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 08/11] KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (6 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 07/11] KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 09/11] KVM: selftests: Verify that L0's TPR doesn't get clobbered Paolo Bonzini
                   ` (4 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson

From: Sean Christopherson <seanjc@google.com>

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-4-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 .../selftests/kvm/x86/nested_x2apic_test.c    | 25 ++++++++++++++++---
 1 file changed, 21 insertions(+), 4 deletions(-)

diff --git a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
index e77b4c347272..a1072bf499ee 100644
--- a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
+++ b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
@@ -15,6 +15,7 @@
 #define POSTED_INTR_WAKEUP_VECTOR	0xf1
 #define POSTED_INTR_NESTED_VECTOR	0xf0
 
+static bool inhibit_apicv;
 static volatile unsigned int nr_irqs;
 
 static void guest_irq_handler(struct ex_regs *regs)
@@ -25,7 +26,8 @@ static void guest_irq_handler(struct ex_regs *regs)
 
 static void l2_guest_code(void)
 {
-	wrmsr(MSR_IA32_APICBASE, rdmsr(MSR_IA32_APICBASE) & GENMASK_ULL(11, 0));
+	if (inhibit_apicv)
+		wrmsr(MSR_IA32_APICBASE, rdmsr(MSR_IA32_APICBASE) & GENMASK_ULL(11, 0));
 	asm volatile("cpuid" ::: "eax", "ebx", "ecx", "edx");
 }
 
@@ -78,20 +80,20 @@ static void l1_guest_code(void *test_data)
 	GUEST_DONE();
 }
 
-int main(int argc, char *argv[])
+static void __test_x2apic_intercepts(void)
 {
 	gva_t nested_test_data_gva;
 	struct kvm_vcpu *vcpu;
 	struct kvm_vm *vm;
 	struct ucall uc;
 
-	TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM) || kvm_cpu_has(X86_FEATURE_VMX));
-
 	vm = vm_create_with_one_vcpu(&vcpu, l1_guest_code);
 	vm_install_exception_handler(vm, POSTED_INTR_VECTOR, guest_irq_handler);
 	vm_install_exception_handler(vm, POSTED_INTR_WAKEUP_VECTOR, guest_irq_handler);
 	vm_install_exception_handler(vm, POSTED_INTR_NESTED_VECTOR, guest_irq_handler);
 
+	sync_global_to_guest(vm, inhibit_apicv);
+
 	if (kvm_cpu_has(X86_FEATURE_SVM))
 		vcpu_alloc_svm(vm, &nested_test_data_gva);
 	else
@@ -115,3 +117,18 @@ int main(int argc, char *argv[])
 
 	kvm_vm_free(vm);
 }
+
+#define test_x2apic_intercepts(inhibit_apic_setting)	\
+do {							\
+	inhibit_apic_setting;				\
+							\
+	__test_x2apic_intercepts();			\
+} while (0)
+
+int main(int argc, char *argv[])
+{
+	TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM) || kvm_cpu_has(X86_FEATURE_VMX));
+
+	test_x2apic_intercepts(inhibit_apicv = true);
+	test_x2apic_intercepts(inhibit_apicv = false);
+}
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 09/11] KVM: selftests: Verify that L0's TPR doesn't get clobbered
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (7 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 08/11] KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2 Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 10/11] KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt Paolo Bonzini
                   ` (3 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson

From: Sean Christopherson <seanjc@google.com>

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-5-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 .../selftests/kvm/x86/nested_x2apic_test.c       | 16 ++++++++++++++++
 1 file changed, 16 insertions(+)

diff --git a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
index a1072bf499ee..3b59ba3e3342 100644
--- a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
+++ b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
@@ -28,6 +28,10 @@ static void l2_guest_code(void)
 {
 	if (inhibit_apicv)
 		wrmsr(MSR_IA32_APICBASE, rdmsr(MSR_IA32_APICBASE) & GENMASK_ULL(11, 0));
+
+	x2apic_write_reg(APIC_TASKPRI, 0xf0);
+	GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0xf0);
+
 	asm volatile("cpuid" ::: "eax", "ebx", "ecx", "edx");
 }
 
@@ -73,10 +77,22 @@ static void l1_guest_code(void *test_data)
 
 	sti_nop();
 
+	GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0xf0);
+
 	x2apic_write_reg(APIC_ICR, APIC_DEST_SELF | APIC_INT_ASSERT | POSTED_INTR_VECTOR);
+	GUEST_ASSERT_EQ(nr_irqs, 0);
+
+	x2apic_write_reg(APIC_TASKPRI, 0xff);
+	GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0xff);
 	x2apic_write_reg(APIC_ICR, APIC_DEST_SELF | APIC_INT_ASSERT | POSTED_INTR_WAKEUP_VECTOR);
+	GUEST_ASSERT_EQ(nr_irqs, 0);
+
+	x2apic_write_reg(APIC_TASKPRI, 0);
+	GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0);
+
 	x2apic_write_reg(APIC_ICR, APIC_DEST_SELF | APIC_INT_ASSERT | POSTED_INTR_NESTED_VECTOR);
 	GUEST_ASSERT_EQ(nr_irqs, 3);
+
 	GUEST_DONE();
 }
 
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 10/11] KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (8 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 09/11] KVM: selftests: Verify that L0's TPR doesn't get clobbered Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  5:32 ` [PATCH 11/11] KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12 Paolo Bonzini
                   ` (2 subsequent siblings)
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson

From: Sean Christopherson <seanjc@google.com>

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-6-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 .../selftests/kvm/x86/nested_x2apic_test.c    | 74 ++++++++++++++++---
 1 file changed, 64 insertions(+), 10 deletions(-)

diff --git a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
index 3b59ba3e3342..e94d4e77256b 100644
--- a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
+++ b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
@@ -29,10 +29,12 @@ static void l2_guest_code(void)
 	if (inhibit_apicv)
 		wrmsr(MSR_IA32_APICBASE, rdmsr(MSR_IA32_APICBASE) & GENMASK_ULL(11, 0));
 
-	x2apic_write_reg(APIC_TASKPRI, 0xf0);
-	GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0xf0);
+	for (;;) {
+		x2apic_write_reg(APIC_TASKPRI, 0xf0);
+		GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0xf0);
 
-	asm volatile("cpuid" ::: "eax", "ebx", "ecx", "edx");
+		asm volatile("cpuid" ::: "eax", "ebx", "ecx", "edx");
+	}
 }
 
 static void l1_svm_code(struct svm_test_data *svm)
@@ -58,22 +60,52 @@ static void l1_vmx_code(struct vmx_pages *vmx)
 	prepare_vmcs(vmx, NULL);
 	GUEST_ASSERT_EQ(vmwrite(GUEST_RIP, (unsigned long)l2_guest_code), 0);
 
+	control = vmreadz(PIN_BASED_VM_EXEC_CONTROL);
+	control |= PIN_BASED_EXT_INTR_MASK;
+	vmwrite(PIN_BASED_VM_EXEC_CONTROL, control);
+
 	control = vmreadz(CPU_BASED_VM_EXEC_CONTROL);
-	control |= CPU_BASED_USE_MSR_BITMAPS;
+	control |= CPU_BASED_USE_MSR_BITMAPS | CPU_BASED_TPR_SHADOW;
 	GUEST_ASSERT_EQ(vmwrite(CPU_BASED_VM_EXEC_CONTROL, control), 0);
 
+	if (control & CPU_BASED_ACTIVATE_SECONDARY_CONTROLS) {
+		control = vmreadz(SECONDARY_VM_EXEC_CONTROL);
+		control |= SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
+			   SECONDARY_EXEC_APIC_REGISTER_VIRT |
+			   SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY;
+		control &= (rdmsr(MSR_IA32_VMX_PROCBASED_CTLS2) >> 32);
+		GUEST_ASSERT_EQ(vmwrite(SECONDARY_VM_EXEC_CONTROL, control), 0);
+	}
+
 	GUEST_ASSERT(!vmlaunch());
 	GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), EXIT_REASON_CPUID);
+	GUEST_ASSERT_EQ(vmwrite(GUEST_RIP,
+			vmreadz(GUEST_RIP) + vmreadz(VM_EXIT_INSTRUCTION_LEN)), 0);
 }
 
-static void l1_guest_code(void *test_data)
+static void l1_vmx_code_part2(void)
 {
-	x2apic_enable();
+	u64 control;
 
-	if (this_cpu_has(X86_FEATURE_SVM))
-		l1_svm_code(test_data);
-	else
-		l1_vmx_code(test_data);
+	control = vmreadz(CPU_BASED_VM_EXEC_CONTROL);
+	control &= ~CPU_BASED_TPR_SHADOW;
+	GUEST_ASSERT_EQ(vmwrite(CPU_BASED_VM_EXEC_CONTROL, control), 0);
+
+	control = vmread(SECONDARY_VM_EXEC_CONTROL, &control);
+	control &= ~(SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
+			SECONDARY_EXEC_APIC_REGISTER_VIRT |
+			SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY);
+	GUEST_ASSERT_EQ(vmwrite(SECONDARY_VM_EXEC_CONTROL, control), 0);
+
+	GUEST_ASSERT(!vmresume());
+	GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), EXIT_REASON_CPUID);
+	GUEST_ASSERT_EQ(vmwrite(GUEST_RIP,
+			vmreadz(GUEST_RIP) + vmreadz(VM_EXIT_INSTRUCTION_LEN)), 0);
+}
+
+static void l1_test_x2apic_intercepts(void)
+{
+	GUEST_ASSERT_EQ(nr_irqs, 0);
 
 	sti_nop();
 
@@ -93,6 +125,28 @@ static void l1_guest_code(void *test_data)
 	x2apic_write_reg(APIC_ICR, APIC_DEST_SELF | APIC_INT_ASSERT | POSTED_INTR_NESTED_VECTOR);
 	GUEST_ASSERT_EQ(nr_irqs, 3);
 
+	nr_irqs = 0;
+}
+
+static void l1_guest_code(void *test_data)
+{
+	x2apic_enable();
+
+	if (this_cpu_has(X86_FEATURE_SVM))
+		l1_svm_code(test_data);
+	else
+		l1_vmx_code(test_data);
+
+	GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0);
+	x2apic_write_reg(APIC_TASKPRI, 0xf0);
+
+	l1_test_x2apic_intercepts();
+
+	if (this_cpu_has(X86_FEATURE_VMX))
+		l1_vmx_code_part2();
+
+	l1_test_x2apic_intercepts();
+
 	GUEST_DONE();
 }
 
-- 
2.52.0



^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH 11/11] KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (9 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 10/11] KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt Paolo Bonzini
@ 2026-09-26  5:32 ` Paolo Bonzini
  2026-09-26  6:11 ` [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
  2026-09-26  6:12 ` Paolo Bonzini
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  5:32 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: Sean Christopherson

From: Sean Christopherson <seanjc@google.com>

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260813223610.2043560-7-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 .../selftests/kvm/x86/nested_x2apic_test.c    | 53 +++++++++++++++----
 1 file changed, 42 insertions(+), 11 deletions(-)

diff --git a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
index e94d4e77256b..eb89d5bfc0fe 100644
--- a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
+++ b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
@@ -50,12 +50,24 @@ static void l1_svm_code(struct svm_test_data *svm)
 	stgi();
 }
 
-static void l1_vmx_code(struct vmx_pages *vmx)
+static void l1_vmx_code(struct vmx_pages *vmx, struct hyperv_test_pages *hv_pages)
 {
 	u64 control;
 
+	if (hv_pages) {
+		wrmsr(HV_X64_MSR_GUEST_OS_ID, HYPERV_LINUX_OS_ID);
+		enable_vp_assist(hv_pages->vp_assist_gpa, hv_pages->vp_assist);
+		evmcs_enable();
+	}
+
 	GUEST_ASSERT_EQ(prepare_for_vmx_operation(vmx), true);
-	GUEST_ASSERT_EQ(load_vmcs(vmx), true);
+
+	if (hv_pages) {
+		GUEST_ASSERT(load_evmcs(hv_pages));
+		current_evmcs->hv_enlightenments_control.msr_bitmap = 1;
+	} else {
+		GUEST_ASSERT(load_vmcs(vmx));
+	}
 
 	prepare_vmcs(vmx, NULL);
 	GUEST_ASSERT_EQ(vmwrite(GUEST_RIP, (unsigned long)l2_guest_code), 0);
@@ -128,14 +140,14 @@ static void l1_test_x2apic_intercepts(void)
 	nr_irqs = 0;
 }
 
-static void l1_guest_code(void *test_data)
+static void l1_guest_code(void *test_data, void *hv_pages)
 {
 	x2apic_enable();
 
 	if (this_cpu_has(X86_FEATURE_SVM))
 		l1_svm_code(test_data);
 	else
-		l1_vmx_code(test_data);
+		l1_vmx_code(test_data, hv_pages);
 
 	GUEST_ASSERT_EQ(x2apic_read_reg(APIC_TASKPRI), 0);
 	x2apic_write_reg(APIC_TASKPRI, 0xf0);
@@ -150,9 +162,9 @@ static void l1_guest_code(void *test_data)
 	GUEST_DONE();
 }
 
-static void __test_x2apic_intercepts(void)
+static void __test_x2apic_intercepts(bool use_evmcs)
 {
-	gva_t nested_test_data_gva;
+	gva_t nested_test_data_gva, hv_pages_gva = 0;
 	struct kvm_vcpu *vcpu;
 	struct kvm_vm *vm;
 	struct ucall uc;
@@ -169,7 +181,14 @@ static void __test_x2apic_intercepts(void)
 	else
 		vcpu_alloc_vmx(vm, &nested_test_data_gva);
 
-	vcpu_args_set(vcpu, 1, nested_test_data_gva);
+	if (use_evmcs) {
+		vcpu_set_hv_cpuid(vcpu);
+		vcpu_enable_evmcs(vcpu);
+
+		vcpu_alloc_hyperv_test_pages(vm, &hv_pages_gva);
+	}
+
+	vcpu_args_set(vcpu, 2, nested_test_data_gva, hv_pages_gva);
 
 	vcpu_run(vcpu);
 
@@ -188,17 +207,29 @@ static void __test_x2apic_intercepts(void)
 	kvm_vm_free(vm);
 }
 
-#define test_x2apic_intercepts(inhibit_apic_setting)	\
+#define _test_x2apic_intercepts(inhibit_apic_setting)	\
 do {							\
+							\
 	inhibit_apic_setting;				\
 							\
-	__test_x2apic_intercepts();			\
+	__test_x2apic_intercepts(use_evmcs);		\
 } while (0)
 
+#define test_x2apic_intercepts(use_evmcs_setting)	\
+do {							\
+	bool use_evmcs_setting;				\
+							\
+	_test_x2apic_intercepts(inhibit_apicv = true);	\
+	_test_x2apic_intercepts(inhibit_apicv = false);	\
+} while (0)
+
+
 int main(int argc, char *argv[])
 {
 	TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM) || kvm_cpu_has(X86_FEATURE_VMX));
 
-	test_x2apic_intercepts(inhibit_apicv = true);
-	test_x2apic_intercepts(inhibit_apicv = false);
+	test_x2apic_intercepts(use_evmcs = false);
+
+	if (kvm_has_cap(KVM_CAP_HYPERV_ENLIGHTENED_VMCS))
+		test_x2apic_intercepts(use_evmcs = true);
 }
-- 
2.52.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH 00/11] KVM: fix issues with stale control fields
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (10 preceding siblings ...)
  2026-09-26  5:32 ` [PATCH 11/11] KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12 Paolo Bonzini
@ 2026-09-26  6:11 ` Paolo Bonzini
  2026-09-28 17:01   ` Sean Christopherson
  2026-09-26  6:12 ` Paolo Bonzini
  12 siblings, 1 reply; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  6:11 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: seanjc

> Fix two bugs where the guest could do stupid things on purpose to
> cause problems in the host.
>
> Patches 1-5 cover cases where actions done through VMCB control fields
> have to be redone if VMRUN fails.  In particular, failed VMRUNs can
> cause pending TLB flushes to be dropped.
>
> Patch 6 fixes a case where eVMCS execution controls can cause the
> host to use a stale MSR permission bitmap.  Patches 7-11 are tests
> for nested x2APIC; don't run them on an unpatched kernel.

In addition to what was reported by Sashiko, the test does not pass on
SVM.  Fixed as follows:

diff --git a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
index eb89d5bfc0fe..ce204ce29a9c 100644
--- a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
+++ b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
@@ -48,6 +48,7 @@ static void l1_svm_code(struct svm_test_data *svm)
 	GUEST_ASSERT_EQ(ctrl->exit_code, SVM_EXIT_CPUID);

 	stgi();
+	x2apic_write_reg(APIC_TASKPRI, 0);
 }

 static void l1_vmx_code(struct vmx_pages *vmx, struct hyperv_test_pages *hv_pages)
@@ -80,14 +81,12 @@ static void l1_vmx_code(struct vmx_pages *vmx, struct hyperv_test_pages *hv_page
 	control |= CPU_BASED_USE_MSR_BITMAPS | CPU_BASED_TPR_SHADOW;
 	GUEST_ASSERT_EQ(vmwrite(CPU_BASED_VM_EXEC_CONTROL, control), 0);

-	if (control & CPU_BASED_ACTIVATE_SECONDARY_CONTROLS) {
-		control = vmreadz(SECONDARY_VM_EXEC_CONTROL);
-		control |= SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
-			   SECONDARY_EXEC_APIC_REGISTER_VIRT |
-			   SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY;
-		control &= (rdmsr(MSR_IA32_VMX_PROCBASED_CTLS2) >> 32);
-		GUEST_ASSERT_EQ(vmwrite(SECONDARY_VM_EXEC_CONTROL, control), 0);
-	}
+	control = vmreadz(SECONDARY_VM_EXEC_CONTROL);
+	control |= SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
+		   SECONDARY_EXEC_APIC_REGISTER_VIRT |
+		   SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY;
+	control &= (rdmsr(MSR_IA32_VMX_PROCBASED_CTLS2) >> 32);
+	GUEST_ASSERT_EQ(vmwrite(SECONDARY_VM_EXEC_CONTROL, control), 0);

 	GUEST_ASSERT(!vmlaunch());
 	GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), EXIT_REASON_CPUID);
@@ -103,7 +102,7 @@ static void l1_vmx_code_part2(void)
 	control &= ~CPU_BASED_TPR_SHADOW;
 	GUEST_ASSERT_EQ(vmwrite(CPU_BASED_VM_EXEC_CONTROL, control), 0);

-	control = vmread(SECONDARY_VM_EXEC_CONTROL, &control);
+	control = vmreadz(SECONDARY_VM_EXEC_CONTROL);
 	control &= ~(SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
 			SECONDARY_EXEC_APIC_REGISTER_VIRT |
 			SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY);
@@ -154,21 +153,23 @@ static void l1_guest_code(void *test_data, void *hv_pages)

 	l1_test_x2apic_intercepts();

-	if (this_cpu_has(X86_FEATURE_VMX))
+	if (this_cpu_has(X86_FEATURE_VMX)) {
 		l1_vmx_code_part2();
-
-	l1_test_x2apic_intercepts();
+		l1_test_x2apic_intercepts();
+	}

 	GUEST_DONE();
 }

-static void __test_x2apic_intercepts(bool use_evmcs)
+static void test_x2apic_intercepts(bool with_inhibit_apicv, bool use_evmcs)
 {
 	gva_t nested_test_data_gva, hv_pages_gva = 0;
 	struct kvm_vcpu *vcpu;
 	struct kvm_vm *vm;
 	struct ucall uc;

+	inhibit_apicv = with_inhibit_apicv;
+
 	vm = vm_create_with_one_vcpu(&vcpu, l1_guest_code);
 	vm_install_exception_handler(vm, POSTED_INTR_VECTOR, guest_irq_handler);
 	vm_install_exception_handler(vm, POSTED_INTR_WAKEUP_VECTOR, guest_irq_handler);
@@ -207,29 +208,15 @@ static void __test_x2apic_intercepts(bool use_evmcs)
 	kvm_vm_free(vm);
 }

-#define _test_x2apic_intercepts(inhibit_apic_setting)	\
-do {							\
-							\
-	inhibit_apic_setting;				\
-							\
-	__test_x2apic_intercepts(use_evmcs);		\
-} while (0)
-
-#define test_x2apic_intercepts(use_evmcs_setting)	\
-do {							\
-	bool use_evmcs_setting;				\
-							\
-	_test_x2apic_intercepts(inhibit_apicv = true);	\
-	_test_x2apic_intercepts(inhibit_apicv = false);	\
-} while (0)
-
-
 int main(int argc, char *argv[])
 {
 	TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM) || kvm_cpu_has(X86_FEATURE_VMX));

-	test_x2apic_intercepts(use_evmcs = false);
+	test_x2apic_intercepts(true, false);
+	test_x2apic_intercepts(false, false);

-	if (kvm_has_cap(KVM_CAP_HYPERV_ENLIGHTENED_VMCS))
-		test_x2apic_intercepts(use_evmcs = true);
+	if (kvm_has_cap(KVM_CAP_HYPERV_ENLIGHTENED_VMCS)) {
+		test_x2apic_intercepts(true, true);
+		test_x2apic_intercepts(false, true);
+	}
 }

Paolo


^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH 00/11] KVM: fix issues with stale control fields
  2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
                   ` (11 preceding siblings ...)
  2026-09-26  6:11 ` [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
@ 2026-09-26  6:12 ` Paolo Bonzini
  12 siblings, 0 replies; 15+ messages in thread
From: Paolo Bonzini @ 2026-09-26  6:12 UTC (permalink / raw)
  To: linux-kernel, kvm; +Cc: seanjc

> Fix two bugs where the guest could do stupid things on purpose to
> cause problems in the host.
>
> Patches 1-5 cover cases where actions done through VMCB control fields
> have to be redone if VMRUN fails.  In particular, failed VMRUNs can
> cause pending TLB flushes to be dropped.
>
> Patch 6 fixes a case where eVMCS execution controls can cause the
> host to use a stale MSR permission bitmap.  Patches 7-11 are tests
> for nested x2APIC; don't run them on an unpatched kernel.

In addition to what was reported by Sashiko, the test does not pass on
SVM.  Fixed as follows:

diff --git a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
index eb89d5bfc0fe..ce204ce29a9c 100644
--- a/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
+++ b/tools/testing/selftests/kvm/x86/nested_x2apic_test.c
@@ -48,6 +48,7 @@ static void l1_svm_code(struct svm_test_data *svm)
 	GUEST_ASSERT_EQ(ctrl->exit_code, SVM_EXIT_CPUID);

 	stgi();
+	x2apic_write_reg(APIC_TASKPRI, 0);
 }

 static void l1_vmx_code(struct vmx_pages *vmx, struct hyperv_test_pages *hv_pages)
@@ -80,14 +81,12 @@ static void l1_vmx_code(struct vmx_pages *vmx, struct hyperv_test_pages *hv_page
 	control |= CPU_BASED_USE_MSR_BITMAPS | CPU_BASED_TPR_SHADOW;
 	GUEST_ASSERT_EQ(vmwrite(CPU_BASED_VM_EXEC_CONTROL, control), 0);

-	if (control & CPU_BASED_ACTIVATE_SECONDARY_CONTROLS) {
-		control = vmreadz(SECONDARY_VM_EXEC_CONTROL);
-		control |= SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
-			   SECONDARY_EXEC_APIC_REGISTER_VIRT |
-			   SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY;
-		control &= (rdmsr(MSR_IA32_VMX_PROCBASED_CTLS2) >> 32);
-		GUEST_ASSERT_EQ(vmwrite(SECONDARY_VM_EXEC_CONTROL, control), 0);
-	}
+	control = vmreadz(SECONDARY_VM_EXEC_CONTROL);
+	control |= SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
+		   SECONDARY_EXEC_APIC_REGISTER_VIRT |
+		   SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY;
+	control &= (rdmsr(MSR_IA32_VMX_PROCBASED_CTLS2) >> 32);
+	GUEST_ASSERT_EQ(vmwrite(SECONDARY_VM_EXEC_CONTROL, control), 0);

 	GUEST_ASSERT(!vmlaunch());
 	GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), EXIT_REASON_CPUID);
@@ -103,7 +102,7 @@ static void l1_vmx_code_part2(void)
 	control &= ~CPU_BASED_TPR_SHADOW;
 	GUEST_ASSERT_EQ(vmwrite(CPU_BASED_VM_EXEC_CONTROL, control), 0);

-	control = vmread(SECONDARY_VM_EXEC_CONTROL, &control);
+	control = vmreadz(SECONDARY_VM_EXEC_CONTROL);
 	control &= ~(SECONDARY_EXEC_VIRTUALIZE_X2APIC_MODE |
 			SECONDARY_EXEC_APIC_REGISTER_VIRT |
 			SECONDARY_EXEC_VIRTUAL_INTR_DELIVERY);
@@ -154,21 +153,23 @@ static void l1_guest_code(void *test_data, void *hv_pages)

 	l1_test_x2apic_intercepts();

-	if (this_cpu_has(X86_FEATURE_VMX))
+	if (this_cpu_has(X86_FEATURE_VMX)) {
 		l1_vmx_code_part2();
-
-	l1_test_x2apic_intercepts();
+		l1_test_x2apic_intercepts();
+	}

 	GUEST_DONE();
 }

-static void __test_x2apic_intercepts(bool use_evmcs)
+static void test_x2apic_intercepts(bool with_inhibit_apicv, bool use_evmcs)
 {
 	gva_t nested_test_data_gva, hv_pages_gva = 0;
 	struct kvm_vcpu *vcpu;
 	struct kvm_vm *vm;
 	struct ucall uc;

+	inhibit_apicv = with_inhibit_apicv;
+
 	vm = vm_create_with_one_vcpu(&vcpu, l1_guest_code);
 	vm_install_exception_handler(vm, POSTED_INTR_VECTOR, guest_irq_handler);
 	vm_install_exception_handler(vm, POSTED_INTR_WAKEUP_VECTOR, guest_irq_handler);
@@ -207,29 +208,15 @@ static void __test_x2apic_intercepts(bool use_evmcs)
 	kvm_vm_free(vm);
 }

-#define _test_x2apic_intercepts(inhibit_apic_setting)	\
-do {							\
-							\
-	inhibit_apic_setting;				\
-							\
-	__test_x2apic_intercepts(use_evmcs);		\
-} while (0)
-
-#define test_x2apic_intercepts(use_evmcs_setting)	\
-do {							\
-	bool use_evmcs_setting;				\
-							\
-	_test_x2apic_intercepts(inhibit_apicv = true);	\
-	_test_x2apic_intercepts(inhibit_apicv = false);	\
-} while (0)
-
-
 int main(int argc, char *argv[])
 {
 	TEST_REQUIRE(kvm_cpu_has(X86_FEATURE_SVM) || kvm_cpu_has(X86_FEATURE_VMX));

-	test_x2apic_intercepts(use_evmcs = false);
+	test_x2apic_intercepts(true, false);
+	test_x2apic_intercepts(false, false);

-	if (kvm_has_cap(KVM_CAP_HYPERV_ENLIGHTENED_VMCS))
-		test_x2apic_intercepts(use_evmcs = true);
+	if (kvm_has_cap(KVM_CAP_HYPERV_ENLIGHTENED_VMCS)) {
+		test_x2apic_intercepts(true, true);
+		test_x2apic_intercepts(false, true);
+	}
 }

Paolo


^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [PATCH 00/11] KVM: fix issues with stale control fields
  2026-09-26  6:11 ` [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
@ 2026-09-28 17:01   ` Sean Christopherson
  0 siblings, 0 replies; 15+ messages in thread
From: Sean Christopherson @ 2026-09-28 17:01 UTC (permalink / raw)
  To: Paolo Bonzini; +Cc: linux-kernel, kvm

On Sat, Sep 26, 2026, Paolo Bonzini wrote:
> > Fix two bugs where the guest could do stupid things on purpose to
> > cause problems in the host.
> >
> > Patches 1-5 cover cases where actions done through VMCB control fields
> > have to be redone if VMRUN fails.  In particular, failed VMRUNs can
> > cause pending TLB flushes to be dropped.
> >
> > Patch 6 fixes a case where eVMCS execution controls can cause the
> > host to use a stale MSR permission bitmap.  Patches 7-11 are tests
> > for nested x2APIC; don't run them on an unpatched kernel.
> 
> In addition to what was reported by Sashiko, the test does not pass on
> SVM.  Fixed as follows:

Thanks for cleaning up the mess!  I was definitely getting too greedy trying to
get bonus coverage on AMD (and I've been ignoring the SVM failures in my local
testing for an embarrasingly long time).

^ permalink raw reply	[flat|nested] 15+ messages in thread

end of thread, other threads:[~2026-09-28 17:01 UTC | newest]

Thread overview: 15+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-26  5:32 [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
2026-09-26  5:32 ` [PATCH 01/11] KVM: SVM: Preserve TLB control (i.e. pending TLB flush) on failed VMRUN Paolo Bonzini
2026-09-26  5:32 ` [PATCH 02/11] KVM: SVM: Update control fields on #VMEXIT if and only if VMRUN succeeded Paolo Bonzini
2026-09-26  5:32 ` [PATCH 03/11] KVM: SVM: Don't mark ASID fields as dirty when setting control.tlb_ctl Paolo Bonzini
2026-09-26  5:32 ` [PATCH 04/11] KVM: SVM: Sync guest's PERF_CNTR_GLOBAL_CTL from h/w only on successful VMRUN Paolo Bonzini
2026-09-26  5:32 ` [PATCH 05/11] KVM: SVM: Use the active VMCB's MSR bitmap when checking if MSR is intercepted Paolo Bonzini
2026-09-26  5:32 ` [PATCH 06/11] KVM: nVMX: Force MSR bitmap refresh if runtime eVMCS controls are modified Paolo Bonzini
2026-09-26  5:32 ` [PATCH 07/11] KVM: selftests: Add x2APIC MSR test for inhibiting APICv while nested Paolo Bonzini
2026-09-26  5:32 ` [PATCH 08/11] KVM: selftests: Run the nested x2APIC with and without APICv being inhibited in L2 Paolo Bonzini
2026-09-26  5:32 ` [PATCH 09/11] KVM: selftests: Verify that L0's TPR doesn't get clobbered Paolo Bonzini
2026-09-26  5:32 ` [PATCH 10/11] KVM: selftests: Extend nested x2APIC test to validate disabling x2APIC virt Paolo Bonzini
2026-09-26  5:32 ` [PATCH 11/11] KVM: selftests: Extend nested x2APIC test to validate using eVMCS for vmcs12 Paolo Bonzini
2026-09-26  6:11 ` [PATCH 00/11] KVM: fix issues with stale control fields Paolo Bonzini
2026-09-28 17:01   ` Sean Christopherson
2026-09-26  6:12 ` Paolo Bonzini

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®