* [PATCH v2 0/2] KVM: nVMX: fix nested APICv emulation for windows guests
@ 2026-10-11 0:13 Maxim Levitsky
2026-10-11 0:13 ` [PATCH v2 1/2] KVM: nVMX: don't check PIR.ON when processing nested posted interrupts Maxim Levitsky
2026-10-11 0:13 ` [PATCH v2 2/2] KVM: nVMX: avoid losing the posted notification interrupt when exiting L2 Maxim Levitsky
0 siblings, 2 replies; 3+ messages in thread
From: Maxim Levitsky @ 2026-10-11 0:13 UTC (permalink / raw)
To: kvm; +Cc: Sean Christopherson, x86, linux-kernel, Paolo Bonzini, Maxim Levitsky
Fix two problems that were recently discovered that cause nested APICv
to malfunction on windows with hyperv enabled.
The first problem is that unlike KVM, windows doesn't set PIR.ON when
sending posted interrupts. X86 spec actually allows this.
The second problem is that unlilke KVM, which mostly ignores them,
windows depends on accurate delivery of the posted notification
interrupts, and currently it is possible for a posted notification
to get lost it it arrives at the same time as the nested VM exit.
V2: to avoid spurious notification in L1, process the posted
notification in L2, just before VM exit.
PS: I am not 100% sure if I got the barriers right, plus I am not sure
about what to do if vmx_complete_nested_posted_interrupt can't write
to the nested PIR, in this it calls kvm_handle_memory_failure, and
I don't know 100% its interactions with nested VM exit.
Best regards,
Maxim Levitsky
Maxim Levitsky (2):
KVM: nVMX: don't check PIR.ON when processing nested posted interrupts
KVM: nVMX: avoid losing the posted notification interrupt when exiting
L2
arch/x86/kvm/vmx/nested.c | 22 ++++++++++++++++++----
arch/x86/kvm/vmx/vmx.c | 5 +++++
2 files changed, 23 insertions(+), 4 deletions(-)
--
2.54.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v2 1/2] KVM: nVMX: don't check PIR.ON when processing nested posted interrupts
2026-10-11 0:13 [PATCH v2 0/2] KVM: nVMX: fix nested APICv emulation for windows guests Maxim Levitsky
@ 2026-10-11 0:13 ` Maxim Levitsky
2026-10-11 0:13 ` [PATCH v2 2/2] KVM: nVMX: avoid losing the posted notification interrupt when exiting L2 Maxim Levitsky
1 sibling, 0 replies; 3+ messages in thread
From: Maxim Levitsky @ 2026-10-11 0:13 UTC (permalink / raw)
To: kvm; +Cc: Sean Christopherson, x86, linux-kernel, Paolo Bonzini, Maxim Levitsky
Despite the suggestion to test and set the PIR.ON and avoid sending
a posted interrupt if it is already set, there is no requirement
for this bit to be set for posted interrupt processing to happen
when the target CPU receives the posted notification vector.
See section 30.6 POSTED-INTERRUPT PROCESSING:
"The processor clears the outstanding-notification bit in
the posted-interrupt descriptor. This is done atomically so as to leave
the remainder of the descriptor unmodified
(e.g., with a locked AND operation)."
Apparently Windows' implementation of APICv does exactly this:
it sets bits in PIR, without bothering to also set the PIR.ON.
Signed-off-by: Maxim Levitsky <mlevitsk@redhat.com>
---
arch/x86/kvm/vmx/nested.c | 15 +++++++++++----
1 file changed, 11 insertions(+), 4 deletions(-)
diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
index 151873407abd..8504a2c12d9d 100644
--- a/arch/x86/kvm/vmx/nested.c
+++ b/arch/x86/kvm/vmx/nested.c
@@ -4041,8 +4041,16 @@ static int vmx_complete_nested_posted_interrupt(struct kvm_vcpu *vcpu)
vmx->nested.pi_pending = false;
- if (!pi_test_and_clear_on(vmx->nested.pi_desc))
- return 0;
+ /*
+ * Don't test the value of PID.ON.
+ * It is valid to trigger a posted interrupt without setting this bit
+ */
+ pi_clear_on(vmx->nested.pi_desc);
+
+ /* Ensure that guests sees the change to PIR.ON before KVM harvest
+ * the PIR
+ */
+ __smp_mb__after_atomic();
max_irr = pi_find_highest_vector(vmx->nested.pi_desc);
if (max_irr > 0) {
@@ -4209,8 +4217,7 @@ static bool vmx_has_nested_events(struct kvm_vcpu *vcpu, bool for_injection)
if ((max_irr & 0xf0) > (vppr & 0xf0))
return true;
- if (vmx->nested.pi_pending && vmx->nested.pi_desc &&
- pi_test_on(vmx->nested.pi_desc)) {
+ if (vmx->nested.pi_pending && vmx->nested.pi_desc) {
max_irr = pi_find_highest_vector(vmx->nested.pi_desc);
if (max_irr > 0 && (max_irr & 0xf0) > (vppr & 0xf0))
return true;
--
2.54.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v2 2/2] KVM: nVMX: avoid losing the posted notification interrupt when exiting L2
2026-10-11 0:13 [PATCH v2 0/2] KVM: nVMX: fix nested APICv emulation for windows guests Maxim Levitsky
2026-10-11 0:13 ` [PATCH v2 1/2] KVM: nVMX: don't check PIR.ON when processing nested posted interrupts Maxim Levitsky
@ 2026-10-11 0:13 ` Maxim Levitsky
1 sibling, 0 replies; 3+ messages in thread
From: Maxim Levitsky @ 2026-10-11 0:13 UTC (permalink / raw)
To: kvm; +Cc: Sean Christopherson, x86, linux-kernel, Paolo Bonzini, Maxim Levitsky
While delivering a nested posted interrupt notification
(see vmx_deliver_nested_posted_interrupt), KVM assumes that,
as long as the target vCPU is in the guest mode, KVM can send the
special POSTED_INTR_NESTED_VECTOR, which will either trigger APICv ucode
to inject all interrupts into L2 or cause a VM exit, after which
vmx_complete_nested_posted_interrupt is supposed to finish the job.
However if the target vCPU was about to exit the L2 when
nested posted interrupt notification was delivered in this way,
it is possible that it will be lost.
Fix this by running vmx_complete_nested_posted_interrupt just before
nested VM exit.
Detect this case in the nested VM exit path and act accordingly.
Signed-off-by: Maxim Levitsky <mlevitsk@redhat.com>
---
arch/x86/kvm/vmx/nested.c | 7 +++++++
arch/x86/kvm/vmx/vmx.c | 5 +++++
2 files changed, 12 insertions(+)
diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
index 8504a2c12d9d..1b2f55a13993 100644
--- a/arch/x86/kvm/vmx/nested.c
+++ b/arch/x86/kvm/vmx/nested.c
@@ -5114,8 +5114,15 @@ void __nested_vmx_vmexit(struct kvm_vcpu *vcpu, u32 vm_exit_reason,
if (enable_ept && is_pae_paging(vcpu))
vmx_ept_load_pdptrs(vcpu);
+
leave_guest_mode(vcpu);
+ /* pairs with barrier in vmx_deliver_nested_posted_interrupt */
+ smp_wmb();
+
+ if (vmx->nested.pi_pending)
+ vmx_complete_nested_posted_interrupt(vcpu);
+
if (nested_cpu_has_preemption_timer(vmcs12))
hrtimer_cancel(&to_vmx(vcpu)->nested.preemption_timer);
diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
index 612ab07d4100..21f2f8e4c4f0 100644
--- a/arch/x86/kvm/vmx/vmx.c
+++ b/arch/x86/kvm/vmx/vmx.c
@@ -4381,10 +4381,15 @@ static int vmx_deliver_nested_posted_interrupt(struct kvm_vcpu *vcpu,
*/
if (is_guest_mode(vcpu) &&
vector == vmx->nested.posted_intr_nv) {
+
+ /* pairs with barrier in __nested_vmx_vmexit */
+ smp_rmb();
+
/*
* If a posted intr is not recognized by hardware,
* we will accomplish it in the next vmentry.
*/
+
vmx->nested.pi_pending = true;
kvm_make_request(KVM_REQ_EVENT, vcpu);
--
2.54.0
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-11 0:13 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-11 0:13 [PATCH v2 0/2] KVM: nVMX: fix nested APICv emulation for windows guests Maxim Levitsky
2026-10-11 0:13 ` [PATCH v2 1/2] KVM: nVMX: don't check PIR.ON when processing nested posted interrupts Maxim Levitsky
2026-10-11 0:13 ` [PATCH v2 2/2] KVM: nVMX: avoid losing the posted notification interrupt when exiting L2 Maxim Levitsky
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®