From: Sean Christopherson <seanjc@google.com>
To: Keqiang Duan <duankeqiangcym@gmail.com>
Cc: pbonzini@redhat.com, kvm@vger.kernel.org,
linux-kernel@vger.kernel.org, stable@vger.kernel.org,
Qinguang Chen <chenqinguang@kuaishou.com>,
Zhiping Du <duzhiping@kuaishou.com>
Subject: Re: [PATCH v2] KVM: x86: Clear hardware HLT state when userspace makes a vCPU not-halted
Date: Mon, 28 Sep 2026 15:54:56 -0700 [thread overview]
Message-ID: <arrwQKQJpFZtARmN@google.com> (raw)
In-Reply-To: <20260820123512.87236-1-duankeqiangcym@gmail.com>
On Thu, Aug 20, 2026, Keqiang Duan wrote:
> diff --git a/arch/x86/kvm/vmx/x86_ops.h b/arch/x86/kvm/vmx/x86_ops.h
> index cdb38d940cfb..45e47f502a37 100644
> --- a/arch/x86/kvm/vmx/x86_ops.h
> +++ b/arch/x86/kvm/vmx/x86_ops.h
> @@ -95,6 +95,7 @@ int vmx_interrupt_allowed(struct kvm_vcpu *vcpu, bool for_injection);
> int vmx_nmi_allowed(struct kvm_vcpu *vcpu, bool for_injection);
> bool vmx_get_nmi_mask(struct kvm_vcpu *vcpu);
> void vmx_set_nmi_mask(struct kvm_vcpu *vcpu, bool masked);
> +void vmx_clear_hlt(struct kvm_vcpu *vcpu);
> void vmx_enable_nmi_window(struct kvm_vcpu *vcpu);
> void vmx_enable_irq_window(struct kvm_vcpu *vcpu);
> void vmx_update_cr8_intercept(struct kvm_vcpu *vcpu, int tpr, int irr);
> diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
> index d94b59140c45..3b224c3bbe42 100644
> --- a/arch/x86/kvm/x86.c
> +++ b/arch/x86/kvm/x86.c
> @@ -9058,6 +9058,21 @@ int kvm_arch_vcpu_ioctl_set_mpstate(struct kvm_vcpu *vcpu,
> }
>
> kvm_set_mp_state(vcpu, mp_state->mp_state);
> +
> + /*
> + * Force the vCPU out of any hardware-tracked halted state, e.g. VMX's
> + * GUEST_ACTIVITY_STATE=HLT, when userspace puts the vCPU into a state
> + * other than HALTED. The hardware state is sticky across VM-Exit and
> + * VM-Enter and is not touched by any other ioctl, so a vCPU that halted
> + * with HLT-exiting disabled stays wedged even after userspace rewrites
> + * its registers, e.g. when a VMM emulates a machine reset. Waking from
> + * HLT is architecturally allowed to be spurious, so clearing it is
Everything looks good except this claim that spurious wakeups is architecturally
allowed. I'm 99% certain that is straight up wrong. The SDM explicitly states
what will break HLT:
An enabled interrupt (including NMI and SMI), a debug exception, the BINIT#
signal, the INIT# signal, or the RESET# signal will resume execution.
and the APM goes a step further, and in addition to listing the wake events:
Execution resumes when an unmasked hardware interrupt (INTR), non-maskable
interrupt (NMI), system management interrupt (SMI), RESET, or INIT occurs.
very clearly states that doing HLT with RFLAGS.IF=0 means:
If rFLAGS.IF = 0, the system will remain in a HALT state until an NMI, SMI,
RESET, or INIT occurs.
AFAIK, nothing in either the SDM or APM suggests spurious HLT wakeups are allowed.
And FWIW, this is not a theoretical issue. A few years back we had a customer
issue where a spurious HLT wakeup due to a KVM bug crashed the guest (IIRC, the
guest offlined CPUs and put them in HLT, then kexec'd into a new kernel which
unmapped the code containing the HLT loop).
Anyways, unless someone cares enough to want to back up the claim that spurious
wakeups are ok, I'll just drop that line when applying. I don't see any reason
to mention spurious wakes: the vCPU is clearly being moved out of HALTED state,
it's on userspace not to screw up (for this particular case; there are other live
migration issues that userspace can't solve).
> + * always safe.
> + */
> + if (kvm_hlt_in_guest(vcpu->kvm) &&
> + mp_state->mp_state != KVM_MP_STATE_HALTED)
> + kvm_x86_call(clear_hlt)(vcpu);
> +
> kvm_make_request(KVM_REQ_EVENT, vcpu);
>
> ret = 0;
>
> base-commit: 1b731e5ded480bd1e5546aed35584238661ce72e
> --
> 2.24.3
>
>
next prev parent reply other threads:[~2026-09-28 22:54 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-20 12:35 Keqiang Duan
2026-08-20 12:44 ` keqiang duan
2026-09-09 8:00 ` keqiang duan
2026-09-15 2:24 ` Chao Gao
2026-09-28 22:54 ` Sean Christopherson [this message]
2026-09-29 2:46 ` Keqiang Duan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arrwQKQJpFZtARmN@google.com \
--to=seanjc@google.com \
--cc=chenqinguang@kuaishou.com \
--cc=duankeqiangcym@gmail.com \
--cc=duzhiping@kuaishou.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=pbonzini@redhat.com \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®