From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f197.google.com (mail-pl1-f197.google.com [209.85.214.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 203DC33E37C for ; Mon, 28 Sep 2026 22:54:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.197 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790636099; cv=none; b=JCIoWA1OGOW64BY0ulHrZEOJq/VCzmoUGL/lY/ZIajznOQHHnnlSMmx87fNqAMGdVMrzWYg7+YQ9JsXAgJCXydXkscU4Vge0rycT0+xOW6DJmjJzIeuJFXlCr5rDkOldzl8saPoK0vBz4UveWlxBRXVuqVxdyJDhvG1PyB6GmOM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790636099; c=relaxed/simple; bh=m74lr8LoG/ylJgenev9YkvH/ytiCSiyudyVpiSIKlw4=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=rM5ze5LT/Jair1lYSRv63UOuRKFUqJPFOXhvmJ6CwkoF7HGUW7b708Tnhj6owb3MWAnPhs5F3o+Y1GRjeo+jN5V2gHxyTlcqpBjI0jykoI0EApLPCVh2rOn3Xgz+B5sZd0E75Te+WNdeuxlAkVLKdIpvwtAihrDxL7bF1O9NMn0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=a8nWqzJO; arc=none smtp.client-ip=209.85.214.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="a8nWqzJO" Received: by mail-pl1-f197.google.com with SMTP id d9443c01a7336-2db8c152622so23204465ad.2 for ; Mon, 28 Sep 2026 15:54:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790636097; x=1791240897; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=WHF3bdTRwFplOoq2cs8Ws3l+4hhDs1tM2IjzYePrw0I=; b=a8nWqzJOLpL2HnthUZhB2raZ7a6xkjsfEjkUcZQQ8GpiXSNy98xcn/aNFuxbPZZ6wZ UA/Gwy2XFvhZW9vmt+g2vKIhj4fs29HYPV+SDBBa+bQUEXy3RypNbBFDfEQC9W5Q9Use Y/TkoC2NtPWaO8BtebWvpofE4gLfkWwawd1UFxO7TscvdmYirll7PqcdZDlAOuoKF7Dp B6D30wQunajB1UBwX7QrHL4dHU5sNVhbno+FocwFeJxEaPMLlinEY56jcF4Ol37VD+hb mS89V3vOnNzCcNiNmRRpxQmlL4w6EnBUmqG0VuTTkDgVJZOuAZGeiORk5ShkshxlSmSb eDSQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790636097; x=1791240897; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=WHF3bdTRwFplOoq2cs8Ws3l+4hhDs1tM2IjzYePrw0I=; b=YU4qz6RneEivNcoQ3W6TYuP/Y5DxeIE+BEy8MtgFooIiUkocPSSMKt2/5kDgKbeE4p UcQUsCeXrAsoaoJyvWCjuYHd3c4hWxHWbmvdwfzD2A36KwKLD59Qh1oM8xw+Debg/p92 sspaYz8e4rQcjV1vfKJNa6ENxVRW6UQvGn8Q3J3LFk7ZtWZfIwWhAz5kqn8Pc+xOVIQM lm+ybO0KfK1XPNVZhd5YdagOxNFZTJZyhjkahIpjDW1WdmtcJ8MM2Jvsw4LgorHWr0tK PvgRrp8FxY0c3BlUnHxKCysr+YZ8RwagJcLiQ4R+J4GV8oh+4Ghr+WJvG9Y92GJiiZ7p 2nqw== X-Forwarded-Encrypted: i=1; AKwUvByLTmtRliOnpRrage04FDDE12K48QdCDlcigIwHxh7YS+SfSmchk+PW4FX2htlt1nwtkMV8/ebVRQWqDds=@vger.kernel.org X-Gm-Message-State: AFuF++mystSo3zTqMd84lZYHSd7PNZ849jstUAaE0asfauGwmdO942El 9ouSGXJiTzV/3GW1oLnYLgU9R3yYGnTVAxAanRsIQTY1xHlpoT9uB9WwwTbxLtUGUdQB4hrX7pk OSqJlxA== X-Received: from pldm6.prod.google.com ([2002:a17:902:db86:b0:2dd:63c:daa3]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:8c5:b0:2df:8b09:93c2 with SMTP id d9443c01a7336-2df8b09982fmr96875675ad.43.1790636097191; Mon, 28 Sep 2026 15:54:57 -0700 (PDT) Date: Mon, 28 Sep 2026 15:54:56 -0700 In-Reply-To: <20260820123512.87236-1-duankeqiangcym@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260820123512.87236-1-duankeqiangcym@gmail.com> Message-ID: Subject: Re: [PATCH v2] KVM: x86: Clear hardware HLT state when userspace makes a vCPU not-halted From: Sean Christopherson To: Keqiang Duan Cc: pbonzini@redhat.com, kvm@vger.kernel.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org, Qinguang Chen , Zhiping Du Content-Type: text/plain; charset="us-ascii" On Thu, Aug 20, 2026, Keqiang Duan wrote: > diff --git a/arch/x86/kvm/vmx/x86_ops.h b/arch/x86/kvm/vmx/x86_ops.h > index cdb38d940cfb..45e47f502a37 100644 > --- a/arch/x86/kvm/vmx/x86_ops.h > +++ b/arch/x86/kvm/vmx/x86_ops.h > @@ -95,6 +95,7 @@ int vmx_interrupt_allowed(struct kvm_vcpu *vcpu, bool for_injection); > int vmx_nmi_allowed(struct kvm_vcpu *vcpu, bool for_injection); > bool vmx_get_nmi_mask(struct kvm_vcpu *vcpu); > void vmx_set_nmi_mask(struct kvm_vcpu *vcpu, bool masked); > +void vmx_clear_hlt(struct kvm_vcpu *vcpu); > void vmx_enable_nmi_window(struct kvm_vcpu *vcpu); > void vmx_enable_irq_window(struct kvm_vcpu *vcpu); > void vmx_update_cr8_intercept(struct kvm_vcpu *vcpu, int tpr, int irr); > diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c > index d94b59140c45..3b224c3bbe42 100644 > --- a/arch/x86/kvm/x86.c > +++ b/arch/x86/kvm/x86.c > @@ -9058,6 +9058,21 @@ int kvm_arch_vcpu_ioctl_set_mpstate(struct kvm_vcpu *vcpu, > } > > kvm_set_mp_state(vcpu, mp_state->mp_state); > + > + /* > + * Force the vCPU out of any hardware-tracked halted state, e.g. VMX's > + * GUEST_ACTIVITY_STATE=HLT, when userspace puts the vCPU into a state > + * other than HALTED. The hardware state is sticky across VM-Exit and > + * VM-Enter and is not touched by any other ioctl, so a vCPU that halted > + * with HLT-exiting disabled stays wedged even after userspace rewrites > + * its registers, e.g. when a VMM emulates a machine reset. Waking from > + * HLT is architecturally allowed to be spurious, so clearing it is Everything looks good except this claim that spurious wakeups is architecturally allowed. I'm 99% certain that is straight up wrong. The SDM explicitly states what will break HLT: An enabled interrupt (including NMI and SMI), a debug exception, the BINIT# signal, the INIT# signal, or the RESET# signal will resume execution. and the APM goes a step further, and in addition to listing the wake events: Execution resumes when an unmasked hardware interrupt (INTR), non-maskable interrupt (NMI), system management interrupt (SMI), RESET, or INIT occurs. very clearly states that doing HLT with RFLAGS.IF=0 means: If rFLAGS.IF = 0, the system will remain in a HALT state until an NMI, SMI, RESET, or INIT occurs. AFAIK, nothing in either the SDM or APM suggests spurious HLT wakeups are allowed. And FWIW, this is not a theoretical issue. A few years back we had a customer issue where a spurious HLT wakeup due to a KVM bug crashed the guest (IIRC, the guest offlined CPUs and put them in HLT, then kexec'd into a new kernel which unmapped the code containing the HLT loop). Anyways, unless someone cares enough to want to back up the claim that spurious wakeups are ok, I'll just drop that line when applying. I don't see any reason to mention spurious wakes: the vCPU is clearly being moved out of HALTED state, it's on userspace not to screw up (for this particular case; there are other live migration issues that userspace can't solve). > + * always safe. > + */ > + if (kvm_hlt_in_guest(vcpu->kvm) && > + mp_state->mp_state != KVM_MP_STATE_HALTED) > + kvm_x86_call(clear_hlt)(vcpu); > + > kvm_make_request(KVM_REQ_EVENT, vcpu); > > ret = 0; > > base-commit: 1b731e5ded480bd1e5546aed35584238661ce72e > -- > 2.24.3 > >