From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f202.google.com (mail-pf1-f202.google.com [209.85.210.202]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 592354D8DB1 for ; Wed, 3 Jun 2026 22:34:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.202 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780526067; cv=none; b=i+tQS7d2J3g22v/cOaWditaQbr7idS2UI0Lfu1Mfs0KpWsAU5+eYtMagKEkfiB+95ww6xZPxh87yiIA6rK65vabFwQVOYV9tt63TXPH+H3XHDl/KZmnD0lsxUOfn7MyMsbQfbvQykDT/egTTC2IxISF8+1sxRA2d8mkzLL0VbwY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780526067; c=relaxed/simple; bh=KtdlDyfbm3TdN3Bu5jvCzKeLtKQzA5ao31bl8Pxom84=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=srhAQBXN19bg0FPsy1x4+2h7dSIqdoLOaJT+HRNyFbjyMXYlsMTWdCQ5+GltwI/PsbD29abzCRW4dI0RYJNjTkBV8nI5IP7Qknoyl4k5CnGgetHkaFQc+chukRzrYNwWSI8j9c91vdw/vDuVy4dxLJVNJHdnPhaaX8O2Y71Oc+E= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=tyvlDC+Y; arc=none smtp.client-ip=209.85.210.202 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="tyvlDC+Y" Received: by mail-pf1-f202.google.com with SMTP id d2e1a72fcca58-8421ffff8a3so58958b3a.2 for ; Wed, 03 Jun 2026 15:34:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1780526063; x=1781130863; darn=vger.kernel.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:reply-to:from:to:cc:subject:date:message-id:reply-to; bh=YT22xFPe+uvBINsx4QHwFFDf4DhD63vVETjx/ZAWt2E=; b=tyvlDC+Y99qjHpEnqftZV/RRwHqb8Q9tEauYZvCGXZzV17cGV7f0HOatReGEsRGUNm gEZxmwcsxQyhnwKWsBgQRO1iBk4RkEV4KW/MpGu5Srolxl/5huAkzLkhgU7MNIIAgcXL OCmzamR2LHvzmdxkwoErfXIT3IAgvSOzOdb0rGdncLiUIeju5/IVcbLue1022R7B3HnW 599Z5fsRgqr36q74aD0Vl6xq8JcmkWmfbjHFaUAVXBMx0vPo7HmWK/V67vi1g3mRn9j3 o5TybdGV9snAgIlQostN4TYsTPQhKBNvWrC7DUAAUyq2yfKSGHD5bSHk0Xa6m0fZxVTy hvKA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1780526063; x=1781130863; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:reply-to:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=YT22xFPe+uvBINsx4QHwFFDf4DhD63vVETjx/ZAWt2E=; b=JpjEG7a3HsMO7Vgay8WEUtoj8j38QLCF8qYvmtKEqSPOTQAiVvgRAjYQyb+A6NFgmX t80DAzWt74MdZmPg9MUZV8ENw2+ekDH3oOjx/2NlsAIl7yauDZUxCJ/6qZePTwPOPHcK Hups3OckHjJHQCx46RLNpNVRjbA2uzBiJifQ0y+4PJIAqG16g2vfv1B6L7W4m/Q7dx5n 2jKBIi6Hhw1VjRlmG4M0F+2C2ZDzepsuc+qXQmyW83MPTF8X7bVJ3YHYvcj48OyoYc0W 45aJQqUWBMR5O6iTjsG7LfsZDe5IJEBQlrSr+NOHMB0KEnxeQYh8jVySQjkOqC6/lOpr xW8w== X-Forwarded-Encrypted: i=1; AFNElJ9xSWf3lSWWG5JRC27G52Z/zyPbBLGFyBWZbWvF8JvIIr/PvfW/ehX9ZQukJukAsoXvlsTAOiT/B5y/3C0=@vger.kernel.org X-Gm-Message-State: AOJu0YxDDL7ELgtJDdkYnqDWPpMLPDocYX8W12l1Rsx4L1WMNbdxLgjX eTlC2s00y3Uns97aBEPU2JWIk0Z41FbdfwYoUe0Aew2YozbwydRdDrFwB0g0APtjMt+cVJVSE5n 12X9mwg== X-Received: from pfm4.prod.google.com ([2002:a05:6a00:724:b0:842:2c74:b8ce]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:b483:b0:842:2ddb:e303 with SMTP id d2e1a72fcca58-84284da5a08mr5238397b3a.12.1780526063099; Wed, 03 Jun 2026 15:34:23 -0700 (PDT) Reply-To: Sean Christopherson Date: Wed, 3 Jun 2026 15:34:18 -0700 In-Reply-To: <20260603223418.1720035-1-seanjc@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260603223418.1720035-1-seanjc@google.com> X-Mailer: git-send-email 2.54.0.1032.g2f8565e1d1-goog Message-ID: <20260603223418.1720035-3-seanjc@google.com> Subject: [PATCH 2/2] KVM: nVMX: Don't use vmcs01.GUEST_CR3 to snapshot L1's CR3 when EPT is disabled From: Sean Christopherson To: Sean Christopherson , Paolo Bonzini Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org, Jim Mattson Content-Type: text/plain; charset="UTF-8" Add a dedicated field in "struct nested_vmx" to track L1's pre-VM-Enter CR3 instead of using vmcs01.GUEST_CR3, which isn't anywhere near as safe as the comment purports it to be. E.g. in addition to the warn_on_missed_cc bug (that was fixed by relocating the consistency check), if getting vmcs12 pages (during actual nested VM-Entry) fails and EPT is disabled (in KVM), KVM will return control to userspace with vmcs01.GUEST_CR3 holding a guest- controlled value. Alternatively, KVM could force a reload of vmcs01.GUEST_CR3 by resetting the MMU context in the error path, but as above, the safety of the vmcs01 approach is extremely questionable, e.g. it took all of ~4 months for the code to break. Fixes: 671ddc700fd0 ("KVM: nVMX: Don't leak L1 MMIO regions to L2") Cc: stable@vger.kernel.org Cc: Jim Mattson Signed-off-by: Sean Christopherson --- arch/x86/kvm/vmx/nested.c | 21 ++++++++------------- arch/x86/kvm/vmx/vmx.h | 7 +++++++ 2 files changed, 15 insertions(+), 13 deletions(-) diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c index 039e234e7d2b..772b8090d06a 100644 --- a/arch/x86/kvm/vmx/nested.c +++ b/arch/x86/kvm/vmx/nested.c @@ -3668,19 +3668,14 @@ enum nvmx_vmentry_status nested_vmx_enter_non_root_mode(struct kvm_vcpu *vcpu, &vmx->nested.pre_vmenter_ssp_tbl); /* - * Overwrite vmcs01.GUEST_CR3 with L1's CR3 if EPT is disabled. In the - * event of a "late" VM-Fail, i.e. a VM-Fail detected by hardware but - * not KVM, KVM must unwind its software model to the pre-VM-Entry host - * state. When EPT is disabled, GUEST_CR3 holds KVM's shadow CR3, not - * L1's "real" CR3, which causes nested_vmx_restore_host_state() to - * corrupt vcpu->arch.cr3. Stuffing vmcs01.GUEST_CR3 results in the - * unwind naturally setting arch.cr3 to the correct value. Smashing - * vmcs01.GUEST_CR3 is safe because nested VM-Exits, and the unwind, - * reset KVM's MMU, i.e. vmcs01.GUEST_CR3 is guaranteed to be - * overwritten with a shadow CR3 prior to re-entering L1. + * Stash L1's CR3, so that in the event of a "late" VM-Fail, i.e. a + * VM-Fail detected by hardware but not KVM, KVM can unwind its + * software model to the pre-VM-Entry host state. When EPT is + * disabled, GUEST_CR3 holds KVM's shadow CR3, not L1's "real" CR3, + * and so simply restoring from vmcs01.GUEST_CR3 would corrupt + * vcpu->arch.cr3. */ - if (!enable_ept) - vmcs_writel(GUEST_CR3, vcpu->arch.cr3); + vmx->nested.pre_vmenter_cr3 = vcpu->arch.cr3; vmx_switch_vmcs(vcpu, &vmx->nested.vmcs02); @@ -4992,7 +4987,7 @@ static void nested_vmx_restore_host_state(struct kvm_vcpu *vcpu) vmx_set_cr4(vcpu, vmcs_readl(CR4_READ_SHADOW)); nested_ept_uninit_mmu_context(vcpu); - vcpu->arch.cr3 = vmcs_readl(GUEST_CR3); + vcpu->arch.cr3 = vmx->nested.pre_vmenter_cr3; kvm_register_mark_available(vcpu, VCPU_REG_CR3); /* diff --git a/arch/x86/kvm/vmx/vmx.h b/arch/x86/kvm/vmx/vmx.h index de9de0d2016c..dc8517f15bc4 100644 --- a/arch/x86/kvm/vmx/vmx.h +++ b/arch/x86/kvm/vmx/vmx.h @@ -159,6 +159,13 @@ struct nested_vmx { bool has_preemption_timer_deadline; bool preemption_timer_expired; + /* + * Used to restore L1's CR3 if hardware detects a VM-Fail Consistency + * Check that KVM does not, in which case KVM needs to unwind CR3 back + * to its pre-VM-Enter state, NOT to vmcs01.HOST_CR3. + */ + unsigned long pre_vmenter_cr3; + /* * Used to snapshot MSRs that are conditionally loaded on VM-Enter in * order to propagate the guest's pre-VM-Enter value into vmcs02. For -- 2.54.0.1032.g2f8565e1d1-goog