mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Sean Christopherson <seanjc@google.com>
To: Yosry Ahmed <yosry@kernel.org>
Cc: Paolo Bonzini <pbonzini@redhat.com>,
	kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
	 Stefan Teodorescu <fane@google.com>
Subject: Re: [PATCH 3/4] KVM: x86/mmu: Bug the VM if KVM calcs a CPU role with EFER.LMA=1 && CR4.PAE=0
Date: Thu, 27 Aug 2026 07:57:38 -0700	[thread overview]
Message-ID: <apBQYmmBC8lE-KiQ@google.com> (raw)
In-Reply-To: <CAO9r8zM7KSRX5r-PazJ87QaZze3wkB5WmJL4e6az5akmJLterQ@mail.gmail.com>

On Thu, Aug 27, 2026, Yosry Ahmed wrote:
> On Wed, Aug 26, 2026 at 2:18 PM Sean Christopherson <seanjc@google.com> wrote:
> >
> > Bug the VM if KVM attempts to construct a CPU role with the should-be-
> > impossible combination of long mode being active without PAE paging being
> > enabled.  KVM's MMU construction assumes that EFER.LMA can be set if and
> > only CR4.PAE is set, and will create a completely invalid MMU if that
> > assumption fails.  FNAME(walk_addr_generic) already has sanity checks to
> > try and mitigate the fallout, but attempt to catch such bugs earlier, as
> > this is (at least) the second time KVM has had bugs that escaped into
> > FNAME(walk_addr_generic), and it's entirely possible the bad state could
> > cause problems elsewhere.
> >
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Sean Christopherson <seanjc@google.com>
> > ---
> >  arch/x86/kvm/mmu/mmu.c | 3 +++
> >  1 file changed, 3 insertions(+)
> >
> > diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
> > index 064ecc33b926..81c30e2c74f3 100644
> > --- a/arch/x86/kvm/mmu/mmu.c
> > +++ b/arch/x86/kvm/mmu/mmu.c
> > @@ -5910,6 +5910,9 @@ static union kvm_cpu_role kvm_calc_cpu_role(struct kvm_vcpu *vcpu,
> >                 return role;
> >         }
> >
> > +       if (KVM_BUG_ON(____is_efer_lma(regs) && !____is_cr4_pae(regs), vcpu->kvm))
> 
> Can we shove this into the existing if (____is_efer_lma(regs)) below?

No, because there are three more checks on EFER.LMA:

	role.ext.cr4_pke = ____is_efer_lma(regs) && ____is_cr4_pke(regs);
	role.ext.cr4_la57 = ____is_efer_lma(regs) && ____is_cr4_la57(regs);
	role.ext.efer_lma = ____is_efer_lma(regs);

and I don't want to have to condition them all on something that shouldn't happen.

OMG, I hate SVM.  I resurrected the selftest hack I used to verify this bug, to
demonstrate that Sashiko's "technically that's undefined behavior and this is
useless" complaint is wrong, because even though it's undefined behavior and the
compiler *could* ignore the change, in practice the compiler probably won't ignore
the change.  And since this is defense-in-depth, it's "fine" if the paranoid
hardening only isn't guaranteed to kick in.

And in doing so managed to trip this KVM_BUG_ON() in *L0* when running the test
in L1, because as you kinda sorta noted in patch 1, KVM doesn't ignore EFER.LMA
when loading L2 state.

I had actually tried to do exactly that, by having nested_vmcb_check_save() clear
EFER.LMA if EFER.LME=0, but that doesn't work because svm_set_nested_state() uses
the "cache" only for the checks, not for the actual loading of state.  *sigh*

So in addition to patch 1, we also need this to guard against configuring L2's
walk_mmu with bad state.

diff --git a/arch/x86/kvm/svm/nested.c b/arch/x86/kvm/svm/nested.c
index 49fb10ad1f9f..23d29597d6bf 100644
--- a/arch/x86/kvm/svm/nested.c
+++ b/arch/x86/kvm/svm/nested.c
@@ -789,6 +789,10 @@ static void nested_vmcb02_prepare_save(struct vcpu_svm *svm)
 
 	kvm_set_rflags(vcpu, save->rflags | X86_EFLAGS_FIXED);
 
+	/* SVM ignores EFER.LMA if EFER.LME=0 (instead of failing VMRUN). */
+	if (!(svm->nested.save.efer & EFER_LME))
+		svm->nested.save.efer &= ~EFER_LMA;
+
 	svm_set_efer(vcpu, svm->nested.save.efer);
 
 	svm_set_cr0(vcpu, svm->nested.save.cr0);

Anyways, back to Sashiko's "technically this is wrong" statement, I confirmed
that tweaking the code to do this does NOT trigger the KVM_BUG_ON() with at least
clang-21.  I.e. my assertion that clearing regs->efer.LMA could be useful holds
true.

diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index b515a49c5e86..cab8690d0fa0 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -5932,8 +5932,10 @@ static union kvm_cpu_role kvm_calc_cpu_role(struct kvm_vcpu *vcpu,
 		return role;
 	}
 
-	if (KVM_BUG_ON(____is_efer_lma(regs) && !____is_cr4_pae(regs), vcpu->kvm))
+	if (____is_efer_lma(regs) && !____is_cr4_pae(regs)) {
+		pr_warn("Forcing EFER.LMA=0 in calc CPU role\n");
 		*(u64 *)&regs->efer &= ~EFER_LMA;
+	}
 
 	role.base.efer_nx = ____is_efer_nx(regs);
 	role.base.cr0_wp = ____is_cr0_wp(regs);
@@ -5941,6 +5943,8 @@ static union kvm_cpu_role kvm_calc_cpu_role(struct kvm_vcpu *vcpu,
 	role.base.smap_andnot_wp = ____is_cr4_smap(regs) && !____is_cr0_wp(regs);
 	role.base.has_4_byte_gpte = !____is_cr4_pae(regs);
 
+	KVM_BUG_ON(____is_efer_lma(regs) && !____is_cr4_pae(regs), vcpu->kvm);
+
 	if (____is_efer_lma(regs))
 		role.base.level = ____is_cr4_la57(regs) ? PT64_ROOT_5LEVEL
 							: PT64_ROOT_4LEVEL;

Side topic, I also (inadvertantly) somewhat justified keeping the

	if (KVM_BUG_ON(is_long_mode(vcpu) && !is_pae(vcpu), vcpu->kvm) ||

check in FNAME(walk_addr_generic) when testing the above.  If KVM manages to
configure a sane MMU, but still has a vCPU with the above state, then we still
want to WARN and bail.

> > +               *(u64 *)&regs->efer &= ~EFER_LMA;
> 
> Why do this if we will crash the VM anyway (and Sashiko doesn't like it)?

Because there's a lot of code between here and checking KVM_VM_DEAD in
vcpu_enter_guest().  And has been proven far too many times this year, detecting
a flaw doesn't automagically mitigate true badness.

> > +
> >         role.base.efer_nx = ____is_efer_nx(regs);
> >         role.base.cr0_wp = ____is_cr0_wp(regs);
> >         role.base.cr4_smep = ____is_cr4_smep(regs);
> > --
> > 2.55.0.887.g758fc8c411-goog
> >

  reply	other threads:[~2026-08-27 14:57 UTC|newest]

Thread overview: 23+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-26 21:18 [PATCH 0/4] KVM: nSVM: Disallow bad L1 EFER for KVM_SET_NESTED_STATE Sean Christopherson
2026-08-26 21:18 ` [PATCH 1/4] KVM: nSVM: Reject KVM_SET_NESTED_STATE if L1 has EFER.LMA=1 && EFER.LME=0 Sean Christopherson
2026-08-27  7:02   ` Yosry Ahmed
2026-08-27 13:36     ` Sean Christopherson
2026-08-27 16:29       ` Yosry Ahmed
2026-08-27 17:33         ` Sean Christopherson
2026-08-27 17:55           ` Yosry Ahmed
2026-08-27 18:24             ` Sean Christopherson
2026-08-27 21:41           ` Paolo Bonzini
2026-08-27 22:01             ` Yosry Ahmed
2026-08-28  5:26               ` Paolo Bonzini
2026-08-26 21:18 ` [PATCH 2/4] KVM: x86/mmu: Bug the VM if KVM attempts to walk more levels than the MMU has Sean Christopherson
2026-08-27  7:05   ` Yosry Ahmed
2026-08-27 13:48     ` Sean Christopherson
2026-08-26 21:18 ` [PATCH 3/4] KVM: x86/mmu: Bug the VM if KVM calcs a CPU role with EFER.LMA=1 && CR4.PAE=0 Sean Christopherson
2026-08-27  7:08   ` Yosry Ahmed
2026-08-27 14:57     ` Sean Christopherson [this message]
2026-08-27 16:38       ` Yosry Ahmed
2026-08-27 17:29         ` Sean Christopherson
2026-08-27 17:48           ` Yosry Ahmed
2026-08-27 18:13             ` Sean Christopherson
2026-08-26 21:18 ` [PATCH 4/4] KVM: x86/mmu: Convert MMU walker's bounds check from BUG_ON() to KVM_BUG_ON() Sean Christopherson
2026-08-27  7:11   ` Yosry Ahmed

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=apBQYmmBC8lE-KiQ@google.com \
    --to=seanjc@google.com \
    --cc=fane@google.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=yosry@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®