From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C43743BE144 for ; Tue, 2 Jun 2026 14:24:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780410272; cv=none; b=iO4qgE/7Fs82l9eH+DYLiWt2Yuwb5D/M7ek5vbfUPCMD6E57OPrXcVZzTuRXEix7m5IjKgzG1Pt1XMO6YxF9uUFd1hMS2rK48j8ChPf6IKlEKGjdY45bju7ZNc7boECI1Os6yIs+EyvqMKK5Kb8BPZQRXgxq4EgKTfW3pPPc+VI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780410272; c=relaxed/simple; bh=yHWawHcRNy+Mlf/3GbgmyH5P0mKzD9qwHAo72JxYNeY=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=JHWYHtOQxFe5ZWwrlqPIU5lb9J89JiAMYQ5jHYLURHDIDVQgRfqBW06/sK3xnELGWGIx/cv2cGEWFghfNSf4BwYY3pVQjnlyNJtfZxWeZaAGCo2Jt5ZQEhXMn7UIRAkIvb63Xa2p2FGGpI5ungFiIPpFKAVnv5G2Jw6dalypFy4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Dbtr3ac3; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=PFyl721d; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Dbtr3ac3"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="PFyl721d" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1780410270; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=K6qeqqoJSoBMzW4sniDDXZPSZAPEaf893MDZNg+z0GY=; b=Dbtr3ac3z5XVuWN1D3XifLFXftNPj6Qc0+uyd/KmW+TpsbSxkgOyN1m5fNNnZMHw6q9kHx /TGOAZzF2P7AcmuF4WQ7nNxDmdWk0CRiKrNlXKuvUG1tMDi9nHn6WX6jScmJ2SljcgNhZt cGcesty48HwXxp51k8NVqhCX63z3a+o= Received: from mail-qt1-f197.google.com (mail-qt1-f197.google.com [209.85.160.197]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-376-9XtT-NkAPe63XIa8M9aQNA-1; Tue, 02 Jun 2026 10:24:26 -0400 X-MC-Unique: 9XtT-NkAPe63XIa8M9aQNA-1 X-Mimecast-MFC-AGG-ID: 9XtT-NkAPe63XIa8M9aQNA_1780410266 Received: by mail-qt1-f197.google.com with SMTP id d75a77b69052e-514551d5f2aso57666851cf.2 for ; Tue, 02 Jun 2026 07:24:26 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1780410266; x=1781015066; darn=vger.kernel.org; h=mime-version:user-agent:content-transfer-encoding:references :in-reply-to:date:cc:to:from:subject:message-id:from:to:cc:subject :date:message-id:reply-to; bh=K6qeqqoJSoBMzW4sniDDXZPSZAPEaf893MDZNg+z0GY=; b=PFyl721d+XHUMKcyZIUPIQuMt0wJc/KqKSYz7tqUokp4a4J4HKN0Fs+oPd5YWf58V5 m7dq4Y/v0iEeL493/G3OAyBHrJ8brjMabHIAu/j0veurV/mxFvDMkxV7AjC2vpRpFZ9A MP1FcwrLPM1nQ9JBwhK93DQd/G5B5mzE1CVG/o4CIKahcFc6xYu8Q5IjKfy830mJvM6n gX/+Y25oWUawrlyFDRJYkCP4B9zWzMPW7LAvkGYCJixZ/Oyy4D9CrDM30by4NGDVGHt7 8gbltB77zb8ZlBNVwFCJyjm0f3T8w43ENVKh+u9nrQAe7fYG3ldO1ENDeEDdT7KT7HLY ePwA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1780410266; x=1781015066; h=mime-version:user-agent:content-transfer-encoding:references :in-reply-to:date:cc:to:from:subject:message-id:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=K6qeqqoJSoBMzW4sniDDXZPSZAPEaf893MDZNg+z0GY=; b=Hqdm/zAUfzg0DbXzX04xpxc5+sHc86vlmTMKPf2LAVhaxb/i6llrlZNl+J359JbJ8v rfus0rdy4wpkIcrSqCZfn5jAoUsooJuBh6We7CyJIJMtAVY59QvrSfv5rS1kMn+nf2Bl YDieDDsTc9P0qXRc7YnAxbA2lx55TKW/DdxDYf5M3dGnstpYM83xfO8qHVS36rENsWYZ JLcnKETfrqf59cVyjwt/mHq7BkePbfjpJEbn7N4P64DL4y/AZXddI+ydo3SF8WSXa5S3 rRpL74en1bJsbnLrQQKWZ/Evw4PE5LPXJIkYLfQpOHwvE0q/zFHvyVu9kKY7bzSAFTSL a1aA== X-Forwarded-Encrypted: i=1; AFNElJ9ExmpAUepqOeSYehC13X9BZKfO/SGdrVPnghajLorQenkfl5NgLLoNryE6EM54cOmPPX8Qj5cOIAaWW2o=@vger.kernel.org X-Gm-Message-State: AOJu0Yx/74R7ipajcTuzksOtagiKXQ3Xk/IX5hp0MZcX5AFr/0x3z42D gSMuOc7/WSJP37UZgNbFpS3NVTL4yt8VtId1q1GMv8XW3qyHe93zPumzqJiXABQ+zyoazS5adXp kHcPUNeTJSYp7xQ87ke6dpQrzBtJ8xbVNYMJths1+rzfXJ35+vB7pSs4qioy8at4WqQ== X-Gm-Gg: Acq92OGpzszeFkWejWniJT2UIApGdZ8w1jIqF6oawxh668kPvntb9AbqBqn7M7GzCJo By2+oW0l6+/w3814tjNf7OkrlcZ9+KBB3QERGmYAA8rux5ptUhg3t8o918KnLBmfSAV1Qpynid5 CvNDbbCCWVhCgFSQtxOC53FgMu0pFpGGKTOyUeMlh3SN98PTdg9PzXR5Ci2oBLaDYQOBT61mKhR V89qm/7eWGkFzB3+LwhwydZN45c2G1iEVuIvq1v7yPyC+U5vYXaxg+8dPUu5KqGXZMNdLlDN2Uo YiFEVxOYrnTLfjue59a/s+buO8I2FkPP/CnRbrniQU7waa7BBSPt7541lKihX+Y8f76a8sKVBqM D+qP22AcpYQuujm6YfZovjuIFrIS75+1GArHy1P8= X-Received: by 2002:a05:622a:8d09:b0:516:d678:5337 with SMTP id d75a77b69052e-5173a821e2bmr199534171cf.28.1780410266013; Tue, 02 Jun 2026 07:24:26 -0700 (PDT) X-Received: by 2002:a05:622a:8d09:b0:516:d678:5337 with SMTP id d75a77b69052e-5173a821e2bmr199533631cf.28.1780410265401; Tue, 02 Jun 2026 07:24:25 -0700 (PDT) Received: from intellaptop.lan ([2607:fea8:fc01:88aa:f1de:f35:7935:804f]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-51757c02ba9sm43829141cf.22.2026.06.02.07.24.23 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 02 Jun 2026 07:24:24 -0700 (PDT) Message-ID: <9a2b9700a835798fdc090035936be3e8680f3566.camel@redhat.com> Subject: Re: [PATCH 12/28] KVM: x86: make translate_nested_gpa vendor-specific From: mlevitsk@redhat.com To: Paolo Bonzini , linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: d.riley@proxmox.com, jon@nutanix.com Date: Tue, 02 Jun 2026 10:24:23 -0400 In-Reply-To: <20260505195226.563317-13-pbonzini@redhat.com> References: <20260505195226.563317-1-pbonzini@redhat.com> <20260505195226.563317-13-pbonzini@redhat.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.52.4 (3.52.4-2.fc40) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 On Tue, 2026-05-05 at 21:52 +0200, Paolo Bonzini wrote: > EPT and NPT have different rules for passing PFERR_USER_MASK to the > nested page table walk.=C2=A0 In particular, for final addresses EPT > uses the U bit of the guest (nGVA->nGPA) walk. >=20 > While at it, remove PFERR_USER_MASK from the VMX version of the > function, since it is actually ignored by the tables that > update_permission_bitmask() generates for EPT. Hi! Since this used to not to be the case, it means that in theory=C2=A0 this patch series fixes a theoretical bug in which kvm_translate_gpa would = fail on nested EPT=C2=A0 if requested to do exec-only access,=C2=A0because it used to ask incorrectl= y for PFERR_USER_MASK=C2=A0access which is used to be translated to a read ac= cess,=C2=A0 and therefore=C2=A0 a guest PTE having only exec permission would fail the = check. I am not sure it is worth mentioning this, it is likely impossible to hit this, it is just something that came up my mind when reviewing later patche= s. >=20 > Tested-by: David Riley > Signed-off-by: Paolo Bonzini > --- > =C2=A0arch/x86/include/asm/kvm_host.h |=C2=A0 4 ++++ > =C2=A0arch/x86/kvm/hyperv.c=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2= =A0=C2=A0=C2=A0 |=C2=A0 3 ++- > =C2=A0arch/x86/kvm/mmu.h=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 |=C2=A0 9 +++------ > =C2=A0arch/x86/kvm/svm/nested.c=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 | 15 = +++++++++++++++ > =C2=A0arch/x86/kvm/vmx/nested.c=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 | 12 = ++++++++++++ > =C2=A0arch/x86/kvm/x86.c=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0=C2=A0= =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 | 16 ---------------- > =C2=A06 files changed, 36 insertions(+), 23 deletions(-) >=20 > diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_h= ost.h > index 8f2a1b915df9..62dc782b2dd3 100644 > --- a/arch/x86/include/asm/kvm_host.h > +++ b/arch/x86/include/asm/kvm_host.h > @@ -2010,6 +2010,10 @@ struct kvm_x86_nested_ops { > =C2=A0 struct kvm_nested_state *kvm_state); > =C2=A0 bool (*get_nested_state_pages)(struct kvm_vcpu *vcpu); > =C2=A0 int (*write_log_dirty)(struct kvm_vcpu *vcpu, gpa_t l2_gpa); > + gpa_t (*translate_nested_gpa)(struct kvm_vcpu *vcpu, gpa_t gpa, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 u64 access, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 struct x86_exception *exception, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 u64 pte_access); > =C2=A0 > =C2=A0 int (*enable_evmcs)(struct kvm_vcpu *vcpu, > =C2=A0 =C2=A0=C2=A0=C2=A0 uint16_t *vmcs_version); > diff --git a/arch/x86/kvm/hyperv.c b/arch/x86/kvm/hyperv.c > index 53688f7b76eb..f35fae3a7b3d 100644 > --- a/arch/x86/kvm/hyperv.c > +++ b/arch/x86/kvm/hyperv.c > @@ -2041,7 +2041,8 @@ static u64 kvm_hv_flush_tlb(struct kvm_vcpu *vcpu, = struct kvm_hv_hcall *hc) > =C2=A0 * read with kvm_read_guest(). > =C2=A0 */ > =C2=A0 if (!hc->fast && is_guest_mode(vcpu)) { > - hc->ingpa =3D translate_nested_gpa(vcpu, hc->ingpa, > + hc->ingpa =3D kvm_x86_ops.nested_ops->translate_nested_gpa( > + vcpu, hc->ingpa, > =C2=A0 PFERR_GUEST_FINAL_MASK, NULL, 0); > =C2=A0 if (unlikely(hc->ingpa =3D=3D INVALID_GPA)) > =C2=A0 return HV_STATUS_INVALID_HYPERCALL_INPUT; > diff --git a/arch/x86/kvm/mmu.h b/arch/x86/kvm/mmu.h > index 635c2e5d8513..63be5c5efed9 100644 > --- a/arch/x86/kvm/mmu.h > +++ b/arch/x86/kvm/mmu.h > @@ -294,10 +294,6 @@ static inline void kvm_update_page_stats(struct kvm = *kvm, int level, int count) > =C2=A0 atomic64_add(count, &kvm->stat.pages[level - 1]); > =C2=A0} > =C2=A0 > -gpa_t translate_nested_gpa(struct kvm_vcpu *vcpu, gpa_t gpa, u64 access, > - =C2=A0=C2=A0 struct x86_exception *exception, > - =C2=A0=C2=A0 u64 pte_access); > - > =C2=A0static inline gpa_t kvm_translate_gpa(struct kvm_vcpu *vcpu, > =C2=A0 =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 struct kvm_mmu *mmu, > =C2=A0 =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 gpa_t gpa, u64 access, > @@ -306,8 +302,9 @@ static inline gpa_t kvm_translate_gpa(struct kvm_vcpu= *vcpu, > =C2=A0{ > =C2=A0 if (mmu !=3D &vcpu->arch.nested_mmu) > =C2=A0 return gpa; > - return translate_nested_gpa(vcpu, gpa, access, exception, > - =C2=A0=C2=A0=C2=A0 pte_access); > + return kvm_x86_ops.nested_ops->translate_nested_gpa(vcpu, gpa, access, > + =C2=A0=C2=A0=C2=A0 exception, > + =C2=A0=C2=A0=C2=A0 pte_access); > =C2=A0} > =C2=A0 > =C2=A0static inline bool kvm_has_mirrored_tdp(const struct kvm *kvm) > diff --git a/arch/x86/kvm/svm/nested.c b/arch/x86/kvm/svm/nested.c > index 961804df5f45..df232153eb24 100644 > --- a/arch/x86/kvm/svm/nested.c > +++ b/arch/x86/kvm/svm/nested.c > @@ -2071,8 +2071,23 @@ static bool svm_get_nested_state_pages(struct kvm_= vcpu *vcpu) > =C2=A0 return true; > =C2=A0} > =C2=A0 > +static gpa_t svm_translate_nested_gpa(struct kvm_vcpu *vcpu, gpa_t gpa, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 u64 access, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 struct x86_exception *exception, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 u64 pte_access) > +{ > + struct kvm_mmu *mmu =3D vcpu->arch.mmu; > + > + BUG_ON(!mmu_is_nested(vcpu)); > + > + /* NPT walks are always user-walks */ Tiny nitpick: Maybe we can extend the comment above to explicitly mention t= hat even for the last, actual guest memory access, the CPU pretends to do a use= r access to the NPT? (NPT is really weird...) Something like that: /* NPT walks are always user-walks, even for the actual guest linear memory= access, regardless of the actual guest access (user/supervisor) */ Unrelated, If I understand it correctly, without GMET, NPT *does* check the= U bit, but for all NPT walks, even for the last one, it requires the U bit to be set. And with GMET, U bit is only used for execute access which can happen only on the last level, and for reads/writes U bit is ignored. To be honest, I haven't found this mentioned in the APM. Is this mentioned somewhere in the APM? > + access |=3D PFERR_USER_MASK; > + return mmu->gva_to_gpa(vcpu, mmu, gpa, access, exception); > +} > + > =C2=A0struct kvm_x86_nested_ops svm_nested_ops =3D { > =C2=A0 .leave_nested =3D svm_leave_nested, > + .translate_nested_gpa =3D svm_translate_nested_gpa, > =C2=A0 .is_exception_vmexit =3D nested_svm_is_exception_vmexit, > =C2=A0 .check_events =3D svm_check_nested_events, > =C2=A0 .triple_fault =3D nested_svm_triple_fault, > diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c > index 3fe88f29be7a..cd1924c6e075 100644 > --- a/arch/x86/kvm/vmx/nested.c > +++ b/arch/x86/kvm/vmx/nested.c > @@ -7438,8 +7438,20 @@ __init int nested_vmx_hardware_setup(int (*exit_ha= ndlers[])(struct kvm_vcpu *)) > =C2=A0 return 0; > =C2=A0} > =C2=A0 > +static gpa_t vmx_translate_nested_gpa(struct kvm_vcpu *vcpu, gpa_t gpa, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 u64 access, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 struct x86_exception *exception, > + =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 u64 pte_access) > +{ > + struct kvm_mmu *mmu =3D vcpu->arch.mmu; > + > + BUG_ON(!mmu_is_nested(vcpu)); > + return mmu->gva_to_gpa(vcpu, mmu, gpa, access, exception); > +} > + > =C2=A0struct kvm_x86_nested_ops vmx_nested_ops =3D { > =C2=A0 .leave_nested =3D vmx_leave_nested, > + .translate_nested_gpa =3D vmx_translate_nested_gpa, > =C2=A0 .is_exception_vmexit =3D nested_vmx_is_exception_vmexit, > =C2=A0 .check_events =3D vmx_check_nested_events, > =C2=A0 .has_events =3D vmx_has_nested_events, > diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c > index 67979b7de5d6..7c6942afae81 100644 > --- a/arch/x86/kvm/x86.c > +++ b/arch/x86/kvm/x86.c > @@ -7848,22 +7848,6 @@ void kvm_get_segment(struct kvm_vcpu *vcpu, > =C2=A0 kvm_x86_call(get_segment)(vcpu, var, seg); > =C2=A0} > =C2=A0 > -gpa_t translate_nested_gpa(struct kvm_vcpu *vcpu, gpa_t gpa, u64 access, > - =C2=A0=C2=A0 struct x86_exception *exception, > - =C2=A0=C2=A0 u64 pte_access) > -{ > - struct kvm_mmu *mmu =3D vcpu->arch.mmu; > - gpa_t t_gpa; > - > - BUG_ON(!mmu_is_nested(vcpu)); > - > - /* NPT walks are always user-walks */ > - access |=3D PFERR_USER_MASK; > - t_gpa=C2=A0 =3D mmu->gva_to_gpa(vcpu, mmu, gpa, access, exception); > - > - return t_gpa; > -} > - > =C2=A0gpa_t kvm_mmu_gva_to_gpa_read(struct kvm_vcpu *vcpu, gva_t gva, > =C2=A0 =C2=A0=C2=A0=C2=A0=C2=A0=C2=A0 struct x86_exception *exception) > =C2=A0{ Reviewed-by: Maxim Levitsky Best regards, Maxim Levitsky