From: Yan Zhao <yan.y.zhao@intel.com>
To: seanjc@google.com, pbonzini@redhat.com, dave.hansen@intel.com
Cc: linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
x86@kernel.org, rick.p.edgecombe@intel.com, kas@kernel.org,
tabba@google.com, ackerleytng@google.com, michael.roth@amd.com,
david@kernel.org, vannapurve@google.com, sagis@google.com,
vbabka@suse.cz, thomas.lendacky@amd.com, nik.borisov@suse.com,
pgonda@google.com, fan.du@intel.com, jun.miao@intel.com,
francescolavra.fl@gmail.com, jgross@suse.com,
xiaoyao.li@intel.com, kai.huang@intel.com,
binbin.wu@linux.intel.com, chao.p.peng@intel.com,
chao.gao@intel.com, farrah.chen@intel.com, yan.y.zhao@intel.com
Subject: [PATCH v4 11/17] KVM: TDX: Honor the guest's accept level contained in an EPT violation
Date: Mon, 28 Sep 2026 17:11:16 +0800 [thread overview]
Message-ID: <20260928091116.15647-1-yan.y.zhao@intel.com> (raw)
In-Reply-To: <20260928090729.15468-1-yan.y.zhao@intel.com>
TDX requires guests to accept S-EPT mappings created by the host KVM. Due
to the current implementation of the TDX module, if a guest accepts a GFN
at a lower level after KVM maps it at a higher level, the TDX module will
synthesize an EPT Violation VM-Exit to KVM instead of returning a size
mismatch error to the guest. If KVM fails to perform page splitting in the
EPT Violation handler, the guest's ACCEPT operation will be triggered
again upon re-entering the guest, causing a repeated EPT Violation VM-Exit.
To ensure forward progress, honor the guest's accept level if an EPT
Violation VM-Exit contains the guest accept level (the TDX module provides
the level when synthesizing a VM-Exit in response to a failed guest ACCEPT,
e.g., due to accept level < host mapping level or due to intermediate
paging structure missing or inaccessible).
(1) Set the guest inhibit bit in the lpage info to prevent KVM's MMU
from mapping at a higher level than the guest's accept level.
(2) Split any existing mapping higher than the guest's accept level.
For now, take mmu_lock for write across the entire operation to keep things
simple. This can/will be revisited when the TDX module adds support for
NON-BLOCKING-RESIZE, at which point KVM can split the huge page without
needing to handle UNBLOCK failure if the DEMOTE fails.
To avoid unnecessarily contending mmu_lock, check if the inhibit flag is
already set before acquiring mmu_lock, e.g. so that vCPUs doing ACCEPT
on a region of memory aren't completely serialized. Note, this relies on
(a) setting the inhibit after performing the split, and (b) never clearing
the flag, e.g., to avoid false positives and potentially triggering the
zero-step mitigation.
Note: EPT Violation VM-Exits without the guest's accept level are *never*
caused by the guest's ACCEPT operation, but instead occur if the guest
accesses memory before said memory is accepted. Since KVM can't obtain
the guest accept level info from such EPT Violations (the ACCEPT operation
hasn't occurred yet), KVM may still map at a higher level than the guest's
later ACCEPT level.
So, the typical guest/KVM interaction flow is:
- If guest accesses private memory without first accepting it,
(like non-Linux guests):
1. Guest accesses a private memory.
2. KVM finds it can map the GFN at 2MB. So, AUG at 2MB.
3. Guest accepts the GFN at 4KB.
4. KVM receives an EPT violation with eeq_type of ACCEPT + 4KB level.
5. KVM splits the 2MB mapping.
6. Guest accepts successfully and accesses the page.
- If guest first accepts private memory before accessing it,
(like Linux guests):
1. Guest accepts a private memory at 4KB.
2. KVM receives an EPT violation with eeq_type of ACCEPT + 4KB level.
3. KVM AUG at 4KB.
4. Guest accepts successfully and accesses the page.
Link: https://lore.kernel.org/all/a6ffe23fb97e64109f512fa43e9f6405236ed40a.camel@intel.com
Suggested-by: Rick Edgecombe <rick.p.edgecombe@intel.com>
Suggested-by: Sean Christopherson <seanjc@google.com>
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- In [1], Sean renamed tdx_honor_guest_accept_level() to
tdx_handle_mismatched_accept(). In v4, Yan further renamed it to
tdx_handle_guest_accept_ept_violation(), because EPT violations caused by
a guest's ACCEPT operation are not necessarily due to the guest's ACCEPT
level mismatching KVM's mapping level. In the case where a guest accepts
memory before ever accessing it, EPT violations caused by the guest's
ACCEPT operation occur when there are no present KVM mappings.
- Introduced helper tdx_is_mismatched_accepted() (Yan renamed it to
tdx_is_guest_accept_ept_violation()). (Sean)
- Introduced helper tdx_get_ept_violation_level(). (Sean)
- Invoked kvm_tdp_mmu_mirrors_split_huge_pages() instead of
kvm_split_cross_boundary_leafs(). (Sean)
[1] https://lore.kernel.org/all/20260129011517.3545883-42-seanjc@google.com
---
arch/x86/kvm/vmx/tdx.c | 75 +++++++++++++++++++++++++++++++++++++
arch/x86/kvm/vmx/tdx_arch.h | 3 ++
2 files changed, 78 insertions(+)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 397a308b834c..f97b76bd8fe9 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -14,6 +14,7 @@
#include "tdx.h"
#include "vmx.h"
#include "mmu/spte.h"
+#include "mmu/tdp_mmu.h"
#include "common.h"
#include "posted_intr.h"
#include "irq.h"
@@ -2049,6 +2050,76 @@ static inline bool tdx_is_sept_violation_unexpected_pending(struct kvm_vcpu *vcp
return !(eq & EPT_VIOLATION_PROT_MASK);
}
+static bool tdx_is_guest_accept_ept_violation(struct kvm_vcpu *vcpu)
+{
+ return (to_tdx(vcpu)->ext_exit_qualification & TDX_EXT_EXIT_QUAL_TYPE_MASK) ==
+ TDX_EXT_EXIT_QUAL_TYPE_ACCEPT;
+}
+
+static int tdx_get_ept_violation_level(struct kvm_vcpu *vcpu)
+{
+ u64 ext_exit_qual = to_tdx(vcpu)->ext_exit_qualification;
+
+ return (((ext_exit_qual & TDX_EXT_EXIT_QUAL_INFO_MASK) >>
+ TDX_EXT_EXIT_QUAL_INFO_SHIFT) & GENMASK(2, 0)) + 1;
+}
+
+/*
+ * An EPT violation can be either due to the guest's ACCEPT operation or
+ * due to the guest's access of memory before the guest accepts the
+ * memory.
+ *
+ * Type TDX_EXT_EXIT_QUAL_TYPE_ACCEPT in the extended exit qualification
+ * identifies EPT violations caused by the guest's ACCEPT operation, which must
+ * contain a valid guest accept level. For such EPT violations, honor guest's
+ * accept level by setting guest inhibit bit on levels above the guest accept
+ * level and split the existing mapping for the faulting GFN if it's with a
+ * higher level than the guest accept level.
+ *
+ * Do nothing if the EPT violation is not caused by the guest's ACCEPT
+ * operation. KVM will map the GFN without considering the guest's accept level
+ * (unless the guest inhibit bit is already set).
+ */
+static int tdx_handle_guest_accept_ept_violation(struct kvm_vcpu *vcpu, gfn_t gfn)
+{
+ struct kvm_memory_slot *slot = kvm_vcpu_gfn_to_memslot(vcpu, gfn);
+ struct kvm *kvm = vcpu->kvm;
+ gfn_t start, end;
+ int level, r;
+
+ if (!slot || !tdx_is_guest_accept_ept_violation(vcpu))
+ return 0;
+
+ if (WARN_ON_ONCE(!VALID_PAGE(vcpu->arch.mmu->mirror_root_hpa)))
+ return 0;
+
+ level = tdx_get_ept_violation_level(vcpu);
+ if (level > PG_LEVEL_2M)
+ return 0;
+
+ if (hugepage_test_guest_inhibit(slot, gfn, level + 1))
+ return 0;
+
+ guard(write_lock)(&kvm->mmu_lock);
+
+ start = gfn_round_for_level(gfn, level);
+ end = start + KVM_PAGES_PER_HPAGE(level);
+
+ r = kvm_tdp_mmu_mirrors_split_huge_pages(kvm, start, end, level);
+ if (r)
+ return r;
+
+ /*
+ * No TLB flush is required, as the "BLOCK + TRACK + kick off vCPUs"
+ * sequence required by the TDX-Module includes a TLB flush.
+ */
+ hugepage_set_guest_inhibit(slot, gfn, level + 1);
+ if (level == PG_LEVEL_4K)
+ hugepage_set_guest_inhibit(slot, gfn, level + 2);
+
+ return 0;
+}
+
static int tdx_handle_ept_violation(struct kvm_vcpu *vcpu)
{
unsigned long exit_qual;
@@ -2074,6 +2145,10 @@ static int tdx_handle_ept_violation(struct kvm_vcpu *vcpu)
*/
exit_qual = EPT_VIOLATION_ACC_WRITE;
+ ret = tdx_handle_guest_accept_ept_violation(vcpu, gpa_to_gfn(gpa));
+ if (ret)
+ return ret;
+
/* Only private GPA triggers zero-step mitigation */
local_retry = true;
} else {
diff --git a/arch/x86/kvm/vmx/tdx_arch.h b/arch/x86/kvm/vmx/tdx_arch.h
index 350143b9b145..6b8b18b0689a 100644
--- a/arch/x86/kvm/vmx/tdx_arch.h
+++ b/arch/x86/kvm/vmx/tdx_arch.h
@@ -76,7 +76,10 @@ struct tdx_cpuid_value {
} __packed;
#define TDX_EXT_EXIT_QUAL_TYPE_MASK GENMASK(3, 0)
+#define TDX_EXT_EXIT_QUAL_TYPE_ACCEPT 1
#define TDX_EXT_EXIT_QUAL_TYPE_PENDING_EPT_VIOLATION 6
+#define TDX_EXT_EXIT_QUAL_INFO_MASK GENMASK(63, 32)
+#define TDX_EXT_EXIT_QUAL_INFO_SHIFT 32
/*
* TD_PARAMS is provided as an input to TDH_MNG_INIT, the size of which is 1024B.
*/
--
2.43.2
next prev parent reply other threads:[~2026-09-28 9:11 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
2026-09-28 9:08 ` [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages Yan Zhao
2026-09-28 9:08 ` [PATCH v4 02/17] x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page Yan Zhao
2026-09-28 9:08 ` [PATCH v4 03/17] KVM: TDX: Reset private huge pages after S-EPT page removal Yan Zhao
2026-09-28 9:09 ` [PATCH v4 04/17] KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault path Yan Zhao
2026-09-28 9:09 ` [PATCH v4 05/17] KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table splitting Yan Zhao
2026-09-28 9:10 ` [PATCH v4 06/17] KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split Yan Zhao
2026-09-28 9:10 ` [PATCH v4 07/17] KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings to 4KB Yan Zhao
2026-09-28 9:10 ` [PATCH v4 08/17] KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting S-EPT Yan Zhao
2026-09-28 9:10 ` [PATCH v4 09/17] KVM: x86/mmu: Introduce hugepage_set_guest_inhibit() Yan Zhao
2026-09-28 9:11 ` [PATCH v4 10/17] KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots Yan Zhao
2026-09-28 9:11 ` Yan Zhao [this message]
2026-09-28 9:11 ` [PATCH v4 12/17] KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU context Yan Zhao
2026-09-28 9:11 ` [PATCH v4 13/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns give gmem+slot+pgoff Yan Zhao
2026-09-28 9:11 ` [PATCH v4 14/17] [GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start() to start() and zap() Yan Zhao
2026-09-28 9:12 ` [PATCH v4 15/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap() Yan Zhao
2026-09-28 9:12 ` [PATCH v4 16/17] [GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split S-EPT Yan Zhao
2026-09-28 9:12 ` [PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M Yan Zhao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928091116.15647-1-yan.y.zhao@intel.com \
--to=yan.y.zhao@intel.com \
--cc=ackerleytng@google.com \
--cc=binbin.wu@linux.intel.com \
--cc=chao.gao@intel.com \
--cc=chao.p.peng@intel.com \
--cc=dave.hansen@intel.com \
--cc=david@kernel.org \
--cc=fan.du@intel.com \
--cc=farrah.chen@intel.com \
--cc=francescolavra.fl@gmail.com \
--cc=jgross@suse.com \
--cc=jun.miao@intel.com \
--cc=kai.huang@intel.com \
--cc=kas@kernel.org \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=michael.roth@amd.com \
--cc=nik.borisov@suse.com \
--cc=pbonzini@redhat.com \
--cc=pgonda@google.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sagis@google.com \
--cc=seanjc@google.com \
--cc=tabba@google.com \
--cc=thomas.lendacky@amd.com \
--cc=vannapurve@google.com \
--cc=vbabka@suse.cz \
--cc=x86@kernel.org \
--cc=xiaoyao.li@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®