mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Yan Zhao <yan.y.zhao@intel.com>
To: seanjc@google.com, pbonzini@redhat.com, dave.hansen@intel.com
Cc: linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
	x86@kernel.org, rick.p.edgecombe@intel.com, kas@kernel.org,
	tabba@google.com, ackerleytng@google.com, michael.roth@amd.com,
	david@kernel.org, vannapurve@google.com, sagis@google.com,
	vbabka@suse.cz, thomas.lendacky@amd.com, nik.borisov@suse.com,
	pgonda@google.com, fan.du@intel.com, jun.miao@intel.com,
	francescolavra.fl@gmail.com, jgross@suse.com,
	xiaoyao.li@intel.com, kai.huang@intel.com,
	binbin.wu@linux.intel.com, chao.p.peng@intel.com,
	chao.gao@intel.com, farrah.chen@intel.com, yan.y.zhao@intel.com
Subject: [PATCH v4 12/17] KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU context
Date: Mon, 28 Sep 2026 17:11:31 +0800	[thread overview]
Message-ID: <20260928091131.15663-1-yan.y.zhao@intel.com> (raw)
In-Reply-To: <20260928090729.15468-1-yan.y.zhao@intel.com>

Add support for splitting S-EPT entries under a non-vCPU context. This
prepares for zapping a subset of a huge mapping in S-EPT caused by
private-to-shared conversions or guest_memfd reclaiming of physical memory.

KVM must precisely zap/remove S-EPT entries to avoid clobbering guest
memory (the lifetime of guest private memory is tied to the S-EPT). So, KVM
needs to first split a huge mapping so that small mappings can be zapped
precisely.

Since there's no vCPU context, introduce a per-VM PAMT cache of
pre-allocated pages used to populate the Dynamic PAMT. Add a helper
tdx_get_pamt_cache() to select the per-VM PAMT cache when there's no vCPU
context.  Add the "kvm" arg to .topup_external_cache() and its caller
tdp_mmu_alloc_sp_for_split() for the purpose of passing the "kvm" arg to
tdx_get_pamt_cache().

Use a mutex to guard the entire cycle from the per-VM PAMT cache topup to
drawing pages from the cache. Using a mutex (e.g., versus a spinlock) is
important as it allows KVM to only drop and re-aquire the mmu_lock (a
spinlock) while continuing holding the mutex for memory allocation.

Introduce a local static function tdx_sept_split_huge_pages(), which
acquires the mutex before triggering the S-EPT entries splitting under a
non-vCPU context. This function is intended to be invoked by guest_memfd
via an arch hook in a later patch. Though functions related to dirty page
tracking can also trigger splitting under a non-vCPU context, they do not
yet involve mirror roots. So, how those functions should acquire the mutex
is deferred to a later consideration.

tdx_sept_split_huge_pages() internally invokes API
kvm_tdp_mmu_mirrors_split_huge_pages() to split mirror roots. To avoid
unnecessary work, explicitly detect unaligned head and tail pages relative
to the max page size supported by KVM (currently 2MB for private memory),
and split only those pages, as only unaligned head/tail pages will undergo
partial zapping.

Signed-off-by: Sean Christopherson <seanjc@google.com>
[Yan: Tweak patch log/function names, split out .gmem_prezap() hook]
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- New patch.
- Split out the registration of the .gmem_prezap hook into a later patch to
  isolate gmem-related changes. (Yan)
---
 arch/x86/include/asm/kvm_host.h |  2 +-
 arch/x86/kvm/mmu/mmu.c          |  2 +-
 arch/x86/kvm/mmu/tdp_mmu.c      |  7 +--
 arch/x86/kvm/vmx/tdx.c          | 94 +++++++++++++++++++++++++++++----
 arch/x86/kvm/vmx/tdx.h          |  5 ++
 5 files changed, 96 insertions(+), 14 deletions(-)

diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index 373559a7cca5..5dd1db64562f 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -1650,7 +1650,7 @@ struct kvm_x86_ops {
 	/* Update external page tables for page table about to be freed. */
 	void (*free_external_spt)(struct kvm *kvm, struct kvm_mmu_page *sp);
 
-	int (*topup_external_cache)(struct kvm_vcpu *vcpu, int min_nr_spts);
+	int (*topup_external_cache)(struct kvm *kvm, struct kvm_vcpu *vcpu, int min_nr_spts);
 
 	bool (*has_wbinvd_exit)(void);
 
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 0725b92b7008..d92170a34728 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -618,7 +618,7 @@ static int mmu_topup_memory_caches(struct kvm_vcpu *vcpu, bool maybe_indirect)
 		if (r)
 			return r;
 
-		r = kvm_x86_call(topup_external_cache)(vcpu, PT64_ROOT_MAX_LEVEL);
+		r = kvm_x86_call(topup_external_cache)(vcpu->kvm, vcpu, PT64_ROOT_MAX_LEVEL);
 		if (r)
 			return r;
 	}
diff --git a/arch/x86/kvm/mmu/tdp_mmu.c b/arch/x86/kvm/mmu/tdp_mmu.c
index 9c783accb9ed..4829ddcd3b55 100644
--- a/arch/x86/kvm/mmu/tdp_mmu.c
+++ b/arch/x86/kvm/mmu/tdp_mmu.c
@@ -1466,7 +1466,8 @@ bool kvm_tdp_mmu_wrprot_slot(struct kvm *kvm,
 	return spte_set;
 }
 
-static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(bool is_mirror_sp)
+static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(struct kvm *kvm,
+						       bool is_mirror_sp)
 {
 	struct kvm_mmu_page *sp;
 
@@ -1483,7 +1484,7 @@ static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(bool is_mirror_sp)
 		if (!sp->external_spt)
 			goto err_external_spt;
 
-		if (kvm_x86_call(topup_external_cache)(kvm_get_running_vcpu(), 1))
+		if (kvm_x86_call(topup_external_cache)(kvm, kvm_get_running_vcpu(), 1))
 			goto err_external_split;
 	}
 
@@ -1575,7 +1576,7 @@ static int tdp_mmu_split_huge_pages_root(struct kvm *kvm,
 			else
 				write_unlock(&kvm->mmu_lock);
 
-			sp = tdp_mmu_alloc_sp_for_split(is_mirror_root);
+			sp = tdp_mmu_alloc_sp_for_split(kvm, is_mirror_root);
 
 			if (shared)
 				read_lock(&kvm->mmu_lock);
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index f97b76bd8fe9..5c5919b76bdd 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -585,6 +585,8 @@ void tdx_vm_destroy(struct kvm *kvm)
 {
 	struct kvm_tdx *kvm_tdx = to_kvm_tdx(kvm);
 
+	tdx_free_pamt_cache(&kvm_tdx->pamt_cache);
+
 	tdx_reclaim_td_control_pages(kvm);
 
 	kvm_tdx->state = TD_STATE_UNINITIALIZED;
@@ -650,6 +652,9 @@ int tdx_vm_init(struct kvm *kvm)
 
 	kvm_tdx->state = TD_STATE_UNINITIALIZED;
 
+	tdx_init_pamt_cache(&kvm_tdx->pamt_cache);
+	mutex_init(&kvm_tdx->pamt_cache_lock);
+
 	return 0;
 }
 
@@ -1629,15 +1634,31 @@ void tdx_load_mmu_pgd(struct kvm_vcpu *vcpu, hpa_t root_hpa, int pgd_level)
 	td_vmcs_write64(to_tdx(vcpu), SHARED_EPT_POINTER, root_hpa);
 }
 
-static int tdx_topup_external_pamt_cache(struct kvm_vcpu *vcpu, int min_nr_spts)
+static struct tdx_pamt_cache *tdx_get_pamt_cache(struct kvm *kvm,
+						 struct kvm_vcpu *vcpu)
 {
+	if (KVM_BUG_ON(vcpu && vcpu->kvm != kvm, kvm))
+		return NULL;
+
+	if (vcpu)
+		return &to_tdx(vcpu)->pamt_cache;
+
+	lockdep_assert_held(&to_kvm_tdx(kvm)->pamt_cache_lock);
+	return &to_kvm_tdx(kvm)->pamt_cache;
+}
+
+static int tdx_topup_external_pamt_cache(struct kvm *kvm, struct kvm_vcpu *vcpu,
+					 int min_nr_spts)
+{
+	struct tdx_pamt_cache *pamt_cache;
 	int dpamt_pairs;
 
-	if (WARN_ON_ONCE(!vcpu))
+	pamt_cache = tdx_get_pamt_cache(kvm, vcpu);
+	if (!pamt_cache)
 		return -EIO;
 
 	/* Exclude the root SPT, as its DPAMT page pair is already installed */
-	if (min_nr_spts == vcpu->kvm->arch.mirror_root_level)
+	if (min_nr_spts == kvm->arch.mirror_root_level)
 		min_nr_spts -= 1;
 
 	/*
@@ -1658,7 +1679,7 @@ static int tdx_topup_external_pamt_cache(struct kvm_vcpu *vcpu, int min_nr_spts)
 	 */
 	dpamt_pairs += 1;
 
-	return tdx_topup_pamt_cache(&to_tdx(vcpu)->pamt_cache, dpamt_pairs);
+	return tdx_topup_pamt_cache(pamt_cache, dpamt_pairs);
 }
 
 static int tdx_mem_page_add(struct kvm *kvm, gfn_t gfn, enum pg_level level,
@@ -1909,8 +1930,8 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
 static int tdx_sept_split_leaf_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
 				    u64 new_spte, enum pg_level level)
 {
-	struct kvm_vcpu *vcpu = kvm_get_running_vcpu();
 	struct kvm_tdx *kvm_tdx = to_kvm_tdx(kvm);
+	struct tdx_pamt_cache *pamt_cache;
 	gpa_t gpa = gfn_to_gpa(gfn);
 	u64 err, entry, level_state;
 	struct page *sept_pt;
@@ -1925,10 +1946,11 @@ static int tdx_sept_split_leaf_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
 	if (!sept_pt)
 		return -EIO;
 
-	if (KVM_BUG_ON(!vcpu || vcpu->kvm != kvm, kvm))
+	pamt_cache = tdx_get_pamt_cache(kvm, kvm_get_running_vcpu());
+	if (!pamt_cache)
 		return -EIO;
 
-	r = tdx_pamt_get(page_to_pfn(sept_pt), PG_LEVEL_4K, &to_tdx(vcpu)->pamt_cache);
+	r = tdx_pamt_get(page_to_pfn(sept_pt), PG_LEVEL_4K, pamt_cache);
 	if (KVM_BUG_ON(r, kvm))
 		return r;
 
@@ -1941,8 +1963,8 @@ static int tdx_sept_split_leaf_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
 
 	tdx_track(kvm);
 	err = tdh_do_no_vcpus(tdh_mem_page_demote, kvm, &kvm_tdx->td, gpa,
-			      level, spte_to_pfn(old_spte), sept_pt,
-			      &to_tdx(vcpu)->pamt_cache, &entry, &level_state);
+			      level, spte_to_pfn(old_spte), sept_pt, pamt_cache,
+			      &entry, &level_state);
 	if (TDX_BUG_ON_2(err, TDH_MEM_PAGE_DEMOTE, entry, level_state, kvm)) {
 		r = -EIO;
 		goto err;
@@ -2027,6 +2049,60 @@ static void tdx_sept_free_private_spt(struct kvm *kvm, struct kvm_mmu_page *sp)
 	sp->external_spt = NULL;
 }
 
+static int tdx_sept_split_huge_page_at(struct kvm *kvm, gfn_t gfn, int target_level)
+{
+	gfn_t end = gfn + KVM_PAGES_PER_HPAGE(target_level + 1);
+
+	return kvm_tdp_mmu_mirrors_split_huge_pages(kvm, gfn, end, target_level);
+}
+
+static int tdx_sept_split_straddling_huge_pages_to_level(struct kvm *kvm, gfn_t start,
+							 gfn_t end, int target_level)
+{
+	gfn_t head = gfn_round_for_level(start, target_level + 1);
+	gfn_t tail = gfn_round_for_level(end, target_level + 1);
+	int r;
+
+	if (head != start) {
+		r = tdx_sept_split_huge_page_at(kvm, head, target_level);
+		if (r)
+			return r;
+	}
+
+	if (tail != end && (head != tail || head == start)) {
+		r = tdx_sept_split_huge_page_at(kvm, tail, target_level);
+		if (r)
+			return r;
+	}
+
+	return 0;
+}
+
+/*
+ * Split S-EPT huge mappings that straddle [start, end) under non-vCPU context.
+ *
+ * Split potential huge mappings at the head and tail of the to-be-zapped range
+ * so that KVM doesn't overzap due to dropping a hugepage that doesn't fall
+ * wholly inside the range.
+ *
+ * Acquire the external cache lock, a.k.a. the Dynamic PAMT lock, to protect the
+ * per-VM cache of pre-allocated pages used to populate the Dynamic PAMT when
+ * splitting S-EPT huge pages.
+ */
+static int __maybe_unused tdx_sept_split_huge_pages(struct kvm *kvm, gfn_t start,
+						    gfn_t end)
+{
+	guard(mutex)(&to_kvm_tdx(kvm)->pamt_cache_lock);
+
+	guard(write_lock)(&kvm->mmu_lock);
+
+	/*
+	 * TODO: Also split from PG_LEVEL_1G => PG_LEVEL_2M when KVM supports
+	 *       1GB S-EPT pages.
+	 */
+	return tdx_sept_split_straddling_huge_pages_to_level(kvm, start, end, PG_LEVEL_4K);
+}
+
 void tdx_deliver_interrupt(struct kvm_lapic *apic, int delivery_mode,
 			   int trig_mode, int vector)
 {
diff --git a/arch/x86/kvm/vmx/tdx.h b/arch/x86/kvm/vmx/tdx.h
index fd368e3ee060..602d1c289c2d 100644
--- a/arch/x86/kvm/vmx/tdx.h
+++ b/arch/x86/kvm/vmx/tdx.h
@@ -47,6 +47,11 @@ struct kvm_tdx {
 	 * Set/unset is protected with kvm->mmu_lock.
 	 */
 	bool wait_for_sept_zap;
+
+	/* The per-VM cache for DPAMT pages for S-EPT pages and guest pages */
+	struct tdx_pamt_cache pamt_cache;
+	/* Protect the per-VM cache for DPAMT pages */
+	struct mutex pamt_cache_lock;
 };
 
 /* TDX module vCPU states */
-- 
2.43.2


  parent reply	other threads:[~2026-09-28  9:12 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
2026-09-28  9:08 ` [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages Yan Zhao
2026-09-28  9:08 ` [PATCH v4 02/17] x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page Yan Zhao
2026-09-28  9:08 ` [PATCH v4 03/17] KVM: TDX: Reset private huge pages after S-EPT page removal Yan Zhao
2026-09-28  9:09 ` [PATCH v4 04/17] KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault path Yan Zhao
2026-09-28  9:09 ` [PATCH v4 05/17] KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table splitting Yan Zhao
2026-09-28  9:10 ` [PATCH v4 06/17] KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split Yan Zhao
2026-09-28  9:10 ` [PATCH v4 07/17] KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings to 4KB Yan Zhao
2026-09-28  9:10 ` [PATCH v4 08/17] KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting S-EPT Yan Zhao
2026-09-28  9:10 ` [PATCH v4 09/17] KVM: x86/mmu: Introduce hugepage_set_guest_inhibit() Yan Zhao
2026-09-28  9:11 ` [PATCH v4 10/17] KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots Yan Zhao
2026-09-28  9:11 ` [PATCH v4 11/17] KVM: TDX: Honor the guest's accept level contained in an EPT violation Yan Zhao
2026-09-28  9:11 ` Yan Zhao [this message]
2026-09-28  9:11 ` [PATCH v4 13/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns give gmem+slot+pgoff Yan Zhao
2026-09-28  9:11 ` [PATCH v4 14/17] [GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start() to start() and zap() Yan Zhao
2026-09-28  9:12 ` [PATCH v4 15/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap() Yan Zhao
2026-09-28  9:12 ` [PATCH v4 16/17] [GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split S-EPT Yan Zhao
2026-09-28  9:12 ` [PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M Yan Zhao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260928091131.15663-1-yan.y.zhao@intel.com \
    --to=yan.y.zhao@intel.com \
    --cc=ackerleytng@google.com \
    --cc=binbin.wu@linux.intel.com \
    --cc=chao.gao@intel.com \
    --cc=chao.p.peng@intel.com \
    --cc=dave.hansen@intel.com \
    --cc=david@kernel.org \
    --cc=fan.du@intel.com \
    --cc=farrah.chen@intel.com \
    --cc=francescolavra.fl@gmail.com \
    --cc=jgross@suse.com \
    --cc=jun.miao@intel.com \
    --cc=kai.huang@intel.com \
    --cc=kas@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=michael.roth@amd.com \
    --cc=nik.borisov@suse.com \
    --cc=pbonzini@redhat.com \
    --cc=pgonda@google.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=sagis@google.com \
    --cc=seanjc@google.com \
    --cc=tabba@google.com \
    --cc=thomas.lendacky@amd.com \
    --cc=vannapurve@google.com \
    --cc=vbabka@suse.cz \
    --cc=x86@kernel.org \
    --cc=xiaoyao.li@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®