mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Yan Zhao <yan.y.zhao@intel.com>
To: seanjc@google.com, pbonzini@redhat.com, dave.hansen@intel.com
Cc: linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
	x86@kernel.org, rick.p.edgecombe@intel.com, kas@kernel.org,
	tabba@google.com, ackerleytng@google.com, michael.roth@amd.com,
	david@kernel.org, vannapurve@google.com, sagis@google.com,
	vbabka@suse.cz, thomas.lendacky@amd.com, nik.borisov@suse.com,
	pgonda@google.com, fan.du@intel.com, jun.miao@intel.com,
	francescolavra.fl@gmail.com, jgross@suse.com,
	xiaoyao.li@intel.com, kai.huang@intel.com,
	binbin.wu@linux.intel.com, chao.p.peng@intel.com,
	chao.gao@intel.com, farrah.chen@intel.com, yan.y.zhao@intel.com
Subject: [PATCH v4 02/17] x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page
Date: Mon, 28 Sep 2026 17:08:31 +0800	[thread overview]
Message-ID: <20260928090831.15502-1-yan.y.zhao@intel.com> (raw)
In-Reply-To: <20260928090729.15468-1-yan.y.zhao@intel.com>

tl;dr: Introduce a SEAMCALL wrapper to invoke the TDH_MEM_PAGE_DEMOTE
SEAMCALL for splitting a 2MB page mapped in S-EPT, and export it as an API.

The following assumptions are made by this wrapper:
1. The huge page is mapped at 2MB level in the S-EPT before invoking this
   wrapper.
2. DPAMT is enabled.
3. The TDX module supports uninterruptible demote.

The only expected error returned from this wrapper is TDX_OPERAND_BUSY.
Callers may need to kick vCPUs out of guest mode to avoid this error.
Callers must provide a pamt_cache for drawing DPAMT pages for guest memory
and ensure that pages can be drawn from the pamt_cache locklessly in this
wrapper.
Callers must ensure no concurrent invocations of this wrapper on the same
GPA.

Long version:

Huge pages offer improved TLB efficiency, faster page walks, better memory
contiguity, and reduced page fault and memory overhead (fewer intermediate
page table structures, and no 4KB DPAMT pages). However, splitting/demoting
a huge page is required in the following cases:
- when the guest accepts the private memory at a lower level,
- when part of the huge page is converted from private to shared, or
- when VMM wants to reclaim part of the huge page.

The TDH_MEM_PAGE_DEMOTE SEAMCALL demotes a guest private huge page by
locating the corresponding leaf huge S-EPT entry and replacing it with a
non-leaf S-EPT entry pointing to a new page table page, which holds the
split leaf S-EPT entries. When DPAMT is enabled, upon successfully demoting
a 2MB page to 4KB pages, the SEAMCALL performs DPAMT pages installation for
the 2MB physical memory range, similar to the DPAMT pages installation
performed by the TDH_PHYMEM_PAMT_ADD SEAMCALL.

The wrapper args "gpa", "level" specify the original huge leaf S-EPT entry;
"new_sept_pt" specifies the newly added page table page; "pamt_cache"
provides a pair of pages to gift to the TDX module as DPAMT pages; "pfn" is
used to locate the dpamt_refcount; and the output args "ext_err1" and
"ext_err2" allow the caller to retrieve SEAMCALL failure information.

Three assumptions are made by this wrapper (as in "tl;dr").
Assumption 1 holds because the VMM can only create mappings up to 2MB level
in the S-EPT via the TDH_MEM_PAGE_AUG SEAMCALL. 1GB mappings in the S-EPT
are only possible via the TDH_MEM_PAGE_PROMOTE SEAMCALL, which is not yet
supported.

Assumptions 1 and 2 together simplify the wrapper implementation by
eliminating the need to check whether DPAMT-related handling is required.

Assumption 3 is a tradeoff reached after research and discussion: a TDX
module that does not support uninterruptible demote returns the error
TDX_INTERRUPTED_RESTARTABLE if any pending host interrupts are detected
during the TDH_MEM_PAGE_DEMOTE SEAMCALL. Since the TDX module cannot
guarantee a maximum retry count to ensure forward progress of the demotion,
interrupt storms could result in a DoS if the host retries endlessly.
Disabling interrupts before invoking the TDH_MEM_PAGE_DEMOTE SEAMCALL also
does not work, as the TDX module also checks for pending NMIs. Therefore,
the tradeoff is to disable huge pages when the TDX module does not support
uninterruptible demote. This is acceptable for basic TDX huge page
enabling, since the SEAMCALL execution time under uninterruptible demote
mode remains reasonable [1][2].

Later patches in KVM will enforce theses assumptions by setting the maximum
mapping level to 2MB, disallowing page merging, and disabling TDX huge
pages if assumption 2 or 3 is not met. This wrapper warns and returns
TDX_SW_ERROR if any assumption is violated [6].

Due to assumptions 1 and 2, the wrapper always performs DPAMT-related
handling before and after invoking the core helper __tdh_mem_page_demote():
- Draw a pair of pages from pamt_cache and gift them to the TDX module for
  DPAMT pages installation. The pamt_cache is a list of pre-allocated pages
  provided by the caller. The caller ensures that drawing pages from the
  list is contention-free, either by using a per-vCPU thread-local list or
  by holding a per-VM lock. If the SEAMCALL fails, the pair of pages is
  freed directly rather than being re-inserted into the list, for
  simplicity.

- Use the global spinlock dpamt_lock to avoid potential contention between
  the TDH_MEM_PAGE_DEMOTE and TDH_PHYMEM_PAMT_{ADD/REMOVE} SEAMCALLs [3].
  The contention here can be cross-VM, similar to that between
  TDH_PHYMEM_PAMT_ADD and TDH_PHYMEM_PAMT_{ADD/REMOVE} [4]. dpamt_lock is
  acquired directly without any pre-testing, since a 2MB huge page mapped
  at the 2MB level in the S-EPT must have no existing dpamt_refcount before
  demotion unless there is a bug. For defensive programming, warn and
  return an error if dpamt_refcount is non-zero after acquiring dpamt_lock.

- Set dpamt_refcount to 512 upon successful demotion so that future
  invocations of tdx_pamt_put() when removing the split 4KB pages work
  correctly.

Note: The new_sept_pt page, which will be added as the S-EPT page table
      page upon successful demotion, also requires DPAMT pages installation
      before invoking the TDH_MEM_PAGE_DEMOTE SEAMCALL. This wrapper relies
      on the caller to invoke tdx_pamt_get() separately for the new_sept_pt
      page prior to calling this wrapper.

Additionally, caller must ensure no concurrent invocations of this wrapper
on the same GPA (for example, by holding KVM's write mmu_lock). Otherwise,
a second invocation on the same GPA may encounter an incorrect
dpamt_refcount or a SEAMCALL error caused by attempting to have the TDX
module demote a 4KB page. This not only simplifies the wrapper
implementation, but also prevents potential issues in the caller (for
example, avoiding a level mismatch between the KVM mirror EPT entry and the
corresponding S-EPT entry).

The only expected error returned from this wrapper is TDX_OPERAND_BUSY,
caused by potential contention between the TDH_MEM_PAGE_DEMOTE SEAMCALL and
TDH_VP_ENTER SEAMALL or guest TDCALLs operating on the S-EPT entry being
split, due to lack of lock protection. This error is returned directly to
the caller without retrying; the caller may handle it by kicking vCPUs and
preventing them from re-entering guest mode to resolve the contention.

For defensive programming, the wrapper still takes arg "level" (though it's
now always expected to be 2MB); a warning is issued before returning any
unexpected errors, regardless of whether the caller would issue a similar
warning [5][6].

Link: https://lore.kernel.org/kvm/99f5585d759328db973403be0713f68e492b492a.camel@intel.com [1]
Link: https://lore.kernel.org/all/fbf04b09f13bc2ce004ac97ee9c1f2c965f44fdf.camel@intel.com [2]
Link: https://lore.kernel.org/kvm/aip1eO7wvJDxKtBX@yzhao56-desk.sh.intel.com [3]
Link: https://lore.kernel.org/kvm/aNX6V6OSIwly1hu4@yzhao56-desk.sh.intel.com [4]
Link: https://lore.kernel.org/all/aYoS45AZNY0rUJQD@google.com [5]
Link: https://lore.kernel.org/all/aXzPIO2qZwuwaeLi@google.com [6]

This patch is based on earlier work from Xiaoyao Li, Isaku Yamahata, and
Kiryl Shutsemau.

Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- Rebased to DPAMT v11.
- Merged in DPAMT-related handling. (Sean).
- Only allow demotion when DPAMT is enabled. (Rick)
- Protect DEMOTE and pamt_refcount with dpamt_lock. (Yan)
- ENHANCE to ENHANCED in TDX_FEATURES0_ENHANCED_DEMOTE_INTERRUPTIBILITY.
  (Kai)
- tdx_supports_demote_nointerrupt() -->
  tdx_huge_page_demote_uninterruptible().  (Kai)
- Updated the patch log/SoB to match tip's preference.

v3:
- Use a var name that clearly tell that the page is used as a page table
  page. (Binbin).
- Check if TDX module supports feature ENHANCE_DEMOTE_INTERRUPTIBILITY.
  (Kai).

RFC v2:
- Refine the patch log (Rick).
- Do not handle TDX_INTERRUPTED_RESTARTABLE as the new TDX modules in
  planning do not check interrupts for basic TDX.

RFC v1:
- Rebased and split patch. Updated patch log.
---
 arch/x86/include/asm/tdx.h  |  9 ++++
 arch/x86/virt/vmx/tdx/tdx.c | 82 +++++++++++++++++++++++++++++++++++--
 arch/x86/virt/vmx/tdx/tdx.h |  1 +
 3 files changed, 89 insertions(+), 3 deletions(-)

diff --git a/arch/x86/include/asm/tdx.h b/arch/x86/include/asm/tdx.h
index e2790ac2c521..56549cc50651 100644
--- a/arch/x86/include/asm/tdx.h
+++ b/arch/x86/include/asm/tdx.h
@@ -37,6 +37,7 @@
 #define TDX_FEATURES0_TD_PRESERVING	BIT_ULL(1)
 #define TDX_FEATURES0_NO_RBP_MOD	BIT_ULL(18)
 #define TDX_FEATURES0_DYNAMIC_PAMT	BIT_ULL(36)
+#define TDX_FEATURES0_ENHANCED_DEMOTE_INTERRUPTIBILITY	BIT_ULL(51)
 
 #ifndef __ASSEMBLER__
 
@@ -119,6 +120,11 @@ static inline bool tdx_supports_runtime_update(const struct tdx_sys_info *sysinf
 	return sysinfo->features.tdx_features0 & TDX_FEATURES0_TD_PRESERVING;
 }
 
+static inline bool tdx_huge_page_demote_uninterruptible(const struct tdx_sys_info *sysinfo)
+{
+	return sysinfo->features.tdx_features0 & TDX_FEATURES0_ENHANCED_DEMOTE_INTERRUPTIBILITY;
+}
+
 bool tdx_supports_dynamic_pamt(const struct tdx_sys_info *sysinfo);
 
 /* Simple structure for pre-allocating DPAMT pages outside of spinlocks. */
@@ -184,6 +190,9 @@ u64 tdh_mng_key_config(struct tdx_td *td);
 u64 tdh_mng_create(struct tdx_td *td, u16 hkid);
 u64 tdh_vp_create(struct tdx_td *td, struct tdx_vp *vp);
 u64 tdh_mng_rd(struct tdx_td *td, u64 field, u64 *data);
+u64 tdh_mem_page_demote(struct tdx_td *td, u64 gpa, enum pg_level level, kvm_pfn_t pfn,
+			struct page *new_sept_pt, struct tdx_pamt_cache *pamt_cache,
+			u64 *ext_err1, u64 *ext_err2);
 u64 tdh_mr_extend(struct tdx_td *td, u64 gpa, u64 *ext_err1, u64 *ext_err2);
 u64 tdh_mr_finalize(struct tdx_td *td);
 u64 tdh_vp_flush(struct tdx_vp *vp);
diff --git a/arch/x86/virt/vmx/tdx/tdx.c b/arch/x86/virt/vmx/tdx/tdx.c
index 43f813afc5b1..1f121c24b9d9 100644
--- a/arch/x86/virt/vmx/tdx/tdx.c
+++ b/arch/x86/virt/vmx/tdx/tdx.c
@@ -75,6 +75,9 @@ static struct tdmr_info_list tdx_tdmr_list;
  */
 static atomic_t *dpamt_refcounts;
 
+/* Serializes adding/removing DPAMT memory */
+static DEFINE_SPINLOCK(dpamt_lock);
+
 /* All TDX-usable memory regions.  Protected by mem_hotplug_lock. */
 static LIST_HEAD(tdx_memlist);
 
@@ -82,6 +85,9 @@ static struct tdx_sys_info tdx_sysinfo;
 
 static DEFINE_RAW_SPINLOCK(sysinit_lock);
 
+static int alloc_pamt_array(struct page **pamt_pages, struct tdx_pamt_cache *cache);
+static void free_pamt_array(struct page **pamt_pages);
+
 /*
  * Do the module global initialization once and return its result.
  * It can be done on any cpu, and from task or IRQ context.
@@ -1827,6 +1833,79 @@ u64 tdh_mng_rd(struct tdx_td *td, u64 field, u64 *data)
 }
 EXPORT_SYMBOL_FOR_KVM(tdh_mng_rd);
 
+static u64 __tdh_mem_page_demote(struct tdx_td *td, u64 gpa, enum pg_level level,
+				 struct page *new_sept_pt, struct page **dpamt_pages,
+				 u64 *ext_err1, u64 *ext_err2)
+{
+	struct tdx_module_args args = {
+		.rcx = gpa | pg_level_to_tdx_sept_level(level),
+		.rdx = tdx_tdr_pa(td),
+		.r8 = page_to_phys(new_sept_pt),
+		.r12 = page_to_phys(dpamt_pages[0]),
+		.r13 = page_to_phys(dpamt_pages[1]),
+	};
+	u64 ret;
+
+	ret = seamcall_saved_ret(TDH_MEM_PAGE_DEMOTE, &args);
+
+	*ext_err1 = args.rcx;
+	*ext_err2 = args.rdx;
+	return ret;
+}
+
+u64 tdh_mem_page_demote(struct tdx_td *td, u64 gpa, enum pg_level level, kvm_pfn_t pfn,
+			struct page *new_sept_pt, struct tdx_pamt_cache *pamt_cache,
+			u64 *ext_err1, u64 *ext_err2)
+{
+	atomic_t *dpamt_refcount = tdx_find_dpamt_refcount(pfn);
+	struct page *dpamt_pages[TDX_DPAMT_ENTRY_PAGE_CNT];
+	u64 ret;
+
+	*ext_err1 = 0;
+	*ext_err2 = 0;
+
+	if (WARN_ON_ONCE(!tdx_huge_page_demote_uninterruptible(&tdx_sysinfo) ||
+			 !tdx_supports_dynamic_pamt(&tdx_sysinfo) ||
+			 level != PG_LEVEL_2M))
+		return TDX_SW_ERROR;
+
+	if (WARN_ON_ONCE(alloc_pamt_array(dpamt_pages, pamt_cache)))
+		return TDX_SW_ERROR;
+
+	/*
+	 * The dpamt_lock is used to avoid contention between the
+	 * TDH_MEM_PAGE_DEMOTE and TDH_PHYMEM_PAMT_{ADD/REMOVE} SEAMCALLs.
+	 */
+	spin_lock(&dpamt_lock);
+
+	/*
+	 * The caller ensures that DEMOTE is only invoked on a 2MB mapping, so
+	 * dpamt_refcount must be 0.
+	 */
+	if (WARN_ON_ONCE(atomic_read(dpamt_refcount))) {
+		ret = TDX_SW_ERROR;
+		goto out_free;
+	}
+
+	ret = __tdh_mem_page_demote(td, gpa, level, new_sept_pt, dpamt_pages,
+				    ext_err1, ext_err2);
+	if (ret != TDX_SUCCESS) {
+		WARN_ON_ONCE((ret & TDX_SEAMCALL_STATUS_MASK) != TDX_OPERAND_BUSY);
+		goto out_free;
+	}
+
+	atomic_set(dpamt_refcount, PMD_SIZE / PAGE_SIZE);
+	spin_unlock(&dpamt_lock);
+
+	return TDX_SUCCESS;
+
+out_free:
+	spin_unlock(&dpamt_lock);
+	free_pamt_array(dpamt_pages);
+	return ret;
+}
+EXPORT_SYMBOL_FOR_KVM(tdh_mem_page_demote);
+
 u64 tdh_mr_extend(struct tdx_td *td, u64 gpa, u64 *ext_err1, u64 *ext_err2)
 {
 	struct tdx_module_args args = {
@@ -2108,9 +2187,6 @@ static u64 tdh_phymem_pamt_remove(kvm_pfn_t pfn, struct page **pamt_pages)
 	return 0;
 }
 
-/* Serializes adding/removing DPAMT memory */
-static DEFINE_SPINLOCK(dpamt_lock);
-
 /* Bump DPAMT refcount for the given pfn and allocate DPAMT backing if needed. */
 int tdx_pamt_get(kvm_pfn_t pfn, enum pg_level level, struct tdx_pamt_cache *cache)
 {
diff --git a/arch/x86/virt/vmx/tdx/tdx.h b/arch/x86/virt/vmx/tdx/tdx.h
index 94b2333e5f7e..8157f7b954d3 100644
--- a/arch/x86/virt/vmx/tdx/tdx.h
+++ b/arch/x86/virt/vmx/tdx/tdx.h
@@ -24,6 +24,7 @@
 #define TDH_MNG_KEY_CONFIG		8
 #define TDH_MNG_CREATE			9
 #define TDH_MNG_RD			11
+#define TDH_MEM_PAGE_DEMOTE		15
 #define TDH_MR_EXTEND			16
 #define TDH_MR_FINALIZE			17
 #define TDH_VP_FLUSH			18
-- 
2.43.2


  parent reply	other threads:[~2026-09-28  9:09 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-28  9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
2026-09-28  9:08 ` [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages Yan Zhao
2026-09-28  9:08 ` Yan Zhao [this message]
2026-09-28  9:08 ` [PATCH v4 03/17] KVM: TDX: Reset private huge pages after S-EPT page removal Yan Zhao
2026-09-28  9:09 ` [PATCH v4 04/17] KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault path Yan Zhao
2026-09-28  9:09 ` [PATCH v4 05/17] KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table splitting Yan Zhao
2026-09-28  9:10 ` [PATCH v4 06/17] KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split Yan Zhao
2026-09-28  9:10 ` [PATCH v4 07/17] KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings to 4KB Yan Zhao
2026-09-28  9:10 ` [PATCH v4 08/17] KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting S-EPT Yan Zhao
2026-09-28  9:10 ` [PATCH v4 09/17] KVM: x86/mmu: Introduce hugepage_set_guest_inhibit() Yan Zhao
2026-09-28  9:11 ` [PATCH v4 10/17] KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots Yan Zhao
2026-09-28  9:11 ` [PATCH v4 11/17] KVM: TDX: Honor the guest's accept level contained in an EPT violation Yan Zhao
2026-09-28  9:11 ` [PATCH v4 12/17] KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU context Yan Zhao
2026-09-28  9:11 ` [PATCH v4 13/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns give gmem+slot+pgoff Yan Zhao
2026-09-28  9:11 ` [PATCH v4 14/17] [GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start() to start() and zap() Yan Zhao
2026-09-28  9:12 ` [PATCH v4 15/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap() Yan Zhao
2026-09-28  9:12 ` [PATCH v4 16/17] [GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split S-EPT Yan Zhao
2026-09-28  9:12 ` [PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M Yan Zhao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260928090831.15502-1-yan.y.zhao@intel.com \
    --to=yan.y.zhao@intel.com \
    --cc=ackerleytng@google.com \
    --cc=binbin.wu@linux.intel.com \
    --cc=chao.gao@intel.com \
    --cc=chao.p.peng@intel.com \
    --cc=dave.hansen@intel.com \
    --cc=david@kernel.org \
    --cc=fan.du@intel.com \
    --cc=farrah.chen@intel.com \
    --cc=francescolavra.fl@gmail.com \
    --cc=jgross@suse.com \
    --cc=jun.miao@intel.com \
    --cc=kai.huang@intel.com \
    --cc=kas@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=michael.roth@amd.com \
    --cc=nik.borisov@suse.com \
    --cc=pbonzini@redhat.com \
    --cc=pgonda@google.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=sagis@google.com \
    --cc=seanjc@google.com \
    --cc=tabba@google.com \
    --cc=thomas.lendacky@amd.com \
    --cc=vannapurve@google.com \
    --cc=vbabka@suse.cz \
    --cc=x86@kernel.org \
    --cc=xiaoyao.li@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®