* [PATCH v4 00/17] KVM: TDX huge page support for private memory
@ 2026-09-28 9:07 Yan Zhao
2026-09-28 9:08 ` [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages Yan Zhao
` (16 more replies)
0 siblings, 17 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:07 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
This is v4 of the TDX huge page series. This revision incorporates changes
from Sean's combined DPAMT + Huge Pages series [1], addresses review
feedback from v3 [0], fixes bugs, refines commit logs and code comments,
isolates gmem-dependent glue code, and aligns SoB tags with tip tree
preference.
As with prior versions, TDX huge page support requires the TDX module to
report the ENHANCED_DEMOTE_INTERRUPTIBILITY feature. Additionally, this
revision enables TDX huge pages only when both DPAMT and guest_memfd
in-place conversion are enabled. (See the 'Base' section below for the
dependency stack. A working branch is at [2]).
TDX huge page support relies on guest_memfd (gmem) as the backing huge page
allocator. Pending the availability of upstream HugeTLB-based gmem, an
out-of-tree allocator allocating 2MB folios from the buddy is used for
testing. However, the core of this TDX huge page series is gmem-agnostic,
independent of the specific gmem allocator.
In previous upstream planning discussions [3], Sean noted openness to
landing this series (e.g., via a topic branch) ahead of in-tree huge page
support in guest_memfd. To facilitate this, the gmem-dependent glue code is
isolated into 4 patches tagged "GMEM-DEPENDENT" at the tail of this series.
The primary goals of this revision are to:
- Collect Reviewed-by tags (especially for tip-tree patches 1-2).
- Seek feedback and alignment from Sean on the following KVM/gmem changes:
* Adjusting the topup count of page pairs in PAMT cache for splitting
S-EPT entries (patch 8).
* Splitting guest_memfd invalidation into stages: invalidate_start,
prezap, zap, and invalidate_end (patches 14-15).
* Renaming the hook .gmem_convert() to .gmem_prezap() (referencing prior
discussions at [4]) (patch 15).
Earlier settled items
---------------------
1. TDX huge pages series depends on DPAMT series, and patches the TDX huge
page series should not be divided based on whether they are
DPAMT-related.
2. Only enable TDX huge pages when gmem in-place conversion is enabled.
3. Only enable TDX huge pages when the TDX module supports the
ENHANCED_DEMOTE_INTERRUPTIBILITY feature.
4. Do not support page merging for the initial huge page support.
5. Hook splitting S-EPT entries to .set_external_spte(), and use PFN
instead of folio/page to manage guest pages.
6. Handle the installation and uninstallation of external_spt's DPAMT pages
when mapping, splitting, and unmapping S-EPT entries, rather than
during page allocation and freeing.
7. Reuse the .topup_external_cache() hook to allocate PAMT pages for
splitting, and use a mutex instead of a spinlock to protect the per-VM
PAMT cache.
8. Add and export an API kvm_tdp_mmu_mirrors_split_huge_pages() to split
mirror page tables for a given range.
9. Honor the guest's accept level by splitting any existing mapping higher
than the guest's accept level, setting the guest inhibit bit after
performing the split, and never clearing the flag. Take mmu_lock for
write during the process to keep things simple.
10. Have gmem invoke arch code, which further invokes a gmem hook to
trigger S-EPT entry splitting under a non-vCPU context. (Note: the gmem
hook name and invocation sequence under a specific gmem implementation
are not yet finalized.)
Main changes from v3
--------------------
1. Reorganized patches so that they are no longer divided by whether they
are DPAMT-related.
2. Hooked to .set_external_spte() for splitting.
3. Added a dedicated API kvm_tdp_mmu_mirrors_split_huge_pages() for
splitting mirror roots.
4. Reused the .topup_external_cache() hook to top up DPAMT pages used
during S-EPT entry splitting. Extended the hook to work under both vCPU
and non-vCPU contexts. Used a mutex instead of a spinlock to protect
the per-VM PAMT cache under a non-vCPU context.
5. Had gmem trigger S-EPT splitting via the .gmem_prezap() hook.
6. Moved considerations of not supporting promotion from the cover letter
to the patch log.
7. Dropped patches for working with per-VM memory attributes. Only enabled
huge pages when gmem in-place conversions are enabled.
8. Based TDX huge pages on DPAMT. Simplified the DEMOTE SEAMCALL
implementation by only enabling TDX huge pages when DPAMT is enabled.
9. Split out gmem glue code into patches tagged with "GMEM-DEPENDENT".
10. Fixed bugs:
- adjusted the topup PAMT page pair count for splitting.
- used a global spinlock to avoid contention between DEMOTE and
PAMT.{ADD/REMOVE} SEAMCALLs.
11. Dropped cache-flush-related handling.
Patches Layout
--------------
Patches 1-2: Enhance/Introduce SEAMCALL wrappers/helpers to support huge
pages. (tip tree patches)
Patches 3-4: Enhance private pages reset and disallow page merging in the
mirror page table (KVM patches).
Patches 5-11: Support for splitting under vCPU context. (KVM patches)
Patches 12-16: Support for splitting under non-vCPU context. (KVM patches)
Patch 12: Core support to split S-EPT under non-vCPU context.
Patches 13-14: Gmem refactors.
Patches 15-16: Have gmem trigger splitting under a non-vCPU
context.
Patch 17: Turns on TDX huge page. (KVM patch)
Base
----
This revision is based on kvm-x86-next-2026.09.22 (containing gmem stop
returning struct page in pfn lookup series [7] and gmem in-place conversion
series [8]), plus the following series in the stack (refer to branch [2]):
- DPAMT v11 [5].
- Drop unneeded cache flushing [6]
- gmem 2MB buddy [9] + an alignment fix and an out-of-place conversion
workaround. *
*: For testing purposes, this revision uses the out-of-tree gmem 2MB buddy
solution as the huge page allocator.
The gmem 2MB buddy allocator allocates 2MB folios from the buddy for
private memory, while the shared memory is allocated from a different
backend (specified by HVA). To avoid fragmentation, private-to-shared
conversions only split private mappings without splitting or reclaiming
the backing 2MB folios. The 2MB folios are always retained in the gmem
inode filemap cache without splitting even though part of them are not
mapped as private memory.
Since the shared memory is not allocated from gmem (ensured by having
gmem 2MB buddy allocator mode intentionally incompatible with the MMAP
flag), commit 640d75f52828 ("KVM: guest_memfd: Always fault from
guest_memfd if in-place conversion is enabled") is reverted in the
testing branch [2] as a workaround to allow out-of-place conversions.
The gmem 2MB buddy solution and the workaround to allow out-of-place
conversions do not affect the code of the TDX private huge page series.
Both of them can be dropped once the final HugeTLB-based gmem is
available (with some rebasing effort for GMEM-DEPENDENT patches).
Testing
-------
A TDX module enumerating the ENHANCED_DEMOTE_INTERRUPTIBILITY feature is
required (AFAIK, TDX module 1.5.28 is the earliest version that reports
this feature; For testing on versions prior to 1.5.28, a workaround patch
to force enable TDX huge page can be applied [10]).
To test TDX huge pages, set the following module parameters at boot/load:
kvm.gmem_in_place_conversion=on
kvm.gmem_2MB_private_mem_from_buddy_for_testing=on
kvm_intel.tdx=on
kvm_intel.tdx_huge_page=on
The count of active 2MB mappings can be verified via:
/sys/kernel/debug/kvm/pages_2m
QEMU and TDX KVM selftests enhanced with in-place conversion uABI support
are required.
Farrah Chen tested this stack on EMR platform with TDX module version
1.5.41.00.1153:
- 2MB huge pages were allocated and mapped as expected.
- Stress testing running 5 concurrent TDs alongside 5 standard VMs
completed successfully with no issues observed.
- Stress testing running 63 concurrent TDs completed successfully with no
issues observed.
Thanks
Yan
[0] v3: https://lore.kernel.org/all/20260106101646.24809-1-yan.y.zhao@intel.com
[1] DPAMT + Hugepage: https://lore.kernel.org/all/20260129011517.3545883-1-seanjc@google.com
[2] A working branch: https://github.com/intel-staging/tdx/tree/huge_page_v4
[3] Upstream planning: https://lore.kernel.org/all/aYS85EsXu_xuQXSI@google.com
[4] gmem_convert() discussion: https://lore.kernel.org/all/anLrGmbZwgWnNUkp@yzhao56-desk.sh.intel.com
[5] DPAMT: https://lore.kernel.org/all/20260904215841.303070-1-rick.p.edgecombe@intel.com/
[6] drop cache flush: https://lore.kernel.org/all/20260922205215.870563-1-rick.p.edgecombe@intel.com
[7] gmem pfn lookup: https://lore.kernel.org/all/20260826-gmem-no-return-page-v4-0-3bb9c1ddb4e3@google.com
[8] gmem in-place conversion: https://lore.kernel.org/all/20260910-gmem-inplace-conversion-v13-0-dd6fbf94f4e1@google.com
[9] gmem 2MB buddy: https://github.com/intel-staging/tdx/commit/5ba004e1b4af4d928acb7fde006195bbe84035a3
[10] workaround on old TDX module: https://github.com/intel-staging/tdx/commit/f1fe677b82e11fc97679a04e078b2fd05da55b34
Isaku Yamahata (1):
KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table
splitting
Rick Edgecombe (1):
KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault
path
Sean Christopherson (4):
KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split
KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots
[GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns
give gmem+slot+pgoff
[GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap()
Yan Zhao (11):
x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages
x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page
KVM: TDX: Reset private huge pages after S-EPT page removal
KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings
to 4KB
KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting
S-EPT
KVM: x86/mmu: Introduce hugepage_set_guest_inhibit()
KVM: TDX: Honor the guest's accept level contained in an EPT violation
KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU
context
[GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start()
to start() and zap()
[GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split
S-EPT
KVM: TDX: Turn on PG_LEVEL_2M
arch/x86/include/asm/kvm-x86-ops.h | 3 +
arch/x86/include/asm/kvm_host.h | 18 +-
arch/x86/include/asm/tdx.h | 13 +-
arch/x86/kvm/Kconfig | 1 +
arch/x86/kvm/mmu.h | 4 +
arch/x86/kvm/mmu/mmu.c | 28 ++-
arch/x86/kvm/mmu/tdp_mmu.c | 79 +++++--
arch/x86/kvm/mmu/tdp_mmu.h | 2 +
arch/x86/kvm/vmx/tdx.c | 322 +++++++++++++++++++++++++++--
arch/x86/kvm/vmx/tdx.h | 5 +
arch/x86/kvm/vmx/tdx_arch.h | 3 +
arch/x86/kvm/x86.c | 8 +
arch/x86/virt/vmx/tdx/tdx.c | 94 ++++++++-
arch/x86/virt/vmx/tdx/tdx.h | 1 +
include/linux/kvm_host.h | 5 +
include/linux/kvm_types.h | 1 +
virt/kvm/Kconfig | 4 +
virt/kvm/guest_memfd.c | 135 +++++++++++-
18 files changed, 662 insertions(+), 64 deletions(-)
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
@ 2026-09-28 9:08 ` Yan Zhao
2026-09-28 9:08 ` [PATCH v4 02/17] x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page Yan Zhao
` (15 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:08 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
The TDX module uses PAMT to track metadata for every physical page (within
TDMR) across three levels: 1GB, 2MB, and 4KB. The VMM is responsible for
providing backing pages for the PAMT. Without Dynamic PAMT (DPAMT), all
PAMT backing pages are allocated statically at system boot time. With DPAMT
enabled, PAMT backing pages tracking physical pages at the 4KB level
(referred to as DPAMT pages) are dynamically installed and uninstalled at
runtime.
Currently, the VMM invokes tdx_pamt_get()/tdx_pamt_put() to install and
uninstall DPAMT pages under the assumption that all physical pages are
tracked at 4KB level. However, when a guest page is mapped as a huge page
in the S-EPT, the TDX module tracks it directly at the huge page level,
meaning installation and uninstallation of DPAMT pages for huge pages is
neither needed nor applicable.
Add a "level" parameter to tdx_pamt_get()/tdx_pamt_put() so the helpers
can skip DPAMT page installation and uninstallation for huge pages.
TD control pages (including S-EPT page table pages) are managed by the TDX
module at the 4KB level. Therefore, pass PG_LEVEL_4K to
tdx_pamt_get()/tdx_pamt_put() for those pages.
This patch is based on earlier work from Kiryl Shutsemau.
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
arch/x86/include/asm/tdx.h | 4 ++--
arch/x86/kvm/vmx/tdx.c | 14 +++++++++-----
arch/x86/virt/vmx/tdx/tdx.c | 12 ++++++------
3 files changed, 17 insertions(+), 13 deletions(-)
diff --git a/arch/x86/include/asm/tdx.h b/arch/x86/include/asm/tdx.h
index 48efd3eac375..e2790ac2c521 100644
--- a/arch/x86/include/asm/tdx.h
+++ b/arch/x86/include/asm/tdx.h
@@ -135,8 +135,8 @@ static inline void tdx_init_pamt_cache(struct tdx_pamt_cache *cache)
void tdx_free_pamt_cache(struct tdx_pamt_cache *cache);
int tdx_topup_pamt_cache(struct tdx_pamt_cache *cache, unsigned long npages);
-int tdx_pamt_get(kvm_pfn_t pfn, struct tdx_pamt_cache *cache);
-void tdx_pamt_put(kvm_pfn_t pfn);
+int tdx_pamt_get(kvm_pfn_t pfn, enum pg_level level, struct tdx_pamt_cache *cache);
+void tdx_pamt_put(kvm_pfn_t pfn, enum pg_level level);
int tdx_guest_keyid_alloc(void);
u32 tdx_get_nr_guest_keyids(void);
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 1f0dba322640..410c149bf776 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -1710,14 +1710,18 @@ static int tdx_sept_map_nonleaf_spte(struct kvm *kvm, gfn_t gfn,
if (!sept_pt)
return -EIO;
- ret = tdx_pamt_get(page_to_pfn(sept_pt), &to_tdx(vcpu)->pamt_cache);
+ /*
+ * An S-EPT page table page is tracked by the TDX module at 4KB level,
+ * so pass 4KB instead of the target S-EPT level.
+ */
+ ret = tdx_pamt_get(page_to_pfn(sept_pt), PG_LEVEL_4K, &to_tdx(vcpu)->pamt_cache);
if (KVM_BUG_ON(ret, kvm))
return ret;
err = tdh_mem_sept_add(&to_kvm_tdx(kvm)->td, gpa, level, sept_pt,
&entry, &level_state);
if (err)
- tdx_pamt_put(page_to_pfn(sept_pt));
+ tdx_pamt_put(page_to_pfn(sept_pt), PG_LEVEL_4K);
if (unlikely(tdx_operand_busy(err)))
return -EBUSY;
@@ -1745,7 +1749,7 @@ static int tdx_sept_map_leaf_spte(struct kvm *kvm, gfn_t gfn, enum pg_level leve
WARN_ON_ONCE((new_spte & VMX_EPT_RWX_MASK) != VMX_EPT_RWX_MASK);
- ret = tdx_pamt_get(pfn, &to_tdx(vcpu)->pamt_cache);
+ ret = tdx_pamt_get(pfn, level, &to_tdx(vcpu)->pamt_cache);
if (KVM_BUG_ON(ret, kvm))
return ret;
@@ -1767,7 +1771,7 @@ static int tdx_sept_map_leaf_spte(struct kvm *kvm, gfn_t gfn, enum pg_level leve
ret = tdx_mem_page_add(kvm, gfn, level, pfn);
if (ret)
- tdx_pamt_put(pfn);
+ tdx_pamt_put(pfn, level);
return ret;
}
@@ -1862,7 +1866,7 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
return -EIO;
tdx_quirk_reset_paddr(PFN_PHYS(pfn), PAGE_SIZE);
- tdx_pamt_put(pfn);
+ tdx_pamt_put(pfn, level);
return 0;
}
diff --git a/arch/x86/virt/vmx/tdx/tdx.c b/arch/x86/virt/vmx/tdx/tdx.c
index 33253869e2fd..43f813afc5b1 100644
--- a/arch/x86/virt/vmx/tdx/tdx.c
+++ b/arch/x86/virt/vmx/tdx/tdx.c
@@ -2112,14 +2112,14 @@ static u64 tdh_phymem_pamt_remove(kvm_pfn_t pfn, struct page **pamt_pages)
static DEFINE_SPINLOCK(dpamt_lock);
/* Bump DPAMT refcount for the given pfn and allocate DPAMT backing if needed. */
-int tdx_pamt_get(kvm_pfn_t pfn, struct tdx_pamt_cache *cache)
+int tdx_pamt_get(kvm_pfn_t pfn, enum pg_level level, struct tdx_pamt_cache *cache)
{
struct page *pamt_pages[TDX_DPAMT_ENTRY_PAGE_CNT];
atomic_t *dpamt_refcount;
u64 tdx_status;
int ret;
- if (!tdx_supports_dynamic_pamt(&tdx_sysinfo))
+ if (!tdx_supports_dynamic_pamt(&tdx_sysinfo) || level != PG_LEVEL_4K)
return 0;
ret = alloc_pamt_array(pamt_pages, cache);
@@ -2157,13 +2157,13 @@ int tdx_pamt_get(kvm_pfn_t pfn, struct tdx_pamt_cache *cache)
EXPORT_SYMBOL_FOR_KVM(tdx_pamt_get);
/* Drop DPAMT refcount for the given pfn and free DPAMT backing if needed. */
-void tdx_pamt_put(kvm_pfn_t pfn)
+void tdx_pamt_put(kvm_pfn_t pfn, enum pg_level level)
{
struct page *pamt_pages[TDX_DPAMT_ENTRY_PAGE_CNT] = {};
atomic_t *dpamt_refcount;
u64 tdx_status;
- if (!tdx_supports_dynamic_pamt(&tdx_sysinfo))
+ if (!tdx_supports_dynamic_pamt(&tdx_sysinfo) || level != PG_LEVEL_4K)
return;
dpamt_refcount = tdx_find_dpamt_refcount(pfn);
@@ -2243,7 +2243,7 @@ struct page *tdx_alloc_control_page(void)
if (!page)
return NULL;
- if (tdx_pamt_get(page_to_pfn(page), NULL)) {
+ if (tdx_pamt_get(page_to_pfn(page), PG_LEVEL_4K, NULL)) {
__free_page(page);
return NULL;
}
@@ -2261,7 +2261,7 @@ void tdx_free_control_page(struct page *page)
if (!page)
return;
- tdx_pamt_put(page_to_pfn(page));
+ tdx_pamt_put(page_to_pfn(page), PG_LEVEL_4K);
__free_page(page);
}
EXPORT_SYMBOL_FOR_KVM(tdx_free_control_page);
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 02/17] x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
2026-09-28 9:08 ` [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages Yan Zhao
@ 2026-09-28 9:08 ` Yan Zhao
2026-09-28 9:08 ` [PATCH v4 03/17] KVM: TDX: Reset private huge pages after S-EPT page removal Yan Zhao
` (14 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:08 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
tl;dr: Introduce a SEAMCALL wrapper to invoke the TDH_MEM_PAGE_DEMOTE
SEAMCALL for splitting a 2MB page mapped in S-EPT, and export it as an API.
The following assumptions are made by this wrapper:
1. The huge page is mapped at 2MB level in the S-EPT before invoking this
wrapper.
2. DPAMT is enabled.
3. The TDX module supports uninterruptible demote.
The only expected error returned from this wrapper is TDX_OPERAND_BUSY.
Callers may need to kick vCPUs out of guest mode to avoid this error.
Callers must provide a pamt_cache for drawing DPAMT pages for guest memory
and ensure that pages can be drawn from the pamt_cache locklessly in this
wrapper.
Callers must ensure no concurrent invocations of this wrapper on the same
GPA.
Long version:
Huge pages offer improved TLB efficiency, faster page walks, better memory
contiguity, and reduced page fault and memory overhead (fewer intermediate
page table structures, and no 4KB DPAMT pages). However, splitting/demoting
a huge page is required in the following cases:
- when the guest accepts the private memory at a lower level,
- when part of the huge page is converted from private to shared, or
- when VMM wants to reclaim part of the huge page.
The TDH_MEM_PAGE_DEMOTE SEAMCALL demotes a guest private huge page by
locating the corresponding leaf huge S-EPT entry and replacing it with a
non-leaf S-EPT entry pointing to a new page table page, which holds the
split leaf S-EPT entries. When DPAMT is enabled, upon successfully demoting
a 2MB page to 4KB pages, the SEAMCALL performs DPAMT pages installation for
the 2MB physical memory range, similar to the DPAMT pages installation
performed by the TDH_PHYMEM_PAMT_ADD SEAMCALL.
The wrapper args "gpa", "level" specify the original huge leaf S-EPT entry;
"new_sept_pt" specifies the newly added page table page; "pamt_cache"
provides a pair of pages to gift to the TDX module as DPAMT pages; "pfn" is
used to locate the dpamt_refcount; and the output args "ext_err1" and
"ext_err2" allow the caller to retrieve SEAMCALL failure information.
Three assumptions are made by this wrapper (as in "tl;dr").
Assumption 1 holds because the VMM can only create mappings up to 2MB level
in the S-EPT via the TDH_MEM_PAGE_AUG SEAMCALL. 1GB mappings in the S-EPT
are only possible via the TDH_MEM_PAGE_PROMOTE SEAMCALL, which is not yet
supported.
Assumptions 1 and 2 together simplify the wrapper implementation by
eliminating the need to check whether DPAMT-related handling is required.
Assumption 3 is a tradeoff reached after research and discussion: a TDX
module that does not support uninterruptible demote returns the error
TDX_INTERRUPTED_RESTARTABLE if any pending host interrupts are detected
during the TDH_MEM_PAGE_DEMOTE SEAMCALL. Since the TDX module cannot
guarantee a maximum retry count to ensure forward progress of the demotion,
interrupt storms could result in a DoS if the host retries endlessly.
Disabling interrupts before invoking the TDH_MEM_PAGE_DEMOTE SEAMCALL also
does not work, as the TDX module also checks for pending NMIs. Therefore,
the tradeoff is to disable huge pages when the TDX module does not support
uninterruptible demote. This is acceptable for basic TDX huge page
enabling, since the SEAMCALL execution time under uninterruptible demote
mode remains reasonable [1][2].
Later patches in KVM will enforce theses assumptions by setting the maximum
mapping level to 2MB, disallowing page merging, and disabling TDX huge
pages if assumption 2 or 3 is not met. This wrapper warns and returns
TDX_SW_ERROR if any assumption is violated [6].
Due to assumptions 1 and 2, the wrapper always performs DPAMT-related
handling before and after invoking the core helper __tdh_mem_page_demote():
- Draw a pair of pages from pamt_cache and gift them to the TDX module for
DPAMT pages installation. The pamt_cache is a list of pre-allocated pages
provided by the caller. The caller ensures that drawing pages from the
list is contention-free, either by using a per-vCPU thread-local list or
by holding a per-VM lock. If the SEAMCALL fails, the pair of pages is
freed directly rather than being re-inserted into the list, for
simplicity.
- Use the global spinlock dpamt_lock to avoid potential contention between
the TDH_MEM_PAGE_DEMOTE and TDH_PHYMEM_PAMT_{ADD/REMOVE} SEAMCALLs [3].
The contention here can be cross-VM, similar to that between
TDH_PHYMEM_PAMT_ADD and TDH_PHYMEM_PAMT_{ADD/REMOVE} [4]. dpamt_lock is
acquired directly without any pre-testing, since a 2MB huge page mapped
at the 2MB level in the S-EPT must have no existing dpamt_refcount before
demotion unless there is a bug. For defensive programming, warn and
return an error if dpamt_refcount is non-zero after acquiring dpamt_lock.
- Set dpamt_refcount to 512 upon successful demotion so that future
invocations of tdx_pamt_put() when removing the split 4KB pages work
correctly.
Note: The new_sept_pt page, which will be added as the S-EPT page table
page upon successful demotion, also requires DPAMT pages installation
before invoking the TDH_MEM_PAGE_DEMOTE SEAMCALL. This wrapper relies
on the caller to invoke tdx_pamt_get() separately for the new_sept_pt
page prior to calling this wrapper.
Additionally, caller must ensure no concurrent invocations of this wrapper
on the same GPA (for example, by holding KVM's write mmu_lock). Otherwise,
a second invocation on the same GPA may encounter an incorrect
dpamt_refcount or a SEAMCALL error caused by attempting to have the TDX
module demote a 4KB page. This not only simplifies the wrapper
implementation, but also prevents potential issues in the caller (for
example, avoiding a level mismatch between the KVM mirror EPT entry and the
corresponding S-EPT entry).
The only expected error returned from this wrapper is TDX_OPERAND_BUSY,
caused by potential contention between the TDH_MEM_PAGE_DEMOTE SEAMCALL and
TDH_VP_ENTER SEAMALL or guest TDCALLs operating on the S-EPT entry being
split, due to lack of lock protection. This error is returned directly to
the caller without retrying; the caller may handle it by kicking vCPUs and
preventing them from re-entering guest mode to resolve the contention.
For defensive programming, the wrapper still takes arg "level" (though it's
now always expected to be 2MB); a warning is issued before returning any
unexpected errors, regardless of whether the caller would issue a similar
warning [5][6].
Link: https://lore.kernel.org/kvm/99f5585d759328db973403be0713f68e492b492a.camel@intel.com [1]
Link: https://lore.kernel.org/all/fbf04b09f13bc2ce004ac97ee9c1f2c965f44fdf.camel@intel.com [2]
Link: https://lore.kernel.org/kvm/aip1eO7wvJDxKtBX@yzhao56-desk.sh.intel.com [3]
Link: https://lore.kernel.org/kvm/aNX6V6OSIwly1hu4@yzhao56-desk.sh.intel.com [4]
Link: https://lore.kernel.org/all/aYoS45AZNY0rUJQD@google.com [5]
Link: https://lore.kernel.org/all/aXzPIO2qZwuwaeLi@google.com [6]
This patch is based on earlier work from Xiaoyao Li, Isaku Yamahata, and
Kiryl Shutsemau.
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- Rebased to DPAMT v11.
- Merged in DPAMT-related handling. (Sean).
- Only allow demotion when DPAMT is enabled. (Rick)
- Protect DEMOTE and pamt_refcount with dpamt_lock. (Yan)
- ENHANCE to ENHANCED in TDX_FEATURES0_ENHANCED_DEMOTE_INTERRUPTIBILITY.
(Kai)
- tdx_supports_demote_nointerrupt() -->
tdx_huge_page_demote_uninterruptible(). (Kai)
- Updated the patch log/SoB to match tip's preference.
v3:
- Use a var name that clearly tell that the page is used as a page table
page. (Binbin).
- Check if TDX module supports feature ENHANCE_DEMOTE_INTERRUPTIBILITY.
(Kai).
RFC v2:
- Refine the patch log (Rick).
- Do not handle TDX_INTERRUPTED_RESTARTABLE as the new TDX modules in
planning do not check interrupts for basic TDX.
RFC v1:
- Rebased and split patch. Updated patch log.
---
arch/x86/include/asm/tdx.h | 9 ++++
arch/x86/virt/vmx/tdx/tdx.c | 82 +++++++++++++++++++++++++++++++++++--
arch/x86/virt/vmx/tdx/tdx.h | 1 +
3 files changed, 89 insertions(+), 3 deletions(-)
diff --git a/arch/x86/include/asm/tdx.h b/arch/x86/include/asm/tdx.h
index e2790ac2c521..56549cc50651 100644
--- a/arch/x86/include/asm/tdx.h
+++ b/arch/x86/include/asm/tdx.h
@@ -37,6 +37,7 @@
#define TDX_FEATURES0_TD_PRESERVING BIT_ULL(1)
#define TDX_FEATURES0_NO_RBP_MOD BIT_ULL(18)
#define TDX_FEATURES0_DYNAMIC_PAMT BIT_ULL(36)
+#define TDX_FEATURES0_ENHANCED_DEMOTE_INTERRUPTIBILITY BIT_ULL(51)
#ifndef __ASSEMBLER__
@@ -119,6 +120,11 @@ static inline bool tdx_supports_runtime_update(const struct tdx_sys_info *sysinf
return sysinfo->features.tdx_features0 & TDX_FEATURES0_TD_PRESERVING;
}
+static inline bool tdx_huge_page_demote_uninterruptible(const struct tdx_sys_info *sysinfo)
+{
+ return sysinfo->features.tdx_features0 & TDX_FEATURES0_ENHANCED_DEMOTE_INTERRUPTIBILITY;
+}
+
bool tdx_supports_dynamic_pamt(const struct tdx_sys_info *sysinfo);
/* Simple structure for pre-allocating DPAMT pages outside of spinlocks. */
@@ -184,6 +190,9 @@ u64 tdh_mng_key_config(struct tdx_td *td);
u64 tdh_mng_create(struct tdx_td *td, u16 hkid);
u64 tdh_vp_create(struct tdx_td *td, struct tdx_vp *vp);
u64 tdh_mng_rd(struct tdx_td *td, u64 field, u64 *data);
+u64 tdh_mem_page_demote(struct tdx_td *td, u64 gpa, enum pg_level level, kvm_pfn_t pfn,
+ struct page *new_sept_pt, struct tdx_pamt_cache *pamt_cache,
+ u64 *ext_err1, u64 *ext_err2);
u64 tdh_mr_extend(struct tdx_td *td, u64 gpa, u64 *ext_err1, u64 *ext_err2);
u64 tdh_mr_finalize(struct tdx_td *td);
u64 tdh_vp_flush(struct tdx_vp *vp);
diff --git a/arch/x86/virt/vmx/tdx/tdx.c b/arch/x86/virt/vmx/tdx/tdx.c
index 43f813afc5b1..1f121c24b9d9 100644
--- a/arch/x86/virt/vmx/tdx/tdx.c
+++ b/arch/x86/virt/vmx/tdx/tdx.c
@@ -75,6 +75,9 @@ static struct tdmr_info_list tdx_tdmr_list;
*/
static atomic_t *dpamt_refcounts;
+/* Serializes adding/removing DPAMT memory */
+static DEFINE_SPINLOCK(dpamt_lock);
+
/* All TDX-usable memory regions. Protected by mem_hotplug_lock. */
static LIST_HEAD(tdx_memlist);
@@ -82,6 +85,9 @@ static struct tdx_sys_info tdx_sysinfo;
static DEFINE_RAW_SPINLOCK(sysinit_lock);
+static int alloc_pamt_array(struct page **pamt_pages, struct tdx_pamt_cache *cache);
+static void free_pamt_array(struct page **pamt_pages);
+
/*
* Do the module global initialization once and return its result.
* It can be done on any cpu, and from task or IRQ context.
@@ -1827,6 +1833,79 @@ u64 tdh_mng_rd(struct tdx_td *td, u64 field, u64 *data)
}
EXPORT_SYMBOL_FOR_KVM(tdh_mng_rd);
+static u64 __tdh_mem_page_demote(struct tdx_td *td, u64 gpa, enum pg_level level,
+ struct page *new_sept_pt, struct page **dpamt_pages,
+ u64 *ext_err1, u64 *ext_err2)
+{
+ struct tdx_module_args args = {
+ .rcx = gpa | pg_level_to_tdx_sept_level(level),
+ .rdx = tdx_tdr_pa(td),
+ .r8 = page_to_phys(new_sept_pt),
+ .r12 = page_to_phys(dpamt_pages[0]),
+ .r13 = page_to_phys(dpamt_pages[1]),
+ };
+ u64 ret;
+
+ ret = seamcall_saved_ret(TDH_MEM_PAGE_DEMOTE, &args);
+
+ *ext_err1 = args.rcx;
+ *ext_err2 = args.rdx;
+ return ret;
+}
+
+u64 tdh_mem_page_demote(struct tdx_td *td, u64 gpa, enum pg_level level, kvm_pfn_t pfn,
+ struct page *new_sept_pt, struct tdx_pamt_cache *pamt_cache,
+ u64 *ext_err1, u64 *ext_err2)
+{
+ atomic_t *dpamt_refcount = tdx_find_dpamt_refcount(pfn);
+ struct page *dpamt_pages[TDX_DPAMT_ENTRY_PAGE_CNT];
+ u64 ret;
+
+ *ext_err1 = 0;
+ *ext_err2 = 0;
+
+ if (WARN_ON_ONCE(!tdx_huge_page_demote_uninterruptible(&tdx_sysinfo) ||
+ !tdx_supports_dynamic_pamt(&tdx_sysinfo) ||
+ level != PG_LEVEL_2M))
+ return TDX_SW_ERROR;
+
+ if (WARN_ON_ONCE(alloc_pamt_array(dpamt_pages, pamt_cache)))
+ return TDX_SW_ERROR;
+
+ /*
+ * The dpamt_lock is used to avoid contention between the
+ * TDH_MEM_PAGE_DEMOTE and TDH_PHYMEM_PAMT_{ADD/REMOVE} SEAMCALLs.
+ */
+ spin_lock(&dpamt_lock);
+
+ /*
+ * The caller ensures that DEMOTE is only invoked on a 2MB mapping, so
+ * dpamt_refcount must be 0.
+ */
+ if (WARN_ON_ONCE(atomic_read(dpamt_refcount))) {
+ ret = TDX_SW_ERROR;
+ goto out_free;
+ }
+
+ ret = __tdh_mem_page_demote(td, gpa, level, new_sept_pt, dpamt_pages,
+ ext_err1, ext_err2);
+ if (ret != TDX_SUCCESS) {
+ WARN_ON_ONCE((ret & TDX_SEAMCALL_STATUS_MASK) != TDX_OPERAND_BUSY);
+ goto out_free;
+ }
+
+ atomic_set(dpamt_refcount, PMD_SIZE / PAGE_SIZE);
+ spin_unlock(&dpamt_lock);
+
+ return TDX_SUCCESS;
+
+out_free:
+ spin_unlock(&dpamt_lock);
+ free_pamt_array(dpamt_pages);
+ return ret;
+}
+EXPORT_SYMBOL_FOR_KVM(tdh_mem_page_demote);
+
u64 tdh_mr_extend(struct tdx_td *td, u64 gpa, u64 *ext_err1, u64 *ext_err2)
{
struct tdx_module_args args = {
@@ -2108,9 +2187,6 @@ static u64 tdh_phymem_pamt_remove(kvm_pfn_t pfn, struct page **pamt_pages)
return 0;
}
-/* Serializes adding/removing DPAMT memory */
-static DEFINE_SPINLOCK(dpamt_lock);
-
/* Bump DPAMT refcount for the given pfn and allocate DPAMT backing if needed. */
int tdx_pamt_get(kvm_pfn_t pfn, enum pg_level level, struct tdx_pamt_cache *cache)
{
diff --git a/arch/x86/virt/vmx/tdx/tdx.h b/arch/x86/virt/vmx/tdx/tdx.h
index 94b2333e5f7e..8157f7b954d3 100644
--- a/arch/x86/virt/vmx/tdx/tdx.h
+++ b/arch/x86/virt/vmx/tdx/tdx.h
@@ -24,6 +24,7 @@
#define TDH_MNG_KEY_CONFIG 8
#define TDH_MNG_CREATE 9
#define TDH_MNG_RD 11
+#define TDH_MEM_PAGE_DEMOTE 15
#define TDH_MR_EXTEND 16
#define TDH_MR_FINALIZE 17
#define TDH_VP_FLUSH 18
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 03/17] KVM: TDX: Reset private huge pages after S-EPT page removal
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
2026-09-28 9:08 ` [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages Yan Zhao
2026-09-28 9:08 ` [PATCH v4 02/17] x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page Yan Zhao
@ 2026-09-28 9:08 ` Yan Zhao
2026-09-28 9:09 ` [PATCH v4 04/17] KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault path Yan Zhao
` (13 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:08 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
"Resetting" TDX pages is a workaround for the TDX_PW_MCE erratum. Once a
leaf SPTE can map to a huge page in the S-EPT, the entire huge physical
range, rather than just the first 4KB, needs to be reset before the host
can reuse it.
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
arch/x86/kvm/vmx/tdx.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 410c149bf776..11792a490330 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -1865,7 +1865,7 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
if (TDX_BUG_ON_2(err, TDH_MEM_PAGE_REMOVE, entry, level_state, kvm))
return -EIO;
- tdx_quirk_reset_paddr(PFN_PHYS(pfn), PAGE_SIZE);
+ tdx_quirk_reset_paddr(PFN_PHYS(pfn), KVM_HPAGE_SIZE(level));
tdx_pamt_put(pfn, level);
return 0;
}
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 04/17] KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault path
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (2 preceding siblings ...)
2026-09-28 9:08 ` [PATCH v4 03/17] KVM: TDX: Reset private huge pages after S-EPT page removal Yan Zhao
@ 2026-09-28 9:09 ` Yan Zhao
2026-09-28 9:09 ` [PATCH v4 05/17] KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table splitting Yan Zhao
` (12 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:09 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
From: Rick Edgecombe <rick.p.edgecombe@intel.com>
Disallow huge page promotion in the TDP MMU for mirror roots as KVM doesn't
currently support promoting S-EPT entries due to the complexity incurred
by the TDX module's rules for huge page promotion.
- The current TDX module requires all 4KB leafs to be either all PENDING
or all ACCEPTED before a successful promotion to 2MB. This requirement
prevents successful page merging after partially converting a 2MB
range from private to shared and then back to private, which is the
primary scenario necessitating page promotion.
- The TDX module effectively requires a break-before-make sequence (to
satisfy its TLB flushing rules), i.e., creates a window of time where a
different vCPU can encounter faults on a SPTE that KVM is trying to
promote to a huge page. To avoid unexpected BUSY errors, KVM would need
to FREEZE the non-leaf SPTE before replacing it with a huge SPTE.
Disable huge page promotion for all map() operations, as supporting page
promotion when building the initial image is still non-trivial, and the
vast majority of images are ~4MB or less, i.e., the benefit of creating
huge pages during TD build time is minimal.
Signed-off-by: Rick Edgecombe <rick.p.edgecombe@intel.com>
[sean: check root, add comment, rewrite changelog]
Signed-off-by: Sean Christopherson <seanjc@google.com>
Co-developed-by: Yan Zhao <yan.y.zhao@intel.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
arch/x86/kvm/mmu/mmu.c | 3 ++-
arch/x86/kvm/mmu/tdp_mmu.c | 12 +++++++++++-
2 files changed, 13 insertions(+), 2 deletions(-)
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 266109d18536..a9a7c015256b 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -3533,7 +3533,8 @@ void disallowed_hugepage_adjust(struct kvm_page_fault *fault, u64 spte, int cur_
cur_level == fault->goal_level &&
is_shadow_present_pte(spte) &&
!is_large_pte(spte) &&
- spte_to_child_sp(spte)->nx_huge_page_disallowed) {
+ ((spte_to_child_sp(spte)->nx_huge_page_disallowed) ||
+ is_mirror_sp(spte_to_child_sp(spte)))) {
/*
* A small SPTE exists for this pfn, but FNAME(fetch),
* direct_map(), or kvm_tdp_mmu_map() would like to create a
diff --git a/arch/x86/kvm/mmu/tdp_mmu.c b/arch/x86/kvm/mmu/tdp_mmu.c
index 44dad106fad1..001449142d32 100644
--- a/arch/x86/kvm/mmu/tdp_mmu.c
+++ b/arch/x86/kvm/mmu/tdp_mmu.c
@@ -1233,7 +1233,17 @@ int kvm_tdp_mmu_map(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault)
for_each_tdp_pte(iter, kvm, root, fault->gfn, fault->gfn + 1) {
int r;
- if (fault->nx_huge_page_workaround_enabled)
+ /*
+ * Don't replace a page table (non-leaf) SPTE with a huge SPTE
+ * (a.k.a. hugepage promotion) if the NX hugepage workaround is
+ * enabled, as doing so will cause significant thrashing if one
+ * or more leaf SPTEs need to be executable.
+ *
+ * Disallow hugepage promotion for mirror roots as KVM doesn't
+ * (yet) support promoting S-EPT entries while holding mmu_lock
+ * for read (due to complexity induced by the TDX-Module APIs).
+ */
+ if (fault->nx_huge_page_workaround_enabled || is_mirror_sp(root))
disallowed_hugepage_adjust(fault, iter.old_spte, iter.level);
/*
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 05/17] KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table splitting
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (3 preceding siblings ...)
2026-09-28 9:09 ` [PATCH v4 04/17] KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault path Yan Zhao
@ 2026-09-28 9:09 ` Yan Zhao
2026-09-28 9:10 ` [PATCH v4 06/17] KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split Yan Zhao
` (11 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:09 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
From: Isaku Yamahata <isaku.yamahata@intel.com>
Enhance tdp_mmu_alloc_sp_for_split() to allocate a page table page for the
external page table in preparation for splitting the mirror page table.
When the mirror page table is split in tdp_mmu_split_huge_page(), the
corresponding external page table also needs to be split. Therefore,
allocate external_spt in tdp_mmu_alloc_sp_for_split() to prepare for the
splitting.
The external_spt will be gifted to the TDX module and mapped as a page
table page in the S-EPT during splitting. Since the TDX module will
initialize the page content in the DEMOTE SEAMCALL, there is no need to
zero external_spt.
Signed-off-by: Isaku Yamahata <isaku.yamahata@intel.com>
[sean: use __get_free_page(), let is_mirror_root be const and called once]
Signed-off-by: Sean Christopherson <seanjc@google.com>
Co-developed-by: Yan Zhao <yan.y.zhao@intel.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- "bool mirror" --> "bool is_mirror_sp". (Sean)
- Use __get_free_page() instead of get_zeroed_page(). (Sean)
- let is_mirror_root be const and called once. (Sean)
v3:
- Removed unnecessary declaration of tdp_mmu_alloc_sp_for_split(). (Kai)
- Fixed a typo in the patch log. (Kai)
RFC v2:
- NO change.
RFC v1:
- Rebased and simplified the code.
---
arch/x86/kvm/mmu/tdp_mmu.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
diff --git a/arch/x86/kvm/mmu/tdp_mmu.c b/arch/x86/kvm/mmu/tdp_mmu.c
index 001449142d32..f3311317a63a 100644
--- a/arch/x86/kvm/mmu/tdp_mmu.c
+++ b/arch/x86/kvm/mmu/tdp_mmu.c
@@ -1466,7 +1466,7 @@ bool kvm_tdp_mmu_wrprot_slot(struct kvm *kvm,
return spte_set;
}
-static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(void)
+static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(bool is_mirror_sp)
{
struct kvm_mmu_page *sp;
@@ -1480,6 +1480,15 @@ static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(void)
return NULL;
}
+ if (is_mirror_sp) {
+ sp->external_spt = (void *)__get_free_page(GFP_KERNEL_ACCOUNT);
+ if (!sp->external_spt) {
+ free_page((unsigned long)sp->spt);
+ kmem_cache_free(mmu_page_header_cache, sp);
+ return NULL;
+ }
+ }
+
return sp;
}
@@ -1527,6 +1536,7 @@ static int tdp_mmu_split_huge_pages_root(struct kvm *kvm,
gfn_t start, gfn_t end,
int target_level, bool shared)
{
+ const bool is_mirror_root = is_mirror_sp(root);
struct kvm_mmu_page *sp = NULL;
struct tdp_iter iter;
@@ -1559,7 +1569,7 @@ static int tdp_mmu_split_huge_pages_root(struct kvm *kvm,
else
write_unlock(&kvm->mmu_lock);
- sp = tdp_mmu_alloc_sp_for_split();
+ sp = tdp_mmu_alloc_sp_for_split(is_mirror_root);
if (shared)
read_lock(&kvm->mmu_lock);
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 06/17] KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (4 preceding siblings ...)
2026-09-28 9:09 ` [PATCH v4 05/17] KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table splitting Yan Zhao
@ 2026-09-28 9:10 ` Yan Zhao
2026-09-28 9:10 ` [PATCH v4 07/17] KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings to 4KB Yan Zhao
` (10 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:10 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
From: Sean Christopherson <seanjc@google.com>
Extend the TDP MMU to allocate Dynamic PAMT backing pages (DPAMT pages) for
vCPU-induced huge page splits in mirror roots when DPAMT is enabled.
Leverage the .topup_external_cache() interface to topup the DPAMT cache
when allocating a new child page table for splitting. The DPAMT cache is
currently a per-vCPU thread-local list. When a vCPU-induced page split
occurs, DPAMT pages can be drawn locklessly from the list.
Pass min_nr_spts as 1 to .topup_external_cache(), indicating there's one
new S-EPT page table page. So, tdx_topup_external_pamt_cache() will
allocate DPAMT page pairs for both the newly added S-EPT page table page
and the demoted guest private page.
tdp_mmu_alloc_sp_for_split() is currently not reachable from a non-vCPU
context for mirror roots, since dirty page tracking is not yet allowed on
mirror roots. So, simply add a WARN if tdx_topup_external_pamt_cache() is
invoked under a non-vCPU context.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4: new patch.
---
arch/x86/kvm/mmu/tdp_mmu.c | 24 +++++++++++++++---------
arch/x86/kvm/vmx/tdx.c | 3 +++
2 files changed, 18 insertions(+), 9 deletions(-)
diff --git a/arch/x86/kvm/mmu/tdp_mmu.c b/arch/x86/kvm/mmu/tdp_mmu.c
index f3311317a63a..472419963a19 100644
--- a/arch/x86/kvm/mmu/tdp_mmu.c
+++ b/arch/x86/kvm/mmu/tdp_mmu.c
@@ -1475,21 +1475,27 @@ static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(bool is_mirror_sp)
return NULL;
sp->spt = (void *)__get_free_page(GFP_KERNEL_ACCOUNT);
- if (!sp->spt) {
- kmem_cache_free(mmu_page_header_cache, sp);
- return NULL;
- }
+ if (!sp->spt)
+ goto err_spt;
if (is_mirror_sp) {
sp->external_spt = (void *)__get_free_page(GFP_KERNEL_ACCOUNT);
- if (!sp->external_spt) {
- free_page((unsigned long)sp->spt);
- kmem_cache_free(mmu_page_header_cache, sp);
- return NULL;
- }
+ if (!sp->external_spt)
+ goto err_external_spt;
+
+ if (kvm_x86_call(topup_external_cache)(kvm_get_running_vcpu(), 1))
+ goto err_external_split;
}
return sp;
+
+err_external_split:
+ free_page((unsigned long)sp->external_spt);
+err_external_spt:
+ free_page((unsigned long)sp->spt);
+err_spt:
+ kmem_cache_free(mmu_page_header_cache, sp);
+ return NULL;
}
/* Note, the caller is responsible for initializing @sp. */
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 11792a490330..3dcddf1b48c5 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -1630,6 +1630,9 @@ void tdx_load_mmu_pgd(struct kvm_vcpu *vcpu, hpa_t root_hpa, int pgd_level)
static int tdx_topup_external_pamt_cache(struct kvm_vcpu *vcpu, int min_nr_spts)
{
+ if (WARN_ON_ONCE(!vcpu))
+ return -EIO;
+
/*
* Minus one page to exclude the root SPT, but plus one page for a
* possible 4KB private mapping.
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 07/17] KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings to 4KB
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (5 preceding siblings ...)
2026-09-28 9:10 ` [PATCH v4 06/17] KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split Yan Zhao
@ 2026-09-28 9:10 ` Yan Zhao
2026-09-28 9:10 ` [PATCH v4 08/17] KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting S-EPT Yan Zhao
` (9 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:10 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
Add support for splitting, a.k.a. demoting, a 2MB S-EPT leaf mapping to 512
smaller 4KB leaf mappings. As per the TDX module rules, first invoke
MEM.RANGE.BLOCK to put the huge S-EPT leaf entry into a splittable state,
then do MEM.TRACK and kick all vCPUs outside of guest mode to flush TLBs,
and finally do MEM.PAGE.DEMOTE to demote/split the huge S-EPT leaf mapping.
Assert the mmu_lock is held for write, as the BLOCK => TRACK => DEMOTE
sequence needs to be "atomic" to guarantee success (and because mmu_lock
must be held for write to use tdh_do_no_vcpus()).
Note, even with kvm->mmu_lock held for write, tdh_mem_page_demote() may
contend with tdh_vp_enter() and potentially with the guest's S-EPT entry
operations. Therefore, wrap the call with tdh_do_no_vcpus() to kick other
vCPUs out of the guest and prevent tdh_vp_enter() to ensure success.
Invoke tdx_pamt_get() before invoking tdh_mem_page_demote() so that DPAMT
pages for the new S-EPT page table page are installed before the DEMOTE
SEAMCALL when DPAMT is enabled. DPAMT pages for the guest memory must be
installed inside the DEMOTE SEAMCALL, since it is impossible to do so
before a successful demotion.
Instead of allocating and freeing DPAMT pages for guest pages on the KVM
side, pass pamt_cache to tdh_mem_page_demote() and let it draw DPAMT pages
from pamt_cache before the DEMOTE SEAMCALL. This prevents KVM from having
to manage DPAMT pages directly via alloc_pamt_array() and
free_pamt_array(), or having knowledge of DPAMT-specific details such as
TDX_DPAMT_ENTRY_PAGE_CNT.
Signed-off-by: Xiaoyao Li <xiaoyao.li@intel.com>
Signed-off-by: Isaku Yamahata <isaku.yamahata@intel.com>
[sean: wire up via op set_external_spte(), merge in DPAMT-related code,
massage changelog]
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- Hooked tdx_sept_split_leaf_spte() in x86 op set_external_spte() instead
of in x86 op split_external_spte() which was no longer introduced in v4.
(Sean).
- Renamed tdx_sept_split_private_spte() --> tdx_sept_split_leaf_spte().
- Merged in DPAMT-related code (i.e., passing to_tdx(vcpu)->pamt_cache to
tdh_mem_page_demote(). (Sean).
- Assert new_spte is non-leaf. (Yan)
v3:
- Rebased on top of Sean's cleanup series.
- Call out UNBLOCK is not required after DEMOTE. (Kai)
- tdx_sept_split_private_spt() --> tdx_sept_split_private_spte().
RFC v2:
- Split out the code to handle the error TDX_INTERRUPTED_RESTARTABLE.
- Rebased to 6.16.0-rc6 (the way of defining TDX hook changes).
RFC v1:
- Split patch for exclusive mmu_lock only,
- Invoke tdx_sept_zap_private_spte() and tdx_track() for splitting.
- Handled busy error of tdh_mem_page_demote() by kicking off vCPUs.
---
arch/x86/kvm/vmx/tdx.c | 69 +++++++++++++++++++++++++++++++++++++++++-
1 file changed, 68 insertions(+), 1 deletion(-)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 3dcddf1b48c5..3186c4808cae 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -1873,11 +1873,74 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
return 0;
}
+/*
+ * Split a huge mapping into smaller mappings at a lower level. Currently only
+ * supports splitting 2MB mappings (KVM doesn't yet support 1GB mappings for TDX
+ * guests).
+ *
+ * Invoke "BLOCK + TRACK + kick off vCPUs (inside tdx_track())" since the TDX
+ * module does not yet support the NON-BLOCKING-RESIZE feature for DEMOTE.
+ *
+ * No UNBLOCK is needed after a successful DEMOTE.
+ *
+ * Under write mmu_lock, kick off all vCPUs and disallow vCPUs from entering to
+ * ensure DEMOTE will succeed on the second invocation if the first invocation
+ * returns BUSY.
+ */
+static int tdx_sept_split_leaf_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
+ u64 new_spte, enum pg_level level)
+{
+ struct kvm_vcpu *vcpu = kvm_get_running_vcpu();
+ struct kvm_tdx *kvm_tdx = to_kvm_tdx(kvm);
+ gpa_t gpa = gfn_to_gpa(gfn);
+ u64 err, entry, level_state;
+ struct page *sept_pt;
+ int r;
+
+ lockdep_assert_held_write(&kvm->mmu_lock);
+
+ if (KVM_BUG_ON(!is_last_spte(old_spte, level) || is_last_spte(new_spte, level), kvm))
+ return -EIO;
+
+ sept_pt = tdx_spte_to_sept_pt(kvm, gfn, new_spte, level);
+ if (!sept_pt)
+ return -EIO;
+
+ if (KVM_BUG_ON(!vcpu || vcpu->kvm != kvm, kvm))
+ return -EIO;
+
+ r = tdx_pamt_get(page_to_pfn(sept_pt), PG_LEVEL_4K, &to_tdx(vcpu)->pamt_cache);
+ if (KVM_BUG_ON(r, kvm))
+ return r;
+
+ err = tdh_do_no_vcpus(tdh_mem_range_block, kvm, &kvm_tdx->td, gpa,
+ level, &entry, &level_state);
+ if (TDX_BUG_ON_2(err, TDH_MEM_RANGE_BLOCK, entry, level_state, kvm)) {
+ r = -EIO;
+ goto err;
+ }
+
+ tdx_track(kvm);
+ err = tdh_do_no_vcpus(tdh_mem_page_demote, kvm, &kvm_tdx->td, gpa,
+ level, spte_to_pfn(old_spte), sept_pt,
+ &to_tdx(vcpu)->pamt_cache, &entry, &level_state);
+ if (TDX_BUG_ON_2(err, TDH_MEM_PAGE_DEMOTE, entry, level_state, kvm)) {
+ r = -EIO;
+ goto err;
+ }
+
+ return 0;
+err:
+ tdx_pamt_put(page_to_pfn(sept_pt), PG_LEVEL_4K);
+ return r;
+}
+
/*
* Handle changes for
* (1) leaf SPTEs from non-present to present
* (2) non-leaf SPTEs from non-present to present
* (3) leaf SPTEs from present to non-present
+ * (4) present leaf SPTEs to present non-leaf SPTEs (splitting)
*
* - (1) and (2) must be under shared mmu_lock. If (1) and (2) are under
* exclusive mmu_lock (currently impossible), contention errors may lead to
@@ -1888,13 +1951,17 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
* (currently impossible), warnings will be generated due to
* lockdep_assert_held_write() or TDX_BUG_ON() caused by concurrent BLOCK,
* TRACK, REMOVE.
- * - Promotion/demotion is not yet supported.
+ * - (4) must be under write mmu_lock currently.
+ * - Promotion is not yet supported.
*/
static int tdx_sept_set_private_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
u64 new_spte, enum pg_level level)
{
lockdep_assert_held(&kvm->mmu_lock);
+ if (is_shadow_present_pte(old_spte) && is_shadow_present_pte(new_spte))
+ return tdx_sept_split_leaf_spte(kvm, gfn, old_spte, new_spte, level);
+
if (is_shadow_present_pte(old_spte))
return tdx_sept_remove_leaf_spte(kvm, gfn, level, old_spte);
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 08/17] KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting S-EPT
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (6 preceding siblings ...)
2026-09-28 9:10 ` [PATCH v4 07/17] KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings to 4KB Yan Zhao
@ 2026-09-28 9:10 ` Yan Zhao
2026-09-28 9:10 ` [PATCH v4 09/17] KVM: x86/mmu: Introduce hugepage_set_guest_inhibit() Yan Zhao
` (8 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:10 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
KVM needs to allocate enough DPAMT page pairs in the pamt_cache for
consumption by both page table pages and guest pages. Since the DPAMT page
pair for the S-EPT root page is already allocated during TD initialization,
there is no need to allocate the DPAMT page pair for the S-EPT root page.
Therefore, previously the DPAMT page pairs required equals
"min_nr_spts - 1 + 1".
When splitting S-EPT, min_nr_spts does not include the root SPT. So, limit
the -1 calculation to when min_nr_spts equals root_level, though this will
cause one pair over-allocation in the normal page fault path because
PT64_ROOT_MAX_LEVEL is always passed even when launching a 4-level TD.
Additionally, KVM may need to retry tdh_mem_page_demote() a second time,
causing the DPAMT page pair for the guest private pages to be drawn from
the pamt_cache twice in the worst case. Therefore, add an extra +1 to cover
this worst-case scenario.
The slight over-allocation is acceptable since KVM already pre-allocates
more pages than needed (e.g., when mapping huge pages) in case of the
worst-case scenario.
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
arch/x86/kvm/vmx/tdx.c | 26 ++++++++++++++++++++++----
1 file changed, 22 insertions(+), 4 deletions(-)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 3186c4808cae..397a308b834c 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -1630,16 +1630,34 @@ void tdx_load_mmu_pgd(struct kvm_vcpu *vcpu, hpa_t root_hpa, int pgd_level)
static int tdx_topup_external_pamt_cache(struct kvm_vcpu *vcpu, int min_nr_spts)
{
+ int dpamt_pairs;
+
if (WARN_ON_ONCE(!vcpu))
return -EIO;
+ /* Exclude the root SPT, as its DPAMT page pair is already installed */
+ if (min_nr_spts == vcpu->kvm->arch.mirror_root_level)
+ min_nr_spts -= 1;
+
+ /*
+ * Each S-EPT page table page + 4KB guest private page needs a pair of
+ * DPAMT pages.
+ */
+ dpamt_pairs = min_nr_spts + 1;
+
/*
- * Minus one page to exclude the root SPT, but plus one page for a
- * possible 4KB private mapping.
+ * After each topup, KVM may invoke DEMOTE at most twice. The first
+ * DEMOTE invocation draws two pairs of pages from the cache: one for
+ * the S-EPT page table page and one for the guest private memory.
+ * Since these pages are not returned to the cache, the second DEMOTE
+ * invocation still needs to consume one additional pair for the guest
+ * private memory (the pair for the S-EPT page table page is reused for
+ * the 2nd invocation). Increase the topup count to account for this
+ * worst-case scenario.
*/
- min_nr_spts += -1 + 1;
+ dpamt_pairs += 1;
- return tdx_topup_pamt_cache(&to_tdx(vcpu)->pamt_cache, min_nr_spts);
+ return tdx_topup_pamt_cache(&to_tdx(vcpu)->pamt_cache, dpamt_pairs);
}
static int tdx_mem_page_add(struct kvm *kvm, gfn_t gfn, enum pg_level level,
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 09/17] KVM: x86/mmu: Introduce hugepage_set_guest_inhibit()
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (7 preceding siblings ...)
2026-09-28 9:10 ` [PATCH v4 08/17] KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting S-EPT Yan Zhao
@ 2026-09-28 9:10 ` Yan Zhao
2026-09-28 9:11 ` [PATCH v4 10/17] KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots Yan Zhao
` (7 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:10 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
TDX requires guests to accept S-EPT mappings created by the host KVM. Due
to the current implementation of the TDX module, if a guest accepts a GFN
at a lower level after KVM maps it at a higher level, the TDX module will
emulate an EPT violation VMExit to KVM instead of returning a size mismatch
error to the guest. If KVM fails to perform page splitting in the VMExit
handler, the guest's accept operation will be triggered again upon
re-entering the guest, causing a repeated EPT violation VMExit.
To facilitate passing the guest's accept level information to the KVM MMU
core and to prevent the repeated mapping of a GFN at different levels due
to different accept levels specified by different vCPUs, introduce the
interface hugepage_set_guest_inhibit(). This interface specifies across
vCPUs that mapping at a certain level is inhibited from the guest.
Intentionally don't provide an API to clear KVM_LPAGE_GUEST_INHIBIT_FLAG
for the time being, as detecting that it's ok to (re)install a huge page is
tricky (and costly if KVM wants to be 100% accurate), and KVM doesn't
currently support huge page promotion (only direct installation of
huge pages) for S-EPT.
As a result, the only scenario where clearing the flag would likely allow
KVM to install a huge page is when an entire 2MB / 1GB range is converted
to shared or private. But if the guest is accepting at 4KB granularity,
odds are good the guest is using the memory for something "special" and
will never convert the entire range to shared (and/or back to private).
Punt that optimization to the future, if it's ever needed.
Unlike hugepage_set/test_mixed(), which are placed under
CONFIG_KVM_VM_MEMORY_ATTRIBUTES since they are unneeded when per-VM memory
attributes is disabled, hugepage_set/test_guest_inhibit() are invoked when
per-gmem memory attributes is enabled, since TDX huge page support is
intentionally enabled only when per-gmem memory attributes is enabled.
Link: https://lore.kernel.org/all/a6ffe23fb97e64109f512fa43e9f6405236ed40a.camel@intel.com [1]
Suggested-by: Rick Edgecombe <rick.p.edgecombe@intel.com>
Suggested-by: Sean Christopherson <seanjc@google.com>
[sean: explain *why* the flag is never cleared]
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- Explain *why* the flag is never cleared. (Sean)
- Explain why hugepage_set/test_guest_inhibit() are not placed near to
hugepage_set/test_mixed().
v3:
- Use EXPORT_SYMBOL_FOR_KVM_INTERNAL().
RFC v2:
- new in RFC v2
---
arch/x86/kvm/mmu.h | 4 ++++
arch/x86/kvm/mmu/mmu.c | 23 +++++++++++++++++++----
2 files changed, 23 insertions(+), 4 deletions(-)
diff --git a/arch/x86/kvm/mmu.h b/arch/x86/kvm/mmu.h
index 2ae7f9ed4cf8..84a96ac94ada 100644
--- a/arch/x86/kvm/mmu.h
+++ b/arch/x86/kvm/mmu.h
@@ -410,4 +410,8 @@ static inline bool kvm_is_gfn_alias(struct kvm *kvm, gfn_t gfn)
{
return gfn & kvm_gfn_direct_bits(kvm);
}
+
+void hugepage_set_guest_inhibit(struct kvm_memory_slot *slot, gfn_t gfn, int level);
+bool hugepage_test_guest_inhibit(struct kvm_memory_slot *slot, gfn_t gfn, int level);
+
#endif
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index a9a7c015256b..0725b92b7008 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -750,12 +750,14 @@ static bool kvm_gfn_is_lpage_allowed(struct kvm *kvm,
}
/*
- * The most significant bit in disallow_lpage tracks whether or not memory
- * attributes are mixed, i.e. not identical for all gfns at the current level.
+ * The 2 most significant bits in disallow_lpage track whether or not memory
+ * attributes are mixed, i.e. not identical for all gfns at the current level,
+ * or whether or not guest inhibits the current level of hugepage at the gfn.
* The lower order bits are used to refcount other cases where a hugepage is
* disallowed, e.g. if KVM has shadow a page table at the gfn.
*/
-#define KVM_LPAGE_MIXED_FLAG BIT(31)
+#define KVM_LPAGE_MIXED_FLAG BIT(31)
+#define KVM_LPAGE_GUEST_INHIBIT_FLAG BIT(30)
static void update_gfn_disallow_lpage_count(const struct kvm_memory_slot *slot,
gfn_t gfn, int count)
@@ -768,7 +770,8 @@ static void update_gfn_disallow_lpage_count(const struct kvm_memory_slot *slot,
old = linfo->disallow_lpage;
linfo->disallow_lpage += count;
- WARN_ON_ONCE((old ^ linfo->disallow_lpage) & KVM_LPAGE_MIXED_FLAG);
+ WARN_ON_ONCE((old ^ linfo->disallow_lpage) &
+ (KVM_LPAGE_MIXED_FLAG | KVM_LPAGE_GUEST_INHIBIT_FLAG));
}
}
@@ -782,6 +785,18 @@ void kvm_mmu_gfn_allow_lpage(const struct kvm_memory_slot *slot, gfn_t gfn)
update_gfn_disallow_lpage_count(slot, gfn, -1);
}
+bool hugepage_test_guest_inhibit(struct kvm_memory_slot *slot, gfn_t gfn, int level)
+{
+ return lpage_info_slot(gfn, slot, level)->disallow_lpage & KVM_LPAGE_GUEST_INHIBIT_FLAG;
+}
+EXPORT_SYMBOL_FOR_KVM_INTERNAL(hugepage_test_guest_inhibit);
+
+void hugepage_set_guest_inhibit(struct kvm_memory_slot *slot, gfn_t gfn, int level)
+{
+ lpage_info_slot(gfn, slot, level)->disallow_lpage |= KVM_LPAGE_GUEST_INHIBIT_FLAG;
+}
+EXPORT_SYMBOL_FOR_KVM_INTERNAL(hugepage_set_guest_inhibit);
+
static void account_shadowed(struct kvm *kvm, struct kvm_mmu_page *sp)
{
struct kvm_memslots *slots;
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 10/17] KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (8 preceding siblings ...)
2026-09-28 9:10 ` [PATCH v4 09/17] KVM: x86/mmu: Introduce hugepage_set_guest_inhibit() Yan Zhao
@ 2026-09-28 9:11 ` Yan Zhao
2026-09-28 9:11 ` [PATCH v4 11/17] KVM: TDX: Honor the guest's accept level contained in an EPT violation Yan Zhao
` (6 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:11 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
From: Sean Christopherson <seanjc@google.com>
Add an exported API to split huge pages in mirror roots for a given gfn
range. TDX will use the API to split huge pages in preparation for
partially zapping a private huge page, e.g., for converting a huge page
from private to shared, or for splitting a huge page to match the guest's
ACCEPT level.
For all intents and purposes, no functional change intended.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- new patch.
---
arch/x86/kvm/mmu/tdp_mmu.c | 40 ++++++++++++++++++++++++++++----------
arch/x86/kvm/mmu/tdp_mmu.h | 2 ++
2 files changed, 32 insertions(+), 10 deletions(-)
diff --git a/arch/x86/kvm/mmu/tdp_mmu.c b/arch/x86/kvm/mmu/tdp_mmu.c
index 472419963a19..9c783accb9ed 100644
--- a/arch/x86/kvm/mmu/tdp_mmu.c
+++ b/arch/x86/kvm/mmu/tdp_mmu.c
@@ -1616,6 +1616,26 @@ static int tdp_mmu_split_huge_pages_root(struct kvm *kvm,
return 0;
}
+static int tdp_mmu_split_huge_pages(struct kvm *kvm, int as_id,
+ enum kvm_tdp_mmu_root_types type,
+ gfn_t start, gfn_t end,
+ int target_level, bool shared)
+{
+ struct kvm_mmu_page *root;
+ int r;
+
+ kvm_lockdep_assert_mmu_lock_held(kvm, shared);
+
+ __for_each_tdp_mmu_root_yield_safe(kvm, root, as_id, type) {
+ r = tdp_mmu_split_huge_pages_root(kvm, root, start, end,
+ target_level, shared);
+ if (r) {
+ kvm_tdp_mmu_put_root(kvm, root);
+ return r;
+ }
+ }
+ return 0;
+}
/*
* Try to split all huge pages mapped by the TDP MMU down to the target level.
@@ -1625,18 +1645,18 @@ void kvm_tdp_mmu_try_split_huge_pages(struct kvm *kvm,
gfn_t start, gfn_t end,
int target_level, bool shared)
{
- struct kvm_mmu_page *root;
- int r = 0;
+ tdp_mmu_split_huge_pages(kvm, slot->as_id, KVM_VALID_ROOTS, start, end,
+ target_level, shared);
+}
- kvm_lockdep_assert_mmu_lock_held(kvm, shared);
- for_each_valid_tdp_mmu_root_yield_safe(kvm, root, slot->as_id) {
- r = tdp_mmu_split_huge_pages_root(kvm, root, start, end, target_level, shared);
- if (r) {
- kvm_tdp_mmu_put_root(kvm, root);
- break;
- }
- }
+int kvm_tdp_mmu_mirrors_split_huge_pages(struct kvm *kvm, gfn_t start,
+ gfn_t end, int target_level)
+{
+ /* The as_id for mirror roots can only be 0 */
+ return tdp_mmu_split_huge_pages(kvm, 0, KVM_MIRROR_ROOTS, start, end,
+ target_level, false);
}
+EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_tdp_mmu_mirrors_split_huge_pages);
static bool tdp_mmu_need_write_protect(struct kvm *kvm, struct kvm_mmu_page *sp)
{
diff --git a/arch/x86/kvm/mmu/tdp_mmu.h b/arch/x86/kvm/mmu/tdp_mmu.h
index bd62977c9199..a6919de10ca2 100644
--- a/arch/x86/kvm/mmu/tdp_mmu.h
+++ b/arch/x86/kvm/mmu/tdp_mmu.h
@@ -97,6 +97,8 @@ void kvm_tdp_mmu_try_split_huge_pages(struct kvm *kvm,
const struct kvm_memory_slot *slot,
gfn_t start, gfn_t end,
int target_level, bool shared);
+int kvm_tdp_mmu_mirrors_split_huge_pages(struct kvm *kvm, gfn_t start,
+ gfn_t end, int target_level);
static inline void kvm_tdp_mmu_walk_lockless_begin(void)
{
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 11/17] KVM: TDX: Honor the guest's accept level contained in an EPT violation
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (9 preceding siblings ...)
2026-09-28 9:11 ` [PATCH v4 10/17] KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots Yan Zhao
@ 2026-09-28 9:11 ` Yan Zhao
2026-09-28 9:11 ` [PATCH v4 12/17] KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU context Yan Zhao
` (5 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:11 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
TDX requires guests to accept S-EPT mappings created by the host KVM. Due
to the current implementation of the TDX module, if a guest accepts a GFN
at a lower level after KVM maps it at a higher level, the TDX module will
synthesize an EPT Violation VM-Exit to KVM instead of returning a size
mismatch error to the guest. If KVM fails to perform page splitting in the
EPT Violation handler, the guest's ACCEPT operation will be triggered
again upon re-entering the guest, causing a repeated EPT Violation VM-Exit.
To ensure forward progress, honor the guest's accept level if an EPT
Violation VM-Exit contains the guest accept level (the TDX module provides
the level when synthesizing a VM-Exit in response to a failed guest ACCEPT,
e.g., due to accept level < host mapping level or due to intermediate
paging structure missing or inaccessible).
(1) Set the guest inhibit bit in the lpage info to prevent KVM's MMU
from mapping at a higher level than the guest's accept level.
(2) Split any existing mapping higher than the guest's accept level.
For now, take mmu_lock for write across the entire operation to keep things
simple. This can/will be revisited when the TDX module adds support for
NON-BLOCKING-RESIZE, at which point KVM can split the huge page without
needing to handle UNBLOCK failure if the DEMOTE fails.
To avoid unnecessarily contending mmu_lock, check if the inhibit flag is
already set before acquiring mmu_lock, e.g. so that vCPUs doing ACCEPT
on a region of memory aren't completely serialized. Note, this relies on
(a) setting the inhibit after performing the split, and (b) never clearing
the flag, e.g., to avoid false positives and potentially triggering the
zero-step mitigation.
Note: EPT Violation VM-Exits without the guest's accept level are *never*
caused by the guest's ACCEPT operation, but instead occur if the guest
accesses memory before said memory is accepted. Since KVM can't obtain
the guest accept level info from such EPT Violations (the ACCEPT operation
hasn't occurred yet), KVM may still map at a higher level than the guest's
later ACCEPT level.
So, the typical guest/KVM interaction flow is:
- If guest accesses private memory without first accepting it,
(like non-Linux guests):
1. Guest accesses a private memory.
2. KVM finds it can map the GFN at 2MB. So, AUG at 2MB.
3. Guest accepts the GFN at 4KB.
4. KVM receives an EPT violation with eeq_type of ACCEPT + 4KB level.
5. KVM splits the 2MB mapping.
6. Guest accepts successfully and accesses the page.
- If guest first accepts private memory before accessing it,
(like Linux guests):
1. Guest accepts a private memory at 4KB.
2. KVM receives an EPT violation with eeq_type of ACCEPT + 4KB level.
3. KVM AUG at 4KB.
4. Guest accepts successfully and accesses the page.
Link: https://lore.kernel.org/all/a6ffe23fb97e64109f512fa43e9f6405236ed40a.camel@intel.com
Suggested-by: Rick Edgecombe <rick.p.edgecombe@intel.com>
Suggested-by: Sean Christopherson <seanjc@google.com>
Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- In [1], Sean renamed tdx_honor_guest_accept_level() to
tdx_handle_mismatched_accept(). In v4, Yan further renamed it to
tdx_handle_guest_accept_ept_violation(), because EPT violations caused by
a guest's ACCEPT operation are not necessarily due to the guest's ACCEPT
level mismatching KVM's mapping level. In the case where a guest accepts
memory before ever accessing it, EPT violations caused by the guest's
ACCEPT operation occur when there are no present KVM mappings.
- Introduced helper tdx_is_mismatched_accepted() (Yan renamed it to
tdx_is_guest_accept_ept_violation()). (Sean)
- Introduced helper tdx_get_ept_violation_level(). (Sean)
- Invoked kvm_tdp_mmu_mirrors_split_huge_pages() instead of
kvm_split_cross_boundary_leafs(). (Sean)
[1] https://lore.kernel.org/all/20260129011517.3545883-42-seanjc@google.com
---
arch/x86/kvm/vmx/tdx.c | 75 +++++++++++++++++++++++++++++++++++++
arch/x86/kvm/vmx/tdx_arch.h | 3 ++
2 files changed, 78 insertions(+)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 397a308b834c..f97b76bd8fe9 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -14,6 +14,7 @@
#include "tdx.h"
#include "vmx.h"
#include "mmu/spte.h"
+#include "mmu/tdp_mmu.h"
#include "common.h"
#include "posted_intr.h"
#include "irq.h"
@@ -2049,6 +2050,76 @@ static inline bool tdx_is_sept_violation_unexpected_pending(struct kvm_vcpu *vcp
return !(eq & EPT_VIOLATION_PROT_MASK);
}
+static bool tdx_is_guest_accept_ept_violation(struct kvm_vcpu *vcpu)
+{
+ return (to_tdx(vcpu)->ext_exit_qualification & TDX_EXT_EXIT_QUAL_TYPE_MASK) ==
+ TDX_EXT_EXIT_QUAL_TYPE_ACCEPT;
+}
+
+static int tdx_get_ept_violation_level(struct kvm_vcpu *vcpu)
+{
+ u64 ext_exit_qual = to_tdx(vcpu)->ext_exit_qualification;
+
+ return (((ext_exit_qual & TDX_EXT_EXIT_QUAL_INFO_MASK) >>
+ TDX_EXT_EXIT_QUAL_INFO_SHIFT) & GENMASK(2, 0)) + 1;
+}
+
+/*
+ * An EPT violation can be either due to the guest's ACCEPT operation or
+ * due to the guest's access of memory before the guest accepts the
+ * memory.
+ *
+ * Type TDX_EXT_EXIT_QUAL_TYPE_ACCEPT in the extended exit qualification
+ * identifies EPT violations caused by the guest's ACCEPT operation, which must
+ * contain a valid guest accept level. For such EPT violations, honor guest's
+ * accept level by setting guest inhibit bit on levels above the guest accept
+ * level and split the existing mapping for the faulting GFN if it's with a
+ * higher level than the guest accept level.
+ *
+ * Do nothing if the EPT violation is not caused by the guest's ACCEPT
+ * operation. KVM will map the GFN without considering the guest's accept level
+ * (unless the guest inhibit bit is already set).
+ */
+static int tdx_handle_guest_accept_ept_violation(struct kvm_vcpu *vcpu, gfn_t gfn)
+{
+ struct kvm_memory_slot *slot = kvm_vcpu_gfn_to_memslot(vcpu, gfn);
+ struct kvm *kvm = vcpu->kvm;
+ gfn_t start, end;
+ int level, r;
+
+ if (!slot || !tdx_is_guest_accept_ept_violation(vcpu))
+ return 0;
+
+ if (WARN_ON_ONCE(!VALID_PAGE(vcpu->arch.mmu->mirror_root_hpa)))
+ return 0;
+
+ level = tdx_get_ept_violation_level(vcpu);
+ if (level > PG_LEVEL_2M)
+ return 0;
+
+ if (hugepage_test_guest_inhibit(slot, gfn, level + 1))
+ return 0;
+
+ guard(write_lock)(&kvm->mmu_lock);
+
+ start = gfn_round_for_level(gfn, level);
+ end = start + KVM_PAGES_PER_HPAGE(level);
+
+ r = kvm_tdp_mmu_mirrors_split_huge_pages(kvm, start, end, level);
+ if (r)
+ return r;
+
+ /*
+ * No TLB flush is required, as the "BLOCK + TRACK + kick off vCPUs"
+ * sequence required by the TDX-Module includes a TLB flush.
+ */
+ hugepage_set_guest_inhibit(slot, gfn, level + 1);
+ if (level == PG_LEVEL_4K)
+ hugepage_set_guest_inhibit(slot, gfn, level + 2);
+
+ return 0;
+}
+
static int tdx_handle_ept_violation(struct kvm_vcpu *vcpu)
{
unsigned long exit_qual;
@@ -2074,6 +2145,10 @@ static int tdx_handle_ept_violation(struct kvm_vcpu *vcpu)
*/
exit_qual = EPT_VIOLATION_ACC_WRITE;
+ ret = tdx_handle_guest_accept_ept_violation(vcpu, gpa_to_gfn(gpa));
+ if (ret)
+ return ret;
+
/* Only private GPA triggers zero-step mitigation */
local_retry = true;
} else {
diff --git a/arch/x86/kvm/vmx/tdx_arch.h b/arch/x86/kvm/vmx/tdx_arch.h
index 350143b9b145..6b8b18b0689a 100644
--- a/arch/x86/kvm/vmx/tdx_arch.h
+++ b/arch/x86/kvm/vmx/tdx_arch.h
@@ -76,7 +76,10 @@ struct tdx_cpuid_value {
} __packed;
#define TDX_EXT_EXIT_QUAL_TYPE_MASK GENMASK(3, 0)
+#define TDX_EXT_EXIT_QUAL_TYPE_ACCEPT 1
#define TDX_EXT_EXIT_QUAL_TYPE_PENDING_EPT_VIOLATION 6
+#define TDX_EXT_EXIT_QUAL_INFO_MASK GENMASK(63, 32)
+#define TDX_EXT_EXIT_QUAL_INFO_SHIFT 32
/*
* TD_PARAMS is provided as an input to TDH_MNG_INIT, the size of which is 1024B.
*/
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 12/17] KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU context
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (10 preceding siblings ...)
2026-09-28 9:11 ` [PATCH v4 11/17] KVM: TDX: Honor the guest's accept level contained in an EPT violation Yan Zhao
@ 2026-09-28 9:11 ` Yan Zhao
2026-09-28 9:11 ` [PATCH v4 13/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns give gmem+slot+pgoff Yan Zhao
` (4 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:11 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
Add support for splitting S-EPT entries under a non-vCPU context. This
prepares for zapping a subset of a huge mapping in S-EPT caused by
private-to-shared conversions or guest_memfd reclaiming of physical memory.
KVM must precisely zap/remove S-EPT entries to avoid clobbering guest
memory (the lifetime of guest private memory is tied to the S-EPT). So, KVM
needs to first split a huge mapping so that small mappings can be zapped
precisely.
Since there's no vCPU context, introduce a per-VM PAMT cache of
pre-allocated pages used to populate the Dynamic PAMT. Add a helper
tdx_get_pamt_cache() to select the per-VM PAMT cache when there's no vCPU
context. Add the "kvm" arg to .topup_external_cache() and its caller
tdp_mmu_alloc_sp_for_split() for the purpose of passing the "kvm" arg to
tdx_get_pamt_cache().
Use a mutex to guard the entire cycle from the per-VM PAMT cache topup to
drawing pages from the cache. Using a mutex (e.g., versus a spinlock) is
important as it allows KVM to only drop and re-aquire the mmu_lock (a
spinlock) while continuing holding the mutex for memory allocation.
Introduce a local static function tdx_sept_split_huge_pages(), which
acquires the mutex before triggering the S-EPT entries splitting under a
non-vCPU context. This function is intended to be invoked by guest_memfd
via an arch hook in a later patch. Though functions related to dirty page
tracking can also trigger splitting under a non-vCPU context, they do not
yet involve mirror roots. So, how those functions should acquire the mutex
is deferred to a later consideration.
tdx_sept_split_huge_pages() internally invokes API
kvm_tdp_mmu_mirrors_split_huge_pages() to split mirror roots. To avoid
unnecessary work, explicitly detect unaligned head and tail pages relative
to the max page size supported by KVM (currently 2MB for private memory),
and split only those pages, as only unaligned head/tail pages will undergo
partial zapping.
Signed-off-by: Sean Christopherson <seanjc@google.com>
[Yan: Tweak patch log/function names, split out .gmem_prezap() hook]
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- New patch.
- Split out the registration of the .gmem_prezap hook into a later patch to
isolate gmem-related changes. (Yan)
---
arch/x86/include/asm/kvm_host.h | 2 +-
arch/x86/kvm/mmu/mmu.c | 2 +-
arch/x86/kvm/mmu/tdp_mmu.c | 7 +--
arch/x86/kvm/vmx/tdx.c | 94 +++++++++++++++++++++++++++++----
arch/x86/kvm/vmx/tdx.h | 5 ++
5 files changed, 96 insertions(+), 14 deletions(-)
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index 373559a7cca5..5dd1db64562f 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -1650,7 +1650,7 @@ struct kvm_x86_ops {
/* Update external page tables for page table about to be freed. */
void (*free_external_spt)(struct kvm *kvm, struct kvm_mmu_page *sp);
- int (*topup_external_cache)(struct kvm_vcpu *vcpu, int min_nr_spts);
+ int (*topup_external_cache)(struct kvm *kvm, struct kvm_vcpu *vcpu, int min_nr_spts);
bool (*has_wbinvd_exit)(void);
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 0725b92b7008..d92170a34728 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -618,7 +618,7 @@ static int mmu_topup_memory_caches(struct kvm_vcpu *vcpu, bool maybe_indirect)
if (r)
return r;
- r = kvm_x86_call(topup_external_cache)(vcpu, PT64_ROOT_MAX_LEVEL);
+ r = kvm_x86_call(topup_external_cache)(vcpu->kvm, vcpu, PT64_ROOT_MAX_LEVEL);
if (r)
return r;
}
diff --git a/arch/x86/kvm/mmu/tdp_mmu.c b/arch/x86/kvm/mmu/tdp_mmu.c
index 9c783accb9ed..4829ddcd3b55 100644
--- a/arch/x86/kvm/mmu/tdp_mmu.c
+++ b/arch/x86/kvm/mmu/tdp_mmu.c
@@ -1466,7 +1466,8 @@ bool kvm_tdp_mmu_wrprot_slot(struct kvm *kvm,
return spte_set;
}
-static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(bool is_mirror_sp)
+static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(struct kvm *kvm,
+ bool is_mirror_sp)
{
struct kvm_mmu_page *sp;
@@ -1483,7 +1484,7 @@ static struct kvm_mmu_page *tdp_mmu_alloc_sp_for_split(bool is_mirror_sp)
if (!sp->external_spt)
goto err_external_spt;
- if (kvm_x86_call(topup_external_cache)(kvm_get_running_vcpu(), 1))
+ if (kvm_x86_call(topup_external_cache)(kvm, kvm_get_running_vcpu(), 1))
goto err_external_split;
}
@@ -1575,7 +1576,7 @@ static int tdp_mmu_split_huge_pages_root(struct kvm *kvm,
else
write_unlock(&kvm->mmu_lock);
- sp = tdp_mmu_alloc_sp_for_split(is_mirror_root);
+ sp = tdp_mmu_alloc_sp_for_split(kvm, is_mirror_root);
if (shared)
read_lock(&kvm->mmu_lock);
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index f97b76bd8fe9..5c5919b76bdd 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -585,6 +585,8 @@ void tdx_vm_destroy(struct kvm *kvm)
{
struct kvm_tdx *kvm_tdx = to_kvm_tdx(kvm);
+ tdx_free_pamt_cache(&kvm_tdx->pamt_cache);
+
tdx_reclaim_td_control_pages(kvm);
kvm_tdx->state = TD_STATE_UNINITIALIZED;
@@ -650,6 +652,9 @@ int tdx_vm_init(struct kvm *kvm)
kvm_tdx->state = TD_STATE_UNINITIALIZED;
+ tdx_init_pamt_cache(&kvm_tdx->pamt_cache);
+ mutex_init(&kvm_tdx->pamt_cache_lock);
+
return 0;
}
@@ -1629,15 +1634,31 @@ void tdx_load_mmu_pgd(struct kvm_vcpu *vcpu, hpa_t root_hpa, int pgd_level)
td_vmcs_write64(to_tdx(vcpu), SHARED_EPT_POINTER, root_hpa);
}
-static int tdx_topup_external_pamt_cache(struct kvm_vcpu *vcpu, int min_nr_spts)
+static struct tdx_pamt_cache *tdx_get_pamt_cache(struct kvm *kvm,
+ struct kvm_vcpu *vcpu)
{
+ if (KVM_BUG_ON(vcpu && vcpu->kvm != kvm, kvm))
+ return NULL;
+
+ if (vcpu)
+ return &to_tdx(vcpu)->pamt_cache;
+
+ lockdep_assert_held(&to_kvm_tdx(kvm)->pamt_cache_lock);
+ return &to_kvm_tdx(kvm)->pamt_cache;
+}
+
+static int tdx_topup_external_pamt_cache(struct kvm *kvm, struct kvm_vcpu *vcpu,
+ int min_nr_spts)
+{
+ struct tdx_pamt_cache *pamt_cache;
int dpamt_pairs;
- if (WARN_ON_ONCE(!vcpu))
+ pamt_cache = tdx_get_pamt_cache(kvm, vcpu);
+ if (!pamt_cache)
return -EIO;
/* Exclude the root SPT, as its DPAMT page pair is already installed */
- if (min_nr_spts == vcpu->kvm->arch.mirror_root_level)
+ if (min_nr_spts == kvm->arch.mirror_root_level)
min_nr_spts -= 1;
/*
@@ -1658,7 +1679,7 @@ static int tdx_topup_external_pamt_cache(struct kvm_vcpu *vcpu, int min_nr_spts)
*/
dpamt_pairs += 1;
- return tdx_topup_pamt_cache(&to_tdx(vcpu)->pamt_cache, dpamt_pairs);
+ return tdx_topup_pamt_cache(pamt_cache, dpamt_pairs);
}
static int tdx_mem_page_add(struct kvm *kvm, gfn_t gfn, enum pg_level level,
@@ -1909,8 +1930,8 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
static int tdx_sept_split_leaf_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
u64 new_spte, enum pg_level level)
{
- struct kvm_vcpu *vcpu = kvm_get_running_vcpu();
struct kvm_tdx *kvm_tdx = to_kvm_tdx(kvm);
+ struct tdx_pamt_cache *pamt_cache;
gpa_t gpa = gfn_to_gpa(gfn);
u64 err, entry, level_state;
struct page *sept_pt;
@@ -1925,10 +1946,11 @@ static int tdx_sept_split_leaf_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
if (!sept_pt)
return -EIO;
- if (KVM_BUG_ON(!vcpu || vcpu->kvm != kvm, kvm))
+ pamt_cache = tdx_get_pamt_cache(kvm, kvm_get_running_vcpu());
+ if (!pamt_cache)
return -EIO;
- r = tdx_pamt_get(page_to_pfn(sept_pt), PG_LEVEL_4K, &to_tdx(vcpu)->pamt_cache);
+ r = tdx_pamt_get(page_to_pfn(sept_pt), PG_LEVEL_4K, pamt_cache);
if (KVM_BUG_ON(r, kvm))
return r;
@@ -1941,8 +1963,8 @@ static int tdx_sept_split_leaf_spte(struct kvm *kvm, gfn_t gfn, u64 old_spte,
tdx_track(kvm);
err = tdh_do_no_vcpus(tdh_mem_page_demote, kvm, &kvm_tdx->td, gpa,
- level, spte_to_pfn(old_spte), sept_pt,
- &to_tdx(vcpu)->pamt_cache, &entry, &level_state);
+ level, spte_to_pfn(old_spte), sept_pt, pamt_cache,
+ &entry, &level_state);
if (TDX_BUG_ON_2(err, TDH_MEM_PAGE_DEMOTE, entry, level_state, kvm)) {
r = -EIO;
goto err;
@@ -2027,6 +2049,60 @@ static void tdx_sept_free_private_spt(struct kvm *kvm, struct kvm_mmu_page *sp)
sp->external_spt = NULL;
}
+static int tdx_sept_split_huge_page_at(struct kvm *kvm, gfn_t gfn, int target_level)
+{
+ gfn_t end = gfn + KVM_PAGES_PER_HPAGE(target_level + 1);
+
+ return kvm_tdp_mmu_mirrors_split_huge_pages(kvm, gfn, end, target_level);
+}
+
+static int tdx_sept_split_straddling_huge_pages_to_level(struct kvm *kvm, gfn_t start,
+ gfn_t end, int target_level)
+{
+ gfn_t head = gfn_round_for_level(start, target_level + 1);
+ gfn_t tail = gfn_round_for_level(end, target_level + 1);
+ int r;
+
+ if (head != start) {
+ r = tdx_sept_split_huge_page_at(kvm, head, target_level);
+ if (r)
+ return r;
+ }
+
+ if (tail != end && (head != tail || head == start)) {
+ r = tdx_sept_split_huge_page_at(kvm, tail, target_level);
+ if (r)
+ return r;
+ }
+
+ return 0;
+}
+
+/*
+ * Split S-EPT huge mappings that straddle [start, end) under non-vCPU context.
+ *
+ * Split potential huge mappings at the head and tail of the to-be-zapped range
+ * so that KVM doesn't overzap due to dropping a hugepage that doesn't fall
+ * wholly inside the range.
+ *
+ * Acquire the external cache lock, a.k.a. the Dynamic PAMT lock, to protect the
+ * per-VM cache of pre-allocated pages used to populate the Dynamic PAMT when
+ * splitting S-EPT huge pages.
+ */
+static int __maybe_unused tdx_sept_split_huge_pages(struct kvm *kvm, gfn_t start,
+ gfn_t end)
+{
+ guard(mutex)(&to_kvm_tdx(kvm)->pamt_cache_lock);
+
+ guard(write_lock)(&kvm->mmu_lock);
+
+ /*
+ * TODO: Also split from PG_LEVEL_1G => PG_LEVEL_2M when KVM supports
+ * 1GB S-EPT pages.
+ */
+ return tdx_sept_split_straddling_huge_pages_to_level(kvm, start, end, PG_LEVEL_4K);
+}
+
void tdx_deliver_interrupt(struct kvm_lapic *apic, int delivery_mode,
int trig_mode, int vector)
{
diff --git a/arch/x86/kvm/vmx/tdx.h b/arch/x86/kvm/vmx/tdx.h
index fd368e3ee060..602d1c289c2d 100644
--- a/arch/x86/kvm/vmx/tdx.h
+++ b/arch/x86/kvm/vmx/tdx.h
@@ -47,6 +47,11 @@ struct kvm_tdx {
* Set/unset is protected with kvm->mmu_lock.
*/
bool wait_for_sept_zap;
+
+ /* The per-VM cache for DPAMT pages for S-EPT pages and guest pages */
+ struct tdx_pamt_cache pamt_cache;
+ /* Protect the per-VM cache for DPAMT pages */
+ struct mutex pamt_cache_lock;
};
/* TDX module vCPU states */
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 13/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns give gmem+slot+pgoff
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (11 preceding siblings ...)
2026-09-28 9:11 ` [PATCH v4 12/17] KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU context Yan Zhao
@ 2026-09-28 9:11 ` Yan Zhao
2026-09-28 9:11 ` [PATCH v4 14/17] [GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start() to start() and zap() Yan Zhao
` (3 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:11 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
From: Sean Christopherson <seanjc@google.com>
Add helpers for getting a gfn given a gmem slot+pgoff, and for getting a
gfn given a starting or ending pgoff, i.e. an offset that may be beyond
the range of the memslot binding. Providing helpers will avoid duplicate
boilerplate code "if" future code also needs to iterate over gfn ranges.
No functional change intended.
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
virt/kvm/guest_memfd.c | 21 +++++++++++++++++----
1 file changed, 17 insertions(+), 4 deletions(-)
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 5bf7c8b4509f..fea460e5e503 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -80,6 +80,21 @@ static pgoff_t kvm_gmem_get_index(struct kvm_memory_slot *slot, gfn_t gfn)
return gfn - slot->base_gfn + slot->gmem.pgoff;
}
+static gfn_t kvm_gmem_get_gfn(struct kvm_memory_slot *slot, pgoff_t pgoff)
+{
+ return slot->base_gfn + pgoff - slot->gmem.pgoff;
+}
+
+static gfn_t kvm_gmem_get_start_gfn(struct kvm_memory_slot *slot, pgoff_t start)
+{
+ return kvm_gmem_get_gfn(slot, max(slot->gmem.pgoff, start));
+}
+
+static gfn_t kvm_gmem_get_end_gfn(struct kvm_memory_slot *slot, pgoff_t end)
+{
+ return kvm_gmem_get_gfn(slot, min(slot->gmem.pgoff + slot->npages, end));
+}
+
static u64 kvm_gmem_get_default_attributes(struct inode *inode)
{
bool init_shared = GMEM_I(inode)->flags & GUEST_MEMFD_FLAG_INIT_SHARED;
@@ -212,11 +227,9 @@ static void __kvm_gmem_invalidate_start(struct gmem_file *f, pgoff_t start,
unsigned long index;
xa_for_each_range(&f->bindings, index, slot, start, end - 1) {
- pgoff_t pgoff = slot->gmem.pgoff;
-
struct kvm_gfn_range gfn_range = {
- .start = slot->base_gfn + max(pgoff, start) - pgoff,
- .end = slot->base_gfn + min(pgoff + slot->npages, end) - pgoff,
+ .start = kvm_gmem_get_start_gfn(slot, start),
+ .end = kvm_gmem_get_end_gfn(slot, end),
.slot = slot,
.may_block = true,
.attr_filter = attr_filter,
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 14/17] [GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start() to start() and zap()
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (12 preceding siblings ...)
2026-09-28 9:11 ` [PATCH v4 13/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns give gmem+slot+pgoff Yan Zhao
@ 2026-09-28 9:11 ` Yan Zhao
2026-09-28 9:12 ` [PATCH v4 15/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap() Yan Zhao
` (2 subsequent siblings)
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:11 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
Split the kvm_gmem_invalidate_start() to two parts:
1) kvm_gmem_invalidate_start(): notifies KVM that MMU invalidation starts,
and updates the invalidate range.
2) kvm_gmem_zap(): triggers the actual zapping of KVM secondary MMU
mappings.
This prepares adding and invoking a .prezap() hook to guest_memfd.
No functional changes expected.
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
virt/kvm/guest_memfd.c | 49 ++++++++++++++++++++++++++++++++++++------
1 file changed, 43 insertions(+), 6 deletions(-)
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index fea460e5e503..60417008b21e 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -217,9 +217,8 @@ static enum kvm_gfn_range_filter kvm_gmem_get_all_gfns_filter(struct inode *inod
return KVM_FILTER_PRIVATE;
}
-static void __kvm_gmem_invalidate_start(struct gmem_file *f, pgoff_t start,
- pgoff_t end,
- enum kvm_gfn_range_filter attr_filter)
+static void __kvm_gmem_zap(struct gmem_file *f, pgoff_t start, pgoff_t end,
+ enum kvm_gfn_range_filter attr_filter)
{
bool flush = false, found_memslot = false;
struct kvm_memory_slot *slot;
@@ -239,7 +238,6 @@ static void __kvm_gmem_invalidate_start(struct gmem_file *f, pgoff_t start,
found_memslot = true;
KVM_MMU_LOCK(kvm);
- kvm_mmu_invalidate_start(kvm);
}
flush |= kvm_mmu_unmap_gfn_range(kvm, &gfn_range);
@@ -256,6 +254,41 @@ static void __kvm_gmem_invalidate_start(struct gmem_file *f, pgoff_t start,
KVM_MMU_UNLOCK(kvm);
}
+static void kvm_gmem_zap(struct inode *inode, pgoff_t start, pgoff_t end,
+ enum kvm_gfn_range_filter filter)
+{
+ struct gmem_file *f;
+
+ kvm_gmem_for_each_file(f, inode)
+ __kvm_gmem_zap(f, start, end, filter);
+}
+
+static void __kvm_gmem_invalidate_start(struct gmem_file *f, pgoff_t start,
+ pgoff_t end,
+ enum kvm_gfn_range_filter attr_filter)
+{
+ bool found_memslot = false;
+ struct kvm_memory_slot *slot;
+ struct kvm *kvm = f->kvm;
+ unsigned long index;
+
+ xa_for_each_range(&f->bindings, index, slot, start, end - 1) {
+ gfn_t invalidate_start = kvm_gmem_get_start_gfn(slot, start);
+ gfn_t invalidate_end = kvm_gmem_get_end_gfn(slot, end);
+
+ if (!found_memslot) {
+ found_memslot = true;
+
+ KVM_MMU_LOCK(kvm);
+ kvm_mmu_invalidate_start(kvm);
+ }
+ kvm_mmu_invalidate_range_add(kvm, invalidate_start, invalidate_end);
+ }
+
+ if (found_memslot)
+ KVM_MMU_UNLOCK(kvm);
+}
+
static void kvm_gmem_invalidate_start(struct inode *inode, pgoff_t start,
pgoff_t end,
enum kvm_gfn_range_filter filter)
@@ -309,6 +342,7 @@ static long kvm_gmem_punch_hole(struct inode *inode, loff_t offset, loff_t len)
filemap_invalidate_lock(inode->i_mapping);
kvm_gmem_invalidate_start(inode, start, end, filter);
+ kvm_gmem_zap(inode, start, end, filter);
truncate_inode_pages_range(inode->i_mapping, offset, offset + len - 1);
@@ -392,6 +426,7 @@ static long kvm_gmem_fallocate(struct file *file, int mode, loff_t offset,
static int kvm_gmem_release(struct inode *inode, struct file *file)
{
+ enum kvm_gfn_range_filter filter = kvm_gmem_get_all_gfns_filter(inode);
struct gmem_file *f = file->private_data;
struct kvm_memory_slot *slot;
struct kvm *kvm = f->kvm;
@@ -461,8 +496,8 @@ static int kvm_gmem_release(struct inode *inode, struct file *file)
* Zap all SPTEs pointed at by this file. Do not free the backing
* memory, as its lifetime is associated with the inode, not the file.
*/
- __kvm_gmem_invalidate_start(f, 0, -1ul,
- kvm_gmem_get_all_gfns_filter(inode));
+ __kvm_gmem_invalidate_start(f, 0, -1ul, filter);
+ __kvm_gmem_zap(f, 0, -1ul, filter);
__kvm_gmem_invalidate_end(f, 0, -1ul);
list_del(&f->entry);
@@ -789,6 +824,7 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE;
kvm_gmem_invalidate_start(inode, start, end, filter);
+ kvm_gmem_zap(inode, start, end, filter);
if (!to_private && kvm_arch_has_gmem_convert())
kvm_gmem_make_shared(inode, start, end);
@@ -890,6 +926,7 @@ static int kvm_gmem_error_folio(struct address_space *mapping, struct folio *fol
filter = kvm_gmem_get_all_gfns_filter(inode);
kvm_gmem_invalidate_start(inode, start, end, filter);
+ kvm_gmem_zap(inode, start, end, filter);
/*
* Do not truncate the range, what action is taken in response to the
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 15/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap()
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (13 preceding siblings ...)
2026-09-28 9:11 ` [PATCH v4 14/17] [GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start() to start() and zap() Yan Zhao
@ 2026-09-28 9:12 ` Yan Zhao
2026-09-28 9:12 ` [PATCH v4 16/17] [GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split S-EPT Yan Zhao
2026-09-28 9:12 ` [PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M Yan Zhao
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:12 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
From: Sean Christopherson <seanjc@google.com>
Add a gmem 'pre-zap' hook to allow arch code to take action before a zap,
e.g., for shared<=>private conversion, and just as importantly, to let arch
code reject performing the actual zap, e.g., if the conversion requires new
page tables and KVM hits an OOM situation.
The arch code and hook will be used by TDX to split huge mappings as
necessary to avoid over-zapping PTEs, which for all intents and purposes
corrupts guest data for TDX VMs (memory is wiped when private PTEs are
removed).
The hook is allowed to fail, however, there is no rollback when an error
occurs. Therefore, the hook implementation is expected to be safe without
any rollback on error. For example, in TDX, the hook splits huge mappings
as necessary to avoid over-zapping PTEs. It is safe to leave the preceding
successfully split mappings as-is rather than merging them back.
Currently, the pre-zap hook is invoked before zaps for memory attribute
conversions and punch hole operations. The invocation in punch hole should
be a no-op if the punch hole range is aligned to the huge page size. There
is no need to trigger the pre-zap hook before releasing gmem, as all
mappings will be gone anyway. The pre-zap hook is not invoked in
kvm_gmem_error_folio() to avoid introducing additional failure points, and
since when kvm_gmem_error_folio() is invoked, the VM is about to be killed,
over-zapping is not a concern in that case.
Not-Yet-Signed-off-by: Sean Christopherson <seanjc@google.com>
[Yan: Renamed to .gmem_prezap(), used attr_filter for private/shared info]
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
- Rebased to gmem in-place conversion v13.
- Renamed .gmem_convert() in [1] to .gmem_prezap() as previously noted in
[2].
- Added a comment for .gmem_prezap() noting that the hook is allowed to
fail and must be safe without any rollback in the event of an error.(Yan)
[1] https://lore.kernel.org/all/20260129011517.3545883-44-seanjc@google.com
[2] https://lore.kernel.org/all/anLrGmbZwgWnNUkp@yzhao56-desk.sh.intel.com
---
arch/x86/include/asm/kvm-x86-ops.h | 3 ++
arch/x86/include/asm/kvm_host.h | 16 ++++++++
arch/x86/kvm/x86.c | 8 ++++
include/linux/kvm_host.h | 5 +++
include/linux/kvm_types.h | 1 +
virt/kvm/Kconfig | 4 ++
virt/kvm/guest_memfd.c | 65 +++++++++++++++++++++++++++++-
7 files changed, 101 insertions(+), 1 deletion(-)
diff --git a/arch/x86/include/asm/kvm-x86-ops.h b/arch/x86/include/asm/kvm-x86-ops.h
index a2eec24fb326..ada0590b4a1b 100644
--- a/arch/x86/include/asm/kvm-x86-ops.h
+++ b/arch/x86/include/asm/kvm-x86-ops.h
@@ -157,6 +157,9 @@ KVM_X86_OP_OPTIONAL(gmem_make_shared)
#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_INVALIDATE
KVM_X86_OP_OPTIONAL(gmem_invalidate_range)
#endif
+#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREZAP
+KVM_X86_OP_OPTIONAL_RET0(gmem_prezap)
+#endif
KVM_X86_OP_OPTIONAL_RET0(gmem_max_mapping_level)
#endif
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index 5dd1db64562f..ac1af1a083a6 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -1744,6 +1744,19 @@ struct kvm_x86_ops {
#endif
#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_INVALIDATE
void (*gmem_invalidate_range)(struct kvm *kvm, struct kvm_gfn_range *range);
+#endif
+#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREZAP
+ /*
+ * Preparation before gmem triggering MMU zap, e.g., splitting huge
+ * mappings in S-EPT to prevent over-zapping in TDX.
+ * Note: Though the preparation is allowed to fail, it must be safe to
+ * proceed without any rollback when an error occurs. For example, if an
+ * error occurs while splitting a huge mapping, it is safe to leave the
+ * preceding successfully split mappings as-is rather than merging them
+ * back.
+ */
+ int (*gmem_prezap)(struct kvm *kvm, gfn_t start, gfn_t end,
+ enum kvm_gfn_range_filter attr_filter);
#endif
int (*gmem_max_mapping_level)(struct kvm *kvm, kvm_pfn_t pfn, bool is_private);
};
@@ -1866,6 +1879,9 @@ enum kvm_intr_type {
#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_CONVERT
#define kvm_arch_has_gmem_convert() (!!kvm_x86_ops.gmem_make_private)
#endif
+#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREZAP
+#define kvm_arch_has_gmem_prezap() (!!kvm_x86_ops.gmem_prezap)
+#endif
#define kvm_arch_has_readonly_mem(kvm) (!(kvm)->arch.has_protected_state)
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index 578d624aea28..f28549b3ae73 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -10667,6 +10667,14 @@ void kvm_arch_gmem_invalidate_range(struct kvm *kvm, struct kvm_gfn_range *range
kvm_x86_call(gmem_invalidate_range)(kvm, range);
}
#endif
+
+#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREZAP
+int kvm_arch_gmem_prezap(struct kvm *kvm, gfn_t start, gfn_t end,
+ enum kvm_gfn_range_filter attr_filter)
+{
+ return kvm_x86_call(gmem_prezap)(kvm, start, end, attr_filter);
+}
+#endif
#endif
void kvm_fixup_and_inject_pf_error(struct kvm_vcpu *vcpu, gva_t gva, u16 error_code)
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 284fc7d68c60..1debafa18d77 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -2622,6 +2622,11 @@ void kvm_arch_gmem_make_shared(kvm_pfn_t pfn, kvm_pfn_t nr_pages);
#define kvm_arch_has_gmem_convert() false
#endif
+#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREZAP
+int kvm_arch_gmem_prezap(struct kvm *kvm, gfn_t start, gfn_t end,
+ enum kvm_gfn_range_filter attr_filter);
+#endif
+
#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_POPULATE
/**
* kvm_gmem_populate() - Populate/prepare a GPA range with guest data
diff --git a/include/linux/kvm_types.h b/include/linux/kvm_types.h
index a568d8e6f4e8..edd19584a3a1 100644
--- a/include/linux/kvm_types.h
+++ b/include/linux/kvm_types.h
@@ -49,6 +49,7 @@ struct kvm_vcpu_init;
struct kvm_memslots;
enum kvm_mr_change;
+enum kvm_gfn_range_filter;
/*
* Address types:
diff --git a/virt/kvm/Kconfig b/virt/kvm/Kconfig
index a0678ef8ee3f..564ea066ed1f 100644
--- a/virt/kvm/Kconfig
+++ b/virt/kvm/Kconfig
@@ -116,6 +116,10 @@ config HAVE_KVM_ARCH_GMEM_INVALIDATE
bool
depends on KVM_GUEST_MEMFD
+config HAVE_KVM_ARCH_GMEM_PREZAP
+ bool
+ depends on KVM_GUEST_MEMFD
+
config HAVE_KVM_ARCH_GMEM_POPULATE
bool
depends on KVM_GUEST_MEMFD
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 60417008b21e..0b49c2215183 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -217,6 +217,53 @@ static enum kvm_gfn_range_filter kvm_gmem_get_all_gfns_filter(struct inode *inod
return KVM_FILTER_PRIVATE;
}
+#ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREZAP
+static int __kvm_gmem_prezap(struct gmem_file *f, pgoff_t start, pgoff_t end,
+ enum kvm_gfn_range_filter filter)
+{
+ struct kvm_memory_slot *slot;
+ unsigned long index;
+ int r;
+
+ /*
+ * Since kvm_arch_gmem_prezap() internally holds mutex, no need to hold
+ * mmu_lock here. Let kvm_arch_gmem_prezap() acquire mmu_lock by itself.
+ */
+ xa_for_each_range(&f->bindings, index, slot, start, end - 1) {
+ r = kvm_arch_gmem_prezap(f->kvm,
+ kvm_gmem_get_start_gfn(slot, start),
+ kvm_gmem_get_end_gfn(slot, end),
+ filter);
+ if (r)
+ return r;
+ }
+ return 0;
+}
+
+static int kvm_gmem_prezap(struct inode *inode, pgoff_t start, pgoff_t end,
+ enum kvm_gfn_range_filter filter)
+{
+ struct gmem_file *f;
+ int r;
+
+ if (!kvm_arch_has_gmem_prezap())
+ return 0;
+
+ kvm_gmem_for_each_file(f, inode) {
+ r = __kvm_gmem_prezap(f, start, end, filter);
+ if (r)
+ return r;
+ }
+ return 0;
+}
+#else
+static int kvm_gmem_prezap(struct inode *inode, pgoff_t start, pgoff_t end,
+ enum kvm_gfn_range_filter filter)
+{
+ return 0;
+}
+#endif
+
static void __kvm_gmem_zap(struct gmem_file *f, pgoff_t start, pgoff_t end,
enum kvm_gfn_range_filter attr_filter)
{
@@ -326,6 +373,7 @@ static long kvm_gmem_punch_hole(struct inode *inode, loff_t offset, loff_t len)
pgoff_t start = offset >> PAGE_SHIFT;
pgoff_t end = (offset + len) >> PAGE_SHIFT;
struct gmem_inode *gi = GMEM_I(inode);
+ int r = 0;
/*
* gi->page_order is 0 by default and is set to PMD order only when the
@@ -342,15 +390,20 @@ static long kvm_gmem_punch_hole(struct inode *inode, loff_t offset, loff_t len)
filemap_invalidate_lock(inode->i_mapping);
kvm_gmem_invalidate_start(inode, start, end, filter);
+ r = kvm_gmem_prezap(inode, start, end, filter);
+ if (r)
+ goto out;
+
kvm_gmem_zap(inode, start, end, filter);
truncate_inode_pages_range(inode->i_mapping, offset, offset + len - 1);
+out:
kvm_gmem_invalidate_end(inode, start, end);
filemap_invalidate_unlock(inode->i_mapping);
- return 0;
+ return r;
}
static long kvm_gmem_allocate(struct inode *inode, loff_t offset, loff_t len)
@@ -497,6 +550,8 @@ static int kvm_gmem_release(struct inode *inode, struct file *file)
* memory, as its lifetime is associated with the inode, not the file.
*/
__kvm_gmem_invalidate_start(f, 0, -1ul, filter);
+
+ /* No need to prezap since all mappings will be gone */
__kvm_gmem_zap(f, 0, -1ul, filter);
__kvm_gmem_invalidate_end(f, 0, -1ul);
@@ -824,6 +879,14 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
filter = to_private ? KVM_FILTER_SHARED : KVM_FILTER_PRIVATE;
kvm_gmem_invalidate_start(inode, start, end, filter);
+ r = kvm_gmem_prezap(inode, start, end, filter);
+ if (r) {
+ *err_index = start;
+ mas_destroy(&mas);
+ kvm_gmem_invalidate_end(inode, start, end);
+ goto out;
+ }
+
kvm_gmem_zap(inode, start, end, filter);
if (!to_private && kvm_arch_has_gmem_convert())
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 16/17] [GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split S-EPT
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (14 preceding siblings ...)
2026-09-28 9:12 ` [PATCH v4 15/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap() Yan Zhao
@ 2026-09-28 9:12 ` Yan Zhao
2026-09-28 9:12 ` [PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M Yan Zhao
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:12 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
Implement the .gmem_prezap() hook to split S-EPT huge mappings before
guest_memfd zaps S-EPT, in case KVM zaps only a subset of an S-EPT huge
mapping's range. This allows guest_memfd to precisely zap/remove S-EPT
entries to avoid clobbering guest memory, as the lifetime of guest private
memory is tied to the S-EPT. That is, KVM must first split a huge mapping
so that it does not over-zap irrelevant subranges of the huge mapping,
which could cause TD termination.
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
arch/x86/kvm/Kconfig | 1 +
arch/x86/kvm/vmx/tdx.c | 18 ++++++++++++++++--
2 files changed, 17 insertions(+), 2 deletions(-)
diff --git a/arch/x86/kvm/Kconfig b/arch/x86/kvm/Kconfig
index 2c3c22aeafa5..9ac79639dfba 100644
--- a/arch/x86/kvm/Kconfig
+++ b/arch/x86/kvm/Kconfig
@@ -147,6 +147,7 @@ config KVM_INTEL_TDX
default y
depends on INTEL_TDX_HOST
select HAVE_KVM_ARCH_GMEM_POPULATE
+ select HAVE_KVM_ARCH_GMEM_PREZAP
help
Provides support for launching Intel Trust Domain Extensions (TDX)
confidential VMs on Intel processors.
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 5c5919b76bdd..0ecfc17e6235 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -2089,8 +2089,7 @@ static int tdx_sept_split_straddling_huge_pages_to_level(struct kvm *kvm, gfn_t
* per-VM cache of pre-allocated pages used to populate the Dynamic PAMT when
* splitting S-EPT huge pages.
*/
-static int __maybe_unused tdx_sept_split_huge_pages(struct kvm *kvm, gfn_t start,
- gfn_t end)
+static int tdx_sept_split_huge_pages(struct kvm *kvm, gfn_t start, gfn_t end)
{
guard(mutex)(&to_kvm_tdx(kvm)->pamt_cache_lock);
@@ -2103,6 +2102,20 @@ static int __maybe_unused tdx_sept_split_huge_pages(struct kvm *kvm, gfn_t start
return tdx_sept_split_straddling_huge_pages_to_level(kvm, start, end, PG_LEVEL_4K);
}
+/*
+ * Before zapping private mappings, KVM must first split any huge mappings that
+ * straddle the zapping range, as KVM mustn't overzap private mappings for TDX
+ * guests, i.e. KVM must zap _exactly_ [start, end).
+ */
+static int tdx_gmem_prezap(struct kvm *kvm, gfn_t start, gfn_t end,
+ enum kvm_gfn_range_filter attr_filter)
+{
+ if (!(attr_filter & KVM_FILTER_PRIVATE) || !kvm_has_mirrored_tdp(kvm))
+ return 0;
+
+ return tdx_sept_split_huge_pages(kvm, start, end);
+}
+
void tdx_deliver_interrupt(struct kvm_lapic *apic, int delivery_mode,
int trig_mode, int vector)
{
@@ -3799,6 +3812,7 @@ int __init tdx_hardware_setup(void)
vt_x86_ops.set_external_spte = tdx_sept_set_private_spte;
vt_x86_ops.free_external_spt = tdx_sept_free_private_spt;
+ vt_x86_ops.gmem_prezap = tdx_gmem_prezap;
if (tdx_supports_dynamic_pamt(tdx_sysinfo))
vt_x86_ops.topup_external_cache = tdx_topup_external_pamt_cache;
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
* [PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
` (15 preceding siblings ...)
2026-09-28 9:12 ` [PATCH v4 16/17] [GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split S-EPT Yan Zhao
@ 2026-09-28 9:12 ` Yan Zhao
16 siblings, 0 replies; 18+ messages in thread
From: Yan Zhao @ 2026-09-28 9:12 UTC (permalink / raw)
To: seanjc, pbonzini, dave.hansen
Cc: linux-kernel, kvm, x86, rick.p.edgecombe, kas, tabba,
ackerleytng, michael.roth, david, vannapurve, sagis, vbabka,
thomas.lendacky, nik.borisov, pgonda, fan.du, jun.miao,
francescolavra.fl, jgross, xiaoyao.li, kai.huang, binbin.wu,
chao.p.peng, chao.gao, farrah.chen, yan.y.zhao
Introduce a module parameter named "tdx_huge_page" for kvm-intel.ko and
enable TDX huge pages if the module parameter is true and if:
(1) KVM supports gmem in-place conversions,
(2) the TDX module supports uninterruptible demote, and
(3) the TDX module supports DPAMT.
Condition (1) forces TDX huge pages to be the first user of gmem in-place
conversion. Condition (2) is required because the current splitting
implementation depends on the TDX module's non-interruptible demote
feature. Condition (3) simplifies the DEMOTE SEAMCALL's implementation.
When TDX huge page support is enabled, report to KVM MMU that the maximum
allowed mapping level for private memory is 2MB when the TD is RUNNABLE,
while forcing it to 4KB during TD build time. This is because KVM can only
create mappings up to 2MB level in the S-EPT via the TDH_MEM_PAGE_AUG
SEAMCALL, while TDH_MEM_PAGE_ADD mandates the mapping level must be 4KB.
1GB mappings in the S-EPT are only possible via promotion, which is not yet
supported due to the complexity incurred by the TDX module's rules for huge
page promotion.
Signed-off-by: Xiaoyao Li <xiaoyao.li@intel.com>
Signed-off-by: Isaku Yamahata <isaku.yamahata@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yan Zhao <yan.y.zhao@intel.com>
---
v4:
- Disallowed huge page if gmem_in_place_conversion is false. (Sean)
- Updated map part due to MMU refactor. (Sean)
- Disallowed huge page if DPAMT is not enabled. (Rick)
v3:
- Introduce the module param enable_tdx_huge_page and disable to toggle TDX
huge page support.
- Disable TDX huge page if TDX module does not support
TDX_FEATURES0_ENHANCE_DEMOTE_INTERRUPTIBILITY. (Kai).
- Explain why not allow 2M before TD is RUNNABLE in patch log.(Kai)
- Add comment to explain the relationship between returning PG_LEVEL_2M
and guest accept level. (Kai)
- Dropped some KVM_BUG_ON()s due to rebasing. Updated KVM_BUG_ON()s on
mapping levels to take into account of enable_tdx_huge_page.
RFC v2:
- Merged RFC v1's patch 4 (forcing PG_LEVEL_4K before TD runnable) with
patch 9 (allowing PG_LEVEL_2M after TD runnable).
---
arch/x86/kvm/vmx/tdx.c | 41 ++++++++++++++++++++++++++++++++++-------
1 file changed, 34 insertions(+), 7 deletions(-)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 0ecfc17e6235..ee17e50bf439 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -57,6 +57,9 @@
bool enable_tdx __ro_after_init;
module_param_named(tdx, enable_tdx, bool, 0444);
+static bool __read_mostly enable_tdx_huge_page = true;
+module_param_named(tdx_huge_page, enable_tdx_huge_page, bool, 0444);
+
static const struct tdx_sys_info *tdx_sysinfo;
void tdh_vp_rd_failed(struct vcpu_tdx *tdx, char *uclass, u32 field, u64 err)
@@ -1786,8 +1789,9 @@ static int tdx_sept_map_leaf_spte(struct kvm *kvm, gfn_t gfn, enum pg_level leve
if (KVM_BUG_ON(!vcpu, kvm))
return -EIO;
- /* TODO: handle large pages. */
- if (KVM_BUG_ON(level != PG_LEVEL_4K, kvm))
+ /* TODO: Support hugepages when building the initial TD image. */
+ if (KVM_BUG_ON(level != PG_LEVEL_4K &&
+ to_kvm_tdx(kvm)->state != TD_STATE_RUNNABLE, kvm))
return -EIO;
WARN_ON_ONCE((new_spte & VMX_EPT_RWX_MASK) != VMX_EPT_RWX_MASK);
@@ -1883,10 +1887,6 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
if (KVM_BUG_ON(!is_hkid_assigned(to_kvm_tdx(kvm)), kvm))
return -EIO;
- /* TODO: handle large pages. */
- if (KVM_BUG_ON(level != PG_LEVEL_4K, kvm))
- return -EIO;
-
err = tdh_do_no_vcpus(tdh_mem_range_block, kvm, &kvm_tdx->td, gpa,
level, &entry, &level_state);
if (TDX_BUG_ON_2(err, TDH_MEM_RANGE_BLOCK, entry, level_state, kvm))
@@ -3665,12 +3665,34 @@ int tdx_vcpu_ioctl(struct kvm_vcpu *vcpu, void __user *argp)
return ret;
}
+/*
+ * For private pages:
+ *
+ * Force KVM to map at 4KB level when !enable_tdx_huge_page (e.g., due to
+ * incompatible TDX module) or before TD state is RUNNABLE.
+ *
+ * Always allow KVM to map at 2MB level in other cases, though KVM may still map
+ * the page at 4KB (i.e., passing in PG_LEVEL_4K to AUG) due to
+ * (1) the backend folio is 4KB,
+ * (2) disallow_lpage restrictions:
+ * - mixed private/shared pages in the 2MB range
+ * - level misalignment due to slot base_gfn, slot size, and ugfn
+ * - guest_inhibit bit set due to guest's 4KB accept level
+ * (3) page merging is disallowed (e.g., when part of a 2MB range has been
+ * mapped at 4KB level during TD build time).
+ */
int tdx_gmem_max_mapping_level(struct kvm *kvm, kvm_pfn_t pfn, bool is_private)
{
if (!is_private)
return 0;
- return PG_LEVEL_4K;
+ if (!enable_tdx_huge_page)
+ return PG_LEVEL_4K;
+
+ if (unlikely(to_kvm_tdx(kvm)->state != TD_STATE_RUNNABLE))
+ return PG_LEVEL_4K;
+
+ return PG_LEVEL_2M;
}
void tdx_hardware_unsetup(void)
@@ -3749,6 +3771,11 @@ static int __init __tdx_hardware_setup(void)
if (misc_cg_set_capacity(MISC_CG_RES_TDX, tdx_get_nr_guest_keyids()))
return -EINVAL;
+ if (enable_tdx_huge_page && (!gmem_in_place_conversion ||
+ !tdx_huge_page_demote_uninterruptible(tdx_sysinfo) ||
+ !tdx_supports_dynamic_pamt(tdx_sysinfo)))
+ enable_tdx_huge_page = false;
+
return 0;
}
--
2.43.2
^ permalink raw reply [flat|nested] 18+ messages in thread
end of thread, other threads:[~2026-09-28 9:13 UTC | newest]
Thread overview: 18+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 9:07 [PATCH v4 00/17] KVM: TDX huge page support for private memory Yan Zhao
2026-09-28 9:08 ` [PATCH v4 01/17] x86/virt/tdx: Enhance tdx_pamt_get/put() to support huge pages Yan Zhao
2026-09-28 9:08 ` [PATCH v4 02/17] x86/virt/tdx: Add a SEAMCALL wrapper to demote a 2MB huge page Yan Zhao
2026-09-28 9:08 ` [PATCH v4 03/17] KVM: TDX: Reset private huge pages after S-EPT page removal Yan Zhao
2026-09-28 9:09 ` [PATCH v4 04/17] KVM: x86/mmu: Prevent huge page promotion for mirror roots in fault path Yan Zhao
2026-09-28 9:09 ` [PATCH v4 05/17] KVM: x86/tdp_mmu: Alloc external_spt page for mirror page table splitting Yan Zhao
2026-09-28 9:10 ` [PATCH v4 06/17] KVM: x86/mmu: Allocate DPAMT pages for vCPU-induced page split Yan Zhao
2026-09-28 9:10 ` [PATCH v4 07/17] KVM: TDX: Add core support for splitting/demoting 2MB S-EPT mappings to 4KB Yan Zhao
2026-09-28 9:10 ` [PATCH v4 08/17] KVM: TDX: Adjust the topup count of DPAMT page pairs for splitting S-EPT Yan Zhao
2026-09-28 9:10 ` [PATCH v4 09/17] KVM: x86/mmu: Introduce hugepage_set_guest_inhibit() Yan Zhao
2026-09-28 9:11 ` [PATCH v4 10/17] KVM: x86/mmu: Add a TDP MMU API to split huge pages for mirror roots Yan Zhao
2026-09-28 9:11 ` [PATCH v4 11/17] KVM: TDX: Honor the guest's accept level contained in an EPT violation Yan Zhao
2026-09-28 9:11 ` [PATCH v4 12/17] KVM: x86/mmu: Add support for splitting S-EPT entry under non-vCPU context Yan Zhao
2026-09-28 9:11 ` [PATCH v4 13/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add helpers to get start/end gfns give gmem+slot+pgoff Yan Zhao
2026-09-28 9:11 ` [PATCH v4 14/17] [GMEM-DEPENDENT] KVM: guest_memfd: Split kvm_gmem_invalidate_start() to start() and zap() Yan Zhao
2026-09-28 9:12 ` [PATCH v4 15/17] [GMEM-DEPENDENT] KVM: guest_memfd: Add a pre-zap hook .gmem_prezap() Yan Zhao
2026-09-28 9:12 ` [PATCH v4 16/17] [GMEM-DEPENDENT] KVM: TDX: Implement .gmem_prezap() hook to split S-EPT Yan Zhao
2026-09-28 9:12 ` [PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M Yan Zhao
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®