* [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup
@ 2026-09-29 4:02 Zack Rusin
2026-09-29 4:02 ` [PATCH v2 1/6] percpu: Page-align decrypted data in UP kernels Zack Rusin
` (6 more replies)
0 siblings, 7 replies; 9+ messages in thread
From: Zack Rusin @ 2026-09-29 4:02 UTC (permalink / raw)
To: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov
Cc: Ajay Kaher, Alexey Makhalov, Thomas Gleixner, Ingo Molnar,
Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
VMware publishes its per-CPU steal-time GPA without first sharing the
storage in encrypted guests. Following Kiryl's review, this series moves
sharing into early per-CPU initialization for both KVM and VMware.
This replaces the VMware fix from Bo Gan and Alexey Makhalov:
https://lore.kernel.org/r/20260309235250.2611115-5-alexey.makhalov@broadcom.com
The common helper uses boot-time page-table allocation and preserves
initial contents on AMD and TDX. Conversion precedes buffer registration
on SMP and UP; failures stop boot. Encrypted guests use the embedded
allocator, warn on percpu_alloc=page, and never fall back to page mode.
There is no driver conversion loop or readiness flag. .bss..decrypted
remains outside this series.
I interpreted the allocator requirement as guest-specific, so both it
and per-CPU conversion use CC_ATTR_GUEST_MEM_ENCRYPT and leave bare-metal
SME behavior unchanged.
Two corrections to v1: its isolation claim missed UP, where ordinary data
could share the decrypted objects' page; patch 1 fixes that separately.
TDX's vmalloc constraint concerns private aliases and adjacent accesses
through load_unaligned_zeropad(), not just GPA lookup. This series removes
the unused declaration macro and adds explicit SMP/UP section boundaries.
I have not carried the v1 Ack onto the revised patches; renewed review
would be appreciated.
Tested on VMware ESXi with SEV-SNP and TDX guests, and on KVM with
SEV and SEV-SNP guests, using SMP and UP kernels.
v1: https://lore.kernel.org/r/cover.1789488039.git.zack.rusin@broadcom.com
Zack Rusin (6):
percpu: Page-align decrypted data in UP kernels
percpu: Bound decrypted storage for all x86 encrypted guests
x86/percpu: Require embedded allocation in encrypted guests
x86/mm: Provide common early memory decryption
x86/tdx: Support early sharing of kernel data
x86/percpu: Share decrypted storage before guest CPU setup
.../admin-guide/kernel-parameters.txt | 3 +
arch/x86/coco/sev/core.c | 16 ++
arch/x86/coco/tdx/tdx.c | 32 ++++
arch/x86/hyperv/ivm.c | 4 +
arch/x86/include/asm/mem_encrypt.h | 9 +-
arch/x86/include/asm/x86_init.h | 4 +
arch/x86/kernel/cpu/vmware.c | 2 +-
arch/x86/kernel/kvm.c | 35 -----
arch/x86/kernel/setup.c | 2 +
arch/x86/kernel/setup_percpu.c | 14 +-
arch/x86/mm/mem_encrypt.c | 139 ++++++++++++++++++
arch/x86/mm/mem_encrypt_amd.c | 12 +-
include/asm-generic/vmlinux.lds.h | 19 ++-
include/linux/percpu-defs.h | 7 +-
14 files changed, 247 insertions(+), 51 deletions(-)
base-commit: 93f51579e7df248780214094418f205253383cc5
--
2.53.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v2 1/6] percpu: Page-align decrypted data in UP kernels
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
@ 2026-09-29 4:02 ` Zack Rusin
2026-09-29 4:02 ` [PATCH v2 2/6] percpu: Bound decrypted storage for all x86 encrypted guests Zack Rusin
` (5 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Zack Rusin @ 2026-09-29 4:02 UTC (permalink / raw)
To: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov
Cc: Ajay Kaher, Alexey Makhalov, Thomas Gleixner, Ingo Molnar,
Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
Without SMP, DEFINE_PER_CPU_DECRYPTED() places variables in
.data..decrypted. Align both ends of that section so that changing a
variable's encryption attribute cannot expose unrelated kernel data.
Keep the section in permanent data, where UP variables are instantiated.
Fixes: 000f8870a47b ("vmlinux.lds.h: Fix placement of '.data..decrypted' section")
Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
---
include/asm-generic/vmlinux.lds.h | 11 ++++++++++-
1 file changed, 10 insertions(+), 1 deletion(-)
diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h
index b2988aa12f66..64bc2bfdd2ec 100644
--- a/include/asm-generic/vmlinux.lds.h
+++ b/include/asm-generic/vmlinux.lds.h
@@ -368,10 +368,19 @@
/*
* .data section
*/
+#if defined(CONFIG_AMD_MEM_ENCRYPT) && !defined(CONFIG_SMP)
+#define DATA_DECRYPTED \
+ . = ALIGN(PAGE_SIZE); \
+ *(.data..decrypted) \
+ . = ALIGN(PAGE_SIZE);
+#else
+#define DATA_DECRYPTED *(.data..decrypted)
+#endif
+
#define DATA_DATA \
*(.xiptext) \
*(DATA_MAIN) \
- *(.data..decrypted) \
+ DATA_DECRYPTED \
*(.ref.data) \
*(.data..shared_aligned) /* percpu related */ \
*(.data..unlikely) \
--
2.53.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v2 2/6] percpu: Bound decrypted storage for all x86 encrypted guests
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
2026-09-29 4:02 ` [PATCH v2 1/6] percpu: Page-align decrypted data in UP kernels Zack Rusin
@ 2026-09-29 4:02 ` Zack Rusin
2026-09-29 7:54 ` Peter Zijlstra
2026-09-29 4:02 ` [PATCH v2 3/6] x86/percpu: Require embedded allocation in " Zack Rusin
` (4 subsequent siblings)
6 siblings, 1 reply; 9+ messages in thread
From: Zack Rusin @ 2026-09-29 4:02 UTC (permalink / raw)
To: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov
Cc: Ajay Kaher, Alexey Makhalov, Thomas Gleixner, Ingo Molnar,
Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
TDX also needs shared per-CPU buffers. Use X86_MEM_ENCRYPT for their
definition and placement, and provide page-aligned boundaries so the
architecture can convert each CPU's whole section before registration.
Define the boundaries in the SMP template or UP data as appropriate.
Drop the unused DECLARE_PER_CPU_DECRYPTED() macro.
Suggested-by: Kiryl Shutsemau <kas@kernel.org>
Link: https://lore.kernel.org/r/aqqGUAX65s4LdJkr@thinkstation
Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
---
include/asm-generic/vmlinux.lds.h | 12 ++++++++----
include/linux/percpu-defs.h | 7 ++-----
2 files changed, 10 insertions(+), 9 deletions(-)
diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h
index 64bc2bfdd2ec..145fcdbbe9db 100644
--- a/include/asm-generic/vmlinux.lds.h
+++ b/include/asm-generic/vmlinux.lds.h
@@ -368,11 +368,13 @@
/*
* .data section
*/
-#if defined(CONFIG_AMD_MEM_ENCRYPT) && !defined(CONFIG_SMP)
+#if defined(CONFIG_X86_MEM_ENCRYPT) && !defined(CONFIG_SMP)
#define DATA_DECRYPTED \
. = ALIGN(PAGE_SIZE); \
+ __start_percpu_decrypted = .; \
*(.data..decrypted) \
- . = ALIGN(PAGE_SIZE);
+ . = ALIGN(PAGE_SIZE); \
+ __end_percpu_decrypted = .;
#else
#define DATA_DECRYPTED *(.data..decrypted)
#endif
@@ -1022,11 +1024,13 @@
* Note: We use a separate section so that only this section gets
* decrypted to avoid exposing more than we wish.
*/
-#ifdef CONFIG_AMD_MEM_ENCRYPT
+#if defined(CONFIG_X86_MEM_ENCRYPT) && defined(CONFIG_SMP)
#define PERCPU_DECRYPTED_SECTION \
. = ALIGN(PAGE_SIZE); \
+ __start_percpu_decrypted = .; \
*(.data..percpu..decrypted) \
- . = ALIGN(PAGE_SIZE);
+ . = ALIGN(PAGE_SIZE); \
+ __end_percpu_decrypted = .;
#else
#define PERCPU_DECRYPTED_SECTION
#endif
diff --git a/include/linux/percpu-defs.h b/include/linux/percpu-defs.h
index dbe3267a0a13..fdb666a1b4de 100644
--- a/include/linux/percpu-defs.h
+++ b/include/linux/percpu-defs.h
@@ -169,13 +169,10 @@
DEFINE_PER_CPU_SECTION(type, name, "..read_mostly")
/*
- * Declaration/definition used for per-CPU variables that should be accessed
+ * Definition used for built-in per-CPU variables that should be accessed
* as decrypted when memory encryption is enabled in the guest.
*/
-#ifdef CONFIG_AMD_MEM_ENCRYPT
-#define DECLARE_PER_CPU_DECRYPTED(type, name) \
- DECLARE_PER_CPU_SECTION(type, name, "..decrypted")
-
+#ifdef CONFIG_X86_MEM_ENCRYPT
#define DEFINE_PER_CPU_DECRYPTED(type, name) \
DEFINE_PER_CPU_SECTION(type, name, "..decrypted")
#else
--
2.53.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v2 3/6] x86/percpu: Require embedded allocation in encrypted guests
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
2026-09-29 4:02 ` [PATCH v2 1/6] percpu: Page-align decrypted data in UP kernels Zack Rusin
2026-09-29 4:02 ` [PATCH v2 2/6] percpu: Bound decrypted storage for all x86 encrypted guests Zack Rusin
@ 2026-09-29 4:02 ` Zack Rusin
2026-09-29 4:02 ` [PATCH v2 4/6] x86/mm: Provide common early memory decryption Zack Rusin
` (3 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Zack Rusin @ 2026-09-29 4:02 UTC (permalink / raw)
To: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov
Cc: Ajay Kaher, Alexey Makhalov, Thomas Gleixner, Ingo Molnar,
Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
Early sharing needs direct-mapped per-CPU storage. In TDX, a private
direct-map alias of a shared page can terminate the guest, including
through load_unaligned_zeropad(). Require embed rather than adding early
vmalloc conversion and alias synchronization.
Warn when overriding percpu_alloc=page and let embed failures reach the
existing panic. Restrict this policy to encrypted guests, leaving host
SME allocation unchanged.
Suggested-by: Kiryl Shutsemau <kas@kernel.org>
Link: https://lore.kernel.org/r/aqvxQoIoYhJTZpAC@thinkstation
Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
---
Documentation/admin-guide/kernel-parameters.txt | 3 +++
arch/x86/kernel/setup_percpu.c | 12 ++++++++++--
2 files changed, 13 insertions(+), 2 deletions(-)
diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index 33cd30996e47..8ba881af3520 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -5347,6 +5347,9 @@ Kernel parameters
See comments in mm/percpu.c for details on each
allocator. This parameter is primarily for debugging
and performance comparison.
+ On x86 encrypted guests, only "embed" is supported;
+ "page" is ignored with a warning. An embed allocation
+ failure is fatal instead of falling back to "page".
pirq= [SMP,APIC] Manual mp-table setup
See Documentation/arch/x86/i386/IO-APIC.rst.
diff --git a/arch/x86/kernel/setup_percpu.c b/arch/x86/kernel/setup_percpu.c
index bfa48e7a32a2..c83c61e0b20a 100644
--- a/arch/x86/kernel/setup_percpu.c
+++ b/arch/x86/kernel/setup_percpu.c
@@ -8,6 +8,7 @@
#include <linux/percpu.h>
#include <linux/kexec.h>
#include <linux/crash_dump.h>
+#include <linux/cc_platform.h>
#include <linux/smp.h>
#include <linux/topology.h>
#include <linux/pfn.h>
@@ -112,6 +113,7 @@ void __init setup_per_cpu_areas(void)
{
unsigned int cpu;
unsigned long delta;
+ bool encrypted = cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT);
int rc;
pr_info("NR_CPUS:%d nr_cpumask_bits:%d nr_cpu_ids:%u nr_node_ids:%u\n",
@@ -127,6 +129,12 @@ void __init setup_per_cpu_areas(void)
if (pcpu_chosen_fc == PCPU_FC_AUTO && pcpu_need_numa())
pcpu_chosen_fc = PCPU_FC_PAGE;
#endif
+ if (encrypted) {
+ if (pcpu_chosen_fc == PCPU_FC_PAGE)
+ pr_warn("Ignoring percpu_alloc=page in an encrypted guest\n");
+ pcpu_chosen_fc = PCPU_FC_EMBED;
+ }
+
rc = -EINVAL;
if (pcpu_chosen_fc != PCPU_FC_PAGE) {
const size_t dyn_size = PERCPU_MODULE_RESERVE +
@@ -149,11 +157,11 @@ void __init setup_per_cpu_areas(void)
dyn_size, atom_size,
pcpu_cpu_distance,
pcpu_cpu_to_node);
- if (rc < 0)
+ if (rc < 0 && !encrypted)
pr_warn("%s allocator failed (%d), falling back to page size\n",
pcpu_fc_names[pcpu_chosen_fc], rc);
}
- if (rc < 0)
+ if (rc < 0 && !encrypted)
rc = pcpu_page_first_chunk(PERCPU_FIRST_CHUNK_RESERVE,
pcpu_cpu_to_node);
if (rc < 0)
--
2.53.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v2 4/6] x86/mm: Provide common early memory decryption
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
` (2 preceding siblings ...)
2026-09-29 4:02 ` [PATCH v2 3/6] x86/percpu: Require embedded allocation in " Zack Rusin
@ 2026-09-29 4:02 ` Zack Rusin
2026-09-29 4:02 ` [PATCH v2 5/6] x86/tdx: Support early sharing of kernel data Zack Rusin
` (2 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Zack Rusin @ 2026-09-29 4:02 UTC (permalink / raw)
To: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov
Cc: Ajay Kaher, Alexey Makhalov, Thomas Gleixner, Ingo Molnar,
Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
Static per-CPU buffers must be shared before the page allocator is ready.
Provide a page conversion callback with common early page-table splitting,
including the kernel-image alias used by UP data. Flush large translations
before changing attributes and retain the AMD content and SNP transitions.
Drop UP image aliases before SNP kexec makes the backing pages private.
Suggested-by: Kiryl Shutsemau <kas@kernel.org>
Link: https://lore.kernel.org/r/aqqGUAX65s4LdJkr@thinkstation
Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
---
arch/x86/coco/sev/core.c | 16 +++++
arch/x86/include/asm/mem_encrypt.h | 7 ++-
arch/x86/include/asm/x86_init.h | 2 +
arch/x86/mm/mem_encrypt.c | 117 +++++++++++++++++++++++++++++++++++++
arch/x86/mm/mem_encrypt_amd.c | 12 +++-
5 files changed, 149 insertions(+), 5 deletions(-)
diff --git a/arch/x86/coco/sev/core.c b/arch/x86/coco/sev/core.c
index cc292d7c6fd1..bc1afa366c33 100644
--- a/arch/x86/coco/sev/core.c
+++ b/arch/x86/coco/sev/core.c
@@ -541,6 +541,10 @@ static void set_pte_enc(pte_t *kpte, int level, void *va)
set_pte_enc_mask(kpte, d.pfn, d.new_pgprot);
}
+#ifndef CONFIG_SMP
+extern char __start_percpu_decrypted[], __end_percpu_decrypted[];
+#endif
+
static void unshare_all_memory(void)
{
unsigned long addr, end, size, ghcb;
@@ -550,6 +554,18 @@ static void unshare_all_memory(void)
pte_t *pte;
int cpu;
+#ifndef CONFIG_SMP
+ /* Drop image aliases before the direct-map walk makes pages private. */
+ addr = (unsigned long)__start_percpu_decrypted;
+ end = (unsigned long)__end_percpu_decrypted;
+ for (; addr < end; addr += PAGE_SIZE) {
+ pte = lookup_address(addr, &level);
+ if (pte && pte_decrypted(*pte))
+ set_pte(pte, __pte(0));
+ }
+ __flush_tlb_all();
+#endif
+
/* Unshare the direct mapping. */
addr = PAGE_OFFSET;
end = PAGE_OFFSET + get_max_mapped();
diff --git a/arch/x86/include/asm/mem_encrypt.h b/arch/x86/include/asm/mem_encrypt.h
index ea6494628cb0..4d81f693b1d3 100644
--- a/arch/x86/include/asm/mem_encrypt.h
+++ b/arch/x86/include/asm/mem_encrypt.h
@@ -21,9 +21,13 @@ struct boot_params;
#ifdef CONFIG_X86_MEM_ENCRYPT
void __init mem_encrypt_init(void);
void __init mem_encrypt_setup_arch(void);
+int __init early_set_memory_decrypted(unsigned long vaddr, unsigned long size);
+void __init early_set_page_decrypted(unsigned long addr, unsigned long alias);
#else
static inline void mem_encrypt_init(void) { }
static inline void __init mem_encrypt_setup_arch(void) { }
+static inline int __init
+early_set_memory_decrypted(unsigned long vaddr, unsigned long size) { return 0; }
#endif
#ifdef CONFIG_AMD_MEM_ENCRYPT
@@ -50,7 +54,6 @@ void __init sme_early_init(void);
void sme_encrypt_kernel(struct boot_params *bp);
void sme_enable(struct boot_params *bp);
-int __init early_set_memory_decrypted(unsigned long vaddr, unsigned long size);
int __init early_set_memory_encrypted(unsigned long vaddr, unsigned long size);
void __init early_set_mem_enc_dec_hypercall(unsigned long vaddr,
unsigned long size, bool enc);
@@ -86,8 +89,6 @@ static inline void sme_enable(struct boot_params *bp) { }
static inline void sev_es_init_vc_handling(void) { }
-static inline int __init
-early_set_memory_decrypted(unsigned long vaddr, unsigned long size) { return 0; }
static inline int __init
early_set_memory_encrypted(unsigned long vaddr, unsigned long size) { return 0; }
static inline void __init
diff --git a/arch/x86/include/asm/x86_init.h b/arch/x86/include/asm/x86_init.h
index 953d3199408a..e4131402c783 100644
--- a/arch/x86/include/asm/x86_init.h
+++ b/arch/x86/include/asm/x86_init.h
@@ -75,9 +75,11 @@ struct x86_init_oem {
* the kernel pagetables and prepare accessors functions.
* Callback must call paging_init(). Called once after the
* direct mapping for phys memory is available.
+ * @early_decrypt_page: Share a direct-mapped page and its optional image alias
*/
struct x86_init_paging {
void (*pagetable_init)(void);
+ int (*early_decrypt_page)(unsigned long addr, unsigned long alias);
};
/**
diff --git a/arch/x86/mm/mem_encrypt.c b/arch/x86/mm/mem_encrypt.c
index 3aefdef5bcb0..c3e239a47b66 100644
--- a/arch/x86/mm/mem_encrypt.c
+++ b/arch/x86/mm/mem_encrypt.c
@@ -12,10 +12,127 @@
#include <linux/swiotlb.h>
#include <linux/cc_platform.h>
#include <linux/mem_encrypt.h>
+#include <linux/pgalloc.h>
#include <linux/virtio_anchor.h>
#include <linux/iommu-dma.h>
+#include <asm/sections.h>
+#include <asm/set_memory.h>
#include <asm/sev.h>
+#include <asm/tlbflush.h>
+#include <asm/x86_init.h>
+
+#include "mm_internal.h"
+
+static pte_t * __init early_lookup_pte(unsigned long addr)
+{
+ unsigned long pfn, step;
+ unsigned int level, i;
+ pgprot_t prot;
+ pte_t *pte, *table;
+
+ for (;;) {
+ pte = lookup_address(addr, &level);
+ if (!pte || !pte_present(*pte))
+ return NULL;
+ if (level == PG_LEVEL_4K)
+ return pte;
+
+ if (level == PG_LEVEL_2M) {
+ pfn = pmd_pfn(*(pmd_t *)pte);
+ prot = pgprot_large_2_4k(pmd_pgprot(*(pmd_t *)pte));
+ step = 1;
+ } else if (level == PG_LEVEL_1G) {
+ pfn = pud_pfn(*(pud_t *)pte);
+ prot = pud_pgprot(*(pud_t *)pte);
+ step = PMD_SIZE >> PAGE_SHIFT;
+ } else {
+ return NULL;
+ }
+
+ table = alloc_low_page();
+ if (!table)
+ return NULL;
+ for (i = 0; i < PTRS_PER_PTE; i++, pfn += step)
+ set_pte(&table[i], pfn_pte(pfn, prot));
+
+ if (level == PG_LEVEL_2M)
+ pmd_populate_kernel(&init_mm, (pmd_t *)pte, table);
+ else
+ pud_populate(&init_mm, (pud_t *)pte, (pmd_t *)table);
+
+ if (addr - PAGE_OFFSET < get_max_mapped()) {
+ update_page_count(level, -1);
+ update_page_count(level - 1, PTRS_PER_PTE);
+ }
+
+ /* Flush the large translation before changing any attributes. */
+ __flush_tlb_all();
+ }
+}
+
+/* The caller has split both mappings before starting the page transition. */
+void __init early_set_page_decrypted(unsigned long addr, unsigned long alias)
+{
+ unsigned int level;
+ pte_t *pte;
+
+ pte = lookup_address(addr, &level);
+ set_pte(pte, __pte(cc_mkdec(pte_val(*pte))));
+ if (alias) {
+ pte = lookup_address(alias, &level);
+ set_pte(pte, __pte(cc_mkdec(pte_val(*pte))));
+ }
+ __flush_tlb_all();
+}
+
+/* Boot CPU only; size is in bytes and the contents are preserved. */
+int __init early_set_memory_decrypted(unsigned long vaddr, unsigned long size)
+{
+ unsigned long end, addr, alias, pa;
+ pte_t *pte;
+ int ret;
+
+ if (!size || !cc_platform_has(CC_ATTR_MEM_ENCRYPT))
+ return 0;
+ if (!x86_init.paging.early_decrypt_page)
+ return -EOPNOTSUPP;
+ if (size > ULONG_MAX - vaddr)
+ return -EINVAL;
+ end = PAGE_ALIGN(vaddr + size);
+ if (end < vaddr)
+ return -EINVAL;
+
+ for (vaddr &= PAGE_MASK; vaddr < end; vaddr += PAGE_SIZE) {
+ if (vaddr - PAGE_OFFSET >= get_max_mapped() &&
+ (vaddr < (unsigned long)_text || vaddr >= _brk_end))
+ return -EINVAL;
+
+ pa = __pa(vaddr);
+ addr = (unsigned long)__va(pa);
+ pte = early_lookup_pte(addr);
+ if (!pte || pte_pfn(*pte) != PHYS_PFN(pa))
+ return -EFAULT;
+
+ alias = 0;
+ if (pa >= __pa_symbol(_text) &&
+ pa <= __pa_symbol(roundup(_brk_end, PMD_SIZE) - 1)) {
+ alias = (unsigned long)_text + pa - __pa_symbol(_text);
+ if (!early_lookup_pte(alias))
+ return -EFAULT;
+ }
+
+ if (pte_decrypted(*pte)) {
+ early_set_page_decrypted(addr, alias);
+ continue;
+ }
+ ret = x86_init.paging.early_decrypt_page(addr, alias);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
/* Override for DMA direct allocation check - ARCH_HAS_FORCE_DMA_UNENCRYPTED */
bool force_dma_unencrypted(struct device *dev)
diff --git a/arch/x86/mm/mem_encrypt_amd.c b/arch/x86/mm/mem_encrypt_amd.c
index 2f8c32173972..e854f3a039f7 100644
--- a/arch/x86/mm/mem_encrypt_amd.c
+++ b/arch/x86/mm/mem_encrypt_amd.c
@@ -459,9 +459,15 @@ static int __init early_set_memory_enc_dec(unsigned long vaddr,
return ret;
}
-int __init early_set_memory_decrypted(unsigned long vaddr, unsigned long size)
+static int __init amd_early_decrypt_page(unsigned long addr, unsigned long alias)
{
- return early_set_memory_enc_dec(vaddr, size, false);
+ unsigned int level;
+ pte_t *pte = lookup_address(addr, &level);
+
+ __set_clr_pte_enc(pte, PG_LEVEL_4K, false);
+ early_set_page_decrypted(addr, alias);
+ early_set_mem_enc_dec_hypercall(addr, PAGE_SIZE, false);
+ return 0;
}
int __init early_set_memory_encrypted(unsigned long vaddr, unsigned long size)
@@ -479,6 +485,8 @@ void __init sme_early_init(void)
if (!sme_me_mask)
return;
+ x86_init.paging.early_decrypt_page = amd_early_decrypt_page;
+
early_pmd_flags = __sme_set(early_pmd_flags);
__supported_pte_mask = __sme_set(__supported_pte_mask);
--
2.53.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v2 5/6] x86/tdx: Support early sharing of kernel data
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
` (3 preceding siblings ...)
2026-09-29 4:02 ` [PATCH v2 4/6] x86/mm: Provide common early memory decryption Zack Rusin
@ 2026-09-29 4:02 ` Zack Rusin
2026-09-29 4:02 ` [PATCH v2 6/6] x86/percpu: Share decrypted storage before guest CPU setup Zack Rusin
2026-09-29 5:39 ` [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Borislav Petkov
6 siblings, 0 replies; 9+ messages in thread
From: Zack Rusin @ 2026-09-29 4:02 UTC (permalink / raw)
To: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov
Cc: Ajay Kaher, Alexey Makhalov, Thomas Gleixner, Ingo Molnar,
Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
Share pages before the page allocator is available, after both their
direct-map and kernel-image mappings have been made shared. Preserve
initial contents in a private scratch page, matching the AMD early helper,
and use the existing MapGPA path and shared-page accounting.
UP per-CPU data has an image alias. Drop shared image mappings before
kexec converts their backing pages to private through the direct map.
Suggested-by: Kiryl Shutsemau <kas@kernel.org>
Link: https://lore.kernel.org/r/aqqGUAX65s4LdJkr@thinkstation
Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
---
arch/x86/coco/tdx/tdx.c | 32 ++++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/x86/coco/tdx/tdx.c b/arch/x86/coco/tdx/tdx.c
index f904a636d449..4a683e9eafc3 100644
--- a/arch/x86/coco/tdx/tdx.c
+++ b/arch/x86/coco/tdx/tdx.c
@@ -4,6 +4,7 @@
#undef pr_fmt
#define pr_fmt(fmt) "tdx: " fmt
+#include <linux/cacheflush.h>
#include <linux/cpufeature.h>
#include <linux/export.h>
#include <linux/io.h>
@@ -18,6 +19,7 @@
#include <asm/paravirt_types.h>
#include <asm/pgtable.h>
+#include <asm/sections.h>
#include <asm/set_memory.h>
#include <asm/traps.h>
/* MMIO direction */
@@ -1006,6 +1008,23 @@ static int tdx_enc_status_change_finish(unsigned long vaddr, int numpages,
return 0;
}
+static char tdx_early_buffer[PAGE_SIZE] __initdata __aligned(PAGE_SIZE);
+
+static int __init tdx_early_decrypt_page(unsigned long addr, unsigned long alias)
+{
+ void *buffer = __va(__pa_symbol(tdx_early_buffer));
+ int ret;
+
+ memcpy(buffer, (void *)addr, PAGE_SIZE);
+ clflush_cache_range((void *)addr, PAGE_SIZE);
+ early_set_page_decrypted(addr, alias);
+ ret = tdx_enc_status_change_finish(addr, 1, false);
+ if (ret)
+ return ret;
+ memcpy((void *)addr, buffer, PAGE_SIZE);
+ return 0;
+}
+
/* Stop new private<->shared conversions */
static void tdx_kexec_begin(void)
{
@@ -1033,6 +1052,18 @@ static void tdx_kexec_finish(void)
lockdep_assert_irqs_disabled();
+ /* Drop image aliases before the direct-map walk makes pages private. */
+ addr = (unsigned long)_text;
+ while (addr < _brk_end) {
+ unsigned int level;
+ pte_t *pte = lookup_address(addr, &level);
+
+ if (pte && pte_decrypted(*pte))
+ set_pte(pte, __pte(0));
+ addr = (addr & page_level_mask(level)) + page_level_size(level);
+ }
+ __flush_tlb_all();
+
addr = PAGE_OFFSET;
end = PAGE_OFFSET + get_max_mapped();
@@ -1160,6 +1191,7 @@ void __init tdx_early_init(void)
*/
x86_platform.guest.enc_status_change_prepare = tdx_enc_status_change_prepare;
x86_platform.guest.enc_status_change_finish = tdx_enc_status_change_finish;
+ x86_init.paging.early_decrypt_page = tdx_early_decrypt_page;
x86_platform.guest.enc_cache_flush_required = tdx_cache_flush_required;
x86_platform.guest.enc_tlb_flush_required = tdx_tlb_flush_required;
--
2.53.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v2 6/6] x86/percpu: Share decrypted storage before guest CPU setup
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
` (4 preceding siblings ...)
2026-09-29 4:02 ` [PATCH v2 5/6] x86/tdx: Support early sharing of kernel data Zack Rusin
@ 2026-09-29 4:02 ` Zack Rusin
2026-09-29 5:39 ` [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Borislav Petkov
6 siblings, 0 replies; 9+ messages in thread
From: Zack Rusin @ 2026-09-29 4:02 UTC (permalink / raw)
To: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov
Cc: Ajay Kaher, Alexey Makhalov, Thomas Gleixner, Ingo Molnar,
Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
Convert every possible CPU's decrypted section after per-CPU setup and
before the boot CPU registers its buffers. This replaces KVM's object
loop and shares VMware steal-time storage without a driver conversion
path or readiness state.
UP KVM registers inside setup_arch(), so convert before guest_late_init()
there and move VMware's UP registration to that hook. Stop boot on a
conversion failure: inconsistent page state must not be published to the
hypervisor. Host SME keeps its existing mappings. Skip Hyper-V vTOM
per-CPU storage: its visibility callbacks require later Hyper-V
initialization.
Suggested-by: Kiryl Shutsemau <kas@kernel.org>
Link: https://lore.kernel.org/r/aqqGUAX65s4LdJkr@thinkstation
Link: https://lore.kernel.org/r/aqvxQoIoYhJTZpAC@thinkstation
Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
---
arch/x86/hyperv/ivm.c | 4 ++++
arch/x86/include/asm/mem_encrypt.h | 2 ++
arch/x86/include/asm/x86_init.h | 2 ++
arch/x86/kernel/cpu/vmware.c | 2 +-
arch/x86/kernel/kvm.c | 35 -----------------------------------
arch/x86/kernel/setup.c | 2 ++
arch/x86/kernel/setup_percpu.c | 2 ++
arch/x86/mm/mem_encrypt.c | 22 ++++++++++++++++++++++
8 files changed, 35 insertions(+), 36 deletions(-)
diff --git a/arch/x86/hyperv/ivm.c b/arch/x86/hyperv/ivm.c
index 2ce4dfe53472..4a5c735c9c52 100644
--- a/arch/x86/hyperv/ivm.c
+++ b/arch/x86/hyperv/ivm.c
@@ -887,6 +887,10 @@ void __init hv_vtom_init(void)
cc_set_mask(ms_hyperv.shared_gpa_boundary);
physical_mask &= ms_hyperv.shared_gpa_boundary - 1;
+ /* vTOM has no early per-CPU consumers and needs the Hyper-V setup. */
+ x86_init.paging.skip_percpu_decryption = true;
+ x86_init.paging.early_decrypt_page = NULL;
+
x86_platform.hyper.is_private_mmio = hv_is_private_mmio;
x86_platform.guest.enc_cache_flush_required = hv_vtom_cache_flush_required;
x86_platform.guest.enc_tlb_flush_required = hv_vtom_tlb_flush_required;
diff --git a/arch/x86/include/asm/mem_encrypt.h b/arch/x86/include/asm/mem_encrypt.h
index 4d81f693b1d3..05cf395407ef 100644
--- a/arch/x86/include/asm/mem_encrypt.h
+++ b/arch/x86/include/asm/mem_encrypt.h
@@ -21,11 +21,13 @@ struct boot_params;
#ifdef CONFIG_X86_MEM_ENCRYPT
void __init mem_encrypt_init(void);
void __init mem_encrypt_setup_arch(void);
+void __init mem_encrypt_init_percpu(void);
int __init early_set_memory_decrypted(unsigned long vaddr, unsigned long size);
void __init early_set_page_decrypted(unsigned long addr, unsigned long alias);
#else
static inline void mem_encrypt_init(void) { }
static inline void __init mem_encrypt_setup_arch(void) { }
+static inline void __init mem_encrypt_init_percpu(void) { }
static inline int __init
early_set_memory_decrypted(unsigned long vaddr, unsigned long size) { return 0; }
#endif
diff --git a/arch/x86/include/asm/x86_init.h b/arch/x86/include/asm/x86_init.h
index e4131402c783..8d1597372eb6 100644
--- a/arch/x86/include/asm/x86_init.h
+++ b/arch/x86/include/asm/x86_init.h
@@ -76,10 +76,12 @@ struct x86_init_oem {
* Callback must call paging_init(). Called once after the
* direct mapping for phys memory is available.
* @early_decrypt_page: Share a direct-mapped page and its optional image alias
+ * @skip_percpu_decryption: Platform does not use early shared per-CPU data
*/
struct x86_init_paging {
void (*pagetable_init)(void);
int (*early_decrypt_page)(unsigned long addr, unsigned long alias);
+ bool skip_percpu_decryption;
};
/**
diff --git a/arch/x86/kernel/cpu/vmware.c b/arch/x86/kernel/cpu/vmware.c
index 34b73573b108..b477cc027b18 100644
--- a/arch/x86/kernel/cpu/vmware.c
+++ b/arch/x86/kernel/cpu/vmware.c
@@ -366,7 +366,7 @@ static void __init vmware_paravirt_ops_setup(void)
vmware_cpu_down_prepare) < 0)
pr_err("vmware_guest: Failed to install cpu hotplug callbacks\n");
#else
- vmware_guest_cpu_init();
+ x86_init.hyper.guest_late_init = vmware_guest_cpu_init;
#endif
}
}
diff --git a/arch/x86/kernel/kvm.c b/arch/x86/kernel/kvm.c
index 6b0a5861ccb8..acb3b7b18ebe 100644
--- a/arch/x86/kernel/kvm.c
+++ b/arch/x86/kernel/kvm.c
@@ -429,34 +429,6 @@ static u64 kvm_steal_clock(int cpu)
return steal;
}
-static inline __init void __set_percpu_decrypted(void *ptr, unsigned long size)
-{
- early_set_memory_decrypted((unsigned long) ptr, size);
-}
-
-/*
- * Iterate through all possible CPUs and map the memory region pointed
- * by apf_reason, steal_time and kvm_apic_eoi as decrypted at once.
- *
- * Note: we iterate through all possible CPUs to ensure that CPUs
- * hotplugged will have their per-cpu variable already mapped as
- * decrypted.
- */
-static void __init sev_map_percpu_data(void)
-{
- int cpu;
-
- if (cc_vendor != CC_VENDOR_AMD ||
- !cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT))
- return;
-
- for_each_possible_cpu(cpu) {
- __set_percpu_decrypted(&per_cpu(apf_reason, cpu), sizeof(apf_reason));
- __set_percpu_decrypted(&per_cpu(steal_time, cpu), sizeof(steal_time));
- __set_percpu_decrypted(&per_cpu(kvm_apic_eoi, cpu), sizeof(kvm_apic_eoi));
- }
-}
-
static void kvm_guest_cpu_offline(bool shutdown)
{
kvm_disable_steal_time();
@@ -709,12 +681,6 @@ arch_initcall(kvm_alloc_cpumask);
static void __init kvm_smp_prepare_boot_cpu(void)
{
- /*
- * Map the per-cpu variables as decrypted before kvm_guest_cpu_init()
- * shares the guest physical address with the hypervisor.
- */
- sev_map_percpu_data();
-
kvm_guest_cpu_init();
native_smp_prepare_boot_cpu();
kvm_spinlock_init();
@@ -868,7 +834,6 @@ static void __init kvm_guest_init(void)
kvm_cpu_online, kvm_cpu_down_prepare) < 0)
pr_err("failed to install cpu hotplug callbacks\n");
#else
- sev_map_percpu_data();
kvm_guest_cpu_init();
#endif
diff --git a/arch/x86/kernel/setup.c b/arch/x86/kernel/setup.c
index cda6adb9f69c..8eebd85e59af 100644
--- a/arch/x86/kernel/setup.c
+++ b/arch/x86/kernel/setup.c
@@ -1251,6 +1251,8 @@ void __init setup_arch(char **cmdline_p)
io_apic_init_mappings();
+ if (!IS_ENABLED(CONFIG_SMP))
+ mem_encrypt_init_percpu();
x86_init.hyper.guest_late_init();
e820__reserve_resources();
diff --git a/arch/x86/kernel/setup_percpu.c b/arch/x86/kernel/setup_percpu.c
index c83c61e0b20a..8526ad37a81b 100644
--- a/arch/x86/kernel/setup_percpu.c
+++ b/arch/x86/kernel/setup_percpu.c
@@ -5,6 +5,7 @@
#include <linux/export.h>
#include <linux/init.h>
#include <linux/memblock.h>
+#include <linux/mem_encrypt.h>
#include <linux/percpu.h>
#include <linux/kexec.h>
#include <linux/crash_dump.h>
@@ -234,4 +235,5 @@ void __init setup_per_cpu_areas(void)
* this call?
*/
sync_initial_page_table();
+ mem_encrypt_init_percpu();
}
diff --git a/arch/x86/mm/mem_encrypt.c b/arch/x86/mm/mem_encrypt.c
index c3e239a47b66..d3224607170d 100644
--- a/arch/x86/mm/mem_encrypt.c
+++ b/arch/x86/mm/mem_encrypt.c
@@ -13,6 +13,7 @@
#include <linux/cc_platform.h>
#include <linux/mem_encrypt.h>
#include <linux/pgalloc.h>
+#include <linux/percpu.h>
#include <linux/virtio_anchor.h>
#include <linux/iommu-dma.h>
@@ -24,6 +25,8 @@
#include "mm_internal.h"
+extern char __percpu __start_percpu_decrypted[], __end_percpu_decrypted[];
+
static pte_t * __init early_lookup_pte(unsigned long addr)
{
unsigned long pfn, step;
@@ -134,6 +137,25 @@ int __init early_set_memory_decrypted(unsigned long vaddr, unsigned long size)
return 0;
}
+void __init mem_encrypt_init_percpu(void)
+{
+ unsigned long size = __end_percpu_decrypted - __start_percpu_decrypted;
+ int cpu, ret;
+
+ if (!cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT) ||
+ x86_init.paging.skip_percpu_decryption)
+ return;
+
+ for_each_possible_cpu(cpu) {
+ unsigned long addr = (unsigned long)
+ per_cpu_ptr(__start_percpu_decrypted, cpu);
+
+ ret = early_set_memory_decrypted(addr, size);
+ if (ret)
+ panic("Cannot share CPU %d per-CPU data (err=%d)", cpu, ret);
+ }
+}
+
/* Override for DMA direct allocation check - ARCH_HAS_FORCE_DMA_UNENCRYPTED */
bool force_dma_unencrypted(struct device *dev)
{
--
2.53.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
` (5 preceding siblings ...)
2026-09-29 4:02 ` [PATCH v2 6/6] x86/percpu: Share decrypted storage before guest CPU setup Zack Rusin
@ 2026-09-29 5:39 ` Borislav Petkov
6 siblings, 0 replies; 9+ messages in thread
From: Borislav Petkov @ 2026-09-29 5:39 UTC (permalink / raw)
To: Zack Rusin
Cc: Kiryl Shutsemau, x86, Dennis Zhou, Tejun Heo, Arnd Bergmann,
Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui, Paolo Bonzini,
Vitaly Kuznetsov, Ajay Kaher, Alexey Makhalov, Thomas Gleixner,
Ingo Molnar, Dave Hansen, H. Peter Anvin, virtualization,
bcm-kernel-feedback-list, linux-kernel, Christoph Lameter,
Andrew Morton, Bo Gan, linux-mm, linux-arch, linux-coco, kvm,
Jonathan Corbet, K. Y. Srinivasan, Haiyang Zhang, Long Li,
Andy Lutomirski, Peter Zijlstra, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
On Tue, Sep 29, 2026 at 12:02:49AM -0400, Zack Rusin wrote:
> VMware publishes its per-CPU steal-time GPA without first sharing the
> storage in encrypted guests.
So this one liner is the only explanation why this patchset exists, AFAICT. So
I asked AI. I'm pasting what it said below, at the end.
Is any of it true?
If so, how much and is that the reason this patchset exists?
I.e., you want for vmware monitoring tools to work with CoCo guests.
Yes, no? Anything else?
And looking at your commit messages, they don't really talk about why the
patches exist.
And I don't think you know your audience - you're sending a bunch of patches
touching arch/x86/ and you throw all that virt gibberish around like it ain't
no tomorrow and you're thinking that tip people can follow. And I think we can
follow only bits and pieces but all of us will be mostly head-scratching here
and probably ignore the whole set.
So perhaps you should lay off your virt hat for a minute and try to explain in
an approachable way what you're trying to do so that even non-virt people can
follow.
Right now concepts are flying around left and right in that text and I have no
clue what's going on here.
Thx.
P.S., here's the AI gunk:
"In this context, “VMware per‑CPU steal‑time GPA” is referring to **Guest
Physical Address (GPA)**–based accounting of **CPU steal time per vCPU**
inside a VMware virtual machine.
Let’s break the components down:
- **Steal time**
In virtualization, “steal time” is the amount of time a guest CPU *wanted*
to run but **was not scheduled by the hypervisor** because the physical CPU
was busy running other VMs or host tasks.
Per‑CPU steal time means this is measured **for each vCPU** separately.
- **GPA (Guest Physical Address)**
GPA is the CPU’s view of “physical” memory inside the guest. It’s:
- The address space the guest OS uses for RAM, MMIO, etc.
- Translated by the hypervisor to real host physical memory (or nested translations with EPT/NPT).
- **Why GPA for steal time?**
VMware (and other hypervisors) can expose performance statistics (like steal time, idle time, run time) via:
- Per‑vCPU counters/timers
- Memory‑mapped (GPA‑mapped) structures that the guest can read
- Paravirtual interfaces (e.g., hypercalls, PV drivers) which are often described or backed by GPA regions
So when someone says “VMware per‑CPU steal-time GPA,” they usually mean:
A VMware mechanism where **per‑vCPU steal‑time metrics** are exposed to the
guest via data structures located in **guest physical address space**, so the
guest OS or monitoring tools can read those counters and attribute lost CPU
time per vCPU."
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v2 2/6] percpu: Bound decrypted storage for all x86 encrypted guests
2026-09-29 4:02 ` [PATCH v2 2/6] percpu: Bound decrypted storage for all x86 encrypted guests Zack Rusin
@ 2026-09-29 7:54 ` Peter Zijlstra
0 siblings, 0 replies; 9+ messages in thread
From: Peter Zijlstra @ 2026-09-29 7:54 UTC (permalink / raw)
To: Zack Rusin
Cc: Kiryl Shutsemau, Borislav Petkov, x86, Dennis Zhou, Tejun Heo,
Arnd Bergmann, Rick Edgecombe, Tom Lendacky, Wei Liu, Dexuan Cui,
Paolo Bonzini, Vitaly Kuznetsov, Ajay Kaher, Alexey Makhalov,
Thomas Gleixner, Ingo Molnar, Dave Hansen, H. Peter Anvin,
virtualization, bcm-kernel-feedback-list, linux-kernel,
Christoph Lameter, Andrew Morton, Bo Gan, linux-mm, linux-arch,
linux-coco, kvm, Jonathan Corbet, K. Y. Srinivasan,
Haiyang Zhang, Long Li, Andy Lutomirski, linux-doc, linux-hyperv,
Nathan Chancellor, Kees Cook, Ashish Kalra
On Tue, Sep 29, 2026 at 12:02:51AM -0400, Zack Rusin wrote:
> TDX also needs shared per-CPU buffers. Use X86_MEM_ENCRYPT for their
> definition and placement, and provide page-aligned boundaries so the
> architecture can convert each CPU's whole section before registration.
>
> Define the boundaries in the SMP template or UP data as appropriate.
> Drop the unused DECLARE_PER_CPU_DECRYPTED() macro.
>
> Suggested-by: Kiryl Shutsemau <kas@kernel.org>
> Link: https://lore.kernel.org/r/aqqGUAX65s4LdJkr@thinkstation
> Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
> ---
> include/asm-generic/vmlinux.lds.h | 12 ++++++++----
> include/linux/percpu-defs.h | 7 ++-----
> 2 files changed, 10 insertions(+), 9 deletions(-)
>
> diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinux.lds.h
> index 64bc2bfdd2ec..145fcdbbe9db 100644
> --- a/include/asm-generic/vmlinux.lds.h
> +++ b/include/asm-generic/vmlinux.lds.h
> @@ -1022,11 +1024,13 @@
> * Note: We use a separate section so that only this section gets
> * decrypted to avoid exposing more than we wish.
> */
> -#ifdef CONFIG_AMD_MEM_ENCRYPT
> +#if defined(CONFIG_X86_MEM_ENCRYPT) && defined(CONFIG_SMP)
> #define PERCPU_DECRYPTED_SECTION \
> . = ALIGN(PAGE_SIZE); \
> + __start_percpu_decrypted = .; \
> *(.data..percpu..decrypted) \
> - . = ALIGN(PAGE_SIZE);
> + . = ALIGN(PAGE_SIZE); \
> + __end_percpu_decrypted = .;
> #else
> #define PERCPU_DECRYPTED_SECTION
> #endif
So you're page aligning something that will get different protection and
will thus shatter large pages?
That is somewhat uncool. We like large pages, large pages are good.
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-09-29 7:54 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
2026-09-29 4:02 ` [PATCH v2 1/6] percpu: Page-align decrypted data in UP kernels Zack Rusin
2026-09-29 4:02 ` [PATCH v2 2/6] percpu: Bound decrypted storage for all x86 encrypted guests Zack Rusin
2026-09-29 7:54 ` Peter Zijlstra
2026-09-29 4:02 ` [PATCH v2 3/6] x86/percpu: Require embedded allocation in " Zack Rusin
2026-09-29 4:02 ` [PATCH v2 4/6] x86/mm: Provide common early memory decryption Zack Rusin
2026-09-29 4:02 ` [PATCH v2 5/6] x86/tdx: Support early sharing of kernel data Zack Rusin
2026-09-29 4:02 ` [PATCH v2 6/6] x86/percpu: Share decrypted storage before guest CPU setup Zack Rusin
2026-09-29 5:39 ` [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Borislav Petkov
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®