* [PATCH v4 0/4] KVM: TDX: Validate directly configurable CPUID bits
@ 2026-09-17 7:25 Binbin Wu
2026-09-17 7:25 ` [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Binbin Wu
` (3 more replies)
0 siblings, 4 replies; 27+ messages in thread
From: Binbin Wu @ 2026-09-17 7:25 UTC (permalink / raw)
To: linux-kernel, kvm
Cc: seanjc, pbonzini, dave.hansen, andrew.cooper3, nik.borisov, kas,
rick.p.edgecombe, xiaoyao.li, chao.gao, tony.lindgren,
kishen.maloor, dedekind1, binbin.wu
The TDX module reports a set of CPUID bits that the VMM can directly
configure for a TD, but KVM cannot blindly expose all module-supported
bits to userspace. Certain features imply additional architectural state
(such as one or more MSRs) that KVM must explicitly manage across
host/guest transitions to prevent host state corruption.
Today, KVM hardcodes a denylist to filter out host state clobbering
features that KVM doesn't support. However, a denylist doesn't scale as
the TDX module evolves, i.e. host state clobbering features added by a
future TDX module are not covered by the denylist, and userspace could
enable a feature that KVM doesn't support.
This series builds an allowlist for directly configurable CPUID bits, and
filters and validates userspace input against it, so that newly introduced
TDX directly configurable CPUID feature bits can only be enabled by
userspace after KVM explicitly opts in. Except for a few fixed-1 bits
required for basic TDX support, host state clobbering features are either
directly configurable or gated by TD ATTRIBUTES/XFAM. KVM already
validates ATTRIBUTES/XFAM, so the allowlist only needs to cover directly
configurable bits.
For this version, I'm hoping to collect some RBs and Acks. The key
discussions from v3 [1] are summarized below. Note, FRED KVM support is
in progress upstream [2], and no TDX module with FRED support exists yet,
so there is no immediate issue. But it would be nice to get this series
merged sooner rather than later.
Discussions from v3:
- Whether to restrict the TDX feature bits to kvm_cpu_caps[].
Xiaoyao and Sean preferred not to restrict the TDX feature bits to
kvm_cpu_caps[], to avoid an unnecessary dependency on feature enabling
for non-TDX VMs [3][4]. However, Rick pointed out that ideally the TDX
arch would not be finalized before the normal VM KVM design reaches some
level of maturity [5]. Also, it's safer for TDX to respect KVM's logic
around disabling certain features, since the reasons for disabling them
could apply to TDX as well [6].
This version still restricts the TDX feature bits to kvm_cpu_caps[], with
a few exceptions. If TDX feature enabling turns out to be consistently
blocked by non-TDX VM enabling, a solution can be worked out then.
Note that the final allowlist is the same either way, i.e. whether or
not the TDX feature bits are restricted to kvm_cpu_caps[].
- New opt-in interface vs. backporting.
There are two options for old KVM versions that don't do the filtering
and validation:
1) Add a new opt-in interface to expose features newly introduced by the
TDX module.
2) Backport this series to LTS/stable branches.
The opt-in design vs. backport decision is deferred, since this series is
correct and an improvement either way [7].
- Potential backward compatibility issue when an effective fixed-1 bit
becomes directly configurable due to x86 feature deprecation.
This is left for the future, i.e. it can be addressed if and when such a
deprecation actually happens, e.g. by adding an opt-in interface to let
users/admins choose between migration flexibility and backwards
compatibility on old platforms [8].
Changes since v3:
- Audited the directly configurable CPUID bits outside of kvm_cpu_caps[],
added XTPR and HT to the allowlist, in addition to MWAIT and
CORE_CAPABILITIES. The remaining bits are left out of the allowlist
because they are either not supported or can't be properly virtualized.
- Improved comments/changelogs. (Xiaoyao, Kishen)
- Updated documentation/comment about the ABI change. (Xiaoyao, Rick)
- Renamed functions for better readability. (Tony)
- Collected two RBs from Tony.
The series follows the CSV version of the CPUID virtualization document
from "Intel TDX Module ABI Definitions" [9], updated in June 2026.
AI tools were used to:
- Polish the cover letter and commit messages.
- Do code review.
[1] v3: https://lore.kernel.org/all/20260827031837.2863609-1-binbin.wu@linux.intel.com
v2: https://lore.kernel.org/all/20260604023314.3907511-1-binbin.wu@linux.intel.com
v1: https://lore.kernel.org/all/20260417073610.3246316-1-binbin.wu@linux.intel.com
[2] https://lore.kernel.org/all/20260911213659.2025974-1-sohil.mehta@intel.com
[3] https://lore.kernel.org/all/9abeed40-bcdc-48d9-8cbb-cb87154bb9f6@intel.com
[4] https://lore.kernel.org/all/aqHdwkOFVR4XoS2M@google.com
[5] https://lore.kernel.org/all/c967dc3eb6b128055e415bae45d1c14a52dc4acd.camel@intel.com
[6] https://lore.kernel.org/all/4c379b90da893c5d1ffea46dc7e40303e7923872.camel@intel.com
[7] https://lore.kernel.org/all/c6b240bd-3c45-4abe-b17c-72a1a7da582f@linux.intel.com
[8] https://lore.kernel.org/all/1deddc78d14326b378ddd62ad99c21d66da401e6.camel@intel.com
[9] https://cdrdv2.intel.com/v1/dl/getContent/795381
Binbin Wu (4):
KVM: TDX: Track configurable CPUID bits allowed by KVM
KVM: TDX: Report CORE_CAPABILITIES as configurable
KVM: TDX: Filter configurable CPUID bits
KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM
Documentation/virt/kvm/x86/intel-tdx.rst | 10 +
arch/x86/include/uapi/asm/kvm.h | 2 +
arch/x86/kvm/vmx/tdx.c | 279 ++++++++++++++++++++---
3 files changed, 262 insertions(+), 29 deletions(-)
base-commit: 70c944caf570fda2d79baa71435589a8db39f048
--
2.46.0
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-17 7:25 [PATCH v4 0/4] KVM: TDX: Validate directly configurable CPUID bits Binbin Wu
@ 2026-09-17 7:25 ` Binbin Wu
2026-09-23 0:01 ` Edgecombe, Rick P
2026-09-24 6:36 ` Xiaoyao Li
2026-09-17 7:25 ` [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Binbin Wu
` (2 subsequent siblings)
3 siblings, 2 replies; 27+ messages in thread
From: Binbin Wu @ 2026-09-17 7:25 UTC (permalink / raw)
To: linux-kernel, kvm
Cc: seanjc, pbonzini, dave.hansen, andrew.cooper3, nik.borisov, kas,
rick.p.edgecombe, xiaoyao.li, chao.gao, tony.lindgren,
kishen.maloor, dedekind1, binbin.wu
Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable
CPUID feature bits that KVM supports, and build the masks during TDX
hardware setup via tdx_initialize_cpu_cfg_caps().
The TDX module reports the CPUID bits that the VMM can directly configure
for a TD, but KVM cannot blindly expose all reported bits to userspace.
Certain features imply additional architectural state, e.g. one or more
MSRs, that KVM must explicitly manage across host/guest transitions to
prevent host state corruption. The existing hardcoded denylist cannot
account for new host state clobbering features introduced by future TDX
modules.
Except for a few fixed-1 bits required for basic TDX support, host state
clobbering features are either directly configurable or gated by TD
ATTRIBUTES/XFAM, which KVM already validates. Tracking only the directly
configurable bits to build an allowlist is therefore sufficient.
Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can
be built with similar feature-name based initializers. Directly
configurable non-feature bits will be handled separately.
The allowlist is prepared to be consumed by later patches to filter
KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to
KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID
feature bits stay hidden from userspace until KVM explicitly opts in.
By default, intersect the allowlist with kvm_cpu_caps[] via TDX_CFG_F().
Requiring support for non-TDX VMs avoids committing to TDX-specific
behavior before general KVM support is established, and respects KVM's
logic around disabling certain features, since the reasons for disabling
them could apply to TDX as well. Allow exceptions through
TDX_CFG_EXTRA_F() only with sufficient justification.
Add comments as placeholders for HLE, RTM, WAITPKG and FRED, which KVM
doesn't support for TDX yet.
Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
not advertised in kvm_cpu_caps[]. The remaining directly configurable
feature bits outside kvm_cpu_caps[] are left out of the allowlist:
- Features forced to zero when #VE is reduced, or lacking KVM support
for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
in a subsequent patch.
- Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
- Features that can clobber host state and lack KVM support for TDX:
FRED.
- Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
platform), and RAO_INT (defined only for future processors).
Filtering KVM_TDX_CAPABILITIES in a subsequent patch will intentionally
stop advertising the excluded bits as configurable, as the corresponding
features are unsupported or cannot be properly virtualized.
Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
---
v4:
- Add XTPR and HT to the allowlist.
- Explain in the changelog why some directly configurable bits are not
added to the allowlist.
- Improve comments/changelog. (Xiaoyao)
- Fix the bug that AMX_COMPLEX should be added for CPUID_1E_1_EAX instead
of CPUID_7_1_EDX.
v3:
- Drop the new data structure in v2 and only track feature bits by
following the organization of kvm_cpu_caps[], handle non-feature
bits separately. (Sean)
- Use two versions of macros (TDX_CFG_F() vs. TDX_CFG_EXTRA_F()) to
distinguish whether a supported TDX configurable CPUID bit should be
checked against KVM's common cpu capabilities.
- Add AMX_COMPLEX since it has been defined in the CPUID virtualization doc.
---
arch/x86/kvm/vmx/tdx.c | 164 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 164 insertions(+)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 8a3f88b79b83d..2b51a85c998e8 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -52,6 +52,168 @@
__TDX_BUG_ON(__err, #__fn, __kvm, ", " #a1 " 0x%llx, " #a2 ", 0x%llx, " #a3 " 0x%llx", \
a1, a2, a3)
+static u32 tdx_cpu_cfg_caps[NR_KVM_CPU_CAPS] __ro_after_init;
+static_assert(ARRAY_SIZE(tdx_cpu_cfg_caps) == ARRAY_SIZE(kvm_cpu_caps));
+
+#define TDX_VALIDATE_CPU_CAP_USAGE(name) \
+ BUILD_BUG_ON(__feature_leaf(X86_FEATURE_##name) != \
+ tdx_cpu_cap_init_in_progress)
+
+/* For a feature bit that needs to be cap'ed by kvm_cpu_caps[]. */
+#define TDX_CFG_F(name) \
+({ \
+ TDX_VALIDATE_CPU_CAP_USAGE(name); \
+ tdx_cfg_caps |= feature_bit(name); \
+})
+
+/*
+ * For a feature bit that KVM allows for TDX guests even though it isn't
+ * advertised through kvm_cpu_caps[], e.g. MWAIT. Use this version only when
+ * there is a justification.
+ */
+#define TDX_CFG_EXTRA_F(name) \
+({ \
+ TDX_VALIDATE_CPU_CAP_USAGE(name); \
+ tdx_cfg_caps_extra |= feature_bit(name); \
+})
+
+#define tdx_cpu_cfg_cap_init(leaf, feature_initializers...) \
+do { \
+ const u32 __maybe_unused tdx_cpu_cap_init_in_progress = leaf; \
+ u32 tdx_cfg_caps_extra = 0; \
+ u32 tdx_cfg_caps = 0; \
+ \
+ feature_initializers \
+ tdx_cpu_cfg_caps[leaf] = (tdx_cfg_caps & kvm_cpu_caps[leaf]) | \
+ tdx_cfg_caps_extra; \
+} while (0)
+
+/*
+ * Initialize tdx_cpu_cfg_caps[], the list of CPUID features that KVM
+ * supports for TDX guests. It covers only the directly configurable CPUID
+ * bits reported by the TDX module; features controlled by XFAM and
+ * ATTRIBUTES are handled separately.
+ */
+static void __init tdx_initialize_cpu_cfg_caps(void)
+{
+ tdx_cpu_cfg_cap_init(CPUID_1_ECX,
+ /*
+ * KVM allows userspace to enumerate MONITOR+MWAIT support to
+ * the guest, but the MWAIT feature flag is never advertised
+ * to userspace for non-TDX VMs.
+ */
+ TDX_CFG_EXTRA_F(MWAIT),
+ /*
+ * XTPR can be exposed to a TD, but it never takes effect in
+ * the underlying hardware when the TD changes
+ * IA32_MISC_ENABLE[23].
+ */
+ TDX_CFG_EXTRA_F(XTPR),
+ TDX_CFG_F(TSC_DEADLINE_TIMER),
+ TDX_CFG_F(AVX),
+ TDX_CFG_F(F16C),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_1_EDX,
+ TDX_CFG_F(MCE),
+ TDX_CFG_F(MTRR),
+ TDX_CFG_F(MCA),
+ TDX_CFG_F(SELFSNOOP),
+ /*
+ * HT is a topology enumeration bit that KVM doesn't care
+ * about, but userspace may want to expose it to the guests.
+ */
+ TDX_CFG_EXTRA_F(HT),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_7_0_EBX,
+ TDX_CFG_F(BMI1),
+ /* HLE */
+ TDX_CFG_F(BMI2),
+ TDX_CFG_F(ERMS),
+ /* RTM */
+ TDX_CFG_F(AVX512F),
+ TDX_CFG_F(AVX512DQ),
+ TDX_CFG_F(ADX),
+ TDX_CFG_F(AVX512IFMA),
+ TDX_CFG_F(AVX512PF),
+ TDX_CFG_F(AVX512ER),
+ TDX_CFG_F(AVX512CD),
+ TDX_CFG_F(AVX512BW),
+ TDX_CFG_F(AVX512VL),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_7_ECX,
+ TDX_CFG_F(UMIP),
+ /* WAITPKG */
+ TDX_CFG_F(AVX512_VBMI2),
+ TDX_CFG_F(GFNI),
+ TDX_CFG_F(VAES),
+ TDX_CFG_F(VPCLMULQDQ),
+ TDX_CFG_F(AVX512_VNNI),
+ TDX_CFG_F(AVX512_BITALG),
+ TDX_CFG_F(AVX512_VPOPCNTDQ),
+ TDX_CFG_F(LA57),
+ TDX_CFG_F(RDPID),
+ TDX_CFG_F(CLDEMOTE),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_7_EDX,
+ TDX_CFG_F(AVX512_4VNNIW),
+ TDX_CFG_F(AVX512_4FMAPS),
+ TDX_CFG_F(FSRM),
+ TDX_CFG_F(AVX512_VP2INTERSECT),
+ TDX_CFG_F(SERIALIZE),
+ TDX_CFG_F(TSXLDTRK),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_7_1_EAX,
+ TDX_CFG_F(SHA512),
+ TDX_CFG_F(SM3),
+ TDX_CFG_F(SM4),
+ TDX_CFG_F(AVX_VNNI),
+ TDX_CFG_F(AVX512_BF16),
+ TDX_CFG_F(CMPCCXADD),
+ TDX_CFG_F(FZRM),
+ TDX_CFG_F(FSRS),
+ TDX_CFG_F(FSRC),
+ /* FRED */
+ TDX_CFG_F(LKGS),
+ TDX_CFG_F(WRMSRNS),
+ TDX_CFG_F(AMX_FP16),
+ TDX_CFG_F(AVX_IFMA),
+ TDX_CFG_F(LAM),
+ TDX_CFG_F(MOVRS),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_7_1_EDX,
+ TDX_CFG_F(AVX_VNNI_INT8),
+ TDX_CFG_F(AVX_NE_CONVERT),
+ TDX_CFG_F(AVX_VNNI_INT16),
+ TDX_CFG_F(PREFETCHITI),
+ TDX_CFG_F(AVX10),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_7_2_EDX,
+ TDX_CFG_F(DDPD_U),
+ TDX_CFG_F(MCDT_NO),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_1E_1_EAX,
+ TDX_CFG_F(AMX_COMPLEX_ALIAS),
+ TDX_CFG_F(AMX_FP8),
+ TDX_CFG_F(AMX_TF32),
+ TDX_CFG_F(AMX_AVX512),
+ TDX_CFG_F(AMX_MOVRS),
+ );
+
+ tdx_cpu_cfg_cap_init(CPUID_8000_0008_EBX,
+ TDX_CFG_F(WBNOINVD),
+ );
+}
+
+#undef TDX_CFG_F
+#undef TDX_CFG_EXTRA_F
bool enable_tdx __ro_after_init;
module_param_named(tdx, enable_tdx, bool, 0444);
@@ -3487,6 +3649,8 @@ int __init tdx_hardware_setup(void)
return r;
}
+ tdx_initialize_cpu_cfg_caps();
+
KVM_SANITY_CHECK_VM_STRUCT_SIZE(kvm_tdx);
vt_x86_ops.vm_size = max_t(unsigned int, vt_x86_ops.vm_size, sizeof(struct kvm_tdx));
--
2.46.0
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable
2026-09-17 7:25 [PATCH v4 0/4] KVM: TDX: Validate directly configurable CPUID bits Binbin Wu
2026-09-17 7:25 ` [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Binbin Wu
@ 2026-09-17 7:25 ` Binbin Wu
2026-09-22 21:11 ` Edgecombe, Rick P
2026-09-24 6:44 ` Xiaoyao Li
2026-09-17 7:25 ` [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits Binbin Wu
2026-09-17 7:25 ` [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM Binbin Wu
3 siblings, 2 replies; 27+ messages in thread
From: Binbin Wu @ 2026-09-17 7:25 UTC (permalink / raw)
To: linux-kernel, kvm
Cc: seanjc, pbonzini, dave.hansen, andrew.cooper3, nik.borisov, kas,
rick.p.edgecombe, xiaoyao.li, chao.gao, tony.lindgren,
kishen.maloor, dedekind1, binbin.wu
Add CORE_CAPABILITIES (CPUID.0x7.0.EDX[30]) to KVM's allowlist of TDX
directly configurable CPUID feature bits, even though KVM doesn't support
MSR_IA32_CORE_CAPS for TDX guests, to accommodate the legacy TDX module
definition and userspace's stale knowledge of it.
Older TDX specifications define the CORE_CAPABILITIES CPUID bit as
fixed-1, so userspace may expect the bit to be enabled for TDs. #VE
reduction turns it into a directly configurable bit, so leaving it out of
the allowlist would make the bit impossible to enable once KVM starts
validating userspace's CPUID input, i.e. would be a surprising behavior
change for such userspace.
Reporting CORE_CAPABILITIES as directly configurable also lets userspace
detect that the bit is no longer fixed-1, and thus correct its stale
knowledge.
Keep MSR_IA32_CORE_CAPS unsupported for TDs, as no existing TDX user needs
guest access to the MSR.
Note, CORE_CAPABILITIES is the only bit that is unsupported by KVM *and*
changed from fixed-1 to directly configurable by #VE reduction, and no
further #VE reductions are expected.
Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
Reviewed-by: Tony Lindgren <tony.lindgren@linux.intel.com>
---
v4:
- Add #VE reduction related background to the changelog. (Kishen)
- Add RB from Tony.
v3:
- Drop the code for MSR_IA32_CORE_CAPS access.
---
arch/x86/kvm/vmx/tdx.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 2b51a85c998e8..b34afc52b714e 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -165,6 +165,13 @@ static void __init tdx_initialize_cpu_cfg_caps(void)
TDX_CFG_F(AVX512_VP2INTERSECT),
TDX_CFG_F(SERIALIZE),
TDX_CFG_F(TSXLDTRK),
+ /*
+ * KVM doesn't support MSR_IA32_CORE_CAPS, but older TDX specs
+ * define this bit as fixed-1. Report it as configurable to
+ * accommodate the legacy TDX module definition, and to let
+ * userspace detect that the bit is no longer fixed-1.
+ */
+ TDX_CFG_EXTRA_F(CORE_CAPABILITIES),
);
tdx_cpu_cfg_cap_init(CPUID_7_1_EAX,
--
2.46.0
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits
2026-09-17 7:25 [PATCH v4 0/4] KVM: TDX: Validate directly configurable CPUID bits Binbin Wu
2026-09-17 7:25 ` [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Binbin Wu
2026-09-17 7:25 ` [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Binbin Wu
@ 2026-09-17 7:25 ` Binbin Wu
2026-09-23 0:16 ` Edgecombe, Rick P
2026-09-17 7:25 ` [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM Binbin Wu
3 siblings, 1 reply; 27+ messages in thread
From: Binbin Wu @ 2026-09-17 7:25 UTC (permalink / raw)
To: linux-kernel, kvm
Cc: seanjc, pbonzini, dave.hansen, andrew.cooper3, nik.borisov, kas,
rick.p.edgecombe, xiaoyao.li, chao.gao, tony.lindgren,
kishen.maloor, dedekind1, binbin.wu
Filter the directly configurable CPUID bits reported through
KVM_TDX_CAPABILITIES against KVM's TDX allowlist, and drop the hardcoded
denylist based filtering.
The TDX module reports directly configurable CPUID bits that it supports
for a TD, but KVM must not expose bits that it doesn't support, as blindly
exposing a host state clobbering feature can lead to host state corruption.
The existing denylist, which clears only TSX and WAITPKG, is not fail-safe.
Add tdx_get_cpuid_cfg_mask() to get the mask of directly configurable bits
allowed by KVM for a given CPUID register, covering both feature bits,
which come from tdx_cpu_cfg_caps[], and non-feature bits, which are
enumerated at runtime. Apply the mask to every CPUID entry reported
through KVM_TDX_CAPABILITIES.
With the allowlist in place, newly introduced TDX directly configurable
CPUID bits stay hidden from userspace until KVM explicitly opts in.
Note, filtering KVM_TDX_CAPABILITIES intentionally stops advertising the
directly configurable bits that aren't in the allowlist, as the
corresponding features are unsupported or cannot be properly virtualized.
Update the documentation to reflect the ABI change.
Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
---
v4:
- Rename functions. (Tony)
tdx_cfg_non_feature_mask() -> tdx_get_cpuid_cfg_non_feature_mask()
tdx_cfg_feature_mask() -> tdx_get_cpuid_cfg_feature_mask()
tdx_get_allowed_cfg_cpuid_mask() -> tdx_get_cpuid_cfg_mask()
- Update documentation. (Xiaoyao, Rick)
v3:
- Handle non-feature leafs at runtime. (Sean)
- Add CPUID.0x24.0.EBX[7:0] into allowlist. There is a mismatch of the
description about CPUID.0x24.0.EBX[7:0], which is listed as
"XFAM & CPUID_Enabled & Native" but should be "XFAM & CPUID_Enabled &
Configured & Native".
---
Documentation/virt/kvm/x86/intel-tdx.rst | 8 ++
arch/x86/kvm/vmx/tdx.c | 95 ++++++++++++++++++------
2 files changed, 80 insertions(+), 23 deletions(-)
diff --git a/Documentation/virt/kvm/x86/intel-tdx.rst b/Documentation/virt/kvm/x86/intel-tdx.rst
index 6a222e9d09541..6beeb89d7f057 100644
--- a/Documentation/virt/kvm/x86/intel-tdx.rst
+++ b/Documentation/virt/kvm/x86/intel-tdx.rst
@@ -69,6 +69,14 @@ Return the TDX capabilities that current KVM supports with the specific TDX
module loaded in the system. It reports what features/capabilities are allowed
to be configured to the TDX guest.
+Note, a CPUID feature bit that the TDX module reports as directly configurable
+is not necessarily reported as configurable by KVM. KVM omits features it
+doesn't support. Generally speaking, KVM reports a feature as configurable
+only if KVM supports the feature for both TDX and non-TDX VMs, though there are
+a handful of exceptions where KVM allows a feature for TDX guests that it never
+advertises for non-TDX VMs. Userspace must rely on KVM_TDX_CAPABILITIES to
+determine what can be passed in cpuid to KVM_TDX_INIT_VM.
+
- id: KVM_TDX_CAPABILITIES
- flags: must be 0
- data: pointer to struct kvm_tdx_capabilities
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index b34afc52b714e..ad7d70b36ffe5 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -294,36 +294,79 @@ static bool has_tsx(const struct kvm_cpuid_entry2 *entry)
(entry->ebx & TDX_FEATURE_TSX);
}
-static void clear_tsx(struct kvm_cpuid_entry2 *entry)
-{
- entry->ebx &= ~TDX_FEATURE_TSX;
-}
-
static bool has_waitpkg(const struct kvm_cpuid_entry2 *entry)
{
return entry->function == 7 && entry->index == 0 &&
(entry->ecx & __feature_bit(X86_FEATURE_WAITPKG));
}
-static void clear_waitpkg(struct kvm_cpuid_entry2 *entry)
-{
- entry->ecx &= ~__feature_bit(X86_FEATURE_WAITPKG);
-}
-
-static void tdx_clear_unsupported_cpuid(struct kvm_cpuid_entry2 *entry)
-{
- if (has_tsx(entry))
- clear_tsx(entry);
-
- if (has_waitpkg(entry))
- clear_waitpkg(entry);
-}
-
static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry)
{
return has_tsx(entry) || has_waitpkg(entry);
}
+#define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0)
+
+static u32 tdx_get_cpuid_cfg_non_feature_mask(u32 function, u32 index, int reg)
+{
+ /*
+ * For a leaf/subleaf/register that will never be repurposed to hold
+ * feature bits, it's safe to return TDX_CPUID_ALL_ALLOWED_MASK, i.e.
+ * leave the TDX module's CPUID config mask intact.
+ */
+ switch (function) {
+ case 1:
+ if (reg == CPUID_EAX || reg == CPUID_EBX)
+ return TDX_CPUID_ALL_ALLOWED_MASK;
+ return 0;
+ case 4:
+ case 0x18:
+ case 0x1f:
+ return TDX_CPUID_ALL_ALLOWED_MASK;
+ case 0x24:
+ if (index == 0 && reg == CPUID_EBX)
+ return GENMASK_U32(7, 0);
+ return 0;
+ case 0x80000008:
+ if (reg == CPUID_EAX)
+ return TDX_CPUID_ALL_ALLOWED_MASK;
+ return 0;
+ default:
+ return 0;
+ }
+}
+
+static u32 tdx_get_cpuid_cfg_feature_mask(u32 function, u32 index, int reg)
+{
+ for (int i = 0; i < NR_KVM_CPU_CAPS; i++) {
+ const struct cpuid_reg *cpuid = &reverse_cpuid[i];
+
+ if (!cpuid->function)
+ continue;
+
+ if (cpuid->function == function && cpuid->index == index &&
+ cpuid->reg == reg)
+ return tdx_cpu_cfg_caps[i];
+ }
+
+ return 0;
+}
+
+static u32 tdx_get_cpuid_cfg_mask(u32 function, u32 index, int reg)
+{
+ u32 non_feature_mask = tdx_get_cpuid_cfg_non_feature_mask(function, index, reg);
+
+ /* Skip reverse_cpuid[] walk if it's already all allowed. */
+ if (non_feature_mask == TDX_CPUID_ALL_ALLOWED_MASK)
+ return TDX_CPUID_ALL_ALLOWED_MASK;
+
+ /*
+ * It's possible that a CPUID register contains both feature and
+ * non-feature bits.
+ */
+ return non_feature_mask | tdx_get_cpuid_cfg_feature_mask(function, index, reg);
+}
+
#define KVM_TDX_CPUID_NO_SUBLEAF ((__u32)-1)
static void td_init_cpuid_entry2(struct kvm_cpuid_entry2 *entry, unsigned char idx)
@@ -347,8 +390,6 @@ static void td_init_cpuid_entry2(struct kvm_cpuid_entry2 *entry, unsigned char i
*/
if (entry->function == 0x80000008)
entry->eax = tdx_set_guest_phys_addr_bits(entry->eax, 0xff);
-
- tdx_clear_unsupported_cpuid(entry);
}
#define TDVMCALLINFO_SETUP_EVENT_NOTIFY_INTERRUPT BIT(1)
@@ -371,8 +412,16 @@ static int init_kvm_tdx_caps(const struct tdx_sys_info_td_conf *td_conf,
caps->user_tdvmcallinfo_1_r11 =
TDVMCALLINFO_SETUP_EVENT_NOTIFY_INTERRUPT;
- for (i = 0; i < td_conf->num_cpuid_config; i++)
- td_init_cpuid_entry2(&caps->cpuid.entries[i], i);
+ for (i = 0; i < td_conf->num_cpuid_config; i++) {
+ struct kvm_cpuid_entry2 *e = &caps->cpuid.entries[i];
+
+ td_init_cpuid_entry2(e, i);
+ /* Only report the configurable bits allowed by KVM. */
+ e->eax &= tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_EAX);
+ e->ebx &= tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_EBX);
+ e->ecx &= tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_ECX);
+ e->edx &= tdx_get_cpuid_cfg_mask(e->function, e->index, CPUID_EDX);
+ }
return 0;
}
--
2.46.0
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM
2026-09-17 7:25 [PATCH v4 0/4] KVM: TDX: Validate directly configurable CPUID bits Binbin Wu
` (2 preceding siblings ...)
2026-09-17 7:25 ` [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits Binbin Wu
@ 2026-09-17 7:25 ` Binbin Wu
2026-09-23 0:16 ` Edgecombe, Rick P
3 siblings, 1 reply; 27+ messages in thread
From: Binbin Wu @ 2026-09-17 7:25 UTC (permalink / raw)
To: linux-kernel, kvm
Cc: seanjc, pbonzini, dave.hansen, andrew.cooper3, nik.borisov, kas,
rick.p.edgecombe, xiaoyao.li, chao.gao, tony.lindgren,
kishen.maloor, dedekind1, binbin.wu
Validate the CPUID configuration provided by userspace through
KVM_TDX_INIT_VM against KVM's TDX allowlist, and drop the hardcoded
denylist based check.
The TDX module lets the VMM configure certain CPUID features for a TD at
initialization time, but KVM must strictly govern which of them userspace
can actually enable, otherwise a host state clobbering feature could be
enabled behind KVM's back. The existing check only rejects TSX and
WAITPKG, i.e. it is not fail-safe, as any bit that a future TDX module
makes configurable would be accepted even if KVM has no idea about the
feature.
Add tdx_has_unsupported_cpuid_cfg_bit() and reject KVM_TDX_INIT_VM if
userspace sets any bit outside the mask returned by
tdx_get_cpuid_cfg_mask(). There is no need to first mask the userspace
input with the bits the TDX module reports as directly configurable, as
anything outside that set is rejected by the TDX module itself.
Also reject CPUID entries whose index differs from the value expected by
the TDX module, as kvm_find_cpuid_entry2() ignores the index when
KVM_CPUID_FLAG_SIGNIFCANT_INDEX is cleared, i.e. a mismatching entry could
otherwise be applied to the wrong subleaf.
Update the comments for KVM_TDX_INIT_VM in the uapi header and the TDX
documentation accordingly.
Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
Reviewed-by: Tony Lindgren <tony.lindgren@linux.intel.com>
---
v4:
- Update the comments/documentation.
- Collect RB tag from Tony.
v3:
- Check CPUID entry index mismatch b/t userspace input and TDX sysinfo
configuration. (Sashiko)
- No need to mask the userspace input with TDX module reported directly
configurable bits first, since the userspace input should be subset of
directly configurable bits. Otherwise, the input will be rejected by
the TDX module.
---
Documentation/virt/kvm/x86/intel-tdx.rst | 2 ++
arch/x86/include/uapi/asm/kvm.h | 2 ++
arch/x86/kvm/vmx/tdx.c | 41 ++++++++++++------------
3 files changed, 25 insertions(+), 20 deletions(-)
diff --git a/Documentation/virt/kvm/x86/intel-tdx.rst b/Documentation/virt/kvm/x86/intel-tdx.rst
index 6beeb89d7f057..36bd0fee2480d 100644
--- a/Documentation/virt/kvm/x86/intel-tdx.rst
+++ b/Documentation/virt/kvm/x86/intel-tdx.rst
@@ -135,6 +135,8 @@ KVM_CREATE_VM and before creating any VCPUs.
/*
* Call KVM_TDX_INIT_VM before vcpu creation, thus before
* KVM_SET_CPUID2.
+ * KVM validates @cpuid, i.e. setting a bit that KVM_TDX_CAPABILITIES
+ * doesn't report as configurable fails the ioctl.
* This configuration supersedes KVM_SET_CPUID2s for VCPUs because the
* TDX module directly virtualizes those CPUIDs without VMM. The user
* space VMM, e.g. qemu, should make KVM_SET_CPUID2 consistent with
diff --git a/arch/x86/include/uapi/asm/kvm.h b/arch/x86/include/uapi/asm/kvm.h
index 1585ec8040666..784cb5cd7f664 100644
--- a/arch/x86/include/uapi/asm/kvm.h
+++ b/arch/x86/include/uapi/asm/kvm.h
@@ -1032,6 +1032,8 @@ struct kvm_tdx_init_vm {
/*
* Call KVM_TDX_INIT_VM before vcpu creation, thus before
* KVM_SET_CPUID2.
+ * KVM validates @cpuid, i.e. setting a bit that KVM_TDX_CAPABILITIES
+ * doesn't report as configurable fails the ioctl.
* This configuration supersedes KVM_SET_CPUID2s for VCPUs because the
* TDX module directly virtualizes those CPUIDs without VMM. The user
* space VMM, e.g. qemu, should make KVM_SET_CPUID2 consistent with
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index ad7d70b36ffe5..5426191905d0d 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -286,25 +286,6 @@ static u32 tdx_set_guest_phys_addr_bits(const u32 eax, int addr_bits)
return (eax & ~GENMASK(23, 16)) | (addr_bits & 0xff) << 16;
}
-#define TDX_FEATURE_TSX (__feature_bit(X86_FEATURE_HLE) | __feature_bit(X86_FEATURE_RTM))
-
-static bool has_tsx(const struct kvm_cpuid_entry2 *entry)
-{
- return entry->function == 7 && entry->index == 0 &&
- (entry->ebx & TDX_FEATURE_TSX);
-}
-
-static bool has_waitpkg(const struct kvm_cpuid_entry2 *entry)
-{
- return entry->function == 7 && entry->index == 0 &&
- (entry->ecx & __feature_bit(X86_FEATURE_WAITPKG));
-}
-
-static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry)
-{
- return has_tsx(entry) || has_waitpkg(entry);
-}
-
#define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0)
static u32 tdx_get_cpuid_cfg_non_feature_mask(u32 function, u32 index, int reg)
@@ -2548,6 +2529,17 @@ static int setup_tdparams_eptp_controls(struct kvm_cpuid2 *cpuid,
return 0;
}
+static bool tdx_has_unsupported_cpuid_cfg_bit(const struct kvm_cpuid_entry2 *entry)
+{
+ u32 function = entry->function;
+ u32 index = entry->index;
+
+ return (entry->eax & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_EAX)) ||
+ (entry->ebx & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_EBX)) ||
+ (entry->ecx & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_ECX)) ||
+ (entry->edx & ~tdx_get_cpuid_cfg_mask(function, index, CPUID_EDX));
+}
+
static int setup_tdparams_cpuids(struct kvm_cpuid2 *cpuid,
struct td_params *td_params)
{
@@ -2571,7 +2563,16 @@ static int setup_tdparams_cpuids(struct kvm_cpuid2 *cpuid,
if (!entry)
continue;
- if (tdx_unsupported_cpuid(entry))
+ /*
+ * Reject entries whose index doesn't match the expected one.
+ * This catches userspace passing a CPUID entry with the
+ * KVM_CPUID_FLAG_SIGNIFCANT_INDEX flag cleared when the index
+ * is significant.
+ */
+ if (entry->index != tmp.index)
+ return -EINVAL;
+
+ if (tdx_has_unsupported_cpuid_cfg_bit(entry))
return -EINVAL;
copy_cnt++;
--
2.46.0
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable
2026-09-17 7:25 ` [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Binbin Wu
@ 2026-09-22 21:11 ` Edgecombe, Rick P
2026-09-23 0:03 ` Binbin Wu
2026-09-24 6:44 ` Xiaoyao Li
1 sibling, 1 reply; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-22 21:11 UTC (permalink / raw)
To: kvm, linux-kernel, binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, Maloor, Kishen,
dedekind1, tony.lindgren, pbonzini, andrew.cooper3, nik.borisov
On Thu, 2026-09-17 at 15:25 +0800, Binbin Wu wrote:
> Add CORE_CAPABILITIES (CPUID.0x7.0.EDX[30]) to KVM's allowlist of TDX
> directly configurable CPUID feature bits, even though KVM doesn't support
> MSR_IA32_CORE_CAPS for TDX guests, to accommodate the legacy TDX module
> definition and userspace's stale knowledge of it.
>
> Older TDX specifications define the CORE_CAPABILITIES CPUID bit as
> fixed-1, so userspace may expect the bit to be enabled for TDs. #VE
> reduction turns it into a directly configurable bit, so leaving it out of
> the allowlist would make the bit impossible to enable once KVM starts
> validating userspace's CPUID input, i.e. would be a surprising behavior
> change for such userspace.
Since #VE reduction is controlled from the guest side, this is a bit confusing.
Turning on #VE reduction can't change CORE_CAPABILITIES configurability status
because it's too late. You mean that this was changed from fixed-1 to
configurable to support #VE reduction arch? (I'm not sure why though).
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-17 7:25 ` [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Binbin Wu
@ 2026-09-23 0:01 ` Edgecombe, Rick P
2026-09-23 0:14 ` Binbin Wu
2026-09-24 6:36 ` Xiaoyao Li
1 sibling, 1 reply; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-23 0:01 UTC (permalink / raw)
To: kvm, linux-kernel, binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, Maloor, Kishen,
dedekind1, tony.lindgren, pbonzini, andrew.cooper3, nik.borisov
On Thu, 2026-09-17 at 15:25 +0800, Binbin Wu wrote:
> Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable
> CPUID feature bits that KVM supports, and build the masks during TDX
> hardware setup via tdx_initialize_cpu_cfg_caps().
>
> The TDX module reports the CPUID bits that the VMM can directly configure
> for a TD, but KVM cannot blindly expose all reported bits to userspace.
> Certain features imply additional architectural state, e.g. one or more
> MSRs, that KVM must explicitly manage across host/guest transitions to
> prevent host state corruption. The existing hardcoded denylist cannot
> account for new host state clobbering features introduced by future TDX
> modules.
>
> Except for a few fixed-1 bits required for basic TDX support, host state
> clobbering features are either directly configurable or gated by TD
> ATTRIBUTES/XFAM, which KVM already validates. Tracking only the directly
> configurable bits to build an allowlist is therefore sufficient.
>
> Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can
> be built with similar feature-name based initializers. Directly
> configurable non-feature bits will be handled separately.
>
> The allowlist is prepared to be consumed by later patches to filter
> KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to
> KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID
> feature bits stay hidden from userspace until KVM explicitly opts in.
>
> By default, intersect the allowlist with kvm_cpu_caps[] via TDX_CFG_F().
> Requiring support for non-TDX VMs avoids committing to TDX-specific
> behavior before general KVM support is established, and respects KVM's
> logic around disabling certain features, since the reasons for disabling
> them could apply to TDX as well. Allow exceptions through
> TDX_CFG_EXTRA_F() only with sufficient justification.
>
> Add comments as placeholders for HLE, RTM, WAITPKG and FRED, which KVM
> doesn't support for TDX yet.
>
> Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
> not advertised in kvm_cpu_caps[]. The remaining directly configurable
> feature bits outside kvm_cpu_caps[] are left out of the allowlist:
>
> - Features forced to zero when #VE is reduced,
>
Sorry, I'm not following this logic exactly. The guest can control it's own view
of CPUID. Why do we need to filter the host setting them via direct
configuration, just because the guest can change it's view to exclude them?
> or lacking KVM support
> for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
> RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
> in a subsequent patch.
>
> - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
> TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
>
> - Features that can clobber host state and lack KVM support for TDX:
> FRED.
I think we can't say this quite yet. The arch isn't finalized.
>
> - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
> TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
> platform), and RAO_INT (defined only for future processors).
>
> Filtering KVM_TDX_CAPABILITIES in a subsequent patch will intentionally
Nit: "patch" -> "change"
This is drilled into my head working on the tip side. I'm not sure if Sean has
the same allergy though.
> stop advertising the excluded bits as configurable, as the corresponding
> features are unsupported or cannot be properly virtualized.
>
> Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
Overall it looks very good to me.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable
2026-09-22 21:11 ` Edgecombe, Rick P
@ 2026-09-23 0:03 ` Binbin Wu
0 siblings, 0 replies; 27+ messages in thread
From: Binbin Wu @ 2026-09-23 0:03 UTC (permalink / raw)
To: Edgecombe, Rick P
Cc: kvm, linux-kernel, Gao, Chao, seanjc, dave.hansen, kas, Li,
Xiaoyao, Maloor, Kishen, dedekind1, tony.lindgren, pbonzini,
andrew.cooper3, nik.borisov
On 9/23/2026 5:11 AM, Edgecombe, Rick P wrote:
> On Thu, 2026-09-17 at 15:25 +0800, Binbin Wu wrote:
>> Add CORE_CAPABILITIES (CPUID.0x7.0.EDX[30]) to KVM's allowlist of TDX
>> directly configurable CPUID feature bits, even though KVM doesn't support
>> MSR_IA32_CORE_CAPS for TDX guests, to accommodate the legacy TDX module
>> definition and userspace's stale knowledge of it.
>>
>> Older TDX specifications define the CORE_CAPABILITIES CPUID bit as
>> fixed-1, so userspace may expect the bit to be enabled for TDs. #VE
>> reduction turns it into a directly configurable bit, so leaving it out of
>> the allowlist would make the bit impossible to enable once KVM starts
>> validating userspace's CPUID input, i.e. would be a surprising behavior
>> change for such userspace.
>
> Since #VE reduction is controlled from the guest side, this is a bit confusing.
> Turning on #VE reduction can't change CORE_CAPABILITIES configurability status
> because it's too late. You mean that this was changed from fixed-1 to
> configurable to support #VE reduction arch? (I'm not sure why though).
Yes, a fixed-1 bit was changed to directly configurable bit if it is a
#VE reduction bit, regardless the guest enables #VE reduction or not.
I.e. CORE_CAPABILITIES was fixed-1 bit, but now it's directly configurable
by the definition of the specs.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 0:01 ` Edgecombe, Rick P
@ 2026-09-23 0:14 ` Binbin Wu
2026-09-23 0:45 ` Edgecombe, Rick P
0 siblings, 1 reply; 27+ messages in thread
From: Binbin Wu @ 2026-09-23 0:14 UTC (permalink / raw)
To: Edgecombe, Rick P
Cc: kvm, linux-kernel, Gao, Chao, seanjc, dave.hansen, kas, Li,
Xiaoyao, Maloor, Kishen, dedekind1, tony.lindgren, pbonzini,
andrew.cooper3, nik.borisov
On 9/23/2026 8:01 AM, Edgecombe, Rick P wrote:
> On Thu, 2026-09-17 at 15:25 +0800, Binbin Wu wrote:
>>
>> Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
>> not advertised in kvm_cpu_caps[]. The remaining directly configurable
>> feature bits outside kvm_cpu_caps[] are left out of the allowlist:
>>
>> - Features forced to zero when #VE is reduced,
>>
>
> Sorry, I'm not following this logic exactly. The guest can control it's own view
> of CPUID. Why do we need to filter the host setting them via direct
> configuration, just because the guest can change it's view to exclude them?
These feature neither supported by the TDX module nor supported by the KVM.
If the guest enabled the #VE reduction, the access to the related MSRs will
cause #GP. If the guest doesn't enabled the #VE redcution, #VE cannot be
supported by KVM since the related MSRs are not support by KVM.
So the guest cannot use these features anyway.
>
>> or lacking KVM support
>> for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
>> RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
>> in a subsequent patch.
>>
>> - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
>> TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
>>
>> - Features that can clobber host state and lack KVM support for TDX:
>> FRED.
>
> I think we can't say this quite yet. The arch isn't finalized
Oh, right. I should be more rigorous.
>
>>
>> - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
>> TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
>> platform), and RAO_INT (defined only for future processors).
>>
>> Filtering KVM_TDX_CAPABILITIES in a subsequent patch will intentionally
>
> Nit: "patch" -> "change"
>
> This is drilled into my head working on the tip side. I'm not sure if Sean has
> the same allergy though.
Will update it.
>
>> stop advertising the excluded bits as configurable, as the corresponding
>> features are unsupported or cannot be properly virtualized.
>>
>> Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
>
> Overall it looks very good to me.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits
2026-09-17 7:25 ` [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits Binbin Wu
@ 2026-09-23 0:16 ` Edgecombe, Rick P
2026-09-23 0:28 ` Binbin Wu
0 siblings, 1 reply; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-23 0:16 UTC (permalink / raw)
To: kvm, linux-kernel, binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, Maloor, Kishen,
dedekind1, tony.lindgren, pbonzini, andrew.cooper3, nik.borisov
On Thu, 2026-09-17 at 15:25 +0800, Binbin Wu wrote:
> Filter the directly configurable CPUID bits reported through
> KVM_TDX_CAPABILITIES against KVM's TDX allowlist, and drop the hardcoded
> denylist based filtering.
>
> The TDX module reports directly configurable CPUID bits that it supports
> for a TD, but KVM must not expose bits that it doesn't support, as blindly
> exposing a host state clobbering feature can lead to host state corruption.
> The existing denylist, which clears only TSX and WAITPKG, is not fail-safe.
>
> Add tdx_get_cpuid_cfg_mask() to get the mask of directly configurable bits
> allowed by KVM for a given CPUID register, covering both feature bits,
> which come from tdx_cpu_cfg_caps[], and non-feature bits, which are
> enumerated at runtime. Apply the mask to every CPUID entry reported
> through KVM_TDX_CAPABILITIES.
>
> With the allowlist in place, newly introduced TDX directly configurable
> CPUID bits stay hidden from userspace until KVM explicitly opts in.
>
> Note, filtering KVM_TDX_CAPABILITIES intentionally stops advertising the
> directly configurable bits that aren't in the allowlist, as the
> corresponding features are unsupported or cannot be properly virtualized.
I think we need to more loudly call out that this is an ABI change and we expect
it not to impact other VMMs because...
>
> Update the documentation to reflect the ABI change.
>
> Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
> ---
> v4:
> - Rename functions. (Tony)
> tdx_cfg_non_feature_mask() -> tdx_get_cpuid_cfg_non_feature_mask()
> tdx_cfg_feature_mask() -> tdx_get_cpuid_cfg_feature_mask()
> tdx_get_allowed_cfg_cpuid_mask() -> tdx_get_cpuid_cfg_mask()
> - Update documentation. (Xiaoyao, Rick)
>
> v3:
> - Handle non-feature leafs at runtime. (Sean)
> - Add CPUID.0x24.0.EBX[7:0] into allowlist. There is a mismatch of the
> description about CPUID.0x24.0.EBX[7:0], which is listed as
> "XFAM & CPUID_Enabled & Native" but should be "XFAM & CPUID_Enabled &
> Configured & Native".
> ---
> Documentation/virt/kvm/x86/intel-tdx.rst | 8 ++
> arch/x86/kvm/vmx/tdx.c | 95 ++++++++++++++++++------
> 2 files changed, 80 insertions(+), 23 deletions(-)
>
> diff --git a/Documentation/virt/kvm/x86/intel-tdx.rst b/Documentation/virt/kvm/x86/intel-tdx.rst
> index 6a222e9d09541..6beeb89d7f057 100644
> --- a/Documentation/virt/kvm/x86/intel-tdx.rst
> +++ b/Documentation/virt/kvm/x86/intel-tdx.rst
> @@ -69,6 +69,14 @@ Return the TDX capabilities that current KVM supports with the specific TDX
> module loaded in the system. It reports what features/capabilities are allowed
> to be configured to the TDX guest.
>
> +Note, a CPUID feature bit that the TDX module reports as directly configurable
> +is not necessarily reported as configurable by KVM. KVM omits features it
> +doesn't support. Generally speaking, KVM reports a feature as configurable
> +only if KVM supports the feature for both TDX and non-TDX VMs, though there are
> +a handful of exceptions where KVM allows a feature for TDX guests that it never
> +advertises for non-TDX VMs. Userspace must rely on KVM_TDX_CAPABILITIES to
> +determine what can be passed in cpuid to KVM_TDX_INIT_VM.
> +
> - id: KVM_TDX_CAPABILITIES
> - flags: must be 0
> - data: pointer to struct kvm_tdx_capabilities
> diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
> index b34afc52b714e..ad7d70b36ffe5 100644
> --- a/arch/x86/kvm/vmx/tdx.c
> +++ b/arch/x86/kvm/vmx/tdx.c
> @@ -294,36 +294,79 @@ static bool has_tsx(const struct kvm_cpuid_entry2 *entry)
> (entry->ebx & TDX_FEATURE_TSX);
> }
>
> -static void clear_tsx(struct kvm_cpuid_entry2 *entry)
> -{
> - entry->ebx &= ~TDX_FEATURE_TSX;
> -}
> -
> static bool has_waitpkg(const struct kvm_cpuid_entry2 *entry)
> {
> return entry->function == 7 && entry->index == 0 &&
> (entry->ecx & __feature_bit(X86_FEATURE_WAITPKG));
> }
>
> -static void clear_waitpkg(struct kvm_cpuid_entry2 *entry)
> -{
> - entry->ecx &= ~__feature_bit(X86_FEATURE_WAITPKG);
> -}
> -
> -static void tdx_clear_unsupported_cpuid(struct kvm_cpuid_entry2 *entry)
> -{
> - if (has_tsx(entry))
> - clear_tsx(entry);
> -
> - if (has_waitpkg(entry))
> - clear_waitpkg(entry);
> -}
> -
> static bool tdx_unsupported_cpuid(const struct kvm_cpuid_entry2 *entry)
> {
> return has_tsx(entry) || has_waitpkg(entry);
> }
>
> +#define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0)
> +
> +static u32 tdx_get_cpuid_cfg_non_feature_mask(u32 function, u32 index, int reg)
> +{
> + /*
> + * For a leaf/subleaf/register that will never be repurposed to hold
> + * feature bits, it's safe to return TDX_CPUID_ALL_ALLOWED_MASK, i.e.
> + * leave the TDX module's CPUID config mask intact.
> + */
> + switch (function) {
> + case 1:
> + if (reg == CPUID_EAX || reg == CPUID_EBX)
> + return TDX_CPUID_ALL_ALLOWED_MASK;
> + return 0;
> + case 4:
> + case 0x18:
> + case 0x1f:
> + return TDX_CPUID_ALL_ALLOWED_MASK;
> + case 0x24:
> + if (index == 0 && reg == CPUID_EBX)
> + return GENMASK_U32(7, 0);
> + return 0;
> + case 0x80000008:
> + if (reg == CPUID_EAX)
> + return TDX_CPUID_ALL_ALLOWED_MASK;
There are reserved bits in some of these. Are these "no feature bit" assumptions
checked somehow?
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM
2026-09-17 7:25 ` [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM Binbin Wu
@ 2026-09-23 0:16 ` Edgecombe, Rick P
0 siblings, 0 replies; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-23 0:16 UTC (permalink / raw)
To: kvm, linux-kernel, binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, Maloor, Kishen,
dedekind1, tony.lindgren, pbonzini, andrew.cooper3, nik.borisov
On Thu, 2026-09-17 at 15:25 +0800, Binbin Wu wrote:
> Validate the CPUID configuration provided by userspace through
> KVM_TDX_INIT_VM against KVM's TDX allowlist, and drop the hardcoded
> denylist based check.
>
> The TDX module lets the VMM configure certain CPUID features for a TD at
> initialization time, but KVM must strictly govern which of them userspace
> can actually enable, otherwise a host state clobbering feature could be
> enabled behind KVM's back. The existing check only rejects TSX and
> WAITPKG, i.e. it is not fail-safe, as any bit that a future TDX module
> makes configurable would be accepted even if KVM has no idea about the
> feature.
>
> Add tdx_has_unsupported_cpuid_cfg_bit() and reject KVM_TDX_INIT_VM if
> userspace sets any bit outside the mask returned by
> tdx_get_cpuid_cfg_mask(). There is no need to first mask the userspace
> input with the bits the TDX module reports as directly configurable, as
> anything outside that set is rejected by the TDX module itself.
>
> Also reject CPUID entries whose index differs from the value expected by
> the TDX module, as kvm_find_cpuid_entry2() ignores the index when
> KVM_CPUID_FLAG_SIGNIFCANT_INDEX is cleared, i.e. a mismatching entry could
> otherwise be applied to the wrong subleaf.
>
> Update the comments for KVM_TDX_INIT_VM in the uapi header and the TDX
> documentation accordingly.
Same comment on this one as patch 3, this removes some previously configurable
features right? And you checked qemu and reasoned on the rest that this won't
break anything. We don't need the full bit by bit detail here, but if someone
bisects this patch it would be good to have some info about what the ABI change
was.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits
2026-09-23 0:16 ` Edgecombe, Rick P
@ 2026-09-23 0:28 ` Binbin Wu
2026-09-23 0:34 ` Edgecombe, Rick P
0 siblings, 1 reply; 27+ messages in thread
From: Binbin Wu @ 2026-09-23 0:28 UTC (permalink / raw)
To: Edgecombe, Rick P
Cc: kvm, linux-kernel, Gao, Chao, seanjc, dave.hansen, kas, Li,
Xiaoyao, Maloor, Kishen, dedekind1, tony.lindgren, pbonzini,
andrew.cooper3, nik.borisov
On 9/23/2026 8:16 AM, Edgecombe, Rick P wrote:
> On Thu, 2026-09-17 at 15:25 +0800, Binbin Wu wrote:
>> Filter the directly configurable CPUID bits reported through
>> KVM_TDX_CAPABILITIES against KVM's TDX allowlist, and drop the hardcoded
>> denylist based filtering.
>>
>> The TDX module reports directly configurable CPUID bits that it supports
>> for a TD, but KVM must not expose bits that it doesn't support, as blindly
>> exposing a host state clobbering feature can lead to host state corruption.
>> The existing denylist, which clears only TSX and WAITPKG, is not fail-safe.
>>
>> Add tdx_get_cpuid_cfg_mask() to get the mask of directly configurable bits
>> allowed by KVM for a given CPUID register, covering both feature bits,
>> which come from tdx_cpu_cfg_caps[], and non-feature bits, which are
>> enumerated at runtime. Apply the mask to every CPUID entry reported
>> through KVM_TDX_CAPABILITIES.
>>
>> With the allowlist in place, newly introduced TDX directly configurable
>> CPUID bits stay hidden from userspace until KVM explicitly opts in.
>>
>> Note, filtering KVM_TDX_CAPABILITIES intentionally stops advertising the
>> directly configurable bits that aren't in the allowlist, as the
>> corresponding features are unsupported or cannot be properly virtualized.
>
> I think we need to more loudly call out that this is an ABI change and we expect
> it not to impact other VMMs because...
OK, will do this for this patch and patch 4.
Thanks!
>
>>
>> Update the documentation to reflect the ABI change.
>>
[...]
>> +#define TDX_CPUID_ALL_ALLOWED_MASK GENMASK_U32(31, 0)
>> +
>> +static u32 tdx_get_cpuid_cfg_non_feature_mask(u32 function, u32 index, int reg)
>> +{
>> + /*
>> + * For a leaf/subleaf/register that will never be repurposed to hold
>> + * feature bits, it's safe to return TDX_CPUID_ALL_ALLOWED_MASK, i.e.
>> + * leave the TDX module's CPUID config mask intact.
>> + */
>> + switch (function) {
>> + case 1:
>> + if (reg == CPUID_EAX || reg == CPUID_EBX)
>> + return TDX_CPUID_ALL_ALLOWED_MASK;
>> + return 0;
>> + case 4:
>> + case 0x18:
>> + case 0x1f:
>> + return TDX_CPUID_ALL_ALLOWED_MASK;
>> + case 0x24:
>> + if (index == 0 && reg == CPUID_EBX)
>> + return GENMASK_U32(7, 0);
>> + return 0;
>> + case 0x80000008:
>> + if (reg == CPUID_EAX)
>> + return TDX_CPUID_ALL_ALLOWED_MASK;
>
> There are reserved bits in some of these. Are these "no feature bit" assumptions
> checked somehow?
Xiaoyao mentioned it in v3 too.
Sean suggested in v2 that "Realistically, CPUID.0x1.E{A,B}X are never going to be
repurposed to hold feature bits, and so generating a mask of allowed bits adds
unnecessary cognitive load and maintenance. Ditto for CPUID 0x4, 0x18, and 0x1F."
https://lore.kernel.org/kvm/aj1fi_0SBxMK5WOB@google.com/
If we really have concerns, I can change it to exclude reserved bits in the next
version.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits
2026-09-23 0:28 ` Binbin Wu
@ 2026-09-23 0:34 ` Edgecombe, Rick P
0 siblings, 0 replies; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-23 0:34 UTC (permalink / raw)
To: binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, linux-kernel,
Maloor, Kishen, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On Wed, 2026-09-23 at 08:28 +0800, Binbin Wu wrote:
> Xiaoyao mentioned it in v3 too.
>
> Sean suggested in v2 that "Realistically, CPUID.0x1.E{A,B}X are never going to be
> repurposed to hold feature bits, and so generating a mask of allowed bits adds
> unnecessary cognitive load and maintenance. Ditto for CPUID 0x4, 0x18, and 0x1F."
> https://lore.kernel.org/kvm/aj1fi_0SBxMK5WOB@google.com/
>
> If we really have concerns, I can change it to exclude reserved bits in the next
> version.
Could we check the x86 arch folks real quick? If Sean is ok betting on this
though, I'm ok too.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 0:14 ` Binbin Wu
@ 2026-09-23 0:45 ` Edgecombe, Rick P
2026-09-23 0:57 ` Binbin Wu
0 siblings, 1 reply; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-23 0:45 UTC (permalink / raw)
To: binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, linux-kernel,
Maloor, Kishen, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On Wed, 2026-09-23 at 08:14 +0800, Binbin Wu wrote:
> > Sorry, I'm not following this logic exactly. The guest can control it's own
> > view
> > of CPUID. Why do we need to filter the host setting them via direct
> > configuration, just because the guest can change it's view to exclude them?
>
>
> These feature neither supported by the TDX module nor supported by the KVM.
> If the guest enabled the #VE reduction, the access to the related MSRs will
> cause #GP. If the guest doesn't enabled the #VE redcution, #VE cannot be
> supported by KVM since the related MSRs are not support by KVM.
>
> So the guest cannot use these features anyway.
But directly configurable bits don't have any direct conntion to tdcall queried
bits. Userspace can still set those to whatever it wants regardless of the allow
list.
So the argument is basically there is not expected to be any reason to allow
them, so save the code in the allow list. It's a "there is no point" reason.
Makes sense, but I couldn't get that from explanation.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 0:45 ` Edgecombe, Rick P
@ 2026-09-23 0:57 ` Binbin Wu
2026-09-23 1:09 ` Edgecombe, Rick P
0 siblings, 1 reply; 27+ messages in thread
From: Binbin Wu @ 2026-09-23 0:57 UTC (permalink / raw)
To: Edgecombe, Rick P
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, linux-kernel,
Maloor, Kishen, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On 9/23/2026 8:45 AM, Edgecombe, Rick P wrote:
> On Wed, 2026-09-23 at 08:14 +0800, Binbin Wu wrote:
>>> Sorry, I'm not following this logic exactly. The guest can control it's own
>>> view
>>> of CPUID. Why do we need to filter the host setting them via direct
>>> configuration, just because the guest can change it's view to exclude them?
>>
>>
>> These feature neither supported by the TDX module nor supported by the KVM.
>> If the guest enabled the #VE reduction, the access to the related MSRs will
>> cause #GP. If the guest doesn't enabled the #VE redcution, #VE cannot be
>> supported by KVM since the related MSRs are not support by KVM.
>>
>> So the guest cannot use these features anyway.
>
> But directly configurable bits don't have any direct conntion to tdcall queried
> bits. Userspace can still set those to whatever it wants regardless of the allow
> list.
I didn't quite get this.
>
> So the argument is basically there is not expected to be any reason to allow
> them, so save the code in the allow list. It's a "there is no point" reason.
> Makes sense, but I couldn't get that from explanation.
I think "not just save the code"?
One argument is that if these bits are reported as allowed to userspace and userspace
enabled them during init, and the guest doesn't opt-in the #VE reduction (KVM doesn't
know the real setting), it would cause problem in the guest.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 0:57 ` Binbin Wu
@ 2026-09-23 1:09 ` Edgecombe, Rick P
2026-09-23 1:17 ` Binbin Wu
0 siblings, 1 reply; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-23 1:09 UTC (permalink / raw)
To: binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, Maloor, Kishen,
linux-kernel, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On Wed, 2026-09-23 at 08:57 +0800, Binbin Wu wrote:
> > But directly configurable bits don't have any direct conntion to tdcall
> > queried bits. Userspace can still set those to whatever it wants regardless
> > of the allow list.
>
> I didn't quite get this.
VE bits don't query the directly configurable bits in any case unless userspace
sets it up that way. So they are not involved in the problem we are trying to
fix here.
>
> >
> > So the argument is basically there is not expected to be any reason to allow
> > them, so save the code in the allow list. It's a "there is no point" reason.
> > Makes sense, but I couldn't get that from explanation.
>
> I think "not just save the code"?
> One argument is that if these bits are reported as allowed to userspace and
> userspace enabled them during init, and the guest doesn't opt-in the #VE
> reduction (KVM doesn't know the real setting), it would cause problem in the
> guest.
Userspace is allowed to cause problems to the guest. We shouldn't try to prevent
it. But here, I think we are saving lines of allowlist in preventing it. So the
reason is to have less allow list lines, not to help userspace. And that is
worth it IMO.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 1:09 ` Edgecombe, Rick P
@ 2026-09-23 1:17 ` Binbin Wu
2026-09-23 11:21 ` Xiaoyao Li
0 siblings, 1 reply; 27+ messages in thread
From: Binbin Wu @ 2026-09-23 1:17 UTC (permalink / raw)
To: Edgecombe, Rick P
Cc: Gao, Chao, seanjc, dave.hansen, kas, Li, Xiaoyao, Maloor, Kishen,
linux-kernel, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On 9/23/2026 9:09 AM, Edgecombe, Rick P wrote:
> On Wed, 2026-09-23 at 08:57 +0800, Binbin Wu wrote:
>>> But directly configurable bits don't have any direct conntion to tdcall
>>> queried bits. Userspace can still set those to whatever it wants regardless
>>> of the allow list.
>>
>> I didn't quite get this.
>
> VE bits don't query the directly configurable bits in any case unless userspace
> sets it up that way. So they are not involved in the problem we are trying to
> fix here.
>
>>
>>>
>>> So the argument is basically there is not expected to be any reason to allow
>>> them, so save the code in the allow list. It's a "there is no point" reason.
>>> Makes sense, but I couldn't get that from explanation.
>>
>> I think "not just save the code"?
>> One argument is that if these bits are reported as allowed to userspace and
>> userspace enabled them during init, and the guest doesn't opt-in the #VE
>> reduction (KVM doesn't know the real setting), it would cause problem in the
>> guest.
>
> Userspace is allowed to cause problems to the guest. We shouldn't try to prevent
> it.
But userspace doesn't know these bits are not supported by TDX module and KVM.
If KVM report them as supported, and it causes problems, we cannot say it's
userspace's fault.
> But here, I think we are saving lines of allowlist in preventing it. So the
> reason is to have less allow list lines, not to help userspace. And that is
> worth it IMO.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 1:17 ` Binbin Wu
@ 2026-09-23 11:21 ` Xiaoyao Li
2026-09-23 14:25 ` Edgecombe, Rick P
0 siblings, 1 reply; 27+ messages in thread
From: Xiaoyao Li @ 2026-09-23 11:21 UTC (permalink / raw)
To: Binbin Wu, Edgecombe, Rick P
Cc: Gao, Chao, seanjc, dave.hansen, kas, Maloor, Kishen,
linux-kernel, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On 9/23/2026 9:17 AM, Binbin Wu wrote:
> On 9/23/2026 9:09 AM, Edgecombe, Rick P wrote:
>> On Wed, 2026-09-23 at 08:57 +0800, Binbin Wu wrote:
>>>> But directly configurable bits don't have any direct conntion to tdcall
>>>> queried bits. Userspace can still set those to whatever it wants regardless
>>>> of the allow list.
>>>
>>> I didn't quite get this.
>>
>> VE bits don't query the directly configurable bits in any case unless userspace
>> sets it up that way. So they are not involved in the problem we are trying to
>> fix here.
>>
>>>
>>>>
>>>> So the argument is basically there is not expected to be any reason to allow
>>>> them, so save the code in the allow list. It's a "there is no point" reason.
>>>> Makes sense, but I couldn't get that from explanation.
>>>
>>> I think "not just save the code"?
>>> One argument is that if these bits are reported as allowed to userspace and
>>> userspace enabled them during init, and the guest doesn't opt-in the #VE
>>> reduction (KVM doesn't know the real setting), it would cause problem in the
>>> guest.
>>
>> Userspace is allowed to cause problems to the guest. We shouldn't try to prevent
>> it.
>
> But userspace doesn't know these bits are not supported by TDX module and KVM.
> If KVM report them as supported, and it causes problems, we cannot say it's
> userspace's fault.
It sounds like a bug fix to existing KVM behavior.
Current KVM doesn't allow HLE/RTM/WAITPKG because it might cause bad effect on
the host. But for those #VE related, though KVM doesn't support virtualizing the
features, it doesn't cause any harm to the host when KVM allows them to be
configured to TDs. It only causes problems to TDs.
For normal VMs, KVM_GET_SUPPORTED_CPUID and other KVM CAPs serve as the
interface to report to userspace if a feature is support or not. If userspace
exposes a feature to guest when KVM reports the feature is not supported, and it
causes problem to guests, we can say it's userspace's fault. But for TDX, there
is not interface to report KVM's support capabilities. After this series,
KVM_TDX_CAPABILITIES can serve as the interface. So it looks like a bug fix to me.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 11:21 ` Xiaoyao Li
@ 2026-09-23 14:25 ` Edgecombe, Rick P
2026-09-24 2:05 ` Xiaoyao Li
0 siblings, 1 reply; 27+ messages in thread
From: Edgecombe, Rick P @ 2026-09-23 14:25 UTC (permalink / raw)
To: Li, Xiaoyao, binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Maloor, Kishen,
linux-kernel, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On Wed, 2026-09-23 at 19:21 +0800, Xiaoyao Li wrote:
> > But userspace doesn't know these bits are not supported by TDX module and
> > KVM. If KVM report them as supported, and it causes problems, we cannot say
> > it's userspace's fault.
>
> It sounds like a bug fix to existing KVM behavior.
>
> Current KVM doesn't allow HLE/RTM/WAITPKG because it might cause bad effect on
> the host. But for those #VE related, though KVM doesn't support virtualizing
> the features, it doesn't cause any harm to the host when KVM allows them to be
> configured to TDs. It only causes problems to TDs.
>
> For normal VMs, KVM_GET_SUPPORTED_CPUID and other KVM CAPs serve as the
> interface to report to userspace if a feature is support or not. If userspace
> exposes a feature to guest when KVM reports the feature is not supported, and
> it causes problem to guests, we can say it's userspace's fault.
Yes.
> But for TDX,
> there is not interface to report KVM's support capabilities. After this
> series, KVM_TDX_CAPABILITIES can serve as the interface. So it looks like a
> bug fix to me.
KVM needs to support all the bits that are 1 in the TD. Either directly
configurable or fixed-1 or whatever. Userspace doesn't find out which bits these
are until it does the KVM_TDX_GET_CPUID, at which point it finds out the
"support" in a heavy handed way. So I think userspace does learn everything it
needs already. And as long we don't have new fixed-1 bits, we will avoid KVM
being forced to support things too early.
Except for exposing TD KVM PV bit support, which we already talked about doing
as follow on.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-23 14:25 ` Edgecombe, Rick P
@ 2026-09-24 2:05 ` Xiaoyao Li
0 siblings, 0 replies; 27+ messages in thread
From: Xiaoyao Li @ 2026-09-24 2:05 UTC (permalink / raw)
To: Edgecombe, Rick P, binbin.wu
Cc: Gao, Chao, seanjc, dave.hansen, kas, Maloor, Kishen,
linux-kernel, tony.lindgren, kvm, pbonzini, nik.borisov,
dedekind1, andrew.cooper3
On 9/23/2026 10:25 PM, Edgecombe, Rick P wrote:
>> But for TDX,
>> there is not interface to report KVM's support capabilities. After this
>> series, KVM_TDX_CAPABILITIES can serve as the interface. So it looks like a
>> bug fix to me.
> KVM needs to support all the bits that are 1 in the TD. Either directly
> configurable or fixed-1 or whatever. Userspace doesn't find out which bits these
> are until it does the KVM_TDX_GET_CPUID, at which point it finds out the
> "support" in a heavy handed way. So I think userspace does learn everything it
> needs already. And as long we don't have new fixed-1 bits, we will avoid KVM
> being forced to support things too early.
Do you mean after this series of before this series?
Before this series, i.e. with current KVM, KVM_TDX_GET_CPUID can tell userspace
the "support" bit for a given TD, but some of the "support" bits might not be
supported by KVM. For example, current userspace can set RDT_A and RDT_M, and
they will be reported as 1 on KVM_TDX_GET_CPUID. However, KVM doesn't support them.
After this series, KVM will not report such bits like RDT_A and RDT_M any more,
and userspace will no longer be able to set them for a TD neither.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-17 7:25 ` [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Binbin Wu
2026-09-23 0:01 ` Edgecombe, Rick P
@ 2026-09-24 6:36 ` Xiaoyao Li
2026-09-24 7:49 ` Binbin Wu
1 sibling, 1 reply; 27+ messages in thread
From: Xiaoyao Li @ 2026-09-24 6:36 UTC (permalink / raw)
To: Binbin Wu, linux-kernel, kvm
Cc: seanjc, pbonzini, dave.hansen, andrew.cooper3, nik.borisov, kas,
rick.p.edgecombe, chao.gao, tony.lindgren, kishen.maloor,
dedekind1
On 9/17/2026 3:25 PM, Binbin Wu wrote:
> Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable
> CPUID feature bits that KVM supports, and build the masks during TDX
> hardware setup via tdx_initialize_cpu_cfg_caps().
>
> The TDX module reports the CPUID bits that the VMM can directly configure
> for a TD, but KVM cannot blindly expose all reported bits to userspace.
> Certain features imply additional architectural state, e.g. one or more
> MSRs, that KVM must explicitly manage across host/guest transitions to
> prevent host state corruption. The existing hardcoded denylist cannot
> account for new host state clobbering features introduced by future TDX
> modules.
>
> Except for a few fixed-1 bits required for basic TDX support, host state
> clobbering features are either directly configurable or gated by TD
> ATTRIBUTES/XFAM, which KVM already validates. Tracking only the directly
> configurable bits to build an allowlist is therefore sufficient.
>
> Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can
> be built with similar feature-name based initializers. Directly
> configurable non-feature bits will be handled separately.
>
> The allowlist is prepared to be consumed by later patches to filter
> KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to
> KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID
> feature bits stay hidden from userspace until KVM explicitly opts in.
>
> By default, intersect the allowlist with kvm_cpu_caps[] via TDX_CFG_F().
> Requiring support for non-TDX VMs avoids committing to TDX-specific
> behavior before general KVM support is established, and respects KVM's
> logic around disabling certain features, since the reasons for disabling
> them could apply to TDX as well. Allow exceptions through
> TDX_CFG_EXTRA_F() only with sufficient justification.
One name just crosses my mind. How about TDX_CFG_INDEP_F()? where INDEP stands
for independent.
- TDX_CFG_F() for features that are dependent on kvm_cpu_caps[]
- TDX_CFG_INDEP_F() for features not dependent on kvm_cpu_caps[]
> Add comments as placeholders for HLE, RTM, WAITPKG and FRED, which KVM
> doesn't support for TDX yet.
1. why add comments for them? Because implementation wise, they are possible to
be supported by KVM but just current KVM doesn't have code to support them yet?
If so, are the others the features that are concluded to be impossible to be
supported by KVM?
2. We don't need to talk about FRED, which is even not suppported for normal VMs
yet. Unless we are trying to enable FRED for TDs specifically.
> Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
> not advertised in kvm_cpu_caps[]. The remaining directly configurable
> feature bits outside kvm_cpu_caps[] are left out of the allowlist:
>
> - Features forced to zero when #VE is reduced, or lacking KVM support
> for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
> RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
> in a subsequent patch.
why not split into two category? Because they are overlapped?
- Features forced to zero when #VE is reduced
- Features lack KVM support
As for "features forced to zero when #VE is reduced", do you mean when TDX
module supports "#VE reduction" feature, the features becomes fixed0? or when
guest enables "reduce #VE", the guest see a 0 value even though the feature is
still configurable to host userspace and host userspace configure it to 1?
> - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
> TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
What is CID?
and Is PBE related to any bit of IA32_MISC_ENABLE? Did you mean PEBS instead?
I don't understand what this category. It talks about something when
TDCS.TD_CTLS.REDUCE_VE is set, which depends on guest TD's behavior. So what if
the guest TD doesn't set TDCS.TD_CTLS.REDUCE_VE?
> - Features that can clobber host state and lack KVM support for TDX:
> FRED.
>
> - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
> TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
> platform), and RAO_INT (defined only for future processors).
PREFETCHWT1 (if the name is correct) is not supported by the kernel right? For
such unknown features, it's straightforward to not support it.
AMX-TRANSPOSE has been turned into reserved, and kernel never supports it. I
think we don't need to mention it at all.
RAO_INT is fixed 0 for TDX 1.5?
At last, I think we can just classify all of the above as 3 types:
- Features that are directly configurable by TDX and in kvm_cpu_caps[].
use TDX_CFG_F()
- Features that are directly configurable by TDX and KVM actually supports them
or KVM can support them but not in kvm_cpu_caps[]. Use TDX_CFG_EXTRA_F().
For example, MWAIT, XTPR, and HT.
- Features that are directly configurable by TDX but KVM or kenrel doesn't
support them.
For example, RDT_A, RDT_M, PREFETCHWT1, PSN,
> Filtering KVM_TDX_CAPABILITIES in a subsequent patch will intentionally
> stop advertising the excluded bits as configurable, as the corresponding
> features are unsupported or cannot be properly virtualized.
>
> Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
> ---
> v4:
> - Add XTPR and HT to the allowlist.
> - Explain in the changelog why some directly configurable bits are not
> added to the allowlist.
> - Improve comments/changelog. (Xiaoyao)
> - Fix the bug that AMX_COMPLEX should be added for CPUID_1E_1_EAX instead
> of CPUID_7_1_EDX.
>
> v3:
> - Drop the new data structure in v2 and only track feature bits by
> following the organization of kvm_cpu_caps[], handle non-feature
> bits separately. (Sean)
> - Use two versions of macros (TDX_CFG_F() vs. TDX_CFG_EXTRA_F()) to
> distinguish whether a supported TDX configurable CPUID bit should be
> checked against KVM's common cpu capabilities.
> - Add AMX_COMPLEX since it has been defined in the CPUID virtualization doc.
> ---
> arch/x86/kvm/vmx/tdx.c | 164 +++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 164 insertions(+)
>
> diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
> index 8a3f88b79b83d..2b51a85c998e8 100644
> --- a/arch/x86/kvm/vmx/tdx.c
> +++ b/arch/x86/kvm/vmx/tdx.c
> @@ -52,6 +52,168 @@
> __TDX_BUG_ON(__err, #__fn, __kvm, ", " #a1 " 0x%llx, " #a2 ", 0x%llx, " #a3 " 0x%llx", \
> a1, a2, a3)
>
> +static u32 tdx_cpu_cfg_caps[NR_KVM_CPU_CAPS] __ro_after_init;
> +static_assert(ARRAY_SIZE(tdx_cpu_cfg_caps) == ARRAY_SIZE(kvm_cpu_caps));
> +
> +#define TDX_VALIDATE_CPU_CAP_USAGE(name) \
> + BUILD_BUG_ON(__feature_leaf(X86_FEATURE_##name) != \
> + tdx_cpu_cap_init_in_progress)
> +
> +/* For a feature bit that needs to be cap'ed by kvm_cpu_caps[]. */
> +#define TDX_CFG_F(name) \
> +({ \
> + TDX_VALIDATE_CPU_CAP_USAGE(name); \
> + tdx_cfg_caps |= feature_bit(name); \
> +})
> +
> +/*
> + * For a feature bit that KVM allows for TDX guests even though it isn't
> + * advertised through kvm_cpu_caps[], e.g. MWAIT. Use this version only when
> + * there is a justification.
> + */
> +#define TDX_CFG_EXTRA_F(name) \
> +({ \
> + TDX_VALIDATE_CPU_CAP_USAGE(name); \
> + tdx_cfg_caps_extra |= feature_bit(name); \
> +})
> +
> +#define tdx_cpu_cfg_cap_init(leaf, feature_initializers...) \
> +do { \
> + const u32 __maybe_unused tdx_cpu_cap_init_in_progress = leaf; \
> + u32 tdx_cfg_caps_extra = 0; \
> + u32 tdx_cfg_caps = 0; \
> + \
> + feature_initializers \
> + tdx_cpu_cfg_caps[leaf] = (tdx_cfg_caps & kvm_cpu_caps[leaf]) | \
> + tdx_cfg_caps_extra; \
> +} while (0)
> +
> +/*
> + * Initialize tdx_cpu_cfg_caps[], the list of CPUID features that KVM
> + * supports for TDX guests. It covers only the directly configurable CPUID
> + * bits reported by the TDX module; features controlled by XFAM and
> + * ATTRIBUTES are handled separately.
> + */
> +static void __init tdx_initialize_cpu_cfg_caps(void)
> +{
> + tdx_cpu_cfg_cap_init(CPUID_1_ECX,
> + /*
> + * KVM allows userspace to enumerate MONITOR+MWAIT support to
> + * the guest, but the MWAIT feature flag is never advertised
> + * to userspace for non-TDX VMs.
> + */
> + TDX_CFG_EXTRA_F(MWAIT),
> + /*
> + * XTPR can be exposed to a TD, but it never takes effect in
> + * the underlying hardware when the TD changes
> + * IA32_MISC_ENABLE[23].
> + */
> + TDX_CFG_EXTRA_F(XTPR),
> + TDX_CFG_F(TSC_DEADLINE_TIMER),
> + TDX_CFG_F(AVX),
> + TDX_CFG_F(F16C),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_1_EDX,
> + TDX_CFG_F(MCE),
> + TDX_CFG_F(MTRR),
> + TDX_CFG_F(MCA),
> + TDX_CFG_F(SELFSNOOP),
> + /*
> + * HT is a topology enumeration bit that KVM doesn't care
> + * about, but userspace may want to expose it to the guests.
> + */
> + TDX_CFG_EXTRA_F(HT),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_7_0_EBX,
> + TDX_CFG_F(BMI1),
> + /* HLE */
> + TDX_CFG_F(BMI2),
> + TDX_CFG_F(ERMS),
> + /* RTM */
> + TDX_CFG_F(AVX512F),
> + TDX_CFG_F(AVX512DQ),
> + TDX_CFG_F(ADX),
> + TDX_CFG_F(AVX512IFMA),
> + TDX_CFG_F(AVX512PF),
> + TDX_CFG_F(AVX512ER),
> + TDX_CFG_F(AVX512CD),
> + TDX_CFG_F(AVX512BW),
> + TDX_CFG_F(AVX512VL),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_7_ECX,
> + TDX_CFG_F(UMIP),
> + /* WAITPKG */
> + TDX_CFG_F(AVX512_VBMI2),
> + TDX_CFG_F(GFNI),
> + TDX_CFG_F(VAES),
> + TDX_CFG_F(VPCLMULQDQ),
> + TDX_CFG_F(AVX512_VNNI),
> + TDX_CFG_F(AVX512_BITALG),
> + TDX_CFG_F(AVX512_VPOPCNTDQ),
> + TDX_CFG_F(LA57),
> + TDX_CFG_F(RDPID),
> + TDX_CFG_F(CLDEMOTE),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_7_EDX,
> + TDX_CFG_F(AVX512_4VNNIW),
> + TDX_CFG_F(AVX512_4FMAPS),
> + TDX_CFG_F(FSRM),
> + TDX_CFG_F(AVX512_VP2INTERSECT),
> + TDX_CFG_F(SERIALIZE),
> + TDX_CFG_F(TSXLDTRK),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_7_1_EAX,
> + TDX_CFG_F(SHA512),
> + TDX_CFG_F(SM3),
> + TDX_CFG_F(SM4),
> + TDX_CFG_F(AVX_VNNI),
> + TDX_CFG_F(AVX512_BF16),
> + TDX_CFG_F(CMPCCXADD),
> + TDX_CFG_F(FZRM),
> + TDX_CFG_F(FSRS),
> + TDX_CFG_F(FSRC),
> + /* FRED */
> + TDX_CFG_F(LKGS),
> + TDX_CFG_F(WRMSRNS),
> + TDX_CFG_F(AMX_FP16),
> + TDX_CFG_F(AVX_IFMA),
> + TDX_CFG_F(LAM),
> + TDX_CFG_F(MOVRS),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_7_1_EDX,
> + TDX_CFG_F(AVX_VNNI_INT8),
> + TDX_CFG_F(AVX_NE_CONVERT),
> + TDX_CFG_F(AVX_VNNI_INT16),
> + TDX_CFG_F(PREFETCHITI),
> + TDX_CFG_F(AVX10),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_7_2_EDX,
> + TDX_CFG_F(DDPD_U),
> + TDX_CFG_F(MCDT_NO),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_1E_1_EAX,
> + TDX_CFG_F(AMX_COMPLEX_ALIAS),
> + TDX_CFG_F(AMX_FP8),
> + TDX_CFG_F(AMX_TF32),
> + TDX_CFG_F(AMX_AVX512),
> + TDX_CFG_F(AMX_MOVRS),
> + );
> +
> + tdx_cpu_cfg_cap_init(CPUID_8000_0008_EBX,
> + TDX_CFG_F(WBNOINVD),
> + );
> +}
> +
> +#undef TDX_CFG_F
> +#undef TDX_CFG_EXTRA_F
>
> bool enable_tdx __ro_after_init;
> module_param_named(tdx, enable_tdx, bool, 0444);
> @@ -3487,6 +3649,8 @@ int __init tdx_hardware_setup(void)
> return r;
> }
>
> + tdx_initialize_cpu_cfg_caps();
> +
> KVM_SANITY_CHECK_VM_STRUCT_SIZE(kvm_tdx);
>
> vt_x86_ops.vm_size = max_t(unsigned int, vt_x86_ops.vm_size, sizeof(struct kvm_tdx));
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable
2026-09-17 7:25 ` [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Binbin Wu
2026-09-22 21:11 ` Edgecombe, Rick P
@ 2026-09-24 6:44 ` Xiaoyao Li
1 sibling, 0 replies; 27+ messages in thread
From: Xiaoyao Li @ 2026-09-24 6:44 UTC (permalink / raw)
To: Binbin Wu, linux-kernel, kvm
Cc: seanjc, pbonzini, dave.hansen, andrew.cooper3, nik.borisov, kas,
rick.p.edgecombe, chao.gao, tony.lindgren, kishen.maloor,
dedekind1
On 9/17/2026 3:25 PM, Binbin Wu wrote:
> Add CORE_CAPABILITIES (CPUID.0x7.0.EDX[30]) to KVM's allowlist of TDX
> directly configurable CPUID feature bits, even though KVM doesn't support
> MSR_IA32_CORE_CAPS for TDX guests, to accommodate the legacy TDX module
> definition and userspace's stale knowledge of it.
>
> Older TDX specifications define the CORE_CAPABILITIES CPUID bit as
> fixed-1, so userspace may expect the bit to be enabled for TDs. #VE
> reduction turns it into a directly configurable bit, so leaving it out of
> the allowlist would make the bit impossible to enable once KVM starts
> validating userspace's CPUID input, i.e. would be a surprising behavior
> change for such userspace.
>
> Reporting CORE_CAPABILITIES as directly configurable also lets userspace
> detect that the bit is no longer fixed-1, and thus correct its stale
> knowledge.
>
> Keep MSR_IA32_CORE_CAPS unsupported for TDs, as no existing TDX user needs
> guest access to the MSR.
>
> Note, CORE_CAPABILITIES is the only bit that is unsupported by KVM *and*
> changed from fixed-1 to directly configurable by #VE reduction, and no
> further #VE reductions are expected.
>
> Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
> Reviewed-by: Tony Lindgren <tony.lindgren@linux.intel.com>
Reviewed-by: Xiaoyao Li <xiaoyao.li@intel.com>
> ---
> v4:
> - Add #VE reduction related background to the changelog. (Kishen)
> - Add RB from Tony.
>
> v3:
> - Drop the code for MSR_IA32_CORE_CAPS access.
> ---
> arch/x86/kvm/vmx/tdx.c | 7 +++++++
> 1 file changed, 7 insertions(+)
>
> diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
> index 2b51a85c998e8..b34afc52b714e 100644
> --- a/arch/x86/kvm/vmx/tdx.c
> +++ b/arch/x86/kvm/vmx/tdx.c
> @@ -165,6 +165,13 @@ static void __init tdx_initialize_cpu_cfg_caps(void)
> TDX_CFG_F(AVX512_VP2INTERSECT),
> TDX_CFG_F(SERIALIZE),
> TDX_CFG_F(TSXLDTRK),
> + /*
> + * KVM doesn't support MSR_IA32_CORE_CAPS, but older TDX specs
> + * define this bit as fixed-1. Report it as configurable to
> + * accommodate the legacy TDX module definition, and to let
> + * userspace detect that the bit is no longer fixed-1.
> + */
> + TDX_CFG_EXTRA_F(CORE_CAPABILITIES),
> );
>
> tdx_cpu_cfg_cap_init(CPUID_7_1_EAX,
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-24 6:36 ` Xiaoyao Li
@ 2026-09-24 7:49 ` Binbin Wu
2026-09-24 9:30 ` Xiaoyao Li
0 siblings, 1 reply; 27+ messages in thread
From: Binbin Wu @ 2026-09-24 7:49 UTC (permalink / raw)
To: Xiaoyao Li
Cc: linux-kernel, kvm, seanjc, pbonzini, dave.hansen, andrew.cooper3,
nik.borisov, kas, rick.p.edgecombe, chao.gao, tony.lindgren,
kishen.maloor, dedekind1
On 9/24/2026 2:36 PM, Xiaoyao Li wrote:
Thanks a lot for the detailed review!
> On 9/17/2026 3:25 PM, Binbin Wu wrote:
>> Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable
>> CPUID feature bits that KVM supports, and build the masks during TDX
>> hardware setup via tdx_initialize_cpu_cfg_caps().
>>
>> The TDX module reports the CPUID bits that the VMM can directly configure
>> for a TD, but KVM cannot blindly expose all reported bits to userspace.
>> Certain features imply additional architectural state, e.g. one or more
>> MSRs, that KVM must explicitly manage across host/guest transitions to
>> prevent host state corruption. The existing hardcoded denylist cannot
>> account for new host state clobbering features introduced by future TDX
>> modules.
>>
>> Except for a few fixed-1 bits required for basic TDX support, host state
>> clobbering features are either directly configurable or gated by TD
>> ATTRIBUTES/XFAM, which KVM already validates. Tracking only the directly
>> configurable bits to build an allowlist is therefore sufficient.
>>
>> Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can
>> be built with similar feature-name based initializers. Directly
>> configurable non-feature bits will be handled separately.
>>
>> The allowlist is prepared to be consumed by later patches to filter
>> KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to
>> KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID
>> feature bits stay hidden from userspace until KVM explicitly opts in.
>>
>> By default, intersect the allowlist with kvm_cpu_caps[] via TDX_CFG_F().
>> Requiring support for non-TDX VMs avoids committing to TDX-specific
>> behavior before general KVM support is established, and respects KVM's
>> logic around disabling certain features, since the reasons for disabling
>> them could apply to TDX as well. Allow exceptions through
>> TDX_CFG_EXTRA_F() only with sufficient justification.
>
> One name just crosses my mind. How about TDX_CFG_INDEP_F()? where INDEP stands
> for independent.
>
> - TDX_CFG_F() for features that are dependent on kvm_cpu_caps[]
> - TDX_CFG_INDEP_F() for features not dependent on kvm_cpu_caps[]
I am not sure if it's better.
But if people prefer independent to extra, I would use TDX_CFG_INDEPENDENT_F()
instead.
>
>> Add comments as placeholders for HLE, RTM, WAITPKG and FRED, which KVM
>> doesn't support for TDX yet.
>
> 1. why add comments for them? Because implementation wise, they are possible to
> be supported by KVM but just current KVM doesn't have code to support them yet?
>
> If so, are the others the features that are concluded to be impossible to be
> supported by KVM?
>
> 2. We don't need to talk about FRED, which is even not suppported for normal VMs
> yet. Unless we are trying to enable FRED for TDs specifically.
OK, I will drop them to avoid confusion.
>
>> Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
>> not advertised in kvm_cpu_caps[]. The remaining directly configurable
>> feature bits outside kvm_cpu_caps[] are left out of the allowlist:
>>
>> - Features forced to zero when #VE is reduced, or lacking KVM support
>> for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
>> RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
>> in a subsequent patch.
>
> why not split into two category? Because they are overlapped?
>
> - Features forced to zero when #VE is reduced
> - Features lack KVM support
>
Yes, there is overlap between them.
If split into two categories, it would be:
- Features forced to zero when #VE is reduced, which also lack KVM support
- Features unrelated to #VE reduction that lack KVM support
Because the changelog was already getting quite long, I combined them into a
single category.
> As for "features forced to zero when #VE is reduced", do you mean when TDX
> module supports "#VE reduction" feature, the features becomes fixed0? or when
> guest enables "reduce #VE", the guest see a 0 value even though the feature is
> still configurable to host userspace and host userspace configure it to 1?
#VE is reduced means the guest reduced the related #VE, i.e. TDCS.TD_CTRL.REDUCE_VE
is 1 and the related bit in TDCS.FEATURE_PARAVIRT_CTRL is 0.
>
>> - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
>> TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
>
> What is CID?
X86_FEATURE_CID, Context ID, more specifically, L1 CONTEXT ID
> and Is PBE related to any bit of IA32_MISC_ENABLE? Did you mean PEBS instead?
X86_FEATURE_PBE, Pending Break Enable.
Since kernel already has X86_FEATURE_XXX for them, so I use them directly without
other description.
>
> I don't understand what this category. It talks about something when
> TDCS.TD_CTLS.REDUCE_VE is set, which depends on guest TD's behavior. So what if
> the guest TD doesn't set TDCS.TD_CTLS.REDUCE_VE?
If the guest set TDCS.TD_CTLS.REDUCE_VE, write 1 to the the related bit trigger #GP.
If the guest doesn't set TDCS.TD_CTLS.REDUCE_VE, the value is passed to KVM, and
KVM will save the value without do corresponding virtualization.
Considering TDCS.TD_CTLS.REDUCE_VE is set by the Linux guest by default, it should
not allow these features.
>
>> - Features that can clobber host state and lack KVM support for TDX:
>> FRED.
>>
>> - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
>> TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
>> platform), and RAO_INT (defined only for future processors).
>
> PREFETCHWT1 (if the name is correct) is not supported by the kernel right? For
> such unknown features, it's straightforward to not support it.
>
> AMX-TRANSPOSE has been turned into reserved, and kernel never supports it. I
> think we don't need to mention it at all.
The guest could be other OS, not just Linux.
>
> RAO_INT is fixed 0 for TDX 1.5?
I was doing this patch series based on the latest CSV version of the CPUID
virtualization document from "Intel TDX Module ABI Definitions" updated in June 2026.
https://cdrdv2.intel.com/v1/dl/getContent/795381
>
> At last, I think we can just classify all of the above as 3 types:
>
> - Features that are directly configurable by TDX and in kvm_cpu_caps[].
> use TDX_CFG_F()
>
> - Features that are directly configurable by TDX and KVM actually supports them
> or KVM can support them but not in kvm_cpu_caps[]. Use TDX_CFG_EXTRA_F().
> For example, MWAIT, XTPR, and HT.
>
> - Features that are directly configurable by TDX but KVM or kenrel doesn't
> support them.
> For example, RDT_A, RDT_M, PREFETCHWT1, PSN,
Because excluding some bits changes the ABI, I listed all these directly configurable
features out of kvm_cpu_caps[] to explain why we must exclude some bits.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-24 7:49 ` Binbin Wu
@ 2026-09-24 9:30 ` Xiaoyao Li
2026-09-24 11:41 ` Binbin Wu
0 siblings, 1 reply; 27+ messages in thread
From: Xiaoyao Li @ 2026-09-24 9:30 UTC (permalink / raw)
To: Binbin Wu
Cc: linux-kernel, kvm, seanjc, pbonzini, dave.hansen, andrew.cooper3,
nik.borisov, kas, rick.p.edgecombe, chao.gao, tony.lindgren,
kishen.maloor, dedekind1
On 9/24/2026 3:49 PM, Binbin Wu wrote:
> On 9/24/2026 2:36 PM, Xiaoyao Li wrote:
>
> Thanks a lot for the detailed review!
>
>> On 9/17/2026 3:25 PM, Binbin Wu wrote:
>>> Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable
>>> CPUID feature bits that KVM supports, and build the masks during TDX
>>> hardware setup via tdx_initialize_cpu_cfg_caps().
>>>
>>> The TDX module reports the CPUID bits that the VMM can directly configure
>>> for a TD, but KVM cannot blindly expose all reported bits to userspace.
>>> Certain features imply additional architectural state, e.g. one or more
>>> MSRs, that KVM must explicitly manage across host/guest transitions to
>>> prevent host state corruption. The existing hardcoded denylist cannot
>>> account for new host state clobbering features introduced by future TDX
>>> modules.
>>>
>>> Except for a few fixed-1 bits required for basic TDX support, host state
>>> clobbering features are either directly configurable or gated by TD
>>> ATTRIBUTES/XFAM, which KVM already validates. Tracking only the directly
>>> configurable bits to build an allowlist is therefore sufficient.
>>>
>>> Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can
>>> be built with similar feature-name based initializers. Directly
>>> configurable non-feature bits will be handled separately.
>>>
>>> The allowlist is prepared to be consumed by later patches to filter
>>> KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to
>>> KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID
>>> feature bits stay hidden from userspace until KVM explicitly opts in.
>>>
>>> By default, intersect the allowlist with kvm_cpu_caps[] via TDX_CFG_F().
>>> Requiring support for non-TDX VMs avoids committing to TDX-specific
>>> behavior before general KVM support is established, and respects KVM's
>>> logic around disabling certain features, since the reasons for disabling
>>> them could apply to TDX as well. Allow exceptions through
>>> TDX_CFG_EXTRA_F() only with sufficient justification.
>>
>> One name just crosses my mind. How about TDX_CFG_INDEP_F()? where INDEP stands
>> for independent.
>>
>> - TDX_CFG_F() for features that are dependent on kvm_cpu_caps[]
>> - TDX_CFG_INDEP_F() for features not dependent on kvm_cpu_caps[]
>
> I am not sure if it's better.
> But if people prefer independent to extra, I would use TDX_CFG_INDEPENDENT_F()
> instead.
>
>>
>>> Add comments as placeholders for HLE, RTM, WAITPKG and FRED, which KVM
>>> doesn't support for TDX yet.
>>
>> 1. why add comments for them? Because implementation wise, they are possible to
>> be supported by KVM but just current KVM doesn't have code to support them yet?
>>
>> If so, are the others the features that are concluded to be impossible to be
>> supported by KVM?
>>
>> 2. We don't need to talk about FRED, which is even not suppported for normal VMs
>> yet. Unless we are trying to enable FRED for TDs specifically.
>
> OK, I will drop them to avoid confusion.
>
>>
>>> Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
>>> not advertised in kvm_cpu_caps[]. The remaining directly configurable
>>> feature bits outside kvm_cpu_caps[] are left out of the allowlist:
>>>
>>> - Features forced to zero when #VE is reduced, or lacking KVM support
>>> for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
>>> RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
>>> in a subsequent patch.
>>
>> why not split into two category? Because they are overlapped?
>>
>> - Features forced to zero when #VE is reduced
>> - Features lack KVM support
>>
>
> Yes, there is overlap between them.
> If split into two categories, it would be:
> - Features forced to zero when #VE is reduced, which also lack KVM support
> - Features unrelated to #VE reduction that lack KVM support
>
> Because the changelog was already getting quite long, I combined them into a
> single category.
I think that should be OK. It only adds 2-3 lines?
>> As for "features forced to zero when #VE is reduced", do you mean when TDX
>> module supports "#VE reduction" feature, the features becomes fixed0? or when
>> guest enables "reduce #VE", the guest see a 0 value even though the feature is
>> still configurable to host userspace and host userspace configure it to 1?
>
>
> #VE is reduced means the guest reduced the related #VE, i.e. TDCS.TD_CTRL.REDUCE_VE
> is 1 and the related bit in TDCS.FEATURE_PARAVIRT_CTRL is 0.
I think "features forced to zero when #VE is reduced" doesn't matter at all
here. KVM cannot predict TD's behavior and KVM cannot make decision just based
on one of the possibility of TD's behavior. i.e., KVM still needs to consider if
it can support the features when TD doesn't enable reduce_ve. So what really
matters is KVM doesn't/cannot support these features correctly.
So it's just one category: features not supported by KVM.
We don't need to talk #VE reduction at all.
>>
>>> - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
>>> TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
>>
>> What is CID?
>
> X86_FEATURE_CID, Context ID, more specifically, L1 CONTEXT ID
My bad. I only search 'CID' in arch/x86/kvm/cpuid.c where uses CNXT-ID instead
of CID.
>> and Is PBE related to any bit of IA32_MISC_ENABLE? Did you mean PEBS instead?
>
> X86_FEATURE_PBE, Pending Break Enable.
>
> Since kernel already has X86_FEATURE_XXX for them, so I use them directly without
> other description.
>
>>
>> I don't understand what this category. It talks about something when
>> TDCS.TD_CTLS.REDUCE_VE is set, which depends on guest TD's behavior. So what if
>> the guest TD doesn't set TDCS.TD_CTLS.REDUCE_VE?
>
> If the guest set TDCS.TD_CTLS.REDUCE_VE, write 1 to the the related bit trigger #GP.
> If the guest doesn't set TDCS.TD_CTLS.REDUCE_VE, the value is passed to KVM, and
> KVM will save the value without do corresponding virtualization.
>
> Considering TDCS.TD_CTLS.REDUCE_VE is set by the Linux guest by default, it should
> not allow these features.
After some search, I find that X86_FEATURE_CID (CPUID.0x1:ECX[10]) is related to
MSR_IA32_MISC_ENABLE[24], which seems to be a model specific bit.
And on SDM, vol3. 14.5.6 L1 Data Cache Context Mode
L1 data cache context mode is a feature of processors based on the Intel
NetBurst microarchitecture that support Intel Hyper-Threading Technology.
When CPUID.01H:ECX[10] =1, the processor supports setting L1 data cache
context mode using the L1 data cache context mode flag (IA32_MISC_ENABLE[bit
24]). Selectable modes are adaptive mode (default) and shared mode.
CPUID.0x1:ECX[10] is 0 on SPR and EMR (I didn't check more). It looks this CPUID
bit will not show up on any TDX capable CPUs. I think we can report it to TDX
architects and just make the bit fixed0 as a spec fix?
For X86_FEATURE_PBE (CPUID.0x1:EDX[31]), it's related to bit 10 of MSR
IA32_MISC_ENABLE. The CPUID is 1 on SPR and EMR. But bit 10 of MSR
IA32_MISC_ENABLE is always 0 and setting bit 10 to 1 fails. It seems though
CPUID.0x1:EDX[31] is set to 1, but it's not related to bit 10 of MSR
IA32_MISC_ENABLE?
>>
>>> - Features that can clobber host state and lack KVM support for TDX:
>>> FRED.
>>>
>>> - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
>>> TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
>>> platform), and RAO_INT (defined only for future processors).
>>
>> PREFETCHWT1 (if the name is correct) is not supported by the kernel right? For
>> such unknown features, it's straightforward to not support it.
>>
>> AMX-TRANSPOSE has been turned into reserved, and kernel never supports it. I
>> think we don't need to mention it at all.
>
> The guest could be other OS, not just Linux.
why guest OS matters? Are you considering the compatibility that it will break
the case where host userspace configures them to TDs and the TD OS (other than
Linux) is using them? If from this perspective, other bits will be a concern as
well.
>>
>> RAO_INT is fixed 0 for TDX 1.5?
>
> I was doing this patch series based on the latest CSV version of the CPUID
> virtualization document from "Intel TDX Module ABI Definitions" updated in June 2026.
> https://cdrdv2.intel.com/v1/dl/getContent/795381
I see, it might be defined for TDX 2.0/3.0.
>>
>> At last, I think we can just classify all of the above as 3 types:
>>
>> - Features that are directly configurable by TDX and in kvm_cpu_caps[].
>> use TDX_CFG_F()
>>
>> - Features that are directly configurable by TDX and KVM actually supports them
>> or KVM can support them but not in kvm_cpu_caps[]. Use TDX_CFG_EXTRA_F().
>> For example, MWAIT, XTPR, and HT.
>>
>> - Features that are directly configurable by TDX but KVM or kenrel doesn't
>> support them.
>> For example, RDT_A, RDT_M, PREFETCHWT1, PSN,
>
> Because excluding some bits changes the ABI, I listed all these directly configurable
> features out of kvm_cpu_caps[] to explain why we must exclude some bits.
I think all the bits that are directly configurable but out of kvm_cpu_caps[]
and not supported by KVM fall into the last category. You can just list all of
them instead of my "for example".
As for the bits like FRED/RAO_INT, we don't need to consider them since they are
not generally supported by KVM. They are just the new features to KVM, which
can be just treated as unknown though the TDX spec has defined the
virtualization type of them.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-24 9:30 ` Xiaoyao Li
@ 2026-09-24 11:41 ` Binbin Wu
2026-09-24 13:55 ` Xiaoyao Li
2026-09-24 14:33 ` Xiaoyao Li
0 siblings, 2 replies; 27+ messages in thread
From: Binbin Wu @ 2026-09-24 11:41 UTC (permalink / raw)
To: Xiaoyao Li
Cc: linux-kernel, kvm, seanjc, pbonzini, dave.hansen, andrew.cooper3,
nik.borisov, kas, rick.p.edgecombe, chao.gao, tony.lindgren,
kishen.maloor, dedekind1
[...]
>>>> Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
>>>> not advertised in kvm_cpu_caps[]. The remaining directly configurable
>>>> feature bits outside kvm_cpu_caps[] are left out of the allowlist:
>>>>
>>>> - Features forced to zero when #VE is reduced, or lacking KVM support
>>>> for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
>>>> RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
>>>> in a subsequent patch.
>>>
>>> why not split into two category? Because they are overlapped?
>>>
>>> - Features forced to zero when #VE is reduced
>>> - Features lack KVM support
>>>
>>
>> Yes, there is overlap between them.
>> If split into two categories, it would be:
>> - Features forced to zero when #VE is reduced, which also lack KVM support
>> - Features unrelated to #VE reduction that lack KVM support
>>
>> Because the changelog was already getting quite long, I combined them into a
>> single category.
>
> I think that should be OK. It only adds 2-3 lines?
OK.
>
>>> As for "features forced to zero when #VE is reduced", do you mean when TDX
>>> module supports "#VE reduction" feature, the features becomes fixed0? or when
>>> guest enables "reduce #VE", the guest see a 0 value even though the feature is
>>> still configurable to host userspace and host userspace configure it to 1?
>>
>>
>> #VE is reduced means the guest reduced the related #VE, i.e. TDCS.TD_CTRL.REDUCE_VE
>> is 1 and the related bit in TDCS.FEATURE_PARAVIRT_CTRL is 0.
>
> I think "features forced to zero when #VE is reduced" doesn't matter at all
> here.
I think it still matters.
If there is a feature that is virtualized by the TDX module itself, but KVM doesn't
support it for non-TDX VMs, I think it should be allowed.
These features can't be added to the allow list because neither TDX module nor
KVM do the virtualization for them.
KVM cannot predict TD's behavior and KVM cannot make decision just based
> on one of the possibility of TD's behavior. i.e., KVM still needs to consider if
> it can support the features when TD doesn't enable reduce_ve. So what really
> matters is KVM doesn't/cannot support these features correctly.
>
> So it's just one category: features not supported by KVM.
> We don't need to talk #VE reduction at all.
>
>>>
>>>> - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
>>>> TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
>>>
>>> What is CID?
>>
>> X86_FEATURE_CID, Context ID, more specifically, L1 CONTEXT ID
>
> My bad. I only search 'CID' in arch/x86/kvm/cpuid.c where uses CNXT-ID instead
> of CID.
>
>>> and Is PBE related to any bit of IA32_MISC_ENABLE? Did you mean PEBS instead?
>>
>> X86_FEATURE_PBE, Pending Break Enable.
>>
>> Since kernel already has X86_FEATURE_XXX for them, so I use them directly without
>> other description.
>>
>>>
>>> I don't understand what this category. It talks about something when
>>> TDCS.TD_CTLS.REDUCE_VE is set, which depends on guest TD's behavior. So what if
>>> the guest TD doesn't set TDCS.TD_CTLS.REDUCE_VE?
>>
>> If the guest set TDCS.TD_CTLS.REDUCE_VE, write 1 to the the related bit trigger #GP.
>> If the guest doesn't set TDCS.TD_CTLS.REDUCE_VE, the value is passed to KVM, and
>> KVM will save the value without do corresponding virtualization.
>>
>> Considering TDCS.TD_CTLS.REDUCE_VE is set by the Linux guest by default, it should
>> not allow these features.
>
> After some search, I find that X86_FEATURE_CID (CPUID.0x1:ECX[10]) is related to
> MSR_IA32_MISC_ENABLE[24], which seems to be a model specific bit.
>
> And on SDM, vol3. 14.5.6 L1 Data Cache Context Mode
>
> L1 data cache context mode is a feature of processors based on the Intel
> NetBurst microarchitecture that support Intel Hyper-Threading Technology.
> When CPUID.01H:ECX[10] =1, the processor supports setting L1 data cache
> context mode using the L1 data cache context mode flag (IA32_MISC_ENABLE[bit
> 24]). Selectable modes are adaptive mode (default) and shared mode.
>
> CPUID.0x1:ECX[10] is 0 on SPR and EMR (I didn't check more). It looks this CPUID
> bit will not show up on any TDX capable CPUs. I think we can report it to TDX
> architects and just make the bit fixed0 as a spec fix?
I think we can confirm with the TDX module team. If it will never show up on TDX
capable CPUs, then this is similar to PREFETCHWT1.
But I am not sure we want new fixed-0 bits.
>
> For X86_FEATURE_PBE (CPUID.0x1:EDX[31]), it's related to bit 10 of MSR
> IA32_MISC_ENABLE. The CPUID is 1 on SPR and EMR. But bit 10 of MSR
> IA32_MISC_ENABLE is always 0 and setting bit 10 to 1 fails. It seems though
> CPUID.0x1:EDX[31] is set to 1, but it's not related to bit 10 of MSR
> IA32_MISC_ENABLE?
In the ISE (319433-062, table 1-6), it says:
PBE:
Pending Break Enable. The processor supports the use of the FERR#/PBE# pin
when the processor is in the stop-clock state (STPCLK# is asserted) to signal
the processor that an interrupt is pending and that the processor should
return to normal operation to handle the interrupt. Bit 10 (PBE enable) in the
IA32_MISC_ENABLE MSR enables this capability.
TDX module doesn't allow IA32_MISC_ENABLE[10] to be 1 is since it's reserved
TDCS.TD_CTLS.REDUCE_VE is set, so my understanding was the CPUID bit should be 0.
It seems bit 10 is phased out in IA32_MISC_ENABLE on modern processors like SPR
and EMR, since the bit 10 is reserved in "IA-32 Architectural MSRs" table. But
somehow, CPUID still reports PBE as supported.
I don't know why the CPUID is disconnected with the MSR.
But this bit is present on the host, I think we can add it to the allowlist.
>
>>>
>>>> - Features that can clobber host state and lack KVM support for TDX:
>>>> FRED.
>>>>
>>>> - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
>>>> TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
>>>> platform), and RAO_INT (defined only for future processors).
>>>
>>> PREFETCHWT1 (if the name is correct) is not supported by the kernel right? For
>>> such unknown features, it's straightforward to not support it.
>>>
>>> AMX-TRANSPOSE has been turned into reserved, and kernel never supports it. I
>>> think we don't need to mention it at all.
>>
>> The guest could be other OS, not just Linux.
>
> why guest OS matters? Are you considering the compatibility that it will break
> the case where host userspace configures them to TDs and the TD OS (other than
> Linux) is using them? If from this perspective, other bits will be a concern as
> well.
That's the reason I explained in the commit message why other bits should not be
added to the allow list because if they are used, since they are can't be properly
virtualized, there should have been issues already before this patch series.
>
>>>
>>> RAO_INT is fixed 0 for TDX 1.5?
>>
>> I was doing this patch series based on the latest CSV version of the CPUID
>> virtualization document from "Intel TDX Module ABI Definitions" updated in June 2026.
>> https://cdrdv2.intel.com/v1/dl/getContent/795381
>
> I see, it might be defined for TDX 2.0/3.0.
>
>>>
>>> At last, I think we can just classify all of the above as 3 types:
>>>
>>> - Features that are directly configurable by TDX and in kvm_cpu_caps[].
>>> use TDX_CFG_F()
>>>
>>> - Features that are directly configurable by TDX and KVM actually supports them
>>> or KVM can support them but not in kvm_cpu_caps[]. Use TDX_CFG_EXTRA_F().
>>> For example, MWAIT, XTPR, and HT.
>>>
>>> - Features that are directly configurable by TDX but KVM or kenrel doesn't
>>> support them.
>>> For example, RDT_A, RDT_M, PREFETCHWT1, PSN,
>>
>> Because excluding some bits changes the ABI, I listed all these directly configurable
>> features out of kvm_cpu_caps[] to explain why we must exclude some bits.
>
> I think all the bits that are directly configurable but out of kvm_cpu_caps[]
> and not supported by KVM fall into the last category. You can just list all of
> them instead of my "for example".
As I mentioned above, if there is a feature that is virtualized by the TDX module itself,
but KVM doesn't support it for non-TDX VMs, it should be allowed. Although I don't have
a concrete example. I think we still need to mention #VE reduction to some degree, but
let me see how to simplify the description.
>
> As for the bits like FRED/RAO_INT, we don't need to consider them since they are
> not generally supported by KVM. They are just the new features to KVM, which
> can be just treated as unknown though the TDX spec has defined the
> virtualization type of them.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-24 11:41 ` Binbin Wu
@ 2026-09-24 13:55 ` Xiaoyao Li
2026-09-24 14:33 ` Xiaoyao Li
1 sibling, 0 replies; 27+ messages in thread
From: Xiaoyao Li @ 2026-09-24 13:55 UTC (permalink / raw)
To: Binbin Wu
Cc: linux-kernel, kvm, seanjc, pbonzini, dave.hansen, andrew.cooper3,
nik.borisov, kas, rick.p.edgecombe, chao.gao, tony.lindgren,
kishen.maloor, dedekind1
On 9/24/2026 7:41 PM, Binbin Wu wrote:
...
>>>> As for "features forced to zero when #VE is reduced", do you mean when TDX
>>>> module supports "#VE reduction" feature, the features becomes fixed0? or when
>>>> guest enables "reduce #VE", the guest see a 0 value even though the feature is
>>>> still configurable to host userspace and host userspace configure it to 1?
>>>
>>>
>>> #VE is reduced means the guest reduced the related #VE, i.e. TDCS.TD_CTRL.REDUCE_VE
>>> is 1 and the related bit in TDCS.FEATURE_PARAVIRT_CTRL is 0.
>>
>> I think "features forced to zero when #VE is reduced" doesn't matter at all
>> here.
>
> I think it still matters.
>
> If there is a feature that is virtualized by the TDX module itself, but KVM doesn't
> support it for non-TDX VMs, I think it should be allowed.
>
> These features can't be added to the allow list because neither TDX module nor
> KVM do the virtualization for them.
... <snip>
>> I think all the bits that are directly configurable but out of kvm_cpu_caps[]
>> and not supported by KVM fall into the last category. You can just list all of
>> them instead of my "for example".
>
> As I mentioned above, if there is a feature that is virtualized by the TDX module itself,
> but KVM doesn't support it for non-TDX VMs, it should be allowed. Although I don't have
> a concrete example. I think we still need to mention #VE reduction to some degree, but
> let me see how to simplify the description.
(This reply only concentrates on the discussion of "if there is a feature that
is virtualized by the TDX module itself")
If you are talking about the features related to #VE, I think there is no.
Because if TDX module itself can virtualize the feature, then it shouldn't be
#VE at all. #VE means it needs para-virt support from host VMM.
If you are talking about the features unrelated to #VE. I think it's just the
case I argued for in previous v3, that we can support a feature for TDX first
before non-TDX VMs. I think we all agreed that such case can happen and we
decided to support such case when there is a real case.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
2026-09-24 11:41 ` Binbin Wu
2026-09-24 13:55 ` Xiaoyao Li
@ 2026-09-24 14:33 ` Xiaoyao Li
1 sibling, 0 replies; 27+ messages in thread
From: Xiaoyao Li @ 2026-09-24 14:33 UTC (permalink / raw)
To: Binbin Wu
Cc: linux-kernel, kvm, seanjc, pbonzini, dave.hansen, andrew.cooper3,
nik.borisov, kas, rick.p.edgecombe, chao.gao, tony.lindgren,
kishen.maloor, dedekind1
On 9/24/2026 7:41 PM, Binbin Wu wrote:
>> After some search, I find that X86_FEATURE_CID (CPUID.0x1:ECX[10]) is related to
>> MSR_IA32_MISC_ENABLE[24], which seems to be a model specific bit.
>>
>> And on SDM, vol3. 14.5.6 L1 Data Cache Context Mode
>>
>> L1 data cache context mode is a feature of processors based on the Intel
>> NetBurst microarchitecture that support Intel Hyper-Threading Technology.
>> When CPUID.01H:ECX[10] =1, the processor supports setting L1 data cache
>> context mode using the L1 data cache context mode flag (IA32_MISC_ENABLE[bit
>> 24]). Selectable modes are adaptive mode (default) and shared mode.
>>
>> CPUID.0x1:ECX[10] is 0 on SPR and EMR (I didn't check more). It looks this CPUID
>> bit will not show up on any TDX capable CPUs. I think we can report it to TDX
>> architects and just make the bit fixed0 as a spec fix?
> I think we can confirm with the TDX module team. If it will never show up on TDX
> capable CPUs, then this is similar to PREFETCHWT1.
>
> But I am not sure we want new fixed-0 bits.
Leave it as current "Configured & Native" is also OK. Since no TDX capable CPU
has this bit as 1, it's effectively a fixed-0 bit.
>> For X86_FEATURE_PBE (CPUID.0x1:EDX[31]), it's related to bit 10 of MSR
>> IA32_MISC_ENABLE. The CPUID is 1 on SPR and EMR. But bit 10 of MSR
>> IA32_MISC_ENABLE is always 0 and setting bit 10 to 1 fails. It seems though
>> CPUID.0x1:EDX[31] is set to 1, but it's not related to bit 10 of MSR
>> IA32_MISC_ENABLE?
> In the ISE (319433-062, table 1-6), it says:
> PBE:
> Pending Break Enable. The processor supports the use of the FERR#/PBE# pin
> when the processor is in the stop-clock state (STPCLK# is asserted) to signal
> the processor that an interrupt is pending and that the processor should
> return to normal operation to handle the interrupt. Bit 10 (PBE enable) in the
> IA32_MISC_ENABLE MSR enables this capability.
>
> TDX module doesn't allow IA32_MISC_ENABLE[10] to be 1 is since it's reserved
> TDCS.TD_CTLS.REDUCE_VE is set, so my understanding was the CPUID bit should be 0.
>
> It seems bit 10 is phased out in IA32_MISC_ENABLE on modern processors like SPR
> and EMR, since the bit 10 is reserved in "IA-32 Architectural MSRs" table. But
> somehow, CPUID still reports PBE as supported.
yeah. arch/x86/include/asm/msr-index.h divides the bits of MSR IA32_MISC_ENABLE
into two groups: architectural and model-specific. Bit 10 falls into model-specific.
Maybe we can ask internally why CPUID still reports 1.
> I don't know why the CPUID is disconnected with the MSR.
> But this bit is present on the host, I think we can add it to the allowlist.
I have opposite opinion. Since 1) this feature is related to hardware processor
PIN, 2) KVM never advertise it to normal VMs, and 3) no usage in kernel for this
feature. I think no TD relies on it and don't allow it should be OK.
^ permalink raw reply [flat|nested] 27+ messages in thread
end of thread, other threads:[~2026-09-24 14:33 UTC | newest]
Thread overview: 27+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-17 7:25 [PATCH v4 0/4] KVM: TDX: Validate directly configurable CPUID bits Binbin Wu
2026-09-17 7:25 ` [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Binbin Wu
2026-09-23 0:01 ` Edgecombe, Rick P
2026-09-23 0:14 ` Binbin Wu
2026-09-23 0:45 ` Edgecombe, Rick P
2026-09-23 0:57 ` Binbin Wu
2026-09-23 1:09 ` Edgecombe, Rick P
2026-09-23 1:17 ` Binbin Wu
2026-09-23 11:21 ` Xiaoyao Li
2026-09-23 14:25 ` Edgecombe, Rick P
2026-09-24 2:05 ` Xiaoyao Li
2026-09-24 6:36 ` Xiaoyao Li
2026-09-24 7:49 ` Binbin Wu
2026-09-24 9:30 ` Xiaoyao Li
2026-09-24 11:41 ` Binbin Wu
2026-09-24 13:55 ` Xiaoyao Li
2026-09-24 14:33 ` Xiaoyao Li
2026-09-17 7:25 ` [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Binbin Wu
2026-09-22 21:11 ` Edgecombe, Rick P
2026-09-23 0:03 ` Binbin Wu
2026-09-24 6:44 ` Xiaoyao Li
2026-09-17 7:25 ` [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits Binbin Wu
2026-09-23 0:16 ` Edgecombe, Rick P
2026-09-23 0:28 ` Binbin Wu
2026-09-23 0:34 ` Edgecombe, Rick P
2026-09-17 7:25 ` [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM Binbin Wu
2026-09-23 0:16 ` Edgecombe, Rick P
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®