From: Binbin Wu <binbin.wu@linux.intel.com>
To: Xiaoyao Li <xiaoyao.li@intel.com>
Cc: linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
seanjc@google.com, pbonzini@redhat.com,
dave.hansen@linux.intel.com, andrew.cooper3@citrix.com,
nik.borisov@suse.com, kas@kernel.org, rick.p.edgecombe@intel.com,
chao.gao@intel.com, tony.lindgren@linux.intel.com,
kishen.maloor@intel.com, dedekind1@gmail.com
Subject: Re: [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
Date: Thu, 24 Sep 2026 15:49:45 +0800 [thread overview]
Message-ID: <062f5365-7eb5-4e04-9c84-f5ecd566f019@linux.intel.com> (raw)
In-Reply-To: <39047bbc-757d-4131-a25c-5a7743fd591a@intel.com>
On 9/24/2026 2:36 PM, Xiaoyao Li wrote:
Thanks a lot for the detailed review!
> On 9/17/2026 3:25 PM, Binbin Wu wrote:
>> Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable
>> CPUID feature bits that KVM supports, and build the masks during TDX
>> hardware setup via tdx_initialize_cpu_cfg_caps().
>>
>> The TDX module reports the CPUID bits that the VMM can directly configure
>> for a TD, but KVM cannot blindly expose all reported bits to userspace.
>> Certain features imply additional architectural state, e.g. one or more
>> MSRs, that KVM must explicitly manage across host/guest transitions to
>> prevent host state corruption. The existing hardcoded denylist cannot
>> account for new host state clobbering features introduced by future TDX
>> modules.
>>
>> Except for a few fixed-1 bits required for basic TDX support, host state
>> clobbering features are either directly configurable or gated by TD
>> ATTRIBUTES/XFAM, which KVM already validates. Tracking only the directly
>> configurable bits to build an allowlist is therefore sufficient.
>>
>> Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can
>> be built with similar feature-name based initializers. Directly
>> configurable non-feature bits will be handled separately.
>>
>> The allowlist is prepared to be consumed by later patches to filter
>> KVM_TDX_CAPABILITIES and to reject unsupported CPUID input to
>> KVM_TDX_INIT_VM, so that newly introduced TDX directly configurable CPUID
>> feature bits stay hidden from userspace until KVM explicitly opts in.
>>
>> By default, intersect the allowlist with kvm_cpu_caps[] via TDX_CFG_F().
>> Requiring support for non-TDX VMs avoids committing to TDX-specific
>> behavior before general KVM support is established, and respects KVM's
>> logic around disabling certain features, since the reasons for disabling
>> them could apply to TDX as well. Allow exceptions through
>> TDX_CFG_EXTRA_F() only with sufficient justification.
>
> One name just crosses my mind. How about TDX_CFG_INDEP_F()? where INDEP stands
> for independent.
>
> - TDX_CFG_F() for features that are dependent on kvm_cpu_caps[]
> - TDX_CFG_INDEP_F() for features not dependent on kvm_cpu_caps[]
I am not sure if it's better.
But if people prefer independent to extra, I would use TDX_CFG_INDEPENDENT_F()
instead.
>
>> Add comments as placeholders for HLE, RTM, WAITPKG and FRED, which KVM
>> doesn't support for TDX yet.
>
> 1. why add comments for them? Because implementation wise, they are possible to
> be supported by KVM but just current KVM doesn't have code to support them yet?
>
> If so, are the others the features that are concluded to be impossible to be
> supported by KVM?
>
> 2. We don't need to talk about FRED, which is even not suppported for normal VMs
> yet. Unless we are trying to enable FRED for TDs specifically.
OK, I will drop them to avoid confusion.
>
>> Allow MWAIT, XTPR, and HT through TDX_CFG_EXTRA_F(), as these bits are
>> not advertised in kvm_cpu_caps[]. The remaining directly configurable
>> feature bits outside kvm_cpu_caps[] are left out of the allowlist:
>>
>> - Features forced to zero when #VE is reduced, or lacking KVM support
>> for the associated MSRs: EST, TM2, SDBG, DCA, ACPI, ACC (TM), RDT_A,
>> RDT_M, TME, PCONFIG, and CORE_CAPABILITIES. Handle CORE_CAPABILITIES
>> in a subsequent patch.
>
> why not split into two category? Because they are overlapped?
>
> - Features forced to zero when #VE is reduced
> - Features lack KVM support
>
Yes, there is overlap between them.
If split into two categories, it would be:
- Features forced to zero when #VE is reduced, which also lack KVM support
- Features unrelated to #VE reduction that lack KVM support
Because the changelog was already getting quite long, I combined them into a
single category.
> As for "features forced to zero when #VE is reduced", do you mean when TDX
> module supports "#VE reduction" feature, the features becomes fixed0? or when
> guest enables "reduce #VE", the guest see a 0 value even though the feature is
> still configurable to host userspace and host userspace configure it to 1?
#VE is reduced means the guest reduced the related #VE, i.e. TDCS.TD_CTRL.REDUCE_VE
is 1 and the related bit in TDCS.FEATURE_PARAVIRT_CTRL is 0.
>
>> - Features tied to IA32_MISC_ENABLE bits that a TD cannot set when
>> TDCS.TD_CTLS.REDUCE_VE is set: CID and PBE.
>
> What is CID?
X86_FEATURE_CID, Context ID, more specifically, L1 CONTEXT ID
> and Is PBE related to any bit of IA32_MISC_ENABLE? Did you mean PEBS instead?
X86_FEATURE_PBE, Pending Break Enable.
Since kernel already has X86_FEATURE_XXX for them, so I use them directly without
other description.
>
> I don't understand what this category. It talks about something when
> TDCS.TD_CTLS.REDUCE_VE is set, which depends on guest TD's behavior. So what if
> the guest TD doesn't set TDCS.TD_CTLS.REDUCE_VE?
If the guest set TDCS.TD_CTLS.REDUCE_VE, write 1 to the the related bit trigger #GP.
If the guest doesn't set TDCS.TD_CTLS.REDUCE_VE, the value is passed to KVM, and
KVM will save the value without do corresponding virtualization.
Considering TDCS.TD_CTLS.REDUCE_VE is set by the Linux guest by default, it should
not allow these features.
>
>> - Features that can clobber host state and lack KVM support for TDX:
>> FRED.
>>
>> - Unsupported features: PREFETCHWT1 (Xeon Phi only), PSN (absent from
>> TDX-capable CPUs), AMX-TRANSPOSE (never implemented on an Intel
>> platform), and RAO_INT (defined only for future processors).
>
> PREFETCHWT1 (if the name is correct) is not supported by the kernel right? For
> such unknown features, it's straightforward to not support it.
>
> AMX-TRANSPOSE has been turned into reserved, and kernel never supports it. I
> think we don't need to mention it at all.
The guest could be other OS, not just Linux.
>
> RAO_INT is fixed 0 for TDX 1.5?
I was doing this patch series based on the latest CSV version of the CPUID
virtualization document from "Intel TDX Module ABI Definitions" updated in June 2026.
https://cdrdv2.intel.com/v1/dl/getContent/795381
>
> At last, I think we can just classify all of the above as 3 types:
>
> - Features that are directly configurable by TDX and in kvm_cpu_caps[].
> use TDX_CFG_F()
>
> - Features that are directly configurable by TDX and KVM actually supports them
> or KVM can support them but not in kvm_cpu_caps[]. Use TDX_CFG_EXTRA_F().
> For example, MWAIT, XTPR, and HT.
>
> - Features that are directly configurable by TDX but KVM or kenrel doesn't
> support them.
> For example, RDT_A, RDT_M, PREFETCHWT1, PSN,
Because excluding some bits changes the ABI, I listed all these directly configurable
features out of kvm_cpu_caps[] to explain why we must exclude some bits.
next prev parent reply other threads:[~2026-09-24 7:49 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 7:25 [PATCH v4 0/4] KVM: TDX: Validate directly configurable CPUID bits Binbin Wu
2026-09-17 7:25 ` [PATCH v4 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM Binbin Wu
2026-09-23 0:01 ` Edgecombe, Rick P
2026-09-23 0:14 ` Binbin Wu
2026-09-23 0:45 ` Edgecombe, Rick P
2026-09-23 0:57 ` Binbin Wu
2026-09-23 1:09 ` Edgecombe, Rick P
2026-09-23 1:17 ` Binbin Wu
2026-09-23 11:21 ` Xiaoyao Li
2026-09-23 14:25 ` Edgecombe, Rick P
2026-09-24 2:05 ` Xiaoyao Li
2026-09-24 6:36 ` Xiaoyao Li
2026-09-24 7:49 ` Binbin Wu [this message]
2026-09-24 9:30 ` Xiaoyao Li
2026-09-24 11:41 ` Binbin Wu
2026-09-17 7:25 ` [PATCH v4 2/4] KVM: TDX: Report CORE_CAPABILITIES as configurable Binbin Wu
2026-09-22 21:11 ` Edgecombe, Rick P
2026-09-23 0:03 ` Binbin Wu
2026-09-24 6:44 ` Xiaoyao Li
2026-09-17 7:25 ` [PATCH v4 3/4] KVM: TDX: Filter configurable CPUID bits Binbin Wu
2026-09-23 0:16 ` Edgecombe, Rick P
2026-09-23 0:28 ` Binbin Wu
2026-09-23 0:34 ` Edgecombe, Rick P
2026-09-17 7:25 ` [PATCH v4 4/4] KVM: TDX: Validate userspace CPUID input for KVM_TDX_INIT_VM Binbin Wu
2026-09-23 0:16 ` Edgecombe, Rick P
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=062f5365-7eb5-4e04-9c84-f5ecd566f019@linux.intel.com \
--to=binbin.wu@linux.intel.com \
--cc=andrew.cooper3@citrix.com \
--cc=chao.gao@intel.com \
--cc=dave.hansen@linux.intel.com \
--cc=dedekind1@gmail.com \
--cc=kas@kernel.org \
--cc=kishen.maloor@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=nik.borisov@suse.com \
--cc=pbonzini@redhat.com \
--cc=rick.p.edgecombe@intel.com \
--cc=seanjc@google.com \
--cc=tony.lindgren@linux.intel.com \
--cc=xiaoyao.li@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®