Re: [PATCH v3 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM

From: Binbin Wu

Date: Tue Sep 01 2026 - 20:35:26 EST




On 9/1/2026 10:35 PM, Xiaoyao Li wrote:
> On 8/27/2026 11:18 AM, Binbin Wu wrote:
>> Add tdx_cpu_cfg_caps[] to track the subset of TDX directly configurable
>> CPUID feature bits that KVM supports, and build the masks during TDX
>> hardware setup via tdx_initialize_cpu_cfg_caps().
>>
>> The TDX module reports the CPUID bits that the VMM can directly configure
>> for a TD, but KVM cannot blindly expose all reported bits to userspace.
>> Certain features imply additional architectural state, e.g. one or more
>> MSRs, that KVM must explicitly manage across host/guest transitions to
>> prevent host state corruption.
>>
>> Today KVM relies on a hardcoded denylist, i.e. it clears a few known
>> problematic bits, e.g. TSX and WAITPKG, and passes everything else through.
>> A denylist is fundamentally fragile while an allowlist inverts the default,
>> i.e. unknown configurable bits are hidden and not allowed to be enabled
>> until KVM explicitly opts in.
>>
>> Except for a few fixed-1 bits required for basic TDX support, host state
>> clobbering features are either directly configurable or gated by TD
>> ATTRIBUTES/XFAM. 
>
>> Tracking only the directly configurable feature bits is
>> therefore sufficient to serve the purpose while keeping the code footprint
>> small.
>
> I'm not clear how it is therefore sufficient. We at least need to explain that ATTRIBUTS/XFAM are validated separately by KVM already?

I was trying to say this by "gated by TD ATTRIBUTES/XFAM", I will describe it
more clearly.

>> Organize tdx_cpu_cfg_caps[] following kvm_cpu_caps[] so that the masks can
>> be built with the similar feature-name based initializers.  CPUID registers
>> that hold directly configurable non-feature (multi-bit) fields are handled
>> separately.
>>
>> The allowlist is consumed by later patches to filter KVM_TDX_CAPABILITIES
>> and to reject unsupported CPUID input to KVM_TDX_INIT_VM, so that newly
>> introduced TDX directly configurable CPUID feature bits stay hidden from
>> userspace until KVM explicitly opts in.
>>
>> Add comments as placeholders for HLE, RTM and WAITPKG, which KVM doesn't
>> support for TDX yet.
>>
>> Signed-off-by: Binbin Wu <binbin.wu@xxxxxxxxxxxxxxx>
>> ---
>> v3:
>> - Drop the new data structure in v2 and only track feature bits by
>>    following the organization of kvm_cpu_caps[], handle non-feature
>>    bits separately. (Sean)
>> - Use two versions of macros (TDX_CFG_F() VS. TDX_CFG_EXTRA_F()) to
>>    distinguish whether a supported TDX configurable CPUID bit should be
>>    checked against KVM's common cpu capabilities.
>> - Add AMX_COMPLEX since it has been defined in the CPUID virtualization doc.
>> ---
>>   arch/x86/kvm/vmx/tdx.c | 145 +++++++++++++++++++++++++++++++++++++++++
>>   1 file changed, 145 insertions(+)
>>
>> diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
>> index b272c20586a7..d4a3a42cfd9d 100644
>> --- a/arch/x86/kvm/vmx/tdx.c
>> +++ b/arch/x86/kvm/vmx/tdx.c
>> @@ -52,6 +52,149 @@
>>       __TDX_BUG_ON(__err, #__fn, __kvm, ", " #a1 " 0x%llx, " #a2 ", 0x%llx, " #a3 " 0x%llx", \
>>                a1, a2, a3)
>>   +static u32 tdx_cpu_cfg_caps[NR_KVM_CPU_CAPS] __ro_after_init;
>> +static_assert(ARRAY_SIZE(tdx_cpu_cfg_caps) == ARRAY_SIZE(kvm_cpu_caps));
>> +
>> +#define TDX_VALIDATE_CPU_CAP_USAGE(name)            \
>> +    BUILD_BUG_ON(__feature_leaf(X86_FEATURE_##name) !=    \
>> +             tdx_cpu_cap_init_in_progress)
>> +
>> +/* For feature bit that KVM advertised through kvm_cpu_caps[]. */
>
> I would say it
>
> For feature bit that needs to be cap'ed by kvm_cpu_caps[]
>
>> +#define TDX_CFG_F(name)                    \
>> +({                            \
>> +    TDX_VALIDATE_CPU_CAP_USAGE(name);        \
>> +    tdx_cfg_caps |= feature_bit(name);        \
>> +})
>> +
>> +/*
>> + * For feature bit KVM allows for TDX guests even though it is not advertised
>> + * through kvm_cpu_caps[], e.g. MWAIT.
>> + */
>> +#define TDX_CFG_EXTRA_F(name)                \
>
> EXTRA doesn't sound like a fit name, though

I also struggled with naming this macro and couldn't come up with a better one.
Any suggestion?

>
>> +({                            \
>> +    TDX_VALIDATE_CPU_CAP_USAGE(name);        \
>> +    tdx_cfg_extra_caps |= feature_bit(name);    \
>> +})
>> +
>> +#define tdx_cpu_cfg_cap_init(leaf, feature_initializers...)        \
>> +do {                                    \
>> +    const u32 __maybe_unused tdx_cpu_cap_init_in_progress = leaf;    \
>> +    u32 tdx_cfg_extra_caps = 0;                    \
>> +    u32 tdx_cfg_caps = 0;                        \
>> +                                    \
>> +    feature_initializers                        \
>> +    tdx_cpu_cfg_caps[leaf] = (tdx_cfg_caps & kvm_cpu_caps[leaf]) |    \
>> +                 tdx_cfg_extra_caps;            \
>> +} while (0)
>> +
>> +/*
>> + * Track only CPUID feature bits that are directly configurable by userspace.
>
> the "by userspace" is misleading. It's just the directly configurable CPUID bits reported by TDX module.
>
>> + * Features controlled by XFAM or ATTRIBUTES are excluded; userspace cannot
>> + * enable them until KVM adds support for the corresponding control.
>> + */
>
> I don't like the comments. How about somthing
>
> /*
>  * Intialize tdx_cpu_cfg_caps[], which is list of CPUID features that
>  * KVM supports for TDX. It only covers the directly configurable CPIUD
>  * bits reported by TDX module. Features controlled by XFAM and
>  * ATTRIBUTES are maintained separately.
>  */
>
Thanks, it reads better.

>> +static void __init tdx_initialize_cpu_cfg_caps(void)
>> +{
>> +    tdx_cpu_cfg_cap_init(CPUID_1_ECX,
>> +        TDX_CFG_EXTRA_F(MWAIT),
>> +        TDX_CFG_F(TSC_DEADLINE_TIMER),
>> +        TDX_CFG_F(AVX),
>> +        TDX_CFG_F(F16C),
>> +    );
>
> TDX 1.5.24 on SPR report configurable bits of CPUID_1_ECX as
> 0x31044988, which have
>
> - bit 3        MWAIT
> - bit 7        EST
> - bit 8        TM2
> - bit 11    SDBG
> - bit 14    XTPR
> - bit 18    DCA
> - bit 24    TSC_DEADLINE_TIMER
> - bit 28    AVX
> - bit 29    F16C
>
> but EST/TM2/SDBG/XTPR/DCA are not list here. I guess the reason is kvm_cpu_cap[] doesn't support it. If so, it seems to guard twice:
> 1. mentally/manually check if it a feature is supported in kvm_cpu_caps[]
>
> 2. kvm_cpu_caps guarding in tdx_cpu_cfg_cap_init().
>
> I think 1) is not necessary, we can rely on 2)

In general, if a feature is not supported by the common KVM CPU caps, I prefer not
to add it to the list to save a few lines of code, which probably is dead code,
unless people find it too confusing.
I can add a comment to clarify this.

>
> BTW, this seems also breaks the current userspace after this series.
> - Before, EST/TM2/SDBG/XTPR/DCA are allowed to be exposed to TD
> - After, they are not.

This does change the values returned by KVM_TDX_CAPABILITIES. However, my
understanding is that userspace is generally expected to only configure
features supported by both KVM and TDX, i.e. except for the features initialized
via TDX_CFG_EXTRA_F(), userspace is not expected to configure features not advertised
by kvm_cpu_caps[].
I can call this out in the changelog, and maybe also the doc for KVM_TDX_CAPABILITIES.

>
> If we cares CORE_CAPABILITIES in patch 2, why EST/TM2/SDBG/XTPR/DCA don't matter?

Because CORE_CAPABILITIES was previously defined as fixed-1 in some old spec and
the QEMU marks it as fixed1. EST/TM2/SDBG/XTPR/DCA are not the case.