Re: [PATCH v3 1/4] KVM: TDX: Track configurable CPUID bits allowed by KVM
From: Binbin Wu
Date: Thu Sep 10 2026 - 20:55:47 EST
On 9/11/2026 5:29 AM, Edgecombe, Rick P wrote:
> On Thu, 2026-09-10 at 10:39 +0800, Binbin Wu wrote:
>>> I think this actually surfaces another problem with TD-first enabling.
>>> KVM_TDX_CAPABILITIES only returns the directly configurable bits. Then
>>> recall, KVM_TDX_GET_CPUID returns the actual TDX module's view of CPUID bits
>>> to userspace. Then userspace calls KVM_SET_CPUID to actually put them on
>>> KVM's vcpu so they can match between Qemu, KVM and TDX
>>
>> That brings up a point..
>>
>> Today, vcpu->arch.cpu_caps[] is capped by kvm_cpu_caps[] (plus a few special
>> cases). As mentioned in the cover letter, this patch series doesn't enforce
>> consistency between KVM's view and the guest's view of vCPU capabilities
>> because KVM doesn't currently use its own view to make decisions for TDs (e.g.
>> saving/restoring feature-related MSRs).
>
> Not sure if I'm missing your point here. I don't think we ever want to have KVM
> enforce consistency between KVM's view and guests. We just need to provide
> enough info to userspace such that it can make them consistent.
Consistency check prevents malicious userspace VMMs from lying to the KVM about
some host state clobbering if KVM uses guest_cpu_cap_has() to management the state
for TDX in the future.
>
>>
>> However, if KVM starts making decisions for TDX based on vcpu-
>>> arch.cpu_caps[], intersecting userspace input with kvm_cpu_caps[] will not
>> work for TDX.
>
> vcpu->arch.cpu_caps are actually already consulted for TDX. I remember seeing a
> bunch of the the guest cpuid feature checks during the base enabling, probably
> working on this problem. Let me what we have today.
>
> From a Linux guest boot, guest_cpu_cap_has() returns true for:
> xsave
> smep
> smap
> fsgsbase
> pku
> la57
> umip
> vmx
> pcid
> lam
> unknown
> ibt
> x2apic
>
> Since we share code with normal VMs (and manage shared EPT in KVM), some checks
> are going to happen. If there is some new feature foo we enable for TDX. And
> later KVM adds new logic around vcpu->arch.cpu_caps for it, then there is a
> small risk of being pinned down when we want to add new guest_cpu_cap_has()
> logic for normal VMs. Since we already are hitting these checks for TDX, the
> general case is not theoretical.
... Yes, this was my concern.
If in the future KVM adds new guest_cpu_cap_has() for a host state clobbering
feature and it is used TDX, it would cause problem if there is a mismatch between
guest_cpu_cap_has() and the real value exposed to the TD.
>
>> I think this is probably needed in the future? If so, allowing features
>> outside of kvm_cpu_caps[] for TDX means
>
> Yea, I think allowing TDX features outside of kvm_cpu_caps is for special cases.
> And filtering like you have is good.
>
>> we will need TDX-specific handling to construct KVM's view of vCPU
>> capabilities. That likely implies tracking all known/supported TDX features,
>> which is doable, but it will make the allow list bigger.
>
> In this thread we have been talking about what "normal VMs" support, but in the
> code and uAPI it really is about what KVM supports. If we let TDX use a feature
> that *KVM* doesn't support, it is the risky zone.
>
> I say we punt on this. Let's remember it's dicey and if we find TDX feature
> enabling is being blocked all the time by normal VM enabling, we can work on a
> solution. Does anyone see any big risk of this being harder later than it is
> today?
>
> I prefer to at least start filtering ASAP.
+1