[PATCH v4 17/17] KVM: TDX: Turn on PG_LEVEL_2M
From: Yan Zhao
Date: Mon Sep 28 2026 - 05:25:54 EST
Introduce a module parameter named "tdx_huge_page" for kvm-intel.ko and
enable TDX huge pages if the module parameter is true and if:
(1) KVM supports gmem in-place conversions,
(2) the TDX module supports uninterruptible demote, and
(3) the TDX module supports DPAMT.
Condition (1) forces TDX huge pages to be the first user of gmem in-place
conversion. Condition (2) is required because the current splitting
implementation depends on the TDX module's non-interruptible demote
feature. Condition (3) simplifies the DEMOTE SEAMCALL's implementation.
When TDX huge page support is enabled, report to KVM MMU that the maximum
allowed mapping level for private memory is 2MB when the TD is RUNNABLE,
while forcing it to 4KB during TD build time. This is because KVM can only
create mappings up to 2MB level in the S-EPT via the TDH_MEM_PAGE_AUG
SEAMCALL, while TDH_MEM_PAGE_ADD mandates the mapping level must be 4KB.
1GB mappings in the S-EPT are only possible via promotion, which is not yet
supported due to the complexity incurred by the TDX module's rules for huge
page promotion.
Signed-off-by: Xiaoyao Li <xiaoyao.li@xxxxxxxxx>
Signed-off-by: Isaku Yamahata <isaku.yamahata@xxxxxxxxx>
Signed-off-by: Sean Christopherson <seanjc@xxxxxxxxxx>
Signed-off-by: Yan Zhao <yan.y.zhao@xxxxxxxxx>
---
v4:
- Disallowed huge page if gmem_in_place_conversion is false. (Sean)
- Updated map part due to MMU refactor. (Sean)
- Disallowed huge page if DPAMT is not enabled. (Rick)
v3:
- Introduce the module param enable_tdx_huge_page and disable to toggle TDX
huge page support.
- Disable TDX huge page if TDX module does not support
TDX_FEATURES0_ENHANCE_DEMOTE_INTERRUPTIBILITY. (Kai).
- Explain why not allow 2M before TD is RUNNABLE in patch log.(Kai)
- Add comment to explain the relationship between returning PG_LEVEL_2M
and guest accept level. (Kai)
- Dropped some KVM_BUG_ON()s due to rebasing. Updated KVM_BUG_ON()s on
mapping levels to take into account of enable_tdx_huge_page.
RFC v2:
- Merged RFC v1's patch 4 (forcing PG_LEVEL_4K before TD runnable) with
patch 9 (allowing PG_LEVEL_2M after TD runnable).
---
arch/x86/kvm/vmx/tdx.c | 41 ++++++++++++++++++++++++++++++++++-------
1 file changed, 34 insertions(+), 7 deletions(-)
diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c
index 0ecfc17e6235..ee17e50bf439 100644
--- a/arch/x86/kvm/vmx/tdx.c
+++ b/arch/x86/kvm/vmx/tdx.c
@@ -57,6 +57,9 @@
bool enable_tdx __ro_after_init;
module_param_named(tdx, enable_tdx, bool, 0444);
+static bool __read_mostly enable_tdx_huge_page = true;
+module_param_named(tdx_huge_page, enable_tdx_huge_page, bool, 0444);
+
static const struct tdx_sys_info *tdx_sysinfo;
void tdh_vp_rd_failed(struct vcpu_tdx *tdx, char *uclass, u32 field, u64 err)
@@ -1786,8 +1789,9 @@ static int tdx_sept_map_leaf_spte(struct kvm *kvm, gfn_t gfn, enum pg_level leve
if (KVM_BUG_ON(!vcpu, kvm))
return -EIO;
- /* TODO: handle large pages. */
- if (KVM_BUG_ON(level != PG_LEVEL_4K, kvm))
+ /* TODO: Support hugepages when building the initial TD image. */
+ if (KVM_BUG_ON(level != PG_LEVEL_4K &&
+ to_kvm_tdx(kvm)->state != TD_STATE_RUNNABLE, kvm))
return -EIO;
WARN_ON_ONCE((new_spte & VMX_EPT_RWX_MASK) != VMX_EPT_RWX_MASK);
@@ -1883,10 +1887,6 @@ static int tdx_sept_remove_leaf_spte(struct kvm *kvm, gfn_t gfn,
if (KVM_BUG_ON(!is_hkid_assigned(to_kvm_tdx(kvm)), kvm))
return -EIO;
- /* TODO: handle large pages. */
- if (KVM_BUG_ON(level != PG_LEVEL_4K, kvm))
- return -EIO;
-
err = tdh_do_no_vcpus(tdh_mem_range_block, kvm, &kvm_tdx->td, gpa,
level, &entry, &level_state);
if (TDX_BUG_ON_2(err, TDH_MEM_RANGE_BLOCK, entry, level_state, kvm))
@@ -3665,12 +3665,34 @@ int tdx_vcpu_ioctl(struct kvm_vcpu *vcpu, void __user *argp)
return ret;
}
+/*
+ * For private pages:
+ *
+ * Force KVM to map at 4KB level when !enable_tdx_huge_page (e.g., due to
+ * incompatible TDX module) or before TD state is RUNNABLE.
+ *
+ * Always allow KVM to map at 2MB level in other cases, though KVM may still map
+ * the page at 4KB (i.e., passing in PG_LEVEL_4K to AUG) due to
+ * (1) the backend folio is 4KB,
+ * (2) disallow_lpage restrictions:
+ * - mixed private/shared pages in the 2MB range
+ * - level misalignment due to slot base_gfn, slot size, and ugfn
+ * - guest_inhibit bit set due to guest's 4KB accept level
+ * (3) page merging is disallowed (e.g., when part of a 2MB range has been
+ * mapped at 4KB level during TD build time).
+ */
int tdx_gmem_max_mapping_level(struct kvm *kvm, kvm_pfn_t pfn, bool is_private)
{
if (!is_private)
return 0;
- return PG_LEVEL_4K;
+ if (!enable_tdx_huge_page)
+ return PG_LEVEL_4K;
+
+ if (unlikely(to_kvm_tdx(kvm)->state != TD_STATE_RUNNABLE))
+ return PG_LEVEL_4K;
+
+ return PG_LEVEL_2M;
}
void tdx_hardware_unsetup(void)
@@ -3749,6 +3771,11 @@ static int __init __tdx_hardware_setup(void)
if (misc_cg_set_capacity(MISC_CG_RES_TDX, tdx_get_nr_guest_keyids()))
return -EINVAL;
+ if (enable_tdx_huge_page && (!gmem_in_place_conversion ||
+ !tdx_huge_page_demote_uninterruptible(tdx_sysinfo) ||
+ !tdx_supports_dynamic_pamt(tdx_sysinfo)))
+ enable_tdx_huge_page = false;
+
return 0;
}
--
2.43.2