Re: [PATCH] mm/huge_memory: let special huge VMAs bypass the THP policy check
From: David Hildenbrand (Arm)
Date: Thu Aug 06 2026 - 10:26:31 EST
On 8/5/26 07:55, Cédric Le Goater wrote:
> From: Cedric Le Goater <clg@xxxxxxxxxx>
>
> The global THP sysfs policy (transparent_hugepage=never/madvise/always)
> gates the huge fault dispatch path in __thp_vma_allowable_orders() for
> all non-anonymous VMAs, including PFN-mapped device BARs (VM_PFNMAP).
>
> DAX VMAs already bypass this check via an early return:
>
> if (vma_is_dax(vma))
> return in_pf ? orders : 0;
>
> But "special huge" VMAs -- identified by vma_is_special_huge() -- do not
> get this early return, even though they share the same fundamental
> property: they map physical addresses directly into page tables and
> involve no memory allocation, no compaction, no splitting, and no
> reclaim. The THP policy has no meaningful effect on them.
>
> This matters for VFIO PCI passthrough of large-BAR devices such as
> NVIDIA H200 NVL GPUs (256 GB BAR each). The VFIO driver registers a
> .huge_fault handler (vfio_pci_mmap_huge_fault) that dispatches to
> vmf_insert_pfn_pmd/pud, and QEMU's vfio_region_mmap() aligns the BAR
> mappings for huge page table entries. Both prerequisites are met, but
> with THP=never or THP=madvise, __thp_vma_allowable_orders() returns 0
> before reaching the "trust huge_fault handlers" code.
>
> The result: each 256 GB BAR is mapped at 4 KiB granularity -- 67 million
> page faults per GPU instead of a few thousand PMD/PUD faults. On hosts
> with 8 GPUs (2 TB of BAR space), this causes VM boot times to degrade
> severely, with 99.98% of CPU time spent in the VFIO BAR mapping path.
>
> Configurations that trigger this:
> - transparent_hugepage=never on the kernel command line
> - The tuned cpu-partitioning profile (inherits network-latency, which
> sets transparent_hugepages=never via sysfs)
> - transparent_hugepage=madvise (the RHEL default), since VFIO VMAs
> lack VM_HUGEPAGE and QEMU does not call madvise(MADV_HUGEPAGE) on
> BAR mmap regions
>
> Extend the existing DAX early return to also cover vma_is_special_huge()
> VMAs. This is consistent with how vma_is_special_huge() is already
> treated for supported_orders (grouped with DAX). The mm/Kconfig TODO
> comment "Allow to be enabled without THP" also acknowledges this
> coupling is wrong.
>
> Cc: Peter Xu <peterx@xxxxxxxxxx>
> Cc: Andrew Morton <akpm@xxxxxxxxxxxxxxxxxxxx>
> Cc: Lorenzo Stoakes <ljs@xxxxxxxxxx>
> Cc: David Hildenbrand <david@xxxxxxxxxx>
> Cc: Alex Williamson <alex.williamson@xxxxxxxxxx>
> Cc: Jason Gunthorpe <jgg@xxxxxxxxxx>
> Cc: Zi Yan <ziy@xxxxxxxxxx>
> Fixes: 5dd40721f147 ("mm: allow THP orders for PFNMAPs")
> Cc: stable@xxxxxxxxxxxxxxx
> Assisted-by: Claude:claude-opus-4
> Signed-off-by: Cedric Le Goater <clg@xxxxxxxxxx>
> ---
> mm/huge_memory.c | 8 ++++++--
> 1 file changed, 6 insertions(+), 2 deletions(-)
>
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index 58cabe6af33d031e48250e21db51506bc46c97b2..6dfef5500a054f09f9ece6df8bf7a0194624350f 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -139,8 +139,12 @@ unsigned long __thp_vma_allowable_orders(struct vm_area_struct *vma,
> if (thp_disabled_by_hw() || vma_thp_disabled(vma, vm_flags, forced_collapse))
> return 0;
>
> - /* khugepaged doesn't collapse DAX vma, but page fault is fine. */
> - if (vma_is_dax(vma))
> + /*
> + * khugepaged doesn't collapse DAX or special huge VMAs, but page
> + * fault is fine. These map physical addresses directly — the THP
> + * policy is irrelevant for them.
emdash in a code comment?
Then I spot
Assisted-by: Claude:claude-opus-4
and really have to shake my head.
--
Cheers,
David