Re: [PATCH v4] mm/hugetlb: fix overbroad MMU notifiers for unshared PMDs

From: Muchun Song

Date: Sun Sep 27 2026 - 23:09:44 EST




> On Sep 28, 2026, at 10:47, Li Zhe <lizhe.67@xxxxxxxxxxxxx> wrote:
>
> Hugetlb currently expands MMU notifier ranges to PUD boundaries whenever
> PMD sharing is possible. That is only needed when huge_pmd_unshare()
> actually detaches a shared PMD page table, because clearing the PUD
> invalidates the whole PUD-sized virtual address range.
>
> For hugetlbfs hole punch, and similarly for other hugetlb unmap paths,
> a shared mapping can pass the "PMD sharing is possible" range test in
> adjust_range_if_pmd_sharing_possible() even when the hugetlbfs file does
> not currently have any shared PMD page tables. KVM then receives a 1G
> invalidation for a 2M operation and zaps unrelated secondary mappings,
> so the guest has to fault them back in.
>
> Avoid this by tracking active PMD-sharing attachments per hugetlbfs
> inode. The count is incremented only after huge_pmd_share() successfully
> installs a shared PMD table, and decremented when __huge_pmd_unshare()
> actually detaches one. Since huge_pmd_share() can run concurrently under
> i_mmap_lock_read(), use a 64-bit atomic counter. A zero count is used to
> skip the conservative notifier range expansion only after excluding
> concurrent PMD sharing with the mapping write lock.
>
> On a Redis-in-VM workload that punches cold 2M hugetlb pages, this patch
> improves P99 QPS stability while punching pages, reducing the QPS
> degradation ratio from 7.09% to 1.45%.
>
> Reported-by: aiqi.i7 <aiqi.i7@xxxxxxxxxxxxx>
> Signed-off-by: Li Zhe <lizhe.67@xxxxxxxxxxxxx>

Acked-by: Muchun Song <muchun.song@xxxxxxxxx>