Re: [PATCH v2] mm/hugetlb: fix overbroad MMU notifiers for unshared PMDs

From: Li Zhe

Date: Wed Sep 23 2026 - 03:41:50 EST


On 9/23/26 10:00 AM, Andrew Morton wrote:
> On Tue, 22 Sep 2026 17:07:49 +0800 "Li Zhe" <lizhe.67@xxxxxxxxxxxxx> wrote:
>
>> Hugetlb currently expands MMU notifier ranges to PUD boundaries whenever
>> PMD sharing is possible. That is only needed when huge_pmd_unshare()
>> actually detaches a shared PMD page table, because clearing the PUD
>> invalidates the whole PUD-sized virtual address range.
>>
>> For hugetlbfs hole punch, and similarly for other hugetlb unmap paths,
>> a shared mapping can pass the "PMD sharing is possible" range test in
>> adjust_range_if_pmd_sharing_possible() even when the hugetlbfs file has
>> never actually had any shared PMD page tables. KVM then receives a 1G
>> invalidation for a 2M operation and zaps unrelated secondary mappings,
>> so the guest has to fault them back in.
>>
>> Avoid this by remembering, per hugetlbfs inode, whether PMD sharing was
>> ever established for the file. Set the flag when huge_pmd_share()
>> successfully populates a shared PMD table. For the hugetlb unmap paths,
>> skip the conservative PUD-sized notifier expansion while the file has
>> never seen PMD sharing.
>>
>> The state is intentionally sticky and file-wide. Once PMD sharing has
>> ever happened for the file, the unmap paths keep the existing
>> conservative expansion. This avoids the no-sharing case without adding a
>> page-table walk to every unmap.
> Oh. It's not feasible to figure out when PMD sharing has ended and go
> back to never-seen-sharing state?


Yes, it is feasible.


I did consider a refcount-based approach before sending v2, but my
initial version looked more complicated than I was comfortable with.  So
I used the sticky state in v2 to keep the change simple and avoid adding
page-table walks to the unmap path.

After looking at this again, I think the refcounting can be kept
reasonably small.  The updated version uses a per-hugetlbfs-inode
counter for active PMD-sharing attachments.  The counter is incremented
only after huge_pmd_share() successfully installs a shared PMD table, and
is decremented when __huge_pmd_unshare() actually detaches one.

The share side can run concurrently under i_mmap_lock_read(), so the
counter uses atomic operations.  A zero count is used to skip the
PUD-sized notifier expansion only after excluding concurrent PMD sharing
with the mapping write lock.

Sorry for the churn; I will send v3 with this refcounting approach.

>> On a Redis-in-VM workload that punches cold 2M hugetlb pages, this patch
>> improves P99 QPS stability while punching pages, reducing the QPS
>> degradation ratio from 7.09% to 1.45%.
> How realistic is this test? IOW, how much benefit can people expect to
> see in real-world usage?


This was measured on a KVM workload using hugetlbfs-backed guest memory.

The workload performs policy-driven 2M hugetlbfs unmaps while the guest
keeps running.  The important part is not the policy itself, but that a
2M hugetlbfs unmap can currently be reported to KVM as a 1G invalidation
when PMD sharing is only possible, but not actually active.

So the benefit is expected for workloads with secondary MMU users, such
as KVM, that see sub-PUD hugetlbfs unmaps on files without active PMD
sharing.  In that case, avoiding the unnecessary PUD-sized notifier
prevents KVM from zapping unrelated secondary mappings in the same PUD
range.

For workloads without secondary MMU mappings, without sub-PUD hugetlbfs
unmaps, or with active PMD sharing, the patch should mostly preserve the
existing behavior.

>
>> --- a/fs/hugetlbfs/inode.c
>> +++ b/fs/hugetlbfs/inode.c
>> @@ -921,6 +921,9 @@ static struct inode *hugetlbfs_get_inode(struct super_block *sb,
>> simple_inode_init_ts(inode);
>> info->resv_map = resv_map;
>> info->seals = F_SEAL_SEAL;
>> +#ifdef CONFIG_HUGETLB_PMD_PAGE_TABLE_SHARING
>> + info->pmd_sharing_seen = false;
>> +#endif
> This could use a slightly modified hugetlbfs_set_pmd_sharing_seen() and
> remove the ifdefs.
>
> hugetlbfs_set_pmd_sharing_seen(inode, false);


Thanks for pointing this out.  I will fix it in v3.

Thanks,
Zhe

>
> Not very important.