Re: [PATCH v8 10/30] mm: make PMD migration-entry splitting explicit
From: Zi Yan
Date: Wed Oct 07 2026 - 22:54:11 EST
On Fri Oct 2, 2026 at 5:52 AM EDT, Usama Arif wrote:
> __split_huge_pmd() and friends take a "freeze" boolean that every caller
> has to pass and almost every caller passes as false. The name says nothing
> about what it selects, and the one thing it does select - PTE migration
> entries instead of PTE mappings - is only ever wanted by the rmap migration
> path.
>
> Rename it to to_migration_entries, keep it private to mm/huge_memory.c,
> and add split_pmd_to_migration_entries() for try_to_migrate_one(), the only
> caller that wants it.
>
> migrate_vma_split_unmapped_folio() also passed freeze=true, but only ever
> runs on a PMD that is already a migration entry, which the generic helper
Is it always a PMD migration entry? Can a concurrent MADV_DONTNEED zap
the PMD migration entry or split it with a partial zap? Although it
does not affect the correctness of the patch.
If that concurrent zap is possible, the existing code can leak a ref,
causing the folio to be not freed, since
split_huge_pmd_address(freeze=true) does not drop the folio refcount
when the PMD is zapped or becomes a PTE page.
We might want a separate fix for it.
Actually, Mika Penttilä's "migrate on fault for device pages"[1] also
talks about the race.
[1] https://lore.kernel.org/r/20260924065313.899730-1-mpenttil@xxxxxxxxxx
> expands into PTE migration entries either way. Its folio_get() only existed
> to balance the put_page() that freeze=true performs, so both go.
>
> split_pmd_to_migration_entries() is only ever handed a present or
> device-private PMD, so assert that in __split_huge_pmd_locked() instead of
> silently skipping anything else.
>
> Other than that assertion, no functional change intended.
>
> Suggested-by: David Hildenbrand (Arm) <david@xxxxxxxxxx>
> Signed-off-by: Usama Arif <usama.arif@xxxxxxxxx>
> Reviewed-by: Kiryl Shutsemau (Meta) <kas@xxxxxxxxxx>
> ---
> include/linux/huge_mm.h | 21 ++++++------
> mm/huge_memory.c | 74 ++++++++++++++++++++++++++---------------
> mm/memory.c | 4 +--
> mm/migrate_device.c | 7 +---
> mm/mprotect.c | 2 +-
> mm/rmap.c | 8 ++---
> 6 files changed, 65 insertions(+), 51 deletions(-)
>
<snip>
> diff --git a/mm/migrate_device.c b/mm/migrate_device.c
> index 0c437004329d9..4a0b61d50d222 100644
> --- a/mm/migrate_device.c
> +++ b/mm/migrate_device.c
> @@ -918,12 +918,7 @@ static int migrate_vma_split_unmapped_folio(struct migrate_vma *migrate,
> unsigned long flags;
> int ret = 0;
>
> - /*
> - * take a reference, since split_huge_pmd_address() with freeze = true
> - * drops a reference at the end.
> - */
> - folio_get(folio);
> - split_huge_pmd_address(migrate->vma, addr, true);
> + split_huge_pmd_address(migrate->vma, addr);
> ret = folio_split_unmapped(folio, 0);
> if (ret)
> return ret;
This part can be a separate fix for the refcount leak.
Otherwise, LGTM.
--
Best Regards,
Yan, Zi