Re: [PATCH v8 13/30] mm: handle PMD swap entries in fork path
From: Zi Yan
Date: Fri Oct 09 2026 - 13:40:31 EST
On Fri Oct 2, 2026 at 5:52 AM EDT, Usama Arif wrote:
> copy_huge_pmd() only knows about migration and device-private PMDs, so a
> PMD swap entry would fall through to the present-PMD path and fork() would
> duplicate it without taking a reference on the slots it points at.
>
> Copy it the way copy_nonpresent_pte() copies a PTE swap entry: duplicate
> the swap references, clear the exclusive marker on the source, put the
> destination mm on mmlist, and account the child's slots to MM_SWAPENTS.
>
> The GFP_ATOMIC extend-table allocation inside the dup can fail. Report that
> as -EIO and let copy_pmd_range() retry with GFP_KERNEL, as
> copy_nonpresent_pte() and copy_pte_range() already do for a PTE swap entry.
> copy_huge_pmd() hands the entry back so the caller knows which range to
> allocate for.
>
> Only -ENOMEM is reported that way. The other failures mean the entry itself
> is bad, and swap_retry_table_alloc_nr() returns 0 for those, so collapsing
> them into -EIO as the PTE path does would spin in the caller's retry rather
> than failing the fork.
>
> While here, move the mm counter update into each entry-type arm, as the PTE
> version does, so the swap arm can account MM_SWAPENTS instead of
> MM_ANONPAGES.
>
> Signed-off-by: Usama Arif <usama.arif@xxxxxxxxx>
> ---
> include/linux/huge_mm.h | 3 +-
> mm/huge_memory.c | 62 ++++++++++++++++++++++++++++++-----------
> mm/memory.c | 12 +++++++-
> 3 files changed, 58 insertions(+), 19 deletions(-)
>
> diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h
> index 8205e83f27771..7aa63d982af07 100644
> --- a/include/linux/huge_mm.h
> +++ b/include/linux/huge_mm.h
> @@ -10,7 +10,8 @@
> vm_fault_t do_huge_pmd_anonymous_page(struct vm_fault *vmf);
> int copy_huge_pmd(struct mm_struct *dst_mm, struct mm_struct *src_mm,
> pmd_t *dst_pmd, pmd_t *src_pmd, unsigned long addr,
> - struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma);
> + struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma,
> + softleaf_t *entryp);
> bool huge_pmd_set_accessed(struct vm_fault *vmf);
> int copy_huge_pud(struct mm_struct *dst_mm, struct mm_struct *src_mm,
> pud_t *dst_pud, pud_t *src_pud, unsigned long addr,
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index 24d116ae1fc30..80d18ca972ecf 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -1894,7 +1894,7 @@ bool touch_pmd(struct vm_area_struct *vma, unsigned long addr,
> return false;
> }
>
> -static void copy_huge_non_present_pmd(
> +static int copy_huge_non_present_pmd(
> struct mm_struct *dst_mm, struct mm_struct *src_mm,
> pmd_t *dst_pmd, pmd_t *src_pmd, unsigned long addr,
> struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma,
> @@ -1902,18 +1902,41 @@ static void copy_huge_non_present_pmd(
> {
> softleaf_t entry = softleaf_from_pmd(pmd);
> struct folio *src_folio;
> + int err;
>
> VM_WARN_ON_ONCE(!pmd_is_valid_softleaf(pmd));
>
> - if (softleaf_is_migration_write(entry) ||
> - softleaf_is_migration_read_exclusive(entry)) {
> - entry = make_readable_migration_entry(swp_offset(entry));
> - pmd = softleaf_to_pmd(entry);
> - if (pmd_swp_soft_dirty(*src_pmd))
> - pmd = pmd_swp_mksoft_dirty(pmd);
> - if (pmd_swp_uffd(*src_pmd))
> - pmd = pmd_swp_mkuffd(pmd);
> - set_pmd_at(src_mm, addr, src_pmd, pmd);
> + if (softleaf_is_swap(entry)) {
> + /*
> + * A PMD swap entry only exists under CONFIG_THP_SWAP, where
> + * SWAPFILE_CLUSTER == HPAGE_PMD_NR, and it is cluster aligned,
> + * so these HPAGE_PMD_NR slots are exactly one cluster - which
> + * is what swap_dup_entries_direct() requires.
> + */
> + err = swap_dup_entries_direct(entry, HPAGE_PMD_NR);
> + if (err)
> + /* Only -ENOMEM is worth a GFP_KERNEL retry. */
Is it better to say "if swap_dup_entries_direct() cannot allocate
memory, return -EIO so that the caller can retry with GFP_KERNEL. Fail
the fork on any other error"? It clarifies why EIO is needed.
> + return err == -ENOMEM ? -EIO : -ENOMEM;
> +
Otherwise, LGTM.
Reviewed-by: Zi Yan <ziy@xxxxxxxxxx>
--
Best Regards,
Yan, Zi