Re: [PATCH v8 26/30] mm: handle PMD swap entries in UFFDIO_MOVE
From: Lance Yang
Date: Sun Oct 04 2026 - 04:19:34 EST
Hey Usama,
On Fri, Oct 02, 2026 at 02:52:40AM -0700, Usama Arif wrote:
[...]
>+static int move_swap_pmd(struct mm_struct *mm, struct vm_area_struct *dst_vma,
>+ unsigned long dst_addr, unsigned long src_addr,
>+ pmd_t *dst_pmd, pmd_t *src_pmd,
>+ pmd_t orig_dst_pmd, pmd_t orig_src_pmd,
>+ spinlock_t *dst_ptl, spinlock_t *src_ptl,
>+ struct folio *src_folio, swp_entry_t entry)
>+{
[...]
>+
>+ moved_pmd = pmdp_huge_get_and_clear(mm, src_addr, src_pmd);
>+ if (pgtable_supports_soft_dirty())
>+ moved_pmd = pmd_swp_mksoft_dirty(moved_pmd);
Shouldn't we clear the source UFFD marker here?
moved_pmd = pmd_swp_clear_uffd(moved_pmd);
just like move_swap_pte(), no?
>+ /* Re-arm RWP on the moved swap entry if dst_vma is RWP-registered. */
>+ if (userfaultfd_rwp(dst_vma))
>+ moved_pmd = pmd_swp_mkuffd(moved_pmd);
>+ set_pmd_at(mm, dst_addr, dst_pmd, moved_pmd);
static int move_swap_pte(struct mm_struct *mm, struct vm_area_struct *dst_vma,
unsigned long dst_addr, unsigned long src_addr,
pte_t *dst_pte, pte_t *src_pte,
pte_t orig_dst_pte, pte_t orig_src_pte,
pmd_t *dst_pmd, pmd_t dst_pmdval,
spinlock_t *dst_ptl, spinlock_t *src_ptl,
struct folio *src_folio,
struct swap_info_struct *si, swp_entry_t entry)
{
...
orig_src_pte = ptep_get_and_clear(mm, src_addr, src_pte);
...
orig_src_pte = pte_swp_clear_uffd(orig_src_pte);
/* Re-arm RWP on the moved swap entry if dst_vma is RWP-registered. */
if (userfaultfd_rwp(dst_vma))
orig_src_pte = pte_swp_mkuffd(orig_src_pte);
set_pte_at(mm, dst_addr, dst_pte, orig_src_pte);
...
}
Suppose the source is an exclusive swap PMD carrying a UFFD-WP marker, and
the destination is an empty hole registered for synchronous UFFD-WP without
RWP. We haven't write-protected the destination, but a successful whole-PMD
MOVE carries the source's UFFD-WP marker into it.
On swap-in, that marker is restored in the present PMD, leaving it
write-protected. The write fault then reaches
handle_userfault(vmf, VM_UFFD_WP), even though we never write-protected
the destination ...
That would be quite a surprise for userspace, and could leave the writer
stuck if the unexpected WP event goes unhandled :)
vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf)
{
struct vm_area_struct *vma = vmf->vma;
...
bool exclusive, stable_writes, rwp_restore = false;
bool write = vmf->flags & FAULT_FLAG_WRITE;
...
pmd = folio_mk_pmd(folio, vma->vm_page_prot);
...
if (pmd_swp_uffd(vmf->orig_pmd))
pmd = pmd_mkuffd(pmd);
if (pmd_swp_uffd(vmf->orig_pmd) && userfaultfd_rwp(vma)) {
pmd = pmd_modify(pmd, PAGE_NONE);
rwp_restore = true;
}
...
if (exclusive) {
if (!rwp_restore && (vma->vm_flags & VM_WRITE) &&
!userfaultfd_huge_pmd_wp(vma, pmd) &&
!pmd_needs_soft_dirty_wp(vma, pmd)) {
pmd = pmd_mkwrite(pmd, vma);
if (write)
pmd = pmd_mkdirty(pmd);
}
rmap_flags |= RMAP_EXCLUSIVE;
}
...
vmf->orig_pmd = pmd;
...
if (write && !pmd_write(pmd) && !rwp_restore) {
vm_fault_t wp_ret = wp_huge_pmd(vmf);
...
}
...
}
vm_fault_t wp_huge_pmd(struct vm_fault *vmf)
{
struct vm_area_struct *vma = vmf->vma;
const bool unshare = vmf->flags & FAULT_FLAG_UNSHARE;
...
if (vma_is_anonymous(vma)) {
if (likely(!unshare) &&
userfaultfd_huge_pmd_wp(vma, vmf->orig_pmd)) {
if (userfaultfd_wp_async(vmf->vma))
goto split;
return handle_userfault(vmf, VM_UFFD_WP);
}
return do_huge_pmd_wp_page(vmf);
}
...
}
So, clearing the source UFFD-WP marker before the destination RWP check
would match the PTE helper and keep the RWP re-arm :)
Am I missing something?
>+ src_pgtable = pgtable_trans_huge_withdraw(mm, src_pmd);
>+ pgtable_trans_huge_deposit(mm, dst_pmd, src_pgtable);
>+
>+ double_pt_unlock(dst_ptl, src_ptl);
>+ return 0;
>+}
>+#endif /* CONFIG_THP_SWAP */
>+
[...]
Cheers, Lance