[PATCH] mm/huge_memory: avoid transient none PMDs during lazyfree reclaim

From: Kyle Zeng

Date: Fri Oct 09 2026 - 12:57:47 EST


__discard_anon_folio_pmd_locked() clears and flushes a huge PMD before
checking whether the folio can be discarded. If the folio was redirtied
or has unexpected references, it restores the original PMD. Reclaim holds
the PMD lock and the anon_vma read lock, but not the owning mmap_lock.

Both zap_pmd_range() and mremap's get_old_pmd() can skip a none PMD
without taking its lock. A concurrent munmap() or whole-VMA
MREMAP_DONTUNMAP can therefore miss the PMD and later unlink the source
VMA from its anon_vma, leaving a restored mapping behind. The anon_vma
write lock taken by unlink_anon_vmas() waits for reclaim to finish, but
does not repeat the skipped page-table walk. Once the remaining VMA
links are removed, the folio's positive mapcount no longer guarantees a
live anon_vma. Racing lazyfree reclaim against MREMAP_DONTUNMAP as an
unprivileged user reproduces a KASAN use-after-free in
folio_lock_anon_vma_read().

Use pmdp_invalidate() to keep a recognizable, non-none huge PMD while
the discard can still fail. Concurrent unmap and move operations then
have to synchronize on the PMD lock instead of skipping the mapping.
Clear the invalidated PMD only after the dirty and reference checks have
succeeded, before removing the rmap and withdrawing the deposited page
table. The full invalidation retains the TLB flush required for the
dirty and GUP-fast checks.

Fixes: 735ecdfaf4e8 ("mm/vmscan: avoid split lazyfree THP during shrink_folio_list()")
Cc: stable@xxxxxxxxxxxxxxx
Assisted-by: Codex:gpt-6-astra
Signed-off-by: Kyle Zeng <kylebot@xxxxxxxxxx>
---
mm/huge_memory.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)

diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 1e5d68acf62a..20262195385a 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -3569,11 +3569,17 @@ static bool __discard_anon_folio_pmd_locked(struct vm_area_struct *vma,
return false;
}

- orig_pmd = pmdp_huge_clear_flush(vma, addr, pmdp);
+ /*
+ * The discard may fail, so keep the PMD non-none until we're
+ * committed to discarding it. Otherwise, concurrent munmap() or
+ * mremap() can skip the PMD without taking the PTL and later unlink
+ * the VMA from its anon_vma despite a restored mapping.
+ */
+ orig_pmd = pmdp_invalidate(vma, addr, pmdp);

/*
* Syncing against concurrent GUP-fast:
- * - clear PMD; barrier; read refcount
+ * - invalidate PMD; barrier; read refcount
* - inc refcount; barrier; read PMD
*/
smp_mb();
@@ -3607,6 +3613,7 @@ static bool __discard_anon_folio_pmd_locked(struct vm_area_struct *vma,
return false;
}

+ pmdp_huge_get_and_clear(mm, addr, pmdp);
folio_remove_rmap_pmd(folio, pmd_page(orig_pmd), vma);
zap_deposited_table(mm, pmdp);
add_mm_counter(mm, MM_ANONPAGES, -HPAGE_PMD_NR);