Re: [PATCH] mm/huge_memory: avoid transient none PMDs during lazyfree reclaim
From: Lance Yang
Date: Fri Oct 09 2026 - 22:04:42 EST
On 2026/10/10 09:56, Zi Yan wrote:
On Fri Oct 9, 2026 at 9:12 PM EDT, Lance Yang wrote:
On 2026/10/10 04:44, David Hildenbrand (Arm) wrote:
On 10/9/26 22:27, Zi Yan wrote:
On 9 Oct 2026, at 16:23, David Hildenbrand (Arm) wrote:
On 10/9/26 20:17, Zi Yan wrote:
There are some gaps we need to close before getting this fix in:
1. GUP-fast cannot follow the invalidated PMD: riscv and LoongArch need
pmd_access_permitted() that requires _PAGE_PRESENT; s390's
pmdp_invalidate() needs a change.
Ugh. We really need pmdp_invalidate() to have reasonable semantics. This whole
PMD locking is a mess :(
2. sparc64's thp_pte_count can be imbalanced with this change (based on
Lance's offlist feedback).
Ack.
In addition, powerpc has a page table check issue similr to riscv, where riscv
fixed it with commit 9f4a88b8d01a6. This is not related to this issue
but discovered along with the investigation.
Zi, do you have the capacity to take over this patch?
Yes, I can take over it. My plan is to send a series including
patches for 1 and this patch as is. powerpc fix can be a separate one.
Best Regards,
Yan, Zi
If we need a quick stable fix we could temporarily disable the whole thing until
it is fixed, just a thought.
+1 we could do that for now. A quick fix with minimal churn (correctness
comes first).
How about the patch below? unmap_huge_pmd_locked() and
__discard_anon_folio_pmd_locked() will be dead code to keep the patch
small. Later, with arch fixes, the function can be re-enabled.
Looks good to me! Thanks for the quick fix!
From 2204b9d08530796847a586428b1f782b987d071c Mon Sep 17 00:00:00 2001
From: Zi Yan <ziy@xxxxxxxxxx>
Date: Fri, 9 Oct 2026 21:28:10 -0400
Subject: [PATCH] mm/rmap: don't discard lazyfree THPs at PMD level
__discard_anon_folio_pmd_locked() clears a lazyfree THP PMD before it knows
whether the folio can be discarded, and restores it if the folio was
redirtied or has extra references. A concurrent munmap() or
MREMAP_DONTUNMAP skips the temporary none PMD and unlinks the VMA from its
anon_vma, so the folio stays mapped after the anon_vma is freed and a later
rmap walk uses the freed anon_vma.
Using an invalidated PMD instead of a cleared one requires additional arch
code fixes. Instead, disable the PMD level discard of lazyfree THPs, as
before commit 735ecdfaf4e8 ("mm/vmscan: avoid split lazyfree THP during
shrink_folio_list()").
Fixes: 735ecdfaf4e8 ("mm/vmscan: avoid split lazyfree THP during shrink_folio_list()")
Reported-by: Kyle Zeng <kylebot@xxxxxxxxxx>
Closes: https://lore.kernel.org/r/20261009165214.40212-2-kylebot@xxxxxxxxxx
Cc: stable@xxxxxxxxxxxxxxx
Assisted-by: LLM
Signed-off-by: Zi Yan <ziy@xxxxxxxxxx>
---
Reviewed-by: Lance Yang <lance.yang@xxxxxxxxx>
mm/rmap.c | 11 -----------
1 file changed, 11 deletions(-)
diff --git a/mm/rmap.c b/mm/rmap.c
index 805db93fe0428..1131b76bbbc28 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -2275,17 +2275,6 @@ static bool try_to_unmap_one(struct folio *folio, struct vm_area_struct *vma,
}
if (!pvmw.pte) {
- if (folio_test_lazyfree(folio)) {
- if (unmap_huge_pmd_locked(vma, pvmw.address, pvmw.pmd, folio))
- goto walk_done;
- /*
- * unmap_huge_pmd_locked has either already marked
- * the folio as swap-backed or decided to retain it
- * due to GUP or speculative references.
- */
- goto walk_abort;
- }
-
if (flags & TTU_SPLIT_HUGE_PMD) {
/*
* We temporarily have to drop the PTL and