Re: [PATCH] mm/huge_memory: avoid transient none PMDs during lazyfree reclaim
From: Zi Yan
Date: Fri Oct 09 2026 - 16:27:33 EST
On 9 Oct 2026, at 16:23, David Hildenbrand (Arm) wrote:
> On 10/9/26 20:17, Zi Yan wrote:
>> On Fri Oct 9, 2026 at 12:52 PM EDT, Kyle Zeng wrote:
>>> __discard_anon_folio_pmd_locked() clears and flushes a huge PMD before
>>> checking whether the folio can be discarded. If the folio was redirtied
>>> or has unexpected references, it restores the original PMD. Reclaim holds
>>> the PMD lock and the anon_vma read lock, but not the owning mmap_lock.
>>>
>>> Both zap_pmd_range() and mremap's get_old_pmd() can skip a none PMD
>>> without taking its lock. A concurrent munmap() or whole-VMA
>>> MREMAP_DONTUNMAP can therefore miss the PMD and later unlink the source
>>> VMA from its anon_vma, leaving a restored mapping behind. The anon_vma
>>> write lock taken by unlink_anon_vmas() waits for reclaim to finish, but
>>> does not repeat the skipped page-table walk. Once the remaining VMA
>>> links are removed, the folio's positive mapcount no longer guarantees a
>>> live anon_vma. Racing lazyfree reclaim against MREMAP_DONTUNMAP as an
>>> unprivileged user reproduces a KASAN use-after-free in
>>> folio_lock_anon_vma_read().
>>>
>>> Use pmdp_invalidate() to keep a recognizable, non-none huge PMD while
>>> the discard can still fail. Concurrent unmap and move operations then
>>> have to synchronize on the PMD lock instead of skipping the mapping.
>>> Clear the invalidated PMD only after the dirty and reference checks have
>>> succeeded, before removing the rmap and withdrawing the deposited page
>>> table. The full invalidation retains the TLB flush required for the
>>> dirty and GUP-fast checks.
>>>
>>> Fixes: 735ecdfaf4e8 ("mm/vmscan: avoid split lazyfree THP during shrink_folio_list()")
>>> Cc: stable@xxxxxxxxxxxxxxx
>>> Assisted-by: Codex:gpt-6-astra
>>> Signed-off-by: Kyle Zeng <kylebot@xxxxxxxxxx>
>>> ---
>>> mm/huge_memory.c | 11 +++++++++--
>>> 1 file changed, 9 insertions(+), 2 deletions(-)
>>>
>>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
>>> index 1e5d68acf62a..20262195385a 100644
>>> --- a/mm/huge_memory.c
>>> +++ b/mm/huge_memory.c
>>> @@ -3569,11 +3569,17 @@ static bool __discard_anon_folio_pmd_locked(struct vm_area_struct *vma,
>>> return false;
>>> }
>>>
>>> - orig_pmd = pmdp_huge_clear_flush(vma, addr, pmdp);
>>> + /*
>>> + * The discard may fail, so keep the PMD non-none until we're
>>> + * committed to discarding it. Otherwise, concurrent munmap() or
>>> + * mremap() can skip the PMD without taking the PTL and later unlink
>>> + * the VMA from its anon_vma despite a restored mapping.
>>> + */
>>> + orig_pmd = pmdp_invalidate(vma, addr, pmdp);
>>
>> There are some gaps we need to close before getting this fix in:
>>
>> 1. GUP-fast cannot follow the invalidated PMD: riscv and LoongArch need
>> pmd_access_permitted() that requires _PAGE_PRESENT; s390's
>> pmdp_invalidate() needs a change.
>
> Ugh. We really need pmdp_invalidate() to have reasonable semantics. This whole
> PMD locking is a mess :(
>
>>
>> 2. sparc64's thp_pte_count can be imbalanced with this change (based on
>> Lance's offlist feedback).
>
> Ack.
>
>>
>> In addition, powerpc has a page table check issue similr to riscv, where riscv
>> fixed it with commit 9f4a88b8d01a6. This is not related to this issue
>> but discovered along with the investigation.
>
> Zi, do you have the capacity to take over this patch?
Yes, I can take over it. My plan is to send a series including
patches for 1 and this patch as is. powerpc fix can be a separate one.
Best Regards,
Yan, Zi