Re: [PATCH v6 03/12] mm: add PMD swap entry splitting support

From: David Hildenbrand (Arm)

Date: Wed Aug 19 2026 - 12:10:27 EST


On 8/19/26 17:16, Kiryl Shutsemau wrote:
> On Tue, Aug 18, 2026 at 07:53:48PM +0200, David Hildenbrand (Arm) wrote:
>> On 8/18/26 15:09, Usama Arif wrote:
>>> Add a swap branch in __split_huge_pmd_locked() that splits a PMD swap
>>> entry into 512 PTE swap entries. No folio reference is needed because
>>> swap entries point to swap slots rather than pages. Each PTE inherits
>>> the correct sub-slot offset and preserves soft_dirty, uffd_wp, and
>>> exclusive flags.
>>>
>>> The folio_remove_rmap_pmd() gate at the end must inspect old_pmd
>>> rather than *pmd: for a present THP split, *pmd has already been
>>> cleared by pmdp_invalidate(), and that invalidated bit pattern can
>>> decode as a plausible swap entry.
>>>
>>> This branch is reached from the explicit __split_huge_pmd() callers
>>> that hit a non-present PMD: partial-range mprotect / munmap, the
>>> wp_huge_pmd() PMD-COW fallback, and the swap-in / swapoff fallbacks
>>> added in later patches when the cached folio is no longer PMD-sized.
>>> page_vma_mapped_walk() does not iterate PMD swap entries, so
>>> try_to_unmap_one() and try_to_migrate_one() do not reach this branch
>>> and freeze=true cannot occur in this branch today. page and folio
>>> are therefore left uninitialized in the swap branch; a
>>> VM_WARN_ON_ONCE(freeze) catches any future caller that breaks this
>>> invariant before the freeze path dereferences page_to_pfn(page + i)
>>> or put_page(page).
>>>
>>> Signed-off-by: Usama Arif <usama.arif@xxxxxxxxx>
>>> ---
>>> mm/huge_memory.c | 29 ++++++++++++++++++++++++++++-
>>> 1 file changed, 28 insertions(+), 1 deletion(-)
>>>
>>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
>>> index 1b6b0aa2baa3b..a473e85d30f51 100644
>>> --- a/mm/huge_memory.c
>>> +++ b/mm/huge_memory.c
>>> @@ -3252,6 +3252,14 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
>>> folio_add_anon_rmap_ptes(folio, page, HPAGE_PMD_NR,
>>> vma, haddr, rmap_flags);
>>> }
>>> + } else if (pmd_is_swap_entry(*pmd)) {
>>> + VM_WARN_ON_ONCE(freeze);
>>> + /* Swap entries have no page for the migration freeze path. */
>>> + freeze = false;
>>
>> It's odd to VM_WARN_ON_ONCE() and then set freeze=false;
>>
>> I'd just add the comment above the VM_WARN_ON_ONCE() and drop the =false.
>
> VM_WARN_ON_ONCE() is compiled out without DEBUG_VM. The freeze = false is
> what actually keeps the swap entry out of
>
> if (freeze || pmd_is_migration_entry(old_pmd)) {
> ...
> make_writable_migration_entry(page_to_pfn(page + i));
>
> and out of the put_page(page) below it, where page is uninitialized in the
> swap branch. Dropping it leaves nothing on production builds.

Right, because it's a condition you shouldn't hit on production builds.

And I raised how we can restructure the code to make it clear that we are being
overly cautions here.

Let's not add unnecessary code just because the existing code is hard to follow.

--
Cheers,

David