Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
From: Lance Yang
Date: Wed Oct 07 2026 - 08:47:10 EST
On Wed, Oct 07, 2026 at 01:01:14PM +0200, David Hildenbrand (Arm) wrote:
>On 10/5/26 15:46, Lance Yang wrote:
>>
>> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>>> On 10/4/26 05:03, Lance Yang wrote:
>>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>>> a problem for PMD migration entries, since pmd_present() also checks
>>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>>> cleared during splitting.
>>>
>>> I'm curious: why do we have to set leaf indications for non-present things? The
>>> HW sure will ignore it, right?
>>
>> Yeah, that surprised me too :) Still wrapping my head around the details
>> ...
>>
>>> Is this a sw problem? Who needs that?
>>
>> IIUC, it's for software during a PMD split.
>>
>> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
>> before installing the PTE table. Software still needs pmd_present() and
>> pmd_trans_huge() to recognize the THP in between.
>>
>> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
>> to hardware but still identifiable as a THP by software.
>>
>> Hopefully I didn't miss something.
>
>I thought some change we performed to GUP-fast wouldn't require that anymore.
>But my memory is a bit vague ... :)
Yeah, I'll have another look at that :)
>>
>>>>
>>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>>> migration PMD as a present THP! The fault handler skips
>>>> pmd_migration_entry_wait(), and a write fault can end up in
>>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>>> as a mapped PFN.
>>>
>>> That sounds bad.
>>
>
>[...]
>
>> YES, looks a bit off ...
>>> That reduces the effective swap size (and PFN we can store). Could that be a
>>> problem?
>>
>> We don't need all 52 bits of the swap offset.
>>
>> RV64 PFNs only need 44 bits for migration entries, and actual swap is
>> already limited to about 16 TiB per area with 4 KiB pages by
>> last_page (__u32) and swap_info_struct.max (unsigned int).
>>
>> So there's room to reserve a bit without reducing the supported swap
>> size or PFN range.
>>
>> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
>> offset.
>
>Good, can you spell that out in the patch description?
Sure, will add that to the commit message in v2 once I've heard back
from maintainers :)
Cheers, Lance