Re: [PATCH] mm/gup: batch PTE-mapped large folios in gup_fast_pte_range()

From: David Hildenbrand (Arm)

Date: Wed Oct 07 2026 - 07:47:07 EST


On 10/7/26 12:07, Rik van Riel wrote:
> On Thu, 2026-09-17 at 10:42 +0200, David Hildenbrand (Arm) wrote:
>> On 9/17/26 10:37, Yuan-Hao Hsu wrote:
>>> GUP-fast grabs a PTE-mapped large folio one page at a time.  Every
>>> PTE
>>> costs a try_grab_folio_fast() (a refcount cmpxchg, plus the
>>> pincount
>>> atomic, a full barrier and a node stat update for FOLL_PIN), a
>>> gup_fast_folio_allowed() and a folio_set_referenced(), so a 64 kB
>>> mTHP
>>> pays sixteen of each and a PTE-mapped 2 MB THP pays 512.  The PMD
>>> and
>>> PUD leaf paths already take the whole range with one
>>> try_grab_folio_fast() call, and the slow path is getting the same
>>> treatment for PTEs in Rik's follow_page_mask() series.
>>>
>>> IORING_REGISTER_BUFFERS of 1 GB, median of 15:
>>>
>>>   64 kB mTHP           6179/5998 us  ->  1599 us
>>>   1 MB mTHP            5790/5801 us  ->   893 us
>>>   4 kB pages           7173/6741 us  ->  6763 us
>>>
>>
>> I think Rik was already working on this and sent some patches.
>> Anyhow, there is
>> quite some GUP review backlog I have t go through, so this will have
>> to wait.
>
> My patches are for the regular gup path, in
> follow_pages_mask()
>
> This is for the gup_fast path.
>
> Unfortunately there is some nearly duplicated
> code between them. Last I looked I did not
> find any clear ways to merge them, but I
> can take another look if you want.

We'll likely have to keep it separate.

I looked at the patch here and it needs quite some work do do it cleanly. Too
much for me to comment on this, so I'll likely end up doing it myself or
delegate it to someone internally.

--
Cheers,

David