Re: [PATCH] mm/gup: batch PTE-mapped large folios in gup_fast_pte_range()

From: Rik van Riel

Date: Wed Oct 07 2026 - 06:23:02 EST


On Thu, 2026-09-17 at 10:42 +0200, David Hildenbrand (Arm) wrote:
> On 9/17/26 10:37, Yuan-Hao Hsu wrote:
> > GUP-fast grabs a PTE-mapped large folio one page at a time.  Every
> > PTE
> > costs a try_grab_folio_fast() (a refcount cmpxchg, plus the
> > pincount
> > atomic, a full barrier and a node stat update for FOLL_PIN), a
> > gup_fast_folio_allowed() and a folio_set_referenced(), so a 64 kB
> > mTHP
> > pays sixteen of each and a PTE-mapped 2 MB THP pays 512.  The PMD
> > and
> > PUD leaf paths already take the whole range with one
> > try_grab_folio_fast() call, and the slow path is getting the same
> > treatment for PTEs in Rik's follow_page_mask() series.
> >
> > IORING_REGISTER_BUFFERS of 1 GB, median of 15:
> >
> >   64 kB mTHP           6179/5998 us  ->  1599 us
> >   1 MB mTHP            5790/5801 us  ->   893 us
> >   4 kB pages           7173/6741 us  ->  6763 us
> >
>
> I think Rik was already working on this and sent some patches.
> Anyhow, there is
> quite some GUP review backlog I have t go through, so this will have
> to wait.

My patches are for the regular gup path, in
follow_pages_mask()

This is for the gup_fast path.

Unfortunately there is some nearly duplicated
code between them. Last I looked I did not
find any clear ways to merge them, but I
can take another look if you want.

--
All Rights Reversed.