Re: [PATCH] mm/gup: batch PTE-mapped large folios in gup_fast_pte_range()
From: Rik van Riel
Date: Wed Oct 07 2026 - 06:23:02 EST
On Thu, 2026-09-17 at 10:42 +0200, David Hildenbrand (Arm) wrote:
> On 9/17/26 10:37, Yuan-Hao Hsu wrote:
> > GUP-fast grabs a PTE-mapped large folio one page at a time. Every
> > PTE
> > costs a try_grab_folio_fast() (a refcount cmpxchg, plus the
> > pincount
> > atomic, a full barrier and a node stat update for FOLL_PIN), a
> > gup_fast_folio_allowed() and a folio_set_referenced(), so a 64 kB
> > mTHP
> > pays sixteen of each and a PTE-mapped 2 MB THP pays 512. The PMD
> > and
> > PUD leaf paths already take the whole range with one
> > try_grab_folio_fast() call, and the slow path is getting the same
> > treatment for PTEs in Rik's follow_page_mask() series.
> >
> > IORING_REGISTER_BUFFERS of 1 GB, median of 15:
> >
> > 64 kB mTHP 6179/5998 us -> 1599 us
> > 1 MB mTHP 5790/5801 us -> 893 us
> > 4 kB pages 7173/6741 us -> 6763 us
> >
>
> I think Rik was already working on this and sent some patches.
> Anyhow, there is
> quite some GUP review backlog I have t go through, so this will have
> to wait.
My patches are for the regular gup path, in
follow_pages_mask()
This is for the gup_fast path.
Unfortunately there is some nearly duplicated
code between them. Last I looked I did not
find any clear ways to merge them, but I
can take another look if you want.
--
All Rights Reversed.