Re: [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array()
From: David Hildenbrand (Arm)
Date: Fri Oct 02 2026 - 14:56:41 EST
On 10/2/26 03:25, Qiliang Yuan wrote:
> move_pages() with a NULL node list reports the node of each page. RDMA
> and KV-cache transfer engines use it to find where large registered
> buffers live, querying every 4K page of buffers that span hundreds of
> gigabytes.
>
> do_pages_stat_array() looks up the VMA and walks the page tables from
> the top for every address, taking and dropping the PTE lock each time.
> That costs about 105 ns per page, so a 16 GiB buffer takes 486 ms.
>
> Callers almost always pass consecutive addresses. Group them into runs
> and walk each run with walk_page_range(), which looks up each VMA and
> PTE table once and answers every page under it while holding the lock.
> Report pages as folio_walk_start() with FW_ZEROPAGE found them: the node
> of a normal folio, -EFAULT for the zero page or an address outside any
> VMA, and -ENOENT otherwise. Handle PUD and hugetlb leaves in their own
> callbacks so that the walk never splits them.
>
> On 7.3-rc5 in a 16-vCPU VM, querying every page of a populated 4 GiB
> buffer:
>
> before after
> 4K pages 105 ns 31.9 ns
> THP 90 ns 23.1 ns
>
> Suggested-by: Zi Yan <ziy@xxxxxxxxxx>
> Signed-off-by: Qiliang Yuan <odys.yuan@xxxxxxxxx>
> ---
> mm/migrate.c | 188 ++++++++++++++++++++++++++++++++++++++++++++++++++---------
> 1 file changed, 162 insertions(+), 26 deletions(-)
That's a lot of churn ... which is really a shame, because all we want to walk
is folio ranges.
Oscar was working on a better page table walker API (but I was too busy to
provide review so far :( ), which sounds strategically like the better long-term
solution.
Is there particular need to optimize this in the ns range for 4 KiB of memory
immediately?
--
Cheers,
David