Re: [PATCH v2 1/2] mm/migrate: walk runs of consecutive pages in do_pages_stat_array()

From: Qiliang Yuan

Date: Wed Oct 07 2026 - 22:27:58 EST


Hi Ying, David,

On Mon, 05 Oct 2026 21:24:07 +0800, Huang, Ying wrote:
> As pointed out by David, nanosecond-level optimization for a not-so-hot
> path isn't very attractive. I understand the target of your
> optimization is not the performance of a single page but that of a large
> number of pages (such as 16 GiB). So, please describe more clearly why
> your change is necessary, for example, by providing the performance
> improvement of querying 16 GiB memory.

Right, the target is large buffers, which also answers David's question
about the need. Mooncake, the KV-cache store used for Kimi serving,
calls move_pages() with a NULL node list on every 4K page of the
buffers it registers for RDMA, to find which NUMA node each one lives
on. Its issue tracker reports 216 ms for a 4 GiB buffer, and it
registers hundreds of GiB per instance, so this query alone takes
seconds of every startup.

Time to query every page of a populated buffer, with the two v2
patches applied, median of 5 runs:

before after change
1 GiB, 4K pages 27.3 ms 3.2 ms -88%
4 GiB, 4K pages 109.5 ms 13.0 ms -88%
16 GiB, 4K pages 486.1 ms 52.3 ms -89%
1 GiB, THP 23.5 ms 0.5 ms -98%
4 GiB, THP 95.0 ms 2.1 ms -98%
16 GiB, THP 381.0 ms 9.1 ms -98%

> Additionally, the raw performance number depends on the system under
> test. Please provide a little more information about your testing
> system, for example, the CPU architecture, generation, physical core
> count, etc. For comparison, the performance improvement percentage
> would also be helpful.

The host is a 2-socket AMD EPYC 9654 (Zen 4, 96 cores per socket). The
test runs in a KVM guest on 7.3-rc5 with 16 vCPUs pinned to one host
NUMA node and 32 GiB of memory, split into two guest NUMA nodes. The
test program is bound to CPU 0, and the before and after kernels were
measured back to back in the same guest.

Thanks,
Qiliang