Re: [PATCH v8 2/2] mm/memory_hotplug: optimize zone contiguous check when changing pfn range

From: David Hildenbrand (Arm)

Date: Thu Sep 10 2026 - 12:25:48 EST


On 9/1/26 07:29, Yuan Liu wrote:
> When move_pfn_range_to_zone() or remove_pfn_range_from_zone() updates a
> zone, set_zone_contiguous() rescans the entire zone pageblock-by-pageblock
> to rebuild zone->contiguous. For large zones this is a significant cost
> during memory hotplug and hot-unplug.
>
> Add a new zone member, pages_with_online_memmap, that tracks the
> number of pages within the zone span that have an online memory map,
> including present pages and memory holes whose memory map has been
> initialized and for which pfn_to_online_page() succeeds.
>
> For early boot memory, pages_with_online_memmap is calculated in
> memmap_init_zone_range(). PFNs initialized by memmap_init_range() are
> included in pages_with_online_memmap, and hole PFNs for which
> pfn_to_online_page() succeeds are also counted in
> init_unavailable_range(). For hotplugged memory,
> pages_with_online_memmap is updated through adjust_present_page_count(),
> which is called during memory online and offline operations. When
> spanned_pages == pages_with_online_memmap, every PFN in the zone span
> has a valid memmap entry, so pfn_to_page() can be called for any PFN
> within the zone span without an additional pfn_valid() check.
>
> The counter may temporarily undercount when pages with an online
> memory map exist outside the current zone span. This can only happen
> during boot, when initializing the memory map of pages that do not
> fall into any zone span. Growing the zone to cover such pages and
> later shrinking it back may result in a value that is too small.
> This is safe, as it merely prevents detecting a contiguous zone.
>
> The contiguity check using pages_with_online_memmap is stricter than
> the old pageblock-by-pageblock scan. The old set_zone_contiguous()
> iterated at pageblock granularity via pageblock_pfn_to_page(), so a
> zone could be marked contiguous even if a subsection-sized hole
> existed within a pageblock. The new check requires
> spanned_pages == pages_with_online_memmap, meaning every PFN in the
> zone span must satisfy pfn_to_online_page().
>
> The following test cases of memory hotplug for a VM [1], tested in the
> environment [2], show that this optimization can significantly reduce the
> memory hotplug time [3].
>
> +----------------+------+---------------+--------------+----------------+
> | | Size | Time (before) | Time (after) | Time Reduction |
> | +------+---------------+--------------+----------------+
> | Plug Memory | 256G | 10s | 3s | 70% |
> | +------+---------------+--------------+----------------+
> | | 512G | 36s | 7s | 81% |
> +----------------+------+---------------+--------------+----------------+
>
> +----------------+------+---------------+--------------+----------------+
> | | Size | Time (before) | Time (after) | Time Reduction |
> | +------+---------------+--------------+----------------+
> | Unplug Memory | 256G | 11s | 4s | 64% |
> | +------+---------------+--------------+----------------+
> | | 512G | 36s | 9s | 75% |
> +----------------+------+---------------+--------------+----------------+
>
> [1] Qemu commands to hotplug 256G/512G memory for a VM:
> object_add memory-backend-ram,id=hotmem0,size=256G/512G,share=on
> device_add virtio-mem-pci,id=vmem1,memdev=hotmem0,bus=port1
> qom-set vmem1 requested-size 256G/512G (Plug Memory)
> qom-set vmem1 requested-size 0G (Unplug Memory)
>
> [2] Hardware : Intel Icelake server
> Guest Kernel : v7.3-rc1
> Qemu : v9.0.0
>
> Launch VM :
> qemu-system-x86_64 -accel kvm -cpu host \
> -drive file=./Centos10_cloud.qcow2,format=qcow2,if=virtio \
> -drive file=./seed.img,format=raw,if=virtio \
> -smp 3,cores=3,threads=1,sockets=1,maxcpus=3 \
> -m 2G,slots=10,maxmem=2052472M \
> -device pcie-root-port,id=port1,bus=pcie.0,slot=1,multifunction=on \
> -device pcie-root-port,id=port2,bus=pcie.0,slot=2 \
> -nographic -machine q35 \
> -nic user,hostfwd=tcp::3000-:22
>
> Guest kernel auto-onlines newly added memory blocks:
> echo online > /sys/devices/system/memory/auto_online_blocks
>
> [3] The time from typing the QEMU commands in [1] to when the output of
> 'grep MemTotal /proc/meminfo' on Guest reflects that all hotplugged
> memory is recognized.
>
> Reported-by: Nanhai Zou <nanhai.zou@xxxxxxxxx>
> Reported-by: Chen Zhang <zhangchen.kidd@xxxxxx>
> Tested-by: Yuan Liu <yuan1.liu@xxxxxxxxx>
> Reviewed-by: Jason Zeng <jason.zeng@xxxxxxxxx>
> Reviewed-by: Chen Yu <yu.c.chen@xxxxxxxxx>
> Reviewed-by: Pan Deng <pan.deng@xxxxxxxxx>
> Co-developed-by: Tianyou Li <tianyou.li@xxxxxxxxx>
> Signed-off-by: Tianyou Li <tianyou.li@xxxxxxxxx>
> Signed-off-by: Yuan Liu <yuan1.liu@xxxxxxxxx>
> ---
> Documentation/mm/physical_memory.rst | 6 +++
> drivers/base/memory.c | 7 +++-
> include/linux/mmzone.h | 48 +++++++++++++++++++++
> mm/memory_hotplug.c | 12 +-----
> mm/mm_init.c | 63 +++++++++++++++-------------
> mm/mm_init.h | 6 ---
> mm/page_alloc.h | 2 +-
> 7 files changed, 98 insertions(+), 46 deletions(-)
>
> diff --git a/Documentation/mm/physical_memory.rst b/Documentation/mm/physical_memory.rst
> index a09407d72973..2e67e8b23a99 100644
> --- a/Documentation/mm/physical_memory.rst
> +++ b/Documentation/mm/physical_memory.rst
> @@ -480,6 +480,12 @@ General
> ``present_pages`` should use ``get_online_mems()`` to get a stable value. It
> is initialized by ``calculate_node_totalpages()``.
>
> +``pages_with_online_memmap``
> + Pages within the zone that have an online memory map: present pages and
> + memory holes whose memory map has been initialized and
> + ``pfn_to_online_page()`` succeeds. See the comment for
> + ``pages_with_online_memmap`` in ``include/linux/mmzone.h`` for more details.
> +
> ``present_early_pages``
> The present pages existing within the zone located on memory available since
> early boot, excluding hotplugged memory. Defined only when
> diff --git a/drivers/base/memory.c b/drivers/base/memory.c
> index 5eead3346f1e..28f9503f6a71 100644
> --- a/drivers/base/memory.c
> +++ b/drivers/base/memory.c
> @@ -255,6 +255,7 @@ static int memory_block_online(struct memory_block *mem)
> nr_vmemmap_pages = mem->altmap->free;
>
> mem_hotplug_begin();
> + clear_zone_contiguous(zone);
> if (nr_vmemmap_pages) {
> ret = mhp_init_memmap_on_memory(start_pfn, nr_vmemmap_pages, zone);
> if (ret)
> @@ -279,6 +280,7 @@ static int memory_block_online(struct memory_block *mem)
>
> mem->zone = zone;
> out:
> + set_zone_contiguous(zone);
> mem_hotplug_done();
> return ret;
> }
> @@ -304,6 +306,7 @@ static int memory_block_offline(struct memory_block *mem)
> nr_vmemmap_pages = mem->altmap->free;
>
> mem_hotplug_begin();
> + clear_zone_contiguous(mem->zone);
> if (nr_vmemmap_pages)
> adjust_present_page_count(pfn_to_page(start_pfn), mem->group,
> -nr_vmemmap_pages);
> @@ -321,8 +324,10 @@ static int memory_block_offline(struct memory_block *mem)
> if (nr_vmemmap_pages)
> mhp_deinit_memmap_on_memory(start_pfn, nr_vmemmap_pages);
>
> - mem->zone = NULL;
> out:
> + set_zone_contiguous(mem->zone);
> + if (!ret)
> + mem->zone = NULL;
> mem_hotplug_done();
> return ret;
> }
> diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
> index 94f9c3ff5416..7bfb871d6344 100644
> --- a/include/linux/mmzone.h
> +++ b/include/linux/mmzone.h
> @@ -1043,6 +1043,21 @@ struct zone {
> * cma pages is present pages that are assigned for CMA use
> * (MIGRATE_CMA).
> *
> + * pages_with_online_memmap tracks pages within the zone that have
> + * an online memory map: present pages and memory holes whose
> + * memory map has been initialized and pfn_to_online_page()
> + * succeeds. When spanned_pages == pages_with_online_memmap,
> + * pfn_to_page() can be performed without further checks on any
> + * PFN within the zone span.
> + *
> + * Note: this counter may temporarily undercount when pages with an
> + * online memory map exist outside the current zone span. This can
> + * only happen during boot, when initializing the memory map of
> + * pages that do not fall into any zone span. Growing the zone to
> + * cover such pages and later shrinking it back may result in a
> + * "too small" value. This is safe: it merely prevents detecting a
> + * contiguous zone.

It's actually:

[ zone span ]
[ zone pages ]

e.g., spanned 10, initialized 15, online 10

online == spanned -> contiguous

Growing after hotplug (hotplug 5):

[ zone span ]
[ zone pages ] [ zone pages ]

e.g., spanned 30, initialized 20, online 15

online != spanned -> not contiguous

Shrinking after hotunplug (hotunplug 5 again):

[ zone span ]
[ zone pages ]

e.g., spanned 15, initialized 15, online 10

online != spanned -> not contiguous although contiguous


So this can only happen after boot (during memory hotplug) when we grow and then
shrink the zone.


Should be clarified in here and in the patch description, maybe even stating an
example as above (obviously we cannot hotplug/hotunplug 5 pages ;) )

--
Cheers,

David