RE: [PATCH v6 1/2] mm/memory_hotplug: optimize zone contiguous check when changing pfn range
From: Liu, Yuan1
Date: Fri Aug 07 2026 - 08:18:17 EST
> -----Original Message-----
> From: David Hildenbrand (Arm) <david@xxxxxxxxxx>
> Sent: Friday, August 7, 2026 7:30 PM
> To: Liu, Yuan1 <yuan1.liu@xxxxxxxxx>; Oscar Salvador <osalvador@xxxxxxx>;
> Mike Rapoport <rppt@xxxxxxxxxx>; Wei Yang <richard.weiyang@xxxxxxxxx>
> Cc: linux-mm@xxxxxxxxx; Zou, Nanhai <nanhai.zou@xxxxxxxxx>; Deng, Pan
> <pan.deng@xxxxxxxxx>; Li, Tianyou <tianyou.li@xxxxxxxxx>; Chen Zhang
> <zhangchen.kidd@xxxxxx>; Zeng, Jason <jason.zeng@xxxxxxxxx>; linux-
> kernel@xxxxxxxxxxxxxxx
> Subject: Re: [PATCH v6 1/2] mm/memory_hotplug: optimize zone contiguous
> check when changing pfn range
>
> On 8/6/26 11:52, Liu, Yuan1 wrote:
> >> -----Original Message-----
> >> From: David Hildenbrand (Arm) <david@xxxxxxxxxx>
> >> Sent: Thursday, August 6, 2026 4:46 PM
> >> To: Liu, Yuan1 <yuan1.liu@xxxxxxxxx>; Oscar Salvador
> <osalvador@xxxxxxx>;
> >> Mike Rapoport <rppt@xxxxxxxxxx>; Wei Yang <richard.weiyang@xxxxxxxxx>
> >> Cc: linux-mm@xxxxxxxxx; Zou, Nanhai <nanhai.zou@xxxxxxxxx>; Deng, Pan
> >> <pan.deng@xxxxxxxxx>; Li, Tianyou <tianyou.li@xxxxxxxxx>; Chen Zhang
> >> <zhangchen.kidd@xxxxxx>; Zeng, Jason <jason.zeng@xxxxxxxxx>; linux-
> >> kernel@xxxxxxxxxxxxxxx
> >> Subject: Re: [PATCH v6 1/2] mm/memory_hotplug: optimize zone contiguous
> >> check when changing pfn range
> >>
> >>
> >>>
> >>> Will do.
> >>>
> >>>
> >>> pages_with_online_memmap counts all PFNs where pfn_to_online_page() is
> >>> valid. With CONFIG_SPARSEMEM_VMEMMAP, pfn_section_valid() operates at
> >>> PAGES_PER_SUBSECTION granularity — when any page in a subsection has
> >>> memory, the entire subsection is valid/online. So we align to
> subsection
> >>> boundaries to include hole pages within partially-populated
> subsections.
> >>
> >> But we must never account exceeding the zone range. So I don't
> understand
> >> why we
> >> would have to care about PAGES_PER_SUBSECTION here at all?
> >
> > We never account beyond the zone range, because `sub_start` and
> > `sub_end` are still clamped to the zone boundaries after the
> > alignment.
> >
> > The subsection alignment is needed for hole PFNs between memblocks
> > within the zone. These hole PFNs are valid for
> > `pfn_to_online_page()`, because they sit in a subsection that has
> > memory, so the whole subsection's memmap is online.
> >
> > |----------- zone range -----------|
> >
> > +-----------+---------+------------+
> > + memblock 1| hole | memblock 2 |
> > +-----------+---------+------------+
> > ^
> > |
> > |
> > subsection boundary
>
> Right, but init_unavailable_range() can just return how many were actually
> initalized? Why can't we piggy-back on that?
>
> I think we had something similar previously, why can't we use that?
>
> We know the zone span, so we can just account all the mmap in the zone
> span
> that we initialize.
>
> What is the problem with that?
Hi David
My understanding is that init_unavailable_range() initializes all
PFNs that satisfy pfn_valid(), but not all of them satisfy
pfn_to_online_page(), since some PFNs belong to subsections that are
not online.
You previously mentioned:
pfn_valid() says early sections always have a full memmap, so even invalid
subsections have a memmap. pfn_to_online_page() says an invalid subsection
cannot be online and its content must be stale. for_each_valid_pfn() follows
pfn_valid() semantics, and we use it to initialize memmap that is not going
to be online and account it as pages_with_online_memmap, which is wrong.
The cleanest approach is to avoid allocating memmap for subsections, which
also removes the special early-section handling from pfn_valid() and
for_each_valid_pfn().
I also share the concerns raised by Sashiko in the analysis below [1]:
Scanners like isolate_migratepages_block() will then blindly iterate through
the pageblock and access the completely uninitialized struct pages of the hole,
leading to functional errors or kernel panics when reading these zero-filled
structures via macros like PageHuge() or page_zone().
That's why we went with the current approach in v6 instead of your
earlier suggestion. I'd really appreciate your guidance on which
direction you think would be more appropriate.
[1] https://sashiko.dev/#/patchset/20260520093457.3719960-1-yuan1.liu%40intel.com
Best Regards,
Liu, Yuan
> Cheers,
>
> David