Re: [PATCH v2 03/17] mm/mm_init: skip initializing shared vmemmap tail pages
From: Muchun Song
Date: Thu Jul 30 2026 - 09:26:52 EST
> On Jul 30, 2026, at 18:48, Mike Rapoport <rppt@xxxxxxxxxx> wrote:
>
> Hi Muchun,
Hi,
>
> On 2026-07-26 12:30:17+08:00, Muchun Song wrote:
>>> On Jul 20, 2026, at 17:31, Muchun Song <songmuchun@xxxxxxxxxxxxx> wrote:
>>>
>>> memmap_init_range() initializes every struct page in the target range.
>>> For compound pages with vmemmap optimization, the tail struct pages are
>>> backed by a shared vmemmap page.
>>>
>>> Initializing those tail struct pages would overwrite the shared
>>> vmemmap page contents, requiring users such as HugeTLB to restore the
>>> metadata afterwards.
>>>
>>> Track the compound order for HVO-backed sections and use that metadata
>>> to detect struct pages that fall into the shared tail vmemmap range.
>>> Skip those shared tail pages in memmap_init_range(), then initialize
>>> pageblock migratetypes for the processed range with a helper after the
>>> per-page initialization loop.
>>>
>>> The !SPARSEMEM __pfn_to_section() stub is needed only for the build:
>>> memmap_init_range() references __pfn_to_section() after checking
>>> pfn_vmemmap_optimizable(), and !SPARSEMEM builds still have to compile
>>> that code even though pfn_vmemmap_optimizable() folds to false.
>>>
>>> This is a preparatory change for consolidating handling across users of
>>> vmemmap optimization, and it also avoids redundant initialization of
>>> shared tail vmemmap pages during early boot.
>>>
>>> Signed-off-by: Muchun Song <songmuchun@xxxxxxxxxxxxx>
>>> ---
>>> v2:
>>> - Fold section order tracking into the first user instead of keeping a
>>> standalone API-only patch (suggested by Mike Rapoport)
>>> - Rename page_vmemmap_optimizable() to pfn_vmemmap_optimizable() and
>>> pass a PFN directly (suggested by Mike Rapoport)
>>> - Initialize pageblock migratetypes from a helper after the per-page
>>> loop (suggested by Mike Rapoport)
>>> - Use a 1G PFN chunk for cond_resched() in the pageblock helper
>>> (suggested by Mike Rapoport)
>>> - Guard section_order() with CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP so
>>> it returns 0 when HVO is disabled and lets the compiler optimize the
>>> code as much as possible (suggested by Mike Rapoport)
>>> - Explain why the !SPARSEMEM __pfn_to_section() stub belongs here
>>> (suggested by Mike Rapoport)
>>> ---
>>> include/linux/mmzone.h | 14 ++++++++++++++
>>> mm/mm_init.c | 33 +++++++++++++++++----------------
>>> mm/sparse.h | 23 +++++++++++++++++++++++
>>> 3 files changed, 54 insertions(+), 16 deletions(-)
>>>
>>> diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
>>> index 82b0155d886f..2a32101d55e6 100644
>>> --- a/include/linux/mmzone.h
>>> +++ b/include/linux/mmzone.h
>>> @@ -2011,6 +2011,14 @@ struct mem_section {
>>> unsigned long section_mem_map;
>>>
>>> struct mem_section_usage *usage;
>>> +#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
>>> + /*
>>> + * Normally, sections hold regular (order-0) pages. However, for
>>> + * sections with HVO enabled, this tracks the compound page order
>>> + * to enable deduplication of redundant vmemmap pages.
>>> + */
>>> + unsigned int order;
>>> +#endif
>>> #ifdef CONFIG_PAGE_EXTENSION
>>> /*
>>> * If SPARSEMEM, pgdat doesn't have page_ext pointer. We use
>>> @@ -2365,8 +2373,14 @@ static inline unsigned long next_present_section_nr(unsigned long section_nr)
>>> #endif
>>>
>>> #else
>>> +struct mem_section;
>>> +
>>> #define sparse_vmemmap_init_nid_early(_nid) do {} while (0)
>>> #define pfn_in_present_section pfn_valid
>>> +static inline struct mem_section *__pfn_to_section(unsigned long pfn)
>>> +{
>>> + return NULL;
>>> +}
>>
>> I'd like to propose an alternative implementation that doesn't require
>> exposing the mem_section. The idea is to add a new helper function,
>> pfn_to_section_order(), so that for non-sparse-memory configurations,
>> the mem_section concept stays hidden internally. I'd really appreciate
>> any thoughts or concerns — if everyone is comfortable with it, I can go
>> ahead and implement this in the next version.
>
> A helper that keeps mem_section hidden from !SPARSMEM makes perfect
> sense to me.
>
> I'd even take it one step further and make it return how many pfns
> should be skipped in pfn_vmemmap_optimizable case.
To make sure we're on the same page, let me walk you through the specific
changes I have in mind. My initial plan is to introduce pfn_to_section_order,
and the expected diff changes are as follow to keep mem_sectionhidden from
!SPARSEMEM.
diff --git a/mm/mm_init.c b/mm/mm_init.c
index dcb757b36902..0b0c2996d080 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -884,7 +884,7 @@ void __meminit memmap_init_range(unsigned long size, int nid, unsigned long zone
}
if (pfn_vmemmap_optimizable(pfn)) {
- unsigned int order = section_order(__pfn_to_section(pfn));
+ unsigned int order = pfn_to_section_order(pfn);
pfn = min(ALIGN(pfn, 1UL << order), end_pfn);
continue;
diff --git a/mm/sparse.h b/mm/sparse.h
index 030248030dc7..c5fbcdde3cee 100644
--- a/mm/sparse.h
+++ b/mm/sparse.h
@@ -47,6 +47,11 @@ static inline void __section_mark_present(struct mem_section *ms,
ms->section_mem_map |= SECTION_MARKED_PRESENT;
}
+
+static inline unsigned int pfn_to_section_order(unsigned long pfn)
+{
+ return section_order(__pfn_to_section(pfn));
+}
#else
static inline void sparse_init(void) {}
#endif /* CONFIG_SPARSEMEM */
Since we also use __pfn_to_section in the patch 14 in this series for
!SPARSEMEM, we need to make corresponding adjustments—specifically, by using
pfn_to_section_order to determine whether the vmemmap of a given section is
optimizable. This new helper will be called from several places, so I'm afraid
its introduction is unavoidable.
That said, I've also considered an alternative: introducing another helper that
returns the exact number of PFNs to skip, and using it solely within
memmap_init_range(). However, that approach doesn't seem to offer much in terms
of code simplification. If I'm missing something or if my reasoning doesn't align
with your expectations, I would really appreciate your guidance. Thank you for
your patience!
Thanks,
Muchun