Re: [PATCH v2 03/17] mm/mm_init: skip initializing shared vmemmap tail pages
From: Mike Rapoport
Date: Thu Jul 30 2026 - 06:50:33 EST
Hi Muchun,
On 2026-07-26 12:30:17+08:00, Muchun Song wrote:
> > On Jul 20, 2026, at 17:31, Muchun Song <songmuchun@xxxxxxxxxxxxx> wrote:
> >
> > memmap_init_range() initializes every struct page in the target range.
> > For compound pages with vmemmap optimization, the tail struct pages are
> > backed by a shared vmemmap page.
> >
> > Initializing those tail struct pages would overwrite the shared
> > vmemmap page contents, requiring users such as HugeTLB to restore the
> > metadata afterwards.
> >
> > Track the compound order for HVO-backed sections and use that metadata
> > to detect struct pages that fall into the shared tail vmemmap range.
> > Skip those shared tail pages in memmap_init_range(), then initialize
> > pageblock migratetypes for the processed range with a helper after the
> > per-page initialization loop.
> >
> > The !SPARSEMEM __pfn_to_section() stub is needed only for the build:
> > memmap_init_range() references __pfn_to_section() after checking
> > pfn_vmemmap_optimizable(), and !SPARSEMEM builds still have to compile
> > that code even though pfn_vmemmap_optimizable() folds to false.
> >
> > This is a preparatory change for consolidating handling across users of
> > vmemmap optimization, and it also avoids redundant initialization of
> > shared tail vmemmap pages during early boot.
> >
> > Signed-off-by: Muchun Song <songmuchun@xxxxxxxxxxxxx>
> > ---
> > v2:
> > - Fold section order tracking into the first user instead of keeping a
> > standalone API-only patch (suggested by Mike Rapoport)
> > - Rename page_vmemmap_optimizable() to pfn_vmemmap_optimizable() and
> > pass a PFN directly (suggested by Mike Rapoport)
> > - Initialize pageblock migratetypes from a helper after the per-page
> > loop (suggested by Mike Rapoport)
> > - Use a 1G PFN chunk for cond_resched() in the pageblock helper
> > (suggested by Mike Rapoport)
> > - Guard section_order() with CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP so
> > it returns 0 when HVO is disabled and lets the compiler optimize the
> > code as much as possible (suggested by Mike Rapoport)
> > - Explain why the !SPARSEMEM __pfn_to_section() stub belongs here
> > (suggested by Mike Rapoport)
> > ---
> > include/linux/mmzone.h | 14 ++++++++++++++
> > mm/mm_init.c | 33 +++++++++++++++++----------------
> > mm/sparse.h | 23 +++++++++++++++++++++++
> > 3 files changed, 54 insertions(+), 16 deletions(-)
> >
> > diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
> > index 82b0155d886f..2a32101d55e6 100644
> > --- a/include/linux/mmzone.h
> > +++ b/include/linux/mmzone.h
> > @@ -2011,6 +2011,14 @@ struct mem_section {
> > unsigned long section_mem_map;
> >
> > struct mem_section_usage *usage;
> > +#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
> > + /*
> > + * Normally, sections hold regular (order-0) pages. However, for
> > + * sections with HVO enabled, this tracks the compound page order
> > + * to enable deduplication of redundant vmemmap pages.
> > + */
> > + unsigned int order;
> > +#endif
> > #ifdef CONFIG_PAGE_EXTENSION
> > /*
> > * If SPARSEMEM, pgdat doesn't have page_ext pointer. We use
> > @@ -2365,8 +2373,14 @@ static inline unsigned long next_present_section_nr(unsigned long section_nr)
> > #endif
> >
> > #else
> > +struct mem_section;
> > +
> > #define sparse_vmemmap_init_nid_early(_nid) do {} while (0)
> > #define pfn_in_present_section pfn_valid
> > +static inline struct mem_section *__pfn_to_section(unsigned long pfn)
> > +{
> > + return NULL;
> > +}
>
> I'd like to propose an alternative implementation that doesn't require
> exposing the mem_section. The idea is to add a new helper function,
> pfn_to_section_order(), so that for non-sparse-memory configurations,
> the mem_section concept stays hidden internally. I'd really appreciate
> any thoughts or concerns — if everyone is comfortable with it, I can go
> ahead and implement this in the next version.
A helper that keeps mem_section hidden from !SPARSMEM makes perfect
sense to me.
I'd even take it one step further and make it return how many pfns
should be skipped in pfn_vmemmap_optimizable case.
> Muchun,
> Thanks.