Re: [PATCH v5 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization
From: Qi Zheng
Date: Tue Aug 25 2026 - 23:26:13 EST
On 8/26/26 10:38 AM, Muchun Song wrote:
On Aug 25, 2026, at 21:12, Qi Zheng <qi.zheng@xxxxxxxxx> wrote:
On 8/25/26 4:46 PM, Muchun Song wrote:
HugeTLB bootmem vmemmap optimization still carries its own early setup
path, including pre-populating optimized mappings before the generic
sparse-vmemmap code runs.
Now that the section-based vmemmap optimization can derive HugeTLB
vmemmap deduplication from section metadata, HugeTLB only needs to mark
the bootmem huge page range with the appropriate order. The generic
sparse-vmemmap population path can then allocate and map the shared tail
vmemmap pages without any HugeTLB-specific early population code.
Do that by setting the section order when a bootmem huge page is
allocated and dropping the dedicated pre-HVO helpers and related
special-casing.
This removes duplicate early setup logic and switches HugeTLB to the
section-based vmemmap optimization path.
Signed-off-by: Muchun Song <songmuchun@xxxxxxxxxxxxx>
Acked-by: Mike Rapoport (Microsoft) <rppt@xxxxxxxxxx>
---
v3:
- Use the order-based helper for the bootmem vmemmap-optimized check
v2:
- Collect Acked-by from Mike Rapoport
---
include/linux/hugetlb.h | 1 -
include/linux/mm.h | 3 --
mm/hugetlb.c | 30 ++------------
mm/hugetlb_vmemmap.c | 90 +++--------------------------------------
mm/hugetlb_vmemmap.h | 14 +++----
mm/sparse-vmemmap.c | 31 --------------
mm/sparse.h | 27 +++++++++++++
7 files changed, 42 insertions(+), 154 deletions(-)
diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
index 16c4c4caa126..fe28f98e1b22 100644
--- a/include/linux/hugetlb.h
+++ b/include/linux/hugetlb.h
@@ -171,7 +171,6 @@ struct address_space *hugetlb_folio_mapping_lock_write(struct folio *folio);
extern int movable_gigantic_pages __read_mostly;
extern int sysctl_hugetlb_shm_group __read_mostly;
-extern struct list_head huge_boot_pages[MAX_NUMNODES];
void hugetlb_bootmem_struct_page_init(void);
void hugetlb_bootmem_alloc(void);
diff --git a/include/linux/mm.h b/include/linux/mm.h
index dd09c438fa23..441bd39eab73 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -5159,9 +5159,6 @@ int vmemmap_populate_hugepages(unsigned long start, unsigned long end,
int node, struct vmem_altmap *altmap);
int vmemmap_populate(unsigned long start, unsigned long end, int node,
struct vmem_altmap *altmap);
-int vmemmap_populate_hvo(unsigned long start, unsigned long end,
- unsigned int order, struct zone *zone,
- unsigned long headsize);
void vmemmap_wrprotect_hvo(unsigned long start, unsigned long end, int node,
unsigned long headsize);
void vmemmap_populate_print_last(void);
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index 04e6c4244cd6..fbb0c83bea79 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -52,6 +52,7 @@
#include "hugetlb_cma.h"
#include "hugetlb_internal.h"
#include "mm_init.h"
+#include "sparse.h"
#include <linux/page-isolation.h>
int hugetlb_max_hstate __read_mostly;
@@ -59,7 +60,7 @@ unsigned int default_hstate_idx;
struct hstate hstates[HUGE_MAX_HSTATE];
__initdata nodemask_t hugetlb_bootmem_nodes;
-__initdata struct list_head huge_boot_pages[MAX_NUMNODES];
+static struct list_head huge_boot_pages[MAX_NUMNODES] __initdata;
/*
* Due to ordering constraints across the init code for various
@@ -3139,6 +3140,7 @@ static bool __init alloc_bootmem_huge_page(struct hstate *h, int nid)
} else {
list_add_tail(&m->list, &huge_boot_pages[nid]);
m->flags |= HUGE_BOOTMEM_ZONES_VALID;
+ hugetlb_vmemmap_optimize_bootmem_page(m);
/*
* Only initialize the head struct page in memmap_init_reserved_pages,
* rest of the struct pages will be initialized by the HugeTLB
@@ -3299,6 +3301,7 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid)
* this folio.
*/
folio_set_hugetlb_vmemmap_optimized(folio);
+ section_set_order_range(folio_pfn(folio), folio_nr_pages(folio), 0);
So section->order is only used during initialization. Is it ever
accessed later at runtime? If not, can we just skip zeroing it out?
You have keenly noticed a detail. This was actually done deliberately,
because in patch 14, HUGE_BOOTMEM_HVO was removed and replaced with
section->order. The ->order field may store a value that is less than
VMEMMAP_OPTIMIZATION_MIN_ORDER, so clearing it to 0 here is to prevent
potential issues with this memory region during the hotplug/hotremove
process in the future (in my future series).
However, is there also a way to avoid clearing it to 0? There is. We
could add an extra check in hugetlb_vmemmap_optimize_bootmem_page to
only call section_set_order_range when the current hstate->order is
greater than or equal to VMEMMAP_OPTIMIZATION_MIN_ORDER. I just feel
that this would add a bit more code. And then during the boot phase,
doing one extra zeroing-out doesn't introduce much overhead anyway.
Therefore, I chose the simpler implementation approach.
Got it. Thanks for the detailed explanation!