Re: [PATCH] riscv/mm: use physical alignment for vmemmap_start_pfn
From: Muchun Song
Date: Mon Jul 20 2026 - 05:33:51 EST
> On Jul 16, 2026, at 19:53, Jiakai Xu <xujiakai2025@xxxxxxxxxxx> wrote:
>
> RISC-V computes vmemmap_start_pfn by rounding phys_ram_base down to
> VMEMMAP_ADDR_ALIGN. That alignment must therefore be expressed in the
> physical-address domain.
>
> Commit 476849b0fba4 ("riscv/mm: align vmemmap to maximal folio size")
> attempted to account for the maximal folio alignment by feeding
> MAX_FOLIO_VMEMMAP_ALIGN directly into VMEMMAP_ADDR_ALIGN. However,
> MAX_FOLIO_VMEMMAP_ALIGN is measured in bytes of struct page storage,
> whereas VMEMMAP_ADDR_ALIGN is used to align a physical address.
>
> The mask-based compound_info encoding requires pfn_to_page(0) to be
> naturally aligned to MAX_FOLIO_VMEMMAP_ALIGN. Commit 9f94db4c7eaa
> ("mm/sparse: check memmap alignment for compound_info_has_mask()")
> added a check for that requirement and exposed the unit mismatch on
> systems such as QEMU virt, where the DRAM base is not aligned to
> MAX_FOLIO_NR_PAGES * PAGE_SIZE.
>
> Convert MAX_FOLIO_VMEMMAP_ALIGN to the equivalent physical alignment
> before using it in VMEMMAP_ADDR_ALIGN. This keeps the existing
> round_down() logic while making the resulting vmemmap base satisfy the
> mask-alignment requirement.
>
> Fixes: 476849b0fba4 ("riscv/mm: align vmemmap to maximal folio size")
> Signed-off-by: Jiakai Xu <xujiakai2025@xxxxxxxxxxx>
> Assisted-by: YuanSheng:DeepSeek-V4-Flash
I've always wondered why RISC-V uses a complex logic to calculate the
mapping relationship between vmemmap and PFN. We could easily follow the
x86 approach to make it much simpler.
I previously ran an experiment and found that the system boots perfectly
fine using the following diff code. Granted, I wrote this quite a while ago,
so it might not be directly compatible with the current codebase, but it
should work with some minor adjustments.
In this scenario, the overall processing logic would become significantly
more straightforward, and handling alignment would be much easier as well.
I'm not sure if the following direction is correct (Perhaps this was a
deliberate choice by RISC-V; I'm just not aware of the background behind it),
but from my perspective, it makes the code much simpler. We might need to
reach out to the relevant RISC-V maintainers to confirm whether this aligns
with what's expected.
Thanks,
Muchun
diff --git a/arch/riscv/include/asm/page.h b/arch/riscv/include/asm/page.h
index ffe213ad65a4..f2f67007f482 100644
--- a/arch/riscv/include/asm/page.h
+++ b/arch/riscv/include/asm/page.h
@@ -119,7 +119,6 @@ struct kernel_mapping {
extern struct kernel_mapping kernel_map;
extern phys_addr_t phys_ram_base;
-extern unsigned long vmemmap_start_pfn;
#define is_kernel_mapping(x) \
((x) >= kernel_map.virt_addr && (x) < (kernel_map.virt_addr + kernel_map.size))
diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
index 8bd36ac842eb..067c87062bcf 100644
--- a/arch/riscv/include/asm/pgtable.h
+++ b/arch/riscv/include/asm/pgtable.h
@@ -91,7 +91,7 @@
* Define vmemmap for pfn_to_page & page_to_pfn calls. Needed if kernel
* is configured with CONFIG_SPARSEMEM_VMEMMAP enabled.
*/
-#define vmemmap ((struct page *)VMEMMAP_START - vmemmap_start_pfn)
+#define vmemmap ((struct page *)VMEMMAP_START)
#define PCI_IO_SIZE SZ_16M
#define PCI_IO_END VMEMMAP_START
diff --git a/arch/riscv/mm/init.c b/arch/riscv/mm/init.c
index addb8a9305be..8f86a6d07a53 100644
--- a/arch/riscv/mm/init.c
+++ b/arch/riscv/mm/init.c
@@ -62,13 +62,6 @@ EXPORT_SYMBOL(pgtable_l5_enabled);
phys_addr_t phys_ram_base __ro_after_init;
EXPORT_SYMBOL(phys_ram_base);
-#ifdef CONFIG_SPARSEMEM_VMEMMAP
-#define VMEMMAP_ADDR_ALIGN (1ULL << SECTION_SIZE_BITS)
-
-unsigned long vmemmap_start_pfn __ro_after_init;
-EXPORT_SYMBOL(vmemmap_start_pfn);
-#endif
-
unsigned long empty_zero_page[PAGE_SIZE / sizeof(unsigned long)]
__page_aligned_bss;
EXPORT_SYMBOL(empty_zero_page);
@@ -246,12 +239,8 @@ static void __init setup_bootmem(void)
* Make sure we align the start of the memory on a PMD boundary so that
* at worst, we map the linear mapping with PMD mappings.
*/
- if (!IS_ENABLED(CONFIG_XIP_KERNEL)) {
+ if (!IS_ENABLED(CONFIG_XIP_KERNEL))
phys_ram_base = memblock_start_of_DRAM() & PMD_MASK;
-#ifdef CONFIG_SPARSEMEM_VMEMMAP
- vmemmap_start_pfn = round_down(phys_ram_base, VMEMMAP_ADDR_ALIGN) >> PAGE_SHIFT;
-#endif
- }
/*
* In 64-bit, any use of __va/__pa before this point is wrong as we
@@ -1121,9 +1110,6 @@ asmlinkage void __init setup_vm(uintptr_t dtb_pa)
kernel_map.xiprom_sz = (uintptr_t)(&_exiprom) - (uintptr_t)(&_xiprom);
phys_ram_base = CONFIG_PHYS_RAM_BASE;
-#ifdef CONFIG_SPARSEMEM_VMEMMAP
- vmemmap_start_pfn = round_down(phys_ram_base, VMEMMAP_ADDR_ALIGN) >> PAGE_SHIFT;
-#endif
kernel_map.phys_addr = (uintptr_t)CONFIG_PHYS_RAM_BASE;
kernel_map.size = (uintptr_t)(&_end) - (uintptr_t)(&_start);