Re: [PATCH v5 00/12] mm: Switch device DAX to section-based vmemmap optimization

From: Muchun Song

Date: Sun Sep 27 2026 - 06:52:44 EST




> On Sep 27, 2026, at 13:51, Andrew Morton <akpm@xxxxxxxxxxxxxxxxxxxx> wrote:
>
> On Sun, 27 Sep 2026 10:54:29 +0800 Muchun Song <songmuchun@xxxxxxxxxxxxx> wrote:
>
>> After the HugeTLB conversion, optimized vmemmap state is described by
>> the memory section and the sparse-vmemmap population path can allocate or
>> reuse shared tail vmemmap pages based on that metadata. Device DAX still
>> uses the older DAX-specific population model, including a separate tail
>> vmemmap page reservation and architecture-specific logic to locate or
>> populate reusable tail pages.
>>
>> This series makes device DAX use the same section-based model. Device DAX
>> records the compound page order from pgmap->vmemmap_shift in section
>> metadata before vmemmap population, uses the common per-zone shared tail
>> vmemmap page, and drops the extra reserved tail page. The powerpc radix
>> path is updated to use the same shared tail-page helper, so the generic
>> and powerpc DAX paths follow the same reservation model.
>
> Thanks, I've updated mm-unstable to this version.

Thanks.

>
> Sashiko asked a thing:
> https://sashiko.dev/#/patchset/20260927025441.741633-1-songmuchun@xxxxxxxxxxxxx

Sashiko said page->refcount can overflow by incrementing it over 2.14 billion
times when mapping more than **524 TB** of DEV-DAX memory on a single NUMA
node, where the pages share the same node, order, and zone.

I am not aware of any practical hardware configuration approaching this
topology today.

Handling that theoretical limit would add non-trivial lifetime or
architecture-specific teardown complexity. Without a concrete hardware
requirement, I prefer not to over-engineer the current series. We can revisit
it when such a system or use case becomes realistic.

Thanks.