Re: [PATCH RFC] mm/huge_memory: eliminate reference-counted huge zero folio
From: Lorenzo Stoakes (ARM)
Date: Thu Jul 30 2026 - 06:29:59 EST
On Tue, Jul 28, 2026 at 09:16:08PM +0200, David Hildenbrand (Arm) wrote:
> On 7/28/26 15:06, Lorenzo Stoakes (ARM) wrote:
> > The reference-counted huge zero folio is problematic - it's prone to subtle
> > races, it can introduce allocation latency and it is confusing and
> > complicated code.
>
> Just a note that the subject is misleading: the huge zero folio is always
> refcounted (e.g., GUP must take references), it's just that the persistent huge
> zero folio never drops the refcount to 0 to free it.
I meant huge_zero_refcount not folio refcount :) but maybe could have been
clearer.
>
> What you probably mean is "mm/huge_memory: eliminate dynamically allocated huge
> zero folio".
Well that's not really getting to the core of it, the problem is the
huge_zero_refcount meaning it can come and go and you have races.
But agreed refcount is imprecise + clearly confusing vs. folio refcount given
you immediately assuemd that.
So 'mm/huge_memory: eliminiate dynmaically alloc'd/freed huge zero folio' maybe?
>
> If we go down that path here, we can also teach GUP to not take a reference on
> the huge zero folio anymore on the FOLL_PIN path, similar to how we handle the
> zero folio.
>
> See the occurrences of is_zero_folio() in mm/gup.c.
Ah yeah. Can do that on a respin, thanks.
>
> The reason we do it for now for FOLL_PIN is that a remote pin might outlive the
> process, meaning the huge folio could otherwise get freed by the shrinker even
> though someone still holds a reference. Bad. So that would no longer apply if
> the huge zero folio can no longer get freed.
Incredible that we have to think about this...
>
>
> I recall there was a discussion (dropping dynamically allocated huge zero folio)
> on this topic while upstreaming the persistent huge zero folio.
>
> I don't have time right now to dig, but it would be worthwhile to look at that
> history if there was a good reason to keep it buried in that discussion.
Maybe this:
https://lore.kernel.org/all/20250811084113.647267-4-kernel@xxxxxxxxxxxxxxxx/T/#m4100481232bae8a7b661e8f6e86ecc74532c69a4
Generally seems 'smaller systems blah blah'. But smaller systems that set
CONFIG_THP get bigger reserves etc. so the argument doesn't really work.
I suppose they might default CONFIG_THP on but not enable THP but at that point
are you really so tiny a system that 2 MiB matters?
Obviously it's more problematic for 64 KiB page size, but then if you have THP
at all your reserves will already be problematic, and any read fault on a THP
page gets you the exact same issue.
>
> --
> Cheers,
>
> David
Cheers, Lorenzo