Re: [PATCH mm-hotfixes v2 1/2] mm/huge_memory: fix huge_zero_pfn race
From: David Hildenbrand (Arm)
Date: Thu Jul 30 2026 - 10:44:18 EST
On 7/30/26 12:55, Lorenzo Stoakes (ARM) wrote:
> If !CONFIG_PERSISTENT_HUGE_ZERO_FOLIO, the huge_zero_folio is refcounted by
> huge_zero_refcount and returned by mm_get_huge_zero_folio().
>
> When the caller is done with the huge zero page, its reference count is
> decremented. Only a shrinker can set the reference count to zero.
>
> A race can unfortunately occur between a shrinker decrementing the
> reference count to zero and a concurrent page fault.
>
> This is because shrink_huge_zero_folio_scan() might, if very
> unlucky, be preempted between setting huge_zero_refcount to zero
> and writing an invalid value.
>
> During this time get_huge_zero_folio() could write to huge_zero_pfn
> before shrink_huge_zero_folio_scan() resumes.
>
> In this event the huge zero folio will be persistently
> misidentified causing the THP code path to be entered
> inappropriately for the huge zero folio:
>
> CPU 0 CPU 1
> =======================================|=================================
> shrink_huge_zero_folio_scan() |
> atomic_cmpxchg() sets refcount to 0 |
> xchg() sets huge_zero_folio to NULL | get_huge_zero_folio()
> | | atomic_inc_not_zero() -> zero
> preempted for a long time | Allocate new huge zero folio
> | | Write valid huge_zero_folio
> v | Write valid huge_zero_pfn
> Overwrite huge_zero_pfn with ~0UL <--- Invalid overwrite!
>
> This results in is_huge_zero_pfn() and is_huge_zero_pmd() incorrectly
> returning false for a huge zero page which could result in issues like the
> huge zero folio being incorrectly split.
>
> Note that the issue is with huge_zero_pfn not huge_zero_folio, as
> get_huge_zero_folio() uses cmpxchg() gated on huge_zero_folio being NULL
> with a retry loop and shrink_huge_zero_folio_scan() uses xchg() to set
> huge_zero_folio.
>
> Fix the issue by introducing a spinlock, huge_zero_lock, to prevent
> concurrent write of huge_zero_folio, huge_zero_pfn and huge_zero_refcount.
>
> There needs to be significant care taken here to ensure correctness:
>
> The fast path in get_huge_zero_folio() uses atomic_inc_not_zero(), which is
> outside of the critical section, and means huge zero allocation is gated on
> zero huge_zero_refcount.
>
> The fast path doesn't use huge_zero_lock, so the critical section is
> irrelevant to it.
>
> So invariants are required - huge_zero_refcount MUST:
>
> * Only be set in the huge_zero_lock critical section to ensure
> serialisation of huge_zero_pfn, huge_zero_folio and huge_zero_refcount
> writes.
>
> * Be set non-zero only AFTER huge_zero_[pfn, folio] are set to valid values
> so installation of the huge zero folio on read page fault ensures
> concurrent is_huge_zero_*() calls correctly identify the huge zero folio.
>
> * Be set zero only BEFORE huge_zero_[pfn, folio] are set to NULL and ~0UL
> respectively, and atomically.
>
> Establish these by:
>
> * Only setting huge_zero_refcount to zero or an absolute value in the
> huge_zero_lock critical section in get_huge_zero_folio() and
> shrink_huge_zero_folio_scan(), and always updating atomically there
> and elsewhere.
>
> * Using atomic_set_release(&huge_zero_refcount) in get_huge_zero_folio()
> after huge_zero_[pfn, folio] are set. This is paired with
> atomic_inc_not_zero() to ensure atomic_inc_not_zero() only observes a
> non-zero value if huge_zero_[pfn, folio] are set.
>
> * Using atomic_cmpxchg() in shrink_huge_zero_folio_scan() (as before) to
> ensure that it is set zero only when equal to 1 and set atomically.
>
> * atomic_cmpxchg() being fully ordered ensures this is done prior to
> huge_zero_[folio, pfn] being set to NULL and ~0UL respectively.
>
> Eliminate the retry loop in get_huge_zero_folio() as the atomic_cmpxchg()
> in shrink_huge_zero_folio_scan() is now performed under the lock, and
> replace with an equally locked atomic_inc() to set the reference count
> should the caller be raced on huge zero folio installation.
>
> folio_put() naturally implies a full memory barrier so its ordering is
> maintained correctly.
>
> The huge zero folio also cannot be released except when the shrinker does
> so as it is non-LRU and non-rmappable.
>
> Note that only the huge zero shrinker (via shrink_huge_zero_folio_scan())
> can actually set huge_zero_refcount to zero, which is the count of mm's
> which have at least one huge zero folio installed plus one shrinker pin.
>
> Additionally convert a BUG_ON() to a VM_WARN_ON_ONCE().
>
> Suggested-by: David Hildenbrand (Arm) <david@xxxxxxxxxx>
> Reported-by: Hengbin Zhang <uqbarz@xxxxxxxxx>
> Closes: https://lore.kernel.org/linux-mm/20260727154001.4102341-1-uqbarz@xxxxxxxxx/
> Fixes: 3b77e8c8cde5 ("mm/thp: make is_huge_zero_pmd() safe and quicker")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Lorenzo Stoakes (ARM) <ljs@xxxxxxxxxx>
> ---
Thanks!
Acked-by: David Hildenbrand (Arm) <david@xxxxxxxxxx>
--
Cheers,
David