Re: [PATCH mm-hotfixes 2/2] mm/huge_memory: fix huge_zero_pfn race

From: David Hildenbrand (Arm)

Date: Thu Jul 30 2026 - 05:44:48 EST


[...]

> So invariants are required - huge_zero_refcount MUST:
>
> * Only be set in the huge_zero_lock critical section to ensure
> serialisation of huge_zero_pfn, huge_zero_folio and huge_zero_refcount
> writes.
>
> * Be set non-zero only AFTER huge_zero_[pfn, folio] are set to valid values
> so installation of the huge zero folio on read page fault ensures
> concurrent is_huge_zero_*() calls correctly identify the huge zero folio.
>
> * Be set zero only BEFORE huge_zero_[pfn, folio] are set to NULL and ~0UL
> respectively, and atomically.
>
> Establish these by:
>
> * Only updating huge_zero_refcount in the huge_zero_lock critical section
> in get_huge_zero_folio() and shrink_huge_zero_folio_scan().

That is imprecise. huge_zero_refcount is updated (incremented) outside of
huge_zero_lock in get_huge_zero_folio().

>
> * Using atomic_set_release(&huge_zero_refcount) in get_huge_zero_folio()
> after huge_zero_[pfn, folio] are set. This is paired with
> atomic_inc_not_zero() to ensure atomic_inc_not_zero() only observes a
> non-zero value if huge_zero_[pfn, folio] are set.

That makes sense, yes.

>
> * Using atomic_cmpxchg() in shrink_huge_zero_folio_scan() to ensure that it
> is set zero only when equal to 1 and set atomically.

What we already do, yes.

>
> * atomic_cmpxchg() being fully ordered ensures this is done prior to
> huge_zero_[folio, pfn] being set to NULL and ~0UL respectively.

Which is always the case for atomic_cmpxchg() I think, yes.

[...]

>
> @@ -270,7 +272,8 @@ void mm_put_huge_zero_folio(struct mm_struct *mm)
> static bool get_huge_zero_folio(void)
> {
> struct folio *zero_folio;
> -retry:
> +
> + /* Paired with atomic_set_release(). */
> if (likely(atomic_inc_not_zero(&huge_zero_refcount)))
> return true;
>
> @@ -278,17 +281,21 @@ static bool get_huge_zero_folio(void)
> if (unlikely(!zero_folio))
> return false;
>
> - preempt_disable();
> - if (cmpxchg(&huge_zero_folio, NULL, zero_folio)) {
> - preempt_enable();
> + /* Paired with critical section in shrink_huge_zero_folio_scan(). */
> + spin_lock(&huge_zero_lock);
> + if (huge_zero_folio) {
> + /* Somebody else already installed it. */
> + atomic_inc(&huge_zero_refcount);

You replace the retry loop by an atomic_inc() here.

That works now by moving the atomic_cmpxchg() in shrink_huge_zero_folio_scan()
under the lock as well, so it cannot go away concurrently.

Worth spelling that out in the patch description (unless I missed it).

> + spin_unlock(&huge_zero_lock);
> folio_put(zero_folio);
> - goto retry;
> + return true;
> }
> + WRITE_ONCE(huge_zero_folio, zero_folio);
> WRITE_ONCE(huge_zero_pfn, folio_pfn(zero_folio));
> + /* Paired with atomic_inc_not_zero(). +1 for shrinker pin. */
> + atomic_set_release(&huge_zero_refcount, 2);

Ah, that must now go into the lock as well, otherwise the concurrent
atomic_inc() would be problematic as well.

> + spin_unlock(&huge_zero_lock);
>
> - /* We take additional reference here. It will be put back by shrinker */
> - atomic_set(&huge_zero_refcount, 2);
> - preempt_enable();
> count_vm_event(THP_ZERO_PAGE_ALLOC);
> return true;
> }
> @@ -312,15 +319,22 @@ static unsigned long shrink_huge_zero_folio_count(struct shrinker *shrink,
> static unsigned long shrink_huge_zero_folio_scan(struct shrinker *shrink,
> struct shrink_control *sc)
> {
> - if (atomic_cmpxchg(&huge_zero_refcount, 1, 0) == 1) {
> - struct folio *zero_folio = xchg(&huge_zero_folio, NULL);
> - BUG_ON(zero_folio == NULL);
> + struct folio *zero_folio;
> +
> + /* Paired with critical section in get_huge_zero_folio(). */
> + scoped_guard(spinlock, &huge_zero_lock) {
> + /* Paired with atomic_inc_not_zero() in get_huge_zero_folio(). */
> + if (atomic_cmpxchg(&huge_zero_refcount, 1, 0) != 1)
> + return 0;
> +
> + zero_folio = huge_zero_folio;
> + VM_WARN_ON_ONCE(!huge_zero_folio);
> + WRITE_ONCE(huge_zero_folio, NULL);
> WRITE_ONCE(huge_zero_pfn, HUGE_ZERO_UNSET_PFN);
> - folio_put(zero_folio);
> - return HPAGE_PMD_NR;
> }

One important point: The folio can't get freed and reallocated before we updated
huge_zero_folio/huge_zero_pfn I think.

The folio_put() should involve a full memory barrier (atomic RMW) and not get
reordered.

So we cannot misclassify the folio in new context as still being the zero folio.

LGTM, thanks!

--
Cheers,

David