[PATCH] mm/page_alloc: skip shuffling and reporting for no-lock frees
From: Karl Mehltretter
Date: Mon Oct 05 2026 - 02:35:41 EST
free_pages_nolock() may run in NMI context or while a raw spinlock is
held. When its zone trylock succeeds, __free_one_page() still reaches
optional work written for ordinary frees.
shuffle_pick_tail() can refill its random-bit cache through
get_random_u64(), which takes a per-CPU local lock.
page_reporting_notify_free() can queue delayed work. Both operations can
enter locking outside the trylock-or-defer mechanism used by no-lock
frees.
An instrumented PREEMPT_RT lockdep run forced a random-bit cache refill
while calling free_pages_nolock() under a raw spinlock. The unpatched
allocator reported:
BUG: Invalid wait context
The attempted lock was batched_entropy_u64.lock in get_random_u64().
Lockdep showed the raw test lock and zone->lock both held.
Skip allocator shuffling and page-reporting notification for FPI_NOLOCK.
The freed page remains on the buddy list and can be collected by a later
reporting pass.
Trees from v6.15 through v7.2 call this flag FPI_TRYLOCK. Use
FPI_TRYLOCK in place of FPI_NOLOCK when backporting this change.
Fixes: 8c57b687e833 ("mm, bpf: Introduce free_pages_nolock()")
Cc: <stable@xxxxxxxxxxxxxxx> # see patch description, needs adjustments for <= 7.2
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@xxxxxxxxx>
---
This was found while auditing the no-lock paths discussed at:
https://lore.kernel.org/r/20260919171443.90512-1-kmehltretter@xxxxxxxxx
Tested on base e767a4ea70a3 in four-vCPU x86-64 QEMU TCG with
PREEMPT_RT, PROVE_LOCKING, PROVE_RAW_LOCK_NESTING,
SHUFFLE_PAGE_ALLOCATOR, PAGE_REPORTING and page_alloc.shuffle=1.
A temporary initcall registered page reporting and freed 65 maximum-order
allocations with free_pages_nolock() under a raw spinlock. This exhausts
shuffle_pick_tail()'s cache and forces a refill.
unpatched patched
free_pages_nolock() calls 65 65
__free_one_page() calls 130 130
shuffle_pick_tail() calls 65 0
page_reporting_notify_free() calls 130 0
reporting work pending 1 0
invalid-wait-context reports 1 0
Maximum-order splitting accounts for the 130 __free_one_page() calls.
Both guests powered off. The patched guest emitted no diagnostics. This
forced test does not measure normal frequency.
A second A/B used no test-only MM caller. A BPF arena program attached to
mm_page_free_batched requested 2,048 pages from a memcg-limited map. Partial
failures rolled allocated pages back through free_pages_nolock() while the
tracepoint held the per-CPU page-list lock. A virtio-balloon device provided
page reporting.
unpatched patched
arena allocation attempts 70 70
arena allocation failures 70 70
shuffle calls from rollback 9 0
get_random_u64() calls from shuffle 1 0
reporting notifications 12 0
delayed-work queues 5 0
Both guests powered off without a diagnostic. This confirms that an in-tree
caller reaches the optional operations. It did not reproduce a deadlock or
lockdep warning.
The upstream form applies to v7.3-rc1, mainline and next-20261002. The
FPI_TRYLOCK backport applies to v6.15 through v7.2. mm/page_alloc.o builds
on v6.18 and v7.2.7. Stable backports were not runtime tested.
mm/page_alloc.c | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 12fac9084c48..3e21dc90b858 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -89,7 +89,10 @@ typedef int __bitwise fpi_t;
*/
#define FPI_TO_TAIL ((__force fpi_t)BIT(1))
-/* Free the page without taking locks. Rely on trylock only. */
+/*
+ * Request the no-lock free path. Use trylocks and skip optional operations
+ * which can take other locks or queue work.
+ */
#define FPI_NOLOCK ((__force fpi_t)BIT(2))
/* free_pages_prepare() has already been called for page(s) being freed. */
@@ -1013,7 +1016,7 @@ static inline void __free_one_page(struct page *page,
if (fpi_flags & FPI_TO_TAIL)
to_tail = true;
- else if (is_shuffle_order(order))
+ else if (!(fpi_flags & FPI_NOLOCK) && is_shuffle_order(order))
to_tail = shuffle_pick_tail();
else
to_tail = buddy_merge_likely(pfn, buddy_pfn, page, order);
@@ -1021,7 +1024,7 @@ static inline void __free_one_page(struct page *page,
__add_to_free_list(page, zone, order, migratetype, to_tail);
/* Notify page reporting subsystem of freed page */
- if (!(fpi_flags & FPI_SKIP_REPORT_NOTIFY))
+ if (!(fpi_flags & (FPI_SKIP_REPORT_NOTIFY | FPI_NOLOCK)))
page_reporting_notify_free(order);
}
--
2.53.0