Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order

From: Kiryl Shutsemau

Date: Wed Sep 30 2026 - 09:49:49 EST


On Tue, Sep 29, 2026 at 07:39:41PM +0100, Harry Yoo wrote:
> On Tue, Sep 29, 2026 at 06:45:51PM +0100, Kiryl Shutsemau wrote:
> > From: "Kiryl Shutsemau (Meta)" <kas@xxxxxxxxxx>
> >
> > Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> > storm in defrag_mode"), direct reclaim and compaction for non-movable
> > requests under defrag_mode run at pageblock_order, to produce the whole
> > blocks that ALLOC_NOFRAGMENT needs.
>
> > The retry decisions that follow still use the request order.
>
> Indeed, good catch!
>
> > An order-0 request can therefore retry
> > indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
> >
> > - Reclaim at pageblock_order gives up after one pass as soon as a zone
> > looks compaction_ready(), and do_try_to_free_pages() then returns 1
> > even though nothing was reclaimed. It returns before the retry that
> > would reclaim memory.low-protected cgroups, so when most memory is
> > protected, the pass that did run finds next to nothing.
> >
> > - Compaction at pageblock_order fails or is deferred.
> >
> > - should_reclaim_retry() takes the reported progress as progress for
> > the order-0 request and resets no_progress_loops. The request
> > retries.
>
> Makes sense to me.
>
> > Order 1-3 requests loop the same way, and should_compact_retry() also
> > checks their pageblock_order compaction result against the request
> > order.
> >
> > On a production host (64G, defrag_mode, memory.low covering most of the
> > workload), 95% of direct reclaim runs were order-9 runs that returned 1
> > with nothing reclaimed, at up to 60k runs per second. Across ~200M
> > should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> > The spinning allocations were SLUB slab refills for inode and dentry
> > caches. The time spent registers as memory pressure, and pressure-based
> > OOM killing takes down both workloads and system services.
> >
> > Treat promoted requests like costly orders:
> >
> > - Reclaim progress does not reset no_progress_loops for them.
> >
> > - should_compact_retry() checks the compaction result at the promoted
> > order. It does not retry COMPACT_SKIPPED, since the request can fall
> > back, and it does not escalate compaction to COMPACT_PRIO_SYNC_FULL.
> >
> > When the fallback is taken, reset the retry counters, so that the
> > fallback attempt gets a full retry budget before the OOM killer is
> > considered.
> >
> > In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
> >
> > before after
> > should_reclaim_retry() calls 63M 293k
> > peak memory pressure (PSI some avg10) 99% 12%
> >
> > File creation runs 5.7x faster.
> >
> > Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
> > Cc: <stable@xxxxxxxxxxxxxxx>
> > Assisted-by: LLM
> > Signed-off-by: Kiryl Shutsemau (Meta) <kas@xxxxxxxxxx>
> > ---
> > mm/page_alloc.c | 85 ++++++++++++++++++++++++++++++++-----------------
> > 1 file changed, 56 insertions(+), 29 deletions(-)
> >
> > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > index 12fac9084c48..608487672d93 100644
> > --- a/mm/page_alloc.c
> > +++ b/mm/page_alloc.c
> > @@ -4127,6 +4127,31 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> > return page;
> > }
> >
> > +/*
> > + * If fallbacks are not permitted (defrag_mode), we either need to
> > + * reclaim space in a block of matching type, or clear out an entire
> > + * block to allow __rmqueue_claim() to convert.
> > + *
> > + * Reclaim by itself is primarily freeing space in movable blocks,
> > + * since that's where the LRU pages live. So this works for movable
> > + * requests, but not for others.
> > + *
> > + * For those, promote the order of reclaim and compaction to help make
> > + * blocks, instead of spinning in reclaim alone unproductively. Retry
> > + * decisions based on the outcome of that work - reclaim progress and
> > + * compaction results - must account for the promotion as well, see
> > + * should_reclaim_retry() and should_compact_retry().
> > + */
> > +static inline unsigned int nofrag_promote_order(unsigned int order,
> > + unsigned int alloc_flags,
> > + const struct alloc_context *ac)
> > +{
> > + if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
> > + return max(order, pageblock_order);
> > +
> > + return order;
> > +}
>
> I think we should start distinguishing order and compact/reclaim_order
> in __alloc_pages_slowpath(). Silently overriding it makes it harder to
> follow and easy to make a mistake.

Agreed, four callers recomputing the same thing is asking for a
mismatch.

I would rather not grow this patch, it has to go to stable.

I will look into a cleanup on top: __alloc_pages_slowpath() computes the
promoted order once per iteration and passes it to direct
reclaim/compaction and the two retry helpers next to the request order,
so the helpers stop knowing about defrag_mode.

> > @@ -4299,7 +4316,7 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
> > /*
> > * Compaction failed. Retry with increasing priority.
> > */
> > - min_priority = (order > PAGE_ALLOC_COSTLY_ORDER) ?
> > + min_priority = (compact_order > PAGE_ALLOC_COSTLY_ORDER) ?
> > MIN_COMPACT_COSTLY_PRIORITY : MIN_COMPACT_PRIORITY;
>
> This would change how hard we try to compact with defrag_mode in direct
> compaction as it won't try compaction with MIN_COMPACT_PRIORITY anymore.
>
> It doesn't make much sense to change that as part of this fix?

It is a choice between compacting harder and falling back, which
fragments a block.

It is a judgement call on what defrag_mode means.

It would also mean that order-0 allocation request promoted to pageblock
can trigger SYNC_FULL compaction. I cannot say I understand the
implications. Will give it a try with the reproducer.

Johannes, Vlastimil, any comments here?

--
Kiryl Shutsemau / Kirill A. Shutemov