Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order

From: Kiryl Shutsemau

Date: Wed Sep 30 2026 - 08:41:42 EST


On Tue, Sep 29, 2026 at 12:58:16PM -0700, Andrew Morton wrote:
> On Tue, 29 Sep 2026 18:45:51 +0100 Kiryl Shutsemau <kirill@xxxxxxxxxxxxx> wrote:
>
> > From: "Kiryl Shutsemau (Meta)" <kas@xxxxxxxxxx>
> >
> > Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> > storm in defrag_mode"), direct reclaim and compaction for non-movable
> > requests under defrag_mode run at pageblock_order, to produce the whole
> > blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
> > still use the request order. An order-0 request can therefore retry
> > indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
>
> 7e8756d7ad22 is new in 7.3-rcX, so no cc:stable needed.

The commit itself has cc:stable. And it is in 6.18.50 and 7.2.4.
The fix needs to follow it there.

> > - Reclaim at pageblock_order gives up after one pass as soon as a zone
> > looks compaction_ready(), and do_try_to_free_pages() then returns 1
> > even though nothing was reclaimed. It returns before the retry that
> > would reclaim memory.low-protected cgroups, so when most memory is
> > protected, the pass that did run finds next to nothing.
> >
> > - Compaction at pageblock_order fails or is deferred.
> >
> > - should_reclaim_retry() takes the reported progress as progress for
> > the order-0 request and resets no_progress_loops. The request
> > retries.
> >
> > Order 1-3 requests loop the same way, and should_compact_retry() also
> > checks their pageblock_order compaction result against the request
> > order.
> >
> > On a production host (64G, defrag_mode, memory.low covering most of the
> > workload), 95% of direct reclaim runs were order-9 runs that returned 1
> > with nothing reclaimed, at up to 60k runs per second. Across ~200M
> > should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> > The spinning allocations were SLUB slab refills for inode and dentry
> > caches. The time spent registers as memory pressure, and pressure-based
> > OOM killing takes down both workloads and system services.
>
> A production host running latest -rc?

We are running 7.1 with 7e8756d7ad22 backported. We found defrag mode
crucial if we want to use large folios in page cache.

But nobody besides us seems to test it :/

Can we make it default pretty please? :P

--
Kiryl Shutsemau / Kirill A. Shutemov