Re: [PATCH v3] mm/page_alloc: avoid direct compaction for costly __GFP_NORETRY allocations

From: Vlastimil Babka (SUSE)

Date: Wed Jul 22 2026 - 08:36:08 EST


On 7/14/26 14:02, Salvatore Dipietro wrote:
>
> Hi Johannes, Matthew,
>
> Thank you both for the alternative proposals. I've tested both
> approaches on the same test environment used for v3.
>
> Results (4 runs each):
>
> Config Avg TPS % vs Baseline
> --------------------------------------------------------
> baseline (no patch) 70,735 -
> Johannes' approach 156,908 +121.8%
> Matthew's approach 70,145 -0.8% (within noise)
>
>
> Johannes' approach (clearing __GFP_DIRECT_RECLAIM early in the
> slowpath for costly __GFP_NORETRY) delivers the same ~2.2x speedup
> as v3, as expected - it prevents the entire direct reclaim and
> compaction machinery from running for these opportunistic allocations.

AFAICS both your v3 and the implemented Johannes' suggestion effectively
make all the compact_first logic dead code - at least for the intended THP
use case based on costly_order. It will only continue to do something for
the non-movable order>0 case. So that's a significant change which will not
show up in your results, but might affect other workloads (I'd expect mainly
by having fewer THPs).

> Matthew's filemap.c approach does not help in this workload. The
> reason is that the first allocation attempt at max order still carries
> __GFP_DIRECT_RECLAIM and enters the slowpath with direct compaction
> enabled. The __GFP_DIRECT_RECLAIM clearing only takes effect for
> subsequent lower-order attempts in the fallback loop, but the costly
> compaction has already executed on the first try.
>
> Let me know if you have any other variant you want me to test,
> or if I should prepare a v4 based on Johannes' suggestion.
>
> This is what I tested for Johannes' approach:
>
>
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index a63733dac659..6e960c969e67 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -4733,10 +4733,10 @@ static inline struct page *
> __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
> struct alloc_context *ac)
> {
> - bool can_direct_reclaim = gfp_mask & __GFP_DIRECT_RECLAIM;
> - bool can_compact = can_direct_reclaim && gfp_compaction_allowed(gfp_mask);
> - bool nofail = gfp_mask & __GFP_NOFAIL;
> const bool costly_order = order > PAGE_ALLOC_COSTLY_ORDER;
> + bool can_direct_reclaim;
> + bool can_compact;
> + bool nofail;
> struct page *page = NULL;
> unsigned int alloc_flags;
> unsigned long did_some_progress;
> @@ -4751,6 +4751,20 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
> bool can_retry_reserves = true;
> unsigned long alloc_start_time = jiffies;
>
> + /*
> + * Costly __GFP_NORETRY allocations are opportunistic: the caller
> + * can fall back to smaller orders. Don't stall on direct reclaim
> + * or compaction; clearing __GFP_DIRECT_RECLAIM makes the entire
> + * slowpath treat this as a non-blocking request. kswapd will wake
> + * kcompactd as needed for background defragmentation.
> + */
> + if (costly_order && (gfp_mask & __GFP_NORETRY))
> + gfp_mask &= ~__GFP_DIRECT_RECLAIM;
> +
> + can_direct_reclaim = gfp_mask & __GFP_DIRECT_RECLAIM;
> + can_compact = can_direct_reclaim && gfp_compaction_allowed(gfp_mask);
> + nofail = gfp_mask & __GFP_NOFAIL;
> +
> if (unlikely(nofail)) {
> /*
> * Also we don't support __GFP_NOFAIL without __GFP_DIRECT_RECLAIM,
>
>
> Thanks,
> Salvatore
>
>
>
>
>
> AMAZON DEVELOPMENT CENTER ITALY SRL, viale Monte Grappa 3/5, 20124 Milano, Italia, Registro delle Imprese di Milano Monza Brianza Lodi REA n. 2504859, Capitale Sociale: 10.000 EUR i.v., Cod. Fisc. e P.IVA 10100050961, Societa con Socio Unico
>
>
>