[PATCH 2/2] mm/compaction: defer failed async direct compaction
From: Qiliang Yuan
Date: Thu Oct 01 2026 - 12:32:48 EST
A THP fault on a MADV_HUGEPAGE VMA first tries the local node only,
with __GFP_THISNODE | __GFP_NORETRY, and runs a single round of async
direct compaction there before falling back to other nodes. A failed
async run is never deferred.
When the local node has plenty of free memory but no free pageblock,
and what sits between the free pages can't be migrated, such as memory
long-term pinned for RDMA, each of these faults scans the zone again
and fails. Populating a large buffer pays for one failed compaction
per 2M fault. A KV-cache store that registers hundreds of GiB for RDMA
reports registration growing from tens of seconds to tens of minutes,
and works around it with MPOL_INTERLEAVE.
Defer the async state when an async run fails, and check it before
further async runs. Sync compaction keeps its own state and behaves as
before, and a successful compaction or allocation resets both. Keep
resetting the pageblock skip hints only when sync compaction restarts,
as async relies on them.
Populating a 4 GiB MADV_HUGEPAGE buffer from node 0 of a two-node VM,
with node 0 (24 GiB) fragmented by long-term pinning every other page,
median of 3 runs:
before after
direct compactions 735 15
time in compaction 122.7 ms 3.6 ms
population time 380 ms 273 ms
THPs on node 0 1313 1126
The THPs that no longer land on node 0 come from node 1. With the same
fragmentation but movable memory, compaction still succeeds and the
median run still gets all 2048 THPs on node 0, as before.
Signed-off-by: Qiliang Yuan <odys.yuan@xxxxxxxxx>
---
mm/compaction.c | 21 ++++++++++++++-------
1 file changed, 14 insertions(+), 7 deletions(-)
diff --git a/mm/compaction.c b/mm/compaction.c
index 7f8845d1990aa..25f758bc54f06 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -2600,7 +2600,9 @@ compact_zone(struct compact_control *cc, struct capture_control *capc)
/*
* Clear pageblock skip if there were failures recently and compaction
- * is about to be retried after being deferred.
+ * is about to be retried after being deferred. Only do it when sync
+ * compaction restarts: async compaction relies on the skip hints, and
+ * clearing them on every async retry would rescan the whole zone.
*/
if (compaction_restarting(cc->zone, cc->order))
__reset_isolation_suitable(cc->zone);
@@ -2858,8 +2860,10 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
!__cpuset_zone_allowed(zone, gfp_mask))
continue;
- if (prio > MIN_COMPACT_PRIORITY
- && compaction_deferred(zone, order, true)) {
+ if (prio > MIN_COMPACT_PRIORITY &&
+ (compaction_deferred(zone, order, true) ||
+ (prio == COMPACT_PRIO_ASYNC &&
+ compaction_deferred(zone, order, false)))) {
rc = max_t(enum compact_result, COMPACT_DEFERRED, rc);
continue;
}
@@ -2890,14 +2894,17 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
break;
}
- if (prio != COMPACT_PRIO_ASYNC && (status == COMPACT_COMPLETE ||
- status == COMPACT_PARTIAL_SKIPPED))
+ if (status == COMPACT_COMPLETE ||
+ status == COMPACT_PARTIAL_SKIPPED)
/*
* We think that allocation won't succeed in this zone
* so we defer compaction there. If it ends up
- * succeeding after all, it will be reset.
+ * succeeding after all, it will be reset. A failed
+ * async run only defers further async runs, as sync
+ * compaction may succeed on pageblocks it skipped.
*/
- defer_compaction(zone, order, true);
+ defer_compaction(zone, order,
+ prio != COMPACT_PRIO_ASYNC);
/*
* We might have stopped compacting due to need_resched() in
--
2.43.0