[RFC PATCH v2 0/4] mm: compaction: mTHP-friendly memory compaction
From: Bo Zhang
Date: Wed Sep 30 2026 - 20:14:35 EST
Hi all,
This series improves memory compaction to better serve mTHP (multi-size
THP) allocations, particularly for small orders like order-2 (16KB).
The changes focus on the proactive (kcompactd) compaction path.
Problem:
The current proactive compaction targets COMPACTION_HPAGE_ORDER (order-9,
2MB), which is unfriendly to mTHP in two ways:
1. It migrates folios that already satisfy mTHP allocation needs. For
example, an order-2 folio is a valid mTHP page, yet compaction still
moves it around trying to form order-9 blocks. This is unnecessary
work and wastes energy.
2. Compaction designed for 2MB huge pages is too heavyweight for mTHP.
mTHP allocations are frequent and only require small contiguous blocks
(e.g., 4 pages for order-2). A lighter-weight, mTHP-aware compaction
strategy is needed to reduce overhead and power consumption.
Approach:
This series makes four changes:
1. Generalize the fragmentation score functions to accept an order
parameter. This is a pure refactor with no functional change; all
callers still pass COMPACTION_HPAGE_ORDER.
2. Use the highest always-enabled mTHP order as the proactive compaction
target. Because a fragmentation score for a small order is naturally
much lower than for order-9 under the same fragmentation state, the
proactive watermark is scaled down in proportion to the target order;
otherwise a small order would almost never reach the trigger threshold.
3. During proactive compaction (compact_memory), skip isolating folios
that already satisfy mTHP requirements, avoiding unnecessary migration
overhead.
4. For non-costly mTHP orders, restrict proactive compaction to movable
(and CMA) pageblocks. Unmovable/reclaimable blocks are dominated by
non-migratable kernel allocations. scanning them wastes cycles because
a small order can be satisfied from movable blocks alone. Costly orders
keep scanning all blocks, since they are harder to form and benefit
from any migratable page available.
Test setup:
x86 host, arm64 cross-build (defconfig + CONFIG_DEBUG_INFO=y,
CONFIG_DEBUG_INFO_DWARF5=y), booted with mem=4G. order-3 (32KB) mTHP is
set to "always"; ext4 root mounted data=journal. The kernel build
naturally fragments order-3 free memory during the run.
Runtime (two defrag modes were tested, see Results):
echo always > .../hugepages-32kB/enabled
echo madvise > .../transparent_hugepage/defrag # or: always
echo 20 > /proc/sys/vm/compaction_proactiveness
A background loop keeps proactiveness poked every 100ms so kcompactd
re-evaluates promptly (its default 500ms interval is too coarse to catch
the transient order-3 fragmentation peaks under this workload):
while true; do
echo 20 > /proc/sys/vm/compaction_proactiveness
sleep 0.1
done
Workload: make ARCH=arm64 CROSS_COMPILE=aarch64-linux-gnu- vmlinux -j28
5 rounds per configuration; each round: make clean + drop_caches, then
counter deltas around the build. The baseline is the unmodified kernel.
Results (5-round means). anon_fault_alloc is within 1% between kernels,
so the fallback numbers are directly comparable.
defrag=madvise:
w/o patch w/ patch change
-------------------------- ----------- ----------- --------
anon_fault_alloc 14,307,597 14,401,699 +1%
anon_fault_fallback 97,701 40,975 -58%
fallback ratio 0.68% 0.28%
compact_isolated 8,479,773 1,458,949 -83%
compact_migrate_scanned 78,198,828 22,900,359 -71%
kcompactd cpu (s) 43 7 -84%
sys (kernel build) 6m34s 6m34s ~0%
defrag=always:
w/o patch w/ patch change
-------------------------- ----------- ----------- --------
anon_fault_alloc 14,370,984 14,404,234 ~0%
anon_fault_fallback 619 823
compact_stall 979 585 -40%
compact_isolated 6,390,940 2,431,800 -62%
compact_migrate_scanned 86,171,095 46,339,911 -46%
kcompactd cpu (s) 36 13 -62%
sys (kernel build) 6m34s 6m27s -1.8%
1. With defrag=madvise, the patch cuts mTHP fallback by 58% (ratio
0.68% -> 0.28%).
2. With defrag=always, direct compaction already keeps fallback near
zero, and the patch instead cuts direct compaction stalls by 40%.
In both cases the compaction is done far more cheaply (kcompactd CPU
-62..84%, compact_isolated/scanned -46..83%).
Open questions:
1. In skip_isolation_on_order() and elsewhere, the target order is read
from huge_anon_orders_always directly, since during proactive
compaction cc->order is -1 (via compact_memory) and there is no clean
way to plumb the mTHP order down. This variable can change under
userspace at any time, making the semantics fragile. Ideas on how to
pass the target order through the proactive path cleanly are welcome.
2. The proactive watermark is scaled linearly by the target order
(wmark * order / COMPACTION_HPAGE_ORDER). This is empirical; a more
principled cross-order normalization of the fragmentation score would
be welcome.
Changes since v1
================
The series is now scoped to the proactive compaction path only. Two
patches from v1 were dropped, and two new mechanisms were added:
- Dropped v1 patch 3 ("don't skip proactive compaction for non-costly
mTHP", i.e. running proactive compaction concurrently with kswapd).
Follow-up testing showed it lowered fallback mainly by driving extra
anonymous reclaim/swap rather than by better compaction; keeping the
original kswapd back-off avoids trading workingset for mTHP fallback.
- Dropped v1 patch 4 (mTHP-aware zone_effective_free_pages() watermark):
it added reclaim pressure without a clear net benefit in this scope.
- Split the former patch 1 into a pure refactor (parameterize the
fragmentation score by order) plus the mTHP-targeting change, so the
mechanical and behavioural parts can be reviewed independently.
- Added the proactive watermark scaling by target order, so a small mTHP
order can actually reach the trigger threshold.
- Added patch 4 (skip unmovable blocks for non-costly proactive mTHP
compaction). Without it, proactive compaction wastes most of its
scanning on unmovable blocks full of kernel allocations; with it,
kcompactd's CPU and the scanned/isolated counters drop by an order of
magnitude while fallback still improves.
- compact_hpage_order() uses __fls() (highest always-enabled mTHP order)
guarded by #ifdef CONFIG_TRANSPARENT_HUGEPAGE (m68k allmodconfig build
fix from the kernel test robot).
- Refreshed test data on an mem=4G, order-3, data=journal kernel-build
workload; see Results above.
Bo Zhang (4):
mm: compaction: parameterize fragmentation score by order
mm: compaction: target mTHP order for proactive compaction
mm: compaction: skip isolating large folios that already satisfy mTHP
mm: compaction: skip unmovable blocks for non-costly proactive mTHP
compaction
mm/compaction.c | 85 +++++++++++++++++++++++++++++++++++++++----------
1 file changed, 68 insertions(+), 17 deletions(-)
--
2.34.1