Re: [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs
From: Yosry Ahmed
Date: Thu Aug 06 2026 - 19:31:35 EST
On Sun, Jul 26, 2026 at 10:22:57PM +0000, Brendan Jackman wrote:
> This is the minimal solution for ensuring that compaction can service
> unmapped allocations. Without this, it's possible for compaction to just
> check watermarks and see plenty of free pages, without being aware of
> the direct map state, and thereby cause an ALLOC_UNMAPPED allocation to
> fail unnecessarily.
>
> Instead, with this change, promote compact_order to pageblock order for
> unmapped allocations, much like defrag_mode. Then, check specifically in
> compaction for the presence of wholly mapped blocks that can be unmapped
> once direct compact is complete.
>
> This all takes advantage of a major simplification: since unmapped
> blocks are currently always unmovable, this can be asymmetric. There is
> never a need to promote a !ALLOC_UNMAPPED allocation to compacting at
> pageblock_order, because compaction would be trying to generate a
> currently-unmapped block to map; that will always fail because it would
> require migrating unmapped pages, which is not supported at the moment.
>
> Signed-off-by: Brendan Jackman <jackmanb@xxxxxxxxxx>
> ---
> mm/compaction.c | 22 ++++++++++++++++++----
> mm/page_alloc.c | 9 +++++++++
> 2 files changed, 27 insertions(+), 4 deletions(-)
>
> diff --git a/mm/compaction.c b/mm/compaction.c
> index ed12d2fc6fad3..fe1aaf293bbce 100644
> --- a/mm/compaction.c
> +++ b/mm/compaction.c
> @@ -2531,12 +2531,25 @@ bool compaction_zonelist_suitable(struct alloc_context *ac, int order,
> static enum compact_result
> compaction_suit_allocation_order(struct zone *zone, unsigned int order,
> int highest_zoneidx, unsigned int alloc_flags,
> - bool async, bool kcompactd)
> + bool unmapped, bool async, bool kcompactd)
> {
> unsigned long free_pages;
> unsigned long watermark;
>
> - if (kcompactd && defrag_mode)
> + /*
> + * When trying to generate an unmapped block, check the counter for
> + * direct-mapped blocks specifically, since we'll need to unmap the
> + * whole block to service the allocation.
> + *
> + * Why doesn't this apply to the other way around too? (Mightn't we need
> + * to _map_ a whole block, to service a !ALLOC_UNMAPPED allocation?) No,
> + * because of a likely-temporary simplification: currently, unmapped
> + * blocks never contain movable pages, so compaction isn't going to free
> + * up one of those.
> + */
> + if (unmapped)
> + free_pages = zone_page_state(zone, NR_FREE_PAGES_BLOCKS_MAPPED);
> + else if (kcompactd && defrag_mode)
> free_pages = zone_free_pages_blocks(zone);
> else
> free_pages = zone_page_state(zone, NR_FREE_PAGES);
(Sorry in advance for the wall of text, while trying to understand why
we need this I ended up spending a lot of time staring at the compaction
code and coming up with even more questions)
Hmm why do we need to do this here?
free pages is used in this check below:
watermark = wmark_pages(zone, alloc_flags & ALLOC_WMARK_MASK);
if (__zone_watermark_ok(zone, order, watermark, highest_zoneidx,
alloc_flags, free_pages))
return COMPACT_SUCCESS;
IIUC, this basically checks if we can skip compaction because we already
have enough free pages, both in terms of zone watermarks as well as the
availability of pages in the right order.
free_pages is used for the watermark checks, so using
NR_FREE_PAGES_BLOCKS_MAPPED will result in __zone_watermark_ok()
returning false if all free memory is above the watermark, but free
mapped memory specifically is above the watermark.
This kinda makes sense, but:
1. Shouldn't we be checking all free page blocks (i.e.
zone_free_pages_blocks() like defrag_mode), as entirely free unmapped
blocks can also serve the allocation?
2. More importantly, I think for this check the more important part is
checking if we have pages from the correct order (i.e. the loop at the
end of __zone_watermark_ok()). Since we promote compaction requests for
unmapped allocations to pageblock_order, this will essentially check if
we have any free pageblocks, which is ultimately what we want.
Checking the watermark against mapped pages only here seems arbitrary
tbh, since all other callers of __zone_watermark_ok() are oblivious to
mapped vs. unmapped, which seems to be a bigger issue.
For example, compaction_suitable(), which IIUC actually checks if we
have enough free memory as scratch space for compaction will check all
free memory against the watermark, even though it cannot use unmapped
memory. Same probably applies for other callers in
allocation/reclaim/compaction paths.
Since the watermark checks are ignorant of unmapped memory, we can end
up serving allocations when the actual free mapped memory we have is
below watermarks, dipping into reserves and causing OOM kills if we run
out of mapped memory.
Other than the watermark checks, the loop at the end of
__zone_watermark_ok() checking for available free pages of the required
order is also oblivious to unmapped memory:
for (o = order; o < NR_PAGE_ORDERS; o++) {
struct free_area *area = &z->free_area[o];
if (!area->nr_free)
continue;
for (int ft_idx = 0; ft_idx < NR_FREETYPE_IDXS; ft_idx++) {
freetype_t ft = freetype_from_idx(ft_idx);
int mt = free_to_migratetype(ft);
if (list_empty(&area->free_list[ft_idx]))
continue;
if (mt < MIGRATE_PCPTYPES)
return true;
...
}
}
AFAICT free_to_migratetype() will return MIRATE_UNMOVABLE for unmapped
pages, so __zone_watermark_ok() will return true for mapped allocations
if there's a free unmapped page of the same order, even though it cannot
actually be used (unless order >= pageblock_order).
So it seems like __zone_watermark_ok() needs to be reworked to account
for unmapped pages, and potentially other watermark checks.
I wonder if this problem can be side-stepped if we allow sharing
pageblocks between mapped and unmapped pages as a fallback (as I
mentioned in another comment). With this, the checks in
__zone_watermark_ok() can probably stay as-is since all free unmapped
memory can be used for all allocations.
Of course, this is not ideal and would cause performance regressions so
we need this to be the last option, probably by generating free
pageblocks more aggressively (like defrag mode). Perhaps we can somehow
detect the presence of unmapped allocations and treat it like defrag
mode?
> @@ -2599,6 +2612,7 @@ compact_zone(struct compact_control *cc, struct capture_control *capc)
> ret = compaction_suit_allocation_order(cc->zone, cc->order,
> cc->highest_zoneidx,
> cc->alloc_flags,
> + freetype_unmapped(cc->freetype),
> cc->mode == MIGRATE_ASYNC,
> !cc->direct_compaction);
> if (ret != COMPACT_CONTINUE)
> @@ -3084,7 +3098,7 @@ static bool kcompactd_node_suitable(pg_data_t *pgdat)
> ret = compaction_suit_allocation_order(zone,
> pgdat->kcompactd_max_order,
> highest_zoneidx, alloc_flags,
> - false, true);
> + false, false, true);
> if (ret == COMPACT_CONTINUE)
> return true;
> }
> @@ -3127,7 +3141,7 @@ static void kcompactd_do_work(pg_data_t *pgdat)
>
> ret = compaction_suit_allocation_order(zone,
> cc.order, zoneid, cc.alloc_flags,
> - false, true);
> + false, false, true);
> if (ret != COMPACT_CONTINUE)
> continue;
>
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index d12ce84662ab7..5f1dea7eee15b 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -827,6 +827,9 @@ compaction_capture(struct capture_control *capc, struct page *page,
> capc_mt != MIGRATE_MOVABLE)
> return false;
>
> + if (freetype_flags(freetype) != freetype_flags(capc->freetype))
> + return false;
> +
> if (migratetype != capc_mt)
> trace_mm_page_alloc_extfrag(page, capc->order, order,
> capc_mt, migratetype);
> @@ -4523,6 +4526,12 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
> if ((alloc_flags & ALLOC_NOFRAGMENT) &&
> free_to_migratetype(ac->freetype) != MIGRATE_MOVABLE)
> compact_order = max(order, pageblock_order);
> + /*
> + * Unmapped allocations benefit from compaction even at order 0, because the
> + * allocator will actually grab a whole block.
> + */
> + if (freetype_flags(ac->freetype) & FREETYPE_UNMAPPED)
> + compact_order = max(order, pageblock_order);
>
> if (!compact_order)
> return NULL;
>
> --
> 2.54.0
>