Re: [PATCH] mm/page_alloc: do not boost watermarks in kdump capture kernels

From: Vlastimil Babka (SUSE)

Date: Tue Sep 15 2026 - 08:24:33 EST


On 9/14/26 15:11, Yuanhe Shu wrote:
> A watermark boost is not confined to one watermark: wmark_pages() adds
> it to min, low and high alike, so every watermark check sees it,
> including should_reclaim_retry() and the last ditch ALLOC_WMARK_HIGH
> attempt in __alloc_pages_may_oom(). Once the boost exceeds the memory
> still free the allocator gives up and invokes the OOM killer, and a
> capture kernel that is still booting has nothing to kill: the boot
> panics and the vmcore is lost.

Sounds like an oversight and we should deboost watermarks first before going
for an oom kill? But I guess at the same time not boosting in the first
place in kdump capture kernels makes sense and it's simpler to do.

> Seen on an arm64 machine with 64K pages, CONFIG_PAGE_BLOCK_MAX_ORDER=10
> (pageblock = 64M) and crashkernel=512M, running a distribution kernel
> based on 7.0.14. A high order UNMOVABLE allocation fell back to a
> MOVABLE pageblock while the capture kernel was still in do_initcalls():
>
> Node 0 DMA free:68096kB boost:65536kB min:68160kB
> low:68800kB high:69440kB managed:479168kB
> Out of memory and no killable processes...
> Kernel panic - not syncing: System is deadlocked on memory
>
> The zone was not short of memory. Subtracting the boost gives
> min:2624kB low:3264kB high:3904kB, so the 68096kB still free sat 17
> times above the high watermark and the allocator would not even have
> entered its slow path. The boost supplied 65536kB of the 68160kB min
> and by itself put the zone 64kB under water. It is that large because
> boost_watermark() clamps it with max(pageblock_nr_pages, max_boost);
> watermark_boost_factor alone would have allowed 5824kB.
>
> Commit 14f69140ff9c ("mm: limit boost_watermark on small zones") already
> tried to protect capture kernels, but it infers them from the zone size
> and skips the boost only below four pageblocks. arm64 64K pageblocks
> were 512M then, so the guard reached zones up to 2G;
> CONFIG_PAGE_BLOCK_MAX_ORDER can cap them at 64M, which shrinks the guard
> to zones under 256M and lets this 468M zone through.

Hm while the large pageblocks on 64kb kernels are source of various
surprises, this at least seems consistent to me. The check together with the
clamp means we limit the boost to 1/4 of the zone regardless of pageblock
size, right?

> kdump is a property of the kernel, not of the zone, so test for it
> directly. A capture kernel exits within seconds and never uses the
> fragmentation avoidance the boost buys. Normal kernels are unaffected:
> the size based check still covers their genuinely tiny zones.
>
> Passing sysctl.vm.watermark_boost_factor=0 to the capture kernel does
> not cover this window: sysctl.* parameters are written through procfs
> by do_sysctl_args(), which runs after do_initcalls() where the panic
> above happened, and watermark_boost_factor has no early_param of its
> own.
>
> Fixes: 1c30844d2dfe ("mm: reclaim small amounts of memory when an external fragmentation event occurs")
> Cc: stable@xxxxxxxxxxxxxxx # v5.0
> Cc: Henry Willard <henry.willard@xxxxxxxxxx>
> Cc: David Hildenbrand <david@xxxxxxxxxx>
> Signed-off-by: Yuanhe Shu <xiangzao@xxxxxxxxxxxxxxxxx>
> ---
> Build tested with CONFIG_CRASH_DUMP=y and =n.
>
> Tested on the affected machine with the original crashkernel=512M and
> capture kernel command line. Without this patch both capture boots
> panicked at 7.9s inside do_initcalls() with a 64M boost in place. With
> it, and with boost_watermark() instrumented, every call was suppressed -
> including those inside do_initcalls(), where the panics happened - free
> memory came down to within 64kB of the min watermark, where a single
> boost would have put it 65472kB under, and the dump completed.
> ---
> mm/page_alloc.c | 9 +++++++++
> 1 file changed, 9 insertions(+)
>
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index 12fac9084c48..276fa7169b99 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -37,6 +37,7 @@
> #include <linux/vmstat.h>
> #include <linux/fault-inject.h>
> #include <linux/compaction.h>
> +#include <linux/crash_dump.h>
> #include <trace/events/kmem.h>
> #include <trace/events/oom.h>
> #include <linux/prefetch.h>
> @@ -2178,6 +2179,14 @@ static inline bool boost_watermark(struct zone *zone)
>
> if (!watermark_boost_factor)
> return false;
> +
> + /*
> + * A kdump capture kernel exits before a boost can pay off, while
> + * the raised watermark can exceed the memory left for the dump.
> + */
> + if (is_kdump_kernel())
> + return false;

I was going to argue for making watermark_boost_factor zero with
is_kdump_kernel() but it would be more code to handle and this is not a
fastpath so I guess it's fine.

> +
> /*
> * Don't bother in zones that are unlikely to produce results.
> * On small machines, including kdump capture kernels running

We should stop mentioning kdump capture kernels here then? They can't reach
here anymore.

>
> base-commit: 704340f1cd0dcef829eb62f5b48ae95a2ce17bdf