Re: [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()
From: Andrew Morton
Date: Tue Aug 11 2026 - 15:12:35 EST
On Tue, 11 Aug 2026 21:36:55 +0800 Vernon Yang <vernon2gm@xxxxxxxxx> wrote:
> When the swap entries found exceed max_ptes_swap, the loop is left via
> break with folio still holding the xarray value that encodes the swap
> entry, not valid folio pointer.
>
> That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
> to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
> plain pointer arithmetic, so the trace event merely prints bogus
> scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
> dereferencing the tiny encoded integer and oopsing khugepaged whenever
> the trace event is enabled.
>
> So set folio to NULL before breaking out, the tracepoint maps NULL to
> scan_pfn of -1, just like exhausted scan naturally.
>
> Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
Added in 2022. Why so long - do people not use tracing?
Sashiko might have a found a couple of other tracing bugs in this code,
which I suggest are on-topic for your patch:
https://sashiko.dev/#/patchset/20260811133655.267739-1-vernon2gm@xxxxxxxxx
Also a possible bug mapping large folios which straddle i_size, which
is a separate thing.