Re: [PATCH] mm: khugepaged: don't pass swap entry value to trace_mm_khugepaged_scan_file()

From: Vernon Yang

Date: Wed Aug 12 2026 - 10:09:00 EST


On Tue, Aug 11, 2026 at 05:19:38PM +0200, David Hildenbrand (Arm) wrote:
> On 8/11/26 15:36, Vernon Yang wrote:
> > From: Vernon Yang <yanglincheng@xxxxxxxxxx>
> >
> > When the swap entries found exceed max_ptes_swap, the loop is left via
> > break with folio still holding the xarray value that encodes the swap
> > entry, not valid folio pointer.
> >
> > That value is passed to trace_mm_khugepaged_scan_file(), which feeds it
> > to folio_pfn(). On FLATMEM and SPARSEMEM_VMEMMAP, the page_to_pfn() is
> > plain pointer arithmetic, so the trace event merely prints bogus
> > scan_pfn. On classic SPARSEMEM, the page_to_pfn() reads page->flags,
> > dereferencing the tiny encoded integer and oopsing khugepaged whenever
> > the trace event is enabled.
> >
> > So set folio to NULL before breaking out, the tracepoint maps NULL to
> > scan_pfn of -1, just like exhausted scan naturally.
> >
> > Fixes: d41fd2016ed0 ("mm/khugepaged: add tracepoint to hpage_collapse_scan_file()")
> > Cc: stable@xxxxxxxxxxxxxxx
> > Signed-off-by: Vernon Yang <yanglincheng@xxxxxxxxxx>
> > ---
> > mm/khugepaged.c | 1 +
> > 1 file changed, 1 insertion(+)
> >
> > diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> > index 617bca76db49..bc0d04c9162d 100644
> > --- a/mm/khugepaged.c
> > +++ b/mm/khugepaged.c
> > @@ -2696,6 +2696,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
> > if (xa_is_value(folio)) {
> > swap += 1 << xas_get_order(&xas);
> > if (swap > max_ptes_swap) {
> > + folio = NULL;
> > result = SCAN_EXCEED_SWAP_PTE;
> > count_vm_event(THP_SCAN_EXCEED_SWAP_PTE);
> > break;
>
> Yes, we'll do a folio_pfn(), and used to do a page_to_pfn().
>
> Using the folio after dropping the reference is rather nasty.
>
> Instead of passing the folio, should we just pass the pfn directly?

Yes, LGTM.

Would similar modifications like the following match the effect you want?
If so, I'll make these changes in the next version.

diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 617bca76db49..e7830761d3a2 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -2683,6 +2683,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
int present, swap;
int node = NUMA_NO_NODE;
enum scan_result result = SCAN_SUCCEED;
+ unsigned long pfn;

present = 0;
swap = 0;
@@ -2720,27 +2721,23 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
* PMD-sized THP implies that we can only try
* retracting the PTE table.
*/
- folio_put(folio);
break;
}

node = folio_nid(folio);
if (collapse_scan_abort(node, cc)) {
result = SCAN_SCAN_ABORT;
- folio_put(folio);
break;
}
cc->node_load[node]++;

if (!folio_test_lru(folio)) {
result = SCAN_PAGE_LRU;
- folio_put(folio);
break;
}

if (folio_expected_ref_count(folio) + 1 != folio_ref_count(folio)) {
result = SCAN_PAGE_COUNT;
- folio_put(folio);
break;
}

@@ -2759,7 +2756,14 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
cond_resched_rcu();
}
}
+ if (!folio || xa_is_value(folio)) {
+ pfn = -1;
+ } else {
+ pfn = folio_pfn(folio);
+ folio_put(folio);
+ }
rcu_read_unlock();
+
if (result == SCAN_PTE_MAPPED_HUGEPAGE)
cc->progress++;
else
@@ -2774,7 +2778,7 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
}
}

- trace_mm_khugepaged_scan_file(mm, folio, file, present, swap, result);
+ trace_mm_khugepaged_scan_file(mm, pfn, file, present, swap, result);
return result;
}

--
Cheers,
Vernon