Re: [PATCH v4 5/7] mm/khugepaged: Refactor the PTE state checks into a helper

From: David Hildenbrand (Arm)

Date: Wed Aug 12 2026 - 04:47:56 EST


On 8/11/26 14:48, Nico Pache (Red Hat) wrote:
> For anonymous collapse, the collapse_scan_pmd() and
> __collapse_huge_page_isolate() functions share a large portion of their
> logic. These functions both check the state of the PTEs and verify the
> following:
> - max_pte_* values are not exceeded
> - uffd is not active
> - lazyfree properties
> - non-anonymous
>
> Merge these checks into a helper collapse_check_pte() to reduce code
> duplication. We also add a helper struct for this function called
> pte_check_context which allows us to pass the required parameters in a
> clean and elegant manner.
>
> A helper function is also introduced pte_check_fail() to provide a clean
> interface to set the pte_check_context failure results and return
> PTE_CHECK_FAIL state. This helps reduce code duplications across the new
> collapse_check_pte function.
>
> Two slight modifications are done to the original functionality. We now
> warn (instead of crash) if the anon test fails, and we leverage the
> vm_normal_folio function instead of page->folio, this should be
> functionally equivalent.
>
> No other functional changes intended.
>
> This patch is heavily based off work done by Lance Yang, but modified to
> deal with conflicts and feedback received during the review cycle [1].
>

TL;DR, I think this patch here needs some more work, and we should not fast
track it at this point.

@Andrew, can we delay this patch here for this merge window? Removing it from
mm-unstable shouldn't conflict with any other patch in this series.


> [1] https://lore.kernel.org/linux-mm/20251008043748.45554-1-lance.yang@xxxxxxxxx/
> Suggested-by: David Hildenbrand <david@xxxxxxxxxx>
> Signed-off-by: Nico Pache (Red Hat) <nico.pache@xxxxxxxxx>
> ---
> mm/khugepaged.c | 298 +++++++++++++++++++++++++++++---------------------------
> 1 file changed, 157 insertions(+), 141 deletions(-)
>
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index 90d6e595d282..b7372aba4417 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -65,6 +65,12 @@ enum scan_result {
> SCAN_PAGE_DIRTY_OR_WRITEBACK,
> };
>
> +enum pte_check_result {
> + PTE_CHECK_SUCCEED,
> + PTE_CHECK_FAIL,
> + PTE_CHECK_CONTINUE,
> +};

I don't love this. "pte_check_result" is a bit too generic for my taste. What is
the difference between "success" and "continue"? Unclear.

Likely, "continue" should actually be something like "skip". But it sounds like
we are mixing two things that shouldn't be mixed (a check that can do more than
just succeed or fail).

Not sure if this was suggested during earlier review, the cover letter doesn't
spell it out. Ideally we'd avoid this completely and just rely on existing error
codes. Like scan_result.


> +
> #define CREATE_TRACE_POINTS
> #include <trace/events/huge_memory.h>
>
> @@ -119,6 +125,20 @@ struct collapse_control {
> DECLARE_BITMAP(mthp_present_ptes, MAX_PTRS_PER_PTE);
> };
>
> +struct pte_check_context {
> + struct collapse_control *cc;
> + struct vm_area_struct *vma;
> + unsigned int order;
> + struct folio *folio;
> + int none_or_zero;
> + int shared;
> + int unmapped;
> + enum scan_result result;
> + unsigned int max_ptes_none;
> + unsigned int max_ptes_swap;
> + unsigned int max_ptes_shared;
> +};
> +
> /**
> * struct khugepaged_scan - cursor for scanning
> * @mm_head: the head of the mm list to scan
> @@ -696,74 +716,131 @@ static void count_collapse_event(unsigned int order, enum vm_event_item vm_event
> count_mthp_stat(order, mthp_event);
> }
>
> +/*
> + * pte_check_fail() - A simple helper to set the pte_check_context result and
> + * return PTE_CHECK_FAIL.
> + */
> +static enum pte_check_result pte_check_fail(struct pte_check_context *ctx,
> + enum scan_result result)
> +{
> + ctx->result = result;
> + return PTE_CHECK_FAIL;
> +}

Looks a bit over-engineered and the function doc is just unnecessary.

> +
> +/*
> + * collapse_check_pte() - Check if a PTE is suitable for collapse
> + *
> + * Check if a PTE is suitable for collapse based on the following criteria:
> + * - max_pte_* values are not exceeded
> + * - uffd is not active
> + * - lazyfree properties are not present
> + * - only anonymous pages are present
> + *
> + * a helper struct pte_check_context is used to pass and store relevant
> + * information between the collapse_check_pte() function and the caller.
> + *
> + * Return: PTE_CHECK_SUCCEED if the PTE is suitable for collapse,
> + * PTE_CHECK_FAIL if the PTE is not suitable for collapse,
> + * PTE_CHECK_CONTINUE if the scan should continue to check the next PTE.
> + */

Why is this doc required?

> +static enum pte_check_result collapse_check_pte(pte_t pteval,
> + unsigned long addr, struct pte_check_context *ctx)
> +{
> + if (pte_none_or_zero(pteval)) {
> + if (++ctx->none_or_zero > ctx->max_ptes_none) {
> + count_collapse_event(ctx->order, THP_SCAN_EXCEED_NONE_PTE,
> + MTHP_STAT_COLLAPSE_EXCEED_NONE);
> + return pte_check_fail(ctx, SCAN_EXCEED_NONE_PTE);
> + }
> + return PTE_CHECK_CONTINUE;
> + }
> + if (!pte_present(pteval)) {
> + if (ctx->unmapped == -1)
> + return pte_check_fail(ctx, SCAN_PTE_NON_PRESENT);
> + if (++ctx->unmapped > ctx->max_ptes_swap) {
> + count_collapse_event(ctx->order, THP_SCAN_EXCEED_SWAP_PTE,
> + MTHP_STAT_COLLAPSE_EXCEED_SWAP);
> + return pte_check_fail(ctx, SCAN_EXCEED_SWAP_PTE);
> + }
> + if (pte_swp_uffd_any(pteval))
> + return pte_check_fail(ctx, SCAN_PTE_UFFD);
> + return PTE_CHECK_CONTINUE;
> + }
> + /*
> + * Don't collapse if any of the small PTEs are armed with uffd
> + * write protection. Marking the new huge pmd as write protected
> + * could bring userfault messages that fall outside of the
> + * registered range.
> + */
> + if (pte_uffd(pteval))
> + return pte_check_fail(ctx, SCAN_PTE_UFFD);
> +
> + ctx->folio = vm_normal_folio(ctx->vma, addr, pteval);
> + if (unlikely(!ctx->folio) || unlikely(folio_is_zone_device(ctx->folio)))
> + return pte_check_fail(ctx, SCAN_PAGE_NULL);
> +
> + /*
> + * If the vma has the VM_DROPPABLE flag, the collapse will
> + * preserve the lazyfree property without needing to skip.
> + */
> + if (ctx->cc->is_khugepaged && !(ctx->vma->vm_flags & VM_DROPPABLE) &&
> + folio_test_lazyfree(ctx->folio) && !pte_dirty(pteval))
> + return pte_check_fail(ctx, SCAN_PAGE_LAZYFREE);
> +
> + if (!folio_test_anon(ctx->folio)) {
> + VM_WARN_ON_FOLIO(!folio_test_anon(ctx->folio), ctx->folio);

Huh, that looks odd.

That should just be a VM_WARN_ON_FOLIO(true, ..) or sth like that.

But in collapse_scan_pmd() that warning never existed? So this raises eyebrows.

[...]

I'll play with it to see if we can do better and will reply here later.

--
Cheers,

David