Re: [PATCH mm-stable] mm/khugepaged: avoid unnecessary checking for swap entries when collapsing a mTHP

From: Pedro Demarchi Gomes

Date: Thu Aug 27 2026 - 23:25:40 EST


On Thu, Aug 27, 2026 at 08:11:51PM +0800, Lance Yang wrote:
>
> On Thu, Aug 27, 2026 at 01:02:14PM +0200, David Hildenbrand (Arm) wrote:
> >On 8/27/26 10:06, Baolin Wang wrote:
> >>
> >>
> >> On 8/26/26 3:24 AM, Pedro Demarchi Gomes wrote:
> >>> mthp_collapse() tries to swap in PTEs when collapsing a mTHP if there are any
> >>> swap PTEs in the PMD range, even if none of those swap PTEs are
> >>> actually part of the mTHP's range.
> >>
> >> Are you sure? I wonder how you tested your patch? Because we never swapin PTEs
> >> for mTHP collapse, see the code in __collapse_huge_page_swapin():
> >
> >I'm confused as well, this doesn't really make sense.
>
> Well ... the change does remove a redundant PTE walk, IIUC ...
>
> I think Pedro's wording is causing the confusion :)

Sorry for the confusion. My patch description was not clear enough.
Thanks for taking the time to clarify the intent of the patch and explain the
code path in more detail.

>
> Yeah, Baolin is right that __collapse_huge_page_swapin() never reaches
> do_swap_page() for a mTHP. For an otherwise eligible lower-order
> candidate, current code still calls it and walks the candidate's PTE range
> when an unrelated swap PTE exists elsewhere in the same PMD :)
>
> Assume an otherwise eligible lower-order candidate has no swap PTE, while
> another subrange in the PMD has one. collapse_scan_pmd() counts unmapped
> over the full PMD and passes that PMD-wide value into mthp_collapse():
>
> static enum scan_result collapse_scan_pmd(struct mm_struct *mm,
> struct vm_area_struct *vma, unsigned long start_addr,
> bool *lock_dropped, struct collapse_control *cc)
> {
> ...
> int node = NUMA_NO_NODE, unmapped = 0;
> ...
> for (i = 0; i < HPAGE_PMD_NR; i++) {
> _pte = pte + i;
> addr = start_addr + i * PAGE_SIZE;
> pteval = ptep_get(_pte);
> ...
> if (pte_none_or_zero(pteval)) {
> ...
> continue;
> }
> if (!pte_present(pteval)) {
> if (++unmapped > max_ptes_swap) {
> ...
> }
> ...
> if (pte_swp_uffd_any(pteval)) {
> result = SCAN_PTE_UFFD;
> goto out_unmap;
> }
> continue;
> }
> ...
> }
> if (cc->is_khugepaged &&
> (!referenced ||
> (unmapped && referenced < HPAGE_PMD_NR / 2))) {
> result = SCAN_LACK_REFERENCED_PAGE;
> } else {
> result = SCAN_SUCCEED;
> }
> ...
> if (result == SCAN_SUCCEED) {
> ...
> result = mthp_collapse(mm, start_addr, referenced,
> unmapped, cc, enabled_orders);
> ...
> }
> ...
> return result;
> }
>
> unmapped is PMD-wide above. mthp_collapse() then passes the same value to
> every attempted candidate:
>
> static enum scan_result mthp_collapse(struct mm_struct *mm,
> unsigned long address, int referenced, int unmapped,
> struct collapse_control *cc, unsigned long enabled_orders)
> {
> unsigned int nr_occupied_ptes, nr_ptes, max_ptes_none;
> ...
> unsigned int order = HPAGE_PMD_ORDER;
>
> while (offset < HPAGE_PMD_NR) {
> nr_ptes = 1UL << order;
>
> if (!test_bit(order, &enabled_orders))
> goto next_order;
>
> max_ptes_none = collapse_max_ptes_none(cc, NULL, order);
> nr_occupied_ptes = bitmap_weight_from(cc->mthp_present_ptes, offset,
> offset + nr_ptes);
>
> /*
> * Swap PTEs accepted during the scan are counted in @unmapped,
> * not in the present-PTE bitmap. Account them for the PMD-order
> * candidate.
> */
> if (is_pmd_order(order))
> nr_occupied_ptes += unmapped;
>
> if (nr_occupied_ptes >= nr_ptes - max_ptes_none) {
> enum scan_result ret;
>
> collapse_address = address + offset * PAGE_SIZE;
> ret = collapse_huge_page(mm, collapse_address, referenced,
> unmapped, cc, order);
> ...
> }
>
> next_order:
> ...
> if (order > KHUGEPAGED_MIN_MTHP_ORDER &&
> (enabled_orders & GENMASK(order - 1, 0))) {
> order--;
> continue;
> }
> next_offset:
> ...
> offset += nr_ptes;
> order = max_order_from_offset(offset);
> }
> ...
> }
>
> Once it reaches the lower-order candidate with no swap PTE, unmapped is
> still nonzero and collapse_huge_page() calls the swapin helper:
>
> static enum scan_result collapse_huge_page(struct mm_struct *mm, unsigned long start_addr,
> int referenced, int unmapped, struct collapse_control *cc,
> unsigned int order)
> {
> ...
> if (unmapped) {
> ...
> result = __collapse_huge_page_swapin(mm, vma, start_addr, pmd,
> referenced, order);
> ...
> }
> ...
> }
>
> __collapse_huge_page_swapin() only calls do_swap_page() after finding a
> non-present, non-none PTE, and lower orders return before that call:
>
> static enum scan_result __collapse_huge_page_swapin(struct mm_struct *mm,
> struct vm_area_struct *vma, unsigned long start_addr,
> pmd_t *pmd, int referenced, unsigned int order)
> {
> ...
> unsigned long addr, end = start_addr + (PAGE_SIZE << order);
> enum scan_result result;
> pte_t *pte = NULL;
> spinlock_t *ptl;
>
> for (addr = start_addr; addr < end; addr += PAGE_SIZE) {
> ...
> vmf.orig_pte = ptep_get_lockless(pte);
> if (pte_none(vmf.orig_pte) ||
> pte_present(vmf.orig_pte))
> continue;
> ...
> if (!is_pmd_order(order)) {
> ...
> result = SCAN_EXCEED_SWAP_PTE;
> goto out;
> }
>
> vmf.pte = pte;
> vmf.ptl = ptl;
> ret = do_swap_page(&vmf);
> ...
> }
> ...
> result = SCAN_SUCCEED;
> out:
> ...
> return result;
> }
>
> For the candidate above, the loop only reads its PTEs and returns
> SCAN_SUCCEED. Pedro's bitmap lets mthp_collapse() determine that upfront
> and skip the walk.
>
> Emm ... as I asked before[1], any numbers showing how much the extra scan
> costs?

I tested this on my machine with a 2 MB PMD size. In my measurements, removing
the extra PTE walk did not show a significant reduction in the overall scan
time. Maybe for machines with a bigger PMD size this can make a difference.

>
> [1] https://lore.kernel.org/lkml/20260825192433.3185880-1-pedrodemargomes@xxxxxxxxx/
>
> Cheers, Lance
>