Re: [PATCH v8 15/30] mm: swap in PMD swap entries as whole THPs during swapoff

From: Zi Yan

Date: Fri Oct 09 2026 - 15:19:19 EST


On Fri Oct 2, 2026 at 5:52 AM EDT, Usama Arif wrote:
> swapoff walks every mm and faults every slot of the device back in.
> unuse_pte_range() only understands PTEs, so a PMD swap entry would never be
> found and swapoff would never complete.
>
> A PMD swap entry is a compact encoding for HPAGE_PMD_NR slots, not a
> promise that the swap cache holds one folio for them. Add
> swap_pmd_cache_lookup() to classify the covered range as empty, one
> PMD-sized folio, or already split, and unuse_pmd() to map the first two
> cases back in as one THP, preserving soft-dirty, exclusive and UFFD state.
>
> Everything else falls back to PTEs: a split cache, per-page zswap state, a
> failed PMD-order allocation or read, or a poisoned subpage. Check
> PageHWPoison on every subpage rather than the folio-level flag, which
> memory_failure() only sets after taking the folio lock.
>
> All the fallback reasons are observed without the PMD lock and possibly
> after sleeping, so they share one exit that re-checks the PMD is still the
> entry we were called for before splitting it. That exit also drops a folio
> that is not uptodate, or that has never been mapped, from the swap cache:
> the PTE path cannot re-read the first, and would add a single-page rmap to
> the second.
>
> Signed-off-by: Usama Arif <usama.arif@xxxxxxxxx>
> ---
> mm/internal.h | 16 ++++
> mm/swap.h | 17 +++++
> mm/swap_state.c | 44 +++++++++++
> mm/swapfile.c | 189 ++++++++++++++++++++++++++++++++++++++++++++++++
> 4 files changed, 266 insertions(+)
>
> diff --git a/mm/internal.h b/mm/internal.h
> index 05179c4b2090e..ec7f007bc2c0d 100644
> --- a/mm/internal.h
> +++ b/mm/internal.h
> @@ -24,6 +24,22 @@
>
> struct folio_batch;
>
> +/*

A high level description like below might be better.

Check every page in a folio for HWPoison flag. Return true if there is
any.

> + * Unlike folio_contain_hwpoisoned_page(), this does not rely on the folio-level
> + * PG_has_hwpoisoned, which memory_failure() only sets after taking the folio
> + * lock and so can lag a tail-page poison.
> + */
> +static inline bool folio_has_hwpoisoned_subpage(const struct folio *folio)

s/subpage/page, IIRC, we are moving away from subpage.

Maybe folio_has_any_hwpoisoned_page()?

> +{
> + long nr = folio_nr_pages(folio);
> + long i;
> +
> + for (i = 0; i < nr; i++)
> + if (PageHWPoison(folio_page(folio, i)))
> + return true;
> + return false;
> +}
> +


--
Best Regards,
Yan, Zi