[RFC PATCH 4/6] mm/memory-failure: handle swapcached hugetlb folios

From: Zongkun Lei

Date: Tue Sep 29 2026 - 04:07:34 EST


Anonymous hugetlb folios can now sit in the swap cache (swap-out
writeback window, cached copy after swap-in), a state memory failure
never had to handle before. Two gaps:

1. try_to_unmap() routes every hugetlb folio to
try_to_unmap_poisoned_hugetlb_one(), which requires TTU_HWPOISON.
But unmap_poisoned_folio() clears TTU_HWPOISON for a dirty
swapcache folio -- its mappings must be replaced with swap entries,
exactly like for a 4K swapcache folio, so the swap accounting stays
intact and the fault path gets to kill the owner. Routing such a
folio to the poisoned handler fires its VM_WARN_ON_ONCE and would
install hwpoison entries on top of live swap state. Route hugetlb
folios by flag instead: TTU_HWPOISON set -> the poisoned handler;
cleared -> try_to_unmap_swap_hugetlb_one().

2. me_huge_page() has no swapcache case: a clean poisoned hugetlb
folio in the swap cache falls into the truncate path, which evicts
it (the swap address space has no error_remove_folio). The folio
is never returned to the hstate pool (a silent 2M pool leak) and
the poison marker is lost, so a later fault would swap possibly
corrupt data back in. Mirror me_swapcache_dirty(): keep the
poisoned folio in the swap cache and return MF_DELAYED, so the
swap-in path sees folio_test_hwpoison() and kills the accessor.

Signed-off-by: Zongkun Lei <leizongkun@xxxxxx>
---
mm/memory-failure.c | 19 +++++++++++++++++++
mm/rmap.c | 32 ++++++++++++++++++++++++++------
2 files changed, 45 insertions(+), 6 deletions(-)

diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index a8b03e2920ba..0c31c8ab5542 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -1151,6 +1151,25 @@ static int me_huge_page(struct page_state *ps, struct page *p)
struct address_space *mapping;
bool extra_pins = false;

+ /*
+ * A hugetlb folio reaches the swap cache only via hugetlb
+ * swap-out. Keep a poisoned one there so that the swap-in
+ * fault path intercepts folio_test_hwpoison() and kills the
+ * accessor. The truncate path below would evict a clean one
+ * (the swap address space has no error_remove_folio), losing
+ * both the folio (it is never returned to the hstate pool) and
+ * the poison marker, so a later fault would silently swap
+ * possibly-corrupt data back in. Mirror me_swapcache_dirty().
+ */
+ if (folio_test_swapcache(folio)) {
+ folio_clear_dirty(folio);
+ folio_unlock(folio);
+ /* The swap cache pin is intentionally retained. */
+ if (has_extra_refcount(ps, p, true))
+ return MF_FAILED;
+ return MF_DELAYED;
+ }
+
mapping = folio_mapping(folio);
if (mapping) {
res = truncate_error_folio(folio, page_to_pfn(p), mapping);
diff --git a/mm/rmap.c b/mm/rmap.c
index 2ce54cb5ff20..a88c911e8463 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -1991,8 +1991,9 @@ static bool try_to_unmap_poisoned_hugetlb_one(struct folio *folio,
pte_t pteval;

/*
- * The try_to_unmap() is only passed a hugetlb folio in the case
- * where the hugetlb folio is poisoned.
+ * try_to_unmap() routes a hugetlb folio here only for memory
+ * failure with TTU_HWPOISON set; a poisoned folio kept in the
+ * swap cache is routed to try_to_unmap_swap_hugetlb_one() instead.
*/
VM_WARN_ON_ONCE_FOLIO(!folio_test_hwpoison(folio), folio);
VM_WARN_ON_ONCE(!(flags & TTU_HWPOISON));
@@ -2199,8 +2200,10 @@ static bool ttu_anon_folio(struct vm_area_struct *vma, struct folio *folio,
* with a swap entry covering the whole folio; for a file-backed folio the
* PTE is simply cleared, the swap anchor lives in the hugetlbfs page
* cache instead (shmem-style). Called from the hugetlb reclaim path
- * (hugetlb_reclaim_pages()) with the folio lock held; an anonymous folio
- * is swapbacked with its swap slots allocated.
+ * (hugetlb_reclaim_pages()) and, via try_to_unmap(), from memory failure
+ * for a poisoned folio that is kept in the swap cache (see
+ * unmap_poisoned_folio()); the folio lock is held in all cases, and an
+ * anonymous folio is swapbacked with its swap slots allocated.
*
* Keep in sync with ttu_anon_swapbacked_folio().
*/
@@ -2567,13 +2570,30 @@ static int folio_not_mapped(struct folio *folio)
void try_to_unmap(struct folio *folio, enum ttu_flags flags)
{
struct rmap_walk_control rwc = {
- .rmap_one = folio_test_hugetlb(folio) ?
- try_to_unmap_poisoned_hugetlb_one : try_to_unmap_one,
+ .rmap_one = try_to_unmap_one,
.arg = (void *)flags,
.done = folio_not_mapped,
.anon_lock = folio_lock_anon_vma_read,
};

+ /*
+ * try_to_unmap() is passed a hugetlb folio only by memory failure.
+ * With TTU_HWPOISON the mappings are replaced with hwpoison
+ * entries. Without it the folio is a poisoned folio that is kept
+ * in the swap cache (unmap_poisoned_folio() cleared TTU_HWPOISON),
+ * so the mappings are replaced with swap entries, exactly like for
+ * a 4K swapcache folio: the swap accounting stays intact and the
+ * fault path gets to kill the owner.
+ */
+ if (folio_test_hugetlb(folio)) {
+ if (flags & TTU_HWPOISON)
+ rwc.rmap_one = try_to_unmap_poisoned_hugetlb_one;
+ else if (WARN_ON_ONCE(!folio_test_swapcache(folio)))
+ return;
+ else
+ rwc.rmap_one = try_to_unmap_swap_hugetlb_one;
+ }
+
if (flags & TTU_RMAP_LOCKED)
rmap_walk_locked(folio, &rwc);
else
--
2.53.0