Re: [PATCH 1/2] mm/compaction: skip folios containing hwpoisoned pages
From: David Hildenbrand (Arm)
Date: Mon Sep 28 2026 - 08:39:40 EST
On 9/28/26 12:58, Yuanhe Shu wrote:
> memory_failure() sets PG_hwpoison on the error page before it manages
> to pin the folio via get_hwpoison_page(). Compaction running on
> another CPU can isolate that same folio in the meantime, so
> HWPoisonHandlable() then fails on !PageLRU(), the retry loop in
> get_any_page() exhausts itself and memory_failure() gives up with
> MF_IGNORED - while compaction still has the folio on its
> migratepages list, flagged PG_hwpoison.
>
> Nothing stops compaction from migrating such a folio afterwards.
> The corrupted data is copied into a fresh folio that carries no
> poison marker, and no hwpoison PTE entry is installed on the way
> back: the userspace mapping is silently redirected to the corrupted
> copy and no SIGBUS is ever delivered.
>
> Note that a large folio with a PG_hwpoisoned subpage can also sit on
> the LRU without any race at all, via the THP-split-failure path of
> memory_failure() (kill_procs_now() + MF_MSG_UNSPLIT_THP). Guarding
> compaction is therefore needed regardless of how the race in
> memory_failure() itself might be addressed.
>
> folio_mc_copy() does not close this hole: it only detects corruption
> at copy time and only where ARCH_HAS_COPY_MC is implemented (x86_64
> and PPC64; elsewhere copy_mc_highpage() degrades to a plain copy).
Then they should implement it.
> Software-injected poison - what MADV_HWPOISON and the hwpoison-inject
> interface produce, and what tests and fuzzers exercise - never traps
> during the copy on any architecture.
And these are debug interfaces, why do we care?
> An isolation-time check avoids
> the migration entirely, on all architectures and for both real and
> simulated poison.
It's racy. See the link below.
> The check itself is two flag tests on the head page (PG_hwpoison
> and the folio-level has_hwpoisoned marker), not a walk over
> subpages.
>
> Neither page reclaim nor memory hotplug lets such a folio be
> copied into a fresh one: reclaim unmaps the poisoned order-0 page,
> and skips the hwpoisoned large folios it cannot safely unmap, in
> shrink_folio_list()
> commit 1b0449544c64 ("mm/vmscan: don't try to reclaim hwpoison folio")
> commit 9f1e8cd0b7c4 ("mm/vmscan: fix hwpoisoned large folio handling in shrink_folio_list")
> while memory hotplug refuses to migrate them in do_migrate_range()
> commit 5f5ee52d4f58 ("mm/hwpoison: introduce folio_contain_hwpoisoned_page() helper").
> Skip them in compaction as well, for the same reason reclaim settled
> on skipping: the UCE is rare and a race with compaction is rarer
> still, so skipping is enough, and a later memory_failure() will
> handle the folio if the UCE is triggered again - while a migrated
> folio would silently propagate the corruption instead.
>
> Deterministic validation on v7.3-rc4-75-g62f4c998b297: order-4
> mTHP folios were brought into the stable "large, on LRU,
> PG_hwpoisoned subpage" state by pinning a sibling subpage with
> vmsplice and injecting MADV_HWPOISON on another subpage, which
> makes memory_failure() take the THP-split-failure path; after
> triggering compaction, /proc/kpageflags shows which poisoned
> folios moved. The unpatched kernel migrated 16/16 poisoned
> folios and 16/16 clean controls in the same pageblocks; with
> this patch, 0/16 poisoned folios were migrated while their
> controls still migrated 16/16.
Was any of this written by an LLM?
>
> Soft offline is not affected: it only sets the poison marker after
> its own migration has succeeded.
>
> Signed-off-by: Yuanhe Shu <xiangzao@xxxxxxxxxxxxxxxxx>
> ---
> mm/compaction.c | 8 ++++++++
> 1 file changed, 8 insertions(+)
>
> diff --git a/mm/compaction.c b/mm/compaction.c
> index a049415512c6..491adcbc5313 100644
> --- a/mm/compaction.c
> +++ b/mm/compaction.c
> @@ -1093,6 +1093,14 @@ isolate_migratepages_block(struct compact_control *cc, unsigned long low_pfn,
> if (unlikely(!folio))
> goto isolate_fail;
>
> + /*
> + * Migrating would copy the corrupted data into a fresh
> + * folio with no poison marker; skip, as memory hotplug
> + * refuses to migrate such folios for the same reason.
> + */
> + if (folio_contain_hwpoisoned_page(folio))
> + goto isolate_fail_put;
> +
> /*
> * Migration will fail if an anonymous page is pinned in memory,
> * so avoid taking lru_lock and isolating it unnecessarily in an
See
https://lore.kernel.org/linux-mm/20260707090136.52904-1-kaitao.cheng@xxxxxxxxx/
--
Cheers,
David