Re: [PATCH v2 0/3] mm: reject zone device folios in more folio walkers

From: Lance Yang

Date: Sun Aug 30 2026 - 01:18:50 EST




On 2026/8/30 08:18, Andrew Morton wrote:
On Tue, 18 Aug 2026 12:31:43 +0800 Lance Yang <lance.yang@xxxxxxxxx> wrote:


On Mon, Aug 17, 2026 at 06:08:07PM -0400, Gregory Price wrote:
Several LRU-oriented mm walkers resolve the folio backing a PMD entry
(or a physical pfn) and then reclaim, age, migrate, or lazyfree it
without ever checking for ZONE_DEVICE memory.

This series adds missing folio_is_zone_device() rejections, matching
the checks that comparable walkers already perform.

- mm/huge_memory, mm/madvise: the !pmd_present branch above these sites
only filters device-private entries (which are non-present).

A present zone device PMD (e.g. device-coherent) would still reach the
folio and be lazyfreed / aged / paged out. Add an explicit check.

- mm/mempolicy: queue_folios_pmd() can see a present zone device PMD
(e.g. device-coherent) and queue it for migration.

No crash reproducer - this is a correctness/hardening cleanup found by
inspection. All checks are placed after the folio is resolved and before
it is acted upon, on paths that already hold the relevant page-table lock,
so no locking or refcount changes are involved.

Cool!

Gave the whole series a spin on x86_64 QEMU with a PMD-mapped
device-coherent THP. Without these patches, partial MADV_FREE and
MADV_COLD reliably hit a kernel panic in remove_migration_pte(), while
mbind(MPOL_MF_MOVE | MPOL_MF_STRICT) returned -EIO.

With v2, all three worked fine, PMD mapping stayed intact, and data
checked out :)

Note that both kernels used the same small change to the in-kernel HMM
test driver, allowing its coherent device memory to be allocated as 2 MB
folios so the PMD-mapped test case could be exercised.

Tested-by: Lance Yang <lance.yang@xxxxxxxxx>

Thanks Lance, you're so diligent.

I'm wondering what to do here. Gregory told us

: No crash reproducer - this is a correctness/hardening cleanup found by
: inspection. All checks are placed after the folio is resolved and before
: it is acted upon, on paths that already hold the relevant page-table lock,
: so no locking or refcount changes are involved.

And you had to tweak the hmm-test driver to reproduce the bug(s).

So when do we push this series out to -stable? As a hair-on-fire
hotfix, or as a leisurely next-merge-window thing?

Thanks, Andrew :) Yeah, I'd say next merge window should be fine :)

The crash is real once the mapping exists, but I had to tweak test_hmm
to create that PMD-mapped device-coherent folio, and I couldn't find
any in-tree production driver doing that today.

So no need to rush this one, I guess.


And Sashiko was clearly having a bad day, able to find only nine
pre-existing things to shout about:
https://sashiko.dev/#/patchset/20260817220810.1175596-1-gourry@xxxxxxxxxx