Re: [PATCH v2 0/3] mm: reject zone device folios in more folio walkers
From: Andrew Morton
Date: Sat Aug 29 2026 - 20:18:27 EST
On Tue, 18 Aug 2026 12:31:43 +0800 Lance Yang <lance.yang@xxxxxxxxx> wrote:
>
> On Mon, Aug 17, 2026 at 06:08:07PM -0400, Gregory Price wrote:
> >Several LRU-oriented mm walkers resolve the folio backing a PMD entry
> >(or a physical pfn) and then reclaim, age, migrate, or lazyfree it
> >without ever checking for ZONE_DEVICE memory.
> >
> >This series adds missing folio_is_zone_device() rejections, matching
> >the checks that comparable walkers already perform.
> >
> >- mm/huge_memory, mm/madvise: the !pmd_present branch above these sites
> > only filters device-private entries (which are non-present).
> >
> > A present zone device PMD (e.g. device-coherent) would still reach the
> > folio and be lazyfreed / aged / paged out. Add an explicit check.
> >
> >- mm/mempolicy: queue_folios_pmd() can see a present zone device PMD
> > (e.g. device-coherent) and queue it for migration.
> >
> >No crash reproducer - this is a correctness/hardening cleanup found by
> >inspection. All checks are placed after the folio is resolved and before
> >it is acted upon, on paths that already hold the relevant page-table lock,
> >so no locking or refcount changes are involved.
>
> Cool!
>
> Gave the whole series a spin on x86_64 QEMU with a PMD-mapped
> device-coherent THP. Without these patches, partial MADV_FREE and
> MADV_COLD reliably hit a kernel panic in remove_migration_pte(), while
> mbind(MPOL_MF_MOVE | MPOL_MF_STRICT) returned -EIO.
>
> With v2, all three worked fine, PMD mapping stayed intact, and data
> checked out :)
>
> Note that both kernels used the same small change to the in-kernel HMM
> test driver, allowing its coherent device memory to be allocated as 2 MB
> folios so the PMD-mapped test case could be exercised.
>
> Tested-by: Lance Yang <lance.yang@xxxxxxxxx>
Thanks Lance, you're so diligent.
I'm wondering what to do here. Gregory told us
: No crash reproducer - this is a correctness/hardening cleanup found by
: inspection. All checks are placed after the folio is resolved and before
: it is acted upon, on paths that already hold the relevant page-table lock,
: so no locking or refcount changes are involved.
And you had to tweak the hmm-test driver to reproduce the bug(s).
So when do we push this series out to -stable? As a hair-on-fire
hotfix, or as a leisurely next-merge-window thing?
And Sashiko was clearly having a bad day, able to find only nine
pre-existing things to shout about:
https://sashiko.dev/#/patchset/20260817220810.1175596-1-gourry@xxxxxxxxxx