Re: [PATCH] x86/mm: Fix pmd_modify() dropping the dirty bit
From: Pedro Falcato
Date: Mon Sep 07 2026 - 06:04:09 EST
On Thu, Sep 03, 2026 at 11:16:08AM +0800, Vernon Yang wrote:
> From: Vernon Yang <yanglincheng@xxxxxxxxxx>
>
> pmd_modify() masks the old value with (_HPAGE_CHG_MASK & ~_PAGE_DIRTY),
> silently discarding the hardware dirty bit. The subsequent
> pmd_mksaveddirty() call is supposed to transfer _PAGE_DIRTY into
> _PAGE_SAVED_DIRTY when write-protecting, but the dirty bit was already
> stripped from the value, so there is nothing left to transfer.
>
> Contrast with pte_modify(), which keeps _PAGE_DIRTY_BITS in its mask,
> and pud_modify(), which keeps _HPAGE_CHG_MASK untouched: pmd_modify()
> is the odd one out. Any pmd_modify() on a writable, dirty PMD loses
> the dirty state.
>
> One visible consequence is data loss with MADV_FREE on PMD-mapped THP:
>
> memset(buf, 0x5A, size); // PMD-mapped THP, PMD dirty
> madvise(buf, size, MADV_FREE); // PMD cleaned but left writable,
> // folio marked lazyfree
> memset(buf, 0x5A, size); // hardware sets _PAGE_DIRTY again
> mprotect(buf, size, PROT_READ); // pmd_modify() drops the dirty bit
> mprotect(buf, size, PROT_READ|PROT_WRITE);
> // ... memory pressure ...
>
> Reclaim (e.g. under memcg pressure) then finds the lazyfree folio with
> no dirty bit set anywhere and frees it in
> __discard_anon_folio_pmd_locked(), even though the data was rewritten
> after MADV_FREE; subsequent reads fault in fresh zero pages. NUMA
> hinting alone can trigger the same loss, as do_huge_pmd_numa_page()
> restores the PMD through pmd_modify() as well.
>
> PMD-mapped file THPs are affected too: mprotect()/NUMA hinting dropping
> the dirty bit means rewritten data is never written back.
>
> Fix it by keeping _PAGE_DIRTY in the preserved mask, exactly like
> pte_modify() and pud_modify() do. The existing
> pmd_mksaveddirty()/pmd_clear_saveddirty() pair then performs the
> hardware-dirty <-> saved-dirty transition based on the write bit,
> preserving the shadow-stack encoding rules.
>
> Closes: https://lore.kernel.org/r/CAJxLxMUGu1-L+O_nAONOwOXnS=cNbNApCWqdthRjd76LThtSPg@xxxxxxxxxxxxxx/
> Fixes: bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Vernon Yang <yanglincheng@xxxxxxxxxx>
FYI we just had a customer case for our downstream kernel for this exact same
issue. This fixed it beautifully (I'm wondering if polars started doing
something interesting lately, that nobody else does, and that's why...).
Reviewed-by: Pedro Falcato <pfalcato@xxxxxxx>
Tested-by: Pedro Falcato <pfalcato@xxxxxxx>
--
Pedro