Re: [PATCH] mm/huge_memory: transfer the pmd dirty bit to the folio on zap
From: Hugh Dickins
Date: Wed Aug 19 2026 - 16:34:09 EST
On Wed, 19 Aug 2026, Kiryl Shutsemau wrote:
> On Wed, Aug 19, 2026 at 03:12:22AM -0700, Usama Arif wrote:
> > zap_huge_pmd_folio() propagates the pmd young bit to the folio for the
> > file case, but not the dirty bit. The pte path does propagate it, in
> > zap_present_folio_ptes() and so does the pmd split path, in
> > __split_huge_pmd_locked().
> >
> > For most file mappings the omission is harmless, because writing to a
> > shared file mapping goes through page_mkwrite(), which dirties the
> > folio. tmpfs is different: it has no page_mkwrite(), and
> > vma_wants_writenotify() is false for it, so a *read* fault on a
> > MAP_SHARED tmpfs mapping installs a writable pmd via do_read_fault().
> > do_read_fault() does not call fault_dirty_shared_page(), so subsequent
> > stores through that mapping set only the hardware dirty bit in the pmd
> > and never call folio_mark_dirty().
> >
> > A shmem folio allocated by a fault
> > is marked uptodate but not dirty (see the clear: block in
> > shmem_get_folio_gfp()), so PG_dirty is never set at all.
> >
> > Unmapping such a folio - munmap(), or exit_mmap() when the process dies
> > - then loses the only record that it was written, because zap_huge_pmd()
> > drops the pmd without transferring the dirty bit. Reclaim afterwards
> > sees a clean shmem folio: the whole swap-out block in
> > shrink_folio_list() is inside "if (folio_test_dirty(folio))", so
> > pageout() is skipped and the folio falls into __remove_mapping().
> > There, folio_is_file_lru() is false for a swapbacked folio, so no shadow
> > entry is created and __filemap_remove_folio(folio, NULL) simply empties
> > the i_pages slot. The data is freed without ever being written to swap,
> > and the next fault on that index returns a freshly zeroed folio.
> >
> > This is silent data loss for any process that keeps state in a
> > MAP_SHARED tmpfs segment across an unmap - for example a cache handed
> > from one process generation to the next through /dev/shm. It requires
> > the folio to be PMD-mapped, so it only shows up once shmem THP is
> > enabled (which is what we did in Meta fleet and started noticing crashes);
> > with THP off the pte path transfers the dirty bit correctly.
> > It also only becomes visible when swap is enabled, because with no swap
> > device shmem folios (which are on the anon LRU) are not scanned by
> > reclaim at all, so the clean folio is never dropped.
> >
> > Reproduced on x86_64 with a tmpfs mounted huge=within_size: read-fault a
> > 2MB-backed region, write a known pattern through the resulting mapping,
> > munmap, force reclaim of the cgroup, then re-map and read back. Without
> > this patch the region reads back as zeros and vmstat shows zswpout 0 -
> > the data was discarded rather than swapped. With this patch the region
> > reads back correctly and the pages are swapped out as expected. With
> > huge=never, or when the first touch is a write, the test passes either
> > way.
>
> +Hugh.
>
> Oopsie.
How ghastly! Thanks for finding and fixing, Usama.
>
> I'm confused why it took a decade to discover the bug...
> Maybe read ahead of write for shmem is too rare, I donno.
It isn't entirely clear from the report, but this is all about modifying
a 2MiB+ *hole* in a shmem file through an mmap thereof, with first fault
a read fault not a write fault. I suppose only a few proceed in that way
(though truncating an empty file to some size and then mmap'ing that size
is very normal).
Or everybody who tried to report this bug, wrote their report into a
2MiB hole in a shmem file through an mmap thereof.
Other than holes, all shmem folios are dirty throughout (and re-marked
dirty as soon as brought back from swap): so for most, it doesn't matter
what the pmd says.
This raised a dim memory, took a while to locate what I was remembering:
e1f1b1572e8d ("mm/huge_memory.c: fix data loss when splitting a file pmd")
from 2018.
Not quite the same; but what a pity that one didn't prompt any of us
to look further and find what Usama now has.
>
> >
> > Fixes: 800d8c63b2e9 ("shmem: add huge pages support")
>
> This would be more precise: b5072380eb61 ("thp: support file pages in zap_huge_pmd()")
>
> Reviewed-by: Kiryl Shutsemau <kas@xxxxxxxxxx>
>
> > Cc: <stable@xxxxxxxxxxxxxxx>
> > Signed-off-by: Usama Arif <usama.arif@xxxxxxxxx>
Acked-by: Hugh Dickins <hughd@xxxxxxxxxx>