Re: [BUG] shmem: FALLOC_FL_PUNCH_HOLE vs fault-around race corrupts page cache / rss counters

From: Jan Kara

Date: Thu Sep 24 2026 - 05:20:05 EST


On Thu 24-09-26 09:34:24, Pedro Falcato wrote:
> On Thu, Sep 24, 2026 at 06:16:21AM +0000, Ayush Ranjan wrote:
> > Here is my best understanding of the race -- corrections welcome:
> >
> > shmem guards faults against an in-progress hole punch with
> > inode->i_private: shmem_fault() -> shmem_falloc_wait() waits while
> > shmem_fallocate(PUNCH_HOLE) holds i_private. But shmem's .map_pages
> > is the generic filemap_map_pages() (shmem_vm_ops /
> > shmem_anon_vm_ops), which does not consult i_private and does not
> > take invalidate_lock, and shmem does not use invalidate_lock to
> > serialize faults against truncation the way regular filesystems do --
> > the i_private + waitq scheme stands in for it, but only shmem_fault()
> > participates in that scheme.
> >
> > So while shmem_fallocate(PUNCH_HOLE) is between
> > unmap_mapping_range() and shmem_truncate_range(), a concurrent
> > fault-around can (re-)install PTEs for folios that are about to be
> > truncated:
> >
> > - filemap_map_pages() samples mm_counter_file(folio) once per batch
> > and applies it with add_mm_counter() after mapping; if the
> > folio's swapbacked state changes while it is concurrently torn
>
> But that cannot happen? We hold the folio lock in filemap_map_pages().
> The folio (naturally) cannot be torn down while we have the folio lock.
>
> > down, the map-time counter (MM_FILEPAGES) and the zap-time
> > counter (MM_SHMEMPAGES) disagree by exactly one folio -- the
> > +/-512 imbalance above.
> >
> > - a folio re-mapped in this window (by fault-around directly, or
> > via a child VMA whose PTEs copy_page_range() installs after
> > unmap_mapping_range() has already walked the i_mmap tree -- the
> > dup_mmap() variant discussed in the earlier thread) can be
> > deleted from the page cache while still mapped -> "still mapped
> > when deleted".
>
> No, I don't think this paragraph is true. Page cache truncation (via
> truncate, or fallocate PUNCH_HOLE) takes the folio lock for each folio
> that is about to be truncated out. Mapping folios takes the folio lock
> as well, except in the fork() case where a myriad of weird interval tree
> + PTE lock interactions make it safe (AIUI).

Can this be perhaps somehow related to the fixes in partial large folio
truncation Zhang Yi is working on, possibly even the tmpfs bug in handling
of folio split I've found [1]? It seems large folios are used here so that
matches, I just don't immediately see how those bugs would lead to the
errors reported here...

Honza

[1] https://lore.kernel.org/all/5pthbyxtn7q6xi4fmkofvksmcjzfnujcw2g4fxmxjzfin5pbgf@zui3vcimb4cv

--
Jan Kara <jack@xxxxxxxx>
SUSE Labs, CR