Re: [PATCH] mm: thp: pin the inode across a file folio split

From: Kiryl Shutsemau

Date: Mon Jul 13 2026 - 19:45:57 EST


On Mon, Jul 13, 2026 at 04:13:41PM -0700, Andrew Morton wrote:
> On Mon, 13 Jul 2026 18:09:15 +0100 Kiryl Shutsemau <kirill@xxxxxxxxxxxxx> wrote:
>
> > From: "Kiryl Shutsemau (Meta)" <kirill@xxxxxxxxxxxxx>
> >
> > __folio_split() looks up mapping = folio->mapping for a file-backed
> > folio and keeps dereferencing it after the split completes:
> > shmem_uncharge(mapping->host) for folios dropped beyond EOF and
> > i_mmap_unlock_read(mapping) on the way out.
> >
> > Nothing holds an inode reference for that duration. The split relies on
> > the folio the caller keeps locked (@lock_at) to pin the inode through
> > the page cache: while it is locked and present,
> > truncate_inode_pages_final() in evict() cannot make progress. But the
> > split drops @lock_at from the page cache when it falls beyond EOF (the
> > @end handling in __folio_freeze_and_split_unmapped()), while keeping it
> > locked for the caller. That removes the last pin, and a concurrent final
> > iput() can then evict and RCU-free the inode before __folio_split() is
> > done touching mapping.
> >
> > This is reachable from memory_failure(): poisoning a tail page of a
> > shmem THP that straddles EOF makes try_to_split_thp_page() split at that
> > page, so the dropped @lock_at is the folio returned locked. The result
> > is a use-after-free, e.g.:
> >
> > BUG: KASAN: slab-use-after-free in __up_read+0x634/0x790
> > i_mmap_unlock_read include/linux/fs.h:537 [inline]
> > __folio_split+0x732/0x1640 mm/huge_memory.c:4100
> > try_to_split_thp_page+0xab/0x390 mm/memory-failure.c:1675
> > memory_failure+0x1394/0x26e0 mm/memory-failure.c:2470
> >
> > Freed by task 4601:
> > shmem_free_in_core_inode+0x54/0xb0 mm/shmem.c:5177
> > i_callback+0x4c/0xa0 fs/inode.c:326
> > destroy_inode+0x144/0x1e0 fs/inode.c:402
> > evict+0x57f/0xac0 fs/inode.c:870
> >
> > Pin the inode with igrab() before the split and drop the reference with
> > iput() after the last mapping dereference. igrab() returns NULL only if
> > the inode is already being evicted (i_count 0 and I_FREEING set), which
> > a split racing eviction can observe; there is nothing safe to split
> > then, so return -EBUSY, which callers already handle.
>
> Sashiko is worried about iput() while holding folio_lock():
>
> https://sashiko.dev/#/patchset/20260713170915.239819-1-kirill@xxxxxxxxxxxxx

Sashiko is right. If iput() drops the last reference and
@lock_at is still in page cache we would self-deadlock.

I don't see an obvious solution. Will think more tomorrow.

--
Kiryl Shutsemau / Kirill A. Shutemov