Re: [RFC PATCH 0/5] mm: sub-folio dirty tracking for PTE-mapped mmap writes

From: Kiryl Shutsemau

Date: Wed Sep 09 2026 - 06:46:58 EST


On Wed, Sep 09, 2026 at 03:02:39AM -0700, Usama Arif wrote:
> On Thu, 3 Sep 2026 19:29:38 +0100 Kiryl Shutsemau <kirill@xxxxxxxxxxxxx> wrote:
>
> > From: "Kiryl Shutsemau (Meta)" <kas@xxxxxxxxxx>
> >
> > A store through a shared file mapping dirties the whole folio. With large
> > page cache folios that turns a 4K store into 2M of writeback: one dirty
> > bit per folio, and writeback has no way to know which part changed.
> >
> > XFS already knows better. iomap tracks dirty state per block and
> > iomap_writeback_folio() submits only the dirty ranges, and the buffered
> > write path sets just the range it copied. Only the mmap path throws that
> > away, because iomap_dirty_folio() covers the whole folio.
> >
> > Narrowing the dirtying at page_mkwrite() time does not work on its own:
> > set_pte_range() batch-maps a whole folio writable on the first shared
> > write fault, so the stores that follow never fault and never reach the
> > filesystem.
> >
> > So harvest the hardware instead. folio_clear_dirty_for_io() already calls
> > folio_mkclean(), whose rmap walk reads pte_dirty() for every entry of the
> > folio and throws it away. Those bits are the only record of which parts
> > of a large folio were written through a mapping. Collect them there and
> > hand the filesystem the runs that were dirty, through a new
> > a_ops->dirty_folio_range().
> >
> > All of this is about PTE-mapped folios. A PMD-mapped folio has a single
> > dirty bit for the 2M it maps, so there is nothing finer to harvest, and
> > it keeps writing back whole. Keeping shared write faults off PMDs is a
> > separate patch and not part of this posting.
> >
> > On a 512M file in 2M folios on XFS, storing one byte per folio and
> > calling msync() wrote 512M before and writes 1M after, with identical
> > minor fault counts.
>
> Hi Kiryl,
>
> The motivation makes sense to me. I will look into the patches.
>
> Just wanted to check, the above xfs example, is that on an ARM host?

No, that was x86. But the math is the same on any arch with 2M THP/mTHP
in page cache and 4k PAGE_SIZE.

--
Kiryl Shutsemau / Kirill A. Shutemov