Re: [PATCH v6 22/31] ext4: submit and wait for pending disksize-grow I/O on writeback

From: Ojaswin Mujoo

Date: Thu Oct 08 2026 - 05:20:05 EST


On Thu, Oct 08, 2026 at 04:53:24PM +0800, Zhang Yi wrote:
> On 10/5/2026 5:00 PM, Ojaswin Mujoo wrote:
> > On Thu, Sep 03, 2026 at 08:35:34PM +0800, Zhang Yi wrote:
> > > From: Zhang Yi <yi.zhang@xxxxxxxxxx>
> > >
> > > When the current writeback pass begins beyond the disksize-grow-pending
> > > zeroed EOF block, the ioend worker would otherwise have to wait for the
> > > pending EOF block to complete before it can advance i_disksize.
> > > Otherwise the old EOF block could be exposed as stale data once
> > > i_disksize advances past it.
> > >
> > > Therefore, introduce the ioend mechanism for the pending range, tag
> > > ioends that cover the pending zeroed EOF block which straddles
> > > i_disksize with EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO in
> > > ext4_iomap_writeback_submit(), and clear the bit and wake up all waiters
> > > in ext4_iomap_end_bio() when such an ioend completes.
> > >
> > > Clearing EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO does not depend on whether
> > > the disksize grow I/O succeeds. That is, even if the I/O fails, we still
> > > allow subsequent writes in the range to update i_disksize. This is
> > > consistent with the previous behavior, and we rely on data_err=abort to
> > > prevent metadata updates when data write failures occur.
> > >
> > > In order to avoid the ioend that passes the pending range waiting for a
> > > long time, proactively submit the pending range first in
> > > ext4_iomap_writepages() so it completes in parallel with the rest of the
> > > writeback.
> > >
> > > Note that the handling of discarding the zeroed EOF folio will be
> > > processed later, otherwise the bit will be set forever.
> > > EXT4_STATE_DISKSIZE_GROW_PENDING will be set after everthing is done.
> > >
> > > Signed-off-by: Zhang Yi <yi.zhang@xxxxxxxxxx>
> > > ---
> > > fs/ext4/ext4.h | 6 ++++++
> > > fs/ext4/inode.c | 53 ++++++++++++++++++++++++++++++++++++++++++++++-
> > > fs/ext4/page-io.c | 40 +++++++++++++++++++++++++++++++++++
> > > 3 files changed, 98 insertions(+), 1 deletion(-)
> > >
> > > diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h
> > > index 1c3d736fb700..089dbd39c5c2 100644
> > > --- a/fs/ext4/ext4.h
> > > +++ b/fs/ext4/ext4.h
> > > @@ -3986,6 +3986,12 @@ extern int ext4_move_extents(struct file *o_filp, struct file *d_filp,
> > > __u64 len, __u64 *moved_len);
> > > /* page-io.c */
> > > +/*
> > > + * The I/O range covers the zeroed EOF block that straddles i_disksize
> > > + * and will advance it upon completion.
> > > + */
> > > +#define EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO 1UL
> > > +
> > > extern int __init ext4_init_pageio(void);
> > > extern void ext4_exit_pageio(void);
> > > extern ext4_io_end_t *ext4_init_io_end(struct inode *inode, gfp_t flags);
> > > diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
> > > index 05dd4ee805fb..4239be5a769f 100644
> > > --- a/fs/ext4/inode.c
> > > +++ b/fs/ext4/inode.c
> > > @@ -4366,7 +4366,10 @@ static int ext4_iomap_writeback_submit(struct iomap_writepage_ctx *wpc,
> > > int error)
> > > {
> > > struct iomap_ioend *ioend = wpc->wb_ctx;
> > > - struct ext4_inode_info *ei = EXT4_I(ioend->io_inode);
> > > + struct inode *inode = ioend->io_inode;
> > > + struct ext4_inode_info *ei = EXT4_I(inode);
> > > + unsigned int blocksize = i_blocksize(inode);
> > > + loff_t pstart, plen;
> > > /*
> > > * After I/O completion, a worker needs to be scheduled when:
> > > @@ -4379,6 +4382,21 @@ static int ext4_iomap_writeback_submit(struct iomap_writepage_ctx *wpc,
> > > test_opt(ioend->io_inode->i_sb, DATA_ERR_ABORT))
> > > ioend->io_bio.bi_end_io = ext4_iomap_end_bio;
> > > + /*
> > > + * Mark the I/O as DISKSIZE_GROW_IO by setting io_private to
> > > + * EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO if it covers the pending range.
> > > + * Such I/O will allow or trigger i_disksize advancement in the
> > > + * ioend worker.
> > > + */
> > > + plen = ext4_iomap_get_disksize_pending_range(inode, &pstart);
> > > + if (plen &&
> > > + round_down(ioend->io_offset, blocksize) <= pstart &&
> > > + round_up(ioend->io_offset + ioend->io_size, blocksize) >=
> > > + pstart + plen) {
> > > + ioend->io_bio.bi_end_io = ext4_iomap_end_bio;
> > > + ioend->io_private = (void *)EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO;
> > > + }
> > > +
> > > /*
> > > * ext4_iomap_end_bio() always defers endio processing, disable
> > > * generic BIO in task to avoid double deferral since we will use
> > > @@ -4398,6 +4416,33 @@ static const struct iomap_writeback_ops ext4_writeback_ops = {
> > > .writeback_submit = ext4_iomap_writeback_submit,
> > > };
> > > +/*
> > > + * If the current writeback range begins after the pending zeroed EOF
> > > + * block range which straddles i_disksize, issue a separate writeback to
> > > + * flush it first, so as to avoid prolonged waiting.
> > > + */
> > > +static void ext4_iomap_wb_submit_zeroed_eof(struct inode *inode,
> > > + struct writeback_control *wbc)
> > > +{
> > > + struct address_space *mapping = inode->i_mapping;
> > > + loff_t pstart, plen, range_start;
> > > +
> > > + if (wbc->range_cyclic)
> > > + range_start = (loff_t)mapping->writeback_index << PAGE_SHIFT;
> > > + else
> > > + range_start = wbc->range_start;
> > > +
> > > + plen = ext4_iomap_get_disksize_pending_range(inode, &pstart);
> > > + if (!plen || range_start < pstart + plen)
> > > + return;
> > Hi Zhang,
> >
> > Maybe we can check here if pstart lies withing the same folio as the
> > writeback range, then we don't need an explicit flush as we will anyways
> > flush out the whole folio. We can check against the
> > mapping_min_folio_nrbytes perhaps.
> >
>
> Hi Ojaswin,
>
> Thank you for the suggestion. IIUC, do you mean to change it somewhat
> like the following:
>
> if (!plen ||
> round_down(range_start, mapping_min_folio_nrbytes(mapping)) <
> pstart + plen)
> return;

Hi Zhang, yep exactly, it's a small optimization for bs < ps on machines
like powerpc or arm16/64 which have larger page sizes.

Regards,
ojaswin
>
> The benefit is that it covers the case of blocksize < PAGE_SIZE, where
> the pending range occupies only part of the folio and range_start lands
> in the middle or later part of that same folio. It does not help when
> the pending range happens to be in a large folio, though, because
> mapping_min_folio_nrbytes() is only a lower bound on the folio size.
> That said, the pending range is the old EOF of the file and it is
> unaligned, so that folio can more likely to be a single page than a
> large one.
>
> Or do you want it to be very precise, able to know exactly the size of
> the folio containing the pending range and its positional relationship
> with the writeback range? If the latter, that would require an exact
> folio lookup, which seems a bit expensive and perhaps over-optimization
> to me.
>
> Thanks,
> Yi.
>