Re: [PATCH v6 22/31] ext4: submit and wait for pending disksize-grow I/O on writeback
From: Zhang Yi
Date: Thu Oct 08 2026 - 08:23:08 EST
On 10/8/2026 5:15 PM, Ojaswin Mujoo wrote:
> On Thu, Oct 08, 2026 at 04:53:24PM +0800, Zhang Yi wrote:
>> On 10/5/2026 5:00 PM, Ojaswin Mujoo wrote:
>>> On Thu, Sep 03, 2026 at 08:35:34PM +0800, Zhang Yi wrote:
>>>> From: Zhang Yi <yi.zhang@xxxxxxxxxx>
>>>>
>>>> When the current writeback pass begins beyond the disksize-grow-pending
>>>> zeroed EOF block, the ioend worker would otherwise have to wait for the
>>>> pending EOF block to complete before it can advance i_disksize.
>>>> Otherwise the old EOF block could be exposed as stale data once
>>>> i_disksize advances past it.
>>>>
>>>> Therefore, introduce the ioend mechanism for the pending range, tag
>>>> ioends that cover the pending zeroed EOF block which straddles
>>>> i_disksize with EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO in
>>>> ext4_iomap_writeback_submit(), and clear the bit and wake up all waiters
>>>> in ext4_iomap_end_bio() when such an ioend completes.
>>>>
>>>> Clearing EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO does not depend on whether
>>>> the disksize grow I/O succeeds. That is, even if the I/O fails, we still
>>>> allow subsequent writes in the range to update i_disksize. This is
>>>> consistent with the previous behavior, and we rely on data_err=abort to
>>>> prevent metadata updates when data write failures occur.
>>>>
>>>> In order to avoid the ioend that passes the pending range waiting for a
>>>> long time, proactively submit the pending range first in
>>>> ext4_iomap_writepages() so it completes in parallel with the rest of the
>>>> writeback.
>>>>
>>>> Note that the handling of discarding the zeroed EOF folio will be
>>>> processed later, otherwise the bit will be set forever.
>>>> EXT4_STATE_DISKSIZE_GROW_PENDING will be set after everthing is done.
>>>>
>>>> Signed-off-by: Zhang Yi <yi.zhang@xxxxxxxxxx>
>>>> ---
>>>> fs/ext4/ext4.h | 6 ++++++
>>>> fs/ext4/inode.c | 53 ++++++++++++++++++++++++++++++++++++++++++++++-
>>>> fs/ext4/page-io.c | 40 +++++++++++++++++++++++++++++++++++
>>>> 3 files changed, 98 insertions(+), 1 deletion(-)
>>>>
>>>> diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h
>>>> index 1c3d736fb700..089dbd39c5c2 100644
>>>> --- a/fs/ext4/ext4.h
>>>> +++ b/fs/ext4/ext4.h
>>>> @@ -3986,6 +3986,12 @@ extern int ext4_move_extents(struct file *o_filp, struct file *d_filp,
>>>> __u64 len, __u64 *moved_len);
>>>> /* page-io.c */
>>>> +/*
>>>> + * The I/O range covers the zeroed EOF block that straddles i_disksize
>>>> + * and will advance it upon completion.
>>>> + */
>>>> +#define EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO 1UL
>>>> +
>>>> extern int __init ext4_init_pageio(void);
>>>> extern void ext4_exit_pageio(void);
>>>> extern ext4_io_end_t *ext4_init_io_end(struct inode *inode, gfp_t flags);
>>>> diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
>>>> index 05dd4ee805fb..4239be5a769f 100644
>>>> --- a/fs/ext4/inode.c
>>>> +++ b/fs/ext4/inode.c
>>>> @@ -4366,7 +4366,10 @@ static int ext4_iomap_writeback_submit(struct iomap_writepage_ctx *wpc,
>>>> int error)
>>>> {
>>>> struct iomap_ioend *ioend = wpc->wb_ctx;
>>>> - struct ext4_inode_info *ei = EXT4_I(ioend->io_inode);
>>>> + struct inode *inode = ioend->io_inode;
>>>> + struct ext4_inode_info *ei = EXT4_I(inode);
>>>> + unsigned int blocksize = i_blocksize(inode);
>>>> + loff_t pstart, plen;
>>>> /*
>>>> * After I/O completion, a worker needs to be scheduled when:
>>>> @@ -4379,6 +4382,21 @@ static int ext4_iomap_writeback_submit(struct iomap_writepage_ctx *wpc,
>>>> test_opt(ioend->io_inode->i_sb, DATA_ERR_ABORT))
>>>> ioend->io_bio.bi_end_io = ext4_iomap_end_bio;
>>>> + /*
>>>> + * Mark the I/O as DISKSIZE_GROW_IO by setting io_private to
>>>> + * EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO if it covers the pending range.
>>>> + * Such I/O will allow or trigger i_disksize advancement in the
>>>> + * ioend worker.
>>>> + */
>>>> + plen = ext4_iomap_get_disksize_pending_range(inode, &pstart);
>>>> + if (plen &&
>>>> + round_down(ioend->io_offset, blocksize) <= pstart &&
>>>> + round_up(ioend->io_offset + ioend->io_size, blocksize) >=
>>>> + pstart + plen) {
>>>> + ioend->io_bio.bi_end_io = ext4_iomap_end_bio;
>>>> + ioend->io_private = (void *)EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO;
>>>> + }
>>>> +
>>>> /*
>>>> * ext4_iomap_end_bio() always defers endio processing, disable
>>>> * generic BIO in task to avoid double deferral since we will use
>>>> @@ -4398,6 +4416,33 @@ static const struct iomap_writeback_ops ext4_writeback_ops = {
>>>> .writeback_submit = ext4_iomap_writeback_submit,
>>>> };
>>>> +/*
>>>> + * If the current writeback range begins after the pending zeroed EOF
>>>> + * block range which straddles i_disksize, issue a separate writeback to
>>>> + * flush it first, so as to avoid prolonged waiting.
>>>> + */
>>>> +static void ext4_iomap_wb_submit_zeroed_eof(struct inode *inode,
>>>> + struct writeback_control *wbc)
>>>> +{
>>>> + struct address_space *mapping = inode->i_mapping;
>>>> + loff_t pstart, plen, range_start;
>>>> +
>>>> + if (wbc->range_cyclic)
>>>> + range_start = (loff_t)mapping->writeback_index << PAGE_SHIFT;
>>>> + else
>>>> + range_start = wbc->range_start;
>>>> +
>>>> + plen = ext4_iomap_get_disksize_pending_range(inode, &pstart);
>>>> + if (!plen || range_start < pstart + plen)
>>>> + return;
>>> Hi Zhang,
>>>
>>> Maybe we can check here if pstart lies withing the same folio as the
>>> writeback range, then we don't need an explicit flush as we will anyways
>>> flush out the whole folio. We can check against the
>>> mapping_min_folio_nrbytes perhaps.
>>>
>>
>> Hi Ojaswin,
>>
>> Thank you for the suggestion. IIUC, do you mean to change it somewhat
>> like the following:
>>
>> if (!plen ||
>> round_down(range_start, mapping_min_folio_nrbytes(mapping)) <
>> pstart + plen)
>> return;
>
> Hi Zhang, yep exactly, it's a small optimization for bs < ps on machines
> like powerpc or arm16/64 which have larger page sizes.
>
> Regards,
> ojaswin
>>
Sure, I can fold this optimization into my next iteration.
Thanks,
Yi.
>> The benefit is that it covers the case of blocksize < PAGE_SIZE, where
>> the pending range occupies only part of the folio and range_start lands
>> in the middle or later part of that same folio. It does not help when
>> the pending range happens to be in a large folio, though, because
>> mapping_min_folio_nrbytes() is only a lower bound on the folio size.
>> That said, the pending range is the old EOF of the file and it is
>> unaligned, so that folio can more likely to be a single page than a
>> large one.
>>
>> Or do you want it to be very precise, able to know exactly the size of
>> the folio containing the pending range and its positional relationship
>> with the writeback range? If the latter, that would require an exact
>> folio lookup, which seems a bit expensive and perhaps over-optimization
>> to me.
>>
>> Thanks,
>> Yi.
>>