Re: [PATCH v6 27/31] ext4: set DISKSIZE_GROW_PENDING after zeroing unaligned EOF block

From: Zhang Yi

Date: Thu Oct 08 2026 - 08:21:28 EST


On 10/5/2026 4:55 PM, Ojaswin Mujoo wrote:
> On Thu, Sep 03, 2026 at 08:35:39PM +0800, Zhang Yi wrote:
>> From: Zhang Yi <yi.zhang@xxxxxxxxxx>
>>
>> In the iomap buffered I/O path, data=ordered mode is not used, so the
>> zeroed EOF block has no implicit ordering with later i_disksize updates.
>> Without the pending state being set, i_disksize can be advanced past the
>> zeroed block before writeback completes, exposing stale data after a
>> crash.
>>
>> Previous patches added the consumer side of the
>> disksize-grow-pending mechanism: the state bit, clear and wait helpers,
>> and ioend tagging. Now add ext4_iomap_mark_disksize_pending() and call
>> it from ext4_block_zero_eof() after zeroing the tail of the block that
>> straddles i_disksize.
>>
>> The helper locks the folio, waits for any in-flight writeback on it to
>> complete, then sets EXT4_STATE_DISKSIZE_GROW_PENDING only if the folio
>> is still dirty. Waiting for writeback prevents folio_test_dirty() from
>> returning false mid-writeback, which would cause us to skip the pending
>> state while zeroed data is still in flight. The dirty check then avoids
>> setting the bit when the data has already been written back.
>>
>> Suggested-by: Jan Kara <jack@xxxxxxx>
>> Signed-off-by: Zhang Yi <yi.zhang@xxxxxxxxxx>
>
> Hey Zhang, so I've been looking at the last few patches
> (DISKSIZE_GROW_PENDING) handling as a whole, and I have a race that I
> think might hit. (This is assuming we will revert 3/31 patch here as its
> buggy). Imagine the following set of operations:

Thank you for the review! Though I don't think this race can actually
happen. See below.

>
> Initial state: i_size = i_disksize = 2k
> pwrite(4k,6k)
> pwrite(8k,10k)
>
> The following order might cause an issue:
>
> 1. pwrite(4k,6k)
> - ext4_block_zero_eof(2k,4k)
> set inode state DISK_SIZE_GROW_PENDING
> - i_size=6k, i_disksize=2k
> 2. writeback(2k,4k)
> - submits a GROW_IO ioend for 0,4k
> - i_size=6k, i_disksize=2k (unchanged)
> 3. writeback(4k,6k)
> - Imainge here a parallel mmap has written beyond 6k so we have
> data past EOF
> - submits IO for 4k,8k
> - i_size=6k, i_disksize=2k (unchanged)
> 4. IO for 3. completes (disk has data past 6kb now)
> - completion not called yet.
> 5. pwrite(8k,10k)
> - ext4_block_zero_eof(6k,8k)

Please note that ext4_iomap_mark_disksize_pending() will call
folio_wait_writeback() before setting bit, so this will wait for I/O for
3 to complete. Once it's done, we can obtain a fresh, stable i_disksize
and re-set the DISK_SIZE_GROW_PENDING flag. So the 6k - 8K data shouldn't
be exposed below.

Thanks,
Yi.

> DISK_SIZE_GROW_PENDING already set
> - i_size=10k, i_disksize=2k (updated)
> 6. endio handler for 2.
> - Sees GROW_IO ioend, clears DISK_SIZE_GROW_PENDING
> - i_size=10k i_disksize=10k (updated)
> 7. writeback (8k, 10k) for 5.
> - submits IO for 8k,12k
> - i_size=10k i_disksize=10k (unchanged)
> 8. endio handler for 7.
> - Sees DISK_SIZE_GROW_PENDING cleared.
> - converts 8k,12k written.
> - i_size=10k i_disksize=10k (unchanged)
> 9. end handler for 3.
> - Sees DISK_SIZE_GROW_PENDING cleared.
> - converts 4k,8k to written. This exposes 6k to 8k data
> which is zeroed in folio but not on disk yet.
> - i_size=10k i_disksize=10k (unchanged)
> 10. **crash**
> 11. We now have i_disksize of 10k with 6k to 8k written data that was
> supposed to be zeroed but is not
>
> In this case, ideally 6. should not clear the DISKSIZE_GROW_PENDING
> state and update the i_disksize. That should happen once the zero doen
> in 5. is complete. I think, just having an inode state might not be
> enough to handle such cases, we either might need an atomic counter or a
> linked list of ranges.
>
> But I'd first like to to hear your opinion on it to see if the race is
> possible or did I miss something in the logic.
>
> Regards,
> ojaswin
>
>> ---
>> fs/ext4/inode.c | 85 ++++++++++++++++++++++++++++++++++++++++++-------
>> 1 file changed, 73 insertions(+), 12 deletions(-)
>>
>> diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
>> index 8e8859c8f5b0..4cb0ea82f58c 100644
>> --- a/fs/ext4/inode.c
>> +++ b/fs/ext4/inode.c
>> @@ -4838,6 +4838,68 @@ static int ext4_block_zero_range(struct inode *inode,
>> zero_written);
>> }
>>
>> +/*
>> + * Inodes using the iomap buffered I/O path do not use data=ordered mode.
>> + * Therefore, we mark the inode as disksize-grow-pending after zeroing the
>> + * EOF block. The zeroed block will be submitted before any subsequent
>> + * data.
>> + *
>> + * In the I/O completion path, ext4_iomap_wb_disksize_pending_wait() will
>> + * wait for I/O completion before advancing i_disksize if the write
>> + * extends beyond the zeroed boundary.
>> + *
>> + * When zeroed I/O is in progress, operations that extend i_disksize are
>> + * handled as follows:
>> + *
>> + * - Truncate up, append fallocate and zero_range:
>> + * Defer the update. The file size will be updated to i_size by the
>> + * end_io handler once the ongoing pending I/O completes.
>> + *
>> + * - Insert range and collapse range operations:
>> + * Wait synchronously for the relevant I/O to complete before updating
>> + * i_disksize.
>> + */
>> +static int ext4_iomap_mark_disksize_pending(struct inode *inode, loff_t from)
>> +{
>> + struct folio *folio;
>> +
>> + folio = filemap_lock_folio(inode->i_mapping, from >> PAGE_SHIFT);
>> + if (IS_ERR(folio))
>> + /* Already in writeback and cleared? */
>> + return PTR_ERR(folio) == -ENOENT ? 0 : PTR_ERR(folio);
>> +
>> + /*
>> + * Ensure that in-flight writeback, possibly started after
>> + * iomap_zero_range() unlocked the folio, has completed. Without
>> + * this wait folio_test_dirty() below may miss the zeroed data
>> + * (writeback clears PG_dirty), causing us to skip the
>> + * disksize-grow-pending tracking and potentially expose stale
>> + * on-disk data.
>> + */
>> + folio_wait_writeback(folio);
>> + WARN_ON_ONCE(folio_test_writeback(folio));
>> +
>> + /*
>> + * Mark the inode as disksize-grow-pending. The zeroed block will
>> + * be written out by the generic writepages cycle or any other
>> + * syncing operation.
>> + *
>> + * Multiple overlapping unaligned EOF writes should not happen,
>> + * because we only mark the pending state after zeroing the on-disk
>> + * EOF block, and i_disksize can only be updated after the previous
>> + * zeroed pending block has been written back or the dirty folio
>> + * has been discared.
>> + */
>> + if (likely(folio_test_dirty(folio) &&
>> + !ext4_test_inode_state(inode,
>> + EXT4_STATE_DISKSIZE_GROW_PENDING)))
>> + ext4_set_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING);
>> +
>> + folio_unlock(folio);
>> + folio_put(folio);
>> + return 0;
>> +}
>> +
>> /*
>> * Submit and wait for the pending zeroed EOF block range to complete
>> * if the given range [@offset, @end) fully covers it. Must be called
>> @@ -4926,22 +4988,21 @@ int ext4_block_zero_eof(struct inode *inode, loff_t from, loff_t end)
>> * concurrent writeback that updates i_disksize. And if such a race
>> * occurs, it means the previous unaligned EOF block has already been
>> * zeroed (if needed) and persisted to disk.
>> - *
>> - * TODO: In the iomap path, handle this by tracking the ordered range
>> - * and updating i_disksize to i_size after the zeroed data has been
>> - * written back.
>> */
>> - if (ext4_should_order_data(inode) &&
>> - did_zero && zero_written && !IS_DAX(inode) &&
>> + if (did_zero && zero_written && !IS_DAX(inode) &&
>> from < round_up(READ_ONCE(EXT4_I(inode)->i_disksize), blocksize)) {
>> - handle_t *handle;
>> + if (ext4_should_order_data(inode)) {
>> + handle_t *handle;
>>
>> - handle = ext4_journal_start(inode, EXT4_HT_MISC, 1);
>> - if (IS_ERR(handle))
>> - return PTR_ERR(handle);
>> + handle = ext4_journal_start(inode, EXT4_HT_MISC, 1);
>> + if (IS_ERR(handle))
>> + return PTR_ERR(handle);
>>
>> - err = ext4_jbd2_inode_add_write(handle, inode, from, length);
>> - ext4_journal_stop(handle);
>> + err = ext4_jbd2_inode_add_write(handle, inode, from,
>> + length);
>> + ext4_journal_stop(handle);
>> + } else if (ext4_inode_buffered_iomap(inode))
>> + err = ext4_iomap_mark_disksize_pending(inode, from);
>> if (err)
>> return err;
>> }
>> --
>> 2.52.0
>>