[PATCH v2 0/2] ext4: fix fast commit and extent status after unwritten extent zero-out

From: Daejun Park via B4 Relay

Date: Wed Oct 07 2026 - 21:25:48 EST


When ext4_ext_convert_to_initialized() converts part of an unwritten
extent, it may zero out the blocks around the range and convert them
too, and ext4_split_extent() does the same with a whole extent when a
split fails. Two things go wrong around that.

1/2: if the zeroout fails, the blocks on that side stay unwritten in the
extent tree, but the extent status tree caches them as written. Reads
return old disk contents, and writes go in place without converting the
extent, so they read back as zeroes once the entry is dropped.

2/2: when the zeroout succeeds, fast commit is told only about the range
that was written. A later in-place overwrite of a zeroed block is not
tracked either, and after a crash the fsynced data reads back as zeroes.

Tested in QEMU on ext4 dev 9091c97be340 plus this series:
- the reproducer described in 2/2, with -o nodelalloc and with
-o dioread_lock, also on a PROVE_LOCKING and DEBUG_ATOMIC_SLEEP kernel
(no report), and the dm-error reproducer described in 1/2;
- FALLOC_FL_WRITE_ZEROES on scsi_debug (lbpws=1 lbprz=1) over a whole
unwritten extent, over one block in the middle of one, and over a
hole, each followed by an in-place write, fsync and
EXT4_IOC_SHUTDOWN: the written block survives replay, with and
without the series;
- xfstests ext4/044 ext4/045 generic/455 generic/456 generic/482 with
-O fast_commit, with and without -o nodelalloc: the same results as
without the series. generic/455 fails on both, at different marks.
generic/482 fails now and then on both (base 3 of 53 runs, with the
series 2 of 46), with e2fsck reporting i_blocks one block short after
replay;
- the ext4 KUnit tests, 72 of 72. They are the only ones to reach the
zeroout after a failed split, with the zeroout stubbed and without
fast commit, so the tracking there is covered by review only.

For stable: the fast commit part goes back to 5.10. 2/2 applies as is
from 7.3-rc1 on, and I will send backports for older trees.

---
Changes in v2:
- 2/2: track the range in ext4_issue_zeroout(), once the zeroout has
succeeded, instead of in ext4_zeroout_es() and
ext4_split_extent_zeroout() (Jan). ext4_issue_zeroout() and
ext4_ext_zeroout() take the handle for that.
ext4_alloc_file_blocks() zeroes out without a handle and passes NULL.
The ext4_convert_unwritten_extents() call after it tracks the blocks
it converts through ext4_map_blocks(), blocks that were already
written keep their mapping, and the FALLOC_FL_WRITE_ZEROES test above
keeps the block on the base kernel too, so I found no gap there.
- 1/2: add Jan's Reviewed-by.
- 2/2: Cc stable without "# 7.0.x"; the fix is in ext4_issue_zeroout()
now, which older trees have as well.
- Link to v1: https://lore.kernel.org/r/20261007-ext4-fc-zeroout-v1-0-b86de6439431@xxxxxxxxxxx

---
Daejun Park (2):
ext4: don't cache unzeroed blocks as written after a failed zeroout
ext4: track zeroed out blocks of unwritten extents for fast commit

fs/ext4/ext4.h | 5 +++--
fs/ext4/ext4_extents.h | 3 ++-
fs/ext4/extents-test.c | 8 +++++---
fs/ext4/extents.c | 39 ++++++++++++++++++++++++++-------------
fs/ext4/inode.c | 25 +++++++++++++++++--------
5 files changed, 53 insertions(+), 27 deletions(-)
---
base-commit: 9091c97be34083587a75db174aab51551d8e8543
change-id: 20261006-ext4-fc-zeroout-e429007c72a0

Best regards,
--
Daejun Park <daejun7.park@xxxxxxxxxxx>