Re: [PATCH 2/8] xfs: refuse a direct allocation whose block reservation would wrap

From: Darrick J. Wong

Date: Thu Oct 08 2026 - 11:44:42 EST


On Thu, Oct 08, 2026 at 10:39:31AM +0900, Daejun Park via B4 Relay wrote:
> From: Daejun Park <daejun7.park@xxxxxxxxxxx>
>
> xfs_iomap_write_direct() maps one extent, but it reserves blocks for the
> whole count it is given, in an unsigned int. A count of about 2^32
> blocks or more wraps the reservation. If it wraps to a few blocks, the
> allocation uses more blocks than the transaction reserved, and
> xfs_trans_mod_sb() shuts the filesystem down. If it wraps to nearly 2^32
> blocks, the allocation fails with -ENOSPC.
>
> Direct I/O, DAX and buffered writes with an extent size hint never ask
> for that much, as xfs_direct_write_iomap_begin() limits them to
> 1024 pages to keep the count below 32 bits. xfs_fs_map_blocks() has no
> such limit: it asks for everything from the start of a hole to the end
> of the range of a pNFS layout, which a client chooses, and an RW
> LAYOUTGET at offset 0 of an empty file for 16 TiB does it.
>
> Fail such a count with -ENOSPC before anything is reserved, also on a
> filesystem with that much free space, as the transaction cannot take
> it. A count that fits is reserved in full, as before, so a hole longer
> than the free space still fails before it is allocated.
>
> With pynfs as the client and a 32 GiB XFS, an RW LAYOUTGET at offset 0
> of an empty file for 16 TiB shut the filesystem down: "Corruption of
> in-memory data (0x8) detected at xfs_trans_mod_sb". With this patch it
> gets NFS4ERR_NOSPC and nothing is allocated, as with 100 GiB, and one
> for 16 GiB, which fits, is still granted, in three extents.
>
> Fixes: 527851124d10 ("xfs: implement pNFS export operations")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Daejun Park <daejun7.park@xxxxxxxxxxx>
> ---
> fs/xfs/xfs_iomap.c | 7 +++++++
> 1 file changed, 7 insertions(+)
>
> diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
> index 7c6238fed6..1917e49166 100644
> --- a/fs/xfs/xfs_iomap.c
> +++ b/fs/xfs/xfs_iomap.c
> @@ -287,6 +287,13 @@ xfs_iomap_write_direct(
>
> resaligned = xfs_aligned_fsb_count(offset_fsb, count_fsb,
> xfs_get_extsz_hint(ip));
> + /*
> + * The transaction takes the block reservation as an unsigned int.
> + * Refuse a count that does not fit rather than reserve too little;
> + * only a pNFS layout for a huge range asks for that much.
> + */
> + if (resaligned > UINT_MAX - XFS_DIOSTRAT_SPACE_RES(mp, 0))
> + return -ENOSPC;

Uh... seeing as extent records can only map 2^21 blocks maximum and
iomap/pnfs can handle xfs returning a mapping of whatever length we
want, why don't we constrain count_fsb to XFS_BMBT_MAX_EXTLEN?

--D

> if (unlikely(XFS_IS_REALTIME_INODE(ip))) {
> dblocks = XFS_DIOSTRAT_SPACE_RES(mp, 0);
> rblocks = resaligned;
>
> --
> 2.43.0
>
>
>