Re: [PATCH 8/8] nfsd: get all extents of a block layout in one ->map_blocks call
From: Chuck Lever
Date: Thu Oct 08 2026 - 12:45:15 EST
Daejun Park <daejun7.park@xxxxxxxxxxx> wrote:
> @@ -146,42 +124,67 @@ nfsd4_block_proc_layoutget(struct svc_rqst *rqstp, struct inode *inode,
> goto out_error;
> bl->nr_extents = nr_extents_max;
> args->lg_content = bl;
> + iomaps = kmalloc_array(nr_extents_max, sizeof(*iomaps), GFP_KERNEL);
> + if (!iomaps)
> + goto out_error;
nr_extents_max is bounded by PAGE_SIZE so that the per-request buffer
stays small. But a struct iomap is nearly twice the size of an encoded
extent, so this array is nearly twice the layout buffer it sits next
to. On 64 KiB pages that is well over 100 KiB of physically contiguous
memory per LAYOUTGET. A failure here turns into NFS4ERR_DELAY for
every request under fragmentation.
Using kvmalloc_array() would cover the large case, or you could cap
nr_extents_max independently of PAGE_SIZE.
> + /*
> + * Get all extents in one call, so that the filesystem maps them
> + * together and they fit.
> + */
> + nr_iomaps = nr_extents_max;
> + error = sb->s_export_op->block_ops->map_blocks(inode, offset, length,
> + iomaps, &nr_iomaps, seg->iomode != IOMODE_READ,
> + &device_generation);
An RW LAYOUTGET with loga_minlength of 0 is refused by
nfsd4_block_iomap_to_extent() on the first unwritten extent, so
currently no write mapping is ever handed out for it. Yet this call
still passes write, so XFS allocates every hole in the range, updates
the inode and forces the log before the result is thrown away. With
one call per request this looks like up to nr_extents_max allocation
transactions instead of one.
Passing write only when loga_minlength is non-zero, and having
nfsd4_block_iomap_to_extent() refuse IOMAP_HOLE for IOMODE_RW the way
it already refuses IOMAP_UNWRITTEN, gives the client the same answer
without allocating anything. Admittedly this is a pre-existing
issue -- perhaps a pre-requisite fix for this series is best, so
that the fix can be backported to LTS.
A note about process: the series touches XFS and exportfs. It will
need to go through the VFS tree, I think, and might conflict with
what's already in nfsd-testing.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)