[PATCH 6/8] xfs: take the invalidate lock while mapping a pNFS layout
From: Daejun Park via B4 Relay
Date: Wed Oct 07 2026 - 21:42:46 EST
From: Daejun Park <daejun7.park@xxxxxxxxxxx>
xfs_fs_map_blocks() flushes and invalidates the page cache of the file
under the iolock and then maps the range, but a write fault takes only
the invalidate lock (XFS_MMAPLOCK_SHARED in __xfs_write_fault()), not
the iolock. A fault on a shared mapping of the file can therefore add a
delalloc extent to the range between the flush and the mapping, which
trips the ASSERT on DELAYSTARTBLOCK, or is handed to nfsd, which warns
about a filesystem that returned a delalloc extent and refuses the
layout.
Take the invalidate lock in exclusive mode together with the iolock, as
__xfs_file_fallocate() does, so that page faults wait until the layout
has been mapped. The next patch maps all extents of a layout under these
locks, which makes the window between the flush and the last mapping
wider.
Not marked for stable: the race needs a process on the server that
writes to a shared mapping of the file while a client gets a layout for
the same range, and the comment above xfs_break_leased_layouts() already
leaves such writers unsynchronized with pNFS clients.
Suggested-by: Darrick J. Wong <djwong@xxxxxxxxxx>
Link: https://lore.kernel.org/r/20260519145949.GH9555@frogsfrogsfrogs
Signed-off-by: Daejun Park <daejun7.park@xxxxxxxxxxx>
---
fs/xfs/xfs_pnfs.c | 11 +++++++----
1 file changed, 7 insertions(+), 4 deletions(-)
diff --git a/fs/xfs/xfs_pnfs.c b/fs/xfs/xfs_pnfs.c
index 81e1ae0351..bd6f7d5093 100644
--- a/fs/xfs/xfs_pnfs.c
+++ b/fs/xfs/xfs_pnfs.c
@@ -116,7 +116,7 @@ xfs_fs_map_update_inode(
/*
* Map the extent of a pNFS layout that starts at @offset_fsb, allocating it
* first if it is a hole in a write layout. The mapping may end before or
- * after @end. Called with the iolock held.
+ * after @end. Called with the iolock and the invalidate lock held.
*/
static int
xfs_fs_map_extent(
@@ -225,8 +225,11 @@ xfs_fs_map_blocks(
* similar to direct I/O, except that the synchronization is much more
* complicated. See the comment near xfs_break_leased_layouts
* for a detailed explanation.
+ *
+ * Take the invalidate lock as well, so that page faults cannot add
+ * delalloc extents to the range while it is being mapped.
*/
- xfs_ilock(ip, XFS_IOLOCK_EXCL);
+ xfs_ilock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
error = -EINVAL;
limit = mp->m_super->s_maxbytes;
@@ -262,14 +265,14 @@ xfs_fs_map_blocks(
if (error)
goto out_unlock;
}
- xfs_iunlock(ip, XFS_IOLOCK_EXCL);
+ xfs_iunlock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
error = xfs_bmbt_to_iomap(ip, iomaps, &imap, 0, 0, seq);
*nr_iomaps = 1;
*device_generation = mp->m_generation;
return error;
out_unlock:
- xfs_iunlock(ip, XFS_IOLOCK_EXCL);
+ xfs_iunlock(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL);
return error;
}
--
2.43.0