RE:(2) [PATCH 0/6] nfsd: layout recall and lease break fixes
From: Daejun Park
Date: Tue Oct 06 2026 - 22:11:53 EST
On Tue, Oct 06, 2026 at 03:46:40PM -0400, Chuck Lever wrote:
> I think the fix has to come first. Before the WRITE or SETATTR, call
> break_layout() non-blocking to start the recall, and return
> NFS4ERR_DELAY on -EWOULDBLOCK, as we do for delegations. Then the
> fence worker and the poll loop can take as long as they need and no
> nfsd thread waits on them.
Agreed, I will put that patch first in v2. Besides WRITE and SETATTR
with a size, XFS breaks layouts in fallocate and in remap, where
xfs_iolock_two_inodes_and_break_layout() waits on the source file as
well, so ALLOCATE, DEALLOCATE, CLONE and COPY will be checked too, the
last two on both files.
That also covers a wait that is unbounded today. For the SCSI layout,
when fencing keeps failing, the fence worker retries without limit and
nfsd4_layout_lm_breaker_timedout() returns false while it is in flight,
so the thread in __break_lease() keeps waiting (read from the code, not
run).
Two things I will test and report with v2:
- The check in nfsd and the break_layout() in XFS are not atomic. A
LAYOUTGET granted between them still puts the thread to sleep in
xfs_break_leased_layouts(). No new layout lease is granted while a
recall is in progress, so this needs a layout granted when none was
held at the check.
- A writer going through the MDS now retries after NFS4ERR_DELAY
instead of waking up when the layout is returned, so a client that
keeps getting new layouts could starve it.
Daejun