[PATCH 0/2] SUNRPC/NFS: limit transport disconnects when cancelling flexfiles I/O

From: Tim Menninger

Date: Thu Sep 24 2026 - 16:50:15 EST


This series narrows the transport disconnects performed when flexfiles
cancels I/O for a recalled or revoked layout segment.

Today ff_layout_cancel_io() first cancels matching RPC tasks and then
calls rpc_clnt_disconnect() if any tasks matched. rpc_clnt_disconnect()
operates on the entire rpc_clnt, so for a data server client using
multiple transports it disconnects transports that may have had no I/O
associated with the affected layout segment.

I encountered this while testing pNFS over RPC/RDMA during data server
recovery. Restarting the NFS service on a data server while sustained
direct-read I/O was active could cause repeated reconnect activity and a
large throughput drop. The client-wide disconnect from
ff_layout_cancel_io() amplified the disruption by tearing down otherwise
unaffected transports.

The first patch adds a SUNRPC helper that combines task cancellation with
request-scoped conditional disconnects. Matching tasks are cancelled as
before. A subsequent transport-queue pass marks matching requests
belonging to the same rpc_clnt that remain on the transmit or receive
queues so that, when released, xprt_conditional_disconnect() is called
using the request's saved connection cookie.

The disconnect state is stored on struct rpc_rqst rather than struct
rpc_task. A task can release one request and later acquire another, so a
task-scoped indication could otherwise be consumed by a successor
request and applied to the wrong transport generation.

The request-marking pass also preserves the rpc_clnt selection used by
rpc_cancel_tasks(). An rpc_xprt may be shared by multiple rpc_clnt
instances, so requests belonging to other clients sharing the same
transport are not marked.

The second patch switches flexfiles to the new helper, preserving its
I/O cancellation behavior while avoiding client-wide transport
disconnects.

Testing:

The code in this series does not depend on:

https://lore.kernel.org/all/20260916232906.4101987-1-tmenninger@xxxxxxxxxxxxxxxx/

However, I used that change while collecting the results below because,
without it, the client enters the earlier recovery livelock before the
transport-disconnect behavior can be observed cleanly.

As a control experiment, I also tested skipping the
ff_layout_cancel_io() task-cancellation and client-disconnect path
entirely for RPC/RDMA. Doing so eliminates the reconnect pathology. This
shows that the pathology depends on the cancellation/disconnect path
introduced by commit b739a5bd9d9f ("NFSv4/flexfiles: Cancel I/O if the
layout is recalled or revoked"), but simply opting RPC/RDMA out of that
behavior would also remove the cancellation that is intended to drain
I/O associated with a recalled or revoked layout segment.

This series instead preserves that cancellation while limiting transport
disconnects to matching requests that remain in flight.

I have repeated the data-server restart test many times and the behavior
is consistent. The counters below are intentionally collected after the
restart event, once throughput has partially recovered and the client has
entered the post-recovery state in which it remains until the workload
completes.

This distinction matters because some transport activity during the
data-server restart itself is expected. The problem addressed here is
the continued reconnect activity after that transition has completed:
once recovery has settled, the client should not continue cycling
otherwise healthy RPC/RDMA connections.

Over a 30-second post-recovery sampling window, this series reduces
sunrpc events from 4,526 to 289, a roughly
16-fold reduction. RPC/RDMA disconnect and connect events fall by about
6.6-fold. During the same interval, read throughput improves from
approximately 0.4-1.1 GB/s to 1.9-2.2 GB/s.

Before (throughput: approximately 0.4-1.1 GB/s):

Performance counter stats for 'system wide':

4,526 sunrpc:xprt_disconnect_force
1,911 rpcrdma:xprtrdma_disconnect
1,909 rpcrdma:xprtrdma_connect

30.021267692 seconds time elapsed

After (throughput: approximately 1.9-2.2 GB/s):

Performance counter stats for 'system wide':

289 sunrpc:xprt_disconnect_force
289 rpcrdma:xprtrdma_disconnect
291 rpcrdma:xprtrdma_connect

30.080790959 seconds time elapsed

Tim Menninger (2):
SUNRPC: add request-scoped disconnects when cancelling tasks
NFSv4/flexfiles: limit transport disconnects when cancelling I/O

fs/nfs/flexfilelayout/flexfilelayout.c | 7 +-
include/linux/sunrpc/sched.h | 4 ++
include/linux/sunrpc/xprt.h | 5 +-
net/sunrpc/clnt.c | 93 +++++++++++++++++++++-----
net/sunrpc/sunrpc.h | 6 ++
net/sunrpc/xprt.c | 75 +++++++++++++++++++++
6 files changed, 168 insertions(+), 22 deletions(-)

--
2.34.1