[PATCH RFC v2 0/8] Reduce lock contention in the NFS client
From: Chuck Lever
Date: Wed Sep 02 2026 - 15:41:30 EST
Under a 4KB NFSv3 workload on 100GbE RDMA, roughly 150 RPC worker
threads drive the client, and lock contention dominates its CPU
profile: up to 53% of non-idle cycles are spent in
native_queued_spin_lock_slowpath.
Three locks account for that: reserve_lock on every XID
allocation, queue_lock on every submit and completion, and the
unbound worker pool lock on every enqueue and dequeue for rpciod,
nfsiod, and xprtiod.
An earlier RFC put this fix in the workqueue core, falling back
automatically to a finer scope when the configured scope
degenerates to a single pod. Tejun rejected a finer global default
for unbound workqueues and suggested two alternatives: a sharded
scope between CACHE and SMT, and a scope change on the NFS
workqueues themselves [1]. The first became the
WQ_AFFN_CACHE_SHARD default, whose 8-core shards still leave a
12-core single-LLC system with only two pools. This series does
the second.
Dispatching each new async task through rpciod costs one enqueue
and one dequeue per RPC, and this series leaves that hop in
place. Running an async task's first states in the submitter's
context lets its completion callbacks run there as well. A pNFS
read resent from rpc_release then waits for a layout segment that
its submitter still holds. That is the deadlock commit
54e4a0dfa25d ("pNFS: Fix a deadlock between read resends and
layoutreturn") fixed.
The scope change is applied after alloc_workqueue() has
registered the queue in sysfs, so an administrator's write that
lands in between is overwritten. Closing that window needs
workqueue_sysfs_register() exported and WQ_SYSFS dropped from the
alloc calls. This series leaves it open.
[1] https://lore.kernel.org/all/aYUVVuIidMpuYy3j@xxxxxxxxxxxxxxx/
---
Changes in v2:
- Fix send bvec use-after-free in xprt_request_dequeue_xprt() (sashiko)
- Replace the three workqueue exports with workqueue_set_affn_scope()
- Split the WQ_SYSFS patch into SUNRPC and NFS patches
- Drop v1 patch 2: async completions in the submitter can deadlock
- Correct the pool cost and SMT group wording in the scope patches
- Link to v1: https://patch.msgid.link/20260831-performance-v1-0-8d9fd9b67f96@xxxxxxxxxx
---
Chuck Lever (8):
SUNRPC: Use atomic_t for XID allocation
SUNRPC: Split recv_lock out of xprt->queue_lock
SUNRPC: Set WQ_SYSFS on rpciod and xprtiod
NFS: Set WQ_SYSFS on nfsiod
workqueue: add workqueue_set_affn_scope()
SUNRPC: Reduce rpciod workqueue contention
NFS: Reduce nfsiod workqueue contention
SUNRPC: Reduce xprtiod workqueue contention
fs/nfs/inode.c | 8 ++-
include/linux/sunrpc/xprt.h | 8 ++-
include/linux/workqueue.h | 2 +
kernel/workqueue.c | 37 +++++++++++++
net/sunrpc/sched.c | 18 ++++++-
net/sunrpc/svcsock.c | 6 +--
net/sunrpc/xprt.c | 85 ++++++++++++++++++------------
net/sunrpc/xprtrdma/rpc_rdma.c | 14 ++---
net/sunrpc/xprtrdma/svc_rdma_backchannel.c | 8 +--
net/sunrpc/xprtsock.c | 18 +++----
10 files changed, 142 insertions(+), 62 deletions(-)
---
base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
change-id: 20260831-performance-e465e621c1c0
Best regards,
--
Chuck Lever <cel@xxxxxxxxxx>