Re: [PATCH v2] nvme-tcp: pin io_cpu to submitter cpu
From: Saravanan D
Date: Sun Aug 23 2026 - 13:44:01 EST
On Sun, 23 Aug 2026 00:20:21 +0300 Sagi Grimberg <sagi@xxxxxxxxxxx> wrote:
> Can you share a concrete example?
> 1. lscpu
> 2. grep -r '' /sys/block/nvmeXnY/mq/*/cpu_list
> Also, what would be the VMs cpusets on such a server?
Very grateful for your review and suggestions.
Concrete example from one of these hosts, where there are more cpus than
the io queues NVMe-oF targets offer.
lscpu (trimmed): AMD EPYC 9654, 2 sockets, 2 threads/core, 384 CPUs
NUMA node0 CPUs: 0-95,192-287
NUMA node1 CPUs: 96-191,288-383
Each nvme-tcp controller exposes 128 io queues (queue_count 129) against
384 CPUs, so blk_mq_map_queues() folds three CPUs into every hctx. cpu_list for the namespace's path device, around one tenant's boundary:
hctx67: 6, 7, 198
hctx68: 8, 9, 200
hctx69: 10, 11, 202
hctx75: 22, 23, 214
hctx76: 24, 25, 216
hctx113: 199, 201, 203
hctx114: 205, 207, 209
hctx115: 211, 213, 215
Two VMs are live on this same shared controller:
VM A cpuset: 8-23,200-215
VM B cpuset: 104-191,296-383
VM A owns 200-215, so hctx113's group {199,201,203} has 201 and 203 inside
VM A and 199 outside it. The connect time pick can select 199 as that
queue's io_cpu, so VM A's I/O through hctx113 runs its socket work on 199,
outside its cpuset for its entire lifetime. Adoption will pull it back
to 201 or 203 when VM A submits I/O.
> In an ideal setting, the nvme controller would support $NR_CPUS worth of
> io queues [...] map cpu-queue with perfect alignment.
Right, and that is the case we cannot reach here. The controller tops out
at 128 io queues while the host has 384 CPUs, so the three to one folding
is unavoidable and some groups end up straddling a tenant boundary.
> I think that for controllers with low queue count this approach would not
> work very well. Plus the cpu<->queue mapping by the block layer is not
> guaranteed to align with how workloads/VMs are scheduled on the host...
Agreed on both. A queue whose cpu group straddles two busy tenants is only
partially addressed, since adoption then selects which tenant hosts that
shared queue rather than removing the sharing. It fixes the common case
where cpu group sits within one VM's cpuset and the case where only one
tenant on the queue is active, and it is strictly better than the connect
time pick, which chooses that queue's io_cpu before any I/O and can land on
any member. Solving the common case also lets the control plane, which owns
the VM cpusets, affinitize each VM's virtio-blk submitter threads away from
CPUs whose mapped queue would land socket work on a neighbor.
> To me the problem statement sounds to me like you are looking to get a
> "private" controller (set of IO queues) for each VM. [...]
> nvme connect -t tcp -a <traddr> --cpu-groups=0-63,64-127,128-191,192-255
A private controller or a cpu-groups partition fits workloads where the CPU
partition is known and stable at connect time. Our production constraint
does not afford us that luxury. The nvme-tcp controllers are host wide and
shared, the tenant VMs are namespaces under one multipath subsystem, and the
control plane creates, destroys and bin packs VMs continuously on these core
dense compute hosts. A cpu-groups mask fixed at connect would be stale as
soon as the tenant mix changes, and re-partitioning would mean reconnecting
a controller that many live tenants depend on. We also cannot give each
tenant its own controller, because the overhead on the host adds up and we
run the risk of exceeding the target's connection limits. Adopting the
submitting CPU needs no per tenant configuration and follows the tenant
churn on its own, which is why it fits a long lived shared connection whose
tenant set is not known at connect time.
Next the fleet operators who care about VM steal time already steer NIC IRQs,
NAPI to a dedicated housekeeping CPU pool outside the tenant cpusets, so the
common network processing overhead Nilay raised does not run on tenant CPUs.
With that in place the remaining contributor that still lands on a
tenant's CPU is nvme_tcp_io_work, which this patch pins to the submitting
cpu and keeps within the tenant's cpuset.
Thanks,
Saravanan D.