Re: [PATCH v16 0/9] blk: honor isolcpus configuration
From: Ionut Nechita (Wind River)
Date: Tue Oct 06 2026 - 10:04:02 EST
Hi Aaron,
Follow-up with the controlled data I promised, and a correction. I was
wrong to point at your series in my first mail. After a clean A/B the
control-plane stall is a shared blk-mq queue-contention problem on a
shared disk, reproducible on *plain* managed_irq. Your series does not
introduce it -- but it does remove the per-CPU-queue separation that
otherwise keeps it from happening, so I think there is still a real
I/O-QoS question for you. Details below, each claim backed by a debugfs
capture.
Setup
-----
Kernel: 6.18.15-rt (PREEMPT_RT), v16 applied
CPU: 2x Intel Xeon Gold 6338N, 128 CPUs, 2 NUMA nodes
112 isolated, 16 housekeeping (8 carrying managed IRQs)
Storage: Dell PERC / MegaRAID (megaraid_sas) fronting a SATA SSD
(Samsung PM893), passthrough. etcd and the fio target share
the SAME physical disk (sda4 / cgts-vg).
Workload: one fio pod, libaio iodepth=256 numjobs=4 direct=1,
SCHED_OTHER (NOT RT -- RT is not required to reproduce),
pinned to one isolated CPU (CPU 2) via the StarlingX
isolcpus device-plugin label (verified with ps -eLo psr).
Method: per-hctx in-flight snapshot from
/sys/kernel/debug/block/sda/hctx*/busy during the run,
cross-referenced with each queue's mq/*/cpu_list and with
etcd's CPU set (ps psr).
The decisive observable is not "etcd alive/dead" but which single
blk-mq queue the isolated pod's I/O lands on, and whether one of etcd's
CPUs is mapped to that same queue.
The clean A/B (identical kernel = plain managed_irq, identical fio 4k
randread, identical nr_requests=5089, fio always on isolated CPU 2;
only the device's blk-mq queue count differs)
----------------------------------------------------------------------
(B) default MSI-X -> 120 blk-mq queues, ~1 CPU per queue
All ~35 in-flight land on hctx61, a queue mapped to the isolated
submitter and NOT used by any etcd CPU.
-> no collision. etcd healthy. Device load ~82k IOPS 4k,
sda aqu-sz ~1020, r_await ~12.3 ms.
(C) megaraid_sas.msix_vectors=16 -> 15 blk-mq queues, ~8 CPUs each
All ~33 in-flight land on hctx7, mapped to CPUs {0,4,64,68}.
etcd runs on CPU0 and CPU68 -> same queue.
-> collision. etcd apply 2-7 s, "request timed out",
"context deadline exceeded", /registry/health failing.
Control plane down -- on plain managed_irq, WITHOUT your series.
Same kernel, same disk, same fio, same depth; the only thing that
changed is whether there were enough hardware queues for the isolated
CPU to get one that etcd does not also use. That is the whole bug.
Supporting data point on the strict kernel
------------------------------------------
I also have one capture on the managed_irq_strict image (earlier run,
so not parameter-matched to B/C: it was a 64 KB read profile, and
/sys/block/sda/queue/nr_requests was already capped to 64 -- that cap
was my doing, the only way I could keep etcd / kube-apiserver alive
enough for the node to stay up and for me to collect the capture at
all; the default 5089 took the control plane down hard). It is still
informative directionally: with 0006 active, sda is confined to 16
blk-mq queues on the HK CPUs, fio on isolated CPU 2 funnels all ~30
in-flight onto hctx0 == CPU1, and CPU1 is an etcd CPU -> even with that
nr_requests=64 cap in place, etcd apply still blows out to multi-second
with context-deadline / health-probe failures. So strict reproduces the
same collision by construction, and the nr_requests knob only partially
masks it there.
Conclusion
----------
The failure is not strict-vs-non-strict. It is whether the isolated
pod's submission queue coincides with a queue an etcd CPU also uses:
- On a device with enough hardware queues, plain managed_irq gives
the isolated CPU its own blk-mq queue, disjoint from etcd's, and
the heavy pod I/O does not head-of-line-block etcd (B).
- Shrink the queue count until CPUs must share a queue -- via a low
MSI-X count (C), or via your 0006 "use HK CPUs only" confinement
(strict) -- and the isolated pod's deep queue lands on a queue etcd
uses, and etcd starves.
So this is pre-existing shared-queue contention on a shared slow disk,
not a bug your series introduces. I retract the implication in my
previous mail. (It is also device-level, not CPU: during the stall 7 of
8 IRQ-HK cores are idle and the backlog is pure block-layer queueing,
aqu-sz ~1020 at ~70% util.)
Where I think the series is still relevant
------------------------------------------
0006 confines blk-mq to the HK CPUs unconditionally. On a
well-provisioned device (B, 120 queues) that is exactly the config that
*removes* the per-CPU-queue separation which was keeping isolated bulk
I/O off etcd's queues. In other words, strict turns a
"only on queue-starved devices" problem into "always, because every
isolated submission is forced onto the HK queue set the control plane
also lives on." That is a deliberate and correct part of the design --
isolated CPUs must not host the queues -- but it makes an I/O-QoS gap
unavoidable rather than incidental.
Mitigation
----------
Capping /sys/block/sda/queue/nr_requests on this device (megaraid_sas
with the SATA SSD exposed transparently / passthrough, e.g. 5089 -> a
small value in the 32-64 range -- lower for a slower disk) shortens the
head-of-line backlog and helps, but note it is weaker under strict/low-
queue configs: when all submissions pile on one etcd-shared queue, even
a depth of 32-64 on that single queue still hurts etcd. For production I'd
prefer cgroup v2 io.max / io.latency to protect control-plane I/O, or a
dedicated device for etcd.
Questions
---------
1. When 0006 forces all isolated-pod block I/O onto the HK queue set,
do you consider protecting co-located latency-critical services
(etcd/kube-apiserver) from that I/O an operator concern
(io.latency/io.weight, dedicated device), or is there a case for
the series to leave a per-CPU submission path / reserved HK queue
so bulk isolated I/O cannot head-of-line-block them?
2. Is there an assumed minimum HK queue count relative to the
aggregate block I/O of isolated pods? 8 IRQ-HK vs 112 isolated on
one shared disk makes the HK queue set trivial to saturate on
depth even while it is CPU-idle.
Raw debugfs hctx snapshots + iostat/etcd logs available on request.
Net: not a v16 regression; a shared-disk I/O-QoS gap that v16 makes
unavoidable by design, which may or may not be something you want to
address in the series rather than leave to operators.
Thanks, and sorry for the premature pointer in the first mail,
Ionut