Re: [PATCH v16 0/9] blk: honor isolcpus configuration

From: Ionut Nechita (Wind River)

Date: Mon Oct 05 2026 - 11:32:04 EST



Hi Aaron,

An early heads-up while I continue testing v16, and a question. On a
small-housekeeping / large-isolated box I can drive the Kubernetes
control plane (etcd / kube-apiserver) into timeouts with heavy block I/O
from an isolated pod. I want to flag it now for context, but I have to
be upfront about the state of the evidence: so far I have only
reproduced it on *plain* isolcpus=managed_irq, i.e. WITHOUT your series.
I am reinstalling with managed_irq_strict to run the controlled A/B and
will follow up with that data. So this is not (yet) a claim about your
series -- it may well turn out to be a pre-existing shared-disk
contention issue in our config. I would value your read either way.

Platform
--------
Kernel: 6.18.15-rt (PREEMPT_RT)
CPU: 2x Intel Xeon Gold 6338N (Ice Lake-SP), 32C/64T each,
128 CPUs total, 2 NUMA nodes
Storage: Dell PERC / Broadcom MegaRAID (megaraid_sas), MR(8192MB),
fronting SATA SSDs (Samsung PM893, 6.0 Gb/s); ~0.5 GB/s per
link, not NVMe-class. Capped to 16 MSI-X / 15 blk-mq queues
via megaraid_sas.msix_vectors=16 for a legible repro.
Layout: etcd and the fio target share the SAME physical disk (sda):
etcd on cgts-vg/etcd-lv and the pod's /tmp overlay on
cgts-vg/kubelet-lv, both on sda4.
cmdline (relevant bits, CURRENT boot = non-strict):
nohz_full=2-3,6-57,60-61,66-67,70-121,124-125
isolcpus=nohz,domain,managed_irq,2-3,6-57,60-61,66-67,70-121,124-125
kthread_cpus=0-1,4-5,64-65,68-69
irqaffinity=58-59,62-63,122-123,126-127

=> isolated = 112 CPUs; housekeeping = 16, of which 8 carry IRQs.

Workload
--------
A single fio pod pinned to one isolated CPU via a device-plugin
extended resource (isolcpus kubernetes label), libaio, iodepth=256,
numjobs=4, direct=1, OS storage sweep (rand/seq read, mixed rw).

What I observe (on plain managed_irq)
-------------------------------------
During the heavy-throughput / write profiles the control plane times
out hard. From etcd (host service, /var/log/daemon.log):

apply request took too long took:"13.08s" expected-duration:"100ms"
... error:"etcdserver: request timed out"
timed out waiting for read index response ... timeout:"7s"
Failed to check current member's leadership: context deadline exceeded

kube-apiserver (static pod log) in the same window:

etcdserver: request timed out
apiserver was unable to write a ... response: http: Handler timeout
retrying of unary invoker failed ... context deadline exceeded

The failure is device-bound, not CPU-bound. mpstat over the 8 IRQ-HK
cores during the run shows 7 of 8 ~100% idle and exactly one core at
100% %iowait (not %sys/%soft) -- i.e. a single blk-mq queue draining to
the slow shared SATA SSD, with its submitter blocked, while the rest of
HK is idle. So this is not softirq/CPU saturation of the HK set; it is
I/O-level contention on the one disk etcd also uses.

Why I am mailing this thread specifically
-----------------------------------------
Your series changes which CPUs/queues serve isolated-core I/O, so it is
the natural place to ask the question even though my current data is on
plain managed_irq. My working hypothesis for the A/B is:

- Without strict, isolated-pod I/O uses the local per-CPU queue and
etcd uses its own; they hit the same SATA SSD but through different
blk-mq queues.
- With strict, all queues collapse onto the HK set, so isolated-pod
bulk I/O and etcd's fsync/read-index may share the same queue(s)
and head-of-line block each other, pushing the control plane over
the edge at lower load.

That is a hypothesis, not a result. The A/B (same fio, same disk,
only the isolcpus flag changing) will tell us whether strict actually
makes it worse, leaves it unchanged, or is irrelevant.

Questions for you, independent of the A/B outcome
-------------------------------------------------
1. Is there an assumed minimum housekeeping size relative to the
aggregate block I/O of isolated pods? With 8 IRQ-HK vs 112
isolated and a shared slow disk, the HK pool / its queues look
easy to overwhelm.
2. Have you considered an I/O-QoS angle (e.g. protecting
control-plane I/O via blkio io.latency/io.weight, or a reserved
housekeeping queue) so that confining isolated-pod I/O onto HK
cannot starve co-located latency-critical services? Or do you
consider that purely an operator/config concern?

I will reply to this thread with the strict-vs-non-strict A/B (etcd
apply-latency + mpstat + fio IOPS) once the reinstall is done, plus the
full fio sweep and irq-monitor captures. Happy to be told this is a
config problem on our side rather than anything to do with v16.

Thanks,
Ionut