[PATCH v16 8/9] genirq/affinity: Restrict managed IRQ affinity to housekeeping CPUs
From: Aaron Tomlin
Date: Thu Sep 10 2026 - 12:52:22 EST
At present, the managed interrupt spreading algorithm distributes vectors
across all available CPUs within a given node or system. On systems
employing CPU isolation (e.g. "isolcpus=managed_irq_strict"), this
behaviour defeats the primary purpose of isolation by routing hardware
interrupts (such as NVMe completion queues) directly to isolated cores.
Update irq_create_affinity_masks() to respect the housekeeping CPU mask.
By passing the HK_TYPE_MANAGED_IRQ_STRICT mask directly to the
topological distribution function (group_mask_cpus_evenly()), we ensure
that managed interrupts are kept strictly off isolated CPUs.
This patch additionally addresses the architectural constraints of
restricted vector distribution:
1. Vector limits and multi-set scaling
Updated irq_calc_affinity_vectors() to bound the maximum number
of allocated vectors to the weight of the housekeeping mask for
single-set drivers. For drivers providing a calc_sets()
callback, vector calculations continue to scale with the
driver's requested set sizes (maxvec - resv), preventing
unnecessary queue contention across distinct functional sets
while irq_create_affinity_masks() guarantees that all allocated
vectors remain strictly restricted to housekeeping CPUs.
2. Multi-set alignment and leak prevention
When isolation constraints result in fewer available masks than
requested vectors for a given set, the remaining vector slots
are padded with the housekeeping mask. This replaces the
historical irq_default_affinity padding, ensuring excess managed
queues do not leak interrupts onto isolated CPUs.
3. Minimum vector safety net
To prevent fatal -ENOSPC device probe aborts on heavily isolated
systems (where the housekeeping CPU count might be lower than a
device's structural minimum), the final vector calculation is
safeguarded to never drop below minvec. Queues will safely share
the available housekeeping CPUs instead of failing the probe.
4. Zero overhead
The housekeeping mask is conditionally assigned via a direct
pointer, completely avoiding temporary mask allocations (e.g.
alloc_cpumask_var) and bitwise operations when CPU isolation is
disabled. This guarantees zero performance or memory overhead
for standard configurations.
Signed-off-by: Aaron Tomlin <atomlin@xxxxxxxxxxx>
---
kernel/irq/affinity.c | 29 ++++++++++++++++++++++-------
1 file changed, 22 insertions(+), 7 deletions(-)
diff --git a/kernel/irq/affinity.c b/kernel/irq/affinity.c
index 78f2418a8925..7796882a567a 100644
--- a/kernel/irq/affinity.c
+++ b/kernel/irq/affinity.c
@@ -8,6 +8,7 @@
#include <linux/slab.h>
#include <linux/cpu.h>
#include <linux/group_cpus.h>
+#include <linux/sched/isolation.h>
static void default_calc_sets(struct irq_affinity *affd, unsigned int affvecs)
{
@@ -25,8 +26,10 @@ static void default_calc_sets(struct irq_affinity *affd, unsigned int affvecs)
struct irq_affinity_desc *
irq_create_affinity_masks(unsigned int nvecs, struct irq_affinity *affd)
{
- unsigned int affvecs, curvec, usedvecs, i;
+ unsigned int affvecs, curvec, usedvecs, i, j;
struct irq_affinity_desc *masks = NULL;
+ const struct cpumask *hk_mask = housekeeping_cpumask(HK_TYPE_MANAGED_IRQ_STRICT);
+ bool hk_enabled = housekeeping_enabled(HK_TYPE_MANAGED_IRQ_STRICT);
/*
* Determine the number of vectors which need interrupt affinities
@@ -70,19 +73,29 @@ irq_create_affinity_masks(unsigned int nvecs, struct irq_affinity *affd)
*/
for (i = 0, usedvecs = 0; i < affd->nr_sets; i++) {
unsigned int nr_masks, this_vecs = affd->set_size[i];
- struct cpumask *result = group_cpus_evenly(this_vecs, &nr_masks);
+ struct cpumask *result;
+ const struct cpumask *mask;
+ if (hk_enabled)
+ mask = hk_mask;
+ else
+ mask = cpu_possible_mask;
+
+ result = group_mask_cpus_evenly(this_vecs, mask,
+ &nr_masks);
if (!result) {
kfree(masks);
return NULL;
}
-
- for (int j = 0; j < nr_masks; j++)
+ for (j = 0; j < nr_masks; j++)
cpumask_copy(&masks[curvec + j].mask, &result[j]);
+ for (j = nr_masks; j < this_vecs; j++)
+ cpumask_copy(&masks[curvec + j].mask, mask);
+
kfree(result);
- curvec += nr_masks;
- usedvecs += nr_masks;
+ curvec += this_vecs;
+ usedvecs += this_vecs;
}
/* Fill out vectors at the end that don't need affinity */
@@ -117,8 +130,10 @@ unsigned int irq_calc_affinity_vectors(unsigned int minvec, unsigned int maxvec,
if (affd->calc_sets)
set_vecs = maxvec - resv;
+ else if (housekeeping_enabled(HK_TYPE_MANAGED_IRQ_STRICT))
+ set_vecs = cpumask_weight(housekeeping_cpumask(HK_TYPE_MANAGED_IRQ_STRICT));
else
set_vecs = cpumask_weight(cpu_possible_mask);
- return resv + min(set_vecs, maxvec - resv);
+ return max(minvec, resv + min(set_vecs, maxvec - resv));
}
--
2.55.0