Re: [RFC] PCI_IRQ_AFFINITY limits MSI-X allocation on 384 CPU / 1000+ NVMe system

From: Thomas Gleixner

Date: Thu Jul 30 2026 - 09:32:22 EST


On Wed, Jul 29 2026 at 11:59, santhosh kumar wrote:
> On Wed, Jul 29, 2026 at 4:22 AM Thomas Gleixner <tglx@xxxxxxxxxxxxx> wrote:
>> Nothing to see here. It's simply resource exhaustion.
>>
>> If you want that odd setup to be supported you have to talk to the NVME
>> people and not to a random list of folks which have absolutely nothing
>> to do with NVME.

I've cleaned up the CC list for you because your AI buddy ignored that
request.

> I am trying to create 1024 nvme devices , each device requesting 2 msix vectors.
> Total we need 1024*2=2048 msix vectors
> Total number of available device vectors: ~ (200 * 384) = 76800
> But still seeing following messages:
> [ 9166.370948] Interrupt reservation exceeds available resources
> [ 9167.455029] Interrupt reservation exceeds available resources
> [ 9167.455572] Interrupt reservation exceeds available resources
> [ 9167.455888] Interrupt reservation exceeds available resources
>
> ● Summary: Affinity Mask Generation in Linux 6.13

Copying and pasting the output of your AI buddy is a pretty useless
exercise simply because your AI buddy does not understand how all of
this works. Neither did you actually validate that his slop makes any
sense.

Aside of that reports want to be against the latest upstream kernel and
not against something which is 8 revisions behind and eventually lacks a
ton of updates.

But let me explain you why your AI buddy gets it wrong.

> The affinity mask is generated in multiple layers, here's the exact flow:

We all know how that works even without the wisdom of your AI buddy.

> The Problem Code:
> // lib/group_cpus.c - Creates restricted CPU masks
> grp_spread_init_one()
> {
> cpu = cpumask_first(nmsk); // Pick CPU 68 for nvme0q0
> cpumask_set_cpu(cpu, irqmsk); // Mask = {68} only!
> }

See below.

> // arch/x86/kernel/apic/vector.c - Can only use that mask
> assign_irq_vector_any_locked()
> {
> affmsk = irq_data_get_affinity_mask(irqd); // Gets {68}
> assign_vector_locked(irqd, affmsk); // ONLY tries CPU 68!
> // ❌ No fallback to CPU 200-383 even if they have free vectors

Again your AI buddy is wrong.

assign_irq_vector_any_locked() has a fallback if the mask does not
result in an allocatable interrupt. It tries to get a close vector and
falls through to the very end of the function, which does:

/* Try the full online mask */
return assign_vector_locked(irqd, cpu_online_mask);

That's why the function has _any_ in the name. This function has a full
fallback. and that function is irrelevant for managed interrupts. It
only is invoked for non-managed interrupts.

Managed interrupts go through a different code path and that has no
fallback by design because that's the fundamental property of managed
interrupts.

So if group_spread_one() inits the mask for the managed interrupt with
only CPU 68 set then there is no fallback because there is no other
choice. But that also means that the driver did allocate more than two
vectors. Why?

If it really only allocates only two, then one is non-managed and the
other one is managed. In that case both end up with the a cpumask of
0-383, because spreading _ONE_ interrupt over 384 CPUs simply results in
a affinity mask with all CPUs set.

Just for illustration with a system with 64 CPUs and two nodes, where
CPU 0-15 and 32-47 are on node 0, CPU 16-31 and 48-63 are on node 1, the
spreading results in:

nr interrupts | nr_masks | masks
-----------------------------------------------------------------------
1 | 1 | [0-63]
2 | 2 | [0-15,32-47] [16-31,48-63]
4 | 4 | [0-7,32-39] [8-15,40-47] [16-23,48-55] ...
...
32 | 32 | [0,32] [1,33] ... [16,48] [17,49] ...
64 | 64 | [0] [1] [2] ....

That's the basic principle of the spreading mechanism. It's pretty
obvious, no?

Can you now explain me how that mechanism ends up creating a affinity
mask with a single CPU set if there is only _ONE_ interrupt to be
spreaded out?

I doubt it, but I can explain to you what happens in principle with NVME
and managed interrupts independent of the number of queues per device.

Managed interrupts ensure on startup, that for each CPU in the affinity
mask of each managed interrupt there is a vector guaranteed
available. That means:

nr interrupts | nr_masks | CPUs per mask | vectors | vectors
| | | per CPU | total
------------------------------------------------------------------------
1 | 1 | 64 | 1 | 64
2 | 2 | 32 | 1 | 64
4 | 4 | 16 | 1 | 64
...
32 | 32 | 2 | 1 | 64
64 | 64 | 1 | 1 | 64

So depending on the other interrupts allocated this will fail somewhere
around 200 devices for sure.

I'm sure it's the same on your side because that's independent of the
number of CPUs, even if you failed to provide that information.

As I told you before it's simple math and resource exhaustion due to the
way how NVME is implemented and the x86 vector space limitations.

If you want that to behave differently, then you have to talk to the
NVME people as I told you before.

Thanks,

tglx