Re: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure

From: Luigi Rizzo

Date: Fri Oct 09 2026 - 07:35:01 EST


On Fri, Oct 9, 2026 at 12:17 PM Naman Jain <namjain@xxxxxxxxxxxxxxxxxxx> wrote:
>
>
>
> On 10/9/2026 12:21 PM, Naman Jain wrote:
> >
> >
> > On 10/9/2026 11:24 AM, Michael Kelley wrote:
> >> From: Naman Jain <namjain@xxxxxxxxxxxxxxxxxxx> Sent: Thursday, October
> >> 8, 2026 10:06 PM
> >>>
> >>> On systems with several fast NVMe controllers, completion interrupts can
> >>> keep returning to the same CPUs faster than scheduled work can run. Each
> >>> handler may drain only a small number of completions, but the combined
> >>> interrupt stream can still prevent scheduler and watchdog progress.
> >>
> >> See this recent proposal [1] that sounds like it is addressing the
> >> same or a
> >> similar issue. And there is this [2] more global approach. It's
> >> worthwhile to read
> >> through the discussion on both threads. I haven't done a detailed
> >> comparison
> >> of either vs. your proposal.
> >>
> >> Michael
> >>
> >> [1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1-
> >> changfengnan@xxxxxxxxxxxxx/
> >> [2] https://lore.kernel.org/lkml/20260819124341.4185621-1-
> >> lrizzo@xxxxxxxxxx/
> >>
> >
> >
> > Thanks for sharing these Michael. I'll check more on these, and try it out.
> >
> > Regards,
> > Naman
>
> ++ authors of these two series, for awareness and if there is some
> configuration in their patches I should be trying to fix these lockup
> issues.

I see your thorough analysis below, thanks for the details.
Is there any reason why you used nvme.use_threaded_interrupts=0 ?

Setting to 1 moves the work to a kernel thread and seems to be the
best way to prevent too much work in the hardirq and the soft lockup.

I am surprised that GSIM fails to prevent the soft lockup, though:
under high load the sequence of events (starting from unmoderated)
should be the following:

1. HW sends MSIx interrupt, not blocked by anything
2. SW calls handle_fasteoi_irq() --> handle_irq_event() --> ... -->
action->handler() which starts processing
3. HW possibly sends another MSIx interrupt
4. SW handle_irq_event() completes
5. SW irq_start_moderation() starts the timer, __disable_irq() and
sets IRQD_IRQ_INPROGRESS | IRQD_MODERATED
6. if #3 happened, or another HW interrupt comes before the timer expires:
SW handle_fasteoi_irq() finds IRQD_IRQ_INPROGRESS | IRQD_MODERATED both set,
and does not call handle_irq_event(), postponing the call
7. when the timer expires, the pending interrupt is reinjected.

The above suggests that we should see some pauses between runs of
action->handler(),
at least for a single queue per CPU.

Do you have any way to look with perf and bpftrace at the CPU
processing interrupts,
to see what keeps it busy (the pause between 6 and 7 should keep it
well below 100%)
and try to verify whether the calls to hndle_irq_event() are actually spaced
as the sequence above describes?

Having multiple interrupts on the same CPU mean that the gap
between interrupts for one can be used by others, and that may
definitely make the CPU 100% busy in hardirq.

Again to address these cases, nvme.use_threaded_interrupts=1 would
greatly reduce the time spent in hardirq, moving the bulk of the work to softirq
hence at least becoming interruptible.

cheers
luigi