RE: [RFC PATCH 0/3] nvme-pci: yield completions under scheduler pressure

From: Michael Kelley

Date: Fri Oct 09 2026 - 01:54:43 EST


From: Naman Jain <namjain@xxxxxxxxxxxxxxxxxxx> Sent: Thursday, October 8, 2026 10:06 PM
>
> On systems with several fast NVMe controllers, completion interrupts can
> keep returning to the same CPUs faster than scheduled work can run. Each
> handler may drain only a small number of completions, but the combined
> interrupt stream can still prevent scheduler and watchdog progress.

See this recent proposal [1] that sounds like it is addressing the same or a
similar issue. And there is this [2] more global approach. It's worthwhile to read
through the discussion on both threads. I haven't done a detailed comparison
of either vs. your proposal.

Michael

[1] https://lore.kernel.org/linux-nvme/20260818033846.53790-1-changfengnan@xxxxxxxxxxxxx/
[2] https://lore.kernel.org/lkml/20260819124341.4185621-1-lrizzo@xxxxxxxxxx/

>
> This series keeps the existing hardirq completion path for normal I/O.
> For an I/O queue with a dedicated MSI-X vector, if an interrupt arrives
> with a completion pending while need_resched() is set, mask that vector
> and hand the queue to irq_poll. The poller drains bounded batches and
> returns the queue to interrupt mode as soon as it catches up.
>
> The completion-path transitions are:
>
> +-----------------+
> | normal IRQ mode |
> +-----------------+
> |
> CQE pending and need_resched()
> |
> v
> mask vector, schedule irq_poll
> |
> v
> +--------------------------------+
> +---->| drain up to one poll budget |
> | +--------------------------------+
> | | |
> | full budget short batch
> | | |
> | v v
> | remain on poll list irq_poll_complete()
> | | mark handoff allowance
> | generic irq_poll clear poll ownership
> | invokes next pass enable vector
> | | |
> +--------------+ v
> +-----------------+
> | normal IRQ mode |
> +-----------------+
>
> A full-budget return deliberately stays on the irq_poll list. The generic
> irq_poll core invokes another bounded pass. Eventually a pass consumes
> less than its budget, takes the short-batch path, and restores interrupt
> mode.
>
> An MSI-X message may have been latched while the vector was masked, but
> not every MSI-X hierarchy can report pending state. The handoff allowance
> therefore does not claim that an interrupt is known to be stale. It lets
> exactly one empty interrupt after unmasking return IRQ_HANDLED. If a new
> completion is present, the normal completion path handles it.
>
> The core idea is simple: when a hardirq sees need_resched(), it masks the
> interrupt and offloads completion processing to irq_poll. The interrupt is
> unmasked once the CQ catches up. Everything else handles queue and
> controller state changes while irq_poll is active, because its asynchronous
> callback can outlive the hardirq that scheduled it.
>
> The lifecycle transitions are:
>
> normal IRQ mode / irq_poll
> |
> | set PAUSED, join IRQ and irq_poll, mask vector
> v
> +--------+
> | paused |---- temporary ----> resume after controller is ready
> +--------+ clear PAUSED, restore IRQ mode
> |
> | permanent stop
> | clear INIT, balance owned mask, synchronize replay
> v
> free registered IRQ
>
> INIT means the queue has an initialized irq_poll instance and a dedicated
> MSI-X vector. IRQ_POLL owns completion processing, while MASKED records the
> matching disable_irq_nosync(). PAUSED blocks IRQ handling and new admission.
> HANDOFF allows one empty interrupt after unmasking. A temporary pause keeps
> MASKED set for resume; permanent stop keeps PAUSED set through IRQ removal.
>
> Admin, shared, legacy, and threaded interrupt paths are unchanged. Explicit
> polling never enters irq_poll; it only disables bottom halves while holding
> the shared CQ polling lock. The series adds no timer, rate sampler,
> workqueue, sysfs knob, module parameter, or IRQ thread.
>
> Why not use nvme.use_threaded_interrupts=1 instead?
>
> Threaded mode installs nvme_irq_check() as the primary handler and
> nvme_irq() as the IRQ thread. When the primary handler finds a pending
> completion, it returns IRQ_WAKE_THREAD, so completion queue draining runs
> through a schedulable task rather than the normal hardirq fast path.
>
> That remains a useful opt-in mode and this series does not change it. On
> the affected VM it prevented the lockup, but it also reduced peak
> throughput. This series has a narrower goal: retain direct hardirq
> completion while the CPU is keeping up, and use bounded irq_poll work
> only after the scheduler has already asserted need_resched().
>
> The distinction is not that irq_poll is always faster than an IRQ thread.
> It is that the normal path is left alone, while the fallback creates a
> scheduling boundary only after a reschedule request is observed.
>
> need_resched() is used here as a best-effort pressure signal, not as an
> interrupt-flood detector or an elapsed-time guarantee. It is transient,
> and irq_poll may execute one bounded callback before the scheduler runs.
> The narrower claim is that, once an NVMe hardirq observes an already
> pending reschedule request, it stops unbounded CQ draining in hardirq.
>
> Previous NVMe attempts made the fallback the normal high-load path. An
> all-irq_poll RFC in 2016 reported an 8-10% single-core IOPS loss and a
> measurable low-queue-depth latency cost. A hardirq-first, threaded
> overflow series in 2019 fixed the Azure lockup, but Azure testing
> reported a substantial throughput loss.
>
> Relevant discussions from the past around this problem:
>
>
> https://lore.kernel.org/
> %2Flinux-nvme%2F1475660534-16681-1-git-send-email-
> sagi%40grimberg.me%2F&data=05%7C02%7C%7Cbbdfaf4318614114c57808df25c30
> 4c3%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C639271191764118549
> %7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwM
> CIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=1
> mVDHNgaLqX48N6F4S5ny1KC76E4Yv6CM4k%2BOI7V3ks%3D&reserved=0
>
> https://lore.kernel.org/
> %2Flkml%2F1566281669-48212-1-git-send-email-
> longli%40linuxonhyperv.com%2F&data=05%7C02%7C%7Cbbdfaf4318614114c57808
> df25c304c3%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C639271191764
> 147973%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAu
> MDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C
> &sdata=8LgU81UBs5FTTkkc7zbQo7Gx828IjfJO6Un9aebcdkM%3D&reserved=0
>
> https://lore.kernel.org/
> %2Flinux-nvme%2F20191209175622.1964-1-
> kbusch%40kernel.org%2F&data=05%7C02%7C%7Cbbdfaf4318614114c57808df25c3
> 04c3%7C84df9e7fe9f640afb435aaaaaaaaaaaa%7C1%7C0%7C639271191764169164
> %7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwM
> CIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=N
> IMpebg%2B3xIyJvnIk7FbUJANUxhUSFmljodgIAqswTY%3D&reserved=0
>
> The first patch selects IRQ_POLL. The second makes existing task-context
> CQ polling safe against the new softirq user. The third adds the NVMe
> queue handoff and lifecycle handling.
>
> The RFC starts with a per-queue weight of 256, matching the existing
> global irq_poll budget. This is a starting point, not a claim that 256 is
> the optimal NVMe value.
>
> Testing included:
>
> - full x86_64 kernel build and strict checkpatch
> - four-controller QEMU fio, full-disable S3, controller reset, and
> post-reset I/O
> - QEMU lockdep, prove-locking, and atomic-context validation
> - ARM64 Hyper-V VM with four 14-queue NVMe data controllers
> - natural irq_poll activation with threaded interrupts disabled
> - matching default-hardirq, IRQ-poll fallback, and threaded-IRQ comparison
> - four controller resets followed by successful reads
>
> Read-only 4 KiB random-read results follow. "Default hardirq" uses the
> unpatched base. The other modes use the same kernel built from this series:
> "IRQ-poll fallback" uses the default nvme.use_threaded_interrupts=0, while
> "threaded IRQs" uses nvme.use_threaded_interrupts=1. Low-load values are
> medians of two 30-second runs. High-load fallback and threaded values are
> medians of two and three 90-second runs respectively; the default-hardirq
> run locked up before it could complete.
>
> default IRQ-poll threaded
> hardirq fallback IRQs
> 1 disk, QD1
> IOPS 24.42K 24.96K 21.24K
> mean completion latency 38.94 us 38.17 us 44.60 us
> p99 completion latency 90.6 us 90.1 us 94.2 us
> p99.9 completion latency 94.7 us 94.7 us 103.9 us
>
> 4 disks, 128 jobs, QD256
> status lockup (26s) pass pass
> IOPS N/A 4.986M 3.165M
> bandwidth N/A 19.0 GiB/s 12.1 GiB/s
> mean completion latency N/A 6.544 ms 10.332 ms
> p99 completion latency N/A 14.483 ms 22.151 ms
> p99.9 completion latency N/A 22.544 ms 24.773 ms
>
> IRQ-poll fallback prevents hardirq starvation while avoiding the 15-37%
> IOPS loss and higher latency observed with threaded interrupts, by
> activating only under scheduler pressure.
>
> Naman Jain (3):
> nvme-pci: select IRQ_POLL
> nvme-pci: make completion queue polling softirq-safe
> nvme-pci: defer completions when rescheduling is needed
>
> drivers/nvme/host/Kconfig | 1 +
> drivers/nvme/host/pci.c | 295
> ++++++++++++++++++++++++++++++++++++++++++++--
> 2 files changed, 283 insertions(+), 13 deletions(-)
>
> base-commit: eea3fef32a9cf36abcb5975a5a594e4135a6b026
> --
> 2.43.0