Re: [RFC PATCH] nvme-pci: adaptively poll completions on busy queues

From: changfengnan

Date: Thu Jul 23 2026 - 08:15:59 EST


Two additional points: 
1. The current handling of `nvme_adaptive_mask_irq` is a bit complex.
 I don’t think it’s necessary to add a separate logic path specifically for
MSI; using `disable_irq_nosync` directly should suffice. 
Keith, what do you think? 
2. More detailed testing is currently underway, covering a wider
range of workloads and different drives. 
I’ll update the test data later.

Thanks.
Fengnan

> From: "Fengnan Chang"<changfengnan@xxxxxxxxxxxxx>
> Date:  Thu, Jul 23, 2026, 20:06
> Subject:  [RFC PATCH] nvme-pci: adaptively poll completions on busy queues
> To: <linux-nvme@xxxxxxxxxxxxxxxxxxx>, "Keith Busch"<kbusch@xxxxxxxxxx>, "Jens Axboe"<axboe@xxxxxxxxx>, "Christoph Hellwig"<hch@xxxxxx>, "Sagi Grimberg"<sagi@xxxxxxxxxxx>, "Bart Van Assche"<bvanassche@xxxxxxx>, "Andy Shevchenko"<andriy.shevchenko@xxxxxxxxx>, "Thomas Gleixner"<tglx@xxxxxxxxxx>, "Jun Zeng"<jun1.zeng@xxxxxxxxx>, "Gang Cao"<gang.cao@xxxxxxxxx>, "Jun I Jin"<jun.i.jin@xxxxxxxxx>, "Liang A Fang"<liang.a.fang@xxxxxxxxx>, "Yong Hu"<yong.hu@xxxxxxxxx>
> Cc: <linux-kernel@xxxxxxxxxxxxxxx>, "Fengnan Chang"<changfengnan@xxxxxxxxxxxxx>, "Guzebing"<guzebing@xxxxxxxxxxxxx>
> In high-IOPS scenarios, relying on interrupts to handle I/O operations
> can limit performance. This issue becomes particularly pronounced in
> multi-disk environments, where performance is constrained by the CPU’s
> interrupt-handling capacity.
> 
> Each Solidigm SB5PH27X038T device used for testing can deliver about 3.2M
> 4 KiB random-read IOPS.  Four devices therefore have about 12.8M IOPS of
> aggregate capability, but interrupt-driven completion topped out at 7.81M
> IOPS, or about 61% of that capability.
> 
> Add optional adaptive polling for busy interrupt-driven I/O queues.  After
> an interrupt drains the CQ, use the ring distance between last_sq_tail and
> cq_head as an approximate host-side inflight count.  If the count reaches
> a per-controller high watermark, let the threaded handler keep draining
> completions instead of returning to interrupt-driven completion right away.
> 
> The thread always checks for pending CQEs first.  When the CQ is empty, it
> compares the inflight count with a per-controller low watermark and sleeps
> for a configurable interval before checking again.  It returns to
> interrupt-driven completion after three consecutive empty-CQ checks below
> the low watermark or when the loop limit is reached.
> 
> The mode is disabled by default. It can be enabled at module load time.
> The high and low watermarks, empty-CQ interval, and loop limit are
> available through sysfs.  The defaults are 32 commands, 3 commands,
> 30 us, and 1000 iterations.
> 
> Tested with the default settings and 4 KiB random reads on four Solidigm
> SB5PH27X038T devices.  With adaptive polling off and on:
> 
>   - t/io_uring, depth 1024 and four workers per device:
>     7.810M and 12.780M aggregate IOPS (+63.64%).
>   - fio/libaio, iodepth 1024, numjobs 4 and completion batches of 32:
>     7.699M and 12.785M aggregate IOPS (+66.05%).
> 
> In a paired libaio and io_uring boundary sweep, QD8, QD16, QD32, and
> QD16/numjobs=4 changed by -0.15% to +0.48%.  QD64 improved by 28.61% to
> 59.75%.
> 
> Link: https://lore.kernel.org/linux-nvme/20260720145927.GA17986@xxxxxx/
> Signed-off-by: Guzebing <guzebing@xxxxxxxxxxxxx>
> Signed-off-by: Fengnan Chang <changfengnan@xxxxxxxxxxxxx>
> ---
>  Documentation/ABI/testing/sysfs-nvme |  63 +++++++
>  drivers/nvme/host/pci.c              | 266 +++++++++++++++++++++++++++
>  2 files changed, 329 insertions(+)
> 
> diff --git a/Documentation/ABI/testing/sysfs-nvme b/Documentation/ABI/testing/sysfs-nvme
> index 499d5f843cd43..9e22a51a0f5bb 100644
> --- a/Documentation/ABI/testing/sysfs-nvme
> +++ b/Documentation/ABI/testing/sysfs-nvme
> @@ -11,3 +11,66 @@ Description:
>                  (REPLACETLSPSK) with the target. After a reauthentication
>                  the value returned by tls_configured_key will be the new
>                  serial.
> +
> +What:                /sys/class/nvme/nvmeX/adaptive_poll_inflight_high
> +Date:                July 2026
> +KernelVersion:        7.3
> +Contact:        Linux NVMe mailing list <linux-nvme@xxxxxxxxxxxxxxxxxxx>
> +Description:
> +                The number of outstanding commands at which an interrupt-driven
> +                PCI I/O queue wakes its adaptive completion-drain thread. The value
> +                must be greater than adaptive_poll_inflight_low and smaller than
> +                the queue depth.
> +
> +                The default value is 32.
> +
> +What:                /sys/class/nvme/nvmeX/adaptive_poll_inflight_low
> +Date:                July 2026
> +KernelVersion:        7.3
> +Contact:        Linux NVMe mailing list <linux-nvme@xxxxxxxxxxxxxxxxxxx>
> +Description:
> +                The number of outstanding commands below which an adaptive
> +                completion-drain thread starts counting empty-CQ checks. The
> +                thread returns to interrupt-driven completion after three such
> +                checks in a row. The counter is reset whenever the number of
> +                outstanding commands reaches this value. The value must be nonzero
> +                and smaller than adaptive_poll_inflight_high.
> +
> +                The default value is 3.
> +
> +What:                /sys/class/nvme/nvmeX/adaptive_poll_interval_us
> +Date:                July 2026
> +KernelVersion:        7.3
> +Contact:        Linux NVMe mailing list <linux-nvme@xxxxxxxxxxxxxxxxxxx>
> +Description:
> +                The number of microseconds that an adaptive completion-drain
> +                thread waits before rechecking an empty CQ. The value must be
> +                nonzero.
> +
> +                The default value is 30.
> +
> +What:                /sys/class/nvme/nvmeX/adaptive_poll_max_loops
> +Date:                July 2026
> +KernelVersion:        7.3
> +Contact:        Linux NVMe mailing list <linux-nvme@xxxxxxxxxxxxxxxxxxx>
> +Description:
> +                The maximum number of iterations performed each time an adaptive
> +                completion-drain thread is woken. An iteration is counted whether
> +                or not the CQ has pending completions. The value must be nonzero
> +                and is sampled when the thread starts, so a new value applies to
> +                subsequent thread activations.
> +
> +                The thread may exit earlier when the low-inflight condition is met.
> +
> +                The default value is 1000.
> +
> +What:                /sys/module/nvme/parameters/use_adaptive_polling
> +Date:                July 2026
> +KernelVersion:        7.3
> +Contact:        Linux NVMe mailing list <linux-nvme@xxxxxxxxxxxxxxxxxxx>
> +Description:
> +                Global module parameter for adaptive polling on PCI I/O queues.
> +                Set to Y when loading the module, or on the kernel command line when
> +                the driver is built in, to enable adaptive polling for all NVMe PCI
> +                controllers. The default is N and the value cannot be changed after
> +                the driver is loaded.
> diff --git a/drivers/nvme/host/pci.c b/drivers/nvme/host/pci.c
> index 69932d640b537..296eadaf43199 100644
> --- a/drivers/nvme/host/pci.c
> +++ b/drivers/nvme/host/pci.c
> @@ -9,6 +9,7 @@
>  #include <linux/blkdev.h>
>  #include <linux/blk-mq-dma.h>
>  #include <linux/blk-integrity.h>
> +#include <linux/delay.h>
>  #include <linux/dmi.h>
>  #include <linux/init.h>
>  #include <linux/interrupt.h>
> @@ -21,6 +22,7 @@
>  #include <linux/nodemask.h>
>  #include <linux/once.h>
>  #include <linux/pci.h>
> +#include <linux/sched.h>
>  #include <linux/suspend.h>
>  #include <linux/t10-pi.h>
>  #include <linux/types.h>
> @@ -82,6 +84,17 @@ struct quirk_entry {
>  static int use_threaded_interrupts;
>  module_param(use_threaded_interrupts, int, 0444);
>  
> +static bool use_adaptive_polling;
> +module_param(use_adaptive_polling, bool, 0444);
> +MODULE_PARM_DESC(use_adaptive_polling,
> +                 "enable adaptive polling on busy interrupt-driven I/O queues");
> +
> +#define NVME_ADAPTIVE_POLL_INFLIGHT_HIGH_DEFAULT        32
> +#define NVME_ADAPTIVE_POLL_INFLIGHT_LOW_DEFAULT                3
> +#define NVME_ADAPTIVE_POLL_INTERVAL_US_DEFAULT        30
> +#define NVME_ADAPTIVE_POLL_MAX_LOOPS_DEFAULT        1000
> +#define NVME_ADAPTIVE_POLL_LOW_INFLIGHT_CHECKS        3
> +
>  static bool use_cmb_sqes = true;
>  module_param(use_cmb_sqes, bool, 0444);
>  MODULE_PARM_DESC(use_cmb_sqes, "use controller's memory buffer for I/O SQes");
> @@ -302,6 +315,11 @@ struct nvme_dev {
>          unsigned io_queues[HCTX_MAX_TYPES];
>          unsigned int num_vecs;
>          u32 q_depth;
> +        unsigned int adaptive_poll_inflight_high;
> +        unsigned int adaptive_poll_inflight_low;
> +        unsigned int adaptive_poll_interval_us;
> +        unsigned int adaptive_poll_max_loops;
> +        struct mutex adaptive_poll_config_lock; /* serializes watermark updates */
>          int io_sqes;
>          u32 db_stride;
>          void __iomem *bar;
> @@ -386,6 +404,7 @@ struct nvme_queue {
>  #define NVMEQ_SQ_CMB                1
>  #define NVMEQ_DELETE_ERROR        2
>  #define NVMEQ_POLLED                3
> +#define NVMEQ_ADAPTIVE_POLLING        4
>          __le32 *dbbuf_sq_db;
>          __le32 *dbbuf_cq_db;
>          __le32 *dbbuf_sq_ei;
> @@ -1649,6 +1668,99 @@ static irqreturn_t nvme_irq_check(int irq, void *data)
>          return IRQ_NONE;
>  }
>  
> +/*
> + * PCI I/O queues have one SQ and one CQ of the same depth.  last_sq_tail
> + * counts commands published to the controller while cq_head counts commands
> + * completed and consumed by the host, so their ring distance is the number
> + * of commands that can still produce completions.
> + */
> +static unsigned int nvme_adaptive_poll_inflight(struct nvme_queue *nvmeq)
> +{
> +        int inflight = READ_ONCE(nvmeq->last_sq_tail) -
> +                        READ_ONCE(nvmeq->cq_head);
> +
> +        if (inflight < 0)
> +                inflight += nvmeq->q_depth;
> +        return inflight;
> +}
> +
> +static void nvme_adaptive_mask_irq(struct nvme_queue *nvmeq, int irq)
> +{
> +        struct nvme_dev *dev = nvmeq->dev;
> +        struct pci_dev *pdev = to_pci_dev(dev->dev);
> +
> +        if (pdev->msi_enabled) {
> +                writel(BIT(nvmeq->cq_vector), dev->bar + NVME_REG_INTMS);
> +                readl(dev->bar + NVME_REG_INTMS);
> +        } else {
> +                disable_irq_nosync(irq);
> +        }
> +}
> +
> +static void nvme_adaptive_unmask_irq(struct nvme_queue *nvmeq, int irq)
> +{
> +        struct nvme_dev *dev = nvmeq->dev;
> +        struct pci_dev *pdev = to_pci_dev(dev->dev);
> +
> +        clear_bit(NVMEQ_ADAPTIVE_POLLING, &nvmeq->flags);
> +        if (pdev->msi_enabled) {
> +                writel(BIT(nvmeq->cq_vector), dev->bar + NVME_REG_INTMC);
> +                readl(dev->bar + NVME_REG_INTMS);
> +        } else {
> +                enable_irq(irq);
> +        }
> +}
> +
> +static irqreturn_t nvme_irq_adaptive_poll(int irq, void *data)
> +{
> +        struct nvme_queue *nvmeq = data;
> +        unsigned int max_loops =
> +                READ_ONCE(nvmeq->dev->adaptive_poll_max_loops);
> +        unsigned int low_inflight_checks = 0;
> +
> +        do {
> +                if (nvme_cqe_pending(nvmeq)) {
> +                        nvme_irq(irq, data);
> +                        low_inflight_checks = 0;
> +                        cond_resched();
> +                        continue;
> +                }
> +                if (nvme_adaptive_poll_inflight(nvmeq) <
> +                    READ_ONCE(nvmeq->dev->adaptive_poll_inflight_low)) {
> +                        if (++low_inflight_checks >=
> +                            NVME_ADAPTIVE_POLL_LOW_INFLIGHT_CHECKS)
> +                                break;
> +                } else {
> +                        low_inflight_checks = 0;
> +                }
> +                fsleep(READ_ONCE(nvmeq->dev->adaptive_poll_interval_us));
> +        } while (--max_loops);
> +
> +        nvme_adaptive_unmask_irq(nvmeq, irq);
> +        return IRQ_HANDLED;
> +}
> +
> +static irqreturn_t nvme_irq_adaptive_check(int irq, void *data)
> +{
> +        struct nvme_queue *nvmeq = data;
> +        irqreturn_t ret;
> +
> +        if (test_bit(NVMEQ_ADAPTIVE_POLLING, &nvmeq->flags))
> +                return IRQ_HANDLED;
> +
> +        ret = nvme_irq(irq, data);
> +        if (ret != IRQ_HANDLED)
> +                return ret;
> +
> +        if (nvme_adaptive_poll_inflight(nvmeq) <
> +            READ_ONCE(nvmeq->dev->adaptive_poll_inflight_high))
> +                return ret;
> +
> +        set_bit(NVMEQ_ADAPTIVE_POLLING, &nvmeq->flags);
> +        nvme_adaptive_mask_irq(nvmeq, irq);
> +        return IRQ_WAKE_THREAD;
> +}
> +
>  /*
>   * Poll for completions for any interrupt driven queue
>   * Can be called from any context.
> @@ -2171,9 +2283,22 @@ static int queue_request_irq(struct nvme_queue *nvmeq)
>          struct pci_dev *pdev = to_pci_dev(nvmeq->dev->dev);
>          int nr = nvmeq->dev->ctrl.instance;
>  
> +        /*
> +         * INTMS/INTMC mask MSI vectors and only cover vectors 0-31.  They must
> +         * not be accessed while MSI-X is enabled, so mask MSI-X vectors through
> +         * the PCI IRQ chip.
> +         */
>          if (use_threaded_interrupts) {
>                  return pci_request_irq(pdev, nvmeq->cq_vector, nvme_irq_check,
>                                  nvme_irq, nvmeq, "nvme%dq%d", nr, nvmeq->qid);
> +        } else if (use_adaptive_polling && nvmeq->qid &&
> +            nvmeq->dev->num_vecs > 1 &&
> +            (pdev->msix_enabled ||
> +             (pdev->msi_enabled && nvmeq->cq_vector < 32))) {
> +                return pci_request_irq(pdev, nvmeq->cq_vector,
> +                                nvme_irq_adaptive_check, nvme_irq_adaptive_poll,
> +                                nvmeq,
> +                                "nvme%dq%d", nr, nvmeq->qid);
>          } else {
>                  return pci_request_irq(pdev, nvmeq->cq_vector, nvme_irq,
>                                  NULL, nvmeq, "nvme%dq%d", nr, nvmeq->qid);
> @@ -2808,6 +2933,134 @@ static ssize_t hmb_store(struct device *dev, struct device_attribute *attr,
>  }
>  static DEVICE_ATTR_RW(hmb);
>  
> +static ssize_t adaptive_poll_inflight_high_show(struct device *dev,
> +                                                struct device_attribute *attr,
> +                                                char *buf)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +
> +        return sysfs_emit(buf, "%u\n",
> +                          READ_ONCE(ndev->adaptive_poll_inflight_high));
> +}
> +
> +static ssize_t adaptive_poll_inflight_high_store(struct device *dev,
> +                                                 struct device_attribute *attr,
> +                                                 const char *buf, size_t count)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +        unsigned int value;
> +        int ret;
> +
> +        ret = kstrtouint(buf, 10, &value);
> +        if (ret)
> +                return ret;
> +        if (!value || value >= ndev->q_depth)
> +                return -EINVAL;
> +
> +        mutex_lock(&ndev->adaptive_poll_config_lock);
> +        if (value <= READ_ONCE(ndev->adaptive_poll_inflight_low)) {
> +                mutex_unlock(&ndev->adaptive_poll_config_lock);
> +                return -EINVAL;
> +        }
> +        WRITE_ONCE(ndev->adaptive_poll_inflight_high, value);
> +        mutex_unlock(&ndev->adaptive_poll_config_lock);
> +        return count;
> +}
> +static DEVICE_ATTR_RW(adaptive_poll_inflight_high);
> +
> +static ssize_t adaptive_poll_inflight_low_show(struct device *dev,
> +                                               struct device_attribute *attr,
> +                                               char *buf)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +
> +        return sysfs_emit(buf, "%u\n",
> +                          READ_ONCE(ndev->adaptive_poll_inflight_low));
> +}
> +
> +static ssize_t adaptive_poll_inflight_low_store(struct device *dev,
> +                                                struct device_attribute *attr,
> +                                                const char *buf, size_t count)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +        unsigned int value;
> +        int ret;
> +
> +        ret = kstrtouint(buf, 10, &value);
> +        if (ret)
> +                return ret;
> +        if (!value || value >= ndev->q_depth)
> +                return -EINVAL;
> +
> +        mutex_lock(&ndev->adaptive_poll_config_lock);
> +        if (value >= READ_ONCE(ndev->adaptive_poll_inflight_high)) {
> +                mutex_unlock(&ndev->adaptive_poll_config_lock);
> +                return -EINVAL;
> +        }
> +        WRITE_ONCE(ndev->adaptive_poll_inflight_low, value);
> +        mutex_unlock(&ndev->adaptive_poll_config_lock);
> +        return count;
> +}
> +static DEVICE_ATTR_RW(adaptive_poll_inflight_low);
> +
> +static ssize_t adaptive_poll_interval_us_show(struct device *dev,
> +                                              struct device_attribute *attr,
> +                                              char *buf)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +
> +        return sysfs_emit(buf, "%u\n",
> +                          READ_ONCE(ndev->adaptive_poll_interval_us));
> +}
> +
> +static ssize_t adaptive_poll_interval_us_store(struct device *dev,
> +                                               struct device_attribute *attr,
> +                                               const char *buf, size_t count)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +        unsigned int value;
> +        int ret;
> +
> +        ret = kstrtouint(buf, 10, &value);
> +        if (ret)
> +                return ret;
> +        if (!value)
> +                return -EINVAL;
> +
> +        WRITE_ONCE(ndev->adaptive_poll_interval_us, value);
> +        return count;
> +}
> +static DEVICE_ATTR_RW(adaptive_poll_interval_us);
> +
> +static ssize_t adaptive_poll_max_loops_show(struct device *dev,
> +                                            struct device_attribute *attr,
> +                                            char *buf)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +
> +        return sysfs_emit(buf, "%u\n",
> +                          READ_ONCE(ndev->adaptive_poll_max_loops));
> +}
> +
> +static ssize_t adaptive_poll_max_loops_store(struct device *dev,
> +                                             struct device_attribute *attr,
> +                                             const char *buf, size_t count)
> +{
> +        struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev));
> +        unsigned int value;
> +        int ret;
> +
> +        ret = kstrtouint(buf, 10, &value);
> +        if (ret)
> +                return ret;
> +        if (!value)
> +                return -EINVAL;
> +
> +        WRITE_ONCE(ndev->adaptive_poll_max_loops, value);
> +        return count;
> +}
> +static DEVICE_ATTR_RW(adaptive_poll_max_loops);
> +
>  static umode_t nvme_pci_attrs_are_visible(struct kobject *kobj,
>                  struct attribute *a, int n)
>  {
> @@ -2832,6 +3085,10 @@ static struct attribute *nvme_pci_attrs[] = {
>          &dev_attr_cmbloc.attr,
>          &dev_attr_cmbsz.attr,
>          &dev_attr_hmb.attr,
> +        &dev_attr_adaptive_poll_inflight_high.attr,
> +        &dev_attr_adaptive_poll_inflight_low.attr,
> +        &dev_attr_adaptive_poll_interval_us.attr,
> +        &dev_attr_adaptive_poll_max_loops.attr,
>          NULL,
>  };
>  
> @@ -3685,6 +3942,15 @@ static struct nvme_dev *nvme_pci_alloc_dev(struct pci_dev *pdev,
>                  return ERR_PTR(-ENOMEM);
>          INIT_WORK(&dev->ctrl.reset_work, nvme_reset_work);
>          mutex_init(&dev->shutdown_lock);
> +        mutex_init(&dev->adaptive_poll_config_lock);
> +        dev->adaptive_poll_inflight_high =
> +                NVME_ADAPTIVE_POLL_INFLIGHT_HIGH_DEFAULT;
> +        dev->adaptive_poll_inflight_low =
> +                NVME_ADAPTIVE_POLL_INFLIGHT_LOW_DEFAULT;
> +        dev->adaptive_poll_interval_us =
> +                NVME_ADAPTIVE_POLL_INTERVAL_US_DEFAULT;
> +        dev->adaptive_poll_max_loops =
> +                NVME_ADAPTIVE_POLL_MAX_LOOPS_DEFAULT;
>  
>          dev->nr_write_queues = write_queues;
>          dev->nr_poll_queues = poll_queues;
> -- 
> 2.39.5 (Apple Git-154)
>