[PATCH v4 0/2] nvme-multipath: fail I/O when no usable path exists
From: Krishna Iyer
Date: Thu Oct 01 2026 - 06:01:54 EST
On virtualization hosts we run NVMe/TCP targets with ctrl_loss_tmo=-1.
During a long fabric outage, I/O on a multipath namespace is queued
until a path returns, so a SIGKILLed VM process cannot exit while it
drains I/O to an unreachable target and is stuck in D state.
Patch 1 fixes nvme_available_path() so it stops queueing I/O that should
be failed, by default and without any flag:
- respect fast_io_fail_tmo: when a path exists but its failfast timer
has expired, fail instead of falling through to queue_if_no_path
- a LIVE path whose ANA state is inaccessible or persistent-loss no
longer counts as available (ANA change stays transient and queues)
- kick the requeue work from nvme_update_ns_ana_state() so parked I/O
is re-evaluated when an ANA transition leaves a namespace inaccessible
Patch 2 adds the fail_if_no_path attribute, an opt-in per-namespace
policy that fails parked and newly arriving I/O instead of queueing it
when no usable path exists. It is the namespace-scoped counterpart to
the controller-scoped fast_io_fail_tmo, needed because sibling
namespaces behind the same controllers must keep queueing while one
namespace whose consumer is gone releases its parked I/O. It sits
alongside delayed_removal_secs as a per-namespace policy in
nvme-multipath sysfs. Controller state is untouched and reconnects
continue.
This answers the question from the v3 review. The intent is exactly per
namespace differentiation, some namespaces stop waiting while others
keep waiting for the same controllers to reconnect. Patch 1 fails the
conditions the kernel can detect, an expired failfast timer or an
inaccessible ANA state. It cannot detect that the consumer of a
namespace has exited while its paths are still connecting, so that case
still queues by default and only an explicit per namespace opt in can
release it.
Changes since v3 [3]:
- Split the single v3 patch into a two-patch series: the
default-behavior fixes first, the opt-in attribute on top (Hannes,
Keith).
- The ANA inaccessible and persistent-loss handling is now the default
and no longer gated behind the flag (Keith).
- New: respect fast_io_fail_tmo. When a path is present but its failfast
timer has expired, nvme_available_path() now fails instead of falling
through to the queue_if_no_path window (Keith).
- New: kick the requeue work from nvme_update_ns_ana_state() so parked
I/O is re-evaluated when an ANA transition leaves a namespace
inaccessible (Keith).
- Patch 2 now carries only the fail_if_no_path attribute, rebased on
patch 1's corrected defaults: it gates the CONNECTING case and
overrides delayed_removal_secs, taking precedence over
queue_if_no_path when both are set.
- Document the fail_if_no_path vs delayed_removal_secs (queue-if-no-path)
interaction in the ABI entry (Hannes).
Changes since v2 [2]:
- Use the nvme_state_is_live() helper for the ANA state check, keeping
the explicit NVME_ANA_CHANGE carve-out so transient ANA transitions
still queue (Nilay).
- Return early when the stored value matches the current setting,
skipping synchronize_srcu() and the requeue kick (Nilay).
Changes since v1 [1]:
- Rename the attribute from fail_io_now to fail_if_no_path (Nilay).
- Make the policy persistent instead of self-clearing. Drop the clear
in nvme_mpath_set_live() so it stays set until userspace clears it.
- Enforce it inside nvme_available_path() by controller state. A
CONNECTING controller and a LIVE controller whose path ANA state is
inaccessible or persistent-loss stop counting as available. RESETTING
and ANA change keep queueing. With no controllers left it overrides
the delayed_removal_secs window.
Validated on hardware with a 6.17 backport of this series: parked I/O on
a SIGKILLed VM process failed within a second of enabling the policy and
the process was reaped, new I/O failed fast during the outage, the
policy persisted across path recovery, and disabling it restored
queueing.
[1] v1: https://lore.kernel.org/linux-nvme/20260904032605.65758-1-kiyer@xxxxxxxxx/
[2] v2: https://lore.kernel.org/linux-nvme/20260917231647.79956-1-kiyer@xxxxxxxxx/
[3] v3: https://lore.kernel.org/linux-nvme/20260923004959.88440-1-kiyer@xxxxxxxxx/
Krishna Iyer (2):
nvme-multipath: fix path state evaluation for failfast and ANA
nvme-multipath: add fail_if_no_path sysfs attribute
Documentation/ABI/stable/sysfs-nvme | 16 +++++++
drivers/nvme/host/multipath.c | 74 +++++++++++++++++++++++++----
drivers/nvme/host/nvme.h | 2 +
drivers/nvme/host/sysfs.c | 4 +-
4 files changed, 87 insertions(+), 9 deletions(-)
base-commit: 2ee54f01f07c0307deaf90ca8691a4643ae0357b
--
2.54.0