Re: [PATCH v2 1/2] nvmet: defer setting ns->enabled to false in nvmet_ns_disable()

From: Nilay Shroff

Date: Thu Oct 01 2026 - 07:49:23 EST


On 10/1/26 11:25 AM, Shin'ichiro Kawasaki wrote:
On Sep 18, 2026 / 21:57, Nilay Shroff wrote:
nvmet_ns_disable() currently clears ns->enabled before draining
in-flight I/O references. This allows namespace configuration to be
changed while existing I/O requests can still hold a reference to the
namespace.

This can race with configuration of namespace attributes such as the
device path, UUID, NGUID etc. These attributes can be accessed by I/O
requests without holding subsys->lock and must not be modified while
such requests are still using the namespace.

In nvmet_ns_disable(), keep ns->enabled set while existing namespace
references are being drained, so namespace configuration remains blocked
until all in-flight I/O has completed. Set ns->enabled to false only
after the namespace references have been drained and the namespace
device has been disabled.

Introduce the NVMET_NS_IO_LIVE flag, which is set after the namespace
is successfully enabled in nvmet_ns_enable(). When nvmet_ns_disable()
starts, clear NVMET_NS_IO_LIVE so that the I/O path stops admitting new
I/O once the flag is cleared. Using test_and_clear_bit() in
nvmet_ns_disable() also prevents concurrent callers from starting a
second disable operation.

Signed-off-by: Nilay Shroff<nilay@xxxxxxxxxxxxx>
Recent blktests-ci trial runs for nvme-7.3 branch reported failure of nvme/052
[*]. It is required to repeat the test case 3 to 20 times to recreate the
failure on my test system. I bisected and find this patch is the trigger of the
failure.

Nilay, may I ask you to take a look in the failure? I'm not sure if this should
be addressed in kernel side of blktests side.

Thanks for the report! I looked into the failure and I think I found the root cause.
I couldn't reproduce it on my system, but I see there is a narrow race window that can
potentially trigger the observed failure.

I don't think this is a new race introduced by this patch. However, the changes in this
patch may have altered the timing enough to make the race easier to trigger.

In nvmet_ns_enable(), we currently queue the asynchronous namespace-change event
before setting ns->enabled and the corresponding NVMET_NS_ENABLED xarray mark.
Therefore, if the asynchronous namespace-change event is received and processed
by the host before the namespace is fully enabled on the target, the host may not
find the namespace while scanning the active namespace list.

To close this race, I think we should send the namespace-change notification only
after the namespace has been fully enabled in nvmet_ns_enable(). In particular,
we can move nvmet_ns_changed() after setting ns->enabled, the NVMET_NS_ENABLED
xarray mark, and NVMET_NS_IO_LIVE.

Could you please try the following change on your test system and let me know if
it fixes the failure? If this resolves the issue, I'll send a formal patch:

diff --git a/drivers/nvme/target/core.c b/drivers/nvme/target/core.c
index 8eea0a504308..da98e5cae7f6 100644
--- a/drivers/nvme/target/core.c
+++ b/drivers/nvme/target/core.c
@@ -626,11 +626,11 @@ int nvmet_ns_enable(struct nvmet_ns *ns)
if (ret)
goto out_pr_exit;

- nvmet_ns_changed(subsys, ns->nsid);
ns->enabled = true;
xa_set_mark(&subsys->namespaces, ns->nsid, NVMET_NS_ENABLED);
nvmet_debugfs_ns_setup(ns);
set_bit(NVMET_NS_IO_LIVE, &ns->flags);
+ nvmet_ns_changed(subsys, ns->nsid);
ret = 0;
out_unlock:
mutex_unlock(&subsys->lock);