[PATCH] ublk: keep a canceled FETCH round canceling until the server is gone

From: Ming Lei

Date: Tue Sep 29 2026 - 09:25:05 EST


A cancel completes a fetched io command and sets io->cmd to NULL, but
the io still counts as ready. Only ubq->canceling stops ublk_queue_rq()
from using the NULL io->cmd. Two paths clear that flag too early:

partial FETCH round recovery, two queues
------------------- --------------------
task A: FETCH tags 0-2 q0 ready: q0->canceling = false
task A exits: q0 task exits:
start_cancel(): mark queues start_cancel(): ub->canceling
tags 0-2: io->cmd = NULL is still set -> q0 not marked
task B: FETCH tag 3 q0: io->cmd = NULL
queue ready: canceling = false q1 ready
START_DEV END_USER_RECOVERY
read -> ublk_queue_cmd(NULL) read on q0 -> ublk_queue_cmd(NULL)

A FETCH round runs from the open of /dev/ublkcN until
ublk_reset_ch_dev() resets the queues. A canceled command can't be
fetched again in the same round. So the ready check only needs one
fact: did this round see a cancel? ub->canceling already records it:
ublk_start_cancel() sets it under cancel_mutex before any command is
taken. It is only cleared too early, when all queues are ready.

Fix:

- clear ub->canceling only in ublk_reset_ch_dev()
- when a queue gets ready, clear ubq->canceling only if ub->canceling
is not set, and check it under cancel_mutex

The check and the marking are ordered by cancel_mutex:

check first: queue cleared -> start_cancel() marks it again
-> command taken
cancel first: check sees ub->canceling -> queue stays canceling

While ub->canceling is set, no queue clears its flag. So the "all
queues marked" shortcut in ublk_start_cancel() is correct again.

Fixes: 728cbac5fe21 ("ublk: move device reset into ublk_ch_release()")
Fixes: 3f3850785594 ("ublk: fix batch I/O recovery -ENODEV error")
Cc: stable@xxxxxxxxxxxxxxx # v6.17+: needs cancel_mutex
Reported-by: Josef Bacik <josef@xxxxxxxxxxxxxx>
Closes: https://lore.kernel.org/linux-block/20260928-b4-ublk-cancel-stop-v1-0-4a4360232a46@xxxxxxxxxxxxxx/
Signed-off-by: Ming Lei <tom.leiming@xxxxxxxxx>
Assisted-by: LLM
---
drivers/block/ublk_drv.c | 54 ++++++++++++++++++++++++----------------
1 file changed, 33 insertions(+), 21 deletions(-)

diff --git a/drivers/block/ublk_drv.c b/drivers/block/ublk_drv.c
index 66eb55e7162e..38ed7d0e3979 100644
--- a/drivers/block/ublk_drv.c
+++ b/drivers/block/ublk_drv.c
@@ -334,6 +334,11 @@ struct ublk_device {
u16 nr_queue_ready;
bool unprivileged_daemons;
struct mutex cancel_mutex;
+ /*
+ * A cancel started in this FETCH round. Set by ublk_set_canceling(),
+ * cleared only by ublk_reset_ch_dev() when a new round starts. While
+ * it is set, no queue clears its ->canceling.
+ */
bool canceling;
pid_t ublksrv_tgid;
struct delayed_work exit_work;
@@ -2415,6 +2420,11 @@ static void ublk_reset_ch_dev(struct ublk_device *ub)
spin_unlock(&ubq->cancel_lock);
}

+ /* a new FETCH round starts, the queues stay canceling until ready */
+ mutex_lock(&ub->cancel_mutex);
+ ub->canceling = false;
+ mutex_unlock(&ub->cancel_mutex);
+
/* set to NULL, otherwise new tasks cannot mmap io_cmd_buf */
ub->mm = NULL;
ub->nr_queue_ready = 0;
@@ -3024,11 +3034,19 @@ static void ublk_reset_io_flags(struct ublk_queue *ubq, struct ublk_io *io)
}

/* reset per-queue io flags */
-static void ublk_queue_reset_io_flags(struct ublk_queue *ubq)
+static void ublk_queue_reset_io_flags(struct ublk_device *ub,
+ struct ublk_queue *ubq)
{
- spin_lock(&ubq->cancel_lock);
- ubq->canceling = false;
- spin_unlock(&ubq->cancel_lock);
+ /*
+ * A cancel in this FETCH round took a command which still counts as
+ * ready, so the queue has to stay canceling. ub->canceling is set
+ * under cancel_mutex before any command is taken: either we see it
+ * here, or the cancel marks this queue again later.
+ */
+ mutex_lock(&ub->cancel_mutex);
+ if (!ub->canceling)
+ ubq->canceling = false;
+ mutex_unlock(&ub->cancel_mutex);
ubq->fail_io = false;
ubq->force_abort = false;
}
@@ -3051,24 +3069,17 @@ static void ublk_mark_io_ready(struct ublk_device *ub, u16 q_id,
ub->nr_queue_ready++;

/*
- * Reset queue flags as soon as this queue is ready.
- * This clears the canceling flag, allowing batch FETCH commands
- * to succeed during recovery without waiting for all queues.
+ * Reset queue flags as soon as this queue is ready. Unless
+ * this round saw a cancel, this clears the canceling flag,
+ * allowing batch FETCH commands to succeed during recovery
+ * without waiting for all queues.
*/
- ublk_queue_reset_io_flags(ubq);
+ ublk_queue_reset_io_flags(ub, ubq);
}

- /* Check if all queues are ready */
- if (ublk_dev_ready(ub)) {
- /*
- * All queues ready - clear device-level canceling flag
- * and wake ublk_dev_ready() waiters.
- */
- mutex_lock(&ub->cancel_mutex);
- ub->canceling = false;
- mutex_unlock(&ub->cancel_mutex);
+ /* All queues ready - wake ublk_dev_ready() waiters */
+ if (ublk_dev_ready(ub))
wake_up_var(&ub->nr_queue_ready);
- }
}

static inline int ublk_check_cmd_op(u32 cmd_op)
@@ -4433,9 +4444,10 @@ static bool ublk_validate_user_pid(struct ublk_device *ub, pid_t ublksrv_pid)

/*
* Wait until all queues have fetched their I/O commands, and return with
- * ub->mutex held and readiness guaranteed: then every queue's ->canceling
- * is cleared. Ready may regress between wakeup and mutex_lock() (F_BATCH
- * UNPREP, daemon death), so re-check it under the mutex and wait again.
+ * ub->mutex held and readiness guaranteed. The queues stay canceling if
+ * this round saw a cancel, see ublk_queue_reset_io_flags(). Ready may
+ * regress between wakeup and mutex_lock() (F_BATCH UNPREP, daemon death),
+ * so re-check it under the mutex and wait again.
*/
static int ublk_wait_dev_ready_and_lock(struct ublk_device *ub)
{
--
2.55.0




Thanks,
Ming