Re: [PATCH 0/9] ublk: fix dispatch to canceled io commands

From: Ming Lei

Date: Wed Sep 30 2026 - 13:11:20 EST


On Wed, Sep 30, 2026 at 02:17:31PM +0000, Josef Bacik wrote:
> On Tue, Sep 29, 2026 at 09:44:16AM -0500, Ming Lei wrote:
> > It looks two races: STOP_DEV vs. START_DEV, STOP_DEV vs. FETCH.
> >
> > Looks fast io path shouldn't be touched for fixing the races.
> >
> > > 2. A partial FETCH round whose task exits, once another task
> > > completes the round.
> > > 3. During recovery, the task of a queue which is ready already
> > > exiting before the last queue is ready.
> >
> > 2 and 3 could be solved in single simpler patch by making use of the
> > ub->canceling flag, and it is easier for backport.
>
> Agreed, yours is much simpler, and keeping the flag set for the whole
> FETCH round is the right model. I ran it on top of for-next (d70609a2f68c)
> with KASAN and lockdep through my reproducers and the ublk selftests. The
> oopses for 2 and 3 are gone, and recover_01-04, batch_01-03, generic_17,
> stress_01/02/05 and 60 batch QUIESCE_DEV/recover cycles pass.
>
> What's left for 2 and 3 is that the device still comes up. For 2,
> START_DEV returns 0 and the new disk fails every request. For 3,
> END_USER_RECOVERY returns 0, the device is LIVE, and every read on the
> queue whose task exited sits requeued forever, since the queue stays
> canceling and nothing kicks the requeue list. With ub->canceling
> covering the whole round that's a small check: return -ENODEV from
> START_DEV and END_USER_RECOVERY when ub->canceling is set, checked under
> cancel_mutex against publishing ub->ub_disk. The server can't fetch
> those commands again anyway. Patch 9 of my series did that on the old
> model, I'll redo it on top of yours.
>
> For 1, your patch alone still oopses in ublk_queue_rq() from the
> partition scan when START_DEV follows STOP_DEV, same as before. I'll
> respin my series as just that, on top of your patch and without touching
> the commit path: STOP_DEV marks the queues canceling and takes the
> fetched commands under ub->mutex, and FETCH marks its command cancelable
> before it publishes it, so a cancel from the control path never
> completes a command io_uring doesn't have on its cancelable list yet.
>
> For your patch:
>
> Tested-by: Josef Bacik <josef@xxxxxxxxxxxxxx>

Thanks for the test!

For STOP_DEV related races with STOP_DEV, START_DEV and FETCH, one simple
idea is to add internal device state of UB_STATE_STOPPING, which is set
in ublk_stop_dev() in case of any pending uring_cmd, and cleared in
ublk_reset_ch_dev() when the char dev is closed.

Then we can fail STOP_DEV, START_DEV and FETCH if UB_STATE_STOPPING is set.

I have written patches towards this direction, so far so good, pass all
selftests and survive in races of your reports, will post out for review
further.


Thanks,
Ming