Re: [PATCH 6.18.y] scsi: ufs: core: Re-arm the device command completion before submitting
From: Alice Chao
Date: Tue Oct 06 2026 - 05:56:03 EST
On Mon, 2026-10-05 at 13:54 +0200, Bean Huo wrote:
> The late CQE can come after the re-arm:
>
> timeout -> cleanup -> retry B -> re-arm -> send B -> A's CQE arrives,
> and then B
> is woken up early and fails, if B's own CQE in turn comes after C's
> re-arm, C
> fails the same way, and C may even read B's response as its own.
>
> Whether this stops depends on device latency against our retry path,
> not on the
> patch.
>
You are right. Re-arming only covers a late CQE that arrives while no
device command is in flight. The query retry wrappers resubmit right
away, so the next re-arm will usually happen before the previous CQE
lands, and the skew can cascade exactly as you describe. It is also
worse than an early wakeup: every device command shares the reserved
slot's response UPIU, so C can parse B's response as its own and
return a wrong value rather than an error.
> The root cause is that all device commands share hba->reserved_slot
> and the CQE
> carries only the tag. Since a successful SQ cleanup must post an
> ABORTED CQE,
> could the MCQ timeout path wait for and consume that CQE before
> returning -
> EAGAIN?
>
Agreed, that closes the window instead of narrowing it. In v2 the MCQ
timeout path will:
- after the SQ cleanup, wait (bounded) for the reserved slot's CQE -
the ABORTED one, or the regular one if the command completed
anyway - before returning;
- if no CQE shows up in time (e.g. cleanup failed, or
UFSHCD_QUIRK_MCQ_BROKEN_RTC), force a host reset and refuse device
commands outside the error handler until it has happened, so the
slot is not reused while that CQE can still arrive;
- keep reinit_completion() at submission time as a safety net.
> does this patch only covers a late CQE arriving while no device
> command is in
> flight. Did you test it with injected timeouts to see whether it
> recovers?
>
Yes, v1 only covers that case. And no, I have not tested it with
injected timeouts yet.
Before posting v2 I will run it with injected device command timeouts
in MCQ mode: periodic fake timeouts under a descriptor/attribute read
loop, with the values checked against known-good ones, and the case
where the CQE never arrives, to check that the host gets reset and
device commands recover. This will also show whether our controller
actually posts the ABORTED CQE after the SQ cleanup. I will run the
same injection on v1 and on the unpatched kernel for comparison and
put the results in the v2 changelog.
Thanks for the careful review.
Alice