[PATCH] ALSA: hda: Reset pending verb count on response timeout

From: songxiebing

Date: Thu Oct 08 2026 - 23:41:13 EST


From: Bob Song <songxiebing@xxxxxxxxxx>

On some controllers the HDA link reports a verb timeout and the driver
falls back to polling mode. After that, every subsequent verb sent to
that codec address keeps timing out (roughly one second per verb), while
the hardware itself looks perfectly healthy: the RIRB write pointer keeps
advancing, interrupts are still delivered, and verbs that carry no reply
payload -- notably writes -- still take effect on the codec.

The reason is that snd_hdac_bus_send_cmd() increments
bus->rirb.cmds[addr] for every verb, but the counter is only decremented
by snd_hdac_bus_update_rirb() when a matching RIRB response is read, with
no rollback if the response never arrives.

When a controller cannot fall back to single-command mode, i.e.
chip->fallback_to_single_cmd is zero, azx_rirb_get_response() bails out
with -EIO immediately and never reaches the bus-reset / single_cmd
recovery that reinitializes CORB/RIRB and clears the counters (that path
is guarded by the same flag). The ACPI platform, Tegra and CIX
controllers never set that flag, so they are all exposed to this.

A single lost response therefore poisons bus->rirb.cmds[addr]
permanently: each later verb still gets one response that only brings the
count back down to the stale baseline of 1, so
snd_hdac_bus_get_response() keeps timing out even though the codec
answers normally.

Fix this by dropping the pending count for the codec address once the
verb is known to be dead. A late response is then handled as a spurious
response exactly as before, and the following verbs recover.

Signed-off-by: Bob Song <songxiebing@xxxxxxxxxx>
---
sound/hda/common/controller.c | 20 ++++++++++++++++----
1 file changed, 16 insertions(+), 4 deletions(-)

diff --git a/sound/hda/common/controller.c b/sound/hda/common/controller.c
index afec5c5546ec..7a0b641fba14 100644
--- a/sound/hda/common/controller.c
+++ b/sound/hda/common/controller.c
@@ -773,7 +773,7 @@ static int azx_rirb_get_response(struct hdac_bus *bus, unsigned int addr,
return 0;

if (hbus->no_response_fallback)
- return -EIO;
+ goto error;

if (!bus->polling_mode) {
dev_warn(chip->card->dev,
@@ -789,7 +789,7 @@ static int azx_rirb_get_response(struct hdac_bus *bus, unsigned int addr,
bus->last_cmd[addr]);
if (chip->ops->disable_msi_reset_irq &&
chip->ops->disable_msi_reset_irq(chip) < 0)
- return -EIO;
+ goto error;
goto again;
}

@@ -798,12 +798,12 @@ static int azx_rirb_get_response(struct hdac_bus *bus, unsigned int addr,
* phase, this is likely an access to a non-existing codec
* slot. Better to return an error and reset the system.
*/
- return -EIO;
+ goto error;
}

/* no fallback mechanism? */
if (!chip->fallback_to_single_cmd)
- return -EIO;
+ goto error;

/* a fatal communication error; need either to reset or to fallback
* to the single_cmd mode
@@ -823,6 +823,18 @@ static int azx_rirb_get_response(struct hdac_bus *bus, unsigned int addr,
hbus->response_reset = 0;
snd_hdac_bus_stop_cmd_io(bus);
return -EIO;
+
+ error:
+ /*
+ * The command will not get any response. Drop its pending count,
+ * otherwise bus->rirb.cmds[addr] stays non-zero forever and every
+ * later verb to this codec address keeps timing out, even though its
+ * own response is delivered normally.
+ */
+ scoped_guard(spinlock_irq, &bus->reg_lock) {
+ bus->rirb.cmds[addr] = 0;
+ }
+ return -EIO;
}

/*
--
2.25.1