[PATCH v4 05/15] iommu/arm-smmu-v3: Flush in-flight fault work on domain detach

From: Nicolin Chen

Date: Thu Sep 10 2026 - 19:20:12 EST


After the hardware queue is drained, an event may still be moving from the
IRQ thread to the IOPF workqueue, while earlier IOPF work is still running.

Synchronize the EVTQ and combined IRQs, then call iopf_queue_flush_dev().
This finishes all old-domain work before the IOMMU core frees the domain.
Skip synchronize_irq() after a drain timeout because a stuck consumer can
otherwise leave it waiting forever.

If arm_smmu_wait_for_queue_drained() times out, fault work may still be in
flight, and iopf_queue_remove_device() would free iopf groups that the work
also references. Skip the iopf teardown and leak the master_domain, rather
than risk a use-after-free.

The skip also leaks the iopf refcount, keeping the device enrolled on the
IOPF queue, which would strand its fault parameter on the queue list once
the device teardown frees dev->iommu, crashing a later iopf_queue_free().
Reclaim the enrollment in arm_smmu_release_device(), where all the attach
handles are gone so a straggler report cannot queue a new fault group.

Note that a residual race window remains between an iopf_queue_flush_dev()
and iopf_queue_remove_device(): a fault arriving in between still resolves
to the old attach handle, as the IOMMU core publishes a handle change only
after the driver ops return. This window predates the drain narrowing it,
and is only closable by an ordering fix in the IOMMU core. Furthermore, a
timed-out drain shares exactly the same window, given that it must keep the
device enrolled on the IOPF queue, where iopf_queue_remove_device() would
free the iopf groups that any in-flight fault work still references.

Fixes: cfea71aea921 ("iommu/arm-smmu-v3: Put iopf enablement in the domain attach path")
Cc: stable@xxxxxxxxxxxxxxx # v6.16
Co-developed-by: Barak Biber <bbiber@xxxxxxxxxx>
Signed-off-by: Barak Biber <bbiber@xxxxxxxxxx>
Co-developed-by: Stefan Kaestle <skaestle@xxxxxxxxxx>
Signed-off-by: Stefan Kaestle <skaestle@xxxxxxxxxx>
Signed-off-by: Malak Marrid <mmarrid@xxxxxxxxxx>
Assisted-by: Claude:claude-fable-5
Signed-off-by: Nicolin Chen <nicolinc@xxxxxxxxxx>
---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 46 +++++++++++++++++++--
1 file changed, 43 insertions(+), 3 deletions(-)

diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
index ef1fddad7868e..a915d8b0baf69 100644
--- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
+++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c
@@ -3402,6 +3402,7 @@ void arm_smmu_attach_release(struct arm_smmu_attach_state *state)
struct arm_smmu_master_domain *master_domain = state->old_master_domain;
struct arm_smmu_master *master = state->master;
struct arm_smmu_device *smmu = master->smmu;
+ bool timed_out = false;

lockdep_assert_not_held(&arm_smmu_asid_lock);
iommu_group_mutex_assert(master->dev);
@@ -3414,8 +3415,38 @@ void arm_smmu_attach_release(struct arm_smmu_attach_state *state)
* which the IOMMU core might free once this returns. Drain the hardware
* eventq, so that a pending event cannot turn into new fault work.
*/
- if (master_domain->using_iopf && master->stall_enabled)
- arm_smmu_wait_for_queue_drained(smmu, &smmu->evtq.q, false);
+ if (master_domain->using_iopf && master->stall_enabled) {
+ timed_out = arm_smmu_wait_for_queue_drained(smmu, &smmu->evtq.q,
+ false);
+ /*
+ * Ensure pending events have reached the IOPF queue, unless
+ * the drain timed out: a stuck consumer would also block an
+ * unbounded wait_event() inside the synchronize_irq().
+ */
+ if (!timed_out) {
+ if (smmu->evtq.q.irq)
+ synchronize_irq(smmu->evtq.q.irq);
+ /* Pending events might be in the combined_irq handler */
+ if (smmu->combined_irq)
+ synchronize_irq(smmu->combined_irq);
+ }
+ }
+
+ /* Lastly, flush the fault work that the drained events queued */
+ if (master_domain->using_iopf) {
+ iopf_queue_flush_dev(master->dev);
+
+ /*
+ * A timed-out drain may leave fault work in flight, and
+ * iopf_queue_remove_device() would free iopf groups that
+ * such work still references. Skip the iopf teardown and
+ * leak master_domain, rather than risk a UAF.
+ */
+ if (WARN_ON(timed_out)) {
+ state->old_master_domain = NULL;
+ return;
+ }
+ }

arm_smmu_disable_iopf(master, master_domain);
kfree(master_domain);
@@ -4399,7 +4430,16 @@ static void arm_smmu_release_device(struct device *dev)
{
struct arm_smmu_master *master = dev_iommu_priv_get(dev);

- WARN_ON(master->iopf_refcount);
+ /*
+ * A timed-out drain in arm_smmu_attach_release() leaks the refcount,
+ * keeping the device on the IOPF queue. Reclaim it here, since every
+ * attach handle is gone: a straggler fault can no longer queue a new
+ * fault group, so the queue turns stable once flushed.
+ */
+ if (WARN_ON(master->iopf_refcount)) {
+ iopf_queue_flush_dev(dev);
+ iopf_queue_remove_device(master->smmu->evtq.iopf, dev);
+ }

arm_smmu_disable_pasid(master);
arm_smmu_remove_master(master);
--
2.43.0