[PATCH] RDMA/cma: wait for addr_handler() to finish in rdma_destroy_id()
From: Palla Raghunath
Date: Sun Oct 04 2026 - 07:26:32 EST
syzbot hit a use-after-free of id_priv in addr_handler(). The bad read
is in debug_mutex_unlock(), from the last mutex_unlock() of
handler_mutex, and the memory was freed by ucma_close() ->
rdma_destroy_id().
addr_handler() moves the state from RDMA_CM_ADDR_QUERY to
RDMA_CM_ADDR_RESOLVED (or RDMA_CM_ADDR_BOUND on error) under
handler_mutex. If rdma_destroy_id() runs at that point, it waits on
handler_mutex. When addr_handler() unlocks, the destroying task can take
the mutex before mutex_unlock() has returned. It then sees a state other
than RDMA_CM_ADDR_QUERY, so cma_cancel_operation() skips
rdma_addr_cancel(), and _destroy_id() frees id_priv. mutex_unlock() in
the work then touches the freed lock:
ib_addr work close()
addr_handler()
mutex_lock(handler_mutex)
ADDR_QUERY -> ADDR_RESOLVED
... rdma_destroy_id()
mutex_lock(handler_mutex)
mutex_unlock(handler_mutex)
owner cleared gets the mutex
state != ADDR_QUERY, no cancel
kfree(id_priv)
debug_mutex_unlock(lock) <- use-after-free
Documentation/locking/mutex-design.rst says mutex_unlock() may still
touch the mutex after another task has acquired it, so handler_mutex
can't be what keeps id_priv alive here. The comment in
cma_cancel_operation() assumes it can. Before commit 722c7b2bfead
("RDMA/{cma, core}: Avoid callback on rdma_addr_cancel()"),
addr_handler() held a reference on id_priv until after mutex_unlock(),
which covered this.
So in rdma_destroy_id(), call rdma_addr_cancel() before taking
handler_mutex if a resolve was ever started on this id. The req stays on
req_list until the callback returns, so rdma_addr_cancel() finds it and
cancel_delayed_work_sync() waits until addr_handler() is really done.
This doesn't deadlock with the work itself. When addr_handler() destroys
the id because the event handler returned non-zero, it uses
destroy_id_handler_unlock(), not rdma_destroy_id(). Event handlers also
run with handler_mutex held, so they can't call rdma_destroy_id() on
their own id anyway.
There's no reproducer from syzbot, so I made the window bigger with a
debug-only mdelay(1000) after the mutex_unlock() in addr_handler(),
followed by a read of handler_mutex.magic. A small test program creates
an id, resolves an address on an rxe device and closes the fd while the
work sits in that delay. Without this patch KASAN reports the same
slab-use-after-free as syzbot (1048 bytes into a kmalloc-2k object,
freed by ucma_close()). With it the test runs clean, and close() just
waits for the work to finish.
Fixes: 722c7b2bfead ("RDMA/{cma, core}: Avoid callback on rdma_addr_cancel()")
Reported-by: syzbot+ae549381b4daac2895b1@xxxxxxxxxxxxxxxxxxxxxxxxx
Closes: https://syzkaller.appspot.com/bug?extid=ae549381b4daac2895b1
Signed-off-by: Palla Raghunath <raghunathpalla.0209@xxxxxxxxx>
---
drivers/infiniband/core/cma.c | 19 +++++++++++++++----
1 file changed, 15 insertions(+), 4 deletions(-)
diff --git a/drivers/infiniband/core/cma.c b/drivers/infiniband/core/cma.c
index 337a49d1acf7..24235eb0c49b 100644
--- a/drivers/infiniband/core/cma.c
+++ b/drivers/infiniband/core/cma.c
@@ -1968,10 +1968,10 @@ static void cma_cancel_operation(struct rdma_id_private *id_priv,
/*
* We can avoid doing the rdma_addr_cancel() based on state,
* only RDMA_CM_ADDR_QUERY has a work that could still execute.
- * Notice that the addr_handler work could still be exiting
- * outside this state, however due to the interaction with the
- * handler_mutex the work is guaranteed not to touch id_priv
- * during exit.
+ * The addr_handler work can still be finishing its
+ * mutex_unlock() after it has left this state.
+ * rdma_destroy_id() waits for that before it takes
+ * handler_mutex.
*/
rdma_addr_cancel(&id_priv->id.route.addr.dev_addr);
break;
@@ -2121,6 +2121,17 @@ void rdma_destroy_id(struct rdma_cm_id *id)
struct rdma_id_private *id_priv =
container_of(id, struct rdma_id_private, id);
+ /*
+ * addr_handler() can still be in mutex_unlock(&handler_mutex) after
+ * it has moved the state on from RDMA_CM_ADDR_QUERY, and
+ * mutex_unlock() may touch the mutex even after we have taken it.
+ * Wait for the work to finish before we free id_priv. The req stays
+ * on req_list until the callback returns, so rdma_addr_cancel() will
+ * find it.
+ */
+ if (id_priv->used_resolve_ip)
+ rdma_addr_cancel(&id->route.addr.dev_addr);
+
mutex_lock(&id_priv->handler_mutex);
destroy_id_handler_unlock(id_priv);
}
--
2.34.1