Re: [PATCH 0/2] smbdirect: don't hang on netdev reconfiguration

From: Stefan Metzmacher

Date: Tue Sep 08 2026 - 12:37:34 EST


Hi Ammar,

thanks for the fixes!
I let the rdma maintainers comment on the patches in detail,
but it's good to fix these problems where they happen.

When I mount a CIFS share using SMB Direct over RoCEv2, I observe a kernel
thread hang if I "configure" its slave network device by taking its link
down and bringing it back up. Attempting to just `ls` the mount point in
this state gives EHOSTDOWN.

I believe the following is the root-cause of the hang: When the link is
taken down, the corresponding GID Table Entry is marked as pending deletion
and has its slave ndev set to NULL. All sends on RDMA connections still
using that GID Table Entry fail at MAD creation. SMB Direct eventually
detects this and tries to disconnect / reconnect to recover. Unfortunately,
disconnecting requires successfully sending either a DREQ or a DREP. Since
neither of them even post, the connection remains in the RDMA_CM_CONNECT
state, and no callback moves it out. The SMB Direct layer never gets the
RDMA_CM_EVENT_DISCONNECTED it's waiting for, and hangs.

Fix this in the CMA. If we call `rdma_disconnect` on a connected connection
and we fail to send both a DREP and a DREQ; disconnect, and thereby send
the `RDMA_CM_EVENT_DISCONNECTED` event to SMB Direct.

I tested this change in QEMU using RXE. I ran Ubuntu 26.04.1 with a
mainline kernel. On commit 9f0346dcbea3 ("Merge tag 'driver-core-7.3-rc2'
of git://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core"),
I reproduce the hang. With this patch applied, SMB Direct immediately
disconnects and reconnects when the slave device is brought down then up.
Listing and reading files from the mount point also work afterwards.

Artifacts for reproducing the hang and testing my fix are available at

https://github.com/ammrat13/linux-cifs
Given you have some automation to reproduce it I'm
wondering if you could also test the iwarp case, see
https://lore.kernel.org/linux-rdma/20260805000159.321645-2-yunseong.kim@xxxxxxxx/

That was reported as fix for ksmbd, but I guess it
will also happen for the case your're seeing with rxe.
And your fixes are likely also good for ksmbd.

Thanks!
metze