[PATCH v1.1] can: j1939: cancel pending address claim timers on rx release
From: Tetsuo Handa
Date: Tue Sep 29 2026 - 04:30:33 EST
syzbot is reporting "struct j1939_ecu" refcount leak, for
j1939_ecu_get(ecu);
priv->ents[ecu->addr] = ecu;
in j1939_ecu_map_locked() from j1939_ecu_timer_handler() can succeed
even after
priv->ents[ecu->addr] = NULL;
j1939_ecu_put(ecu);
in j1939_ecu_unmap_locked() from j1939_ecu_unmap_all() from
__j1939_rx_release() from j1939_netdev_stop() has completed.
unregister_netdevice: waiting for vxcan1 to become free. Usage count = 3
ref_tracker: netdev@ffff8880710f0700 has 1/2 users at
__netdev_tracker_alloc include/linux/netdevice.h:4496 [inline]
netdev_hold include/linux/netdevice.h:4525 [inline]
j1939_ecu_create_locked+0x1c9/0x400 net/can/j1939/bus.c:159
j1939_local_ecu_get+0xeb/0x220 net/can/j1939/bus.c:293
j1939_sk_bind+0x70a/0xc60 net/can/j1939/socket.c:529
__sys_bind_socket net/socket.c:1920 [inline]
__sys_bind+0x2e3/0x410 net/socket.c:1951
__do_sys_bind net/socket.c:1956 [inline]
__se_sys_bind net/socket.c:1954 [inline]
__x64_sys_bind+0x7a/0x90 net/socket.c:1954
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
ref_tracker: netdev@ffff8880710f0700 has 1/2 users at
__netdev_tracker_alloc include/linux/netdevice.h:4496 [inline]
netdev_hold include/linux/netdevice.h:4525 [inline]
j1939_priv_create net/can/j1939/main.c:140 [inline]
j1939_netdev_start+0x387/0xb20 net/can/j1939/main.c:268
j1939_sk_bind+0x946/0xc60 net/can/j1939/socket.c:506
__sys_bind_socket net/socket.c:1920 [inline]
__sys_bind+0x2e3/0x410 net/socket.c:1951
__do_sys_bind net/socket.c:1956 [inline]
__se_sys_bind net/socket.c:1954 [inline]
__x64_sys_bind+0x7a/0x90 net/socket.c:1954
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Fix this race condition by calling j1939_ecu_timer_cancel() before calling
j1939_ecu_unmap_all() from __j1939_rx_release() from j1939_netdev_stop().
Calling j1939_ecu_timer_cancel() here is safe without holding priv->lock
because the CAN RX path has already been unregistered via
j1939_can_rx_unregister() followed by synchronize_rcu(), meaning
no concurrent network events can mutate the priv->ecus list.
Reported-by: syzbot+e2af46126e0644cbebdd@xxxxxxxxxxxxxxxxxxxxxxxxx
Closes: https://syzkaller.appspot.com/bug?extid=e2af46126e0644cbebdd
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928193312.553632-1-mkl%40pengutronix.de # [PATCH net 06/22]
Assisted-by: Gemini-Pro gpt-6-astra opus-5-5
Signed-off-by: Tetsuo Handa <penguin-kernel@xxxxxxxxxxxxxxxxxxx>
---
Changes in v1.1:
Call synchronize_rcu() after j1939_can_rx_unregister().
net/can/j1939/main.c | 15 +++++++++++++++
1 file changed, 15 insertions(+)
diff --git a/net/can/j1939/main.c b/net/can/j1939/main.c
index 5e5e6c228f22..8706680e2458 100644
--- a/net/can/j1939/main.c
+++ b/net/can/j1939/main.c
@@ -212,8 +212,23 @@ static void __j1939_rx_release(struct kref *kref)
{
struct j1939_priv *priv = container_of(kref, struct j1939_priv,
rx_kref);
+ struct j1939_ecu *ecu, *tmp;
j1939_can_rx_unregister(priv);
+
+ /* can_rx_unregister() uses call_rcu() internally and is asynchronous.
+ * We must wait for an RCU grace period to ensure that any in-flight
+ * j1939_can_recv() instances on other CPUs have fully completed
+ * before we safely traverse the priv->ecus list without locks.
+ */
+ synchronize_rcu();
+ /* Cancel all pending address claim timers before unmapping the ECUs.
+ * This prevents an orphaned timer from re-mapping an ECU after the
+ * rx path has been completely torn down.
+ */
+ list_for_each_entry_safe(ecu, tmp, &priv->ecus, list)
+ j1939_ecu_timer_cancel(ecu);
+
j1939_ecu_unmap_all(priv);
j1939_priv_set(priv->ndev, NULL);
mutex_unlock(&j1939_netdev_lock);
--
2.52.0