[PATCH] Bluetooth: hci_sync: don't drain cmd_sync backlog on unregister
From: Nguyen Ngoc Thang
Date: Sun Sep 27 2026 - 11:18:35 EST
hci_unregister_dev() disables cmd_work and cmd_timer, then calls
hci_cmd_sync_clear(), whose cancel_work_sync() waits for
hci_cmd_sync_work() to return. That worker dequeues and runs every
entry on cmd_sync_work_list. With cmd_work disabled nothing reaches
the controller, so each queued HCI command waits the full
HCI_CMD_TIMEOUT (2s).
The backlog has no bound. Userspace can keep queuing MGMT commands
such as MGMT_OP_GET_CLOCK_INFO against a controller that doesn't
answer. Closing /dev/vhci then blocks in vhci_release() for
backlog * 2s:
INFO: task syz-executor:5749 blocked for more than 143 seconds.
cancel_work_sync
hci_cmd_sync_clear
hci_unregister_dev
vhci_release
In a local reproduction the backlog held more than 5000 entries,
which comes to hours of hang.
Stop the worker from taking new entries once HCI_UNREGISTER is set,
and wake any request still waiting with -ENODEV before cancelling
the work. The entries left over are destroyed with -ECANCELED by the
existing sweep in hci_cmd_sync_clear(). They stay on the list until
the work has stopped, so hci_cmd_sync_dequeue() and friends still see
them. A callback already running may still issue another command,
which delays unregister by at most one timeout per remaining command,
not by the whole backlog.
Fixes: 008ee9eb8a11 ("Bluetooth: hci_sync: Fix not processing all entries on cmd_sync_work")
Reported-by: syzbot+217e3f1283cafe80586e@xxxxxxxxxxxxxxxxxxxxxxxxx
Closes: https://syzkaller.appspot.com/bug?extid=217e3f1283cafe80586e
Signed-off-by: Nguyen Ngoc Thang <ngocthang2710.1999@xxxxxxxxx>
---
Notes (not for the changelog):
Two reproducers were tested on QEMU x86_64 with KASAN and PROVE_LOCKING,
on master @28c5139761e5 (v7.3-rc2-1134). The patch also applies cleanly
to bluetooth-next/master.
- syzbot C repro (converted from the syz program): it loops on
MGMT_OP_GET_CLOCK_INFO, which queues HCI_OP_READ_CLOCK.
- The repro from the cause bisection: it emits LE Create BIG Complete
events for an unknown BIG, and each event queues HCI_OP_LE_TERM_BIG.
That queueing path is why the bisection landed on the ISO BIS commit.
Each run lets the repro go for N seconds, then SIGKILLs it and times
how long it takes to be reaped, i.e. how long vhci_release() takes.
unpatched, both repros: still blocked after more than 10 minutes,
hung-task reports every 20s, same lock
state as the syzbot report
patched, both repros, N=8/30/60s: reaped in 0-1s
patched + debug print: backlog of 2641 entries (N=30s) and 5260
entries (N=60s); hci_cmd_sync_clear() took
1-2ms
No KASAN, lockdep or hung-task reports on the patched kernel.
net/bluetooth/hci_sync.c | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/net/bluetooth/hci_sync.c b/net/bluetooth/hci_sync.c
index 74e2b04c84b2..e41fc1427d34 100644
--- a/net/bluetooth/hci_sync.c
+++ b/net/bluetooth/hci_sync.c
@@ -312,6 +312,10 @@ static void hci_cmd_sync_work(struct work_struct *work)
while (1) {
struct hci_cmd_sync_work_entry *entry;
+ /* Leave the backlog to hci_cmd_sync_clear() */
+ if (hci_dev_test_flag(hdev, HCI_UNREGISTER))
+ break;
+
mutex_lock(&hdev->cmd_sync_work_lock);
entry = list_first_entry_or_null(&hdev->cmd_sync_work_list,
struct hci_cmd_sync_work_entry,
@@ -658,6 +662,8 @@ void hci_cmd_sync_clear(struct hci_dev *hdev)
{
struct hci_cmd_sync_work_entry *entry, *tmp;
+ /* cmd_work is disabled, the pending request can only time out */
+ hci_cmd_sync_cancel_sync(hdev, ENODEV);
cancel_work_sync(&hdev->cmd_sync_work);
cancel_work_sync(&hdev->reenable_adv_work);
--
2.43.0