Re: [PATCH v1] tty: n_tty: use kvzalloc/kvfree for line discipline data
From: Greg KH
Date: Mon Aug 17 2026 - 11:42:13 EST
On Mon, Aug 17, 2026 at 04:16:53PM +0200, Greg KH wrote:
> On Mon, Aug 17, 2026 at 09:55:26PM +0800, Xin Chen wrote:
> > BT enable fails intermittently with -ETIMEDOUT (-110). The kernel log
> > shows the HCI Read Local Version command was sent and the firmware
> > replied with status 0x00 (logged by hci_req_cmd_complete() BT_DBG),
> > but the waiter in __hci_cmd_sync_sk() never woke up and timed out
> > after 10 s:
> >
> > bluetooth hci0: Opcode 0xfc00 // __hci_cmd_sync_sk
> > bluetooth hci0: opcode 0xfc00 plen 1 // hci_cmd_sync_add
> > bluetooth hci0: skb len 4 // hci_cmd_sync_alloc
> > bluetooth hci0: length 1 // hci_req_sync_run
> > Bluetooth: hci0 cmd_cnt 1 cmd queued 1 // hci_cmd_work
> > Bluetooth: hci0 type 1 len 4 // hci_send_frame
> > Bluetooth: opcode 0xfc00 status 0x00 // hci_req_cmd_complete
> > <-- req_skb NULL: req_complete_skb not set,
> > hci_cmd_sync_complete() never called,
> > req_status stays HCI_REQ_PEND -->
> > <-- 10 s later: wait_event_interruptible_timeout expires -->
> > bluetooth hci0: end: err -110 // __hci_cmd_sync_sk
> >
> > The root cause is that hci_send_cmd_sync() clones the sent command
> > into hdev->req_skb so that hci_req_cmd_complete() can locate the
> > registered completion callback. Under memory pressure this
> > skb_clone() fails, leaving hdev->req_skb NULL. The firmware reply
> > is received and processed, but hci_req_cmd_complete() finds NULL
> > req_skb, so hci_cmd_sync_complete() is never called, req_status
> > stays HCI_REQ_PEND, and the waiter times out with -ETIMEDOUT.
> >
> > The memory pressure is caused by n_tty_open(). When a BT UART
> > transport is opened, serdev_device_open() may be called multiple
> > times in quick succession, each triggering n_tty_open(). n_tty_open()
> > uses vzalloc() for the ~10 KB n_tty_data structure, which always
> > allocates page-by-page from the buddy order-0 free list. Repeated
> > vzalloc() calls drain enough order-0 pages that the subsequent
> > skb_clone(GFP_KERNEL) in hci_send_cmd_sync() cannot get a page.
>
> So you run out of memory? That feels wrong.
Also, you are papering over the real problem here. If this one
allocation is failing, what keeps the next one from failing and then the
skb will not be able to be allocated?
Why is the system so out of memory in this slab that this is happening?
What changed in the tty layer to cause this? Or did it happen
elsewhere?
And no cc: stable or Fixes: tag?
thanks,
greg k-h