[PATCH net v2 0/2] netpoll: fix PREEMPT_RT deferred transmit locking

From: Karl Mehltretter

Date: Wed Sep 30 2026 - 17:36:11 EST


netpoll's deferred transmit path can acquire two sleeping locks with hard
interrupts disabled on PREEMPT_RT. One is the sk_buff_head lock used by
the sender and worker. The other is the device transmit lock used by the
worker. They produce independent atomic-sleep reports.

Patch 1 gives the deferred skb queue a dedicated raw lock.

Patch 2 makes the worker try the device transmit lock and defer the skb
again when the lock is busy. It uses the queue helper added by patch 1.

Testing:

- Arm64 and x86_64 PREEMPT_RT W=1 builds passed.
- An unpatched Pi 400 emitted the queue-lock report five times during a
one-hour run. Ten boots with the complete fix passed, including one
3,606-second run. Every trigger and post-trigger marker arrived, Wi-Fi
remained usable, and neither atomic-sleep warning occurred.
- An arm64 PREEMPT_RT QEMU control reproduced both warnings under forced
transmit backpressure. Three boots with the complete fix passed.
- An x86_64 QEMU control reproduced the queue-lock warning on PREEMPT_RT.
Three fixed RT boots passed; the non-RT control and treatment each
passed three boots.
- An arm64 QEMU A/B test held the device transmit lock while queuing 64
skbs through a synthetic CYW43455 SDIO device. It used HZ=250 and ran
with PREEMPT_RT enabled and disabled. After releasing the lock, the
HZ / 10 version completed in 128-178 ms on RT and 108-109 ms on
non-RT. The one-jiffy version completed in 59-114 ms on RT and
28-55 ms on non-RT. All 64 skbs completed in every run.
- The exact v2 series and pending brcmfmac fix passed 10/10 counted
physical boots across a Pi 400 and Pi 500+: three RT and two non-RT
boots per board. Every lock-contention run freed all 64 skbs without a
timeout. Each board and kernel flavor completed a 1,800-second stream,
ten TXHI stop/drain/wake cycles and a Wi-Fi reconnect. No counted boot
reported an atomic-sleep warning, new lockdep splat, stall or lockup.

The A/B test applied the same pending brcmfmac netpoll fix to both
variants. The driver recorded zero or one dropped packet per run, and
host capture lost the same number; both netpoll variants showed this.

v1: https://lore.kernel.org/r/20260928064239.32456-1-kmehltretter@xxxxxxxxx/

Changes in v2:

- Retry device transmit lock contention on the next tick instead of
after HZ / 10. Keep HZ / 10 for a stopped transmit queue or a driver
which returns busy.
- Add the arm64 QEMU A/B results and combined Pi 400/Pi 500+ RT and
non-RT hardware results.

Karl Mehltretter (2):
netpoll: use a raw lock for the deferred transmit queue
netpoll: avoid blocking on the transmit lock in queue_process

include/linux/netpoll.h | 1 +
net/core/netpoll.c | 83 +++++++++++++++++++++++++++++++++++++----
2 files changed, 77 insertions(+), 7 deletions(-)

Range-diff against v1:

1: b95097f73f29 = 1: b95097f73f29 netpoll: use a raw lock for the deferred transmit queue
2: bb9c0adcd8a9 ! 2: 30f5c80cea9c netpoll: avoid blocking on the transmit lock in queue_process
@@ Commit message
process_one_work

Use HARD_TX_TRYLOCK() instead. If another CPU owns the lock, put the skb
- back at the head of the deferred queue, restore interrupts and retry
- after the existing HZ / 10 delay.
+ back at the head of the deferred queue, restore interrupts and retry on
+ the next tick. Keep the longer HZ / 10 backoff for a stopped transmit
+ queue or a driver which returns busy.

Fixes: 3640543df26f ("[PATCH] netpoll: fix netpoll lockup")
Cc: stable@xxxxxxxxxxxxxxx # 6.12+
@@ net/core/netpoll.c: static void queue_process(struct work_struct *work)
+ netpoll_txq_queue_head(npinfo, skb);
+ local_irq_restore(flags);
+
-+ schedule_delayed_work(&npinfo->tx_work, HZ / 10);
++ schedule_delayed_work(&npinfo->tx_work, 1);
+ return;
+ }
if (netif_xmit_frozen_or_stopped(txq) ||

base-commit: a7bfaba4823e3c165bb2004c74eff7c096672bc7
--
2.53.0