[PATCH net v2] net: yield the CPU on every exit of the threaded NAPI poll loop
From: Vitaliy Sochnev
Date: Wed Sep 02 2026 - 15:10:18 EST
napi_threaded_poll_loop() reaches cond_resched() only when it is about to
iterate. When __napi_poll() clears repoll the loop breaks first, and
napi_thread_wait() returns without scheduling if work is already pending.
Under a receive load arriving as fast as it is drained the kthread never
yields, and on CONFIG_PREEMPT_NONE nothing else on that CPU runs.
Everything waiting for deferred work on that CPU then blocks. Deleting a
netdev hits three such waits - synchronize_net(), flush_all_backlogs() ->
flush_work() and rcu_barrier() from netdev_run_todo() - which is how this
was found. run_backlog_napi() runs the same loop, so the backlog kthread
can be held off the same way.
rcu_softirq_qs_periodic() does not cover it: it reports a quiescent state
but does not schedule, so the work items and the callbacks still wait.
Moving only that call is not enough either - the delay then migrates from
synchronize_net() to flush_work() and rcu_barrier().
Without the patch the kernel reports the thread holding the CPU:
rcu: INFO: rcu_sched self-detected stall on CPU
rcu: 0-....: (5999 ticks this GP) ... (t=6000 jiffies g=913 q=1218)
CPU: 0 UID: 0 PID: 203 Comm: napi/qdma_eth-0 Not tainted 6.18.44 #0
Hardware name: Nokia XG-040G-MD (UBI) (DT)
pc : __dma_sync_single_for_device+0x8/0xfc
Measured on that board (Airoha AN7581, quad core Cortex-A53, PREEMPT_NONE,
HZ=100, airoha_eth with threaded NAPI) while it terminates a 950 Mbit/s TCP
receive load. 30 minute runs, timing "ip link del" of a dummy interface:
before after
mean 31.09 s 0.20 s
worst 151.92 s 1.00 s
over 1 s 13 of 37 0 of 89
RCU stalls, classic 14 0
RCU stalls, expedited 43 0
packet rate 79363 p/s 79520 p/s
The packet rate is the control: the same work is done in both runs, so the
difference is not a lighter load. Three of the four CPUs sat around 65%
idle throughout the first run and did not help - the deferred work the
delete waits for is tied to the CPU the poll loop holds.
On preemption, raised in v1: same board and load, unpatched, two halves
differing only in that choice - worst "ip link del" 248.48 s with 8 stalls
under PREEMPT_NONE against 0.39 s and none under PREEMPT_LAZY. So lazy
preemption does hide the symptom, and PREEMPT_NONE and PREEMPT_VOLUNTARY
builds are what is left. It is not the NAPI thread being preempted more -
nonvoluntary_ctxt_switches on it is 4.6/s under LAZY against 13.9/s under
PREEMPT_NONE - so that is a measurement, not a mechanism. The missing yield
is there under either model.
The loop is unchanged in Linus's tree - net/core/dev.c at v7.3-rc1 is
identical here to net/main. The numbers come from 6.18 because that is the
only kernel this board runs: mainline carries en7581-evb alone, while the
SoC dtsi, the board DTS and the airoha_eth changes it needs are still out
of tree. On 6.18 the loop has no busy_poll_last_qs parameter, but with that
pointer NULL the two are the same code, so the change under test is this
one. The busy-poll path is unaffected; cond_resched() was already reached
there.
Reproducing needs the loop re-entered tens of thousands of times a second,
which depends on the driver and the shape of the load rather than on this
board: threaded NAPI can be turned on for any driver through
/sys/class/net/<dev>/threaded, and what it takes on top is little or no
interrupt coalescing. It did not reproduce on mtk_eth_soc, whose net_dim
moderation folds the same packet rate into far fewer interrupts.
Fixes: 29863d41bb6e ("net: implement threaded-able napi poll loop support")
Signed-off-by: Vitaliy Sochnev <sochnev.v.74@xxxxxxxxx>
---
v2:
- answered the tree question: the loop is unchanged in Linus's tree at
v7.3-rc1, and said why the numbers have to come from 6.18
- re-ran the A/B on a kernel with no out-of-tree module, so the splat and
the numbers now come from an untainted 6.18.44 build
- added the preemption-model measurement, and why PREEMPT_NONE and
PREEMPT_VOLUNTARY are still the exposed configs
- noted the repro is not board-specific
- shortened the comment; no other code change
v1: https://lore.kernel.org/netdev/20260814220427.623427-1-sochnev.v.74@xxxxxxxxx/
net/core/dev.c | 10 ++++++----
1 file changed, 6 insertions(+), 4 deletions(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index 38336858c168..5c7f8cdf8443 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -7924,11 +7924,13 @@ static void napi_threaded_poll_loop(struct napi_struct *napi,
gro_flush_normal(&napi->gro, HZ >= 1000);
local_bh_enable();
- /* Call cond_resched here to avoid watchdog warnings. */
- if (repoll || busy_poll_last_qs) {
+ if (repoll || busy_poll_last_qs)
rcu_softirq_qs_periodic(last_qs);
- cond_resched();
- }
+
+ /* napi_thread_wait() can return without scheduling, so yield on
+ * every exit, not only when the loop iterates.
+ */
+ cond_resched();
if (!repoll)
break;
--
2.55.0