[PATCH net] net: do not let defer_count outlive an empty defer_list
From: Joseph Kain
Date: Wed Oct 07 2026 - 14:03:26 EST
__skb_defer_free_flush() can leave defer_count at or above skb_defer_max
with an empty defer_list. That state is absorbing: every producer then
takes the nodefer path without queueing, so the list never refills, and
the flush keeps returning early without clearing. Deferred freeing stays
off for that (cpu, node) until reboot, and cross-CPU frees fall back to
the page allocator, which is the contention deferral exists to avoid.
Two changes from commit 844c9db7f7f5 ("net: use llist for sd->defer_list")
combine to produce it.
skb_attempt_defer_free() now increments before testing the limit and does
not undo the increment when it rejects, so a rejected free leaves the
counter raised with nothing queued.
The flush clears defer_count before llist_del_all() rather than together
with it under defer_lock as it did before. A remote producer that queues
between the two has its skb drained by that same llist_del_all() but its
increment retained, so the counter drifts upward until it reaches the
limit.
Neither needs PREEMPT_RT. The drift needs no preemption at all: the
flush normally runs in NAPI poll on the CPU that owns the list, while
skb_attempt_defer_free() only ever defers to a remote CPU, so producers
cannot be excluded by disabling preemption on the flusher. PREEMPT_RT
only widens the window, because local_bh_disable() does not disable
preemption, so the flusher can be scheduled out between the clear and
the drain.
The latch is also reachable with no race at all:
sysctl -w net.core.skb_defer_max=1 # frees increment, then reject
sysctl -w net.core.skb_defer_max=128 # does not recover
skb_defer_max has .extra1 = SYSCTL_ZERO and no .extra2, so 1 is settable.
0 is safe, because proc_do_skb_defer_max() flips skb_defer_disable_key
and the producer returns before incrementing, but that guard covers 0
only. Recovering from the latched state needs a skb_defer_max above the
accumulated count, which an operator has no way to determine.
Drain before clearing, so the error can only under-read, is bounded by
the producers in flight, and self-heals on the next flush. Clear the
counter on the empty-list path as well, so no count can outlive an empty
list and an already-stranded counter recovers.
llist_empty() has already loaded that cache line, so the added read is
free; the unlikely() keeps the store off the hot path, where it would
otherwise bounce the line against remote producers.
Found on an arm64 PREEMPT_RT system running the two commits backported to
6.1. With skb_defer_max lowered, defer_count reached 788,888 on one host
and 122,399 on another with the list empty throughout, and restoring the
default did not re-enable deferral.
Fixes: 844c9db7f7f5 ("net: use llist for sd->defer_list")
Signed-off-by: Joseph Kain <jkain@xxxxxxx>
---
net/core/dev.c | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index 18dc88990510..043a7f9166a9 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -6938,10 +6938,13 @@ static void __skb_defer_free_flush(struct skb_defer_node *sdn, int budget)
struct llist_node *free_list;
struct sk_buff *skb, *next;
- if (llist_empty(&sdn->defer_list))
+ if (llist_empty(&sdn->defer_list)) {
+ if (unlikely(atomic_long_read(&sdn->defer_count)))
+ atomic_long_set(&sdn->defer_count, 0);
return;
- atomic_long_set(&sdn->defer_count, 0);
+ }
free_list = llist_del_all(&sdn->defer_list);
+ atomic_long_set(&sdn->defer_count, 0);
llist_for_each_entry_safe(skb, next, free_list, ll_node) {
prefetch(next);
--
2.55.0