[PATCH net v2] net: do not let defer_count outlive an empty defer_list

From: Joseph Kain

Date: Thu Oct 08 2026 - 18:04:50 EST


__skb_defer_free_flush() can leave defer_count at or above skb_defer_max
with an empty defer_list. That state is absorbing: every producer then
takes the nodefer path without queueing, so the list never refills, and
the flush keeps returning early without clearing. Deferred freeing stays
off for that (cpu, node) until reboot, and cross-CPU frees fall back to
the page allocator, which is the contention deferral exists to avoid.

Two changes from commit 844c9db7f7f5 ("net: use llist for sd->defer_list")
combine to produce it.

skb_attempt_defer_free() increments before testing the limit and does not
undo the increment when it rejects, so a rejected free leaves the counter
raised with nothing queued.

The flush clears defer_count before llist_del_all() rather than together
with it under defer_lock as it did before. A remote producer that queues
between the two has its skb drained by that same llist_del_all() but its
increment retained.

That part is not cumulative: any later flush finding a non-empty list
clears the counter again, so lost increments do not build up across
flushes. Reaching the limit takes skb_defer_max - 1 queued increments
inside a single clear-to-drain window. At a small skb_defer_max that
needs nothing special. At the default of 128 the flusher has to be
interrupted or preempted between those two lines for long enough, which
is where PREEMPT_RT matters: local_bh_disable() does not disable
preemption there, so the flusher can be scheduled out at that point.

Once the counter is stuck at or above the limit with an empty list,
nothing lowers it again, and that half needs no race at all:

sysctl -w net.core.skb_defer_max=1 # frees increment, then reject
sysctl -w net.core.skb_defer_max=128 # does not recover

skb_defer_max has .extra1 = SYSCTL_ZERO and no .extra2, so 1 is settable.
>From v7.1 on, 0 is safe because proc_do_skb_defer_max() flips
skb_defer_disable_key and the producer returns before incrementing. On
6.18 through 7.0, which carry the llist conversion but not commit
08dc30de1a40 ("net: add skb_defer_disable_key static key"), setting 0 and
restoring the previous value strands the counter the same way -- and 0 is
the documented way to turn the feature off. Recovering needs a
skb_defer_max above the accumulated count, which an operator has no way
to determine.

Drain before clearing, so the error can only under-read, is bounded by
the producers in flight, and self-heals on the next flush. Clear the
counter on the empty-list path as well, so no count can outlive an empty
list and an already-stranded counter recovers.

llist_empty() has already loaded that cache line, so the added read is
free; the unlikely() keeps the store off the hot path, where it would
otherwise bounce the line against remote producers.

Fixes: 844c9db7f7f5 ("net: use llist for sd->defer_list")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Joseph Kain <jkain@xxxxxxx>
---
v2:
- Reworded the drift description: it is not cumulative, since any later
flush finding a non-empty list clears the counter again. Reaching the
limit takes skb_defer_max - 1 queued increments inside a single
clear-to-drain window, and said where PREEMPT_RT actually matters
(Eric Dumazet).
- Noted that 6.18..7.0 strand the counter via skb_defer_max=0, the
documented way to disable the feature, since they carry the llist
conversion but not 08dc30de1a40 (Eric Dumazet).
- Dropped the paragraph describing where this was first reproduced
(Eric Dumazet).
- Added Assisted-by.

Built for x86_64 and arm64, with and without CONFIG_PREEMPT_RT. On
v7.3-rc6 with skb_defer_max=1 and a cross-CPU TCP load, defer_count
reached 1918118 unpatched and did not recover when the default was
restored; with this patch it peaked at 1.

v1: https://lore.kernel.org/netdev/20261007180056.1394868-1-jkain@xxxxxxx/
net/core/dev.c | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index 18dc88990510..043a7f9166a9 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -6938,10 +6938,13 @@ static void __skb_defer_free_flush(struct skb_defer_node *sdn, int budget)
struct llist_node *free_list;
struct sk_buff *skb, *next;

- if (llist_empty(&sdn->defer_list))
+ if (llist_empty(&sdn->defer_list)) {
+ if (unlikely(atomic_long_read(&sdn->defer_count)))
+ atomic_long_set(&sdn->defer_count, 0);
return;
- atomic_long_set(&sdn->defer_count, 0);
+ }
free_list = llist_del_all(&sdn->defer_list);
+ atomic_long_set(&sdn->defer_count, 0);

llist_for_each_entry_safe(skb, next, free_list, ll_node) {
prefetch(next);
--
2.55.0