Re: [PATCH net] net: do not let defer_count outlive an empty defer_list
From: Eric Dumazet
Date: Wed Oct 07 2026 - 16:25:57 EST
Le mer. 7 oct. 2026 à 20:03, Joseph Kain <jkain@xxxxxxx> a écrit :
>
> __skb_defer_free_flush() can leave defer_count at or above skb_defer_max
> with an empty defer_list. That state is absorbing: every producer then
> takes the nodefer path without queueing, so the list never refills, and
> the flush keeps returning early without clearing. Deferred freeing stays
> off for that (cpu, node) until reboot, and cross-CPU frees fall back to
> the page allocator, which is the contention deferral exists to avoid.
>
> Two changes from commit 844c9db7f7f5 ("net: use llist for sd->defer_list")
> combine to produce it.
>
> skb_attempt_defer_free() now increments before testing the limit and does
> not undo the increment when it rejects, so a rejected free leaves the
> counter raised with nothing queued.
>
> The flush clears defer_count before llist_del_all() rather than together
> with it under defer_lock as it did before. A remote producer that queues
> between the two has its skb drained by that same llist_del_all() but its
> increment retained, so the counter drifts upward until it reaches the
> limit.
This is not cumulative: the next flush finding a non-empty list resets
the counter. To get stuck, (skb_defer_max - 1) increments must land
between the clear and the drain of a single flush. With small limits this
needs nothing special, but with the default of 128 the flusher has to be
interrupted or preempted at that point. Please reword this part and the
"needs no preemption at all" paragraph.
>
> Neither needs PREEMPT_RT. The drift needs no preemption at all: the
> flush normally runs in NAPI poll on the CPU that owns the list, while
> skb_attempt_defer_free() only ever defers to a remote CPU, so producers
> cannot be excluded by disabling preemption on the flusher. PREEMPT_RT
> only widens the window, because local_bh_disable() does not disable
> preemption, so the flusher can be scheduled out between the clear and
> the drain.
>
> The latch is also reachable with no race at all:
>
> sysctl -w net.core.skb_defer_max=1 # frees increment, then reject
> sysctl -w net.core.skb_defer_max=128 # does not recover
>
> skb_defer_max has .extra1 = SYSCTL_ZERO and no .extra2, so 1 is settable.
> 0 is safe, because proc_do_skb_defer_max() flips skb_defer_disable_key
> and the producer returns before incrementing, but that guard covers 0
> only. Recovering from the latched state needs a skb_defer_max above the
> accumulated count, which an operator has no way to determine.
>
> Drain before clearing, so the error can only under-read, is bounded by
> the producers in flight, and self-heals on the next flush. Clear the
> counter on the empty-list path as well, so no count can outlive an empty
> list and an already-stranded counter recovers.
>
> llist_empty() has already loaded that cache line, so the added read is
> free; the unlikely() keeps the store off the hot path, where it would
> otherwise bounce the line against remote producers.
>
> Found on an arm64 PREEMPT_RT system running the two commits backported to
> 6.1. With skb_defer_max lowered, defer_count reached 788,888 on one host
> and 122,399 on another with the list empty throughout, and restoring the
> default did not re-enable deferral.
Not sure if this paragraph is needed?
Lowered to which value? Have you seen this with the default?
If the value was 0, note that kernels without commit 08dc30de1a40
("net: add skb_defer_disable_key static key"), i.e. 6.18 to 7.0, have the
same issue as your skb_defer_max=1 example: setting 0 and restoring the
previous value leaves deferral disabled. This is worth mentioning, since
0 is the documented way to turn the feature off.
Please send a v2 with these changes (but please wait ~24 hours before V2)
>
> Fixes: 844c9db7f7f5 ("net: use llist for sd->defer_list")
> Signed-off-by: Joseph Kain <jkain@xxxxxxx>
This is your first linux contribution, are you sure you have not used an LLM?
( See Assisted-by: LLM) tag in Documentation/process/coding-assistants.rst
pw-bot: cr