Re: [PATCH net-next] net: bcmgenet: complete Tx NAPI after one reclaim pass

From: Eric Dumazet

Date: Mon Sep 28 2026 - 09:26:29 EST


On Mon, Sep 28, 2026 at 2:15 PM Nicolai Buchwitz <nb@xxxxxxxxxxx> wrote:
>
> After reclaiming, the Tx poll asks to be polled again. That extra poll
> finds nothing to do and does not reduce the interrupt rate.
>
> Complete the NAPI and unmask the ring interrupt right after the reclaim.
>
> On a CM4 this raises 60 byte pktgen Tx from about 187k to 224k pps and
> saves about 5% of one core at TCP line rate.

Reviewed-by: Eric Dumazet <edumazet@xxxxxxxxxx>

As a follow-up, you can get rid of ring->lock in the TX
fastpath altogether (it looks like a leftover from before TX reclaim
was moved to NAPI in commit 4092e6acf5cb ("net: bcmgenet: use NAPI for
Tx completion")).

bcmgenet_xmit() is already serialized by the netdev queue lock, and
bcmgenet_tx_poll() by NAPI. To make them lockless with respect to each
other:

1. Drop ring->free_bds and compute available space from
(READ_ONCE(ring->prod_index) - READ_ONCE(ring->c_index)) & DMA_C_INDEX_MASK.
2. Use the lockless queue stop/wake helpers from <net/pkt_sched.h>
(netif_txq_maybe_stop() and __netif_txq_completed_wake()).
3. Ensure proper memory barriers between populating tx_cb_ptr /
writing TDMA_PROD_INDEX in bcmgenet_xmit() and reading
TDMA_CONS_INDEX / freeing tx_cbs in __bcmgenet_tx_reclaim() (since
bcmgenet uses readl_relaxed/writel_relaxed).
4. Avoid sharing ring->stats64.syncp between bcmgenet_add_tsb()
(tx_dropped) and __bcmgenet_tx_reclaim(), and serialize
bcmgenet_timeout() with napi_disable() + __netif_tx_lock_bh().

Thanks.