Re: [PATCH v2] net: macb: add TX stall timeout callback to recover from lost TSTART write

From: Théo Lebrun

Date: Fri Sep 18 2026 - 11:50:16 EST


Hello all,

On Tue Jun 16, 2026 at 3:23 PM CEST, Andrea della Porta wrote:
> From: Lukasz Raczylo <lukasz@xxxxxxxxxxx>
>
> The MACB found in the Raspberry Pi RP1 suffers from sporadic stalls on
> the TX queue.
> While the exact root cause is not yet fully understood, it is likely
> related to a hardware issue where a TSTART write to the NCR register
> is missed, preventing the transmission from being kicked off.
>
> Implement a timeout callback to handle TX queue stalls, triggering the
> existing restart mechanism to recover.

Any news on this topic?

There was a guess that a "flush PCIe posted write after TSTART
doorbell" patch [0] could solve it but I was sceptical [1]. We never
landed that and only took the Tx timeout implementation. If it ever
triggers then the kernel log gets a "NETDEV WATCHDOG: ..." critical
line appended [2].

I got reminded because I came across on this wiki page [3] about the
issue, whose author is in Cc. As an aside, I landed there by testing out
the Marginalia search engine and queried "macb driver" (as one does).

Main discovery of the day: it happens on EyeQ5 & latest net/main
(46bc52d13594) as well. Thanks to Andrea for the quick reproducer.

# udhcpc -i eth1
...
# iperf3 -c $IP -P10 -t3000
...
^C
# dmesg | grep eth1
[ 2.182191] macb 2b00000.ethernet eth1: Cadence GEM rev 0x00070200 at 0x02b00000 irq 35 (00:28:f8:94:24:69)
[ 23.570335] macb 2b00000.ethernet eth1: PHY [2b00000.ethernet-ffffffff:0e] driver [Marvell 88E1510] (irq=POLL)
[ 23.570981] macb 2b00000.ethernet eth1: configuring for phy/rgmii-id link mode
[ 27.783260] macb 2b00000.ethernet eth1: Link is Up - 1Gbps/Full - flow control tx
[ 42.182370] macb 2b00000.ethernet eth1: NETDEV WATCHDOG: CPU: 0: transmit queue 0 timed out 5088 ms
[ 51.158331] macb 2b00000.ethernet eth1: NETDEV WATCHDOG: CPU: 0: transmit queue 0 timed out 5004 ms
[ 56.262381] macb 2b00000.ethernet eth1: NETDEV WATCHDOG: CPU: 0: transmit queue 0 timed out 10108 ms
[ 61.126317] macb 2b00000.ethernet eth1: NETDEV WATCHDOG: CPU: 3: transmit queue 0 timed out 14972 ms
[ 67.206322] macb 2b00000.ethernet eth1: NETDEV WATCHDOG: CPU: 0: transmit queue 0 timed out 5004 ms

Note: the initial reproducer was `iperf3 -c $IP -P10 -t3000 -w4M`.
I can reproduce without -w4M but -P10 is required even though we reach
line rate with a single stream. This points to a race/mb issue in
macb_start_xmit?

I'm posting to see if anyone has theories. I'll be posting some race
fixes soon but they don't fix it (not surprising as they are unrelated).
I have many ideas but I need more time for testing; it'll be much easier
now with a reproduction setup.

[0]: https://lore.kernel.org/netdev/20260514215459.36109-2-lukasz@xxxxxxxxxxx/
[1]: https://lore.kernel.org/netdev/DIK002QFFNBY.31C3KUX2SQC6W@xxxxxxxxxxx/
[2]: https://elixir.bootlin.com/linux/v7.2.5/source/net/sched/sch_generic.c#L563-L569
[3]: https://dtype.org/wiki/Cm5_macb_network_hang

Thanks,

--
Théo Lebrun, Bootlin
Embedded Linux and Kernel engineering
https://bootlin.com