Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
From: anmory
Date: Fri Oct 09 2026 - 00:25:11 EST
Hi Joe, Eric, Michael, Pavan,
quick update from my testing.
I backported Eric's patch bnxt_en: fix DMA mapping length for padded small packets onto the exact Debian 6.12.111-1 source and built a Debian test kernel from it.
The affected host has now been running the patched kernel under the same normal workload and network configuration that reproduced the problem with stock 6.12.111-1.
The relevant setup is:
Broadcom BCM57504 / bnxt_en
affected interface is an 802.1Q trunk with multiple VLAN subinterfaces
Keepalived VRRP with use_vmac
forwarding/conntrack/SNAT traffic through the host
25 Gbit/s link
So far the patched kernel has been stable. It ran for the entire day yesterday and is still running without reproducing the failure.
With stock 6.12.111-1 we had repeatedly seen the following sequence:
AMD-Vi IO_PAGE_FAULT
NETDEV WATCHDOG / TX timeout
hwrm_ring_free / hwrm_ring_alloc failures
bnxt_init_nic failure
On the patched kernel, the capture has so far only recorded the normal bnxt_en initialization and link-up messages. No IOMMU fault, TX timeout, or HWRM recovery failure has occurred.
For comparison:
6.12.107-1 stable
6.12.111-1 stock reproduces the issue
6.12.111-1 + Eric's patch stable so far
On this particular host, stock 6.12.111-1 reproduced once after roughly 20 minutes, although we have also seen occurrences take several hours. Given that the patched kernel has now survived the full day yesterday under normal traffic, this is looking increasingly consistent with Eric's DMA padding fix addressing the issue.
I will keep the patched kernel running and continue monitoring it. If the problem does reproduce, I also have the kernel capture running and will collect the NIC coredump as requested.
Thanks, R.
-------- Original Message --------
On Wednesday, 10/07/26 at 00:37 Joe Damato <joe@xxxxxxx> wrote:
On Tue, Oct 06, 2026 at 10:22:51PM +0000, anmory wrote:
>
> Hi Joe,
> one important detail I omitted from my previous mail: the affected eth1 interface is an 802.1Q trunk with multiple VLAN subinterfaces.
> We also use Keepalived VRRP with use_vmac, so virtual MAC interfaces are layered on the VLAN interfaces. Traffic is routed/forwarded through the host and uses conntrack/SNAT before leaving through another physical interface.
> We are not deliberately toggling VLAN offload during the test. The reproduction so far is simply to boot Debian 6.12.111-1 with this normal network configuration and workload. On one host it reproduced after about 20 minutes; on another it reproduced almost immediately. We have also seen a previous occurrence after several hours, so we do not yet have a deterministic packet-level trigger.
> Given the VLAN configuration, Eric's patch looks especially relevant. I'll test 6.12.111 with that patch first as you suggested and report whether the failure still reproduces.
Seems likely you are hitting the bug Eric just fixed, IMHO. If his patch
resolves your issue, please let me know.