Re: [PATCH net v3 1/2] net: ngbe: propagate resume errors to the PM core

From: Joe Damato

Date: Wed Sep 30 2026 - 18:12:17 EST


On Wed, Sep 30, 2026 at 05:47:47PM +0800, Zhang Yunfei wrote:
> After a failed resume from suspend, the device stays detached from
> the networking stack: every attempt to bring the interface up fails
> the netif_device_present() check in __dev_open() with -ENODEV,
> while the PM core is told that resume succeeded.
>
> ngbe_resume() declares err as u32 and unconditionally returns 0:
> the return value of ngbe_reset_hw() is ignored entirely, and
> failures of wx_init_interrupt_scheme() and ngbe_open() are silently
> swallowed. The interrupt scheme torn down at suspend is never
> rebuilt, so the device cannot self-heal.
>
> Fix the type to int and propagate the errors, making the whole tail
> of the resume path consistent with the pci_enable_device_mem()
> failure path at the top. If the hardware reset fails, return early:
> the remaining resume steps cannot succeed. The suspend path may
> already have torn the interface down (ngbe_close() and
> wx_clear_interrupt_scheme()), leaving freed rings and IRQs behind
> while netif_running() still reports true, so set WX_STATE_RES_FREED
> on every failing return, the same mark ngbe_down_suspend() uses for
> the PCI error recovery path: a later ngbe_close() skips the
> teardown of the already-freed state, and ngbe_up_complete() clears
> the bit again so later opens are unaffected. This is safe on the
> wx_init_interrupt_scheme() failure path too, as it cleans up after
> itself.
>
> Found by manual code inspection of the PM error paths; the missing
> ngbe_reset_hw() propagation was reported by Sashiko in its review
> of v1. The failure paths are unreachable without fault injection:
> a loadable test module injects failures into ngbe_reset_hw(),
> wx_init_interrupt_scheme() and ngbe_open(). In a QEMU VM the
> unfixed driver hits the kernel "Trying to free already-free IRQ"
> warning on the post-suspend close; with the fix, the error is
> reported and no warning appears. No physical ngbe device is
> involved.
>
> Fixes: 6963e463256e ("net: ngbe: add Wake on Lan support")
> Reported-by: Sashiko <sashiko-bot@xxxxxxxxxx>
> Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260917090050.1927999-1-zhangyunfei1%40kylinos.cn
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Zhang Yunfei <zhangyunfei1@xxxxxxxxxx>
> ---
> Changes in v3:
> - set WX_STATE_RES_FREED on every failing return of ngbe_resume()
> (the pci_enable_device_mem() failure, the reset failure early return
> and the wx_init_interrupt_scheme()/ngbe_open() failures), so that a
> later ngbe_close() skips re-running the teardown on the already-freed
> post-suspend state (Sashiko review of v2);
>
> Changes in v2:
> - also propagate the ngbe_reset_hw() failure, so the whole tail of
> ngbe_resume() reports errors to the PM core;
> - drop the inaccurate "device can be re-probed" claim: the PM core
> records and logs the failure, there is no re-probe.
>
> drivers/net/ethernet/wangxun/ngbe/ngbe_main.c | 14 +++++++++++---
> 1 file changed, 11 insertions(+), 3 deletions(-)

propagating the error through seems sensible, so:

Reviewed-by: Joe Damato <joe@xxxxxxx>