Re: [PATCH iwl-next 6/8] e1000e: disable NAPI while interface is down
From: Matt Vollrath
Date: Thu Sep 03 2026 - 13:20:15 EST
On 9/3/26 06:27, Simon Horman wrote:
This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
Full review at: https://sashiko.dev/#/patchset/20260830232146.36948-2-tactii@xxxxxxxxx?part=6
---
e1000e: disable NAPI while interface is down
This commit moves napi_disable into e1000e_down and adds napi_enable to
e1000e_up to ensure NAPI is disabled during reconfiguration windows.
diff --git a/drivers/net/ethernet/intel/e1000e/netdev.c b/drivers/net/ethernet/intel/e1000e/netdev.c[ ... ]
index 194166531bc84..6d846cacb026e 100644
--- a/drivers/net/ethernet/intel/e1000e/netdev.c
+++ b/drivers/net/ethernet/intel/e1000e/netdev.c
@@ -3601,7 +3603,7 @@ void e1000e_down(struct e1000_adapter *adapter, bool reset)
e1000_irq_disable(adapter);
- napi_synchronize(&adapter->napi);
+ napi_disable(&adapter->napi);
[Severity: High]
Does this change introduce an infinite hang in napi_disable() due to
concurrent or duplicate e1000e_down() calls?
A runtime PM suspend callback (e1000e_pm_runtime_suspend) can execute
concurrently with an ethtool operation like e1000_set_ringparam. The ethtool
operation sets the __E1000_RESETTING bit and holds it while allocating
resources.
No, ethtool and runtime PM suspend are synchronized:
* ethtool core calls pm_runtime_get_sync() before any operation. This
increments dev->power.usage_count and waits for any ongoing suspend to
finish.
* PM core rpm_suspend exits without retry (rpm_check_suspend_allowed()).
The PM suspend callback waits on this bit but uses a bounded timeout (300ms).
If the ethtool operation takes longer than 300ms, the PM callback times out,
prints a WARN_ON, and unconditionally calls e1000e_down():
e1000e_pm_runtime_suspend() {
while (test_bit(__E1000_RESETTING, &adapter->state) && count--)
usleep_range(10000, 11000);
WARN_ON(test_bit(__E1000_RESETTING, &adapter->state));
/* Down the device without resetting the hardware */
e1000e_down(adapter, false);
}
Because the ethtool operation also called (or is calling) e1000e_down(),
napi_disable() is executed either concurrently or sequentially on an
already-disabled NAPI struct, leading to an infinite hang waiting for
NAPI_STATE_SCHED.
[Severity: High]
Could this also lead to a sequential double napi_disable() hang due to state
mismatches where e1000e_down() is called but netif_running() remains true?
The e1000e_pm_thaw() function attempts to bring the device back up after a
freeze/suspend. If e1000_request_irq() fails, it exits via the err_irq label
without calling e1000e_up(), leaving NAPI disabled:
e1000e_pm_thaw() {
rc = e1000_request_irq(adapter);
if (rc)
goto err_irq;
e1000e_up(adapter);
}
However, the device's IFF_UP flag is untouched, so netif_running() remains
true. Subsequently, if an administrator changes the MTU via
e1000_change_mtu(), the driver checks netif_running() and blindly calls
e1000e_down():
e1000_change_mtu() {
if (netif_running(netdev))
e1000e_down(adapter, true);
}
Core does not call ndo_change_mtu on detached devices, it checks
netif_device_present() in the typical path. Bonding and team paths will
close the device before changing MTU.
Same netif_device_present() check in ethtool core and e1000e_pm_freeze().
However, the PM runtime suspend and resume ops only check IFF_UP and not
netif_device_present(). PM core does not check this or gate it on a known
thaw failure (by design). That is a real pre-existing bug made more
consequential by this change.
This invokes napi_disable() a second time sequentially, which hangs
indefinitely because the NAPI instance was never re-enabled.
timer_delete_sync(&adapter->watchdog_timer);
timer_delete_sync(&adapter->phy_info_timer);