[PATCH iwl-net v2 1/2] ice: detach the VF representor when ice_start_vfs() fails
From: Linkui Xiao
Date: Mon Sep 28 2026 - 02:57:26 EST
From: Linkui Xiao <xiaolinkui@xxxxxxxxxx>
ice_start_vfs() attaches every VF it brings up to the eswitch with
ice_eswitch_attach_vf(), but the teardown path only undoes the queue
mappings and the VF VSI. Nothing calls ice_eswitch_detach_vf() for the
VFs that were attached before the failure, and the caller, ice_ena_vfs(),
goes straight to ice_free_vf_entries(), which drops its reference to
every VF.
The port representors created for those VFs therefore outlive the
failed VF creation:
- the representor netdev stays registered and its devlink port stays
registered too. That port is embedded in struct ice_vf, so it ends
up pointing into the memory that ice_sriov_free_vf() releases;
- repr->vf keeps pointing at the freed struct ice_vf, and repr->src_vsi
at the VF VSI that ice_vf_vsi_release() tore down. The leftover
netdev is still visible to the user, so even a plain
"ip -s link show" of it reaches ice_repr_get_stats64(), which calls
repr->ops.ready() -> ice_check_vf_ready_for_cfg(repr->vf) and then
reads repr->src_vsi through ice_update_eth_stats();
- the virtchnl ops of that VF, which ice_repr_add_vf() replaced with
ice_virtchnl_set_repr_ops(), are never handed back to
ice_virtchnl_set_dflt_ops();
- pf->eswitch.reprs never becomes empty, so ice_eswitch_detach() never
calls ice_eswitch_disable_switchdev(). pf->eswitch.is_running stays
true, with the bridge offloads and the devlink rate topology still
up, and ice_eswitch_release_env() is skipped, so the uplink VSI is
left in the switchdev configuration that ice_eswitch_setup_env()
gave it.
Detach the representor in the teardown loop the way ice_free_vfs() does,
ahead of ice_vf_vsi_release(), because ice_repr_rem_vf() and
ice_eswitch_release_repr() both need repr->src_vsi to still be valid.
Every VF the teardown loop walks completed ice_eswitch_attach_vf()
successfully, and ice_eswitch_detach_vf() already returns early for a VF
without a representor, so no extra condition is needed.
Hold vf->cfg_lock across the teardown of each VF as well, like
ice_free_vfs() does. The VFs unwound here are the ones that already
reached set_bit(ICE_VF_STATE_INIT), which is exactly what
ice_check_vf_ready_for_cfg() checks, so a host administrator can still
run "ip link set dev <pf> vf N ..." and a VF can still send a mailbox
message while the loop walks them. Both paths take cfg_lock and then run
ice_reset_vf(), which gets to ice_eswitch_update_repr() and writes
through the representor that is being freed, or reach
ice_vc_process_vf_msg() reading vf->virtchnl_ops while
ice_virtchnl_set_dflt_ops() hands them back.
Fixes: fff292b47ac1 ("ice: add VF representors one by one")
Cc: stable@xxxxxxxxxxxxxxx
Signed-off-by: Linkui Xiao <xiaolinkui@xxxxxxxxxx>
---
v1:
- Link: https://lore.kernel.org/netdev/20260921031616.3390259-1-xiaolinkui@xxxxxxx/
Changes in v2:
- Hold vf->cfg_lock across the teardown of each VF, the way ice_free_vfs()
does, so that a concurrent VF reconfiguration cannot walk through the
representor and the virtchnl ops that are being torn down.
(Sashiko AI review)
- Patch 2/2 is new and returns the VF MSI-X window that the same failure path
reserves. That is a separate, pre-existing bug, so it is not folded in here.
(Sashiko AI review)
- Not carrying over the Reviewed-by from Aleksandr Loktionov, as the code
changed after his review.
drivers/net/ethernet/intel/ice/ice_sriov.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index e04de0215596..95abc6704820 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -508,8 +508,13 @@ static int ice_start_vfs(struct ice_pf *pf)
if (it_cnt == 0)
break;
+ mutex_lock(&vf->cfg_lock);
+
+ ice_eswitch_detach_vf(pf, vf);
ice_dis_vf_mappings(vf);
ice_vf_vsi_release(vf);
+ mutex_unlock(&vf->cfg_lock);
+
it_cnt--;
}
--
2.25.1