Re: [PATCH net] ice: detach the VF representor when ice_start_vfs() fails
From: Linkui Xiao
Date: Mon Sep 28 2026 - 02:47:19 EST
> [High] The newly added ice_eswitch_detach_vf(pf, vf) in the ice_start_vfs()
> teardown loop -- Should this call be made with vf->cfg_lock held?
Agreed, and thanks for the exact interleaving. Both existing callers,
ice_free_vfs() and ice_reset_all_vfs(), hold vf->cfg_lock across the detach,
and the VFs this loop unwinds are precisely the ones that already reached
set_bit(ICE_VF_STATE_INIT) -- which is what ice_check_vf_ready_for_cfg()
checks -- so nothing keeps a concurrent "ip link set dev <pf> vf N ..." or
a VF mailbox message out of them any more.
v2 holds vf->cfg_lock across the whole teardown of each VF, the way
ice_free_vfs() does, which also covers ice_dis_vf_mappings() and
ice_vf_vsi_release(). Both writers in your interleaving are excluded by
that: ice_reset_vf() (reached from __ice_set_vf_mac(), ice_set_vf_trust(),
ice_set_vf_port_vlan() and the other ndo handlers) and
ice_vc_process_vf_msg() take the same lock. Attach stays lockless on
purpose: on the way up ICE_VF_STATE_INIT is not set yet, so
ice_check_vf_init() keeps the configuration paths out.
> [High] AB-BA lock ordering inversion between pf->vfs.table_lock and the
> devlink instance lock.
I do not think this one is added by the patch, and the fix for it is much
larger than this UAF fix.
The pair table_lock -> devl_lock is already taken in this very loop before
the failure: whenever the teardown loop has anything to unwind (it_cnt != 0),
at least one ice_eswitch_attach_vf() has already run in the forward loop
above, under the same table_lock. ice_eswitch_detach_vf() in ice_free_vfs()
and ice_reset_all_vfs() does the same. The call added here goes through the
same helper, inside the region that is already under table_lock, so it does
not establish a new ordering pair.
Converging on one hierarchy, as you suggest, means taking devl_lock outside
table_lock in ice_ena_vfs()/ice_free_vfs()/ice_reset_all_vfs() and adding
devl_lock-held variants of ice_eswitch_attach_vf()/ice_eswitch_detach_vf().
That is a lock-hierarchy change across three callers plus the eswitch API,
and it does not belong in a -net patch that has to be backported to stable.
I would rather propose it as a separate series if you think it is worth
doing.
> [Medium] ... does this loop also need to return the VF MSI-X window?
Agreed, this one is real. ice_virt_get_irqs() does bitmap_set() on the
PF-wide pf->virt_irq_tracker.bm and nothing on this path calls
ice_virt_free_irqs(): not the release_vsi label of ice_init_vf_vsi_res(),
not the ice_eswitch_attach_vf() failure branch, and not the teardown loop.
The caller does not help either. Since the tracker lives for the lifetime of
the PF and ice_set_per_vf_res() sizes the VFs from
pf->virt_irq_tracker.num_entries rather than from the free area, every
failed "echo N > sriov_numvfs" leaks a little more of the range.
As you say, it is pre-existing, and its root cause is different from the
representor bug, so v2 sends it as patch 2/2 instead of folding it into
patch 1/2, with its own Fixes: tag. It returns the window on every failure
path, in the order ice_free_vfs() uses.
v2 is two patches on the same baseline as v1:
1/2 ice: detach the VF representor when ice_start_vfs() fails
2/2 ice: release the VF MSI-X window when ice_start_vfs() fails (new)
Code of 1/2 changed to add the cfg_lock, so the Reviewed-by from Aleksandr
Loktionov is not carried over.
pw-bot: cr