From: Linkui Xiao <xiaolinkui@126.com>
To: netdev-bot+sashiko@kernel.org
Cc: anthony.l.nguyen@intel.com, przemyslaw.kitszel@intel.com,
andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
kuba@kernel.org, pabeni@redhat.com,
intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org,
linux-kernel@vger.kernel.org, xiaolinkui@kylinos.cn,
stable@vger.kernel.org
Subject: Re: [PATCH net] ice: detach the VF representor when ice_start_vfs() fails
Date: Mon, 28 Sep 2026 14:44:54 +0800 [thread overview]
Message-ID: <8182d29f-d945-41f6-9c7f-e52eb7699ca6@126.com> (raw)
In-Reply-To: <179009028142.2160803.9147341738754348504@kernel.org>
> [High] The newly added ice_eswitch_detach_vf(pf, vf) in the
ice_start_vfs()
> teardown loop -- Should this call be made with vf->cfg_lock held?
Agreed, and thanks for the exact interleaving. Both existing callers,
ice_free_vfs() and ice_reset_all_vfs(), hold vf->cfg_lock across the detach,
and the VFs this loop unwinds are precisely the ones that already reached
set_bit(ICE_VF_STATE_INIT) -- which is what ice_check_vf_ready_for_cfg()
checks -- so nothing keeps a concurrent "ip link set dev <pf> vf N ..." or
a VF mailbox message out of them any more.
v2 holds vf->cfg_lock across the whole teardown of each VF, the way
ice_free_vfs() does, which also covers ice_dis_vf_mappings() and
ice_vf_vsi_release(). Both writers in your interleaving are excluded by
that: ice_reset_vf() (reached from __ice_set_vf_mac(), ice_set_vf_trust(),
ice_set_vf_port_vlan() and the other ndo handlers) and
ice_vc_process_vf_msg() take the same lock. Attach stays lockless on
purpose: on the way up ICE_VF_STATE_INIT is not set yet, so
ice_check_vf_init() keeps the configuration paths out.
> [High] AB-BA lock ordering inversion between pf->vfs.table_lock and the
> devlink instance lock.
I do not think this one is added by the patch, and the fix for it is much
larger than this UAF fix.
The pair table_lock -> devl_lock is already taken in this very loop before
the failure: whenever the teardown loop has anything to unwind (it_cnt
!= 0),
at least one ice_eswitch_attach_vf() has already run in the forward loop
above, under the same table_lock. ice_eswitch_detach_vf() in ice_free_vfs()
and ice_reset_all_vfs() does the same. The call added here goes through the
same helper, inside the region that is already under table_lock, so it does
not establish a new ordering pair.
Converging on one hierarchy, as you suggest, means taking devl_lock outside
table_lock in ice_ena_vfs()/ice_free_vfs()/ice_reset_all_vfs() and adding
devl_lock-held variants of ice_eswitch_attach_vf()/ice_eswitch_detach_vf().
That is a lock-hierarchy change across three callers plus the eswitch API,
and it does not belong in a -net patch that has to be backported to stable.
I would rather propose it as a separate series if you think it is worth
doing.
> [Medium] ... does this loop also need to return the VF MSI-X window?
Agreed, this one is real. ice_virt_get_irqs() does bitmap_set() on the
PF-wide pf->virt_irq_tracker.bm and nothing on this path calls
ice_virt_free_irqs(): not the release_vsi label of ice_init_vf_vsi_res(),
not the ice_eswitch_attach_vf() failure branch, and not the teardown loop.
The caller does not help either. Since the tracker lives for the lifetime of
the PF and ice_set_per_vf_res() sizes the VFs from
pf->virt_irq_tracker.num_entries rather than from the free area, every
failed "echo N > sriov_numvfs" leaks a little more of the range.
As you say, it is pre-existing, and its root cause is different from the
representor bug, so v2 sends it as patch 2/2 instead of folding it into
patch 1/2, with its own Fixes: tag. It returns the window on every failure
path, in the order ice_free_vfs() uses.
v2 is two patches on the same baseline as v1:
1/2 ice: detach the VF representor when ice_start_vfs() fails
2/2 ice: release the VF MSI-X window when ice_start_vfs() fails (new)
Code of 1/2 changed to add the cfg_lock, so the Reviewed-by from Aleksandr
Loktionov is not carried over.
pw-bot: cr
prev parent reply other threads:[~2026-09-28 6:46 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 3:16 Linkui Xiao
2026-09-21 15:29 ` Loktionov, Aleksandr
2026-09-22 15:18 ` netdev-bot+sashiko
2026-09-28 6:44 ` Linkui Xiao [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=8182d29f-d945-41f6-9c7f-e52eb7699ca6@126.com \
--to=xiaolinkui@126.com \
--cc=andrew+netdev@lunn.ch \
--cc=anthony.l.nguyen@intel.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=intel-wired-lan@lists.osuosl.org \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev-bot+sashiko@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=przemyslaw.kitszel@intel.com \
--cc=stable@vger.kernel.org \
--cc=xiaolinkui@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®