From: Linkui Xiao <xiaolinkui@126.com>
To: anthony.l.nguyen@intel.com, przemyslaw.kitszel@intel.com,
andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
kuba@kernel.org, pabeni@redhat.com
Cc: intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org,
linux-kernel@vger.kernel.org, Linkui Xiao <xiaolinkui@kylinos.cn>,
stable@vger.kernel.org
Subject: [PATCH iwl-net v3 1/3] ice: attach and detach VF representors outside of vf->cfg_lock
Date: Thu, 8 Oct 2026 20:57:52 +0800 [thread overview]
Message-ID: <20261008125754.3520773-2-xiaolinkui@126.com> (raw)
In-Reply-To: <20261008125754.3520773-1-xiaolinkui@126.com>
From: Linkui Xiao <xiaolinkui@kylinos.cn>
ice_free_vfs() and ice_reset_all_vfs() call ice_eswitch_detach_vf()
and ice_eswitch_attach_vf() with vf->cfg_lock held. Both take the
devlink instance lock and then register or unregister the representor
netdev, which takes RTNL. That gives
cfg_lock -> devlink instance lock -> RTNL
while three other paths take cfg_lock with RTNL already held:
- __ice_set_vf_mac(), from ndo_set_vf_mac()
- ice_set_vf_port_vlan(), from ndo_set_vf_vlan()
- ice_repr_ethtool_reset(), via ice_reset_vf() with
ICE_VF_RESET_LOCK
which is the opposite order, RTNL -> cfg_lock. An "ip link set ... vf
N mac" on one CPU against "echo 0 > sriov_numvfs" or a PF reset on
another can therefore deadlock.
Move the detach and the attach outside of cfg_lock. Representors are
created and removed under pf->vfs.table_lock only, which both callers
already hold, and ice_start_vfs() attaches without cfg_lock too.
Nothing that the detach reads is torn down by moving it: repr->src_vsi
still points at the VF VSI, which is only released later, under
cfg_lock.
Both callers raise ICE_VF_DIS in pf->state before entering the loop,
so a reset that starts after that leaves through ice_is_vf_disabled()
before it reaches ice_eswitch_update_repr(), and
ice_check_vf_ready_for_cfg() keeps the ndo handlers and the ethtool
reset away as well. ICE_VF_DIS does not cover a reset that is already
inside cfg_lock when the loop gets there, so take cfg_lock once before
the detach to wait that one out while its representor is still
attached. Without the wait it could xa_load() the representor that
ice_repr_destroy() is freeing, and that is free_netdev() + kfree(),
with no RCU grace period in between.
In ice_reset_all_vfs() the attach moves after mutex_unlock(), so the
"VSI rebuild failed" path still leaves the VF detached, exactly as it
did before.
Found by code inspection of the VF setup and teardown paths; the same
inversion has also been hit out of tree. It was not triggered here
and no stack trace was captured. Compile-tested only, not run on
hardware.
Fixes: fff292b47ac1 ("ice: add VF representors one by one")
Fixes: c9663f79cd82 ("ice: adjust switchdev rebuild path")
Cc: stable@vger.kernel.org
Suggested-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
---
Changes in v3:
- New patch. Move the ice_eswitch_detach_vf() and ice_eswitch_attach_vf() calls
in ice_free_vfs() and ice_reset_all_vfs() out of vf->cfg_lock, which is where
the inversion is: both take the devlink instance lock and then RTNL, while
ndo_set_vf_mac(), ndo_set_vf_vlan() and the representor ethtool reset take
vf->cfg_lock with RTNL already held.
(Przemek Kitszel)
- Take cfg_lock once before the detach as well, so that a reset that is already
inside the lock is waited out instead of being left with a window between the
detach and the lock acquisition that follows it.
- Include how the issue was found, that it has not been triggered, and that the
change is compile tested only, as netdev-bot asked for.
drivers/net/ethernet/intel/ice/ice_sriov.c | 12 ++++++++++++
drivers/net/ethernet/intel/ice/ice_vf_lib.c | 16 ++++++++++++++--
2 files changed, 26 insertions(+), 2 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index e04de0215596..471c1e29a865 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -154,9 +154,21 @@ void ice_free_vfs(struct ice_pf *pf)
mutex_lock(&vfs->table_lock);
ice_for_each_vf(pf, bkt, vf) {
+ /* Detach the representor before cfg_lock: it takes the
+ * devlink instance lock and then RTNL, while
+ * __ice_set_vf_mac(), ice_set_vf_port_vlan() and the
+ * representor ethtool reset take cfg_lock under RTNL.
+ * Take cfg_lock once ahead of it to wait out a reset that
+ * is already inside the lock; ICE_VF_DIS, raised above,
+ * keeps later ones out.
+ */
mutex_lock(&vf->cfg_lock);
+ mutex_unlock(&vf->cfg_lock);
ice_eswitch_detach_vf(pf, vf);
+
+ mutex_lock(&vf->cfg_lock);
+
ice_dis_vf_qs(vf);
ice_virt_free_irqs(pf, vf->first_vector_idx, vf->num_msix);
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.c b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
index a54cb2b8d3c7..b91eec0adf51 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
@@ -789,9 +789,21 @@ void ice_reset_all_vfs(struct ice_pf *pf)
/* free VF resources to begin resetting the VSI state */
ice_for_each_vf(pf, bkt, vf) {
+ /* Detach the representor before cfg_lock and attach it after
+ * releasing it again: both take the devlink instance lock and
+ * then RTNL, while __ice_set_vf_mac(), ice_set_vf_port_vlan()
+ * and the representor ethtool reset take cfg_lock under RTNL.
+ * Take cfg_lock once ahead of the detach to wait out a reset
+ * that is already inside the lock; ICE_VF_DIS, raised above,
+ * keeps later ones out.
+ */
mutex_lock(&vf->cfg_lock);
+ mutex_unlock(&vf->cfg_lock);
ice_eswitch_detach_vf(pf, vf);
+
+ mutex_lock(&vf->cfg_lock);
+
vf->driver_caps = 0;
ice_vc_set_default_allowlist(vf);
@@ -812,10 +824,10 @@ void ice_reset_all_vfs(struct ice_pf *pf)
}
ice_vf_post_vsi_rebuild(vf);
+ mutex_unlock(&vf->cfg_lock);
+
if (ice_is_eswitch_mode_switchdev(pf))
ice_eswitch_attach_vf(pf, vf);
-
- mutex_unlock(&vf->cfg_lock);
}
ice_flush(hw);
--
2.25.1
next prev parent reply other threads:[~2026-10-08 12:58 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 12:57 [PATCH iwl-net v3 0/3] ice: fix VF representor lock ordering and teardown error paths Linkui Xiao
2026-10-08 12:57 ` Linkui Xiao [this message]
2026-10-08 12:57 ` [PATCH iwl-net v3 2/3] ice: detach the VF representor when ice_start_vfs() fails Linkui Xiao
2026-10-08 12:57 ` [PATCH iwl-net v3 3/3] ice: free the VF MSI-X vectors when VF start fails Linkui Xiao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261008125754.3520773-2-xiaolinkui@126.com \
--to=xiaolinkui@126.com \
--cc=andrew+netdev@lunn.ch \
--cc=anthony.l.nguyen@intel.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=intel-wired-lan@lists.osuosl.org \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=przemyslaw.kitszel@intel.com \
--cc=stable@vger.kernel.org \
--cc=xiaolinkui@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®