mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Linkui Xiao <xiaolinkui@126.com>
To: anthony.l.nguyen@intel.com, przemyslaw.kitszel@intel.com,
	andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
	kuba@kernel.org, pabeni@redhat.com
Cc: intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org,
	linux-kernel@vger.kernel.org, Linkui Xiao <xiaolinkui@kylinos.cn>,
	stable@vger.kernel.org
Subject: [PATCH iwl-net v3 1/3] ice: attach and detach VF representors outside of vf->cfg_lock
Date: Thu,  8 Oct 2026 20:57:52 +0800	[thread overview]
Message-ID: <20261008125754.3520773-2-xiaolinkui@126.com> (raw)
In-Reply-To: <20261008125754.3520773-1-xiaolinkui@126.com>

From: Linkui Xiao <xiaolinkui@kylinos.cn>

ice_free_vfs() and ice_reset_all_vfs() call ice_eswitch_detach_vf()
and ice_eswitch_attach_vf() with vf->cfg_lock held. Both take the
devlink instance lock and then register or unregister the representor
netdev, which takes RTNL. That gives

	cfg_lock -> devlink instance lock -> RTNL

while three other paths take cfg_lock with RTNL already held:

  - __ice_set_vf_mac(), from ndo_set_vf_mac()
  - ice_set_vf_port_vlan(), from ndo_set_vf_vlan()
  - ice_repr_ethtool_reset(), via ice_reset_vf() with
    ICE_VF_RESET_LOCK

which is the opposite order, RTNL -> cfg_lock. An "ip link set ... vf
N mac" on one CPU against "echo 0 > sriov_numvfs" or a PF reset on
another can therefore deadlock.

Move the detach and the attach outside of cfg_lock. Representors are
created and removed under pf->vfs.table_lock only, which both callers
already hold, and ice_start_vfs() attaches without cfg_lock too.
Nothing that the detach reads is torn down by moving it: repr->src_vsi
still points at the VF VSI, which is only released later, under
cfg_lock.

Both callers raise ICE_VF_DIS in pf->state before entering the loop,
so a reset that starts after that leaves through ice_is_vf_disabled()
before it reaches ice_eswitch_update_repr(), and
ice_check_vf_ready_for_cfg() keeps the ndo handlers and the ethtool
reset away as well. ICE_VF_DIS does not cover a reset that is already
inside cfg_lock when the loop gets there, so take cfg_lock once before
the detach to wait that one out while its representor is still
attached. Without the wait it could xa_load() the representor that
ice_repr_destroy() is freeing, and that is free_netdev() + kfree(),
with no RCU grace period in between.

In ice_reset_all_vfs() the attach moves after mutex_unlock(), so the
"VSI rebuild failed" path still leaves the VF detached, exactly as it
did before.

Found by code inspection of the VF setup and teardown paths; the same
inversion has also been hit out of tree. It was not triggered here
and no stack trace was captured. Compile-tested only, not run on
hardware.

Fixes: fff292b47ac1 ("ice: add VF representors one by one")
Fixes: c9663f79cd82 ("ice: adjust switchdev rebuild path")
Cc: stable@vger.kernel.org
Suggested-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
---
Changes in v3:
- New patch. Move the ice_eswitch_detach_vf() and ice_eswitch_attach_vf() calls
  in ice_free_vfs() and ice_reset_all_vfs() out of vf->cfg_lock, which is where
  the inversion is: both take the devlink instance lock and then RTNL, while
  ndo_set_vf_mac(), ndo_set_vf_vlan() and the representor ethtool reset take
  vf->cfg_lock with RTNL already held.
  (Przemek Kitszel)
- Take cfg_lock once before the detach as well, so that a reset that is already
  inside the lock is waited out instead of being left with a window between the
  detach and the lock acquisition that follows it.
- Include how the issue was found, that it has not been triggered, and that the
  change is compile tested only, as netdev-bot asked for.
 drivers/net/ethernet/intel/ice/ice_sriov.c  | 12 ++++++++++++
 drivers/net/ethernet/intel/ice/ice_vf_lib.c | 16 ++++++++++++++--
 2 files changed, 26 insertions(+), 2 deletions(-)

diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index e04de0215596..471c1e29a865 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -154,9 +154,21 @@ void ice_free_vfs(struct ice_pf *pf)
 	mutex_lock(&vfs->table_lock);
 
 	ice_for_each_vf(pf, bkt, vf) {
+		/* Detach the representor before cfg_lock: it takes the
+		 * devlink instance lock and then RTNL, while
+		 * __ice_set_vf_mac(), ice_set_vf_port_vlan() and the
+		 * representor ethtool reset take cfg_lock under RTNL.
+		 * Take cfg_lock once ahead of it to wait out a reset that
+		 * is already inside the lock; ICE_VF_DIS, raised above,
+		 * keeps later ones out.
+		 */
 		mutex_lock(&vf->cfg_lock);
+		mutex_unlock(&vf->cfg_lock);
 
 		ice_eswitch_detach_vf(pf, vf);
+
+		mutex_lock(&vf->cfg_lock);
+
 		ice_dis_vf_qs(vf);
 		ice_virt_free_irqs(pf, vf->first_vector_idx, vf->num_msix);
 
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.c b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
index a54cb2b8d3c7..b91eec0adf51 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
@@ -789,9 +789,21 @@ void ice_reset_all_vfs(struct ice_pf *pf)
 
 	/* free VF resources to begin resetting the VSI state */
 	ice_for_each_vf(pf, bkt, vf) {
+		/* Detach the representor before cfg_lock and attach it after
+		 * releasing it again: both take the devlink instance lock and
+		 * then RTNL, while __ice_set_vf_mac(), ice_set_vf_port_vlan()
+		 * and the representor ethtool reset take cfg_lock under RTNL.
+		 * Take cfg_lock once ahead of the detach to wait out a reset
+		 * that is already inside the lock; ICE_VF_DIS, raised above,
+		 * keeps later ones out.
+		 */
 		mutex_lock(&vf->cfg_lock);
+		mutex_unlock(&vf->cfg_lock);
 
 		ice_eswitch_detach_vf(pf, vf);
+
+		mutex_lock(&vf->cfg_lock);
+
 		vf->driver_caps = 0;
 		ice_vc_set_default_allowlist(vf);
 
@@ -812,10 +824,10 @@ void ice_reset_all_vfs(struct ice_pf *pf)
 		}
 		ice_vf_post_vsi_rebuild(vf);
 
+		mutex_unlock(&vf->cfg_lock);
+
 		if (ice_is_eswitch_mode_switchdev(pf))
 			ice_eswitch_attach_vf(pf, vf);
-
-		mutex_unlock(&vf->cfg_lock);
 	}
 
 	ice_flush(hw);
-- 
2.25.1


  reply	other threads:[~2026-10-08 12:58 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08 12:57 [PATCH iwl-net v3 0/3] ice: fix VF representor lock ordering and teardown error paths Linkui Xiao
2026-10-08 12:57 ` Linkui Xiao [this message]
2026-10-08 12:57 ` [PATCH iwl-net v3 2/3] ice: detach the VF representor when ice_start_vfs() fails Linkui Xiao
2026-10-08 12:57 ` [PATCH iwl-net v3 3/3] ice: free the VF MSI-X vectors when VF start fails Linkui Xiao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261008125754.3520773-2-xiaolinkui@126.com \
    --to=xiaolinkui@126.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=anthony.l.nguyen@intel.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=intel-wired-lan@lists.osuosl.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=przemyslaw.kitszel@intel.com \
    --cc=stable@vger.kernel.org \
    --cc=xiaolinkui@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®