From: Cosmin Ratiu <cratiu@nvidia.com>
To: Tariq Toukan <tariqt@nvidia.com>,
"netdev-bot+sashiko@kernel.org" <netdev-bot+sashiko@kernel.org>
Cc: Boris Pismenny <borisp@nvidia.com>,
"andrew+netdev@lunn.ch" <andrew+netdev@lunn.ch>,
"daniel.zahka@gmail.com" <daniel.zahka@gmail.com>,
"davem@davemloft.net" <davem@davemloft.net>,
Jianbo Liu <jianbol@nvidia.com>,
"sd@queasysnail.net" <sd@queasysnail.net>,
Emeel Hakim <ehakim@nvidia.com>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"edumazet@google.com" <edumazet@google.com>,
"pabeni@redhat.com" <pabeni@redhat.com>,
"linux-rdma@vger.kernel.org" <linux-rdma@vger.kernel.org>,
Raed Salem <raeds@nvidia.com>,
"leon@kernel.org" <leon@kernel.org>,
Dragos Tatulea <dtatulea@nvidia.com>,
Carolina Jubran <cjubran@nvidia.com>, Roi Dayan <roid@nvidia.com>,
"kuba@kernel.org" <kuba@kernel.org>,
Mark Bloch <mbloch@nvidia.com>,
Lior Nahmanson <liorna@nvidia.com>,
Parav Pandit <parav@nvidia.com>,
"netdev@vger.kernel.org" <netdev@vger.kernel.org>,
Rahul Rameshbabu <rrameshbabu@nvidia.com>,
Gal Pressman <gal@nvidia.com>,
"steffen.klassert@secunet.com" <steffen.klassert@secunet.com>,
Saeed Mahameed <saeedm@nvidia.com>
Subject: Re: [PATCH net 1/7] net/mlx5e: ipsec: Block eswitch mode changes during policy creation
Date: Wed, 23 Sep 2026 07:27:29 +0000 [thread overview]
Message-ID: <7a5d133411ce1f9b4821472f0210d3d08d561344.camel@nvidia.com> (raw)
In-Reply-To: <179001654155.2160803.10642575868860131608@kernel.org>
On Mon, 2026-09-21 at 18:49 +0000, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0 (1)
>
> If you believe none of the issues are legitimate, please reply to
> each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let
> patchwork
> know by including "pw-bot: cr" as a separate line at the end of your
> reply
> (one such reply per series is enough).
>
> - [Medium] `mlx5e_xfrm_add_policy()`
> (drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec.c:1268-
> 1274)…
>
> Pre-existing issues:
> - [Low] Incomplete-fix concern: at this commit
> `mlx5e_xfrm_add_state()`
> (drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec.c:781-813)…
>
> --
>
> net/mlx5e: ipsec: Block eswitch mode changes during policy creation
>
> This moves the eswitch mode block in mlx5e_xfrm_add_policy() to the
> top of
> the function, before priv->ipsec is checked and the policy is
> validated, and
> releases it once the flow table holds its own block or after
> unwinding.
> A check_users argument is added to mlx5_eswitch_block_mode() /
> mlx5_esw_try_lock() so the new temporary block does not reject
> policies while
> TC users are present.
>
> > diff --git
> > a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec.c
> > b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec.c
> > index db260e3d1412f..841ecdc2c4d9e 100644
> > --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec.c
> > +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec.c
> > @@ -806,7 +806,7 @@ static int mlx5e_xfrm_add_state(struct
> > net_device *dev,
> > goto err_xfrm;
> > }
> >
> > - err = mlx5_eswitch_block_mode(priv->mdev);
> > + err = mlx5_eswitch_block_mode(priv->mdev, true);
> > if (err)
> > goto unblock_ipsec;
>
> [Severity: Low]
> This isn't a bug introduced by this patch, but as of this commit
> mlx5e_xfrm_add_state() still has the ordering the changelog describes
> as
> unsafe: priv->ipsec is tested, cached into ipsec and sa_entry->ipsec,
> and
> dereferenced through mlx5e_ipsec_fs_tunnel_allowed(sa_entry) before
> the
> mode block above is taken.
>
> priv = netdev_priv(dev);
> if (!priv->ipsec)
> return -EOPNOTSUPP;
>
> ipsec = priv->ipsec;
> ...
> err = mlx5_eswitch_block_mode(priv->mdev, true);
>
> The immediately following patch in the series, "net/mlx5e: ipsec:
> Block
> eswitch mode changes before accessing priv->ipsec", moves the block
> to the
> top of mlx5e_xfrm_add_state() and adds the matching unblock_mode:
> unwind, so
> the ordering is consistent again at the end of the series. Would it
> be worth
> mentioning in this changelog that the SA path is handled in the next
> patch,
> so this intermediate state is not read as an oversight?
Will do.
>
> > @@ -1266,18 +1266,28 @@ static int mlx5e_xfrm_add_policy(struct
> > xfrm_policy *x,
> > int err;
> >
> > priv = netdev_priv(netdev);
> > + /* Block esw mode changes until the policy holds its own
> > block. */
> > + err = mlx5_eswitch_block_mode(priv->mdev, false);
> > + if (err) {
> > + NL_SET_ERR_MSG_MOD(extack, "Eswitch busy, can't
> > add policy");
> > + return err;
> > + }
> > +
>
> [Severity: Medium]
> Can this make policy adds fail spuriously with -EINVAL?
>
> The new block is taken on every policy add, and mlx5_esw_try_lock()
> bails out
> before the check_users test:
>
> drivers/net/ethernet/mellanox/mlx5/core/eswitch.c:mlx5_esw_try_lock()
> {
> if (down_write_trylock(&esw->mode_lock) == 0)
> return -EINVAL;
> ...
> }
>
> So passing check_users=false does not help when the mode_lock is
> simply held
> by someone else at that moment. A concrete reader is the TC path,
> which holds
> the read side across the whole flow creation:
>
> drivers/net/ethernet/mellanox/mlx5/core/en_tc.c:mlx5e_configure_flowe
> r() {
> if (!mlx5_esw_hold(priv->mdev))
> return -EBUSY;
> ...
> err = mlx5e_tc_add_flow(priv, f, flags, dev, &flow);
> ...
> mlx5_esw_release(priv->mdev);
> }
>
> mlx5e_ipsec_block_tc_offload(), mlx5_eswitch_unblock_mode() and
> mlx5_devlink_switchdev_active_mode_change() take the write side with
> a
> blocking down_write().
>
> The errno is passed straight to userspace:
>
> net/xfrm/xfrm_device.c:xfrm_dev_policy_add() {
> err = dev->xfrmdev_ops->xdo_dev_policy_add(xp, extack);
> if (err) {
> ...
> return err;
> }
> }
>
> Before this change, a policy add that reused already-created tables
> never ran
> a trylock at all:
>
> drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c:rx_get()
> {
> if (rx->ft.refcnt)
> goto skip;
>
> err = mlx5_eswitch_block_mode(mdev, true);
> ...
> }
>
> which is the same "TC rule on a VF representor" case the changelog
> wants to
> keep working. Would an ip xfrm policy add now fail immediately, with
> no wait
> and no retry, whenever it overlaps a tc filter add on a representor?
>
> Also, the extack text says "Eswitch busy" while the returned errno
> can be
> -EINVAL, which the kernel-doc updated by this patch defines as "lock
> is
> already held" rather than a busy/retriable condition. Should the
> transient
> trylock collision be reported as -EBUSY so userspace can tell it
> apart from a
> permanent validation rejection?
Will address.
>
> > if (!priv->ipsec) {
> > NL_SET_ERR_MSG_MOD(extack, "Device doesn't support
> > IPsec packet offload");
> > - return -EOPNOTSUPP;
> > + err = -EOPNOTSUPP;
> > + goto unblock_mode;
> > }
> >
> > err = mlx5e_xfrm_validate_policy(priv->mdev, x, extack);
> > if (err)
> > - return err;
> > + goto unblock_mode;
> >
> > pol_entry = kzalloc_obj(*pol_entry);
> > - if (!pol_entry)
> > - return -ENOMEM;
> > + if (!pol_entry) {
> > + err = -ENOMEM;
> > + goto unblock_mode;
> > + }
> >
> > pol_entry->x = x;
> > pol_entry->ipsec = priv->ipsec;
> > @@ -1293,6 +1303,7 @@ static int mlx5e_xfrm_add_policy(struct
> > xfrm_policy *x,
> > goto err_fs;
> >
> > x->xdo.offload_handle = (unsigned long)pol_entry;
> > + mlx5_eswitch_unblock_mode(priv->mdev);
> > return 0;
> >
> > err_fs:
> > @@ -1300,6 +1311,8 @@ static int mlx5e_xfrm_add_policy(struct
> > xfrm_policy *x,
> > ipsec_busy:
> > kfree(pol_entry);
> > NL_SET_ERR_MSG_MOD(extack, "Device failed to offload this
> > policy");
> > +unblock_mode:
> > + mlx5_eswitch_unblock_mode(priv->mdev);
> > return err;
> > }
>
> [ ... ]
>
> > diff --git a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
> > b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
> > index fc197d1dc9dfe..0b48cc7a67343 100644
> > --- a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
> > +++ b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
> > @@ -3005,21 +3005,22 @@ void mlx5_esw_put(struct mlx5_core_dev
> > *mdev)
> > /**
> > * mlx5_esw_try_lock() - Take a write lock on esw mode lock.
> > * @esw: eswitch device.
> > + * @check_users: reject the lock if eswitch users exist.
> > *
> > * Should be called by esw mode change routine.
> > *
> > * Return:
> > - * * 0 - esw mode if successfully locked and refcount is 0.
> > - * * -EBUSY - refcount is not 0.
> > - * * -EINVAL - In the middle of switching mode or lock is already
> > held.
> > + * * >= 0 - esw mode if successfully locked.
> > + * * -EBUSY - mode change in progress or users exist with
> > check_users set.
> > + * * -EINVAL - lock is already held.
> > */
> > -int mlx5_esw_try_lock(struct mlx5_eswitch *esw)
> > +int mlx5_esw_try_lock(struct mlx5_eswitch *esw, bool check_users)
> > {
> > if (down_write_trylock(&esw->mode_lock) == 0)
> > return -EINVAL;
> >
> > if (esw->eswitch_operation_in_progress ||
> > - atomic64_read(&esw->user_count) > 0) {
> > + (check_users && atomic64_read(&esw->user_count) > 0))
> > {
> > up_write(&esw->mode_lock);
> > return -EBUSY;
> > }
>
> [ ... ]
>
> > diff --git
> > a/drivers/net/ethernet/mellanox/mlx5/core/eswitch_offloads.c
> > b/drivers/net/ethernet/mellanox/mlx5/core/eswitch_offloads.c
> > index eb74b62601682..996490e116253 100644
> > --- a/drivers/net/ethernet/mellanox/mlx5/core/eswitch_offloads.c
> > +++ b/drivers/net/ethernet/mellanox/mlx5/core/eswitch_offloads.c
> > @@ -4351,8 +4351,7 @@ int mlx5_eswitch_block_mode(struct
> > mlx5_core_dev *dev)
> > if (!mlx5_esw_allowed(esw))
> > return 0;
> >
> > - /* Take TC into account */
> > - err = mlx5_esw_try_lock(esw);
> > + err = mlx5_esw_try_lock(esw, check_users);
> > if (err < 0)
> > return err;
>
> [ ... ]
>
next prev parent reply other threads:[~2026-09-23 7:27 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 17:54 [PATCH net 0/7] net/mlx5e: Fix offload lifetime, cleanup and exclusion bugs Tariq Toukan
2026-09-17 17:54 ` [PATCH net 1/7] net/mlx5e: ipsec: Block eswitch mode changes during policy creation Tariq Toukan
2026-09-21 18:49 ` netdev-bot+sashiko
2026-09-23 7:27 ` Cosmin Ratiu [this message]
2026-09-17 17:54 ` [PATCH net 2/7] net/mlx5e: ipsec: Block eswitch mode changes before accessing priv->ipsec Tariq Toukan
2026-09-21 18:49 ` netdev-bot+sashiko
2026-09-23 7:52 ` Cosmin Ratiu
2026-09-17 17:54 ` [PATCH net 3/7] net/mlx5e: tc: Tie esw & accel blocking refs to the flow's lifetime Tariq Toukan
2026-09-21 18:49 ` netdev-bot+sashiko
2026-09-23 7:59 ` Cosmin Ratiu
2026-09-17 17:54 ` [PATCH net 4/7] net/mlx5e: macsec: Track hardware object ownership for SA teardown Tariq Toukan
2026-09-21 18:49 ` netdev-bot+sashiko
2026-09-23 8:32 ` Cosmin Ratiu
2026-09-17 17:54 ` [PATCH net 5/7] net/mlx5e: macsec: Delete remaining SecYs during cleanup Tariq Toukan
2026-09-21 18:49 ` netdev-bot+sashiko
2026-09-23 8:49 ` Cosmin Ratiu
2026-09-17 17:54 ` [PATCH net 6/7] net/mlx5e: Serialize TC and IPsec offload exclusion counters Tariq Toukan
2026-09-17 17:54 ` [PATCH net 7/7] net/mlx5e: shampo: Do not merge PSP packets Tariq Toukan
2026-09-21 18:49 ` netdev-bot+sashiko
2026-09-23 8:50 ` Cosmin Ratiu
2026-09-22 10:44 ` [PATCH net 0/7] net/mlx5e: Fix offload lifetime, cleanup and exclusion bugs Paolo Abeni
2026-09-23 10:47 ` Cosmin Ratiu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7a5d133411ce1f9b4821472f0210d3d08d561344.camel@nvidia.com \
--to=cratiu@nvidia.com \
--cc=andrew+netdev@lunn.ch \
--cc=borisp@nvidia.com \
--cc=cjubran@nvidia.com \
--cc=daniel.zahka@gmail.com \
--cc=davem@davemloft.net \
--cc=dtatulea@nvidia.com \
--cc=edumazet@google.com \
--cc=ehakim@nvidia.com \
--cc=gal@nvidia.com \
--cc=jianbol@nvidia.com \
--cc=kuba@kernel.org \
--cc=leon@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=liorna@nvidia.com \
--cc=mbloch@nvidia.com \
--cc=netdev-bot+sashiko@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=parav@nvidia.com \
--cc=raeds@nvidia.com \
--cc=roid@nvidia.com \
--cc=rrameshbabu@nvidia.com \
--cc=saeedm@nvidia.com \
--cc=sd@queasysnail.net \
--cc=steffen.klassert@secunet.com \
--cc=tariqt@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®