* [PATCH 6.6.y] net/mlx5e: Fix netif state handling
@ 2026-09-30 1:21 Artem Dinaburg
2026-09-30 1:23 ` netdev-bot+sinfo
0 siblings, 1 reply; 2+ messages in thread
From: Artem Dinaburg @ 2026-09-30 1:21 UTC (permalink / raw)
To: stable
Cc: Artem Dinaburg, Greg Kroah-Hartman, Sasha Levin, Shay Drory,
Tariq Toukan, Simon Horman, Jakub Kicinski, Saeed Mahameed,
Leon Romanovsky, David S. Miller, Eric Dumazet, Paolo Abeni,
netdev, linux-rdma, linux-kernel, Mark Bloch, Leon Romanovsky,
Andrew Lunn, erezsh
From: Shay Drory <shayd@nvidia.com>
[ Upstream commit 3d5918477f94e4c2f064567875c475468e264644 ]
mlx5e_suspend cleans resources only if netif_device_present() returns
true. However, mlx5e_resume changes the state of netif, via
mlx5e_nic_enable, only if reg_state == NETREG_REGISTERED.
In the below case, the above leads to NULL-ptr Oops[1] and memory
leaks:
mlx5e_probe
_mlx5e_resume
mlx5e_attach_netdev
mlx5e_nic_enable <-- netdev not reg, not calling netif_device_attach()
register_netdev <-- failed for some reason.
ERROR_FLOW:
_mlx5e_suspend <-- netif_device_present return false, resources aren't freed :(
Hence, clean resources in this case as well.
[1]
BUG: kernel NULL pointer dereference, address: 0000000000000000
PGD 0 P4D 0
Oops: 0010 [#1] SMP
CPU: 2 PID: 9345 Comm: test-ovs-ct-gen Not tainted 6.5.0_for_upstream_min_debug_2023_09_05_16_01 #1
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.13.0-0-gf21b5a4aeb02-prebuilt.qemu.org 04/01/2014
RIP: 0010:0x0
Code: Unable to access opcode bytes at0xffffffffffffffd6.
RSP: 0018:ffff888178aaf758 EFLAGS: 00010246
Call Trace:
<TASK>
? __die+0x20/0x60
? page_fault_oops+0x14c/0x3c0
? exc_page_fault+0x75/0x140
? asm_exc_page_fault+0x22/0x30
notifier_call_chain+0x35/0xb0
blocking_notifier_call_chain+0x3d/0x60
mlx5_blocking_notifier_call_chain+0x22/0x30 [mlx5_core]
mlx5_core_uplink_netdev_event_replay+0x3e/0x60 [mlx5_core]
mlx5_mdev_netdev_track+0x53/0x60 [mlx5_ib]
mlx5_ib_roce_init+0xc3/0x340 [mlx5_ib]
__mlx5_ib_add+0x34/0xd0 [mlx5_ib]
mlx5r_probe+0xe1/0x210 [mlx5_ib]
? auxiliary_match_id+0x6a/0x90
auxiliary_bus_probe+0x38/0x80
? driver_sysfs_add+0x51/0x80
really_probe+0xc9/0x3e0
? driver_probe_device+0x90/0x90
__driver_probe_device+0x80/0x160
driver_probe_device+0x1e/0x90
__device_attach_driver+0x7d/0x100
bus_for_each_drv+0x80/0xd0
__device_attach+0xbc/0x1f0
bus_probe_device+0x86/0xa0
device_add+0x637/0x840
__auxiliary_device_add+0x3b/0xa0
add_adev+0xc9/0x140 [mlx5_core]
mlx5_rescan_drivers_locked+0x22a/0x310 [mlx5_core]
mlx5_register_device+0x53/0xa0 [mlx5_core]
mlx5_init_one_devl_locked+0x5c4/0x9c0 [mlx5_core]
mlx5_init_one+0x3b/0x60 [mlx5_core]
probe_one+0x44c/0x730 [mlx5_core]
local_pci_probe+0x3e/0x90
pci_device_probe+0xbf/0x210
? kernfs_create_link+0x5d/0xa0
? sysfs_do_create_link_sd+0x60/0xc0
really_probe+0xc9/0x3e0
? driver_probe_device+0x90/0x90
__driver_probe_device+0x80/0x160
driver_probe_device+0x1e/0x90
__device_attach_driver+0x7d/0x100
bus_for_each_drv+0x80/0xd0
__device_attach+0xbc/0x1f0
pci_bus_add_device+0x54/0x80
pci_iov_add_virtfn+0x2e6/0x320
sriov_enable+0x208/0x420
mlx5_core_sriov_configure+0x9e/0x200 [mlx5_core]
sriov_numvfs_store+0xae/0x1a0
kernfs_fop_write_iter+0x10c/0x1a0
vfs_write+0x291/0x3c0
ksys_write+0x5f/0xe0
do_syscall_64+0x3d/0x90
entry_SYSCALL_64_after_hwframe+0x46/0xb0
CR2: 0000000000000000
---[ end trace 0000000000000000 ]---
[ Backport to 6.6.y: added an internal suspend helper so probe unwind
cleans resources without changing the registered-device suspend path. ]
Fixes: 2c3b5beec46a ("net/mlx5e: More generic netdev management API")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://lore.kernel.org/r/20240509112951.590184-2-tariqt@nvidia.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
Assisted-by: LLM
Signed-off-by: Artem Dinaburg <artem@trailofbits.com>
---
Hi Greg, Sasha, and net maintainers,
I am working through the small CVE backports still missing from 6.6.y.
This one addresses CVE-2024-38608. It releases probe resources even when
register_netdev() failed before device-present state.
The fix is already present in 6.12.y, 6.18.y, and 7.2.y, but not in 6.6.y.
This fix also affects 6.1.y, which will need a separate backport; this
submission contains only the 6.6.y patch.
The target-specific adjustment is recorded in the bracketed note above.
Could you please queue it for 6.6.y?
CVE: CVE-2024-38608
Upstream: 3d5918477f94e4c2f064567875c475468e264644
AI assistance: An LLM helped identify, adapt, and validate this backport; I
reviewed the resulting code and validation evidence.
Thanks,
Artem Dinaburg
drivers/net/ethernet/mellanox/mlx5/core/en_main.c | 13 +++++++++----
1 file changed, 9 insertions(+), 4 deletions(-)
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
index 97fa2095498cc8..bfccdcd7928be0 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
@@ -6018,14 +6018,15 @@ static int mlx5e_resume(struct auxiliary_device *adev)
return 0;
}
-static int mlx5e_suspend(struct auxiliary_device *adev, pm_message_t state)
+static int mlx5e_suspend_common(struct auxiliary_device *adev,
+ bool pre_netdev_reg)
{
struct mlx5e_dev *mlx5e_dev = auxiliary_get_drvdata(adev);
struct mlx5e_priv *priv = mlx5e_dev->priv;
struct net_device *netdev = priv->netdev;
struct mlx5_core_dev *mdev = priv->mdev;
- if (!netif_device_present(netdev)) {
+ if (!pre_netdev_reg && !netif_device_present(netdev)) {
if (test_bit(MLX5E_STATE_DESTROYING, &priv->state))
mlx5e_destroy_mdev_resources(mdev);
return -ENODEV;
@@ -6036,6 +6037,11 @@ static int mlx5e_suspend(struct auxiliary_device *adev, pm_message_t state)
return 0;
}
+static int mlx5e_suspend(struct auxiliary_device *adev, pm_message_t state)
+{
+ return mlx5e_suspend_common(adev, false);
+}
+
static int mlx5e_probe(struct auxiliary_device *adev,
const struct auxiliary_device_id *id)
{
@@ -6044,7 +6050,6 @@ static int mlx5e_probe(struct auxiliary_device *adev,
struct mlx5_core_dev *mdev = edev->mdev;
struct mlx5e_dev *mlx5e_dev;
struct net_device *netdev;
- pm_message_t state = {};
struct mlx5e_priv *priv;
int err;
@@ -6107,7 +6112,7 @@ static int mlx5e_probe(struct auxiliary_device *adev,
return 0;
err_resume:
- mlx5e_suspend(adev, state);
+ mlx5e_suspend_common(adev, true);
err_profile_cleanup:
profile->cleanup(priv);
err_destroy_netdev:
--
2.39.5
^ permalink raw reply [flat|nested] 2+ messages in thread* Re: [PATCH 6.6.y] net/mlx5e: Fix netif state handling
2026-09-30 1:21 [PATCH 6.6.y] net/mlx5e: Fix netif state handling Artem Dinaburg
@ 2026-09-30 1:23 ` netdev-bot+sinfo
0 siblings, 0 replies; 2+ messages in thread
From: netdev-bot+sinfo @ 2026-09-30 1:23 UTC (permalink / raw)
To: Artem Dinaburg
Cc: stable, Greg Kroah-Hartman, Sasha Levin, Shay Drory,
Tariq Toukan, Simon Horman, Jakub Kicinski, Saeed Mahameed,
Leon Romanovsky, David S. Miller, Eric Dumazet, Paolo Abeni,
netdev, linux-rdma, linux-kernel, Mark Bloch, Leon Romanovsky,
Andrew Lunn, erezsh
Hi!
This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:
- What hardware the change was tested on. For driver fixes please
mention the device (and if relevant firmware version) used for
testing, or say that the change was not tested on real hardware.
Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.
The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-30 1:23 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 1:21 [PATCH 6.6.y] net/mlx5e: Fix netif state handling Artem Dinaburg
2026-09-30 1:23 ` netdev-bot+sinfo
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®