* [PATCH net-next v5 01/19] net: stmmac: unwind the WoL IRQ after a safety IRQ request failure
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 02/19] net: stmmac: request the MDIO reset GPIO only once James Hilliard
` (19 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
The IRQ setup paths request the MAC IRQ, then the optional WoL IRQ, then
the common safety IRQ. If requesting the safety IRQ fails, cleanup must
release the WoL and MAC IRQs, but not the failed safety IRQ.
The REQ_IRQ_ERR_SFTY case instead frees the safety IRQ and skips the WoL
IRQ. Allocation fault injection during live XDP reopening reproduces a
"Trying to free already-free IRQ" warning and leaves the WoL handler
registered after the datapath resources have been released. A subsequent
open can then fail to request that still-owned IRQ.
Move safety IRQ cleanup before REQ_IRQ_ERR_SFTY and WoL IRQ cleanup
after it. This restores reverse acquisition order for both shared and
MSI IRQ setup, including unwind after failures in later per-queue IRQ
requests.
Fixes: 5c2215167d12 ("net: stmmac: Add driver support for common safety IRQ")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 3ad9252bf6ae..894de213b92f 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -3828,13 +3828,13 @@ static void stmmac_free_irq(struct net_device *dev,
free_irq(msi->sfty_ce_irq, dev);
fallthrough;
case REQ_IRQ_ERR_SFTY_CE:
- if (priv->wol_irq > 0 && priv->wol_irq != dev->irq)
- free_irq(priv->wol_irq, dev);
- fallthrough;
- case REQ_IRQ_ERR_SFTY:
if (priv->sfty_irq > 0 && priv->sfty_irq != dev->irq)
free_irq(priv->sfty_irq, dev);
fallthrough;
+ case REQ_IRQ_ERR_SFTY:
+ if (priv->wol_irq > 0 && priv->wol_irq != dev->irq)
+ free_irq(priv->wol_irq, dev);
+ fallthrough;
case REQ_IRQ_ERR_WOL:
free_irq(dev->irq, dev);
fallthrough;
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 02/19] net: stmmac: request the MDIO reset GPIO only once
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 01/19] net: stmmac: unwind the WoL IRQ after a safety IRQ request failure James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 23:35 ` Linus Walleij
2026-09-27 21:59 ` [PATCH net-next v5 03/19] net: phylink: allow stopping a suspended instance James Hilliard
` (18 subsequent siblings)
20 siblings, 1 reply; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard, stable
From: Linkui Xiao <xiaolinkui@kylinos.cn>
stmmac_mdio_reset() calls devm_gpiod_get_optional() every time it runs.
A GPIO line can only be requested once, so from the second call on
gpiod_request_commit() returns -EBUSY. devm_gpiod_get_optional() only
turns -ENOENT into NULL, hence the error is passed straight back and
stmmac_mdio_reset() bails out before pulsing "snps,reset" and before
running the STE101P MDC workaround.
The first call, made by of_mdiobus_register(), succeeds, so the failure
is only visible later on: every resume that does not use WoL goes
through stmmac_resume() -> stmmac_mdio_reset(), and that caller ignores
the return value, so the PHY silently stays un-reset.
The descriptor used to be requested exactly once: stmmac_mdio_reset()
resolved "snps,reset-gpio" itself and cached the GPIO number in
stmmac_mdio_bus_data::reset_gpio, and commit ae26c1c6cb9b ("stmmac: fix
PHY reset during resume") relies on that cache to reuse the line on
every call. commit 7c86f20d15b7 ("net: stmmac: use GPIO descriptors in
stmmac_mdio_reset") replaced it with a devm_gpiod_get_optional() that
caches nothing, so the request is repeated on every call and fails from
the second one on.
Parse the whole reset description, the GPIO and "snps,reset-delays-us",
in stmmac_mdio_register() at probe time, and keep it in struct
stmmac_priv. This is where devm-gpiod is meant to be used: the line is
acquired with the device and released with it, and any failure to
acquire it is reported during probe instead of being ignored by
stmmac_resume(). stmmac_mdio_reset() then only pulses the cached line,
with the delays that were read once and for all at probe time.
Cache the request and delays for DT devices regardless of
mdio_bus_data->needs_reset. That flag controls the registration-time
bus reset callback, but system resume calls stmmac_mdio_reset()
directly. The reset routine no longer looks at the device tree: where
the description is absent the cached descriptor is NULL and the delays
are zero, so the pulse remains a no-op.
Keep acquisition conditional on CONFIG_STMMAC_PLATFORM, matching the
reset callback, so non-platform configurations do not request an unused
GPIO. Also skip acquisition for a disabled MDIO child: registering that
bus returns -ENODEV without calling its reset callback, and the driver
must retain the existing disabled-bus success path even if the unused
GPIO is unavailable. Remove the unnecessary gpio_desc forward declaration.
Fixes: 7c86f20d15b7 ("net: stmmac: use GPIO descriptors in stmmac_mdio_reset")
Cc: stable@vger.kernel.org
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
Co-developed-by: James Hilliard <james.hilliard1@gmail.com>
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 2 +
drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c | 50 +++++++++++------------
2 files changed, 27 insertions(+), 25 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 4fc96b317d79..83c30b39f704 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -287,6 +287,8 @@ struct stmmac_priv {
unsigned int pause_time;
struct mii_bus *mii;
+ struct gpio_desc *mdio_reset_gpio;
+ u32 mdio_reset_delays[3];
struct stmmac_pcs *integrated_pcs;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c
index afe98ff5bdcb..63287ad9652f 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c
@@ -384,33 +384,16 @@ int stmmac_mdio_reset(struct mii_bus *bus)
struct stmmac_priv *priv = netdev_priv(bus->priv);
unsigned int mii_address = priv->hw->mii.addr;
-#ifdef CONFIG_OF
- if (priv->device->of_node) {
- struct gpio_desc *reset_gpio;
- u32 delays[3] = { 0, 0, 0 };
+ if (priv->mdio_reset_delays[0])
+ msleep(DIV_ROUND_UP(priv->mdio_reset_delays[0], 1000));
- reset_gpio = devm_gpiod_get_optional(priv->device,
- "snps,reset",
- GPIOD_OUT_LOW);
- if (IS_ERR(reset_gpio))
- return PTR_ERR(reset_gpio);
+ gpiod_set_value_cansleep(priv->mdio_reset_gpio, 1);
+ if (priv->mdio_reset_delays[1])
+ msleep(DIV_ROUND_UP(priv->mdio_reset_delays[1], 1000));
- device_property_read_u32_array(priv->device,
- "snps,reset-delays-us",
- delays, ARRAY_SIZE(delays));
-
- if (delays[0])
- msleep(DIV_ROUND_UP(delays[0], 1000));
-
- gpiod_set_value_cansleep(reset_gpio, 1);
- if (delays[1])
- msleep(DIV_ROUND_UP(delays[1], 1000));
-
- gpiod_set_value_cansleep(reset_gpio, 0);
- if (delays[2])
- msleep(DIV_ROUND_UP(delays[2], 1000));
- }
-#endif
+ gpiod_set_value_cansleep(priv->mdio_reset_gpio, 0);
+ if (priv->mdio_reset_delays[2])
+ msleep(DIV_ROUND_UP(priv->mdio_reset_delays[2], 1000));
/* This is a workaround for problems with the STE101P PHY.
* It doesn't complete its reset until at least one clock cycle
@@ -608,6 +591,23 @@ int stmmac_mdio_register(struct net_device *ndev)
if (!mdio_bus_data)
return 0;
+ /* Resume calls stmmac_mdio_reset() even when registration does not
+ * install a bus reset callback, so cache its resources in both cases.
+ */
+ if (IS_ENABLED(CONFIG_STMMAC_PLATFORM) && dev_of_node(priv->device) &&
+ (!mdio_node || of_device_is_available(mdio_node))) {
+ priv->mdio_reset_gpio =
+ devm_gpiod_get_optional(priv->device, "snps,reset",
+ GPIOD_OUT_LOW);
+ if (IS_ERR(priv->mdio_reset_gpio))
+ return PTR_ERR(priv->mdio_reset_gpio);
+
+ device_property_read_u32_array(priv->device,
+ "snps,reset-delays-us",
+ priv->mdio_reset_delays,
+ ARRAY_SIZE(priv->mdio_reset_delays));
+ }
+
stmmac_mdio_bus_config(priv);
new_bus = mdiobus_alloc();
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* Re: [PATCH net-next v5 02/19] net: stmmac: request the MDIO reset GPIO only once
2026-09-27 21:59 ` [PATCH net-next v5 02/19] net: stmmac: request the MDIO reset GPIO only once James Hilliard
@ 2026-09-27 23:35 ` Linus Walleij
2026-09-27 23:49 ` James Hilliard
0 siblings, 1 reply; 27+ messages in thread
From: Linus Walleij @ 2026-09-27 23:35 UTC (permalink / raw)
To: James Hilliard
Cc: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Martin Blumenstingl,
Magnus Karlsson, Maciej Fijalkowski, Simon Horman,
Björn Töpel, Thierry Reding, Jonathan Hunter,
Chen-Yu Tsai, Jernej Skrabec, Samuel Holland, Jose Abreu, Yao Zi,
Philipp Zabel, Richard Genoud, Alastair D'Silva,
Maxime Ripard, netdev, linux-kernel, linux-stm32,
linux-arm-kernel, bpf, ZhaoJinming, Lorenzo Bianconi, Ding Hui,
Linkui Xiao, Linkui Xiao, linux-tegra, linux-sunxi, stable
Hi James,
On Sun, Sep 27, 2026 at 11:59 PM James Hilliard
<james.hilliard1@gmail.com> wrote:
> From: Linkui Xiao <xiaolinkui@kylinos.cn>
>
> stmmac_mdio_reset() calls devm_gpiod_get_optional() every time it runs.
> A GPIO line can only be requested once, so from the second call on
> gpiod_request_commit() returns -EBUSY. devm_gpiod_get_optional() only
> turns -ENOENT into NULL, hence the error is passed straight back and
> stmmac_mdio_reset() bails out before pulsing "snps,reset" and before
> running the STE101P MDC workaround.
>
> The first call, made by of_mdiobus_register(), succeeds, so the failure
> is only visible later on: every resume that does not use WoL goes
> through stmmac_resume() -> stmmac_mdio_reset(), and that caller ignores
> the return value, so the PHY silently stays un-reset.
>
> The descriptor used to be requested exactly once: stmmac_mdio_reset()
> resolved "snps,reset-gpio" itself and cached the GPIO number in
> stmmac_mdio_bus_data::reset_gpio, and commit ae26c1c6cb9b ("stmmac: fix
> PHY reset during resume") relies on that cache to reuse the line on
> every call. commit 7c86f20d15b7 ("net: stmmac: use GPIO descriptors in
> stmmac_mdio_reset") replaced it with a devm_gpiod_get_optional() that
> caches nothing, so the request is repeated on every call and fails from
> the second one on.
>
> Parse the whole reset description, the GPIO and "snps,reset-delays-us",
> in stmmac_mdio_register() at probe time, and keep it in struct
> stmmac_priv. This is where devm-gpiod is meant to be used: the line is
> acquired with the device and released with it, and any failure to
> acquire it is reported during probe instead of being ignored by
> stmmac_resume(). stmmac_mdio_reset() then only pulses the cached line,
> with the delays that were read once and for all at probe time.
>
> Cache the request and delays for DT devices regardless of
> mdio_bus_data->needs_reset. That flag controls the registration-time
> bus reset callback, but system resume calls stmmac_mdio_reset()
> directly. The reset routine no longer looks at the device tree: where
> the description is absent the cached descriptor is NULL and the delays
> are zero, so the pulse remains a no-op.
>
> Keep acquisition conditional on CONFIG_STMMAC_PLATFORM, matching the
> reset callback, so non-platform configurations do not request an unused
> GPIO. Also skip acquisition for a disabled MDIO child: registering that
> bus returns -ENODEV without calling its reset callback, and the driver
> must retain the existing disabled-bus success path even if the unused
> GPIO is unavailable. Remove the unnecessary gpio_desc forward declaration.
>
> Fixes: 7c86f20d15b7 ("net: stmmac: use GPIO descriptors in stmmac_mdio_reset")
> Cc: stable@vger.kernel.org
> Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
> Co-developed-by: James Hilliard <james.hilliard1@gmail.com>
> Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
Dostoyevsky commit message, I didn't read it. Ask the agent
to be terse.
I looked at the code and from a GPIO PoV it does the right
thing: use 1 as asserted and 0 as de-asserted RESET line.
Reviewed-by: Linus Walleij <linusw@kernel.org>
Yours,
Linus Walleij
^ permalink raw reply [flat|nested] 27+ messages in thread* Re: [PATCH net-next v5 02/19] net: stmmac: request the MDIO reset GPIO only once
2026-09-27 23:35 ` Linus Walleij
@ 2026-09-27 23:49 ` James Hilliard
0 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 23:49 UTC (permalink / raw)
To: Linus Walleij
Cc: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Martin Blumenstingl,
Magnus Karlsson, Maciej Fijalkowski, Simon Horman,
Björn Töpel, Thierry Reding, Jonathan Hunter,
Chen-Yu Tsai, Jernej Skrabec, Samuel Holland, Jose Abreu, Yao Zi,
Philipp Zabel, Richard Genoud, Alastair D'Silva,
Maxime Ripard, netdev, linux-kernel, linux-stm32,
linux-arm-kernel, bpf, ZhaoJinming, Lorenzo Bianconi, Ding Hui,
Linkui Xiao, Linkui Xiao, linux-tegra, linux-sunxi, stable
On Sun, Sep 27, 2026 at 5:35 PM Linus Walleij <linusw@kernel.org> wrote:
>
> Hi James,
>
> On Sun, Sep 27, 2026 at 11:59 PM James Hilliard
> <james.hilliard1@gmail.com> wrote:
>
> > From: Linkui Xiao <xiaolinkui@kylinos.cn>
> >
> > stmmac_mdio_reset() calls devm_gpiod_get_optional() every time it runs.
> > A GPIO line can only be requested once, so from the second call on
> > gpiod_request_commit() returns -EBUSY. devm_gpiod_get_optional() only
> > turns -ENOENT into NULL, hence the error is passed straight back and
> > stmmac_mdio_reset() bails out before pulsing "snps,reset" and before
> > running the STE101P MDC workaround.
> >
> > The first call, made by of_mdiobus_register(), succeeds, so the failure
> > is only visible later on: every resume that does not use WoL goes
> > through stmmac_resume() -> stmmac_mdio_reset(), and that caller ignores
> > the return value, so the PHY silently stays un-reset.
> >
> > The descriptor used to be requested exactly once: stmmac_mdio_reset()
> > resolved "snps,reset-gpio" itself and cached the GPIO number in
> > stmmac_mdio_bus_data::reset_gpio, and commit ae26c1c6cb9b ("stmmac: fix
> > PHY reset during resume") relies on that cache to reuse the line on
> > every call. commit 7c86f20d15b7 ("net: stmmac: use GPIO descriptors in
> > stmmac_mdio_reset") replaced it with a devm_gpiod_get_optional() that
> > caches nothing, so the request is repeated on every call and fails from
> > the second one on.
> >
> > Parse the whole reset description, the GPIO and "snps,reset-delays-us",
> > in stmmac_mdio_register() at probe time, and keep it in struct
> > stmmac_priv. This is where devm-gpiod is meant to be used: the line is
> > acquired with the device and released with it, and any failure to
> > acquire it is reported during probe instead of being ignored by
> > stmmac_resume(). stmmac_mdio_reset() then only pulses the cached line,
> > with the delays that were read once and for all at probe time.
> >
> > Cache the request and delays for DT devices regardless of
> > mdio_bus_data->needs_reset. That flag controls the registration-time
> > bus reset callback, but system resume calls stmmac_mdio_reset()
> > directly. The reset routine no longer looks at the device tree: where
> > the description is absent the cached descriptor is NULL and the delays
> > are zero, so the pulse remains a no-op.
> >
> > Keep acquisition conditional on CONFIG_STMMAC_PLATFORM, matching the
> > reset callback, so non-platform configurations do not request an unused
> > GPIO. Also skip acquisition for a disabled MDIO child: registering that
> > bus returns -ENODEV without calling its reset callback, and the driver
> > must retain the existing disabled-bus success path even if the unused
> > GPIO is unavailable. Remove the unnecessary gpio_desc forward declaration.
> >
> > Fixes: 7c86f20d15b7 ("net: stmmac: use GPIO descriptors in stmmac_mdio_reset")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
> > Co-developed-by: James Hilliard <james.hilliard1@gmail.com>
> > Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
>
> Dostoyevsky commit message, I didn't read it. Ask the agent
> to be terse.
That commit message style mostly just came from the imported patch,
I'll change it to be more terse in the next revision:
https://lore.kernel.org/all/20260921015727.2643540-1-xiaolinkui@126.com/
> I looked at the code and from a GPIO PoV it does the right
> thing: use 1 as asserted and 0 as de-asserted RESET line.
>
> Reviewed-by: Linus Walleij <linusw@kernel.org>
>
> Yours,
> Linus Walleij
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH net-next v5 03/19] net: phylink: allow stopping a suspended instance
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 01/19] net: stmmac: unwind the WoL IRQ after a safety IRQ request failure James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 02/19] net: stmmac: request the MDIO reset GPIO only once James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 04/19] xsk: freeze deferred pool teardown during system sleep James Hilliard
` (17 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
If a network driver's system resume fails before phylink_resume(), the
network device can remain administratively up with phylink suspended.
Closing the interface then needs to terminate that suspended instance.
Calling phylink_resume() merely to make phylink_stop() work is not a
safe substitute: resume reconfigures the MAC and restarts link
resolution, although the driver has not successfully restored the MAC.
This is a missing suspend-to-stop transition, independent of the reason
hardware restoration failed. No MAC recovery policy belongs in phylink;
the driver still decides whether to retry resume or wait for an ordinary
administrative down/up cycle.
Without MAC Wake-on-LAN, phylink_suspend() has already called
phylink_stop(). Do not repeat PHY, SFP and PCS shutdown. However,
phylink_prepare_resume() may have powered that stopped PHY back up to
provide the receive clock for MAC reset. Suspend it again if WoL
permits, without repeating phy_stop() on a PHY which is already halted.
With MAC Wake-on-LAN, suspend deliberately defers mac_link_down() and
sets PHYLINK_DISABLE_MAC_WOL. Finish that deferred link-down, drain
resolution work and clear the WoL disable bit while retaining
PHYLINK_DISABLE_STOPPED. Otherwise a subsequent start cannot resolve the
link.
Also undo PHY speed control performed by phylink_suspend(). As Andrew
Lunn pointed out, phylink_start() does not restore the advertised speeds
that phylink_resume() normally restores. Track suspend-owned speed
control and restore the saved advertisement from either resume or
suspended stop, including when PHY shutdown has already completed.
An explicit driver speed-down request, such as stmmac's close-time power
saving, must remain in effect until its matching speed-up. Restore any
suspend-owned reduction before applying that request, so it cannot save
the reduced advertisement over the original one. The following stop must
not undo the driver's new reduction.
Document that a suspended instance can be stopped directly. This neither
resumes the PHY nor reconfigures or brings up the MAC.
Fixes: f97493657c63 ("net: phylink: add suspend/resume support")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/phy/phylink.c | 51 ++++++++++++++++++++++++++++++++++++++++++++---
1 file changed, 48 insertions(+), 3 deletions(-)
diff --git a/drivers/net/phy/phylink.c b/drivers/net/phy/phylink.c
index a7d086cdc9b2..f97fee908df6 100644
--- a/drivers/net/phy/phylink.c
+++ b/drivers/net/phy/phylink.c
@@ -78,6 +78,7 @@ struct phylink {
bool link_failed;
bool suspend_link_up;
+ bool suspend_speed_down;
bool force_major_config;
bool major_config_failed;
bool mac_supports_eee_ops;
@@ -2618,6 +2619,14 @@ void phylink_start(struct phylink *pl)
}
EXPORT_SYMBOL_GPL(phylink_start);
+static void phylink_restore_suspend_speed(struct phylink *pl)
+{
+ if (pl->suspend_speed_down) {
+ phylink_speed_up(pl);
+ pl->suspend_speed_down = false;
+ }
+}
+
/**
* phylink_stop() - stop a phylink instance
* @pl: a pointer to a &struct phylink returned from phylink_create()
@@ -2629,11 +2638,30 @@ EXPORT_SYMBOL_GPL(phylink_start);
*
* This will synchronously bring down the link if the link is not already
* down (in other words, it will trigger a mac_link_down() method call.)
+ * A suspended instance may be stopped without first calling phylink_resume().
+ * In particular, closing a device after a failed resume must not restart the
+ * link or reconfigure the MAC just to finish shutting it down.
+ * Any PHY advertisement reduced by phylink_suspend() is restored as part
+ * of this transition.
+ * If phylink_prepare_resume() powered up an already stopped PHY, suspend
+ * it again when Wake-on-LAN permits.
*/
void phylink_stop(struct phylink *pl)
{
ASSERT_RTNL();
+ /* Also undo PHY speed control when terminating a suspended instance. */
+ phylink_restore_suspend_speed(pl);
+
+ if (test_bit(PHYLINK_DISABLE_STOPPED, &pl->phylink_disable_state)) {
+ /* A failed MAC resume may have called phylink_prepare_resume()
+ * and powered the stopped PHY back up to supply its RX clock.
+ */
+ if (pl->phydev)
+ phy_suspend(pl->phydev);
+ return;
+ }
+
if (pl->sfp_bus)
sfp_upstream_stop(pl->sfp_bus);
if (pl->phydev)
@@ -2646,6 +2674,16 @@ void phylink_stop(struct phylink *pl)
phylink_run_resolve_and_disable(pl, PHYLINK_DISABLE_STOPPED);
+ if (test_bit(PHYLINK_DISABLE_MAC_WOL, &pl->phylink_disable_state)) {
+ /* Finish the link-down deferred by MAC WoL, without restarting. */
+ flush_work(&pl->resolve);
+ mutex_lock(&pl->state_mutex);
+ if (pl->suspend_link_up)
+ phylink_link_down(pl);
+ __clear_bit(PHYLINK_DISABLE_MAC_WOL, &pl->phylink_disable_state);
+ mutex_unlock(&pl->state_mutex);
+ }
+
pl->pcs_state = PCS_STATE_DOWN;
phylink_pcs_disable(pl->pcs);
@@ -2777,8 +2815,10 @@ void phylink_suspend(struct phylink *pl, bool mac_wol)
phylink_stop(pl);
}
- if (phylink_phy_pm_speed_ctrl(pl))
+ if (phylink_phy_pm_speed_ctrl(pl)) {
phylink_speed_down(pl, false);
+ pl->suspend_speed_down = true;
+ }
}
EXPORT_SYMBOL_GPL(phylink_suspend);
@@ -2818,8 +2858,7 @@ void phylink_resume(struct phylink *pl)
{
ASSERT_RTNL();
- if (phylink_phy_pm_speed_ctrl(pl))
- phylink_speed_up(pl);
+ phylink_restore_suspend_speed(pl);
if (test_bit(PHYLINK_DISABLE_MAC_WOL, &pl->phylink_disable_state)) {
/* Wake-on-Lan enabled, MAC handling */
@@ -3687,6 +3726,12 @@ int phylink_speed_down(struct phylink *pl, bool sync)
ASSERT_RTNL();
+ /* An explicit request takes over from suspend-time speed control.
+ * Restore the original advertisement before saving it again, so a
+ * repeated speed-down cannot replace it with the reduced advertisement.
+ */
+ phylink_restore_suspend_speed(pl);
+
if (!pl->sfp_bus && pl->phydev)
ret = phy_speed_down(pl->phydev, sync);
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 04/19] xsk: freeze deferred pool teardown during system sleep
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (2 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 03/19] net: phylink: allow stopping a suspended instance James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-28 12:16 ` Björn Töpel
2026-09-27 21:59 ` [PATCH net-next v5 05/19] net: stmmac: embed struct stmmac_est in stmmac_priv struct James Hilliard
` (16 subsequent siblings)
20 siblings, 1 reply; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Pool destruction calls the driver under RTNL from system_wq. That queue
is not frozen during system sleep, so ndo_bpf() can run after the device
suspend callback has gated its clocks or after the noirq phase.
Use the freezable workqueue for deferred pool release. Work already
running completes before device suspend, and newly queued destruction
waits until device resume. The pool, UMEM and netdev references remain
owned by the work until then. This does not replace driver error
handling after a failed resume.
Fixes: 1c1efc2af158 ("xsk: Create and free buffer pool independently from umem")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
net/xdp/xsk_buff_pool.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
index 9d2d94f1fb75..c58f56f24a9c 100644
--- a/net/xdp/xsk_buff_pool.c
+++ b/net/xdp/xsk_buff_pool.c
@@ -337,7 +337,10 @@ bool xp_put_pool(struct xsk_buff_pool *pool)
if (refcount_dec_and_test(&pool->users)) {
INIT_WORK(&pool->work, xp_release_deferred);
- schedule_work(&pool->work);
+ /* Teardown calls ndo_bpf(), which may need powered hardware.
+ * RTNL alone does not exclude the device's system PM callbacks.
+ */
+ queue_work(system_freezable_wq, &pool->work);
return true;
}
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* Re: [PATCH net-next v5 04/19] xsk: freeze deferred pool teardown during system sleep
2026-09-27 21:59 ` [PATCH net-next v5 04/19] xsk: freeze deferred pool teardown during system sleep James Hilliard
@ 2026-09-28 12:16 ` Björn Töpel
0 siblings, 0 replies; 27+ messages in thread
From: Björn Töpel @ 2026-09-28 12:16 UTC (permalink / raw)
To: James Hilliard, Russell King, Andrew Lunn, Heiner Kallweit,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Thierry Reding, Jonathan Hunter, Chen-Yu Tsai,
Jernej Skrabec, Samuel Holland, Jose Abreu, Yao Zi,
Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
James Hilliard <james.hilliard1@gmail.com> writes:
> Pool destruction calls the driver under RTNL from system_wq. That queue
> is not frozen during system sleep, so ndo_bpf() can run after the device
> suspend callback has gated its clocks or after the noirq phase.
>
> Use the freezable workqueue for deferred pool release. Work already
> running completes before device suspend, and newly queued destruction
> waits until device resume. The pool, UMEM and netdev references remain
> owned by the work until then. This does not replace driver error
> handling after a failed resume.
>
> Fixes: 1c1efc2af158 ("xsk: Create and free buffer pool independently from umem")
> Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
This fix should go as a separate net fix for AF_XDP (potentially paired
with a minimal stmac-fix).
> diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
> index 9d2d94f1fb75..c58f56f24a9c 100644
> --- a/net/xdp/xsk_buff_pool.c
> +++ b/net/xdp/xsk_buff_pool.c
> @@ -337,7 +337,10 @@ bool xp_put_pool(struct xsk_buff_pool *pool)
>
> if (refcount_dec_and_test(&pool->users)) {
> INIT_WORK(&pool->work, xp_release_deferred);
> - schedule_work(&pool->work);
> + /* Teardown calls ndo_bpf(), which may need powered hardware.
> + * RTNL alone does not exclude the device's system PM callbacks.
> + */
Remove the comment, please. The commit message is enough in this case.
Björn
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH net-next v5 05/19] net: stmmac: embed struct stmmac_est in stmmac_priv struct
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (3 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 04/19] xsk: freeze deferred pool teardown during system sleep James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 06/19] net: stmmac: pass the desired EST enable state to est_configure() James Hilliard
` (15 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
From: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
This is a preliminary change to fix EST reconfiguration in the open
and resume paths: the taprio offload must be re-applied after the DMA
soft reset clears the MTL_EST registers, but the current layout makes
that fragile.
priv->est is currently a pointer allocated with devm_kzalloc() on the
first taprio REPLACE, and the mutex guarding it (priv->est_lock) is
initialized at the same time. That ties the lock's validity to whether
taprio has ever been configured, so the EST parameters can not be
read under the lock (e.g. to check priv->est->enable in the open and
resume paths) before the first offload setup.
Embed struct stmmac_est into struct stmmac_priv and initialize the mutex
in probe(). This makes the code simpler (no logical changes added).
Moreover, the lock is now unconditionally valid, so the enable flag can
be inspected under the lock from any control path.
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 2 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 17 +++----
drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c | 22 ++++----
drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c | 61 ++++++++++-------------
4 files changed, 45 insertions(+), 57 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 83c30b39f704..c65d95fc4760 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -300,7 +300,7 @@ struct stmmac_priv {
struct plat_stmmacenet_data *plat;
/* Protect est parameters */
struct mutex est_lock;
- struct stmmac_est *est;
+ struct stmmac_est est;
struct dma_features dma_cap;
struct stmmac_counters mmc;
int hw_cap_support;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 894de213b92f..b233105cb449 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -2745,9 +2745,8 @@ static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
if (!xsk_tx_peek_desc(pool, &xdp_desc))
break;
- if (priv->est && priv->est->enable &&
- priv->est->max_sdu[queue] &&
- xdp_desc.len > priv->est->max_sdu[queue]) {
+ if (priv->est.enable && priv->est.max_sdu[queue] &&
+ xdp_desc.len > priv->est.max_sdu[queue]) {
priv->xstats.max_sdu_txq_drop[queue]++;
continue;
}
@@ -4843,13 +4842,12 @@ static netdev_tx_t stmmac_xmit(struct sk_buff *skb, struct net_device *dev)
if (skb_is_gso(skb))
return stmmac_tso_xmit(skb, dev);
- if (priv->est && priv->est->enable &&
- priv->est->max_sdu[queue]) {
+ if (priv->est.enable && priv->est.max_sdu[queue]) {
sdu_len = skb->len;
/* Add VLAN tag length if VLAN tag insertion offload is requested */
if (priv->dma_cap.vlins && skb_vlan_tag_present(skb))
sdu_len += VLAN_HLEN;
- if (sdu_len > priv->est->max_sdu[queue]) {
+ if (sdu_len > priv->est.max_sdu[queue]) {
priv->xstats.max_sdu_txq_drop[queue]++;
goto max_sdu_err;
}
@@ -5253,9 +5251,8 @@ static int stmmac_xdp_xmit_xdpf(struct stmmac_priv *priv, int queue,
if (stmmac_tx_avail(priv, queue) < STMMAC_TX_THRESH(priv))
return STMMAC_XDP_CONSUMED;
- if (priv->est && priv->est->enable &&
- priv->est->max_sdu[queue] &&
- xdpf->len > priv->est->max_sdu[queue]) {
+ if (priv->est.enable && priv->est.max_sdu[queue] &&
+ xdpf->len > priv->est.max_sdu[queue]) {
priv->xstats.max_sdu_txq_drop[queue]++;
return STMMAC_XDP_CONSUMED;
}
@@ -8106,6 +8103,7 @@ static int __stmmac_dvr_probe(struct device *device,
stmmac_napi_add(ndev);
mutex_init(&priv->lock);
+ mutex_init(&priv->est_lock);
rwlock_init(&priv->ptp_lock);
stmmac_fpe_init(priv);
@@ -8238,6 +8236,7 @@ void stmmac_dvr_remove(struct device *dev)
stmmac_mdio_unregister(ndev);
destroy_workqueue(priv->wq);
+ mutex_destroy(&priv->est_lock);
mutex_destroy(&priv->lock);
bitmap_free(priv->af_xdp_zc_qps);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
index 3bfcc9760dce..2a4099fe470c 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
@@ -69,11 +69,11 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
nsec = reminder;
/* If EST is enabled, disabled it before adjust ptp time. */
- if (priv->est && priv->est->enable) {
+ if (priv->est.enable) {
est_rst = true;
mutex_lock(&priv->est_lock);
- priv->est->enable = false;
- stmmac_est_configure(priv, priv, priv->est,
+ priv->est.enable = false;
+ stmmac_est_configure(priv, priv, &priv->est,
priv->plat->clk_ptp_rate);
mutex_unlock(&priv->est_lock);
}
@@ -91,19 +91,19 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
mutex_lock(&priv->est_lock);
priv->ptp_clock_ops.gettime64(&priv->ptp_clock_ops, ¤t_time);
current_time_ns = timespec64_to_ktime(current_time);
- time.tv_nsec = priv->est->btr_reserve[0];
- time.tv_sec = priv->est->btr_reserve[1];
+ time.tv_nsec = priv->est.btr_reserve[0];
+ time.tv_sec = priv->est.btr_reserve[1];
basetime = timespec64_to_ktime(time);
- cycle_time = (u64)priv->est->ctr[1] * NSEC_PER_SEC +
- priv->est->ctr[0];
+ cycle_time = (u64)priv->est.ctr[1] * NSEC_PER_SEC +
+ priv->est.ctr[0];
time = stmmac_calc_tas_basetime(basetime,
current_time_ns,
cycle_time);
- priv->est->btr[0] = (u32)time.tv_nsec;
- priv->est->btr[1] = (u32)time.tv_sec;
- priv->est->enable = true;
- ret = stmmac_est_configure(priv, priv, priv->est,
+ priv->est.btr[0] = (u32)time.tv_nsec;
+ priv->est.btr[1] = (u32)time.tv_sec;
+ priv->est.enable = true;
+ ret = stmmac_est_configure(priv, priv, &priv->est,
priv->plat->clk_ptp_rate);
mutex_unlock(&priv->est_lock);
if (ret)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
index 42a00446e9b4..357d1eaf0d7d 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
@@ -959,7 +959,7 @@ static void tc_taprio_map_maxsdu_txq(struct stmmac_priv *priv,
count = qopt->mqprio.qopt.count[i];
for (j = offset; j < offset + count; j++)
- priv->est->max_sdu[j] = qopt->max_sdu[i] + ETH_HLEN - ETH_TLEN;
+ priv->est.max_sdu[j] = qopt->max_sdu[i] + ETH_HLEN - ETH_TLEN;
}
}
@@ -1023,24 +1023,15 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
if (qopt->cycle_time_extension >= BIT(wid + 7))
return -ERANGE;
- if (!priv->est) {
- priv->est = devm_kzalloc(priv->device, sizeof(*priv->est),
- GFP_KERNEL);
- if (!priv->est)
- return -ENOMEM;
-
- mutex_init(&priv->est_lock);
- } else {
- mutex_lock(&priv->est_lock);
- memset(priv->est, 0, sizeof(*priv->est));
- mutex_unlock(&priv->est_lock);
- }
+ mutex_lock(&priv->est_lock);
+ memset(&priv->est, 0, sizeof(priv->est));
+ mutex_unlock(&priv->est_lock);
size = qopt->num_entries;
mutex_lock(&priv->est_lock);
- priv->est->gcl_size = size;
- priv->est->enable = qopt->cmd == TAPRIO_CMD_REPLACE;
+ priv->est.gcl_size = size;
+ priv->est.enable = qopt->cmd == TAPRIO_CMD_REPLACE;
mutex_unlock(&priv->est_lock);
for (i = 0; i < size; i++) {
@@ -1065,7 +1056,7 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
return -EOPNOTSUPP;
}
- priv->est->gcl[i] = delta_ns | (gates << wid);
+ priv->est.gcl[i] = delta_ns | (gates << wid);
}
mutex_lock(&priv->est_lock);
@@ -1075,22 +1066,22 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
time = stmmac_calc_tas_basetime(qopt->base_time, current_time_ns,
qopt->cycle_time);
- priv->est->btr[0] = (u32)time.tv_nsec;
- priv->est->btr[1] = (u32)time.tv_sec;
+ priv->est.btr[0] = (u32)time.tv_nsec;
+ priv->est.btr[1] = (u32)time.tv_sec;
qopt_time = ktime_to_timespec64(qopt->base_time);
- priv->est->btr_reserve[0] = (u32)qopt_time.tv_nsec;
- priv->est->btr_reserve[1] = (u32)qopt_time.tv_sec;
+ priv->est.btr_reserve[0] = (u32)qopt_time.tv_nsec;
+ priv->est.btr_reserve[1] = (u32)qopt_time.tv_sec;
ctr = qopt->cycle_time;
- priv->est->ctr[0] = do_div(ctr, NSEC_PER_SEC);
- priv->est->ctr[1] = (u32)ctr;
+ priv->est.ctr[0] = do_div(ctr, NSEC_PER_SEC);
+ priv->est.ctr[1] = (u32)ctr;
- priv->est->ter = qopt->cycle_time_extension;
+ priv->est.ter = qopt->cycle_time_extension;
tc_taprio_map_maxsdu_txq(priv, qopt);
- ret = stmmac_est_configure(priv, priv, priv->est,
+ ret = stmmac_est_configure(priv, priv, &priv->est,
priv->plat->clk_ptp_rate);
mutex_unlock(&priv->est_lock);
if (ret) {
@@ -1106,19 +1097,17 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
return 0;
disable:
- if (priv->est) {
- mutex_lock(&priv->est_lock);
- priv->est->enable = false;
- stmmac_est_configure(priv, priv, priv->est,
- priv->plat->clk_ptp_rate);
- /* Reset taprio status */
- for (i = 0; i < priv->plat->tx_queues_to_use; i++) {
- priv->xstats.max_sdu_txq_drop[i] = 0;
- priv->xstats.mtl_est_txq_hlbf[i] = 0;
- priv->xstats.mtl_est_txq_hlbs[i] = 0;
- }
- mutex_unlock(&priv->est_lock);
+ mutex_lock(&priv->est_lock);
+ priv->est.enable = false;
+ stmmac_est_configure(priv, priv, &priv->est,
+ priv->plat->clk_ptp_rate);
+ /* Reset taprio status */
+ for (i = 0; i < priv->plat->tx_queues_to_use; i++) {
+ priv->xstats.max_sdu_txq_drop[i] = 0;
+ priv->xstats.mtl_est_txq_hlbf[i] = 0;
+ priv->xstats.mtl_est_txq_hlbs[i] = 0;
}
+ mutex_unlock(&priv->est_lock);
err = stmmac_fpe_map_preemption_class(priv, priv->dev, extack, 0);
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 06/19] net: stmmac: pass the desired EST enable state to est_configure()
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (4 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 05/19] net: stmmac: embed struct stmmac_est in stmmac_priv struct James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 07/19] net: stmmac: re-apply taprio offload in __stmmac_open() and stmmac_resume() James Hilliard
` (14 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
From: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Pass the desired EST enable state explicitly to est_configure() instead
of having it derive the EEST/EST_INT_EN bits from cfg->enable. This
decouples the hardware programming state from the priv->est.enable
flag, which records whether the taprio offload is attached.
No functional change intended: the callers keep toggling
priv->est.enable around the EST programming, as before.
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/hwif.h | 2 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_est.c | 6 +++---
drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c | 4 ++--
drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c | 4 ++--
4 files changed, 8 insertions(+), 8 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/hwif.h b/drivers/net/ethernet/stmicro/stmmac/hwif.h
index a8a5c8fdd5ed..1fd9f1ab316e 100644
--- a/drivers/net/ethernet/stmicro/stmmac/hwif.h
+++ b/drivers/net/ethernet/stmicro/stmmac/hwif.h
@@ -617,7 +617,7 @@ struct stmmac_mmc_ops {
struct stmmac_est_ops {
int (*configure)(struct stmmac_priv *priv, struct stmmac_est *cfg,
- unsigned int ptp_rate);
+ unsigned int ptp_rate, bool enable);
void (*irq_status)(struct stmmac_priv *priv, struct net_device *dev,
struct stmmac_extra_stats *x, u32 txqcnt);
};
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
index afc516059b89..f15d4d046aa7 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
@@ -26,7 +26,7 @@ static int est_write(void __iomem *est_addr, u32 reg, u32 val, bool gcl)
}
static int est_configure(struct stmmac_priv *priv, struct stmmac_est *cfg,
- unsigned int ptp_rate)
+ unsigned int ptp_rate, bool enable)
{
void __iomem *est_addr = priv->estaddr;
int i, ret = 0;
@@ -62,7 +62,7 @@ static int est_configure(struct stmmac_priv *priv, struct stmmac_est *cfg,
ctrl |= ((NSEC_PER_SEC / ptp_rate) * EST_GMAC5_PTOV_MUL) <<
EST_GMAC5_PTOV_SHIFT;
}
- if (cfg->enable)
+ if (enable)
ctrl |= EST_EEST | EST_SSWL | EST_DFBS;
else
ctrl &= ~EST_EEST;
@@ -70,7 +70,7 @@ static int est_configure(struct stmmac_priv *priv, struct stmmac_est *cfg,
writel(ctrl, est_addr + EST_CONTROL);
/* Configure EST interrupt */
- if (cfg->enable)
+ if (enable)
ctrl = EST_IECGCE | EST_IEHS | EST_IEHF | EST_IEBE | EST_IECC;
else
ctrl = 0;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
index 2a4099fe470c..64d890664421 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
@@ -74,7 +74,7 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
mutex_lock(&priv->est_lock);
priv->est.enable = false;
stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate);
+ priv->plat->clk_ptp_rate, false);
mutex_unlock(&priv->est_lock);
}
@@ -104,7 +104,7 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
priv->est.btr[1] = (u32)time.tv_sec;
priv->est.enable = true;
ret = stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate);
+ priv->plat->clk_ptp_rate, true);
mutex_unlock(&priv->est_lock);
if (ret)
netdev_err(priv->dev, "failed to configure EST\n");
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
index 357d1eaf0d7d..67c6fc32d0ea 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
@@ -1082,7 +1082,7 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
tc_taprio_map_maxsdu_txq(priv, qopt);
ret = stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate);
+ priv->plat->clk_ptp_rate, true);
mutex_unlock(&priv->est_lock);
if (ret) {
netdev_err(priv->dev, "failed to configure EST\n");
@@ -1100,7 +1100,7 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
mutex_lock(&priv->est_lock);
priv->est.enable = false;
stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate);
+ priv->plat->clk_ptp_rate, false);
/* Reset taprio status */
for (i = 0; i < priv->plat->tx_queues_to_use; i++) {
priv->xstats.max_sdu_txq_drop[i] = 0;
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 07/19] net: stmmac: re-apply taprio offload in __stmmac_open() and stmmac_resume()
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (5 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 06/19] net: stmmac: pass the desired EST enable state to est_configure() James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 08/19] net: stmmac: serialize and retain PHC configuration across reset James Hilliard
` (13 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard, Rayagond Kokatanur
From: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
The core soft reset issued in stmmac_init_dma_engine() clears the
MTL_EST registers, but nothing re-applies the taprio offload after it:
priv->est.enable stays true while the hardware EST block is left
disabled. The TX/XDP paths then keep dropping frames larger than
priv->est.max_sdu[] and taprio is reported as offloaded, although the
EST block is not programmed.
Re-apply the taprio offload in __stmmac_open() and stmmac_resume()
after PTP is up. The base time is recomputed from the reserved base
time and the current PTP time, since the timestamp counter has been
re-initialized and the previously programmed base time is stale.
Introduce the stmmac_setup_est utility routine.
While at it, make the EST enable state transition atomic with the
hardware programming: pass the desired enable state explicitly to
est_configure() and hold est_lock across the whole configure sequence
in tc_taprio_configure(), so that the software enable flag and the
hardware state transition atomically. Moreover, priv->est.enable is
only set once the hardware has been programmed successfully, and any
failure path disables EST before returning. This also fixes the stale
HW schedule left behind when a REPLACE fails validation, since priv->est
is no longer wiped before the new schedule is known to be valid.
Fixes: b60189e0392f ("net: stmmac: Integrate EST with TAPRIO scheduler API")
Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com>
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac_est.c | 30 ++++++
drivers/net/ethernet/stmicro/stmmac/stmmac_est.h | 13 +++
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 13 +++
drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c | 35 +------
drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c | 107 ++++++++++++----------
5 files changed, 120 insertions(+), 78 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
index f15d4d046aa7..49edfebbc39e 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
@@ -80,6 +80,36 @@ static int est_configure(struct stmmac_priv *priv, struct stmmac_est *cfg,
return 0;
}
+int __stmmac_setup_est(struct stmmac_priv *priv)
+{
+ struct timespec64 current_time, time;
+ ktime_t current_time_ns, basetime;
+ u64 cycle_time;
+ int err;
+
+ lockdep_assert_held(&priv->est_lock);
+
+ priv->ptp_clock_ops.gettime64(&priv->ptp_clock_ops, ¤t_time);
+ current_time_ns = timespec64_to_ktime(current_time);
+
+ time.tv_nsec = priv->est.btr_reserve[0];
+ time.tv_sec = priv->est.btr_reserve[1];
+ basetime = timespec64_to_ktime(time);
+
+ cycle_time = (u64)priv->est.ctr[1] * NSEC_PER_SEC + priv->est.ctr[0];
+
+ time = stmmac_calc_tas_basetime(basetime, current_time_ns, cycle_time);
+ priv->est.btr[0] = (u32)time.tv_nsec;
+ priv->est.btr[1] = (u32)time.tv_sec;
+
+ err = stmmac_est_configure(priv, priv, &priv->est,
+ priv->plat->clk_ptp_rate, true);
+ if (err)
+ netdev_err(priv->dev, "failed to re-configure EST\n");
+
+ return err;
+}
+
static void est_irq_status(struct stmmac_priv *priv, struct net_device *dev,
struct stmmac_extra_stats *x, u32 txqcnt)
{
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h
index f70221c9c84a..b4d1a1f04f10 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h
@@ -65,3 +65,16 @@
#define EST_GCL_DATA 0x00000034
extern const struct stmmac_est_ops dwmac510_est_ops;
+
+int __stmmac_setup_est(struct stmmac_priv *priv);
+static inline int stmmac_setup_est(struct stmmac_priv *priv)
+{
+ int ret = 0;
+
+ mutex_lock(&priv->est_lock);
+ if (priv->est.enable)
+ ret = __stmmac_setup_est(priv);
+ mutex_unlock(&priv->est_lock);
+
+ return ret;
+}
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index b233105cb449..4e80548ecfe9 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -48,6 +48,7 @@
#include "stmmac_ptp.h"
#include "stmmac_fpe.h"
#include "stmmac.h"
+#include "stmmac_est.h"
#include "stmmac_pcs.h"
#include "stmmac_xdp.h"
#include <linux/reset.h>
@@ -4200,6 +4201,13 @@ static int __stmmac_open(struct net_device *dev,
if (ret)
goto ptp_error;
+ /* The core soft reset in stmmac_hw_setup() clears the MTL_EST
+ * registers, so re-apply the taprio offload after PTP is up.
+ */
+ ret = stmmac_setup_est(priv);
+ if (ret < 0)
+ goto est_error;
+
stmmac_init_coalesce(priv);
phylink_start(priv->phylink);
@@ -4224,6 +4232,7 @@ static int __stmmac_open(struct net_device *dev,
for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
hrtimer_cancel(&priv->dma_conf.tx_queue[chan].txtimer);
+est_error:
stmmac_release_ptp(priv);
ptp_error:
stmmac_stop_all_dma(priv);
@@ -8420,6 +8429,10 @@ int stmmac_resume(struct device *dev)
}
init_coalesce:
+ ret = stmmac_setup_est(priv);
+ if (ret < 0)
+ goto error_stop_dma;
+
stmmac_init_coalesce(priv);
phylink_rx_clk_stop_block(priv->phylink);
stmmac_set_rx_mode(ndev);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
index 64d890664421..db5610f0fbab 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
@@ -8,6 +8,7 @@
Author: Rayagond Kokatanur <rayagond@vayavyalabs.com>
*******************************************************************************/
#include "stmmac.h"
+#include "stmmac_est.h"
#include "stmmac_ptp.h"
#define PTP_SAFE_TIME_OFFSET_NS 500000
@@ -55,7 +56,6 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
u32 quotient, reminder;
int neg_adj = 0;
bool xmac, est_rst = false;
- int ret;
xmac = dwmac_is_xmac(priv->plat->core_type);
@@ -69,46 +69,21 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
nsec = reminder;
/* If EST is enabled, disabled it before adjust ptp time. */
+ mutex_lock(&priv->est_lock);
if (priv->est.enable) {
est_rst = true;
- mutex_lock(&priv->est_lock);
- priv->est.enable = false;
stmmac_est_configure(priv, priv, &priv->est,
priv->plat->clk_ptp_rate, false);
- mutex_unlock(&priv->est_lock);
}
+ mutex_unlock(&priv->est_lock);
write_lock_irqsave(&priv->ptp_lock, flags);
stmmac_adjust_systime(priv, priv->ptpaddr, sec, nsec, neg_adj, xmac);
write_unlock_irqrestore(&priv->ptp_lock, flags);
/* Calculate new basetime and re-configured EST after PTP time adjust. */
- if (est_rst) {
- struct timespec64 current_time, time;
- ktime_t current_time_ns, basetime;
- u64 cycle_time;
-
- mutex_lock(&priv->est_lock);
- priv->ptp_clock_ops.gettime64(&priv->ptp_clock_ops, ¤t_time);
- current_time_ns = timespec64_to_ktime(current_time);
- time.tv_nsec = priv->est.btr_reserve[0];
- time.tv_sec = priv->est.btr_reserve[1];
- basetime = timespec64_to_ktime(time);
- cycle_time = (u64)priv->est.ctr[1] * NSEC_PER_SEC +
- priv->est.ctr[0];
- time = stmmac_calc_tas_basetime(basetime,
- current_time_ns,
- cycle_time);
-
- priv->est.btr[0] = (u32)time.tv_nsec;
- priv->est.btr[1] = (u32)time.tv_sec;
- priv->est.enable = true;
- ret = stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate, true);
- mutex_unlock(&priv->est_lock);
- if (ret)
- netdev_err(priv->dev, "failed to configure EST\n");
- }
+ if (est_rst)
+ stmmac_setup_est(priv);
return 0;
}
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
index 67c6fc32d0ea..70d623f7083b 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
@@ -10,6 +10,7 @@
#include "dwmac4.h"
#include "dwmac5.h"
#include "stmmac.h"
+#include "stmmac_est.h"
static void tc_fill_all_pass_entry(struct stmmac_tc_entry *entry)
{
@@ -968,10 +969,10 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
{
u32 size, wid = priv->dma_cap.estwid, dep = priv->dma_cap.estdep;
struct netlink_ext_ack *extack = qopt->mqprio.extack;
- struct timespec64 time, current_time, qopt_time;
- ktime_t current_time_ns;
- int err, i, ret = 0;
- u64 ctr;
+ u64 ctr = qopt->cycle_time;
+ struct timespec64 time;
+ u32 *gcl = NULL;
+ int i, ret = 0;
if (qopt->base_time < 0)
return -ERANGE;
@@ -979,6 +980,9 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
if (!priv->dma_cap.estsel)
return -EOPNOTSUPP;
+ if (ctr > (u64)U32_MAX * NSEC_PER_SEC)
+ return -ERANGE;
+
switch (wid) {
case 0x1:
wid = 16;
@@ -1013,35 +1017,45 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
return -EOPNOTSUPP;
}
+ mutex_lock(&priv->est_lock);
+
if (qopt->cmd == TAPRIO_CMD_DESTROY)
goto disable;
- if (qopt->num_entries > dep)
- return -EINVAL;
- if (!qopt->cycle_time)
- return -ERANGE;
- if (qopt->cycle_time_extension >= BIT(wid + 7))
- return -ERANGE;
+ if (qopt->num_entries > dep) {
+ ret = -EINVAL;
+ goto unlock;
+ }
- mutex_lock(&priv->est_lock);
- memset(&priv->est, 0, sizeof(priv->est));
- mutex_unlock(&priv->est_lock);
+ if (!qopt->cycle_time) {
+ ret = -ERANGE;
+ goto unlock;
+ }
- size = qopt->num_entries;
+ if (qopt->cycle_time_extension >= BIT(wid + 7)) {
+ ret = -ERANGE;
+ goto unlock;
+ }
- mutex_lock(&priv->est_lock);
- priv->est.gcl_size = size;
- priv->est.enable = qopt->cmd == TAPRIO_CMD_REPLACE;
- mutex_unlock(&priv->est_lock);
+ gcl = kzalloc(sizeof(*gcl) * EST_GCL, GFP_KERNEL);
+ if (!gcl) {
+ ret = -ENOMEM;
+ goto unlock;
+ }
+ size = qopt->num_entries;
for (i = 0; i < size; i++) {
s64 delta_ns = qopt->entries[i].interval;
u32 gates = qopt->entries[i].gate_mask;
- if (delta_ns > GENMASK(wid - 1, 0))
- return -ERANGE;
- if (gates > GENMASK(31 - wid, 0))
- return -ERANGE;
+ if (delta_ns > GENMASK(wid - 1, 0)) {
+ ret = -ERANGE;
+ goto free_gcl;
+ }
+ if (gates > GENMASK(31 - wid, 0)) {
+ ret = -ERANGE;
+ goto free_gcl;
+ }
switch (qopt->entries[i].command) {
case TC_TAPRIO_CMD_SET_GATES:
@@ -1053,51 +1067,44 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
gates &= ~BIT(0);
break;
default:
- return -EOPNOTSUPP;
+ ret = -EOPNOTSUPP;
+ goto free_gcl;
}
- priv->est.gcl[i] = delta_ns | (gates << wid);
+ gcl[i] = delta_ns | (gates << wid);
}
- mutex_lock(&priv->est_lock);
- /* Adjust for real system time */
- priv->ptp_clock_ops.gettime64(&priv->ptp_clock_ops, ¤t_time);
- current_time_ns = timespec64_to_ktime(current_time);
- time = stmmac_calc_tas_basetime(qopt->base_time, current_time_ns,
- qopt->cycle_time);
-
- priv->est.btr[0] = (u32)time.tv_nsec;
- priv->est.btr[1] = (u32)time.tv_sec;
+ memset(&priv->est, 0, sizeof(priv->est));
+ memcpy(priv->est.gcl, gcl, sizeof(priv->est.gcl));
+ priv->est.gcl_size = size;
- qopt_time = ktime_to_timespec64(qopt->base_time);
- priv->est.btr_reserve[0] = (u32)qopt_time.tv_nsec;
- priv->est.btr_reserve[1] = (u32)qopt_time.tv_sec;
+ time = ktime_to_timespec64(qopt->base_time);
+ priv->est.btr_reserve[0] = (u32)time.tv_nsec;
+ priv->est.btr_reserve[1] = (u32)time.tv_sec;
- ctr = qopt->cycle_time;
priv->est.ctr[0] = do_div(ctr, NSEC_PER_SEC);
priv->est.ctr[1] = (u32)ctr;
priv->est.ter = qopt->cycle_time_extension;
-
tc_taprio_map_maxsdu_txq(priv, qopt);
- ret = stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate, true);
- mutex_unlock(&priv->est_lock);
- if (ret) {
- netdev_err(priv->dev, "failed to configure EST\n");
+ ret = __stmmac_setup_est(priv);
+ if (ret)
goto disable;
- }
ret = stmmac_fpe_map_preemption_class(priv, priv->dev, extack,
qopt->mqprio.preemptible_tcs);
if (ret)
goto disable;
+ priv->est.enable = true;
+ kfree(gcl);
+
+ mutex_unlock(&priv->est_lock);
+
return 0;
disable:
- mutex_lock(&priv->est_lock);
priv->est.enable = false;
stmmac_est_configure(priv, priv, &priv->est,
priv->plat->clk_ptp_rate, false);
@@ -1107,11 +1114,15 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
priv->xstats.mtl_est_txq_hlbf[i] = 0;
priv->xstats.mtl_est_txq_hlbs[i] = 0;
}
+ i = stmmac_fpe_map_preemption_class(priv, priv->dev, extack, 0);
+ if (qopt->cmd == TAPRIO_CMD_DESTROY)
+ ret = i;
+free_gcl:
+ kfree(gcl);
+unlock:
mutex_unlock(&priv->est_lock);
- err = stmmac_fpe_map_preemption_class(priv, priv->dev, extack, 0);
-
- return qopt->cmd == TAPRIO_CMD_DESTROY ? err : ret;
+ return ret;
}
static void tc_taprio_stats(struct stmmac_priv *priv,
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 08/19] net: stmmac: serialize and retain PHC configuration across reset
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (6 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 07/19] net: stmmac: re-apply taprio offload in __stmmac_open() and stmmac_resume() James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 09/19] net: stmmac: leave the datapath running for normal-size MTU changes James Hilliard
` (12 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Keep the frequency correction, PEROUT requests and EXTTS selection
independently of timestamp register contents. Cache PPS requests before
converting units so they can be replayed after reset.
Serialize timestamp writers and devlink timestamp-mode updates with a
mutex, and let atomic gettime callers observe a reset gate under
ptp_lock. Provide a common replay helper for later reset transactions,
without changing PHC registration lifetime. Reapply the cached frequency
correction when recalculating the timestamp increment, including a
devlink change back to fine adjustment mode.
Hold est_lock across the whole PHC time adjustment, including EST
disable and replay. Return hardware errors instead of reporting a
successful clock step with EST left disabled. If the clock update
fails after disabling EST, still try to restore the installed schedule
and preserve the first error.
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
Changes in v5:
- Keep the programmed addend consistent with the cached frequency
correction when changing timestamp increment or adjustment mode.
- Propagate EST disable, clock-update and replay failures from adjtime,
keeping est_lock held throughout the transition.
---
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 7 +
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 23 ++-
drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c | 163 ++++++++++++++++++----
3 files changed, 164 insertions(+), 29 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index c65d95fc4760..23fd883d0509 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -340,6 +340,12 @@ struct stmmac_priv {
int use_riwt;
int irq_wake;
rwlock_t ptp_lock;
+ /* Serialize PHC changes with a hardware reset; gettime uses ptp_lock. */
+ struct mutex ptp_mutex;
+ bool ptp_blocked;
+ long ptp_scaled_ppm;
+ u32 ptp_perout;
+ u32 ptp_extts;
/* Protects auxiliary snapshot registers from concurrent access. */
struct mutex aux_ts_lock;
wait_queue_head_t tstamp_busy_wait;
@@ -407,6 +413,7 @@ void stmmac_set_ethtool_ops(struct net_device *netdev);
void stmmac_ptp_register(struct stmmac_priv *priv);
void stmmac_ptp_unregister(struct stmmac_priv *priv);
+int stmmac_ptp_restore(struct stmmac_priv *priv);
int stmmac_xdp_open(struct net_device *dev);
void stmmac_xdp_release(struct net_device *dev);
int stmmac_get_phy_intf_sel(phy_interface_t interface);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 4e80548ecfe9..a8e5e86e0ead 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -607,6 +607,7 @@ static void stmmac_update_subsecond_increment(struct stmmac_priv *priv)
bool xmac = dwmac_is_xmac(priv->plat->core_type);
u32 sec_inc = 0;
u64 temp = 0;
+ u32 addend;
stmmac_config_hw_tstamping(priv, priv->ptpaddr, priv->systime_flags);
@@ -626,7 +627,8 @@ static void stmmac_update_subsecond_increment(struct stmmac_priv *priv)
*/
temp = (u64)(temp << 32);
priv->default_addend = div_u64(temp, priv->plat->clk_ptp_rate);
- stmmac_config_addend(priv, priv->ptpaddr, priv->default_addend);
+ addend = adjust_by_scaled_ppm(priv->default_addend, priv->ptp_scaled_ppm);
+ stmmac_config_addend(priv, priv->ptpaddr, addend);
}
/**
@@ -954,6 +956,11 @@ static int stmmac_setup_ptp(struct stmmac_priv *priv)
return 0;
}
+ priv->ptp_scaled_ppm = 0;
+ priv->ptp_perout = 0;
+ priv->ptp_extts = 0;
+ priv->ptp_blocked = false;
+
ret = clk_prepare_enable(priv->plat->clk_ptp_ref);
if (ret < 0) {
netdev_warn(priv->dev,
@@ -6329,7 +6336,8 @@ static void stmmac_common_interrupt(struct stmmac_priv *priv)
for (queue = 0; queue < queues_count; queue++)
stmmac_host_mtl_irq_status(priv, priv->hw, queue);
- stmmac_timestamp_interrupt(priv, priv);
+ if (!READ_ONCE(priv->ptp_blocked))
+ stmmac_timestamp_interrupt(priv, priv);
}
}
@@ -7762,6 +7770,14 @@ static int stmmac_dl_ts_coarse_set(struct devlink *dl, u32 id,
{
struct stmmac_devlink_priv *dl_priv = devlink_priv(dl);
struct stmmac_priv *priv = dl_priv->stmmac_priv;
+ unsigned long flags;
+
+ mutex_lock(&priv->ptp_mutex);
+ if (priv->ptp_blocked) {
+ mutex_unlock(&priv->ptp_mutex);
+ return -EBUSY;
+ }
+ write_lock_irqsave(&priv->ptp_lock, flags);
priv->tsfupdt_coarse = ctx->val.vbool;
@@ -7774,6 +7790,8 @@ static int stmmac_dl_ts_coarse_set(struct devlink *dl, u32 id,
* reconfigure the systime, subsecond increment and addend.
*/
stmmac_update_subsecond_increment(priv);
+ write_unlock_irqrestore(&priv->ptp_lock, flags);
+ mutex_unlock(&priv->ptp_mutex);
return 0;
}
@@ -8114,6 +8132,7 @@ static int __stmmac_dvr_probe(struct device *device,
mutex_init(&priv->lock);
mutex_init(&priv->est_lock);
rwlock_init(&priv->ptp_lock);
+ mutex_init(&priv->ptp_mutex);
stmmac_fpe_init(priv);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
index db5610f0fbab..b34711231227 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
@@ -13,6 +13,16 @@
#define PTP_SAFE_TIME_OFFSET_NS 500000
+static int stmmac_ptp_begin(struct stmmac_priv *priv)
+{
+ mutex_lock(&priv->ptp_mutex);
+ if (priv->ptp_blocked) {
+ mutex_unlock(&priv->ptp_mutex);
+ return -EBUSY;
+ }
+ return 0;
+}
+
/**
* stmmac_adjust_freq
*
@@ -29,14 +39,21 @@ static int stmmac_adjust_freq(struct ptp_clock_info *ptp, long scaled_ppm)
container_of(ptp, struct stmmac_priv, ptp_clock_ops);
unsigned long flags;
u32 addend;
+ int ret;
- addend = adjust_by_scaled_ppm(priv->default_addend, scaled_ppm);
+ ret = stmmac_ptp_begin(priv);
+ if (ret)
+ return ret;
write_lock_irqsave(&priv->ptp_lock, flags);
- stmmac_config_addend(priv, priv->ptpaddr, addend);
+ addend = adjust_by_scaled_ppm(priv->default_addend, scaled_ppm);
+ ret = stmmac_config_addend(priv, priv->ptpaddr, addend);
+ if (!ret)
+ priv->ptp_scaled_ppm = scaled_ppm;
write_unlock_irqrestore(&priv->ptp_lock, flags);
+ mutex_unlock(&priv->ptp_mutex);
- return 0;
+ return ret;
}
/**
@@ -55,7 +72,12 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
u32 sec, nsec;
u32 quotient, reminder;
int neg_adj = 0;
- bool xmac, est_rst = false;
+ bool xmac;
+ int ret, err;
+
+ ret = stmmac_ptp_begin(priv);
+ if (ret)
+ return ret;
xmac = dwmac_is_xmac(priv->plat->core_type);
@@ -68,24 +90,32 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
sec = quotient;
nsec = reminder;
- /* If EST is enabled, disabled it before adjust ptp time. */
+ /* Keep the installed schedule stable across the entire clock step. */
mutex_lock(&priv->est_lock);
if (priv->est.enable) {
- est_rst = true;
- stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate, false);
+ ret = stmmac_est_configure(priv, priv, &priv->est,
+ priv->plat->clk_ptp_rate, false);
+ if (ret)
+ goto out_unlock;
}
- mutex_unlock(&priv->est_lock);
write_lock_irqsave(&priv->ptp_lock, flags);
- stmmac_adjust_systime(priv, priv->ptpaddr, sec, nsec, neg_adj, xmac);
+ ret = stmmac_adjust_systime(priv, priv->ptpaddr, sec, nsec, neg_adj, xmac);
write_unlock_irqrestore(&priv->ptp_lock, flags);
- /* Calculate new basetime and re-configured EST after PTP time adjust. */
- if (est_rst)
- stmmac_setup_est(priv);
+ /* Also try to restore EST after a failed clock update. Keep the first
+ * error, but do not report success if only schedule replay failed.
+ */
+ if (priv->est.enable) {
+ err = __stmmac_setup_est(priv);
+ if (!ret)
+ ret = err;
+ }
- return 0;
+out_unlock:
+ mutex_unlock(&priv->est_lock);
+ mutex_unlock(&priv->ptp_mutex);
+ return ret;
}
/**
@@ -103,14 +133,18 @@ static int stmmac_get_time(struct ptp_clock_info *ptp, struct timespec64 *ts)
container_of(ptp, struct stmmac_priv, ptp_clock_ops);
unsigned long flags;
u64 ns = 0;
+ int ret = 0;
read_lock_irqsave(&priv->ptp_lock, flags);
- stmmac_get_systime(priv, priv->ptpaddr, &ns);
+ if (priv->ptp_blocked)
+ ret = -EBUSY;
+ else
+ stmmac_get_systime(priv, priv->ptpaddr, &ns);
read_unlock_irqrestore(&priv->ptp_lock, flags);
*ts = ns_to_timespec64(ns);
- return 0;
+ return ret;
}
/**
@@ -128,25 +162,34 @@ static int stmmac_set_time(struct ptp_clock_info *ptp,
struct stmmac_priv *priv =
container_of(ptp, struct stmmac_priv, ptp_clock_ops);
unsigned long flags;
+ int ret;
+
+ ret = stmmac_ptp_begin(priv);
+ if (ret)
+ return ret;
write_lock_irqsave(&priv->ptp_lock, flags);
- stmmac_init_systime(priv, priv->ptpaddr, ts->tv_sec, ts->tv_nsec);
+ ret = stmmac_init_systime(priv, priv->ptpaddr, ts->tv_sec, ts->tv_nsec);
write_unlock_irqrestore(&priv->ptp_lock, flags);
+ mutex_unlock(&priv->ptp_mutex);
- return 0;
+ return ret;
}
-static int stmmac_enable(struct ptp_clock_info *ptp,
- struct ptp_clock_request *rq, int on)
+static int __stmmac_enable(struct ptp_clock_info *ptp,
+ struct ptp_clock_request *rq, int on)
{
struct stmmac_priv *priv =
container_of(ptp, struct stmmac_priv, ptp_clock_ops);
void __iomem *ptpaddr = priv->ptpaddr;
- struct stmmac_pps_cfg *cfg;
+ struct stmmac_pps_cfg pps, saved, *cfg = &pps;
int ret = -EOPNOTSUPP;
unsigned long flags;
u32 acr_value;
+ if (priv->plat->core_type == DWMAC_CORE_GMAC)
+ return dwmac1000_ptp_enable(ptp, rq, on);
+
switch (rq->type) {
case PTP_CLK_REQ_PEROUT: {
struct timespec64 curr_time;
@@ -157,8 +200,6 @@ static int stmmac_enable(struct ptp_clock_info *ptp,
if (rq->perout.flags)
return -EOPNOTSUPP;
- cfg = &priv->pps[rq->perout.index];
-
cfg->start.tv_sec = rq->perout.start.sec;
cfg->start.tv_nsec = rq->perout.start.nsec;
@@ -188,6 +229,7 @@ static int stmmac_enable(struct ptp_clock_info *ptp,
cfg->period.tv_sec = rq->perout.period.sec;
cfg->period.tv_nsec = rq->perout.period.nsec;
+ saved = *cfg;
write_lock_irqsave(&priv->ptp_lock, flags);
ret = stmmac_flex_pps_config(priv, priv->ioaddr,
@@ -195,6 +237,9 @@ static int stmmac_enable(struct ptp_clock_info *ptp,
priv->sub_second_inc,
priv->systime_flags);
write_unlock_irqrestore(&priv->ptp_lock, flags);
+ /* Some cores convert cfg->start to binary rollover units. */
+ if (!ret)
+ priv->pps[rq->perout.index] = saved;
break;
}
case PTP_CLK_REQ_EXTTS: {
@@ -240,6 +285,64 @@ static int stmmac_enable(struct ptp_clock_info *ptp,
return ret;
}
+static int stmmac_enable(struct ptp_clock_info *ptp,
+ struct ptp_clock_request *rq, int on)
+{
+ struct stmmac_priv *priv =
+ container_of(ptp, struct stmmac_priv, ptp_clock_ops);
+ int ret;
+
+ ret = stmmac_ptp_begin(priv);
+ if (ret)
+ return ret;
+ ret = __stmmac_enable(ptp, rq, on);
+ if (!ret) {
+ if (rq->type == PTP_CLK_REQ_PEROUT) {
+ if (on)
+ priv->ptp_perout |= BIT(rq->perout.index);
+ else
+ priv->ptp_perout &= ~BIT(rq->perout.index);
+ } else if (rq->type == PTP_CLK_REQ_EXTTS) {
+ priv->ptp_extts = on ? BIT(rq->extts.index) : 0;
+ }
+ }
+ mutex_unlock(&priv->ptp_mutex);
+ return ret;
+}
+
+/* Called with ptp_mutex held and PHC access blocked across the MAC reset. */
+int stmmac_ptp_restore(struct stmmac_priv *priv)
+{
+ struct ptp_clock_request rq = { .type = PTP_CLK_REQ_EXTTS };
+ unsigned long flags;
+ u64 ns = 0, period;
+ u32 addend;
+ int i, ret;
+
+ write_lock_irqsave(&priv->ptp_lock, flags);
+ addend = adjust_by_scaled_ppm(priv->default_addend, priv->ptp_scaled_ppm);
+ ret = stmmac_config_addend(priv, priv->ptpaddr, addend);
+ for (i = 0; !ret && i < STMMAC_PPS_MAX; i++) {
+ struct stmmac_pps_cfg cfg = priv->pps[i];
+
+ if (!(priv->ptp_perout & BIT(i)))
+ continue;
+ stmmac_get_systime(priv, priv->ptpaddr, &ns);
+ period = timespec64_to_ns(&cfg.period);
+ /* Retain phase, but move an expired target into the future. */
+ cfg.start = stmmac_calc_tas_basetime(timespec64_to_ktime(cfg.start),
+ ns + PTP_SAFE_TIME_OFFSET_NS, period);
+ ret = stmmac_flex_pps_config(priv, priv->ioaddr, i, &cfg, true,
+ priv->sub_second_inc, priv->systime_flags);
+ }
+ write_unlock_irqrestore(&priv->ptp_lock, flags);
+ if (ret || !priv->ptp_extts)
+ return ret;
+
+ rq.extts.index = __ffs(priv->ptp_extts);
+ return __stmmac_enable(&priv->ptp_clock_ops, &rq, 1);
+}
+
/**
* stmmac_get_syncdevicetime
* @device: current device time
@@ -262,9 +365,15 @@ static int stmmac_getcrosststamp(struct ptp_clock_info *ptp,
{
struct stmmac_priv *priv =
container_of(ptp, struct stmmac_priv, ptp_clock_ops);
-
- return get_device_system_crosststamp(stmmac_get_syncdevicetime,
- priv, NULL, xtstamp);
+ int ret;
+
+ ret = stmmac_ptp_begin(priv);
+ if (ret)
+ return ret;
+ ret = get_device_system_crosststamp(stmmac_get_syncdevicetime,
+ priv, NULL, xtstamp);
+ mutex_unlock(&priv->ptp_mutex);
+ return ret;
}
/* structure describing a PTP hardware clock */
@@ -298,7 +407,7 @@ const struct ptp_clock_info dwmac1000_ptp_clock_ops = {
.adjtime = stmmac_adjust_time,
.gettime64 = stmmac_get_time,
.settime64 = stmmac_set_time,
- .enable = dwmac1000_ptp_enable,
+ .enable = stmmac_enable,
};
/**
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 09/19] net: stmmac: leave the datapath running for normal-size MTU changes
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (7 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 08/19] net: stmmac: serialize and retain PHC configuration across reset James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 10/19] net: stmmac: fix error path cleanup in DMA descriptor ring allocation James Hilliard
` (11 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Changing an MTU at or below ETH_DATA_LEN does not change the receive
buffer size or the MAC receive limit when the previous MTU was also in
that range. Do not release and reopen the datapath for those changes.
Besides avoiding unnecessary hardware resets and their failure paths,
this keeps live AF_XDP pool bindings intact. Preparing replacement rings
before stopping the old rings otherwise binds the pool to a temporary
RXQ and consumes fill-ring entries while its current RXQ is still
active. XDP already rejects jumbo MTUs, so all supported live XDP MTU
changes can use this path without preparing replacement rings.
Keep jumbo transitions on the existing reinitialization path for now.
The later ownership and rollback changes address that path separately.
Fixes: 3470079687448 ("net: ethernet: stmicro: stmmac: permit MTU change with interface up")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index a8e5e86e0ead..d9d676ae1c82 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -6212,7 +6212,12 @@ static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
if ((txfifosz < new_mtu) || (new_mtu > BUF_SIZE_16KiB))
return -EINVAL;
- if (netif_running(dev)) {
+ /* Normal-size frames use the same buffers and MAC receive limits.
+ * In particular, do not disturb a live AF_XDP pool: XDP does not
+ * support jumbo frames, so it never needs the ring replacement below.
+ */
+ if (netif_running(dev) &&
+ (dev->mtu > ETH_DATA_LEN || mtu > ETH_DATA_LEN)) {
netdev_dbg(priv->dev, "restarting interface to change its MTU\n");
/* Try to allocate the new DMA conf with the new mtu */
dma_conf = stmmac_setup_dma_desc(priv, mtu);
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 10/19] net: stmmac: fix error path cleanup in DMA descriptor ring allocation
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (8 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 09/19] net: stmmac: leave the datapath running for normal-size MTU changes James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 11/19] net: stmmac: complete DMA configuration allocation unwind James Hilliard
` (10 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard, Ding Hui
From: Ding Hui <dinghui@lixiang.com>
__alloc_dma_rx_desc_resources() and __alloc_dma_tx_desc_resources()
allocate resources in multiple steps but return early on failure
without cleaning up what they have already allocated. The outer
error paths then call the free helpers on partially-initialized
queues, which dereference pointers that were never allocated:
- dma_free_rx_skbufs() and dma_free_rx_xskbufs() dereference
rx_q->buf_pool[i] via stmmac_free_rx_buffer(), but buf_pool
may be NULL if its kzalloc_objs() failed.
- dma_free_tx_skbufs() dereferences tx_q->tx_skbuff_dma[i] via
stmmac_free_tx_buffer(), but tx_skbuff_dma may be NULL if its
kzalloc_objs() failed.
- stmmac_free_tx_buffer() dereferences tx_q->xdpf[i] and
tx_q->tx_skbuff[i] (aliased through a union), but tx_skbuff
may be NULL if its allocation failed while tx_skbuff_dma
succeeded.
Fix this by making each allocation function responsible for undoing
its own allocations on error, following the standard kernel error
handling pattern of cleaning up in reverse order. Also add NULL
checks in the free helpers as a defensive measure, since they may
be called on partially-initialized queues.
Additionally, make __free_dma_rx_desc_resources() and
__free_dma_tx_desc_resources() clear the pointers they free, so that
the NULL guards in the free helpers hold reliably when the long-lived
priv->dma_conf is reused across XDP open/release cycles.
Signed-off-by: Ding Hui <dinghui@lixiang.com>
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 70 +++++++++++++++++++----
1 file changed, 60 insertions(+), 10 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index d9d676ae1c82..258de45d122c 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -1767,7 +1767,7 @@ static void stmmac_free_tx_buffer(struct stmmac_priv *priv,
DMA_TO_DEVICE);
}
- if (tx_q->xdpf[i] &&
+ if (tx_q->xdpf && tx_q->xdpf[i] &&
(tx_q->tx_skbuff_dma[i].buf_type == STMMAC_TXBUF_T_XDP_TX ||
tx_q->tx_skbuff_dma[i].buf_type == STMMAC_TXBUF_T_XDP_NDO)) {
xdp_return_frame(tx_q->xdpf[i]);
@@ -1777,7 +1777,7 @@ static void stmmac_free_tx_buffer(struct stmmac_priv *priv,
if (tx_q->tx_skbuff_dma[i].buf_type == STMMAC_TXBUF_T_XSK_TX)
tx_q->xsk_frames_done++;
- if (tx_q->tx_skbuff[i] &&
+ if (tx_q->tx_skbuff && tx_q->tx_skbuff[i] &&
tx_q->tx_skbuff_dma[i].buf_type == STMMAC_TXBUF_T_SKB) {
dev_kfree_skb_any(tx_q->tx_skbuff[i]);
tx_q->tx_skbuff[i] = NULL;
@@ -1800,6 +1800,10 @@ static void dma_free_rx_skbufs(struct stmmac_priv *priv,
struct stmmac_rx_queue *rx_q = &dma_conf->rx_queue[queue];
int i;
+ /* buf_pool may not be allocated if alloc failed early */
+ if (!rx_q->buf_pool)
+ return;
+
for (i = 0; i < dma_conf->dma_rx_size; i++)
stmmac_free_rx_buffer(priv, rx_q, i);
}
@@ -1841,6 +1845,10 @@ static void dma_free_rx_xskbufs(struct stmmac_priv *priv,
struct stmmac_rx_queue *rx_q = &dma_conf->rx_queue[queue];
int i;
+ /* buf_pool may not be allocated if alloc failed early */
+ if (!rx_q->buf_pool)
+ return;
+
for (i = 0; i < dma_conf->dma_rx_size; i++) {
struct stmmac_rx_buffer *buf = &rx_q->buf_pool[i];
@@ -2136,6 +2144,10 @@ static void dma_free_tx_skbufs(struct stmmac_priv *priv,
struct stmmac_tx_queue *tx_q = &dma_conf->tx_queue[queue];
int i;
+ /* tx_skbuff_dma may not be allocated if alloc failed early */
+ if (!tx_q->tx_skbuff_dma)
+ return;
+
tx_q->xsk_frames_done = 0;
for (i = 0; i < dma_conf->dma_tx_size; i++)
@@ -2193,13 +2205,20 @@ static void __free_dma_rx_desc_resources(struct stmmac_priv *priv,
size = stmmac_get_rx_desc_size(priv) * dma_conf->dma_rx_size;
dma_free_coherent(priv->device, size, addr, rx_q->dma_rx_phy);
+ rx_q->dma_erx = NULL;
+ rx_q->dma_rx = NULL;
+ rx_q->dma_rx_phy = 0;
if (xdp_rxq_info_is_reg(&rx_q->xdp_rxq))
xdp_rxq_info_unreg(&rx_q->xdp_rxq);
kfree(rx_q->buf_pool);
- if (rx_q->page_pool)
+ rx_q->buf_pool = NULL;
+
+ if (rx_q->page_pool) {
page_pool_destroy(rx_q->page_pool);
+ rx_q->page_pool = NULL;
+ }
}
static void free_dma_rx_desc_resources(struct stmmac_priv *priv,
@@ -2241,9 +2260,16 @@ static void __free_dma_tx_desc_resources(struct stmmac_priv *priv,
size = stmmac_get_tx_desc_size(priv, tx_q) * dma_conf->dma_tx_size;
dma_free_coherent(priv->device, size, addr, tx_q->dma_tx_phy);
+ tx_q->dma_etx = NULL;
+ tx_q->dma_entx = NULL;
+ tx_q->dma_tx = NULL;
+ tx_q->dma_tx_phy = 0;
kfree(tx_q->tx_skbuff_dma);
+ tx_q->tx_skbuff_dma = NULL;
+
kfree(tx_q->tx_skbuff);
+ tx_q->tx_skbuff = NULL;
}
static void free_dma_tx_desc_resources(struct stmmac_priv *priv,
@@ -2311,15 +2337,19 @@ static int __alloc_dma_rx_desc_resources(struct stmmac_priv *priv,
}
rx_q->buf_pool = kzalloc_objs(*rx_q->buf_pool, dma_conf->dma_rx_size);
- if (!rx_q->buf_pool)
- return -ENOMEM;
+ if (!rx_q->buf_pool) {
+ ret = -ENOMEM;
+ goto err_destroy_pool;
+ }
size = stmmac_get_rx_desc_size(priv) * dma_conf->dma_rx_size;
addr = dma_alloc_coherent(priv->device, size, &rx_q->dma_rx_phy,
GFP_KERNEL);
- if (!addr)
- return -ENOMEM;
+ if (!addr) {
+ ret = -ENOMEM;
+ goto err_free_buf_pool;
+ }
if (priv->extend_desc)
rx_q->dma_erx = addr;
@@ -2335,10 +2365,22 @@ static int __alloc_dma_rx_desc_resources(struct stmmac_priv *priv,
ret = xdp_rxq_info_reg(&rx_q->xdp_rxq, priv->dev, queue, napi_id);
if (ret) {
netdev_err(priv->dev, "Failed to register xdp rxq info\n");
- return -EINVAL;
+ goto err_free_dma;
}
return 0;
+
+err_free_dma:
+ dma_free_coherent(priv->device, size, addr, rx_q->dma_rx_phy);
+ rx_q->dma_erx = NULL;
+ rx_q->dma_rx = NULL;
+err_free_buf_pool:
+ kfree(rx_q->buf_pool);
+ rx_q->buf_pool = NULL;
+err_destroy_pool:
+ page_pool_destroy(rx_q->page_pool);
+ rx_q->page_pool = NULL;
+ return ret;
}
static int alloc_dma_rx_desc_resources(struct stmmac_priv *priv,
@@ -2391,14 +2433,14 @@ static int __alloc_dma_tx_desc_resources(struct stmmac_priv *priv,
tx_q->tx_skbuff = kzalloc_objs(struct sk_buff *, dma_conf->dma_tx_size);
if (!tx_q->tx_skbuff)
- return -ENOMEM;
+ goto err_free_skbuff_dma;
size = stmmac_get_tx_desc_size(priv, tx_q) * dma_conf->dma_tx_size;
addr = dma_alloc_coherent(priv->device, size,
&tx_q->dma_tx_phy, GFP_KERNEL);
if (!addr)
- return -ENOMEM;
+ goto err_free_skbuff;
if (priv->extend_desc)
tx_q->dma_etx = addr;
@@ -2408,6 +2450,14 @@ static int __alloc_dma_tx_desc_resources(struct stmmac_priv *priv,
tx_q->dma_tx = addr;
return 0;
+
+err_free_skbuff:
+ kfree(tx_q->tx_skbuff);
+ tx_q->tx_skbuff = NULL;
+err_free_skbuff_dma:
+ kfree(tx_q->tx_skbuff_dma);
+ tx_q->tx_skbuff_dma = NULL;
+ return -ENOMEM;
}
static int alloc_dma_tx_desc_resources(struct stmmac_priv *priv,
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 11/19] net: stmmac: complete DMA configuration allocation unwind
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (9 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 10/19] net: stmmac: fix error path cleanup in DMA descriptor ring allocation James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 12/19] net: stmmac: keep DMA configurations at stable addresses James Hilliard
` (9 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Build on Ding Hui's per-queue allocation cleanup. The combined RX/TX
allocator still needs to release the successful RX allocation if TX
allocation fails. Guard coherent frees after the per-queue unwind has
already emptied a queue.
Propagate RXQ memory-model registration errors instead of continuing
with an unusable RXQ. Clear XSK RXQ bindings before the RXQ goes away and
release any saved partial packet. The MTU transaction added later relies
on preparation failures being fully unwound without touching the active
configuration.
Take ownership of saved partial RX state at poll entry by clearing the
saved flag and skb pointer immediately. Preserve incomplete state if the
next descriptor is still DMA-owned. A budget-one completion must not
leave an already delivered or freed skb reachable by the new teardown
cleanup.
Fixes: 71fedb0198cb ("net: stmmac: break some functions into RX and TX scopes")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 54 ++++++++++++++++-------
1 file changed, 37 insertions(+), 17 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 258de45d122c..bc19f8c19bb8 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -1931,17 +1931,19 @@ static int __init_dma_rx_desc_rings(struct stmmac_priv *priv,
rx_q->xsk_pool = stmmac_get_xsk_pool(priv, queue);
if (rx_q->xsk_pool) {
- WARN_ON(xdp_rxq_info_reg_mem_model(&rx_q->xdp_rxq,
- MEM_TYPE_XSK_BUFF_POOL,
- NULL));
+ ret = xdp_rxq_info_reg_mem_model(&rx_q->xdp_rxq,
+ MEM_TYPE_XSK_BUFF_POOL, NULL);
+ if (ret)
+ return ret;
netdev_info(priv->dev,
"Register MEM_TYPE_XSK_BUFF_POOL RxQ-%d\n",
queue);
xsk_pool_set_rxq_info(rx_q->xsk_pool, &rx_q->xdp_rxq);
} else {
- WARN_ON(xdp_rxq_info_reg_mem_model(&rx_q->xdp_rxq,
- MEM_TYPE_PAGE_POOL,
- rx_q->page_pool));
+ ret = xdp_rxq_info_reg_mem_model(&rx_q->xdp_rxq,
+ MEM_TYPE_PAGE_POOL, rx_q->page_pool);
+ if (ret)
+ return ret;
netdev_info(priv->dev,
"Register MEM_TYPE_PAGE_POOL RxQ-%d\n",
queue);
@@ -2003,6 +2005,8 @@ static int init_dma_rx_desc_rings(struct net_device *dev,
dma_free_rx_skbufs(priv, dma_conf, queue);
rx_q->buf_alloc_num = 0;
+ if (rx_q->xsk_pool)
+ xsk_pool_set_rxq_info(rx_q->xsk_pool, NULL);
rx_q->xsk_pool = NULL;
queue--;
@@ -2188,10 +2192,16 @@ static void __free_dma_rx_desc_resources(struct stmmac_priv *priv,
void *addr;
/* Release the DMA RX socket buffers */
- if (rx_q->xsk_pool)
+ if (rx_q->xsk_pool) {
dma_free_rx_xskbufs(priv, dma_conf, queue);
- else
+ xsk_pool_set_rxq_info(rx_q->xsk_pool, NULL);
+ } else {
dma_free_rx_skbufs(priv, dma_conf, queue);
+ }
+ if (rx_q->state_saved)
+ dev_kfree_skb_any(rx_q->state.skb);
+ rx_q->state.skb = NULL;
+ rx_q->state_saved = 0;
rx_q->buf_alloc_num = 0;
rx_q->xsk_pool = NULL;
@@ -2204,7 +2214,8 @@ static void __free_dma_rx_desc_resources(struct stmmac_priv *priv,
size = stmmac_get_rx_desc_size(priv) * dma_conf->dma_rx_size;
- dma_free_coherent(priv->device, size, addr, rx_q->dma_rx_phy);
+ if (addr)
+ dma_free_coherent(priv->device, size, addr, rx_q->dma_rx_phy);
rx_q->dma_erx = NULL;
rx_q->dma_rx = NULL;
rx_q->dma_rx_phy = 0;
@@ -2259,7 +2270,8 @@ static void __free_dma_tx_desc_resources(struct stmmac_priv *priv,
size = stmmac_get_tx_desc_size(priv, tx_q) * dma_conf->dma_tx_size;
- dma_free_coherent(priv->device, size, addr, tx_q->dma_tx_phy);
+ if (addr)
+ dma_free_coherent(priv->device, size, addr, tx_q->dma_tx_phy);
tx_q->dma_etx = NULL;
tx_q->dma_entx = NULL;
tx_q->dma_tx = NULL;
@@ -2500,6 +2512,8 @@ static int alloc_dma_desc_resources(struct stmmac_priv *priv,
return ret;
ret = alloc_dma_tx_desc_resources(priv, dma_conf);
+ if (ret)
+ free_dma_rx_desc_resources(priv, dma_conf);
return ret;
}
@@ -5822,6 +5836,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
struct sk_buff *skb = NULL;
struct stmmac_xdp_buff ctx;
int xdp_status = 0;
+ bool in_progress = rx_q->state_saved;
int bufsz;
dma_dir = page_pool_get_dma_dir(rx_q->page_pool);
@@ -5836,6 +5851,14 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
stmmac_display_ring(priv, rx_head, priv->dma_conf.dma_rx_size, true,
rx_q->dma_rx_phy, desc_size);
}
+ if (in_progress) {
+ skb = rx_q->state.skb;
+ error = rx_q->state.error;
+ len = rx_q->state.len;
+ rx_q->state.skb = NULL;
+ rx_q->state_saved = false;
+ }
+
while (count < limit) {
unsigned int buf1_len = 0, buf2_len = 0;
enum pkt_hash_types hash_type;
@@ -5844,12 +5867,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
int entry;
u32 hash;
- if (!count && rx_q->state_saved) {
- skb = rx_q->state.skb;
- error = rx_q->state.error;
- len = rx_q->state.len;
- } else {
- rx_q->state_saved = false;
+ if (!in_progress) {
skb = NULL;
error = 0;
len = 0;
@@ -5883,6 +5901,8 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
prefetch(np);
+ in_progress = status & rx_not_ls;
+
if (priv->extend_desc)
stmmac_rx_extended_status(priv, &priv->xstats, rx_q->dma_erx + entry);
if (unlikely(status == discard_frame)) {
@@ -6057,7 +6077,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
count++;
}
- if (status & rx_not_ls || skb) {
+ if (in_progress || skb) {
rx_q->state_saved = true;
rx_q->state.skb = skb;
rx_q->state.error = error;
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 12/19] net: stmmac: keep DMA configurations at stable addresses
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (10 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 11/19] net: stmmac: complete DMA configuration allocation unwind James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 13/19] net: stmmac: track datapath and power ownership across failed reopening James Hilliard
` (8 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
The allocated DMA configuration contains RXQ metadata registered with
XDP and referenced by AF_XDP pools. Copying the configuration into priv
and freeing its original allocation leaves those references pointing at
the old storage. Retain the allocated object instead and keep a pointer
in priv, preserving ring sizes and per-queue settings while down.
Use persistent channel objects for per-queue IRQ contexts instead of
deriving priv from an embedded DMA configuration. XSK wakeup must also
avoid accessing replaceable queue objects. Drain transmitters and NAPI
poll tails before cancelling TX timers so a late rearm cannot outlive
the configuration containing the timer.
Update both open callers with the ownership change. Successful open
keeps the new configuration and frees the empty old object. Failed open
restores the old pointer before freeing the failed replacement. Detach
around the existing MTU reopen so XDP transmit cannot enter while that
pointer is being replaced, and reattach after success. The later MTU
transaction replaces this reopen path with retained-resource rollback.
The empty configuration remains allocated while down because ethtool, TC
and the next open still use its ring sizes and per-queue settings. The
probe-managed action frees the current object after netdev teardown.
Cancel the software EEE timer after the poll/transmit drain as well: a
TX completion may have passed its enable check before phylink cancelled
it.
GSO feature checks are not excluded by stopped queues or netdev detach.
Read the immutable platform TBS capability instead of dereferencing the
replaceable configuration from ndo_features_check().
Fixes: ba39b344e924 ("net: ethernet: stmicro: stmmac: generate stmmac dma conf before open")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/chain_mode.c | 6 +-
drivers/net/ethernet/stmicro/stmmac/ring_mode.c | 4 +-
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 2 +-
.../net/ethernet/stmicro/stmmac/stmmac_ethtool.c | 4 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 324 +++++++++++----------
.../net/ethernet/stmicro/stmmac/stmmac_selftests.c | 8 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c | 6 +-
7 files changed, 189 insertions(+), 165 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/chain_mode.c b/drivers/net/ethernet/stmicro/stmmac/chain_mode.c
index 66025e2509e9..65243c5e539e 100644
--- a/drivers/net/ethernet/stmicro/stmmac/chain_mode.c
+++ b/drivers/net/ethernet/stmicro/stmmac/chain_mode.c
@@ -48,7 +48,7 @@ static int jumbo_frm(struct stmmac_tx_queue *tx_q, struct sk_buff *skb,
while (len != 0) {
tx_q->tx_skbuff[entry] = NULL;
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_tx_size);
desc = tx_q->dma_tx + entry;
if (len > bmax) {
@@ -137,7 +137,7 @@ static void refill_desc3(struct stmmac_rx_queue *rx_q, struct dma_desc *p)
*/
p->des3 = cpu_to_le32((unsigned int)(rx_q->dma_rx_phy +
(((rx_q->dirty_rx) + 1) %
- priv->dma_conf.dma_rx_size) *
+ priv->dma_conf->dma_rx_size) *
sizeof(struct dma_desc)));
}
@@ -154,7 +154,7 @@ static void clean_desc3(struct stmmac_tx_queue *tx_q, struct dma_desc *p)
*/
p->des3 = cpu_to_le32((unsigned int)((tx_q->dma_tx_phy +
((tx_q->dirty_tx + 1) %
- priv->dma_conf.dma_tx_size))
+ priv->dma_conf->dma_tx_size))
* sizeof(struct dma_desc)));
}
diff --git a/drivers/net/ethernet/stmicro/stmmac/ring_mode.c b/drivers/net/ethernet/stmicro/stmmac/ring_mode.c
index d2f0c321661d..0299d6a6c32b 100644
--- a/drivers/net/ethernet/stmicro/stmmac/ring_mode.c
+++ b/drivers/net/ethernet/stmicro/stmmac/ring_mode.c
@@ -52,7 +52,7 @@ static int jumbo_frm(struct stmmac_tx_queue *tx_q, struct sk_buff *skb,
stmmac_prepare_tx_desc(priv, desc, 1, bmax, csum,
STMMAC_RING_MODE, 0, false, skb->len);
tx_q->tx_skbuff[entry] = NULL;
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_tx_size);
if (priv->extend_desc)
desc = (struct dma_desc *)(tx_q->dma_etx + entry);
@@ -102,7 +102,7 @@ static void refill_desc3(struct stmmac_rx_queue *rx_q, struct dma_desc *p)
struct stmmac_priv *priv = rx_q->priv_data;
/* Fill DES3 in case of RING mode */
- if (priv->dma_conf.dma_buf_sz == BUF_SIZE_16KiB)
+ if (priv->dma_conf->dma_buf_sz == BUF_SIZE_16KiB)
p->des3 = cpu_to_le32(le32_to_cpu(p->des2) + BUF_SIZE_8KiB);
}
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 23fd883d0509..fa01070cbef3 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -280,7 +280,7 @@ struct stmmac_priv {
int (*hwif_quirks)(struct stmmac_priv *priv);
struct mutex lock;
- struct stmmac_dma_conf dma_conf;
+ struct stmmac_dma_conf *dma_conf;
/* Generic channel for NAPI */
struct stmmac_channel channel[STMMAC_CH_MAX];
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
index 1cf0f8820b33..1fdd63d6bba1 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
@@ -396,8 +396,8 @@ static void stmmac_get_ringparam(struct net_device *netdev,
ring->rx_max_pending = DMA_MAX_RX_SIZE;
ring->tx_max_pending = DMA_MAX_TX_SIZE;
- ring->rx_pending = priv->dma_conf.dma_rx_size;
- ring->tx_pending = priv->dma_conf.dma_tx_size;
+ ring->rx_pending = priv->dma_conf->dma_rx_size;
+ ring->tx_pending = priv->dma_conf->dma_tx_size;
}
static int stmmac_set_ringparam(struct net_device *netdev,
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index bc19f8c19bb8..212bca73b9cf 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -80,7 +80,7 @@ static int phyaddr = -1;
module_param(phyaddr, int, 0444);
MODULE_PARM_DESC(phyaddr, "Physical device address");
-#define STMMAC_TX_THRESH(x) ((x)->dma_conf.dma_tx_size / 4)
+#define STMMAC_TX_THRESH(x) ((x)->dma_conf->dma_tx_size / 4)
/* Limit to make sure XDP TX and slow path can coexist */
#define STMMAC_XSK_TX_BUDGET_MAX 256
@@ -298,7 +298,7 @@ static void stmmac_disable_all_queues(struct stmmac_priv *priv)
/* synchronize_rcu() needed for pending XDP buffers to drain */
for (queue = 0; queue < rx_queues_cnt; queue++) {
- rx_q = &priv->dma_conf.rx_queue[queue];
+ rx_q = &priv->dma_conf->rx_queue[queue];
if (rx_q->xsk_pool) {
synchronize_rcu();
break;
@@ -357,10 +357,10 @@ static void print_pkt(unsigned char *buf, int len)
static inline u32 stmmac_tx_avail(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
return CIRC_SPACE(tx_q->cur_tx, tx_q->dirty_tx,
- priv->dma_conf.dma_tx_size);
+ priv->dma_conf->dma_tx_size);
}
static size_t stmmac_get_tx_desc_size(struct stmmac_priv *priv,
@@ -439,7 +439,7 @@ static void stmmac_set_queue_rx_buf_size(struct stmmac_priv *priv,
if (rx_q->xsk_pool && rx_q->buf_alloc_num)
buf_size = xsk_pool_get_rx_frame_size(rx_q->xsk_pool);
else
- buf_size = priv->dma_conf.dma_buf_sz;
+ buf_size = priv->dma_conf->dma_buf_sz;
stmmac_set_dma_bfsize(priv, priv->ioaddr, buf_size, chan);
}
@@ -451,10 +451,10 @@ static void stmmac_set_queue_rx_buf_size(struct stmmac_priv *priv,
*/
static inline u32 stmmac_rx_dirty(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
return CIRC_CNT(rx_q->cur_rx, rx_q->dirty_rx,
- priv->dma_conf.dma_rx_size);
+ priv->dma_conf->dma_rx_size);
}
static bool stmmac_eee_tx_busy(struct stmmac_priv *priv)
@@ -464,7 +464,7 @@ static bool stmmac_eee_tx_busy(struct stmmac_priv *priv)
/* check if all TX queues have the work finished */
for (queue = 0; queue < tx_cnt; queue++) {
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
if (tx_q->dirty_tx != tx_q->cur_tx)
return true; /* still unfinished work */
@@ -2174,7 +2174,7 @@ static void stmmac_free_tx_skbufs(struct stmmac_priv *priv)
u8 queue;
for (queue = 0; queue < tx_queue_cnt; queue++)
- dma_free_tx_skbufs(priv, &priv->dma_conf, queue);
+ dma_free_tx_skbufs(priv, priv->dma_conf, queue);
}
/**
@@ -2712,7 +2712,7 @@ static void stmmac_dma_operation_mode(struct stmmac_priv *priv)
/* configure all channels */
for (chan = 0; chan < rx_channels_count; chan++) {
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[chan];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[chan];
qmode = priv->plat->rx_queues_cfg[chan].mode_to_use;
@@ -2784,7 +2784,7 @@ static const struct xsk_tx_metadata_ops stmmac_xsk_tx_metadata_ops = {
static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
{
struct netdev_queue *nq = netdev_get_tx_queue(priv->dev, queue);
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
struct stmmac_txq_stats *txq_stats = &priv->xstats.txq_stats[queue];
bool csum = !priv->plat->tx_queues_cfg[queue].coe_unsupported;
struct xsk_buff_pool *pool = tx_q->xsk_pool;
@@ -2872,7 +2872,7 @@ static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
xsk_tx_metadata_to_compl(meta,
&tx_q->tx_skbuff_dma[entry].xsk_meta);
- tx_q->cur_tx = STMMAC_NEXT_ENTRY(tx_q->cur_tx, priv->dma_conf.dma_tx_size);
+ tx_q->cur_tx = STMMAC_NEXT_ENTRY(tx_q->cur_tx, priv->dma_conf->dma_tx_size);
entry = tx_q->cur_tx;
}
u64_stats_update_begin(&txq_stats->napi_syncp);
@@ -2920,7 +2920,7 @@ static void stmmac_bump_dma_threshold(struct stmmac_priv *priv, u32 chan)
static int stmmac_tx_clean(struct stmmac_priv *priv, int budget, u32 queue,
bool *pending_packets)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
struct stmmac_txq_stats *txq_stats = &priv->xstats.txq_stats[queue];
unsigned int bytes_compl = 0, pkts_compl = 0;
unsigned int entry, xmits = 0, count = 0;
@@ -2933,7 +2933,7 @@ static int stmmac_tx_clean(struct stmmac_priv *priv, int budget, u32 queue,
entry = tx_q->dirty_tx;
/* Try to clean all TX complete frame in 1 shot */
- while ((entry != tx_q->cur_tx) && count < priv->dma_conf.dma_tx_size) {
+ while ((entry != tx_q->cur_tx) && count < priv->dma_conf->dma_tx_size) {
struct xdp_frame *xdpf;
struct sk_buff *skb;
struct dma_desc *p;
@@ -3040,7 +3040,7 @@ static int stmmac_tx_clean(struct stmmac_priv *priv, int budget, u32 queue,
stmmac_release_tx_desc(priv, p, priv->descriptor_mode);
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_tx_size);
}
tx_q->dirty_tx = entry;
@@ -3108,13 +3108,13 @@ static int stmmac_tx_clean(struct stmmac_priv *priv, int budget, u32 queue,
*/
static void stmmac_tx_err(struct stmmac_priv *priv, u32 chan)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[chan];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[chan];
netif_tx_stop_queue(netdev_get_tx_queue(priv->dev, chan));
stmmac_stop_tx_dma(priv, chan);
- dma_free_tx_skbufs(priv, &priv->dma_conf, chan);
- stmmac_clear_tx_descriptors(priv, &priv->dma_conf, chan);
+ dma_free_tx_skbufs(priv, priv->dma_conf, chan);
+ stmmac_clear_tx_descriptors(priv, priv->dma_conf, chan);
stmmac_reset_tx_queue(priv, chan);
stmmac_init_tx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
tx_q->dma_tx_phy, chan);
@@ -3175,8 +3175,8 @@ static int stmmac_napi_check(struct stmmac_priv *priv, u32 chan, u32 dir)
{
int status = stmmac_dma_interrupt_status(priv, priv->ioaddr,
&priv->xstats, chan, dir);
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[chan];
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[chan];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[chan];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[chan];
struct stmmac_channel *ch = &priv->channel[chan];
struct napi_struct *rx_napi;
struct napi_struct *tx_napi;
@@ -3396,7 +3396,7 @@ static int stmmac_init_dma_engine(struct stmmac_priv *priv)
/* DMA RX Channel Configuration */
for (chan = 0; chan < rx_channels_count; chan++) {
- rx_q = &priv->dma_conf.rx_queue[chan];
+ rx_q = &priv->dma_conf->rx_queue[chan];
stmmac_init_rx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
rx_q->dma_rx_phy, chan);
@@ -3407,7 +3407,7 @@ static int stmmac_init_dma_engine(struct stmmac_priv *priv)
/* DMA TX Channel Configuration */
for (chan = 0; chan < tx_channels_count; chan++) {
- tx_q = &priv->dma_conf.tx_queue[chan];
+ tx_q = &priv->dma_conf->tx_queue[chan];
stmmac_init_tx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
tx_q->dma_tx_phy, chan);
@@ -3420,7 +3420,7 @@ static int stmmac_init_dma_engine(struct stmmac_priv *priv)
static void stmmac_tx_timer_arm(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
u32 tx_coal_timer = priv->tx_coal_timer[queue];
struct stmmac_channel *ch;
struct napi_struct *napi;
@@ -3488,7 +3488,7 @@ static void stmmac_init_coalesce(struct stmmac_priv *priv)
u8 chan;
for (chan = 0; chan < tx_channel_count; chan++) {
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[chan];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[chan];
priv->tx_coal_frames[chan] = STMMAC_TX_FRAMES;
priv->tx_coal_timer[chan] = STMMAC_COAL_TX_TIMER;
@@ -3509,12 +3509,12 @@ static void stmmac_set_rings_length(struct stmmac_priv *priv)
/* set TX ring length */
for (chan = 0; chan < tx_channels_count; chan++)
stmmac_set_tx_ring_len(priv, priv->ioaddr,
- (priv->dma_conf.dma_tx_size - 1), chan);
+ (priv->dma_conf->dma_tx_size - 1), chan);
/* set RX ring length */
for (chan = 0; chan < rx_channels_count; chan++)
stmmac_set_rx_ring_len(priv, priv->ioaddr,
- (priv->dma_conf.dma_rx_size - 1), chan);
+ (priv->dma_conf->dma_rx_size - 1), chan);
}
/**
@@ -3722,8 +3722,10 @@ static void stmmac_safety_feat_configuration(struct stmmac_priv *priv)
static bool stmmac_tso_channel_permitted(struct stmmac_priv *priv,
unsigned int chan)
{
- /* TSO and TBS cannot co-exist */
- return !(priv->dma_conf.tx_queue[chan].tbs & STMMAC_TBS_AVAIL);
+ /* TSO and TBS cannot co-exist. Feature checks also run while the
+ * datapath is detached, so use the lifetime-stable platform setting.
+ */
+ return !priv->plat->tx_queues_cfg[chan].tbs_en;
}
/**
@@ -3841,7 +3843,7 @@ static int stmmac_hw_setup(struct net_device *dev)
/* TBS */
for (chan = 0; chan < tx_cnt; chan++) {
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[chan];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[chan];
int enable = tx_q->tbs & STMMAC_TBS_AVAIL;
stmmac_enable_tbs(priv, priv->ioaddr, enable, chan);
@@ -3877,7 +3879,7 @@ static void stmmac_free_irq(struct net_device *dev,
if (msi->tx_irq[j] > 0) {
irq_set_affinity_hint(msi->tx_irq[j], NULL);
free_irq(msi->tx_irq[j],
- &priv->dma_conf.tx_queue[j]);
+ &priv->channel[j]);
}
}
irq_idx = priv->plat->rx_queues_to_use;
@@ -3887,7 +3889,7 @@ static void stmmac_free_irq(struct net_device *dev,
if (msi->rx_irq[j] > 0) {
irq_set_affinity_hint(msi->rx_irq[j], NULL);
free_irq(msi->rx_irq[j],
- &priv->dma_conf.rx_queue[j]);
+ &priv->channel[j]);
}
}
@@ -4041,7 +4043,7 @@ static int stmmac_request_irq_multi_msi(struct net_device *dev)
sprintf(int_name, "%s:%s-%d", dev->name, "rx", i);
ret = request_irq(msi->rx_irq[i],
stmmac_msi_intr_rx,
- 0, int_name, &priv->dma_conf.rx_queue[i]);
+ 0, int_name, &priv->channel[i]);
if (unlikely(ret < 0)) {
netdev_err(priv->dev,
"%s: alloc rx-%d MSI %d (error: %d)\n",
@@ -4065,7 +4067,7 @@ static int stmmac_request_irq_multi_msi(struct net_device *dev)
sprintf(int_name, "%s:%s-%d", dev->name, "tx", i);
ret = request_irq(msi->tx_irq[i],
stmmac_msi_intr_tx,
- 0, int_name, &priv->dma_conf.tx_queue[i]);
+ 0, int_name, &priv->channel[i]);
if (unlikely(ret < 0)) {
netdev_err(priv->dev,
"%s: alloc tx-%d MSI %d (error: %d)\n",
@@ -4189,8 +4191,8 @@ stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu)
/* Chose the tx/rx size from the already defined one in the
* priv struct. (if defined)
*/
- dma_conf->dma_tx_size = priv->dma_conf.dma_tx_size;
- dma_conf->dma_rx_size = priv->dma_conf.dma_rx_size;
+ dma_conf->dma_tx_size = priv->dma_conf->dma_tx_size;
+ dma_conf->dma_rx_size = priv->dma_conf->dma_rx_size;
if (!dma_conf->dma_tx_size)
dma_conf->dma_tx_size = DMA_DEFAULT_TX_SIZE;
@@ -4204,6 +4206,7 @@ stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu)
/* Setup per-TXQ tbs flag before TX descriptor alloc */
tx_q->tbs |= tbs_en ? STMMAC_TBS_AVAIL : 0;
+ tx_q->tbs |= priv->dma_conf->tx_queue[chan].tbs & STMMAC_TBS_EN;
}
ret = alloc_dma_desc_resources(priv, dma_conf);
@@ -4246,10 +4249,8 @@ static int __stmmac_open(struct net_device *dev,
u8 chan;
int ret;
- for (int i = 0; i < priv->plat->tx_queues_to_use; i++)
- if (priv->dma_conf.tx_queue[i].tbs & STMMAC_TBS_EN)
- dma_conf->tx_queue[i].tbs = priv->dma_conf.tx_queue[i].tbs;
- memcpy(&priv->dma_conf, dma_conf, sizeof(*dma_conf));
+ /* Keep RXQ metadata and timers at their registered addresses. */
+ priv->dma_conf = dma_conf;
/* The PHY is suspended when the interface is reopened without
* disconnecting the PHY, e.g. on MTU change. IEEE 802.3 allows PHYs
@@ -4301,7 +4302,7 @@ static int __stmmac_open(struct net_device *dev,
stmmac_stop_all_dma(priv);
for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
- hrtimer_cancel(&priv->dma_conf.tx_queue[chan].txtimer);
+ hrtimer_cancel(&priv->dma_conf->tx_queue[chan].txtimer);
est_error:
stmmac_release_ptp(priv);
@@ -4315,6 +4316,7 @@ static int __stmmac_open(struct net_device *dev,
static int stmmac_open(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ struct stmmac_dma_conf *old_conf = priv->dma_conf;
struct stmmac_dma_conf *dma_conf;
int ret;
@@ -4344,7 +4346,7 @@ static int stmmac_open(struct net_device *dev)
if (ret)
goto err_serdes;
- kfree(dma_conf);
+ kfree(old_conf);
/* We may have called phylink_speed_down before */
phylink_speed_up(priv->phylink);
@@ -4358,34 +4360,54 @@ static int stmmac_open(struct net_device *dev)
err_runtime_pm:
pm_runtime_put(priv->device);
err_dma_resources:
+ priv->dma_conf = old_conf;
free_dma_desc_resources(priv, dma_conf);
kfree(dma_conf);
return ret;
}
-static void __stmmac_release(struct net_device *dev)
+static void stmmac_stop_tx_queues(struct stmmac_priv *priv)
{
- struct stmmac_priv *priv = netdev_priv(dev);
u8 chan;
- /* Stop and disconnect the PHY */
- phylink_stop(priv->phylink);
+ netif_tx_disable(priv->dev);
- stmmac_disable_all_queues(priv);
+ /* A poll function can still arm a timer after napi_complete_done().
+ * Drain those poll tails and in-flight transmitters before cancelling
+ * the timers, so none can be rearmed after their final cancellation.
+ */
+ synchronize_net();
for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
- hrtimer_cancel(&priv->dma_conf.tx_queue[chan].txtimer);
+ hrtimer_cancel(&priv->dma_conf->tx_queue[chan].txtimer);
+}
+
+/* Quiesce NAPI and transmit queues without releasing their resources. */
+static void stmmac_quiesce(struct stmmac_priv *priv)
+{
+ stmmac_disable_all_queues(priv);
+ stmmac_stop_tx_queues(priv);
+ /* A TX completion may have rearmed this after phylink stopped EEE. */
+ timer_delete_sync(&priv->eee_ctrl_timer);
+}
- netif_tx_disable(dev);
+static void __stmmac_release(struct net_device *dev)
+{
+ struct stmmac_priv *priv = netdev_priv(dev);
+
+ phylink_stop(priv->phylink);
+ stmmac_quiesce(priv);
/* Free the IRQ lines */
stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
+ /* TX error IRQs can restart a queue after the first quiescence. */
+ stmmac_stop_tx_queues(priv);
/* Stop TX/RX DMA and clear the descriptors */
stmmac_stop_all_dma(priv);
/* Release and free the Rx/Tx resources */
- free_dma_desc_resources(priv, &priv->dma_conf);
+ free_dma_desc_resources(priv, priv->dma_conf);
stmmac_release_ptp(priv);
@@ -4439,7 +4461,7 @@ static bool stmmac_vlan_insert(struct stmmac_priv *priv, struct sk_buff *skb,
return false;
stmmac_set_tx_owner(priv, p);
- tx_q->cur_tx = STMMAC_NEXT_ENTRY(tx_q->cur_tx, priv->dma_conf.dma_tx_size);
+ tx_q->cur_tx = STMMAC_NEXT_ENTRY(tx_q->cur_tx, priv->dma_conf->dma_tx_size);
return true;
}
@@ -4459,7 +4481,7 @@ static void stmmac_tso_allocator(struct stmmac_priv *priv, u32 *entry,
dma_addr_t des, int total_len,
bool last_segment, u32 queue)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
struct dma_desc *desc;
u32 buff_size;
int tmp_len;
@@ -4469,7 +4491,7 @@ static void stmmac_tso_allocator(struct stmmac_priv *priv, u32 *entry,
while (tmp_len > 0) {
dma_addr_t curr_addr;
- *entry = STMMAC_NEXT_ENTRY(*entry, priv->dma_conf.dma_tx_size);
+ *entry = STMMAC_NEXT_ENTRY(*entry, priv->dma_conf->dma_tx_size);
WARN_ON(tx_q->tx_skbuff[*entry]);
if (tx_q->tbs & STMMAC_TBS_AVAIL)
@@ -4493,7 +4515,7 @@ static void stmmac_tso_allocator(struct stmmac_priv *priv, u32 *entry,
static void stmmac_flush_tx_descriptors(struct stmmac_priv *priv, int queue)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
/* The own bit must be the latest setting done when prepare the
* descriptor and then barrier is needed to make sure that
@@ -4646,7 +4668,7 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
nfrags = skb_shinfo(skb)->nr_frags;
queue = skb_get_queue_mapping(skb);
- tx_q = &priv->dma_conf.tx_queue[queue];
+ tx_q = &priv->dma_conf->tx_queue[queue];
txq_stats = &priv->xstats.txq_stats[queue];
first_tx = tx_q->cur_tx;
@@ -4684,7 +4706,7 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
stmmac_set_mss(priv, mss_desc, mss);
tx_q->mss = mss;
tx_q->cur_tx = STMMAC_NEXT_ENTRY(tx_q->cur_tx,
- priv->dma_conf.dma_tx_size);
+ priv->dma_conf->dma_tx_size);
WARN_ON(tx_q->tx_skbuff[tx_q->cur_tx]);
}
@@ -4755,7 +4777,7 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
/* Manage tx mitigation */
tx_packets = CIRC_CNT(tx_q->cur_tx + 1, first_tx,
- priv->dma_conf.dma_tx_size);
+ priv->dma_conf->dma_tx_size);
tx_q->tx_count_frames += tx_packets;
if ((skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP) && priv->hwts_tx_en)
@@ -4787,7 +4809,7 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
* ndo_start_xmit will fill this descriptor the next time it's
* called and stmmac_tx_clean may clean up to this descriptor.
*/
- tx_q->cur_tx = STMMAC_NEXT_ENTRY(tx_q->cur_tx, priv->dma_conf.dma_tx_size);
+ tx_q->cur_tx = STMMAC_NEXT_ENTRY(tx_q->cur_tx, priv->dma_conf->dma_tx_size);
if (unlikely(stmmac_tx_avail(priv, queue) <= (MAX_SKB_FRAGS + 1))) {
netif_dbg(priv, hw, priv->dev, "%s: stop transmitted packets\n",
@@ -4817,7 +4839,7 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
* segment.
*/
is_last_segment = CIRC_CNT(tx_q->cur_tx, first_entry,
- priv->dma_conf.dma_tx_size) == 1;
+ priv->dma_conf->dma_tx_size) == 1;
/* Complete the first descriptor before granting the DMA */
stmmac_prepare_tso_tx_desc(priv, first, 1, proto_hdr_len, 0, 1,
@@ -4855,13 +4877,13 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
for (;;) {
desc = stmmac_get_tx_desc(priv, tx_q, first_entry);
stmmac_release_tx_desc(priv, desc, priv->descriptor_mode);
- stmmac_free_tx_buffer(priv, &priv->dma_conf, queue,
+ stmmac_free_tx_buffer(priv, priv->dma_conf, queue,
first_entry);
if (first_entry == entry)
break;
first_entry = STMMAC_NEXT_ENTRY(first_entry,
- priv->dma_conf.dma_tx_size);
+ priv->dma_conf->dma_tx_size);
}
error:
dev_err(priv->device, "Tx dma map failed\n");
@@ -4945,7 +4967,7 @@ static netdev_tx_t stmmac_xmit(struct sk_buff *skb, struct net_device *dev)
return NETDEV_TX_BUSY;
}
- tx_q = &priv->dma_conf.tx_queue[queue];
+ tx_q = &priv->dma_conf->tx_queue[queue];
first_tx = tx_q->cur_tx;
/* Check if VLAN can be inserted by HW */
@@ -5021,7 +5043,7 @@ static netdev_tx_t stmmac_xmit(struct sk_buff *skb, struct net_device *dev)
unsigned int frag_size = skb_frag_size(frag);
bool last_segment = (i == (nfrags - 1));
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_tx_size);
WARN_ON(tx_q->tx_skbuff[entry]);
desc = stmmac_get_tx_desc(priv, tx_q, entry);
@@ -5051,7 +5073,7 @@ static netdev_tx_t stmmac_xmit(struct sk_buff *skb, struct net_device *dev)
* This approach takes care about the fragments: desc is the first
* element in case of no SG.
*/
- tx_packets = CIRC_CNT(entry + 1, first_tx, priv->dma_conf.dma_tx_size);
+ tx_packets = CIRC_CNT(entry + 1, first_tx, priv->dma_conf->dma_tx_size);
tx_q->tx_count_frames += tx_packets;
if ((skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP) && priv->hwts_tx_en)
@@ -5079,7 +5101,7 @@ static netdev_tx_t stmmac_xmit(struct sk_buff *skb, struct net_device *dev)
* ndo_start_xmit will fill this descriptor the next time it's
* called and stmmac_tx_clean may clean up to this descriptor.
*/
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_tx_size);
tx_q->cur_tx = entry;
if (netif_msg_pktdata(priv)) {
@@ -5194,7 +5216,7 @@ static void stmmac_rx_vlan(struct net_device *dev, struct sk_buff *skb)
*/
static inline void stmmac_rx_refill(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
int dirty = stmmac_rx_dirty(priv, queue);
unsigned int entry = rx_q->dirty_rx;
gfp_t gfp = (GFP_ATOMIC | __GFP_NOWARN);
@@ -5245,7 +5267,7 @@ static inline void stmmac_rx_refill(struct stmmac_priv *priv, u32 queue)
dma_wmb();
stmmac_set_rx_owner(priv, p, use_rx_wd);
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_rx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_rx_size);
}
rx_q->dirty_rx = entry;
stmmac_set_queue_rx_tail_ptr(priv, rx_q, queue, rx_q->dirty_rx);
@@ -5273,12 +5295,12 @@ static unsigned int stmmac_rx_buf1_len(struct stmmac_priv *priv,
/* First descriptor, not last descriptor and not split header */
if (status & rx_not_ls)
- return priv->dma_conf.dma_buf_sz;
+ return priv->dma_conf->dma_buf_sz;
plen = stmmac_get_rx_frame_len(priv, p, coe);
/* First descriptor and last descriptor and not split header */
- return min_t(unsigned int, priv->dma_conf.dma_buf_sz, plen);
+ return min_t(unsigned int, priv->dma_conf->dma_buf_sz, plen);
}
static unsigned int stmmac_rx_buf2_len(struct stmmac_priv *priv,
@@ -5308,7 +5330,7 @@ static unsigned int stmmac_rx_buf2_len(struct stmmac_priv *priv,
/* Not GMAC4 and not last descriptor */
if (priv->plat->core_type != DWMAC_CORE_GMAC4 && (status & rx_not_ls))
- return priv->dma_conf.dma_buf_sz;
+ return priv->dma_conf->dma_buf_sz;
/* GMAC4 or last descriptor */
plen = stmmac_get_rx_frame_len(priv, p, coe);
@@ -5320,7 +5342,7 @@ static int stmmac_xdp_xmit_xdpf(struct stmmac_priv *priv, int queue,
struct xdp_frame *xdpf, bool dma_map)
{
struct stmmac_txq_stats *txq_stats = &priv->xstats.txq_stats[queue];
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
bool csum = !priv->plat->tx_queues_cfg[queue].coe_unsupported;
unsigned int entry = tx_q->cur_tx;
enum stmmac_txbuf_type buf_type;
@@ -5385,7 +5407,7 @@ static int stmmac_xdp_xmit_xdpf(struct stmmac_priv *priv, int queue,
stmmac_enable_dma_transmission(priv, priv->ioaddr, queue);
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_tx_size);
tx_q->cur_tx = entry;
return STMMAC_XDP_TX;
@@ -5576,7 +5598,7 @@ static void stmmac_dispatch_skb_zc(struct stmmac_priv *priv, u32 queue,
static bool stmmac_rx_refill_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
{
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
unsigned int entry = rx_q->dirty_rx;
struct dma_desc *rx_desc = NULL;
bool ret = true;
@@ -5616,7 +5638,7 @@ static bool stmmac_rx_refill_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
dma_wmb();
stmmac_set_rx_owner(priv, rx_desc, use_rx_wd);
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_rx_size);
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf->dma_rx_size);
}
if (rx_desc) {
@@ -5640,7 +5662,7 @@ static struct stmmac_xdp_buff *xsk_buff_to_stmmac_ctx(struct xdp_buff *xdp)
static int stmmac_rx_zc(struct stmmac_priv *priv, int limit, u32 queue)
{
struct stmmac_rxq_stats *rxq_stats = &priv->xstats.rxq_stats[queue];
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
unsigned int count = 0, error = 0, len = 0;
int dirty = stmmac_rx_dirty(priv, queue);
unsigned int next_entry = rx_q->cur_rx;
@@ -5657,7 +5679,7 @@ static int stmmac_rx_zc(struct stmmac_priv *priv, int limit, u32 queue)
netdev_dbg(priv->dev, "%s: descriptor ring:\n", __func__);
desc_size = stmmac_get_rx_desc_size(priv);
- stmmac_display_ring(priv, rx_head, priv->dma_conf.dma_rx_size, true,
+ stmmac_display_ring(priv, rx_head, priv->dma_conf->dma_rx_size, true,
rx_q->dma_rx_phy, desc_size);
}
while (count < limit) {
@@ -5701,7 +5723,7 @@ static int stmmac_rx_zc(struct stmmac_priv *priv, int limit, u32 queue)
/* Prefetch the next RX descriptor */
next_entry = STMMAC_NEXT_ENTRY(rx_q->cur_rx,
- priv->dma_conf.dma_rx_size);
+ priv->dma_conf->dma_rx_size);
if (unlikely(next_entry == rx_q->dirty_rx))
break;
@@ -5826,7 +5848,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
{
u32 rx_errors = 0, rx_dropped = 0, rx_bytes = 0, rx_packets = 0;
struct stmmac_rxq_stats *rxq_stats = &priv->xstats.rxq_stats[queue];
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
struct stmmac_channel *ch = &priv->channel[queue];
unsigned int count = 0, error = 0, len = 0;
int status = 0, coe = priv->hw->rx_csum;
@@ -5840,7 +5862,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
int bufsz;
dma_dir = page_pool_get_dma_dir(rx_q->page_pool);
- bufsz = DIV_ROUND_UP(priv->dma_conf.dma_buf_sz, PAGE_SIZE) * PAGE_SIZE;
+ bufsz = DIV_ROUND_UP(priv->dma_conf->dma_buf_sz, PAGE_SIZE) * PAGE_SIZE;
if (netif_msg_rx_status(priv)) {
void *rx_head = stmmac_get_rx_desc(priv, rx_q, 0);
@@ -5848,7 +5870,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
netdev_dbg(priv->dev, "%s: descriptor ring:\n", __func__);
desc_size = stmmac_get_rx_desc_size(priv);
- stmmac_display_ring(priv, rx_head, priv->dma_conf.dma_rx_size, true,
+ stmmac_display_ring(priv, rx_head, priv->dma_conf->dma_rx_size, true,
rx_q->dma_rx_phy, desc_size);
}
if (in_progress) {
@@ -5891,7 +5913,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
break;
next_entry = STMMAC_NEXT_ENTRY(rx_q->cur_rx,
- priv->dma_conf.dma_rx_size);
+ priv->dma_conf->dma_rx_size);
if (unlikely(next_entry == rx_q->dirty_rx))
break;
@@ -6027,7 +6049,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
buf1_len, dma_dir);
skb_add_rx_frag(skb, skb_shinfo(skb)->nr_frags,
buf->page, buf->page_offset, buf1_len,
- priv->dma_conf.dma_buf_sz);
+ priv->dma_conf->dma_buf_sz);
buf->page = NULL;
}
@@ -6036,7 +6058,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue)
buf2_len, dma_dir);
skb_add_rx_frag(skb, skb_shinfo(skb)->nr_frags,
buf->sec_page, 0, buf2_len,
- priv->dma_conf.dma_buf_sz);
+ priv->dma_conf->dma_buf_sz);
buf->sec_page = NULL;
}
@@ -6261,6 +6283,7 @@ static void stmmac_set_rx_mode(struct net_device *dev)
static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ struct stmmac_dma_conf *old_conf = priv->dma_conf;
int txfifosz = priv->plat->tx_fifo_size;
struct stmmac_dma_conf *dma_conf;
const int mtu = new_mtu;
@@ -6297,19 +6320,22 @@ static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
return PTR_ERR(dma_conf);
}
+ netif_device_detach(dev);
__stmmac_release(dev);
ret = __stmmac_open(dev, dma_conf);
if (ret) {
+ priv->dma_conf = old_conf;
free_dma_desc_resources(priv, dma_conf);
kfree(dma_conf);
netdev_err(priv->dev, "failed reopening the interface after MTU change\n");
return ret;
}
- kfree(dma_conf);
+ kfree(old_conf);
stmmac_set_rx_mode(dev);
+ netif_device_attach(dev);
}
WRITE_ONCE(dev->mtu, mtu);
@@ -6481,15 +6507,11 @@ static irqreturn_t stmmac_safety_interrupt(int irq, void *dev_id)
static irqreturn_t stmmac_msi_intr_tx(int irq, void *data)
{
- struct stmmac_tx_queue *tx_q = (struct stmmac_tx_queue *)data;
- struct stmmac_dma_conf *dma_conf;
- int chan = tx_q->queue_index;
- struct stmmac_priv *priv;
+ struct stmmac_channel *ch = data;
+ struct stmmac_priv *priv = ch->priv_data;
+ int chan = ch->index;
int status;
- dma_conf = container_of(tx_q, struct stmmac_dma_conf, tx_queue[chan]);
- priv = container_of(dma_conf, struct stmmac_priv, dma_conf);
-
/* Check if adapter is up */
if (test_bit(STMMAC_DOWN, &priv->state))
return IRQ_HANDLED;
@@ -6508,13 +6530,9 @@ static irqreturn_t stmmac_msi_intr_tx(int irq, void *data)
static irqreturn_t stmmac_msi_intr_rx(int irq, void *data)
{
- struct stmmac_rx_queue *rx_q = (struct stmmac_rx_queue *)data;
- struct stmmac_dma_conf *dma_conf;
- int chan = rx_q->queue_index;
- struct stmmac_priv *priv;
-
- dma_conf = container_of(rx_q, struct stmmac_dma_conf, rx_queue[chan]);
- priv = container_of(dma_conf, struct stmmac_priv, dma_conf);
+ struct stmmac_channel *ch = data;
+ struct stmmac_priv *priv = ch->priv_data;
+ int chan = ch->index;
/* Check if adapter is up */
if (test_bit(STMMAC_DOWN, &priv->state))
@@ -6651,45 +6669,48 @@ static int stmmac_rings_status_show(struct seq_file *seq, void *v)
{
struct net_device *dev = seq->private;
struct stmmac_priv *priv = netdev_priv(dev);
- u8 rx_count = priv->plat->rx_queues_to_use;
- u8 tx_count = priv->plat->tx_queues_to_use;
- u8 queue;
+ u8 rx_count, tx_count, queue;
+ rtnl_lock();
if ((dev->flags & IFF_UP) == 0)
- return 0;
+ goto out_unlock;
+ rx_count = priv->plat->rx_queues_to_use;
+ tx_count = priv->plat->tx_queues_to_use;
for (queue = 0; queue < rx_count; queue++) {
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
seq_printf(seq, "RX Queue %d:\n", queue);
if (priv->extend_desc) {
seq_printf(seq, "Extended descriptor ring:\n");
sysfs_display_ring((void *)rx_q->dma_erx,
- priv->dma_conf.dma_rx_size, 1, seq, rx_q->dma_rx_phy);
+ priv->dma_conf->dma_rx_size, 1, seq, rx_q->dma_rx_phy);
} else {
seq_printf(seq, "Descriptor ring:\n");
sysfs_display_ring((void *)rx_q->dma_rx,
- priv->dma_conf.dma_rx_size, 0, seq, rx_q->dma_rx_phy);
+ priv->dma_conf->dma_rx_size, 0, seq, rx_q->dma_rx_phy);
}
}
for (queue = 0; queue < tx_count; queue++) {
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
seq_printf(seq, "TX Queue %d:\n", queue);
if (priv->extend_desc) {
seq_printf(seq, "Extended descriptor ring:\n");
sysfs_display_ring((void *)tx_q->dma_etx,
- priv->dma_conf.dma_tx_size, 1, seq, tx_q->dma_tx_phy);
+ priv->dma_conf->dma_tx_size, 1, seq, tx_q->dma_tx_phy);
} else if (!(tx_q->tbs & STMMAC_TBS_AVAIL)) {
seq_printf(seq, "Descriptor ring:\n");
sysfs_display_ring((void *)tx_q->dma_tx,
- priv->dma_conf.dma_tx_size, 0, seq, tx_q->dma_tx_phy);
+ priv->dma_conf->dma_tx_size, 0, seq, tx_q->dma_tx_phy);
}
}
+out_unlock:
+ rtnl_unlock();
return 0;
}
DEFINE_SHOW_ATTRIBUTE(stmmac_rings_status);
@@ -7113,6 +7134,10 @@ static int stmmac_xdp_xmit(struct net_device *dev, int num_frames,
nq = netdev_get_tx_queue(priv->dev, queue);
__netif_tx_lock(nq, cpu);
+ if (unlikely(!netif_device_present(dev) || netif_tx_queue_stopped(nq))) {
+ __netif_tx_unlock(nq);
+ return -ENETDOWN;
+ }
/* Avoids TX time-out as we are sharing with slow path */
txq_trans_cond_update(nq);
@@ -7146,31 +7171,31 @@ void stmmac_disable_rx_queue(struct stmmac_priv *priv, u32 queue)
spin_unlock_irqrestore(&ch->lock, flags);
stmmac_stop_rx_dma(priv, queue);
- __free_dma_rx_desc_resources(priv, &priv->dma_conf, queue);
+ __free_dma_rx_desc_resources(priv, priv->dma_conf, queue);
}
void stmmac_enable_rx_queue(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
struct stmmac_channel *ch = &priv->channel[queue];
unsigned long flags;
int ret;
- ret = __alloc_dma_rx_desc_resources(priv, &priv->dma_conf, queue);
+ ret = __alloc_dma_rx_desc_resources(priv, priv->dma_conf, queue);
if (ret) {
netdev_err(priv->dev, "Failed to alloc RX desc.\n");
return;
}
- ret = __init_dma_rx_desc_rings(priv, &priv->dma_conf, queue, GFP_KERNEL);
+ ret = __init_dma_rx_desc_rings(priv, priv->dma_conf, queue, GFP_KERNEL);
if (ret) {
- __free_dma_rx_desc_resources(priv, &priv->dma_conf, queue);
+ __free_dma_rx_desc_resources(priv, priv->dma_conf, queue);
netdev_err(priv->dev, "Failed to init RX desc.\n");
return;
}
stmmac_reset_rx_queue(priv, queue);
- stmmac_clear_rx_descriptors(priv, &priv->dma_conf, queue);
+ stmmac_clear_rx_descriptors(priv, priv->dma_conf, queue);
stmmac_init_rx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
rx_q->dma_rx_phy, queue);
@@ -7196,31 +7221,31 @@ void stmmac_disable_tx_queue(struct stmmac_priv *priv, u32 queue)
spin_unlock_irqrestore(&ch->lock, flags);
stmmac_stop_tx_dma(priv, queue);
- __free_dma_tx_desc_resources(priv, &priv->dma_conf, queue);
+ __free_dma_tx_desc_resources(priv, priv->dma_conf, queue);
}
void stmmac_enable_tx_queue(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
struct stmmac_channel *ch = &priv->channel[queue];
unsigned long flags;
int ret;
- ret = __alloc_dma_tx_desc_resources(priv, &priv->dma_conf, queue);
+ ret = __alloc_dma_tx_desc_resources(priv, priv->dma_conf, queue);
if (ret) {
netdev_err(priv->dev, "Failed to alloc TX desc.\n");
return;
}
- ret = __init_dma_tx_desc_rings(priv, &priv->dma_conf, queue);
+ ret = __init_dma_tx_desc_rings(priv, priv->dma_conf, queue);
if (ret) {
- __free_dma_tx_desc_resources(priv, &priv->dma_conf, queue);
+ __free_dma_tx_desc_resources(priv, priv->dma_conf, queue);
netdev_err(priv->dev, "Failed to init TX desc.\n");
return;
}
stmmac_reset_tx_queue(priv, queue);
- stmmac_clear_tx_descriptors(priv, &priv->dma_conf, queue);
+ stmmac_clear_tx_descriptors(priv, priv->dma_conf, queue);
stmmac_init_tx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
tx_q->dma_tx_phy, queue);
@@ -7240,25 +7265,18 @@ void stmmac_enable_tx_queue(struct stmmac_priv *priv, u32 queue)
void stmmac_xdp_release(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
- u8 chan;
-
- /* Ensure tx function is not running */
- netif_tx_disable(dev);
- /* Disable NAPI process */
- stmmac_disable_all_queues(priv);
-
- for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
- hrtimer_cancel(&priv->dma_conf.tx_queue[chan].txtimer);
+ stmmac_quiesce(priv);
/* Free the IRQ lines */
stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
+ stmmac_stop_tx_queues(priv);
/* Stop TX/RX DMA channels */
stmmac_stop_all_dma(priv);
/* Release and free the Rx/Tx resources */
- free_dma_desc_resources(priv, &priv->dma_conf);
+ free_dma_desc_resources(priv, priv->dma_conf);
/* Disable the MAC Rx/Tx */
stmmac_mac_set(priv, priv->ioaddr, false);
@@ -7282,14 +7300,14 @@ int stmmac_xdp_open(struct net_device *dev)
u8 chan;
int ret;
- ret = alloc_dma_desc_resources(priv, &priv->dma_conf);
+ ret = alloc_dma_desc_resources(priv, priv->dma_conf);
if (ret < 0) {
netdev_err(dev, "%s: DMA descriptors allocation failed\n",
__func__);
goto dma_desc_error;
}
- ret = init_dma_desc_rings(dev, &priv->dma_conf, GFP_KERNEL);
+ ret = init_dma_desc_rings(dev, priv->dma_conf, GFP_KERNEL);
if (ret < 0) {
netdev_err(dev, "%s: DMA descriptors initialization failed\n",
__func__);
@@ -7309,7 +7327,7 @@ int stmmac_xdp_open(struct net_device *dev)
/* DMA RX Channel Configuration */
for (chan = 0; chan < rx_cnt; chan++) {
- rx_q = &priv->dma_conf.rx_queue[chan];
+ rx_q = &priv->dma_conf->rx_queue[chan];
stmmac_init_rx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
rx_q->dma_rx_phy, chan);
@@ -7324,7 +7342,7 @@ int stmmac_xdp_open(struct net_device *dev)
/* DMA TX Channel Configuration */
for (chan = 0; chan < tx_cnt; chan++) {
- tx_q = &priv->dma_conf.tx_queue[chan];
+ tx_q = &priv->dma_conf->tx_queue[chan];
stmmac_init_tx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
tx_q->dma_tx_phy, chan);
@@ -7356,10 +7374,10 @@ int stmmac_xdp_open(struct net_device *dev)
stmmac_stop_all_dma(priv);
for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
- hrtimer_cancel(&priv->dma_conf.tx_queue[chan].txtimer);
+ hrtimer_cancel(&priv->dma_conf->tx_queue[chan].txtimer);
init_error:
- free_dma_desc_resources(priv, &priv->dma_conf);
+ free_dma_desc_resources(priv, priv->dma_conf);
dma_desc_error:
return ret;
}
@@ -7367,8 +7385,6 @@ int stmmac_xdp_open(struct net_device *dev)
int stmmac_xsk_wakeup(struct net_device *dev, u32 queue, u32 flags)
{
struct stmmac_priv *priv = netdev_priv(dev);
- struct stmmac_rx_queue *rx_q;
- struct stmmac_tx_queue *tx_q;
struct stmmac_channel *ch;
if (test_bit(STMMAC_DOWN, &priv->state) ||
@@ -7382,11 +7398,9 @@ int stmmac_xsk_wakeup(struct net_device *dev, u32 queue, u32 flags)
queue >= priv->plat->tx_queues_to_use)
return -EINVAL;
- rx_q = &priv->dma_conf.rx_queue[queue];
- tx_q = &priv->dma_conf.tx_queue[queue];
ch = &priv->channel[queue];
- if (!rx_q->xsk_pool && !tx_q->xsk_pool)
+ if (!test_bit(queue, priv->af_xdp_zc_qps))
return -EINVAL;
if (!napi_if_scheduled_mark_missed(&ch->rxtx_napi)) {
@@ -7799,8 +7813,8 @@ int stmmac_reinit_ringparam(struct net_device *dev, u32 rx_size, u32 tx_size)
if (netif_running(dev))
stmmac_release(dev);
- priv->dma_conf.dma_rx_size = rx_size;
- priv->dma_conf.dma_tx_size = tx_size;
+ priv->dma_conf->dma_rx_size = rx_size;
+ priv->dma_conf->dma_tx_size = tx_size;
if (netif_running(dev))
ret = stmmac_open(dev);
@@ -7974,6 +7988,13 @@ struct plat_stmmacenet_data *stmmac_plat_dat_alloc(struct device *dev)
}
EXPORT_SYMBOL_GPL(stmmac_plat_dat_alloc);
+static void stmmac_free_dma_conf(void *data)
+{
+ struct stmmac_priv *priv = data;
+
+ kfree(priv->dma_conf);
+}
+
static int __stmmac_dvr_probe(struct device *device,
struct plat_stmmacenet_data *plat_dat,
struct stmmac_resources *res)
@@ -7998,6 +8019,13 @@ static int __stmmac_dvr_probe(struct device *device,
priv = netdev_priv(ndev);
priv->device = device;
priv->dev = ndev;
+ /* Keep ring sizes and per-queue settings even while the device is down. */
+ priv->dma_conf = kzalloc_obj(*priv->dma_conf);
+ if (!priv->dma_conf)
+ return -ENOMEM;
+ ret = devm_add_action_or_reset(device, stmmac_free_dma_conf, priv);
+ if (ret)
+ return ret;
for (i = 0; i < MTL_MAX_RX_QUEUES; i++)
u64_stats_init(&priv->xstats.rxq_stats[i].napi_syncp);
@@ -8362,7 +8390,6 @@ int stmmac_suspend(struct device *dev)
{
struct net_device *ndev = dev_get_drvdata(dev);
struct stmmac_priv *priv = netdev_priv(ndev);
- u8 chan;
if (!ndev || !netif_running(ndev))
goto suspend_bsp;
@@ -8371,10 +8398,7 @@ int stmmac_suspend(struct device *dev)
netif_device_detach(ndev);
- stmmac_disable_all_queues(priv);
-
- for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
- hrtimer_cancel(&priv->dma_conf.tx_queue[chan].txtimer);
+ stmmac_quiesce(priv);
if (priv->eee_sw_timer_en) {
priv->tx_path_in_lpi_mode = false;
@@ -8414,7 +8438,7 @@ EXPORT_SYMBOL_GPL(stmmac_suspend);
static void stmmac_reset_rx_queue(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_rx_queue *rx_q = &priv->dma_conf.rx_queue[queue];
+ struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
rx_q->cur_rx = 0;
rx_q->dirty_rx = 0;
@@ -8422,7 +8446,7 @@ static void stmmac_reset_rx_queue(struct stmmac_priv *priv, u32 queue)
static void stmmac_reset_tx_queue(struct stmmac_priv *priv, u32 queue)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
+ struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
tx_q->cur_tx = 0;
tx_q->dirty_tx = 0;
@@ -8505,7 +8529,7 @@ int stmmac_resume(struct device *dev)
stmmac_reset_queues_param(priv);
stmmac_free_tx_skbufs(priv);
- stmmac_clear_descriptors(priv, &priv->dma_conf);
+ stmmac_clear_descriptors(priv, priv->dma_conf);
ret = stmmac_hw_setup(ndev);
if (ret < 0) {
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c
index 6097f312fce4..1adf634292d3 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c
@@ -869,8 +869,8 @@ static int stmmac_test_flowctrl(struct stmmac_priv *priv)
struct stmmac_channel *ch = &priv->channel[i];
u32 tail;
- tail = priv->dma_conf.rx_queue[i].dma_rx_phy +
- (priv->dma_conf.dma_rx_size * sizeof(struct dma_desc));
+ tail = priv->dma_conf->rx_queue[i].dma_rx_phy +
+ (priv->dma_conf->dma_rx_size * sizeof(struct dma_desc));
stmmac_set_rx_tail_ptr(priv, priv->ioaddr, tail, i);
stmmac_start_rx(priv, priv->ioaddr, i);
@@ -1678,7 +1678,7 @@ static int stmmac_test_l4filt_sa_udp(struct stmmac_priv *priv)
static int __stmmac_test_jumbo(struct stmmac_priv *priv, u16 queue)
{
struct stmmac_packet_attrs attr = { };
- int size = priv->dma_conf.dma_buf_sz;
+ int size = priv->dma_conf->dma_buf_sz;
if (!dwmac_is_xmac(priv->plat->core_type))
size -= NET_IP_ALIGN;
@@ -1764,7 +1764,7 @@ static int stmmac_test_tbs(struct stmmac_priv *priv)
/* Find first TBS enabled Queue, if any */
for (i = 0; i < priv->plat->tx_queues_to_use; i++)
- if (priv->dma_conf.tx_queue[i].tbs & STMMAC_TBS_AVAIL)
+ if (priv->dma_conf->tx_queue[i].tbs & STMMAC_TBS_AVAIL)
break;
if (i >= priv->plat->tx_queues_to_use)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
index 70d623f7083b..839a93166520 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
@@ -1197,13 +1197,13 @@ static int tc_setup_etf(struct stmmac_priv *priv,
return -EOPNOTSUPP;
if (qopt->queue >= priv->plat->tx_queues_to_use)
return -EINVAL;
- if (!(priv->dma_conf.tx_queue[qopt->queue].tbs & STMMAC_TBS_AVAIL))
+ if (!(priv->dma_conf->tx_queue[qopt->queue].tbs & STMMAC_TBS_AVAIL))
return -EINVAL;
if (qopt->enable)
- priv->dma_conf.tx_queue[qopt->queue].tbs |= STMMAC_TBS_EN;
+ priv->dma_conf->tx_queue[qopt->queue].tbs |= STMMAC_TBS_EN;
else
- priv->dma_conf.tx_queue[qopt->queue].tbs &= ~STMMAC_TBS_EN;
+ priv->dma_conf->tx_queue[qopt->queue].tbs &= ~STMMAC_TBS_EN;
netdev_info(priv->dev, "%s ETF for Queue %d\n",
qopt->enable ? "enabled" : "disabled", qopt->queue);
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 13/19] net: stmmac: track datapath and power ownership across failed reopening
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (11 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 12/19] net: stmmac: keep DMA configurations at stable addresses James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 14/19] net: stmmac: use the tracked datapath restart for XSK pool changes James Hilliard
` (7 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Track allocated rings, IRQs and NAPI independently of IFF_UP and of
system sleep. Leave failed live reopening detached while the successful
ndo_open still owns its PHY attachment and runtime-PM reference;
ordinary down/up can recover without duplicate NAPI shutdown or PM puts.
Convert every live reopen consumer together: MTU, XDP program changes
and ethtool ring/channel changes. Restore old settings on error and
publish queue counts only after setup succeeds. XDP removal must release
its program even if rebuilding the non-XDP datapath fails. Use phylink
link replay for successful XDP rebuilding so the PHY is not stopped or
renegotiated.
Keep core sleep cleanup separate from platform and clock restoration.
Serialize noirq clock ownership and retry pending power transitions
before accessing MMIO. If restoration fails, block MAC, MDIO, PCS, IRQ
and PHC register accesses and finish only software cleanup. Check MAC
WoL viability before skipping an already completed core sleep sequence.
A level PMT interrupt can arrive before the IRQ subsystem suspends the
handlers, or when it reenables them before normal device resume. Keep
platform MAC-WoL register, receive and timestamp clocks powered throughout
those windows, instead of returning without acknowledging an asserted
interrupt. Separate optional wake-resource hooks from BSP power removal;
STM32 enables its stop-mode clock without dropping the live clocks.
Tegra retains its private clocks and reset state, and tracks actual clock
ownership across failed recovery. Publish wake handling before arming PMT.
Acknowledge timestamp status even while the PHC is blocked, without
delivering clock events; a new timestamp can arrive after PMT is disarmed
but before the reset or recovery completes.
PCI keeps its existing PME/D3 policy, but runs the BSP power transition
inside the noirq window. Restore register access before handlers run.
If restoration fails, mask the function's INTx assertion through PCI
configuration space, leaving a shared controller line usable by other
devices. Restore INTx after a successful datapath restart. Balance pending
wake resources during resume or close, and undo core sleep if a suspend
callback itself fails.
Close must reattach the netdev so ndo_open can retry power restoration.
Reject ethtool, timestamp, MAC-address, feature and qdisc installation
requests while registers remain inaccessible. Drain receive-filter work
under the address lock before power removal. VLAN and qdisc teardown
still clear their software state without touching unpowered registers.
Permit flower and u32 destruction while down or detached, clearing cached
rules without disabling already-stopped NAPI or accessing registers. TC
forgets deleted filters even when the driver returns an error, so rejecting
destruction would leave stale rules to be replayed on recovery. Keep
installation-only feature and RSS checks out of the destruction path.
Guard non-netdev-detach consumers, including TC, descriptor readback,
XDP transmission and reset work. Freeze deferred XSK teardown before
sleep through the prerequisite XSK change. Preserve PHY ownership on
failed ethtool reopen and stop a PHY temporarily resumed to supply the
MAC reset clock.
Fixes: 3470079687448 ("net: ethernet: stmicro: stmmac: permit MTU change with interface up")
Fixes: 6896c2449a18 ("net: stmmac: Check stmmac_hw_setup() in stmmac_resume()")
Fixes: ac746c8520d9 ("net: stmmac: enhance XDP ZC driver level switching performance")
Fixes: aa042f60e496 ("net: stmmac: Add support to Ethtool get/set ring parameters")
Fixes: 0366f7e06a6b ("net: stmmac: add ethtool support for get/set channels")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/dwmac-intel.c | 2 +-
.../net/ethernet/stmicro/stmmac/dwmac-loongson.c | 2 +-
.../net/ethernet/stmicro/stmmac/dwmac-motorcomm.c | 2 +-
drivers/net/ethernet/stmicro/stmmac/dwmac-stm32.c | 18 +
drivers/net/ethernet/stmicro/stmmac/dwmac-tegra.c | 27 +-
.../net/ethernet/stmicro/stmmac/dwmac1000_core.c | 4 +
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 33 +
.../net/ethernet/stmicro/stmmac/stmmac_ethtool.c | 11 +
.../net/ethernet/stmicro/stmmac/stmmac_hwtstamp.c | 10 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 697 ++++++++++++++++++---
drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c | 15 +
drivers/net/ethernet/stmicro/stmmac/stmmac_pci.c | 2 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_pcs.c | 3 +-
.../net/ethernet/stmicro/stmmac/stmmac_platform.c | 40 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c | 98 +--
drivers/net/ethernet/stmicro/stmmac/stmmac_vlan.c | 6 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c | 39 +-
include/linux/stmmac.h | 7 +
18 files changed, 831 insertions(+), 185 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-intel.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-intel.c
index f5f9fa67ecd7..ba1231852e10 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac-intel.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-intel.c
@@ -1420,7 +1420,7 @@ static struct pci_driver intel_eth_pci_driver = {
.probe = intel_eth_pci_probe,
.remove = intel_eth_pci_remove,
.driver = {
- .pm = &stmmac_simple_pm_ops,
+ .pm = &stmmac_pci_pm_ops,
},
};
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-loongson.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-loongson.c
index eb14c197d6ae..b91877d3d5fe 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac-loongson.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-loongson.c
@@ -607,7 +607,7 @@ static struct pci_driver loongson_dwmac_driver = {
.probe = loongson_dwmac_probe,
.remove = loongson_dwmac_remove,
.driver = {
- .pm = &stmmac_simple_pm_ops,
+ .pm = &stmmac_pci_pm_ops,
},
};
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-motorcomm.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-motorcomm.c
index 124fcd55a429..b888b87bfd90 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac-motorcomm.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-motorcomm.c
@@ -376,7 +376,7 @@ static struct pci_driver dwmac_motorcomm_pci_driver = {
.probe = motorcomm_probe,
.remove = motorcomm_remove,
.driver = {
- .pm = &stmmac_simple_pm_ops,
+ .pm = &stmmac_pci_pm_ops,
},
};
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-stm32.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-stm32.c
index e1b260ed4790..9544925bc36f 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac-stm32.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-stm32.c
@@ -522,6 +522,22 @@ static int stm32_dwmac_resume(struct device *dev, void *bsp_priv)
return stm32_dwmac_init(priv->plat);
}
+static int stm32_dwmac_suspend_wol(struct device *dev, void *bsp_priv)
+{
+ struct stm32_dwmac *dwmac = bsp_priv;
+
+ /* Enable the stop-mode wake clock without dropping the live clocks. */
+ return dwmac->ops->suspend ? dwmac->ops->suspend(dwmac) : 0;
+}
+
+static void stm32_dwmac_resume_wol(struct device *dev, void *bsp_priv)
+{
+ struct stm32_dwmac *dwmac = bsp_priv;
+
+ if (dwmac->ops->resume)
+ dwmac->ops->resume(dwmac);
+}
+
static int stm32_dwmac_probe(struct platform_device *pdev)
{
struct plat_stmmacenet_data *plat_dat;
@@ -561,6 +577,8 @@ static int stm32_dwmac_probe(struct platform_device *pdev)
plat_dat->bsp_priv = dwmac;
plat_dat->suspend = stm32_dwmac_suspend;
plat_dat->resume = stm32_dwmac_resume;
+ plat_dat->suspend_wol = stm32_dwmac_suspend_wol;
+ plat_dat->resume_wol = stm32_dwmac_resume_wol;
ret = stm32_dwmac_init(plat_dat);
if (ret)
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-tegra.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-tegra.c
index 4ede4420c93b..43d9f6119165 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac-tegra.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-tegra.c
@@ -16,6 +16,7 @@ struct tegra_mgbe {
struct device *dev;
struct clk_bulk_data *clks;
+ bool suspended;
struct reset_control *rst_mac;
struct reset_control *rst_pcs;
@@ -54,18 +55,29 @@ struct tegra_mgbe {
#define MAC_SBD_INTR BIT(2)
#define MGBE_WRAP_AXI_ASID0_CTRL 0x8400
+static int __maybe_unused tegra_mgbe_resume(struct device *dev);
+
static int __maybe_unused tegra_mgbe_suspend(struct device *dev)
{
struct tegra_mgbe *mgbe = get_stmmac_bsp_priv(dev);
+ struct stmmac_priv *priv = netdev_priv(dev_get_drvdata(dev));
int err;
err = stmmac_suspend(dev);
if (err)
return err;
+ /* MAC WoL needs the receiver and the register interface kept alive. */
+ if (priv->irq_wake || mgbe->suspended)
+ return 0;
+
clk_bulk_disable_unprepare(ARRAY_SIZE(mgbe_clks), mgbe->clks);
+ mgbe->suspended = true;
- return reset_control_assert(mgbe->rst_mac);
+ err = reset_control_assert(mgbe->rst_mac);
+ if (err && tegra_mgbe_resume(dev))
+ dev_err(dev, "failed to restore device after suspend failure\n");
+ return err;
}
static int __maybe_unused tegra_mgbe_resume(struct device *dev)
@@ -74,13 +86,18 @@ static int __maybe_unused tegra_mgbe_resume(struct device *dev)
u32 value;
int err;
+ if (!mgbe->suspended)
+ return stmmac_resume(dev);
+
err = clk_bulk_prepare_enable(ARRAY_SIZE(mgbe_clks), mgbe->clks);
if (err < 0)
return err;
err = reset_control_deassert(mgbe->rst_mac);
- if (err < 0)
+ if (err < 0) {
+ clk_bulk_disable_unprepare(ARRAY_SIZE(mgbe_clks), mgbe->clks);
return err;
+ }
/* Enable common interrupt at wrapper level */
writel(MAC_SBD_INTR, mgbe->regs + MGBE_WRAP_COMMON_INTR_ENABLE);
@@ -104,9 +121,11 @@ static int __maybe_unused tegra_mgbe_resume(struct device *dev)
return err;
}
+ mgbe->suspended = false;
err = stmmac_resume(dev);
- if (err < 0)
- clk_bulk_disable_unprepare(ARRAY_SIZE(mgbe_clks), mgbe->clks);
+ /* Core resume failure retains the suspended datapath for retry or
+ * close. Keep its register interface powered until that cleanup.
+ */
return err;
}
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac1000_core.c b/drivers/net/ethernet/stmicro/stmmac/dwmac1000_core.c
index d4ace3924891..8945d93c6a05 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac1000_core.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac1000_core.c
@@ -531,6 +531,10 @@ void dwmac1000_timestamp_interrupt(struct stmmac_priv *priv)
/* Clears the timestamp interrupt */
ts_status = readl(priv->ptpaddr + GMAC3_X_TIMESTAMP_STATUS);
+ /* A blocked PHC still needs its powered interrupt source cleared. */
+ if (READ_ONCE(priv->ptp_blocked))
+ return;
+
if (!(priv->plat->flags & STMMAC_FLAG_EXT_SNAPSHOT_EN))
return;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index fa01070cbef3..6927afd01175 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -258,6 +258,15 @@ struct stmmac_msi {
char int_name_tx_irq[MTL_MAX_TX_QUEUES][IFNAMSIZ + 18];
};
+enum stmmac_datapath_state {
+ /* No IRQs or DMA allocations owned by a successful open. */
+ STMMAC_DATAPATH_DOWN,
+ /* Resources allocated, NAPI enabled. */
+ STMMAC_DATAPATH_RUNNING,
+ /* Resources retained, NAPI and DMA stopped; also after failed resume. */
+ STMMAC_DATAPATH_SUSPENDED,
+};
+
struct stmmac_priv {
/* Frequently used values are kept adjacent for cache effect */
u32 tx_coal_frames[MTL_MAX_TX_QUEUES];
@@ -281,6 +290,21 @@ struct stmmac_priv {
struct mutex lock;
struct stmmac_dma_conf *dma_conf;
+ /* IRQ/DMA ownership and NAPI state, serialized by RTNL. */
+ enum stmmac_datapath_state datapath;
+ /* Core sleep sequence completed, independently of datapath ownership. */
+ bool hw_suspended;
+ /* System PM blocks MMIO until power restoration has completed. */
+ bool hw_unavailable;
+ bool bsp_suspended;
+ bool wol_suspended;
+ /* Function-local INTx mask after a failed PCI noirq power restoration. */
+ bool pci_intx_masked;
+ bool bus_clks_suspended;
+ bool ptp_clock_enabled;
+ bool ptp_clock_suspended;
+ /* Clock ownership can also change in noirq PM, without RTNL. */
+ struct mutex pm_mutex;
/* Generic channel for NAPI */
struct stmmac_channel channel[STMMAC_CH_MAX];
@@ -401,10 +425,12 @@ enum stmmac_state {
};
extern const struct dev_pm_ops stmmac_simple_pm_ops;
+extern const struct dev_pm_ops stmmac_pci_pm_ops;
int stmmac_mdio_unregister(struct net_device *ndev);
int stmmac_mdio_register(struct net_device *ndev);
int stmmac_mdio_reset(struct mii_bus *mii);
+int stmmac_resume_clocks(struct stmmac_priv *priv);
void stmmac_mdio_lock(struct stmmac_priv *priv);
void stmmac_mdio_unlock(struct stmmac_priv *priv);
int stmmac_pcs_setup(struct net_device *ndev);
@@ -440,6 +466,13 @@ static inline bool stmmac_xdp_is_enabled(struct stmmac_priv *priv)
return !!priv->xdp_prog;
}
+/* RTNL serializes TC callbacks with datapath and power transitions. */
+static inline bool stmmac_tc_active(struct stmmac_priv *priv)
+{
+ return priv->datapath == STMMAC_DATAPATH_RUNNING &&
+ netif_device_present(priv->dev) && !priv->hw_unavailable;
+}
+
void stmmac_disable_rx_queue(struct stmmac_priv *priv, u32 queue);
void stmmac_enable_rx_queue(struct stmmac_priv *priv, u32 queue);
void stmmac_disable_tx_queue(struct stmmac_priv *priv, u32 queue);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
index 1fdd63d6bba1..ffbb00bc6fc0 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
@@ -1089,7 +1089,18 @@ static void stmmac_get_mm_stats(struct net_device *ndev,
s->MACMergeHoldCount = mmc->mmc_tx_hold_req_cntr;
}
+static int stmmac_ethtool_begin(struct net_device *dev)
+{
+ struct stmmac_priv *priv = netdev_priv(dev);
+
+ /* Close reattaches the netdev so open can retry power restoration.
+ * Presence alone does not make registers accessible after failed resume.
+ */
+ return priv->hw_unavailable ? -EHOSTDOWN : 0;
+}
+
static const struct ethtool_ops stmmac_ethtool_ops = {
+ .begin = stmmac_ethtool_begin,
.supported_coalesce_params = ETHTOOL_COALESCE_USECS |
ETHTOOL_COALESCE_MAX_FRAMES,
.get_drvinfo = stmmac_ethtool_getdrvinfo,
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_hwtstamp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_hwtstamp.c
index b9a985fa772c..710bb6ad06d3 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_hwtstamp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_hwtstamp.c
@@ -220,7 +220,8 @@ static void timestamp_interrupt(struct stmmac_priv *priv)
u64 ptp_time;
int i;
- if (priv->plat->flags & STMMAC_FLAG_INT_SNAPSHOT_EN) {
+ if (!READ_ONCE(priv->ptp_blocked) &&
+ (priv->plat->flags & STMMAC_FLAG_INT_SNAPSHOT_EN)) {
wake_up(&priv->tstamp_busy_wait);
return;
}
@@ -235,6 +236,13 @@ static void timestamp_interrupt(struct stmmac_priv *priv)
*/
ts_status = readl(priv->ioaddr + GMAC_TIMESTAMP_STATUS);
+ /* The powered register interface can still signal a level interrupt
+ * while the PHC is unavailable. Acknowledge it, but do not report an
+ * event from a clock which is being reset or shut down.
+ */
+ if (READ_ONCE(priv->ptp_blocked))
+ return;
+
if (!(priv->plat->flags & STMMAC_FLAG_EXT_SNAPSHOT_EN))
return;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 212bca73b9cf..ad5ed2c95af7 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -39,6 +39,7 @@
#endif /* CONFIG_DEBUG_FS */
#include <linux/net_tstamp.h>
#include <linux/phylink.h>
+#include <linux/pci.h>
#include <linux/udp.h>
#include <linux/bpf_trace.h>
#include <net/devlink.h>
@@ -656,6 +657,9 @@ static int stmmac_hwtstamp_set(struct net_device *dev,
u32 ts_master_en = 0;
u32 ts_event_en = 0;
+ if (priv->hw_unavailable)
+ return -EHOSTDOWN;
+
if (!priv->plat->clk_ptp_rate ||
!(priv->dma_cap.time_stamp || priv->adv_ts)) {
NL_SET_ERR_MSG_MOD(extack, "No support for HW time stamping");
@@ -961,7 +965,10 @@ static int stmmac_setup_ptp(struct stmmac_priv *priv)
priv->ptp_extts = 0;
priv->ptp_blocked = false;
+ mutex_lock(&priv->pm_mutex);
ret = clk_prepare_enable(priv->plat->clk_ptp_ref);
+ priv->ptp_clock_enabled = !ret;
+ mutex_unlock(&priv->pm_mutex);
if (ret < 0) {
netdev_warn(priv->dev,
"failed to enable PTP reference clock: %pe\n",
@@ -970,20 +977,25 @@ static int stmmac_setup_ptp(struct stmmac_priv *priv)
}
if (stmmac_init_ptp_clk_freq(priv)) {
- clk_disable_unprepare(priv->plat->clk_ptp_ref);
- return 0;
+ ret = 0;
+ goto disable_clock;
}
ret = stmmac_init_timestamping(priv);
- if (ret) {
- clk_disable_unprepare(priv->plat->clk_ptp_ref);
- return ret;
- }
+ if (ret)
+ goto disable_clock;
stmmac_ptp_register(priv);
priv->ptp_enabled = true;
return 0;
+
+disable_clock:
+ mutex_lock(&priv->pm_mutex);
+ clk_disable_unprepare(priv->plat->clk_ptp_ref);
+ priv->ptp_clock_enabled = false;
+ mutex_unlock(&priv->pm_mutex);
+ return ret;
}
static void stmmac_release_ptp(struct stmmac_priv *priv)
@@ -992,8 +1004,27 @@ static void stmmac_release_ptp(struct stmmac_priv *priv)
return;
stmmac_ptp_unregister(priv);
- clk_disable_unprepare(priv->plat->clk_ptp_ref);
priv->ptp_enabled = false;
+ mutex_lock(&priv->pm_mutex);
+ if (priv->ptp_clock_enabled) {
+ clk_disable_unprepare(priv->plat->clk_ptp_ref);
+ priv->ptp_clock_enabled = false;
+ }
+ /* A later noirq resume must not reacquire a released reference. */
+ priv->ptp_clock_suspended = false;
+ mutex_unlock(&priv->pm_mutex);
+}
+
+/* ptp_mutex excludes configuration and crosstimestamp operations. The
+ * spinlock also excludes atomic clock reads while changing this gate.
+ */
+static void stmmac_block_ptp(struct stmmac_priv *priv, bool block)
+{
+ unsigned long flags;
+
+ write_lock_irqsave(&priv->ptp_lock, flags);
+ WRITE_ONCE(priv->ptp_blocked, block);
+ write_unlock_irqrestore(&priv->ptp_lock, flags);
}
static void stmmac_legacy_serdes_power_down(struct stmmac_priv *priv)
@@ -1098,6 +1129,9 @@ static void stmmac_mac_link_down(struct phylink_config *config,
{
struct stmmac_priv *priv = netdev_priv(to_net_dev(config->dev));
+ if (READ_ONCE(priv->hw_unavailable))
+ return;
+
stmmac_mac_set(priv, priv->ioaddr, false);
if (priv->dma_cap.eee)
stmmac_set_eee_pls(priv, priv->hw, false);
@@ -1222,10 +1256,12 @@ static void stmmac_mac_disable_tx_lpi(struct phylink_config *config)
netdev_dbg(priv->dev, "disable EEE\n");
priv->eee_sw_timer_en = false;
timer_delete_sync(&priv->eee_ctrl_timer);
- stmmac_set_lpi_mode(priv, priv->hw, STMMAC_LPI_DISABLE, false, 0);
priv->tx_path_in_lpi_mode = false;
- stmmac_set_eee_timer(priv, priv->hw, 0, STMMAC_DEFAULT_TWT_LS);
+ if (!READ_ONCE(priv->hw_unavailable)) {
+ stmmac_set_lpi_mode(priv, priv->hw, STMMAC_LPI_DISABLE, false, 0);
+ stmmac_set_eee_timer(priv, priv->hw, 0, STMMAC_DEFAULT_TWT_LS);
+ }
mutex_unlock(&priv->lock);
}
@@ -3849,10 +3885,6 @@ static int stmmac_hw_setup(struct net_device *dev)
stmmac_enable_tbs(priv, priv->ioaddr, enable, chan);
}
- /* Configure real RX and TX queues */
- netif_set_real_num_rx_queues(dev, priv->plat->rx_queues_to_use);
- netif_set_real_num_tx_queues(dev, priv->plat->tx_queues_to_use);
-
/* Start the ball rolling... */
stmmac_start_all_dma(priv);
@@ -4140,6 +4172,16 @@ static int stmmac_request_irq_single(struct net_device *dev)
return ret;
}
+static void stmmac_unmask_pci_irq(struct stmmac_priv *priv)
+{
+#ifdef CONFIG_PCI
+ if (priv->pci_intx_masked) {
+ priv->pci_intx_masked = false;
+ pci_intx(to_pci_dev(priv->device), true);
+ }
+#endif
+}
+
static int stmmac_request_irq(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
@@ -4151,9 +4193,38 @@ static int stmmac_request_irq(struct net_device *dev)
else
ret = stmmac_request_irq_single(dev);
+ if (!ret)
+ stmmac_unmask_pci_irq(priv);
+
return ret;
}
+/* Drain registered handlers without disabling lines shared by other devices. */
+static void stmmac_synchronize_irq(struct stmmac_priv *priv)
+{
+ struct stmmac_msi *msi = priv->msi;
+ int irq = priv->dev->irq;
+ u32 i;
+
+ synchronize_irq(irq);
+ if (priv->wol_irq > 0 && priv->wol_irq != irq)
+ synchronize_irq(priv->wol_irq);
+ if (priv->sfty_irq > 0 && priv->sfty_irq != irq)
+ synchronize_irq(priv->sfty_irq);
+ if (!msi)
+ return;
+ if (msi->sfty_ce_irq > 0 && msi->sfty_ce_irq != irq)
+ synchronize_irq(msi->sfty_ce_irq);
+ if (msi->sfty_ue_irq > 0 && msi->sfty_ue_irq != irq)
+ synchronize_irq(msi->sfty_ue_irq);
+ for (i = 0; i < priv->plat->rx_queues_to_use; i++)
+ if (msi->rx_irq[i] > 0)
+ synchronize_irq(msi->rx_irq[i]);
+ for (i = 0; i < priv->plat->tx_queues_to_use; i++)
+ if (msi->tx_irq[i] > 0)
+ synchronize_irq(msi->tx_irq[i]);
+}
+
/**
* stmmac_setup_dma_desc - Generate a dma_conf and allocate DMA queue
* @priv: driver private structure
@@ -4232,6 +4303,97 @@ stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu)
return ERR_PTR(ret);
}
+/* The freezer excludes pool teardown and userspace from noirq callbacks.
+ * Outside system sleep these flags are changed under RTNL.
+ */
+int stmmac_resume_clocks(struct stmmac_priv *priv)
+{
+ int ret = 0;
+
+ mutex_lock(&priv->pm_mutex);
+ if (priv->bus_clks_suspended) {
+ ret = pm_runtime_force_resume(priv->device);
+ if (ret)
+ goto out;
+ priv->bus_clks_suspended = false;
+ }
+ if (priv->ptp_clock_suspended) {
+ ret = clk_prepare_enable(priv->plat->clk_ptp_ref);
+ if (ret)
+ goto out;
+ priv->ptp_clock_enabled = true;
+ priv->ptp_clock_suspended = false;
+ }
+out:
+ mutex_unlock(&priv->pm_mutex);
+ return ret;
+}
+EXPORT_SYMBOL_GPL(stmmac_resume_clocks);
+
+static int stmmac_resume_power(struct stmmac_priv *priv, bool system_resume)
+{
+ bool pending = priv->bus_clks_suspended || priv->bsp_suspended ||
+ priv->ptp_clock_suspended;
+ int ret;
+
+ ret = stmmac_resume_clocks(priv);
+ if (ret)
+ return ret;
+ if (priv->wol_suspended) {
+ priv->plat->resume_wol(priv->device, priv->plat->bsp_priv);
+ priv->wol_suspended = false;
+ }
+ if (priv->bsp_suspended && priv->plat->resume) {
+ ret = priv->plat->resume(priv->device, priv->plat->bsp_priv);
+ if (ret)
+ return ret;
+ priv->bsp_suspended = false;
+ }
+ /* A wrapper which powers the device outside these callbacks must
+ * complete its own resume before core register access is possible.
+ */
+ if (system_resume || pending)
+ WRITE_ONCE(priv->hw_unavailable, false);
+ return priv->hw_unavailable ? -EHOSTDOWN : 0;
+}
+
+/* Finish core sleep state only after its power dependencies are restored. */
+static int stmmac_resume_hw(struct stmmac_priv *priv)
+{
+ int ret;
+
+ if (!priv->hw_suspended)
+ return 0;
+ if (priv->hw_unavailable)
+ return -EHOSTDOWN;
+
+ /* Use the state installed by suspend, not a subsequently changed WoL
+ * setting. Clear PMT even when a different device caused the wakeup.
+ */
+ if (priv->irq_wake) {
+ mutex_lock(&priv->lock);
+ stmmac_pmt(priv, priv->hw, 0);
+ mutex_unlock(&priv->lock);
+ WRITE_ONCE(priv->irq_wake, 0);
+ /* Finish any timestamp handler admitted by the retained wake
+ * clocks before resume resets the PHC or close unregisters it.
+ */
+ stmmac_synchronize_irq(priv);
+ } else {
+ ret = pinctrl_pm_select_default_state(priv->device);
+ if (ret)
+ return ret;
+ if (priv->mii) {
+ ret = stmmac_mdio_reset(priv->mii);
+ if (ret)
+ return ret;
+ }
+ }
+ priv->hw_suspended = false;
+
+ return 0;
+}
+
/**
* __stmmac_open - open entry point of the driver
* @dev : pointer to the device structure.
@@ -4266,12 +4428,12 @@ static int __stmmac_open(struct net_device *dev,
ret = stmmac_hw_setup(dev);
if (ret < 0) {
netdev_err(priv->dev, "%s: Hw setup failed\n", __func__);
- return ret;
+ goto init_error;
}
ret = stmmac_setup_ptp(priv);
if (ret)
- goto ptp_error;
+ goto init_error;
/* The core soft reset in stmmac_hw_setup() clears the MTL_EST
* registers, so re-apply the taprio offload after PTP is up.
@@ -4282,34 +4444,43 @@ static int __stmmac_open(struct net_device *dev,
stmmac_init_coalesce(priv);
- phylink_start(priv->phylink);
-
stmmac_vlan_restore(priv);
ret = stmmac_request_irq(dev);
if (ret)
goto irq_error;
+ /* Publish the topology only when no other fallible setup remains.
+ * The combined setter restores the old counts if an increase fails.
+ */
+ ret = netif_set_real_num_queues(dev, priv->plat->tx_queues_to_use,
+ priv->plat->rx_queues_to_use);
+ if (ret) {
+ stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
+ goto irq_error;
+ }
+
stmmac_enable_all_queues(priv);
netif_tx_start_all_queues(priv->dev);
stmmac_enable_all_dma_irq(priv);
+ priv->datapath = STMMAC_DATAPATH_RUNNING;
+ phylink_start(priv->phylink);
return 0;
irq_error:
- phylink_stop(priv->phylink);
-
- stmmac_stop_all_dma(priv);
-
for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
hrtimer_cancel(&priv->dma_conf->tx_queue[chan].txtimer);
est_error:
stmmac_release_ptp(priv);
-ptp_error:
+init_error:
+ /* Undo phylink_prepare_resume() even if hardware setup failed before
+ * phylink_start(). Keep the PHY attachment and outer PM ownership.
+ */
+ phylink_stop(priv->phylink);
stmmac_stop_all_dma(priv);
stmmac_mac_set(priv, priv->ioaddr, false);
-
return ret;
}
@@ -4332,6 +4503,13 @@ static int stmmac_open(struct net_device *dev)
if (ret < 0)
goto err_dma_resources;
+ ret = stmmac_resume_power(priv, false);
+ if (ret)
+ goto err_runtime_pm;
+ ret = stmmac_resume_hw(priv);
+ if (ret)
+ goto err_runtime_pm;
+
ret = stmmac_init_phy(dev);
if (ret)
goto err_runtime_pm;
@@ -4395,23 +4573,36 @@ static void __stmmac_release(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ /* A failed MTU reopen has already released the data path. */
+ if (priv->datapath == STMMAC_DATAPATH_DOWN)
+ return;
+
phylink_stop(priv->phylink);
- stmmac_quiesce(priv);
+
+ /* Suspend retains the resources, but has already stopped activity. */
+ if (priv->datapath == STMMAC_DATAPATH_RUNNING)
+ stmmac_quiesce(priv);
+ priv->datapath = STMMAC_DATAPATH_DOWN;
/* Free the IRQ lines */
stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
+
/* TX error IRQs can restart a queue after the first quiescence. */
stmmac_stop_tx_queues(priv);
- /* Stop TX/RX DMA and clear the descriptors */
- stmmac_stop_all_dma(priv);
+ /* Stop TX/RX DMA after draining IRQ handlers which can restart it. */
+ if (!priv->hw_unavailable) {
+ stmmac_stop_all_dma(priv);
+ /* Link resolution need not have reached mac_link_up() yet. */
+ stmmac_mac_set(priv, priv->ioaddr, false);
+ }
/* Release and free the Rx/Tx resources */
free_dma_desc_resources(priv, priv->dma_conf);
stmmac_release_ptp(priv);
- if (stmmac_fpe_supported(priv))
+ if (!priv->hw_unavailable && stmmac_fpe_supported(priv))
ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
}
@@ -4424,6 +4615,17 @@ static void __stmmac_release(struct net_device *dev)
static int stmmac_release(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ int ret;
+
+ /* Resume may have failed before restoring pins or disabling MAC wake.
+ * Complete that cleanup without restarting the link or the datapath.
+ * If it fails, keep hw_suspended set so a fresh open can retry it.
+ */
+ ret = stmmac_resume_power(priv, false);
+ if (!ret)
+ ret = stmmac_resume_hw(priv);
+ if (ret)
+ netdev_err(dev, "failed to restore hardware sleep state: %d\n", ret);
/* If the PHY or MAC has WoL enabled, then the PHY will not be
* suspended when phylink_stop() is called below. Set the PHY
@@ -4437,6 +4639,8 @@ static int stmmac_release(struct net_device *dev)
stmmac_legacy_serdes_power_down(priv);
phylink_disconnect_phy(priv->phylink);
pm_runtime_put(priv->device);
+ /* Allow a fresh open after a failed MTU reopen or resume. */
+ netif_device_attach(dev);
return 0;
}
@@ -6266,6 +6470,9 @@ static void stmmac_set_rx_mode(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ if (READ_ONCE(priv->hw_unavailable))
+ return;
+
stmmac_set_filter(priv, priv->hw, dev);
}
@@ -6328,6 +6535,11 @@ static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
priv->dma_conf = old_conf;
free_dma_desc_resources(priv, dma_conf);
kfree(dma_conf);
+ /*
+ * Keep the administrative state and PHY/PM ownership until
+ * ndo_stop(), but prevent use of the released data path.
+ */
+ netif_device_detach(dev);
netdev_err(priv->dev, "failed reopening the interface after MTU change\n");
return ret;
}
@@ -6371,6 +6583,9 @@ static int stmmac_set_features(struct net_device *netdev,
{
struct stmmac_priv *priv = netdev_priv(netdev);
+ if (priv->hw_unavailable)
+ return -EHOSTDOWN;
+
/* Keep the COE Type in case of csum is supporting */
if (features & NETIF_F_RXCSUM)
priv->hw->rx_csum = priv->plat->rx_coe;
@@ -6412,7 +6627,7 @@ static void stmmac_common_interrupt(struct stmmac_priv *priv)
xmac = dwmac_is_xmac(priv->plat->core_type);
queues_count = (rx_cnt > tx_cnt) ? rx_cnt : tx_cnt;
- if (priv->irq_wake)
+ if (READ_ONCE(priv->irq_wake))
pm_wakeup_event(priv->device, 0);
if (priv->dma_cap.estsel)
@@ -6437,8 +6652,10 @@ static void stmmac_common_interrupt(struct stmmac_priv *priv)
for (queue = 0; queue < queues_count; queue++)
stmmac_host_mtl_irq_status(priv, priv->hw, queue);
- if (!READ_ONCE(priv->ptp_blocked))
- stmmac_timestamp_interrupt(priv, priv);
+ /* Powered handlers must acknowledge timestamp sources even while
+ * PHC access is blocked. The callback then skips event delivery.
+ */
+ stmmac_timestamp_interrupt(priv, priv);
}
}
@@ -6458,6 +6675,12 @@ static irqreturn_t stmmac_interrupt(int irq, void *dev_id)
struct net_device *dev = (struct net_device *)dev_id;
struct stmmac_priv *priv = netdev_priv(dev);
+ if (READ_ONCE(priv->hw_unavailable)) {
+ if (priv->irq_wake)
+ pm_wakeup_event(priv->device, 0);
+ return IRQ_NONE;
+ }
+
/* Check if adapter is up */
if (test_bit(STMMAC_DOWN, &priv->state))
return IRQ_HANDLED;
@@ -6480,6 +6703,12 @@ static irqreturn_t stmmac_mac_interrupt(int irq, void *dev_id)
struct net_device *dev = (struct net_device *)dev_id;
struct stmmac_priv *priv = netdev_priv(dev);
+ if (READ_ONCE(priv->hw_unavailable)) {
+ if (priv->irq_wake)
+ pm_wakeup_event(priv->device, 0);
+ return IRQ_NONE;
+ }
+
/* Check if adapter is up */
if (test_bit(STMMAC_DOWN, &priv->state))
return IRQ_HANDLED;
@@ -6495,6 +6724,9 @@ static irqreturn_t stmmac_safety_interrupt(int irq, void *dev_id)
struct net_device *dev = (struct net_device *)dev_id;
struct stmmac_priv *priv = netdev_priv(dev);
+ if (READ_ONCE(priv->hw_unavailable))
+ return IRQ_NONE;
+
/* Check if adapter is up */
if (test_bit(STMMAC_DOWN, &priv->state))
return IRQ_HANDLED;
@@ -6512,6 +6744,9 @@ static irqreturn_t stmmac_msi_intr_tx(int irq, void *data)
int chan = ch->index;
int status;
+ if (READ_ONCE(priv->hw_unavailable))
+ return IRQ_NONE;
+
/* Check if adapter is up */
if (test_bit(STMMAC_DOWN, &priv->state))
return IRQ_HANDLED;
@@ -6534,6 +6769,9 @@ static irqreturn_t stmmac_msi_intr_rx(int irq, void *data)
struct stmmac_priv *priv = ch->priv_data;
int chan = ch->index;
+ if (READ_ONCE(priv->hw_unavailable))
+ return IRQ_NONE;
+
/* Check if adapter is up */
if (test_bit(STMMAC_DOWN, &priv->state))
return IRQ_HANDLED;
@@ -6565,13 +6803,40 @@ static int stmmac_ioctl(struct net_device *dev, struct ifreq *rq, int cmd)
static int stmmac_setup_tc_block_cb(enum tc_setup_type type, void *type_data,
void *cb_priv)
{
+ struct flow_cls_common_offload *common = type_data;
struct stmmac_priv *priv = cb_priv;
+ bool active = stmmac_tc_active(priv);
int ret = -EOPNOTSUPP;
+ bool destroy;
- if (!tc_cls_can_offload_and_chain0(priv->dev, type_data))
+ switch (type) {
+ case TC_SETUP_CLSU32:
+ destroy = ((struct tc_cls_u32_offload *)type_data)->command ==
+ TC_CLSU32_DELETE_KNODE;
+ break;
+ case TC_SETUP_CLSFLOWER:
+ destroy = ((struct flow_cls_offload *)type_data)->command ==
+ FLOW_CLS_DESTROY;
+ break;
+ default:
return ret;
+ }
- __stmmac_disable_all_queues(priv);
+ if (common->chain_index) {
+ NL_SET_ERR_MSG(common->extack, "Driver supports only offload of chain 0");
+ return ret;
+ }
+ if (!destroy && !tc_can_offload_extack(priv->dev, common->extack))
+ return ret;
+ if (!destroy && !active)
+ return -ENETDOWN;
+
+ /* TC discards deleted filters regardless of the callback result. Clear
+ * their cached state even while down or detached, without touching NAPI
+ * or inaccessible registers. Recovery must not replay deleted rules.
+ */
+ if (active)
+ __stmmac_disable_all_queues(priv);
switch (type) {
case TC_SETUP_CLSU32:
@@ -6584,7 +6849,8 @@ static int stmmac_setup_tc_block_cb(enum tc_setup_type type, void *type_data,
break;
}
- stmmac_enable_all_queues(priv);
+ if (active)
+ stmmac_enable_all_queues(priv);
return ret;
}
@@ -6621,6 +6887,9 @@ static int stmmac_set_mac_address(struct net_device *ndev, void *addr)
struct stmmac_priv *priv = netdev_priv(ndev);
int ret = 0;
+ if (priv->hw_unavailable)
+ return -EHOSTDOWN;
+
ret = pm_runtime_resume_and_get(priv->device);
if (ret < 0)
return ret;
@@ -6672,8 +6941,9 @@ static int stmmac_rings_status_show(struct seq_file *seq, void *v)
u8 rx_count, tx_count, queue;
rtnl_lock();
- if ((dev->flags & IFF_UP) == 0)
+ if (priv->datapath == STMMAC_DATAPATH_DOWN)
goto out_unlock;
+
rx_count = priv->plat->rx_queues_to_use;
tx_count = priv->plat->tx_queues_to_use;
@@ -6999,7 +7269,7 @@ static int stmmac_vlan_update(struct stmmac_priv *priv, bool is_double)
hash = 0;
}
- if (!netif_running(priv->dev))
+ if (!netif_running(priv->dev) || priv->hw_unavailable)
return 0;
return stmmac_update_vlan_hash(priv, priv->hw, hash, pmatch, is_double);
@@ -7015,6 +7285,9 @@ static int stmmac_vlan_rx_add_vid(struct net_device *ndev, __be16 proto, u16 vid
bool is_double = false;
int ret;
+ if (priv->hw_unavailable)
+ return -EHOSTDOWN;
+
ret = pm_runtime_resume_and_get(priv->device);
if (ret < 0)
return ret;
@@ -7057,9 +7330,12 @@ static int stmmac_vlan_rx_kill_vid(struct net_device *ndev, __be16 proto, u16 vi
bool is_double = false;
int ret;
- ret = pm_runtime_resume_and_get(priv->device);
- if (ret < 0)
- return ret;
+ /* Removal must update the cached filters even if power cannot return. */
+ if (!priv->hw_unavailable) {
+ ret = pm_runtime_resume_and_get(priv->device);
+ if (ret < 0)
+ return ret;
+ }
if (be16_to_cpu(proto) == ETH_P_8021AD)
is_double = true;
@@ -7084,7 +7360,8 @@ static int stmmac_vlan_rx_kill_vid(struct net_device *ndev, __be16 proto, u16 vi
priv->num_double_vlans = num_double_vlans;
del_vlan_error:
- pm_runtime_put(priv->device);
+ if (!priv->hw_unavailable)
+ pm_runtime_put(priv->device);
return ret;
}
@@ -7104,6 +7381,18 @@ static int stmmac_bpf(struct net_device *dev, struct netdev_bpf *bpf)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ if (bpf->command != XDP_SETUP_PROG &&
+ bpf->command != XDP_SETUP_XSK_POOL)
+ return -EOPNOTSUPP;
+
+ /*
+ * Pool removal must succeed even after a failed resume. Release the
+ * suspended rings before their pool or XDP buffer layout can change.
+ * Leave the interface detached until it is closed and reopened.
+ */
+ if (priv->datapath == STMMAC_DATAPATH_SUSPENDED)
+ __stmmac_release(dev);
+
switch (bpf->command) {
case XDP_SETUP_PROG:
return stmmac_xdp_set_prog(priv, bpf->prog, bpf->extack);
@@ -7266,7 +7555,13 @@ void stmmac_xdp_release(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ netif_device_detach(dev);
+ /* Pause MAC/PCS resolution, but keep the PHY and negotiation running.
+ * stmmac_xdp_open() completes this replay under the same RTNL lock.
+ */
+ phylink_replay_link_begin(priv->phylink);
stmmac_quiesce(priv);
+ priv->datapath = STMMAC_DATAPATH_DOWN;
/* Free the IRQ lines */
stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
@@ -7285,7 +7580,13 @@ void stmmac_xdp_release(struct net_device *dev)
* watchdogs during reset
*/
netif_trans_update(dev);
- netif_carrier_off(dev);
+
+ if (stmmac_fpe_supported(priv))
+ ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
+
+ /* Keep PTP across the immediately following stmmac_xdp_open(). That
+ * function releases it if reopening fails, before returning DOWN.
+ */
}
int stmmac_xdp_open(struct net_device *dev)
@@ -7364,21 +7665,25 @@ int stmmac_xdp_open(struct net_device *dev)
/* Enable NAPI process*/
stmmac_enable_all_queues(priv);
- netif_carrier_on(dev);
- netif_tx_start_all_queues(dev);
stmmac_enable_all_dma_irq(priv);
+ priv->datapath = STMMAC_DATAPATH_RUNNING;
+ phylink_replay_link_end(priv->phylink);
+ netif_device_attach(dev);
return 0;
irq_error:
+ stmmac_stop_tx_queues(priv);
stmmac_stop_all_dma(priv);
-
- for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
- hrtimer_cancel(&priv->dma_conf->tx_queue[chan].txtimer);
+ stmmac_mac_set(priv, priv->ioaddr, false);
init_error:
free_dma_desc_resources(priv, priv->dma_conf);
dma_desc_error:
+ /* STOPPED keeps replay_end() from reconfiguring a failed MAC. */
+ phylink_stop(priv->phylink);
+ phylink_replay_link_end(priv->phylink);
+ stmmac_release_ptp(priv);
return ret;
}
@@ -7500,6 +7805,9 @@ static void stmmac_reset_subtask(struct stmmac_priv *priv)
netdev_err(priv->dev, "Reset adapter.\n");
rtnl_lock();
+ if (!netif_device_present(priv->dev))
+ goto out_unlock;
+
netif_trans_update(priv->dev);
while (test_and_set_bit(STMMAC_RESETING, &priv->state))
usleep_range(1000, 2000);
@@ -7509,6 +7817,7 @@ static void stmmac_reset_subtask(struct stmmac_priv *priv)
dev_open(priv->dev, NULL);
clear_bit(STMMAC_DOWN, &priv->state);
clear_bit(STMMAC_RESETING, &priv->state);
+out_unlock:
rtnl_unlock();
}
@@ -7780,13 +8089,37 @@ static void stmmac_napi_del(struct net_device *dev)
}
}
-int stmmac_reinit_queues(struct net_device *dev, u8 rx_cnt, u8 tx_cnt)
+/* Rebuild only the datapath. The administratively-up device still owns its
+ * PHY attachment and runtime-PM reference, even if this reopen fails.
+ */
+static int stmmac_reopen(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
- int ret = 0, i;
+ struct stmmac_dma_conf *old_conf = priv->dma_conf;
+ struct stmmac_dma_conf *dma_conf;
+ int ret;
- if (netif_running(dev))
- stmmac_release(dev);
+ dma_conf = stmmac_setup_dma_desc(priv, dev->mtu);
+ if (IS_ERR(dma_conf))
+ return PTR_ERR(dma_conf);
+
+ ret = __stmmac_open(dev, dma_conf);
+ if (ret) {
+ priv->dma_conf = old_conf;
+ free_dma_desc_resources(priv, dma_conf);
+ kfree(dma_conf);
+ return ret;
+ }
+
+ kfree(old_conf);
+ netif_device_attach(dev);
+ return 0;
+}
+
+static void stmmac_set_queues(struct net_device *dev, u8 rx_cnt, u8 tx_cnt)
+{
+ struct stmmac_priv *priv = netdev_priv(dev);
+ int i;
stmmac_napi_del(dev);
@@ -7798,9 +8131,31 @@ int stmmac_reinit_queues(struct net_device *dev, u8 rx_cnt, u8 tx_cnt)
rx_cnt);
stmmac_napi_add(dev);
+}
+
+int stmmac_reinit_queues(struct net_device *dev, u8 rx_cnt, u8 tx_cnt)
+{
+ struct stmmac_priv *priv = netdev_priv(dev);
+ u8 old_rx = priv->plat->rx_queues_to_use;
+ u8 old_tx = priv->plat->tx_queues_to_use;
+ int ret = 0;
+
+ if (netif_running(dev)) {
+ if (!netif_device_present(dev))
+ return -ENETDOWN;
+ netif_device_detach(dev);
+ __stmmac_release(dev);
+ }
+
+ stmmac_set_queues(dev, rx_cnt, tx_cnt);
if (netif_running(dev))
- ret = stmmac_open(dev);
+ ret = stmmac_reopen(dev);
+ if (ret) {
+ stmmac_set_queues(dev, old_rx, old_tx);
+ netdev_err(dev, "failed reopening after channel change: %pe; interface remains detached\n",
+ ERR_PTR(ret));
+ }
return ret;
}
@@ -7808,16 +8163,28 @@ int stmmac_reinit_queues(struct net_device *dev, u8 rx_cnt, u8 tx_cnt)
int stmmac_reinit_ringparam(struct net_device *dev, u32 rx_size, u32 tx_size)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ u32 old_rx = priv->dma_conf->dma_rx_size;
+ u32 old_tx = priv->dma_conf->dma_tx_size;
int ret = 0;
- if (netif_running(dev))
- stmmac_release(dev);
+ if (netif_running(dev)) {
+ if (!netif_device_present(dev))
+ return -ENETDOWN;
+ netif_device_detach(dev);
+ __stmmac_release(dev);
+ }
priv->dma_conf->dma_rx_size = rx_size;
priv->dma_conf->dma_tx_size = tx_size;
if (netif_running(dev))
- ret = stmmac_open(dev);
+ ret = stmmac_reopen(dev);
+ if (ret) {
+ priv->dma_conf->dma_rx_size = old_rx;
+ priv->dma_conf->dma_tx_size = old_tx;
+ netdev_err(dev, "failed reopening after ring change: %pe; interface remains detached\n",
+ ERR_PTR(ret));
+ }
return ret;
}
@@ -8234,6 +8601,7 @@ static int __stmmac_dvr_probe(struct device *device,
mutex_init(&priv->lock);
mutex_init(&priv->est_lock);
+ mutex_init(&priv->pm_mutex);
rwlock_init(&priv->ptp_lock);
mutex_init(&priv->ptp_mutex);
@@ -8379,41 +8747,71 @@ void stmmac_dvr_remove(struct device *dev)
}
EXPORT_SYMBOL_GPL(stmmac_dvr_remove);
+static void stmmac_block_mmio(struct stmmac_priv *priv)
+{
+ /* Drain MDIO transactions and receive-filter updates before removing
+ * register power. Both paths can run independently of RTNL.
+ */
+ if (priv->mii)
+ mutex_lock(&priv->mii->mdio_lock);
+ netif_addr_lock_bh(priv->dev);
+ WRITE_ONCE(priv->hw_unavailable, true);
+ netif_addr_unlock_bh(priv->dev);
+ if (priv->mii)
+ mutex_unlock(&priv->mii->mdio_lock);
+}
+
/**
- * stmmac_suspend - suspend callback
+ * __stmmac_suspend - suspend callback
* @dev: device pointer
- * Description: this is the function to suspend the device and it is called
- * by the platform driver to stop the network queue, release the resources,
- * program the PMT register (for WoL), clean and release driver resources.
+ * @bsp_noirq: defer PCI power removal until IRQ handlers are suspended
+ * Description: stop network activity and program hardware for system sleep,
+ * preserving any datapath resources still owned for resume or close.
*/
-int stmmac_suspend(struct device *dev)
+static int __stmmac_suspend(struct device *dev, bool bsp_noirq)
{
struct net_device *ndev = dev_get_drvdata(dev);
struct stmmac_priv *priv = netdev_priv(ndev);
+ bool accessible;
+ int ret = 0;
- if (!ndev || !netif_running(ndev))
+ rtnl_lock();
+ if (!netif_running(ndev))
+ goto suspend_bsp;
+
+ /* A failed datapath cannot provide a working MAC wake path. It may
+ * even have released its wake IRQ. Do not silently suspend without WoL.
+ */
+ if (priv->wolopts && priv->datapath != STMMAC_DATAPATH_RUNNING) {
+ netdev_err(ndev, "cannot suspend failed datapath with MAC WoL enabled\n");
+ rtnl_unlock();
+ return -EBUSY;
+ }
+ if (priv->hw_suspended)
goto suspend_bsp;
mutex_lock(&priv->lock);
netif_device_detach(ndev);
- stmmac_quiesce(priv);
+ if (priv->datapath == STMMAC_DATAPATH_RUNNING)
+ stmmac_quiesce(priv);
if (priv->eee_sw_timer_en) {
priv->tx_path_in_lpi_mode = false;
timer_delete_sync(&priv->eee_ctrl_timer);
}
- /* Stop TX/RX DMA */
stmmac_stop_all_dma(priv);
- stmmac_legacy_serdes_power_down(priv);
+ if (!priv->wolopts)
+ stmmac_legacy_serdes_power_down(priv);
/* Enable Power down mode by programming the PMT regs */
if (priv->wolopts) {
+ /* An interrupt can arrive as soon as PMT is armed. */
+ WRITE_ONCE(priv->irq_wake, 1);
stmmac_pmt(priv, priv->hw, priv->wolopts);
- priv->irq_wake = 1;
} else {
stmmac_mac_set(priv, priv->ioaddr, false);
pinctrl_pm_select_sleep_state(priv->device);
@@ -8421,18 +8819,58 @@ int stmmac_suspend(struct device *dev)
mutex_unlock(&priv->lock);
- rtnl_lock();
phylink_suspend(priv->phylink, !!priv->wolopts);
- rtnl_unlock();
+ if (priv->datapath == STMMAC_DATAPATH_RUNNING)
+ priv->datapath = STMMAC_DATAPATH_SUSPENDED;
+ priv->hw_suspended = true;
if (stmmac_fpe_supported(priv))
ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
suspend_bsp:
- if (priv->plat->suspend)
- return priv->plat->suspend(dev, priv->plat->bsp_priv);
+ accessible = !priv->hw_unavailable;
+ mutex_lock(&priv->ptp_mutex);
+ stmmac_block_ptp(priv, true);
+ mutex_unlock(&priv->ptp_mutex);
+ /* A level PMT interrupt must remain serviceable before suspend_noirq
+ * and after resume_noirq. Platform MAC WoL retains power; PCI defers
+ * its power transition to the noirq callbacks instead.
+ */
+ if (!priv->irq_wake)
+ stmmac_block_mmio(priv);
+ if (priv->datapath == STMMAC_DATAPATH_SUSPENDED) {
+ stmmac_synchronize_irq(priv);
+ stmmac_stop_tx_queues(priv);
+ if (accessible)
+ stmmac_stop_all_dma(priv);
+ }
+ if (bsp_noirq)
+ goto out_unlock;
+ if (priv->irq_wake) {
+ if (priv->plat->suspend_wol && !priv->wol_suspended) {
+ ret = priv->plat->suspend_wol(dev, priv->plat->bsp_priv);
+ if (!ret)
+ priv->wol_suspended = true;
+ }
+ } else if (priv->plat->suspend && !priv->bsp_suspended) {
+ priv->bsp_suspended = true;
+ ret = priv->plat->suspend(dev, priv->plat->bsp_priv);
+ }
+out_unlock:
+ rtnl_unlock();
- return 0;
+ /* The PM core will not resume a device whose suspend callback failed.
+ * A failed rollback remains detached for ordinary down/up recovery.
+ */
+ if (ret && stmmac_resume(dev))
+ netdev_err(ndev, "failed to restore device after suspend failure\n");
+
+ return ret;
+}
+
+int stmmac_suspend(struct device *dev)
+{
+ return __stmmac_suspend(dev, false);
}
EXPORT_SYMBOL_GPL(stmmac_suspend);
@@ -8484,41 +8922,39 @@ int stmmac_resume(struct device *dev)
struct stmmac_priv *priv = netdev_priv(ndev);
int ret;
- if (priv->plat->resume) {
- ret = priv->plat->resume(dev, priv->plat->bsp_priv);
- if (ret)
- return ret;
+ rtnl_lock();
+ ret = stmmac_resume_power(priv, true);
+ if (ret)
+ goto out_unlock;
+
+ if (!netif_running(ndev)) {
+ ret = 0;
+ goto out_unlock;
}
- if (!netif_running(ndev))
- return 0;
+ if (priv->hw_suspended) {
+ ret = stmmac_resume_hw(priv);
+ if (ret)
+ goto out_unlock;
- /* Power Down bit, into the PM register, is cleared
- * automatically as soon as a magic packet or a Wake-up frame
- * is received. Anyway, it's better to manually clear
- * this bit because it can generate problems while resuming
- * from another devices (e.g. serial console).
- */
- if (priv->wolopts) {
- mutex_lock(&priv->lock);
- stmmac_pmt(priv, priv->hw, 0);
- mutex_unlock(&priv->lock);
- priv->irq_wake = 0;
- } else {
- pinctrl_pm_select_default_state(priv->device);
- /* reset the phy so that it's ready */
- if (priv->mii)
- stmmac_mdio_reset(priv->mii);
+ /* Terminate PM speed control without restarting a datapath
+ * whose IRQs or rings were released before system sleep.
+ */
+ if (priv->datapath != STMMAC_DATAPATH_SUSPENDED)
+ phylink_stop(priv->phylink);
+ }
+
+ if (priv->datapath != STMMAC_DATAPATH_SUSPENDED) {
+ ret = 0;
+ goto out_unlock;
}
if (!(priv->plat->flags & STMMAC_FLAG_SERDES_UP_AFTER_PHY_LINKUP)) {
ret = stmmac_legacy_serdes_power_up(priv);
if (ret < 0)
- return ret;
+ goto out_unlock;
}
- rtnl_lock();
-
/* Prepare the PHY to resume, ensuring that its clocks which are
* necessary for the MAC DMA reset to complete are running
*/
@@ -8534,7 +8970,7 @@ int stmmac_resume(struct device *dev)
ret = stmmac_hw_setup(ndev);
if (ret < 0) {
netdev_err(priv->dev, "%s: Hw setup failed\n", __func__);
- goto error_unlock;
+ goto error_stop_dma;
}
if (priv->ptp_enabled) {
@@ -8547,6 +8983,9 @@ int stmmac_resume(struct device *dev)
}
init_coalesce:
+ mutex_lock(&priv->ptp_mutex);
+ stmmac_block_ptp(priv, false);
+ mutex_unlock(&priv->ptp_mutex);
ret = stmmac_setup_est(priv);
if (ret < 0)
goto error_stop_dma;
@@ -8560,6 +8999,7 @@ int stmmac_resume(struct device *dev)
stmmac_enable_all_queues(priv);
stmmac_enable_all_dma_irq(priv);
+ stmmac_unmask_pci_irq(priv);
mutex_unlock(&priv->lock);
@@ -8568,18 +9008,25 @@ int stmmac_resume(struct device *dev)
* workqueue thread, which will race with initialisation.
*/
phylink_resume(priv->phylink);
- rtnl_unlock();
-
+ priv->datapath = STMMAC_DATAPATH_RUNNING;
netif_device_attach(ndev);
+ rtnl_unlock();
return 0;
error_stop_dma:
+ mutex_lock(&priv->ptp_mutex);
+ stmmac_block_ptp(priv, true);
+ mutex_unlock(&priv->ptp_mutex);
stmmac_stop_all_dma(priv);
stmmac_mac_set(priv, priv->ioaddr, false);
-error_unlock:
stmmac_legacy_serdes_power_down(priv);
mutex_unlock(&priv->lock);
+ /*
+ * Keep the suspended data path detached. A later resume may retry, or
+ * ndo_stop() can release its resources without disabling NAPI again.
+ */
+out_unlock:
rtnl_unlock();
return ret;
@@ -8590,6 +9037,62 @@ EXPORT_SYMBOL_GPL(stmmac_resume);
DEFINE_SIMPLE_DEV_PM_OPS(stmmac_simple_pm_ops, stmmac_suspend, stmmac_resume);
EXPORT_SYMBOL_GPL(stmmac_simple_pm_ops);
+/* PCI config-space power transitions do not need device IRQ handlers. Keep
+ * them inside the noirq window, including PME-based WoL from D3, so the MAC
+ * interrupt can always be acknowledged whenever its handler can run.
+ */
+static int __maybe_unused stmmac_pci_suspend(struct device *dev)
+{
+ return __stmmac_suspend(dev, true);
+}
+
+static int __maybe_unused stmmac_pci_suspend_noirq(struct device *dev)
+{
+ struct stmmac_priv *priv = netdev_priv(dev_get_drvdata(dev));
+ bool unavailable = priv->hw_unavailable;
+ int ret = 0;
+
+ stmmac_block_mmio(priv);
+ if (priv->plat->suspend && !priv->bsp_suspended) {
+ ret = priv->plat->suspend(dev, priv->plat->bsp_priv);
+ if (!ret)
+ priv->bsp_suspended = true;
+ else
+ WRITE_ONCE(priv->hw_unavailable, unavailable);
+ }
+ return ret;
+}
+
+static int __maybe_unused stmmac_pci_resume_noirq(struct device *dev)
+{
+ struct stmmac_priv *priv = netdev_priv(dev_get_drvdata(dev));
+ int ret;
+
+ ret = stmmac_resume_power(priv, true);
+#ifdef CONFIG_PCI
+ if (ret || priv->pci_intx_masked) {
+ struct pci_dev *pdev = to_pci_dev(dev);
+
+ /* Failed restoration must not expose a level PMT interrupt whose
+ * registers cannot be acknowledged. Mask only this PCI function,
+ * not the shared controller line, until the datapath recovers.
+ * MSI/MSI-X are message interrupts, not asserted level lines.
+ */
+ if (!pdev->msi_enabled && !pdev->msix_enabled) {
+ pci_intx(pdev, false);
+ priv->pci_intx_masked = true;
+ }
+ }
+#endif
+ return ret;
+}
+
+const struct dev_pm_ops stmmac_pci_pm_ops = {
+ SET_SYSTEM_SLEEP_PM_OPS(stmmac_pci_suspend, stmmac_resume)
+ SET_NOIRQ_SYSTEM_SLEEP_PM_OPS(stmmac_pci_suspend_noirq, stmmac_pci_resume_noirq)
+};
+EXPORT_SYMBOL_GPL(stmmac_pci_pm_ops);
+
#ifndef MODULE
static int __init stmmac_cmdline_opt(char *str)
{
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c
index 63287ad9652f..092b244aa04c 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_mdio.c
@@ -135,6 +135,9 @@ static int stmmac_xgmac2_mdio_read_c22(struct mii_bus *bus, int phyaddr,
struct stmmac_priv *priv = netdev_priv(bus->priv);
u32 addr;
+ if (READ_ONCE(priv->hw_unavailable))
+ return -EHOSTDOWN;
+
/* Until ver 2.20 XGMAC does not support C22 addr >= 4 */
if (priv->synopsys_id < DWXGMAC_CORE_2_20 &&
phyaddr > MII_XGMAC_MAX_C22ADDR)
@@ -151,6 +154,9 @@ static int stmmac_xgmac2_mdio_read_c45(struct mii_bus *bus, int phyaddr,
struct stmmac_priv *priv = netdev_priv(bus->priv);
u32 addr;
+ if (READ_ONCE(priv->hw_unavailable))
+ return -EHOSTDOWN;
+
stmmac_xgmac2_c45_format(priv, phyaddr, devad, phyreg, &addr);
return stmmac_xgmac2_mdio_read(priv, addr, MII_XGMAC_BUSY);
@@ -198,6 +204,9 @@ static int stmmac_xgmac2_mdio_write_c22(struct mii_bus *bus, int phyaddr,
struct stmmac_priv *priv = netdev_priv(bus->priv);
u32 addr;
+ if (READ_ONCE(priv->hw_unavailable))
+ return -EHOSTDOWN;
+
/* Until ver 2.20 XGMAC does not support C22 addr >= 4 */
if (priv->synopsys_id < DWXGMAC_CORE_2_20 &&
phyaddr > MII_XGMAC_MAX_C22ADDR)
@@ -215,6 +224,9 @@ static int stmmac_xgmac2_mdio_write_c45(struct mii_bus *bus, int phyaddr,
struct stmmac_priv *priv = netdev_priv(bus->priv);
u32 addr;
+ if (READ_ONCE(priv->hw_unavailable))
+ return -EHOSTDOWN;
+
stmmac_xgmac2_c45_format(priv, phyaddr, devad, phyreg, &addr);
return stmmac_xgmac2_mdio_write(priv, addr, MII_XGMAC_BUSY,
@@ -248,6 +260,9 @@ static int stmmac_mdio_access(struct stmmac_priv *priv, unsigned int pa,
u32 addr;
int ret;
+ if (READ_ONCE(priv->hw_unavailable))
+ return -EHOSTDOWN;
+
ret = pm_runtime_resume_and_get(priv->device);
if (ret < 0)
return ret;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_pci.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_pci.c
index d584fd2daa6f..c5fcd4e433a7 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_pci.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_pci.c
@@ -215,7 +215,7 @@ static struct pci_driver stmmac_pci_driver = {
.probe = stmmac_pci_probe,
.remove = stmmac_pci_remove,
.driver = {
- .pm = &stmmac_simple_pm_ops,
+ .pm = &stmmac_pci_pm_ops,
},
};
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_pcs.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_pcs.c
index df37af5ab837..7de83fb8f06f 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_pcs.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_pcs.c
@@ -60,7 +60,8 @@ static void dwmac_integrated_pcs_disable(struct phylink_pcs *pcs)
{
struct stmmac_pcs *spcs = phylink_pcs_to_stmmac_pcs(pcs);
- stmmac_mac_irq_modify(spcs->priv, spcs->int_mask, 0);
+ if (!READ_ONCE(spcs->priv->hw_unavailable))
+ stmmac_mac_irq_modify(spcs->priv, spcs->int_mask, 0);
}
static void dwmac_integrated_pcs_get_state(struct phylink_pcs *pcs,
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_platform.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_platform.c
index 3ea16d618215..2f98df3f7147 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_platform.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_platform.c
@@ -968,15 +968,23 @@ static int __maybe_unused stmmac_pltfr_noirq_suspend(struct device *dev)
if (!netif_running(ndev))
return 0;
- if (!priv->wolopts) {
- /* Disable clock in case of PWM is off */
- if (priv->ptp_enabled)
+ mutex_lock(&priv->pm_mutex);
+ if (!priv->irq_wake) {
+ /* A detached datapath may already have released its PTP clock. */
+ if (priv->ptp_clock_enabled) {
clk_disable_unprepare(priv->plat->clk_ptp_ref);
+ priv->ptp_clock_enabled = false;
+ priv->ptp_clock_suspended = true;
+ }
+ priv->bus_clks_suspended = true;
ret = pm_runtime_force_suspend(dev);
- if (ret)
+ if (ret) {
+ mutex_unlock(&priv->pm_mutex);
return ret;
+ }
}
+ mutex_unlock(&priv->pm_mutex);
return 0;
}
@@ -985,30 +993,8 @@ static int __maybe_unused stmmac_pltfr_noirq_resume(struct device *dev)
{
struct net_device *ndev = dev_get_drvdata(dev);
struct stmmac_priv *priv = netdev_priv(ndev);
- int ret;
- if (!netif_running(ndev))
- return 0;
-
- if (!priv->wolopts) {
- /* enable the clk previously disabled */
- ret = pm_runtime_force_resume(dev);
- if (ret)
- return ret;
-
- if (!priv->ptp_enabled)
- return 0;
-
- ret = clk_prepare_enable(priv->plat->clk_ptp_ref);
- if (ret < 0) {
- netdev_warn(priv->dev,
- "failed to enable PTP reference clock: %pe\n",
- ERR_PTR(ret));
- return ret;
- }
- }
-
- return 0;
+ return stmmac_resume_clocks(priv);
}
const struct dev_pm_ops stmmac_pltfr_pm_ops = {
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
index 839a93166520..cac2b6e34c94 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
@@ -12,6 +12,17 @@
#include "stmmac.h"
#include "stmmac_est.h"
+static int tc_config_preemption(struct stmmac_priv *priv,
+ struct netlink_ext_ack *extack, u32 preemptible_tcs)
+{
+ /* Qdisc teardown must not access unpowered registers. */
+ if (priv->hw_unavailable)
+ return 0;
+
+ return stmmac_fpe_map_preemption_class(priv, priv->dev, extack,
+ preemptible_tcs);
+}
+
static void tc_fill_all_pass_entry(struct stmmac_tc_entry *entry)
{
memset(entry, 0, sizeof(*entry));
@@ -172,18 +183,16 @@ static int tc_fill_entry(struct stmmac_priv *priv,
static void tc_unfill_entry(struct stmmac_priv *priv,
struct tc_cls_u32_offload *cls)
{
- struct stmmac_tc_entry *entry;
+ struct stmmac_tc_entry *entry, *frag;
entry = tc_find_entry(priv, cls, false);
if (!entry)
return;
- entry->in_use = false;
- if (entry->frag_ptr) {
- entry = entry->frag_ptr;
- entry->is_frag = false;
- entry->in_use = false;
- }
+ frag = entry->frag_ptr;
+ if (frag)
+ memset(frag, 0, sizeof(*frag));
+ memset(entry, 0, sizeof(*entry));
}
static int tc_config_knode(struct stmmac_priv *priv,
@@ -213,6 +222,9 @@ static int tc_delete_knode(struct stmmac_priv *priv,
/* Set entry and fragments as not used */
tc_unfill_entry(priv, cls);
+ if (!stmmac_tc_active(priv))
+ return 0;
+
return stmmac_rxp_config(priv, priv->hw->pcsr, priv->tc_entries,
priv->tc_entries_max);
}
@@ -346,6 +358,8 @@ static int tc_setup_cbs(struct stmmac_priv *priv,
return -EINVAL;
if (!priv->dma_cap.av)
return -EOPNOTSUPP;
+ if (priv->hw_unavailable && qopt->enable)
+ return -EHOSTDOWN;
port_transmit_rate_kbps = qopt->idleslope - qopt->sendslope;
@@ -381,10 +395,12 @@ static int tc_setup_cbs(struct stmmac_priv *priv,
priv->plat->tx_queues_cfg[queue].mode_to_use = MTL_QUEUE_AVB;
} else if (!qopt->enable) {
- ret = stmmac_dma_qmode(priv, priv->ioaddr, queue,
- MTL_QUEUE_DCB);
- if (ret)
- return ret;
+ if (!priv->hw_unavailable) {
+ ret = stmmac_dma_qmode(priv, priv->ioaddr, queue,
+ MTL_QUEUE_DCB);
+ if (ret)
+ return ret;
+ }
priv->plat->tx_queues_cfg[queue].mode_to_use = MTL_QUEUE_DCB;
return 0;
@@ -646,23 +662,21 @@ static int tc_del_flow(struct stmmac_priv *priv,
struct flow_cls_offload *cls)
{
struct stmmac_flow_entry *entry = tc_find_flow(priv, cls, false);
- int ret;
+ int ret = 0;
if (!entry || !entry->in_use)
return -ENOENT;
- if (entry->is_l4) {
- ret = stmmac_config_l4_filter(priv, priv->hw, entry->idx, false,
- false, false, false, 0);
- } else {
- ret = stmmac_config_l3_filter(priv, priv->hw, entry->idx, false,
- false, false, false, 0);
+ if (stmmac_tc_active(priv)) {
+ if (entry->is_l4)
+ ret = stmmac_config_l4_filter(priv, priv->hw, entry->idx,
+ false, false, false, false, 0);
+ else
+ ret = stmmac_config_l3_filter(priv, priv->hw, entry->idx,
+ false, false, false, false, 0);
}
- entry->in_use = false;
- entry->cookie = 0;
- entry->is_l4 = false;
- entry->action = 0;
+ *entry = (struct stmmac_flow_entry) { .idx = entry->idx };
return ret;
}
@@ -745,7 +759,8 @@ static int tc_del_vlan_flow(struct stmmac_priv *priv,
if (!entry || !entry->in_use || entry->type != STMMAC_RFS_T_VLAN)
return -ENOENT;
- stmmac_rx_queue_prio(priv, priv->hw, 0, entry->tc);
+ if (stmmac_tc_active(priv))
+ stmmac_rx_queue_prio(priv, priv->hw, 0, entry->tc);
entry->in_use = false;
entry->cookie = 0;
@@ -841,13 +856,13 @@ static int tc_del_ethtype_flow(struct stmmac_priv *priv,
switch (entry->etype) {
case ETH_P_LLDP:
- stmmac_rx_queue_routing(priv, priv->hw,
- PACKET_DCBCPQ, 0);
+ if (stmmac_tc_active(priv))
+ stmmac_rx_queue_routing(priv, priv->hw, PACKET_DCBCPQ, 0);
priv->rfs_entries_cnt[STMMAC_RFS_T_LLDP]--;
break;
case ETH_P_1588:
- stmmac_rx_queue_routing(priv, priv->hw,
- PACKET_PTPQ, 0);
+ if (stmmac_tc_active(priv))
+ stmmac_rx_queue_routing(priv, priv->hw, PACKET_PTPQ, 0);
priv->rfs_entries_cnt[STMMAC_RFS_T_1588]--;
break;
default:
@@ -902,12 +917,12 @@ static int tc_setup_cls(struct stmmac_priv *priv,
{
int ret = 0;
- /* When RSS is enabled, the filtering will be bypassed */
- if (priv->rss.enable)
- return -EBUSY;
-
switch (cls->command) {
case FLOW_CLS_REPLACE:
+ /* When RSS is enabled, the filtering will be bypassed. */
+ if (priv->rss.enable)
+ return -EBUSY;
+
ret = tc_add_flow_cls(priv, cls);
break;
case FLOW_CLS_DESTROY:
@@ -1021,6 +1036,10 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
if (qopt->cmd == TAPRIO_CMD_DESTROY)
goto disable;
+ if (priv->hw_unavailable) {
+ ret = -EHOSTDOWN;
+ goto unlock;
+ }
if (qopt->num_entries > dep) {
ret = -EINVAL;
@@ -1092,8 +1111,7 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
if (ret)
goto disable;
- ret = stmmac_fpe_map_preemption_class(priv, priv->dev, extack,
- qopt->mqprio.preemptible_tcs);
+ ret = tc_config_preemption(priv, extack, qopt->mqprio.preemptible_tcs);
if (ret)
goto disable;
@@ -1106,15 +1124,16 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
disable:
priv->est.enable = false;
- stmmac_est_configure(priv, priv, &priv->est,
- priv->plat->clk_ptp_rate, false);
+ if (!priv->hw_unavailable)
+ stmmac_est_configure(priv, priv, &priv->est,
+ priv->plat->clk_ptp_rate, false);
/* Reset taprio status */
for (i = 0; i < priv->plat->tx_queues_to_use; i++) {
priv->xstats.max_sdu_txq_drop[i] = 0;
priv->xstats.mtl_est_txq_hlbf[i] = 0;
priv->xstats.mtl_est_txq_hlbs[i] = 0;
}
- i = stmmac_fpe_map_preemption_class(priv, priv->dev, extack, 0);
+ i = tc_config_preemption(priv, extack, 0);
if (qopt->cmd == TAPRIO_CMD_DESTROY)
ret = i;
free_gcl:
@@ -1269,7 +1288,7 @@ static int stmmac_reset_tc_mqprio(struct net_device *ndev,
netdev_reset_tc(ndev);
netif_set_real_num_tx_queues(ndev, priv->plat->tx_queues_to_use);
- return stmmac_fpe_map_preemption_class(priv, ndev, extack, 0);
+ return tc_config_preemption(priv, extack, 0);
}
static int tc_setup_dwmac510_mqprio(struct stmmac_priv *priv,
@@ -1286,6 +1305,8 @@ static int tc_setup_dwmac510_mqprio(struct stmmac_priv *priv,
if (!qopt->num_tc)
return stmmac_reset_tc_mqprio(ndev, extack);
+ if (priv->hw_unavailable)
+ return -EHOSTDOWN;
if (qopt->num_tc > ARRAY_SIZE(tc_to_txq))
return -EINVAL;
@@ -1315,8 +1336,7 @@ static int tc_setup_dwmac510_mqprio(struct stmmac_priv *priv,
if (err)
goto error_reset_tc;
- err = stmmac_fpe_map_preemption_class(priv, ndev, extack,
- mqprio->preemptible_tcs);
+ err = tc_config_preemption(priv, extack, mqprio->preemptible_tcs);
if (err)
goto error_reset_num_tx_queues;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_vlan.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_vlan.c
index e24efe3bfedb..006882c14a8d 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_vlan.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_vlan.c
@@ -114,6 +114,8 @@ static int vlan_del_hw_rx_fltr(struct net_device *dev,
struct mac_device_info *hw,
__be16 proto, u16 vid)
{
+ struct stmmac_priv *priv = netdev_priv(dev);
+ bool update_hw = netif_running(dev) && !priv->hw_unavailable;
int i, ret = 0;
/* Single Rx VLAN Filter */
@@ -121,7 +123,7 @@ static int vlan_del_hw_rx_fltr(struct net_device *dev,
if ((hw->vlan_filter[0] & VLAN_TAG_VID) == vid) {
hw->vlan_filter[0] = 0;
- if (netif_running(dev))
+ if (update_hw)
vlan_write_single(dev, 0);
}
return 0;
@@ -132,7 +134,7 @@ static int vlan_del_hw_rx_fltr(struct net_device *dev,
if ((hw->vlan_filter[i] & VLAN_TAG_DATA_VEN) &&
((hw->vlan_filter[i] & VLAN_TAG_DATA_VID) == vid)) {
- if (netif_running(dev)) {
+ if (update_hw) {
ret = vlan_write_filter(dev, hw, i, 0);
if (ret)
return ret;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
index d7e4db7224b0..7ecb7addd2ea 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
@@ -31,7 +31,8 @@ static int stmmac_xdp_enable_pool(struct stmmac_priv *priv,
return err;
}
- need_update = netif_running(priv->dev) && stmmac_xdp_is_enabled(priv);
+ need_update = priv->datapath == STMMAC_DATAPATH_RUNNING &&
+ stmmac_xdp_is_enabled(priv);
if (need_update) {
napi_disable(&ch->rx_napi);
@@ -69,7 +70,8 @@ static int stmmac_xdp_disable_pool(struct stmmac_priv *priv, u16 queue)
if (!pool)
return -EINVAL;
- need_update = netif_running(priv->dev) && stmmac_xdp_is_enabled(priv);
+ need_update = priv->datapath == STMMAC_DATAPATH_RUNNING &&
+ stmmac_xdp_is_enabled(priv);
if (need_update) {
napi_disable(&ch->rxtx_napi);
@@ -106,8 +108,9 @@ int stmmac_xdp_set_prog(struct stmmac_priv *priv, struct bpf_prog *prog,
struct bpf_prog *old_prog;
bool need_update;
bool if_running;
+ int ret;
- if_running = netif_running(dev);
+ if_running = priv->datapath == STMMAC_DATAPATH_RUNNING;
if (prog && dev->mtu > ETH_DATA_LEN) {
/* For now, the driver doesn't support XDP functionality with
@@ -117,25 +120,41 @@ int stmmac_xdp_set_prog(struct stmmac_priv *priv, struct bpf_prog *prog,
return -EOPNOTSUPP;
}
- if (!prog)
- xdp_features_clear_redirect_target(dev);
-
need_update = !!priv->xdp_prog != !!prog;
if (if_running && need_update)
stmmac_xdp_release(dev);
old_prog = xchg(&priv->xdp_prog, prog);
- if (old_prog)
- bpf_prog_put(old_prog);
/* Disable RX SPH for XDP operation */
priv->sph_active = priv->sph_capable && !stmmac_xdp_is_enabled(priv);
- if (if_running && need_update)
- stmmac_xdp_open(dev);
+ if (if_running && need_update) {
+ ret = stmmac_xdp_open(dev);
+ if (ret) {
+ netdev_err(dev, "failed reopening after XDP change: %pe; interface remains detached\n",
+ ERR_PTR(ret));
+ if (prog) {
+ /* The core retains the old program on error and drops
+ * the reference it passed for the proposed program.
+ */
+ xchg(&priv->xdp_prog, old_prog);
+ priv->sph_active = priv->sph_capable && !old_prog;
+ return ret;
+ }
+ /* Uninstalling a BPF link must release its program even
+ * if the non-XDP datapath cannot be restarted.
+ */
+ }
+ }
+
+ if (old_prog)
+ bpf_prog_put(old_prog);
if (prog)
xdp_features_set_redirect_target(dev, false);
+ else
+ xdp_features_clear_redirect_target(dev);
return 0;
}
diff --git a/include/linux/stmmac.h b/include/linux/stmmac.h
index 00be2df63d22..17c54afc415a 100644
--- a/include/linux/stmmac.h
+++ b/include/linux/stmmac.h
@@ -291,6 +291,13 @@ struct plat_stmmacenet_data {
void (*exit)(struct device *dev, void *priv);
int (*suspend)(struct device *dev, void *priv);
int (*resume)(struct device *dev, void *priv);
+ /* MAC WoL retains register/receive/timestamp clocks and power. These
+ * hooks only prepare/undo additional wake resources, without gating
+ * register access. Set both hooks together; a failed preparation must
+ * unwind its own resources.
+ */
+ int (*suspend_wol)(struct device *dev, void *priv);
+ void (*resume_wol)(struct device *dev, void *priv);
int (*mac_setup)(void *priv, struct mac_device_info *mac);
int (*clks_config)(void *priv, bool enabled);
int (*crosststamp)(ktime_t *device, struct system_counterval_t *system,
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 14/19] net: stmmac: use the tracked datapath restart for XSK pool changes
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (12 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 13/19] net: stmmac: track datapath and power ownership across failed reopening James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 15/19] net: stmmac: restore TC offloads before restarting DMA James Hilliard
` (6 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Replace the void per-queue enable helpers with the tracked XDP restart.
Pause all queues and MAC link resolution while the pool bitmap and rings
change, leaving the PHY running through phylink replay.
Unwind a failed pool attachment without leaving NAPI over missing
buffers. Pool removal must complete even if ordinary-ring rebuilding
fails, after retiring all references to the departing pool. Preserve TBS
configuration and return a detached interface to the ordinary down/up
recovery path.
Fixes: bba2556efad6 ("net: stmmac: Enable RX via AF_XDP zero-copy")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 5 -
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 116 +++-------------------
drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c | 44 ++++----
3 files changed, 32 insertions(+), 133 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 6927afd01175..38370f2cbe87 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -473,11 +473,6 @@ static inline bool stmmac_tc_active(struct stmmac_priv *priv)
netif_device_present(priv->dev) && !priv->hw_unavailable;
}
-void stmmac_disable_rx_queue(struct stmmac_priv *priv, u32 queue);
-void stmmac_enable_rx_queue(struct stmmac_priv *priv, u32 queue);
-void stmmac_disable_tx_queue(struct stmmac_priv *priv, u32 queue);
-void stmmac_enable_tx_queue(struct stmmac_priv *priv, u32 queue);
-int stmmac_xsk_wakeup(struct net_device *dev, u32 queue, u32 flags);
struct timespec64 stmmac_calc_tas_basetime(ktime_t old_base_time,
ktime_t current_time,
u64 cycle_time);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index ad5ed2c95af7..95757f3cbc64 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -2318,6 +2318,7 @@ static void __free_dma_tx_desc_resources(struct stmmac_priv *priv,
kfree(tx_q->tx_skbuff);
tx_q->tx_skbuff = NULL;
+ tx_q->xsk_pool = NULL;
}
static void free_dma_tx_desc_resources(struct stmmac_priv *priv,
@@ -2830,6 +2831,12 @@ static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
bool work_done = true;
u32 tx_set_ic_bit = 0;
+ /* Nothing can be submitted while the link is down. Let NAPI complete;
+ * userspace can retry ndo_xsk_wakeup() once carrier has returned.
+ */
+ if (!netif_carrier_ok(priv->dev))
+ return true;
+
/* Avoids TX time-out as we are sharing with slow path */
txq_trans_cond_update(nq);
@@ -2844,8 +2851,7 @@ static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
/* We are sharing with slow path and stop XSK TX desc submission when
* available TX ring is less than threshold.
*/
- if (unlikely(stmmac_tx_avail(priv, queue) < STMMAC_TX_XSK_AVAIL) ||
- !netif_carrier_ok(priv->dev)) {
+ if (unlikely(stmmac_tx_avail(priv, queue) < STMMAC_TX_XSK_AVAIL)) {
work_done = false;
break;
}
@@ -7450,107 +7456,6 @@ static int stmmac_xdp_xmit(struct net_device *dev, int num_frames,
return nxmit;
}
-void stmmac_disable_rx_queue(struct stmmac_priv *priv, u32 queue)
-{
- struct stmmac_channel *ch = &priv->channel[queue];
- unsigned long flags;
-
- spin_lock_irqsave(&ch->lock, flags);
- stmmac_disable_dma_irq(priv, priv->ioaddr, queue, 1, 0);
- spin_unlock_irqrestore(&ch->lock, flags);
-
- stmmac_stop_rx_dma(priv, queue);
- __free_dma_rx_desc_resources(priv, priv->dma_conf, queue);
-}
-
-void stmmac_enable_rx_queue(struct stmmac_priv *priv, u32 queue)
-{
- struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[queue];
- struct stmmac_channel *ch = &priv->channel[queue];
- unsigned long flags;
- int ret;
-
- ret = __alloc_dma_rx_desc_resources(priv, priv->dma_conf, queue);
- if (ret) {
- netdev_err(priv->dev, "Failed to alloc RX desc.\n");
- return;
- }
-
- ret = __init_dma_rx_desc_rings(priv, priv->dma_conf, queue, GFP_KERNEL);
- if (ret) {
- __free_dma_rx_desc_resources(priv, priv->dma_conf, queue);
- netdev_err(priv->dev, "Failed to init RX desc.\n");
- return;
- }
-
- stmmac_reset_rx_queue(priv, queue);
- stmmac_clear_rx_descriptors(priv, priv->dma_conf, queue);
-
- stmmac_init_rx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
- rx_q->dma_rx_phy, queue);
-
- stmmac_set_queue_rx_tail_ptr(priv, rx_q, queue, rx_q->buf_alloc_num);
-
- stmmac_set_queue_rx_buf_size(priv, rx_q, queue);
-
- stmmac_start_rx_dma(priv, queue);
-
- spin_lock_irqsave(&ch->lock, flags);
- stmmac_enable_dma_irq(priv, priv->ioaddr, queue, 1, 0);
- spin_unlock_irqrestore(&ch->lock, flags);
-}
-
-void stmmac_disable_tx_queue(struct stmmac_priv *priv, u32 queue)
-{
- struct stmmac_channel *ch = &priv->channel[queue];
- unsigned long flags;
-
- spin_lock_irqsave(&ch->lock, flags);
- stmmac_disable_dma_irq(priv, priv->ioaddr, queue, 0, 1);
- spin_unlock_irqrestore(&ch->lock, flags);
-
- stmmac_stop_tx_dma(priv, queue);
- __free_dma_tx_desc_resources(priv, priv->dma_conf, queue);
-}
-
-void stmmac_enable_tx_queue(struct stmmac_priv *priv, u32 queue)
-{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[queue];
- struct stmmac_channel *ch = &priv->channel[queue];
- unsigned long flags;
- int ret;
-
- ret = __alloc_dma_tx_desc_resources(priv, priv->dma_conf, queue);
- if (ret) {
- netdev_err(priv->dev, "Failed to alloc TX desc.\n");
- return;
- }
-
- ret = __init_dma_tx_desc_rings(priv, priv->dma_conf, queue);
- if (ret) {
- __free_dma_tx_desc_resources(priv, priv->dma_conf, queue);
- netdev_err(priv->dev, "Failed to init TX desc.\n");
- return;
- }
-
- stmmac_reset_tx_queue(priv, queue);
- stmmac_clear_tx_descriptors(priv, priv->dma_conf, queue);
-
- stmmac_init_tx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
- tx_q->dma_tx_phy, queue);
-
- if (tx_q->tbs & STMMAC_TBS_AVAIL)
- stmmac_enable_tbs(priv, priv->ioaddr, 1, queue);
-
- stmmac_set_queue_tx_tail_ptr(priv, tx_q, queue, 0);
-
- stmmac_start_tx_dma(priv, queue);
-
- spin_lock_irqsave(&ch->lock, flags);
- stmmac_enable_dma_irq(priv, priv->ioaddr, queue, 0, 1);
- spin_unlock_irqrestore(&ch->lock, flags);
-}
-
void stmmac_xdp_release(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
@@ -7650,6 +7555,9 @@ int stmmac_xdp_open(struct net_device *dev)
stmmac_set_queue_tx_tail_ptr(priv, tx_q, chan, 0);
+ if (tx_q->tbs & STMMAC_TBS_AVAIL)
+ stmmac_enable_tbs(priv, priv->ioaddr, 1, chan);
+
hrtimer_setup(&tx_q->txtimer, stmmac_tx_timer, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
}
@@ -7687,7 +7595,7 @@ int stmmac_xdp_open(struct net_device *dev)
return ret;
}
-int stmmac_xsk_wakeup(struct net_device *dev, u32 queue, u32 flags)
+static int stmmac_xsk_wakeup(struct net_device *dev, u32 queue, u32 flags)
{
struct stmmac_priv *priv = netdev_priv(dev);
struct stmmac_channel *ch;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
index 7ecb7addd2ea..907ac49a1b76 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
@@ -9,7 +9,6 @@
static int stmmac_xdp_enable_pool(struct stmmac_priv *priv,
struct xsk_buff_pool *pool, u16 queue)
{
- struct stmmac_channel *ch = &priv->channel[queue];
bool need_update;
u32 frame_size;
int err;
@@ -34,23 +33,23 @@ static int stmmac_xdp_enable_pool(struct stmmac_priv *priv,
need_update = priv->datapath == STMMAC_DATAPATH_RUNNING &&
stmmac_xdp_is_enabled(priv);
- if (need_update) {
- napi_disable(&ch->rx_napi);
- napi_disable(&ch->tx_napi);
- stmmac_disable_rx_queue(priv, queue);
- stmmac_disable_tx_queue(priv, queue);
- }
+ if (need_update)
+ stmmac_xdp_release(priv->dev);
set_bit(queue, priv->af_xdp_zc_qps);
if (need_update) {
- stmmac_enable_rx_queue(priv, queue);
- stmmac_enable_tx_queue(priv, queue);
- napi_enable(&ch->rxtx_napi);
-
- err = stmmac_xsk_wakeup(priv->dev, queue, XDP_WAKEUP_RX);
- if (err)
+ err = stmmac_xdp_open(priv->dev);
+ if (err) {
+ clear_bit(queue, priv->af_xdp_zc_qps);
+ xsk_pool_dma_unmap(pool, STMMAC_RX_DMA_ATTR);
+ netdev_err(priv->dev, "failed reopening after XSK pool attach: %pe; interface remains detached\n",
+ ERR_PTR(err));
return err;
+ }
+
+ /* The pool is installed even if link resolution is still pending. */
+ napi_schedule(&priv->channel[queue].rxtx_napi);
}
return 0;
@@ -58,9 +57,9 @@ static int stmmac_xdp_enable_pool(struct stmmac_priv *priv,
static int stmmac_xdp_disable_pool(struct stmmac_priv *priv, u16 queue)
{
- struct stmmac_channel *ch = &priv->channel[queue];
struct xsk_buff_pool *pool;
bool need_update;
+ int err;
if (queue >= priv->plat->rx_queues_to_use ||
queue >= priv->plat->tx_queues_to_use)
@@ -73,24 +72,21 @@ static int stmmac_xdp_disable_pool(struct stmmac_priv *priv, u16 queue)
need_update = priv->datapath == STMMAC_DATAPATH_RUNNING &&
stmmac_xdp_is_enabled(priv);
- if (need_update) {
- napi_disable(&ch->rxtx_napi);
- stmmac_disable_rx_queue(priv, queue);
- stmmac_disable_tx_queue(priv, queue);
- synchronize_rcu();
- }
+ if (need_update)
+ stmmac_xdp_release(priv->dev);
xsk_pool_dma_unmap(pool, STMMAC_RX_DMA_ATTR);
clear_bit(queue, priv->af_xdp_zc_qps);
if (need_update) {
- stmmac_enable_rx_queue(priv, queue);
- stmmac_enable_tx_queue(priv, queue);
- napi_enable(&ch->rx_napi);
- napi_enable(&ch->tx_napi);
+ err = stmmac_xdp_open(priv->dev);
+ if (err)
+ netdev_err(priv->dev, "failed reopening after XSK pool removal: %pe; interface remains detached\n",
+ ERR_PTR(err));
}
+ /* Socket teardown must be able to unmap and free the removed pool. */
return 0;
}
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 15/19] net: stmmac: restore TC offloads before restarting DMA
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (13 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 14/19] net: stmmac: use the tracked datapath restart for XSK pool changes James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 16/19] xsk: allow drivers to retain DMA mappings independently of pools James Hilliard
` (5 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
A DMA software reset loses MAC filters and the MTL gate schedule, not
just the ring addresses. Build on Lorenzo Bianconi's EST replay helper
and embedded schedule state to restore the remaining offloads during
ordinary hardware setup and resume. The later live-XDP reset fallback
and retained-ring MTU transaction use the same restoration.
Keep the parsed L3/L4 rule rather than only its cookie. Program from
that saved definition both when installing a rule and after reset. Do
not publish a partially programmed replacement; restore the previous
slot when installation fails. Keep runtime VLAN priorities in the
existing queue configuration and replay EtherType steering and the
configured preemption/TC mapping. Preserve the additional fragment size.
Use stmmac_setup_est() after timestamp initialization, with the PTP mutex
held. Read the counter internally under the PTP lock so reset replay can
run while public PHC operations remain blocked. Advance the saved base
time by whole cycles when necessary. Delay DMA start until filters,
timestamp state and the gate schedule have been restored. Propagate a
replay error through the existing rollback or detached recovery path.
Build and validate TAPRIO replacements separately, including the PHC time
read, and publish the saved schedule only after hardware setup succeeds.
Keep the previous schedule on rejection and attempt to restore it after
a programming error. A rejected first install must not leave an enabled
zero-cycle cache for PHC adjustment or reset replay. Serialize schedule
publication with those consumers under the PTP mutex.
Pass the unpublished schedule to the same EST base-time and hardware
programming helper used by open, resume and PHC time adjustment. Keep
public TAPRIO requests gated while the PHC is blocked, without blocking
internal reset replay.
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 8 +
drivers/net/ethernet/stmicro/stmmac/stmmac_est.c | 29 ++-
drivers/net/ethernet/stmicro/stmmac/stmmac_est.h | 4 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_fpe.c | 1 +
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 32 +--
drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c | 2 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c | 275 +++++++++++++---------
7 files changed, 216 insertions(+), 135 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 38370f2cbe87..04d08b2c1e3f 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -151,6 +151,9 @@ struct stmmac_fpe_cfg {
struct ethtool_mmsv mmsv;
const struct stmmac_fpe_reg *reg;
u32 fpe_csr; /* MAC_FPE_CTRL_STS reg cache */
+ u32 preemptible_tcs;
+ u32 add_frag_size;
+ bool mapping_configured;
};
struct stmmac_tc_entry {
@@ -194,6 +197,10 @@ struct stmmac_flow_entry {
unsigned long cookie;
unsigned long action;
u8 ip_proto;
+ u32 ip4_src;
+ u32 ip4_dst;
+ u16 port_src;
+ u16 port_dst;
int in_use;
int idx;
int is_l4;
@@ -440,6 +447,7 @@ void stmmac_set_ethtool_ops(struct net_device *netdev);
void stmmac_ptp_register(struct stmmac_priv *priv);
void stmmac_ptp_unregister(struct stmmac_priv *priv);
int stmmac_ptp_restore(struct stmmac_priv *priv);
+int stmmac_tc_restore_filters(struct stmmac_priv *priv);
int stmmac_xdp_open(struct net_device *dev);
void stmmac_xdp_release(struct net_device *dev);
int stmmac_get_phy_intf_sel(phy_interface_t interface);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
index 49edfebbc39e..059ef60c13d0 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.c
@@ -80,29 +80,42 @@ static int est_configure(struct stmmac_priv *priv, struct stmmac_est *cfg,
return 0;
}
-int __stmmac_setup_est(struct stmmac_priv *priv)
+/* Program either the installed schedule or an unpublished replacement. */
+int __stmmac_setup_est(struct stmmac_priv *priv, struct stmmac_est *est)
{
struct timespec64 current_time, time;
ktime_t current_time_ns, basetime;
+ unsigned long flags;
+ u64 now;
u64 cycle_time;
int err;
lockdep_assert_held(&priv->est_lock);
+ lockdep_assert_held(&priv->ptp_mutex);
- priv->ptp_clock_ops.gettime64(&priv->ptp_clock_ops, ¤t_time);
+ if (!priv->ptp_enabled)
+ return -EOPNOTSUPP;
+
+ /* Reset replay owns ptp_mutex while public PHC reads are blocked. */
+ read_lock_irqsave(&priv->ptp_lock, flags);
+ err = stmmac_get_systime(priv, priv->ptpaddr, &now);
+ read_unlock_irqrestore(&priv->ptp_lock, flags);
+ if (err)
+ return err;
+ current_time = ns_to_timespec64(now);
current_time_ns = timespec64_to_ktime(current_time);
- time.tv_nsec = priv->est.btr_reserve[0];
- time.tv_sec = priv->est.btr_reserve[1];
+ time.tv_nsec = est->btr_reserve[0];
+ time.tv_sec = est->btr_reserve[1];
basetime = timespec64_to_ktime(time);
- cycle_time = (u64)priv->est.ctr[1] * NSEC_PER_SEC + priv->est.ctr[0];
+ cycle_time = (u64)est->ctr[1] * NSEC_PER_SEC + est->ctr[0];
time = stmmac_calc_tas_basetime(basetime, current_time_ns, cycle_time);
- priv->est.btr[0] = (u32)time.tv_nsec;
- priv->est.btr[1] = (u32)time.tv_sec;
+ est->btr[0] = (u32)time.tv_nsec;
+ est->btr[1] = (u32)time.tv_sec;
- err = stmmac_est_configure(priv, priv, &priv->est,
+ err = stmmac_est_configure(priv, priv, est,
priv->plat->clk_ptp_rate, true);
if (err)
netdev_err(priv->dev, "failed to re-configure EST\n");
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h
index b4d1a1f04f10..5665e7a53994 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_est.h
@@ -66,14 +66,14 @@
extern const struct stmmac_est_ops dwmac510_est_ops;
-int __stmmac_setup_est(struct stmmac_priv *priv);
+int __stmmac_setup_est(struct stmmac_priv *priv, struct stmmac_est *est);
static inline int stmmac_setup_est(struct stmmac_priv *priv)
{
int ret = 0;
mutex_lock(&priv->est_lock);
if (priv->est.enable)
- ret = __stmmac_setup_est(priv);
+ ret = __stmmac_setup_est(priv, &priv->est);
mutex_unlock(&priv->est_lock);
return ret;
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_fpe.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_fpe.c
index c889204a7aa5..067ea1f5134b 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_fpe.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_fpe.c
@@ -195,6 +195,7 @@ void stmmac_fpe_set_add_frag_size(struct stmmac_priv *priv, u32 add_frag_size)
value = readl(ioaddr + reg->mtl_fpe_reg);
writel(u32_replace_bits(value, add_frag_size, FPE_MTL_ADD_FRAG_SZ),
ioaddr + reg->mtl_fpe_reg);
+ priv->fpe_cfg.add_frag_size = add_frag_size;
}
#define ALG_ERR_MSG "TX algorithm SP is not suitable for one-to-many mapping"
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 95757f3cbc64..f5060924dae8 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -3829,6 +3829,11 @@ static int stmmac_hw_setup(struct net_device *dev)
if (ret)
return ret;
}
+ ret = stmmac_tc_restore_filters(priv);
+ if (ret)
+ return ret;
+ if (stmmac_fpe_supported(priv))
+ stmmac_fpe_set_add_frag_size(priv, priv->fpe_cfg.add_frag_size);
/* Initialize Safety Features */
stmmac_safety_feat_configuration(priv);
@@ -3891,9 +3896,6 @@ static int stmmac_hw_setup(struct net_device *dev)
stmmac_enable_tbs(priv, priv->ioaddr, enable, chan);
}
- /* Start the ball rolling... */
- stmmac_start_all_dma(priv);
-
phylink_rx_clk_stop_block(priv->phylink);
stmmac_set_hw_vlan_mode(priv, priv->hw);
phylink_rx_clk_stop_unblock(priv->phylink);
@@ -4441,16 +4443,17 @@ static int __stmmac_open(struct net_device *dev,
if (ret)
goto init_error;
- /* The core soft reset in stmmac_hw_setup() clears the MTL_EST
- * registers, so re-apply the taprio offload after PTP is up.
- */
- ret = stmmac_setup_est(priv);
- if (ret < 0)
- goto est_error;
-
stmmac_init_coalesce(priv);
stmmac_vlan_restore(priv);
+ mutex_lock(&priv->ptp_mutex);
+ ret = stmmac_setup_est(priv);
+ mutex_unlock(&priv->ptp_mutex);
+ if (ret)
+ goto irq_error;
+
+ /* All reset-sensitive offloads must be installed before DMA runs. */
+ stmmac_start_all_dma(priv);
ret = stmmac_request_irq(dev);
if (ret)
@@ -4478,7 +4481,6 @@ static int __stmmac_open(struct net_device *dev,
for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
hrtimer_cancel(&priv->dma_conf->tx_queue[chan].txtimer);
-est_error:
stmmac_release_ptp(priv);
init_error:
/* Undo phylink_prepare_resume() even if hardware setup failed before
@@ -8892,10 +8894,11 @@ int stmmac_resume(struct device *dev)
init_coalesce:
mutex_lock(&priv->ptp_mutex);
- stmmac_block_ptp(priv, false);
- mutex_unlock(&priv->ptp_mutex);
ret = stmmac_setup_est(priv);
- if (ret < 0)
+ if (!ret)
+ stmmac_block_ptp(priv, false);
+ mutex_unlock(&priv->ptp_mutex);
+ if (ret)
goto error_stop_dma;
stmmac_init_coalesce(priv);
@@ -8905,6 +8908,7 @@ int stmmac_resume(struct device *dev)
stmmac_vlan_restore(priv);
+ stmmac_start_all_dma(priv);
stmmac_enable_all_queues(priv);
stmmac_enable_all_dma_irq(priv);
stmmac_unmask_pci_irq(priv);
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
index b34711231227..f27c6ff210a3 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ptp.c
@@ -107,7 +107,7 @@ static int stmmac_adjust_time(struct ptp_clock_info *ptp, s64 delta)
* error, but do not report success if only schedule replay failed.
*/
if (priv->est.enable) {
- err = __stmmac_setup_est(priv);
+ err = __stmmac_setup_est(priv, &priv->est);
if (!ret)
ret = err;
}
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
index cac2b6e34c94..06089d47f373 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_tc.c
@@ -15,12 +15,17 @@
static int tc_config_preemption(struct stmmac_priv *priv,
struct netlink_ext_ack *extack, u32 preemptible_tcs)
{
- /* Qdisc teardown must not access unpowered registers. */
- if (priv->hw_unavailable)
- return 0;
+ int ret = 0;
- return stmmac_fpe_map_preemption_class(priv, priv->dev, extack,
- preemptible_tcs);
+ /* Qdisc teardown still clears the saved mapping after failed resume. */
+ if (!priv->hw_unavailable)
+ ret = stmmac_fpe_map_preemption_class(priv, priv->dev, extack,
+ preemptible_tcs);
+ if (!ret) {
+ priv->fpe_cfg.preemptible_tcs = preemptible_tcs;
+ priv->fpe_cfg.mapping_configured = true;
+ }
+ return ret;
}
static void tc_fill_all_pass_entry(struct stmmac_tc_entry *entry)
@@ -520,31 +525,15 @@ static int tc_add_ip4_flow(struct stmmac_priv *priv,
{
struct flow_rule *rule = flow_cls_offload_flow_rule(cls);
struct flow_dissector *dissector = rule->match.dissector;
- bool inv = entry->action & STMMAC_FLOW_ACTION_DROP;
struct flow_match_ipv4_addrs match;
- u32 hw_match;
- int ret;
/* Nothing to do here */
if (!dissector_uses_key(dissector, FLOW_DISSECTOR_KEY_IPV4_ADDRS))
return -EINVAL;
flow_rule_match_ipv4_addrs(rule, &match);
- hw_match = ntohl(match.key->src) & ntohl(match.mask->src);
- if (hw_match) {
- ret = stmmac_config_l3_filter(priv, priv->hw, entry->idx, true,
- false, true, inv, hw_match);
- if (ret)
- return ret;
- }
-
- hw_match = ntohl(match.key->dst) & ntohl(match.mask->dst);
- if (hw_match) {
- ret = stmmac_config_l3_filter(priv, priv->hw, entry->idx, true,
- false, false, inv, hw_match);
- if (ret)
- return ret;
- }
+ entry->ip4_src = ntohl(match.key->src) & ntohl(match.mask->src);
+ entry->ip4_dst = ntohl(match.key->dst) & ntohl(match.mask->dst);
return 0;
}
@@ -555,11 +544,7 @@ static int tc_add_ports_flow(struct stmmac_priv *priv,
{
struct flow_rule *rule = flow_cls_offload_flow_rule(cls);
struct flow_dissector *dissector = rule->match.dissector;
- bool inv = entry->action & STMMAC_FLOW_ACTION_DROP;
struct flow_match_ports match;
- u32 hw_match;
- bool is_udp;
- int ret;
/* Nothing to do here */
if (!dissector_uses_key(dissector, FLOW_DISSECTOR_KEY_PORTS))
@@ -567,10 +552,7 @@ static int tc_add_ports_flow(struct stmmac_priv *priv,
switch (entry->ip_proto) {
case IPPROTO_TCP:
- is_udp = false;
- break;
case IPPROTO_UDP:
- is_udp = true;
break;
default:
return -EINVAL;
@@ -578,23 +560,46 @@ static int tc_add_ports_flow(struct stmmac_priv *priv,
flow_rule_match_ports(rule, &match);
- hw_match = ntohs(match.key->src) & ntohs(match.mask->src);
- if (hw_match) {
- ret = stmmac_config_l4_filter(priv, priv->hw, entry->idx, true,
- is_udp, true, inv, hw_match);
+ entry->port_src = ntohs(match.key->src) & ntohs(match.mask->src);
+ entry->port_dst = ntohs(match.key->dst) & ntohs(match.mask->dst);
+
+ entry->is_l4 = true;
+ return 0;
+}
+
+static int tc_config_flow(struct stmmac_priv *priv,
+ const struct stmmac_flow_entry *entry)
+{
+ bool inv = entry->action & STMMAC_FLOW_ACTION_DROP;
+ bool udp = entry->ip_proto == IPPROTO_UDP;
+ int ret;
+
+ /* Clear the whole slot, including matches removed by a replacement. */
+ ret = stmmac_config_l3_filter(priv, priv->hw, entry->idx, false,
+ false, false, false, 0);
+ if (ret || !entry->in_use)
+ return ret;
+ if (entry->ip4_src) {
+ ret = stmmac_config_l3_filter(priv, priv->hw, entry->idx, true,
+ false, true, inv, entry->ip4_src);
if (ret)
return ret;
}
-
- hw_match = ntohs(match.key->dst) & ntohs(match.mask->dst);
- if (hw_match) {
+ if (entry->ip4_dst) {
+ ret = stmmac_config_l3_filter(priv, priv->hw, entry->idx, true,
+ false, false, inv, entry->ip4_dst);
+ if (ret)
+ return ret;
+ }
+ if (entry->port_src) {
ret = stmmac_config_l4_filter(priv, priv->hw, entry->idx, true,
- is_udp, false, inv, hw_match);
+ udp, true, inv, entry->port_src);
if (ret)
return ret;
}
-
- entry->is_l4 = true;
+ if (entry->port_dst)
+ return stmmac_config_l4_filter(priv, priv->hw, entry->idx, true,
+ udp, false, inv, entry->port_dst);
return 0;
}
@@ -630,6 +635,7 @@ static int tc_add_flow(struct stmmac_priv *priv,
{
struct stmmac_flow_entry *entry = tc_find_flow(priv, cls, false);
struct flow_rule *rule = flow_cls_offload_flow_rule(cls);
+ struct stmmac_flow_entry new = {};
int i, ret;
if (!entry) {
@@ -638,23 +644,33 @@ static int tc_add_flow(struct stmmac_priv *priv,
return -ENOENT;
}
- ret = tc_parse_flow_actions(priv, &rule->action, entry,
+ new.idx = entry->idx;
+ ret = tc_parse_flow_actions(priv, &rule->action, &new,
cls->common.extack);
if (ret)
return ret;
for (i = 0; i < ARRAY_SIZE(tc_flow_parsers); i++) {
- ret = tc_flow_parsers[i].fn(priv, cls, entry);
+ ret = tc_flow_parsers[i].fn(priv, cls, &new);
if (!ret)
- entry->in_use = true;
+ new.in_use = true;
else if (ret == -EOPNOTSUPP)
return ret;
}
- if (!entry->in_use)
+ if (!new.in_use)
return -EINVAL;
- entry->cookie = cls->cookie;
+ ret = tc_config_flow(priv, &new);
+ if (ret) {
+ /* Do not publish a rule that was only partially programmed. */
+ if (tc_config_flow(priv, entry))
+ netdev_err(priv->dev, "failed to restore flower filter %d\n",
+ entry->idx);
+ return ret;
+ }
+ new.cookie = cls->cookie;
+ *entry = new;
return 0;
}
@@ -740,6 +756,8 @@ static int tc_add_vlan_flow(struct stmmac_priv *priv,
prio = BIT(match.key->vlan_priority);
stmmac_rx_queue_prio(priv, priv->hw, prio, tc);
+ priv->plat->rx_queues_cfg[tc].prio = prio;
+ priv->plat->rx_queues_cfg[tc].use_prio = true;
entry->in_use = true;
entry->cookie = cls->cookie;
@@ -761,6 +779,8 @@ static int tc_del_vlan_flow(struct stmmac_priv *priv,
if (stmmac_tc_active(priv))
stmmac_rx_queue_prio(priv, priv->hw, 0, entry->tc);
+ priv->plat->rx_queues_cfg[entry->tc].prio = 0;
+ priv->plat->rx_queues_cfg[entry->tc].use_prio = true;
entry->in_use = false;
entry->cookie = 0;
@@ -935,6 +955,45 @@ static int tc_setup_cls(struct stmmac_priv *priv,
return ret;
}
+int stmmac_tc_restore_filters(struct stmmac_priv *priv)
+{
+ int i, ret;
+
+ for (i = 0; i < priv->flow_entries_max; i++) {
+ struct stmmac_flow_entry *entry = &priv->flow_entries[i];
+
+ if (!entry->in_use)
+ continue;
+ ret = tc_config_flow(priv, entry);
+ if (ret)
+ return ret;
+ }
+ for (i = 0; i < priv->rfs_entries_total; i++) {
+ struct stmmac_rfs_entry *entry = &priv->rfs_entries[i];
+
+ if (!entry->in_use)
+ continue;
+ switch (entry->type) {
+ /* VLAN priorities are replayed by stmmac_mtl_configuration(). */
+ case STMMAC_RFS_T_VLAN:
+ break;
+ case STMMAC_RFS_T_LLDP:
+ stmmac_rx_queue_routing(priv, priv->hw, PACKET_DCBCPQ,
+ entry->tc);
+ break;
+ case STMMAC_RFS_T_1588:
+ stmmac_rx_queue_routing(priv, priv->hw, PACKET_PTPQ,
+ entry->tc);
+ break;
+ }
+ }
+ /* The XGMAC callback also restores the runtime TC-to-queue mapping. */
+ if (priv->fpe_cfg.mapping_configured)
+ return tc_config_preemption(priv, NULL,
+ priv->fpe_cfg.preemptible_tcs);
+ return 0;
+}
+
struct timespec64 stmmac_calc_tas_basetime(ktime_t old_base_time,
ktime_t current_time,
u64 cycle_time)
@@ -958,7 +1017,7 @@ struct timespec64 stmmac_calc_tas_basetime(ktime_t old_base_time,
return time;
}
-static void tc_taprio_map_maxsdu_txq(struct stmmac_priv *priv,
+static void tc_taprio_map_maxsdu_txq(struct stmmac_est *est,
struct tc_taprio_qopt_offload *qopt)
{
u32 num_tc = qopt->mqprio.qopt.num_tc;
@@ -975,7 +1034,7 @@ static void tc_taprio_map_maxsdu_txq(struct stmmac_priv *priv,
count = qopt->mqprio.qopt.count[i];
for (j = offset; j < offset + count; j++)
- priv->est.max_sdu[j] = qopt->max_sdu[i] + ETH_HLEN - ETH_TLEN;
+ est->max_sdu[j] = qopt->max_sdu[i] + ETH_HLEN - ETH_TLEN;
}
}
@@ -984,10 +1043,10 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
{
u32 size, wid = priv->dma_cap.estwid, dep = priv->dma_cap.estdep;
struct netlink_ext_ack *extack = qopt->mqprio.extack;
+ struct timespec64 qopt_time;
u64 ctr = qopt->cycle_time;
- struct timespec64 time;
- u32 *gcl = NULL;
- int i, ret = 0;
+ struct stmmac_est *est;
+ int i, ret;
if (qopt->base_time < 0)
return -ERANGE;
@@ -1032,48 +1091,43 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
return -EOPNOTSUPP;
}
- mutex_lock(&priv->est_lock);
-
if (qopt->cmd == TAPRIO_CMD_DESTROY)
goto disable;
- if (priv->hw_unavailable) {
- ret = -EHOSTDOWN;
- goto unlock;
- }
-
- if (qopt->num_entries > dep) {
- ret = -EINVAL;
- goto unlock;
- }
-
- if (!qopt->cycle_time) {
- ret = -ERANGE;
- goto unlock;
- }
+ if (priv->hw_unavailable)
+ return -EHOSTDOWN;
+ if (!priv->ptp_enabled || !priv->ptp_clock_ops.gettime64)
+ return -EOPNOTSUPP;
+ /* Unlike reset replay, a new schedule requires an accessible PHC. */
+ if (priv->ptp_blocked)
+ return -EBUSY;
- if (qopt->cycle_time_extension >= BIT(wid + 7)) {
- ret = -ERANGE;
- goto unlock;
- }
+ if (qopt->num_entries > dep)
+ return -EINVAL;
+ if (!qopt->cycle_time)
+ return -ERANGE;
+ if (qopt->cycle_time_extension >= BIT(wid + 7))
+ return -ERANGE;
- gcl = kzalloc(sizeof(*gcl) * EST_GCL, GFP_KERNEL);
- if (!gcl) {
- ret = -ENOMEM;
- goto unlock;
- }
+ /* Build the replacement without changing the installed schedule. An
+ * entry rejected below must not leave an enabled, zero-cycle cache for
+ * PHC adjustment or reset replay to consume.
+ */
+ est = kzalloc_obj(*est);
+ if (!est)
+ return -ENOMEM;
size = qopt->num_entries;
+ est->gcl_size = size;
+ est->enable = true;
+
for (i = 0; i < size; i++) {
s64 delta_ns = qopt->entries[i].interval;
u32 gates = qopt->entries[i].gate_mask;
- if (delta_ns > GENMASK(wid - 1, 0)) {
+ if (delta_ns > GENMASK(wid - 1, 0) ||
+ gates > GENMASK(31 - wid, 0)) {
ret = -ERANGE;
- goto free_gcl;
- }
- if (gates > GENMASK(31 - wid, 0)) {
- ret = -ERANGE;
- goto free_gcl;
+ goto free_est;
}
switch (qopt->entries[i].command) {
@@ -1087,61 +1141,59 @@ static int tc_taprio_configure(struct stmmac_priv *priv,
break;
default:
ret = -EOPNOTSUPP;
- goto free_gcl;
+ goto free_est;
}
- gcl[i] = delta_ns | (gates << wid);
+ est->gcl[i] = delta_ns | (gates << wid);
}
- memset(&priv->est, 0, sizeof(priv->est));
- memcpy(priv->est.gcl, gcl, sizeof(priv->est.gcl));
- priv->est.gcl_size = size;
-
- time = ktime_to_timespec64(qopt->base_time);
- priv->est.btr_reserve[0] = (u32)time.tv_nsec;
- priv->est.btr_reserve[1] = (u32)time.tv_sec;
+ qopt_time = ktime_to_timespec64(qopt->base_time);
+ est->btr_reserve[0] = (u32)qopt_time.tv_nsec;
+ est->btr_reserve[1] = (u32)qopt_time.tv_sec;
+ est->ctr[0] = do_div(ctr, NSEC_PER_SEC);
+ est->ctr[1] = (u32)ctr;
+ est->ter = qopt->cycle_time_extension;
- priv->est.ctr[0] = do_div(ctr, NSEC_PER_SEC);
- priv->est.ctr[1] = (u32)ctr;
+ tc_taprio_map_maxsdu_txq(est, qopt);
- priv->est.ter = qopt->cycle_time_extension;
- tc_taprio_map_maxsdu_txq(priv, qopt);
-
- ret = __stmmac_setup_est(priv);
+ mutex_lock(&priv->est_lock);
+ ret = __stmmac_setup_est(priv, est);
if (ret)
- goto disable;
+ goto restore;
ret = tc_config_preemption(priv, extack, qopt->mqprio.preemptible_tcs);
if (ret)
- goto disable;
-
- priv->est.enable = true;
- kfree(gcl);
+ goto restore;
+ priv->est = *est;
mutex_unlock(&priv->est_lock);
+free_est:
+ kfree(est);
+ return ret;
- return 0;
+restore:
+ /* A failed hardware update must not publish the rejected schedule. */
+ if (stmmac_est_configure(priv, priv, &priv->est,
+ priv->plat->clk_ptp_rate, priv->est.enable))
+ netdev_err(priv->dev, "failed to restore EST\n");
+ mutex_unlock(&priv->est_lock);
+ goto free_est;
disable:
+ mutex_lock(&priv->est_lock);
priv->est.enable = false;
if (!priv->hw_unavailable)
stmmac_est_configure(priv, priv, &priv->est,
priv->plat->clk_ptp_rate, false);
- /* Reset taprio status */
+ /* Reset taprio stats */
for (i = 0; i < priv->plat->tx_queues_to_use; i++) {
priv->xstats.max_sdu_txq_drop[i] = 0;
priv->xstats.mtl_est_txq_hlbf[i] = 0;
priv->xstats.mtl_est_txq_hlbs[i] = 0;
}
- i = tc_config_preemption(priv, extack, 0);
- if (qopt->cmd == TAPRIO_CMD_DESTROY)
- ret = i;
-free_gcl:
- kfree(gcl);
-unlock:
mutex_unlock(&priv->est_lock);
- return ret;
+ return tc_config_preemption(priv, extack, 0);
}
static void tc_taprio_stats(struct stmmac_priv *priv,
@@ -1182,7 +1234,10 @@ static int tc_setup_taprio(struct stmmac_priv *priv,
switch (qopt->cmd) {
case TAPRIO_CMD_REPLACE:
case TAPRIO_CMD_DESTROY:
+ /* Serialize cache publication with PHC adjustment and reset replay. */
+ mutex_lock(&priv->ptp_mutex);
err = tc_taprio_configure(priv, qopt);
+ mutex_unlock(&priv->ptp_mutex);
break;
case TAPRIO_CMD_STATS:
tc_taprio_stats(priv, qopt);
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 16/19] xsk: allow drivers to retain DMA mappings independently of pools
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (14 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 15/19] net: stmmac: restore TC offloads before restarting DMA James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-28 12:22 ` Björn Töpel
2026-09-27 21:59 ` [PATCH net-next v5 17/19] net: stmmac: retain DMA memory until hardware shutdown completes James Hilliard
` (4 subsequent siblings)
20 siblings, 1 reply; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
An unsuccessful device shutdown does not make its DMA memory safe to
unmap. AF_XDP pool removal must nevertheless complete: the last socket
release invokes the driver detach callback and then destroys the pool,
regardless of the callback's return value.
Provide an independent reference to the existing DMA mapping and its
UMEM pages. A driver can take it while installing rings and release it
after DMA has actually stopped. Retain the pool and buffer-head storage
separately from the pool users reference which triggers teardown. This
lets a driver keep DMA-owned frames out of the reusable free list even
after socket teardown has released its fill and completion rings.
Return retained buffers only after DMA shutdown, before dropping the DMA
reference. Keeping pages pinned alone does not prevent an active pool
from recycling a frame still reachable by hardware. The retained metadata
does not permit other pool operations after detach.
Keep the DMA device and the mapping's netdev lookup key allocated,
without taking a netdev usage reference that would prevent unregister.
Save the mapping attributes and unmap before dropping the last retained
UMEM reference. Mapping reference operations are serialized by RTNL,
like the existing mapping list operations.
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
Changes in v5:
- Retain pool/head storage independently of socket users so DMA-owned
frames need not be returned to the free list during pool removal.
- Keep the saved mapping in a separate reference object, since detach
clears the pool's device and DMA lookup state.
---
include/net/xdp_sock_drv.h | 25 +++++++++++++++++
include/net/xsk_buff_pool.h | 7 +++++
net/xdp/xsk_buff_pool.c | 67 +++++++++++++++++++++++++++++++++++++++++++--
3 files changed, 96 insertions(+), 3 deletions(-)
diff --git a/include/net/xdp_sock_drv.h b/include/net/xdp_sock_drv.h
index d94aeb506379..a39768b1c00f 100644
--- a/include/net/xdp_sock_drv.h
+++ b/include/net/xdp_sock_drv.h
@@ -95,6 +95,22 @@ static inline void xsk_pool_dma_unmap(struct xsk_buff_pool *pool,
xp_dma_unmap(pool, attrs);
}
+/* RTNL must be held. Retain mappings, pages and buffer metadata without
+ * postponing the socket's detach callback. Do not return DMA-owned buffers
+ * with xsk_buff_free() until hardware has stopped, even after pool removal.
+ * Afterwards, free those buffers before putting the reference. This does not
+ * retain the FILL/COMPLETION rings or allow other pool operations after detach.
+ */
+static inline struct xsk_dma_ref *xsk_pool_dma_get(struct xsk_buff_pool *pool)
+{
+ return xp_dma_get(pool);
+}
+
+static inline void xsk_pool_dma_put(struct xsk_dma_ref *ref)
+{
+ xp_dma_put(ref);
+}
+
static inline int xsk_pool_dma_map(struct xsk_buff_pool *pool,
struct device *dev, unsigned long attrs)
{
@@ -432,6 +448,15 @@ static inline void xsk_pool_dma_unmap(struct xsk_buff_pool *pool,
{
}
+static inline struct xsk_dma_ref *xsk_pool_dma_get(struct xsk_buff_pool *pool)
+{
+ return NULL;
+}
+
+static inline void xsk_pool_dma_put(struct xsk_dma_ref *ref)
+{
+}
+
static inline int xsk_pool_dma_map(struct xsk_buff_pool *pool,
struct device *dev, unsigned long attrs)
{
diff --git a/include/net/xsk_buff_pool.h b/include/net/xsk_buff_pool.h
index a7df573784fd..df7e20fe5ae3 100644
--- a/include/net/xsk_buff_pool.h
+++ b/include/net/xsk_buff_pool.h
@@ -11,6 +11,7 @@
#include <net/xdp.h>
struct xsk_buff_pool;
+struct xsk_dma_ref;
struct xdp_rxq_info;
struct xsk_cb_desc;
struct xsk_queue;
@@ -38,6 +39,8 @@ struct xsk_dma_map {
dma_addr_t *dma_pages;
struct device *dev;
struct net_device *netdev;
+ struct xdp_umem *umem;
+ unsigned long attrs;
refcount_t users;
struct list_head list; /* Protected by the RTNL_LOCK */
u32 dma_pages_cnt;
@@ -52,6 +55,8 @@ struct xsk_buff_pool {
spinlock_t xsk_tx_list_lock;
refcount_t users;
struct xdp_umem *umem;
+ /* Pool/head storage; DMA references must not postpone socket teardown. */
+ refcount_t refs;
struct work_struct work;
/* Protects generic receive in shared and non-shared umem mode. */
spinlock_t rx_lock;
@@ -143,6 +148,8 @@ void xp_fill_cb(struct xsk_buff_pool *pool, struct xsk_cb_desc *desc);
int xp_dma_map(struct xsk_buff_pool *pool, struct device *dev,
unsigned long attrs, struct page **pages, u32 nr_pages);
void xp_dma_unmap(struct xsk_buff_pool *pool, unsigned long attrs);
+struct xsk_dma_ref *xp_dma_get(struct xsk_buff_pool *pool);
+void xp_dma_put(struct xsk_dma_ref *ref);
struct xdp_buff *xp_alloc(struct xsk_buff_pool *pool);
u32 xp_alloc_batch(struct xsk_buff_pool *pool, struct xdp_buff **xdp, u32 max);
bool xp_can_alloc(struct xsk_buff_pool *pool, u32 count);
diff --git a/net/xdp/xsk_buff_pool.c b/net/xdp/xsk_buff_pool.c
index c58f56f24a9c..8c71974494c9 100644
--- a/net/xdp/xsk_buff_pool.c
+++ b/net/xdp/xsk_buff_pool.c
@@ -12,6 +12,11 @@
#define ETH_PAD_LEN (ETH_HLEN + 2 * VLAN_HLEN + ETH_FCS_LEN)
+struct xsk_dma_ref {
+ struct xsk_dma_map *dma_map;
+ struct xsk_buff_pool *pool;
+};
+
void xp_add_xsk(struct xsk_buff_pool *pool, struct xdp_sock *xs)
{
if (!xs->tx)
@@ -34,7 +39,7 @@ void xp_del_xsk(struct xsk_buff_pool *pool, struct xdp_sock *xs)
void xp_destroy(struct xsk_buff_pool *pool)
{
- if (!pool)
+ if (!pool || !refcount_dec_and_test(&pool->refs))
return;
kvfree(pool->tx_descs);
@@ -68,6 +73,7 @@ struct xsk_buff_pool *xp_create_and_assign_umem(struct xdp_sock *xs,
pool = kvzalloc_flex(*pool, free_heads, entries);
if (!pool)
goto out;
+ refcount_set(&pool->refs, 1);
pool->heads = kvzalloc_objs(*pool->heads, umem->chunks);
if (!pool->heads)
@@ -360,7 +366,8 @@ static struct xsk_dma_map *xp_find_dma_map(struct xsk_buff_pool *pool)
}
static struct xsk_dma_map *xp_create_dma_map(struct device *dev, struct net_device *netdev,
- u32 nr_pages, struct xdp_umem *umem)
+ u32 nr_pages, struct xdp_umem *umem,
+ unsigned long attrs)
{
struct xsk_dma_map *dma_map;
@@ -376,6 +383,8 @@ static struct xsk_dma_map *xp_create_dma_map(struct device *dev, struct net_devi
dma_map->netdev = netdev;
dma_map->dev = dev;
+ dma_map->umem = umem;
+ dma_map->attrs = attrs;
dma_map->dma_pages_cnt = nr_pages;
refcount_set(&dma_map->users, 1);
list_add(&dma_map->list, &umem->xsk_dma_list);
@@ -430,6 +439,58 @@ void xp_dma_unmap(struct xsk_buff_pool *pool, unsigned long attrs)
}
EXPORT_SYMBOL(xp_dma_unmap);
+struct xsk_dma_ref *xp_dma_get(struct xsk_buff_pool *pool)
+{
+ struct xsk_dma_map *dma_map;
+ struct xsk_dma_ref *ref;
+
+ ASSERT_RTNL();
+ if (!pool->dma_pages)
+ return NULL;
+ dma_map = xp_find_dma_map(pool);
+ if (WARN_ON_ONCE(!dma_map))
+ return NULL;
+
+ ref = kmalloc_obj(*ref);
+ if (!ref)
+ return NULL;
+ ref->dma_map = dma_map;
+ ref->pool = pool;
+ refcount_inc(&pool->refs);
+ refcount_inc(&dma_map->users);
+ xdp_get_umem(dma_map->umem);
+ get_device(dma_map->dev);
+ /* Keep the mapping's lookup key alive without preventing unregister. */
+ get_device(&dma_map->netdev->dev);
+ return ref;
+}
+EXPORT_SYMBOL_GPL(xp_dma_get);
+
+void xp_dma_put(struct xsk_dma_ref *ref)
+{
+ struct xsk_dma_map *dma_map;
+ struct net_device *netdev;
+ struct xdp_umem *umem;
+ struct device *dev;
+
+ ASSERT_RTNL();
+ if (!ref)
+ return;
+ dma_map = ref->dma_map;
+ dev = dma_map->dev;
+ netdev = dma_map->netdev;
+ umem = dma_map->umem;
+ if (refcount_dec_and_test(&dma_map->users))
+ __xp_dma_unmap(dma_map, dma_map->attrs);
+ /* Unmap before the final reference can unpin the UMEM pages. */
+ xdp_put_umem(umem, false);
+ put_device(&netdev->dev);
+ put_device(dev);
+ xp_destroy(ref->pool);
+ kfree(ref);
+}
+EXPORT_SYMBOL_GPL(xp_dma_put);
+
static void xp_check_dma_contiguity(struct xsk_dma_map *dma_map)
{
u32 i;
@@ -487,7 +548,7 @@ int xp_dma_map(struct xsk_buff_pool *pool, struct device *dev,
return 0;
}
- dma_map = xp_create_dma_map(dev, pool->netdev, nr_pages, pool->umem);
+ dma_map = xp_create_dma_map(dev, pool->netdev, nr_pages, pool->umem, attrs);
if (!dma_map)
return -ENOMEM;
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* Re: [PATCH net-next v5 16/19] xsk: allow drivers to retain DMA mappings independently of pools
2026-09-27 21:59 ` [PATCH net-next v5 16/19] xsk: allow drivers to retain DMA mappings independently of pools James Hilliard
@ 2026-09-28 12:22 ` Björn Töpel
0 siblings, 0 replies; 27+ messages in thread
From: Björn Töpel @ 2026-09-28 12:22 UTC (permalink / raw)
To: James Hilliard, Russell King, Andrew Lunn, Heiner Kallweit,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Thierry Reding, Jonathan Hunter, Chen-Yu Tsai,
Jernej Skrabec, Samuel Holland, Jose Abreu, Yao Zi,
Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
James Hilliard <james.hilliard1@gmail.com> writes:
> An unsuccessful device shutdown does not make its DMA memory safe to
> unmap.
Why does this need an AF_XDP API? Can stmmac fence DMA before releasing
the pool? Please show a case where stop and reset fail, DMA remains
active, and the driver cannot isolate it.
Björn
^ permalink raw reply [flat|nested] 27+ messages in thread
* [PATCH net-next v5 17/19] net: stmmac: retain DMA memory until hardware shutdown completes
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (15 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 16/19] xsk: allow drivers to retain DMA mappings independently of pools James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 18/19] net: stmmac: prepare device-local DMA interrupt quiescence James Hilliard
` (3 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Clearing the DMA start bits requests a stop but need not complete an
in-flight frame or descriptor writeback. Ordinary release, live XDP/XSK
replacement and late open failures currently free the rings and buffers
immediately afterwards. An XSK socket can then unmap and unpin the UMEM
while hardware still has its addresses.
Wait for stopped process states on the legacy and Allwinner DMA engines.
For GMAC4 configurations represented by DSR0, also require its bus-busy
bits to clear. Fall back to a completed global reset if the idle wait
times out or the integration has no supported idle indication, including
XGMAC and GMAC4 configurations with more than three channels. Prepare
the PHY receive clock for that reset and restore PHC configuration
before a live XDP restart which required it. As with MTU reset,
continuous PHC time is not preserved on this fallback.
Track configurations exposed to DMA separately from software datapath
ownership. If both idle and reset fail, keep the rings, DMA mappings and
backing memory until a subsequent successful reset. Do not overwrite
retained buffers in an XDP reopen or change their ring/channel geometry.
Record failed open replacements too, including ones which are no longer
the active configuration pointer. A successful down/up reset retires
them.
Take independent XSK DMA/UMEM and buffer metadata references when
initializing RX rings. On failed shutdown, remove active pool pointers
but retain the RX buffer heads without returning them to the free list.
Socket teardown still completes. A later successful reset releases the
buffers and their references; pinning pages alone would not prevent an
active pool from reusing a frame that hardware could still overwrite.
Defer hard TX error recovery to process context instead of rewriting a
ring in the IRQ handler immediately after clearing ST. Likewise, move
resume-time TX cleanup and descriptor rebuilding after the reset
succeeds. Do not let queued recovery work reopen an administratively
closed device.
There is no generic, guaranteed isolation mechanism across all stmmac
integrations. If hardware still cannot stop or reset at removal,
deliberately retain the DMA allocations and report the quarantine rather
than expose recycled memory to DMA. Such an unrecoverable device can
therefore retain memory, including pinned UMEM, until reboot.
Rebuild retained RX descriptors after reset with buffer addresses and
chain links written before ownership. GMAC4 and XGMAC secondary-address
programming overwrites des3, so publishing OWN first would lose it.
Use this rebuild on system resume too: writeback status is not a valid
read-format descriptor. Fill missing page-pool buffers after reset, and
leave a failed refill detached. Recycle XSK buffers only after the reset
fence, accept empty or partial FILL rings, and publish ownership only for
populated descriptors.
Fixes: ac746c8520d9 ("net: stmmac: enhance XDP ZC driver level switching performance")
Fixes: bba2556efad6 ("net: stmmac: Enable RX via AF_XDP zero-copy")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
Changes in v5:
- Keep quarantined XSK RX buffers outside the pool free list until a
successful shutdown/reset, including after socket teardown.
- Rebuild read-format RX descriptors on resume, not only MTU rollback.
Refill page-pool holes and handle empty/partial XSK FILL rings without
applying the page-pool-only rollback helper to XSK buffers.
---
drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c | 17 +
.../net/ethernet/stmicro/stmmac/dwmac1000_dma.c | 1 +
drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c | 1 +
drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c | 2 +
drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h | 6 +
drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c | 18 +
drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h | 2 +
drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c | 20 +
drivers/net/ethernet/stmicro/stmmac/hwif.h | 4 +
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 11 +-
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 459 ++++++++++++++++-----
11 files changed, 445 insertions(+), 96 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c
index 38d7e71de925..c9145441aab0 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c
@@ -163,6 +163,7 @@ static const struct emac_variant emac_variant_h6 = {
#define EMAC_TX_CUR_DESC 0xB4
#define EMAC_TX_CUR_BUF 0xB8
#define EMAC_RX_DMA_STA 0xC0
+#define EMAC_DMA_STATE_MASK GENMASK(2, 0)
#define EMAC_RX_CUR_DESC 0xC4
#define EMAC_RX_CUR_BUF 0xC8
@@ -425,6 +426,21 @@ static void sun8i_dwmac_dma_stop_rx(struct stmmac_priv *priv,
writel(v, ioaddr + EMAC_RX_CTL1);
}
+static int sun8i_dwmac_dma_wait_idle(struct stmmac_priv *priv,
+ void __iomem *ioaddr)
+{
+ u32 value;
+ int ret;
+
+ /* STOP (0) follows the frame transfer and descriptor close states. */
+ ret = readl_poll_timeout(ioaddr + EMAC_TX_DMA_STA, value,
+ !(value & EMAC_DMA_STATE_MASK), 100, 100000);
+ if (ret)
+ return ret;
+ return readl_poll_timeout(ioaddr + EMAC_RX_DMA_STA, value,
+ !(value & EMAC_DMA_STATE_MASK), 100, 100000);
+}
+
static int sun8i_dwmac_dma_interrupt(struct stmmac_priv *priv,
void __iomem *ioaddr,
struct stmmac_extra_stats *x, u32 chan,
@@ -553,6 +569,7 @@ static void sun8i_dwmac_dma_operation_mode_tx(struct stmmac_priv *priv,
static const struct stmmac_dma_ops sun8i_dwmac_dma_ops = {
.reset = sun8i_dwmac_dma_reset,
+ .wait_idle = sun8i_dwmac_dma_wait_idle,
.init = sun8i_dwmac_dma_init,
.init_rx_chan = sun8i_dwmac_dma_init_rx,
.init_tx_chan = sun8i_dwmac_dma_init_tx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c b/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c
index 3ac7a7949529..4cb7e6c16bdd 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c
@@ -252,6 +252,7 @@ static void dwmac1000_rx_watchdog(struct stmmac_priv *priv,
const struct stmmac_dma_ops dwmac1000_dma_ops = {
.reset = dwmac_dma_reset,
+ .wait_idle = dwmac_dma_wait_idle,
.init_chan = dwmac1000_dma_init_channel,
.init_rx_chan = dwmac1000_dma_init_rx,
.init_tx_chan = dwmac1000_dma_init_tx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c b/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c
index 12b2bf2d739a..5ffd3c1471c4 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c
@@ -108,6 +108,7 @@ static void dwmac100_dma_diagnostic_fr(struct stmmac_extra_stats *x,
const struct stmmac_dma_ops dwmac100_dma_ops = {
.reset = dwmac_dma_reset,
+ .wait_idle = dwmac_dma_wait_idle,
.init = dwmac100_dma_init,
.init_rx_chan = dwmac100_dma_init_rx,
.init_tx_chan = dwmac100_dma_init_tx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c
index 14ac3f0e51f7..d7928678dee1 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c
@@ -570,6 +570,7 @@ static int dwmac4_enable_tbs(struct stmmac_priv *priv, void __iomem *ioaddr,
const struct stmmac_dma_ops dwmac4_dma_ops = {
.reset = dwmac4_dma_reset,
+ .wait_idle = dwmac4_dma_wait_idle,
.init = dwmac4_dma_init,
.init_chan = dwmac4_dma_init_channel,
.deinit_chan = dwmac4_dma_deinit_channel,
@@ -600,6 +601,7 @@ const struct stmmac_dma_ops dwmac4_dma_ops = {
const struct stmmac_dma_ops dwmac410_dma_ops = {
.reset = dwmac4_dma_reset,
+ .wait_idle = dwmac4_dma_wait_idle,
.init = dwmac4_dma_init,
.init_chan = dwmac410_dma_init_channel,
.deinit_chan = dwmac410_dma_deinit_channel,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h
index 43b036d4e95b..9352107204eb 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h
@@ -10,6 +10,12 @@
#ifndef __DWMAC4_DMA_H__
#define __DWMAC4_DMA_H__
+int dwmac4_dma_wait_idle(struct stmmac_priv *priv, void __iomem *ioaddr);
+
+#define DMA_DEBUG_STATUS0 0x0000100c
+#define DMA_DEBUG_BUS_BUSY GENMASK(1, 0)
+#define DMA_DEBUG_CH_STATE(ch) (GENMASK(15, 8) << ((ch) * 8))
+
/* Define the max channel number used for tx (also rx).
* dwmac4 accepts up to 8 channels for TX (and also 8 channels for RX
*/
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c b/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c
index a0249715fafa..9af0565a9bca 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c
@@ -26,6 +26,24 @@ int dwmac4_dma_reset(void __iomem *ioaddr)
10000, 1000000);
}
+int dwmac4_dma_wait_idle(struct stmmac_priv *priv, void __iomem *ioaddr)
+{
+ u32 channels = max(priv->plat->rx_queues_to_use,
+ priv->plat->tx_queues_to_use);
+ u32 mask = DMA_DEBUG_BUS_BUSY;
+ u32 value, chan;
+
+ /* DSR0 describes channels 0..2 and outstanding AXI transactions.
+ * Other debug-register layouts require a successful reset instead.
+ */
+ if (channels > 3)
+ return -EOPNOTSUPP;
+ for (chan = 0; chan < channels; chan++)
+ mask |= DMA_DEBUG_CH_STATE(chan);
+ return readl_poll_timeout(ioaddr + DMA_DEBUG_STATUS0, value,
+ !(value & mask), 100, 100000);
+}
+
void dwmac4_set_rx_tail_ptr(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 tail_ptr, u32 chan)
{
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h b/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h
index e1c37ac2c99d..970495bccfd2 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h
@@ -11,6 +11,8 @@
#ifndef __DWMAC_DMA_H__
#define __DWMAC_DMA_H__
+int dwmac_dma_wait_idle(struct stmmac_priv *priv, void __iomem *ioaddr);
+
/* DMA CRS Control and Status Register Mapping */
#define DMA_BUS_MODE 0x00001000 /* Bus Mode */
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c b/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c
index a0383f9486c2..bb907db8fca1 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c
@@ -27,6 +27,26 @@ int dwmac_dma_reset(void __iomem *ioaddr)
10000, 200000);
}
+int dwmac_dma_wait_idle(struct stmmac_priv *priv, void __iomem *ioaddr)
+{
+ u32 channels = max(priv->plat->rx_queues_to_use,
+ priv->plat->tx_queues_to_use);
+ u32 value, chan;
+ int ret;
+
+ /* CSR5 process states, not the latched process-stopped interrupts.
+ * Stopped is reached after the outstanding descriptor writeback.
+ */
+ for (chan = 0; chan < channels; chan++) {
+ ret = readl_poll_timeout(ioaddr + DMA_CHAN_STATUS(chan), value,
+ !(value & (DMA_STATUS_TS_MASK | DMA_STATUS_RS_MASK)),
+ 100, 100000);
+ if (ret)
+ return ret;
+ }
+ return 0;
+}
+
/* CSR1 enables the transmit DMA to check for new descriptor */
void dwmac_enable_dma_transmission(void __iomem *ioaddr, u32 chan)
{
diff --git a/drivers/net/ethernet/stmicro/stmmac/hwif.h b/drivers/net/ethernet/stmicro/stmmac/hwif.h
index 1fd9f1ab316e..4b7381a6fcce 100644
--- a/drivers/net/ethernet/stmicro/stmmac/hwif.h
+++ b/drivers/net/ethernet/stmicro/stmmac/hwif.h
@@ -205,6 +205,8 @@ struct stmmac_dma_ops {
u32 chan);
void (*stop_rx)(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan);
+ /* Called after stopping every channel; must also drain bus accesses. */
+ int (*wait_idle)(struct stmmac_priv *priv, void __iomem *ioaddr);
int (*dma_interrupt)(struct stmmac_priv *priv, void __iomem *ioaddr,
struct stmmac_extra_stats *x, u32 chan, u32 dir);
/* If supported then get the optional core features */
@@ -269,6 +271,8 @@ struct stmmac_dma_ops {
stmmac_do_void_callback(__priv, dma, start_rx, __priv, __args)
#define stmmac_stop_rx(__priv, __args...) \
stmmac_do_void_callback(__priv, dma, stop_rx, __priv, __args)
+#define stmmac_dma_wait_idle(__priv, __args...) \
+ stmmac_do_callback(__priv, dma, wait_idle, __priv, __args)
#define stmmac_dma_interrupt_status(__priv, __args...) \
stmmac_do_callback(__priv, dma, dma_interrupt, __priv, __args)
#define stmmac_get_hw_feature(__priv, __args...) \
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 04d08b2c1e3f..090d79aeb2ad 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -120,6 +120,7 @@ struct stmmac_rx_queue {
u32 queue_index;
struct xdp_rxq_info xdp_rxq;
struct xsk_buff_pool *xsk_pool;
+ struct xsk_dma_ref *xsk_dma;
struct page_pool *page_pool;
struct stmmac_rx_buffer *buf_pool;
struct stmmac_priv *priv_data;
@@ -223,6 +224,10 @@ struct stmmac_rfs_entry {
};
struct stmmac_dma_conf {
+ /* RTNL: all configurations exposed to DMA survive until stop/reset. */
+ struct list_head list;
+ bool dma_owned;
+ bool retired;
unsigned int dma_buf_sz;
/* RX Queue */
@@ -266,11 +271,11 @@ struct stmmac_msi {
};
enum stmmac_datapath_state {
- /* No IRQs or DMA allocations owned by a successful open. */
+ /* No IRQs or enabled NAPI; failed DMA shutdown may retain memory. */
STMMAC_DATAPATH_DOWN,
/* Resources allocated, NAPI enabled. */
STMMAC_DATAPATH_RUNNING,
- /* Resources retained, NAPI and DMA stopped; also after failed resume. */
+ /* Resources retained, NAPI disabled, DMA stop requested. */
STMMAC_DATAPATH_SUSPENDED,
};
@@ -297,6 +302,8 @@ struct stmmac_priv {
struct mutex lock;
struct stmmac_dma_conf *dma_conf;
+ struct list_head dma_confs;
+ bool dma_reset_needed;
/* IRQ/DMA ownership and NAPI state, serialized by RTNL. */
enum stmmac_datapath_state datapath;
/* Core sleep sequence completed, independently of datapath ownership. */
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index f5060924dae8..98dbc873e1c8 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -1027,6 +1027,25 @@ static void stmmac_block_ptp(struct stmmac_priv *priv, bool block)
write_unlock_irqrestore(&priv->ptp_lock, flags);
}
+static int stmmac_restore_timestamping(struct stmmac_priv *priv)
+{
+ int ret;
+
+ if (!priv->ptp_enabled)
+ return 0;
+
+ ret = stmmac_init_ptp_clk_freq(priv);
+ if (ret)
+ return ret;
+
+ ret = stmmac_init_tstamp_counter(priv, priv->systime_flags);
+ if (ret)
+ return ret;
+ if (priv->plat->flags & STMMAC_FLAG_HWTSTAMP_CORRECT_LATENCY)
+ stmmac_hwtstamp_correct_latency(priv, priv);
+ return stmmac_ptp_restore(priv);
+}
+
static void stmmac_legacy_serdes_power_down(struct stmmac_priv *priv)
{
if (priv->plat->serdes_powerdown && priv->legacy_serdes_is_powered)
@@ -1704,24 +1723,10 @@ static void stmmac_clear_descriptors(struct stmmac_priv *priv,
stmmac_clear_tx_descriptors(priv, dma_conf, queue);
}
-/**
- * stmmac_init_rx_buffers - init the RX descriptor buffer.
- * @priv: driver private structure
- * @dma_conf: structure to take the dma data
- * @p: descriptor pointer
- * @i: descriptor index
- * @flags: gfp flag
- * @queue: RX queue index
- * Description: this function is called to allocate a receive buffer, perform
- * the DMA mapping and init the descriptor.
- */
-static int stmmac_init_rx_buffers(struct stmmac_priv *priv,
- struct stmmac_dma_conf *dma_conf,
- struct dma_desc *p,
- int i, gfp_t flags, u32 queue)
+static int stmmac_alloc_rx_buffer(struct stmmac_priv *priv,
+ struct stmmac_rx_queue *rx_q,
+ struct stmmac_rx_buffer *buf)
{
- struct stmmac_rx_queue *rx_q = &dma_conf->rx_queue[queue];
- struct stmmac_rx_buffer *buf = &rx_q->buf_pool[i];
gfp_t gfp = (GFP_ATOMIC | __GFP_NOWARN);
if (priv->dma_cap.host_dma_width <= 32)
@@ -1738,19 +1743,49 @@ static int stmmac_init_rx_buffers(struct stmmac_priv *priv,
buf->sec_page = page_pool_alloc_pages(rx_q->page_pool, gfp);
if (!buf->sec_page)
return -ENOMEM;
-
buf->sec_addr = page_pool_get_dma_addr(buf->sec_page);
- stmmac_set_desc_sec_addr(priv, p, buf->sec_addr, true);
- } else {
- buf->sec_page = NULL;
- stmmac_set_desc_sec_addr(priv, p, buf->sec_addr, false);
}
+ return 0;
+}
+
+static void stmmac_init_rx_buffer_desc(struct stmmac_priv *priv,
+ struct stmmac_dma_conf *dma_conf,
+ struct dma_desc *p,
+ struct stmmac_rx_buffer *buf)
+{
+ if (buf->sec_page)
+ buf->sec_addr = page_pool_get_dma_addr(buf->sec_page);
+ stmmac_set_desc_sec_addr(priv, p, buf->sec_addr, !!buf->sec_page);
buf->addr = page_pool_get_dma_addr(buf->page) + buf->page_offset;
stmmac_set_desc_addr(priv, p, buf->addr);
if (dma_conf->dma_buf_sz == BUF_SIZE_16KiB)
stmmac_init_desc3(priv, p);
+}
+
+/**
+ * stmmac_init_rx_buffers - allocate a receive buffer and init its descriptor
+ * @priv: driver private structure
+ * @dma_conf: structure to take the dma data
+ * @p: descriptor pointer
+ * @i: descriptor index
+ * @flags: gfp flag
+ * @queue: RX queue index
+ */
+static int stmmac_init_rx_buffers(struct stmmac_priv *priv,
+ struct stmmac_dma_conf *dma_conf,
+ struct dma_desc *p,
+ int i, gfp_t flags, u32 queue)
+{
+ struct stmmac_rx_queue *rx_q = &dma_conf->rx_queue[queue];
+ struct stmmac_rx_buffer *buf = &rx_q->buf_pool[i];
+ int ret;
+
+ ret = stmmac_alloc_rx_buffer(priv, rx_q, buf);
+ if (ret)
+ return ret;
+ stmmac_init_rx_buffer_desc(priv, dma_conf, p, buf);
return 0;
}
@@ -1967,6 +2002,9 @@ static int __init_dma_rx_desc_rings(struct stmmac_priv *priv,
rx_q->xsk_pool = stmmac_get_xsk_pool(priv, queue);
if (rx_q->xsk_pool) {
+ rx_q->xsk_dma = xsk_pool_dma_get(rx_q->xsk_pool);
+ if (!rx_q->xsk_dma)
+ return -EINVAL;
ret = xdp_rxq_info_reg_mem_model(&rx_q->xdp_rxq,
MEM_TYPE_XSK_BUFF_POOL, NULL);
if (ret)
@@ -2213,6 +2251,70 @@ static void stmmac_free_tx_skbufs(struct stmmac_priv *priv)
dma_free_tx_skbufs(priv, priv->dma_conf, queue);
}
+/* Only after a successful DMA reset. MTU rollback has already filled holes;
+ * resume may need new page-pool buffers or a partially populated XSK ring.
+ */
+static int stmmac_reinit_dma_desc(struct stmmac_priv *priv)
+{
+ struct stmmac_dma_conf *dma_conf = priv->dma_conf;
+ u32 queue, i;
+ int ret;
+
+ stmmac_free_tx_skbufs(priv);
+ stmmac_reset_queues_param(priv);
+ init_dma_tx_desc_rings(priv->dev, dma_conf);
+ for (queue = 0; queue < priv->plat->tx_queues_to_use; queue++)
+ stmmac_clear_tx_descriptors(priv, dma_conf, queue);
+
+ for (queue = 0; queue < priv->plat->rx_queues_to_use; queue++) {
+ struct stmmac_rx_queue *rx_q = &dma_conf->rx_queue[queue];
+
+ if (rx_q->state_saved)
+ dev_kfree_skb_any(rx_q->state.skb);
+ rx_q->state.skb = NULL;
+ rx_q->state_saved = 0;
+ rx_q->rx_count_frames = 0;
+ rx_q->buf_alloc_num = 0;
+
+ /* Writeback format contains status, not buffer addresses. Rebuild
+ * read format from software ownership before publishing any OWN.
+ */
+ memset(stmmac_get_rx_desc(priv, rx_q, 0), 0,
+ stmmac_get_rx_desc_size(priv) * dma_conf->dma_rx_size);
+ if (rx_q->xsk_pool) {
+ dma_free_rx_xskbufs(priv, dma_conf, queue);
+ /* Empty FILL rings are valid, including TX-only sockets. */
+ stmmac_alloc_rx_buffers_zc(priv, dma_conf, queue);
+ } else {
+ for (i = 0; i < dma_conf->dma_rx_size; i++) {
+ struct dma_desc *p = stmmac_get_rx_desc(priv, rx_q, i);
+
+ ret = stmmac_alloc_rx_buffer(priv, rx_q,
+ &rx_q->buf_pool[i]);
+ if (ret)
+ return ret;
+ stmmac_init_rx_buffer_desc(priv, dma_conf, p,
+ &rx_q->buf_pool[i]);
+ rx_q->buf_alloc_num++;
+ }
+ }
+
+ if (priv->descriptor_mode == STMMAC_CHAIN_MODE)
+ stmmac_mode_init(priv, stmmac_get_rx_desc(priv, rx_q, 0),
+ rx_q->dma_rx_phy, dma_conf->dma_rx_size,
+ priv->extend_desc);
+
+ dma_wmb();
+ for (i = 0; i < rx_q->buf_alloc_num; i++)
+ stmmac_init_rx_desc(priv, stmmac_get_rx_desc(priv, rx_q, i),
+ priv->use_riwt, priv->descriptor_mode,
+ i == dma_conf->dma_rx_size - 1,
+ dma_conf->dma_buf_sz);
+ }
+
+ return 0;
+}
+
/**
* __free_dma_rx_desc_resources - free RX dma desc resources (per queue)
* @priv: private structure
@@ -2228,9 +2330,10 @@ static void __free_dma_rx_desc_resources(struct stmmac_priv *priv,
void *addr;
/* Release the DMA RX socket buffers */
- if (rx_q->xsk_pool) {
+ if (rx_q->xsk_pool || rx_q->xsk_dma) {
dma_free_rx_xskbufs(priv, dma_conf, queue);
- xsk_pool_set_rxq_info(rx_q->xsk_pool, NULL);
+ if (rx_q->xsk_pool)
+ xsk_pool_set_rxq_info(rx_q->xsk_pool, NULL);
} else {
dma_free_rx_skbufs(priv, dma_conf, queue);
}
@@ -2260,6 +2363,9 @@ static void __free_dma_rx_desc_resources(struct stmmac_priv *priv,
xdp_rxq_info_unreg(&rx_q->xdp_rxq);
kfree(rx_q->buf_pool);
+ if (rx_q->xsk_dma)
+ xsk_pool_dma_put(rx_q->xsk_dma);
+ rx_q->xsk_dma = NULL;
rx_q->buf_pool = NULL;
if (rx_q->page_pool) {
@@ -2271,11 +2377,10 @@ static void __free_dma_rx_desc_resources(struct stmmac_priv *priv,
static void free_dma_rx_desc_resources(struct stmmac_priv *priv,
struct stmmac_dma_conf *dma_conf)
{
- u8 rx_count = priv->plat->rx_queues_to_use;
u8 queue;
/* Free RX queue resources */
- for (queue = 0; queue < rx_count; queue++)
+ for (queue = 0; queue < MTL_MAX_RX_QUEUES; queue++)
__free_dma_rx_desc_resources(priv, dma_conf, queue);
}
@@ -2324,11 +2429,10 @@ static void __free_dma_tx_desc_resources(struct stmmac_priv *priv,
static void free_dma_tx_desc_resources(struct stmmac_priv *priv,
struct stmmac_dma_conf *dma_conf)
{
- u8 tx_count = priv->plat->tx_queues_to_use;
u8 queue;
/* Free TX queue resources */
- for (queue = 0; queue < tx_count; queue++)
+ for (queue = 0; queue < MTL_MAX_TX_QUEUES; queue++)
__free_dma_tx_desc_resources(priv, dma_conf, queue);
}
@@ -2555,14 +2659,35 @@ static int alloc_dma_desc_resources(struct stmmac_priv *priv,
return ret;
}
-/**
- * free_dma_desc_resources - free dma desc resources
- * @priv: private structure
- * @dma_conf: structure to take the dma data
- */
+static void stmmac_detach_xsk_buffers(struct stmmac_priv *priv,
+ struct stmmac_dma_conf *dma_conf)
+{
+ u32 queue;
+
+ /* Socket teardown can complete, but the DMA references must retain both
+ * the mapping and the buffer metadata. Do not put these frames on the
+ * pool's free list while hardware can still overwrite them.
+ */
+ for (queue = 0; queue < MTL_MAX_RX_QUEUES; queue++) {
+ struct stmmac_rx_queue *rx_q = &dma_conf->rx_queue[queue];
+
+ if (!rx_q->xsk_pool)
+ continue;
+ xsk_pool_set_rxq_info(rx_q->xsk_pool, NULL);
+ rx_q->xsk_pool = NULL;
+ }
+ for (queue = 0; queue < MTL_MAX_TX_QUEUES; queue++)
+ dma_conf->tx_queue[queue].xsk_pool = NULL;
+}
+
static void free_dma_desc_resources(struct stmmac_priv *priv,
struct stmmac_dma_conf *dma_conf)
{
+ if (dma_conf->dma_owned) {
+ stmmac_detach_xsk_buffers(priv, dma_conf);
+ return;
+ }
+
/* Release the DMA TX socket buffers */
free_dma_tx_desc_resources(priv, dma_conf);
@@ -2572,6 +2697,42 @@ static void free_dma_desc_resources(struct stmmac_priv *priv,
free_dma_rx_desc_resources(priv, dma_conf);
}
+static void stmmac_put_dma_conf(struct stmmac_priv *priv,
+ struct stmmac_dma_conf *dma_conf)
+{
+ free_dma_desc_resources(priv, dma_conf);
+ if (dma_conf->dma_owned) {
+ dma_conf->retired = true;
+ return;
+ }
+ list_del(&dma_conf->list);
+ kfree(dma_conf);
+}
+
+/* A successful global reset is also the retirement fence for configurations
+ * retained by a previous failed close, open, or MTU rollback.
+ */
+static void stmmac_dma_reset_complete(struct stmmac_priv *priv)
+{
+ struct stmmac_dma_conf *dma_conf, *next;
+
+ list_for_each_entry_safe(dma_conf, next, &priv->dma_confs, list) {
+ dma_conf->dma_owned = false;
+ if (dma_conf->retired)
+ stmmac_put_dma_conf(priv, dma_conf);
+ }
+}
+
+static bool stmmac_dma_busy(struct stmmac_priv *priv)
+{
+ struct stmmac_dma_conf *dma_conf;
+
+ list_for_each_entry(dma_conf, &priv->dma_confs, list)
+ if (dma_conf->dma_owned)
+ return true;
+ return false;
+}
+
/**
* stmmac_mac_enable_rx_queues - Enable MAC rx queues
* @priv: driver private structure
@@ -2598,6 +2759,7 @@ static void stmmac_mac_enable_rx_queues(struct stmmac_priv *priv)
*/
static void stmmac_start_rx_dma(struct stmmac_priv *priv, u32 chan)
{
+ priv->dma_conf->dma_owned = true;
netdev_dbg(priv->dev, "DMA RX processes started in channel %d\n", chan);
stmmac_start_rx(priv, priv->ioaddr, chan);
}
@@ -2611,6 +2773,7 @@ static void stmmac_start_rx_dma(struct stmmac_priv *priv, u32 chan)
*/
static void stmmac_start_tx_dma(struct stmmac_priv *priv, u32 chan)
{
+ priv->dma_conf->dma_owned = true;
netdev_dbg(priv->dev, "DMA TX processes started in channel %d\n", chan);
stmmac_start_tx(priv, priv->ioaddr, chan);
}
@@ -3145,25 +3308,17 @@ static int stmmac_tx_clean(struct stmmac_priv *priv, int budget, u32 queue,
* stmmac_tx_err - to manage the tx error
* @priv: driver private structure
* @chan: channel index
- * Description: it cleans the descriptors and restarts the transmission
- * in case of transmission errors.
+ * Description: stop submissions and request process-context DMA recovery.
*/
static void stmmac_tx_err(struct stmmac_priv *priv, u32 chan)
{
- struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[chan];
-
netif_tx_stop_queue(netdev_get_tx_queue(priv->dev, chan));
-
stmmac_stop_tx_dma(priv, chan);
- dma_free_tx_skbufs(priv, priv->dma_conf, chan);
- stmmac_clear_tx_descriptors(priv, priv->dma_conf, chan);
- stmmac_reset_tx_queue(priv, chan);
- stmmac_init_tx_chan(priv, priv->ioaddr, priv->plat->dma_cfg,
- tx_q->dma_tx_phy, chan);
- stmmac_start_tx_dma(priv, chan);
-
priv->xstats.tx_errors++;
- netif_tx_wake_queue(netdev_get_tx_queue(priv->dev, chan));
+ /* Recovery must wait for DMA before freeing or rewriting descriptors.
+ * Use the process-context reset path, not teardown in hard IRQ context.
+ */
+ stmmac_global_err(priv);
}
/**
@@ -3399,12 +3554,13 @@ static int stmmac_prereset_configure(struct stmmac_priv *priv)
/**
* stmmac_init_dma_engine - DMA init.
* @priv: driver private structure
+ * @reinit: rebuild the retained rings after a successful reset
* Description:
* It inits the DMA invoking the specific MAC/GMAC callback.
* Some DMA parameters can be passed from the platform;
* in case of these are not passed a default is kept for the MAC or GMAC.
*/
-static int stmmac_init_dma_engine(struct stmmac_priv *priv)
+static int stmmac_init_dma_engine(struct stmmac_priv *priv, bool reinit)
{
u8 rx_channels_count = priv->plat->rx_queues_to_use;
u8 tx_channels_count = priv->plat->tx_queues_to_use;
@@ -3423,6 +3579,17 @@ static int stmmac_init_dma_engine(struct stmmac_priv *priv)
netdev_err(priv->dev, "Failed to reset the dma\n");
return ret;
}
+ stmmac_dma_reset_complete(priv);
+ priv->dma_reset_needed = false;
+
+ if (reinit || priv->datapath == STMMAC_DATAPATH_SUSPENDED) {
+ /* Suspend only requested a stop. Do not modify its descriptors
+ * or release pending TX buffers until this reset has completed.
+ */
+ ret = stmmac_reinit_dma_desc(priv);
+ if (ret)
+ return ret;
+ }
/* DMA Configuration */
stmmac_dma_init(priv, priv->ioaddr, priv->plat->dma_cfg);
@@ -3773,6 +3940,8 @@ static bool stmmac_tso_channel_permitted(struct stmmac_priv *priv,
/**
* stmmac_hw_setup - setup mac in a usable state.
* @dev : pointer to the device structure.
+ * @reinit: rebuild retained descriptor rings after the DMA reset
+ * @keep_ptp: restore the registered PHC's configuration before starting DMA
* Description:
* this is the main function to setup the HW in a usable state because the
* dma engine is reset, the core registers are configured (e.g. AXI,
@@ -3782,7 +3951,7 @@ static bool stmmac_tso_channel_permitted(struct stmmac_priv *priv,
* 0 on success and an appropriate (-)ve integer as defined in errno.h
* file on failure.
*/
-static int stmmac_hw_setup(struct net_device *dev)
+static int stmmac_hw_setup(struct net_device *dev, bool reinit, bool keep_ptp)
{
struct stmmac_priv *priv = netdev_priv(dev);
u8 rx_cnt = priv->plat->rx_queues_to_use;
@@ -3804,7 +3973,7 @@ static int stmmac_hw_setup(struct net_device *dev)
phylink_rx_clk_stop_block(priv->phylink);
/* DMA initialization and SW reset */
- ret = stmmac_init_dma_engine(priv);
+ ret = stmmac_init_dma_engine(priv, reinit);
if (ret < 0) {
phylink_rx_clk_stop_unblock(priv->phylink);
netdev_err(priv->dev, "%s: DMA engine initialization failed\n",
@@ -3896,6 +4065,15 @@ static int stmmac_hw_setup(struct net_device *dev)
stmmac_enable_tbs(priv, priv->ioaddr, enable, chan);
}
+ if (keep_ptp) {
+ ret = stmmac_restore_timestamping(priv);
+ if (ret)
+ return ret;
+ ret = stmmac_setup_est(priv);
+ if (ret)
+ return ret;
+ }
+
phylink_rx_clk_stop_block(priv->phylink);
stmmac_set_hw_vlan_mode(priv, priv->hw);
phylink_rx_clk_stop_unblock(priv->phylink);
@@ -4255,6 +4433,7 @@ stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu)
__func__);
return ERR_PTR(-ENOMEM);
}
+ list_add_tail(&dma_conf->list, &priv->dma_confs);
len = mtu + ETH_HLEN + 2 * VLAN_HLEN + ETH_FCS_LEN;
@@ -4305,9 +4484,8 @@ stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu)
return dma_conf;
init_error:
- free_dma_desc_resources(priv, dma_conf);
alloc_error:
- kfree(dma_conf);
+ stmmac_put_dma_conf(priv, dma_conf);
return ERR_PTR(ret);
}
@@ -4402,6 +4580,57 @@ static int stmmac_resume_hw(struct stmmac_priv *priv)
return 0;
}
+/* NAPI, transmitters and IRQ handlers have already been drained. Clearing
+ * ST/SR only requests a stop: the current frame may still access memory.
+ * Keep the configuration DMA-owned unless hardware acknowledges idle/reset.
+ */
+static void stmmac_drain_dma(struct stmmac_priv *priv)
+{
+ struct stmmac_dma_conf *dma_conf;
+ int ret;
+
+ if (priv->hw_unavailable)
+ return;
+ stmmac_stop_all_dma(priv);
+ stmmac_mac_set(priv, priv->ioaddr, false);
+ if (!stmmac_dma_busy(priv))
+ return;
+
+ /* A failed replacement may have programmed a different topology. Only
+ * a global reset can acknowledge all of those retired configurations.
+ */
+ list_for_each_entry(dma_conf, &priv->dma_confs, list)
+ if (dma_conf != priv->dma_conf && dma_conf->dma_owned)
+ goto reset;
+
+ ret = stmmac_dma_wait_idle(priv, priv->ioaddr);
+ if (!ret) {
+ priv->dma_conf->dma_owned = false;
+ return;
+ }
+
+ /* Some integrations do not expose a usable idle indication. Reset is
+ * also the fallback after a stop timeout. It needs the PHY RX clock,
+ * even though phylink has already stopped link resolution.
+ */
+reset:
+ phylink_prepare_resume(priv->phylink);
+ mutex_lock(&priv->ptp_mutex);
+ stmmac_block_ptp(priv, true);
+ priv->dma_reset_needed = true;
+ phylink_rx_clk_stop_block(priv->phylink);
+ ret = stmmac_prereset_configure(priv);
+ if (!ret)
+ ret = stmmac_reset(priv);
+ phylink_rx_clk_stop_unblock(priv->phylink);
+ if (!ret)
+ stmmac_dma_reset_complete(priv);
+ else
+ netdev_err(priv->dev, "DMA shutdown failed: %pe; retaining DMA memory\n",
+ ERR_PTR(ret));
+ mutex_unlock(&priv->ptp_mutex);
+}
+
/**
* __stmmac_open - open entry point of the driver
* @dev : pointer to the device structure.
@@ -4433,7 +4662,7 @@ static int __stmmac_open(struct net_device *dev,
stmmac_reset_queues_param(priv);
- ret = stmmac_hw_setup(dev);
+ ret = stmmac_hw_setup(dev, false, false);
if (ret < 0) {
netdev_err(priv->dev, "%s: Hw setup failed\n", __func__);
goto init_error;
@@ -4487,8 +4716,9 @@ static int __stmmac_open(struct net_device *dev,
* phylink_start(). Keep the PHY attachment and outer PM ownership.
*/
phylink_stop(priv->phylink);
- stmmac_stop_all_dma(priv);
- stmmac_mac_set(priv, priv->ioaddr, false);
+ stmmac_drain_dma(priv);
+ /* Reset fallback may have powered the stopped PHY up for its clock. */
+ phylink_stop(priv->phylink);
return ret;
}
@@ -4532,7 +4762,7 @@ static int stmmac_open(struct net_device *dev)
if (ret)
goto err_serdes;
- kfree(old_conf);
+ stmmac_put_dma_conf(priv, old_conf);
/* We may have called phylink_speed_down before */
phylink_speed_up(priv->phylink);
@@ -4547,8 +4777,7 @@ static int stmmac_open(struct net_device *dev)
pm_runtime_put(priv->device);
err_dma_resources:
priv->dma_conf = old_conf;
- free_dma_desc_resources(priv, dma_conf);
- kfree(dma_conf);
+ stmmac_put_dma_conf(priv, dma_conf);
return ret;
}
@@ -4595,23 +4824,19 @@ static void __stmmac_release(struct net_device *dev)
/* Free the IRQ lines */
stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
- /* TX error IRQs can restart a queue after the first quiescence. */
+ /* Drain any final IRQ-triggered network activity before DMA shutdown. */
stmmac_stop_tx_queues(priv);
+ if (!priv->hw_unavailable && stmmac_fpe_supported(priv))
+ ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
- /* Stop TX/RX DMA after draining IRQ handlers which can restart it. */
- if (!priv->hw_unavailable) {
- stmmac_stop_all_dma(priv);
- /* Link resolution need not have reached mac_link_up() yet. */
- stmmac_mac_set(priv, priv->ioaddr, false);
- }
+ /* Only confirmed hardware shutdown permits releasing DMA memory. */
+ stmmac_drain_dma(priv);
+ phylink_stop(priv->phylink);
/* Release and free the Rx/Tx resources */
free_dma_desc_resources(priv, priv->dma_conf);
stmmac_release_ptp(priv);
-
- if (!priv->hw_unavailable && stmmac_fpe_supported(priv))
- ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
}
/**
@@ -6541,8 +6766,7 @@ static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
ret = __stmmac_open(dev, dma_conf);
if (ret) {
priv->dma_conf = old_conf;
- free_dma_desc_resources(priv, dma_conf);
- kfree(dma_conf);
+ stmmac_put_dma_conf(priv, dma_conf);
/*
* Keep the administrative state and PHY/PM ownership until
* ndo_stop(), but prevent use of the released data path.
@@ -6552,7 +6776,7 @@ static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
return ret;
}
- kfree(old_conf);
+ stmmac_put_dma_conf(priv, old_conf);
stmmac_set_rx_mode(dev);
netif_device_attach(dev);
@@ -7473,24 +7697,19 @@ void stmmac_xdp_release(struct net_device *dev)
/* Free the IRQ lines */
stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
stmmac_stop_tx_queues(priv);
+ if (stmmac_fpe_supported(priv))
+ ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
- /* Stop TX/RX DMA channels */
- stmmac_stop_all_dma(priv);
+ stmmac_drain_dma(priv);
/* Release and free the Rx/Tx resources */
free_dma_desc_resources(priv, priv->dma_conf);
- /* Disable the MAC Rx/Tx */
- stmmac_mac_set(priv, priv->ioaddr, false);
-
/* set trans_start so we don't get spurious
* watchdogs during reset
*/
netif_trans_update(dev);
- if (stmmac_fpe_supported(priv))
- ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
-
/* Keep PTP across the immediately following stmmac_xdp_open(). That
* function releases it if reopening fails, before returning DOWN.
*/
@@ -7508,6 +7727,14 @@ int stmmac_xdp_open(struct net_device *dev)
u8 chan;
int ret;
+ /* The old rings cannot be overwritten after a failed shutdown. Pool
+ * removal still completes, with their mappings held independently.
+ */
+ if (stmmac_dma_busy(priv)) {
+ ret = -EBUSY;
+ goto dma_desc_error;
+ }
+
ret = alloc_dma_desc_resources(priv, priv->dma_conf);
if (ret < 0) {
netdev_err(dev, "%s: DMA descriptors allocation failed\n",
@@ -7523,6 +7750,21 @@ int stmmac_xdp_open(struct net_device *dev)
}
stmmac_reset_queues_param(priv);
+ if (priv->dma_reset_needed) {
+ phylink_prepare_resume(priv->phylink);
+ mutex_lock(&priv->ptp_mutex);
+ ret = stmmac_hw_setup(dev, false, true);
+ if (!ret)
+ stmmac_block_ptp(priv, false);
+ mutex_unlock(&priv->ptp_mutex);
+ if (ret) {
+ stmmac_drain_dma(priv);
+ goto init_error;
+ }
+ stmmac_set_rx_mode(dev);
+ stmmac_vlan_restore(priv);
+ goto setup_timers;
+ }
/* DMA CSR Channel configuration */
for (chan = 0; chan < dma_csr_ch; chan++) {
@@ -7559,15 +7801,17 @@ int stmmac_xdp_open(struct net_device *dev)
if (tx_q->tbs & STMMAC_TBS_AVAIL)
stmmac_enable_tbs(priv, priv->ioaddr, 1, chan);
-
- hrtimer_setup(&tx_q->txtimer, stmmac_tx_timer, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
}
/* Enable the MAC Rx/Tx */
stmmac_mac_set(priv, priv->ioaddr, true);
- /* Start Rx & Tx DMA Channels */
+setup_timers:
+ /* The reset path has also restored filters, PTP and the EST schedule. */
stmmac_start_all_dma(priv);
+ for (chan = 0; chan < tx_cnt; chan++)
+ hrtimer_setup(&priv->dma_conf->tx_queue[chan].txtimer,
+ stmmac_tx_timer, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
ret = stmmac_request_irq(dev);
if (ret)
@@ -7584,8 +7828,7 @@ int stmmac_xdp_open(struct net_device *dev)
irq_error:
stmmac_stop_tx_queues(priv);
- stmmac_stop_all_dma(priv);
- stmmac_mac_set(priv, priv->ioaddr, false);
+ stmmac_drain_dma(priv);
init_error:
free_dma_desc_resources(priv, priv->dma_conf);
@@ -7715,7 +7958,7 @@ static void stmmac_reset_subtask(struct stmmac_priv *priv)
netdev_err(priv->dev, "Reset adapter.\n");
rtnl_lock();
- if (!netif_device_present(priv->dev))
+ if (!netif_device_present(priv->dev) || !netif_running(priv->dev))
goto out_unlock;
netif_trans_update(priv->dev);
@@ -8016,12 +8259,11 @@ static int stmmac_reopen(struct net_device *dev)
ret = __stmmac_open(dev, dma_conf);
if (ret) {
priv->dma_conf = old_conf;
- free_dma_desc_resources(priv, dma_conf);
- kfree(dma_conf);
+ stmmac_put_dma_conf(priv, dma_conf);
return ret;
}
- kfree(old_conf);
+ stmmac_put_dma_conf(priv, old_conf);
netif_device_attach(dev);
return 0;
}
@@ -8056,6 +8298,8 @@ int stmmac_reinit_queues(struct net_device *dev, u8 rx_cnt, u8 tx_cnt)
netif_device_detach(dev);
__stmmac_release(dev);
}
+ if (stmmac_dma_busy(priv))
+ return -EBUSY;
stmmac_set_queues(dev, rx_cnt, tx_cnt);
@@ -8083,6 +8327,8 @@ int stmmac_reinit_ringparam(struct net_device *dev, u32 rx_size, u32 tx_size)
netif_device_detach(dev);
__stmmac_release(dev);
}
+ if (stmmac_dma_busy(priv))
+ return -EBUSY;
priv->dma_conf->dma_rx_size = rx_size;
priv->dma_conf->dma_tx_size = tx_size;
@@ -8268,8 +8514,20 @@ EXPORT_SYMBOL_GPL(stmmac_plat_dat_alloc);
static void stmmac_free_dma_conf(void *data)
{
struct stmmac_priv *priv = data;
+ struct stmmac_dma_conf *dma_conf, *next;
- kfree(priv->dma_conf);
+ list_for_each_entry_safe(dma_conf, next, &priv->dma_confs, list) {
+ list_del(&dma_conf->list);
+ /* A permanently unresponsive device must not DMA into recycled
+ * memory, even on unbind. There is no generic isolation mechanism
+ * for all stmmac integrations. Deliberately retain these allocations.
+ */
+ if (dma_conf->dma_owned) {
+ dev_err(priv->device, "DMA still active on removal; DMA memory quarantined\n");
+ continue;
+ }
+ kfree(dma_conf);
+ }
}
static int __stmmac_dvr_probe(struct device *device,
@@ -8296,10 +8554,12 @@ static int __stmmac_dvr_probe(struct device *device,
priv = netdev_priv(ndev);
priv->device = device;
priv->dev = ndev;
+ INIT_LIST_HEAD(&priv->dma_confs);
/* Keep ring sizes and per-queue settings even while the device is down. */
priv->dma_conf = kzalloc_obj(*priv->dma_conf);
if (!priv->dma_conf)
return -ENOMEM;
+ list_add_tail(&priv->dma_conf->list, &priv->dma_confs);
ret = devm_add_action_or_reset(device, stmmac_free_dma_conf, priv);
if (ret)
return ret;
@@ -8624,12 +8884,28 @@ void stmmac_dvr_remove(struct device *dev)
{
struct net_device *ndev = dev_get_drvdata(dev);
struct stmmac_priv *priv = netdev_priv(ndev);
+ struct stmmac_dma_conf *dma_conf;
+ u32 queue;
netdev_info(priv->dev, "%s: removing driver", __func__);
pm_runtime_get_sync(dev);
unregister_netdev(ndev);
+ rtnl_lock();
+ /* A failed ndo_open has no matching ndo_stop. Its retained resources
+ * still need retirement, or software-only disconnection on timeout.
+ */
+ list_for_each_entry(dma_conf, &priv->dma_confs, list) {
+ free_dma_desc_resources(priv, dma_conf);
+ for (queue = 0; queue < MTL_MAX_RX_QUEUES; queue++) {
+ struct xdp_rxq_info *rxq = &dma_conf->rx_queue[queue].xdp_rxq;
+
+ if (xdp_rxq_info_is_reg(rxq))
+ xdp_rxq_info_unreg(rxq);
+ }
+ }
+ rtnl_unlock();
#ifdef CONFIG_DEBUG_FS
stmmac_exit_fs(ndev);
@@ -8872,12 +9148,7 @@ int stmmac_resume(struct device *dev)
mutex_lock(&priv->lock);
- stmmac_reset_queues_param(priv);
-
- stmmac_free_tx_skbufs(priv);
- stmmac_clear_descriptors(priv, priv->dma_conf);
-
- ret = stmmac_hw_setup(ndev);
+ ret = stmmac_hw_setup(ndev, false, false);
if (ret < 0) {
netdev_err(priv->dev, "%s: Hw setup failed\n", __func__);
goto error_stop_dma;
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 18/19] net: stmmac: prepare device-local DMA interrupt quiescence
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (16 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 17/19] net: stmmac: retain DMA memory until hardware shutdown completes James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 21:59 ` [PATCH net-next v5 19/19] net: stmmac: retain DMA resources across MTU changes James Hilliard
` (2 subsequent siblings)
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Add DMA interrupt-mask accessors for the supported cores and a
per-channel gate protected by the channel lock. A handler invoked on a
shared IRQ can then mask newly enabled device sources without
acknowledging pending events or touching rings being replaced.
The retained-ring MTU transaction will save and restore these masks and
synchronize the registered handlers. No interrupt-controller line needs
to be disabled, so other devices sharing the IRQ remain serviceable.
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c | 11 ++++++++
.../net/ethernet/stmicro/stmmac/dwmac1000_dma.c | 1 +
drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c | 1 +
drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c | 2 ++
drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h | 2 ++
drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c | 11 ++++++++
drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h | 2 ++
drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c | 10 ++++++++
drivers/net/ethernet/stmicro/stmmac/dwxgmac2_dma.c | 11 ++++++++
drivers/net/ethernet/stmicro/stmmac/hwif.h | 5 ++++
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 2 ++
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 30 ++++++++++++++++------
12 files changed, 80 insertions(+), 8 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c
index c9145441aab0..8a3f15134402 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-sun8i.c
@@ -374,6 +374,16 @@ static void sun8i_dwmac_disable_dma_irq(struct stmmac_priv *priv,
writel(value, ioaddr + EMAC_INT_EN);
}
+static u32 sun8i_dwmac_set_dma_irq_mask(struct stmmac_priv *priv,
+ void __iomem *ioaddr, u32 chan, u32 mask)
+{
+ u32 old_mask = readl(ioaddr + EMAC_INT_EN);
+
+ writel(mask, ioaddr + EMAC_INT_EN);
+ readl(ioaddr + EMAC_INT_EN);
+ return old_mask;
+}
+
static void sun8i_dwmac_dma_start_tx(struct stmmac_priv *priv,
void __iomem *ioaddr, u32 chan)
{
@@ -579,6 +589,7 @@ static const struct stmmac_dma_ops sun8i_dwmac_dma_ops = {
.enable_dma_transmission = sun8i_dwmac_enable_dma_transmission,
.enable_dma_irq = sun8i_dwmac_enable_dma_irq,
.disable_dma_irq = sun8i_dwmac_disable_dma_irq,
+ .set_irq_mask = sun8i_dwmac_set_dma_irq_mask,
.start_tx = sun8i_dwmac_dma_start_tx,
.stop_tx = sun8i_dwmac_dma_stop_tx,
.start_rx = sun8i_dwmac_dma_start_rx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c b/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c
index 4cb7e6c16bdd..2285eac69071 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac1000_dma.c
@@ -264,6 +264,7 @@ const struct stmmac_dma_ops dwmac1000_dma_ops = {
.enable_dma_reception = dwmac_enable_dma_reception,
.enable_dma_irq = dwmac_enable_dma_irq,
.disable_dma_irq = dwmac_disable_dma_irq,
+ .set_irq_mask = dwmac_set_dma_irq_mask,
.start_tx = dwmac_dma_start_tx,
.stop_tx = dwmac_dma_stop_tx,
.start_rx = dwmac_dma_start_rx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c b/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c
index 5ffd3c1471c4..41579d10af3c 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac100_dma.c
@@ -118,6 +118,7 @@ const struct stmmac_dma_ops dwmac100_dma_ops = {
.enable_dma_transmission = dwmac_enable_dma_transmission,
.enable_dma_irq = dwmac_enable_dma_irq,
.disable_dma_irq = dwmac_disable_dma_irq,
+ .set_irq_mask = dwmac_set_dma_irq_mask,
.start_tx = dwmac_dma_start_tx,
.stop_tx = dwmac_dma_stop_tx,
.start_rx = dwmac_dma_start_rx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c
index d7928678dee1..d5f0cc8b851d 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.c
@@ -582,6 +582,7 @@ const struct stmmac_dma_ops dwmac4_dma_ops = {
.dma_tx_mode = dwmac4_dma_tx_chan_op_mode,
.enable_dma_irq = dwmac4_enable_dma_irq,
.disable_dma_irq = dwmac4_disable_dma_irq,
+ .set_irq_mask = dwmac4_set_dma_irq_mask,
.start_tx = dwmac4_dma_start_tx,
.stop_tx = dwmac4_dma_stop_tx,
.start_rx = dwmac4_dma_start_rx,
@@ -613,6 +614,7 @@ const struct stmmac_dma_ops dwmac410_dma_ops = {
.dma_tx_mode = dwmac4_dma_tx_chan_op_mode,
.enable_dma_irq = dwmac4_enable_dma_irq,
.disable_dma_irq = dwmac4_disable_dma_irq,
+ .set_irq_mask = dwmac4_set_dma_irq_mask,
.start_tx = dwmac4_dma_start_tx,
.stop_tx = dwmac4_dma_stop_tx,
.start_rx = dwmac4_dma_start_rx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h
index 9352107204eb..edccc09f0b03 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_dma.h
@@ -184,6 +184,8 @@ void dwmac4_enable_dma_irq(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan, bool rx, bool tx);
void dwmac4_disable_dma_irq(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan, bool rx, bool tx);
+u32 dwmac4_set_dma_irq_mask(struct stmmac_priv *priv, void __iomem *ioaddr,
+ u32 chan, u32 mask);
void dwmac4_dma_start_tx(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan);
void dwmac4_dma_stop_tx(struct stmmac_priv *priv, void __iomem *ioaddr,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c b/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c
index 9af0565a9bca..477bdb081c52 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_lib.c
@@ -155,6 +155,17 @@ void dwmac4_disable_dma_irq(struct stmmac_priv *priv, void __iomem *ioaddr,
writel(value, ioaddr + DMA_CHAN_INTR_ENA(dwmac4_addrs, chan));
}
+u32 dwmac4_set_dma_irq_mask(struct stmmac_priv *priv, void __iomem *ioaddr,
+ u32 chan, u32 mask)
+{
+ const struct dwmac4_addrs *dwmac4_addrs = priv->plat->dwmac4_addrs;
+ u32 old_mask = readl(ioaddr + DMA_CHAN_INTR_ENA(dwmac4_addrs, chan));
+
+ writel(mask, ioaddr + DMA_CHAN_INTR_ENA(dwmac4_addrs, chan));
+ readl(ioaddr + DMA_CHAN_INTR_ENA(dwmac4_addrs, chan));
+ return old_mask;
+}
+
int dwmac4_dma_interrupt(struct stmmac_priv *priv, void __iomem *ioaddr,
struct stmmac_extra_stats *x, u32 chan, u32 dir)
{
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h b/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h
index 970495bccfd2..4726801253f5 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac_dma.h
@@ -150,6 +150,8 @@ void dwmac_enable_dma_irq(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan, bool rx, bool tx);
void dwmac_disable_dma_irq(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan, bool rx, bool tx);
+u32 dwmac_set_dma_irq_mask(struct stmmac_priv *priv, void __iomem *ioaddr,
+ u32 chan, u32 mask);
void dwmac_dma_start_tx(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan);
void dwmac_dma_stop_tx(struct stmmac_priv *priv, void __iomem *ioaddr,
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c b/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c
index bb907db8fca1..88d904ed4685 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac_lib.c
@@ -84,6 +84,16 @@ void dwmac_disable_dma_irq(struct stmmac_priv *priv, void __iomem *ioaddr,
writel(value, ioaddr + DMA_CHAN_INTR_ENA(chan));
}
+u32 dwmac_set_dma_irq_mask(struct stmmac_priv *priv, void __iomem *ioaddr,
+ u32 chan, u32 mask)
+{
+ u32 old_mask = readl(ioaddr + DMA_CHAN_INTR_ENA(chan));
+
+ writel(mask, ioaddr + DMA_CHAN_INTR_ENA(chan));
+ readl(ioaddr + DMA_CHAN_INTR_ENA(chan));
+ return old_mask;
+}
+
void dwmac_dma_start_tx(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan)
{
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwxgmac2_dma.c b/drivers/net/ethernet/stmicro/stmmac/dwxgmac2_dma.c
index ff83858ebc1f..30915d3f5230 100644
--- a/drivers/net/ethernet/stmicro/stmmac/dwxgmac2_dma.c
+++ b/drivers/net/ethernet/stmicro/stmmac/dwxgmac2_dma.c
@@ -256,6 +256,16 @@ static void dwxgmac2_disable_dma_irq(struct stmmac_priv *priv,
writel(value, ioaddr + XGMAC_DMA_CH_INT_EN(chan));
}
+static u32 dwxgmac2_set_dma_irq_mask(struct stmmac_priv *priv,
+ void __iomem *ioaddr, u32 chan, u32 mask)
+{
+ u32 old_mask = readl(ioaddr + XGMAC_DMA_CH_INT_EN(chan));
+
+ writel(mask, ioaddr + XGMAC_DMA_CH_INT_EN(chan));
+ readl(ioaddr + XGMAC_DMA_CH_INT_EN(chan));
+ return old_mask;
+}
+
static void dwxgmac2_dma_start_tx(struct stmmac_priv *priv,
void __iomem *ioaddr, u32 chan)
{
@@ -604,6 +614,7 @@ const struct stmmac_dma_ops dwxgmac210_dma_ops = {
.dma_tx_mode = dwxgmac2_dma_tx_mode,
.enable_dma_irq = dwxgmac2_enable_dma_irq,
.disable_dma_irq = dwxgmac2_disable_dma_irq,
+ .set_irq_mask = dwxgmac2_set_dma_irq_mask,
.start_tx = dwxgmac2_dma_start_tx,
.stop_tx = dwxgmac2_dma_stop_tx,
.start_rx = dwxgmac2_dma_start_rx,
diff --git a/drivers/net/ethernet/stmicro/stmmac/hwif.h b/drivers/net/ethernet/stmicro/stmmac/hwif.h
index 4b7381a6fcce..20de97ac011a 100644
--- a/drivers/net/ethernet/stmicro/stmmac/hwif.h
+++ b/drivers/net/ethernet/stmicro/stmmac/hwif.h
@@ -197,6 +197,9 @@ struct stmmac_dma_ops {
u32 chan, bool rx, bool tx);
void (*disable_dma_irq)(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan, bool rx, bool tx);
+ /* Replace and flush the full interrupt enable mask; return the old mask. */
+ u32 (*set_irq_mask)(struct stmmac_priv *priv, void __iomem *ioaddr,
+ u32 chan, u32 mask);
void (*start_tx)(struct stmmac_priv *priv, void __iomem *ioaddr,
u32 chan);
void (*stop_tx)(struct stmmac_priv *priv, void __iomem *ioaddr,
@@ -263,6 +266,8 @@ struct stmmac_dma_ops {
stmmac_do_void_callback(__priv, dma, enable_dma_irq, __priv, __args)
#define stmmac_disable_dma_irq(__priv, __args...) \
stmmac_do_void_callback(__priv, dma, disable_dma_irq, __priv, __args)
+#define stmmac_set_dma_irq_mask(__priv, __args...) \
+ stmmac_do_callback(__priv, dma, set_irq_mask, __priv, __args)
#define stmmac_start_tx(__priv, __args...) \
stmmac_do_void_callback(__priv, dma, start_tx, __priv, __args)
#define stmmac_stop_tx(__priv, __args...) \
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 090d79aeb2ad..06fe750624b6 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -146,6 +146,8 @@ struct stmmac_channel {
struct stmmac_priv *priv_data;
spinlock_t lock;
u32 index;
+ /* Protected by lock; IRQ handlers must not access the DMA rings. */
+ bool irq_quiesced;
};
struct stmmac_fpe_cfg {
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 98dbc873e1c8..435c76b7db30 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -3370,14 +3370,30 @@ static bool stmmac_safety_feat_interrupt(struct stmmac_priv *priv)
static int stmmac_napi_check(struct stmmac_priv *priv, u32 chan, u32 dir)
{
- int status = stmmac_dma_interrupt_status(priv, priv->ioaddr,
- &priv->xstats, chan, dir);
- struct stmmac_rx_queue *rx_q = &priv->dma_conf->rx_queue[chan];
- struct stmmac_tx_queue *tx_q = &priv->dma_conf->tx_queue[chan];
struct stmmac_channel *ch = &priv->channel[chan];
+ struct stmmac_rx_queue *rx_q;
+ struct stmmac_tx_queue *tx_q;
struct napi_struct *rx_napi;
struct napi_struct *tx_napi;
unsigned long flags;
+ int status;
+
+ spin_lock_irqsave(&ch->lock, flags);
+ if (unlikely(ch->irq_quiesced)) {
+ /* A shared IRQ may still invoke us, and DMA initialization can
+ * restore interrupt enables. Mask them again without acknowledging
+ * pending events or accessing the configuration being replaced.
+ */
+ stmmac_set_dma_irq_mask(priv, priv->ioaddr, chan, 0);
+ spin_unlock_irqrestore(&ch->lock, flags);
+ return 0;
+ }
+ spin_unlock_irqrestore(&ch->lock, flags);
+
+ status = stmmac_dma_interrupt_status(priv, priv->ioaddr,
+ &priv->xstats, chan, dir);
+ rx_q = &priv->dma_conf->rx_queue[chan];
+ tx_q = &priv->dma_conf->tx_queue[chan];
rx_napi = rx_q->xsk_pool ? &ch->rxtx_napi : &ch->rx_napi;
tx_napi = tx_q->xsk_pool ? &ch->rxtx_napi : &ch->tx_napi;
@@ -4096,8 +4112,7 @@ static void stmmac_free_irq(struct net_device *dev,
for (j = irq_idx - 1; msi && j >= 0; j--) {
if (msi->tx_irq[j] > 0) {
irq_set_affinity_hint(msi->tx_irq[j], NULL);
- free_irq(msi->tx_irq[j],
- &priv->channel[j]);
+ free_irq(msi->tx_irq[j], &priv->channel[j]);
}
}
irq_idx = priv->plat->rx_queues_to_use;
@@ -4106,8 +4121,7 @@ static void stmmac_free_irq(struct net_device *dev,
for (j = irq_idx - 1; msi && j >= 0; j--) {
if (msi->rx_irq[j] > 0) {
irq_set_affinity_hint(msi->rx_irq[j], NULL);
- free_irq(msi->rx_irq[j],
- &priv->channel[j]);
+ free_irq(msi->rx_irq[j], &priv->channel[j]);
}
}
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* [PATCH net-next v5 19/19] net: stmmac: retain DMA resources across MTU changes
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (17 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 18/19] net: stmmac: prepare device-local DMA interrupt quiescence James Hilliard
@ 2026-09-27 21:59 ` James Hilliard
2026-09-27 22:10 ` [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures Jakub Kicinski
2026-09-28 6:50 ` Maxime Chevallier
20 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 21:59 UTC (permalink / raw)
To: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, James Hilliard
Prepare the replacement configuration before quiescing the old datapath.
Retain its rings and IRQ registrations until setup succeeds so rollback
needs no new allocations or IRQ requests. Fill RX buffer holes before
reset without altering active descriptors; rebuild the retained rings
only after a successful reset.
Mask the device DMA interrupt sources, gate shared-IRQ handlers and
drain all registered handlers before the final transmitter and timer
cancellation. Restore saved interrupt masks only after the selected
rings and NAPI are ready. This leaves interrupt-controller lines
available to unrelated devices.
Program receive limits using the prospective MTU and restore the old MTU
before rollback. Reapply PHC and TC state before starting DMA,
preserving the PHC registration and packet timestamp filters. Continuous
PHC time across the reset is not preserved.
If rollback fails, leave the administratively-up interface detached in a
distinct HALTED state: rings retained, NAPI disabled and IRQ
registrations released. Close or a later down/up can finish cleanup and
recovery without freeing IRQs twice. Retain potentially active DMA
memory until hardware shutdown is confirmed.
Fixes: 3470079687448 ("net: ethernet: stmicro: stmmac: permit MTU change with interface up")
Signed-off-by: James Hilliard <james.hilliard1@gmail.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac.h | 2 +
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 216 +++++++++++++++++-----
2 files changed, 173 insertions(+), 45 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac.h b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
index 06fe750624b6..a65253d309d0 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac.h
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac.h
@@ -279,6 +279,8 @@ enum stmmac_datapath_state {
STMMAC_DATAPATH_RUNNING,
/* Resources retained, NAPI disabled, DMA stop requested. */
STMMAC_DATAPATH_SUSPENDED,
+ /* Failed MTU rollback: rings retained, but no IRQs or running NAPI. */
+ STMMAC_DATAPATH_HALTED,
};
struct stmmac_priv {
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index 435c76b7db30..7951b3e60d84 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -1016,7 +1016,7 @@ static void stmmac_release_ptp(struct stmmac_priv *priv)
}
/* ptp_mutex excludes configuration and crosstimestamp operations. The
- * spinlock also excludes atomic clock reads while changing this gate.
+ * spinlock also excludes atomic gettime callers while changing this gate.
*/
static void stmmac_block_ptp(struct stmmac_priv *priv, bool block)
{
@@ -2251,6 +2251,30 @@ static void stmmac_free_tx_skbufs(struct stmmac_priv *priv)
dma_free_tx_skbufs(priv, priv->dma_conf, queue);
}
+/* NAPI is stopped, but DMA may still be using the old rings. Fill holes in
+ * the software buffer array without changing any descriptors. If allocation
+ * fails, the old rings can continue unchanged. Otherwise rollback after a
+ * reset will not need to allocate buffers.
+ */
+static int stmmac_prepare_rx_buffers(struct stmmac_priv *priv)
+{
+ struct stmmac_dma_conf *dma_conf = priv->dma_conf;
+ u32 queue, i;
+ int ret;
+
+ for (queue = 0; queue < priv->plat->rx_queues_to_use; queue++) {
+ struct stmmac_rx_queue *rx_q = &dma_conf->rx_queue[queue];
+
+ for (i = 0; i < dma_conf->dma_rx_size; i++) {
+ ret = stmmac_alloc_rx_buffer(priv, rx_q, &rx_q->buf_pool[i]);
+ if (ret)
+ return ret;
+ }
+ }
+
+ return 0;
+}
+
/* Only after a successful DMA reset. MTU rollback has already filled holes;
* resume may need new page-pool buffers or a partially populated XSK ring.
*/
@@ -4425,14 +4449,38 @@ static void stmmac_synchronize_irq(struct stmmac_priv *priv)
synchronize_irq(msi->tx_irq[i]);
}
+/* Keep the IRQ registrations, but prevent DMA handlers from using the rings.
+ * The caller drains handlers after quiescing every channel and restores their
+ * masks only once the active DMA configuration is ready again.
+ */
+static void stmmac_set_dma_irq_state(struct stmmac_priv *priv, bool enable,
+ u32 *irq_mask)
+{
+ u32 channels = max(priv->plat->rx_queues_to_use,
+ priv->plat->tx_queues_to_use);
+ u32 chan;
+
+ for (chan = 0; chan < channels; chan++) {
+ struct stmmac_channel *ch = &priv->channel[chan];
+ unsigned long flags;
+
+ spin_lock_irqsave(&ch->lock, flags);
+ ch->irq_quiesced = !enable;
+ if (enable)
+ stmmac_set_dma_irq_mask(priv, priv->ioaddr, chan,
+ irq_mask[chan]);
+ else
+ irq_mask[chan] = stmmac_set_dma_irq_mask(priv, priv->ioaddr,
+ chan, 0);
+ spin_unlock_irqrestore(&ch->lock, flags);
+ }
+}
+
/**
- * stmmac_setup_dma_desc - Generate a dma_conf and allocate DMA queue
- * @priv: driver private structure
- * @mtu: MTU to setup the dma queue and buf with
- * Description: Allocate and generate a dma_conf based on the provided MTU.
- * Allocate the Tx/Rx DMA queue and init them.
- * Return value:
- * the dma_conf allocated struct on success and an appropriate ERR_PTR on failure.
+ * stmmac_setup_dma_desc - allocate and initialize a DMA configuration
+ * @priv: driver private structure
+ * @mtu: MTU to size the receive buffers for
+ * Return: the allocated configuration, or an ERR_PTR on failure
*/
static struct stmmac_dma_conf *
stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu)
@@ -4666,7 +4714,7 @@ static int __stmmac_open(struct net_device *dev,
priv->dma_conf = dma_conf;
/* The PHY is suspended when the interface is reopened without
- * disconnecting the PHY, e.g. on MTU change. IEEE 802.3 allows PHYs
+ * disconnecting the PHY, e.g. on an ethtool change. IEEE 802.3 allows PHYs
* to stop their receive clock while powered down, but the DMA
* software reset in stmmac_hw_setup() requires a running receive
* clock, and phylink_start() below resumes the PHY only after the
@@ -4823,20 +4871,22 @@ static void stmmac_quiesce(struct stmmac_priv *priv)
static void __stmmac_release(struct net_device *dev)
{
struct stmmac_priv *priv = netdev_priv(dev);
+ enum stmmac_datapath_state state = priv->datapath;
- /* A failed MTU reopen has already released the data path. */
+ /* There may be no resources left after detached XDP reconfiguration. */
if (priv->datapath == STMMAC_DATAPATH_DOWN)
return;
phylink_stop(priv->phylink);
- /* Suspend retains the resources, but has already stopped activity. */
+ /* SUSPENDED and HALTED retain rings with NAPI already disabled. */
if (priv->datapath == STMMAC_DATAPATH_RUNNING)
stmmac_quiesce(priv);
priv->datapath = STMMAC_DATAPATH_DOWN;
/* Free the IRQ lines */
- stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
+ if (state != STMMAC_DATAPATH_HALTED)
+ stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
/* Drain any final IRQ-triggered network activity before DMA shutdown. */
stmmac_stop_tx_queues(priv);
@@ -6723,6 +6773,109 @@ static void stmmac_set_rx_mode(struct net_device *dev)
stmmac_set_filter(priv, priv->hw, dev);
}
+static int stmmac_reconfigure_mtu(struct net_device *dev, int mtu)
+{
+ struct stmmac_priv *priv = netdev_priv(dev);
+ struct stmmac_dma_conf *old_conf = priv->dma_conf;
+ struct stmmac_dma_conf *new_conf;
+ int old_mtu = dev->mtu;
+ int ret, restore_ret;
+ u32 irq_mask[STMMAC_CH_MAX];
+ u32 chan;
+
+ new_conf = stmmac_setup_dma_desc(priv, mtu);
+ if (IS_ERR(new_conf))
+ return PTR_ERR(new_conf);
+
+ mutex_lock(&priv->ptp_mutex);
+ stmmac_block_ptp(priv, true);
+ netif_device_detach(dev);
+ phylink_stop(priv->phylink);
+ stmmac_quiesce(priv);
+ if (stmmac_fpe_supported(priv))
+ ethtool_mmsv_stop(&priv->fpe_cfg.mmsv);
+
+ /* Drain handlers before the final TX stop and configuration swap,
+ * and keep the registrations for rollback.
+ */
+ stmmac_set_dma_irq_state(priv, false, irq_mask);
+ stmmac_synchronize_irq(priv);
+ stmmac_stop_tx_queues(priv);
+
+ ret = stmmac_prepare_rx_buffers(priv);
+ if (ret)
+ goto restart;
+
+ stmmac_stop_all_dma(priv);
+ phylink_prepare_resume(priv->phylink);
+
+ /* MAC receive limits must be programmed for the prospective MTU. */
+ WRITE_ONCE(dev->mtu, mtu);
+ priv->dma_conf = new_conf;
+ stmmac_reset_queues_param(priv);
+ ret = stmmac_hw_setup(dev, false, true);
+ if (ret) {
+ stmmac_stop_all_dma(priv);
+ stmmac_mac_set(priv, priv->ioaddr, false);
+ priv->dma_conf = old_conf;
+ WRITE_ONCE(dev->mtu, old_mtu);
+
+ /* Reuse the retained rings. Reinitialize them only after reset
+ * has completed, not merely after clearing the DMA enable bits.
+ */
+ restore_ret = stmmac_hw_setup(dev, true, true);
+ if (restore_ret) {
+ stmmac_stop_all_dma(priv);
+ stmmac_mac_set(priv, priv->ioaddr, false);
+ /* Setup may have restored DMA interrupt enables. */
+ stmmac_set_dma_irq_state(priv, false, irq_mask);
+ stmmac_free_irq(dev, REQ_IRQ_ERR_ALL, 0);
+ stmmac_stop_tx_queues(priv);
+ stmmac_stop_all_dma(priv);
+ memset(irq_mask, 0, sizeof(irq_mask));
+ stmmac_set_dma_irq_state(priv, true, irq_mask);
+ priv->datapath = STMMAC_DATAPATH_HALTED;
+ netdev_err(dev, "MTU rollback failed: %pe; interface remains detached\n",
+ ERR_PTR(restore_ret));
+ goto free_new;
+ }
+ } else {
+ /* Hardware setup completed its reset before using the new rings.
+ * The old DMA allocations can now be released safely.
+ */
+ stmmac_put_dma_conf(priv, old_conf);
+ for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
+ hrtimer_setup(&new_conf->tx_queue[chan].txtimer,
+ stmmac_tx_timer, CLOCK_MONOTONIC,
+ HRTIMER_MODE_REL);
+ }
+
+ stmmac_set_rx_mode(dev);
+ stmmac_vlan_restore(priv);
+ stmmac_start_all_dma(priv);
+
+restart:
+ stmmac_block_ptp(priv, false);
+ mutex_unlock(&priv->ptp_mutex);
+ stmmac_enable_all_queues(priv);
+ stmmac_set_dma_irq_state(priv, true, irq_mask);
+ stmmac_enable_all_dma_irq(priv);
+ phylink_start(priv->phylink);
+ netif_device_attach(dev);
+ for (chan = 0; chan < priv->plat->tx_queues_to_use; chan++)
+ stmmac_tx_timer_arm(priv, chan);
+ if (!ret)
+ return 0;
+ goto free_conf;
+
+free_new:
+ /* Failed rollback leaves the registered PHC inaccessible as well. */
+ mutex_unlock(&priv->ptp_mutex);
+free_conf:
+ stmmac_put_dma_conf(priv, new_conf);
+ return ret;
+}
+
/**
* stmmac_change_mtu - entry point to change MTU size for the device.
* @dev : device pointer.
@@ -6737,9 +6890,7 @@ static void stmmac_set_rx_mode(struct net_device *dev)
static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
{
struct stmmac_priv *priv = netdev_priv(dev);
- struct stmmac_dma_conf *old_conf = priv->dma_conf;
int txfifosz = priv->plat->tx_fifo_size;
- struct stmmac_dma_conf *dma_conf;
const int mtu = new_mtu;
int ret;
@@ -6765,35 +6916,9 @@ static int stmmac_change_mtu(struct net_device *dev, int new_mtu)
*/
if (netif_running(dev) &&
(dev->mtu > ETH_DATA_LEN || mtu > ETH_DATA_LEN)) {
- netdev_dbg(priv->dev, "restarting interface to change its MTU\n");
- /* Try to allocate the new DMA conf with the new mtu */
- dma_conf = stmmac_setup_dma_desc(priv, mtu);
- if (IS_ERR(dma_conf)) {
- netdev_err(priv->dev, "failed allocating new dma conf for new MTU %d\n",
- mtu);
- return PTR_ERR(dma_conf);
- }
-
- netif_device_detach(dev);
- __stmmac_release(dev);
-
- ret = __stmmac_open(dev, dma_conf);
- if (ret) {
- priv->dma_conf = old_conf;
- stmmac_put_dma_conf(priv, dma_conf);
- /*
- * Keep the administrative state and PHY/PM ownership until
- * ndo_stop(), but prevent use of the released data path.
- */
- netif_device_detach(dev);
- netdev_err(priv->dev, "failed reopening the interface after MTU change\n");
+ ret = stmmac_reconfigure_mtu(dev, mtu);
+ if (ret)
return ret;
- }
-
- stmmac_put_dma_conf(priv, old_conf);
-
- stmmac_set_rx_mode(dev);
- netif_device_attach(dev);
}
WRITE_ONCE(dev->mtu, mtu);
@@ -7632,11 +7757,12 @@ static int stmmac_bpf(struct net_device *dev, struct netdev_bpf *bpf)
return -EOPNOTSUPP;
/*
- * Pool removal must succeed even after a failed resume. Release the
- * suspended rings before their pool or XDP buffer layout can change.
+ * Pool removal must succeed after failed resume or MTU rollback. Release
+ * retained rings before their pool or XDP buffer layout can change.
* Leave the interface detached until it is closed and reopened.
*/
- if (priv->datapath == STMMAC_DATAPATH_SUSPENDED)
+ if (priv->datapath == STMMAC_DATAPATH_SUSPENDED ||
+ priv->datapath == STMMAC_DATAPATH_HALTED)
__stmmac_release(dev);
switch (bpf->command) {
--
2.53.0
^ permalink raw reply [flat|nested] 27+ messages in thread* Re: [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (18 preceding siblings ...)
2026-09-27 21:59 ` [PATCH net-next v5 19/19] net: stmmac: retain DMA resources across MTU changes James Hilliard
@ 2026-09-27 22:10 ` Jakub Kicinski
2026-09-27 23:15 ` James Hilliard
2026-09-28 6:50 ` Maxime Chevallier
20 siblings, 1 reply; 27+ messages in thread
From: Jakub Kicinski @ 2026-09-27 22:10 UTC (permalink / raw)
To: James Hilliard
Cc: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel, Richard Genoud,
Alastair D'Silva, Maxime Ripard, netdev, linux-kernel,
linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, Rayagond Kokatanur, Ding Hui
On Sun, 27 Sep 2026 15:59:35 -0600 James Hilliard wrote:
> Keep the stmmac datapath coherent after failed MTU changes or hardware
> resume without changing the interface's administrative state. Retain the
> working MTU configuration for rollback, and allow ordinary down/up
> recovery when hardware cannot be restored.
Please read:
https://www.kernel.org/doc/html/latest/process/maintainer-netdev.html
The limit is 15 patches per tree, you have 22 outstanding to
net-next right now.
Violating the limit again will result in a temporary ban.
^ permalink raw reply [flat|nested] 27+ messages in thread* Re: [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures
2026-09-27 22:10 ` [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures Jakub Kicinski
@ 2026-09-27 23:15 ` James Hilliard
0 siblings, 0 replies; 27+ messages in thread
From: James Hilliard @ 2026-09-27 23:15 UTC (permalink / raw)
To: Jakub Kicinski
Cc: Russell King, Andrew Lunn, Heiner Kallweit, David S. Miller,
Eric Dumazet, Paolo Abeni, Russell King (Oracle),
Maxime Chevallier, Andrew Lunn, Maxime Coquelin,
Alexandre Torgue, Christian Marangi, Tiezhu Yang, Huacai Chen,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Serge Semin, Suraj Jaiswal,
Richard Cochran, Joao Pinto, Vladimir Oltean, Ong Boon Leong,
Voon Weifeng, Song, Yoong Siang, Linus Walleij,
Martin Blumenstingl, Magnus Karlsson, Maciej Fijalkowski,
Simon Horman, Björn Töpel, Thierry Reding,
Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec, Samuel Holland,
Jose Abreu, Yao Zi, Philipp Zabel, Richard Genoud,
Alastair D'Silva, Maxime Ripard, netdev, linux-kernel,
linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, Rayagond Kokatanur, Ding Hui
On Sun, Sep 27, 2026 at 4:10 PM Jakub Kicinski <kuba@kernel.org> wrote:
>
> On Sun, 27 Sep 2026 15:59:35 -0600 James Hilliard wrote:
> > Keep the stmmac datapath coherent after failed MTU changes or hardware
> > resume without changing the interface's administrative state. Retain the
> > working MTU configuration for rollback, and allow ordinary down/up
> > recovery when hardware cannot be restored.
>
> Please read:
> https://www.kernel.org/doc/html/latest/process/maintainer-netdev.html
>
> The limit is 15 patches per tree, you have 22 outstanding to
> net-next right now.
Oh, I had bundled some prerequisite patches to avoid breaking the review
bot since my fix for prerequisite handling for the review bot doesn't appear
to be live quite yet(it was merged but the live server looks to be on an older
tag still).
https://github.com/sashiko-dev/sashiko/pull/389
That just got merged so that shouldn't be an issue going forwards as I
assume it should go live soon.
This was originally on the net tree but I had switched to net-next since
some additional prerequisite commits were only in net-next and I was
trying to avoid conflicts or having to bundle the patches to avoid breaking
the review bot.
https://lore.kernel.org/all/8ed1298a-4c29-4db3-b76e-c8d34f91b4c9@bootlin.com/
I guess I could move this back to net and specify the prerequisites with b4
metadata. It isn't entirely clear to me which branch is appropriate for this
however.
^ permalink raw reply [flat|nested] 27+ messages in thread
* Re: [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures
2026-09-27 21:59 [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures James Hilliard
` (19 preceding siblings ...)
2026-09-27 22:10 ` [PATCH net-next v5 00/19] net: stmmac: preserve datapath state across MTU and resume failures Jakub Kicinski
@ 2026-09-28 6:50 ` Maxime Chevallier
20 siblings, 0 replies; 27+ messages in thread
From: Maxime Chevallier @ 2026-09-28 6:50 UTC (permalink / raw)
To: James Hilliard, Russell King, Andrew Lunn, Heiner Kallweit,
David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Russell King (Oracle),
Andrew Lunn, Maxime Coquelin, Alexandre Torgue,
Christian Marangi, Tiezhu Yang, Huacai Chen, Alexei Starovoitov,
Daniel Borkmann, Jesper Dangaard Brouer, John Fastabend,
Stanislav Fomichev, Suraj Jaiswal, Richard Cochran, Joao Pinto,
Vladimir Oltean, Ong Boon Leong, Voon Weifeng, Song, Yoong Siang,
Linus Walleij, Martin Blumenstingl, Magnus Karlsson,
Maciej Fijalkowski, Simon Horman, Björn Töpel,
Thierry Reding, Jonathan Hunter, Chen-Yu Tsai, Jernej Skrabec,
Samuel Holland, Jose Abreu, Yao Zi, Philipp Zabel
Cc: Richard Genoud, Alastair D'Silva, Maxime Ripard, netdev,
linux-kernel, linux-stm32, linux-arm-kernel, bpf, ZhaoJinming,
Lorenzo Bianconi, Ding Hui, Linkui Xiao, Linkui Xiao,
linux-tegra, linux-sunxi, stable, Rayagond Kokatanur, Ding Hui
Hi James,
On 9/27/26 23:59, James Hilliard wrote:
> Keep the stmmac datapath coherent after failed MTU changes or hardware
> resume without changing the interface's administrative state. Retain the
> working MTU configuration for rollback, and allow ordinary down/up
> recovery when hardware cannot be restored.
>
> This revision incorporates the following previously posted fixes as
> prerequisites:
>
> - Linkui Xiao's v3 MDIO reset GPIO fix:
> https://lore.kernel.org/r/20260921015727.2643540-1-xiaolinkui@126.com
> - Ding Hui's v3 DMA descriptor allocation cleanup:
> https://lore.kernel.org/r/20260919121436.1642724-1-dinghui1111@163.com
> - Lorenzo Bianconi's v3 EST series, patches 2-4:
> https://lore.kernel.org/r/20260902-stmmac-est-reapply-after-open-v3-0-e72a6df5a7ef@oss.qualcomm.com
Then please wait for them to be merged.
Also why are you targetting net-next now ?
And most important, no more than 15 outstanding patches,
this series alone goes over the limit.
You need to split this up, tackle the issues one at a time, even
if the overall goal is "MTU and resume failures".
As you can see there's a lot of backlog in this driver, and if you try
to address sashiko's review by adding more patches and account for
all "pre-existing issues", this won't work.
This series already went from a 3-ish series that was
sent alongside the allwinner glue to 19 patches in about a week.
Please think about the people who have to review that :
> 39 files changed, 2390 insertions(+), 808 deletions(-)
Not only the quantity, but the quality of the commit logs, the changelogs
and the comments. This really feels like the raw output of an LLM and it's
very hard to read.
Maxime
^ permalink raw reply [flat|nested] 27+ messages in thread