* [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA
@ 2026-09-09 13:00 Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 01/14] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
` (13 more replies)
0 siblings, 14 replies; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When lan966x operates as a PCIe endpoint, the driver currently uses
register-based I/O for frame injection and extraction. This approach is
functional but slow, topping out at around 33 Mbps on an Intel x86 host
with a lan966x PCIe card.
This series adds FDMA (Frame DMA) support for the PCIe path. When
operating as a PCIe endpoint, the internal FDMA engine on lan966x cannot
directly access host memory, so DMA buffers are allocated as contiguous
coherent memory and mapped through the PCIe Address Translation Unit
(ATU). The ATU provides outbound windows that translate internal FDMA
addresses to PCIe bus addresses, allowing the FDMA engine to read and
write host memory. Because the ATU requires contiguous address regions,
page_pool and normal per-page DMA mappings cannot be used. Instead,
frames are transferred using memcpy between the ATU-mapped buffers and
the network stack. With this, throughput increases from ~33 Mbps to
~620 Mbps for default MTU.
Patch 1 adds the shared drivers/net/ethernet/microchip/fdma/ directory
to the Sparx5 SoC MAINTAINERS entry.
Patches 2-3 prepare the shared FDMA library: patch 2 renames the
contiguous dataptr helpers for clarity, and patch 3 adds PCIe ATU
region management and coherent DMA allocation with ATU mapping.
Patches 4-7 refactor the lan966x FDMA code to support both platform
and PCIe paths: extracting the LLP register write into a helper,
exporting shared functions, introducing a dedicated device for DMA
operations, and adding an ops dispatch table selected at probe time.
Patches 8-9 harden the existing FDMA path for the PCIe endpoint
lifecycle: patch 8 clears latched FDMA error/interrupt stickies after
the switch reset so they don't assert as soon as interrupts are
enabled, and patch 9 adds a shutdown() callback that quiesces the
FDMA engine on host warm reboot (on the PCIe card the FDMA survives
host reset and would otherwise keep the shared INTx asserted into
the next probe).
Patch 10 adds the core PCIe FDMA implementation with RX/TX using
contiguous ATU-mapped buffers. Patches 11 and 12 extend it with MTU
change and XDP support respectively. XDP_PASS, XDP_TX, XDP_DROP and
XDP_ABORTED are supported; XDP_REDIRECT is deliberately not, because
the PCIe data path does not use page_pool.
Patches 13-14 update the lan966x PCI device tree overlay to extend the
cpu register mapping to cover the ATU register space and add the FDMA
interrupt.
Patches 1-12 touch MAINTAINERS and drivers/net/ethernet/microchip/, and
are for the netdev tree.
Patches 13-14 touch drivers/misc/lan966x_pci.dtso, and are for the
char-misc tree.
To: Andrew Lunn <andrew+netdev@lunn.ch>
To: David S. Miller <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Horatiu Vultur <horatiu.vultur@microchip.com>
To: Steen Hegelund <steen.hegelund@microchip.com>
To: UNGLinuxDriver@microchip.com
To: Alexei Starovoitov <ast@kernel.org>
To: Daniel Borkmann <daniel@iogearbox.net>
To: Jesper Dangaard Brouer <hawk@kernel.org>
To: John Fastabend <john.fastabend@gmail.com>
To: Stanislav Fomichev <sdf@fomichev.me>
To: Herve Codina <herve.codina@bootlin.com>
To: Arnd Bergmann <arnd@arndb.de>
To: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
To: Mohsin Bashir <mohsin.bashr@gmail.com>
Cc: Richard Cochran <richardcochran@gmail.com>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: bpf@vger.kernel.org
Cc: linux-arm-kernel@lists.infradead.org
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
Changes in v6:
The main change is a DMA bug on the PCIe path, found after enabling the
IOMMU on the test host. The rest is hardening and cleanup.
- New patch 6: use a dedicated device for DMA operations. On the PCIe
path lan966x->dev is the platform device created by
of_platform_default_populate(), which carries no iommus/dma-ranges of
its own, so it is not the device the IOMMU has a domain for. DMA
mapped against it lands outside the domain the endpoint's requester ID
is associated with. With the IOMMU enabled, AMD-Vi reported
IO_PAGE_FAULT for the endpoint and no traffic passed. All FDMA DMA now
targets the real PCIe endpoint device. Retested with the IOMMU in
translated (DMA-FQ) mode.
- ATU: serialise region allocation/free and ATU register access with a
mutex.
- ATU: program the region limit per mapping (base_addr + size - 1)
instead of the full region bound, and disable every region in
fdma_pci_atu_init().
- ATU: return -ERANGE instead of -E2BIG when the requested size exceeds
the region size, and clear fdma->atu_region on unmap.
- PCIe FDMA: check fdma_dcbs_init() and unwind the coherent allocation
and the ATU mapping on failure.
- PCIe FDMA: drop the WARN_ON() on an out-of-range src_port; a malformed
IFH should not be able to splat the host.
- Quiesce consistently in shutdown, deinit and the MTU reload: stop the
netdev queues and NAPI before disabling the FDMA channels, and share
one lan966x_fdma_tx_disable_netdev() helper instead of open-coding it
per path.
- shutdown: also mask ANA_ANAINTR, which shares the PCIe INTx with the
FDMA sources.
- MTU change: reject an MTU whose ring would not fit in a single
MAX_PAGE_ORDER coherent block, and reject one whose db_size would be
truncated by the 16-bit DCB DATAL field, rather than programming a
wrong buffer length.
- Assign lan966x->ops directly in the ops dispatch patch, so that patch
no longer adds a helper returning a single constant; PCI detection now
derives from dma_dev and arrives with the PCIe implementation.
- ATU: reject a target address that is not aligned to the 64KB ATU region
granularity, instead of programming a translation offset by the
misalignment.
- Drop the CONFIG_MCHP_LAN966X_PCI guard around the fdma_pci.h include in
fdma_api.h (Jakub).
- Use ifneq ($(CONFIG_MCHP_LAN966X_PCI),) instead of ifdef in both the fdma
and lan966x Makefiles; the symbol is tristate, so the condition has to
cover both y and m (Jakub).
- Reprogram the FDMA LLP registers on the platform reload restore path,
which stopped happening when the LLP write moved into the allocation
functions.
- XDP: drop the napi_synchronize() before freeing the old program on the
PCIe path; READ_ONCE() on port->xdp_prog plus the RCU-deferred
bpf_prog_put() already cover the in-flight poll.
- XDP: note that the ETH_ZLEN floor is deliberately not enforced on XDP_TX.
- PCIe FDMA: count tx_dropped when the DCB ring is full.
- PCIe FDMA: guard the deinit napi_disable() on fdma_ndev, as shutdown()
already does; NAPI is only initialised on the first ndo_open.
- ATU: program the region translation before enabling the region.
- ATU: pad the mapped allocation to the 64KB outbound region granularity,
so the window does not extend past the coherent allocation.
- ATU: reject mapping an fdma that already holds a region, instead of
overwriting the handle and leaking the old mapping. The MTU reload clears
the handle on the live rings first, since it keeps the old mapping alive
until the new rings are in place.
- PCIe FDMA: declare the RX DCB DATAL as the space actually available to the
extraction engine (db_size - XDP_PACKET_HEADROOM), since the data block
pointer starts that far into the slot.
- Move the XDP-unsupported guards from the XDP patch into the PCIe FDMA
patch, so that no intermediate commit advertises NETDEV_XDP_ACT_NDO_XMIT
for a PCIe instance while tx->dcbs_buf is NULL, and none lets an XDP
attach reach the platform page_pool reload. The final tree is unchanged.
- shutdown: take rtnl, so it cannot interleave with the rtnl-only reload
paths and double-disable the NAPI, which would spin forever in
napi_disable().
- Link to v5: https://lore.kernel.org/r/20260520-lan966x-pci-fdma-v5-0-ca56197ae05b@microchip.com
Changes in v5:
This version fixes a single AI review issue, flagged by Paolo. Other AI
issues for v4 has been classified as pre-existing or changes for
follow-ups.
- Fix premature napi_complete_done() in lan966x_fdma_pci_napi_poll() on
FDMA_ERROR and napi_alloc_skb() failure. Bailing out left DONE=1 DCBs
in the ring with no IRQ to drain them. Drop the frame and continue
the poll loop instead. Bump rx_dropped on memory-pressure drop.
(Paolo)
- Link to v4: https://lore.kernel.org/r/20260508-lan966x-pci-fdma-v4-0-14e0c89d8d63@microchip.com
Changes in v4:
- Consolidate rx size checks into lan966x_fdma_pci_rx_size_fits().
Subtract XDP_PACKET_HEADROOM on the max size check, and add ETH_HLEN
on the min size check. This fixes potential OOB reads/writes.
- On xdp_prepare_buff(), update comment to clarify that data is already
offset by XDP_PACKET_HEADROOM.
- Link to v3:
https://lore.kernel.org/r/20260504-lan966x-pci-fdma-v3-0-a56f5740d870@microchip.com
Changes in v3:
Version 3 fixes a number of issues reported by sashiko - mostly
hardening.
- Fix double use of XDP_PACKET_HEADROOM.
- Fix ERR_PTR persistence in fdma->atu_region and add missing
NULL/ERR_PTR guard in fdma_pci_atu_region_unmap().
- Reject size <= 0 in fdma_pci_atu_region_map() and return
-ENOSPC (was -ENOMEM) when no region is free.
- Introduce lan966x_fdma_pci_tx_size_fits() that accounts for
XDP_PACKET_HEADROOM; use it from both xmit paths to keep
bpf_xdp_adjust_tail from writing past the TX slot.
- Validate BLOCKL in rx_check_frame() (reject < IFH+FCS or
> db_size) before it feeds memcpy/XDP sizes.
- READ_ONCE(port->xdp_prog) inside lan966x_xdp_pci_run() to close
a TOCTOU on XDP detach that could deref NULL in
bpf_prog_run_xdp().
- Strip IFH and FCS pre-XDP in rx_check_frame(). After BPF runs
the driver cannot tell whether the tail was modified; drop the
unconditional skb_pull/skb_trim in rx_get_frame().
- Account tx_bytes/tx_packets on XDP_TX success and tx_dropped on
XDP_TX size reject.
- Add dma_wmb()/dma_rmb() around DCB status writes and reads in
xmit, xmit_xdpf, and napi_poll.
- Collected Tested-by: Hervé Codina.
- Link to v2: https://lore.kernel.org/r/20260428-lan966x-pci-fdma-v2-0-d3ec66e06202@microchip.com
Changes in v2:
Version 2 primarily addresses issues with module unload/load, where
traffic would stop working (Hervé), and XDP head/tail adjust that would be
discarded (Mohsin).
Apart from that, I ran through issues reported by Sashiko, and fixed a
number of other issues.
- New patch 1: add drivers/net/ethernet/microchip/fdma/ to the Sparx5
SoC MAINTAINERS entry.
- New patch 7: clear latched FDMA error/interrupt stickies after the
switch reset so they don't fire as soon as interrupts are enabled.
- New patch 8: shutdown() callback, quiescing FDMA on host warm reboot.
- Replaced the depth-2 dev_is_pci(parent->parent) backend selector
with a parent-chain walk.
- XDP: use xdp.data/xdp.data_end for the post-XDP frame length so that
bpf_xdp_adjust_head/tail are respected (Mohsin Bashir)
- MTU change: drain in-flight xmits with netif_tx_disable() on every
port before reallocating rings, waking them again on completion.
- MTU change: cap the PCIe DCB ring at 256 entries so a full-ring
coherent DMA allocation fits in a single MAX_PAGE_ORDER block at
jumbo MTU.
- PCIe ATU: disable the region before clearing its translation on
unmap.
- PCIe FDMA: hold tx_lock in napi_poll around the free-DCB check used
to wake stopped netdev queues.
- PCIe FDMA: return -ENOSPC (not -1) when the DCB ring is exhausted.
- Link to v1: https://lore.kernel.org/r/20260320-lan966x-pci-fdma-v1-0-ef54cb9b0c4b@microchip.com
---
Daniel Machon (14):
MAINTAINERS: add FDMA library to Sparx5 SoC entry
net: microchip: fdma: rename contiguous dataptr helpers
net: microchip: fdma: add PCIe ATU support
net: lan966x: add FDMA LLP register write helper
net: lan966x: export FDMA helpers for reuse
net: lan966x: use a dedicated device for DMA operations
net: lan966x: add FDMA ops dispatch for PCIe support
net: lan966x: clear FDMA interrupt stickies after switch reset
net: lan966x: add shutdown callback to stop FDMA on reboot
net: lan966x: add PCIe FDMA support
net: lan966x: add PCIe FDMA MTU change support
net: lan966x: add PCIe FDMA XDP support
misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
misc: lan966x-pci: dts: add fdma interrupt to overlay
MAINTAINERS | 1 +
drivers/misc/lan966x_pci.dtso | 5 +-
drivers/net/ethernet/microchip/fdma/Makefile | 4 +
drivers/net/ethernet/microchip/fdma/fdma_api.c | 44 ++
drivers/net/ethernet/microchip/fdma/fdma_api.h | 22 +-
drivers/net/ethernet/microchip/fdma/fdma_pci.c | 203 ++++++
drivers/net/ethernet/microchip/fdma/fdma_pci.h | 52 ++
drivers/net/ethernet/microchip/lan966x/Makefile | 4 +
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 97 +--
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 692 +++++++++++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.c | 88 ++-
.../net/ethernet/microchip/lan966x/lan966x_main.h | 42 ++
.../net/ethernet/microchip/lan966x/lan966x_regs.h | 25 +
.../net/ethernet/microchip/lan966x/lan966x_xdp.c | 6 +
14 files changed, 1229 insertions(+), 56 deletions(-)
---
base-commit: 548b86839f7fb819a4d6c83b71c73ec378d24275
change-id: 20260313-lan966x-pci-fdma-94ed485d23fa
Best regards,
--
Daniel Machon <daniel.machon@microchip.com>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 01/14] MAINTAINERS: add FDMA library to Sparx5 SoC entry
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 02/14] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
` (12 subsequent siblings)
13 siblings, 0 replies; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
The FDMA library under drivers/net/ethernet/microchip/fdma/ is shared by
the lan966x, sparx5 and lan969x drivers, but is not covered by an entry
in the MAINTAINERS file. A subsequent patch will add new files to the
FDMA library, so let's make sure it's covered.
Add drivers/net/ethernet/microchip/fdma/ to the Sparx5 SoC entry, since
I am already listed there.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
MAINTAINERS | 1 +
1 file changed, 1 insertion(+)
diff --git a/MAINTAINERS b/MAINTAINERS
index b23fb6f2f4ef..e41d4337136d 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -3214,6 +3214,7 @@ M: UNGLinuxDriver@microchip.com
L: linux-arm-kernel@lists.infradead.org (moderated for non-subscribers)
S: Supported
F: arch/arm64/boot/dts/microchip/sparx*
+F: drivers/net/ethernet/microchip/fdma/
F: drivers/net/ethernet/microchip/vcap/
F: drivers/pinctrl/pinctrl-microchip-sgpio.c
N: sparx5
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 02/14] net: microchip: fdma: rename contiguous dataptr helpers
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 01/14] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 03/14] net: microchip: fdma: add PCIe ATU support Daniel Machon
` (11 subsequent siblings)
13 siblings, 0 replies; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When the FDMA library was introduced [1], two helpers to get the DMA and
virtual address of a data buffer, in contiguous memory, were added. These
helpers have had no callers until this series. I found the naming I
initially used confusing and inconsistent.
Rename fdma_dataptr_get_contiguous() and
fdma_dataptr_virt_get_contiguous() to fdma_dataptr_dma_addr_contiguous()
and fdma_dataptr_virt_addr_contiguous(). This makes the pair symmetric
and clarifies what type of address each returns.
[1]: commit 30e48a75df9c ("net: microchip: add FDMA library")
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/net/ethernet/microchip/fdma/fdma_api.h | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index d91affe8bd98..dea8e3cc155a 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -197,8 +197,8 @@ static inline int fdma_nextptr_cb(struct fdma *fdma, int dcb_idx, u64 *nextptr)
* if the dataptr addresses and DCB's are in contiguous memory and the driver
* supports XDP.
*/
-static inline u64 fdma_dataptr_get_contiguous(struct fdma *fdma, int dcb_idx,
- int db_idx)
+static inline u64 fdma_dataptr_dma_addr_contiguous(struct fdma *fdma,
+ int dcb_idx, int db_idx)
{
return fdma->dma + (sizeof(struct fdma_dcb) * fdma->n_dcbs) +
(dcb_idx * fdma->n_dbs + db_idx) * fdma->db_size +
@@ -209,8 +209,8 @@ static inline u64 fdma_dataptr_get_contiguous(struct fdma *fdma, int dcb_idx,
* applicable if the dataptr addresses and DCB's are in contiguous memory and
* the driver supports XDP.
*/
-static inline void *fdma_dataptr_virt_get_contiguous(struct fdma *fdma,
- int dcb_idx, int db_idx)
+static inline void *fdma_dataptr_virt_addr_contiguous(struct fdma *fdma,
+ int dcb_idx, int db_idx)
{
return (u8 *)fdma->dcbs + (sizeof(struct fdma_dcb) * fdma->n_dcbs) +
(dcb_idx * fdma->n_dbs + db_idx) * fdma->db_size +
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 03/14] net: microchip: fdma: add PCIe ATU support
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 01/14] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 02/14] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 04/14] net: lan966x: add FDMA LLP register write helper Daniel Machon
` (10 subsequent siblings)
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When lan966x or lan969x operates as a PCIe endpoint, the internal FDMA
engine cannot directly access host memory. Instead, DMA addresses must
be translated through the PCIe Address Translation Unit (ATU). The ATU
provides outbound windows that map internal addresses to PCIe bus
addresses.
The ATU outbound address space (0x10000000-0x1fffffff) is divided into
six equally-sized regions (~42MB each). When FDMA buffers are allocated,
a free ATU region is claimed and programmed with the DMA target address.
The FDMA engine then uses the region's base address in its descriptors,
and the ATU translates these to the actual DMA addresses on the PCIe bus.
Add the required functions and helpers that combine the DMA allocation
with the ATU region mapping, effectively adding support for PCIe FDMA.
The ATU cannot express a limit finer than its 64KB region granularity,
so pad the mapped allocation to that boundary; otherwise the outbound
window would extend past the memory the host allocated for DMA.
This implementation will also be used by the lan969x, when PCIe FDMA is
added for that platform in the future.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/net/ethernet/microchip/fdma/Makefile | 4 +
drivers/net/ethernet/microchip/fdma/fdma_api.c | 44 ++++++
drivers/net/ethernet/microchip/fdma/fdma_api.h | 14 ++
drivers/net/ethernet/microchip/fdma/fdma_pci.c | 203 +++++++++++++++++++++++++
drivers/net/ethernet/microchip/fdma/fdma_pci.h | 52 +++++++
5 files changed, 317 insertions(+)
diff --git a/drivers/net/ethernet/microchip/fdma/Makefile b/drivers/net/ethernet/microchip/fdma/Makefile
index cc9a736be357..910a10b33fb9 100644
--- a/drivers/net/ethernet/microchip/fdma/Makefile
+++ b/drivers/net/ethernet/microchip/fdma/Makefile
@@ -5,3 +5,7 @@
obj-$(CONFIG_FDMA) += fdma.o
fdma-y += fdma_api.o
+
+ifneq ($(CONFIG_MCHP_LAN966X_PCI),)
+fdma-y += fdma_pci.o
+endif
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.c b/drivers/net/ethernet/microchip/fdma/fdma_api.c
index e78c3590da9e..a3c9e3097c5c 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.c
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.c
@@ -127,6 +127,50 @@ void fdma_free_phys(struct fdma *fdma)
}
EXPORT_SYMBOL_GPL(fdma_free_phys);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+/* Allocate coherent DMA memory and map it in the ATU. */
+int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
+ struct fdma_pci_atu *atu)
+{
+ struct fdma_pci_atu_region *region;
+ int err;
+
+ if (WARN_ON(fdma->atu_region))
+ return -EBUSY;
+
+ /* The ATU cannot express a limit finer than the region granularity, so
+ * the hardware widens the programmed limit to that boundary. Pad the
+ * allocation to match, or the outbound window would extend past the
+ * memory we own.
+ */
+ fdma->size = ALIGN(fdma->size, FDMA_PCI_ATU_REGION_ALIGN);
+
+ err = fdma_alloc_coherent(dev, fdma);
+ if (err)
+ return err;
+
+ region = fdma_pci_atu_region_map(atu, fdma->dma, fdma->size);
+ if (IS_ERR(region)) {
+ fdma_free_coherent(dev, fdma);
+ return PTR_ERR(region);
+ }
+
+ fdma->atu_region = region;
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(fdma_alloc_coherent_and_map);
+
+/* Free coherent DMA memory and unmap the memory in the ATU. */
+void fdma_free_coherent_and_unmap(struct device *dev, struct fdma *fdma)
+{
+ fdma_pci_atu_region_unmap(fdma->atu_region);
+ fdma->atu_region = NULL;
+ fdma_free_coherent(dev, fdma);
+}
+EXPORT_SYMBOL_GPL(fdma_free_coherent_and_unmap);
+#endif
+
/* Get the size of the FDMA memory */
u32 fdma_get_size(struct fdma *fdma)
{
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index dea8e3cc155a..ccc30d506e89 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -7,6 +7,8 @@
#include <linux/etherdevice.h>
#include <linux/types.h>
+#include "fdma_pci.h"
+
/* This provides a common set of functions and data structures for interacting
* with the Frame DMA engine on multiple Microchip switchcores.
*
@@ -109,6 +111,11 @@ struct fdma {
u32 channel_id;
struct fdma_ops ops;
+
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+ /* PCI ATU region for this FDMA instance. */
+ struct fdma_pci_atu_region *atu_region;
+#endif
};
/* Advance the DCB index and wrap if required. */
@@ -233,9 +240,16 @@ int __fdma_dcb_add(struct fdma *fdma, int dcb_idx, u64 info, u64 status,
int fdma_alloc_coherent(struct device *dev, struct fdma *fdma);
int fdma_alloc_phys(struct fdma *fdma);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
+ struct fdma_pci_atu *atu);
+#endif
void fdma_free_coherent(struct device *dev, struct fdma *fdma);
void fdma_free_phys(struct fdma *fdma);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+void fdma_free_coherent_and_unmap(struct device *dev, struct fdma *fdma);
+#endif
u32 fdma_get_size(struct fdma *fdma);
u32 fdma_get_size_contiguous(struct fdma *fdma);
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.c b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
new file mode 100644
index 000000000000..bbfd3c67e03b
--- /dev/null
+++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
@@ -0,0 +1,203 @@
+// SPDX-License-Identifier: GPL-2.0+
+
+#include <linux/align.h>
+#include <linux/errno.h>
+#include <linux/io.h>
+#include <linux/kernel.h>
+#include <linux/types.h>
+
+#include "fdma_pci.h"
+
+/* When the switch operates as a PCIe endpoint, the FDMA engine needs to
+ * DMA to/from host memory. The FDMA writes to addresses within the endpoint's
+ * internal Outbound (OB) address space, and the PCIe ATU translates these to
+ * DMA addresses on the PCIe bus, targeting host memory.
+ *
+ * The ATU supports up to six outbound regions. This implementation divides
+ * the OB address space into six equally sized chunks.
+ *
+ * +-------------+------------+------------+-----+------------+
+ * | Index | Region 0 | Region 1 | ... | Region 5 |
+ * +-------------+------------+------------+-----+------------+
+ * | Base addr | 0x10000000 | 0x12aa0000 | ... | 0x1d520000 |
+ * | Limit addr | 0x12a9ffff | 0x1553ffff | ... | 0x1ffbffff |
+ * | Target addr | host dma | host dma | ... | host dma |
+ * +-------------+------------+------------+-----+------------+
+ *
+ * Base addr is the start address of the region within the OB address space.
+ * Limit addr is each region's own upper bound. The value actually
+ * programmed is set per-mapping as base_addr + mapped size - 1, and is
+ * usually smaller than this.
+ * Target addr is the host DMA address that the base addr translates to.
+ */
+
+#define FDMA_PCI_ATU_OB_START 0x10000000
+#define FDMA_PCI_ATU_OB_END 0x1fffffff
+
+#define FDMA_PCI_ATU_ADDR 0x300000
+#define FDMA_PCI_ATU_IDX_SIZE 0x200
+#define FDMA_PCI_ATU_ENA_REG 0x4
+#define FDMA_PCI_ATU_ENA_BIT BIT(31)
+#define FDMA_PCI_ATU_LWR_BASE_ADDR 0x8
+#define FDMA_PCI_ATU_UPP_BASE_ADDR 0xc
+#define FDMA_PCI_ATU_LIMIT_ADDR 0x10
+#define FDMA_PCI_ATU_LWR_TARGET_ADDR 0x14
+#define FDMA_PCI_ATU_UPP_TARGET_ADDR 0x18
+
+static u32 fdma_pci_atu_region_size(void)
+{
+ return round_down((FDMA_PCI_ATU_OB_END - FDMA_PCI_ATU_OB_START) /
+ FDMA_PCI_ATU_REGION_MAX, FDMA_PCI_ATU_REGION_ALIGN);
+}
+
+static void __iomem *fdma_pci_atu_addr_get(void __iomem *addr, int offset,
+ int idx)
+{
+ return addr + FDMA_PCI_ATU_ADDR + FDMA_PCI_ATU_IDX_SIZE * idx + offset;
+}
+
+static void fdma_pci_atu_region_enable(struct fdma_pci_atu_region *region)
+{
+ writel(FDMA_PCI_ATU_ENA_BIT,
+ fdma_pci_atu_addr_get(region->atu->addr, FDMA_PCI_ATU_ENA_REG,
+ region->idx));
+}
+
+static void fdma_pci_atu_region_disable(struct fdma_pci_atu_region *region)
+{
+ writel(0, fdma_pci_atu_addr_get(region->atu->addr, FDMA_PCI_ATU_ENA_REG,
+ region->idx));
+}
+
+/* Configure the address translation in the ATU. */
+static void
+fdma_pci_atu_configure_translation(struct fdma_pci_atu_region *region)
+{
+ struct fdma_pci_atu *atu = region->atu;
+ int idx = region->idx;
+
+ writel(lower_32_bits(region->base_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_LWR_BASE_ADDR, idx));
+
+ writel(upper_32_bits(region->base_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_UPP_BASE_ADDR, idx));
+
+ /* Upper limit register only needed with REGION_SIZE > 4GB. */
+ writel(region->limit_addr,
+ fdma_pci_atu_addr_get(atu->addr, FDMA_PCI_ATU_LIMIT_ADDR, idx));
+
+ writel(lower_32_bits(region->target_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_LWR_TARGET_ADDR, idx));
+
+ writel(upper_32_bits(region->target_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_UPP_TARGET_ADDR, idx));
+}
+
+/* Find an unused ATU region. */
+static struct fdma_pci_atu_region *
+fdma_pci_atu_region_get_free(struct fdma_pci_atu *atu)
+{
+ struct fdma_pci_atu_region *regions = atu->regions;
+
+ for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
+ if (regions[i].in_use)
+ continue;
+
+ return ®ions[i];
+ }
+
+ return ERR_PTR(-ENOSPC);
+}
+
+/* Unmap an ATU region, clearing its translation and disabling it. */
+void fdma_pci_atu_region_unmap(struct fdma_pci_atu_region *region)
+{
+ if (IS_ERR_OR_NULL(region))
+ return;
+
+ mutex_lock(®ion->atu->lock);
+
+ region->target_addr = 0;
+ region->in_use = false;
+
+ fdma_pci_atu_region_disable(region);
+ fdma_pci_atu_configure_translation(region);
+
+ mutex_unlock(®ion->atu->lock);
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_region_unmap);
+
+/* Map a host DMA address into a free outbound region. */
+struct fdma_pci_atu_region *
+fdma_pci_atu_region_map(struct fdma_pci_atu *atu, u64 target_addr, int size)
+{
+ struct fdma_pci_atu_region *region;
+
+ if (!atu)
+ return ERR_PTR(-EINVAL);
+
+ if (size <= 0)
+ return ERR_PTR(-EINVAL);
+
+ if (size > fdma_pci_atu_region_size())
+ return ERR_PTR(-ERANGE);
+
+ /* The ATU region base is only ever aligned to FDMA_PCI_ATU_REGION_ALIGN;
+ * require the same alignment of the host target address, since the ATU
+ * translates addr - target_addr + base_addr and any misalignment here
+ * would shift every translated address by the same amount.
+ */
+ if (!IS_ALIGNED(target_addr, FDMA_PCI_ATU_REGION_ALIGN))
+ return ERR_PTR(-EINVAL);
+
+ mutex_lock(&atu->lock);
+
+ region = fdma_pci_atu_region_get_free(atu);
+ if (IS_ERR(region)) {
+ mutex_unlock(&atu->lock);
+ return region;
+ }
+
+ region->target_addr = target_addr;
+ region->limit_addr = region->base_addr + size - 1;
+ region->in_use = true;
+
+ fdma_pci_atu_configure_translation(region);
+ fdma_pci_atu_region_enable(region);
+
+ mutex_unlock(&atu->lock);
+
+ return region;
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_region_map);
+
+/* Translate a host DMA address to the corresponding OB address. */
+u64 fdma_pci_atu_translate_addr(struct fdma_pci_atu_region *region, u64 addr)
+{
+ return region->base_addr + (addr - region->target_addr);
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_translate_addr);
+
+/* Initialize ATU, dividing the OB space into equally sized regions. */
+void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr)
+{
+ struct fdma_pci_atu_region *regions = atu->regions;
+ u32 region_size = fdma_pci_atu_region_size();
+
+ atu->addr = addr;
+ mutex_init(&atu->lock);
+
+ for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
+ regions[i].base_addr =
+ FDMA_PCI_ATU_OB_START + (i * region_size);
+ regions[i].idx = i;
+ regions[i].atu = atu;
+
+ fdma_pci_atu_region_disable(®ions[i]);
+ }
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_init);
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.h b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
new file mode 100644
index 000000000000..60aa2d2a9af3
--- /dev/null
+++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
@@ -0,0 +1,52 @@
+/* SPDX-License-Identifier: GPL-2.0+ */
+
+#ifndef _FDMA_PCI_H_
+#define _FDMA_PCI_H_
+
+#include <linux/align.h>
+#include <linux/bits.h>
+#include <linux/mutex.h>
+#include <linux/types.h>
+
+#define FDMA_PCI_ATU_REGION_MAX 6
+
+/* Outbound regions are 64KB granular (datasheet section 3.24.7.4.1), so both
+ * the region base and the mapped size must be aligned to this.
+ */
+#define FDMA_PCI_ATU_REGION_ALIGN BIT(16)
+
+#define FDMA_PCI_DB_ALIGN 128
+#define FDMA_PCI_DB_SIZE(mtu) ALIGN(mtu, FDMA_PCI_DB_ALIGN)
+
+struct fdma_pci_atu;
+
+struct fdma_pci_atu_region {
+ struct fdma_pci_atu *atu;
+ u64 base_addr; /* Base addr of the OB window */
+ u64 limit_addr; /* End addr of the active mapping (base_addr + size - 1) */
+ u64 target_addr; /* Host DMA address this region maps to */
+ int idx;
+ bool in_use;
+};
+
+struct fdma_pci_atu {
+ void __iomem *addr;
+ struct mutex lock; /* Protects region alloc/free and ATU register access */
+ struct fdma_pci_atu_region regions[FDMA_PCI_ATU_REGION_MAX];
+};
+
+/* Initialize ATU, dividing OB space into regions. */
+void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr);
+
+/* Unmap an ATU region, clearing its translation and disabling it. */
+void fdma_pci_atu_region_unmap(struct fdma_pci_atu_region *region);
+
+/* Map a host DMA address into a free ATU region. */
+struct fdma_pci_atu_region *fdma_pci_atu_region_map(struct fdma_pci_atu *atu,
+ u64 target_addr,
+ int size);
+
+/* Translate a host DMA address to the OB address space. */
+u64 fdma_pci_atu_translate_addr(struct fdma_pci_atu_region *region, u64 addr);
+
+#endif
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 04/14] net: lan966x: add FDMA LLP register write helper
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (2 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 03/14] net: microchip: fdma: add PCIe ATU support Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 05/14] net: lan966x: export FDMA helpers for reuse Daniel Machon
` (9 subsequent siblings)
13 siblings, 0 replies; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
The FDMA Link List Pointer (LLP) register points to the first DCB in the
chain and must be written before the channel is activated. This tells
the FDMA engine where to begin DMA transfers.
Move the LLP register writes from the channel start/activate functions
into the allocation functions and introduce a shared
lan966x_fdma_llp_configure() helper. This is needed because the upcoming
PCIe FDMA path writes ATU-translated addresses to the LLP registers
instead of DMA addresses. Keeping the writes in the shared
start/activate path would overwrite these translated addresses.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 30 ++++++++++------------
1 file changed, 14 insertions(+), 16 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 41d4ec7f2f57..b8344fd5e5ad 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -109,6 +109,13 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
return 0;
}
+static void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+ u8 channel_id)
+{
+ lan_wr(lower_32_bits(addr), lan966x, FDMA_DCB_LLP(channel_id));
+ lan_wr(upper_32_bits(addr), lan966x, FDMA_DCB_LLP1(channel_id));
+}
+
static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
@@ -128,6 +135,8 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
fdma_dcbs_init(fdma, FDMA_DCB_INFO_DATAL(fdma->db_size),
FDMA_DCB_STATUS_INTR);
+ lan966x_fdma_llp_configure(lan966x, fdma->dma, fdma->channel_id);
+
return 0;
}
@@ -137,14 +146,6 @@ static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
struct fdma *fdma = &rx->fdma;
u32 mask;
- /* When activating a channel, first is required to write the first DCB
- * address and then to activate it
- */
- lan_wr(lower_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP(fdma->channel_id));
- lan_wr(upper_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP1(fdma->channel_id));
-
lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
@@ -215,6 +216,8 @@ static int lan966x_fdma_tx_alloc(struct lan966x_tx *tx)
fdma_dcbs_init(fdma, 0, 0);
+ lan966x_fdma_llp_configure(lan966x, fdma->dma, fdma->channel_id);
+
return 0;
out:
@@ -236,14 +239,6 @@ static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
struct fdma *fdma = &tx->fdma;
u32 mask;
- /* When activating a channel, first is required to write the first DCB
- * address and then to activate it
- */
- lan_wr(lower_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP(fdma->channel_id));
- lan_wr(upper_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP1(fdma->channel_id));
-
lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
@@ -876,6 +871,9 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
MEM_TYPE_PAGE_POOL, page_pool);
}
+ lan966x_fdma_llp_configure(lan966x, lan966x->rx.fdma.dma,
+ lan966x->rx.fdma.channel_id);
+
lan966x_fdma_rx_start(&lan966x->rx);
lan966x_fdma_wakeup_netdev(lan966x);
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 05/14] net: lan966x: export FDMA helpers for reuse
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (3 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 04/14] net: lan966x: add FDMA LLP register write helper Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 06/14] net: lan966x: use a dedicated device for DMA operations Daniel Machon
` (8 subsequent siblings)
13 siblings, 0 replies; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Make shared FDMA helpers non-static, so they can be reused by the PCIe
FDMA implementation.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 22 +++++++++++-----------
.../net/ethernet/microchip/lan966x/lan966x_main.h | 11 +++++++++++
2 files changed, 22 insertions(+), 11 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index b8344fd5e5ad..1311b79bd11f 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -109,8 +109,8 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
return 0;
}
-static void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
- u8 channel_id)
+void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+ u8 channel_id)
{
lan_wr(lower_32_bits(addr), lan966x, FDMA_DCB_LLP(channel_id));
lan_wr(upper_32_bits(addr), lan966x, FDMA_DCB_LLP1(channel_id));
@@ -140,7 +140,7 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
return 0;
}
-static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
+void lan966x_fdma_rx_start(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
struct fdma *fdma = &rx->fdma;
@@ -171,7 +171,7 @@ static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
lan966x, FDMA_CH_ACTIVATE);
}
-static void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
+void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
struct fdma *fdma = &rx->fdma;
@@ -191,7 +191,7 @@ static void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
lan966x, FDMA_CH_DB_DISCARD);
}
-static void lan966x_fdma_rx_reload(struct lan966x_rx *rx)
+void lan966x_fdma_rx_reload(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
@@ -264,7 +264,7 @@ static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
lan966x, FDMA_CH_ACTIVATE);
}
-static void lan966x_fdma_tx_disable(struct lan966x_tx *tx)
+void lan966x_fdma_tx_disable(struct lan966x_tx *tx)
{
struct lan966x *lan966x = tx->lan966x;
struct fdma *fdma = &tx->fdma;
@@ -296,7 +296,7 @@ static void lan966x_fdma_tx_reload(struct lan966x_tx *tx)
lan966x, FDMA_CH_RELOAD);
}
-static void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x)
+void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x)
{
struct lan966x_port *port;
int i;
@@ -470,7 +470,7 @@ static struct sk_buff *lan966x_fdma_rx_get_frame(struct lan966x_rx *rx,
return NULL;
}
-static int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight)
+int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight)
{
struct lan966x *lan966x = container_of(napi, struct lan966x, napi);
struct lan966x_rx *rx = &lan966x->rx;
@@ -583,7 +583,7 @@ static int lan966x_fdma_get_next_dcb(struct lan966x_tx *tx)
return -1;
}
-static void lan966x_fdma_tx_start(struct lan966x_tx *tx)
+void lan966x_fdma_tx_start(struct lan966x_tx *tx)
{
struct lan966x *lan966x = tx->lan966x;
@@ -801,7 +801,7 @@ static int lan966x_fdma_get_max_mtu(struct lan966x *lan966x)
return max_mtu;
}
-static int lan966x_qsys_sw_status(struct lan966x *lan966x)
+int lan966x_qsys_sw_status(struct lan966x *lan966x)
{
return lan_rd(lan966x, QSYS_SW_STATUS(CPU_PORT));
}
@@ -883,7 +883,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
return err;
}
-static int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
+int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
{
return lan966x_fdma_get_max_mtu(lan966x) +
IFH_LEN_BYTES +
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index eea286c29474..83c361abb789 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -561,6 +561,17 @@ int lan966x_fdma_init(struct lan966x *lan966x);
void lan966x_fdma_deinit(struct lan966x *lan966x);
irqreturn_t lan966x_fdma_irq_handler(int irq, void *args);
int lan966x_fdma_reload_page_pool(struct lan966x *lan966x);
+int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight);
+void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+ u8 channel_id);
+void lan966x_fdma_rx_start(struct lan966x_rx *rx);
+void lan966x_fdma_rx_disable(struct lan966x_rx *rx);
+void lan966x_fdma_rx_reload(struct lan966x_rx *rx);
+void lan966x_fdma_tx_start(struct lan966x_tx *tx);
+void lan966x_fdma_tx_disable(struct lan966x_tx *tx);
+void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
+int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
+int lan966x_qsys_sw_status(struct lan966x *lan966x);
int lan966x_lag_port_join(struct lan966x_port *port,
struct net_device *brport_dev,
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 06/14] net: lan966x: use a dedicated device for DMA operations
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (4 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 05/14] net: lan966x: export FDMA helpers for reuse Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 07/14] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
` (7 subsequent siblings)
13 siblings, 0 replies; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
In preparation for the PCIe FDMA implementation, add lan966x->dma_dev,
resolved once at probe time, and use it for every DMA operation the
FDMA path performs: coherent allocations, streaming mappings and the RX
page pool. Natively, dma_dev is lan966x->dev, so there is no functional
change.
Since dma_dev differs from dev only on the PCIe path, add
lan966x_is_pci() on top of it, for the later PCIe patches to key off.
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 30 +++++++++++-----------
.../net/ethernet/microchip/lan966x/lan966x_main.c | 16 ++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.h | 10 ++++++++
3 files changed, 41 insertions(+), 15 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 1311b79bd11f..15fc59cb4ad7 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -80,7 +80,7 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
.flags = PP_FLAG_DMA_MAP | PP_FLAG_DMA_SYNC_DEV,
.pool_size = rx->fdma.n_dcbs,
.nid = NUMA_NO_NODE,
- .dev = lan966x->dev,
+ .dev = lan966x->dma_dev,
.dma_dir = DMA_FROM_DEVICE,
.offset = XDP_PACKET_HEADROOM,
.max_len = rx->max_mtu -
@@ -126,7 +126,7 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
if (err)
return err;
- err = fdma_alloc_coherent(lan966x->dev, fdma);
+ err = fdma_alloc_coherent(lan966x->dma_dev, fdma);
if (err) {
page_pool_destroy(rx->page_pool);
return err;
@@ -210,7 +210,7 @@ static int lan966x_fdma_tx_alloc(struct lan966x_tx *tx)
if (!tx->dcbs_buf)
return -ENOMEM;
- err = fdma_alloc_coherent(lan966x->dev, fdma);
+ err = fdma_alloc_coherent(lan966x->dma_dev, fdma);
if (err)
goto out;
@@ -230,7 +230,7 @@ static void lan966x_fdma_tx_free(struct lan966x_tx *tx)
struct lan966x *lan966x = tx->lan966x;
kfree(tx->dcbs_buf);
- fdma_free_coherent(lan966x->dev, &tx->fdma);
+ fdma_free_coherent(lan966x->dma_dev, &tx->fdma);
}
static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
@@ -355,7 +355,7 @@ static void lan966x_fdma_tx_clear_buf(struct lan966x *lan966x, int weight)
dcb_buf->used = false;
if (dcb_buf->use_skb) {
- dma_unmap_single(lan966x->dev,
+ dma_unmap_single(lan966x->dma_dev,
dcb_buf->dma_addr,
dcb_buf->len,
DMA_TO_DEVICE);
@@ -364,7 +364,7 @@ static void lan966x_fdma_tx_clear_buf(struct lan966x *lan966x, int weight)
napi_consume_skb(dcb_buf->data.skb, weight);
} else {
if (dcb_buf->xdp_ndo)
- dma_unmap_single(lan966x->dev,
+ dma_unmap_single(lan966x->dma_dev,
dcb_buf->dma_addr,
dcb_buf->len,
DMA_TO_DEVICE);
@@ -400,7 +400,7 @@ static int lan966x_fdma_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
if (unlikely(!page))
return FDMA_ERROR;
- dma_sync_single_for_cpu(lan966x->dev,
+ dma_sync_single_for_cpu(lan966x->dma_dev,
(dma_addr_t)db->dataptr + XDP_PACKET_HEADROOM,
FDMA_DCB_STATUS_BLOCKL(db->status),
DMA_FROM_DEVICE);
@@ -635,11 +635,11 @@ int lan966x_fdma_xmit_xdpf(struct lan966x_port *port, void *ptr, u32 len)
lan966x_ifh_set_bypass(ifh, 1);
lan966x_ifh_set_port(ifh, BIT_ULL(port->chip_port));
- dma_addr = dma_map_single(lan966x->dev,
+ dma_addr = dma_map_single(lan966x->dma_dev,
xdpf->data - IFH_LEN_BYTES,
xdpf->len + IFH_LEN_BYTES,
DMA_TO_DEVICE);
- if (dma_mapping_error(lan966x->dev, dma_addr)) {
+ if (dma_mapping_error(lan966x->dma_dev, dma_addr)) {
ret = NETDEV_TX_OK;
goto out;
}
@@ -655,7 +655,7 @@ int lan966x_fdma_xmit_xdpf(struct lan966x_port *port, void *ptr, u32 len)
lan966x_ifh_set_port(ifh, BIT_ULL(port->chip_port));
dma_addr = page_pool_get_dma_addr(page);
- dma_sync_single_for_device(lan966x->dev,
+ dma_sync_single_for_device(lan966x->dma_dev,
dma_addr + XDP_PACKET_HEADROOM,
len + IFH_LEN_BYTES,
DMA_TO_DEVICE);
@@ -734,9 +734,9 @@ int lan966x_fdma_xmit(struct sk_buff *skb, __be32 *ifh, struct net_device *dev)
memcpy(skb->data, ifh, IFH_LEN_BYTES);
skb_put(skb, 4);
- dma_addr = dma_map_single(lan966x->dev, skb->data, skb->len,
+ dma_addr = dma_map_single(lan966x->dma_dev, skb->data, skb->len,
DMA_TO_DEVICE);
- if (dma_mapping_error(lan966x->dev, dma_addr)) {
+ if (dma_mapping_error(lan966x->dma_dev, dma_addr)) {
dev->stats.tx_dropped++;
err = NETDEV_TX_OK;
goto release;
@@ -842,7 +842,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
page_pool_put_full_page(page_pool,
old_pages[i][j], false);
- fdma_free_coherent(lan966x->dev, &fdma_rx_old);
+ fdma_free_coherent(lan966x->dma_dev, &fdma_rx_old);
page_pool_destroy(page_pool);
@@ -992,7 +992,7 @@ int lan966x_fdma_init(struct lan966x *lan966x)
err = lan966x_fdma_tx_alloc(&lan966x->tx);
if (err) {
- fdma_free_coherent(lan966x->dev, &lan966x->rx.fdma);
+ fdma_free_coherent(lan966x->dma_dev, &lan966x->rx.fdma);
page_pool_destroy(lan966x->rx.page_pool);
return err;
}
@@ -1014,7 +1014,7 @@ void lan966x_fdma_deinit(struct lan966x *lan966x)
napi_disable(&lan966x->napi);
lan966x_fdma_rx_free_pages(&lan966x->rx);
- fdma_free_coherent(lan966x->dev, &lan966x->rx.fdma);
+ fdma_free_coherent(lan966x->dma_dev, &lan966x->rx.fdma);
page_pool_destroy(lan966x->rx.page_pool);
lan966x_fdma_tx_free(&lan966x->tx);
}
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 1179a6e127c5..2741f7c9fa4c 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -7,6 +7,7 @@
#include <linux/ip.h>
#include <linux/of.h>
#include <linux/of_net.h>
+#include <linux/pci.h>
#include <linux/phy/phy.h>
#include <linux/platform_device.h>
#include <linux/reset.h>
@@ -1081,6 +1082,20 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
return 0;
}
+/* When enumerated over PCIe, dev is a platform device with no
+ * iommus/dma-ranges of its own, so DMA must target the PCIe endpoint instead.
+ * The result differs from dev only in that case; lan966x_is_pci() relies on it.
+ */
+static struct device *lan966x_get_dma_dev(struct device *dev)
+{
+ for (struct device *p = dev->parent; p; p = p->parent) {
+ if (dev_is_pci(p))
+ return p;
+ }
+
+ return dev;
+}
+
static int lan966x_probe(struct platform_device *pdev)
{
struct fwnode_handle *ports, *portnp;
@@ -1094,6 +1109,7 @@ static int lan966x_probe(struct platform_device *pdev)
platform_set_drvdata(pdev, lan966x);
lan966x->dev = &pdev->dev;
+ lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
if (!device_get_mac_address(&pdev->dev, mac_addr)) {
ether_addr_copy(lan966x->base_mac, mac_addr);
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 83c361abb789..3f09f8ddf620 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -270,6 +270,9 @@ struct lan966x_skb_cb {
struct lan966x {
struct device *dev;
+ /* Device used for DMA; the PCIe endpoint when enumerated over PCIe. */
+ struct device *dma_dev;
+
u8 num_phys_ports;
struct lan966x_port **ports;
@@ -573,6 +576,13 @@ void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
int lan966x_qsys_sw_status(struct lan966x *lan966x);
+/* dma_dev differs from dev only on the PCIe path. */
+static inline bool lan966x_is_pci(struct lan966x *lan966x)
+{
+ return IS_ENABLED(CONFIG_MCHP_LAN966X_PCI) &&
+ lan966x->dma_dev != lan966x->dev;
+}
+
int lan966x_lag_port_join(struct lan966x_port *port,
struct net_device *brport_dev,
struct net_device *bond,
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 07/14] net: lan966x: add FDMA ops dispatch for PCIe support
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (5 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 06/14] net: lan966x: use a dedicated device for DMA operations Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 08/14] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
` (6 subsequent siblings)
13 siblings, 0 replies; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Introduce lan966x_fdma_ops to support different FDMA implementations
for platform and PCIe. Plumb fdma_init, fdma_deinit, fdma_xmit,
fdma_poll and fdma_resize through the ops table, and select the
implementation at probe time. Only the platform implementation exists
at this point; the PCIe implementation is added in a later patch.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 2 +-
.../net/ethernet/microchip/lan966x/lan966x_main.c | 20 +++++++++++++++-----
.../net/ethernet/microchip/lan966x/lan966x_main.h | 13 +++++++++++++
3 files changed, 29 insertions(+), 6 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 15fc59cb4ad7..2695bc41e52a 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -947,7 +947,7 @@ void lan966x_fdma_netdev_init(struct lan966x *lan966x, struct net_device *dev)
return;
lan966x->fdma_ndev = dev;
- netif_napi_add(dev, &lan966x->napi, lan966x_fdma_napi_poll);
+ netif_napi_add(dev, &lan966x->napi, lan966x->ops->fdma_poll);
napi_enable(&lan966x->napi);
}
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 2741f7c9fa4c..6e6c08bb8eea 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -27,6 +27,14 @@
#define IO_RANGES 2
+static const struct lan966x_fdma_ops lan966x_fdma_ops = {
+ .fdma_init = &lan966x_fdma_init,
+ .fdma_deinit = &lan966x_fdma_deinit,
+ .fdma_xmit = &lan966x_fdma_xmit,
+ .fdma_poll = &lan966x_fdma_napi_poll,
+ .fdma_resize = &lan966x_fdma_change_mtu,
+};
+
static const struct of_device_id lan966x_match[] = {
{ .compatible = "microchip,lan966x-switch" },
{ }
@@ -392,7 +400,7 @@ static netdev_tx_t lan966x_port_xmit(struct sk_buff *skb,
spin_lock(&lan966x->tx_lock);
if (port->lan966x->fdma)
- err = lan966x_fdma_xmit(skb, ifh, dev);
+ err = lan966x->ops->fdma_xmit(skb, ifh, dev);
else
err = lan966x_port_ifh_xmit(skb, ifh, dev);
spin_unlock(&lan966x->tx_lock);
@@ -414,7 +422,7 @@ static int lan966x_port_change_mtu(struct net_device *dev, int new_mtu)
if (!lan966x->fdma)
return 0;
- err = lan966x_fdma_change_mtu(lan966x);
+ err = lan966x->ops->fdma_resize(lan966x);
if (err) {
lan_wr(DEV_MAC_MAXLEN_CFG_MAX_LEN_SET(LAN966X_HW_MTU(old_mtu)),
lan966x, DEV_MAC_MAXLEN_CFG(port->chip_port));
@@ -1111,6 +1119,8 @@ static int lan966x_probe(struct platform_device *pdev)
lan966x->dev = &pdev->dev;
lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
+ lan966x->ops = &lan966x_fdma_ops;
+
if (!device_get_mac_address(&pdev->dev, mac_addr)) {
ether_addr_copy(lan966x->base_mac, mac_addr);
} else {
@@ -1250,7 +1260,7 @@ static int lan966x_probe(struct platform_device *pdev)
if (err)
goto cleanup_fdb;
- err = lan966x_fdma_init(lan966x);
+ err = lan966x->ops->fdma_init(lan966x);
if (err)
goto cleanup_ptp;
@@ -1263,7 +1273,7 @@ static int lan966x_probe(struct platform_device *pdev)
return 0;
cleanup_fdma:
- lan966x_fdma_deinit(lan966x);
+ lan966x->ops->fdma_deinit(lan966x);
cleanup_ptp:
lan966x_ptp_deinit(lan966x);
@@ -1291,7 +1301,7 @@ static void lan966x_remove(struct platform_device *pdev)
lan966x_taprio_deinit(lan966x);
lan966x_vcap_deinit(lan966x);
- lan966x_fdma_deinit(lan966x);
+ lan966x->ops->fdma_deinit(lan966x);
lan966x_cleanup_ports(lan966x);
cancel_delayed_work_sync(&lan966x->stats_work);
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 3f09f8ddf620..aab5e53ed059 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -193,6 +193,17 @@ enum vcap_is1_port_sel_rt {
VCAP_IS1_PS_RT_FOLLOW_OTHER = 7,
};
+struct lan966x;
+
+struct lan966x_fdma_ops {
+ int (*fdma_init)(struct lan966x *lan966x);
+ void (*fdma_deinit)(struct lan966x *lan966x);
+ int (*fdma_xmit)(struct sk_buff *skb, __be32 *ifh,
+ struct net_device *dev);
+ int (*fdma_poll)(struct napi_struct *napi, int weight);
+ int (*fdma_resize)(struct lan966x *lan966x);
+};
+
struct lan966x_port;
struct lan966x_rx {
@@ -273,6 +284,8 @@ struct lan966x {
/* Device used for DMA; the PCIe endpoint when enumerated over PCIe. */
struct device *dma_dev;
+ const struct lan966x_fdma_ops *ops;
+
u8 num_phys_ports;
struct lan966x_port **ports;
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 08/14] net: lan966x: clear FDMA interrupt stickies after switch reset
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (6 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 07/14] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 09/14] net: lan966x: add shutdown callback to stop FDMA on reboot Daniel Machon
` (5 subsequent siblings)
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When in PCI mode, the GCB soft reset issued by the reset controller
can latch spurious bits in the FDMA error stickies. The latched bits
sit in FDMA_INTR_ERR until the FDMA IRQ is requested later in probe,
at which point the handler fires immediately and WARNs.
Clear FDMA_ERRORS, FDMA_INTR_ERR and FDMA_INTR_DB right after the
switch reset so the FDMA comes out clean and the IRQ handler does not
see ghost errors on probe.
The clear runs on both the PCI and platform paths. On the platform
path it has no effect — there are no spurious stickies to clear — but
keeping it unconditional avoids a PCI-specific code path here.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/net/ethernet/microchip/lan966x/lan966x_main.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 6e6c08bb8eea..11094a381ec2 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -1067,6 +1067,15 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
reset_control_reset(switch_reset);
+ /* When in PCI mode, the GCB soft reset issued by the reset
+ * controller can latch spurious bits in the FDMA error stickies.
+ * Clear them before request_irq hooks up the FDMA IRQ line,
+ * otherwise the handler fires immediately on probe.
+ */
+ lan_wr(lan_rd(lan966x, FDMA_ERRORS), lan966x, FDMA_ERRORS);
+ lan_wr(lan_rd(lan966x, FDMA_INTR_ERR), lan966x, FDMA_INTR_ERR);
+ lan_wr(lan_rd(lan966x, FDMA_INTR_DB), lan966x, FDMA_INTR_DB);
+
/* Don't reinitialize the switch core, if it is already initialized. In
* case it is initialized twice, some pointers inside the queue system
* in HW will get corrupted and then after a while the queue system gets
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 09/14] net: lan966x: add shutdown callback to stop FDMA on reboot
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (7 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 08/14] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 10/14] net: lan966x: add PCIe FDMA support Daniel Machon
` (4 subsequent siblings)
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When lan966x is used as a PCIe endpoint, the FDMA engine runs on the
card and survives a host reboot. Without a shutdown callback, channels
stay active and interrupt sources stay armed across the reset, causing
the shared PCIe INTx to assert before the driver has re-probed.
Add a shutdown callback, shared by the platform and PCI paths, that
disables NAPI and stops the netdev TX queues, then disables the RX and
TX channels, and finally masks the interrupt sources that share the PCIe
INTx: FDMA (FDMA_INTR_ENA and FDMA_INTR_DB_ENA) and the analyzer
(ANA_ANAINTR), which lan966x_init() arms unconditionally.
Masking is only required on this path: on reboot the kernel calls
device_shutdown(), so shutdown() is the only callback that runs. On
unbind, fdma_deinit() disables both channels and waits for them to go
idle before lan966x_cleanup_ports() releases the FDMA interrupt handler,
so the engine cannot assert the shared line once the handler is gone.
FDMA_INTR_ENA persists on the card across a warm reboot, so also
restore the full enable in lan966x_fdma_rx_start() to re-arm interrupts
after a previous shutdown(). rx_start() runs after both the RX and TX
rings are allocated, so the same single-site re-arm works for both the
platform and PCIe backends.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 19 ++++++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.c | 26 ++++++++++++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.h | 1 +
.../net/ethernet/microchip/lan966x/lan966x_regs.h | 15 +++++++++++++
4 files changed, 61 insertions(+)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 2695bc41e52a..5793a83268fc 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -146,6 +146,10 @@ void lan966x_fdma_rx_start(struct lan966x_rx *rx)
struct fdma *fdma = &rx->fdma;
u32 mask;
+ lan_wr(FDMA_INTR_ENA_INTR_PORT_ENA_SET(GENMASK(1, 0)) |
+ FDMA_INTR_ENA_INTR_CH_ENA_SET(GENMASK(7, 0)),
+ lan966x, FDMA_INTR_ENA);
+
lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
@@ -325,6 +329,21 @@ static void lan966x_fdma_stop_netdev(struct lan966x *lan966x)
}
}
+/* Drain in-flight xmit callers and stop all TX queues on every port. */
+void lan966x_fdma_tx_disable_netdev(struct lan966x *lan966x)
+{
+ struct lan966x_port *port;
+ int i;
+
+ for (i = 0; i < lan966x->num_phys_ports; ++i) {
+ port = lan966x->ports[i];
+ if (!port)
+ continue;
+
+ netif_tx_disable(port->dev);
+ }
+}
+
static void lan966x_fdma_tx_clear_buf(struct lan966x *lan966x, int weight)
{
struct lan966x_tx *tx = &lan966x->tx;
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 11094a381ec2..d6ce1e3e373f 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -1324,9 +1324,35 @@ static void lan966x_remove(struct platform_device *pdev)
debugfs_remove_recursive(lan966x->debugfs_root);
}
+static void lan966x_shutdown(struct platform_device *pdev)
+{
+ struct lan966x *lan966x = platform_get_drvdata(pdev);
+
+ if (!lan966x->fdma)
+ return;
+
+ /* The reload paths disable this NAPI under rtnl; serialize with them. */
+ rtnl_lock();
+
+ if (lan966x->fdma_ndev)
+ napi_disable(&lan966x->napi);
+
+ lan966x_fdma_tx_disable_netdev(lan966x);
+
+ lan966x_fdma_rx_disable(&lan966x->rx);
+ lan966x_fdma_tx_disable(&lan966x->tx);
+
+ lan_wr(0, lan966x, FDMA_INTR_ENA);
+ lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
+ lan_wr(0, lan966x, ANA_ANAINTR);
+
+ rtnl_unlock();
+}
+
static struct platform_driver lan966x_driver = {
.probe = lan966x_probe,
.remove = lan966x_remove,
+ .shutdown = lan966x_shutdown,
.driver = {
.name = "lan966x-switch",
.of_match_table = lan966x_match,
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index aab5e53ed059..b7e3ce4f0355 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -586,6 +586,7 @@ void lan966x_fdma_rx_reload(struct lan966x_rx *rx);
void lan966x_fdma_tx_start(struct lan966x_tx *tx);
void lan966x_fdma_tx_disable(struct lan966x_tx *tx);
void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
+void lan966x_fdma_tx_disable_netdev(struct lan966x *lan966x);
int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
int lan966x_qsys_sw_status(struct lan966x *lan966x);
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
index 4b553927d2e0..aba0d36ae6b5 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
@@ -1039,6 +1039,21 @@ enum lan966x_target {
/* FDMA:FDMA:FDMA_INTR_ERR */
#define FDMA_INTR_ERR __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 400, 0, 1, 4)
+/* FDMA:FDMA:FDMA_INTR_ENA */
+#define FDMA_INTR_ENA __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 404, 0, 1, 4)
+
+#define FDMA_INTR_ENA_INTR_PORT_ENA GENMASK(9, 8)
+#define FDMA_INTR_ENA_INTR_PORT_ENA_SET(x)\
+ FIELD_PREP(FDMA_INTR_ENA_INTR_PORT_ENA, x)
+#define FDMA_INTR_ENA_INTR_PORT_ENA_GET(x)\
+ FIELD_GET(FDMA_INTR_ENA_INTR_PORT_ENA, x)
+
+#define FDMA_INTR_ENA_INTR_CH_ENA GENMASK(7, 0)
+#define FDMA_INTR_ENA_INTR_CH_ENA_SET(x)\
+ FIELD_PREP(FDMA_INTR_ENA_INTR_CH_ENA, x)
+#define FDMA_INTR_ENA_INTR_CH_ENA_GET(x)\
+ FIELD_GET(FDMA_INTR_ENA_INTR_CH_ENA, x)
+
/* FDMA:FDMA:FDMA_ERRORS */
#define FDMA_ERRORS __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 412, 0, 1, 4)
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 10/14] net: lan966x: add PCIe FDMA support
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (8 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 09/14] net: lan966x: add shutdown callback to stop FDMA on reboot Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 11/14] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
` (3 subsequent siblings)
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add PCIe FDMA support for lan966x. The PCIe FDMA path uses contiguous
DMA buffers mapped through the endpoint's ATU, with memcpy-based frame
transfer instead of per-page DMA mappings.
With PCIe FDMA, throughput increases from ~33 Mbps (register-based I/O)
to ~620 Mbps on an Intel x86 host with a lan966x PCIe card.
XDP is not supported on this path yet, so do not advertise xdp_features
and reject program attach for PCIe instances.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/net/ethernet/microchip/lan966x/Makefile | 4 +
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 404 +++++++++++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.c | 11 +-
.../net/ethernet/microchip/lan966x/lan966x_main.h | 7 +
.../net/ethernet/microchip/lan966x/lan966x_regs.h | 10 +
.../net/ethernet/microchip/lan966x/lan966x_xdp.c | 6 +
6 files changed, 439 insertions(+), 3 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/Makefile b/drivers/net/ethernet/microchip/lan966x/Makefile
index 4cdbe263502c..02db4d4d9826 100644
--- a/drivers/net/ethernet/microchip/lan966x/Makefile
+++ b/drivers/net/ethernet/microchip/lan966x/Makefile
@@ -18,6 +18,10 @@ lan966x-switch-objs := lan966x_main.o lan966x_phylink.o lan966x_port.o \
lan966x-switch-$(CONFIG_LAN966X_DCB) += lan966x_dcb.o
lan966x-switch-$(CONFIG_DEBUG_FS) += lan966x_vcap_debugfs.o
+ifneq ($(CONFIG_MCHP_LAN966X_PCI),)
+lan966x-switch-y += lan966x_fdma_pci.o
+endif
+
# Provide include files
ccflags-y += -I$(srctree)/drivers/net/ethernet/microchip/vcap
ccflags-y += -I$(srctree)/drivers/net/ethernet/microchip/fdma
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
new file mode 100644
index 000000000000..f1f3c789d3a6
--- /dev/null
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -0,0 +1,404 @@
+// SPDX-License-Identifier: GPL-2.0+
+
+#include "fdma_api.h"
+#include "lan966x_main.h"
+
+static int lan966x_fdma_pci_dataptr_cb(struct fdma *fdma, int dcb, int db,
+ u64 *dataptr)
+{
+ u64 addr;
+
+ addr = fdma_dataptr_dma_addr_contiguous(fdma, dcb, db);
+
+ *dataptr = fdma_pci_atu_translate_addr(fdma->atu_region, addr);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_nextptr_cb(struct fdma *fdma, int dcb, u64 *nextptr)
+{
+ u64 addr;
+
+ fdma_nextptr_cb(fdma, dcb, &addr);
+
+ *nextptr = fdma_pci_atu_translate_addr(fdma->atu_region, addr);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_rx_alloc(struct lan966x_rx *rx)
+{
+ struct lan966x *lan966x = rx->lan966x;
+ struct fdma *fdma = &rx->fdma;
+ int err;
+
+ err = fdma_alloc_coherent_and_map(lan966x->dma_dev, fdma,
+ &lan966x->atu);
+ if (err)
+ return err;
+
+ err = fdma_dcbs_init(fdma,
+ FDMA_DCB_INFO_DATAL(fdma->db_size - XDP_PACKET_HEADROOM),
+ FDMA_DCB_STATUS_INTR);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, fdma);
+ return err;
+ }
+
+ lan966x_fdma_llp_configure(lan966x,
+ fdma->atu_region->base_addr,
+ fdma->channel_id);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_tx_alloc(struct lan966x_tx *tx)
+{
+ struct lan966x *lan966x = tx->lan966x;
+ struct fdma *fdma = &tx->fdma;
+ int err;
+
+ err = fdma_alloc_coherent_and_map(lan966x->dma_dev, fdma,
+ &lan966x->atu);
+ if (err)
+ return err;
+
+ err = fdma_dcbs_init(fdma,
+ FDMA_DCB_INFO_DATAL(fdma->db_size),
+ FDMA_DCB_STATUS_DONE);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, fdma);
+ return err;
+ }
+
+ lan966x_fdma_llp_configure(lan966x,
+ fdma->atu_region->base_addr,
+ fdma->channel_id);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_get_next_dcb(struct fdma *fdma)
+{
+ struct fdma_db *db;
+
+ for (int i = 0; i < fdma->n_dcbs; i++) {
+ db = fdma_db_get(fdma, i, 0);
+
+ if (!fdma_db_is_done(db))
+ continue;
+ if (fdma_is_last(fdma, &fdma->dcbs[i]))
+ continue;
+
+ return i;
+ }
+
+ return -ENOSPC;
+}
+
+/* TX slot layout (sizes in bytes):
+ *
+ * +---------------------+-----+---------+-----+
+ * | XDP_PACKET_HEADROOM | IFH | payload | FCS |
+ * | 256 | 28 | len | 4 |
+ * +---------------------+-----+---------+-----+
+ * |<---------------- db_size ----------------->|
+ *
+ * Return true if the frame plus required overhead fits.
+ */
+static bool lan966x_fdma_pci_tx_size_fits(struct fdma *fdma, u32 len)
+{
+ return XDP_PACKET_HEADROOM + IFH_LEN_BYTES + len + ETH_FCS_LEN <=
+ fdma->db_size;
+}
+
+/* Return true if blockl is a valid RX frame size. */
+static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
+{
+ return blockl >= IFH_LEN_BYTES + ETH_HLEN + ETH_FCS_LEN &&
+ blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
+}
+
+static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
+{
+ struct lan966x *lan966x = rx->lan966x;
+ struct fdma *fdma = &rx->fdma;
+ struct lan966x_port *port;
+ struct fdma_db *db;
+ void *virt_addr;
+ u32 blockl;
+
+ /* virt_addr points to the IFH. */
+ virt_addr = fdma_dataptr_virt_addr_contiguous(fdma,
+ fdma->dcb_index,
+ fdma->db_index);
+
+ lan966x_ifh_get_src_port(virt_addr, src_port);
+
+ if (*src_port >= lan966x->num_phys_ports)
+ return FDMA_ERROR;
+
+ port = lan966x->ports[*src_port];
+ if (!port)
+ return FDMA_ERROR;
+
+ db = fdma_db_next_get(fdma);
+
+ /* BLOCKL is a 16-bit HW-populated field; reject obviously-bad
+ * values before they feed memcpy/XDP sizes.
+ */
+ blockl = FDMA_DCB_STATUS_BLOCKL(db->status);
+ if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
+ return FDMA_ERROR;
+
+ return FDMA_PASS;
+}
+
+static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
+ u64 src_port)
+{
+ struct lan966x *lan966x = rx->lan966x;
+ struct fdma *fdma = &rx->fdma;
+ struct sk_buff *skb;
+ struct fdma_db *db;
+ u32 data_len;
+
+ /* Get the received frame and create an SKB for it. */
+ db = fdma_db_next_get(fdma);
+ data_len = FDMA_DCB_STATUS_BLOCKL(db->status);
+
+ skb = napi_alloc_skb(&lan966x->napi, data_len);
+ if (unlikely(!skb))
+ return NULL;
+
+ memcpy(skb->data,
+ fdma_dataptr_virt_addr_contiguous(fdma,
+ fdma->dcb_index,
+ fdma->db_index),
+ data_len);
+
+ skb_put(skb, data_len);
+
+ skb->dev = lan966x->ports[src_port]->dev;
+ skb_pull(skb, IFH_LEN_BYTES);
+
+ skb_trim(skb, skb->len - ETH_FCS_LEN);
+
+ skb->protocol = eth_type_trans(skb, skb->dev);
+
+ if (lan966x->bridge_mask & BIT(src_port)) {
+ skb->offload_fwd_mark = 1;
+
+ skb_reset_network_header(skb);
+ if (!lan966x_hw_offload(lan966x, src_port, skb))
+ skb->offload_fwd_mark = 0;
+ }
+
+ skb->dev->stats.rx_bytes += skb->len;
+ skb->dev->stats.rx_packets++;
+
+ return skb;
+}
+
+static int lan966x_fdma_pci_xmit(struct sk_buff *skb, __be32 *ifh,
+ struct net_device *dev)
+{
+ struct lan966x_port *port = netdev_priv(dev);
+ struct lan966x *lan966x = port->lan966x;
+ struct lan966x_tx *tx = &lan966x->tx;
+ struct fdma *fdma = &tx->fdma;
+ int next_to_use;
+ void *virt_addr;
+
+ next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
+
+ if (next_to_use < 0) {
+ netif_stop_queue(dev);
+ return NETDEV_TX_BUSY;
+ }
+
+ if (skb_put_padto(skb, ETH_ZLEN)) {
+ dev->stats.tx_dropped++;
+ return NETDEV_TX_OK;
+ }
+
+ if (!lan966x_fdma_pci_tx_size_fits(fdma, skb->len)) {
+ dev_kfree_skb_any(skb);
+ dev->stats.tx_dropped++;
+ return NETDEV_TX_OK;
+ }
+
+ skb_tx_timestamp(skb);
+
+ /* virt_addr points to the IFH. */
+ virt_addr = fdma_dataptr_virt_addr_contiguous(fdma, next_to_use, 0);
+ memcpy(virt_addr, ifh, IFH_LEN_BYTES);
+ memcpy(virt_addr + IFH_LEN_BYTES, skb->data, skb->len);
+
+ /* Order frame write before DCB status write below. */
+ dma_wmb();
+
+ fdma_dcb_add(fdma,
+ next_to_use,
+ 0,
+ FDMA_DCB_STATUS_INTR |
+ FDMA_DCB_STATUS_SOF |
+ FDMA_DCB_STATUS_EOF |
+ FDMA_DCB_STATUS_BLOCKO(0) |
+ FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + skb->len + ETH_FCS_LEN));
+
+ /* Start the transmission. */
+ lan966x_fdma_tx_start(tx);
+
+ dev->stats.tx_bytes += skb->len;
+ dev->stats.tx_packets++;
+
+ /* Safe to free: PTP is not supported on the PCIe path yet,
+ * so lan966x->ptp is always 0 here.
+ */
+ dev_consume_skb_any(skb);
+
+ return NETDEV_TX_OK;
+}
+
+static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
+{
+ struct lan966x *lan966x = container_of(napi, struct lan966x, napi);
+ struct lan966x_rx *rx = &lan966x->rx;
+ struct fdma *fdma = &rx->fdma;
+ int dcb_reload, old_dcb;
+ struct sk_buff *skb;
+ int counter = 0;
+ u64 src_port;
+
+ /* Wake any stopped TX queues if a TX DCB is available. */
+ spin_lock(&lan966x->tx_lock);
+ if (lan966x_fdma_pci_get_next_dcb(&lan966x->tx.fdma) >= 0)
+ lan966x_fdma_wakeup_netdev(lan966x);
+ spin_unlock(&lan966x->tx_lock);
+
+ dcb_reload = fdma->dcb_index;
+
+ /* Get all received skbs. */
+ while (counter < weight) {
+ if (!fdma_has_frames(fdma))
+ break;
+ /* Order DONE read before DCB/frame reads below. */
+ dma_rmb();
+ counter++;
+ switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
+ case FDMA_PASS:
+ break;
+ case FDMA_ERROR:
+ /* No rx_dropped increment here because src_port is
+ * invalid.
+ */
+ fdma_dcb_advance(fdma);
+ continue;
+ }
+ skb = lan966x_fdma_pci_rx_get_frame(rx, src_port);
+ fdma_dcb_advance(fdma);
+ if (!skb) {
+ lan966x->ports[src_port]->dev->stats.rx_dropped++;
+ continue;
+ }
+
+ napi_gro_receive(&lan966x->napi, skb);
+ }
+ while (dcb_reload != fdma->dcb_index) {
+ old_dcb = dcb_reload;
+ dcb_reload++;
+ dcb_reload &= fdma->n_dcbs - 1;
+
+ fdma_dcb_add(fdma,
+ old_dcb,
+ FDMA_DCB_INFO_DATAL(fdma->db_size - XDP_PACKET_HEADROOM),
+ FDMA_DCB_STATUS_INTR);
+
+ lan966x_fdma_rx_reload(rx);
+ }
+
+ if (counter < weight && napi_complete_done(napi, counter))
+ lan_wr(0xff, lan966x, FDMA_INTR_DB_ENA);
+
+ return counter;
+}
+
+static int lan966x_fdma_pci_init(struct lan966x *lan966x)
+{
+ struct fdma *rx_fdma = &lan966x->rx.fdma;
+ struct fdma *tx_fdma = &lan966x->tx.fdma;
+ int err;
+
+ if (!lan966x->fdma)
+ return 0;
+
+ lan_wr(FDMA_CTRL_NRESET_SET(0), lan966x, FDMA_CTRL);
+ lan_wr(FDMA_CTRL_NRESET_SET(1), lan966x, FDMA_CTRL);
+
+ fdma_pci_atu_init(&lan966x->atu, lan966x->regs[TARGET_PCIE_DBI]);
+
+ lan966x->rx.lan966x = lan966x;
+ lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
+ rx_fdma->channel_id = FDMA_XTR_CHANNEL;
+ rx_fdma->n_dcbs = FDMA_DCB_MAX;
+ rx_fdma->n_dbs = FDMA_RX_DCB_MAX_DBS;
+ rx_fdma->priv = lan966x;
+ rx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ rx_fdma->size = fdma_get_size_contiguous(rx_fdma);
+ rx_fdma->ops.nextptr_cb = &lan966x_fdma_pci_nextptr_cb;
+ rx_fdma->ops.dataptr_cb = &lan966x_fdma_pci_dataptr_cb;
+
+ lan966x->tx.lan966x = lan966x;
+ tx_fdma->channel_id = FDMA_INJ_CHANNEL;
+ tx_fdma->n_dcbs = FDMA_DCB_MAX;
+ tx_fdma->n_dbs = FDMA_TX_DCB_MAX_DBS;
+ tx_fdma->priv = lan966x;
+ tx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ tx_fdma->size = fdma_get_size_contiguous(tx_fdma);
+ tx_fdma->ops.nextptr_cb = &lan966x_fdma_pci_nextptr_cb;
+ tx_fdma->ops.dataptr_cb = &lan966x_fdma_pci_dataptr_cb;
+
+ err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
+ if (err)
+ return err;
+
+ err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, rx_fdma);
+ return err;
+ }
+
+ lan966x_fdma_rx_start(&lan966x->rx);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
+{
+ return -EOPNOTSUPP;
+}
+
+static void lan966x_fdma_pci_deinit(struct lan966x *lan966x)
+{
+ if (!lan966x->fdma)
+ return;
+
+ if (lan966x->fdma_ndev)
+ napi_disable(&lan966x->napi);
+
+ lan966x_fdma_tx_disable_netdev(lan966x);
+ lan966x_fdma_rx_disable(&lan966x->rx);
+ lan966x_fdma_tx_disable(&lan966x->tx);
+
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->rx.fdma);
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->tx.fdma);
+}
+
+const struct lan966x_fdma_ops lan966x_fdma_pci_ops = {
+ .fdma_init = &lan966x_fdma_pci_init,
+ .fdma_deinit = &lan966x_fdma_pci_deinit,
+ .fdma_xmit = &lan966x_fdma_pci_xmit,
+ .fdma_poll = &lan966x_fdma_pci_napi_poll,
+ .fdma_resize = &lan966x_fdma_pci_resize,
+};
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index d6ce1e3e373f..e089c08b9c37 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -50,6 +50,7 @@ struct lan966x_main_io_resource {
static const struct lan966x_main_io_resource lan966x_main_iomap[] = {
{ TARGET_CPU, 0xc0000, 0 }, /* 0xe00c0000 */
{ TARGET_FDMA, 0xc0400, 0 }, /* 0xe00c0400 */
+ { TARGET_PCIE_DBI, 0x400000, 0 }, /* 0xe0400000 */
{ TARGET_ORG, 0, 1 }, /* 0xe2000000 */
{ TARGET_GCB, 0x4000, 1 }, /* 0xe2004000 */
{ TARGET_QS, 0x8000, 1 }, /* 0xe2008000 */
@@ -873,7 +874,8 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
port->phylink = phylink;
- if (lan966x->fdma)
+ /* XDP is not supported on the PCIe FDMA path. */
+ if (lan966x->fdma && !lan966x_is_pci(lan966x))
dev->xdp_features = NETDEV_XDP_ACT_BASIC |
NETDEV_XDP_ACT_REDIRECT |
NETDEV_XDP_ACT_NDO_XMIT;
@@ -1128,7 +1130,8 @@ static int lan966x_probe(struct platform_device *pdev)
lan966x->dev = &pdev->dev;
lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
- lan966x->ops = &lan966x_fdma_ops;
+ lan966x->ops = lan966x_is_pci(lan966x) ? &lan966x_fdma_pci_ops :
+ &lan966x_fdma_ops;
if (!device_get_mac_address(&pdev->dev, mac_addr)) {
ether_addr_copy(lan966x->base_mac, mac_addr);
@@ -1187,7 +1190,9 @@ static int lan966x_probe(struct platform_device *pdev)
if (err)
return dev_err_probe(&pdev->dev, err, "Unable to use ptp irq");
- lan966x->ptp = 1;
+ /* PTP is not supported on the PCIe path yet. */
+ if (!lan966x_is_pci(lan966x))
+ lan966x->ptp = 1;
}
lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index b7e3ce4f0355..48c7527845fb 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -17,6 +17,7 @@
#include <net/xdp.h>
#include <fdma_api.h>
+#include <fdma_pci.h>
#include <vcap_api.h>
#include <vcap_api_client.h>
@@ -291,6 +292,10 @@ struct lan966x {
void __iomem *regs[NUM_TARGETS];
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+ struct fdma_pci_atu atu;
+#endif
+
int shared_queue_sz;
u8 base_mac[ETH_ALEN];
@@ -590,6 +595,8 @@ void lan966x_fdma_tx_disable_netdev(struct lan966x *lan966x);
int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
int lan966x_qsys_sw_status(struct lan966x *lan966x);
+extern const struct lan966x_fdma_ops lan966x_fdma_pci_ops;
+
/* dma_dev differs from dev only on the PCIe path. */
static inline bool lan966x_is_pci(struct lan966x *lan966x)
{
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
index aba0d36ae6b5..4778ea217673 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
@@ -20,6 +20,7 @@ enum lan966x_target {
TARGET_FDMA = 21,
TARGET_GCB = 27,
TARGET_ORG = 36,
+ TARGET_PCIE_DBI = 40,
TARGET_PTP = 41,
TARGET_QS = 42,
TARGET_QSYS = 46,
@@ -1009,6 +1010,15 @@ enum lan966x_target {
#define FDMA_CH_CFG_CH_MEM_GET(x)\
FIELD_GET(FDMA_CH_CFG_CH_MEM, x)
+/* FDMA:FDMA:FDMA_CTRL */
+#define FDMA_CTRL __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 424, 0, 1, 4)
+
+#define FDMA_CTRL_NRESET BIT(0)
+#define FDMA_CTRL_NRESET_SET(x)\
+ FIELD_PREP(FDMA_CTRL_NRESET, x)
+#define FDMA_CTRL_NRESET_GET(x)\
+ FIELD_GET(FDMA_CTRL_NRESET, x)
+
/* FDMA:FDMA:FDMA_PORT_CTRL */
#define FDMA_PORT_CTRL(r) __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 376, r, 2, 4)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
index 9ee61db8690b..e63634887f64 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
@@ -20,6 +20,12 @@ static int lan966x_xdp_setup(struct net_device *dev, struct netdev_bpf *xdp)
return -EOPNOTSUPP;
}
+ if (lan966x_is_pci(lan966x)) {
+ NL_SET_ERR_MSG_MOD(xdp->extack,
+ "XDP is not supported on the PCIe FDMA path");
+ return -EOPNOTSUPP;
+ }
+
old_xdp = lan966x_xdp_present(lan966x);
old_prog = xchg(&port->xdp_prog, xdp->prog);
new_xdp = lan966x_xdp_present(lan966x);
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 11/14] net: lan966x: add PCIe FDMA MTU change support
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (9 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 10/14] net: lan966x: add PCIe FDMA support Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 12/14] net: lan966x: add PCIe FDMA XDP support Daniel Machon
` (2 subsequent siblings)
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add MTU change support for the PCIe FDMA path. When the MTU changes,
the contiguous ATU-mapped RX and TX buffers are reallocated with the
new size. On allocation failure, the existing buffers are reused
after being reset.
Cap the PCIe DCB ring at 256 (FDMA_PCI_DCB_MAX) to keep the entire
contiguous allocation under MAX_PAGE_ORDER at jumbo MTU, which 512
DCBs would overflow.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 168 ++++++++++++++++++++-
1 file changed, 165 insertions(+), 3 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
index f1f3c789d3a6..6cabbb8b47f2 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -1,8 +1,15 @@
// SPDX-License-Identifier: GPL-2.0+
+#include <linux/mmzone.h>
+
#include "fdma_api.h"
#include "lan966x_main.h"
+/* Ring must fit in one MAX_PAGE_ORDER DMA block; 512 DCBs overflows
+ * at jumbo MTU.
+ */
+#define FDMA_PCI_DCB_MAX 256
+
static int lan966x_fdma_pci_dataptr_cb(struct fdma *fdma, int dcb, int db,
u64 *dataptr)
{
@@ -341,7 +348,7 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
lan966x->rx.lan966x = lan966x;
lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
rx_fdma->channel_id = FDMA_XTR_CHANNEL;
- rx_fdma->n_dcbs = FDMA_DCB_MAX;
+ rx_fdma->n_dcbs = FDMA_PCI_DCB_MAX;
rx_fdma->n_dbs = FDMA_RX_DCB_MAX_DBS;
rx_fdma->priv = lan966x;
rx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
@@ -351,7 +358,7 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
lan966x->tx.lan966x = lan966x;
tx_fdma->channel_id = FDMA_INJ_CHANNEL;
- tx_fdma->n_dcbs = FDMA_DCB_MAX;
+ tx_fdma->n_dcbs = FDMA_PCI_DCB_MAX;
tx_fdma->n_dbs = FDMA_TX_DCB_MAX_DBS;
tx_fdma->priv = lan966x;
tx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
@@ -374,9 +381,164 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
return 0;
}
+/* Reset existing rx and tx buffers. */
+static void lan966x_fdma_pci_reset_mem(struct lan966x *lan966x)
+{
+ struct lan966x_rx *rx = &lan966x->rx;
+ struct lan966x_tx *tx = &lan966x->tx;
+
+ memset(rx->fdma.dcbs, 0, rx->fdma.size);
+ memset(tx->fdma.dcbs, 0, tx->fdma.size);
+
+ fdma_dcbs_init(&rx->fdma,
+ FDMA_DCB_INFO_DATAL(rx->fdma.db_size - XDP_PACKET_HEADROOM),
+ FDMA_DCB_STATUS_INTR);
+
+ fdma_dcbs_init(&tx->fdma,
+ FDMA_DCB_INFO_DATAL(tx->fdma.db_size),
+ FDMA_DCB_STATUS_DONE);
+
+ lan966x_fdma_llp_configure(lan966x,
+ tx->fdma.atu_region->base_addr,
+ tx->fdma.channel_id);
+ lan966x_fdma_llp_configure(lan966x,
+ rx->fdma.atu_region->base_addr,
+ rx->fdma.channel_id);
+}
+
+/* Wake all TX queues on every port (undoes lan966x_fdma_tx_disable_netdev). */
+static void lan966x_fdma_pci_wakeup_netdev(struct lan966x *lan966x)
+{
+ for (int i = 0; i < lan966x->num_phys_ports; ++i) {
+ struct lan966x_port *port = lan966x->ports[i];
+
+ if (port)
+ netif_tx_wake_all_queues(port->dev);
+ }
+}
+
+static int lan966x_fdma_pci_reload(struct lan966x *lan966x, int new_mtu)
+{
+ struct fdma tx_fdma_old = lan966x->tx.fdma;
+ struct fdma rx_fdma_old = lan966x->rx.fdma;
+ u32 old_mtu = lan966x->rx.max_mtu;
+ int err;
+
+ napi_disable(&lan966x->napi);
+ lan966x_fdma_tx_disable_netdev(lan966x);
+ lan966x_fdma_rx_disable(&lan966x->rx);
+ lan966x_fdma_tx_disable(&lan966x->tx);
+
+ lan966x->rx.max_mtu = new_mtu;
+
+ /* Must be NULL'ed in order to realloc them. */
+ lan966x->rx.fdma.atu_region = NULL;
+ lan966x->tx.fdma.atu_region = NULL;
+
+ lan966x->tx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ lan966x->tx.fdma.size = fdma_get_size_contiguous(&lan966x->tx.fdma);
+ lan966x->rx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ lan966x->rx.fdma.size = fdma_get_size_contiguous(&lan966x->rx.fdma);
+
+ err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
+ if (err)
+ goto restore;
+
+ err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev,
+ &lan966x->rx.fdma);
+ goto restore;
+ }
+
+ /* Free and unmap old memory. */
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &tx_fdma_old);
+
+ napi_enable(&lan966x->napi);
+ lan966x_fdma_rx_start(&lan966x->rx);
+ lan966x_fdma_pci_wakeup_netdev(lan966x);
+
+ return err;
+restore:
+
+ /* No new buffers are allocated at this point. Use the old buffers,
+ * but reset them before starting the FDMA again.
+ */
+
+ memcpy(&lan966x->tx.fdma, &tx_fdma_old, sizeof(struct fdma));
+ memcpy(&lan966x->rx.fdma, &rx_fdma_old, sizeof(struct fdma));
+
+ lan966x->rx.max_mtu = old_mtu;
+
+ lan966x_fdma_pci_reset_mem(lan966x);
+
+ napi_enable(&lan966x->napi);
+ lan966x_fdma_rx_start(&lan966x->rx);
+ lan966x_fdma_pci_wakeup_netdev(lan966x);
+
+ return err;
+}
+
+static int __lan966x_fdma_pci_reload(struct lan966x *lan966x, int max_mtu)
+{
+ int err;
+ u32 val;
+
+ /* Disable the CPU port. */
+ lan_rmw(QSYS_SW_PORT_MODE_PORT_ENA_SET(0),
+ QSYS_SW_PORT_MODE_PORT_ENA,
+ lan966x, QSYS_SW_PORT_MODE(CPU_PORT));
+
+ /* Flush the CPU queues. */
+ readx_poll_timeout(lan966x_qsys_sw_status,
+ lan966x,
+ val,
+ !(QSYS_SW_STATUS_EQ_AVAIL_GET(val)),
+ READL_SLEEP_US, READL_TIMEOUT_US);
+
+ /* Add a sleep in case there are frames between the queues and the CPU
+ * port
+ */
+ usleep_range(USEC_PER_MSEC, 2 * USEC_PER_MSEC);
+
+ err = lan966x_fdma_pci_reload(lan966x, max_mtu);
+
+ /* Enable back the CPU port. */
+ lan_rmw(QSYS_SW_PORT_MODE_PORT_ENA_SET(1),
+ QSYS_SW_PORT_MODE_PORT_ENA,
+ lan966x, QSYS_SW_PORT_MODE(CPU_PORT));
+
+ return err;
+}
+
static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
{
- return -EOPNOTSUPP;
+ struct fdma rx_fdma;
+ int max_mtu;
+
+ max_mtu = lan966x_fdma_get_max_frame(lan966x);
+ if (max_mtu == lan966x->rx.max_mtu)
+ return 0;
+
+ /* rx and tx have n_dbs == 1, so both rings need the same contiguous
+ * dma_alloc_coherent() block, which can't exceed MAX_PAGE_ORDER. The
+ * allocation is padded to the ATU region granularity, so test the
+ * padded size.
+ */
+ rx_fdma = lan966x->rx.fdma;
+ rx_fdma.db_size = FDMA_PCI_DB_SIZE(max_mtu);
+ if (ALIGN(fdma_get_size_contiguous(&rx_fdma),
+ FDMA_PCI_ATU_REGION_ALIGN) > (PAGE_SIZE << MAX_PAGE_ORDER))
+ return -ERANGE;
+
+ /* db_size is also handed to the FDMA in the 16-bit DCB DATAL field,
+ * where a larger value would be silently truncated.
+ */
+ if (rx_fdma.db_size > GENMASK(15, 0))
+ return -ERANGE;
+
+ return __lan966x_fdma_pci_reload(lan966x, max_mtu);
}
static void lan966x_fdma_pci_deinit(struct lan966x *lan966x)
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 12/14] net: lan966x: add PCIe FDMA XDP support
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (10 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 11/14] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 13/14] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add XDP support for the PCIe FDMA path. The implementation operates on
contiguous ATU-mapped buffers with memcpy-based XDP_TX, unlike the
platform path which uses page_pool.
XDP sees the frame with IFH and FCS stripped. These are removed in
lan966x_fdma_pci_rx_check_frame() before the BPF program runs, because
after the program returns the driver cannot tell whether the tail
region was modified. The skb_pull/skb_trim previously done in
lan966x_fdma_pci_rx_get_frame() are removed for the same reason; the
frame pointer and length are pre-computed by rx_check_frame() and
passed through rx_get_frame() and lan966x_xdp_pci_run() to the caller.
lan966x_fdma_pci_xmit_xdpf() handles XDP_TX: it rebuilds a fresh IFH
in the TX slot, copies the post-XDP frame after it, and lets HW insert
a new FCS.
lan966x_xdp_setup() is extended so the PCIe path skips the page_pool
reload that the platform path needs.
Only XDP_ACT_BASIC is supported.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 166 ++++++++++++++++++---
.../net/ethernet/microchip/lan966x/lan966x_main.c | 12 +-
.../net/ethernet/microchip/lan966x/lan966x_xdp.c | 12 +-
3 files changed, 159 insertions(+), 31 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
index 6cabbb8b47f2..02c021617a3e 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: GPL-2.0+
+#include <linux/bpf_trace.h>
#include <linux/mmzone.h>
#include "fdma_api.h"
@@ -126,7 +127,123 @@ static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
}
-static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
+static int lan966x_fdma_pci_xmit_xdpf(struct lan966x_port *port,
+ void *ptr, u32 len)
+{
+ struct lan966x *lan966x = port->lan966x;
+ struct lan966x_tx *tx = &lan966x->tx;
+ struct fdma *fdma = &tx->fdma;
+ int next_to_use, ret = 0;
+ void *virt_addr;
+
+ spin_lock(&lan966x->tx_lock);
+
+ next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
+
+ if (next_to_use < 0) {
+ netif_stop_queue(port->dev);
+ port->dev->stats.tx_dropped++;
+ ret = NETDEV_TX_BUSY;
+ goto out;
+ }
+
+ /* Only the upper bound is enforced: XDP owns the frame contents and
+ * length, so a program that shrinks below ETH_ZLEN gets what it asked
+ * for.
+ */
+ if (!lan966x_fdma_pci_tx_size_fits(fdma, len)) {
+ port->dev->stats.tx_dropped++;
+ ret = -EINVAL;
+ goto out;
+ }
+
+ /* virt_addr points to the IFH. */
+ virt_addr = fdma_dataptr_virt_addr_contiguous(fdma, next_to_use, 0);
+
+ /* Construct a fresh IFH. */
+ memset(virt_addr, 0, IFH_LEN_BYTES);
+ lan966x_ifh_set_bypass(virt_addr, 1);
+ lan966x_ifh_set_port(virt_addr, BIT_ULL(port->chip_port));
+
+ /* Copy the (post-XDP) frame after the IFH. */
+ memcpy(virt_addr + IFH_LEN_BYTES, ptr, len);
+
+ /* Order frame write before DCB status write below. */
+ dma_wmb();
+
+ /* Reserve ETH_FCS_LEN for the HW-inserted FCS (len is FCS-stripped). */
+ fdma_dcb_add(fdma,
+ next_to_use,
+ 0,
+ FDMA_DCB_STATUS_INTR |
+ FDMA_DCB_STATUS_SOF |
+ FDMA_DCB_STATUS_EOF |
+ FDMA_DCB_STATUS_BLOCKO(0) |
+ FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + len + ETH_FCS_LEN));
+
+ /* Start the transmission. */
+ lan966x_fdma_tx_start(tx);
+
+ port->dev->stats.tx_bytes += len;
+ port->dev->stats.tx_packets++;
+
+out:
+ spin_unlock(&lan966x->tx_lock);
+
+ return ret;
+}
+
+static int lan966x_xdp_pci_run(struct lan966x_port *port, void *data,
+ u32 data_len, void **xdp_data, u32 *xdp_len)
+{
+ /* Read once so the NULL check and bpf_prog_run_xdp() see the same
+ * pointer.
+ */
+ struct bpf_prog *xdp_prog = READ_ONCE(port->xdp_prog);
+ struct lan966x *lan966x = port->lan966x;
+ struct fdma *fdma = &lan966x->rx.fdma;
+ struct xdp_buff xdp;
+ u32 act;
+
+ if (!xdp_prog)
+ return FDMA_PASS;
+
+ xdp_init_buff(&xdp, fdma->db_size, &port->xdp_rxq);
+
+ /* hard_start is set to slot start (virt_addr is XDP_PACKET_HEADROOM
+ * into the slot). Headroom includes the IFH; BPF may grow into it
+ * via adjust_head. IFH is rebuilt on XDP_TX and unread on XDP_PASS.
+ */
+ xdp_prepare_buff(&xdp,
+ data - XDP_PACKET_HEADROOM,
+ XDP_PACKET_HEADROOM + IFH_LEN_BYTES,
+ data_len,
+ false);
+
+ act = bpf_prog_run_xdp(xdp_prog, &xdp);
+
+ *xdp_data = xdp.data;
+ *xdp_len = xdp.data_end - xdp.data;
+
+ switch (act) {
+ case XDP_PASS:
+ return FDMA_PASS;
+ case XDP_TX:
+ return lan966x_fdma_pci_xmit_xdpf(port, *xdp_data, *xdp_len) ?
+ FDMA_DROP : FDMA_TX;
+ default:
+ bpf_warn_invalid_xdp_action(port->dev, xdp_prog, act);
+ fallthrough;
+ case XDP_ABORTED:
+ trace_xdp_exception(port->dev, xdp_prog, act);
+ fallthrough;
+ case XDP_DROP:
+ return FDMA_DROP;
+ }
+}
+
+static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port,
+ void **data, u32 *data_len)
{
struct lan966x *lan966x = rx->lan966x;
struct fdma *fdma = &rx->fdma;
@@ -158,38 +275,33 @@ static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
return FDMA_ERROR;
- return FDMA_PASS;
+ /* Present the Ethernet frame (no IFH, no FCS). HW re-inserts the
+ * FCS on TX; see lan966x_fdma_pci_xmit_xdpf(). May be overridden
+ * by XDP. The FCS strip is unconditional because NETIF_F_RXFCS
+ * is not advertised in hw_features.
+ */
+ *data = virt_addr + IFH_LEN_BYTES;
+ *data_len = blockl - IFH_LEN_BYTES - ETH_FCS_LEN;
+
+ return lan966x_xdp_pci_run(port, virt_addr, *data_len, data, data_len);
}
static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
- u64 src_port)
+ u64 src_port, void *data,
+ u32 data_len)
{
struct lan966x *lan966x = rx->lan966x;
- struct fdma *fdma = &rx->fdma;
struct sk_buff *skb;
- struct fdma_db *db;
- u32 data_len;
-
- /* Get the received frame and create an SKB for it. */
- db = fdma_db_next_get(fdma);
- data_len = FDMA_DCB_STATUS_BLOCKL(db->status);
skb = napi_alloc_skb(&lan966x->napi, data_len);
if (unlikely(!skb))
return NULL;
- memcpy(skb->data,
- fdma_dataptr_virt_addr_contiguous(fdma,
- fdma->dcb_index,
- fdma->db_index),
- data_len);
+ memcpy(skb->data, data, data_len);
skb_put(skb, data_len);
skb->dev = lan966x->ports[src_port]->dev;
- skb_pull(skb, IFH_LEN_BYTES);
-
- skb_trim(skb, skb->len - ETH_FCS_LEN);
skb->protocol = eth_type_trans(skb, skb->dev);
@@ -277,6 +389,8 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
struct sk_buff *skb;
int counter = 0;
u64 src_port;
+ u32 data_len;
+ void *data;
/* Wake any stopped TX queues if a TX DCB is available. */
spin_lock(&lan966x->tx_lock);
@@ -293,7 +407,10 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
/* Order DONE read before DCB/frame reads below. */
dma_rmb();
counter++;
- switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
+ switch (lan966x_fdma_pci_rx_check_frame(rx,
+ &src_port,
+ &data,
+ &data_len)) {
case FDMA_PASS:
break;
case FDMA_ERROR:
@@ -302,8 +419,17 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
*/
fdma_dcb_advance(fdma);
continue;
+ case FDMA_TX:
+ fdma_dcb_advance(fdma);
+ continue;
+ case FDMA_DROP:
+ fdma_dcb_advance(fdma);
+ continue;
}
- skb = lan966x_fdma_pci_rx_get_frame(rx, src_port);
+ skb = lan966x_fdma_pci_rx_get_frame(rx,
+ src_port,
+ data,
+ data_len);
fdma_dcb_advance(fdma);
if (!skb) {
lan966x->ports[src_port]->dev->stats.rx_dropped++;
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index e089c08b9c37..9ada497767c6 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -874,11 +874,13 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
port->phylink = phylink;
- /* XDP is not supported on the PCIe FDMA path. */
- if (lan966x->fdma && !lan966x_is_pci(lan966x))
- dev->xdp_features = NETDEV_XDP_ACT_BASIC |
- NETDEV_XDP_ACT_REDIRECT |
- NETDEV_XDP_ACT_NDO_XMIT;
+ if (lan966x->fdma) {
+ dev->xdp_features = NETDEV_XDP_ACT_BASIC;
+
+ if (!lan966x_is_pci(lan966x))
+ dev->xdp_features |= NETDEV_XDP_ACT_REDIRECT |
+ NETDEV_XDP_ACT_NDO_XMIT;
+ }
err = register_netdev(dev);
if (err) {
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
index e63634887f64..b98426afc785 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
@@ -20,16 +20,16 @@ static int lan966x_xdp_setup(struct net_device *dev, struct netdev_bpf *xdp)
return -EOPNOTSUPP;
}
- if (lan966x_is_pci(lan966x)) {
- NL_SET_ERR_MSG_MOD(xdp->extack,
- "XDP is not supported on the PCIe FDMA path");
- return -EOPNOTSUPP;
- }
-
old_xdp = lan966x_xdp_present(lan966x);
old_prog = xchg(&port->xdp_prog, xdp->prog);
new_xdp = lan966x_xdp_present(lan966x);
+ /* PCIe FDMA uses contiguous buffers, so no page_pool reload
+ * is needed.
+ */
+ if (lan966x_is_pci(lan966x))
+ goto out;
+
if (old_xdp == new_xdp)
goto out;
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 13/14] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (11 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 12/14] net: lan966x: add PCIe FDMA XDP support Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
The ATU outbound windows used by the FDMA engine are programmed through
registers at offset 0x400000+, which falls outside the current cpu reg
mapping. Extend the cpu reg size from 0x100000 (1MB) to 0x800000 (8MB)
to cover the full PCIE DBI and iATU register space.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/misc/lan966x_pci.dtso | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
index 7b196b0a0eb6..7bb726550caf 100644
--- a/drivers/misc/lan966x_pci.dtso
+++ b/drivers/misc/lan966x_pci.dtso
@@ -135,7 +135,7 @@ lan966x_phy1: ethernet-lan966x_phy@2 {
switch: switch@e0000000 {
compatible = "microchip,lan966x-switch";
- reg = <0xe0000000 0x0100000>,
+ reg = <0xe0000000 0x0800000>,
<0xe2000000 0x0800000>;
reg-names = "cpu", "gcb";
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
` (12 preceding siblings ...)
2026-09-09 13:00 ` [PATCH net-next v6 13/14] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
@ 2026-09-09 13:00 ` Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
13 siblings, 1 reply; 23+ messages in thread
From: Daniel Machon @ 2026-09-09 13:00 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add the fdma interrupt (OIC interrupt 14) to the lan966x PCI device
tree overlay, enabling FDMA-based frame injection/extraction when
the switch is connected over PCIe.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/misc/lan966x_pci.dtso | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
index 7bb726550caf..5bb12dbc0843 100644
--- a/drivers/misc/lan966x_pci.dtso
+++ b/drivers/misc/lan966x_pci.dtso
@@ -141,8 +141,9 @@ switch: switch@e0000000 {
interrupt-parent = <&oic>;
interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
+ <14 IRQ_TYPE_LEVEL_HIGH>,
<9 IRQ_TYPE_LEVEL_HIGH>;
- interrupt-names = "xtr", "ana";
+ interrupt-names = "xtr", "fdma", "ana";
resets = <&reset 0>;
reset-names = "switch";
--
2.34.1
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 03/14] net: microchip: fdma: add PCIe ATU support
2026-09-09 13:00 ` [PATCH net-next v6 03/14] net: microchip: fdma: add PCIe ATU support Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 6 potential
issue(s) to consider:
Critical: 0 · High: 0 · Medium: 3 · Low: 3
- [Medium] fdma_pci_atu_region_map() (EXPORT_SYMBOL_GPL) is the primitive
that programs the ATU limit register (region->limit_addr =…
- [Medium] fdma_pci_atu_region_enable() writes the region enable bit
(BIT(31) at offset 0x4) and returns immediately, and…
- [Medium] fdma_pci_atu_configure_translation() programs lower/upper base
(0x8/0xc), limit (0x10), lower/upper target (0x14/0x18) and the enable…
- [Low] The commit message says the patch "effectively add[s] support for
PCIe FDMA", but at this commit nothing in the tree calls…
- [Low] fdma_pci_atu_init() is the designated initializer of struct
fdma_pci_atu (it even programs hardware, disabling all six outbound…
- [Low] fdma_pci_atu_translate_addr() reads region->base_addr and
region->target_addr with no lock, no READ_ONCE() and no in_use…
--
net: microchip: fdma: add PCIe ATU support
This adds a new fdma_pci.c/fdma_pci.h library that carves the PCIe endpoint
outbound address space (0x10000000-0x1fffffff) into six equally sized ATU
regions, and adds fdma_alloc_coherent_and_map()/fdma_free_coherent_and_unmap()
helpers that pair a coherent DMA allocation with an ATU region mapping.
A few questions below.
> Add the required functions and helpers that combine the DMA allocation
> with the ATU region mapping, effectively adding support for PCIe FDMA.
[Severity: Low]
Is "effectively adding support for PCIe FDMA" the right wording for this
patch on its own?
At this commit nothing in the tree calls fdma_pci_atu_init(),
fdma_pci_atu_region_map(), fdma_pci_atu_translate_addr(),
fdma_alloc_coherent_and_map() or fdma_free_coherent_and_unmap(), and
FDMA_PCI_DB_ALIGN/FDMA_PCI_DB_SIZE() have no user either.
The wiring appears later in the series, in "net: lan966x: add PCIe FDMA
support" and its follow-ups (lan966x_fdma_pci.c calls fdma_pci_atu_init(),
fdma_alloc_coherent_and_map(), fdma_free_coherent_and_unmap(),
fdma_pci_atu_translate_addr() and FDMA_PCI_DB_SIZE()).
Would "add the infrastructure for PCIe FDMA" describe this patch more
accurately?
> diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.c b/drivers/net/ethernet/microchip/fdma/fdma_api.c
> index e78c3590da9e..a3c9e3097c5c 100644
> --- a/drivers/net/ethernet/microchip/fdma/fdma_api.c
> +++ b/drivers/net/ethernet/microchip/fdma/fdma_api.c
> @@ -127,6 +127,50 @@ void fdma_free_phys(struct fdma *fdma)
> }
> EXPORT_SYMBOL_GPL(fdma_free_phys);
>
> +#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
> +/* Allocate coherent DMA memory and map it in the ATU. */
> +int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
> + struct fdma_pci_atu *atu)
> +{
> + struct fdma_pci_atu_region *region;
> + int err;
> +
> + if (WARN_ON(fdma->atu_region))
> + return -EBUSY;
> +
> + /* The ATU cannot express a limit finer than the region granularity, so
> + * the hardware widens the programmed limit to that boundary. Pad the
> + * allocation to match, or the outbound window would extend past the
> + * memory we own.
> + */
> + fdma->size = ALIGN(fdma->size, FDMA_PCI_ATU_REGION_ALIGN);
> +
> + err = fdma_alloc_coherent(dev, fdma);
> + if (err)
> + return err;
> +
> + region = fdma_pci_atu_region_map(atu, fdma->dma, fdma->size);
> + if (IS_ERR(region)) {
> + fdma_free_coherent(dev, fdma);
> + return PTR_ERR(region);
> + }
[ ... ]
> diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.c b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
> new file mode 100644
> index 000000000000..bbfd3c67e03b
> --- /dev/null
> +++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
> @@ -0,0 +1,203 @@
[ ... ]
> +#define FDMA_PCI_ATU_ADDR 0x300000
> +#define FDMA_PCI_ATU_IDX_SIZE 0x200
> +#define FDMA_PCI_ATU_ENA_REG 0x4
> +#define FDMA_PCI_ATU_ENA_BIT BIT(31)
> +#define FDMA_PCI_ATU_LWR_BASE_ADDR 0x8
> +#define FDMA_PCI_ATU_UPP_BASE_ADDR 0xc
> +#define FDMA_PCI_ATU_LIMIT_ADDR 0x10
> +#define FDMA_PCI_ATU_LWR_TARGET_ADDR 0x14
> +#define FDMA_PCI_ATU_UPP_TARGET_ADDR 0x18
[ ... ]
> +static void fdma_pci_atu_region_enable(struct fdma_pci_atu_region *region)
> +{
> + writel(FDMA_PCI_ATU_ENA_BIT,
> + fdma_pci_atu_addr_get(region->atu->addr, FDMA_PCI_ATU_ENA_REG,
> + region->idx));
> +}
[Severity: Medium]
Does the enable bit take effect by the time this writel() retires?
The register block looks like the DesignWare unrolled iATU: the 0x300000
base matches DEFAULT_DBI_ATU_OFFSET, the 0x200 stride matches
PCIE_ATU_UNROLL_BASE(dir, index) = index << 9, and 0x0/0x4/0x8/0xc/0x10/
0x14/0x18 match PCIE_ATU_UNR_REGION_CTRL1/CTRL2/LOWER_BASE/UPPER_BASE/
LOWER_LIMIT/LOWER_TARGET/UPPER_TARGET with PCIE_ATU_ENABLE = BIT(31) in
CTRL2.
For that IP, dw_pcie_prog_outbound_atu() in
drivers/pci/controller/dwc/pcie-designware.c polls until the enable bit
reads back:
for (retries = 0; retries < LINK_WAIT_MAX_IATU_RETRIES; retries++) {
val = dw_pcie_readl_atu(pci, dir, index, PCIE_ATU_REGION_CTRL2);
if (val & PCIE_ATU_ENABLE)
return 0;
mdelay(LINK_WAIT_IATU);
}
dev_err(pci->dev, "Outbound iATU is not being enabled\n");
return -ETIMEDOUT;
with the comment "Make sure ATU enable takes effect before any subsequent
config and I/O accesses".
Here fdma_pci_atu_region_map() returns the region as usable right after
this write, and the caller programs FDMA descriptors with
region->base_addr. Should the enable be read back (and a failure to arm
reported to the caller) instead of being fire-and-forget?
> +/* Configure the address translation in the ATU. */
> +static void
> +fdma_pci_atu_configure_translation(struct fdma_pci_atu_region *region)
> +{
> + struct fdma_pci_atu *atu = region->atu;
> + int idx = region->idx;
> +
> + writel(lower_32_bits(region->base_addr),
> + fdma_pci_atu_addr_get(atu->addr,
> + FDMA_PCI_ATU_LWR_BASE_ADDR, idx));
> +
> + writel(upper_32_bits(region->base_addr),
> + fdma_pci_atu_addr_get(atu->addr,
> + FDMA_PCI_ATU_UPP_BASE_ADDR, idx));
> +
> + /* Upper limit register only needed with REGION_SIZE > 4GB. */
> + writel(region->limit_addr,
> + fdma_pci_atu_addr_get(atu->addr, FDMA_PCI_ATU_LIMIT_ADDR, idx));
[Severity: Medium]
Should region control 1 (offset 0x0) be programmed here too?
Everything except CTRL1 is written: base, limit, target and the enable bit
in CTRL2. On the DWC unrolled iATU, CTRL1 holds the outbound TLP TYPE (MEM
= 0), TD, PCIE_ATU_FUNC_NUM(pf) and PCIE_ATU_INCREASE_REGION_SIZE
(BIT(13)). dw_pcie_prog_outbound_atu() always writes CTRL1 before enabling
CTRL2:
val = type | PCIE_ATU_FUNC_NUM(func_no);
...
dw_pcie_writel_atu(pci, dir, index, PCIE_ATU_REGION_CTRL1, val);
dw_pcie_writel_atu(pci, dir, index, PCIE_ATU_REGION_CTRL2,
PCIE_ATU_ENABLE);
There is no define for offset 0x0 in this file, and fdma_pci_atu_init()
only clears the enable bit, so each window is armed with whatever CTRL1
value reset, the endpoint bootloader, or a previous OS instance left in
place. With a residual TYPE the FDMA writes go out as the wrong TLP type,
with a stale FUNC_NUM they carry the wrong requester, and with a stale
INCREASE_REGION_SIZE the window end is taken from a register this code
never writes.
Related, is the comment "Upper limit register only needed with REGION_SIZE
> 4GB" accurate here? That mode is selected by CTRL1 BIT(13), which is
neither set nor cleared, and the upper limit register is never written.
[ ... ]
> +/* Find an unused ATU region. */
> +static struct fdma_pci_atu_region *
> +fdma_pci_atu_region_get_free(struct fdma_pci_atu *atu)
> +{
> + struct fdma_pci_atu_region *regions = atu->regions;
> +
> + for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
> + if (regions[i].in_use)
> + continue;
> +
> + return ®ions[i];
> + }
> +
> + return ERR_PTR(-ENOSPC);
> +}
[ ... ]
> +/* Map a host DMA address into a free outbound region. */
> +struct fdma_pci_atu_region *
> +fdma_pci_atu_region_map(struct fdma_pci_atu *atu, u64 target_addr, int size)
> +{
> + struct fdma_pci_atu_region *region;
> +
> + if (!atu)
> + return ERR_PTR(-EINVAL);
> +
> + if (size <= 0)
> + return ERR_PTR(-EINVAL);
> +
> + if (size > fdma_pci_atu_region_size())
> + return ERR_PTR(-ERANGE);
> +
> + /* The ATU region base is only ever aligned to FDMA_PCI_ATU_REGION_ALIGN;
> + * require the same alignment of the host target address, since the ATU
> + * translates addr - target_addr + base_addr and any misalignment here
> + * would shift every translated address by the same amount.
> + */
> + if (!IS_ALIGNED(target_addr, FDMA_PCI_ATU_REGION_ALIGN))
> + return ERR_PTR(-EINVAL);
[Severity: Medium]
Should this validation block also reject a size that is not a multiple of
FDMA_PCI_ATU_REGION_ALIGN?
fdma_pci.h states the rule for both halves:
/* Outbound regions are 64KB granular (datasheet section 3.24.7.4.1), so both
* the region base and the mapped size must be aligned to this.
*/
#define FDMA_PCI_ATU_REGION_ALIGN BIT(16)
and the commit message names the consequence of not honouring it: "the
outbound window would extend past the memory the host allocated for DMA".
The size given here becomes the programmed limit:
region->limit_addr = region->base_addr + size - 1;
so with, say, a PAGE_SIZE-aligned 0x21000 from fdma_get_size() the hardware
widens the limit to the next 64KB boundary and the enabled outbound window
covers up to 64KB-1 of host memory beyond the buffer, which the FDMA engine
can then write into.
The 64KB padding is only done in the convenience wrapper
fdma_alloc_coherent_and_map(), not in this exported primitive that
programs the register, and the commit message advertises the primitive for
reuse by lan969x. Would an IS_ALIGNED(size, FDMA_PCI_ATU_REGION_ALIGN)
check (or rounding here) be better placed alongside the target_addr check?
Also, the header comment for fdma_pci_atu_region_map() is just "Map a host
DMA address into a free ATU region", with no mention that a target_addr
that is not 64KB aligned is rejected with -EINVAL. Could that requirement
be documented at the prototype?
> +
> + mutex_lock(&atu->lock);
> +
> + region = fdma_pci_atu_region_get_free(atu);
> + if (IS_ERR(region)) {
> + mutex_unlock(&atu->lock);
> + return region;
> + }
> +
> + region->target_addr = target_addr;
> + region->limit_addr = region->base_addr + size - 1;
> + region->in_use = true;
> +
> + fdma_pci_atu_configure_translation(region);
> + fdma_pci_atu_region_enable(region);
> +
> + mutex_unlock(&atu->lock);
> +
> + return region;
> +}
> +EXPORT_SYMBOL_GPL(fdma_pci_atu_region_map);
> +
> +/* Translate a host DMA address to the corresponding OB address. */
> +u64 fdma_pci_atu_translate_addr(struct fdma_pci_atu_region *region, u64 addr)
> +{
> + return region->base_addr + (addr - region->target_addr);
> +}
> +EXPORT_SYMBOL_GPL(fdma_pci_atu_translate_addr);
[Severity: Low]
What serializes these reads against the writers?
fdma_pci_atu_region_map() and fdma_pci_atu_region_unmap() mutate
target_addr/in_use and the hardware window under the sleeping atu->lock,
while this helper reads base_addr/target_addr with no lock, no READ_ONCE()
and no in_use check. Its intended callers are descriptor callbacks
(lan966x_fdma_pci_dataptr_cb()/nextptr_cb(), reached from NAPI poll and
from lan966x_fdma_pci_xmit()), which cannot take a mutex.
Looking at the end of the series, the writers are
lan966x_fdma_pci_reload() and lan966x_fdma_pci_deinit(), both of which
first do napi_disable() and lan966x_fdma_tx_disable_netdev() (netif_tx_disable()
drains in-flight ndo_start_xmit callers), and both are under rtnl, so there
does not seem to be a reachable race today.
Given that, would it help to document at the prototype that the caller must
quiesce DMA and the descriptor paths before unmapping a region, since this
is an exported helper?
> +
> +/* Initialize ATU, dividing the OB space into equally sized regions. */
> +void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr)
> +{
> + struct fdma_pci_atu_region *regions = atu->regions;
> + u32 region_size = fdma_pci_atu_region_size();
> +
> + atu->addr = addr;
> + mutex_init(&atu->lock);
> +
> + for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
> + regions[i].base_addr =
> + FDMA_PCI_ATU_OB_START + (i * region_size);
> + regions[i].idx = i;
> + regions[i].atu = atu;
> +
> + fdma_pci_atu_region_disable(®ions[i]);
> + }
> +}
> +EXPORT_SYMBOL_GPL(fdma_pci_atu_init);
[Severity: Low]
Should this also initialize in_use, target_addr and limit_addr?
This is the designated initializer for struct fdma_pci_atu and it even
touches hardware (disabling all six windows), but it only sets base_addr,
idx and atu. The allocation state that fdma_pci_atu_region_get_free()
consumes is in_use, which is never assigned here, so the helper silently
depends on the embedding structure having been zeroed. If in_use held
garbage, every region would look busy and mapping would fail with -ENOSPC.
The only in-tree caller later in the series passes &lan966x->atu, and
lan966x comes from devm_kzalloc() in lan966x_probe(), so it works today.
Would it be worth clearing the per-region software state here (or noting
the "caller must zero the struct" precondition at the prototype)?
There is also no counterpart to this function: nothing resets the software
state, unmaps still-mapped regions, or calls mutex_destroy(). Is a
fdma_pci_atu_deinit() wanted for symmetry?
> diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.h b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
> new file mode 100644
> index 000000000000..60aa2d2a9af3
> --- /dev/null
> +++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
> @@ -0,0 +1,52 @@
[ ... ]
> +/* Outbound regions are 64KB granular (datasheet section 3.24.7.4.1), so both
> + * the region base and the mapped size must be aligned to this.
> + */
> +#define FDMA_PCI_ATU_REGION_ALIGN BIT(16)
[ ... ]
> +/* Initialize ATU, dividing OB space into regions. */
> +void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr);
> +
> +/* Unmap an ATU region, clearing its translation and disabling it. */
> +void fdma_pci_atu_region_unmap(struct fdma_pci_atu_region *region);
> +
> +/* Map a host DMA address into a free ATU region. */
> +struct fdma_pci_atu_region *fdma_pci_atu_region_map(struct fdma_pci_atu *atu,
> + u64 target_addr,
> + int size);
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 08/14] net: lan966x: clear FDMA interrupt stickies after switch reset
2026-09-09 13:00 ` [PATCH net-next v6 08/14] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:
Critical: 0 · High: 0 · Medium: 0 · Low: 1
- [Low] Commit-message/comment vs. reachability at this revision: the
added FDMA sticky clear in lan966x_reset_switch()…
--
net: lan966x: clear FDMA interrupt stickies after switch reset
The patch acks FDMA_ERRORS, FDMA_INTR_ERR and FDMA_INTR_DB right after
reset_control_reset() in lan966x_reset_switch(), so that spurious sticky
bits latched by the GCB soft reset in PCI mode do not make the FDMA IRQ
handler fire as soon as the IRQ is requested. The clear is unconditional and
is described as a no-op on the platform path.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index 6e6c08bb8eea..11094a381ec2 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -1067,6 +1067,15 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
>
> reset_control_reset(switch_reset);
>
> + /* When in PCI mode, the GCB soft reset issued by the reset
> + * controller can latch spurious bits in the FDMA error stickies.
> + * Clear them before request_irq hooks up the FDMA IRQ line,
> + * otherwise the handler fires immediately on probe.
> + */
> + lan_wr(lan_rd(lan966x, FDMA_ERRORS), lan966x, FDMA_ERRORS);
> + lan_wr(lan_rd(lan966x, FDMA_INTR_ERR), lan966x, FDMA_INTR_ERR);
> + lan_wr(lan_rd(lan966x, FDMA_INTR_DB), lan966x, FDMA_INTR_DB);
[Severity: Low]
This isn't a bug in the register writes themselves, but is the failure
described in the changelog reachable at this point in the series?
The commit message says the latched bits "sit in FDMA_INTR_ERR until the FDMA
IRQ is requested later in probe, at which point the handler fires immediately
and WARNs". In lan966x_probe() the FDMA handler is only installed when a named
"fdma" interrupt exists:
lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
if (lan966x->fdma_irq > 0) {
err = devm_request_irq(&pdev->dev, lan966x->fdma_irq,
lan966x_fdma_irq_handler, 0,
"fdma irq", lan966x);
...
lan966x->fdma = true;
}
At this commit the PCI overlay in drivers/misc/lan966x_pci.dtso does not
describe that interrupt:
interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
<9 IRQ_TYPE_LEVEL_HIGH>;
interrupt-names = "xtr", "ana";
so platform_get_irq_byname(pdev, "fdma") fails and lan966x_fdma_irq_handler()
is never requested on the PCI path.
The overlay also disagrees with the positional list in
Documentation/devicetree/bindings/net/microchip,lan966x-switch.yaml, where the
second entry is reserved for "fdma":
interrupt-names:
minItems: 1
items:
- const: xtr
- const: fdma
- const: ana
Both points are resolved later in the same series by "misc: lan966x-pci: dts:
add fdma interrupt to overlay", which switches the overlay to
interrupts <12>,<14>,<9> and interrupt-names = "xtr", "fdma", "ana", so this is
only an observation about the ordering of the changelog against the code state
here rather than an actionable defect.
Would it help to mention in the commit message that the WARN only becomes
reachable once the overlay gains the "fdma" interrupt?
> /* Don't reinitialize the switch core, if it is already initialized. In
> * case it is initialized twice, some pointers inside the queue system
> * in HW will get corrupted and then after a while the queue system gets
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 09/14] net: lan966x: add shutdown callback to stop FDMA on reboot
2026-09-09 13:00 ` [PATCH net-next v6 09/14] net: lan966x: add shutdown callback to stop FDMA on reboot Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider:
Critical: 0 · High: 0 · Medium: 4 · Low: 1
- [Medium] lan966x_shutdown() leaves the shared NAPI permanently disabled
(napi_disable() with no napi_enable() and no state marker) while the…
- [Medium] lan966x_shutdown() quiesces TX only with napi_disable() +
netif_tx_disable() and then manipulates FDMA TX state…
- [Medium] lan966x_shutdown() masks the analyzer source with a plain
lan_wr(0, ANA_ANAINTR) while the ANA interrupt is still requested and…
- [Medium] The new shutdown path relies on
lan966x_fdma_rx_disable()/lan966x_fdma_tx_disable() to prove the…
- [Low] Claim: the `if (!lan966x->fdma) return;` guard in
lan966x_shutdown() skips the ANA_ANAINTR mask, which lan966x_init()…
--
net: lan966x: add shutdown callback to stop FDMA on reboot
Adds a platform shutdown() callback that disables NAPI, stops the netdev
TX queues, disables the RX and TX FDMA channels and masks FDMA_INTR_ENA,
FDMA_INTR_DB_ENA and ANA_ANAINTR so a card that survives a host reboot
does not assert the shared PCIe INTx before re-probe. FDMA_INTR_ENA is
also re-armed in lan966x_fdma_rx_start().
A few questions about the new shutdown path below.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
> index 2695bc41e52a..5793a83268fc 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
[ ... ]
> @@ -325,6 +329,21 @@ static void lan966x_fdma_stop_netdev(struct lan966x *lan966x)
> }
> }
>
> +/* Drain in-flight xmit callers and stop all TX queues on every port. */
> +void lan966x_fdma_tx_disable_netdev(struct lan966x *lan966x)
> +{
> + struct lan966x_port *port;
> + int i;
> +
> + for (i = 0; i < lan966x->num_phys_ports; ++i) {
> + port = lan966x->ports[i];
> + if (!port)
> + continue;
> +
> + netif_tx_disable(port->dev);
> + }
> +}
[Severity: Medium]
Does netif_tx_disable() really drain every in-flight xmit caller here?
lan966x also registers .ndo_xdp_xmit = lan966x_xdp_xmit in
lan966x_port_netdev_ops, and that entry point is dispatched from a foreign
device's NAPI/softirq (bq_xmit_all()) without the netdev txq locks and
without regard for the txq stopped state or lan966x's NAPI state:
lan966x_xdp_xmit()
lan966x_fdma_xmit_xdpf()
spin_lock(&lan966x->tx_lock);
...
/* Start the transmission */
lan966x_fdma_tx_start(tx);
lan966x_fdma_tx_start() calls lan966x_fdma_tx_activate() when
tx->activated is false, and that re-sets BIT(channel_id) in
FDMA_INTR_DB_ENA and writes FDMA_CH_ACTIVATE. lan966x_fdma_tx_disable()
in the new shutdown path sets tx->activated = false, and lan966x_shutdown()
touches FDMA_CH_DISABLE, tx->activated and FDMA_INTR_DB_ENA without taking
lan966x->tx_lock.
If an XDP program on another interface redirects into a lan966x port in the
reboot window, can the TX channel end up re-activated (and the DB enable
bit re-set) after shutdown() returns? The shutdown path deliberately keeps
the DCB rings mapped, so a re-activated channel would keep fetching DCBs
from and writing DCB status into host memory that the next kernel reuses,
which is the state the commit message says it eliminates.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index 11094a381ec2..d6ce1e3e373f 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -1324,9 +1324,35 @@ static void lan966x_remove(struct platform_device *pdev)
> debugfs_remove_recursive(lan966x->debugfs_root);
> }
>
> +static void lan966x_shutdown(struct platform_device *pdev)
> +{
> + struct lan966x *lan966x = platform_get_drvdata(pdev);
> +
> + if (!lan966x->fdma)
> + return;
> +
> + /* The reload paths disable this NAPI under rtnl; serialize with them. */
> + rtnl_lock();
> +
> + if (lan966x->fdma_ndev)
> + napi_disable(&lan966x->napi);
[Severity: Medium]
The NAPI is disabled here and never re-enabled, and no state is recorded to
say the device has been shut down. After rtnl_unlock() the netdevs are
still registered and their ndo callbacks are still live.
Can a later MTU change on any port hang?
lan966x_port_change_mtu() /* .ndo_change_mtu */
lan966x->ops->fdma_resize()
lan966x_fdma_change_mtu()
__lan966x_fdma_reload()
lan966x_fdma_reload()
napi_synchronize(&lan966x->napi);
napi_disable(&lan966x->napi);
The second napi_disable() lands in napi_disable_locked(), which spins:
while (val & (NAPIF_STATE_SCHED | NAPIF_STATE_NPSVC)) {
usleep_range(20, 200);
val = READ_ONCE(n->state);
}
NAPIF_STATE_SCHED was set by the napi_disable() in lan966x_shutdown() and
is only cleared by napi_enable(), which shutdown() never calls.
Wouldn't the netlink task then loop with no timeout and no alternative wake
condition, while holding rtnl_mutex, blocking any later rtnl user in the
reboot sequence (including a second instance's lan966x_shutdown(), which
also takes rtnl_lock())? device_shutdown() runs from
kernel_restart_prepare() with userspace still runnable, and an XDP attach
reaches the same napi_disable() through lan966x_xdp_setup() ->
lan966x_fdma_reload_page_pool() -> __lan966x_fdma_reload().
> +
> + lan966x_fdma_tx_disable_netdev(lan966x);
> +
> + lan966x_fdma_rx_disable(&lan966x->rx);
> + lan966x_fdma_tx_disable(&lan966x->tx);
[Severity: Medium]
The commit message says "fdma_deinit() disables both channels and waits for
them to go idle", and shutdown() relies on the same two helpers for that
guarantee. Both of them drop the poll result though:
lan966x_fdma_rx_disable() and lan966x_fdma_tx_disable() {
readx_poll_timeout_atomic(lan966x_fdma_channel_active, lan966x,
val, !(val & BIT(fdma->channel_id)),
READL_SLEEP_US, READL_TIMEOUT_US);
}
Is a timeout on FDMA_CH_ACTIVE distinguishable from a channel that really
went idle here? lan966x_fdma_tx_disable() sets tx->activated = false
regardless, and shutdown() then masks FDMA_INTR_ENA/FDMA_INTR_DB_ENA, which
hides the error sources that would report the still-active channel, and
returns.
Unlike the unbind path, shutdown() keeps the RX pages, the coherent DCB
rings and (on the PCIe backend) the outbound ATU window mapped, so a
channel that never reached the idle state keeps write access to memory the
next kernel will reuse. Would it make sense to check the poll return and
at least report it?
> +
> + lan_wr(0, lan966x, FDMA_INTR_ENA);
> + lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
> + lan_wr(0, lan966x, ANA_ANAINTR);
[Severity: Medium]
Can the ANA_ANAINTR mask be undone right after it is written? The ANA irq
is still requested and enabled at this point, and it is a threaded handler
(devm_request_threaded_irq() with lan966x_ana_irq_handler and IRQF_ONESHOT
in lan966x_probe()), whose tail does an unlocked read-modify-write of the
same register:
lan966x_mac_irq_handler() {
...
lan_rmw(ANA_ANAINTR_INTR_SET(0),
ANA_ANAINTR_INTR,
lan966x, ANA_ANAINTR);
}
lan_rmw() is readl() -> modify -> writel() and preserves ANA_ANAINTR_INTR_ENA
from the value it read. If the handler thread has already done its readl()
(with INTR_ENA still 1) when shutdown() writes 0, wouldn't its write-back
restore ANA_ANAINTR_INTR_ENA = 1?
shutdown() holds only rtnl, which the threaded handler never takes, and it
neither disable_irq()s nor frees the ANA irq, so the source stays mapped
and unmasked. The switch keeps learning and forwarding after the host
reboots, so MAC-table sync events keep arriving. Would a
disable_irq()/synchronize_irq() on lan966x->ana_irq before the mask, or a
shared lock with the handler, close that window?
> +
> + rtnl_unlock();
> +}
> +
> static struct platform_driver lan966x_driver = {
> .probe = lan966x_probe,
> .remove = lan966x_remove,
> + .shutdown = lan966x_shutdown,
> .driver = {
> .name = "lan966x-switch",
> .of_match_table = lan966x_match,
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 10/14] net: lan966x: add PCIe FDMA support
2026-09-09 13:00 ` [PATCH net-next v6 10/14] net: lan966x: add PCIe FDMA support Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider:
Critical: 0 · High: 0 (1) · Medium: 1 (1) · Low: 3
- [Medium] DISCARDED — fixed later in the series.
- [Low] DISCARDED — fixed later in the series.
- [Low] DISCARDED — fixed later in the series.
- [Low] DISCARDED — fixed later in the series.
Pre-existing issues:
- [High] Use-after-free in the new PCIe receive path:
`lan966x_fdma_pci_rx_get_frame()`…
- [Medium] DISCARDED — pre-existing shared-library ordering weakness.
--
net: lan966x: add PCIe FDMA support
Adds a second FDMA backend for lan966x used when the switch is reached over
PCIe, using contiguous ATU-mapped DMA buffers and memcpy-based frame
transfer instead of per-page mappings. XDP and PTP are not advertised or
allowed on this path.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> new file mode 100644
> index 000000000000..f1f3c789d3a6
> --- /dev/null
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> @@ -0,0 +1,404 @@
[ ... ]
> +/* Return true if blockl is a valid RX frame size. */
> +static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
> +{
> + return blockl >= IFH_LEN_BYTES + ETH_HLEN + ETH_FCS_LEN &&
> + blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
> +}
[ ... ]
> + skb->protocol = eth_type_trans(skb, skb->dev);
> +
> + if (lan966x->bridge_mask & BIT(src_port)) {
> + skb->offload_fwd_mark = 1;
> +
> + skb_reset_network_header(skb);
> + if (!lan966x_hw_offload(lan966x, src_port, skb))
> + skb->offload_fwd_mark = 0;
> + }
> +
> + skb->dev->stats.rx_bytes += skb->len;
> + skb->dev->stats.rx_packets++;
> +
> + return skb;
> +}
[Severity: High]
This isn't a bug introduced by this patch - lan966x_fdma_rx_get_frame() in
lan966x_fdma.c has the same caller pattern - but can lan966x_fdma_pci_rx_get_frame()
touch skb after lan966x_hw_offload() has already freed it?
lan966x_hw_offload() replaces the buffer and reports only a boolean:
drivers/net/ethernet/microchip/lan966x/lan966x_main.c:lan966x_hw_offload() {
...
if (eth_type_vlan(skb->protocol)) {
skb = skb_vlan_untag(skb);
if (unlikely(!skb))
return false;
}
...
}
and skb_vlan_untag() frees the caller's skb on failure:
net/core/skbuff.c:skb_vlan_untag() {
...
if (unlikely(!pskb_may_pull(skb, VLAN_HLEN + sizeof(unsigned short))))
goto err_free;
...
err_free:
kfree_skb(skb);
return NULL;
}
So on the false return the caller's pointer is dangling, and
lan966x_fdma_pci_rx_get_frame() then does:
skb->offload_fwd_mark = 0;
skb->dev->stats.rx_bytes += skb->len;
skb->dev->stats.rx_packets++;
return skb;
with the returned skb going to napi_gro_receive() in
lan966x_fdma_pci_napi_poll().
Is the new lower bound in lan966x_fdma_pci_rx_size_fits() wide enough to let
this happen? IFH_LEN_BYTES + ETH_HLEN + ETH_FCS_LEN is 46, so a BLOCKL of
46..51 is accepted, and after skb_pull(IFH_LEN_BYTES),
skb_trim(skb->len - ETH_FCS_LEN) and eth_type_trans() there are only 0..5
bytes left. For a frame with h_proto 0x8100 on a bridged port that has IGMP
or MLD redirection enabled, the pskb_may_pull(VLAN_HLEN + 2) inside
skb_vlan_untag() then fails and the buffer is freed.
Would it make sense to have lan966x_hw_offload() take a struct sk_buff **,
or to have the caller stop using skb once it returns false?
> static int lan966x_fdma_pci_xmit(struct sk_buff *skb, __be32 *ifh,
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 11/14] net: lan966x: add PCIe FDMA MTU change support
2026-09-09 13:00 ` [PATCH net-next v6 11/14] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider:
Critical: 0 · High: 2 · Medium: 1 · Low: 1
- [High] NULL pointer dereference in the new lan966x_fdma_pci_reset_mem()
restore path.
- [High] Unbalanced/unguarded napi_disable()/napi_enable() in the new MTU
reload path.
- [Medium] Netdev MTU contract mismatch on the PCIe FDMA path.
- [Low] Inverted restart ordering in lan966x_fdma_pci_reload(): both the
success path (lines 458-460) and the restore path (lines 476-478) call…
--
net: lan966x: add PCIe FDMA MTU change support
This adds MTU change support to the PCIe FDMA path by reallocating the
contiguous ATU-mapped RX and TX buffers, falling back to resetting and
reusing the existing buffers if the reallocation fails. It also caps the
PCIe DCB ring at 256 entries so the contiguous allocation stays below
MAX_PAGE_ORDER at jumbo MTU.
A few questions about the new reload and resize paths below.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> index f1f3c789d3a6..6cabbb8b47f2 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> @@ -1,8 +1,15 @@
> // SPDX-License-Identifier: GPL-2.0+
>
> +#include <linux/mmzone.h>
> +
> #include "fdma_api.h"
> #include "lan966x_main.h"
>
> +/* Ring must fit in one MAX_PAGE_ORDER DMA block; 512 DCBs overflows
> + * at jumbo MTU.
> + */
> +#define FDMA_PCI_DCB_MAX 256
> +
[ ... ]
> @@ -374,9 +381,164 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
> return 0;
> }
>
> +/* Reset existing rx and tx buffers. */
> +static void lan966x_fdma_pci_reset_mem(struct lan966x *lan966x)
> +{
> + struct lan966x_rx *rx = &lan966x->rx;
> + struct lan966x_tx *tx = &lan966x->tx;
> +
> + memset(rx->fdma.dcbs, 0, rx->fdma.size);
> + memset(tx->fdma.dcbs, 0, tx->fdma.size);
> +
> + fdma_dcbs_init(&rx->fdma,
> + FDMA_DCB_INFO_DATAL(rx->fdma.db_size - XDP_PACKET_HEADROOM),
> + FDMA_DCB_STATUS_INTR);
> +
> + fdma_dcbs_init(&tx->fdma,
> + FDMA_DCB_INFO_DATAL(tx->fdma.db_size),
> + FDMA_DCB_STATUS_DONE);
> +
> + lan966x_fdma_llp_configure(lan966x,
> + tx->fdma.atu_region->base_addr,
> + tx->fdma.channel_id);
[Severity: High]
Can tx->fdma.atu_region be NULL here?
lan966x_fdma_pci_reset_mem() is only reached from the restore label of
lan966x_fdma_pci_reload(), and that label is taken when the ring
allocation failed. Probe registers the netdevs from
lan966x_probe_port() -> register_netdev() inside the ports loop, while
lan966x->fdma is already true, and ops->fdma_init() (which sets
n_dcbs/db_size/size and maps the ATU regions) only runs after that loop.
An MTU change in that window passes the !lan966x->fdma guard in
lan966x_port_change_mtu() and reaches lan966x_fdma_pci_resize() with a
zeroed fdma:
rx.max_mtu == 0, so the "max_mtu == lan966x->rx.max_mtu" early return
does not fire
n_dcbs == 0, so fdma_get_size_contiguous() returns ALIGN(0, PAGE_SIZE)
== 0 and both -ERANGE guards pass
lan966x_fdma_pci_reload() then recomputes size, which stays 0, and:
lan966x_fdma_pci_rx_alloc()
fdma_alloc_coherent_and_map()
fdma_pci_atu_region_map()
if (size <= 0)
return ERR_PTR(-EINVAL);
so control reaches restore, memcpy's the still-zeroed fdma structs back
(atu_region == NULL, dcbs == NULL, size == 0), and calls reset_mem().
memset(NULL, 0, 0) and fdma_dcbs_init() with n_dcbs == 0 are both
no-ops, so nothing stops execution before tx->fdma.atu_region->base_addr
is evaluated. Would a NULL check on atu_region (or an early bail in
resize() when the FDMA is not initialized yet) be appropriate here?
> + lan966x_fdma_llp_configure(lan966x,
> + rx->fdma.atu_region->base_addr,
> + rx->fdma.channel_id);
> +}
> +
> +/* Wake all TX queues on every port (undoes lan966x_fdma_tx_disable_netdev). */
> +static void lan966x_fdma_pci_wakeup_netdev(struct lan966x *lan966x)
> +{
> + for (int i = 0; i < lan966x->num_phys_ports; ++i) {
> + struct lan966x_port *port = lan966x->ports[i];
> +
> + if (port)
> + netif_tx_wake_all_queues(port->dev);
> + }
> +}
> +
> +static int lan966x_fdma_pci_reload(struct lan966x *lan966x, int new_mtu)
> +{
> + struct fdma tx_fdma_old = lan966x->tx.fdma;
> + struct fdma rx_fdma_old = lan966x->rx.fdma;
> + u32 old_mtu = lan966x->rx.max_mtu;
> + int err;
> +
> + napi_disable(&lan966x->napi);
[Severity: High]
Should this napi_disable() be guarded the way the other users of
lan966x->napi in this driver are?
lan966x_fdma_pci_deinit() does:
if (lan966x->fdma_ndev)
napi_disable(&lan966x->napi);
and lan966x_shutdown() has the same guard, with a comment noting that the
reload paths disable this NAPI under rtnl. Two states look problematic
for the unguarded call:
The NAPI may not have been added yet. netif_napi_add() only runs from
lan966x_fdma_netdev_init(), called by lan966x_port_init(), which happens
after lan966x_probe_port() already did register_netdev(). An MTU change
in that window reaches napi_disable() with n->dev == NULL (lan966x is
devm_kzalloc'ed), and napi_disable() does netdev_lock(n->dev).
The NAPI may already be disabled. lan966x_remove() calls
ops->fdma_deinit() (which disables the NAPI and frees/unmaps both rings)
before lan966x_cleanup_ports() unregisters the netdevs, and
lan966x_shutdown() disables the NAPI without clearing fdma_ndev. A
concurrent MTU change then calls napi_disable() a second time and
napi_disable_locked() spins:
net/core/dev.c:napi_disable_locked() {
...
while (val & (NAPIF_STATE_SCHED | NAPIF_STATE_NPSVC)) {
usleep_range(20, 200);
val = READ_ONCE(n->state);
}
...
}
There is no timeout and no other wake condition, and ndo_change_mtu
holds rtnl throughout, which also blocks the unregister_netdev() that
would end the window. The matching napi_enable() calls on both exit
paths below have the same issue.
> + lan966x_fdma_tx_disable_netdev(lan966x);
> + lan966x_fdma_rx_disable(&lan966x->rx);
> + lan966x_fdma_tx_disable(&lan966x->tx);
> +
> + lan966x->rx.max_mtu = new_mtu;
> +
> + /* Must be NULL'ed in order to realloc them. */
> + lan966x->rx.fdma.atu_region = NULL;
> + lan966x->tx.fdma.atu_region = NULL;
> +
> + lan966x->tx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
> + lan966x->tx.fdma.size = fdma_get_size_contiguous(&lan966x->tx.fdma);
> + lan966x->rx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
> + lan966x->rx.fdma.size = fdma_get_size_contiguous(&lan966x->rx.fdma);
> +
> + err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
> + if (err)
> + goto restore;
> +
> + err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
> + if (err) {
> + fdma_free_coherent_and_unmap(lan966x->dma_dev,
> + &lan966x->rx.fdma);
> + goto restore;
> + }
> +
> + /* Free and unmap old memory. */
> + fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
> + fdma_free_coherent_and_unmap(lan966x->dma_dev, &tx_fdma_old);
> +
> + napi_enable(&lan966x->napi);
> + lan966x_fdma_rx_start(&lan966x->rx);
> + lan966x_fdma_pci_wakeup_netdev(lan966x);
[Severity: Low]
This isn't a bug as such, but the restart order here is the reverse of
the pre-existing non-PCI path. lan966x_fdma_reload() does:
lan966x_fdma_rx_start(&lan966x->rx);
...
lan966x_fdma_wakeup_netdev(lan966x);
napi_enable(&lan966x->napi);
Once napi_enable() clears NAPI_STATE_DISABLE, a pending FDMA DB
interrupt reaching napi_schedule() from lan966x_fdma_irq_handler() can
run lan966x_fdma_pci_napi_poll() on another CPU concurrently with
lan966x_fdma_rx_start(). Both then do unsynchronized read-modify-write
on FDMA_INTR_DB_ENA:
lan966x_fdma_pci_napi_poll():
if (counter < weight && napi_complete_done(napi, counter))
lan_wr(0xff, lan966x, FDMA_INTR_DB_ENA);
lan966x_fdma_rx_start() instead does lan_rd() followed by lan_rmw() on
the same register, and the poll can also issue FDMA_CH_RELOAD for a
channel rx_start() has not activated yet. Would matching the ordering
used by lan966x_fdma_reload() be preferable?
> +
> + return err;
> +restore:
> +
> + /* No new buffers are allocated at this point. Use the old buffers,
> + * but reset them before starting the FDMA again.
> + */
> +
> + memcpy(&lan966x->tx.fdma, &tx_fdma_old, sizeof(struct fdma));
> + memcpy(&lan966x->rx.fdma, &rx_fdma_old, sizeof(struct fdma));
> +
> + lan966x->rx.max_mtu = old_mtu;
> +
> + lan966x_fdma_pci_reset_mem(lan966x);
> +
> + napi_enable(&lan966x->napi);
> + lan966x_fdma_rx_start(&lan966x->rx);
> + lan966x_fdma_pci_wakeup_netdev(lan966x);
> +
> + return err;
> +}
> +
[ ... ]
> static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
> {
> - return -EOPNOTSUPP;
> + struct fdma rx_fdma;
> + int max_mtu;
> +
> + max_mtu = lan966x_fdma_get_max_frame(lan966x);
> + if (max_mtu == lan966x->rx.max_mtu)
> + return 0;
> +
> + /* rx and tx have n_dbs == 1, so both rings need the same contiguous
> + * dma_alloc_coherent() block, which can't exceed MAX_PAGE_ORDER. The
> + * allocation is padded to the ATU region granularity, so test the
> + * padded size.
> + */
> + rx_fdma = lan966x->rx.fdma;
> + rx_fdma.db_size = FDMA_PCI_DB_SIZE(max_mtu);
> + if (ALIGN(fdma_get_size_contiguous(&rx_fdma),
> + FDMA_PCI_ATU_REGION_ALIGN) > (PAGE_SIZE << MAX_PAGE_ORDER))
> + return -ERANGE;
[Severity: Medium]
Does the MTU range the netdev advertises still match what this path can
actually do?
lan966x_probe_port() sets dev->max_mtu = ETH_MAX_MTU for the PCIe
instantiation too, but with n_dcbs == FDMA_PCI_DCB_MAX (256),
n_dbs == FDMA_RX_DCB_MAX_DBS (1) and db_size == ALIGN(max_frame, 128),
the padded contiguous allocation crosses PAGE_SIZE << MAX_PAGE_ORDER at
roughly 15.7 KiB of MTU on a 4 KiB-page build, so larger requests get a
bare -ERANGE that lan966x_port_change_mtu() passes straight to
userspace, and the cut-off moves with PAGE_SIZE and MAX_PAGE_ORDER.
There is a second effect for MTUs in [65518, 65535], which the
advertised max_mtu still permits. lan966x_port_change_mtu() programs
the register before calling the resize callback:
lan_wr(DEV_MAC_MAXLEN_CFG_MAX_LEN_SET(LAN966X_HW_MTU(new_mtu)),
lan966x, DEV_MAC_MAXLEN_CFG(port->chip_port));
DEV_MAC_MAXLEN_CFG_MAX_LEN is GENMASK(15, 0) and LAN966X_HW_MTU(mtu) is
mtu + 18, so FIELD_PREP() masks the runtime value, e.g. mtu 65535
programs MAX_LEN = 65553 & 0xffff = 17. lan966x_fdma_pci_resize() then
derives max_mtu from that already-truncated register via
lan966x_fdma_get_max_frame() -> lan966x_fdma_get_max_mtu() ->
lan_rd(DEV_MAC_MAXLEN_CFG), so it either returns 0 (multi-port case,
where the max over ports is unchanged) or reloads with a tiny db_size,
and the MTU change is reported as successful while the MAC is configured
for a 17-byte frame.
Before this patch the PCIe callback returned -EOPNOTSUPP for every
resize, so lan966x_port_change_mtu() always rolled the register back;
this patch is what makes the path reachable. Would capping dev->max_mtu
for the PCIe path (to both the ring limit and the 16-bit MAC field) be
the better fix, so the resize-time rejection becomes unreachable?
> +
> + /* db_size is also handed to the FDMA in the 16-bit DCB DATAL field,
> + * where a larger value would be silently truncated.
> + */
> + if (rx_fdma.db_size > GENMASK(15, 0))
> + return -ERANGE;
> +
> + return __lan966x_fdma_pci_reload(lan966x, max_mtu);
> }
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 12/14] net: lan966x: add PCIe FDMA XDP support
2026-09-09 13:00 ` [PATCH net-next v6 12/14] net: lan966x: add PCIe FDMA XDP support Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:
Critical: 0 · High: 0 · Medium: 1 · Low: 0
- [Medium] Inconsistent XDP frame presentation between the two RX hooks
of the same driver: after this patch lan966x has two in-tree XDP RX…
--
net: lan966x: add PCIe FDMA XDP support
This adds XDP support to the lan966x PCIe FDMA path, using contiguous
ATU-mapped buffers and a memcpy-based XDP_TX instead of the page_pool
scheme used by the platform path. The IFH and FCS are stripped before
the BPF program runs, and lan966x_xdp_setup() now skips the page_pool
reload for the PCIe case. Only XDP_ACT_BASIC is advertised for PCIe.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> index 6cabbb8b47f2..02c021617a3e 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
[ ... ]
> @@ -126,7 +127,123 @@ static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
[ ... ]
> +static int lan966x_xdp_pci_run(struct lan966x_port *port, void *data,
> + u32 data_len, void **xdp_data, u32 *xdp_len)
> +{
[ ... ]
> + xdp_prepare_buff(&xdp,
> + data - XDP_PACKET_HEADROOM,
> + XDP_PACKET_HEADROOM + IFH_LEN_BYTES,
> + data_len,
> + false);
> +
> + act = bpf_prog_run_xdp(xdp_prog, &xdp);
[ ... ]
> @@ -158,38 +275,33 @@ static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
> if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
> return FDMA_ERROR;
>
> - return FDMA_PASS;
> + /* Present the Ethernet frame (no IFH, no FCS). HW re-inserts the
> + * FCS on TX; see lan966x_fdma_pci_xmit_xdpf(). May be overridden
> + * by XDP. The FCS strip is unconditional because NETIF_F_RXFCS
> + * is not advertised in hw_features.
> + */
> + *data = virt_addr + IFH_LEN_BYTES;
> + *data_len = blockl - IFH_LEN_BYTES - ETH_FCS_LEN;
> +
> + return lan966x_xdp_pci_run(port, virt_addr, *data_len, data, data_len);
> }
[Severity: Medium]
Should the two lan966x XDP receive hooks agree on whether the FCS is part
of the window handed to the BPF program?
On the PCIe path, lan966x_fdma_pci_rx_check_frame() subtracts the FCS
before calling into the program:
*data_len = blockl - IFH_LEN_BYTES - ETH_FCS_LEN;
return lan966x_xdp_pci_run(port, virt_addr, *data_len, data, data_len);
On the platform path, lan966x_fdma_rx_check_frame() passes the raw BLOCKL
value and lan966x_xdp_run() only removes the IFH:
drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c:lan966x_fdma_rx_check_frame() {
...
return lan966x_xdp_run(port, page, FDMA_DCB_STATUS_BLOCKL(db->status));
}
drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c:lan966x_xdp_run() {
...
xdp_prepare_buff(&xdp, page_address(page),
IFH_LEN_BYTES + XDP_PACKET_HEADROOM,
data_len - IFH_LEN_BYTES, false);
}
That BLOCKL value still contains the FCS, as shown by the trim done later
on the same path:
drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c:lan966x_fdma_rx_get_frame() {
...
skb_put(skb, FDMA_DCB_STATUS_BLOCKL(db->status));
...
skb_pull(skb, IFH_LEN_BYTES);
if (likely(!(skb->dev->features & NETIF_F_RXFCS)))
skb_trim(skb, skb->len - ETH_FCS_LEN);
}
So xdp.data_end points four bytes past the end of the Ethernet payload on
the platform path, and exactly at the end of the payload on the new PCIe
path. Both windows stay inside the DMA buffer, so this looks like a
behavioural difference rather than an out-of-bounds access.
A program that parses backwards from data_end, compares data_end - data
against an L3 length field, or trims via bpf_xdp_adjust_tail() would get
different results depending on whether the switch is platform- or
PCIe-attached. The PCIe framing matches the usual XDP convention of no
FCS, which would make the platform hook the one that differs from
convention.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index e089c08b9c37..9ada497767c6 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -874,11 +874,13 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
>
> port->phylink = phylink;
>
> - /* XDP is not supported on the PCIe FDMA path. */
> - if (lan966x->fdma && !lan966x_is_pci(lan966x))
> - dev->xdp_features = NETDEV_XDP_ACT_BASIC |
> - NETDEV_XDP_ACT_REDIRECT |
> - NETDEV_XDP_ACT_NDO_XMIT;
> + if (lan966x->fdma) {
> + dev->xdp_features = NETDEV_XDP_ACT_BASIC;
> +
> + if (!lan966x_is_pci(lan966x))
> + dev->xdp_features |= NETDEV_XDP_ACT_REDIRECT |
> + NETDEV_XDP_ACT_NDO_XMIT;
> + }
[Severity: Medium]
Both paths now report the same NETDEV_XDP_ACT_BASIC bit from
lan966x_probe_port(), so is there anything in the uAPI that lets a
program distinguish the FCS-included window of lan966x_xdp_run() from
the FCS-stripped window of lan966x_xdp_pci_run()?
Would it make sense to make one of the two hooks match the other so that
the frame presentation is the same for identically-advertised netdevs?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 13/14] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
2026-09-09 13:00 ` [PATCH net-next v6 13/14] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:
Critical: 0 · High: 0 · Medium: 0 · Low: 1
- [Low] Claimed patch-ordering / bisectability bug: the "cpu" reg window
is enlarged from 0x100000 to 0x800000 only in this commit, while the…
--
misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
This grows the "cpu" register window of the switch node in the lan966x PCI
device tree overlay from 0x100000 to 0x800000 so that the PCIE DBI and iATU
registers at offset 0x400000+ are covered. Those registers are programmed
when the FDMA engine sets up the ATU outbound windows.
No Link: tags in the commit header.
> diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> index 7b196b0a0eb6..7bb726550caf 100644
> --- a/drivers/misc/lan966x_pci.dtso
> +++ b/drivers/misc/lan966x_pci.dtso
> @@ -135,7 +135,7 @@ lan966x_phy1: ethernet-lan966x_phy@2 {
>
> switch: switch@e0000000 {
> compatible = "microchip,lan966x-switch";
> - reg = <0xe0000000 0x0100000>,
> + reg = <0xe0000000 0x0800000>,
> <0xe2000000 0x0800000>;
> reg-names = "cpu", "gcb";
[Severity: Low]
This isn't a bug, just a question about the ordering of the series.
The earlier commits "net: lan966x: add PCIe FDMA support", "net: lan966x: add
PCIe FDMA MTU change support" and "net: lan966x: add PCIe FDMA XDP support"
already derive regs[TARGET_PCIE_DBI] from the "cpu" base plus 0x400000, and
fdma_pci_atu_init() writes at DBI + 0x300000 + idx * 0x200:
drivers/net/ethernet/microchip/fdma/fdma_pci.c
static void __iomem *fdma_pci_atu_addr_get(void __iomem *addr, int offset, int idx)
{
return addr + FDMA_PCI_ATU_ADDR + FDMA_PCI_ATU_IDX_SIZE * idx + offset;
}
That lands roughly 0x700000 from the "cpu" base, i.e. outside the 0x100000
mapping that exists until this commit. Would it be clearer to move this
overlay change before the driver commits that compute the DBI offset, so
that the mapping is never smaller than the offsets the driver derives from
it?
For what it is worth, the intermediate commits do look safe in practice,
since the ATU access in lan966x_fdma_pci_init() is gated:
drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
if (!lan966x->fdma)
return 0;
...
fdma_pci_atu_init(&lan966x->atu, lan966x->regs[TARGET_PCIE_DBI]);
and lan966x->fdma is only set when the node names an "fdma" interrupt in
lan966x_probe():
lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
if (lan966x->fdma_irq > 0) {
...
lan966x->fdma = true;
}
The overlay here still carries only interrupt-names = "xtr", "ana", and the
"fdma" interrupt is added by the following commit, so the enlarged window is
in place before the FDMA path can be enabled. Is that ordering intentional?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay
2026-09-09 13:00 ` [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
@ 2026-09-10 13:05 ` netdev-bot+sashiko
0 siblings, 0 replies; 23+ messages in thread
From: netdev-bot+sashiko @ 2026-09-10 13:05 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
richardcochran, netdev, linux-kernel, bpf, linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:
Critical: 0 · High: 0 · Medium: 1 · Low: 0
- [Medium] The FDMA interrupt is hooked up (and thereby unmasked in the
OIC) long before any FDMA state exists.
--
misc: lan966x-pci: dts: add fdma interrupt to overlay
This adds OIC interrupt 14 as the "fdma" named interrupt to the lan966x
PCI device tree overlay, so FDMA-based injection/extraction can be used
when the switch is reached over PCIe.
> diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> index 7bb726550caf..5bb12dbc0843 100644
> --- a/drivers/misc/lan966x_pci.dtso
> +++ b/drivers/misc/lan966x_pci.dtso
> @@ -141,8 +141,9 @@ switch: switch@e0000000 {
>
> interrupt-parent = <&oic>;
> interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
> + <14 IRQ_TYPE_LEVEL_HIGH>,
> <9 IRQ_TYPE_LEVEL_HIGH>;
> - interrupt-names = "xtr", "ana";
> + interrupt-names = "xtr", "fdma", "ana";
[Severity: Medium]
With source 14 now named "fdma", platform_get_irq_byname(pdev, "fdma")
starts succeeding on the PCIe path, and lan966x_probe() requests the
handler right there:
drivers/net/ethernet/microchip/lan966x/lan966x_main.c:lan966x_probe() {
lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
if (lan966x->fdma_irq > 0) {
err = devm_request_irq(&pdev->dev, lan966x->fdma_irq,
lan966x_fdma_irq_handler, 0,
"fdma irq", lan966x);
...
lan966x->fdma = true;
}
}
Is this ordering safe? The OIC maps and unmasks source 14 at
devm_request_irq() time, but at that point no FDMA state exists yet:
- netif_napi_add()/napi_enable() only run later, from
lan966x_port_init() -> lan966x_fdma_netdev_init() in the per-port
probe loop, so lan966x->napi is still the zeroed devm_kzalloc()
memory and napi.poll is NULL.
- the DCB rings, the LLP/ATU windows and the FDMA_CTRL NRESET pulse are
all done even later, from lan966x->ops->fdma_init() ->
lan966x_fdma_pci_init().
If the OIC delivers an FDMA interrupt in that window, the handler
schedules NAPI unconditionally:
drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c:lan966x_fdma_irq_handler() {
db = lan_rd(lan966x, FDMA_INTR_DB);
err = lan_rd(lan966x, FDMA_INTR_ERR);
if (db) {
lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
lan_wr(db, lan966x, FDMA_INTR_DB);
napi_schedule(&lan966x->napi);
}
}
Can this then oops in __napi_poll(), either calling n->poll() while it
is still NULL, or running lan966x_fdma_pci_napi_poll() against
rx->fdma.dcbs == NULL and n_dcbs == 0, where fdma_has_frames() ->
fdma_db_get() dereferences the dcbs array?
How reachable is that window in practice? The series itself describes
the card arriving with FDMA state from a previous boot. The commit
"net: lan966x: clear FDMA interrupt stickies after switch reset" says
the stickies "sit in FDMA_INTR_ERR until the FDMA IRQ is requested later
in probe, at which point the handler fires immediately and WARNs", and
"net: lan966x: add shutdown callback to stop FDMA on reboot" says that
on the endpoint "channels stay active and interrupt sources stay armed
across the reset" and that "FDMA_INTR_ENA persists on the card across a
warm reboot".
The sticky clear in lan966x_reset_switch() clears the latched bits at
one instant, but does not mask FDMA_INTR_ENA/FDMA_INTR_DB_ENA and does
not stop the channels. After a reboot that skips device_shutdown()
(panic reboot, kexec/kdump, watchdog warm reset with the card still
powered), could a still-active channel latch FDMA_INTR_DB again after
the clear and have it delivered as soon as source 14 is unmasked?
Would it be more robust to reset and mask the FDMA block (the
FDMA_CTRL NRESET pulse that already exists in lan966x_fdma_pci_init(),
plus clearing FDMA_INTR_ENA/FDMA_INTR_DB_ENA) before
devm_request_irq(), or to request the FDMA IRQ only after fdma_init()
has built the rings and added the NAPI instance?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671%40microchip.com
^ permalink raw reply [flat|nested] 23+ messages in thread
end of thread, other threads:[~2026-09-10 13:05 UTC | newest]
Thread overview: 23+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-09 13:00 [PATCH net-next v6 00/14] net: lan966x: add support for PCIe FDMA Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 01/14] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 02/14] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 03/14] net: microchip: fdma: add PCIe ATU support Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 04/14] net: lan966x: add FDMA LLP register write helper Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 05/14] net: lan966x: export FDMA helpers for reuse Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 06/14] net: lan966x: use a dedicated device for DMA operations Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 07/14] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
2026-09-09 13:00 ` [PATCH net-next v6 08/14] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 09/14] net: lan966x: add shutdown callback to stop FDMA on reboot Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 10/14] net: lan966x: add PCIe FDMA support Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 11/14] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 12/14] net: lan966x: add PCIe FDMA XDP support Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 13/14] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
2026-09-09 13:00 ` [PATCH net-next v6 14/14] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
2026-09-10 13:05 ` netdev-bot+sashiko
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®