* [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA
@ 2026-09-28 19:32 Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
` (14 more replies)
0 siblings, 15 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When lan966x operates as a PCIe endpoint, the driver currently uses
register-based I/O for frame injection and extraction. This approach is
functional but slow, topping out at around 33 Mbps on an Intel x86 host
with a lan966x PCIe card.
This series adds FDMA (Frame DMA) support for the PCIe path. When
operating as a PCIe endpoint, the internal FDMA engine on lan966x cannot
directly access host memory, so DMA buffers are allocated as contiguous
coherent memory and mapped through the PCIe Address Translation Unit
(ATU). The ATU provides outbound windows that translate internal FDMA
addresses to PCIe bus addresses, allowing the FDMA engine to read and
write host memory. Because the ATU requires contiguous address regions,
page_pool and normal per-page DMA mappings cannot be used. Instead,
frames are transferred using memcpy between the ATU-mapped buffers and
the network stack. With this, throughput increases from ~33 Mbps to
~620 Mbps for default MTU.
Patch 1 adds the shared drivers/net/ethernet/microchip/fdma/ directory
to the Sparx5 SoC MAINTAINERS entry.
Patches 2-4 prepare the shared FDMA library: patch 2 renames the
contiguous dataptr helpers for clarity, patch 3 adds PCIe ATU region
management and coherent DMA allocation with ATU mapping, and patch 4
gives the descriptor fields an explicit little-endian type, since on the
PCIe path the descriptors are written by the host CPU.
Patches 5-8 refactor the lan966x FDMA code to support both platform
and PCIe paths: extracting the LLP register write into a helper,
exporting shared functions, introducing a dedicated device for DMA
operations, and adding an ops dispatch table selected at probe time.
Patches 9-10 harden the existing FDMA path for the PCIe endpoint
lifecycle: patch 9 clears latched FDMA error/interrupt stickies after
the switch reset so they don't assert as soon as interrupts are
enabled, and patch 10 adds a shutdown() callback that quiesces the
FDMA engine on host warm reboot (on the PCIe card the FDMA survives
host reset and would otherwise keep the shared INTx asserted into
the next probe).
Patch 11 adds the core PCIe FDMA implementation with RX/TX using
contiguous ATU-mapped buffers. Patches 12 and 13 extend it with MTU
change and XDP support respectively. XDP_PASS, XDP_TX, XDP_DROP and
XDP_ABORTED are supported; XDP_REDIRECT is deliberately not, because
the PCIe data path does not use page_pool.
Patches 14-15 update the lan966x PCI device tree overlay to extend the
cpu register mapping to cover the ATU register space and add the FDMA
interrupt.
Patches 1-13 touch MAINTAINERS and drivers/net/ethernet/microchip/, and
are for the netdev tree.
Patches 14-15 touch drivers/misc/lan966x_pci.dtso, and are for the
char-misc tree.
To: Andrew Lunn <andrew+netdev@lunn.ch>
To: David S. Miller <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Horatiu Vultur <horatiu.vultur@microchip.com>
To: Steen Hegelund <steen.hegelund@microchip.com>
To: UNGLinuxDriver@microchip.com
To: Alexei Starovoitov <ast@kernel.org>
To: Daniel Borkmann <daniel@iogearbox.net>
To: Jesper Dangaard Brouer <hawk@kernel.org>
To: John Fastabend <john.fastabend@gmail.com>
To: Stanislav Fomichev <sdf@fomichev.me>
To: Herve Codina <herve.codina@bootlin.com>
To: Arnd Bergmann <arnd@arndb.de>
To: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
To: Mohsin Bashir <mohsin.bashr@gmail.com>
To: Simon Horman <horms@kernel.org>
Cc: Richard Cochran <richardcochran@gmail.com>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: bpf@vger.kernel.org
Cc: linux-arm-kernel@lists.infradead.org
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
Changes in v9:
- PCIe: use a private copy of lan966x_hw_offload() that reports when
skb_vlan_untag() frees the skb, fixing a use-after-free reachable via
XDP shrinking the frame.
- Link to v8: https://lore.kernel.org/r/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b@microchip.com
Changes in v8:
- New patch 4: use little-endian types for the FDMA descriptor fields.
On the PCIe path the descriptors live in host memory and are written
by the host CPU, which may be big-endian, while the FDMA engine reads
them little-endian. The fields become __le64 and are converted at the
library boundary; the dataptr/nextptr callbacks keep their u64
signatures, so sparx5 and lan969x need no changes. Checked with sparse
(-D__CHECK_ENDIAN__). No functional change on little-endian hosts.
Tested on lan966x (arm), sparx5 and lan969x (arm64), and lan966x over
PCIe on an x86 host, with no regressions. (Simon)
- PCIe: read the RX block length through fdma_db_len_get() instead of
open-coding the status field, following patch 4.
- Stickies: extend the comment to also mention the data-block sticky,
which is cleared alongside the error stickies.
- Shutdown: reword the commit message to describe the tree at that
commit, noting where it relies on the PCIe FDMA backend added later
in the series.
- Link to v7: https://lore.kernel.org/r/20260918-lan966x-pci-fdma-v7-0-0ecc179c8a2c@microchip.com
Changes in v7:
Mostly a response to the automated review of v6. The notable change is
that the resize-time MTU guards added in v6 are replaced by an advertised
limit. v6 rejected an oversized MTU inside lan966x_fdma_pci_resize() with
-ERANGE, which is late, and leaves dev->max_mtu advertising a range the
PCIe path cannot honour. v7 derives the maximum from the same constants
and publishes it in dev->max_mtu instead, so the core rejects the request
up front and both -ERANGE returns become unreachable and are removed.
Also, the shutdown() was refactored to not race with ndo's.
- PCIe: cap dev->max_mtu on the PCIe path (FDMA_PCI_MAX_MTU), derived
from the contiguous ring having to fit one MAX_PAGE_ORDER block after
padding to the ATU region granularity, and from db_size having to fit
the 16-bit DCB DATAL field. On a 4KB-page, MAX_PAGE_ORDER=10 build that
is MTU 15498. Both resize-time -ERANGE checks are removed with it.
- PCIe: factor the frame overhead into LAN966X_FDMA_OVERHEAD and use it
for both the cap and lan966x_fdma_get_max_frame(), so it is not
open-coded twice.
- PCIe: skip the resize while the FDMA is not initialised. The rings are
built by lan966x_fdma_pci_init(), which runs after the netdevs are
registered, so an MTU change in that window would dereference NULL.
- PCIe: comment why NAPI has to be enabled before the queues are woken in
the MTU reload path.
- shutdown(): rework the quiesce ordering. Free the xtr, ana and FDMA
irqs first so no source can assert INTx, mask ANA_ANAINTR whether or not
the FDMA is in use, then disable NAPI, stop and detach the netdevs,
disable the FDMA channels, clear the FDMA interrupt enables and unmap
the ATU windows.
- ATU: document the alignment requirement and -EINVAL at
fdma_pci_atu_region_map(), document at fdma_pci_atu_translate_addr()
that the caller must quiesce DMA and the descriptor paths before
unmapping a region, initialise the region bookkeeping explicitly in
fdma_pci_atu_init(), and fix the comment on the upper limit register.
- Patch 3 commit message: drop the claim that the patch adds PCIe FDMA
support; it adds the helpers the lan966x path later uses.
Reported by the v6 review and deliberately not fixed here, because the
code is pre-existing and the fixes want a Fixes: tag:
- register_netdev() and the FDMA irq request both happen before
ops->fdma_init(), so an MTU change or a stale interrupt can reach an
uninitialised FDMA. Pre-existing on the platform path; the fix is a
reorder of the shared probe, which I will send against net. The skip
above closes the MTU-change half on PCIe in the meantime.
- lan966x_hw_offload() frees the caller's skb and reports only a bool.
Pre-existing, all three callers affected. Goes to net with
Fixes: 6476f90aefaf.
- The platform XDP hook leaves the FCS in the window handed to the BPF
program, while the PCIe hook strips it. Pre-existing. Goes to net with
Fixes: 6a2159be7604.
- Link to v6: https://lore.kernel.org/r/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671@microchip.com
Changes in v6:
The main change is a DMA bug on the PCIe path, found after enabling the
IOMMU on the test host. The rest is hardening and cleanup.
- New patch 6: use a dedicated device for DMA operations. On the PCIe
path lan966x->dev is the platform device created by
of_platform_default_populate(), which carries no iommus/dma-ranges of
its own, so it is not the device the IOMMU has a domain for. DMA
mapped against it lands outside the domain the endpoint's requester ID
is associated with. With the IOMMU enabled, AMD-Vi reported
IO_PAGE_FAULT for the endpoint and no traffic passed. All FDMA DMA now
targets the real PCIe endpoint device. Retested with the IOMMU in
translated (DMA-FQ) mode.
- ATU: serialise region allocation/free and ATU register access with a
mutex.
- ATU: program the region limit per mapping (base_addr + size - 1)
instead of the full region bound, and disable every region in
fdma_pci_atu_init().
- ATU: return -ERANGE instead of -E2BIG when the requested size exceeds
the region size, and clear fdma->atu_region on unmap.
- PCIe FDMA: check fdma_dcbs_init() and unwind the coherent allocation
and the ATU mapping on failure.
- PCIe FDMA: drop the WARN_ON() on an out-of-range src_port; a malformed
IFH should not be able to splat the host.
- Quiesce consistently in shutdown, deinit and the MTU reload: stop the
netdev queues and NAPI before disabling the FDMA channels, and share
one lan966x_fdma_tx_disable_netdev() helper instead of open-coding it
per path.
- shutdown: also mask ANA_ANAINTR, which shares the PCIe INTx with the
FDMA sources.
- MTU change: reject an MTU whose ring would not fit in a single
MAX_PAGE_ORDER coherent block, and reject one whose db_size would be
truncated by the 16-bit DCB DATAL field, rather than programming a
wrong buffer length.
- Assign lan966x->ops directly in the ops dispatch patch, so that patch
no longer adds a helper returning a single constant; PCI detection now
derives from dma_dev and arrives with the PCIe implementation.
- ATU: reject a target address that is not aligned to the 64KB ATU region
granularity, instead of programming a translation offset by the
misalignment.
- Drop the CONFIG_MCHP_LAN966X_PCI guard around the fdma_pci.h include in
fdma_api.h (Jakub).
- Use ifneq ($(CONFIG_MCHP_LAN966X_PCI),) instead of ifdef in both the fdma
and lan966x Makefiles; the symbol is tristate, so the condition has to
cover both y and m (Jakub).
- Reprogram the FDMA LLP registers on the platform reload restore path,
which stopped happening when the LLP write moved into the allocation
functions.
- XDP: drop the napi_synchronize() before freeing the old program on the
PCIe path; READ_ONCE() on port->xdp_prog plus the RCU-deferred
bpf_prog_put() already cover the in-flight poll.
- XDP: note that the ETH_ZLEN floor is deliberately not enforced on XDP_TX.
- PCIe FDMA: count tx_dropped when the DCB ring is full.
- PCIe FDMA: guard the deinit napi_disable() on fdma_ndev, as shutdown()
already does; NAPI is only initialised on the first ndo_open.
- ATU: program the region translation before enabling the region.
- ATU: pad the mapped allocation to the 64KB outbound region granularity,
so the window does not extend past the coherent allocation.
- ATU: reject mapping an fdma that already holds a region, instead of
overwriting the handle and leaking the old mapping. The MTU reload clears
the handle on the live rings first, since it keeps the old mapping alive
until the new rings are in place.
- PCIe FDMA: declare the RX DCB DATAL as the space actually available to the
extraction engine (db_size - XDP_PACKET_HEADROOM), since the data block
pointer starts that far into the slot.
- Move the XDP-unsupported guards from the XDP patch into the PCIe FDMA
patch, so that no intermediate commit advertises NETDEV_XDP_ACT_NDO_XMIT
for a PCIe instance while tx->dcbs_buf is NULL, and none lets an XDP
attach reach the platform page_pool reload. The final tree is unchanged.
- shutdown: take rtnl, so it cannot interleave with the rtnl-only reload
paths and double-disable the NAPI, which would spin forever in
napi_disable().
- Link to v5: https://lore.kernel.org/r/20260520-lan966x-pci-fdma-v5-0-ca56197ae05b@microchip.com
Changes in v5:
This version fixes a single AI review issue, flagged by Paolo. Other AI
issues for v4 has been classified as pre-existing or changes for
follow-ups.
- Fix premature napi_complete_done() in lan966x_fdma_pci_napi_poll() on
FDMA_ERROR and napi_alloc_skb() failure. Bailing out left DONE=1 DCBs
in the ring with no IRQ to drain them. Drop the frame and continue
the poll loop instead. Bump rx_dropped on memory-pressure drop.
(Paolo)
- Link to v4: https://lore.kernel.org/r/20260508-lan966x-pci-fdma-v4-0-14e0c89d8d63@microchip.com
Changes in v4:
- Consolidate rx size checks into lan966x_fdma_pci_rx_size_fits().
Subtract XDP_PACKET_HEADROOM on the max size check, and add ETH_HLEN
on the min size check. This fixes potential OOB reads/writes.
- On xdp_prepare_buff(), update comment to clarify that data is already
offset by XDP_PACKET_HEADROOM.
- Link to v3:
https://lore.kernel.org/r/20260504-lan966x-pci-fdma-v3-0-a56f5740d870@microchip.com
Changes in v3:
Version 3 fixes a number of issues reported by sashiko - mostly
hardening.
- Fix double use of XDP_PACKET_HEADROOM.
- Fix ERR_PTR persistence in fdma->atu_region and add missing
NULL/ERR_PTR guard in fdma_pci_atu_region_unmap().
- Reject size <= 0 in fdma_pci_atu_region_map() and return
-ENOSPC (was -ENOMEM) when no region is free.
- Introduce lan966x_fdma_pci_tx_size_fits() that accounts for
XDP_PACKET_HEADROOM; use it from both xmit paths to keep
bpf_xdp_adjust_tail from writing past the TX slot.
- Validate BLOCKL in rx_check_frame() (reject < IFH+FCS or
> db_size) before it feeds memcpy/XDP sizes.
- READ_ONCE(port->xdp_prog) inside lan966x_xdp_pci_run() to close
a TOCTOU on XDP detach that could deref NULL in
bpf_prog_run_xdp().
- Strip IFH and FCS pre-XDP in rx_check_frame(). After BPF runs
the driver cannot tell whether the tail was modified; drop the
unconditional skb_pull/skb_trim in rx_get_frame().
- Account tx_bytes/tx_packets on XDP_TX success and tx_dropped on
XDP_TX size reject.
- Add dma_wmb()/dma_rmb() around DCB status writes and reads in
xmit, xmit_xdpf, and napi_poll.
- Collected Tested-by: Hervé Codina.
- Link to v2: https://lore.kernel.org/r/20260428-lan966x-pci-fdma-v2-0-d3ec66e06202@microchip.com
Changes in v2:
Version 2 primarily addresses issues with module unload/load, where
traffic would stop working (Hervé), and XDP head/tail adjust that would be
discarded (Mohsin).
Apart from that, I ran through issues reported by Sashiko, and fixed a
number of other issues.
- New patch 1: add drivers/net/ethernet/microchip/fdma/ to the Sparx5
SoC MAINTAINERS entry.
- New patch 7: clear latched FDMA error/interrupt stickies after the
switch reset so they don't fire as soon as interrupts are enabled.
- New patch 8: shutdown() callback, quiescing FDMA on host warm reboot.
- Replaced the depth-2 dev_is_pci(parent->parent) backend selector
with a parent-chain walk.
- XDP: use xdp.data/xdp.data_end for the post-XDP frame length so that
bpf_xdp_adjust_head/tail are respected (Mohsin Bashir)
- MTU change: drain in-flight xmits with netif_tx_disable() on every
port before reallocating rings, waking them again on completion.
- MTU change: cap the PCIe DCB ring at 256 entries so a full-ring
coherent DMA allocation fits in a single MAX_PAGE_ORDER block at
jumbo MTU.
- PCIe ATU: disable the region before clearing its translation on
unmap.
- PCIe FDMA: hold tx_lock in napi_poll around the free-DCB check used
to wake stopped netdev queues.
- PCIe FDMA: return -ENOSPC (not -1) when the DCB ring is exhausted.
- Link to v1: https://lore.kernel.org/r/20260320-lan966x-pci-fdma-v1-0-ef54cb9b0c4b@microchip.com
---
Daniel Machon (15):
MAINTAINERS: add FDMA library to Sparx5 SoC entry
net: microchip: fdma: rename contiguous dataptr helpers
net: microchip: fdma: add PCIe ATU support
net: microchip: fdma: use little-endian types for descriptor fields
net: lan966x: add FDMA LLP register write helper
net: lan966x: export FDMA helpers for reuse
net: lan966x: use a dedicated device for DMA operations
net: lan966x: add FDMA ops dispatch for PCIe support
net: lan966x: clear FDMA interrupt stickies after switch reset
net: lan966x: add shutdown callback to stop the FDMA on reboot
net: lan966x: add PCIe FDMA support
net: lan966x: add PCIe FDMA MTU change support
net: lan966x: add PCIe FDMA XDP support
misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
misc: lan966x-pci: dts: add fdma interrupt to overlay
MAINTAINERS | 1 +
drivers/misc/lan966x_pci.dtso | 5 +-
drivers/net/ethernet/microchip/fdma/Makefile | 4 +
drivers/net/ethernet/microchip/fdma/fdma_api.c | 65 +-
drivers/net/ethernet/microchip/fdma/fdma_api.h | 42 +-
drivers/net/ethernet/microchip/fdma/fdma_pci.c | 208 ++++++
drivers/net/ethernet/microchip/fdma/fdma_pci.h | 57 ++
drivers/net/ethernet/microchip/lan966x/Makefile | 4 +
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 97 ++-
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 731 +++++++++++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.c | 121 +++-
.../net/ethernet/microchip/lan966x/lan966x_main.h | 69 ++
.../net/ethernet/microchip/lan966x/lan966x_regs.h | 25 +
.../net/ethernet/microchip/lan966x/lan966x_xdp.c | 6 +
14 files changed, 1357 insertions(+), 78 deletions(-)
---
base-commit: 014d795c73837ea2339a4ea8e8f82c6e959b845d
change-id: 20260313-lan966x-pci-fdma-94ed485d23fa
Best regards,
--
Daniel Machon <daniel.machon@microchip.com>
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 02/15] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
` (13 subsequent siblings)
14 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
The FDMA library under drivers/net/ethernet/microchip/fdma/ is shared by
the lan966x, sparx5 and lan969x drivers, but is not covered by an entry
in the MAINTAINERS file. A subsequent patch will add new files to the
FDMA library, so let's make sure it's covered.
Add drivers/net/ethernet/microchip/fdma/ to the Sparx5 SoC entry, since
I am already listed there.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
MAINTAINERS | 1 +
1 file changed, 1 insertion(+)
diff --git a/MAINTAINERS b/MAINTAINERS
index 6de1ff058db6..740f362e03f8 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -3214,6 +3214,7 @@ M: UNGLinuxDriver@microchip.com
L: linux-arm-kernel@lists.infradead.org (moderated for non-subscribers)
S: Supported
F: arch/arm64/boot/dts/microchip/sparx*
+F: drivers/net/ethernet/microchip/fdma/
F: drivers/net/ethernet/microchip/vcap/
F: drivers/pinctrl/pinctrl-microchip-sgpio.c
N: sparx5
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 02/15] net: microchip: fdma: rename contiguous dataptr helpers
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
` (12 subsequent siblings)
14 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When the FDMA library was introduced [1], two helpers to get the DMA and
virtual address of a data buffer, in contiguous memory, were added. These
helpers have had no callers until this series. I found the naming I
initially used confusing and inconsistent.
Rename fdma_dataptr_get_contiguous() and
fdma_dataptr_virt_get_contiguous() to fdma_dataptr_dma_addr_contiguous()
and fdma_dataptr_virt_addr_contiguous(). This makes the pair symmetric
and clarifies what type of address each returns.
[1]: commit 30e48a75df9c ("net: microchip: add FDMA library")
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
drivers/net/ethernet/microchip/fdma/fdma_api.h | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index d91affe8bd98..dea8e3cc155a 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -197,8 +197,8 @@ static inline int fdma_nextptr_cb(struct fdma *fdma, int dcb_idx, u64 *nextptr)
* if the dataptr addresses and DCB's are in contiguous memory and the driver
* supports XDP.
*/
-static inline u64 fdma_dataptr_get_contiguous(struct fdma *fdma, int dcb_idx,
- int db_idx)
+static inline u64 fdma_dataptr_dma_addr_contiguous(struct fdma *fdma,
+ int dcb_idx, int db_idx)
{
return fdma->dma + (sizeof(struct fdma_dcb) * fdma->n_dcbs) +
(dcb_idx * fdma->n_dbs + db_idx) * fdma->db_size +
@@ -209,8 +209,8 @@ static inline u64 fdma_dataptr_get_contiguous(struct fdma *fdma, int dcb_idx,
* applicable if the dataptr addresses and DCB's are in contiguous memory and
* the driver supports XDP.
*/
-static inline void *fdma_dataptr_virt_get_contiguous(struct fdma *fdma,
- int dcb_idx, int db_idx)
+static inline void *fdma_dataptr_virt_addr_contiguous(struct fdma *fdma,
+ int dcb_idx, int db_idx)
{
return (u8 *)fdma->dcbs + (sizeof(struct fdma_dcb) * fdma->n_dcbs) +
(dcb_idx * fdma->n_dbs + db_idx) * fdma->db_size +
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 03/15] net: microchip: fdma: add PCIe ATU support
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 02/15] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:32 ` [PATCH net-next v9 04/15] net: microchip: fdma: use little-endian types for descriptor fields Daniel Machon
` (11 subsequent siblings)
14 siblings, 1 reply; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When lan966x or lan969x operates as a PCIe endpoint, the internal FDMA
engine cannot directly access host memory. Instead, DMA addresses must
be translated through the PCIe Address Translation Unit (ATU). The ATU
provides outbound windows that map internal addresses to PCIe bus
addresses.
The ATU outbound address space (0x10000000-0x1fffffff) is divided into
six equally-sized regions (~42MB each). When FDMA buffers are allocated,
a free ATU region is claimed and programmed with the DMA target address.
The FDMA engine then uses the region's base address in its descriptors,
and the ATU translates these to the actual DMA addresses on the PCIe bus.
Add the required functions and helpers that combine the DMA allocation
with the ATU region mapping. These are used by the lan966x PCIe FDMA
path.
The ATU cannot express a limit finer than its 64KB region granularity,
so pad the mapped allocation to that boundary; otherwise the outbound
window would extend past the memory the host allocated for DMA.
This implementation will also be used by the lan969x, when PCIe FDMA is
added for that platform in the future.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
drivers/net/ethernet/microchip/fdma/Makefile | 4 +
drivers/net/ethernet/microchip/fdma/fdma_api.c | 44 ++++++
drivers/net/ethernet/microchip/fdma/fdma_api.h | 14 ++
drivers/net/ethernet/microchip/fdma/fdma_pci.c | 208 +++++++++++++++++++++++++
drivers/net/ethernet/microchip/fdma/fdma_pci.h | 57 +++++++
5 files changed, 327 insertions(+)
diff --git a/drivers/net/ethernet/microchip/fdma/Makefile b/drivers/net/ethernet/microchip/fdma/Makefile
index cc9a736be357..910a10b33fb9 100644
--- a/drivers/net/ethernet/microchip/fdma/Makefile
+++ b/drivers/net/ethernet/microchip/fdma/Makefile
@@ -5,3 +5,7 @@
obj-$(CONFIG_FDMA) += fdma.o
fdma-y += fdma_api.o
+
+ifneq ($(CONFIG_MCHP_LAN966X_PCI),)
+fdma-y += fdma_pci.o
+endif
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.c b/drivers/net/ethernet/microchip/fdma/fdma_api.c
index e78c3590da9e..a3c9e3097c5c 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.c
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.c
@@ -127,6 +127,50 @@ void fdma_free_phys(struct fdma *fdma)
}
EXPORT_SYMBOL_GPL(fdma_free_phys);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+/* Allocate coherent DMA memory and map it in the ATU. */
+int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
+ struct fdma_pci_atu *atu)
+{
+ struct fdma_pci_atu_region *region;
+ int err;
+
+ if (WARN_ON(fdma->atu_region))
+ return -EBUSY;
+
+ /* The ATU cannot express a limit finer than the region granularity, so
+ * the hardware widens the programmed limit to that boundary. Pad the
+ * allocation to match, or the outbound window would extend past the
+ * memory we own.
+ */
+ fdma->size = ALIGN(fdma->size, FDMA_PCI_ATU_REGION_ALIGN);
+
+ err = fdma_alloc_coherent(dev, fdma);
+ if (err)
+ return err;
+
+ region = fdma_pci_atu_region_map(atu, fdma->dma, fdma->size);
+ if (IS_ERR(region)) {
+ fdma_free_coherent(dev, fdma);
+ return PTR_ERR(region);
+ }
+
+ fdma->atu_region = region;
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(fdma_alloc_coherent_and_map);
+
+/* Free coherent DMA memory and unmap the memory in the ATU. */
+void fdma_free_coherent_and_unmap(struct device *dev, struct fdma *fdma)
+{
+ fdma_pci_atu_region_unmap(fdma->atu_region);
+ fdma->atu_region = NULL;
+ fdma_free_coherent(dev, fdma);
+}
+EXPORT_SYMBOL_GPL(fdma_free_coherent_and_unmap);
+#endif
+
/* Get the size of the FDMA memory */
u32 fdma_get_size(struct fdma *fdma)
{
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index dea8e3cc155a..ccc30d506e89 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -7,6 +7,8 @@
#include <linux/etherdevice.h>
#include <linux/types.h>
+#include "fdma_pci.h"
+
/* This provides a common set of functions and data structures for interacting
* with the Frame DMA engine on multiple Microchip switchcores.
*
@@ -109,6 +111,11 @@ struct fdma {
u32 channel_id;
struct fdma_ops ops;
+
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+ /* PCI ATU region for this FDMA instance. */
+ struct fdma_pci_atu_region *atu_region;
+#endif
};
/* Advance the DCB index and wrap if required. */
@@ -233,9 +240,16 @@ int __fdma_dcb_add(struct fdma *fdma, int dcb_idx, u64 info, u64 status,
int fdma_alloc_coherent(struct device *dev, struct fdma *fdma);
int fdma_alloc_phys(struct fdma *fdma);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
+ struct fdma_pci_atu *atu);
+#endif
void fdma_free_coherent(struct device *dev, struct fdma *fdma);
void fdma_free_phys(struct fdma *fdma);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+void fdma_free_coherent_and_unmap(struct device *dev, struct fdma *fdma);
+#endif
u32 fdma_get_size(struct fdma *fdma);
u32 fdma_get_size_contiguous(struct fdma *fdma);
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.c b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
new file mode 100644
index 000000000000..dd1dc46cbc9d
--- /dev/null
+++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
@@ -0,0 +1,208 @@
+// SPDX-License-Identifier: GPL-2.0+
+
+#include <linux/align.h>
+#include <linux/errno.h>
+#include <linux/io.h>
+#include <linux/kernel.h>
+#include <linux/types.h>
+
+#include "fdma_pci.h"
+
+/* When the switch operates as a PCIe endpoint, the FDMA engine needs to
+ * DMA to/from host memory. The FDMA writes to addresses within the endpoint's
+ * internal Outbound (OB) address space, and the PCIe ATU translates these to
+ * DMA addresses on the PCIe bus, targeting host memory.
+ *
+ * The ATU supports up to six outbound regions. This implementation divides
+ * the OB address space into six equally sized chunks.
+ *
+ * +-------------+------------+------------+-----+------------+
+ * | Index | Region 0 | Region 1 | ... | Region 5 |
+ * +-------------+------------+------------+-----+------------+
+ * | Base addr | 0x10000000 | 0x12aa0000 | ... | 0x1d520000 |
+ * | Limit addr | 0x12a9ffff | 0x1553ffff | ... | 0x1ffbffff |
+ * | Target addr | host dma | host dma | ... | host dma |
+ * +-------------+------------+------------+-----+------------+
+ *
+ * Base addr is the start address of the region within the OB address space.
+ * Limit addr is each region's own upper bound. The value actually
+ * programmed is set per-mapping as base_addr + mapped size - 1, and is
+ * usually smaller than this.
+ * Target addr is the host DMA address that the base addr translates to.
+ */
+
+#define FDMA_PCI_ATU_OB_START 0x10000000
+#define FDMA_PCI_ATU_OB_END 0x1fffffff
+
+#define FDMA_PCI_ATU_ADDR 0x300000
+#define FDMA_PCI_ATU_IDX_SIZE 0x200
+#define FDMA_PCI_ATU_ENA_REG 0x4
+#define FDMA_PCI_ATU_ENA_BIT BIT(31)
+#define FDMA_PCI_ATU_LWR_BASE_ADDR 0x8
+#define FDMA_PCI_ATU_UPP_BASE_ADDR 0xc
+#define FDMA_PCI_ATU_LIMIT_ADDR 0x10
+#define FDMA_PCI_ATU_LWR_TARGET_ADDR 0x14
+#define FDMA_PCI_ATU_UPP_TARGET_ADDR 0x18
+
+static u32 fdma_pci_atu_region_size(void)
+{
+ return round_down((FDMA_PCI_ATU_OB_END - FDMA_PCI_ATU_OB_START) /
+ FDMA_PCI_ATU_REGION_MAX, FDMA_PCI_ATU_REGION_ALIGN);
+}
+
+static void __iomem *fdma_pci_atu_addr_get(void __iomem *addr, int offset,
+ int idx)
+{
+ return addr + FDMA_PCI_ATU_ADDR + FDMA_PCI_ATU_IDX_SIZE * idx + offset;
+}
+
+static void fdma_pci_atu_region_enable(struct fdma_pci_atu_region *region)
+{
+ writel(FDMA_PCI_ATU_ENA_BIT,
+ fdma_pci_atu_addr_get(region->atu->addr, FDMA_PCI_ATU_ENA_REG,
+ region->idx));
+}
+
+static void fdma_pci_atu_region_disable(struct fdma_pci_atu_region *region)
+{
+ writel(0, fdma_pci_atu_addr_get(region->atu->addr, FDMA_PCI_ATU_ENA_REG,
+ region->idx));
+}
+
+/* Configure the address translation in the ATU. */
+static void
+fdma_pci_atu_configure_translation(struct fdma_pci_atu_region *region)
+{
+ struct fdma_pci_atu *atu = region->atu;
+ int idx = region->idx;
+
+ writel(lower_32_bits(region->base_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_LWR_BASE_ADDR, idx));
+
+ writel(upper_32_bits(region->base_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_UPP_BASE_ADDR, idx));
+
+ /* The OB address space lies entirely below 4GB, so the limit always
+ * fits the lower limit register and the upper one is left alone.
+ */
+ writel(region->limit_addr,
+ fdma_pci_atu_addr_get(atu->addr, FDMA_PCI_ATU_LIMIT_ADDR, idx));
+
+ writel(lower_32_bits(region->target_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_LWR_TARGET_ADDR, idx));
+
+ writel(upper_32_bits(region->target_addr),
+ fdma_pci_atu_addr_get(atu->addr,
+ FDMA_PCI_ATU_UPP_TARGET_ADDR, idx));
+}
+
+/* Find an unused ATU region. */
+static struct fdma_pci_atu_region *
+fdma_pci_atu_region_get_free(struct fdma_pci_atu *atu)
+{
+ struct fdma_pci_atu_region *regions = atu->regions;
+
+ for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
+ if (regions[i].in_use)
+ continue;
+
+ return ®ions[i];
+ }
+
+ return ERR_PTR(-ENOSPC);
+}
+
+/* Unmap an ATU region, clearing its translation and disabling it. */
+void fdma_pci_atu_region_unmap(struct fdma_pci_atu_region *region)
+{
+ if (IS_ERR_OR_NULL(region))
+ return;
+
+ mutex_lock(®ion->atu->lock);
+
+ region->target_addr = 0;
+ region->in_use = false;
+
+ fdma_pci_atu_region_disable(region);
+ fdma_pci_atu_configure_translation(region);
+
+ mutex_unlock(®ion->atu->lock);
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_region_unmap);
+
+/* Map a host DMA address into a free outbound region. */
+struct fdma_pci_atu_region *
+fdma_pci_atu_region_map(struct fdma_pci_atu *atu, u64 target_addr, int size)
+{
+ struct fdma_pci_atu_region *region;
+
+ if (!atu)
+ return ERR_PTR(-EINVAL);
+
+ if (size <= 0)
+ return ERR_PTR(-EINVAL);
+
+ if (size > fdma_pci_atu_region_size())
+ return ERR_PTR(-ERANGE);
+
+ /* The ATU region base is only ever aligned to FDMA_PCI_ATU_REGION_ALIGN;
+ * require the same alignment of the host target address, since the ATU
+ * translates addr - target_addr + base_addr and any misalignment here
+ * would shift every translated address by the same amount.
+ */
+ if (!IS_ALIGNED(target_addr, FDMA_PCI_ATU_REGION_ALIGN))
+ return ERR_PTR(-EINVAL);
+
+ mutex_lock(&atu->lock);
+
+ region = fdma_pci_atu_region_get_free(atu);
+ if (IS_ERR(region)) {
+ mutex_unlock(&atu->lock);
+ return region;
+ }
+
+ region->target_addr = target_addr;
+ region->limit_addr = region->base_addr + size - 1;
+ region->in_use = true;
+
+ fdma_pci_atu_configure_translation(region);
+ fdma_pci_atu_region_enable(region);
+
+ mutex_unlock(&atu->lock);
+
+ return region;
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_region_map);
+
+/* Translate a host DMA address to the corresponding OB address. */
+u64 fdma_pci_atu_translate_addr(struct fdma_pci_atu_region *region, u64 addr)
+{
+ return region->base_addr + (addr - region->target_addr);
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_translate_addr);
+
+/* Initialize ATU, dividing the OB space into equally sized regions. */
+void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr)
+{
+ struct fdma_pci_atu_region *regions = atu->regions;
+ u32 region_size = fdma_pci_atu_region_size();
+
+ atu->addr = addr;
+ mutex_init(&atu->lock);
+
+ for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
+ regions[i].base_addr =
+ FDMA_PCI_ATU_OB_START + (i * region_size);
+ regions[i].limit_addr = 0;
+ regions[i].target_addr = 0;
+ regions[i].idx = i;
+ regions[i].atu = atu;
+ regions[i].in_use = false;
+
+ fdma_pci_atu_region_disable(®ions[i]);
+ }
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_init);
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.h b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
new file mode 100644
index 000000000000..010bdb4d50b0
--- /dev/null
+++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
@@ -0,0 +1,57 @@
+/* SPDX-License-Identifier: GPL-2.0+ */
+
+#ifndef _FDMA_PCI_H_
+#define _FDMA_PCI_H_
+
+#include <linux/align.h>
+#include <linux/bits.h>
+#include <linux/mutex.h>
+#include <linux/types.h>
+
+#define FDMA_PCI_ATU_REGION_MAX 6
+
+/* Outbound regions are 64KB granular (datasheet section 3.24.7.4.1), so both
+ * the region base and the mapped size must be aligned to this.
+ */
+#define FDMA_PCI_ATU_REGION_ALIGN BIT(16)
+
+#define FDMA_PCI_DB_ALIGN 128
+#define FDMA_PCI_DB_SIZE(mtu) ALIGN(mtu, FDMA_PCI_DB_ALIGN)
+
+struct fdma_pci_atu;
+
+struct fdma_pci_atu_region {
+ struct fdma_pci_atu *atu;
+ u64 base_addr; /* Base addr of the OB window */
+ u64 limit_addr; /* End addr of the active mapping (base_addr + size - 1) */
+ u64 target_addr; /* Host DMA address this region maps to */
+ int idx;
+ bool in_use;
+};
+
+struct fdma_pci_atu {
+ void __iomem *addr;
+ struct mutex lock; /* Protects region alloc/free and ATU register access */
+ struct fdma_pci_atu_region regions[FDMA_PCI_ATU_REGION_MAX];
+};
+
+/* Initialize ATU, dividing OB space into regions. */
+void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr);
+
+/* Unmap an ATU region, clearing its translation and disabling it. */
+void fdma_pci_atu_region_unmap(struct fdma_pci_atu_region *region);
+
+/* Map a host DMA address into a free ATU region. target_addr and size must be
+ * FDMA_PCI_ATU_REGION_ALIGN aligned; a misaligned target_addr returns -EINVAL.
+ */
+struct fdma_pci_atu_region *fdma_pci_atu_region_map(struct fdma_pci_atu *atu,
+ u64 target_addr,
+ int size);
+
+/* Translate a host DMA address to the OB address space. Reads the region
+ * unlocked, so the caller must quiesce DMA and the descriptor paths before
+ * unmapping the region.
+ */
+u64 fdma_pci_atu_translate_addr(struct fdma_pci_atu_region *region, u64 addr);
+
+#endif
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 04/15] net: microchip: fdma: use little-endian types for descriptor fields
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (2 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-10-02 13:25 ` Simon Horman
2026-09-28 19:32 ` [PATCH net-next v9 05/15] net: lan966x: add FDMA LLP register write helper Daniel Machon
` (10 subsequent siblings)
14 siblings, 1 reply; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
The FDMA engine reads and writes the DCB and DB descriptors in
little-endian byte order. So far the descriptors have only been produced
by the little-endian SoC itself, so plain u64 fields were fine. With the
PCIe FDMA path, the descriptors live in host memory and are written by
the host CPU, which may be big-endian.
Change the descriptor fields to __le64 and convert at the library
boundary: in __fdma_db_add() and __fdma_dcb_add() on write, and in the
fdma_db_*() accessors on read. The dataptr and nextptr callbacks keep
their u64 signatures, so their implementations are unchanged. Add
fdma_db_dataptr_get(), and convert the lan966x sites that read the
descriptor fields directly to use the accessors.
No functional change on little-endian hosts.
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/net/ethernet/microchip/fdma/fdma_api.c | 21 ++++++++++++++++-----
drivers/net/ethernet/microchip/fdma/fdma_api.h | 20 +++++++++++++-------
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 9 +++++----
3 files changed, 34 insertions(+), 16 deletions(-)
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.c b/drivers/net/ethernet/microchip/fdma/fdma_api.c
index a3c9e3097c5c..f7a42348e932 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.c
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.c
@@ -12,10 +12,18 @@ static int __fdma_db_add(struct fdma *fdma, int dcb_idx, int db_idx, u64 status,
int db_idx, u64 *dataptr))
{
struct fdma_db *db = fdma_db_get(fdma, dcb_idx, db_idx);
+ u64 dataptr;
+ int err;
+
+ db->status = cpu_to_le64(status);
- db->status = status;
+ err = cb(fdma, dcb_idx, db_idx, &dataptr);
+ if (unlikely(err))
+ return err;
- return cb(fdma, dcb_idx, db_idx, &db->dataptr);
+ db->dataptr = cpu_to_le64(dataptr);
+
+ return 0;
}
/* Add a DB to a DCB, using the callback set in the fdma_ops struct. */
@@ -35,6 +43,7 @@ int __fdma_dcb_add(struct fdma *fdma, int dcb_idx, u64 info, u64 status,
u64 *dataptr))
{
struct fdma_dcb *dcb = fdma_dcb_get(fdma, dcb_idx);
+ u64 nextptr;
int i, err;
for (i = 0; i < fdma->n_dbs; i++) {
@@ -43,14 +52,16 @@ int __fdma_dcb_add(struct fdma *fdma, int dcb_idx, u64 info, u64 status,
return err;
}
- err = dcb_cb(fdma, dcb_idx, &fdma->last_dcb->nextptr);
+ err = dcb_cb(fdma, dcb_idx, &nextptr);
if (unlikely(err))
return err;
+ fdma->last_dcb->nextptr = cpu_to_le64(nextptr);
+
fdma->last_dcb = dcb;
- dcb->nextptr = FDMA_DCB_INVALID_DATA;
- dcb->info = info;
+ dcb->nextptr = cpu_to_le64(FDMA_DCB_INVALID_DATA);
+ dcb->info = cpu_to_le64(info);
return 0;
}
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index ccc30d506e89..4e4f009b77cb 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -66,13 +66,13 @@
struct fdma;
struct fdma_db {
- u64 dataptr;
- u64 status;
+ __le64 dataptr;
+ __le64 status;
};
struct fdma_dcb {
- u64 nextptr;
- u64 info;
+ __le64 nextptr;
+ __le64 info;
struct fdma_db db[FDMA_DB_MAX];
};
@@ -147,19 +147,25 @@ static inline bool fdma_dcb_is_reusable(struct fdma *fdma)
/* Check if the FDMA has marked this DB as done. */
static inline bool fdma_db_is_done(struct fdma_db *db)
{
- return db->status & FDMA_DCB_STATUS_DONE;
+ return le64_to_cpu(db->status) & FDMA_DCB_STATUS_DONE;
}
/* Get the length of a DB. */
static inline int fdma_db_len_get(struct fdma_db *db)
{
- return FDMA_DCB_STATUS_BLOCKL(db->status);
+ return FDMA_DCB_STATUS_BLOCKL(le64_to_cpu(db->status));
+}
+
+/* Get the dataptr of a DB. */
+static inline u64 fdma_db_dataptr_get(struct fdma_db *db)
+{
+ return le64_to_cpu(db->dataptr);
}
/* Set the length of a DB. */
static inline void fdma_dcb_len_set(struct fdma_dcb *dcb, u32 len)
{
- dcb->info = FDMA_DCB_INFO_DATAL(len);
+ dcb->info = cpu_to_le64(FDMA_DCB_INFO_DATAL(len));
}
/* Get a DB by index. */
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 41d4ec7f2f57..68fd454ebc98 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -406,8 +406,9 @@ static int lan966x_fdma_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
return FDMA_ERROR;
dma_sync_single_for_cpu(lan966x->dev,
- (dma_addr_t)db->dataptr + XDP_PACKET_HEADROOM,
- FDMA_DCB_STATUS_BLOCKL(db->status),
+ (dma_addr_t)fdma_db_dataptr_get(db) +
+ XDP_PACKET_HEADROOM,
+ fdma_db_len_get(db),
DMA_FROM_DEVICE);
lan966x_ifh_get_src_port(page_address(page) + XDP_PACKET_HEADROOM,
@@ -419,7 +420,7 @@ static int lan966x_fdma_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
if (!lan966x_xdp_port_present(port))
return FDMA_PASS;
- return lan966x_xdp_run(port, page, FDMA_DCB_STATUS_BLOCKL(db->status));
+ return lan966x_xdp_run(port, page, fdma_db_len_get(db));
}
static struct sk_buff *lan966x_fdma_rx_get_frame(struct lan966x_rx *rx,
@@ -443,7 +444,7 @@ static struct sk_buff *lan966x_fdma_rx_get_frame(struct lan966x_rx *rx,
skb_mark_for_recycle(skb);
skb_reserve(skb, XDP_PACKET_HEADROOM);
- skb_put(skb, FDMA_DCB_STATUS_BLOCKL(db->status));
+ skb_put(skb, fdma_db_len_get(db));
lan966x_ifh_get_timestamp(skb->data, ×tamp);
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 05/15] net: lan966x: add FDMA LLP register write helper
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (3 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 04/15] net: microchip: fdma: use little-endian types for descriptor fields Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 06/15] net: lan966x: export FDMA helpers for reuse Daniel Machon
` (9 subsequent siblings)
14 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
The FDMA Link List Pointer (LLP) register points to the first DCB in the
chain and must be written before the channel is activated. This tells
the FDMA engine where to begin DMA transfers.
Move the LLP register writes from the channel start/activate functions
into the allocation functions and introduce a shared
lan966x_fdma_llp_configure() helper. This is needed because the upcoming
PCIe FDMA path writes ATU-translated addresses to the LLP registers
instead of DMA addresses. Keeping the writes in the shared
start/activate path would overwrite these translated addresses.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 30 ++++++++++------------
1 file changed, 14 insertions(+), 16 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 68fd454ebc98..1c5484a506ff 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -109,6 +109,13 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
return 0;
}
+static void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+ u8 channel_id)
+{
+ lan_wr(lower_32_bits(addr), lan966x, FDMA_DCB_LLP(channel_id));
+ lan_wr(upper_32_bits(addr), lan966x, FDMA_DCB_LLP1(channel_id));
+}
+
static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
@@ -128,6 +135,8 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
fdma_dcbs_init(fdma, FDMA_DCB_INFO_DATAL(fdma->db_size),
FDMA_DCB_STATUS_INTR);
+ lan966x_fdma_llp_configure(lan966x, fdma->dma, fdma->channel_id);
+
return 0;
}
@@ -137,14 +146,6 @@ static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
struct fdma *fdma = &rx->fdma;
u32 mask;
- /* When activating a channel, first is required to write the first DCB
- * address and then to activate it
- */
- lan_wr(lower_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP(fdma->channel_id));
- lan_wr(upper_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP1(fdma->channel_id));
-
lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
@@ -215,6 +216,8 @@ static int lan966x_fdma_tx_alloc(struct lan966x_tx *tx)
fdma_dcbs_init(fdma, 0, 0);
+ lan966x_fdma_llp_configure(lan966x, fdma->dma, fdma->channel_id);
+
return 0;
out:
@@ -236,14 +239,6 @@ static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
struct fdma *fdma = &tx->fdma;
u32 mask;
- /* When activating a channel, first is required to write the first DCB
- * address and then to activate it
- */
- lan_wr(lower_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP(fdma->channel_id));
- lan_wr(upper_32_bits((u64)fdma->dma), lan966x,
- FDMA_DCB_LLP1(fdma->channel_id));
-
lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
@@ -877,6 +872,9 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
MEM_TYPE_PAGE_POOL, page_pool);
}
+ lan966x_fdma_llp_configure(lan966x, lan966x->rx.fdma.dma,
+ lan966x->rx.fdma.channel_id);
+
lan966x_fdma_rx_start(&lan966x->rx);
lan966x_fdma_wakeup_netdev(lan966x);
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 06/15] net: lan966x: export FDMA helpers for reuse
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (4 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 05/15] net: lan966x: add FDMA LLP register write helper Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 07/15] net: lan966x: use a dedicated device for DMA operations Daniel Machon
` (8 subsequent siblings)
14 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Make shared FDMA helpers non-static, so they can be reused by the PCIe
FDMA implementation.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 22 +++++++++++-----------
.../net/ethernet/microchip/lan966x/lan966x_main.h | 11 +++++++++++
2 files changed, 22 insertions(+), 11 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 1c5484a506ff..651ea03eede7 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -109,8 +109,8 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
return 0;
}
-static void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
- u8 channel_id)
+void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+ u8 channel_id)
{
lan_wr(lower_32_bits(addr), lan966x, FDMA_DCB_LLP(channel_id));
lan_wr(upper_32_bits(addr), lan966x, FDMA_DCB_LLP1(channel_id));
@@ -140,7 +140,7 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
return 0;
}
-static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
+void lan966x_fdma_rx_start(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
struct fdma *fdma = &rx->fdma;
@@ -171,7 +171,7 @@ static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
lan966x, FDMA_CH_ACTIVATE);
}
-static void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
+void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
struct fdma *fdma = &rx->fdma;
@@ -191,7 +191,7 @@ static void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
lan966x, FDMA_CH_DB_DISCARD);
}
-static void lan966x_fdma_rx_reload(struct lan966x_rx *rx)
+void lan966x_fdma_rx_reload(struct lan966x_rx *rx)
{
struct lan966x *lan966x = rx->lan966x;
@@ -264,7 +264,7 @@ static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
lan966x, FDMA_CH_ACTIVATE);
}
-static void lan966x_fdma_tx_disable(struct lan966x_tx *tx)
+void lan966x_fdma_tx_disable(struct lan966x_tx *tx)
{
struct lan966x *lan966x = tx->lan966x;
struct fdma *fdma = &tx->fdma;
@@ -296,7 +296,7 @@ static void lan966x_fdma_tx_reload(struct lan966x_tx *tx)
lan966x, FDMA_CH_RELOAD);
}
-static void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x)
+void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x)
{
struct lan966x_port *port;
int i;
@@ -471,7 +471,7 @@ static struct sk_buff *lan966x_fdma_rx_get_frame(struct lan966x_rx *rx,
return NULL;
}
-static int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight)
+int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight)
{
struct lan966x *lan966x = container_of(napi, struct lan966x, napi);
struct lan966x_rx *rx = &lan966x->rx;
@@ -584,7 +584,7 @@ static int lan966x_fdma_get_next_dcb(struct lan966x_tx *tx)
return -1;
}
-static void lan966x_fdma_tx_start(struct lan966x_tx *tx)
+void lan966x_fdma_tx_start(struct lan966x_tx *tx)
{
struct lan966x *lan966x = tx->lan966x;
@@ -802,7 +802,7 @@ static int lan966x_fdma_get_max_mtu(struct lan966x *lan966x)
return max_mtu;
}
-static int lan966x_qsys_sw_status(struct lan966x *lan966x)
+int lan966x_qsys_sw_status(struct lan966x *lan966x)
{
return lan_rd(lan966x, QSYS_SW_STATUS(CPU_PORT));
}
@@ -884,7 +884,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
return err;
}
-static int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
+int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
{
return lan966x_fdma_get_max_mtu(lan966x) +
IFH_LEN_BYTES +
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index eea286c29474..83c361abb789 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -561,6 +561,17 @@ int lan966x_fdma_init(struct lan966x *lan966x);
void lan966x_fdma_deinit(struct lan966x *lan966x);
irqreturn_t lan966x_fdma_irq_handler(int irq, void *args);
int lan966x_fdma_reload_page_pool(struct lan966x *lan966x);
+int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight);
+void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+ u8 channel_id);
+void lan966x_fdma_rx_start(struct lan966x_rx *rx);
+void lan966x_fdma_rx_disable(struct lan966x_rx *rx);
+void lan966x_fdma_rx_reload(struct lan966x_rx *rx);
+void lan966x_fdma_tx_start(struct lan966x_tx *tx);
+void lan966x_fdma_tx_disable(struct lan966x_tx *tx);
+void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
+int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
+int lan966x_qsys_sw_status(struct lan966x *lan966x);
int lan966x_lag_port_join(struct lan966x_port *port,
struct net_device *brport_dev,
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 07/15] net: lan966x: use a dedicated device for DMA operations
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (5 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 06/15] net: lan966x: export FDMA helpers for reuse Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 08/15] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
` (7 subsequent siblings)
14 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
In preparation for the PCIe FDMA implementation, add lan966x->dma_dev,
resolved once at probe time, and use it for every DMA operation the
FDMA path performs: coherent allocations, streaming mappings and the RX
page pool. Natively, dma_dev is lan966x->dev, so there is no functional
change.
Since dma_dev differs from dev only on the PCIe path, add
lan966x_is_pci() on top of it, for the later PCIe patches to key off.
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 30 +++++++++++-----------
.../net/ethernet/microchip/lan966x/lan966x_main.c | 16 ++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.h | 10 ++++++++
3 files changed, 41 insertions(+), 15 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 651ea03eede7..ae6e5aec5e13 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -80,7 +80,7 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
.flags = PP_FLAG_DMA_MAP | PP_FLAG_DMA_SYNC_DEV,
.pool_size = rx->fdma.n_dcbs,
.nid = NUMA_NO_NODE,
- .dev = lan966x->dev,
+ .dev = lan966x->dma_dev,
.dma_dir = DMA_FROM_DEVICE,
.offset = XDP_PACKET_HEADROOM,
.max_len = rx->max_mtu -
@@ -126,7 +126,7 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
if (err)
return err;
- err = fdma_alloc_coherent(lan966x->dev, fdma);
+ err = fdma_alloc_coherent(lan966x->dma_dev, fdma);
if (err) {
page_pool_destroy(rx->page_pool);
return err;
@@ -210,7 +210,7 @@ static int lan966x_fdma_tx_alloc(struct lan966x_tx *tx)
if (!tx->dcbs_buf)
return -ENOMEM;
- err = fdma_alloc_coherent(lan966x->dev, fdma);
+ err = fdma_alloc_coherent(lan966x->dma_dev, fdma);
if (err)
goto out;
@@ -230,7 +230,7 @@ static void lan966x_fdma_tx_free(struct lan966x_tx *tx)
struct lan966x *lan966x = tx->lan966x;
kfree(tx->dcbs_buf);
- fdma_free_coherent(lan966x->dev, &tx->fdma);
+ fdma_free_coherent(lan966x->dma_dev, &tx->fdma);
}
static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
@@ -355,7 +355,7 @@ static void lan966x_fdma_tx_clear_buf(struct lan966x *lan966x, int weight)
dcb_buf->used = false;
if (dcb_buf->use_skb) {
- dma_unmap_single(lan966x->dev,
+ dma_unmap_single(lan966x->dma_dev,
dcb_buf->dma_addr,
dcb_buf->len,
DMA_TO_DEVICE);
@@ -364,7 +364,7 @@ static void lan966x_fdma_tx_clear_buf(struct lan966x *lan966x, int weight)
napi_consume_skb(dcb_buf->data.skb, weight);
} else {
if (dcb_buf->xdp_ndo)
- dma_unmap_single(lan966x->dev,
+ dma_unmap_single(lan966x->dma_dev,
dcb_buf->dma_addr,
dcb_buf->len,
DMA_TO_DEVICE);
@@ -400,7 +400,7 @@ static int lan966x_fdma_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
if (unlikely(!page))
return FDMA_ERROR;
- dma_sync_single_for_cpu(lan966x->dev,
+ dma_sync_single_for_cpu(lan966x->dma_dev,
(dma_addr_t)fdma_db_dataptr_get(db) +
XDP_PACKET_HEADROOM,
fdma_db_len_get(db),
@@ -636,11 +636,11 @@ int lan966x_fdma_xmit_xdpf(struct lan966x_port *port, void *ptr, u32 len)
lan966x_ifh_set_bypass(ifh, 1);
lan966x_ifh_set_port(ifh, BIT_ULL(port->chip_port));
- dma_addr = dma_map_single(lan966x->dev,
+ dma_addr = dma_map_single(lan966x->dma_dev,
xdpf->data - IFH_LEN_BYTES,
xdpf->len + IFH_LEN_BYTES,
DMA_TO_DEVICE);
- if (dma_mapping_error(lan966x->dev, dma_addr)) {
+ if (dma_mapping_error(lan966x->dma_dev, dma_addr)) {
ret = NETDEV_TX_OK;
goto out;
}
@@ -656,7 +656,7 @@ int lan966x_fdma_xmit_xdpf(struct lan966x_port *port, void *ptr, u32 len)
lan966x_ifh_set_port(ifh, BIT_ULL(port->chip_port));
dma_addr = page_pool_get_dma_addr(page);
- dma_sync_single_for_device(lan966x->dev,
+ dma_sync_single_for_device(lan966x->dma_dev,
dma_addr + XDP_PACKET_HEADROOM,
len + IFH_LEN_BYTES,
DMA_TO_DEVICE);
@@ -735,9 +735,9 @@ int lan966x_fdma_xmit(struct sk_buff *skb, __be32 *ifh, struct net_device *dev)
memcpy(skb->data, ifh, IFH_LEN_BYTES);
skb_put(skb, 4);
- dma_addr = dma_map_single(lan966x->dev, skb->data, skb->len,
+ dma_addr = dma_map_single(lan966x->dma_dev, skb->data, skb->len,
DMA_TO_DEVICE);
- if (dma_mapping_error(lan966x->dev, dma_addr)) {
+ if (dma_mapping_error(lan966x->dma_dev, dma_addr)) {
dev->stats.tx_dropped++;
err = NETDEV_TX_OK;
goto release;
@@ -843,7 +843,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
page_pool_put_full_page(page_pool,
old_pages[i][j], false);
- fdma_free_coherent(lan966x->dev, &fdma_rx_old);
+ fdma_free_coherent(lan966x->dma_dev, &fdma_rx_old);
page_pool_destroy(page_pool);
@@ -993,7 +993,7 @@ int lan966x_fdma_init(struct lan966x *lan966x)
err = lan966x_fdma_tx_alloc(&lan966x->tx);
if (err) {
- fdma_free_coherent(lan966x->dev, &lan966x->rx.fdma);
+ fdma_free_coherent(lan966x->dma_dev, &lan966x->rx.fdma);
page_pool_destroy(lan966x->rx.page_pool);
return err;
}
@@ -1015,7 +1015,7 @@ void lan966x_fdma_deinit(struct lan966x *lan966x)
napi_disable(&lan966x->napi);
lan966x_fdma_rx_free_pages(&lan966x->rx);
- fdma_free_coherent(lan966x->dev, &lan966x->rx.fdma);
+ fdma_free_coherent(lan966x->dma_dev, &lan966x->rx.fdma);
page_pool_destroy(lan966x->rx.page_pool);
lan966x_fdma_tx_free(&lan966x->tx);
}
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 1179a6e127c5..2741f7c9fa4c 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -7,6 +7,7 @@
#include <linux/ip.h>
#include <linux/of.h>
#include <linux/of_net.h>
+#include <linux/pci.h>
#include <linux/phy/phy.h>
#include <linux/platform_device.h>
#include <linux/reset.h>
@@ -1081,6 +1082,20 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
return 0;
}
+/* When enumerated over PCIe, dev is a platform device with no
+ * iommus/dma-ranges of its own, so DMA must target the PCIe endpoint instead.
+ * The result differs from dev only in that case; lan966x_is_pci() relies on it.
+ */
+static struct device *lan966x_get_dma_dev(struct device *dev)
+{
+ for (struct device *p = dev->parent; p; p = p->parent) {
+ if (dev_is_pci(p))
+ return p;
+ }
+
+ return dev;
+}
+
static int lan966x_probe(struct platform_device *pdev)
{
struct fwnode_handle *ports, *portnp;
@@ -1094,6 +1109,7 @@ static int lan966x_probe(struct platform_device *pdev)
platform_set_drvdata(pdev, lan966x);
lan966x->dev = &pdev->dev;
+ lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
if (!device_get_mac_address(&pdev->dev, mac_addr)) {
ether_addr_copy(lan966x->base_mac, mac_addr);
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 83c361abb789..3f09f8ddf620 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -270,6 +270,9 @@ struct lan966x_skb_cb {
struct lan966x {
struct device *dev;
+ /* Device used for DMA; the PCIe endpoint when enumerated over PCIe. */
+ struct device *dma_dev;
+
u8 num_phys_ports;
struct lan966x_port **ports;
@@ -573,6 +576,13 @@ void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
int lan966x_qsys_sw_status(struct lan966x *lan966x);
+/* dma_dev differs from dev only on the PCIe path. */
+static inline bool lan966x_is_pci(struct lan966x *lan966x)
+{
+ return IS_ENABLED(CONFIG_MCHP_LAN966X_PCI) &&
+ lan966x->dma_dev != lan966x->dev;
+}
+
int lan966x_lag_port_join(struct lan966x_port *port,
struct net_device *brport_dev,
struct net_device *bond,
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 08/15] net: lan966x: add FDMA ops dispatch for PCIe support
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (6 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 07/15] net: lan966x: use a dedicated device for DMA operations Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
` (6 subsequent siblings)
14 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Introduce lan966x_fdma_ops to support different FDMA implementations
for platform and PCIe. Plumb fdma_init, fdma_deinit, fdma_xmit,
fdma_poll and fdma_resize through the ops table, and select the
implementation at probe time. Only the platform implementation exists
at this point; the PCIe implementation is added in a later patch.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 2 +-
.../net/ethernet/microchip/lan966x/lan966x_main.c | 20 +++++++++++++++-----
.../net/ethernet/microchip/lan966x/lan966x_main.h | 13 +++++++++++++
3 files changed, 29 insertions(+), 6 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index ae6e5aec5e13..5f46166cbfe4 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -948,7 +948,7 @@ void lan966x_fdma_netdev_init(struct lan966x *lan966x, struct net_device *dev)
return;
lan966x->fdma_ndev = dev;
- netif_napi_add(dev, &lan966x->napi, lan966x_fdma_napi_poll);
+ netif_napi_add(dev, &lan966x->napi, lan966x->ops->fdma_poll);
napi_enable(&lan966x->napi);
}
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 2741f7c9fa4c..6e6c08bb8eea 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -27,6 +27,14 @@
#define IO_RANGES 2
+static const struct lan966x_fdma_ops lan966x_fdma_ops = {
+ .fdma_init = &lan966x_fdma_init,
+ .fdma_deinit = &lan966x_fdma_deinit,
+ .fdma_xmit = &lan966x_fdma_xmit,
+ .fdma_poll = &lan966x_fdma_napi_poll,
+ .fdma_resize = &lan966x_fdma_change_mtu,
+};
+
static const struct of_device_id lan966x_match[] = {
{ .compatible = "microchip,lan966x-switch" },
{ }
@@ -392,7 +400,7 @@ static netdev_tx_t lan966x_port_xmit(struct sk_buff *skb,
spin_lock(&lan966x->tx_lock);
if (port->lan966x->fdma)
- err = lan966x_fdma_xmit(skb, ifh, dev);
+ err = lan966x->ops->fdma_xmit(skb, ifh, dev);
else
err = lan966x_port_ifh_xmit(skb, ifh, dev);
spin_unlock(&lan966x->tx_lock);
@@ -414,7 +422,7 @@ static int lan966x_port_change_mtu(struct net_device *dev, int new_mtu)
if (!lan966x->fdma)
return 0;
- err = lan966x_fdma_change_mtu(lan966x);
+ err = lan966x->ops->fdma_resize(lan966x);
if (err) {
lan_wr(DEV_MAC_MAXLEN_CFG_MAX_LEN_SET(LAN966X_HW_MTU(old_mtu)),
lan966x, DEV_MAC_MAXLEN_CFG(port->chip_port));
@@ -1111,6 +1119,8 @@ static int lan966x_probe(struct platform_device *pdev)
lan966x->dev = &pdev->dev;
lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
+ lan966x->ops = &lan966x_fdma_ops;
+
if (!device_get_mac_address(&pdev->dev, mac_addr)) {
ether_addr_copy(lan966x->base_mac, mac_addr);
} else {
@@ -1250,7 +1260,7 @@ static int lan966x_probe(struct platform_device *pdev)
if (err)
goto cleanup_fdb;
- err = lan966x_fdma_init(lan966x);
+ err = lan966x->ops->fdma_init(lan966x);
if (err)
goto cleanup_ptp;
@@ -1263,7 +1273,7 @@ static int lan966x_probe(struct platform_device *pdev)
return 0;
cleanup_fdma:
- lan966x_fdma_deinit(lan966x);
+ lan966x->ops->fdma_deinit(lan966x);
cleanup_ptp:
lan966x_ptp_deinit(lan966x);
@@ -1291,7 +1301,7 @@ static void lan966x_remove(struct platform_device *pdev)
lan966x_taprio_deinit(lan966x);
lan966x_vcap_deinit(lan966x);
- lan966x_fdma_deinit(lan966x);
+ lan966x->ops->fdma_deinit(lan966x);
lan966x_cleanup_ports(lan966x);
cancel_delayed_work_sync(&lan966x->stats_work);
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 3f09f8ddf620..aab5e53ed059 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -193,6 +193,17 @@ enum vcap_is1_port_sel_rt {
VCAP_IS1_PS_RT_FOLLOW_OTHER = 7,
};
+struct lan966x;
+
+struct lan966x_fdma_ops {
+ int (*fdma_init)(struct lan966x *lan966x);
+ void (*fdma_deinit)(struct lan966x *lan966x);
+ int (*fdma_xmit)(struct sk_buff *skb, __be32 *ifh,
+ struct net_device *dev);
+ int (*fdma_poll)(struct napi_struct *napi, int weight);
+ int (*fdma_resize)(struct lan966x *lan966x);
+};
+
struct lan966x_port;
struct lan966x_rx {
@@ -273,6 +284,8 @@ struct lan966x {
/* Device used for DMA; the PCIe endpoint when enumerated over PCIe. */
struct device *dma_dev;
+ const struct lan966x_fdma_ops *ops;
+
u8 num_phys_ports;
struct lan966x_port **ports;
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (7 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 08/15] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
` (5 subsequent siblings)
14 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
When in PCI mode, the GCB soft reset issued by the reset controller
can latch spurious bits in the FDMA error stickies. The latched bits
sit in FDMA_INTR_ERR until the FDMA IRQ is requested later in probe,
at which point the handler fires immediately and WARNs.
Clear FDMA_ERRORS, FDMA_INTR_ERR and FDMA_INTR_DB right after the
switch reset so the FDMA comes out clean and the IRQ handler does not
see ghost errors on probe.
The clear runs on both the PCI and platform paths. On the platform
path it has no effect — there are no spurious stickies to clear — but
keeping it unconditional avoids a PCI-specific code path here.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
drivers/net/ethernet/microchip/lan966x/lan966x_main.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 6e6c08bb8eea..259d81e75907 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -1067,6 +1067,15 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
reset_control_reset(switch_reset);
+ /* When in PCI mode, the GCB soft reset issued by the reset
+ * controller can latch spurious bits in the FDMA error and
+ * data-block stickies. Clear them before request_irq hooks up the
+ * FDMA IRQ line, otherwise the handler fires immediately on probe.
+ */
+ lan_wr(lan_rd(lan966x, FDMA_ERRORS), lan966x, FDMA_ERRORS);
+ lan_wr(lan_rd(lan966x, FDMA_INTR_ERR), lan966x, FDMA_INTR_ERR);
+ lan_wr(lan_rd(lan966x, FDMA_INTR_DB), lan966x, FDMA_INTR_DB);
+
/* Don't reinitialize the switch core, if it is already initialized. In
* case it is initialized twice, some pointers inside the queue system
* in HW will get corrupted and then after a while the queue system gets
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (8 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:32 ` [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
` (4 subsequent siblings)
14 siblings, 1 reply; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
As a PCIe endpoint, lan966x is not reset by a host reboot: its FDMA
channels and interrupt sources stay armed, and the OIC ORs every
source into the shared PCIe INTx, asserted before the driver has
re-probed. A still-active channel also keeps write access to host
memory the next kernel will reuse.
Add a shutdown callback that:
- frees the ana, xtr and FDMA irqs, masking and unmapping them at
the OIC (disable_irq() would leave both set - the OIC has no
irq_disable())
- masks the analyzer source, armed unconditionally by lan966x_init()
and re-armed by the MAC table's age timer
- stops and detaches the netdevs, draining in-flight xmit and
clearing netif_device_present() so ndo_open/ndo_change_mtu cannot
re-enter the FDMA against a disabled NAPI
- disables both FDMA channels and masks their interrupts
- unmaps the outbound ATU windows, leaving none armed (the windows
are mapped by the PCIe FDMA backend added later in this series)
NAPI is skipped when fdma_ndev is unset (a probed switch with no
usable port never adds one), and XDP attach cannot re-enter the FDMA
either: on PCIe, lan966x_xdp_setup() returns before the page pool
reload, which is the only point where it touches the FDMA.
Only the PCIe instantiation needs this - the SoC one resets with the
chip - so the callback returns early on a platform device; the check
is at runtime since .shutdown belongs to the driver, and a
PCIe-enabled kernel binds both.
FDMA_INTR_ENA persists across a warm reboot, so also restore the
full enable in lan966x_fdma_rx_start(), run after both rings are
allocated. The PCIe FDMA backend added later in this series starts
RX through the same function, so both backends are re-armed from
one site.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 4 ++
.../net/ethernet/microchip/lan966x/lan966x_main.c | 56 ++++++++++++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_regs.h | 15 ++++++
3 files changed, 75 insertions(+)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 5f46166cbfe4..2e8f786d6fee 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -146,6 +146,10 @@ void lan966x_fdma_rx_start(struct lan966x_rx *rx)
struct fdma *fdma = &rx->fdma;
u32 mask;
+ lan_wr(FDMA_INTR_ENA_INTR_PORT_ENA_SET(GENMASK(1, 0)) |
+ FDMA_INTR_ENA_INTR_CH_ENA_SET(GENMASK(7, 0)),
+ lan966x, FDMA_INTR_ENA);
+
lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 259d81e75907..024ce9f9916c 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -1324,9 +1324,65 @@ static void lan966x_remove(struct platform_device *pdev)
debugfs_remove_recursive(lan966x->debugfs_root);
}
+static void lan966x_shutdown(struct platform_device *pdev)
+{
+ struct lan966x *lan966x = platform_get_drvdata(pdev);
+
+ /* As a PCIe endpoint the switch is not reset by the host reboot, so it
+ * has to be quiesced here:
+ *
+ * Free the irqs and mask the sources: no source can assert INTx.
+ * Disable NAPI: the teardown must not race a poll.
+ * Stop and detach the netdevs: drains xmit, closes ndo_open and MTU.
+ * Stop the FDMA channels: waits for the engine to go idle.
+ * Unmap the ATU windows: revokes the engine's access to host memory.
+ */
+ if (!lan966x_is_pci(lan966x))
+ return;
+
+ if (lan966x->xtr_irq > 0)
+ devm_free_irq(lan966x->dev, lan966x->xtr_irq, lan966x);
+ if (lan966x->ana_irq > 0)
+ devm_free_irq(lan966x->dev, lan966x->ana_irq, lan966x);
+ if (lan966x->fdma_irq > 0)
+ devm_free_irq(lan966x->dev, lan966x->fdma_irq, lan966x);
+
+ lan_wr(0, lan966x, ANA_ANAINTR);
+
+ if (!lan966x->fdma)
+ return;
+
+ rtnl_lock();
+
+ if (lan966x->fdma_ndev)
+ napi_disable(&lan966x->napi);
+
+ for (int p = 0; p < lan966x->num_phys_ports; p++) {
+ if (!lan966x->ports[p] || !lan966x->ports[p]->dev)
+ continue;
+
+ netif_tx_disable(lan966x->ports[p]->dev);
+ netif_device_detach(lan966x->ports[p]->dev);
+ }
+
+ lan966x_fdma_rx_disable(&lan966x->rx);
+ lan966x_fdma_tx_disable(&lan966x->tx);
+
+ lan_wr(0, lan966x, FDMA_INTR_ENA);
+ lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
+
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+ fdma_pci_atu_region_unmap(lan966x->rx.fdma.atu_region);
+ fdma_pci_atu_region_unmap(lan966x->tx.fdma.atu_region);
+#endif
+
+ rtnl_unlock();
+}
+
static struct platform_driver lan966x_driver = {
.probe = lan966x_probe,
.remove = lan966x_remove,
+ .shutdown = lan966x_shutdown,
.driver = {
.name = "lan966x-switch",
.of_match_table = lan966x_match,
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
index 4b553927d2e0..aba0d36ae6b5 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
@@ -1039,6 +1039,21 @@ enum lan966x_target {
/* FDMA:FDMA:FDMA_INTR_ERR */
#define FDMA_INTR_ERR __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 400, 0, 1, 4)
+/* FDMA:FDMA:FDMA_INTR_ENA */
+#define FDMA_INTR_ENA __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 404, 0, 1, 4)
+
+#define FDMA_INTR_ENA_INTR_PORT_ENA GENMASK(9, 8)
+#define FDMA_INTR_ENA_INTR_PORT_ENA_SET(x)\
+ FIELD_PREP(FDMA_INTR_ENA_INTR_PORT_ENA, x)
+#define FDMA_INTR_ENA_INTR_PORT_ENA_GET(x)\
+ FIELD_GET(FDMA_INTR_ENA_INTR_PORT_ENA, x)
+
+#define FDMA_INTR_ENA_INTR_CH_ENA GENMASK(7, 0)
+#define FDMA_INTR_ENA_INTR_CH_ENA_SET(x)\
+ FIELD_PREP(FDMA_INTR_ENA_INTR_CH_ENA, x)
+#define FDMA_INTR_ENA_INTR_CH_ENA_GET(x)\
+ FIELD_GET(FDMA_INTR_ENA_INTR_CH_ENA, x)
+
/* FDMA:FDMA:FDMA_ERRORS */
#define FDMA_ERRORS __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 412, 0, 1, 4)
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (9 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
@ 2026-09-28 19:32 ` Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 14:14 ` Simon Horman
2026-09-28 19:33 ` [PATCH net-next v9 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
` (3 subsequent siblings)
14 siblings, 2 replies; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:32 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add PCIe FDMA support for lan966x. The PCIe FDMA path uses contiguous
DMA buffers mapped through the endpoint's ATU, with memcpy-based frame
transfer instead of per-page DMA mappings.
With PCIe FDMA, throughput increases from ~33 Mbps (register-based I/O)
to ~620 Mbps on an Intel x86 host with a lan966x PCIe card.
The RX path uses its own copy of lan966x_hw_offload() that reports back
when skb_vlan_untag() frees the skb, so the caller can drop the frame
instead of touching it.
XDP is not supported on this path yet, so do not advertise xdp_features
and reject program attach for PCIe instances.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
drivers/net/ethernet/microchip/lan966x/Makefile | 4 +
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 458 +++++++++++++++++++++
.../net/ethernet/microchip/lan966x/lan966x_main.c | 11 +-
.../net/ethernet/microchip/lan966x/lan966x_main.h | 7 +
.../net/ethernet/microchip/lan966x/lan966x_regs.h | 10 +
.../net/ethernet/microchip/lan966x/lan966x_xdp.c | 6 +
6 files changed, 493 insertions(+), 3 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/Makefile b/drivers/net/ethernet/microchip/lan966x/Makefile
index 4cdbe263502c..02db4d4d9826 100644
--- a/drivers/net/ethernet/microchip/lan966x/Makefile
+++ b/drivers/net/ethernet/microchip/lan966x/Makefile
@@ -18,6 +18,10 @@ lan966x-switch-objs := lan966x_main.o lan966x_phylink.o lan966x_port.o \
lan966x-switch-$(CONFIG_LAN966X_DCB) += lan966x_dcb.o
lan966x-switch-$(CONFIG_DEBUG_FS) += lan966x_vcap_debugfs.o
+ifneq ($(CONFIG_MCHP_LAN966X_PCI),)
+lan966x-switch-y += lan966x_fdma_pci.o
+endif
+
# Provide include files
ccflags-y += -I$(srctree)/drivers/net/ethernet/microchip/vcap
ccflags-y += -I$(srctree)/drivers/net/ethernet/microchip/fdma
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
new file mode 100644
index 000000000000..f511e7061314
--- /dev/null
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -0,0 +1,458 @@
+// SPDX-License-Identifier: GPL-2.0+
+
+#include <linux/ip.h>
+#include <net/addrconf.h>
+
+#include "fdma_api.h"
+#include "lan966x_main.h"
+
+static int lan966x_fdma_pci_dataptr_cb(struct fdma *fdma, int dcb, int db,
+ u64 *dataptr)
+{
+ u64 addr;
+
+ addr = fdma_dataptr_dma_addr_contiguous(fdma, dcb, db);
+
+ *dataptr = fdma_pci_atu_translate_addr(fdma->atu_region, addr);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_nextptr_cb(struct fdma *fdma, int dcb, u64 *nextptr)
+{
+ u64 addr;
+
+ fdma_nextptr_cb(fdma, dcb, &addr);
+
+ *nextptr = fdma_pci_atu_translate_addr(fdma->atu_region, addr);
+
+ return 0;
+}
+
+/* Stop the TX queues on every port, so nothing feeds the injection channel
+ * while it is torn down or resized.
+ */
+static void lan966x_fdma_tx_disable_netdev(struct lan966x *lan966x)
+{
+ struct lan966x_port *port;
+ int i;
+
+ for (i = 0; i < lan966x->num_phys_ports; ++i) {
+ port = lan966x->ports[i];
+ if (!port)
+ continue;
+
+ netif_tx_disable(port->dev);
+ }
+}
+
+static int lan966x_fdma_pci_rx_alloc(struct lan966x_rx *rx)
+{
+ struct lan966x *lan966x = rx->lan966x;
+ struct fdma *fdma = &rx->fdma;
+ int err;
+
+ err = fdma_alloc_coherent_and_map(lan966x->dma_dev, fdma,
+ &lan966x->atu);
+ if (err)
+ return err;
+
+ err = fdma_dcbs_init(fdma,
+ FDMA_DCB_INFO_DATAL(fdma->db_size - XDP_PACKET_HEADROOM),
+ FDMA_DCB_STATUS_INTR);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, fdma);
+ return err;
+ }
+
+ lan966x_fdma_llp_configure(lan966x,
+ fdma->atu_region->base_addr,
+ fdma->channel_id);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_tx_alloc(struct lan966x_tx *tx)
+{
+ struct lan966x *lan966x = tx->lan966x;
+ struct fdma *fdma = &tx->fdma;
+ int err;
+
+ err = fdma_alloc_coherent_and_map(lan966x->dma_dev, fdma,
+ &lan966x->atu);
+ if (err)
+ return err;
+
+ err = fdma_dcbs_init(fdma,
+ FDMA_DCB_INFO_DATAL(fdma->db_size),
+ FDMA_DCB_STATUS_DONE);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, fdma);
+ return err;
+ }
+
+ lan966x_fdma_llp_configure(lan966x,
+ fdma->atu_region->base_addr,
+ fdma->channel_id);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_get_next_dcb(struct fdma *fdma)
+{
+ struct fdma_db *db;
+
+ for (int i = 0; i < fdma->n_dcbs; i++) {
+ db = fdma_db_get(fdma, i, 0);
+
+ if (!fdma_db_is_done(db))
+ continue;
+ if (fdma_is_last(fdma, &fdma->dcbs[i]))
+ continue;
+
+ return i;
+ }
+
+ return -ENOSPC;
+}
+
+/* TX slot layout (sizes in bytes):
+ *
+ * +---------------------+-----+---------+-----+
+ * | XDP_PACKET_HEADROOM | IFH | payload | FCS |
+ * | 256 | 28 | len | 4 |
+ * +---------------------+-----+---------+-----+
+ * |<---------------- db_size ----------------->|
+ *
+ * Return true if the frame plus required overhead fits.
+ */
+static bool lan966x_fdma_pci_tx_size_fits(struct fdma *fdma, u32 len)
+{
+ return XDP_PACKET_HEADROOM + IFH_LEN_BYTES + len + ETH_FCS_LEN <=
+ fdma->db_size;
+}
+
+/* Return true if blockl is a valid RX frame size. */
+static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
+{
+ return blockl >= IFH_LEN_BYTES + ETH_HLEN + ETH_FCS_LEN &&
+ blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
+}
+
+static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
+{
+ struct lan966x *lan966x = rx->lan966x;
+ struct fdma *fdma = &rx->fdma;
+ struct lan966x_port *port;
+ struct fdma_db *db;
+ void *virt_addr;
+ u32 blockl;
+
+ /* virt_addr points to the IFH. */
+ virt_addr = fdma_dataptr_virt_addr_contiguous(fdma,
+ fdma->dcb_index,
+ fdma->db_index);
+
+ lan966x_ifh_get_src_port(virt_addr, src_port);
+
+ if (*src_port >= lan966x->num_phys_ports)
+ return FDMA_ERROR;
+
+ port = lan966x->ports[*src_port];
+ if (!port)
+ return FDMA_ERROR;
+
+ db = fdma_db_next_get(fdma);
+
+ /* BLOCKL is a 16-bit HW-populated field; reject obviously-bad
+ * values before they feed memcpy/XDP sizes.
+ */
+ blockl = fdma_db_len_get(db);
+ if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
+ return FDMA_ERROR;
+
+ return FDMA_PASS;
+}
+
+static bool lan966x_fdma_pci_hw_offload(struct lan966x *lan966x, u32 port,
+ struct sk_buff **pskb)
+{
+ struct sk_buff *skb = *pskb;
+ u32 val;
+
+ val = lan_rd(lan966x, ANA_CPU_FWD_CFG(port));
+ if (!(val & (ANA_CPU_FWD_CFG_IGMP_REDIR_ENA |
+ ANA_CPU_FWD_CFG_MLD_REDIR_ENA)))
+ return true;
+
+ if (eth_type_vlan(skb->protocol)) {
+ skb = skb_vlan_untag(skb);
+ *pskb = skb;
+ if (unlikely(!skb))
+ return false;
+ }
+
+ if (skb->protocol == htons(ETH_P_IP) &&
+ ip_hdr(skb)->protocol == IPPROTO_IGMP)
+ return false;
+
+ if (IS_ENABLED(CONFIG_IPV6) &&
+ skb->protocol == htons(ETH_P_IPV6) &&
+ ipv6_addr_is_multicast(&ipv6_hdr(skb)->daddr) &&
+ !ipv6_mc_check_mld(skb))
+ return false;
+
+ return true;
+}
+
+static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
+ u64 src_port)
+{
+ struct lan966x *lan966x = rx->lan966x;
+ struct fdma *fdma = &rx->fdma;
+ struct sk_buff *skb;
+ struct fdma_db *db;
+ u32 data_len;
+
+ /* Get the received frame and create an SKB for it. */
+ db = fdma_db_next_get(fdma);
+ data_len = fdma_db_len_get(db);
+
+ skb = napi_alloc_skb(&lan966x->napi, data_len);
+ if (unlikely(!skb))
+ return NULL;
+
+ memcpy(skb->data,
+ fdma_dataptr_virt_addr_contiguous(fdma,
+ fdma->dcb_index,
+ fdma->db_index),
+ data_len);
+
+ skb_put(skb, data_len);
+
+ skb->dev = lan966x->ports[src_port]->dev;
+ skb_pull(skb, IFH_LEN_BYTES);
+
+ skb_trim(skb, skb->len - ETH_FCS_LEN);
+
+ skb->protocol = eth_type_trans(skb, skb->dev);
+
+ if (lan966x->bridge_mask & BIT(src_port)) {
+ skb->offload_fwd_mark = 1;
+
+ skb_reset_network_header(skb);
+ if (!lan966x_fdma_pci_hw_offload(lan966x, src_port, &skb)) {
+ if (unlikely(!skb))
+ return NULL;
+ skb->offload_fwd_mark = 0;
+ }
+ }
+
+ skb->dev->stats.rx_bytes += skb->len;
+ skb->dev->stats.rx_packets++;
+
+ return skb;
+}
+
+static int lan966x_fdma_pci_xmit(struct sk_buff *skb, __be32 *ifh,
+ struct net_device *dev)
+{
+ struct lan966x_port *port = netdev_priv(dev);
+ struct lan966x *lan966x = port->lan966x;
+ struct lan966x_tx *tx = &lan966x->tx;
+ struct fdma *fdma = &tx->fdma;
+ int next_to_use;
+ void *virt_addr;
+
+ next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
+
+ if (next_to_use < 0) {
+ netif_stop_queue(dev);
+ return NETDEV_TX_BUSY;
+ }
+
+ if (skb_put_padto(skb, ETH_ZLEN)) {
+ dev->stats.tx_dropped++;
+ return NETDEV_TX_OK;
+ }
+
+ if (!lan966x_fdma_pci_tx_size_fits(fdma, skb->len)) {
+ dev_kfree_skb_any(skb);
+ dev->stats.tx_dropped++;
+ return NETDEV_TX_OK;
+ }
+
+ skb_tx_timestamp(skb);
+
+ /* virt_addr points to the IFH. */
+ virt_addr = fdma_dataptr_virt_addr_contiguous(fdma, next_to_use, 0);
+ memcpy(virt_addr, ifh, IFH_LEN_BYTES);
+ memcpy(virt_addr + IFH_LEN_BYTES, skb->data, skb->len);
+
+ /* Order frame write before DCB status write below. */
+ dma_wmb();
+
+ fdma_dcb_add(fdma,
+ next_to_use,
+ 0,
+ FDMA_DCB_STATUS_INTR |
+ FDMA_DCB_STATUS_SOF |
+ FDMA_DCB_STATUS_EOF |
+ FDMA_DCB_STATUS_BLOCKO(0) |
+ FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + skb->len + ETH_FCS_LEN));
+
+ /* Start the transmission. */
+ lan966x_fdma_tx_start(tx);
+
+ dev->stats.tx_bytes += skb->len;
+ dev->stats.tx_packets++;
+
+ /* Safe to free: PTP is not supported on the PCIe path yet,
+ * so lan966x->ptp is always 0 here.
+ */
+ dev_consume_skb_any(skb);
+
+ return NETDEV_TX_OK;
+}
+
+static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
+{
+ struct lan966x *lan966x = container_of(napi, struct lan966x, napi);
+ struct lan966x_rx *rx = &lan966x->rx;
+ struct fdma *fdma = &rx->fdma;
+ int dcb_reload, old_dcb;
+ struct sk_buff *skb;
+ int counter = 0;
+ u64 src_port;
+
+ /* Wake any stopped TX queues if a TX DCB is available. */
+ spin_lock(&lan966x->tx_lock);
+ if (lan966x_fdma_pci_get_next_dcb(&lan966x->tx.fdma) >= 0)
+ lan966x_fdma_wakeup_netdev(lan966x);
+ spin_unlock(&lan966x->tx_lock);
+
+ dcb_reload = fdma->dcb_index;
+
+ /* Get all received skbs. */
+ while (counter < weight) {
+ if (!fdma_has_frames(fdma))
+ break;
+ /* Order DONE read before DCB/frame reads below. */
+ dma_rmb();
+ counter++;
+ switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
+ case FDMA_PASS:
+ break;
+ case FDMA_ERROR:
+ /* No rx_dropped increment here because src_port is
+ * invalid.
+ */
+ fdma_dcb_advance(fdma);
+ continue;
+ }
+ skb = lan966x_fdma_pci_rx_get_frame(rx, src_port);
+ fdma_dcb_advance(fdma);
+ if (!skb) {
+ lan966x->ports[src_port]->dev->stats.rx_dropped++;
+ continue;
+ }
+
+ napi_gro_receive(&lan966x->napi, skb);
+ }
+ while (dcb_reload != fdma->dcb_index) {
+ old_dcb = dcb_reload;
+ dcb_reload++;
+ dcb_reload &= fdma->n_dcbs - 1;
+
+ fdma_dcb_add(fdma,
+ old_dcb,
+ FDMA_DCB_INFO_DATAL(fdma->db_size - XDP_PACKET_HEADROOM),
+ FDMA_DCB_STATUS_INTR);
+
+ lan966x_fdma_rx_reload(rx);
+ }
+
+ if (counter < weight && napi_complete_done(napi, counter))
+ lan_wr(0xff, lan966x, FDMA_INTR_DB_ENA);
+
+ return counter;
+}
+
+static int lan966x_fdma_pci_init(struct lan966x *lan966x)
+{
+ struct fdma *rx_fdma = &lan966x->rx.fdma;
+ struct fdma *tx_fdma = &lan966x->tx.fdma;
+ int err;
+
+ if (!lan966x->fdma)
+ return 0;
+
+ lan_wr(FDMA_CTRL_NRESET_SET(0), lan966x, FDMA_CTRL);
+ lan_wr(FDMA_CTRL_NRESET_SET(1), lan966x, FDMA_CTRL);
+
+ fdma_pci_atu_init(&lan966x->atu, lan966x->regs[TARGET_PCIE_DBI]);
+
+ lan966x->rx.lan966x = lan966x;
+ lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
+ rx_fdma->channel_id = FDMA_XTR_CHANNEL;
+ rx_fdma->n_dcbs = FDMA_DCB_MAX;
+ rx_fdma->n_dbs = FDMA_RX_DCB_MAX_DBS;
+ rx_fdma->priv = lan966x;
+ rx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ rx_fdma->size = fdma_get_size_contiguous(rx_fdma);
+ rx_fdma->ops.nextptr_cb = &lan966x_fdma_pci_nextptr_cb;
+ rx_fdma->ops.dataptr_cb = &lan966x_fdma_pci_dataptr_cb;
+
+ lan966x->tx.lan966x = lan966x;
+ tx_fdma->channel_id = FDMA_INJ_CHANNEL;
+ tx_fdma->n_dcbs = FDMA_DCB_MAX;
+ tx_fdma->n_dbs = FDMA_TX_DCB_MAX_DBS;
+ tx_fdma->priv = lan966x;
+ tx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ tx_fdma->size = fdma_get_size_contiguous(tx_fdma);
+ tx_fdma->ops.nextptr_cb = &lan966x_fdma_pci_nextptr_cb;
+ tx_fdma->ops.dataptr_cb = &lan966x_fdma_pci_dataptr_cb;
+
+ err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
+ if (err)
+ return err;
+
+ err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, rx_fdma);
+ return err;
+ }
+
+ lan966x_fdma_rx_start(&lan966x->rx);
+
+ return 0;
+}
+
+static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
+{
+ return -EOPNOTSUPP;
+}
+
+static void lan966x_fdma_pci_deinit(struct lan966x *lan966x)
+{
+ if (!lan966x->fdma)
+ return;
+
+ if (lan966x->fdma_ndev)
+ napi_disable(&lan966x->napi);
+
+ lan966x_fdma_tx_disable_netdev(lan966x);
+ lan966x_fdma_rx_disable(&lan966x->rx);
+ lan966x_fdma_tx_disable(&lan966x->tx);
+
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->rx.fdma);
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->tx.fdma);
+}
+
+const struct lan966x_fdma_ops lan966x_fdma_pci_ops = {
+ .fdma_init = &lan966x_fdma_pci_init,
+ .fdma_deinit = &lan966x_fdma_pci_deinit,
+ .fdma_xmit = &lan966x_fdma_pci_xmit,
+ .fdma_poll = &lan966x_fdma_pci_napi_poll,
+ .fdma_resize = &lan966x_fdma_pci_resize,
+};
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 024ce9f9916c..de2202786826 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -50,6 +50,7 @@ struct lan966x_main_io_resource {
static const struct lan966x_main_io_resource lan966x_main_iomap[] = {
{ TARGET_CPU, 0xc0000, 0 }, /* 0xe00c0000 */
{ TARGET_FDMA, 0xc0400, 0 }, /* 0xe00c0400 */
+ { TARGET_PCIE_DBI, 0x400000, 0 }, /* 0xe0400000 */
{ TARGET_ORG, 0, 1 }, /* 0xe2000000 */
{ TARGET_GCB, 0x4000, 1 }, /* 0xe2004000 */
{ TARGET_QS, 0x8000, 1 }, /* 0xe2008000 */
@@ -873,7 +874,8 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
port->phylink = phylink;
- if (lan966x->fdma)
+ /* XDP is not supported on the PCIe FDMA path. */
+ if (lan966x->fdma && !lan966x_is_pci(lan966x))
dev->xdp_features = NETDEV_XDP_ACT_BASIC |
NETDEV_XDP_ACT_REDIRECT |
NETDEV_XDP_ACT_NDO_XMIT;
@@ -1128,7 +1130,8 @@ static int lan966x_probe(struct platform_device *pdev)
lan966x->dev = &pdev->dev;
lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
- lan966x->ops = &lan966x_fdma_ops;
+ lan966x->ops = lan966x_is_pci(lan966x) ? &lan966x_fdma_pci_ops :
+ &lan966x_fdma_ops;
if (!device_get_mac_address(&pdev->dev, mac_addr)) {
ether_addr_copy(lan966x->base_mac, mac_addr);
@@ -1187,7 +1190,9 @@ static int lan966x_probe(struct platform_device *pdev)
if (err)
return dev_err_probe(&pdev->dev, err, "Unable to use ptp irq");
- lan966x->ptp = 1;
+ /* PTP is not supported on the PCIe path yet. */
+ if (!lan966x_is_pci(lan966x))
+ lan966x->ptp = 1;
}
lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index aab5e53ed059..16bc28c8f11f 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -17,6 +17,7 @@
#include <net/xdp.h>
#include <fdma_api.h>
+#include <fdma_pci.h>
#include <vcap_api.h>
#include <vcap_api_client.h>
@@ -291,6 +292,10 @@ struct lan966x {
void __iomem *regs[NUM_TARGETS];
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+ struct fdma_pci_atu atu;
+#endif
+
int shared_queue_sz;
u8 base_mac[ETH_ALEN];
@@ -589,6 +594,8 @@ void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
int lan966x_qsys_sw_status(struct lan966x *lan966x);
+extern const struct lan966x_fdma_ops lan966x_fdma_pci_ops;
+
/* dma_dev differs from dev only on the PCIe path. */
static inline bool lan966x_is_pci(struct lan966x *lan966x)
{
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
index aba0d36ae6b5..4778ea217673 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
@@ -20,6 +20,7 @@ enum lan966x_target {
TARGET_FDMA = 21,
TARGET_GCB = 27,
TARGET_ORG = 36,
+ TARGET_PCIE_DBI = 40,
TARGET_PTP = 41,
TARGET_QS = 42,
TARGET_QSYS = 46,
@@ -1009,6 +1010,15 @@ enum lan966x_target {
#define FDMA_CH_CFG_CH_MEM_GET(x)\
FIELD_GET(FDMA_CH_CFG_CH_MEM, x)
+/* FDMA:FDMA:FDMA_CTRL */
+#define FDMA_CTRL __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 424, 0, 1, 4)
+
+#define FDMA_CTRL_NRESET BIT(0)
+#define FDMA_CTRL_NRESET_SET(x)\
+ FIELD_PREP(FDMA_CTRL_NRESET, x)
+#define FDMA_CTRL_NRESET_GET(x)\
+ FIELD_GET(FDMA_CTRL_NRESET, x)
+
/* FDMA:FDMA:FDMA_PORT_CTRL */
#define FDMA_PORT_CTRL(r) __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 376, r, 2, 4)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
index 9ee61db8690b..e63634887f64 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
@@ -20,6 +20,12 @@ static int lan966x_xdp_setup(struct net_device *dev, struct netdev_bpf *xdp)
return -EOPNOTSUPP;
}
+ if (lan966x_is_pci(lan966x)) {
+ NL_SET_ERR_MSG_MOD(xdp->extack,
+ "XDP is not supported on the PCIe FDMA path");
+ return -EOPNOTSUPP;
+ }
+
old_xdp = lan966x_xdp_present(lan966x);
old_prog = xchg(&port->xdp_prog, xdp->prog);
new_xdp = lan966x_xdp_present(lan966x);
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 12/15] net: lan966x: add PCIe FDMA MTU change support
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (10 preceding siblings ...)
2026-09-28 19:32 ` [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
@ 2026-09-28 19:33 ` Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:33 ` [PATCH net-next v9 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
` (2 subsequent siblings)
14 siblings, 1 reply; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:33 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add MTU change support for the PCIe FDMA path: on an MTU change, the
contiguous ATU-mapped RX and TX buffers are reallocated at the new
size, falling back to the existing buffers on failure.
Cap the PCIe DCB ring at 256 (FDMA_PCI_DCB_MAX): 512 DCBs would
overflow MAX_PAGE_ORDER at jumbo MTU.
The ring must fit one MAX_PAGE_ORDER block after ATU padding, and
db_size is handed to the FDMA in the 16-bit DCB DATAL field.
Advertise the resulting limit in dev->max_mtu (FDMA_PCI_MAX_MTU)
when the FDMA is in use - a switch with no "fdma" interrupt named
uses the unconstrained register-based path instead, so the cap is
gated on lan966x->fdma, not lan966x_is_pci() alone. On a 4KB-page,
MAX_PAGE_ORDER=10 build this is MTU 15498; the overhead is shared
with lan966x_fdma_get_max_frame() via FDMA_OVERHEAD.
Skip the resize until lan966x_fdma_pci_init() has built the rings;
it runs after the netdevs register and sizes them from
DEV_MAC_MAXLEN_CFG, already programmed by the caller.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
.../net/ethernet/microchip/lan966x/lan966x_fdma.c | 6 +-
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 153 ++++++++++++++++++++-
.../net/ethernet/microchip/lan966x/lan966x_main.c | 3 +-
.../net/ethernet/microchip/lan966x/lan966x_main.h | 28 ++++
4 files changed, 181 insertions(+), 9 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 2e8f786d6fee..a7940eca5df3 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -890,11 +890,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
{
- return lan966x_fdma_get_max_mtu(lan966x) +
- IFH_LEN_BYTES +
- SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) +
- VLAN_HLEN * 2 +
- XDP_PACKET_HEADROOM;
+ return lan966x_fdma_get_max_mtu(lan966x) + FDMA_OVERHEAD;
}
static int __lan966x_fdma_reload(struct lan966x *lan966x, int max_mtu)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
index f511e7061314..758554c951c5 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -395,7 +395,7 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
lan966x->rx.lan966x = lan966x;
lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
rx_fdma->channel_id = FDMA_XTR_CHANNEL;
- rx_fdma->n_dcbs = FDMA_DCB_MAX;
+ rx_fdma->n_dcbs = FDMA_PCI_DCB_MAX;
rx_fdma->n_dbs = FDMA_RX_DCB_MAX_DBS;
rx_fdma->priv = lan966x;
rx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
@@ -405,7 +405,7 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
lan966x->tx.lan966x = lan966x;
tx_fdma->channel_id = FDMA_INJ_CHANNEL;
- tx_fdma->n_dcbs = FDMA_DCB_MAX;
+ tx_fdma->n_dcbs = FDMA_PCI_DCB_MAX;
tx_fdma->n_dbs = FDMA_TX_DCB_MAX_DBS;
tx_fdma->priv = lan966x;
tx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
@@ -428,9 +428,156 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
return 0;
}
+/* Reset existing rx and tx buffers. */
+static void lan966x_fdma_pci_reset_mem(struct lan966x *lan966x)
+{
+ struct lan966x_rx *rx = &lan966x->rx;
+ struct lan966x_tx *tx = &lan966x->tx;
+
+ memset(rx->fdma.dcbs, 0, rx->fdma.size);
+ memset(tx->fdma.dcbs, 0, tx->fdma.size);
+
+ fdma_dcbs_init(&rx->fdma,
+ FDMA_DCB_INFO_DATAL(rx->fdma.db_size - XDP_PACKET_HEADROOM),
+ FDMA_DCB_STATUS_INTR);
+
+ fdma_dcbs_init(&tx->fdma,
+ FDMA_DCB_INFO_DATAL(tx->fdma.db_size),
+ FDMA_DCB_STATUS_DONE);
+
+ lan966x_fdma_llp_configure(lan966x,
+ tx->fdma.atu_region->base_addr,
+ tx->fdma.channel_id);
+ lan966x_fdma_llp_configure(lan966x,
+ rx->fdma.atu_region->base_addr,
+ rx->fdma.channel_id);
+}
+
+/* Wake all TX queues on every port (undoes lan966x_fdma_tx_disable_netdev). */
+static void lan966x_fdma_pci_wakeup_netdev(struct lan966x *lan966x)
+{
+ for (int i = 0; i < lan966x->num_phys_ports; ++i) {
+ struct lan966x_port *port = lan966x->ports[i];
+
+ if (port)
+ netif_tx_wake_all_queues(port->dev);
+ }
+}
+
+static int lan966x_fdma_pci_reload(struct lan966x *lan966x, int new_mtu)
+{
+ struct fdma tx_fdma_old = lan966x->tx.fdma;
+ struct fdma rx_fdma_old = lan966x->rx.fdma;
+ u32 old_mtu = lan966x->rx.max_mtu;
+ int err;
+
+ napi_disable(&lan966x->napi);
+ lan966x_fdma_tx_disable_netdev(lan966x);
+ lan966x_fdma_rx_disable(&lan966x->rx);
+ lan966x_fdma_tx_disable(&lan966x->tx);
+
+ lan966x->rx.max_mtu = new_mtu;
+
+ /* Must be NULL'ed in order to realloc them. */
+ lan966x->rx.fdma.atu_region = NULL;
+ lan966x->tx.fdma.atu_region = NULL;
+
+ lan966x->tx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ lan966x->tx.fdma.size = fdma_get_size_contiguous(&lan966x->tx.fdma);
+ lan966x->rx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+ lan966x->rx.fdma.size = fdma_get_size_contiguous(&lan966x->rx.fdma);
+
+ err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
+ if (err)
+ goto restore;
+
+ err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
+ if (err) {
+ fdma_free_coherent_and_unmap(lan966x->dma_dev,
+ &lan966x->rx.fdma);
+ goto restore;
+ }
+
+ /* Free and unmap old memory. */
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
+ fdma_free_coherent_and_unmap(lan966x->dma_dev, &tx_fdma_old);
+
+ /* Order matters: napi_enable() must precede the wakes, or a TX that
+ * completes first clears FDMA_INTR_DB_ENA with nothing scheduled to
+ * restore it, leaving RX dead until the next reload.
+ */
+ napi_enable(&lan966x->napi);
+ lan966x_fdma_rx_start(&lan966x->rx);
+ lan966x_fdma_pci_wakeup_netdev(lan966x);
+
+ return err;
+restore:
+
+ /* No new buffers are allocated at this point. Use the old buffers,
+ * but reset them before starting the FDMA again.
+ */
+
+ memcpy(&lan966x->tx.fdma, &tx_fdma_old, sizeof(struct fdma));
+ memcpy(&lan966x->rx.fdma, &rx_fdma_old, sizeof(struct fdma));
+
+ lan966x->rx.max_mtu = old_mtu;
+
+ lan966x_fdma_pci_reset_mem(lan966x);
+
+ napi_enable(&lan966x->napi);
+ lan966x_fdma_rx_start(&lan966x->rx);
+ lan966x_fdma_pci_wakeup_netdev(lan966x);
+
+ return err;
+}
+
+static int __lan966x_fdma_pci_reload(struct lan966x *lan966x, int max_mtu)
+{
+ int err;
+ u32 val;
+
+ /* Disable the CPU port. */
+ lan_rmw(QSYS_SW_PORT_MODE_PORT_ENA_SET(0),
+ QSYS_SW_PORT_MODE_PORT_ENA,
+ lan966x, QSYS_SW_PORT_MODE(CPU_PORT));
+
+ /* Flush the CPU queues. */
+ readx_poll_timeout(lan966x_qsys_sw_status,
+ lan966x,
+ val,
+ !(QSYS_SW_STATUS_EQ_AVAIL_GET(val)),
+ READL_SLEEP_US, READL_TIMEOUT_US);
+
+ /* Add a sleep in case there are frames between the queues and the CPU
+ * port
+ */
+ usleep_range(USEC_PER_MSEC, 2 * USEC_PER_MSEC);
+
+ err = lan966x_fdma_pci_reload(lan966x, max_mtu);
+
+ /* Enable back the CPU port. */
+ lan_rmw(QSYS_SW_PORT_MODE_PORT_ENA_SET(1),
+ QSYS_SW_PORT_MODE_PORT_ENA,
+ lan966x, QSYS_SW_PORT_MODE(CPU_PORT));
+
+ return err;
+}
+
static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
{
- return -EOPNOTSUPP;
+ int max_mtu;
+
+ /* Nothing to resize until fdma_pci_init() has built the rings; it
+ * sizes them from DEV_MAC_MAXLEN_CFG, which the caller already set.
+ */
+ if (!lan966x->rx.lan966x)
+ return 0;
+
+ max_mtu = lan966x_fdma_get_max_frame(lan966x);
+ if (max_mtu == lan966x->rx.max_mtu)
+ return 0;
+
+ return __lan966x_fdma_pci_reload(lan966x, max_mtu);
}
static void lan966x_fdma_pci_deinit(struct lan966x *lan966x)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index de2202786826..c3afc4cc597f 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -823,7 +823,8 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
port->chip_port = p;
lan966x->ports[p] = port;
- dev->max_mtu = ETH_MAX_MTU;
+ dev->max_mtu = lan966x_is_pci(lan966x) && lan966x->fdma ?
+ FDMA_PCI_MAX_MTU : ETH_MAX_MTU;
dev->netdev_ops = &lan966x_port_netdev_ops;
dev->ethtool_ops = &lan966x_ethtool_ops;
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 16bc28c8f11f..1877f1916d71 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -7,6 +7,7 @@
#include <linux/etherdevice.h>
#include <linux/if_vlan.h>
#include <linux/jiffies.h>
+#include <linux/mmzone.h>
#include <linux/phy.h>
#include <linux/phylink.h>
#include <linux/ptp_clock_kernel.h>
@@ -87,6 +88,33 @@
#define FDMA_INJ_CHANNEL 0
#define FDMA_DCB_MAX 512
+/* Ring must fit in one MAX_PAGE_ORDER DMA block; 512 DCBs overflows
+ * at jumbo MTU.
+ */
+#define FDMA_PCI_DCB_MAX 256
+
+#define FDMA_OVERHEAD \
+ (IFH_LEN_BYTES + \
+ SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) + \
+ VLAN_HLEN * 2 + \
+ XDP_PACKET_HEADROOM)
+
+/* Largest db_size keeping the ATU-padded ring inside one MAX_PAGE_ORDER
+ * block and within the 16-bit DCB DATAL field. Inverts ALIGN(x, R) <= L
+ * into x <= ALIGN_DOWN(L, R) to bound x directly.
+ */
+#define FDMA_PCI_DB_SIZE_MAX \
+ MIN_T(u32, \
+ (ALIGN_DOWN(PAGE_SIZE << MAX_PAGE_ORDER, \
+ FDMA_PCI_ATU_REGION_ALIGN) - \
+ FDMA_PCI_DCB_MAX * sizeof(struct fdma_dcb)) / \
+ (FDMA_PCI_DCB_MAX * FDMA_RX_DCB_MAX_DBS), \
+ ALIGN_DOWN(GENMASK(15, 0), FDMA_PCI_DB_ALIGN))
+
+#define FDMA_PCI_MAX_MTU \
+ (FDMA_PCI_DB_SIZE_MAX - FDMA_OVERHEAD - \
+ (ETH_HLEN + ETH_FCS_LEN))
+
#define SE_IDX_QUEUE 0 /* 0-79 : Queue scheduler elements */
#define SE_IDX_PORT 80 /* 80-89 : Port schedular elements */
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 13/15] net: lan966x: add PCIe FDMA XDP support
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (11 preceding siblings ...)
2026-09-28 19:33 ` [PATCH net-next v9 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
@ 2026-09-28 19:33 ` Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:33 ` [PATCH net-next v9 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
2026-09-28 19:33 ` [PATCH net-next v9 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
14 siblings, 1 reply; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:33 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add XDP support for the PCIe FDMA path. The implementation operates on
contiguous ATU-mapped buffers with memcpy-based XDP_TX, unlike the
platform path which uses page_pool.
XDP sees the frame with IFH and FCS stripped. These are removed in
lan966x_fdma_pci_rx_check_frame() before the BPF program runs, because
after the program returns the driver cannot tell whether the tail
region was modified. The skb_pull/skb_trim previously done in
lan966x_fdma_pci_rx_get_frame() are removed for the same reason; the
frame pointer and length are pre-computed by rx_check_frame() and
passed through rx_get_frame() and lan966x_xdp_pci_run() to the caller.
lan966x_fdma_pci_xmit_xdpf() handles XDP_TX: it rebuilds a fresh IFH
in the TX slot, copies the post-XDP frame after it, and lets HW insert
a new FCS.
lan966x_xdp_setup() is extended so the PCIe path skips the page_pool
reload that the platform path needs.
Only XDP_ACT_BASIC is supported.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
.../ethernet/microchip/lan966x/lan966x_fdma_pci.c | 166 ++++++++++++++++++---
.../net/ethernet/microchip/lan966x/lan966x_main.c | 12 +-
.../net/ethernet/microchip/lan966x/lan966x_xdp.c | 12 +-
3 files changed, 159 insertions(+), 31 deletions(-)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
index 758554c951c5..949994874ed9 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: GPL-2.0+
+#include <linux/bpf_trace.h>
#include <linux/ip.h>
#include <net/addrconf.h>
@@ -139,7 +140,123 @@ static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
}
-static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
+static int lan966x_fdma_pci_xmit_xdpf(struct lan966x_port *port,
+ void *ptr, u32 len)
+{
+ struct lan966x *lan966x = port->lan966x;
+ struct lan966x_tx *tx = &lan966x->tx;
+ struct fdma *fdma = &tx->fdma;
+ int next_to_use, ret = 0;
+ void *virt_addr;
+
+ spin_lock(&lan966x->tx_lock);
+
+ next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
+
+ if (next_to_use < 0) {
+ netif_stop_queue(port->dev);
+ port->dev->stats.tx_dropped++;
+ ret = NETDEV_TX_BUSY;
+ goto out;
+ }
+
+ /* Only the upper bound is enforced: XDP owns the frame contents and
+ * length, so a program that shrinks below ETH_ZLEN gets what it asked
+ * for.
+ */
+ if (!lan966x_fdma_pci_tx_size_fits(fdma, len)) {
+ port->dev->stats.tx_dropped++;
+ ret = -EINVAL;
+ goto out;
+ }
+
+ /* virt_addr points to the IFH. */
+ virt_addr = fdma_dataptr_virt_addr_contiguous(fdma, next_to_use, 0);
+
+ /* Construct a fresh IFH. */
+ memset(virt_addr, 0, IFH_LEN_BYTES);
+ lan966x_ifh_set_bypass(virt_addr, 1);
+ lan966x_ifh_set_port(virt_addr, BIT_ULL(port->chip_port));
+
+ /* Copy the (post-XDP) frame after the IFH. */
+ memcpy(virt_addr + IFH_LEN_BYTES, ptr, len);
+
+ /* Order frame write before DCB status write below. */
+ dma_wmb();
+
+ /* Reserve ETH_FCS_LEN for the HW-inserted FCS (len is FCS-stripped). */
+ fdma_dcb_add(fdma,
+ next_to_use,
+ 0,
+ FDMA_DCB_STATUS_INTR |
+ FDMA_DCB_STATUS_SOF |
+ FDMA_DCB_STATUS_EOF |
+ FDMA_DCB_STATUS_BLOCKO(0) |
+ FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + len + ETH_FCS_LEN));
+
+ /* Start the transmission. */
+ lan966x_fdma_tx_start(tx);
+
+ port->dev->stats.tx_bytes += len;
+ port->dev->stats.tx_packets++;
+
+out:
+ spin_unlock(&lan966x->tx_lock);
+
+ return ret;
+}
+
+static int lan966x_xdp_pci_run(struct lan966x_port *port, void *data,
+ u32 data_len, void **xdp_data, u32 *xdp_len)
+{
+ /* Read once so the NULL check and bpf_prog_run_xdp() see the same
+ * pointer.
+ */
+ struct bpf_prog *xdp_prog = READ_ONCE(port->xdp_prog);
+ struct lan966x *lan966x = port->lan966x;
+ struct fdma *fdma = &lan966x->rx.fdma;
+ struct xdp_buff xdp;
+ u32 act;
+
+ if (!xdp_prog)
+ return FDMA_PASS;
+
+ xdp_init_buff(&xdp, fdma->db_size, &port->xdp_rxq);
+
+ /* hard_start is set to slot start (virt_addr is XDP_PACKET_HEADROOM
+ * into the slot). Headroom includes the IFH; BPF may grow into it
+ * via adjust_head. IFH is rebuilt on XDP_TX and unread on XDP_PASS.
+ */
+ xdp_prepare_buff(&xdp,
+ data - XDP_PACKET_HEADROOM,
+ XDP_PACKET_HEADROOM + IFH_LEN_BYTES,
+ data_len,
+ false);
+
+ act = bpf_prog_run_xdp(xdp_prog, &xdp);
+
+ *xdp_data = xdp.data;
+ *xdp_len = xdp.data_end - xdp.data;
+
+ switch (act) {
+ case XDP_PASS:
+ return FDMA_PASS;
+ case XDP_TX:
+ return lan966x_fdma_pci_xmit_xdpf(port, *xdp_data, *xdp_len) ?
+ FDMA_DROP : FDMA_TX;
+ default:
+ bpf_warn_invalid_xdp_action(port->dev, xdp_prog, act);
+ fallthrough;
+ case XDP_ABORTED:
+ trace_xdp_exception(port->dev, xdp_prog, act);
+ fallthrough;
+ case XDP_DROP:
+ return FDMA_DROP;
+ }
+}
+
+static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port,
+ void **data, u32 *data_len)
{
struct lan966x *lan966x = rx->lan966x;
struct fdma *fdma = &rx->fdma;
@@ -171,7 +288,15 @@ static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
return FDMA_ERROR;
- return FDMA_PASS;
+ /* Present the Ethernet frame (no IFH, no FCS). HW re-inserts the
+ * FCS on TX; see lan966x_fdma_pci_xmit_xdpf(). May be overridden
+ * by XDP. The FCS strip is unconditional because NETIF_F_RXFCS
+ * is not advertised in hw_features.
+ */
+ *data = virt_addr + IFH_LEN_BYTES;
+ *data_len = blockl - IFH_LEN_BYTES - ETH_FCS_LEN;
+
+ return lan966x_xdp_pci_run(port, virt_addr, *data_len, data, data_len);
}
static bool lan966x_fdma_pci_hw_offload(struct lan966x *lan966x, u32 port,
@@ -206,34 +331,21 @@ static bool lan966x_fdma_pci_hw_offload(struct lan966x *lan966x, u32 port,
}
static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
- u64 src_port)
+ u64 src_port, void *data,
+ u32 data_len)
{
struct lan966x *lan966x = rx->lan966x;
- struct fdma *fdma = &rx->fdma;
struct sk_buff *skb;
- struct fdma_db *db;
- u32 data_len;
-
- /* Get the received frame and create an SKB for it. */
- db = fdma_db_next_get(fdma);
- data_len = fdma_db_len_get(db);
skb = napi_alloc_skb(&lan966x->napi, data_len);
if (unlikely(!skb))
return NULL;
- memcpy(skb->data,
- fdma_dataptr_virt_addr_contiguous(fdma,
- fdma->dcb_index,
- fdma->db_index),
- data_len);
+ memcpy(skb->data, data, data_len);
skb_put(skb, data_len);
skb->dev = lan966x->ports[src_port]->dev;
- skb_pull(skb, IFH_LEN_BYTES);
-
- skb_trim(skb, skb->len - ETH_FCS_LEN);
skb->protocol = eth_type_trans(skb, skb->dev);
@@ -324,6 +436,8 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
struct sk_buff *skb;
int counter = 0;
u64 src_port;
+ u32 data_len;
+ void *data;
/* Wake any stopped TX queues if a TX DCB is available. */
spin_lock(&lan966x->tx_lock);
@@ -340,7 +454,10 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
/* Order DONE read before DCB/frame reads below. */
dma_rmb();
counter++;
- switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
+ switch (lan966x_fdma_pci_rx_check_frame(rx,
+ &src_port,
+ &data,
+ &data_len)) {
case FDMA_PASS:
break;
case FDMA_ERROR:
@@ -349,8 +466,17 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
*/
fdma_dcb_advance(fdma);
continue;
+ case FDMA_TX:
+ fdma_dcb_advance(fdma);
+ continue;
+ case FDMA_DROP:
+ fdma_dcb_advance(fdma);
+ continue;
}
- skb = lan966x_fdma_pci_rx_get_frame(rx, src_port);
+ skb = lan966x_fdma_pci_rx_get_frame(rx,
+ src_port,
+ data,
+ data_len);
fdma_dcb_advance(fdma);
if (!skb) {
lan966x->ports[src_port]->dev->stats.rx_dropped++;
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index c3afc4cc597f..e4c9c5812cca 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -875,11 +875,13 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
port->phylink = phylink;
- /* XDP is not supported on the PCIe FDMA path. */
- if (lan966x->fdma && !lan966x_is_pci(lan966x))
- dev->xdp_features = NETDEV_XDP_ACT_BASIC |
- NETDEV_XDP_ACT_REDIRECT |
- NETDEV_XDP_ACT_NDO_XMIT;
+ if (lan966x->fdma) {
+ dev->xdp_features = NETDEV_XDP_ACT_BASIC;
+
+ if (!lan966x_is_pci(lan966x))
+ dev->xdp_features |= NETDEV_XDP_ACT_REDIRECT |
+ NETDEV_XDP_ACT_NDO_XMIT;
+ }
err = register_netdev(dev);
if (err) {
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
index e63634887f64..b98426afc785 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
@@ -20,16 +20,16 @@ static int lan966x_xdp_setup(struct net_device *dev, struct netdev_bpf *xdp)
return -EOPNOTSUPP;
}
- if (lan966x_is_pci(lan966x)) {
- NL_SET_ERR_MSG_MOD(xdp->extack,
- "XDP is not supported on the PCIe FDMA path");
- return -EOPNOTSUPP;
- }
-
old_xdp = lan966x_xdp_present(lan966x);
old_prog = xchg(&port->xdp_prog, xdp->prog);
new_xdp = lan966x_xdp_present(lan966x);
+ /* PCIe FDMA uses contiguous buffers, so no page_pool reload
+ * is needed.
+ */
+ if (lan966x_is_pci(lan966x))
+ goto out;
+
if (old_xdp == new_xdp)
goto out;
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (12 preceding siblings ...)
2026-09-28 19:33 ` [PATCH net-next v9 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
@ 2026-09-28 19:33 ` Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:33 ` [PATCH net-next v9 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
14 siblings, 1 reply; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:33 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
The ATU outbound windows used by the FDMA engine are programmed through
registers at offset 0x400000+, which falls outside the current cpu reg
mapping. Extend the cpu reg size from 0x100000 (1MB) to 0x800000 (8MB)
to cover the full PCIE DBI and iATU register space.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
drivers/misc/lan966x_pci.dtso | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
index 7b196b0a0eb6..7bb726550caf 100644
--- a/drivers/misc/lan966x_pci.dtso
+++ b/drivers/misc/lan966x_pci.dtso
@@ -135,7 +135,7 @@ lan966x_phy1: ethernet-lan966x_phy@2 {
switch: switch@e0000000 {
compatible = "microchip,lan966x-switch";
- reg = <0xe0000000 0x0100000>,
+ reg = <0xe0000000 0x0800000>,
<0xe2000000 0x0800000>;
reg-names = "cpu", "gcb";
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v9 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
` (13 preceding siblings ...)
2026-09-28 19:33 ` [PATCH net-next v9 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
@ 2026-09-28 19:33 ` Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
14 siblings, 1 reply; 29+ messages in thread
From: Daniel Machon @ 2026-09-28 19:33 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel
Add the fdma interrupt (OIC interrupt 14) to the lan966x PCI device
tree overlay, enabling FDMA-based frame injection/extraction when
the switch is connected over PCIe.
Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
drivers/misc/lan966x_pci.dtso | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
index 7bb726550caf..5bb12dbc0843 100644
--- a/drivers/misc/lan966x_pci.dtso
+++ b/drivers/misc/lan966x_pci.dtso
@@ -141,8 +141,9 @@ switch: switch@e0000000 {
interrupt-parent = <&oic>;
interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
+ <14 IRQ_TYPE_LEVEL_HIGH>,
<9 IRQ_TYPE_LEVEL_HIGH>;
- interrupt-names = "xtr", "ana";
+ interrupt-names = "xtr", "fdma", "ana";
resets = <&reset 0>;
reset-names = "switch";
--
2.34.1
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 03/15] net: microchip: fdma: add PCIe ATU support
2026-09-28 19:32 ` [PATCH net-next v9 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
@ 2026-10-02 7:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 7:33 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] fdma_pci_atu_region_map() in
drivers/net/ethernet/microchip/fdma/fdma_pci.c is exported.
- [Low] The alignment comment in fdma_pci_atu_region_map()
(drivers/net/ethernet/microchip/fdma/fdma_pci.c, lines 151-155) says…
- [Low] fdma_pci_atu_region_enable() writes REGION_CTRL2 = BIT(31) with a
single posted writel().
- [Low] fdma_pci_atu_configure_translation() programs LOWER/UPPER_BASE,
LIMIT and LOWER/UPPER_TARGET, and fdma_pci_atu_init() only clears…
--
net: microchip: fdma: add PCIe ATU support
This commit adds PCIe ATU outbound window support to the Microchip FDMA
library. When the switch runs as a PCIe endpoint, the lan966x (and later
lan969x) FDMA engine can then reach host memory. The outbound space is
split into six regions, and new helpers combine a coherent DMA allocation
with claiming and programming a region, padding the allocation to the
64KB ATU granularity.
> diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.c b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
> new file mode 100644
> index 000000000000..dd1dc46cbc9d
> --- /dev/null
> +++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
[ ... ]
> +/* Configure the address translation in the ATU. */
> +static void
> +fdma_pci_atu_configure_translation(struct fdma_pci_atu_region *region)
> +{
> + struct fdma_pci_atu *atu = region->atu;
> + int idx = region->idx;
> +
> + writel(lower_32_bits(region->base_addr),
> + fdma_pci_atu_addr_get(atu->addr,
> + FDMA_PCI_ATU_LWR_BASE_ADDR, idx));
> +
> + writel(upper_32_bits(region->base_addr),
> + fdma_pci_atu_addr_get(atu->addr,
> + FDMA_PCI_ATU_UPP_BASE_ADDR, idx));
> +
> + /* The OB address space lies entirely below 4GB, so the limit always
> + * fits the lower limit register and the upper one is left alone.
> + */
> + writel(region->limit_addr,
> + fdma_pci_atu_addr_get(atu->addr, FDMA_PCI_ATU_LIMIT_ADDR, idx));
[Severity: Low]
Should REGION_CTRL1 (offset 0x0) and the upper limit register (offset
0x20) be written explicitly, either here or in fdma_pci_atu_init()?
Neither fdma_pci_atu_configure_translation() nor fdma_pci_atu_init()
writes them. fdma_pci_atu_init() only clears the enable bit through
fdma_pci_atu_region_disable(). As a result, the TLP type, function
number, INCREASE_REGION_SIZE and the upper limit all depend on reset
values.
For comparison, dw_pcie_prog_outbound_atu() in
drivers/pci/controller/dwc/pcie-designware.c always writes CTRL1. On
4.60a and later it also writes the upper limit:
if (dw_pcie_ver_is_ge(pci, 460A))
dw_pcie_writel_atu_ob(pci, atu->index, PCIE_ATU_UPPER_LIMIT,
upper_32_bits(limit_addr));
...
dw_pcie_writel_atu_ob(pci, atu->index, PCIE_ATU_REGION_CTRL1, val);
This driver never changes those registers, so reloading it can't leave
them stale. It would only matter if something else, such as endpoint-side
firmware, had programmed the ATU before the host driver bound. No in-tree
path doing that was found.
Would writing CTRL1 = 0 at init be reasonable hardening?
[ ... ]
> +/* Map a host DMA address into a free outbound region. */
> +struct fdma_pci_atu_region *
> +fdma_pci_atu_region_map(struct fdma_pci_atu *atu, u64 target_addr, int size)
> +{
[ ... ]
> + /* The ATU region base is only ever aligned to FDMA_PCI_ATU_REGION_ALIGN;
> + * require the same alignment of the host target address, since the ATU
> + * translates addr - target_addr + base_addr and any misalignment here
> + * would shift every translated address by the same amount.
> + */
[Severity: Low]
Is the translation direction in this comment reversed?
In struct fdma_pci_atu_region, base_addr is the OB window address and
target_addr is the host DMA address. The file header says the ATU turns
FDMA OB addresses into host DMA addresses, so the hardware computes
target_addr + (addr - base_addr).
The formula in the comment is the host to OB direction. That is what the
software helper fdma_pci_atu_translate_addr() computes:
return region->base_addr + (addr - region->target_addr);
This does not affect the code's behaviour. Could the comment instead
describe the OB to host translation that the ATU performs?
> + if (!IS_ALIGNED(target_addr, FDMA_PCI_ATU_REGION_ALIGN))
> + return ERR_PTR(-EINVAL);
[Severity: Low]
Should size also be checked for FDMA_PCI_ATU_REGION_ALIGN alignment here?
The prototype comment in fdma_pci.h says:
/* Map a host DMA address into a free ATU region. target_addr and size must be
* FDMA_PCI_ATU_REGION_ALIGN aligned; a misaligned target_addr returns -EINVAL.
*/
However, fdma_pci_atu_region_map() only rejects a misaligned target_addr.
Any size in the range 0 < size <= fdma_pci_atu_region_size() is accepted
and programmed directly:
region->limit_addr = region->base_addr + size - 1;
The commit message says the hardware widens the limit to the 64KB
boundary. The padding that keeps the widened window inside owned memory
is applied only in fdma_alloc_coherent_and_map():
fdma->size = ALIGN(fdma->size, FDMA_PCI_ATU_REGION_ALIGN);
That wrapper is the only caller in the series, so no current path passes
an unaligned size. A future direct caller of the exported
fdma_pci_atu_region_map() could, though, and that would open a window
past the end of its buffer.
Would an IS_ALIGNED(size, FDMA_PCI_ATU_REGION_ALIGN) check next to the
target_addr check make the function match its documented contract?
> +
> + mutex_lock(&atu->lock);
[ ... ]
> + region->target_addr = target_addr;
> + region->limit_addr = region->base_addr + size - 1;
> + region->in_use = true;
> +
> + fdma_pci_atu_configure_translation(region);
> + fdma_pci_atu_region_enable(region);
[Severity: Low]
Should the enable bit be read back before fdma_pci_atu_region_map()
returns success?
fdma_pci_atu_region_enable() is a single posted writel() of BIT(31) to
REGION_CTRL2, and nothing checks whether it took effect.
dw_pcie_prog_outbound_atu() programs the same register layout. It polls
for PCIE_ATU_ENABLE and returns -ETIMEDOUT if the bit never appears:
/*
* Make sure ATU enable takes effect before any subsequent config
* and I/O accesses.
*/
Unmap is similar. fdma_pci_atu_region_unmap() posts the disable with no
flushing read, and fdma_free_coherent_and_unmap() then frees the buffer
right away through fdma_free_coherent().
In the lan966x user, fdma_pci_atu_region_map() is followed by
non-posted reads through the same BAR before FDMA_CH_ACTIVATE, and the
channels are polled until inactive before every unmap. That probably
hides this in practice.
Would a readback after the enable and the disable writes make the
ordering explicit?
> +
> + mutex_unlock(&atu->lock);
> +
> + return region;
> +}
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot
2026-09-28 19:32 ` [PATCH net-next v9 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
@ 2026-10-02 7:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 7:33 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] Concern: lan966x_shutdown() disables lan966x->napi and then
releases rtnl.
- [Low] Concern: lan966x_fdma_rx_disable() and lan966x_fdma_tx_disable()
ignore the return value of readx_poll_timeout_atomic() on…
- [Low] Concern: after lan966x_shutdown(), a devmap or bpf_redirect from
another NIC could reach lan966x_xdp_xmit() -> lan966x_fdma_xmit_xdpf()…
--
net: lan966x: add shutdown callback to stop the FDMA on reboot
This commit adds a platform shutdown callback for the PCIe instance of
lan966x. The callback frees the irqs, masks the analyzer interrupt,
disables NAPI, detaches the netdevs, stops both FDMA channels and unmaps
the outbound ATU windows. It also restores the full FDMA_INTR_ENA enable
in lan966x_fdma_rx_start(), because that register persists across a warm
reboot.
> NAPI is skipped when fdma_ndev is unset (a probed switch with no
> usable port never adds one), and XDP attach cannot re-enter the FDMA
> either: on PCIe, lan966x_xdp_setup() returns before the page pool
> reload, which is the only point where it touches the FDMA.
[Severity: Low]
Is this accurate at this point in the series? At this commit,
lan966x_xdp_setup() checks only lan966x->fdma before the reload:
lan966x_xdp.c:lan966x_xdp_setup() {
if (!lan966x->fdma) {
NL_SET_ERR_MSG_MOD(xdp->extack,
"Allow to set xdp only when using fdma");
return -EOPNOTSUPP;
}
...
err = lan966x_fdma_reload_page_pool(lan966x);
...
}
Suppose lan966x_shutdown() has called napi_disable() and dropped rtnl.
An XDP attach or detach would then reach:
lan966x_xdp_setup()
lan966x_fdma_reload_page_pool()
lan966x_fdma_reload()
napi_synchronize(&lan966x->napi);
napi_disable(&lan966x->napi);
That would hang on the already disabled NAPI with rtnl held. It could
also re-arm RX through lan966x_fdma_rx_start().
In practice this doesn't look reachable at this revision.
lan966x_probe() sets lan966x->fdma only when an "fdma" interrupt exists.
The in-tree lan966x_pci.dtso overlay doesn't declare one until the last
patch of the series ("misc: lan966x-pci: dts: add fdma interrupt to
overlay"). So lan966x_shutdown() returns before napi_disable(), and XDP
attach returns -EOPNOTSUPP.
The lan966x_is_pci() early return in lan966x_xdp_setup() is added later
by "net: lan966x: add PCIe FDMA support" and "net: lan966x: add PCIe
FDMA XDP support". Could the commit message describe the state at this
revision, or say that the early return is added later in the series?
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index 259d81e75907..024ce9f9916c 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -1324,9 +1324,65 @@ static void lan966x_remove(struct platform_device *pdev)
> debugfs_remove_recursive(lan966x->debugfs_root);
> }
>
> +static void lan966x_shutdown(struct platform_device *pdev)
> +{
> + struct lan966x *lan966x = platform_get_drvdata(pdev);
> +
[ ... ]
> + if (!lan966x->fdma)
> + return;
> +
> + rtnl_lock();
> +
> + if (lan966x->fdma_ndev)
> + napi_disable(&lan966x->napi);
> +
> + for (int p = 0; p < lan966x->num_phys_ports; p++) {
> + if (!lan966x->ports[p] || !lan966x->ports[p]->dev)
> + continue;
> +
> + netif_tx_disable(lan966x->ports[p]->dev);
> + netif_device_detach(lan966x->ports[p]->dev);
> + }
> +
> + lan966x_fdma_rx_disable(&lan966x->rx);
> + lan966x_fdma_tx_disable(&lan966x->tx);
[Severity: Low]
Can anything re-arm the TX channel after this point?
lan966x_fdma_tx_disable() clears tx->activated without taking tx_lock.
A later devmap or bpf_redirect from another NIC would then go:
lan966x_xdp_xmit()
lan966x_fdma_xmit_xdpf()
lan966x_fdma_tx_start()
lan966x->tx.activated = true;
lan966x_fdma_tx_activate(tx);
As far as I can tell this path is blocked:
- At this commit, PCIe instances have lan966x->fdma == false with the
in-tree overlay, so shutdown never reaches this code.
- At the end of the series, lan966x_probe_port() adds
NETDEV_XDP_ACT_REDIRECT and NETDEV_XDP_ACT_NDO_XMIT only when
!lan966x_is_pci(). __xdp_enqueue() in kernel/bpf/devmap.c rejects
devices that lack NETDEV_XDP_ACT_NDO_XMIT.
- Regular xmit is drained by netif_tx_disable() above. The PCIe XDP_TX
path runs only from NAPI, which is disabled first.
Is the xdp_features gate the intended guarantee that nothing re-arms
this channel after shutdown?
> +
> + lan_wr(0, lan966x, FDMA_INTR_ENA);
> + lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
> +
> +#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
> + fdma_pci_atu_region_unmap(lan966x->rx.fdma.atu_region);
> + fdma_pci_atu_region_unmap(lan966x->tx.fdma.atu_region);
> +#endif
[Severity: Low]
This isn't a bug, but lan966x_fdma_rx_disable() and
lan966x_fdma_tx_disable() both discard the result of the FDMA_CH_ACTIVE
poll:
readx_poll_timeout_atomic(lan966x_fdma_channel_active, lan966x,
val, !(val & BIT(fdma->channel_id)),
READL_SLEEP_US, READL_TIMEOUT_US);
So the ATU windows are unmapped here even if a channel never went idle,
and nothing is logged. READL_TIMEOUT_US is 100 s, so a timeout means the
engine is stuck. Unmapping the windows is still the right step in that
case.
A stale FDMA error carried into the next kernel is not a concern either.
lan966x_reset_switch() clears FDMA_ERRORS, FDMA_INTR_ERR and
FDMA_INTR_DB before lan966x_probe() requests the fdma irq.
Would a warning on timeout help make a stuck FDMA visible at shutdown?
The disable helpers themselves predate this patch.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support
2026-09-28 19:32 ` [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
@ 2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:02 ` Daniel Machon
2026-10-02 14:14 ` Simon Horman
1 sibling, 1 reply; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 7:33 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 8 potential
issue(s) to consider.
Critical: 0 · High: 0 (1) · Medium: 3 (1) · Low: 5 (1)
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] Double fetch of the RX BLOCKL.
- [Medium] lan966x_fdma_pci_napi_poll() limits RX consumption only by
`counter < weight` and tracks refill progress only through the wrapped…
- [Medium] When lan966x_fdma_pci_get_next_dcb() finds no free TX DCB,
lan966x_fdma_pci_xmit() calls netif_stop_queue(dev) and returns…
- [Low] At this commit the in-tree PCIe overlay
(drivers/misc/lan966x_pci.dtso) has no "fdma" interrupt.
- [Low] The iomap entry { TARGET_PCIE_DBI, 0x400000, 0 } derives
regs[TARGET_PCIE_DBI] 4 MB into IO range 0.
- [Low] lan966x_fdma_pci_resize() always returns -EOPNOTSUPP, so once
PCIe FDMA is active every ndo_change_mtu call is rolled back.
- [Low] In lan966x_fdma_pci_napi_poll(), the FDMA_ERROR branch says "No
rx_dropped increment here because src_port is invalid".
- [Low] With lan966x->ptp left at 0 on PCIe, lan966x_ptp_init() never
initialises port->tx_skbs or the PTP locks, while the "ptp" IRQ…
Pre-existing issues:
- [High] This problem wasn't introduced by this patch, but the new PCIe
backend inherits it.
- [Medium] __fdma_dcb_add() writes fdma->last_dcb->nextptr (publishing
the new DCB) before it writes the new DCB's own nextptr =…
- [Low] lan966x_fdma_pci_hw_offload() reads ip_hdr(skb)->protocol and
ipv6_hdr(skb)->daddr without checking that skb->len covers an IP…
--
net: lan966x: add PCIe FDMA support
This commit adds a PCIe FDMA backend for lan966x. It uses contiguous
coherent DMA buffers mapped through the endpoint's ATU and copies frames
with memcpy. The new ops are selected when lan966x_is_pci() is true, and
XDP and PTP are disabled on the PCIe path.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> new file mode 100644
> index 000000000000..f511e7061314
> --- /dev/null
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
[ ... ]
> +static bool lan966x_fdma_pci_hw_offload(struct lan966x *lan966x, u32 port,
> + struct sk_buff **pskb)
> +{
[ ... ]
> + if (skb->protocol == htons(ETH_P_IP) &&
> + ip_hdr(skb)->protocol == IPPROTO_IGMP)
> + return false;
> +
> + if (IS_ENABLED(CONFIG_IPV6) &&
> + skb->protocol == htons(ETH_P_IPV6) &&
> + ipv6_addr_is_multicast(&ipv6_hdr(skb)->daddr) &&
> + !ipv6_mc_check_mld(skb))
> + return false;
[Severity: Low]
This isn't a bug introduced by this patch, because the same code already
exists in lan966x_hw_offload() in lan966x_main.c. Still,
ip_hdr(skb)->protocol and ipv6_hdr(skb)->daddr are read here without
checking that skb->len covers an IP header.
On a runt frame, or on a frame shrunk by an XDP program once XDP support
lands later in the series, could this read uninitialized skb tailroom?
The reads stay inside the skb head allocation, and any minimum-size
Ethernet frame covers both fields. So only the offload_fwd_mark decision
for an already malformed packet is affected.
> +static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
> + u64 src_port)
> +{
[ ... ]
> + /* Get the received frame and create an SKB for it. */
> + db = fdma_db_next_get(fdma);
> + data_len = fdma_db_len_get(db);
> +
> + skb = napi_alloc_skb(&lan966x->napi, data_len);
> + if (unlikely(!skb))
> + return NULL;
> +
> + memcpy(skb->data,
> + fdma_dataptr_virt_addr_contiguous(fdma,
> + fdma->dcb_index,
> + fdma->db_index),
> + data_len);
[Severity: Medium]
BLOCKL is validated in lan966x_fdma_pci_rx_check_frame() through
lan966x_fdma_pci_rx_size_fits(). Here it is read again from the DCB
status in coherent DMA memory.
Can the value passed to napi_alloc_skb() and memcpy() differ from the
value that was checked?
It looks like "net: lan966x: add PCIe FDMA XDP support" later in the
series fixes this. That patch reads blockl once in rx_check_frame() and
passes data and data_len to rx_get_frame(). Would it make sense to fold
that change into this patch?
At this commit the path isn't enabled in-tree yet, because the overlay
has no "fdma" interrupt.
[ ... ]
> +static int lan966x_fdma_pci_xmit(struct sk_buff *skb, __be32 *ifh,
> + struct net_device *dev)
> +{
[ ... ]
> + next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
> +
> + if (next_to_use < 0) {
> + netif_stop_queue(dev);
> + return NETDEV_TX_BUSY;
> + }
[Severity: Medium]
netif_stop_queue() only stops TX queue 0. Each port netdev is created in
lan966x_probe_port() with 8 TX queues:
dev = devm_alloc_etherdev_mqs(lan966x->dev,
sizeof(struct lan966x_port),
NUM_PRIO_QUEUES, 1);
There is no ndo_select_queue, so skbs are spread over all 8 queues.
lan966x_fdma_wakeup_netdev(), called from the PCIe NAPI poll, also only
checks and wakes queue 0.
When the shared TX ring is full and an skb arrives on one of queues 1 to
7, does that queue ever get stopped?
It looks like sch_direct_xmit() would requeue the skb and reschedule the
qdisc. net_tx_action would then keep retrying, taking tx_lock and
scanning the whole DCB ring each time, until the hardware completes a
DCB.
Would netif_tx_stop_all_queues() and netif_tx_wake_all_queues() be a
better fit here? The platform lan966x_fdma_xmit() has the same pattern,
and lan966x_fdma_pci_xmit_xdpf() from "net: lan966x: add PCIe FDMA XDP
support" repeats it.
[ ... ]
> + /* Order frame write before DCB status write below. */
> + dma_wmb();
> +
> + fdma_dcb_add(fdma,
> + next_to_use,
> + 0,
> + FDMA_DCB_STATUS_INTR |
> + FDMA_DCB_STATUS_SOF |
> + FDMA_DCB_STATUS_EOF |
> + FDMA_DCB_STATUS_BLOCKO(0) |
> + FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + skb->len + ETH_FCS_LEN));
[Severity: Medium]
This is a pre-existing issue in the shared fdma_api.c helper and was not
introduced by this patch, but the new backend depends on it.
__fdma_dcb_add() links the new DCB into the chain before terminating it,
and there is no DMA write barrier between the two steps:
drivers/net/ethernet/microchip/fdma/fdma_api.c:__fdma_dcb_add() {
...
fdma->last_dcb->nextptr = cpu_to_le64(nextptr);
fdma->last_dcb = dcb;
dcb->nextptr = cpu_to_le64(FDMA_DCB_INVALID_DATA);
dcb->info = cpu_to_le64(info);
...
}
If the FDMA channel is still walking the chain, could it follow the new
link and read a stale nextptr or info from the new DCB?
The window is short, and whether the hardware acts on it depends on its
prefetch behavior. When the channel is stopped, the writel() doorbell in
lan966x_fdma_tx_start() or lan966x_fdma_rx_reload() orders the earlier
stores.
[ ... ]
> +static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
> +{
[ ... ]
> + dcb_reload = fdma->dcb_index;
> +
> + /* Get all received skbs. */
> + while (counter < weight) {
> + if (!fdma_has_frames(fdma))
> + break;
> + /* Order DONE read before DCB/frame reads below. */
> + dma_rmb();
> + counter++;
> + switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
> + case FDMA_PASS:
> + break;
> + case FDMA_ERROR:
> + /* No rx_dropped increment here because src_port is
> + * invalid.
> + */
[Severity: Low]
Is this comment accurate for every FDMA_ERROR return?
lan966x_fdma_pci_rx_check_frame() also returns FDMA_ERROR after src_port
and ports[src_port] have already been validated:
blockl = fdma_db_len_get(db);
if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
return FDMA_ERROR;
In that case a frame with a bad BLOCKL on a valid port is dropped without
being counted in rx_dropped or rx_length_errors. This code is unchanged
at the end of the series.
> + fdma_dcb_advance(fdma);
> + continue;
> + }
[ ... ]
> + while (dcb_reload != fdma->dcb_index) {
> + old_dcb = dcb_reload;
> + dcb_reload++;
> + dcb_reload &= fdma->n_dcbs - 1;
[Severity: Medium]
The RX loop is bounded only by counter < weight. Refill progress is
tracked only through the wrapped dcb_reload and dcb_index values.
What happens if weight >= n_dcbs and every DCB is DONE? n_dcbs is
FDMA_DCB_MAX, which is 512 here and 256 after "net: lan966x: add PCIe
FDMA MTU change support".
fdma_dcb_advance() would wrap dcb_index back to dcb_reload. DONE bits are
only cleared by this refill loop, so the RX loop would then reprocess
stale DONE descriptors and pass the same frames to napi_gro_receive()
again.
When weight is a multiple of n_dcbs, dcb_reload equals fdma->dcb_index at
this point. The refill loop then doesn't run at all, and the ring is left
un-armed.
The normal NAPI weight of 64 can't reach this, but busy polling can.
__napi_busy_loop() passes the socket's budget straight to napi->poll, and
SO_BUSY_POLL_BUDGET accepts values up to U16_MAX with CAP_NET_ADMIN.
The native lan966x_fdma_napi_poll() has the same structure. This is
still present at the end of the series.
[ ... ]
> +static int lan966x_fdma_pci_init(struct lan966x *lan966x)
> +{
> + struct fdma *rx_fdma = &lan966x->rx.fdma;
> + struct fdma *tx_fdma = &lan966x->tx.fdma;
> + int err;
> +
> + if (!lan966x->fdma)
> + return 0;
[Severity: Low]
At this commit the in-tree PCIe overlay (drivers/misc/lan966x_pci.dtso)
only declares:
interrupt-names = "xtr", "ana";
As a result lan966x->fdma stays false on the PCIe instance, and this
function returns right away. Is it intended that the new backend is
dormant at this commit?
The "fdma" interrupt is added later in the series by "misc: lan966x-pci:
dts: add fdma interrupt to overlay". Until then the driver keeps using
register-based I/O.
[ ... ]
> + lan966x->tx.lan966x = lan966x;
> + tx_fdma->channel_id = FDMA_INJ_CHANNEL;
> + tx_fdma->n_dcbs = FDMA_DCB_MAX;
[Severity: High]
This isn't a bug introduced by this patch, because the platform
lan966x_fdma_init() uses the same ordering, but the new PCIe backend
inherits it.
lan966x_probe() sets lan966x->fdma = true, registers every port netdev,
and enables NAPI (lan966x_port_init() -> lan966x_fdma_netdev_init()), all
before it calls lan966x->ops->fdma_init().
Here tx_fdma->n_dcbs is set without holding tx_lock. tx_fdma->dcbs is
only allocated later, in lan966x_fdma_pci_tx_alloc(). That happens after
the GFP_KERNEL coherent allocation in lan966x_fdma_pci_rx_alloc(), which
can sleep.
If a port is opened and transmits during that window, can the following
happen?
lan966x_port_xmit()
spin_lock(&lan966x->tx_lock);
if (port->lan966x->fdma)
lan966x->ops->fdma_xmit()
lan966x_fdma_pci_xmit()
lan966x_fdma_pci_get_next_dcb()
fdma_db_get(fdma, i, 0) <- fdma->dcbs is NULL
Later in the same window, xmit could also pick DCBs that
fdma_dcbs_init() has already marked DONE while last_dcb is still being
set up. It could also start the injection channel before
lan966x_fdma_llp_configure() has programmed the TX LLP.
Would either of these close the window?
- Calling fdma_init before the ports are registered.
- Publishing n_dcbs under tx_lock, only after the rings are allocated.
[ ... ]
> +static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
> +{
> + return -EOPNOTSUPP;
> +}
[Severity: Low]
lan966x_port_change_mtu() rolls back the MTU whenever
lan966x->ops->fdma_resize() fails. With this stub, won't every
ndo_change_mtu call fail once PCIe FDMA is active?
The next patch, "net: lan966x: add PCIe FDMA MTU change support",
replaces the stub with a real implementation. At this commit the path is
dormant, because the overlay has no "fdma" interrupt.
[ ... ]
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index 024ce9f9916c..de2202786826 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -50,6 +50,7 @@ struct lan966x_main_io_resource {
> static const struct lan966x_main_io_resource lan966x_main_iomap[] = {
> { TARGET_CPU, 0xc0000, 0 }, /* 0xe00c0000 */
> { TARGET_FDMA, 0xc0400, 0 }, /* 0xe00c0400 */
> + { TARGET_PCIE_DBI, 0x400000, 0 }, /* 0xe0400000 */
[Severity: Low]
This entry places regs[TARGET_PCIE_DBI] 4 MB into IO range 0. At this
commit the overlay's "cpu" reg is only 1 MB:
reg = <0xe0000000 0x0100000>,
Wouldn't the iATU accesses from fdma_pci_atu_init(), and the region map
and unmap, fall outside the ioremap window?
This can't be reached at this commit, because lan966x_fdma_pci_init()
returns early without an "fdma" interrupt. Later in the series, "misc:
lan966x-pci: dts: extend cpu reg to cover PCIE DBI space" enlarges the
reg to 0x0800000, and it comes before the "fdma" interrupt is added. On
the SoC instance the pointer is computed but never used.
[ ... ]
> @@ -1187,7 +1190,9 @@ static int lan966x_probe(struct platform_device *pdev)
> if (err)
> return dev_err_probe(&pdev->dev, err, "Unable to use ptp irq");
>
> - lan966x->ptp = 1;
> + /* PTP is not supported on the PCIe path yet. */
> + if (!lan966x_is_pci(lan966x))
> + lan966x->ptp = 1;
> }
[Severity: Low]
With lan966x->ptp left at 0 on PCIe, lan966x_ptp_init() never
initializes port->tx_skbs or the PTP locks. If a "ptp" interrupt exists,
lan966x_ptp_irq_handler has already been registered just above. Could
that handler then touch uninitialized PTP state?
This can't happen in-tree, because drivers/misc/lan966x_pci.dtso doesn't
declare a "ptp" interrupt. Would it be more robust to also skip the ptp
IRQ request on PCIe?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 12/15] net: lan966x: add PCIe FDMA MTU change support
2026-09-28 19:33 ` [PATCH net-next v9 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
@ 2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:08 ` Daniel Machon
0 siblings, 1 reply; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 7:33 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 1 (1) · Medium: 0 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] lan966x_fdma_pci_resize() uses `lan966x->rx.lan966x` to decide
whether the rings exist, and that field does not track ring state.
- [Low] lan966x_fdma_pci_reload() frees and ATU-unmaps the old rings and
reprograms LLP right after lan966x_fdma_rx_disable() and…
Pre-existing issues:
- [High] This problem wasn't introduced by this patch, but the patch
explicitly acknowledges the hazard and fixes only the PCIe side.
--
net: lan966x: add PCIe FDMA MTU change support
This adds MTU change support to the lan966x PCIe FDMA path. On an MTU
change, the contiguous ATU-mapped RX and TX rings are reallocated at the
new size, and the old buffers are kept if allocation fails. The PCIe DCB
ring is capped at 256, FDMA_PCI_MAX_MTU is advertised as max_mtu when the
FDMA is in use, and the resize is skipped until lan966x_fdma_pci_init()
has built the rings.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> index f511e70613141..758554c951c52 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
[ ... ]
> +static int lan966x_fdma_pci_reload(struct lan966x *lan966x, int new_mtu)
> +{
> + struct fdma tx_fdma_old = lan966x->tx.fdma;
> + struct fdma rx_fdma_old = lan966x->rx.fdma;
> + u32 old_mtu = lan966x->rx.max_mtu;
> + int err;
> +
> + napi_disable(&lan966x->napi);
> + lan966x_fdma_tx_disable_netdev(lan966x);
> + lan966x_fdma_rx_disable(&lan966x->rx);
> + lan966x_fdma_tx_disable(&lan966x->tx);
[Severity: Low]
lan966x_fdma_rx_disable() and lan966x_fdma_tx_disable() both throw away
the result of the FDMA_CH_ACTIVE poll:
readx_poll_timeout_atomic(lan966x_fdma_channel_active, lan966x,
val, !(val & BIT(fdma->channel_id)),
READL_SLEEP_US, READL_TIMEOUT_US);
__lan966x_fdma_pci_reload() also ignores the result of the CPU queue drain
poll. After that, the old rings are freed and ATU-unmapped straight away.
If a channel is still active when the poll times out, could it follow the
old DCB chain into memory that has already been unmapped and freed?
The timeout is READL_TIMEOUT_US (100 seconds), so this would need wedged
hardware. The platform lan966x_fdma_reload() uses the same sequence. It
might still be worth checking the poll results and bailing out before
freeing the old rings.
> +
> + lan966x->rx.max_mtu = new_mtu;
[ ... ]
> + /* Free and unmap old memory. */
> + fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
> + fdma_free_coherent_and_unmap(lan966x->dma_dev, &tx_fdma_old);
[ ... ]
> static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
> {
> - return -EOPNOTSUPP;
> + int max_mtu;
> +
> + /* Nothing to resize until fdma_pci_init() has built the rings; it
> + * sizes them from DEV_MAC_MAXLEN_CFG, which the caller already set.
> + */
> + if (!lan966x->rx.lan966x)
> + return 0;
[Severity: High]
Does rx.lan966x actually show whether the rings have been built?
In lan966x_fdma_pci_init() it is set before either ring is allocated or
ATU-mapped:
lan966x->rx.lan966x = lan966x;
lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
...
err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
if (err)
return err;
It is not cleared on the init error paths or in
lan966x_fdma_pci_deinit(). Init and deinit both run without RTNL while the
port netdevs are registered, so ndo_change_mtu can reach this check during
those windows.
First, lan966x_remove() tears down the FDMA before the netdevs are
unregistered:
lan966x->ops->fdma_deinit(lan966x);
lan966x_cleanup_ports(lan966x);
lan966x_fdma_pci_deinit() calls napi_disable() and frees both rings. An
MTU change in that window passes this check and reaches
lan966x_fdma_pci_reload(), which calls napi_disable() a second time.
Would napi_disable_locked() then spin forever on NAPIF_STATE_SCHED |
NAPIF_STATE_NPSVC with RTNL held? That would leave unregister_netdev() and
every other RTNL user blocked. The probe unwind at cleanup_fdma has the
same ordering.
Second, if lan966x_fdma_pci_tx_alloc() fails in init, the RX ring is freed.
However, rx.fdma.dcbs still points at it and rx.lan966x stays set. The
probe unwind jumps to cleanup_ptp and skips fdma_deinit, so the netdevs
stay registered until lan966x_cleanup_ports(). An MTU change in that
window runs lan966x_fdma_pci_reload(), which calls:
fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
Can this free the stale RX buffer a second time? The newly allocated,
ATU-mapped rings would then leak while the FDMA runs on them.
The same thing looks possible when the ATU map fails inside
lan966x_fdma_pci_rx_alloc(), because fdma_alloc_coherent_and_map() frees
dcbs without clearing the pointer.
Third, lan966x_probe() registers every port netdev in lan966x_probe_port()
before ops->fdma_init() runs. A concurrent MTU change, for example udev
applying MTUBytes, could pass this check while init is still sleeping in
dma_alloc_coherent() inside rx_alloc or tx_alloc.
Reload would then snapshot a half-built struct fdma, clear atu_region, and
rewrite db_size and size while init is still filling them in. Could that
leak or double-map ATU regions, or leave LLP pointing at freed memory?
The commit message says:
Skip the resize until lan966x_fdma_pci_init() has built the rings;
The check does not seem to do that. The resize, init and deinit paths are
also unchanged at the end of the series.
> +
> + max_mtu = lan966x_fdma_get_max_frame(lan966x);
> + if (max_mtu == lan966x->rx.max_mtu)
> + return 0;
> +
> + return __lan966x_fdma_pci_reload(lan966x, max_mtu);
> }
[Severity: High]
This isn't a bug introduced by this patch, but the platform backend has
the same pre-init window, and the guard added here only covers PCIe.
lan966x_fdma_change_mtu() in lan966x_fdma.c has no equivalent check:
max_mtu = lan966x_fdma_get_max_frame(lan966x);
if (max_mtu == lan966x->rx.max_mtu)
return 0;
return __lan966x_fdma_reload(lan966x, max_mtu);
rx.max_mtu is still 0 before lan966x_fdma_init() runs. So an MTU change
after register_netdev() in lan966x_probe_port() always continues down:
lan966x_port_change_mtu()->lan966x_fdma_change_mtu()->
__lan966x_fdma_reload()->lan966x_fdma_reload()
>From there, lan966x_fdma_rx_disable() does:
struct lan966x *lan966x = rx->lan966x;
...
lan_rmw(FDMA_CH_DISABLE_CH_DISABLE_SET(BIT(fdma->channel_id)),
FDMA_CH_DISABLE_CH_DISABLE,
lan966x, FDMA_CH_DISABLE);
rx->lan966x is still NULL at that point. Can this oops under RTNL?
Would a check in the common caller, or an equivalent guard in
lan966x_fdma_change_mtu(), cover both backends?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 13/15] net: lan966x: add PCIe FDMA XDP support
2026-09-28 19:33 ` [PATCH net-next v9 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
@ 2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:11 ` Daniel Machon
0 siblings, 1 reply; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 7:33 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] This patch removes the PCIe rejection from lan966x_xdp_setup()
(drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c), so a bond can…
- [Low] In lan966x_xdp_pci_run(), xdp_init_buff() is called with frame_sz
= fdma->db_size, and hard_start is the slot start (virt_addr -…
--
net: lan966x: add PCIe FDMA XDP support
This commit adds XDP support to the lan966x PCIe FDMA path. The IFH and
FCS are stripped before the program runs, and XDP_TX is done by copying
the frame into a TX slot behind a newly built IFH. lan966x_xdp_setup()
no longer rejects PCIe and skips the page_pool reload on that path.
> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> index 758554c951c52..949994874ed96 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
[ ... ]
> @@ -139,7 +140,123 @@ static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
> blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
> }
>
[ ... ]
> +static int lan966x_xdp_pci_run(struct lan966x_port *port, void *data,
> + u32 data_len, void **xdp_data, u32 *xdp_len)
> +{
[ ... ]
> + xdp_init_buff(&xdp, fdma->db_size, &port->xdp_rxq);
> +
> + /* hard_start is set to slot start (virt_addr is XDP_PACKET_HEADROOM
> + * into the slot). Headroom includes the IFH; BPF may grow into it
> + * via adjust_head. IFH is rebuilt on XDP_TX and unread on XDP_PASS.
> + */
> + xdp_prepare_buff(&xdp,
> + data - XDP_PACKET_HEADROOM,
> + XDP_PACKET_HEADROOM + IFH_LEN_BYTES,
> + data_len,
> + false);
[Severity: Low]
Can data_end end up past xdp_data_hard_end() here?
frame_sz is the full db_size and hard_start is the start of the slot, so
the XDP core computes the data limit as:
include/net/xdp.h:xdp_data_hard_end()
((xdp)->data_hard_start + (xdp)->frame_sz -
SKB_DATA_ALIGN(sizeof(struct skb_shared_info)))
The upper bound in lan966x_fdma_pci_rx_size_fits(), shown in the context
above, only reserves the headroom:
blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
lan966x_fdma_pci_rx_check_frame() sets data to slot + XDP_PACKET_HEADROOM
+ IFH_LEN_BYTES and data_len to blockl - IFH_LEN_BYTES - ETH_FCS_LEN. At
the largest accepted blockl, data_end is then slot + db_size -
ETH_FCS_LEN, roughly 316 bytes past xdp_data_hard_end().
For such a buffer, bpf_xdp_adjust_tail() returns -EINVAL for any offset,
shrinks included, because it checks this first:
if (unlikely(data_end > data_hard_end))
return -EINVAL;
Reaching this needs a blockl that eats into the shinfo reserve, for
example a malformed BLOCKL. It could also be an oversized frame truncated
to DATAL while lan966x_port_change_mtu() has raised MAXLEN but
fdma_resize() has not yet reallocated the rings.
Memory safety is not affected, since data_end stays inside the slot the
driver owns.
Should lan966x_fdma_pci_rx_size_fits() also reserve
SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) so it matches the frame_sz
given to xdp_init_buff()?
> +
> + act = bpf_prog_run_xdp(xdp_prog, &xdp);
> +
> + *xdp_data = xdp.data;
> + *xdp_len = xdp.data_end - xdp.data;
> +
> + switch (act) {
> + case XDP_PASS:
> + return FDMA_PASS;
> + case XDP_TX:
> + return lan966x_fdma_pci_xmit_xdpf(port, *xdp_data, *xdp_len) ?
> + FDMA_DROP : FDMA_TX;
> + default:
> + bpf_warn_invalid_xdp_action(port->dev, xdp_prog, act);
> + fallthrough;
> + case XDP_ABORTED:
> + trace_xdp_exception(port->dev, xdp_prog, act);
> + fallthrough;
> + case XDP_DROP:
> + return FDMA_DROP;
> + }
> +}
[Severity: Medium]
Now that lan966x_xdp_setup() no longer rejects PCIe, can a bond attach a
program to lan966x PCIe ports and then have its XDP_TX frames silently
dropped here?
The ports share lan966x_port_netdev_ops, which provides both .ndo_bpf and
.ndo_xdp_xmit. That is all bond_xdp_set() checks, so it accepts these
ports. It then pushes the program down:
bond_xdp_set()
dev_xdp_propagate()
lan966x_xdp_setup() <- now succeeds on PCIe
bond_xdp_set() also calls
static_branch_inc(&bpf_master_redirect_enabled_key). After that,
bpf_prog_run_xdp() rewrites XDP_TX on a bond slave:
if (act == XDP_TX && netif_is_bond_slave(xdp->rxq->dev))
act = xdp_master_redirect(xdp);
In round-robin, XOR and 802.3ad modes, xdp_master_redirect() returns
XDP_REDIRECT whenever the bond picks a transmit slave other than the
receiving port. That action falls into the default case above:
bpf_warn_invalid_xdp_action() -> trace_xdp_exception() -> FDMA_DROP
A bond program that only returns XDP_TX attaches without error, but a
hash- or round-robin-dependent share of its packets is dropped. With two
slaves in round-robin, that is about half. Before this patch the attach
failed with -EOPNOTSUPP.
Should the PCIe path handle XDP_REDIRECT, or keep refusing the attach
when the port is a bond slave?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
2026-09-28 19:33 ` [PATCH net-next v9 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
@ 2026-10-02 7:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 7:33 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] At this commit, the switch@e0000000 node in
drivers/misc/lan966x_pci.dtso gets an 8MB "cpu" reg for the PCIe…
--
misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
This commit grows the "cpu" reg region of the switch@e0000000 node in the
lan966x PCI overlay from 1MB to 8MB. The FDMA engine programs its ATU
outbound windows through registers at offset 0x400000 and above, and the
larger mapping covers the PCIe DBI and iATU space that holds them.
> diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> index 7b196b0a0eb65..7bb726550cafd 100644
> --- a/drivers/misc/lan966x_pci.dtso
> +++ b/drivers/misc/lan966x_pci.dtso
> @@ -135,7 +135,7 @@ lan966x_phy1: ethernet-lan966x_phy@2 {
>
> switch: switch@e0000000 {
> compatible = "microchip,lan966x-switch";
> - reg = <0xe0000000 0x0100000>,
> + reg = <0xe0000000 0x0800000>,
> <0xe2000000 0x0800000>;
> reg-names = "cpu", "gcb";
[Severity: Low]
This isn't a bug, but is the larger mapping used at this point in the
series?
At this commit the same node still has only two interrupts:
interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
<9 IRQ_TYPE_LEVEL_HIGH>;
interrupt-names = "xtr", "ana";
lan966x_probe() enables FDMA only when the named "fdma" interrupt exists:
lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
Without it, lan966x->fdma stays false and lan966x_fdma_pci_init() returns
before the ATU is programmed:
if (!lan966x->fdma)
return 0;
So nothing touches the extra DBI/iATU space until the next patch in the
series, "misc: lan966x-pci: dts: add fdma interrupt to overlay". That
patch adds <14 IRQ_TYPE_LEVEL_HIGH> as "fdma" between "xtr" and "ana",
which matches the binding order.
The in-between state is not a regression, because the driver keeps using
register-based extraction. This is only about how the series is split.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay
2026-09-28 19:33 ` [PATCH net-next v9 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
@ 2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:16 ` Daniel Machon
0 siblings, 1 reply; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 7:33 UTC (permalink / raw)
To: daniel.machon
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] Once this overlay adds the "fdma" interrupt, PCIe instances
use lan966x_fdma_pci_ops, and a teardown ordering bug in that backend…
--
misc: lan966x-pci: dts: add fdma interrupt to overlay
This commit adds the fdma interrupt (OIC interrupt 14) to the lan966x PCI
device tree overlay. This turns on FDMA-based frame injection and extraction
when the switch is connected over PCIe.
> diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> index 7bb726550cafd..5bb12dbc08431 100644
> --- a/drivers/misc/lan966x_pci.dtso
> +++ b/drivers/misc/lan966x_pci.dtso
> @@ -141,8 +141,9 @@ switch: switch@e0000000 {
>
> interrupt-parent = <&oic>;
> interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
> + <14 IRQ_TYPE_LEVEL_HIGH>,
> <9 IRQ_TYPE_LEVEL_HIGH>;
> - interrupt-names = "xtr", "fdma", "ana";
> + interrupt-names = "xtr", "fdma", "ana";
[Severity: Medium]
Once the "fdma" interrupt is present, PCIe instances use
lan966x_fdma_pci_ops. Does that make a teardown ordering problem in the PCIe
FDMA backend reachable?
lan966x_remove() tears down the FDMA before it unregisters the netdevs:
lan966x_main.c:lan966x_remove() {
...
lan966x->ops->fdma_deinit(lan966x);
lan966x_cleanup_ports(lan966x);
...
}
lan966x_fdma_pci_deinit() disables NAPI and frees both rings. It does not
clear lan966x->rx.lan966x, it does not detach the netdevs, and rtnl is not
held at this point:
lan966x_fdma_pci.c:lan966x_fdma_pci_deinit() {
if (lan966x->fdma_ndev)
napi_disable(&lan966x->napi);
...
fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->rx.fdma);
fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->tx.fdma);
}
fdma_free_coherent_and_unmap() only clears atu_region, so dcbs and dma
still hold their old values.
netif_set_mtu_ext() only checks netif_device_present(). That means an MTU
change that alters the max frame size can still run during this window:
lan966x_port_change_mtu()
lan966x_fdma_pci_resize()
__lan966x_fdma_pci_reload()
lan966x_fdma_pci_reload()
napi_disable(&lan966x->napi)
The resize guard is only this check:
if (!lan966x->rx.lan966x)
return 0;
rx.lan966x is still non-NULL after deinit. Would the second napi_disable()
on an already disabled NAPI spin forever while holding rtnl? That would also
block the unregister_netdev() that lan966x_remove() calls next.
The probe error path in lan966x_fdma_pci_init() looks like it has a similar
gap:
err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
if (err) {
fdma_free_coherent_and_unmap(lan966x->dma_dev, rx_fdma);
return err;
}
rx.lan966x has already been set earlier in lan966x_fdma_pci_init(). The
netdevs registered by lan966x_probe_port() stay live until the cleanup_ports
label.
If an MTU change arrives in that window, NAPI is still enabled from
lan966x_port_init(). lan966x_fdma_pci_reload() then allocates new rings and
does:
fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
Is this a second dma_free_coherent() of the RX ring that the error path
already freed?
After that, reload re-arms the FDMA on the new rings. The probe error path
never calls fdma_deinit for them. Would the coherent memory for the new rings
leak, with the outbound ATU windows still mapped and DMA active after probe
fails?
The resize guard added in "net: lan966x: add PCIe FDMA MTU change support"
is meant to skip the resize until lan966x_fdma_pci_init() has built the
rings. It doesn't seem to cover rings that were torn down or only partly
built.
The shutdown callback from "net: lan966x: add shutdown callback to stop the
FDMA on reboot" avoids this by taking rtnl and calling netif_device_detach()
so that ndo_open and ndo_change_mtu can't re-enter the FDMA. Should the
remove and probe error paths get the same protection?
The platform backend already has the same remove ordering. The PCIe deinit,
resize and reload code is new in this series, though, and this overlay change
is what first makes it reachable.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support
2026-10-02 7:33 ` netdev-bot+sashiko
@ 2026-10-02 9:02 ` Daniel Machon
0 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-10-02 9:02 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
> Thank you for your contribution! Sashiko AI review found 8 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 (1) · Medium: 3 (1) · Low: 5 (1)
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] Double fetch of the RX BLOCKL.
> - [Medium] lan966x_fdma_pci_napi_poll() limits RX consumption only by
> `counter < weight` and tracks refill progress only through the wrapped…
> - [Medium] When lan966x_fdma_pci_get_next_dcb() finds no free TX DCB,
> lan966x_fdma_pci_xmit() calls netif_stop_queue(dev) and returns…
> - [Low] At this commit the in-tree PCIe overlay
> (drivers/misc/lan966x_pci.dtso) has no "fdma" interrupt.
> - [Low] The iomap entry { TARGET_PCIE_DBI, 0x400000, 0 } derives
> regs[TARGET_PCIE_DBI] 4 MB into IO range 0.
> - [Low] lan966x_fdma_pci_resize() always returns -EOPNOTSUPP, so once
> PCIe FDMA is active every ndo_change_mtu call is rolled back.
> - [Low] In lan966x_fdma_pci_napi_poll(), the FDMA_ERROR branch says "No
> rx_dropped increment here because src_port is invalid".
> - [Low] With lan966x->ptp left at 0 on PCIe, lan966x_ptp_init() never
> initialises port->tx_skbs or the PTP locks, while the "ptp" IRQ…
>
> Pre-existing issues:
> - [High] This problem wasn't introduced by this patch, but the new PCIe
> backend inherits it.
> - [Medium] __fdma_dcb_add() writes fdma->last_dcb->nextptr (publishing
> the new DCB) before it writes the new DCB's own nextptr =…
> - [Low] lan966x_fdma_pci_hw_offload() reads ip_hdr(skb)->protocol and
> ipv6_hdr(skb)->daddr without checking that skb->len covers an IP…
>
> --
>
> net: lan966x: add PCIe FDMA support
>
> This commit adds a PCIe FDMA backend for lan966x. It uses contiguous
> coherent DMA buffers mapped through the endpoint's ATU and copies frames
> with memcpy. The new ops are selected when lan966x_is_pci() is true, and
> XDP and PTP are disabled on the PCIe path.
>
> > diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> > new file mode 100644
> > index 000000000000..f511e7061314
> > --- /dev/null
> > +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
>
> [ ... ]
>
> > +static bool lan966x_fdma_pci_hw_offload(struct lan966x *lan966x, u32 port,
> > + struct sk_buff **pskb)
> > +{
>
> [ ... ]
>
> > + if (skb->protocol == htons(ETH_P_IP) &&
> > + ip_hdr(skb)->protocol == IPPROTO_IGMP)
> > + return false;
> > +
> > + if (IS_ENABLED(CONFIG_IPV6) &&
> > + skb->protocol == htons(ETH_P_IPV6) &&
> > + ipv6_addr_is_multicast(&ipv6_hdr(skb)->daddr) &&
> > + !ipv6_mc_check_mld(skb))
> > + return false;
>
> [Severity: Low]
> This isn't a bug introduced by this patch, because the same code already
> exists in lan966x_hw_offload() in lan966x_main.c. Still,
> ip_hdr(skb)->protocol and ipv6_hdr(skb)->daddr are read here without
> checking that skb->len covers an IP header.
>
> On a runt frame, or on a frame shrunk by an XDP program once XDP support
> lands later in the series, could this read uninitialized skb tailroom?
>
> The reads stay inside the skb head allocation, and any minimum-size
> Ethernet frame covers both fields. So only the offload_fwd_mark decision
> for an already malformed packet is affected.
>
> > +static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
> > + u64 src_port)
> > +{
>
> [ ... ]
>
> > + /* Get the received frame and create an SKB for it. */
> > + db = fdma_db_next_get(fdma);
> > + data_len = fdma_db_len_get(db);
> > +
> > + skb = napi_alloc_skb(&lan966x->napi, data_len);
> > + if (unlikely(!skb))
> > + return NULL;
> > +
> > + memcpy(skb->data,
> > + fdma_dataptr_virt_addr_contiguous(fdma,
> > + fdma->dcb_index,
> > + fdma->db_index),
> > + data_len);
>
> [Severity: Medium]
> BLOCKL is validated in lan966x_fdma_pci_rx_check_frame() through
> lan966x_fdma_pci_rx_size_fits(). Here it is read again from the DCB
> status in coherent DMA memory.
>
> Can the value passed to napi_alloc_skb() and memcpy() differ from the
> value that was checked?
>
> It looks like "net: lan966x: add PCIe FDMA XDP support" later in the
> series fixes this. That patch reads blockl once in rx_check_frame() and
> passes data and data_len to rx_get_frame(). Would it make sense to fold
> that change into this patch?
>
> At this commit the path isn't enabled in-tree yet, because the overlay
> has no "fdma" interrupt.
Not enabled, yes. Non-issue.
>
> [ ... ]
>
> > +static int lan966x_fdma_pci_xmit(struct sk_buff *skb, __be32 *ifh,
> > + struct net_device *dev)
> > +{
>
> [ ... ]
>
> > + next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
> > +
> > + if (next_to_use < 0) {
> > + netif_stop_queue(dev);
> > + return NETDEV_TX_BUSY;
> > + }
>
> [Severity: Medium]
> netif_stop_queue() only stops TX queue 0. Each port netdev is created in
> lan966x_probe_port() with 8 TX queues:
>
> dev = devm_alloc_etherdev_mqs(lan966x->dev,
> sizeof(struct lan966x_port),
> NUM_PRIO_QUEUES, 1);
>
> There is no ndo_select_queue, so skbs are spread over all 8 queues.
> lan966x_fdma_wakeup_netdev(), called from the PCIe NAPI poll, also only
> checks and wakes queue 0.
>
> When the shared TX ring is full and an skb arrives on one of queues 1 to
> 7, does that queue ever get stopped?
>
> It looks like sch_direct_xmit() would requeue the skb and reschedule the
> qdisc. net_tx_action would then keep retrying, taking tx_lock and
> scanning the whole DCB ring each time, until the hardware completes a
> DCB.
>
> Would netif_tx_stop_all_queues() and netif_tx_wake_all_queues() be a
> better fit here? The platform lan966x_fdma_xmit() has the same pattern,
> and lan966x_fdma_pci_xmit_xdpf() from "net: lan966x: add PCIe FDMA XDP
> support" repeats it.
>
This is pre-existing behaviour. The lan966x platform path, sparx5 and lan969x
all do it the same way. I dont think this is a bug, so if anything, it should go
to net-next with at patchset for all platforms together
> [ ... ]
>
> > + /* Order frame write before DCB status write below. */
> > + dma_wmb();
> > +
> > + fdma_dcb_add(fdma,
> > + next_to_use,
> > + 0,
> > + FDMA_DCB_STATUS_INTR |
> > + FDMA_DCB_STATUS_SOF |
> > + FDMA_DCB_STATUS_EOF |
> > + FDMA_DCB_STATUS_BLOCKO(0) |
> > + FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + skb->len + ETH_FCS_LEN));
>
> [Severity: Medium]
> This is a pre-existing issue in the shared fdma_api.c helper and was not
> introduced by this patch, but the new backend depends on it.
> __fdma_dcb_add() links the new DCB into the chain before terminating it,
> and there is no DMA write barrier between the two steps:
>
> drivers/net/ethernet/microchip/fdma/fdma_api.c:__fdma_dcb_add() {
> ...
> fdma->last_dcb->nextptr = cpu_to_le64(nextptr);
>
> fdma->last_dcb = dcb;
>
> dcb->nextptr = cpu_to_le64(FDMA_DCB_INVALID_DATA);
> dcb->info = cpu_to_le64(info);
> ...
> }
>
> If the FDMA channel is still walking the chain, could it follow the new
> link and read a stale nextptr or info from the new DCB?
>
> The window is short, and whether the hardware acts on it depends on its
> prefetch behavior. When the channel is stopped, the writel() doorbell in
> lan966x_fdma_tx_start() or lan966x_fdma_rx_reload() orders the earlier
> stores.
>
> [ ... ]
>
> > +static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
> > +{
>
> [ ... ]
>
> > + dcb_reload = fdma->dcb_index;
> > +
> > + /* Get all received skbs. */
> > + while (counter < weight) {
> > + if (!fdma_has_frames(fdma))
> > + break;
> > + /* Order DONE read before DCB/frame reads below. */
> > + dma_rmb();
> > + counter++;
> > + switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
> > + case FDMA_PASS:
> > + break;
> > + case FDMA_ERROR:
> > + /* No rx_dropped increment here because src_port is
> > + * invalid.
> > + */
>
> [Severity: Low]
> Is this comment accurate for every FDMA_ERROR return?
> lan966x_fdma_pci_rx_check_frame() also returns FDMA_ERROR after src_port
> and ports[src_port] have already been validated:
>
> blockl = fdma_db_len_get(db);
> if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
> return FDMA_ERROR;
>
> In that case a frame with a bad BLOCKL on a valid port is dropped without
> being counted in rx_dropped or rx_length_errors. This code is unchanged
> at the end of the series.
>
> > + fdma_dcb_advance(fdma);
> > + continue;
> > + }
>
> [ ... ]
>
> > + while (dcb_reload != fdma->dcb_index) {
> > + old_dcb = dcb_reload;
> > + dcb_reload++;
> > + dcb_reload &= fdma->n_dcbs - 1;
>
> [Severity: Medium]
> The RX loop is bounded only by counter < weight. Refill progress is
> tracked only through the wrapped dcb_reload and dcb_index values.
>
> What happens if weight >= n_dcbs and every DCB is DONE? n_dcbs is
> FDMA_DCB_MAX, which is 512 here and 256 after "net: lan966x: add PCIe
> FDMA MTU change support".
>
> fdma_dcb_advance() would wrap dcb_index back to dcb_reload. DONE bits are
> only cleared by this refill loop, so the RX loop would then reprocess
> stale DONE descriptors and pass the same frames to napi_gro_receive()
> again.
>
> When weight is a multiple of n_dcbs, dcb_reload equals fdma->dcb_index at
> this point. The refill loop then doesn't run at all, and the ring is left
> un-armed.
>
> The normal NAPI weight of 64 can't reach this, but busy polling can.
> __napi_busy_loop() passes the socket's budget straight to napi->poll, and
> SO_BUSY_POLL_BUDGET accepts values up to U16_MAX with CAP_NET_ADMIN.
>
> The native lan966x_fdma_napi_poll() has the same structure. This is
> still present at the end of the series.
This is a re-raise of v7 finding that was answered and dismissed.
>
> [ ... ]
>
> > +static int lan966x_fdma_pci_init(struct lan966x *lan966x)
> > +{
> > + struct fdma *rx_fdma = &lan966x->rx.fdma;
> > + struct fdma *tx_fdma = &lan966x->tx.fdma;
> > + int err;
> > +
> > + if (!lan966x->fdma)
> > + return 0;
>
> [Severity: Low]
> At this commit the in-tree PCIe overlay (drivers/misc/lan966x_pci.dtso)
> only declares:
>
> interrupt-names = "xtr", "ana";
>
> As a result lan966x->fdma stays false on the PCIe instance, and this
> function returns right away. Is it intended that the new backend is
> dormant at this commit?
>
> The "fdma" interrupt is added later in the series by "misc: lan966x-pci:
> dts: add fdma interrupt to overlay". Until then the driver keeps using
> register-based I/O.
>
> [ ... ]
>
> > + lan966x->tx.lan966x = lan966x;
> > + tx_fdma->channel_id = FDMA_INJ_CHANNEL;
> > + tx_fdma->n_dcbs = FDMA_DCB_MAX;
>
> [Severity: High]
> This isn't a bug introduced by this patch, because the platform
> lan966x_fdma_init() uses the same ordering, but the new PCIe backend
> inherits it.
>
> lan966x_probe() sets lan966x->fdma = true, registers every port netdev,
> and enables NAPI (lan966x_port_init() -> lan966x_fdma_netdev_init()), all
> before it calls lan966x->ops->fdma_init().
>
> Here tx_fdma->n_dcbs is set without holding tx_lock. tx_fdma->dcbs is
> only allocated later, in lan966x_fdma_pci_tx_alloc(). That happens after
> the GFP_KERNEL coherent allocation in lan966x_fdma_pci_rx_alloc(), which
> can sleep.
>
> If a port is opened and transmits during that window, can the following
> happen?
>
> lan966x_port_xmit()
> spin_lock(&lan966x->tx_lock);
> if (port->lan966x->fdma)
> lan966x->ops->fdma_xmit()
> lan966x_fdma_pci_xmit()
> lan966x_fdma_pci_get_next_dcb()
> fdma_db_get(fdma, i, 0) <- fdma->dcbs is NULL
>
> Later in the same window, xmit could also pick DCBs that
> fdma_dcbs_init() has already marked DONE while last_dcb is still being
> set up. It could also start the injection channel before
> lan966x_fdma_llp_configure() has programmed the TX LLP.
>
> Would either of these close the window?
>
> - Calling fdma_init before the ports are registered.
> - Publishing n_dcbs under tx_lock, only after the rings are allocated.
>
This is a re-raise of v7 finding that was answered and dismissed. Probe order is
pre-existing and not something this series can fix.
> [ ... ]
>
> > +static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
> > +{
> > + return -EOPNOTSUPP;
> > +}
>
> [Severity: Low]
> lan966x_port_change_mtu() rolls back the MTU whenever
> lan966x->ops->fdma_resize() fails. With this stub, won't every
> ndo_change_mtu call fail once PCIe FDMA is active?
>
> The next patch, "net: lan966x: add PCIe FDMA MTU change support",
> replaces the stub with a real implementation. At this commit the path is
> dormant, because the overlay has no "fdma" interrupt.
>
> [ ... ]
>
> > diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> > index 024ce9f9916c..de2202786826 100644
> > --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> > +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> > @@ -50,6 +50,7 @@ struct lan966x_main_io_resource {
> > static const struct lan966x_main_io_resource lan966x_main_iomap[] = {
> > { TARGET_CPU, 0xc0000, 0 }, /* 0xe00c0000 */
> > { TARGET_FDMA, 0xc0400, 0 }, /* 0xe00c0400 */
> > + { TARGET_PCIE_DBI, 0x400000, 0 }, /* 0xe0400000 */
>
> [Severity: Low]
> This entry places regs[TARGET_PCIE_DBI] 4 MB into IO range 0. At this
> commit the overlay's "cpu" reg is only 1 MB:
>
> reg = <0xe0000000 0x0100000>,
>
> Wouldn't the iATU accesses from fdma_pci_atu_init(), and the region map
> and unmap, fall outside the ioremap window?
>
> This can't be reached at this commit, because lan966x_fdma_pci_init()
> returns early without an "fdma" interrupt. Later in the series, "misc:
> lan966x-pci: dts: extend cpu reg to cover PCIE DBI space" enlarges the
> reg to 0x0800000, and it comes before the "fdma" interrupt is added. On
> the SoC instance the pointer is computed but never used.
>
> [ ... ]
>
> > @@ -1187,7 +1190,9 @@ static int lan966x_probe(struct platform_device *pdev)
> > if (err)
> > return dev_err_probe(&pdev->dev, err, "Unable to use ptp irq");
> >
> > - lan966x->ptp = 1;
> > + /* PTP is not supported on the PCIe path yet. */
> > + if (!lan966x_is_pci(lan966x))
> > + lan966x->ptp = 1;
> > }
>
> [Severity: Low]
> With lan966x->ptp left at 0 on PCIe, lan966x_ptp_init() never
> initializes port->tx_skbs or the PTP locks. If a "ptp" interrupt exists,
> lan966x_ptp_irq_handler has already been registered just above. Could
> that handler then touch uninitialized PTP state?
>
> This can't happen in-tree, because drivers/misc/lan966x_pci.dtso doesn't
> declare a "ptp" interrupt. Would it be more robust to also skip the ptp
> IRQ request on PCIe?
>
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 12/15] net: lan966x: add PCIe FDMA MTU change support
2026-10-02 7:33 ` netdev-bot+sashiko
@ 2026-10-02 9:08 ` Daniel Machon
0 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-10-02 9:08 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
> Thank you for your contribution! Sashiko AI review found 2 potential
> issue(s) to consider.
>
> Critical: 0 · High: 1 (1) · Medium: 0 · Low: 1
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [High] lan966x_fdma_pci_resize() uses `lan966x->rx.lan966x` to decide
> whether the rings exist, and that field does not track ring state.
> - [Low] lan966x_fdma_pci_reload() frees and ATU-unmaps the old rings and
> reprograms LLP right after lan966x_fdma_rx_disable() and…
>
> Pre-existing issues:
> - [High] This problem wasn't introduced by this patch, but the patch
> explicitly acknowledges the hazard and fixes only the PCIe side.
>
> --
>
> net: lan966x: add PCIe FDMA MTU change support
>
> This adds MTU change support to the lan966x PCIe FDMA path. On an MTU
> change, the contiguous ATU-mapped RX and TX rings are reallocated at the
> new size, and the old buffers are kept if allocation fails. The PCIe DCB
> ring is capped at 256, FDMA_PCI_MAX_MTU is advertised as max_mtu when the
> FDMA is in use, and the resize is skipped until lan966x_fdma_pci_init()
> has built the rings.
>
> > diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> > index f511e70613141..758554c951c52 100644
> > --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> > +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
>
> [ ... ]
>
> > +static int lan966x_fdma_pci_reload(struct lan966x *lan966x, int new_mtu)
> > +{
> > + struct fdma tx_fdma_old = lan966x->tx.fdma;
> > + struct fdma rx_fdma_old = lan966x->rx.fdma;
> > + u32 old_mtu = lan966x->rx.max_mtu;
> > + int err;
> > +
> > + napi_disable(&lan966x->napi);
> > + lan966x_fdma_tx_disable_netdev(lan966x);
> > + lan966x_fdma_rx_disable(&lan966x->rx);
> > + lan966x_fdma_tx_disable(&lan966x->tx);
>
> [Severity: Low]
> lan966x_fdma_rx_disable() and lan966x_fdma_tx_disable() both throw away
> the result of the FDMA_CH_ACTIVE poll:
>
> readx_poll_timeout_atomic(lan966x_fdma_channel_active, lan966x,
> val, !(val & BIT(fdma->channel_id)),
> READL_SLEEP_US, READL_TIMEOUT_US);
>
> __lan966x_fdma_pci_reload() also ignores the result of the CPU queue drain
> poll. After that, the old rings are freed and ATU-unmapped straight away.
>
> If a channel is still active when the poll times out, could it follow the
> old DCB chain into memory that has already been unmapped and freed?
>
> The timeout is READL_TIMEOUT_US (100 seconds), so this would need wedged
> hardware. The platform lan966x_fdma_reload() uses the same sequence. It
> might still be worth checking the poll results and bailing out before
> freeing the old rings.
>
> > +
> > + lan966x->rx.max_mtu = new_mtu;
>
> [ ... ]
>
> > + /* Free and unmap old memory. */
> > + fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
> > + fdma_free_coherent_and_unmap(lan966x->dma_dev, &tx_fdma_old);
>
> [ ... ]
>
> > static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
> > {
> > - return -EOPNOTSUPP;
> > + int max_mtu;
> > +
> > + /* Nothing to resize until fdma_pci_init() has built the rings; it
> > + * sizes them from DEV_MAC_MAXLEN_CFG, which the caller already set.
> > + */
> > + if (!lan966x->rx.lan966x)
> > + return 0;
>
> [Severity: High]
> Does rx.lan966x actually show whether the rings have been built?
>
> In lan966x_fdma_pci_init() it is set before either ring is allocated or
> ATU-mapped:
>
> lan966x->rx.lan966x = lan966x;
> lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
> ...
> err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
> if (err)
> return err;
>
> It is not cleared on the init error paths or in
> lan966x_fdma_pci_deinit(). Init and deinit both run without RTNL while the
> port netdevs are registered, so ndo_change_mtu can reach this check during
> those windows.
>
> First, lan966x_remove() tears down the FDMA before the netdevs are
> unregistered:
>
> lan966x->ops->fdma_deinit(lan966x);
> lan966x_cleanup_ports(lan966x);
>
> lan966x_fdma_pci_deinit() calls napi_disable() and frees both rings. An
> MTU change in that window passes this check and reaches
> lan966x_fdma_pci_reload(), which calls napi_disable() a second time.
>
> Would napi_disable_locked() then spin forever on NAPIF_STATE_SCHED |
> NAPIF_STATE_NPSVC with RTNL held? That would leave unregister_netdev() and
> every other RTNL user blocked. The probe unwind at cleanup_fdma has the
> same ordering.
>
> Second, if lan966x_fdma_pci_tx_alloc() fails in init, the RX ring is freed.
> However, rx.fdma.dcbs still points at it and rx.lan966x stays set. The
> probe unwind jumps to cleanup_ptp and skips fdma_deinit, so the netdevs
> stay registered until lan966x_cleanup_ports(). An MTU change in that
> window runs lan966x_fdma_pci_reload(), which calls:
>
> fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
>
> Can this free the stale RX buffer a second time? The newly allocated,
> ATU-mapped rings would then leak while the FDMA runs on them.
>
> The same thing looks possible when the ATU map fails inside
> lan966x_fdma_pci_rx_alloc(), because fdma_alloc_coherent_and_map() frees
> dcbs without clearing the pointer.
>
> Third, lan966x_probe() registers every port netdev in lan966x_probe_port()
> before ops->fdma_init() runs. A concurrent MTU change, for example udev
> applying MTUBytes, could pass this check while init is still sleeping in
> dma_alloc_coherent() inside rx_alloc or tx_alloc.
>
> Reload would then snapshot a half-built struct fdma, clear atu_region, and
> rewrite db_size and size while init is still filling them in. Could that
> leak or double-map ATU regions, or leave LLP pointing at freed memory?
>
> The commit message says:
>
> Skip the resize until lan966x_fdma_pci_init() has built the rings;
>
> The check does not seem to do that. The resize, init and deinit paths are
> also unchanged at the end of the series.
I dismissed this issue in v7, and Simon gave his RB. In short, all these windows
come from the probe init order and the netdev registration.
>
> > +
> > + max_mtu = lan966x_fdma_get_max_frame(lan966x);
> > + if (max_mtu == lan966x->rx.max_mtu)
> > + return 0;
> > +
> > + return __lan966x_fdma_pci_reload(lan966x, max_mtu);
> > }
>
> [Severity: High]
> This isn't a bug introduced by this patch, but the platform backend has
> the same pre-init window, and the guard added here only covers PCIe.
>
> lan966x_fdma_change_mtu() in lan966x_fdma.c has no equivalent check:
>
> max_mtu = lan966x_fdma_get_max_frame(lan966x);
> if (max_mtu == lan966x->rx.max_mtu)
> return 0;
>
> return __lan966x_fdma_reload(lan966x, max_mtu);
>
> rx.max_mtu is still 0 before lan966x_fdma_init() runs. So an MTU change
> after register_netdev() in lan966x_probe_port() always continues down:
>
> lan966x_port_change_mtu()->lan966x_fdma_change_mtu()->
> __lan966x_fdma_reload()->lan966x_fdma_reload()
>
> >From there, lan966x_fdma_rx_disable() does:
>
> struct lan966x *lan966x = rx->lan966x;
> ...
> lan_rmw(FDMA_CH_DISABLE_CH_DISABLE_SET(BIT(fdma->channel_id)),
> FDMA_CH_DISABLE_CH_DISABLE,
> lan966x, FDMA_CH_DISABLE);
>
> rx->lan966x is still NULL at that point. Can this oops under RTNL?
>
> Would a check in the common caller, or an equivalent guard in
> lan966x_fdma_change_mtu(), cover both backends?
>
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 13/15] net: lan966x: add PCIe FDMA XDP support
2026-10-02 7:33 ` netdev-bot+sashiko
@ 2026-10-02 9:11 ` Daniel Machon
0 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-10-02 9:11 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
> Thank you for your contribution! Sashiko AI review found 2 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 1
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] This patch removes the PCIe rejection from lan966x_xdp_setup()
> (drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c), so a bond can…
> - [Low] In lan966x_xdp_pci_run(), xdp_init_buff() is called with frame_sz
> = fdma->db_size, and hard_start is the slot start (virt_addr -…
>
> --
>
> net: lan966x: add PCIe FDMA XDP support
>
> This commit adds XDP support to the lan966x PCIe FDMA path. The IFH and
> FCS are stripped before the program runs, and XDP_TX is done by copying
> the frame into a TX slot behind a newly built IFH. lan966x_xdp_setup()
> no longer rejects PCIe and skips the page_pool reload on that path.
>
> > diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> > index 758554c951c52..949994874ed96 100644
> > --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> > +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
>
> [ ... ]
>
> > @@ -139,7 +140,123 @@ static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
> > blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
> > }
> >
>
> [ ... ]
>
> > +static int lan966x_xdp_pci_run(struct lan966x_port *port, void *data,
> > + u32 data_len, void **xdp_data, u32 *xdp_len)
> > +{
>
> [ ... ]
>
> > + xdp_init_buff(&xdp, fdma->db_size, &port->xdp_rxq);
> > +
> > + /* hard_start is set to slot start (virt_addr is XDP_PACKET_HEADROOM
> > + * into the slot). Headroom includes the IFH; BPF may grow into it
> > + * via adjust_head. IFH is rebuilt on XDP_TX and unread on XDP_PASS.
> > + */
> > + xdp_prepare_buff(&xdp,
> > + data - XDP_PACKET_HEADROOM,
> > + XDP_PACKET_HEADROOM + IFH_LEN_BYTES,
> > + data_len,
> > + false);
>
> [Severity: Low]
> Can data_end end up past xdp_data_hard_end() here?
>
> frame_sz is the full db_size and hard_start is the start of the slot, so
> the XDP core computes the data limit as:
>
> include/net/xdp.h:xdp_data_hard_end()
> ((xdp)->data_hard_start + (xdp)->frame_sz -
> SKB_DATA_ALIGN(sizeof(struct skb_shared_info)))
>
> The upper bound in lan966x_fdma_pci_rx_size_fits(), shown in the context
> above, only reserves the headroom:
>
> blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
>
> lan966x_fdma_pci_rx_check_frame() sets data to slot + XDP_PACKET_HEADROOM
> + IFH_LEN_BYTES and data_len to blockl - IFH_LEN_BYTES - ETH_FCS_LEN. At
> the largest accepted blockl, data_end is then slot + db_size -
> ETH_FCS_LEN, roughly 316 bytes past xdp_data_hard_end().
>
> For such a buffer, bpf_xdp_adjust_tail() returns -EINVAL for any offset,
> shrinks included, because it checks this first:
>
> if (unlikely(data_end > data_hard_end))
> return -EINVAL;
>
> Reaching this needs a blockl that eats into the shinfo reserve, for
> example a malformed BLOCKL. It could also be an oversized frame truncated
> to DATAL while lan966x_port_change_mtu() has raised MAXLEN but
> fdma_resize() has not yet reallocated the rings.
>
> Memory safety is not affected, since data_end stays inside the slot the
> driver owns.
>
> Should lan966x_fdma_pci_rx_size_fits() also reserve
> SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) so it matches the frame_sz
> given to xdp_init_buff()?
>
> > +
> > + act = bpf_prog_run_xdp(xdp_prog, &xdp);
> > +
> > + *xdp_data = xdp.data;
> > + *xdp_len = xdp.data_end - xdp.data;
> > +
> > + switch (act) {
> > + case XDP_PASS:
> > + return FDMA_PASS;
> > + case XDP_TX:
> > + return lan966x_fdma_pci_xmit_xdpf(port, *xdp_data, *xdp_len) ?
> > + FDMA_DROP : FDMA_TX;
> > + default:
> > + bpf_warn_invalid_xdp_action(port->dev, xdp_prog, act);
> > + fallthrough;
> > + case XDP_ABORTED:
> > + trace_xdp_exception(port->dev, xdp_prog, act);
> > + fallthrough;
> > + case XDP_DROP:
> > + return FDMA_DROP;
> > + }
> > +}
>
> [Severity: Medium]
> Now that lan966x_xdp_setup() no longer rejects PCIe, can a bond attach a
> program to lan966x PCIe ports and then have its XDP_TX frames silently
> dropped here?
>
> The ports share lan966x_port_netdev_ops, which provides both .ndo_bpf and
> .ndo_xdp_xmit. That is all bond_xdp_set() checks, so it accepts these
> ports. It then pushes the program down:
>
> bond_xdp_set()
> dev_xdp_propagate()
> lan966x_xdp_setup() <- now succeeds on PCIe
>
> bond_xdp_set() also calls
> static_branch_inc(&bpf_master_redirect_enabled_key). After that,
> bpf_prog_run_xdp() rewrites XDP_TX on a bond slave:
>
> if (act == XDP_TX && netif_is_bond_slave(xdp->rxq->dev))
> act = xdp_master_redirect(xdp);
>
> In round-robin, XOR and 802.3ad modes, xdp_master_redirect() returns
> XDP_REDIRECT whenever the bond picks a transmit slave other than the
> receiving port. That action falls into the default case above:
>
> bpf_warn_invalid_xdp_action() -> trace_xdp_exception() -> FDMA_DROP
>
> A bond program that only returns XDP_TX attaches without error, but a
> hash- or round-robin-dependent share of its packets is dropped. With two
> slaves in round-robin, that is about half. Before this patch the attach
> failed with -EOPNOTSUPP.
>
> Should the PCIe path handle XDP_REDIRECT, or keep refusing the attach
> when the port is a bond slave?
Re-raise of issue dismissed in v7. xdp_features doesn't advertise REDIRECT on
PCIe, and the warning path is the intended fallback
>
> [ ... ]
>
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay
2026-10-02 7:33 ` netdev-bot+sashiko
@ 2026-10-02 9:16 ` Daniel Machon
0 siblings, 0 replies; 29+ messages in thread
From: Daniel Machon @ 2026-10-02 9:16 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
horms, richardcochran, netdev, linux-kernel, bpf,
linux-arm-kernel
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] Once this overlay adds the "fdma" interrupt, PCIe instances
> use lan966x_fdma_pci_ops, and a teardown ordering bug in that backend…
>
> --
>
> misc: lan966x-pci: dts: add fdma interrupt to overlay
>
> This commit adds the fdma interrupt (OIC interrupt 14) to the lan966x PCI
> device tree overlay. This turns on FDMA-based frame injection and extraction
> when the switch is connected over PCIe.
>
> > diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> > index 7bb726550cafd..5bb12dbc08431 100644
> > --- a/drivers/misc/lan966x_pci.dtso
> > +++ b/drivers/misc/lan966x_pci.dtso
> > @@ -141,8 +141,9 @@ switch: switch@e0000000 {
> >
> > interrupt-parent = <&oic>;
> > interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
> > + <14 IRQ_TYPE_LEVEL_HIGH>,
> > <9 IRQ_TYPE_LEVEL_HIGH>;
> > - interrupt-names = "xtr", "fdma", "ana";
> > + interrupt-names = "xtr", "fdma", "ana";
>
> [Severity: Medium]
> Once the "fdma" interrupt is present, PCIe instances use
> lan966x_fdma_pci_ops. Does that make a teardown ordering problem in the PCIe
> FDMA backend reachable?
>
> lan966x_remove() tears down the FDMA before it unregisters the netdevs:
>
> lan966x_main.c:lan966x_remove() {
> ...
> lan966x->ops->fdma_deinit(lan966x);
> lan966x_cleanup_ports(lan966x);
> ...
> }
>
> lan966x_fdma_pci_deinit() disables NAPI and frees both rings. It does not
> clear lan966x->rx.lan966x, it does not detach the netdevs, and rtnl is not
> held at this point:
>
> lan966x_fdma_pci.c:lan966x_fdma_pci_deinit() {
> if (lan966x->fdma_ndev)
> napi_disable(&lan966x->napi);
> ...
> fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->rx.fdma);
> fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->tx.fdma);
> }
>
> fdma_free_coherent_and_unmap() only clears atu_region, so dcbs and dma
> still hold their old values.
>
> netif_set_mtu_ext() only checks netif_device_present(). That means an MTU
> change that alters the max frame size can still run during this window:
>
> lan966x_port_change_mtu()
> lan966x_fdma_pci_resize()
> __lan966x_fdma_pci_reload()
> lan966x_fdma_pci_reload()
> napi_disable(&lan966x->napi)
>
> The resize guard is only this check:
>
> if (!lan966x->rx.lan966x)
> return 0;
>
> rx.lan966x is still non-NULL after deinit. Would the second napi_disable()
> on an already disabled NAPI spin forever while holding rtnl? That would also
> block the unregister_netdev() that lan966x_remove() calls next.
>
> The probe error path in lan966x_fdma_pci_init() looks like it has a similar
> gap:
>
> err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
> if (err) {
> fdma_free_coherent_and_unmap(lan966x->dma_dev, rx_fdma);
> return err;
> }
>
> rx.lan966x has already been set earlier in lan966x_fdma_pci_init(). The
> netdevs registered by lan966x_probe_port() stay live until the cleanup_ports
> label.
>
> If an MTU change arrives in that window, NAPI is still enabled from
> lan966x_port_init(). lan966x_fdma_pci_reload() then allocates new rings and
> does:
>
> fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
>
> Is this a second dma_free_coherent() of the RX ring that the error path
> already freed?
>
> After that, reload re-arms the FDMA on the new rings. The probe error path
> never calls fdma_deinit for them. Would the coherent memory for the new rings
> leak, with the outbound ATU windows still mapped and DMA active after probe
> fails?
>
> The resize guard added in "net: lan966x: add PCIe FDMA MTU change support"
> is meant to skip the resize until lan966x_fdma_pci_init() has built the
> rings. It doesn't seem to cover rings that were torn down or only partly
> built.
>
> The shutdown callback from "net: lan966x: add shutdown callback to stop the
> FDMA on reboot" avoids this by taking rtnl and calling netif_device_detach()
> so that ndo_open and ndo_change_mtu can't re-enter the FDMA. Should the
> remove and probe error paths get the same protection?
>
> The platform backend already has the same remove ordering. The PCIe deinit,
> resize and reload code is new in this series, though, and this overlay change
> is what first makes it reachable.
>
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928-lan966x-pci-fdma-v9-0-a17caf60f89a%40microchip.com
Pre-existing. Lan966x_remove() is a platform function. It's a re-raise of v7
11/14 / v8 15/15, the probe/remove ordering. The issue is real, but out-of-scope
for this series.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 04/15] net: microchip: fdma: use little-endian types for descriptor fields
2026-09-28 19:32 ` [PATCH net-next v9 04/15] net: microchip: fdma: use little-endian types for descriptor fields Daniel Machon
@ 2026-10-02 13:25 ` Simon Horman
0 siblings, 0 replies; 29+ messages in thread
From: Simon Horman @ 2026-10-02 13:25 UTC (permalink / raw)
To: Daniel Machon
Cc: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Richard Cochran, netdev,
linux-kernel, bpf, linux-arm-kernel
On Mon, Sep 28, 2026 at 09:32:52PM +0200, Daniel Machon wrote:
> The FDMA engine reads and writes the DCB and DB descriptors in
> little-endian byte order. So far the descriptors have only been produced
> by the little-endian SoC itself, so plain u64 fields were fine. With the
> PCIe FDMA path, the descriptors live in host memory and are written by
> the host CPU, which may be big-endian.
>
> Change the descriptor fields to __le64 and convert at the library
> boundary: in __fdma_db_add() and __fdma_dcb_add() on write, and in the
> fdma_db_*() accessors on read. The dataptr and nextptr callbacks keep
> their u64 signatures, so their implementations are unchanged. Add
> fdma_db_dataptr_get(), and convert the lan966x sites that read the
> descriptor fields directly to use the accessors.
>
> No functional change on little-endian hosts.
>
> Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support
2026-09-28 19:32 ` [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
@ 2026-10-02 14:14 ` Simon Horman
1 sibling, 0 replies; 29+ messages in thread
From: Simon Horman @ 2026-10-02 14:14 UTC (permalink / raw)
To: Daniel Machon
Cc: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
Greg Kroah-Hartman, Mohsin Bashir, Richard Cochran, netdev,
linux-kernel, bpf, linux-arm-kernel
On Mon, Sep 28, 2026 at 09:32:59PM +0200, Daniel Machon wrote:
> Add PCIe FDMA support for lan966x. The PCIe FDMA path uses contiguous
> DMA buffers mapped through the endpoint's ATU, with memcpy-based frame
> transfer instead of per-page DMA mappings.
>
> With PCIe FDMA, throughput increases from ~33 Mbps (register-based I/O)
> to ~620 Mbps on an Intel x86 host with a lan966x PCIe card.
>
> The RX path uses its own copy of lan966x_hw_offload() that reports back
> when skb_vlan_untag() frees the skb, so the caller can drop the frame
> instead of touching it.
>
> XDP is not supported on this path yet, so do not advertise xdp_features
> and reject program attach for PCIe instances.
>
> Tested-by: Herve Codina <herve.codina@bootlin.com>
> Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
^ permalink raw reply [flat|nested] 29+ messages in thread
end of thread, other threads:[~2026-10-02 14:15 UTC | newest]
Thread overview: 29+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 19:32 [PATCH net-next v9 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 02/15] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:32 ` [PATCH net-next v9 04/15] net: microchip: fdma: use little-endian types for descriptor fields Daniel Machon
2026-10-02 13:25 ` Simon Horman
2026-09-28 19:32 ` [PATCH net-next v9 05/15] net: lan966x: add FDMA LLP register write helper Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 06/15] net: lan966x: export FDMA helpers for reuse Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 07/15] net: lan966x: use a dedicated device for DMA operations Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 08/15] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
2026-09-28 19:32 ` [PATCH net-next v9 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:32 ` [PATCH net-next v9 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:02 ` Daniel Machon
2026-10-02 14:14 ` Simon Horman
2026-09-28 19:33 ` [PATCH net-next v9 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:08 ` Daniel Machon
2026-09-28 19:33 ` [PATCH net-next v9 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:11 ` Daniel Machon
2026-09-28 19:33 ` [PATCH net-next v9 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-09-28 19:33 ` [PATCH net-next v9 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
2026-10-02 7:33 ` netdev-bot+sashiko
2026-10-02 9:16 ` Daniel Machon
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®