mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA
@ 2026-09-24 19:56 Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
                   ` (14 more replies)
  0 siblings, 15 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

When lan966x operates as a PCIe endpoint, the driver currently uses
register-based I/O for frame injection and extraction. This approach is
functional but slow, topping out at around 33 Mbps on an Intel x86 host
with a lan966x PCIe card.

This series adds FDMA (Frame DMA) support for the PCIe path. When
operating as a PCIe endpoint, the internal FDMA engine on lan966x cannot
directly access host memory, so DMA buffers are allocated as contiguous
coherent memory and mapped through the PCIe Address Translation Unit
(ATU). The ATU provides outbound windows that translate internal FDMA
addresses to PCIe bus addresses, allowing the FDMA engine to read and
write host memory. Because the ATU requires contiguous address regions,
page_pool and normal per-page DMA mappings cannot be used. Instead,
frames are transferred using memcpy between the ATU-mapped buffers and
the network stack. With this, throughput increases from ~33 Mbps to
~620 Mbps for default MTU.

Patch 1 adds the shared drivers/net/ethernet/microchip/fdma/ directory
to the Sparx5 SoC MAINTAINERS entry.

Patches 2-4 prepare the shared FDMA library: patch 2 renames the
contiguous dataptr helpers for clarity, patch 3 adds PCIe ATU region
management and coherent DMA allocation with ATU mapping, and patch 4
gives the descriptor fields an explicit little-endian type, since on the
PCIe path the descriptors are written by the host CPU.

Patches 5-8 refactor the lan966x FDMA code to support both platform
and PCIe paths: extracting the LLP register write into a helper,
exporting shared functions, introducing a dedicated device for DMA
operations, and adding an ops dispatch table selected at probe time.

Patches 9-10 harden the existing FDMA path for the PCIe endpoint
lifecycle: patch 9 clears latched FDMA error/interrupt stickies after
the switch reset so they don't assert as soon as interrupts are
enabled, and patch 10 adds a shutdown() callback that quiesces the
FDMA engine on host warm reboot (on the PCIe card the FDMA survives
host reset and would otherwise keep the shared INTx asserted into
the next probe).

Patch 11 adds the core PCIe FDMA implementation with RX/TX using
contiguous ATU-mapped buffers. Patches 12 and 13 extend it with MTU
change and XDP support respectively. XDP_PASS, XDP_TX, XDP_DROP and
XDP_ABORTED are supported; XDP_REDIRECT is deliberately not, because
the PCIe data path does not use page_pool.

Patches 14-15 update the lan966x PCI device tree overlay to extend the
cpu register mapping to cover the ATU register space and add the FDMA
interrupt.

Patches 1-13 touch MAINTAINERS and drivers/net/ethernet/microchip/, and
are for the netdev tree.

Patches 14-15 touch drivers/misc/lan966x_pci.dtso, and are for the
char-misc tree.

To: Andrew Lunn <andrew+netdev@lunn.ch>
To: David S. Miller <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Horatiu Vultur <horatiu.vultur@microchip.com>
To: Steen Hegelund <steen.hegelund@microchip.com>
To: UNGLinuxDriver@microchip.com
To: Alexei Starovoitov <ast@kernel.org>
To: Daniel Borkmann <daniel@iogearbox.net>
To: Jesper Dangaard Brouer <hawk@kernel.org>
To: John Fastabend <john.fastabend@gmail.com>
To: Stanislav Fomichev <sdf@fomichev.me>
To: Herve Codina <herve.codina@bootlin.com>
To: Arnd Bergmann <arnd@arndb.de>
To: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
To: Mohsin Bashir <mohsin.bashr@gmail.com>
To: Simon Horman <horms@kernel.org>
Cc: Richard Cochran <richardcochran@gmail.com>
Cc: netdev@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: bpf@vger.kernel.org
Cc: linux-arm-kernel@lists.infradead.org

Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
Changes in v8:
- New patch 4: use little-endian types for the FDMA descriptor fields.
  On the PCIe path the descriptors live in host memory and are written
  by the host CPU, which may be big-endian, while the FDMA engine reads
  them little-endian. The fields become __le64 and are converted at the
  library boundary; the dataptr/nextptr callbacks keep their u64
  signatures, so sparx5 and lan969x need no changes. Checked with sparse
  (-D__CHECK_ENDIAN__). No functional change on little-endian hosts.
  Tested on lan966x (arm), sparx5 and lan969x (arm64), and lan966x over
  PCIe on an x86 host, with no regressions. (Simon)
- PCIe: read the RX block length through fdma_db_len_get() instead of
  open-coding the status field, following patch 4.
- Stickies: extend the comment to also mention the data-block sticky,
  which is cleared alongside the error stickies.
- Shutdown: reword the commit message to describe the tree at that
  commit, noting where it relies on the PCIe FDMA backend added later
  in the series.
- Link to v7: https://lore.kernel.org/r/20260918-lan966x-pci-fdma-v7-0-0ecc179c8a2c@microchip.com

Changes in v7:

Mostly a response to the automated review of v6. The notable change is
that the resize-time MTU guards added in v6 are replaced by an advertised
limit. v6 rejected an oversized MTU inside lan966x_fdma_pci_resize() with
-ERANGE, which is late, and leaves dev->max_mtu advertising a range the
PCIe path cannot honour. v7 derives the maximum from the same constants
and publishes it in dev->max_mtu instead, so the core rejects the request
up front and both -ERANGE returns become unreachable and are removed.
Also, the shutdown() was refactored to not race with ndo's.

- PCIe: cap dev->max_mtu on the PCIe path (FDMA_PCI_MAX_MTU), derived
  from the contiguous ring having to fit one MAX_PAGE_ORDER block after
  padding to the ATU region granularity, and from db_size having to fit
  the 16-bit DCB DATAL field. On a 4KB-page, MAX_PAGE_ORDER=10 build that
  is MTU 15498. Both resize-time -ERANGE checks are removed with it.
- PCIe: factor the frame overhead into LAN966X_FDMA_OVERHEAD and use it
  for both the cap and lan966x_fdma_get_max_frame(), so it is not
  open-coded twice.
- PCIe: skip the resize while the FDMA is not initialised. The rings are
  built by lan966x_fdma_pci_init(), which runs after the netdevs are
  registered, so an MTU change in that window would dereference NULL.
- PCIe: comment why NAPI has to be enabled before the queues are woken in
  the MTU reload path.
- shutdown(): rework the quiesce ordering. Free the xtr, ana and FDMA
  irqs first so no source can assert INTx, mask ANA_ANAINTR whether or not
  the FDMA is in use, then disable NAPI, stop and detach the netdevs,
  disable the FDMA channels, clear the FDMA interrupt enables and unmap
  the ATU windows.
- ATU: document the alignment requirement and -EINVAL at
  fdma_pci_atu_region_map(), document at fdma_pci_atu_translate_addr()
  that the caller must quiesce DMA and the descriptor paths before
  unmapping a region, initialise the region bookkeeping explicitly in
  fdma_pci_atu_init(), and fix the comment on the upper limit register.
- Patch 3 commit message: drop the claim that the patch adds PCIe FDMA
  support; it adds the helpers the lan966x path later uses.

Reported by the v6 review and deliberately not fixed here, because the
code is pre-existing and the fixes want a Fixes: tag:

- register_netdev() and the FDMA irq request both happen before
  ops->fdma_init(), so an MTU change or a stale interrupt can reach an
  uninitialised FDMA. Pre-existing on the platform path; the fix is a
  reorder of the shared probe, which I will send against net. The skip
  above closes the MTU-change half on PCIe in the meantime.
- lan966x_hw_offload() frees the caller's skb and reports only a bool.
  Pre-existing, all three callers affected. Goes to net with
  Fixes: 6476f90aefaf.
- The platform XDP hook leaves the FCS in the window handed to the BPF
  program, while the PCIe hook strips it. Pre-existing. Goes to net with
  Fixes: 6a2159be7604.
- Link to v6: https://lore.kernel.org/r/20260909-lan966x-pci-fdma-v6-0-6f48dab9d671@microchip.com

Changes in v6:

The main change is a DMA bug on the PCIe path, found after enabling the
IOMMU on the test host. The rest is hardening and cleanup.

- New patch 6: use a dedicated device for DMA operations. On the PCIe
  path lan966x->dev is the platform device created by
  of_platform_default_populate(), which carries no iommus/dma-ranges of
  its own, so it is not the device the IOMMU has a domain for. DMA
  mapped against it lands outside the domain the endpoint's requester ID
  is associated with. With the IOMMU enabled, AMD-Vi reported
  IO_PAGE_FAULT for the endpoint and no traffic passed. All FDMA DMA now
  targets the real PCIe endpoint device. Retested with the IOMMU in
  translated (DMA-FQ) mode.
- ATU: serialise region allocation/free and ATU register access with a
  mutex.
- ATU: program the region limit per mapping (base_addr + size - 1)
  instead of the full region bound, and disable every region in
  fdma_pci_atu_init().
- ATU: return -ERANGE instead of -E2BIG when the requested size exceeds
  the region size, and clear fdma->atu_region on unmap.
- PCIe FDMA: check fdma_dcbs_init() and unwind the coherent allocation
  and the ATU mapping on failure.
- PCIe FDMA: drop the WARN_ON() on an out-of-range src_port; a malformed
  IFH should not be able to splat the host.
- Quiesce consistently in shutdown, deinit and the MTU reload: stop the
  netdev queues and NAPI before disabling the FDMA channels, and share
  one lan966x_fdma_tx_disable_netdev() helper instead of open-coding it
  per path.
- shutdown: also mask ANA_ANAINTR, which shares the PCIe INTx with the
  FDMA sources.
- MTU change: reject an MTU whose ring would not fit in a single
  MAX_PAGE_ORDER coherent block, and reject one whose db_size would be
  truncated by the 16-bit DCB DATAL field, rather than programming a
  wrong buffer length.
- Assign lan966x->ops directly in the ops dispatch patch, so that patch
  no longer adds a helper returning a single constant; PCI detection now
  derives from dma_dev and arrives with the PCIe implementation.
- ATU: reject a target address that is not aligned to the 64KB ATU region
  granularity, instead of programming a translation offset by the
  misalignment.
- Drop the CONFIG_MCHP_LAN966X_PCI guard around the fdma_pci.h include in
  fdma_api.h (Jakub).
- Use ifneq ($(CONFIG_MCHP_LAN966X_PCI),) instead of ifdef in both the fdma
  and lan966x Makefiles; the symbol is tristate, so the condition has to
  cover both y and m (Jakub).
- Reprogram the FDMA LLP registers on the platform reload restore path,
  which stopped happening when the LLP write moved into the allocation
  functions.
- XDP: drop the napi_synchronize() before freeing the old program on the
  PCIe path; READ_ONCE() on port->xdp_prog plus the RCU-deferred
  bpf_prog_put() already cover the in-flight poll.
- XDP: note that the ETH_ZLEN floor is deliberately not enforced on XDP_TX.
- PCIe FDMA: count tx_dropped when the DCB ring is full.
- PCIe FDMA: guard the deinit napi_disable() on fdma_ndev, as shutdown()
  already does; NAPI is only initialised on the first ndo_open.
- ATU: program the region translation before enabling the region.
- ATU: pad the mapped allocation to the 64KB outbound region granularity,
  so the window does not extend past the coherent allocation.
- ATU: reject mapping an fdma that already holds a region, instead of
  overwriting the handle and leaking the old mapping. The MTU reload clears
  the handle on the live rings first, since it keeps the old mapping alive
  until the new rings are in place.
- PCIe FDMA: declare the RX DCB DATAL as the space actually available to the
  extraction engine (db_size - XDP_PACKET_HEADROOM), since the data block
  pointer starts that far into the slot.
- Move the XDP-unsupported guards from the XDP patch into the PCIe FDMA
  patch, so that no intermediate commit advertises NETDEV_XDP_ACT_NDO_XMIT
  for a PCIe instance while tx->dcbs_buf is NULL, and none lets an XDP
  attach reach the platform page_pool reload. The final tree is unchanged.
- shutdown: take rtnl, so it cannot interleave with the rtnl-only reload
  paths and double-disable the NAPI, which would spin forever in
  napi_disable().
- Link to v5: https://lore.kernel.org/r/20260520-lan966x-pci-fdma-v5-0-ca56197ae05b@microchip.com

Changes in v5:

This version fixes a single AI review issue, flagged by Paolo. Other AI
issues  for v4 has been classified as pre-existing or changes for
follow-ups.

- Fix premature napi_complete_done() in lan966x_fdma_pci_napi_poll() on
  FDMA_ERROR and napi_alloc_skb() failure. Bailing out left DONE=1 DCBs
  in the ring with no IRQ to drain them. Drop the frame and continue
  the poll loop instead. Bump rx_dropped on memory-pressure drop.
  (Paolo)
- Link to v4: https://lore.kernel.org/r/20260508-lan966x-pci-fdma-v4-0-14e0c89d8d63@microchip.com

Changes in v4:
- Consolidate rx size checks into lan966x_fdma_pci_rx_size_fits().
  Subtract XDP_PACKET_HEADROOM on the max size check, and add ETH_HLEN
  on the min size check. This fixes potential OOB reads/writes.
- On xdp_prepare_buff(), update comment to clarify that data is already
  offset by XDP_PACKET_HEADROOM.
- Link to v3:
  https://lore.kernel.org/r/20260504-lan966x-pci-fdma-v3-0-a56f5740d870@microchip.com

Changes in v3:

Version 3 fixes a number of issues reported by sashiko - mostly
hardening.

- Fix double use of XDP_PACKET_HEADROOM.
- Fix ERR_PTR persistence in fdma->atu_region and add missing
  NULL/ERR_PTR guard in fdma_pci_atu_region_unmap().
- Reject size <= 0 in fdma_pci_atu_region_map() and return
  -ENOSPC (was -ENOMEM) when no region is free.
- Introduce lan966x_fdma_pci_tx_size_fits() that accounts for
  XDP_PACKET_HEADROOM; use it from both xmit paths to keep
  bpf_xdp_adjust_tail from writing past the TX slot.
- Validate BLOCKL in rx_check_frame() (reject < IFH+FCS or
  > db_size) before it feeds memcpy/XDP sizes.
- READ_ONCE(port->xdp_prog) inside lan966x_xdp_pci_run() to close
  a TOCTOU on XDP detach that could deref NULL in
  bpf_prog_run_xdp().
- Strip IFH and FCS pre-XDP in rx_check_frame(). After BPF runs
  the driver cannot tell whether the tail was modified; drop the
  unconditional skb_pull/skb_trim in rx_get_frame().
- Account tx_bytes/tx_packets on XDP_TX success and tx_dropped on
  XDP_TX size reject.
- Add dma_wmb()/dma_rmb() around DCB status writes and reads in
  xmit, xmit_xdpf, and napi_poll.
- Collected Tested-by: Hervé Codina.
- Link to v2: https://lore.kernel.org/r/20260428-lan966x-pci-fdma-v2-0-d3ec66e06202@microchip.com

Changes in v2:

Version 2 primarily addresses issues with module unload/load, where
traffic would stop working (Hervé), and XDP head/tail adjust that would be
discarded (Mohsin).

Apart from that, I ran through issues reported by Sashiko, and fixed a
number of other issues.

- New patch 1: add drivers/net/ethernet/microchip/fdma/ to the Sparx5
  SoC MAINTAINERS entry.
- New patch 7: clear latched FDMA error/interrupt stickies after the
  switch reset so they don't fire as soon as interrupts are enabled.
- New patch 8: shutdown() callback, quiescing FDMA on host warm reboot.
- Replaced the depth-2 dev_is_pci(parent->parent) backend selector
  with a parent-chain walk.
- XDP: use xdp.data/xdp.data_end for the post-XDP frame length so that
  bpf_xdp_adjust_head/tail are respected (Mohsin Bashir)
- MTU change: drain in-flight xmits with netif_tx_disable() on every
  port before reallocating rings, waking them again on completion.
- MTU change: cap the PCIe DCB ring at 256 entries so a full-ring
  coherent DMA allocation fits in a single MAX_PAGE_ORDER block at
  jumbo MTU.
- PCIe ATU: disable the region before clearing its translation on
  unmap.
- PCIe FDMA: hold tx_lock in napi_poll around the free-DCB check used
  to wake stopped netdev queues.
- PCIe FDMA: return -ENOSPC (not -1) when the DCB ring is exhausted.
- Link to v1: https://lore.kernel.org/r/20260320-lan966x-pci-fdma-v1-0-ef54cb9b0c4b@microchip.com

---
Daniel Machon (15):
      MAINTAINERS: add FDMA library to Sparx5 SoC entry
      net: microchip: fdma: rename contiguous dataptr helpers
      net: microchip: fdma: add PCIe ATU support
      net: microchip: fdma: use little-endian types for descriptor fields
      net: lan966x: add FDMA LLP register write helper
      net: lan966x: export FDMA helpers for reuse
      net: lan966x: use a dedicated device for DMA operations
      net: lan966x: add FDMA ops dispatch for PCIe support
      net: lan966x: clear FDMA interrupt stickies after switch reset
      net: lan966x: add shutdown callback to stop the FDMA on reboot
      net: lan966x: add PCIe FDMA support
      net: lan966x: add PCIe FDMA MTU change support
      net: lan966x: add PCIe FDMA XDP support
      misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
      misc: lan966x-pci: dts: add fdma interrupt to overlay

 MAINTAINERS                                        |   1 +
 drivers/misc/lan966x_pci.dtso                      |   5 +-
 drivers/net/ethernet/microchip/fdma/Makefile       |   4 +
 drivers/net/ethernet/microchip/fdma/fdma_api.c     |  65 +-
 drivers/net/ethernet/microchip/fdma/fdma_api.h     |  42 +-
 drivers/net/ethernet/microchip/fdma/fdma_pci.c     | 208 ++++++
 drivers/net/ethernet/microchip/fdma/fdma_pci.h     |  57 ++
 drivers/net/ethernet/microchip/lan966x/Makefile    |   4 +
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c  |  97 ++-
 .../ethernet/microchip/lan966x/lan966x_fdma_pci.c  | 695 +++++++++++++++++++++
 .../net/ethernet/microchip/lan966x/lan966x_main.c  | 121 +++-
 .../net/ethernet/microchip/lan966x/lan966x_main.h  |  69 ++
 .../net/ethernet/microchip/lan966x/lan966x_regs.h  |  25 +
 .../net/ethernet/microchip/lan966x/lan966x_xdp.c   |   6 +
 14 files changed, 1321 insertions(+), 78 deletions(-)
---
base-commit: 447cb143d024d97257cefcdb2192e4b7d2fe47ac
change-id: 20260313-lan966x-pci-fdma-94ed485d23fa

Best regards,
-- 
Daniel Machon <daniel.machon@microchip.com>


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 02/15] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
                   ` (13 subsequent siblings)
  14 siblings, 0 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

The FDMA library under drivers/net/ethernet/microchip/fdma/ is shared by
the lan966x, sparx5 and lan969x drivers, but is not covered by an entry
in the MAINTAINERS file. A subsequent patch will add new files to the
FDMA library, so let's make sure it's covered.

Add drivers/net/ethernet/microchip/fdma/ to the Sparx5 SoC entry, since
I am already listed there.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 MAINTAINERS | 1 +
 1 file changed, 1 insertion(+)

diff --git a/MAINTAINERS b/MAINTAINERS
index e3ce77c839b0..e5da9d0378b6 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -3214,6 +3214,7 @@ M:	UNGLinuxDriver@microchip.com
 L:	linux-arm-kernel@lists.infradead.org (moderated for non-subscribers)
 S:	Supported
 F:	arch/arm64/boot/dts/microchip/sparx*
+F:	drivers/net/ethernet/microchip/fdma/
 F:	drivers/net/ethernet/microchip/vcap/
 F:	drivers/pinctrl/pinctrl-microchip-sgpio.c
 N:	sparx5

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 02/15] net: microchip: fdma: rename contiguous dataptr helpers
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
                   ` (12 subsequent siblings)
  14 siblings, 0 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

When the FDMA library was introduced [1], two helpers to get the DMA and
virtual address of a data buffer, in contiguous memory, were added. These
helpers have had no callers until this series. I found the naming I
initially used confusing and inconsistent.

Rename fdma_dataptr_get_contiguous() and
fdma_dataptr_virt_get_contiguous() to fdma_dataptr_dma_addr_contiguous()
and fdma_dataptr_virt_addr_contiguous(). This makes the pair symmetric
and clarifies what type of address each returns.

[1]: commit 30e48a75df9c ("net: microchip: add FDMA library")

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 drivers/net/ethernet/microchip/fdma/fdma_api.h | 8 ++++----
 1 file changed, 4 insertions(+), 4 deletions(-)

diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index d91affe8bd98..dea8e3cc155a 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -197,8 +197,8 @@ static inline int fdma_nextptr_cb(struct fdma *fdma, int dcb_idx, u64 *nextptr)
  * if the dataptr addresses and DCB's are in contiguous memory and the driver
  * supports XDP.
  */
-static inline u64 fdma_dataptr_get_contiguous(struct fdma *fdma, int dcb_idx,
-					      int db_idx)
+static inline u64 fdma_dataptr_dma_addr_contiguous(struct fdma *fdma,
+						   int dcb_idx, int db_idx)
 {
 	return fdma->dma + (sizeof(struct fdma_dcb) * fdma->n_dcbs) +
 	       (dcb_idx * fdma->n_dbs + db_idx) * fdma->db_size +
@@ -209,8 +209,8 @@ static inline u64 fdma_dataptr_get_contiguous(struct fdma *fdma, int dcb_idx,
  * applicable if the dataptr addresses and DCB's are in contiguous memory and
  * the driver supports XDP.
  */
-static inline void *fdma_dataptr_virt_get_contiguous(struct fdma *fdma,
-						     int dcb_idx, int db_idx)
+static inline void *fdma_dataptr_virt_addr_contiguous(struct fdma *fdma,
+						      int dcb_idx, int db_idx)
 {
 	return (u8 *)fdma->dcbs + (sizeof(struct fdma_dcb) * fdma->n_dcbs) +
 	       (dcb_idx * fdma->n_dbs + db_idx) * fdma->db_size +

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 03/15] net: microchip: fdma: add PCIe ATU support
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 02/15] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  2026-09-24 19:56 ` [PATCH net-next v8 04/15] net: microchip: fdma: use little-endian types for descriptor fields Daniel Machon
                   ` (11 subsequent siblings)
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

When lan966x or lan969x operates as a PCIe endpoint, the internal FDMA
engine cannot directly access host memory. Instead, DMA addresses must
be translated through the PCIe Address Translation Unit (ATU). The ATU
provides outbound windows that map internal addresses to PCIe bus
addresses.

The ATU outbound address space (0x10000000-0x1fffffff) is divided into
six equally-sized regions (~42MB each). When FDMA buffers are allocated,
a free ATU region is claimed and programmed with the DMA target address.
The FDMA engine then uses the region's base address in its descriptors,
and the ATU translates these to the actual DMA addresses on the PCIe bus.

Add the required functions and helpers that combine the DMA allocation
with the ATU region mapping. These are used by the lan966x PCIe FDMA
path.

The ATU cannot express a limit finer than its 64KB region granularity,
so pad the mapped allocation to that boundary; otherwise the outbound
window would extend past the memory the host allocated for DMA.

This implementation will also be used by the lan969x, when PCIe FDMA is
added for that platform in the future.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 drivers/net/ethernet/microchip/fdma/Makefile   |   4 +
 drivers/net/ethernet/microchip/fdma/fdma_api.c |  44 ++++++
 drivers/net/ethernet/microchip/fdma/fdma_api.h |  14 ++
 drivers/net/ethernet/microchip/fdma/fdma_pci.c | 208 +++++++++++++++++++++++++
 drivers/net/ethernet/microchip/fdma/fdma_pci.h |  57 +++++++
 5 files changed, 327 insertions(+)

diff --git a/drivers/net/ethernet/microchip/fdma/Makefile b/drivers/net/ethernet/microchip/fdma/Makefile
index cc9a736be357..910a10b33fb9 100644
--- a/drivers/net/ethernet/microchip/fdma/Makefile
+++ b/drivers/net/ethernet/microchip/fdma/Makefile
@@ -5,3 +5,7 @@
 
 obj-$(CONFIG_FDMA) += fdma.o
 fdma-y += fdma_api.o
+
+ifneq ($(CONFIG_MCHP_LAN966X_PCI),)
+fdma-y += fdma_pci.o
+endif
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.c b/drivers/net/ethernet/microchip/fdma/fdma_api.c
index e78c3590da9e..a3c9e3097c5c 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.c
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.c
@@ -127,6 +127,50 @@ void fdma_free_phys(struct fdma *fdma)
 }
 EXPORT_SYMBOL_GPL(fdma_free_phys);
 
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+/* Allocate coherent DMA memory and map it in the ATU. */
+int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
+				struct fdma_pci_atu *atu)
+{
+	struct fdma_pci_atu_region *region;
+	int err;
+
+	if (WARN_ON(fdma->atu_region))
+		return -EBUSY;
+
+	/* The ATU cannot express a limit finer than the region granularity, so
+	 * the hardware widens the programmed limit to that boundary. Pad the
+	 * allocation to match, or the outbound window would extend past the
+	 * memory we own.
+	 */
+	fdma->size = ALIGN(fdma->size, FDMA_PCI_ATU_REGION_ALIGN);
+
+	err = fdma_alloc_coherent(dev, fdma);
+	if (err)
+		return err;
+
+	region = fdma_pci_atu_region_map(atu, fdma->dma, fdma->size);
+	if (IS_ERR(region)) {
+		fdma_free_coherent(dev, fdma);
+		return PTR_ERR(region);
+	}
+
+	fdma->atu_region = region;
+
+	return 0;
+}
+EXPORT_SYMBOL_GPL(fdma_alloc_coherent_and_map);
+
+/* Free coherent DMA memory and unmap the memory in the ATU. */
+void fdma_free_coherent_and_unmap(struct device *dev, struct fdma *fdma)
+{
+	fdma_pci_atu_region_unmap(fdma->atu_region);
+	fdma->atu_region = NULL;
+	fdma_free_coherent(dev, fdma);
+}
+EXPORT_SYMBOL_GPL(fdma_free_coherent_and_unmap);
+#endif
+
 /* Get the size of the FDMA memory */
 u32 fdma_get_size(struct fdma *fdma)
 {
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index dea8e3cc155a..ccc30d506e89 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -7,6 +7,8 @@
 #include <linux/etherdevice.h>
 #include <linux/types.h>
 
+#include "fdma_pci.h"
+
 /* This provides a common set of functions and data structures for interacting
  * with the Frame DMA engine on multiple Microchip switchcores.
  *
@@ -109,6 +111,11 @@ struct fdma {
 	u32 channel_id;
 
 	struct fdma_ops ops;
+
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+	/* PCI ATU region for this FDMA instance. */
+	struct fdma_pci_atu_region *atu_region;
+#endif
 };
 
 /* Advance the DCB index and wrap if required. */
@@ -233,9 +240,16 @@ int __fdma_dcb_add(struct fdma *fdma, int dcb_idx, u64 info, u64 status,
 
 int fdma_alloc_coherent(struct device *dev, struct fdma *fdma);
 int fdma_alloc_phys(struct fdma *fdma);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
+				struct fdma_pci_atu *atu);
+#endif
 
 void fdma_free_coherent(struct device *dev, struct fdma *fdma);
 void fdma_free_phys(struct fdma *fdma);
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+void fdma_free_coherent_and_unmap(struct device *dev, struct fdma *fdma);
+#endif
 
 u32 fdma_get_size(struct fdma *fdma);
 u32 fdma_get_size_contiguous(struct fdma *fdma);
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.c b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
new file mode 100644
index 000000000000..dd1dc46cbc9d
--- /dev/null
+++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
@@ -0,0 +1,208 @@
+// SPDX-License-Identifier: GPL-2.0+
+
+#include <linux/align.h>
+#include <linux/errno.h>
+#include <linux/io.h>
+#include <linux/kernel.h>
+#include <linux/types.h>
+
+#include "fdma_pci.h"
+
+/* When the switch operates as a PCIe endpoint, the FDMA engine needs to
+ * DMA to/from host memory. The FDMA writes to addresses within the endpoint's
+ * internal Outbound (OB) address space, and the PCIe ATU translates these to
+ * DMA addresses on the PCIe bus, targeting host memory.
+ *
+ * The ATU supports up to six outbound regions. This implementation divides
+ * the OB address space into six equally sized chunks.
+ *
+ * +-------------+------------+------------+-----+------------+
+ * | Index       | Region 0   | Region 1   | ... | Region 5   |
+ * +-------------+------------+------------+-----+------------+
+ * | Base addr   | 0x10000000 | 0x12aa0000 | ... | 0x1d520000 |
+ * | Limit addr  | 0x12a9ffff | 0x1553ffff | ... | 0x1ffbffff |
+ * | Target addr | host dma   | host dma   | ... | host dma   |
+ * +-------------+------------+------------+-----+------------+
+ *
+ * Base addr is the start address of the region within the OB address space.
+ * Limit addr is each region's own upper bound. The value actually
+ * programmed is set per-mapping as base_addr + mapped size - 1, and is
+ * usually smaller than this.
+ * Target addr is the host DMA address that the base addr translates to.
+ */
+
+#define FDMA_PCI_ATU_OB_START        0x10000000
+#define FDMA_PCI_ATU_OB_END          0x1fffffff
+
+#define FDMA_PCI_ATU_ADDR            0x300000
+#define FDMA_PCI_ATU_IDX_SIZE        0x200
+#define FDMA_PCI_ATU_ENA_REG         0x4
+#define FDMA_PCI_ATU_ENA_BIT         BIT(31)
+#define FDMA_PCI_ATU_LWR_BASE_ADDR   0x8
+#define FDMA_PCI_ATU_UPP_BASE_ADDR   0xc
+#define FDMA_PCI_ATU_LIMIT_ADDR      0x10
+#define FDMA_PCI_ATU_LWR_TARGET_ADDR 0x14
+#define FDMA_PCI_ATU_UPP_TARGET_ADDR 0x18
+
+static u32 fdma_pci_atu_region_size(void)
+{
+	return round_down((FDMA_PCI_ATU_OB_END - FDMA_PCI_ATU_OB_START) /
+			  FDMA_PCI_ATU_REGION_MAX, FDMA_PCI_ATU_REGION_ALIGN);
+}
+
+static void __iomem *fdma_pci_atu_addr_get(void __iomem *addr, int offset,
+					   int idx)
+{
+	return addr + FDMA_PCI_ATU_ADDR + FDMA_PCI_ATU_IDX_SIZE * idx + offset;
+}
+
+static void fdma_pci_atu_region_enable(struct fdma_pci_atu_region *region)
+{
+	writel(FDMA_PCI_ATU_ENA_BIT,
+	       fdma_pci_atu_addr_get(region->atu->addr, FDMA_PCI_ATU_ENA_REG,
+				     region->idx));
+}
+
+static void fdma_pci_atu_region_disable(struct fdma_pci_atu_region *region)
+{
+	writel(0, fdma_pci_atu_addr_get(region->atu->addr, FDMA_PCI_ATU_ENA_REG,
+					region->idx));
+}
+
+/* Configure the address translation in the ATU. */
+static void
+fdma_pci_atu_configure_translation(struct fdma_pci_atu_region *region)
+{
+	struct fdma_pci_atu *atu = region->atu;
+	int idx = region->idx;
+
+	writel(lower_32_bits(region->base_addr),
+	       fdma_pci_atu_addr_get(atu->addr,
+				     FDMA_PCI_ATU_LWR_BASE_ADDR, idx));
+
+	writel(upper_32_bits(region->base_addr),
+	       fdma_pci_atu_addr_get(atu->addr,
+				     FDMA_PCI_ATU_UPP_BASE_ADDR, idx));
+
+	/* The OB address space lies entirely below 4GB, so the limit always
+	 * fits the lower limit register and the upper one is left alone.
+	 */
+	writel(region->limit_addr,
+	       fdma_pci_atu_addr_get(atu->addr, FDMA_PCI_ATU_LIMIT_ADDR, idx));
+
+	writel(lower_32_bits(region->target_addr),
+	       fdma_pci_atu_addr_get(atu->addr,
+				     FDMA_PCI_ATU_LWR_TARGET_ADDR, idx));
+
+	writel(upper_32_bits(region->target_addr),
+	       fdma_pci_atu_addr_get(atu->addr,
+				     FDMA_PCI_ATU_UPP_TARGET_ADDR, idx));
+}
+
+/* Find an unused ATU region. */
+static struct fdma_pci_atu_region *
+fdma_pci_atu_region_get_free(struct fdma_pci_atu *atu)
+{
+	struct fdma_pci_atu_region *regions = atu->regions;
+
+	for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
+		if (regions[i].in_use)
+			continue;
+
+		return &regions[i];
+	}
+
+	return ERR_PTR(-ENOSPC);
+}
+
+/* Unmap an ATU region, clearing its translation and disabling it. */
+void fdma_pci_atu_region_unmap(struct fdma_pci_atu_region *region)
+{
+	if (IS_ERR_OR_NULL(region))
+		return;
+
+	mutex_lock(&region->atu->lock);
+
+	region->target_addr = 0;
+	region->in_use = false;
+
+	fdma_pci_atu_region_disable(region);
+	fdma_pci_atu_configure_translation(region);
+
+	mutex_unlock(&region->atu->lock);
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_region_unmap);
+
+/* Map a host DMA address into a free outbound region. */
+struct fdma_pci_atu_region *
+fdma_pci_atu_region_map(struct fdma_pci_atu *atu, u64 target_addr, int size)
+{
+	struct fdma_pci_atu_region *region;
+
+	if (!atu)
+		return ERR_PTR(-EINVAL);
+
+	if (size <= 0)
+		return ERR_PTR(-EINVAL);
+
+	if (size > fdma_pci_atu_region_size())
+		return ERR_PTR(-ERANGE);
+
+	/* The ATU region base is only ever aligned to FDMA_PCI_ATU_REGION_ALIGN;
+	 * require the same alignment of the host target address, since the ATU
+	 * translates addr - target_addr + base_addr and any misalignment here
+	 * would shift every translated address by the same amount.
+	 */
+	if (!IS_ALIGNED(target_addr, FDMA_PCI_ATU_REGION_ALIGN))
+		return ERR_PTR(-EINVAL);
+
+	mutex_lock(&atu->lock);
+
+	region = fdma_pci_atu_region_get_free(atu);
+	if (IS_ERR(region)) {
+		mutex_unlock(&atu->lock);
+		return region;
+	}
+
+	region->target_addr = target_addr;
+	region->limit_addr = region->base_addr + size - 1;
+	region->in_use = true;
+
+	fdma_pci_atu_configure_translation(region);
+	fdma_pci_atu_region_enable(region);
+
+	mutex_unlock(&atu->lock);
+
+	return region;
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_region_map);
+
+/* Translate a host DMA address to the corresponding OB address. */
+u64 fdma_pci_atu_translate_addr(struct fdma_pci_atu_region *region, u64 addr)
+{
+	return region->base_addr + (addr - region->target_addr);
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_translate_addr);
+
+/* Initialize ATU, dividing the OB space into equally sized regions. */
+void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr)
+{
+	struct fdma_pci_atu_region *regions = atu->regions;
+	u32 region_size = fdma_pci_atu_region_size();
+
+	atu->addr = addr;
+	mutex_init(&atu->lock);
+
+	for (int i = 0; i < FDMA_PCI_ATU_REGION_MAX; i++) {
+		regions[i].base_addr =
+			FDMA_PCI_ATU_OB_START + (i * region_size);
+		regions[i].limit_addr = 0;
+		regions[i].target_addr = 0;
+		regions[i].idx = i;
+		regions[i].atu = atu;
+		regions[i].in_use = false;
+
+		fdma_pci_atu_region_disable(&regions[i]);
+	}
+}
+EXPORT_SYMBOL_GPL(fdma_pci_atu_init);
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.h b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
new file mode 100644
index 000000000000..010bdb4d50b0
--- /dev/null
+++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.h
@@ -0,0 +1,57 @@
+/* SPDX-License-Identifier: GPL-2.0+ */
+
+#ifndef _FDMA_PCI_H_
+#define _FDMA_PCI_H_
+
+#include <linux/align.h>
+#include <linux/bits.h>
+#include <linux/mutex.h>
+#include <linux/types.h>
+
+#define FDMA_PCI_ATU_REGION_MAX 6
+
+/* Outbound regions are 64KB granular (datasheet section 3.24.7.4.1), so both
+ * the region base and the mapped size must be aligned to this.
+ */
+#define FDMA_PCI_ATU_REGION_ALIGN BIT(16)
+
+#define FDMA_PCI_DB_ALIGN 128
+#define FDMA_PCI_DB_SIZE(mtu) ALIGN(mtu, FDMA_PCI_DB_ALIGN)
+
+struct fdma_pci_atu;
+
+struct fdma_pci_atu_region {
+	struct fdma_pci_atu *atu;
+	u64 base_addr; /* Base addr of the OB window */
+	u64 limit_addr; /* End addr of the active mapping (base_addr + size - 1) */
+	u64 target_addr; /* Host DMA address this region maps to */
+	int idx;
+	bool in_use;
+};
+
+struct fdma_pci_atu {
+	void __iomem *addr;
+	struct mutex lock; /* Protects region alloc/free and ATU register access */
+	struct fdma_pci_atu_region regions[FDMA_PCI_ATU_REGION_MAX];
+};
+
+/* Initialize ATU, dividing OB space into regions. */
+void fdma_pci_atu_init(struct fdma_pci_atu *atu, void __iomem *addr);
+
+/* Unmap an ATU region, clearing its translation and disabling it. */
+void fdma_pci_atu_region_unmap(struct fdma_pci_atu_region *region);
+
+/* Map a host DMA address into a free ATU region. target_addr and size must be
+ * FDMA_PCI_ATU_REGION_ALIGN aligned; a misaligned target_addr returns -EINVAL.
+ */
+struct fdma_pci_atu_region *fdma_pci_atu_region_map(struct fdma_pci_atu *atu,
+						    u64 target_addr,
+						    int size);
+
+/* Translate a host DMA address to the OB address space. Reads the region
+ * unlocked, so the caller must quiesce DMA and the descriptor paths before
+ * unmapping the region.
+ */
+u64 fdma_pci_atu_translate_addr(struct fdma_pci_atu_region *region, u64 addr);
+
+#endif

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 04/15] net: microchip: fdma: use little-endian types for descriptor fields
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (2 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 05/15] net: lan966x: add FDMA LLP register write helper Daniel Machon
                   ` (10 subsequent siblings)
  14 siblings, 0 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

The FDMA engine reads and writes the DCB and DB descriptors in
little-endian byte order. So far the descriptors have only been produced
by the little-endian SoC itself, so plain u64 fields were fine. With the
PCIe FDMA path, the descriptors live in host memory and are written by
the host CPU, which may be big-endian.

Change the descriptor fields to __le64 and convert at the library
boundary: in __fdma_db_add() and __fdma_dcb_add() on write, and in the
fdma_db_*() accessors on read. The dataptr and nextptr callbacks keep
their u64 signatures, so their implementations are unchanged. Add
fdma_db_dataptr_get(), and convert the lan966x sites that read the
descriptor fields directly to use the accessors.

No functional change on little-endian hosts.

Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
 drivers/net/ethernet/microchip/fdma/fdma_api.c      | 21 ++++++++++++++++-----
 drivers/net/ethernet/microchip/fdma/fdma_api.h      | 20 +++++++++++++-------
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c   |  9 +++++----
 3 files changed, 34 insertions(+), 16 deletions(-)

diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.c b/drivers/net/ethernet/microchip/fdma/fdma_api.c
index a3c9e3097c5c..f7a42348e932 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.c
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.c
@@ -12,10 +12,18 @@ static int __fdma_db_add(struct fdma *fdma, int dcb_idx, int db_idx, u64 status,
 				   int db_idx, u64 *dataptr))
 {
 	struct fdma_db *db = fdma_db_get(fdma, dcb_idx, db_idx);
+	u64 dataptr;
+	int err;
+
+	db->status = cpu_to_le64(status);
 
-	db->status = status;
+	err = cb(fdma, dcb_idx, db_idx, &dataptr);
+	if (unlikely(err))
+		return err;
 
-	return cb(fdma, dcb_idx, db_idx, &db->dataptr);
+	db->dataptr = cpu_to_le64(dataptr);
+
+	return 0;
 }
 
 /* Add a DB to a DCB, using the callback set in the fdma_ops struct. */
@@ -35,6 +43,7 @@ int __fdma_dcb_add(struct fdma *fdma, int dcb_idx, u64 info, u64 status,
 				u64 *dataptr))
 {
 	struct fdma_dcb *dcb = fdma_dcb_get(fdma, dcb_idx);
+	u64 nextptr;
 	int i, err;
 
 	for (i = 0; i < fdma->n_dbs; i++) {
@@ -43,14 +52,16 @@ int __fdma_dcb_add(struct fdma *fdma, int dcb_idx, u64 info, u64 status,
 			return err;
 	}
 
-	err = dcb_cb(fdma, dcb_idx, &fdma->last_dcb->nextptr);
+	err = dcb_cb(fdma, dcb_idx, &nextptr);
 	if (unlikely(err))
 		return err;
 
+	fdma->last_dcb->nextptr = cpu_to_le64(nextptr);
+
 	fdma->last_dcb = dcb;
 
-	dcb->nextptr = FDMA_DCB_INVALID_DATA;
-	dcb->info = info;
+	dcb->nextptr = cpu_to_le64(FDMA_DCB_INVALID_DATA);
+	dcb->info = cpu_to_le64(info);
 
 	return 0;
 }
diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.h b/drivers/net/ethernet/microchip/fdma/fdma_api.h
index ccc30d506e89..4e4f009b77cb 100644
--- a/drivers/net/ethernet/microchip/fdma/fdma_api.h
+++ b/drivers/net/ethernet/microchip/fdma/fdma_api.h
@@ -66,13 +66,13 @@
 struct fdma;
 
 struct fdma_db {
-	u64 dataptr;
-	u64 status;
+	__le64 dataptr;
+	__le64 status;
 };
 
 struct fdma_dcb {
-	u64 nextptr;
-	u64 info;
+	__le64 nextptr;
+	__le64 info;
 	struct fdma_db db[FDMA_DB_MAX];
 };
 
@@ -147,19 +147,25 @@ static inline bool fdma_dcb_is_reusable(struct fdma *fdma)
 /* Check if the FDMA has marked this DB as done. */
 static inline bool fdma_db_is_done(struct fdma_db *db)
 {
-	return db->status & FDMA_DCB_STATUS_DONE;
+	return le64_to_cpu(db->status) & FDMA_DCB_STATUS_DONE;
 }
 
 /* Get the length of a DB. */
 static inline int fdma_db_len_get(struct fdma_db *db)
 {
-	return FDMA_DCB_STATUS_BLOCKL(db->status);
+	return FDMA_DCB_STATUS_BLOCKL(le64_to_cpu(db->status));
+}
+
+/* Get the dataptr of a DB. */
+static inline u64 fdma_db_dataptr_get(struct fdma_db *db)
+{
+	return le64_to_cpu(db->dataptr);
 }
 
 /* Set the length of a DB. */
 static inline void fdma_dcb_len_set(struct fdma_dcb *dcb, u32 len)
 {
-	dcb->info = FDMA_DCB_INFO_DATAL(len);
+	dcb->info = cpu_to_le64(FDMA_DCB_INFO_DATAL(len));
 }
 
 /* Get a DB by index. */
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 41d4ec7f2f57..68fd454ebc98 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -406,8 +406,9 @@ static int lan966x_fdma_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
 		return FDMA_ERROR;
 
 	dma_sync_single_for_cpu(lan966x->dev,
-				(dma_addr_t)db->dataptr + XDP_PACKET_HEADROOM,
-				FDMA_DCB_STATUS_BLOCKL(db->status),
+				(dma_addr_t)fdma_db_dataptr_get(db) +
+				XDP_PACKET_HEADROOM,
+				fdma_db_len_get(db),
 				DMA_FROM_DEVICE);
 
 	lan966x_ifh_get_src_port(page_address(page) + XDP_PACKET_HEADROOM,
@@ -419,7 +420,7 @@ static int lan966x_fdma_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
 	if (!lan966x_xdp_port_present(port))
 		return FDMA_PASS;
 
-	return lan966x_xdp_run(port, page, FDMA_DCB_STATUS_BLOCKL(db->status));
+	return lan966x_xdp_run(port, page, fdma_db_len_get(db));
 }
 
 static struct sk_buff *lan966x_fdma_rx_get_frame(struct lan966x_rx *rx,
@@ -443,7 +444,7 @@ static struct sk_buff *lan966x_fdma_rx_get_frame(struct lan966x_rx *rx,
 	skb_mark_for_recycle(skb);
 
 	skb_reserve(skb, XDP_PACKET_HEADROOM);
-	skb_put(skb, FDMA_DCB_STATUS_BLOCKL(db->status));
+	skb_put(skb, fdma_db_len_get(db));
 
 	lan966x_ifh_get_timestamp(skb->data, &timestamp);
 

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 05/15] net: lan966x: add FDMA LLP register write helper
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (3 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 04/15] net: microchip: fdma: use little-endian types for descriptor fields Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 06/15] net: lan966x: export FDMA helpers for reuse Daniel Machon
                   ` (9 subsequent siblings)
  14 siblings, 0 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

The FDMA Link List Pointer (LLP) register points to the first DCB in the
chain and must be written before the channel is activated. This tells
the FDMA engine where to begin DMA transfers.

Move the LLP register writes from the channel start/activate functions
into the allocation functions and introduce a shared
lan966x_fdma_llp_configure() helper. This is needed because the upcoming
PCIe FDMA path writes ATU-translated addresses to the LLP registers
instead of DMA addresses. Keeping the writes in the shared
start/activate path would overwrite these translated addresses.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c  | 30 ++++++++++------------
 1 file changed, 14 insertions(+), 16 deletions(-)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 68fd454ebc98..1c5484a506ff 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -109,6 +109,13 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
 	return 0;
 }
 
+static void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+				       u8 channel_id)
+{
+	lan_wr(lower_32_bits(addr), lan966x, FDMA_DCB_LLP(channel_id));
+	lan_wr(upper_32_bits(addr), lan966x, FDMA_DCB_LLP1(channel_id));
+}
+
 static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
 {
 	struct lan966x *lan966x = rx->lan966x;
@@ -128,6 +135,8 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
 	fdma_dcbs_init(fdma, FDMA_DCB_INFO_DATAL(fdma->db_size),
 		       FDMA_DCB_STATUS_INTR);
 
+	lan966x_fdma_llp_configure(lan966x, fdma->dma, fdma->channel_id);
+
 	return 0;
 }
 
@@ -137,14 +146,6 @@ static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
 	struct fdma *fdma = &rx->fdma;
 	u32 mask;
 
-	/* When activating a channel, first is required to write the first DCB
-	 * address and then to activate it
-	 */
-	lan_wr(lower_32_bits((u64)fdma->dma), lan966x,
-	       FDMA_DCB_LLP(fdma->channel_id));
-	lan_wr(upper_32_bits((u64)fdma->dma), lan966x,
-	       FDMA_DCB_LLP1(fdma->channel_id));
-
 	lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
 	       FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
 	       FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
@@ -215,6 +216,8 @@ static int lan966x_fdma_tx_alloc(struct lan966x_tx *tx)
 
 	fdma_dcbs_init(fdma, 0, 0);
 
+	lan966x_fdma_llp_configure(lan966x, fdma->dma, fdma->channel_id);
+
 	return 0;
 
 out:
@@ -236,14 +239,6 @@ static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
 	struct fdma *fdma = &tx->fdma;
 	u32 mask;
 
-	/* When activating a channel, first is required to write the first DCB
-	 * address and then to activate it
-	 */
-	lan_wr(lower_32_bits((u64)fdma->dma), lan966x,
-	       FDMA_DCB_LLP(fdma->channel_id));
-	lan_wr(upper_32_bits((u64)fdma->dma), lan966x,
-	       FDMA_DCB_LLP1(fdma->channel_id));
-
 	lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
 	       FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
 	       FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
@@ -877,6 +872,9 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
 					   MEM_TYPE_PAGE_POOL, page_pool);
 	}
 
+	lan966x_fdma_llp_configure(lan966x, lan966x->rx.fdma.dma,
+				   lan966x->rx.fdma.channel_id);
+
 	lan966x_fdma_rx_start(&lan966x->rx);
 
 	lan966x_fdma_wakeup_netdev(lan966x);

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 06/15] net: lan966x: export FDMA helpers for reuse
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (4 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 05/15] net: lan966x: add FDMA LLP register write helper Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 07/15] net: lan966x: use a dedicated device for DMA operations Daniel Machon
                   ` (8 subsequent siblings)
  14 siblings, 0 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

Make shared FDMA helpers non-static, so they can be reused by the PCIe
FDMA implementation.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c  | 22 +++++++++++-----------
 .../net/ethernet/microchip/lan966x/lan966x_main.h  | 11 +++++++++++
 2 files changed, 22 insertions(+), 11 deletions(-)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 1c5484a506ff..651ea03eede7 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -109,8 +109,8 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
 	return 0;
 }
 
-static void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
-				       u8 channel_id)
+void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+				u8 channel_id)
 {
 	lan_wr(lower_32_bits(addr), lan966x, FDMA_DCB_LLP(channel_id));
 	lan_wr(upper_32_bits(addr), lan966x, FDMA_DCB_LLP1(channel_id));
@@ -140,7 +140,7 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
 	return 0;
 }
 
-static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
+void lan966x_fdma_rx_start(struct lan966x_rx *rx)
 {
 	struct lan966x *lan966x = rx->lan966x;
 	struct fdma *fdma = &rx->fdma;
@@ -171,7 +171,7 @@ static void lan966x_fdma_rx_start(struct lan966x_rx *rx)
 		lan966x, FDMA_CH_ACTIVATE);
 }
 
-static void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
+void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
 {
 	struct lan966x *lan966x = rx->lan966x;
 	struct fdma *fdma = &rx->fdma;
@@ -191,7 +191,7 @@ static void lan966x_fdma_rx_disable(struct lan966x_rx *rx)
 		lan966x, FDMA_CH_DB_DISCARD);
 }
 
-static void lan966x_fdma_rx_reload(struct lan966x_rx *rx)
+void lan966x_fdma_rx_reload(struct lan966x_rx *rx)
 {
 	struct lan966x *lan966x = rx->lan966x;
 
@@ -264,7 +264,7 @@ static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
 		lan966x, FDMA_CH_ACTIVATE);
 }
 
-static void lan966x_fdma_tx_disable(struct lan966x_tx *tx)
+void lan966x_fdma_tx_disable(struct lan966x_tx *tx)
 {
 	struct lan966x *lan966x = tx->lan966x;
 	struct fdma *fdma = &tx->fdma;
@@ -296,7 +296,7 @@ static void lan966x_fdma_tx_reload(struct lan966x_tx *tx)
 		lan966x, FDMA_CH_RELOAD);
 }
 
-static void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x)
+void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x)
 {
 	struct lan966x_port *port;
 	int i;
@@ -471,7 +471,7 @@ static struct sk_buff *lan966x_fdma_rx_get_frame(struct lan966x_rx *rx,
 	return NULL;
 }
 
-static int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight)
+int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight)
 {
 	struct lan966x *lan966x = container_of(napi, struct lan966x, napi);
 	struct lan966x_rx *rx = &lan966x->rx;
@@ -584,7 +584,7 @@ static int lan966x_fdma_get_next_dcb(struct lan966x_tx *tx)
 	return -1;
 }
 
-static void lan966x_fdma_tx_start(struct lan966x_tx *tx)
+void lan966x_fdma_tx_start(struct lan966x_tx *tx)
 {
 	struct lan966x *lan966x = tx->lan966x;
 
@@ -802,7 +802,7 @@ static int lan966x_fdma_get_max_mtu(struct lan966x *lan966x)
 	return max_mtu;
 }
 
-static int lan966x_qsys_sw_status(struct lan966x *lan966x)
+int lan966x_qsys_sw_status(struct lan966x *lan966x)
 {
 	return lan_rd(lan966x, QSYS_SW_STATUS(CPU_PORT));
 }
@@ -884,7 +884,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
 	return err;
 }
 
-static int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
+int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
 {
 	return lan966x_fdma_get_max_mtu(lan966x) +
 	       IFH_LEN_BYTES +
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index eea286c29474..83c361abb789 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -561,6 +561,17 @@ int lan966x_fdma_init(struct lan966x *lan966x);
 void lan966x_fdma_deinit(struct lan966x *lan966x);
 irqreturn_t lan966x_fdma_irq_handler(int irq, void *args);
 int lan966x_fdma_reload_page_pool(struct lan966x *lan966x);
+int lan966x_fdma_napi_poll(struct napi_struct *napi, int weight);
+void lan966x_fdma_llp_configure(struct lan966x *lan966x, u64 addr,
+				u8 channel_id);
+void lan966x_fdma_rx_start(struct lan966x_rx *rx);
+void lan966x_fdma_rx_disable(struct lan966x_rx *rx);
+void lan966x_fdma_rx_reload(struct lan966x_rx *rx);
+void lan966x_fdma_tx_start(struct lan966x_tx *tx);
+void lan966x_fdma_tx_disable(struct lan966x_tx *tx);
+void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
+int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
+int lan966x_qsys_sw_status(struct lan966x *lan966x);
 
 int lan966x_lag_port_join(struct lan966x_port *port,
 			  struct net_device *brport_dev,

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 07/15] net: lan966x: use a dedicated device for DMA operations
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (5 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 06/15] net: lan966x: export FDMA helpers for reuse Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 08/15] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
                   ` (7 subsequent siblings)
  14 siblings, 0 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

In preparation for the PCIe FDMA implementation, add lan966x->dma_dev,
resolved once at probe time, and use it for every DMA operation the
FDMA path performs: coherent allocations, streaming mappings and the RX
page pool. Natively, dma_dev is lan966x->dev, so there is no functional
change.

Since dma_dev differs from dev only on the PCIe path, add
lan966x_is_pci() on top of it, for the later PCIe patches to key off.

Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c  | 30 +++++++++++-----------
 .../net/ethernet/microchip/lan966x/lan966x_main.c  | 16 ++++++++++++
 .../net/ethernet/microchip/lan966x/lan966x_main.h  | 10 ++++++++
 3 files changed, 41 insertions(+), 15 deletions(-)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 651ea03eede7..ae6e5aec5e13 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -80,7 +80,7 @@ static int lan966x_fdma_rx_alloc_page_pool(struct lan966x_rx *rx)
 		.flags = PP_FLAG_DMA_MAP | PP_FLAG_DMA_SYNC_DEV,
 		.pool_size = rx->fdma.n_dcbs,
 		.nid = NUMA_NO_NODE,
-		.dev = lan966x->dev,
+		.dev = lan966x->dma_dev,
 		.dma_dir = DMA_FROM_DEVICE,
 		.offset = XDP_PACKET_HEADROOM,
 		.max_len = rx->max_mtu -
@@ -126,7 +126,7 @@ static int lan966x_fdma_rx_alloc(struct lan966x_rx *rx)
 	if (err)
 		return err;
 
-	err = fdma_alloc_coherent(lan966x->dev, fdma);
+	err = fdma_alloc_coherent(lan966x->dma_dev, fdma);
 	if (err) {
 		page_pool_destroy(rx->page_pool);
 		return err;
@@ -210,7 +210,7 @@ static int lan966x_fdma_tx_alloc(struct lan966x_tx *tx)
 	if (!tx->dcbs_buf)
 		return -ENOMEM;
 
-	err = fdma_alloc_coherent(lan966x->dev, fdma);
+	err = fdma_alloc_coherent(lan966x->dma_dev, fdma);
 	if (err)
 		goto out;
 
@@ -230,7 +230,7 @@ static void lan966x_fdma_tx_free(struct lan966x_tx *tx)
 	struct lan966x *lan966x = tx->lan966x;
 
 	kfree(tx->dcbs_buf);
-	fdma_free_coherent(lan966x->dev, &tx->fdma);
+	fdma_free_coherent(lan966x->dma_dev, &tx->fdma);
 }
 
 static void lan966x_fdma_tx_activate(struct lan966x_tx *tx)
@@ -355,7 +355,7 @@ static void lan966x_fdma_tx_clear_buf(struct lan966x *lan966x, int weight)
 
 		dcb_buf->used = false;
 		if (dcb_buf->use_skb) {
-			dma_unmap_single(lan966x->dev,
+			dma_unmap_single(lan966x->dma_dev,
 					 dcb_buf->dma_addr,
 					 dcb_buf->len,
 					 DMA_TO_DEVICE);
@@ -364,7 +364,7 @@ static void lan966x_fdma_tx_clear_buf(struct lan966x *lan966x, int weight)
 				napi_consume_skb(dcb_buf->data.skb, weight);
 		} else {
 			if (dcb_buf->xdp_ndo)
-				dma_unmap_single(lan966x->dev,
+				dma_unmap_single(lan966x->dma_dev,
 						 dcb_buf->dma_addr,
 						 dcb_buf->len,
 						 DMA_TO_DEVICE);
@@ -400,7 +400,7 @@ static int lan966x_fdma_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
 	if (unlikely(!page))
 		return FDMA_ERROR;
 
-	dma_sync_single_for_cpu(lan966x->dev,
+	dma_sync_single_for_cpu(lan966x->dma_dev,
 				(dma_addr_t)fdma_db_dataptr_get(db) +
 				XDP_PACKET_HEADROOM,
 				fdma_db_len_get(db),
@@ -636,11 +636,11 @@ int lan966x_fdma_xmit_xdpf(struct lan966x_port *port, void *ptr, u32 len)
 		lan966x_ifh_set_bypass(ifh, 1);
 		lan966x_ifh_set_port(ifh, BIT_ULL(port->chip_port));
 
-		dma_addr = dma_map_single(lan966x->dev,
+		dma_addr = dma_map_single(lan966x->dma_dev,
 					  xdpf->data - IFH_LEN_BYTES,
 					  xdpf->len + IFH_LEN_BYTES,
 					  DMA_TO_DEVICE);
-		if (dma_mapping_error(lan966x->dev, dma_addr)) {
+		if (dma_mapping_error(lan966x->dma_dev, dma_addr)) {
 			ret = NETDEV_TX_OK;
 			goto out;
 		}
@@ -656,7 +656,7 @@ int lan966x_fdma_xmit_xdpf(struct lan966x_port *port, void *ptr, u32 len)
 		lan966x_ifh_set_port(ifh, BIT_ULL(port->chip_port));
 
 		dma_addr = page_pool_get_dma_addr(page);
-		dma_sync_single_for_device(lan966x->dev,
+		dma_sync_single_for_device(lan966x->dma_dev,
 					   dma_addr + XDP_PACKET_HEADROOM,
 					   len + IFH_LEN_BYTES,
 					   DMA_TO_DEVICE);
@@ -735,9 +735,9 @@ int lan966x_fdma_xmit(struct sk_buff *skb, __be32 *ifh, struct net_device *dev)
 	memcpy(skb->data, ifh, IFH_LEN_BYTES);
 	skb_put(skb, 4);
 
-	dma_addr = dma_map_single(lan966x->dev, skb->data, skb->len,
+	dma_addr = dma_map_single(lan966x->dma_dev, skb->data, skb->len,
 				  DMA_TO_DEVICE);
-	if (dma_mapping_error(lan966x->dev, dma_addr)) {
+	if (dma_mapping_error(lan966x->dma_dev, dma_addr)) {
 		dev->stats.tx_dropped++;
 		err = NETDEV_TX_OK;
 		goto release;
@@ -843,7 +843,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
 			page_pool_put_full_page(page_pool,
 						old_pages[i][j], false);
 
-	fdma_free_coherent(lan966x->dev, &fdma_rx_old);
+	fdma_free_coherent(lan966x->dma_dev, &fdma_rx_old);
 
 	page_pool_destroy(page_pool);
 
@@ -993,7 +993,7 @@ int lan966x_fdma_init(struct lan966x *lan966x)
 
 	err = lan966x_fdma_tx_alloc(&lan966x->tx);
 	if (err) {
-		fdma_free_coherent(lan966x->dev, &lan966x->rx.fdma);
+		fdma_free_coherent(lan966x->dma_dev, &lan966x->rx.fdma);
 		page_pool_destroy(lan966x->rx.page_pool);
 		return err;
 	}
@@ -1015,7 +1015,7 @@ void lan966x_fdma_deinit(struct lan966x *lan966x)
 	napi_disable(&lan966x->napi);
 
 	lan966x_fdma_rx_free_pages(&lan966x->rx);
-	fdma_free_coherent(lan966x->dev, &lan966x->rx.fdma);
+	fdma_free_coherent(lan966x->dma_dev, &lan966x->rx.fdma);
 	page_pool_destroy(lan966x->rx.page_pool);
 	lan966x_fdma_tx_free(&lan966x->tx);
 }
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 1179a6e127c5..2741f7c9fa4c 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -7,6 +7,7 @@
 #include <linux/ip.h>
 #include <linux/of.h>
 #include <linux/of_net.h>
+#include <linux/pci.h>
 #include <linux/phy/phy.h>
 #include <linux/platform_device.h>
 #include <linux/reset.h>
@@ -1081,6 +1082,20 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
 	return 0;
 }
 
+/* When enumerated over PCIe, dev is a platform device with no
+ * iommus/dma-ranges of its own, so DMA must target the PCIe endpoint instead.
+ * The result differs from dev only in that case; lan966x_is_pci() relies on it.
+ */
+static struct device *lan966x_get_dma_dev(struct device *dev)
+{
+	for (struct device *p = dev->parent; p; p = p->parent) {
+		if (dev_is_pci(p))
+			return p;
+	}
+
+	return dev;
+}
+
 static int lan966x_probe(struct platform_device *pdev)
 {
 	struct fwnode_handle *ports, *portnp;
@@ -1094,6 +1109,7 @@ static int lan966x_probe(struct platform_device *pdev)
 
 	platform_set_drvdata(pdev, lan966x);
 	lan966x->dev = &pdev->dev;
+	lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
 
 	if (!device_get_mac_address(&pdev->dev, mac_addr)) {
 		ether_addr_copy(lan966x->base_mac, mac_addr);
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 83c361abb789..3f09f8ddf620 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -270,6 +270,9 @@ struct lan966x_skb_cb {
 struct lan966x {
 	struct device *dev;
 
+	/* Device used for DMA; the PCIe endpoint when enumerated over PCIe. */
+	struct device *dma_dev;
+
 	u8 num_phys_ports;
 	struct lan966x_port **ports;
 
@@ -573,6 +576,13 @@ void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
 int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
 int lan966x_qsys_sw_status(struct lan966x *lan966x);
 
+/* dma_dev differs from dev only on the PCIe path. */
+static inline bool lan966x_is_pci(struct lan966x *lan966x)
+{
+	return IS_ENABLED(CONFIG_MCHP_LAN966X_PCI) &&
+	       lan966x->dma_dev != lan966x->dev;
+}
+
 int lan966x_lag_port_join(struct lan966x_port *port,
 			  struct net_device *brport_dev,
 			  struct net_device *bond,

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 08/15] net: lan966x: add FDMA ops dispatch for PCIe support
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (6 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 07/15] net: lan966x: use a dedicated device for DMA operations Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-24 19:56 ` [PATCH net-next v8 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
                   ` (6 subsequent siblings)
  14 siblings, 0 replies; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

Introduce lan966x_fdma_ops to support different FDMA implementations
for platform and PCIe. Plumb fdma_init, fdma_deinit, fdma_xmit,
fdma_poll and fdma_resize through the ops table, and select the
implementation at probe time. Only the platform implementation exists
at this point; the PCIe implementation is added in a later patch.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c    |  2 +-
 .../net/ethernet/microchip/lan966x/lan966x_main.c    | 20 +++++++++++++++-----
 .../net/ethernet/microchip/lan966x/lan966x_main.h    | 13 +++++++++++++
 3 files changed, 29 insertions(+), 6 deletions(-)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index ae6e5aec5e13..5f46166cbfe4 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -948,7 +948,7 @@ void lan966x_fdma_netdev_init(struct lan966x *lan966x, struct net_device *dev)
 		return;
 
 	lan966x->fdma_ndev = dev;
-	netif_napi_add(dev, &lan966x->napi, lan966x_fdma_napi_poll);
+	netif_napi_add(dev, &lan966x->napi, lan966x->ops->fdma_poll);
 	napi_enable(&lan966x->napi);
 }
 
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 2741f7c9fa4c..6e6c08bb8eea 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -27,6 +27,14 @@
 
 #define IO_RANGES 2
 
+static const struct lan966x_fdma_ops lan966x_fdma_ops = {
+	.fdma_init = &lan966x_fdma_init,
+	.fdma_deinit = &lan966x_fdma_deinit,
+	.fdma_xmit = &lan966x_fdma_xmit,
+	.fdma_poll = &lan966x_fdma_napi_poll,
+	.fdma_resize = &lan966x_fdma_change_mtu,
+};
+
 static const struct of_device_id lan966x_match[] = {
 	{ .compatible = "microchip,lan966x-switch" },
 	{ }
@@ -392,7 +400,7 @@ static netdev_tx_t lan966x_port_xmit(struct sk_buff *skb,
 
 	spin_lock(&lan966x->tx_lock);
 	if (port->lan966x->fdma)
-		err = lan966x_fdma_xmit(skb, ifh, dev);
+		err = lan966x->ops->fdma_xmit(skb, ifh, dev);
 	else
 		err = lan966x_port_ifh_xmit(skb, ifh, dev);
 	spin_unlock(&lan966x->tx_lock);
@@ -414,7 +422,7 @@ static int lan966x_port_change_mtu(struct net_device *dev, int new_mtu)
 	if (!lan966x->fdma)
 		return 0;
 
-	err = lan966x_fdma_change_mtu(lan966x);
+	err = lan966x->ops->fdma_resize(lan966x);
 	if (err) {
 		lan_wr(DEV_MAC_MAXLEN_CFG_MAX_LEN_SET(LAN966X_HW_MTU(old_mtu)),
 		       lan966x, DEV_MAC_MAXLEN_CFG(port->chip_port));
@@ -1111,6 +1119,8 @@ static int lan966x_probe(struct platform_device *pdev)
 	lan966x->dev = &pdev->dev;
 	lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
 
+	lan966x->ops = &lan966x_fdma_ops;
+
 	if (!device_get_mac_address(&pdev->dev, mac_addr)) {
 		ether_addr_copy(lan966x->base_mac, mac_addr);
 	} else {
@@ -1250,7 +1260,7 @@ static int lan966x_probe(struct platform_device *pdev)
 	if (err)
 		goto cleanup_fdb;
 
-	err = lan966x_fdma_init(lan966x);
+	err = lan966x->ops->fdma_init(lan966x);
 	if (err)
 		goto cleanup_ptp;
 
@@ -1263,7 +1273,7 @@ static int lan966x_probe(struct platform_device *pdev)
 	return 0;
 
 cleanup_fdma:
-	lan966x_fdma_deinit(lan966x);
+	lan966x->ops->fdma_deinit(lan966x);
 
 cleanup_ptp:
 	lan966x_ptp_deinit(lan966x);
@@ -1291,7 +1301,7 @@ static void lan966x_remove(struct platform_device *pdev)
 
 	lan966x_taprio_deinit(lan966x);
 	lan966x_vcap_deinit(lan966x);
-	lan966x_fdma_deinit(lan966x);
+	lan966x->ops->fdma_deinit(lan966x);
 	lan966x_cleanup_ports(lan966x);
 
 	cancel_delayed_work_sync(&lan966x->stats_work);
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 3f09f8ddf620..aab5e53ed059 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -193,6 +193,17 @@ enum vcap_is1_port_sel_rt {
 	VCAP_IS1_PS_RT_FOLLOW_OTHER = 7,
 };
 
+struct lan966x;
+
+struct lan966x_fdma_ops {
+	int (*fdma_init)(struct lan966x *lan966x);
+	void (*fdma_deinit)(struct lan966x *lan966x);
+	int (*fdma_xmit)(struct sk_buff *skb, __be32 *ifh,
+			 struct net_device *dev);
+	int (*fdma_poll)(struct napi_struct *napi, int weight);
+	int (*fdma_resize)(struct lan966x *lan966x);
+};
+
 struct lan966x_port;
 
 struct lan966x_rx {
@@ -273,6 +284,8 @@ struct lan966x {
 	/* Device used for DMA; the PCIe endpoint when enumerated over PCIe. */
 	struct device *dma_dev;
 
+	const struct lan966x_fdma_ops *ops;
+
 	u8 num_phys_ports;
 	struct lan966x_port **ports;
 

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (7 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 08/15] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  2026-09-24 19:56 ` [PATCH net-next v8 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
                   ` (5 subsequent siblings)
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

When in PCI mode, the GCB soft reset issued by the reset controller
can latch spurious bits in the FDMA error stickies. The latched bits
sit in FDMA_INTR_ERR until the FDMA IRQ is requested later in probe,
at which point the handler fires immediately and WARNs.

Clear FDMA_ERRORS, FDMA_INTR_ERR and FDMA_INTR_DB right after the
switch reset so the FDMA comes out clean and the IRQ handler does not
see ghost errors on probe.

The clear runs on both the PCI and platform paths. On the platform
path it has no effect — there are no spurious stickies to clear — but
keeping it unconditional avoids a PCI-specific code path here.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 drivers/net/ethernet/microchip/lan966x/lan966x_main.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 6e6c08bb8eea..259d81e75907 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -1067,6 +1067,15 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
 
 	reset_control_reset(switch_reset);
 
+	/* When in PCI mode, the GCB soft reset issued by the reset
+	 * controller can latch spurious bits in the FDMA error and
+	 * data-block stickies. Clear them before request_irq hooks up the
+	 * FDMA IRQ line, otherwise the handler fires immediately on probe.
+	 */
+	lan_wr(lan_rd(lan966x, FDMA_ERRORS), lan966x, FDMA_ERRORS);
+	lan_wr(lan_rd(lan966x, FDMA_INTR_ERR), lan966x, FDMA_INTR_ERR);
+	lan_wr(lan_rd(lan966x, FDMA_INTR_DB), lan966x, FDMA_INTR_DB);
+
 	/* Don't reinitialize the switch core, if it is already initialized. In
 	 * case it is initialized twice, some pointers inside the queue system
 	 * in HW will get corrupted and then after a while the queue system gets

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (8 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  2026-09-24 19:56 ` [PATCH net-next v8 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
                   ` (4 subsequent siblings)
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

As a PCIe endpoint, lan966x is not reset by a host reboot: its FDMA
channels and interrupt sources stay armed, and the OIC ORs every
source into the shared PCIe INTx, asserted before the driver has
re-probed. A still-active channel also keeps write access to host
memory the next kernel will reuse.

Add a shutdown callback that:
- frees the ana, xtr and FDMA irqs, masking and unmapping them at
  the OIC (disable_irq() would leave both set - the OIC has no
  irq_disable())
- masks the analyzer source, armed unconditionally by lan966x_init()
  and re-armed by the MAC table's age timer
- stops and detaches the netdevs, draining in-flight xmit and
  clearing netif_device_present() so ndo_open/ndo_change_mtu cannot
  re-enter the FDMA against a disabled NAPI
- disables both FDMA channels and masks their interrupts
- unmaps the outbound ATU windows, leaving none armed (the windows
  are mapped by the PCIe FDMA backend added later in this series)

NAPI is skipped when fdma_ndev is unset (a probed switch with no
usable port never adds one), and XDP attach cannot re-enter the FDMA
either: on PCIe, lan966x_xdp_setup() returns before the page pool
reload, which is the only point where it touches the FDMA.

Only the PCIe instantiation needs this - the SoC one resets with the
chip - so the callback returns early on a platform device; the check
is at runtime since .shutdown belongs to the driver, and a
PCIe-enabled kernel binds both.

FDMA_INTR_ENA persists across a warm reboot, so also restore the
full enable in lan966x_fdma_rx_start(), run after both rings are
allocated. The PCIe FDMA backend added later in this series starts
RX through the same function, so both backends are re-armed from
one site.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c  |  4 ++
 .../net/ethernet/microchip/lan966x/lan966x_main.c  | 56 ++++++++++++++++++++++
 .../net/ethernet/microchip/lan966x/lan966x_regs.h  | 15 ++++++
 3 files changed, 75 insertions(+)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 5f46166cbfe4..2e8f786d6fee 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -146,6 +146,10 @@ void lan966x_fdma_rx_start(struct lan966x_rx *rx)
 	struct fdma *fdma = &rx->fdma;
 	u32 mask;
 
+	lan_wr(FDMA_INTR_ENA_INTR_PORT_ENA_SET(GENMASK(1, 0)) |
+	       FDMA_INTR_ENA_INTR_CH_ENA_SET(GENMASK(7, 0)),
+	       lan966x, FDMA_INTR_ENA);
+
 	lan_wr(FDMA_CH_CFG_CH_DCB_DB_CNT_SET(fdma->n_dbs) |
 	       FDMA_CH_CFG_CH_INTR_DB_EOF_ONLY_SET(1) |
 	       FDMA_CH_CFG_CH_INJ_PORT_SET(0) |
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 259d81e75907..024ce9f9916c 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -1324,9 +1324,65 @@ static void lan966x_remove(struct platform_device *pdev)
 	debugfs_remove_recursive(lan966x->debugfs_root);
 }
 
+static void lan966x_shutdown(struct platform_device *pdev)
+{
+	struct lan966x *lan966x = platform_get_drvdata(pdev);
+
+	/* As a PCIe endpoint the switch is not reset by the host reboot, so it
+	 * has to be quiesced here:
+	 *
+	 *   Free the irqs and mask the sources: no source can assert INTx.
+	 *   Disable NAPI: the teardown must not race a poll.
+	 *   Stop and detach the netdevs: drains xmit, closes ndo_open and MTU.
+	 *   Stop the FDMA channels: waits for the engine to go idle.
+	 *   Unmap the ATU windows: revokes the engine's access to host memory.
+	 */
+	if (!lan966x_is_pci(lan966x))
+		return;
+
+	if (lan966x->xtr_irq > 0)
+		devm_free_irq(lan966x->dev, lan966x->xtr_irq, lan966x);
+	if (lan966x->ana_irq > 0)
+		devm_free_irq(lan966x->dev, lan966x->ana_irq, lan966x);
+	if (lan966x->fdma_irq > 0)
+		devm_free_irq(lan966x->dev, lan966x->fdma_irq, lan966x);
+
+	lan_wr(0, lan966x, ANA_ANAINTR);
+
+	if (!lan966x->fdma)
+		return;
+
+	rtnl_lock();
+
+	if (lan966x->fdma_ndev)
+		napi_disable(&lan966x->napi);
+
+	for (int p = 0; p < lan966x->num_phys_ports; p++) {
+		if (!lan966x->ports[p] || !lan966x->ports[p]->dev)
+			continue;
+
+		netif_tx_disable(lan966x->ports[p]->dev);
+		netif_device_detach(lan966x->ports[p]->dev);
+	}
+
+	lan966x_fdma_rx_disable(&lan966x->rx);
+	lan966x_fdma_tx_disable(&lan966x->tx);
+
+	lan_wr(0, lan966x, FDMA_INTR_ENA);
+	lan_wr(0, lan966x, FDMA_INTR_DB_ENA);
+
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+	fdma_pci_atu_region_unmap(lan966x->rx.fdma.atu_region);
+	fdma_pci_atu_region_unmap(lan966x->tx.fdma.atu_region);
+#endif
+
+	rtnl_unlock();
+}
+
 static struct platform_driver lan966x_driver = {
 	.probe = lan966x_probe,
 	.remove = lan966x_remove,
+	.shutdown = lan966x_shutdown,
 	.driver = {
 		.name = "lan966x-switch",
 		.of_match_table = lan966x_match,
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
index 4b553927d2e0..aba0d36ae6b5 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
@@ -1039,6 +1039,21 @@ enum lan966x_target {
 /*      FDMA:FDMA:FDMA_INTR_ERR */
 #define FDMA_INTR_ERR             __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 400, 0, 1, 4)
 
+/*      FDMA:FDMA:FDMA_INTR_ENA */
+#define FDMA_INTR_ENA             __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 404, 0, 1, 4)
+
+#define FDMA_INTR_ENA_INTR_PORT_ENA              GENMASK(9, 8)
+#define FDMA_INTR_ENA_INTR_PORT_ENA_SET(x)\
+	FIELD_PREP(FDMA_INTR_ENA_INTR_PORT_ENA, x)
+#define FDMA_INTR_ENA_INTR_PORT_ENA_GET(x)\
+	FIELD_GET(FDMA_INTR_ENA_INTR_PORT_ENA, x)
+
+#define FDMA_INTR_ENA_INTR_CH_ENA                GENMASK(7, 0)
+#define FDMA_INTR_ENA_INTR_CH_ENA_SET(x)\
+	FIELD_PREP(FDMA_INTR_ENA_INTR_CH_ENA, x)
+#define FDMA_INTR_ENA_INTR_CH_ENA_GET(x)\
+	FIELD_GET(FDMA_INTR_ENA_INTR_CH_ENA, x)
+
 /*      FDMA:FDMA:FDMA_ERRORS */
 #define FDMA_ERRORS               __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 412, 0, 1, 4)
 

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 11/15] net: lan966x: add PCIe FDMA support
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (9 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
@ 2026-09-24 19:56 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  2026-09-24 19:57 ` [PATCH net-next v8 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
                   ` (3 subsequent siblings)
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:56 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

Add PCIe FDMA support for lan966x. The PCIe FDMA path uses contiguous
DMA buffers mapped through the endpoint's ATU, with memcpy-based frame
transfer instead of per-page DMA mappings.

With PCIe FDMA, throughput increases from ~33 Mbps (register-based I/O)
to ~620 Mbps on an Intel x86 host with a lan966x PCIe card.

XDP is not supported on this path yet, so do not advertise xdp_features
and reject program attach for PCIe instances.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
---
 drivers/net/ethernet/microchip/lan966x/Makefile    |   4 +
 .../ethernet/microchip/lan966x/lan966x_fdma_pci.c  | 421 +++++++++++++++++++++
 .../net/ethernet/microchip/lan966x/lan966x_main.c  |  11 +-
 .../net/ethernet/microchip/lan966x/lan966x_main.h  |   7 +
 .../net/ethernet/microchip/lan966x/lan966x_regs.h  |  10 +
 .../net/ethernet/microchip/lan966x/lan966x_xdp.c   |   6 +
 6 files changed, 456 insertions(+), 3 deletions(-)

diff --git a/drivers/net/ethernet/microchip/lan966x/Makefile b/drivers/net/ethernet/microchip/lan966x/Makefile
index 4cdbe263502c..02db4d4d9826 100644
--- a/drivers/net/ethernet/microchip/lan966x/Makefile
+++ b/drivers/net/ethernet/microchip/lan966x/Makefile
@@ -18,6 +18,10 @@ lan966x-switch-objs  := lan966x_main.o lan966x_phylink.o lan966x_port.o \
 lan966x-switch-$(CONFIG_LAN966X_DCB) += lan966x_dcb.o
 lan966x-switch-$(CONFIG_DEBUG_FS) += lan966x_vcap_debugfs.o
 
+ifneq ($(CONFIG_MCHP_LAN966X_PCI),)
+lan966x-switch-y += lan966x_fdma_pci.o
+endif
+
 # Provide include files
 ccflags-y += -I$(srctree)/drivers/net/ethernet/microchip/vcap
 ccflags-y += -I$(srctree)/drivers/net/ethernet/microchip/fdma
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
new file mode 100644
index 000000000000..bccd1b8590d7
--- /dev/null
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -0,0 +1,421 @@
+// SPDX-License-Identifier: GPL-2.0+
+
+#include "fdma_api.h"
+#include "lan966x_main.h"
+
+static int lan966x_fdma_pci_dataptr_cb(struct fdma *fdma, int dcb, int db,
+				       u64 *dataptr)
+{
+	u64 addr;
+
+	addr = fdma_dataptr_dma_addr_contiguous(fdma, dcb, db);
+
+	*dataptr = fdma_pci_atu_translate_addr(fdma->atu_region, addr);
+
+	return 0;
+}
+
+static int lan966x_fdma_pci_nextptr_cb(struct fdma *fdma, int dcb, u64 *nextptr)
+{
+	u64 addr;
+
+	fdma_nextptr_cb(fdma, dcb, &addr);
+
+	*nextptr = fdma_pci_atu_translate_addr(fdma->atu_region, addr);
+
+	return 0;
+}
+
+/* Stop the TX queues on every port, so nothing feeds the injection channel
+ * while it is torn down or resized.
+ */
+static void lan966x_fdma_tx_disable_netdev(struct lan966x *lan966x)
+{
+	struct lan966x_port *port;
+	int i;
+
+	for (i = 0; i < lan966x->num_phys_ports; ++i) {
+		port = lan966x->ports[i];
+		if (!port)
+			continue;
+
+		netif_tx_disable(port->dev);
+	}
+}
+
+static int lan966x_fdma_pci_rx_alloc(struct lan966x_rx *rx)
+{
+	struct lan966x *lan966x = rx->lan966x;
+	struct fdma *fdma = &rx->fdma;
+	int err;
+
+	err = fdma_alloc_coherent_and_map(lan966x->dma_dev, fdma,
+					  &lan966x->atu);
+	if (err)
+		return err;
+
+	err = fdma_dcbs_init(fdma,
+			     FDMA_DCB_INFO_DATAL(fdma->db_size - XDP_PACKET_HEADROOM),
+			     FDMA_DCB_STATUS_INTR);
+	if (err) {
+		fdma_free_coherent_and_unmap(lan966x->dma_dev, fdma);
+		return err;
+	}
+
+	lan966x_fdma_llp_configure(lan966x,
+				   fdma->atu_region->base_addr,
+				   fdma->channel_id);
+
+	return 0;
+}
+
+static int lan966x_fdma_pci_tx_alloc(struct lan966x_tx *tx)
+{
+	struct lan966x *lan966x = tx->lan966x;
+	struct fdma *fdma = &tx->fdma;
+	int err;
+
+	err = fdma_alloc_coherent_and_map(lan966x->dma_dev, fdma,
+					  &lan966x->atu);
+	if (err)
+		return err;
+
+	err = fdma_dcbs_init(fdma,
+			     FDMA_DCB_INFO_DATAL(fdma->db_size),
+			     FDMA_DCB_STATUS_DONE);
+	if (err) {
+		fdma_free_coherent_and_unmap(lan966x->dma_dev, fdma);
+		return err;
+	}
+
+	lan966x_fdma_llp_configure(lan966x,
+				   fdma->atu_region->base_addr,
+				   fdma->channel_id);
+
+	return 0;
+}
+
+static int lan966x_fdma_pci_get_next_dcb(struct fdma *fdma)
+{
+	struct fdma_db *db;
+
+	for (int i = 0; i < fdma->n_dcbs; i++) {
+		db = fdma_db_get(fdma, i, 0);
+
+		if (!fdma_db_is_done(db))
+			continue;
+		if (fdma_is_last(fdma, &fdma->dcbs[i]))
+			continue;
+
+		return i;
+	}
+
+	return -ENOSPC;
+}
+
+/* TX slot layout (sizes in bytes):
+ *
+ *  +---------------------+-----+---------+-----+
+ *  | XDP_PACKET_HEADROOM | IFH | payload | FCS |
+ *  |         256         |  28 |   len   |   4 |
+ *  +---------------------+-----+---------+-----+
+ *  |<---------------- db_size ----------------->|
+ *
+ * Return true if the frame plus required overhead fits.
+ */
+static bool lan966x_fdma_pci_tx_size_fits(struct fdma *fdma, u32 len)
+{
+	return XDP_PACKET_HEADROOM + IFH_LEN_BYTES + len + ETH_FCS_LEN <=
+	       fdma->db_size;
+}
+
+/* Return true if blockl is a valid RX frame size. */
+static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
+{
+	return blockl >= IFH_LEN_BYTES + ETH_HLEN + ETH_FCS_LEN &&
+	       blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
+}
+
+static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
+{
+	struct lan966x *lan966x = rx->lan966x;
+	struct fdma *fdma = &rx->fdma;
+	struct lan966x_port *port;
+	struct fdma_db *db;
+	void *virt_addr;
+	u32 blockl;
+
+	/* virt_addr points to the IFH. */
+	virt_addr = fdma_dataptr_virt_addr_contiguous(fdma,
+						      fdma->dcb_index,
+						      fdma->db_index);
+
+	lan966x_ifh_get_src_port(virt_addr, src_port);
+
+	if (*src_port >= lan966x->num_phys_ports)
+		return FDMA_ERROR;
+
+	port = lan966x->ports[*src_port];
+	if (!port)
+		return FDMA_ERROR;
+
+	db = fdma_db_next_get(fdma);
+
+	/* BLOCKL is a 16-bit HW-populated field; reject obviously-bad
+	 * values before they feed memcpy/XDP sizes.
+	 */
+	blockl = fdma_db_len_get(db);
+	if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
+		return FDMA_ERROR;
+
+	return FDMA_PASS;
+}
+
+static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
+						     u64 src_port)
+{
+	struct lan966x *lan966x = rx->lan966x;
+	struct fdma *fdma = &rx->fdma;
+	struct sk_buff *skb;
+	struct fdma_db *db;
+	u32 data_len;
+
+	/* Get the received frame and create an SKB for it. */
+	db = fdma_db_next_get(fdma);
+	data_len = fdma_db_len_get(db);
+
+	skb = napi_alloc_skb(&lan966x->napi, data_len);
+	if (unlikely(!skb))
+		return NULL;
+
+	memcpy(skb->data,
+	       fdma_dataptr_virt_addr_contiguous(fdma,
+						 fdma->dcb_index,
+						 fdma->db_index),
+						 data_len);
+
+	skb_put(skb, data_len);
+
+	skb->dev = lan966x->ports[src_port]->dev;
+	skb_pull(skb, IFH_LEN_BYTES);
+
+	skb_trim(skb, skb->len - ETH_FCS_LEN);
+
+	skb->protocol = eth_type_trans(skb, skb->dev);
+
+	if (lan966x->bridge_mask & BIT(src_port)) {
+		skb->offload_fwd_mark = 1;
+
+		skb_reset_network_header(skb);
+		if (!lan966x_hw_offload(lan966x, src_port, skb))
+			skb->offload_fwd_mark = 0;
+	}
+
+	skb->dev->stats.rx_bytes += skb->len;
+	skb->dev->stats.rx_packets++;
+
+	return skb;
+}
+
+static int lan966x_fdma_pci_xmit(struct sk_buff *skb, __be32 *ifh,
+				 struct net_device *dev)
+{
+	struct lan966x_port *port = netdev_priv(dev);
+	struct lan966x *lan966x = port->lan966x;
+	struct lan966x_tx *tx = &lan966x->tx;
+	struct fdma *fdma = &tx->fdma;
+	int next_to_use;
+	void *virt_addr;
+
+	next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
+
+	if (next_to_use < 0) {
+		netif_stop_queue(dev);
+		return NETDEV_TX_BUSY;
+	}
+
+	if (skb_put_padto(skb, ETH_ZLEN)) {
+		dev->stats.tx_dropped++;
+		return NETDEV_TX_OK;
+	}
+
+	if (!lan966x_fdma_pci_tx_size_fits(fdma, skb->len)) {
+		dev_kfree_skb_any(skb);
+		dev->stats.tx_dropped++;
+		return NETDEV_TX_OK;
+	}
+
+	skb_tx_timestamp(skb);
+
+	/* virt_addr points to the IFH. */
+	virt_addr = fdma_dataptr_virt_addr_contiguous(fdma, next_to_use, 0);
+	memcpy(virt_addr, ifh, IFH_LEN_BYTES);
+	memcpy(virt_addr + IFH_LEN_BYTES, skb->data, skb->len);
+
+	/* Order frame write before DCB status write below. */
+	dma_wmb();
+
+	fdma_dcb_add(fdma,
+		     next_to_use,
+		     0,
+		     FDMA_DCB_STATUS_INTR |
+		     FDMA_DCB_STATUS_SOF |
+		     FDMA_DCB_STATUS_EOF |
+		     FDMA_DCB_STATUS_BLOCKO(0) |
+		     FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + skb->len + ETH_FCS_LEN));
+
+	/* Start the transmission. */
+	lan966x_fdma_tx_start(tx);
+
+	dev->stats.tx_bytes += skb->len;
+	dev->stats.tx_packets++;
+
+	/* Safe to free: PTP is not supported on the PCIe path yet,
+	 * so lan966x->ptp is always 0 here.
+	 */
+	dev_consume_skb_any(skb);
+
+	return NETDEV_TX_OK;
+}
+
+static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
+{
+	struct lan966x *lan966x = container_of(napi, struct lan966x, napi);
+	struct lan966x_rx *rx = &lan966x->rx;
+	struct fdma *fdma = &rx->fdma;
+	int dcb_reload, old_dcb;
+	struct sk_buff *skb;
+	int counter = 0;
+	u64 src_port;
+
+	/* Wake any stopped TX queues if a TX DCB is available. */
+	spin_lock(&lan966x->tx_lock);
+	if (lan966x_fdma_pci_get_next_dcb(&lan966x->tx.fdma) >= 0)
+		lan966x_fdma_wakeup_netdev(lan966x);
+	spin_unlock(&lan966x->tx_lock);
+
+	dcb_reload = fdma->dcb_index;
+
+	/* Get all received skbs. */
+	while (counter < weight) {
+		if (!fdma_has_frames(fdma))
+			break;
+		/* Order DONE read before DCB/frame reads below. */
+		dma_rmb();
+		counter++;
+		switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
+		case FDMA_PASS:
+			break;
+		case FDMA_ERROR:
+			/* No rx_dropped increment here because src_port is
+			 * invalid.
+			 */
+			fdma_dcb_advance(fdma);
+			continue;
+		}
+		skb = lan966x_fdma_pci_rx_get_frame(rx, src_port);
+		fdma_dcb_advance(fdma);
+		if (!skb) {
+			lan966x->ports[src_port]->dev->stats.rx_dropped++;
+			continue;
+		}
+
+		napi_gro_receive(&lan966x->napi, skb);
+	}
+	while (dcb_reload != fdma->dcb_index) {
+		old_dcb = dcb_reload;
+		dcb_reload++;
+		dcb_reload &= fdma->n_dcbs - 1;
+
+		fdma_dcb_add(fdma,
+			     old_dcb,
+			     FDMA_DCB_INFO_DATAL(fdma->db_size - XDP_PACKET_HEADROOM),
+			     FDMA_DCB_STATUS_INTR);
+
+		lan966x_fdma_rx_reload(rx);
+	}
+
+	if (counter < weight && napi_complete_done(napi, counter))
+		lan_wr(0xff, lan966x, FDMA_INTR_DB_ENA);
+
+	return counter;
+}
+
+static int lan966x_fdma_pci_init(struct lan966x *lan966x)
+{
+	struct fdma *rx_fdma = &lan966x->rx.fdma;
+	struct fdma *tx_fdma = &lan966x->tx.fdma;
+	int err;
+
+	if (!lan966x->fdma)
+		return 0;
+
+	lan_wr(FDMA_CTRL_NRESET_SET(0), lan966x, FDMA_CTRL);
+	lan_wr(FDMA_CTRL_NRESET_SET(1), lan966x, FDMA_CTRL);
+
+	fdma_pci_atu_init(&lan966x->atu, lan966x->regs[TARGET_PCIE_DBI]);
+
+	lan966x->rx.lan966x = lan966x;
+	lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
+	rx_fdma->channel_id = FDMA_XTR_CHANNEL;
+	rx_fdma->n_dcbs = FDMA_DCB_MAX;
+	rx_fdma->n_dbs = FDMA_RX_DCB_MAX_DBS;
+	rx_fdma->priv = lan966x;
+	rx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+	rx_fdma->size = fdma_get_size_contiguous(rx_fdma);
+	rx_fdma->ops.nextptr_cb = &lan966x_fdma_pci_nextptr_cb;
+	rx_fdma->ops.dataptr_cb = &lan966x_fdma_pci_dataptr_cb;
+
+	lan966x->tx.lan966x = lan966x;
+	tx_fdma->channel_id = FDMA_INJ_CHANNEL;
+	tx_fdma->n_dcbs = FDMA_DCB_MAX;
+	tx_fdma->n_dbs = FDMA_TX_DCB_MAX_DBS;
+	tx_fdma->priv = lan966x;
+	tx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+	tx_fdma->size = fdma_get_size_contiguous(tx_fdma);
+	tx_fdma->ops.nextptr_cb = &lan966x_fdma_pci_nextptr_cb;
+	tx_fdma->ops.dataptr_cb = &lan966x_fdma_pci_dataptr_cb;
+
+	err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
+	if (err)
+		return err;
+
+	err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
+	if (err) {
+		fdma_free_coherent_and_unmap(lan966x->dma_dev, rx_fdma);
+		return err;
+	}
+
+	lan966x_fdma_rx_start(&lan966x->rx);
+
+	return 0;
+}
+
+static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
+{
+	return -EOPNOTSUPP;
+}
+
+static void lan966x_fdma_pci_deinit(struct lan966x *lan966x)
+{
+	if (!lan966x->fdma)
+		return;
+
+	if (lan966x->fdma_ndev)
+		napi_disable(&lan966x->napi);
+
+	lan966x_fdma_tx_disable_netdev(lan966x);
+	lan966x_fdma_rx_disable(&lan966x->rx);
+	lan966x_fdma_tx_disable(&lan966x->tx);
+
+	fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->rx.fdma);
+	fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->tx.fdma);
+}
+
+const struct lan966x_fdma_ops lan966x_fdma_pci_ops = {
+	.fdma_init = &lan966x_fdma_pci_init,
+	.fdma_deinit = &lan966x_fdma_pci_deinit,
+	.fdma_xmit = &lan966x_fdma_pci_xmit,
+	.fdma_poll = &lan966x_fdma_pci_napi_poll,
+	.fdma_resize = &lan966x_fdma_pci_resize,
+};
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index 024ce9f9916c..de2202786826 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -50,6 +50,7 @@ struct lan966x_main_io_resource {
 static const struct lan966x_main_io_resource lan966x_main_iomap[] =  {
 	{ TARGET_CPU,                   0xc0000, 0 }, /* 0xe00c0000 */
 	{ TARGET_FDMA,                  0xc0400, 0 }, /* 0xe00c0400 */
+	{ TARGET_PCIE_DBI,             0x400000, 0 }, /* 0xe0400000 */
 	{ TARGET_ORG,                         0, 1 }, /* 0xe2000000 */
 	{ TARGET_GCB,                    0x4000, 1 }, /* 0xe2004000 */
 	{ TARGET_QS,                     0x8000, 1 }, /* 0xe2008000 */
@@ -873,7 +874,8 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
 
 	port->phylink = phylink;
 
-	if (lan966x->fdma)
+	/* XDP is not supported on the PCIe FDMA path. */
+	if (lan966x->fdma && !lan966x_is_pci(lan966x))
 		dev->xdp_features = NETDEV_XDP_ACT_BASIC |
 				    NETDEV_XDP_ACT_REDIRECT |
 				    NETDEV_XDP_ACT_NDO_XMIT;
@@ -1128,7 +1130,8 @@ static int lan966x_probe(struct platform_device *pdev)
 	lan966x->dev = &pdev->dev;
 	lan966x->dma_dev = lan966x_get_dma_dev(lan966x->dev);
 
-	lan966x->ops = &lan966x_fdma_ops;
+	lan966x->ops = lan966x_is_pci(lan966x) ? &lan966x_fdma_pci_ops :
+						 &lan966x_fdma_ops;
 
 	if (!device_get_mac_address(&pdev->dev, mac_addr)) {
 		ether_addr_copy(lan966x->base_mac, mac_addr);
@@ -1187,7 +1190,9 @@ static int lan966x_probe(struct platform_device *pdev)
 		if (err)
 			return dev_err_probe(&pdev->dev, err, "Unable to use ptp irq");
 
-		lan966x->ptp = 1;
+		/* PTP is not supported on the PCIe path yet. */
+		if (!lan966x_is_pci(lan966x))
+			lan966x->ptp = 1;
 	}
 
 	lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index aab5e53ed059..16bc28c8f11f 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -17,6 +17,7 @@
 #include <net/xdp.h>
 
 #include <fdma_api.h>
+#include <fdma_pci.h>
 #include <vcap_api.h>
 #include <vcap_api_client.h>
 
@@ -291,6 +292,10 @@ struct lan966x {
 
 	void __iomem *regs[NUM_TARGETS];
 
+#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
+	struct fdma_pci_atu atu;
+#endif
+
 	int shared_queue_sz;
 
 	u8 base_mac[ETH_ALEN];
@@ -589,6 +594,8 @@ void lan966x_fdma_wakeup_netdev(struct lan966x *lan966x);
 int lan966x_fdma_get_max_frame(struct lan966x *lan966x);
 int lan966x_qsys_sw_status(struct lan966x *lan966x);
 
+extern const struct lan966x_fdma_ops lan966x_fdma_pci_ops;
+
 /* dma_dev differs from dev only on the PCIe path. */
 static inline bool lan966x_is_pci(struct lan966x *lan966x)
 {
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
index aba0d36ae6b5..4778ea217673 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_regs.h
@@ -20,6 +20,7 @@ enum lan966x_target {
 	TARGET_FDMA = 21,
 	TARGET_GCB = 27,
 	TARGET_ORG = 36,
+	TARGET_PCIE_DBI = 40,
 	TARGET_PTP = 41,
 	TARGET_QS = 42,
 	TARGET_QSYS = 46,
@@ -1009,6 +1010,15 @@ enum lan966x_target {
 #define FDMA_CH_CFG_CH_MEM_GET(x)\
 	FIELD_GET(FDMA_CH_CFG_CH_MEM, x)
 
+/*      FDMA:FDMA:FDMA_CTRL */
+#define FDMA_CTRL                 __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 424, 0, 1, 4)
+
+#define FDMA_CTRL_NRESET                         BIT(0)
+#define FDMA_CTRL_NRESET_SET(x)\
+	FIELD_PREP(FDMA_CTRL_NRESET, x)
+#define FDMA_CTRL_NRESET_GET(x)\
+	FIELD_GET(FDMA_CTRL_NRESET, x)
+
 /*      FDMA:FDMA:FDMA_PORT_CTRL */
 #define FDMA_PORT_CTRL(r)         __REG(TARGET_FDMA, 0, 1, 8, 0, 1, 428, 376, r, 2, 4)
 
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
index 9ee61db8690b..e63634887f64 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
@@ -20,6 +20,12 @@ static int lan966x_xdp_setup(struct net_device *dev, struct netdev_bpf *xdp)
 		return -EOPNOTSUPP;
 	}
 
+	if (lan966x_is_pci(lan966x)) {
+		NL_SET_ERR_MSG_MOD(xdp->extack,
+				   "XDP is not supported on the PCIe FDMA path");
+		return -EOPNOTSUPP;
+	}
+
 	old_xdp = lan966x_xdp_present(lan966x);
 	old_prog = xchg(&port->xdp_prog, xdp->prog);
 	new_xdp = lan966x_xdp_present(lan966x);

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 12/15] net: lan966x: add PCIe FDMA MTU change support
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (10 preceding siblings ...)
  2026-09-24 19:56 ` [PATCH net-next v8 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
@ 2026-09-24 19:57 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  2026-09-24 19:57 ` [PATCH net-next v8 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
                   ` (2 subsequent siblings)
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:57 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

Add MTU change support for the PCIe FDMA path: on an MTU change, the
contiguous ATU-mapped RX and TX buffers are reallocated at the new
size, falling back to the existing buffers on failure.

Cap the PCIe DCB ring at 256 (FDMA_PCI_DCB_MAX): 512 DCBs would
overflow MAX_PAGE_ORDER at jumbo MTU.

The ring must fit one MAX_PAGE_ORDER block after ATU padding, and
db_size is handed to the FDMA in the 16-bit DCB DATAL field.
Advertise the resulting limit in dev->max_mtu (FDMA_PCI_MAX_MTU)
when the FDMA is in use - a switch with no "fdma" interrupt named
uses the unconstrained register-based path instead, so the cap is
gated on lan966x->fdma, not lan966x_is_pci() alone. On a 4KB-page,
MAX_PAGE_ORDER=10 build this is MTU 15498; the overhead is shared
with lan966x_fdma_get_max_frame() via FDMA_OVERHEAD.

Skip the resize until lan966x_fdma_pci_init() has built the rings;
it runs after the netdevs register and sizes them from
DEV_MAC_MAXLEN_CFG, already programmed by the caller.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 .../net/ethernet/microchip/lan966x/lan966x_fdma.c  |   6 +-
 .../ethernet/microchip/lan966x/lan966x_fdma_pci.c  | 153 ++++++++++++++++++++-
 .../net/ethernet/microchip/lan966x/lan966x_main.c  |   3 +-
 .../net/ethernet/microchip/lan966x/lan966x_main.h  |  28 ++++
 4 files changed, 181 insertions(+), 9 deletions(-)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
index 2e8f786d6fee..a7940eca5df3 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
@@ -890,11 +890,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
 
 int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
 {
-	return lan966x_fdma_get_max_mtu(lan966x) +
-	       IFH_LEN_BYTES +
-	       SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) +
-	       VLAN_HLEN * 2 +
-	       XDP_PACKET_HEADROOM;
+	return lan966x_fdma_get_max_mtu(lan966x) + FDMA_OVERHEAD;
 }
 
 static int __lan966x_fdma_reload(struct lan966x *lan966x, int max_mtu)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
index bccd1b8590d7..7185e65dda43 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -358,7 +358,7 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
 	lan966x->rx.lan966x = lan966x;
 	lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
 	rx_fdma->channel_id = FDMA_XTR_CHANNEL;
-	rx_fdma->n_dcbs = FDMA_DCB_MAX;
+	rx_fdma->n_dcbs = FDMA_PCI_DCB_MAX;
 	rx_fdma->n_dbs = FDMA_RX_DCB_MAX_DBS;
 	rx_fdma->priv = lan966x;
 	rx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
@@ -368,7 +368,7 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
 
 	lan966x->tx.lan966x = lan966x;
 	tx_fdma->channel_id = FDMA_INJ_CHANNEL;
-	tx_fdma->n_dcbs = FDMA_DCB_MAX;
+	tx_fdma->n_dcbs = FDMA_PCI_DCB_MAX;
 	tx_fdma->n_dbs = FDMA_TX_DCB_MAX_DBS;
 	tx_fdma->priv = lan966x;
 	tx_fdma->db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
@@ -391,9 +391,156 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)
 	return 0;
 }
 
+/* Reset existing rx and tx buffers. */
+static void lan966x_fdma_pci_reset_mem(struct lan966x *lan966x)
+{
+	struct lan966x_rx *rx = &lan966x->rx;
+	struct lan966x_tx *tx = &lan966x->tx;
+
+	memset(rx->fdma.dcbs, 0, rx->fdma.size);
+	memset(tx->fdma.dcbs, 0, tx->fdma.size);
+
+	fdma_dcbs_init(&rx->fdma,
+		       FDMA_DCB_INFO_DATAL(rx->fdma.db_size - XDP_PACKET_HEADROOM),
+		       FDMA_DCB_STATUS_INTR);
+
+	fdma_dcbs_init(&tx->fdma,
+		       FDMA_DCB_INFO_DATAL(tx->fdma.db_size),
+		       FDMA_DCB_STATUS_DONE);
+
+	lan966x_fdma_llp_configure(lan966x,
+				   tx->fdma.atu_region->base_addr,
+				   tx->fdma.channel_id);
+	lan966x_fdma_llp_configure(lan966x,
+				   rx->fdma.atu_region->base_addr,
+				   rx->fdma.channel_id);
+}
+
+/* Wake all TX queues on every port (undoes lan966x_fdma_tx_disable_netdev). */
+static void lan966x_fdma_pci_wakeup_netdev(struct lan966x *lan966x)
+{
+	for (int i = 0; i < lan966x->num_phys_ports; ++i) {
+		struct lan966x_port *port = lan966x->ports[i];
+
+		if (port)
+			netif_tx_wake_all_queues(port->dev);
+	}
+}
+
+static int lan966x_fdma_pci_reload(struct lan966x *lan966x, int new_mtu)
+{
+	struct fdma tx_fdma_old = lan966x->tx.fdma;
+	struct fdma rx_fdma_old = lan966x->rx.fdma;
+	u32 old_mtu = lan966x->rx.max_mtu;
+	int err;
+
+	napi_disable(&lan966x->napi);
+	lan966x_fdma_tx_disable_netdev(lan966x);
+	lan966x_fdma_rx_disable(&lan966x->rx);
+	lan966x_fdma_tx_disable(&lan966x->tx);
+
+	lan966x->rx.max_mtu = new_mtu;
+
+	/* Must be NULL'ed in order to realloc them. */
+	lan966x->rx.fdma.atu_region = NULL;
+	lan966x->tx.fdma.atu_region = NULL;
+
+	lan966x->tx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+	lan966x->tx.fdma.size = fdma_get_size_contiguous(&lan966x->tx.fdma);
+	lan966x->rx.fdma.db_size = FDMA_PCI_DB_SIZE(lan966x->rx.max_mtu);
+	lan966x->rx.fdma.size = fdma_get_size_contiguous(&lan966x->rx.fdma);
+
+	err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
+	if (err)
+		goto restore;
+
+	err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
+	if (err) {
+		fdma_free_coherent_and_unmap(lan966x->dma_dev,
+					     &lan966x->rx.fdma);
+		goto restore;
+	}
+
+	/* Free and unmap old memory. */
+	fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
+	fdma_free_coherent_and_unmap(lan966x->dma_dev, &tx_fdma_old);
+
+	/* Order matters: napi_enable() must precede the wakes, or a TX that
+	 * completes first clears FDMA_INTR_DB_ENA with nothing scheduled to
+	 * restore it, leaving RX dead until the next reload.
+	 */
+	napi_enable(&lan966x->napi);
+	lan966x_fdma_rx_start(&lan966x->rx);
+	lan966x_fdma_pci_wakeup_netdev(lan966x);
+
+	return err;
+restore:
+
+	/* No new buffers are allocated at this point. Use the old buffers,
+	 * but reset them before starting the FDMA again.
+	 */
+
+	memcpy(&lan966x->tx.fdma, &tx_fdma_old, sizeof(struct fdma));
+	memcpy(&lan966x->rx.fdma, &rx_fdma_old, sizeof(struct fdma));
+
+	lan966x->rx.max_mtu = old_mtu;
+
+	lan966x_fdma_pci_reset_mem(lan966x);
+
+	napi_enable(&lan966x->napi);
+	lan966x_fdma_rx_start(&lan966x->rx);
+	lan966x_fdma_pci_wakeup_netdev(lan966x);
+
+	return err;
+}
+
+static int __lan966x_fdma_pci_reload(struct lan966x *lan966x, int max_mtu)
+{
+	int err;
+	u32 val;
+
+	/* Disable the CPU port. */
+	lan_rmw(QSYS_SW_PORT_MODE_PORT_ENA_SET(0),
+		QSYS_SW_PORT_MODE_PORT_ENA,
+		lan966x, QSYS_SW_PORT_MODE(CPU_PORT));
+
+	/* Flush the CPU queues. */
+	readx_poll_timeout(lan966x_qsys_sw_status,
+			   lan966x,
+			   val,
+			   !(QSYS_SW_STATUS_EQ_AVAIL_GET(val)),
+			   READL_SLEEP_US, READL_TIMEOUT_US);
+
+	/* Add a sleep in case there are frames between the queues and the CPU
+	 * port
+	 */
+	usleep_range(USEC_PER_MSEC, 2 * USEC_PER_MSEC);
+
+	err = lan966x_fdma_pci_reload(lan966x, max_mtu);
+
+	/* Enable back the CPU port. */
+	lan_rmw(QSYS_SW_PORT_MODE_PORT_ENA_SET(1),
+		QSYS_SW_PORT_MODE_PORT_ENA,
+		lan966x, QSYS_SW_PORT_MODE(CPU_PORT));
+
+	return err;
+}
+
 static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
 {
-	return -EOPNOTSUPP;
+	int max_mtu;
+
+	/* Nothing to resize until fdma_pci_init() has built the rings; it
+	 * sizes them from DEV_MAC_MAXLEN_CFG, which the caller already set.
+	 */
+	if (!lan966x->rx.lan966x)
+		return 0;
+
+	max_mtu = lan966x_fdma_get_max_frame(lan966x);
+	if (max_mtu == lan966x->rx.max_mtu)
+		return 0;
+
+	return __lan966x_fdma_pci_reload(lan966x, max_mtu);
 }
 
 static void lan966x_fdma_pci_deinit(struct lan966x *lan966x)
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index de2202786826..c3afc4cc597f 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -823,7 +823,8 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
 	port->chip_port = p;
 	lan966x->ports[p] = port;
 
-	dev->max_mtu = ETH_MAX_MTU;
+	dev->max_mtu = lan966x_is_pci(lan966x) && lan966x->fdma ?
+		       FDMA_PCI_MAX_MTU : ETH_MAX_MTU;
 
 	dev->netdev_ops = &lan966x_port_netdev_ops;
 	dev->ethtool_ops = &lan966x_ethtool_ops;
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
index 16bc28c8f11f..1877f1916d71 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
@@ -7,6 +7,7 @@
 #include <linux/etherdevice.h>
 #include <linux/if_vlan.h>
 #include <linux/jiffies.h>
+#include <linux/mmzone.h>
 #include <linux/phy.h>
 #include <linux/phylink.h>
 #include <linux/ptp_clock_kernel.h>
@@ -87,6 +88,33 @@
 #define FDMA_INJ_CHANNEL		0
 #define FDMA_DCB_MAX			512
 
+/* Ring must fit in one MAX_PAGE_ORDER DMA block; 512 DCBs overflows
+ * at jumbo MTU.
+ */
+#define FDMA_PCI_DCB_MAX		256
+
+#define FDMA_OVERHEAD							\
+	(IFH_LEN_BYTES +						\
+	 SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) +		\
+	 VLAN_HLEN * 2 +						\
+	 XDP_PACKET_HEADROOM)
+
+/* Largest db_size keeping the ATU-padded ring inside one MAX_PAGE_ORDER
+ * block and within the 16-bit DCB DATAL field. Inverts ALIGN(x, R) <= L
+ * into x <= ALIGN_DOWN(L, R) to bound x directly.
+ */
+#define FDMA_PCI_DB_SIZE_MAX						\
+	MIN_T(u32,							\
+	      (ALIGN_DOWN(PAGE_SIZE << MAX_PAGE_ORDER,			\
+			  FDMA_PCI_ATU_REGION_ALIGN) -			\
+	       FDMA_PCI_DCB_MAX * sizeof(struct fdma_dcb)) /		\
+	      (FDMA_PCI_DCB_MAX * FDMA_RX_DCB_MAX_DBS),			\
+	      ALIGN_DOWN(GENMASK(15, 0), FDMA_PCI_DB_ALIGN))
+
+#define FDMA_PCI_MAX_MTU						\
+	(FDMA_PCI_DB_SIZE_MAX - FDMA_OVERHEAD -				\
+	 (ETH_HLEN + ETH_FCS_LEN))
+
 #define SE_IDX_QUEUE			0  /* 0-79 : Queue scheduler elements */
 #define SE_IDX_PORT			80 /* 80-89 : Port schedular elements */
 

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 13/15] net: lan966x: add PCIe FDMA XDP support
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (11 preceding siblings ...)
  2026-09-24 19:57 ` [PATCH net-next v8 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
@ 2026-09-24 19:57 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  2026-09-24 19:57 ` [PATCH net-next v8 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
  2026-09-24 19:57 ` [PATCH net-next v8 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:57 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

Add XDP support for the PCIe FDMA path. The implementation operates on
contiguous ATU-mapped buffers with memcpy-based XDP_TX, unlike the
platform path which uses page_pool.

XDP sees the frame with IFH and FCS stripped. These are removed in
lan966x_fdma_pci_rx_check_frame() before the BPF program runs, because
after the program returns the driver cannot tell whether the tail
region was modified. The skb_pull/skb_trim previously done in
lan966x_fdma_pci_rx_get_frame() are removed for the same reason; the
frame pointer and length are pre-computed by rx_check_frame() and
passed through rx_get_frame() and lan966x_xdp_pci_run() to the caller.

lan966x_fdma_pci_xmit_xdpf() handles XDP_TX: it rebuilds a fresh IFH
in the TX slot, copies the post-XDP frame after it, and lets HW insert
a new FCS.

lan966x_xdp_setup() is extended so the PCIe path skips the page_pool
reload that the platform path needs.

Only XDP_ACT_BASIC is supported.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 .../ethernet/microchip/lan966x/lan966x_fdma_pci.c  | 167 ++++++++++++++++++---
 .../net/ethernet/microchip/lan966x/lan966x_main.c  |  12 +-
 .../net/ethernet/microchip/lan966x/lan966x_xdp.c   |  12 +-
 3 files changed, 160 insertions(+), 31 deletions(-)

diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
index 7185e65dda43..216e9cbcd158 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
@@ -1,5 +1,7 @@
 // SPDX-License-Identifier: GPL-2.0+
 
+#include <linux/bpf_trace.h>
+
 #include "fdma_api.h"
 #include "lan966x_main.h"
 
@@ -136,7 +138,123 @@ static bool lan966x_fdma_pci_rx_size_fits(struct fdma *fdma, u32 blockl)
 	       blockl <= fdma->db_size - XDP_PACKET_HEADROOM;
 }
 
-static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
+static int lan966x_fdma_pci_xmit_xdpf(struct lan966x_port *port,
+				      void *ptr, u32 len)
+{
+	struct lan966x *lan966x = port->lan966x;
+	struct lan966x_tx *tx = &lan966x->tx;
+	struct fdma *fdma = &tx->fdma;
+	int next_to_use, ret = 0;
+	void *virt_addr;
+
+	spin_lock(&lan966x->tx_lock);
+
+	next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
+
+	if (next_to_use < 0) {
+		netif_stop_queue(port->dev);
+		port->dev->stats.tx_dropped++;
+		ret = NETDEV_TX_BUSY;
+		goto out;
+	}
+
+	/* Only the upper bound is enforced: XDP owns the frame contents and
+	 * length, so a program that shrinks below ETH_ZLEN gets what it asked
+	 * for.
+	 */
+	if (!lan966x_fdma_pci_tx_size_fits(fdma, len)) {
+		port->dev->stats.tx_dropped++;
+		ret = -EINVAL;
+		goto out;
+	}
+
+	/* virt_addr points to the IFH. */
+	virt_addr = fdma_dataptr_virt_addr_contiguous(fdma, next_to_use, 0);
+
+	/* Construct a fresh IFH. */
+	memset(virt_addr, 0, IFH_LEN_BYTES);
+	lan966x_ifh_set_bypass(virt_addr, 1);
+	lan966x_ifh_set_port(virt_addr, BIT_ULL(port->chip_port));
+
+	/* Copy the (post-XDP) frame after the IFH. */
+	memcpy(virt_addr + IFH_LEN_BYTES, ptr, len);
+
+	/* Order frame write before DCB status write below. */
+	dma_wmb();
+
+	/* Reserve ETH_FCS_LEN for the HW-inserted FCS (len is FCS-stripped). */
+	fdma_dcb_add(fdma,
+		     next_to_use,
+		     0,
+		     FDMA_DCB_STATUS_INTR |
+		     FDMA_DCB_STATUS_SOF |
+		     FDMA_DCB_STATUS_EOF |
+		     FDMA_DCB_STATUS_BLOCKO(0) |
+		     FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + len + ETH_FCS_LEN));
+
+	/* Start the transmission. */
+	lan966x_fdma_tx_start(tx);
+
+	port->dev->stats.tx_bytes += len;
+	port->dev->stats.tx_packets++;
+
+out:
+	spin_unlock(&lan966x->tx_lock);
+
+	return ret;
+}
+
+static int lan966x_xdp_pci_run(struct lan966x_port *port, void *data,
+			       u32 data_len, void **xdp_data, u32 *xdp_len)
+{
+	/* Read once so the NULL check and bpf_prog_run_xdp() see the same
+	 * pointer.
+	 */
+	struct bpf_prog *xdp_prog = READ_ONCE(port->xdp_prog);
+	struct lan966x *lan966x = port->lan966x;
+	struct fdma *fdma = &lan966x->rx.fdma;
+	struct xdp_buff xdp;
+	u32 act;
+
+	if (!xdp_prog)
+		return FDMA_PASS;
+
+	xdp_init_buff(&xdp, fdma->db_size, &port->xdp_rxq);
+
+	/* hard_start is set to slot start (virt_addr is XDP_PACKET_HEADROOM
+	 * into the slot). Headroom includes the IFH; BPF may grow into it
+	 * via adjust_head. IFH is rebuilt on XDP_TX and unread on XDP_PASS.
+	 */
+	xdp_prepare_buff(&xdp,
+			 data - XDP_PACKET_HEADROOM,
+			 XDP_PACKET_HEADROOM + IFH_LEN_BYTES,
+			 data_len,
+			 false);
+
+	act = bpf_prog_run_xdp(xdp_prog, &xdp);
+
+	*xdp_data = xdp.data;
+	*xdp_len = xdp.data_end - xdp.data;
+
+	switch (act) {
+	case XDP_PASS:
+		return FDMA_PASS;
+	case XDP_TX:
+		return lan966x_fdma_pci_xmit_xdpf(port, *xdp_data, *xdp_len) ?
+		       FDMA_DROP : FDMA_TX;
+	default:
+		bpf_warn_invalid_xdp_action(port->dev, xdp_prog, act);
+		fallthrough;
+	case XDP_ABORTED:
+		trace_xdp_exception(port->dev, xdp_prog, act);
+		fallthrough;
+	case XDP_DROP:
+		return FDMA_DROP;
+	}
+}
+
+static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port,
+					   void **data, u32 *data_len)
 {
 	struct lan966x *lan966x = rx->lan966x;
 	struct fdma *fdma = &rx->fdma;
@@ -168,38 +286,33 @@ static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
 	if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
 		return FDMA_ERROR;
 
-	return FDMA_PASS;
+	/* Present the Ethernet frame (no IFH, no FCS). HW re-inserts the
+	 * FCS on TX; see lan966x_fdma_pci_xmit_xdpf(). May be overridden
+	 * by XDP. The FCS strip is unconditional because NETIF_F_RXFCS
+	 * is not advertised in hw_features.
+	 */
+	*data = virt_addr + IFH_LEN_BYTES;
+	*data_len = blockl - IFH_LEN_BYTES - ETH_FCS_LEN;
+
+	return lan966x_xdp_pci_run(port, virt_addr, *data_len, data, data_len);
 }
 
 static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
-						     u64 src_port)
+						     u64 src_port, void *data,
+						     u32 data_len)
 {
 	struct lan966x *lan966x = rx->lan966x;
-	struct fdma *fdma = &rx->fdma;
 	struct sk_buff *skb;
-	struct fdma_db *db;
-	u32 data_len;
-
-	/* Get the received frame and create an SKB for it. */
-	db = fdma_db_next_get(fdma);
-	data_len = fdma_db_len_get(db);
 
 	skb = napi_alloc_skb(&lan966x->napi, data_len);
 	if (unlikely(!skb))
 		return NULL;
 
-	memcpy(skb->data,
-	       fdma_dataptr_virt_addr_contiguous(fdma,
-						 fdma->dcb_index,
-						 fdma->db_index),
-						 data_len);
+	memcpy(skb->data, data, data_len);
 
 	skb_put(skb, data_len);
 
 	skb->dev = lan966x->ports[src_port]->dev;
-	skb_pull(skb, IFH_LEN_BYTES);
-
-	skb_trim(skb, skb->len - ETH_FCS_LEN);
 
 	skb->protocol = eth_type_trans(skb, skb->dev);
 
@@ -287,6 +400,8 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
 	struct sk_buff *skb;
 	int counter = 0;
 	u64 src_port;
+	u32 data_len;
+	void *data;
 
 	/* Wake any stopped TX queues if a TX DCB is available. */
 	spin_lock(&lan966x->tx_lock);
@@ -303,7 +418,10 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
 		/* Order DONE read before DCB/frame reads below. */
 		dma_rmb();
 		counter++;
-		switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
+		switch (lan966x_fdma_pci_rx_check_frame(rx,
+							&src_port,
+							&data,
+							&data_len)) {
 		case FDMA_PASS:
 			break;
 		case FDMA_ERROR:
@@ -312,8 +430,17 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
 			 */
 			fdma_dcb_advance(fdma);
 			continue;
+		case FDMA_TX:
+			fdma_dcb_advance(fdma);
+			continue;
+		case FDMA_DROP:
+			fdma_dcb_advance(fdma);
+			continue;
 		}
-		skb = lan966x_fdma_pci_rx_get_frame(rx, src_port);
+		skb = lan966x_fdma_pci_rx_get_frame(rx,
+						    src_port,
+						    data,
+						    data_len);
 		fdma_dcb_advance(fdma);
 		if (!skb) {
 			lan966x->ports[src_port]->dev->stats.rx_dropped++;
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
index c3afc4cc597f..e4c9c5812cca 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
@@ -875,11 +875,13 @@ static int lan966x_probe_port(struct lan966x *lan966x, u32 p,
 
 	port->phylink = phylink;
 
-	/* XDP is not supported on the PCIe FDMA path. */
-	if (lan966x->fdma && !lan966x_is_pci(lan966x))
-		dev->xdp_features = NETDEV_XDP_ACT_BASIC |
-				    NETDEV_XDP_ACT_REDIRECT |
-				    NETDEV_XDP_ACT_NDO_XMIT;
+	if (lan966x->fdma) {
+		dev->xdp_features = NETDEV_XDP_ACT_BASIC;
+
+		if (!lan966x_is_pci(lan966x))
+			dev->xdp_features |= NETDEV_XDP_ACT_REDIRECT |
+					     NETDEV_XDP_ACT_NDO_XMIT;
+	}
 
 	err = register_netdev(dev);
 	if (err) {
diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
index e63634887f64..b98426afc785 100644
--- a/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
+++ b/drivers/net/ethernet/microchip/lan966x/lan966x_xdp.c
@@ -20,16 +20,16 @@ static int lan966x_xdp_setup(struct net_device *dev, struct netdev_bpf *xdp)
 		return -EOPNOTSUPP;
 	}
 
-	if (lan966x_is_pci(lan966x)) {
-		NL_SET_ERR_MSG_MOD(xdp->extack,
-				   "XDP is not supported on the PCIe FDMA path");
-		return -EOPNOTSUPP;
-	}
-
 	old_xdp = lan966x_xdp_present(lan966x);
 	old_prog = xchg(&port->xdp_prog, xdp->prog);
 	new_xdp = lan966x_xdp_present(lan966x);
 
+	/* PCIe FDMA uses contiguous buffers, so no page_pool reload
+	 * is needed.
+	 */
+	if (lan966x_is_pci(lan966x))
+		goto out;
+
 	if (old_xdp == new_xdp)
 		goto out;
 

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (12 preceding siblings ...)
  2026-09-24 19:57 ` [PATCH net-next v8 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
@ 2026-09-24 19:57 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  2026-09-24 19:57 ` [PATCH net-next v8 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:57 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

The ATU outbound windows used by the FDMA engine are programmed through
registers at offset 0x400000+, which falls outside the current cpu reg
mapping. Extend the cpu reg size from 0x100000 (1MB) to 0x800000 (8MB)
to cover the full PCIE DBI and iATU register space.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 drivers/misc/lan966x_pci.dtso | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
index 7b196b0a0eb6..7bb726550caf 100644
--- a/drivers/misc/lan966x_pci.dtso
+++ b/drivers/misc/lan966x_pci.dtso
@@ -135,7 +135,7 @@ lan966x_phy1: ethernet-lan966x_phy@2 {
 
 				switch: switch@e0000000 {
 					compatible = "microchip,lan966x-switch";
-					reg = <0xe0000000 0x0100000>,
+					reg = <0xe0000000 0x0800000>,
 					      <0xe2000000 0x0800000>;
 					reg-names = "cpu", "gcb";
 

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* [PATCH net-next v8 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay
  2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
                   ` (13 preceding siblings ...)
  2026-09-24 19:57 ` [PATCH net-next v8 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
@ 2026-09-24 19:57 ` Daniel Machon
  2026-09-25 20:52   ` netdev-bot+sashiko
  14 siblings, 1 reply; 24+ messages in thread
From: Daniel Machon @ 2026-09-24 19:57 UTC (permalink / raw)
  To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Horatiu Vultur, Steen Hegelund, UNGLinuxDriver,
	Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
	John Fastabend, Stanislav Fomichev, Herve Codina, Arnd Bergmann,
	Greg Kroah-Hartman, Mohsin Bashir, Simon Horman
  Cc: Richard Cochran, netdev, linux-kernel, bpf, linux-arm-kernel

Add the fdma interrupt (OIC interrupt 14) to the lan966x PCI device
tree overlay, enabling FDMA-based frame injection/extraction when
the switch is connected over PCIe.

Tested-by: Herve Codina <herve.codina@bootlin.com>
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Reviewed-by: Simon Horman <horms@kernel.org>
---
 drivers/misc/lan966x_pci.dtso | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
index 7bb726550caf..5bb12dbc0843 100644
--- a/drivers/misc/lan966x_pci.dtso
+++ b/drivers/misc/lan966x_pci.dtso
@@ -141,8 +141,9 @@ switch: switch@e0000000 {
 
 					interrupt-parent = <&oic>;
 					interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
+						     <14 IRQ_TYPE_LEVEL_HIGH>,
 						     <9 IRQ_TYPE_LEVEL_HIGH>;
-					interrupt-names = "xtr", "ana";
+					interrupt-names = "xtr", "fdma", "ana";
 
 					resets = <&reset 0>;
 					reset-names = "switch";

-- 
2.34.1


^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 03/15] net: microchip: fdma: add PCIe ATU support
  2026-09-24 19:56 ` [PATCH net-next v8 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] fdma_pci_atu_region_map()
  (drivers/net/ethernet/microchip/fdma/fdma_pci.c) is exported.
- [Low] The ATU region claimed in fdma_alloc_coherent_and_map() is
  released only by fdma_free_coherent_and_unmap().
- [Low] If fdma_pci_atu_region_map() fails, fdma_alloc_coherent_and_map()
  calls fdma_free_coherent() (fdma_api.c:154) but leaves fdma->dcbs and…
- [Low] fdma_pci_atu_region_map() sets the iATU enable bit (CTRL2, offset
  0x4) with a bare writel() and never reads it back.
- [Low] fdma_pci_atu_configure_translation() and fdma_pci_atu_init()
  never write iATU REGION_CTRL1 (offset 0x0: TLP type, function number,…

--

net: microchip: fdma: add PCIe ATU support

This patch adds PCIe ATU support to the shared Microchip FDMA library. It
lets lan966x (and later lan969x) DMA to host memory when running as a PCIe
endpoint. The outbound address space is split into six regions. New
helpers combine coherent DMA allocation with claiming and programming an
ATU region, and the mapped allocation is padded to the 64KB ATU
granularity.

> diff --git a/drivers/net/ethernet/microchip/fdma/fdma_api.c b/drivers/net/ethernet/microchip/fdma/fdma_api.c
> index e78c3590da9e6..a3c9e3097c5c2 100644
> --- a/drivers/net/ethernet/microchip/fdma/fdma_api.c
> +++ b/drivers/net/ethernet/microchip/fdma/fdma_api.c
> @@ -127,6 +127,50 @@ void fdma_free_phys(struct fdma *fdma)
>  }
>  EXPORT_SYMBOL_GPL(fdma_free_phys);
>  
> +#if IS_ENABLED(CONFIG_MCHP_LAN966X_PCI)
> +/* Allocate coherent DMA memory and map it in the ATU. */
> +int fdma_alloc_coherent_and_map(struct device *dev, struct fdma *fdma,
> +				struct fdma_pci_atu *atu)
> +{
> +	struct fdma_pci_atu_region *region;
> +	int err;
> +
> +	if (WARN_ON(fdma->atu_region))
> +		return -EBUSY;
> +

[ ... ]

> +	err = fdma_alloc_coherent(dev, fdma);
> +	if (err)
> +		return err;
> +
> +	region = fdma_pci_atu_region_map(atu, fdma->dma, fdma->size);
> +	if (IS_ERR(region)) {
> +		fdma_free_coherent(dev, fdma);
> +		return PTR_ERR(region);
> +	}

[Severity: Low]
If fdma_pci_atu_region_map() fails here, the buffer is freed, but
fdma->dcbs and fdma->dma still point at it. A later
fdma_free_coherent_and_unmap() or fdma_free_coherent() on the same struct
fdma would then call dma_free_coherent() on this buffer a second time.
Could that happen?

No caller later in the series seems to do this:
lan966x_fdma_pci_rx_alloc()/tx_alloc() just return the error.
lan966x_fdma_pci_init() frees only the rx struct when the tx allocation
fails. lan966x_fdma_pci_reload() puts the old structs back with memcpy()
on failure.

Would clearing fdma->dcbs and fdma->dma after the rollback free still be
worth doing?

> +
> +	fdma->atu_region = region;
> +
> +	return 0;
> +}
> +EXPORT_SYMBOL_GPL(fdma_alloc_coherent_and_map);
> +
> +/* Free coherent DMA memory and unmap the memory in the ATU. */
> +void fdma_free_coherent_and_unmap(struct device *dev, struct fdma *fdma)
> +{
> +	fdma_pci_atu_region_unmap(fdma->atu_region);
> +	fdma->atu_region = NULL;
> +	fdma_free_coherent(dev, fdma);
> +}

[Severity: Low]
This is the only path that releases the region claimed in
fdma_alloc_coherent_and_map(). The plain fdma_free_coherent() is still
exported and does not touch fdma->atu_region.

What happens if a by-value copy of struct fdma is freed with
fdma_free_coherent(), as the platform lan966x_fdma_reload() does? The
region would stay in_use, with an enabled outbound window pointing at
freed memory. The next map could also hit WARN_ON(fdma->atu_region).

Later in the series the PCIe path looks correct. lan966x_fdma_pci_reload()
sets rx.fdma.atu_region and tx.fdma.atu_region to NULL before
reallocating, then frees the old copies with
fdma_free_coherent_and_unmap(). The platform lan966x_fdma_reload() only
runs through lan966x_fdma_ops, which never maps ATU regions.

Would it help to document that fdma_free_coherent() must not be used on
a mapped struct fdma?

[ ... ]

> diff --git a/drivers/net/ethernet/microchip/fdma/fdma_pci.c b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
> new file mode 100644
> index 0000000000000..dd1dc46cbc9dc
> --- /dev/null
> +++ b/drivers/net/ethernet/microchip/fdma/fdma_pci.c
> @@ -0,0 +1,208 @@

[ ... ]

> +/* Configure the address translation in the ATU. */
> +static void
> +fdma_pci_atu_configure_translation(struct fdma_pci_atu_region *region)
> +{

[ ... ]

> +	/* The OB address space lies entirely below 4GB, so the limit always
> +	 * fits the lower limit register and the upper one is left alone.
> +	 */
> +	writel(region->limit_addr,
> +	       fdma_pci_atu_addr_get(atu->addr, FDMA_PCI_ATU_LIMIT_ADDR, idx));

[Severity: Low]
This writes base, limit and target, but it never writes REGION_CTRL1
(offset 0x0: TLP type, function number, INCREASE_REGION_SIZE) or
UPPER_LIMIT (offset 0x20). fdma_pci_atu_init() doesn't write them either;
it only clears the enable bit.

Does this code depend on those registers still holding their reset values
(memory request type, function 0, no increased region size)? For
comparison, dw_pcie_prog_outbound_atu() always writes CTRL1 explicitly.

No other agent that programs these registers has been identified, and
this driver never writes them. Would setting CTRL1 and UPPER_LIMIT to 0
during setup be reasonable hardening?

[ ... ]

> +/* Map a host DMA address into a free outbound region. */
> +struct fdma_pci_atu_region *
> +fdma_pci_atu_region_map(struct fdma_pci_atu *atu, u64 target_addr, int size)
> +{

[ ... ]

> +	if (!IS_ALIGNED(target_addr, FDMA_PCI_ATU_REGION_ALIGN))
> +		return ERR_PTR(-EINVAL);

[ ... ]

> +	region->target_addr = target_addr;
> +	region->limit_addr = region->base_addr + size - 1;

[Severity: Low]
target_addr is checked for FDMA_PCI_ATU_REGION_ALIGN alignment, but size
is not, and size goes straight into limit_addr here.

The commit message says the hardware widens a limit that is not 64KB
aligned. fdma_pci_atu_region_map() is exported. If a caller passed a size
such as 0x1234, would the outbound window then extend to
target_addr + 0xffff, past the end of the caller's buffer?

This isn't reachable today. The fdma_pci.h comment says "target_addr and
size must be FDMA_PCI_ATU_REGION_ALIGN aligned". The only in-tree caller,
fdma_alloc_coherent_and_map(), pads fdma->size with ALIGN() first.

Could an IS_ALIGNED(size, FDMA_PCI_ATU_REGION_ALIGN) check that returns
-EINVAL go next to the target_addr check, so the documented precondition
is enforced?

> +	region->in_use = true;
> +
> +	fdma_pci_atu_configure_translation(region);
> +	fdma_pci_atu_region_enable(region);

[Severity: Low]
fdma_pci_atu_region_enable() sets the enable bit in CTRL2 (offset 0x4)
with a bare writel(), and success is returned without reading the bit
back. On the other side, fdma_pci_atu_region_unmap() clears the bit with
a posted write, and fdma_free_coherent_and_unmap() frees the buffer right
after.

dw_pcie_prog_outbound_atu() polls CTRL2 after enabling a region, with the
comment "Make sure ATU enable takes effect before any subsequent config
and I/O accesses". Does the lan966x iATU need a similar read-back?

In the series as posted, lan966x_fdma_rx_start() reads registers on the
same endpoint before it activates a channel, and those reads would flush
the enable write. Every unmap caller also first waits for the channels to
go idle via lan966x_fdma_rx_disable()/tx_disable(). So this only matters
if the hardware needs settle time beyond that flush.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset
  2026-09-24 19:56 ` [PATCH net-next v8 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The commit message and the new comment in lan966x_reset_switch()
  (drivers/net/ethernet/microchip/lan966x/lan966x_main.c) describe a…

--

net: lan966x: clear FDMA interrupt stickies after switch reset

This commit clears the FDMA_ERRORS, FDMA_INTR_ERR and FDMA_INTR_DB sticky
registers in lan966x_reset_switch() right after the switch reset. The aim
is to stop spurious bits that the GCB soft reset latches in PCI mode from
triggering the FDMA IRQ handler, and its WARN, during probe.

> When in PCI mode, the GCB soft reset issued by the reset controller
> can latch spurious bits in the FDMA error stickies. The latched bits
> sit in FDMA_INTR_ERR until the FDMA IRQ is requested later in probe,
> at which point the handler fires immediately and WARNs.

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index 6e6c08bb8eeab..259d81e759076 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -1067,6 +1067,15 @@ static int lan966x_reset_switch(struct lan966x *lan966x)
>  
>  	reset_control_reset(switch_reset);
>  
> +	/* When in PCI mode, the GCB soft reset issued by the reset
> +	 * controller can latch spurious bits in the FDMA error and
> +	 * data-block stickies. Clear them before request_irq hooks up the
> +	 * FDMA IRQ line, otherwise the handler fires immediately on probe.
> +	 */

[Severity: Low]
This isn't a bug, but can the WARN described in the commit message and in
this comment actually be hit in PCI mode at this commit?

The PCI overlay describes only two interrupts for the switch node:

drivers/misc/lan966x_pci.dtso:
	interrupt-names = "xtr", "ana";

lan966x_probe() requests the FDMA IRQ only when an "fdma" interrupt is
described:

	lan966x->fdma_irq = platform_get_irq_byname(pdev, "fdma");
	if (lan966x->fdma_irq > 0) {
		err = devm_request_irq(&pdev->dev, lan966x->fdma_irq,
				       lan966x_fdma_irq_handler, 0,
	...

So in PCI mode lan966x_fdma_irq_handler() is never registered, and its
WARN() cannot fire. That changes only with later patches in the series,
"net: lan966x: add PCIe FDMA support" and "misc: lan966x-pci: dts: add fdma
interrupt to overlay", which add "fdma" to the overlay.

The series ordering itself looks fine. Would it be clearer to say that the
clear is done in preparation for enabling the FDMA IRQ in PCI mode? As
written, it reads as a fix for a WARN that can be hit today.

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot
  2026-09-24 19:56 ` [PATCH net-next v8 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.

Critical: 0 · High: 2 · Medium: 0 · Low: 2

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] Baseline decision: false positive; not reachable at this commit,
  and closed later in the series.
- [High] Baseline decision: false positive; not reachable at this commit,
  and closed later in the series.
- [Low] The part of lan966x_shutdown() that actually runs on PCIe at this
  commit fixes behaviour that has existed since PCIe support for lan966x…
- [Low] Baseline decision: false positive.

--

net: lan966x: add shutdown callback to stop the FDMA on reboot

This adds a .shutdown callback to the lan966x switch driver. The callback
only acts on the PCIe instantiation, where it frees the xtr, ana and FDMA
irqs, masks the analyzer interrupt, detaches the netdevs, disables both
FDMA channels and unmaps the outbound ATU windows.

It also restores the full FDMA_INTR_ENA enable in lan966x_fdma_rx_start(),
because that register persists across a warm reboot.

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index 259d81e759076..024ce9f9916c7 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -1324,9 +1324,65 @@ static void lan966x_remove(struct platform_device *pdev)
>  	debugfs_remove_recursive(lan966x->debugfs_root);
>  }
>
> +static void lan966x_shutdown(struct platform_device *pdev)
> +{
> +	struct lan966x *lan966x = platform_get_drvdata(pdev);
> +

[ ... ]

> +	if (!lan966x_is_pci(lan966x))
> +		return;
> +
> +	if (lan966x->xtr_irq > 0)
> +		devm_free_irq(lan966x->dev, lan966x->xtr_irq, lan966x);
> +	if (lan966x->ana_irq > 0)
> +		devm_free_irq(lan966x->dev, lan966x->ana_irq, lan966x);
> +	if (lan966x->fdma_irq > 0)
> +		devm_free_irq(lan966x->dev, lan966x->fdma_irq, lan966x);
> +
> +	lan_wr(0, lan966x, ANA_ANAINTR);

[Severity: Low]
The code here looks correct, but should this change carry a Fixes: tag?

At this point in the series, the only part of lan966x_shutdown() that runs
on PCIe is the part that frees the xtr and ana irqs and clears ANA_ANAINTR.

The xtr and ana interrupts were wired up by commit 185686beb464 ("misc: Add
support for LAN966x PCI device"), in drivers/misc/lan966x_pci.dtso:

    interrupt-names = "xtr", "ana";

Since that commit, a warm host reboot leaves these OIC sources enabled on an
endpoint that is not reset. The same is true of the analyzer interrupt
armed by lan966x_init(). This matches what the commit message describes:

  "As a PCIe endpoint, lan966x is not reset by a host reboot: its FDMA
  channels and interrupt sources stay armed, and the OIC ORs every
  source into the shared PCIe INTx, asserted before the driver has
  re-probed."

Could the xtr/ana part be split out as a standalone fix with:

  Fixes: 185686beb464 ("misc: Add support for LAN966x PCI device")

As written, it depends on lan966x_is_pci() and lan966x->dma_dev, which were
added earlier in this series. That means it can't be identified for stable
or backported on its own.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 11/15] net: lan966x: add PCIe FDMA support
  2026-09-24 19:56 ` [PATCH net-next v8 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found 11 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 6 (1) · Low: 5

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] lan966x_fdma_pci_napi_poll() limits RX processing only by the
  caller's budget and decides which descriptors to re-arm with `while…
- [Medium] Port netdevs are allocated with NUM_PRIO_QUEUES TX queues
  (devm_alloc_etherdev_mqs(..., NUM_PRIO_QUEUES, 1)).
- [Medium] lan966x_fdma_pci_init() sets tx_fdma->n_dcbs = FDMA_DCB_MAX
  (FDMA_PCI_DCB_MAX at the series end) before allocating the ring.
- [Medium] lan966x->fdma is set to true, and the port netdevs are
  registered, before lan966x->ops->fdma_init() runs.
- [Medium] At this commit, lan966x_fdma_pci_rx_check_frame() reads and
  validates BLOCKL, but lan966x_fdma_pci_rx_get_frame() re-reads it from…
- [Medium] lan966x_fdma_pci_get_next_dcb() treats a TX DCB as free when
  DB0 has DONE set and the DCB is not fdma->last_dcb, and it always…
- [Low] At this commit the in-tree PCIe overlay
  (drivers/misc/lan966x_pci.dtso) declares only the "xtr" and "ana"…
- [Low] The patch adds TARGET_PCIE_DBI at offset 0x400000 in IO range 0
  ("cpu").
- [Low] At this commit lan966x_fdma_pci_resize() always returns
  -EOPNOTSUPP, so lan966x_port_change_mtu() rolls back every MTU change…
- [Low] In lan966x_fdma_pci_napi_poll(), the FDMA_ERROR case says "No
  rx_dropped increment here because src_port is invalid".
- [Low] lan966x_main.h documents FDMA_ERROR as "something went wrong,
  stop getting more frames", and the platform lan966x_fdma_napi_poll()…

Pre-existing issues:
- [Medium] __fdma_dcb_add() writes the DB status and dataptr, links the
  DCB into the live chain (fdma->last_dcb->nextptr = new), and only then…

--

net: lan966x: add PCIe FDMA support

This adds a PCIe FDMA backend for lan966x. The backend uses contiguous
coherent DMA buffers mapped through the endpoint ATU and copies frames with
memcpy. The patch also maps the PCIe DBI register target, disables PTP on
the PCIe path and rejects XDP on PCIe instances.

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> new file mode 100644
> index 0000000000000..bccd1b8590d74
> --- /dev/null
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> @@ -0,0 +1,421 @@

[ ... ]

> +static int lan966x_fdma_pci_get_next_dcb(struct fdma *fdma)
> +{
> +	struct fdma_db *db;
> +
> +	for (int i = 0; i < fdma->n_dcbs; i++) {
> +		db = fdma_db_get(fdma, i, 0);
> +
> +		if (!fdma_db_is_done(db))
> +			continue;
> +		if (fdma_is_last(fdma, &fdma->dcbs[i]))
> +			continue;
> +
> +		return i;
> +	}
> +
> +	return -ENOSPC;
> +}

[Severity: Medium]
Can this hand out the DCB that the TX engine is still parked on?

A DCB counts as free here once DB0 has DONE set and it is not
fdma->last_dcb. The scan always starts at index 0.

Suppose the engine is parked on a completed DCB0 whose NEXTPTR is INVALID:

xmit B
  lan966x_fdma_pci_get_next_dcb() returns 1
  __fdma_dcb_add(1): DCB0->nextptr = DCB1, last_dcb = DCB1
  lan966x_fdma_tx_start() writes CH_RELOAD

xmit C (right after, serialized only by tx_lock)
  lan966x_fdma_pci_get_next_dcb() returns 0 (DONE, not last)
  __fdma_dcb_add(0): clears DCB0 DONE, DCB1->nextptr = DCB0,
                     DCB0->nextptr = INVALID

What if the engine re-reads DCB0.NEXTPTR after the posted RELOAD write
and after C's stores? Would it then see INVALID and stay parked?

In that case frames B and C would never be sent, and DCB0 and DCB1 would
never see DONE again. The next xmit would link DCB0 to DCB2 and strand
them. If this repeats, the ring could drain until xmit returns
NETDEV_TX_BUSY for good.

This depends on the FDMA RELOAD semantics, which I could not confirm. The
platform path uses a used flag for similar selection. The same logic
appears in lan966x_fdma_pci_xmit_xdpf() later in the series.

[ ... ]

> +static int lan966x_fdma_pci_rx_check_frame(struct lan966x_rx *rx, u64 *src_port)
> +{

[ ... ]

> +	blockl = fdma_db_len_get(db);
> +	if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
> +		return FDMA_ERROR;
> +
> +	return FDMA_PASS;
> +}
> +
> +static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
> +						     u64 src_port)
> +{

[ ... ]

> +	/* Get the received frame and create an SKB for it. */
> +	db = fdma_db_next_get(fdma);
> +	data_len = fdma_db_len_get(db);
> +
> +	skb = napi_alloc_skb(&lan966x->napi, data_len);

[Severity: Medium]
Is it safe to read BLOCKL a second time here?

lan966x_fdma_pci_rx_check_frame() validates BLOCKL but then drops the
value. This function reads it again from coherent memory that the device
can write. The second, unchecked value then drives napi_alloc_skb(),
memcpy(), skb_put(), skb_pull() and skb_trim().

If the endpoint rewrites BLOCKL between the two loads, could this read
past the buffer or underflow skb->len?

The later patch "net: lan966x: add PCIe FDMA XDP support" appears to fix
this. It computes data_len once in rx_check_frame() and passes it into
rx_get_frame(). The path is also not reachable at this commit, because the
overlay has no "fdma" irq.

[ ... ]

> +static int lan966x_fdma_pci_xmit(struct sk_buff *skb, __be32 *ifh,
> +				 struct net_device *dev)
> +{

[ ... ]

> +	next_to_use = lan966x_fdma_pci_get_next_dcb(fdma);
> +
> +	if (next_to_use < 0) {
> +		netif_stop_queue(dev);
> +		return NETDEV_TX_BUSY;
> +	}

[Severity: Medium]
Does netif_stop_queue() stop the queue this skb came from?

lan966x_probe_port() creates port netdevs with NUM_PRIO_QUEUES TX queues:

    dev = devm_alloc_etherdev_mqs(lan966x->dev,
                                  sizeof(struct lan966x_port),
                                  NUM_PRIO_QUEUES, 1);

There is no ndo_select_queue, so traffic is hashed over all 8 queues.
netif_stop_queue(dev) stops only txq 0. lan966x_fdma_wakeup_netdev(),
called from the new poll, also checks and wakes only txq 0.

For an skb on txq 1-7, the qdisc requeues it, calls __netif_schedule() and
retries at once, because that queue was never stopped. Could
net_tx_action then spin in softirq, contending on tx_lock, until a DCB
frees up?

The same pattern shows up in lan966x_fdma_pci_xmit_xdpf() later in the
series.

[ ... ]

> +	/* Order frame write before DCB status write below. */
> +	dma_wmb();
> +
> +	fdma_dcb_add(fdma,
> +		     next_to_use,
> +		     0,
> +		     FDMA_DCB_STATUS_INTR |
> +		     FDMA_DCB_STATUS_SOF |
> +		     FDMA_DCB_STATUS_EOF |
> +		     FDMA_DCB_STATUS_BLOCKO(0) |
> +		     FDMA_DCB_STATUS_BLOCKL(IFH_LEN_BYTES + skb->len + ETH_FCS_LEN));
> +
> +	/* Start the transmission. */
> +	lan966x_fdma_tx_start(tx);

[Severity: Medium]
This is a pre-existing issue in __fdma_dcb_add() and was not introduced
by this patch, but the new PCIe TX and RX re-arm paths depend on it.

__fdma_dcb_add() links the new DCB into the live chain before it has
finished initializing it, with no barrier in between:

drivers/net/ethernet/microchip/fdma/fdma_api.c:__fdma_dcb_add() {
    ...
	fdma->last_dcb->nextptr = cpu_to_le64(nextptr);
	fdma->last_dcb = dcb;
	dcb->nextptr = cpu_to_le64(FDMA_DCB_INVALID_DATA);
	dcb->info = cpu_to_le64(info);
    ...
}

The dma_wmb() here orders only the frame data against the descriptor
stores. Both callers then issue a CH_RELOAD writel(), which orders the
earlier stores. That leaves a problem only if the engine follows a
freshly written nextptr without a RELOAD.

Can this FDMA do that? If it can, could the device fetch a DCB whose
nextptr or info is still stale?

[ ... ]

> +static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
> +{

[ ... ]

> +		switch (lan966x_fdma_pci_rx_check_frame(rx, &src_port)) {
> +		case FDMA_PASS:
> +			break;
> +		case FDMA_ERROR:
> +			/* No rx_dropped increment here because src_port is
> +			 * invalid.
> +			 */
> +			fdma_dcb_advance(fdma);
> +			continue;
> +		}

[Severity: Low]
Is this comment accurate? lan966x_fdma_pci_rx_check_frame() also returns
FDMA_ERROR when BLOCKL fails lan966x_fdma_pci_rx_size_fits(). By that
point it has already confirmed that src_port is in range and
ports[src_port] is non-NULL:

	blockl = fdma_db_len_get(db);
	if (!lan966x_fdma_pci_rx_size_fits(fdma, blockl))
		return FDMA_ERROR;

So a frame dropped for its length on a valid port is counted in neither
rx_dropped nor rx_length_errors. This is still the case at the end of the
series.

[Severity: Low]
Should FDMA_ERROR stop the loop here? lan966x_main.h documents it as:

 * FDMA_ERROR, something went wrong, stop getting more frames

and lan966x_fdma_napi_poll() on the platform path follows that:

	case FDMA_ERROR:
		lan966x_fdma_rx_free_page(rx);
		fdma_dcb_advance(fdma);
		goto allocate_new;

The PCIe poll instead advances and keeps pulling frames up to the budget.
The two backends therefore handle the same action code differently. The
code is unchanged at the end of the series.

> +		skb = lan966x_fdma_pci_rx_get_frame(rx, src_port);
> +		fdma_dcb_advance(fdma);
> +		if (!skb) {
> +			lan966x->ports[src_port]->dev->stats.rx_dropped++;
> +			continue;
> +		}
> +
> +		napi_gro_receive(&lan966x->napi, skb);
> +	}
> +	while (dcb_reload != fdma->dcb_index) {
> +		old_dcb = dcb_reload;
> +		dcb_reload++;
> +		dcb_reload &= fdma->n_dcbs - 1;

[Severity: Medium]
What happens here when the budget is at least n_dcbs?

The RX loop above is bounded only by weight, not by the ring size. If
every descriptor is DONE and weight >= n_dcbs, fdma_dcb_advance() wraps
dcb_index back to where it started. This loop then runs zero times, no DCB
goes back to the hardware and the DONE bits stay set.

Won't the next poll then deliver the same stale frames again?

With budgets above n_dcbs, descriptors are also revisited within a single
call, and only the modulo remainder is re-armed.

__napi_busy_loop() passes the busy-poll budget straight to napi->poll().
SO_BUSY_POLL_BUDGET accepts up to U16_MAX with CAP_NET_ADMIN. n_dcbs is
512 here and 256 (FDMA_PCI_DCB_MAX) at the end of the series. The platform
poll has the same pattern.

[ ... ]

> +static int lan966x_fdma_pci_init(struct lan966x *lan966x)
> +{
> +	struct fdma *rx_fdma = &lan966x->rx.fdma;
> +	struct fdma *tx_fdma = &lan966x->tx.fdma;
> +	int err;
> +
> +	if (!lan966x->fdma)
> +		return 0;

[Severity: Low]
Can any of the new PCIe FDMA code run at this commit?

lan966x_probe() sets lan966x->fdma only when an "fdma" irq exists. The
in-tree overlay drivers/misc/lan966x_pci.dtso declares only:

	interrupt-names = "xtr", "ana";

so on PCIe this function returns immediately. The "~620 Mbps" figure in
the commit message can't be reproduced from this commit alone.

The later patch "misc: lan966x-pci: dts: add fdma interrupt to overlay"
adds the interrupt, so this is resolved once the series is applied.

[ ... ]

> +	lan966x->tx.lan966x = lan966x;
> +	tx_fdma->channel_id = FDMA_INJ_CHANNEL;
> +	tx_fdma->n_dcbs = FDMA_DCB_MAX;

[Severity: Medium]
Can a transmit race with this initialization?

When lan966x_probe() calls lan966x->ops->fdma_init(), lan966x->fdma is
already true and the port netdevs are already registered.

This function publishes tx_fdma->n_dcbs, db_size and the ops callbacks
without holding lan966x->tx_lock. It then sleeps in dma_alloc_coherent(),
via fdma_alloc_coherent_and_map(). Only after that does it set
tx.fdma.dcbs and atu_region.

A transmit on a port that is up with carrier during that window would
take this path:

lan966x_port_xmit()
  spin_lock(&lan966x->tx_lock)
  lan966x->ops->fdma_xmit()
    lan966x_fdma_pci_xmit()
      lan966x_fdma_pci_get_next_dcb()
        fdma_db_get(fdma, i, 0)   <- dcbs is still NULL

It could also pick a DCB whose DONE bit fdma_dcbs_init() has just set.
fdma_dcb_add() would then run with a NULL atu_region in
lan966x_fdma_pci_dataptr_cb().

Before this point n_dcbs is 0 from kzalloc, so the window opens exactly at
this assignment. The platform backend has the same probe ordering.

[ ... ]

> +	err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
> +	if (err)
> +		return err;
> +
> +	err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
> +	if (err) {
> +		fdma_free_coherent_and_unmap(lan966x->dma_dev, rx_fdma);
> +		return err;
> +	}

[Severity: Medium]
Does this error path leave tx.fdma and rx.fdma in a state that xmit can
still use?

tx_fdma->n_dcbs stays set. tx.fdma.dcbs is then either NULL, if
dma_alloc_coherent() failed, or points at freed coherent memory, because
fdma_alloc_coherent_and_map() frees the buffer without clearing the
pointer:

	region = fdma_pci_atu_region_map(atu, fdma->dma, fdma->size);
	if (IS_ERR(region)) {
		fdma_free_coherent(dev, fdma);
		return PTR_ERR(region);
	}

In the same way, rx.fdma.dcbs is left dangling after
fdma_free_coherent_and_unmap(rx_fdma).

Probe then unwinds through cleanup_ptp, cleanup_fdb and cleanup_ports.
lan966x->fdma stays true until lan966x_cleanup_ports() unregisters the
netdevs. Could a transmit on an up port in that window reach
lan966x_fdma_pci_get_next_dcb() and dereference NULL or read freed memory?

If a stale DONE bit is seen, the memcpy() in lan966x_fdma_pci_xmit() would
write into the freed buffer. fdma_dcb_add() would then use the NULL
atu_region.

[ ... ]

> +static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
> +{
> +	return -EOPNOTSUPP;
> +}

[Severity: Low]
With this stub, won't lan966x_port_change_mtu() roll back every MTU
change on a PCIe FDMA instance?

	err = lan966x->ops->fdma_resize(lan966x);
	if (err) {
		...
		dev->mtu = old_mtu;
	}

The later patch "net: lan966x: add PCIe FDMA MTU change support" replaces
this stub with a real resize. The PCIe FDMA path also can't be reached at
this commit, because the overlay has no "fdma" irq.

[ ... ]

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> index 024ce9f9916c7..de22027868261 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.c
> @@ -50,6 +50,7 @@ struct lan966x_main_io_resource {
>  static const struct lan966x_main_io_resource lan966x_main_iomap[] =  {
>  	{ TARGET_CPU,                   0xc0000, 0 }, /* 0xe00c0000 */
>  	{ TARGET_FDMA,                  0xc0400, 0 }, /* 0xe00c0400 */
> +	{ TARGET_PCIE_DBI,             0x400000, 0 }, /* 0xe0400000 */

[Severity: Low]
Is range 0 large enough for this offset at this commit?

The in-tree overlay sizes the "cpu" resource at 1 MB:

	reg = <0xe0000000 0x0100000>, ...

lan966x_create_targets() ioremaps only resource_size() and does not check
the offsets against it. lan966x_fdma_pci_init() passes
regs[TARGET_PCIE_DBI] to fdma_pci_atu_init(), which writes at
+0x300000 + 0x200 * idx, so those writes would fall outside the mapping.

This code runs only when lan966x->fdma is true. The later patch
"misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space" grows
the reg to 0x800000 before the "fdma" irq is added, so the out-of-range
write can't happen at any point in the series.

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 12/15] net: lan966x: add PCIe FDMA MTU change support
  2026-09-24 19:57 ` [PATCH net-next v8 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found 3 potential
issue(s) to consider.

Critical: 0 · High: 2 (1) · Medium: 0 · Low: 1 (1)

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] lan966x_fdma_pci_resize() skips the resize only while
  `lan966x->rx.lan966x` is NULL.
- [High] lan966x_fdma_pci_resize() reads the plain field
  `lan966x->rx.lan966x` without any synchronization with…
- [Low] The comment says FDMA_PCI_DB_SIZE_MAX (lan966x_main.h:106) is the
  'Largest db_size keeping the ATU-padded ring inside one MAX_PAGE_ORDER…

Pre-existing issues:
- [High] This problem wasn't introduced by this patch, but the patch
  works around it only on the PCIe path.
- [Low] lan966x_fdma_pci_reload() frees the old RX/TX coherent rings and
  unmaps their ATU regions (lines 465-466) after calling…

--

net: lan966x: add PCIe FDMA MTU change support

This commit adds MTU change support to the lan966x PCIe FDMA path. On an
MTU change it reallocates the contiguous ATU-mapped RX and TX rings at the
new size, and falls back to the old rings if that fails. It also caps the
PCIe DCB ring at 256 entries, advertises FDMA_PCI_MAX_MTU in dev->max_mtu
when the FDMA is in use, and skips the resize until
lan966x_fdma_pci_init() has run.

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
> index 2e8f786d6feec..a7940eca5df31 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma.c
> @@ -890,11 +890,7 @@ static int lan966x_fdma_reload(struct lan966x *lan966x, int new_mtu)
>  
>  int lan966x_fdma_get_max_frame(struct lan966x *lan966x)
>  {
> -	return lan966x_fdma_get_max_mtu(lan966x) +
> -	       IFH_LEN_BYTES +
> -	       SKB_DATA_ALIGN(sizeof(struct skb_shared_info)) +
> -	       VLAN_HLEN * 2 +
> -	       XDP_PACKET_HEADROOM;
> +	return lan966x_fdma_get_max_mtu(lan966x) + FDMA_OVERHEAD;
>  }

[Severity: High]
This isn't a bug introduced by this patch, but the commit message notes
that the FDMA init "runs after the netdevs register". The platform FDMA
has the same ordering. Does the platform path need the same guard?

lan966x_probe_port() calls register_netdev(). lan966x_fdma_init() runs
only later, from lan966x_probe() via lan966x->ops->fdma_init(). The
platform resize callback has no check for that window:

lan966x_fdma_change_mtu() {
	max_mtu = lan966x_fdma_get_max_frame(lan966x);
	if (max_mtu == lan966x->rx.max_mtu)
		return 0;

	return __lan966x_fdma_reload(lan966x, max_mtu);
}

Before init, rx.max_mtu is still 0, so an MTU change in that window goes
through:

__lan966x_fdma_reload()->lan966x_fdma_reload()->lan966x_fdma_rx_disable()

lan966x_fdma_rx_disable() uses rx->lan966x, which is still NULL at that
point. If NAPI has not been added yet, napi_disable() on the
uninitialized napi_struct could crash first.

A udev .link MTUBytes= setting or NetworkManager can apply an MTU at
netdev registration, which can land in this window.

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> index bccd1b8590d74..7185e65dda435 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c

[ ... ]

> @@ -391,9 +391,156 @@ static int lan966x_fdma_pci_init(struct lan966x *lan966x)

[ ... ]

> +static int lan966x_fdma_pci_reload(struct lan966x *lan966x, int new_mtu)
> +{
> +	struct fdma tx_fdma_old = lan966x->tx.fdma;
> +	struct fdma rx_fdma_old = lan966x->rx.fdma;
> +	u32 old_mtu = lan966x->rx.max_mtu;
> +	int err;
> +
> +	napi_disable(&lan966x->napi);
> +	lan966x_fdma_tx_disable_netdev(lan966x);
> +	lan966x_fdma_rx_disable(&lan966x->rx);
> +	lan966x_fdma_tx_disable(&lan966x->tx);

[ ... ]

> +	err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
> +	if (err) {
> +		fdma_free_coherent_and_unmap(lan966x->dma_dev,
> +					     &lan966x->rx.fdma);
> +		goto restore;
> +	}
> +
> +	/* Free and unmap old memory. */
> +	fdma_free_coherent_and_unmap(lan966x->dma_dev, &rx_fdma_old);
> +	fdma_free_coherent_and_unmap(lan966x->dma_dev, &tx_fdma_old);

[Severity: Low]
This is a pre-existing issue, but lan966x_fdma_rx_disable() and
lan966x_fdma_tx_disable() both ignore the return value of their poll on
FDMA_CH_ACTIVE:

	readx_poll_timeout_atomic(lan966x_fdma_channel_active, lan966x,
				  val, !(val & BIT(fdma->channel_id)),
				  READL_SLEEP_US, READL_TIMEOUT_US);

If a channel never goes idle, the old rings are freed here and their ATU
regions unmapped while the FDMA could still be using them. Should the
timeout be propagated, and the free skipped when it fires?

The timeout is READL_TIMEOUT_US (100 seconds), so this only matters for a
wedged engine. The platform lan966x_fdma_reload() and both deinit paths
already free rings after the same unchecked poll, so this patch just adds
another caller.

[ ... ]

>  static int lan966x_fdma_pci_resize(struct lan966x *lan966x)
>  {
> -	return -EOPNOTSUPP;
> +	int max_mtu;
> +
> +	/* Nothing to resize until fdma_pci_init() has built the rings; it
> +	 * sizes them from DEV_MAC_MAXLEN_CFG, which the caller already set.
> +	 */
> +	if (!lan966x->rx.lan966x)
> +		return 0;

[Severity: High]
Does a non-NULL rx.lan966x really mean the rings have been built?
lan966x_fdma_pci_init() sets it before either ring is allocated, and it
is not cleared when init fails:

lan966x_fdma_pci_init() {
	...
	lan966x->rx.lan966x = lan966x;
	...
	err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
	if (err)
		return err;

	err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);
	if (err) {
		fdma_free_coherent_and_unmap(lan966x->dma_dev, rx_fdma);
		return err;
	}
	...
}

lan966x_fdma_pci_deinit() disables NAPI and frees both rings, but it
does not clear rx.lan966x or rx.max_mtu either.

The netdevs are still registered in both cases. lan966x_remove() calls
fdma_deinit before lan966x_cleanup_ports(), and it does so without RTNL.
When fdma_init fails, lan966x_probe() jumps to cleanup_ptp and reaches
lan966x_cleanup_ports() only later.

An MTU change in either window passes this check. Because rx.max_mtu is
stale, it then goes into lan966x_fdma_pci_reload().

After deinit, can the unconditional napi_disable() in
lan966x_fdma_pci_reload() spin forever in napi_disable_locked() while
holding RTNL? The first disable left SCHED|NPSVC set. lan966x_remove()
would then block in unregister_netdev() waiting for RTNL. If deinit and
the reload run at the same time, both would disable NAPI and free the
same coherent buffers and ATU regions.

After a tx_alloc failure in init, rx.fdma.dcbs still points at the
freed buffer. The reload copies it into rx_fdma_old and, if the new
allocation succeeds, frees it a second time with
fdma_free_coherent_and_unmap(&rx_fdma_old). The new rings are then left
running with nothing to free them, because the probe error path never
calls fdma_deinit.

If the reload's own allocation fails instead,
lan966x_fdma_pci_reset_mem() memsets the freed buffer and dereferences
the NULL atu_region.

Before this patch the PCIe resize returned -EOPNOTSUPP, so these paths
could not be reached. The later XDP patch in the series does not change
this check, init or deinit either.

Would this need a real ring-state flag, with teardown serialized against
MTU changes?

[Severity: High]
Can this check also race with lan966x_fdma_pci_init() itself?

lan966x_probe() calls fdma_init after register_netdev(), and without
RTNL. lan966x_fdma_pci_init() sets rx.lan966x first and tx.lan966x
later. It then allocates both rings with calls that can sleep
(dma_alloc_coherent(), the ATU mutex).

An ndo_change_mtu under RTNL can pass this check as soon as rx.lan966x
is set:

CPU1 (probe)                        CPU2 (RTNL)
lan966x_fdma_pci_init()
  lan966x->rx.lan966x = lan966x;
                                    lan966x_port_change_mtu()
                                      lan966x_fdma_pci_resize()
                                        __lan966x_fdma_pci_reload()
  lan966x_fdma_pci_rx_alloc()             lan966x_fdma_pci_reload()
                                            copies half-built rx/tx fdma
                                            allocates new rings
                                            frees "old" dcbs/ATU region

Both contexts then write the same rx/tx struct fdma. Init can overwrite
atu_region and dcbs. It can also hit WARN_ON(fdma->atu_region) in
fdma_alloc_coherent_and_map() and free rx_fdma, which by then may be the
buffer the reload just started the FDMA on.

Could this end in DMA into freed coherent memory, double frees, or
leaked ATU regions?

If the probe thread is preempted after setting rx.lan966x but before
setting tx.lan966x, would lan966x_fdma_tx_disable() and
lan966x_fdma_pci_tx_alloc() in the reload dereference a NULL
tx->lan966x?

> +
> +	max_mtu = lan966x_fdma_get_max_frame(lan966x);
> +	if (max_mtu == lan966x->rx.max_mtu)
> +		return 0;
> +
> +	return __lan966x_fdma_pci_reload(lan966x, max_mtu);
>  }

[ ... ]

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
> index 16bc28c8f11f8..1877f1916d716 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_main.h
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_main.h

[ ... ]

> @@ -87,6 +88,33 @@

[ ... ]

> +/* Largest db_size keeping the ATU-padded ring inside one MAX_PAGE_ORDER
> + * block and within the 16-bit DCB DATAL field. Inverts ALIGN(x, R) <= L
> + * into x <= ALIGN_DOWN(L, R) to bound x directly.
> + */
> +#define FDMA_PCI_DB_SIZE_MAX						\
> +	MIN_T(u32,							\
> +	      (ALIGN_DOWN(PAGE_SIZE << MAX_PAGE_ORDER,			\
> +			  FDMA_PCI_ATU_REGION_ALIGN) -			\
> +	       FDMA_PCI_DCB_MAX * sizeof(struct fdma_dcb)) /		\
> +	      (FDMA_PCI_DCB_MAX * FDMA_RX_DCB_MAX_DBS),			\
> +	      ALIGN_DOWN(GENMASK(15, 0), FDMA_PCI_DB_ALIGN))
> +
> +#define FDMA_PCI_MAX_MTU						\
> +	(FDMA_PCI_DB_SIZE_MAX - FDMA_OVERHEAD -				\
> +	 (ETH_HLEN + ETH_FCS_LEN))

[Severity: Low]
Can this arithmetic wrap on builds with a small MAX_PAGE_ORDER?

sizeof(struct fdma_dcb) is 256, so the subtracted term is 65536. If
PAGE_SIZE << MAX_PAGE_ORDER is below 64K, the unsigned subtraction
wraps. MIN_T(u32) truncates the result to 0xFFFFFF00, so the 16-bit
bound (65408) is picked, even though the ring would then be far larger
than one MAX_PAGE_ORDER block.

At 64K or 128K, FDMA_PCI_DB_SIZE_MAX is 0 or 256. That is below
FDMA_OVERHEAD, so FDMA_PCI_MAX_MTU wraps to roughly 4G, and
lan966x_probe_port() writes that value into dev->max_mtu. For blocks
smaller than about 1MB, the result is below the default MTU of 1500.

The value is correct for the tested configuration (4K pages,
MAX_PAGE_ORDER=10). arm32 ARCH_FORCE_MAX_ORDER has no range limit,
though, so the smaller configurations can be built.

Would a BUILD_BUG_ON() on this formula be worth adding?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 13/15] net: lan966x: add PCIe FDMA XDP support
  2026-09-24 19:57 ` [PATCH net-next v8 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found 3 potential
issue(s) to consider.

Critical: 1 · High: 0 · Medium: 2 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Critical] Use-after-free and double free on the PCIe RX path when an
  XDP program shrinks a frame.
- [Medium] On the PCIe path the driver accepts XDP but cannot handle
  XDP_REDIRECT, and the bond driver cannot tell.
- [Medium] In lan966x_fdma_pci_napi_poll(), a poll that consumes all
  FDMA_PCI_DCB_MAX (256) RX descriptors never re-arms any of them.

--

net: lan966x: add PCIe FDMA XDP support

This commit adds basic XDP support (XDP_PASS, XDP_DROP, XDP_TX) to the
lan966x PCIe FDMA path. The IFH and FCS are stripped before the program
runs, and the pointer and length from after XDP are passed on to skb
construction or to a new memcpy-based XDP_TX routine.
lan966x_xdp_setup() now accepts programs on PCIe and skips the page_pool
reload there.

> diff --git a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> index 7185e65dda435..216e9cbcd1580 100644
> --- a/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c
> +++ b/drivers/net/ethernet/microchip/lan966x/lan966x_fdma_pci.c

[ ... ]

> +static int lan966x_xdp_pci_run(struct lan966x_port *port, void *data,
> +			       u32 data_len, void **xdp_data, u32 *xdp_len)
> +{

[ ... ]

> +	act = bpf_prog_run_xdp(xdp_prog, &xdp);
> +
> +	*xdp_data = xdp.data;
> +	*xdp_len = xdp.data_end - xdp.data;
> +
> +	switch (act) {
> +	case XDP_PASS:
> +		return FDMA_PASS;
> +	case XDP_TX:
> +		return lan966x_fdma_pci_xmit_xdpf(port, *xdp_data, *xdp_len) ?
> +		       FDMA_DROP : FDMA_TX;
> +	default:
> +		bpf_warn_invalid_xdp_action(port->dev, xdp_prog, act);
> +		fallthrough;

[Severity: Medium]
What happens here when the PCIe port is a bond slave?

The PCIe ports use the same lan966x_port_netdev_ops as the platform
ports, so both .ndo_bpf and .ndo_xdp_xmit are set. bond_xdp_set() looks
only at those pointers and never reads xdp_features:

drivers/net/bonding/bond_main.c:bond_xdp_set() {
    ...
		if (!slave_dev->netdev_ops->ndo_bpf ||
		    !slave_dev->netdev_ops->ndo_xdp_xmit) {
    ...
}

This patch removes the -EOPNOTSUPP return for PCIe in
lan966x_xdp_setup(), so a bond with a lan966x PCIe slave can now install
its program on that slave. For bond slaves, bpf_prog_run_xdp() can then
turn XDP_TX into XDP_REDIRECT:

net/core/filter.c:xdp_master_redirect() {
    ...
	if (slave && slave != xdp->rxq->dev) {
		ri->tgt_index = slave->ifindex;
		ri->map_id = INT_MAX;
		ri->map_type = BPF_MAP_TYPE_UNSPEC;
		return XDP_REDIRECT;
    ...
}

It can also return XDP_ABORTED when the master is down.

This switch has no XDP_REDIRECT case. Would those frames go to the
default branch, log a "Driver unsupported" warning and be dropped?
Nothing calls xdp_do_redirect(), so the redirect info set by
xdp_master_redirect() would never be used.

[ ... ]

>  static struct sk_buff *lan966x_fdma_pci_rx_get_frame(struct lan966x_rx *rx,
> -						     u64 src_port)
> +						     u64 src_port, void *data,
> +						     u32 data_len)
>  {

[ ... ]

> -	memcpy(skb->data,
> -	       fdma_dataptr_virt_addr_contiguous(fdma,
> -						 fdma->dcb_index,
> -						 fdma->db_index),
> -						 data_len);
> +	memcpy(skb->data, data, data_len);
>  
>  	skb_put(skb, data_len);
>  
>  	skb->dev = lan966x->ports[src_port]->dev;
> -	skb_pull(skb, IFH_LEN_BYTES);
> -
> -	skb_trim(skb, skb->len - ETH_FCS_LEN);
>  
>  	skb->protocol = eth_type_trans(skb, skb->dev);

[Severity: Critical]
Can this lead to a use-after-free when an XDP program shrinks the frame?

data_len is now the length from after XDP, returned by
lan966x_xdp_pci_run(). bpf_xdp_adjust_tail() and bpf_xdp_adjust_head()
let it shrink to ETH_HLEN. Take a 14-byte frame with h_proto 0x8100
that returns XDP_PASS. eth_type_trans() then leaves skb->len == 0 and
skb->protocol == ETH_P_8021Q.

Now suppose src_port is in bridge_mask and IGMP/MLD redirect is enabled
in ANA_CPU_FWD_CFG. lan966x_port_mc_set() enables it when multicast
snooping is on, and snooping is on by default. In that case the code
after this point reaches:

drivers/net/ethernet/microchip/lan966x/lan966x_main.c:lan966x_hw_offload() {
    ...
	if (eth_type_vlan(skb->protocol)) {
		skb = skb_vlan_untag(skb);
		if (unlikely(!skb))
			return false;
	}
    ...
}

pskb_may_pull(skb, VLAN_HLEN + 2) fails in skb_vlan_untag(), so it
frees the skb with kfree_skb() and returns NULL. lan966x_hw_offload()
only returns a bool, so the caller still holds the freed pointer:

		if (!lan966x_hw_offload(lan966x, src_port, skb))
			skb->offload_fwd_mark = 0;
	}

	skb->dev->stats.rx_bytes += skb->len;

The freed skb is then returned to lan966x_fdma_pci_napi_poll(), which
passes it to napi_gro_receive(). Would that be a use-after-free followed
by a double free?

The problem in lan966x_hw_offload() is older. Before this patch it could
not be reached, because skbs were built from the hardware length and the
MAC drops runts. The frame contents can also come from the network, for
example when a decap program uses adjust_head on a minimum-size outer
frame.

[ ... ]

> @@ -312,8 +430,17 @@ static int lan966x_fdma_pci_napi_poll(struct napi_struct *napi, int weight)
>  			 */
>  			fdma_dcb_advance(fdma);
>  			continue;
> +		case FDMA_TX:
> +			fdma_dcb_advance(fdma);
> +			continue;
> +		case FDMA_DROP:
> +			fdma_dcb_advance(fdma);
> +			continue;
>  		}

[Severity: Medium]
What happens if one poll consumes all 256 (FDMA_PCI_DCB_MAX) RX
descriptors?

Every result path calls fdma_dcb_advance(), which wraps at n_dcbs. After
256 frames, dcb_index is back at the value saved in dcb_reload. The
re-arm loop at the end of lan966x_fdma_pci_napi_poll() then runs zero
times:

	while (dcb_reload != fdma->dcb_index) {
		old_dcb = dcb_reload;
		dcb_reload++;
		dcb_reload &= fdma->n_dcbs - 1;

So none of the consumed DCBs are re-armed, and they all keep
FDMA_DCB_STATUS_DONE. With a budget above 256, fdma_has_frames() stays
true and the same frames are processed again in the same call. With a
budget of exactly 256, the next poll processes them again.

Would this deliver stale frames twice, run XDP on them again, and resend
the XDP_TX frames?

Only the budget bounds the loop. __napi_busy_loop() passes the socket's
SO_BUSY_POLL_BUDGET straight to napi_poll(), and that value can be up to
U16_MAX with CAP_NET_ADMIN. The ring can fill completely, because the
hardware fills the last DCB before it stops on the invalid nextptr.

This loop was added by the earlier "net: lan966x: add PCIe FDMA support"
commit in this series. This patch extends it with the XDP_TX and
XDP_DROP results, and it does not seem to be fixed later in the series.
The platform lan966x_fdma_napi_poll() has the same pattern with 512
DCBs.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space
  2026-09-24 19:57 ` [PATCH net-next v8 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Concern: nothing enforces the contract between the DT "cpu" reg
  size and the driver's hard-coded TARGET_PCIE_DBI offset (0x400000)…

--

misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space

This patch grows the "cpu" reg of the switch node in the lan966x PCI
overlay from 1MB to 8MB. The FDMA engine programs its ATU outbound windows
through registers at offset 0x400000 and above, and the old mapping did
not reach them.

> diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> index 7b196b0a0eb65..7bb726550cafd 100644
> --- a/drivers/misc/lan966x_pci.dtso
> +++ b/drivers/misc/lan966x_pci.dtso
> @@ -135,7 +135,7 @@ lan966x_phy1: ethernet-lan966x_phy@2 {
>  
>  				switch: switch@e0000000 {
>  					compatible = "microchip,lan966x-switch";
> -					reg = <0xe0000000 0x0100000>,
> +					reg = <0xe0000000 0x0800000>,

[Severity: Low]
This isn't a bug, but the driver only works if the size here is large
enough, and nothing in the driver checks that.

The driver computes the DBI target from a fixed offset into the "cpu"
resource:

drivers/net/ethernet/microchip/lan966x/lan966x_main.c:
	{ TARGET_PCIE_DBI,             0x400000, 0 }, /* 0xe0400000 */

lan966x_create_targets() then adds that offset to the ioremapped base
without comparing it to resource_size():

	lan966x->regs[iomap->id] = begin[iomap->range] + iomap->offset;

fdma_pci_atu_addr_get() in drivers/net/ethernet/microchip/fdma/fdma_pci.c
adds a further FDMA_PCI_ATU_ADDR (0x300000) + idx * 0x200 + offset:

	return addr + FDMA_PCI_ATU_ADDR + FDMA_PCI_ATU_IDX_SIZE * idx + offset;

With this patch, the in-tree overlay covers the highest register touched
(about 0x700A18 into the cpu range). The native lan966x.dtsi still uses a
1MB cpu reg, but it never reaches this path, because lan966x_is_pci() is
false there and lan966x_shutdown() returns early.

So there is no in-tree problem. Would it still make sense to check in
lan966x_create_targets() that TARGET_PCIE_DBI plus the ATU range fits
within the "cpu" resource on PCI, rather than depending on the DT size
alone?

>  					      <0xe2000000 0x0800000>;
>  					reg-names = "cpu", "gcb";
>

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH net-next v8 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay
  2026-09-24 19:57 ` [PATCH net-next v8 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
@ 2026-09-25 20:52   ` netdev-bot+sashiko
  0 siblings, 0 replies; 24+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 20:52 UTC (permalink / raw)
  To: daniel.machon
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horatiu.vultur,
	steen.hegelund, UNGLinuxDriver, ast, daniel, hawk,
	john.fastabend, sdf, herve.codina, arnd, gregkh, mohsin.bashr,
	horms, richardcochran, netdev, linux-kernel, bpf,
	linux-arm-kernel

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 2 · Medium: 0 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] PCIe FDMA init races with a concurrent MTU change on netdevices
  that are already registered.
- [High] Removing a PCIe lan966x can deadlock against a concurrent MTU
  change.

--

misc: lan966x-pci: dts: add fdma interrupt to overlay

This patch adds the fdma interrupt (OIC interrupt 14) to the lan966x PCI
device tree overlay. With the "fdma" named interrupt present,
platform_get_irq_byname() succeeds and lan966x->fdma becomes true for the
PCIe card, which turns on the new PCIe FDMA init, resize and deinit paths.

> diff --git a/drivers/misc/lan966x_pci.dtso b/drivers/misc/lan966x_pci.dtso
> index 7bb726550cafd..5bb12dbc08431 100644
> --- a/drivers/misc/lan966x_pci.dtso
> +++ b/drivers/misc/lan966x_pci.dtso
> @@ -141,8 +141,9 @@ switch: switch@e0000000 {
>  
>  					interrupt-parent = <&oic>;
>  					interrupts = <12 IRQ_TYPE_LEVEL_HIGH>,
> +						     <14 IRQ_TYPE_LEVEL_HIGH>,
>  						     <9 IRQ_TYPE_LEVEL_HIGH>;
> -					interrupt-names = "xtr", "ana";
> +					interrupt-names = "xtr", "fdma", "ana";

[Severity: High]
This overlay change makes the PCIe FDMA path reachable. Can
lan966x_fdma_pci_init() race with an MTU change on a port that is already
registered?

In lan966x_probe(), the port netdevs are registered and NAPI is enabled
from inside the port loop:

lan966x_probe()
  lan966x_probe_port()
    register_netdev(dev)
  lan966x_port_init()
    lan966x_fdma_netdev_init()

Only after that loop does probe call lan966x->ops->fdma_init(), and it
does so without RTNL.

The only thing guarding the resize path is a plain pointer check in
lan966x_fdma_pci_resize():

	if (!lan966x->rx.lan966x)
		return 0;

However, lan966x_fdma_pci_init() sets that pointer first, before
rx.max_mtu, before tx.lan966x and before the rings are allocated:

	lan966x->rx.lan966x = lan966x;
	lan966x->rx.max_mtu = lan966x_fdma_get_max_frame(lan966x);
	...
	lan966x->tx.lan966x = lan966x;
	...
	err = lan966x_fdma_pci_rx_alloc(&lan966x->rx);
	...
	err = lan966x_fdma_pci_tx_alloc(&lan966x->tx);

If an RTM_SETLINK MTU change lands in that window (for example a network
manager configuring the new netdev), it would take this path:

lan966x_port_change_mtu()
  lan966x_fdma_pci_resize()
    __lan966x_fdma_pci_reload()
      lan966x_fdma_pci_reload()

lan966x_fdma_pci_reload() snapshots the half-built rings into
rx_fdma_old/tx_fdma_old. It then disables NAPI and the channels, and
allocates and frees rings in the same lan966x->rx.fdma/tx.fdma fields that
probe is still filling in.

If probe has not yet set tx.lan966x, lan966x_fdma_tx_disable() would pass
a NULL lan966x to lan_rmw(). If probe is past that point, both threads
would allocate, overwrite and free the same ring and ATU region fields at
the same time. Could that leak the rings and ATU windows, or free memory
that probe or the FDMA engine is still using?

The commit that adds PCIe FDMA MTU change support says the resize is
skipped "until lan966x_fdma_pci_init() has built the rings". The flag,
though, is set at the start of init rather than after the rings exist.
Would it be better to set it last, or to serialize init against
ndo_change_mtu, for example by holding RTNL?

[Severity: High]
Also enabled by this change: can lan966x_remove() deadlock against a
concurrent MTU change?

lan966x_remove() tears down the FDMA before the netdevs are unregistered.
It does not take RTNL and does not call netif_device_detach():

	lan966x->ops->fdma_deinit(lan966x);
	lan966x_cleanup_ports(lan966x);

lan966x_fdma_pci_deinit() disables NAPI and frees the rings. It leaves
rx.lan966x set:

	if (lan966x->fdma_ndev)
		napi_disable(&lan966x->napi);
	...
	fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->rx.fdma);
	fdma_free_coherent_and_unmap(lan966x->dma_dev, &lan966x->tx.fdma);

The ports stay registered and present until each unregister_netdev() in
lan966x_cleanup_ports() finishes. During that time, an MTU change that
alters the max frame size would still pass the check in
lan966x_fdma_pci_resize() and reach lan966x_fdma_pci_reload(), which
calls napi_disable() a second time with RTNL held:

	napi_disable(&lan966x->napi);

The first disable leaves NAPIF_STATE_SCHED set. napi_disable_locked()
therefore waits forever:

	while (val & (NAPIF_STATE_SCHED | NAPIF_STATE_NPSVC)) {
		usleep_range(20, 200);
		val = READ_ONCE(n->state);
	}

In that case the remove task would then block in unregister_netdev() on
rtnl_lock(), and so would every other RTNL user. If the second
napi_disable() did return, the reload would go on to free
rx_fdma_old/tx_fdma_old, which deinit has already freed.

lan966x_shutdown() in this series already guards against this. It takes
rtnl_lock() and calls netif_device_detach() on each port. The shutdown
commit says this is done "so ndo_open/ndo_change_mtu cannot re-enter the
FDMA against a disabled NAPI".

Does lan966x_remove() need the same protection before calling
fdma_deinit(), or should deinit clear rx.lan966x so that later resizes
are skipped?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924-lan966x-pci-fdma-v8-0-201c7b707d8b%40microchip.com

^ permalink raw reply	[flat|nested] 24+ messages in thread

end of thread, other threads:[~2026-09-25 20:52 UTC | newest]

Thread overview: 24+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-24 19:56 [PATCH net-next v8 00/15] net: lan966x: add support for PCIe FDMA Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 01/15] MAINTAINERS: add FDMA library to Sparx5 SoC entry Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 02/15] net: microchip: fdma: rename contiguous dataptr helpers Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 03/15] net: microchip: fdma: add PCIe ATU support Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko
2026-09-24 19:56 ` [PATCH net-next v8 04/15] net: microchip: fdma: use little-endian types for descriptor fields Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 05/15] net: lan966x: add FDMA LLP register write helper Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 06/15] net: lan966x: export FDMA helpers for reuse Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 07/15] net: lan966x: use a dedicated device for DMA operations Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 08/15] net: lan966x: add FDMA ops dispatch for PCIe support Daniel Machon
2026-09-24 19:56 ` [PATCH net-next v8 09/15] net: lan966x: clear FDMA interrupt stickies after switch reset Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko
2026-09-24 19:56 ` [PATCH net-next v8 10/15] net: lan966x: add shutdown callback to stop the FDMA on reboot Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko
2026-09-24 19:56 ` [PATCH net-next v8 11/15] net: lan966x: add PCIe FDMA support Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko
2026-09-24 19:57 ` [PATCH net-next v8 12/15] net: lan966x: add PCIe FDMA MTU change support Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko
2026-09-24 19:57 ` [PATCH net-next v8 13/15] net: lan966x: add PCIe FDMA XDP support Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko
2026-09-24 19:57 ` [PATCH net-next v8 14/15] misc: lan966x-pci: dts: extend cpu reg to cover PCIE DBI space Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko
2026-09-24 19:57 ` [PATCH net-next v8 15/15] misc: lan966x-pci: dts: add fdma interrupt to overlay Daniel Machon
2026-09-25 20:52   ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®