mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs
@ 2026-09-28 12:32 illusion.wang
  2026-09-28 12:32 ` [PATCH v30 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
                   ` (7 more replies)
  0 siblings, 8 replies; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

This series continues and refines the previous v28 NBL NIC driver patch
set for Nebula-matrix m18110/m18000 (SNIC s1000) Ethernet adapters.

The first two foundational patches of v28 have been merged into net-next
main branch:
- net/nebula-matrix: add minimum nbl build framework (cffde9e49e86)
- net/nebula-matrix: add core driver architecture and HW layer
  initialization (825e6d230163)

Following maintainer feedback to shorten the patch series, this
iteration drops the two merged commits, reorganizes and polishes the
remaining functionality into a standalone 8-patch sequence. No new
features are added; this series only completes the pending core
infrastructure.

The overall driver development plan remains consistent with the v28
two-phase roadmap:
1) Complete hardware initialization, mailbox/channel communication and
   control-plane infrastructure (this series)
2) Implement netdev registration and TX/RX data-path functionality
   (follow-up series)

This patch set supplements the merged foundational code, covering
channel layer abstraction, resource management, interrupt handling,
chip-level initialization, control dispatch routing, channel RPC
framework, and device lifecycle init/remove/start/stop logic.

changes v29->v30
Link to v29:https://lore.kernel.org/netdev/20260922120311.86593-1-illusion.wang@nebula-matrix.com/
1.AI review issues
changes v28->v29
Link to v28:https://lore.kernel.org/netdev/20260914123429.56596-1-illusion.wang@nebula-matrix.com/
1.AI review issues
2.remove the hash table implementation

illusion wang (8):
  net/nebula-matrix: add channel layer
  net/nebula-matrix: add common resource implementation
  net/nebula-matrix: add intr resource implementation
  net/nebula-matrix: add chip-wide hardware init/deinit implementation
  net/nebula-matrix: dispatch: add control-level routing core
    infrastructure
  net/nebula-matrix: dispatch: implement channel RPC framework and
    serialize hardware ops
  net/nebula-matrix: add common/ctrl dev init/remove operation
  net/nebula-matrix: add common dev start/stop operation

 .../net/ethernet/nebula-matrix/nbl/Makefile   |   10 +-
 .../nbl/nbl_channel/nbl_channel.c             | 1467 +++++++++++++++++
 .../nbl/nbl_channel/nbl_channel.h             |  181 ++
 .../nebula-matrix/nbl/nbl_common/nbl_common.c |   52 +
 .../nebula-matrix/nbl/nbl_common/nbl_common.h |   14 +
 .../net/ethernet/nebula-matrix/nbl/nbl_core.h |   20 +
 .../nebula-matrix/nbl/nbl_core/nbl_dev.c      |  625 +++++++
 .../nebula-matrix/nbl/nbl_core/nbl_dev.h      |   55 +
 .../nebula-matrix/nbl/nbl_core/nbl_dispatch.c |  680 ++++++++
 .../nebula-matrix/nbl/nbl_core/nbl_dispatch.h |   25 +
 .../nebula-matrix/nbl/nbl_hw/nbl_chip.c       |   32 +
 .../nebula-matrix/nbl/nbl_hw/nbl_chip.h       |   12 +
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c  | 1094 ++++++++++++
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h  |  320 ++++
 .../nbl_hw_leonis/nbl_resource_leonis.c       |  359 ++++
 .../nbl_hw_leonis/nbl_resource_leonis.h       |   12 +
 .../nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h     |   33 +
 .../nebula-matrix/nbl/nbl_hw/nbl_interrupt.c  |  703 ++++++++
 .../nebula-matrix/nbl/nbl_hw/nbl_interrupt.h  |   21 +
 .../nebula-matrix/nbl/nbl_hw/nbl_resource.c   |  157 ++
 .../nebula-matrix/nbl/nbl_hw/nbl_resource.h   |  121 ++
 .../nbl/nbl_include/nbl_def_channel.h         |  174 ++
 .../nbl/nbl_include/nbl_def_common.h          |   23 +
 .../nbl/nbl_include/nbl_def_dev.h             |   16 +
 .../nbl/nbl_include/nbl_def_dispatch.h        |   61 +
 .../nbl/nbl_include/nbl_def_hw.h              |   62 +
 .../nbl/nbl_include/nbl_def_resource.h        |   38 +
 .../nbl/nbl_include/nbl_include.h             |   32 +
 .../net/ethernet/nebula-matrix/nbl/nbl_main.c |  138 ++
 29 files changed, 6536 insertions(+), 1 deletion(-)
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h

-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 1/8] net/nebula-matrix: add channel layer
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  2026-09-28 12:32 ` [PATCH v30 net-next 2/8] net/nebula-matrix: add common resource implementation illusion.wang
                   ` (6 subsequent siblings)
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Add dedicated mailbox-based generic channel management layer for Nebula Matrix
Ethernet adapter, implementing core inter-PF control communication capability
between PF0 and other auxiliary PFs.
This patch builds fundamental inter-PF communication infrastructure for
multi-PF control plane interaction scenarios. It decouples hardware register
operations from business control logic, with complete hardware abstraction
and robust error recovery capabilities.
Key features of the new channel layer include:
- Thread-safe message handling based on xarray, supporting O(1) message
  lookup and duplicate handler protection. Both fire-and-forget and
  synchronous ACK-based transmission modes are supported.
- TX slot concurrency control, returning -EAGAIN when slots are full to
  avoid congestion and resource exhaustion. Supports 16B inline small
  payload and 4KB max DMA-based large payload transmission.
- Devm-managed mailbox queue resources, with one-time DMA allocation
  during probe. Full hardware queue lifecycle management covers init,
  runtime configuration, start, stop and teardown.
- Dual interrupt/polling RX processing, offloading cleanup work to
  dedicated workqueue to reduce IRQ latency. Reliable teardown logic
  prevents invalid DMA access and message loss.
- Isolated hw_ops layer decouples low-level hardware implementation from
  upper channel logic. Fine-grained register locking optimizes BAR0/BAR2
  access for stability under high stress.
- Complete auxiliary infrastructure including per-channel state control,
  timeout/retry recovery, and unified message pack/parse helpers for
  consistent upper-layer calling.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../net/ethernet/nebula-matrix/nbl/Makefile   |    4 +-
 .../nbl/nbl_channel/nbl_channel.c             | 1467 +++++++++++++++++
 .../nbl/nbl_channel/nbl_channel.h             |  181 ++
 .../nebula-matrix/nbl/nbl_common/nbl_common.c |   31 +
 .../nebula-matrix/nbl/nbl_common/nbl_common.h |   14 +
 .../net/ethernet/nebula-matrix/nbl/nbl_core.h |    7 +
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c  |  234 +++
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h  |   41 +
 .../nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h     |   32 +
 .../nbl/nbl_include/nbl_def_channel.h         |  125 ++
 .../nbl/nbl_include/nbl_def_common.h          |    4 +
 .../nbl/nbl_include/nbl_def_hw.h              |   35 +
 .../nbl/nbl_include/nbl_include.h             |    3 +
 .../net/ethernet/nebula-matrix/nbl/nbl_main.c |    7 +
 14 files changed, 2184 insertions(+), 1 deletion(-)
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index cc060cf8bf75..04e1aa1fb4bd 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -3,5 +3,7 @@
 
 obj-$(CONFIG_NBL) := nbl.o
 
-nbl-objs +=	nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
+nbl-objs +=	nbl_common/nbl_common.o \
+		nbl_channel/nbl_channel.o \
+		nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
 		nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
new file mode 100644
index 000000000000..d2c8182c72b6
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
@@ -0,0 +1,1467 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/delay.h>
+#include <linux/device.h>
+#include <linux/mutex.h>
+#include <linux/bitfield.h>
+#include <linux/pci.h>
+#include <linux/bits.h>
+#include <linux/dma-mapping.h>
+#include <linux/atomic.h>
+#include <linux/wait.h>
+#include "nbl_channel.h"
+
+static int nbl_chan_add_msg_handler(struct nbl_channel_mgt *chan_mgt,
+				    u16 msg_type, nbl_chan_resp func,
+				    void *priv)
+{
+	struct nbl_chan_msg_node_data *handler;
+	int ret;
+
+	handler = kzalloc_obj(*handler, GFP_KERNEL);
+	if (!handler)
+		return -ENOMEM;
+
+	handler->func = func;
+	handler->priv = priv;
+
+	/* Each msg_type is registered at most once; reject duplicates. */
+	mutex_lock(&chan_mgt->handler_lock);
+	ret = xa_insert(&chan_mgt->handler_xa, msg_type, handler, GFP_KERNEL);
+	mutex_unlock(&chan_mgt->handler_lock);
+	if (ret)
+		kfree(handler);
+
+	return ret;
+}
+
+static int nbl_chan_init_msg_handler(struct nbl_channel_mgt *chan_mgt)
+{
+	int ret;
+
+	ret = devm_mutex_init(chan_mgt->common->dev, &chan_mgt->handler_lock);
+	if (ret)
+		return ret;
+
+	xa_init(&chan_mgt->handler_xa);
+
+	return 0;
+}
+
+static void nbl_chan_remove_msg_handler(struct nbl_channel_mgt *chan_mgt)
+{
+	struct nbl_chan_msg_node_data *handler;
+	unsigned long msg_type;
+
+	mutex_lock(&chan_mgt->handler_lock);
+	xa_for_each(&chan_mgt->handler_xa, msg_type, handler) {
+		xa_erase(&chan_mgt->handler_xa, msg_type);
+		kfree(handler);
+	}
+	mutex_unlock(&chan_mgt->handler_lock);
+	xa_destroy(&chan_mgt->handler_xa);
+}
+
+static void nbl_chan_init_queue_param(struct nbl_chan_info *chan_info,
+				      u16 num_txq_entries, u16 num_rxq_entries,
+				      u16 txq_buf_size, u16 rxq_buf_size)
+{
+	chan_info->num_txq_entries = num_txq_entries;
+	chan_info->num_rxq_entries = num_rxq_entries;
+	chan_info->txq_buf_size = txq_buf_size;
+	chan_info->rxq_buf_size = rxq_buf_size;
+	atomic_set(&chan_info->inflight_tx_cnt, 0);
+	WRITE_ONCE(chan_info->shutdn, false);
+	WRITE_ONCE(chan_info->active, false);
+	WRITE_ONCE(chan_info->wait_head_index, 0);
+	memset(chan_info->state, 0, sizeof(chan_info->state));
+	init_waitqueue_head(&chan_info->inflight_wait);
+}
+
+static int nbl_chan_init_tx_queue(struct nbl_common_info *common,
+				  struct nbl_chan_info *chan_info)
+{
+	struct nbl_chan_ring *txq = &chan_info->txq;
+	struct device *dev = common->dev;
+	size_t size =
+		chan_info->num_txq_entries * sizeof(struct nbl_chan_tx_desc);
+	u16 i;
+
+	txq->desc.tx_desc =
+		dmam_alloc_coherent(dev, size, &txq->dma, GFP_KERNEL);
+	if (!txq->desc.tx_desc)
+		return -ENOMEM;
+
+	chan_info->wait = devm_kcalloc(dev, chan_info->num_txq_entries,
+				       sizeof(*chan_info->wait), GFP_KERNEL);
+	if (!chan_info->wait)
+		return -ENOMEM;
+	for (i = 0; i < chan_info->num_txq_entries; i++) {
+		init_waitqueue_head(&chan_info->wait[i].wait_queue);
+		WRITE_ONCE(chan_info->wait[i].status, NBL_MBX_STATUS_IDLE);
+		WRITE_ONCE(chan_info->wait[i].acked, 0);
+		WRITE_ONCE(chan_info->wait[i].ack_data, NULL);
+		WRITE_ONCE(chan_info->wait[i].ack_data_len, 0);
+		WRITE_ONCE(chan_info->wait[i].ack_err, 0);
+		WRITE_ONCE(chan_info->wait[i].msg_type, 0);
+		WRITE_ONCE(chan_info->wait[i].msg_index, 0);
+		WRITE_ONCE(chan_info->wait[i].dstid, 0);
+	}
+
+	txq->buf = devm_kcalloc(dev, chan_info->num_txq_entries,
+				sizeof(*txq->buf), GFP_KERNEL);
+	if (!txq->buf)
+		return -ENOMEM;
+
+	return 0;
+}
+
+static int nbl_chan_init_rx_queue(struct nbl_common_info *common,
+				  struct nbl_chan_info *chan_info)
+{
+	struct nbl_chan_ring *rxq = &chan_info->rxq;
+	struct device *dev = common->dev;
+	size_t size =
+		chan_info->num_rxq_entries * sizeof(struct nbl_chan_rx_desc);
+
+	rxq->desc.rx_desc =
+		dmam_alloc_coherent(dev, size, &rxq->dma, GFP_KERNEL);
+	if (!rxq->desc.rx_desc) {
+		dev_err_ratelimited(dev,
+				    "Allocate DMA for chan rx descriptor ring failed\n");
+		return -ENOMEM;
+	}
+
+	rxq->buf = devm_kcalloc(dev, chan_info->num_rxq_entries,
+				sizeof(*rxq->buf), GFP_KERNEL);
+	if (!rxq->buf)
+		return -ENOMEM;
+
+	return 0;
+}
+
+static int nbl_chan_init_queue(struct nbl_common_info *common,
+			       struct nbl_chan_info *chan_info)
+{
+	int err;
+
+	err = nbl_chan_init_tx_queue(common, chan_info);
+	if (err)
+		return err;
+
+	err = nbl_chan_init_rx_queue(common, chan_info);
+
+	return err;
+}
+
+static void nbl_chan_config_queue(struct nbl_channel_mgt *chan_mgt,
+				  struct nbl_chan_info *chan_info, bool tx)
+{
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+	struct nbl_hw_mgt *p = chan_mgt->hw_ops_tbl->priv;
+	struct nbl_chan_ring *ring;
+	dma_addr_t addr;
+	int size_bwid;
+
+	if (tx)
+		ring = &chan_info->txq;
+	else
+		ring = &chan_info->rxq;
+	addr = ring->dma;
+	if (tx) {
+		size_bwid = ilog2(chan_info->num_txq_entries);
+		hw_ops->config_mailbox_txq(p, addr, size_bwid);
+	} else {
+		size_bwid = ilog2(chan_info->num_rxq_entries);
+		hw_ops->config_mailbox_rxq(p, addr, size_bwid);
+	}
+}
+
+static int nbl_chan_alloc_all_tx_bufs(struct nbl_channel_mgt *chan_mgt,
+				      struct nbl_chan_info *chan_info)
+{
+	struct nbl_chan_ring *txq = &chan_info->txq;
+	struct device *dev = chan_mgt->common->dev;
+	struct nbl_chan_buf *buf;
+	u16 i;
+
+	for (i = 0; i < chan_info->num_txq_entries; i++) {
+		buf = &txq->buf[i];
+		buf->va = dmam_alloc_coherent(dev, chan_info->txq_buf_size,
+					      &buf->pa, GFP_KERNEL);
+		if (!buf->va) {
+			dev_err_ratelimited(dev,
+					    "Allocate buffer for chan tx queue failed\n");
+			return -ENOMEM;
+		}
+	}
+
+	txq->next_to_clean = 0;
+	txq->next_to_use = 0;
+	txq->tail_ptr = 0;
+
+	return 0;
+}
+
+static void nbl_chan_cfg_qinfo_map_table(struct nbl_channel_mgt *chan_mgt,
+					 u8 bus, u8 devid)
+{
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+	struct nbl_hw_mgt *p = chan_mgt->hw_ops_tbl->priv;
+	u32 pf_mask = 0;
+	u8 func_id;
+
+	/*
+	 * k_pf_mask rule: bit N == 0 means PF#N enabled, bit N == 1 masked out.
+	 * Program mailbox QINFO entry for each hardware-active PF func_id.
+	 *
+	 * Note: This loop iterates over raw hardware PF func_id.
+	 * Product constraints limit supported PF counts to contiguous sets:
+	 * PF0 only, PF0~1, or PF0~3. Non-contiguous or unsupported PF count
+	 * will be rejected in resource initialization logic.
+	 */
+	hw_ops->get_host_pf_mask(p, &pf_mask);
+	for (func_id = 0; func_id < NBL_MAX_PF; func_id++) {
+		if (!(pf_mask & (1 << func_id)))
+			hw_ops->cfg_mailbox_qinfo(p, func_id, bus,
+						  devid, func_id);
+	}
+}
+
+static int nbl_chan_alloc_all_rx_bufs(struct nbl_channel_mgt *chan_mgt,
+				      struct nbl_chan_info *chan_info)
+{
+	struct nbl_chan_ring *rxq = &chan_info->rxq;
+	struct device *dev = chan_mgt->common->dev;
+	struct nbl_chan_rx_desc *desc;
+	struct nbl_chan_buf *buf;
+	u16 i;
+
+	for (i = 0; i < chan_info->num_rxq_entries; i++) {
+		buf = &rxq->buf[i];
+		buf->va = dmam_alloc_coherent(dev, chan_info->rxq_buf_size,
+					      &buf->pa, GFP_KERNEL);
+		if (!buf->va) {
+			dev_err_ratelimited(dev,
+					    "Allocate buffer for chan rx queue failed\n");
+			goto err;
+		}
+	}
+
+	desc = rxq->desc.rx_desc;
+	/*
+	 * Initially leave one RX descriptor unused so that
+	 * next_to_clean and next_to_use can distinguish an empty
+	 * ring from a full ring.
+	 *
+	 * The unused slot is replenished as RX descriptors are
+	 * consumed and recycled.
+	 */
+	for (i = 0; i < chan_info->num_rxq_entries - 1; i++) {
+		buf = &rxq->buf[i];
+		desc[i].buf_addr = cpu_to_le64(buf->pa);
+		desc[i].buf_len = cpu_to_le32(chan_info->rxq_buf_size);
+		desc[i].flags = cpu_to_le16(BIT(NBL_CHAN_RX_DESC_AVAIL));
+	}
+
+	rxq->next_to_clean = 0;
+	rxq->next_to_use = chan_info->num_rxq_entries - 1;
+	rxq->tail_ptr = chan_info->num_rxq_entries - 1;
+
+	return 0;
+err:
+	return -ENOMEM;
+}
+
+static int nbl_chan_alloc_all_bufs(struct nbl_channel_mgt *chan_mgt,
+				   struct nbl_chan_info *chan_info)
+{
+	int err;
+
+	err = nbl_chan_alloc_all_tx_bufs(chan_mgt, chan_info);
+	if (err)
+		return err;
+	err = nbl_chan_alloc_all_rx_bufs(chan_mgt, chan_info);
+
+	return err;
+}
+
+/*
+ * The RX QINFO region is written only by config/stop paths, never by
+ * senders (which only touch the TX QINFO region and the doorbell), so
+ * stopping RX does not require txq_lock.
+ */
+static void nbl_chan_stop_rx_queue(struct nbl_channel_mgt *chan_mgt)
+{
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+
+	hw_ops->stop_mailbox_rxq(chan_mgt->hw_ops_tbl->priv);
+}
+
+/*
+ * The TX QINFO region is also written by nbl_chan_quiesce_and_reclaim_tx()
+ * under txq_lock, so TX stop must serialize against it via txq_lock.
+ */
+static void nbl_chan_stop_tx_queue(struct nbl_channel_mgt *chan_mgt)
+{
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+
+	hw_ops->stop_mailbox_txq(chan_mgt->hw_ops_tbl->priv);
+}
+
+static void nbl_chan_reset_wait_head(struct nbl_chan_info *chan_info,
+				     struct nbl_chan_waitqueue_head *wait_head)
+{
+	lockdep_assert_held(&chan_info->pending_lock);
+
+	WRITE_ONCE(wait_head->acked, 0);
+	WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_IDLE);
+	WRITE_ONCE(wait_head->ack_data, NULL);
+	WRITE_ONCE(wait_head->ack_data_len, 0);
+	WRITE_ONCE(wait_head->ack_err, 0);
+	WRITE_ONCE(wait_head->msg_type, 0);
+	WRITE_ONCE(wait_head->dstid, 0);
+}
+
+static int nbl_chan_teardown_queue(struct nbl_channel_mgt *chan_mgt,
+				   u8 chan_type)
+{
+	struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+	struct nbl_chan_waitqueue_head *wait_head;
+	struct work_struct *task;
+	u16 i;
+
+	if (!READ_ONCE(chan_info->active)) {
+		dev_warn(chan_mgt->common->dev, "channel not active, skip duplicate teardown\n");
+		return 0;
+	}
+	/*
+	 * Step1:
+	 * block new sender
+	 */
+	mutex_lock(&chan_info->state_lock);
+	WRITE_ONCE(chan_info->shutdn, true);
+	task = READ_ONCE(chan_info->clean_task);
+	WRITE_ONCE(chan_info->clean_task, NULL);
+	mutex_unlock(&chan_info->state_lock);
+	/*
+	 * Step2:
+	 * abort pending ACK waiters
+	 *
+	 * Poke only slots with a live waiter (status == WAITING): publish a
+	 * synthetic (-EIO, len 0) completion and wake the owner. The slot is
+	 * deliberately NOT reset to IDLE here - ownership stays with the
+	 * sender, which consumes the completion and resets its own slot in
+	 * nbl_chan_send_msg():out_clear_wait_slot under pending_lock.
+	 *
+	 * Resetting an owned slot here would corrupt two things:
+	 *   - an ACKD slot whose sender is consuming the completion
+	 *     locklessly (smp_rmb()-ordered reads) could observe torn
+	 *     ack_data_len/ack_err pairs;
+	 *   - nbl_chan_get_msg_id() hands out IDLE/TIMEOUT slots, so a slot
+	 *     reset to IDLE under its live owner could be reallocated to a
+	 *     sender admitted before shutdn was set.
+	 *
+	 * IDLE/TIMEOUT slots have no waiter and need no action. A late ACK
+	 * racing this poke is serialized by pending_lock: either the real
+	 * completion lands first (status becomes ACKD and is skipped) or
+	 * this synthetic failure does, both are valid outcomes.
+	 */
+	mutex_lock(&chan_info->pending_lock);
+	for (i = 0; i < chan_info->num_txq_entries; i++) {
+		wait_head = &chan_info->wait[i];
+		if (READ_ONCE(wait_head->status) != NBL_MBX_STATUS_WAITING)
+			continue;
+		WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_ACKD);
+		WRITE_ONCE(wait_head->ack_data_len, 0);
+		WRITE_ONCE(wait_head->ack_err, (s32)-EIO);
+		/* Publish completion fields before the acked flag */
+		smp_wmb();
+		WRITE_ONCE(wait_head->acked, 1);
+		wake_up(&wait_head->wait_queue);
+	}
+	mutex_unlock(&chan_info->pending_lock);
+
+	/*
+	 * Step3:
+	 * stop the RX queue unconditionally. The RX QINFO region is written
+	 * only by config/stop paths, never by senders (which only touch the
+	 * TX QINFO region and the doorbell), so no txq_lock is needed. The
+	 * RX engine must be quiesced before the rings are freed, regardless
+	 * of TX-side state - skipping it would leave the device DMAing peer
+	 * messages and descriptor writebacks into released coherent memory.
+	 */
+	nbl_chan_stop_rx_queue(chan_mgt);
+
+	/*
+	 * Drain strategy mirrors mlx5 command interface teardown:
+	 * set shutdown flag first, abort all pending waiters, then
+	 * block until inflight_tx_cnt reaches zero.
+	 *
+	 * After shutdn is set every sender exits promptly at its next
+	 * checkpoint:
+	 *   - interrupt-driven senders wake on shutdn immediately
+	 *     (it is part of the wait_event condition);
+	 *   - ACK polling senders check shutdn each iteration before
+	 *     sleeping 1000-1200us. This per-iteration check avoids
+	 *     waiting the full 5.0-6.0s worst-case ACK timeout.
+	 *
+	 * The wait is deliberately unbounded: a surviving sender still
+	 * holds txq_lock or touches chan_info (state_lock, wait[]) on its
+	 * exit path, so returning early would let devres free the rings
+	 * and chan_info underneath it. In practice the TX polling loop is
+	 * bounded (NBL_CHAN_TX_WAIT_TIMES) and wedged-link MMIO reads
+	 * eventually complete with all-1s data, so this loop terminates;
+	 * warn periodically to keep a genuinely stuck sender diagnosable.
+	 */
+	while (wait_event_timeout(chan_info->inflight_wait,
+				  atomic_read(&chan_info->inflight_tx_cnt) == 0,
+				  msecs_to_jiffies(5000)) == 0)
+		dev_warn(chan_mgt->common->dev,
+			 "teardown: still waiting for %d inflight sender(s)\n",
+			 atomic_read(&chan_info->inflight_tx_cnt));
+
+	/*
+	 * Join the last exiting sender through state_lock: it drops
+	 * inflight_tx_cnt and wakes this drain while still holding
+	 * state_lock, so observing the counter reach zero does not prove
+	 * the sender has finished its mutex_unlock(). Acquiring the lock
+	 * once here guarantees no sender still holds it - mutex_unlock()
+	 * is the sender's last touch of chan_info - before returning lets
+	 * devres free chan_info. shutdn blocks new senders, so one
+	 * acquire/release pair is sufficient.
+	 */
+	mutex_lock(&chan_info->state_lock);
+	mutex_unlock(&chan_info->state_lock);
+
+	/*
+	 * All senders have exited, so txq_lock is uncontended: plain lock.
+	 * Take it so this is the sole writer of the TX QINFO block
+	 * (sender-side stop/config/doorbell sequences all run under
+	 * txq_lock).
+	 */
+	mutex_lock(&chan_info->txq_lock);
+	nbl_chan_stop_tx_queue(chan_mgt);
+	mutex_unlock(&chan_info->txq_lock);
+
+	/*
+	 * Join the RX clean work before active=false and before devres
+	 * releases the rings. IRQ is already freed and clean_task is
+	 * cleared under state_lock, so no new instance can be queued.
+	 */
+	if (task)
+		cancel_work_sync(task);
+	WRITE_ONCE(chan_info->active, false);
+	return 0;
+}
+
+static int nbl_chan_setup_queue(struct nbl_channel_mgt *chan_mgt, u8 chan_type)
+{
+	struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+	struct nbl_common_info *common = chan_mgt->common;
+	struct nbl_chan_ring *rxq = &chan_info->rxq;
+	int err;
+
+	if (READ_ONCE(chan_info->active)) {
+		dev_warn(common->dev, "channel already active, reject duplicate setup\n");
+		return -EBUSY;
+	}
+	/*
+	 * DMA resources are allocated once and released only by devres
+	 * at device detach.  A teardown does not free them, so a second
+	 * setup would orphan ~2 MiB of coherent DMA per channel.
+	 */
+	if (chan_info->dma_allocated) {
+		dev_warn(common->dev,
+			 "channel DMA already allocated, re-setup not supported\n");
+		return -EBUSY;
+	}
+	nbl_chan_init_queue_param(chan_info, NBL_CHAN_QUEUE_LEN,
+				  NBL_CHAN_QUEUE_LEN, NBL_CHAN_BUF_LEN,
+				  NBL_CHAN_BUF_LEN);
+	err = nbl_chan_init_queue(common, chan_info);
+	if (err)
+		return err;
+	err = nbl_chan_alloc_all_bufs(chan_mgt, chan_info);
+	if (err)
+		return err;
+	nbl_chan_config_queue(chan_mgt, chan_info, true); /* tx */
+	nbl_chan_config_queue(chan_mgt, chan_info, false); /* rx */
+	nbl_chan_update_tail_ptr(hw_ops, chan_mgt->hw_ops_tbl->priv,
+				 rxq->tail_ptr, NBL_MB_RX_QID);
+	WRITE_ONCE(chan_info->active, true);
+	chan_info->dma_allocated = true;
+	return 0;
+}
+
+static bool nbl_chan_txq_full(struct nbl_chan_ring *txq,
+			      u16 num_entries)
+{
+	return NBL_NEXT_ID(txq->next_to_use, num_entries - 1) ==
+	       txq->next_to_clean;
+}
+
+static int nbl_chan_update_txqueue(struct nbl_channel_mgt *chan_mgt,
+				   struct nbl_chan_info *chan_info,
+				   struct nbl_chan_tx_param *param)
+{
+	struct nbl_chan_ring *txq = &chan_info->txq;
+	struct nbl_chan_tx_desc *tx_desc;
+	struct nbl_chan_buf *tx_buf;
+
+	if (nbl_chan_txq_full(txq, chan_info->num_txq_entries))
+		return -EBUSY;
+	if (param->arg_len > NBL_CHAN_BUF_LEN - sizeof(*tx_desc))
+		return -EINVAL;
+	tx_desc =
+		NBL_CHAN_TX_RING_TO_DESC(txq, txq->next_to_use);
+	tx_buf =
+		NBL_CHAN_TX_RING_TO_BUF(txq, txq->next_to_use);
+	tx_desc->dstid = cpu_to_le16(param->dstid);
+	tx_desc->msg_type = cpu_to_le16(param->msg_type);
+	tx_desc->msgid = cpu_to_le16(param->msgid);
+
+	/*
+	 * srcid field is filled by mailbox hardware after peer receives this
+	 * packet, driver producer never writes srcid; reused descriptor slots
+	 * will contain stale srcid value temporarily until hardware overwrites
+	 * it.
+	 */
+	if (param->arg_len > NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN) {
+		if (param->arg)
+			memcpy(tx_buf->va, param->arg, param->arg_len);
+		tx_desc->buf_addr = cpu_to_le64(tx_buf->pa);
+		tx_desc->buf_len = cpu_to_le16(param->arg_len);
+		tx_desc->data_len = 0;
+		memset(tx_desc->data, 0, sizeof(tx_desc->data));
+	} else {
+		memset(tx_desc->data, 0, sizeof(tx_desc->data));
+		memset(&tx_desc->buf_addr, 0, sizeof(tx_desc->buf_addr));
+		if (param->arg && param->arg_len > 0)
+			memcpy(tx_desc->data, param->arg, param->arg_len);
+		tx_desc->buf_len = 0;
+		tx_desc->data_len = cpu_to_le16(param->arg_len);
+	}
+	/* Ensure descriptor data visible to device before AVAIL flag */
+	dma_wmb();
+	tx_desc->flags = cpu_to_le16(BIT(NBL_CHAN_TX_DESC_AVAIL));
+
+	txq->next_to_use =
+		NBL_NEXT_ID(txq->next_to_use, chan_info->num_txq_entries - 1);
+	txq->tail_ptr++;
+
+	return 0;
+}
+
+/*
+ * Quiesce the TX mailbox queue and reclaim all outstanding
+ * descriptors.  Called from the timeout path of nbl_chan_kick_tx_ring()
+ * with txq_lock held.
+ *
+ * The device failed to fetch/complete the current descriptor within the
+ * polling window.  We assert QUEUE_RST to stop further DMA fetches,
+ * reclaim every descriptor between next_to_clean and next_to_use,
+ * reset the software tail_ptr counter to match the hardware reset state,
+ * and re-enable the queue so subsequent sends can proceed.
+ */
+static void nbl_chan_quiesce_and_reclaim_tx(struct nbl_channel_mgt *chan_mgt,
+					    struct nbl_chan_info *chan_info)
+{
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+	struct nbl_hw_mgt *hw_priv = chan_mgt->hw_ops_tbl->priv;
+	struct nbl_chan_ring *txq = &chan_info->txq;
+	struct nbl_chan_tx_desc *tx_desc;
+
+	/*
+	 * Assert QUEUE_RST to stop hardware fetching new descriptors.
+	 * stop_mailbox_txq() flushes through the mailbox BAR itself, so
+	 * the reset has reached the device when it returns.
+	 */
+	hw_ops->stop_mailbox_txq(hw_priv);
+
+	/*
+	 * Reclaim all outstanding descriptors between next_to_clean and
+	 * next_to_use.  Under txq_lock there is at most one in-flight
+	 * descriptor, but iterate the full range for robustness.
+	 */
+	while (txq->next_to_clean != txq->next_to_use) {
+		tx_desc = NBL_CHAN_TX_RING_TO_DESC(txq,
+						   txq->next_to_clean);
+		WRITE_ONCE(tx_desc->flags, 0);
+		txq->next_to_clean =
+			NBL_NEXT_ID(txq->next_to_clean,
+				    chan_info->num_txq_entries - 1);
+	}
+
+	/*
+	 * Hardware tail_ptr counter is cleared by QUEUE_RST.  Reset
+	 * software counter to match so the next doorbell update does
+	 * not produce a false 16-bit wrap delta.
+	 */
+	txq->tail_ptr = 0;
+	txq->next_to_use = 0;
+	txq->next_to_clean = 0;
+
+	/*
+	 * Authoritative re-arm gate, evaluated while txq_lock is held:
+	 * teardown sets shutdn before draining senders and then owns
+	 * hardware state. If teardown started while we were stopping/
+	 * reclaiming, leave the queue in QUEUE_RST and never program
+	 * QUEUE_EN, otherwise the queue could be armed again pointing
+	 * at rings that are about to be freed.
+	 */
+	if (READ_ONCE(chan_info->shutdn))
+		return;
+
+	/* Re-enable queue with current ring base and size */
+	nbl_chan_config_queue(chan_mgt, chan_info, true);
+}
+
+static int nbl_chan_kick_tx_ring(struct nbl_channel_mgt *chan_mgt,
+				 struct nbl_chan_info *chan_info)
+{
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+	struct nbl_chan_ring *txq = &chan_info->txq;
+	struct device *dev = chan_mgt->common->dev;
+	int max_retries = NBL_CHAN_TX_WAIT_TIMES;
+	struct nbl_chan_tx_desc *tx_desc;
+	int retry_count = 0;
+	u16 msg_type;
+
+	nbl_chan_update_tail_ptr(hw_ops, chan_mgt->hw_ops_tbl->priv,
+				 txq->tail_ptr, NBL_MB_TX_QID);
+
+	tx_desc = NBL_CHAN_TX_RING_TO_DESC(txq, txq->next_to_clean);
+	/*
+	 * Poll for HW to mark descriptor as USED.
+	 * Mailbox is a low-speed control channel for management commands.
+	 * We avoid enabling dedicated per-TX interrupt for single control
+	 * message to reduce interrupt overhead, so use bounded polling
+	 * with small delay instead.
+	 */
+	while (retry_count < max_retries) {
+		if (READ_ONCE(chan_info->shutdn))
+			return -ESHUTDOWN;
+
+		if (le16_to_cpu(READ_ONCE(tx_desc->flags)) &
+		    BIT(NBL_CHAN_TX_DESC_USED)) {
+			 /*
+			  * Order reads of other device-written descriptor
+			  * fields after observing USED.  Matches the RX side
+			  * pattern in nbl_chan_clean_queue().
+			  */
+			dma_rmb();
+			break;
+		}
+
+		retry_count++;
+		if (retry_count == max_retries) {
+			msg_type = le16_to_cpu(READ_ONCE(tx_desc->msg_type));
+			dev_err_ratelimited(dev, "chan send msg type: %d timeout\n",
+					    msg_type);
+			/*
+			 * Teardown may have started after the loop-top
+			 * check on this final iteration. Do not run lock
+			 * recovery then: it would stop/reclaim the queue
+			 * and, without the in-quiesce gate, re-arm it.
+			 * nbl_chan_teardown_queue() owns hardware state.
+			 */
+			if (READ_ONCE(chan_info->shutdn))
+				return -ESHUTDOWN;
+			/*
+			 * Device failed to complete this descriptor.
+			 * Quiesce the queue, reclaim the timed-out
+			 * descriptor, and re-enable so future sends can
+			 * proceed instead of stalling the ring full.
+			 */
+			nbl_chan_quiesce_and_reclaim_tx(chan_mgt,
+							chan_info);
+			return -ETIMEDOUT;
+		}
+		usleep_range(NBL_CHAN_TX_WAIT_US, NBL_CHAN_TX_WAIT_US_MAX);
+	}
+
+	txq->next_to_clean = txq->next_to_use;
+
+	return 0;
+}
+
+static void nbl_chan_recv_ack_msg(void *priv, u16 srcid, u16 msgid, void *data,
+				  u32 data_len)
+{
+	struct nbl_channel_mgt *chan_mgt = (struct nbl_channel_mgt *)priv;
+	struct nbl_chan_waitqueue_head *wait_head = NULL;
+	struct device *dev = chan_mgt->common->dev;
+	struct nbl_chan_info *chan_info =
+		chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX];
+	u16 w_dstid, w_msgtype, w_msgidx;
+	u32 *payload = data;
+	u16 ack_msgtype = 0;
+	u16 ack_msgid = 0;
+	u32 ack_datalen;
+	void *ack_data;
+	u32 copy_len;
+	int w_status;
+	s32 raw_err;
+
+	if (READ_ONCE(chan_info->shutdn))
+		return;
+	if (data_len > NBL_CHAN_BUF_LEN ||
+	    data_len < NBL_CHAN_ACK_HEAD_LEN * sizeof(u32)) {
+		dev_err_ratelimited(dev, "Invalid ACK data_len: %u\n",
+				    data_len);
+		return;
+	}
+	ack_datalen = data_len - NBL_CHAN_ACK_HEAD_LEN * sizeof(u32);
+	ack_msgtype = le16_to_cpu(*(__le16 *)(payload + NBL_CHAN_MSG_TYPE_POS));
+	ack_msgid = le16_to_cpu(*(__le16 *)(payload + NBL_CHAN_MSG_ID_POS));
+	if (FIELD_GET(NBL_CHAN_MSGID_LOC_MASK, ack_msgid) >=
+	    chan_info->num_txq_entries) {
+		dev_err_ratelimited(dev, "chan recv msg id: %u err\n",
+				    ack_msgid);
+		return;
+	}
+	wait_head =
+		&chan_info->wait[FIELD_GET(NBL_CHAN_MSGID_LOC_MASK, ack_msgid)];
+
+	mutex_lock(&chan_info->pending_lock);
+
+	/* Cache repeated READ_ONCE values */
+	w_dstid = READ_ONCE(wait_head->dstid);
+	w_status = READ_ONCE(wait_head->status);
+	w_msgtype = READ_ONCE(wait_head->msg_type);
+	w_msgidx = READ_ONCE(wait_head->msg_index);
+
+	if (srcid != w_dstid) {
+		mutex_unlock(&chan_info->pending_lock);
+		dev_err_ratelimited(dev, "ACK srcid=%u != dstid=%u, rejecting\n",
+				    srcid, w_dstid);
+		return;
+	}
+	if (w_status != NBL_MBX_STATUS_WAITING) {
+		mutex_unlock(&chan_info->pending_lock);
+		dev_err_ratelimited(dev,
+				    "Skip ack invalid status, wait msgtype:%u idx:%u status:%d ack msgtype:%u msgid:%u datalen:%u\n",
+				    w_msgtype, w_msgidx, w_status,
+				    ack_msgtype, ack_msgid, ack_datalen);
+		return;
+	}
+
+	if (w_msgtype != ack_msgtype) {
+		mutex_unlock(&chan_info->pending_lock);
+		dev_err_ratelimited(dev,
+				    "Skip ack msgtype mismatch, wait msgtype:%u idx:%u ack msgtype:%u msgid:%u\n",
+				    w_msgtype, w_msgidx, ack_msgtype,
+				    ack_msgid);
+		return;
+	}
+	if (FIELD_GET(NBL_CHAN_MSGID_INDEX_MASK, ack_msgid) != w_msgidx) {
+		mutex_unlock(&chan_info->pending_lock);
+		dev_err_ratelimited(dev,
+				    "Stale ACK: expected index=%u, got msgid=%u\n",
+				    w_msgidx, ack_msgid);
+		return;
+	}
+
+	raw_err = (s32)le32_to_cpu(*(__le32 *)&payload[NBL_CHAN_ACK_RET_POS]);
+	if (raw_err > 0 || raw_err < -MAX_ERRNO)
+		raw_err = -EREMOTEIO;
+
+	WRITE_ONCE(wait_head->ack_err, raw_err);
+
+	copy_len = min_t(u32, READ_ONCE(wait_head->ack_data_len), ack_datalen);
+	if (READ_ONCE(wait_head->ack_err) >= 0 && copy_len > 0) {
+		ack_data = READ_ONCE(wait_head->ack_data);
+		if (!ack_data) {
+			dev_err_ratelimited(dev, "ACK payload dropped: ack_data is NULL\n");
+			WRITE_ONCE(wait_head->ack_data_len, 0);
+			goto ack_done;
+		}
+		memcpy((char *)ack_data,
+		       payload + NBL_CHAN_ACK_HEAD_LEN, copy_len);
+		WRITE_ONCE(wait_head->ack_data_len, (u16)copy_len);
+	} else {
+		WRITE_ONCE(wait_head->ack_data_len, 0);
+	}
+ack_done:
+	/* Guarantee payload data finished before acked flag visible */
+	smp_wmb();
+	WRITE_ONCE(wait_head->acked, 1);
+	WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_ACKD);
+	mutex_unlock(&chan_info->pending_lock);
+	wake_up(&wait_head->wait_queue);
+}
+
+static void nbl_chan_recv_msg(struct nbl_channel_mgt *chan_mgt, void *data)
+{
+	struct device *dev = chan_mgt->common->dev;
+	struct nbl_chan_msg_node_data *msg_handler;
+	u16 msg_type, payload_len, srcid, msgid;
+	struct nbl_chan_info *chan_info =
+		chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX];
+	struct nbl_chan_tx_desc *tx_desc;
+	void *payload;
+	size_t avail_space;
+	u16 data_len_fw;
+
+	if (READ_ONCE(chan_info->shutdn))
+		return;
+
+	tx_desc = data;
+	msg_type = le16_to_cpu(READ_ONCE(tx_desc->msg_type));
+	dev_dbg(dev, "recv msg_type: %d\n", msg_type);
+
+	srcid = le16_to_cpu(READ_ONCE(tx_desc->srcid));
+	msgid = le16_to_cpu(READ_ONCE(tx_desc->msgid));
+
+	if (msg_type >= NBL_CHAN_MSG_MAILBOX_MAX)
+		return;
+
+	data_len_fw = le16_to_cpu(READ_ONCE(tx_desc->data_len));
+	if (data_len_fw) {
+		payload_len = data_len_fw;
+
+		if (payload_len > NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN) {
+			dev_err_ratelimited(dev,
+					    "data_len=%u exceeds embedded buffer size=%u\n",
+					    payload_len,
+					    NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN);
+			return;
+		}
+		/* Small pkt: payload stored inside descriptor data[] array */
+		payload = tx_desc->data;
+	} else {
+		payload_len = le16_to_cpu(READ_ONCE(tx_desc->buf_len));
+
+		avail_space = NBL_CHAN_BUF_LEN - sizeof(*tx_desc);
+		if (payload_len > avail_space) {
+			dev_err_ratelimited(dev,
+					    "buf_len=%u exceeds external buffer size=%zu\n",
+					    payload_len, avail_space);
+			return;
+		}
+		/* Large pkt: payload follows immediately after tx_desc */
+		payload = tx_desc + 1;
+	}
+
+	msg_handler = xa_load(&chan_mgt->handler_xa, msg_type);
+	if (!msg_handler || !msg_handler->func) {
+		dev_err_ratelimited(dev,
+				    "No handler for msg_type: %u (srcid=%u, msgid=%u)\n",
+				    msg_type, srcid, msgid);
+		return;
+	}
+
+	msg_handler->func(msg_handler->priv, srcid, msgid, payload,
+			  payload_len);
+}
+
+static void nbl_chan_advance_rx_ring(struct nbl_channel_mgt *chan_mgt,
+				     struct nbl_chan_info *chan_info,
+				     struct nbl_chan_ring *rxq)
+{
+	struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+	struct nbl_chan_rx_desc *rx_desc;
+	struct nbl_chan_buf *rx_buf;
+	u16 next_to_use;
+
+	next_to_use = rxq->next_to_use;
+	rx_desc = NBL_CHAN_RX_RING_TO_DESC(rxq, next_to_use);
+	rx_buf = NBL_CHAN_RX_RING_TO_BUF(rxq, next_to_use);
+
+	/*
+	 * Recycle the RX descriptor at next_to_use. The initial
+	 * unused slot is intentionally recycled after the first
+	 * RX descriptor is consumed, allowing the ring to become
+	 * fully populated while next_to_clean tracks the consumer.
+	 */
+	rx_desc->buf_addr = cpu_to_le64(rx_buf->pa);
+	rx_desc->buf_len = cpu_to_le32(chan_info->rxq_buf_size);
+
+	/*
+	 * DMA Write Memory Barrier:
+	 * Ensures all previous DMA-mapped writes (buffer address/length)
+	 * are completed before the descriptor flags are updated.
+	 * This prevents hardware from seeing a partially updated descriptor
+	 * where flags are set but buffer info isn't ready yet.
+	 */
+	dma_wmb();
+
+	rx_desc->flags = cpu_to_le16(BIT(NBL_CHAN_RX_DESC_AVAIL));
+
+	rxq->next_to_use++;
+	if (rxq->next_to_use == chan_info->num_rxq_entries)
+		rxq->next_to_use = 0;
+	rxq->tail_ptr++;
+
+	nbl_chan_update_tail_ptr(hw_ops, chan_mgt->hw_ops_tbl->priv,
+				 rxq->tail_ptr, NBL_MB_RX_QID);
+}
+
+static void nbl_chan_clean_queue(struct nbl_channel_mgt *chan_mgt,
+				 struct nbl_chan_info *chan_info)
+{
+	struct nbl_common_info *common = chan_mgt->common;
+	struct nbl_chan_ring *rxq = &chan_info->rxq;
+	struct device *dev = chan_mgt->common->dev;
+	u32 budget = NBL_CHAN_RX_CLEAN_BUDGET;
+	struct nbl_chan_rx_desc *rx_desc;
+	struct nbl_chan_buf *rx_buf;
+	struct work_struct *task;
+	bool more_work = false;
+	u16 next_to_clean;
+	u16 flags;
+
+	next_to_clean = rxq->next_to_clean;
+	rx_desc = NBL_CHAN_RX_RING_TO_DESC(rxq, next_to_clean);
+	rx_buf = NBL_CHAN_RX_RING_TO_BUF(rxq, next_to_clean);
+	while (le16_to_cpu(READ_ONCE(rx_desc->flags)) &
+	       BIT(NBL_CHAN_RX_DESC_USED)) {
+		flags = le16_to_cpu(READ_ONCE(rx_desc->flags));
+
+		if (READ_ONCE(chan_info->shutdn))
+			break;
+		if (!(flags & BIT(NBL_CHAN_RX_DESC_WRITE)))
+			dev_dbg(dev,
+				"mailbox rx flag 0x%x missing NBL_CHAN_RX_DESC_WRITE\n",
+				flags);
+
+		/* Make sure hardware written descriptor visible to CPU */
+		dma_rmb();
+		nbl_chan_recv_msg(chan_mgt, rx_buf->va);
+		nbl_chan_advance_rx_ring(chan_mgt, chan_info, rxq);
+		next_to_clean++;
+		if (next_to_clean == chan_info->num_rxq_entries)
+			next_to_clean = 0;
+		rx_desc = NBL_CHAN_RX_RING_TO_DESC(rxq, next_to_clean);
+		rx_buf = NBL_CHAN_RX_RING_TO_BUF(rxq, next_to_clean);
+		if (--budget == 0) {
+			more_work = true;
+			break;
+		}
+		cond_resched();
+	}
+	rxq->next_to_clean = next_to_clean;
+
+	mutex_lock(&chan_info->state_lock);
+	/* Prevent queue_work after teardown clears clean_task */
+	if (READ_ONCE(chan_info->shutdn)) {
+		mutex_unlock(&chan_info->state_lock);
+		return;
+	}
+	if (common->wq && more_work) {
+		task = READ_ONCE(chan_info->clean_task);
+		if (task)
+			queue_work(common->wq, task);
+	}
+	mutex_unlock(&chan_info->state_lock);
+}
+
+static void nbl_chan_clean_queue_subtask(struct nbl_channel_mgt *chan_mgt,
+					 u8 chan_type)
+{
+	struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+
+	nbl_chan_clean_queue(chan_mgt, chan_info);
+}
+
+static int nbl_chan_get_msg_id(struct nbl_chan_info *chan_info,
+			       u16 *msgid)
+{
+	int search_loc = READ_ONCE(chan_info->wait_head_index), i;
+	struct nbl_chan_waitqueue_head *wait = NULL;
+	int status;
+	int next;
+
+	lockdep_assert_held(&chan_info->pending_lock);
+	for (i = 0; i < chan_info->num_txq_entries; i++) {
+		wait = &chan_info->wait[search_loc];
+		status = READ_ONCE(wait->status);
+		if (status == NBL_MBX_STATUS_IDLE ||
+		    status == NBL_MBX_STATUS_TIMEOUT) {
+			WRITE_ONCE(wait->msg_index,
+				   NBL_NEXT_ID(wait->msg_index,
+					       NBL_CHAN_MSG_INDEX_MAX));
+
+			*msgid = FIELD_PREP(NBL_CHAN_MSGID_INDEX_MASK,
+					    wait->msg_index) |
+				 FIELD_PREP(NBL_CHAN_MSGID_LOC_MASK,
+					    search_loc);
+
+			/* Advance starting search position for next caller */
+			next = NBL_NEXT_ID(search_loc,
+					   chan_info->num_txq_entries - 1);
+			WRITE_ONCE(chan_info->wait_head_index, next);
+			return 0;
+		}
+
+		search_loc = NBL_NEXT_ID(search_loc,
+					 chan_info->num_txq_entries - 1);
+	}
+
+	/*
+	 * All tx slots are occupied. May happen under high transmit load
+	 * or delayed remote ACK responses. Caller should retry later.
+	 */
+	return -EAGAIN;
+}
+
+static int nbl_chan_send_msg(struct nbl_channel_mgt *chan_mgt,
+			     struct nbl_chan_send_info *chan_send)
+{
+	struct nbl_common_info *common = chan_mgt->common;
+	struct nbl_chan_waitqueue_head *wait_head = NULL;
+	struct nbl_chan_tx_param tx_param = { 0 };
+	int i = NBL_CHAN_TX_WAIT_ACK_TIMES;
+	struct nbl_chan_info *chan_info =
+		chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX];
+	struct device *dev = common->dev;
+	struct work_struct *task;
+	u16 msgid = 0;
+	int ret;
+
+	if (chan_send->resp_len > NBL_CHAN_BUF_LEN) {
+		dev_err_ratelimited(dev, "resp_len %zu exceeds max %d\n",
+				    chan_send->resp_len, NBL_CHAN_BUF_LEN);
+		return -EINVAL;
+	}
+
+	mutex_lock(&chan_info->state_lock);
+	if (READ_ONCE(chan_info->shutdn)) {
+		mutex_unlock(&chan_info->state_lock);
+		return -ESHUTDOWN;
+	}
+	atomic_inc(&chan_info->inflight_tx_cnt);
+	mutex_unlock(&chan_info->state_lock);
+
+	tx_param.msg_type = chan_send->msg_type;
+	tx_param.arg = chan_send->arg;
+	tx_param.arg_len = chan_send->arg_len;
+	tx_param.dstid = chan_send->dstid;
+	tx_param.msgid = msgid;
+	if (chan_send->ack) {
+		mutex_lock(&chan_info->pending_lock);
+
+		ret = nbl_chan_get_msg_id(chan_info, &msgid);
+		if (ret) {
+			mutex_unlock(&chan_info->pending_lock);
+			dev_err_ratelimited(dev,
+					    "Channel tx wait head full, send msgtype:%u to dstid:%u failed\n",
+					    chan_send->msg_type,
+					    chan_send->dstid);
+			goto out_clean_inflight;
+		}
+		wait_head =
+			&chan_info->wait[FIELD_GET(NBL_CHAN_MSGID_LOC_MASK,
+						   msgid)];
+		WRITE_ONCE(wait_head->acked, 0);
+		WRITE_ONCE(wait_head->ack_data, chan_send->resp);
+		WRITE_ONCE(wait_head->ack_data_len, chan_send->resp_len);
+		WRITE_ONCE(wait_head->msg_type, chan_send->msg_type);
+		WRITE_ONCE(wait_head->msg_index,
+			   FIELD_GET(NBL_CHAN_MSGID_INDEX_MASK, msgid));
+		WRITE_ONCE(wait_head->dstid, chan_send->dstid);
+
+		WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_WAITING);
+		mutex_unlock(&chan_info->pending_lock);
+
+		tx_param.msgid = msgid;
+	}
+
+	mutex_lock(&chan_info->txq_lock);
+	ret = nbl_chan_update_txqueue(chan_mgt, chan_info, &tx_param);
+	if (ret) {
+		mutex_unlock(&chan_info->txq_lock);
+		dev_err_ratelimited(dev,
+				    "Channel tx queue full, send msgtype:%u to dstid:%u failed\n",
+				    chan_send->msg_type, chan_send->dstid);
+		if (wait_head)
+			goto out_clear_wait_slot;
+		goto out_clean_inflight;
+	}
+
+	ret = nbl_chan_kick_tx_ring(chan_mgt, chan_info);
+	mutex_unlock(&chan_info->txq_lock);
+	if (ret) {
+		if (wait_head)
+			goto out_clear_wait_slot;
+		goto out_clean_inflight;
+	}
+
+	if (!chan_send->ack) {
+		ret = 0;
+		goto out_clean_inflight;
+	}
+
+	if (test_bit(NBL_CHAN_IRQ_RDY, chan_info->state)) {
+		while (!READ_ONCE(wait_head->acked)) {
+			/*
+			 * avoids long task blocking when interrupt mode is
+			 * disabled mid-wait. Cannot guarantee subsequent ACK
+			 * delivery after interrupt mask off, only prevents
+			 * infinite blocking. Spurious timeout is possible.
+			 */
+			ret = wait_event_timeout(wait_head->wait_queue,
+						 READ_ONCE(wait_head->acked) ||
+						 READ_ONCE(chan_info->shutdn) ||
+						 !test_bit(NBL_CHAN_IRQ_RDY,
+							   chan_info->state),
+						 NBL_CHAN_ACK_WAIT_TIME);
+
+			if (READ_ONCE(chan_info->shutdn)) {
+				ret = -ESHUTDOWN;
+				goto out_clear_wait_slot;
+			}
+			if (!test_bit(NBL_CHAN_IRQ_RDY, chan_info->state)) {
+				ret = -EIO;
+				goto out_clear_wait_slot;
+			}
+			if (ret == 0) {
+				mutex_lock(&chan_info->pending_lock);
+				if (READ_ONCE(wait_head->status) ==
+				    NBL_MBX_STATUS_WAITING) {
+					WRITE_ONCE(wait_head->status,
+						   NBL_MBX_STATUS_TIMEOUT);
+					WRITE_ONCE(wait_head->acked, 0);
+					WRITE_ONCE(wait_head->ack_data, NULL);
+					WRITE_ONCE(wait_head->ack_data_len, 0);
+					/*
+					 * Ensure all status/ack slot
+					 * updates are visible before subsequent
+					 * readers observe acked == 0
+					 */
+					smp_wmb();
+					mutex_unlock(&chan_info->pending_lock);
+					dev_err_ratelimited(dev,
+							    "Channel waiting ack failed, message type: %d, msg id: %u\n",
+							    chan_send->msg_type,
+							    msgid);
+					ret = -ETIMEDOUT;
+					/*
+					 * TIMEOUT slots can be reused by
+					 * another sender. A late ACK is
+					 * rejected by the receive path,
+					 * which only accepts WAITING slots
+					 * (and by the msg_index generation
+					 * check after reallocation). The
+					 * current sender no longer owns the
+					 * slot, so skip the IDLE reset.
+					 */
+					goto out_clean_inflight;
+				}
+				/*
+				 * ACK won the timeout race: between the
+				 * final condition check and this lock the
+				 * receiver published acked=1/ACKD and
+				 * copied the payload into our response
+				 * buffer. The slot is still exclusively
+				 * ours (only IDLE/TIMEOUT slots are ever
+				 * reallocated), so consume the completion:
+				 * unlock and continue to the loop-bottom
+				 * acked check, which enters the normal
+				 * success readout and IDLE reset. Leaving
+				 * here via the timeout path would strand
+				 * the slot in ACKD forever, permanently
+				 * shrinking the 256-entry pool, and would
+				 * report a false -ETIMEDOUT for a request
+				 * whose response already arrived.
+				 */
+				mutex_unlock(&chan_info->pending_lock);
+			}
+
+			if (READ_ONCE(wait_head->acked))
+				break;
+		}
+		if (READ_ONCE(wait_head->acked)) {
+			/*
+			 * Load ordering: observe acked flag before
+			 * reading ACK payload metadata.
+			 */
+			smp_rmb();
+			chan_send->ack_len = READ_ONCE(wait_head->ack_data_len);
+			ret = READ_ONCE(wait_head->ack_err);
+		}
+	} else {
+		/* Polling path for synchronous ACK */
+		while (i--) {
+			if (READ_ONCE(chan_info->shutdn)) {
+				ret = -ESHUTDOWN;
+				goto out_clear_wait_slot;
+			}
+
+			mutex_lock(&chan_info->state_lock);
+			task = READ_ONCE(chan_info->clean_task);
+			if (common->wq && task &&
+			    !READ_ONCE(chan_info->shutdn) &&
+			    !work_pending(task))
+				queue_work(common->wq, task);
+			mutex_unlock(&chan_info->state_lock);
+			if (READ_ONCE(wait_head->acked)) {
+				/*
+				 * Guarantee load order: observe acked
+				 * flag before reading ack payload metadata.
+				 */
+				smp_rmb();
+				chan_send->ack_len =
+					READ_ONCE(wait_head->ack_data_len);
+				ret = READ_ONCE(wait_head->ack_err);
+				goto out_clear_wait_slot;
+			}
+
+			usleep_range(NBL_CHAN_TX_WAIT_ACK_US_MIN,
+				     NBL_CHAN_TX_WAIT_ACK_US_MAX);
+			cond_resched();
+		}
+		mutex_lock(&chan_info->pending_lock);
+		if (READ_ONCE(wait_head->status) == NBL_MBX_STATUS_ACKD) {
+			chan_send->ack_len = READ_ONCE(wait_head->ack_data_len);
+			ret = READ_ONCE(wait_head->ack_err);
+			mutex_unlock(&chan_info->pending_lock);
+			goto out_clear_wait_slot;
+		}
+		/*
+		 * Polling timed out without receiving an ACK.  Transition
+		 * the slot to TIMEOUT so nbl_chan_get_msg_id() can reuse
+		 * it.  Must not proceed to out_clean_inflight with
+		 * the slot still in WAITING state — that would leak the
+		 * slot permanently and eventually exhaust the 256-entry
+		 * channel queue.
+		 */
+		if (READ_ONCE(wait_head->status) == NBL_MBX_STATUS_WAITING) {
+			WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_TIMEOUT);
+			WRITE_ONCE(wait_head->acked, 0);
+			WRITE_ONCE(wait_head->ack_data, NULL);
+			WRITE_ONCE(wait_head->ack_data_len, 0);
+		}
+		mutex_unlock(&chan_info->pending_lock);
+		dev_err_ratelimited(dev,
+				    "Channel polling ack failed, message type: %d msg id: %u\n",
+				    chan_send->msg_type, msgid);
+		ret = -ETIMEDOUT;
+		goto out_clean_inflight;
+	}
+
+out_clear_wait_slot:
+	mutex_lock(&chan_info->pending_lock);
+	nbl_chan_reset_wait_head(chan_info, wait_head);
+	mutex_unlock(&chan_info->pending_lock);
+
+out_clean_inflight:
+	mutex_lock(&chan_info->state_lock);
+	if (atomic_dec_and_test(&chan_info->inflight_tx_cnt))
+		wake_up(&chan_info->inflight_wait);
+	mutex_unlock(&chan_info->state_lock);
+	return ret;
+}
+
+static int nbl_chan_send_ack(struct nbl_channel_mgt *chan_mgt,
+			     struct nbl_chan_ack_info *chan_ack)
+{
+	size_t head_len = NBL_CHAN_ACK_HEAD_LEN * sizeof(u32);
+	size_t data_len = chan_ack->data_len;
+	struct nbl_chan_send_info chan_send;
+	__le32 *tmp;
+	size_t len;
+	int ret;
+
+	if (data_len >
+	    NBL_CHAN_BUF_LEN - sizeof(struct nbl_chan_tx_desc) - head_len)
+		return -EINVAL;
+
+	len = head_len + data_len;
+	tmp = kzalloc(len, GFP_KERNEL);
+	if (!tmp)
+		return -ENOMEM;
+
+	*(__le16 *)&tmp[NBL_CHAN_MSG_TYPE_POS] =
+		cpu_to_le16(chan_ack->msg_type);
+	*(__le16 *)&tmp[NBL_CHAN_MSG_ID_POS] = cpu_to_le16(chan_ack->msgid);
+	tmp[NBL_CHAN_ACK_RET_POS] = cpu_to_le32(chan_ack->err);
+	if (chan_ack->data && chan_ack->data_len)
+		memcpy(&tmp[NBL_CHAN_ACK_HEAD_LEN], chan_ack->data,
+		       chan_ack->data_len);
+
+	nbl_chan_fill_send_info(&chan_send, chan_ack->dstid, NBL_CHAN_MSG_ACK,
+				tmp, len, NULL, 0, 0);
+	ret = nbl_chan_send_msg(chan_mgt, &chan_send);
+	kfree(tmp);
+
+	return ret;
+}
+
+static int nbl_chan_register_msg(struct nbl_channel_mgt *chan_mgt, u16 msg_type,
+				 nbl_chan_resp func, void *callback)
+{
+	return nbl_chan_add_msg_handler(chan_mgt, msg_type, func, callback);
+}
+
+static bool nbl_chan_check_queue_exist(struct nbl_channel_mgt *chan_mgt,
+				       u8 chan_type)
+{
+	struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+
+	return chan_info ? true : false;
+}
+
+static void nbl_chan_register_chan_task(struct nbl_channel_mgt *chan_mgt,
+					u8 chan_type, struct work_struct *task)
+{
+	struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+
+	mutex_lock(&chan_info->state_lock);
+	if (!READ_ONCE(chan_info->shutdn))
+		WRITE_ONCE(chan_info->clean_task, task);
+	mutex_unlock(&chan_info->state_lock);
+}
+
+static void nbl_chan_set_queue_state(struct nbl_channel_mgt *chan_mgt,
+				     enum nbl_chan_state state, u8 chan_type,
+				     u8 set)
+{
+	struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+	int i;
+
+	if (set)
+		set_bit(state, chan_info->state);
+	else
+		clear_bit(state, chan_info->state);
+	/*
+	 * When clearing IRQ_RDY, wake all per-slot wait queues so
+	 * sleeping senders observe the condition immediately and
+	 * return -EIO instead of waiting out the 3s timeout and
+	 * reporting -ETIMEDOUT.
+	 */
+	if (!set && state == NBL_CHAN_IRQ_RDY) {
+		for (i = 0; i < chan_info->num_txq_entries; i++)
+			wake_up_all(&chan_info->wait[i].wait_queue);
+	}
+}
+
+static struct nbl_channel_ops chan_ops = {
+	.send_msg			= nbl_chan_send_msg,
+	.send_ack			= nbl_chan_send_ack,
+	.register_msg			= nbl_chan_register_msg,
+	.cfg_chan_qinfo_map_table	= nbl_chan_cfg_qinfo_map_table,
+	.check_queue_exist		= nbl_chan_check_queue_exist,
+	.setup_queue			= nbl_chan_setup_queue,
+	.teardown_queue			= nbl_chan_teardown_queue,
+	.clean_queue_subtask		= nbl_chan_clean_queue_subtask,
+	.register_chan_task		= nbl_chan_register_chan_task,
+	.set_queue_state		= nbl_chan_set_queue_state,
+};
+
+static struct nbl_channel_mgt *
+nbl_chan_setup_chan_mgt(struct nbl_adapter *adapter)
+{
+	struct nbl_hw_ops_tbl *hw_ops_tbl = adapter->intf.hw_ops_tbl;
+	struct nbl_common_info *common = &adapter->common;
+	struct device *dev = &adapter->pdev->dev;
+	struct nbl_channel_mgt *chan_mgt;
+	struct nbl_chan_info *mailbox;
+	int ret;
+
+	chan_mgt = devm_kzalloc(dev, sizeof(*chan_mgt), GFP_KERNEL);
+	if (!chan_mgt)
+		return ERR_PTR(-ENOMEM);
+
+	chan_mgt->common = common;
+	chan_mgt->hw_ops_tbl = hw_ops_tbl;
+
+	mailbox = devm_kzalloc(dev, sizeof(*mailbox), GFP_KERNEL);
+	if (!mailbox)
+		return ERR_PTR(-ENOMEM);
+	mailbox->chan_type = NBL_CHAN_TYPE_MAILBOX;
+	chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX] = mailbox;
+
+	ret = nbl_chan_init_msg_handler(chan_mgt);
+	if (ret)
+		return ERR_PTR(ret);
+	ret = devm_mutex_init(common->dev, &mailbox->txq_lock);
+	if (ret)
+		return ERR_PTR(ret);
+	ret = devm_mutex_init(common->dev, &mailbox->state_lock);
+	if (ret)
+		return ERR_PTR(ret);
+	ret = devm_mutex_init(common->dev, &mailbox->pending_lock);
+	if (ret)
+		return ERR_PTR(ret);
+	return chan_mgt;
+}
+
+static struct nbl_channel_ops_tbl *
+nbl_chan_setup_ops(struct device *dev, struct nbl_channel_mgt *chan_mgt)
+{
+	struct nbl_channel_ops_tbl *chan_ops_tbl;
+	int ret;
+
+	chan_ops_tbl = devm_kzalloc(dev, sizeof(*chan_ops_tbl), GFP_KERNEL);
+	if (!chan_ops_tbl)
+		return ERR_PTR(-ENOMEM);
+	if (!chan_ops.send_msg || !chan_ops.send_ack ||
+	    !chan_ops.register_msg || !chan_ops.cfg_chan_qinfo_map_table ||
+	    !chan_ops.check_queue_exist || !chan_ops.setup_queue ||
+	    !chan_ops.teardown_queue || !chan_ops.clean_queue_subtask ||
+	    !chan_ops.register_chan_task || !chan_ops.set_queue_state)
+		return ERR_PTR(-EINVAL);
+
+	chan_ops_tbl->ops = &chan_ops;
+	chan_ops_tbl->priv = chan_mgt;
+
+	ret = nbl_chan_register_msg(chan_mgt, NBL_CHAN_MSG_ACK,
+				    nbl_chan_recv_ack_msg, chan_mgt);
+	if (ret)
+		return ERR_PTR(ret);
+
+	return chan_ops_tbl;
+}
+
+int nbl_chan_init_common(struct nbl_adapter *adap)
+{
+	struct nbl_channel_ops_tbl *chan_ops_tbl;
+	struct device *dev = &adap->pdev->dev;
+	struct nbl_channel_mgt *chan_mgt;
+	int ret;
+
+	chan_mgt = nbl_chan_setup_chan_mgt(adap);
+	if (IS_ERR(chan_mgt)) {
+		ret = PTR_ERR(chan_mgt);
+		goto exit;
+	}
+
+	chan_ops_tbl = nbl_chan_setup_ops(dev, chan_mgt);
+	if (IS_ERR(chan_ops_tbl)) {
+		ret = PTR_ERR(chan_ops_tbl);
+		goto cleanup_mgt;
+	}
+
+	adap->intf.channel_ops_tbl = chan_ops_tbl;
+	adap->core.chan_mgt = chan_mgt;
+	ret = nbl_common_create_wq(&adap->common);
+	if (ret)
+		goto cleanup_mgt;
+	return 0;
+
+cleanup_mgt:
+	nbl_chan_remove_msg_handler(chan_mgt);
+exit:
+	return ret;
+}
+
+void nbl_chan_remove_common(struct nbl_adapter *adap)
+{
+	struct nbl_channel_mgt *chan_mgt = adap->core.chan_mgt;
+
+	if (!chan_mgt)
+		return;
+	nbl_common_destroy_wq(&adap->common);
+	/*
+	 * All channel queues shall be torn down earlier in remove path
+	 * to drain inflight tx workers and stop hardware before destroying
+	 * message handler xarray.
+	 */
+	nbl_chan_remove_msg_handler(chan_mgt);
+	adap->core.chan_mgt = NULL;
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
new file mode 100644
index 000000000000..21cbb21946cf
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
@@ -0,0 +1,181 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_CHANNEL_H_
+#define _NBL_CHANNEL_H_
+
+#include <linux/types.h>
+#include <linux/xarray.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_hw.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+#define NBL_CHAN_TX_RING_TO_DESC(tx_ring, i) \
+	(&((((tx_ring)->desc.tx_desc))[i]))
+#define NBL_CHAN_RX_RING_TO_DESC(rx_ring, i) \
+	(&((((rx_ring)->desc.rx_desc))[i]))
+#define NBL_CHAN_TX_RING_TO_BUF(tx_ring, i) (&(((tx_ring)->buf)[i]))
+#define NBL_CHAN_RX_RING_TO_BUF(rx_ring, i) (&(((rx_ring)->buf)[i]))
+
+#define NBL_CHAN_TX_WAIT_US			100
+#define NBL_CHAN_TX_WAIT_US_MAX			120
+#define NBL_CHAN_TX_WAIT_TIMES			100
+#define NBL_CHAN_TX_WAIT_ACK_US_MIN		1000
+#define NBL_CHAN_TX_WAIT_ACK_US_MAX		1200
+#define NBL_CHAN_TX_WAIT_ACK_TIMES		5000
+#define NBL_CHAN_QUEUE_LEN			256
+#define NBL_CHAN_BUF_LEN			4096
+#define NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN	16
+
+#define NBL_CHAN_TX_DESC_AVAIL			0
+#define NBL_CHAN_TX_DESC_USED			1
+#define NBL_CHAN_RX_DESC_WRITE			1
+#define NBL_CHAN_RX_DESC_AVAIL			3
+#define NBL_CHAN_RX_DESC_USED			4
+
+#define NBL_CHAN_ACK_HEAD_LEN			3
+#define NBL_CHAN_ACK_RET_POS			2
+#define NBL_CHAN_MSG_ID_POS			1
+#define NBL_CHAN_MSG_TYPE_POS			0
+
+#define NBL_CHAN_ACK_WAIT_TIME			(3 * HZ)
+#define NBL_CHAN_RX_CLEAN_BUDGET		64
+
+enum {
+	NBL_MB_RX_QID = 0,
+	NBL_MB_TX_QID = 1,
+};
+
+enum {
+	NBL_MBX_STATUS_IDLE = 0,
+	NBL_MBX_STATUS_WAITING,
+	NBL_MBX_STATUS_ACKD,
+	NBL_MBX_STATUS_TIMEOUT,
+};
+
+struct nbl_chan_tx_param {
+	enum nbl_chan_msg_type msg_type;
+	void *arg;
+	size_t arg_len;
+	u16 dstid;
+	u16 msgid;
+};
+
+struct nbl_chan_buf {
+	void *va;
+	dma_addr_t pa;
+	size_t size;
+};
+
+struct nbl_chan_tx_desc {
+	__le16 flags;
+	__le16 srcid;
+	__le16 dstid;
+	__le16 data_len;
+	__le16 buf_len;
+	__le64 buf_addr;
+	__le16 msg_type;
+	u8 data[NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN];
+	__le16 msgid;
+	u8 rsv[26];
+} __packed;
+
+struct nbl_chan_rx_desc {
+	__le16 flags;
+	__le32 buf_len;
+	__le16 buf_id;
+	__le64 buf_addr;
+} __packed;
+
+union nbl_chan_desc_ptr {
+	struct nbl_chan_tx_desc *tx_desc;
+	struct nbl_chan_rx_desc *rx_desc;
+};
+
+struct nbl_chan_ring {
+	union nbl_chan_desc_ptr desc;
+	struct nbl_chan_buf *buf;
+	u16 next_to_use;
+	u16 tail_ptr; /* hardware does modulo ring size internally */
+	u16 next_to_clean;
+	dma_addr_t dma;
+};
+
+#define NBL_CHAN_MSG_INDEX_MAX 63
+
+#define NBL_CHAN_MSGID_INDEX_MASK GENMASK(5, 0)
+#define NBL_CHAN_MSGID_LOC_MASK GENMASK(13, 6)
+
+static inline void nbl_chan_update_tail_ptr(struct nbl_hw_ops *hw_ops,
+					    void *hw_priv, u32 tail_ptr, u8 qid)
+{
+	hw_ops->update_mailbox_queue_tail_ptr(hw_priv, tail_ptr, qid);
+}
+
+struct nbl_chan_waitqueue_head {
+	struct wait_queue_head wait_queue;
+	char *ack_data;
+	int acked;
+	s32 ack_err;
+	u16 ack_data_len;
+	u16 msg_type;
+	int status;
+	u8 msg_index;
+	u16 dstid;
+};
+
+struct nbl_chan_info {
+	wait_queue_head_t inflight_wait;
+	struct nbl_chan_ring txq;
+	struct nbl_chan_ring rxq;
+	struct nbl_chan_waitqueue_head *wait;
+	/*
+	 *Protects access to the TX queue (txq) and related metadata.
+	 *This mutex ensures exclusive access when updating the TX queue
+	 */
+	struct mutex txq_lock;
+	/* Guards channel state bitmap and shutdn flag */
+	struct mutex state_lock;
+	/* Guards pending requests and pending work list operations */
+	struct mutex pending_lock;
+	struct work_struct *clean_task;
+	u16 wait_head_index;
+	u16 num_txq_entries;
+	u16 num_rxq_entries;
+	u16 txq_buf_size;
+	u16 rxq_buf_size;
+	DECLARE_BITMAP(state, NBL_CHAN_STATE_NBITS);
+	u8 chan_type;
+	atomic_t inflight_tx_cnt;
+	bool shutdn;
+	bool active;
+	/*
+	 * One-shot lifecycle flag: DMA ring/buffer allocations use devm/
+	 * dmam and are only released at device detach.  Teardown only
+	 * stops the hardware and clears active; it does not free memory.
+	 * setup must never run twice on the same chan_info or the first
+	 * allocation set would be orphaned.
+	 */
+	bool dma_allocated;
+};
+
+struct nbl_chan_msg_node_data {
+	nbl_chan_resp func;
+	void *priv;
+};
+
+struct nbl_channel_mgt {
+	struct nbl_common_info *common;
+	struct nbl_hw_ops_tbl *hw_ops_tbl;
+	struct nbl_chan_info *chan_info[NBL_CHAN_TYPE_MAX];
+	/* Serializes handler_xa registration and teardown */
+	struct mutex handler_lock;
+	struct xarray handler_xa;
+};
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
new file mode 100644
index 000000000000..ce0ed2869af4
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
@@ -0,0 +1,31 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#include <linux/device.h>
+#include "nbl_common.h"
+
+void nbl_common_destroy_wq(struct nbl_common_info *common)
+{
+	if (!common || !common->wq)
+		return;
+
+	destroy_workqueue(common->wq);
+	common->wq = NULL;
+}
+
+int nbl_common_create_wq(struct nbl_common_info *common)
+{
+	char wq_name[32];
+
+	snprintf(wq_name, sizeof(wq_name), "nbl_wq_%s", pci_name(common->pdev));
+	common->wq = alloc_workqueue(wq_name, WQ_UNBOUND, 0);
+	if (!common->wq) {
+		dev_err(common->dev, "Failed to alloc workqueue %s\n", wq_name);
+		return -ENOMEM;
+	}
+
+	return 0;
+}
+
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h
new file mode 100644
index 000000000000..129073b08eb9
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h
@@ -0,0 +1,14 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_COMMON_H_
+#define _NBL_COMMON_H_
+
+#include <linux/types.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_common.h"
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index 1cd6587a8fcb..f998a2b44e5c 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -14,13 +14,20 @@ enum {
 	NBL_CAP_HAS_NET_BIT,
 };
 
+struct nbl_interface {
+	struct nbl_hw_ops_tbl *hw_ops_tbl;
+	struct nbl_channel_ops_tbl *channel_ops_tbl;
+};
+
 struct nbl_core {
 	struct nbl_hw_mgt *hw_mgt;
+	struct nbl_channel_mgt *chan_mgt;
 };
 
 struct nbl_adapter {
 	struct pci_dev *pdev;
 	struct nbl_core core;
+	struct nbl_interface intf;
 	struct nbl_common_info common;
 };
 
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index cf40ddc45192..eeff6216e4aa 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -6,8 +6,213 @@
 #include <linux/pci.h>
 #include <linux/bits.h>
 #include <linux/io.h>
+#include <linux/spinlock.h>
+#include <linux/bitfield.h>
 #include "nbl_hw_leonis.h"
 
+static void nbl_hw_read_mbx_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
+				 u32 len)
+{
+	u32 i;
+
+	if (len % 4)
+		return;
+	if (reg >= (u64)hw_mgt->mailbox_bar_size ||
+	    reg + len > (u64)hw_mgt->mailbox_bar_size) {
+		dev_err_once(hw_mgt->common->dev,
+			     "mbx read out of range: reg=0x%llx len=%u bar_size=%pa\n",
+			     reg, len, &hw_mgt->mailbox_bar_size);
+		return;
+	}
+	for (i = 0; i < len / 4; i++)
+		data[i] = nbl_mbx_rd32(hw_mgt, reg + i * sizeof(u32));
+}
+
+static void nbl_hw_write_mbx_regs(struct nbl_hw_mgt *hw_mgt, u64 reg,
+				  const u32 *data, u32 len)
+{
+	u32 i;
+
+	if (len % 4)
+		return;
+	if (reg >= (u64)hw_mgt->mailbox_bar_size ||
+	    reg + len > (u64)hw_mgt->mailbox_bar_size) {
+		dev_err_once(hw_mgt->common->dev,
+			     "mbx write out of range: reg=0x%llx len=%u bar_size=%pa\n",
+			     reg, len, &hw_mgt->mailbox_bar_size);
+		return;
+	}
+	for (i = 0; i < len / 4; i++)
+		nbl_mbx_wr32(hw_mgt, reg + i * sizeof(u32), data[i]);
+}
+
+/*
+ * Flush posted mailbox-BAR writes by reading back through the same
+ * BAR.
+ */
+static void nbl_hw_flush_mbx_write(struct nbl_hw_mgt *hw_mgt, u64 reg)
+{
+	u32 data;
+
+	nbl_hw_read_mbx_regs(hw_mgt, reg, &data, sizeof(data));
+}
+
+static void nbl_hw_rd_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
+			   u32 len)
+{
+	u32 size = len / 4;
+	u32 i;
+
+	if (len % 4)
+		return;
+	for (i = 0; i < size; i++)
+		data[i] = rd32(hw_mgt->hw_addr, reg + i * sizeof(u32));
+}
+
+static void nbl_hw_wr_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, const u32 *data,
+			   u32 len)
+{
+	u32 size = len / 4;
+	u32 i;
+
+	if (len % 4)
+		return;
+	for (i = 0; i < size; i++)
+		wr32(hw_mgt->hw_addr, reg + i * sizeof(u32), data[i]);
+}
+
+static void nbl_hw_rd_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
+				u32 len)
+{
+	u32 size = len / 4;
+	u32 i;
+
+	if (len % 4)
+		return;
+
+	spin_lock(&hw_mgt->reg_lock);
+
+	for (i = 0; i < size; i++)
+		data[i] = rd32(hw_mgt->hw_addr, reg + i * sizeof(u32));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_hw_update_mailbox_queue_tail_ptr(struct nbl_hw_mgt *hw_mgt,
+						 u16 tail_ptr, u8 txrx)
+{
+	/* local_qid 0 and 1 denote rx and tx queue respectively */
+	u32 local_qid = txrx;
+	u32 value = ((u32)tail_ptr << 16) | local_qid;
+
+	/* wmb for doorbell */
+	wmb();
+	nbl_mbx_wr32(hw_mgt, NBL_MAILBOX_NOTIFY_ADDR, value);
+}
+
+static void nbl_hw_config_mailbox_rxq(struct nbl_hw_mgt *hw_mgt,
+				      dma_addr_t dma_addr, int size_bwid)
+{
+	struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+	memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+	cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+	nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR,
+			      cfg_tbl.data, sizeof(cfg_tbl));
+
+	cfg_tbl.data[0] = lower_32_bits(dma_addr);
+	cfg_tbl.data[1] = upper_32_bits(dma_addr);
+	cfg_tbl.data[2] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_SIZE_BWID_MASK,
+				     size_bwid);
+	cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 0) |
+			  FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_EN_MASK, 1);
+	nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR,
+			      cfg_tbl.data, sizeof(cfg_tbl));
+}
+
+static void nbl_hw_config_mailbox_txq(struct nbl_hw_mgt *hw_mgt,
+				      dma_addr_t dma_addr, int size_bwid)
+{
+	struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+	memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+	cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+	nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR,
+			      cfg_tbl.data, sizeof(cfg_tbl));
+
+	cfg_tbl.data[0] = lower_32_bits(dma_addr);
+	cfg_tbl.data[1] = upper_32_bits(dma_addr);
+	cfg_tbl.data[2] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_SIZE_BWID_MASK,
+				     size_bwid);
+	cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 0) |
+			  FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_EN_MASK, 1);
+	nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR,
+			      cfg_tbl.data, sizeof(cfg_tbl));
+}
+
+static void nbl_hw_stop_mailbox_rxq(struct nbl_hw_mgt *hw_mgt)
+{
+	struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+	memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+	cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+	nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR,
+			      cfg_tbl.data, sizeof(cfg_tbl));
+	/* Ensure QUEUE_RST has reached the device before caller proceeds */
+	nbl_hw_flush_mbx_write(hw_mgt,
+			       NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR);
+}
+
+static void nbl_hw_stop_mailbox_txq(struct nbl_hw_mgt *hw_mgt)
+{
+	struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+	memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+	cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+	nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR,
+			      cfg_tbl.data, sizeof(cfg_tbl));
+	/* Ensure QUEUE_RST has reached the device before caller proceeds */
+	nbl_hw_flush_mbx_write(hw_mgt,
+			       NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR);
+}
+
+static void nbl_hw_get_host_pf_mask(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask)
+{
+	nbl_hw_rd_regs_lock(hw_mgt, NBL_PCIE_HOST_K_PF_MASK_REG, pf_mask,
+			    sizeof(*pf_mask));
+}
+
+static void nbl_hw_cfg_mailbox_qinfo(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+				     u8 bus, u8 devid, u8 function)
+{
+	u32 data = 0;
+
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id),
+		       &data, sizeof(data));
+	data &= ~(NBL_MAILBOX_QINFO_MAP_FUNCTION_MASK |
+		  NBL_MAILBOX_QINFO_MAP_DEVID_MASK |
+		  NBL_MAILBOX_QINFO_MAP_BUS_MASK |
+		  NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK |
+		  NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK);
+	data |= FIELD_PREP(NBL_MAILBOX_QINFO_MAP_FUNCTION_MASK, function) |
+	       FIELD_PREP(NBL_MAILBOX_QINFO_MAP_DEVID_MASK, devid) |
+	       FIELD_PREP(NBL_MAILBOX_QINFO_MAP_BUS_MASK, bus);
+	nbl_hw_wr_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id),
+		       &data, sizeof(data));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static struct nbl_hw_ops hw_ops = {
+	.update_mailbox_queue_tail_ptr = nbl_hw_update_mailbox_queue_tail_ptr,
+	.config_mailbox_rxq = nbl_hw_config_mailbox_rxq,
+	.config_mailbox_txq = nbl_hw_config_mailbox_txq,
+	.stop_mailbox_rxq = nbl_hw_stop_mailbox_rxq,
+	.stop_mailbox_txq = nbl_hw_stop_mailbox_txq,
+	.get_host_pf_mask = nbl_hw_get_host_pf_mask,
+	.cfg_mailbox_qinfo = nbl_hw_cfg_mailbox_qinfo,
+
+};
+
 /* Structure starts here, adding an op should not modify anything below */
 static struct nbl_hw_mgt *nbl_hw_setup_hw_mgt(struct nbl_common_info *common)
 {
@@ -23,6 +228,27 @@ static struct nbl_hw_mgt *nbl_hw_setup_hw_mgt(struct nbl_common_info *common)
 	return hw_mgt;
 }
 
+static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
+					       struct nbl_hw_mgt *hw_mgt)
+{
+	struct nbl_hw_ops_tbl *hw_ops_tbl;
+	struct device *dev;
+
+	dev = common->dev;
+	hw_ops_tbl = devm_kzalloc(dev, sizeof(*hw_ops_tbl), GFP_KERNEL);
+	if (!hw_ops_tbl)
+		return ERR_PTR(-ENOMEM);
+	if (!hw_ops.update_mailbox_queue_tail_ptr ||
+	    !hw_ops.config_mailbox_rxq || !hw_ops.config_mailbox_txq ||
+	    !hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
+	    !hw_ops.get_host_pf_mask || !hw_ops.cfg_mailbox_qinfo)
+		return ERR_PTR(-EINVAL);
+	hw_ops_tbl->ops = &hw_ops;
+	hw_ops_tbl->priv = hw_mgt;
+
+	return hw_ops_tbl;
+}
+
 static int nbl_pcim_request_selected_bars(struct pci_dev *pdev, u32 mask,
 					  const char *name)
 {
@@ -43,6 +269,7 @@ int nbl_hw_init_leonis(struct nbl_adapter *adapter)
 {
 	resource_size_t expect_sz = NBL_MEM_BAR_TOTAL_SIZE;
 	struct nbl_common_info *common = &adapter->common;
+	struct nbl_hw_ops_tbl *hw_ops_tbl = NULL;
 	struct pci_dev *pdev = common->pdev;
 	struct nbl_hw_mgt *hw_mgt = NULL;
 	resource_size_t bar_len;
@@ -136,7 +363,14 @@ int nbl_hw_init_leonis(struct nbl_adapter *adapter)
 	}
 
 	hw_mgt->mailbox_bar_size = bar_len;
+	spin_lock_init(&hw_mgt->reg_lock);
 
+	hw_ops_tbl = nbl_hw_setup_ops(common, hw_mgt);
+	if (IS_ERR(hw_ops_tbl)) {
+		ret = PTR_ERR(hw_ops_tbl);
+		goto setup_mgt_fail;
+	}
+	adapter->intf.hw_ops_tbl = hw_ops_tbl;
 	adapter->core.hw_mgt = hw_mgt;
 
 	return 0;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index 1f9e509dc631..93e5c7518288 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -11,5 +11,46 @@
 #include "../../nbl_include/nbl_include.h"
 #include "../nbl_hw_reg.h"
 
+/*  ----------  REG BASE ADDR  ----------  */
+/* Interface modules base addr */
+#define NBL_INTF_HOST_PCOMPLETER_BASE		0x00f08000
+#define NBL_INTF_HOST_PADPT_BASE		0x00f4c000
+#define NBL_INTF_HOST_MAILBOX_BASE		0x00fb0000
+#define NBL_INTF_HOST_PCIE_BASE			0X01504000
+/*  --------  MAILBOX BAR2 -----  */
+#define NBL_MAILBOX_NOTIFY_ADDR			0x00000000
+#define NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR	0x10
+#define NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR	0x20
+
+/*  --------  MAILBOX  --------  */
+
+/* mailbox BAR qinfo_cfg_table */
+#define MAILBOX_QINFO_CFG_TABLE_DWLEN	4
+/* data[2] */
+#define NBL_MAILBOX_QINFO_CFG_QUEUE_SIZE_BWID_MASK	GENMASK(3, 0)
+/* data[3] */
+#define NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK		BIT(0)
+#define NBL_MAILBOX_QINFO_CFG_QUEUE_EN_MASK		BIT(1)
+#define NBL_MAILBOX_QINFO_CFG_DIF_ERR_MASK		BIT(2)
+#define NBL_MAILBOX_QINFO_CFG_PTR_ERR_MASK		BIT(3)
+struct nbl_mailbox_qinfo_cfg_table {
+	u32 data[MAILBOX_QINFO_CFG_TABLE_DWLEN];
+};
+
+/*  --------  MAILBOX BAR0 -----  */
+/* mailbox qinfo_map_table */
+#define NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id) \
+	(NBL_INTF_HOST_MAILBOX_BASE + 0x00001000 + (func_id) * sizeof(u32))
+
+/* MAILBOX qinfo_map_table */
+#define NBL_MAILBOX_QINFO_MAP_FUNCTION_MASK		GENMASK(2, 0)
+#define NBL_MAILBOX_QINFO_MAP_DEVID_MASK		GENMASK(7, 3)
+#define NBL_MAILBOX_QINFO_MAP_BUS_MASK			GENMASK(15, 8)
+#define NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK		GENMASK(28, 16)
+#define NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK	BIT(29)
+
+/*  --------  HOST_PCIE  --------  */
+#define NBL_PCIE_HOST_K_PF_MASK_REG (NBL_INTF_HOST_PCIE_BASE + 0x00001004)
+
 #define NBL_BAR2_MAX_LEN		0x300
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
index e281109f502e..4c1bb789c465 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
@@ -8,6 +8,7 @@
 
 #include <linux/types.h>
 
+#include "../nbl_include/nbl_def_channel.h"
 #include "../nbl_include/nbl_def_hw.h"
 #include "../nbl_include/nbl_def_common.h"
 #include "../nbl_core.h"
@@ -26,6 +27,37 @@ struct nbl_hw_mgt {
 	u8 __iomem *hw_addr;
 	u8 __iomem *mailbox_bar_hw_addr;
 	resource_size_t mailbox_bar_size;
+	spinlock_t reg_lock; /* Protect reg access */
 };
 
+static inline u32 rd32(u8 __iomem *addr, u64 reg)
+{
+	return readl(addr + reg);
+}
+
+static inline void wr32(u8 __iomem *addr, u64 reg, u32 value)
+{
+	writel(value, addr + reg);
+}
+
+static inline void nbl_hw_wr32(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 value)
+{
+	wr32(hw_mgt->hw_addr, reg, value);
+}
+
+static inline u32 nbl_hw_rd32(struct nbl_hw_mgt *hw_mgt, u64 reg)
+{
+	return rd32(hw_mgt->hw_addr, reg);
+}
+
+static inline void nbl_mbx_wr32(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 value)
+{
+	writel(value, hw_mgt->mailbox_bar_hw_addr + reg);
+}
+
+static inline u32 nbl_mbx_rd32(struct nbl_hw_mgt *hw_mgt, u64 reg)
+{
+	return readl(hw_mgt->mailbox_bar_hw_addr + reg);
+}
+
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
new file mode 100644
index 000000000000..d6faa0bc4026
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
@@ -0,0 +1,125 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_CHANNEL_H_
+#define _NBL_DEF_CHANNEL_H_
+
+#include <linux/types.h>
+
+struct nbl_channel_mgt;
+struct nbl_adapter;
+
+typedef void (*nbl_chan_resp)(void *, u16, u16, void *, u32);
+
+/*
+ * Mailbox wire opcodes, stable wire ABI shared between driver and firmware.
+ * Each opcode has a fixed assigned number to preserve compatibility.
+ * ABI compatibility rules:
+ * 1. New opcodes shall only be appended before NBL_CHAN_MSG_MAILBOX_MAX;
+ * 2. Reordering, inserting or deleting existing enumerators breaks driver-
+ *    firmware interoperability and must be avoided;
+ * 3. Modifications to existing opcodes require synchronized firmware ABI
+ *    updates.
+ *
+ * Only opcodes currently used by in-tree driver logic are defined here.
+ * Unimplemented feature opcodes (KTLS, IPsec, vDPA, mirror etc.) will be
+ * added incrementally together with their corresponding driver
+ * implementation patches.
+ */
+enum nbl_chan_msg_type {
+	NBL_CHAN_MSG_ACK = 0,
+	/* mailbox msg end */
+	NBL_CHAN_MSG_MAILBOX_MAX,
+};
+
+enum nbl_chan_state {
+	NBL_CHAN_IRQ_RDY,
+	NBL_CHAN_STATE_NBITS
+};
+
+struct nbl_chan_send_info {
+	void *arg;
+	size_t arg_len;
+	void *resp;
+	size_t resp_len;
+	u16 dstid;
+	u16 msg_type;
+	u16 ack;
+	u16 ack_len;
+};
+
+struct nbl_chan_ack_info {
+	void *data;
+	int err;
+	u32 data_len;
+	u16 dstid;
+	u16 msg_type;
+	u16 msgid;
+};
+
+enum nbl_channel_type {
+	NBL_CHAN_TYPE_MAILBOX,
+	NBL_CHAN_TYPE_MAX
+};
+
+static inline void
+nbl_chan_fill_send_info(struct nbl_chan_send_info *info,
+			u16 dst_id, u16 msg_type,
+			void *argument, u32 arg_length,
+			void *response, u32 resp_length,
+			bool need_ack)
+{
+	info->dstid = dst_id;
+	info->msg_type = msg_type;
+	info->arg = argument;
+	info->arg_len = arg_length;
+	info->resp = response;
+	info->resp_len = resp_length;
+	info->ack = need_ack;
+}
+
+static inline void
+nbl_chan_fill_ack_info(struct nbl_chan_ack_info *info,
+		       u16 dst_id, u16 msg_type, u16 msg_id,
+		       int err_code, void *ack_data, u32 data_length)
+{
+	info->dstid = dst_id;
+	info->msg_type = msg_type;
+	info->msgid = msg_id;
+	info->err = err_code;
+	info->data = ack_data;
+	info->data_len = data_length;
+}
+
+struct nbl_channel_ops {
+	int (*send_msg)(struct nbl_channel_mgt *chan_mgt,
+			struct nbl_chan_send_info *chan_send);
+	int (*send_ack)(struct nbl_channel_mgt *chan_mgt,
+			struct nbl_chan_ack_info *chan_ack);
+	int (*register_msg)(struct nbl_channel_mgt *chan_mgt, u16 msg_type,
+			    nbl_chan_resp func, void *callback_priv);
+	void (*cfg_chan_qinfo_map_table)(struct nbl_channel_mgt *chan_mgt,
+					 u8 bus, u8 devid);
+	bool (*check_queue_exist)(struct nbl_channel_mgt *chan_mgt,
+				  u8 chan_type);
+	int (*setup_queue)(struct nbl_channel_mgt *chan_mgt, u8 chan_type);
+	int (*teardown_queue)(struct nbl_channel_mgt *chan_mgt, u8 chan_type);
+	void (*clean_queue_subtask)(struct nbl_channel_mgt *chan_mgt,
+				    u8 chan_type);
+	void (*register_chan_task)(struct nbl_channel_mgt *chan_mgt,
+				   u8 chan_type, struct work_struct *task);
+	void (*set_queue_state)(struct nbl_channel_mgt *chan_mgt,
+				enum nbl_chan_state state, u8 chan_type,
+				u8 set);
+};
+
+struct nbl_channel_ops_tbl {
+	struct nbl_channel_ops *ops;
+	struct nbl_channel_mgt *priv;
+};
+
+int nbl_chan_init_common(struct nbl_adapter *adapter);
+void nbl_chan_remove_common(struct nbl_adapter *adapter);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
index ea646d1efc2d..26a364e3a194 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
@@ -12,6 +12,7 @@
 #include "nbl_include.h"
 
 struct nbl_common_info {
+	struct workqueue_struct *wq;
 	struct pci_dev *pdev;
 	struct device *dev;
 	u16 vsi_id;
@@ -28,4 +29,7 @@ struct nbl_common_info {
 	u8 has_net;
 };
 
+void nbl_common_destroy_wq(struct nbl_common_info *common);
+int nbl_common_create_wq(struct nbl_common_info *common);
+
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index ecbf440e4366..d87bf9d41a24 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -10,6 +10,41 @@
 
 struct nbl_hw_mgt;
 struct nbl_adapter;
+struct nbl_hw_ops {
+	void (*update_mailbox_queue_tail_ptr)(struct nbl_hw_mgt *hw_mgt,
+					      u16 tail_ptr, u8 txrx);
+	void (*config_mailbox_rxq)(struct nbl_hw_mgt *hw_mgt,
+				   dma_addr_t dma_addr, int size_bwid);
+	void (*config_mailbox_txq)(struct nbl_hw_mgt *hw_mgt,
+				   dma_addr_t dma_addr, int size_bwid);
+	void (*stop_mailbox_rxq)(struct nbl_hw_mgt *hw_mgt);
+	void (*stop_mailbox_txq)(struct nbl_hw_mgt *hw_mgt);
+	/**
+	 * get_host_pf_mask - Fetch host PF mask from firmware k_pf_mask reg
+	 * @hw_mgt: hardware management context
+	 * @pf_mask: output pointer for PF mask value
+	 *
+	 * k_pf_mask register rule:
+	 *   bit N == 0 -> PF#N enabled; bit N == 1 -> PF#N masked out.
+	 *   bit0 is PF0's mask bit (not reserved); PF0 can be masked but
+	 *   the driver requires at least PF0 enabled.
+	 *   Only 1/2/4 PFs are supported:
+	 *     1 PF  (PF0):     mask = 0xfe
+	 *     2 PFs (PF0,PF1): mask = 0xfc
+	 *     4 PFs (PF0~PF3): mask = 0xf0
+	 *   All-zero mask (0x00) means all 8 PFs enabled, which is
+	 *   unsupported by the driver and rejected with -EINVAL.
+	 */
+	void (*get_host_pf_mask)(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask);
+
+	void (*cfg_mailbox_qinfo)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+				  u8 bus, u8 devid, u8 function);
+};
+
+struct nbl_hw_ops_tbl {
+	struct nbl_hw_ops *ops;
+	struct nbl_hw_mgt *priv;
+};
 
 int nbl_hw_init_leonis(struct nbl_adapter *adapter);
 void nbl_hw_remove_leonis(struct nbl_adapter *adapter);
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
index 14e7b19f9a4c..f2d802397d98 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
@@ -10,6 +10,9 @@
 
 /*  ------  Basic definitions  -------  */
 #define NBL_DRIVER_NAME					"nbl"
+#define NBL_MAX_PF					8
+#define NBL_NEXT_ID(id, max) (((id) + 1) % ((max) + 1))
+
 struct nbl_func_caps {
 	u32 has_ctrl:1;
 	u32 has_net:1;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index f2552bc73293..b7c80ea54c8d 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -8,6 +8,7 @@
 #include <linux/module.h>
 #include <linux/bits.h>
 #include "nbl_include/nbl_include.h"
+#include "nbl_include/nbl_def_channel.h"
 #include "nbl_include/nbl_def_hw.h"
 #include "nbl_include/nbl_def_common.h"
 #include "nbl_core.h"
@@ -38,13 +39,19 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 	if (ret)
 		goto hw_init_fail;
 
+	ret = nbl_chan_init_common(adapter);
+	if (ret)
+		goto chan_init_fail;
 	return adapter;
+chan_init_fail:
+	nbl_hw_remove_leonis(adapter);
 hw_init_fail:
 	return ERR_PTR(ret);
 }
 
 void nbl_core_remove(struct nbl_adapter *adapter)
 {
+	nbl_chan_remove_common(adapter);
 	nbl_hw_remove_leonis(adapter);
 }
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 2/8] net/nebula-matrix: add common resource implementation
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
  2026-09-28 12:32 ` [PATCH v30 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  2026-09-28 12:32 ` [PATCH v30 net-next 3/8] net/nebula-matrix: add intr " illusion.wang
                   ` (5 subsequent siblings)
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Add chip-agnostic resource layer to manage PF topology, SR-IOV BDF, Ethernet
port and VSI identity mappings, providing reusable conversion helpers for
the Nebula Matrix driver control plane.
This patch implements core resource initialization and lookup logic,
running exclusively on the control PF and building read-only resource
tables once during probe:
- nbl_res_init_pf_num(): parse firmware PF mask and validate supported
  topologies. Only contiguous 1/2/4 PFs starting from PF0 are permitted,
  rejecting invalid/sparse configurations with early probe failure.
- nbl_common_func_id_to_rel_pf_id(): convert absolute PF function ID
  to relative PF index for consistent resource table indexing.
- nbl_res_ctrl_dev_sriov_info_init(): calculate and store per-PF BDF
  entries based on hardware PCI bus number for MSI-X programming.
- nbl_res_ctrl_dev_setup_eth_info(): verify firmware port count and
  Ethernet bitmap consistency, then construct per-PF eth_id lookup table.
  logic_eth_id is computed on-the-fly from relative PF ID instead of
  table lookup.
- nbl_res_ctrl_dev_vsi_info_init(): assign staggered per-PF VSI base
  IDs with gaps of 1024, 512 and 256 for 1, 2 and 4 port modes.
- Implement VSI/PF/Eth ID conversion helpers with strict control PF
  validity and range guards.
All resource tables are initialized once at control PF probe time and
remain read-only afterwards, so lookup helpers require no internal locking.
Upper-layer serialization via mutex and non-control PF mailbox RPC routing
will be introduced in later series patches.
All resource memory allocations use devm managed semantics, requiring
no explicit cleanup in the remove path.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../net/ethernet/nebula-matrix/nbl/Makefile   |   2 +
 .../nebula-matrix/nbl/nbl_common/nbl_common.c |  21 ++
 .../net/ethernet/nebula-matrix/nbl/nbl_core.h |   2 +
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c  |  67 +++-
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h  |  15 +
 .../nbl_hw_leonis/nbl_resource_leonis.c       | 329 ++++++++++++++++++
 .../nbl_hw_leonis/nbl_resource_leonis.h       |  10 +
 .../nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h     |   1 +
 .../nebula-matrix/nbl/nbl_hw/nbl_resource.c   | 125 +++++++
 .../nebula-matrix/nbl/nbl_hw/nbl_resource.h   |  69 ++++
 .../nbl/nbl_include/nbl_def_channel.h         |  10 +
 .../nbl/nbl_include/nbl_def_common.h          |  21 +-
 .../nbl/nbl_include/nbl_def_hw.h              |  15 +
 .../nbl/nbl_include/nbl_def_resource.h        |  29 ++
 .../nbl/nbl_include/nbl_include.h             |   6 +
 .../net/ethernet/nebula-matrix/nbl/nbl_main.c |   9 +
 16 files changed, 728 insertions(+), 3 deletions(-)
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index 04e1aa1fb4bd..3dab9519a277 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -6,4 +6,6 @@ obj-$(CONFIG_NBL) := nbl.o
 nbl-objs +=	nbl_common/nbl_common.o \
 		nbl_channel/nbl_channel.o \
 		nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
+		nbl_hw/nbl_hw_leonis/nbl_resource_leonis.o \
+		nbl_hw/nbl_resource.o \
 		nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
index ce0ed2869af4..302c3c48dc4d 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
@@ -29,3 +29,24 @@ int nbl_common_create_wq(struct nbl_common_info *common)
 	return 0;
 }
 
+/**
+ * nbl_common_func_id_to_rel_pf_id - convert absolute PF id to relative PF id
+ * @common: common device info
+ * @pf_id: absolute PF identifier
+ * @rel_pf_id: output relative pf id
+ *
+ * Leonis uses fixed mgt_pf = 0. Support future non-zero management PF.
+ *
+ * Return: 0 on success, -EINVAL on invalid arguments.
+ */
+int nbl_common_func_id_to_rel_pf_id(struct nbl_common_info *common, u32 pf_id,
+				    u32 *rel_pf_id)
+{
+	if (!rel_pf_id)
+		return -EINVAL;
+
+	if (pf_id < common->mgt_pf)
+		return -EINVAL;
+	*rel_pf_id = pf_id - common->mgt_pf;
+	return 0;
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index f998a2b44e5c..dd24ebec0171 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -16,11 +16,13 @@ enum {
 
 struct nbl_interface {
 	struct nbl_hw_ops_tbl *hw_ops_tbl;
+	struct nbl_resource_ops_tbl *resource_ops_tbl;
 	struct nbl_channel_ops_tbl *channel_ops_tbl;
 };
 
 struct nbl_core {
 	struct nbl_hw_mgt *hw_mgt;
+	struct nbl_resource_mgt *res_mgt;
 	struct nbl_channel_mgt *chan_mgt;
 };
 
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index eeff6216e4aa..4d3477f70bcc 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -97,6 +97,32 @@ static void nbl_hw_rd_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
 	spin_unlock(&hw_mgt->reg_lock);
 }
 
+/*
+ * Only call this when has_ctrl=true, which maps enough space
+ * (bar_len - 8192) to cover NBL_HW_DUMMY_REG (0x1300904).
+ * The flow/design guarantees this is only called in the
+ * has_ctrl path.
+ */
+static void nbl_flush_writes(struct nbl_hw_mgt *hw_mgt)
+{
+	nbl_hw_rd32(hw_mgt, NBL_HW_DUMMY_REG);
+}
+
+/*
+ * Registers reset to zero after cold boot / FLR / bus reset. Firmware
+ * programs valid values before driver probe, so zero is only seen on
+ * hardware fault or register read failure. Initialize data=0 to guard
+ * against nbl_hw_read_mbx_regs() early-return on bounds-check failure.
+ */
+static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)
+{
+	u32 data = 0;
+
+	nbl_hw_read_mbx_regs(hw_mgt, NBL_FW_BOARD_DW6_OFFSET, &data,
+			     sizeof(data));
+	*eth_map = FIELD_GET(NBL_FW_BOARD_DW6_ETH_BITMAP_MASK, data);
+}
+
 static void nbl_hw_update_mailbox_queue_tail_ptr(struct nbl_hw_mgt *hw_mgt,
 						 u16 tail_ptr, u8 txrx)
 {
@@ -181,6 +207,15 @@ static void nbl_hw_get_host_pf_mask(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask)
 			    sizeof(*pf_mask));
 }
 
+static void nbl_hw_get_real_bus(struct nbl_hw_mgt *hw_mgt, u8 *bus)
+{
+	u32 data = 0;
+
+	nbl_hw_rd_regs_lock(hw_mgt, NBL_PCIE_HOST_TL_CFG_BUSDEV, &data,
+			    sizeof(data));
+	*bus = FIELD_GET(NBL_PCIE_BUS_MASK, data);
+}
+
 static void nbl_hw_cfg_mailbox_qinfo(struct nbl_hw_mgt *hw_mgt, u16 func_id,
 				     u8 bus, u8 devid, u8 function)
 {
@@ -202,15 +237,41 @@ static void nbl_hw_cfg_mailbox_qinfo(struct nbl_hw_mgt *hw_mgt, u16 func_id,
 	spin_unlock(&hw_mgt->reg_lock);
 }
 
+/*
+ * Registers reset to zero after cold boot / FLR / bus reset. Firmware
+ * programs valid values before driver probe, so zero is only seen on
+ * hardware fault or register read failure. Initialize data=0 to guard
+ * against nbl_hw_read_mbx_regs() early-return on bounds-check failure.
+ */
+static void nbl_hw_get_board_info(struct nbl_hw_mgt *hw_mgt,
+				  struct nbl_board_port_info *board_info)
+{
+	u32 data = 0;
+
+	nbl_hw_read_mbx_regs(hw_mgt, NBL_FW_BOARD_DW3_OFFSET, &data,
+			     sizeof(data));
+	board_info->eth_num = FIELD_GET(NBL_FW_BOARD_DW3_PORT_NUM_MASK, data);
+	board_info->eth_speed =
+		FIELD_GET(NBL_FW_BOARD_DW3_PORT_SPEED_MASK, data);
+	board_info->p4_version =
+		FIELD_GET(NBL_FW_BOARD_DW3_P4_VERSION_MASK, data);
+}
+
 static struct nbl_hw_ops hw_ops = {
+	.flush_write = nbl_flush_writes,
+
 	.update_mailbox_queue_tail_ptr = nbl_hw_update_mailbox_queue_tail_ptr,
 	.config_mailbox_rxq = nbl_hw_config_mailbox_rxq,
 	.config_mailbox_txq = nbl_hw_config_mailbox_txq,
 	.stop_mailbox_rxq = nbl_hw_stop_mailbox_rxq,
 	.stop_mailbox_txq = nbl_hw_stop_mailbox_txq,
 	.get_host_pf_mask = nbl_hw_get_host_pf_mask,
+	.get_real_bus = nbl_hw_get_real_bus,
+
 	.cfg_mailbox_qinfo = nbl_hw_cfg_mailbox_qinfo,
 
+	.get_fw_eth_map = nbl_hw_get_fw_eth_map,
+	.get_board_info = nbl_hw_get_board_info,
 };
 
 /* Structure starts here, adding an op should not modify anything below */
@@ -238,10 +299,12 @@ static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
 	hw_ops_tbl = devm_kzalloc(dev, sizeof(*hw_ops_tbl), GFP_KERNEL);
 	if (!hw_ops_tbl)
 		return ERR_PTR(-ENOMEM);
-	if (!hw_ops.update_mailbox_queue_tail_ptr ||
+	if (!hw_ops.flush_write || !hw_ops.update_mailbox_queue_tail_ptr ||
 	    !hw_ops.config_mailbox_rxq || !hw_ops.config_mailbox_txq ||
 	    !hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
-	    !hw_ops.get_host_pf_mask || !hw_ops.cfg_mailbox_qinfo)
+	    !hw_ops.get_host_pf_mask || !hw_ops.get_real_bus ||
+	    !hw_ops.cfg_mailbox_qinfo ||
+	    !hw_ops.get_fw_eth_map || !hw_ops.get_board_info)
 		return ERR_PTR(-EINVAL);
 	hw_ops_tbl->ops = &hw_ops;
 	hw_ops_tbl->priv = hw_mgt;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index 93e5c7518288..251dd68d0721 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -51,6 +51,21 @@ struct nbl_mailbox_qinfo_cfg_table {
 
 /*  --------  HOST_PCIE  --------  */
 #define NBL_PCIE_HOST_K_PF_MASK_REG (NBL_INTF_HOST_PCIE_BASE + 0x00001004)
+#define NBL_PCIE_HOST_TL_CFG_BUSDEV (NBL_INTF_HOST_PCIE_BASE + 0x11040)
+
+#define NBL_PCIE_BUS_MASK	GENMASK(12, 5)
+#define NBL_FW_BOARD_CONFIG			0x200
+#define NBL_FW_BOARD_DW3_OFFSET			(NBL_FW_BOARD_CONFIG + 12)
+#define NBL_FW_BOARD_DW6_OFFSET			(NBL_FW_BOARD_CONFIG + 24)
+
+#define NBL_FW_BOARD_DW3_PORT_TYPE_MASK BIT(0)
+#define NBL_FW_BOARD_DW3_PORT_NUM_MASK GENMASK(7, 1)
+#define NBL_FW_BOARD_DW3_PORT_SPEED_MASK GENMASK(9, 8)
+#define NBL_FW_BOARD_DW3_GPIO_TYPE_MASK GENMASK(12, 10)
+#define NBL_FW_BOARD_DW3_P4_VERSION_MASK GENMASK(13, 13)
+
+#define NBL_FW_BOARD_DW6_LANE_BITMAP_MASK GENMASK(7, 0)
+#define NBL_FW_BOARD_DW6_ETH_BITMAP_MASK GENMASK(15, 8)
 
 #define NBL_BAR2_MAX_LEN		0x300
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
new file mode 100644
index 000000000000..7804762a96e0
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
@@ -0,0 +1,329 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/pci.h>
+#include <linux/bits.h>
+#include "nbl_resource_leonis.h"
+
+static struct nbl_resource_ops res_ops = {
+	.get_vsi_id = nbl_res_func_id_to_vsi_id,
+	.get_eth_id = nbl_res_get_eth_id,
+};
+
+static struct nbl_resource_mgt *
+nbl_res_setup_res_mgt(struct nbl_common_info *common)
+{
+	struct nbl_resource_info *resource_info;
+	struct nbl_resource_mgt *res_mgt;
+	struct device *dev = common->dev;
+
+	res_mgt = devm_kzalloc(dev, sizeof(*res_mgt), GFP_KERNEL);
+	if (!res_mgt)
+		return ERR_PTR(-ENOMEM);
+	res_mgt->common = common;
+
+	resource_info =
+		devm_kzalloc(dev, sizeof(*resource_info), GFP_KERNEL);
+	if (!resource_info)
+		return ERR_PTR(-ENOMEM);
+	res_mgt->resource_info = resource_info;
+
+	return res_mgt;
+}
+
+static struct nbl_resource_ops_tbl *
+nbl_res_setup_ops(struct device *dev, struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_resource_ops_tbl *res_ops_tbl;
+
+	res_ops_tbl = devm_kzalloc(dev, sizeof(*res_ops_tbl), GFP_KERNEL);
+	if (!res_ops_tbl)
+		return ERR_PTR(-ENOMEM);
+	if (!res_ops.get_vsi_id || !res_ops.get_eth_id)
+		return ERR_PTR(-EINVAL);
+	res_ops_tbl->ops = &res_ops;
+	res_ops_tbl->priv = res_mgt;
+
+	return res_ops_tbl;
+}
+
+static int nbl_res_ctrl_dev_setup_eth_info(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	struct device *dev = res_mgt->common->dev;
+	struct nbl_eth_info *eth_info;
+	u32 eth_bitmap = 0;
+	u32 eth_num = 0;
+	u32 fw_port_num;
+	int i;
+
+	eth_info = devm_kzalloc(dev, sizeof(*eth_info), GFP_KERNEL);
+	if (!eth_info)
+		return -ENOMEM;
+
+	res_mgt->resource_info->eth_info = eth_info;
+
+	fw_port_num = res_mgt->resource_info->board_info.eth_num;
+
+	hw_ops->get_fw_eth_map(res_mgt->hw_ops_tbl->priv, &eth_bitmap);
+	if (eth_bitmap & ~((1 << NBL_MAX_ETHERNET) - 1)) {
+		dev_err(dev, "FW reported invalid eth_bitmap 0x%x\n",
+			eth_bitmap);
+		return -EINVAL;
+	}
+	if (fw_port_num != hweight32(eth_bitmap)) {
+		dev_err(dev, "FW inconsistency: port_num=%u, bitmap=0x%x\n",
+			fw_port_num, eth_bitmap);
+		return -EINVAL;
+	}
+	/*
+	 * Firmware is ready before probe. Valid port counts are 1/2/4;
+	 * 0 (invalid config), 3 (unsupported topology), and >4 (exceeds
+	 * hardware max) are all rejected with -EINVAL.
+	 */
+	if (fw_port_num == 0 || fw_port_num == 3 ||
+	    fw_port_num > NBL_MAX_ETHERNET) {
+		dev_err(dev, "FW reports %u Ethernet ports, unsupported (valid: 1/2/4)\n",
+			fw_port_num);
+		return -EINVAL;
+	}
+	eth_info->eth_num = fw_port_num;
+	/* Intentional design constraint: each PF maps to exactly one
+	 * Ethernet port. This couples PF identity to port identity
+	 * and is required by nbl_res_get_eth_id() which indexes
+	 * eth_info->eth_id[] by relative PF id.
+	 */
+	if (res_mgt->common->max_pf != eth_info->eth_num) {
+		dev_err(dev, "Invalid PF-to-port topology: max_pf=%u, eth_num=%u\n",
+			res_mgt->common->max_pf, eth_info->eth_num);
+		return -EINVAL;
+	}
+
+	/*
+	 * Any subset of valid bitmap bits is accepted (e.g. 0/1, 0/2,
+	 * 1/3, etc.).  Firmware only needs to report the correct count
+	 * of active ports; no hard-coded fixed bit positions required.
+	 * eth_id[] is filled in ascending bitmap-bit order, so the Nth
+	 * relative PF owns the Nth logical port (max_pf == eth_num is
+	 * enforced above).
+	 */
+	for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+		if ((1 << i) & eth_bitmap) {
+			eth_info->eth_id[eth_num] = i;
+			eth_num++;
+		}
+	}
+
+	return 0;
+}
+
+static int nbl_res_ctrl_dev_sriov_info_init(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	struct nbl_hw_mgt *p = res_mgt->hw_ops_tbl->priv;
+	struct nbl_common_info *common = res_mgt->common;
+	struct nbl_sriov_info *sriov_info;
+	struct device *dev = common->dev;
+	u8 hw_bus = 0;
+	u16 function;
+	u16 func_id;
+
+	hw_ops->get_real_bus(p, &hw_bus);
+	if (common->function + common->max_pf > NBL_MAX_PF) {
+		dev_err(dev, "PF count exceeds available function space\n");
+		return -EINVAL;
+	}
+	sriov_info = devm_kcalloc(dev, common->max_pf,
+				  sizeof(*sriov_info), GFP_KERNEL);
+	if (!sriov_info)
+		return -ENOMEM;
+
+	res_mgt->resource_info->sriov_info = sriov_info;
+	/*
+	 * Real bus number of the control PF;
+	 * the BDF table below is built from it.
+	 */
+	common->hw_bus = hw_bus;
+
+	for (func_id = 0; func_id < common->max_pf; func_id++) {
+		sriov_info = res_mgt->resource_info->sriov_info + func_id;
+		function = common->function + func_id;
+		sriov_info->bdf = PCI_DEVID(common->hw_bus,
+					    PCI_DEVFN(common->devid, function));
+	}
+
+	return 0;
+}
+
+static int nbl_res_ctrl_dev_vsi_info_init(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_eth_info *eth_info = res_mgt->resource_info->eth_info;
+	struct nbl_common_info *common = res_mgt->common;
+	struct device *dev = common->dev;
+	struct nbl_vsi_info *vsi_info;
+	int i;
+
+	vsi_info = devm_kzalloc(dev, sizeof(*vsi_info), GFP_KERNEL);
+	if (!vsi_info)
+		return -ENOMEM;
+
+	res_mgt->resource_info->vsi_info = vsi_info;
+	/*
+	 * case 1 one port(1pf)
+	 * pf0 (NBL_VSI_SERV_PF_DATA_TYPE) vsi is 0
+	 * case 2 two port(2pf)
+	 * pf0,pf1(NBL_VSI_SERV_PF_DATA_TYPE) vsi is 0,512
+	 * case 3 four port(4pf)
+	 * pf0,pf1,pf2,pf3(NBL_VSI_SERV_PF_DATA_TYPE) vsi is 0,256,512,768
+	 */
+
+	vsi_info->num = eth_info->eth_num;
+	/*
+	 * eth_num can be 1/2/4:
+	 * - 2/4 ports use dedicated gap constants;
+	 * - 1 port falls back to NBL_DEFAULT_VSI_ID_GAP (1024).
+	 * All three values produce valid base_id offsets.
+	 */
+	for (i = 0; i < vsi_info->num; i++) {
+		vsi_info->serv_info[i][NBL_VSI_SERV_PF_DATA_TYPE].base_id =
+			i * nbl_vsi_id_gap(vsi_info->num);
+		vsi_info->serv_info[i][NBL_VSI_SERV_PF_DATA_TYPE].num = 1;
+	}
+
+	return 0;
+}
+
+static int nbl_res_init_pf_num(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	u32 exp_contiguous_mask = 0;
+	u32 pf_mask = 0;
+	u32 pf_num = 0;
+	int i;
+
+	hw_ops->get_host_pf_mask(res_mgt->hw_ops_tbl->priv, &pf_mask);
+
+	/*
+	 * k_pf_mask register rule:
+	 * bit N == 0  -> PF#N enabled; bit N == 1 -> PF#N masked out.
+	 * Hardware constraint: bit0 is PF0's mask bit; driver requires
+	 * PF0 enabled as management PF, so bit0 must be clear.
+	 * All-zero pf_mask means all PF0~PF7 are enabled, which is unsupported
+	 * by the driver
+	 *
+	 * Product firmware constraint: only 3 valid configurations supported:
+	 * 1 PF  (PF0 only): pf_num = 1, mask = 0xfe
+	 * 2 PFs (PF0,PF1):  pf_num = 2, mask = 0xfc
+	 * 4 PFs (PF0~PF3): pf_num = 4, mask = 0xf0
+	 * No other PF count or sparse/non-contiguous PF layout is allowed.
+	 */
+	for (i = 0; i < NBL_MAX_PF; i++) {
+		if (!(pf_mask & (1 << i)))
+			pf_num++;
+	}
+
+	/*
+	 * Sanity check: enabled PFs must be contiguous starting from PF0.
+	 * Current resource framework uses relative PF id, sparse PF layout
+	 * will cause mismatch between resource layer and hardware func_id.
+	 */
+	for (i = 0; i < pf_num; i++)
+		exp_contiguous_mask |= BIT(i);
+	if ((pf_mask & exp_contiguous_mask) != 0) {
+		dev_err(res_mgt->common->dev,
+			"pf_mask 0x%08x: non-contiguous enabled PF, unsupported\n",
+			pf_mask);
+		return -EINVAL;
+	}
+
+	/* Only allow product-specified PF count: 1 / 2 / 4 */
+	if (pf_num != 1 && pf_num != 2 && pf_num != 4) {
+		dev_err(res_mgt->common->dev,
+			"Invalid pf_num=%u (mask=0x%08x), only 1/2/4 PFs supported\n",
+			pf_num, pf_mask);
+		return -EINVAL;
+	}
+
+	res_mgt->common->max_pf = pf_num;
+
+	return 0;
+}
+
+static void nbl_res_init_board_info(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+
+	hw_ops->get_board_info(res_mgt->hw_ops_tbl->priv,
+			       &res_mgt->resource_info->board_info);
+}
+
+static int nbl_res_start(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_common_info *common = res_mgt->common;
+	int ret = 0;
+
+	if (common->has_ctrl) {
+		nbl_res_init_board_info(res_mgt);
+
+		ret = nbl_res_init_pf_num(res_mgt);
+		if (ret)
+			return ret;
+
+		ret = nbl_res_ctrl_dev_sriov_info_init(res_mgt);
+		if (ret)
+			return ret;
+
+		ret = nbl_res_ctrl_dev_setup_eth_info(res_mgt);
+		if (ret)
+			return ret;
+
+		ret = nbl_res_ctrl_dev_vsi_info_init(res_mgt);
+		if (ret)
+			return ret;
+	}
+
+	return 0;
+}
+
+int nbl_res_init_leonis(struct nbl_adapter *adap)
+{
+	struct nbl_channel_ops_tbl *chan_ops_tbl = adap->intf.channel_ops_tbl;
+	struct nbl_hw_ops_tbl *hw_ops_tbl = adap->intf.hw_ops_tbl;
+	struct nbl_common_info *common = &adap->common;
+	struct nbl_resource_ops_tbl *res_ops_tbl;
+	struct device *dev = &adap->pdev->dev;
+	struct nbl_resource_mgt *res_mgt;
+	int ret;
+
+	res_mgt = nbl_res_setup_res_mgt(common);
+	if (IS_ERR(res_mgt)) {
+		ret = PTR_ERR(res_mgt);
+		return ret;
+	}
+	res_mgt->chan_ops_tbl = chan_ops_tbl;
+	res_mgt->hw_ops_tbl = hw_ops_tbl;
+
+	ret = nbl_res_start(res_mgt);
+	if (ret)
+		return ret;
+
+	res_ops_tbl = nbl_res_setup_ops(dev, res_mgt);
+	if (IS_ERR(res_ops_tbl)) {
+		ret = PTR_ERR(res_ops_tbl);
+		return ret;
+	}
+	adap->intf.resource_ops_tbl = res_ops_tbl;
+	adap->core.res_mgt = res_mgt;
+
+	return 0;
+}
+
+void nbl_res_remove_leonis(struct nbl_adapter *adap)
+{
+	/*
+	 * No resource release here because all memory uses devm managed
+	 * allocation
+	 */
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
new file mode 100644
index 000000000000..b9355262c00d
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
@@ -0,0 +1,10 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_RESOURCE_LEONIS_H_
+#define _NBL_RESOURCE_LEONIS_H_
+
+#include "../nbl_resource.h"
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
index 4c1bb789c465..00fa33eeabaa 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
@@ -17,6 +17,7 @@
 #define NBL_MAILBOX_BAR				2
 #define NBL_RDMA_NOTIFY_LEN			(8ULL << 10)
 #define NBL_REG_NET_ONLY_LEN			(8ULL << 10)
+#define NBL_HW_DUMMY_REG			0x1300904
 /*
  * PCI MEMORY BAR total size: 64MiB.
  */
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
new file mode 100644
index 000000000000..b316fb8e7051
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
@@ -0,0 +1,125 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#include <linux/pci.h>
+#include "nbl_resource.h"
+
+int nbl_res_func_id_to_vsi_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			      u16 type, u16 *vsi_id)
+{
+	struct nbl_vsi_info *vsi_info = res_mgt->resource_info->vsi_info;
+	enum nbl_vsi_serv_type dst_type = NBL_VSI_SERV_PF_DATA_TYPE;
+	struct nbl_common_info *common = res_mgt->common;
+	struct device *dev = res_mgt->common->dev;
+	int pfid = func_id;
+	u32 rel_pf_id;
+	int ret;
+
+	if (!common->has_ctrl || !vsi_id) {
+		dev_dbg(dev, "No control plane or null vsi output ptr\n");
+		return -EINVAL;
+	}
+	ret = nbl_common_func_id_to_rel_pf_id(common, pfid, &rel_pf_id);
+	if (ret)
+		return ret;
+	if (rel_pf_id >= vsi_info->num) {
+		dev_err(dev, "PF %d (diff=%u) exceeds vsi_info->num (%u)\n",
+			pfid, rel_pf_id, vsi_info->num);
+		return -EINVAL;
+	}
+
+	ret = nbl_res_pf_dev_vsi_type_to_hw_vsi_type(res_mgt, type, &dst_type);
+	if (ret) {
+		dev_err(dev, "Invalid vsi type %u func_id %u\n", type, func_id);
+		return ret;
+	}
+	*vsi_id = vsi_info->serv_info[rel_pf_id][dst_type].base_id;
+	return 0;
+}
+
+int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id)
+{
+	struct nbl_vsi_info *vsi_info = res_mgt->resource_info->vsi_info;
+	struct nbl_common_info *common = res_mgt->common;
+	struct device *dev = res_mgt->common->dev;
+	int j = NBL_VSI_SERV_PF_DATA_TYPE;
+	int pf_id, i;
+
+	if (!common->has_ctrl) {
+		dev_dbg(dev, "No control plane available\n");
+		return -EINVAL;
+	}
+	for (i = 0; i < vsi_info->num; i++) {
+		if (vsi_id >= vsi_info->serv_info[i][j].base_id &&
+		    (vsi_id < vsi_info->serv_info[i][j].base_id +
+					vsi_info->serv_info[i][j].num)) {
+			pf_id = i + common->mgt_pf;
+			if (pf_id >= NBL_MAX_PF) {
+				dev_err(dev, "PF ID overflow\n");
+				return -ERANGE;
+			}
+			return pf_id;
+		}
+	}
+
+	dev_dbg(dev, "VSI ID %u not found\n", vsi_id);
+	return -ENOENT;
+}
+
+int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+		       u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id)
+{
+	struct nbl_eth_info *eth_info = res_mgt->resource_info->eth_info;
+	struct nbl_common_info *common = res_mgt->common;
+	struct device *dev = res_mgt->common->dev;
+	int pfid = func_id;
+	int rel_pf_id;
+	int abs_pf_id;
+
+	if (!common->has_ctrl || !eth_num || !eth_id || !logic_eth_id)
+		return -EINVAL;
+	abs_pf_id = nbl_res_vsi_id_to_pf_id(res_mgt, vsi_id);
+	if (abs_pf_id < 0) {
+		dev_err(dev, "Failed to get PF ID from VSI ID %u\n", vsi_id);
+		return -EINVAL;
+	}
+	if (abs_pf_id != pfid) {
+		dev_err(dev, "func_id %u does not match pf derived from vsi_id %u\n",
+			pfid, vsi_id);
+		return -EINVAL;
+	}
+	rel_pf_id = abs_pf_id - common->mgt_pf;
+
+	if (rel_pf_id >= eth_info->eth_num) {
+		dev_err(dev, "rel_pf_id %d out of range [0, %u)\n",
+			rel_pf_id, eth_info->eth_num);
+		return -ERANGE;
+	}
+
+	*eth_num = eth_info->eth_num;
+	*eth_id = eth_info->eth_id[rel_pf_id];
+	/*
+	 * Logical eth id equals the relative PF id: eth_info setup enforces
+	 * max_pf == eth_num and fills eth_id[] in ascending bitmap-bit
+	 * order, so the Nth PF owns the Nth logical port.
+	 */
+	*logic_eth_id = rel_pf_id;
+	return 0;
+}
+
+int nbl_res_pf_dev_vsi_type_to_hw_vsi_type(struct nbl_resource_mgt *res_mgt,
+					   u16 src_type,
+					   enum nbl_vsi_serv_type *dst_type)
+{
+	switch (src_type) {
+	case NBL_VSI_DATA:
+		*dst_type = NBL_VSI_SERV_PF_DATA_TYPE;
+		return 0;
+	default:
+		dev_err_once(res_mgt->common->dev,
+			     "Unsupported vsi src_type %u\n", src_type);
+		return -EINVAL;
+	}
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
new file mode 100644
index 000000000000..ae0a3d33198d
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
@@ -0,0 +1,69 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_RESOURCE_H_
+#define _NBL_RESOURCE_H_
+
+#include <linux/types.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_hw.h"
+#include "../nbl_include/nbl_def_resource.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+struct nbl_resource_mgt;
+
+/* --------- INFO ---------- */
+struct nbl_sriov_info {
+	unsigned int bdf;
+};
+
+struct nbl_eth_info {
+	u8 eth_num;
+	u8 resv[3];
+	u8 eth_id[NBL_MAX_ETHERNET];
+};
+
+enum nbl_vsi_serv_type {
+	NBL_VSI_SERV_PF_DATA_TYPE,
+	NBL_VSI_SERV_MAX_TYPE,
+};
+
+struct nbl_vsi_serv_info {
+	u16 base_id;
+	u16 num;
+};
+
+struct nbl_vsi_info {
+	u16 num;
+	struct nbl_vsi_serv_info serv_info[NBL_MAX_ETHERNET]
+					  [NBL_VSI_SERV_MAX_TYPE];
+};
+
+struct nbl_resource_info {
+	struct nbl_sriov_info *sriov_info;
+	struct nbl_eth_info *eth_info;
+	struct nbl_vsi_info *vsi_info;
+	struct nbl_board_port_info board_info;
+};
+
+struct nbl_resource_mgt {
+	struct nbl_common_info *common;
+	struct nbl_resource_info *resource_info;
+	struct nbl_channel_ops_tbl *chan_ops_tbl;
+	struct nbl_hw_ops_tbl *hw_ops_tbl;
+};
+
+int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id);
+int nbl_res_func_id_to_vsi_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			      u16 type, u16 *vsi_id);
+int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+		       u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id);
+int nbl_res_pf_dev_vsi_type_to_hw_vsi_type(struct nbl_resource_mgt *res_mgt,
+					   u16 src_type,
+					   enum nbl_vsi_serv_type *dst_type);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
index d6faa0bc4026..bf971121d2ec 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
@@ -39,6 +39,16 @@ enum nbl_chan_state {
 	NBL_CHAN_STATE_NBITS
 };
 
+struct nbl_board_port_info {
+	u8 eth_num;
+	u8 eth_speed;
+	u8 p4_version;
+	u8 rsv[5];
+};
+
+static_assert(sizeof(struct nbl_board_port_info) == 8,
+	      "nbl_board_port_info size must be 8 bytes");
+
 struct nbl_chan_send_info {
 	void *arg;
 	size_t arg_len;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
index 26a364e3a194..3e18a2a97850 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
@@ -11,6 +11,22 @@
 #include <linux/device.h>
 #include "nbl_include.h"
 
+#define NBL_TWO_ETHERNET_PORT			2
+#define NBL_FOUR_ETHERNET_PORT			4
+#define NBL_DEFAULT_VSI_ID_GAP			1024
+#define NBL_TWO_ETHERNET_VSI_ID_GAP		512
+#define NBL_FOUR_ETHERNET_VSI_ID_GAP		256
+
+static inline u32 nbl_vsi_id_gap(u32 m)
+{
+	if (m == NBL_FOUR_ETHERNET_PORT)
+		return NBL_FOUR_ETHERNET_VSI_ID_GAP;
+	else if (m == NBL_TWO_ETHERNET_PORT)
+		return NBL_TWO_ETHERNET_VSI_ID_GAP;
+
+	return NBL_DEFAULT_VSI_ID_GAP;
+}
+
 struct nbl_common_info {
 	struct workqueue_struct *wq;
 	struct pci_dev *pdev;
@@ -24,12 +40,15 @@ struct nbl_common_info {
 	u8 devid;
 	u8 bus;
 	u8 hw_bus;
+	u16 mgt_pf;
 
 	u8 has_ctrl;
 	u8 has_net;
+	u8 max_pf;
 };
 
 void nbl_common_destroy_wq(struct nbl_common_info *common);
 int nbl_common_create_wq(struct nbl_common_info *common);
-
+int nbl_common_func_id_to_rel_pf_id(struct nbl_common_info *common, u32 pf_id,
+				    u32 *rel_pf_id);
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index d87bf9d41a24..28c3366aa01e 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -8,9 +8,11 @@
 
 #include <linux/types.h>
 
+struct nbl_board_port_info;
 struct nbl_hw_mgt;
 struct nbl_adapter;
 struct nbl_hw_ops {
+	void (*flush_write)(struct nbl_hw_mgt *hw_mgt);
 	void (*update_mailbox_queue_tail_ptr)(struct nbl_hw_mgt *hw_mgt,
 					      u16 tail_ptr, u8 txrx);
 	void (*config_mailbox_rxq)(struct nbl_hw_mgt *hw_mgt,
@@ -36,9 +38,22 @@ struct nbl_hw_ops {
 	 *   unsupported by the driver and rejected with -EINVAL.
 	 */
 	void (*get_host_pf_mask)(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask);
+	void (*get_real_bus)(struct nbl_hw_mgt *hw_mgt, u8 *bus);
 
 	void (*cfg_mailbox_qinfo)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
 				  u8 bus, u8 devid, u8 function);
+	void (*get_fw_eth_map)(struct nbl_hw_mgt *hw_mgt, u32 *eth_map);
+	/**
+	 * get_board_info - Fetch board info from firmware
+	 * @hw_mgt: hardware management context
+	 * @board_info: output pointer for board info structure
+	 *
+	 * Firmware contract: board_info.eth_num MUST equal the number of
+	 * unmasked PFs from get_host_pf_mask(). See get_host_pf_mask for
+	 * details.
+	 */
+	void (*get_board_info)(struct nbl_hw_mgt *hw_mgt,
+			       struct nbl_board_port_info *board_info);
 };
 
 struct nbl_hw_ops_tbl {
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
new file mode 100644
index 000000000000..7136b282fb80
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
@@ -0,0 +1,29 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_RESOURCE_H_
+#define _NBL_DEF_RESOURCE_H_
+
+#include <linux/types.h>
+
+struct nbl_resource_mgt;
+struct nbl_adapter;
+
+struct nbl_resource_ops {
+	int (*get_vsi_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			  u16 type, u16 *vsi_id);
+	int (*get_eth_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			  u16 vsi_id, u8 *eth_num, u8 *eth_id,
+			  u8 *logic_eth_id);
+};
+
+struct nbl_resource_ops_tbl {
+	struct nbl_resource_ops *ops;
+	struct nbl_resource_mgt *priv;
+};
+
+int nbl_res_init_leonis(struct nbl_adapter *adapter);
+void nbl_res_remove_leonis(struct nbl_adapter *adapter);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
index f2d802397d98..59e44feab44f 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
@@ -13,6 +13,12 @@
 #define NBL_MAX_PF					8
 #define NBL_NEXT_ID(id, max) (((id) + 1) % ((max) + 1))
 
+#define NBL_MAX_ETHERNET				4
+
+enum {
+	NBL_VSI_DATA = 0,
+};
+
 struct nbl_func_caps {
 	u32 has_ctrl:1;
 	u32 has_net:1;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index b7c80ea54c8d..1aafed2d46d7 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -10,6 +10,7 @@
 #include "nbl_include/nbl_include.h"
 #include "nbl_include/nbl_def_channel.h"
 #include "nbl_include/nbl_def_hw.h"
+#include "nbl_include/nbl_def_resource.h"
 #include "nbl_include/nbl_def_common.h"
 #include "nbl_core.h"
 
@@ -27,6 +28,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 	adapter->pdev = pdev;
 	common = &adapter->common;
 
+	common->mgt_pf = 0;
 	common->pdev = pdev;
 	common->dev = &pdev->dev;
 	common->has_ctrl = param->caps.has_ctrl;
@@ -42,7 +44,13 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 	ret = nbl_chan_init_common(adapter);
 	if (ret)
 		goto chan_init_fail;
+
+	ret = nbl_res_init_leonis(adapter);
+	if (ret)
+		goto res_init_fail;
 	return adapter;
+res_init_fail:
+	nbl_chan_remove_common(adapter);
 chan_init_fail:
 	nbl_hw_remove_leonis(adapter);
 hw_init_fail:
@@ -51,6 +59,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 
 void nbl_core_remove(struct nbl_adapter *adapter)
 {
+	nbl_res_remove_leonis(adapter);
 	nbl_chan_remove_common(adapter);
 	nbl_hw_remove_leonis(adapter);
 }
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 3/8] net/nebula-matrix: add intr resource implementation
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
  2026-09-28 12:32 ` [PATCH v30 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
  2026-09-28 12:32 ` [PATCH v30 net-next 2/8] net/nebula-matrix: add common resource implementation illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  2026-09-28 12:32 ` [PATCH v30 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
                   ` (4 subsequent siblings)
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Add dedicated nbl_interrupt module to manage chip-internal MSI-X interrupt resource
and hardware mapping for Nebula Matrix Ethernet driver.

This module manages driver-wide global hardware MSI-X vector index space,
split into independent network and control interrupt bitmaps (intr_net_bmap/
intr_other_bmap), and handles programming of chip-internal MSI-X mapping
registers. It explicitly does not manage physical PCI MSI-X entries;
physical vector allocation via pci_alloc_irq_vectors() will be implemented
in a follow-up patch via nbl_dev_init_interrupt_scheme().

Core functional interfaces:
1. cfg_msix_map: Allocate global hardware MSI-X vectors from separate net/other
interrupt bitmaps. Reuse per-function coherent DMA map table on reconfig,
eliminating free/realloc cycles, DMA race windows and redundant quiesce
sleeps. Only tear down old hardware state after new allocation succeeds
to avoid interrupt loss, and program MSI-X table DMA address + control-PF
BDF into NBL_PCOMPLETER_FUNCTION_MSIX_MAP.

2. destroy_msix_map: Recycle global vector indices, clear hardware MSI-X mappings,
and release DMA/descriptor resources. Implements two-stage teardown:
disable mailbox IRQ routing first, mask vectors, clear hardware VALID bits
while retaining live DMA addresses, wait ~1ms for hardware DMA quiescence,
then zero table entries and free coherent memory to prevent torn reads.

3. set_mailbox_irq: Toggle PF-specific mailbox MSI-X routing by updating
NBL_MAILBOX_QINFO_MAP_REG_ARR. The disable path works without a configured
MSI-X map, enabling safe routing cleanup before vector release.

4. cfg_msix_info: Program PADPT_HOST_MSIX_INFO and PCOMPLETER_HOST_MSIX_FID_TABLE
with strict hardware-defined programming order (forward for enable, reverse
for teardown) to avoid inconsistent hardware state.

Key design & safety features:
- Self-contained intr_mgt->lock protects global bitmaps and per-function
state; all public APIs internally hold the lock, no upper-layer locking
required by callers.
- Global cleanup via nbl_intr_mgt_stop(): iterate all function IDs to
clean up leftover MSI-X maps. Remote PF maps are originally created via
mailbox RPC in later patches; cleanup uses local MMIO writes.
- PF-only support: explicitly reject VF function IDs with -EOPNOTSUPP.

Instantiate the interrupt manager via nbl_intr_mgt_start() during device
resource initialization, and attach it to the resource management context.

Add corresponding hardware register definitions, helper functions, resource
ops callbacks, and Makefile entries to wire up the new module.
The new resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq)
are hooked into resource_ops but have no in-tree callers within this patch;
invocation will be added in subsequent patches in this series.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../net/ethernet/nebula-matrix/nbl/Makefile   |   1 +
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c  | 155 +++-
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h  |  42 ++
 .../nbl_hw_leonis/nbl_resource_leonis.c       |  33 +-
 .../nbl_hw_leonis/nbl_resource_leonis.h       |   1 +
 .../nebula-matrix/nbl/nbl_hw/nbl_interrupt.c  | 703 ++++++++++++++++++
 .../nebula-matrix/nbl/nbl_hw/nbl_interrupt.h  |  21 +
 .../nebula-matrix/nbl/nbl_hw/nbl_resource.c   |  32 +
 .../nebula-matrix/nbl/nbl_hw/nbl_resource.h   |  52 ++
 .../nbl/nbl_include/nbl_def_hw.h              |   9 +
 .../nbl/nbl_include/nbl_def_resource.h        |   6 +
 .../nbl/nbl_include/nbl_include.h             |   4 +
 12 files changed, 1054 insertions(+), 5 deletions(-)
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index 3dab9519a277..5aec8e44f5d7 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -8,4 +8,5 @@ nbl-objs +=	nbl_common/nbl_common.o \
 		nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
 		nbl_hw/nbl_hw_leonis/nbl_resource_leonis.o \
 		nbl_hw/nbl_resource.o \
+		nbl_hw/nbl_interrupt.o \
 		nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index 4d3477f70bcc..acd4f3dd0757 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -97,6 +97,20 @@ static void nbl_hw_rd_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
 	spin_unlock(&hw_mgt->reg_lock);
 }
 
+static void nbl_hw_wr_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg,
+				const u32 *data, u32 len)
+{
+	u32 size = len / 4;
+	u32 i;
+
+	if (len % 4)
+		return;
+	spin_lock(&hw_mgt->reg_lock);
+	for (i = 0; i < size; i++)
+		wr32(hw_mgt->hw_addr, reg + i * sizeof(u32), data[i]);
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
 /*
  * Only call this when has_ctrl=true, which maps enough space
  * (bar_len - 8192) to cover NBL_HW_DUMMY_REG (0x1300904).
@@ -123,6 +137,139 @@ static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)
 	*eth_map = FIELD_GET(NBL_FW_BOARD_DW6_ETH_BITMAP_MASK, data);
 }
 
+/*
+ * nbl_hw_set_mailbox_irq - read-modify-write NBL_MAILBOX_QINFO_MAP_REG_ARR
+ *
+ * The full RMW sequence is wrapped by reg_lock, so concurrent register
+ * access from different CPUs is already serialized safely.
+ * nbl_hw_cfg_mailbox_qinfo() programs the BDF fields during control-PF
+ * init and clears MSIX_IDX/MSIX_IDX_VALID at the same time (they survive
+ * kexec/forced unload without FLR), so mailbox MSIX routing for a PF
+ * starts disarmed at init and is armed only by an explicit en_msix=true
+ * call here.
+ */
+static void nbl_hw_set_mailbox_irq(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+				   bool en_msix, u16 gvec)
+{
+	u32 data = 0;
+
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id), &data,
+		       sizeof(data));
+	data &= ~(NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK |
+		  NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK);
+	if (en_msix)
+		data |= FIELD_PREP(NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK,
+				   gvec) |
+			FIELD_PREP(NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK,
+				   1);
+
+	nbl_hw_wr_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id), &data,
+		       sizeof(data));
+	spin_unlock(&hw_mgt->reg_lock);
+	nbl_flush_writes(hw_mgt);
+}
+
+static void nbl_hw_cfg_msix_map(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+				bool valid, dma_addr_t dma_addr, u8 bus,
+				u8 devid, u8 function)
+{
+	struct nbl_function_msix_map function_msix_map;
+
+	memset(&function_msix_map, 0, sizeof(function_msix_map));
+	if (valid) {
+		/* clear VALID first, prevent torn read of partial entry */
+		function_msix_map.data[0] = 0;
+		function_msix_map.data[1] = 0;
+		function_msix_map.data[2] = 0;
+		nbl_hw_wr_regs_lock(hw_mgt,
+				    NBL_PCOMPLETER_FUNCTION_MSIX_MAP(func_id),
+				    function_msix_map.data,
+				    sizeof(function_msix_map));
+
+		/* program full entry and set VALID */
+		function_msix_map.data[0] = lower_32_bits(dma_addr);
+		function_msix_map.data[1] = upper_32_bits(dma_addr);
+		function_msix_map.data[2] =
+			FIELD_PREP(NBL_FUNCTION_MSIX_MAP_FUNCTION_MASK,
+				   function) |
+			FIELD_PREP(NBL_FUNCTION_MSIX_MAP_DEVID_MASK, devid) |
+			FIELD_PREP(NBL_FUNCTION_MSIX_MAP_BUS_MASK, bus) |
+			FIELD_PREP(NBL_FUNCTION_MSIX_MAP_VALID_MASK, 1);
+
+		nbl_hw_wr_regs_lock(hw_mgt,
+				    NBL_PCOMPLETER_FUNCTION_MSIX_MAP(func_id),
+				    function_msix_map.data,
+				    sizeof(function_msix_map));
+	} else {
+		/*
+		 * reg_lock prevents concurrent CPU writes to the same
+		 * function's MSIX entry, but cannot synchronize hardware DMA
+		 * reads. Upper layer uses two-stage destruction + sync sleep
+		 * to avoid torn hardware read of partial MSIX entry.
+		 * Keep valid live dma address here, only clear VALID flag.
+		 */
+		function_msix_map.data[0] = lower_32_bits(dma_addr);
+		function_msix_map.data[1] = upper_32_bits(dma_addr);
+		function_msix_map.data[2] = 0;
+
+		nbl_hw_wr_regs_lock(hw_mgt,
+				    NBL_PCOMPLETER_FUNCTION_MSIX_MAP(func_id),
+				    function_msix_map.data,
+				    sizeof(function_msix_map));
+	}
+}
+
+static void nbl_hw_cfg_msix_info(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+				 bool valid, u16 interrupt_id, u8 bus,
+				 u8 devid, u8 function, bool msix_mask_en)
+{
+	u32 host_msix_fid = 0;
+	struct nbl_host_msix_info msix_info;
+
+	memset(&msix_info, 0, sizeof(msix_info));
+	if (valid) {
+		host_msix_fid =
+			FIELD_PREP(NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_FID_MASK,
+				   func_id) |
+			FIELD_PREP(NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_VLD_MASK,
+				   1);
+
+		msix_info.data[1] =
+			FIELD_PREP(NBL_HOST_MSIX_INFO_FUNCTION_MASK, function) |
+			FIELD_PREP(NBL_HOST_MSIX_INFO_DEVID_MASK, devid) |
+			FIELD_PREP(NBL_HOST_MSIX_INFO_BUS_MASK, bus) |
+			FIELD_PREP(NBL_HOST_MSIX_INFO_VALID_MASK, 1);
+
+		if (msix_mask_en)
+			msix_info.data[1] |=
+			FIELD_PREP(NBL_HOST_MSIX_INFO_MSIX_MASK_EN_MASK, 1);
+	}
+	spin_lock(&hw_mgt->reg_lock);
+	/*
+	 * Programming order rule:
+	 * Enable: PADPT_HOST_MSIX_INFO -> PCOMPLETER_HOST_MSIX_FID_TABLE
+	 * Teardown: reverse order, clear FID VLD first to avoid inconsistent
+	 * state
+	 */
+	if (valid) {
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_PADPT_HOST_MSIX_INFO_REG_ARR(interrupt_id),
+			       msix_info.data, sizeof(msix_info));
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_PCOMPLETER_HOST_MSIX_FID_TABLE(interrupt_id),
+			       &host_msix_fid, sizeof(host_msix_fid));
+	} else {
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_PCOMPLETER_HOST_MSIX_FID_TABLE(interrupt_id),
+			       &host_msix_fid, sizeof(host_msix_fid));
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_PADPT_HOST_MSIX_INFO_REG_ARR(interrupt_id),
+			       msix_info.data, sizeof(msix_info));
+	}
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
 static void nbl_hw_update_mailbox_queue_tail_ptr(struct nbl_hw_mgt *hw_mgt,
 						 u16 tail_ptr, u8 txrx)
 {
@@ -258,6 +405,8 @@ static void nbl_hw_get_board_info(struct nbl_hw_mgt *hw_mgt,
 }
 
 static struct nbl_hw_ops hw_ops = {
+	.cfg_msix_map = nbl_hw_cfg_msix_map,
+	.cfg_msix_info = nbl_hw_cfg_msix_info,
 	.flush_write = nbl_flush_writes,
 
 	.update_mailbox_queue_tail_ptr = nbl_hw_update_mailbox_queue_tail_ptr,
@@ -269,6 +418,7 @@ static struct nbl_hw_ops hw_ops = {
 	.get_real_bus = nbl_hw_get_real_bus,
 
 	.cfg_mailbox_qinfo = nbl_hw_cfg_mailbox_qinfo,
+	.set_mailbox_irq = nbl_hw_set_mailbox_irq,
 
 	.get_fw_eth_map = nbl_hw_get_fw_eth_map,
 	.get_board_info = nbl_hw_get_board_info,
@@ -299,11 +449,12 @@ static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
 	hw_ops_tbl = devm_kzalloc(dev, sizeof(*hw_ops_tbl), GFP_KERNEL);
 	if (!hw_ops_tbl)
 		return ERR_PTR(-ENOMEM);
-	if (!hw_ops.flush_write || !hw_ops.update_mailbox_queue_tail_ptr ||
+	if (!hw_ops.cfg_msix_map || !hw_ops.cfg_msix_info ||
+	    !hw_ops.flush_write || !hw_ops.update_mailbox_queue_tail_ptr ||
 	    !hw_ops.config_mailbox_rxq || !hw_ops.config_mailbox_txq ||
 	    !hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
 	    !hw_ops.get_host_pf_mask || !hw_ops.get_real_bus ||
-	    !hw_ops.cfg_mailbox_qinfo ||
+	    !hw_ops.cfg_mailbox_qinfo || !hw_ops.set_mailbox_irq ||
 	    !hw_ops.get_fw_eth_map || !hw_ops.get_board_info)
 		return ERR_PTR(-EINVAL);
 	hw_ops_tbl->ops = &hw_ops;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index 251dd68d0721..5ca74b63ef42 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -54,6 +54,48 @@ struct nbl_mailbox_qinfo_cfg_table {
 #define NBL_PCIE_HOST_TL_CFG_BUSDEV (NBL_INTF_HOST_PCIE_BASE + 0x11040)
 
 #define NBL_PCIE_BUS_MASK	GENMASK(12, 5)
+
+/*  --------  HOST_PADPT  --------  */
+/* host_padpt host_msix_info */
+#define NBL_PADPT_HOST_MSIX_INFO_REG_ARR(vector_id) \
+	(NBL_INTF_HOST_PADPT_BASE + 0x00010000 +    \
+	 (vector_id) * sizeof(struct nbl_host_msix_info))
+
+#define NBL_HOST_MSIX_INFO_DWLEN	2
+/* data[0] */
+#define NBL_HOST_MSIX_INFO_INTRL_PNUM_MASK GENMASK(15, 0)
+#define NBL_HOST_MSIX_INFO_INTRL_RATE_MASK GENMASK(31, 16)
+/* data[1] */
+#define NBL_HOST_MSIX_INFO_FUNCTION_MASK GENMASK(2, 0)
+#define NBL_HOST_MSIX_INFO_DEVID_MASK GENMASK(7, 3)
+#define NBL_HOST_MSIX_INFO_BUS_MASK GENMASK(15, 8)
+#define NBL_HOST_MSIX_INFO_VALID_MASK BIT(16)
+#define NBL_HOST_MSIX_INFO_MSIX_MASK_EN_MASK BIT(17)
+struct nbl_host_msix_info {
+	u32 data[NBL_HOST_MSIX_INFO_DWLEN];
+};
+
+/*  --------  HOST_PCOMPLETER  --------  */
+/* pcompleter_host function_msix_map_table */
+#define NBL_PCOMPLETER_FUNCTION_MSIX_MAP(i)   \
+	(NBL_INTF_HOST_PCOMPLETER_BASE + 0x00004000 + \
+	 (i) * sizeof(struct nbl_function_msix_map))
+#define NBL_PCOMPLETER_HOST_MSIX_FID_TABLE(i) \
+	(NBL_INTF_HOST_PCOMPLETER_BASE + 0x0003a000 + (i) * sizeof(u32))
+
+#define NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_FID_MASK  GENMASK(9, 0)
+#define NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_VLD_MASK  BIT(10)
+
+#define NBL_FUNC_MSIX_MAP_DWLEN		4
+/* data[2] */
+#define NBL_FUNCTION_MSIX_MAP_FUNCTION_MASK GENMASK(2, 0)
+#define NBL_FUNCTION_MSIX_MAP_DEVID_MASK GENMASK(7, 3)
+#define NBL_FUNCTION_MSIX_MAP_BUS_MASK GENMASK(15, 8)
+#define NBL_FUNCTION_MSIX_MAP_VALID_MASK BIT(16)
+struct nbl_function_msix_map {
+	u32 data[NBL_FUNC_MSIX_MAP_DWLEN];
+};
+
 #define NBL_FW_BOARD_CONFIG			0x200
 #define NBL_FW_BOARD_DW3_OFFSET			(NBL_FW_BOARD_CONFIG + 12)
 #define NBL_FW_BOARD_DW6_OFFSET			(NBL_FW_BOARD_CONFIG + 24)
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
index 7804762a96e0..9b53e70ae4af 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
@@ -10,6 +10,9 @@
 static struct nbl_resource_ops res_ops = {
 	.get_vsi_id = nbl_res_func_id_to_vsi_id,
 	.get_eth_id = nbl_res_get_eth_id,
+	.cfg_msix_map = nbl_res_intr_cfg_msix_map,
+	.destroy_msix_map = nbl_res_intr_destroy_msix_map,
+	.set_mailbox_irq = nbl_res_intr_set_mailbox_irq,
 };
 
 static struct nbl_resource_mgt *
@@ -41,7 +44,9 @@ nbl_res_setup_ops(struct device *dev, struct nbl_resource_mgt *res_mgt)
 	res_ops_tbl = devm_kzalloc(dev, sizeof(*res_ops_tbl), GFP_KERNEL);
 	if (!res_ops_tbl)
 		return ERR_PTR(-ENOMEM);
-	if (!res_ops.get_vsi_id || !res_ops.get_eth_id)
+	if (!res_ops.get_vsi_id || !res_ops.get_eth_id ||
+	    !res_ops.cfg_msix_map || !res_ops.destroy_msix_map ||
+	    !res_ops.set_mailbox_irq)
 		return ERR_PTR(-EINVAL);
 	res_ops_tbl->ops = &res_ops;
 	res_ops_tbl->priv = res_mgt;
@@ -282,6 +287,10 @@ static int nbl_res_start(struct nbl_resource_mgt *res_mgt)
 		ret = nbl_res_ctrl_dev_vsi_info_init(res_mgt);
 		if (ret)
 			return ret;
+
+		ret = nbl_intr_mgt_start(res_mgt);
+		if (ret)
+			return ret;
 	}
 
 	return 0;
@@ -322,8 +331,26 @@ int nbl_res_init_leonis(struct nbl_adapter *adap)
 
 void nbl_res_remove_leonis(struct nbl_adapter *adap)
 {
+	struct nbl_resource_mgt *res_mgt = adap->core.res_mgt;
+	struct nbl_common_info *common = &adap->common;
+
+	if (!res_mgt)
+		return;
+
 	/*
-	 * No resource release here because all memory uses devm managed
-	 * allocation
+	 * Tear down all MSI-X maps before destroying coherent tables.
+	 * This is critical on the control PF, which may hold
+	 * maps for remote PFs that are still bound.
+	 */
+	if (common->has_ctrl && res_mgt->intr_mgt)
+		nbl_intr_mgt_stop(res_mgt);
+
+	/* Note:
+	 * per-function interrupts arrays (kcalloc) are freed by
+	 * nbl_intr_mgt_stop().
+	 * MSIX coherent tables are explicitly freed by dma_free_coherent()
+	 * inside the intr destroy path, before nbl_intr_mgt_stop() returns.
+	 * intr_mgt itself (devm_kzalloc) is released by devres after this
+	 * function returns
 	 */
 }
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
index b9355262c00d..6eb4dc9e695a 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
@@ -7,4 +7,5 @@
 #define _NBL_RESOURCE_LEONIS_H_
 
 #include "../nbl_resource.h"
+#include "../nbl_interrupt.h"
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
new file mode 100644
index 000000000000..c41cfa14f90f
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
@@ -0,0 +1,703 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/delay.h>
+#include <linux/dma-mapping.h>
+#include <linux/bitfield.h>
+#include "nbl_interrupt.h"
+
+#define NBL_MSIX_DMA_SYNC_MIN_US	1000 /* us */
+#define NBL_MSIX_DMA_SYNC_MAX_US	1200 /* us */
+
+/*
+ * Release global vector IDs back to intr_net_bmap / intr_other_bmap.
+ * Caller must hold intr_mgt->lock.
+ */
+static void nbl_intr_release_bitmap(struct nbl_resource_mgt *res_mgt,
+				    u16 *vec_buf, u16 cnt)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	u16 bit;
+	u16 i;
+
+	lockdep_assert_held(&intr_mgt->lock);
+
+	if (!vec_buf || cnt == 0)
+		return;
+
+	for (i = 0; i < cnt; i++) {
+		u16 intr_index = vec_buf[i];
+
+		if (intr_index >= NBL_NET_INTR_BASE) {
+			bit = intr_index - NBL_NET_INTR_BASE;
+			if (bit < NBL_MAX_NET_INTERRUPT)
+				clear_bit(bit, intr_mgt->intr_net_bmap);
+			else
+				dev_warn(res_mgt->common->dev,
+					 "invalid net intr index %u\n",
+					 intr_index);
+		} else {
+			if (intr_index < NBL_MAX_OTHER_INTERRUPT)
+				clear_bit(intr_index,
+					  intr_mgt->intr_other_bmap);
+			else
+				dev_warn(res_mgt->common->dev,
+					 "invalid other intr index %u\n",
+					 intr_index);
+		}
+	}
+}
+
+/*
+ * Internal (unlocked) mailbox IRQ bind.  Caller must hold
+ * intr_mgt->lock.  The disable path does not require a configured
+ * MSI-X map because the hardware op ignores gvec when
+ * en_msix=false.
+ */
+static int __nbl_res_intr_set_mailbox_irq(struct nbl_resource_mgt *res_mgt,
+					  u16 func_id, u16 vector_id,
+					  bool en_msix)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	struct nbl_common_info *common = res_mgt->common;
+	struct device *dev = common->dev;
+	struct nbl_func_interrupt_resource_mng *func_res;
+	u16 gvec;
+
+	lockdep_assert_held(&intr_mgt->lock);
+
+	if (func_id >= NBL_MAX_FUNC) {
+		dev_err(dev, "func_id %u out of range\n", func_id);
+		return -EINVAL;
+	}
+
+	if (!en_msix) {
+		hw_ops->set_mailbox_irq(res_mgt->hw_ops_tbl->priv,
+					func_id, false, 0);
+		hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+		return 0;
+	}
+
+	/*
+	 * Enable path: the map must be live and not under teardown,
+	 * otherwise routing would point at a vector that the complete
+	 * phase is about to release and never re-disables.
+	 */
+	if (intr_mgt->stopping)
+		return -ESHUTDOWN;
+
+	func_res = &intr_mgt->func_intr_res[func_id];
+	if (func_res->state != NBL_INTR_FUNC_CONFIGURED) {
+		dev_err(dev, "func %u MSIX map not configured (state %u)\n",
+			func_id, func_res->state);
+		return -ENODEV;
+	}
+
+	if (vector_id >= func_res->num_interrupts) {
+		dev_err(dev, "vector_id %u out of range (max %u)\n",
+			vector_id, func_res->num_interrupts - 1);
+		return -EINVAL;
+	}
+
+	gvec = func_res->interrupts[vector_id];
+	hw_ops->set_mailbox_irq(res_mgt->hw_ops_tbl->priv, func_id,
+				en_msix, gvec);
+	hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+	return 0;
+}
+
+/*
+ * Internal (unlocked) MSI-X map teardown prepare phase: only hardware
+ * register operations. The DMA address is retained and only the VALID
+ * bit is cleared; zeroing the address (Stage 2) is deferred to the
+ * complete phase after the hardware-DMA quiesce window.
+ *
+ * Caller must hold intr_mgt->lock.
+ */
+static int
+__nbl_res_intr_prepare_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
+					u16 func)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	struct nbl_func_interrupt_resource_mng *func_res;
+	u16 *interrupts;
+	u16 intr_num, i;
+	int ret;
+
+	lockdep_assert_held(&intr_mgt->lock);
+
+	if (func >= NBL_MAX_FUNC) {
+		dev_err(res_mgt->common->dev, "Invalid func_id %u\n", func);
+		return -EINVAL;
+	}
+
+	func_res = &intr_mgt->func_intr_res[func];
+	if (func_res->state != NBL_INTR_FUNC_CONFIGURED)
+		return 0;
+
+	interrupts = func_res->interrupts;
+	intr_num = func_res->num_interrupts;
+
+	/* Step 0: disable mailbox IRQ routing before tearing down map */
+	ret = __nbl_res_intr_set_mailbox_irq(res_mgt, func, 0, false);
+	if (ret) {
+		dev_err(res_mgt->common->dev,
+			"disable mailbox irq failed, func=%u ret=%d\n",
+			func, ret);
+		return ret;
+	}
+
+	/* Step 1: invalidate each MSIX info entry in hardware first */
+	for (i = 0; i < intr_num; i++) {
+		hw_ops->cfg_msix_info(res_mgt->hw_ops_tbl->priv,
+				      func, false, interrupts[i],
+				      0, 0, 0, false);
+	}
+	hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+	/*
+	 * Stage 1: retain the DMA address, only clear the VALID bit.
+	 * Stage 2 runs after the quiesce window in the complete phase.
+	 */
+	hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func,
+			     false, func_res->msix_map_table.dma,
+			     0, 0, 0);
+	hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+	func_res->state = NBL_INTR_FUNC_DESTROYING;
+
+	return 0;
+}
+
+/*
+ * __nbl_res_intr_complete_destroy_msix_map - finish hardware teardown and
+ * release vector bitmap, DMA memory and interrupt buffer after the
+ * hardware quiesce window has elapsed.
+ *
+ * Caller must hold intr_mgt->lock.
+ */
+static int
+__nbl_res_intr_complete_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
+					 u16 func_id)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	struct nbl_func_interrupt_resource_mng *func_res;
+	struct nbl_msix_map_table *msix_map_table;
+	struct device *dev = res_mgt->common->dev;
+	u16 *interrupts;
+	u16 intr_num;
+
+	lockdep_assert_held(&intr_mgt->lock);
+
+	if (func_id >= NBL_MAX_FUNC) {
+		dev_err(dev, "Invalid func_id %u\n", func_id);
+		return -EINVAL;
+	}
+
+	func_res = &intr_mgt->func_intr_res[func_id];
+	if (func_res->state != NBL_INTR_FUNC_DESTROYING)
+		return 0;
+
+	/*
+	 * Stage 2: the quiesce window has elapsed, it is now safe to
+	 * zero the DMA base address in the hardware map register.
+	 */
+	hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
+			     false, 0, 0, 0, 0);
+	hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+	interrupts = func_res->interrupts;
+	intr_num = func_res->num_interrupts;
+	msix_map_table = &func_res->msix_map_table;
+
+	if (interrupts) {
+		nbl_intr_release_bitmap(res_mgt, interrupts, intr_num);
+		kfree(interrupts);
+	}
+
+	/*
+	 * Release the coherent table independently of interrupts so a
+	 * partially built config (table allocated, vectors never
+	 * published) cannot leak coherent DMA memory.
+	 */
+	if (msix_map_table->base_addr) {
+		dma_free_coherent(dev, msix_map_table->size,
+				  msix_map_table->base_addr,
+				  msix_map_table->dma);
+		msix_map_table->base_addr = NULL;
+		msix_map_table->dma = 0;
+		msix_map_table->size = 0;
+	}
+
+	func_res->interrupts = NULL;
+	func_res->num_interrupts = 0;
+	func_res->num_net_interrupts = 0;
+	func_res->state = NBL_INTR_FUNC_IDLE;
+
+	return 0;
+}
+
+/*
+ * Internal (unlocked) MSI-X map teardown.  Caller must hold
+ * intr_mgt->lock for the whole sequence, including the hardware-DMA
+ * quiesce window: dropping the lock would let a concurrent caller (or
+ * nbl_intr_mgt_stop()) install/free state against this teardown.
+ *
+ * This is used for the single function synchronous destroy path.
+ */
+static int __nbl_res_intr_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
+					   u16 func_id)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	int ret;
+
+	lockdep_assert_held(&intr_mgt->lock);
+
+	if (intr_mgt->stopping)
+		return -ESHUTDOWN;
+
+	ret = __nbl_res_intr_prepare_destroy_msix_map(res_mgt, func_id);
+	if (ret)
+		return ret;
+	/*
+	 * prepare() only transitions CONFIGURED functions; an IDLE func
+	 * has nothing to wait for or complete.
+	 */
+	if (intr_mgt->func_intr_res[func_id].state !=
+	    NBL_INTR_FUNC_DESTROYING)
+		return 0;
+
+	usleep_range(NBL_MSIX_DMA_SYNC_MIN_US, NBL_MSIX_DMA_SYNC_MAX_US);
+
+	return __nbl_res_intr_complete_destroy_msix_map(res_mgt, func_id);
+}
+
+int nbl_res_intr_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
+				  u16 func_id)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	int ret;
+
+	if (!intr_mgt)
+		return -EINVAL;
+
+	mutex_lock(&intr_mgt->lock);
+	ret = __nbl_res_intr_destroy_msix_map(res_mgt, func_id);
+	mutex_unlock(&intr_mgt->lock);
+
+	return ret;
+}
+
+/**
+ * nbl_res_intr_cfg_msix_map - allocate & program MSI-X mapping table
+ * @res_mgt: resource management instance
+ * @func_id: target function identifier
+ * @num_net_msix: required net data interrupt vectors
+ * @num_others_msix: required control interrupt vectors
+ * @net_msix_mask_en: enable mask for net interrupt entries
+ *
+ * Allocate interrupt vectors; MSIX coherent DMA table is allocated once
+ * per function on first configuration, entries are rewritten while the
+ * map is invalidated on subsequent reconfigurations. No free/realloc of
+ * DMA table on vector count changes. This removes the DMA table
+ * free/realloc cycle. On reconfiguration the map VALID bit is cleared
+ * and the hardware-DMA quiesce window is observed (lock held) before old
+ * vectors are recycled and the table is rewritten.
+ *
+ * Serialization: this function takes intr_mgt->lock internally to
+ * protect the global vector bitmaps and per-function state against
+ * concurrent callers.
+ *
+ * Return: 0 on success, negative errno on failure
+ */
+int nbl_res_intr_cfg_msix_map(struct nbl_resource_mgt *res_mgt,
+			      u16 func_id, u16 num_net_msix,
+			      u16 num_others_msix,
+			      bool net_msix_mask_en)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	struct nbl_common_info *common = res_mgt->common;
+	struct nbl_msix_map_table *official_tbl;
+	struct nbl_msix_map *msix_map_entries;
+	struct device *dev = common->dev;
+	u16 requested, intr_index;
+	u8 bus, devid, function;
+	bool entry_masked = false;
+	u16 *tmp_interrupts = NULL;
+	u16 allocated_cnt = 0;
+	u16 *old_interrupts;
+	u16 old_num;
+	bool had_config;
+	int ret = 0;
+	u16 gvec;
+	u16 i, j;
+
+	if (!intr_mgt)
+		return -EINVAL;
+
+	if (!common->has_ctrl)
+		return -EINVAL;
+
+	if (func_id >= NBL_MAX_FUNC) {
+		dev_err(dev, "Invalid func_id %u\n", func_id);
+		return -EINVAL;
+	}
+
+	if (num_net_msix == 0 && num_others_msix == 0) {
+		dev_err(dev, "MSI-X vector count cannot both be zero\n");
+		return -EINVAL;
+	}
+
+	if (num_net_msix > NBL_MSIX_MAP_TABLE_MAX_ENTRIES ||
+	    num_others_msix > NBL_MSIX_MAP_TABLE_MAX_ENTRIES) {
+		dev_err(dev, "MSI-X count out of limit: net=%u, others=%u\n",
+			num_net_msix, num_others_msix);
+		return -EINVAL;
+	}
+
+	if (check_add_overflow(num_net_msix, num_others_msix, &requested) ||
+	    requested > NBL_MSIX_MAP_TABLE_MAX_ENTRIES) {
+		dev_err(dev, "Total MSI-X vectors %u exceeds maximum %u\n",
+			requested, NBL_MSIX_MAP_TABLE_MAX_ENTRIES);
+		return -EINVAL;
+	}
+
+	ret = nbl_res_func_id_to_bdf(res_mgt, func_id, &bus, &devid, &function);
+	if (ret) {
+		if (ret == -EOPNOTSUPP)
+			dev_err(dev,
+				"MSI-X mapping for VF func_id=%u is not supported\n",
+				func_id);
+		return ret;
+	}
+
+	mutex_lock(&intr_mgt->lock);
+	official_tbl = &intr_mgt->func_intr_res[func_id].msix_map_table;
+
+	/* Reject new configs during teardown or while func is mid-destroy */
+	if (intr_mgt->stopping) {
+		ret = -ESHUTDOWN;
+		goto out_unlock;
+	}
+	if (intr_mgt->func_intr_res[func_id].state ==
+	    NBL_INTR_FUNC_DESTROYING) {
+		ret = -EBUSY;
+		goto out_unlock;
+	}
+
+	had_config = intr_mgt->func_intr_res[func_id].state ==
+		     NBL_INTR_FUNC_CONFIGURED;
+
+	/*
+	 * Phase1: allocate global vector array first.
+	 * Allocate the fixed-size MSIX DMA table only ONCE for this function.
+	 */
+	tmp_interrupts = kcalloc(requested, sizeof(*tmp_interrupts),
+				 GFP_KERNEL);
+	if (!tmp_interrupts) {
+		ret = -ENOMEM;
+		goto out_unlock;
+	}
+	/* Allocate MSIX DMA table once per function */
+	if (!official_tbl->base_addr) {
+		official_tbl->size =
+			sizeof(struct nbl_msix_map) *
+			NBL_MSIX_MAP_TABLE_MAX_ENTRIES;
+		official_tbl->base_addr = dma_alloc_coherent(dev,
+							     official_tbl->size,
+							     &official_tbl->dma,
+							     GFP_KERNEL);
+		if (!official_tbl->base_addr) {
+			dev_err(dev, "Failed to allocate DMA memory for MSIX table\n");
+			ret = -ENOMEM;
+			goto release_vecs_unlock;
+		}
+	}
+
+	/* Allocate net interrupt vectors */
+	for (i = 0; i < num_net_msix; i++) {
+		intr_index = find_first_zero_bit(intr_mgt->intr_net_bmap,
+						 NBL_MAX_NET_INTERRUPT);
+		if (intr_index == NBL_MAX_NET_INTERRUPT) {
+			dev_err(dev, "No free net interrupt vectors left\n");
+			ret = -EAGAIN;
+			goto release_vecs_unlock;
+		}
+		tmp_interrupts[i] = intr_index + NBL_NET_INTR_BASE;
+		set_bit(intr_index, intr_mgt->intr_net_bmap);
+		allocated_cnt++;
+	}
+
+	/* Allocate other interrupt vectors */
+	for (; i < requested; i++) {
+		intr_index =
+			find_first_zero_bit(intr_mgt->intr_other_bmap,
+					    NBL_MAX_OTHER_INTERRUPT);
+		if (intr_index == NBL_MAX_OTHER_INTERRUPT) {
+			dev_err(dev, "No free control interrupt vectors left\n");
+			ret = -EAGAIN;
+			goto release_vecs_unlock;
+		}
+		tmp_interrupts[i] = intr_index;
+		set_bit(intr_index, intr_mgt->intr_other_bmap);
+		allocated_cnt++;
+	}
+
+	/*
+	 * Phase2: quiesce the old hardware MSIX config before touching
+	 * the live DMA table. Same sequence as destroy:
+	 *   disable mailbox routing -> invalidate per-vector INFO ->
+	 *   clear map VALID -> flush -> wait for in-flight table fetches.
+	 * The lock stays held across the wait, so no concurrent caller
+	 * can program the quiesced function. Only then are old vectors
+	 * recycled.
+	 * NOTE: NO DMA table free here.
+	 */
+	if (had_config) {
+		old_interrupts =
+			intr_mgt->func_intr_res[func_id].interrupts;
+		old_num = intr_mgt->func_intr_res[func_id].num_interrupts;
+
+		ret = __nbl_res_intr_set_mailbox_irq(res_mgt, func_id, 0,
+						     false);
+		if (ret) {
+			dev_err(dev, "%s: disable old mailbox irq failed, keep old config\n",
+				__func__);
+			goto release_vecs_unlock;
+		}
+		for (j = 0; j < old_num; j++) {
+			hw_ops->cfg_msix_info(res_mgt->hw_ops_tbl->priv,
+					      func_id, false,
+					      old_interrupts[j],
+					      0, 0, 0, false);
+		}
+		hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
+				     false, official_tbl->dma,
+				     0, 0, 0);
+		hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+		usleep_range(NBL_MSIX_DMA_SYNC_MIN_US,
+			     NBL_MSIX_DMA_SYNC_MAX_US);
+
+		nbl_intr_release_bitmap(res_mgt, old_interrupts, old_num);
+		kfree(old_interrupts);
+		intr_mgt->func_intr_res[func_id].interrupts = NULL;
+		intr_mgt->func_intr_res[func_id].num_interrupts = 0;
+		intr_mgt->func_intr_res[func_id].num_net_interrupts = 0;
+	}
+
+	/* Swap new vector array into func state */
+	intr_mgt->func_intr_res[func_id].interrupts = tmp_interrupts;
+	intr_mgt->func_intr_res[func_id].num_interrupts = requested;
+	intr_mgt->func_intr_res[func_id].num_net_interrupts = num_net_msix;
+	tmp_interrupts = NULL;
+
+	/*
+	 * Rewrite the table in the pre-allocated DMA buffer while the
+	 * map is invalid (on reconfig) or not yet valid (on fresh config),
+	 * so the device cannot observe a torn old/new mix. Only entries
+	 * beyond requested count need explicit zeroing.
+	 */
+	msix_map_entries = official_tbl->base_addr;
+	memset(msix_map_entries + requested, 0,
+	       (NBL_MSIX_MAP_TABLE_MAX_ENTRIES - requested) *
+	       sizeof(*msix_map_entries));
+
+	for (i = 0; i < requested; i++) {
+		gvec = intr_mgt->func_intr_res[func_id].interrupts[i];
+		msix_map_entries[i].data =
+			cpu_to_le16(FIELD_PREP(NBL_MSIX_MAP_VALID_MASK, 1) |
+				    FIELD_PREP(NBL_MSIX_MAP_INDEX_MASK,
+					       gvec));
+	}
+
+	/* Ensure coherent table writes are visible before HW fetch/enable */
+	dma_wmb();
+
+	/* Enable per-vector INFO entries after the table is published */
+	for (i = 0; i < requested; i++) {
+		gvec = intr_mgt->func_intr_res[func_id].interrupts[i];
+		entry_masked = (i < num_net_msix && net_msix_mask_en);
+		hw_ops->cfg_msix_info(res_mgt->hw_ops_tbl->priv,
+				      func_id, true, gvec,
+				      bus, devid, function,
+				      entry_masked);
+	}
+	hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+	/*
+	 * Point the map at the table last and set VALID.
+	 *
+	 * cfg_msix_map uses the control PF's own BDF (common->hw_bus etc.),
+	 * not the target function's BDF.  This BDF tags the pcompler DMA
+	 * read of the MSI-X map table as originating from the control PF.
+	 * The target function's BDF (bus/devid/function from
+	 * nbl_res_func_id_to_bdf) is used only in cfg_msix_info for the
+	 * host_msix_ctrl table entry BDF filtering.
+	 */
+	hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
+			     true, official_tbl->dma, common->hw_bus,
+			     common->devid, common->function);
+	hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+	intr_mgt->func_intr_res[func_id].state = NBL_INTR_FUNC_CONFIGURED;
+	mutex_unlock(&intr_mgt->lock);
+	return 0;
+
+release_vecs_unlock:
+	nbl_intr_release_bitmap(res_mgt, tmp_interrupts, allocated_cnt);
+	kfree(tmp_interrupts);
+	/*
+	 * On a failed fresh configuration, release the DMA table
+	 * allocated during this call. On a failed reconfiguration the
+	 * old configuration is still intact (it is only torn down
+	 * after all vector allocations succeed) and owns the table.
+	 */
+	if (!had_config && official_tbl->base_addr) {
+		dma_free_coherent(dev, official_tbl->size,
+				  official_tbl->base_addr,
+				  official_tbl->dma);
+		official_tbl->base_addr = NULL;
+		official_tbl->dma = 0;
+		official_tbl->size = 0;
+	}
+out_unlock:
+	mutex_unlock(&intr_mgt->lock);
+	return ret;
+}
+
+/**
+ * nbl_res_intr_set_mailbox_irq - bind mailbox IRQ to specified vector
+ * @res_mgt: resource management instance
+ * @func_id: target function identifier
+ * @vector_id: index inside local interrupt array
+ * @en_msix: enable/disable mailbox interrupt
+ *
+ * Serialization: takes intr_mgt->lock internally.
+ *
+ * Return: 0 on success, negative errno on parameter or state check
+ * failure.  The hardware op is void and cannot report failure.
+ */
+int nbl_res_intr_set_mailbox_irq(struct nbl_resource_mgt *res_mgt,
+				 u16 func_id, u16 vector_id,
+				 bool en_msix)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	struct nbl_common_info *common = res_mgt->common;
+	int ret;
+
+	if (!intr_mgt)
+		return -EINVAL;
+
+	if (!common->has_ctrl)
+		return -EINVAL;
+
+	mutex_lock(&intr_mgt->lock);
+	ret = __nbl_res_intr_set_mailbox_irq(res_mgt, func_id,
+					     vector_id, en_msix);
+	mutex_unlock(&intr_mgt->lock);
+
+	return ret;
+}
+
+static struct nbl_interrupt_mgt *nbl_intr_setup_mgt(struct device *dev)
+{
+	struct nbl_interrupt_mgt *intr_mgt;
+	int err;
+
+	intr_mgt = devm_kzalloc(dev, sizeof(*intr_mgt), GFP_KERNEL);
+	if (!intr_mgt)
+		return ERR_PTR(-ENOMEM);
+
+	err = devm_mutex_init(dev, &intr_mgt->lock);
+	if (err)
+		return ERR_PTR(err);
+
+	intr_mgt->stopping = false;
+	bitmap_zero(intr_mgt->intr_net_bmap, NBL_MAX_NET_INTERRUPT);
+	bitmap_zero(intr_mgt->intr_other_bmap, NBL_MAX_OTHER_INTERRUPT);
+
+	return intr_mgt;
+}
+
+int nbl_intr_mgt_start(struct nbl_resource_mgt *res_mgt)
+{
+	struct device *dev = res_mgt->common->dev;
+	struct nbl_interrupt_mgt *intr_mgt;
+	int ret;
+
+	intr_mgt = nbl_intr_setup_mgt(dev);
+	if (IS_ERR(intr_mgt)) {
+		ret = PTR_ERR(intr_mgt);
+		return ret;
+	}
+
+	res_mgt->intr_mgt = intr_mgt;
+	return 0;
+}
+
+void nbl_intr_mgt_stop(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+	u16 func_id;
+	int ret;
+
+	if (!intr_mgt)
+		return;
+
+	/*
+	 * Phase 1: batch invalidate all hardware MSIX map entries.
+	 * stopping is set under the lock, so any caller racing with the
+	 * quiesce window below either holds the lock and sees stopping
+	 * at its next checkpoint, or acquires it after this phase and
+	 * fails (-ESHUTDOWN/-EBUSY/-ENODEV) before issuing MMIO.
+	 */
+	mutex_lock(&intr_mgt->lock);
+	intr_mgt->stopping = true;
+	for (func_id = 0; func_id < NBL_MAX_FUNC; func_id++) {
+		if (intr_mgt->func_intr_res[func_id].state ==
+		    NBL_INTR_FUNC_CONFIGURED) {
+			dev_info(res_mgt->common->dev,
+				 "intr_mgt_stop: preparing destroy map for func %u\n",
+				 func_id);
+			ret = __nbl_res_intr_prepare_destroy_msix_map(res_mgt,
+								      func_id);
+			if (ret)
+				dev_warn(res_mgt->common->dev,
+					 "intr_mgt_stop: prepare destroy map for func %u failed: %d\n",
+					 func_id, ret);
+		}
+	}
+	mutex_unlock(&intr_mgt->lock);
+
+	/*
+	 * Global quiesce: wait for straggler DMA table reads after all
+	 * MSIX map entries have been invalidated in hardware, before
+	 * freeing coherent memory. Best-effort only.
+	 */
+	usleep_range(NBL_MSIX_DMA_SYNC_MIN_US, NBL_MSIX_DMA_SYNC_MAX_US);
+
+	/* Phase2: safely release MSIX coherent memory and intr resources */
+	mutex_lock(&intr_mgt->lock);
+	for (func_id = 0; func_id < NBL_MAX_FUNC; func_id++) {
+		if (intr_mgt->func_intr_res[func_id].state ==
+		    NBL_INTR_FUNC_DESTROYING) {
+			ret = __nbl_res_intr_complete_destroy_msix_map(res_mgt,
+								       func_id);
+			if (ret)
+				dev_warn(res_mgt->common->dev,
+					 "intr_mgt_stop: complete destroy map for func %u failed: %d\n",
+					 func_id, ret);
+		}
+	}
+	/* Clear the published pointer under the lock, last */
+	res_mgt->intr_mgt = NULL;
+	mutex_unlock(&intr_mgt->lock);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h
new file mode 100644
index 000000000000..9f66f5e19c98
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h
@@ -0,0 +1,21 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_INTERRUPT_H_
+#define _NBL_INTERRUPT_H_
+
+#include "nbl_resource.h"
+
+#define NBL_MSIX_MAP_TABLE_MAX_ENTRIES	1024
+int nbl_res_intr_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
+				  u16 func_id);
+int nbl_res_intr_cfg_msix_map(struct nbl_resource_mgt *res_mgt,
+			      u16 func_id, u16 num_net_msix,
+			      u16 num_others_msix,
+			      bool net_msix_mask_en);
+int nbl_res_intr_set_mailbox_irq(struct nbl_resource_mgt *res_mgt,
+				 u16 func_id, u16 vector_id,
+				 bool en_msix);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
index b316fb8e7051..635f34312c56 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
@@ -68,6 +68,38 @@ int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id)
 	return -ENOENT;
 }
 
+int nbl_res_func_id_to_bdf(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			   u8 *bus, u8 *dev, u8 *function)
+{
+	struct nbl_common_info *common = res_mgt->common;
+	struct nbl_sriov_info *sriov_info;
+	int pfid = func_id;
+	u8 pf_bus, devfn;
+	u32 rel_pf_id;
+	int ret;
+
+	if (!common->has_ctrl || !bus || !dev || !function)
+		return -EINVAL;
+	ret = nbl_common_func_id_to_rel_pf_id(common, pfid, &rel_pf_id);
+	if (ret)
+		return ret;
+	if (rel_pf_id >= common->max_pf) {
+		dev_err(common->dev,
+			"func_id=%u rel_pf_id=%u exceeds max_pf=%u, VF BDF unsupported\n",
+			pfid, rel_pf_id,
+			common->max_pf);
+		return -EOPNOTSUPP;
+	}
+	sriov_info = res_mgt->resource_info->sriov_info + rel_pf_id;
+	pf_bus = PCI_BUS_NUM(sriov_info->bdf);
+	devfn = sriov_info->bdf & 0xff;
+	*bus = pf_bus;
+	*dev = PCI_SLOT(devfn);
+	*function = PCI_FUNC(devfn);
+
+	return 0;
+}
+
 int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
 		       u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id)
 {
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
index ae0a3d33198d..c14a78a47c98 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
@@ -17,6 +17,53 @@
 
 struct nbl_resource_mgt;
 
+/* --------- INTERRUPT ---------- */
+#define NBL_MAX_OTHER_INTERRUPT			1024
+#define NBL_MAX_NET_INTERRUPT			4096
+#define NBL_NET_INTR_BASE		NBL_MAX_OTHER_INTERRUPT
+
+#define NBL_MSIX_MAP_VALID_MASK		BIT(0)
+#define NBL_MSIX_MAP_INDEX_MASK		GENMASK(13, 1)
+#define NBL_MSIX_MAP_RSV_MASK		GENMASK(15, 14)
+
+struct nbl_msix_map {
+	__le16 data;
+};
+
+struct nbl_msix_map_table {
+	struct nbl_msix_map *base_addr;
+	dma_addr_t dma;
+	size_t size;
+};
+
+/*
+ * Per-function MSI-X resource state.  CONFIGURED -> DESTROYING spans
+ * the hardware-DMA quiesce window of a teardown; the lock is held across
+ * that window so a concurrent configuration cannot install a map that
+ * the in-flight teardown would free.
+ */
+enum nbl_intr_func_state {
+	NBL_INTR_FUNC_IDLE = 0,
+	NBL_INTR_FUNC_CONFIGURED,
+	NBL_INTR_FUNC_DESTROYING,
+};
+
+struct nbl_func_interrupt_resource_mng {
+	u16 num_interrupts;
+	u16 num_net_interrupts;
+	u16 *interrupts;
+	struct nbl_msix_map_table msix_map_table;
+	u8 state; /* enum nbl_intr_func_state */
+};
+
+struct nbl_interrupt_mgt {
+	struct mutex lock; /* Protects bitmap + func_intr_res[] */
+	DECLARE_BITMAP(intr_net_bmap, NBL_MAX_NET_INTERRUPT);
+	DECLARE_BITMAP(intr_other_bmap, NBL_MAX_OTHER_INTERRUPT);
+	bool stopping; /* set on teardown, rejects new configurations */
+	struct nbl_func_interrupt_resource_mng func_intr_res[NBL_MAX_FUNC];
+};
+
 /* --------- INFO ---------- */
 struct nbl_sriov_info {
 	unsigned int bdf;
@@ -56,14 +103,19 @@ struct nbl_resource_mgt {
 	struct nbl_resource_info *resource_info;
 	struct nbl_channel_ops_tbl *chan_ops_tbl;
 	struct nbl_hw_ops_tbl *hw_ops_tbl;
+	struct nbl_interrupt_mgt *intr_mgt;
 };
 
 int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id);
 int nbl_res_func_id_to_vsi_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
 			      u16 type, u16 *vsi_id);
+int nbl_res_func_id_to_bdf(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			   u8 *bus, u8 *dev, u8 *function);
 int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
 		       u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id);
+int nbl_intr_mgt_start(struct nbl_resource_mgt *res_mgt);
 int nbl_res_pf_dev_vsi_type_to_hw_vsi_type(struct nbl_resource_mgt *res_mgt,
 					   u16 src_type,
 					   enum nbl_vsi_serv_type *dst_type);
+void nbl_intr_mgt_stop(struct nbl_resource_mgt *res_mgt);
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index 28c3366aa01e..9a9642f52af2 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -12,6 +12,13 @@ struct nbl_board_port_info;
 struct nbl_hw_mgt;
 struct nbl_adapter;
 struct nbl_hw_ops {
+	void (*cfg_msix_map)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+			     bool valid, dma_addr_t dma_addr, u8 bus,
+			     u8 devid, u8 function);
+	void (*cfg_msix_info)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+			      bool valid, u16 interrupt_id, u8 bus,
+			      u8 devid, u8 function,
+			      bool net_msix_mask_en);
 	void (*flush_write)(struct nbl_hw_mgt *hw_mgt);
 	void (*update_mailbox_queue_tail_ptr)(struct nbl_hw_mgt *hw_mgt,
 					      u16 tail_ptr, u8 txrx);
@@ -42,6 +49,8 @@ struct nbl_hw_ops {
 
 	void (*cfg_mailbox_qinfo)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
 				  u8 bus, u8 devid, u8 function);
+	void (*set_mailbox_irq)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+				bool en_msix, u16 gvec);
 	void (*get_fw_eth_map)(struct nbl_hw_mgt *hw_mgt, u32 *eth_map);
 	/**
 	 * get_board_info - Fetch board info from firmware
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
index 7136b282fb80..e718ea41a816 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
@@ -12,6 +12,12 @@ struct nbl_resource_mgt;
 struct nbl_adapter;
 
 struct nbl_resource_ops {
+	int (*cfg_msix_map)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			    u16 num_net_msix, u16 num_others_msix,
+			    bool net_msix_mask_en);
+	int (*destroy_msix_map)(struct nbl_resource_mgt *res_mgt, u16 func_id);
+	int (*set_mailbox_irq)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+			       u16 vector_id, bool en_msix);
 	int (*get_vsi_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
 			  u16 type, u16 *vsi_id);
 	int (*get_eth_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
index 59e44feab44f..34abf20ebf5a 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
@@ -11,8 +11,12 @@
 /*  ------  Basic definitions  -------  */
 #define NBL_DRIVER_NAME					"nbl"
 #define NBL_MAX_PF					8
+/* Chip-wide VF budget shared across all PFs */
+#define NBL_MAX_VF					512
 #define NBL_NEXT_ID(id, max) (((id) + 1) % ((max) + 1))
 
+/* Total PCI functions: PFs plus the chip-wide VF budget */
+#define NBL_MAX_FUNC			(NBL_MAX_PF + NBL_MAX_VF)
 #define NBL_MAX_ETHERNET				4
 
 enum {
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
                   ` (2 preceding siblings ...)
  2026-09-28 12:32 ` [PATCH v30 net-next 3/8] net/nebula-matrix: add intr " illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  2026-09-28 12:32 ` [PATCH v30 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
                   ` (3 subsequent siblings)
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Add Leonis chip-wide hardware init/deinit support for the Nebula Matrix Ethernet
driver, implementing full datapath pipeline initialization and ordered chip-level
resource teardown. This patch establishes complete chip hardware lifecycle
management and driver/firmware state synchronization via the driver_status flag.

Core changes:
1. Add hw_ops.init_module/deinit_module and corresponding resource_ops entry
points for chip-level lifecycle management (caller hooks added later).
2. Enforce deterministic teardown order by adding device link dependencies:
all non-management PFs depend on function 0 management PF, ensuring safe
chip-global firmware deinit without race conditions.
3. Implement firmware quirk parsing to enable hardware-specific tuning,
such as UVN descriptor prefetch alignment optimization.

Initialize all core datapath modules with speed/port-count tailored settings:
- dped/uped: packet engine config, L4 checksum offload, IPv4/IPv6 TCP profiles
- dsch: scheduler quanta and host QID max threshold configuration
- ustore/dstore: Tx/Rx buffer drop thresholds and speed-adaptive flow control
- dvn/uvn: descriptor timeout, PCIe relaxed ordering and prefetch parameters
- uqm: queue counter reset and default queue mode initialization
- shaping: per-port rate shaping and PSHA scheduler configuration

Standard chip init flow:
1. Full datapath initialization via unified nbl_dp_init()
2. Host PADPT layer flow control initialization via nbl_intf_init()
3. Mark driver_status active and flush all hardware register writes

Safety & design rules:
- All chip-global operations are restricted to management PF via has_ctrl
- Driver programs registers only; full chip reset is firmware-managed
- Explicit register flushing synchronizes with asynchronous firmware cleanup
- Teardown ordering blocks new DMA and in-flight mailbox transactions
- Strict validation for port speed and ethernet port count parameters

Deinit note: firmware does not restore chip-wide datapath registers on deinit.
The deinit path only clears driver_status and flushes writes. Hardware state
persists until chip reset or explicit reconfiguration to avoid redundant reinit.

Also add required register definitions, firmware ABI enums, helper routines,
resource callbacks and Makefile entries for the new chip management logic.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../net/ethernet/nebula-matrix/nbl/Makefile   |   1 +
 .../nebula-matrix/nbl/nbl_hw/nbl_chip.c       |  32 +
 .../nebula-matrix/nbl/nbl_hw/nbl_chip.h       |  12 +
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c  | 648 +++++++++++++++++-
 .../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h  | 222 ++++++
 .../nbl_hw_leonis/nbl_resource_leonis.c       |   5 +-
 .../nbl_hw_leonis/nbl_resource_leonis.h       |   1 +
 .../nbl/nbl_include/nbl_def_hw.h              |   3 +
 .../nbl/nbl_include/nbl_def_resource.h        |   3 +
 .../nbl/nbl_include/nbl_include.h             |  19 +
 .../net/ethernet/nebula-matrix/nbl/nbl_main.c |  87 +++
 11 files changed, 1031 insertions(+), 2 deletions(-)
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index 5aec8e44f5d7..be314b909d66 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -9,4 +9,5 @@ nbl-objs +=	nbl_common/nbl_common.o \
 		nbl_hw/nbl_hw_leonis/nbl_resource_leonis.o \
 		nbl_hw/nbl_resource.o \
 		nbl_hw/nbl_interrupt.o \
+		nbl_hw/nbl_chip.o \
 		nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
new file mode 100644
index 000000000000..419eb6392ada
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
@@ -0,0 +1,32 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include "nbl_chip.h"
+
+void nbl_res_chip_deinit_module(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+	struct nbl_common_info *common = res_mgt->common;
+
+	if (!common->has_ctrl)
+		return;
+	hw_ops->deinit_module(res_mgt->hw_ops_tbl->priv);
+}
+
+int nbl_res_chip_init_module(struct nbl_resource_mgt *res_mgt)
+{
+	struct nbl_common_info *common = res_mgt->common;
+	struct nbl_hw_ops *hw_ops;
+	u8 eth_speed, eth_num;
+	struct nbl_hw_mgt *p;
+
+	if (!common->has_ctrl)
+		return -EINVAL;
+	eth_speed = res_mgt->resource_info->board_info.eth_speed;
+	eth_num = res_mgt->resource_info->board_info.eth_num;
+	hw_ops = res_mgt->hw_ops_tbl->ops;
+	p = res_mgt->hw_ops_tbl->priv;
+	return hw_ops->init_module(p, eth_speed, eth_num);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h
new file mode 100644
index 000000000000..d14093ba916c
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h
@@ -0,0 +1,12 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_CHIP_H_
+#define _NBL_CHIP_H_
+
+#include "nbl_resource.h"
+int nbl_res_chip_init_module(struct nbl_resource_mgt *res_mgt);
+void nbl_res_chip_deinit_module(struct nbl_resource_mgt *res_mgt);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index acd4f3dd0757..47ec995e3ac8 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -5,11 +5,23 @@
 #include <linux/device.h>
 #include <linux/pci.h>
 #include <linux/bits.h>
+#include <linux/delay.h>
 #include <linux/io.h>
 #include <linux/spinlock.h>
 #include <linux/bitfield.h>
 #include "nbl_hw_leonis.h"
 
+/*
+ * Firmware cleanup after driver_status=false is asynchronous and the
+ * current hardware revision exposes no cleanup-complete status bit.
+ * Wait a bounded window so the firmware pass finishes before this
+ * function returns, establishing an explicit boundary against a later
+ * init_module() that would otherwise reprogram per-PF/chip-wide
+ * registers while firmware is still wiping them. Best-effort only.
+ */
+#define NBL_FW_CLEANUP_SYNC_MIN_US	2000
+#define NBL_FW_CLEANUP_SYNC_MAX_US	3000
+
 static void nbl_hw_read_mbx_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
 				 u32 len)
 {
@@ -137,6 +149,636 @@ static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)
 	*eth_map = FIELD_GET(NBL_FW_BOARD_DW6_ETH_BITMAP_MASK, data);
 }
 
+static u32 nbl_hw_get_quirks(struct nbl_hw_mgt *hw_mgt)
+{
+	u32 quirks = 0;
+
+	/*
+	 * Read quirk bits from mailbox register.
+	 * All supported firmware implement the quirk ABI,
+	 * firmware always populates NBL_LEONIS_QUIRKS_OFFSET.
+	 * Value ~0U indicates no active quirks.
+	 */
+	nbl_hw_read_mbx_regs(hw_mgt, NBL_LEONIS_QUIRKS_OFFSET, &quirks,
+			     sizeof(u32));
+
+	if (quirks == ~0u)
+		return 0;
+
+	return quirks;
+}
+
+static void nbl_configure_dped_checksum(struct nbl_hw_mgt *hw_mgt)
+{
+	u32 data = 0;
+
+	/* DPED dped_l4_ck_cmd_40 for sctp */
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_DPED_L4_CK_CMD_40_ADDR, &data, sizeof(data));
+	data |= FIELD_PREP(NBL_DPED_L4_CK_CMD_40_EN_MASK, 1);
+	nbl_hw_wr_regs(hw_mgt, NBL_DPED_L4_CK_CMD_40_ADDR, &data, sizeof(data));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_dped_init(struct nbl_hw_mgt *hw_mgt)
+{
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_wr32(hw_mgt, NBL_DPED_VLAN_OFFSET, 0xC);
+	nbl_hw_wr32(hw_mgt, NBL_DPED_DSCP_OFFSET_0, 0x8);
+	nbl_hw_wr32(hw_mgt, NBL_DPED_DSCP_OFFSET_1, 0x4);
+	spin_unlock(&hw_mgt->reg_lock);
+	/* dped checksum offload */
+	nbl_configure_dped_checksum(hw_mgt);
+}
+
+static void nbl_uped_init(struct nbl_hw_mgt *hw_mgt)
+{
+	u32 hw_edit = 0;
+
+	/* V4 TCP: l3_len = 0 */
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V4_TCP_IDX),
+		       &hw_edit, sizeof(hw_edit));
+	hw_edit &= ~NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK;
+	nbl_hw_wr_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V4_TCP_IDX),
+		       &hw_edit, sizeof(hw_edit));
+
+	/* V6 TCP: l3_len = 1 */
+	nbl_hw_rd_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V6_TCP_IDX),
+		       &hw_edit, sizeof(hw_edit));
+	hw_edit = (hw_edit & ~NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK) |
+		  FIELD_PREP(NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK, 1);
+	nbl_hw_wr_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V6_TCP_IDX),
+		       &hw_edit, sizeof(hw_edit));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static int nbl_shaping_eth_init(struct nbl_hw_mgt *hw_mgt, u8 eth_id, u8 speed)
+{
+	struct nbl_shaping_dvn_dport_u dvn_dport = { 0 };
+	struct nbl_shaping_dport_u dport = { 0 };
+	u32 rate, half_rate;
+	u32 depth;
+	u64 low_val, high_val;
+
+	switch (speed) {
+	case NBL_FW_PORT_SPEED_100G:
+		rate = 100000;
+		break;
+	case NBL_FW_PORT_SPEED_50G:
+		rate = 50000;
+		break;
+	case NBL_FW_PORT_SPEED_25G:
+		rate = 25000;
+		break;
+	case NBL_FW_PORT_SPEED_10G:
+		rate = 10000;
+		break;
+	default:
+		dev_err(hw_mgt->common->dev,
+			"Unsupported port speed %u for eth%u\n", speed, eth_id);
+		return -EINVAL;
+	}
+
+	half_rate = rate / 2;
+	depth = max_t(u32, rate * 2, NBL_LR_LEONIS_NET_BUCKET_DEPTH);
+
+	/* 1. clear valid first
+	 * dport and dvn_dport are zero-initialised above, so VALID=0 already
+	 */
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DPORT_REG(eth_id), dport.data,
+		       sizeof(dport));
+	nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DVN_DPORT_REG(eth_id),
+		       dvn_dport.data, sizeof(dvn_dport));
+
+	/* 2. write config words (valid=0, safe) */
+	low_val = FIELD_PREP(NBL_DPORT_CIR_MASK, rate) |
+		  FIELD_PREP(NBL_DPORT_PIR_MASK, rate) |
+		  FIELD_PREP(NBL_DPORT_DEPTH_MASK, depth) |
+		  FIELD_PREP(NBL_DPORT_CBS_MASK_LOW, depth & 0x3F);
+	high_val = FIELD_PREP(NBL_DPORT_CBS_MASK_HIGH, depth >> 6) |
+		   FIELD_PREP(NBL_DPORT_PBS_MASK, depth);
+	/* Fixed split, independent of host endian */
+	dport.data[0] = lower_32_bits(low_val);
+	dport.data[1] = upper_32_bits(low_val);
+	dport.data[2] = lower_32_bits(high_val);
+	dport.data[3] = upper_32_bits(high_val);
+
+	low_val = FIELD_PREP(NBL_DPORT_CIR_MASK, half_rate) |
+		  FIELD_PREP(NBL_DPORT_PIR_MASK, rate) |
+		  FIELD_PREP(NBL_DPORT_DEPTH_MASK, depth) |
+		  FIELD_PREP(NBL_DPORT_CBS_MASK_LOW, depth & 0x3F);
+	high_val = FIELD_PREP(NBL_DPORT_CBS_MASK_HIGH, depth >> 6) |
+		   FIELD_PREP(NBL_DPORT_PBS_MASK, depth);
+	dvn_dport.data[0] = lower_32_bits(low_val);
+	dvn_dport.data[1] = upper_32_bits(low_val);
+	dvn_dport.data[2] = lower_32_bits(high_val);
+	dvn_dport.data[3] = upper_32_bits(high_val);
+
+	nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DPORT_REG(eth_id), dport.data,
+		       sizeof(dport));
+	nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DVN_DPORT_REG(eth_id),
+		       dvn_dport.data, sizeof(dvn_dport));
+
+	/* 3. commit: set valid last */
+	low_val = FIELD_PREP(NBL_DPORT_VALID_MASK, 1);
+	dport.data[0] |= lower_32_bits(low_val);
+
+	low_val = FIELD_PREP(NBL_DPORT_VALID_MASK, 1);
+	dvn_dport.data[0] |= lower_32_bits(low_val);
+
+	nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DPORT_REG(eth_id), dport.data,
+		       sizeof(dport));
+	nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DVN_DPORT_REG(eth_id),
+		       dvn_dport.data, sizeof(dvn_dport));
+	spin_unlock(&hw_mgt->reg_lock);
+	return 0;
+}
+
+static int nbl_shaping_init(struct nbl_hw_mgt *hw_mgt, u8 speed)
+{
+#define NBL_SHAPING_FLUSH_INTERVAL 128
+	struct nbl_shaping_net_u net_shaping = { 0 };
+	u32 eth_bitmap = 0;
+	u32 reg_val;
+	int ret;
+	int i;
+
+	nbl_hw_get_fw_eth_map(hw_mgt, &eth_bitmap);
+	for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+		if (!(eth_bitmap & BIT(i)))
+			continue;
+		ret = nbl_shaping_eth_init(hw_mgt, i, speed);
+		if (ret)
+			return ret;
+	}
+	nbl_hw_rd_regs_lock(hw_mgt, NBL_DSCH_PSHA_EN_ADDR, &reg_val,
+			    sizeof(reg_val));
+	reg_val &= ~NBL_DSCH_PSHA_EN_MASK;
+	reg_val |= FIELD_PREP(NBL_DSCH_PSHA_EN_MASK,
+			      eth_bitmap & GENMASK(3, 0));
+	nbl_hw_wr_regs_lock(hw_mgt, NBL_DSCH_PSHA_EN_ADDR, &reg_val,
+			    sizeof(reg_val));
+
+	for (i = 0; i < NBL_MAX_FUNC; i++) {
+		nbl_hw_wr_regs_lock(hw_mgt, NBL_SHAPING_NET_REG(i),
+				    net_shaping.data,
+				    sizeof(net_shaping));
+		if ((i + 1) % NBL_SHAPING_FLUSH_INTERVAL == 0)
+			nbl_flush_writes(hw_mgt);
+	}
+	nbl_flush_writes(hw_mgt);
+	return 0;
+}
+
+static void nbl_dsch_qid_max_init(struct nbl_hw_mgt *hw_mgt)
+{
+	u32 quanta = 0;
+	u32 qid_max = 0;
+
+	quanta = FIELD_PREP(NBL_DSCH_VN_QUANTA_H_QUA_MASK, NBL_HOST_QUANTA) |
+		 FIELD_PREP(NBL_DSCH_VN_QUANTA_E_QUA_MASK, NBL_ECPU_QUANTA);
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_wr_regs(hw_mgt, NBL_DSCH_VN_QUANTA_ADDR, &quanta,
+		       sizeof(quanta));
+	nbl_hw_rd_regs(hw_mgt, NBL_DSCH_HOST_QID_MAX, &qid_max,
+		       sizeof(qid_max));
+	qid_max &= ~NBL_DSCH_HOST_QID_MAX_MASK;
+	qid_max |= FIELD_PREP(NBL_DSCH_HOST_QID_MAX_MASK, NBL_MAX_QUEUE_ID);
+	nbl_hw_wr_regs(hw_mgt, NBL_DSCH_HOST_QID_MAX, &qid_max,
+		       sizeof(qid_max));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static int nbl_ustore_init(struct nbl_hw_mgt *hw_mgt, u8 eth_num)
+{
+	u32 eth_bitmap = 0;
+	u32 drop_th = 0;
+	u32 pkt_len = 0;
+	u32 reg_val = 0;
+	int i;
+
+	/*
+	 * eth_num is validated in the resource layer:
+	 * nbl_res_init_pf_num() requires 1/2/4 PFs, and
+	 * nbl_res_ctrl_dev_setup_eth_info() requires max_pf == eth_num.
+	 * This is a defensive check only; if it triggers, the resource
+	 * layer validation was bypassed, which is a bug.
+	 */
+	if (WARN_ON(eth_num != 1 && eth_num != 2 && eth_num != 4))
+		return -EINVAL;
+	/* Read current packet length config
+	 *(to preserve other fields while updating 'min')
+	 */
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_USTORE_PKT_LEN_ADDR, &pkt_len,
+		       sizeof(pkt_len));
+	/* min arp packet length 42 (14 + 28) */
+	pkt_len &= ~NBL_USTORE_PKT_LEN_MIN_MASK;
+	pkt_len |= FIELD_PREP(NBL_USTORE_PKT_LEN_MIN_MASK, 42);
+	nbl_hw_wr_regs(hw_mgt, NBL_USTORE_PKT_LEN_ADDR, &pkt_len,
+		       sizeof(pkt_len));
+
+	drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_EN_MASK, 1);
+	if (eth_num == 1)
+		drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK,
+				      NBL_USTORE_SINGLE_ETH_DROP_TH);
+	else if (eth_num == 2)
+		drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK,
+				      NBL_USTORE_DUAL_ETH_DROP_TH);
+	else
+		drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK,
+				      NBL_USTORE_QUAD_ETH_DROP_TH);
+	nbl_hw_get_fw_eth_map(hw_mgt, &eth_bitmap);
+	for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+		if (!(eth_bitmap & BIT(i)))
+			continue;
+		nbl_hw_rd_regs(hw_mgt, NBL_USTORE_PORT_DROP_TH_REG_ARR(i),
+			       &reg_val, sizeof(reg_val));
+		reg_val &= ~(NBL_USTORE_PORT_DROP_TH_EN_MASK |
+			    NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK);
+		reg_val |= drop_th;
+		nbl_hw_wr_regs(hw_mgt, NBL_USTORE_PORT_DROP_TH_REG_ARR(i),
+			       &reg_val, sizeof(reg_val));
+	}
+
+	/* Clear port drop/truncate counters by reading them
+	 * (hardware has read-to-clear behavior for these registers)
+	 */
+	for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+		if (!(eth_bitmap & BIT(i)))
+			continue;
+		nbl_hw_rd32(hw_mgt, NBL_USTORE_BUF_PORT_DROP_PKT(i));
+		nbl_hw_rd32(hw_mgt, NBL_USTORE_BUF_PORT_TRUN_PKT(i));
+	}
+	spin_unlock(&hw_mgt->reg_lock);
+	return 0;
+}
+
+static void nbl_dstore_init(struct nbl_hw_mgt *hw_mgt, u8 speed)
+{
+	u32 eth_bitmap = 0;
+	u32 drop_th = 0;
+	u32 fc_th = 0;
+	u32 bp_th = 0;
+	int i;
+
+	for (i = 0; i < NBL_DSTORE_PORT_DROP_TH_DEPTH; i++) {
+		spin_lock(&hw_mgt->reg_lock);
+		nbl_hw_rd_regs(hw_mgt, NBL_DSTORE_PORT_DROP_TH_REG(i), &drop_th,
+			       sizeof(drop_th));
+		drop_th &= ~NBL_DSTORE_PORT_DROP_EN_MASK;
+		nbl_hw_wr_regs(hw_mgt, NBL_DSTORE_PORT_DROP_TH_REG(i), &drop_th,
+			       sizeof(drop_th));
+		spin_unlock(&hw_mgt->reg_lock);
+	}
+
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_DSTORE_DISC_BP_TH, &bp_th, sizeof(bp_th));
+	bp_th |= FIELD_PREP(NBL_DSTORE_DISC_BP_TH_EN_MASK, 1);
+	nbl_hw_wr_regs(hw_mgt, NBL_DSTORE_DISC_BP_TH, &bp_th, sizeof(bp_th));
+	spin_unlock(&hw_mgt->reg_lock);
+
+	nbl_hw_get_fw_eth_map(hw_mgt, &eth_bitmap);
+	for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+		if (!(eth_bitmap & BIT(i)))
+			continue;
+		spin_lock(&hw_mgt->reg_lock);
+		nbl_hw_rd_regs(hw_mgt, NBL_DSTORE_D_DPORT_FC_TH_REG(i), &fc_th,
+			       sizeof(fc_th));
+		fc_th &= ~(NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK |
+			   NBL_DSTORE_D_DPORT_FC_XON_TH_MASK);
+		if (speed == NBL_FW_PORT_SPEED_100G) {
+			fc_th |=
+				FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK,
+					   NBL_DSTORE_DROP_XOFF_TH_100G) |
+				FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XON_TH_MASK,
+					   NBL_DSTORE_DROP_XON_TH_100G);
+		} else {
+			fc_th |=
+				FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK,
+					   NBL_DSTORE_DROP_XOFF_TH) |
+				FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XON_TH_MASK,
+					   NBL_DSTORE_DROP_XON_TH);
+		}
+
+		fc_th |= FIELD_PREP(NBL_DSTORE_D_DPORT_FC_FC_EN_MASK, 1);
+		nbl_hw_wr_regs(hw_mgt, NBL_DSTORE_D_DPORT_FC_TH_REG(i), &fc_th,
+			       sizeof(fc_th));
+		spin_unlock(&hw_mgt->reg_lock);
+	}
+}
+
+static void nbl_dvn_descreq_num_cfg(struct nbl_hw_mgt *hw_mgt, u8 descreq_num)
+{
+	u8 split_ring_num = (descreq_num >> 3) & 0x1;
+	u8 ring_num = descreq_num & 0x7;
+	u32 num_cfg;
+	u32 reg_val;
+
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_DVN_DESCREQ_NUM_CFG, &reg_val,
+		       sizeof(reg_val));
+
+	num_cfg = FIELD_PREP(NBL_DVN_DESCREQ_NUM_CFG_AVRING_DESREQ_NUM_CFG_MASK,
+			     split_ring_num) |
+		  FIELD_PREP(NBL_DVN_DESCREQ_NUM_CFG_PACKED_L1_NUM_MASK,
+			     ring_num);
+	reg_val &= ~(NBL_DVN_DESCREQ_NUM_CFG_AVRING_DESREQ_NUM_CFG_MASK |
+		NBL_DVN_DESCREQ_NUM_CFG_PACKED_L1_NUM_MASK);
+	reg_val |= num_cfg;
+	nbl_hw_wr_regs(hw_mgt, NBL_DVN_DESCREQ_NUM_CFG, &reg_val,
+		       sizeof(reg_val));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_dvn_init(struct nbl_hw_mgt *hw_mgt, u8 speed)
+{
+	u32 timeout = 0;
+	u32 ro_flag = 0;
+
+	nbl_hw_wr32(hw_mgt, NBL_DVN_ECPU_QUEUE_NUM, 0);
+	timeout = FIELD_PREP(NBL_DVN_DESC_WR_MERGE_TIMEOUT_CFG_CYCLE_MASK,
+			     DEFAULT_DVN_DESC_WR_MERGE_TIMEOUT_MAX);
+	nbl_hw_wr_regs_lock(hw_mgt, NBL_DVN_DESC_WR_MERGE_TIMEOUT, &timeout,
+			    sizeof(timeout));
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_DVN_DIF_REQ_RD_RO_FLAG, &ro_flag,
+		       sizeof(ro_flag));
+	if (pcie_relaxed_ordering_enabled(hw_mgt->common->pdev)) {
+		ro_flag |=
+			FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DESC_RO_EN_MASK,
+				   1) |
+			FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DATA_RO_EN_MASK,
+				   1) |
+			FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_AVRING_RO_EN_MASK,
+				   1);
+	} else {
+		ro_flag &=
+			~(FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DESC_RO_EN_MASK,
+				     1) |
+			FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DATA_RO_EN_MASK,
+				   1) |
+			FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_AVRING_RO_EN_MASK,
+				   1));
+	}
+	nbl_hw_wr_regs(hw_mgt, NBL_DVN_DIF_REQ_RD_RO_FLAG, &ro_flag,
+		       sizeof(ro_flag));
+	spin_unlock(&hw_mgt->reg_lock);
+	if (speed == NBL_FW_PORT_SPEED_100G)
+		nbl_dvn_descreq_num_cfg(hw_mgt,
+					DEFAULT_DVN_100G_DESCREQ_NUMCFG);
+	else
+		nbl_dvn_descreq_num_cfg(hw_mgt, DEFAULT_DVN_DESCREQ_NUMCFG);
+}
+
+static void nbl_uvn_init(struct nbl_hw_mgt *hw_mgt)
+{
+	u16 wr_timeout = NBL_UVN_DESC_WR_TIMEOUT_VAL;
+	u32 timeout = NBL_UVN_DESC_RD_WAIT_TICKS;
+	u32 prefetch_init = 0;
+	bool ro_enabled;
+	u32 flag = 0;
+	u32 mask = 0;
+	u32 quirks;
+	u32 reg_val;
+
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_wr32(hw_mgt, NBL_UVN_ECPU_QUEUE_NUM, 0);
+	nbl_hw_wr32(hw_mgt, NBL_UVN_DESC_RD_WAIT, timeout);
+	nbl_hw_rd_regs(hw_mgt, NBL_UVN_DESC_WR_TIMEOUT,
+		       &reg_val, sizeof(reg_val));
+	reg_val &= ~NBL_UVN_DESC_WR_TIMEOUT_NUM_MASK;
+	reg_val |= FIELD_PREP(NBL_UVN_DESC_WR_TIMEOUT_NUM_MASK, wr_timeout);
+	nbl_hw_wr_regs(hw_mgt, NBL_UVN_DESC_WR_TIMEOUT, &reg_val,
+		       sizeof(reg_val));
+	ro_enabled = pcie_relaxed_ordering_enabled(hw_mgt->common->pdev);
+
+	nbl_hw_rd_regs(hw_mgt, NBL_UVN_DIF_REQ_RO_FLAG, &flag, sizeof(flag));
+	if (ro_enabled) {
+		flag |= FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_AVAIL_RD_MASK, 1) |
+		FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_RD_MASK, 1) |
+		FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_PKT_WR_MASK, 1);
+		flag &= ~FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_WR_MASK, 1);
+	} else {
+		flag &= ~(FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_AVAIL_RD_MASK, 1) |
+			  FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_RD_MASK, 1) |
+			  FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_PKT_WR_MASK, 1) |
+			  FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_WR_MASK, 1));
+	}
+	nbl_hw_wr_regs(hw_mgt, NBL_UVN_DIF_REQ_RO_FLAG, &flag, sizeof(flag));
+
+	nbl_hw_rd_regs(hw_mgt, NBL_UVN_QUEUE_ERR_MASK, &mask, sizeof(mask));
+	mask |= FIELD_PREP(NBL_UVN_QUEUE_ERR_MASK_DIF_ERR_MASK, 1);
+
+	nbl_hw_wr_regs(hw_mgt, NBL_UVN_QUEUE_ERR_MASK, &mask, sizeof(mask));
+
+	spin_unlock(&hw_mgt->reg_lock);
+	quirks = nbl_hw_get_quirks(hw_mgt);
+	/*
+	 * sel=0: use configured num; sel=1: use internal calc (max 32)
+	 * Default is sel=1, unless NBL_QUIRK_UVN_PREFETCH_ALIGN is set,
+	 * in which case override to sel=0.
+	 */
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_rd_regs(hw_mgt, NBL_UVN_DESC_PREFETCH_INIT, &reg_val,
+		       sizeof(reg_val));
+	prefetch_init =
+		FIELD_PREP(NBL_UVN_DESC_PREFETCH_INIT_NUM_MASK,
+			   NBL_UVN_DESC_PREFETCH_NUM) |
+		FIELD_PREP(NBL_UVN_DESC_PREFETCH_INIT_SEL_MASK,
+			   (quirks & NBL_QUIRK_UVN_PREFETCH_ALIGN) ? 0 : 1);
+	reg_val &= ~(NBL_UVN_DESC_PREFETCH_INIT_NUM_MASK |
+		NBL_UVN_DESC_PREFETCH_INIT_SEL_MASK);
+	reg_val |= prefetch_init;
+	nbl_hw_wr_regs(hw_mgt, NBL_UVN_DESC_PREFETCH_INIT, &reg_val,
+		       sizeof(reg_val));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_uqm_init(struct nbl_hw_mgt *hw_mgt)
+{
+	u32 que_type = 0;
+	u32 cnt = 0;
+	int i;
+
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_wr_regs(hw_mgt, NBL_UQM_FWD_DROP_CNT, &cnt, sizeof(cnt));
+
+	nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_PKT_CNT, &cnt, sizeof(cnt));
+	nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_PKT_SLICE_CNT, &cnt, sizeof(cnt));
+	nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_PKT_LEN_ADD_CNT, &cnt, sizeof(cnt));
+	nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_HEAD_PNTR_ADD_CNT, &cnt,
+		       sizeof(cnt));
+	nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_WEIGHT_ADD_CNT, &cnt, sizeof(cnt));
+
+	for (i = 0; i < NBL_UQM_PORT_DROP_DEPTH; i++) {
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_UQM_PORT_DROP_PKT_CNT + (sizeof(cnt) * i),
+			       &cnt, sizeof(cnt));
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_UQM_PORT_DROP_PKT_SLICE_CNT +
+				       (sizeof(cnt) * i),
+			       &cnt, sizeof(cnt));
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_UQM_PORT_DROP_PKT_LEN_ADD_CNT +
+				       (sizeof(cnt) * i),
+			       &cnt, sizeof(cnt));
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_UQM_PORT_DROP_HEAD_PNTR_ADD_CNT +
+				       (sizeof(cnt) * i),
+			       &cnt, sizeof(cnt));
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_UQM_PORT_DROP_WEIGHT_ADD_CNT +
+				       (sizeof(cnt) * i),
+			       &cnt, sizeof(cnt));
+	}
+
+	for (i = 0; i < NBL_UQM_DPORT_DROP_DEPTH; i++)
+		nbl_hw_wr_regs(hw_mgt,
+			       NBL_UQM_DPORT_DROP_CNT + (sizeof(cnt) * i), &cnt,
+			       sizeof(cnt));
+	/* bit0: 0=bp mode, 1=drop mode, resv bit1-31 */
+	nbl_hw_wr_regs(hw_mgt, NBL_UQM_QUE_TYPE, &que_type, sizeof(que_type));
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static int nbl_dp_init(struct nbl_hw_mgt *hw_mgt, u8 speed, u8 eth_num)
+{
+	int ret;
+
+	nbl_dped_init(hw_mgt);
+	nbl_uped_init(hw_mgt);
+	ret = nbl_shaping_init(hw_mgt, speed);
+	if (ret)
+		return ret;
+	nbl_dsch_qid_max_init(hw_mgt);
+	ret = nbl_ustore_init(hw_mgt, eth_num);
+	if (ret)
+		return ret;
+	nbl_dstore_init(hw_mgt, speed);
+	nbl_dvn_init(hw_mgt, speed);
+	nbl_uvn_init(hw_mgt);
+	nbl_uqm_init(hw_mgt);
+	return 0;
+}
+
+static void nbl_host_padpt_init(struct nbl_hw_mgt *hw_mgt)
+{
+	/* padpt flow  control register */
+	spin_lock(&hw_mgt->reg_lock);
+	nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_CPLH_UP,
+		    NBL_HOST_PADPT_CFG_FC_CPLH_UP_VAL);
+	nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_PD_DN,
+		    NBL_HOST_PADPT_CFG_FC_PD_DN_VAL);
+	nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_PH_DN,
+		    NBL_HOST_PADPT_CFG_FC_PH_DN_VAL);
+	nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_NPH_DN,
+		    NBL_HOST_PADPT_CFG_FC_NPH_DN_VAL);
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_intf_init(struct nbl_hw_mgt *hw_mgt)
+{
+	nbl_host_padpt_init(hw_mgt);
+}
+
+static void nbl_hw_set_driver_status(struct nbl_hw_mgt *hw_mgt, bool active)
+{
+	u32 status;
+
+	spin_lock(&hw_mgt->reg_lock);
+	status = nbl_hw_rd32(hw_mgt, NBL_DRIVER_STATUS_REG);
+
+	status &= ~BIT(NBL_DRIVER_STATUS_BIT);
+	status |= FIELD_PREP(BIT(NBL_DRIVER_STATUS_BIT), active);
+
+	nbl_hw_wr32(hw_mgt, NBL_DRIVER_STATUS_REG, status);
+	spin_unlock(&hw_mgt->reg_lock);
+}
+
+/*
+ * Setting driver status to false notifies firmware to clean up per-PF
+ * hardware state such as qinfo registers.
+ *
+ * Note: firmware does NOT automatically revert chip-wide registers
+ * configured in this init flow. Those chip-wide settings remain valid
+ * until chip reset or explicitly overwritten by driver.
+ *
+ * This deinit_module only clears driver active status and flush writes.
+ * It does NOT reset or restore chip-wide datapath registers.
+ *
+ * Firmware cleanup is asynchronous with no completion status register,
+ * so a bounded sync delay follows the flush. On return the firmware
+ * pass has settled, so a subsequent init_module() cannot race it.
+ *
+ * Caller must ensure no new DMA is initiated after this point.
+ * The mailbox channel is stopped by nbl_chan_teardown_queue()
+ * before this function is called, so no in-flight mailbox DMA
+ * remains.
+ */
+static void nbl_hw_deinit_module(struct nbl_hw_mgt *hw_mgt)
+{
+	nbl_hw_set_driver_status(hw_mgt, false);
+	/* ensure driver_status reaches the chip */
+	nbl_flush_writes(hw_mgt);
+	/*
+	 * Allow the asynchronous firmware cleanup pass to finish before
+	 * returning, so this function is the boundary between firmware
+	 * per-PF teardown and any later driver reprogramming.
+	 */
+	usleep_range(NBL_FW_CLEANUP_SYNC_MIN_US, NBL_FW_CLEANUP_SYNC_MAX_US);
+}
+
+static bool nbl_hw_eth_speed_valid(u8 speed)
+{
+	switch (speed) {
+	case NBL_FW_PORT_SPEED_10G:
+	case NBL_FW_PORT_SPEED_25G:
+	case NBL_FW_PORT_SPEED_50G:
+	case NBL_FW_PORT_SPEED_100G:
+		return true;
+	default:
+		return false;
+	}
+}
+
+static bool nbl_hw_eth_num_valid(u8 eth_num)
+{
+	return eth_num == 1 || eth_num == 2 || eth_num == 4;
+}
+
+/*
+ * Full chip hardware initialization is handled by firmware.
+ * This function only configures driver-level table entries and registers.
+ */
+static int nbl_hw_init_module(struct nbl_hw_mgt *hw_mgt, u8 eth_speed,
+			      u8 eth_num)
+{
+	int ret;
+
+	if (!nbl_hw_eth_speed_valid(eth_speed)) {
+		dev_err(hw_mgt->common->dev, "Invalid eth_speed %u\n",
+			eth_speed);
+		return -EINVAL;
+	}
+	if (!nbl_hw_eth_num_valid(eth_num)) {
+		dev_err(hw_mgt->common->dev, "Invalid eth_num %u\n", eth_num);
+		return -EINVAL;
+	}
+
+	ret = nbl_dp_init(hw_mgt, eth_speed, eth_num);
+	if (ret)
+		return ret;
+	nbl_intf_init(hw_mgt);
+	nbl_hw_set_driver_status(hw_mgt, true);
+	/* ensure registers written */
+	nbl_flush_writes(hw_mgt);
+
+	return 0;
+}
+
 /*
  * nbl_hw_set_mailbox_irq - read-modify-write NBL_MAILBOX_QINFO_MAP_REG_ARR
  *
@@ -405,6 +1047,9 @@ static void nbl_hw_get_board_info(struct nbl_hw_mgt *hw_mgt,
 }
 
 static struct nbl_hw_ops hw_ops = {
+	.init_module = nbl_hw_init_module,
+	.deinit_module = nbl_hw_deinit_module,
+
 	.cfg_msix_map = nbl_hw_cfg_msix_map,
 	.cfg_msix_info = nbl_hw_cfg_msix_info,
 	.flush_write = nbl_flush_writes,
@@ -455,7 +1100,8 @@ static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
 	    !hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
 	    !hw_ops.get_host_pf_mask || !hw_ops.get_real_bus ||
 	    !hw_ops.cfg_mailbox_qinfo || !hw_ops.set_mailbox_irq ||
-	    !hw_ops.get_fw_eth_map || !hw_ops.get_board_info)
+	    !hw_ops.get_fw_eth_map || !hw_ops.get_board_info ||
+	    !hw_ops.init_module || !hw_ops.deinit_module)
 		return ERR_PTR(-EINVAL);
 	hw_ops_tbl->ops = &hw_ops;
 	hw_ops_tbl->priv = hw_mgt;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index 5ca74b63ef42..396ba910ccab 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -11,12 +11,25 @@
 #include "../../nbl_include/nbl_include.h"
 #include "../nbl_hw_reg.h"
 
+#define NBL_DRIVER_STATUS_REG			0x1300444
+#define NBL_DRIVER_STATUS_BIT			16
+
 /*  ----------  REG BASE ADDR  ----------  */
 /* Interface modules base addr */
 #define NBL_INTF_HOST_PCOMPLETER_BASE		0x00f08000
 #define NBL_INTF_HOST_PADPT_BASE		0x00f4c000
 #define NBL_INTF_HOST_MAILBOX_BASE		0x00fb0000
 #define NBL_INTF_HOST_PCIE_BASE			0X01504000
+/* DP modules base addr */
+#define NBL_DP_USTORE_BASE			0x00104000
+#define NBL_DP_UQM_BASE				0x00114000
+#define NBL_DP_UPED_BASE			0x0015c000
+#define NBL_DP_UVN_BASE				0x00244000
+#define NBL_DP_DSCH_BASE			0x00404000
+#define NBL_DP_SHAPING_BASE			0x00504000
+#define NBL_DP_DVN_BASE				0x00514000
+#define NBL_DP_DSTORE_BASE			0x00704000
+#define NBL_DP_DPED_BASE			0x0075c000
 /*  --------  MAILBOX BAR2 -----  */
 #define NBL_MAILBOX_NOTIFY_ADDR			0x00000000
 #define NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR	0x10
@@ -56,6 +69,17 @@ struct nbl_mailbox_qinfo_cfg_table {
 #define NBL_PCIE_BUS_MASK	GENMASK(12, 5)
 
 /*  --------  HOST_PADPT  --------  */
+#define NBL_HOST_PADPT_HOST_CFG_FC_PD_DN (NBL_INTF_HOST_PADPT_BASE + 0x00000160)
+#define NBL_HOST_PADPT_HOST_CFG_FC_PH_DN (NBL_INTF_HOST_PADPT_BASE + 0x00000164)
+#define NBL_HOST_PADPT_HOST_CFG_FC_NPH_DN \
+	(NBL_INTF_HOST_PADPT_BASE + 0x0000016C)
+#define NBL_HOST_PADPT_HOST_CFG_FC_CPLH_UP \
+	(NBL_INTF_HOST_PADPT_BASE + 0x00000170)
+
+#define NBL_HOST_PADPT_CFG_FC_CPLH_UP_VAL      0x10400
+#define NBL_HOST_PADPT_CFG_FC_PD_DN_VAL        0x10080
+#define NBL_HOST_PADPT_CFG_FC_PH_DN_VAL        0x10010
+#define NBL_HOST_PADPT_CFG_FC_NPH_DN_VAL       0x10010
 /* host_padpt host_msix_info */
 #define NBL_PADPT_HOST_MSIX_INFO_REG_ARR(vector_id) \
 	(NBL_INTF_HOST_PADPT_BASE + 0x00010000 +    \
@@ -96,6 +120,203 @@ struct nbl_function_msix_map {
 	u32 data[NBL_FUNC_MSIX_MAP_DWLEN];
 };
 
+/*  ----------  DPED  ----------  */
+#define NBL_DPED_VLAN_OFFSET		(NBL_DP_DPED_BASE + 0x000003F4)
+#define NBL_DPED_DSCP_OFFSET_0		(NBL_DP_DPED_BASE + 0x000003F8)
+#define NBL_DPED_DSCP_OFFSET_1		(NBL_DP_DPED_BASE + 0x000003FC)
+/* DPED hw_edt_prof/ UPED hw_edt_prof */
+
+#define NBL_DPED_L4_CK_CMD_40_ADDR	(NBL_DP_DPED_BASE + 0x00000338)
+
+#define NBL_DPED_L4_CK_CMD_40_EN_MASK		BIT(31)
+
+/*  ----------  UPED  ----------  */
+/* UPED uped_hw_edt_prof */
+#define NBL_UPED_HW_EDT_PROF_TABLE(i) \
+	(NBL_DP_UPED_BASE + 0x00001000 + (i) * sizeof(u32))
+#define NBL_UPED_V4_TCP_IDX		5
+#define NBL_UPED_V6_TCP_IDX		6
+#define NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK GENMASK(3, 2)
+
+/*  ----------  DSCH  ----------  */
+#define NBL_DSCH_PSHA_EN_MASK  GENMASK(3, 0)
+/* DSCH dsch maxqid */
+#define NBL_DSCH_HOST_QID_MAX (NBL_DP_DSCH_BASE + 0x00000118)
+#define NBL_DSCH_HOST_QID_MAX_MASK GENMASK(10, 0)
+#define NBL_DSCH_VN_QUANTA_ADDR (NBL_DP_DSCH_BASE + 0x00000134)
+
+#define NBL_MAX_QUEUE_ID	0x7ff
+#define NBL_HOST_QUANTA		0x8000
+#define NBL_ECPU_QUANTA		0x1000
+
+#define NBL_DSCH_VN_QUANTA_H_QUA_MASK GENMASK(15, 0)
+#define NBL_DSCH_VN_QUANTA_E_QUA_MASK GENMASK(31, 16)
+
+/*  ----------  DVN  ----------  */
+/* DVN dvn_queue_table */
+#define NBL_DVN_ECPU_QUEUE_NUM			(NBL_DP_DVN_BASE + 0x0000041C)
+#define NBL_DVN_DESCREQ_NUM_CFG			(NBL_DP_DVN_BASE + 0x00000430)
+#define NBL_DVN_DESC_WR_MERGE_TIMEOUT		(NBL_DP_DVN_BASE + 0x00000480)
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG		(NBL_DP_DVN_BASE + 0x0000045C)
+
+#define DEFAULT_DVN_DESCREQ_NUMCFG		0x03
+#define DEFAULT_DVN_100G_DESCREQ_NUMCFG		0x07
+
+#define DEFAULT_DVN_DESC_WR_MERGE_TIMEOUT_MAX	0x3FF
+
+/* spilit ring descreq_num 0:8,1:16 */
+#define NBL_DVN_DESCREQ_NUM_CFG_AVRING_DESREQ_NUM_CFG_MASK BIT(0)
+/* packet ring descreq_num
+ * 0:8,1:12,2:16;3:20,4:24,5:26;6:32,7:32
+ */
+#define NBL_DVN_DESCREQ_NUM_CFG_PACKED_L1_NUM_MASK GENMASK(6, 4)
+
+#define NBL_DVN_DESC_WR_MERGE_TIMEOUT_CFG_CYCLE_MASK GENMASK(9, 0)
+
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG_DESC_RO_EN_MASK BIT(0)
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG_DATA_RO_EN_MASK BIT(1)
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG_AVRING_RO_EN_MASK BIT(2)
+
+/*  ----------  UVN  ----------  */
+/* UVN uvn_queue_table */
+
+#define NBL_UVN_DESC_RD_WAIT			(NBL_DP_UVN_BASE + 0x0000020C)
+#define NBL_UVN_QUEUE_ERR_MASK			(NBL_DP_UVN_BASE + 0x00000224)
+#define NBL_UVN_ECPU_QUEUE_NUM			(NBL_DP_UVN_BASE + 0x0000023C)
+#define NBL_UVN_DESC_WR_TIMEOUT			(NBL_DP_UVN_BASE + 0x00000214)
+#define NBL_UVN_DIF_REQ_RO_FLAG			(NBL_DP_UVN_BASE + 0x00000250)
+#define NBL_UVN_DESC_PREFETCH_INIT		(NBL_DP_UVN_BASE + 0x00000204)
+#define NBL_UVN_DESC_PREFETCH_NUM		4
+
+#define NBL_UVN_DIF_REQ_RO_FLAG_AVAIL_RD_MASK BIT(0)
+#define NBL_UVN_DIF_REQ_RO_FLAG_DESC_RD_MASK BIT(1)
+#define NBL_UVN_DIF_REQ_RO_FLAG_PKT_WR_MASK BIT(2)
+#define NBL_UVN_DIF_REQ_RO_FLAG_DESC_WR_MASK BIT(3)
+
+#define NBL_UVN_DESC_WR_TIMEOUT_NUM_MASK GENMASK(14, 0)
+
+#define NBL_UVN_QUEUE_ERR_MASK_DIF_ERR_MASK BIT(5)
+
+#define NBL_UVN_DESC_PREFETCH_INIT_NUM_MASK GENMASK(7, 0)
+#define NBL_UVN_DESC_PREFETCH_INIT_SEL_MASK BIT(16)
+
+#define NBL_UVN_DESC_WR_TIMEOUT_VAL 0x12c
+/* 200us = 200000ns / 1.67ns per tick = 119760 ticks */
+#define NBL_UVN_DESC_RD_WAIT_TICKS 119760
+
+/*  --------  USTORE  --------  */
+#define NBL_USTORE_PKT_LEN_ADDR (NBL_DP_USTORE_BASE + 0x00000108)
+#define NBL_USTORE_PORT_DROP_TH_REG_ARR(port_id) \
+	(NBL_DP_USTORE_BASE + 0x00000150 + (port_id) * sizeof(u32))
+#define NBL_USTORE_BUF_PORT_DROP_PKT(eth_id) \
+	(NBL_DP_USTORE_BASE + 0x00002500 + (eth_id) * sizeof(u32))
+#define NBL_USTORE_BUF_PORT_TRUN_PKT(eth_id) \
+	(NBL_DP_USTORE_BASE + 0x00002540 + (eth_id) * sizeof(u32))
+
+#define NBL_USTORE_SINGLE_ETH_DROP_TH		0xC80
+#define NBL_USTORE_DUAL_ETH_DROP_TH		0x640
+#define NBL_USTORE_QUAD_ETH_DROP_TH		0x320
+
+/* USTORE pkt_len */
+#define NBL_USTORE_PKT_LEN_MIN_MASK GENMASK(6, 0)
+
+/* USTORE port_drop_th */
+#define NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK GENMASK(11, 0)
+#define NBL_USTORE_PORT_DROP_TH_EN_MASK BIT(31)
+
+/* UQM*/
+#define NBL_UQM_QUE_TYPE			(NBL_DP_UQM_BASE + 0x0000013c)
+#define NBL_UQM_DROP_PKT_CNT			(NBL_DP_UQM_BASE + 0x000009C0)
+#define NBL_UQM_DROP_PKT_SLICE_CNT		(NBL_DP_UQM_BASE + 0x000009C4)
+#define NBL_UQM_DROP_PKT_LEN_ADD_CNT		(NBL_DP_UQM_BASE + 0x000009C8)
+#define NBL_UQM_DROP_HEAD_PNTR_ADD_CNT		(NBL_DP_UQM_BASE + 0x000009CC)
+#define NBL_UQM_DROP_WEIGHT_ADD_CNT		(NBL_DP_UQM_BASE + 0x000009D0)
+#define NBL_UQM_PORT_DROP_PKT_CNT		(NBL_DP_UQM_BASE + 0x000009D4)
+#define NBL_UQM_PORT_DROP_PKT_SLICE_CNT		(NBL_DP_UQM_BASE + 0x000009F4)
+#define NBL_UQM_PORT_DROP_PKT_LEN_ADD_CNT	(NBL_DP_UQM_BASE + 0x00000A14)
+#define NBL_UQM_PORT_DROP_HEAD_PNTR_ADD_CNT	(NBL_DP_UQM_BASE + 0x00000A34)
+#define NBL_UQM_PORT_DROP_WEIGHT_ADD_CNT	(NBL_DP_UQM_BASE + 0x00000A54)
+#define NBL_UQM_FWD_DROP_CNT			(NBL_DP_UQM_BASE + 0x00000A80)
+#define NBL_UQM_DPORT_DROP_CNT			(NBL_DP_UQM_BASE + 0x00000B74)
+
+#define NBL_UQM_PORT_DROP_DEPTH			6
+#define NBL_UQM_DPORT_DROP_DEPTH		16
+
+/*  ---------  SHAPING  ---------  */
+
+/* Shaping rate unit: 1 = 1 Mbit/s.
+ * e.g. 100000 = 100 Gbit/s, 25000 = 25 Gbit/s.
+ */
+#define NBL_LR_LEONIS_NET_BUCKET_DEPTH		9600
+#define NBL_SHAPING_DPORT_ADDR (NBL_DP_SHAPING_BASE + 0x700)
+#define NBL_SHAPING_DPORT_DWLEN 4
+#define NBL_SHAPING_DPORT_REG(r) \
+	(NBL_SHAPING_DPORT_ADDR + (NBL_SHAPING_DPORT_DWLEN * 4) * (r))
+#define NBL_SHAPING_DVN_DPORT_ADDR (NBL_DP_SHAPING_BASE + 0x750)
+#define NBL_SHAPING_DVN_DPORT_DWLEN 4
+#define NBL_SHAPING_DVN_DPORT_REG(r) \
+	(NBL_SHAPING_DVN_DPORT_ADDR + (NBL_SHAPING_DVN_DPORT_DWLEN * 4) * (r))
+#define NBL_DSCH_PSHA_EN_ADDR (NBL_DP_DSCH_BASE + 0x00000314)
+#define NBL_SHAPING_NET_ADDR (NBL_DP_SHAPING_BASE + 0x1800)
+#define NBL_SHAPING_NET_DWLEN 4
+#define NBL_SHAPING_NET_REG(r) \
+	(NBL_SHAPING_NET_ADDR + (NBL_SHAPING_NET_DWLEN * 4) * (r))
+
+#define NBL_DPORT_VALID_MASK		GENMASK_ULL(0, 0)
+#define NBL_DPORT_DEPTH_MASK		GENMASK_ULL(19, 1)
+#define NBL_DPORT_CIR_MASK		GENMASK_ULL(38, 20)
+#define NBL_DPORT_PIR_MASK		GENMASK_ULL(57, 39)
+#define NBL_DPORT_CBS_MASK_LOW		GENMASK_ULL(63, 58)
+#define NBL_DPORT_CBS_MASK_HIGH		GENMASK_ULL(14, 0)
+#define NBL_DPORT_PBS_MASK		GENMASK_ULL(35, 15)
+
+/* SHAPING shaping_net */
+struct nbl_shaping_net_u {
+	u32 data[NBL_SHAPING_NET_DWLEN];
+};
+
+struct nbl_shaping_dport_u {
+	u32 data[NBL_SHAPING_DPORT_DWLEN];
+};
+
+struct nbl_shaping_dvn_dport_u {
+	u32 data[NBL_SHAPING_DVN_DPORT_DWLEN];
+};
+
+/*  --------  DSTORE  --------  */
+#define NBL_DSTORE_D_DPORT_FC_TH_ADDR  (NBL_DP_DSTORE_BASE + 0x00000600)
+#define NBL_DSTORE_D_DPORT_FC_TH_DEPTH 5
+#define NBL_DSTORE_D_DPORT_FC_TH_WIDTH 32
+#define NBL_DSTORE_D_DPORT_FC_TH_DWLEN 1
+
+#define NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK GENMASK(10, 0)
+#define NBL_DSTORE_D_DPORT_FC_XON_TH_MASK GENMASK(26, 16)
+#define NBL_DSTORE_D_DPORT_FC_FC_EN_MASK BIT(31)
+
+#define NBL_DSTORE_D_DPORT_FC_TH_REG(r)  \
+	(NBL_DSTORE_D_DPORT_FC_TH_ADDR + \
+	 (NBL_DSTORE_D_DPORT_FC_TH_DWLEN * 4) * (r))
+#define NBL_DSTORE_PORT_DROP_TH_ADDR (NBL_DP_DSTORE_BASE + 0x00000150)
+#define NBL_DSTORE_PORT_DROP_TH_DEPTH 6
+#define NBL_DSTORE_PORT_DROP_TH_WIDTH 32
+#define NBL_DSTORE_PORT_DROP_TH_DWLEN 1
+
+#define NBL_DSTORE_PORT_DROP_EN_MASK BIT(31)
+
+#define NBL_DSTORE_DROP_XOFF_TH			0xC8
+#define NBL_DSTORE_DROP_XON_TH			0x64
+
+#define NBL_DSTORE_DROP_XOFF_TH_100G		0x1F4
+#define NBL_DSTORE_DROP_XON_TH_100G		0x12C
+
+#define NBL_DSTORE_DISC_BP_TH (NBL_DP_DSTORE_BASE + 0x00000630)
+
+#define NBL_DSTORE_DISC_BP_TH_EN_MASK BIT(31)
+
+#define NBL_DSTORE_PORT_DROP_TH_REG(r)  \
+	(NBL_DSTORE_PORT_DROP_TH_ADDR + \
+	 (NBL_DSTORE_PORT_DROP_TH_DWLEN * 4) * (r))
+
 #define NBL_FW_BOARD_CONFIG			0x200
 #define NBL_FW_BOARD_DW3_OFFSET			(NBL_FW_BOARD_CONFIG + 12)
 #define NBL_FW_BOARD_DW6_OFFSET			(NBL_FW_BOARD_CONFIG + 24)
@@ -109,5 +330,6 @@ struct nbl_function_msix_map {
 #define NBL_FW_BOARD_DW6_LANE_BITMAP_MASK GENMASK(7, 0)
 #define NBL_FW_BOARD_DW6_ETH_BITMAP_MASK GENMASK(15, 8)
 
+#define NBL_LEONIS_QUIRKS_OFFSET	0x00000140
 #define NBL_BAR2_MAX_LEN		0x300
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
index 9b53e70ae4af..88a1478e858d 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
@@ -13,6 +13,8 @@ static struct nbl_resource_ops res_ops = {
 	.cfg_msix_map = nbl_res_intr_cfg_msix_map,
 	.destroy_msix_map = nbl_res_intr_destroy_msix_map,
 	.set_mailbox_irq = nbl_res_intr_set_mailbox_irq,
+	.init_module = nbl_res_chip_init_module,
+	.deinit_module = nbl_res_chip_deinit_module,
 };
 
 static struct nbl_resource_mgt *
@@ -46,7 +48,8 @@ nbl_res_setup_ops(struct device *dev, struct nbl_resource_mgt *res_mgt)
 		return ERR_PTR(-ENOMEM);
 	if (!res_ops.get_vsi_id || !res_ops.get_eth_id ||
 	    !res_ops.cfg_msix_map || !res_ops.destroy_msix_map ||
-	    !res_ops.set_mailbox_irq)
+	    !res_ops.set_mailbox_irq || !res_ops.init_module ||
+	    !res_ops.deinit_module)
 		return ERR_PTR(-EINVAL);
 	res_ops_tbl->ops = &res_ops;
 	res_ops_tbl->priv = res_mgt;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
index 6eb4dc9e695a..f1cb0f23240b 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
@@ -8,4 +8,5 @@
 
 #include "../nbl_resource.h"
 #include "../nbl_interrupt.h"
+#include "../nbl_chip.h"
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index 9a9642f52af2..6bd388c7cd43 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -12,6 +12,9 @@ struct nbl_board_port_info;
 struct nbl_hw_mgt;
 struct nbl_adapter;
 struct nbl_hw_ops {
+	int (*init_module)(struct nbl_hw_mgt *hw_mgt, u8 eth_speed, u8 eth_num);
+	void (*deinit_module)(struct nbl_hw_mgt *hw_mgt);
+
 	void (*cfg_msix_map)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
 			     bool valid, dma_addr_t dma_addr, u8 bus,
 			     u8 devid, u8 function);
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
index e718ea41a816..8dc64e806c1e 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
@@ -12,6 +12,9 @@ struct nbl_resource_mgt;
 struct nbl_adapter;
 
 struct nbl_resource_ops {
+	int (*init_module)(struct nbl_resource_mgt *res_mgt);
+	void (*deinit_module)(struct nbl_resource_mgt *res_mgt);
+
 	int (*cfg_msix_map)(struct nbl_resource_mgt *res_mgt, u16 func_id,
 			    u16 num_net_msix, u16 num_others_msix,
 			    bool net_msix_mask_en);
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
index 34abf20ebf5a..bb68085a7054 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
@@ -33,4 +33,23 @@ struct nbl_init_param {
 	struct nbl_func_caps caps;
 };
 
+/*
+ * Firmware ABI defines port speed enum fixed, value 0 represents 10G, cannot
+ * reassign 0 to INVALID for compatibility
+ */
+enum nbl_fw_port_speed {
+	NBL_FW_PORT_SPEED_10G,
+	NBL_FW_PORT_SPEED_25G,
+	NBL_FW_PORT_SPEED_50G,
+	NBL_FW_PORT_SPEED_100G,
+};
+
+/*
+ * Firmware quirk word @ NBL_LEONIS_QUIRKS_OFFSET (0x140)
+ * Sentinel value: ~0U (0xFFFFFFFF) = firmware reports no active quirks
+ * BIT(1): NBL_QUIRK_UVN_PREFETCH_ALIGN – control UVN descriptor prefetch
+ * selection
+ */
+#define NBL_QUIRK_UVN_PREFETCH_ALIGN	BIT(1)
+
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index 1aafed2d46d7..e1a30b5ba0cd 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -83,6 +83,87 @@ static void nbl_get_func_param(struct pci_dev *pdev, kernel_ulong_t driver_data,
 		param->caps.has_ctrl = 1;
 }
 
+/*
+ * Establish chip-wide dependencies for this PF:
+ *  - for every non-management PF, add a consumer->management PF device
+ *    link. The driver core then guarantees (sysfs unbind, driver
+ *    unregister, hot-unplug alike) that this consumer is released
+ *    BEFORE the func 0 supplier, which is the only teardown order in
+ *    which the chip-global firmware deinit is safe.
+ *
+ * The management PF is addressed by the deterministic identity
+ * (domain, bus, slot, func 0) instead of any name-based scan: hardware
+ * guarantees PFs are contiguous from func 0 in the same slot.
+ *
+ * The func 0 driver must already be fully bound. A managed link created
+ * from inside this probe while the supplier has no driver starts DORMANT
+ * and is force-promoted (with a WARN backtrace) when the probe ends, so
+ * it never delivers the unbind-order guarantee; a link created while
+ * func 0 is still probing can race its chip-global init and is not
+ * unwound if func 0's probe fails. Probing is deferred until func 0 is
+ * bound, per the consumer responsibility rule in
+ * Documentation/driver-api/device_link.rst.
+ *
+ * The link uses DL_FLAG_AUTOREMOVE_CONSUMER, so it is purged by the
+ * driver core when this PF fails to probe or later detaches; it must
+ * not be removed manually.
+ *
+ * Return: 0 on success, negative errno on failure.
+ */
+static int nbl_probe_chip_deps(struct pci_dev *pdev, bool has_ctrl)
+{
+	struct device_link *link;
+	struct pci_dev *mgt;
+
+	if (has_ctrl)
+		return 0;
+
+	mgt = pci_get_domain_bus_and_slot(pci_domain_nr(pdev->bus),
+					  pdev->bus->number,
+					  PCI_DEVFN(PCI_SLOT(pdev->devfn), 0));
+	if (!mgt) {
+		dev_err(&pdev->dev,
+			"management PF (func 0) not found on this chip\n");
+		return -ENODEV;
+	}
+
+	/*
+	 * The management PF must have COMPLETED probing, not merely
+	 * have a driver assigned.  While func 0 is still PROBING, its
+	 * chip-global init (mailbox QINFO map programming) can race a
+	 * sibling's mailbox setup, and if func 0 then fails its probe,
+	 * the core only downgrades an already-ACTIVE sibling link to
+	 * DORMANT without unbinding it, leaving an orphaned consumer.
+	 * Defer until func 0 is bound; this also covers the unbound
+	 * (DORMANT-link) case.
+	 */
+	if (READ_ONCE(mgt->dev.links.status) != DL_DEV_DRIVER_BOUND) {
+		pci_dev_put(mgt);
+		return -EPROBE_DEFER;
+	}
+
+	link = device_link_add(&pdev->dev, &mgt->dev,
+			       DL_FLAG_AUTOREMOVE_CONSUMER);
+	pci_dev_put(mgt);
+	if (!link) {
+		dev_err(&pdev->dev,
+			"failed to create device link to management PF\n");
+		return -ENOMEM;
+	}
+
+	/*
+	 * If func 0 began unbinding between the driver check above and
+	 * link creation, the core's unbind-consumers pass has already
+	 * run and missed this new link, which starts SUPPLIER_UNBIND.
+	 * Fail the probe: the managed link is auto-purged and the core
+	 * retries after func 0 is bound again.
+	 */
+	if (READ_ONCE(link->status) == DL_STATE_SUPPLIER_UNBIND)
+		return -EPROBE_DEFER;
+
+	return 0;
+}
+
 static int nbl_probe(struct pci_dev *pdev,
 		     const struct pci_device_id *id)
 {
@@ -103,6 +184,10 @@ static int nbl_probe(struct pci_dev *pdev,
 
 	pci_set_master(pdev);
 
+	err = nbl_probe_chip_deps(pdev, param.caps.has_ctrl);
+	if (err)
+		goto chip_deps_err;
+
 	adapter = nbl_core_init(pdev, &param);
 	if (IS_ERR(adapter)) {
 		dev_err(dev, "Nbl adapter init fail: %pe\n", adapter);
@@ -112,6 +197,7 @@ static int nbl_probe(struct pci_dev *pdev,
 	pci_set_drvdata(pdev, adapter);
 	return 0;
 adapter_init_err:
+chip_deps_err:
 	pci_clear_master(pdev);
 	return err;
 }
@@ -122,6 +208,7 @@ static void nbl_remove(struct pci_dev *pdev)
 
 	if (!adapter)
 		return;
+
 	pci_set_drvdata(pdev, NULL);
 	nbl_core_remove(adapter);
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
                   ` (3 preceding siblings ...)
  2026-09-28 12:32 ` [PATCH v30 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  2026-09-28 12:32 ` [PATCH v30 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops illusion.wang
                   ` (2 subsequent siblings)
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Allocate dispatch management state and ops table, add ctrl_lvl bitmap
to track control privileges, and hook init_module/deinit_module wrappers
to resource ops.

MGT level is only enabled for Control PF. Add kerneldoc for caller
permission constraints. Add wire ABI structures with static_assert for
upcoming mailbox RPC.

Dispatch objects are devm-managed and require no explicit remove logic.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../net/ethernet/nebula-matrix/nbl/Makefile   |   1 +
 .../net/ethernet/nebula-matrix/nbl/nbl_core.h |   2 +
 .../nebula-matrix/nbl/nbl_core/nbl_dispatch.c | 117 ++++++++++++++++++
 .../nebula-matrix/nbl/nbl_core/nbl_dispatch.h |  23 ++++
 .../nbl/nbl_include/nbl_def_dispatch.h        |  45 +++++++
 .../net/ethernet/nebula-matrix/nbl/nbl_main.c |   8 ++
 6 files changed, 196 insertions(+)
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index be314b909d66..b7eebd89b4d1 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -10,4 +10,5 @@ nbl-objs +=	nbl_common/nbl_common.o \
 		nbl_hw/nbl_resource.o \
 		nbl_hw/nbl_interrupt.o \
 		nbl_hw/nbl_chip.o \
+		nbl_core/nbl_dispatch.o \
 		nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index dd24ebec0171..4d8cea8d8ab3 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -17,12 +17,14 @@ enum {
 struct nbl_interface {
 	struct nbl_hw_ops_tbl *hw_ops_tbl;
 	struct nbl_resource_ops_tbl *resource_ops_tbl;
+	struct nbl_dispatch_ops_tbl *dispatch_ops_tbl;
 	struct nbl_channel_ops_tbl *channel_ops_tbl;
 };
 
 struct nbl_core {
 	struct nbl_hw_mgt *hw_mgt;
 	struct nbl_resource_mgt *res_mgt;
+	struct nbl_dispatch_mgt *disp_mgt;
 	struct nbl_channel_mgt *chan_mgt;
 };
 
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
new file mode 100644
index 000000000000..966fee2dec8b
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
@@ -0,0 +1,117 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/pci.h>
+#include "nbl_dispatch.h"
+
+static void nbl_disp_deinit_module(struct nbl_dispatch_mgt *disp_mgt)
+{
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+
+	if (res_ops->deinit_module)
+		res_ops->deinit_module(p);
+}
+
+static int nbl_disp_init_module(struct nbl_dispatch_mgt *disp_mgt)
+{
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+
+	if (res_ops->init_module)
+		return res_ops->init_module(p);
+	return -EOPNOTSUPP;
+}
+
+static void nbl_disp_set_ctrl_bit(struct nbl_dispatch_mgt *disp_mgt, u32 lvl)
+{
+	set_bit(lvl, disp_mgt->ctrl_lvl);
+}
+
+static void nbl_disp_refresh_ctrl_ops(struct nbl_dispatch_mgt *disp_mgt)
+{
+	struct nbl_dispatch_ops *disp_ops = disp_mgt->disp_ops_tbl->ops;
+
+	if (test_bit(NBL_DISP_CTRL_LVL_MGT, disp_mgt->ctrl_lvl)) {
+		disp_ops->init_module = nbl_disp_init_module;
+		disp_ops->deinit_module = nbl_disp_deinit_module;
+	}
+}
+
+static struct nbl_dispatch_mgt *
+nbl_disp_setup_disp_mgt(struct nbl_common_info *common)
+{
+	struct nbl_dispatch_mgt *disp_mgt;
+	struct device *dev = common->dev;
+
+	disp_mgt = devm_kzalloc(dev, sizeof(*disp_mgt), GFP_KERNEL);
+	if (!disp_mgt)
+		return ERR_PTR(-ENOMEM);
+
+	disp_mgt->common = common;
+	return disp_mgt;
+}
+
+static struct nbl_dispatch_ops_tbl *
+nbl_disp_setup_ops(struct device *dev, struct nbl_dispatch_mgt *disp_mgt)
+{
+	struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+	struct nbl_dispatch_ops *disp_ops;
+
+	disp_ops_tbl = devm_kzalloc(dev, sizeof(*disp_ops_tbl), GFP_KERNEL);
+	if (!disp_ops_tbl)
+		return ERR_PTR(-ENOMEM);
+
+	disp_ops = devm_kzalloc(dev, sizeof(*disp_ops), GFP_KERNEL);
+	if (!disp_ops)
+		return ERR_PTR(-ENOMEM);
+
+	disp_ops_tbl->ops = disp_ops;
+	disp_ops_tbl->priv = disp_mgt;
+
+	return disp_ops_tbl;
+}
+
+int nbl_disp_init(struct nbl_adapter *adapter)
+{
+	struct nbl_common_info *common = &adapter->common;
+	struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+	struct nbl_resource_ops_tbl *res_ops_tbl =
+		adapter->intf.resource_ops_tbl;
+	struct nbl_channel_ops_tbl *chan_ops_tbl =
+		adapter->intf.channel_ops_tbl;
+	struct device *dev = &adapter->pdev->dev;
+	struct nbl_dispatch_mgt *disp_mgt;
+	int ret;
+
+	disp_mgt = nbl_disp_setup_disp_mgt(common);
+	if (IS_ERR(disp_mgt)) {
+		ret = PTR_ERR(disp_mgt);
+		return ret;
+	}
+
+	disp_ops_tbl = nbl_disp_setup_ops(dev, disp_mgt);
+	if (IS_ERR(disp_ops_tbl)) {
+		ret = PTR_ERR(disp_ops_tbl);
+		return ret;
+	}
+
+	disp_mgt->res_ops_tbl = res_ops_tbl;
+	disp_mgt->chan_ops_tbl = chan_ops_tbl;
+	disp_mgt->disp_ops_tbl = disp_ops_tbl;
+	adapter->core.disp_mgt = disp_mgt;
+	adapter->intf.dispatch_ops_tbl = disp_ops_tbl;
+
+	if (common->has_ctrl)
+		nbl_disp_set_ctrl_bit(disp_mgt, NBL_DISP_CTRL_LVL_MGT);
+
+	nbl_disp_refresh_ctrl_ops(disp_mgt);
+	return 0;
+}
+
+void nbl_disp_remove(struct nbl_adapter *adapter)
+{
+	/* Dispatch structures are allocated via devm */
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
new file mode 100644
index 000000000000..a7e5802344b4
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
@@ -0,0 +1,23 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DISPATCH_H_
+#define _NBL_DISPATCH_H_
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_resource.h"
+#include "../nbl_include/nbl_def_dispatch.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+struct nbl_dispatch_mgt {
+	struct nbl_common_info *common;
+	struct nbl_resource_ops_tbl *res_ops_tbl;
+	struct nbl_channel_ops_tbl *chan_ops_tbl;
+	struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+	DECLARE_BITMAP(ctrl_lvl, NBL_DISP_CTRL_LVL_MAX);
+};
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
new file mode 100644
index 000000000000..b3398591035b
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
@@ -0,0 +1,45 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_DISPATCH_H_
+#define _NBL_DEF_DISPATCH_H_
+
+#include <linux/types.h>
+
+struct nbl_dispatch_mgt;
+struct nbl_adapter;
+enum {
+	NBL_DISP_CTRL_LVL_MGT,
+	NBL_DISP_CTRL_LVL_MAX,
+};
+
+/**
+ * struct nbl_dispatch_ops - dispatch control plane operation callbacks
+ * @init_module: dispatch layer initialization, control-PF exclusive,
+ *               caller must check has_ctrl guard
+ * @deinit_module: dispatch layer cleanup, control-PF exclusive,
+ *                 caller must check has_ctrl guard
+ *
+ * Warning: init_module/deinit_module are control-PF exclusive. The five
+ * resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq,
+ * get_vsi_id, get_eth_id) are PF-only and resolve to either a local
+ * resource call (control PF) or a mailbox RPC (non-control PF with
+ * has_net). A function with neither has_ctrl nor has_net leaves these
+ * pointers NULL; callers must not invoke them on such functions. VFs are
+ * rejected by the responders with -EPERM.
+ */
+struct nbl_dispatch_ops {
+	int (*init_module)(struct nbl_dispatch_mgt *disp_mgt);
+	void (*deinit_module)(struct nbl_dispatch_mgt *disp_mgt);
+};
+
+struct nbl_dispatch_ops_tbl {
+	struct nbl_dispatch_ops *ops;
+	struct nbl_dispatch_mgt *priv;
+};
+
+int nbl_disp_init(struct nbl_adapter *adapter);
+void nbl_disp_remove(struct nbl_adapter *adapter);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index e1a30b5ba0cd..bf4a5ea3e410 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -11,6 +11,7 @@
 #include "nbl_include/nbl_def_channel.h"
 #include "nbl_include/nbl_def_hw.h"
 #include "nbl_include/nbl_def_resource.h"
+#include "nbl_include/nbl_def_dispatch.h"
 #include "nbl_include/nbl_def_common.h"
 #include "nbl_core.h"
 
@@ -48,7 +49,13 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 	ret = nbl_res_init_leonis(adapter);
 	if (ret)
 		goto res_init_fail;
+
+	ret = nbl_disp_init(adapter);
+	if (ret)
+		goto disp_init_fail;
 	return adapter;
+disp_init_fail:
+	nbl_res_remove_leonis(adapter);
 res_init_fail:
 	nbl_chan_remove_common(adapter);
 chan_init_fail:
@@ -59,6 +66,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 
 void nbl_core_remove(struct nbl_adapter *adapter)
 {
+	nbl_disp_remove(adapter);
 	nbl_res_remove_leonis(adapter);
 	nbl_chan_remove_common(adapter);
 	nbl_hw_remove_leonis(adapter);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
                   ` (4 preceding siblings ...)
  2026-09-28 12:32 ` [PATCH v30 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  2026-09-28 12:32 ` [PATCH v30 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
  2026-09-28 12:32 ` [PATCH v30 net-next 8/8] net/nebula-matrix: add common dev start/stop operation illusion.wang
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Implement bidirectional mailbox RPC framework for cross-PF resource management,
adding request/response handlers for five core dispatch operations:
configure_msix_map, destroy_msix_map, set_mailbox_irq, get_vsi_id, get_eth_id.

Dispatch operations are dynamically resolved based on PF capability level:
- Management PF: invokes local hardware resource ops directly.
- Net-capable non-management PF: forwards operations to manager PF via RPC.
- Other functions: all dispatch ops are null and return -EOPNOTSUPP.

Add per-dispatch mutex to serialize mutable hardware operations including
MSIX map configuration/destruction and mailbox IRQ setup. The mutex covers
both local PF call paths and remote RPC response handlers to eliminate
concurrent modification race conditions. Read-only VSI/ETH ID lookup APIs
access static init metadata without serialization requirements.

Register five non-contiguous mailbox message handlers for the new RPC
operations, preserving backward compatibility with reserved message IDs.
All responders enforce strict runtime validation:
- Reject VF/out-of-range PF IDs with -EPERM;
- Validate incoming payload length before parsing;
- Check function pointer existence before invoking resource ops.

Fix error propagation by forwarding native Linux errnos from remote
resource operations instead of unconditionally returning -EREMOTEIO.
Only truncated ACK responses return -EREMOTEIO to distinguish protocol
errors from legitimate operation failures. Transport layer errors are
passed through unchanged from channel send routines.

For MSIX map reconfiguration, pre-allocate DMA buffers and vector indices
prior to tearing down old configurations to prevent interrupt loss.
Destroy path clears mailbox MSIX routing entries upfront to avoid stale
interrupt triggers from recycled hardware vectors.

Document RPC caller preconditions for MSIX map operations: disable mailbox
IRQ_RDY flag and use polling send path, as responders modify mailbox MSIX
routing and cannot rely on interrupt-based ACK wakeup. Extend the same
implicit precondition to set_mailbox_irq for future caller safety.

Clarify module teardown semantics in comment: dispatch message handlers
are owned by the channel layer and unregistered in nbl_chan_remove_common(),
with strict stop ordering preventing use-after-free or stale execution.
Narrow kernel-doc warning to clarify new ops are PF-capability bounded,
not available for VF functions.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../nebula-matrix/nbl/nbl_core/nbl_dispatch.c | 565 +++++++++++++++++-
 .../nebula-matrix/nbl/nbl_core/nbl_dispatch.h |   2 +
 .../nbl/nbl_include/nbl_def_channel.h         |  39 ++
 .../nbl/nbl_include/nbl_def_dispatch.h        |  16 +
 4 files changed, 621 insertions(+), 1 deletion(-)

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
index 966fee2dec8b..e6a5e53e9d2d 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
@@ -3,9 +3,178 @@
  * Copyright (c) 2026 Nebula Matrix Limited.
  */
 #include <linux/device.h>
+#include <linux/mutex.h>
 #include <linux/pci.h>
 #include "nbl_dispatch.h"
 
+static int nbl_disp_chan_get_vsi_id_req(struct nbl_dispatch_mgt *disp_mgt,
+					u16 type, u16 *vsi_id)
+{
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_common_info *common = disp_mgt->common;
+	struct nbl_chan_param_get_vsi_id result = { 0 };
+	struct nbl_chan_param_get_vsi_id param = { 0 };
+	struct nbl_chan_send_info chan_send = {0};
+	int ret;
+
+	param.type = cpu_to_le16(type);
+
+	nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+				NBL_CHAN_MSG_GET_VSI_ID,
+				&param, sizeof(param), &result,
+				sizeof(result), 1);
+	ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+	if (ret)
+		return ret;
+	if (chan_send.ack_len != sizeof(result)) {
+		dev_err(disp_mgt->common->dev,
+			"get_vsi_id: short ACK, ack_len=%u expected %zu\n",
+			chan_send.ack_len, sizeof(result));
+		return -EBADMSG;
+	}
+	*vsi_id = le16_to_cpu(result.vsi_id);
+	return 0;
+}
+
+static void nbl_disp_chan_get_vsi_id_resp(void *priv, u16 src_id, u16 msg_id,
+					  void *data, u32 data_len)
+{
+	struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct device *dev = disp_mgt->common->dev;
+	struct nbl_chan_param_get_vsi_id result = { 0 };
+	struct nbl_chan_param_get_vsi_id param = { 0 };
+	struct nbl_chan_ack_info chan_ack;
+	int err = 0;
+	u16 vsi_id = 0;
+	u32 rel_pf_id;
+	int ret;
+
+	ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+					      &rel_pf_id);
+	if (ret) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (rel_pf_id >= disp_mgt->common->max_pf) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (data_len < sizeof(param)) {
+		err = -EBADMSG;
+		goto ack_out;
+	}
+	memcpy(&param, data, sizeof(param));
+
+	if (res_ops->get_vsi_id) {
+		ret = res_ops->get_vsi_id(p, src_id, le16_to_cpu(param.type),
+					  &vsi_id);
+		/* Forward the resource errno verbatim to the requester */
+		if (ret)
+			err = ret;
+	} else {
+		err = -EOPNOTSUPP;
+	}
+
+	result.vsi_id = cpu_to_le16(vsi_id);
+ack_out:
+	nbl_chan_fill_ack_info(&chan_ack, src_id,
+			       NBL_CHAN_MSG_GET_VSI_ID, msg_id, err,
+			       &result, sizeof(result));
+	ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+	if (ret)
+		dev_err(dev,
+			"channel send ack failed with ret: %d, msg_type: %d\n",
+			ret, NBL_CHAN_MSG_GET_VSI_ID);
+}
+
+static int nbl_disp_chan_get_eth_id_req(struct nbl_dispatch_mgt *disp_mgt,
+					u16 vsi_id, u8 *eth_num, u8 *eth_id,
+					u8 *logic_eth_id)
+{
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_common_info *common = disp_mgt->common;
+	struct nbl_chan_param_get_eth_id result = { 0 };
+	struct nbl_chan_param_get_eth_id param = { 0 };
+	struct nbl_chan_send_info chan_send = {0};
+	int ret;
+
+	param.vsi_id = cpu_to_le16(vsi_id);
+
+	nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+				NBL_CHAN_MSG_GET_ETH_ID,
+				&param, sizeof(param), &result,
+				sizeof(result), 1);
+	ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+	if (ret)
+		return ret;
+	if (chan_send.ack_len != sizeof(result)) {
+		dev_err(disp_mgt->common->dev,
+			"get_eth_id: short ACK, ack_len=%u expected %zu\n",
+			chan_send.ack_len, sizeof(result));
+		return -EBADMSG;
+	}
+	*eth_num = result.eth_num;
+	*eth_id = result.eth_id;
+	*logic_eth_id = result.logic_eth_id;
+
+	return 0;
+}
+
+static void nbl_disp_chan_get_eth_id_resp(void *priv, u16 src_id, u16 msg_id,
+					  void *data, u32 data_len)
+{
+	struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct nbl_chan_param_get_eth_id result = { 0 };
+	struct nbl_chan_param_get_eth_id param = { 0 };
+	struct device *dev = disp_mgt->common->dev;
+	struct nbl_chan_ack_info chan_ack;
+	int err = 0;
+	u32 rel_pf_id;
+	int ret;
+
+	ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+					      &rel_pf_id);
+	if (ret) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (rel_pf_id >= disp_mgt->common->max_pf) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (data_len < sizeof(param)) {
+		err = -EBADMSG;
+		goto ack_out;
+	}
+	memcpy(&param, data, sizeof(param));
+
+	if (res_ops->get_eth_id) {
+		ret = res_ops->get_eth_id(p, src_id, le16_to_cpu(param.vsi_id),
+					  &result.eth_num, &result.eth_id,
+					  &result.logic_eth_id);
+		/* Forward the resource errno verbatim to the requester */
+		if (ret)
+			err = ret;
+	} else {
+		err = -EOPNOTSUPP;
+	}
+ack_out:
+	nbl_chan_fill_ack_info(&chan_ack, src_id,
+			       NBL_CHAN_MSG_GET_ETH_ID, msg_id, err,
+			       &result, sizeof(result));
+	ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+	if (ret)
+		dev_err(dev,
+			"channel send ack failed with ret: %d, msg_type: %d\n",
+			ret, NBL_CHAN_MSG_GET_ETH_ID);
+}
+
 static void nbl_disp_deinit_module(struct nbl_dispatch_mgt *disp_mgt)
 {
 	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
@@ -25,6 +194,357 @@ static int nbl_disp_init_module(struct nbl_dispatch_mgt *disp_mgt)
 	return -EOPNOTSUPP;
 }
 
+static int nbl_disp_cfg_msix_map(struct nbl_dispatch_mgt *disp_mgt,
+				 u16 num_net_msix, u16 num_others_msix,
+				 bool net_msix_mask_en)
+{
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct nbl_common_info *common = disp_mgt->common;
+	int ret;
+
+	if (!res_ops->cfg_msix_map)
+		return -EOPNOTSUPP;
+	mutex_lock(&disp_mgt->ops_mutex_lock);
+	ret = res_ops->cfg_msix_map(p, common->mgt_pf, num_net_msix,
+					  num_others_msix, net_msix_mask_en);
+	mutex_unlock(&disp_mgt->ops_mutex_lock);
+	return ret;
+}
+
+/*
+ * Precondition: caller must disable mailbox IRQ_RDY and switch send_msg
+ * to the polling path before issuing this RPC. The responder rewrites
+ * the requester's mailbox MSI-X routing during the resource op, so the
+ * ACK cannot rely on interrupt wakeup while routing is in flux.
+ */
+static int
+nbl_disp_chan_cfg_msix_map_req(struct nbl_dispatch_mgt *disp_mgt,
+			       u16 num_net_msix, u16 num_others_msix,
+			       bool net_msix_mask_en)
+{
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_common_info *common = disp_mgt->common;
+	struct nbl_chan_param_cfg_msix_map param = { 0 };
+	struct nbl_chan_send_info chan_send = {0};
+	int ret;
+
+	param.num_net_msix = cpu_to_le16(num_net_msix);
+	param.num_others_msix = cpu_to_le16(num_others_msix);
+	param.msix_mask_en = cpu_to_le16(!!net_msix_mask_en);
+
+	nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+				NBL_CHAN_MSG_CONFIGURE_MSIX_MAP,
+				&param, sizeof(param),
+				NULL, 0, 1);
+	ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+	if (ret)
+		return ret;
+	return 0;
+}
+
+static void nbl_disp_chan_cfg_msix_map_resp(void *priv, u16 src_id, u16 msg_id,
+					    void *data, u32 data_len)
+{
+	struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct device *dev = disp_mgt->common->dev;
+	struct nbl_chan_param_cfg_msix_map param = { 0 };
+	struct nbl_chan_ack_info chan_ack;
+	int err = 0;
+	u32 rel_pf_id;
+	int ret;
+
+	ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+					      &rel_pf_id);
+	if (ret) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (rel_pf_id >= disp_mgt->common->max_pf) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (data_len < sizeof(param)) {
+		err = -EBADMSG;
+		goto ack_out;
+	}
+	memcpy(&param, data, sizeof(param));
+
+	if (res_ops->cfg_msix_map) {
+		mutex_lock(&disp_mgt->ops_mutex_lock);
+		ret = res_ops->cfg_msix_map(p, src_id,
+					    le16_to_cpu(param.num_net_msix),
+					    le16_to_cpu(param.num_others_msix),
+					    !!le16_to_cpu(param.msix_mask_en));
+		mutex_unlock(&disp_mgt->ops_mutex_lock);
+		/* Forward the resource errno verbatim to the requester */
+		if (ret)
+			err = ret;
+	} else {
+		err = -EOPNOTSUPP;
+	}
+ack_out:
+	nbl_chan_fill_ack_info(&chan_ack, src_id,
+			       NBL_CHAN_MSG_CONFIGURE_MSIX_MAP, msg_id,
+			       err, NULL, 0);
+	ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+	if (ret)
+		dev_err(dev,
+			"channel send ack failed with ret: %d, msg_type: %d\n",
+			ret, NBL_CHAN_MSG_CONFIGURE_MSIX_MAP);
+}
+
+/*
+ * Precondition: caller must disable mailbox IRQ_RDY and switch send_msg
+ * to the polling path before issuing this RPC. The responder retargets
+ * the requester's mailbox MSI-X routing during the resource op, so the
+ * ACK cannot rely on interrupt wakeup while routing is in flux.
+ */
+static int nbl_disp_chan_destroy_msix_map_req(struct nbl_dispatch_mgt *disp_mgt)
+{
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_common_info *common = disp_mgt->common;
+	struct nbl_chan_send_info chan_send = {0};
+	int ret;
+
+	nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+				NBL_CHAN_MSG_DESTROY_MSIX_MAP,
+				NULL, 0, NULL, 0, 1);
+	ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+	if (ret)
+		return ret;
+	return 0;
+}
+
+static void nbl_disp_chan_destroy_msix_map_resp(void *priv, u16 src_id,
+						u16 msg_id, void *data,
+						u32 data_len)
+{
+	struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct device *dev = disp_mgt->common->dev;
+	struct nbl_chan_ack_info chan_ack;
+	int err = 0;
+	u32 rel_pf_id;
+	int ret;
+
+	ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+					      &rel_pf_id);
+	if (ret) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (rel_pf_id >= disp_mgt->common->max_pf) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (res_ops->destroy_msix_map) {
+		mutex_lock(&disp_mgt->ops_mutex_lock);
+		ret = res_ops->destroy_msix_map(p, src_id);
+		mutex_unlock(&disp_mgt->ops_mutex_lock);
+		/* Forward the resource errno verbatim to the requester */
+		if (ret)
+			err = ret;
+	} else {
+		err = -EOPNOTSUPP;
+	}
+ack_out:
+	nbl_chan_fill_ack_info(&chan_ack, src_id,
+			       NBL_CHAN_MSG_DESTROY_MSIX_MAP, msg_id,
+			       err, NULL, 0);
+	ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+	if (ret)
+		dev_err(dev,
+			"channel send ack failed with ret: %d, msg_type: %d\n",
+			ret, NBL_CHAN_MSG_DESTROY_MSIX_MAP);
+}
+
+/*
+ * Precondition: caller must disable mailbox IRQ_RDY and switch send_msg
+ * to the polling path before issuing this RPC. The responder rewrites
+ * the requester's own mailbox MSI-X routing (MSIX_IDX / MSIX_IDX_VALID)
+ * before the ACK is sent, so the ACK cannot rely on interrupt wakeup
+ * while routing is in flux.
+ */
+static int nbl_disp_chan_set_mailbox_irq_req(struct nbl_dispatch_mgt *disp_mgt,
+					     u16 vector_id, bool en_msix)
+{
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_chan_param_set_mailbox_irq param = { 0 };
+	struct nbl_common_info *common = disp_mgt->common;
+	struct nbl_chan_send_info chan_send = {0};
+	int ret;
+
+	param.vector_id = cpu_to_le16(vector_id);
+	param.en_msix = !!en_msix;
+
+	nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+				NBL_CHAN_MSG_MAILBOX_SET_IRQ,
+				&param, sizeof(param), NULL, 0, 1);
+	ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+	if (ret)
+		return ret;
+	return 0;
+}
+
+static void nbl_disp_chan_set_mailbox_irq_resp(void *priv, u16 src_id,
+					       u16 msg_id, void *data,
+					       u32 data_len)
+{
+	struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct nbl_chan_param_set_mailbox_irq param = { 0 };
+	struct device *dev = disp_mgt->common->dev;
+	struct nbl_chan_ack_info chan_ack;
+	int err = 0;
+	u16 vector_id;
+	u32 rel_pf_id;
+	bool en_msix;
+	int ret;
+
+	ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+					      &rel_pf_id);
+	if (ret) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (rel_pf_id >= disp_mgt->common->max_pf) {
+		err = -EPERM;
+		goto ack_out;
+	}
+	if (data_len < sizeof(param)) {
+		err = -EBADMSG;
+		goto ack_out;
+	}
+	memcpy(&param, data, sizeof(param));
+	vector_id = le16_to_cpu(param.vector_id);
+	en_msix = !!param.en_msix;
+
+	if (res_ops->set_mailbox_irq) {
+		mutex_lock(&disp_mgt->ops_mutex_lock);
+		ret = res_ops->set_mailbox_irq(p, src_id, vector_id, en_msix);
+		mutex_unlock(&disp_mgt->ops_mutex_lock);
+		/* Forward the resource errno verbatim to the requester */
+		if (ret)
+			err = ret;
+	} else {
+		err = -EOPNOTSUPP;
+	}
+
+ack_out:
+	nbl_chan_fill_ack_info(&chan_ack, src_id,
+			       NBL_CHAN_MSG_MAILBOX_SET_IRQ, msg_id,
+			       err, NULL, 0);
+	ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+	if (ret)
+		dev_err(dev,
+			"channel send ack failed with ret: %d, msg_type: %d\n",
+			ret, NBL_CHAN_MSG_MAILBOX_SET_IRQ);
+}
+
+static int nbl_disp_destroy_msix_map(struct nbl_dispatch_mgt *disp_mgt)
+{
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct nbl_common_info *common = disp_mgt->common;
+	int ret;
+
+	if (!res_ops->destroy_msix_map)
+		return -EOPNOTSUPP;
+	mutex_lock(&disp_mgt->ops_mutex_lock);
+	ret = res_ops->destroy_msix_map(p, common->mgt_pf);
+	mutex_unlock(&disp_mgt->ops_mutex_lock);
+	return ret;
+}
+
+static int nbl_disp_set_mailbox_irq(struct nbl_dispatch_mgt *disp_mgt,
+				    u16 vector_id, bool en_msix)
+{
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct nbl_common_info *common = disp_mgt->common;
+	int ret;
+
+	if (!res_ops->set_mailbox_irq)
+		return -EOPNOTSUPP;
+	mutex_lock(&disp_mgt->ops_mutex_lock);
+	ret = res_ops->set_mailbox_irq(p, common->mgt_pf, vector_id, en_msix);
+	mutex_unlock(&disp_mgt->ops_mutex_lock);
+	return ret;
+}
+
+static int nbl_disp_get_vsi_id(struct nbl_dispatch_mgt *disp_mgt, u16 type,
+			       u16 *vsi_id)
+{
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct nbl_common_info *common = disp_mgt->common;
+
+	if (res_ops->get_vsi_id)
+		return res_ops->get_vsi_id(p, common->mgt_pf, type, vsi_id);
+	return -EOPNOTSUPP;
+}
+
+static int nbl_disp_get_eth_id(struct nbl_dispatch_mgt *disp_mgt, u16 vsi_id,
+			       u8 *eth_num, u8 *eth_id, u8 *logic_eth_id)
+{
+	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+	struct nbl_common_info *common = disp_mgt->common;
+
+	if (res_ops->get_eth_id)
+		return res_ops->get_eth_id(p, common->mgt_pf, vsi_id,
+					   eth_num, eth_id, logic_eth_id);
+	return -EOPNOTSUPP;
+}
+
+static int nbl_disp_setup_msg(struct nbl_dispatch_mgt *disp_mgt)
+{
+	struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+	struct nbl_channel_mgt *p = disp_mgt->chan_ops_tbl->priv;
+	int ret = 0;
+	int _ret;
+
+	_ret = chan_ops->register_msg(p, NBL_CHAN_MSG_CONFIGURE_MSIX_MAP,
+				      nbl_disp_chan_cfg_msix_map_resp,
+				      disp_mgt);
+	if (_ret < 0 && !ret)
+		ret = _ret;
+
+	_ret = chan_ops->register_msg(p, NBL_CHAN_MSG_DESTROY_MSIX_MAP,
+				      nbl_disp_chan_destroy_msix_map_resp,
+				      disp_mgt);
+	if (_ret < 0 && !ret)
+		ret = _ret;
+
+	_ret = chan_ops->register_msg(p, NBL_CHAN_MSG_MAILBOX_SET_IRQ,
+				      nbl_disp_chan_set_mailbox_irq_resp,
+				      disp_mgt);
+	if (_ret < 0 && !ret)
+		ret = _ret;
+
+	_ret = chan_ops->register_msg(p, NBL_CHAN_MSG_GET_VSI_ID,
+				      nbl_disp_chan_get_vsi_id_resp,
+				      disp_mgt);
+	if (_ret < 0 && !ret)
+		ret = _ret;
+
+	_ret = chan_ops->register_msg(p, NBL_CHAN_MSG_GET_ETH_ID,
+				      nbl_disp_chan_get_eth_id_resp,
+				      disp_mgt);
+	if (_ret < 0 && !ret)
+		ret = _ret;
+
+	return ret;
+}
+
 static void nbl_disp_set_ctrl_bit(struct nbl_dispatch_mgt *disp_mgt, u32 lvl)
 {
 	set_bit(lvl, disp_mgt->ctrl_lvl);
@@ -34,9 +554,22 @@ static void nbl_disp_refresh_ctrl_ops(struct nbl_dispatch_mgt *disp_mgt)
 {
 	struct nbl_dispatch_ops *disp_ops = disp_mgt->disp_ops_tbl->ops;
 
+	memset(disp_ops, 0, sizeof(*disp_ops));
 	if (test_bit(NBL_DISP_CTRL_LVL_MGT, disp_mgt->ctrl_lvl)) {
 		disp_ops->init_module = nbl_disp_init_module;
 		disp_ops->deinit_module = nbl_disp_deinit_module;
+		disp_ops->cfg_msix_map = nbl_disp_cfg_msix_map;
+		disp_ops->destroy_msix_map = nbl_disp_destroy_msix_map;
+		disp_ops->set_mailbox_irq = nbl_disp_set_mailbox_irq;
+		disp_ops->get_vsi_id = nbl_disp_get_vsi_id;
+		disp_ops->get_eth_id = nbl_disp_get_eth_id;
+	} else if (test_bit(NBL_DISP_CTRL_LVL_NET, disp_mgt->ctrl_lvl)) {
+		disp_ops->cfg_msix_map =
+			nbl_disp_chan_cfg_msix_map_req;
+		disp_ops->destroy_msix_map = nbl_disp_chan_destroy_msix_map_req;
+		disp_ops->set_mailbox_irq = nbl_disp_chan_set_mailbox_irq_req;
+		disp_ops->get_vsi_id = nbl_disp_chan_get_vsi_id_req;
+		disp_ops->get_eth_id = nbl_disp_chan_get_eth_id_req;
 	}
 }
 
@@ -45,12 +578,16 @@ nbl_disp_setup_disp_mgt(struct nbl_common_info *common)
 {
 	struct nbl_dispatch_mgt *disp_mgt;
 	struct device *dev = common->dev;
+	int err;
 
 	disp_mgt = devm_kzalloc(dev, sizeof(*disp_mgt), GFP_KERNEL);
 	if (!disp_mgt)
 		return ERR_PTR(-ENOMEM);
 
 	disp_mgt->common = common;
+	err = devm_mutex_init(common->dev, &disp_mgt->ops_mutex_lock);
+	if (err)
+		return ERR_PTR(err);
 	return disp_mgt;
 }
 
@@ -104,14 +641,40 @@ int nbl_disp_init(struct nbl_adapter *adapter)
 	adapter->core.disp_mgt = disp_mgt;
 	adapter->intf.dispatch_ops_tbl = disp_ops_tbl;
 
+	ret = nbl_disp_setup_msg(disp_mgt);
+	if (ret)
+		return ret;
+
 	if (common->has_ctrl)
 		nbl_disp_set_ctrl_bit(disp_mgt, NBL_DISP_CTRL_LVL_MGT);
 
+	if (common->has_net)
+		nbl_disp_set_ctrl_bit(disp_mgt, NBL_DISP_CTRL_LVL_NET);
 	nbl_disp_refresh_ctrl_ops(disp_mgt);
 	return 0;
 }
 
 void nbl_disp_remove(struct nbl_adapter *adapter)
 {
-	/* Dispatch structures are allocated via devm */
+	/*
+	 * Dispatch structures are devm-allocated and freed at detach.
+	 *
+	 * The five responders registered by nbl_disp_setup_msg() are
+	 * owned by the channel layer (xarray of handlers) and are never
+	 * unregistered here. This is safe because the teardown order
+	 * guarantees no responder can run after this point:
+	 *
+	 *   mailbox teardown
+	 *     -> cancel_work_sync(clean_mbx_task)   // drain RX work
+	 *     -> nbl_chan_teardown_queue()          // stop HW queue,
+	 *                                                 // join clean task,
+	 *                                                 // active=false
+	 *     -> nbl_chan_remove_common()
+	 *          -> destroy_wq()                       // no new work
+	 *          -> nbl_chan_remove_msg_handler()      // free handler nodes
+	 *
+	 * By the time devres frees disp_mgt, the mailbox queue is stopped
+	 * and the handler xarray is empty, so no responder can touch
+	 * res_mgt->intr_mgt after nbl_intr_mgt_stop() has cleared it.
+	 */
 }
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
index a7e5802344b4..8549048f76e9 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
@@ -18,6 +18,8 @@ struct nbl_dispatch_mgt {
 	struct nbl_channel_ops_tbl *chan_ops_tbl;
 	struct nbl_dispatch_ops_tbl *disp_ops_tbl;
 	DECLARE_BITMAP(ctrl_lvl, NBL_DISP_CTRL_LVL_MAX);
+	/* use for the caller not in interrupt */
+	struct mutex ops_mutex_lock;
 };
 
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
index bf971121d2ec..5583c465ac16 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
@@ -30,6 +30,11 @@ typedef void (*nbl_chan_resp)(void *, u16, u16, void *, u32);
  */
 enum nbl_chan_msg_type {
 	NBL_CHAN_MSG_ACK = 0,
+	NBL_CHAN_MSG_CONFIGURE_MSIX_MAP = 17,
+	NBL_CHAN_MSG_DESTROY_MSIX_MAP = 18,
+	NBL_CHAN_MSG_MAILBOX_SET_IRQ = 19,
+	NBL_CHAN_MSG_GET_VSI_ID = 21,
+	NBL_CHAN_MSG_GET_ETH_ID = 67,
 	/* mailbox msg end */
 	NBL_CHAN_MSG_MAILBOX_MAX,
 };
@@ -39,6 +44,32 @@ enum nbl_chan_state {
 	NBL_CHAN_STATE_NBITS
 };
 
+struct nbl_chan_param_cfg_msix_map {
+	__le16 num_net_msix;
+	__le16 num_others_msix;
+	__le16 msix_mask_en;
+	__le16 rsvd;
+};
+
+struct nbl_chan_param_set_mailbox_irq {
+	__le16 vector_id;
+	u8 en_msix;
+	u8 rsvd;
+};
+
+struct nbl_chan_param_get_vsi_id {
+	__le16 vsi_id;
+	__le16 type;
+};
+
+struct nbl_chan_param_get_eth_id {
+	__le16 vsi_id;
+	u8 eth_num;
+	u8 eth_id;
+	u8 logic_eth_id;
+	u8 rsvd[3];
+};
+
 struct nbl_board_port_info {
 	u8 eth_num;
 	u8 eth_speed;
@@ -46,6 +77,14 @@ struct nbl_board_port_info {
 	u8 rsv[5];
 };
 
+static_assert(sizeof(struct nbl_chan_param_cfg_msix_map) == 8,
+	      "nbl_chan_param_cfg_msix_map size must be 8 bytes");
+static_assert(sizeof(struct nbl_chan_param_set_mailbox_irq) == 4,
+	      "nbl_chan_param_set_mailbox_irq size must be 4 bytes");
+static_assert(sizeof(struct nbl_chan_param_get_vsi_id) == 4,
+	      "nbl_chan_param_get_vsi_id size must be 4 bytes");
+static_assert(sizeof(struct nbl_chan_param_get_eth_id) == 8,
+	      "nbl_chan_param_get_eth_id size must be 8 bytes");
 static_assert(sizeof(struct nbl_board_port_info) == 8,
 	      "nbl_board_port_info size must be 8 bytes");
 
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
index b3398591035b..74ff69fe8512 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
@@ -12,6 +12,7 @@ struct nbl_dispatch_mgt;
 struct nbl_adapter;
 enum {
 	NBL_DISP_CTRL_LVL_MGT,
+	NBL_DISP_CTRL_LVL_NET,
 	NBL_DISP_CTRL_LVL_MAX,
 };
 
@@ -21,6 +22,11 @@ enum {
  *               caller must check has_ctrl guard
  * @deinit_module: dispatch layer cleanup, control-PF exclusive,
  *                 caller must check has_ctrl guard
+ * @cfg_msix_map: configure function msix mapping table
+ * @destroy_msix_map: tear down msix mapping resource
+ * @set_mailbox_irq: bind mailbox interrupt to specified msix vector
+ * @get_vsi_id: resolve VSI ID by type
+ * @get_eth_id: resolve eth port info from VSI ID
  *
  * Warning: init_module/deinit_module are control-PF exclusive. The five
  * resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq,
@@ -33,6 +39,16 @@ enum {
 struct nbl_dispatch_ops {
 	int (*init_module)(struct nbl_dispatch_mgt *disp_mgt);
 	void (*deinit_module)(struct nbl_dispatch_mgt *disp_mgt);
+	int (*cfg_msix_map)(struct nbl_dispatch_mgt *disp_mgt,
+			    u16 num_net_msix, u16 num_others_msix,
+			    bool net_msix_mask_en);
+	int (*destroy_msix_map)(struct nbl_dispatch_mgt *disp_mgt);
+	int (*set_mailbox_irq)(struct nbl_dispatch_mgt *disp_mgt,
+			       u16 vector_id, bool en_msix);
+	int (*get_vsi_id)(struct nbl_dispatch_mgt *disp_mgt, u16 type,
+			  u16 *vsi_id);
+	int (*get_eth_id)(struct nbl_dispatch_mgt *disp_mgt, u16 vsi_id,
+			  u8 *eth_num, u8 *eth_id, u8 *logic_eth_id);
 };
 
 struct nbl_dispatch_ops_tbl {
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
                   ` (5 preceding siblings ...)
  2026-09-28 12:32 ` [PATCH v30 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  2026-09-28 12:32 ` [PATCH v30 net-next 8/8] net/nebula-matrix: add common dev start/stop operation illusion.wang
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Add nbl_dev core device infrastructure for nebula-matrix NIC lifecycle
management. Create nbl_dev.c, nbl_dev.h and nbl_def_dev.h to host
device-level initialization and teardown logic, and wire the new
nbl_dev_init() / nbl_dev_remove() entry points into core init/remove.

Implement common device setup helper to allocate per-device state,
initialize mailbox cleanup work, create mailbox channel queue, and
pre-populate MSI-X service vector counts. Defer VSI/ETH identity
lookup and MSI-X vector allocation to the subsequent start routine,
to guarantee a fully ready control PF mailbox responder before any
cross-PF RPC access and keep intermediate commits bisectable.

Implement control-PF-only device setup to invoke chip-level module
initialization and program mailbox QINFO routing table entries for
PF bus/devid mapping.

Enforce hardware-aware init/teardown ordering: initialize common
mailbox channel state before control device hardware setup, and
teardown mailbox channel and drain in-flight DMA prior to firmware
de-initialization notification, to avoid invalid DMA writes.

The firmware handles global hardware cleanup and QINFO routing state
reclamation asynchronously after driver status is marked inactive.
Device link dependency established during non-control PF probe ensures
sibling PFs are unbound before the control PF is removed.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../net/ethernet/nebula-matrix/nbl/Makefile   |   1 +
 .../net/ethernet/nebula-matrix/nbl/nbl_core.h |   1 +
 .../nebula-matrix/nbl/nbl_core/nbl_dev.c      | 253 ++++++++++++++++++
 .../nebula-matrix/nbl/nbl_core/nbl_dev.h      |  55 ++++
 .../nbl/nbl_include/nbl_def_dev.h             |  14 +
 .../net/ethernet/nebula-matrix/nbl/nbl_main.c |   9 +
 6 files changed, 333 insertions(+)
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h
 create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index b7eebd89b4d1..71fbe3ee7e62 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -11,4 +11,5 @@ nbl-objs +=	nbl_common/nbl_common.o \
 		nbl_hw/nbl_interrupt.o \
 		nbl_hw/nbl_chip.o \
 		nbl_core/nbl_dispatch.o \
+		nbl_core/nbl_dev.o \
 		nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index 4d8cea8d8ab3..c3c4dd685bf6 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -25,6 +25,7 @@ struct nbl_core {
 	struct nbl_hw_mgt *hw_mgt;
 	struct nbl_resource_mgt *res_mgt;
 	struct nbl_dispatch_mgt *disp_mgt;
+	struct nbl_dev_mgt *dev_mgt;
 	struct nbl_channel_mgt *chan_mgt;
 };
 
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
new file mode 100644
index 000000000000..e094b97acdfb
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
@@ -0,0 +1,253 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/pci.h>
+#include "nbl_dev.h"
+
+static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+	struct nbl_msix_info *msix_info = &dev_common->msix_info;
+
+	/* mailbox vector allocated in nbl_dev_start() via
+	 * nbl_dev_init_interrupt_scheme(); nbl_dev_request_mailbox_irq()
+	 * only attaches the irq handler to pre-allocated vectors.
+	 */
+	msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num = 1;
+}
+
+/* ----------  Channel config  ---------- */
+static void nbl_dev_setup_chan_qinfo(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
+{
+	struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+	struct nbl_channel_mgt *priv = dev_mgt->chan_ops_tbl->priv;
+	struct nbl_common_info *common = dev_mgt->common;
+
+	if (!chan_ops->check_queue_exist(priv, chan_type))
+		return;
+
+	/*
+	 * common->hw_bus is the control PF's real bus number, captured in
+	 * nbl_res_ctrl_dev_sriov_info_init() during nbl_res_init_leonis().
+	 * nbl_core_init() runs resource init before nbl_dev_init(), so the
+	 * value is always initialized when this control-PF-only path runs;
+	 * nbl_res_intr_cfg_msix_map() consumes it for cfg_msix_map() the
+	 * same way.
+	 */
+	chan_ops->cfg_chan_qinfo_map_table(priv, common->hw_bus, common->devid);
+}
+
+static int nbl_dev_setup_chan_queue(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
+{
+	struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+	struct nbl_channel_mgt *priv = dev_mgt->chan_ops_tbl->priv;
+	int ret = 0;
+
+	if (chan_ops->check_queue_exist(priv, chan_type))
+		ret = chan_ops->setup_queue(priv, chan_type);
+
+	return ret;
+}
+
+static int nbl_dev_remove_chan_queue(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
+{
+	struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+	struct nbl_channel_mgt *priv = dev_mgt->chan_ops_tbl->priv;
+	int ret = 0;
+
+	if (chan_ops->check_queue_exist(priv, chan_type))
+		ret = chan_ops->teardown_queue(priv, chan_type);
+
+	return ret;
+}
+
+static void nbl_dev_register_chan_task(struct nbl_dev_mgt *dev_mgt,
+				       u8 chan_type, struct work_struct *task)
+{
+	struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+
+	if (chan_ops->check_queue_exist(dev_mgt->chan_ops_tbl->priv, chan_type))
+		chan_ops->register_chan_task(dev_mgt->chan_ops_tbl->priv,
+					     chan_type, task);
+}
+
+/* ----------  Tasks config  ---------- */
+static void nbl_dev_clean_mailbox_task(struct work_struct *work)
+{
+	struct nbl_dev_common *common_dev =
+		container_of(work, struct nbl_dev_common, clean_mbx_task);
+	struct nbl_dev_mgt *dev_mgt = common_dev->dev_mgt;
+	struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+
+	chan_ops->clean_queue_subtask(dev_mgt->chan_ops_tbl->priv,
+				      NBL_CHAN_TYPE_MAILBOX);
+}
+
+/* ----------  Dev init process  ---------- */
+static int nbl_dev_setup_common_dev(struct nbl_adapter *adapter)
+{
+	struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+	struct nbl_dev_common *common_dev;
+	int ret;
+
+	common_dev = devm_kzalloc(&adapter->pdev->dev, sizeof(*common_dev),
+				  GFP_KERNEL);
+	if (!common_dev)
+		return -ENOMEM;
+	common_dev->dev_mgt = dev_mgt;
+
+	/*
+	 * Initialize clean_mbx_task before any operation that could jump
+	 * to err_cleanup, so cancel_work_sync() is always safe there.
+	 */
+	INIT_WORK(&common_dev->clean_mbx_task, nbl_dev_clean_mailbox_task);
+
+	ret = nbl_dev_setup_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
+	if (ret)
+		goto err_cleanup;
+
+	nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX,
+				   &common_dev->clean_mbx_task);
+	/*
+	 * VSI/ETH identity fetch moved to nbl_dev_start().
+	 * This avoids cross-PF probe race when manager PF is not ready.
+	 */
+	dev_mgt->common_dev = common_dev;
+	nbl_dev_init_msix_cnt(dev_mgt);
+
+	return 0;
+err_cleanup:
+	cancel_work_sync(&common_dev->clean_mbx_task);
+	nbl_dev_remove_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
+	nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX, NULL);
+	return ret;
+}
+
+static void nbl_dev_remove_common_dev(struct nbl_adapter *adapter)
+{
+	struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+	struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+	int ret;
+
+	if (!common_dev)
+		return;
+	cancel_work_sync(&common_dev->clean_mbx_task);
+	ret = nbl_dev_remove_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
+	if (ret)
+		dev_warn(&adapter->pdev->dev, "mailbox queue teardown failed, inflight DMA may exist: %d\n",
+			 ret);
+	nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX, NULL);
+}
+
+static int nbl_dev_setup_ctrl_dev(struct nbl_adapter *adapter)
+{
+	struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+	struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+
+	return disp_ops->init_module(dev_mgt->disp_ops_tbl->priv);
+}
+
+/*
+ * Tear down control device: deinit_module sets driver_status=false
+ * to notify firmware to clean all per-PF hardware state (including
+ * qinfo registers).  The qinfo map programmed in nbl_dev_init() via
+ * nbl_dev_setup_chan_qinfo() is not explicitly cleared; firmware
+ * handles it on driver_status change.
+ *
+ * Teardown ordering guarantee: every non-management PF creates a
+ * consumer->control PF device link in its probe path, so the driver
+ * core always unbinds all siblings before allowing the control PF to
+ * be detached (sysfs unbind, driver unregister and hot-unplug alike).
+ * nbl_probe_chip_deps() defers a sibling's probe until the control PF
+ * is fully bound (DL_DEV_DRIVER_BOUND), so no sibling can race its
+ * mailbox setup against this deinit.
+ *
+ * Direct control-PF FLR (which never runs driver teardown) cannot be
+ * guarded here.
+ */
+static void nbl_dev_remove_ctrl_dev(struct nbl_adapter *adapter)
+{
+	struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+	struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+
+	disp_ops->deinit_module(dev_mgt->disp_ops_tbl->priv);
+}
+
+static struct nbl_dev_mgt *nbl_dev_setup_dev_mgt(struct nbl_common_info *common)
+{
+	struct nbl_dev_mgt *dev_mgt;
+
+	dev_mgt = devm_kzalloc(common->dev, sizeof(*dev_mgt), GFP_KERNEL);
+	if (!dev_mgt)
+		return ERR_PTR(-ENOMEM);
+
+	dev_mgt->common = common;
+	return dev_mgt;
+}
+
+int nbl_dev_init(struct nbl_adapter *adapter)
+{
+	struct nbl_common_info *common = &adapter->common;
+	struct nbl_dispatch_ops_tbl *disp_ops_tbl =
+		adapter->intf.dispatch_ops_tbl;
+	struct nbl_channel_ops_tbl *chan_ops_tbl =
+		adapter->intf.channel_ops_tbl;
+	struct nbl_dev_mgt *dev_mgt;
+	int ret;
+
+	dev_mgt = nbl_dev_setup_dev_mgt(common);
+	if (IS_ERR(dev_mgt)) {
+		ret = PTR_ERR(dev_mgt);
+		return ret;
+	}
+
+	dev_mgt->disp_ops_tbl = disp_ops_tbl;
+	dev_mgt->chan_ops_tbl = chan_ops_tbl;
+	adapter->core.dev_mgt = dev_mgt;
+	if (common->has_ctrl)
+		nbl_dev_setup_chan_qinfo(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
+	/*
+	 * Chip hardware initialization is completed by firmware at power-up.
+	 * Only driver functional table/register config follows here, safe to
+	 * access hardware registers before ctrl dev setup.
+	 */
+	ret = nbl_dev_setup_common_dev(adapter);
+	if (ret)
+		goto setup_err;
+
+	if (common->has_ctrl) {
+		ret = nbl_dev_setup_ctrl_dev(adapter);
+		if (ret)
+			goto setup_ctrl_dev_fail;
+	}
+
+	return 0;
+setup_ctrl_dev_fail:
+	nbl_dev_remove_common_dev(adapter);
+setup_err:
+	return ret;
+}
+
+/*
+ * Teardown order: Stop mailbox channel and drain all inflight DMA first,
+ * then invoke deinit_module to notify firmware.
+ *
+ * This intentionally breaks strict init/teardown mirror symmetry due to
+ * hardware constraint: firmware may perform asynchronous global hardware
+ * cleanup once driver_status=false is set. We must guarantee no ongoing
+ * mailbox DMA before deinit_module to avoid invalid DMA write.
+ *
+ * Init order: program mailbox QINFO routing (ctrl PF only) ->
+ *            create mailbox queue/common_dev -> ctrl dev init_module
+ * Teardown order: destroy mailbox queue/common_dev -> ctrl dev deinit_module
+ */
+void nbl_dev_remove(struct nbl_adapter *adapter)
+{
+	struct nbl_common_info *common = &adapter->common;
+
+	nbl_dev_remove_common_dev(adapter);
+	if (common->has_ctrl)
+		nbl_dev_remove_ctrl_dev(adapter);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h
new file mode 100644
index 000000000000..24e890fd8987
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h
@@ -0,0 +1,55 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEV_H_
+#define _NBL_DEV_H_
+
+#include <linux/types.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_hw.h"
+#include "../nbl_include/nbl_def_resource.h"
+#include "../nbl_include/nbl_def_dispatch.h"
+#include "../nbl_include/nbl_def_dev.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+#define NBL_STRING_NAME_LEN			32
+
+enum nbl_msix_serv_type {
+	NBL_MSIX_NET_TYPE,
+	NBL_MSIX_MAILBOX_TYPE,
+	NBL_MSIX_TYPE_MAX
+};
+
+struct nbl_msix_serv_info {
+	char irq_name[NBL_STRING_NAME_LEN];
+	u16 num;
+	u16 base_vector_id;
+	/* true: hw report msix, hw need to mask actively */
+	bool hw_self_mask_en;
+};
+
+struct nbl_msix_info {
+	struct nbl_msix_serv_info serv_info[NBL_MSIX_TYPE_MAX];
+};
+
+struct nbl_dev_common {
+	struct nbl_dev_mgt *dev_mgt;
+	struct nbl_msix_info msix_info;
+	char mailbox_name[NBL_STRING_NAME_LEN];
+	/* for ctrl-dev/net-dev mailbox recv msg */
+	struct work_struct clean_mbx_task;
+};
+
+struct nbl_dev_mgt {
+	struct nbl_common_info *common;
+	struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+	struct nbl_channel_ops_tbl *chan_ops_tbl;
+	struct nbl_dev_common *common_dev;
+};
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
new file mode 100644
index 000000000000..51cf04e4c552
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
@@ -0,0 +1,14 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_DEV_H_
+#define _NBL_DEF_DEV_H_
+
+struct nbl_adapter;
+
+int nbl_dev_init(struct nbl_adapter *adapter);
+void nbl_dev_remove(struct nbl_adapter *adapter);
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index bf4a5ea3e410..af430d3dfb70 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -12,6 +12,7 @@
 #include "nbl_include/nbl_def_hw.h"
 #include "nbl_include/nbl_def_resource.h"
 #include "nbl_include/nbl_def_dispatch.h"
+#include "nbl_include/nbl_def_dev.h"
 #include "nbl_include/nbl_def_common.h"
 #include "nbl_core.h"
 
@@ -53,7 +54,14 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 	ret = nbl_disp_init(adapter);
 	if (ret)
 		goto disp_init_fail;
+
+	ret = nbl_dev_init(adapter);
+	if (ret)
+		goto dev_init_fail;
 	return adapter;
+
+dev_init_fail:
+	nbl_disp_remove(adapter);
 disp_init_fail:
 	nbl_res_remove_leonis(adapter);
 res_init_fail:
@@ -66,6 +74,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 
 void nbl_core_remove(struct nbl_adapter *adapter)
 {
+	nbl_dev_remove(adapter);
 	nbl_disp_remove(adapter);
 	nbl_res_remove_leonis(adapter);
 	nbl_chan_remove_common(adapter);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH v30 net-next 8/8] net/nebula-matrix: add common dev start/stop operation
  2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
                   ` (6 preceding siblings ...)
  2026-09-28 12:32 ` [PATCH v30 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
@ 2026-09-28 12:32 ` illusion.wang
  2026-10-02  3:35   ` netdev-bot+sashiko
  7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-09-28 12:32 UTC (permalink / raw)
  To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
  Cc: kuba, edumazet, horms, open list

From: illusion wang <illusion.wang@nebula-matrix.com>

Implement nbl_dev_start() and nbl_dev_stop() to manage MSI-X hardware
mapping and mailbox interrupt setup/teardown.

nbl_dev_start() performs device startup in strict order: configure hardware
MSI-X mapping table, fetch VSI/ETH identity via dispatch RPC, allocate
MSI-X IRQ vectors with mailbox vector affinity isolation, request and
attach mailbox interrupt handler, then enable hardware mailbox interrupt
and mark channel interrupt ready.

Convert mailbox RPC timeouts to -EPROBE_DEFER for non-control PFs on
all startup RPC paths to trigger deferred probing when the control PF
is not ready. Isolate mailbox administrative vector from CPU affinity
management to avoid mailbox RPC stalls caused by CPU offline.

Skip redundant MSI-X map destroy RPC on unconfigured state during probe
rollback to eliminate unnecessary polling timeouts. Adjust interrupt
teardown ordering to clear channel IRQ ready state before masking
hardware interrupts, preventing in-flight ACK discard and "Channel
waiting ack failed" errors.

nbl_dev_stop() tears down resources in reverse startup order. Switch
channel to polling mode and clear IRQ ready state before masking
hardware interrupts to preserve ACK integrity. Release mailbox IRQ
handlers and destroy device-side MSI-X mapping while leaving kernel
MSI-X vectors managed by devres/pcim to avoid double-free.
Drain pending mailbox cleanup work to ensure consistent teardown.

Teardown is best-effort: failed MSI-X destroy RPC leaves stale hardware
entries which are reclaimed by firmware on chip reset, with
pci_clear_master() serving as the final safety net. The start/stop
pair is single-shot and non-repeatable, tied strictly to PCI probe/remove
device lifecycle.

Add thin nbl_core_start()/nbl_core_stop() wrappers and hook them into
PCI probe and remove paths to complete the device lifecycle management
series.

Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
 .../net/ethernet/nebula-matrix/nbl/nbl_core.h |   8 +
 .../nebula-matrix/nbl/nbl_core/nbl_dev.c      | 372 ++++++++++++++++++
 .../nbl/nbl_include/nbl_def_dev.h             |   2 +
 .../net/ethernet/nebula-matrix/nbl/nbl_main.c |  18 +
 4 files changed, 400 insertions(+)

diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index c3c4dd685bf6..655dfb43e365 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -40,4 +40,12 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 				  struct nbl_init_param *param);
 void nbl_core_remove(struct nbl_adapter *adapter);
 
+/*
+ * Single-shot start/stop pair, called once each from PCI probe/remove.
+ * Not repeatable: MSI-X vectors stay allocated until device detach, so
+ * a second start on a bound device is rejected by the MSI-X core.
+ */
+int nbl_core_start(struct nbl_adapter *adapter);
+void nbl_core_stop(struct nbl_adapter *adapter);
+
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
index e094b97acdfb..79c62c169d70 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
@@ -6,6 +6,17 @@
 #include <linux/pci.h>
 #include "nbl_dev.h"
 
+static void nbl_dev_clean_mailbox_schedule(struct nbl_dev_mgt *dev_mgt);
+
+/* ----------  Interrupt config  ---------- */
+static irqreturn_t nbl_dev_clean_mailbox(int irq __always_unused, void *data)
+{
+	struct nbl_dev_mgt *dev_mgt = (struct nbl_dev_mgt *)data;
+
+	nbl_dev_clean_mailbox_schedule(dev_mgt);
+	return IRQ_HANDLED;
+}
+
 static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
 {
 	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
@@ -18,6 +29,233 @@ static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
 	msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num = 1;
 }
 
+static int nbl_dev_request_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+	struct nbl_msix_info *msix_info = &dev_common->msix_info;
+	struct nbl_common_info *common = dev_mgt->common;
+	u16 lvec;
+	int irq_num;
+	int err;
+
+	if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+		return 0;
+
+	lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+	irq_num = pci_irq_vector(common->pdev, lvec);
+	if (irq_num < 0) {
+		dev_err(common->dev, "Failed to get mailbox IRQ vector: %d\n",
+			irq_num);
+		return irq_num;
+	}
+
+	snprintf(dev_common->mailbox_name, sizeof(dev_common->mailbox_name),
+		 "nbl_mailbox@pci:%s", pci_name(common->pdev));
+	err = request_irq(irq_num, nbl_dev_clean_mailbox, 0,
+			  dev_common->mailbox_name, dev_mgt);
+	if (err)
+		return err;
+
+	return 0;
+}
+
+static void nbl_dev_free_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+	struct nbl_msix_info *msix_info = &dev_common->msix_info;
+	struct nbl_common_info *common = dev_mgt->common;
+	u16 lvec;
+	int irq_num;
+
+	if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+		return;
+
+	lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+	irq_num = pci_irq_vector(common->pdev, lvec);
+	if (irq_num >= 0)
+		free_irq(irq_num, dev_mgt);
+}
+
+/*
+ * Translate a mailbox RPC timeout on a non-control PF into deferred
+ * probing: the management PF/firmware is not responsive yet and the
+ * driver core should retry once PF0 is ready. Other errors, and all
+ * errors on the control PF itself, pass through unchanged.
+ */
+static int nbl_dev_rpc_timeout(struct nbl_common_info *common, int ret)
+{
+	if (!common->has_ctrl && ret == -ETIMEDOUT)
+		return -EPROBE_DEFER;
+	return ret;
+}
+
+static int nbl_dev_enable_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+	struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+	struct nbl_msix_info *msix_info = &dev_common->msix_info;
+	u16 lvec;
+	int ret;
+
+	if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+		return 0;
+
+	lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+	/*
+	 * Enable sequence: perform set_mailbox_irq RPC in polling mode first.
+	 * Only set NBL_CHAN_IRQ_RDY after RPC succeeds, mirroring disable path.
+	 * This avoids waiting for an interrupt which has not been armed yet.
+	 */
+	ret = disp_ops->set_mailbox_irq(dev_mgt->disp_ops_tbl->priv,
+					lvec, true);
+	if (ret)
+		return nbl_dev_rpc_timeout(dev_mgt->common, ret);
+	chan_ops->set_queue_state(dev_mgt->chan_ops_tbl->priv,
+				  NBL_CHAN_IRQ_RDY,
+				  NBL_CHAN_TYPE_MAILBOX, true);
+	return 0;
+}
+
+static int nbl_dev_disable_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+	struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+	struct nbl_msix_info *msix_info = &dev_common->msix_info;
+	u16 lvec;
+
+	if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+		return 0;
+
+	lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+	/*
+	 * Disable sequence invariant: update software state first, then mask
+	 * hardware interrupt. Must not reverse the order.
+	 *
+	 * If hardware interrupt is masked before clearing INTERRUPT_READY,
+	 * the hardware may still transmit outstanding ACK packets for in-flight
+	 * messages. Subsequent switch to polling mode discards pending ACK
+	 * processing, triggering "Channel waiting ack failed" and "Skip ack
+	 * with invalid status" errors.
+	 *
+	 * By entering polling mode first, any late hardware interrupts are
+	 * ignored without pending ACK expectations, then hardware interrupt
+	 * can be safely disabled.
+	 *
+	 * This helper is invoked in two paths:
+	 * 1. Error unwind path of nbl_dev_start(): followed immediately by
+	 * nbl_dev_free_mailbox_irq() and full channel teardown. No new mailbox
+	 * interrupts can fire afterwards, and subsequent cancel_work_sync()
+	 * drains pending cleanup work before resources are released.
+	 * 2. Normal device stop path nbl_dev_stop(): free_irq() blocks until
+	 * any in-flight hardirq handler completes and prevents new interrupts.
+	 * cancel_work_sync() then waits for any already running mailbox cleanup
+	 * work to finish, or cancels queued but unstarted work items before
+	 * final channel destruction. No stuck descriptors linger in either
+	 * scenario.
+	 */
+	chan_ops->set_queue_state(dev_mgt->chan_ops_tbl->priv,
+				  NBL_CHAN_IRQ_RDY,
+				  NBL_CHAN_TYPE_MAILBOX, false);
+
+	return disp_ops->set_mailbox_irq(dev_mgt->disp_ops_tbl->priv,
+					 lvec, false);
+}
+
+static int nbl_dev_cfg_msix_map(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+	struct nbl_msix_info *msix_info = &dev_common->msix_info;
+	bool mask_en = msix_info->serv_info[NBL_MSIX_NET_TYPE].hw_self_mask_en;
+	u16 msix_net_num = msix_info->serv_info[NBL_MSIX_NET_TYPE].num;
+	u16 msix_not_net_num = 0;
+	int err, i;
+
+	msix_info->serv_info[NBL_MSIX_NET_TYPE].base_vector_id = 0;
+	/*
+	 * Calculate base_vector_id for each MSIX service type.
+	 * This relies on NBL_MSIX_TYPE enum being ordered sequentially,
+	 * starting from NBL_MSIX_NET_TYPE.
+	 */
+	for (i = NBL_MSIX_NET_TYPE + 1; i < NBL_MSIX_TYPE_MAX; i++)
+		msix_info->serv_info[i].base_vector_id =
+			msix_info->serv_info[i - 1].base_vector_id +
+			msix_info->serv_info[i - 1].num;
+
+	for (i = 0; i < NBL_MSIX_TYPE_MAX; i++) {
+		if (i == NBL_MSIX_NET_TYPE)
+			continue;
+		msix_not_net_num += msix_info->serv_info[i].num;
+	}
+
+	err = disp_ops->cfg_msix_map(dev_mgt->disp_ops_tbl->priv,
+				     msix_net_num, msix_not_net_num,
+				     mask_en);
+
+	return err;
+}
+
+static int nbl_dev_destroy_msix_map(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+
+	return disp_ops->destroy_msix_map(dev_mgt->disp_ops_tbl->priv);
+}
+
+static int nbl_dev_init_interrupt_scheme(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+	struct nbl_msix_info *msix_info = &dev_common->msix_info;
+	struct nbl_common_info *common = dev_mgt->common;
+	struct irq_affinity affd = { 0 };
+	int needed = 0;
+	int err;
+	int i;
+
+	for (i = 0; i < NBL_MSIX_TYPE_MAX; i++)
+		needed += msix_info->serv_info[i].num;
+
+	/*
+	 * The mailbox vector is the trailing administrative vector;
+	 * reserve it via post_vectors so it is excluded from managed
+	 * affinity spreading. A managed vector can be shut down when
+	 * its last assigned CPU goes offline, which would stall every
+	 * mailbox RPC while NBL_CHAN_IRQ_RDY stays set.
+	 */
+	affd.post_vectors = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num;
+
+	err = pci_alloc_irq_vectors_affinity(common->pdev, needed, needed,
+					     PCI_IRQ_MSIX | PCI_IRQ_AFFINITY,
+					     &affd);
+	if (err < 0) {
+		dev_err(common->dev,
+			"pci_alloc_irq_vectors failed, err = %d\n", err);
+		return err;
+	}
+	if (err != needed) {
+		dev_err(common->dev, "pci_alloc_irq_vectors got %d vecs, need %d\n",
+			err, needed);
+		return -ENOSPC;
+	}
+	return 0;
+}
+
+/*
+ * Kernel-side MSI-X vectors are deliberately NOT freed here. They are
+ * released at device detach by the pcim_msi_release devres callback
+ * registered by pci_alloc_irq_vectors() (probe uses pcim_enable_device);
+ * freeing them here would double-free that callback.
+ *
+ * Consequently start/stop is a single-shot pair tied to probe/remove:
+ * vectors stay allocated until detach, and a second start against a
+ * device with msix_enabled set is rejected by __pci_enable_msix_range().
+ */
+static void nbl_dev_clear_interrupt_scheme(struct nbl_dev_mgt *dev_mgt)
+{
+}
+
 /* ----------  Channel config  ---------- */
 static void nbl_dev_setup_chan_qinfo(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
 {
@@ -85,6 +323,14 @@ static void nbl_dev_clean_mailbox_task(struct work_struct *work)
 				      NBL_CHAN_TYPE_MAILBOX);
 }
 
+static void nbl_dev_clean_mailbox_schedule(struct nbl_dev_mgt *dev_mgt)
+{
+	struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+	struct nbl_common_info *common = dev_mgt->common;
+
+	queue_work(common->wq, &common_dev->clean_mbx_task);
+}
+
 /* ----------  Dev init process  ---------- */
 static int nbl_dev_setup_common_dev(struct nbl_adapter *adapter)
 {
@@ -251,3 +497,129 @@ void nbl_dev_remove(struct nbl_adapter *adapter)
 	if (common->has_ctrl)
 		nbl_dev_remove_ctrl_dev(adapter);
 }
+
+/* ----------  Dev start process  ---------- */
+int nbl_dev_start(struct nbl_adapter *adapter)
+{
+	struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+	struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+	struct nbl_dispatch_mgt *priv = dev_mgt->disp_ops_tbl->priv;
+	struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+	struct nbl_common_info *common = dev_mgt->common;
+	bool map_ready = false;
+	int cleanup_ret;
+	int ret;
+
+	ret = nbl_dev_rpc_timeout(common, nbl_dev_cfg_msix_map(dev_mgt));
+	if (ret)
+		goto err_destroy_map;
+	map_ready = true;
+
+	/* Fetch VSI/ETH identity after cfg_msix_map */
+	ret = disp_ops->get_vsi_id(priv, NBL_VSI_DATA, &common->vsi_id);
+	if (ret) {
+		ret = nbl_dev_rpc_timeout(common, ret);
+		goto err_destroy_map;
+	}
+	ret = disp_ops->get_eth_id(priv, common->vsi_id, &common->eth_num,
+				   &common->eth_id, &common->logic_eth_id);
+	if (ret) {
+		ret = nbl_dev_rpc_timeout(common, ret);
+		goto err_destroy_map;
+	}
+
+	ret = nbl_dev_init_interrupt_scheme(dev_mgt);
+	if (ret)
+		goto err_destroy_map;
+
+	ret = nbl_dev_request_mailbox_irq(dev_mgt);
+	if (ret)
+		goto err_destroy_map;
+
+	ret = nbl_dev_enable_mailbox_irq(dev_mgt);
+	if (ret)
+		goto err_disable_irq;
+
+	return 0;
+
+err_disable_irq:
+	cleanup_ret = nbl_dev_disable_mailbox_irq(dev_mgt);
+	if (cleanup_ret)
+		dev_err(dev_mgt->common->dev,
+			"rollback: disable mailbox IRQ failed: %d\n",
+			cleanup_ret);
+	nbl_dev_free_mailbox_irq(dev_mgt);
+err_destroy_map:
+	/*
+	 * Destroy the device-side MSI-X map only when it was configured.
+	 * On a cfg RPC failure there is no known-good remote map; when
+	 * the failure is a timeout against an unready/unresponsive
+	 * management PF, issuing the destroy RPC would just burn another
+	 * multi-second ACK timeout. Partial remote state is best-effort
+	 * and reclaimed by firmware on chip reset.
+	 *
+	 * The destroy runs before detach even though kernel vectors stay
+	 * allocated until then (see nbl_dev_clear_interrupt_scheme): it
+	 * masks all hardware vectors and clears the pcompleter map entry,
+	 * so no MSI-X message can be delivered after vectors are released
+	 * at detach.
+	 *
+	 * For non-control PFs this is a polling-mode mailbox RPC
+	 * (IRQ_RDY already cleared by disable above, or never set).
+	 *
+	 * Note: pci_clear_master() runs in nbl_probe() core_start_err path;
+	 * on a non-control PF it cannot stop DMA using the management PF's
+	 * BDF, this is a known limitation.
+	 */
+	if (map_ready) {
+		cleanup_ret = nbl_dev_destroy_msix_map(dev_mgt);
+		if (cleanup_ret)
+			dev_err(dev_mgt->common->dev,
+				"rollback: destroy MSI-X map failed: %d\n",
+				cleanup_ret);
+	}
+	nbl_dev_clear_interrupt_scheme(dev_mgt);
+
+	cancel_work_sync(&common_dev->clean_mbx_task);
+	return ret;
+}
+
+void nbl_dev_stop(struct nbl_adapter *adapter)
+{
+	struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+	struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+	int ret;
+
+	ret = nbl_dev_disable_mailbox_irq(dev_mgt);
+	if (ret)
+		dev_err(dev_mgt->common->dev,
+			"Failed to disable mailbox IRQ: %d\n", ret);
+	nbl_dev_free_mailbox_irq(dev_mgt);
+
+	/*
+	 * Destroy the hardware MSI-X map now, before the kernel vectors
+	 * are released at device detach (they deliberately stay
+	 * allocated until then - see nbl_dev_clear_interrupt_scheme).
+	 * Masks all device vectors and clears the pcompleter map entry.
+	 *
+	 * Best-effort: if the destroy RPC fails the hardware map stays
+	 * valid until vectors are freed at detach; stale entries rely on
+	 * firmware cleanup on chip reset.
+	 *
+	 * pci_clear_master() on a non-control PF cannot stop DMA using the
+	 * management PF's BDF.
+	 */
+	ret = nbl_dev_destroy_msix_map(dev_mgt);
+	if (ret)
+		dev_err(dev_mgt->common->dev,
+			"Failed to destroy MSI-X map: %d\n", ret);
+
+	nbl_dev_clear_interrupt_scheme(dev_mgt);
+
+	/*
+	 * destroy_msix_map() sends ack-requested messages which may
+	 * requeue clean_mbx_task via polling send path.  Drain work
+	 * after the operation.
+	 */
+	cancel_work_sync(&common_dev->clean_mbx_task);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
index 51cf04e4c552..a66c633a0e7a 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
@@ -10,5 +10,7 @@ struct nbl_adapter;
 
 int nbl_dev_init(struct nbl_adapter *adapter);
 void nbl_dev_remove(struct nbl_adapter *adapter);
+int nbl_dev_start(struct nbl_adapter *adapter);
+void nbl_dev_stop(struct nbl_adapter *adapter);
 
 #endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index af430d3dfb70..f16d79504e5d 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -16,6 +16,16 @@
 #include "nbl_include/nbl_def_common.h"
 #include "nbl_core.h"
 
+int nbl_core_start(struct nbl_adapter *adapter)
+{
+	return nbl_dev_start(adapter);
+}
+
+void nbl_core_stop(struct nbl_adapter *adapter)
+{
+	nbl_dev_stop(adapter);
+}
+
 struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
 				  struct nbl_init_param *param)
 {
@@ -212,7 +222,13 @@ static int nbl_probe(struct pci_dev *pdev,
 		goto adapter_init_err;
 	}
 	pci_set_drvdata(pdev, adapter);
+	err = nbl_core_start(adapter);
+	if (err)
+		goto core_start_err;
+
 	return 0;
+core_start_err:
+	nbl_core_remove(adapter);
 adapter_init_err:
 chip_deps_err:
 	pci_clear_master(pdev);
@@ -227,6 +243,8 @@ static void nbl_remove(struct pci_dev *pdev)
 		return;
 
 	pci_set_drvdata(pdev, NULL);
+
+	nbl_core_stop(adapter);
 	nbl_core_remove(adapter);
 
 	pci_clear_master(pdev);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 1/8] net/nebula-matrix: add channel layer
  2026-09-28 12:32 ` [PATCH v30 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found 7 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 6

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] nbl_chan_alloc_all_tx_bufs() and nbl_chan_alloc_all_rx_bufs()
  make a separate dmam_alloc_coherent(dev, 4096, ...) call for each of…
- [Low] Most of the channel layer (setup_queue, teardown_queue, send_msg,
  cfg_chan_qinfo_map_table, clean_queue_subtask, register_chan_task,…
- [Low] Some specific claims in the commit message don't match the code.
- [Low] The get_host_pf_mask kernel-doc says unsupported masks are
  'rejected with -EINVAL', and nbl_chan_cfg_qinfo_map_table() defers…
- [Low] The lock comments in struct nbl_chan_info are wrong.
- [Low] send_msg, clean_queue and set_queue_state do not check
  chan_info->active.
- [Low] If setup fails partway, its devres/dmam allocations are orphaned
  when setup is retried.

--

net/nebula-matrix: add channel layer

This adds a mailbox-based channel layer to the nbl driver. It carries
control messages between PF0 and the other PFs. The patch adds the
channel ops table, TX/RX ring handling with ACK tracking, a workqueue for
RX cleanup, and the BAR0/BAR2 hw_ops that program the mailbox queues.

> - TX slot concurrency control, returning -EAGAIN when slots are full to
>   avoid congestion and resource exhaustion. Supports 16B inline small
>   payload and 4KB max DMA-based large payload transmission.

[ ... ]

> - Isolated hw_ops layer decouples low-level hardware implementation from
>   upper channel logic. Fine-grained register locking optimizes BAR0/BAR2
>   access for stability under high stress.

[Severity: Low]
Do these claims match the code?

nbl_chan_get_msg_id() returns -EAGAIN when the ACK wait slots run out.
When the descriptor ring is full, though, nbl_chan_update_txqueue()
returns -EBUSY:

	if (nbl_chan_txq_full(txq, chan_info->num_txq_entries))
		return -EBUSY;
	if (param->arg_len > NBL_CHAN_BUF_LEN - sizeof(*tx_desc))
		return -EINVAL;

That same check caps the payload at 4096 - 64 = 4032 bytes, not 4KB.
nbl_chan_send_ack() allows 12 bytes less than that.

On locking, the patch adds only one reg_lock spinlock, and only the BAR0
helpers take it (nbl_hw_rd_regs_lock() and nbl_hw_cfg_mailbox_qinfo()).
nbl_hw_read_mbx_regs(), nbl_hw_write_mbx_regs() and the doorbell write in
nbl_hw_update_mailbox_queue_tail_ptr() go through BAR2 without any
register lock.

Could the commit message be updated to match?

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
> new file mode 100644
> index 000000000000..d2c8182c72b6
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
> @@ -0,0 +1,1467 @@

[ ... ]

> +static int nbl_chan_alloc_all_tx_bufs(struct nbl_channel_mgt *chan_mgt,
> +				      struct nbl_chan_info *chan_info)
> +{
> +	struct nbl_chan_ring *txq = &chan_info->txq;
> +	struct device *dev = chan_mgt->common->dev;
> +	struct nbl_chan_buf *buf;
> +	u16 i;
> +
> +	for (i = 0; i < chan_info->num_txq_entries; i++) {
> +		buf = &txq->buf[i];
> +		buf->va = dmam_alloc_coherent(dev, chan_info->txq_buf_size,
> +					      &buf->pa, GFP_KERNEL);

[Severity: Medium]
Is one coherent allocation per ring entry intended here?
nbl_chan_alloc_all_rx_bufs() does the same, so each PF makes 512 separate
dmam_alloc_coherent(dev, 4096, ...) calls.

Coherent allocations are rounded up to whole pages. On 16K or 64K page
kernels (arm64, ppc64le), one PF's mailbox uses 8 MiB or 32 MiB of
coherent memory, not the ~2 MiB the comment in nbl_chan_setup_queue()
assumes. The driver can build for these, since Kconfig only requires
64BIT and !CPU_BIG_ENDIAN. A 4-PF card with 64K pages would use 128 MiB.
Each call also adds its own devres node and IOMMU mapping.

Could this be one contiguous allocation split into 4K slots, or a
dma_pool?

[ ... ]

> +	/*
> +	 * DMA resources are allocated once and released only by devres
> +	 * at device detach.  A teardown does not free them, so a second
> +	 * setup would orphan ~2 MiB of coherent DMA per channel.
> +	 */
> +	if (chan_info->dma_allocated) {
> +		dev_warn(common->dev,
> +			 "channel DMA already allocated, re-setup not supported\n");
> +		return -EBUSY;
> +	}
> +	nbl_chan_init_queue_param(chan_info, NBL_CHAN_QUEUE_LEN,
> +				  NBL_CHAN_QUEUE_LEN, NBL_CHAN_BUF_LEN,
> +				  NBL_CHAN_BUF_LEN);
> +	err = nbl_chan_init_queue(common, chan_info);
> +	if (err)
> +		return err;
> +	err = nbl_chan_alloc_all_bufs(chan_mgt, chan_info);
> +	if (err)
> +		return err;
> +	nbl_chan_config_queue(chan_mgt, chan_info, true); /* tx */
> +	nbl_chan_config_queue(chan_mgt, chan_info, false); /* rx */
> +	nbl_chan_update_tail_ptr(hw_ops, chan_mgt->hw_ops_tbl->priv,
> +				 rxq->tail_ptr, NBL_MB_RX_QID);
> +	WRITE_ONCE(chan_info->active, true);
> +	chan_info->dma_allocated = true;
> +	return 0;
> +}

[Severity: Low]
What happens if nbl_chan_setup_queue() fails partway, for example in
nbl_chan_alloc_all_rx_bufs()? At that point active and dma_allocated are
both still false. A second call would get past both checks and allocate
everything again. The first set of dmam buffers would stay allocated
until detach.

In the rest of the series, setup_queue runs once per probe from
nbl_dev_setup_common_dev(), and a failure fails the probe. So no retry
path exists today.

Should dma_allocated be set before the first allocation, so the guard
does what its comment says?

[ ... ]

> +static int nbl_chan_send_msg(struct nbl_channel_mgt *chan_mgt,
> +			     struct nbl_chan_send_info *chan_send)
> +{

[ ... ]

> +	mutex_lock(&chan_info->state_lock);
> +	if (READ_ONCE(chan_info->shutdn)) {
> +		mutex_unlock(&chan_info->state_lock);
> +		return -ESHUTDOWN;
> +	}
> +	atomic_inc(&chan_info->inflight_tx_cnt);
> +	mutex_unlock(&chan_info->state_lock);

[Severity: Low]
Can nbl_chan_send_msg() be called on a channel that hasn't gone through
setup_queue? Nothing here checks chan_info->active.

Before setup, num_txq_entries is 0, so this path divides by zero:

nbl_chan_update_txqueue()->nbl_chan_txq_full()
    NBL_NEXT_ID(next_to_use, num_entries - 1)

nbl_chan_clean_queue() has a similar problem: before setup it reads
through rxq->desc.rx_desc, which is still NULL. nbl_chan_set_queue_state()
walks wait[], which can be NULL after a partial setup failure.

In the rest of the series, every caller runs after
nbl_dev_setup_common_dev(), so this doesn't look reachable today. Would
it be worth checking active in these entry points anyway?

[ ... ]

> +static struct nbl_channel_ops chan_ops = {
> +	.send_msg			= nbl_chan_send_msg,
> +	.send_ack			= nbl_chan_send_ack,
> +	.register_msg			= nbl_chan_register_msg,
> +	.cfg_chan_qinfo_map_table	= nbl_chan_cfg_qinfo_map_table,
> +	.check_queue_exist		= nbl_chan_check_queue_exist,
> +	.setup_queue			= nbl_chan_setup_queue,
> +	.teardown_queue			= nbl_chan_teardown_queue,
> +	.clean_queue_subtask		= nbl_chan_clean_queue_subtask,
> +	.register_chan_task		= nbl_chan_register_chan_task,
> +	.set_queue_state		= nbl_chan_set_queue_state,
> +};

[Severity: Low]
This isn't a bug on its own, but most of these ops have no caller in
this patch: setup_queue, teardown_queue, send_msg,
cfg_chan_qinfo_map_table, clean_queue_subtask, register_chan_task and
set_queue_state. nbl_core_init() only calls nbl_chan_init_common().

The commit message describes probe-time DMA allocation and IRQ/polling
RX as if they already work here. Later patches in the series add the
callers in nbl_dev.c and nbl_dispatch.c.

Could the commit message say that queue setup and the RX paths are
wired up in later patches?

[ ... ]

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
> new file mode 100644
> index 000000000000..21cbb21946cf
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
> @@ -0,0 +1,181 @@

[ ... ]

> +struct nbl_chan_info {

[ ... ]

> +	/* Guards channel state bitmap and shutdn flag */
> +	struct mutex state_lock;
> +	/* Guards pending requests and pending work list operations */
> +	struct mutex pending_lock;

[Severity: Low]
Are these lock comments accurate? The state bitmap is never accessed
under state_lock. nbl_chan_set_queue_state() changes it without the lock:

	if (set)
		set_bit(state, chan_info->state);
	else
		clear_bit(state, chan_info->state);

nbl_chan_send_msg() also tests NBL_CHAN_IRQ_RDY without the lock. The
code relies on atomic bitops instead.

There is also no pending work list. pending_lock seems to guard the
wait[] slot array.

Could the comments be updated to say what each lock actually protects?

[ ... ]

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
> index ecbf440e4366..d87bf9d41a24 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
> @@ -10,6 +10,41 @@

[ ... ]

> +	/**
> +	 * get_host_pf_mask - Fetch host PF mask from firmware k_pf_mask reg
> +	 * @hw_mgt: hardware management context
> +	 * @pf_mask: output pointer for PF mask value
> +	 *
> +	 * k_pf_mask register rule:
> +	 *   bit N == 0 -> PF#N enabled; bit N == 1 -> PF#N masked out.
> +	 *   bit0 is PF0's mask bit (not reserved); PF0 can be masked but
> +	 *   the driver requires at least PF0 enabled.
> +	 *   Only 1/2/4 PFs are supported:
> +	 *     1 PF  (PF0):     mask = 0xfe
> +	 *     2 PFs (PF0,PF1): mask = 0xfc
> +	 *     4 PFs (PF0~PF3): mask = 0xf0
> +	 *   All-zero mask (0x00) means all 8 PFs enabled, which is
> +	 *   unsupported by the driver and rejected with -EINVAL.
> +	 */
> +	void (*get_host_pf_mask)(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask);

[Severity: Low]
Where does the -EINVAL rejection described here happen? The comment in
nbl_chan_cfg_qinfo_map_table() also defers to a check this patch doesn't
contain:

	 * Non-contiguous or unsupported PF count
	 * will be rejected in resource initialization logic.

Later in the series, "net/nebula-matrix: add common resource
implementation" adds nbl_res_init_pf_num(), which does this check. It
runs before cfg_chan_qinfo_map_table is called, so at the end of the
series only these comments are affected.

Could the comments point to where the check lives, or be added in the
same patch as the check?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 2/8] net/nebula-matrix: add common resource implementation
  2026-09-28 12:32 ` [PATCH v30 net-next 2/8] net/nebula-matrix: add common resource implementation illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Baseline: the out-of-bounds MMIO hazard is not reachable, but the
  comment is wrong.

--

net/nebula-matrix: add common resource implementation

This patch adds a chip-agnostic resource layer for the Nebula Matrix
driver. It builds read-only PF topology, SR-IOV BDF, Ethernet port and
VSI tables on the control PF at probe time. It also adds several new
hw_ops, including flush_write, get_real_bus, get_fw_eth_map and
get_board_info.

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> index eeff6216e4aa..4d3477f70bcc 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> @@ -97,6 +97,32 @@ static void nbl_hw_rd_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
>  	spin_unlock(&hw_mgt->reg_lock);
>  }
>  
> +/*
> + * Only call this when has_ctrl=true, which maps enough space
> + * (bar_len - 8192) to cover NBL_HW_DUMMY_REG (0x1300904).
> + * The flow/design guarantees this is only called in the
> + * has_ctrl path.
> + */
> +static void nbl_flush_writes(struct nbl_hw_mgt *hw_mgt)
> +{
> +	nbl_hw_rd32(hw_mgt, NBL_HW_DUMMY_REG);
> +}

[Severity: Low]
Is the "(bar_len - 8192)" part of this comment accurate? The has_ctrl
path in nbl_hw_init_leonis() does not size the mapping from bar_len:

    hw_size = PAGE_ALIGN_DOWN(NBL_MEM_BAR_TOTAL_SIZE -
                              NBL_RDMA_NOTIFY_LEN);
    hw_mgt->hw_addr =
        pcim_iomap(pdev, NBL_MEMORY_BAR,
                   hw_size);

The size is fixed and does not depend on bar_len. It is also smaller
than 64 MiB - 8 KiB when PAGE_SIZE is larger than 8 KiB.

The has_ctrl-only rule is also written only in this comment. The
flush_write member of struct nbl_hw_ops in nbl_def_hw.h says nothing
about it, and nbl_hw_setup_ops() installs the op for every PF. That
includes net-only PFs, which map only NBL_REG_NET_ONLY_LEN (8 KiB). For
those PFs, NBL_HW_DUMMY_REG (0x1300904) is well outside the mapping.

This patch adds no callers. Later in the series, the flush_write callers
in nbl_interrupt.c all seem to sit behind has_ctrl checks, so the
out-of-bounds read does not look reachable. Could the comment be
corrected? Could the has_ctrl requirement also be documented on the
flush_write op in nbl_def_hw.h, or enforced in the code?

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 3/8] net/nebula-matrix: add intr resource implementation
  2026-09-28 12:32 ` [PATCH v30 net-next 3/8] net/nebula-matrix: add intr " illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found 7 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 6

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] Nothing clears the chip-internal MSI-X routing tables at init:
  NBL_PCOMPLETER_FUNCTION_MSIX_MAP (per-function VALID, table DMA…
- [Low] In nbl_res_intr_cfg_msix_map(), a reconfiguration allocates the
  whole new vector set from intr_net_bmap/intr_other_bmap (the…
- [Low] In nbl_hw_cfg_msix_map(), the valid branch has the comment "clear
  VALID first, prevent torn read of partial entry", but the code writes…
- [Low] Several comments describe teardown synchronization the code does
  not implement.
- [Low] struct nbl_interrupt_mgt (nbl_resource.h) embeds
  func_intr_res[NBL_MAX_FUNC] with NBL_MAX_FUNC = NBL_MAX_PF +…
- [Low] The public entry points nbl_res_intr_destroy_msix_map() and
  nbl_res_intr_set_mailbox_irq() read res_mgt->intr_mgt without a lock,…
- [Low] Before dma_free_coherent() of the MSI-X map table
  (__nbl_res_intr_destroy_msix_map ->…

--

net/nebula-matrix: add intr resource implementation

This patch adds an nbl_interrupt module. For the control PF, it manages
the chip-internal MSI-X vector index space with separate net and other
bitmaps, and it programs the chip MSI-X map, info and FID tables. It also
adds the cfg_msix_map, destroy_msix_map and set_mailbox_irq resource ops,
and a global nbl_intr_mgt_stop() teardown called from
nbl_res_remove_leonis().

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> index 4d3477f70bcc..acd4f3dd0757 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c

[ ... ]

> @@ -123,6 +137,139 @@ static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)

[ ... ]

> +static void nbl_hw_cfg_msix_map(struct nbl_hw_mgt *hw_mgt, u16 func_id,
> +				bool valid, dma_addr_t dma_addr, u8 bus,
> +				u8 devid, u8 function)
> +{
> +	struct nbl_function_msix_map function_msix_map;
> +
> +	memset(&function_msix_map, 0, sizeof(function_msix_map));
> +	if (valid) {
> +		/* clear VALID first, prevent torn read of partial entry */
> +		function_msix_map.data[0] = 0;
> +		function_msix_map.data[1] = 0;
> +		function_msix_map.data[2] = 0;
> +		nbl_hw_wr_regs_lock(hw_mgt,
> +				    NBL_PCOMPLETER_FUNCTION_MSIX_MAP(func_id),
> +				    function_msix_map.data,
> +				    sizeof(function_msix_map));

[Severity: Low]
Does this write actually clear VALID first? nbl_hw_wr_regs_lock() writes
data[0] through data[3] in ascending order, and the VALID bit is in
data[2]:

nbl_hw_leonis.h:
/* data[2] */
...
#define NBL_FUNCTION_MSIX_MAP_VALID_MASK BIT(16)

If the entry is already VALID when this branch runs, the two DMA address
words get zeroed while the old VALID bit is still set. VALID is cleared
third, not first.

That leaves a window where the hardware could see VALID=1 with a zero or
half-written address, which is the torn state the comment says it
prevents. The teardown branch below avoids this by keeping the live
address while it clears VALID.

This needs the entry to be VALID on entry. After prepare/Phase 2 that
only happens with stale state from a previous kernel (kexec or forced
unload).

Would the comment hold if data[2] were cleared on its own first, or if
this first write kept the old address?

> +
> +		/* program full entry and set VALID */

[ ... ]

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
> new file mode 100644
> index 000000000000..c41cfa14f90f
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
> @@ -0,0 +1,703 @@

[ ... ]

> +static int __nbl_res_intr_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
> +					   u16 func_id)
> +{
> +	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
> +	int ret;
> +
> +	lockdep_assert_held(&intr_mgt->lock);
> +
> +	if (intr_mgt->stopping)
> +		return -ESHUTDOWN;
> +
> +	ret = __nbl_res_intr_prepare_destroy_msix_map(res_mgt, func_id);
> +	if (ret)
> +		return ret;
> +	/*
> +	 * prepare() only transitions CONFIGURED functions; an IDLE func
> +	 * has nothing to wait for or complete.
> +	 */
> +	if (intr_mgt->func_intr_res[func_id].state !=
> +	    NBL_INTR_FUNC_DESTROYING)
> +		return 0;
> +
> +	usleep_range(NBL_MSIX_DMA_SYNC_MIN_US, NBL_MSIX_DMA_SYNC_MAX_US);
> +
> +	return __nbl_res_intr_complete_destroy_msix_map(res_mgt, func_id);

[Severity: Low]
Is a fixed 1ms sleep enough to guarantee that the pcompleter has stopped
fetching from the table before
__nbl_res_intr_complete_destroy_msix_map() calls dma_free_coherent() on
it?

The only ordering here is: clear FUNCTION_MSIX_MAP VALID, call
flush_write(), then call usleep_range(1000, 1200). Nothing polls a
hardware idle or completion status.

nbl_intr_mgt_stop() uses the same wait and calls it "Best-effort only".
The reconfig path in nbl_res_intr_cfg_msix_map() also relies on it before
recycling vectors and rewriting the live table.

If a table fetch is still in flight after the window, could it hit an
unmapped IOVA or reused memory? Is there a status register that could be
polled instead?

> +}
> +
> +int nbl_res_intr_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
> +				  u16 func_id)
> +{
> +	struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
> +	int ret;
> +
> +	if (!intr_mgt)
> +		return -EINVAL;
> +
> +	mutex_lock(&intr_mgt->lock);
> +	ret = __nbl_res_intr_destroy_msix_map(res_mgt, func_id);

[Severity: Low]
This reads res_mgt->intr_mgt without a lock and then locks the captured
object. __nbl_res_intr_destroy_msix_map() above does not use that
pointer; it loads res_mgt->intr_mgt again, then dereferences it for
lockdep_assert_held() and intr_mgt->stopping.
__nbl_res_intr_set_mailbox_irq() does the same, and its enable path
dereferences intr_mgt->stopping.

nbl_intr_mgt_stop() clears the pointer while it holds the mutex:

	res_mgt->intr_mgt = NULL;
	mutex_unlock(&intr_mgt->lock);

Suppose a caller was blocked on the mutex during stop's Phase 2. Wouldn't
it take the lock, read NULL in the helper, and dereference it?

And if a caller captured the pointer before nbl_res_remove_leonis()
returned, wouldn't it lock a mutex in devres-freed memory?

With the teardown ordering at the end of the series, no concurrent caller
seems possible. nbl_dev_remove() tears down the mailbox and cancels
clean_mbx_task before nbl_res_remove_leonis() runs, and device links
unbind sibling PFs first. The stop() comment does expect racing callers,
though.

Could the helpers take intr_mgt as a parameter from the locked caller?

> +	mutex_unlock(&intr_mgt->lock);
> +
> +	return ret;
> +}

[ ... ]

> +	/* Allocate net interrupt vectors */
> +	for (i = 0; i < num_net_msix; i++) {
> +		intr_index = find_first_zero_bit(intr_mgt->intr_net_bmap,
> +						 NBL_MAX_NET_INTERRUPT);
> +		if (intr_index == NBL_MAX_NET_INTERRUPT) {
> +			dev_err(dev, "No free net interrupt vectors left\n");
> +			ret = -EAGAIN;
> +			goto release_vecs_unlock;
> +		}

[Severity: Low]
On reconfiguration, the whole new vector set is taken from intr_net_bmap
and intr_other_bmap here while the function's old vectors are still set.
The old ones are only released later, in the had_config block:

		nbl_intr_release_bitmap(res_mgt, old_interrupts, old_num);

A reconfig therefore needs free room for the whole new request on top of
the old one. When the pool is nearly full, even a reconfig of the same
size or smaller fails with -EAGAIN, although the old config stays intact.

This follows from the "Only tear down old hardware state after new
allocation succeeds" design. In-tree callers at the end of the series ask
for one other vector per PF (nbl_dev_init_msix_cnt()), so they can't use
up the pool. Is this capacity limit intended?

[ ... ]

> +static struct nbl_interrupt_mgt *nbl_intr_setup_mgt(struct device *dev)
> +{
> +	struct nbl_interrupt_mgt *intr_mgt;
> +	int err;
> +
> +	intr_mgt = devm_kzalloc(dev, sizeof(*intr_mgt), GFP_KERNEL);
> +	if (!intr_mgt)
> +		return ERR_PTR(-ENOMEM);
> +
> +	err = devm_mutex_init(dev, &intr_mgt->lock);
> +	if (err)
> +		return ERR_PTR(err);
> +
> +	intr_mgt->stopping = false;
> +	bitmap_zero(intr_mgt->intr_net_bmap, NBL_MAX_NET_INTERRUPT);
> +	bitmap_zero(intr_mgt->intr_other_bmap, NBL_MAX_OTHER_INTERRUPT);

[Severity: Medium]
Only software state is reset here. Does anything clear the chip-internal
tables at init, namely NBL_PCOMPLETER_FUNCTION_MSIX_MAP,
NBL_PADPT_HOST_MSIX_INFO and NBL_PCOMPLETER_HOST_MSIX_FID_TABLE?

The comment above nbl_hw_set_mailbox_irq() says this chip state survives
kexec/forced unload without FLR. The pci_driver also has no .shutdown
callback.

After kexec or kdump, the old kernel's VALID map entries would still
point at DMA table addresses the new kernel no longer owns. INFO/FID
entries would also remain for gvecs that the new bitmaps treat as free.

nbl_intr_mgt_stop() only cleans functions whose software state is
CONFIGURED, so this stale state is never cleaned up.

On a fresh config (had_config == false), nbl_res_intr_cfg_msix_map() also
arms the new INFO/FID entries before it overwrites the map entry, which
may still be VALID from the old kernel:

	for (i = 0; i < requested; i++) {
		...
		hw_ops->cfg_msix_info(res_mgt->hw_ops_tbl->priv,
				      func_id, true, gvec, ...);
	}
	...
	hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
			     true, official_tbl->dma, ...);

That doesn't match the "not yet valid (on fresh config)" comment above
it.

Could the pcompleter then DMA-read a map table at an old-kernel address,
or route a stale map into gvecs that were just given to another function?

In this series, nbl_hw_cfg_mailbox_qinfo() disarms mailbox routing, so
whether this triggers depends on how the hardware fetches and on future
interrupt sources.

Would it make sense to invalidate these tables for all functions during
control-PF init, the same way the mailbox MSIX fields are scrubbed?

> +
> +	return intr_mgt;
> +}

[ ... ]

> +	/*
> +	 * Phase 1: batch invalidate all hardware MSIX map entries.
> +	 * stopping is set under the lock, so any caller racing with the
> +	 * quiesce window below either holds the lock and sees stopping
> +	 * at its next checkpoint, or acquires it after this phase and
> +	 * fails (-ESHUTDOWN/-EBUSY/-ENODEV) before issuing MMIO.
> +	 */
> +	mutex_lock(&intr_mgt->lock);
> +	intr_mgt->stopping = true;

[ ... ]

> +	mutex_unlock(&intr_mgt->lock);
> +
> +	/*
> +	 * Global quiesce: wait for straggler DMA table reads after all
> +	 * MSIX map entries have been invalidated in hardware, before
> +	 * freeing coherent memory. Best-effort only.
> +	 */
> +	usleep_range(NBL_MSIX_DMA_SYNC_MIN_US, NBL_MSIX_DMA_SYNC_MAX_US);
> +
> +	/* Phase2: safely release MSIX coherent memory and intr resources */
> +	mutex_lock(&intr_mgt->lock);

[ ... ]

> +	/* Clear the published pointer under the lock, last */
> +	res_mgt->intr_mgt = NULL;
> +	mutex_unlock(&intr_mgt->lock);
> +}

[Severity: Low]
Some comments here and in nbl_resource.h don't seem to match the code.

The enum nbl_intr_func_state comment says:

 * the hardware-DMA quiesce window of a teardown; the lock is held across
 * that window so a concurrent configuration cannot install a map that

However, nbl_intr_mgt_stop() drops the lock after Phase 1, sleeps, and
takes it again for Phase 2. In that window, correctness depends on the
stopping flag, which the comment doesn't mention.

The Phase 1 comment says racing callers fail with -ESHUTDOWN/-EBUSY/-ENODEV
"before issuing MMIO". But the en_msix=false branch of
__nbl_res_intr_set_mailbox_irq() runs before any stopping check:

	if (!en_msix) {
		hw_ops->set_mailbox_irq(res_mgt->hw_ops_tbl->priv,
					func_id, false, 0);
		hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
		return 0;
	}

So that branch writes NBL_MAILBOX_QINFO_MAP_REG_ARR and returns 0. After
stop completes, callers get -EINVAL, which the comment doesn't list
either.

"Clear the published pointer under the lock, last" suggests that readers
synchronize on intr_mgt->lock. Every public entry point reads
res_mgt->intr_mgt without holding it, though.

Could these comments be updated to describe the actual contract?

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
> index ae0a3d33198d..c14a78a47c98 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
> @@ -17,6 +17,53 @@

[ ... ]

> +struct nbl_interrupt_mgt {
> +	struct mutex lock; /* Protects bitmap + func_intr_res[] */
> +	DECLARE_BITMAP(intr_net_bmap, NBL_MAX_NET_INTERRUPT);
> +	DECLARE_BITMAP(intr_other_bmap, NBL_MAX_OTHER_INTERRUPT);
> +	bool stopping; /* set on teardown, rejects new configurations */
> +	struct nbl_func_interrupt_resource_mng func_intr_res[NBL_MAX_FUNC];
> +};

[Severity: Low]
This isn't a bug, but func_intr_res[] is sized by NBL_MAX_FUNC
(NBL_MAX_PF + NBL_MAX_VF = 520). Each entry is about 48 bytes on 64-bit,
so struct nbl_interrupt_mgt comes to about 25 KB. That makes the
devm_kzalloc() in nbl_intr_setup_mgt() a physically contiguous order-3
allocation.

Only PF entries can ever be used. nbl_res_func_id_to_bdf() returns
-EOPNOTSUPP for rel_pf_id >= max_pf before func_intr_res[] is touched.
nbl_intr_mgt_stop() also walks all 520 entries twice.

Could this be sized by NBL_MAX_PF, or allocated with devm_kvzalloc()?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation
  2026-09-28 12:32 ` [PATCH v30 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 4 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] nbl_probe_chip_deps() (drivers/net/ethernet/nebula-
  matrix/nbl/nbl_main.c:113) adds a managed device link from each non-…
- [Medium] nbl_hw_deinit_module() (nbl_hw_leonis.c:721) clears
  driver_status, which starts asynchronous firmware cleanup, then…
- [Medium] nbl_probe_chip_deps() checks the supplier state without
  holding the device-links lock (READ_ONCE(mgt->dev.links.status) !=…
- [Medium] nbl_probe_chip_deps() dereferences the managed link returned
  by device_link_add() (READ_ONCE(link->status), nbl_main.c:161) without…
- [Low] nbl_probe_chip_deps() treats func 0 as a ready nbl management PF
  as soon as READ_ONCE(mgt->dev.links.status) == DL_DEV_DRIVER_BOUND…

--

net/nebula-matrix: add chip-wide hardware init/deinit implementation

This commit adds Leonis chip-wide datapath init/deinit hooks
(init_module/deinit_module) in the hw and resource layers. The hooks use
a driver_status flag handshake with firmware. It also adds a device link
from each non-management PF to func 0, so that sibling PFs are unbound
before the management PF.

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> index acd4f3dd0757..47ec995e3ac8 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> @@ -5,11 +5,23 @@
>  #include <linux/device.h>
>  #include <linux/pci.h>
>  #include <linux/bits.h>
> +#include <linux/delay.h>
>  #include <linux/io.h>
>  #include <linux/spinlock.h>
>  #include <linux/bitfield.h>
>  #include "nbl_hw_leonis.h"
>  
> +/*
> + * Firmware cleanup after driver_status=false is asynchronous and the
> + * current hardware revision exposes no cleanup-complete status bit.
> + * Wait a bounded window so the firmware pass finishes before this
> + * function returns, establishing an explicit boundary against a later
> + * init_module() that would otherwise reprogram per-PF/chip-wide
> + * registers while firmware is still wiping them. Best-effort only.
> + */
> +#define NBL_FW_CLEANUP_SYNC_MIN_US	2000
> +#define NBL_FW_CLEANUP_SYNC_MAX_US	3000

[ ... ]

> +/*
> + * Setting driver status to false notifies firmware to clean up per-PF
> + * hardware state such as qinfo registers.
> + *

[ ... ]

> + * Firmware cleanup is asynchronous with no completion status register,
> + * so a bounded sync delay follows the flush. On return the firmware
> + * pass has settled, so a subsequent init_module() cannot race it.
> + *

[ ... ]

> +static void nbl_hw_deinit_module(struct nbl_hw_mgt *hw_mgt)
> +{
> +	nbl_hw_set_driver_status(hw_mgt, false);
> +	/* ensure driver_status reaches the chip */
> +	nbl_flush_writes(hw_mgt);
> +	/*
> +	 * Allow the asynchronous firmware cleanup pass to finish before
> +	 * returning, so this function is the boundary between firmware
> +	 * per-PF teardown and any later driver reprogramming.
> +	 */
> +	usleep_range(NBL_FW_CLEANUP_SYNC_MIN_US, NBL_FW_CLEANUP_SYNC_MAX_US);
> +}

[Severity: Medium]

Is the guarantee in the nbl_hw_deinit_module() kernel-doc accurate? It
says "On return the firmware pass has settled, so a subsequent
init_module() cannot race it". But the comment above
NBL_FW_CLEANUP_SYNC_MIN_US says there is no cleanup-complete status bit
and calls the wait "Best-effort only".

The only synchronisation here is a fixed 2-3 ms sleep. There is no
firmware acknowledgement.

If firmware cleanup takes longer than that, could a quick PF0
unbind/rebind reprogram qinfo in nbl_dev_setup_chan_qinfo() and the
datapath in nbl_hw_init_module(), and then have firmware wipe that state?

nbl_hw_init_module() also sets driver_status again without checking that
firmware is idle. It does not handle a driver_status bit left set by a
previous instance that crashed or was kexec'd.

The commit message also says:

  - Explicit register flushing synchronizes with asynchronous firmware cleanup

and:

  The deinit path only clears driver_status and flushes writes.

The read of NBL_HW_DUMMY_REG only ensures the posted write reached the
device. It can't observe firmware completion. The deinit path also
sleeps, which the message doesn't mention.

Could the commit message and both comments match what the code actually
guarantees? This is still present at the end of the series.

[ ... ]

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> index 1aafed2d46d7..e1a30b5ba0cd 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> @@ -83,6 +83,87 @@ static void nbl_get_func_param(struct pci_dev *pdev, kernel_ulong_t driver_data,
>  		param->caps.has_ctrl = 1;
>  }
>  
> +/*
> + * Establish chip-wide dependencies for this PF:
> + *  - for every non-management PF, add a consumer->management PF device
> + *    link. The driver core then guarantees (sysfs unbind, driver
> + *    unregister, hot-unplug alike) that this consumer is released
> + *    BEFORE the func 0 supplier, which is the only teardown order in
> + *    which the chip-global firmware deinit is safe.
> + *

[ ... ]

> +static int nbl_probe_chip_deps(struct pci_dev *pdev, bool has_ctrl)
> +{
> +	struct device_link *link;
> +	struct pci_dev *mgt;
> +
> +	if (has_ctrl)
> +		return 0;
> +
> +	mgt = pci_get_domain_bus_and_slot(pci_domain_nr(pdev->bus),
> +					  pdev->bus->number,
> +					  PCI_DEVFN(PCI_SLOT(pdev->devfn), 0));
> +	if (!mgt) {
> +		dev_err(&pdev->dev,
> +			"management PF (func 0) not found on this chip\n");
> +		return -ENODEV;
> +	}
> +
> +	/*
> +	 * The management PF must have COMPLETED probing, not merely
> +	 * have a driver assigned.  While func 0 is still PROBING, its

[ ... ]

> +	 */
> +	if (READ_ONCE(mgt->dev.links.status) != DL_DEV_DRIVER_BOUND) {
> +		pci_dev_put(mgt);
> +		return -EPROBE_DEFER;
> +	}

[Severity: Low]

Does this check confirm that func 0 is bound to nbl specifically?
links.status is DL_DEV_DRIVER_BOUND for any bound driver. If func 0 is
bound to vfio-pci or pci-stub through driver_override, the check passes
and the sibling creates a link to that other driver.

nbl_core_init() would then run on the non-management PF even though nbl
never did the func 0 chip-global setup. That includes the mailbox QINFO
map from nbl_dev_setup_chan_qinfo() under has_ctrl.

The comment above says "The management PF must have COMPLETED probing",
which reads as the nbl probe. The check doesn't enforce that.

Later in the series, nbl_dev_start()->nbl_dev_rpc_timeout() turns the
mailbox RPC timeout into -EPROBE_DEFER. The sibling would then likely
keep deferring, with an RPC timeout on every retry.

> +
> +	link = device_link_add(&pdev->dev, &mgt->dev,
> +			       DL_FLAG_AUTOREMOVE_CONSUMER);

[Severity: Medium]

What happens to the sibling PFs when func 0 is unbound and then bound
again?

On a func 0 unbind, device_links_unbind_consumers() releases every
sibling PF driver. device_links_driver_cleanup() then deletes the
AUTOREMOVE_CONSUMER link.

When func 0 binds again, device_links_driver_bound() re-queues consumers
only through AUTOPROBE links:

drivers/base/core.c:device_links_driver_bound() {
    ...
	if (device_link_test(link, DL_FLAG_AUTOPROBE_CONSUMER))
		driver_deferred_probe_add(link->consumer);
    ...
}

device_link_add() also returns NULL when DL_FLAG_AUTOPROBE_CONSUMER is
combined with DL_FLAG_AUTOREMOVE_CONSUMER.

So would PF1..PF3 and their Ethernet ports stay unbound until an admin
rebinds each one by hand?

The commit message ("ensuring safe chip-global firmware deinit") and the
comment above nbl_probe_chip_deps() describe only the teardown ordering.
Neither mentions that the ports are not restored. Would a persistent
managed link with DL_FLAG_AUTOPROBE_CONSUMER work better here?

At this commit, nbl_res_chip_deinit_module() has no caller yet. The later
commit "net/nebula-matrix: add common/ctrl dev init/remove operation"
wires deinit_module into nbl_dev_remove_ctrl_dev(). The missing re-probe
is still there at the end of the series.

> +	pci_dev_put(mgt);
> +	if (!link) {
> +		dev_err(&pdev->dev,
> +			"failed to create device link to management PF\n");
> +		return -ENOMEM;
> +	}
> +
> +	/*
> +	 * If func 0 began unbinding between the driver check above and
> +	 * link creation, the core's unbind-consumers pass has already
> +	 * run and missed this new link, which starts SUPPLIER_UNBIND.
> +	 * Fail the probe: the managed link is auto-purged and the core
> +	 * retries after func 0 is bound again.
> +	 */
> +	if (READ_ONCE(link->status) == DL_STATE_SUPPLIER_UNBIND)
> +		return -EPROBE_DEFER;

[Severity: Medium]

Can func 0 change state between the lockless check of
mgt->dev.links.status and device_link_add()? Func 0 unbind and func N
probe hold different device locks. In that window func 0 can go
UNBINDING -> NO_DRIVER, and possibly back to PROBING.

device_link_init_status() would then set the new link to one of:

  supplier NO_DRIVER                   -> DL_STATE_DORMANT
  supplier PROBING, consumer probing   -> DL_STATE_CONSUMER_PROBE

This check rejects neither state.

In the DORMANT case, the sibling finishes probe with no bound management
PF. device_links_driver_bound() then hits:

	WARN_ON(link->status != DL_STATE_CONSUMER_PROBE);
	WRITE_ONCE(link->status, DL_STATE_ACTIVE);

In the CONSUMER_PROBE case, the sibling's mailbox setup races func 0's
chip-global init. The comment above the links.status check is meant to
prevent exactly that.

If func 0's probe then fails, device_links_no_driver() moves the link to
DORMANT without unbinding the sibling. A later func 0 bind makes the link
AVAILABLE. device_links_unbind_consumers() only force-unbinds ACTIVE
consumers, so a later func 0 unbind would run deinit_module while the
sibling is still bound and doing mailbox DMA.

Is there a way to close this window, rather than rejecting only
DL_STATE_SUPPLIER_UNBIND after the fact? This is still present at the end
of the series.

[Severity: Medium]

Is it safe to dereference link here? The device_link_add() kernel-doc
says that for a managed link the caller "is handing the management of the
link over to the driver core entirely and its return value can only be
used to check whether or not the link is present".

Suppose func 0 is being removed and has already passed
device_links_busy(). The new link starts in SUPPLIER_UNBIND and does not
block the removal.

If this probe thread is preempted after device_link_add() returns, func 0
removal can continue:

device_del()
  device_links_purge()
    __device_link_del()
...
device_link_release_fn()
  device_link_synchronize_removal()
  kfree(link)

Would READ_ONCE(link->status) then read freed memory? Nothing here holds
the device-links SRCU read lock or a reference on the link. This is still
present at the end of the series.

> +
> +	return 0;
> +}

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure
  2026-09-28 12:32 ` [PATCH v30 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The commit message of "net/nebula-matrix: dispatch: add control-
  level routing core infrastructure" says the patch will "Add wire ABI…
- [Low] At this commit, the kernel-doc Warning paragraph for struct
  nbl_dispatch_ops (nbl_def_dispatch.h lines 25-31) describes five…
- [Low] The kernel-doc for struct nbl_dispatch_ops (nbl_def_dispatch.h)
  calls @init_module "dispatch layer initialization" and @deinit_module…

--

net/nebula-matrix: dispatch: add control-level routing core infrastructure

This commit adds the nbl dispatch layer. It allocates a devm-managed
nbl_dispatch_mgt, an ops table and a ctrl_lvl bitmap. When has_ctrl is
set, it enables the MGT level and installs init_module/deinit_module
wrappers that forward to the resource ops.

> MGT level is only enabled for Control PF. Add kerneldoc for caller
> permission constraints. Add wire ABI structures with static_assert for
> upcoming mailbox RPC.

[Severity: Low]
Is the last sentence accurate for this patch? The diff touches the
Makefile, nbl_core.h, nbl_dispatch.c, nbl_dispatch.h, nbl_def_dispatch.h
and nbl_main.c. None of these files gets a mailbox wire-format struct or
a static_assert.

At this commit, the driver has one static_assert, for struct
nbl_board_port_info in nbl_def_channel.h. It was already there, and this
patch does not change it.

The RPC payload structs and their size checks (nbl_chan_param_cfg_msix_map,
nbl_chan_param_set_mailbox_irq, nbl_chan_param_get_vsi_id,
nbl_chan_param_get_eth_id) first appear in the next patch,
"net/nebula-matrix: dispatch: implement channel RPC framework and
serialize hardware ops".

Could this sentence be dropped here, or moved to that patch's commit
message?

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
> new file mode 100644
> index 000000000000..b3398591035b
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
> @@ -0,0 +1,45 @@

[ ... ]

> +/**
> + * struct nbl_dispatch_ops - dispatch control plane operation callbacks
> + * @init_module: dispatch layer initialization, control-PF exclusive,
> + *               caller must check has_ctrl guard
> + * @deinit_module: dispatch layer cleanup, control-PF exclusive,
> + *                 caller must check has_ctrl guard

[Severity: Low]
Do "dispatch layer initialization" and "dispatch layer cleanup" describe
these callbacks correctly? All dispatch-layer setup happens in
nbl_disp_init(), and nbl_disp_remove() is empty. The installed callbacks
only forward to the resource ops:

nbl_disp_init_module()
    res_ops->init_module(p)          /* nbl_res_chip_init_module() */
        hw_ops->init_module(p, eth_speed, eth_num)

nbl_disp_deinit_module()
    res_ops->deinit_module(p)        /* nbl_res_chip_deinit_module() */
        hw_ops->deinit_module(res_mgt->hw_ops_tbl->priv)

On leonis, these calls program the chip-wide datapath and set or clear
driver_status. Clearing driver_status starts asynchronous firmware
cleanup. nbl_dev.c notes that mailbox DMA must be drained before
deinit_module is called.

Would it be clearer to call these chip/firmware init and deinit, and to
mention the ordering requirement? That way later callers won't treat
deinit_module as software-only cleanup. The same wording is still there
at the end of the series.

> + *
> + * Warning: init_module/deinit_module are control-PF exclusive. The five
> + * resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq,
> + * get_vsi_id, get_eth_id) are PF-only and resolve to either a local
> + * resource call (control PF) or a mailbox RPC (non-control PF with
> + * has_net). A function with neither has_ctrl nor has_net leaves these
> + * pointers NULL; callers must not invoke them on such functions. VFs are
> + * rejected by the responders with -EPERM.
> + */
> +struct nbl_dispatch_ops {
> +	int (*init_module)(struct nbl_dispatch_mgt *disp_mgt);
> +	void (*deinit_module)(struct nbl_dispatch_mgt *disp_mgt);
> +};

[Severity: Low]
Is this paragraph describing code that isn't in this patch yet? Here,
struct nbl_dispatch_ops has only init_module and deinit_module, and the
enum has only NBL_DISP_CTRL_LVL_MGT. nbl_disp_init() never checks
has_net, and nbl_disp_refresh_ctrl_ops() installs only the MGT pointers:

	if (test_bit(NBL_DISP_CTRL_LVL_MGT, disp_mgt->ctrl_lvl)) {
		disp_ops->init_module = nbl_disp_init_module;
		disp_ops->deinit_module = nbl_disp_deinit_module;
	}

The five ops, NBL_DISP_CTRL_LVL_NET, the has_net routing and the -EPERM
responders all come in the next patch, "net/nebula-matrix: dispatch:
implement channel RPC framework and serialize hardware ops". By the end
of the series, the doc matches the code.

Could this paragraph be added in that patch instead, so the documentation
in each commit matches its code?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops
  2026-09-28 12:32 ` [PATCH v30 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 4

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The commit message describes three changes that are not in
  this patch.
- [Low] The commit message contradicts the error codes the patch actually
  returns.
- [Low] Concern (baseline: not a valid finding at end of series):
  nbl_disp_remove()'s new comment describes a teardown sequence that…
- [Low] The new disp_mgt->ops_mutex_lock is redundant with the resource
  layer's own locking, which contradicts the commit message's claim that…
- [Low] Concern (baseline: not a valid finding):
  nbl_disp_chan_cfg_msix_map_resp() passes the peer's num_net_msix and…

--

net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops

This patch adds mailbox RPC request and response handlers for five
dispatch operations: cfg_msix_map, destroy_msix_map, set_mailbox_irq,
get_vsi_id and get_eth_id. The management PF calls the resource ops
directly, and net-capable non-management PFs forward them to the
management PF over the channel. It also adds a per-dispatch mutex around
the MSI-X and mailbox IRQ operations, which change hardware state.

> Fix error propagation by forwarding native Linux errnos from remote
> resource operations instead of unconditionally returning -EREMOTEIO.
> Only truncated ACK responses return -EREMOTEIO to distinguish protocol
> errors from legitimate operation failures. Transport layer errors are
> passed through unchanged from channel send routines.

[Severity: Low]
Does this paragraph match the code? Both new ACK length checks, in
nbl_disp_chan_get_vsi_id_req() and nbl_disp_chan_get_eth_id_req(),
return -EBADMSG, not -EREMOTEIO:

    if (chan_send.ack_len != sizeof(result)) {
        ...
        return -EBADMSG;
    }

The only -EREMOTEIO in the driver appears to be in
nbl_chan_recv_ack_msg() in nbl_channel/nbl_channel.c. There it replaces
an out-of-range remote errno, so it has nothing to do with a truncated
ACK:

    if (raw_err > 0 || raw_err < -MAX_ERRNO)
        raw_err = -EREMOTEIO;

The parent commit also has no RPC request or response code. So there is
no earlier "unconditionally returning -EREMOTEIO" behaviour for this
patch to fix.

Was this paragraph left over from an earlier revision?

> - Other functions: all dispatch ops are null and return -EOPNOTSUPP.

[ ... ]

> For MSIX map reconfiguration, pre-allocate DMA buffers and vector indices
> prior to tearing down old configurations to prevent interrupt loss.
> Destroy path clears mailbox MSIX routing entries upfront to avoid stale
> interrupt triggers from recycled hardware vectors.

[ ... ]

> Narrow kernel-doc warning to clarify new ops are PF-capability bounded,
> not available for VF functions.

[Severity: Medium]
Several parts of the commit message describe changes that this patch
does not contain.

For a function with neither the MGT bit nor the NET bit set,
nbl_disp_refresh_ctrl_ops() memsets disp_ops and assigns nothing, so the
pointers stay NULL. Calling one of those ops is a NULL function pointer
call, not an -EOPNOTSUPP return. The kernel-doc in nbl_def_dispatch.h
already says callers must not invoke them, which matches the code.

The MSI-X pre-allocation, and the clearing of the mailbox routing at the
start of destroy, live in nbl_res_intr_cfg_msix_map() and
__nbl_res_intr_prepare_destroy_msix_map() in nbl_hw/nbl_interrupt.c. The
earlier commit "net/nebula-matrix: add intr resource implementation"
added that code, and this patch does not touch nbl_hw/.

The "Warning:" paragraph in the struct nbl_dispatch_ops kernel-doc is
identical in the parent commit, which already listed the five ops and
the VF restriction. This patch only adds the @member lines.

Could the commit message describe only what this patch changes?

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
> index 966fee2dec8b..e6a5e53e9d2d 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c

[ ... ]

> @@ -25,6 +194,357 @@ static int nbl_disp_init_module(struct nbl_dispatch_mgt *disp_mgt)
>  	return -EOPNOTSUPP;
>  }
>  
> +static int nbl_disp_cfg_msix_map(struct nbl_dispatch_mgt *disp_mgt,
> +				 u16 num_net_msix, u16 num_others_msix,
> +				 bool net_msix_mask_en)
> +{
> +	struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
> +	struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
> +	struct nbl_common_info *common = disp_mgt->common;
> +	int ret;
> +
> +	if (!res_ops->cfg_msix_map)
> +		return -EOPNOTSUPP;
> +	mutex_lock(&disp_mgt->ops_mutex_lock);
> +	ret = res_ops->cfg_msix_map(p, common->mgt_pf, num_net_msix,
> +					  num_others_msix, net_msix_mask_en);
> +	mutex_unlock(&disp_mgt->ops_mutex_lock);
> +	return ret;
> +}

[Severity: Low]
Is ops_mutex_lock needed here? The resource ops already serialize
themselves under intr_mgt->lock. The documentation of
nbl_res_intr_cfg_msix_map() says:

 * Serialization: this function takes intr_mgt->lock internally to
 * protect the global vector bitmaps and per-function state against
 * concurrent callers.

nbl_res_intr_set_mailbox_irq() also takes intr_mgt->lock internally.
nbl_res_intr_destroy_msix_map() takes it around
__nbl_res_intr_destroy_msix_map(). The "add intr resource
implementation" commit also says callers need no upper-layer locking.

This wrapper makes a single resource call. So do
nbl_disp_destroy_msix_map(), nbl_disp_set_mailbox_irq() and the three
matching *_resp handlers. None of them has a multi-step sequence that
needs an outer lock.

The field comment in nbl_dispatch.h, "use for the caller not in
interrupt", does not say what the mutex protects. That makes the commit
message's claim that the mutex eliminates concurrent modification races
hard to square with the code.

Could the mutex be dropped? If it covers state that intr_mgt->lock does
not, could that state be documented?

[ ... ]

> +static void nbl_disp_chan_cfg_msix_map_resp(void *priv, u16 src_id, u16 msg_id,
> +					    void *data, u32 data_len)
> +{

[ ... ]

> +	if (res_ops->cfg_msix_map) {
> +		mutex_lock(&disp_mgt->ops_mutex_lock);
> +		ret = res_ops->cfg_msix_map(p, src_id,
> +					    le16_to_cpu(param.num_net_msix),
> +					    le16_to_cpu(param.num_others_msix),
> +					    !!le16_to_cpu(param.msix_mask_en));
> +		mutex_unlock(&disp_mgt->ops_mutex_lock);

[Severity: Low]
num_net_msix and num_others_msix come straight from the peer PF and go
to nbl_res_intr_cfg_msix_map(). The only bound there appears to be the
table size:

    if (num_net_msix > NBL_MSIX_MAP_TABLE_MAX_ENTRIES ||
        num_others_msix > NBL_MSIX_MAP_TABLE_MAX_ENTRIES) {

NBL_MSIX_MAP_TABLE_MAX_ENTRIES (1024) equals the size of the other
interrupt pool (NBL_MAX_OTHER_INTERRUPT). Could one PF request nearly
all of the other vectors and leave the remaining PFs with -EAGAIN?

This probably can't be reached in practice:

- The rel_pf_id < max_pf check rejects VFs.
- nbl_get_func_param() in nbl_main.c says the hardware and firmware
  forbid PF passthrough to guests. Every peer PF therefore runs this
  same host driver.
- The in-tree requester nbl_dev_cfg_msix_map() only asks for the counts
  from nbl_dev_init_msix_cnt(): one mailbox vector and no net vectors.

Is a per-function limit worth adding anyway, or is a peer PF always
treated as trusted?

[ ... ]

> @@ -104,14 +641,40 @@ int nbl_disp_init(struct nbl_adapter *adapter)

[ ... ]

>  void nbl_disp_remove(struct nbl_adapter *adapter)
>  {
> -	/* Dispatch structures are allocated via devm */
> +	/*
> +	 * Dispatch structures are devm-allocated and freed at detach.
> +	 *
> +	 * The five responders registered by nbl_disp_setup_msg() are
> +	 * owned by the channel layer (xarray of handlers) and are never
> +	 * unregistered here. This is safe because the teardown order
> +	 * guarantees no responder can run after this point:
> +	 *
> +	 *   mailbox teardown
> +	 *     -> cancel_work_sync(clean_mbx_task)   // drain RX work
> +	 *     -> nbl_chan_teardown_queue()          // stop HW queue,
> +	 *                                                 // join clean task,
> +	 *                                                 // active=false
> +	 *     -> nbl_chan_remove_common()
> +	 *          -> destroy_wq()                       // no new work
> +	 *          -> nbl_chan_remove_msg_handler()      // free handler nodes
> +	 *
> +	 * By the time devres frees disp_mgt, the mailbox queue is stopped
> +	 * and the handler xarray is empty, so no responder can touch
> +	 * res_mgt->intr_mgt after nbl_intr_mgt_stop() has cleared it.
> +	 */
>  }

[Severity: Low]
Does this teardown sequence exist at this commit? clean_mbx_task is not
defined yet, and nothing calls the channel setup_queue or teardown_queue
ops. So the "cancel_work_sync(clean_mbx_task) ->
nbl_chan_teardown_queue()" chain refers to code that doesn't exist yet
at this point in the series.

This is not a reachable race here, because the mailbox queue is never
set up and the responders can't run.

The next patch, "net/nebula-matrix: add common/ctrl dev init/remove
operation", makes the comment accurate. It adds nbl_dev_remove() ahead
of nbl_disp_remove() in nbl_core_remove(). Its
nbl_dev_remove_common_dev() calls
cancel_work_sync(&common_dev->clean_mbx_task) and then
nbl_chan_teardown_queue(), both before nbl_res_remove_leonis() calls
nbl_intr_mgt_stop().

Would it make sense to move this comment into that patch, so the comment
matches the code at each point in the series?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation
  2026-09-28 12:32 ` [PATCH v30 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 3

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The control-PF teardown and init added here rely on
  nbl_probe_chip_deps() to guarantee that no sibling PF probes unless…
- [Low] The commit message says 'The firmware handles global hardware
  cleanup and QINFO routing state reclamation asynchronously after…
- [Low] The new callers disagree with how nbl_chan_teardown_queue()
  actually behaves.
- [Low] Two comments on the same control-PF remove sequence contradict
  each other.

--

net/nebula-matrix: add common/ctrl dev init/remove operation

This commit adds nbl_dev.c with nbl_dev_init() and nbl_dev_remove(). They
set up the mailbox channel queue and cleanup work, program mailbox QINFO
routing on the control PF, and call chip-level init_module/deinit_module.
Teardown destroys the mailbox queue before notifying firmware through
deinit_module.

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> new file mode 100644
> index 000000000000..e094b97acdfb
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> @@ -0,0 +1,253 @@

[ ... ]

> +/* ----------  Dev init process  ---------- */
> +static int nbl_dev_setup_common_dev(struct nbl_adapter *adapter)
> +{

[ ... ]

> +	ret = nbl_dev_setup_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
> +	if (ret)
> +		goto err_cleanup;

[ ... ]

> +	return 0;
> +err_cleanup:
> +	cancel_work_sync(&common_dev->clean_mbx_task);
> +	nbl_dev_remove_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
> +	nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX, NULL);
> +	return ret;
> +}
> +
> +static void nbl_dev_remove_common_dev(struct nbl_adapter *adapter)
> +{
> +	struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
> +	struct nbl_dev_common *common_dev = dev_mgt->common_dev;
> +	int ret;
> +
> +	if (!common_dev)
> +		return;
> +	cancel_work_sync(&common_dev->clean_mbx_task);
> +	ret = nbl_dev_remove_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
> +	if (ret)
> +		dev_warn(&adapter->pdev->dev, "mailbox queue teardown failed, inflight DMA may exist: %d\n",
> +			 ret);

[Severity: Low]
Can nbl_dev_remove_chan_queue() return non-zero here? Both exits of
nbl_chan_teardown_queue() return 0: the early return when the channel is
not active, and the normal path after draining. That makes this dev_warn()
unreachable.

The message also suggests teardown can fail with DMA still in flight. If
it could, nbl_dev_remove() would still go on to call deinit_module, and
nbl_hw_deinit_module() assumes no mailbox DMA is left at that point. Does
this warning cover a real failure mode, or can it be dropped?

There is a related issue on the err_cleanup path of
nbl_dev_setup_common_dev() above. That path only runs when setup_queue
failed, and chan_info->active is only set to true when setup succeeds. So
the teardown lands here:

nbl_chan_teardown_queue() {
	if (!READ_ONCE(chan_info->active)) {
		dev_warn(chan_mgt->common->dev, "channel not active, skip duplicate teardown\n");
		return 0;
	}
	...
}

Won't this log "skip duplicate teardown" on the first cleanup after a
failed setup, when no duplicate teardown has happened?

[ ... ]

> +/*
> + * Tear down control device: deinit_module sets driver_status=false
> + * to notify firmware to clean all per-PF hardware state (including
> + * qinfo registers).  The qinfo map programmed in nbl_dev_init() via
> + * nbl_dev_setup_chan_qinfo() is not explicitly cleared; firmware
> + * handles it on driver_status change.
> + *
> + * Teardown ordering guarantee: every non-management PF creates a
> + * consumer->control PF device link in its probe path, so the driver
> + * core always unbinds all siblings before allowing the control PF to
> + * be detached (sysfs unbind, driver unregister and hot-unplug alike).

[Severity: Low]
This comment conflicts with the one in nbl_res_remove_leonis(), which
nbl_core_remove() calls right after nbl_dev_remove():

nbl_res_remove_leonis() {
	...
	/*
	 * Tear down all MSI-X maps before destroying coherent tables.
	 * This is critical on the control PF, which may hold
	 * maps for remote PFs that are still bound.
	 */
	...
}

If the driver core always unbinds all siblings before the control PF, no
remote PF can still be bound at that point. Which of these does the
teardown actually rely on? One of the two comments probably needs
updating.

This patch relies on the stronger claim to justify calling deinit_module
unconditionally. The next comment describes a case where it does not
fully hold.

> + * nbl_probe_chip_deps() defers a sibling's probe until the control PF
> + * is fully bound (DL_DEV_DRIVER_BOUND), so no sibling can race its
> + * mailbox setup against this deinit.

[Severity: Medium]
Is this exclusion guaranteed? nbl_probe_chip_deps() in nbl_main.c checks
the supplier state without a lock, then creates the link, and rejects only
one link state:

	if (READ_ONCE(mgt->dev.links.status) != DL_DEV_DRIVER_BOUND) {
		pci_dev_put(mgt);
		return -EPROBE_DEFER;
	}

	link = device_link_add(&pdev->dev, &mgt->dev,
			       DL_FLAG_AUTOREMOVE_CONSUMER);
	...
	if (READ_ONCE(link->status) == DL_STATE_SUPPLIER_UNBIND)
		return -EPROBE_DEFER;

device_link_init_status() in drivers/base/core.c sets
DL_STATE_SUPPLIER_UNBIND only while the supplier is DL_DEV_UNBINDING:

	case DL_DEV_UNBINDING:
		link->status = DL_STATE_SUPPLIER_UNBIND;
		break;
	default:
		link->status = DL_STATE_DORMANT;

Two cases seem to get past the check:

(a) The control PF finishes unbinding, including nbl_dev_remove() ->
deinit_module, between the check and device_link_add(). It is then
DL_DEV_NO_DRIVER, so the link starts as DL_STATE_DORMANT and passes.

(b) The control PF has started re-probing (DL_DEV_PROBING) while the
sibling is probing. The link starts as DL_STATE_CONSUMER_PROBE and also
passes.

In both cases the sibling goes on to nbl_core_init() -> nbl_dev_init()
and sets up its mailbox queue. In (a) this happens right after the
control PF set driver_status=false, while the firmware's asynchronous
per-PF cleanup may still be running. In (b) it races the control PF's
nbl_dev_setup_chan_qinfo() and init_module programming.

The sibling and the control PF hold separate device_lock()s, so the
driver core does not serialize these paths.

The commit message makes the same claim:

  Device link dependency established during non-control PF probe ensures
  sibling PFs are unbound before the control PF is removed.

Should nbl_probe_chip_deps() re-check the supplier state after
device_link_add() and also reject DORMANT links and links to a supplier
that is still probing?

[ ... ]

> +/*
> + * Teardown order: Stop mailbox channel and drain all inflight DMA first,
> + * then invoke deinit_module to notify firmware.
> + *
> + * This intentionally breaks strict init/teardown mirror symmetry due to
> + * hardware constraint: firmware may perform asynchronous global hardware
> + * cleanup once driver_status=false is set. We must guarantee no ongoing
> + * mailbox DMA before deinit_module to avoid invalid DMA write.

[Severity: Low]
The commit message says:

  The firmware handles global hardware cleanup and QINFO routing state
  reclamation asynchronously after driver status is marked inactive.

This comment likewise refers to "asynchronous global hardware cleanup".
Is that accurate? This path reaches nbl_hw_deinit_module() in
nbl_hw_leonis.c through deinit_module, and that function documents a
narrower contract:

 * Setting driver status to false notifies firmware to clean up per-PF
 * hardware state such as qinfo registers.
 *
 * Note: firmware does NOT automatically revert chip-wide registers
 * configured in this init flow. Those chip-wide settings remain valid
 * until chip reset or explicitly overwritten by driver.

The earlier "add chip-wide hardware init/deinit implementation" commit
also says that firmware does not restore chip-wide datapath registers on
deinit.

So after the control PF is removed, the registers programmed by
nbl_dp_init() and nbl_intf_init() stay in place until chip reset. Could
the commit message and this comment say per-PF cleanup instead of global
cleanup?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH v30 net-next 8/8] net/nebula-matrix: add common dev start/stop operation
  2026-09-28 12:32 ` [PATCH v30 net-next 8/8] net/nebula-matrix: add common dev start/stop operation illusion.wang
@ 2026-10-02  3:35   ` netdev-bot+sashiko
  0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  3:35 UTC (permalink / raw)
  To: illusion.wang
  Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
	linux-kernel

Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 2 · Low: 3

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] nbl_dev_rpc_timeout() in drivers/net/ethernet/nebula-
  matrix/nbl/nbl_core/nbl_dev.c (lines 85-90) turns every -ETIMEDOUT…
- [Medium] In nbl_dev_start() (nbl_dev.c:513-516), map_ready is set only
  after cfg_msix_map returns success.
- [Low] The commit message says teardown is best-effort 'with
  pci_clear_master() serving as the final safety net' when the MSI-X…
- [Low] Two problems here.
- [Low] The commit message says clearing NBL_CHAN_IRQ_RDY before masking
  the hardware interrupt is done 'preventing in-flight ACK discard and…

--

net/nebula-matrix: add common dev start/stop operation

This patch adds nbl_dev_start() and nbl_dev_stop(). Start configures the
device MSI-X map, fetches the VSI/ETH ids, allocates MSI-X vectors and sets
up the mailbox interrupt; stop does the matching teardown. Startup RPC
timeouts on non-control PFs become -EPROBE_DEFER, and both functions are
called from nbl_probe()/nbl_remove() through nbl_core_start()/nbl_core_stop().

> Teardown is best-effort: failed MSI-X destroy RPC leaves stale hardware
> entries which are reclaimed by firmware on chip reset, with
> pci_clear_master() serving as the final safety net. The start/stop
> pair is single-shot and non-repeatable, tied strictly to PCI probe/remove
> device lifecycle.

[Severity: Low]
Is the pci_clear_master() safety net claim accurate for non-control PFs?
The comments this patch adds in nbl_dev_start() and nbl_dev_stop() say the
opposite:

 * pci_clear_master() on a non-control PF cannot stop DMA using the
 * management PF's BDF.

nbl_res_intr_cfg_msix_map() programs the PCOMPLETER map with the control
PF's common->hw_bus/devid/function. If destroy_msix_map fails, clearing bus
mastering on the departing PF would not stop those map-table fetches.

Could the commit message be changed to match the code comments?

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> index e094b97acdfb..79c62c169d70 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c

[ ... ]

> @@ -18,6 +29,233 @@ static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)

[ ... ]

> +/*
> + * Translate a mailbox RPC timeout on a non-control PF into deferred
> + * probing: the management PF/firmware is not responsive yet and the
> + * driver core should retry once PF0 is ready. Other errors, and all
> + * errors on the control PF itself, pass through unchanged.
> + */
> +static int nbl_dev_rpc_timeout(struct nbl_common_info *common, int ret)
> +{
> +	if (!common->has_ctrl && ret == -ETIMEDOUT)
> +		return -EPROBE_DEFER;
> +	return ret;
> +}

[Severity: Medium]
Can the "management PF/firmware is not responsive yet" case still happen
when these RPCs are sent? nbl_probe() calls nbl_probe_chip_deps() before
nbl_core_init()/nbl_core_start(), and that function already defers until
func 0 is bound:

nbl_probe_chip_deps() {
    ...
	if (READ_ONCE(mgt->dev.links.status) != DL_DEV_DRIVER_BOUND) {
		pci_dev_put(mgt);
		return -EPROBE_DEFER;
	}
    ...
}

DL_DEV_DRIVER_BOUND is set only after PF0's nbl_probe() returns 0. With
this patch, that probe includes PF0's own nbl_core_start() (cfg_msix_map,
request_irq and set_mailbox_irq(true)). The managed device link also stops
PF0 from unbinding while a consumer is probing.

So on a sibling PF, a -ETIMEDOUT from cfg_msix_map, get_vsi_id, get_eth_id
or set_mailbox_irq(true) comes after a 5-6 second polling wait. That looks
like a real failure, not PF0 still starting up.

Will returning -EPROBE_DEFER here actually cause a retry? really_probe()
only reports the deferral with dev_dbg(). Deferred devices are retried from
driver_bound()->driver_deferred_probe_trigger() when some other device
binds, and PF0's bind trigger has already fired by this point.

If nothing else binds later (the last sibling PF, or a module loaded after
boot), could the PF stay unbound with no probe failure printed at the
default log level? If other devices do bind, does each retry pay another
multi-second timeout for every startup RPC?

If this changes, the commit message wording "to trigger deferred probing
when the control PF is not ready" would need updating too.

[ ... ]

> +static int nbl_dev_disable_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
> +{

[ ... ]

> +	/*
> +	 * Disable sequence invariant: update software state first, then mask
> +	 * hardware interrupt. Must not reverse the order.
> +	 *
> +	 * If hardware interrupt is masked before clearing INTERRUPT_READY,
> +	 * the hardware may still transmit outstanding ACK packets for in-flight
> +	 * messages. Subsequent switch to polling mode discards pending ACK
> +	 * processing, triggering "Channel waiting ack failed" and "Skip ack
> +	 * with invalid status" errors.
> +	 *
> +	 * By entering polling mode first, any late hardware interrupts are
> +	 * ignored without pending ACK expectations, then hardware interrupt
> +	 * can be safely disabled.
> +	 *

[ ... ]

> +	 */
> +	chan_ops->set_queue_state(dev_mgt->chan_ops_tbl->priv,
> +				  NBL_CHAN_IRQ_RDY,
> +				  NBL_CHAN_TYPE_MAILBOX, false);
> +
> +	return disp_ops->set_mailbox_irq(dev_mgt->disp_ops_tbl->priv,
> +					 lvec, false);
> +}

[Severity: Low]
Does clearing NBL_CHAN_IRQ_RDY first actually keep in-flight ACKs? When
IRQ_RDY is cleared, nbl_chan_set_queue_state() wakes every waiter:

nbl_chan_set_queue_state() {
    ...
	if (!set && state == NBL_CHAN_IRQ_RDY) {
		for (i = 0; i < chan_info->num_txq_entries; i++)
			wake_up_all(&chan_info->wait[i].wait_queue);
	}
}

An interrupt-mode waiter in nbl_chan_send_msg() then gives up:

			if (!test_bit(NBL_CHAN_IRQ_RDY, chan_info->state)) {
				ret = -EIO;
				goto out_clear_wait_slot;
			}

The slot goes back to IDLE. When the peer's ACK arrives later,
nbl_chan_recv_ack_msg() sees w_status != NBL_MBX_STATUS_WAITING, logs
"Skip ack invalid status" and drops it. The comment says this ordering
prevents exactly that; the only difference is that the sender gets -EIO
instead of -ETIMEDOUT.

In this patch every ACK-requesting sender runs from probe/remove, and the
clean task only sends replies without ACKs, so nothing hits this today.

Should the comment and the commit message ("preventing in-flight ACK
discard and "Channel waiting ack failed" errors", "to preserve ACK
integrity") be corrected before later patches rely on this invariant?
Only draining in-flight waiters before clearing IRQ_RDY would provide that
guarantee.

[ ... ]

> @@ -251,3 +497,129 @@ void nbl_dev_remove(struct nbl_adapter *adapter)

[ ... ]

> +/* ----------  Dev start process  ---------- */
> +int nbl_dev_start(struct nbl_adapter *adapter)
> +{

[ ... ]

> +	ret = nbl_dev_rpc_timeout(common, nbl_dev_cfg_msix_map(dev_mgt));
> +	if (ret)
> +		goto err_destroy_map;
> +	map_ready = true;

[ ... ]

> +	ret = nbl_dev_enable_mailbox_irq(dev_mgt);
> +	if (ret)
> +		goto err_disable_irq;
> +
> +	return 0;
> +
> +err_disable_irq:
> +	cleanup_ret = nbl_dev_disable_mailbox_irq(dev_mgt);
> +	if (cleanup_ret)
> +		dev_err(dev_mgt->common->dev,
> +			"rollback: disable mailbox IRQ failed: %d\n",
> +			cleanup_ret);
> +	nbl_dev_free_mailbox_irq(dev_mgt);

[Severity: Low]
Is the set_mailbox_irq(false) RPC needed on this rollback path? If
nbl_dev_enable_mailbox_irq() failed, NBL_CHAN_IRQ_RDY was never set, yet
nbl_dev_disable_mailbox_irq() still sends a polling-mode RPC that requests
an ACK.

Control then falls through to err_destroy_map. On the control PF, the
destroy path already turns off mailbox routing:

__nbl_res_intr_prepare_destroy_msix_map() {
    ...
	/* Step 0: disable mailbox IRQ routing before tearing down map */
	ret = __nbl_res_intr_set_mailbox_irq(res_mgt, func, 0, false);
    ...
}

With an unresponsive peer, this adds another polling timeout of about 5-6
seconds, and every deferred retry repeats it. That works against the stated
goal of removing unnecessary polling timeouts during probe rollback.

Separately, some commit message wording seems to describe changes from an
earlier revision:

  "Skip redundant MSI-X map destroy RPC on unconfigured state during probe
  rollback to eliminate unnecessary polling timeouts. Adjust interrupt
  teardown ordering to ..."

This patch introduces nbl_dev_start()/nbl_dev_stop(), so should these
version notes go below the --- line instead?

> +err_destroy_map:
> +	/*
> +	 * Destroy the device-side MSI-X map only when it was configured.
> +	 * On a cfg RPC failure there is no known-good remote map; when
> +	 * the failure is a timeout against an unready/unresponsive
> +	 * management PF, issuing the destroy RPC would just burn another
> +	 * multi-second ACK timeout. Partial remote state is best-effort
> +	 * and reclaimed by firmware on chip reset.

[ ... ]

> +	 */
> +	if (map_ready) {
> +		cleanup_ret = nbl_dev_destroy_msix_map(dev_mgt);

[Severity: Medium]
If cfg_msix_map times out, can skipping destroy here leave control-PF
resources held for a function that has no driver? A -ETIMEDOUT from the
polling send only means the ACK did not arrive in time.
nbl_chan_send_msg() explicitly handles ACKs that arrive after the sender
has timed out, and the request may still be in the control PF's RX ring.

On the control PF, nbl_disp_chan_cfg_msix_map_resp()->
nbl_res_intr_cfg_msix_map() finishes all of its work before it sends the
ACK:

- kcalloc() for the interrupts array
- dma_alloc_coherent() for the map table
- set_bit() in intr_net_bmap/intr_other_bmap
- setting the PCOMPLETER map to VALID

and finally:

	intr_mgt->func_intr_res[func_id].state = NBL_INTR_FUNC_CONFIGURED;

With map_ready false, the sibling returns -EPROBE_DEFER without sending
destroy. The resources are then freed only when the same PF later
re-probes successfully (the had_config path) or when the control PF
unbinds (nbl_res_remove_leonis()->nbl_intr_mgt_stop()). A re-probe after
deferral is not guaranteed.

The comment "Partial remote state is best-effort and reclaimed by firmware
on chip reset" also looks inaccurate, as does the similar comment in
nbl_dev_stop(). This is host memory and bitmap state owned by the control
PF driver, and nothing requests a chip reset.

Would it be safer to always send destroy after a cfg timeout?
__nbl_res_intr_prepare_destroy_msix_map() returns 0 when the function is
not NBL_INTR_FUNC_CONFIGURED, so it looks harmless in the unconfigured
case.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928123237.88994-1-illusion.wang%40nebula-matrix.com

^ permalink raw reply	[flat|nested] 17+ messages in thread

end of thread, other threads:[~2026-10-02  3:35 UTC | newest]

Thread overview: 17+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 12:32 [PATCH v30 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
2026-09-28 12:32 ` [PATCH v30 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko
2026-09-28 12:32 ` [PATCH v30 net-next 2/8] net/nebula-matrix: add common resource implementation illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko
2026-09-28 12:32 ` [PATCH v30 net-next 3/8] net/nebula-matrix: add intr " illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko
2026-09-28 12:32 ` [PATCH v30 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko
2026-09-28 12:32 ` [PATCH v30 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko
2026-09-28 12:32 ` [PATCH v30 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko
2026-09-28 12:32 ` [PATCH v30 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko
2026-09-28 12:32 ` [PATCH v30 net-next 8/8] net/nebula-matrix: add common dev start/stop operation illusion.wang
2026-10-02  3:35   ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®