* [PATCH v31 net-next 1/8] net/nebula-matrix: add channel layer
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
2026-10-10 9:59 ` [PATCH v31 net-next 2/8] net/nebula-matrix: add common resource implementation illusion.wang
` (6 subsequent siblings)
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Add a mailbox-based channel layer that carries control messages between
PF0 and the other PFs, providing the inter-PF control transport the rest
of the driver builds on.
The channel ops table is defined here; queue setup, RX interrupt/polling
processing and the upper-layer callers are wired up in later patches of
this series. This patch also hooks channel init/cleanup into the PF
probe/remove paths in nbl_main.c.
Implementation notes:
- Message handlers live in an xarray, giving lockless lookup and
duplicate-registration rejection (xa_insert returns -EBUSY). Both
fire-and-forget and synchronous ACK-based sends are supported.
- ACK wait-slot exhaustion returns -EAGAIN; a full TX descriptor ring
returns -EBUSY. Payloads are inline (up to 16 B) or DMA-based (up to
4032 B).
- Each ring uses one contiguous coherent DMA block sliced into
fixed-size slots, so a channel costs one devres node and one IOMMU
mapping per ring instead of one per slot.
- Mailbox queue resources are devm-managed and the DMA is allocated once
during setup; the queue lifecycle covers init, configuration, start,
stop and teardown.
- RX is driven either from the interrupt or by polling, with cleanup
offloaded to a workqueue. Teardown quiesces both queues before the
rings can be freed.
- Register access goes through hw_ops. The reg_lock spinlock covers BAR0
register access only; BAR2 mailbox registers and the doorbell write do
not take it.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../net/ethernet/nebula-matrix/nbl/Makefile | 4 +-
.../nbl/nbl_channel/nbl_channel.c | 1486 +++++++++++++++++
.../nbl/nbl_channel/nbl_channel.h | 196 +++
.../nebula-matrix/nbl/nbl_common/nbl_common.c | 31 +
.../nebula-matrix/nbl/nbl_common/nbl_common.h | 14 +
.../net/ethernet/nebula-matrix/nbl/nbl_core.h | 7 +
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c | 234 +++
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h | 41 +
.../nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h | 32 +
.../nbl/nbl_include/nbl_def_channel.h | 125 ++
.../nbl/nbl_include/nbl_def_common.h | 4 +
.../nbl/nbl_include/nbl_def_hw.h | 41 +
.../nbl/nbl_include/nbl_include.h | 3 +
.../net/ethernet/nebula-matrix/nbl/nbl_main.c | 115 ++
14 files changed, 2332 insertions(+), 1 deletion(-)
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index cc060cf8bf75..04e1aa1fb4bd 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -3,5 +3,7 @@
obj-$(CONFIG_NBL) := nbl.o
-nbl-objs += nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
+nbl-objs += nbl_common/nbl_common.o \
+ nbl_channel/nbl_channel.o \
+ nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
new file mode 100644
index 000000000000..ac3115a3ab55
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
@@ -0,0 +1,1486 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/delay.h>
+#include <linux/device.h>
+#include <linux/mutex.h>
+#include <linux/bitfield.h>
+#include <linux/pci.h>
+#include <linux/bits.h>
+#include <linux/dma-mapping.h>
+#include <linux/atomic.h>
+#include <linux/wait.h>
+#include "nbl_channel.h"
+
+static int nbl_chan_add_msg_handler(struct nbl_channel_mgt *chan_mgt,
+ u16 msg_type, nbl_chan_resp func,
+ void *priv)
+{
+ struct nbl_chan_msg_node_data *handler;
+ int ret;
+
+ handler = kzalloc_obj(*handler, GFP_KERNEL);
+ if (!handler)
+ return -ENOMEM;
+
+ handler->func = func;
+ handler->priv = priv;
+
+ /* Each msg_type is registered at most once; reject duplicates. */
+ mutex_lock(&chan_mgt->handler_lock);
+ ret = xa_insert(&chan_mgt->handler_xa, msg_type, handler, GFP_KERNEL);
+ mutex_unlock(&chan_mgt->handler_lock);
+ if (ret)
+ kfree(handler);
+
+ return ret;
+}
+
+static int nbl_chan_init_msg_handler(struct nbl_channel_mgt *chan_mgt)
+{
+ int ret;
+
+ ret = devm_mutex_init(chan_mgt->common->dev, &chan_mgt->handler_lock);
+ if (ret)
+ return ret;
+
+ xa_init(&chan_mgt->handler_xa);
+
+ return 0;
+}
+
+static void nbl_chan_remove_msg_handler(struct nbl_channel_mgt *chan_mgt)
+{
+ struct nbl_chan_msg_node_data *handler;
+ unsigned long msg_type;
+
+ mutex_lock(&chan_mgt->handler_lock);
+ xa_for_each(&chan_mgt->handler_xa, msg_type, handler) {
+ xa_erase(&chan_mgt->handler_xa, msg_type);
+ kfree(handler);
+ }
+ mutex_unlock(&chan_mgt->handler_lock);
+ xa_destroy(&chan_mgt->handler_xa);
+}
+
+static void nbl_chan_init_queue_param(struct nbl_chan_info *chan_info,
+ u16 num_txq_entries, u16 num_rxq_entries,
+ u16 txq_buf_size, u16 rxq_buf_size)
+{
+ chan_info->num_txq_entries = num_txq_entries;
+ chan_info->num_rxq_entries = num_rxq_entries;
+ chan_info->txq_buf_size = txq_buf_size;
+ chan_info->rxq_buf_size = rxq_buf_size;
+ atomic_set(&chan_info->inflight_tx_cnt, 0);
+ WRITE_ONCE(chan_info->shutdn, false);
+ WRITE_ONCE(chan_info->active, false);
+ WRITE_ONCE(chan_info->wait_head_index, 0);
+ memset(chan_info->state, 0, sizeof(chan_info->state));
+ init_waitqueue_head(&chan_info->inflight_wait);
+}
+
+static int nbl_chan_init_tx_queue(struct nbl_common_info *common,
+ struct nbl_chan_info *chan_info)
+{
+ struct nbl_chan_ring *txq = &chan_info->txq;
+ struct device *dev = common->dev;
+ dma_addr_t base_dma;
+ size_t size =
+ chan_info->num_txq_entries * sizeof(struct nbl_chan_tx_desc);
+ void *base;
+ u16 i;
+
+ txq->desc.tx_desc = dmam_alloc_coherent(dev, size, &txq->dma,
+ GFP_KERNEL);
+ if (!txq->desc.tx_desc)
+ return -ENOMEM;
+
+ /*
+ * All per-entry buffers come from one contiguous coherent block, so a
+ * ring costs exactly one allocation, one devres node and one IOMMU
+ * mapping whatever the page size is. Per-entry dmam_alloc_coherent()
+ * instead rounds every NBL_CHAN_BUF_LEN buffer up to a whole page
+ * (8 MiB per PF on 16 KiB pages, 32 MiB on 64 KiB pages), and a
+ * dma_pool only fixes the large-page case: with NBL_CHAN_BUF_LEN equal
+ * to PAGE_SIZE it still degenerates to one allocation per buffer on
+ * 4 KiB pages. NBL_CHAN_BUF_LEN is a power of two, so the block
+ * always divides evenly into per-entry slots.
+ */
+ base = dmam_alloc_coherent(dev,
+ chan_info->num_txq_entries *
+ chan_info->txq_buf_size,
+ &base_dma, GFP_KERNEL);
+ if (!base)
+ return -ENOMEM;
+
+ chan_info->wait = devm_kcalloc(dev, chan_info->num_txq_entries,
+ sizeof(*chan_info->wait), GFP_KERNEL);
+ if (!chan_info->wait)
+ return -ENOMEM;
+ for (i = 0; i < chan_info->num_txq_entries; i++) {
+ init_waitqueue_head(&chan_info->wait[i].wait_queue);
+ WRITE_ONCE(chan_info->wait[i].status, NBL_MBX_STATUS_IDLE);
+ WRITE_ONCE(chan_info->wait[i].acked, 0);
+ WRITE_ONCE(chan_info->wait[i].ack_data, NULL);
+ WRITE_ONCE(chan_info->wait[i].ack_data_len, 0);
+ WRITE_ONCE(chan_info->wait[i].ack_err, 0);
+ WRITE_ONCE(chan_info->wait[i].msg_type, 0);
+ WRITE_ONCE(chan_info->wait[i].msg_index, 0);
+ WRITE_ONCE(chan_info->wait[i].dstid, 0);
+ }
+
+ txq->buf = devm_kcalloc(dev, chan_info->num_txq_entries,
+ sizeof(*txq->buf), GFP_KERNEL);
+ if (!txq->buf)
+ return -ENOMEM;
+
+ /* Carve the contiguous block into fixed-size per-entry slots */
+ for (i = 0; i < chan_info->num_txq_entries; i++) {
+ txq->buf[i].va = (u8 *)base +
+ (size_t)i * chan_info->txq_buf_size;
+ txq->buf[i].pa = base_dma +
+ (dma_addr_t)i * chan_info->txq_buf_size;
+ txq->buf[i].size = chan_info->txq_buf_size;
+ }
+
+ txq->next_to_clean = 0;
+ txq->next_to_use = 0;
+ txq->tail_ptr = 0;
+
+ return 0;
+}
+
+static int nbl_chan_init_rx_queue(struct nbl_common_info *common,
+ struct nbl_chan_info *chan_info)
+{
+ struct nbl_chan_ring *rxq = &chan_info->rxq;
+ struct device *dev = common->dev;
+ struct nbl_chan_rx_desc *desc;
+ dma_addr_t base_dma;
+ size_t size =
+ chan_info->num_rxq_entries * sizeof(struct nbl_chan_rx_desc);
+ void *base;
+ u16 i;
+
+ rxq->desc.rx_desc = dmam_alloc_coherent(dev, size, &rxq->dma,
+ GFP_KERNEL);
+ if (!rxq->desc.rx_desc) {
+ dev_err_ratelimited(dev,
+ "Allocate DMA for chan rx descriptor ring failed\n");
+ return -ENOMEM;
+ }
+
+ /* See nbl_chan_init_tx_queue() for why one contiguous block is used */
+ base = dmam_alloc_coherent(dev,
+ chan_info->num_rxq_entries *
+ chan_info->rxq_buf_size,
+ &base_dma, GFP_KERNEL);
+ if (!base)
+ return -ENOMEM;
+
+ rxq->buf = devm_kcalloc(dev, chan_info->num_rxq_entries,
+ sizeof(*rxq->buf), GFP_KERNEL);
+ if (!rxq->buf)
+ return -ENOMEM;
+
+ /* Carve the contiguous block into fixed-size per-entry slots */
+ for (i = 0; i < chan_info->num_rxq_entries; i++) {
+ rxq->buf[i].va = (u8 *)base +
+ (size_t)i * chan_info->rxq_buf_size;
+ rxq->buf[i].pa = base_dma +
+ (dma_addr_t)i * chan_info->rxq_buf_size;
+ rxq->buf[i].size = chan_info->rxq_buf_size;
+ }
+
+ desc = rxq->desc.rx_desc;
+ /*
+ * Initially leave one RX descriptor unused so that
+ * next_to_clean and next_to_use can distinguish an empty
+ * ring from a full ring.
+ *
+ * The unused slot is replenished as RX descriptors are
+ * consumed and recycled.
+ */
+ for (i = 0; i < chan_info->num_rxq_entries - 1; i++) {
+ desc[i].buf_addr = cpu_to_le64(rxq->buf[i].pa);
+ desc[i].buf_len = cpu_to_le32(chan_info->rxq_buf_size);
+ desc[i].flags = cpu_to_le16(BIT(NBL_CHAN_RX_DESC_AVAIL));
+ }
+
+ rxq->next_to_clean = 0;
+ rxq->next_to_use = chan_info->num_rxq_entries - 1;
+ rxq->tail_ptr = chan_info->num_rxq_entries - 1;
+
+ return 0;
+}
+
+static int nbl_chan_init_queue(struct nbl_common_info *common,
+ struct nbl_chan_info *chan_info)
+{
+ int err;
+
+ err = nbl_chan_init_tx_queue(common, chan_info);
+ if (err)
+ return err;
+
+ err = nbl_chan_init_rx_queue(common, chan_info);
+
+ return err;
+}
+
+static void nbl_chan_config_queue(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_info *chan_info, bool tx)
+{
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+ struct nbl_hw_mgt *p = chan_mgt->hw_ops_tbl->priv;
+ struct nbl_chan_ring *ring;
+ dma_addr_t addr;
+ int size_bwid;
+
+ if (tx)
+ ring = &chan_info->txq;
+ else
+ ring = &chan_info->rxq;
+ addr = ring->dma;
+ if (tx) {
+ size_bwid = ilog2(chan_info->num_txq_entries);
+ hw_ops->config_mailbox_txq(p, addr, size_bwid);
+ } else {
+ size_bwid = ilog2(chan_info->num_rxq_entries);
+ hw_ops->config_mailbox_rxq(p, addr, size_bwid);
+ }
+}
+
+static void nbl_chan_cfg_qinfo_map_table(struct nbl_channel_mgt *chan_mgt,
+ u8 bus, u8 devid)
+{
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+ struct nbl_hw_mgt *p = chan_mgt->hw_ops_tbl->priv;
+ u32 pf_mask = 0;
+ u8 func_id;
+
+ /*
+ * k_pf_mask rule: bit N == 0 means PF#N enabled, bit N == 1 masked out.
+ * Program mailbox QINFO entry for each hardware-active PF func_id.
+ *
+ * Note: This loop iterates over the raw hardware PF func_id.
+ * Product constraints limit supported PF counts to the contiguous
+ * sets PF0, PF0~1 and PF0~3. The mask is validated by
+ * nbl_res_init_pf_num() in the resource layer, which runs before
+ * this call and rejects any other mask with -EINVAL, so pf_mask
+ * here is always one of 0xfe/0xfc/0xf0.
+ */
+ hw_ops->get_host_pf_mask(p, &pf_mask);
+ for (func_id = 0; func_id < NBL_MAX_PF; func_id++) {
+ if (!(pf_mask & (1 << func_id)))
+ hw_ops->cfg_mailbox_qinfo(p, func_id, bus,
+ devid, func_id);
+ }
+}
+
+/*
+ * The RX QINFO region is written only by config/stop paths, never by
+ * senders (which only touch the TX QINFO region and the doorbell), so
+ * stopping RX does not require txq_lock.
+ */
+static void nbl_chan_stop_rx_queue(struct nbl_channel_mgt *chan_mgt)
+{
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+
+ hw_ops->stop_mailbox_rxq(chan_mgt->hw_ops_tbl->priv);
+}
+
+/*
+ * The TX QINFO region is also written by nbl_chan_quiesce_and_reclaim_tx()
+ * under txq_lock, so TX stop must serialize against it via txq_lock.
+ */
+static void nbl_chan_stop_tx_queue(struct nbl_channel_mgt *chan_mgt)
+{
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+
+ hw_ops->stop_mailbox_txq(chan_mgt->hw_ops_tbl->priv);
+}
+
+static void nbl_chan_reset_wait_head(struct nbl_chan_info *chan_info,
+ struct nbl_chan_waitqueue_head *wait_head)
+{
+ lockdep_assert_held(&chan_info->pending_lock);
+
+ WRITE_ONCE(wait_head->acked, 0);
+ WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_IDLE);
+ WRITE_ONCE(wait_head->ack_data, NULL);
+ WRITE_ONCE(wait_head->ack_data_len, 0);
+ WRITE_ONCE(wait_head->ack_err, 0);
+ WRITE_ONCE(wait_head->msg_type, 0);
+ WRITE_ONCE(wait_head->dstid, 0);
+}
+
+static int nbl_chan_teardown_queue(struct nbl_channel_mgt *chan_mgt,
+ u8 chan_type)
+{
+ struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+ struct nbl_chan_waitqueue_head *wait_head;
+ struct work_struct *task;
+ u16 i;
+
+ if (!READ_ONCE(chan_info->active)) {
+ dev_warn(chan_mgt->common->dev, "channel not active, skip duplicate teardown\n");
+ return 0;
+ }
+ /*
+ * Step1:
+ * block new sender
+ */
+ mutex_lock(&chan_info->state_lock);
+ WRITE_ONCE(chan_info->shutdn, true);
+ task = READ_ONCE(chan_info->clean_task);
+ WRITE_ONCE(chan_info->clean_task, NULL);
+ mutex_unlock(&chan_info->state_lock);
+ /*
+ * Step2:
+ * abort pending ACK waiters
+ *
+ * Poke only slots with a live waiter (status == WAITING): publish a
+ * synthetic (-EIO, len 0) completion and wake the owner. The slot is
+ * deliberately NOT reset to IDLE here - ownership stays with the
+ * sender, which consumes the completion and resets its own slot in
+ * nbl_chan_send_msg():out_clear_wait_slot under pending_lock.
+ *
+ * Resetting an owned slot here would corrupt two things:
+ * - an ACKD slot whose sender is consuming the completion
+ * locklessly (smp_rmb()-ordered reads) could observe torn
+ * ack_data_len/ack_err pairs;
+ * - nbl_chan_get_msg_id() hands out IDLE/TIMEOUT slots, so a slot
+ * reset to IDLE under its live owner could be reallocated to a
+ * sender admitted before shutdn was set.
+ *
+ * IDLE/TIMEOUT slots have no waiter and need no action. A late ACK
+ * racing this poke is serialized by pending_lock: either the real
+ * completion lands first (status becomes ACKD and is skipped) or
+ * this synthetic failure does, both are valid outcomes.
+ */
+ mutex_lock(&chan_info->pending_lock);
+ for (i = 0; i < chan_info->num_txq_entries; i++) {
+ wait_head = &chan_info->wait[i];
+ if (READ_ONCE(wait_head->status) != NBL_MBX_STATUS_WAITING)
+ continue;
+ WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_ACKD);
+ WRITE_ONCE(wait_head->ack_data_len, 0);
+ WRITE_ONCE(wait_head->ack_err, (s32)-EIO);
+ /* Publish completion fields before the acked flag */
+ smp_wmb();
+ WRITE_ONCE(wait_head->acked, 1);
+ wake_up(&wait_head->wait_queue);
+ }
+ mutex_unlock(&chan_info->pending_lock);
+
+ /*
+ * Step3:
+ * stop the RX queue unconditionally. The RX QINFO region is written
+ * only by config/stop paths, never by senders (which only touch the
+ * TX QINFO region and the doorbell), so no txq_lock is needed. The
+ * RX engine must be quiesced before the rings are freed, regardless
+ * of TX-side state - skipping it would leave the device DMAing peer
+ * messages and descriptor writebacks into released coherent memory.
+ */
+ nbl_chan_stop_rx_queue(chan_mgt);
+
+ /*
+ * Drain strategy mirrors mlx5 command interface teardown:
+ * set shutdown flag first, abort all pending waiters, then
+ * block until inflight_tx_cnt reaches zero.
+ *
+ * After shutdn is set every sender exits promptly at its next
+ * checkpoint:
+ * - interrupt-driven senders wake on shutdn immediately
+ * (it is part of the wait_event condition);
+ * - ACK polling senders check shutdn each iteration before
+ * sleeping 1000-1200us. This per-iteration check avoids
+ * waiting the full 5.0-6.0s worst-case ACK timeout.
+ *
+ * The wait is deliberately unbounded: a surviving sender still
+ * holds txq_lock or touches chan_info (state_lock, wait[]) on its
+ * exit path, so returning early would let devres free the rings
+ * and chan_info underneath it. In practice the TX polling loop is
+ * bounded (NBL_CHAN_TX_WAIT_TIMES) and wedged-link MMIO reads
+ * eventually complete with all-1s data, so this loop terminates;
+ * warn periodically to keep a genuinely stuck sender diagnosable.
+ */
+ while (wait_event_timeout(chan_info->inflight_wait,
+ atomic_read(&chan_info->inflight_tx_cnt) == 0,
+ msecs_to_jiffies(5000)) == 0)
+ dev_warn(chan_mgt->common->dev,
+ "teardown: still waiting for %d inflight sender(s)\n",
+ atomic_read(&chan_info->inflight_tx_cnt));
+
+ /*
+ * Join the last exiting sender through state_lock: it drops
+ * inflight_tx_cnt and wakes this drain while still holding
+ * state_lock, so observing the counter reach zero does not prove
+ * the sender has finished its mutex_unlock(). Acquiring the lock
+ * once here guarantees no sender still holds it - mutex_unlock()
+ * is the sender's last touch of chan_info - before returning lets
+ * devres free chan_info. shutdn blocks new senders, so one
+ * acquire/release pair is sufficient.
+ */
+ mutex_lock(&chan_info->state_lock);
+ mutex_unlock(&chan_info->state_lock);
+
+ /*
+ * All senders have exited, so txq_lock is uncontended: plain lock.
+ * Take it so this is the sole writer of the TX QINFO block
+ * (sender-side stop/config/doorbell sequences all run under
+ * txq_lock).
+ */
+ mutex_lock(&chan_info->txq_lock);
+ nbl_chan_stop_tx_queue(chan_mgt);
+ mutex_unlock(&chan_info->txq_lock);
+
+ /*
+ * Join the RX clean work before active=false and before devres
+ * releases the rings. IRQ is already freed and clean_task is
+ * cleared under state_lock, so no new instance can be queued.
+ */
+ if (task)
+ cancel_work_sync(task);
+
+ WRITE_ONCE(chan_info->active, false);
+ return 0;
+}
+
+static int nbl_chan_setup_queue(struct nbl_channel_mgt *chan_mgt, u8 chan_type)
+{
+ struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+ struct nbl_common_info *common = chan_mgt->common;
+ struct nbl_chan_ring *rxq = &chan_info->rxq;
+ int err;
+
+ if (READ_ONCE(chan_info->active)) {
+ dev_warn(common->dev, "channel already active, reject duplicate setup\n");
+ return -EBUSY;
+ }
+ /*
+ * The coherent descriptor rings and per-ring contiguous buffer
+ * blocks are allocated once and released only by devres at device
+ * detach; teardown never frees them. Re-setup would orphan them, so
+ * it is rejected.
+ */
+ if (chan_info->dma_allocated) {
+ dev_warn(common->dev,
+ "channel DMA already allocated, re-setup not supported\n");
+ return -EBUSY;
+ }
+
+ nbl_chan_init_queue_param(chan_info, NBL_CHAN_QUEUE_LEN,
+ NBL_CHAN_QUEUE_LEN, NBL_CHAN_BUF_LEN,
+ NBL_CHAN_BUF_LEN);
+ /*
+ * Set the flag before the first allocation: if setup fails halfway,
+ * the buffers allocated so far already live until detach, so a retry
+ * must be rejected instead of allocating a second set and orphaning
+ * the first one.
+ */
+ chan_info->dma_allocated = true;
+ err = nbl_chan_init_queue(common, chan_info);
+ if (err) {
+ chan_info->dma_allocated = false;
+ return err;
+ }
+ nbl_chan_config_queue(chan_mgt, chan_info, true); /* tx */
+ nbl_chan_config_queue(chan_mgt, chan_info, false); /* rx */
+ nbl_chan_update_tail_ptr(hw_ops, chan_mgt->hw_ops_tbl->priv,
+ rxq->tail_ptr, NBL_MB_RX_QID);
+ WRITE_ONCE(chan_info->active, true);
+ return 0;
+}
+
+static bool nbl_chan_txq_full(struct nbl_chan_ring *txq,
+ u16 num_entries)
+{
+ return NBL_NEXT_ID(txq->next_to_use, num_entries - 1) ==
+ txq->next_to_clean;
+}
+
+static int nbl_chan_update_txqueue(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_info *chan_info,
+ struct nbl_chan_tx_param *param)
+{
+ struct nbl_chan_ring *txq = &chan_info->txq;
+ struct nbl_chan_tx_desc *tx_desc;
+ struct nbl_chan_buf *tx_buf;
+
+ if (nbl_chan_txq_full(txq, chan_info->num_txq_entries))
+ return -EBUSY;
+ if (param->arg_len > NBL_CHAN_BUF_LEN - sizeof(*tx_desc))
+ return -EINVAL;
+ tx_desc =
+ NBL_CHAN_TX_RING_TO_DESC(txq, txq->next_to_use);
+ tx_buf =
+ NBL_CHAN_TX_RING_TO_BUF(txq, txq->next_to_use);
+ tx_desc->dstid = cpu_to_le16(param->dstid);
+ tx_desc->msg_type = cpu_to_le16(param->msg_type);
+ tx_desc->msgid = cpu_to_le16(param->msgid);
+
+ /*
+ * srcid field is filled by mailbox hardware after peer receives this
+ * packet, driver producer never writes srcid; reused descriptor slots
+ * will contain stale srcid value temporarily until hardware overwrites
+ * it.
+ */
+ if (param->arg_len > NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN) {
+ if (param->arg)
+ memcpy(tx_buf->va, param->arg, param->arg_len);
+ tx_desc->buf_addr = cpu_to_le64(tx_buf->pa);
+ tx_desc->buf_len = cpu_to_le16(param->arg_len);
+ tx_desc->data_len = 0;
+ memset(tx_desc->data, 0, sizeof(tx_desc->data));
+ } else {
+ memset(tx_desc->data, 0, sizeof(tx_desc->data));
+ memset(&tx_desc->buf_addr, 0, sizeof(tx_desc->buf_addr));
+ if (param->arg && param->arg_len > 0)
+ memcpy(tx_desc->data, param->arg, param->arg_len);
+ tx_desc->buf_len = 0;
+ tx_desc->data_len = cpu_to_le16(param->arg_len);
+ }
+ /* Ensure descriptor data visible to device before AVAIL flag */
+ dma_wmb();
+ tx_desc->flags = cpu_to_le16(BIT(NBL_CHAN_TX_DESC_AVAIL));
+
+ txq->next_to_use =
+ NBL_NEXT_ID(txq->next_to_use, chan_info->num_txq_entries - 1);
+ txq->tail_ptr++;
+
+ return 0;
+}
+
+/*
+ * Quiesce the TX mailbox queue and reclaim all outstanding
+ * descriptors. Called from the timeout path of nbl_chan_kick_tx_ring()
+ * with txq_lock held.
+ *
+ * The device failed to fetch/complete the current descriptor within the
+ * polling window. We assert QUEUE_RST to stop further DMA fetches,
+ * reclaim every descriptor between next_to_clean and next_to_use,
+ * reset the software tail_ptr counter to match the hardware reset state,
+ * and re-enable the queue so subsequent sends can proceed.
+ */
+static void nbl_chan_quiesce_and_reclaim_tx(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_info *chan_info)
+{
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+ struct nbl_hw_mgt *hw_priv = chan_mgt->hw_ops_tbl->priv;
+ struct nbl_chan_ring *txq = &chan_info->txq;
+ struct nbl_chan_tx_desc *tx_desc;
+
+ /*
+ * Assert QUEUE_RST to stop hardware fetching new descriptors.
+ * stop_mailbox_txq() flushes through the mailbox BAR itself, so
+ * the reset has reached the device when it returns.
+ */
+ hw_ops->stop_mailbox_txq(hw_priv);
+
+ /*
+ * Reclaim all outstanding descriptors between next_to_clean and
+ * next_to_use. Under txq_lock there is at most one in-flight
+ * descriptor, but iterate the full range for robustness.
+ */
+ while (txq->next_to_clean != txq->next_to_use) {
+ tx_desc = NBL_CHAN_TX_RING_TO_DESC(txq,
+ txq->next_to_clean);
+ WRITE_ONCE(tx_desc->flags, 0);
+ txq->next_to_clean =
+ NBL_NEXT_ID(txq->next_to_clean,
+ chan_info->num_txq_entries - 1);
+ }
+
+ /*
+ * Hardware tail_ptr counter is cleared by QUEUE_RST. Reset
+ * software counter to match so the next doorbell update does
+ * not produce a false 16-bit wrap delta.
+ */
+ txq->tail_ptr = 0;
+ txq->next_to_use = 0;
+ txq->next_to_clean = 0;
+
+ /*
+ * Authoritative re-arm gate, evaluated while txq_lock is held:
+ * teardown sets shutdn before draining senders and then owns
+ * hardware state. If teardown started while we were stopping/
+ * reclaiming, leave the queue in QUEUE_RST and never program
+ * QUEUE_EN, otherwise the queue could be armed again pointing
+ * at rings that are about to be freed.
+ */
+ if (READ_ONCE(chan_info->shutdn))
+ return;
+
+ /* Re-enable queue with current ring base and size */
+ nbl_chan_config_queue(chan_mgt, chan_info, true);
+}
+
+static int nbl_chan_kick_tx_ring(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_info *chan_info)
+{
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+ struct nbl_chan_ring *txq = &chan_info->txq;
+ struct device *dev = chan_mgt->common->dev;
+ int max_retries = NBL_CHAN_TX_WAIT_TIMES;
+ struct nbl_chan_tx_desc *tx_desc;
+ int retry_count = 0;
+ u16 msg_type;
+
+ nbl_chan_update_tail_ptr(hw_ops, chan_mgt->hw_ops_tbl->priv,
+ txq->tail_ptr, NBL_MB_TX_QID);
+
+ tx_desc = NBL_CHAN_TX_RING_TO_DESC(txq, txq->next_to_clean);
+ /*
+ * Poll for HW to mark descriptor as USED.
+ * Mailbox is a low-speed control channel for management commands.
+ * We avoid enabling dedicated per-TX interrupt for single control
+ * message to reduce interrupt overhead, so use bounded polling
+ * with small delay instead.
+ */
+ while (retry_count < max_retries) {
+ if (READ_ONCE(chan_info->shutdn))
+ return -ESHUTDOWN;
+
+ if (le16_to_cpu(READ_ONCE(tx_desc->flags)) &
+ BIT(NBL_CHAN_TX_DESC_USED)) {
+ /*
+ * Order reads of other device-written descriptor
+ * fields after observing USED. Matches the RX side
+ * pattern in nbl_chan_clean_queue().
+ */
+ dma_rmb();
+ break;
+ }
+
+ retry_count++;
+ if (retry_count == max_retries) {
+ msg_type = le16_to_cpu(READ_ONCE(tx_desc->msg_type));
+ dev_err_ratelimited(dev, "chan send msg type: %d timeout\n",
+ msg_type);
+ /*
+ * Teardown may have started after the loop-top
+ * check on this final iteration. Do not run lock
+ * recovery then: it would stop/reclaim the queue
+ * and, without the in-quiesce gate, re-arm it.
+ * nbl_chan_teardown_queue() owns hardware state.
+ */
+ if (READ_ONCE(chan_info->shutdn))
+ return -ESHUTDOWN;
+ /*
+ * Device failed to complete this descriptor.
+ * Quiesce the queue, reclaim the timed-out
+ * descriptor, and re-enable so future sends can
+ * proceed instead of stalling the ring full.
+ */
+ nbl_chan_quiesce_and_reclaim_tx(chan_mgt,
+ chan_info);
+ return -ETIMEDOUT;
+ }
+ usleep_range(NBL_CHAN_TX_WAIT_US, NBL_CHAN_TX_WAIT_US_MAX);
+ }
+
+ txq->next_to_clean = txq->next_to_use;
+
+ return 0;
+}
+
+static void nbl_chan_recv_ack_msg(void *priv, u16 srcid, u16 msgid, void *data,
+ u32 data_len)
+{
+ struct nbl_channel_mgt *chan_mgt = (struct nbl_channel_mgt *)priv;
+ struct nbl_chan_waitqueue_head *wait_head = NULL;
+ struct device *dev = chan_mgt->common->dev;
+ struct nbl_chan_info *chan_info =
+ chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX];
+ u16 w_dstid, w_msgtype, w_msgidx;
+ u32 *payload = data;
+ u16 ack_msgtype = 0;
+ u16 ack_msgid = 0;
+ u32 ack_datalen;
+ void *ack_data;
+ u32 copy_len;
+ int w_status;
+ s32 raw_err;
+
+ if (READ_ONCE(chan_info->shutdn))
+ return;
+ if (data_len > NBL_CHAN_BUF_LEN ||
+ data_len < NBL_CHAN_ACK_HEAD_LEN * sizeof(u32)) {
+ dev_err_ratelimited(dev, "Invalid ACK data_len: %u\n",
+ data_len);
+ return;
+ }
+ ack_datalen = data_len - NBL_CHAN_ACK_HEAD_LEN * sizeof(u32);
+ ack_msgtype = le16_to_cpu(*(__le16 *)(payload + NBL_CHAN_MSG_TYPE_POS));
+ ack_msgid = le16_to_cpu(*(__le16 *)(payload + NBL_CHAN_MSG_ID_POS));
+ if (FIELD_GET(NBL_CHAN_MSGID_LOC_MASK, ack_msgid) >=
+ chan_info->num_txq_entries) {
+ dev_err_ratelimited(dev, "chan recv msg id: %u err\n",
+ ack_msgid);
+ return;
+ }
+ wait_head =
+ &chan_info->wait[FIELD_GET(NBL_CHAN_MSGID_LOC_MASK, ack_msgid)];
+
+ mutex_lock(&chan_info->pending_lock);
+
+ /* Cache repeated READ_ONCE values */
+ w_dstid = READ_ONCE(wait_head->dstid);
+ w_status = READ_ONCE(wait_head->status);
+ w_msgtype = READ_ONCE(wait_head->msg_type);
+ w_msgidx = READ_ONCE(wait_head->msg_index);
+
+ if (srcid != w_dstid) {
+ mutex_unlock(&chan_info->pending_lock);
+ dev_err_ratelimited(dev, "ACK srcid=%u != dstid=%u, rejecting\n",
+ srcid, w_dstid);
+ return;
+ }
+ if (w_status != NBL_MBX_STATUS_WAITING) {
+ mutex_unlock(&chan_info->pending_lock);
+ dev_err_ratelimited(dev,
+ "Skip ack invalid status, wait msgtype:%u idx:%u status:%d ack msgtype:%u msgid:%u datalen:%u\n",
+ w_msgtype, w_msgidx, w_status,
+ ack_msgtype, ack_msgid, ack_datalen);
+ return;
+ }
+
+ if (w_msgtype != ack_msgtype) {
+ mutex_unlock(&chan_info->pending_lock);
+ dev_err_ratelimited(dev,
+ "Skip ack msgtype mismatch, wait msgtype:%u idx:%u ack msgtype:%u msgid:%u\n",
+ w_msgtype, w_msgidx, ack_msgtype,
+ ack_msgid);
+ return;
+ }
+ if (FIELD_GET(NBL_CHAN_MSGID_INDEX_MASK, ack_msgid) != w_msgidx) {
+ mutex_unlock(&chan_info->pending_lock);
+ dev_err_ratelimited(dev,
+ "Stale ACK: expected index=%u, got msgid=%u\n",
+ w_msgidx, ack_msgid);
+ return;
+ }
+
+ raw_err = (s32)le32_to_cpu(*(__le32 *)&payload[NBL_CHAN_ACK_RET_POS]);
+ if (raw_err > 0 || raw_err < -MAX_ERRNO)
+ raw_err = -EREMOTEIO;
+
+ WRITE_ONCE(wait_head->ack_err, raw_err);
+
+ copy_len = min_t(u32, READ_ONCE(wait_head->ack_data_len), ack_datalen);
+ if (READ_ONCE(wait_head->ack_err) >= 0 && copy_len > 0) {
+ ack_data = READ_ONCE(wait_head->ack_data);
+ if (!ack_data) {
+ dev_err_ratelimited(dev, "ACK payload dropped: ack_data is NULL\n");
+ WRITE_ONCE(wait_head->ack_data_len, 0);
+ goto ack_done;
+ }
+ memcpy((char *)ack_data,
+ payload + NBL_CHAN_ACK_HEAD_LEN, copy_len);
+ WRITE_ONCE(wait_head->ack_data_len, (u16)copy_len);
+ } else {
+ WRITE_ONCE(wait_head->ack_data_len, 0);
+ }
+ack_done:
+ /* Guarantee payload data finished before acked flag visible */
+ smp_wmb();
+ WRITE_ONCE(wait_head->acked, 1);
+ WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_ACKD);
+ mutex_unlock(&chan_info->pending_lock);
+ wake_up(&wait_head->wait_queue);
+}
+
+static void nbl_chan_recv_msg(struct nbl_channel_mgt *chan_mgt, void *data)
+{
+ struct device *dev = chan_mgt->common->dev;
+ struct nbl_chan_msg_node_data *msg_handler;
+ u16 msg_type, payload_len, srcid, msgid;
+ struct nbl_chan_info *chan_info =
+ chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX];
+ struct nbl_chan_tx_desc *tx_desc;
+ void *payload;
+ size_t avail_space;
+ u16 data_len_fw;
+
+ if (READ_ONCE(chan_info->shutdn))
+ return;
+
+ tx_desc = data;
+ msg_type = le16_to_cpu(READ_ONCE(tx_desc->msg_type));
+ dev_dbg(dev, "recv msg_type: %d\n", msg_type);
+
+ srcid = le16_to_cpu(READ_ONCE(tx_desc->srcid));
+ msgid = le16_to_cpu(READ_ONCE(tx_desc->msgid));
+
+ if (msg_type >= NBL_CHAN_MSG_MAILBOX_MAX)
+ return;
+
+ data_len_fw = le16_to_cpu(READ_ONCE(tx_desc->data_len));
+ if (data_len_fw) {
+ payload_len = data_len_fw;
+
+ if (payload_len > NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN) {
+ dev_err_ratelimited(dev,
+ "data_len=%u exceeds embedded buffer size=%u\n",
+ payload_len,
+ NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN);
+ return;
+ }
+ /* Small pkt: payload stored inside descriptor data[] array */
+ payload = tx_desc->data;
+ } else {
+ payload_len = le16_to_cpu(READ_ONCE(tx_desc->buf_len));
+
+ avail_space = NBL_CHAN_BUF_LEN - sizeof(*tx_desc);
+ if (payload_len > avail_space) {
+ dev_err_ratelimited(dev,
+ "buf_len=%u exceeds external buffer size=%zu\n",
+ payload_len, avail_space);
+ return;
+ }
+ /* Large pkt: payload follows immediately after tx_desc */
+ payload = tx_desc + 1;
+ }
+
+ msg_handler = xa_load(&chan_mgt->handler_xa, msg_type);
+ if (!msg_handler || !msg_handler->func) {
+ dev_err_ratelimited(dev,
+ "No handler for msg_type: %u (srcid=%u, msgid=%u)\n",
+ msg_type, srcid, msgid);
+ return;
+ }
+
+ msg_handler->func(msg_handler->priv, srcid, msgid, payload,
+ payload_len);
+}
+
+static void nbl_chan_advance_rx_ring(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_info *chan_info,
+ struct nbl_chan_ring *rxq)
+{
+ struct nbl_hw_ops *hw_ops = chan_mgt->hw_ops_tbl->ops;
+ struct nbl_chan_rx_desc *rx_desc;
+ struct nbl_chan_buf *rx_buf;
+ u16 next_to_use;
+
+ next_to_use = rxq->next_to_use;
+ rx_desc = NBL_CHAN_RX_RING_TO_DESC(rxq, next_to_use);
+ rx_buf = NBL_CHAN_RX_RING_TO_BUF(rxq, next_to_use);
+
+ /*
+ * Recycle the RX descriptor at next_to_use. The initial
+ * unused slot is intentionally recycled after the first
+ * RX descriptor is consumed, allowing the ring to become
+ * fully populated while next_to_clean tracks the consumer.
+ */
+ rx_desc->buf_addr = cpu_to_le64(rx_buf->pa);
+ rx_desc->buf_len = cpu_to_le32(chan_info->rxq_buf_size);
+
+ /*
+ * DMA Write Memory Barrier:
+ * Ensures all previous DMA-mapped writes (buffer address/length)
+ * are completed before the descriptor flags are updated.
+ * This prevents hardware from seeing a partially updated descriptor
+ * where flags are set but buffer info isn't ready yet.
+ */
+ dma_wmb();
+
+ rx_desc->flags = cpu_to_le16(BIT(NBL_CHAN_RX_DESC_AVAIL));
+
+ rxq->next_to_use++;
+ if (rxq->next_to_use == chan_info->num_rxq_entries)
+ rxq->next_to_use = 0;
+ rxq->tail_ptr++;
+
+ nbl_chan_update_tail_ptr(hw_ops, chan_mgt->hw_ops_tbl->priv,
+ rxq->tail_ptr, NBL_MB_RX_QID);
+}
+
+static void nbl_chan_clean_queue(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_info *chan_info)
+{
+ struct nbl_common_info *common = chan_mgt->common;
+ struct nbl_chan_ring *rxq = &chan_info->rxq;
+ struct device *dev = chan_mgt->common->dev;
+ u32 budget = NBL_CHAN_RX_CLEAN_BUDGET;
+ struct nbl_chan_rx_desc *rx_desc;
+ struct nbl_chan_buf *rx_buf;
+ struct work_struct *task;
+ bool more_work = false;
+ u16 next_to_clean;
+ u16 flags;
+
+ /* Nothing to clean before setup: the RX ring is still NULL */
+ if (!READ_ONCE(chan_info->active))
+ return;
+
+ next_to_clean = rxq->next_to_clean;
+ rx_desc = NBL_CHAN_RX_RING_TO_DESC(rxq, next_to_clean);
+ rx_buf = NBL_CHAN_RX_RING_TO_BUF(rxq, next_to_clean);
+ while (le16_to_cpu(READ_ONCE(rx_desc->flags)) &
+ BIT(NBL_CHAN_RX_DESC_USED)) {
+ flags = le16_to_cpu(READ_ONCE(rx_desc->flags));
+
+ if (READ_ONCE(chan_info->shutdn))
+ break;
+ if (!(flags & BIT(NBL_CHAN_RX_DESC_WRITE)))
+ dev_dbg(dev,
+ "mailbox rx flag 0x%x missing NBL_CHAN_RX_DESC_WRITE\n",
+ flags);
+
+ /* Make sure hardware written descriptor visible to CPU */
+ dma_rmb();
+ nbl_chan_recv_msg(chan_mgt, rx_buf->va);
+ nbl_chan_advance_rx_ring(chan_mgt, chan_info, rxq);
+ next_to_clean++;
+ if (next_to_clean == chan_info->num_rxq_entries)
+ next_to_clean = 0;
+ rx_desc = NBL_CHAN_RX_RING_TO_DESC(rxq, next_to_clean);
+ rx_buf = NBL_CHAN_RX_RING_TO_BUF(rxq, next_to_clean);
+ if (--budget == 0) {
+ more_work = true;
+ break;
+ }
+ cond_resched();
+ }
+ rxq->next_to_clean = next_to_clean;
+
+ mutex_lock(&chan_info->state_lock);
+ /* Prevent queue_work after teardown clears clean_task */
+ if (READ_ONCE(chan_info->shutdn)) {
+ mutex_unlock(&chan_info->state_lock);
+ return;
+ }
+ if (common->wq && more_work) {
+ task = READ_ONCE(chan_info->clean_task);
+ if (task)
+ queue_work(common->wq, task);
+ }
+ mutex_unlock(&chan_info->state_lock);
+}
+
+static void nbl_chan_clean_queue_subtask(struct nbl_channel_mgt *chan_mgt,
+ u8 chan_type)
+{
+ struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+
+ nbl_chan_clean_queue(chan_mgt, chan_info);
+}
+
+static int nbl_chan_get_msg_id(struct nbl_chan_info *chan_info,
+ u16 *msgid)
+{
+ int search_loc = READ_ONCE(chan_info->wait_head_index), i;
+ struct nbl_chan_waitqueue_head *wait = NULL;
+ int status;
+ int next;
+
+ lockdep_assert_held(&chan_info->pending_lock);
+ for (i = 0; i < chan_info->num_txq_entries; i++) {
+ wait = &chan_info->wait[search_loc];
+ status = READ_ONCE(wait->status);
+ if (status == NBL_MBX_STATUS_IDLE ||
+ status == NBL_MBX_STATUS_TIMEOUT) {
+ WRITE_ONCE(wait->msg_index,
+ NBL_NEXT_ID(wait->msg_index,
+ NBL_CHAN_MSG_INDEX_MAX));
+
+ *msgid = FIELD_PREP(NBL_CHAN_MSGID_INDEX_MASK,
+ wait->msg_index) |
+ FIELD_PREP(NBL_CHAN_MSGID_LOC_MASK,
+ search_loc);
+
+ /* Advance starting search position for next caller */
+ next = NBL_NEXT_ID(search_loc,
+ chan_info->num_txq_entries - 1);
+ WRITE_ONCE(chan_info->wait_head_index, next);
+ return 0;
+ }
+
+ search_loc = NBL_NEXT_ID(search_loc,
+ chan_info->num_txq_entries - 1);
+ }
+
+ /*
+ * All tx slots are occupied. May happen under high transmit load
+ * or delayed remote ACK responses. Caller should retry later.
+ */
+ return -EAGAIN;
+}
+
+static int nbl_chan_send_msg(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_send_info *chan_send)
+{
+ struct nbl_common_info *common = chan_mgt->common;
+ struct nbl_chan_waitqueue_head *wait_head = NULL;
+ struct nbl_chan_tx_param tx_param = { 0 };
+ int i = NBL_CHAN_TX_WAIT_ACK_TIMES;
+ struct nbl_chan_info *chan_info =
+ chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX];
+ struct device *dev = common->dev;
+ struct work_struct *task;
+ u16 msgid = 0;
+ int ret;
+
+ if (chan_send->resp_len > NBL_CHAN_BUF_LEN) {
+ dev_err_ratelimited(dev, "resp_len %zu exceeds max %d\n",
+ chan_send->resp_len, NBL_CHAN_BUF_LEN);
+ return -EINVAL;
+ }
+
+ /*
+ * Entry points are only reached after a successful setup_queue()
+ * (probe path), but the channel is a dynamic object: fail cleanly
+ * instead of dereferencing rings/wait[] that were never allocated
+ * (num_txq_entries == 0). active is published last in setup and
+ * cleared last in teardown, so a false value here means the channel
+ * was never set up; the teardown window itself is covered by the
+ * shutdn check below.
+ */
+ if (!READ_ONCE(chan_info->active))
+ return -ENODEV;
+
+ mutex_lock(&chan_info->state_lock);
+ if (READ_ONCE(chan_info->shutdn)) {
+ mutex_unlock(&chan_info->state_lock);
+ return -ESHUTDOWN;
+ }
+ atomic_inc(&chan_info->inflight_tx_cnt);
+ mutex_unlock(&chan_info->state_lock);
+
+ tx_param.msg_type = chan_send->msg_type;
+ tx_param.arg = chan_send->arg;
+ tx_param.arg_len = chan_send->arg_len;
+ tx_param.dstid = chan_send->dstid;
+ tx_param.msgid = msgid;
+ if (chan_send->ack) {
+ mutex_lock(&chan_info->pending_lock);
+
+ ret = nbl_chan_get_msg_id(chan_info, &msgid);
+ if (ret) {
+ mutex_unlock(&chan_info->pending_lock);
+ dev_err_ratelimited(dev,
+ "Channel tx wait head full, send msgtype:%u to dstid:%u failed\n",
+ chan_send->msg_type,
+ chan_send->dstid);
+ goto out_clean_inflight;
+ }
+ wait_head =
+ &chan_info->wait[FIELD_GET(NBL_CHAN_MSGID_LOC_MASK,
+ msgid)];
+ WRITE_ONCE(wait_head->acked, 0);
+ WRITE_ONCE(wait_head->ack_data, chan_send->resp);
+ WRITE_ONCE(wait_head->ack_data_len, chan_send->resp_len);
+ WRITE_ONCE(wait_head->msg_type, chan_send->msg_type);
+ WRITE_ONCE(wait_head->msg_index,
+ FIELD_GET(NBL_CHAN_MSGID_INDEX_MASK, msgid));
+ WRITE_ONCE(wait_head->dstid, chan_send->dstid);
+
+ WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_WAITING);
+ mutex_unlock(&chan_info->pending_lock);
+
+ tx_param.msgid = msgid;
+ }
+
+ mutex_lock(&chan_info->txq_lock);
+ ret = nbl_chan_update_txqueue(chan_mgt, chan_info, &tx_param);
+ if (ret) {
+ mutex_unlock(&chan_info->txq_lock);
+ dev_err_ratelimited(dev,
+ "Channel tx queue full, send msgtype:%u to dstid:%u failed\n",
+ chan_send->msg_type, chan_send->dstid);
+ if (wait_head)
+ goto out_clear_wait_slot;
+ goto out_clean_inflight;
+ }
+
+ ret = nbl_chan_kick_tx_ring(chan_mgt, chan_info);
+ mutex_unlock(&chan_info->txq_lock);
+ if (ret) {
+ if (wait_head)
+ goto out_clear_wait_slot;
+ goto out_clean_inflight;
+ }
+
+ if (!chan_send->ack) {
+ ret = 0;
+ goto out_clean_inflight;
+ }
+
+ if (test_bit(NBL_CHAN_IRQ_RDY, chan_info->state)) {
+ while (!READ_ONCE(wait_head->acked)) {
+ /*
+ * avoids long task blocking when interrupt mode is
+ * disabled mid-wait. Cannot guarantee subsequent ACK
+ * delivery after interrupt mask off, only prevents
+ * infinite blocking. Spurious timeout is possible.
+ */
+ ret = wait_event_timeout(wait_head->wait_queue,
+ READ_ONCE(wait_head->acked) ||
+ READ_ONCE(chan_info->shutdn) ||
+ !test_bit(NBL_CHAN_IRQ_RDY,
+ chan_info->state),
+ NBL_CHAN_ACK_WAIT_TIME);
+
+ if (READ_ONCE(chan_info->shutdn)) {
+ ret = -ESHUTDOWN;
+ goto out_clear_wait_slot;
+ }
+ if (!test_bit(NBL_CHAN_IRQ_RDY, chan_info->state)) {
+ ret = -EIO;
+ goto out_clear_wait_slot;
+ }
+ if (ret == 0) {
+ mutex_lock(&chan_info->pending_lock);
+ if (READ_ONCE(wait_head->status) ==
+ NBL_MBX_STATUS_WAITING) {
+ WRITE_ONCE(wait_head->status,
+ NBL_MBX_STATUS_TIMEOUT);
+ WRITE_ONCE(wait_head->acked, 0);
+ WRITE_ONCE(wait_head->ack_data, NULL);
+ WRITE_ONCE(wait_head->ack_data_len, 0);
+ /*
+ * Ensure all status/ack slot
+ * updates are visible before subsequent
+ * readers observe acked == 0
+ */
+ smp_wmb();
+ mutex_unlock(&chan_info->pending_lock);
+ dev_err_ratelimited(dev,
+ "Channel waiting ack failed, message type: %d, msg id: %u\n",
+ chan_send->msg_type,
+ msgid);
+ ret = -ETIMEDOUT;
+ /*
+ * TIMEOUT slots can be reused by
+ * another sender. A late ACK is
+ * rejected by the receive path,
+ * which only accepts WAITING slots
+ * (and by the msg_index generation
+ * check after reallocation). The
+ * current sender no longer owns the
+ * slot, so skip the IDLE reset.
+ */
+ goto out_clean_inflight;
+ }
+ /*
+ * ACK won the timeout race: between the
+ * final condition check and this lock the
+ * receiver published acked=1/ACKD and
+ * copied the payload into our response
+ * buffer. The slot is still exclusively
+ * ours (only IDLE/TIMEOUT slots are ever
+ * reallocated), so consume the completion:
+ * unlock and continue to the loop-bottom
+ * acked check, which enters the normal
+ * success readout and IDLE reset. Leaving
+ * here via the timeout path would strand
+ * the slot in ACKD forever, permanently
+ * shrinking the 256-entry pool, and would
+ * report a false -ETIMEDOUT for a request
+ * whose response already arrived.
+ */
+ mutex_unlock(&chan_info->pending_lock);
+ }
+
+ if (READ_ONCE(wait_head->acked))
+ break;
+ }
+ if (READ_ONCE(wait_head->acked)) {
+ /*
+ * Load ordering: observe acked flag before
+ * reading ACK payload metadata.
+ */
+ smp_rmb();
+ chan_send->ack_len = READ_ONCE(wait_head->ack_data_len);
+ ret = READ_ONCE(wait_head->ack_err);
+ }
+ } else {
+ /* Polling path for synchronous ACK */
+ while (i--) {
+ if (READ_ONCE(chan_info->shutdn)) {
+ ret = -ESHUTDOWN;
+ goto out_clear_wait_slot;
+ }
+
+ mutex_lock(&chan_info->state_lock);
+ task = READ_ONCE(chan_info->clean_task);
+ if (common->wq && task &&
+ !READ_ONCE(chan_info->shutdn) &&
+ !work_pending(task))
+ queue_work(common->wq, task);
+ mutex_unlock(&chan_info->state_lock);
+ if (READ_ONCE(wait_head->acked)) {
+ /*
+ * Guarantee load order: observe acked
+ * flag before reading ack payload metadata.
+ */
+ smp_rmb();
+ chan_send->ack_len =
+ READ_ONCE(wait_head->ack_data_len);
+ ret = READ_ONCE(wait_head->ack_err);
+ goto out_clear_wait_slot;
+ }
+
+ usleep_range(NBL_CHAN_TX_WAIT_ACK_US_MIN,
+ NBL_CHAN_TX_WAIT_ACK_US_MAX);
+ cond_resched();
+ }
+ mutex_lock(&chan_info->pending_lock);
+ if (READ_ONCE(wait_head->status) == NBL_MBX_STATUS_ACKD) {
+ chan_send->ack_len = READ_ONCE(wait_head->ack_data_len);
+ ret = READ_ONCE(wait_head->ack_err);
+ mutex_unlock(&chan_info->pending_lock);
+ goto out_clear_wait_slot;
+ }
+ /*
+ * Polling timed out without receiving an ACK. Transition
+ * the slot to TIMEOUT so nbl_chan_get_msg_id() can reuse
+ * it. Must not proceed to out_clean_inflight with
+ * the slot still in WAITING state — that would leak the
+ * slot permanently and eventually exhaust the 256-entry
+ * channel queue.
+ */
+ if (READ_ONCE(wait_head->status) == NBL_MBX_STATUS_WAITING) {
+ WRITE_ONCE(wait_head->status, NBL_MBX_STATUS_TIMEOUT);
+ WRITE_ONCE(wait_head->acked, 0);
+ WRITE_ONCE(wait_head->ack_data, NULL);
+ WRITE_ONCE(wait_head->ack_data_len, 0);
+ }
+ mutex_unlock(&chan_info->pending_lock);
+ dev_err_ratelimited(dev,
+ "Channel polling ack failed, message type: %d msg id: %u\n",
+ chan_send->msg_type, msgid);
+ ret = -ETIMEDOUT;
+ goto out_clean_inflight;
+ }
+
+out_clear_wait_slot:
+ mutex_lock(&chan_info->pending_lock);
+ nbl_chan_reset_wait_head(chan_info, wait_head);
+ mutex_unlock(&chan_info->pending_lock);
+
+out_clean_inflight:
+ mutex_lock(&chan_info->state_lock);
+ if (atomic_dec_and_test(&chan_info->inflight_tx_cnt))
+ wake_up(&chan_info->inflight_wait);
+ mutex_unlock(&chan_info->state_lock);
+ return ret;
+}
+
+static int nbl_chan_send_ack(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_ack_info *chan_ack)
+{
+ size_t head_len = NBL_CHAN_ACK_HEAD_LEN * sizeof(u32);
+ size_t data_len = chan_ack->data_len;
+ struct nbl_chan_send_info chan_send;
+ __le32 *tmp;
+ size_t len;
+ int ret;
+
+ if (data_len >
+ NBL_CHAN_BUF_LEN - sizeof(struct nbl_chan_tx_desc) - head_len)
+ return -EINVAL;
+
+ len = head_len + data_len;
+ tmp = kzalloc(len, GFP_KERNEL);
+ if (!tmp)
+ return -ENOMEM;
+
+ *(__le16 *)&tmp[NBL_CHAN_MSG_TYPE_POS] =
+ cpu_to_le16(chan_ack->msg_type);
+ *(__le16 *)&tmp[NBL_CHAN_MSG_ID_POS] = cpu_to_le16(chan_ack->msgid);
+ tmp[NBL_CHAN_ACK_RET_POS] = cpu_to_le32(chan_ack->err);
+ if (chan_ack->data && chan_ack->data_len)
+ memcpy(&tmp[NBL_CHAN_ACK_HEAD_LEN], chan_ack->data,
+ chan_ack->data_len);
+
+ nbl_chan_fill_send_info(&chan_send, chan_ack->dstid, NBL_CHAN_MSG_ACK,
+ tmp, len, NULL, 0, 0);
+ ret = nbl_chan_send_msg(chan_mgt, &chan_send);
+ kfree(tmp);
+
+ return ret;
+}
+
+static int nbl_chan_register_msg(struct nbl_channel_mgt *chan_mgt, u16 msg_type,
+ nbl_chan_resp func, void *callback)
+{
+ return nbl_chan_add_msg_handler(chan_mgt, msg_type, func, callback);
+}
+
+static bool nbl_chan_check_queue_exist(struct nbl_channel_mgt *chan_mgt,
+ u8 chan_type)
+{
+ struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+
+ return chan_info ? true : false;
+}
+
+static void nbl_chan_register_chan_task(struct nbl_channel_mgt *chan_mgt,
+ u8 chan_type, struct work_struct *task)
+{
+ struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+
+ mutex_lock(&chan_info->state_lock);
+ if (!READ_ONCE(chan_info->shutdn))
+ WRITE_ONCE(chan_info->clean_task, task);
+ mutex_unlock(&chan_info->state_lock);
+}
+
+static void nbl_chan_set_queue_state(struct nbl_channel_mgt *chan_mgt,
+ enum nbl_chan_state state, u8 chan_type,
+ u8 set)
+{
+ struct nbl_chan_info *chan_info = chan_mgt->chan_info[chan_type];
+ int i;
+
+ if (set)
+ set_bit(state, chan_info->state);
+ else
+ clear_bit(state, chan_info->state);
+ /*
+ * When clearing IRQ_RDY, wake all per-slot wait queues so
+ * sleeping senders observe the condition immediately and
+ * return -EIO instead of waiting out the 3s timeout and
+ * reporting -ETIMEDOUT.
+ *
+ * wait[] is only allocated by a successful nbl_chan_init_tx_queue(),
+ * while num_txq_entries is already set by nbl_chan_init_queue_param();
+ * skip the walk if that allocation never happened.
+ */
+ if (!set && state == NBL_CHAN_IRQ_RDY && chan_info->wait) {
+ for (i = 0; i < chan_info->num_txq_entries; i++)
+ wake_up_all(&chan_info->wait[i].wait_queue);
+ }
+}
+
+static struct nbl_channel_ops chan_ops = {
+ .send_msg = nbl_chan_send_msg,
+ .send_ack = nbl_chan_send_ack,
+ .register_msg = nbl_chan_register_msg,
+ .cfg_chan_qinfo_map_table = nbl_chan_cfg_qinfo_map_table,
+ .check_queue_exist = nbl_chan_check_queue_exist,
+ .setup_queue = nbl_chan_setup_queue,
+ .teardown_queue = nbl_chan_teardown_queue,
+ .clean_queue_subtask = nbl_chan_clean_queue_subtask,
+ .register_chan_task = nbl_chan_register_chan_task,
+ .set_queue_state = nbl_chan_set_queue_state,
+};
+
+static struct nbl_channel_mgt *
+nbl_chan_setup_chan_mgt(struct nbl_adapter *adapter)
+{
+ struct nbl_hw_ops_tbl *hw_ops_tbl = adapter->intf.hw_ops_tbl;
+ struct nbl_common_info *common = &adapter->common;
+ struct device *dev = &adapter->pdev->dev;
+ struct nbl_channel_mgt *chan_mgt;
+ struct nbl_chan_info *mailbox;
+ int ret;
+
+ chan_mgt = devm_kzalloc(dev, sizeof(*chan_mgt), GFP_KERNEL);
+ if (!chan_mgt)
+ return ERR_PTR(-ENOMEM);
+
+ chan_mgt->common = common;
+ chan_mgt->hw_ops_tbl = hw_ops_tbl;
+
+ mailbox = devm_kzalloc(dev, sizeof(*mailbox), GFP_KERNEL);
+ if (!mailbox)
+ return ERR_PTR(-ENOMEM);
+ mailbox->chan_type = NBL_CHAN_TYPE_MAILBOX;
+ chan_mgt->chan_info[NBL_CHAN_TYPE_MAILBOX] = mailbox;
+
+ ret = nbl_chan_init_msg_handler(chan_mgt);
+ if (ret)
+ return ERR_PTR(ret);
+ ret = devm_mutex_init(common->dev, &mailbox->txq_lock);
+ if (ret)
+ return ERR_PTR(ret);
+ ret = devm_mutex_init(common->dev, &mailbox->state_lock);
+ if (ret)
+ return ERR_PTR(ret);
+ ret = devm_mutex_init(common->dev, &mailbox->pending_lock);
+ if (ret)
+ return ERR_PTR(ret);
+ return chan_mgt;
+}
+
+static struct nbl_channel_ops_tbl *
+nbl_chan_setup_ops(struct device *dev, struct nbl_channel_mgt *chan_mgt)
+{
+ struct nbl_channel_ops_tbl *chan_ops_tbl;
+ int ret;
+
+ chan_ops_tbl = devm_kzalloc(dev, sizeof(*chan_ops_tbl), GFP_KERNEL);
+ if (!chan_ops_tbl)
+ return ERR_PTR(-ENOMEM);
+ if (!chan_ops.send_msg || !chan_ops.send_ack ||
+ !chan_ops.register_msg || !chan_ops.cfg_chan_qinfo_map_table ||
+ !chan_ops.check_queue_exist || !chan_ops.setup_queue ||
+ !chan_ops.teardown_queue || !chan_ops.clean_queue_subtask ||
+ !chan_ops.register_chan_task || !chan_ops.set_queue_state)
+ return ERR_PTR(-EINVAL);
+
+ chan_ops_tbl->ops = &chan_ops;
+ chan_ops_tbl->priv = chan_mgt;
+
+ ret = nbl_chan_register_msg(chan_mgt, NBL_CHAN_MSG_ACK,
+ nbl_chan_recv_ack_msg, chan_mgt);
+ if (ret)
+ return ERR_PTR(ret);
+
+ return chan_ops_tbl;
+}
+
+int nbl_chan_init_common(struct nbl_adapter *adap)
+{
+ struct nbl_channel_ops_tbl *chan_ops_tbl;
+ struct device *dev = &adap->pdev->dev;
+ struct nbl_channel_mgt *chan_mgt;
+ int ret;
+
+ chan_mgt = nbl_chan_setup_chan_mgt(adap);
+ if (IS_ERR(chan_mgt)) {
+ ret = PTR_ERR(chan_mgt);
+ goto exit;
+ }
+
+ chan_ops_tbl = nbl_chan_setup_ops(dev, chan_mgt);
+ if (IS_ERR(chan_ops_tbl)) {
+ ret = PTR_ERR(chan_ops_tbl);
+ goto cleanup_mgt;
+ }
+
+ adap->intf.channel_ops_tbl = chan_ops_tbl;
+ adap->core.chan_mgt = chan_mgt;
+ ret = nbl_common_create_wq(&adap->common);
+ if (ret)
+ goto cleanup_mgt;
+ return 0;
+
+cleanup_mgt:
+ nbl_chan_remove_msg_handler(chan_mgt);
+exit:
+ return ret;
+}
+
+void nbl_chan_remove_common(struct nbl_adapter *adap)
+{
+ struct nbl_channel_mgt *chan_mgt = adap->core.chan_mgt;
+
+ if (!chan_mgt)
+ return;
+ nbl_common_destroy_wq(&adap->common);
+ /*
+ * All channel queues shall be torn down earlier in remove path
+ * to drain inflight tx workers and stop hardware before destroying
+ * message handler xarray.
+ */
+ nbl_chan_remove_msg_handler(chan_mgt);
+ adap->core.chan_mgt = NULL;
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
new file mode 100644
index 000000000000..168f5de221bb
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.h
@@ -0,0 +1,196 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_CHANNEL_H_
+#define _NBL_CHANNEL_H_
+
+#include <linux/types.h>
+#include <linux/xarray.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_hw.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+#define NBL_CHAN_TX_RING_TO_DESC(tx_ring, i) \
+ (&((((tx_ring)->desc.tx_desc))[i]))
+#define NBL_CHAN_RX_RING_TO_DESC(rx_ring, i) \
+ (&((((rx_ring)->desc.rx_desc))[i]))
+#define NBL_CHAN_TX_RING_TO_BUF(tx_ring, i) (&(((tx_ring)->buf)[i]))
+#define NBL_CHAN_RX_RING_TO_BUF(rx_ring, i) (&(((rx_ring)->buf)[i]))
+
+#define NBL_CHAN_TX_WAIT_US 100
+#define NBL_CHAN_TX_WAIT_US_MAX 120
+#define NBL_CHAN_TX_WAIT_TIMES 100
+#define NBL_CHAN_TX_WAIT_ACK_US_MIN 1000
+#define NBL_CHAN_TX_WAIT_ACK_US_MAX 1200
+#define NBL_CHAN_TX_WAIT_ACK_TIMES 5000
+#define NBL_CHAN_QUEUE_LEN 256
+#define NBL_CHAN_BUF_LEN 4096
+#define NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN 16
+
+#define NBL_CHAN_TX_DESC_AVAIL 0
+#define NBL_CHAN_TX_DESC_USED 1
+#define NBL_CHAN_RX_DESC_WRITE 1
+#define NBL_CHAN_RX_DESC_AVAIL 3
+#define NBL_CHAN_RX_DESC_USED 4
+
+#define NBL_CHAN_ACK_HEAD_LEN 3
+#define NBL_CHAN_ACK_RET_POS 2
+#define NBL_CHAN_MSG_ID_POS 1
+#define NBL_CHAN_MSG_TYPE_POS 0
+
+#define NBL_CHAN_ACK_WAIT_TIME (3 * HZ)
+#define NBL_CHAN_RX_CLEAN_BUDGET 64
+
+enum {
+ NBL_MB_RX_QID = 0,
+ NBL_MB_TX_QID = 1,
+};
+
+enum {
+ NBL_MBX_STATUS_IDLE = 0,
+ NBL_MBX_STATUS_WAITING,
+ NBL_MBX_STATUS_ACKD,
+ NBL_MBX_STATUS_TIMEOUT,
+};
+
+struct nbl_chan_tx_param {
+ enum nbl_chan_msg_type msg_type;
+ void *arg;
+ size_t arg_len;
+ u16 dstid;
+ u16 msgid;
+};
+
+struct nbl_chan_buf {
+ void *va;
+ dma_addr_t pa;
+ size_t size;
+};
+
+struct nbl_chan_tx_desc {
+ __le16 flags;
+ __le16 srcid;
+ __le16 dstid;
+ __le16 data_len;
+ __le16 buf_len;
+ __le64 buf_addr;
+ __le16 msg_type;
+ u8 data[NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN];
+ __le16 msgid;
+ u8 rsv[26];
+} __packed;
+
+struct nbl_chan_rx_desc {
+ __le16 flags;
+ __le32 buf_len;
+ __le16 buf_id;
+ __le64 buf_addr;
+} __packed;
+
+union nbl_chan_desc_ptr {
+ struct nbl_chan_tx_desc *tx_desc;
+ struct nbl_chan_rx_desc *rx_desc;
+};
+
+struct nbl_chan_ring {
+ union nbl_chan_desc_ptr desc;
+ /* Per-entry slots carved out of one contiguous coherent block */
+ struct nbl_chan_buf *buf;
+ u16 next_to_use;
+ u16 tail_ptr; /* hardware does modulo ring size internally */
+ u16 next_to_clean;
+ dma_addr_t dma;
+};
+
+#define NBL_CHAN_MSG_INDEX_MAX 63
+
+#define NBL_CHAN_MSGID_INDEX_MASK GENMASK(5, 0)
+#define NBL_CHAN_MSGID_LOC_MASK GENMASK(13, 6)
+
+static inline void nbl_chan_update_tail_ptr(struct nbl_hw_ops *hw_ops,
+ void *hw_priv, u32 tail_ptr, u8 qid)
+{
+ hw_ops->update_mailbox_queue_tail_ptr(hw_priv, tail_ptr, qid);
+}
+
+struct nbl_chan_waitqueue_head {
+ struct wait_queue_head wait_queue;
+ char *ack_data;
+ int acked;
+ s32 ack_err;
+ u16 ack_data_len;
+ u16 msg_type;
+ int status;
+ u8 msg_index;
+ u16 dstid;
+};
+
+struct nbl_chan_info {
+ wait_queue_head_t inflight_wait;
+ struct nbl_chan_ring txq;
+ struct nbl_chan_ring rxq;
+ struct nbl_chan_waitqueue_head *wait;
+ /*
+ * Serializes TX ring producer state (next_to_use/next_to_clean/
+ * tail_ptr), descriptor publication and the doorbell, and the TX
+ * queue config/stop sequences against
+ * nbl_chan_quiesce_and_reclaim_tx().
+ */
+ struct mutex txq_lock;
+ /*
+ * Serializes the channel lifecycle: shutdn publication, clean_task
+ * registration/clearing, inflight_tx_cnt transitions and the
+ * inflight_wait wakeup. The state bitmap is not covered by this
+ * lock - it is only touched with atomic bit ops (set_bit(),
+ * clear_bit(), test_bit()).
+ */
+ struct mutex state_lock;
+ /*
+ * Serializes the ACK wait slot array (wait[]): slot allocation in
+ * nbl_chan_get_msg_id(), slot field publication in
+ * nbl_chan_send_msg(), ACK completion in nbl_chan_recv_ack_msg()
+ * and slot reset in nbl_chan_reset_wait_head().
+ */
+ struct mutex pending_lock;
+ struct work_struct *clean_task;
+ u16 wait_head_index;
+ u16 num_txq_entries;
+ u16 num_rxq_entries;
+ u16 txq_buf_size;
+ u16 rxq_buf_size;
+ DECLARE_BITMAP(state, NBL_CHAN_STATE_NBITS);
+ u8 chan_type;
+ atomic_t inflight_tx_cnt;
+ bool shutdn;
+ bool active;
+ /*
+ * One-shot lifecycle flag: DMA allocations use devm (descriptor/buf
+ * arrays) and dmam (coherent descriptor ring, per-ring contiguous
+ * buffer block) and are only released at device detach. Teardown
+ * only stops the hardware and clears active; it does not free
+ * memory. setup must never run twice on the same chan_info or the
+ * first allocation set would be orphaned.
+ */
+ bool dma_allocated;
+};
+
+struct nbl_chan_msg_node_data {
+ nbl_chan_resp func;
+ void *priv;
+};
+
+struct nbl_channel_mgt {
+ struct nbl_common_info *common;
+ struct nbl_hw_ops_tbl *hw_ops_tbl;
+ struct nbl_chan_info *chan_info[NBL_CHAN_TYPE_MAX];
+ /* Serializes handler_xa registration and teardown */
+ struct mutex handler_lock;
+ struct xarray handler_xa;
+};
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
new file mode 100644
index 000000000000..ce0ed2869af4
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
@@ -0,0 +1,31 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#include <linux/device.h>
+#include "nbl_common.h"
+
+void nbl_common_destroy_wq(struct nbl_common_info *common)
+{
+ if (!common || !common->wq)
+ return;
+
+ destroy_workqueue(common->wq);
+ common->wq = NULL;
+}
+
+int nbl_common_create_wq(struct nbl_common_info *common)
+{
+ char wq_name[32];
+
+ snprintf(wq_name, sizeof(wq_name), "nbl_wq_%s", pci_name(common->pdev));
+ common->wq = alloc_workqueue(wq_name, WQ_UNBOUND, 0);
+ if (!common->wq) {
+ dev_err(common->dev, "Failed to alloc workqueue %s\n", wq_name);
+ return -ENOMEM;
+ }
+
+ return 0;
+}
+
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h
new file mode 100644
index 000000000000..129073b08eb9
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.h
@@ -0,0 +1,14 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_COMMON_H_
+#define _NBL_COMMON_H_
+
+#include <linux/types.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_common.h"
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index 1cd6587a8fcb..f998a2b44e5c 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -14,13 +14,20 @@ enum {
NBL_CAP_HAS_NET_BIT,
};
+struct nbl_interface {
+ struct nbl_hw_ops_tbl *hw_ops_tbl;
+ struct nbl_channel_ops_tbl *channel_ops_tbl;
+};
+
struct nbl_core {
struct nbl_hw_mgt *hw_mgt;
+ struct nbl_channel_mgt *chan_mgt;
};
struct nbl_adapter {
struct pci_dev *pdev;
struct nbl_core core;
+ struct nbl_interface intf;
struct nbl_common_info common;
};
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index cf40ddc45192..eeff6216e4aa 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -6,8 +6,213 @@
#include <linux/pci.h>
#include <linux/bits.h>
#include <linux/io.h>
+#include <linux/spinlock.h>
+#include <linux/bitfield.h>
#include "nbl_hw_leonis.h"
+static void nbl_hw_read_mbx_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
+ u32 len)
+{
+ u32 i;
+
+ if (len % 4)
+ return;
+ if (reg >= (u64)hw_mgt->mailbox_bar_size ||
+ reg + len > (u64)hw_mgt->mailbox_bar_size) {
+ dev_err_once(hw_mgt->common->dev,
+ "mbx read out of range: reg=0x%llx len=%u bar_size=%pa\n",
+ reg, len, &hw_mgt->mailbox_bar_size);
+ return;
+ }
+ for (i = 0; i < len / 4; i++)
+ data[i] = nbl_mbx_rd32(hw_mgt, reg + i * sizeof(u32));
+}
+
+static void nbl_hw_write_mbx_regs(struct nbl_hw_mgt *hw_mgt, u64 reg,
+ const u32 *data, u32 len)
+{
+ u32 i;
+
+ if (len % 4)
+ return;
+ if (reg >= (u64)hw_mgt->mailbox_bar_size ||
+ reg + len > (u64)hw_mgt->mailbox_bar_size) {
+ dev_err_once(hw_mgt->common->dev,
+ "mbx write out of range: reg=0x%llx len=%u bar_size=%pa\n",
+ reg, len, &hw_mgt->mailbox_bar_size);
+ return;
+ }
+ for (i = 0; i < len / 4; i++)
+ nbl_mbx_wr32(hw_mgt, reg + i * sizeof(u32), data[i]);
+}
+
+/*
+ * Flush posted mailbox-BAR writes by reading back through the same
+ * BAR.
+ */
+static void nbl_hw_flush_mbx_write(struct nbl_hw_mgt *hw_mgt, u64 reg)
+{
+ u32 data;
+
+ nbl_hw_read_mbx_regs(hw_mgt, reg, &data, sizeof(data));
+}
+
+static void nbl_hw_rd_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
+ u32 len)
+{
+ u32 size = len / 4;
+ u32 i;
+
+ if (len % 4)
+ return;
+ for (i = 0; i < size; i++)
+ data[i] = rd32(hw_mgt->hw_addr, reg + i * sizeof(u32));
+}
+
+static void nbl_hw_wr_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, const u32 *data,
+ u32 len)
+{
+ u32 size = len / 4;
+ u32 i;
+
+ if (len % 4)
+ return;
+ for (i = 0; i < size; i++)
+ wr32(hw_mgt->hw_addr, reg + i * sizeof(u32), data[i]);
+}
+
+static void nbl_hw_rd_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
+ u32 len)
+{
+ u32 size = len / 4;
+ u32 i;
+
+ if (len % 4)
+ return;
+
+ spin_lock(&hw_mgt->reg_lock);
+
+ for (i = 0; i < size; i++)
+ data[i] = rd32(hw_mgt->hw_addr, reg + i * sizeof(u32));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_hw_update_mailbox_queue_tail_ptr(struct nbl_hw_mgt *hw_mgt,
+ u16 tail_ptr, u8 txrx)
+{
+ /* local_qid 0 and 1 denote rx and tx queue respectively */
+ u32 local_qid = txrx;
+ u32 value = ((u32)tail_ptr << 16) | local_qid;
+
+ /* wmb for doorbell */
+ wmb();
+ nbl_mbx_wr32(hw_mgt, NBL_MAILBOX_NOTIFY_ADDR, value);
+}
+
+static void nbl_hw_config_mailbox_rxq(struct nbl_hw_mgt *hw_mgt,
+ dma_addr_t dma_addr, int size_bwid)
+{
+ struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+ memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+ cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+ nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR,
+ cfg_tbl.data, sizeof(cfg_tbl));
+
+ cfg_tbl.data[0] = lower_32_bits(dma_addr);
+ cfg_tbl.data[1] = upper_32_bits(dma_addr);
+ cfg_tbl.data[2] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_SIZE_BWID_MASK,
+ size_bwid);
+ cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 0) |
+ FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_EN_MASK, 1);
+ nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR,
+ cfg_tbl.data, sizeof(cfg_tbl));
+}
+
+static void nbl_hw_config_mailbox_txq(struct nbl_hw_mgt *hw_mgt,
+ dma_addr_t dma_addr, int size_bwid)
+{
+ struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+ memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+ cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+ nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR,
+ cfg_tbl.data, sizeof(cfg_tbl));
+
+ cfg_tbl.data[0] = lower_32_bits(dma_addr);
+ cfg_tbl.data[1] = upper_32_bits(dma_addr);
+ cfg_tbl.data[2] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_SIZE_BWID_MASK,
+ size_bwid);
+ cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 0) |
+ FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_EN_MASK, 1);
+ nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR,
+ cfg_tbl.data, sizeof(cfg_tbl));
+}
+
+static void nbl_hw_stop_mailbox_rxq(struct nbl_hw_mgt *hw_mgt)
+{
+ struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+ memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+ cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+ nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR,
+ cfg_tbl.data, sizeof(cfg_tbl));
+ /* Ensure QUEUE_RST has reached the device before caller proceeds */
+ nbl_hw_flush_mbx_write(hw_mgt,
+ NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR);
+}
+
+static void nbl_hw_stop_mailbox_txq(struct nbl_hw_mgt *hw_mgt)
+{
+ struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
+
+ memset(&cfg_tbl, 0, sizeof(cfg_tbl));
+ cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
+ nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR,
+ cfg_tbl.data, sizeof(cfg_tbl));
+ /* Ensure QUEUE_RST has reached the device before caller proceeds */
+ nbl_hw_flush_mbx_write(hw_mgt,
+ NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR);
+}
+
+static void nbl_hw_get_host_pf_mask(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask)
+{
+ nbl_hw_rd_regs_lock(hw_mgt, NBL_PCIE_HOST_K_PF_MASK_REG, pf_mask,
+ sizeof(*pf_mask));
+}
+
+static void nbl_hw_cfg_mailbox_qinfo(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ u8 bus, u8 devid, u8 function)
+{
+ u32 data = 0;
+
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id),
+ &data, sizeof(data));
+ data &= ~(NBL_MAILBOX_QINFO_MAP_FUNCTION_MASK |
+ NBL_MAILBOX_QINFO_MAP_DEVID_MASK |
+ NBL_MAILBOX_QINFO_MAP_BUS_MASK |
+ NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK |
+ NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK);
+ data |= FIELD_PREP(NBL_MAILBOX_QINFO_MAP_FUNCTION_MASK, function) |
+ FIELD_PREP(NBL_MAILBOX_QINFO_MAP_DEVID_MASK, devid) |
+ FIELD_PREP(NBL_MAILBOX_QINFO_MAP_BUS_MASK, bus);
+ nbl_hw_wr_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id),
+ &data, sizeof(data));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static struct nbl_hw_ops hw_ops = {
+ .update_mailbox_queue_tail_ptr = nbl_hw_update_mailbox_queue_tail_ptr,
+ .config_mailbox_rxq = nbl_hw_config_mailbox_rxq,
+ .config_mailbox_txq = nbl_hw_config_mailbox_txq,
+ .stop_mailbox_rxq = nbl_hw_stop_mailbox_rxq,
+ .stop_mailbox_txq = nbl_hw_stop_mailbox_txq,
+ .get_host_pf_mask = nbl_hw_get_host_pf_mask,
+ .cfg_mailbox_qinfo = nbl_hw_cfg_mailbox_qinfo,
+
+};
+
/* Structure starts here, adding an op should not modify anything below */
static struct nbl_hw_mgt *nbl_hw_setup_hw_mgt(struct nbl_common_info *common)
{
@@ -23,6 +228,27 @@ static struct nbl_hw_mgt *nbl_hw_setup_hw_mgt(struct nbl_common_info *common)
return hw_mgt;
}
+static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
+ struct nbl_hw_mgt *hw_mgt)
+{
+ struct nbl_hw_ops_tbl *hw_ops_tbl;
+ struct device *dev;
+
+ dev = common->dev;
+ hw_ops_tbl = devm_kzalloc(dev, sizeof(*hw_ops_tbl), GFP_KERNEL);
+ if (!hw_ops_tbl)
+ return ERR_PTR(-ENOMEM);
+ if (!hw_ops.update_mailbox_queue_tail_ptr ||
+ !hw_ops.config_mailbox_rxq || !hw_ops.config_mailbox_txq ||
+ !hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
+ !hw_ops.get_host_pf_mask || !hw_ops.cfg_mailbox_qinfo)
+ return ERR_PTR(-EINVAL);
+ hw_ops_tbl->ops = &hw_ops;
+ hw_ops_tbl->priv = hw_mgt;
+
+ return hw_ops_tbl;
+}
+
static int nbl_pcim_request_selected_bars(struct pci_dev *pdev, u32 mask,
const char *name)
{
@@ -43,6 +269,7 @@ int nbl_hw_init_leonis(struct nbl_adapter *adapter)
{
resource_size_t expect_sz = NBL_MEM_BAR_TOTAL_SIZE;
struct nbl_common_info *common = &adapter->common;
+ struct nbl_hw_ops_tbl *hw_ops_tbl = NULL;
struct pci_dev *pdev = common->pdev;
struct nbl_hw_mgt *hw_mgt = NULL;
resource_size_t bar_len;
@@ -136,7 +363,14 @@ int nbl_hw_init_leonis(struct nbl_adapter *adapter)
}
hw_mgt->mailbox_bar_size = bar_len;
+ spin_lock_init(&hw_mgt->reg_lock);
+ hw_ops_tbl = nbl_hw_setup_ops(common, hw_mgt);
+ if (IS_ERR(hw_ops_tbl)) {
+ ret = PTR_ERR(hw_ops_tbl);
+ goto setup_mgt_fail;
+ }
+ adapter->intf.hw_ops_tbl = hw_ops_tbl;
adapter->core.hw_mgt = hw_mgt;
return 0;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index 1f9e509dc631..93e5c7518288 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -11,5 +11,46 @@
#include "../../nbl_include/nbl_include.h"
#include "../nbl_hw_reg.h"
+/* ---------- REG BASE ADDR ---------- */
+/* Interface modules base addr */
+#define NBL_INTF_HOST_PCOMPLETER_BASE 0x00f08000
+#define NBL_INTF_HOST_PADPT_BASE 0x00f4c000
+#define NBL_INTF_HOST_MAILBOX_BASE 0x00fb0000
+#define NBL_INTF_HOST_PCIE_BASE 0X01504000
+/* -------- MAILBOX BAR2 ----- */
+#define NBL_MAILBOX_NOTIFY_ADDR 0x00000000
+#define NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR 0x10
+#define NBL_MAILBOX_QINFO_CFG_TX_TABLE_ADDR 0x20
+
+/* -------- MAILBOX -------- */
+
+/* mailbox BAR qinfo_cfg_table */
+#define MAILBOX_QINFO_CFG_TABLE_DWLEN 4
+/* data[2] */
+#define NBL_MAILBOX_QINFO_CFG_QUEUE_SIZE_BWID_MASK GENMASK(3, 0)
+/* data[3] */
+#define NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK BIT(0)
+#define NBL_MAILBOX_QINFO_CFG_QUEUE_EN_MASK BIT(1)
+#define NBL_MAILBOX_QINFO_CFG_DIF_ERR_MASK BIT(2)
+#define NBL_MAILBOX_QINFO_CFG_PTR_ERR_MASK BIT(3)
+struct nbl_mailbox_qinfo_cfg_table {
+ u32 data[MAILBOX_QINFO_CFG_TABLE_DWLEN];
+};
+
+/* -------- MAILBOX BAR0 ----- */
+/* mailbox qinfo_map_table */
+#define NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id) \
+ (NBL_INTF_HOST_MAILBOX_BASE + 0x00001000 + (func_id) * sizeof(u32))
+
+/* MAILBOX qinfo_map_table */
+#define NBL_MAILBOX_QINFO_MAP_FUNCTION_MASK GENMASK(2, 0)
+#define NBL_MAILBOX_QINFO_MAP_DEVID_MASK GENMASK(7, 3)
+#define NBL_MAILBOX_QINFO_MAP_BUS_MASK GENMASK(15, 8)
+#define NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK GENMASK(28, 16)
+#define NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK BIT(29)
+
+/* -------- HOST_PCIE -------- */
+#define NBL_PCIE_HOST_K_PF_MASK_REG (NBL_INTF_HOST_PCIE_BASE + 0x00001004)
+
#define NBL_BAR2_MAX_LEN 0x300
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
index e281109f502e..4c1bb789c465 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
@@ -8,6 +8,7 @@
#include <linux/types.h>
+#include "../nbl_include/nbl_def_channel.h"
#include "../nbl_include/nbl_def_hw.h"
#include "../nbl_include/nbl_def_common.h"
#include "../nbl_core.h"
@@ -26,6 +27,37 @@ struct nbl_hw_mgt {
u8 __iomem *hw_addr;
u8 __iomem *mailbox_bar_hw_addr;
resource_size_t mailbox_bar_size;
+ spinlock_t reg_lock; /* Protect reg access */
};
+static inline u32 rd32(u8 __iomem *addr, u64 reg)
+{
+ return readl(addr + reg);
+}
+
+static inline void wr32(u8 __iomem *addr, u64 reg, u32 value)
+{
+ writel(value, addr + reg);
+}
+
+static inline void nbl_hw_wr32(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 value)
+{
+ wr32(hw_mgt->hw_addr, reg, value);
+}
+
+static inline u32 nbl_hw_rd32(struct nbl_hw_mgt *hw_mgt, u64 reg)
+{
+ return rd32(hw_mgt->hw_addr, reg);
+}
+
+static inline void nbl_mbx_wr32(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 value)
+{
+ writel(value, hw_mgt->mailbox_bar_hw_addr + reg);
+}
+
+static inline u32 nbl_mbx_rd32(struct nbl_hw_mgt *hw_mgt, u64 reg)
+{
+ return readl(hw_mgt->mailbox_bar_hw_addr + reg);
+}
+
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
new file mode 100644
index 000000000000..d6faa0bc4026
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
@@ -0,0 +1,125 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_CHANNEL_H_
+#define _NBL_DEF_CHANNEL_H_
+
+#include <linux/types.h>
+
+struct nbl_channel_mgt;
+struct nbl_adapter;
+
+typedef void (*nbl_chan_resp)(void *, u16, u16, void *, u32);
+
+/*
+ * Mailbox wire opcodes, stable wire ABI shared between driver and firmware.
+ * Each opcode has a fixed assigned number to preserve compatibility.
+ * ABI compatibility rules:
+ * 1. New opcodes shall only be appended before NBL_CHAN_MSG_MAILBOX_MAX;
+ * 2. Reordering, inserting or deleting existing enumerators breaks driver-
+ * firmware interoperability and must be avoided;
+ * 3. Modifications to existing opcodes require synchronized firmware ABI
+ * updates.
+ *
+ * Only opcodes currently used by in-tree driver logic are defined here.
+ * Unimplemented feature opcodes (KTLS, IPsec, vDPA, mirror etc.) will be
+ * added incrementally together with their corresponding driver
+ * implementation patches.
+ */
+enum nbl_chan_msg_type {
+ NBL_CHAN_MSG_ACK = 0,
+ /* mailbox msg end */
+ NBL_CHAN_MSG_MAILBOX_MAX,
+};
+
+enum nbl_chan_state {
+ NBL_CHAN_IRQ_RDY,
+ NBL_CHAN_STATE_NBITS
+};
+
+struct nbl_chan_send_info {
+ void *arg;
+ size_t arg_len;
+ void *resp;
+ size_t resp_len;
+ u16 dstid;
+ u16 msg_type;
+ u16 ack;
+ u16 ack_len;
+};
+
+struct nbl_chan_ack_info {
+ void *data;
+ int err;
+ u32 data_len;
+ u16 dstid;
+ u16 msg_type;
+ u16 msgid;
+};
+
+enum nbl_channel_type {
+ NBL_CHAN_TYPE_MAILBOX,
+ NBL_CHAN_TYPE_MAX
+};
+
+static inline void
+nbl_chan_fill_send_info(struct nbl_chan_send_info *info,
+ u16 dst_id, u16 msg_type,
+ void *argument, u32 arg_length,
+ void *response, u32 resp_length,
+ bool need_ack)
+{
+ info->dstid = dst_id;
+ info->msg_type = msg_type;
+ info->arg = argument;
+ info->arg_len = arg_length;
+ info->resp = response;
+ info->resp_len = resp_length;
+ info->ack = need_ack;
+}
+
+static inline void
+nbl_chan_fill_ack_info(struct nbl_chan_ack_info *info,
+ u16 dst_id, u16 msg_type, u16 msg_id,
+ int err_code, void *ack_data, u32 data_length)
+{
+ info->dstid = dst_id;
+ info->msg_type = msg_type;
+ info->msgid = msg_id;
+ info->err = err_code;
+ info->data = ack_data;
+ info->data_len = data_length;
+}
+
+struct nbl_channel_ops {
+ int (*send_msg)(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_send_info *chan_send);
+ int (*send_ack)(struct nbl_channel_mgt *chan_mgt,
+ struct nbl_chan_ack_info *chan_ack);
+ int (*register_msg)(struct nbl_channel_mgt *chan_mgt, u16 msg_type,
+ nbl_chan_resp func, void *callback_priv);
+ void (*cfg_chan_qinfo_map_table)(struct nbl_channel_mgt *chan_mgt,
+ u8 bus, u8 devid);
+ bool (*check_queue_exist)(struct nbl_channel_mgt *chan_mgt,
+ u8 chan_type);
+ int (*setup_queue)(struct nbl_channel_mgt *chan_mgt, u8 chan_type);
+ int (*teardown_queue)(struct nbl_channel_mgt *chan_mgt, u8 chan_type);
+ void (*clean_queue_subtask)(struct nbl_channel_mgt *chan_mgt,
+ u8 chan_type);
+ void (*register_chan_task)(struct nbl_channel_mgt *chan_mgt,
+ u8 chan_type, struct work_struct *task);
+ void (*set_queue_state)(struct nbl_channel_mgt *chan_mgt,
+ enum nbl_chan_state state, u8 chan_type,
+ u8 set);
+};
+
+struct nbl_channel_ops_tbl {
+ struct nbl_channel_ops *ops;
+ struct nbl_channel_mgt *priv;
+};
+
+int nbl_chan_init_common(struct nbl_adapter *adapter);
+void nbl_chan_remove_common(struct nbl_adapter *adapter);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
index ea646d1efc2d..26a364e3a194 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
@@ -12,6 +12,7 @@
#include "nbl_include.h"
struct nbl_common_info {
+ struct workqueue_struct *wq;
struct pci_dev *pdev;
struct device *dev;
u16 vsi_id;
@@ -28,4 +29,7 @@ struct nbl_common_info {
u8 has_net;
};
+void nbl_common_destroy_wq(struct nbl_common_info *common);
+int nbl_common_create_wq(struct nbl_common_info *common);
+
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index ecbf440e4366..58d1b669b8cb 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -10,6 +10,47 @@
struct nbl_hw_mgt;
struct nbl_adapter;
+struct nbl_hw_ops {
+ void (*update_mailbox_queue_tail_ptr)(struct nbl_hw_mgt *hw_mgt,
+ u16 tail_ptr, u8 txrx);
+ void (*config_mailbox_rxq)(struct nbl_hw_mgt *hw_mgt,
+ dma_addr_t dma_addr, int size_bwid);
+ void (*config_mailbox_txq)(struct nbl_hw_mgt *hw_mgt,
+ dma_addr_t dma_addr, int size_bwid);
+ void (*stop_mailbox_rxq)(struct nbl_hw_mgt *hw_mgt);
+ void (*stop_mailbox_txq)(struct nbl_hw_mgt *hw_mgt);
+ /**
+ * get_host_pf_mask - Fetch host PF mask from firmware k_pf_mask reg
+ * @hw_mgt: hardware management context
+ * @pf_mask: output pointer for PF mask value
+ *
+ * k_pf_mask register rule:
+ * bit N == 0 -> PF#N enabled; bit N == 1 -> PF#N masked out.
+ * bit0 is PF0's mask bit (not reserved); PF0 can be masked but
+ * the driver requires at least PF0 enabled.
+ * Only 1/2/4 PFs are supported:
+ * 1 PF (PF0): mask = 0xfe
+ * 2 PFs (PF0,PF1): mask = 0xfc
+ * 4 PFs (PF0~PF3): mask = 0xf0
+ * All-zero mask (0x00) means all 8 PFs enabled, which is
+ * unsupported by the driver.
+ *
+ * The mask is validated in the resource layer:
+ * nbl_res_init_pf_num() rejects an all-zero mask and any
+ * non-contiguous layout with -EINVAL, and it always runs before
+ * this op's only consumer, nbl_chan_cfg_qinfo_map_table(). This
+ * op only returns the raw register value.
+ */
+ void (*get_host_pf_mask)(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask);
+
+ void (*cfg_mailbox_qinfo)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ u8 bus, u8 devid, u8 function);
+};
+
+struct nbl_hw_ops_tbl {
+ struct nbl_hw_ops *ops;
+ struct nbl_hw_mgt *priv;
+};
int nbl_hw_init_leonis(struct nbl_adapter *adapter);
void nbl_hw_remove_leonis(struct nbl_adapter *adapter);
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
index 14e7b19f9a4c..f2d802397d98 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
@@ -10,6 +10,9 @@
/* ------ Basic definitions ------- */
#define NBL_DRIVER_NAME "nbl"
+#define NBL_MAX_PF 8
+#define NBL_NEXT_ID(id, max) (((id) + 1) % ((max) + 1))
+
struct nbl_func_caps {
u32 has_ctrl:1;
u32 has_net:1;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index f2552bc73293..62484fbf1fef 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -8,6 +8,7 @@
#include <linux/module.h>
#include <linux/bits.h>
#include "nbl_include/nbl_include.h"
+#include "nbl_include/nbl_def_channel.h"
#include "nbl_include/nbl_def_hw.h"
#include "nbl_include/nbl_def_common.h"
#include "nbl_core.h"
@@ -38,13 +39,19 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
if (ret)
goto hw_init_fail;
+ ret = nbl_chan_init_common(adapter);
+ if (ret)
+ goto chan_init_fail;
return adapter;
+chan_init_fail:
+ nbl_hw_remove_leonis(adapter);
hw_init_fail:
return ERR_PTR(ret);
}
void nbl_core_remove(struct nbl_adapter *adapter)
{
+ nbl_chan_remove_common(adapter);
nbl_hw_remove_leonis(adapter);
}
@@ -67,6 +74,109 @@ static void nbl_get_func_param(struct pci_dev *pdev, kernel_ulong_t driver_data,
param->caps.has_ctrl = 1;
}
+/*
+ * Defined at the end of this file; referenced to verify that func 0 is
+ * bound to this driver before a device link is created to it.
+ */
+static struct pci_driver nbl_driver;
+
+/*
+ * Establish chip-wide dependencies for this PF:
+ * - for every non-management PF, add a consumer->management PF device
+ * link. The driver core then guarantees (sysfs unbind, driver
+ * unregister, hot-unplug alike) that this consumer is released
+ * BEFORE the func 0 supplier, which is the only teardown order in
+ * which the chip-global firmware deinit is safe.
+ *
+ * The management PF is addressed by the deterministic identity
+ * (domain, bus, slot, func 0) instead of any name-based scan: hardware
+ * guarantees PFs are contiguous from func 0 in the same slot.
+ *
+ * func 0 must already be bound to this driver, not merely to some other
+ * driver: links.status == DL_DEV_DRIVER_BOUND alone would also accept
+ * func 0 bound to vfio-pci or pci-stub through driver_override, where the
+ * chip-global setup this PF depends on never ran. The bound driver and
+ * the binding state are read under device_lock, so they cannot change
+ * while they are being checked.
+ *
+ * The link is a persistent managed link with DL_FLAG_AUTOPROBE_CONSUMER:
+ * the core force-unbinds this consumer when func 0 unbinds (keeping the
+ * teardown order) and re-queues it for probe when func 0 binds again, so
+ * the sibling PFs and their ports come back without a manual rebind.
+ * AUTOREMOVE_CONSUMER would delete the link on unbind and cannot be
+ * combined with AUTOPROBE_CONSUMER (device_link_add() returns NULL for
+ * that pair). The link is therefore not auto-purged on a failed probe:
+ * it lives with the two devices and is dropped when either is removed.
+ *
+ * Return: 0 on success, negative errno on failure.
+ */
+static int nbl_probe_chip_deps(struct pci_dev *pdev, bool has_ctrl)
+{
+ struct device_link *link;
+ struct pci_dev *mgt;
+
+ if (has_ctrl)
+ return 0;
+
+ mgt = pci_get_domain_bus_and_slot(pci_domain_nr(pdev->bus),
+ pdev->bus->number,
+ PCI_DEVFN(PCI_SLOT(pdev->devfn), 0));
+ if (!mgt) {
+ dev_err(&pdev->dev,
+ "management PF (func 0) not found on this chip\n");
+ return -ENODEV;
+ }
+
+ /*
+ * The management PF must have COMPLETED probing with this driver.
+ * While func 0 is still PROBING, its chip-global init (mailbox
+ * QINFO map programming) can race a sibling's mailbox setup, and a
+ * func 0 bound to another driver never ran it at all. Defer until
+ * func 0 is bound to nbl; this also covers the unbound case.
+ */
+ device_lock(&mgt->dev);
+ if (mgt->dev.driver != &nbl_driver.driver ||
+ mgt->dev.links.status != DL_DEV_DRIVER_BOUND) {
+ device_unlock(&mgt->dev);
+ pci_dev_put(mgt);
+ return -EPROBE_DEFER;
+ }
+ device_unlock(&mgt->dev);
+
+ link = device_link_add(&pdev->dev, &mgt->dev,
+ DL_FLAG_AUTOPROBE_CONSUMER);
+ if (!link) {
+ dev_err(&pdev->dev,
+ "failed to create device link to management PF\n");
+ pci_dev_put(mgt);
+ return -ENOMEM;
+ }
+
+ /*
+ * func 0 can start unbinding between the check above and link
+ * creation: unbind and this probe take different device locks, so
+ * the new link can land in DORMANT or CONSUMER_PROBE against a
+ * supplier that is going away, and the core's unbind-consumers
+ * pass may already have run. Re-check under device_lock and defer
+ * if func 0 is no longer bound; the AUTOPROBE link makes the core
+ * re-probe this PF once func 0 is bound again.
+ *
+ * The link pointer is deliberately never dereferenced: for a
+ * managed link device_link_add() only reports whether the link
+ * exists, and device 0's removal could free it concurrently.
+ */
+ device_lock(&mgt->dev);
+ if (mgt->dev.links.status != DL_DEV_DRIVER_BOUND) {
+ device_unlock(&mgt->dev);
+ pci_dev_put(mgt);
+ return -EPROBE_DEFER;
+ }
+ device_unlock(&mgt->dev);
+ pci_dev_put(mgt);
+
+ return 0;
+}
+
static int nbl_probe(struct pci_dev *pdev,
const struct pci_device_id *id)
{
@@ -87,6 +197,10 @@ static int nbl_probe(struct pci_dev *pdev,
pci_set_master(pdev);
+ err = nbl_probe_chip_deps(pdev, param.caps.has_ctrl);
+ if (err)
+ goto chip_deps_err;
+
adapter = nbl_core_init(pdev, ¶m);
if (IS_ERR(adapter)) {
dev_err(dev, "Nbl adapter init fail: %pe\n", adapter);
@@ -96,6 +210,7 @@ static int nbl_probe(struct pci_dev *pdev,
pci_set_drvdata(pdev, adapter);
return 0;
adapter_init_err:
+chip_deps_err:
pci_clear_master(pdev);
return err;
}
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 1/8] net/nebula-matrix: add channel layer
2026-10-10 9:59 ` [PATCH v31 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found 9 potential
issue(s) to consider.
Critical: 0 · High: 1 · Medium: 5 · Low: 3
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] Lock-order inversion (ABBA) between nbl_probe_chip_deps() and
the PCI core's bus/slot locking.
- [Medium] nbl_probe_chip_deps() in drivers/net/ethernet/nebula-
matrix/nbl/nbl_main.c changes how every non-function-0 PF probes, but…
- [Medium] nbl_chan_setup_queue() sets chan_info->dma_allocated = true
before allocating, and the comment says this is so that "a retry must…
- [Medium] nbl_probe_chip_deps() can self-deadlock when the probing
device is an SR-IOV VF whose devfn has function number 0.
- [Medium] In nbl_probe_chip_deps(), the PF0 checks and the device-link
creation are not atomic, and the second check does not close the race.
- [Medium] nbl_hw_stop_mailbox_rxq() and nbl_hw_stop_mailbox_txq() write
the whole 4-dword QINFO_CFG entry through nbl_hw_write_mbx_regs(), one…
- [Low] A comment in nbl_chan_send_msg() says `active` "is published last
in setup and cleared last in teardown, so a false value here means the…
- [Low] nbl_chan_teardown_queue() waits for inflight_tx_cnt to reach 0
and then does one mutex_lock()/mutex_unlock() of state_lock.
- [Low] (Baseline: not a defect.) In the external-buffer branch of
nbl_chan_update_txqueue(), buf_addr and buf_len are programmed even…
--
net/nebula-matrix: add channel layer
This adds a mailbox-based channel layer that carries control messages
between PF0 and the other PFs. It has xarray handler lookup, ACK-based
sends, coherent DMA rings and a quiescing teardown. It also hooks
channel init/cleanup into probe/remove, and adds a device link from
each PF other than function 0 to function 0.
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
> new file mode 100644
> index 0000000000000..ac3115a3ab55f
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_channel/nbl_channel.c
> @@ -0,0 +1,1486 @@
[ ... ]
> +static int nbl_chan_teardown_queue(struct nbl_channel_mgt *chan_mgt,
> + u8 chan_type)
> +{
[ ... ]
> + /*
> + * Join the last exiting sender through state_lock: it drops
> + * inflight_tx_cnt and wakes this drain while still holding
> + * state_lock, so observing the counter reach zero does not prove
> + * the sender has finished its mutex_unlock(). Acquiring the lock
> + * once here guarantees no sender still holds it - mutex_unlock()
> + * is the sender's last touch of chan_info - before returning lets
> + * devres free chan_info. shutdn blocks new senders, so one
> + * acquire/release pair is sufficient.
> + */
> + mutex_lock(&chan_info->state_lock);
> + mutex_unlock(&chan_info->state_lock);
[Severity: Low]
Does one lock/unlock pair give the guarantee this comment claims?
Documentation/locking/mutex-design.rst says:
mutex_unlock() may access the mutex structure even after it has
internally released the lock already - so it's not safe for another
context to acquire the mutex and assume that the mutex_unlock()
context is not using the structure anymore.
The last sender exits nbl_chan_send_msg() at out_clean_inflight like
this:
mutex_lock(&chan_info->state_lock);
if (atomic_dec_and_test(&chan_info->inflight_tx_cnt))
wake_up(&chan_info->inflight_wait);
mutex_unlock(&chan_info->state_lock);
The unlock slowpath of that sender can still be touching state_lock
after teardown has acquired the lock, released it and returned.
Today chan_info is devm-allocated and freed only at detach, so this
series cannot reach a use-after-free. Still, the claim would no longer
hold if chan_info were freed, or mutex_destroy() were called, right
after teardown.
[ ... ]
> +static int nbl_chan_setup_queue(struct nbl_channel_mgt *chan_mgt, u8 chan_type)
> +{
[ ... ]
> + if (chan_info->dma_allocated) {
> + dev_warn(common->dev,
> + "channel DMA already allocated, re-setup not supported\n");
> + return -EBUSY;
> + }
> +
> + nbl_chan_init_queue_param(chan_info, NBL_CHAN_QUEUE_LEN,
> + NBL_CHAN_QUEUE_LEN, NBL_CHAN_BUF_LEN,
> + NBL_CHAN_BUF_LEN);
> + /*
> + * Set the flag before the first allocation: if setup fails halfway,
> + * the buffers allocated so far already live until detach, so a retry
> + * must be rejected instead of allocating a second set and orphaning
> + * the first one.
> + */
> + chan_info->dma_allocated = true;
> + err = nbl_chan_init_queue(common, chan_info);
> + if (err) {
> + chan_info->dma_allocated = false;
> + return err;
> + }
[Severity: Medium]
Clearing dma_allocated here seems to allow the very retry that the
comment above says must be rejected. Is that intended?
nbl_chan_init_tx_queue() and nbl_chan_init_rx_queue() allocate with
dmam_alloc_coherent() and devm_kcalloc(). Anything allocated before a
failure part-way through stays allocated until detach.
A second setup_queue() call would then pass the dma_allocated check and
allocate a whole new set: about 2 MiB of coherent DMA, plus the
descriptor rings and wait[]. It would overwrite txq/rxq/wait and leave
the first set unreferenced.
A later patch in the series ("add common/ctrl dev init/remove
operation") depends on this guard. Its comment in
nbl_dev_setup_common_dev() says "chan_info->dma_allocated blocks a
re-setup".
Also, nbl_chan_teardown_queue() never clears dma_allocated, so the
channel cannot be set up again after teardown. Does that fit the "init,
configuration, start, stop and teardown" lifecycle described in the
commit message?
[ ... ]
> + if (param->arg_len > NBL_CHAN_TX_DESC_EMBEDDED_DATA_LEN) {
> + if (param->arg)
> + memcpy(tx_buf->va, param->arg, param->arg_len);
> + tx_desc->buf_addr = cpu_to_le64(tx_buf->pa);
> + tx_desc->buf_len = cpu_to_le16(param->arg_len);
[Severity: Low]
This isn't a bug, but in nbl_chan_update_txqueue() buf_addr and buf_len
are programmed even when param->arg is NULL. A NULL arg with a non-zero
arg_len would send whatever is left in the slot from before. No caller
in the series passes that combination today. Would it be simpler to
reject it up front?
[ ... ]
> + /*
> + * Entry points are only reached after a successful setup_queue()
> + * (probe path), but the channel is a dynamic object: fail cleanly
> + * instead of dereferencing rings/wait[] that were never allocated
> + * (num_txq_entries == 0). active is published last in setup and
> + * cleared last in teardown, so a false value here means the channel
> + * was never set up; the teardown window itself is covered by the
> + * shutdn check below.
> + */
> + if (!READ_ONCE(chan_info->active))
> + return -ENODEV;
[Severity: Low]
Is this comment still true once teardown has completed?
nbl_chan_teardown_queue() ends with WRITE_ONCE(chan_info->active, false)
and leaves shutdn set. dma_allocated then stops setup from running
again.
A sender that arrives after teardown takes this branch and gets -ENODEV
instead of -ESHUTDOWN, even though the channel was set up.
[ ... ]
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> index cf40ddc451921..eeff6216e4aa5 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
[ ... ]
> +static void nbl_hw_stop_mailbox_rxq(struct nbl_hw_mgt *hw_mgt)
> +{
> + struct nbl_mailbox_qinfo_cfg_table cfg_tbl;
> +
> + memset(&cfg_tbl, 0, sizeof(cfg_tbl));
> + cfg_tbl.data[3] = FIELD_PREP(NBL_MAILBOX_QINFO_CFG_QUEUE_RST_MASK, 1);
> + nbl_hw_write_mbx_regs(hw_mgt, NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR,
> + cfg_tbl.data, sizeof(cfg_tbl));
[Severity: Medium]
Is it safe to zero the ring base and size before QUEUE_RST takes effect?
nbl_hw_write_mbx_regs() writes the entry one dword at a time, data[0]
through data[3]:
for (i = 0; i < len / 4; i++)
nbl_mbx_wr32(hw_mgt, reg + i * sizeof(u32), data[i]);
data[0..2] (base low/high and QUEUE_SIZE_BWID) are therefore written as
zero while QUEUE_EN is still set. QUEUE_RST arrives only with data[3].
Teardown calls nbl_chan_stop_rx_queue() while the peer can still send.
nbl_chan_quiesce_and_reclaim_tx() stops TX while a descriptor is still
outstanding.
If the device applies each dword as it is written, could a transfer in
that window DMA to IOVA 0 or raise PTR_ERR/DIF_ERR? The same applies to
nbl_hw_stop_mailbox_txq(), and to the first {0, 0, 0, RST} write in
nbl_hw_config_mailbox_rxq() and nbl_hw_config_mailbox_txq().
If the hardware does not commit the entry on the data[3] write, would
writing only data[3] with QUEUE_RST set first avoid this?
[ ... ]
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> index f2552bc732937..62484fbf1fef6 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
[ ... ]
> +static int nbl_probe_chip_deps(struct pci_dev *pdev, bool has_ctrl)
> +{
> + struct device_link *link;
> + struct pci_dev *mgt;
> +
> + if (has_ctrl)
> + return 0;
> +
> + mgt = pci_get_domain_bus_and_slot(pci_domain_nr(pdev->bus),
> + pdev->bus->number,
> + PCI_DEVFN(PCI_SLOT(pdev->devfn), 0));
> + if (!mgt) {
> + dev_err(&pdev->dev,
> + "management PF (func 0) not found on this chip\n");
> + return -ENODEV;
> + }
[Severity: Medium]
The commit message only says:
This patch also hooks channel init/cleanup into the PF probe/remove
paths in nbl_main.c.
nbl_probe_chip_deps() also changes how every PF with PCI_FUNC != 0
probes:
- if function 0 is missing from the slot, probe fails with -ENODEV
- if function 0 is unbound, or bound to vfio-pci or pci-stub, probe
returns -EPROBE_DEFER indefinitely
- a persistent DL_FLAG_AUTOPROBE_CONSUMER link is created, so
unbinding function 0 force-unbinds every sibling PF
Nothing in this patch uses that ordering yet, because setup_queue and
cfg_chan_qinfo_map_table have no callers. Later patches ("add
common/ctrl dev init/remove operation", "add common dev start/stop
operation") rely on the link but only mention it in passing.
Could this go in its own patch, or at least be described and justified
in this commit message? A setup with PF0 on vfio-pci and PF1..3 on nbl
would now leave the sibling PFs deferred forever.
> +
> + /*
> + * The management PF must have COMPLETED probing with this driver.
> + * While func 0 is still PROBING, its chip-global init (mailbox
> + * QINFO map programming) can race a sibling's mailbox setup, and a
> + * func 0 bound to another driver never ran it at all. Defer until
> + * func 0 is bound to nbl; this also covers the unbound case.
> + */
> + device_lock(&mgt->dev);
[Severity: Medium]
Can this self-deadlock when the probing device is an SR-IOV VF whose
devfn decodes to function 0?
nbl_get_func_param() sets has_ctrl only for a function 0 that is not a
VF:
if ((PCI_FUNC(pdev->devfn) == 0) && !pdev->is_virtfn)
param->caps.has_ctrl = 1;
Such a VF therefore gets has_ctrl == 0 and reaches the lookup. With ARI,
for example, a VF at devfn 0x08 decodes to slot 1, function 0, so
pci_get_domain_bus_and_slot() returns pdev itself.
device_lock(&mgt->dev) then takes the same dev->mutex that the driver
core already holds across really_probe(). The probe task would block
forever on a lock it owns. Later unbind, VF removal when SR-IOV is
disabled, and anything else that takes that device lock would then hang
too.
VFs can be created through another PF driver and bound to nbl with
driver_override. Should there be an is_virtfn or mgt == pdev check
before taking the lock?
[Severity: High]
Is there a lock ordering inversion here with the PCI core?
During funcN's probe the driver core holds device_lock(funcN), and this
code then takes device_lock(func0). That is the order N -> 0.
__pci_bus_lock() in drivers/pci/pci.c takes them in the opposite order.
It walks bus->devices in devfn order with the blocking pci_dev_lock():
if (bridge)
pci_dev_lock(bridge);
list_for_each_entry(dev, &bus->devices, bus_list) {
...
else
pci_dev_lock(dev);
}
AER recovery reaches that blocking path:
aer_root_reset()
pci_bus_error_reset()
pci_reset_bridge(bridge, PCI_RESET_NO_RESTORE)
pci_bus_reset(bus, PCI_RESET_DO_RESET)
pci_bus_lock()
pci_slot_reset()->pci_slot_lock() reaches it as well.
Suppose the AER thread holds func0's lock and waits for funcN's. At the
same time funcN's probe, for example a deferred probe triggered when
func0 binds, holds funcN's lock and waits here for func0's. Wouldn't
both sides block forever?
dev->mutex uses lockdep_set_novalidate_class(), so lockdep would not
report this.
> + if (mgt->dev.driver != &nbl_driver.driver ||
> + mgt->dev.links.status != DL_DEV_DRIVER_BOUND) {
> + device_unlock(&mgt->dev);
> + pci_dev_put(mgt);
> + return -EPROBE_DEFER;
> + }
> + device_unlock(&mgt->dev);
> +
> + link = device_link_add(&pdev->dev, &mgt->dev,
> + DL_FLAG_AUTOPROBE_CONSUMER);
[ ... ]
> + device_lock(&mgt->dev);
> + if (mgt->dev.links.status != DL_DEV_DRIVER_BOUND) {
> + device_unlock(&mgt->dev);
> + pci_dev_put(mgt);
> + return -EPROBE_DEFER;
> + }
[Severity: Medium]
Does this second check close the race described in the comment above
it?
device_lock(func0) is dropped before device_link_add(). If func0
unbinds in that gap, device_link_init_status() creates the managed link
in DL_STATE_DORMANT. If func0 then rebinds before this check,
device_links_driver_bound() on the supplier moves the link to
DL_STATE_AVAILABLE, not DL_STATE_CONSUMER_PROBE.
This check reads only mgt->dev.links.status == DL_DEV_DRIVER_BOUND, so
it accepts that state. device_links_busy() does not treat AVAILABLE as
busy, so another func0 unbind can run during the sibling's probe. When
the consumer finishes binding, device_links_driver_bound() hits:
WARN_ON(link->status != DL_STATE_CONSUMER_PROBE);
In a related interleaving func0 is rebound to pci-stub or vfio-pci. This
check does not re-test mgt->dev.driver against &nbl_driver.driver, so
the sibling carries on against a PF0 that never ran nbl's chip-global
setup.
Should this check also verify the bound driver, as the first check
does?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v31 net-next 2/8] net/nebula-matrix: add common resource implementation
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
2026-10-10 9:59 ` [PATCH v31 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
2026-10-10 9:59 ` [PATCH v31 net-next 3/8] net/nebula-matrix: add intr " illusion.wang
` (5 subsequent siblings)
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Add chip-agnostic resource layer to manage PF topology, SR-IOV BDF, Ethernet
port and VSI identity mappings, providing reusable conversion helpers for
the Nebula Matrix driver control plane.
This patch implements core resource initialization and lookup logic,
running exclusively on the control PF and building read-only resource
tables once during probe:
- nbl_res_init_pf_num(): parse firmware PF mask and validate supported
topologies. Only contiguous 1/2/4 PFs starting from PF0 are permitted,
rejecting invalid/sparse configurations with early probe failure.
- nbl_common_func_id_to_rel_pf_id(): convert absolute PF function ID
to relative PF index for consistent resource table indexing.
- nbl_res_ctrl_dev_sriov_info_init(): calculate and store per-PF BDF
entries based on hardware PCI bus number for MSI-X programming.
- nbl_res_ctrl_dev_setup_eth_info(): verify firmware port count and
Ethernet bitmap consistency, then construct per-PF eth_id lookup table.
logic_eth_id is computed on-the-fly from relative PF ID instead of
table lookup.
- nbl_res_ctrl_dev_vsi_info_init(): assign staggered per-PF VSI base
IDs with gaps of 1024, 512 and 256 for 1, 2 and 4 port modes.
- Implement VSI/PF/Eth ID conversion helpers with strict control PF
validity and range guards.
All resource tables are initialized once at control PF probe time and
remain read-only afterwards, so lookup helpers require no internal locking.
Upper-layer serialization via mutex and non-control PF mailbox RPC routing
will be introduced in later series patches.
All resource memory allocations use devm managed semantics, requiring
no explicit cleanup in the remove path.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../net/ethernet/nebula-matrix/nbl/Makefile | 2 +
.../nebula-matrix/nbl/nbl_common/nbl_common.c | 21 ++
.../net/ethernet/nebula-matrix/nbl/nbl_core.h | 2 +
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c | 80 ++++-
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h | 15 +
.../nbl_hw_leonis/nbl_resource_leonis.c | 329 ++++++++++++++++++
.../nbl_hw_leonis/nbl_resource_leonis.h | 10 +
.../nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h | 1 +
.../nebula-matrix/nbl/nbl_hw/nbl_resource.c | 125 +++++++
.../nebula-matrix/nbl/nbl_hw/nbl_resource.h | 69 ++++
.../nbl/nbl_include/nbl_def_channel.h | 10 +
.../nbl/nbl_include/nbl_def_common.h | 21 +-
.../nbl/nbl_include/nbl_def_hw.h | 26 ++
.../nbl/nbl_include/nbl_def_resource.h | 29 ++
.../nbl/nbl_include/nbl_include.h | 6 +
.../net/ethernet/nebula-matrix/nbl/nbl_main.c | 9 +
16 files changed, 752 insertions(+), 3 deletions(-)
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index 04e1aa1fb4bd..3dab9519a277 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -6,4 +6,6 @@ obj-$(CONFIG_NBL) := nbl.o
nbl-objs += nbl_common/nbl_common.o \
nbl_channel/nbl_channel.o \
nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
+ nbl_hw/nbl_hw_leonis/nbl_resource_leonis.o \
+ nbl_hw/nbl_resource.o \
nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
index ce0ed2869af4..302c3c48dc4d 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_common/nbl_common.c
@@ -29,3 +29,24 @@ int nbl_common_create_wq(struct nbl_common_info *common)
return 0;
}
+/**
+ * nbl_common_func_id_to_rel_pf_id - convert absolute PF id to relative PF id
+ * @common: common device info
+ * @pf_id: absolute PF identifier
+ * @rel_pf_id: output relative pf id
+ *
+ * Leonis uses fixed mgt_pf = 0. Support future non-zero management PF.
+ *
+ * Return: 0 on success, -EINVAL on invalid arguments.
+ */
+int nbl_common_func_id_to_rel_pf_id(struct nbl_common_info *common, u32 pf_id,
+ u32 *rel_pf_id)
+{
+ if (!rel_pf_id)
+ return -EINVAL;
+
+ if (pf_id < common->mgt_pf)
+ return -EINVAL;
+ *rel_pf_id = pf_id - common->mgt_pf;
+ return 0;
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index f998a2b44e5c..dd24ebec0171 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -16,11 +16,13 @@ enum {
struct nbl_interface {
struct nbl_hw_ops_tbl *hw_ops_tbl;
+ struct nbl_resource_ops_tbl *resource_ops_tbl;
struct nbl_channel_ops_tbl *channel_ops_tbl;
};
struct nbl_core {
struct nbl_hw_mgt *hw_mgt;
+ struct nbl_resource_mgt *res_mgt;
struct nbl_channel_mgt *chan_mgt;
};
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index eeff6216e4aa..9db7a2adbcc9 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -97,6 +97,45 @@ static void nbl_hw_rd_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
spin_unlock(&hw_mgt->reg_lock);
}
+/*
+ * Flush posted MEMORY-BAR writes by reading back a register that only
+ * exists in the control-PF mapping.
+ *
+ * The mapped size does not derive from bar_len: nbl_hw_init_leonis()
+ * maps PAGE_ALIGN_DOWN(NBL_MEM_BAR_TOTAL_SIZE - NBL_RDMA_NOTIFY_LEN)
+ * (~64 MiB) when has_ctrl is set, and NBL_REG_NET_ONLY_LEN (8 KiB) when
+ * it is not. NBL_HW_DUMMY_REG (0x1300904) is only covered by the
+ * control-PF mapping, so this must not be called on a net-only PF.
+ *
+ * Every current caller runs on the control PF: init_module()/
+ * deinit_module() are has_ctrl-guarded in the resource layer, and the
+ * MSI-X paths are either has_ctrl-guarded as well or dispatched to the
+ * management PF. The check below is therefore a defensive guard against
+ * a future caller, not a reachable failure path.
+ */
+static void nbl_flush_writes(struct nbl_hw_mgt *hw_mgt)
+{
+ if (WARN_ON_ONCE(!hw_mgt->common->has_ctrl))
+ return;
+
+ nbl_hw_rd32(hw_mgt, NBL_HW_DUMMY_REG);
+}
+
+/*
+ * Registers reset to zero after cold boot / FLR / bus reset. Firmware
+ * programs valid values before driver probe, so zero is only seen on
+ * hardware fault or register read failure. Initialize data=0 to guard
+ * against nbl_hw_read_mbx_regs() early-return on bounds-check failure.
+ */
+static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)
+{
+ u32 data = 0;
+
+ nbl_hw_read_mbx_regs(hw_mgt, NBL_FW_BOARD_DW6_OFFSET, &data,
+ sizeof(data));
+ *eth_map = FIELD_GET(NBL_FW_BOARD_DW6_ETH_BITMAP_MASK, data);
+}
+
static void nbl_hw_update_mailbox_queue_tail_ptr(struct nbl_hw_mgt *hw_mgt,
u16 tail_ptr, u8 txrx)
{
@@ -181,6 +220,15 @@ static void nbl_hw_get_host_pf_mask(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask)
sizeof(*pf_mask));
}
+static void nbl_hw_get_real_bus(struct nbl_hw_mgt *hw_mgt, u8 *bus)
+{
+ u32 data = 0;
+
+ nbl_hw_rd_regs_lock(hw_mgt, NBL_PCIE_HOST_TL_CFG_BUSDEV, &data,
+ sizeof(data));
+ *bus = FIELD_GET(NBL_PCIE_BUS_MASK, data);
+}
+
static void nbl_hw_cfg_mailbox_qinfo(struct nbl_hw_mgt *hw_mgt, u16 func_id,
u8 bus, u8 devid, u8 function)
{
@@ -202,15 +250,41 @@ static void nbl_hw_cfg_mailbox_qinfo(struct nbl_hw_mgt *hw_mgt, u16 func_id,
spin_unlock(&hw_mgt->reg_lock);
}
+/*
+ * Registers reset to zero after cold boot / FLR / bus reset. Firmware
+ * programs valid values before driver probe, so zero is only seen on
+ * hardware fault or register read failure. Initialize data=0 to guard
+ * against nbl_hw_read_mbx_regs() early-return on bounds-check failure.
+ */
+static void nbl_hw_get_board_info(struct nbl_hw_mgt *hw_mgt,
+ struct nbl_board_port_info *board_info)
+{
+ u32 data = 0;
+
+ nbl_hw_read_mbx_regs(hw_mgt, NBL_FW_BOARD_DW3_OFFSET, &data,
+ sizeof(data));
+ board_info->eth_num = FIELD_GET(NBL_FW_BOARD_DW3_PORT_NUM_MASK, data);
+ board_info->eth_speed =
+ FIELD_GET(NBL_FW_BOARD_DW3_PORT_SPEED_MASK, data);
+ board_info->p4_version =
+ FIELD_GET(NBL_FW_BOARD_DW3_P4_VERSION_MASK, data);
+}
+
static struct nbl_hw_ops hw_ops = {
+ .flush_write = nbl_flush_writes,
+
.update_mailbox_queue_tail_ptr = nbl_hw_update_mailbox_queue_tail_ptr,
.config_mailbox_rxq = nbl_hw_config_mailbox_rxq,
.config_mailbox_txq = nbl_hw_config_mailbox_txq,
.stop_mailbox_rxq = nbl_hw_stop_mailbox_rxq,
.stop_mailbox_txq = nbl_hw_stop_mailbox_txq,
.get_host_pf_mask = nbl_hw_get_host_pf_mask,
+ .get_real_bus = nbl_hw_get_real_bus,
+
.cfg_mailbox_qinfo = nbl_hw_cfg_mailbox_qinfo,
+ .get_fw_eth_map = nbl_hw_get_fw_eth_map,
+ .get_board_info = nbl_hw_get_board_info,
};
/* Structure starts here, adding an op should not modify anything below */
@@ -238,10 +312,12 @@ static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
hw_ops_tbl = devm_kzalloc(dev, sizeof(*hw_ops_tbl), GFP_KERNEL);
if (!hw_ops_tbl)
return ERR_PTR(-ENOMEM);
- if (!hw_ops.update_mailbox_queue_tail_ptr ||
+ if (!hw_ops.flush_write || !hw_ops.update_mailbox_queue_tail_ptr ||
!hw_ops.config_mailbox_rxq || !hw_ops.config_mailbox_txq ||
!hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
- !hw_ops.get_host_pf_mask || !hw_ops.cfg_mailbox_qinfo)
+ !hw_ops.get_host_pf_mask || !hw_ops.get_real_bus ||
+ !hw_ops.cfg_mailbox_qinfo ||
+ !hw_ops.get_fw_eth_map || !hw_ops.get_board_info)
return ERR_PTR(-EINVAL);
hw_ops_tbl->ops = &hw_ops;
hw_ops_tbl->priv = hw_mgt;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index 93e5c7518288..251dd68d0721 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -51,6 +51,21 @@ struct nbl_mailbox_qinfo_cfg_table {
/* -------- HOST_PCIE -------- */
#define NBL_PCIE_HOST_K_PF_MASK_REG (NBL_INTF_HOST_PCIE_BASE + 0x00001004)
+#define NBL_PCIE_HOST_TL_CFG_BUSDEV (NBL_INTF_HOST_PCIE_BASE + 0x11040)
+
+#define NBL_PCIE_BUS_MASK GENMASK(12, 5)
+#define NBL_FW_BOARD_CONFIG 0x200
+#define NBL_FW_BOARD_DW3_OFFSET (NBL_FW_BOARD_CONFIG + 12)
+#define NBL_FW_BOARD_DW6_OFFSET (NBL_FW_BOARD_CONFIG + 24)
+
+#define NBL_FW_BOARD_DW3_PORT_TYPE_MASK BIT(0)
+#define NBL_FW_BOARD_DW3_PORT_NUM_MASK GENMASK(7, 1)
+#define NBL_FW_BOARD_DW3_PORT_SPEED_MASK GENMASK(9, 8)
+#define NBL_FW_BOARD_DW3_GPIO_TYPE_MASK GENMASK(12, 10)
+#define NBL_FW_BOARD_DW3_P4_VERSION_MASK GENMASK(13, 13)
+
+#define NBL_FW_BOARD_DW6_LANE_BITMAP_MASK GENMASK(7, 0)
+#define NBL_FW_BOARD_DW6_ETH_BITMAP_MASK GENMASK(15, 8)
#define NBL_BAR2_MAX_LEN 0x300
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
new file mode 100644
index 000000000000..7804762a96e0
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
@@ -0,0 +1,329 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/pci.h>
+#include <linux/bits.h>
+#include "nbl_resource_leonis.h"
+
+static struct nbl_resource_ops res_ops = {
+ .get_vsi_id = nbl_res_func_id_to_vsi_id,
+ .get_eth_id = nbl_res_get_eth_id,
+};
+
+static struct nbl_resource_mgt *
+nbl_res_setup_res_mgt(struct nbl_common_info *common)
+{
+ struct nbl_resource_info *resource_info;
+ struct nbl_resource_mgt *res_mgt;
+ struct device *dev = common->dev;
+
+ res_mgt = devm_kzalloc(dev, sizeof(*res_mgt), GFP_KERNEL);
+ if (!res_mgt)
+ return ERR_PTR(-ENOMEM);
+ res_mgt->common = common;
+
+ resource_info =
+ devm_kzalloc(dev, sizeof(*resource_info), GFP_KERNEL);
+ if (!resource_info)
+ return ERR_PTR(-ENOMEM);
+ res_mgt->resource_info = resource_info;
+
+ return res_mgt;
+}
+
+static struct nbl_resource_ops_tbl *
+nbl_res_setup_ops(struct device *dev, struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_resource_ops_tbl *res_ops_tbl;
+
+ res_ops_tbl = devm_kzalloc(dev, sizeof(*res_ops_tbl), GFP_KERNEL);
+ if (!res_ops_tbl)
+ return ERR_PTR(-ENOMEM);
+ if (!res_ops.get_vsi_id || !res_ops.get_eth_id)
+ return ERR_PTR(-EINVAL);
+ res_ops_tbl->ops = &res_ops;
+ res_ops_tbl->priv = res_mgt;
+
+ return res_ops_tbl;
+}
+
+static int nbl_res_ctrl_dev_setup_eth_info(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ struct device *dev = res_mgt->common->dev;
+ struct nbl_eth_info *eth_info;
+ u32 eth_bitmap = 0;
+ u32 eth_num = 0;
+ u32 fw_port_num;
+ int i;
+
+ eth_info = devm_kzalloc(dev, sizeof(*eth_info), GFP_KERNEL);
+ if (!eth_info)
+ return -ENOMEM;
+
+ res_mgt->resource_info->eth_info = eth_info;
+
+ fw_port_num = res_mgt->resource_info->board_info.eth_num;
+
+ hw_ops->get_fw_eth_map(res_mgt->hw_ops_tbl->priv, ð_bitmap);
+ if (eth_bitmap & ~((1 << NBL_MAX_ETHERNET) - 1)) {
+ dev_err(dev, "FW reported invalid eth_bitmap 0x%x\n",
+ eth_bitmap);
+ return -EINVAL;
+ }
+ if (fw_port_num != hweight32(eth_bitmap)) {
+ dev_err(dev, "FW inconsistency: port_num=%u, bitmap=0x%x\n",
+ fw_port_num, eth_bitmap);
+ return -EINVAL;
+ }
+ /*
+ * Firmware is ready before probe. Valid port counts are 1/2/4;
+ * 0 (invalid config), 3 (unsupported topology), and >4 (exceeds
+ * hardware max) are all rejected with -EINVAL.
+ */
+ if (fw_port_num == 0 || fw_port_num == 3 ||
+ fw_port_num > NBL_MAX_ETHERNET) {
+ dev_err(dev, "FW reports %u Ethernet ports, unsupported (valid: 1/2/4)\n",
+ fw_port_num);
+ return -EINVAL;
+ }
+ eth_info->eth_num = fw_port_num;
+ /* Intentional design constraint: each PF maps to exactly one
+ * Ethernet port. This couples PF identity to port identity
+ * and is required by nbl_res_get_eth_id() which indexes
+ * eth_info->eth_id[] by relative PF id.
+ */
+ if (res_mgt->common->max_pf != eth_info->eth_num) {
+ dev_err(dev, "Invalid PF-to-port topology: max_pf=%u, eth_num=%u\n",
+ res_mgt->common->max_pf, eth_info->eth_num);
+ return -EINVAL;
+ }
+
+ /*
+ * Any subset of valid bitmap bits is accepted (e.g. 0/1, 0/2,
+ * 1/3, etc.). Firmware only needs to report the correct count
+ * of active ports; no hard-coded fixed bit positions required.
+ * eth_id[] is filled in ascending bitmap-bit order, so the Nth
+ * relative PF owns the Nth logical port (max_pf == eth_num is
+ * enforced above).
+ */
+ for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+ if ((1 << i) & eth_bitmap) {
+ eth_info->eth_id[eth_num] = i;
+ eth_num++;
+ }
+ }
+
+ return 0;
+}
+
+static int nbl_res_ctrl_dev_sriov_info_init(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ struct nbl_hw_mgt *p = res_mgt->hw_ops_tbl->priv;
+ struct nbl_common_info *common = res_mgt->common;
+ struct nbl_sriov_info *sriov_info;
+ struct device *dev = common->dev;
+ u8 hw_bus = 0;
+ u16 function;
+ u16 func_id;
+
+ hw_ops->get_real_bus(p, &hw_bus);
+ if (common->function + common->max_pf > NBL_MAX_PF) {
+ dev_err(dev, "PF count exceeds available function space\n");
+ return -EINVAL;
+ }
+ sriov_info = devm_kcalloc(dev, common->max_pf,
+ sizeof(*sriov_info), GFP_KERNEL);
+ if (!sriov_info)
+ return -ENOMEM;
+
+ res_mgt->resource_info->sriov_info = sriov_info;
+ /*
+ * Real bus number of the control PF;
+ * the BDF table below is built from it.
+ */
+ common->hw_bus = hw_bus;
+
+ for (func_id = 0; func_id < common->max_pf; func_id++) {
+ sriov_info = res_mgt->resource_info->sriov_info + func_id;
+ function = common->function + func_id;
+ sriov_info->bdf = PCI_DEVID(common->hw_bus,
+ PCI_DEVFN(common->devid, function));
+ }
+
+ return 0;
+}
+
+static int nbl_res_ctrl_dev_vsi_info_init(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_eth_info *eth_info = res_mgt->resource_info->eth_info;
+ struct nbl_common_info *common = res_mgt->common;
+ struct device *dev = common->dev;
+ struct nbl_vsi_info *vsi_info;
+ int i;
+
+ vsi_info = devm_kzalloc(dev, sizeof(*vsi_info), GFP_KERNEL);
+ if (!vsi_info)
+ return -ENOMEM;
+
+ res_mgt->resource_info->vsi_info = vsi_info;
+ /*
+ * case 1 one port(1pf)
+ * pf0 (NBL_VSI_SERV_PF_DATA_TYPE) vsi is 0
+ * case 2 two port(2pf)
+ * pf0,pf1(NBL_VSI_SERV_PF_DATA_TYPE) vsi is 0,512
+ * case 3 four port(4pf)
+ * pf0,pf1,pf2,pf3(NBL_VSI_SERV_PF_DATA_TYPE) vsi is 0,256,512,768
+ */
+
+ vsi_info->num = eth_info->eth_num;
+ /*
+ * eth_num can be 1/2/4:
+ * - 2/4 ports use dedicated gap constants;
+ * - 1 port falls back to NBL_DEFAULT_VSI_ID_GAP (1024).
+ * All three values produce valid base_id offsets.
+ */
+ for (i = 0; i < vsi_info->num; i++) {
+ vsi_info->serv_info[i][NBL_VSI_SERV_PF_DATA_TYPE].base_id =
+ i * nbl_vsi_id_gap(vsi_info->num);
+ vsi_info->serv_info[i][NBL_VSI_SERV_PF_DATA_TYPE].num = 1;
+ }
+
+ return 0;
+}
+
+static int nbl_res_init_pf_num(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ u32 exp_contiguous_mask = 0;
+ u32 pf_mask = 0;
+ u32 pf_num = 0;
+ int i;
+
+ hw_ops->get_host_pf_mask(res_mgt->hw_ops_tbl->priv, &pf_mask);
+
+ /*
+ * k_pf_mask register rule:
+ * bit N == 0 -> PF#N enabled; bit N == 1 -> PF#N masked out.
+ * Hardware constraint: bit0 is PF0's mask bit; driver requires
+ * PF0 enabled as management PF, so bit0 must be clear.
+ * All-zero pf_mask means all PF0~PF7 are enabled, which is unsupported
+ * by the driver
+ *
+ * Product firmware constraint: only 3 valid configurations supported:
+ * 1 PF (PF0 only): pf_num = 1, mask = 0xfe
+ * 2 PFs (PF0,PF1): pf_num = 2, mask = 0xfc
+ * 4 PFs (PF0~PF3): pf_num = 4, mask = 0xf0
+ * No other PF count or sparse/non-contiguous PF layout is allowed.
+ */
+ for (i = 0; i < NBL_MAX_PF; i++) {
+ if (!(pf_mask & (1 << i)))
+ pf_num++;
+ }
+
+ /*
+ * Sanity check: enabled PFs must be contiguous starting from PF0.
+ * Current resource framework uses relative PF id, sparse PF layout
+ * will cause mismatch between resource layer and hardware func_id.
+ */
+ for (i = 0; i < pf_num; i++)
+ exp_contiguous_mask |= BIT(i);
+ if ((pf_mask & exp_contiguous_mask) != 0) {
+ dev_err(res_mgt->common->dev,
+ "pf_mask 0x%08x: non-contiguous enabled PF, unsupported\n",
+ pf_mask);
+ return -EINVAL;
+ }
+
+ /* Only allow product-specified PF count: 1 / 2 / 4 */
+ if (pf_num != 1 && pf_num != 2 && pf_num != 4) {
+ dev_err(res_mgt->common->dev,
+ "Invalid pf_num=%u (mask=0x%08x), only 1/2/4 PFs supported\n",
+ pf_num, pf_mask);
+ return -EINVAL;
+ }
+
+ res_mgt->common->max_pf = pf_num;
+
+ return 0;
+}
+
+static void nbl_res_init_board_info(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+
+ hw_ops->get_board_info(res_mgt->hw_ops_tbl->priv,
+ &res_mgt->resource_info->board_info);
+}
+
+static int nbl_res_start(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_common_info *common = res_mgt->common;
+ int ret = 0;
+
+ if (common->has_ctrl) {
+ nbl_res_init_board_info(res_mgt);
+
+ ret = nbl_res_init_pf_num(res_mgt);
+ if (ret)
+ return ret;
+
+ ret = nbl_res_ctrl_dev_sriov_info_init(res_mgt);
+ if (ret)
+ return ret;
+
+ ret = nbl_res_ctrl_dev_setup_eth_info(res_mgt);
+ if (ret)
+ return ret;
+
+ ret = nbl_res_ctrl_dev_vsi_info_init(res_mgt);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
+
+int nbl_res_init_leonis(struct nbl_adapter *adap)
+{
+ struct nbl_channel_ops_tbl *chan_ops_tbl = adap->intf.channel_ops_tbl;
+ struct nbl_hw_ops_tbl *hw_ops_tbl = adap->intf.hw_ops_tbl;
+ struct nbl_common_info *common = &adap->common;
+ struct nbl_resource_ops_tbl *res_ops_tbl;
+ struct device *dev = &adap->pdev->dev;
+ struct nbl_resource_mgt *res_mgt;
+ int ret;
+
+ res_mgt = nbl_res_setup_res_mgt(common);
+ if (IS_ERR(res_mgt)) {
+ ret = PTR_ERR(res_mgt);
+ return ret;
+ }
+ res_mgt->chan_ops_tbl = chan_ops_tbl;
+ res_mgt->hw_ops_tbl = hw_ops_tbl;
+
+ ret = nbl_res_start(res_mgt);
+ if (ret)
+ return ret;
+
+ res_ops_tbl = nbl_res_setup_ops(dev, res_mgt);
+ if (IS_ERR(res_ops_tbl)) {
+ ret = PTR_ERR(res_ops_tbl);
+ return ret;
+ }
+ adap->intf.resource_ops_tbl = res_ops_tbl;
+ adap->core.res_mgt = res_mgt;
+
+ return 0;
+}
+
+void nbl_res_remove_leonis(struct nbl_adapter *adap)
+{
+ /*
+ * No resource release here because all memory uses devm managed
+ * allocation
+ */
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
new file mode 100644
index 000000000000..b9355262c00d
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
@@ -0,0 +1,10 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_RESOURCE_LEONIS_H_
+#define _NBL_RESOURCE_LEONIS_H_
+
+#include "../nbl_resource.h"
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
index 4c1bb789c465..00fa33eeabaa 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_reg.h
@@ -17,6 +17,7 @@
#define NBL_MAILBOX_BAR 2
#define NBL_RDMA_NOTIFY_LEN (8ULL << 10)
#define NBL_REG_NET_ONLY_LEN (8ULL << 10)
+#define NBL_HW_DUMMY_REG 0x1300904
/*
* PCI MEMORY BAR total size: 64MiB.
*/
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
new file mode 100644
index 000000000000..b316fb8e7051
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
@@ -0,0 +1,125 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#include <linux/pci.h>
+#include "nbl_resource.h"
+
+int nbl_res_func_id_to_vsi_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 type, u16 *vsi_id)
+{
+ struct nbl_vsi_info *vsi_info = res_mgt->resource_info->vsi_info;
+ enum nbl_vsi_serv_type dst_type = NBL_VSI_SERV_PF_DATA_TYPE;
+ struct nbl_common_info *common = res_mgt->common;
+ struct device *dev = res_mgt->common->dev;
+ int pfid = func_id;
+ u32 rel_pf_id;
+ int ret;
+
+ if (!common->has_ctrl || !vsi_id) {
+ dev_dbg(dev, "No control plane or null vsi output ptr\n");
+ return -EINVAL;
+ }
+ ret = nbl_common_func_id_to_rel_pf_id(common, pfid, &rel_pf_id);
+ if (ret)
+ return ret;
+ if (rel_pf_id >= vsi_info->num) {
+ dev_err(dev, "PF %d (diff=%u) exceeds vsi_info->num (%u)\n",
+ pfid, rel_pf_id, vsi_info->num);
+ return -EINVAL;
+ }
+
+ ret = nbl_res_pf_dev_vsi_type_to_hw_vsi_type(res_mgt, type, &dst_type);
+ if (ret) {
+ dev_err(dev, "Invalid vsi type %u func_id %u\n", type, func_id);
+ return ret;
+ }
+ *vsi_id = vsi_info->serv_info[rel_pf_id][dst_type].base_id;
+ return 0;
+}
+
+int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id)
+{
+ struct nbl_vsi_info *vsi_info = res_mgt->resource_info->vsi_info;
+ struct nbl_common_info *common = res_mgt->common;
+ struct device *dev = res_mgt->common->dev;
+ int j = NBL_VSI_SERV_PF_DATA_TYPE;
+ int pf_id, i;
+
+ if (!common->has_ctrl) {
+ dev_dbg(dev, "No control plane available\n");
+ return -EINVAL;
+ }
+ for (i = 0; i < vsi_info->num; i++) {
+ if (vsi_id >= vsi_info->serv_info[i][j].base_id &&
+ (vsi_id < vsi_info->serv_info[i][j].base_id +
+ vsi_info->serv_info[i][j].num)) {
+ pf_id = i + common->mgt_pf;
+ if (pf_id >= NBL_MAX_PF) {
+ dev_err(dev, "PF ID overflow\n");
+ return -ERANGE;
+ }
+ return pf_id;
+ }
+ }
+
+ dev_dbg(dev, "VSI ID %u not found\n", vsi_id);
+ return -ENOENT;
+}
+
+int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id)
+{
+ struct nbl_eth_info *eth_info = res_mgt->resource_info->eth_info;
+ struct nbl_common_info *common = res_mgt->common;
+ struct device *dev = res_mgt->common->dev;
+ int pfid = func_id;
+ int rel_pf_id;
+ int abs_pf_id;
+
+ if (!common->has_ctrl || !eth_num || !eth_id || !logic_eth_id)
+ return -EINVAL;
+ abs_pf_id = nbl_res_vsi_id_to_pf_id(res_mgt, vsi_id);
+ if (abs_pf_id < 0) {
+ dev_err(dev, "Failed to get PF ID from VSI ID %u\n", vsi_id);
+ return -EINVAL;
+ }
+ if (abs_pf_id != pfid) {
+ dev_err(dev, "func_id %u does not match pf derived from vsi_id %u\n",
+ pfid, vsi_id);
+ return -EINVAL;
+ }
+ rel_pf_id = abs_pf_id - common->mgt_pf;
+
+ if (rel_pf_id >= eth_info->eth_num) {
+ dev_err(dev, "rel_pf_id %d out of range [0, %u)\n",
+ rel_pf_id, eth_info->eth_num);
+ return -ERANGE;
+ }
+
+ *eth_num = eth_info->eth_num;
+ *eth_id = eth_info->eth_id[rel_pf_id];
+ /*
+ * Logical eth id equals the relative PF id: eth_info setup enforces
+ * max_pf == eth_num and fills eth_id[] in ascending bitmap-bit
+ * order, so the Nth PF owns the Nth logical port.
+ */
+ *logic_eth_id = rel_pf_id;
+ return 0;
+}
+
+int nbl_res_pf_dev_vsi_type_to_hw_vsi_type(struct nbl_resource_mgt *res_mgt,
+ u16 src_type,
+ enum nbl_vsi_serv_type *dst_type)
+{
+ switch (src_type) {
+ case NBL_VSI_DATA:
+ *dst_type = NBL_VSI_SERV_PF_DATA_TYPE;
+ return 0;
+ default:
+ dev_err_once(res_mgt->common->dev,
+ "Unsupported vsi src_type %u\n", src_type);
+ return -EINVAL;
+ }
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
new file mode 100644
index 000000000000..ae0a3d33198d
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
@@ -0,0 +1,69 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_RESOURCE_H_
+#define _NBL_RESOURCE_H_
+
+#include <linux/types.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_hw.h"
+#include "../nbl_include/nbl_def_resource.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+struct nbl_resource_mgt;
+
+/* --------- INFO ---------- */
+struct nbl_sriov_info {
+ unsigned int bdf;
+};
+
+struct nbl_eth_info {
+ u8 eth_num;
+ u8 resv[3];
+ u8 eth_id[NBL_MAX_ETHERNET];
+};
+
+enum nbl_vsi_serv_type {
+ NBL_VSI_SERV_PF_DATA_TYPE,
+ NBL_VSI_SERV_MAX_TYPE,
+};
+
+struct nbl_vsi_serv_info {
+ u16 base_id;
+ u16 num;
+};
+
+struct nbl_vsi_info {
+ u16 num;
+ struct nbl_vsi_serv_info serv_info[NBL_MAX_ETHERNET]
+ [NBL_VSI_SERV_MAX_TYPE];
+};
+
+struct nbl_resource_info {
+ struct nbl_sriov_info *sriov_info;
+ struct nbl_eth_info *eth_info;
+ struct nbl_vsi_info *vsi_info;
+ struct nbl_board_port_info board_info;
+};
+
+struct nbl_resource_mgt {
+ struct nbl_common_info *common;
+ struct nbl_resource_info *resource_info;
+ struct nbl_channel_ops_tbl *chan_ops_tbl;
+ struct nbl_hw_ops_tbl *hw_ops_tbl;
+};
+
+int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id);
+int nbl_res_func_id_to_vsi_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 type, u16 *vsi_id);
+int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id);
+int nbl_res_pf_dev_vsi_type_to_hw_vsi_type(struct nbl_resource_mgt *res_mgt,
+ u16 src_type,
+ enum nbl_vsi_serv_type *dst_type);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
index d6faa0bc4026..bf971121d2ec 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
@@ -39,6 +39,16 @@ enum nbl_chan_state {
NBL_CHAN_STATE_NBITS
};
+struct nbl_board_port_info {
+ u8 eth_num;
+ u8 eth_speed;
+ u8 p4_version;
+ u8 rsv[5];
+};
+
+static_assert(sizeof(struct nbl_board_port_info) == 8,
+ "nbl_board_port_info size must be 8 bytes");
+
struct nbl_chan_send_info {
void *arg;
size_t arg_len;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
index 26a364e3a194..3e18a2a97850 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_common.h
@@ -11,6 +11,22 @@
#include <linux/device.h>
#include "nbl_include.h"
+#define NBL_TWO_ETHERNET_PORT 2
+#define NBL_FOUR_ETHERNET_PORT 4
+#define NBL_DEFAULT_VSI_ID_GAP 1024
+#define NBL_TWO_ETHERNET_VSI_ID_GAP 512
+#define NBL_FOUR_ETHERNET_VSI_ID_GAP 256
+
+static inline u32 nbl_vsi_id_gap(u32 m)
+{
+ if (m == NBL_FOUR_ETHERNET_PORT)
+ return NBL_FOUR_ETHERNET_VSI_ID_GAP;
+ else if (m == NBL_TWO_ETHERNET_PORT)
+ return NBL_TWO_ETHERNET_VSI_ID_GAP;
+
+ return NBL_DEFAULT_VSI_ID_GAP;
+}
+
struct nbl_common_info {
struct workqueue_struct *wq;
struct pci_dev *pdev;
@@ -24,12 +40,15 @@ struct nbl_common_info {
u8 devid;
u8 bus;
u8 hw_bus;
+ u16 mgt_pf;
u8 has_ctrl;
u8 has_net;
+ u8 max_pf;
};
void nbl_common_destroy_wq(struct nbl_common_info *common);
int nbl_common_create_wq(struct nbl_common_info *common);
-
+int nbl_common_func_id_to_rel_pf_id(struct nbl_common_info *common, u32 pf_id,
+ u32 *rel_pf_id);
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index 58d1b669b8cb..4b420db91963 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -8,9 +8,22 @@
#include <linux/types.h>
+struct nbl_board_port_info;
struct nbl_hw_mgt;
struct nbl_adapter;
struct nbl_hw_ops {
+ /**
+ * flush_write - Flush posted MEMORY-BAR writes
+ * @hw_mgt: hardware management context
+ *
+ * Only valid on the control PF (common->has_ctrl != 0). The
+ * register read back to order the writes, NBL_HW_DUMMY_REG
+ * (0x1300904), is covered by the ~64 MiB MEMORY-BAR mapping the
+ * control PF gets, but sits outside the NBL_REG_NET_ONLY_LEN
+ * (8 KiB) mapping a net-only PF gets. The leonis implementation
+ * enforces this with a WARN_ON_ONCE().
+ */
+ void (*flush_write)(struct nbl_hw_mgt *hw_mgt);
void (*update_mailbox_queue_tail_ptr)(struct nbl_hw_mgt *hw_mgt,
u16 tail_ptr, u8 txrx);
void (*config_mailbox_rxq)(struct nbl_hw_mgt *hw_mgt,
@@ -42,9 +55,22 @@ struct nbl_hw_ops {
* op only returns the raw register value.
*/
void (*get_host_pf_mask)(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask);
+ void (*get_real_bus)(struct nbl_hw_mgt *hw_mgt, u8 *bus);
void (*cfg_mailbox_qinfo)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
u8 bus, u8 devid, u8 function);
+ void (*get_fw_eth_map)(struct nbl_hw_mgt *hw_mgt, u32 *eth_map);
+ /**
+ * get_board_info - Fetch board info from firmware
+ * @hw_mgt: hardware management context
+ * @board_info: output pointer for board info structure
+ *
+ * Firmware contract: board_info.eth_num MUST equal the number of
+ * unmasked PFs from get_host_pf_mask(). See get_host_pf_mask for
+ * details.
+ */
+ void (*get_board_info)(struct nbl_hw_mgt *hw_mgt,
+ struct nbl_board_port_info *board_info);
};
struct nbl_hw_ops_tbl {
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
new file mode 100644
index 000000000000..7136b282fb80
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
@@ -0,0 +1,29 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_RESOURCE_H_
+#define _NBL_DEF_RESOURCE_H_
+
+#include <linux/types.h>
+
+struct nbl_resource_mgt;
+struct nbl_adapter;
+
+struct nbl_resource_ops {
+ int (*get_vsi_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 type, u16 *vsi_id);
+ int (*get_eth_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 vsi_id, u8 *eth_num, u8 *eth_id,
+ u8 *logic_eth_id);
+};
+
+struct nbl_resource_ops_tbl {
+ struct nbl_resource_ops *ops;
+ struct nbl_resource_mgt *priv;
+};
+
+int nbl_res_init_leonis(struct nbl_adapter *adapter);
+void nbl_res_remove_leonis(struct nbl_adapter *adapter);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
index f2d802397d98..59e44feab44f 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
@@ -13,6 +13,12 @@
#define NBL_MAX_PF 8
#define NBL_NEXT_ID(id, max) (((id) + 1) % ((max) + 1))
+#define NBL_MAX_ETHERNET 4
+
+enum {
+ NBL_VSI_DATA = 0,
+};
+
struct nbl_func_caps {
u32 has_ctrl:1;
u32 has_net:1;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index 62484fbf1fef..cd3774691bf4 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -10,6 +10,7 @@
#include "nbl_include/nbl_include.h"
#include "nbl_include/nbl_def_channel.h"
#include "nbl_include/nbl_def_hw.h"
+#include "nbl_include/nbl_def_resource.h"
#include "nbl_include/nbl_def_common.h"
#include "nbl_core.h"
@@ -27,6 +28,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
adapter->pdev = pdev;
common = &adapter->common;
+ common->mgt_pf = 0;
common->pdev = pdev;
common->dev = &pdev->dev;
common->has_ctrl = param->caps.has_ctrl;
@@ -42,7 +44,13 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
ret = nbl_chan_init_common(adapter);
if (ret)
goto chan_init_fail;
+
+ ret = nbl_res_init_leonis(adapter);
+ if (ret)
+ goto res_init_fail;
return adapter;
+res_init_fail:
+ nbl_chan_remove_common(adapter);
chan_init_fail:
nbl_hw_remove_leonis(adapter);
hw_init_fail:
@@ -51,6 +59,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
void nbl_core_remove(struct nbl_adapter *adapter)
{
+ nbl_res_remove_leonis(adapter);
nbl_chan_remove_common(adapter);
nbl_hw_remove_leonis(adapter);
}
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 2/8] net/nebula-matrix: add common resource implementation
2026-10-10 9:59 ` [PATCH v31 net-next 2/8] net/nebula-matrix: add common resource implementation illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] The `get_host_pf_mask` kernel-doc in `nbl_def_hw.h` says
nbl_res_init_pf_num() "always runs before this op's only consumer,…
--
net/nebula-matrix: add common resource implementation
This commit adds a chip-agnostic resource layer for the Nebula Matrix
driver. On the control PF it builds the PF topology, SR-IOV BDF, Ethernet
port and VSI identity tables once at probe time. It also adds new hw ops
(flush_write, get_real_bus, get_fw_eth_map, get_board_info) and the
VSI/PF/Eth ID conversion helpers.
[ ... ]
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
> index 58d1b669b8cbb..4b420db91963d 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
[ ... ]
> @@ -42,9 +55,22 @@ struct nbl_hw_ops {
> * op only returns the raw register value.
> */
> void (*get_host_pf_mask)(struct nbl_hw_mgt *hw_mgt, u32 *pf_mask);
> + void (*get_real_bus)(struct nbl_hw_mgt *hw_mgt, u8 *bus);
[Severity: Low]
This isn't a bug, but the kernel-doc for get_host_pf_mask just above this
hunk now looks stale:
* The mask is validated in the resource layer:
* nbl_res_init_pf_num() rejects an all-zero mask and any
* non-contiguous layout with -EINVAL, and it always runs before
* this op's only consumer, nbl_chan_cfg_qinfo_map_table(). This
* op only returns the raw register value.
This patch adds nbl_res_init_pf_num() in nbl_resource_leonis.c, and that
function calls the op itself:
hw_ops->get_host_pf_mask(res_mgt->hw_ops_tbl->priv, &pf_mask);
That means nbl_chan_cfg_qinfo_map_table() is no longer the only consumer.
The same is still true at the end of the series, where the op is called
from both nbl_resource_leonis.c and nbl_channel.c.
The ordering claim does hold by the end of the series. In "net/nebula-matrix:
add common/ctrl dev init/remove operation", nbl_dev_init() calls
nbl_dev_setup_chan_qinfo() only under common->has_ctrl, and it does so
after nbl_res_init_leonis().
Could the "only consumer" wording be changed to list both callers?
>
> void (*cfg_mailbox_qinfo)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
> u8 bus, u8 devid, u8 function);
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v31 net-next 3/8] net/nebula-matrix: add intr resource implementation
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
2026-10-10 9:59 ` [PATCH v31 net-next 1/8] net/nebula-matrix: add channel layer illusion.wang
2026-10-10 9:59 ` [PATCH v31 net-next 2/8] net/nebula-matrix: add common resource implementation illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
2026-10-10 9:59 ` [PATCH v31 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
` (4 subsequent siblings)
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Add an nbl_interrupt module that manages the driver-wide, chip-internal
MSI-X vector index space and the chip MSI-X mapping tables. It does not
manage the physical PCI MSI-X entries; those are set up in a follow-up
patch by nbl_dev_init_interrupt_scheme().
The global vector space is split into separate net and control bitmaps
(intr_net_bmap/intr_other_bmap). A single intr_mgt->lock protects the
bitmaps and the per-function state and is taken internally by every
public API, so callers need no extra locking. Only PFs are supported;
VF ids are rejected with -EOPNOTSUPP.
The module provides three resource ops:
- cfg_msix_map: allocate global vectors from the two bitmaps and program
NBL_PCOMPLETER_FUNCTION_MSIX_MAP with the table DMA address and the
control PF's BDF. The per-function coherent table is allocated once and
reused on later reconfigurations, which removes the table free/realloc
cycle. The new vector set is allocated before the old hardware state is
torn down, so a failed reconfiguration leaves the old configuration
intact.
- destroy_msix_map: two-stage teardown. Disable mailbox IRQ routing,
invalidate the per-vector INFO entries, clear the map VALID bit while
keeping the live DMA address, wait ~1 ms for in-flight table fetches,
then zero the entry and free the coherent table. The wait is
best-effort, as the chip exposes no fetch completion status.
- set_mailbox_irq: bind or unbind a PF's mailbox routing via
NBL_MAILBOX_QINFO_MAP_REG_ARR. The disable path works without a
configured map, so routing can be cleaned up before vectors are
released.
The hardware helpers program PADPT_HOST_MSIX_INFO and
PCOMPLETER_HOST_MSIX_FID_TABLE in forward order on enable and reverse
order on teardown, to avoid inconsistent hardware state.
The module also adds register definitions and resource ops callbacks,
and instantiates the manager via nbl_intr_mgt_start() during resource
init. nbl_intr_mgt_stop() is the global teardown, called from
nbl_res_remove_leonis().
The new resource ops are hooked into resource_ops but have no in-tree
caller in this patch; invocation is added in later patches of the series.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../net/ethernet/nebula-matrix/nbl/Makefile | 1 +
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c | 164 +++-
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h | 44 +
.../nbl_hw_leonis/nbl_resource_leonis.c | 38 +-
.../nbl_hw_leonis/nbl_resource_leonis.h | 1 +
.../nebula-matrix/nbl/nbl_hw/nbl_interrupt.c | 801 ++++++++++++++++++
.../nebula-matrix/nbl/nbl_hw/nbl_interrupt.h | 21 +
.../nebula-matrix/nbl/nbl_hw/nbl_resource.c | 32 +
.../nebula-matrix/nbl/nbl_hw/nbl_resource.h | 77 ++
.../nbl/nbl_include/nbl_def_hw.h | 9 +
.../nbl/nbl_include/nbl_def_resource.h | 6 +
11 files changed, 1189 insertions(+), 5 deletions(-)
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index 3dab9519a277..5aec8e44f5d7 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -8,4 +8,5 @@ nbl-objs += nbl_common/nbl_common.o \
nbl_hw/nbl_hw_leonis/nbl_hw_leonis.o \
nbl_hw/nbl_hw_leonis/nbl_resource_leonis.o \
nbl_hw/nbl_resource.o \
+ nbl_hw/nbl_interrupt.o \
nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index 9db7a2adbcc9..1712cdbc5fe7 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -97,6 +97,20 @@ static void nbl_hw_rd_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
spin_unlock(&hw_mgt->reg_lock);
}
+static void nbl_hw_wr_regs_lock(struct nbl_hw_mgt *hw_mgt, u64 reg,
+ const u32 *data, u32 len)
+{
+ u32 size = len / 4;
+ u32 i;
+
+ if (len % 4)
+ return;
+ spin_lock(&hw_mgt->reg_lock);
+ for (i = 0; i < size; i++)
+ wr32(hw_mgt->hw_addr, reg + i * sizeof(u32), data[i]);
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
/*
* Flush posted MEMORY-BAR writes by reading back a register that only
* exists in the control-PF mapping.
@@ -136,6 +150,148 @@ static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)
*eth_map = FIELD_GET(NBL_FW_BOARD_DW6_ETH_BITMAP_MASK, data);
}
+/*
+ * nbl_hw_set_mailbox_irq - read-modify-write NBL_MAILBOX_QINFO_MAP_REG_ARR
+ *
+ * The full RMW sequence is wrapped by reg_lock, so concurrent register
+ * access from different CPUs is already serialized safely.
+ * nbl_hw_cfg_mailbox_qinfo() programs the BDF fields during control-PF
+ * init and clears MSIX_IDX/MSIX_IDX_VALID at the same time (they survive
+ * kexec/forced unload without FLR), so mailbox MSIX routing for a PF
+ * starts disarmed at init and is armed only by an explicit en_msix=true
+ * call here.
+ */
+static void nbl_hw_set_mailbox_irq(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ bool en_msix, u16 gvec)
+{
+ u32 data = 0;
+
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id), &data,
+ sizeof(data));
+ data &= ~(NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK |
+ NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK);
+ if (en_msix)
+ data |= FIELD_PREP(NBL_MAILBOX_QINFO_MAP_MSIX_IDX_MASK,
+ gvec) |
+ FIELD_PREP(NBL_MAILBOX_QINFO_MAP_MSIX_IDX_VALID_MASK,
+ 1);
+
+ nbl_hw_wr_regs(hw_mgt, NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id), &data,
+ sizeof(data));
+ spin_unlock(&hw_mgt->reg_lock);
+ nbl_flush_writes(hw_mgt);
+}
+
+static void nbl_hw_cfg_msix_map(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ bool valid, dma_addr_t dma_addr, u8 bus,
+ u8 devid, u8 function)
+{
+ struct nbl_function_msix_map function_msix_map;
+
+ memset(&function_msix_map, 0, sizeof(function_msix_map));
+
+ /*
+ * Clear VALID on its own first. nbl_hw_wr_regs_lock() writes an
+ * entry in ascending dword order, so a single whole-entry write
+ * would publish the new (or zeroed) address in data[0]/data[1]
+ * while the old VALID bit in data[2] was still set, exposing a
+ * valid entry with a half-written address to the pcompleter's
+ * table fetch. With VALID already zero only the address words
+ * change, and no fetch can be started from a torn entry.
+ */
+ nbl_hw_wr_regs_lock(hw_mgt,
+ NBL_PCOMPLETER_FUNCTION_MSIX_MAP(func_id) +
+ NBL_FUNC_MSIX_MAP_VALID_DW * sizeof(u32),
+ &function_msix_map.data[NBL_FUNC_MSIX_MAP_VALID_DW],
+ sizeof(u32));
+
+ if (valid) {
+ /* program full entry and set VALID */
+ function_msix_map.data[0] = lower_32_bits(dma_addr);
+ function_msix_map.data[1] = upper_32_bits(dma_addr);
+ function_msix_map.data[2] =
+ FIELD_PREP(NBL_FUNCTION_MSIX_MAP_FUNCTION_MASK,
+ function) |
+ FIELD_PREP(NBL_FUNCTION_MSIX_MAP_DEVID_MASK, devid) |
+ FIELD_PREP(NBL_FUNCTION_MSIX_MAP_BUS_MASK, bus) |
+ FIELD_PREP(NBL_FUNCTION_MSIX_MAP_VALID_MASK, 1);
+
+ nbl_hw_wr_regs_lock(hw_mgt,
+ NBL_PCOMPLETER_FUNCTION_MSIX_MAP(func_id),
+ function_msix_map.data,
+ sizeof(function_msix_map));
+ } else {
+ /*
+ * reg_lock prevents concurrent CPU writes to the same
+ * function's MSIX entry, but cannot synchronize hardware DMA
+ * reads. Upper layer uses two-stage destruction + sync sleep
+ * to avoid torn hardware read of partial MSIX entry.
+ * VALID is already cleared above, so the address written
+ * here is either the still-live table (prepare) or zero
+ * (complete); data[2] stays zero, leaving the BDF fields
+ * unused because the entry cannot be fetched.
+ */
+ function_msix_map.data[0] = lower_32_bits(dma_addr);
+ function_msix_map.data[1] = upper_32_bits(dma_addr);
+
+ nbl_hw_wr_regs_lock(hw_mgt,
+ NBL_PCOMPLETER_FUNCTION_MSIX_MAP(func_id),
+ function_msix_map.data,
+ sizeof(function_msix_map));
+ }
+}
+
+static void nbl_hw_cfg_msix_info(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ bool valid, u16 interrupt_id, u8 bus,
+ u8 devid, u8 function, bool msix_mask_en)
+{
+ u32 host_msix_fid = 0;
+ struct nbl_host_msix_info msix_info;
+
+ memset(&msix_info, 0, sizeof(msix_info));
+ if (valid) {
+ host_msix_fid =
+ FIELD_PREP(NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_FID_MASK,
+ func_id) |
+ FIELD_PREP(NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_VLD_MASK,
+ 1);
+
+ msix_info.data[1] =
+ FIELD_PREP(NBL_HOST_MSIX_INFO_FUNCTION_MASK, function) |
+ FIELD_PREP(NBL_HOST_MSIX_INFO_DEVID_MASK, devid) |
+ FIELD_PREP(NBL_HOST_MSIX_INFO_BUS_MASK, bus) |
+ FIELD_PREP(NBL_HOST_MSIX_INFO_VALID_MASK, 1);
+
+ if (msix_mask_en)
+ msix_info.data[1] |=
+ FIELD_PREP(NBL_HOST_MSIX_INFO_MSIX_MASK_EN_MASK, 1);
+ }
+ spin_lock(&hw_mgt->reg_lock);
+ /*
+ * Programming order rule:
+ * Enable: PADPT_HOST_MSIX_INFO -> PCOMPLETER_HOST_MSIX_FID_TABLE
+ * Teardown: reverse order, clear FID VLD first to avoid inconsistent
+ * state
+ */
+ if (valid) {
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_PADPT_HOST_MSIX_INFO_REG_ARR(interrupt_id),
+ msix_info.data, sizeof(msix_info));
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_PCOMPLETER_HOST_MSIX_FID_TABLE(interrupt_id),
+ &host_msix_fid, sizeof(host_msix_fid));
+ } else {
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_PCOMPLETER_HOST_MSIX_FID_TABLE(interrupt_id),
+ &host_msix_fid, sizeof(host_msix_fid));
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_PADPT_HOST_MSIX_INFO_REG_ARR(interrupt_id),
+ msix_info.data, sizeof(msix_info));
+ }
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
static void nbl_hw_update_mailbox_queue_tail_ptr(struct nbl_hw_mgt *hw_mgt,
u16 tail_ptr, u8 txrx)
{
@@ -271,6 +427,8 @@ static void nbl_hw_get_board_info(struct nbl_hw_mgt *hw_mgt,
}
static struct nbl_hw_ops hw_ops = {
+ .cfg_msix_map = nbl_hw_cfg_msix_map,
+ .cfg_msix_info = nbl_hw_cfg_msix_info,
.flush_write = nbl_flush_writes,
.update_mailbox_queue_tail_ptr = nbl_hw_update_mailbox_queue_tail_ptr,
@@ -282,6 +440,7 @@ static struct nbl_hw_ops hw_ops = {
.get_real_bus = nbl_hw_get_real_bus,
.cfg_mailbox_qinfo = nbl_hw_cfg_mailbox_qinfo,
+ .set_mailbox_irq = nbl_hw_set_mailbox_irq,
.get_fw_eth_map = nbl_hw_get_fw_eth_map,
.get_board_info = nbl_hw_get_board_info,
@@ -312,11 +471,12 @@ static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
hw_ops_tbl = devm_kzalloc(dev, sizeof(*hw_ops_tbl), GFP_KERNEL);
if (!hw_ops_tbl)
return ERR_PTR(-ENOMEM);
- if (!hw_ops.flush_write || !hw_ops.update_mailbox_queue_tail_ptr ||
+ if (!hw_ops.cfg_msix_map || !hw_ops.cfg_msix_info ||
+ !hw_ops.flush_write || !hw_ops.update_mailbox_queue_tail_ptr ||
!hw_ops.config_mailbox_rxq || !hw_ops.config_mailbox_txq ||
!hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
!hw_ops.get_host_pf_mask || !hw_ops.get_real_bus ||
- !hw_ops.cfg_mailbox_qinfo ||
+ !hw_ops.cfg_mailbox_qinfo || !hw_ops.set_mailbox_irq ||
!hw_ops.get_fw_eth_map || !hw_ops.get_board_info)
return ERR_PTR(-EINVAL);
hw_ops_tbl->ops = &hw_ops;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index 251dd68d0721..b7c0a87c1dad 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -54,6 +54,50 @@ struct nbl_mailbox_qinfo_cfg_table {
#define NBL_PCIE_HOST_TL_CFG_BUSDEV (NBL_INTF_HOST_PCIE_BASE + 0x11040)
#define NBL_PCIE_BUS_MASK GENMASK(12, 5)
+
+/* -------- HOST_PADPT -------- */
+/* host_padpt host_msix_info */
+#define NBL_PADPT_HOST_MSIX_INFO_REG_ARR(vector_id) \
+ (NBL_INTF_HOST_PADPT_BASE + 0x00010000 + \
+ (vector_id) * sizeof(struct nbl_host_msix_info))
+
+#define NBL_HOST_MSIX_INFO_DWLEN 2
+/* data[0] */
+#define NBL_HOST_MSIX_INFO_INTRL_PNUM_MASK GENMASK(15, 0)
+#define NBL_HOST_MSIX_INFO_INTRL_RATE_MASK GENMASK(31, 16)
+/* data[1] */
+#define NBL_HOST_MSIX_INFO_FUNCTION_MASK GENMASK(2, 0)
+#define NBL_HOST_MSIX_INFO_DEVID_MASK GENMASK(7, 3)
+#define NBL_HOST_MSIX_INFO_BUS_MASK GENMASK(15, 8)
+#define NBL_HOST_MSIX_INFO_VALID_MASK BIT(16)
+#define NBL_HOST_MSIX_INFO_MSIX_MASK_EN_MASK BIT(17)
+struct nbl_host_msix_info {
+ u32 data[NBL_HOST_MSIX_INFO_DWLEN];
+};
+
+/* -------- HOST_PCOMPLETER -------- */
+/* pcompleter_host function_msix_map_table */
+#define NBL_PCOMPLETER_FUNCTION_MSIX_MAP(i) \
+ (NBL_INTF_HOST_PCOMPLETER_BASE + 0x00004000 + \
+ (i) * sizeof(struct nbl_function_msix_map))
+#define NBL_PCOMPLETER_HOST_MSIX_FID_TABLE(i) \
+ (NBL_INTF_HOST_PCOMPLETER_BASE + 0x0003a000 + (i) * sizeof(u32))
+
+#define NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_FID_MASK GENMASK(9, 0)
+#define NBL_PCOMPLETER_HOST_MSIX_FID_TABLE_VLD_MASK BIT(10)
+
+#define NBL_FUNC_MSIX_MAP_DWLEN 4
+/* The VLD/BDF fields live in data[2]; it can be cleared on its own */
+#define NBL_FUNC_MSIX_MAP_VALID_DW 2
+/* data[2] */
+#define NBL_FUNCTION_MSIX_MAP_FUNCTION_MASK GENMASK(2, 0)
+#define NBL_FUNCTION_MSIX_MAP_DEVID_MASK GENMASK(7, 3)
+#define NBL_FUNCTION_MSIX_MAP_BUS_MASK GENMASK(15, 8)
+#define NBL_FUNCTION_MSIX_MAP_VALID_MASK BIT(16)
+struct nbl_function_msix_map {
+ u32 data[NBL_FUNC_MSIX_MAP_DWLEN];
+};
+
#define NBL_FW_BOARD_CONFIG 0x200
#define NBL_FW_BOARD_DW3_OFFSET (NBL_FW_BOARD_CONFIG + 12)
#define NBL_FW_BOARD_DW6_OFFSET (NBL_FW_BOARD_CONFIG + 24)
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
index 7804762a96e0..72b5b889dedb 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
@@ -10,6 +10,9 @@
static struct nbl_resource_ops res_ops = {
.get_vsi_id = nbl_res_func_id_to_vsi_id,
.get_eth_id = nbl_res_get_eth_id,
+ .cfg_msix_map = nbl_res_intr_cfg_msix_map,
+ .destroy_msix_map = nbl_res_intr_destroy_msix_map,
+ .set_mailbox_irq = nbl_res_intr_set_mailbox_irq,
};
static struct nbl_resource_mgt *
@@ -41,7 +44,9 @@ nbl_res_setup_ops(struct device *dev, struct nbl_resource_mgt *res_mgt)
res_ops_tbl = devm_kzalloc(dev, sizeof(*res_ops_tbl), GFP_KERNEL);
if (!res_ops_tbl)
return ERR_PTR(-ENOMEM);
- if (!res_ops.get_vsi_id || !res_ops.get_eth_id)
+ if (!res_ops.get_vsi_id || !res_ops.get_eth_id ||
+ !res_ops.cfg_msix_map || !res_ops.destroy_msix_map ||
+ !res_ops.set_mailbox_irq)
return ERR_PTR(-EINVAL);
res_ops_tbl->ops = &res_ops;
res_ops_tbl->priv = res_mgt;
@@ -282,6 +287,10 @@ static int nbl_res_start(struct nbl_resource_mgt *res_mgt)
ret = nbl_res_ctrl_dev_vsi_info_init(res_mgt);
if (ret)
return ret;
+
+ ret = nbl_intr_mgt_start(res_mgt);
+ if (ret)
+ return ret;
}
return 0;
@@ -322,8 +331,31 @@ int nbl_res_init_leonis(struct nbl_adapter *adap)
void nbl_res_remove_leonis(struct nbl_adapter *adap)
{
+ struct nbl_resource_mgt *res_mgt = adap->core.res_mgt;
+ struct nbl_common_info *common = &adap->common;
+
+ if (!res_mgt)
+ return;
+
/*
- * No resource release here because all memory uses devm managed
- * allocation
+ * Tear down all MSI-X maps before destroying coherent tables.
+ * This is critical on the control PF, whose table carries the map
+ * entries installed for every configured function: they must be
+ * invalidated in hardware before the tables backing them are
+ * freed, otherwise the device keeps DMA-reading released pages.
+ * Sibling PFs are already unbound by the device-link ordering
+ * so no remote PF still has a live map entry or a running
+ * mailbox RPC at this point.
+ */
+ if (common->has_ctrl && res_mgt->intr_mgt)
+ nbl_intr_mgt_stop(res_mgt);
+
+ /* Note:
+ * per-function interrupts arrays (kcalloc) are freed by
+ * nbl_intr_mgt_stop().
+ * MSIX coherent tables are explicitly freed by dma_free_coherent()
+ * inside the intr destroy path, before nbl_intr_mgt_stop() returns.
+ * intr_mgt itself (devm_kzalloc) is released by devres after this
+ * function returns
*/
}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
index b9355262c00d..6eb4dc9e695a 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
@@ -7,4 +7,5 @@
#define _NBL_RESOURCE_LEONIS_H_
#include "../nbl_resource.h"
+#include "../nbl_interrupt.h"
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
new file mode 100644
index 000000000000..3048af5bfeed
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
@@ -0,0 +1,801 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/delay.h>
+#include <linux/dma-mapping.h>
+#include <linux/bitfield.h>
+#include "nbl_interrupt.h"
+
+#define NBL_MSIX_DMA_SYNC_MIN_US 1000 /* us */
+#define NBL_MSIX_DMA_SYNC_MAX_US 1200 /* us */
+
+/*
+ * Bounded best-effort wait for in-flight pcompleter fetches of a MSI-X
+ * map table after that entry's VALID bit has been cleared and flushed.
+ *
+ * The chip exposes no idle/completion status for those fetches, so the
+ * only ordering available is: clear VALID -> flush_write() -> wait. The
+ * window is sized well above the tables' worst-case fetch/response
+ * latency; it is not a guarantee, and it protects two things:
+ * - the coherent table that __nbl_res_intr_complete_destroy_msix_map()
+ * is about to dma_free_coherent();
+ * - the vector set that a reconfiguration is about to hand to another
+ * function.
+ * A future caller that cannot tolerate the residual risk must keep the
+ * table allocated for the lifetime of the function instead of relying on
+ * this delay.
+ */
+static void nbl_intr_quiesce_wait(void)
+{
+ usleep_range(NBL_MSIX_DMA_SYNC_MIN_US, NBL_MSIX_DMA_SYNC_MAX_US);
+}
+
+/*
+ * Release global vector IDs back to intr_net_bmap / intr_other_bmap.
+ * Caller must hold intr_mgt->lock, passing the intr_mgt it locked:
+ * nbl_intr_mgt_stop() clears res_mgt->intr_mgt while a caller may
+ * already be blocked on that mutex, so it must not be re-read here.
+ */
+static void nbl_intr_release_bitmap(struct nbl_interrupt_mgt *intr_mgt,
+ struct nbl_resource_mgt *res_mgt,
+ u16 *vec_buf, u16 cnt)
+{
+ u16 bit;
+ u16 i;
+
+ lockdep_assert_held(&intr_mgt->lock);
+
+ if (!vec_buf || cnt == 0)
+ return;
+
+ for (i = 0; i < cnt; i++) {
+ u16 intr_index = vec_buf[i];
+
+ if (intr_index >= NBL_NET_INTR_BASE) {
+ bit = intr_index - NBL_NET_INTR_BASE;
+ if (bit < NBL_MAX_NET_INTERRUPT)
+ clear_bit(bit, intr_mgt->intr_net_bmap);
+ else
+ dev_warn(res_mgt->common->dev,
+ "invalid net intr index %u\n",
+ intr_index);
+ } else {
+ if (intr_index < NBL_MAX_OTHER_INTERRUPT)
+ clear_bit(intr_index,
+ intr_mgt->intr_other_bmap);
+ else
+ dev_warn(res_mgt->common->dev,
+ "invalid other intr index %u\n",
+ intr_index);
+ }
+ }
+}
+
+/*
+ * Internal (unlocked) mailbox IRQ bind. Caller must hold
+ * intr_mgt->lock, passing the intr_mgt it locked (see
+ * nbl_intr_release_bitmap()).
+ *
+ * The disable path deliberately carries no state check: the teardown
+ * sequence uses it to disarm routing while func_res->state is still
+ * CONFIGURED, and it stays harmless after nbl_intr_mgt_stop() because
+ * the hardware op ignores gvec when en_msix=false. Only the enable
+ * path is gated by the stopping latch and the per-function state.
+ */
+static int __nbl_res_intr_set_mailbox_irq(struct nbl_interrupt_mgt *intr_mgt,
+ struct nbl_resource_mgt *res_mgt,
+ u16 func_id, u16 vector_id,
+ bool en_msix)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ struct nbl_common_info *common = res_mgt->common;
+ struct device *dev = common->dev;
+ struct nbl_func_interrupt_resource_mng *func_res;
+ u16 gvec;
+
+ lockdep_assert_held(&intr_mgt->lock);
+
+ /* func_intr_res[] is PF-indexed, VFs are rejected earlier */
+ if (func_id >= NBL_MAX_PF) {
+ dev_err(dev, "func_id %u out of range\n", func_id);
+ return -EINVAL;
+ }
+
+ if (!en_msix) {
+ hw_ops->set_mailbox_irq(res_mgt->hw_ops_tbl->priv,
+ func_id, false, 0);
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+ return 0;
+ }
+
+ /*
+ * Enable path: the map must be live and not under teardown,
+ * otherwise routing would point at a vector that the complete
+ * phase is about to release and never re-disables.
+ */
+ if (intr_mgt->stopping)
+ return -ESHUTDOWN;
+
+ func_res = &intr_mgt->func_intr_res[func_id];
+ if (func_res->state != NBL_INTR_FUNC_CONFIGURED) {
+ dev_err(dev, "func %u MSIX map not configured (state %u)\n",
+ func_id, func_res->state);
+ return -ENODEV;
+ }
+
+ if (vector_id >= func_res->num_interrupts) {
+ dev_err(dev, "vector_id %u out of range (max %u)\n",
+ vector_id, func_res->num_interrupts - 1);
+ return -EINVAL;
+ }
+
+ gvec = func_res->interrupts[vector_id];
+ hw_ops->set_mailbox_irq(res_mgt->hw_ops_tbl->priv, func_id,
+ en_msix, gvec);
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+ return 0;
+}
+
+/*
+ * Internal (unlocked) MSI-X map teardown prepare phase: only hardware
+ * register operations. The DMA address is retained and only the VALID
+ * bit is cleared; zeroing the address (Stage 2) is deferred to the
+ * complete phase after the hardware-DMA quiesce window.
+ *
+ * Caller must hold intr_mgt->lock.
+ */
+static int
+__nbl_res_intr_prepare_destroy_msix_map(struct nbl_interrupt_mgt *intr_mgt,
+ struct nbl_resource_mgt *res_mgt,
+ u16 func)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ struct nbl_func_interrupt_resource_mng *func_res;
+ u16 *interrupts;
+ u16 intr_num, i;
+ int ret;
+
+ lockdep_assert_held(&intr_mgt->lock);
+
+ if (func >= NBL_MAX_PF) {
+ dev_err(res_mgt->common->dev, "Invalid func_id %u\n", func);
+ return -EINVAL;
+ }
+
+ func_res = &intr_mgt->func_intr_res[func];
+ if (func_res->state != NBL_INTR_FUNC_CONFIGURED)
+ return 0;
+
+ interrupts = func_res->interrupts;
+ intr_num = func_res->num_interrupts;
+
+ /* Step 0: disable mailbox IRQ routing before tearing down map */
+ ret = __nbl_res_intr_set_mailbox_irq(intr_mgt, res_mgt, func, 0, false);
+ if (ret) {
+ dev_err(res_mgt->common->dev,
+ "disable mailbox irq failed, func=%u ret=%d\n",
+ func, ret);
+ return ret;
+ }
+
+ /* Step 1: invalidate each MSIX info entry in hardware first */
+ for (i = 0; i < intr_num; i++) {
+ hw_ops->cfg_msix_info(res_mgt->hw_ops_tbl->priv,
+ func, false, interrupts[i],
+ 0, 0, 0, false);
+ }
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+ /*
+ * Stage 1: retain the DMA address, only clear the VALID bit.
+ * Stage 2 runs after the quiesce window in the complete phase.
+ */
+ hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func,
+ false, func_res->msix_map_table.dma,
+ 0, 0, 0);
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+ func_res->state = NBL_INTR_FUNC_DESTROYING;
+
+ return 0;
+}
+
+/*
+ * __nbl_res_intr_complete_destroy_msix_map - finish hardware teardown and
+ * release vector bitmap, DMA memory and interrupt buffer after the
+ * hardware quiesce window has elapsed.
+ *
+ * Caller must hold intr_mgt->lock.
+ */
+static int
+__nbl_res_intr_complete_destroy_msix_map(struct nbl_interrupt_mgt *intr_mgt,
+ struct nbl_resource_mgt *res_mgt,
+ u16 func_id)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ struct nbl_func_interrupt_resource_mng *func_res;
+ struct nbl_msix_map_table *msix_map_table;
+ struct device *dev = res_mgt->common->dev;
+ u16 *interrupts;
+ u16 intr_num;
+
+ lockdep_assert_held(&intr_mgt->lock);
+
+ if (func_id >= NBL_MAX_PF) {
+ dev_err(dev, "Invalid func_id %u\n", func_id);
+ return -EINVAL;
+ }
+
+ func_res = &intr_mgt->func_intr_res[func_id];
+ if (func_res->state != NBL_INTR_FUNC_DESTROYING)
+ return 0;
+
+ /*
+ * Stage 2: the quiesce window has elapsed, it is now safe to
+ * zero the DMA base address in the hardware map register.
+ */
+ hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
+ false, 0, 0, 0, 0);
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+ interrupts = func_res->interrupts;
+ intr_num = func_res->num_interrupts;
+ msix_map_table = &func_res->msix_map_table;
+
+ if (interrupts) {
+ nbl_intr_release_bitmap(intr_mgt, res_mgt, interrupts,
+ intr_num);
+ kfree(interrupts);
+ }
+
+ /*
+ * Release the coherent table independently of interrupts so a
+ * partially built config (table allocated, vectors never
+ * published) cannot leak coherent DMA memory.
+ */
+ if (msix_map_table->base_addr) {
+ dma_free_coherent(dev, msix_map_table->size,
+ msix_map_table->base_addr,
+ msix_map_table->dma);
+ msix_map_table->base_addr = NULL;
+ msix_map_table->dma = 0;
+ msix_map_table->size = 0;
+ }
+
+ func_res->interrupts = NULL;
+ func_res->num_interrupts = 0;
+ func_res->num_net_interrupts = 0;
+ func_res->state = NBL_INTR_FUNC_IDLE;
+
+ return 0;
+}
+
+/*
+ * Internal (unlocked) MSI-X map teardown. Caller must hold
+ * intr_mgt->lock for the whole sequence, including the hardware-DMA
+ * quiesce window: dropping the lock would let a concurrent caller (or
+ * nbl_intr_mgt_stop()) install/free state against this teardown.
+ *
+ * This is used for the single function synchronous destroy path.
+ */
+static int __nbl_res_intr_destroy_msix_map(struct nbl_interrupt_mgt *intr_mgt,
+ struct nbl_resource_mgt *res_mgt,
+ u16 func_id)
+{
+ int ret;
+
+ lockdep_assert_held(&intr_mgt->lock);
+
+ if (intr_mgt->stopping)
+ return -ESHUTDOWN;
+
+ ret = __nbl_res_intr_prepare_destroy_msix_map(intr_mgt, res_mgt,
+ func_id);
+ if (ret)
+ return ret;
+ /*
+ * prepare() only transitions CONFIGURED functions; an IDLE func
+ * has nothing to wait for or complete.
+ */
+ if (intr_mgt->func_intr_res[func_id].state !=
+ NBL_INTR_FUNC_DESTROYING)
+ return 0;
+
+ nbl_intr_quiesce_wait();
+
+ return __nbl_res_intr_complete_destroy_msix_map(intr_mgt, res_mgt,
+ func_id);
+}
+
+int nbl_res_intr_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
+ u16 func_id)
+{
+ struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+ int ret;
+
+ if (!intr_mgt)
+ return -EINVAL;
+
+ mutex_lock(&intr_mgt->lock);
+ ret = __nbl_res_intr_destroy_msix_map(intr_mgt, res_mgt, func_id);
+ mutex_unlock(&intr_mgt->lock);
+
+ return ret;
+}
+
+/**
+ * nbl_res_intr_cfg_msix_map - allocate & program MSI-X mapping table
+ * @res_mgt: resource management instance
+ * @func_id: target function identifier
+ * @num_net_msix: required net data interrupt vectors
+ * @num_others_msix: required control interrupt vectors
+ * @net_msix_mask_en: enable mask for net interrupt entries
+ *
+ * Allocate interrupt vectors; MSIX coherent DMA table is allocated once
+ * per function on first configuration, entries are rewritten while the
+ * map is invalidated on subsequent reconfigurations. No free/realloc of
+ * DMA table on vector count changes. This removes the DMA table
+ * free/realloc cycle. On reconfiguration the map VALID bit is cleared
+ * and the hardware-DMA quiesce window is observed (lock held) before old
+ * vectors are recycled and the table is rewritten.
+ *
+ * Reconfiguration transiently needs room for the whole new vector set on
+ * top of the old one, because the old set is only released after the new
+ * allocation has succeeded (a failed reconfiguration must leave the old
+ * configuration intact). The bitmap pools are therefore sized for peak,
+ * not maximum, concurrent use: a reconfigure of the same or a smaller
+ * size still fails with -EAGAIN when the pool is nearly exhausted. The
+ * in-tree caller (nbl_dev_init_msix_cnt()) asks for one non-net vector
+ * per PF, far below NBL_MAX_OTHER_INTERRUPT.
+ *
+ * Serialization: this function takes intr_mgt->lock internally to
+ * protect the global vector bitmaps and per-function state against
+ * concurrent callers.
+ *
+ * Return: 0 on success, negative errno on failure
+ */
+int nbl_res_intr_cfg_msix_map(struct nbl_resource_mgt *res_mgt,
+ u16 func_id, u16 num_net_msix,
+ u16 num_others_msix,
+ bool net_msix_mask_en)
+{
+ struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ struct nbl_common_info *common = res_mgt->common;
+ struct nbl_msix_map_table *official_tbl;
+ struct nbl_msix_map *msix_map_entries;
+ struct device *dev = common->dev;
+ u16 requested, intr_index;
+ u8 bus, devid, function;
+ bool entry_masked = false;
+ u16 *tmp_interrupts = NULL;
+ u16 allocated_cnt = 0;
+ u16 *old_interrupts;
+ u16 old_num;
+ bool had_config;
+ int ret = 0;
+ u16 gvec;
+ u16 i, j;
+
+ if (!intr_mgt)
+ return -EINVAL;
+
+ if (!common->has_ctrl)
+ return -EINVAL;
+
+ if (func_id >= NBL_MAX_PF) {
+ dev_err(dev, "Invalid func_id %u\n", func_id);
+ return -EINVAL;
+ }
+
+ if (num_net_msix == 0 && num_others_msix == 0) {
+ dev_err(dev, "MSI-X vector count cannot both be zero\n");
+ return -EINVAL;
+ }
+
+ if (num_net_msix > NBL_MSIX_MAP_TABLE_MAX_ENTRIES ||
+ num_others_msix > NBL_MSIX_MAP_TABLE_MAX_ENTRIES) {
+ dev_err(dev, "MSI-X count out of limit: net=%u, others=%u\n",
+ num_net_msix, num_others_msix);
+ return -EINVAL;
+ }
+
+ if (check_add_overflow(num_net_msix, num_others_msix, &requested) ||
+ requested > NBL_MSIX_MAP_TABLE_MAX_ENTRIES) {
+ dev_err(dev, "Total MSI-X vectors %u exceeds maximum %u\n",
+ requested, NBL_MSIX_MAP_TABLE_MAX_ENTRIES);
+ return -EINVAL;
+ }
+
+ ret = nbl_res_func_id_to_bdf(res_mgt, func_id, &bus, &devid, &function);
+ if (ret) {
+ if (ret == -EOPNOTSUPP)
+ dev_err(dev,
+ "MSI-X mapping for VF func_id=%u is not supported\n",
+ func_id);
+ return ret;
+ }
+
+ mutex_lock(&intr_mgt->lock);
+ official_tbl = &intr_mgt->func_intr_res[func_id].msix_map_table;
+
+ /* Reject new configs during teardown or while func is mid-destroy */
+ if (intr_mgt->stopping) {
+ ret = -ESHUTDOWN;
+ goto out_unlock;
+ }
+ if (intr_mgt->func_intr_res[func_id].state ==
+ NBL_INTR_FUNC_DESTROYING) {
+ ret = -EBUSY;
+ goto out_unlock;
+ }
+
+ had_config = intr_mgt->func_intr_res[func_id].state ==
+ NBL_INTR_FUNC_CONFIGURED;
+
+ /*
+ * Phase1: allocate global vector array first.
+ * Allocate the fixed-size MSIX DMA table only ONCE for this function.
+ */
+ tmp_interrupts = kcalloc(requested, sizeof(*tmp_interrupts),
+ GFP_KERNEL);
+ if (!tmp_interrupts) {
+ ret = -ENOMEM;
+ goto out_unlock;
+ }
+ /* Allocate MSIX DMA table once per function */
+ if (!official_tbl->base_addr) {
+ official_tbl->size =
+ sizeof(struct nbl_msix_map) *
+ NBL_MSIX_MAP_TABLE_MAX_ENTRIES;
+ official_tbl->base_addr = dma_alloc_coherent(dev,
+ official_tbl->size,
+ &official_tbl->dma,
+ GFP_KERNEL);
+ if (!official_tbl->base_addr) {
+ dev_err(dev, "Failed to allocate DMA memory for MSIX table\n");
+ ret = -ENOMEM;
+ goto release_vecs_unlock;
+ }
+ }
+
+ /* Allocate net interrupt vectors */
+ for (i = 0; i < num_net_msix; i++) {
+ intr_index = find_first_zero_bit(intr_mgt->intr_net_bmap,
+ NBL_MAX_NET_INTERRUPT);
+ if (intr_index == NBL_MAX_NET_INTERRUPT) {
+ dev_err(dev, "No free net interrupt vectors left\n");
+ ret = -EAGAIN;
+ goto release_vecs_unlock;
+ }
+ tmp_interrupts[i] = intr_index + NBL_NET_INTR_BASE;
+ set_bit(intr_index, intr_mgt->intr_net_bmap);
+ allocated_cnt++;
+ }
+
+ /* Allocate other interrupt vectors */
+ for (; i < requested; i++) {
+ intr_index =
+ find_first_zero_bit(intr_mgt->intr_other_bmap,
+ NBL_MAX_OTHER_INTERRUPT);
+ if (intr_index == NBL_MAX_OTHER_INTERRUPT) {
+ dev_err(dev, "No free control interrupt vectors left\n");
+ ret = -EAGAIN;
+ goto release_vecs_unlock;
+ }
+ tmp_interrupts[i] = intr_index;
+ set_bit(intr_index, intr_mgt->intr_other_bmap);
+ allocated_cnt++;
+ }
+
+ /*
+ * Phase2: quiesce the old hardware MSIX config before touching
+ * the live DMA table. Same sequence as destroy:
+ * disable mailbox routing -> invalidate per-vector INFO ->
+ * clear map VALID -> flush -> wait for in-flight table fetches.
+ * The lock stays held across the wait, so no concurrent caller
+ * can program the quiesced function. Only then are old vectors
+ * recycled.
+ * NOTE: NO DMA table free here.
+ */
+ if (had_config) {
+ old_interrupts =
+ intr_mgt->func_intr_res[func_id].interrupts;
+ old_num = intr_mgt->func_intr_res[func_id].num_interrupts;
+
+ ret = __nbl_res_intr_set_mailbox_irq(intr_mgt, res_mgt, func_id,
+ 0, false);
+ if (ret) {
+ dev_err(dev, "%s: disable old mailbox irq failed, keep old config\n",
+ __func__);
+ goto release_vecs_unlock;
+ }
+ for (j = 0; j < old_num; j++) {
+ hw_ops->cfg_msix_info(res_mgt->hw_ops_tbl->priv,
+ func_id, false,
+ old_interrupts[j],
+ 0, 0, 0, false);
+ }
+ hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
+ false, official_tbl->dma,
+ 0, 0, 0);
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+ nbl_intr_quiesce_wait();
+
+ nbl_intr_release_bitmap(intr_mgt, res_mgt, old_interrupts,
+ old_num);
+ kfree(old_interrupts);
+ intr_mgt->func_intr_res[func_id].interrupts = NULL;
+ intr_mgt->func_intr_res[func_id].num_interrupts = 0;
+ intr_mgt->func_intr_res[func_id].num_net_interrupts = 0;
+ }
+
+ /* Swap new vector array into func state */
+ intr_mgt->func_intr_res[func_id].interrupts = tmp_interrupts;
+ intr_mgt->func_intr_res[func_id].num_interrupts = requested;
+ intr_mgt->func_intr_res[func_id].num_net_interrupts = num_net_msix;
+ tmp_interrupts = NULL;
+
+ /*
+ * A fresh configuration can find the map entry still VALID with a
+ * DMA address owned by a previous kernel (kexec / forced unload
+ * without FLR: the chip state survives and there is no .shutdown
+ * callback). Invalidate it, and wait out the quiesce window,
+ * before the per-vector INFO/FID entries below are armed:
+ * otherwise the pcompleter could still fetch that stale table and
+ * route it into gvecs this configuration is about to hand out.
+ * The reconfiguration path has already done both above.
+ */
+ if (!had_config) {
+ hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
+ false, 0, 0, 0, 0);
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+ nbl_intr_quiesce_wait();
+ }
+
+ /*
+ * Rewrite the table in the pre-allocated DMA buffer while the
+ * map is invalid, so the device cannot observe a torn old/new
+ * mix. Only entries beyond requested count need explicit zeroing.
+ */
+ msix_map_entries = official_tbl->base_addr;
+ memset(msix_map_entries + requested, 0,
+ (NBL_MSIX_MAP_TABLE_MAX_ENTRIES - requested) *
+ sizeof(*msix_map_entries));
+
+ for (i = 0; i < requested; i++) {
+ gvec = intr_mgt->func_intr_res[func_id].interrupts[i];
+ msix_map_entries[i].data =
+ cpu_to_le16(FIELD_PREP(NBL_MSIX_MAP_VALID_MASK, 1) |
+ FIELD_PREP(NBL_MSIX_MAP_INDEX_MASK,
+ gvec));
+ }
+
+ /* Ensure coherent table writes are visible before HW fetch/enable */
+ dma_wmb();
+
+ /* Enable per-vector INFO entries after the table is published */
+ for (i = 0; i < requested; i++) {
+ gvec = intr_mgt->func_intr_res[func_id].interrupts[i];
+ entry_masked = (i < num_net_msix && net_msix_mask_en);
+ hw_ops->cfg_msix_info(res_mgt->hw_ops_tbl->priv,
+ func_id, true, gvec,
+ bus, devid, function,
+ entry_masked);
+ }
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+ /*
+ * Point the map at the table last and set VALID.
+ *
+ * cfg_msix_map uses the control PF's own BDF (common->hw_bus etc.),
+ * not the target function's BDF. This BDF tags the pcompler DMA
+ * read of the MSI-X map table as originating from the control PF.
+ * The target function's BDF (bus/devid/function from
+ * nbl_res_func_id_to_bdf) is used only in cfg_msix_info for the
+ * host_msix_ctrl table entry BDF filtering.
+ */
+ hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
+ true, official_tbl->dma, common->hw_bus,
+ common->devid, common->function);
+ hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
+
+ intr_mgt->func_intr_res[func_id].state = NBL_INTR_FUNC_CONFIGURED;
+ mutex_unlock(&intr_mgt->lock);
+ return 0;
+
+release_vecs_unlock:
+ nbl_intr_release_bitmap(intr_mgt, res_mgt, tmp_interrupts,
+ allocated_cnt);
+ kfree(tmp_interrupts);
+ /*
+ * On a failed fresh configuration, release the DMA table
+ * allocated during this call. On a failed reconfiguration the
+ * old configuration is still intact (it is only torn down
+ * after all vector allocations succeed) and owns the table.
+ */
+ if (!had_config && official_tbl->base_addr) {
+ dma_free_coherent(dev, official_tbl->size,
+ official_tbl->base_addr,
+ official_tbl->dma);
+ official_tbl->base_addr = NULL;
+ official_tbl->dma = 0;
+ official_tbl->size = 0;
+ }
+out_unlock:
+ mutex_unlock(&intr_mgt->lock);
+ return ret;
+}
+
+/**
+ * nbl_res_intr_set_mailbox_irq - bind mailbox IRQ to specified vector
+ * @res_mgt: resource management instance
+ * @func_id: target function identifier
+ * @vector_id: index inside local interrupt array
+ * @en_msix: enable/disable mailbox interrupt
+ *
+ * Serialization: takes intr_mgt->lock internally and passes the
+ * pointer it read (before taking the lock) down to the helpers, which
+ * must not re-read res_mgt->intr_mgt: nbl_intr_mgt_stop() clears that
+ * pointer while a caller may already be blocked on this mutex.
+ *
+ * Return: 0 on success, negative errno on parameter or state check
+ * failure. The hardware op is void and cannot report failure.
+ */
+int nbl_res_intr_set_mailbox_irq(struct nbl_resource_mgt *res_mgt,
+ u16 func_id, u16 vector_id,
+ bool en_msix)
+{
+ struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+ struct nbl_common_info *common = res_mgt->common;
+ int ret;
+
+ if (!intr_mgt)
+ return -EINVAL;
+
+ if (!common->has_ctrl)
+ return -EINVAL;
+
+ mutex_lock(&intr_mgt->lock);
+ ret = __nbl_res_intr_set_mailbox_irq(intr_mgt, res_mgt, func_id,
+ vector_id, en_msix);
+ mutex_unlock(&intr_mgt->lock);
+
+ return ret;
+}
+
+/*
+ * Only software state is reset here. The chip-internal MSI-X tables
+ * (FUNCTION_MSIX_MAP, PADPT_HOST_MSIX_INFO, HOST_MSIX_FID_TABLE) can
+ * survive a kexec or a forced unload without FLR, and there is no
+ * .shutdown callback to scrub them. They are not cleared here because
+ * the INFO/FID tables are indexed by global vector id, which would mean
+ * rewriting every entry; instead the map entry that the pcompleter
+ * actually fetches is invalidated by nbl_res_intr_cfg_msix_map() before
+ * it is programmed, and by the teardown paths.
+ */
+static struct nbl_interrupt_mgt *nbl_intr_setup_mgt(struct device *dev)
+{
+ struct nbl_interrupt_mgt *intr_mgt;
+ int err;
+
+ intr_mgt = devm_kzalloc(dev, sizeof(*intr_mgt), GFP_KERNEL);
+ if (!intr_mgt)
+ return ERR_PTR(-ENOMEM);
+
+ err = devm_mutex_init(dev, &intr_mgt->lock);
+ if (err)
+ return ERR_PTR(err);
+
+ intr_mgt->stopping = false;
+ bitmap_zero(intr_mgt->intr_net_bmap, NBL_MAX_NET_INTERRUPT);
+ bitmap_zero(intr_mgt->intr_other_bmap, NBL_MAX_OTHER_INTERRUPT);
+
+ return intr_mgt;
+}
+
+int nbl_intr_mgt_start(struct nbl_resource_mgt *res_mgt)
+{
+ struct device *dev = res_mgt->common->dev;
+ struct nbl_interrupt_mgt *intr_mgt;
+ int ret;
+
+ intr_mgt = nbl_intr_setup_mgt(dev);
+ if (IS_ERR(intr_mgt)) {
+ ret = PTR_ERR(intr_mgt);
+ return ret;
+ }
+
+ res_mgt->intr_mgt = intr_mgt;
+ return 0;
+}
+
+/*
+ * nbl_intr_mgt_stop - global control-PF interrupt teardown
+ *
+ * Phase 1 sets the stopping latch and invalidates every configured
+ * function's hardware map entry while holding the lock. The lock is
+ * dropped for the global quiesce window, so a caller that races the
+ * window is rejected by the latch (checked under the lock), not by the
+ * lock being held.
+ *
+ * After this returns res_mgt->intr_mgt is NULL, so the public entry
+ * points report -EINVAL. -ESHUTDOWN/-EBUSY/-ENODEV are only observed
+ * by a caller that latched the pointer before it was cleared.
+ */
+void nbl_intr_mgt_stop(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_interrupt_mgt *intr_mgt = res_mgt->intr_mgt;
+ u16 func_id;
+ int ret;
+
+ if (!intr_mgt)
+ return;
+
+ /*
+ * Phase 1: set the stopping latch and invalidate all hardware
+ * MSIX map entries. A caller racing the quiesce window below is
+ * either still waiting for the lock, in which case it observes
+ * stopping == true at its next checkpoint once it is let in, or
+ * it already holds the lock, in which case this phase waits for
+ * it to finish. The latch is never cleared.
+ *
+ * The en_msix=false mailbox path is exempt from the latch on
+ * purpose: this phase and the per-function teardown use it to
+ * disarm routing, and it only rewrites bits that are already
+ * clear by then.
+ */
+ mutex_lock(&intr_mgt->lock);
+ intr_mgt->stopping = true;
+ for (func_id = 0; func_id < NBL_MAX_PF; func_id++) {
+ if (intr_mgt->func_intr_res[func_id].state ==
+ NBL_INTR_FUNC_CONFIGURED) {
+ dev_info(res_mgt->common->dev,
+ "intr_mgt_stop: preparing destroy map for func %u\n",
+ func_id);
+ ret = __nbl_res_intr_prepare_destroy_msix_map(intr_mgt,
+ res_mgt,
+ func_id);
+ if (ret)
+ dev_warn(res_mgt->common->dev,
+ "intr_mgt_stop: prepare destroy map for func %u failed: %d\n",
+ func_id, ret);
+ }
+ }
+ mutex_unlock(&intr_mgt->lock);
+
+ /*
+ * Global quiesce: wait for straggler DMA table reads after all
+ * MSIX map entries have been invalidated in hardware, before
+ * freeing coherent memory. Best-effort only, see
+ * nbl_intr_quiesce_wait().
+ */
+ nbl_intr_quiesce_wait();
+
+ /* Phase2: safely release MSIX coherent memory and intr resources */
+ mutex_lock(&intr_mgt->lock);
+ for (func_id = 0; func_id < NBL_MAX_PF; func_id++) {
+ if (intr_mgt->func_intr_res[func_id].state ==
+ NBL_INTR_FUNC_DESTROYING) {
+ ret = __nbl_res_intr_complete_destroy_msix_map(intr_mgt,
+ res_mgt,
+ func_id);
+ if (ret)
+ dev_warn(res_mgt->common->dev,
+ "intr_mgt_stop: complete destroy map for func %u failed: %d\n",
+ func_id, ret);
+ }
+ }
+ /*
+ * Unpublish last. Readers do not synchronize on intr_mgt->lock,
+ * so this only has to happen after every hardware access and
+ * every free above; a caller that latched the pointer earlier is
+ * rejected by the stopping latch.
+ */
+ res_mgt->intr_mgt = NULL;
+ mutex_unlock(&intr_mgt->lock);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h
new file mode 100644
index 000000000000..9f66f5e19c98
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.h
@@ -0,0 +1,21 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_INTERRUPT_H_
+#define _NBL_INTERRUPT_H_
+
+#include "nbl_resource.h"
+
+#define NBL_MSIX_MAP_TABLE_MAX_ENTRIES 1024
+int nbl_res_intr_destroy_msix_map(struct nbl_resource_mgt *res_mgt,
+ u16 func_id);
+int nbl_res_intr_cfg_msix_map(struct nbl_resource_mgt *res_mgt,
+ u16 func_id, u16 num_net_msix,
+ u16 num_others_msix,
+ bool net_msix_mask_en);
+int nbl_res_intr_set_mailbox_irq(struct nbl_resource_mgt *res_mgt,
+ u16 func_id, u16 vector_id,
+ bool en_msix);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
index b316fb8e7051..635f34312c56 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.c
@@ -68,6 +68,38 @@ int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id)
return -ENOENT;
}
+int nbl_res_func_id_to_bdf(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u8 *bus, u8 *dev, u8 *function)
+{
+ struct nbl_common_info *common = res_mgt->common;
+ struct nbl_sriov_info *sriov_info;
+ int pfid = func_id;
+ u8 pf_bus, devfn;
+ u32 rel_pf_id;
+ int ret;
+
+ if (!common->has_ctrl || !bus || !dev || !function)
+ return -EINVAL;
+ ret = nbl_common_func_id_to_rel_pf_id(common, pfid, &rel_pf_id);
+ if (ret)
+ return ret;
+ if (rel_pf_id >= common->max_pf) {
+ dev_err(common->dev,
+ "func_id=%u rel_pf_id=%u exceeds max_pf=%u, VF BDF unsupported\n",
+ pfid, rel_pf_id,
+ common->max_pf);
+ return -EOPNOTSUPP;
+ }
+ sriov_info = res_mgt->resource_info->sriov_info + rel_pf_id;
+ pf_bus = PCI_BUS_NUM(sriov_info->bdf);
+ devfn = sriov_info->bdf & 0xff;
+ *bus = pf_bus;
+ *dev = PCI_SLOT(devfn);
+ *function = PCI_FUNC(devfn);
+
+ return 0;
+}
+
int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id)
{
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
index ae0a3d33198d..ee09de854aff 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_resource.h
@@ -17,6 +17,69 @@
struct nbl_resource_mgt;
+/* --------- INTERRUPT ---------- */
+#define NBL_MAX_OTHER_INTERRUPT 1024
+#define NBL_MAX_NET_INTERRUPT 4096
+#define NBL_NET_INTR_BASE NBL_MAX_OTHER_INTERRUPT
+
+#define NBL_MSIX_MAP_VALID_MASK BIT(0)
+#define NBL_MSIX_MAP_INDEX_MASK GENMASK(13, 1)
+#define NBL_MSIX_MAP_RSV_MASK GENMASK(15, 14)
+
+struct nbl_msix_map {
+ __le16 data;
+};
+
+struct nbl_msix_map_table {
+ struct nbl_msix_map *base_addr;
+ dma_addr_t dma;
+ size_t size;
+};
+
+/*
+ * Per-function MSI-X resource state.
+ *
+ * IDLE - nothing configured for this function.
+ * CONFIGURED - map valid in hardware, vectors owned by this function.
+ * DESTROYING - map invalidated, teardown in flight: neither the vector
+ * bitmaps nor the coherent table may be released until the
+ * hardware-DMA quiesce window has elapsed.
+ *
+ * The single-function destroy path holds intr_mgt->lock across that
+ * window. nbl_intr_mgt_stop() cannot, because it quiesces every
+ * function at once: it drops the lock and relies on the stopping latch
+ * to keep a concurrent configuration from installing a map that the
+ * in-flight teardown would free.
+ */
+enum nbl_intr_func_state {
+ NBL_INTR_FUNC_IDLE = 0,
+ NBL_INTR_FUNC_CONFIGURED,
+ NBL_INTR_FUNC_DESTROYING,
+};
+
+struct nbl_func_interrupt_resource_mng {
+ u16 num_interrupts;
+ u16 num_net_interrupts;
+ u16 *interrupts;
+ struct nbl_msix_map_table msix_map_table;
+ u8 state; /* enum nbl_intr_func_state */
+};
+
+struct nbl_interrupt_mgt {
+ struct mutex lock; /* Protects the bitmaps + func_intr_res[] */
+ DECLARE_BITMAP(intr_net_bmap, NBL_MAX_NET_INTERRUPT);
+ DECLARE_BITMAP(intr_other_bmap, NBL_MAX_OTHER_INTERRUPT);
+ /* One-way latch: set on teardown, never cleared */
+ bool stopping;
+ /*
+ * Indexed by absolute PCI function id. Only PF entries are ever
+ * used: nbl_res_func_id_to_bdf() rejects VF ids before any
+ * func_intr_res[] access, and the chip-internal MSI-X vector
+ * space is only managed for PFs.
+ */
+ struct nbl_func_interrupt_resource_mng func_intr_res[NBL_MAX_PF];
+};
+
/* --------- INFO ---------- */
struct nbl_sriov_info {
unsigned int bdf;
@@ -56,14 +119,28 @@ struct nbl_resource_mgt {
struct nbl_resource_info *resource_info;
struct nbl_channel_ops_tbl *chan_ops_tbl;
struct nbl_hw_ops_tbl *hw_ops_tbl;
+ /*
+ * Published interrupt manager, control PF only. The public entry
+ * points read this pointer without the manager lock (the object is
+ * devres-owned, so it outlives them) and then pass the value they
+ * read down to the internal helpers; nbl_intr_mgt_stop() clears it
+ * under the manager lock as its last step. A caller that latches
+ * the pointer before that sees stopping == true and fails with
+ * -ESHUTDOWN; a caller that reads it after gets -EINVAL.
+ */
+ struct nbl_interrupt_mgt *intr_mgt;
};
int nbl_res_vsi_id_to_pf_id(struct nbl_resource_mgt *res_mgt, u16 vsi_id);
int nbl_res_func_id_to_vsi_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
u16 type, u16 *vsi_id);
+int nbl_res_func_id_to_bdf(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u8 *bus, u8 *dev, u8 *function);
int nbl_res_get_eth_id(struct nbl_resource_mgt *res_mgt, u16 func_id,
u16 vsi_id, u8 *eth_num, u8 *eth_id, u8 *logic_eth_id);
+int nbl_intr_mgt_start(struct nbl_resource_mgt *res_mgt);
int nbl_res_pf_dev_vsi_type_to_hw_vsi_type(struct nbl_resource_mgt *res_mgt,
u16 src_type,
enum nbl_vsi_serv_type *dst_type);
+void nbl_intr_mgt_stop(struct nbl_resource_mgt *res_mgt);
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index 4b420db91963..89b979e2cdb5 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -12,6 +12,13 @@ struct nbl_board_port_info;
struct nbl_hw_mgt;
struct nbl_adapter;
struct nbl_hw_ops {
+ void (*cfg_msix_map)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ bool valid, dma_addr_t dma_addr, u8 bus,
+ u8 devid, u8 function);
+ void (*cfg_msix_info)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ bool valid, u16 interrupt_id, u8 bus,
+ u8 devid, u8 function,
+ bool net_msix_mask_en);
/**
* flush_write - Flush posted MEMORY-BAR writes
* @hw_mgt: hardware management context
@@ -59,6 +66,8 @@ struct nbl_hw_ops {
void (*cfg_mailbox_qinfo)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
u8 bus, u8 devid, u8 function);
+ void (*set_mailbox_irq)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
+ bool en_msix, u16 gvec);
void (*get_fw_eth_map)(struct nbl_hw_mgt *hw_mgt, u32 *eth_map);
/**
* get_board_info - Fetch board info from firmware
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
index 7136b282fb80..e718ea41a816 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
@@ -12,6 +12,12 @@ struct nbl_resource_mgt;
struct nbl_adapter;
struct nbl_resource_ops {
+ int (*cfg_msix_map)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 num_net_msix, u16 num_others_msix,
+ bool net_msix_mask_en);
+ int (*destroy_msix_map)(struct nbl_resource_mgt *res_mgt, u16 func_id);
+ int (*set_mailbox_irq)(struct nbl_resource_mgt *res_mgt, u16 func_id,
+ u16 vector_id, bool en_msix);
int (*get_vsi_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
u16 type, u16 *vsi_id);
int (*get_eth_id)(struct nbl_resource_mgt *res_mgt, u16 func_id,
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 3/8] net/nebula-matrix: add intr resource implementation
2026-10-10 9:59 ` [PATCH v31 net-next 3/8] net/nebula-matrix: add intr " illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 3
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The comments in nbl_intr_setup_mgt() and
nbl_res_intr_cfg_msix_map() say the chip MSI-X tables…
- [Low] The commit message says "Only PFs are supported; VF ids are
rejected with -EOPNOTSUPP".
- [Low] Several comments describe error semantics the code does not have.
- [Low] Concern (not confirmed): the public entry points read
res_mgt->intr_mgt without a lock and then take intr_mgt->lock.
--
net/nebula-matrix: add intr resource implementation
This patch adds an nbl_interrupt module. It manages the chip-internal
MSI-X vector index space using separate net and control bitmaps, and it
programs the per-PF chip MSI-X mapping tables. It adds the cfg_msix_map,
destroy_msix_map and set_mailbox_irq resource ops plus the hw helpers
behind them, and calls nbl_intr_mgt_start()/nbl_intr_mgt_stop() from
resource init and remove.
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
> new file mode 100644
> index 0000000000000..3048af5bfeed3
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c
[ ... ]
> +int nbl_res_intr_cfg_msix_map(struct nbl_resource_mgt *res_mgt,
> + u16 func_id, u16 num_net_msix,
> + u16 num_others_msix,
> + bool net_msix_mask_en)
> +{
[ ... ]
> + if (func_id >= NBL_MAX_PF) {
> + dev_err(dev, "Invalid func_id %u\n", func_id);
> + return -EINVAL;
> + }
[ ... ]
> + ret = nbl_res_func_id_to_bdf(res_mgt, func_id, &bus, &devid, &function);
> + if (ret) {
> + if (ret == -EOPNOTSUPP)
> + dev_err(dev,
> + "MSI-X mapping for VF func_id=%u is not supported\n",
> + func_id);
> + return ret;
> + }
[Severity: Low]
The commit message says "Only PFs are supported; VF ids are rejected with
-EOPNOTSUPP." Does the code do that?
NBL_MAX_PF is 8, so any VF id fails the func_id >= NBL_MAX_PF check above.
It gets -EINVAL before nbl_res_func_id_to_bdf() is ever called.
nbl_res_func_id_to_bdf() returns -EOPNOTSUPP only when rel_pf_id >=
common->max_pf. That is a PF that does not exist on this board, for
example PF 3 on a 2-port card. Even so, its message ("VF BDF unsupported")
and the message here ("MSI-X mapping for VF func_id=%u") both call it a
VF.
The comment on func_intr_res[] in struct nbl_interrupt_mgt has a related
problem. It says "nbl_res_func_id_to_bdf() rejects VF ids before any
func_intr_res[] access". However, nbl_res_intr_destroy_msix_map() and
nbl_res_intr_set_mailbox_irq() never call nbl_res_func_id_to_bdf(). They
only check func_id >= NBL_MAX_PF.
So an absent-PF id in [max_pf, NBL_MAX_PF) is accepted on those paths, and
the en_msix=false path writes NBL_MAILBOX_QINFO_MAP_REG_ARR(func_id) for
it. For the same reason, the "VFs are rejected earlier" comment in
__nbl_res_intr_set_mailbox_irq() is inaccurate.
Later in the series, the dispatch responders reject rel_pf_id >= max_pf
with -EPERM before calling these ops, and local callers pass
common->mgt_pf. This therefore looks limited to the documentation and the
diagnostics.
Could the commit message, comments and error strings be updated to match?
Alternatively, should destroy_msix_map and set_mailbox_irq check against
max_pf as well?
[ ... ]
> +/*
> + * Only software state is reset here. The chip-internal MSI-X tables
> + * (FUNCTION_MSIX_MAP, PADPT_HOST_MSIX_INFO, HOST_MSIX_FID_TABLE) can
> + * survive a kexec or a forced unload without FLR, and there is no
> + * .shutdown callback to scrub them. They are not cleared here because
> + * the INFO/FID tables are indexed by global vector id, which would mean
> + * rewriting every entry; instead the map entry that the pcompleter
> + * actually fetches is invalidated by nbl_res_intr_cfg_msix_map() before
> + * it is programmed, and by the teardown paths.
> + */
> +static struct nbl_interrupt_mgt *nbl_intr_setup_mgt(struct device *dev)
> +{
[Severity: Medium]
After a kexec or forced unload without FLR, what happens to the
FUNCTION_MSIX_MAP entries of PFs that the new kernel has not configured
yet?
The only stale-map invalidation is the !had_config block in
nbl_res_intr_cfg_msix_map(), and it covers only the func_id being
configured:
if (!had_config) {
hw_ops->cfg_msix_map(res_mgt->hw_ops_tbl->priv, func_id,
false, 0, 0, 0, 0);
The bitmaps start empty here. The first PF the new kernel configures
therefore gets gvecs from the bottom of each pool and arms INFO/FID for
them.
Meanwhile, a sibling PF's map entry from the previous kernel can still be
VALID. It still points at an old-kernel DMA address and maps into those
same gvecs. Bus mastering is already enabled by pci_set_master() in
nbl_probe() at this point.
Could the pcompleter then fetch that stale table, read memory now owned by
the new kernel, or route an interrupt into a gvec now owned by another
function? This looks like the same hazard the fresh-config comment in
nbl_res_intr_cfg_msix_map() describes, but for a different function.
nbl_intr_mgt_stop() and the destroy path only act on CONFIGURED functions.
A stale entry for a PF that stays IDLE would therefore never be cleared.
FUNCTION_MSIX_MAP is indexed by func_id, not by gvec, so the reason given
above for not scrubbing doesn't seem to apply to it. Would it be possible
to invalidate all NBL_MAX_PF map entries at init, followed by one flush
and one quiesce wait?
[ ... ]
> +/*
> + * nbl_intr_mgt_stop - global control-PF interrupt teardown
> + *
> + * Phase 1 sets the stopping latch and invalidates every configured
> + * function's hardware map entry while holding the lock. The lock is
> + * dropped for the global quiesce window, so a caller that races the
> + * window is rejected by the latch (checked under the lock), not by the
> + * lock being held.
> + *
> + * After this returns res_mgt->intr_mgt is NULL, so the public entry
> + * points report -EINVAL. -ESHUTDOWN/-EBUSY/-ENODEV are only observed
> + * by a caller that latched the pointer before it was cleared.
> + */
[Severity: Low]
This isn't a bug, but are these error code descriptions accurate?
The -EBUSY branch in nbl_res_intr_cfg_msix_map() looks unreachable:
if (intr_mgt->func_intr_res[func_id].state ==
NBL_INTR_FUNC_DESTROYING) {
ret = -EBUSY;
__nbl_res_intr_destroy_msix_map() holds the lock from prepare through
complete, and it always leaves the function IDLE. The only other time
DESTROYING is visible under the lock is the quiesce window in
nbl_intr_mgt_stop(). By then stopping is already true, so the
-ESHUTDOWN check just above fires first.
Also, __nbl_res_intr_set_mailbox_irq() returns -ENODEV whenever en_msix is
true and the function is not CONFIGURED, whether or not stop has run.
The comment on nbl_resource_mgt::intr_mgt in nbl_resource.h has a similar
mismatch. It says a caller that latched the pointer early "sees
stopping == true and fails with -ESHUTDOWN". But the en_msix=false path
returns 0 before any stopping check:
if (!en_msix) {
hw_ops->set_mailbox_irq(res_mgt->hw_ops_tbl->priv,
func_id, false, 0);
hw_ops->flush_write(res_mgt->hw_ops_tbl->priv);
return 0;
}
Could these comments be adjusted, and the dead -EBUSY branch dropped?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v31 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
` (2 preceding siblings ...)
2026-10-10 9:59 ` [PATCH v31 net-next 3/8] net/nebula-matrix: add intr " illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
2026-10-10 9:59 ` [PATCH v31 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
` (3 subsequent siblings)
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Add Leonis chip-wide datapath init/deinit hooks (init_module and
deinit_module) in the hw and resource layers. The hooks use the
driver_status register as a handshake with firmware.
deinit_module clears driver_status, flushes the write so it reaches the
device, then waits a bounded 2-3 ms window for the firmware cleanup pass
to run. Firmware cleanup is asynchronous and the hardware exposes no
completion status, so the flush only guarantees the write was posted and
the wait is best-effort; it does not prove the firmware pass finished.
init_module checks the driver_status bit on entry and, if a previous
instance left it set (kexec or a forced unload), clears it and waits the
same window before re-initializing.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../net/ethernet/nebula-matrix/nbl/Makefile | 1 +
.../nebula-matrix/nbl/nbl_hw/nbl_chip.c | 32 +
.../nebula-matrix/nbl/nbl_hw/nbl_chip.h | 12 +
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c | 672 +++++++++++++++++-
.../nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h | 222 ++++++
.../nbl_hw_leonis/nbl_resource_leonis.c | 5 +-
.../nbl_hw_leonis/nbl_resource_leonis.h | 1 +
.../nbl/nbl_include/nbl_def_hw.h | 3 +
.../nbl/nbl_include/nbl_def_resource.h | 3 +
.../nbl/nbl_include/nbl_include.h | 19 +
10 files changed, 968 insertions(+), 2 deletions(-)
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index 5aec8e44f5d7..be314b909d66 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -9,4 +9,5 @@ nbl-objs += nbl_common/nbl_common.o \
nbl_hw/nbl_hw_leonis/nbl_resource_leonis.o \
nbl_hw/nbl_resource.o \
nbl_hw/nbl_interrupt.o \
+ nbl_hw/nbl_chip.o \
nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
new file mode 100644
index 000000000000..419eb6392ada
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
@@ -0,0 +1,32 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include "nbl_chip.h"
+
+void nbl_res_chip_deinit_module(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
+ struct nbl_common_info *common = res_mgt->common;
+
+ if (!common->has_ctrl)
+ return;
+ hw_ops->deinit_module(res_mgt->hw_ops_tbl->priv);
+}
+
+int nbl_res_chip_init_module(struct nbl_resource_mgt *res_mgt)
+{
+ struct nbl_common_info *common = res_mgt->common;
+ struct nbl_hw_ops *hw_ops;
+ u8 eth_speed, eth_num;
+ struct nbl_hw_mgt *p;
+
+ if (!common->has_ctrl)
+ return -EINVAL;
+ eth_speed = res_mgt->resource_info->board_info.eth_speed;
+ eth_num = res_mgt->resource_info->board_info.eth_num;
+ hw_ops = res_mgt->hw_ops_tbl->ops;
+ p = res_mgt->hw_ops_tbl->priv;
+ return hw_ops->init_module(p, eth_speed, eth_num);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h
new file mode 100644
index 000000000000..d14093ba916c
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.h
@@ -0,0 +1,12 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_CHIP_H_
+#define _NBL_CHIP_H_
+
+#include "nbl_resource.h"
+int nbl_res_chip_init_module(struct nbl_resource_mgt *res_mgt);
+void nbl_res_chip_deinit_module(struct nbl_resource_mgt *res_mgt);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
index 1712cdbc5fe7..bd1216265036 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
@@ -5,11 +5,25 @@
#include <linux/device.h>
#include <linux/pci.h>
#include <linux/bits.h>
+#include <linux/delay.h>
#include <linux/io.h>
#include <linux/spinlock.h>
#include <linux/bitfield.h>
#include "nbl_hw_leonis.h"
+/*
+ * Firmware cleanup after driver_status=false is asynchronous and the
+ * current hardware revision exposes no cleanup-complete status bit, so
+ * this delay is best-effort only: it gives the firmware pass time to
+ * settle before the caller returns, but cannot prove that it has.
+ * Bounding the wait keeps a later init_module() from reprogramming
+ * per-PF/chip-wide registers while firmware is realistically still
+ * wiping them; the hard teardown ordering comes from the device-link
+ * relationship between the sibling PFs and the control PF.
+ */
+#define NBL_FW_CLEANUP_SYNC_MIN_US 2000
+#define NBL_FW_CLEANUP_SYNC_MAX_US 3000
+
static void nbl_hw_read_mbx_regs(struct nbl_hw_mgt *hw_mgt, u64 reg, u32 *data,
u32 len)
{
@@ -150,6 +164,658 @@ static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)
*eth_map = FIELD_GET(NBL_FW_BOARD_DW6_ETH_BITMAP_MASK, data);
}
+static u32 nbl_hw_get_quirks(struct nbl_hw_mgt *hw_mgt)
+{
+ u32 quirks = 0;
+
+ /*
+ * Read quirk bits from mailbox register.
+ * All supported firmware implement the quirk ABI,
+ * firmware always populates NBL_LEONIS_QUIRKS_OFFSET.
+ * Value ~0U indicates no active quirks.
+ */
+ nbl_hw_read_mbx_regs(hw_mgt, NBL_LEONIS_QUIRKS_OFFSET, &quirks,
+ sizeof(u32));
+
+ if (quirks == ~0u)
+ return 0;
+
+ return quirks;
+}
+
+static void nbl_configure_dped_checksum(struct nbl_hw_mgt *hw_mgt)
+{
+ u32 data = 0;
+
+ /* DPED dped_l4_ck_cmd_40 for sctp */
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_DPED_L4_CK_CMD_40_ADDR, &data, sizeof(data));
+ data |= FIELD_PREP(NBL_DPED_L4_CK_CMD_40_EN_MASK, 1);
+ nbl_hw_wr_regs(hw_mgt, NBL_DPED_L4_CK_CMD_40_ADDR, &data, sizeof(data));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_dped_init(struct nbl_hw_mgt *hw_mgt)
+{
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_wr32(hw_mgt, NBL_DPED_VLAN_OFFSET, 0xC);
+ nbl_hw_wr32(hw_mgt, NBL_DPED_DSCP_OFFSET_0, 0x8);
+ nbl_hw_wr32(hw_mgt, NBL_DPED_DSCP_OFFSET_1, 0x4);
+ spin_unlock(&hw_mgt->reg_lock);
+ /* dped checksum offload */
+ nbl_configure_dped_checksum(hw_mgt);
+}
+
+static void nbl_uped_init(struct nbl_hw_mgt *hw_mgt)
+{
+ u32 hw_edit = 0;
+
+ /* V4 TCP: l3_len = 0 */
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V4_TCP_IDX),
+ &hw_edit, sizeof(hw_edit));
+ hw_edit &= ~NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK;
+ nbl_hw_wr_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V4_TCP_IDX),
+ &hw_edit, sizeof(hw_edit));
+
+ /* V6 TCP: l3_len = 1 */
+ nbl_hw_rd_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V6_TCP_IDX),
+ &hw_edit, sizeof(hw_edit));
+ hw_edit = (hw_edit & ~NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK) |
+ FIELD_PREP(NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK, 1);
+ nbl_hw_wr_regs(hw_mgt, NBL_UPED_HW_EDT_PROF_TABLE(NBL_UPED_V6_TCP_IDX),
+ &hw_edit, sizeof(hw_edit));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static int nbl_shaping_eth_init(struct nbl_hw_mgt *hw_mgt, u8 eth_id, u8 speed)
+{
+ struct nbl_shaping_dvn_dport_u dvn_dport = { 0 };
+ struct nbl_shaping_dport_u dport = { 0 };
+ u32 rate, half_rate;
+ u32 depth;
+ u64 low_val, high_val;
+
+ switch (speed) {
+ case NBL_FW_PORT_SPEED_100G:
+ rate = 100000;
+ break;
+ case NBL_FW_PORT_SPEED_50G:
+ rate = 50000;
+ break;
+ case NBL_FW_PORT_SPEED_25G:
+ rate = 25000;
+ break;
+ case NBL_FW_PORT_SPEED_10G:
+ rate = 10000;
+ break;
+ default:
+ dev_err(hw_mgt->common->dev,
+ "Unsupported port speed %u for eth%u\n", speed, eth_id);
+ return -EINVAL;
+ }
+
+ half_rate = rate / 2;
+ depth = max_t(u32, rate * 2, NBL_LR_LEONIS_NET_BUCKET_DEPTH);
+
+ /* 1. clear valid first
+ * dport and dvn_dport are zero-initialised above, so VALID=0 already
+ */
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DPORT_REG(eth_id), dport.data,
+ sizeof(dport));
+ nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DVN_DPORT_REG(eth_id),
+ dvn_dport.data, sizeof(dvn_dport));
+
+ /* 2. write config words (valid=0, safe) */
+ low_val = FIELD_PREP(NBL_DPORT_CIR_MASK, rate) |
+ FIELD_PREP(NBL_DPORT_PIR_MASK, rate) |
+ FIELD_PREP(NBL_DPORT_DEPTH_MASK, depth) |
+ FIELD_PREP(NBL_DPORT_CBS_MASK_LOW, depth & 0x3F);
+ high_val = FIELD_PREP(NBL_DPORT_CBS_MASK_HIGH, depth >> 6) |
+ FIELD_PREP(NBL_DPORT_PBS_MASK, depth);
+ /* Fixed split, independent of host endian */
+ dport.data[0] = lower_32_bits(low_val);
+ dport.data[1] = upper_32_bits(low_val);
+ dport.data[2] = lower_32_bits(high_val);
+ dport.data[3] = upper_32_bits(high_val);
+
+ low_val = FIELD_PREP(NBL_DPORT_CIR_MASK, half_rate) |
+ FIELD_PREP(NBL_DPORT_PIR_MASK, rate) |
+ FIELD_PREP(NBL_DPORT_DEPTH_MASK, depth) |
+ FIELD_PREP(NBL_DPORT_CBS_MASK_LOW, depth & 0x3F);
+ high_val = FIELD_PREP(NBL_DPORT_CBS_MASK_HIGH, depth >> 6) |
+ FIELD_PREP(NBL_DPORT_PBS_MASK, depth);
+ dvn_dport.data[0] = lower_32_bits(low_val);
+ dvn_dport.data[1] = upper_32_bits(low_val);
+ dvn_dport.data[2] = lower_32_bits(high_val);
+ dvn_dport.data[3] = upper_32_bits(high_val);
+
+ nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DPORT_REG(eth_id), dport.data,
+ sizeof(dport));
+ nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DVN_DPORT_REG(eth_id),
+ dvn_dport.data, sizeof(dvn_dport));
+
+ /* 3. commit: set valid last */
+ low_val = FIELD_PREP(NBL_DPORT_VALID_MASK, 1);
+ dport.data[0] |= lower_32_bits(low_val);
+
+ low_val = FIELD_PREP(NBL_DPORT_VALID_MASK, 1);
+ dvn_dport.data[0] |= lower_32_bits(low_val);
+
+ nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DPORT_REG(eth_id), dport.data,
+ sizeof(dport));
+ nbl_hw_wr_regs(hw_mgt, NBL_SHAPING_DVN_DPORT_REG(eth_id),
+ dvn_dport.data, sizeof(dvn_dport));
+ spin_unlock(&hw_mgt->reg_lock);
+ return 0;
+}
+
+static int nbl_shaping_init(struct nbl_hw_mgt *hw_mgt, u8 speed)
+{
+#define NBL_SHAPING_FLUSH_INTERVAL 128
+ struct nbl_shaping_net_u net_shaping = { 0 };
+ u32 eth_bitmap = 0;
+ u32 reg_val;
+ int ret;
+ int i;
+
+ nbl_hw_get_fw_eth_map(hw_mgt, ð_bitmap);
+ for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+ if (!(eth_bitmap & BIT(i)))
+ continue;
+ ret = nbl_shaping_eth_init(hw_mgt, i, speed);
+ if (ret)
+ return ret;
+ }
+ nbl_hw_rd_regs_lock(hw_mgt, NBL_DSCH_PSHA_EN_ADDR, ®_val,
+ sizeof(reg_val));
+ reg_val &= ~NBL_DSCH_PSHA_EN_MASK;
+ reg_val |= FIELD_PREP(NBL_DSCH_PSHA_EN_MASK,
+ eth_bitmap & GENMASK(3, 0));
+ nbl_hw_wr_regs_lock(hw_mgt, NBL_DSCH_PSHA_EN_ADDR, ®_val,
+ sizeof(reg_val));
+
+ for (i = 0; i < NBL_MAX_PF; i++) {
+ nbl_hw_wr_regs_lock(hw_mgt, NBL_SHAPING_NET_REG(i),
+ net_shaping.data,
+ sizeof(net_shaping));
+ if ((i + 1) % NBL_SHAPING_FLUSH_INTERVAL == 0)
+ nbl_flush_writes(hw_mgt);
+ }
+ nbl_flush_writes(hw_mgt);
+ return 0;
+}
+
+static void nbl_dsch_qid_max_init(struct nbl_hw_mgt *hw_mgt)
+{
+ u32 quanta = 0;
+ u32 qid_max = 0;
+
+ quanta = FIELD_PREP(NBL_DSCH_VN_QUANTA_H_QUA_MASK, NBL_HOST_QUANTA) |
+ FIELD_PREP(NBL_DSCH_VN_QUANTA_E_QUA_MASK, NBL_ECPU_QUANTA);
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_wr_regs(hw_mgt, NBL_DSCH_VN_QUANTA_ADDR, &quanta,
+ sizeof(quanta));
+ nbl_hw_rd_regs(hw_mgt, NBL_DSCH_HOST_QID_MAX, &qid_max,
+ sizeof(qid_max));
+ qid_max &= ~NBL_DSCH_HOST_QID_MAX_MASK;
+ qid_max |= FIELD_PREP(NBL_DSCH_HOST_QID_MAX_MASK, NBL_MAX_QUEUE_ID);
+ nbl_hw_wr_regs(hw_mgt, NBL_DSCH_HOST_QID_MAX, &qid_max,
+ sizeof(qid_max));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static int nbl_ustore_init(struct nbl_hw_mgt *hw_mgt, u8 eth_num)
+{
+ u32 eth_bitmap = 0;
+ u32 drop_th = 0;
+ u32 pkt_len = 0;
+ u32 reg_val = 0;
+ int i;
+
+ /*
+ * eth_num is validated in the resource layer:
+ * nbl_res_init_pf_num() requires 1/2/4 PFs, and
+ * nbl_res_ctrl_dev_setup_eth_info() requires max_pf == eth_num.
+ * This is a defensive check only; if it triggers, the resource
+ * layer validation was bypassed, which is a bug.
+ */
+ if (WARN_ON(eth_num != 1 && eth_num != 2 && eth_num != 4))
+ return -EINVAL;
+ /* Read current packet length config
+ *(to preserve other fields while updating 'min')
+ */
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_USTORE_PKT_LEN_ADDR, &pkt_len,
+ sizeof(pkt_len));
+ /* min arp packet length 42 (14 + 28) */
+ pkt_len &= ~NBL_USTORE_PKT_LEN_MIN_MASK;
+ pkt_len |= FIELD_PREP(NBL_USTORE_PKT_LEN_MIN_MASK, 42);
+ nbl_hw_wr_regs(hw_mgt, NBL_USTORE_PKT_LEN_ADDR, &pkt_len,
+ sizeof(pkt_len));
+
+ drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_EN_MASK, 1);
+ if (eth_num == 1)
+ drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK,
+ NBL_USTORE_SINGLE_ETH_DROP_TH);
+ else if (eth_num == 2)
+ drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK,
+ NBL_USTORE_DUAL_ETH_DROP_TH);
+ else
+ drop_th |= FIELD_PREP(NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK,
+ NBL_USTORE_QUAD_ETH_DROP_TH);
+ nbl_hw_get_fw_eth_map(hw_mgt, ð_bitmap);
+ for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+ if (!(eth_bitmap & BIT(i)))
+ continue;
+ nbl_hw_rd_regs(hw_mgt, NBL_USTORE_PORT_DROP_TH_REG_ARR(i),
+ ®_val, sizeof(reg_val));
+ reg_val &= ~(NBL_USTORE_PORT_DROP_TH_EN_MASK |
+ NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK);
+ reg_val |= drop_th;
+ nbl_hw_wr_regs(hw_mgt, NBL_USTORE_PORT_DROP_TH_REG_ARR(i),
+ ®_val, sizeof(reg_val));
+ }
+
+ /* Clear port drop/truncate counters by reading them
+ * (hardware has read-to-clear behavior for these registers)
+ */
+ for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+ if (!(eth_bitmap & BIT(i)))
+ continue;
+ nbl_hw_rd32(hw_mgt, NBL_USTORE_BUF_PORT_DROP_PKT(i));
+ nbl_hw_rd32(hw_mgt, NBL_USTORE_BUF_PORT_TRUN_PKT(i));
+ }
+ spin_unlock(&hw_mgt->reg_lock);
+ return 0;
+}
+
+static void nbl_dstore_init(struct nbl_hw_mgt *hw_mgt, u8 speed)
+{
+ u32 eth_bitmap = 0;
+ u32 drop_th = 0;
+ u32 fc_th = 0;
+ u32 bp_th = 0;
+ int i;
+
+ for (i = 0; i < NBL_DSTORE_PORT_DROP_TH_DEPTH; i++) {
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_DSTORE_PORT_DROP_TH_REG(i), &drop_th,
+ sizeof(drop_th));
+ drop_th &= ~NBL_DSTORE_PORT_DROP_EN_MASK;
+ nbl_hw_wr_regs(hw_mgt, NBL_DSTORE_PORT_DROP_TH_REG(i), &drop_th,
+ sizeof(drop_th));
+ spin_unlock(&hw_mgt->reg_lock);
+ }
+
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_DSTORE_DISC_BP_TH, &bp_th, sizeof(bp_th));
+ bp_th |= FIELD_PREP(NBL_DSTORE_DISC_BP_TH_EN_MASK, 1);
+ nbl_hw_wr_regs(hw_mgt, NBL_DSTORE_DISC_BP_TH, &bp_th, sizeof(bp_th));
+ spin_unlock(&hw_mgt->reg_lock);
+
+ nbl_hw_get_fw_eth_map(hw_mgt, ð_bitmap);
+ for (i = 0; i < NBL_MAX_ETHERNET; i++) {
+ if (!(eth_bitmap & BIT(i)))
+ continue;
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_DSTORE_D_DPORT_FC_TH_REG(i), &fc_th,
+ sizeof(fc_th));
+ fc_th &= ~(NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK |
+ NBL_DSTORE_D_DPORT_FC_XON_TH_MASK);
+ if (speed == NBL_FW_PORT_SPEED_100G) {
+ fc_th |=
+ FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK,
+ NBL_DSTORE_DROP_XOFF_TH_100G) |
+ FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XON_TH_MASK,
+ NBL_DSTORE_DROP_XON_TH_100G);
+ } else {
+ fc_th |=
+ FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK,
+ NBL_DSTORE_DROP_XOFF_TH) |
+ FIELD_PREP(NBL_DSTORE_D_DPORT_FC_XON_TH_MASK,
+ NBL_DSTORE_DROP_XON_TH);
+ }
+
+ fc_th |= FIELD_PREP(NBL_DSTORE_D_DPORT_FC_FC_EN_MASK, 1);
+ nbl_hw_wr_regs(hw_mgt, NBL_DSTORE_D_DPORT_FC_TH_REG(i), &fc_th,
+ sizeof(fc_th));
+ spin_unlock(&hw_mgt->reg_lock);
+ }
+}
+
+static void nbl_dvn_descreq_num_cfg(struct nbl_hw_mgt *hw_mgt, u8 descreq_num)
+{
+ u8 split_ring_num = (descreq_num >> 3) & 0x1;
+ u8 ring_num = descreq_num & 0x7;
+ u32 num_cfg;
+ u32 reg_val;
+
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_DVN_DESCREQ_NUM_CFG, ®_val,
+ sizeof(reg_val));
+
+ num_cfg = FIELD_PREP(NBL_DVN_DESCREQ_NUM_CFG_AVRING_DESREQ_NUM_CFG_MASK,
+ split_ring_num) |
+ FIELD_PREP(NBL_DVN_DESCREQ_NUM_CFG_PACKED_L1_NUM_MASK,
+ ring_num);
+ reg_val &= ~(NBL_DVN_DESCREQ_NUM_CFG_AVRING_DESREQ_NUM_CFG_MASK |
+ NBL_DVN_DESCREQ_NUM_CFG_PACKED_L1_NUM_MASK);
+ reg_val |= num_cfg;
+ nbl_hw_wr_regs(hw_mgt, NBL_DVN_DESCREQ_NUM_CFG, ®_val,
+ sizeof(reg_val));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_dvn_init(struct nbl_hw_mgt *hw_mgt, u8 speed)
+{
+ u32 timeout = 0;
+ u32 ro_flag = 0;
+
+ nbl_hw_wr32(hw_mgt, NBL_DVN_ECPU_QUEUE_NUM, 0);
+ timeout = FIELD_PREP(NBL_DVN_DESC_WR_MERGE_TIMEOUT_CFG_CYCLE_MASK,
+ DEFAULT_DVN_DESC_WR_MERGE_TIMEOUT_MAX);
+ nbl_hw_wr_regs_lock(hw_mgt, NBL_DVN_DESC_WR_MERGE_TIMEOUT, &timeout,
+ sizeof(timeout));
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_DVN_DIF_REQ_RD_RO_FLAG, &ro_flag,
+ sizeof(ro_flag));
+ if (pcie_relaxed_ordering_enabled(hw_mgt->common->pdev)) {
+ ro_flag |=
+ FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DESC_RO_EN_MASK,
+ 1) |
+ FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DATA_RO_EN_MASK,
+ 1) |
+ FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_AVRING_RO_EN_MASK,
+ 1);
+ } else {
+ ro_flag &=
+ ~(FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DESC_RO_EN_MASK,
+ 1) |
+ FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_DATA_RO_EN_MASK,
+ 1) |
+ FIELD_PREP(NBL_DVN_DIF_REQ_RD_RO_FLAG_AVRING_RO_EN_MASK,
+ 1));
+ }
+ nbl_hw_wr_regs(hw_mgt, NBL_DVN_DIF_REQ_RD_RO_FLAG, &ro_flag,
+ sizeof(ro_flag));
+ spin_unlock(&hw_mgt->reg_lock);
+ if (speed == NBL_FW_PORT_SPEED_100G)
+ nbl_dvn_descreq_num_cfg(hw_mgt,
+ DEFAULT_DVN_100G_DESCREQ_NUMCFG);
+ else
+ nbl_dvn_descreq_num_cfg(hw_mgt, DEFAULT_DVN_DESCREQ_NUMCFG);
+}
+
+static void nbl_uvn_init(struct nbl_hw_mgt *hw_mgt)
+{
+ u16 wr_timeout = NBL_UVN_DESC_WR_TIMEOUT_VAL;
+ u32 timeout = NBL_UVN_DESC_RD_WAIT_TICKS;
+ u32 prefetch_init = 0;
+ bool ro_enabled;
+ u32 flag = 0;
+ u32 mask = 0;
+ u32 quirks;
+ u32 reg_val;
+
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_wr32(hw_mgt, NBL_UVN_ECPU_QUEUE_NUM, 0);
+ nbl_hw_wr32(hw_mgt, NBL_UVN_DESC_RD_WAIT, timeout);
+ nbl_hw_rd_regs(hw_mgt, NBL_UVN_DESC_WR_TIMEOUT,
+ ®_val, sizeof(reg_val));
+ reg_val &= ~NBL_UVN_DESC_WR_TIMEOUT_NUM_MASK;
+ reg_val |= FIELD_PREP(NBL_UVN_DESC_WR_TIMEOUT_NUM_MASK, wr_timeout);
+ nbl_hw_wr_regs(hw_mgt, NBL_UVN_DESC_WR_TIMEOUT, ®_val,
+ sizeof(reg_val));
+ ro_enabled = pcie_relaxed_ordering_enabled(hw_mgt->common->pdev);
+
+ nbl_hw_rd_regs(hw_mgt, NBL_UVN_DIF_REQ_RO_FLAG, &flag, sizeof(flag));
+ if (ro_enabled) {
+ flag |= FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_AVAIL_RD_MASK, 1) |
+ FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_RD_MASK, 1) |
+ FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_PKT_WR_MASK, 1);
+ flag &= ~FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_WR_MASK, 1);
+ } else {
+ flag &= ~(FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_AVAIL_RD_MASK, 1) |
+ FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_RD_MASK, 1) |
+ FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_PKT_WR_MASK, 1) |
+ FIELD_PREP(NBL_UVN_DIF_REQ_RO_FLAG_DESC_WR_MASK, 1));
+ }
+ nbl_hw_wr_regs(hw_mgt, NBL_UVN_DIF_REQ_RO_FLAG, &flag, sizeof(flag));
+
+ nbl_hw_rd_regs(hw_mgt, NBL_UVN_QUEUE_ERR_MASK, &mask, sizeof(mask));
+ mask |= FIELD_PREP(NBL_UVN_QUEUE_ERR_MASK_DIF_ERR_MASK, 1);
+
+ nbl_hw_wr_regs(hw_mgt, NBL_UVN_QUEUE_ERR_MASK, &mask, sizeof(mask));
+
+ spin_unlock(&hw_mgt->reg_lock);
+ quirks = nbl_hw_get_quirks(hw_mgt);
+ /*
+ * sel=0: use configured num; sel=1: use internal calc (max 32)
+ * Default is sel=1, unless NBL_QUIRK_UVN_PREFETCH_ALIGN is set,
+ * in which case override to sel=0.
+ */
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_rd_regs(hw_mgt, NBL_UVN_DESC_PREFETCH_INIT, ®_val,
+ sizeof(reg_val));
+ prefetch_init =
+ FIELD_PREP(NBL_UVN_DESC_PREFETCH_INIT_NUM_MASK,
+ NBL_UVN_DESC_PREFETCH_NUM) |
+ FIELD_PREP(NBL_UVN_DESC_PREFETCH_INIT_SEL_MASK,
+ (quirks & NBL_QUIRK_UVN_PREFETCH_ALIGN) ? 0 : 1);
+ reg_val &= ~(NBL_UVN_DESC_PREFETCH_INIT_NUM_MASK |
+ NBL_UVN_DESC_PREFETCH_INIT_SEL_MASK);
+ reg_val |= prefetch_init;
+ nbl_hw_wr_regs(hw_mgt, NBL_UVN_DESC_PREFETCH_INIT, ®_val,
+ sizeof(reg_val));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_uqm_init(struct nbl_hw_mgt *hw_mgt)
+{
+ u32 que_type = 0;
+ u32 cnt = 0;
+ int i;
+
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_wr_regs(hw_mgt, NBL_UQM_FWD_DROP_CNT, &cnt, sizeof(cnt));
+
+ nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_PKT_CNT, &cnt, sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_PKT_SLICE_CNT, &cnt, sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_PKT_LEN_ADD_CNT, &cnt, sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_HEAD_PNTR_ADD_CNT, &cnt,
+ sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt, NBL_UQM_DROP_WEIGHT_ADD_CNT, &cnt, sizeof(cnt));
+
+ for (i = 0; i < NBL_UQM_PORT_DROP_DEPTH; i++) {
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_UQM_PORT_DROP_PKT_CNT + (sizeof(cnt) * i),
+ &cnt, sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_UQM_PORT_DROP_PKT_SLICE_CNT +
+ (sizeof(cnt) * i),
+ &cnt, sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_UQM_PORT_DROP_PKT_LEN_ADD_CNT +
+ (sizeof(cnt) * i),
+ &cnt, sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_UQM_PORT_DROP_HEAD_PNTR_ADD_CNT +
+ (sizeof(cnt) * i),
+ &cnt, sizeof(cnt));
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_UQM_PORT_DROP_WEIGHT_ADD_CNT +
+ (sizeof(cnt) * i),
+ &cnt, sizeof(cnt));
+ }
+
+ for (i = 0; i < NBL_UQM_DPORT_DROP_DEPTH; i++)
+ nbl_hw_wr_regs(hw_mgt,
+ NBL_UQM_DPORT_DROP_CNT + (sizeof(cnt) * i), &cnt,
+ sizeof(cnt));
+ /* bit0: 0=bp mode, 1=drop mode, resv bit1-31 */
+ nbl_hw_wr_regs(hw_mgt, NBL_UQM_QUE_TYPE, &que_type, sizeof(que_type));
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static int nbl_dp_init(struct nbl_hw_mgt *hw_mgt, u8 speed, u8 eth_num)
+{
+ int ret;
+
+ nbl_dped_init(hw_mgt);
+ nbl_uped_init(hw_mgt);
+ ret = nbl_shaping_init(hw_mgt, speed);
+ if (ret)
+ return ret;
+ nbl_dsch_qid_max_init(hw_mgt);
+ ret = nbl_ustore_init(hw_mgt, eth_num);
+ if (ret)
+ return ret;
+ nbl_dstore_init(hw_mgt, speed);
+ nbl_dvn_init(hw_mgt, speed);
+ nbl_uvn_init(hw_mgt);
+ nbl_uqm_init(hw_mgt);
+ return 0;
+}
+
+static void nbl_host_padpt_init(struct nbl_hw_mgt *hw_mgt)
+{
+ /* padpt flow control register */
+ spin_lock(&hw_mgt->reg_lock);
+ nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_CPLH_UP,
+ NBL_HOST_PADPT_CFG_FC_CPLH_UP_VAL);
+ nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_PD_DN,
+ NBL_HOST_PADPT_CFG_FC_PD_DN_VAL);
+ nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_PH_DN,
+ NBL_HOST_PADPT_CFG_FC_PH_DN_VAL);
+ nbl_hw_wr32(hw_mgt, NBL_HOST_PADPT_HOST_CFG_FC_NPH_DN,
+ NBL_HOST_PADPT_CFG_FC_NPH_DN_VAL);
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+static void nbl_intf_init(struct nbl_hw_mgt *hw_mgt)
+{
+ nbl_host_padpt_init(hw_mgt);
+}
+
+static void nbl_hw_set_driver_status(struct nbl_hw_mgt *hw_mgt, bool active)
+{
+ u32 status;
+
+ spin_lock(&hw_mgt->reg_lock);
+ status = nbl_hw_rd32(hw_mgt, NBL_DRIVER_STATUS_REG);
+
+ status &= ~BIT(NBL_DRIVER_STATUS_BIT);
+ status |= FIELD_PREP(BIT(NBL_DRIVER_STATUS_BIT), active);
+
+ nbl_hw_wr32(hw_mgt, NBL_DRIVER_STATUS_REG, status);
+ spin_unlock(&hw_mgt->reg_lock);
+}
+
+/*
+ * Setting driver status to false notifies firmware to clean up per-PF
+ * hardware state such as qinfo registers.
+ *
+ * Note: firmware does NOT automatically revert chip-wide registers
+ * configured in this init flow. Those chip-wide settings remain valid
+ * until chip reset or explicitly overwritten by driver.
+ *
+ * This deinit_module only clears driver active status and flush writes.
+ * It does NOT reset or restore chip-wide datapath registers.
+ *
+ * Firmware cleanup is asynchronous with no completion status register.
+ * The flush only guarantees the status write reached the device; the
+ * bounded delay after it (NBL_FW_CLEANUP_SYNC_MIN_US) is best-effort and
+ * does not guarantee that the firmware pass has finished. The ordering
+ * this relies on is the device-link one: sibling PFs are unbound before
+ * the control PF, so a later init_module() requires an admin (or the
+ * core's AUTOPROBE re-probe) to bind this PF again, which is not
+ * instantaneous - but there is no hard firmware handshake.
+ *
+ * Caller must ensure no new DMA is initiated after this point.
+ * The mailbox channel is stopped by nbl_chan_teardown_queue()
+ * before this function is called, so no in-flight mailbox DMA
+ * remains.
+ */
+static void nbl_hw_deinit_module(struct nbl_hw_mgt *hw_mgt)
+{
+ nbl_hw_set_driver_status(hw_mgt, false);
+ /* ensure driver_status reaches the chip */
+ nbl_flush_writes(hw_mgt);
+ /*
+ * Give the asynchronous firmware cleanup pass time to settle
+ * before returning; best-effort, see NBL_FW_CLEANUP_SYNC_MIN_US.
+ */
+ usleep_range(NBL_FW_CLEANUP_SYNC_MIN_US, NBL_FW_CLEANUP_SYNC_MAX_US);
+}
+
+static bool nbl_hw_eth_speed_valid(u8 speed)
+{
+ switch (speed) {
+ case NBL_FW_PORT_SPEED_10G:
+ case NBL_FW_PORT_SPEED_25G:
+ case NBL_FW_PORT_SPEED_50G:
+ case NBL_FW_PORT_SPEED_100G:
+ return true;
+ default:
+ return false;
+ }
+}
+
+static bool nbl_hw_eth_num_valid(u8 eth_num)
+{
+ return eth_num == 1 || eth_num == 2 || eth_num == 4;
+}
+
+/*
+ * Full chip hardware initialization is handled by firmware.
+ * This function only configures driver-level table entries and registers.
+ *
+ * The driver_status handshake is entered from a known state: if the
+ * active bit is already set, a previous kernel instance left it behind
+ * (kexec, or an unload that never reached deinit_module()) and firmware
+ * never ran the per-PF cleanup it drives off the 1 -> 0 transition.
+ * That stale state is cleared - with the same bounded wait deinit_module()
+ * uses - before initializing on top of it.
+ */
+static int nbl_hw_init_module(struct nbl_hw_mgt *hw_mgt, u8 eth_speed,
+ u8 eth_num)
+{
+ u32 status;
+ int ret;
+
+ if (!nbl_hw_eth_speed_valid(eth_speed)) {
+ dev_err(hw_mgt->common->dev, "Invalid eth_speed %u\n",
+ eth_speed);
+ return -EINVAL;
+ }
+ if (!nbl_hw_eth_num_valid(eth_num)) {
+ dev_err(hw_mgt->common->dev, "Invalid eth_num %u\n", eth_num);
+ return -EINVAL;
+ }
+
+ status = nbl_hw_rd32(hw_mgt, NBL_DRIVER_STATUS_REG);
+ if (status & BIT(NBL_DRIVER_STATUS_BIT)) {
+ dev_warn(hw_mgt->common->dev,
+ "driver_status already set at init, cleanup forced\n");
+ nbl_hw_set_driver_status(hw_mgt, false);
+ nbl_flush_writes(hw_mgt);
+ usleep_range(NBL_FW_CLEANUP_SYNC_MIN_US,
+ NBL_FW_CLEANUP_SYNC_MAX_US);
+ }
+
+ ret = nbl_dp_init(hw_mgt, eth_speed, eth_num);
+ if (ret)
+ return ret;
+ nbl_intf_init(hw_mgt);
+ nbl_hw_set_driver_status(hw_mgt, true);
+ /* ensure registers written */
+ nbl_flush_writes(hw_mgt);
+
+ return 0;
+}
+
/*
* nbl_hw_set_mailbox_irq - read-modify-write NBL_MAILBOX_QINFO_MAP_REG_ARR
*
@@ -427,6 +1093,9 @@ static void nbl_hw_get_board_info(struct nbl_hw_mgt *hw_mgt,
}
static struct nbl_hw_ops hw_ops = {
+ .init_module = nbl_hw_init_module,
+ .deinit_module = nbl_hw_deinit_module,
+
.cfg_msix_map = nbl_hw_cfg_msix_map,
.cfg_msix_info = nbl_hw_cfg_msix_info,
.flush_write = nbl_flush_writes,
@@ -477,7 +1146,8 @@ static struct nbl_hw_ops_tbl *nbl_hw_setup_ops(struct nbl_common_info *common,
!hw_ops.stop_mailbox_rxq || !hw_ops.stop_mailbox_txq ||
!hw_ops.get_host_pf_mask || !hw_ops.get_real_bus ||
!hw_ops.cfg_mailbox_qinfo || !hw_ops.set_mailbox_irq ||
- !hw_ops.get_fw_eth_map || !hw_ops.get_board_info)
+ !hw_ops.get_fw_eth_map || !hw_ops.get_board_info ||
+ !hw_ops.init_module || !hw_ops.deinit_module)
return ERR_PTR(-EINVAL);
hw_ops_tbl->ops = &hw_ops;
hw_ops_tbl->priv = hw_mgt;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
index b7c0a87c1dad..e3f4faa1373c 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.h
@@ -11,12 +11,25 @@
#include "../../nbl_include/nbl_include.h"
#include "../nbl_hw_reg.h"
+#define NBL_DRIVER_STATUS_REG 0x1300444
+#define NBL_DRIVER_STATUS_BIT 16
+
/* ---------- REG BASE ADDR ---------- */
/* Interface modules base addr */
#define NBL_INTF_HOST_PCOMPLETER_BASE 0x00f08000
#define NBL_INTF_HOST_PADPT_BASE 0x00f4c000
#define NBL_INTF_HOST_MAILBOX_BASE 0x00fb0000
#define NBL_INTF_HOST_PCIE_BASE 0X01504000
+/* DP modules base addr */
+#define NBL_DP_USTORE_BASE 0x00104000
+#define NBL_DP_UQM_BASE 0x00114000
+#define NBL_DP_UPED_BASE 0x0015c000
+#define NBL_DP_UVN_BASE 0x00244000
+#define NBL_DP_DSCH_BASE 0x00404000
+#define NBL_DP_SHAPING_BASE 0x00504000
+#define NBL_DP_DVN_BASE 0x00514000
+#define NBL_DP_DSTORE_BASE 0x00704000
+#define NBL_DP_DPED_BASE 0x0075c000
/* -------- MAILBOX BAR2 ----- */
#define NBL_MAILBOX_NOTIFY_ADDR 0x00000000
#define NBL_MAILBOX_QINFO_CFG_RX_TABLE_ADDR 0x10
@@ -56,6 +69,17 @@ struct nbl_mailbox_qinfo_cfg_table {
#define NBL_PCIE_BUS_MASK GENMASK(12, 5)
/* -------- HOST_PADPT -------- */
+#define NBL_HOST_PADPT_HOST_CFG_FC_PD_DN (NBL_INTF_HOST_PADPT_BASE + 0x00000160)
+#define NBL_HOST_PADPT_HOST_CFG_FC_PH_DN (NBL_INTF_HOST_PADPT_BASE + 0x00000164)
+#define NBL_HOST_PADPT_HOST_CFG_FC_NPH_DN \
+ (NBL_INTF_HOST_PADPT_BASE + 0x0000016C)
+#define NBL_HOST_PADPT_HOST_CFG_FC_CPLH_UP \
+ (NBL_INTF_HOST_PADPT_BASE + 0x00000170)
+
+#define NBL_HOST_PADPT_CFG_FC_CPLH_UP_VAL 0x10400
+#define NBL_HOST_PADPT_CFG_FC_PD_DN_VAL 0x10080
+#define NBL_HOST_PADPT_CFG_FC_PH_DN_VAL 0x10010
+#define NBL_HOST_PADPT_CFG_FC_NPH_DN_VAL 0x10010
/* host_padpt host_msix_info */
#define NBL_PADPT_HOST_MSIX_INFO_REG_ARR(vector_id) \
(NBL_INTF_HOST_PADPT_BASE + 0x00010000 + \
@@ -98,6 +122,203 @@ struct nbl_function_msix_map {
u32 data[NBL_FUNC_MSIX_MAP_DWLEN];
};
+/* ---------- DPED ---------- */
+#define NBL_DPED_VLAN_OFFSET (NBL_DP_DPED_BASE + 0x000003F4)
+#define NBL_DPED_DSCP_OFFSET_0 (NBL_DP_DPED_BASE + 0x000003F8)
+#define NBL_DPED_DSCP_OFFSET_1 (NBL_DP_DPED_BASE + 0x000003FC)
+/* DPED hw_edt_prof/ UPED hw_edt_prof */
+
+#define NBL_DPED_L4_CK_CMD_40_ADDR (NBL_DP_DPED_BASE + 0x00000338)
+
+#define NBL_DPED_L4_CK_CMD_40_EN_MASK BIT(31)
+
+/* ---------- UPED ---------- */
+/* UPED uped_hw_edt_prof */
+#define NBL_UPED_HW_EDT_PROF_TABLE(i) \
+ (NBL_DP_UPED_BASE + 0x00001000 + (i) * sizeof(u32))
+#define NBL_UPED_V4_TCP_IDX 5
+#define NBL_UPED_V6_TCP_IDX 6
+#define NBL_PED_HW_EDIT_PROFILE_L3_LEN_MASK GENMASK(3, 2)
+
+/* ---------- DSCH ---------- */
+#define NBL_DSCH_PSHA_EN_MASK GENMASK(3, 0)
+/* DSCH dsch maxqid */
+#define NBL_DSCH_HOST_QID_MAX (NBL_DP_DSCH_BASE + 0x00000118)
+#define NBL_DSCH_HOST_QID_MAX_MASK GENMASK(10, 0)
+#define NBL_DSCH_VN_QUANTA_ADDR (NBL_DP_DSCH_BASE + 0x00000134)
+
+#define NBL_MAX_QUEUE_ID 0x7ff
+#define NBL_HOST_QUANTA 0x8000
+#define NBL_ECPU_QUANTA 0x1000
+
+#define NBL_DSCH_VN_QUANTA_H_QUA_MASK GENMASK(15, 0)
+#define NBL_DSCH_VN_QUANTA_E_QUA_MASK GENMASK(31, 16)
+
+/* ---------- DVN ---------- */
+/* DVN dvn_queue_table */
+#define NBL_DVN_ECPU_QUEUE_NUM (NBL_DP_DVN_BASE + 0x0000041C)
+#define NBL_DVN_DESCREQ_NUM_CFG (NBL_DP_DVN_BASE + 0x00000430)
+#define NBL_DVN_DESC_WR_MERGE_TIMEOUT (NBL_DP_DVN_BASE + 0x00000480)
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG (NBL_DP_DVN_BASE + 0x0000045C)
+
+#define DEFAULT_DVN_DESCREQ_NUMCFG 0x03
+#define DEFAULT_DVN_100G_DESCREQ_NUMCFG 0x07
+
+#define DEFAULT_DVN_DESC_WR_MERGE_TIMEOUT_MAX 0x3FF
+
+/* spilit ring descreq_num 0:8,1:16 */
+#define NBL_DVN_DESCREQ_NUM_CFG_AVRING_DESREQ_NUM_CFG_MASK BIT(0)
+/* packet ring descreq_num
+ * 0:8,1:12,2:16;3:20,4:24,5:26;6:32,7:32
+ */
+#define NBL_DVN_DESCREQ_NUM_CFG_PACKED_L1_NUM_MASK GENMASK(6, 4)
+
+#define NBL_DVN_DESC_WR_MERGE_TIMEOUT_CFG_CYCLE_MASK GENMASK(9, 0)
+
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG_DESC_RO_EN_MASK BIT(0)
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG_DATA_RO_EN_MASK BIT(1)
+#define NBL_DVN_DIF_REQ_RD_RO_FLAG_AVRING_RO_EN_MASK BIT(2)
+
+/* ---------- UVN ---------- */
+/* UVN uvn_queue_table */
+
+#define NBL_UVN_DESC_RD_WAIT (NBL_DP_UVN_BASE + 0x0000020C)
+#define NBL_UVN_QUEUE_ERR_MASK (NBL_DP_UVN_BASE + 0x00000224)
+#define NBL_UVN_ECPU_QUEUE_NUM (NBL_DP_UVN_BASE + 0x0000023C)
+#define NBL_UVN_DESC_WR_TIMEOUT (NBL_DP_UVN_BASE + 0x00000214)
+#define NBL_UVN_DIF_REQ_RO_FLAG (NBL_DP_UVN_BASE + 0x00000250)
+#define NBL_UVN_DESC_PREFETCH_INIT (NBL_DP_UVN_BASE + 0x00000204)
+#define NBL_UVN_DESC_PREFETCH_NUM 4
+
+#define NBL_UVN_DIF_REQ_RO_FLAG_AVAIL_RD_MASK BIT(0)
+#define NBL_UVN_DIF_REQ_RO_FLAG_DESC_RD_MASK BIT(1)
+#define NBL_UVN_DIF_REQ_RO_FLAG_PKT_WR_MASK BIT(2)
+#define NBL_UVN_DIF_REQ_RO_FLAG_DESC_WR_MASK BIT(3)
+
+#define NBL_UVN_DESC_WR_TIMEOUT_NUM_MASK GENMASK(14, 0)
+
+#define NBL_UVN_QUEUE_ERR_MASK_DIF_ERR_MASK BIT(5)
+
+#define NBL_UVN_DESC_PREFETCH_INIT_NUM_MASK GENMASK(7, 0)
+#define NBL_UVN_DESC_PREFETCH_INIT_SEL_MASK BIT(16)
+
+#define NBL_UVN_DESC_WR_TIMEOUT_VAL 0x12c
+/* 200us = 200000ns / 1.67ns per tick = 119760 ticks */
+#define NBL_UVN_DESC_RD_WAIT_TICKS 119760
+
+/* -------- USTORE -------- */
+#define NBL_USTORE_PKT_LEN_ADDR (NBL_DP_USTORE_BASE + 0x00000108)
+#define NBL_USTORE_PORT_DROP_TH_REG_ARR(port_id) \
+ (NBL_DP_USTORE_BASE + 0x00000150 + (port_id) * sizeof(u32))
+#define NBL_USTORE_BUF_PORT_DROP_PKT(eth_id) \
+ (NBL_DP_USTORE_BASE + 0x00002500 + (eth_id) * sizeof(u32))
+#define NBL_USTORE_BUF_PORT_TRUN_PKT(eth_id) \
+ (NBL_DP_USTORE_BASE + 0x00002540 + (eth_id) * sizeof(u32))
+
+#define NBL_USTORE_SINGLE_ETH_DROP_TH 0xC80
+#define NBL_USTORE_DUAL_ETH_DROP_TH 0x640
+#define NBL_USTORE_QUAD_ETH_DROP_TH 0x320
+
+/* USTORE pkt_len */
+#define NBL_USTORE_PKT_LEN_MIN_MASK GENMASK(6, 0)
+
+/* USTORE port_drop_th */
+#define NBL_USTORE_PORT_DROP_TH_DISC_TH_MASK GENMASK(11, 0)
+#define NBL_USTORE_PORT_DROP_TH_EN_MASK BIT(31)
+
+/* UQM*/
+#define NBL_UQM_QUE_TYPE (NBL_DP_UQM_BASE + 0x0000013c)
+#define NBL_UQM_DROP_PKT_CNT (NBL_DP_UQM_BASE + 0x000009C0)
+#define NBL_UQM_DROP_PKT_SLICE_CNT (NBL_DP_UQM_BASE + 0x000009C4)
+#define NBL_UQM_DROP_PKT_LEN_ADD_CNT (NBL_DP_UQM_BASE + 0x000009C8)
+#define NBL_UQM_DROP_HEAD_PNTR_ADD_CNT (NBL_DP_UQM_BASE + 0x000009CC)
+#define NBL_UQM_DROP_WEIGHT_ADD_CNT (NBL_DP_UQM_BASE + 0x000009D0)
+#define NBL_UQM_PORT_DROP_PKT_CNT (NBL_DP_UQM_BASE + 0x000009D4)
+#define NBL_UQM_PORT_DROP_PKT_SLICE_CNT (NBL_DP_UQM_BASE + 0x000009F4)
+#define NBL_UQM_PORT_DROP_PKT_LEN_ADD_CNT (NBL_DP_UQM_BASE + 0x00000A14)
+#define NBL_UQM_PORT_DROP_HEAD_PNTR_ADD_CNT (NBL_DP_UQM_BASE + 0x00000A34)
+#define NBL_UQM_PORT_DROP_WEIGHT_ADD_CNT (NBL_DP_UQM_BASE + 0x00000A54)
+#define NBL_UQM_FWD_DROP_CNT (NBL_DP_UQM_BASE + 0x00000A80)
+#define NBL_UQM_DPORT_DROP_CNT (NBL_DP_UQM_BASE + 0x00000B74)
+
+#define NBL_UQM_PORT_DROP_DEPTH 6
+#define NBL_UQM_DPORT_DROP_DEPTH 16
+
+/* --------- SHAPING --------- */
+
+/* Shaping rate unit: 1 = 1 Mbit/s.
+ * e.g. 100000 = 100 Gbit/s, 25000 = 25 Gbit/s.
+ */
+#define NBL_LR_LEONIS_NET_BUCKET_DEPTH 9600
+#define NBL_SHAPING_DPORT_ADDR (NBL_DP_SHAPING_BASE + 0x700)
+#define NBL_SHAPING_DPORT_DWLEN 4
+#define NBL_SHAPING_DPORT_REG(r) \
+ (NBL_SHAPING_DPORT_ADDR + (NBL_SHAPING_DPORT_DWLEN * 4) * (r))
+#define NBL_SHAPING_DVN_DPORT_ADDR (NBL_DP_SHAPING_BASE + 0x750)
+#define NBL_SHAPING_DVN_DPORT_DWLEN 4
+#define NBL_SHAPING_DVN_DPORT_REG(r) \
+ (NBL_SHAPING_DVN_DPORT_ADDR + (NBL_SHAPING_DVN_DPORT_DWLEN * 4) * (r))
+#define NBL_DSCH_PSHA_EN_ADDR (NBL_DP_DSCH_BASE + 0x00000314)
+#define NBL_SHAPING_NET_ADDR (NBL_DP_SHAPING_BASE + 0x1800)
+#define NBL_SHAPING_NET_DWLEN 4
+#define NBL_SHAPING_NET_REG(r) \
+ (NBL_SHAPING_NET_ADDR + (NBL_SHAPING_NET_DWLEN * 4) * (r))
+
+#define NBL_DPORT_VALID_MASK GENMASK_ULL(0, 0)
+#define NBL_DPORT_DEPTH_MASK GENMASK_ULL(19, 1)
+#define NBL_DPORT_CIR_MASK GENMASK_ULL(38, 20)
+#define NBL_DPORT_PIR_MASK GENMASK_ULL(57, 39)
+#define NBL_DPORT_CBS_MASK_LOW GENMASK_ULL(63, 58)
+#define NBL_DPORT_CBS_MASK_HIGH GENMASK_ULL(14, 0)
+#define NBL_DPORT_PBS_MASK GENMASK_ULL(35, 15)
+
+/* SHAPING shaping_net */
+struct nbl_shaping_net_u {
+ u32 data[NBL_SHAPING_NET_DWLEN];
+};
+
+struct nbl_shaping_dport_u {
+ u32 data[NBL_SHAPING_DPORT_DWLEN];
+};
+
+struct nbl_shaping_dvn_dport_u {
+ u32 data[NBL_SHAPING_DVN_DPORT_DWLEN];
+};
+
+/* -------- DSTORE -------- */
+#define NBL_DSTORE_D_DPORT_FC_TH_ADDR (NBL_DP_DSTORE_BASE + 0x00000600)
+#define NBL_DSTORE_D_DPORT_FC_TH_DEPTH 5
+#define NBL_DSTORE_D_DPORT_FC_TH_WIDTH 32
+#define NBL_DSTORE_D_DPORT_FC_TH_DWLEN 1
+
+#define NBL_DSTORE_D_DPORT_FC_XOFF_TH_MASK GENMASK(10, 0)
+#define NBL_DSTORE_D_DPORT_FC_XON_TH_MASK GENMASK(26, 16)
+#define NBL_DSTORE_D_DPORT_FC_FC_EN_MASK BIT(31)
+
+#define NBL_DSTORE_D_DPORT_FC_TH_REG(r) \
+ (NBL_DSTORE_D_DPORT_FC_TH_ADDR + \
+ (NBL_DSTORE_D_DPORT_FC_TH_DWLEN * 4) * (r))
+#define NBL_DSTORE_PORT_DROP_TH_ADDR (NBL_DP_DSTORE_BASE + 0x00000150)
+#define NBL_DSTORE_PORT_DROP_TH_DEPTH 6
+#define NBL_DSTORE_PORT_DROP_TH_WIDTH 32
+#define NBL_DSTORE_PORT_DROP_TH_DWLEN 1
+
+#define NBL_DSTORE_PORT_DROP_EN_MASK BIT(31)
+
+#define NBL_DSTORE_DROP_XOFF_TH 0xC8
+#define NBL_DSTORE_DROP_XON_TH 0x64
+
+#define NBL_DSTORE_DROP_XOFF_TH_100G 0x1F4
+#define NBL_DSTORE_DROP_XON_TH_100G 0x12C
+
+#define NBL_DSTORE_DISC_BP_TH (NBL_DP_DSTORE_BASE + 0x00000630)
+
+#define NBL_DSTORE_DISC_BP_TH_EN_MASK BIT(31)
+
+#define NBL_DSTORE_PORT_DROP_TH_REG(r) \
+ (NBL_DSTORE_PORT_DROP_TH_ADDR + \
+ (NBL_DSTORE_PORT_DROP_TH_DWLEN * 4) * (r))
+
#define NBL_FW_BOARD_CONFIG 0x200
#define NBL_FW_BOARD_DW3_OFFSET (NBL_FW_BOARD_CONFIG + 12)
#define NBL_FW_BOARD_DW6_OFFSET (NBL_FW_BOARD_CONFIG + 24)
@@ -111,5 +332,6 @@ struct nbl_function_msix_map {
#define NBL_FW_BOARD_DW6_LANE_BITMAP_MASK GENMASK(7, 0)
#define NBL_FW_BOARD_DW6_ETH_BITMAP_MASK GENMASK(15, 8)
+#define NBL_LEONIS_QUIRKS_OFFSET 0x00000140
#define NBL_BAR2_MAX_LEN 0x300
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
index 72b5b889dedb..6cb9302af7bc 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.c
@@ -13,6 +13,8 @@ static struct nbl_resource_ops res_ops = {
.cfg_msix_map = nbl_res_intr_cfg_msix_map,
.destroy_msix_map = nbl_res_intr_destroy_msix_map,
.set_mailbox_irq = nbl_res_intr_set_mailbox_irq,
+ .init_module = nbl_res_chip_init_module,
+ .deinit_module = nbl_res_chip_deinit_module,
};
static struct nbl_resource_mgt *
@@ -46,7 +48,8 @@ nbl_res_setup_ops(struct device *dev, struct nbl_resource_mgt *res_mgt)
return ERR_PTR(-ENOMEM);
if (!res_ops.get_vsi_id || !res_ops.get_eth_id ||
!res_ops.cfg_msix_map || !res_ops.destroy_msix_map ||
- !res_ops.set_mailbox_irq)
+ !res_ops.set_mailbox_irq || !res_ops.init_module ||
+ !res_ops.deinit_module)
return ERR_PTR(-EINVAL);
res_ops_tbl->ops = &res_ops;
res_ops_tbl->priv = res_mgt;
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
index 6eb4dc9e695a..f1cb0f23240b 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_resource_leonis.h
@@ -8,4 +8,5 @@
#include "../nbl_resource.h"
#include "../nbl_interrupt.h"
+#include "../nbl_chip.h"
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
index 89b979e2cdb5..fe6771be1eb8 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_hw.h
@@ -12,6 +12,9 @@ struct nbl_board_port_info;
struct nbl_hw_mgt;
struct nbl_adapter;
struct nbl_hw_ops {
+ int (*init_module)(struct nbl_hw_mgt *hw_mgt, u8 eth_speed, u8 eth_num);
+ void (*deinit_module)(struct nbl_hw_mgt *hw_mgt);
+
void (*cfg_msix_map)(struct nbl_hw_mgt *hw_mgt, u16 func_id,
bool valid, dma_addr_t dma_addr, u8 bus,
u8 devid, u8 function);
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
index e718ea41a816..8dc64e806c1e 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_resource.h
@@ -12,6 +12,9 @@ struct nbl_resource_mgt;
struct nbl_adapter;
struct nbl_resource_ops {
+ int (*init_module)(struct nbl_resource_mgt *res_mgt);
+ void (*deinit_module)(struct nbl_resource_mgt *res_mgt);
+
int (*cfg_msix_map)(struct nbl_resource_mgt *res_mgt, u16 func_id,
u16 num_net_msix, u16 num_others_msix,
bool net_msix_mask_en);
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
index 59e44feab44f..efc4c79a082d 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_include.h
@@ -29,4 +29,23 @@ struct nbl_init_param {
struct nbl_func_caps caps;
};
+/*
+ * Firmware ABI defines port speed enum fixed, value 0 represents 10G, cannot
+ * reassign 0 to INVALID for compatibility
+ */
+enum nbl_fw_port_speed {
+ NBL_FW_PORT_SPEED_10G,
+ NBL_FW_PORT_SPEED_25G,
+ NBL_FW_PORT_SPEED_50G,
+ NBL_FW_PORT_SPEED_100G,
+};
+
+/*
+ * Firmware quirk word @ NBL_LEONIS_QUIRKS_OFFSET (0x140)
+ * Sentinel value: ~0U (0xFFFFFFFF) = firmware reports no active quirks
+ * BIT(1): NBL_QUIRK_UVN_PREFETCH_ALIGN – control UVN descriptor prefetch
+ * selection
+ */
+#define NBL_QUIRK_UVN_PREFETCH_ALIGN BIT(1)
+
#endif
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation
2026-10-10 9:59 ` [PATCH v31 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found 7 potential
issue(s) to consider.
Critical: 0 · High: 1 · Medium: 2 · Low: 4
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] The stale-state recovery in nbl_hw_init_module() can wipe the
per-PF mailbox state the driver has just programmed.
- [Medium] The only synchronization between the driver and firmware's
asynchronous driver_status cleanup is a fixed usleep_range(2000,…
- [Medium] nbl_shaping_init() clears the shaping_net table
(NBL_SHAPING_NET_REG(i)) only for i < NBL_MAX_PF (8), yet the loop…
- [Low] The comment on nbl_hw_deinit_module() (lines 730-734) says the
safety of a later init_module() rests on the device link, and that the…
- [Low] The commit message covers only the driver_status handshake and
the 2-3 ms best-effort wait.
- [Low] Baseline decision: false positive, because the series fixes it.
- [Low] Baseline decision: false positive.
--
net/nebula-matrix: add chip-wide hardware init/deinit implementation
This adds Leonis chip-wide init_module and deinit_module hooks in the hw
and resource layers. deinit_module clears the driver_status bit, flushes
and sleeps 2-3 ms to give the asynchronous firmware cleanup time to run.
init_module programs the datapath registers. If it finds driver_status
already set by a previous instance, it first clears the bit and waits for
the same window.
> Add Leonis chip-wide datapath init/deinit hooks (init_module and
> deinit_module) in the hw and resource layers. The hooks use the
> driver_status register as a handshake with firmware.
[Severity: Low]
This isn't a bug, but the commit message only describes the driver_status
handshake and the 2-3 ms wait. Most of the patch is chip-wide datapath
programming in nbl_dp_init() and nbl_intf_init(), and several of those
writes change device behaviour:
- SCTP L4 checksum enable in DPED
- UPED TCP L3-length edit profiles
- per-port CIR/PIR shaping from the firmware speed (DVN dport CIR at
half rate), and PSHA enable
- DSCH host QID max (0x7ff) and quanta
- USTORE min packet length 42 and drop thresholds
- DSTORE drop-enable clear and XON/XOFF thresholds
- relaxed ordering bits taken from DEVCTL
- UQM counter clearing and que_type forced to backpressure mode
- PADPT flow control credits
The patch also adds a new firmware ABI dependency in nbl_hw_get_quirks(). It
reads a quirk word at mailbox offset 0x140, where ~0U means no quirks and
BIT(1) selects the UVN prefetch mode.
deinit_module also doesn't revert any of these datapath registers. They stay
programmed until chip reset, even if the control PF is later bound to
another driver.
Could the commit message describe these?
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
> new file mode 100644
> index 0000000000000..419eb6392ada1
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
> @@ -0,0 +1,32 @@
[ ... ]
> +int nbl_res_chip_init_module(struct nbl_resource_mgt *res_mgt)
> +{
[ ... ]
> + hw_ops = res_mgt->hw_ops_tbl->ops;
> + p = res_mgt->hw_ops_tbl->priv;
> + return hw_ops->init_module(p, eth_speed, eth_num);
> +}
[Severity: Low]
This isn't a bug, but nothing calls init_module or deinit_module at this
point in the series. nbl_hw_setup_ops() and nbl_res_setup_ops() only check
the new hooks for NULL.
Two later commits wire them up:
- "net/nebula-matrix: dispatch: add control-level routing core
infrastructure" adds nbl_disp_init_module() and
nbl_disp_deinit_module()
- "net/nebula-matrix: add common/ctrl dev init/remove operation" calls
them from nbl_dev_setup_ctrl_dev() and nbl_dev_remove_ctrl_dev()
So the final state of the series looks fine. nbl_hw_init_module() also
validates eth_num with nbl_hw_eth_num_valid().
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> index 1712cdbc5fe74..bd1216265036b 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c
[ ... ]
> @@ -150,6 +164,658 @@ static void nbl_hw_get_fw_eth_map(struct nbl_hw_mgt *hw_mgt, u32 *eth_map)
[ ... ]
> +static int nbl_shaping_init(struct nbl_hw_mgt *hw_mgt, u8 speed)
> +{
> +#define NBL_SHAPING_FLUSH_INTERVAL 128
[ ... ]
> + for (i = 0; i < NBL_MAX_PF; i++) {
> + nbl_hw_wr_regs_lock(hw_mgt, NBL_SHAPING_NET_REG(i),
> + net_shaping.data,
> + sizeof(net_shaping));
> + if ((i + 1) % NBL_SHAPING_FLUSH_INTERVAL == 0)
> + nbl_flush_writes(hw_mgt);
> + }
> + nbl_flush_writes(hw_mgt);
> + return 0;
> +}
[Severity: Medium]
Is NBL_MAX_PF the right bound for the shaping_net table here?
NBL_MAX_PF is 8 and NBL_SHAPING_FLUSH_INTERVAL is 128, so (i + 1) % 128 == 0
is never true and the periodic flush is dead code. Flushing every 128 writes
suggests a table with hundreds of entries.
The other tables this patch sweeps have depth macros
(NBL_DSTORE_PORT_DROP_TH_DEPTH, NBL_UQM_PORT_DROP_DEPTH and
NBL_UQM_DPORT_DROP_DEPTH). shaping_net has none, so the bound can't be
checked against the register map.
If the table is deeper than 8, would entries above index 7 keep whatever a
previous instance or firmware left there across kexec or rebind? That could
include VALID set with stale CIR/PIR.
[ ... ]
> +/*
> + * Setting driver status to false notifies firmware to clean up per-PF
> + * hardware state such as qinfo registers.
> + *
> + * Note: firmware does NOT automatically revert chip-wide registers
> + * configured in this init flow. Those chip-wide settings remain valid
> + * until chip reset or explicitly overwritten by driver.
[ ... ]
> + * does not guarantee that the firmware pass has finished. The ordering
> + * this relies on is the device-link one: sibling PFs are unbound before
> + * the control PF, so a later init_module() requires an admin (or the
> + * core's AUTOPROBE re-probe) to bind this PF again, which is not
> + * instantaneous - but there is no hard firmware handshake.
[Severity: Low]
Is the AUTOPROBE part of this comment accurate?
init_module and deinit_module only run on the control PF, because of the
has_ctrl checks in nbl_chip.c. nbl_probe_chip_deps() in nbl_main.c returns
early when has_ctrl is set. For the siblings it does:
link = device_link_add(&pdev->dev, &mgt->dev,
DL_FLAG_AUTOPROBE_CONSUMER);
This makes the siblings consumers and func 0 the supplier.
DL_FLAG_AUTOPROBE_CONSUMER only re-probes consumers, so it never rebinds the
control PF.
The link also doesn't order the control PF's own deinit against its next
init. Only the 2-3 ms sleep does that.
The two comments in this file also disagree. The one above
NBL_FW_CLEANUP_SYNC_MIN_US says:
Bounding the wait keeps a later init_module() from reprogramming
per-PF/chip-wide registers while firmware is realistically still
wiping them
This comment says firmware does not automatically revert chip-wide
registers. Which one describes what firmware actually cleans?
> + *
> + * Caller must ensure no new DMA is initiated after this point.
> + * The mailbox channel is stopped by nbl_chan_teardown_queue()
> + * before this function is called, so no in-flight mailbox DMA
> + * remains.
> + */
> +static void nbl_hw_deinit_module(struct nbl_hw_mgt *hw_mgt)
> +{
> + nbl_hw_set_driver_status(hw_mgt, false);
> + /* ensure driver_status reaches the chip */
> + nbl_flush_writes(hw_mgt);
> + /*
> + * Give the asynchronous firmware cleanup pass time to settle
> + * before returning; best-effort, see NBL_FW_CLEANUP_SYNC_MIN_US.
> + */
> + usleep_range(NBL_FW_CLEANUP_SYNC_MIN_US, NBL_FW_CLEANUP_SYNC_MAX_US);
> +}
[ ... ]
> +static int nbl_hw_init_module(struct nbl_hw_mgt *hw_mgt, u8 eth_speed,
> + u8 eth_num)
> +{
[ ... ]
> + status = nbl_hw_rd32(hw_mgt, NBL_DRIVER_STATUS_REG);
[Severity: Low]
This probably isn't an issue in practice, but bit 16 is tested here with an
unlocked read. The clear below is a separate read-modify-write under
reg_lock in nbl_hw_set_driver_status().
On the driver side, both only run from the serialized control PF probe and
remove paths. Only the driver sets bit 16, so a stale read at worst causes a
redundant clear and sleep.
Does firmware ever write other bits of NBL_DRIVER_STATUS_REG? If it does,
could the read-modify-write in nbl_hw_set_driver_status() lose those
updates?
> + if (status & BIT(NBL_DRIVER_STATUS_BIT)) {
> + dev_warn(hw_mgt->common->dev,
> + "driver_status already set at init, cleanup forced\n");
> + nbl_hw_set_driver_status(hw_mgt, false);
> + nbl_flush_writes(hw_mgt);
> + usleep_range(NBL_FW_CLEANUP_SYNC_MIN_US,
> + NBL_FW_CLEANUP_SYNC_MAX_US);
> + }
[Severity: High]
Can this forced 1 -> 0 transition wipe the mailbox state the driver has just
programmed?
The deinit_module comment above says firmware reacts to driver_status going
to 0 by cleaning per-PF state such as qinfo registers. Later in the series,
nbl_dev_remove_ctrl_dev() relies on the same behaviour:
The qinfo map programmed in nbl_dev_init() via
nbl_dev_setup_chan_qinfo() is not explicitly cleared; firmware
handles it on driver_status change
In the final state of the series, nbl_dev_init() runs in this order:
nbl_dev_init()
nbl_dev_setup_chan_qinfo()
nbl_chan_cfg_qinfo_map_table()
hw_ops->cfg_mailbox_qinfo() /* for every PF */
nbl_dev_setup_common_dev() /* mailbox queue setup */
nbl_dev_setup_ctrl_dev()
nbl_disp_init_module()
nbl_res_chip_init_module()
nbl_hw_init_module()
nbl_hw_set_driver_status(hw_mgt, false)
So the stale clear happens after the mailbox QINFO map and queue registers
are programmed, and nothing reprograms them afterwards. This function then
sets driver_status back to 1 and returns 0.
The driver has no .shutdown callback, so every kexec reaches this path. The
commit message names kexec as a case init_module handles.
If firmware clears the QINFO routing here, would the inter-PF mailbox stop
working? The sibling PFs' cross-PF requests would then time out and their
probes would fail.
Would it work to do the stale driver_status clear before any per-PF
programming in nbl_dev_init(), or to reprogram the mailbox state after it?
> +
> + ret = nbl_dp_init(hw_mgt, eth_speed, eth_num);
> + if (ret)
> + return ret;
> + nbl_intf_init(hw_mgt);
> + nbl_hw_set_driver_status(hw_mgt, true);
[Severity: Medium]
The only synchronization with the asynchronous firmware cleanup is a fixed
usleep_range(2000, 3000). It is used here and in nbl_hw_deinit_module().
On this stale path, the bit stays low only for that sleep plus the
nbl_dp_init() and nbl_intf_init() register writes, and then it is set again
here. If firmware samples driver_status less often than that, can it miss
the 0 entirely? The per-PF cleanup this path is meant to trigger would then
never run.
If firmware does see the 0 but its cleanup takes longer than the window, can
the cleanup overlap the new instance's programming? That includes
driver_status=1 here and the MSI-X map and mailbox IRQ setup later in
nbl_dev_start(). Neither case would be reported.
On the deinit side, only elapsed time separates a quick unbind and rebind
of the control PF from an unfinished firmware pass.
Does firmware provide an acknowledgement that could be polled here, or a
documented upper bound on its polling period and cleanup latency?
> + /* ensure registers written */
> + nbl_flush_writes(hw_mgt);
> +
> + return 0;
> +}
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v31 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
` (3 preceding siblings ...)
2026-10-10 9:59 ` [PATCH v31 net-next 4/8] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
2026-10-10 9:59 ` [PATCH v31 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops illusion.wang
` (2 subsequent siblings)
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Allocate dispatch management state and ops table, add ctrl_lvl bitmap
to track control privileges, and hook init_module/deinit_module wrappers
to resource ops.
MGT level is only enabled for Control PF. Add kerneldoc for caller
permission constraints. Add wire ABI structures with static_assert for
upcoming mailbox RPC.
Dispatch objects are devm-managed and require no explicit remove logic.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../net/ethernet/nebula-matrix/nbl/Makefile | 1 +
.../net/ethernet/nebula-matrix/nbl/nbl_core.h | 2 +
.../nebula-matrix/nbl/nbl_core/nbl_dispatch.c | 117 ++++++++++++++++++
.../nebula-matrix/nbl/nbl_core/nbl_dispatch.h | 23 ++++
.../nbl/nbl_include/nbl_def_dispatch.h | 56 +++++++++
.../net/ethernet/nebula-matrix/nbl/nbl_main.c | 8 ++
6 files changed, 207 insertions(+)
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index be314b909d66..b7eebd89b4d1 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -10,4 +10,5 @@ nbl-objs += nbl_common/nbl_common.o \
nbl_hw/nbl_resource.o \
nbl_hw/nbl_interrupt.o \
nbl_hw/nbl_chip.o \
+ nbl_core/nbl_dispatch.o \
nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index dd24ebec0171..4d8cea8d8ab3 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -17,12 +17,14 @@ enum {
struct nbl_interface {
struct nbl_hw_ops_tbl *hw_ops_tbl;
struct nbl_resource_ops_tbl *resource_ops_tbl;
+ struct nbl_dispatch_ops_tbl *dispatch_ops_tbl;
struct nbl_channel_ops_tbl *channel_ops_tbl;
};
struct nbl_core {
struct nbl_hw_mgt *hw_mgt;
struct nbl_resource_mgt *res_mgt;
+ struct nbl_dispatch_mgt *disp_mgt;
struct nbl_channel_mgt *chan_mgt;
};
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
new file mode 100644
index 000000000000..966fee2dec8b
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
@@ -0,0 +1,117 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/pci.h>
+#include "nbl_dispatch.h"
+
+static void nbl_disp_deinit_module(struct nbl_dispatch_mgt *disp_mgt)
+{
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+
+ if (res_ops->deinit_module)
+ res_ops->deinit_module(p);
+}
+
+static int nbl_disp_init_module(struct nbl_dispatch_mgt *disp_mgt)
+{
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+
+ if (res_ops->init_module)
+ return res_ops->init_module(p);
+ return -EOPNOTSUPP;
+}
+
+static void nbl_disp_set_ctrl_bit(struct nbl_dispatch_mgt *disp_mgt, u32 lvl)
+{
+ set_bit(lvl, disp_mgt->ctrl_lvl);
+}
+
+static void nbl_disp_refresh_ctrl_ops(struct nbl_dispatch_mgt *disp_mgt)
+{
+ struct nbl_dispatch_ops *disp_ops = disp_mgt->disp_ops_tbl->ops;
+
+ if (test_bit(NBL_DISP_CTRL_LVL_MGT, disp_mgt->ctrl_lvl)) {
+ disp_ops->init_module = nbl_disp_init_module;
+ disp_ops->deinit_module = nbl_disp_deinit_module;
+ }
+}
+
+static struct nbl_dispatch_mgt *
+nbl_disp_setup_disp_mgt(struct nbl_common_info *common)
+{
+ struct nbl_dispatch_mgt *disp_mgt;
+ struct device *dev = common->dev;
+
+ disp_mgt = devm_kzalloc(dev, sizeof(*disp_mgt), GFP_KERNEL);
+ if (!disp_mgt)
+ return ERR_PTR(-ENOMEM);
+
+ disp_mgt->common = common;
+ return disp_mgt;
+}
+
+static struct nbl_dispatch_ops_tbl *
+nbl_disp_setup_ops(struct device *dev, struct nbl_dispatch_mgt *disp_mgt)
+{
+ struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+ struct nbl_dispatch_ops *disp_ops;
+
+ disp_ops_tbl = devm_kzalloc(dev, sizeof(*disp_ops_tbl), GFP_KERNEL);
+ if (!disp_ops_tbl)
+ return ERR_PTR(-ENOMEM);
+
+ disp_ops = devm_kzalloc(dev, sizeof(*disp_ops), GFP_KERNEL);
+ if (!disp_ops)
+ return ERR_PTR(-ENOMEM);
+
+ disp_ops_tbl->ops = disp_ops;
+ disp_ops_tbl->priv = disp_mgt;
+
+ return disp_ops_tbl;
+}
+
+int nbl_disp_init(struct nbl_adapter *adapter)
+{
+ struct nbl_common_info *common = &adapter->common;
+ struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+ struct nbl_resource_ops_tbl *res_ops_tbl =
+ adapter->intf.resource_ops_tbl;
+ struct nbl_channel_ops_tbl *chan_ops_tbl =
+ adapter->intf.channel_ops_tbl;
+ struct device *dev = &adapter->pdev->dev;
+ struct nbl_dispatch_mgt *disp_mgt;
+ int ret;
+
+ disp_mgt = nbl_disp_setup_disp_mgt(common);
+ if (IS_ERR(disp_mgt)) {
+ ret = PTR_ERR(disp_mgt);
+ return ret;
+ }
+
+ disp_ops_tbl = nbl_disp_setup_ops(dev, disp_mgt);
+ if (IS_ERR(disp_ops_tbl)) {
+ ret = PTR_ERR(disp_ops_tbl);
+ return ret;
+ }
+
+ disp_mgt->res_ops_tbl = res_ops_tbl;
+ disp_mgt->chan_ops_tbl = chan_ops_tbl;
+ disp_mgt->disp_ops_tbl = disp_ops_tbl;
+ adapter->core.disp_mgt = disp_mgt;
+ adapter->intf.dispatch_ops_tbl = disp_ops_tbl;
+
+ if (common->has_ctrl)
+ nbl_disp_set_ctrl_bit(disp_mgt, NBL_DISP_CTRL_LVL_MGT);
+
+ nbl_disp_refresh_ctrl_ops(disp_mgt);
+ return 0;
+}
+
+void nbl_disp_remove(struct nbl_adapter *adapter)
+{
+ /* Dispatch structures are allocated via devm */
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
new file mode 100644
index 000000000000..a7e5802344b4
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.h
@@ -0,0 +1,23 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DISPATCH_H_
+#define _NBL_DISPATCH_H_
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_resource.h"
+#include "../nbl_include/nbl_def_dispatch.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+struct nbl_dispatch_mgt {
+ struct nbl_common_info *common;
+ struct nbl_resource_ops_tbl *res_ops_tbl;
+ struct nbl_channel_ops_tbl *chan_ops_tbl;
+ struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+ DECLARE_BITMAP(ctrl_lvl, NBL_DISP_CTRL_LVL_MAX);
+};
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
new file mode 100644
index 000000000000..08fef0f67926
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
@@ -0,0 +1,56 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_DISPATCH_H_
+#define _NBL_DEF_DISPATCH_H_
+
+#include <linux/types.h>
+
+struct nbl_dispatch_mgt;
+struct nbl_adapter;
+enum {
+ NBL_DISP_CTRL_LVL_MGT,
+ NBL_DISP_CTRL_LVL_MAX,
+};
+
+/**
+ * struct nbl_dispatch_ops - dispatch control plane operation callbacks
+ * @init_module: chip-wide datapath init, forwarded to the resource layer
+ * (res_ops->init_module() and then hw_ops->init_module());
+ * on leonis it programs the chip-wide datapath registers and
+ * sets driver_status=true. Returns -EOPNOTSUPP if the
+ * resource layer installs no init_module op. Control-PF
+ * exclusive: the pointer is NULL on a function without
+ * has_ctrl, so the caller must check has_ctrl first.
+ * @deinit_module: chip-wide teardown, forwarded the same way; on leonis
+ * it clears driver_status, which starts an asynchronous
+ * firmware per-PF cleanup pass with no completion status
+ * bit, waited for best-effort only. The caller must have
+ * stopped the mailbox channel and drained all in-flight
+ * mailbox DMA before calling it, see nbl_dev_remove().
+ * Returns void. Control-PF exclusive, same NULL rule as
+ * @init_module.
+ *
+ * Warning: init_module/deinit_module are control-PF exclusive. The five
+ * resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq,
+ * get_vsi_id, get_eth_id) are PF-only and resolve to either a local
+ * resource call (control PF) or a mailbox RPC (non-control PF with
+ * has_net). A function with neither has_ctrl nor has_net leaves these
+ * pointers NULL; callers must not invoke them on such functions. VFs are
+ * rejected by the responders with -EPERM.
+ */
+struct nbl_dispatch_ops {
+ int (*init_module)(struct nbl_dispatch_mgt *disp_mgt);
+ void (*deinit_module)(struct nbl_dispatch_mgt *disp_mgt);
+};
+
+struct nbl_dispatch_ops_tbl {
+ struct nbl_dispatch_ops *ops;
+ struct nbl_dispatch_mgt *priv;
+};
+
+int nbl_disp_init(struct nbl_adapter *adapter);
+void nbl_disp_remove(struct nbl_adapter *adapter);
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index cd3774691bf4..86536d582d92 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -11,6 +11,7 @@
#include "nbl_include/nbl_def_channel.h"
#include "nbl_include/nbl_def_hw.h"
#include "nbl_include/nbl_def_resource.h"
+#include "nbl_include/nbl_def_dispatch.h"
#include "nbl_include/nbl_def_common.h"
#include "nbl_core.h"
@@ -48,7 +49,13 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
ret = nbl_res_init_leonis(adapter);
if (ret)
goto res_init_fail;
+
+ ret = nbl_disp_init(adapter);
+ if (ret)
+ goto disp_init_fail;
return adapter;
+disp_init_fail:
+ nbl_res_remove_leonis(adapter);
res_init_fail:
nbl_chan_remove_common(adapter);
chan_init_fail:
@@ -59,6 +66,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
void nbl_core_remove(struct nbl_adapter *adapter)
{
+ nbl_disp_remove(adapter);
nbl_res_remove_leonis(adapter);
nbl_chan_remove_common(adapter);
nbl_hw_remove_leonis(adapter);
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure
2026-10-10 9:59 ` [PATCH v31 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] The commit message says this patch will "Add wire ABI structures
with static_assert for upcoming mailbox RPC", but the patch adds…
- [Low] At this commit, the kernel-doc on struct nbl_dispatch_ops in
nbl_def_dispatch.h (lines 18-43) describes things that are not in the…
--
net/nebula-matrix: dispatch: add control-level routing core infrastructure
This commit allocates the dispatch management state and ops table, adds a
ctrl_lvl bitmap to track control privileges, and wires init_module and
deinit_module wrappers to the resource ops. The MGT level is enabled only
on the control PF.
> MGT level is only enabled for Control PF. Add kerneldoc for caller
> permission constraints. Add wire ABI structures with static_assert for
> upcoming mailbox RPC.
[Severity: Low]
This patch doesn't seem to add any wire ABI structures or static_assert
checks. The diff only touches the Makefile, nbl_core.h, nbl_dispatch.c,
nbl_dispatch.h, nbl_def_dispatch.h and nbl_main.c. None of those files adds
a mailbox wire structure or a static_assert.
At this commit, enum nbl_chan_msg_type in nbl_def_channel.h still contains
only:
enum nbl_chan_msg_type {
NBL_CHAN_MSG_ACK = 0,
/* mailbox msg end */
NBL_CHAN_MSG_MAILBOX_MAX,
};
The only static_assert under drivers/net/ethernet/nebula-matrix/ is the
existing one on struct nbl_board_port_info, which an earlier patch added.
The param structures and their asserts (nbl_chan_param_cfg_msix_map,
nbl_chan_param_set_mailbox_irq, nbl_chan_param_get_vsi_id and
nbl_chan_param_get_eth_id) first appear in the next patch, "dispatch:
implement channel RPC framework and serialize hardware ops".
Could this sentence be removed from this commit message, or moved to the
patch that actually adds the structures?
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
> new file mode 100644
> index 0000000000000..08fef0f679268
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
> @@ -0,0 +1,56 @@
[ ... ]
> +enum {
> + NBL_DISP_CTRL_LVL_MGT,
> + NBL_DISP_CTRL_LVL_MAX,
> +};
> +
> +/**
> + * struct nbl_dispatch_ops - dispatch control plane operation callbacks
[ ... ]
> + * @deinit_module: chip-wide teardown, forwarded the same way; on leonis
> + * it clears driver_status, which starts an asynchronous
> + * firmware per-PF cleanup pass with no completion status
> + * bit, waited for best-effort only. The caller must have
> + * stopped the mailbox channel and drained all in-flight
> + * mailbox DMA before calling it, see nbl_dev_remove().
> + * Returns void. Control-PF exclusive, same NULL rule as
> + * @init_module.
> + *
> + * Warning: init_module/deinit_module are control-PF exclusive. The five
> + * resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq,
> + * get_vsi_id, get_eth_id) are PF-only and resolve to either a local
> + * resource call (control PF) or a mailbox RPC (non-control PF with
> + * has_net). A function with neither has_ctrl nor has_net leaves these
> + * pointers NULL; callers must not invoke them on such functions. VFs are
> + * rejected by the responders with -EPERM.
> + */
> +struct nbl_dispatch_ops {
> + int (*init_module)(struct nbl_dispatch_mgt *disp_mgt);
> + void (*deinit_module)(struct nbl_dispatch_mgt *disp_mgt);
> +};
[Severity: Low]
This isn't a bug, but at this commit the kernel-doc for struct
nbl_dispatch_ops describes code that isn't in the tree yet.
It refers callers to nbl_dev_remove(). That function doesn't exist yet,
because nbl_dev.c hasn't been added at this point.
It also describes five PF-only ops (cfg_msix_map, destroy_msix_map,
set_mailbox_irq, get_vsi_id, get_eth_id), has_net routing over mailbox
RPC, and -EPERM responders. The struct itself has only init_module and
deinit_module, and the enum has only NBL_DISP_CTRL_LVL_MGT and
NBL_DISP_CTRL_LVL_MAX. nbl_disp_refresh_ctrl_ops() also installs only the
two module ops, and only under the MGT level.
Later patches in the series fill this in. The next patch, "dispatch:
implement channel RPC framework and serialize hardware ops", adds
NBL_DISP_CTRL_LVL_NET, the five ops and the responders. The patch after
that, "add common/ctrl dev init/remove operation", adds nbl_dev_remove()
in nbl_dev.c. At the end of the series the comment matches the code; it
only runs ahead of the code at the intermediate commits.
Would it make sense to split this comment so each part lands with the
code it documents?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v31 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
` (4 preceding siblings ...)
2026-10-10 9:59 ` [PATCH v31 net-next 5/8] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
2026-10-10 9:59 ` [PATCH v31 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
2026-10-10 9:59 ` [PATCH v31 net-next 8/8] net/nebula-matrix: add common dev start/stop operation illusion.wang
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Add mailbox RPC requests and responses for five dispatch operations:
cfg_msix_map, destroy_msix_map, set_mailbox_irq, get_vsi_id and
get_eth_id. Each op is resolved by capability level: a function with
has_ctrl calls the resource op locally, a function with has_net forwards
it to the management PF over the channel, and a function with neither
leaves the op pointer NULL, so the caller must not invoke it.
A request builds the payload, sends it synchronously and checks the ACK
length before parsing; a short ACK returns -EBADMSG. A responder rejects
VFs and out-of-range source PF ids with -EPERM, checks the payload
length before parsing, and forwards the resource layer's errno verbatim
so the requester can tell a protocol error from a failed operation. If
the resource op is not installed the responder answers -EOPNOTSUPP. A
peer PF is treated as trusted: hardware and firmware forbid PF
passthrough to guests, so every peer runs this host driver, and the
resource layer bounds the requested vector counts by the global pool.
The cfg_msix_map, destroy_msix_map and set_mailbox_irq requests carry a
caller precondition: disable the mailbox IRQ_RDY flag and use the
polling send path, because the responder changes mailbox MSI-X routing
and an interrupt-driven ACK wakeup cannot be relied on while routing is
in flux.
No dispatch-layer lock is taken around these ops. The resource layer
serializes concurrent callers itself (intr_mgt->lock) and the dispatch
layer has no mutable state of its own to protect; the read-only
get_vsi_id/get_eth_id paths only touch static init metadata.
nbl_disp_remove() documents why the five responders, which are owned by
the channel layer, need no explicit unregistration.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../nebula-matrix/nbl/nbl_core/nbl_dispatch.c | 556 +++++++++++++++++-
.../nbl/nbl_include/nbl_def_channel.h | 40 ++
.../nbl/nbl_include/nbl_def_dispatch.h | 21 +
3 files changed, 616 insertions(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
index 966fee2dec8b..e2d98b7c391b 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
@@ -6,6 +6,174 @@
#include <linux/pci.h>
#include "nbl_dispatch.h"
+static int nbl_disp_chan_get_vsi_id_req(struct nbl_dispatch_mgt *disp_mgt,
+ u16 type, u16 *vsi_id)
+{
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_common_info *common = disp_mgt->common;
+ struct nbl_chan_param_get_vsi_id result = { 0 };
+ struct nbl_chan_param_get_vsi_id param = { 0 };
+ struct nbl_chan_send_info chan_send = {0};
+ int ret;
+
+ param.type = cpu_to_le16(type);
+
+ nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+ NBL_CHAN_MSG_GET_VSI_ID,
+ ¶m, sizeof(param), &result,
+ sizeof(result), 1);
+ ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+ if (ret)
+ return ret;
+ if (chan_send.ack_len != sizeof(result)) {
+ dev_err(disp_mgt->common->dev,
+ "get_vsi_id: short ACK, ack_len=%u expected %zu\n",
+ chan_send.ack_len, sizeof(result));
+ return -EBADMSG;
+ }
+ *vsi_id = le16_to_cpu(result.vsi_id);
+ return 0;
+}
+
+static void nbl_disp_chan_get_vsi_id_resp(void *priv, u16 src_id, u16 msg_id,
+ void *data, u32 data_len)
+{
+ struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct device *dev = disp_mgt->common->dev;
+ struct nbl_chan_param_get_vsi_id result = { 0 };
+ struct nbl_chan_param_get_vsi_id param = { 0 };
+ struct nbl_chan_ack_info chan_ack;
+ int err = 0;
+ u16 vsi_id = 0;
+ u32 rel_pf_id;
+ int ret;
+
+ ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+ &rel_pf_id);
+ if (ret) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (rel_pf_id >= disp_mgt->common->max_pf) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (data_len < sizeof(param)) {
+ err = -EBADMSG;
+ goto ack_out;
+ }
+ memcpy(¶m, data, sizeof(param));
+
+ if (res_ops->get_vsi_id) {
+ ret = res_ops->get_vsi_id(p, src_id, le16_to_cpu(param.type),
+ &vsi_id);
+ /* Forward the resource errno verbatim to the requester */
+ if (ret)
+ err = ret;
+ } else {
+ err = -EOPNOTSUPP;
+ }
+
+ result.vsi_id = cpu_to_le16(vsi_id);
+ack_out:
+ nbl_chan_fill_ack_info(&chan_ack, src_id,
+ NBL_CHAN_MSG_GET_VSI_ID, msg_id, err,
+ &result, sizeof(result));
+ ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+ if (ret)
+ dev_err(dev,
+ "channel send ack failed with ret: %d, msg_type: %d\n",
+ ret, NBL_CHAN_MSG_GET_VSI_ID);
+}
+
+static int nbl_disp_chan_get_eth_id_req(struct nbl_dispatch_mgt *disp_mgt,
+ u16 vsi_id, u8 *eth_num, u8 *eth_id,
+ u8 *logic_eth_id)
+{
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_common_info *common = disp_mgt->common;
+ struct nbl_chan_param_get_eth_id result = { 0 };
+ struct nbl_chan_param_get_eth_id param = { 0 };
+ struct nbl_chan_send_info chan_send = {0};
+ int ret;
+
+ param.vsi_id = cpu_to_le16(vsi_id);
+
+ nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+ NBL_CHAN_MSG_GET_ETH_ID,
+ ¶m, sizeof(param), &result,
+ sizeof(result), 1);
+ ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+ if (ret)
+ return ret;
+ if (chan_send.ack_len != sizeof(result)) {
+ dev_err(disp_mgt->common->dev,
+ "get_eth_id: short ACK, ack_len=%u expected %zu\n",
+ chan_send.ack_len, sizeof(result));
+ return -EBADMSG;
+ }
+ *eth_num = result.eth_num;
+ *eth_id = result.eth_id;
+ *logic_eth_id = result.logic_eth_id;
+
+ return 0;
+}
+
+static void nbl_disp_chan_get_eth_id_resp(void *priv, u16 src_id, u16 msg_id,
+ void *data, u32 data_len)
+{
+ struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct nbl_chan_param_get_eth_id result = { 0 };
+ struct nbl_chan_param_get_eth_id param = { 0 };
+ struct device *dev = disp_mgt->common->dev;
+ struct nbl_chan_ack_info chan_ack;
+ int err = 0;
+ u32 rel_pf_id;
+ int ret;
+
+ ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+ &rel_pf_id);
+ if (ret) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (rel_pf_id >= disp_mgt->common->max_pf) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (data_len < sizeof(param)) {
+ err = -EBADMSG;
+ goto ack_out;
+ }
+ memcpy(¶m, data, sizeof(param));
+
+ if (res_ops->get_eth_id) {
+ ret = res_ops->get_eth_id(p, src_id, le16_to_cpu(param.vsi_id),
+ &result.eth_num, &result.eth_id,
+ &result.logic_eth_id);
+ /* Forward the resource errno verbatim to the requester */
+ if (ret)
+ err = ret;
+ } else {
+ err = -EOPNOTSUPP;
+ }
+ack_out:
+ nbl_chan_fill_ack_info(&chan_ack, src_id,
+ NBL_CHAN_MSG_GET_ETH_ID, msg_id, err,
+ &result, sizeof(result));
+ ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+ if (ret)
+ dev_err(dev,
+ "channel send ack failed with ret: %d, msg_type: %d\n",
+ ret, NBL_CHAN_MSG_GET_ETH_ID);
+}
+
static void nbl_disp_deinit_module(struct nbl_dispatch_mgt *disp_mgt)
{
struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
@@ -25,6 +193,353 @@ static int nbl_disp_init_module(struct nbl_dispatch_mgt *disp_mgt)
return -EOPNOTSUPP;
}
+static int nbl_disp_cfg_msix_map(struct nbl_dispatch_mgt *disp_mgt,
+ u16 num_net_msix, u16 num_others_msix,
+ bool net_msix_mask_en)
+{
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct nbl_common_info *common = disp_mgt->common;
+
+ if (!res_ops->cfg_msix_map)
+ return -EOPNOTSUPP;
+ /*
+ * No core-layer lock is taken here: the resource layer serializes
+ * concurrent callers itself (intr_mgt->lock), and the dispatch
+ * layer has no mutable state of its own for this op to protect.
+ */
+ return res_ops->cfg_msix_map(p, common->mgt_pf, num_net_msix,
+ num_others_msix, net_msix_mask_en);
+}
+
+/*
+ * Precondition: caller must disable mailbox IRQ_RDY and switch send_msg
+ * to the polling path before issuing this RPC. The responder rewrites
+ * the requester's mailbox MSI-X routing during the resource op, so the
+ * ACK cannot rely on interrupt wakeup while routing is in flux.
+ */
+static int
+nbl_disp_chan_cfg_msix_map_req(struct nbl_dispatch_mgt *disp_mgt,
+ u16 num_net_msix, u16 num_others_msix,
+ bool net_msix_mask_en)
+{
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_common_info *common = disp_mgt->common;
+ struct nbl_chan_param_cfg_msix_map param = { 0 };
+ struct nbl_chan_send_info chan_send = {0};
+ int ret;
+
+ param.num_net_msix = cpu_to_le16(num_net_msix);
+ param.num_others_msix = cpu_to_le16(num_others_msix);
+ param.msix_mask_en = cpu_to_le16(!!net_msix_mask_en);
+
+ nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+ NBL_CHAN_MSG_CONFIGURE_MSIX_MAP,
+ ¶m, sizeof(param),
+ NULL, 0, 1);
+ ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+ if (ret)
+ return ret;
+ return 0;
+}
+
+static void nbl_disp_chan_cfg_msix_map_resp(void *priv, u16 src_id, u16 msg_id,
+ void *data, u32 data_len)
+{
+ struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct device *dev = disp_mgt->common->dev;
+ struct nbl_chan_param_cfg_msix_map param = { 0 };
+ struct nbl_chan_ack_info chan_ack;
+ int err = 0;
+ u32 rel_pf_id;
+ int ret;
+
+ ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+ &rel_pf_id);
+ if (ret) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (rel_pf_id >= disp_mgt->common->max_pf) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (data_len < sizeof(param)) {
+ err = -EBADMSG;
+ goto ack_out;
+ }
+ memcpy(¶m, data, sizeof(param));
+
+ if (res_ops->cfg_msix_map) {
+ /*
+ * The vector counts come from the peer PF. A peer is
+ * trusted: hardware and firmware forbid PF passthrough to
+ * guests (see nbl_get_func_param()), so every peer runs this
+ * same host driver, and the resource layer bounds the counts
+ * by the global vector pool. A peer asking for the whole
+ * pool can starve the others with -EAGAIN, but it cannot
+ * corrupt state.
+ */
+ ret = res_ops->cfg_msix_map(p, src_id,
+ le16_to_cpu(param.num_net_msix),
+ le16_to_cpu(param.num_others_msix),
+ !!le16_to_cpu(param.msix_mask_en));
+ /* Forward the resource errno verbatim to the requester */
+ if (ret)
+ err = ret;
+ } else {
+ err = -EOPNOTSUPP;
+ }
+ack_out:
+ nbl_chan_fill_ack_info(&chan_ack, src_id,
+ NBL_CHAN_MSG_CONFIGURE_MSIX_MAP, msg_id,
+ err, NULL, 0);
+ ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+ if (ret)
+ dev_err(dev,
+ "channel send ack failed with ret: %d, msg_type: %d\n",
+ ret, NBL_CHAN_MSG_CONFIGURE_MSIX_MAP);
+}
+
+/*
+ * Precondition: caller must disable mailbox IRQ_RDY and switch send_msg
+ * to the polling path before issuing this RPC. The responder retargets
+ * the requester's mailbox MSI-X routing during the resource op, so the
+ * ACK cannot rely on interrupt wakeup while routing is in flux.
+ */
+static int nbl_disp_chan_destroy_msix_map_req(struct nbl_dispatch_mgt *disp_mgt)
+{
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_common_info *common = disp_mgt->common;
+ struct nbl_chan_send_info chan_send = {0};
+ int ret;
+
+ nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+ NBL_CHAN_MSG_DESTROY_MSIX_MAP,
+ NULL, 0, NULL, 0, 1);
+ ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+ if (ret)
+ return ret;
+ return 0;
+}
+
+static void nbl_disp_chan_destroy_msix_map_resp(void *priv, u16 src_id,
+ u16 msg_id, void *data,
+ u32 data_len)
+{
+ struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct device *dev = disp_mgt->common->dev;
+ struct nbl_chan_ack_info chan_ack;
+ int err = 0;
+ u32 rel_pf_id;
+ int ret;
+
+ ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+ &rel_pf_id);
+ if (ret) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (rel_pf_id >= disp_mgt->common->max_pf) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (res_ops->destroy_msix_map) {
+ ret = res_ops->destroy_msix_map(p, src_id);
+ /* Forward the resource errno verbatim to the requester */
+ if (ret)
+ err = ret;
+ } else {
+ err = -EOPNOTSUPP;
+ }
+ack_out:
+ nbl_chan_fill_ack_info(&chan_ack, src_id,
+ NBL_CHAN_MSG_DESTROY_MSIX_MAP, msg_id,
+ err, NULL, 0);
+ ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+ if (ret)
+ dev_err(dev,
+ "channel send ack failed with ret: %d, msg_type: %d\n",
+ ret, NBL_CHAN_MSG_DESTROY_MSIX_MAP);
+}
+
+/*
+ * Precondition: caller must disable mailbox IRQ_RDY and switch send_msg
+ * to the polling path before issuing this RPC. The responder rewrites
+ * the requester's own mailbox MSI-X routing (MSIX_IDX / MSIX_IDX_VALID)
+ * before the ACK is sent, so the ACK cannot rely on interrupt wakeup
+ * while routing is in flux.
+ */
+static int nbl_disp_chan_set_mailbox_irq_req(struct nbl_dispatch_mgt *disp_mgt,
+ u16 vector_id, bool en_msix)
+{
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_chan_param_set_mailbox_irq param = { 0 };
+ struct nbl_common_info *common = disp_mgt->common;
+ struct nbl_chan_send_info chan_send = {0};
+ int ret;
+
+ param.vector_id = cpu_to_le16(vector_id);
+ param.en_msix = !!en_msix;
+
+ nbl_chan_fill_send_info(&chan_send, common->mgt_pf,
+ NBL_CHAN_MSG_MAILBOX_SET_IRQ,
+ ¶m, sizeof(param), NULL, 0, 1);
+ ret = chan_ops->send_msg(disp_mgt->chan_ops_tbl->priv, &chan_send);
+ if (ret)
+ return ret;
+ return 0;
+}
+
+static void nbl_disp_chan_set_mailbox_irq_resp(void *priv, u16 src_id,
+ u16 msg_id, void *data,
+ u32 data_len)
+{
+ struct nbl_dispatch_mgt *disp_mgt = (struct nbl_dispatch_mgt *)priv;
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct nbl_chan_param_set_mailbox_irq param = { 0 };
+ struct device *dev = disp_mgt->common->dev;
+ struct nbl_chan_ack_info chan_ack;
+ int err = 0;
+ u16 vector_id;
+ u32 rel_pf_id;
+ bool en_msix;
+ int ret;
+
+ ret = nbl_common_func_id_to_rel_pf_id(disp_mgt->common, src_id,
+ &rel_pf_id);
+ if (ret) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (rel_pf_id >= disp_mgt->common->max_pf) {
+ err = -EPERM;
+ goto ack_out;
+ }
+ if (data_len < sizeof(param)) {
+ err = -EBADMSG;
+ goto ack_out;
+ }
+ memcpy(¶m, data, sizeof(param));
+ vector_id = le16_to_cpu(param.vector_id);
+ en_msix = !!param.en_msix;
+
+ if (res_ops->set_mailbox_irq) {
+ ret = res_ops->set_mailbox_irq(p, src_id, vector_id, en_msix);
+ /* Forward the resource errno verbatim to the requester */
+ if (ret)
+ err = ret;
+ } else {
+ err = -EOPNOTSUPP;
+ }
+
+ack_out:
+ nbl_chan_fill_ack_info(&chan_ack, src_id,
+ NBL_CHAN_MSG_MAILBOX_SET_IRQ, msg_id,
+ err, NULL, 0);
+ ret = chan_ops->send_ack(disp_mgt->chan_ops_tbl->priv, &chan_ack);
+ if (ret)
+ dev_err(dev,
+ "channel send ack failed with ret: %d, msg_type: %d\n",
+ ret, NBL_CHAN_MSG_MAILBOX_SET_IRQ);
+}
+
+static int nbl_disp_destroy_msix_map(struct nbl_dispatch_mgt *disp_mgt)
+{
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct nbl_common_info *common = disp_mgt->common;
+
+ if (!res_ops->destroy_msix_map)
+ return -EOPNOTSUPP;
+ return res_ops->destroy_msix_map(p, common->mgt_pf);
+}
+
+static int nbl_disp_set_mailbox_irq(struct nbl_dispatch_mgt *disp_mgt,
+ u16 vector_id, bool en_msix)
+{
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct nbl_common_info *common = disp_mgt->common;
+
+ if (!res_ops->set_mailbox_irq)
+ return -EOPNOTSUPP;
+ return res_ops->set_mailbox_irq(p, common->mgt_pf, vector_id, en_msix);
+}
+
+static int nbl_disp_get_vsi_id(struct nbl_dispatch_mgt *disp_mgt, u16 type,
+ u16 *vsi_id)
+{
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct nbl_common_info *common = disp_mgt->common;
+
+ if (res_ops->get_vsi_id)
+ return res_ops->get_vsi_id(p, common->mgt_pf, type, vsi_id);
+ return -EOPNOTSUPP;
+}
+
+static int nbl_disp_get_eth_id(struct nbl_dispatch_mgt *disp_mgt, u16 vsi_id,
+ u8 *eth_num, u8 *eth_id, u8 *logic_eth_id)
+{
+ struct nbl_resource_ops *res_ops = disp_mgt->res_ops_tbl->ops;
+ struct nbl_resource_mgt *p = disp_mgt->res_ops_tbl->priv;
+ struct nbl_common_info *common = disp_mgt->common;
+
+ if (res_ops->get_eth_id)
+ return res_ops->get_eth_id(p, common->mgt_pf, vsi_id,
+ eth_num, eth_id, logic_eth_id);
+ return -EOPNOTSUPP;
+}
+
+static int nbl_disp_setup_msg(struct nbl_dispatch_mgt *disp_mgt)
+{
+ struct nbl_channel_ops *chan_ops = disp_mgt->chan_ops_tbl->ops;
+ struct nbl_channel_mgt *p = disp_mgt->chan_ops_tbl->priv;
+ int ret = 0;
+ int _ret;
+
+ _ret = chan_ops->register_msg(p, NBL_CHAN_MSG_CONFIGURE_MSIX_MAP,
+ nbl_disp_chan_cfg_msix_map_resp,
+ disp_mgt);
+ if (_ret < 0 && !ret)
+ ret = _ret;
+
+ _ret = chan_ops->register_msg(p, NBL_CHAN_MSG_DESTROY_MSIX_MAP,
+ nbl_disp_chan_destroy_msix_map_resp,
+ disp_mgt);
+ if (_ret < 0 && !ret)
+ ret = _ret;
+
+ _ret = chan_ops->register_msg(p, NBL_CHAN_MSG_MAILBOX_SET_IRQ,
+ nbl_disp_chan_set_mailbox_irq_resp,
+ disp_mgt);
+ if (_ret < 0 && !ret)
+ ret = _ret;
+
+ _ret = chan_ops->register_msg(p, NBL_CHAN_MSG_GET_VSI_ID,
+ nbl_disp_chan_get_vsi_id_resp,
+ disp_mgt);
+ if (_ret < 0 && !ret)
+ ret = _ret;
+
+ _ret = chan_ops->register_msg(p, NBL_CHAN_MSG_GET_ETH_ID,
+ nbl_disp_chan_get_eth_id_resp,
+ disp_mgt);
+ if (_ret < 0 && !ret)
+ ret = _ret;
+
+ return ret;
+}
+
static void nbl_disp_set_ctrl_bit(struct nbl_dispatch_mgt *disp_mgt, u32 lvl)
{
set_bit(lvl, disp_mgt->ctrl_lvl);
@@ -34,9 +549,22 @@ static void nbl_disp_refresh_ctrl_ops(struct nbl_dispatch_mgt *disp_mgt)
{
struct nbl_dispatch_ops *disp_ops = disp_mgt->disp_ops_tbl->ops;
+ memset(disp_ops, 0, sizeof(*disp_ops));
if (test_bit(NBL_DISP_CTRL_LVL_MGT, disp_mgt->ctrl_lvl)) {
disp_ops->init_module = nbl_disp_init_module;
disp_ops->deinit_module = nbl_disp_deinit_module;
+ disp_ops->cfg_msix_map = nbl_disp_cfg_msix_map;
+ disp_ops->destroy_msix_map = nbl_disp_destroy_msix_map;
+ disp_ops->set_mailbox_irq = nbl_disp_set_mailbox_irq;
+ disp_ops->get_vsi_id = nbl_disp_get_vsi_id;
+ disp_ops->get_eth_id = nbl_disp_get_eth_id;
+ } else if (test_bit(NBL_DISP_CTRL_LVL_NET, disp_mgt->ctrl_lvl)) {
+ disp_ops->cfg_msix_map =
+ nbl_disp_chan_cfg_msix_map_req;
+ disp_ops->destroy_msix_map = nbl_disp_chan_destroy_msix_map_req;
+ disp_ops->set_mailbox_irq = nbl_disp_chan_set_mailbox_irq_req;
+ disp_ops->get_vsi_id = nbl_disp_chan_get_vsi_id_req;
+ disp_ops->get_eth_id = nbl_disp_chan_get_eth_id_req;
}
}
@@ -104,14 +632,40 @@ int nbl_disp_init(struct nbl_adapter *adapter)
adapter->core.disp_mgt = disp_mgt;
adapter->intf.dispatch_ops_tbl = disp_ops_tbl;
+ ret = nbl_disp_setup_msg(disp_mgt);
+ if (ret)
+ return ret;
+
if (common->has_ctrl)
nbl_disp_set_ctrl_bit(disp_mgt, NBL_DISP_CTRL_LVL_MGT);
+ if (common->has_net)
+ nbl_disp_set_ctrl_bit(disp_mgt, NBL_DISP_CTRL_LVL_NET);
nbl_disp_refresh_ctrl_ops(disp_mgt);
return 0;
}
void nbl_disp_remove(struct nbl_adapter *adapter)
{
- /* Dispatch structures are allocated via devm */
+ /*
+ * Dispatch structures are devm-allocated and freed at detach.
+ *
+ * The five responders registered by nbl_disp_setup_msg() are
+ * owned by the channel layer (xarray of handlers) and are never
+ * unregistered here. This is safe because the teardown order
+ * guarantees no responder can run after this point:
+ *
+ * mailbox teardown
+ * -> cancel_work_sync(clean_mbx_task) // drain RX work
+ * -> nbl_chan_teardown_queue() // stop HW queue,
+ * // join clean task,
+ * // active=false
+ * -> nbl_chan_remove_common()
+ * -> destroy_wq() // no new work
+ * -> nbl_chan_remove_msg_handler() // free handler nodes
+ *
+ * By the time devres frees disp_mgt, the mailbox queue is stopped
+ * and the handler xarray is empty, so no responder can touch
+ * res_mgt->intr_mgt after nbl_intr_mgt_stop() has cleared it.
+ */
}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
index bf971121d2ec..5aa16d4cd432 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_channel.h
@@ -7,6 +7,7 @@
#define _NBL_DEF_CHANNEL_H_
#include <linux/types.h>
+#include <linux/build_bug.h>
struct nbl_channel_mgt;
struct nbl_adapter;
@@ -30,6 +31,11 @@ typedef void (*nbl_chan_resp)(void *, u16, u16, void *, u32);
*/
enum nbl_chan_msg_type {
NBL_CHAN_MSG_ACK = 0,
+ NBL_CHAN_MSG_CONFIGURE_MSIX_MAP = 17,
+ NBL_CHAN_MSG_DESTROY_MSIX_MAP = 18,
+ NBL_CHAN_MSG_MAILBOX_SET_IRQ = 19,
+ NBL_CHAN_MSG_GET_VSI_ID = 21,
+ NBL_CHAN_MSG_GET_ETH_ID = 67,
/* mailbox msg end */
NBL_CHAN_MSG_MAILBOX_MAX,
};
@@ -39,6 +45,32 @@ enum nbl_chan_state {
NBL_CHAN_STATE_NBITS
};
+struct nbl_chan_param_cfg_msix_map {
+ __le16 num_net_msix;
+ __le16 num_others_msix;
+ __le16 msix_mask_en;
+ __le16 rsvd;
+};
+
+struct nbl_chan_param_set_mailbox_irq {
+ __le16 vector_id;
+ u8 en_msix;
+ u8 rsvd;
+};
+
+struct nbl_chan_param_get_vsi_id {
+ __le16 vsi_id;
+ __le16 type;
+};
+
+struct nbl_chan_param_get_eth_id {
+ __le16 vsi_id;
+ u8 eth_num;
+ u8 eth_id;
+ u8 logic_eth_id;
+ u8 rsvd[3];
+};
+
struct nbl_board_port_info {
u8 eth_num;
u8 eth_speed;
@@ -46,6 +78,14 @@ struct nbl_board_port_info {
u8 rsv[5];
};
+static_assert(sizeof(struct nbl_chan_param_cfg_msix_map) == 8,
+ "nbl_chan_param_cfg_msix_map size must be 8 bytes");
+static_assert(sizeof(struct nbl_chan_param_set_mailbox_irq) == 4,
+ "nbl_chan_param_set_mailbox_irq size must be 4 bytes");
+static_assert(sizeof(struct nbl_chan_param_get_vsi_id) == 4,
+ "nbl_chan_param_get_vsi_id size must be 4 bytes");
+static_assert(sizeof(struct nbl_chan_param_get_eth_id) == 8,
+ "nbl_chan_param_get_eth_id size must be 8 bytes");
static_assert(sizeof(struct nbl_board_port_info) == 8,
"nbl_board_port_info size must be 8 bytes");
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
index 08fef0f67926..bcd74b9df827 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
@@ -12,6 +12,7 @@ struct nbl_dispatch_mgt;
struct nbl_adapter;
enum {
NBL_DISP_CTRL_LVL_MGT,
+ NBL_DISP_CTRL_LVL_NET,
NBL_DISP_CTRL_LVL_MAX,
};
@@ -32,6 +33,16 @@ enum {
* mailbox DMA before calling it, see nbl_dev_remove().
* Returns void. Control-PF exclusive, same NULL rule as
* @init_module.
+ * @cfg_msix_map: configure function msix mapping table
+ * @destroy_msix_map: tear down msix mapping resource
+ * @set_mailbox_irq: bind mailbox interrupt to specified msix vector
+ * @get_vsi_id: resolve VSI ID by type
+ * @get_eth_id: resolve eth port info from VSI ID
+ *
+ * These callbacks are not dispatch-layer setup/teardown: the dispatch
+ * structures are devm-allocated in nbl_disp_init(), and nbl_disp_remove()
+ * only documents why no explicit teardown is needed. They forward
+ * chip/firmware control to the upper layers instead.
*
* Warning: init_module/deinit_module are control-PF exclusive. The five
* resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq,
@@ -44,6 +55,16 @@ enum {
struct nbl_dispatch_ops {
int (*init_module)(struct nbl_dispatch_mgt *disp_mgt);
void (*deinit_module)(struct nbl_dispatch_mgt *disp_mgt);
+ int (*cfg_msix_map)(struct nbl_dispatch_mgt *disp_mgt,
+ u16 num_net_msix, u16 num_others_msix,
+ bool net_msix_mask_en);
+ int (*destroy_msix_map)(struct nbl_dispatch_mgt *disp_mgt);
+ int (*set_mailbox_irq)(struct nbl_dispatch_mgt *disp_mgt,
+ u16 vector_id, bool en_msix);
+ int (*get_vsi_id)(struct nbl_dispatch_mgt *disp_mgt, u16 type,
+ u16 *vsi_id);
+ int (*get_eth_id)(struct nbl_dispatch_mgt *disp_mgt, u16 vsi_id,
+ u8 *eth_num, u8 *eth_id, u8 *logic_eth_id);
};
struct nbl_dispatch_ops_tbl {
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops
2026-10-10 9:59 ` [PATCH v31 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] Baseline decision is false: in the final state of the series this
is not a bug.
- [Low] The kernel-doc this patch adds for struct nbl_dispatch_ops in
nbl_def_dispatch.h says the new callbacks 'forward chip/firmware…
--
net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops
This patch adds mailbox RPC requests and responders for five dispatch ops:
cfg_msix_map, destroy_msix_map, set_mailbox_irq, get_vsi_id and get_eth_id.
Each op is chosen by capability level. A has_ctrl function calls the
resource layer locally, and a has_net function forwards the request to the
management PF over the channel. nbl_disp_remove() gains a comment that
explains why the registered responders are never unregistered.
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
> index 966fee2dec8b0..e2d98b7c391b4 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
[ ... ]
> void nbl_disp_remove(struct nbl_adapter *adapter)
> {
> - /* Dispatch structures are allocated via devm */
> + /*
> + * Dispatch structures are devm-allocated and freed at detach.
> + *
> + * The five responders registered by nbl_disp_setup_msg() are
> + * owned by the channel layer (xarray of handlers) and are never
> + * unregistered here. This is safe because the teardown order
> + * guarantees no responder can run after this point:
> + *
> + * mailbox teardown
> + * -> cancel_work_sync(clean_mbx_task) // drain RX work
> + * -> nbl_chan_teardown_queue() // stop HW queue,
> + * // join clean task,
> + * // active=false
> + * -> nbl_chan_remove_common()
> + * -> destroy_wq() // no new work
> + * -> nbl_chan_remove_msg_handler() // free handler nodes
[Severity: Low]
This isn't a bug, but does this comment describe code that exists at this
commit?
clean_mbx_task doesn't appear anywhere in the tree yet. Nothing calls
chan_ops->setup_queue at this point either, so the mailbox RX path that
this sequence quiesces is never active.
The ordering the comment relies on is nbl_dev_remove() doing the mailbox
teardown before nbl_disp_remove(), nbl_res_remove_leonis() and
nbl_chan_remove_common(). That ordering only arrives with the later
"net/nebula-matrix: add common/ctrl dev init/remove operation" patch.
Also, destroy_wq() looks like shorthand for nbl_common_destroy_wq().
Could this comment move to the patch that introduces that teardown
ordering, and use the real function names?
> + *
> + * By the time devres frees disp_mgt, the mailbox queue is stopped
> + * and the handler xarray is empty, so no responder can touch
> + * res_mgt->intr_mgt after nbl_intr_mgt_stop() has cleared it.
> + */
> }
[ ... ]
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
> index 08fef0f679268..bcd74b9df8275 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dispatch.h
[ ... ]
> @@ -32,6 +33,16 @@ enum {
> * mailbox DMA before calling it, see nbl_dev_remove().
> * Returns void. Control-PF exclusive, same NULL rule as
> * @init_module.
> + * @cfg_msix_map: configure function msix mapping table
> + * @destroy_msix_map: tear down msix mapping resource
> + * @set_mailbox_irq: bind mailbox interrupt to specified msix vector
> + * @get_vsi_id: resolve VSI ID by type
> + * @get_eth_id: resolve eth port info from VSI ID
> + *
> + * These callbacks are not dispatch-layer setup/teardown: the dispatch
> + * structures are devm-allocated in nbl_disp_init(), and nbl_disp_remove()
> + * only documents why no explicit teardown is needed. They forward
> + * chip/firmware control to the upper layers instead.
[Severity: Low]
Is "upper layers" the right direction here? The new callbacks forward
downward, for example:
nbl_disp_cfg_msix_map()
return res_ops->cfg_msix_map(p, common->mgt_pf, num_net_msix,
num_others_msix, net_msix_mask_en);
The nbl_disp_chan_*_req() variants call chan_ops->send_msg() instead.
The resource and channel layers both sit below dispatch, and nbl_dev is
the caller. Later patches in the series don't change this text, so the
wording stays as it is.
> *
> * Warning: init_module/deinit_module are control-PF exclusive. The five
> * resource ops (cfg_msix_map, destroy_msix_map, set_mailbox_irq,
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v31 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
` (5 preceding siblings ...)
2026-10-10 9:59 ` [PATCH v31 net-next 6/8] net/nebula-matrix: dispatch: implement channel RPC framework and serialize hardware ops illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
2026-10-10 9:59 ` [PATCH v31 net-next 8/8] net/nebula-matrix: add common dev start/stop operation illusion.wang
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Add nbl_dev.c, nbl_dev.h and nbl_def_dev.h to host device-level
initialization and teardown logic, and wire the new nbl_dev_init() /
nbl_dev_remove() entry points into core init/remove.
Common device setup allocates per-device state, initializes the mailbox
cleanup work and creates the mailbox channel queue. VSI/ETH identity
lookup and MSI-X vector allocation are deferred to the subsequent start
routine, so that the control PF mailbox responder is fully ready before
any cross-PF RPC is issued and each commit in the series can be probed
on its own.
Control-PF-only setup invokes chip-level module initialization and
programs the mailbox QINFO routing table entries that map PF bus/devid
to the mailbox queue.
Init/teardown ordering is dictated by a hardware constraint rather than
by strict symmetry. Init: program mailbox QINFO (control PF only),
create the mailbox queue, then run control device init_module. Teardown:
stop the mailbox channel and drain all in-flight DMA first, then call
deinit_module(). This ordering prevents the firmware from reclaiming
the routing under a live transfer, which would leave the device DMAing
into released memory.
The firmware reclaims per-PF state (mailbox QINFO routing, qinfo
registers and the like) asynchronously once driver status is marked
inactive. It does not revert the chip-wide datapath registers
configured by init_module; those stay valid until chip reset (see
nbl_hw_deinit_module()). The device link created during non-control PF
probe makes the driver core unbind all sibling PFs before the control PF
is removed.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../net/ethernet/nebula-matrix/nbl/Makefile | 1 +
.../net/ethernet/nebula-matrix/nbl/nbl_core.h | 1 +
.../nebula-matrix/nbl/nbl_core/nbl_dev.c | 276 ++++++++++++++++++
.../nebula-matrix/nbl/nbl_core/nbl_dev.h | 55 ++++
.../nbl/nbl_include/nbl_def_dev.h | 14 +
.../net/ethernet/nebula-matrix/nbl/nbl_main.c | 9 +
6 files changed, 356 insertions(+)
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h
create mode 100644 drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/Makefile b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
index b7eebd89b4d1..71fbe3ee7e62 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/Makefile
+++ b/drivers/net/ethernet/nebula-matrix/nbl/Makefile
@@ -11,4 +11,5 @@ nbl-objs += nbl_common/nbl_common.o \
nbl_hw/nbl_interrupt.o \
nbl_hw/nbl_chip.o \
nbl_core/nbl_dispatch.o \
+ nbl_core/nbl_dev.o \
nbl_main.o
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index 4d8cea8d8ab3..c3c4dd685bf6 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -25,6 +25,7 @@ struct nbl_core {
struct nbl_hw_mgt *hw_mgt;
struct nbl_resource_mgt *res_mgt;
struct nbl_dispatch_mgt *disp_mgt;
+ struct nbl_dev_mgt *dev_mgt;
struct nbl_channel_mgt *chan_mgt;
};
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
new file mode 100644
index 000000000000..390bc2b177e2
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
@@ -0,0 +1,276 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+#include <linux/device.h>
+#include <linux/pci.h>
+#include "nbl_dev.h"
+
+static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+ struct nbl_msix_info *msix_info = &dev_common->msix_info;
+
+ /* mailbox vector allocated in nbl_dev_start() via
+ * nbl_dev_init_interrupt_scheme(); nbl_dev_request_mailbox_irq()
+ * only attaches the irq handler to pre-allocated vectors.
+ */
+ msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num = 1;
+}
+
+/* ---------- Channel config ---------- */
+static void nbl_dev_setup_chan_qinfo(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
+{
+ struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+ struct nbl_channel_mgt *priv = dev_mgt->chan_ops_tbl->priv;
+ struct nbl_common_info *common = dev_mgt->common;
+
+ if (!chan_ops->check_queue_exist(priv, chan_type))
+ return;
+
+ /*
+ * common->hw_bus is the control PF's real bus number, captured in
+ * nbl_res_ctrl_dev_sriov_info_init() during nbl_res_init_leonis().
+ * nbl_core_init() runs resource init before nbl_dev_init(), so the
+ * value is always initialized when this control-PF-only path runs;
+ * nbl_res_intr_cfg_msix_map() consumes it for cfg_msix_map() the
+ * same way.
+ */
+ chan_ops->cfg_chan_qinfo_map_table(priv, common->hw_bus, common->devid);
+}
+
+static int nbl_dev_setup_chan_queue(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
+{
+ struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+ struct nbl_channel_mgt *priv = dev_mgt->chan_ops_tbl->priv;
+ int ret = 0;
+
+ if (chan_ops->check_queue_exist(priv, chan_type))
+ ret = chan_ops->setup_queue(priv, chan_type);
+
+ return ret;
+}
+
+static int nbl_dev_remove_chan_queue(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
+{
+ struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+ struct nbl_channel_mgt *priv = dev_mgt->chan_ops_tbl->priv;
+ int ret = 0;
+
+ if (chan_ops->check_queue_exist(priv, chan_type))
+ ret = chan_ops->teardown_queue(priv, chan_type);
+
+ return ret;
+}
+
+static void nbl_dev_register_chan_task(struct nbl_dev_mgt *dev_mgt,
+ u8 chan_type, struct work_struct *task)
+{
+ struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+
+ if (chan_ops->check_queue_exist(dev_mgt->chan_ops_tbl->priv, chan_type))
+ chan_ops->register_chan_task(dev_mgt->chan_ops_tbl->priv,
+ chan_type, task);
+}
+
+/* ---------- Tasks config ---------- */
+static void nbl_dev_clean_mailbox_task(struct work_struct *work)
+{
+ struct nbl_dev_common *common_dev =
+ container_of(work, struct nbl_dev_common, clean_mbx_task);
+ struct nbl_dev_mgt *dev_mgt = common_dev->dev_mgt;
+ struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+
+ chan_ops->clean_queue_subtask(dev_mgt->chan_ops_tbl->priv,
+ NBL_CHAN_TYPE_MAILBOX);
+}
+
+/* ---------- Dev init process ---------- */
+static int nbl_dev_setup_common_dev(struct nbl_adapter *adapter)
+{
+ struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+ struct nbl_dev_common *common_dev;
+ int ret;
+
+ common_dev = devm_kzalloc(&adapter->pdev->dev, sizeof(*common_dev),
+ GFP_KERNEL);
+ if (!common_dev)
+ return -ENOMEM;
+ common_dev->dev_mgt = dev_mgt;
+
+ /*
+ * INIT_WORK must precede nbl_dev_register_chan_task() below: the
+ * registration hands the work item to the channel and the IRQ
+ * path may queue it as soon as setup_queue() completes.
+ */
+ INIT_WORK(&common_dev->clean_mbx_task, nbl_dev_clean_mailbox_task);
+
+ ret = nbl_dev_setup_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
+ if (ret) {
+ /*
+ * Nothing to unwind: chan_info->active is only set on a
+ * successful setup_queue(), so a failed one left no
+ * sender, no registered work item and no mailbox
+ * traffic. Calling teardown_queue() here would only
+ * take its "channel not active" early exit and emit a
+ * misleading duplicate-teardown warning. The partially
+ * allocated DMA rings are intentionally kept until devres
+ * releases them at detach (chan_info->dma_allocated
+ * blocks a re-setup).
+ */
+ return ret;
+ }
+
+ nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX,
+ &common_dev->clean_mbx_task);
+ /*
+ * VSI/ETH identity fetch moved to nbl_dev_start().
+ * This avoids cross-PF probe race when manager PF is not ready.
+ */
+ dev_mgt->common_dev = common_dev;
+ nbl_dev_init_msix_cnt(dev_mgt);
+
+ return 0;
+}
+
+static void nbl_dev_remove_common_dev(struct nbl_adapter *adapter)
+{
+ struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+ struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+
+ if (!common_dev)
+ return;
+ cancel_work_sync(&common_dev->clean_mbx_task);
+ /*
+ * nbl_chan_teardown_queue() never fails: it returns 0 both on the
+ * early exit for a channel that never became active and on the
+ * normal path, which drains every inflight sender before
+ * returning. There is therefore no partial-failure state that
+ * could leave mailbox DMA in flight and nothing to report here.
+ */
+ nbl_dev_remove_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
+ nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX, NULL);
+}
+
+static int nbl_dev_setup_ctrl_dev(struct nbl_adapter *adapter)
+{
+ struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+ struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+
+ return disp_ops->init_module(dev_mgt->disp_ops_tbl->priv);
+}
+
+/*
+ * Tear down control device: deinit_module sets driver_status=false
+ * to notify firmware to clean all per-PF hardware state (including
+ * qinfo registers). The qinfo map programmed in nbl_dev_init() via
+ * nbl_dev_setup_chan_qinfo() is not explicitly cleared; firmware
+ * handles it on driver_status change.
+ *
+ * Teardown ordering guarantee: every non-management PF creates a
+ * consumer->control PF device link in its probe path, so the driver
+ * core always unbinds all siblings before allowing the control PF to
+ * be detached (sysfs unbind, driver unregister and hot-unplug alike).
+ * nbl_probe_chip_deps() defers a sibling's probe until the control PF
+ * is fully bound (DL_DEV_DRIVER_BOUND), and re-checks that state after
+ * device_link_add() so a supplier that slipped to DL_DEV_NO_DRIVER or
+ * DL_DEV_PROBING in between also defers. A sibling therefore cannot
+ * reach its mailbox setup before the control PF is fully probed, nor
+ * race this deinit afterwards - the link forces it to be unbound
+ * first. The link is a persistent managed link with
+ * DL_FLAG_AUTOPROBE_CONSUMER, so binding the control PF again makes
+ * the core re-probe the siblings; without it they would stay unbound
+ * until an admin rebound each one by hand.
+ *
+ * Direct control-PF FLR (which never runs driver teardown) cannot be
+ * guarded here.
+ */
+static void nbl_dev_remove_ctrl_dev(struct nbl_adapter *adapter)
+{
+ struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+ struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+
+ disp_ops->deinit_module(dev_mgt->disp_ops_tbl->priv);
+}
+
+static struct nbl_dev_mgt *nbl_dev_setup_dev_mgt(struct nbl_common_info *common)
+{
+ struct nbl_dev_mgt *dev_mgt;
+
+ dev_mgt = devm_kzalloc(common->dev, sizeof(*dev_mgt), GFP_KERNEL);
+ if (!dev_mgt)
+ return ERR_PTR(-ENOMEM);
+
+ dev_mgt->common = common;
+ return dev_mgt;
+}
+
+int nbl_dev_init(struct nbl_adapter *adapter)
+{
+ struct nbl_common_info *common = &adapter->common;
+ struct nbl_dispatch_ops_tbl *disp_ops_tbl =
+ adapter->intf.dispatch_ops_tbl;
+ struct nbl_channel_ops_tbl *chan_ops_tbl =
+ adapter->intf.channel_ops_tbl;
+ struct nbl_dev_mgt *dev_mgt;
+ int ret;
+
+ dev_mgt = nbl_dev_setup_dev_mgt(common);
+ if (IS_ERR(dev_mgt)) {
+ ret = PTR_ERR(dev_mgt);
+ return ret;
+ }
+
+ dev_mgt->disp_ops_tbl = disp_ops_tbl;
+ dev_mgt->chan_ops_tbl = chan_ops_tbl;
+ adapter->core.dev_mgt = dev_mgt;
+ if (common->has_ctrl)
+ nbl_dev_setup_chan_qinfo(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
+ /*
+ * Chip hardware initialization is completed by firmware at power-up.
+ * Only driver functional table/register config follows here, safe to
+ * access hardware registers before ctrl dev setup.
+ */
+ ret = nbl_dev_setup_common_dev(adapter);
+ if (ret)
+ goto setup_err;
+
+ if (common->has_ctrl) {
+ ret = nbl_dev_setup_ctrl_dev(adapter);
+ if (ret)
+ goto setup_ctrl_dev_fail;
+ }
+
+ return 0;
+setup_ctrl_dev_fail:
+ nbl_dev_remove_common_dev(adapter);
+setup_err:
+ return ret;
+}
+
+/*
+ * Teardown order: Stop mailbox channel and drain all inflight DMA first,
+ * then invoke deinit_module to notify firmware.
+ *
+ * This intentionally breaks strict init/teardown mirror symmetry due to
+ * a hardware constraint: firmware performs asynchronous *per-PF* cleanup
+ * (mailbox QINFO routing, qinfo registers and the like) once
+ * driver_status=false is set. It does not revert the chip-wide
+ * datapath registers configured by init_module - those stay valid until
+ * chip reset (see nbl_hw_deinit_module()). We must guarantee no ongoing
+ * mailbox DMA before deinit_module so firmware cannot reclaim the
+ * routing under a live transfer and leave the device DMAing into
+ * released memory.
+ *
+ * Init order: program mailbox QINFO routing (ctrl PF only) ->
+ * create mailbox queue/common_dev -> ctrl dev init_module
+ * Teardown order: destroy mailbox queue/common_dev -> ctrl dev deinit_module
+ */
+void nbl_dev_remove(struct nbl_adapter *adapter)
+{
+ struct nbl_common_info *common = &adapter->common;
+
+ nbl_dev_remove_common_dev(adapter);
+ if (common->has_ctrl)
+ nbl_dev_remove_ctrl_dev(adapter);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h
new file mode 100644
index 000000000000..24e890fd8987
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.h
@@ -0,0 +1,55 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEV_H_
+#define _NBL_DEV_H_
+
+#include <linux/types.h>
+
+#include "../nbl_include/nbl_include.h"
+#include "../nbl_include/nbl_def_channel.h"
+#include "../nbl_include/nbl_def_hw.h"
+#include "../nbl_include/nbl_def_resource.h"
+#include "../nbl_include/nbl_def_dispatch.h"
+#include "../nbl_include/nbl_def_dev.h"
+#include "../nbl_include/nbl_def_common.h"
+#include "../nbl_core.h"
+
+#define NBL_STRING_NAME_LEN 32
+
+enum nbl_msix_serv_type {
+ NBL_MSIX_NET_TYPE,
+ NBL_MSIX_MAILBOX_TYPE,
+ NBL_MSIX_TYPE_MAX
+};
+
+struct nbl_msix_serv_info {
+ char irq_name[NBL_STRING_NAME_LEN];
+ u16 num;
+ u16 base_vector_id;
+ /* true: hw report msix, hw need to mask actively */
+ bool hw_self_mask_en;
+};
+
+struct nbl_msix_info {
+ struct nbl_msix_serv_info serv_info[NBL_MSIX_TYPE_MAX];
+};
+
+struct nbl_dev_common {
+ struct nbl_dev_mgt *dev_mgt;
+ struct nbl_msix_info msix_info;
+ char mailbox_name[NBL_STRING_NAME_LEN];
+ /* for ctrl-dev/net-dev mailbox recv msg */
+ struct work_struct clean_mbx_task;
+};
+
+struct nbl_dev_mgt {
+ struct nbl_common_info *common;
+ struct nbl_dispatch_ops_tbl *disp_ops_tbl;
+ struct nbl_channel_ops_tbl *chan_ops_tbl;
+ struct nbl_dev_common *common_dev;
+};
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
new file mode 100644
index 000000000000..51cf04e4c552
--- /dev/null
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
@@ -0,0 +1,14 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 Nebula Matrix Limited.
+ */
+
+#ifndef _NBL_DEF_DEV_H_
+#define _NBL_DEF_DEV_H_
+
+struct nbl_adapter;
+
+int nbl_dev_init(struct nbl_adapter *adapter);
+void nbl_dev_remove(struct nbl_adapter *adapter);
+
+#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index 86536d582d92..6aacfadd8c0e 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -12,6 +12,7 @@
#include "nbl_include/nbl_def_hw.h"
#include "nbl_include/nbl_def_resource.h"
#include "nbl_include/nbl_def_dispatch.h"
+#include "nbl_include/nbl_def_dev.h"
#include "nbl_include/nbl_def_common.h"
#include "nbl_core.h"
@@ -53,7 +54,14 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
ret = nbl_disp_init(adapter);
if (ret)
goto disp_init_fail;
+
+ ret = nbl_dev_init(adapter);
+ if (ret)
+ goto dev_init_fail;
return adapter;
+
+dev_init_fail:
+ nbl_disp_remove(adapter);
disp_init_fail:
nbl_res_remove_leonis(adapter);
res_init_fail:
@@ -66,6 +74,7 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
void nbl_core_remove(struct nbl_adapter *adapter)
{
+ nbl_dev_remove(adapter);
nbl_disp_remove(adapter);
nbl_res_remove_leonis(adapter);
nbl_chan_remove_common(adapter);
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation
2026-10-10 9:59 ` [PATCH v31 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.
Critical: 0 · High: 1 · Medium: 0 · Low: 3
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] On the control PF, nbl_dev_init() runs its steps in an order
that conflicts with the stale-status recovery in nbl_hw_init_module().
- [Low] On the control PF, nbl_dev_init() writes
NBL_MAILBOX_QINFO_MAP_REG_ARR entries for every enabled PF (via…
- [Low] The error-path comment in nbl_dev_setup_common_dev() says no
unwinding is needed after a failed setup_queue() because the partially…
- [Low] nbl_dev_remove_common_dev() calls
nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX, NULL) right…
--
net/nebula-matrix: add common/ctrl dev init/remove operation
This adds nbl_dev.c with nbl_dev_init() and nbl_dev_remove(). On every
PF they set up the mailbox channel queue and its clean work. On the
control PF they also program the mailbox QINFO routing and run
init_module. The new entry points are called from nbl_core_init() and
nbl_core_remove().
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> new file mode 100644
> index 0000000000000..390bc2b177e21
> --- /dev/null
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
[ ... ]
> +static int nbl_dev_setup_common_dev(struct nbl_adapter *adapter)
> +{
[ ... ]
> + ret = nbl_dev_setup_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
> + if (ret) {
> + /*
> + * Nothing to unwind: chan_info->active is only set on a
> + * successful setup_queue(), so a failed one left no
> + * sender, no registered work item and no mailbox
> + * traffic. Calling teardown_queue() here would only
> + * take its "channel not active" early exit and emit a
> + * misleading duplicate-teardown warning. The partially
> + * allocated DMA rings are intentionally kept until devres
> + * releases them at detach (chan_info->dma_allocated
> + * blocks a re-setup).
> + */
> + return ret;
> + }
[Severity: Low]
Is the "chan_info->dma_allocated blocks a re-setup" part of this comment
accurate? nbl_chan_setup_queue() clears the flag again when
nbl_chan_init_queue() fails:
nbl_channel/nbl_channel.c:nbl_chan_setup_queue() {
...
chan_info->dma_allocated = true;
err = nbl_chan_init_queue(common, chan_info);
if (err) {
chan_info->dma_allocated = false;
return err;
}
...
}
The TX and RX init helpers use dmam_alloc_coherent() and devm_kcalloc().
They free nothing on a partial failure, for example when the TX rings
are allocated and the RX descriptor allocation then fails.
After such a failure, active and dma_allocated are both false. A second
setup_queue() would get past both checks and allocate a second set of
coherent rings, and the first set would be orphaned until detach.
The channel code has its own comment just above that reset ("a retry
must be rejected instead of allocating a second set and orphaning the
first one"), and the reset contradicts it as well.
Nothing retries today, so for now the comment just doesn't match the
code.
> +static void nbl_dev_remove_common_dev(struct nbl_adapter *adapter)
> +{
> + struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
> + struct nbl_dev_common *common_dev = dev_mgt->common_dev;
> +
> + if (!common_dev)
> + return;
> + cancel_work_sync(&common_dev->clean_mbx_task);
[ ... ]
> + nbl_dev_remove_chan_queue(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
> + nbl_dev_register_chan_task(dev_mgt, NBL_CHAN_TYPE_MAILBOX, NULL);
> +}
[Severity: Low]
Does the nbl_dev_register_chan_task(..., NULL) call here have any
effect?
common_dev is only set after setup_queue() succeeds, so the channel is
always active by the time this runs. nbl_chan_teardown_queue() then
sets shutdn and clears clean_task under state_lock. After that,
nbl_chan_register_chan_task() ignores the write:
nbl_channel/nbl_channel.c:nbl_chan_register_chan_task() {
...
mutex_lock(&chan_info->state_lock);
if (!READ_ONCE(chan_info->shutdn))
WRITE_ONCE(chan_info->clean_task, task);
mutex_unlock(&chan_info->state_lock);
}
teardown_queue() does the real unregister and drain: it clears
clean_task and calls cancel_work_sync() on it. The cancel_work_sync()
above runs before teardown, while the sender poll paths can still
re-queue the work.
This has no functional effect today. The problem is that this sequence
reads as if the dev layer does the unregister and drain, when the
channel layer actually does it. Could the NULL registration be dropped,
or a comment added saying which layer is responsible?
> +int nbl_dev_init(struct nbl_adapter *adapter)
> +{
[ ... ]
> + adapter->core.dev_mgt = dev_mgt;
> + if (common->has_ctrl)
> + nbl_dev_setup_chan_qinfo(dev_mgt, NBL_CHAN_TYPE_MAILBOX);
> + /*
> + * Chip hardware initialization is completed by firmware at power-up.
> + * Only driver functional table/register config follows here, safe to
> + * access hardware registers before ctrl dev setup.
> + */
> + ret = nbl_dev_setup_common_dev(adapter);
> + if (ret)
> + goto setup_err;
> +
> + if (common->has_ctrl) {
> + ret = nbl_dev_setup_ctrl_dev(adapter);
> + if (ret)
> + goto setup_ctrl_dev_fail;
> + }
[Severity: High]
Can this ordering make firmware wipe the mailbox state that this probe
has just programmed?
On the control PF the sequence is:
nbl_dev_init()
nbl_dev_setup_chan_qinfo() /* QINFO routing for all PFs */
nbl_dev_setup_common_dev()
nbl_chan_setup_queue() /* TX/RX queue config, active=true */
nbl_dev_setup_ctrl_dev()
init_module -> nbl_hw_init_module()
If the active bit in NBL_DRIVER_STATUS_REG is still set,
nbl_hw_init_module() forces it from 1 to 0:
nbl_hw/nbl_hw_leonis/nbl_hw_leonis.c:nbl_hw_init_module() {
...
status = nbl_hw_rd32(hw_mgt, NBL_DRIVER_STATUS_REG);
if (status & BIT(NBL_DRIVER_STATUS_BIT)) {
dev_warn(hw_mgt->common->dev,
"driver_status already set at init, cleanup forced\n");
nbl_hw_set_driver_status(hw_mgt, false);
nbl_flush_writes(hw_mgt);
usleep_range(NBL_FW_CLEANUP_SYNC_MIN_US,
NBL_FW_CLEANUP_SYNC_MAX_US);
...
}
The bit can still be set after kexec or kdump, because nbl_pci_driver
has no .shutdown callback. It can also be set after an unload that
never reached deinit_module().
The commit message says firmware asynchronously reclaims "mailbox QINFO
routing, qinfo registers and the like" when the bit goes from 1 to 0.
Here that happens after the new routing and the live queue registers are
in place. Nothing later in init, or in the start path, programs them
again.
Wouldn't that leave the control PF mailbox dead, so that sibling-PF RPCs
in nbl_dev_start() time out? Firmware could also reclaim the routing
while a mailbox transfer is in flight, which is the hazard the teardown
ordering is meant to avoid.
Should the stale-status cleanup run before nbl_dev_setup_chan_qinfo()
and setup_queue()?
> +
> + return 0;
> +setup_ctrl_dev_fail:
> + nbl_dev_remove_common_dev(adapter);
> +setup_err:
> + return ret;
> +}
[Severity: Low]
Do these error paths leave the NBL_MAILBOX_QINFO_MAP_REG_ARR entries
written by nbl_dev_setup_chan_qinfo() in hardware?
The comment above nbl_dev_remove_ctrl_dev() says the map "is not
explicitly cleared; firmware handles it on driver_status change". But
neither label clears the entries or calls deinit_module().
This applies in two cases:
- nbl_dev_setup_common_dev() fails.
- nbl_hw_init_module() fails, either in the eth_speed/eth_num checks or
in nbl_dp_init(). Both return before
nbl_hw_set_driver_status(true).
In both cases driver_status never goes from 1 to 0, so the firmware
reclaim never runs. The routing entries stay until a later successful
probe overwrites them or the chip is reset.
The practical impact looks small, since sibling PFs can't bind without
the control PF. The hardware state still no longer matches what the
driver thinks it is.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH v31 net-next 8/8] net/nebula-matrix: add common dev start/stop operation
2026-10-10 9:59 [PATCH v31 net-next 0/8] nbl driver for Nebulamatrix NICs illusion.wang
` (6 preceding siblings ...)
2026-10-10 9:59 ` [PATCH v31 net-next 7/8] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
@ 2026-10-10 9:59 ` illusion.wang
2026-10-11 10:51 ` netdev-bot+sashiko
7 siblings, 1 reply; 17+ messages in thread
From: illusion.wang @ 2026-10-10 9:59 UTC (permalink / raw)
To: dimon.zhao, illusion.wang, alvin.wang, sam.chen, netdev
Cc: kuba, edumazet, horms, open list
From: illusion wang <illusion.wang@nebula-matrix.com>
Add nbl_dev_start() and nbl_dev_stop(), called from nbl_probe() and
nbl_remove() via nbl_core_start() / nbl_core_stop().
Start configures the device MSI-X map, fetches the VSI/ETH ids, allocates
the MSI-X vectors and arms the mailbox interrupt; stop performs the
matching teardown.
The start/stop pair is single-shot and non-repeatable, tied strictly to
the PCI probe/remove lifecycle: kernel vectors are released only by the
pcim_msi_release devres at detach, so a second start is rejected.
Startup RPC errors are propagated unchanged. A -ETIMEDOUT here is a real
failure, not a "supplier not ready yet" condition: nbl_probe() already
deferred this PF in nbl_probe_chip_deps() until the control PF completed
probing, and that completion includes the control PF's own start, whose
mailbox responder has been live since nbl_dev_init(). Converting the
timeout to -EPROBE_DEFER would only hide it behind dev_dbg() and make
every later bind pay the multi-second wait again.
Teardown is best-effort. The MSI-X map is torn down even when the
configure RPC failed: a timeout only means the ACK never arrived, while
the responder runs to completion before ACKing, so the entry, the
intr_net/other_bmap bits and the coherent map table stay owned by the
control PF. Those are host resources of that driver, freed by a later
successful configuration of this function or by nbl_intr_mgt_stop() - not
firmware state. pci_clear_master() is not a safety net for this: on a
non-control PF it cannot stop DMA issued with the management PF's BDF.
Signed-off-by: illusion wang <illusion.wang@nebula-matrix.com>
---
.../net/ethernet/nebula-matrix/nbl/nbl_core.h | 8 +
.../nebula-matrix/nbl/nbl_core/nbl_dev.c | 391 ++++++++++++++++++
.../nbl/nbl_include/nbl_def_dev.h | 2 +
.../net/ethernet/nebula-matrix/nbl/nbl_main.c | 18 +
4 files changed, 419 insertions(+)
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
index c3c4dd685bf6..655dfb43e365 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core.h
@@ -40,4 +40,12 @@ struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
struct nbl_init_param *param);
void nbl_core_remove(struct nbl_adapter *adapter);
+/*
+ * Single-shot start/stop pair, called once each from PCI probe/remove.
+ * Not repeatable: MSI-X vectors stay allocated until device detach, so
+ * a second start on a bound device is rejected by the MSI-X core.
+ */
+int nbl_core_start(struct nbl_adapter *adapter);
+void nbl_core_stop(struct nbl_adapter *adapter);
+
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
index 390bc2b177e2..41754fd06ea3 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
@@ -6,6 +6,17 @@
#include <linux/pci.h>
#include "nbl_dev.h"
+static void nbl_dev_clean_mailbox_schedule(struct nbl_dev_mgt *dev_mgt);
+
+/* ---------- Interrupt config ---------- */
+static irqreturn_t nbl_dev_clean_mailbox(int irq __always_unused, void *data)
+{
+ struct nbl_dev_mgt *dev_mgt = (struct nbl_dev_mgt *)data;
+
+ nbl_dev_clean_mailbox_schedule(dev_mgt);
+ return IRQ_HANDLED;
+}
+
static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
{
struct nbl_dev_common *dev_common = dev_mgt->common_dev;
@@ -18,6 +29,230 @@ static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num = 1;
}
+static int nbl_dev_request_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+ struct nbl_msix_info *msix_info = &dev_common->msix_info;
+ struct nbl_common_info *common = dev_mgt->common;
+ u16 lvec;
+ int irq_num;
+ int err;
+
+ if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+ return 0;
+
+ lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+ irq_num = pci_irq_vector(common->pdev, lvec);
+ if (irq_num < 0) {
+ dev_err(common->dev, "Failed to get mailbox IRQ vector: %d\n",
+ irq_num);
+ return irq_num;
+ }
+
+ snprintf(dev_common->mailbox_name, sizeof(dev_common->mailbox_name),
+ "nbl_mailbox@pci:%s", pci_name(common->pdev));
+ err = request_irq(irq_num, nbl_dev_clean_mailbox, 0,
+ dev_common->mailbox_name, dev_mgt);
+ if (err)
+ return err;
+
+ return 0;
+}
+
+static void nbl_dev_free_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+ struct nbl_msix_info *msix_info = &dev_common->msix_info;
+ struct nbl_common_info *common = dev_mgt->common;
+ u16 lvec;
+ int irq_num;
+
+ if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+ return;
+
+ lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+ irq_num = pci_irq_vector(common->pdev, lvec);
+ if (irq_num >= 0)
+ free_irq(irq_num, dev_mgt);
+}
+
+static int nbl_dev_enable_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+ struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+ struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+ struct nbl_msix_info *msix_info = &dev_common->msix_info;
+ u16 lvec;
+ int ret;
+
+ if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+ return 0;
+
+ lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+ /*
+ * Enable sequence: perform set_mailbox_irq RPC in polling mode first.
+ * Only set NBL_CHAN_IRQ_RDY after RPC succeeds, mirroring disable path.
+ * This avoids waiting for an interrupt which has not been armed yet.
+ */
+ ret = disp_ops->set_mailbox_irq(dev_mgt->disp_ops_tbl->priv,
+ lvec, true);
+ if (ret)
+ return ret;
+ chan_ops->set_queue_state(dev_mgt->chan_ops_tbl->priv,
+ NBL_CHAN_IRQ_RDY,
+ NBL_CHAN_TYPE_MAILBOX, true);
+ return 0;
+}
+
+static int nbl_dev_disable_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+ struct nbl_channel_ops *chan_ops = dev_mgt->chan_ops_tbl->ops;
+ struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+ struct nbl_msix_info *msix_info = &dev_common->msix_info;
+ u16 lvec;
+
+ if (!msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num)
+ return 0;
+
+ lvec = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].base_vector_id;
+ /*
+ * Disable sequence invariant: clear the software READY flag first,
+ * then mask the hardware interrupt. Must not reverse the order.
+ *
+ * Clearing NBL_CHAN_IRQ_RDY is what makes an interrupt-mode sender
+ * fail fast: nbl_chan_set_queue_state() wakes every per-slot wait
+ * queue on the transition, and nbl_chan_send_msg() tests the flag
+ * ahead of its timeout branch, so the sender returns -EIO right
+ * away instead of waiting out NBL_CHAN_ACK_WAIT_TIME (3s) and
+ * reporting a bogus -ETIMEDOUT / "Channel waiting ack failed".
+ * Masking first would leave the sleeper blocked on an interrupt
+ * that is already masked, i.e. guaranteed to time out.
+ *
+ * This does NOT preserve in-flight ACKs, and must not be described
+ * as doing so: the woken sender gives up and resets its slot to
+ * IDLE, so an ACK arriving afterwards finds a non-WAITING slot and
+ * is dropped by the receive path ("Skip ack invalid status").
+ * Interrupt mode is being torn down here; the guarantee is a
+ * bounded, deterministic failure, not delivery. In this driver
+ * every ACK-requesting sender runs from probe/remove and the clean
+ * task only sends ACK-less replies, so nothing depends on the
+ * stronger claim today.
+ *
+ * This helper is invoked in two paths:
+ * 1. Error unwind path of nbl_dev_start(): not reached with
+ * IRQ_RDY set (nbl_dev_enable_mailbox_irq() sets it only after
+ * its RPC succeeded), so only the hardware mask still applies,
+ * followed by nbl_dev_free_mailbox_irq() and full channel
+ * teardown.
+ * 2. Normal device stop path nbl_dev_stop(): free_irq() blocks until
+ * any in-flight hardirq handler completes and prevents new interrupts.
+ * cancel_work_sync() then waits for any already running mailbox cleanup
+ * work to finish, or cancels queued but unstarted work items before
+ * final channel destruction. No stuck descriptors linger in either
+ * scenario.
+ */
+ chan_ops->set_queue_state(dev_mgt->chan_ops_tbl->priv,
+ NBL_CHAN_IRQ_RDY,
+ NBL_CHAN_TYPE_MAILBOX, false);
+
+ return disp_ops->set_mailbox_irq(dev_mgt->disp_ops_tbl->priv,
+ lvec, false);
+}
+
+static int nbl_dev_cfg_msix_map(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+ struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+ struct nbl_msix_info *msix_info = &dev_common->msix_info;
+ bool mask_en = msix_info->serv_info[NBL_MSIX_NET_TYPE].hw_self_mask_en;
+ u16 msix_net_num = msix_info->serv_info[NBL_MSIX_NET_TYPE].num;
+ u16 msix_not_net_num = 0;
+ int err, i;
+
+ msix_info->serv_info[NBL_MSIX_NET_TYPE].base_vector_id = 0;
+ /*
+ * Calculate base_vector_id for each MSIX service type.
+ * This relies on NBL_MSIX_TYPE enum being ordered sequentially,
+ * starting from NBL_MSIX_NET_TYPE.
+ */
+ for (i = NBL_MSIX_NET_TYPE + 1; i < NBL_MSIX_TYPE_MAX; i++)
+ msix_info->serv_info[i].base_vector_id =
+ msix_info->serv_info[i - 1].base_vector_id +
+ msix_info->serv_info[i - 1].num;
+
+ for (i = 0; i < NBL_MSIX_TYPE_MAX; i++) {
+ if (i == NBL_MSIX_NET_TYPE)
+ continue;
+ msix_not_net_num += msix_info->serv_info[i].num;
+ }
+
+ err = disp_ops->cfg_msix_map(dev_mgt->disp_ops_tbl->priv,
+ msix_net_num, msix_not_net_num,
+ mask_en);
+
+ return err;
+}
+
+static int nbl_dev_destroy_msix_map(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+
+ return disp_ops->destroy_msix_map(dev_mgt->disp_ops_tbl->priv);
+}
+
+static int nbl_dev_init_interrupt_scheme(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dev_common *dev_common = dev_mgt->common_dev;
+ struct nbl_msix_info *msix_info = &dev_common->msix_info;
+ struct nbl_common_info *common = dev_mgt->common;
+ struct irq_affinity affd = { 0 };
+ int needed = 0;
+ int err;
+ int i;
+
+ for (i = 0; i < NBL_MSIX_TYPE_MAX; i++)
+ needed += msix_info->serv_info[i].num;
+
+ /*
+ * The mailbox vector is the trailing administrative vector;
+ * reserve it via post_vectors so it is excluded from managed
+ * affinity spreading. A managed vector can be shut down when
+ * its last assigned CPU goes offline, which would stall every
+ * mailbox RPC while NBL_CHAN_IRQ_RDY stays set.
+ */
+ affd.post_vectors = msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num;
+
+ err = pci_alloc_irq_vectors_affinity(common->pdev, needed, needed,
+ PCI_IRQ_MSIX | PCI_IRQ_AFFINITY,
+ &affd);
+ if (err < 0) {
+ dev_err(common->dev,
+ "pci_alloc_irq_vectors failed, err = %d\n", err);
+ return err;
+ }
+ if (err != needed) {
+ dev_err(common->dev, "pci_alloc_irq_vectors got %d vecs, need %d\n",
+ err, needed);
+ return -ENOSPC;
+ }
+ return 0;
+}
+
+/*
+ * Kernel-side MSI-X vectors are deliberately NOT freed here. They are
+ * released at device detach by the pcim_msi_release devres callback
+ * registered by pci_alloc_irq_vectors() (probe uses pcim_enable_device);
+ * freeing them here would double-free that callback.
+ *
+ * Consequently start/stop is a single-shot pair tied to probe/remove:
+ * vectors stay allocated until detach, and a second start against a
+ * device with msix_enabled set is rejected by __pci_enable_msix_range().
+ */
+static void nbl_dev_clear_interrupt_scheme(struct nbl_dev_mgt *dev_mgt)
+{
+}
+
/* ---------- Channel config ---------- */
static void nbl_dev_setup_chan_qinfo(struct nbl_dev_mgt *dev_mgt, u8 chan_type)
{
@@ -85,6 +320,14 @@ static void nbl_dev_clean_mailbox_task(struct work_struct *work)
NBL_CHAN_TYPE_MAILBOX);
}
+static void nbl_dev_clean_mailbox_schedule(struct nbl_dev_mgt *dev_mgt)
+{
+ struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+ struct nbl_common_info *common = dev_mgt->common;
+
+ queue_work(common->wq, &common_dev->clean_mbx_task);
+}
+
/* ---------- Dev init process ---------- */
static int nbl_dev_setup_common_dev(struct nbl_adapter *adapter)
{
@@ -274,3 +517,151 @@ void nbl_dev_remove(struct nbl_adapter *adapter)
if (common->has_ctrl)
nbl_dev_remove_ctrl_dev(adapter);
}
+
+/* ---------- Dev start process ---------- */
+int nbl_dev_start(struct nbl_adapter *adapter)
+{
+ struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+ struct nbl_dispatch_ops *disp_ops = dev_mgt->disp_ops_tbl->ops;
+ struct nbl_dispatch_mgt *priv = dev_mgt->disp_ops_tbl->priv;
+ struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+ struct nbl_common_info *common = dev_mgt->common;
+ int cleanup_ret;
+ int ret;
+
+ /*
+ * Startup RPC failures are reported as they are, including the
+ * -ETIMEDOUT of a 5-6s polling wait against an unresponsive peer.
+ * nbl_probe() already deferred this PF until the control PF had
+ * completed probing (nbl_probe_chip_deps()), and that completion
+ * includes the control PF's own nbl_core_start(); its mailbox
+ * responder has been live since nbl_dev_init(). A timeout here
+ * is therefore a real failure, not a "supplier not ready yet"
+ * condition. Turning it into -EPROBE_DEFER would hide it behind
+ * dev_dbg() and leave the retry to the next unrelated device
+ * bind, paying the multi-second timeout again on every retry.
+ */
+ ret = nbl_dev_cfg_msix_map(dev_mgt);
+ if (ret)
+ goto err_destroy_map;
+
+ /* Fetch VSI/ETH identity after cfg_msix_map */
+ ret = disp_ops->get_vsi_id(priv, NBL_VSI_DATA, &common->vsi_id);
+ if (ret)
+ goto err_destroy_map;
+ ret = disp_ops->get_eth_id(priv, common->vsi_id, &common->eth_num,
+ &common->eth_id, &common->logic_eth_id);
+ if (ret)
+ goto err_destroy_map;
+
+ ret = nbl_dev_init_interrupt_scheme(dev_mgt);
+ if (ret)
+ goto err_destroy_map;
+
+ ret = nbl_dev_request_mailbox_irq(dev_mgt);
+ if (ret)
+ goto err_destroy_map;
+
+ ret = nbl_dev_enable_mailbox_irq(dev_mgt);
+ if (ret)
+ goto err_free_irq;
+
+ return 0;
+
+err_free_irq:
+ /*
+ * nbl_dev_enable_mailbox_irq() only fails before it sets
+ * NBL_CHAN_IRQ_RDY - the flag is set after its RPC returns - so
+ * there is no interrupt-mode sender to release and no software
+ * state to revoke. The peer-side mailbox routing, if the failed
+ * RPC did land, is switched off by the destroy_msix_map() below,
+ * whose prepare phase disables mailbox IRQ routing first. Sending
+ * set_mailbox_irq(false) from here would only add one more full
+ * polling-mode ACK timeout against a peer that just proved
+ * unresponsive.
+ */
+ nbl_dev_free_mailbox_irq(dev_mgt);
+err_destroy_map:
+ /*
+ * Always ask the peer to drop this function's MSI-X map, even when
+ * the cfg RPC above failed: the teardown prepare phase returns 0
+ * for a function that is not NBL_INTR_FUNC_CONFIGURED, so it is a
+ * no-op when nothing landed, but a -ETIMEDOUT only means the ACK
+ * never arrived - the responder runs to completion before it ACKs.
+ * In that case the map entry, the intr_net/other_bmap bits, the
+ * kcalloc()ed vector array and the coherent map table are all
+ * still owned by the control PF, i.e. host resources of its
+ * driver, released only when this function is configured again or
+ * when nbl_intr_mgt_stop() tears the control PF down. They are not
+ * firmware state, so skipping the destroy would strand them for a
+ * PF that has no driver to ever ask for them again.
+ *
+ * The destroy runs before detach even though kernel vectors stay
+ * allocated until then (see nbl_dev_clear_interrupt_scheme): it
+ * masks all hardware vectors and clears the pcompleter map entry,
+ * so no MSI-X message can be delivered after vectors are released
+ * at detach.
+ *
+ * For non-control PFs this is a polling-mode mailbox RPC:
+ * NBL_CHAN_IRQ_RDY is never set on any rollback path, since
+ * enabling it is the last step of nbl_dev_start().
+ *
+ * Note: pci_clear_master() runs in nbl_probe() core_start_err path;
+ * on a non-control PF it cannot stop DMA using the management PF's
+ * BDF, this is a known limitation.
+ */
+ cleanup_ret = nbl_dev_destroy_msix_map(dev_mgt);
+ if (cleanup_ret)
+ dev_err(dev_mgt->common->dev,
+ "rollback: destroy MSI-X map failed: %d\n",
+ cleanup_ret);
+ nbl_dev_clear_interrupt_scheme(dev_mgt);
+
+ cancel_work_sync(&common_dev->clean_mbx_task);
+ return ret;
+}
+
+void nbl_dev_stop(struct nbl_adapter *adapter)
+{
+ struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
+ struct nbl_dev_common *common_dev = dev_mgt->common_dev;
+ int ret;
+
+ ret = nbl_dev_disable_mailbox_irq(dev_mgt);
+ if (ret)
+ dev_err(dev_mgt->common->dev,
+ "Failed to disable mailbox IRQ: %d\n", ret);
+ nbl_dev_free_mailbox_irq(dev_mgt);
+
+ /*
+ * Destroy the hardware MSI-X map now, before the kernel vectors
+ * are released at device detach (they deliberately stay
+ * allocated until then - see nbl_dev_clear_interrupt_scheme).
+ * Masks all device vectors and clears the pcompleter map entry.
+ *
+ * Best-effort: if the destroy RPC fails, the control PF still owns
+ * this function's map entry, its intr_net/other_bmap bits and its
+ * coherent map table. Those are host resources of that driver,
+ * released either by a later successful configuration of this
+ * function (reconfiguration recycles them) or by
+ * nbl_intr_mgt_stop() when the control PF is removed; nothing else
+ * collects them, so the hardware entry stays as it is until chip
+ * reset - and nothing requests one.
+ *
+ * pci_clear_master() on a non-control PF cannot stop DMA using the
+ * management PF's BDF.
+ */
+ ret = nbl_dev_destroy_msix_map(dev_mgt);
+ if (ret)
+ dev_err(dev_mgt->common->dev,
+ "Failed to destroy MSI-X map: %d\n", ret);
+
+ nbl_dev_clear_interrupt_scheme(dev_mgt);
+
+ /*
+ * destroy_msix_map() sends ack-requested messages which may
+ * requeue clean_mbx_task via polling send path. Drain work
+ * after the operation.
+ */
+ cancel_work_sync(&common_dev->clean_mbx_task);
+}
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
index 51cf04e4c552..a66c633a0e7a 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_include/nbl_def_dev.h
@@ -10,5 +10,7 @@ struct nbl_adapter;
int nbl_dev_init(struct nbl_adapter *adapter);
void nbl_dev_remove(struct nbl_adapter *adapter);
+int nbl_dev_start(struct nbl_adapter *adapter);
+void nbl_dev_stop(struct nbl_adapter *adapter);
#endif
diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
index 6aacfadd8c0e..cacf4b680aad 100644
--- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
+++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
@@ -16,6 +16,16 @@
#include "nbl_include/nbl_def_common.h"
#include "nbl_core.h"
+int nbl_core_start(struct nbl_adapter *adapter)
+{
+ return nbl_dev_start(adapter);
+}
+
+void nbl_core_stop(struct nbl_adapter *adapter)
+{
+ nbl_dev_stop(adapter);
+}
+
struct nbl_adapter *nbl_core_init(struct pci_dev *pdev,
struct nbl_init_param *param)
{
@@ -234,7 +244,13 @@ static int nbl_probe(struct pci_dev *pdev,
goto adapter_init_err;
}
pci_set_drvdata(pdev, adapter);
+ err = nbl_core_start(adapter);
+ if (err)
+ goto core_start_err;
+
return 0;
+core_start_err:
+ nbl_core_remove(adapter);
adapter_init_err:
chip_deps_err:
pci_clear_master(pdev);
@@ -248,6 +264,8 @@ static void nbl_remove(struct pci_dev *pdev)
if (!adapter)
return;
pci_set_drvdata(pdev, NULL);
+
+ nbl_core_stop(adapter);
nbl_core_remove(adapter);
pci_clear_master(pdev);
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH v31 net-next 8/8] net/nebula-matrix: add common dev start/stop operation
2026-10-10 9:59 ` [PATCH v31 net-next 8/8] net/nebula-matrix: add common dev start/stop operation illusion.wang
@ 2026-10-11 10:51 ` netdev-bot+sashiko
0 siblings, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-11 10:51 UTC (permalink / raw)
To: illusion.wang
Cc: dimon.zhao, alvin.wang, sam.chen, netdev, kuba, edumazet, horms,
linux-kernel
Thank you for your contribution! Sashiko AI review found 3 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 2
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] nbl_probe() now calls nbl_core_start() -> nbl_dev_start() for
every adapter that nbl_core_init() returns.
- [Low] In nbl_dev_disable_mailbox_irq() (drivers/net/ethernet/nebula-
matrix/nbl/nbl_core/nbl_dev.c), the comment says the helper runs in…
- [Low] The err_free_irq comment in nbl_dev_start() (nbl_dev.c) explains
skipping set_mailbox_irq(false) by saying it 'would only add one more…
--
net/nebula-matrix: add common dev start/stop operation
This patch adds nbl_dev_start() and nbl_dev_stop(). They are called from
nbl_probe() and nbl_remove() through nbl_core_start() and nbl_core_stop().
Start configures the MSI-X map, fetches the VSI/ETH ids, allocates the
MSI-X vectors and arms the mailbox interrupt; stop does the matching
best-effort teardown.
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> index 390bc2b177e21..41754fd06ea36 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
[ ... ]
> @@ -18,6 +29,230 @@ static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
[ ... ]
> +static int nbl_dev_disable_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
> +{
[ ... ]
> + * This helper is invoked in two paths:
> + * 1. Error unwind path of nbl_dev_start(): not reached with
> + * IRQ_RDY set (nbl_dev_enable_mailbox_irq() sets it only after
> + * its RPC succeeded), so only the hardware mask still applies,
> + * followed by nbl_dev_free_mailbox_irq() and full channel
> + * teardown.
[Severity: Low]
Is the description of the first path accurate? nbl_dev_start() never
calls nbl_dev_disable_mailbox_irq(). Its err_free_irq label calls only
nbl_dev_free_mailbox_irq() and then falls through to err_destroy_map:
err_free_irq:
nbl_dev_free_mailbox_irq(dev_mgt);
err_destroy_map:
cleanup_ret = nbl_dev_destroy_msix_map(dev_mgt);
...
nbl_dev_clear_interrupt_scheme(dev_mgt);
cancel_work_sync(&common_dev->clean_mbx_task);
This patch also adds the err_free_irq comment, which says
set_mailbox_irq(false) is not sent from that path on purpose. That
contradicts this comment.
nbl_dev_start() does no channel teardown either. That only happens
later, when nbl_probe() calls nbl_core_remove().
nbl_dev_stop() seems to be the only caller of this helper. The rollback
path depends on the routing disable in
__nbl_res_intr_prepare_destroy_msix_map(), not on this helper.
Could the comment list only nbl_dev_stop()?
> + * 2. Normal device stop path nbl_dev_stop(): free_irq() blocks until
[ ... ]
> +static int nbl_dev_cfg_msix_map(struct nbl_dev_mgt *dev_mgt)
> +{
[ ... ]
> + err = disp_ops->cfg_msix_map(dev_mgt->disp_ops_tbl->priv,
> + msix_net_num, msix_not_net_num,
> + mask_en);
[Severity: Medium]
Can disp_ops->cfg_msix_map be NULL here?
nbl_probe() now calls nbl_core_start()->nbl_dev_start() for every
adapter that nbl_core_init() returns. nbl_dev_start() then calls
cfg_msix_map, get_vsi_id and get_eth_id without checking capabilities or
NULL. On the unwind path it also calls destroy_msix_map.
nbl_disp_refresh_ctrl_ops() memsets the dispatch ops table. It fills the
table only when NBL_DISP_CTRL_LVL_MGT (has_ctrl) or NBL_DISP_CTRL_LVL_NET
(has_net) is set.
The struct nbl_dispatch_ops kerneldoc in nbl_def_dispatch.h says:
A function with neither has_ctrl nor has_net leaves these pointers
NULL; callers must not invoke them on such functions.
That state looks reachable. Suppose root binds a VF, or a function other
than func 0, whose ID is not in nbl_id_table, by writing
driver_override=nbl. pci_match_device() then passes the dummy ID:
drivers/pci/pci-driver.c:pci_match_device() {
...
/* driver_override will always match, send a dummy id */
if (ret > 0)
return &pci_device_id_any;
...
}
The dummy ID has driver_data 0. nbl_get_func_param() therefore sets
has_net to 0, and has_ctrl stays 0:
param->caps.has_net = !!(driver_data & BIT(NBL_CAP_HAS_NET_BIT));
new_id cannot produce this state, because new_id_store() rejects a
driver_data value that matches no table entry.
Now assume func 0 of the same slot is bound to nbl, which
nbl_probe_chip_deps() requires, and the init steps succeed. Wouldn't
probe then call a NULL function pointer?
nbl_probe()
nbl_core_start()
nbl_dev_start()
nbl_dev_cfg_msix_map()
disp_ops->cfg_msix_map() <- NULL
Before this patch, nbl_dev_init() used disp_ops only under has_ctrl, so
this path is new. Would it make sense to reject !has_ctrl && !has_net
early in nbl_probe()?
[ ... ]
> @@ -274,3 +517,151 @@ void nbl_dev_remove(struct nbl_adapter *adapter)
[ ... ]
> + ret = nbl_dev_enable_mailbox_irq(dev_mgt);
> + if (ret)
> + goto err_free_irq;
> +
> + return 0;
> +
> +err_free_irq:
> + /*
> + * nbl_dev_enable_mailbox_irq() only fails before it sets
> + * NBL_CHAN_IRQ_RDY - the flag is set after its RPC returns - so
> + * there is no interrupt-mode sender to release and no software
> + * state to revoke. The peer-side mailbox routing, if the failed
> + * RPC did land, is switched off by the destroy_msix_map() below,
> + * whose prepare phase disables mailbox IRQ routing first. Sending
> + * set_mailbox_irq(false) from here would only add one more full
> + * polling-mode ACK timeout against a peer that just proved
> + * unresponsive.
> + */
> + nbl_dev_free_mailbox_irq(dev_mgt);
[Severity: Low]
This isn't a bug, but does the last sentence of this comment hold for
every failure? It assumes nbl_dev_enable_mailbox_irq() can only fail
through an ACK timeout against an unresponsive peer.
On the control PF, disp_ops->set_mailbox_irq is nbl_disp_set_mailbox_irq().
That is a direct local call: no mailbox, no ACK and no timeout.
if (!res_ops->set_mailbox_irq)
return -EOPNOTSUPP;
return res_ops->set_mailbox_irq(p, common->mgt_pf, vector_id, en_msix);
The resource layer, __nbl_res_intr_set_mailbox_irq(), can fail without
touching hardware:
-ESHUTDOWN if intr_mgt->stopping is set
-ENODEV if the function is not NBL_INTR_FUNC_CONFIGURED
-EINVAL if vector_id is out of range
The responder passes these errors to non-control PFs unchanged.
Skipping set_mailbox_irq(false) still looks correct, because the routing
disable in __nbl_res_intr_prepare_destroy_msix_map() covers a CONFIGURED
function. Could the comment give that as the reason instead?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261010095939.2230-1-illusion.wang%40nebula-matrix.com
^ permalink raw reply [flat|nested] 17+ messages in thread