mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v4 0/3] dinghai: firmware handshake, MSI-X pools and async event queues
@ 2026-09-28 12:20 han.junyang
  2026-09-28 12:24 ` [PATCH net-next v4 1/3] dinghai: add firmware version check and RISC-V readiness polling han.junyang
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: han.junyang @ 2026-09-28 12:20 UTC (permalink / raw)
  To: andrew+netdev, davem, edumazet, kuba, pabeni, horms
  Cc: linux-kernel, netdev, han.junyang, ran.ming, han.chengfei, zhang.yanze

From: Junyang Han <han.junyang@zte.com.cn>

This series continues the DingHai (ZXDH) PF driver bring-up: after PCI
probing, it verifies the firmware version contract, waits for the
RISC-V management core to become ready, sets up the MSI-X interrupt
pools, and creates the async event queues through which the firmware
will report events.

Some notes on the IRQ design:

The vector space is partitioned into per-purpose pools (async, RDMA,
vq). Event queues share a vector through an atomic notifier chain
attached to the IRQ, and the sharing degree is governed by pool
thresholds. The IRQ table hangs off a void *priv in the shared core
device: the PF, MPF and SF core devices each carry a different pool
layout behind that pointer, so a type-specific struct in en_pf.c keeps
the shared header free of PF-only details.

Changes in v4:
- Fold the irq and eq table create() steps into their init()
  counterparts and route every probe error path through the same
  destroy() functions remove() uses, so cleanup has a single
  authority from the start (review feedback on v3).
- Give the pool name a buffer bound of its own (16 bytes) instead of
  sharing the 100-byte IRQ name buffer, so W=1 builds stay clean of
  -Wformat-truncation (review feedback on v3).
- Lower the compat region wait from 200 s to 20 s: the firmware
  publishes the block within 10 s of boot, so twice that leaves
  margin, and the wait now sits far below the 180 s udev event
  window raised on the thread (review feedback on v3).
- Wait for the compat block with a per-field readiness check instead
  of only the first dword: the firmware fills the fields one by one,
  and a dword-granular check can read a half-populated block (the
  "single write" justification given in v3 was wrong).
- Log the legacy-firmware assumption when the compat region wait
  times out, making the intentionally ignored timeout visible.

Changes in v3:
- Unmap the modern config MMIO regions on the probe error paths added
  in this series; they used to leak when the fw compat check or the
  RISC-V readiness wait failed.
- Skip the RISC-V readiness wait for firmware without the compat
  region: the erased patch field read as 0xffff and defeated the skip.
- Make fw_minor unsigned so firmware minor versions >= 128 are not
  rejected through sign wrap.
- Balance the per-CPU IRQ accounting on release, so pool teardown does
  not trip its leftover WARN; set the IRQ affinity for real with
  irq_set_affinity_and_hint() instead of only updating the hint; name
  IRQs after their pool instead of a hardcoded prefix; assert that the
  pool is empty at free instead of force-releasing leftovers.
- Clear eq->irq when the async IRQ request fails, so teardown does not
  treat the ERR_PTR as a live IRQ.

Review findings not taken:

- xa_alloc() with a NULL entry does not fail: __xa_alloc() turns a
  NULL entry into the internal zero entry as a reservation
  (lib/xarray.c), which is the reservation semantics the pool relies
  on.
- BAR 0 length checks against a truncated bar: the bar layout is part
  of the board firmware contract and the driver does not defend
  against a broken device, as settled during the review of the
  earlier device bring-up series (Andrew Lunn).

Changes in v2:
- Convert both firmware readiness polls to readx_poll_timeout()
  (Andrew Lunn).
- Read the fw compat block field by field through ioread*() accessors
  instead of an ioread32_rep() bulk copy, which misplaces the u8/u16
  fields on big-endian; no __le annotations are needed since each
  accessor converts from little-endian (Andrew Lunn).
- Use kref for the IRQ reference count (Andrew Lunn).

Junyang Han (3):
  dinghai: add firmware version check and RISC-V readiness polling
  dinghai: add MSI-X interrupt pools
  dinghai: add async event queue for firmware notifications

 drivers/net/ethernet/zte/dinghai/Makefile   |   2 +-
 drivers/net/ethernet/zte/dinghai/en_pf.c    | 223 ++++++++++
 drivers/net/ethernet/zte/dinghai/en_pf.h    | 56 +++
 drivers/net/ethernet/zte/dinghai/zxdh_eq.c  | 137 +++++++
 drivers/net/ethernet/zte/dinghai/zxdh_eq.h  | 60 +++
 drivers/net/ethernet/zte/dinghai/zxdh_irq.c | 406 ++++++++++++++++++++
 drivers/net/ethernet/zte/dinghai/zxdh_irq.h | 73 ++++
 7 files changed, 956 insertions(+), 1 deletion(-)
 create mode 100644 drivers/net/ethernet/zte/dinghai/zxdh_eq.c
 create mode 100644 drivers/net/ethernet/zxdh_eq.h
 create mode 100644 drivers/net/ethernet/zxdh_irq.c
 create mode 100644 drivers/net/ethernet/zxdh_irq.h

--
2.27.0

^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH net-next v4 1/3] dinghai: add firmware version check and RISC-V  readiness polling
  2026-09-28 12:20 [PATCH net-next v4 0/3] dinghai: firmware handshake, MSI-X pools and async event queues han.junyang
@ 2026-09-28 12:24 ` han.junyang
  2026-10-01 13:06   ` Simon Horman
  2026-09-28 12:27 ` [PATCH net-next v4 2/3] dinghai: add MSI-X interrupt pools han.junyang
  2026-09-28 12:33 ` [PATCH net-next v4 3/3] dinghai: add async event queue for firmware notifications han.junyang
  2 siblings, 1 reply; 5+ messages in thread
From: han.junyang @ 2026-09-28 12:24 UTC (permalink / raw)
  To: andrew+netdev, davem, edumazet, kuba, pabeni, horms
  Cc: linux-kernel, netdev, han.junyang, ran.ming, han.chengfei, zhang.yanze

From: Junyang Han <han.junyang@zte.com.cn>

The DingHai firmware publishes a version compatibility block and a
RISC-V health buffer at fixed offsets within BAR 0.

After the PCI capabilities are mapped, poll the compatibility block
until the firmware populates it (the region reads as all ones until
then) and verify the driver/firmware version contract. Then wait for
the RISC-V management core to set its power-on flag in the health
buffer before the rest of the probe continues.

Firmware images that support neither self-healing (health version
other than 1) nor hot-plug interrupts (patch level below
ZXDH_HPIRQ_PATCH) do not report the power-on flag and skip the
readiness wait.

Signed-off-by: Junyang Han <han.junyang@zte.com.cn>
---
 drivers/net/ethernet/zte/dinghai/en_pf.c | 130 +++++++++++++++++++++++
 drivers/net/ethernet/zte/dinghai/en_pf.h |  46 ++++++++
 2 files changed, 176 insertions(+)

diff --git a/drivers/net/ethernet/zte/dinghai/en_pf.c b/drivers/net/ethernet/zte/dinghai/en_pf.c
index 86d437408820..ddff9b0438fb 100644
--- a/drivers/net/ethernet/zte/dinghai/en_pf.c
+++ b/drivers/net/ethernet/zte/dinghai/en_pf.c
@@ -6,6 +6,8 @@

 #include <linux/module.h>
 #include <linux/pci.h>
+#include <linux/io.h>
+#include <linux/iopoll.h>
 #include <net/devlink.h>
 #include <linux/dma-mapping.h>
 #include "en_pf.h"
@@ -369,6 +371,120 @@ int zxdh_pf_modern_cfg_init(struct zxdh_core_dev *zxdh_dev)
 	return ret;
 }

+/* The firmware fills the compatibility block field by field, so it is
+ * only ready to be read once every field we consume has lost its
+ * erased value.
+ */
+static bool zxdh_pf_fw_compat_ready(struct zxdh_fw_compat __iomem *compat)
+{
+	return ioread8(&compat->module_id) != 0xffU &&
+	       ioread8(&compat->major) != 0xffU &&
+	       ioread8(&compat->fw_minor) != 0xffU &&
+	       ioread8(&compat->drv_minor) != 0xffU &&
+	       ioread16(&compat->patch) != 0xffffU;
+}
+
+/* Read the firmware version block and verify the driver/firmware
+ * version contract.
+ */
+static int zxdh_pf_fw_compat_check(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_pf_dev *pf_dev = zxdh_dev->priv;
+	struct zxdh_fw_compat __iomem *compat;
+	struct zxdh_fw_compat *fw_compat;
+	bool ready;
+	int err;
+
+	fw_compat = &pf_dev->fw_compat;
+	compat = pf_dev->pci_ioremap_addr[0] + ZXDH_FW_COMPAT_OFFSET;
+
+	/* The firmware publishes the block within 10 s of boot; wait twice
+	 * that to leave margin.
+	 */
+	err = readx_poll_timeout(zxdh_pf_fw_compat_ready, compat, ready,
+				 ready, USEC_PER_SEC,
+				 ZXDH_FW_COMPAT_TIMEOUT_SEC * USEC_PER_SEC);
+	if (err)
+		dev_info(zxdh_dev->device,
+			 "compat region not populated, assuming legacy firmware\n");
+
+	/* Firmware predating the compatibility region keeps the erased
+	 * pattern, which fails the module id check below and defers the
+	 * decision to the readiness wait.
+	 */
+	fw_compat->module_id = ioread8(&compat->module_id);
+	fw_compat->major = ioread8(&compat->major);
+	fw_compat->fw_minor = ioread8(&compat->fw_minor);
+	fw_compat->drv_minor = ioread8(&compat->drv_minor);
+	fw_compat->patch = ioread16(&compat->patch);
+
+	if (fw_compat->module_id != ZXDH_MODULE_ID) {
+		dev_info(zxdh_dev->device,
+			 "unknown module id %u, skip fw compat check\n",
+			 fw_compat->module_id);
+		/* Unknown firmware is treated as predating the HPIRQ
+		 * patch, so that the readiness wait is skipped.
+		 */
+		fw_compat->patch = 0;
+		return 0;
+	}
+
+	if (fw_compat->major != ZXDH_MAJOR) {
+		dev_err(zxdh_dev->device,
+			"driver major %u incompatible with firmware major %u\n",
+			ZXDH_MAJOR, fw_compat->major);
+		return -EINVAL;
+	}
+
+	if (fw_compat->fw_minor < ZXDH_FW_MINOR) {
+		dev_err(zxdh_dev->device,
+			"firmware fw_minor %u older than required %u\n",
+			fw_compat->fw_minor, ZXDH_FW_MINOR);
+		return -EINVAL;
+	}
+
+	if (fw_compat->drv_minor > ZXDH_DRV_MINOR) {
+		dev_err(zxdh_dev->device,
+			"driver drv_minor %u older than required by firmware %u\n",
+			ZXDH_DRV_MINOR, fw_compat->drv_minor);
+		return -EINVAL;
+	}
+
+	return 0;
+}
+
+/* Wait for the RISC-V management core of the firmware to finish
+ * booting, so that later probe steps can talk to it.
+ */
+static int zxdh_pf_wait_riscv_ready(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_pf_dev *pf_dev = zxdh_dev->priv;
+	struct zxdh_health_buffer __iomem *hb;
+	u8 health_version;
+	u8 power_on;
+	int err;
+
+	hb = pf_dev->pci_ioremap_addr[0] + ZXDH_RISCV_HB_OFFSET;
+	health_version = ioread8(&hb->health_version);
+
+	/* Firmware that supports neither self-healing nor hot-plug
+	 * interrupts does not report the power-on flag.
+	 */
+	if (health_version != 1 &&
+	    pf_dev->fw_compat.patch < ZXDH_HPIRQ_PATCH)
+		return 0;
+
+	err = readx_poll_timeout(ioread8, &hb->riscv_power_on, power_on,
+				 power_on == 1, USEC_PER_SEC,
+				 ZXDH_RISCV_READY_TIMEOUT_SEC * USEC_PER_SEC);
+	if (err) {
+		dev_err(zxdh_dev->device, "timed out waiting for riscv power on\n");
+		return err;
+	}
+
+	return 0;
+}
+
 static int zxdh_pf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 {
 	struct zxdh_core_dev *zxdh_dev;
@@ -405,10 +521,24 @@ static int zxdh_pf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 		goto err_cfg_init;
 	}

+	ret = zxdh_pf_fw_compat_check(zxdh_dev);
+	if (ret) {
+		dev_err(&pdev->dev, "zxdh_pf_fw_compat_check failed: %d\n", ret);
+		goto err_modern_cfg;
+	}
+
+	ret = zxdh_pf_wait_riscv_ready(zxdh_dev);
+	if (ret) {
+		dev_err(&pdev->dev, "zxdh_pf_wait_riscv_ready failed: %d\n", ret);
+		goto err_modern_cfg;
+	}
+
 	devlink_register(devlink);

 	return 0;

+err_modern_cfg:
+	zxdh_pf_modern_cfg_uninit(zxdh_dev);
 err_cfg_init:
 	zxdh_pf_pci_close(zxdh_dev);
 err_pci_init:
diff --git a/drivers/net/ethernet/zte/dinghai/en_pf.h b/drivers/net/ethernet/zte/dinghai/en_pf.h
index 7373dee8d1a9..7733f41c57f1 100644
--- a/drivers/net/ethernet/zte/dinghai/en_pf.h
+++ b/drivers/net/ethernet/zte/dinghai/en_pf.h
@@ -29,6 +29,51 @@
 #define ZXDH_PF_ALIGN2			2
 #define ZXDH_PF_MAP_MINLEN2		2

+/* Fixed offsets of the firmware interface regions within BAR 0. */
+#define ZXDH_RISCV_HB_OFFSET	0x5300
+#define ZXDH_FW_COMPAT_OFFSET	0x5400
+
+/* Driver/firmware version contract. The firmware publishes its side of
+ * the contract in the region at ZXDH_FW_COMPAT_OFFSET.
+ */
+#define ZXDH_MODULE_ID		1
+#define ZXDH_MAJOR		1
+#define ZXDH_FW_MINOR		0
+#define ZXDH_DRV_MINOR		0
+/* Firmware patch level that introduced hot-plug IRQ support. */
+#define ZXDH_HPIRQ_PATCH	4
+
+#define ZXDH_FW_COMPAT_TIMEOUT_SEC	20
+#define ZXDH_RISCV_READY_TIMEOUT_SEC	40
+
+/* Firmware version compatibility block at ZXDH_FW_COMPAT_OFFSET.
+ * Fields are read through ioread*(), which converts from little-endian.
+ */
+struct zxdh_fw_compat {
+	u8 module_id;
+	u8 major;
+	u8 fw_minor;
+	u8 drv_minor;
+	u16 patch;
+	u16 rsv;
+} __packed;
+
+/* Health buffer at ZXDH_RISCV_HB_OFFSET, maintained by the RISC-V
+ * management core of the firmware. Fields are read through ioread*().
+ */
+struct zxdh_health_buffer {
+	u32 synd;		/* Bitmask of active syndrome flags. */
+	u32 health_counter;	/* Incremented heartbeat counter. */
+	u8 status;
+	u8 rfr;
+	u8 fw_exception;
+	u8 riscv_power_on;	/* Set to 1 once the core finished booting. */
+	u8 fw_version[32];
+	u8 pf_status[5];
+	u8 health_version;	/* 1 when the firmware supports self-healing. */
+	u8 rsv1[30];
+} __packed;
+
 struct zxdh_core_dev {
 	struct device *device;
 	struct pci_dev *pdev;
@@ -54,6 +99,7 @@ struct zxdh_pf_dev {
 	s32 modern_bars;
 	void __iomem *pci_ioremap_addr[6];
 	u32 dev_cfg_bar_off;
+	struct zxdh_fw_compat fw_compat;
 };

 void *zxdh_core_alloc_priv(struct zxdh_core_dev *zxdh_dev, size_t size);
-- 
2.27.0

^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH net-next v4 2/3] dinghai: add MSI-X interrupt pools
  2026-09-28 12:20 [PATCH net-next v4 0/3] dinghai: firmware handshake, MSI-X pools and async event queues han.junyang
  2026-09-28 12:24 ` [PATCH net-next v4 1/3] dinghai: add firmware version check and RISC-V readiness polling han.junyang
@ 2026-09-28 12:27 ` han.junyang
  2026-09-28 12:33 ` [PATCH net-next v4 3/3] dinghai: add async event queue for firmware notifications han.junyang
  2 siblings, 0 replies; 5+ messages in thread
From: han.junyang @ 2026-09-28 12:27 UTC (permalink / raw)
  To: andrew+netdev, davem, edumazet, kuba, pabeni, horms
  Cc: linux-kernel, netdev, han.junyang, ran.ming, han.chengfei, zhang.yanze

From: Junyang Han <han.junyang@zte.com.cn>

Allocate the fixed MSI-X vector layout of the device and manage the
vectors in per-purpose pools: the async event queues first, a range
reserved for RDMA in between and the vq queue pairs last. This series
wires up the async pool; the vq pool comes with the netdev series.

IRQs are reference counted so that several event queues can share one
vector. The pool hands out the least loaded matching IRQ once it passes
the pool minimum threshold, and binds newly created IRQs to the least
loaded CPU of the requested affinity mask. Interrupt delivery fans out
through an atomic notifier chain attached to each IRQ, which the async
event queue setup posted later in this series hooks into.

Signed-off-by: Junyang Han <han.junyang@zte.com.cn>
---
 drivers/net/ethernet/zte/dinghai/Makefile   |   2 +-
 drivers/net/ethernet/zte/dinghai/en_pf.c    |  84 ++++
 drivers/net/ethernet/zte/dinghai/en_pf.h    |   8 +
 drivers/net/ethernet/zte/dinghai/zxdh_irq.c | 406 ++++++++++++++++++++
 drivers/net/ethernet/zte/dinghai/zxdh_irq.h |  73 ++++
 5 files changed, 572 insertions(+), 1 deletion(-)
 create mode 100644 drivers/net/ethernet/zte/dinghai/zxdh_irq.c
 create mode 100644 drivers/net/ethernet/zte/dinghai/zxdh_irq.h

diff --git a/drivers/net/ethernet/zte/dinghai/Makefile b/drivers/net/ethernet/zte/dinghai/Makefile
index e2f4da33df59..f4a8f9480291 100644
--- a/drivers/net/ethernet/zte/dinghai/Makefile
+++ b/drivers/net/ethernet/zte/dinghai/Makefile
@@ -4,4 +4,4 @@
 #

 obj-$(CONFIG_DINGHAI_PF) += dinghai10e.o
-dinghai10e-y := en_pf.o
+dinghai10e-y := en_pf.o zxdh_irq.o
diff --git a/drivers/net/ethernet/zte/dinghai/en_pf.c b/drivers/net/ethernet/zte/dinghai/en_pf.c
index ddff9b0438fb..d5462ac6468f 100644
--- a/drivers/net/ethernet/zte/dinghai/en_pf.c
+++ b/drivers/net/ethernet/zte/dinghai/en_pf.c
@@ -8,6 +8,8 @@
 #include <linux/pci.h>
 #include <linux/io.h>
 #include <linux/iopoll.h>
+#include <linux/mm.h>
+#include <linux/slab.h>
 #include <net/devlink.h>
 #include <linux/dma-mapping.h>
 #include "en_pf.h"
@@ -27,6 +29,14 @@ static const struct pci_device_id zxdh_pf_pci_table[] = {

 MODULE_DEVICE_TABLE(pci, zxdh_pf_pci_table);

+struct zxdh_pf_irq_table {
+	struct zxdh_irq_pool *async_pool;
+};
+
+/* IRQ compaction thresholds of the async pool, in number of queues. */
+#define ZXDH_PF_ASYNC_IRQ_MIN_COMP	0
+#define ZXDH_PF_ASYNC_IRQ_MAX_COMP	7
+
 void *zxdh_core_alloc_priv(struct zxdh_core_dev *zxdh_dev, size_t size)
 {
 	void *priv = kzalloc(size, GFP_KERNEL);
@@ -485,6 +495,71 @@ static int zxdh_pf_wait_riscv_ready(struct zxdh_core_dev *zxdh_dev)
 	return 0;
 }

+static int zxdh_pf_irq_pools_init(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_pf_irq_table *pf_irq_table = zxdh_dev->irq_table.priv;
+	struct zxdh_irq_pool *pool;
+
+	pool = zxdh_irq_pool_alloc(zxdh_dev, 0, ZXDH_ASYNC_CHANNELS_NUM,
+				   "zxdh_pf_async", ZXDH_PF_ASYNC_IRQ_MIN_COMP,
+				   ZXDH_PF_ASYNC_IRQ_MAX_COMP);
+	if (IS_ERR(pool))
+		return PTR_ERR(pool);
+
+	pf_irq_table->async_pool = pool;
+
+	return 0;
+}
+
+static void zxdh_pf_irq_pools_destroy(struct zxdh_pf_irq_table *pf_irq_table)
+{
+	zxdh_irq_pool_free(pf_irq_table->async_pool);
+}
+
+int zxdh_pf_irq_table_init(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_irq_table *table = &zxdh_dev->irq_table;
+	struct zxdh_pf_irq_table *priv;
+	int total_vec;
+
+	priv = kvzalloc_obj(*priv, GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+	table->priv = priv;
+
+	total_vec = ZXDH_VQS_CHANNELS_NUM + ZXDH_ASYNC_CHANNELS_NUM +
+		    ZXDH_RDMA_CHANNELS_NUM;
+	total_vec = pci_alloc_irq_vectors(zxdh_dev->pdev, total_vec, total_vec,
+					  PCI_IRQ_MSIX);
+	if (total_vec < 0) {
+		dev_err(zxdh_dev->device, "pci_alloc_irq_vectors failed: %d\n",
+			total_vec);
+		return total_vec;
+	}
+
+	return zxdh_pf_irq_pools_init(zxdh_dev);
+}
+
+void zxdh_pf_irq_table_destroy(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_irq_table *table = &zxdh_dev->irq_table;
+
+	zxdh_pf_irq_pools_destroy(table->priv);
+	kvfree(table->priv);
+	pci_free_irq_vectors(zxdh_dev->pdev);
+}
+
+struct zxdh_irq *zxdh_pf_async_irq_request(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_irq_table *table = &zxdh_dev->irq_table;
+	struct zxdh_pf_irq_table *pf_irq_table = table->priv;
+
+	if (!pf_irq_table->async_pool)
+		return NULL;
+
+	return zxdh_get_irq_of_pool(pf_irq_table->async_pool);
+}
+
 static int zxdh_pf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 {
 	struct zxdh_core_dev *zxdh_dev;
@@ -533,10 +608,18 @@ static int zxdh_pf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 		goto err_modern_cfg;
 	}

+	ret = zxdh_pf_irq_table_init(zxdh_dev);
+	if (ret) {
+		dev_err(&pdev->dev, "zxdh_pf_irq_table_init failed: %d\n", ret);
+		goto err_irq_table;
+	}
+
 	devlink_register(devlink);

 	return 0;

+err_irq_table:
+	zxdh_pf_irq_table_destroy(zxdh_dev);
 err_modern_cfg:
 	zxdh_pf_modern_cfg_uninit(zxdh_dev);
 err_cfg_init:
@@ -554,6 +637,7 @@ static void zxdh_pf_remove(struct pci_dev *pdev)
 	struct devlink *devlink = priv_to_devlink(zxdh_dev);

 	devlink_unregister(devlink);
+	zxdh_pf_irq_table_destroy(zxdh_dev);
 	zxdh_pf_modern_cfg_uninit(zxdh_dev);
 	zxdh_pf_pci_close(zxdh_dev);
 	zxdh_core_free_priv(zxdh_dev);
diff --git a/drivers/net/ethernet/zte/dinghai/en_pf.h b/drivers/net/ethernet/zte/dinghai/en_pf.h
index 7733f41c57f1..af25b06160ef 100644
--- a/drivers/net/ethernet/zte/dinghai/en_pf.h
+++ b/drivers/net/ethernet/zte/dinghai/en_pf.h
@@ -9,6 +9,8 @@

 #include <linux/types.h>

+#include "zxdh_irq.h"
+
 #define ZXDH_PF_VENDOR_ID	0x1cf2
 #define ZXDH_PF_DEVICE_ID	0x8040
 #define ZXDH_VF_DEVICE_ID	0x8041
@@ -78,6 +80,7 @@ struct zxdh_core_dev {
 	struct device *device;
 	struct pci_dev *pdev;
 	struct devlink *devlink;
+	struct zxdh_irq_table irq_table;
 	void *priv;
 };

@@ -118,4 +121,9 @@ int zxdh_pf_device_cfg_init(struct zxdh_core_dev *zxdh_dev);
 void zxdh_pf_modern_cfg_uninit(struct zxdh_core_dev *zxdh_dev);
 int zxdh_pf_modern_cfg_init(struct zxdh_core_dev *zxdh_dev);

+/* PF IRQ table (en_pf.c) */
+int zxdh_pf_irq_table_init(struct zxdh_core_dev *zxdh_dev);
+void zxdh_pf_irq_table_destroy(struct zxdh_core_dev *zxdh_dev);
+struct zxdh_irq *zxdh_pf_async_irq_request(struct zxdh_core_dev *zxdh_dev);
+
 #endif /* __ZXDH_EN_PF_H__ */
diff --git a/drivers/net/ethernet/zte/dinghai/zxdh_irq.c b/drivers/net/ethernet/zte/dinghai/zxdh_irq.c
new file mode 100644
index 000000000000..bc5348dfa5ba
--- /dev/null
+++ b/drivers/net/ethernet/zte/dinghai/zxdh_irq.c
@@ -0,0 +1,406 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * ZTE DingHai Ethernet driver - IRQ pool management
+ * Copyright (c) 2022-2026, ZTE Corporation.
+ *
+ * Device vectors are managed in per-purpose pools. An IRQ carries a
+ * refcount so that several event queues can share one vector; on
+ * raise, the interrupt is fanned out through an atomic notifier chain
+ * attached to the IRQ.
+ */
+
+#include <linux/cpumask.h>
+#include <linux/device.h>
+#include <linux/interrupt.h>
+#include <linux/kernel.h>
+#include <linux/pci.h>
+#include <linux/slab.h>
+#include <linux/string.h>
+#include <linux/xarray.h>
+
+#include "en_pf.h"
+#include "zxdh_irq.h"
+
+static void zxdh_cpu_put(struct zxdh_irq_pool *pool, int cpu)
+{
+	pool->irqs_per_cpu[cpu]--;
+}
+
+/* Release callback of the IRQ kref, called with pool->lock held. */
+static void zxdh_irq_release(struct kref *kref)
+{
+	struct zxdh_irq *irq = container_of(kref, struct zxdh_irq, refcount);
+	struct zxdh_irq_pool *pool = irq->pool;
+
+	lockdep_assert_held(&pool->lock);
+	xa_erase(&pool->irqs, irq->index);
+	zxdh_cpu_put(pool, cpumask_first(irq->mask));
+	/* free_irq() requires the affinity hint to be cleared before it is
+	 * called; the asymmetry with the set path in zxdh_irq_alloc() is
+	 * intentional.
+	 */
+	irq_update_affinity_hint(irq->irqn, NULL);
+	free_cpumask_var(irq->mask);
+	free_irq(irq->irqn, irq);
+	kfree(irq);
+}
+
+/* Drop one reference; returns 1 if the IRQ was released. */
+static int zxdh_irq_put(struct zxdh_irq *irq)
+{
+	struct zxdh_irq_pool *pool = irq->pool;
+	int released;
+
+	mutex_lock(&pool->lock);
+	released = kref_put(&irq->refcount, zxdh_irq_release);
+	mutex_unlock(&pool->lock);
+
+	return released;
+}
+
+static int zxdh_irq_get(struct zxdh_irq *irq)
+{
+	return kref_get_unless_zero(&irq->refcount);
+}
+
+static irqreturn_t zxdh_irq_int_handler(int irq, void *data)
+{
+	struct zxdh_irq *zxdh_irq = data;
+
+	atomic_notifier_call_chain(&zxdh_irq->nh, 0, NULL);
+
+	return IRQ_HANDLED;
+}
+
+static struct zxdh_irq *zxdh_irq_alloc(struct zxdh_irq_pool *pool, int vecidx,
+				       const struct cpumask *affinity)
+{
+	struct zxdh_core_dev *zxdh_dev = pool->dev;
+	struct zxdh_irq *irq;
+	int err;
+	int cpu;
+
+	irq = kzalloc_obj(*irq, GFP_KERNEL);
+	if (!irq)
+		return ERR_PTR(-ENOMEM);
+
+	irq->pool = pool;
+	irq->irqn = pci_irq_vector(zxdh_dev->pdev, vecidx);
+	if (irq->irqn < 0) {
+		err = irq->irqn;
+		goto err_irqn;
+	}
+
+	ATOMIC_INIT_NOTIFIER_HEAD(&irq->nh);
+	snprintf(irq->name, ZXDH_MAX_IRQ_NAME, "%s_%d@pci:%s", pool->name,
+		 vecidx, pci_name(zxdh_dev->pdev));
+
+	err = request_irq(irq->irqn, zxdh_irq_int_handler, 0, irq->name, irq);
+	if (err) {
+		dev_err(zxdh_dev->device, "request_irq failed: %d\n", err);
+		goto err_irqn;
+	}
+
+	if (!zalloc_cpumask_var(&irq->mask, GFP_KERNEL)) {
+		dev_err(zxdh_dev->device, "zalloc_cpumask_var failed\n");
+		err = -ENOMEM;
+		goto err_cpumask;
+	}
+
+	if (affinity) {
+		cpumask_copy(irq->mask, affinity);
+	} else {
+		/* No preference requested; spread over all online CPUs. */
+		for_each_online_cpu(cpu)
+			cpumask_set_cpu(cpu, irq->mask);
+	}
+	irq_set_affinity_and_hint(irq->irqn, irq->mask);
+
+	kref_init(&irq->refcount);
+	irq->index = vecidx;
+	err = xa_err(xa_store(&pool->irqs, irq->index, irq, GFP_KERNEL));
+	if (err) {
+		dev_err(zxdh_dev->device, "xa_store failed for irq %u: %d\n",
+			irq->index, err);
+		goto err_xa;
+	}
+
+	return irq;
+
+err_xa:
+	irq_update_affinity_hint(irq->irqn, NULL);
+	free_cpumask_var(irq->mask);
+err_cpumask:
+	free_irq(irq->irqn, irq);
+err_irqn:
+	kfree(irq);
+	return ERR_PTR(err);
+}
+
+int zxdh_irq_attach_nb(struct zxdh_irq *irq, struct notifier_block *nb)
+{
+	int err;
+
+	if (!zxdh_irq_get(irq))
+		return -ENOENT;
+
+	err = atomic_notifier_chain_register(&irq->nh, nb);
+	if (err)
+		zxdh_irq_put(irq);
+
+	return err;
+}
+
+int zxdh_irq_detach_nb(struct zxdh_irq *irq, struct notifier_block *nb)
+{
+	int err;
+
+	err = atomic_notifier_chain_unregister(&irq->nh, nb);
+	zxdh_irq_put(irq);
+
+	return err;
+}
+
+/* Release one or more IRQs back to their pool. */
+void zxdh_irqs_release_vectors(struct zxdh_irq **irqs, int nirqs)
+{
+	int i;
+
+	for (i = 0; i < nirqs; i++) {
+		synchronize_irq(irqs[i]->irqn);
+		zxdh_irq_put(irqs[i]);
+	}
+}
+
+static void zxdh_cpu_get(struct zxdh_irq_pool *pool, int cpu)
+{
+	pool->irqs_per_cpu[cpu]++;
+}
+
+/* Find the least loaded CPU, i.e. the CPU with the fewest IRQs bound
+ * to it, within the requested mask.
+ */
+static int zxdh_cpu_get_least_loaded(struct zxdh_irq_pool *pool,
+				     const struct cpumask *req_mask)
+{
+	int best_cpu = -1;
+	int cpu;
+
+	for_each_cpu_and(cpu, req_mask, cpu_online_mask) {
+		/* This CPU has no IRQs yet, no need to look further. */
+		if (!pool->irqs_per_cpu[cpu]) {
+			best_cpu = cpu;
+			break;
+		}
+		if (best_cpu < 0)
+			best_cpu = cpu;
+		if (pool->irqs_per_cpu[cpu] < pool->irqs_per_cpu[best_cpu])
+			best_cpu = cpu;
+	}
+
+	if (best_cpu < 0) {
+		dev_err(pool->dev->device,
+			"no online CPU in affinity mask (%*pbl)\n",
+			cpumask_pr_args(req_mask));
+		best_cpu = cpumask_first(cpu_online_mask);
+	}
+	pool->irqs_per_cpu[best_cpu]++;
+
+	return best_cpu;
+}
+
+/* Create a new IRQ from the pool and bind it to the requested mask. */
+static struct zxdh_irq *zxdh_irq_pool_request_irq(struct zxdh_irq_pool *pool,
+						  const struct cpumask *req_mask)
+{
+	const struct cpumask *affinity = req_mask;
+	cpumask_var_t auto_mask;
+	struct zxdh_irq *irq;
+	u32 irq_index;
+	int err;
+
+	if (!zalloc_cpumask_var(&auto_mask, GFP_KERNEL))
+		return ERR_PTR(-ENOMEM);
+
+	err = xa_alloc(&pool->irqs, &irq_index, NULL, pool->xa_num_irqs,
+		       GFP_KERNEL);
+	if (err) {
+		if (err == -EBUSY)
+			err = -EUSERS;
+		goto err_xa;
+	}
+
+	if (cpumask_weight(req_mask) > 1) {
+		/* Bind to the least loaded CPU of the request. */
+		cpumask_set_cpu(zxdh_cpu_get_least_loaded(pool, req_mask),
+				auto_mask);
+		affinity = auto_mask;
+	} else {
+		zxdh_cpu_get(pool, cpumask_first(req_mask));
+	}
+
+	irq = zxdh_irq_alloc(pool, irq_index, affinity);
+	if (IS_ERR(irq))
+		goto err_alloc;
+
+	free_cpumask_var(auto_mask);
+
+	return irq;
+
+err_alloc:
+	/* Undo the per-CPU accounting done above. */
+	zxdh_cpu_put(pool, cpumask_first(affinity));
+	xa_erase(&pool->irqs, irq_index);
+err_xa:
+	free_cpumask_var(auto_mask);
+	return err ? ERR_PTR(err) : irq;
+}
+
+/* Look for the IRQ with the smallest refcount whose affinity mask is a
+ * subset of the requested mask.
+ */
+static struct zxdh_irq *zxdh_irq_pool_find_least_loaded(struct zxdh_irq_pool *pool,
+							const struct cpumask *req_mask)
+{
+	struct zxdh_irq *irq = NULL;
+	struct zxdh_irq *iter;
+	unsigned long index;
+	int iter_refcount;
+
+	lockdep_assert_held(&pool->lock);
+	xa_for_each_range(&pool->irqs, index, iter,
+			  pool->xa_num_irqs.min, pool->xa_num_irqs.max) {
+		iter_refcount = refcount_read(&iter->refcount.refcount);
+
+		/* Skip IRQs whose mask is not a subset of the request. */
+		if (!cpumask_subset(iter->mask, req_mask))
+			continue;
+		/* IRQ below the minimum threshold found, take it. */
+		if (iter_refcount < pool->min_threshold)
+			return iter;
+		if (!irq ||
+		    iter_refcount < refcount_read(&irq->refcount.refcount))
+			irq = iter;
+	}
+
+	return irq;
+}
+
+/* Request an IRQ for the given mask: share the least loaded existing
+ * matching IRQ once it passes the pool minimum threshold, otherwise
+ * create a new one.
+ */
+static struct zxdh_irq *zxdh_irq_affinity_request(struct zxdh_irq_pool *pool,
+						  const struct cpumask *req_mask)
+{
+	struct zxdh_irq *least_loaded_irq;
+	struct zxdh_irq *new_irq;
+
+	mutex_lock(&pool->lock);
+
+	least_loaded_irq = zxdh_irq_pool_find_least_loaded(pool, req_mask);
+	if (least_loaded_irq &&
+	    refcount_read(&least_loaded_irq->refcount.refcount) < pool->min_threshold)
+		goto out;
+
+	/* No IRQ below the minimum threshold, try to create a new one. */
+	new_irq = zxdh_irq_pool_request_irq(pool, req_mask);
+	if (IS_ERR(new_irq)) {
+		if (!least_loaded_irq) {
+			mutex_unlock(&pool->lock);
+			return new_irq;
+		}
+		/* Could not create a new IRQ for the requested affinity,
+		 * share the existing one.
+		 */
+		dev_warn(pool->dev->device,
+			 "irq pool %s exhausted, sharing irqs\n",
+			 pool->name);
+		goto out;
+	}
+
+	least_loaded_irq = new_irq;
+	goto unlock;
+
+out:
+	/* Holding pool->lock keeps the IRQ alive, so this cannot fail. */
+	zxdh_irq_get(least_loaded_irq);
+	if (refcount_read(&least_loaded_irq->refcount.refcount) > pool->max_threshold)
+		dev_warn(pool->dev->device,
+			 "irq %u overloaded, pool %s, %u refs on this irq\n",
+			 pci_irq_vector(pool->dev->pdev,
+					least_loaded_irq->index),
+			 pool->name,
+			 refcount_read(&least_loaded_irq->refcount.refcount) /
+			 ZXDH_EQ_REFS_PER_IRQ);
+unlock:
+	mutex_unlock(&pool->lock);
+
+	return least_loaded_irq;
+}
+
+/* Request an IRQ from the pool, spread over the online CPUs. */
+struct zxdh_irq *zxdh_get_irq_of_pool(struct zxdh_irq_pool *pool)
+{
+	cpumask_var_t req_mask;
+	struct zxdh_irq *irq;
+
+	if (!zalloc_cpumask_var(&req_mask, GFP_KERNEL))
+		return ERR_PTR(-ENOMEM);
+	cpumask_copy(req_mask, cpu_online_mask);
+
+	irq = zxdh_irq_affinity_request(pool, req_mask);
+
+	free_cpumask_var(req_mask);
+
+	return irq;
+}
+
+struct zxdh_irq_pool *zxdh_irq_pool_alloc(struct zxdh_core_dev *zxdh_dev,
+					  int start, int size, const char *name,
+					  u32 min_threshold, u32 max_threshold)
+{
+	struct zxdh_irq_pool *pool;
+	u16 *irqs_per_cpu;
+
+	irqs_per_cpu = kcalloc(nr_cpu_ids, sizeof(*irqs_per_cpu), GFP_KERNEL);
+	if (!irqs_per_cpu)
+		return ERR_PTR(-ENOMEM);
+
+	pool = kvzalloc_obj(*pool, GFP_KERNEL);
+	if (!pool) {
+		kfree(irqs_per_cpu);
+		return ERR_PTR(-ENOMEM);
+	}
+
+	pool->dev = zxdh_dev;
+	pool->irqs_per_cpu = irqs_per_cpu;
+	mutex_init(&pool->lock);
+	xa_init_flags(&pool->irqs, XA_FLAGS_ALLOC);
+	pool->xa_num_irqs.min = start;
+	pool->xa_num_irqs.max = start + size - 1;
+	strscpy(pool->name, name, sizeof(pool->name));
+
+	/* Thresholds are expressed in queue references, two per IRQ. */
+	pool->min_threshold = min_threshold * ZXDH_EQ_REFS_PER_IRQ;
+	pool->max_threshold = max_threshold * ZXDH_EQ_REFS_PER_IRQ;
+
+	return pool;
+}
+
+void zxdh_irq_pool_free(struct zxdh_irq_pool *pool)
+{
+	u32 cpu;
+
+	WARN_ON_ONCE(!xa_empty(&pool->irqs));
+	xa_destroy(&pool->irqs);
+	mutex_destroy(&pool->lock);
+
+	if (pool->irqs_per_cpu) {
+		for_each_online_cpu(cpu)
+			WARN_ON(pool->irqs_per_cpu[cpu]);
+		kfree(pool->irqs_per_cpu);
+	}
+
+	kvfree(pool);
+}
diff --git a/drivers/net/ethernet/zte/dinghai/zxdh_irq.h b/drivers/net/ethernet/zte/dinghai/zxdh_irq.h
new file mode 100644
index 000000000000..a969e04f3ce4
--- /dev/null
+++ b/drivers/net/ethernet/zte/dinghai/zxdh_irq.h
@@ -0,0 +1,73 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/*
+ * ZTE DingHai Ethernet driver - IRQ pool management
+ * Copyright (c) 2022-2026, ZTE Corporation.
+ */
+
+#ifndef __ZXDH_IRQ_H__
+#define __ZXDH_IRQ_H__
+
+#include <linux/cpumask.h>
+#include <linux/kref.h>
+#include <linux/mutex.h>
+#include <linux/notifier.h>
+#include <linux/types.h>
+#include <linux/xarray.h>
+
+struct zxdh_core_dev;
+
+/* Number of MSI-X vectors per purpose. The layout of the vector space
+ * is a hardware interface:
+ *   [0, 8)    async event queues
+ *   [8, 14)   reserved for RDMA (unused for now)
+ *   [14, 78)  vq rx/tx queue pairs
+ */
+#define ZXDH_ASYNC_CHANNELS_NUM		8
+#define ZXDH_RDMA_CHANNELS_NUM		6
+#define ZXDH_VQS_CHANNELS_NUM		64
+
+#define ZXDH_MAX_IRQ_NAME	100
+#define ZXDH_MAX_POOL_NAME	16
+#define ZXDH_EQ_REFS_PER_IRQ	2
+
+/* Reference counted interrupt. Event queues share an IRQ through the
+ * notifier chain, which is fired from hard IRQ context.
+ */
+struct zxdh_irq {
+	struct atomic_notifier_head nh;
+	cpumask_var_t mask;		/* interrupt affinity */
+	char name[ZXDH_MAX_IRQ_NAME];
+	struct zxdh_irq_pool *pool;
+	u32 index;			/* vector index */
+	int irqn;			/* Linux IRQ number */
+	struct kref refcount;
+};
+
+/* Pool of vectors dedicated to one purpose, identified by the
+ * [xa_num_irqs.min, xa_num_irqs.max] vector range.
+ */
+struct zxdh_irq_pool {
+	char name[ZXDH_MAX_POOL_NAME];
+	struct xa_limit xa_num_irqs;
+	struct mutex lock;		/* serializes IRQ creation */
+	struct xarray irqs;
+	u32 min_threshold;		/* in queue references */
+	u32 max_threshold;		/* in queue references */
+	u16 *irqs_per_cpu;		/* number of IRQs bound per CPU */
+	struct zxdh_core_dev *dev;
+};
+
+struct zxdh_irq_table {
+	void *priv;
+};
+
+struct zxdh_irq *zxdh_get_irq_of_pool(struct zxdh_irq_pool *pool);
+int zxdh_irq_attach_nb(struct zxdh_irq *irq, struct notifier_block *nb);
+int zxdh_irq_detach_nb(struct zxdh_irq *irq, struct notifier_block *nb);
+void zxdh_irqs_release_vectors(struct zxdh_irq **irqs, int nirqs);
+struct zxdh_irq_pool *zxdh_irq_pool_alloc(struct zxdh_core_dev *zxdh_dev,
+					  int start, int size, const char *name,
+					  u32 min_threshold, u32 max_threshold);
+void zxdh_irq_pool_free(struct zxdh_irq_pool *pool);
+
+#endif /* __ZXDH_IRQ_H__ */
-- 
2.27.0

^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH net-next v4 3/3] dinghai: add async event queue for firmware  notifications
  2026-09-28 12:20 [PATCH net-next v4 0/3] dinghai: firmware handshake, MSI-X pools and async event queues han.junyang
  2026-09-28 12:24 ` [PATCH net-next v4 1/3] dinghai: add firmware version check and RISC-V readiness polling han.junyang
  2026-09-28 12:27 ` [PATCH net-next v4 2/3] dinghai: add MSI-X interrupt pools han.junyang
@ 2026-09-28 12:33 ` han.junyang
  2 siblings, 0 replies; 5+ messages in thread
From: han.junyang @ 2026-09-28 12:33 UTC (permalink / raw)
  To: andrew+netdev, davem, edumazet, kuba, pabeni, horms
  Cc: linux-kernel, netdev, han.junyang, ran.ming, han.chengfei, zhang.yanze

From: Junyang Han <han.junyang@zte.com.cn>

Add the event queue table and the first async event queue, fed by the
RISC-V management core of the firmware. The queue owns one MSI-X
vector taken from the async pool and, on interrupt, reads the event_id
from the PF receive sub-channel of the BAR message area, maps it to
an event type and fans it out through the per-event-type notifier
chains of the table. The PF probe creates and tears down the queue
through the existing IRQ pool interfaces.

The BAR message channel handshake that tells the firmware which
vectors to signal comes with the BAR message subsystem in a later
series; the PF side wired here is live as soon as that channel is
enabled. Subscribers for the RISCV_READY event register on the
notifier chains in later series as well.

Signed-off-by: Junyang Han <han.junyang@zte.com.cn>
---
 drivers/net/ethernet/zte/dinghai/Makefile  |   2 +-
 drivers/net/ethernet/zte/dinghai/en_pf.c   |  23 ++--
 drivers/net/ethernet/zte/dinghai/en_pf.h   |   2 +
 drivers/net/ethernet/zte/dinghai/zxdh_eq.c | 137 +++++++++++++++++++++
 drivers/net/ethernet/zte/dinghai/zxdh_eq.h |  60 +++++++++
 5 files changed, 216 insertions(+), 8 deletions(-)
 create mode 100644 drivers/net/ethernet/zte/dinghai/zxdh_eq.c
 create mode 100644 drivers/net/ethernet/zte/dinghai/zxdh_eq.h

diff --git a/drivers/net/ethernet/zte/dinghai/Makefile b/drivers/net/ethernet/zte/dinghai/Makefile
index f4a8f9480291..2dc6d68cc0e1 100644
--- a/drivers/net/ethernet/zte/dinghai/Makefile
+++ b/drivers/net/ethernet/zte/dinghai/Makefile
@@ -4,4 +4,4 @@
 #

 obj-$(CONFIG_DINGHAI_PF) += dinghai10e.o
-dinghai10e-y := en_pf.o zxdh_irq.o
+dinghai10e-y := en_pf.o zxdh_irq.o zxdh_eq.o
diff --git a/drivers/net/ethernet/zte/dinghai/en_pf.c b/drivers/net/ethernet/zte/dinghai/en_pf.c
index d5462ac6468f..f7400d66bc69 100644
--- a/drivers/net/ethernet/zte/dinghai/en_pf.c
+++ b/drivers/net/ethernet/zte/dinghai/en_pf.c
@@ -513,7 +513,8 @@ static int zxdh_pf_irq_pools_init(struct zxdh_core_dev *zxdh_dev)

 static void zxdh_pf_irq_pools_destroy(struct zxdh_pf_irq_table *pf_irq_table)
 {
-	zxdh_irq_pool_free(pf_irq_table->async_pool);
+	if (pf_irq_table->async_pool)
+		zxdh_irq_pool_free(pf_irq_table->async_pool);
 }

 int zxdh_pf_irq_table_init(struct zxdh_core_dev *zxdh_dev)
@@ -544,18 +545,17 @@ void zxdh_pf_irq_table_destroy(struct zxdh_core_dev *zxdh_dev)
 {
 	struct zxdh_irq_table *table = &zxdh_dev->irq_table;

-	zxdh_pf_irq_pools_destroy(table->priv);
+	if (table->priv)
+		zxdh_pf_irq_pools_destroy(table->priv);
 	kvfree(table->priv);
+	table->priv = NULL;
 	pci_free_irq_vectors(zxdh_dev->pdev);
 }

+/* Fetch a vector from the async pool for one async event queue. */
 struct zxdh_irq *zxdh_pf_async_irq_request(struct zxdh_core_dev *zxdh_dev)
 {
-	struct zxdh_irq_table *table = &zxdh_dev->irq_table;
-	struct zxdh_pf_irq_table *pf_irq_table = table->priv;
-
-	if (!pf_irq_table->async_pool)
-		return NULL;
+	struct zxdh_pf_irq_table *pf_irq_table = zxdh_dev->irq_table.priv;

 	return zxdh_get_irq_of_pool(pf_irq_table->async_pool);
 }
@@ -614,10 +614,18 @@ static int zxdh_pf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 		goto err_irq_table;
 	}

+	ret = zxdh_pf_eq_table_init(zxdh_dev);
+	if (ret) {
+		dev_err(&pdev->dev, "zxdh_pf_eq_table_init failed: %d\n", ret);
+		goto err_eq_table;
+	}
+
 	devlink_register(devlink);

 	return 0;

+err_eq_table:
+	zxdh_pf_eq_table_destroy(zxdh_dev);
 err_irq_table:
 	zxdh_pf_irq_table_destroy(zxdh_dev);
 err_modern_cfg:
@@ -637,6 +645,7 @@ static void zxdh_pf_remove(struct pci_dev *pdev)
 	struct devlink *devlink = priv_to_devlink(zxdh_dev);

 	devlink_unregister(devlink);
+	zxdh_pf_eq_table_destroy(zxdh_dev);
 	zxdh_pf_irq_table_destroy(zxdh_dev);
 	zxdh_pf_modern_cfg_uninit(zxdh_dev);
 	zxdh_pf_pci_close(zxdh_dev);
diff --git a/drivers/net/ethernet/zte/dinghai/en_pf.h b/drivers/net/ethernet/zte/dinghai/en_pf.h
index af25b06160ef..8220db579111 100644
--- a/drivers/net/ethernet/zte/dinghai/en_pf.h
+++ b/drivers/net/ethernet/zte/dinghai/en_pf.h
@@ -9,6 +9,7 @@

 #include <linux/types.h>

+#include "zxdh_eq.h"
 #include "zxdh_irq.h"

 #define ZXDH_PF_VENDOR_ID	0x1cf2
@@ -81,6 +82,7 @@ struct zxdh_core_dev {
 	struct pci_dev *pdev;
 	struct devlink *devlink;
 	struct zxdh_irq_table irq_table;
+	struct zxdh_eq_table eq_table;
 	void *priv;
 };

diff --git a/drivers/net/ethernet/zte/dinghai/zxdh_eq.c b/drivers/net/ethernet/zte/dinghai/zxdh_eq.c
new file mode 100644
index 000000000000..3e70f4f3063f
--- /dev/null
+++ b/drivers/net/ethernet/zte/dinghai/zxdh_eq.c
@@ -0,0 +1,137 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * ZTE DingHai Ethernet driver - event queue plumbing
+ * Copyright (c) 2022-2026, ZTE Corporation.
+ *
+ * Async event queues are the delivery path for firmware notifications:
+ * a queue owns one MSI-X vector of the async pool and fans interrupts
+ * out to the per-event-type notifier chains of the event queue table.
+ * The BAR message channel handshake that tells the firmware which
+ * vectors to signal comes with the BAR message subsystem in a later
+ * series; the PF side wired here is live as soon as that channel is
+ * enabled.
+ */
+
+#include <linux/io.h>
+#include <linux/kernel.h>
+#include <linux/mutex.h>
+#include <linux/notifier.h>
+#include <linux/slab.h>
+
+#include "en_pf.h"
+#include "zxdh_eq.h"
+#include "zxdh_irq.h"
+
+/* PF side of the event queue table: one async EQ fed by the RISC-V
+ * management core.
+ */
+struct zxdh_pf_eq_table {
+	struct zxdh_eq_async riscv_eq;
+};
+
+/* Fetch the event_id the RISC-V core posted in the PF receive
+ * sub-channel of the BAR message area. The field lives at byte offset
+ * 2 of the 12-byte message header, so a single 32-bit read of the
+ * sub-channel start covers it.
+ */
+static u16 zxdh_eq_event_id_get(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_pf_dev *pf_dev = zxdh_dev->priv;
+	void __iomem *subchan;
+
+	subchan = pf_dev->pci_ioremap_addr[0] + ZXDH_BAR_MSG_SUBCHAN_RECV;
+
+	return ioread32(subchan) >> 16;
+}
+
+static enum zxdh_eq_event_type zxdh_eq_event_type_get(u16 event_id)
+{
+	if (event_id == ZXDH_MSG_MODULE_RISC_READY)
+		return ZXDH_EQ_EVENT_TYPE_RISCV_READY;
+
+	/* Only the RISC-V ready notification is dispatched in this
+	 * series; the remaining firmware events arrive with the BAR
+	 * message subsystem.
+	 */
+	return ZXDH_EQ_EVENT_TYPE_MAX;
+}
+
+static int zxdh_eq_async_riscv_int(struct notifier_block *nb,
+				   unsigned long action, void *data)
+{
+	struct zxdh_eq_async *eq = container_of(nb, struct zxdh_eq_async, irq_nb);
+	struct zxdh_core_dev *zxdh_dev = eq->priv;
+	enum zxdh_eq_event_type event_type;
+
+	event_type = zxdh_eq_event_type_get(zxdh_eq_event_id_get(zxdh_dev));
+	if (event_type == ZXDH_EQ_EVENT_TYPE_MAX)
+		return NOTIFY_DONE;
+
+	atomic_notifier_call_chain(&zxdh_dev->eq_table.nh[event_type],
+				   event_type, NULL);
+
+	return NOTIFY_OK;
+}
+
+int zxdh_pf_eq_table_init(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_eq_table *table = &zxdh_dev->eq_table;
+	struct zxdh_pf_eq_table *priv;
+	struct zxdh_eq_async *eq;
+	int err;
+	int i;
+
+	priv = kvzalloc_obj(*priv, GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+	table->priv = priv;
+
+	mutex_init(&table->lock);
+	for (i = 0; i < ZXDH_EQ_EVENT_TYPE_MAX; i++)
+		ATOMIC_INIT_NOTIFIER_HEAD(&table->nh[i]);
+
+	eq = &priv->riscv_eq;
+	eq->priv = zxdh_dev;
+	eq->irq = zxdh_pf_async_irq_request(zxdh_dev);
+	if (IS_ERR(eq->irq)) {
+		err = PTR_ERR(eq->irq);
+		goto err_irq;
+	}
+
+	eq->irq_nb.notifier_call = zxdh_eq_async_riscv_int;
+	err = zxdh_irq_attach_nb(eq->irq, &eq->irq_nb);
+	if (err)
+		goto err_attach;
+
+	return 0;
+
+err_attach:
+	zxdh_irqs_release_vectors(&eq->irq, 1);
+err_irq:
+	eq->irq = NULL;
+	return err;
+}
+
+void zxdh_pf_eq_table_destroy(struct zxdh_core_dev *zxdh_dev)
+{
+	struct zxdh_pf_eq_table *pf_eq_table = zxdh_dev->eq_table.priv;
+	struct zxdh_eq_table *table = &zxdh_dev->eq_table;
+	struct zxdh_eq_async *eq;
+
+	if (!pf_eq_table)
+		return;
+
+	eq = &pf_eq_table->riscv_eq;
+
+	mutex_lock(&table->lock);
+	if (eq->irq) {
+		zxdh_irq_detach_nb(eq->irq, &eq->irq_nb);
+		zxdh_irqs_release_vectors(&eq->irq, 1);
+		eq->irq = NULL;
+	}
+	mutex_unlock(&table->lock);
+
+	mutex_destroy(&table->lock);
+	kvfree(table->priv);
+	table->priv = NULL;
+}
diff --git a/drivers/net/ethernet/zte/dinghai/zxdh_eq.h b/drivers/net/ethernet/zte/dinghai/zxdh_eq.h
new file mode 100644
index 000000000000..ac9df4763059
--- /dev/null
+++ b/drivers/net/ethernet/zte/dinghai/zxdh_eq.h
@@ -0,0 +1,60 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/*
+ * ZTE DingHai Ethernet driver - event queue abstraction
+ * Copyright (c) 2022-2026, ZTE Corporation.
+ */
+
+#ifndef __ZXDH_EQ_H__
+#define __ZXDH_EQ_H__
+
+#include <linux/mutex.h>
+#include <linux/notifier.h>
+
+struct zxdh_core_dev;
+struct zxdh_irq;
+
+/* BAR message area: shared-memory mailbox between the driver and the
+ * firmware, split into sub-channels of 2048 bytes. Notifications from
+ * the RISC-V management core are posted to the PF receive sub-channel,
+ * one interval above the area base.
+ */
+#define ZXDH_BAR_MSG_OFFSET		0x2000
+#define ZXDH_BAR_MSG_SUBCHAN_INTERVAL	2048
+#define ZXDH_BAR_MSG_SUBCHAN_RECV	(ZXDH_BAR_MSG_OFFSET +	\
+					 ZXDH_BAR_MSG_SUBCHAN_INTERVAL)
+
+/* event_id of the "RISC-V core ready" firmware notification, carried
+ * at byte offset 2 of the BAR message header.
+ */
+#define ZXDH_MSG_MODULE_RISC_READY	13
+
+/* Event types the upper layers can subscribe to over the per-type
+ * notifier chains of the event queue table. The full catalogue lives
+ * in the firmware; only the ones dispatched in this series are named.
+ */
+enum zxdh_eq_event_type {
+	ZXDH_EQ_EVENT_TYPE_RISCV_READY = 0,
+	ZXDH_EQ_EVENT_TYPE_MAX,
+};
+
+/* Asynchronous event queue delivering firmware notifications to the
+ * driver without a request/reply handshake.
+ */
+struct zxdh_eq_async {
+	struct zxdh_irq *irq;
+	struct notifier_block irq_nb;
+	void *priv;
+};
+
+struct zxdh_eq_table {
+	/* Notifier chains fanned out by event type. */
+	struct atomic_notifier_head nh[ZXDH_EQ_EVENT_TYPE_MAX];
+	/* Serializes async EQ setup/teardown. */
+	struct mutex lock;
+	void *priv;
+};
+
+int zxdh_pf_eq_table_init(struct zxdh_core_dev *zxdh_dev);
+void zxdh_pf_eq_table_destroy(struct zxdh_core_dev *zxdh_dev);
+
+#endif /* __ZXDH_EQ_H__ */
-- 
2.27.0

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net-next v4 1/3] dinghai: add firmware version check and RISC-V readiness polling
  2026-09-28 12:24 ` [PATCH net-next v4 1/3] dinghai: add firmware version check and RISC-V readiness polling han.junyang
@ 2026-10-01 13:06   ` Simon Horman
  0 siblings, 0 replies; 5+ messages in thread
From: Simon Horman @ 2026-10-01 13:06 UTC (permalink / raw)
  To: han.junyang
  Cc: Simon Horman, andrew+netdev, davem, edumazet, kuba, pabeni,
	linux-kernel, netdev, ran.ming, han.chengfei, zhang.yanze

This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
Full review at: https://sashiko.dev/#/patchset/202609282024273579VUWteKdnnbvWlyjERLQ-@zte.com.cn?part=1
---
dinghai: add firmware version check and RISC-V readiness polling

The commit adds firmware version checks and a polling mechanism for RISC-V
readiness on DingHai devices. It maps a compatibility block and a health
buffer in BAR 0 to verify driver and firmware compatibility and wait for
power-on status.

> diff --git a/drivers/net/ethernet/zte/dinghai/en_pf.c b/drivers/net/ethernet/zte/dinghai/en_pf.c
> --- a/drivers/net/ethernet/zte/dinghai/en_pf.c
> +++ b/drivers/net/ethernet/zte/dinghai/en_pf.c
[ ... ]
> @@ -369,6 +371,120 @@ int zxdh_pf_modern_cfg_init(struct zxdh_core_dev *zxdh_dev)
[ ... ]
> +static int zxdh_pf_fw_compat_check(struct zxdh_core_dev *zxdh_dev)
> +{
> +	struct zxdh_pf_dev *pf_dev = zxdh_dev->priv;
> +	struct zxdh_fw_compat __iomem *compat;
> +	struct zxdh_fw_compat *fw_compat;
> +	bool ready;
> +	int err;
> +
> +	fw_compat = &pf_dev->fw_compat;
> +	compat = pf_dev->pci_ioremap_addr[0] + ZXDH_FW_COMPAT_OFFSET;

[Severity: High]
Does this code verify the BAR length before calculating offsets into BAR 0?

The driver computes fixed offsets into BAR 0 (ZXDH_FW_COMPAT_OFFSET at 0x5400,
and later ZXDH_RISCV_HB_OFFSET at 0x5300) without validating that the reported
BAR length is large enough.

If a malformed PCI device, emulated device, or corrupt PCI configuration space
presents a BAR smaller than 0x5400 plus the size of the structure, can this lead
to an out-of-bounds MMIO read and page fault during probe?

> +
> +	/* The firmware publishes the block within 10 s of boot; wait twice
> +	 * that to leave margin.
> +	 */
> +	err = readx_poll_timeout(zxdh_pf_fw_compat_ready, compat, ready,
> +				 ready, USEC_PER_SEC,
> +				 ZXDH_FW_COMPAT_TIMEOUT_SEC * USEC_PER_SEC);

[Severity: High]
Is there a performance regression here where this loop forces an unconditional
20-second delay on legacy firmware?

If the driver probes legacy hardware that predates the compatibility region,
the region permanently contains the erased pattern (0xFF). Since legacy
firmware never writes to this region, zxdh_pf_fw_compat_ready() will not see
the fields lose their 0xFF values, meaning the loop waits for the entire
timeout before proceeding.

Could this delay severely degrade system boot or module load times?

[Severity: Low]
Are we passing 1,000,000 microseconds to usleep_range() in this macro?

The readx_poll_timeout macro expands to use usleep_range() for its sleep
intervals. Passing USEC_PER_SEC exceeds the recommended 20ms maximum for
usleep_range(), exhausting hrtimers unnecessarily.

Would it be better to use a smaller polling interval (e.g., 20ms) or a macro
that utilizes msleep() to avoid this timer regression?

> +	if (err)
> +		dev_info(zxdh_dev->device,
> +			 "compat region not populated, assuming legacy firmware\n");
> +
> +	/* Firmware predating the compatibility region keeps the erased
> +	 * pattern, which fails the module id check below and defers the
> +	 * decision to the readiness wait.
> +	 */
> +	fw_compat->module_id = ioread8(&compat->module_id);
> +	fw_compat->major = ioread8(&compat->major);
> +	fw_compat->fw_minor = ioread8(&compat->fw_minor);
> +	fw_compat->drv_minor = ioread8(&compat->drv_minor);
> +	fw_compat->patch = ioread16(&compat->patch);
> +
> +	if (fw_compat->module_id != ZXDH_MODULE_ID) {
> +		dev_info(zxdh_dev->device,
> +			 "unknown module id %u, skip fw compat check\n",
> +			 fw_compat->module_id);
> +		/* Unknown firmware is treated as predating the HPIRQ
> +		 * patch, so that the readiness wait is skipped.
> +		 */
> +		fw_compat->patch = 0;
> +		return 0;
> +	}

[Severity: Medium]
Could this logic incorrectly return success for a dead or physically removed
PCI device?

If a PCI device experiences a bus error or is disconnected, reading its MMIO
registers typically returns an all-ones value (0xFF). Here, an all-ones read
causes module_id to equal 255.

Since the driver treats this as legacy firmware, it skips the compatibility
check. Later, in zxdh_pf_wait_riscv_ready(), a health_version of 255 is also
allowed to skip synchronization, leading to a successful probe.

Does the driver need a check to distinguish between actual legacy firmware and
a broken PCIe link?

[ ... ]
> +static int zxdh_pf_wait_riscv_ready(struct zxdh_core_dev *zxdh_dev)
> +{
[ ... ]
> +	err = readx_poll_timeout(ioread8, &hb->riscv_power_on, power_on,
> +				 power_on == 1, USEC_PER_SEC,
> +				 ZXDH_RISCV_READY_TIMEOUT_SEC * USEC_PER_SEC);

[Severity: Low]
This is another instance of passing 1,000,000 microseconds to usleep_range()
via readx_poll_timeout. Should this use a smaller sleep interval or an
msleep()-backed polling macro?

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-01 13:07 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-28 12:20 [PATCH net-next v4 0/3] dinghai: firmware handshake, MSI-X pools and async event queues han.junyang
2026-09-28 12:24 ` [PATCH net-next v4 1/3] dinghai: add firmware version check and RISC-V readiness polling han.junyang
2026-10-01 13:06   ` Simon Horman
2026-09-28 12:27 ` [PATCH net-next v4 2/3] dinghai: add MSI-X interrupt pools han.junyang
2026-09-28 12:33 ` [PATCH net-next v4 3/3] dinghai: add async event queue for firmware notifications han.junyang

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®