mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set
@ 2026-10-09 14:41 Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 01/13] net: mana: add queue-set allocation and teardown helpers Wei Hu
                   ` (12 more replies)
  0 siblings, 13 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

I am carrying Long Li's queue-set reconfiguration series forward for this
revision. Long remains the author of twelve patches, and Dipayaan Roy
remains the author of the detach change. Their authorship and signoffs
are retained; I have added my signoff for this revision.

Replace detach/attach reconfiguration with pre-allocation and queue-set
replacement. Allocation failure leaves the running configuration intact.
Publication failure attempts rollback; a failed rollback closes the port
and holds carrier down until a successful reopen.

EQs and statistics belong to the port. Valid user-configured RSS tables
survive queue rebuilds. Channel-count changes reuse surviving queues:
reductions retire the tail and increases allocate only the added queues.
Ring, MTU, private-flag and XDP layout changes still perform full rebuilds.

This series does not depend on the pending HWC concurrency/dynamic-depth
v6 series linked below. Their CQ-publication changes overlap in
gdma_main.c and hw_channel.c; whichever series lands second will need a
rebase preserving both table-publication and per-CQ lifetime guarantees.
Related HWC series:
https://lore.kernel.org/all/cover.1790665894.git.weh@linux.microsoft.com/

Changes since v5:

  - Rebase the independent queue-set series onto net-next base
    47a1446725732cd3996edf607e8739334bbf4d78.
  - Drain ALL old TX queues after closing the TX/XDP gate and waiting for
    readers, but before publishing pointers, counts or RSS. A shared
    120-second drain deadline rejects a replacement by resuming the old
    configuration without reprogramming RSS or resetting the function.
    Retirement does not use FLR. Ordinary mana_dealloc_queues() remains
    byte-identical to this base: existing reset policy and independent
    failed-FLR hardening are deliberately not changed.
  - Complete the shared EQ/CQ lifetime prerequisite in patch 2. Initialize
    EQ state before IRQ publication. After stopping a retiring CQ's NAPI
    and work and tearing down its WQ object, flush its actual parent EQ
    while the CQ callback is still published. Clear its table entry and
    wait for RCU IRQ readers before freeing the CQ and callback owner.
    Pair IRQ lookups with release publication at all CQ publishers;
    Ethernet publication follows NAPI/DIM initialization. Matching HWC
    and RDMA publication stores do not change their recovery policies.
  - Move port-lifetime statistics preparation ahead of the first live
    swap, now patch 3. Introduce RX retirement, writer handoff and folding
    atomically with the first swap in patch 4, rather than relying on a
    later patch to make intermediate statistics readers safe.
  - Keep publication barriers consistent with open/attach. Latch forced
    carrier shutdown regardless of the previous carrier state and clear
    it only after successful reopen. Retain the rollback reader grace
    period after restoring old pointers, before callers discard scratch
    state.
  - Join TX queue selection to the publication gate in patch 4. Acquire
    port_is_up before count/RSS reads and snapshot real_num_tx_queues.
    Retiring tail RQs can outlive a TX-count reduction; when a recorded
    RX index no longer fits the live TX count, use the current RSS mapping
    instead of returning that stale index.
  - Clear DRV_XOFF on inactive allocated TX queues in patch 4's cold-path
    restart helper, using netif_tx_start_queue() while preserving the
    availability checks and wake behavior for active queues. Publication
    stops all allocated queues, and the netdev watchdog scans allocated
    num_tx_queues, not just real_num_tx_queues. Leaving the inactive tail
    stopped after a channel shrink caused a watchdog timeout and recovery
    on the previous release candidate.
  - Keep old RX statistics retired until rollback steering restoration
    succeeds. Restart DIM only for RXQs actually returning from retirement,
    with NAPI disabled, work cancelled and DIM reinitialized; carried RXQs
    and their counters are not reset.
  - Reserve live MTU/XDP replacement with channel_changing under the
    existing vport_mutex through queue, fail-close and scratch cleanup.
    This excludes concurrent RDMA RAW-QP vport admission without changing
    early validation, no-op or down-port paths.
  - Correct the full-page RX shortcut description to include all existing
    early-return conditions.

Resource and control-path costs:

Full rebuilds temporarily retain both sets of SQ/RQ/CQ objects, RX buffers
and page pools. The port-owned EQ pool is shared, not doubled. Device
resource limits can reject this temporary peak; no fixed hardware queue
threshold is assumed. Allocation failure preserves the live configuration.
The pre-publication TX drain can also reject a replacement without changing
the live pointer/count/RSS configuration.

CQ retirement adds parent-EQ marker operations and RCU reader waits on the
control path; these are not free and no latency improvement is claimed.
The IRQ lookup uses an acquire load, not a new per-packet lock. This series
does not introduce a general recovery policy for failed hardware commands.

Patch layout:
  1:     Queue-set helpers, preserving ordinary deallocation policy.
  2:     Shared EQ pool and EQ/CQ publication/retirement prerequisite.
  3:     Port-lifetime statistics preparation.
  4:     First live swap, nonreset drain, RX handoff and safe TX restart.
  5-8:   Ring, private-flag, MTU and XDP reconfiguration.
  9:     Remove the unreachable detach early return.
  10-11: Release unused EQs and preserve user RSS tables.
  12-13: Reuse queues across channel-count reductions and increases.

Overlap with separately submitted net fixes:

The XDP pre-allocation pointer fix is subsumed by patch 8's queue-set path,
which leaves the live program unchanged during allocation:
https://lore.kernel.org/all/20260904202640.3900685-1-longli@microsoft.com/

The user-RSS-table fix overlaps patch 11. Keep the three-argument
mana_rss_table_keep() and its queue-set callers in the overlapping blocks;
the RTNL RSS requirement is shared with the standalone net submission:
https://lore.kernel.org/all/20260905004401.3937066-1-longli@microsoft.com/

Neither submission is resent as a standalone patch here. These merge
notes do not prescribe changes to unrelated net-tree code.

History:
  v5: https://lore.kernel.org/all/20260909222416.884246-1-longli@microsoft.com/
  v4: https://lore.kernel.org/all/20260908032843.397667-1-longli@microsoft.com/
  v3: https://lore.kernel.org/all/20260901014442.2945689-1-longli@microsoft.com/
  v2 repost:
      https://lore.kernel.org/all/20260813050418.2906468-1-longli@microsoft.com/
  v2 earlier posting:
      https://lore.kernel.org/all/20260811063506.2428213-1-longli@microsoft.com/
  Full-page RX predecessor (v12, not a queue-set v1):
      https://lore.kernel.org/all/20260711041415.3008868-1-dipayanroy@linux.microsoft.com/

The link labelled "v1" in the old queue-set covers actually points to the
four-patch full-page RX v12 predecessor, not an earlier 13-patch queue-set
posting.

Testing:

Submission tip d7c25bddfea3 differs from the tested production history
only by removing attribution/session trailers. All thirteen intermediate
source trees are identical; build/runtime identities below are unchanged.

Production 6243b02cc714, kernel 7.3.0-rc4-mana-qset-v6-g6243b02cc714,
passed 39 native functional cases per DUT on VM5 and VM6: >=35-second
reduced-channel watchdog dwells, queue ownership, real IPv4 DF jumbo
traffic at MTU 2000/payload 1972, and native XDP PASS/DROP/TX/REDIRECT.

The separate KASAN/DMA memory/fault fixture,
7.3.0-rc4-mana-qset-memdiag-rethook-g9d9757ecf25c, completed
64 PASS/2 UNSUPPORTED/0 FAIL/0 NOT_RUN per DUT, including all 25 fault
cases with each token consumed exactly once. VM5 also completed twelve
>=180-second warning-reproduction trials with attributed security probes
active and untouched. No unexpected warning, KASAN/DMA-debug error or
peer-health failure was observed. Fault4 recovered locally; its independent
fallback was armed, verified and canceled, not executed.

That fixture's four-line FP64 rethook metadata correction and fault hooks
are test-only, not in these thirteen patches or a production dependency.
The entry window is unchanged; no global unwinding guarantee is claimed.
Full-locking remains baseline-blocked: unmodified 47a reproduced the
platform checker failure, and this fixture has no LOCKDEP/PROVE_RAW/
PROVE_RCU coverage. DIM enable returned EOPNOTSUPP. Native verbs is
UNSUPPORTED in the memory matrix because no approved native baseline
probe/result was supplied; RDMA sanity/link health is covered, not native
verbs transfer. PF/non-host-mode and multiport hardware remain untested.

A final forward-UDP receiver comparison completed five release6243 and
five baseline47a samples with complete raw records, matching identities,
nested windows and valid client/server completion; no rows were rejected.
Adjusted TGID CPU averaged 4056.4 versus 4188.4 USER_HZ ticks. Per iperf
summary application Gbit it was 901.422346 versus 930.754115 (~3.15% lower);
per native full-window path Gbit, 518.165386 versus 543.335320 (~4.63% lower).
The process interval includes omit time; iperf's application summary does
not. These are distinct denominators, not interval-matched application
efficiency. Throughput was 99.9980098 versus 99.9980650 Mbit/s, with loss
0.000689492% versus 0.001422219% (release versus baseline).

This measures guest-scheduler-attributed process runtime, not whole-guest,
driver or physical CPU efficiency. Aggregate /proc/stat is observational
and need not conserve against adjusted process time. A separate earlier
five-plus-five sender crossover did not reproduce the initial ~9% signal;
observer/context and phase confounds remain. These bounded results are
not a no-effect or blanket regression-free guarantee.

Current P4-P13 passed forced GCC W=1 builds, 150/150 target translation
units; unchanged P1-P3 retain their original complete proof. Smatch is
not credited: its frontend fails a valid type assertion on both baseline
and feature. Historical supported Sparse coverage is not relabeled as a
new current-tip run. Full-mail lint was rerun after removing the
Copilot attribution/session trailers at the submitter's request.
Hardware-command failure/recovery policies remain scoped as described
above; these tests do not establish recovery from every device failure.

Full evidence, fixture identities and historical limitations are in the
package's RUNTIME-HANDOFF.txt and audit/final-evidence.json. Public v5
review freshness was checked on 2026-10-02: all 21 findings and 28 thread
messages were unchanged, with no new human replies. No actual Sashiko
approval is claimed.

Signed-off-by: Wei Hu <weh@microsoft.com>

Dipayaan Roy (1):
  net: mana: do not bail out of mana_detach on dealloc failure

Long Li (12):
  net: mana: add queue-set allocation and teardown helpers
  net: mana: share the EQ pool across a queue-set swap
  net: mana: keep per-queue statistics in the port context
  net: mana: swap queue sets in mana_set_channels
  net: mana: swap queue sets in mana_set_ringparam
  net: mana: swap queue sets in mana_set_priv_flags
  net: mana: swap queue sets in mana_change_mtu
  net: mana: swap queue sets in mana_xdp_set
  net: mana: release EQs left idle by a channel-count reduction
  net: mana: keep a user-configured RSS table across a queue rebuild
  net: mana: keep the surviving queues when the channel count is reduced
  net: mana: keep the existing queues when the channel count is raised

 drivers/infiniband/hw/mana/cq.c               |    3 +-
 .../net/ethernet/microsoft/mana/gdma_main.c   |   22 +-
 .../net/ethernet/microsoft/mana/hw_channel.c  |    3 +-
 .../net/ethernet/microsoft/mana/mana_bpf.c    |  108 +-
 drivers/net/ethernet/microsoft/mana/mana_en.c | 1328 +++++++++++++++--
 .../ethernet/microsoft/mana/mana_ethtool.c    |  314 ++--
 include/net/mana/gdma.h                       |    8 +-
 include/net/mana/mana.h                       |   97 +-
 8 files changed, 1610 insertions(+), 273 deletions(-)


base-commit: 47a1446725732cd3996edf607e8739334bbf4d78

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 01/13] net: mana: add queue-set allocation and teardown helpers
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 02/13] net: mana: share the EQ pool across a queue-set swap Wei Hu
                   ` (11 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Add queue-set allocation and teardown helpers using a scratch port
context without releasing the vport. These prepare reconfiguration
to retain its running queues if replacement allocation fails.

Leave ordinary mana_dealloc_queues() byte-identical to the pinned
base. Do not introduce reset generations, sibling-port rebuilds, or
an incomplete independent failed-FLR TX-buffer policy change.

Queue-set retirement itself performs no function reset. Its caller
must drain old published TX queues before publishing a replacement;
unpublished sets have never admitted TX. The first live-swap commit
adds that nonresetting pre-publication wait. These helpers have no
live replacement callers yet.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 .../net/ethernet/microsoft/mana/mana_bpf.c    |  24 +++
 drivers/net/ethernet/microsoft/mana/mana_en.c | 184 +++++++++++++++++-
 include/net/mana/mana.h                       |  31 +++
 3 files changed, 236 insertions(+), 3 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_bpf.c b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
index 9ef42b74048b..ff54f8966825 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_bpf.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
@@ -263,3 +263,27 @@ int mana_bpf(struct net_device *ndev, struct netdev_bpf *bpf)
 		return -EOPNOTSUPP;
 	}
 }
+
+struct bpf_prog *mana_chn_xdp_peek(struct mana_port_context *apc)
+{
+	ASSERT_RTNL();
+
+	if (!apc->rxqs || !apc->rxqs[0])
+		return NULL;
+
+	return rtnl_dereference(apc->rxqs[0]->bpf_prog);
+}
+
+/* Keep the per-queue program pointers until RX polling stops. */
+void mana_chn_xdp_release(struct bpf_prog *prog, unsigned int num_queues)
+{
+	unsigned int i;
+
+	ASSERT_RTNL();
+
+	if (!prog)
+		return;
+
+	for (i = 0; i < num_queues; i++)
+		bpf_prog_put(prog);
+}
diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 591fb4191d90..d4b8bb4e1f53 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -2018,7 +2018,8 @@ static void mana_poll_tx_cq(struct mana_cq *cq)
 	/* Ensure checking txq_stopped before apc->port_is_up. */
 	smp_rmb();
 
-	if (txq_stopped && apc->port_is_up && avail_space >= MAX_TX_WQE_SIZE) {
+	if (txq_stopped && !READ_ONCE(txq->retiring) && apc->port_is_up &&
+	    avail_space >= MAX_TX_WQE_SIZE) {
 		netif_tx_wake_queue(net_txq);
 		apc->eth_stats.wake_queue++;
 	}
@@ -3013,11 +3014,11 @@ static int mana_push_wqe(struct mana_rxq *rxq)
 
 static int mana_create_page_pool(struct mana_rxq *rxq, struct gdma_context *gc)
 {
-	struct mana_port_context *mpc = netdev_priv(rxq->ndev);
 	struct page_pool_params pprm = {};
 	int ret;
 
-	pprm.pool_size = mpc->rx_queue_size / rxq->frag_count + 1;
+	/* Size the pool for this RX queue, not the live configuration. */
+	pprm.pool_size = rxq->num_rx_buf / rxq->frag_count + 1;
 	pprm.nid = gc->numa_node;
 	pprm.napi = &rxq->rx_cq.napi;
 	pprm.netdev = rxq->ndev;
@@ -3767,6 +3768,179 @@ static int mana_dealloc_queues(struct net_device *ndev)
 	return 0;
 }
 
+static void mana_qset_snapshot(const struct mana_port_context *ctx,
+			       struct mana_qset *out)
+{
+	out->eqs		= ctx->eqs;
+	out->tx_qp		= ctx->tx_qp;
+	out->rxqs		= ctx->rxqs;
+	out->indir_table	= ctx->indir_table;
+	out->indir_table_sz	= ctx->indir_table_sz;
+	out->rxobj_table	= ctx->rxobj_table;
+	out->default_rxobj	= ctx->default_rxobj;
+	out->num_queues		= ctx->num_queues;
+	out->rx_queue_size	= ctx->rx_queue_size;
+	out->tx_queue_size	= ctx->tx_queue_size;
+	out->priv_flags		= ctx->priv_flags;
+}
+
+/* Vport identity and port debugfs outlive queue sets. */
+static void mana_qset_install(struct mana_port_context *ctx,
+			      const struct mana_qset *qset)
+{
+	ctx->eqs		= qset->eqs;
+	ctx->tx_qp		= qset->tx_qp;
+	ctx->rxqs		= qset->rxqs;
+	ctx->indir_table	= qset->indir_table;
+	ctx->indir_table_sz	= qset->indir_table_sz;
+	ctx->rxobj_table	= qset->rxobj_table;
+	ctx->default_rxobj	= qset->default_rxobj;
+	ctx->num_queues		= qset->num_queues;
+	ctx->rx_queue_size	= qset->rx_queue_size;
+	ctx->tx_queue_size	= qset->tx_queue_size;
+	ctx->priv_flags		= qset->priv_flags;
+}
+
+/* Copy the vport identity without borrowing the live queues. */
+struct mana_port_context *mana_qset_scratch_alloc(struct mana_port_context *apc)
+{
+	struct mana_port_context *scratch;
+
+	scratch = kvzalloc_obj(*scratch, GFP_KERNEL);
+	if (!scratch)
+		return NULL;
+
+	*scratch = *apc;
+
+	scratch->eqs		= NULL;
+	scratch->tx_qp		= NULL;
+	scratch->rxqs		= NULL;
+	scratch->indir_table	= NULL;
+	scratch->rxobj_table	= NULL;
+	scratch->default_rxobj	= INVALID_MANA_HANDLE;
+	scratch->mana_eqs_debugfs = NULL;
+
+	/* Do not consume the live set's pre-allocated RX buffers. */
+	scratch->rxbufs_pre	= NULL;
+	scratch->das_pre	= NULL;
+	scratch->rxbpre_total	= 0;
+
+	/* Suppress debugfs names that would collide with the live set. */
+	scratch->mana_port_debugfs = ERR_PTR(-ENODEV);
+
+	return scratch;
+}
+
+void mana_qset_scratch_free(struct mana_port_context *scratch)
+{
+	kvfree(scratch);
+}
+
+int mana_alloc_qset(struct mana_port_context *scratch, unsigned int num_queues,
+		    unsigned int rx_queue_size, unsigned int tx_queue_size,
+		    u32 priv_flags, struct mana_qset *out)
+{
+	struct net_device *ndev = scratch->ndev;
+	int err;
+
+	ASSERT_RTNL();
+
+	scratch->num_queues	= num_queues;
+	scratch->rx_queue_size	= rx_queue_size;
+	scratch->tx_queue_size	= tx_queue_size;
+	scratch->priv_flags	= priv_flags;
+
+	err = mana_init_port_context(scratch);
+	if (err)
+		goto out_err;
+
+	err = mana_rss_table_alloc(scratch);
+	if (err)
+		goto cleanup_rxq_array;
+
+	err = mana_create_eq(scratch);
+	if (err)
+		goto cleanup_rss;
+
+	err = mana_create_txq(scratch, ndev);
+	if (err)
+		goto cleanup_eq;
+
+	err = mana_add_rx_queues(scratch, ndev);
+	if (err)
+		goto cleanup_rxq;
+
+	mana_rss_table_init(scratch);
+
+	mana_qset_snapshot(scratch, out);
+	return 0;
+
+cleanup_rxq:
+	mana_destroy_rxqs(scratch);
+	mana_destroy_txq(scratch);
+cleanup_eq:
+	mana_destroy_eq(scratch);
+cleanup_rss:
+	mana_cleanup_indir_table(scratch);
+cleanup_rxq_array:
+	kfree(scratch->rxqs);
+	scratch->rxqs = NULL;
+out_err:
+	netdev_err(ndev, "%s(num_queues=%u) failed: %d\n", __func__,
+		   num_queues, err);
+	return err;
+}
+
+/* Under RTNL, free only queues no longer shared with the installed set. */
+void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset)
+{
+	struct bpf_prog *retiring_prog;
+	unsigned int retiring_queues;
+
+	ASSERT_RTNL();
+
+	if (!qset->rxqs && !qset->tx_qp && !qset->eqs)
+		return;
+
+	if (qset->tx_qp) {
+		unsigned int q;
+
+		for (q = 0; q < qset->num_queues; q++) {
+			if (qset->tx_qp[q])
+				WRITE_ONCE(qset->tx_qp[q]->txq.retiring, true);
+		}
+	}
+
+	/* Keep retired queues and arrays through this grace period; local NAPI
+	 * synchronization does not drain other devices' XDP.
+	 */
+	synchronize_net();
+
+	mana_qset_install(scratch, qset);
+
+	/* Keep retiring RXQs' XDP programs and references until RX teardown. */
+	retiring_prog = mana_chn_xdp_peek(scratch);
+	retiring_queues = scratch->num_queues;
+
+	/* Published queues were drained before the swap; unpublished queues
+	 * have never admitted TX.
+	 */
+	/* Fence RQs before unmapping, but teardown proceeds on errors. */
+	mana_fence_rqs(scratch);
+
+	mana_destroy_rxqs(scratch);
+
+	mana_chn_xdp_release(retiring_prog, retiring_queues);
+
+	mana_destroy_txq(scratch);
+	mana_destroy_eq(scratch);
+	mana_cleanup_indir_table(scratch);
+	kfree(scratch->rxqs);
+	scratch->rxqs = NULL;
+
+	memset(qset, 0, sizeof(*qset));
+}
+
 int mana_detach(struct net_device *ndev, bool from_close)
 {
 	struct mana_port_context *apc = netdev_priv(ndev);
@@ -4245,6 +4419,10 @@ void mana_remove(struct gdma_dev *gd, bool suspending)
 		unregister_netdevice(ndev);
 		mana_cleanup_indir_table(apc);
 
+		/* Remove the port from reset walks before freeing its netdev.
+		 */
+		ac->ports[i] = NULL;
+
 		rtnl_unlock();
 
 		free_netdev(ndev);
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index 83b7eff4646e..6a407b34fd68 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -143,6 +143,9 @@ struct mana_txq {
 
 	bool napi_initialized;
 
+	/* Suppress completion wakeups on the replacement's netdev queue. */
+	bool retiring;
+
 	struct mana_stats_tx stats;
 };
 
@@ -537,6 +540,7 @@ struct mana_context {
 	u8 bm_hostmode;
 
 	struct mana_ethtool_hc_stats hc_stats;
+
 	struct workqueue_struct *per_port_queue_reset_wq;
 	/* Workqueue for querying hardware stats */
 	struct delayed_work gf_stats_work;
@@ -661,6 +665,23 @@ struct mana_port_context {
 	u32 steer_cqe_coalescing;
 };
 
+struct mana_qset {
+	struct mana_eq		*eqs;
+	struct mana_tx_qp	**tx_qp;
+	struct mana_rxq		**rxqs;
+
+	u32			*indir_table;
+	u32			indir_table_sz;
+	mana_handle_t		*rxobj_table;
+	mana_handle_t		default_rxobj;
+
+	unsigned int		num_queues;
+	unsigned int		rx_queue_size;
+	unsigned int		tx_queue_size;
+	u32			priv_flags;
+
+};
+
 netdev_tx_t mana_start_xmit(struct sk_buff *skb, struct net_device *ndev);
 int mana_config_rss(struct mana_port_context *ac, enum TRI_STATE rx,
 		    bool update_hash, bool update_tab);
@@ -670,6 +691,14 @@ int mana_alloc_queues(struct net_device *ndev);
 int mana_attach(struct net_device *ndev);
 int mana_detach(struct net_device *ndev, bool from_close);
 
+struct mana_port_context *
+mana_qset_scratch_alloc(struct mana_port_context *apc);
+void mana_qset_scratch_free(struct mana_port_context *scratch);
+int mana_alloc_qset(struct mana_port_context *scratch, unsigned int num_queues,
+		    unsigned int rx_queue_size, unsigned int tx_queue_size,
+		    u32 priv_flags, struct mana_qset *out);
+void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset);
+
 void mana_dim_change(struct mana_cq *cq, bool enable);
 
 int mana_probe(struct gdma_dev *gd, bool resuming);
@@ -685,6 +714,8 @@ u32 mana_run_xdp(struct net_device *ndev, struct mana_rxq *rxq,
 		 struct xdp_buff *xdp, void *buf_va, uint pkt_len);
 struct bpf_prog *mana_xdp_get(struct mana_port_context *apc);
 void mana_chn_setxdp(struct mana_port_context *apc, struct bpf_prog *prog);
+struct bpf_prog *mana_chn_xdp_peek(struct mana_port_context *apc);
+void mana_chn_xdp_release(struct bpf_prog *prog, unsigned int num_queues);
 int mana_bpf(struct net_device *ndev, struct netdev_bpf *bpf);
 int mana_query_gf_stats(struct mana_context *ac);
 int mana_query_link_cfg(struct mana_port_context *apc);

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 02/13] net: mana: share the EQ pool across a queue-set swap
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 01/13] net: mana: add queue-set allocation and teardown helpers Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 03/13] net: mana: keep per-queue statistics in the port context Wei Hu
                   ` (10 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Make the EQ pool port-owned so overlapping queue sets share EQs instead
of requiring old + new vector allocations. Allocate max_queues slots and
track populated entries with num_eqs.

Grow the pool before creating replacement CQs. Additional EQs survive
allocation failure in this patch and are released at port teardown.

Initialize EQ callbacks, context, owner phase and throttle before IRQ
publication, so a live shared IRQ cannot observe an uninitialized EQ.
Preserve the existing hardware-create and unwind ordering.

A shared EQ also outlives individual CQs. After stopping a CQ's NAPI and
work and tearing down its WQ object, flush its actual parent EQ while the
CQ callback remains published. Remove only that CQ's table entry and wait
for RCU IRQ readers before freeing the CQ and its callback owner.

Pair the IRQ lookup with release publication at all CQ publishers. For
Ethernet, publish only after NAPI and DIM initialization. The HWC and RDMA
sites need the matching stores, not changes to their teardown or recovery
policies.

The EQ markers and reader waits are control-path costs; the IRQ lookup
uses an acquire load without new validity checks. Report failed markers
using the existing error-logging convention. This does not add a reset or
reclamation policy for arbitrary failed hardware commands.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 drivers/infiniband/hw/mana/cq.c               |   3 +-
 .../net/ethernet/microsoft/mana/gdma_main.c   |  22 ++--
 .../net/ethernet/microsoft/mana/hw_channel.c  |   3 +-
 drivers/net/ethernet/microsoft/mana/mana_en.c | 115 ++++++++++++++----
 include/net/mana/mana.h                       |   6 +-
 5 files changed, 113 insertions(+), 36 deletions(-)

diff --git a/drivers/infiniband/hw/mana/cq.c b/drivers/infiniband/hw/mana/cq.c
index d4e5e3f91268..0b6ad2547e0e 100644
--- a/drivers/infiniband/hw/mana/cq.c
+++ b/drivers/infiniband/hw/mana/cq.c
@@ -155,7 +155,8 @@ int mana_ib_install_cq_cb(struct mana_ib_dev *mdev, struct mana_ib_cq *cq)
 	gdma_cq->type = GDMA_CQ;
 	gdma_cq->cq.callback = mana_ib_cq_handler;
 	gdma_cq->id = cq->queue.id;
-	gc->cq_table[cq->queue.id] = gdma_cq;
+	/* Pairs with the acquire load in mana_gd_process_eqe(). */
+	smp_store_release(&gc->cq_table[cq->queue.id], gdma_cq);
 	return 0;
 }
 
diff --git a/drivers/net/ethernet/microsoft/mana/gdma_main.c b/drivers/net/ethernet/microsoft/mana/gdma_main.c
index 8e9bfc1d6a2a..a7a491156f64 100644
--- a/drivers/net/ethernet/microsoft/mana/gdma_main.c
+++ b/drivers/net/ethernet/microsoft/mana/gdma_main.c
@@ -925,7 +925,8 @@ static void mana_gd_process_eqe(struct gdma_queue *eq)
 		if (WARN_ON_ONCE(cq_id >= gc->max_num_cqs))
 			break;
 
-		cq = gc->cq_table[cq_id];
+		/* Match release publication of the CQ and its callback state. */
+		cq = smp_load_acquire(&gc->cq_table[cq_id]);
 		if (WARN_ON_ONCE(!cq || cq->type != GDMA_CQ || cq->id != cq_id))
 			break;
 
@@ -1187,17 +1188,17 @@ static int mana_gd_create_eq(struct gdma_dev *gd,
 		return -EINVAL;
 	}
 
+	queue->eq.callback = spec->eq.callback;
+	queue->eq.context = spec->eq.context;
+	queue->head |= INITIALIZED_OWNER_BIT(log2_num_entries);
+	queue->eq.log2_throttle_limit = spec->eq.log2_throttle_limit ?: 1;
+
 	err = mana_gd_register_irq(queue, spec);
 	if (err) {
 		dev_err(dev, "Failed to register irq: %d\n", err);
 		return err;
 	}
 
-	queue->eq.callback = spec->eq.callback;
-	queue->eq.context = spec->eq.context;
-	queue->head |= INITIALIZED_OWNER_BIT(log2_num_entries);
-	queue->eq.log2_throttle_limit = spec->eq.log2_throttle_limit ?: 1;
-
 	if (create_hwq) {
 		err = mana_gd_create_hw_eq(gc, queue);
 		if (err)
@@ -1232,13 +1233,12 @@ static void mana_gd_destroy_cq(struct gdma_context *gc,
 {
 	u32 id = queue->id;
 
-	if (id >= gc->max_num_cqs)
+	if (id >= gc->max_num_cqs || !gc->cq_table)
 		return;
 
-	if (!gc->cq_table[id])
-		return;
-
-	gc->cq_table[id] = NULL;
+	/* Leave a reused ID alone, but still drain readers of this CQ. */
+	cmpxchg(&gc->cq_table[id], queue, NULL);
+	synchronize_rcu();
 }
 
 int mana_gd_create_hwc_queue(struct gdma_dev *gd,
diff --git a/drivers/net/ethernet/microsoft/mana/hw_channel.c b/drivers/net/ethernet/microsoft/mana/hw_channel.c
index 3bca4b683134..aefb8646cd79 100644
--- a/drivers/net/ethernet/microsoft/mana/hw_channel.c
+++ b/drivers/net/ethernet/microsoft/mana/hw_channel.c
@@ -702,7 +702,8 @@ static int mana_hwc_establish_channel(struct gdma_context *gc, u16 *q_depth,
 	if (!gc->cq_table)
 		return -ENOMEM;
 
-	gc->cq_table[cq->id] = cq;
+	/* Pairs with the acquire load in mana_gd_process_eqe(). */
+	smp_store_release(&gc->cq_table[cq->id], cq);
 
 	return 0;
 }
diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index d4b8bb4e1f53..45cb23717151 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -1733,7 +1733,7 @@ void mana_destroy_eq(struct mana_port_context *apc)
 	debugfs_remove_recursive(apc->mana_eqs_debugfs);
 	apc->mana_eqs_debugfs = NULL;
 
-	for (i = 0; i < apc->num_queues; i++) {
+	for (i = 0; i < apc->num_eqs; i++) {
 		eq = apc->eqs[i].eq;
 		if (!eq)
 			continue;
@@ -1745,6 +1745,7 @@ void mana_destroy_eq(struct mana_port_context *apc)
 
 	kfree(apc->eqs);
 	apc->eqs = NULL;
+	apc->num_eqs = 0;
 }
 EXPORT_SYMBOL_NS(mana_destroy_eq, "NET_MANA");
 
@@ -1773,9 +1774,11 @@ int mana_create_eq(struct mana_port_context *apc)
 
 	if (WARN_ON(apc->eqs))
 		return -EEXIST;
-	apc->eqs = kzalloc_objs(struct mana_eq, apc->num_queues);
+	/* Keep EQ array addresses stable while CQs reference them. */
+	apc->eqs = kzalloc_objs(struct mana_eq, apc->max_queues);
 	if (!apc->eqs)
 		return -ENOMEM;
+	apc->num_eqs = 0;
 
 	spec.type = GDMA_EQ;
 	spec.monitor_avl_buf = false;
@@ -1805,6 +1808,7 @@ int mana_create_eq(struct mana_port_context *apc)
 		}
 		apc->eqs[i].eq->eq.irq = gic->irq;
 		mana_create_eq_debugfs(apc, i);
+		apc->num_eqs = i + 1;
 	}
 
 	return 0;
@@ -1814,6 +1818,61 @@ int mana_create_eq(struct mana_port_context *apc)
 }
 EXPORT_SYMBOL_NS(mana_create_eq, "NET_MANA");
 
+/* Grow the shared EQ pool without replacing live entries. */
+static int mana_grow_eqs(struct mana_port_context *apc, unsigned int need)
+{
+	struct gdma_dev *gd = apc->ac->gdma_dev;
+	struct gdma_context *gc = gd->gdma_context;
+	struct gdma_queue_spec spec = {};
+	struct gdma_irq_context *gic;
+	unsigned int i;
+	int err;
+	int msi;
+
+	if (WARN_ON(!apc->eqs))
+		return -EINVAL;
+
+	if (need > apc->max_queues)
+		return -EINVAL;
+
+	if (need <= apc->num_eqs)
+		return 0;
+
+	spec.type = GDMA_EQ;
+	spec.monitor_avl_buf = false;
+	spec.queue_size = EQ_SIZE;
+	spec.eq.callback = NULL;
+	spec.eq.context = apc->eqs;
+	spec.eq.log2_throttle_limit = LOG2_EQ_THROTTLE;
+
+	for (i = apc->num_eqs; i < need; i++) {
+		msi = (i + 1) % gc->num_msix_usable;
+
+		gic = mana_gd_get_gic(gc, !gc->msi_sharing, &msi);
+		if (IS_ERR(gic)) {
+			err = PTR_ERR(gic);
+			goto out;
+		}
+		spec.eq.msix_index = msi;
+
+		err = mana_gd_create_mana_eq(gd, &spec, &apc->eqs[i].eq);
+		if (err) {
+			dev_err(gc->dev, "Failed to grow EQ %u : %d\n", i, err);
+			mana_gd_put_gic(gc, !gc->msi_sharing, msi);
+			goto out;
+		}
+		apc->eqs[i].eq->eq.irq = gic->irq;
+		mana_create_eq_debugfs(apc, i);
+		apc->num_eqs = i + 1;
+	}
+
+	return 0;
+out:
+	/* Retain partial growth for reuse; the live set still needs this pool.
+	 */
+	return err;
+}
+
 static int mana_fence_rq(struct mana_port_context *apc, struct mana_rxq *rxq)
 {
 	struct mana_fence_rq_resp resp = {};
@@ -2624,12 +2683,25 @@ static void mana_schedule_napi(void *context, struct gdma_queue *gdma_queue)
 
 static void mana_deinit_cq(struct mana_port_context *apc, struct mana_cq *cq)
 {
-	struct gdma_dev *gd = apc->ac->gdma_dev;
+	struct gdma_context *gc = apc->ac->gdma_dev->gdma_context;
+	struct gdma_queue *gdma_cq = cq->gdma_cq;
+	struct gdma_queue *eq;
+	int err;
 
-	if (!cq->gdma_cq)
+	if (!gdma_cq)
 		return;
 
-	mana_gd_destroy_queue(gd->gdma_context, cq->gdma_cq);
+	eq = gdma_cq->cq.parent;
+	if (gdma_cq->id < gc->max_num_cqs && eq &&
+	    eq->id != INVALID_QUEUE_ID) {
+		/* Flush queued events after WQ teardown, before removing this CQ. */
+		err = mana_gd_test_eq(gc, eq);
+		if (err && mana_en_need_log(apc, err))
+			netdev_err(apc->ndev, "Failed to flush EQ %u for CQ %u: %d\n",
+				   eq->id, gdma_cq->id, err);
+	}
+
+	mana_gd_destroy_queue(gc, gdma_cq);
 }
 
 static void mana_deinit_txq(struct mana_port_context *apc, struct mana_txq *txq)
@@ -2824,8 +2896,6 @@ static int mana_create_txq(struct mana_port_context *apc,
 			goto out;
 		}
 
-		gc->cq_table[cq->gdma_id] = cq->gdma_cq;
-
 		mana_create_txq_debugfs(apc, i);
 
 		set_bit(NAPI_STATE_NO_BUSY_POLL, &cq->napi.state);
@@ -2840,6 +2910,9 @@ static int mana_create_txq(struct mana_port_context *apc,
 		napi_enable_locked(&cq->napi);
 		txq->napi_initialized = true;
 
+		/* Publish the initialized NAPI/DIM state to the EQ handler. */
+		smp_store_release(&gc->cq_table[cq->gdma_id], cq->gdma_cq);
+
 		mana_gd_ring_cq(cq->gdma_cq, SET_ARM_BIT);
 	}
 
@@ -3150,8 +3223,6 @@ static struct mana_rxq *mana_create_rxq(struct mana_port_context *apc,
 		goto out;
 	}
 
-	gc->cq_table[cq->gdma_id] = cq->gdma_cq;
-
 	netif_napi_add_weight_locked(ndev, &cq->napi, mana_poll, 1);
 
 	WARN_ON(xdp_rxq_info_reg(&rxq->xdp_rxq, ndev, rxq_idx,
@@ -3167,6 +3238,9 @@ static struct mana_rxq *mana_create_rxq(struct mana_port_context *apc,
 
 	napi_enable_locked(&cq->napi);
 
+	/* Publish the initialized NAPI/DIM state to the EQ handler. */
+	smp_store_release(&gc->cq_table[cq->gdma_id], cq->gdma_cq);
+
 	mana_gd_ring_cq(cq->gdma_cq, SET_ARM_BIT);
 out:
 	if (!err)
@@ -3771,7 +3845,6 @@ static int mana_dealloc_queues(struct net_device *ndev)
 static void mana_qset_snapshot(const struct mana_port_context *ctx,
 			       struct mana_qset *out)
 {
-	out->eqs		= ctx->eqs;
 	out->tx_qp		= ctx->tx_qp;
 	out->rxqs		= ctx->rxqs;
 	out->indir_table	= ctx->indir_table;
@@ -3788,7 +3861,6 @@ static void mana_qset_snapshot(const struct mana_port_context *ctx,
 static void mana_qset_install(struct mana_port_context *ctx,
 			      const struct mana_qset *qset)
 {
-	ctx->eqs		= qset->eqs;
 	ctx->tx_qp		= qset->tx_qp;
 	ctx->rxqs		= qset->rxqs;
 	ctx->indir_table	= qset->indir_table;
@@ -3801,7 +3873,9 @@ static void mana_qset_install(struct mana_port_context *ctx,
 	ctx->priv_flags		= qset->priv_flags;
 }
 
-/* Copy the vport identity without borrowing the live queues. */
+/* Scratch starts without SQs/RQs and borrows the port's EQ pool. Never call
+ * mana_destroy_eq() on it.
+ */
 struct mana_port_context *mana_qset_scratch_alloc(struct mana_port_context *apc)
 {
 	struct mana_port_context *scratch;
@@ -3812,13 +3886,11 @@ struct mana_port_context *mana_qset_scratch_alloc(struct mana_port_context *apc)
 
 	*scratch = *apc;
 
-	scratch->eqs		= NULL;
 	scratch->tx_qp		= NULL;
 	scratch->rxqs		= NULL;
 	scratch->indir_table	= NULL;
 	scratch->rxobj_table	= NULL;
 	scratch->default_rxobj	= INVALID_MANA_HANDLE;
-	scratch->mana_eqs_debugfs = NULL;
 
 	/* Do not consume the live set's pre-allocated RX buffers. */
 	scratch->rxbufs_pre	= NULL;
@@ -3836,7 +3908,8 @@ void mana_qset_scratch_free(struct mana_port_context *scratch)
 	kvfree(scratch);
 }
 
-int mana_alloc_qset(struct mana_port_context *scratch, unsigned int num_queues,
+int mana_alloc_qset(struct mana_port_context *apc,
+		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
 		    u32 priv_flags, struct mana_qset *out)
 {
@@ -3858,13 +3931,16 @@ int mana_alloc_qset(struct mana_port_context *scratch, unsigned int num_queues,
 	if (err)
 		goto cleanup_rxq_array;
 
-	err = mana_create_eq(scratch);
+	err = mana_grow_eqs(apc, num_queues);
 	if (err)
 		goto cleanup_rss;
 
+	scratch->eqs = apc->eqs;
+	scratch->num_eqs = apc->num_eqs;
+
 	err = mana_create_txq(scratch, ndev);
 	if (err)
-		goto cleanup_eq;
+		goto cleanup_rss;
 
 	err = mana_add_rx_queues(scratch, ndev);
 	if (err)
@@ -3878,8 +3954,6 @@ int mana_alloc_qset(struct mana_port_context *scratch, unsigned int num_queues,
 cleanup_rxq:
 	mana_destroy_rxqs(scratch);
 	mana_destroy_txq(scratch);
-cleanup_eq:
-	mana_destroy_eq(scratch);
 cleanup_rss:
 	mana_cleanup_indir_table(scratch);
 cleanup_rxq_array:
@@ -3899,7 +3973,7 @@ void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset)
 
 	ASSERT_RTNL();
 
-	if (!qset->rxqs && !qset->tx_qp && !qset->eqs)
+	if (!qset->rxqs && !qset->tx_qp)
 		return;
 
 	if (qset->tx_qp) {
@@ -3933,7 +4007,6 @@ void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset)
 	mana_chn_xdp_release(retiring_prog, retiring_queues);
 
 	mana_destroy_txq(scratch);
-	mana_destroy_eq(scratch);
 	mana_cleanup_indir_table(scratch);
 	kfree(scratch->rxqs);
 	scratch->rxqs = NULL;
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index 6a407b34fd68..d3a79e13e343 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -563,7 +563,9 @@ struct mana_port_context {
 
 	u8 mac_addr[ETH_ALEN];
 
+	/* Port-owned EQ pool: max_queues slots, num_eqs populated. */
 	struct mana_eq *eqs;
+	unsigned int num_eqs;
 	struct dentry *mana_eqs_debugfs;
 
 	enum TRI_STATE rss_state;
@@ -666,7 +668,6 @@ struct mana_port_context {
 };
 
 struct mana_qset {
-	struct mana_eq		*eqs;
 	struct mana_tx_qp	**tx_qp;
 	struct mana_rxq		**rxqs;
 
@@ -694,7 +695,8 @@ int mana_detach(struct net_device *ndev, bool from_close);
 struct mana_port_context *
 mana_qset_scratch_alloc(struct mana_port_context *apc);
 void mana_qset_scratch_free(struct mana_port_context *scratch);
-int mana_alloc_qset(struct mana_port_context *scratch, unsigned int num_queues,
+int mana_alloc_qset(struct mana_port_context *apc,
+		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
 		    u32 priv_flags, struct mana_qset *out);
 void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset);

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 03/13] net: mana: keep per-queue statistics in the port context
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 01/13] net: mana: add queue-set allocation and teardown helpers Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 02/13] net: mana: share the EQ pool across a queue-set swap Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 04/13] net: mana: swap queue sets in mana_set_channels Wei Hu
                   ` (9 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Prepare port-lifetime RX/TX statistics before introducing live queue
replacement. Allocate max_queues slots before netdev registration,
unwind probe failures, and release them only after unregistering the
netdev. Queue writers use pointers into these arrays.

Make ndo_get_stats64 and ethtool statistics independent of replaceable
queue objects. Preserve counters while down, gating only the PHY
query. Reserve zeroed retired-RX slots for the next patch; no second
live queue generation or RX retirement handoff is enabled here.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 .../net/ethernet/microsoft/mana/mana_bpf.c    |   4 +-
 drivers/net/ethernet/microsoft/mana/mana_en.c | 107 ++++++++++++++----
 .../ethernet/microsoft/mana/mana_ethtool.c    |  48 ++++++--
 include/net/mana/mana.h                       |  15 ++-
 4 files changed, 139 insertions(+), 35 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_bpf.c b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
index ff54f8966825..1905214bec48 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_bpf.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
@@ -68,7 +68,7 @@ int mana_xdp_xmit(struct net_device *ndev, int n, struct xdp_frame **frames,
 		count++;
 	}
 
-	tx_stats = &apc->tx_qp[q_idx]->txq.stats;
+	tx_stats = apc->tx_qp[q_idx]->txq.stats;
 
 	u64_stats_update_begin(&tx_stats->syncp);
 	tx_stats->xdp_xmit += count;
@@ -95,7 +95,7 @@ u32 mana_run_xdp(struct net_device *ndev, struct mana_rxq *rxq,
 
 	act = bpf_prog_run_xdp(prog, xdp);
 
-	rx_stats = &rxq->stats;
+	rx_stats = rxq->stats;
 
 	switch (act) {
 	case XDP_PASS:
diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 45cb23717151..02cd5f7656ed 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -372,7 +372,7 @@ netdev_tx_t mana_start_xmit(struct sk_buff *skb, struct net_device *ndev)
 	txq = &apc->tx_qp[txq_idx]->txq;
 	gdma_sq = txq->gdma_sq;
 	cq = &apc->tx_qp[txq_idx]->tx_cq;
-	tx_stats = &txq->stats;
+	tx_stats = txq->stats;
 
 	BUILD_BUG_ON(MAX_TX_WQE_SGL_ENTRIES != MANA_MAX_TX_WQE_SGL_ENTRIES);
 	if (MAX_SKB_FRAGS + 2 > MAX_TX_WQE_SGL_ENTRIES &&
@@ -551,7 +551,7 @@ netdev_tx_t mana_start_xmit(struct sk_buff *skb, struct net_device *ndev)
 	/* Populated the packet and bytes counters based on post GSO packet
 	 * calculations
 	 */
-	tx_stats = &txq->stats;
+	tx_stats = txq->stats;
 	u64_stats_update_begin(&tx_stats->syncp);
 	tx_stats->packets += num_gso_seg;
 	tx_stats->bytes += len + ((num_gso_seg - 1) * gso_hs);
@@ -597,15 +597,15 @@ static void mana_get_stats64(struct net_device *ndev,
 			     struct rtnl_link_stats64 *st)
 {
 	struct mana_port_context *apc = netdev_priv(ndev);
-	unsigned int num_queues = apc->num_queues;
 	struct mana_stats_rx *rx_stats;
 	struct mana_stats_tx *tx_stats;
+	unsigned int num_queues;
 	unsigned int start;
 	u64 packets, bytes;
 	int q;
 
-	if (!apc->port_is_up)
-		return;
+	/* Report even while down; dev_get_stats() zeroes its output. */
+	num_queues = apc->max_queues;
 
 	netdev_stats_to_stats64(st, &ndev->stats);
 
@@ -615,7 +615,18 @@ static void mana_get_stats64(struct net_device *ndev,
 	st->rx_missed_errors = apc->ac->hc_stats.hc_rx_discards_no_wqe;
 
 	for (q = 0; q < num_queues; q++) {
-		rx_stats = &apc->rxqs[q]->stats;
+		rx_stats = &apc->rxq_stats[q];
+
+		do {
+			start = u64_stats_fetch_begin(&rx_stats->syncp);
+			packets = rx_stats->packets;
+			bytes = rx_stats->bytes;
+		} while (u64_stats_fetch_retry(&rx_stats->syncp, start));
+
+		st->rx_packets += packets;
+		st->rx_bytes += bytes;
+
+		rx_stats = &apc->rxq_stats_ret[q];
 
 		do {
 			start = u64_stats_fetch_begin(&rx_stats->syncp);
@@ -628,7 +639,7 @@ static void mana_get_stats64(struct net_device *ndev,
 	}
 
 	for (q = 0; q < num_queues; q++) {
-		tx_stats = &apc->tx_qp[q]->txq.stats;
+		tx_stats = &apc->txq_stats[q];
 
 		do {
 			start = u64_stats_fetch_begin(&tx_stats->syncp);
@@ -1036,6 +1047,53 @@ static void mana_cleanup_port_context(struct mana_port_context *apc)
 	apc->rxqs = NULL;
 }
 
+/* Port lifetime preserves counters across queue replacement. */
+static int mana_alloc_queue_stats(struct mana_port_context *apc)
+{
+	unsigned int i;
+
+	apc->rxq_stats = kcalloc(apc->max_queues, sizeof(*apc->rxq_stats),
+				 GFP_KERNEL);
+	if (!apc->rxq_stats)
+		return -ENOMEM;
+
+	apc->rxq_stats_ret = kcalloc(apc->max_queues,
+				     sizeof(*apc->rxq_stats_ret), GFP_KERNEL);
+	if (!apc->rxq_stats_ret)
+		goto free_rxq_stats;
+
+	apc->txq_stats = kcalloc(apc->max_queues, sizeof(*apc->txq_stats),
+				 GFP_KERNEL);
+	if (!apc->txq_stats)
+		goto free_rxq_stats_ret;
+
+	for (i = 0; i < apc->max_queues; i++) {
+		u64_stats_init(&apc->rxq_stats[i].syncp);
+		u64_stats_init(&apc->rxq_stats_ret[i].syncp);
+		u64_stats_init(&apc->txq_stats[i].syncp);
+	}
+
+	return 0;
+
+free_rxq_stats_ret:
+	kfree(apc->rxq_stats_ret);
+	apc->rxq_stats_ret = NULL;
+free_rxq_stats:
+	kfree(apc->rxq_stats);
+	apc->rxq_stats = NULL;
+	return -ENOMEM;
+}
+
+static void mana_free_queue_stats(struct mana_port_context *apc)
+{
+	kfree(apc->rxq_stats);
+	apc->rxq_stats = NULL;
+	kfree(apc->rxq_stats_ret);
+	apc->rxq_stats_ret = NULL;
+	kfree(apc->txq_stats);
+	apc->txq_stats = NULL;
+}
+
 static void mana_cleanup_indir_table(struct mana_port_context *apc)
 {
 	apc->indir_table_sz = 0;
@@ -2141,7 +2199,7 @@ static void mana_rx_skb(void *buf_va, bool from_pool,
 			struct mana_rxcomp_oob *cqe, struct mana_rxq *rxq,
 			u32 pkt_len, u32 pkt_hash)
 {
-	struct mana_stats_rx *rx_stats = &rxq->stats;
+	struct mana_stats_rx *rx_stats = rxq->stats;
 	struct net_device *ndev = rxq->ndev;
 	u16 rxq_idx = rxq->rxq_idx;
 	struct napi_struct *napi;
@@ -2374,6 +2432,7 @@ static void mana_process_rx_cqe(struct mana_rxq *rxq, struct mana_cq *cq,
 	struct net_device *ndev = rxq->ndev;
 	struct mana_recv_buf_oob *rxbuf_oob;
 	struct mana_port_context *apc;
+	struct mana_stats_rx *rx_stats;
 	struct device *dev = gc->dev;
 	bool coalesced_8 = false;
 	bool coalesced = false;
@@ -2455,13 +2514,15 @@ static void mana_process_rx_cqe(struct mana_rxq *rxq, struct mana_cq *cq,
 	 * Coalesced CQEs have at least 2 packets, so index is pkt_i - 2.
 	 */
 	if (pkt_i > 1) {
-		u64_stats_update_begin(&rxq->stats.syncp);
-		rxq->stats.coalesced_cqe[pkt_i - 2]++;
-		u64_stats_update_end(&rxq->stats.syncp);
+		rx_stats = rxq->stats;
+		u64_stats_update_begin(&rx_stats->syncp);
+		rx_stats->coalesced_cqe[pkt_i - 2]++;
+		u64_stats_update_end(&rx_stats->syncp);
 	} else if (!pkt_i && !pktlen) {
-		u64_stats_update_begin(&rxq->stats.syncp);
-		rxq->stats.pkt_len0_err++;
-		u64_stats_update_end(&rxq->stats.syncp);
+		rx_stats = rxq->stats;
+		u64_stats_update_begin(&rx_stats->syncp);
+		rx_stats->pkt_len0_err++;
+		u64_stats_update_end(&rx_stats->syncp);
 		netdev_err_once(ndev,
 				"RX pkt len=0, rq=%u, cq=%u, rxobj=0x%llx\n",
 				rxq->gdma_id, cq->gdma_id, rxq->rxobj);
@@ -2593,8 +2654,8 @@ static void mana_update_rx_dim(struct mana_cq *cq)
 	if (!smp_load_acquire(&apc->rx_dim_enabled))
 		return;
 
-	dim_update_sample(READ_ONCE(cq->dim_event_ctr), rxq->stats.packets,
-			  rxq->stats.bytes, &dim_sample);
+	dim_update_sample(READ_ONCE(cq->dim_event_ctr), rxq->stats->packets,
+			  rxq->stats->bytes, &dim_sample);
 	net_dim(&cq->dim, &dim_sample);
 }
 
@@ -2824,7 +2885,7 @@ static int mana_create_txq(struct mana_port_context *apc,
 		/* Create SQ */
 		txq = &apc->tx_qp[i]->txq;
 
-		u64_stats_init(&txq->stats.syncp);
+		txq->stats = &apc->txq_stats[i];
 		txq->ndev = net;
 		txq->net_txq = netdev_get_tx_queue(net, i);
 		txq->vp_offset = apc->tx_vp_offset;
@@ -3140,6 +3201,7 @@ static struct mana_rxq *mana_create_rxq(struct mana_port_context *apc,
 		return ERR_PTR(-ENOMEM);
 
 	rxq->ndev = ndev;
+	rxq->stats = &apc->rxq_stats[rxq_idx];
 	rxq->num_rx_buf = apc->rx_queue_size;
 	rxq->rxq_idx = rxq_idx;
 	rxq->rxobj = INVALID_MANA_HANDLE;
@@ -3290,8 +3352,6 @@ static int mana_add_rx_queues(struct mana_port_context *apc,
 			goto out;
 		}
 
-		u64_stats_init(&rxq->stats.syncp);
-
 		apc->rxqs[i] = rxq;
 
 		mana_create_rxq_debugfs(apc, i);
@@ -4091,6 +4151,10 @@ static int mana_probe_port(struct mana_context *ac, int port_idx,
 		apc->tx_dim_enabled = MANA_ADAPTIVE_TX_DEF;
 	}
 
+	err = mana_alloc_queue_stats(apc);
+	if (err)
+		goto free_net;
+
 	mutex_init(&apc->vport_mutex);
 	apc->vport_use_count = 0;
 
@@ -4113,7 +4177,7 @@ static int mana_probe_port(struct mana_context *ac, int port_idx,
 
 	err = mana_init_port(ndev);
 	if (err)
-		goto free_net;
+		goto free_stats;
 
 	err = mana_rss_table_alloc(apc);
 	if (err)
@@ -4150,6 +4214,8 @@ static int mana_probe_port(struct mana_context *ac, int port_idx,
 	mana_cleanup_indir_table(apc);
 reset_apc:
 	mana_cleanup_port_context(apc);
+free_stats:
+	mana_free_queue_stats(apc);
 free_net:
 	*ndev_storage = NULL;
 	netdev_err(ndev, "Failed to probe vPort %d: %d\n", port_idx, err);
@@ -4491,6 +4557,7 @@ void mana_remove(struct gdma_dev *gd, bool suspending)
 
 		unregister_netdevice(ndev);
 		mana_cleanup_indir_table(apc);
+		mana_free_queue_stats(apc);
 
 		/* Remove the port from reset walks before freeing its netdev.
 		 */
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index ece7ff9cc409..f063462cd549 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -242,6 +242,12 @@ static void mana_get_ethtool_stats(struct net_device *ndev,
 	u64 xdp_tx;
 	u64 pkt_len0_err;
 	u64 coalesced_cqe[MANA_CQE_COAL_PKTS_8 - 1];
+	u64 ret_coalesced_cqe[MANA_CQE_COAL_PKTS_8 - 1];
+	u64 ret_packets, ret_bytes;
+	u64 ret_xdp_redirect;
+	u64 ret_pkt_len0_err;
+	u64 ret_xdp_drop;
+	u64 ret_xdp_tx;
 	u64 tso_packets;
 	u64 tso_bytes;
 	u64 tso_inner_packets;
@@ -252,14 +258,11 @@ static void mana_get_ethtool_stats(struct net_device *ndev,
 	u64 mana_map_err;
 	int q, i = 0, j;
 
-	if (!apc->port_is_up)
-		return;
-
-	/* We call this mana function to get the phy stats from GDMA and includes
-	 * aggregate tx/rx drop counters, Per-TC(Traffic Channel) tx/rx and pause
-	 * counters.
+	/* Counters outlive the queues, but suspend can destroy the HW channel
+	 * while the netdev remains registered. Gate only the PHY query.
 	 */
-	mana_query_phy_stats(apc);
+	if (apc->port_is_up)
+		mana_query_phy_stats(apc);
 
 	for (q = 0; q < ARRAY_SIZE(mana_eth_stats); q++)
 		data[i++] = *(u64 *)(eth_stats + mana_eth_stats[q].offset);
@@ -271,7 +274,7 @@ static void mana_get_ethtool_stats(struct net_device *ndev,
 		data[i++] = *(u64 *)(phy_stats + mana_phy_stats[q].offset);
 
 	for (q = 0; q < num_queues; q++) {
-		rx_stats = &apc->rxqs[q]->stats;
+		rx_stats = &apc->rxq_stats[q];
 
 		do {
 			start = u64_stats_fetch_begin(&rx_stats->syncp);
@@ -285,6 +288,33 @@ static void mana_get_ethtool_stats(struct net_device *ndev,
 				coalesced_cqe[j] = rx_stats->coalesced_cqe[j];
 		} while (u64_stats_fetch_retry(&rx_stats->syncp, start));
 
+		/* Snapshot separately so a retry cannot add retired counters
+		 * twice.
+		 */
+		rx_stats = &apc->rxq_stats_ret[q];
+
+		do {
+			start = u64_stats_fetch_begin(&rx_stats->syncp);
+			ret_packets = rx_stats->packets;
+			ret_bytes = rx_stats->bytes;
+			ret_xdp_drop = rx_stats->xdp_drop;
+			ret_xdp_tx = rx_stats->xdp_tx;
+			ret_xdp_redirect = rx_stats->xdp_redirect;
+			ret_pkt_len0_err = rx_stats->pkt_len0_err;
+			for (j = 0; j < MANA_CQE_COAL_PKTS_8 - 1; j++)
+				ret_coalesced_cqe[j] =
+					rx_stats->coalesced_cqe[j];
+		} while (u64_stats_fetch_retry(&rx_stats->syncp, start));
+
+		packets += ret_packets;
+		bytes += ret_bytes;
+		xdp_drop += ret_xdp_drop;
+		xdp_tx += ret_xdp_tx;
+		xdp_redirect += ret_xdp_redirect;
+		pkt_len0_err += ret_pkt_len0_err;
+		for (j = 0; j < MANA_CQE_COAL_PKTS_8 - 1; j++)
+			coalesced_cqe[j] += ret_coalesced_cqe[j];
+
 		data[i++] = packets;
 		data[i++] = bytes;
 		data[i++] = xdp_drop;
@@ -296,7 +326,7 @@ static void mana_get_ethtool_stats(struct net_device *ndev,
 	}
 
 	for (q = 0; q < num_queues; q++) {
-		tx_stats = &apc->tx_qp[q]->txq.stats;
+		tx_stats = &apc->txq_stats[q];
 
 		do {
 			start = u64_stats_fetch_begin(&tx_stats->syncp);
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index d3a79e13e343..d0cf92ac6fa8 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -102,7 +102,7 @@ struct mana_stats_rx {
 	u64 pkt_len0_err;
 	u64 coalesced_cqe[MANA_CQE_COAL_PKTS_8 - 1];
 	struct u64_stats_sync syncp;
-};
+} ____cacheline_aligned_in_smp;
 
 struct mana_stats_tx {
 	u64 packets;
@@ -117,7 +117,7 @@ struct mana_stats_tx {
 	u64 csum_partial;
 	u64 mana_map_err;
 	struct u64_stats_sync syncp;
-};
+} ____cacheline_aligned_in_smp;
 
 struct mana_txq {
 	struct gdma_queue *gdma_sq;
@@ -146,7 +146,7 @@ struct mana_txq {
 	/* Suppress completion wakeups on the replacement's netdev queue. */
 	bool retiring;
 
-	struct mana_stats_tx stats;
+	struct mana_stats_tx *stats;
 };
 
 /* skb data and frags dma mappings */
@@ -408,7 +408,7 @@ struct mana_rxq {
 
 	u32 buf_index;
 
-	struct mana_stats_rx stats;
+	struct mana_stats_rx *stats;
 
 	struct bpf_prog __rcu *bpf_prog;
 	struct xdp_rxq_info xdp_rxq;
@@ -606,6 +606,13 @@ struct mana_port_context {
 	unsigned int max_queues;
 	unsigned int num_queues;
 
+	/* Port-lifetime arrays with max_queues slots. Readers sum live and
+	 * retired RX counters.
+	 */
+	struct mana_stats_rx *rxq_stats;
+	struct mana_stats_rx *rxq_stats_ret;
+	struct mana_stats_tx *txq_stats;
+
 	unsigned int rx_queue_size;
 	unsigned int tx_queue_size;
 

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 04/13] net: mana: swap queue sets in mana_set_channels
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (2 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 03/13] net: mana: keep per-queue statistics in the port context Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 05/13] net: mana: swap queue sets in mana_set_ringparam Wei Hu
                   ` (8 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Build a replacement queue set before quiescing TX. After closing the
TX/XDP gate and draining its readers, wait up to 120 seconds for ALL
old TX queues before changing pointers, queue counts or RSS. On a
timeout, reject the replacement and resume the unchanged old config
without reprogramming RSS or resetting the function.

Drain every old TX queue, including queues later patches may retain,
so even a later failed-rollback close cannot invoke legacy FLR for
old pending TX. Retirement and unpublished-set cleanup need no FLR.
Allocation failure preserves live queues and configuration. After
publication, failure attempts rollback; failed rollback closes the
port and holds carrier down until a successful reopen.

The preceding statistics preparation removes reader dependence on
queue lifetime. Add private retiring RX counters and writer handoff
with this first live swap. Keep old RX counters retired until RSS
restoration succeeds. Common resume_old handling for drain rejection
and successful rollback restarts DIM only on actually retired RXQs,
with NAPI disabled and DIM work drained, then folds counters after
writer quiescence before reopening the old set.

Publish initialized fields before reopening TX/XDP and drain readers
of transient RSS tables after restoring old pointers. Keep RX queue
indices valid until retiring queues stop delivering them.

The temporary SQ/RQ/CQ peak is old plus new; the port EQ pool is
shared, not doubled. Later patches avoid the extra queue allocation
for channel-count changes; full per-queue rebuilds still need it.

Join queue selection to the publication gate so it cannot read
transient RSS or queue-count state while TX is stopped. Old tail
RX queues can outlive the TX-count reduction; fall back to the
current RSS mapping when a recorded RX index is no longer valid.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 .../net/ethernet/microsoft/mana/mana_bpf.c    |   7 +-
 drivers/net/ethernet/microsoft/mana/mana_en.c | 435 +++++++++++++++++-
 .../ethernet/microsoft/mana/mana_ethtool.c    |  78 +++-
 include/net/mana/mana.h                       |  34 +-
 4 files changed, 509 insertions(+), 45 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_bpf.c b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
index 1905214bec48..80950d5b62c5 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_bpf.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
@@ -59,6 +59,11 @@ int mana_xdp_xmit(struct net_device *ndev, int n, struct xdp_frame **frames,
 	if (unlikely(!apc->port_is_up))
 		return 0;
 
+	/* Pair with the smp_wmb() in mana_publish_qset() before reading queue
+	 * state.
+	 */
+	smp_rmb();
+
 	q_idx = smp_processor_id() % ndev->real_num_tx_queues;
 
 	for (i = 0; i < n; i++) {
@@ -95,7 +100,7 @@ u32 mana_run_xdp(struct net_device *ndev, struct mana_rxq *rxq,
 
 	act = bpf_prog_run_xdp(prog, xdp);
 
-	rx_stats = rxq->stats;
+	rx_stats = mana_rxq_stats(rxq);
 
 	switch (act) {
 	case XDP_PASS:
diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 02cd5f7656ed..10225c09f5b7 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -84,12 +84,25 @@ static int mana_open(struct net_device *ndev)
 		return err;
 	}
 
-	apc->port_is_up = true;
+	/* Publish the queues before opening the TX/XDP gate. */
+	smp_wmb();
+	WRITE_ONCE(apc->port_is_up, true);
 
 	/* Ensure port state updated before txq state */
 	smp_wmb();
 
 	netif_tx_wake_all_queues(ndev);
+
+	/* Undo a forced carrier-off unless a disconnect is pending behind RTNL.
+	 */
+	if (apc->carrier_forced_off) {
+		u32 ev = READ_ONCE(apc->ac->link_event);
+
+		apc->carrier_forced_off = false;
+		if (ev != HWC_DATA_HW_LINK_DISCONNECT)
+			netif_carrier_on(ndev);
+	}
+
 	netdev_dbg(ndev, "%s successful\n", __func__);
 	return 0;
 }
@@ -106,6 +119,7 @@ static int mana_close(struct net_device *ndev)
 
 static void mana_link_state_handle(struct work_struct *w)
 {
+	struct mana_port_context *apc;
 	struct mana_context *ac;
 	struct net_device *ndev;
 	u32 link_event;
@@ -131,7 +145,12 @@ static void mana_link_state_handle(struct work_struct *w)
 		if (!ndev)
 			continue;
 
+		apc = netdev_priv(ndev);
+
 		if (link_up) {
+			if (apc->carrier_forced_off)
+				continue;
+
 			netif_carrier_on(ndev);
 
 			__netdev_notify_peers(ndev);
@@ -312,8 +331,8 @@ static void mana_per_port_queue_reset_work_handler(struct work_struct *work)
 
 	rtnl_lock();
 
-	/* Block RDMA from grabbing the vport during the detach/attach
-	 * window, same as mana_set_channels().
+	/* Exclude RDMA across detach/attach; RTNL serializes channel_changing
+	 * writers.
 	 */
 	mutex_lock(&apc->vport_mutex);
 	apc->channel_changing = true;
@@ -366,6 +385,15 @@ netdev_tx_t mana_start_xmit(struct sk_buff *skb, struct net_device *ndev)
 	if (unlikely(!apc->port_is_up))
 		goto tx_drop;
 
+	/* Pair with mana_publish_qset()'s pre-gate smp_wmb(): observe queue
+	 * fields after reading port_is_up.
+	 */
+	smp_rmb();
+
+	/* Retiring RXQs may use indices beyond the live queue count. */
+	if (unlikely(txq_idx >= apc->num_queues))
+		goto tx_drop_count;
+
 	if (skb_cow_head(skb, MANA_HEADROOM))
 		goto tx_drop_count;
 
@@ -672,15 +700,25 @@ static int mana_get_tx_queue(struct net_device *ndev, struct sk_buff *skb,
 static u16 mana_select_queue(struct net_device *ndev, struct sk_buff *skb,
 			     struct net_device *sb_dev)
 {
+	struct mana_port_context *apc = netdev_priv(ndev);
+	unsigned int num_tx_queues;
 	int txq;
 
-	if (ndev->real_num_tx_queues == 1)
+	/* Queue selection also runs while TX is stopped. Observe the same
+	 * publication gate before reading the queue count or RSS table.
+	 */
+	if (!smp_load_acquire(&apc->port_is_up))
+		return 0;
+
+	num_tx_queues = READ_ONCE(ndev->real_num_tx_queues);
+	if (num_tx_queues == 1)
 		return 0;
 
 	txq = sk_tx_queue_get(skb->sk);
 
-	if (txq < 0 || skb->ooo_okay || txq >= ndev->real_num_tx_queues) {
-		if (skb_rx_queue_recorded(skb))
+	if (txq < 0 || skb->ooo_okay || txq >= num_tx_queues) {
+		if (skb_rx_queue_recorded(skb) &&
+		    skb_get_rx_queue(skb) < num_tx_queues)
 			txq = skb_get_rx_queue(skb);
 		else
 			txq = mana_get_tx_queue(ndev, skb, txq);
@@ -1094,6 +1132,58 @@ static void mana_free_queue_stats(struct mana_port_context *apc)
 	apc->txq_stats = NULL;
 }
 
+/* Fold under RTNL after drain_stats writers quiesce. Clear drain_stats to
+ * prevent double counting on rollback.
+ */
+static void mana_fold_rxq_stats(struct mana_port_context *apc,
+				struct mana_rxq *rxq)
+{
+	struct mana_stats_rx *src = &rxq->drain_stats;
+	struct mana_stats_rx *dst;
+	unsigned int i;
+
+	ASSERT_RTNL();
+
+	if (!apc->rxq_stats_ret || rxq->rxq_idx >= apc->max_queues)
+		return;
+
+	dst = &apc->rxq_stats_ret[rxq->rxq_idx];
+
+	u64_stats_update_begin(&dst->syncp);
+	dst->packets		+= src->packets;
+	dst->bytes		+= src->bytes;
+	dst->xdp_drop		+= src->xdp_drop;
+	dst->xdp_tx		+= src->xdp_tx;
+	dst->xdp_redirect	+= src->xdp_redirect;
+	dst->pkt_len0_err	+= src->pkt_len0_err;
+	for (i = 0; i < ARRAY_SIZE(dst->coalesced_cqe); i++)
+		dst->coalesced_cqe[i] += src->coalesced_cqe[i];
+	u64_stats_update_end(&dst->syncp);
+
+	src->packets		= 0;
+	src->bytes		= 0;
+	src->xdp_drop		= 0;
+	src->xdp_tx		= 0;
+	src->xdp_redirect	= 0;
+	src->pkt_len0_err	= 0;
+	for (i = 0; i < ARRAY_SIZE(src->coalesced_cqe); i++)
+		src->coalesced_cqe[i] = 0;
+}
+
+static void mana_fold_qset_rx_stats(struct mana_port_context *apc,
+				    struct mana_qset *qset)
+{
+	unsigned int q;
+
+	if (!qset->rxqs)
+		return;
+
+	for (q = 0; q < qset->num_queues; q++) {
+		if (qset->rxqs[q])
+			mana_fold_rxq_stats(apc, qset->rxqs[q]);
+	}
+}
+
 static void mana_cleanup_indir_table(struct mana_port_context *apc)
 {
 	apc->indir_table_sz = 0;
@@ -1103,6 +1193,7 @@ static void mana_cleanup_indir_table(struct mana_port_context *apc)
 
 static int mana_init_port_context(struct mana_port_context *apc)
 {
+	kfree(apc->rxqs);
 	apc->rxqs = kzalloc_objs(struct mana_rxq *, apc->num_queues);
 
 	return !apc->rxqs ? -ENOMEM : 0;
@@ -2135,6 +2226,7 @@ static void mana_poll_tx_cq(struct mana_cq *cq)
 	/* Ensure checking txq_stopped before apc->port_is_up. */
 	smp_rmb();
 
+	/* Order the stopped-state read before the retiring read. */
 	if (txq_stopped && !READ_ONCE(txq->retiring) && apc->port_is_up &&
 	    avail_space >= MAX_TX_WQE_SIZE) {
 		netif_tx_wake_queue(net_txq);
@@ -2199,7 +2291,7 @@ static void mana_rx_skb(void *buf_va, bool from_pool,
 			struct mana_rxcomp_oob *cqe, struct mana_rxq *rxq,
 			u32 pkt_len, u32 pkt_hash)
 {
-	struct mana_stats_rx *rx_stats = rxq->stats;
+	struct mana_stats_rx *rx_stats = mana_rxq_stats(rxq);
 	struct net_device *ndev = rxq->ndev;
 	u16 rxq_idx = rxq->rxq_idx;
 	struct napi_struct *napi;
@@ -2514,12 +2606,12 @@ static void mana_process_rx_cqe(struct mana_rxq *rxq, struct mana_cq *cq,
 	 * Coalesced CQEs have at least 2 packets, so index is pkt_i - 2.
 	 */
 	if (pkt_i > 1) {
-		rx_stats = rxq->stats;
+		rx_stats = mana_rxq_stats(rxq);
 		u64_stats_update_begin(&rx_stats->syncp);
 		rx_stats->coalesced_cqe[pkt_i - 2]++;
 		u64_stats_update_end(&rx_stats->syncp);
 	} else if (!pkt_i && !pktlen) {
-		rx_stats = rxq->stats;
+		rx_stats = mana_rxq_stats(rxq);
 		u64_stats_update_begin(&rx_stats->syncp);
 		rx_stats->pkt_len0_err++;
 		u64_stats_update_end(&rx_stats->syncp);
@@ -2604,9 +2696,9 @@ static void mana_tx_dim_work(struct work_struct *work)
 	dim->state = DIM_START_MEASURE;
 }
 
-/* The caller must update apc->rx/tx_dim_enabled before disabling and
- * after enabling. And synchronize_net() before draining the DIM work,
- * so that NAPI cannot observe a stale flag.
+/* The caller must exclude NAPI, either by disabling it or by clearing the
+ * per-port DIM flag and calling synchronize_net(). When using the flag,
+ * publish enable only after reinitializing DIM.
  */
 void mana_dim_change(struct mana_cq *cq, bool enable)
 {
@@ -2654,6 +2746,10 @@ static void mana_update_rx_dim(struct mana_cq *cq)
 	if (!smp_load_acquire(&apc->rx_dim_enabled))
 		return;
 
+	/* Skip retiring RXQs; DIM reads shared per-index counters. */
+	if (READ_ONCE(rxq->retiring))
+		return;
+
 	dim_update_sample(READ_ONCE(cq->dim_event_ctr), rxq->stats->packets,
 			  rxq->stats->bytes, &dim_sample);
 	net_dim(&cq->dim, &dim_sample);
@@ -3012,6 +3108,9 @@ static void mana_destroy_rxq(struct mana_port_context *apc,
 		netif_napi_del_locked(napi);
 	}
 
+	/* NAPI is quiesced, so drain_stats has no remaining writer. */
+	mana_fold_rxq_stats(apc, rxq);
+
 	if (xdp_rxq_info_is_reg(&rxq->xdp_rxq))
 		xdp_rxq_info_unreg(&rxq->xdp_rxq);
 
@@ -3202,6 +3301,7 @@ static struct mana_rxq *mana_create_rxq(struct mana_port_context *apc,
 
 	rxq->ndev = ndev;
 	rxq->stats = &apc->rxq_stats[rxq_idx];
+	u64_stats_init(&rxq->drain_stats.syncp);
 	rxq->num_rx_buf = apc->rx_queue_size;
 	rxq->rxq_idx = rxq_idx;
 	rxq->rxobj = INVALID_MANA_HANDLE;
@@ -3808,7 +3908,9 @@ int mana_attach(struct net_device *ndev)
 		}
 	}
 
-	apc->port_is_up = apc->port_st_save;
+	/* Publish restored queues before opening the TX/XDP gate. */
+	smp_wmb();
+	WRITE_ONCE(apc->port_is_up, apc->port_st_save);
 
 	/* Ensure port state updated before txq state */
 	smp_wmb();
@@ -4025,9 +4127,305 @@ int mana_alloc_qset(struct mana_port_context *apc,
 	return err;
 }
 
+/* Destroy caller-owned CQs before closing this dead-end port: closing also
+ * frees the shared EQ pool. Requires RTNL.
+ */
+void mana_publish_close_if_needed(struct mana_port_context *apc)
+{
+	ASSERT_RTNL();
+
+	if (!apc->publish_dead_end)
+		return;
+
+	apc->publish_dead_end = false;
+
+	if (mana_dealloc_queues(apc->ndev))
+		netdev_err(apc->ndev,
+			   "failed to close the port after a failed rollback\n");
+}
+
+/* Carried-over queues may still have full rings. */
+static void mana_start_txqs(struct mana_port_context *apc)
+{
+	struct net_device *ndev = apc->ndev;
+	unsigned int i;
+
+	if (!apc->tx_qp)
+		return;
+
+	/* Order port_is_up=true before ring reads to avoid a missed wakeup.
+	 * Pair with mana_poll_tx_cq()'s full barrier after its tail update.
+	 */
+	smp_mb();
+
+	for (i = 0; i < apc->num_queues; i++) {
+		if (!apc->tx_qp[i])
+			continue;
+
+		if (mana_can_tx(apc->tx_qp[i]->txq.gdma_sq))
+			netif_tx_wake_queue(netdev_get_tx_queue(ndev, i));
+	}
+
+	/* The watchdog scans allocated queues, including the inactive tail.
+	 * Clear its driver stop bits without scheduling inactive qdiscs.
+	 */
+	for (; i < ndev->num_tx_queues; i++)
+		netif_tx_start_queue(netdev_get_tx_queue(ndev, i));
+}
+
+/* Retiring completions must not wake replacement queues. Mark the leaving set
+ * before unmarking the incoming set.
+ */
+static void mana_qset_set_retiring(struct mana_qset *qset,
+				   const struct mana_qset *keep, bool retiring)
+{
+	unsigned int q;
+
+	for (q = 0; q < qset->num_queues; q++) {
+		if (qset->tx_qp && qset->tx_qp[q])
+			WRITE_ONCE(qset->tx_qp[q]->txq.retiring, retiring);
+
+		if (!qset->rxqs || !qset->rxqs[q])
+			continue;
+
+		/* Carried RXQs remain the sole poll writers of shared slots. */
+		if (retiring && keep && q < keep->num_queues &&
+		    keep->rxqs && keep->rxqs[q] == qset->rxqs[q])
+			continue;
+
+		/* Switch to drain_stats; hand off shared slots after a grace
+		 * period.
+		 */
+		WRITE_ONCE(qset->rxqs[q]->retiring, retiring);
+	}
+}
+
+static void mana_qset_restart_rx_dim(struct mana_qset *qset)
+{
+	struct mana_rxq *rxq;
+	struct mana_cq *cq;
+	unsigned int q;
+
+	if (!qset->rxqs)
+		return;
+
+	for (q = 0; q < qset->num_queues; q++) {
+		rxq = qset->rxqs[q];
+		if (!rxq || !READ_ONCE(rxq->retiring))
+			continue;
+
+		cq = &rxq->rx_cq;
+		napi_disable_locked(&cq->napi);
+		mana_dim_change(cq, false);
+		mana_dim_change(cq, true);
+		napi_enable_locked(&cq->napi);
+
+		/* An event may have arrived while NAPI was disabled. */
+		napi_schedule(&cq->napi);
+	}
+}
+
+/* Leave TX stopped and request RX disable; steering may be unrecoverable. */
+static void mana_publish_give_up(struct mana_port_context *apc)
+{
+	int err;
+
+	apc->rss_state = TRI_STATE_FALSE;
+
+	err = mana_disable_vport_rx(apc);
+	if (err && mana_en_need_log(apc, err))
+		netdev_err(apc->ndev, "failed to disable vPort RX: %d\n", err);
+
+	apc->carrier_forced_off = true;
+	netif_carrier_off(apc->ndev);
+	apc->publish_dead_end = true;
+}
+
+/* A replacement must not reset the function and invalidate its own queues. */
+static int mana_wait_qset_txqs(struct mana_port_context *apc)
+{
+	unsigned long timeout = jiffies + 120 * HZ;
+	struct mana_txq *txq;
+	unsigned int i;
+
+	if (!apc->tx_qp)
+		return 0;
+
+	for (i = 0; i < apc->num_queues; i++) {
+		if (!apc->tx_qp[i])
+			continue;
+
+		txq = &apc->tx_qp[i]->txq;
+		while (atomic_read(&txq->pending_sends)) {
+			if (time_after_eq(jiffies, timeout)) {
+				netdev_err(apc->ndev,
+					   "timed out draining TX queue %u for replacement\n",
+					   txq->gdma_txq_id);
+				return -ETIMEDOUT;
+			}
+
+			usleep_range(1000, 2000);
+		}
+	}
+
+	return 0;
+}
+
+/* Keep the RX count high until retiring RQs stop delivering their indices. */
+static int mana_raise_real_num_rx(struct net_device *ndev, unsigned int count)
+{
+	if (count <= ndev->real_num_rx_queues)
+		return 0;
+
+	return netif_set_real_num_rx_queues(ndev, count);
+}
+
+/* Publish under RTNL with TX gated. An error restores old pointers, not
+ * necessarily service. Free only owned queues.
+ */
+int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
+		      struct mana_qset *out_old)
+{
+	struct net_device *ndev = apc->ndev;
+	int err;
+
+	ASSERT_RTNL();
+
+	/* Close the XDP gate before stopping TX queues. Pair with
+	 * mana_poll_tx_cq()'s smp_rmb() to prevent mid-swap wakeups.
+	 */
+	WRITE_ONCE(apc->port_is_up, false);
+
+	/* Ensure port state updated before txq state */
+	smp_wmb();
+
+	netif_tx_disable(ndev);
+
+	mana_qset_snapshot(apc, out_old);
+
+	/* Mark before the grace period so old completions cannot wake the
+	 * replacement's stopped queue.
+	 */
+	mana_qset_set_retiring(out_old, newq, true);
+
+	/* Drain TX/XDP readers past the gate and polls missing retiring. */
+	synchronize_net();
+
+	err = mana_wait_qset_txqs(apc);
+	if (err)
+		goto resume_old;
+
+	mana_qset_set_retiring(newq, NULL, false);
+
+	mana_qset_install(apc, newq);
+	apc->rss_state = apc->num_queues > 1 ? TRI_STATE_TRUE : TRI_STATE_FALSE;
+
+	err = netif_set_real_num_tx_queues(ndev, apc->num_queues);
+	if (err)
+		goto rollback;
+
+	err = mana_raise_real_num_rx(ndev, apc->num_queues);
+	if (err)
+		goto rollback;
+
+	/* Install XDP and per-RXQ references before steering reaches new
+	 * queues.
+	 */
+	mana_chn_setxdp(apc, mana_xdp_get(apc));
+
+	err = mana_config_rss(apc, TRI_STATE_TRUE, true, true);
+	if (err)
+		goto rollback;
+
+	/* Publish fields before opening the gate; pair with TX/XDP read
+	 * barriers. The post-gate full barrier cannot replace this.
+	 */
+	smp_wmb();
+
+	WRITE_ONCE(apc->port_is_up, true);
+	mana_start_txqs(apc);
+
+	return 0;
+
+rollback:
+	netdev_err(ndev, "%s failed: %d, restoring previous queue set\n",
+		   __func__, err);
+
+	mana_qset_set_retiring(newq, out_old, true);
+
+	/* Quiesce new shared-slot writers before restoring old ones. */
+	synchronize_net();
+
+	mana_qset_install(apc, out_old);
+	apc->rss_state = apc->num_queues > 1 ? TRI_STATE_TRUE : TRI_STATE_FALSE;
+
+	/* Drain network readers after restoring the old pointers, before
+	 * callers can discard the unpublished containers.
+	 */
+	synchronize_net();
+
+	if (netif_set_real_num_tx_queues(ndev, apc->num_queues) ||
+	    mana_raise_real_num_rx(ndev, apc->num_queues)) {
+		/* Inconsistent restored queue counts prohibit TX; leave the
+		 * port stopped.
+		 */
+		netdev_err(ndev, "failed to restore queue counts, closing the port\n");
+		mana_publish_give_up(apc);
+		return err;
+	}
+
+	if (mana_config_rss(apc, TRI_STATE_TRUE, true, true)) {
+		/* Do not reopen TX with mismatched steering; RX disable is
+		 * best-effort.
+		 */
+		netdev_err(ndev, "failed to restore RSS steering, closing the port\n");
+		mana_publish_give_up(apc);
+		return err;
+	}
+
+resume_old:
+	if (apc->rx_dim_enabled)
+		mana_qset_restart_rx_dim(out_old);
+	mana_qset_set_retiring(out_old, NULL, false);
+
+	/* Quiesce old drain_stats writers before folding. */
+	synchronize_net();
+	mana_fold_qset_rx_stats(apc, out_old);
+
+	/* Publish restored fields before reopening the gate, as on success. */
+	smp_wmb();
+
+	WRITE_ONCE(apc->port_is_up, true);
+	mana_start_txqs(apc);
+
+	return err;
+}
+
+/* Create missing debugfs nodes once retiring names are gone. */
+static void mana_qset_debugfs_publish(struct mana_port_context *apc)
+{
+	unsigned int i;
+
+	ASSERT_RTNL();
+
+	if (IS_ERR_OR_NULL(apc->mana_port_debugfs))
+		return;
+
+	for (i = 0; i < apc->num_queues; i++) {
+		if (apc->tx_qp && apc->tx_qp[i] &&
+		    IS_ERR_OR_NULL(apc->tx_qp[i]->mana_tx_debugfs))
+			mana_create_txq_debugfs(apc, i);
+
+		if (apc->rxqs && apc->rxqs[i] &&
+		    IS_ERR_OR_NULL(apc->rxqs[i]->mana_rx_debugfs))
+			mana_create_rxq_debugfs(apc, i);
+	}
+}
+
 /* Under RTNL, free only queues no longer shared with the installed set. */
 void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset)
 {
+	struct mana_port_context *apc = netdev_priv(scratch->ndev);
 	struct bpf_prog *retiring_prog;
 	unsigned int retiring_queues;
 
@@ -4052,7 +4450,9 @@ void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset)
 
 	mana_qset_install(scratch, qset);
 
-	/* Keep retiring RXQs' XDP programs and references until RX teardown. */
+	/* Keep retiring RXQs' XDP programs and references until RX teardown.
+	 * Read the program from the queues, not queue-set metadata.
+	 */
 	retiring_prog = mana_chn_xdp_peek(scratch);
 	retiring_queues = scratch->num_queues;
 
@@ -4072,6 +4472,13 @@ void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset)
 	scratch->rxqs = NULL;
 
 	memset(qset, 0, sizeof(*qset));
+
+	/* Retiring RQs can no longer deliver indices beyond the live queue
+	 * count.
+	 */
+	netif_set_real_num_rx_queues(apc->ndev, apc->num_queues);
+
+	mana_qset_debugfs_publish(apc);
 }
 
 int mana_detach(struct net_device *ndev, bool from_close)
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index f063462cd549..ae9a0288a7c0 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -678,52 +678,82 @@ static int mana_set_coalesce(struct net_device *ndev,
 	return 0;
 }
 
-/* mana_set_channels - change the number of queues on a port
- *
- * Returns -EBUSY if RDMA holds the vport with EQs sized to the
- * current num_queues.
- */
 static int mana_set_channels(struct net_device *ndev,
 			     struct ethtool_channels *channels)
 {
 	struct mana_port_context *apc = netdev_priv(ndev);
 	unsigned int new_count = channels->combined_count;
-	unsigned int old_count = apc->num_queues;
+	struct mana_port_context *scratch;
+	struct mana_qset newq, oldq;
 	int err;
 
-	/* Set channel_changing to block RDMA from grabbing the vport
-	 * during the detach/attach window. mana_cfg_vport() checks
-	 * this flag under vport_mutex and returns -EBUSY if set.
+	if (new_count < 1 || new_count > apc->max_queues) {
+		netdev_err(ndev, "Invalid combined_count %u (max %u)\n",
+			   new_count, apc->max_queues);
+		return -EINVAL;
+	}
+
+	if (new_count == apc->num_queues)
+		return 0;
+
+	/* Resize rxqs while down: mana_open() does not recreate the port
+	 * context. RDMA must not own the vport while num_queues changes.
 	 */
 	mutex_lock(&apc->vport_mutex);
-	if (!apc->port_is_up && apc->vport_use_count) {
+	if (!apc->port_is_up) {
+		struct mana_rxq **rxqs;
+
+		if (apc->vport_use_count) {
+			mutex_unlock(&apc->vport_mutex);
+			return -EBUSY;
+		}
+
+		rxqs = kzalloc_objs(struct mana_rxq *, new_count);
+		if (!rxqs) {
+			mutex_unlock(&apc->vport_mutex);
+			return -ENOMEM;
+		}
+
+		kfree(apc->rxqs);
+		apc->rxqs = rxqs;
+		apc->num_queues = new_count;
+		mutex_unlock(&apc->vport_mutex);
+		return 0;
+	}
+
+	/* The Ethernet port already holds a vport reference; exclude RDMA
+	 * through failure cleanup.
+	 */
+	if (apc->channel_changing) {
 		mutex_unlock(&apc->vport_mutex);
 		return -EBUSY;
 	}
 	apc->channel_changing = true;
 	mutex_unlock(&apc->vport_mutex);
 
-	err = mana_pre_alloc_rxbufs(apc, ndev->mtu, new_count);
-	if (err) {
-		netdev_err(ndev, "Insufficient memory for new allocations");
+	scratch = mana_qset_scratch_alloc(apc);
+	if (!scratch) {
+		err = -ENOMEM;
 		goto clear_flag;
 	}
 
-	err = mana_detach(ndev, false);
-	if (err) {
-		netdev_err(ndev, "mana_detach failed: %d\n", err);
-		goto out;
-	}
+	err = mana_alloc_qset(apc, scratch, new_count, apc->rx_queue_size,
+			      apc->tx_queue_size, apc->priv_flags, &newq);
+	if (err)
+		goto free_scratch;
 
-	apc->num_queues = new_count;
-	err = mana_attach(ndev);
+	err = mana_publish_qset(apc, &newq, &oldq);
 	if (err) {
-		apc->num_queues = old_count;
-		netdev_err(ndev, "mana_attach failed: %d\n", err);
+		mana_free_qset(scratch, &newq);
+		goto free_scratch;
 	}
 
-out:
-	mana_pre_dealloc_rxbufs(apc);
+	mana_free_qset(scratch, &oldq);
+
+free_scratch:
+	/* Release unpublished queues before closing their shared EQ pool. */
+	mana_publish_close_if_needed(apc);
+	mana_qset_scratch_free(scratch);
 clear_flag:
 	mutex_lock(&apc->vport_mutex);
 	apc->channel_changing = false;
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index d0cf92ac6fa8..c66f9dcab407 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -408,8 +408,17 @@ struct mana_rxq {
 
 	u32 buf_index;
 
+	/* Port-owned live slot; use mana_rxq_stats() to select the writer's
+	 * slot.
+	 */
 	struct mana_stats_rx *stats;
 
+	/* Set under RTNL before another queue takes over this index. */
+	bool retiring;
+
+	/* Folded under RTNL after drain-stat writers quiesce. */
+	struct mana_stats_rx drain_stats;
+
 	struct bpf_prog __rcu *bpf_prog;
 	struct xdp_rxq_info xdp_rxq;
 	void *xdp_save_va; /* for reusing */
@@ -606,8 +615,9 @@ struct mana_port_context {
 	unsigned int max_queues;
 	unsigned int num_queues;
 
-	/* Port-lifetime arrays with max_queues slots. Readers sum live and
-	 * retired RX counters.
+	/* Port-lifetime arrays with max_queues slots. Live RX queues write
+	 * rxq_stats[]; teardown and rollback fold drain_stats into
+	 * rxq_stats_ret[] under RTNL. Readers sum both.
 	 */
 	struct mana_stats_rx *rxq_stats;
 	struct mana_stats_rx *rxq_stats_ret;
@@ -623,12 +633,16 @@ struct mana_port_context {
 	struct mutex vport_mutex;
 	int vport_use_count;
 
-	/* Set by mana_set_channels() under vport_mutex to block RDMA
-	 * from grabbing the vport during the detach/attach window.
-	 * Checked by mana_cfg_vport() when called from the RDMA path.
-	 */
+	/* Exclude RDMA during reconfiguration; protected by vport_mutex. */
 	bool channel_changing;
 
+	/* Caller must close the port after releasing the unpublished set. */
+	bool publish_dead_end;
+
+	/* Hold carrier off after failed rollback until a successful reopen.
+	 */
+	bool carrier_forced_off;
+
 	/* Net shaper handle*/
 	struct net_shaper_handle handle;
 
@@ -702,10 +716,18 @@ int mana_detach(struct net_device *ndev, bool from_close);
 struct mana_port_context *
 mana_qset_scratch_alloc(struct mana_port_context *apc);
 void mana_qset_scratch_free(struct mana_port_context *scratch);
+static inline struct mana_stats_rx *mana_rxq_stats(struct mana_rxq *rxq)
+{
+	return READ_ONCE(rxq->retiring) ? &rxq->drain_stats : rxq->stats;
+}
+
 int mana_alloc_qset(struct mana_port_context *apc,
 		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
 		    u32 priv_flags, struct mana_qset *out);
+int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
+		      struct mana_qset *out_old);
+void mana_publish_close_if_needed(struct mana_port_context *apc);
 void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset);
 
 void mana_dim_change(struct mana_cq *cq, bool enable);

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 05/13] net: mana: swap queue sets in mana_set_ringparam
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (3 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 04/13] net: mana: swap queue sets in mana_set_channels Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 06/13] net: mana: swap queue sets in mana_set_priv_flags Wei Hu
                   ` (7 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Replace detach/attach with queue-set allocation and publication. Failed
allocation now leaves the running queues and ring sizes unchanged,
rather than risking a detached port after attach failure.

Skip requests whose rounded sizes already match. Keep RDMA excluded
through failure cleanup, which can release the vport.

This full rebuild temporarily keeps both sets of SQ/RQ/CQ objects,
RX buffers and page pools live. The port-owned EQ pool is shared,
not doubled. Device resource limits may reject this temporary peak;
allocation failure preserves the live queues and configuration.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 .../ethernet/microsoft/mana/mana_ethtool.c    | 68 ++++++++++++-------
 1 file changed, 45 insertions(+), 23 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index ae9a0288a7c0..10d672cc9e60 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -780,13 +780,11 @@ static int mana_set_ringparam(struct net_device *ndev,
 			      struct netlink_ext_ack *extack)
 {
 	struct mana_port_context *apc = netdev_priv(ndev);
+	struct mana_port_context *scratch;
+	struct mana_qset newq, oldq;
 	u32 new_tx, new_rx;
-	u32 old_tx, old_rx;
 	int err;
 
-	old_tx = apc->tx_queue_size;
-	old_rx = apc->rx_queue_size;
-
 	if (ring->tx_pending < MIN_TX_BUFFERS_PER_QUEUE) {
 		NL_SET_ERR_MSG_FMT(extack, "tx:%d less than the min:%d", ring->tx_pending,
 				   MIN_TX_BUFFERS_PER_QUEUE);
@@ -804,32 +802,56 @@ static int mana_set_ringparam(struct net_device *ndev,
 	netdev_info(ndev, "Using nearest power of 2 values for Txq:%d Rxq:%d\n",
 		    new_tx, new_rx);
 
-	/* pre-allocating new buffers to prevent failures in mana_attach() later */
-	apc->rx_queue_size = new_rx;
-	err = mana_pre_alloc_rxbufs(apc, ndev->mtu, apc->num_queues);
-	apc->rx_queue_size = old_rx;
-	if (err) {
-		netdev_err(ndev, "Insufficient memory for new allocations\n");
-		return err;
+	if (new_rx == apc->rx_queue_size && new_tx == apc->tx_queue_size)
+		return 0;
+
+	if (!apc->port_is_up) {
+		apc->rx_queue_size = new_rx;
+		apc->tx_queue_size = new_tx;
+		return 0;
 	}
 
-	err = mana_detach(ndev, false);
-	if (err) {
-		netdev_err(ndev, "mana_detach failed: %d\n", err);
-		goto out;
+	/* Exclude RDMA through failure cleanup, which may release the vport. */
+	mutex_lock(&apc->vport_mutex);
+	if (apc->channel_changing) {
+		mutex_unlock(&apc->vport_mutex);
+		return -EBUSY;
+	}
+	apc->channel_changing = true;
+	mutex_unlock(&apc->vport_mutex);
+
+	scratch = mana_qset_scratch_alloc(apc);
+	if (!scratch) {
+		err = -ENOMEM;
+		goto clear_flag;
 	}
 
-	apc->tx_queue_size = new_tx;
-	apc->rx_queue_size = new_rx;
+	err = mana_alloc_qset(apc, scratch, apc->num_queues, new_rx, new_tx,
+			      apc->priv_flags, &newq);
+	if (err) {
+		NL_SET_ERR_MSG_FMT(extack, "failed to change ring params: %d",
+				   err);
+		goto free_scratch;
+	}
 
-	err = mana_attach(ndev);
+	err = mana_publish_qset(apc, &newq, &oldq);
 	if (err) {
-		netdev_err(ndev, "mana_attach failed: %d\n", err);
-		apc->tx_queue_size = old_tx;
-		apc->rx_queue_size = old_rx;
+		NL_SET_ERR_MSG_FMT(extack, "failed to change ring params: %d",
+				   err);
+		mana_free_qset(scratch, &newq);
+		goto free_scratch;
 	}
-out:
-	mana_pre_dealloc_rxbufs(apc);
+
+	mana_free_qset(scratch, &oldq);
+
+free_scratch:
+	/* Release unpublished queues before closing their shared EQ pool. */
+	mana_publish_close_if_needed(apc);
+	mana_qset_scratch_free(scratch);
+clear_flag:
+	mutex_lock(&apc->vport_mutex);
+	apc->channel_changing = false;
+	mutex_unlock(&apc->vport_mutex);
 	return err;
 }
 

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 06/13] net: mana: swap queue sets in mana_set_priv_flags
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (4 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 05/13] net: mana: swap queue sets in mana_set_ringparam Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 07/13] net: mana: swap queue sets in mana_change_mtu Wei Hu
                   ` (6 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Rebuild queues through the swap path when USE_FULL_PAGE_RXBUF changes
the RX layout. Carry priv_flags with the queue set so allocation failure
leaves the live configuration unchanged and rollback restores the flags.

Retain shortcuts for an unchanged RX buffer flag, a down port, or a
configuration whose MTU or XDP already requires full-page RX. A failed
rollback closes the port.

This full rebuild temporarily keeps both sets of SQ/RQ/CQ objects,
RX buffers and page pools live. The port-owned EQ pool is shared,
not doubled. Device resource limits may reject this temporary peak;
allocation failure preserves the live queues and configuration.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 .../ethernet/microsoft/mana/mana_ethtool.c    | 74 +++++++++----------
 1 file changed, 36 insertions(+), 38 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index 10d672cc9e60..7a1120342208 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -881,7 +881,8 @@ static int mana_set_priv_flags(struct net_device *ndev, u32 priv_flags)
 {
 	struct mana_port_context *apc = netdev_priv(ndev);
 	u32 changed = apc->priv_flags ^ priv_flags;
-	u32 old_priv_flags = apc->priv_flags;
+	struct mana_port_context *scratch;
+	struct mana_qset newq, oldq;
 	int err = 0;
 
 	if (!changed)
@@ -891,54 +892,51 @@ static int mana_set_priv_flags(struct net_device *ndev, u32 priv_flags)
 	if (priv_flags & ~GENMASK(MANA_PRIV_FLAG_MAX - 1, 0))
 		return -EINVAL;
 
-	apc->priv_flags = priv_flags;
-
-	if (changed & BIT(MANA_PRIV_FLAG_USE_FULL_PAGE_RXBUF)) {
-		if (!apc->port_is_up)
-			return 0;
-
-		/* If XDP is attached or MTU is jumbo, single-buffer-per-page
-		 * is already forced regardless of this flag. Skip the
-		 * expensive detach/attach cycle since nothing changes.
-		 */
-		if (ndev->mtu + MANA_RXBUF_PAD > PAGE_SIZE / 2 ||
-		    mana_xdp_get(apc))
-			return 0;
+	/* No rebuild is needed if the RX buffer flag is unchanged, the port is
+	 * down, or full-page RX buffers are already required by MTU or XDP.
+	 */
+	if (!(changed & BIT(MANA_PRIV_FLAG_USE_FULL_PAGE_RXBUF)) ||
+	    !apc->port_is_up ||
+	    ndev->mtu + MANA_RXBUF_PAD > PAGE_SIZE / 2 ||
+	    mana_xdp_get(apc)) {
+		apc->priv_flags = priv_flags;
+		return 0;
+	}
 
-		/* Block RDMA from grabbing the vport during detach/attach */
-		mutex_lock(&apc->vport_mutex);
-		apc->channel_changing = true;
+	mutex_lock(&apc->vport_mutex);
+	if (apc->channel_changing) {
 		mutex_unlock(&apc->vport_mutex);
+		return -EBUSY;
+	}
+	apc->channel_changing = true;
+	mutex_unlock(&apc->vport_mutex);
 
-		err = mana_pre_alloc_rxbufs(apc, ndev->mtu, apc->num_queues);
-		if (err) {
-			netdev_err(ndev,
-				   "Insufficient memory for new allocations\n");
-			apc->priv_flags = old_priv_flags;
-			goto clear_flag;
-		}
+	scratch = mana_qset_scratch_alloc(apc);
+	if (!scratch) {
+		err = -ENOMEM;
+		goto clear_flag;
+	}
 
-		err = mana_detach(ndev, false);
-		if (err) {
-			netdev_err(ndev, "mana_detach failed: %d\n", err);
-			apc->priv_flags = old_priv_flags;
-			goto out;
-		}
+	err = mana_alloc_qset(apc, scratch, apc->num_queues, apc->rx_queue_size,
+			      apc->tx_queue_size, priv_flags, &newq);
+	if (err)
+		goto free_scratch;
 
-		err = mana_attach(ndev);
-		if (err) {
-			netdev_err(ndev, "mana_attach failed: %d\n", err);
-			apc->priv_flags = old_priv_flags;
-		}
+	err = mana_publish_qset(apc, &newq, &oldq);
+	if (err) {
+		mana_free_qset(scratch, &newq);
+		goto free_scratch;
 	}
 
-out:
-	mana_pre_dealloc_rxbufs(apc);
+	mana_free_qset(scratch, &oldq);
+
+free_scratch:
+	mana_publish_close_if_needed(apc);
+	mana_qset_scratch_free(scratch);
 clear_flag:
 	mutex_lock(&apc->vport_mutex);
 	apc->channel_changing = false;
 	mutex_unlock(&apc->vport_mutex);
-
 	return err;
 }
 

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 07/13] net: mana: swap queue sets in mana_change_mtu
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (5 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 06/13] net: mana: swap queue sets in mana_set_priv_flags Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 08/13] net: mana: swap queue sets in mana_xdp_set Wei Hu
                   ` (5 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Carry the MTU in the queue set and size replacement RX buffers from it.
Publish ndev->mtu after RSS configuration succeeds; allocation failure
leaves the live queues and advertised MTU unchanged.

This full rebuild temporarily keeps both sets of SQ/RQ/CQ objects,
RX buffers and page pools live. The port-owned EQ pool is shared,
not doubled. Device resource limits may reject this temporary peak;
allocation failure preserves the live queues and configuration.

Reserve RAW-QP admission under vport_mutex for the entire live
replacement, including unpublished queues, fail-close and scratch
cleanup. Preserve validation, no-op and down-port behavior.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 drivers/net/ethernet/microsoft/mana/mana_en.c | 66 ++++++++++++++-----
 .../ethernet/microsoft/mana/mana_ethtool.c    | 10 +--
 include/net/mana/mana.h                       |  7 +-
 3 files changed, 60 insertions(+), 23 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 10225c09f5b7..4577a473ff42 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -927,32 +927,52 @@ int mana_pre_alloc_rxbufs(struct mana_port_context *mpc, int new_mtu, int num_qu
 static int mana_change_mtu(struct net_device *ndev, int new_mtu)
 {
 	struct mana_port_context *mpc = netdev_priv(ndev);
-	unsigned int old_mtu = ndev->mtu;
+	struct mana_port_context *scratch;
+	struct mana_qset newq, oldq;
 	int err;
 
-	/* Pre-allocate buffers to prevent failure in mana_attach later */
-	err = mana_pre_alloc_rxbufs(mpc, new_mtu, mpc->num_queues);
-	if (err) {
-		netdev_err(ndev, "Insufficient memory for new MTU\n");
-		return err;
+	if (!mpc->port_is_up) {
+		mpc->configured_mtu = new_mtu;
+		WRITE_ONCE(ndev->mtu, new_mtu);
+		return 0;
 	}
 
-	err = mana_detach(ndev, false);
-	if (err) {
-		netdev_err(ndev, "mana_detach failed: %d\n", err);
-		goto out;
+	/* Fail-close may release the Ethernet vport reference. */
+	mutex_lock(&mpc->vport_mutex);
+	if (mpc->channel_changing) {
+		mutex_unlock(&mpc->vport_mutex);
+		return -EBUSY;
+	}
+	mpc->channel_changing = true;
+	mutex_unlock(&mpc->vport_mutex);
+
+	scratch = mana_qset_scratch_alloc(mpc);
+	if (!scratch) {
+		err = -ENOMEM;
+		goto clear_flag;
 	}
 
-	WRITE_ONCE(ndev->mtu, new_mtu);
+	err = mana_alloc_qset(mpc, scratch, mpc->num_queues,
+			      mpc->rx_queue_size, mpc->tx_queue_size,
+			      mpc->priv_flags, new_mtu, &newq);
+	if (err)
+		goto free_scratch;
 
-	err = mana_attach(ndev);
+	err = mana_publish_qset(mpc, &newq, &oldq);
 	if (err) {
-		netdev_err(ndev, "mana_attach failed: %d\n", err);
-		WRITE_ONCE(ndev->mtu, old_mtu);
+		mana_free_qset(scratch, &newq);
+		goto free_scratch;
 	}
 
-out:
-	mana_pre_dealloc_rxbufs(mpc);
+	mana_free_qset(scratch, &oldq);
+
+free_scratch:
+	mana_publish_close_if_needed(mpc);
+	mana_qset_scratch_free(scratch);
+clear_flag:
+	mutex_lock(&mpc->vport_mutex);
+	mpc->channel_changing = false;
+	mutex_unlock(&mpc->vport_mutex);
 	return err;
 }
 
@@ -3306,7 +3326,8 @@ static struct mana_rxq *mana_create_rxq(struct mana_port_context *apc,
 	rxq->rxq_idx = rxq_idx;
 	rxq->rxobj = INVALID_MANA_HANDLE;
 
-	mana_get_rxbuf_cfg(apc, ndev->mtu, &rxq->datasize, &rxq->alloc_size,
+	mana_get_rxbuf_cfg(apc, apc->configured_mtu, &rxq->datasize,
+			   &rxq->alloc_size,
 			   &rxq->headroom, &rxq->frag_count);
 	/* Create page pool for RX queue */
 	err = mana_create_page_pool(rxq, gc);
@@ -4017,6 +4038,7 @@ static void mana_qset_snapshot(const struct mana_port_context *ctx,
 	out->rx_queue_size	= ctx->rx_queue_size;
 	out->tx_queue_size	= ctx->tx_queue_size;
 	out->priv_flags		= ctx->priv_flags;
+	out->mtu		= ctx->configured_mtu;
 }
 
 /* Vport identity and port debugfs outlive queue sets. */
@@ -4033,6 +4055,7 @@ static void mana_qset_install(struct mana_port_context *ctx,
 	ctx->rx_queue_size	= qset->rx_queue_size;
 	ctx->tx_queue_size	= qset->tx_queue_size;
 	ctx->priv_flags		= qset->priv_flags;
+	ctx->configured_mtu	= qset->mtu;
 }
 
 /* Scratch starts without SQs/RQs and borrows the port's EQ pool. Never call
@@ -4073,7 +4096,7 @@ void mana_qset_scratch_free(struct mana_port_context *scratch)
 int mana_alloc_qset(struct mana_port_context *apc,
 		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
-		    u32 priv_flags, struct mana_qset *out)
+		    u32 priv_flags, int mtu, struct mana_qset *out)
 {
 	struct net_device *ndev = scratch->ndev;
 	int err;
@@ -4085,6 +4108,8 @@ int mana_alloc_qset(struct mana_port_context *apc,
 	scratch->tx_queue_size	= tx_queue_size;
 	scratch->priv_flags	= priv_flags;
 
+	scratch->configured_mtu	= mtu;
+
 	err = mana_init_port_context(scratch);
 	if (err)
 		goto out_err;
@@ -4337,6 +4362,8 @@ int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 	if (err)
 		goto rollback;
 
+	WRITE_ONCE(ndev->mtu, apc->configured_mtu);
+
 	/* Publish fields before opening the gate; pair with TX/XDP read
 	 * barriers. The post-gate full barrier cannot replace this.
 	 */
@@ -4392,6 +4419,8 @@ int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 	synchronize_net();
 	mana_fold_qset_rx_stats(apc, out_old);
 
+	WRITE_ONCE(ndev->mtu, apc->configured_mtu);
+
 	/* Publish restored fields before reopening the gate, as on success. */
 	smp_wmb();
 
@@ -4544,6 +4573,7 @@ static int mana_probe_port(struct mana_context *ac, int port_idx,
 	apc->port_handle = INVALID_MANA_HANDLE;
 	apc->pf_filter_handle = INVALID_MANA_HANDLE;
 	apc->port_idx = port_idx;
+	apc->configured_mtu = ndev->mtu;
 	apc->link_cfg_error = 1;
 	apc->cqe_coalescing_enable = 0;
 	apc->cqe8_coalescing_enable = 0;
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index 7a1120342208..2ad0fb4d5008 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -738,7 +738,8 @@ static int mana_set_channels(struct net_device *ndev,
 	}
 
 	err = mana_alloc_qset(apc, scratch, new_count, apc->rx_queue_size,
-			      apc->tx_queue_size, apc->priv_flags, &newq);
+			      apc->tx_queue_size, apc->priv_flags,
+			      apc->configured_mtu, &newq);
 	if (err)
 		goto free_scratch;
 
@@ -827,7 +828,7 @@ static int mana_set_ringparam(struct net_device *ndev,
 	}
 
 	err = mana_alloc_qset(apc, scratch, apc->num_queues, new_rx, new_tx,
-			      apc->priv_flags, &newq);
+			      apc->priv_flags, apc->configured_mtu, &newq);
 	if (err) {
 		NL_SET_ERR_MSG_FMT(extack, "failed to change ring params: %d",
 				   err);
@@ -917,8 +918,9 @@ static int mana_set_priv_flags(struct net_device *ndev, u32 priv_flags)
 		goto clear_flag;
 	}
 
-	err = mana_alloc_qset(apc, scratch, apc->num_queues, apc->rx_queue_size,
-			      apc->tx_queue_size, priv_flags, &newq);
+	err = mana_alloc_qset(apc, scratch, apc->num_queues,
+			      apc->rx_queue_size, apc->tx_queue_size,
+			      priv_flags, apc->configured_mtu, &newq);
 	if (err)
 		goto free_scratch;
 
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index c66f9dcab407..3faad257e97e 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -626,6 +626,10 @@ struct mana_port_context {
 	unsigned int rx_queue_size;
 	unsigned int tx_queue_size;
 
+	/* MTU used to size RX buffers, independent of ndev->mtu during a swap.
+	 */
+	int configured_mtu;
+
 	mana_handle_t port_handle;
 	mana_handle_t pf_filter_handle;
 
@@ -702,6 +706,7 @@ struct mana_qset {
 	unsigned int		tx_queue_size;
 	u32			priv_flags;
 
+	int			mtu;
 };
 
 netdev_tx_t mana_start_xmit(struct sk_buff *skb, struct net_device *ndev);
@@ -724,7 +729,7 @@ static inline struct mana_stats_rx *mana_rxq_stats(struct mana_rxq *rxq)
 int mana_alloc_qset(struct mana_port_context *apc,
 		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
-		    u32 priv_flags, struct mana_qset *out);
+		    u32 priv_flags, int mtu, struct mana_qset *out);
 int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 		      struct mana_qset *out_old);
 void mana_publish_close_if_needed(struct mana_port_context *apc);

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 08/13] net: mana: swap queue sets in mana_xdp_set
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (6 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 07/13] net: mana: swap queue sets in mana_change_mtu Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 09/13] net: mana: do not bail out of mana_detach on dealloc failure Wei Hu
                   ` (4 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Carry the XDP program with the queue set and install its per-queue
references before redirecting RSS. This keeps the RX buffer layout and
program together during publication and rollback.

Do not replace the live program during allocation. This also avoids the
pre-existing failed-preallocation stale-pointer bug; its standalone net
fix is linked below.

This full rebuild temporarily keeps both sets of SQ/RQ/CQ objects,
RX buffers and page pools live. The port-owned EQ pool is shared,
not doubled. Device resource limits may reject this temporary peak;
allocation failure preserves the live queues and configuration.

Reserve RAW-QP admission under vport_mutex for the entire live
replacement, including unpublished queues, fail-close and scratch
cleanup. Preserve validation, no-op and down-port behavior.

Link: https://lore.kernel.org/all/20260904202640.3900685-1-longli@microsoft.com/

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 .../net/ethernet/microsoft/mana/mana_bpf.c    | 75 +++++++++++--------
 drivers/net/ethernet/microsoft/mana/mana_en.c | 12 ++-
 .../ethernet/microsoft/mana/mana_ethtool.c    | 11 +--
 include/net/mana/mana.h                       |  5 +-
 4 files changed, 60 insertions(+), 43 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_bpf.c b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
index 80950d5b62c5..0ad3a0fa6747 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_bpf.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
@@ -177,6 +177,8 @@ static int mana_xdp_set(struct net_device *ndev, struct bpf_prog *prog,
 			struct netlink_ext_ack *extack)
 {
 	struct mana_port_context *apc = netdev_priv(ndev);
+	struct mana_port_context *scratch;
+	struct mana_qset newq, oldq;
 	struct bpf_prog *old_prog;
 	struct gdma_context *gc;
 	int err;
@@ -196,47 +198,59 @@ static int mana_xdp_set(struct net_device *ndev, struct bpf_prog *prog,
 		return -EOPNOTSUPP;
 	}
 
-	/* One refcnt of the prog is hold by the caller already, so
-	 * don't increase refcnt for this one.
-	 */
-	apc->bpf_prog = prog;
-
 	if (apc->port_is_up) {
-		/* Re-create rxq's after xdp prog was loaded or unloaded.
-		 * Ex: re create rxq's to switch from full pages to smaller
-		 * size page fragments when xdp prog is unloaded and
-		 * vice-versa.
-		 */
+		/* Fail-close may release the Ethernet vport reference. */
+		mutex_lock(&apc->vport_mutex);
+		if (apc->channel_changing) {
+			mutex_unlock(&apc->vport_mutex);
+			return -EBUSY;
+		}
+		apc->channel_changing = true;
+		mutex_unlock(&apc->vport_mutex);
 
-		/* Pre-allocate buffers to prevent failure in mana_attach */
-		err = mana_pre_alloc_rxbufs(apc, ndev->mtu, apc->num_queues);
-		if (err) {
+		scratch = mana_qset_scratch_alloc(apc);
+		if (!scratch) {
 			NL_SET_ERR_MSG_MOD(extack,
-					   "XDP: Insufficient memory for tx/rx re-config");
-			apc->bpf_prog = old_prog;
-			return err;
+					   "XDP: Insufficient memory for re-config");
+			err = -ENOMEM;
+			goto clear_flag;
 		}
 
-		err = mana_detach(ndev, false);
+		err = mana_alloc_qset(apc, scratch, apc->num_queues,
+				      apc->rx_queue_size, apc->tx_queue_size,
+				      apc->priv_flags, apc->configured_mtu,
+				      prog, &newq);
 		if (err) {
-			netdev_err(ndev,
-				   "mana_detach failed at xdp set: %d\n", err);
 			NL_SET_ERR_MSG_MOD(extack,
-					   "XDP: Re-config failed at detach");
-			goto err_dealloc_rxbuffs;
+					   "XDP: Re-config failed at alloc");
+			goto free_scratch;
 		}
 
-		err = mana_attach(ndev);
+		err = mana_publish_qset(apc, &newq, &oldq);
 		if (err) {
-			netdev_err(ndev,
-				   "mana_attach failed at xdp set: %d\n", err);
 			NL_SET_ERR_MSG_MOD(extack,
-					   "XDP: Re-config failed at attach");
-			goto err_dealloc_rxbuffs;
+					   "XDP: Re-config failed at publish");
+			mana_free_qset(scratch, &newq);
+			/* Free the queues before closing their shared EQ pool.
+			 */
+			mana_publish_close_if_needed(apc);
+			goto free_scratch;
 		}
 
-		mana_chn_setxdp(apc, prog);
-		mana_pre_dealloc_rxbufs(apc);
+		mana_free_qset(scratch, &oldq);
+free_scratch:
+		mana_qset_scratch_free(scratch);
+clear_flag:
+		mutex_lock(&apc->vport_mutex);
+		apc->channel_changing = false;
+		mutex_unlock(&apc->vport_mutex);
+		if (err)
+			return err;
+	} else {
+		/* Use the caller's program reference; mana_open() installs it
+		 * on queues.
+		 */
+		apc->bpf_prog = prog;
 	}
 
 	if (old_prog)
@@ -249,11 +263,6 @@ static int mana_xdp_set(struct net_device *ndev, struct bpf_prog *prog,
 		ndev->max_mtu = gc->adapter_mtu - ETH_HLEN;
 
 	return 0;
-
-err_dealloc_rxbuffs:
-	apc->bpf_prog = old_prog;
-	mana_pre_dealloc_rxbufs(apc);
-	return err;
 }
 
 int mana_bpf(struct net_device *ndev, struct netdev_bpf *bpf)
diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 4577a473ff42..51f9bfc64c6b 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -952,9 +952,9 @@ static int mana_change_mtu(struct net_device *ndev, int new_mtu)
 		goto clear_flag;
 	}
 
-	err = mana_alloc_qset(mpc, scratch, mpc->num_queues,
-			      mpc->rx_queue_size, mpc->tx_queue_size,
-			      mpc->priv_flags, new_mtu, &newq);
+	err = mana_alloc_qset(mpc, scratch, mpc->num_queues, mpc->rx_queue_size,
+			      mpc->tx_queue_size, mpc->priv_flags, new_mtu,
+			      mpc->bpf_prog, &newq);
 	if (err)
 		goto free_scratch;
 
@@ -4039,6 +4039,7 @@ static void mana_qset_snapshot(const struct mana_port_context *ctx,
 	out->tx_queue_size	= ctx->tx_queue_size;
 	out->priv_flags		= ctx->priv_flags;
 	out->mtu		= ctx->configured_mtu;
+	out->bpf_prog		= ctx->bpf_prog;
 }
 
 /* Vport identity and port debugfs outlive queue sets. */
@@ -4056,6 +4057,7 @@ static void mana_qset_install(struct mana_port_context *ctx,
 	ctx->tx_queue_size	= qset->tx_queue_size;
 	ctx->priv_flags		= qset->priv_flags;
 	ctx->configured_mtu	= qset->mtu;
+	ctx->bpf_prog		= qset->bpf_prog;
 }
 
 /* Scratch starts without SQs/RQs and borrows the port's EQ pool. Never call
@@ -4096,7 +4098,8 @@ void mana_qset_scratch_free(struct mana_port_context *scratch)
 int mana_alloc_qset(struct mana_port_context *apc,
 		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
-		    u32 priv_flags, int mtu, struct mana_qset *out)
+		    u32 priv_flags, int mtu, struct bpf_prog *bpf_prog,
+		    struct mana_qset *out)
 {
 	struct net_device *ndev = scratch->ndev;
 	int err;
@@ -4109,6 +4112,7 @@ int mana_alloc_qset(struct mana_port_context *apc,
 	scratch->priv_flags	= priv_flags;
 
 	scratch->configured_mtu	= mtu;
+	scratch->bpf_prog	= bpf_prog;
 
 	err = mana_init_port_context(scratch);
 	if (err)
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index 2ad0fb4d5008..a438bc6097d6 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -739,7 +739,7 @@ static int mana_set_channels(struct net_device *ndev,
 
 	err = mana_alloc_qset(apc, scratch, new_count, apc->rx_queue_size,
 			      apc->tx_queue_size, apc->priv_flags,
-			      apc->configured_mtu, &newq);
+			      apc->configured_mtu, apc->bpf_prog, &newq);
 	if (err)
 		goto free_scratch;
 
@@ -828,7 +828,8 @@ static int mana_set_ringparam(struct net_device *ndev,
 	}
 
 	err = mana_alloc_qset(apc, scratch, apc->num_queues, new_rx, new_tx,
-			      apc->priv_flags, apc->configured_mtu, &newq);
+			      apc->priv_flags, apc->configured_mtu,
+			      apc->bpf_prog, &newq);
 	if (err) {
 		NL_SET_ERR_MSG_FMT(extack, "failed to change ring params: %d",
 				   err);
@@ -918,9 +919,9 @@ static int mana_set_priv_flags(struct net_device *ndev, u32 priv_flags)
 		goto clear_flag;
 	}
 
-	err = mana_alloc_qset(apc, scratch, apc->num_queues,
-			      apc->rx_queue_size, apc->tx_queue_size,
-			      priv_flags, apc->configured_mtu, &newq);
+	err = mana_alloc_qset(apc, scratch, apc->num_queues, apc->rx_queue_size,
+			      apc->tx_queue_size, priv_flags,
+			      apc->configured_mtu, apc->bpf_prog, &newq);
 	if (err)
 		goto free_scratch;
 
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index 3faad257e97e..5f9461bc16e9 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -707,6 +707,8 @@ struct mana_qset {
 	u32			priv_flags;
 
 	int			mtu;
+	struct bpf_prog		*bpf_prog;
+
 };
 
 netdev_tx_t mana_start_xmit(struct sk_buff *skb, struct net_device *ndev);
@@ -729,7 +731,8 @@ static inline struct mana_stats_rx *mana_rxq_stats(struct mana_rxq *rxq)
 int mana_alloc_qset(struct mana_port_context *apc,
 		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
-		    u32 priv_flags, int mtu, struct mana_qset *out);
+		    u32 priv_flags, int mtu, struct bpf_prog *bpf_prog,
+		    struct mana_qset *out);
 int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 		      struct mana_qset *out_old);
 void mana_publish_close_if_needed(struct mana_port_context *apc);

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 09/13] net: mana: do not bail out of mana_detach on dealloc failure
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (7 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 08/13] net: mana: swap queue sets in mana_xdp_set Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 10/13] net: mana: release EQs left idle by a channel-count reduction Wei Hu
                   ` (3 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Dipayaan Roy <dipayanroy@linux.microsoft.com>

Remove the early return after mana_dealloc_queues() so detach continues
its device and port-context cleanup.

The return is currently unreachable: mana_dealloc_queues() only rejects
an up port, and mana_detach() clears port_is_up before calling it. This
is a robustness cleanup, not a fix for a reachable reset failure.

Signed-off-by: Dipayaan Roy <dipayanroy@linux.microsoft.com>
Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 drivers/net/ethernet/microsoft/mana/mana_en.c | 4 +---
 1 file changed, 1 insertion(+), 3 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 51f9bfc64c6b..9297ce15efce 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -4537,10 +4537,8 @@ int mana_detach(struct net_device *ndev, bool from_close)
 
 	if (apc->port_st_save) {
 		err = mana_dealloc_queues(ndev);
-		if (err) {
+		if (err)
 			netdev_err(ndev, "%s failed to deallocate queues: %d\n", __func__, err);
-			return err;
-		}
 	}
 
 	if (!from_close) {

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 10/13] net: mana: release EQs left idle by a channel-count reduction
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (8 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 09/13] net: mana: do not bail out of mana_detach on dealloc failure Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 11/13] net: mana: keep a user-configured RSS table across a queue rebuild Wei Hu
                   ` (2 subsequent siblings)
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Release EQs above the live queue count after retiring queues are freed
or replacement allocation fails. All CQs using those EQs must be gone.
Return their vector allocations to the pool; IRQ registrations remain.

Store each EQ's debugfs dentry in apc->eqs[] rather than a stack copy so
shrinking can remove individual EQ directories.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 drivers/net/ethernet/microsoft/mana/mana_en.c | 49 ++++++++++++++++---
 .../ethernet/microsoft/mana/mana_ethtool.c    |  1 -
 2 files changed, 41 insertions(+), 9 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 9297ce15efce..86a09828b5f4 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -1910,6 +1910,8 @@ void mana_destroy_eq(struct mana_port_context *apc)
 		msi = eq->eq.msix_index;
 		mana_gd_destroy_queue(gc, eq);
 		mana_gd_put_gic(gc, !gc->msi_sharing, msi);
+		apc->eqs[i].eq = NULL;
+		apc->eqs[i].mana_eq_debugfs = NULL;
 	}
 
 	kfree(apc->eqs);
@@ -1920,15 +1922,16 @@ EXPORT_SYMBOL_NS(mana_destroy_eq, "NET_MANA");
 
 static void mana_create_eq_debugfs(struct mana_port_context *apc, int i)
 {
-	struct mana_eq eq = apc->eqs[i];
+	struct mana_eq *eq = &apc->eqs[i];
 	char eqnum[32];
 
 	sprintf(eqnum, "eq%d", i);
-	eq.mana_eq_debugfs = debugfs_create_dir(eqnum, apc->mana_eqs_debugfs);
-	debugfs_create_u32("head", 0400, eq.mana_eq_debugfs, &eq.eq->head);
-	debugfs_create_u32("tail", 0400, eq.mana_eq_debugfs, &eq.eq->tail);
-	debugfs_create_u32("irq", 0400, eq.mana_eq_debugfs, &eq.eq->eq.irq);
-	debugfs_create_file("eq_dump", 0400, eq.mana_eq_debugfs, eq.eq, &mana_dbg_q_fops);
+	eq->mana_eq_debugfs = debugfs_create_dir(eqnum, apc->mana_eqs_debugfs);
+	debugfs_create_u32("head", 0400, eq->mana_eq_debugfs, &eq->eq->head);
+	debugfs_create_u32("tail", 0400, eq->mana_eq_debugfs, &eq->eq->tail);
+	debugfs_create_u32("irq", 0400, eq->mana_eq_debugfs, &eq->eq->eq.irq);
+	debugfs_create_file("eq_dump", 0400, eq->mana_eq_debugfs, eq->eq,
+			    &mana_dbg_q_fops);
 }
 
 int mana_create_eq(struct mana_port_context *apc)
@@ -2037,11 +2040,37 @@ static int mana_grow_eqs(struct mana_port_context *apc, unsigned int need)
 
 	return 0;
 out:
-	/* Retain partial growth for reuse; the live set still needs this pool.
-	 */
 	return err;
 }
 
+/* All CQs referencing EQs at or above @keep must be destroyed first. */
+static void mana_shrink_eqs(struct mana_port_context *apc, unsigned int keep)
+{
+	struct gdma_context *gc = apc->ac->gdma_dev->gdma_context;
+	struct gdma_queue *eq;
+	unsigned int msi;
+	unsigned int i;
+
+	if (!apc->eqs || keep >= apc->num_eqs)
+		return;
+
+	for (i = keep; i < apc->num_eqs; i++) {
+		eq = apc->eqs[i].eq;
+		if (!eq)
+			continue;
+
+		debugfs_remove_recursive(apc->eqs[i].mana_eq_debugfs);
+		apc->eqs[i].mana_eq_debugfs = NULL;
+
+		msi = eq->eq.msix_index;
+		mana_gd_destroy_queue(gc, eq);
+		mana_gd_put_gic(gc, !gc->msi_sharing, msi);
+		apc->eqs[i].eq = NULL;
+	}
+
+	apc->num_eqs = keep;
+}
+
 static int mana_fence_rq(struct mana_port_context *apc, struct mana_rxq *rxq)
 {
 	struct mana_fence_rq_resp resp = {};
@@ -4151,6 +4180,8 @@ int mana_alloc_qset(struct mana_port_context *apc,
 	kfree(scratch->rxqs);
 	scratch->rxqs = NULL;
 out_err:
+	mana_shrink_eqs(apc, apc->num_queues);
+
 	netdev_err(ndev, "%s(num_queues=%u) failed: %d\n", __func__,
 		   num_queues, err);
 	return err;
@@ -4511,6 +4542,8 @@ void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset)
 	 */
 	netif_set_real_num_rx_queues(apc->ndev, apc->num_queues);
 
+	mana_shrink_eqs(apc, apc->num_queues);
+
 	mana_qset_debugfs_publish(apc);
 }
 
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index a438bc6097d6..611521d2d6d1 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -752,7 +752,6 @@ static int mana_set_channels(struct net_device *ndev,
 	mana_free_qset(scratch, &oldq);
 
 free_scratch:
-	/* Release unpublished queues before closing their shared EQ pool. */
 	mana_publish_close_if_needed(apc);
 	mana_qset_scratch_free(scratch);
 clear_flag:

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 11/13] net: mana: keep a user-configured RSS table across a queue rebuild
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (9 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 10/13] net: mana: release EQs left idle by a channel-count reduction Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 12/13] net: mana: keep the surviving queues when the channel count is reduced Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 13/13] net: mana: keep the existing queues when the channel count is raised Wei Hu
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Preserve a user RSS table whenever all entries fit the requested queue
count. Regenerate driver defaults. On growth, a retained user table does
not steer RSS traffic to the added queues until the user updates it.

Report table loss only after successful queue-set publication. The
non-swap allocation path still replaces invalid tables silently because
its callers do not consistently hold the notification's netdev lock.

Require RTNL for RSS setters to serialize them with reset/resume rebuilds.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 drivers/net/ethernet/microsoft/mana/mana_en.c | 41 +++++++++++++++++--
 .../ethernet/microsoft/mana/mana_ethtool.c    |  3 +-
 include/net/mana/mana.h                       |  2 +
 3 files changed, 42 insertions(+), 4 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 86a09828b5f4..5619c6a763c2 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -3595,6 +3595,27 @@ static void mana_rss_table_init(struct mana_port_context *apc)
 			ethtool_rxfh_indir_default(i, apc->num_queues);
 }
 
+/* Keep user tables with valid indices; defer loss notification. */
+static bool mana_rss_table_keep(struct mana_port_context *apc,
+				unsigned int num_queues, bool *lost)
+{
+	u32 i;
+
+	*lost = false;
+
+	if (!netif_is_rxfh_configured(apc->ndev))
+		return false;
+
+	for (i = 0; i < apc->indir_table_sz; i++) {
+		if (apc->indir_table[i] >= num_queues) {
+			*lost = true;
+			return false;
+		}
+	}
+
+	return true;
+}
+
 int mana_disable_vport_rx(struct mana_port_context *apc)
 {
 	return mana_cfg_vport_steering(apc, TRI_STATE_FALSE, false, false,
@@ -3865,6 +3886,7 @@ int mana_alloc_queues(struct net_device *ndev)
 {
 	struct mana_port_context *apc = netdev_priv(ndev);
 	struct gdma_dev *gd = apc->ac->gdma_dev;
+	bool indir_lost;
 	int err;
 
 	err = mana_create_vport(apc, ndev);
@@ -3910,7 +3932,9 @@ int mana_alloc_queues(struct net_device *ndev)
 		goto destroy_rxq;
 	}
 
-	mana_rss_table_init(apc);
+	/* Loss notification needs a netdev instance lock we may lack. */
+	if (!mana_rss_table_keep(apc, apc->num_queues, &indir_lost))
+		mana_rss_table_init(apc);
 
 	err = mana_config_rss(apc, TRI_STATE_TRUE, true, true);
 	if (err) {
@@ -4069,9 +4093,10 @@ static void mana_qset_snapshot(const struct mana_port_context *ctx,
 	out->priv_flags		= ctx->priv_flags;
 	out->mtu		= ctx->configured_mtu;
 	out->bpf_prog		= ctx->bpf_prog;
+
+	out->rxfh_indir_lost	= false;
 }
 
-/* Vport identity and port debugfs outlive queue sets. */
 static void mana_qset_install(struct mana_port_context *ctx,
 			      const struct mana_qset *qset)
 {
@@ -4131,6 +4156,7 @@ int mana_alloc_qset(struct mana_port_context *apc,
 		    struct mana_qset *out)
 {
 	struct net_device *ndev = scratch->ndev;
+	bool indir_lost;
 	int err;
 
 	ASSERT_RTNL();
@@ -4166,9 +4192,14 @@ int mana_alloc_qset(struct mana_port_context *apc,
 	if (err)
 		goto cleanup_rxq;
 
-	mana_rss_table_init(scratch);
+	if (mana_rss_table_keep(apc, num_queues, &indir_lost))
+		memcpy(scratch->indir_table, apc->indir_table,
+		       apc->indir_table_sz * sizeof(*apc->indir_table));
+	else
+		mana_rss_table_init(scratch);
 
 	mana_qset_snapshot(scratch, out);
+	out->rxfh_indir_lost = indir_lost;
 	return 0;
 
 cleanup_rxq:
@@ -4407,6 +4438,10 @@ int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 	WRITE_ONCE(apc->port_is_up, true);
 	mana_start_txqs(apc);
 
+	/* Report a lost user table only after successful publication. */
+	if (newq->rxfh_indir_lost)
+		ethtool_rxfh_indir_lost(ndev);
+
 	return 0;
 
 rollback:
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index 611521d2d6d1..a53e19b15bae 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -953,7 +953,8 @@ const struct ethtool_ops mana_ethtool_ops = {
 	.op_needs_rtnl		= ETHTOOL_OP_NEEDS_RTNL_SCHANNELS |
 				  ETHTOOL_OP_NEEDS_RTNL_SRINGPARAM |
 				  ETHTOOL_OP_NEEDS_RTNL_SPFLAGS |
-				  ETHTOOL_OP_NEEDS_RTNL_GLINK,
+				  ETHTOOL_OP_NEEDS_RTNL_GLINK |
+				  ETHTOOL_OP_NEEDS_RTNL_RSS,
 	.get_ethtool_stats	= mana_get_ethtool_stats,
 	.get_sset_count		= mana_get_sset_count,
 	.get_strings		= mana_get_strings,
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index 5f9461bc16e9..9dad03150d1b 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -709,6 +709,8 @@ struct mana_qset {
 	int			mtu;
 	struct bpf_prog		*bpf_prog;
 
+	/* Notify the core only after this set is published. */
+	bool			rxfh_indir_lost;
 };
 
 netdev_tx_t mana_start_xmit(struct sk_buff *skb, struct net_device *ndev);

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 12/13] net: mana: keep the surviving queues when the channel count is reduced
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (10 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 11/13] net: mana: keep a user-configured RSS table across a queue rebuild Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  2026-10-09 14:41 ` [PATCH net-next v6 13/13] net: mana: keep the existing queues when the channel count is raised Wei Hu
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Split the live set into a kept prefix and a retiring tail. Reductions
allocate only pointer arrays and steering tables, retaining the kept
queues' page pools, buffers, NAPI state and XDP references.

After publication, wait for TX-selection readers before freeing the old
containers, then retire only the tail. Failed publication discards the
new containers without freeing shared queues.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 drivers/net/ethernet/microsoft/mana/mana_en.c | 112 +++++++++++++++++-
 .../ethernet/microsoft/mana/mana_ethtool.c    |  30 +++++
 include/net/mana/mana.h                       |   4 +
 3 files changed, 143 insertions(+), 3 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 5619c6a763c2..2fe53adda119 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -4149,6 +4149,114 @@ void mana_qset_scratch_free(struct mana_port_context *scratch)
 	kvfree(scratch);
 }
 
+/* Split into kept queues and a retiring tail without changing live ownership.
+ * Queue i retains EQ i.
+ */
+int mana_split_qset(struct mana_port_context *apc,
+		    struct mana_port_context *scratch, unsigned int new_count,
+		    struct mana_qset *out_new, struct mana_qset *out_tail)
+{
+	unsigned int old_count = apc->num_queues;
+	struct mana_tx_qp **new_tx, **tail_tx;
+	struct mana_rxq **new_rx, **tail_rx;
+	unsigned int tail_count;
+	bool indir_lost;
+	unsigned int i;
+	int err;
+
+	ASSERT_RTNL();
+
+	if (WARN_ON(new_count == 0 || new_count >= old_count))
+		return -EINVAL;
+	if (WARN_ON(!apc->tx_qp || !apc->rxqs))
+		return -EINVAL;
+
+	tail_count = old_count - new_count;
+
+	/* Build steering separately so it cannot index beyond the shortened RX
+	 * array.
+	 */
+	scratch->num_queues = new_count;
+	err = mana_rss_table_alloc(scratch);
+	if (err)
+		return err;
+
+	if (mana_rss_table_keep(apc, new_count, &indir_lost))
+		memcpy(scratch->indir_table, apc->indir_table,
+		       apc->indir_table_sz * sizeof(*apc->indir_table));
+	else
+		mana_rss_table_init(scratch);
+
+	new_tx = kzalloc_objs(struct mana_tx_qp *, new_count);
+	new_rx = kzalloc_objs(struct mana_rxq *, new_count);
+	tail_tx = kzalloc_objs(struct mana_tx_qp *, tail_count);
+	tail_rx = kzalloc_objs(struct mana_rxq *, tail_count);
+	if (!new_tx || !new_rx || !tail_tx || !tail_rx) {
+		err = -ENOMEM;
+		goto free_arrays;
+	}
+
+	for (i = 0; i < new_count; i++) {
+		new_tx[i] = apc->tx_qp[i];
+		new_rx[i] = apc->rxqs[i];
+	}
+	for (i = 0; i < tail_count; i++) {
+		tail_tx[i] = apc->tx_qp[new_count + i];
+		tail_rx[i] = apc->rxqs[new_count + i];
+	}
+
+	out_new->tx_qp		= new_tx;
+	out_new->rxqs		= new_rx;
+	out_new->indir_table	= scratch->indir_table;
+	out_new->indir_table_sz	= scratch->indir_table_sz;
+	out_new->rxobj_table	= scratch->rxobj_table;
+	out_new->default_rxobj	= apc->rxqs[0]->rxobj;
+	out_new->num_queues	= new_count;
+	out_new->rx_queue_size	= apc->rx_queue_size;
+	out_new->tx_queue_size	= apc->tx_queue_size;
+	out_new->priv_flags	= apc->priv_flags;
+	out_new->mtu		= apc->configured_mtu;
+	out_new->bpf_prog	= apc->bpf_prog;
+	out_new->rxfh_indir_lost = indir_lost;
+
+	scratch->indir_table	= NULL;
+	scratch->rxobj_table	= NULL;
+
+	memset(out_tail, 0, sizeof(*out_tail));
+	out_tail->tx_qp		= tail_tx;
+	out_tail->rxqs		= tail_rx;
+	out_tail->default_rxobj	= INVALID_MANA_HANDLE;
+	out_tail->num_queues	= tail_count;
+	out_tail->rx_queue_size	= apc->rx_queue_size;
+	out_tail->tx_queue_size	= apc->tx_queue_size;
+	out_tail->priv_flags	= apc->priv_flags;
+	out_tail->mtu		= apc->configured_mtu;
+	out_tail->bpf_prog	= apc->bpf_prog;
+
+	return 0;
+
+free_arrays:
+	kfree(new_tx);
+	kfree(new_rx);
+	kfree(tail_tx);
+	kfree(tail_rx);
+	mana_cleanup_indir_table(scratch);
+	return err;
+}
+
+/* Free containers only; the live port still owns the queues. */
+void mana_discard_split(struct mana_qset *newq, struct mana_qset *tailq)
+{
+	kfree(newq->tx_qp);
+	kfree(newq->rxqs);
+	kfree(newq->indir_table);
+	kfree(newq->rxobj_table);
+	kfree(tailq->tx_qp);
+	kfree(tailq->rxqs);
+	memset(newq, 0, sizeof(*newq));
+	memset(tailq, 0, sizeof(*tailq));
+}
+
 int mana_alloc_qset(struct mana_port_context *apc,
 		    struct mana_port_context *scratch, unsigned int num_queues,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
@@ -4419,9 +4527,7 @@ int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 	if (err)
 		goto rollback;
 
-	/* Install XDP and per-RXQ references before steering reaches new
-	 * queues.
-	 */
+	/* Install XDP before steering reaches the incoming RXQs. */
 	mana_chn_setxdp(apc, mana_xdp_get(apc));
 
 	err = mana_config_rss(apc, TRI_STATE_TRUE, true, true);
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index a53e19b15bae..d3e465ea9428 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -737,6 +737,36 @@ static int mana_set_channels(struct net_device *ndev,
 		goto clear_flag;
 	}
 
+	if (new_count < apc->num_queues) {
+		struct mana_qset tailq;
+
+		err = mana_split_qset(apc, scratch, new_count, &newq, &tailq);
+		if (err)
+			goto free_scratch;
+
+		err = mana_publish_qset(apc, &newq, &oldq);
+		if (err) {
+			/* Discard containers only; their queues still belong to
+			 * the old set.
+			 */
+			mana_discard_split(&newq, &tailq);
+			goto free_scratch;
+		}
+
+		/* Wait for ndo_select_queue() readers of oldq.indir_table. Free
+		 * only containers; the queues belong to the kept set or tail.
+		 */
+		synchronize_net();
+
+		kfree(oldq.tx_qp);
+		kfree(oldq.rxqs);
+		kfree(oldq.indir_table);
+		kfree(oldq.rxobj_table);
+
+		mana_free_qset(scratch, &tailq);
+		goto free_scratch;
+	}
+
 	err = mana_alloc_qset(apc, scratch, new_count, apc->rx_queue_size,
 			      apc->tx_queue_size, apc->priv_flags,
 			      apc->configured_mtu, apc->bpf_prog, &newq);
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index 9dad03150d1b..fef800bea2c0 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -735,6 +735,10 @@ int mana_alloc_qset(struct mana_port_context *apc,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
 		    u32 priv_flags, int mtu, struct bpf_prog *bpf_prog,
 		    struct mana_qset *out);
+int mana_split_qset(struct mana_port_context *apc,
+		    struct mana_port_context *scratch, unsigned int new_count,
+		    struct mana_qset *out_new, struct mana_qset *out_tail);
+void mana_discard_split(struct mana_qset *newq, struct mana_qset *tailq);
 int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 		      struct mana_qset *out_old);
 void mana_publish_close_if_needed(struct mana_port_context *apc);

^ permalink raw reply	[flat|nested] 14+ messages in thread

* [PATCH net-next v6 13/13] net: mana: keep the existing queues when the channel count is raised
  2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
                   ` (11 preceding siblings ...)
  2026-10-09 14:41 ` [PATCH net-next v6 12/13] net: mana: keep the surviving queues when the channel count is reduced Wei Hu
@ 2026-10-09 14:41 ` Wei Hu
  12 siblings, 0 replies; 14+ messages in thread
From: Wei Hu @ 2026-10-09 14:41 UTC (permalink / raw)
  To: longli, kotaranov, kuba, davem, pabeni, edumazet, andrew+netdev,
	jgg, leon, haiyangz, wei.liu, decui, shradhagupta, horms, ernis,
	stephen
  Cc: netdev, linux-rdma, linux-hyperv, linux-kernel, dipayanroy, bpf,
	sdf, daniel, hawk, ast, john.fastabend, weh

From: Long Li <longli@microsoft.com>

Keep existing queues and allocate only the added tail. Growing N to M
now needs M SQ/RQ pairs at peak, rather than N + M.

Track the fresh queues separately for failure cleanup and XDP references.
Wait for TX-selection readers before freeing old containers, and clear
slots during partial teardown. Full rebuilds now keep the current count.

Advertise in-driver resize recovery after converting the live resize
paths. Failed rollback still requires recovery.

Signed-off-by: Long Li <longli@microsoft.com>
Signed-off-by: Wei Hu <weh@microsoft.com>
---
 .../net/ethernet/microsoft/mana/mana_bpf.c    |   2 +-
 drivers/net/ethernet/microsoft/mana/mana_en.c | 209 +++++++++++++++---
 .../ethernet/microsoft/mana/mana_ethtool.c    |  31 ++-
 include/net/mana/gdma.h                       |   8 +-
 include/net/mana/mana.h                       |   7 +-
 5 files changed, 217 insertions(+), 40 deletions(-)

diff --git a/drivers/net/ethernet/microsoft/mana/mana_bpf.c b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
index 0ad3a0fa6747..798f0dc485f8 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_bpf.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_bpf.c
@@ -216,7 +216,7 @@ static int mana_xdp_set(struct net_device *ndev, struct bpf_prog *prog,
 			goto clear_flag;
 		}
 
-		err = mana_alloc_qset(apc, scratch, apc->num_queues,
+		err = mana_alloc_qset(apc, scratch,
 				      apc->rx_queue_size, apc->tx_queue_size,
 				      apc->priv_flags, apc->configured_mtu,
 				      prog, &newq);
diff --git a/drivers/net/ethernet/microsoft/mana/mana_en.c b/drivers/net/ethernet/microsoft/mana/mana_en.c
index 2fe53adda119..5fbadc4dccbc 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_en.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_en.c
@@ -952,7 +952,7 @@ static int mana_change_mtu(struct net_device *ndev, int new_mtu)
 		goto clear_flag;
 	}
 
-	err = mana_alloc_qset(mpc, scratch, mpc->num_queues, mpc->rx_queue_size,
+	err = mana_alloc_qset(mpc, scratch, mpc->rx_queue_size,
 			      mpc->tx_queue_size, mpc->priv_flags, new_mtu,
 			      mpc->bpf_prog, &newq);
 	if (err)
@@ -2920,7 +2920,9 @@ static void mana_deinit_txq(struct mana_port_context *apc, struct mana_txq *txq)
 	mana_gd_destroy_queue(gd->gdma_context, txq->gdma_sq);
 }
 
-static void mana_destroy_txq(struct mana_port_context *apc)
+/* Keep the array and queues below @first; clear freed slots. */
+static void mana_destroy_txq_from(struct mana_port_context *apc,
+				  unsigned int first)
 {
 	struct napi_struct *napi;
 	int i;
@@ -2928,7 +2930,7 @@ static void mana_destroy_txq(struct mana_port_context *apc)
 	if (!apc->tx_qp)
 		return;
 
-	for (i = 0; i < apc->num_queues; i++) {
+	for (i = first; i < apc->num_queues; i++) {
 		if (!apc->tx_qp[i])
 			continue;
 
@@ -2952,7 +2954,16 @@ static void mana_destroy_txq(struct mana_port_context *apc)
 		mana_deinit_txq(apc, &apc->tx_qp[i]->txq);
 
 		kvfree(apc->tx_qp[i]);
+		apc->tx_qp[i] = NULL;
 	}
+}
+
+static void mana_destroy_txq(struct mana_port_context *apc)
+{
+	if (!apc->tx_qp)
+		return;
+
+	mana_destroy_txq_from(apc, 0);
 
 	kfree(apc->tx_qp);
 	apc->tx_qp = NULL;
@@ -2983,8 +2994,11 @@ static void mana_create_txq_debugfs(struct mana_port_context *apc, int idx)
 			    tx_qp->tx_cq.gdma_cq, &mana_dbg_q_fops);
 }
 
+/* With @first nonzero, use the existing array and unwind only new queues on
+ * failure.
+ */
 static int mana_create_txq(struct mana_port_context *apc,
-			   struct net_device *net)
+			   struct net_device *net, unsigned int first)
 {
 	struct mana_context *ac = apc->ac;
 	struct gdma_dev *gd = ac->gdma_dev;
@@ -2999,9 +3013,14 @@ static int mana_create_txq(struct mana_port_context *apc,
 	int err;
 	int i;
 
-	apc->tx_qp = kzalloc_objs(struct mana_tx_qp *, apc->num_queues);
-	if (!apc->tx_qp)
-		return -ENOMEM;
+	if (first) {
+		if (WARN_ON(!apc->tx_qp))
+			return -EINVAL;
+	} else {
+		apc->tx_qp = kzalloc_objs(struct mana_tx_qp *, apc->num_queues);
+		if (!apc->tx_qp)
+			return -ENOMEM;
+	}
 
 	/*  The minimum size of the WQE is 32 bytes, hence
 	 *  apc->tx_queue_size represents the maximum number of WQEs
@@ -3018,7 +3037,7 @@ static int mana_create_txq(struct mana_port_context *apc,
 
 	gc = gd->gdma_context;
 
-	for (i = 0; i < apc->num_queues; i++) {
+	for (i = first; i < apc->num_queues; i++) {
 		apc->tx_qp[i] = kvzalloc_obj(*apc->tx_qp[i]);
 		if (!apc->tx_qp[i]) {
 			err = -ENOMEM;
@@ -3126,7 +3145,10 @@ static int mana_create_txq(struct mana_port_context *apc,
 out:
 	netdev_err(net, "Failed to create %d TX queues, %d\n",
 		   apc->num_queues, err);
-	mana_destroy_txq(apc);
+	if (first)
+		mana_destroy_txq_from(apc, first);
+	else
+		mana_destroy_txq(apc);
 	return err;
 }
 
@@ -3487,14 +3509,15 @@ static void mana_create_rxq_debugfs(struct mana_port_context *apc, int idx)
 			    &mana_dbg_q_fops);
 }
 
+/* The caller must destroy queues added before a failure. */
 static int mana_add_rx_queues(struct mana_port_context *apc,
-			      struct net_device *ndev)
+			      struct net_device *ndev, unsigned int first)
 {
 	struct mana_rxq *rxq;
 	int err = 0;
 	int i;
 
-	for (i = 0; i < apc->num_queues; i++) {
+	for (i = first; i < apc->num_queues; i++) {
 		rxq = mana_create_rxq(apc, i, &apc->eqs[i], ndev);
 		if (IS_ERR(rxq)) {
 			err = PTR_ERR(rxq);
@@ -3512,14 +3535,15 @@ static int mana_add_rx_queues(struct mana_port_context *apc,
 	return err;
 }
 
-static void mana_destroy_rxqs(struct mana_port_context *apc)
+static void mana_destroy_rxqs_from(struct mana_port_context *apc,
+				   unsigned int first)
 {
 	struct mana_rxq *rxq;
 	u32 rxq_idx;
 
 	if (apc->rxqs) {
 
-		for (rxq_idx = 0; rxq_idx < apc->num_queues; rxq_idx++) {
+		for (rxq_idx = first; rxq_idx < apc->num_queues; rxq_idx++) {
 			rxq = apc->rxqs[rxq_idx];
 			if (!rxq)
 				continue;
@@ -3530,6 +3554,11 @@ static void mana_destroy_rxqs(struct mana_port_context *apc)
 	}
 }
 
+static void mana_destroy_rxqs(struct mana_port_context *apc)
+{
+	mana_destroy_rxqs_from(apc, 0);
+}
+
 static void mana_destroy_vport(struct mana_port_context *apc)
 {
 	struct gdma_dev *gd = apc->ac->gdma_dev;
@@ -3903,7 +3932,7 @@ int mana_alloc_queues(struct net_device *ndev)
 		goto destroy_vport;
 	}
 
-	err = mana_create_txq(apc, ndev);
+	err = mana_create_txq(apc, ndev, 0);
 	if (err) {
 		netdev_err(ndev, "Failed to create TXQ on vPort %u: %d\n",
 			   apc->port_idx, err);
@@ -3918,7 +3947,7 @@ int mana_alloc_queues(struct net_device *ndev)
 		goto destroy_txq;
 	}
 
-	err = mana_add_rx_queues(apc, ndev);
+	err = mana_add_rx_queues(apc, ndev, 0);
 	if (err)
 		goto destroy_rxq;
 
@@ -4257,8 +4286,137 @@ void mana_discard_split(struct mana_qset *newq, struct mana_qset *tailq)
 	memset(tailq, 0, sizeof(*tailq));
 }
 
+/* Carry existing queues into @out_new; allocate only the tail. @out_fresh
+ * isolates new queues for cleanup after a failed publish.
+ */
+int mana_grow_qset(struct mana_port_context *apc,
+		   struct mana_port_context *scratch, unsigned int new_count,
+		   struct mana_qset *out_new, struct mana_qset *out_fresh)
+{
+	unsigned int old_count = apc->num_queues;
+	struct mana_tx_qp **new_tx, **fresh_tx;
+	struct mana_rxq **new_rx, **fresh_rx;
+	struct net_device *ndev = apc->ndev;
+	unsigned int fresh_count;
+	bool indir_lost;
+	unsigned int i;
+	int err;
+
+	ASSERT_RTNL();
+
+	if (WARN_ON(new_count <= old_count))
+		return -EINVAL;
+	if (WARN_ON(!apc->tx_qp || !apc->rxqs))
+		return -EINVAL;
+
+	fresh_count = new_count - old_count;
+
+	new_tx = kzalloc_objs(struct mana_tx_qp *, new_count);
+	new_rx = kzalloc_objs(struct mana_rxq *, new_count);
+	fresh_tx = kzalloc_objs(struct mana_tx_qp *, fresh_count);
+	fresh_rx = kzalloc_objs(struct mana_rxq *, fresh_count);
+	if (!new_tx || !new_rx || !fresh_tx || !fresh_rx) {
+		err = -ENOMEM;
+		goto free_arrays;
+	}
+
+	for (i = 0; i < old_count; i++) {
+		new_tx[i] = apc->tx_qp[i];
+		new_rx[i] = apc->rxqs[i];
+	}
+
+	scratch->num_queues = new_count;
+	scratch->tx_qp = new_tx;
+	scratch->rxqs = new_rx;
+
+	err = mana_rss_table_alloc(scratch);
+	if (err)
+		goto free_arrays;
+
+	err = mana_grow_eqs(apc, new_count);
+	if (err)
+		goto cleanup_rss;
+
+	scratch->eqs = apc->eqs;
+	scratch->num_eqs = apc->num_eqs;
+
+	err = mana_create_txq(scratch, ndev, old_count);
+	if (err)
+		goto cleanup_rss;
+
+	err = mana_add_rx_queues(scratch, ndev, old_count);
+	if (err)
+		goto cleanup_rxq;
+
+	if (mana_rss_table_keep(apc, new_count, &indir_lost))
+		memcpy(scratch->indir_table, apc->indir_table,
+		       apc->indir_table_sz * sizeof(*apc->indir_table));
+	else
+		mana_rss_table_init(scratch);
+
+	mana_qset_snapshot(scratch, out_new);
+	out_new->rxfh_indir_lost = indir_lost;
+
+	for (i = 0; i < fresh_count; i++) {
+		fresh_tx[i] = new_tx[old_count + i];
+		fresh_rx[i] = new_rx[old_count + i];
+	}
+
+	memset(out_fresh, 0, sizeof(*out_fresh));
+	out_fresh->tx_qp	= fresh_tx;
+	out_fresh->rxqs		= fresh_rx;
+	out_fresh->default_rxobj = INVALID_MANA_HANDLE;
+	out_fresh->num_queues	= fresh_count;
+	out_fresh->rx_queue_size = apc->rx_queue_size;
+	out_fresh->tx_queue_size = apc->tx_queue_size;
+	out_fresh->priv_flags	= apc->priv_flags;
+	out_fresh->mtu		= apc->configured_mtu;
+	out_fresh->bpf_prog	= apc->bpf_prog;
+
+	/* Take XDP refs on fresh RXQs only. On the merged set,
+	 * mana_chn_setxdp() returns early on the carried rxqs[0].
+	 */
+	mana_qset_install(scratch, out_fresh);
+	mana_chn_setxdp(scratch, mana_xdp_get(apc));
+
+	return 0;
+
+cleanup_rxq:
+	mana_destroy_rxqs_from(scratch, old_count);
+	mana_destroy_txq_from(scratch, old_count);
+cleanup_rss:
+	mana_cleanup_indir_table(scratch);
+free_arrays:
+	/* Free containers only; carried queues remain live. */
+	scratch->tx_qp = NULL;
+	scratch->rxqs = NULL;
+	kfree(new_tx);
+	kfree(new_rx);
+	kfree(fresh_tx);
+	kfree(fresh_rx);
+
+	mana_shrink_eqs(apc, apc->num_queues);
+
+	netdev_err(ndev, "%s(num_queues=%u) failed: %d\n", __func__,
+		   new_count, err);
+	return err;
+}
+
+/* Free merged containers only, not carried queues. The caller must retire fresh
+ * queues separately.
+ */
+void mana_discard_grow(struct mana_qset *newq)
+{
+	kfree(newq->tx_qp);
+	kfree(newq->rxqs);
+	kfree(newq->indir_table);
+	kfree(newq->rxobj_table);
+	memset(newq, 0, sizeof(*newq));
+}
+
+/* Rebuild at the current count; resize uses split/grow. */
 int mana_alloc_qset(struct mana_port_context *apc,
-		    struct mana_port_context *scratch, unsigned int num_queues,
+		    struct mana_port_context *scratch,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
 		    u32 priv_flags, int mtu, struct bpf_prog *bpf_prog,
 		    struct mana_qset *out)
@@ -4269,7 +4427,7 @@ int mana_alloc_qset(struct mana_port_context *apc,
 
 	ASSERT_RTNL();
 
-	scratch->num_queues	= num_queues;
+	scratch->num_queues	= apc->num_queues;
 	scratch->rx_queue_size	= rx_queue_size;
 	scratch->tx_queue_size	= tx_queue_size;
 	scratch->priv_flags	= priv_flags;
@@ -4285,22 +4443,19 @@ int mana_alloc_qset(struct mana_port_context *apc,
 	if (err)
 		goto cleanup_rxq_array;
 
-	err = mana_grow_eqs(apc, num_queues);
-	if (err)
-		goto cleanup_rss;
-
+	/* Reuse the existing EQ pool; the queue count is unchanged. */
 	scratch->eqs = apc->eqs;
 	scratch->num_eqs = apc->num_eqs;
 
-	err = mana_create_txq(scratch, ndev);
+	err = mana_create_txq(scratch, ndev, 0);
 	if (err)
 		goto cleanup_rss;
 
-	err = mana_add_rx_queues(scratch, ndev);
+	err = mana_add_rx_queues(scratch, ndev, 0);
 	if (err)
 		goto cleanup_rxq;
 
-	if (mana_rss_table_keep(apc, num_queues, &indir_lost))
+	if (mana_rss_table_keep(apc, scratch->num_queues, &indir_lost))
 		memcpy(scratch->indir_table, apc->indir_table,
 		       apc->indir_table_sz * sizeof(*apc->indir_table));
 	else
@@ -4319,10 +4474,8 @@ int mana_alloc_qset(struct mana_port_context *apc,
 	kfree(scratch->rxqs);
 	scratch->rxqs = NULL;
 out_err:
-	mana_shrink_eqs(apc, apc->num_queues);
-
 	netdev_err(ndev, "%s(num_queues=%u) failed: %d\n", __func__,
-		   num_queues, err);
+		   apc->num_queues, err);
 	return err;
 }
 
@@ -4607,7 +4760,7 @@ int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 }
 
 /* Create missing debugfs nodes once retiring names are gone. */
-static void mana_qset_debugfs_publish(struct mana_port_context *apc)
+void mana_qset_debugfs_publish(struct mana_port_context *apc)
 {
 	unsigned int i;
 
diff --git a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
index d3e465ea9428..ec76a75268ac 100644
--- a/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
+++ b/drivers/net/ethernet/microsoft/mana/mana_ethtool.c
@@ -684,7 +684,7 @@ static int mana_set_channels(struct net_device *ndev,
 	struct mana_port_context *apc = netdev_priv(ndev);
 	unsigned int new_count = channels->combined_count;
 	struct mana_port_context *scratch;
-	struct mana_qset newq, oldq;
+	struct mana_qset newq, oldq, freshq;
 	int err;
 
 	if (new_count < 1 || new_count > apc->max_queues) {
@@ -767,19 +767,33 @@ static int mana_set_channels(struct net_device *ndev,
 		goto free_scratch;
 	}
 
-	err = mana_alloc_qset(apc, scratch, new_count, apc->rx_queue_size,
-			      apc->tx_queue_size, apc->priv_flags,
-			      apc->configured_mtu, apc->bpf_prog, &newq);
+	err = mana_grow_qset(apc, scratch, new_count, &newq, &freshq);
 	if (err)
 		goto free_scratch;
 
 	err = mana_publish_qset(apc, &newq, &oldq);
 	if (err) {
-		mana_free_qset(scratch, &newq);
+		/* Free only the new queues, then discard the merged containers.
+		 */
+		mana_free_qset(scratch, &freshq);
+		mana_discard_grow(&newq);
 		goto free_scratch;
 	}
 
-	mana_free_qset(scratch, &oldq);
+	/* Wait for ndo_select_queue() readers of oldq.indir_table. All queues
+	 * are now live in newq; free only the old and fresh containers.
+	 */
+	synchronize_net();
+
+	kfree(oldq.tx_qp);
+	kfree(oldq.rxqs);
+	kfree(oldq.indir_table);
+	kfree(oldq.rxobj_table);
+	kfree(freshq.tx_qp);
+	kfree(freshq.rxqs);
+
+	/* No retirement runs to publish the new queues' debugfs nodes. */
+	mana_qset_debugfs_publish(apc);
 
 free_scratch:
 	mana_publish_close_if_needed(apc);
@@ -856,7 +870,7 @@ static int mana_set_ringparam(struct net_device *ndev,
 		goto clear_flag;
 	}
 
-	err = mana_alloc_qset(apc, scratch, apc->num_queues, new_rx, new_tx,
+	err = mana_alloc_qset(apc, scratch, new_rx, new_tx,
 			      apc->priv_flags, apc->configured_mtu,
 			      apc->bpf_prog, &newq);
 	if (err) {
@@ -876,7 +890,6 @@ static int mana_set_ringparam(struct net_device *ndev,
 	mana_free_qset(scratch, &oldq);
 
 free_scratch:
-	/* Release unpublished queues before closing their shared EQ pool. */
 	mana_publish_close_if_needed(apc);
 	mana_qset_scratch_free(scratch);
 clear_flag:
@@ -948,7 +961,7 @@ static int mana_set_priv_flags(struct net_device *ndev, u32 priv_flags)
 		goto clear_flag;
 	}
 
-	err = mana_alloc_qset(apc, scratch, apc->num_queues, apc->rx_queue_size,
+	err = mana_alloc_qset(apc, scratch, apc->rx_queue_size,
 			      apc->tx_queue_size, priv_flags,
 			      apc->configured_mtu, apc->bpf_prog, &newq);
 	if (err)
diff --git a/include/net/mana/gdma.h b/include/net/mana/gdma.h
index 308950f9b54b..666565ffb26a 100644
--- a/include/net/mana/gdma.h
+++ b/include/net/mana/gdma.h
@@ -686,6 +686,11 @@ enum {
 /* Driver supports non-contiguous queue buffers */
 #define GDMA_DRV_CAP_FLAG_1_NON_CONTIGUOUS_BUFFERS BIT(30)
 
+/* Resize failures are handled in-driver; a failed rollback still needs
+ * recovery.
+ */
+#define GDMA_DRV_CAP_FLAG_1_SELF_RECOVERY_ON_QUEUE_RESIZE_FAILURE BIT(31)
+
 #define GDMA_DRV_CAP_FLAGS1 \
 	(GDMA_DRV_CAP_FLAG_1_EQ_SHARING_MULTI_VPORT | \
 	 GDMA_DRV_CAP_FLAG_1_NAPI_WKDONE_FIX | \
@@ -703,7 +708,8 @@ enum {
 	 GDMA_DRV_CAP_FLAG_1_HWC_TIMEOUT_RECOVERY | \
 	 GDMA_DRV_CAP_FLAG_1_EQ_MSI_UNSHARE_MULTI_VPORT | \
 	 GDMA_DRV_CAP_FLAG_1_DYN_INTERRUPT_MODERATION | \
-	 GDMA_DRV_CAP_FLAG_1_NON_CONTIGUOUS_BUFFERS)
+	 GDMA_DRV_CAP_FLAG_1_NON_CONTIGUOUS_BUFFERS | \
+	 GDMA_DRV_CAP_FLAG_1_SELF_RECOVERY_ON_QUEUE_RESIZE_FAILURE)
 
 #define GDMA_DRV_CAP_FLAGS2 0
 
diff --git a/include/net/mana/mana.h b/include/net/mana/mana.h
index fef800bea2c0..2e101730e40c 100644
--- a/include/net/mana/mana.h
+++ b/include/net/mana/mana.h
@@ -731,7 +731,7 @@ static inline struct mana_stats_rx *mana_rxq_stats(struct mana_rxq *rxq)
 }
 
 int mana_alloc_qset(struct mana_port_context *apc,
-		    struct mana_port_context *scratch, unsigned int num_queues,
+		    struct mana_port_context *scratch,
 		    unsigned int rx_queue_size, unsigned int tx_queue_size,
 		    u32 priv_flags, int mtu, struct bpf_prog *bpf_prog,
 		    struct mana_qset *out);
@@ -739,10 +739,15 @@ int mana_split_qset(struct mana_port_context *apc,
 		    struct mana_port_context *scratch, unsigned int new_count,
 		    struct mana_qset *out_new, struct mana_qset *out_tail);
 void mana_discard_split(struct mana_qset *newq, struct mana_qset *tailq);
+int mana_grow_qset(struct mana_port_context *apc,
+		   struct mana_port_context *scratch, unsigned int new_count,
+		   struct mana_qset *out_new, struct mana_qset *out_fresh);
+void mana_discard_grow(struct mana_qset *newq);
 int mana_publish_qset(struct mana_port_context *apc, struct mana_qset *newq,
 		      struct mana_qset *out_old);
 void mana_publish_close_if_needed(struct mana_port_context *apc);
 void mana_free_qset(struct mana_port_context *scratch, struct mana_qset *qset);
+void mana_qset_debugfs_publish(struct mana_port_context *apc);
 
 void mana_dim_change(struct mana_cq *cq, bool enable);
 

^ permalink raw reply	[flat|nested] 14+ messages in thread

end of thread, other threads:[~2026-10-09 14:42 UTC | newest]

Thread overview: 14+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 14:41 [PATCH net-next v6 00/13] net: mana: reconfigure by replacing the queue set Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 01/13] net: mana: add queue-set allocation and teardown helpers Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 02/13] net: mana: share the EQ pool across a queue-set swap Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 03/13] net: mana: keep per-queue statistics in the port context Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 04/13] net: mana: swap queue sets in mana_set_channels Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 05/13] net: mana: swap queue sets in mana_set_ringparam Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 06/13] net: mana: swap queue sets in mana_set_priv_flags Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 07/13] net: mana: swap queue sets in mana_change_mtu Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 08/13] net: mana: swap queue sets in mana_xdp_set Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 09/13] net: mana: do not bail out of mana_detach on dealloc failure Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 10/13] net: mana: release EQs left idle by a channel-count reduction Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 11/13] net: mana: keep a user-configured RSS table across a queue rebuild Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 12/13] net: mana: keep the surviving queues when the channel count is reduced Wei Hu
2026-10-09 14:41 ` [PATCH net-next v6 13/13] net: mana: keep the existing queues when the channel count is raised Wei Hu

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®