From: Wei Hu <weh@linux.microsoft.com>
To: Long Li <longli@kernel.org>,
Konstantin Taranov <kotaranov@microsoft.com>,
Jakub Kicinski <kuba@kernel.org>,
"David S. Miller" <davem@davemloft.net>,
Paolo Abeni <pabeni@redhat.com>,
Eric Dumazet <edumazet@google.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
Jason Gunthorpe <jgg@ziepe.ca>, Leon Romanovsky <leon@kernel.org>,
Haiyang Zhang <haiyangz@microsoft.com>,
"K. Y. Srinivasan" <kys@microsoft.com>,
Wei Liu <wei.liu@kernel.org>, Dexuan Cui <decui@microsoft.com>,
Shradha Gupta <shradhagupta@linux.microsoft.com>,
Simon Horman <horms@kernel.org>,
Erni Sri Satya Vennela <ernis@linux.microsoft.com>,
Stephen Hemminger <stephen@networkplumber.org>,
Shiraz Saleem <shirazsaleem@microsoft.com>
Cc: netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
linux-hyperv@vger.kernel.org, linux-kernel@vger.kernel.org,
Aditya Garg <gargaditya@linux.microsoft.com>,
Dipayaan Roy <dipayanroy@linux.microsoft.com>,
Kees Cook <kees+treewide@kernel.org>, Kees Cook <kees@kernel.org>,
Manish Awasthi <mawasthi@linux.microsoft.com>,
Wei Hu <weh@microsoft.com>
Subject: [PATCH net-next v6 0/4] net: mana: concurrent HWC requests and dynamic queue depth
Date: Wed, 7 Oct 2026 12:53:26 +0000 [thread overview]
Message-ID: <cover.1790665894.git.weh@linux.microsoft.com> (raw)
This series picks up Long Li's v5 MANA Hardware Channel (HWC) work,
addresses its review findings, and retains its four-patch feature lineage.
Safety changes required by concurrency and repeated HWC reinitialization
are folded into the patch which owns each behavior rather than posted as a
separate fixes series.
I have taken over follow-up upstreaming from Long. Long remains the author
of every patch, his existing sign-off is preserved, and I added mine for
the changes and submission custody.
Patch 1 prepares HWC ownership for safe reinitialization. It serializes
complete shared-memory transactions, tracks possible PF ownership,
retains mappings after an unacknowledged destroy, fences the HWC EQ before
releasing CQ state, and safely publishes and withdraws the HWC CQ table.
Patch 2 gives each slot independent completion state, withdraws caller
storage before sender return, handles the response-versus-timeout race,
and prevents depth-one slot reuse after a genuine timeout.
Patch 3 adds concurrent requests with bounded admission, serialized SQ
posting, separate sender/response ownership, stale-response quarantine,
coherent timeout updates, private setup publication, and sender draining.
Patch 4 rebuilds the depth-one bootstrap queues at validated depths from 2
through 128. It validates both reports, message limits and doorbells,
keeps the untouched bootstrap channel for depth one or reports above 128,
and restores fresh bootstrap queues after a failed rebuild only when
teardown was acknowledged.
Generic Ethernet/RDMA CQ callback refcounting and generic service-work
versus device-lifecycle serialization are independent pre-existing work
and are intentionally excluded. Malformed-RX handling, duplicate-success
correlation, and broad device-controlled allocation/timeout policy caps
also remain separate follow-ups.
Changes in v6 (v5 -> v6):
- Keep the four-patch v5 structure while folding review-required HWC
ownership, teardown, timeout, publication and fallback safety into
their owning feature patches.
- Serialize complete ESTABLISH_HWC and DESTROY_HWC shared-memory
transactions. Record possible PF ownership before MMIO submission and
clear it only after DESTROY_HWC is acknowledged.
- Retain PF-visible queues after an indeterminate destroy. Fence the
dedicated HWC EQ and withdraw its CQ from dispatch without freeing the
mappings, then retry teardown before assigning fresh queues.
- Destroy the HWC EQ before CQ and completion state. Release-publish the
CQ table and wait for IRQ-side RCU readers before freeing it, without
adding per-completion refcount operations.
- Withdraw the caller's response buffer before return and quarantine a
timed-out slot until its late response arrives or teardown cancels it.
- Return -EBUSY for healthy slot pressure and reserve -ETIMEDOUT for an
admitted request which receives no response.
- Publish slots under the inflight lock, protect response state with the
per-slot lock, and retain sender/response references until both sides
finish.
- Synchronize hwc_timeout updates and publish the HWC only after queues,
admission state, caller contexts and the direct EQ test are ready.
- Replace the goto-heavy request and channel-creation flows called out in
review with focused admission, submission, wait, build, validation,
rebuild and fallback helpers.
- Treat reported depth as a maximum: rebuild exactly for 2 through 128,
retain untouched bootstrap queues for 1 or above 128, and require the
rebuilt report to repeat the selected depth.
- Enforce directional request/response maxima against both the report
and allocated buffers. Rebuild only for exact 4096-byte reports while
permitting bootstrap operation for other nonzero maxima.
- Validate BAR page geometry and every stored doorbell. Invalidate the
old HWC doorbell before re-establishment and suppress MMIO while the
invalid sentinel is present, without adding full validation to each
data-path doorbell ring.
- Remove the proposed capability bit because driver-version
advertisement occurs after HWC establishment and cannot negotiate
creation-time behavior.
- Preserve current net-next behavior which always uses the
hardware-reported dest_vrq_id and dest_vrcq_id.
Changes in v5 (v4 -> v5):
- Shortened comments and documented the existing error paths in patches
1 and 2.
- Rejected apparent success from a response accepted before request
submission.
- Guarded CQ unpublication when establishment failed before cq_table
allocation.
- Required successful retry teardown before creating fresh bootstrap
queues after an initial destroy failure.
Changes in v4 (return to net-next after the v3 split):
- Reworked the feature series into four standalone patches for handover,
per-slot completion state, concurrency and dynamic depth.
- Added bounded admission, guarded sender accounting, timeout quarantine,
queue-depth bounds, message-buffer sizing and rebuild validation.
- Reset routing IDs for each establish and rejected a missing doorbell.
Changes in v3 (v2 -> net v3, historical fixes-only posting):
- Split the combined series after maintainer feedback, posting fixes to
net and deferring concurrency and dynamic depth.
- Added response-buffer withdrawal and stale-response handling.
- Removed pcie_flr() teardown fallback and retained resources after an
unacknowledged destroy.
Changes in v2 (v1 -> v2):
- Bounds-checked the SGE pointer derived from inline OOB size.
- Protected active-sender accounting and drained senders during
teardown and failure exits.
- Rebased on net-next e354f7d60f14.
v1:
- Combined CQ-table lifetime, queue sizing, RX validation, teardown,
concurrent-request and dynamic-depth work in seven patches.
Validation:
- All four patches pass diff-check and strict checkpatch.
- Every patch passes targeted GCC 13 and Clang 18 W=1 MANA
Ethernet/RDMA builds on base c66d93e68728.
- Sparse and Smatch add no diagnostics over that current-base baseline.
- Release kernel 7.3.0-rc4-mana-hwc-v6-4p-g7065a269bba5 booted via
one-time GRUB entries on two Azure MANA VMs.
- Each VM passed 32-worker control stress, queue resizing from 8 to 4
and back, MTU/link changes, three standalone PCI remove/rescan cycles,
and two further cycles while control queries ran concurrently.
- MANA rebound, peer connectivity recovered, and both RDMA links
returned ACTIVE after every cycle. No new MANA error, timeout, warning,
UAF, refcount, RCU, lockdep, KASAN or KCSAN report appeared.
- Fresh safe-kernel and four-patch 16-stream TCP receive throughput was
24.338 and 24.337 Gbit/s respectively. Concurrent control stress
measured 24.339 Gbit/s.
- Four-stream 64-byte UDP sustained the requested 399.986 Mbit/s on both
kernels; loss was 0.812% on the fresh safe baseline and 0.041% on the
four-patch kernel.
- RDMA devices and links remained active. ib_write_bw reproduced the
environment baseline status 21 / syndrome 0x1c6.
- The request, timeout, rebuild, fallback, message-limit and shared-
memory state machines are unchanged from the previously injected
depth/message/doorbell/destroy/response fault tests. Fresh runtime
additionally exercised the new HWC-specific CQ publication path.
- Both VMs were finally rebooted to the known-good
7.3.0-rc5-mana-vmtest-g72d3fcf802c4 kernel with the safe saved_entry,
empty next_entry, active MANA/RDMA and a 24.347 Gbit/s smoke test.
v5:
https://lore.kernel.org/all/20260908035201.402424-1-longli@microsoft.com/
v4:
https://lore.kernel.org/all/20260901200018.3194525-1-longli@microsoft.com/
v3:
https://lore.kernel.org/all/20260803234355.636038-1-longli@microsoft.com/
v2:
https://lore.kernel.org/all/20260721234339.1476932-1-longli@microsoft.com/
v1:
https://lore.kernel.org/all/20260715032942.3945317-1-longli@microsoft.com/
Long Li (4):
net: mana: prepare HWC ownership for safe reinitialization
net: mana: give each HWC message slot its own completion state
net: mana: support concurrent HWC requests
net: mana: add dynamic HWC queue depth with reinit path
.../net/ethernet/microsoft/mana/gdma_main.c | 129 ++-
.../net/ethernet/microsoft/mana/hw_channel.c | 928 +++++++++++++++---
.../net/ethernet/microsoft/mana/shm_channel.c | 58 +-
include/net/mana/gdma.h | 14 +
include/net/mana/hw_channel.h | 56 +-
include/net/mana/shm_channel.h | 10 +-
6 files changed, 1000 insertions(+), 195 deletions(-)
base-commit: c66d93e68728cfb5f40b40d0f24129d7768faf43
--
2.43.0
next reply other threads:[~2026-10-07 12:53 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 12:53 Wei Hu [this message]
2026-10-07 12:53 ` [PATCH net-next v6 1/4] net: mana: prepare HWC ownership for safe reinitialization Wei Hu
2026-10-07 12:53 ` [PATCH net-next v6 2/4] net: mana: give each HWC message slot its own completion state Wei Hu
2026-10-07 12:53 ` [PATCH net-next v6 3/4] net: mana: support concurrent HWC requests Wei Hu
2026-10-07 12:53 ` [PATCH net-next v6 4/4] net: mana: add dynamic HWC queue depth with reinit path Wei Hu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cover.1790665894.git.weh@linux.microsoft.com \
--to=weh@linux.microsoft.com \
--cc=andrew+netdev@lunn.ch \
--cc=davem@davemloft.net \
--cc=decui@microsoft.com \
--cc=dipayanroy@linux.microsoft.com \
--cc=edumazet@google.com \
--cc=ernis@linux.microsoft.com \
--cc=gargaditya@linux.microsoft.com \
--cc=haiyangz@microsoft.com \
--cc=horms@kernel.org \
--cc=jgg@ziepe.ca \
--cc=kees+treewide@kernel.org \
--cc=kees@kernel.org \
--cc=kotaranov@microsoft.com \
--cc=kuba@kernel.org \
--cc=kys@microsoft.com \
--cc=leon@kernel.org \
--cc=linux-hyperv@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=longli@kernel.org \
--cc=mawasthi@linux.microsoft.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=shirazsaleem@microsoft.com \
--cc=shradhagupta@linux.microsoft.com \
--cc=stephen@networkplumber.org \
--cc=weh@microsoft.com \
--cc=wei.liu@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®