From: Emerson Busson <emersonbusson@gmail.com>
To: mhklinux@outlook.com
Cc: kys@microsoft.com, haiyangz@microsoft.com, wei.liu@kernel.org,
decui@microsoft.com, andrew+netdev@lunn.ch, davem@davemloft.net,
edumazet@google.com, kuba@kernel.org, pabeni@redhat.com,
gregkh@linuxfoundation.org, linux-kernel@vger.kernel.org,
linux-hyperv@vger.kernel.org, netdev@vger.kernel.org
Subject: [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation
Date: Wed, 7 Oct 2026 16:07:38 -0300 [thread overview]
Message-ID: <20261007190752.336426-1-emersonbusson@gmail.com> (raw)
This series moves VMBus ring backing to the chunk allocator and keeps the
allocation, GPADL ownership, and mapping references together. It preserves
the exported four-argument interfaces while allowing the private allocation
paths to carry ownership and sharing policy.
A review of the original vzalloc fallback noted that CCA and TDX guests
without a paravisor need direct-map page transitions. This revision uses
contiguous chunks for buffers that must be host-shared, transitions each
page through page_address(), and combines the chunks with vmap(). Ordinary
guest-private buffers retain vzalloc-style allocation.
Allocation and compatibility
----------------------------
Every ring allocation uses the chunk allocator. Ordinary guests and buffers
kept private by channel policy use vzalloc(). For host-shared buffers in an
encrypted or isolated guest, the allocator uses contiguous chunks and
transitions them page by page. Guest-memory encryption is selected
independently of Hyper-V isolation.
Ring consumers select co_ring_buffer. External buffers and the legacy
allocator select co_external_memory. The exported allocator and caller-
decrypted GPADL prototypes retain four arguments throughout all fourteen
patches. UIO migration follows its release support.
The address, page and chunk arrays, and embedded GPADL belong to one buffer.
The allocation extent, host-described GPADL extent, and retained allocation
extent remain distinct. Interior GPADL duplication preserves the existing
exported interfaces. Prepared owned buffers avoid repeat decryption, and
HV_GPADL_BUFFER_DECRYPTED is removed.
Lifetime and host acknowledgement
---------------------------------
A host rescind alone does not prove GPADL revocation. For a known-live GPADL
whose channel identifier is still held, teardown posts a request and waits
for a matching GPADL_TORNDOWN message before clearing ownership. Synthetic
rescind notification cannot replace this acknowledgement. Unknown creates,
local removal, invalid identifiers, failed posting, and timeouts retain
ownership.
Waiter removal shares the response-list lock with matching replies so a late
acknowledgement cannot use a freed waiter. Released owners retain pages
across mapping references and failed page-state restoration. Reclaimer
shutdown cancels pending timers under the owner lock and requeues only work
that was actually canceled; running callbacks drain through workqueue
destruction without re-executing freed embedded work.
Sysfs mappings install no vm_ops and release the bridge pin on every setup
outcome. UIO mappings hold their pin through VMA close. Page, count, and
protection snapshots share the owner-lock boundary. New tests call the real
sysfs wrapper, insert actual PTEs, and check readable bytes and page
references through success and partial -EBUSY. They do not execute a full
kernfs syscall or full device destruction.
The final four fixes use vzalloc() for guest-private requestor arrays and
bitmaps, kvzalloc_obj() for guest-private RNDIS descriptors, handle NULL
control requests on empty NetVSC completions, and use kvzalloc_obj()/kvfree()
for the large guest-private NetVSC device object.
Qualification on the exact fourteen-patch series
------------------------------------------------
The code mails are pinned by series/SHA256SUMS. Applying all fourteen to
93f51579e7df248780214094418f205253383cc5 produces source tree
e6fe523ca92ce30ea7edb265b8a26bf20a867c47.
The VMBus workflow builds x86_64 and arm64 with W=1 and C=2, passes sparse,
and passes all 69 linked x86_64 QEMU KUnit cases. The distinct WSL backport
passes all 26 of its KUnit cases. CoCo source invariants pass on both
architectures; these are necessary static checks, not confidential-hardware
qualification.
Exact-source Hyper-V runtime, including controlled host rescind:
https://github.com/emersonbusson/WSL2-Linux-Kernel/actions/runs/37650468077
On both windows-2025 and windows-latest, the candidate passes 100 bind/unbind
cycles, bounded buddy fragmentation, channel rebind, UIO mapping teardown,
and the controlled host-origin NIC removal. Each report records
HYPERV_DRILL_HOST_RESCIND status=PASS, target_absent=yes,
matching_events=1. Both report zero splats and faults, restore the measured
kernel settings, and return lifecycle map accounting to the settled operational
baseline. Captured owner accounting reports 722 created, zero unobserved
creates, 715 reclaimed, and seven still active; active owners are not
classified as leaks. Trace capture has no overruns or dropped events, and
the disposable VM, VHD, and switch cleanup passes.
The baseline is an intentional negative control. It reproduces the expected
fragmented-allocation failure without guest faults; the workflow accepts
that exact signature as a passing control. It is not a candidate result.
WSL backport runtime matrix:
https://github.com/emersonbusson/WSL2-Linux-Kernel/actions/runs/37645708153
That separate backport passes the actual Windows candidate matrix. Each
runner completes 64 commands with matching teardown acknowledgements,
zero unacknowledged requests, no backing or tail growth, and all six
shutdown checks. The shutdown retry path is covered by 40 local parser and
refusal regressions; the live candidate runs required zero retries. These
results qualify the WSL backport's measured ordinary behavior, not the
mainline mail series.
Qualification limits
--------------------
The Hyper-V runtime uses the ordinary x86_64 vzalloc path. It does not
execute confidential chunk transitions. Full memory saturation, the
chunked CoCo fallback on hardware, SEV-SNP, TDX without a paravisor, Arm
CCA, and complete VMBus module unload and retained-owner destruction remain
unqualified. The exact mainline series has not been installed or booted on
the daily WSL host; the installed kernel and Windows WSL matrix are a separate
backport. The baseline is one intentional fragmentation negative control,
not a matched same-configuration diagnostic-count comparison for every patch.
No confidential-platform result is inferred from ordinary Hyper-V or static
checks.
---
Emerson Busson (14):
hv: vmbus: convert ring backing through the chunk allocator
hv: vmbus: validate chunk buffer allocation and cleanup
uio: hv_generic: describe buffers for owned allocation
hv: vmbus: add KUnit tests for GPADL post failure injection
hv: vmbus: add KUnit test for order-zero allocation fallback
hv: vmbus: cover all shared-page policy combinations
hv: vmbus: distinguish host rescind from local channel unload
hv: vmbus: retain backing until ownership and references clear
hv: use owned VMBus buffers in NetVSC and UIO
hv: vmbus: pin buffer pages across UIO mmap to close the reclaim race
hv: vmbus: vmalloc requestor metadata
hv: netvsc: allocate RNDIS request descriptors with kvzalloc_obj()
hv: netvsc: handle a NULL request address on empty completions
hv: netvsc: use kvzalloc for device state
next reply other threads:[~2026-10-07 19:08 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 19:07 Emerson Busson [this message]
2026-10-07 19:07 ` [PATCH v2 01/14] hv: vmbus: convert ring backing through the chunk allocator Emerson Busson
2026-10-07 19:07 ` [PATCH v2 02/14] hv: vmbus: validate chunk buffer allocation and cleanup Emerson Busson
2026-10-07 19:07 ` [PATCH v2 03/14] uio: hv_generic: describe buffers for owned allocation Emerson Busson
2026-10-07 19:07 ` [PATCH v2 04/14] hv: vmbus: add KUnit tests for GPADL post failure injection Emerson Busson
2026-10-07 19:07 ` [PATCH v2 05/14] hv: vmbus: add KUnit test for order-zero allocation fallback Emerson Busson
2026-10-07 19:07 ` [PATCH v2 06/14] hv: vmbus: cover all shared-page policy combinations Emerson Busson
2026-10-07 19:07 ` [PATCH v2 07/14] hv: vmbus: distinguish host rescind from local channel unload Emerson Busson
2026-10-07 19:07 ` [PATCH v2 08/14] hv: vmbus: retain backing until ownership and references clear Emerson Busson
2026-10-07 19:07 ` [PATCH v2 09/14] hv: use owned VMBus buffers in NetVSC and UIO Emerson Busson
2026-10-07 19:07 ` [PATCH v2 10/14] hv: vmbus: pin buffer pages across UIO mmap to close the reclaim race Emerson Busson
2026-10-07 19:07 ` [PATCH v2 11/14] hv: vmbus: vmalloc requestor metadata Emerson Busson
2026-10-07 19:07 ` [PATCH v2 12/14] hv: netvsc: allocate RNDIS request descriptors with kvzalloc_obj() Emerson Busson
2026-10-07 19:07 ` [PATCH v2 13/14] hv: netvsc: handle a NULL request address on empty completions Emerson Busson
2026-10-07 19:07 ` [PATCH v2 14/14] hv: netvsc: use kvzalloc for device state Emerson Busson
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261007190752.336426-1-emersonbusson@gmail.com \
--to=emersonbusson@gmail.com \
--cc=andrew+netdev@lunn.ch \
--cc=davem@davemloft.net \
--cc=decui@microsoft.com \
--cc=edumazet@google.com \
--cc=gregkh@linuxfoundation.org \
--cc=haiyangz@microsoft.com \
--cc=kuba@kernel.org \
--cc=kys@microsoft.com \
--cc=linux-hyperv@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mhklinux@outlook.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=wei.liu@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®