* [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation
@ 2026-10-07 19:07 Emerson Busson
2026-10-07 19:07 ` [PATCH v2 01/14] hv: vmbus: convert ring backing through the chunk allocator Emerson Busson
` (13 more replies)
0 siblings, 14 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
This series moves VMBus ring backing to the chunk allocator and keeps the
allocation, GPADL ownership, and mapping references together. It preserves
the exported four-argument interfaces while allowing the private allocation
paths to carry ownership and sharing policy.
A review of the original vzalloc fallback noted that CCA and TDX guests
without a paravisor need direct-map page transitions. This revision uses
contiguous chunks for buffers that must be host-shared, transitions each
page through page_address(), and combines the chunks with vmap(). Ordinary
guest-private buffers retain vzalloc-style allocation.
Allocation and compatibility
----------------------------
Every ring allocation uses the chunk allocator. Ordinary guests and buffers
kept private by channel policy use vzalloc(). For host-shared buffers in an
encrypted or isolated guest, the allocator uses contiguous chunks and
transitions them page by page. Guest-memory encryption is selected
independently of Hyper-V isolation.
Ring consumers select co_ring_buffer. External buffers and the legacy
allocator select co_external_memory. The exported allocator and caller-
decrypted GPADL prototypes retain four arguments throughout all fourteen
patches. UIO migration follows its release support.
The address, page and chunk arrays, and embedded GPADL belong to one buffer.
The allocation extent, host-described GPADL extent, and retained allocation
extent remain distinct. Interior GPADL duplication preserves the existing
exported interfaces. Prepared owned buffers avoid repeat decryption, and
HV_GPADL_BUFFER_DECRYPTED is removed.
Lifetime and host acknowledgement
---------------------------------
A host rescind alone does not prove GPADL revocation. For a known-live GPADL
whose channel identifier is still held, teardown posts a request and waits
for a matching GPADL_TORNDOWN message before clearing ownership. Synthetic
rescind notification cannot replace this acknowledgement. Unknown creates,
local removal, invalid identifiers, failed posting, and timeouts retain
ownership.
Waiter removal shares the response-list lock with matching replies so a late
acknowledgement cannot use a freed waiter. Released owners retain pages
across mapping references and failed page-state restoration. Reclaimer
shutdown cancels pending timers under the owner lock and requeues only work
that was actually canceled; running callbacks drain through workqueue
destruction without re-executing freed embedded work.
Sysfs mappings install no vm_ops and release the bridge pin on every setup
outcome. UIO mappings hold their pin through VMA close. Page, count, and
protection snapshots share the owner-lock boundary. New tests call the real
sysfs wrapper, insert actual PTEs, and check readable bytes and page
references through success and partial -EBUSY. They do not execute a full
kernfs syscall or full device destruction.
The final four fixes use vzalloc() for guest-private requestor arrays and
bitmaps, kvzalloc_obj() for guest-private RNDIS descriptors, handle NULL
control requests on empty NetVSC completions, and use kvzalloc_obj()/kvfree()
for the large guest-private NetVSC device object.
Qualification on the exact fourteen-patch series
------------------------------------------------
The code mails are pinned by series/SHA256SUMS. Applying all fourteen to
93f51579e7df248780214094418f205253383cc5 produces source tree
e6fe523ca92ce30ea7edb265b8a26bf20a867c47.
The VMBus workflow builds x86_64 and arm64 with W=1 and C=2, passes sparse,
and passes all 69 linked x86_64 QEMU KUnit cases. The distinct WSL backport
passes all 26 of its KUnit cases. CoCo source invariants pass on both
architectures; these are necessary static checks, not confidential-hardware
qualification.
Exact-source Hyper-V runtime, including controlled host rescind:
https://github.com/emersonbusson/WSL2-Linux-Kernel/actions/runs/37650468077
On both windows-2025 and windows-latest, the candidate passes 100 bind/unbind
cycles, bounded buddy fragmentation, channel rebind, UIO mapping teardown,
and the controlled host-origin NIC removal. Each report records
HYPERV_DRILL_HOST_RESCIND status=PASS, target_absent=yes,
matching_events=1. Both report zero splats and faults, restore the measured
kernel settings, and return lifecycle map accounting to the settled operational
baseline. Captured owner accounting reports 722 created, zero unobserved
creates, 715 reclaimed, and seven still active; active owners are not
classified as leaks. Trace capture has no overruns or dropped events, and
the disposable VM, VHD, and switch cleanup passes.
The baseline is an intentional negative control. It reproduces the expected
fragmented-allocation failure without guest faults; the workflow accepts
that exact signature as a passing control. It is not a candidate result.
WSL backport runtime matrix:
https://github.com/emersonbusson/WSL2-Linux-Kernel/actions/runs/37645708153
That separate backport passes the actual Windows candidate matrix. Each
runner completes 64 commands with matching teardown acknowledgements,
zero unacknowledged requests, no backing or tail growth, and all six
shutdown checks. The shutdown retry path is covered by 40 local parser and
refusal regressions; the live candidate runs required zero retries. These
results qualify the WSL backport's measured ordinary behavior, not the
mainline mail series.
Qualification limits
--------------------
The Hyper-V runtime uses the ordinary x86_64 vzalloc path. It does not
execute confidential chunk transitions. Full memory saturation, the
chunked CoCo fallback on hardware, SEV-SNP, TDX without a paravisor, Arm
CCA, and complete VMBus module unload and retained-owner destruction remain
unqualified. The exact mainline series has not been installed or booted on
the daily WSL host; the installed kernel and Windows WSL matrix are a separate
backport. The baseline is one intentional fragmentation negative control,
not a matched same-configuration diagnostic-count comparison for every patch.
No confidential-platform result is inferred from ordinary Hyper-V or static
checks.
---
Emerson Busson (14):
hv: vmbus: convert ring backing through the chunk allocator
hv: vmbus: validate chunk buffer allocation and cleanup
uio: hv_generic: describe buffers for owned allocation
hv: vmbus: add KUnit tests for GPADL post failure injection
hv: vmbus: add KUnit test for order-zero allocation fallback
hv: vmbus: cover all shared-page policy combinations
hv: vmbus: distinguish host rescind from local channel unload
hv: vmbus: retain backing until ownership and references clear
hv: use owned VMBus buffers in NetVSC and UIO
hv: vmbus: pin buffer pages across UIO mmap to close the reclaim race
hv: vmbus: vmalloc requestor metadata
hv: netvsc: allocate RNDIS request descriptors with kvzalloc_obj()
hv: netvsc: handle a NULL request address on empty completions
hv: netvsc: use kvzalloc for device state
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 01/14] hv: vmbus: convert ring backing through the chunk allocator
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 02/14] hv: vmbus: validate chunk buffer allocation and cleanup Emerson Busson
` (12 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Use the existing virtually contiguous allocator for ring backing with
the ring confidentiality policy. An internal helper keeps that choice
distinct from the established external-buffer compatibility policy.
Preserve the exported allocator and caller-decrypted GPADL signatures
throughout the series; carry uncertain creation in the descriptor
instead of widening an exported function. The independent
guest-memory-encryption condition is present at the first ring consumer.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 176 +++++++++++++++++++++-----------
drivers/hv/hyperv_vmbus.h | 4 +-
drivers/hv/ring_buffer.c | 9 +-
drivers/net/hyperv/hyperv_net.h | 10 +-
drivers/net/hyperv/netvsc.c | 58 ++++++-----
drivers/uio/uio_hv_generic.c | 37 ++++---
include/linux/hyperv.h | 18 +++-
7 files changed, 196 insertions(+), 116 deletions(-)
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index 7e4cc6f55237..f8feb2a0ec0a 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -12,6 +12,8 @@
#include <linux/sched.h>
#include <linux/wait.h>
#include <linux/mm.h>
+#include <linux/cc_platform.h>
+#include <linux/overflow.h>
#include <linux/slab.h>
#include <linux/log2.h>
#include <linux/module.h>
@@ -26,6 +28,12 @@
#include "hyperv_vmbus.h"
+static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
+ u32 size,
+ bool confidential,
+ struct page ***chunks_out,
+ u32 *chunk_cnt_out);
+
/*
* hv_gpadl_size - Return the real size of a gpadl, the size that Hyper-V uses
*
@@ -42,7 +50,6 @@ static inline u32 hv_gpadl_size(enum hv_gpadl_type type, u32 size)
{
switch (type) {
case HV_GPADL_BUFFER:
- case HV_GPADL_BUFFER_DECRYPTED:
return size;
case HV_GPADL_RING:
/* The size of a ringbuffer must be page-aligned */
@@ -103,7 +110,6 @@ static inline u64 hv_gpadl_hvpfn(enum hv_gpadl_type type, void *kbuffer,
switch (type) {
case HV_GPADL_BUFFER:
- case HV_GPADL_BUFFER_DECRYPTED:
break;
case HV_GPADL_RING:
if (i == 0)
@@ -154,17 +160,15 @@ EXPORT_SYMBOL_GPL(vmbus_setevent);
/* vmbus_free_ring - drop mapping of ring buffer */
void vmbus_free_ring(struct vmbus_channel *channel)
{
+ struct vmbus_buffer *buffer = &channel->ringbuffer;
+
hv_ringbuffer_cleanup(&channel->outbound);
hv_ringbuffer_cleanup(&channel->inbound);
- if (channel->ringbuffer_page) {
- /* In a CoCo VM leak the memory if it didn't get re-encrypted */
- if (!channel->ringbuffer_gpadlhandle.decrypted)
- __free_pages(channel->ringbuffer_page,
- get_order(channel->ringbuffer_pagecount
- << PAGE_SHIFT));
- channel->ringbuffer_page = NULL;
- }
+ if (!buffer->addr)
+ return;
+
+ vmbus_release_buffer(buffer);
}
EXPORT_SYMBOL_GPL(vmbus_free_ring);
@@ -172,26 +176,32 @@ EXPORT_SYMBOL_GPL(vmbus_free_ring);
int vmbus_alloc_ring(struct vmbus_channel *newchannel,
u32 send_size, u32 recv_size)
{
- struct page *page;
- int order;
+ struct vmbus_buffer *buffer = &newchannel->ringbuffer;
+ u32 size;
+ u32 i;
- if (send_size % PAGE_SIZE || recv_size % PAGE_SIZE)
+ if (!send_size || !recv_size ||
+ send_size % PAGE_SIZE || recv_size % PAGE_SIZE ||
+ check_add_overflow(send_size, recv_size, &size))
return -EINVAL;
- /* Allocate the ring buffer */
- order = get_order(send_size + recv_size);
- page = alloc_pages_node(cpu_to_node(newchannel->target_cpu),
- GFP_KERNEL|__GFP_ZERO, order);
-
- if (!page)
- page = alloc_pages(GFP_KERNEL|__GFP_ZERO, order);
-
- if (!page)
+ buffer->addr = __vmbus_alloc_buffer(newchannel, size,
+ newchannel->co_ring_buffer,
+ &buffer->chunks, &buffer->chunk_cnt);
+ if (!buffer->addr)
return -ENOMEM;
- newchannel->ringbuffer_page = page;
- newchannel->ringbuffer_pagecount = (send_size + recv_size) >> PAGE_SHIFT;
+ newchannel->ringbuffer_pagecount = size >> PAGE_SHIFT;
newchannel->ringbuffer_send_offset = send_size >> PAGE_SHIFT;
+ buffer->pages = kvcalloc(newchannel->ringbuffer_pagecount,
+ sizeof(*buffer->pages), GFP_KERNEL);
+ if (!buffer->pages) {
+ vmbus_release_buffer(buffer);
+ return -ENOMEM;
+ }
+
+ for (i = 0; i < newchannel->ringbuffer_pagecount; i++)
+ buffer->pages[i] = vmalloc_to_page(buffer->addr + (i << PAGE_SHIFT));
return 0;
}
@@ -442,7 +452,8 @@ static void vmbus_free_channel_msginfo(struct vmbus_channel_msginfo *msginfo)
*/
static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
enum hv_gpadl_type type, void *kbuffer,
- u32 size, u32 send_offset,
+ u32 size, u32 send_offset, bool memory_prepared,
+ bool *leak,
struct vmbus_gpadl *gpadl)
{
struct vmbus_channel_gpadl_header *gpadlmsg;
@@ -452,8 +463,13 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
struct list_head *curr;
u32 next_gpadl_handle;
unsigned long flags;
+ bool posted = false;
int ret = 0;
+ if (leak)
+ *leak = false;
+ gpadl->leak = false;
+
next_gpadl_handle =
(atomic_inc_return(&vmbus_connection.next_gpadl_handle) - 1);
@@ -463,9 +479,9 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
return ret;
}
- gpadl->decrypted = !((channel->co_external_memory && type == HV_GPADL_BUFFER) ||
- (channel->co_ring_buffer && type == HV_GPADL_RING) ||
- (type == HV_GPADL_BUFFER_DECRYPTED));
+ gpadl->decrypted = !memory_prepared &&
+ !((channel->co_external_memory && type == HV_GPADL_BUFFER) ||
+ (channel->co_ring_buffer && type == HV_GPADL_RING));
if (gpadl->decrypted) {
/*
* The "decrypted" flag being true assumes that set_memory_decrypted() succeeds.
@@ -504,6 +520,8 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
goto cleanup;
}
+ /* A failed post may still have reached the host. */
+ posted = true;
ret = vmbus_post_msg(gpadlmsg, msginfo->msgsize -
sizeof(*msginfo), true);
@@ -534,6 +552,7 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
wait_for_completion(&msginfo->waitevent);
if (msginfo->response.gpadl_created.creation_status != 0) {
+ posted = false;
pr_err("Failed to establish GPADL: err = 0x%x\n",
msginfo->response.gpadl_created.creation_status);
@@ -542,6 +561,7 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
}
if (channel->rescind) {
+ posted = false;
ret = -ENODEV;
goto cleanup;
}
@@ -550,6 +570,7 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
gpadl->gpadl_handle = gpadlmsg->gpadl;
gpadl->buffer = kbuffer;
gpadl->size = size;
+ posted = false;
cleanup:
@@ -559,7 +580,13 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
vmbus_free_channel_msginfo(msginfo);
- if (ret) {
+ if (ret && posted) {
+ gpadl->leak = true;
+ if (leak)
+ *leak = true;
+ }
+
+ if (ret && !posted) {
/*
* If set_memory_encrypted() fails, the decrypted flag is
* left as true so the memory is leaked instead of being
@@ -586,7 +613,7 @@ int vmbus_establish_gpadl(struct vmbus_channel *channel, void *kbuffer,
u32 size, struct vmbus_gpadl *gpadl)
{
return __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER, kbuffer, size,
- 0U, gpadl);
+ 0U, false, &gpadl->leak, gpadl);
}
EXPORT_SYMBOL_GPL(vmbus_establish_gpadl);
@@ -597,6 +624,8 @@ EXPORT_SYMBOL_GPL(vmbus_establish_gpadl);
* @channel: a channel
* @kbuffer: from kmalloc or vmalloc; must already be decrypted by the caller
* @size: page-size multiple
+ * @leak: set when a GPADL message may have reached the host but completion is
+ * uncertain; the caller must retain the backing pages
* @gpadl: output gpadl
*
* The caller is responsible for re-encrypting the buffer before freeing it.
@@ -605,8 +634,8 @@ int vmbus_establish_gpadl_caller_decrypted(struct vmbus_channel *channel,
void *kbuffer, u32 size,
struct vmbus_gpadl *gpadl)
{
- return __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER_DECRYPTED,
- kbuffer, size, 0U, gpadl);
+ return __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER,
+ kbuffer, size, 0U, true, &gpadl->leak, gpadl);
}
EXPORT_SYMBOL_GPL(vmbus_establish_gpadl_caller_decrypted);
@@ -648,11 +677,26 @@ void vmbus_free_buffer(void *addr, struct page **chunks, u32 chunk_cnt)
}
EXPORT_SYMBOL_GPL(vmbus_free_buffer);
+void vmbus_release_buffer(struct vmbus_buffer *buffer)
+{
+ if (!buffer->addr)
+ return;
+
+ kvfree(buffer->pages);
+ if (!buffer->leak && !buffer->gpadl.leak &&
+ !buffer->gpadl.gpadl_handle)
+ vmbus_free_buffer(buffer->addr, buffer->chunks,
+ buffer->chunk_cnt);
+ memset(buffer, 0, sizeof(*buffer));
+}
+EXPORT_SYMBOL_GPL(vmbus_release_buffer);
+
/**
- * vmbus_alloc_buffer - allocate a host-visible, virtually-contiguous buffer.
+ * __vmbus_alloc_buffer - allocate host-visible, virtually-contiguous backing.
*
* @channel: the channel the buffer will be attached to
* @size: requested buffer size in bytes (will be rounded up to PAGE_SIZE)
+ * @confidential: keep the buffer private to the guest
* @chunks_out: on success, set to the array of underlying chunks, or NULL when
* the buffer was allocated with vzalloc()
* @chunk_cnt_out: on success, set to the number of chunks
@@ -667,10 +711,11 @@ EXPORT_SYMBOL_GPL(vmbus_free_buffer);
*
* Return: the buffer's virtual address, or NULL on failure.
*/
-void *vmbus_alloc_buffer(struct vmbus_channel *channel,
- u32 size,
- struct page ***chunks_out,
- u32 *chunk_cnt_out)
+static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
+ u32 size,
+ bool confidential,
+ struct page ***chunks_out,
+ u32 *chunk_cnt_out)
{
unsigned long nr_pages = PFN_UP(size);
unsigned long remaining = nr_pages;
@@ -690,7 +735,8 @@ void *vmbus_alloc_buffer(struct vmbus_channel *channel,
return NULL;
/* If the buffer does not need to be decrypted, just use vzalloc() */
- if (!hv_is_isolation_supported() || channel->co_external_memory)
+ if ((!hv_is_isolation_supported() &&
+ !cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT)) || confidential)
return vzalloc(nr_pages << PAGE_SHIFT);
/* Worst case: every chunk is a single page. */
@@ -760,6 +806,14 @@ void *vmbus_alloc_buffer(struct vmbus_channel *channel,
vmbus_free_buffer(NULL, chunks, chunk_cnt);
return NULL;
}
+
+void *vmbus_alloc_buffer(struct vmbus_channel *channel,
+ u32 size, struct page ***chunks_out,
+ u32 *chunk_cnt_out)
+{
+ return __vmbus_alloc_buffer(channel, size, channel->co_external_memory,
+ chunks_out, chunk_cnt_out);
+}
EXPORT_SYMBOL_GPL(vmbus_alloc_buffer);
/**
@@ -832,7 +886,7 @@ static int __vmbus_open(struct vmbus_channel *newchannel,
{
struct vmbus_channel_open_channel *open_msg;
struct vmbus_channel_msginfo *open_info = NULL;
- struct page *page = newchannel->ringbuffer_page;
+ struct vmbus_buffer *buffer = &newchannel->ringbuffer;
u32 send_pages, recv_pages;
unsigned long flags;
int err;
@@ -860,22 +914,24 @@ static int __vmbus_open(struct vmbus_channel *newchannel,
newchannel->max_pkt_size = VMBUS_DEFAULT_MAX_PKT_SIZE;
/* Establish the gpadl for the ring buffer */
- newchannel->ringbuffer_gpadlhandle.gpadl_handle = 0;
+ buffer->gpadl.gpadl_handle = 0;
err = __vmbus_establish_gpadl(newchannel, HV_GPADL_RING,
- page_address(newchannel->ringbuffer_page),
+ buffer->addr,
(send_pages + recv_pages) << PAGE_SHIFT,
newchannel->ringbuffer_send_offset << PAGE_SHIFT,
- &newchannel->ringbuffer_gpadlhandle);
+ true, &buffer->leak, &buffer->gpadl);
if (err)
goto error_clean_ring;
err = hv_ringbuffer_init(&newchannel->outbound,
- page, send_pages, 0, newchannel->co_ring_buffer);
+ buffer->addr, send_pages, 0,
+ newchannel->co_ring_buffer);
if (err)
goto error_free_gpadl;
- err = hv_ringbuffer_init(&newchannel->inbound, &page[send_pages],
+ err = hv_ringbuffer_init(&newchannel->inbound,
+ buffer->addr + (send_pages << PAGE_SHIFT),
recv_pages, newchannel->max_pkt_size,
newchannel->co_ring_buffer);
if (err)
@@ -897,8 +953,7 @@ static int __vmbus_open(struct vmbus_channel *newchannel,
open_msg->header.msgtype = CHANNELMSG_OPENCHANNEL;
open_msg->openid = newchannel->offermsg.child_relid;
open_msg->child_relid = newchannel->offermsg.child_relid;
- open_msg->ringbuffer_gpadlhandle
- = newchannel->ringbuffer_gpadlhandle.gpadl_handle;
+ open_msg->ringbuffer_gpadlhandle = buffer->gpadl.gpadl_handle;
/*
* The unit of ->downstream_ringbuffer_pageoffset is HV_HYP_PAGE and
* the unit of ->ringbuffer_send_offset (i.e. send_pages) is PAGE, so
@@ -956,7 +1011,8 @@ static int __vmbus_open(struct vmbus_channel *newchannel,
error_free_info:
kfree(open_info);
error_free_gpadl:
- vmbus_teardown_gpadl(newchannel, &newchannel->ringbuffer_gpadlhandle);
+ if (vmbus_teardown_gpadl(newchannel, &buffer->gpadl))
+ buffer->leak = true;
error_clean_ring:
hv_ringbuffer_cleanup(&newchannel->outbound);
hv_ringbuffer_cleanup(&newchannel->inbound);
@@ -1058,15 +1114,20 @@ int vmbus_teardown_gpadl(struct vmbus_channel *channel, struct vmbus_gpadl *gpad
kfree(info);
- if (gpadl->decrypted)
- ret = set_memory_encrypted((unsigned long)gpadl->buffer,
- PFN_UP(gpadl->size));
- else
- ret = 0;
- if (ret)
- pr_warn("Fail to set mem host visibility in GPADL teardown %d.\n", ret);
+ if (!ret && gpadl->decrypted) {
+ int encrypt_ret;
- gpadl->decrypted = ret;
+ encrypt_ret = set_memory_encrypted((unsigned long)gpadl->buffer,
+ PFN_UP(gpadl->size));
+ if (encrypt_ret) {
+ pr_warn("Failed to re-encrypt GPADL buffer: %d\n",
+ encrypt_ret);
+ ret = encrypt_ret;
+ }
+ gpadl->decrypted = !!encrypt_ret;
+ }
+ if (ret)
+ gpadl->leak = true;
return ret;
}
@@ -1142,9 +1203,10 @@ static int vmbus_close_internal(struct vmbus_channel *channel)
}
/* Tear down the gpadl for the channel's ring buffer */
- else if (channel->ringbuffer_gpadlhandle.gpadl_handle) {
- ret = vmbus_teardown_gpadl(channel, &channel->ringbuffer_gpadlhandle);
+ else if (channel->ringbuffer.gpadl.gpadl_handle) {
+ ret = vmbus_teardown_gpadl(channel, &channel->ringbuffer.gpadl);
if (ret) {
+ channel->ringbuffer.leak = true;
pr_err("Close failed: teardown gpadl return %d\n", ret);
/*
* If we failed to teardown gpadl,
diff --git a/drivers/hv/hyperv_vmbus.h b/drivers/hv/hyperv_vmbus.h
index 33923621a5a3..20d023c9735e 100644
--- a/drivers/hv/hyperv_vmbus.h
+++ b/drivers/hv/hyperv_vmbus.h
@@ -204,8 +204,8 @@ extern int hv_synic_cleanup(unsigned int cpu);
void hv_ringbuffer_pre_init(struct vmbus_channel *channel);
int hv_ringbuffer_init(struct hv_ring_buffer_info *ring_info,
- struct page *pages, u32 pagecnt, u32 max_pkt_size,
- bool confidential);
+ void *addr, u32 pagecnt, u32 max_pkt_size,
+ bool confidential);
void hv_ringbuffer_cleanup(struct hv_ring_buffer_info *ring_info);
diff --git a/drivers/hv/ring_buffer.c b/drivers/hv/ring_buffer.c
index 592a9601faaa..16b1c2910789 100644
--- a/drivers/hv/ring_buffer.c
+++ b/drivers/hv/ring_buffer.c
@@ -184,8 +184,8 @@ void hv_ringbuffer_pre_init(struct vmbus_channel *channel)
/* Initialize the ring buffer. */
int hv_ringbuffer_init(struct hv_ring_buffer_info *ring_info,
- struct page *pages, u32 page_cnt, u32 max_pkt_size,
- bool confidential)
+ void *addr, u32 page_cnt, u32 max_pkt_size,
+ bool confidential)
{
struct page **pages_wraparound;
int i;
@@ -200,10 +200,11 @@ int hv_ringbuffer_init(struct hv_ring_buffer_info *ring_info,
if (!pages_wraparound)
return -ENOMEM;
- pages_wraparound[0] = pages;
+ pages_wraparound[0] = vmalloc_to_page(addr);
for (i = 0; i < 2 * (page_cnt - 1); i++)
pages_wraparound[i + 1] =
- &pages[i % (page_cnt - 1) + 1];
+ vmalloc_to_page(addr +
+ ((i % (page_cnt - 1) + 1) << PAGE_SHIFT));
ring_info->ring_buffer = (struct hv_ring_buffer *)
vmap(pages_wraparound, page_cnt * 2 - 1, VM_MAP,
diff --git a/drivers/net/hyperv/hyperv_net.h b/drivers/net/hyperv/hyperv_net.h
index 4841367fdab2..a15cb2460344 100644
--- a/drivers/net/hyperv/hyperv_net.h
+++ b/drivers/net/hyperv/hyperv_net.h
@@ -1158,21 +1158,15 @@ struct netvsc_device {
bool tx_disable; /* if true, do not wake up queue again */
/* Receive buffer allocated by us but manages by NetVSP */
- void *recv_buf;
+ struct vmbus_buffer recv_buffer;
u32 recv_buf_size; /* allocated bytes */
- struct page **recv_buf_chunks;
- u32 recv_buf_chunk_cnt;
- struct vmbus_gpadl recv_buf_gpadl_handle;
u32 recv_section_cnt;
u32 recv_section_size;
u32 recv_completion_cnt;
/* Send buffer allocated by us */
- void *send_buf;
+ struct vmbus_buffer send_buffer;
u32 send_buf_size;
- struct page **send_buf_chunks;
- u32 send_buf_chunk_cnt;
- struct vmbus_gpadl send_buf_gpadl_handle;
u32 send_section_cnt;
u32 send_section_size;
unsigned long *send_section_map;
diff --git a/drivers/net/hyperv/netvsc.c b/drivers/net/hyperv/netvsc.c
index 5cd084e5696c..ffba1443396a 100644
--- a/drivers/net/hyperv/netvsc.c
+++ b/drivers/net/hyperv/netvsc.c
@@ -134,10 +134,8 @@ static void __free_netvsc_device(struct netvsc_device *nvdev)
kfree(nvdev->extension);
- vmbus_free_buffer(nvdev->recv_buf, nvdev->recv_buf_chunks,
- nvdev->recv_buf_chunk_cnt);
- vmbus_free_buffer(nvdev->send_buf, nvdev->send_buf_chunks,
- nvdev->send_buf_chunk_cnt);
+ vmbus_release_buffer(&nvdev->recv_buffer);
+ vmbus_release_buffer(&nvdev->send_buffer);
bitmap_free(nvdev->send_section_map);
for (i = 0; i < VRSS_CHANNEL_MAX; i++) {
@@ -245,6 +243,7 @@ static void netvsc_revoke_recv_buf(struct hv_device *device,
if (ret != 0) {
netdev_err(ndev, "unable to send "
"revoke receive buffer to netvsp\n");
+ net_device->recv_buffer.leak = true;
return;
}
net_device->recv_section_cnt = 0;
@@ -296,6 +295,7 @@ static void netvsc_revoke_send_buf(struct hv_device *device,
if (ret != 0) {
netdev_err(ndev, "unable to send "
"revoke send buffer to netvsp\n");
+ net_device->send_buffer.leak = true;
return;
}
net_device->send_section_cnt = 0;
@@ -308,14 +308,18 @@ static void netvsc_teardown_recv_gpadl(struct hv_device *device,
{
int ret;
- if (net_device->recv_buf_gpadl_handle.gpadl_handle) {
+ if (net_device->recv_buffer.leak)
+ return;
+
+ if (net_device->recv_buffer.gpadl.gpadl_handle) {
ret = vmbus_teardown_gpadl(device->channel,
- &net_device->recv_buf_gpadl_handle);
+ &net_device->recv_buffer.gpadl);
/* If we failed here, we might as well return and have a leak
* rather than continue and a bugchk
*/
if (ret != 0) {
+ net_device->recv_buffer.leak = true;
netdev_err(ndev,
"unable to teardown receive buffer's gpadl\n");
return;
@@ -329,14 +333,18 @@ static void netvsc_teardown_send_gpadl(struct hv_device *device,
{
int ret;
- if (net_device->send_buf_gpadl_handle.gpadl_handle) {
+ if (net_device->send_buffer.leak)
+ return;
+
+ if (net_device->send_buffer.gpadl.gpadl_handle) {
ret = vmbus_teardown_gpadl(device->channel,
- &net_device->send_buf_gpadl_handle);
+ &net_device->send_buffer.gpadl);
/* If we failed here, we might as well return and have a leak
* rather than continue and a bugchk
*/
if (ret != 0) {
+ net_device->send_buffer.leak = true;
netdev_err(ndev,
"unable to teardown send buffer's gpadl\n");
return;
@@ -377,11 +385,11 @@ static int netvsc_init_buf(struct hv_device *device,
buf_size = min_t(unsigned int, buf_size,
NETVSC_RECEIVE_BUFFER_SIZE_LEGACY);
- net_device->recv_buf =
+ net_device->recv_buffer.addr =
vmbus_alloc_buffer(device->channel, buf_size,
- &net_device->recv_buf_chunks,
- &net_device->recv_buf_chunk_cnt);
- if (!net_device->recv_buf) {
+ &net_device->recv_buffer.chunks,
+ &net_device->recv_buffer.chunk_cnt);
+ if (!net_device->recv_buffer.addr) {
netdev_err(ndev,
"unable to allocate receive buffer of size %u\n",
buf_size);
@@ -397,9 +405,10 @@ static int netvsc_init_buf(struct hv_device *device,
* than the channel to establish the gpadl handle.
*/
ret = vmbus_establish_gpadl_caller_decrypted(device->channel,
- net_device->recv_buf,
+ net_device->recv_buffer.addr,
buf_size,
- &net_device->recv_buf_gpadl_handle);
+ &net_device->recv_buffer.gpadl);
+ net_device->recv_buffer.leak |= net_device->recv_buffer.gpadl.leak;
if (ret != 0) {
netdev_err(ndev,
"unable to establish receive buffer's gpadl\n");
@@ -411,7 +420,7 @@ static int netvsc_init_buf(struct hv_device *device,
memset(init_packet, 0, sizeof(struct nvsp_message));
init_packet->hdr.msg_type = NVSP_MSG1_TYPE_SEND_RECV_BUF;
init_packet->msg.v1_msg.send_recv_buf.
- gpadl_handle = net_device->recv_buf_gpadl_handle.gpadl_handle;
+ gpadl_handle = net_device->recv_buffer.gpadl.gpadl_handle;
init_packet->msg.v1_msg.
send_recv_buf.id = NETVSC_RECEIVE_BUFFER_ID;
@@ -487,11 +496,11 @@ static int netvsc_init_buf(struct hv_device *device,
buf_size = device_info->send_sections * device_info->send_section_size;
buf_size = round_up(buf_size, PAGE_SIZE);
- net_device->send_buf =
+ net_device->send_buffer.addr =
vmbus_alloc_buffer(device->channel, buf_size,
- &net_device->send_buf_chunks,
- &net_device->send_buf_chunk_cnt);
- if (!net_device->send_buf) {
+ &net_device->send_buffer.chunks,
+ &net_device->send_buffer.chunk_cnt);
+ if (!net_device->send_buffer.addr) {
netdev_err(ndev, "unable to allocate send buffer of size %u\n",
buf_size);
ret = -ENOMEM;
@@ -504,9 +513,10 @@ static int netvsc_init_buf(struct hv_device *device,
* than the channel to establish the gpadl handle.
*/
ret = vmbus_establish_gpadl_caller_decrypted(device->channel,
- net_device->send_buf,
+ net_device->send_buffer.addr,
buf_size,
- &net_device->send_buf_gpadl_handle);
+ &net_device->send_buffer.gpadl);
+ net_device->send_buffer.leak |= net_device->send_buffer.gpadl.leak;
if (ret != 0) {
netdev_err(ndev,
"unable to establish send buffer's gpadl\n");
@@ -518,7 +528,7 @@ static int netvsc_init_buf(struct hv_device *device,
memset(init_packet, 0, sizeof(struct nvsp_message));
init_packet->hdr.msg_type = NVSP_MSG1_TYPE_SEND_SEND_BUF;
init_packet->msg.v1_msg.send_send_buf.gpadl_handle =
- net_device->send_buf_gpadl_handle.gpadl_handle;
+ net_device->send_buffer.gpadl.gpadl_handle;
init_packet->msg.v1_msg.send_send_buf.id = NETVSC_SEND_BUFFER_ID;
trace_nvsp_send(ndev, init_packet);
@@ -968,7 +978,7 @@ static void netvsc_copy_to_send_buf(struct netvsc_device *net_device,
struct hv_page_buffer *pb,
bool xmit_more)
{
- char *start = net_device->send_buf;
+ char *start = net_device->send_buffer.addr;
char *dest = start + (section_index * net_device->send_section_size)
+ pend_size;
int i;
@@ -1475,7 +1485,7 @@ static int netvsc_receive(struct net_device *ndev,
const struct nvsp_message *nvsp = hv_pkt_data(desc);
u32 msglen = hv_pkt_datalen(desc);
u16 q_idx = channel->offermsg.offer.sub_channel_index;
- char *recv_buf = net_device->recv_buf;
+ char *recv_buf = net_device->recv_buffer.addr;
u32 status = NVSP_STAT_SUCCESS;
int i;
int count = 0;
diff --git a/drivers/uio/uio_hv_generic.c b/drivers/uio/uio_hv_generic.c
index 7b4cc456c453..b3f41ffc74f8 100644
--- a/drivers/uio/uio_hv_generic.c
+++ b/drivers/uio/uio_hv_generic.c
@@ -150,19 +150,21 @@ static void hv_uio_rescind(struct vmbus_channel *channel)
vmbus_device_unregister(channel->device_obj);
}
-/* Function used for mmap of ring buffer sysfs interface.
- * The ring buffer is allocated as contiguous memory by vmbus_open
- */
+/* Function used for mmap of the ring buffer sysfs interface. */
static int
hv_uio_ring_mmap_prepare(struct vmbus_channel *channel, struct vm_area_desc *desc)
{
- void *ring_buffer = page_address(channel->ringbuffer_page);
+ unsigned long pages = vma_desc_pages(desc);
+ pgoff_t offset = desc->pgoff;
if (channel->state != CHANNEL_OPENED_STATE)
return -ENODEV;
+ if (offset >= channel->ringbuffer_pagecount ||
+ pages > channel->ringbuffer_pagecount - offset)
+ return -EINVAL;
- mmap_action_simple_ioremap(desc, virt_to_phys(ring_buffer),
- channel->ringbuffer_pagecount << PAGE_SHIFT);
+ mmap_action_map_kernel_pages(desc, desc->start,
+ channel->ringbuffer.pages + offset, pages);
return 0;
}
@@ -196,14 +198,16 @@ static void
hv_uio_cleanup(struct hv_device *dev, struct hv_uio_private_data *pdata)
{
if (pdata->send_gpadl.gpadl_handle) {
- vmbus_teardown_gpadl(dev->channel, &pdata->send_gpadl);
- if (!pdata->send_gpadl.decrypted)
+ if (vmbus_teardown_gpadl(dev->channel, &pdata->send_gpadl))
+ pdata->send_gpadl.leak = true;
+ if (!pdata->send_gpadl.leak && !pdata->send_gpadl.decrypted)
vfree(pdata->send_buf);
}
if (pdata->recv_gpadl.gpadl_handle) {
- vmbus_teardown_gpadl(dev->channel, &pdata->recv_gpadl);
- if (!pdata->recv_gpadl.decrypted)
+ if (vmbus_teardown_gpadl(dev->channel, &pdata->recv_gpadl))
+ pdata->recv_gpadl.leak = true;
+ if (!pdata->recv_gpadl.leak && !pdata->recv_gpadl.decrypted)
vfree(pdata->recv_buf);
}
}
@@ -283,12 +287,11 @@ hv_uio_probe(struct hv_device *dev,
/* mem resources */
pdata->info.mem[TXRX_RING_MAP].name = "txrx_rings";
- ring_buffer = page_address(channel->ringbuffer_page);
- pdata->info.mem[TXRX_RING_MAP].addr
- = (uintptr_t)virt_to_phys(ring_buffer);
+ ring_buffer = channel->ringbuffer.addr;
+ pdata->info.mem[TXRX_RING_MAP].addr = (uintptr_t)ring_buffer;
pdata->info.mem[TXRX_RING_MAP].size
= channel->ringbuffer_pagecount << PAGE_SHIFT;
- pdata->info.mem[TXRX_RING_MAP].memtype = UIO_MEM_IOVA;
+ pdata->info.mem[TXRX_RING_MAP].memtype = UIO_MEM_VIRTUAL;
pdata->info.mem[INT_PAGE_MAP].name = "int_page";
pdata->info.mem[INT_PAGE_MAP].addr
@@ -312,7 +315,8 @@ hv_uio_probe(struct hv_device *dev,
ret = vmbus_establish_gpadl(channel, pdata->recv_buf,
RECV_BUFFER_SIZE, &pdata->recv_gpadl);
if (ret) {
- if (!pdata->recv_gpadl.decrypted)
+ if (!pdata->recv_gpadl.leak &&
+ !pdata->recv_gpadl.decrypted)
vfree(pdata->recv_buf);
goto fail_close;
}
@@ -334,7 +338,8 @@ hv_uio_probe(struct hv_device *dev,
ret = vmbus_establish_gpadl(channel, pdata->send_buf,
SEND_BUFFER_SIZE, &pdata->send_gpadl);
if (ret) {
- if (!pdata->send_gpadl.decrypted)
+ if (!pdata->send_gpadl.leak &&
+ !pdata->send_gpadl.decrypted)
vfree(pdata->send_buf);
goto fail_close;
}
diff --git a/include/linux/hyperv.h b/include/linux/hyperv.h
index 9e109d91aa14..2878aed14c45 100644
--- a/include/linux/hyperv.h
+++ b/include/linux/hyperv.h
@@ -70,8 +70,7 @@
*/
enum hv_gpadl_type {
HV_GPADL_BUFFER,
- HV_GPADL_RING,
- HV_GPADL_BUFFER_DECRYPTED
+ HV_GPADL_RING
};
/* Single-page buffer */
@@ -782,6 +781,16 @@ struct vmbus_gpadl {
u32 size;
void *buffer;
bool decrypted;
+ bool leak;
+};
+
+struct vmbus_buffer {
+ void *addr;
+ struct page **chunks;
+ struct page **pages;
+ u32 chunk_cnt;
+ struct vmbus_gpadl gpadl;
+ bool leak;
};
struct vmbus_channel {
@@ -803,10 +812,8 @@ struct vmbus_channel {
bool rescind_ref; /* got rescind msg, got channel reference */
struct completion rescind_event;
- struct vmbus_gpadl ringbuffer_gpadlhandle;
-
/* Allocated memory for ring buffer */
- struct page *ringbuffer_page;
+ struct vmbus_buffer ringbuffer;
u32 ringbuffer_pagecount;
u32 ringbuffer_send_offset;
struct hv_ring_buffer_info outbound; /* send to parent */
@@ -1219,6 +1226,7 @@ extern void *vmbus_alloc_buffer(struct vmbus_channel *channel,
u32 *chunk_cnt_out);
extern void vmbus_free_buffer(void *addr, struct page **chunks, u32 chunk_cnt);
+void vmbus_release_buffer(struct vmbus_buffer *buffer);
void vmbus_reset_channel_cb(struct vmbus_channel *channel);
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 02/14] hv: vmbus: validate chunk buffer allocation and cleanup
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
2026-10-07 19:07 ` [PATCH v2 01/14] hv: vmbus: convert ring backing through the chunk allocator Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 03/14] uio: hv_generic: describe buffers for owned allocation Emerson Busson
` (11 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Validate rounded sizes before storing them and preserve allocated
chunks when cleanup or GPADL teardown has an uncertain outcome. Add
focused KUnit coverage for buffer sizing, ownership decisions, and the
ring fallback order descent.
vmbus_alloc_buffer() rounds the requested size up to PAGE_SIZE before it
computes nr_pages. A zero size, and a size within PAGE_SIZE of U32_MAX,
both produce a rounded value that is wrong: zero allocates nothing, and
the near-U32_MAX cases wrap to a small nr_pages, so a small chunk array
is built for a buffer the host is told is large. Reject both before the
rounded size is stored or used.
vmbus_buffer_should_free() gathers the leak decision in one place. A
cleanup whose GPADL create or teardown message may have reached the host
cannot prove the host has released the pages. Handing those pages back
to the allocator, or re-encrypting them, while the host may still map
them is a use-after-free from the host's side and a Confidential
Computing hazard. Such a buffer is marked leaked and its chunks are
retained instead of freed.
The new cases live in drivers/hv/vmbus_buffer_test.c, which is built
into the hv_vmbus object rather than a separate module. That keeps the
cases off the production sources while still letting them call the
internal sizing, order-descent and free-decision helpers. Those helpers
are therefore not static; nothing outside hv_vmbus links against them,
and no symbol is exported for test purposes.
The cases pin both rules, the partial-allocation rollback, and the
order-descent fallback.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/Kconfig | 11 ++++
drivers/hv/Makefile | 1 +
drivers/hv/channel.c | 67 ++++++++++++++++++----
drivers/hv/hyperv_vmbus.h | 12 ++++
drivers/hv/vmbus_buffer_test.c | 102 +++++++++++++++++++++++++++++++++
5 files changed, 181 insertions(+), 12 deletions(-)
create mode 100644 drivers/hv/vmbus_buffer_test.c
diff --git a/drivers/hv/Kconfig b/drivers/hv/Kconfig
index aa11bcefddf2..d44dd60fbc23 100644
--- a/drivers/hv/Kconfig
+++ b/drivers/hv/Kconfig
@@ -65,6 +65,17 @@ config HYPERV_VMBUS
help
Select this option to enable Hyper-V Vmbus driver.
+config HYPERV_VMBUS_KUNIT_TEST
+ bool "Build Hyper-V VMBus buffer KUnit tests"
+ depends on HYPERV_VMBUS && KUNIT
+ default KUNIT_ALL_TESTS
+ help
+ Build the vmbus buffer sizing, GPADL lifetime and reclaim KUnit
+ cases into the hv_vmbus object. The cases reach internal helpers
+ that are not exported, so they cannot live in a separate module.
+
+ If unsure, say N.
+
config MSHV_ROOT
tristate "Microsoft Hyper-V root partition support"
depends on HYPERV && (X86_64 || ARM64)
diff --git a/drivers/hv/Makefile b/drivers/hv/Makefile
index 888a748cc7cb..80d33e9c0e57 100644
--- a/drivers/hv/Makefile
+++ b/drivers/hv/Makefile
@@ -12,6 +12,7 @@ hv_vmbus-y := vmbus_drv.o \
hv.o connection.o channel.o \
channel_mgmt.o ring_buffer.o hv_trace.o
hv_vmbus-$(CONFIG_HYPERV_TESTING) += hv_debugfs.o
+hv_vmbus-$(CONFIG_HYPERV_VMBUS_KUNIT_TEST) += vmbus_buffer_test.o
hv_utils-y := hv_util.o hv_kvp.o hv_snapshot.o hv_utils_transport.o
mshv_root-y := mshv_root_main.o mshv_synic.o mshv_eventfd.o mshv_irq.o \
mshv_root_hv_call.o mshv_portid_table.o mshv_regions.o
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index f8feb2a0ec0a..7d5b99281872 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -28,6 +28,44 @@
#include "hyperv_vmbus.h"
+/*
+ * vmbus_buffer_round_size() and the order-descent helpers are also called
+ * from vmbus_buffer_test.c, which is built into this same object. They are
+ * therefore not static; nothing outside hv_vmbus links against them.
+ */
+int vmbus_buffer_round_size(u32 size, u32 *rounded_size)
+{
+ if (!size)
+ return -EINVAL;
+
+ if (size > U32_MAX - (u32)PAGE_SIZE + 1)
+ return -EOVERFLOW;
+
+ *rounded_size = round_up(size, (u32)PAGE_SIZE);
+ return 0;
+}
+
+unsigned int vmbus_buffer_order(unsigned long remaining,
+ unsigned int max_order)
+{
+ return min_t(unsigned int, max_order, ilog2(remaining));
+}
+
+bool vmbus_buffer_lower_order(unsigned int *order)
+{
+ if (!*order)
+ return false;
+
+ (*order)--;
+ return true;
+}
+
+bool vmbus_buffer_should_free(const struct vmbus_buffer *buffer)
+{
+ return !buffer->leak && !buffer->gpadl.leak &&
+ !buffer->gpadl.gpadl_handle;
+}
+
static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
u32 size,
bool confidential,
@@ -661,7 +699,8 @@ void vmbus_free_buffer(void *addr, struct page **chunks, u32 chunk_cnt)
return;
}
- vunmap(addr);
+ if (addr)
+ vunmap(addr);
for (i = 0; i < chunk_cnt; i++) {
unsigned long vaddr =
@@ -683,8 +722,7 @@ void vmbus_release_buffer(struct vmbus_buffer *buffer)
return;
kvfree(buffer->pages);
- if (!buffer->leak && !buffer->gpadl.leak &&
- !buffer->gpadl.gpadl_handle)
+ if (vmbus_buffer_should_free(buffer))
vmbus_free_buffer(buffer->addr, buffer->chunks,
buffer->chunk_cnt);
memset(buffer, 0, sizeof(*buffer));
@@ -717,12 +755,13 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
struct page ***chunks_out,
u32 *chunk_cnt_out)
{
- unsigned long nr_pages = PFN_UP(size);
- unsigned long remaining = nr_pages;
+ u32 rounded_size;
+ unsigned long nr_pages;
+ unsigned long remaining;
unsigned long page_idx = 0;
struct page **chunks = NULL;
struct page **pages = NULL;
- int order = MAX_PAGE_ORDER;
+ unsigned int order = MAX_PAGE_ORDER;
u32 chunk_cnt = 0;
void *addr;
u32 i;
@@ -731,13 +770,15 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
*chunks_out = NULL;
*chunk_cnt_out = 0;
- if (!nr_pages)
+ if (vmbus_buffer_round_size(size, &rounded_size))
return NULL;
+ nr_pages = rounded_size >> PAGE_SHIFT;
+ remaining = nr_pages;
/* If the buffer does not need to be decrypted, just use vzalloc() */
if ((!hv_is_isolation_supported() &&
!cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT)) || confidential)
- return vzalloc(nr_pages << PAGE_SHIFT);
+ return vzalloc(rounded_size);
/* Worst case: every chunk is a single page. */
chunks = kvmalloc_objs(*chunks, nr_pages, GFP_KERNEL | __GFP_ZERO);
@@ -752,7 +793,7 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
struct page *page;
gfp_t gfp;
- order = min(order, ilog2(remaining));
+ order = vmbus_buffer_order(remaining, order);
/*
* Use __GFP_NORETRY | __GFP_NOWARN to avoid OOM-killing,
@@ -767,7 +808,7 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
page = alloc_pages_node(cpu_to_node(channel->target_cpu),
gfp, order);
if (!page) {
- if (!order--)
+ if (!vmbus_buffer_lower_order(&order))
goto err;
continue;
}
@@ -794,7 +835,7 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
if (!addr)
goto err;
- memset(addr, 0, nr_pages << PAGE_SHIFT);
+ memset(addr, 0, rounded_size);
kvfree(pages);
*chunks_out = chunks;
@@ -1067,8 +1108,10 @@ int vmbus_teardown_gpadl(struct vmbus_channel *channel, struct vmbus_gpadl *gpad
info = kzalloc(sizeof(*info) +
sizeof(struct vmbus_channel_gpadl_teardown), GFP_KERNEL);
- if (!info)
+ if (!info) {
+ gpadl->leak = true;
return -ENOMEM;
+ }
init_completion(&info->waitevent);
info->waiting_channel = channel;
diff --git a/drivers/hv/hyperv_vmbus.h b/drivers/hv/hyperv_vmbus.h
index 20d023c9735e..93df3b35cfd4 100644
--- a/drivers/hv/hyperv_vmbus.h
+++ b/drivers/hv/hyperv_vmbus.h
@@ -551,4 +551,16 @@ int hv_create_ring_sysfs(struct vmbus_channel *channel,
struct vm_area_desc *desc));
int hv_remove_ring_sysfs(struct vmbus_channel *channel);
+/*
+ * vmbus buffer sizing, order-descent and free-decision helpers.
+ *
+ * These are shared with vmbus_buffer_test.c, which is built into the
+ * hv_vmbus object alongside channel.c. They stay unexported.
+ */
+int vmbus_buffer_round_size(u32 size, u32 *rounded_size);
+unsigned int vmbus_buffer_order(unsigned long remaining,
+ unsigned int max_order);
+bool vmbus_buffer_lower_order(unsigned int *order);
+bool vmbus_buffer_should_free(const struct vmbus_buffer *buffer);
+
#endif /* _HYPERV_VMBUS_H */
diff --git a/drivers/hv/vmbus_buffer_test.c b/drivers/hv/vmbus_buffer_test.c
new file mode 100644
index 000000000000..d0a927a46caf
--- /dev/null
+++ b/drivers/hv/vmbus_buffer_test.c
@@ -0,0 +1,102 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * KUnit tests for Hyper-V VMBus buffer allocation and GPADL lifetime.
+ *
+ * Built into the hv_vmbus object rather than a separate module so the
+ * cases can reach the internal helpers declared in hyperv_vmbus.h
+ * without exporting them.
+ */
+#include <kunit/test.h>
+#include <linux/hyperv.h>
+#include <linux/mm.h>
+#include <linux/slab.h>
+#include <linux/vmalloc.h>
+
+#include "hyperv_vmbus.h"
+
+static void vmbus_buffer_size_rounding_test(struct kunit *test)
+{
+ u32 rounded_size;
+
+ KUNIT_EXPECT_EQ(test, vmbus_buffer_round_size(1, &rounded_size), 0);
+ KUNIT_EXPECT_EQ(test, rounded_size, (u32)PAGE_SIZE);
+ KUNIT_EXPECT_EQ(test, vmbus_buffer_round_size(PAGE_SIZE, &rounded_size), 0);
+ KUNIT_EXPECT_EQ(test, rounded_size, (u32)PAGE_SIZE);
+}
+
+static void vmbus_buffer_size_overflow_test(struct kunit *test)
+{
+ u32 rounded_size = 0;
+ int ret;
+
+ KUNIT_EXPECT_EQ(test, vmbus_buffer_round_size(0, &rounded_size), -EINVAL);
+ ret = vmbus_buffer_round_size(U32_MAX, &rounded_size);
+ KUNIT_EXPECT_EQ(test, ret, -EOVERFLOW);
+ ret = vmbus_buffer_round_size(U32_MAX - PAGE_SIZE + 1, &rounded_size);
+ KUNIT_EXPECT_EQ(test, ret, 0);
+ KUNIT_EXPECT_EQ(test, rounded_size,
+ (u32)(U32_MAX - PAGE_SIZE + 1));
+}
+
+static void vmbus_ring_fallback_order_zero_test(struct kunit *test)
+{
+ unsigned int order;
+
+ order = vmbus_buffer_order(1UL << MAX_PAGE_ORDER, MAX_PAGE_ORDER);
+ KUNIT_EXPECT_EQ(test, order, (unsigned int)MAX_PAGE_ORDER);
+ while (order)
+ KUNIT_ASSERT_TRUE(test, vmbus_buffer_lower_order(&order));
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_lower_order(&order));
+ KUNIT_EXPECT_EQ(test, order, 0U);
+ KUNIT_EXPECT_EQ(test, vmbus_buffer_order(3, MAX_PAGE_ORDER), 1U);
+}
+
+static void vmbus_buffer_failed_teardown_leaks_test(struct kunit *test)
+{
+ struct vmbus_buffer buffer = {
+ .addr = (void *)1,
+ .gpadl.gpadl_handle = 1,
+ };
+
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_should_free(&buffer));
+ buffer.gpadl.gpadl_handle = 0;
+ buffer.gpadl.leak = true;
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_should_free(&buffer));
+ buffer.gpadl.leak = false;
+ buffer.leak = true;
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_should_free(&buffer));
+ buffer.leak = false;
+ KUNIT_EXPECT_TRUE(test, vmbus_buffer_should_free(&buffer));
+}
+
+static void vmbus_buffer_partial_allocation_cleanup_test(struct kunit *test)
+{
+ struct page **chunks;
+ struct vmbus_buffer buffer = {};
+
+ chunks = kmalloc_obj(*chunks, GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, chunks);
+ vmbus_free_buffer(NULL, chunks, 0);
+
+ buffer.addr = vzalloc(PAGE_SIZE);
+ KUNIT_ASSERT_NOT_NULL(test, buffer.addr);
+ vmbus_release_buffer(&buffer);
+ KUNIT_EXPECT_PTR_EQ(test, buffer.addr, NULL);
+ vmbus_release_buffer(&buffer);
+}
+
+static struct kunit_case vmbus_buffer_test_cases[] = {
+ KUNIT_CASE(vmbus_buffer_size_rounding_test),
+ KUNIT_CASE(vmbus_buffer_size_overflow_test),
+ KUNIT_CASE(vmbus_ring_fallback_order_zero_test),
+ KUNIT_CASE(vmbus_buffer_failed_teardown_leaks_test),
+ KUNIT_CASE(vmbus_buffer_partial_allocation_cleanup_test),
+ {}
+};
+
+static struct kunit_suite vmbus_buffer_test_suite = {
+ .name = "hyperv-vmbus-buffer",
+ .test_cases = vmbus_buffer_test_cases,
+};
+
+kunit_test_suite(vmbus_buffer_test_suite);
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 03/14] uio: hv_generic: describe buffers for owned allocation
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
2026-10-07 19:07 ` [PATCH v2 01/14] hv: vmbus: convert ring backing through the chunk allocator Emerson Busson
2026-10-07 19:07 ` [PATCH v2 02/14] hv: vmbus: validate chunk buffer allocation and cleanup Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 04/14] hv: vmbus: add KUnit tests for GPADL post failure injection Emerson Busson
` (10 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Group receive and send backing in VMBus descriptors before adopting the
owned allocation and release APIs. Keep the existing vzalloc and
ordinary GPADL path until the lifetime prerequisites land.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/uio/uio_hv_generic.c | 70 +++++++++++++++++-------------------
1 file changed, 32 insertions(+), 38 deletions(-)
diff --git a/drivers/uio/uio_hv_generic.c b/drivers/uio/uio_hv_generic.c
index b3f41ffc74f8..2cb95b4786ca 100644
--- a/drivers/uio/uio_hv_generic.c
+++ b/drivers/uio/uio_hv_generic.c
@@ -56,12 +56,10 @@ struct hv_uio_private_data {
struct hv_device *device;
atomic_t refcnt;
- void *recv_buf;
- struct vmbus_gpadl recv_gpadl;
+ struct vmbus_buffer recv_buffer;
char recv_name[32]; /* "recv_4294967295" */
- void *send_buf;
- struct vmbus_gpadl send_gpadl;
+ struct vmbus_buffer send_buffer;
char send_name[32];
};
@@ -197,19 +195,13 @@ hv_uio_new_channel(struct vmbus_channel *new_sc)
static void
hv_uio_cleanup(struct hv_device *dev, struct hv_uio_private_data *pdata)
{
- if (pdata->send_gpadl.gpadl_handle) {
- if (vmbus_teardown_gpadl(dev->channel, &pdata->send_gpadl))
- pdata->send_gpadl.leak = true;
- if (!pdata->send_gpadl.leak && !pdata->send_gpadl.decrypted)
- vfree(pdata->send_buf);
- }
+ if (pdata->send_buffer.gpadl.gpadl_handle)
+ vmbus_teardown_gpadl(dev->channel, &pdata->send_buffer.gpadl);
+ vmbus_release_buffer(&pdata->send_buffer);
- if (pdata->recv_gpadl.gpadl_handle) {
- if (vmbus_teardown_gpadl(dev->channel, &pdata->recv_gpadl))
- pdata->recv_gpadl.leak = true;
- if (!pdata->recv_gpadl.leak && !pdata->recv_gpadl.decrypted)
- vfree(pdata->recv_buf);
- }
+ if (pdata->recv_buffer.gpadl.gpadl_handle)
+ vmbus_teardown_gpadl(dev->channel, &pdata->recv_buffer.gpadl);
+ vmbus_release_buffer(&pdata->recv_buffer);
}
/* VMBus primary channel is opened on first use */
@@ -306,48 +298,50 @@ hv_uio_probe(struct hv_device *dev,
pdata->info.mem[MON_PAGE_MAP].memtype = UIO_MEM_LOGICAL;
if (channel->device_id == HV_NIC) {
- pdata->recv_buf = vzalloc(RECV_BUFFER_SIZE);
- if (!pdata->recv_buf) {
+ pdata->recv_buffer.addr =
+ vzalloc(RECV_BUFFER_SIZE);
+ if (!pdata->recv_buffer.addr) {
ret = -ENOMEM;
goto fail_free_ring;
}
- ret = vmbus_establish_gpadl(channel, pdata->recv_buf,
- RECV_BUFFER_SIZE, &pdata->recv_gpadl);
- if (ret) {
- if (!pdata->recv_gpadl.leak &&
- !pdata->recv_gpadl.decrypted)
- vfree(pdata->recv_buf);
+ ret = vmbus_establish_gpadl(channel,
+ pdata->recv_buffer.addr,
+ RECV_BUFFER_SIZE,
+ &pdata->recv_buffer.gpadl);
+ pdata->recv_buffer.leak |= pdata->recv_buffer.gpadl.leak;
+ if (ret)
goto fail_close;
- }
/* put Global Physical Address Label in name */
snprintf(pdata->recv_name, sizeof(pdata->recv_name),
- "recv:%u", pdata->recv_gpadl.gpadl_handle);
+ "recv:%u", pdata->recv_buffer.gpadl.gpadl_handle);
pdata->info.mem[RECV_BUF_MAP].name = pdata->recv_name;
- pdata->info.mem[RECV_BUF_MAP].addr = (uintptr_t)pdata->recv_buf;
+ pdata->info.mem[RECV_BUF_MAP].addr =
+ (uintptr_t)pdata->recv_buffer.addr;
pdata->info.mem[RECV_BUF_MAP].size = RECV_BUFFER_SIZE;
pdata->info.mem[RECV_BUF_MAP].memtype = UIO_MEM_VIRTUAL;
- pdata->send_buf = vzalloc(SEND_BUFFER_SIZE);
- if (!pdata->send_buf) {
+ pdata->send_buffer.addr =
+ vzalloc(SEND_BUFFER_SIZE);
+ if (!pdata->send_buffer.addr) {
ret = -ENOMEM;
goto fail_close;
}
- ret = vmbus_establish_gpadl(channel, pdata->send_buf,
- SEND_BUFFER_SIZE, &pdata->send_gpadl);
- if (ret) {
- if (!pdata->send_gpadl.leak &&
- !pdata->send_gpadl.decrypted)
- vfree(pdata->send_buf);
+ ret = vmbus_establish_gpadl(channel,
+ pdata->send_buffer.addr,
+ SEND_BUFFER_SIZE,
+ &pdata->send_buffer.gpadl);
+ pdata->send_buffer.leak |= pdata->send_buffer.gpadl.leak;
+ if (ret)
goto fail_close;
- }
snprintf(pdata->send_name, sizeof(pdata->send_name),
- "send:%u", pdata->send_gpadl.gpadl_handle);
+ "send:%u", pdata->send_buffer.gpadl.gpadl_handle);
pdata->info.mem[SEND_BUF_MAP].name = pdata->send_name;
- pdata->info.mem[SEND_BUF_MAP].addr = (uintptr_t)pdata->send_buf;
+ pdata->info.mem[SEND_BUF_MAP].addr =
+ (uintptr_t)pdata->send_buffer.addr;
pdata->info.mem[SEND_BUF_MAP].size = SEND_BUFFER_SIZE;
pdata->info.mem[SEND_BUF_MAP].memtype = UIO_MEM_VIRTUAL;
}
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 04/14] hv: vmbus: add KUnit tests for GPADL post failure injection
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (2 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 03/14] uio: hv_generic: describe buffers for owned allocation Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 05/14] hv: vmbus: add KUnit test for order-zero allocation fallback Emerson Busson
` (9 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Route GPADL message posting through a private callback so KUnit can
inject header, each body, and teardown post failures without a live
host. Exercise uncertain host ownership after failed posts, host
rejection, and channel rescind response handling through the same
helpers used by production.
The cases join the existing ones in drivers/hv/vmbus_buffer_test.c,
which is built into the hv_vmbus object. That keeps tests out of the
production sources while still letting them drive
vmbus_post_gpadl_messages(), vmbus_gpadl_response_status() and
vmbus_post_gpadl_teardown(). Those three are therefore not static,
and their callback type is declared in hyperv_vmbus.h alongside the
sizing helpers. Nothing outside hv_vmbus links against them, and no
symbol is exported.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 142 ++++++++++++++++----------
drivers/hv/hyperv_vmbus.h | 19 ++++
drivers/hv/vmbus_buffer_test.c | 176 +++++++++++++++++++++++++++++++++
3 files changed, 286 insertions(+), 51 deletions(-)
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index 7d5b99281872..6a1bc9b17869 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -488,6 +488,84 @@ static void vmbus_free_channel_msginfo(struct vmbus_channel_msginfo *msginfo)
* should be 0 for BUFFER type gpadl
* @gpadl_handle: some funky thing
*/
+static int vmbus_gpadl_post_real(void *context, void *buffer,
+ size_t buflen, bool can_sleep)
+{
+ (void)context;
+ return vmbus_post_msg(buffer, buflen, can_sleep);
+}
+
+int vmbus_post_gpadl_messages(struct vmbus_channel_msginfo *msginfo,
+ u32 gpadl, bool *posted,
+ vmbus_gpadl_post_fn post_msg,
+ void *context)
+{
+ struct vmbus_channel_gpadl_header *gpadl_header;
+ struct vmbus_channel_msginfo *submsginfo;
+ struct list_head *curr;
+ int ret;
+
+ gpadl_header = (struct vmbus_channel_gpadl_header *)msginfo->msg;
+ gpadl_header->header.msgtype = CHANNELMSG_GPADL_HEADER;
+ gpadl_header->gpadl = gpadl;
+
+ /* A failed post may still have reached the host. */
+ *posted = true;
+ ret = post_msg(context, gpadl_header,
+ msginfo->msgsize - sizeof(*msginfo), true);
+ trace_vmbus_establish_gpadl_header(gpadl_header, ret);
+ if (ret)
+ return ret;
+
+ list_for_each(curr, &msginfo->submsglist) {
+ struct vmbus_channel_gpadl_body *gpadl_body;
+
+ submsginfo = list_entry(curr, struct vmbus_channel_msginfo,
+ msglistentry);
+ gpadl_body = (struct vmbus_channel_gpadl_body *)submsginfo->msg;
+ gpadl_body->header.msgtype = CHANNELMSG_GPADL_BODY;
+ gpadl_body->gpadl = gpadl;
+
+ ret = post_msg(context, gpadl_body,
+ submsginfo->msgsize - sizeof(*submsginfo), true);
+ trace_vmbus_establish_gpadl_body(gpadl_body, ret);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
+
+int vmbus_gpadl_response_status(u32 creation_status, bool rescind,
+ bool *posted)
+{
+ *posted = false;
+ if (creation_status)
+ return -EDQUOT;
+ if (rescind)
+ return -ENODEV;
+
+ return 0;
+}
+
+int vmbus_post_gpadl_teardown(u32 child_relid,
+ struct vmbus_channel_gpadl_teardown *msg,
+ u32 gpadl,
+ vmbus_gpadl_post_fn post_msg,
+ void *context)
+{
+ int ret;
+
+ msg->header.msgtype = CHANNELMSG_GPADL_TEARDOWN;
+ msg->child_relid = child_relid;
+ msg->gpadl = gpadl;
+
+ ret = post_msg(context, msg, sizeof(*msg), true);
+ trace_vmbus_teardown_gpadl(msg, ret);
+
+ return ret;
+}
+
static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
enum hv_gpadl_type type, void *kbuffer,
u32 size, u32 send_offset, bool memory_prepared,
@@ -495,11 +573,9 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
struct vmbus_gpadl *gpadl)
{
struct vmbus_channel_gpadl_header *gpadlmsg;
- struct vmbus_channel_gpadl_body *gpadl_body;
struct vmbus_channel_msginfo *msginfo = NULL;
- struct vmbus_channel_msginfo *submsginfo;
- struct list_head *curr;
u32 next_gpadl_handle;
+ u32 creation_status;
unsigned long flags;
bool posted = false;
int ret = 0;
@@ -558,49 +634,20 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
goto cleanup;
}
- /* A failed post may still have reached the host. */
- posted = true;
- ret = vmbus_post_msg(gpadlmsg, msginfo->msgsize -
- sizeof(*msginfo), true);
-
- trace_vmbus_establish_gpadl_header(gpadlmsg, ret);
-
- if (ret != 0)
+ ret = vmbus_post_gpadl_messages(msginfo, next_gpadl_handle, &posted,
+ vmbus_gpadl_post_real, NULL);
+ if (ret)
goto cleanup;
- list_for_each(curr, &msginfo->submsglist) {
- submsginfo = (struct vmbus_channel_msginfo *)curr;
- gpadl_body =
- (struct vmbus_channel_gpadl_body *)submsginfo->msg;
-
- gpadl_body->header.msgtype =
- CHANNELMSG_GPADL_BODY;
- gpadl_body->gpadl = next_gpadl_handle;
-
- ret = vmbus_post_msg(gpadl_body,
- submsginfo->msgsize - sizeof(*submsginfo),
- true);
-
- trace_vmbus_establish_gpadl_body(gpadl_body, ret);
-
- if (ret != 0)
- goto cleanup;
-
- }
wait_for_completion(&msginfo->waitevent);
- if (msginfo->response.gpadl_created.creation_status != 0) {
- posted = false;
- pr_err("Failed to establish GPADL: err = 0x%x\n",
- msginfo->response.gpadl_created.creation_status);
-
- ret = -EDQUOT;
- goto cleanup;
- }
-
- if (channel->rescind) {
- posted = false;
- ret = -ENODEV;
+ creation_status = msginfo->response.gpadl_created.creation_status;
+ ret = vmbus_gpadl_response_status(creation_status, channel->rescind,
+ &posted);
+ if (ret) {
+ if (creation_status)
+ pr_err("Failed to establish GPADL: err = 0x%x\n",
+ creation_status);
goto cleanup;
}
@@ -608,8 +655,6 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
gpadl->gpadl_handle = gpadlmsg->gpadl;
gpadl->buffer = kbuffer;
gpadl->size = size;
- posted = false;
-
cleanup:
spin_lock_irqsave(&vmbus_connection.channelmsg_lock, flags);
@@ -1118,9 +1163,6 @@ int vmbus_teardown_gpadl(struct vmbus_channel *channel, struct vmbus_gpadl *gpad
msg = (struct vmbus_channel_gpadl_teardown *)info->msg;
- msg->header.msgtype = CHANNELMSG_GPADL_TEARDOWN;
- msg->child_relid = channel->offermsg.child_relid;
- msg->gpadl = gpadl->gpadl_handle;
spin_lock_irqsave(&vmbus_connection.channelmsg_lock, flags);
list_add_tail(&info->msglistentry,
@@ -1130,10 +1172,8 @@ int vmbus_teardown_gpadl(struct vmbus_channel *channel, struct vmbus_gpadl *gpad
if (channel->rescind)
goto post_msg_err;
- ret = vmbus_post_msg(msg, sizeof(struct vmbus_channel_gpadl_teardown),
- true);
-
- trace_vmbus_teardown_gpadl(msg, ret);
+ ret = vmbus_post_gpadl_teardown(channel->offermsg.child_relid, msg, gpadl->gpadl_handle,
+ vmbus_gpadl_post_real, NULL);
if (ret)
goto post_msg_err;
diff --git a/drivers/hv/hyperv_vmbus.h b/drivers/hv/hyperv_vmbus.h
index 93df3b35cfd4..bb585ac2ceac 100644
--- a/drivers/hv/hyperv_vmbus.h
+++ b/drivers/hv/hyperv_vmbus.h
@@ -563,4 +563,23 @@ unsigned int vmbus_buffer_order(unsigned long remaining,
bool vmbus_buffer_lower_order(unsigned int *order);
bool vmbus_buffer_should_free(const struct vmbus_buffer *buffer);
+/*
+ * GPADL post and response helpers, also shared with vmbus_buffer_test.c.
+ * Same deal as the sizing helpers above: defined in channel.c, built into
+ * the hv_vmbus object, never exported.
+ */
+typedef int (*vmbus_gpadl_post_fn)(void *context, void *buffer,
+ size_t buflen, bool can_sleep);
+int vmbus_post_gpadl_messages(struct vmbus_channel_msginfo *msginfo,
+ u32 gpadl, bool *posted,
+ vmbus_gpadl_post_fn post_msg,
+ void *context);
+int vmbus_gpadl_response_status(u32 creation_status, bool rescind,
+ bool *posted);
+int vmbus_post_gpadl_teardown(u32 child_relid,
+ struct vmbus_channel_gpadl_teardown *msg,
+ u32 gpadl,
+ vmbus_gpadl_post_fn post_msg,
+ void *context);
+
#endif /* _HYPERV_VMBUS_H */
diff --git a/drivers/hv/vmbus_buffer_test.c b/drivers/hv/vmbus_buffer_test.c
index d0a927a46caf..fbf7c4351d64 100644
--- a/drivers/hv/vmbus_buffer_test.c
+++ b/drivers/hv/vmbus_buffer_test.c
@@ -85,12 +85,188 @@ static void vmbus_buffer_partial_allocation_cleanup_test(struct kunit *test)
vmbus_release_buffer(&buffer);
}
+struct vmbus_gpadl_post_test_context {
+ struct kunit *test;
+ unsigned int call_count;
+ unsigned int fail_call;
+ u32 first_msgtype;
+ u32 expected_gpadl;
+};
+
+static int vmbus_gpadl_test_post(void *context, void *buffer,
+ size_t buflen, bool can_sleep)
+{
+ struct vmbus_gpadl_post_test_context *test_context = context;
+ struct vmbus_channel_message_header *header = buffer;
+
+ KUNIT_EXPECT_TRUE(test_context->test, can_sleep);
+ KUNIT_EXPECT_GT(test_context->test, buflen,
+ (size_t)sizeof(*header));
+ test_context->call_count++;
+ KUNIT_EXPECT_EQ(test_context->test, header->msgtype,
+ test_context->call_count == 1 ?
+ test_context->first_msgtype : CHANNELMSG_GPADL_BODY);
+ switch (header->msgtype) {
+ case CHANNELMSG_GPADL_HEADER: {
+ struct vmbus_channel_gpadl_header *gpadl_header = buffer;
+
+ KUNIT_EXPECT_EQ(test_context->test, gpadl_header->gpadl,
+ test_context->expected_gpadl);
+ break;
+ }
+ case CHANNELMSG_GPADL_BODY: {
+ struct vmbus_channel_gpadl_body *gpadl_body = buffer;
+
+ KUNIT_EXPECT_EQ(test_context->test, gpadl_body->gpadl,
+ test_context->expected_gpadl);
+ break;
+ }
+ case CHANNELMSG_GPADL_TEARDOWN: {
+ struct vmbus_channel_gpadl_teardown *teardown = buffer;
+
+ KUNIT_EXPECT_EQ(test_context->test, teardown->gpadl,
+ test_context->expected_gpadl);
+ KUNIT_EXPECT_EQ(test_context->test, teardown->child_relid,
+ 7U);
+ break;
+ }
+ default:
+ KUNIT_FAIL(test_context->test,
+ "unexpected GPADL message type: %d",
+ header->msgtype);
+ }
+ if (test_context->call_count == test_context->fail_call)
+ return -EIO;
+
+ return 0;
+}
+
+static struct vmbus_channel_msginfo *
+vmbus_gpadl_test_msginfo(struct kunit *test, unsigned int body_count)
+{
+ struct vmbus_channel_msginfo *msginfo;
+ struct vmbus_channel_msginfo *body_info;
+ unsigned int i;
+
+ msginfo = kunit_kzalloc(test, sizeof(*msginfo) +
+ sizeof(struct vmbus_channel_gpadl_header),
+ GFP_KERNEL);
+ if (!msginfo)
+ return NULL;
+
+ msginfo->msgsize = sizeof(*msginfo) +
+ sizeof(struct vmbus_channel_gpadl_header);
+ INIT_LIST_HEAD(&msginfo->submsglist);
+
+ for (i = 0; i < body_count; i++) {
+ body_info = kunit_kzalloc(test,
+ sizeof(*body_info) +
+ sizeof(struct vmbus_channel_gpadl_body),
+ GFP_KERNEL);
+ if (!body_info)
+ return NULL;
+
+ body_info->msgsize = sizeof(*body_info) +
+ sizeof(struct vmbus_channel_gpadl_body);
+ INIT_LIST_HEAD(&body_info->msglistentry);
+ list_add_tail(&body_info->msglistentry, &msginfo->submsglist);
+ }
+
+ return msginfo;
+}
+
+static void vmbus_gpadl_post_failure_test(struct kunit *test)
+{
+ unsigned int fail_call;
+
+ for (fail_call = 1; fail_call <= 3; fail_call++) {
+ struct vmbus_gpadl_post_test_context context = {
+ .test = test,
+ .fail_call = fail_call,
+ .first_msgtype = CHANNELMSG_GPADL_HEADER,
+ .expected_gpadl = 17,
+ };
+ struct vmbus_channel_msginfo *msginfo;
+ bool posted = false;
+ int ret;
+
+ msginfo = vmbus_gpadl_test_msginfo(test, 2);
+ KUNIT_ASSERT_NOT_NULL(test, msginfo);
+
+ ret = vmbus_post_gpadl_messages(msginfo, 17, &posted,
+ vmbus_gpadl_test_post, &context);
+ KUNIT_EXPECT_EQ(test, ret, -EIO);
+ KUNIT_EXPECT_EQ(test, context.call_count, fail_call);
+ KUNIT_EXPECT_TRUE(test, posted);
+ }
+}
+
+static void vmbus_gpadl_post_success_test(struct kunit *test)
+{
+ struct vmbus_gpadl_post_test_context context = {
+ .test = test,
+ .first_msgtype = CHANNELMSG_GPADL_HEADER,
+ .expected_gpadl = 17,
+ };
+ struct vmbus_channel_msginfo *msginfo;
+ bool posted = false;
+ int ret;
+
+ msginfo = vmbus_gpadl_test_msginfo(test, 2);
+ KUNIT_ASSERT_NOT_NULL(test, msginfo);
+
+ ret = vmbus_post_gpadl_messages(msginfo, 17, &posted,
+ vmbus_gpadl_test_post, &context);
+ KUNIT_EXPECT_EQ(test, ret, 0);
+ KUNIT_EXPECT_EQ(test, context.call_count, 3U);
+ KUNIT_EXPECT_TRUE(test, posted);
+}
+
+static void vmbus_gpadl_response_state_test(struct kunit *test)
+{
+ bool posted = true;
+
+ KUNIT_EXPECT_EQ(test, vmbus_gpadl_response_status(0, false, &posted), 0);
+ KUNIT_EXPECT_FALSE(test, posted);
+
+ posted = true;
+ KUNIT_EXPECT_EQ(test,
+ vmbus_gpadl_response_status(1, false, &posted), -EDQUOT);
+ KUNIT_EXPECT_FALSE(test, posted);
+
+ posted = true;
+ KUNIT_EXPECT_EQ(test,
+ vmbus_gpadl_response_status(0, true, &posted), -ENODEV);
+ KUNIT_EXPECT_FALSE(test, posted);
+}
+
+static void vmbus_gpadl_teardown_post_failure_test(struct kunit *test)
+{
+ struct vmbus_gpadl_post_test_context context = {
+ .test = test,
+ .fail_call = 1,
+ .first_msgtype = CHANNELMSG_GPADL_TEARDOWN,
+ .expected_gpadl = 17,
+ };
+ struct vmbus_channel_gpadl_teardown msg = {};
+ int ret;
+
+ ret = vmbus_post_gpadl_teardown(7, &msg, 17,
+ vmbus_gpadl_test_post, &context);
+ KUNIT_EXPECT_EQ(test, ret, -EIO);
+ KUNIT_EXPECT_EQ(test, context.call_count, 1U);
+}
+
static struct kunit_case vmbus_buffer_test_cases[] = {
KUNIT_CASE(vmbus_buffer_size_rounding_test),
KUNIT_CASE(vmbus_buffer_size_overflow_test),
KUNIT_CASE(vmbus_ring_fallback_order_zero_test),
KUNIT_CASE(vmbus_buffer_failed_teardown_leaks_test),
KUNIT_CASE(vmbus_buffer_partial_allocation_cleanup_test),
+ KUNIT_CASE(vmbus_gpadl_post_failure_test),
+ KUNIT_CASE(vmbus_gpadl_post_success_test),
+ KUNIT_CASE(vmbus_gpadl_response_state_test),
+ KUNIT_CASE(vmbus_gpadl_teardown_post_failure_test),
{}
};
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 05/14] hv: vmbus: add KUnit test for order-zero allocation fallback
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (3 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 04/14] hv: vmbus: add KUnit tests for GPADL post failure injection Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 06/14] hv: vmbus: cover all shared-page policy combinations Emerson Busson
` (8 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Factor the order-descent allocation loop so KUnit can inject failures
for orders above zero while retaining a real order-zero allocation in
the test path. Also verify clean exhaustion when the order-zero attempt
fails.
The factoring keeps the per-order GFP policy of the loop it replaces.
Speculative attempts above order 0 still carry __GFP_COMP together with
__GFP_NORETRY | __GFP_NOWARN so they cannot OOM-kill, and the order-0
attempt is still the plain GFP_KERNEL | __GFP_ZERO final fallback that
is allowed to try harder. Reusing one gfp across the whole descent
would leave __GFP_NORETRY set on the order-0 attempt and silently
weaken that last resort.
The case joins the others in drivers/hv/vmbus_buffer_test.c, built
into the hv_vmbus object. vmbus_alloc_pages_with_fallback() and its
callback type are therefore declared in hyperv_vmbus.h instead of
staying static in channel.c. Nothing outside hv_vmbus links against
them, and no symbol is exported.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 56 ++++++++++++++++++++++----------
drivers/hv/hyperv_vmbus.h | 13 ++++++++
drivers/hv/vmbus_buffer_test.c | 58 ++++++++++++++++++++++++++++++++++
3 files changed, 111 insertions(+), 16 deletions(-)
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index 6a1bc9b17869..0c025a12808a 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -60,6 +60,31 @@ bool vmbus_buffer_lower_order(unsigned int *order)
return true;
}
+struct page *vmbus_alloc_pages_with_fallback(int nid, gfp_t gfp,
+ unsigned int *order,
+ vmbus_alloc_pages_fn alloc,
+ void *context)
+{
+ struct page *page;
+
+ do {
+ gfp_t try_gfp = gfp;
+
+ /*
+ * Use __GFP_NORETRY | __GFP_NOWARN to avoid OOM-killing,
+ * but try harder at order 0 since that is the final
+ * fallback.
+ * __GFP_COMP stores order information in the page folio.
+ */
+ if (*order)
+ try_gfp |= __GFP_COMP | __GFP_NORETRY | __GFP_NOWARN;
+
+ page = alloc(context, nid, try_gfp, *order);
+ if (page || !vmbus_buffer_lower_order(order))
+ return page;
+ } while (true);
+}
+
bool vmbus_buffer_should_free(const struct vmbus_buffer *buffer)
{
return !buffer->leak && !buffer->gpadl.leak &&
@@ -774,6 +799,14 @@ void vmbus_release_buffer(struct vmbus_buffer *buffer)
}
EXPORT_SYMBOL_GPL(vmbus_release_buffer);
+static struct page *vmbus_alloc_pages_node(void *context, int nid,
+ gfp_t gfp,
+ unsigned int order)
+{
+ (void)context;
+ return alloc_pages_node(nid, gfp, order);
+}
+
/**
* __vmbus_alloc_buffer - allocate host-visible, virtually-contiguous backing.
*
@@ -837,26 +870,17 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
while (remaining) {
struct page *page;
gfp_t gfp;
+ int nid;
order = vmbus_buffer_order(remaining, order);
- /*
- * Use __GFP_NORETRY | __GFP_NOWARN to avoid OOM-killing,
- * but try harder at order 0 since that is the final
- * fallback.
- * __GFP_COMP stores order information in the page folio.
- */
gfp = GFP_KERNEL | __GFP_ZERO;
- if (order)
- gfp |= __GFP_COMP | __GFP_NORETRY | __GFP_NOWARN;
-
- page = alloc_pages_node(cpu_to_node(channel->target_cpu),
- gfp, order);
- if (!page) {
- if (!vmbus_buffer_lower_order(&order))
- goto err;
- continue;
- }
+
+ nid = cpu_to_node(channel->target_cpu);
+ page = vmbus_alloc_pages_with_fallback(nid, gfp, &order,
+ vmbus_alloc_pages_node, NULL);
+ if (!page)
+ goto err;
ret = set_memory_decrypted((unsigned long)page_address(page),
1U << order);
diff --git a/drivers/hv/hyperv_vmbus.h b/drivers/hv/hyperv_vmbus.h
index bb585ac2ceac..e08c472041d5 100644
--- a/drivers/hv/hyperv_vmbus.h
+++ b/drivers/hv/hyperv_vmbus.h
@@ -563,6 +563,19 @@ unsigned int vmbus_buffer_order(unsigned long remaining,
bool vmbus_buffer_lower_order(unsigned int *order);
bool vmbus_buffer_should_free(const struct vmbus_buffer *buffer);
+/*
+ * Order-descent allocation, shared with vmbus_buffer_test.c so the
+ * cases can inject per-order failures. Defined in channel.c, built
+ * into the hv_vmbus object, never exported.
+ */
+typedef struct page *(*vmbus_alloc_pages_fn)(void *context, int nid,
+ gfp_t gfp,
+ unsigned int order);
+struct page *vmbus_alloc_pages_with_fallback(int nid, gfp_t gfp,
+ unsigned int *order,
+ vmbus_alloc_pages_fn alloc,
+ void *context);
+
/*
* GPADL post and response helpers, also shared with vmbus_buffer_test.c.
* Same deal as the sizing helpers above: defined in channel.c, built into
diff --git a/drivers/hv/vmbus_buffer_test.c b/drivers/hv/vmbus_buffer_test.c
index fbf7c4351d64..60fc526af1af 100644
--- a/drivers/hv/vmbus_buffer_test.c
+++ b/drivers/hv/vmbus_buffer_test.c
@@ -257,12 +257,70 @@ static void vmbus_gpadl_teardown_post_failure_test(struct kunit *test)
KUNIT_EXPECT_EQ(test, context.call_count, 1U);
}
+struct vmbus_buffer_page_alloc_test_context {
+ struct kunit *test;
+ unsigned int attempts;
+ unsigned int expected_order;
+ bool fail_order_zero;
+};
+
+static struct page *vmbus_buffer_test_alloc_page(void *context, int nid,
+ gfp_t gfp,
+ unsigned int order)
+{
+ struct vmbus_buffer_page_alloc_test_context *test_context = context;
+
+ (void)nid;
+ KUNIT_EXPECT_EQ(test_context->test, order,
+ test_context->expected_order);
+ test_context->attempts++;
+ if (test_context->expected_order)
+ test_context->expected_order--;
+
+ if (order)
+ return NULL;
+ if (test_context->fail_order_zero)
+ return NULL;
+
+ return alloc_pages(gfp, 0);
+}
+
+static void vmbus_buffer_order_zero_allocation_test(struct kunit *test)
+{
+ struct vmbus_buffer_page_alloc_test_context context = {
+ .test = test,
+ .expected_order = MAX_PAGE_ORDER,
+ };
+ struct page *page;
+ unsigned int order = MAX_PAGE_ORDER;
+
+ page = vmbus_alloc_pages_with_fallback(0, GFP_KERNEL, &order,
+ vmbus_buffer_test_alloc_page, &context);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+ KUNIT_EXPECT_EQ(test, order, 0U);
+ KUNIT_EXPECT_EQ(test, context.expected_order, 0U);
+ KUNIT_EXPECT_EQ(test, context.attempts, (unsigned int)MAX_PAGE_ORDER + 1);
+ __free_pages(page, 0);
+
+ context.attempts = 0;
+ context.expected_order = MAX_PAGE_ORDER;
+ context.fail_order_zero = true;
+ order = MAX_PAGE_ORDER;
+ page = vmbus_alloc_pages_with_fallback(0, GFP_KERNEL, &order,
+ vmbus_buffer_test_alloc_page, &context);
+ KUNIT_EXPECT_PTR_EQ(test, page, NULL);
+ KUNIT_EXPECT_EQ(test, order, 0U);
+ KUNIT_EXPECT_EQ(test, context.expected_order, 0U);
+ KUNIT_EXPECT_EQ(test, context.attempts, (unsigned int)MAX_PAGE_ORDER + 1);
+}
+
static struct kunit_case vmbus_buffer_test_cases[] = {
KUNIT_CASE(vmbus_buffer_size_rounding_test),
KUNIT_CASE(vmbus_buffer_size_overflow_test),
KUNIT_CASE(vmbus_ring_fallback_order_zero_test),
KUNIT_CASE(vmbus_buffer_failed_teardown_leaks_test),
KUNIT_CASE(vmbus_buffer_partial_allocation_cleanup_test),
+ KUNIT_CASE(vmbus_buffer_order_zero_allocation_test),
KUNIT_CASE(vmbus_gpadl_post_failure_test),
KUNIT_CASE(vmbus_gpadl_post_success_test),
KUNIT_CASE(vmbus_gpadl_response_state_test),
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 06/14] hv: vmbus: cover all shared-page policy combinations
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (4 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 05/14] hv: vmbus: add KUnit test for order-zero allocation fallback Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 07/14] hv: vmbus: distinguish host rescind from local channel unload Emerson Busson
` (7 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Name the shared-page selection and enumerate all eight combinations of
Hyper-V isolation, independent guest-memory encryption and channel
confidentiality. Either platform visibility signal selects shared
backing; the confidential-channel override keeps private backing. These
are decision tests, not confidential hardware qualification.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 16 ++++++++++++++--
drivers/hv/hyperv_vmbus.h | 2 ++
drivers/hv/vmbus_buffer_test.c | 29 +++++++++++++++++++++++++++++
3 files changed, 45 insertions(+), 2 deletions(-)
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index 0c025a12808a..8ed155f2669f 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -45,6 +45,13 @@ int vmbus_buffer_round_size(u32 size, u32 *rounded_size)
return 0;
}
+bool
+vmbus_needs_shared_pages(bool hv_isolation, bool guest_mem_encrypted,
+ bool confidential)
+{
+ return !confidential && (hv_isolation || guest_mem_encrypted);
+}
+
unsigned int vmbus_buffer_order(unsigned long remaining,
unsigned int max_order)
{
@@ -837,6 +844,8 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
unsigned long nr_pages;
unsigned long remaining;
unsigned long page_idx = 0;
+ bool hv_isolation;
+ bool guest_mem_encrypted;
struct page **chunks = NULL;
struct page **pages = NULL;
unsigned int order = MAX_PAGE_ORDER;
@@ -853,9 +862,12 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
nr_pages = rounded_size >> PAGE_SHIFT;
remaining = nr_pages;
+ hv_isolation = hv_is_isolation_supported();
+ guest_mem_encrypted = cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT);
+
/* If the buffer does not need to be decrypted, just use vzalloc() */
- if ((!hv_is_isolation_supported() &&
- !cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT)) || confidential)
+ if (!vmbus_needs_shared_pages(hv_isolation, guest_mem_encrypted,
+ confidential))
return vzalloc(rounded_size);
/* Worst case: every chunk is a single page. */
diff --git a/drivers/hv/hyperv_vmbus.h b/drivers/hv/hyperv_vmbus.h
index e08c472041d5..06094000f2f1 100644
--- a/drivers/hv/hyperv_vmbus.h
+++ b/drivers/hv/hyperv_vmbus.h
@@ -557,6 +557,8 @@ int hv_remove_ring_sysfs(struct vmbus_channel *channel);
* These are shared with vmbus_buffer_test.c, which is built into the
* hv_vmbus object alongside channel.c. They stay unexported.
*/
+bool vmbus_needs_shared_pages(bool hv_isolation, bool guest_mem_encrypted,
+ bool confidential);
int vmbus_buffer_round_size(u32 size, u32 *rounded_size);
unsigned int vmbus_buffer_order(unsigned long remaining,
unsigned int max_order);
diff --git a/drivers/hv/vmbus_buffer_test.c b/drivers/hv/vmbus_buffer_test.c
index 60fc526af1af..9b401ef2ceea 100644
--- a/drivers/hv/vmbus_buffer_test.c
+++ b/drivers/hv/vmbus_buffer_test.c
@@ -38,6 +38,34 @@ static void vmbus_buffer_size_overflow_test(struct kunit *test)
(u32)(U32_MAX - PAGE_SIZE + 1));
}
+static void vmbus_buffer_private_shared_selection_test(struct kunit *test)
+{
+ static const struct {
+ bool isolation;
+ bool encrypted;
+ bool confidential;
+ bool shared;
+ } cases[] = {
+ { false, false, false, false },
+ { false, false, true, false },
+ { false, true, false, true },
+ { false, true, true, false },
+ { true, false, false, true },
+ { true, false, true, false },
+ { true, true, false, true },
+ { true, true, true, false },
+ };
+ unsigned int i;
+
+ /* Either platform signal requires visibility; confidential stays private. */
+ for (i = 0; i < ARRAY_SIZE(cases); i++)
+ KUNIT_EXPECT_EQ(test,
+ vmbus_needs_shared_pages(cases[i].isolation,
+ cases[i].encrypted,
+ cases[i].confidential),
+ cases[i].shared);
+}
+
static void vmbus_ring_fallback_order_zero_test(struct kunit *test)
{
unsigned int order;
@@ -317,6 +345,7 @@ static void vmbus_buffer_order_zero_allocation_test(struct kunit *test)
static struct kunit_case vmbus_buffer_test_cases[] = {
KUNIT_CASE(vmbus_buffer_size_rounding_test),
KUNIT_CASE(vmbus_buffer_size_overflow_test),
+ KUNIT_CASE(vmbus_buffer_private_shared_selection_test),
KUNIT_CASE(vmbus_ring_fallback_order_zero_test),
KUNIT_CASE(vmbus_buffer_failed_teardown_leaks_test),
KUNIT_CASE(vmbus_buffer_partial_allocation_cleanup_test),
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 07/14] hv: vmbus: distinguish host rescind from local channel unload
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (5 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 06/14] hv: vmbus: cover all shared-page policy combinations Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 08/14] hv: vmbus: retain backing until ownership and references clear Emerson Busson
` (6 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
A channel can go away because the host revoked the offer, or because
the guest is tearing the channel down itself. Both paths arrive at
vmbus_onoffer_rescind() and set channel->rescind, so a later buffer
consumer cannot tell whether the host has already taken the pages
back or whether the guest still owns them and is about to free them.
Carry the origin through the message layer. vmbus_onmessage() takes a
host_generated flag: the DPC work item sets it for host messages and
vmbus_force_channel_rescinded() clears it for the local unload path.
A small table adapter keeps the dispatch signature unchanged, while
the rescind handler itself records the origin in
channel->rescind_from_host next to the existing rescind flag. Both
flags are cleared when a channel is set up.
The disconnected message path frees its work context instead of
returning without a kfree(); it now owns that allocation from the
moment container_of() runs.
Nothing reads rescind_from_host yet. The buffer-ownership rework
lands in the next patch and is what consumes the flag.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel_mgmt.c | 31 +++++++++++++++++++++++++------
drivers/hv/vmbus_drv.c | 13 ++++++++-----
include/linux/hyperv.h | 4 +++-
3 files changed, 36 insertions(+), 12 deletions(-)
diff --git a/drivers/hv/channel_mgmt.c b/drivers/hv/channel_mgmt.c
index a044fd3b3c4e..93fc105cd179 100644
--- a/drivers/hv/channel_mgmt.c
+++ b/drivers/hv/channel_mgmt.c
@@ -170,14 +170,17 @@ static const struct {
* The rescinded channel may be blocked waiting for a response from the host;
* take care of that.
*/
-static void vmbus_rescind_cleanup(struct vmbus_channel *channel)
+static void vmbus_rescind_cleanup(struct vmbus_channel *channel,
+ bool host_generated)
{
struct vmbus_channel_msginfo *msginfo;
unsigned long flags;
spin_lock_irqsave(&vmbus_connection.channelmsg_lock, flags);
- channel->rescind = true;
+ if (host_generated)
+ WRITE_ONCE(channel->rescind_from_host, true);
+ WRITE_ONCE(channel->rescind, true);
list_for_each_entry(msginfo, &vmbus_connection.chn_msg_list,
msglistentry) {
@@ -955,6 +958,9 @@ EXPORT_SYMBOL_GPL(vmbus_initiate_unload);
static void vmbus_setup_channel_state(struct vmbus_channel *channel,
struct vmbus_channel_offer_channel *offer)
{
+ WRITE_ONCE(channel->rescind, false);
+ WRITE_ONCE(channel->rescind_from_host, false);
+
/*
* Setup state for signalling the host.
*/
@@ -1159,7 +1165,8 @@ static void check_ready_for_suspend_event(void)
*
* We queue a work item to process this offer synchronously
*/
-static void vmbus_onoffer_rescind(struct vmbus_channel_message_header *hdr)
+static void vmbus_onoffer_rescind(struct vmbus_channel_message_header *hdr,
+ bool host_generated)
{
struct vmbus_channel_rescind_offer *rescind;
struct vmbus_channel *channel;
@@ -1238,7 +1245,7 @@ static void vmbus_onoffer_rescind(struct vmbus_channel_message_header *hdr)
/*
* Now wait for offer handling to complete.
*/
- vmbus_rescind_cleanup(channel);
+ vmbus_rescind_cleanup(channel, host_generated);
while (READ_ONCE(channel->probe_done) == false) {
/*
* We wait here until any channel offer is currently
@@ -1555,12 +1562,18 @@ static void vmbus_onversion_response(
}
/* Channel message dispatch table */
+static void
+vmbus_onoffer_rescind_from_table(struct vmbus_channel_message_header *hdr)
+{
+ vmbus_onoffer_rescind(hdr, true);
+}
+
const struct vmbus_channel_message_table_entry
channel_message_table[CHANNELMSG_COUNT] = {
{ CHANNELMSG_INVALID, 0, NULL, 0},
{ CHANNELMSG_OFFERCHANNEL, 0, vmbus_onoffer,
sizeof(struct vmbus_channel_offer_channel)},
- { CHANNELMSG_RESCIND_CHANNELOFFER, 0, vmbus_onoffer_rescind,
+ { CHANNELMSG_RESCIND_CHANNELOFFER, 0, vmbus_onoffer_rescind_from_table,
sizeof(struct vmbus_channel_rescind_offer) },
{ CHANNELMSG_REQUESTOFFERS, 0, NULL, 0},
{ CHANNELMSG_ALLOFFERS_DELIVERED, 1, vmbus_onoffers_delivered, 0},
@@ -1596,7 +1609,8 @@ channel_message_table[CHANNELMSG_COUNT] = {
*
* This is invoked in the vmbus worker thread context.
*/
-void vmbus_onmessage(struct vmbus_channel_message_header *hdr)
+void vmbus_onmessage(struct vmbus_channel_message_header *hdr,
+ bool host_generated)
{
trace_vmbus_on_message(hdr);
@@ -1604,6 +1618,11 @@ void vmbus_onmessage(struct vmbus_channel_message_header *hdr)
* vmbus_on_msg_dpc() makes sure the hdr->msgtype here can not go
* out of bound and the message_handler pointer can not be NULL.
*/
+ if (hdr->msgtype == CHANNELMSG_RESCIND_CHANNELOFFER) {
+ vmbus_onoffer_rescind(hdr, host_generated);
+ return;
+ }
+
channel_message_table[hdr->msgtype].message_handler(hdr);
}
diff --git a/drivers/hv/vmbus_drv.c b/drivers/hv/vmbus_drv.c
index 5ebdbe24b5a1..723252f1b551 100644
--- a/drivers/hv/vmbus_drv.c
+++ b/drivers/hv/vmbus_drv.c
@@ -1022,6 +1022,7 @@ static const struct bus_type hv_bus = {
struct onmessage_work_context {
struct work_struct work;
+ bool host_generated;
struct {
struct hv_message_header header;
u8 payload[];
@@ -1032,14 +1033,14 @@ static void vmbus_onmessage_work(struct work_struct *work)
{
struct onmessage_work_context *ctx;
+ ctx = container_of(work, struct onmessage_work_context, work);
/* Do not process messages if we're in DISCONNECTED state */
- if (vmbus_connection.conn_state == DISCONNECTED)
+ if (vmbus_connection.conn_state == DISCONNECTED) {
+ kfree(ctx);
return;
-
- ctx = container_of(work, struct onmessage_work_context,
- work);
+ }
vmbus_onmessage((struct vmbus_channel_message_header *)
- &ctx->msg.payload);
+ &ctx->msg.payload, ctx->host_generated);
kfree(ctx);
}
@@ -1109,6 +1110,7 @@ static void __vmbus_on_msg_dpc(void *message_page_addr)
return;
INIT_WORK(&ctx->work, vmbus_onmessage_work);
+ ctx->host_generated = true;
ctx->msg.header = msg_copy.header;
memcpy(&ctx->msg.payload, msg_copy.u.payload, payload_size);
@@ -1222,6 +1224,7 @@ static void vmbus_force_channel_rescinded(struct vmbus_channel *channel)
rescind->child_relid = channel->offermsg.child_relid;
INIT_WORK(&ctx->work, vmbus_onmessage_work);
+ ctx->host_generated = false;
queue_work(vmbus_connection.work_queue, &ctx->work);
}
diff --git a/include/linux/hyperv.h b/include/linux/hyperv.h
index 2878aed14c45..096054fa07a3 100644
--- a/include/linux/hyperv.h
+++ b/include/linux/hyperv.h
@@ -809,6 +809,7 @@ struct vmbus_channel {
u8 monitor_bit;
bool rescind; /* got rescind msg */
+ bool rescind_from_host; /* host revocation, not local channel removal */
bool rescind_ref; /* got rescind msg, got channel reference */
struct completion rescind_event;
@@ -1117,7 +1118,8 @@ static inline void set_channel_pending_send_size(struct vmbus_channel *c,
c->outbound.ring_buffer->pending_send_sz = size;
}
-void vmbus_onmessage(struct vmbus_channel_message_header *hdr);
+void vmbus_onmessage(struct vmbus_channel_message_header *hdr,
+ bool host_generated);
int vmbus_request_offers(void);
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 08/14] hv: vmbus: retain backing until ownership and references clear
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (6 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 07/14] hv: vmbus: distinguish host rescind from local channel unload Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 09/14] hv: use owned VMBus buffers in NetVSC and UIO Emerson Busson
` (5 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Keep allocation and GPADL state in one retained owner. Unknown create or
teardown ownership and failed page transitions retain backing. A known
live handle may request teardown after a host rescind, but only an
actual matching reply releases host ownership; timeout and local removal
do not. Remove response waiters under their list lock before freeing
them.
Drain only pending reclaim work successfully acquired by cancellation. A
running callback retains custody and is never queued for another
execution. Native-workqueue and private-connection tests cover those
boundaries without mutating the live connection.
Rollback trigger: revert the lifecycle change if ordinary backing,
response identity or callback custody cannot be qualified; never free
unknown host-owned pages to satisfy a bound.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 757 ++++++++++++++++++++++++++-------
drivers/hv/channel_mgmt.c | 24 +-
drivers/hv/hv_trace.h | 24 ++
drivers/hv/hyperv_vmbus.h | 51 +++
drivers/hv/vmbus_buffer_test.c | 453 +++++++++++++++++++-
drivers/hv/vmbus_drv.c | 1 +
include/linux/hyperv.h | 18 +
7 files changed, 1171 insertions(+), 157 deletions(-)
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index 8ed155f2669f..aa33a0f08a7f 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -23,11 +23,37 @@
#include <linux/set_memory.h>
#include <linux/vmalloc.h>
#include <linux/export.h>
+#include <linux/list.h>
+#include <linux/mutex.h>
+#include <linux/workqueue.h>
#include <asm/page.h>
#include <asm/mshyperv.h>
#include "hyperv_vmbus.h"
+static LIST_HEAD(vmbus_buffer_owners);
+static DEFINE_MUTEX(vmbus_buffer_owners_lock);
+static struct workqueue_struct *vmbus_buffer_reclaim_wq;
+static bool vmbus_buffer_reclaimer_stopping;
+static atomic64_t vmbus_buffer_owner_sequence = ATOMIC64_INIT(0);
+
+/*
+ * Reclaim scheduling. VMBUS_BUFFER_RECLAIM_SCHEDULE_MS is the delay
+ * before the first reclaim attempt after ownership state changes;
+ * it is kept at one jiffy so a completed GPADL teardown frees the
+ * buffer promptly. VMBUS_BUFFER_RECLAIM_RETRY_MS is the steady-state
+ * retry while a guest mapping still holds the pages. The workqueue
+ * is unbound so reclaim never runs in the caller's context, and
+ * WQ_MEM_RECLAIM so the free path is not blocked by the very
+ * pressure it is relieving. max_active stays at one: re-encryption
+ * is serialized and the cost of a second concurrent worker is not
+ * worth the ordering questions it would raise.
+ */
+#define VMBUS_BUFFER_RECLAIM_SCHEDULE_MS 1
+#define VMBUS_BUFFER_RECLAIM_RETRY_MS 1000
+#define VMBUS_BUFFER_RECLAIM_WQ_FLAGS (WQ_UNBOUND | WQ_MEM_RECLAIM)
+#define VMBUS_BUFFER_RECLAIM_WQ_MAX_ACTIVE 1
+
/*
* vmbus_buffer_round_size() and the order-descent helpers are also called
* from vmbus_buffer_test.c, which is built into this same object. They are
@@ -98,11 +124,276 @@ bool vmbus_buffer_should_free(const struct vmbus_buffer *buffer)
!buffer->gpadl.gpadl_handle;
}
-static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
- u32 size,
- bool confidential,
- struct page ***chunks_out,
- u32 *chunk_cnt_out);
+/* Caller holds the owner lock or exclusive custody of this owner. */
+static void vmbus_buffer_trace_owner(const struct vmbus_buffer_retained *owner,
+ const char *action)
+{
+ /* Bit 0: host uncertainty; bit 1: page state; bit 2: permanent retain. */
+ u8 state = owner->host_may_own | (owner->encryption_unknown << 1) |
+ (owner->permanent_leak << 2);
+
+ trace_vmbus_buffer_owner(owner->owner_id, owner->channel_id, action,
+ owner->size, owner->page_cnt, state);
+}
+
+bool
+vmbus_buffer_owner_can_reclaim(const struct vmbus_buffer_retained *owner)
+{
+ return owner->released && !owner->permanent_leak &&
+ !owner->encryption_unknown &&
+ !owner->host_may_own;
+}
+
+bool
+vmbus_buffer_owner_should_schedule(const struct vmbus_buffer_retained *owner,
+ bool stopping, bool queue_live)
+{
+ return !owner->work_active && !owner->reclaiming &&
+ vmbus_buffer_owner_can_reclaim(owner) && !stopping && queue_live;
+}
+
+static void
+vmbus_buffer_schedule_reclaim_locked(struct vmbus_buffer_retained *owner)
+{
+ if (!vmbus_buffer_owner_should_schedule(owner,
+ vmbus_buffer_reclaimer_stopping,
+ vmbus_buffer_reclaim_wq))
+ return;
+
+ owner->work_active = true;
+ mod_delayed_work(vmbus_buffer_reclaim_wq, &owner->reclaim_work,
+ msecs_to_jiffies(VMBUS_BUFFER_RECLAIM_SCHEDULE_MS));
+}
+
+static void
+vmbus_buffer_update_host_ownership(struct vmbus_buffer_retained *owner,
+ bool host_may_own)
+{
+ mutex_lock(&vmbus_buffer_owners_lock);
+ owner->host_may_own = host_may_own;
+ vmbus_buffer_schedule_reclaim_locked(owner);
+ mutex_unlock(&vmbus_buffer_owners_lock);
+}
+
+/*
+ * Declared before vmbus_buffer_owner_alloc() hands the callback pointer
+ * to INIT_DELAYED_WORK(); defined below with the rest of the reclaim
+ * machinery.
+ */
+static void vmbus_buffer_reclaim_work(struct work_struct *work);
+
+struct vmbus_buffer_retained *
+vmbus_buffer_owner_alloc(struct vmbus_channel *channel)
+{
+ struct vmbus_buffer_retained *owner;
+
+ owner = kzalloc_obj(*owner);
+ if (!owner)
+ return NULL;
+
+ INIT_LIST_HEAD(&owner->list);
+ INIT_DELAYED_WORK(&owner->reclaim_work, vmbus_buffer_reclaim_work);
+ owner->channel_id = channel->lifetime_id;
+ owner->owner_id = atomic64_inc_return(&vmbus_buffer_owner_sequence);
+
+ mutex_lock(&vmbus_buffer_owners_lock);
+ if (vmbus_buffer_reclaimer_stopping)
+ goto err_unlock;
+ if (!vmbus_buffer_reclaim_wq) {
+ vmbus_buffer_reclaim_wq =
+ alloc_workqueue("vmbus-buffer-reclaim",
+ VMBUS_BUFFER_RECLAIM_WQ_FLAGS,
+ VMBUS_BUFFER_RECLAIM_WQ_MAX_ACTIVE);
+ if (!vmbus_buffer_reclaim_wq)
+ goto err_unlock;
+ }
+ list_add_tail(&owner->list, &vmbus_buffer_owners);
+ vmbus_buffer_trace_owner(owner, "created");
+ mutex_unlock(&vmbus_buffer_owners_lock);
+
+ return owner;
+
+err_unlock:
+ mutex_unlock(&vmbus_buffer_owners_lock);
+ kfree(owner);
+ return NULL;
+}
+
+static void vmbus_buffer_owner_free(struct vmbus_buffer_retained *owner)
+{
+ mutex_lock(&vmbus_buffer_owners_lock);
+ if (!list_empty(&owner->list))
+ list_del_init(&owner->list);
+ mutex_unlock(&vmbus_buffer_owners_lock);
+
+ kfree(owner);
+}
+
+static void vmbus_buffer_owner_remove(struct vmbus_buffer_retained *owner)
+{
+ /*
+ * The owner embeds a delayed_work and its timer may still be
+ * armed when an external caller drops the last reference.
+ * Freeing the object under an armed timer would let the
+ * callback reach freed memory, so stop the timer first. The
+ * reclaim work itself must not come through here: it would
+ * wait for its own completion.
+ */
+ cancel_delayed_work_sync(&owner->reclaim_work);
+ vmbus_buffer_trace_owner(owner, "discarded");
+ vmbus_buffer_owner_free(owner);
+}
+
+bool vmbus_buffer_pages_busy(struct vmbus_buffer_retained *owner)
+{
+ u32 i;
+
+ for (i = 0; i < owner->page_cnt; i++) {
+ if (WARN_ON_ONCE(!owner->pages || !owner->pages[i]))
+ return true;
+ if (folio_ref_count(page_folio(owner->pages[i])) != 1)
+ return true;
+ }
+
+ return false;
+}
+
+static void vmbus_buffer_reclaim_work(struct work_struct *work)
+{
+ struct vmbus_buffer_retained *owner = container_of(to_delayed_work(work),
+ struct vmbus_buffer_retained,
+ reclaim_work);
+ unsigned long delay = msecs_to_jiffies(VMBUS_BUFFER_RECLAIM_RETRY_MS);
+ u32 i;
+ int ret = 0;
+
+ mutex_lock(&vmbus_buffer_owners_lock);
+ /*
+ * A ready owner must still be reclaimable while the workqueue
+ * is being destroyed: shutdown re-queues exactly those so they
+ * drain before destroy_workqueue() returns. New work is never
+ * scheduled past that point because
+ * vmbus_buffer_owner_should_schedule() refuses once stopping.
+ */
+ if (!vmbus_buffer_owner_can_reclaim(owner)) {
+ owner->work_active = false;
+ mutex_unlock(&vmbus_buffer_owners_lock);
+ return;
+ }
+
+ if (vmbus_buffer_pages_busy(owner)) {
+ if (vmbus_buffer_reclaimer_stopping) {
+ /*
+ * A guest mapping still holds these pages as the
+ * workqueue goes away. Freeing them would be a
+ * use-after-free, so retain and say so.
+ */
+ owner->permanent_leak = true;
+ owner->work_active = false;
+ vmbus_buffer_trace_owner(owner, "retained");
+ } else {
+ mod_delayed_work(vmbus_buffer_reclaim_wq,
+ &owner->reclaim_work, delay);
+ }
+ mutex_unlock(&vmbus_buffer_owners_lock);
+ return;
+ }
+ owner->reclaiming = true;
+ mutex_unlock(&vmbus_buffer_owners_lock);
+
+ if (owner->needs_encrypt) {
+ for (i = 0; i < owner->chunk_cnt; i++) {
+ struct page *page = owner->chunks[i];
+ unsigned int order = folio_order(page_folio(page));
+
+ ret = set_memory_encrypted((unsigned long)page_address(page),
+ 1U << order);
+ if (ret)
+ break;
+ }
+ } else if (owner->raw_decrypted) {
+ ret = set_memory_encrypted((unsigned long)owner->addr,
+ PFN_UP(owner->size));
+ }
+
+ if (ret) {
+ mutex_lock(&vmbus_buffer_owners_lock);
+ owner->permanent_leak = true;
+ owner->work_active = false;
+ owner->reclaiming = false;
+ vmbus_buffer_trace_owner(owner, "retained");
+ mutex_unlock(&vmbus_buffer_owners_lock);
+ pr_warn_ratelimited("VMBus buffer reclaim retained pages after encryption failure: %d\n",
+ ret);
+ return;
+ }
+
+ vmbus_buffer_trace_owner(owner, "reclaimed");
+ if (owner->addr) {
+ if (owner->chunks)
+ vunmap(owner->addr);
+ else
+ vfree(owner->addr);
+ }
+
+ for (i = 0; i < owner->chunk_cnt; i++) {
+ struct page *page = owner->chunks[i];
+ unsigned int order = folio_order(page_folio(page));
+
+ __free_pages(page, order);
+ }
+
+ kvfree(owner->chunks);
+ kvfree(owner->pages);
+ vmbus_buffer_owner_free(owner);
+}
+
+/* Caller holds the owner lock or exclusive custody of this owner. */
+void vmbus_buffer_owner_drain(struct vmbus_buffer_retained *owner,
+ struct workqueue_struct *wq)
+{
+ if (!cancel_delayed_work(&owner->reclaim_work))
+ return;
+
+ owner->work_active = false;
+ if (wq && vmbus_buffer_owner_can_reclaim(owner)) {
+ owner->work_active = true;
+ mod_delayed_work(wq, &owner->reclaim_work, 0);
+ }
+}
+
+void vmbus_buffer_reclaimer_shutdown(void)
+{
+ struct workqueue_struct *wq;
+ struct vmbus_buffer_retained *owner;
+
+ mutex_lock(&vmbus_buffer_owners_lock);
+ vmbus_buffer_reclaimer_stopping = true;
+ wq = vmbus_buffer_reclaim_wq;
+ vmbus_buffer_reclaim_wq = NULL;
+
+ /*
+ * destroy_workqueue() drains work that is queued or already
+ * running. A delayed_work still waiting on its timer is not
+ * on the workqueue yet, so cancel those timers here or they
+ * fire against a workqueue that is about to be freed.
+ * cancel_delayed_work() rather than _sync: reclaim work that
+ * is already running blocks on this same lock, and
+ * destroy_workqueue() below waits for it to finish.
+ *
+ * An owner that is already ready to reclaim must not be lost
+ * with its cancelled timer. Requeue only work that was actually
+ * cancelled: a running callback owns its embedded work item and
+ * may free the owner before a second queued execution can run.
+ * destroy_workqueue() drains that running callback itself.
+ */
+ list_for_each_entry(owner, &vmbus_buffer_owners, list)
+ vmbus_buffer_owner_drain(owner, wq);
+ mutex_unlock(&vmbus_buffer_owners_lock);
+
+ if (wq)
+ destroy_workqueue(wq);
+}
/*
* hv_gpadl_size - Return the real size of a gpadl, the size that Hyper-V uses
@@ -248,30 +539,20 @@ int vmbus_alloc_ring(struct vmbus_channel *newchannel,
{
struct vmbus_buffer *buffer = &newchannel->ringbuffer;
u32 size;
- u32 i;
+ int ret;
if (!send_size || !recv_size ||
send_size % PAGE_SIZE || recv_size % PAGE_SIZE ||
check_add_overflow(send_size, recv_size, &size))
return -EINVAL;
- buffer->addr = __vmbus_alloc_buffer(newchannel, size,
- newchannel->co_ring_buffer,
- &buffer->chunks, &buffer->chunk_cnt);
- if (!buffer->addr)
- return -ENOMEM;
+ ret = vmbus_alloc_buffer_owned(newchannel, size,
+ newchannel->co_ring_buffer, buffer);
+ if (ret)
+ return ret;
newchannel->ringbuffer_pagecount = size >> PAGE_SHIFT;
newchannel->ringbuffer_send_offset = send_size >> PAGE_SHIFT;
- buffer->pages = kvcalloc(newchannel->ringbuffer_pagecount,
- sizeof(*buffer->pages), GFP_KERNEL);
- if (!buffer->pages) {
- vmbus_release_buffer(buffer);
- return -ENOMEM;
- }
-
- for (i = 0; i < newchannel->ringbuffer_pagecount; i++)
- buffer->pages[i] = vmalloc_to_page(buffer->addr + (i << PAGE_SHIFT));
return 0;
}
@@ -571,12 +852,15 @@ int vmbus_post_gpadl_messages(struct vmbus_channel_msginfo *msginfo,
int vmbus_gpadl_response_status(u32 creation_status, bool rescind,
bool *posted)
{
- *posted = false;
- if (creation_status)
+ if (creation_status) {
+ *posted = false;
return -EDQUOT;
+ }
if (rescind)
return -ENODEV;
+ /* A response resolved the request; a successful GPADL is tracked by handle. */
+ *posted = false;
return 0;
}
@@ -599,21 +883,22 @@ int vmbus_post_gpadl_teardown(u32 child_relid,
}
static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
- enum hv_gpadl_type type, void *kbuffer,
- u32 size, u32 send_offset, bool memory_prepared,
- bool *leak,
- struct vmbus_gpadl *gpadl)
+ enum hv_gpadl_type type,
+ struct vmbus_buffer *buffer,
+ u32 send_offset, bool memory_prepared)
{
struct vmbus_channel_gpadl_header *gpadlmsg;
struct vmbus_channel_msginfo *msginfo = NULL;
+ struct vmbus_buffer_retained *owner = buffer->owner;
+ struct vmbus_gpadl *gpadl = &buffer->gpadl;
+ void *kbuffer = buffer->addr;
+ u32 size = buffer->size;
u32 next_gpadl_handle;
u32 creation_status;
unsigned long flags;
bool posted = false;
int ret = 0;
- if (leak)
- *leak = false;
gpadl->leak = false;
next_gpadl_handle =
@@ -641,6 +926,9 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
dev_warn(&channel->device_obj->device,
"Failed to set host visibility for new GPADL %d.\n",
ret);
+ if (owner)
+ owner->encryption_unknown = true;
+ gpadl->leak = true;
vmbus_free_channel_msginfo(msginfo);
return ret;
}
@@ -668,6 +956,8 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
ret = vmbus_post_gpadl_messages(msginfo, next_gpadl_handle, &posted,
vmbus_gpadl_post_real, NULL);
+ if (posted && owner)
+ vmbus_buffer_update_host_ownership(owner, true);
if (ret)
goto cleanup;
@@ -687,6 +977,8 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
gpadl->gpadl_handle = gpadlmsg->gpadl;
gpadl->buffer = kbuffer;
gpadl->size = size;
+ if (owner)
+ vmbus_buffer_update_host_ownership(owner, true);
cleanup:
spin_lock_irqsave(&vmbus_connection.channelmsg_lock, flags);
@@ -695,40 +987,44 @@ static int __vmbus_establish_gpadl(struct vmbus_channel *channel,
vmbus_free_channel_msginfo(msginfo);
- if (ret && posted) {
+ if (ret && posted)
gpadl->leak = true;
- if (leak)
- *leak = true;
- }
-
- if (ret && !posted) {
- /*
- * If set_memory_encrypted() fails, the decrypted flag is
- * left as true so the memory is leaked instead of being
- * put back on the free list.
- */
- if (gpadl->decrypted) {
- if (!set_memory_encrypted((unsigned long)kbuffer, PFN_UP(size)))
- gpadl->decrypted = false;
- }
- }
+ else if (ret && owner)
+ vmbus_buffer_update_host_ownership(owner, false);
return ret;
}
/*
- * vmbus_establish_gpadl - Establish a GPADL for the specified buffer
+ * vmbus_establish_gpadl_owned - Establish a GPADL for an owned buffer whose
+ * memory has already been prepared by the allocator.
*
* @channel: a channel
- * @kbuffer: from kmalloc or vmalloc
- * @size: page-size multiple
- * @gpadl: output gpadl
+ * @buffer: allocated by vmbus_alloc_buffer_owned()
*/
+int vmbus_establish_gpadl_owned(struct vmbus_channel *channel,
+ struct vmbus_buffer *buffer)
+{
+ return __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER, buffer,
+ 0U, true);
+}
+EXPORT_SYMBOL_GPL(vmbus_establish_gpadl_owned);
+
+/* Preserve the established exported API for non-owned caller buffers. */
int vmbus_establish_gpadl(struct vmbus_channel *channel, void *kbuffer,
u32 size, struct vmbus_gpadl *gpadl)
{
- return __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER, kbuffer, size,
- 0U, false, &gpadl->leak, gpadl);
+ struct vmbus_buffer buffer = {
+ .addr = kbuffer,
+ .size = size,
+ .gpadl = *gpadl,
+ };
+ int ret;
+
+ ret = __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER, &buffer,
+ 0U, false);
+ *gpadl = buffer.gpadl;
+ return ret;
}
EXPORT_SYMBOL_GPL(vmbus_establish_gpadl);
@@ -746,11 +1042,20 @@ EXPORT_SYMBOL_GPL(vmbus_establish_gpadl);
* The caller is responsible for re-encrypting the buffer before freeing it.
*/
int vmbus_establish_gpadl_caller_decrypted(struct vmbus_channel *channel,
- void *kbuffer, u32 size,
- struct vmbus_gpadl *gpadl)
+ void *kbuffer, u32 size,
+ struct vmbus_gpadl *gpadl)
{
- return __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER,
- kbuffer, size, 0U, true, &gpadl->leak, gpadl);
+ struct vmbus_buffer buffer = {
+ .addr = kbuffer,
+ .size = size,
+ .gpadl = *gpadl,
+ };
+ int ret;
+
+ ret = __vmbus_establish_gpadl(channel, HV_GPADL_BUFFER, &buffer,
+ 0U, true);
+ *gpadl = buffer.gpadl;
+ return ret;
}
EXPORT_SYMBOL_GPL(vmbus_establish_gpadl_caller_decrypted);
@@ -793,16 +1098,51 @@ void vmbus_free_buffer(void *addr, struct page **chunks, u32 chunk_cnt)
}
EXPORT_SYMBOL_GPL(vmbus_free_buffer);
+/* Retain backing until the host and any guest mappings have released it. */
void vmbus_release_buffer(struct vmbus_buffer *buffer)
{
- if (!buffer->addr)
+ struct vmbus_buffer_retained *owner = buffer->owner;
+
+ if (!buffer->addr && !buffer->chunks && !buffer->pages) {
+ if (owner)
+ vmbus_buffer_owner_remove(owner);
+ memset(buffer, 0, sizeof(*buffer));
return;
+ }
+
+ if (!owner) {
+ if (vmbus_buffer_should_free(buffer)) {
+ if (buffer->chunks)
+ vmbus_free_buffer(buffer->addr, buffer->chunks,
+ buffer->chunk_cnt);
+ else
+ vfree(buffer->addr);
+ } else {
+ pr_warn_ratelimited("VMBus buffer has no ownership record; retaining backing pages\n");
+ }
+ memset(buffer, 0, sizeof(*buffer));
+ return;
+ }
+
+ mutex_lock(&vmbus_buffer_owners_lock);
+ owner->addr = buffer->addr;
+ owner->chunks = buffer->chunks;
+ owner->pages = buffer->pages;
+ owner->chunk_cnt = buffer->chunk_cnt;
+ owner->page_cnt = buffer->page_cnt;
+ owner->size = buffer->size;
+ owner->host_may_own |= buffer->gpadl.gpadl_handle ||
+ buffer->gpadl.leak;
+ owner->raw_decrypted |= buffer->gpadl.decrypted;
+ owner->permanent_leak |= buffer->leak;
+ owner->released = true;
+ vmbus_buffer_trace_owner(owner, "released");
+ if (!vmbus_buffer_owner_can_reclaim(owner))
+ vmbus_buffer_trace_owner(owner, "retained");
- kvfree(buffer->pages);
- if (vmbus_buffer_should_free(buffer))
- vmbus_free_buffer(buffer->addr, buffer->chunks,
- buffer->chunk_cnt);
memset(buffer, 0, sizeof(*buffer));
+ vmbus_buffer_schedule_reclaim_locked(owner);
+ mutex_unlock(&vmbus_buffer_owners_lock);
}
EXPORT_SYMBOL_GPL(vmbus_release_buffer);
@@ -815,14 +1155,13 @@ static struct page *vmbus_alloc_pages_node(void *context, int nid,
}
/**
- * __vmbus_alloc_buffer - allocate host-visible, virtually-contiguous backing.
+ * vmbus_alloc_buffer_owned - allocate a host-visible, virtually-contiguous
+ * buffer with VMBus-managed backing-page lifetime.
*
* @channel: the channel the buffer will be attached to
* @size: requested buffer size in bytes (will be rounded up to PAGE_SIZE)
* @confidential: keep the buffer private to the guest
- * @chunks_out: on success, set to the array of underlying chunks, or NULL when
- * the buffer was allocated with vzalloc()
- * @chunk_cnt_out: on success, set to the number of chunks
+ * @buffer: output descriptor that owns the allocation and its backing pages
*
* Buffers not requiring decryption are allocated with vzalloc().
*
@@ -832,51 +1171,67 @@ static struct page *vmbus_alloc_pages_node(void *context, int nid,
* host-visible via set_memory_decrypted() on its direct-map address, then all
* chunks are combined into a virtually-contiguous range via vmap().
*
- * Return: the buffer's virtual address, or NULL on failure.
+ * Return: 0 on success, or a negative error code.
*/
-static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
- u32 size,
- bool confidential,
- struct page ***chunks_out,
- u32 *chunk_cnt_out)
+int vmbus_alloc_buffer_owned(struct vmbus_channel *channel, u32 size,
+ bool confidential, struct vmbus_buffer *buffer)
{
u32 rounded_size;
unsigned long nr_pages;
unsigned long remaining;
unsigned long page_idx = 0;
+ struct vmbus_buffer_retained *owner;
+ unsigned int order = MAX_PAGE_ORDER;
bool hv_isolation;
bool guest_mem_encrypted;
- struct page **chunks = NULL;
- struct page **pages = NULL;
- unsigned int order = MAX_PAGE_ORDER;
- u32 chunk_cnt = 0;
- void *addr;
u32 i;
int ret;
- *chunks_out = NULL;
- *chunk_cnt_out = 0;
+ memset(buffer, 0, sizeof(*buffer));
- if (vmbus_buffer_round_size(size, &rounded_size))
- return NULL;
+ ret = vmbus_buffer_round_size(size, &rounded_size);
+ if (ret)
+ return ret;
nr_pages = rounded_size >> PAGE_SHIFT;
remaining = nr_pages;
-
+ owner = vmbus_buffer_owner_alloc(channel);
+ if (!owner)
+ return -ENOMEM;
+ buffer->owner = owner;
+ buffer->size = rounded_size;
+ owner->size = rounded_size;
hv_isolation = hv_is_isolation_supported();
guest_mem_encrypted = cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT);
/* If the buffer does not need to be decrypted, just use vzalloc() */
if (!vmbus_needs_shared_pages(hv_isolation, guest_mem_encrypted,
- confidential))
- return vzalloc(rounded_size);
+ confidential)) {
+ buffer->addr = vzalloc(rounded_size);
+ if (!buffer->addr) {
+ vmbus_release_buffer(buffer);
+ return -ENOMEM;
+ }
+ buffer->pages = kvcalloc(nr_pages, sizeof(*buffer->pages), GFP_KERNEL);
+ if (!buffer->pages) {
+ vmbus_release_buffer(buffer);
+ return -ENOMEM;
+ }
+ for (page_idx = 0; page_idx < nr_pages; page_idx++)
+ buffer->pages[page_idx] =
+ vmalloc_to_page(buffer->addr +
+ (page_idx << PAGE_SHIFT));
+ buffer->page_cnt = nr_pages;
+ return 0;
+ }
/* Worst case: every chunk is a single page. */
- chunks = kvmalloc_objs(*chunks, nr_pages, GFP_KERNEL | __GFP_ZERO);
- if (!chunks)
+ buffer->chunks = kvmalloc_array(nr_pages, sizeof(*buffer->chunks),
+ GFP_KERNEL | __GFP_ZERO);
+ if (!buffer->chunks)
goto err;
- pages = kvmalloc_objs(*pages, nr_pages);
- if (!pages)
+ buffer->pages = kvmalloc_array(nr_pages, sizeof(*buffer->pages), GFP_KERNEL);
+ if (!buffer->pages)
goto err;
while (remaining) {
@@ -894,6 +1249,11 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
if (!page)
goto err;
+ buffer->chunks[buffer->chunk_cnt++] = page;
+ for (i = 0; i < (1U << order); i++)
+ buffer->pages[page_idx++] = page + i;
+ buffer->page_cnt = page_idx;
+
ret = set_memory_decrypted((unsigned long)page_address(page),
1U << order);
if (ret) {
@@ -901,40 +1261,56 @@ static void *__vmbus_alloc_buffer(struct vmbus_channel *channel,
* set_memory_decrypted() failed; the page state is
* unknown so it must be leaked rather than freed.
*/
+ owner->encryption_unknown = true;
goto err;
}
-
- chunks[chunk_cnt++] = page;
-
- for (i = 0; i < (1U << order); i++)
- pages[page_idx++] = page + i;
+ owner->needs_encrypt = true;
remaining -= 1U << order;
}
- addr = vmap(pages, nr_pages, VM_MAP, pgprot_decrypted(PAGE_KERNEL));
- if (!addr)
+ buffer->addr = vmap(buffer->pages, nr_pages, VM_MAP,
+ pgprot_decrypted(PAGE_KERNEL));
+ if (!buffer->addr)
goto err;
- memset(addr, 0, rounded_size);
-
- kvfree(pages);
- *chunks_out = chunks;
- *chunk_cnt_out = chunk_cnt;
- return addr;
+ memset(buffer->addr, 0, rounded_size);
+ return 0;
err:
- kvfree(pages);
- vmbus_free_buffer(NULL, chunks, chunk_cnt);
- return NULL;
+ vmbus_release_buffer(buffer);
+ return -ENOMEM;
}
-
-void *vmbus_alloc_buffer(struct vmbus_channel *channel,
- u32 size, struct page ***chunks_out,
- u32 *chunk_cnt_out)
+EXPORT_SYMBOL_GPL(vmbus_alloc_buffer_owned);
+/*
+ * vmbus_alloc_buffer - compatibility allocator for callers managing lifetime.
+ * New callers that need retained GPADL and mmap ownership should use
+ * vmbus_alloc_buffer_owned().
+ */
+void *vmbus_alloc_buffer(struct vmbus_channel *channel, u32 size,
+ struct page ***chunks_out, u32 *chunk_cnt_out)
{
- return __vmbus_alloc_buffer(channel, size, channel->co_external_memory,
- chunks_out, chunk_cnt_out);
+ struct vmbus_buffer buffer = {};
+ void *addr;
+ int ret;
+
+ if (!chunks_out || !chunk_cnt_out)
+ return NULL;
+
+ *chunks_out = NULL;
+ *chunk_cnt_out = 0;
+ ret = vmbus_alloc_buffer_owned(channel, size,
+ channel->co_external_memory, &buffer);
+ if (ret)
+ return NULL;
+
+ addr = buffer.addr;
+ *chunks_out = buffer.chunks;
+ *chunk_cnt_out = buffer.chunk_cnt;
+ kvfree(buffer.pages);
+ vmbus_buffer_owner_remove(buffer.owner);
+
+ return addr;
}
EXPORT_SYMBOL_GPL(vmbus_alloc_buffer);
@@ -1038,11 +1414,9 @@ static int __vmbus_open(struct vmbus_channel *newchannel,
/* Establish the gpadl for the ring buffer */
buffer->gpadl.gpadl_handle = 0;
- err = __vmbus_establish_gpadl(newchannel, HV_GPADL_RING,
- buffer->addr,
- (send_pages + recv_pages) << PAGE_SHIFT,
+ err = __vmbus_establish_gpadl(newchannel, HV_GPADL_RING, buffer,
newchannel->ringbuffer_send_offset << PAGE_SHIFT,
- true, &buffer->leak, &buffer->gpadl);
+ true);
if (err)
goto error_clean_ring;
@@ -1133,8 +1507,7 @@ static int __vmbus_open(struct vmbus_channel *newchannel,
error_free_info:
kfree(open_info);
error_free_gpadl:
- if (vmbus_teardown_gpadl(newchannel, &buffer->gpadl))
- buffer->leak = true;
+ vmbus_teardown_gpadl_owned(newchannel, buffer);
error_clean_ring:
hv_ringbuffer_cleanup(&newchannel->outbound);
hv_ringbuffer_cleanup(&newchannel->inbound);
@@ -1180,62 +1553,155 @@ EXPORT_SYMBOL_GPL(vmbus_open);
/*
* vmbus_teardown_gpadl -Teardown the specified GPADL handle
*/
-int vmbus_teardown_gpadl(struct vmbus_channel *channel, struct vmbus_gpadl *gpadl)
+int vmbus_gpadl_teardown_request(struct vmbus_channel *channel,
+ struct vmbus_buffer *buffer,
+ vmbus_gpadl_info_alloc_fn alloc_info,
+ struct vmbus_connection *connection,
+ vmbus_gpadl_teardown_post_fn post,
+ unsigned long timeout)
{
struct vmbus_channel_gpadl_teardown *msg;
struct vmbus_channel_msginfo *info;
+ struct vmbus_gpadl *gpadl = &buffer->gpadl;
+ struct vmbus_buffer_retained *owner = buffer->owner;
unsigned long flags;
- int ret;
+ int ret = -ENODEV;
+ u32 handle = gpadl->gpadl_handle;
+ u32 relid = channel->offermsg.child_relid;
+
+ if (!handle && !gpadl->leak)
+ return 0;
+
+ /* Host rescind permits a request, but only the actual ACK releases it. */
+ if (relid == INVALID_RELID ||
+ (READ_ONCE(channel->rescind) &&
+ !READ_ONCE(channel->rescind_from_host)))
+ goto retain;
+
+ if (!handle) {
+ ret = -EINPROGRESS;
+ goto retain;
+ }
- info = kzalloc(sizeof(*info) +
- sizeof(struct vmbus_channel_gpadl_teardown), GFP_KERNEL);
+ info = alloc_info();
if (!info) {
- gpadl->leak = true;
- return -ENOMEM;
+ ret = -ENOMEM;
+ goto retain;
}
init_completion(&info->waitevent);
- info->waiting_channel = channel;
+ /* A synthetic rescind completion must not stand in for the host reply. */
+ info->waiting_channel = NULL;
msg = (struct vmbus_channel_gpadl_teardown *)info->msg;
+ msg->header.msgtype = CHANNELMSG_GPADL_TEARDOWN;
+ msg->child_relid = relid;
+ msg->gpadl = handle;
- spin_lock_irqsave(&vmbus_connection.channelmsg_lock, flags);
- list_add_tail(&info->msglistentry,
- &vmbus_connection.chn_msg_list);
- spin_unlock_irqrestore(&vmbus_connection.channelmsg_lock, flags);
+ spin_lock_irqsave(&connection->channelmsg_lock, flags);
+ list_add_tail(&info->msglistentry, &connection->chn_msg_list);
+ spin_unlock_irqrestore(&connection->channelmsg_lock, flags);
- if (channel->rescind)
- goto post_msg_err;
+ if (channel->offermsg.child_relid != relid ||
+ (READ_ONCE(channel->rescind) &&
+ !READ_ONCE(channel->rescind_from_host)))
+ goto cleanup;
- ret = vmbus_post_gpadl_teardown(channel->offermsg.child_relid, msg, gpadl->gpadl_handle,
- vmbus_gpadl_post_real, NULL);
+ ret = post(connection, info);
if (ret)
- goto post_msg_err;
-
- wait_for_completion(&info->waitevent);
+ goto cleanup;
- gpadl->gpadl_handle = 0;
+ /* A lost transport is bounded; expiry retains unknown host ownership. */
+ if (!wait_for_completion_timeout(&info->waitevent, timeout)) {
+ ret = -ETIMEDOUT;
+ goto cleanup;
+ }
-post_msg_err:
- /*
- * If the channel has been rescinded;
- * we will be awakened by the rescind
- * handler; set the error code to zero so we don't leak memory.
- */
- if (channel->rescind)
+ if (info->response.gpadl_torndown.header.msgtype ==
+ CHANNELMSG_GPADL_TORNDOWN &&
+ info->response.gpadl_torndown.gpadl == handle) {
ret = 0;
+ gpadl->gpadl_handle = 0;
+ gpadl->leak = false;
+ if (owner)
+ vmbus_buffer_update_host_ownership(owner, false);
+ goto cleanup;
+ }
- spin_lock_irqsave(&vmbus_connection.channelmsg_lock, flags);
+ ret = -ENODEV;
+
+cleanup:
+ spin_lock_irqsave(&connection->channelmsg_lock, flags);
list_del(&info->msglistentry);
- spin_unlock_irqrestore(&vmbus_connection.channelmsg_lock, flags);
+ spin_unlock_irqrestore(&connection->channelmsg_lock, flags);
kfree(info);
- if (!ret && gpadl->decrypted) {
- int encrypt_ret;
+retain:
+ if (ret) {
+ gpadl->leak = true;
+ if (owner)
+ vmbus_buffer_update_host_ownership(owner, true);
+ }
+
+ return ret;
+}
+static struct vmbus_channel_msginfo *vmbus_alloc_teardown_info(void)
+{
+ return kzalloc(sizeof(struct vmbus_channel_msginfo) +
+ sizeof(struct vmbus_channel_gpadl_teardown), GFP_KERNEL);
+}
+
+static int vmbus_gpadl_teardown_post_real(struct vmbus_connection *connection,
+ struct vmbus_channel_msginfo *info)
+{
+ struct vmbus_channel_gpadl_teardown *msg = (void *)info->msg;
+
+ return vmbus_post_gpadl_teardown(msg->child_relid, msg, msg->gpadl,
+ vmbus_gpadl_post_real, NULL);
+}
+
+static int __vmbus_teardown_gpadl_buffer(struct vmbus_channel *channel,
+ struct vmbus_buffer *buffer)
+{
+ return vmbus_gpadl_teardown_request(channel, buffer,
+ vmbus_alloc_teardown_info,
+ &vmbus_connection,
+ vmbus_gpadl_teardown_post_real, 5 * HZ);
+}
+
+int vmbus_teardown_gpadl_owned(struct vmbus_channel *channel,
+ struct vmbus_buffer *buffer)
+{
+ return __vmbus_teardown_gpadl_buffer(channel, buffer);
+}
+EXPORT_SYMBOL_GPL(vmbus_teardown_gpadl_owned);
+
+int vmbus_teardown_gpadl(struct vmbus_channel *channel,
+ struct vmbus_gpadl *gpadl)
+{
+ struct vmbus_buffer buffer = {
+ .gpadl = *gpadl,
+ };
+ int encrypt_ret;
+ int ret;
+
+ ret = __vmbus_teardown_gpadl_buffer(channel, &buffer);
+ *gpadl = buffer.gpadl;
+
+ /*
+ * The compat entry point builds an owner-less buffer, so the
+ * reclaim path never sees it and cannot re-encrypt on its
+ * behalf. Restore the synchronous re-encryption this symbol
+ * has always performed: the caller is about to release the
+ * range, and returning it decrypted would put shared pages
+ * back on the free list. The owned path defers the same work
+ * to reclaim through owner->raw_decrypted.
+ */
+ if (!ret && gpadl->decrypted) {
encrypt_ret = set_memory_encrypted((unsigned long)gpadl->buffer,
PFN_UP(gpadl->size));
if (encrypt_ret) {
@@ -1245,8 +1711,6 @@ int vmbus_teardown_gpadl(struct vmbus_channel *channel, struct vmbus_gpadl *gpad
}
gpadl->decrypted = !!encrypt_ret;
}
- if (ret)
- gpadl->leak = true;
return ret;
}
@@ -1323,9 +1787,8 @@ static int vmbus_close_internal(struct vmbus_channel *channel)
/* Tear down the gpadl for the channel's ring buffer */
else if (channel->ringbuffer.gpadl.gpadl_handle) {
- ret = vmbus_teardown_gpadl(channel, &channel->ringbuffer.gpadl);
+ ret = vmbus_teardown_gpadl_owned(channel, &channel->ringbuffer);
if (ret) {
- channel->ringbuffer.leak = true;
pr_err("Close failed: teardown gpadl return %d\n", ret);
/*
* If we failed to teardown gpadl,
diff --git a/drivers/hv/channel_mgmt.c b/drivers/hv/channel_mgmt.c
index 93fc105cd179..be0f70a3541c 100644
--- a/drivers/hv/channel_mgmt.c
+++ b/drivers/hv/channel_mgmt.c
@@ -26,6 +26,8 @@
#include "hyperv_vmbus.h"
+static atomic64_t vmbus_channel_lifetime_id = ATOMIC64_INIT(0);
+
static void init_vp_index(struct vmbus_channel *channel);
const struct vmbus_device vmbus_devs[] = {
@@ -958,6 +960,7 @@ EXPORT_SYMBOL_GPL(vmbus_initiate_unload);
static void vmbus_setup_channel_state(struct vmbus_channel *channel,
struct vmbus_channel_offer_channel *offer)
{
+ channel->lifetime_id = atomic64_inc_return(&vmbus_channel_lifetime_id);
WRITE_ONCE(channel->rescind, false);
WRITE_ONCE(channel->rescind_from_host, false);
@@ -1484,26 +1487,23 @@ static void vmbus_onmodifychannel_response(struct vmbus_channel_message_header *
* Find the matching request, copy the response and signal the requesting
* thread.
*/
-static void vmbus_ongpadl_torndown(
- struct vmbus_channel_message_header *hdr)
+void vmbus_complete_gpadl_teardown(struct vmbus_connection *connection,
+ struct vmbus_channel_gpadl_torndown *gpadl_torndown)
{
- struct vmbus_channel_gpadl_torndown *gpadl_torndown;
struct vmbus_channel_msginfo *msginfo;
struct vmbus_channel_message_header *requestheader;
struct vmbus_channel_gpadl_teardown *gpadl_teardown;
unsigned long flags;
- gpadl_torndown = (struct vmbus_channel_gpadl_torndown *)hdr;
-
trace_vmbus_ongpadl_torndown(gpadl_torndown);
/*
* Find the open msg, copy the result and signal/unblock the wait event
*/
- spin_lock_irqsave(&vmbus_connection.channelmsg_lock, flags);
+ spin_lock_irqsave(&connection->channelmsg_lock, flags);
- list_for_each_entry(msginfo, &vmbus_connection.chn_msg_list,
- msglistentry) {
+ list_for_each_entry(msginfo, &connection->chn_msg_list,
+ msglistentry) {
requestheader =
(struct vmbus_channel_message_header *)msginfo->msg;
@@ -1521,7 +1521,13 @@ static void vmbus_ongpadl_torndown(
}
}
}
- spin_unlock_irqrestore(&vmbus_connection.channelmsg_lock, flags);
+ spin_unlock_irqrestore(&connection->channelmsg_lock, flags);
+}
+
+static void vmbus_ongpadl_torndown(struct vmbus_channel_message_header *hdr)
+{
+ vmbus_complete_gpadl_teardown(&vmbus_connection,
+ (struct vmbus_channel_gpadl_torndown *)hdr);
}
/*
diff --git a/drivers/hv/hv_trace.h b/drivers/hv/hv_trace.h
index c02a1719e92f..a3eeb817b00e 100644
--- a/drivers/hv/hv_trace.h
+++ b/drivers/hv/hv_trace.h
@@ -8,6 +8,30 @@
#include <linux/tracepoint.h>
+/* Opaque allocation identities; never expose a kernel pointer. */
+TRACE_EVENT(vmbus_buffer_owner,
+ TP_PROTO(u64 owner_id, u64 channel_id, const char *action,
+ u32 size, u32 pages, u8 state),
+ TP_ARGS(owner_id, channel_id, action, size, pages, state),
+ TP_STRUCT__entry(__field(u64, owner_id)
+ __field(u64, channel_id)
+ __string(action, action)
+ __field(u32, size)
+ __field(u32, pages)
+ __field(u8, state)
+ ),
+ TP_fast_assign(__entry->owner_id = owner_id;
+ __entry->channel_id = channel_id;
+ __assign_str(action);
+ __entry->size = size;
+ __entry->pages = pages;
+ __entry->state = state;
+ ),
+ TP_printk("owner_id=%llu channel_id=%llu action=%s size=%u pages=%u state=%u",
+ __entry->owner_id, __entry->channel_id, __get_str(action),
+ __entry->size, __entry->pages, __entry->state)
+);
+
DECLARE_EVENT_CLASS(vmbus_hdr_msg,
TP_PROTO(const struct vmbus_channel_message_header *hdr),
TP_ARGS(hdr),
diff --git a/drivers/hv/hyperv_vmbus.h b/drivers/hv/hyperv_vmbus.h
index 06094000f2f1..2edeb7988bdc 100644
--- a/drivers/hv/hyperv_vmbus.h
+++ b/drivers/hv/hyperv_vmbus.h
@@ -354,6 +354,7 @@ struct vmbus_channel_message_table_entry {
extern const struct vmbus_channel_message_table_entry
channel_message_table[CHANNELMSG_COUNT];
+void vmbus_buffer_reclaimer_shutdown(void);
/* General vmbus interface */
@@ -551,6 +552,31 @@ int hv_create_ring_sysfs(struct vmbus_channel *channel,
struct vm_area_desc *desc));
int hv_remove_ring_sysfs(struct vmbus_channel *channel);
+/*
+ * Retained buffer owner. One per channel-keyed allocation; freed only
+ * after GPADL, page-state and mapping-reference gates all clear.
+ */
+struct vmbus_buffer_retained {
+ struct list_head list;
+ struct delayed_work reclaim_work;
+ u64 channel_id;
+ u64 owner_id;
+ void *addr;
+ struct page **chunks;
+ struct page **pages;
+ u32 chunk_cnt;
+ u32 page_cnt;
+ u32 size;
+ bool released;
+ bool host_may_own;
+ bool needs_encrypt;
+ bool raw_decrypted;
+ bool encryption_unknown;
+ bool permanent_leak;
+ bool work_active;
+ bool reclaiming;
+};
+
/*
* vmbus buffer sizing, order-descent and free-decision helpers.
*
@@ -578,6 +604,20 @@ struct page *vmbus_alloc_pages_with_fallback(int nid, gfp_t gfp,
vmbus_alloc_pages_fn alloc,
void *context);
+/*
+ * Owner lifetime and reclaim-gate helpers, shared with
+ * vmbus_buffer_test.c for the same reason as the sizing helpers.
+ * Defined in channel.c, unexported.
+ */
+bool vmbus_buffer_owner_can_reclaim(const struct vmbus_buffer_retained *owner);
+bool vmbus_buffer_owner_should_schedule(const struct vmbus_buffer_retained *owner,
+ bool stopping, bool queue_live);
+bool vmbus_buffer_pages_busy(struct vmbus_buffer_retained *owner);
+void vmbus_buffer_owner_drain(struct vmbus_buffer_retained *owner,
+ struct workqueue_struct *wq);
+struct vmbus_buffer_retained *
+vmbus_buffer_owner_alloc(struct vmbus_channel *channel);
+
/*
* GPADL post and response helpers, also shared with vmbus_buffer_test.c.
* Same deal as the sizing helpers above: defined in channel.c, built into
@@ -596,5 +636,16 @@ int vmbus_post_gpadl_teardown(u32 child_relid,
u32 gpadl,
vmbus_gpadl_post_fn post_msg,
void *context);
+typedef struct vmbus_channel_msginfo *(*vmbus_gpadl_info_alloc_fn)(void);
+typedef int (*vmbus_gpadl_teardown_post_fn)(struct vmbus_connection *connection,
+ struct vmbus_channel_msginfo *info);
+int vmbus_gpadl_teardown_request(struct vmbus_channel *channel,
+ struct vmbus_buffer *buffer,
+ vmbus_gpadl_info_alloc_fn alloc_info,
+ struct vmbus_connection *connection,
+ vmbus_gpadl_teardown_post_fn post,
+ unsigned long timeout);
+void vmbus_complete_gpadl_teardown(struct vmbus_connection *connection,
+ struct vmbus_channel_gpadl_torndown *response);
#endif /* _HYPERV_VMBUS_H */
diff --git a/drivers/hv/vmbus_buffer_test.c b/drivers/hv/vmbus_buffer_test.c
index 9b401ef2ceea..5c8e70d861ad 100644
--- a/drivers/hv/vmbus_buffer_test.c
+++ b/drivers/hv/vmbus_buffer_test.c
@@ -7,6 +7,7 @@
* without exporting them.
*/
#include <kunit/test.h>
+#include <linux/completion.h>
#include <linux/hyperv.h>
#include <linux/mm.h>
#include <linux/slab.h>
@@ -66,6 +67,439 @@ static void vmbus_buffer_private_shared_selection_test(struct kunit *test)
cases[i].shared);
}
+static void vmbus_buffer_owner_reclaim_gate_test(struct kunit *test)
+{
+ struct vmbus_buffer_retained owner = {
+ .released = false,
+ };
+
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_owner_can_reclaim(&owner));
+ owner.released = true;
+ KUNIT_EXPECT_TRUE(test, vmbus_buffer_owner_can_reclaim(&owner));
+
+ owner.host_may_own = true;
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_owner_can_reclaim(&owner));
+
+ owner.host_may_own = false; /* GPADL teardown acknowledgment */
+ KUNIT_EXPECT_TRUE(test, vmbus_buffer_owner_can_reclaim(&owner));
+
+ owner.permanent_leak = true;
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_owner_can_reclaim(&owner));
+
+ owner.permanent_leak = false;
+ owner.encryption_unknown = true;
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_owner_can_reclaim(&owner));
+}
+
+static void vmbus_buffer_reclaim_schedule_gate_test(struct kunit *test)
+{
+ struct vmbus_buffer_retained owner = {
+ .released = true,
+ };
+
+ KUNIT_EXPECT_TRUE(test,
+ vmbus_buffer_owner_should_schedule(&owner, false, true));
+ owner.work_active = true;
+ KUNIT_EXPECT_FALSE(test,
+ vmbus_buffer_owner_should_schedule(&owner, false, true));
+ owner.work_active = false;
+ owner.reclaiming = true;
+ KUNIT_EXPECT_FALSE(test,
+ vmbus_buffer_owner_should_schedule(&owner, false, true));
+ owner.reclaiming = false;
+ KUNIT_EXPECT_FALSE(test,
+ vmbus_buffer_owner_should_schedule(&owner, true, true));
+ KUNIT_EXPECT_FALSE(test,
+ vmbus_buffer_owner_should_schedule(&owner, false, false));
+}
+
+struct vmbus_reclaim_test_work {
+ struct vmbus_buffer_retained owner;
+ struct workqueue_struct *wq;
+ struct completion entered;
+ struct completion proceed;
+ atomic_t calls;
+ bool claim_owner;
+};
+
+static void vmbus_reclaim_test_callback(struct work_struct *work)
+{
+ struct vmbus_buffer_retained *owner =
+ container_of(to_delayed_work(work),
+ struct vmbus_buffer_retained, reclaim_work);
+ struct vmbus_reclaim_test_work *ctx =
+ container_of(owner, struct vmbus_reclaim_test_work, owner);
+
+ atomic_inc(&ctx->calls);
+ if (ctx->claim_owner)
+ owner->reclaiming = true;
+ complete(&ctx->entered);
+ wait_for_completion(&ctx->proceed);
+}
+
+static void vmbus_reclaim_test_cleanup(void *data)
+{
+ struct vmbus_reclaim_test_work *ctx = data;
+
+ complete_all(&ctx->proceed);
+ cancel_delayed_work_sync(&ctx->owner.reclaim_work);
+ destroy_workqueue(ctx->wq);
+}
+
+static struct vmbus_reclaim_test_work *
+vmbus_reclaim_test_init(struct kunit *test)
+{
+ struct vmbus_reclaim_test_work *ctx;
+
+ ctx = kunit_kzalloc(test, sizeof(*ctx), GFP_KERNEL);
+ if (!ctx)
+ return NULL;
+ init_completion(&ctx->entered);
+ init_completion(&ctx->proceed);
+ atomic_set(&ctx->calls, 0);
+ ctx->owner.released = true;
+ ctx->owner.work_active = true;
+ INIT_DELAYED_WORK(&ctx->owner.reclaim_work,
+ vmbus_reclaim_test_callback);
+ ctx->wq = alloc_workqueue("vmbus-reclaim-test",
+ WQ_UNBOUND | WQ_MEM_RECLAIM, 1);
+ if (!ctx->wq)
+ return NULL;
+ if (kunit_add_action_or_reset(test, vmbus_reclaim_test_cleanup, ctx))
+ return NULL;
+ return ctx;
+}
+
+static void vmbus_reclaim_running_test(struct kunit *test, bool claim_owner)
+{
+ struct vmbus_reclaim_test_work *ctx = vmbus_reclaim_test_init(test);
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ ctx->claim_owner = claim_owner;
+ KUNIT_ASSERT_TRUE(test,
+ queue_delayed_work(ctx->wq, &ctx->owner.reclaim_work, 0));
+ KUNIT_ASSERT_NE(test,
+ wait_for_completion_timeout(&ctx->entered,
+ msecs_to_jiffies(1000)), 0UL);
+
+ /*
+ * The native workqueue cleared pending before entering the callback.
+ * Test both windows around the callback taking ownership. Keep the
+ * fixture alive so an incorrect second queue fails without a UAF.
+ */
+ vmbus_buffer_owner_drain(&ctx->owner, ctx->wq);
+ KUNIT_EXPECT_FALSE(test, delayed_work_pending(&ctx->owner.reclaim_work));
+ KUNIT_EXPECT_TRUE(test, ctx->owner.work_active);
+ complete_all(&ctx->proceed);
+ flush_workqueue(ctx->wq);
+ KUNIT_EXPECT_EQ(test, atomic_read(&ctx->calls), 1);
+}
+
+static void vmbus_reclaim_shutdown_before_claim_test(struct kunit *test)
+{
+ vmbus_reclaim_running_test(test, false);
+}
+
+static void vmbus_reclaim_shutdown_during_reclaim_test(struct kunit *test)
+{
+ vmbus_reclaim_running_test(test, true);
+}
+
+static void vmbus_reclaim_shutdown_pending_test(struct kunit *test)
+{
+ struct vmbus_reclaim_test_work *ctx = vmbus_reclaim_test_init(test);
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ KUNIT_ASSERT_TRUE(test,
+ queue_delayed_work(ctx->wq, &ctx->owner.reclaim_work,
+ msecs_to_jiffies(60000)));
+ complete_all(&ctx->proceed);
+ vmbus_buffer_owner_drain(&ctx->owner, ctx->wq);
+ flush_workqueue(ctx->wq);
+ KUNIT_EXPECT_EQ(test, atomic_read(&ctx->calls), 1);
+}
+
+static void vmbus_reclaim_shutdown_unsafe_pending_test(struct kunit *test)
+{
+ struct vmbus_reclaim_test_work *ctx = vmbus_reclaim_test_init(test);
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ ctx->owner.host_may_own = true;
+ KUNIT_ASSERT_TRUE(test,
+ queue_delayed_work(ctx->wq, &ctx->owner.reclaim_work,
+ msecs_to_jiffies(60000)));
+ vmbus_buffer_owner_drain(&ctx->owner, ctx->wq);
+ KUNIT_EXPECT_FALSE(test, delayed_work_pending(&ctx->owner.reclaim_work));
+ KUNIT_EXPECT_FALSE(test, ctx->owner.work_active);
+ flush_workqueue(ctx->wq);
+ KUNIT_EXPECT_EQ(test, atomic_read(&ctx->calls), 0);
+}
+
+static struct vmbus_channel_msginfo *vmbus_test_alloc_teardown_fail(void)
+{
+ return NULL;
+}
+
+static void vmbus_buffer_rescind_retains_gpadl_test(struct kunit *test)
+{
+ struct vmbus_channel channel = { .rescind = true };
+ struct vmbus_buffer_retained owner = {
+ .released = true,
+ .host_may_own = true,
+ };
+ struct vmbus_buffer buffer = { .owner = &owner };
+ unsigned int origin, pending;
+ u32 handle;
+ int ret;
+
+ /* Exercise the request core without touching the live connection. */
+ for (origin = 0; origin < 2; origin++) {
+ channel.rescind_from_host = origin;
+ for (pending = 0; pending < 2; pending++) {
+ buffer.gpadl.gpadl_handle = pending ? 0 : 17;
+ buffer.gpadl.leak = pending;
+ handle = buffer.gpadl.gpadl_handle;
+ ret = vmbus_gpadl_teardown_request(&channel, &buffer,
+ vmbus_test_alloc_teardown_fail,
+ NULL, NULL, 1);
+ KUNIT_EXPECT_EQ(test, ret, !origin ? -ENODEV :
+ pending ? -EINPROGRESS : -ENOMEM);
+ KUNIT_EXPECT_EQ(test, buffer.gpadl.gpadl_handle, handle);
+ KUNIT_EXPECT_TRUE(test, buffer.gpadl.leak);
+ KUNIT_EXPECT_TRUE(test, owner.host_may_own);
+ KUNIT_EXPECT_FALSE(test, owner.work_active);
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_should_free(&buffer));
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_owner_can_reclaim(&owner));
+ }
+ }
+}
+
+enum vmbus_test_teardown_reply {
+ VMBUS_TEST_REPLY_ACK,
+ VMBUS_TEST_REPLY_NONE,
+ VMBUS_TEST_REPLY_WRONG_HANDLE,
+ VMBUS_TEST_REPLY_WRONG_TYPE,
+ VMBUS_TEST_REPLY_POST_FAILURE,
+};
+
+struct vmbus_teardown_test_context {
+ struct vmbus_connection connection;
+ struct vmbus_channel channel;
+ struct vmbus_buffer buffer;
+ struct vmbus_buffer_retained owner;
+ struct kunit *test;
+ enum vmbus_test_teardown_reply reply;
+ unsigned int post_calls;
+};
+
+static struct vmbus_channel_msginfo *vmbus_test_alloc_teardown(void)
+{
+ return kzalloc(sizeof(struct vmbus_channel_msginfo) +
+ sizeof(struct vmbus_channel_gpadl_teardown), GFP_KERNEL);
+}
+
+static int vmbus_test_post_teardown(struct vmbus_connection *connection,
+ struct vmbus_channel_msginfo *info)
+{
+ struct vmbus_teardown_test_context *ctx =
+ container_of(connection, struct vmbus_teardown_test_context,
+ connection);
+ struct vmbus_channel_gpadl_teardown *msg = (void *)info->msg;
+ struct vmbus_channel_gpadl_torndown response = {
+ .header.msgtype = CHANNELMSG_GPADL_TORNDOWN,
+ .gpadl = msg->gpadl,
+ };
+
+ ctx->post_calls++;
+ KUNIT_EXPECT_EQ(ctx->test, msg->header.msgtype, CHANNELMSG_GPADL_TEARDOWN);
+ KUNIT_EXPECT_EQ(ctx->test, msg->child_relid, 71U);
+ KUNIT_EXPECT_EQ(ctx->test, msg->gpadl, 51U);
+ KUNIT_EXPECT_PTR_EQ(ctx->test, info->waiting_channel, NULL);
+ if (ctx->reply == VMBUS_TEST_REPLY_POST_FAILURE)
+ return -EIO;
+ if (ctx->reply == VMBUS_TEST_REPLY_NONE)
+ return 0;
+ if (ctx->reply == VMBUS_TEST_REPLY_WRONG_HANDLE)
+ response.gpadl++;
+ if (ctx->reply == VMBUS_TEST_REPLY_WRONG_TYPE)
+ response.header.msgtype = CHANNELMSG_GPADL_CREATED;
+ vmbus_complete_gpadl_teardown(connection, &response);
+ return 0;
+}
+
+static struct vmbus_teardown_test_context *
+vmbus_test_teardown_context(struct kunit *test)
+{
+ struct vmbus_teardown_test_context *ctx;
+
+ ctx = kunit_kzalloc(test, sizeof(*ctx), GFP_KERNEL);
+ if (!ctx)
+ return NULL;
+ ctx->test = test;
+ ctx->buffer.gpadl.gpadl_handle = 51;
+ ctx->buffer.owner = &ctx->owner;
+ ctx->owner.host_may_own = true;
+ ctx->channel.rescind = true;
+ ctx->channel.rescind_from_host = true;
+ ctx->channel.offermsg.child_relid = 71;
+ INIT_LIST_HEAD(&ctx->connection.chn_msg_list);
+ spin_lock_init(&ctx->connection.channelmsg_lock);
+ return ctx;
+}
+
+static int vmbus_test_teardown_request(struct vmbus_teardown_test_context *ctx)
+{
+ return vmbus_gpadl_teardown_request(&ctx->channel, &ctx->buffer,
+ vmbus_test_alloc_teardown,
+ &ctx->connection,
+ vmbus_test_post_teardown, 1);
+}
+
+static void vmbus_host_rescind_teardown_ack_test(struct kunit *test)
+{
+ struct vmbus_teardown_test_context *ctx = vmbus_test_teardown_context(test);
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ KUNIT_EXPECT_EQ(test, vmbus_test_teardown_request(ctx), 0);
+ KUNIT_EXPECT_EQ(test, ctx->post_calls, 1U);
+ KUNIT_EXPECT_EQ(test, ctx->buffer.gpadl.gpadl_handle, 0U);
+ KUNIT_EXPECT_FALSE(test, ctx->buffer.gpadl.leak);
+ KUNIT_EXPECT_FALSE(test, ctx->owner.host_may_own);
+ KUNIT_EXPECT_TRUE(test, vmbus_buffer_should_free(&ctx->buffer));
+ KUNIT_EXPECT_TRUE(test, list_empty(&ctx->connection.chn_msg_list));
+}
+
+static void vmbus_test_teardown_failure(struct kunit *test,
+ enum vmbus_test_teardown_reply reply,
+ int expected)
+{
+ struct vmbus_teardown_test_context *ctx = vmbus_test_teardown_context(test);
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ ctx->reply = reply;
+ KUNIT_EXPECT_EQ(test, vmbus_test_teardown_request(ctx), expected);
+ KUNIT_EXPECT_EQ(test, ctx->post_calls, 1U);
+ KUNIT_EXPECT_EQ(test, ctx->buffer.gpadl.gpadl_handle, 51U);
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_should_free(&ctx->buffer));
+ KUNIT_EXPECT_TRUE(test, ctx->owner.host_may_own);
+ KUNIT_EXPECT_TRUE(test, list_empty(&ctx->connection.chn_msg_list));
+}
+
+static void vmbus_host_rescind_teardown_timeout_test(struct kunit *test)
+{
+ vmbus_test_teardown_failure(test, VMBUS_TEST_REPLY_NONE, -ETIMEDOUT);
+}
+
+static void vmbus_host_rescind_teardown_wrong_handle_test(struct kunit *test)
+{
+ vmbus_test_teardown_failure(test, VMBUS_TEST_REPLY_WRONG_HANDLE,
+ -ETIMEDOUT);
+}
+
+static void vmbus_host_rescind_teardown_wrong_type_test(struct kunit *test)
+{
+ vmbus_test_teardown_failure(test, VMBUS_TEST_REPLY_WRONG_TYPE, -ENODEV);
+}
+
+static void vmbus_host_rescind_teardown_post_failure_test(struct kunit *test)
+{
+ vmbus_test_teardown_failure(test, VMBUS_TEST_REPLY_POST_FAILURE, -EIO);
+}
+
+static void vmbus_local_rescind_skips_teardown_test(struct kunit *test)
+{
+ struct vmbus_teardown_test_context *ctx = vmbus_test_teardown_context(test);
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ ctx->channel.rescind_from_host = false;
+ KUNIT_EXPECT_EQ(test, vmbus_test_teardown_request(ctx), -ENODEV);
+ KUNIT_EXPECT_EQ(test, ctx->post_calls, 0U);
+ KUNIT_EXPECT_TRUE(test, ctx->buffer.gpadl.leak);
+ KUNIT_EXPECT_TRUE(test, ctx->owner.host_may_own);
+ KUNIT_EXPECT_EQ(test, ctx->buffer.gpadl.gpadl_handle, 51U);
+}
+
+static void vmbus_invalid_relid_skips_teardown_test(struct kunit *test)
+{
+ struct vmbus_teardown_test_context *ctx = vmbus_test_teardown_context(test);
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ ctx->channel.offermsg.child_relid = INVALID_RELID;
+ KUNIT_EXPECT_EQ(test, vmbus_test_teardown_request(ctx), -ENODEV);
+ KUNIT_EXPECT_EQ(test, ctx->post_calls, 0U);
+ KUNIT_EXPECT_EQ(test, ctx->buffer.gpadl.gpadl_handle, 51U);
+}
+
+static void vmbus_host_rescind_teardown_late_ack_test(struct kunit *test)
+{
+ struct vmbus_teardown_test_context *ctx = vmbus_test_teardown_context(test);
+ struct vmbus_channel_gpadl_torndown response = {
+ .header.msgtype = CHANNELMSG_GPADL_TORNDOWN,
+ .gpadl = 51,
+ };
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ ctx->reply = VMBUS_TEST_REPLY_NONE;
+ KUNIT_ASSERT_EQ(test, vmbus_test_teardown_request(ctx), -ETIMEDOUT);
+ /* Run the production response matcher after its waiter has been freed. */
+ vmbus_complete_gpadl_teardown(&ctx->connection, &response);
+ KUNIT_EXPECT_TRUE(test, list_empty(&ctx->connection.chn_msg_list));
+ KUNIT_EXPECT_EQ(test, ctx->buffer.gpadl.gpadl_handle, 51U);
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_should_free(&ctx->buffer));
+ KUNIT_EXPECT_TRUE(test, ctx->owner.host_may_own);
+}
+
+static void vmbus_buffer_mapping_reference_test(struct kunit *test)
+{
+ struct vmbus_buffer_retained owner = {};
+ struct page *page;
+ struct page *pages[1];
+
+ page = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+ pages[0] = page;
+ owner.pages = pages;
+ owner.page_cnt = ARRAY_SIZE(pages);
+
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_pages_busy(&owner));
+ get_page(page); /* vm_insert_pages() holds one reference per mapping */
+ KUNIT_EXPECT_TRUE(test, vmbus_buffer_pages_busy(&owner));
+ put_page(page);
+ KUNIT_EXPECT_FALSE(test, vmbus_buffer_pages_busy(&owner));
+ __free_page(page);
+}
+
+static void vmbus_buffer_repeated_owner_release_test(struct kunit *test)
+{
+ struct vmbus_channel channel = {};
+ struct vmbus_buffer buffer = {};
+ struct vmbus_buffer empty = {};
+
+ buffer.addr = vzalloc(PAGE_SIZE);
+ KUNIT_ASSERT_NOT_NULL(test, buffer.addr);
+ buffer.owner = vmbus_buffer_owner_alloc(&channel);
+ if (!buffer.owner) {
+ vfree(buffer.addr);
+ KUNIT_FAIL(test, "failed to allocate a VMBus buffer owner");
+ return;
+ }
+ buffer.size = PAGE_SIZE;
+ empty.owner = vmbus_buffer_owner_alloc(&channel);
+ if (!empty.owner) {
+ vmbus_release_buffer(&buffer);
+ KUNIT_FAIL(test, "failed to allocate a second VMBus buffer owner");
+ return;
+ }
+ KUNIT_EXPECT_TRUE(test, empty.owner->owner_id != buffer.owner->owner_id);
+ vmbus_release_buffer(&empty);
+ KUNIT_EXPECT_PTR_EQ(test, empty.owner, NULL);
+
+ vmbus_release_buffer(&buffer);
+ KUNIT_EXPECT_PTR_EQ(test, buffer.addr, NULL);
+ KUNIT_EXPECT_PTR_EQ(test, buffer.owner, NULL);
+ vmbus_release_buffer(&buffer);
+}
+
static void vmbus_ring_fallback_order_zero_test(struct kunit *test)
{
unsigned int order;
@@ -265,7 +699,7 @@ static void vmbus_gpadl_response_state_test(struct kunit *test)
posted = true;
KUNIT_EXPECT_EQ(test,
vmbus_gpadl_response_status(0, true, &posted), -ENODEV);
- KUNIT_EXPECT_FALSE(test, posted);
+ KUNIT_EXPECT_TRUE(test, posted);
}
static void vmbus_gpadl_teardown_post_failure_test(struct kunit *test)
@@ -348,7 +782,24 @@ static struct kunit_case vmbus_buffer_test_cases[] = {
KUNIT_CASE(vmbus_buffer_private_shared_selection_test),
KUNIT_CASE(vmbus_ring_fallback_order_zero_test),
KUNIT_CASE(vmbus_buffer_failed_teardown_leaks_test),
+ KUNIT_CASE(vmbus_buffer_owner_reclaim_gate_test),
+ KUNIT_CASE(vmbus_buffer_reclaim_schedule_gate_test),
+ KUNIT_CASE(vmbus_reclaim_shutdown_before_claim_test),
+ KUNIT_CASE(vmbus_reclaim_shutdown_during_reclaim_test),
+ KUNIT_CASE(vmbus_reclaim_shutdown_pending_test),
+ KUNIT_CASE(vmbus_reclaim_shutdown_unsafe_pending_test),
+ KUNIT_CASE(vmbus_buffer_rescind_retains_gpadl_test),
+ KUNIT_CASE(vmbus_host_rescind_teardown_ack_test),
+ KUNIT_CASE(vmbus_host_rescind_teardown_timeout_test),
+ KUNIT_CASE(vmbus_host_rescind_teardown_wrong_handle_test),
+ KUNIT_CASE(vmbus_host_rescind_teardown_wrong_type_test),
+ KUNIT_CASE(vmbus_host_rescind_teardown_post_failure_test),
+ KUNIT_CASE(vmbus_local_rescind_skips_teardown_test),
+ KUNIT_CASE(vmbus_invalid_relid_skips_teardown_test),
+ KUNIT_CASE(vmbus_host_rescind_teardown_late_ack_test),
+ KUNIT_CASE(vmbus_buffer_mapping_reference_test),
KUNIT_CASE(vmbus_buffer_partial_allocation_cleanup_test),
+ KUNIT_CASE(vmbus_buffer_repeated_owner_release_test),
KUNIT_CASE(vmbus_buffer_order_zero_allocation_test),
KUNIT_CASE(vmbus_gpadl_post_failure_test),
KUNIT_CASE(vmbus_gpadl_post_success_test),
diff --git a/drivers/hv/vmbus_drv.c b/drivers/hv/vmbus_drv.c
index 723252f1b551..bce835c4a015 100644
--- a/drivers/hv/vmbus_drv.c
+++ b/drivers/hv/vmbus_drv.c
@@ -3075,6 +3075,7 @@ static void __exit vmbus_exit(void)
&hyperv_panic_vmbus_unload_block);
bus_unregister(&hv_bus);
+ vmbus_buffer_reclaimer_shutdown();
cpuhp_remove_state(hyperv_cpuhp_online);
hv_synic_free();
diff --git a/include/linux/hyperv.h b/include/linux/hyperv.h
index 096054fa07a3..90bdbacee054 100644
--- a/include/linux/hyperv.h
+++ b/include/linux/hyperv.h
@@ -784,13 +784,18 @@ struct vmbus_gpadl {
bool leak;
};
+struct vmbus_buffer_retained;
+
struct vmbus_buffer {
void *addr;
struct page **chunks;
struct page **pages;
u32 chunk_cnt;
+ u32 page_cnt;
+ u32 size;
struct vmbus_gpadl gpadl;
bool leak;
+ struct vmbus_buffer_retained *owner;
};
struct vmbus_channel {
@@ -811,6 +816,7 @@ struct vmbus_channel {
bool rescind; /* got rescind msg */
bool rescind_from_host; /* host revocation, not local channel removal */
bool rescind_ref; /* got rescind msg, got channel reference */
+ u64 lifetime_id;
struct completion rescind_event;
/* Allocated memory for ring buffer */
@@ -1228,6 +1234,18 @@ extern void *vmbus_alloc_buffer(struct vmbus_channel *channel,
u32 *chunk_cnt_out);
extern void vmbus_free_buffer(void *addr, struct page **chunks, u32 chunk_cnt);
+
+int vmbus_establish_gpadl_owned(struct vmbus_channel *channel,
+ struct vmbus_buffer *buffer);
+
+int vmbus_teardown_gpadl_owned(struct vmbus_channel *channel,
+ struct vmbus_buffer *buffer);
+
+int vmbus_alloc_buffer_owned(struct vmbus_channel *channel,
+ u32 size,
+ bool confidential,
+ struct vmbus_buffer *buffer);
+
void vmbus_release_buffer(struct vmbus_buffer *buffer);
void vmbus_reset_channel_cb(struct vmbus_channel *channel);
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 09/14] hv: use owned VMBus buffers in NetVSC and UIO
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (7 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 08/14] hv: vmbus: retain backing until ownership and references clear Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 10/14] hv: vmbus: pin buffer pages across UIO mmap to close the reclaim race Emerson Busson
` (4 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Move the remaining buffer consumers to the allocation, GPADL and release
entry points that carry one descriptor throughout their lifetime.
Preserve established exported signatures throughout the series. Ring
confidentiality comes from co_ring_buffer; external-buffer compatibility
uses co_external_memory, while UIO buffers explicitly require
host-visible backing.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 4 +--
drivers/net/hyperv/netvsc.c | 52 +++++++++++-------------------------
drivers/uio/uio_hv_generic.c | 32 +++++++---------------
include/linux/hyperv.h | 8 +++---
4 files changed, 32 insertions(+), 64 deletions(-)
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index aa33a0f08a7f..7eb9ea814ef4 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -1035,11 +1035,10 @@ EXPORT_SYMBOL_GPL(vmbus_establish_gpadl);
* @channel: a channel
* @kbuffer: from kmalloc or vmalloc; must already be decrypted by the caller
* @size: page-size multiple
- * @leak: set when a GPADL message may have reached the host but completion is
- * uncertain; the caller must retain the backing pages
* @gpadl: output gpadl
*
* The caller is responsible for re-encrypting the buffer before freeing it.
+ * Ownership of the backing pages stays with the caller's vmbus_buffer.
*/
int vmbus_establish_gpadl_caller_decrypted(struct vmbus_channel *channel,
void *kbuffer, u32 size,
@@ -1282,6 +1281,7 @@ int vmbus_alloc_buffer_owned(struct vmbus_channel *channel, u32 size,
return -ENOMEM;
}
EXPORT_SYMBOL_GPL(vmbus_alloc_buffer_owned);
+
/*
* vmbus_alloc_buffer - compatibility allocator for callers managing lifetime.
* New callers that need retained GPADL and mmap ownership should use
diff --git a/drivers/net/hyperv/netvsc.c b/drivers/net/hyperv/netvsc.c
index ffba1443396a..fd8aa7a3dcb3 100644
--- a/drivers/net/hyperv/netvsc.c
+++ b/drivers/net/hyperv/netvsc.c
@@ -243,7 +243,6 @@ static void netvsc_revoke_recv_buf(struct hv_device *device,
if (ret != 0) {
netdev_err(ndev, "unable to send "
"revoke receive buffer to netvsp\n");
- net_device->recv_buffer.leak = true;
return;
}
net_device->recv_section_cnt = 0;
@@ -295,7 +294,6 @@ static void netvsc_revoke_send_buf(struct hv_device *device,
if (ret != 0) {
netdev_err(ndev, "unable to send "
"revoke send buffer to netvsp\n");
- net_device->send_buffer.leak = true;
return;
}
net_device->send_section_cnt = 0;
@@ -308,18 +306,14 @@ static void netvsc_teardown_recv_gpadl(struct hv_device *device,
{
int ret;
- if (net_device->recv_buffer.leak)
- return;
-
if (net_device->recv_buffer.gpadl.gpadl_handle) {
- ret = vmbus_teardown_gpadl(device->channel,
- &net_device->recv_buffer.gpadl);
+ ret = vmbus_teardown_gpadl_owned(device->channel,
+ &net_device->recv_buffer);
/* If we failed here, we might as well return and have a leak
* rather than continue and a bugchk
*/
if (ret != 0) {
- net_device->recv_buffer.leak = true;
netdev_err(ndev,
"unable to teardown receive buffer's gpadl\n");
return;
@@ -333,18 +327,14 @@ static void netvsc_teardown_send_gpadl(struct hv_device *device,
{
int ret;
- if (net_device->send_buffer.leak)
- return;
-
if (net_device->send_buffer.gpadl.gpadl_handle) {
- ret = vmbus_teardown_gpadl(device->channel,
- &net_device->send_buffer.gpadl);
+ ret = vmbus_teardown_gpadl_owned(device->channel,
+ &net_device->send_buffer);
/* If we failed here, we might as well return and have a leak
* rather than continue and a bugchk
*/
if (ret != 0) {
- net_device->send_buffer.leak = true;
netdev_err(ndev,
"unable to teardown send buffer's gpadl\n");
return;
@@ -385,15 +375,13 @@ static int netvsc_init_buf(struct hv_device *device,
buf_size = min_t(unsigned int, buf_size,
NETVSC_RECEIVE_BUFFER_SIZE_LEGACY);
- net_device->recv_buffer.addr =
- vmbus_alloc_buffer(device->channel, buf_size,
- &net_device->recv_buffer.chunks,
- &net_device->recv_buffer.chunk_cnt);
- if (!net_device->recv_buffer.addr) {
+ ret = vmbus_alloc_buffer_owned(device->channel, buf_size,
+ device->channel->co_external_memory,
+ &net_device->recv_buffer);
+ if (ret) {
netdev_err(ndev,
"unable to allocate receive buffer of size %u\n",
buf_size);
- ret = -ENOMEM;
goto cleanup;
}
@@ -404,11 +392,8 @@ static int netvsc_init_buf(struct hv_device *device,
* channel. Note: This call uses the vmbus connection rather
* than the channel to establish the gpadl handle.
*/
- ret = vmbus_establish_gpadl_caller_decrypted(device->channel,
- net_device->recv_buffer.addr,
- buf_size,
- &net_device->recv_buffer.gpadl);
- net_device->recv_buffer.leak |= net_device->recv_buffer.gpadl.leak;
+ ret = vmbus_establish_gpadl_owned(device->channel,
+ &net_device->recv_buffer);
if (ret != 0) {
netdev_err(ndev,
"unable to establish receive buffer's gpadl\n");
@@ -496,14 +481,12 @@ static int netvsc_init_buf(struct hv_device *device,
buf_size = device_info->send_sections * device_info->send_section_size;
buf_size = round_up(buf_size, PAGE_SIZE);
- net_device->send_buffer.addr =
- vmbus_alloc_buffer(device->channel, buf_size,
- &net_device->send_buffer.chunks,
- &net_device->send_buffer.chunk_cnt);
- if (!net_device->send_buffer.addr) {
+ ret = vmbus_alloc_buffer_owned(device->channel, buf_size,
+ device->channel->co_external_memory,
+ &net_device->send_buffer);
+ if (ret) {
netdev_err(ndev, "unable to allocate send buffer of size %u\n",
buf_size);
- ret = -ENOMEM;
goto cleanup;
}
net_device->send_buf_size = buf_size;
@@ -512,11 +495,8 @@ static int netvsc_init_buf(struct hv_device *device,
* channel. Note: This call uses the vmbus connection rather
* than the channel to establish the gpadl handle.
*/
- ret = vmbus_establish_gpadl_caller_decrypted(device->channel,
- net_device->send_buffer.addr,
- buf_size,
- &net_device->send_buffer.gpadl);
- net_device->send_buffer.leak |= net_device->send_buffer.gpadl.leak;
+ ret = vmbus_establish_gpadl_owned(device->channel,
+ &net_device->send_buffer);
if (ret != 0) {
netdev_err(ndev,
"unable to establish send buffer's gpadl\n");
diff --git a/drivers/uio/uio_hv_generic.c b/drivers/uio/uio_hv_generic.c
index 2cb95b4786ca..b40e80e19c6c 100644
--- a/drivers/uio/uio_hv_generic.c
+++ b/drivers/uio/uio_hv_generic.c
@@ -196,11 +196,11 @@ static void
hv_uio_cleanup(struct hv_device *dev, struct hv_uio_private_data *pdata)
{
if (pdata->send_buffer.gpadl.gpadl_handle)
- vmbus_teardown_gpadl(dev->channel, &pdata->send_buffer.gpadl);
+ vmbus_teardown_gpadl_owned(dev->channel, &pdata->send_buffer);
vmbus_release_buffer(&pdata->send_buffer);
if (pdata->recv_buffer.gpadl.gpadl_handle)
- vmbus_teardown_gpadl(dev->channel, &pdata->recv_buffer.gpadl);
+ vmbus_teardown_gpadl_owned(dev->channel, &pdata->recv_buffer);
vmbus_release_buffer(&pdata->recv_buffer);
}
@@ -298,18 +298,12 @@ hv_uio_probe(struct hv_device *dev,
pdata->info.mem[MON_PAGE_MAP].memtype = UIO_MEM_LOGICAL;
if (channel->device_id == HV_NIC) {
- pdata->recv_buffer.addr =
- vzalloc(RECV_BUFFER_SIZE);
- if (!pdata->recv_buffer.addr) {
- ret = -ENOMEM;
+ ret = vmbus_alloc_buffer_owned(channel, RECV_BUFFER_SIZE, false,
+ &pdata->recv_buffer);
+ if (ret)
goto fail_free_ring;
- }
- ret = vmbus_establish_gpadl(channel,
- pdata->recv_buffer.addr,
- RECV_BUFFER_SIZE,
- &pdata->recv_buffer.gpadl);
- pdata->recv_buffer.leak |= pdata->recv_buffer.gpadl.leak;
+ ret = vmbus_establish_gpadl_owned(channel, &pdata->recv_buffer);
if (ret)
goto fail_close;
@@ -322,18 +316,12 @@ hv_uio_probe(struct hv_device *dev,
pdata->info.mem[RECV_BUF_MAP].size = RECV_BUFFER_SIZE;
pdata->info.mem[RECV_BUF_MAP].memtype = UIO_MEM_VIRTUAL;
- pdata->send_buffer.addr =
- vzalloc(SEND_BUFFER_SIZE);
- if (!pdata->send_buffer.addr) {
- ret = -ENOMEM;
+ ret = vmbus_alloc_buffer_owned(channel, SEND_BUFFER_SIZE, false,
+ &pdata->send_buffer);
+ if (ret)
goto fail_close;
- }
- ret = vmbus_establish_gpadl(channel,
- pdata->send_buffer.addr,
- SEND_BUFFER_SIZE,
- &pdata->send_buffer.gpadl);
- pdata->send_buffer.leak |= pdata->send_buffer.gpadl.leak;
+ ret = vmbus_establish_gpadl_owned(channel, &pdata->send_buffer);
if (ret)
goto fail_close;
diff --git a/include/linux/hyperv.h b/include/linux/hyperv.h
index 90bdbacee054..0f482ed5c776 100644
--- a/include/linux/hyperv.h
+++ b/include/linux/hyperv.h
@@ -1216,9 +1216,9 @@ extern int vmbus_sendpacket_mpb_desc(struct vmbus_channel *channel,
u64 requestid);
extern int vmbus_establish_gpadl(struct vmbus_channel *channel,
- void *kbuffer,
- u32 size,
- struct vmbus_gpadl *gpadl);
+ void *kbuffer,
+ u32 size,
+ struct vmbus_gpadl *gpadl);
extern int vmbus_establish_gpadl_caller_decrypted(struct vmbus_channel *channel,
void *kbuffer,
@@ -1226,7 +1226,7 @@ extern int vmbus_establish_gpadl_caller_decrypted(struct vmbus_channel *channel,
struct vmbus_gpadl *gpadl);
extern int vmbus_teardown_gpadl(struct vmbus_channel *channel,
- struct vmbus_gpadl *gpadl);
+ struct vmbus_gpadl *gpadl);
extern void *vmbus_alloc_buffer(struct vmbus_channel *channel,
u32 size,
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 10/14] hv: vmbus: pin buffer pages across UIO mmap to close the reclaim race
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (8 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 09/14] hv: use owned VMBus buffers in NetVSC and UIO Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 11/14] hv: vmbus: vmalloc requestor metadata Emerson Busson
` (3 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
UIO maps the channel ring, the send/receive buffers, the interrupt
page and the monitor page directly into userspace. Those maps are
built from page arrays that the reclaim worker can free once the
buffer owner is released: the snapshot taken in the mmap path and
the moment the new VMA holds its own references are not the same
instant, so a reclaim in that window frees pages a mapping is about
to publish.
Add vmbus_buffer_pin_pages() and vmbus_buffer_unpin_pages(). Pinning
takes a reference on every page of a vmbus_buffer and refuses to
start once reclaiming has begun or the owner leaked for good. Every
field of the returned struct vmbus_buffer_pin is read under
vmbus_buffer_owners_lock, the lock that publishes and clears the
owner and the page array. The pin carries the page array, its length
and the mapping-protection state of the pages being pinned, so a
mapper never has to look at the buffer descriptor again: a concurrent
vmbus_release_buffer() can clear that descriptor while the pin still
holds the pages alive.
The pin cannot be dropped when mmap_prepare() returns.
mmap_action_map_kernel_pages() only stores the page-array pointer, and
the mapping takes its own folio references inside insert_pages(),
which runs after the hook. Dropping the pin therefore waits for the
VMA to be established or for the attempt to be abandoned. Release is
one-shot: both can run.
Two surfaces need two ownership contracts, because fs/kernfs/file.c
rejects a VMA whose vm_ops carry .close: kernfs has to replace the set
with its own and cannot wrap close. That check runs after the .mmap
hook has already succeeded, so a successful sysfs mapping cannot carry
an abandon path in vm_ops at all.
The /dev/uioN character device is not a kernfs path. It installs
hv_uio_pin_vm_ops (.mapped and .close) and a struct hv_uio_pin, and
drops the pin on whichever of the two runs first. The "ring" sysfs bin
attribute installs no vm_ops and hands a bare struct vmbus_buffer_pin
to hv_mmap_ring_buffer_wrapper(), a synchronous legacy .mmap that owns
it across __compat_vma_mmap() and releases it exactly once when that
returns: success is safe because insert_pages() has already taken a
folio reference per page, failure is safe because the attempt is over.
The legacy .mmap error path frees the VMA without calling vma_close(),
and a failure in mmap_action_prepare() never reaches
compat_set_vma_from_desc(). The wrapper therefore installs the
per-VMA state before __compat_vma_mmap() when a contract does carry
close(), so that abandon path stays reachable.
The ring-buffer sysfs mmap and the UIO mmap_prepare hook share one
page-selection and validation path. Buffer-backed regions are
described only by their owning buffer at selection time; the page
array and the protection state come from the locked pin snapshot and
from nowhere else. The interrupt and monitor pages are snapshot as
struct page * at probe, where the backing is guaranteed to still
exist, and are not buffer-backed so they take no pin.
The mapping KUnit cases cover the range check, the ring mmap path, the
per-map region selection, the mmap_prepare hook, the pin spanning the
mapped() boundary, close() before mapped() including a second close
and a foreign private_data, the absence of .close on the sysfs ops, a
successful buffer-backed character-device map, pin release on prepare
failure, pin release when the mapping is abandoned after prepare,
pages kept alive by the pin across a concurrent buffer release, and
the protection snapshot surviving a cleared descriptor. They are
included from uio_hv_generic.c so they can reach the static helpers
without exporting them.
Disconnect an open channel after unregistering UIO callbacks. An
eventual fd close cannot call .release after unregister, so removal must
close the channel before releasing its ring and external buffers. Mapped
pages retain their folio references until VMA close; disconnect failure
leaves host ownership uncertain and retained. Test open, already closed,
repeated and failed disconnection through the same helper used by
remove.
Hold the UIO notifier device reference across unregister and channel
disconnect. If the last fd closes in that interval, an in-flight channel
callback can still notify the UIO device until disconnect synchronizes
callbacks.
Use ordinary VMBus rescind device unregistration instead of a private
UIO rescind callback. The callback previously outlived a UIO-to-netvsc
rebind and interpreted netvsc private data as UIO, panicking in
uio_event_notify when the host removed the restored adapter. A callback
reset or null check cannot synchronize that race. Generic unregister
already withdraws UIO info and wakes blocked readers with EIO/POLL_HUP.
Add an actual registered UIO withdrawal test verifying reader wakeup and
info removal while a device reference keeps the notifier alive.
Add actual synchronous-wrapper/MM success and partial-insertion cases
with KUnit-managed memory. Read the inserted bytes, check VMA and bridge
references, then unmap. A wrapper-unpin mutation fails both cases. The
controlled prepare callback uses production pin acquisition; full
retained-owner and confidential integration remain separate.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 92 +++
drivers/hv/vmbus_drv.c | 49 +-
drivers/hv/vmbus_mmap_test.c | 145 ++++
drivers/uio/Kconfig | 11 +
drivers/uio/uio.c | 20 +
drivers/uio/uio_hv_generic.c | 321 ++++++++-
drivers/uio/uio_hv_generic_mmap_test.c | 875 +++++++++++++++++++++++++
include/linux/hyperv.h | 24 +
8 files changed, 1512 insertions(+), 25 deletions(-)
create mode 100644 drivers/hv/vmbus_mmap_test.c
create mode 100644 drivers/uio/uio_hv_generic_mmap_test.c
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index 7eb9ea814ef4..2793ca1b7f32 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -1145,6 +1145,98 @@ void vmbus_release_buffer(struct vmbus_buffer *buffer)
}
EXPORT_SYMBOL_GPL(vmbus_release_buffer);
+/**
+ * vmbus_buffer_pin_pages - snapshot a buffer's pages and hold references
+ * @buffer: buffer whose pages the caller is about to map
+ * @pin: out parameter filled entirely under the owners lock
+ *
+ * Take a reference on every page of @buffer and fill @pin with the page
+ * array, the page count and the mapping-protection state that describe
+ * them. The references close the window between the snapshot and the
+ * point where a new mapping holds its own: the reclaim worker refuses to
+ * free pages whose folio reference count is above one, and refuses to
+ * start at all once it has begun reclaiming. Call
+ * vmbus_buffer_unpin_pages() after the mapping holds its own references.
+ *
+ * Every field of @pin is read under vmbus_buffer_owners_lock so the
+ * caller never has to look at @buffer again. In particular @pin->decrypted
+ * is the protection state of the pages being pinned, not of whatever
+ * @buffer looks like after a concurrent vmbus_release_buffer() has
+ * cleared it.
+ *
+ * Return: 0 on success, -ENODEV if the pages are gone or reclaim
+ * has already taken ownership of them.
+ */
+int vmbus_buffer_pin_pages(struct vmbus_buffer *buffer,
+ struct vmbus_buffer_pin *pin)
+{
+ struct vmbus_buffer_retained *owner;
+ u32 i;
+
+ /*
+ * The owner and page array are published and cleared under
+ * vmbus_buffer_owners_lock (see vmbus_release_buffer()). Reading
+ * either outside that lock lets a concurrent release free the array
+ * while this loop still walks it.
+ */
+ mutex_lock(&vmbus_buffer_owners_lock);
+ owner = buffer->owner;
+ if (!buffer->pages || !buffer->page_cnt) {
+ mutex_unlock(&vmbus_buffer_owners_lock);
+ return -ENODEV;
+ }
+
+ if (owner && (owner->reclaiming || owner->permanent_leak)) {
+ mutex_unlock(&vmbus_buffer_owners_lock);
+ return -ENODEV;
+ }
+
+ for (i = 0; i < buffer->page_cnt; i++) {
+ if (WARN_ON_ONCE(!buffer->pages[i])) {
+ while (i--)
+ put_page(buffer->pages[i]);
+ mutex_unlock(&vmbus_buffer_owners_lock);
+ return -ENODEV;
+ }
+ get_page(buffer->pages[i]);
+ }
+
+ pin->pages = buffer->pages;
+ pin->page_count = buffer->page_cnt;
+ /*
+ * Only the shared-page path populates the chunk array, and that is
+ * the path that decrypts each chunk before mapping it. Read it here
+ * so the protection state cannot disagree with the pages pinned.
+ */
+ pin->decrypted = !!buffer->chunks;
+ mutex_unlock(&vmbus_buffer_owners_lock);
+
+ return 0;
+}
+EXPORT_SYMBOL_GPL(vmbus_buffer_pin_pages);
+
+/**
+ * vmbus_buffer_unpin_pages - drop references taken by vmbus_buffer_pin_pages
+ * @pin: pin filled by vmbus_buffer_pin_pages(), or NULL
+ *
+ * Idempotent: a pin that has already been released, or a NULL pin, is a
+ * no-op. Release is one-shot because a VMA can be established and then
+ * torn down, and the abandon path must not put the same pages twice.
+ */
+void vmbus_buffer_unpin_pages(struct vmbus_buffer_pin *pin)
+{
+ unsigned long i;
+
+ if (!pin || !pin->pages)
+ return;
+
+ for (i = 0; i < pin->page_count; i++)
+ put_page(pin->pages[i]);
+ pin->pages = NULL;
+ pin->page_count = 0;
+}
+EXPORT_SYMBOL_GPL(vmbus_buffer_unpin_pages);
+
static struct page *vmbus_alloc_pages_node(void *context, int nid,
gfp_t gfp,
unsigned int order)
diff --git a/drivers/hv/vmbus_drv.c b/drivers/hv/vmbus_drv.c
index bce835c4a015..198b43d24051 100644
--- a/drivers/hv/vmbus_drv.c
+++ b/drivers/hv/vmbus_drv.c
@@ -1929,6 +1929,7 @@ static int hv_mmap_ring_buffer_wrapper(struct file *filp, struct kobject *kobj,
{
struct vmbus_channel *channel = container_of(kobj, struct vmbus_channel, kobj);
struct vm_area_desc desc;
+ struct vmbus_buffer_pin *pin;
int err;
/*
@@ -1940,9 +1941,55 @@ static int hv_mmap_ring_buffer_wrapper(struct file *filp, struct kobject *kobj,
if (err)
return err;
- return __compat_vma_mmap(&desc, vma);
+ /*
+ * This is a kernfs bin attribute. kernfs_fop_mmap() rejects a VMA
+ * whose vm_ops carry .close, because kernfs has to replace the set
+ * with its own and cannot wrap close. So the prepare callback has
+ * to pick one of two ownership contracts, and this wrapper honors
+ * whichever it installed. Each contract has exactly one release
+ * site, reached on every outcome.
+ *
+ * With vm_ops, the callback handed the per-VMA state to an abandon
+ * path in vm_ops->close and typically to vm_ops->mapped. Install
+ * both before __compat_vma_mmap() so a failure in
+ * mmap_action_prepare() still reaches close(): that path never runs
+ * compat_set_vma_from_desc(), and the legacy .mmap error path frees
+ * the VMA without calling close() on its own. __compat_vma_mmap()
+ * installs the same values again on its way to the action. Note
+ * that such a set is rejected by kernfs after this returns, so a
+ * prepare callback that wants a successful mapping must not choose
+ * this contract.
+ *
+ * Without vm_ops, desc.private_data is a plain struct
+ * vmbus_buffer_pin and this wrapper is its only owner. Release it
+ * when __compat_vma_mmap() returns: on success insert_pages() has
+ * already taken a folio reference for every page the VMA maps, and
+ * on failure the attempt is over. Clear vma->vm_private_data first
+ * so the freed pin is never reachable from the VMA that outlives
+ * this call.
+ */
+ if (desc.vm_ops) {
+ vma->vm_ops = desc.vm_ops;
+ vma->vm_private_data = desc.private_data;
+
+ err = __compat_vma_mmap(&desc, vma);
+ if (err && vma->vm_ops && vma->vm_ops->close)
+ vma->vm_ops->close(vma);
+ return err;
+ }
+
+ pin = desc.private_data;
+ err = __compat_vma_mmap(&desc, vma);
+ vma->vm_private_data = NULL;
+ vmbus_buffer_unpin_pages(pin);
+ kfree(pin);
+ return err;
}
+#if IS_ENABLED(CONFIG_HYPERV_VMBUS_KUNIT_TEST)
+#include "vmbus_mmap_test.c"
+#endif
+
static struct bin_attribute chan_attr_ring_buffer = {
.attr = {
.name = "ring",
diff --git a/drivers/hv/vmbus_mmap_test.c b/drivers/hv/vmbus_mmap_test.c
new file mode 100644
index 000000000000..50b961f2b782
--- /dev/null
+++ b/drivers/hv/vmbus_mmap_test.c
@@ -0,0 +1,145 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Included beside the production sysfs wrapper; no VMBus device is touched. */
+#include <kunit/test.h>
+#include <linux/mman.h>
+#include <linux/uaccess.h>
+
+struct vmbus_mmap_test_context {
+ struct vmbus_channel channel;
+ struct page *pages[3];
+ struct page *blocker;
+};
+
+static void vmbus_mmap_test_put_page(void *page)
+{
+ __free_page(page);
+}
+
+static struct page *vmbus_mmap_test_page(struct kunit *test, u8 value)
+{
+ struct page *page = alloc_page(GFP_KERNEL);
+
+ if (!page)
+ return NULL;
+ if (kunit_add_action_or_reset(test, vmbus_mmap_test_put_page, page))
+ return NULL;
+ memset(page_address(page), value, PAGE_SIZE);
+ return page;
+}
+
+static int vmbus_mmap_test_prepare(struct vmbus_channel *channel,
+ struct vm_area_desc *desc)
+{
+ struct vmbus_buffer_pin *pin;
+ int ret;
+
+ pin = kzalloc_obj(*pin);
+ if (!pin)
+ return -ENOMEM;
+ ret = vmbus_buffer_pin_pages(&channel->ringbuffer, pin);
+ if (ret) {
+ kfree(pin);
+ return ret;
+ }
+ mmap_action_map_kernel_pages(desc, desc->start, pin->pages,
+ pin->page_count);
+ desc->vm_ops = NULL;
+ desc->private_data = pin;
+ return 0;
+}
+
+static struct vmbus_mmap_test_context *
+vmbus_mmap_test_context(struct kunit *test)
+{
+ struct vmbus_mmap_test_context *ctx;
+ unsigned int i;
+
+ ctx = kunit_kzalloc(test, sizeof(*ctx), GFP_KERNEL);
+ if (!ctx)
+ return NULL;
+ for (i = 0; i < ARRAY_SIZE(ctx->pages); i++) {
+ ctx->pages[i] = vmbus_mmap_test_page(test, 0x61 + i);
+ if (!ctx->pages[i])
+ return NULL;
+ }
+ ctx->blocker = vmbus_mmap_test_page(test, 0xcc);
+ if (!ctx->blocker)
+ return NULL;
+ ctx->channel.mmap_prepare_ring_buffer = vmbus_mmap_test_prepare;
+ ctx->channel.ringbuffer.pages = ctx->pages;
+ ctx->channel.ringbuffer.page_cnt = ARRAY_SIZE(ctx->pages);
+ return ctx;
+}
+
+static void vmbus_mmap_test_action(struct kunit *test, bool partial)
+{
+ struct vmbus_mmap_test_context *ctx = vmbus_mmap_test_context(test);
+ struct vm_area_struct *vma;
+ unsigned long start;
+ unsigned int i;
+ u8 value = 0;
+ int ret;
+
+ KUNIT_ASSERT_NOT_NULL(test, ctx);
+ start = kunit_vm_mmap(test, NULL, 0, 3 * PAGE_SIZE, PROT_READ | PROT_WRITE,
+ MAP_SHARED | MAP_ANONYMOUS, 0);
+ KUNIT_ASSERT_NE(test, start, 0UL);
+ KUNIT_ASSERT_LT(test, start, (unsigned long)TASK_SIZE);
+
+ mmap_write_lock(current->mm);
+ vma = find_vma(current->mm, start);
+ if (!vma) {
+ mmap_write_unlock(current->mm);
+ KUNIT_FAIL(test, "managed VMA missing");
+ return;
+ }
+ if (partial) {
+ ret = vm_insert_page(vma, start + PAGE_SIZE, ctx->blocker);
+ if (ret) {
+ mmap_write_unlock(current->mm);
+ KUNIT_FAIL(test, "owned blocking PTE could not be installed");
+ return;
+ }
+ }
+ ret = hv_mmap_ring_buffer_wrapper(vma->vm_file, &ctx->channel.kobj,
+ NULL, vma);
+ KUNIT_EXPECT_EQ(test, ret, partial ? -EBUSY : 0);
+ KUNIT_EXPECT_NULL(test, vma->vm_private_data);
+ KUNIT_EXPECT_NULL(test, vma->vm_ops);
+ for (i = 0; i < ARRAY_SIZE(ctx->pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(ctx->pages[i]),
+ partial && i ? 1 : 2);
+ KUNIT_EXPECT_EQ(test, page_ref_count(ctx->blocker), partial ? 2 : 1);
+ mmap_write_unlock(current->mm);
+
+ /* The first insertion really happened, including on the partial error. */
+ KUNIT_EXPECT_EQ(test, copy_from_user(&value, (void __user *)start, 1), 0UL);
+ KUNIT_EXPECT_EQ(test, value, (u8)0x61);
+ KUNIT_ASSERT_EQ(test, vm_munmap(start, 3 * PAGE_SIZE), 0);
+ for (i = 0; i < ARRAY_SIZE(ctx->pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(ctx->pages[i]), 1);
+ KUNIT_EXPECT_EQ(test, page_ref_count(ctx->blocker), 1);
+}
+
+static void vmbus_sysfs_mmap_success_test(struct kunit *test)
+{
+ vmbus_mmap_test_action(test, false);
+}
+
+static void vmbus_sysfs_mmap_partial_insert_test(struct kunit *test)
+{
+ vmbus_mmap_test_action(test, true);
+}
+
+static struct kunit_case vmbus_mmap_test_cases[] = {
+ KUNIT_CASE(vmbus_sysfs_mmap_success_test),
+ KUNIT_CASE(vmbus_sysfs_mmap_partial_insert_test),
+ {}
+};
+
+static struct kunit_suite vmbus_mmap_test_suite = {
+ .name = "hyperv-vmbus-mmap",
+ .test_cases = vmbus_mmap_test_cases,
+};
+
+kunit_test_suite(vmbus_mmap_test_suite);
diff --git a/drivers/uio/Kconfig b/drivers/uio/Kconfig
index 9242e77385c6..0a01906bd6cd 100644
--- a/drivers/uio/Kconfig
+++ b/drivers/uio/Kconfig
@@ -148,6 +148,17 @@ config UIO_HV_GENERIC
If you compile this as a module, it will be called uio_hv_generic.
+config UIO_HV_GENERIC_KUNIT_TEST
+ bool "Tests for the UIO Hyper-V generic mmap" if !KUNIT_ALL_TESTS
+ depends on UIO_HV_GENERIC && KUNIT=y
+ default KUNIT_ALL_TESTS
+ help
+ Enable KUnit tests for the UIO Hyper-V generic mmap paths.
+ The tests are included from uio_hv_generic.c so they can
+ reach the static page-selection helpers. Select this option
+ only if you will boot the kernel for the purpose of running
+ unit tests (e.g. under UML or qemu). If unsure, say N.
+
config UIO_DFL
tristate "Generic driver for DFL (Device Feature List) bus"
depends on FPGA_DFL
diff --git a/drivers/uio/uio.c b/drivers/uio/uio.c
index f8fa20522660..4767ed0661f7 100644
--- a/drivers/uio/uio.c
+++ b/drivers/uio/uio.c
@@ -857,7 +857,27 @@ static int uio_mmap(struct file *filep, struct vm_area_struct *vma)
ret = idev->info->mmap_prepare(idev->info, &desc);
if (ret)
goto out;
+
+ /*
+ * Install the per-VMA state before __compat_vma_mmap() so a
+ * failure in mmap_action_prepare() still reaches close(). That
+ * path never runs compat_set_vma_from_desc(), and the legacy
+ * .mmap error path frees the VMA without calling close() on
+ * its own. __compat_vma_mmap() installs the same values again
+ * on its way to the action.
+ *
+ * close() must be safe to call without a prior mapped() when
+ * it tears down state mmap_prepare() installed. Drivers that
+ * leave vm_ops alone have nothing to tear down here.
+ */
+ if (desc.vm_ops) {
+ vma->vm_ops = desc.vm_ops;
+ vma->vm_private_data = desc.private_data;
+ }
+
ret = __compat_vma_mmap(&desc, vma);
+ if (ret && vma->vm_ops && vma->vm_ops->close)
+ vma->vm_ops->close(vma);
goto out;
}
diff --git a/drivers/uio/uio_hv_generic.c b/drivers/uio/uio_hv_generic.c
index b40e80e19c6c..91cb25d019be 100644
--- a/drivers/uio/uio_hv_generic.c
+++ b/drivers/uio/uio_hv_generic.c
@@ -25,6 +25,7 @@
#include <linux/uio_driver.h>
#include <linux/netdevice.h>
#include <linux/if_ether.h>
+#include <linux/mm.h>
#include <linux/skbuff.h>
#include <linux/hyperv.h>
#include <linux/vmalloc.h>
@@ -55,6 +56,8 @@ struct hv_uio_private_data {
struct uio_info info;
struct hv_device *device;
atomic_t refcnt;
+ struct page *int_pages[1];
+ struct page *monitor_pages[1];
struct vmbus_buffer recv_buffer;
char recv_name[32]; /* "recv_4294967295" */
@@ -121,51 +124,310 @@ static void hv_uio_channel_cb(void *context)
uio_event_notify(&pdata->info);
}
+/* Function used for mmap of the ring buffer sysfs interface. */
+static bool hv_uio_mmap_range_valid(unsigned long map_pages,
+ pgoff_t offset,
+ unsigned long pages)
+{
+ return pages && offset < map_pages && pages <= map_pages - offset;
+}
+
+struct hv_uio_mmap_region {
+ struct page **pages;
+ unsigned long page_count;
+ struct vmbus_buffer *buffer;
+ bool decrypted;
+};
+
+static int hv_uio_mmap_get_region(struct hv_uio_private_data *pdata,
+ unsigned int map_index,
+ struct hv_uio_mmap_region *region)
+{
+ struct hv_device *dev = pdata->device;
+ struct vmbus_channel *channel = dev->channel;
+
+ if (map_index >= MAX_UIO_MAPS || !pdata->info.mem[map_index].size)
+ return -EINVAL;
+
+ /*
+ * Buffer-backed regions are not described here. Their page array
+ * and protection state are only meaningful as the locked snapshot
+ * from vmbus_buffer_pin_pages(): reading the descriptor outside
+ * that lock can describe a buffer a concurrent release has already
+ * cleared. Only the owning buffer is selected at this stage.
+ */
+ region->buffer = NULL;
+ region->pages = NULL;
+ region->page_count = 0;
+ region->decrypted = false;
+
+ switch (map_index) {
+ case TXRX_RING_MAP:
+ if (channel->state != CHANNEL_OPENED_STATE)
+ return -ENODEV;
+ region->buffer = &channel->ringbuffer;
+ return 0;
+ case INT_PAGE_MAP:
+ region->pages = pdata->int_pages;
+ region->page_count = ARRAY_SIZE(pdata->int_pages);
+ region->decrypted = false;
+ break;
+ case MON_PAGE_MAP:
+ region->pages = pdata->monitor_pages;
+ region->page_count = ARRAY_SIZE(pdata->monitor_pages);
+ region->decrypted = true;
+ break;
+ case RECV_BUF_MAP:
+ region->buffer = &pdata->recv_buffer;
+ return 0;
+ case SEND_BUF_MAP:
+ region->buffer = &pdata->send_buffer;
+ return 0;
+ default:
+ return -EINVAL;
+ }
+
+ if (!region->pages || !region->page_count)
+ return -EINVAL;
+
+ return 0;
+}
+
+static int hv_uio_mmap_prepare_pages(struct vm_area_desc *desc,
+ struct page **pages,
+ unsigned long page_count,
+ pgoff_t offset,
+ bool decrypted)
+{
+ unsigned long nr_pages = vma_desc_pages(desc);
+
+ if (!vma_desc_test(desc, VMA_SHARED_BIT))
+ return -EINVAL;
+
+ if (!pages || !hv_uio_mmap_range_valid(page_count, offset, nr_pages))
+ return -EINVAL;
+
+ vma_desc_set_flags(desc, VMA_DONTEXPAND_BIT, VMA_DONTDUMP_BIT);
+ if (decrypted)
+ desc->page_prot = pgprot_decrypted(desc->page_prot);
+
+ mmap_action_map_kernel_pages(desc, desc->start, pages + offset,
+ nr_pages);
+ return 0;
+}
+
/*
- * Callback from vmbus_event when channel is rescinded.
- * It is meant for rescind of primary channels only.
+ * mmap_action_map_kernel_pages() only stores the page-array pointer. The
+ * mapping takes its own folio references inside insert_pages(), which runs
+ * after the prepare callback has returned. Dropping the pin therefore has
+ * to wait for the VMA to be established (mapped) or for the attempt to be
+ * abandoned. Release is one-shot: the establish and abandon paths can both
+ * be reached for the same attempt.
+ *
+ * Two surfaces, two ownership contracts. /dev/uioN is a character device,
+ * so kernfs is not in the path and vm_ops->close is a legal abandon path.
+ * The sysfs "ring" bin attribute is kernfs: fs/kernfs/file.c rejects any
+ * VMA whose vm_ops carry .close, because kernfs has to wrap the operations
+ * and cannot wrap close. That surface installs no vm_ops at all and its
+ * synchronous legacy .mmap wrapper owns the pin for the whole window.
*/
-static void hv_uio_rescind(struct vmbus_channel *channel)
+#define HV_UIO_PIN_MAGIC 0x48565049UL /* "HVPI" */
+
+struct hv_uio_pin {
+ unsigned long magic;
+ struct vmbus_buffer_pin snap;
+};
+
+static void hv_uio_pin_release(void *data)
{
- struct hv_device *hv_dev = channel->device_obj;
- struct hv_uio_private_data *pdata = hv_get_drvdata(hv_dev);
+ struct hv_uio_pin *pin = data;
+ if (!pin)
+ return;
/*
- * Turn off the interrupt file handle
- * Next read for event will return -EIO
+ * close() can be reached with a private_data that this driver did
+ * not install. A foreign pointer must not be freed or unpinned.
*/
- pdata->info.irq = 0;
+ if (pin->magic != HV_UIO_PIN_MAGIC)
+ return;
+ pin->magic = 0;
+ vmbus_buffer_unpin_pages(&pin->snap);
+ kfree(pin);
+}
- /* Wake up reader */
- uio_event_notify(&pdata->info);
+static int hv_uio_pin_vma_mapped(unsigned long start, unsigned long end,
+ pgoff_t pgoff, const struct file *file,
+ void **vm_private_data)
+{
+ /*
+ * insert_pages() has taken a folio reference for every page the
+ * VMA now maps, so the reclaim worker can no longer free them
+ * under the mapping. The bridge pin is done.
+ */
+ hv_uio_pin_release(*vm_private_data);
+ *vm_private_data = NULL;
+ return 0;
+}
+static void hv_uio_pin_vma_close(struct vm_area_struct *vma)
+{
/*
- * With rescind callback registered, rescind path will not unregister the device
- * from vmbus when the primary channel is rescinded.
- * Without it, rescind handling is incomplete and next onoffer msg does not come.
- * Unregister the device from vmbus here.
+ * Reached with the pin still held when the VMA is torn down before
+ * mapped() ran (an insert_pages() failure, or a merge that never
+ * calls mapped), and with nothing left to do when it already ran.
*/
- vmbus_device_unregister(channel->device_obj);
+ hv_uio_pin_release(vma->vm_private_data);
+ vma->vm_private_data = NULL;
+}
+
+/*
+ * Character-device surface only. The sysfs ring must never adopt this set:
+ * see the kernfs contract above.
+ */
+static const struct vm_operations_struct hv_uio_pin_vm_ops = {
+ .mapped = hv_uio_pin_vma_mapped,
+ .close = hv_uio_pin_vma_close,
+};
+
+/*
+ * Install @snap as the VMA's per-mapping state under @ops. Consumes @snap:
+ * returns 0 with the pin owned by @desc and @snap cleared, or -ENOMEM with
+ * the references dropped here. The caller must not drop the references on
+ * the success path: mapped() and close() own them from here.
+ */
+static int hv_uio_pin_install(struct vm_area_desc *desc,
+ const struct vm_operations_struct *ops,
+ struct vmbus_buffer_pin *snap)
+{
+ struct hv_uio_pin *pin;
+
+ pin = kzalloc_obj(*pin);
+ if (!pin) {
+ vmbus_buffer_unpin_pages(snap);
+ return -ENOMEM;
+ }
+ pin->magic = HV_UIO_PIN_MAGIC;
+ pin->snap = *snap;
+ snap->pages = NULL;
+ snap->page_count = 0;
+
+ desc->vm_ops = ops;
+ desc->private_data = pin;
+ return 0;
+}
+
+static int hv_uio_mmap_prepare(struct uio_info *info,
+ struct vm_area_desc *desc)
+{
+ struct hv_uio_private_data *pdata = info->priv;
+ struct hv_uio_mmap_region region;
+ struct vmbus_buffer_pin pin = {};
+ struct page **pages;
+ unsigned long page_count;
+ bool decrypted;
+ int ret;
+
+ /* UIO encodes the map index in pgoff; mmap offsets start at the map. */
+ if (desc->pgoff >= MAX_UIO_MAPS)
+ return -EINVAL;
+
+ ret = hv_uio_mmap_get_region(pdata, desc->pgoff, ®ion);
+ if (ret)
+ return ret;
+
+ /*
+ * Buffer-backed maps race with the reclaim worker: the page
+ * array and the pages themselves can be freed between the
+ * selection above and the mapping below. Hold references until
+ * the VMA holds its own. The pin is also the only safe source of
+ * the protection state: it is taken in the same critical section
+ * that reads the page array.
+ */
+ if (region.buffer) {
+ ret = vmbus_buffer_pin_pages(region.buffer, &pin);
+ if (ret)
+ return ret;
+ pages = pin.pages;
+ page_count = pin.page_count;
+ decrypted = pin.decrypted;
+ } else {
+ pages = region.pages;
+ page_count = region.page_count;
+ decrypted = region.decrypted;
+ }
+
+ ret = hv_uio_mmap_prepare_pages(desc, pages, page_count, 0, decrypted);
+ if (ret) {
+ if (region.buffer)
+ vmbus_buffer_unpin_pages(&pin);
+ return ret;
+ }
+
+ if (!region.buffer)
+ return 0;
+
+ return hv_uio_pin_install(desc, &hv_uio_pin_vm_ops, &pin);
}
-/* Function used for mmap of the ring buffer sysfs interface. */
static int
-hv_uio_ring_mmap_prepare(struct vmbus_channel *channel, struct vm_area_desc *desc)
+hv_uio_ring_mmap_prepare(struct vmbus_channel *channel,
+ struct vm_area_desc *desc)
{
- unsigned long pages = vma_desc_pages(desc);
+ struct vmbus_buffer *buffer = &channel->ringbuffer;
+ struct vmbus_buffer_pin *pin;
pgoff_t offset = desc->pgoff;
+ int ret;
if (channel->state != CHANNEL_OPENED_STATE)
return -ENODEV;
- if (offset >= channel->ringbuffer_pagecount ||
- pages > channel->ringbuffer_pagecount - offset)
- return -EINVAL;
- mmap_action_map_kernel_pages(desc, desc->start,
- channel->ringbuffer.pages + offset, pages);
+ pin = kzalloc_obj(*pin);
+ if (!pin)
+ return -ENOMEM;
+
+ ret = vmbus_buffer_pin_pages(buffer, pin);
+ if (ret) {
+ kfree(pin);
+ return ret;
+ }
+
+ ret = hv_uio_mmap_prepare_pages(desc, pin->pages, pin->page_count,
+ offset, pin->decrypted);
+ if (ret) {
+ vmbus_buffer_unpin_pages(pin);
+ kfree(pin);
+ return ret;
+ }
+
+ /*
+ * Install no vm_ops. kernfs_fop_mmap() returns -EINVAL for a VMA
+ * whose vm_ops carry .close, and a successful mapping still has to
+ * survive that check, so this surface cannot carry an abandon path
+ * in vm_ops at all. The caller is hv_mmap_ring_buffer_wrapper(),
+ * a synchronous legacy .mmap: it holds @pin across
+ * __compat_vma_mmap() and frees it when that returns, exactly once,
+ * on every outcome. desc->private_data is a plain
+ * struct vmbus_buffer_pin for that wrapper to release.
+ */
+ desc->vm_ops = NULL;
+ desc->private_data = pin;
return 0;
}
+/* Caller has withdrawn UIO callbacks before removing its backing memory. */
+static int hv_uio_disconnect_if_open(struct vmbus_channel *channel,
+ int (*disconnect)(struct vmbus_channel *))
+{
+ if (channel->state != CHANNEL_OPENED_STATE)
+ return 0;
+
+ return disconnect(channel);
+}
+
+#if defined(CONFIG_UIO_HV_GENERIC_KUNIT_TEST)
+#include "uio_hv_generic_mmap_test.c"
+#endif
+
/* Callback from VMBUS subsystem when new channel created. */
static void
hv_uio_new_channel(struct vmbus_channel *new_sc)
@@ -216,7 +478,6 @@ hv_uio_open(struct uio_info *info, struct inode *inode)
if (atomic_inc_return(&pdata->refcnt) != 1)
return 0;
- vmbus_set_chn_rescind_callback(dev->channel, hv_uio_rescind);
vmbus_set_sc_create_callback(dev->channel, hv_uio_new_channel);
ret = vmbus_connect_ring(dev->channel,
@@ -272,6 +533,7 @@ hv_uio_probe(struct hv_device *dev,
pdata->info.name = "uio_hv_generic";
pdata->info.version = DRIVER_VERSION;
pdata->info.irqcontrol = hv_uio_irqcontrol;
+ pdata->info.mmap_prepare = hv_uio_mmap_prepare;
pdata->info.open = hv_uio_open;
pdata->info.release = hv_uio_release;
pdata->info.irq = UIO_IRQ_CUSTOM;
@@ -290,12 +552,15 @@ hv_uio_probe(struct hv_device *dev,
= (uintptr_t)vmbus_connection.int_page;
pdata->info.mem[INT_PAGE_MAP].size = HV_HYP_PAGE_SIZE;
pdata->info.mem[INT_PAGE_MAP].memtype = UIO_MEM_LOGICAL;
+ pdata->int_pages[0] = virt_to_page(vmbus_connection.int_page);
pdata->info.mem[MON_PAGE_MAP].name = "monitor_page";
pdata->info.mem[MON_PAGE_MAP].addr
= (uintptr_t)vmbus_connection.monitor_pages[1];
pdata->info.mem[MON_PAGE_MAP].size = HV_HYP_PAGE_SIZE;
pdata->info.mem[MON_PAGE_MAP].memtype = UIO_MEM_LOGICAL;
+ pdata->monitor_pages[0] =
+ virt_to_page(vmbus_connection.monitor_pages[1]);
if (channel->device_id == HV_NIC) {
ret = vmbus_alloc_buffer_owned(channel, RECV_BUFFER_SIZE, false,
@@ -372,12 +637,20 @@ static void
hv_uio_remove(struct hv_device *dev)
{
struct hv_uio_private_data *pdata = hv_get_drvdata(dev);
+ int ret;
if (!pdata)
return;
hv_remove_ring_sysfs(dev->channel);
+ /* Keep event notification alive until channel callbacks are stopped. */
+ get_device(&pdata->info.uio_dev->dev);
uio_unregister_device(&pdata->info);
+ /* unregister prevents the eventual fd close from calling .release. */
+ ret = hv_uio_disconnect_if_open(dev->channel, vmbus_disconnect_ring);
+ if (ret)
+ dev_err(&dev->device, "channel disconnect failed: %d\n", ret);
+ put_device(&pdata->info.uio_dev->dev);
hv_uio_cleanup(dev, pdata);
vmbus_free_ring(dev->channel);
diff --git a/drivers/uio/uio_hv_generic_mmap_test.c b/drivers/uio/uio_hv_generic_mmap_test.c
new file mode 100644
index 000000000000..615df78b266e
--- /dev/null
+++ b/drivers/uio/uio_hv_generic_mmap_test.c
@@ -0,0 +1,875 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * KUnit tests for the UIO Hyper-V generic mmap paths.
+ *
+ * This file is included from uio_hv_generic.c so the cases can reach the
+ * static helpers that select, validate and pin the mapped pages.
+ *
+ * Two ownership contracts are under test:
+ * - the /dev/uioN character-device path installs hv_uio_pin_vm_ops
+ * (.mapped + .close) and a struct hv_uio_pin;
+ * - the sysfs "ring" bin attribute cannot, because kernfs_fop_mmap()
+ * rejects a vm_operations_struct that carries .close. That path
+ * installs no vm_ops and hands a bare struct vmbus_buffer_pin to its
+ * synchronous .mmap wrapper, which is the single release site.
+ */
+#include <kunit/test.h>
+
+/*
+ * Release exactly what hv_mmap_ring_buffer_wrapper() releases after
+ * __compat_vma_mmap() returns: the bare pin's references, then the pin.
+ */
+static void sysfs_pin_release(struct vmbus_buffer_pin *pin)
+{
+ vmbus_buffer_unpin_pages(pin);
+ kfree(pin);
+}
+
+static void hv_uio_ring_mmap_range_test(struct kunit *test)
+{
+ KUNIT_EXPECT_TRUE(test, hv_uio_mmap_range_valid(8, 0, 8));
+ KUNIT_EXPECT_TRUE(test, hv_uio_mmap_range_valid(8, 4, 4));
+ KUNIT_EXPECT_FALSE(test, hv_uio_mmap_range_valid(8, 8, 1));
+ KUNIT_EXPECT_FALSE(test, hv_uio_mmap_range_valid(8, 7, 2));
+ KUNIT_EXPECT_FALSE(test, hv_uio_mmap_range_valid(8, 0, 0));
+ KUNIT_EXPECT_FALSE(test,
+ hv_uio_mmap_range_valid(8, U64_MAX, 1));
+}
+
+static void hv_uio_ring_mmap_prepare_test(struct kunit *test)
+{
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ .ringbuffer = {
+ .page_cnt = 8,
+ },
+ };
+ struct page *pages[8] = {};
+ struct page *chunks[1] = {};
+ struct vmbus_buffer_pin *pin;
+ struct vm_area_desc desc = {
+ .start = PAGE_SIZE,
+ .end = 5 * PAGE_SIZE,
+ .pgoff = 2,
+ .page_prot = PAGE_SHARED,
+ };
+ unsigned int i;
+ int ret;
+
+ /* Mapping pins real pages; the array cannot hold NULL entries. */
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+
+ channel.ringbuffer.pages = pages;
+ channel.ringbuffer.chunks = chunks;
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ ret = hv_uio_ring_mmap_prepare(&channel, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_EQ(test, desc.action.type, MMAP_MAP_KERNEL_PAGES);
+ KUNIT_EXPECT_PTR_EQ(test, desc.action.map_kernel.pages, &pages[2]);
+ KUNIT_EXPECT_EQ(test, desc.action.map_kernel.nr_pages, 4UL);
+ KUNIT_EXPECT_TRUE(test, vma_desc_test(&desc, VMA_DONTEXPAND_BIT));
+ KUNIT_EXPECT_TRUE(test, vma_desc_test(&desc, VMA_DONTDUMP_BIT));
+ KUNIT_EXPECT_EQ(test, pgprot_val(desc.page_prot),
+ pgprot_val(pgprot_decrypted(PAGE_SHARED)));
+
+ /*
+ * The sysfs surface installs no vm_ops at all: see the kernfs
+ * contract in the comment at the top of this file. The wrapper
+ * holds the pin and is the only thing that releases it.
+ */
+ KUNIT_EXPECT_NULL(test, desc.vm_ops);
+ pin = desc.private_data;
+ KUNIT_ASSERT_NOT_NULL(test, pin);
+ KUNIT_EXPECT_TRUE(test, pin->decrypted);
+ KUNIT_EXPECT_EQ(test, pin->page_count, ARRAY_SIZE(pages));
+ sysfs_pin_release(pin);
+
+ desc.vma_flags = EMPTY_VMA_FLAGS;
+ KUNIT_EXPECT_EQ(test, hv_uio_ring_mmap_prepare(&channel, &desc), -EINVAL);
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+
+ desc.start = PAGE_SIZE;
+ desc.end = 3 * PAGE_SIZE;
+ desc.pgoff = 7;
+ KUNIT_EXPECT_EQ(test, hv_uio_ring_mmap_prepare(&channel, &desc), -EINVAL);
+
+ channel.state = CHANNEL_OPEN_STATE;
+ KUNIT_EXPECT_EQ(test, hv_uio_ring_mmap_prepare(&channel, &desc), -ENODEV);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+}
+
+static void hv_uio_mmap_region_select_test(struct kunit *test)
+{
+ struct hv_uio_private_data pdata = {};
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ };
+ struct hv_device dev = {
+ .channel = &channel,
+ };
+ struct hv_uio_mmap_region region;
+ struct page *ring_pages[4] = {};
+ struct page *recv_pages[3] = {};
+ struct page *send_pages[2] = {};
+ struct page *chunks[1] = {};
+ int ret;
+
+ pdata.device = &dev;
+ pdata.info.mem[TXRX_RING_MAP].size =
+ ARRAY_SIZE(ring_pages) * PAGE_SIZE;
+ pdata.info.mem[INT_PAGE_MAP].size = HV_HYP_PAGE_SIZE;
+ pdata.info.mem[MON_PAGE_MAP].size = HV_HYP_PAGE_SIZE;
+ pdata.info.mem[RECV_BUF_MAP].size =
+ ARRAY_SIZE(recv_pages) * PAGE_SIZE;
+ pdata.info.mem[SEND_BUF_MAP].size =
+ ARRAY_SIZE(send_pages) * PAGE_SIZE;
+ channel.ringbuffer.pages = ring_pages;
+ channel.ringbuffer.page_cnt = ARRAY_SIZE(ring_pages);
+ channel.ringbuffer.chunks = chunks;
+ pdata.int_pages[0] = ring_pages[0];
+ pdata.monitor_pages[0] = ring_pages[1];
+ pdata.recv_buffer.pages = recv_pages;
+ pdata.recv_buffer.page_cnt = ARRAY_SIZE(recv_pages);
+ pdata.recv_buffer.chunks = chunks;
+ pdata.send_buffer.pages = send_pages;
+ pdata.send_buffer.page_cnt = ARRAY_SIZE(send_pages);
+
+ /*
+ * Buffer-backed maps are described only by their owning buffer.
+ * The page array and the protection state are deliberately absent
+ * here: reading them outside vmbus_buffer_owners_lock can describe
+ * a buffer a concurrent release has already cleared. The locked
+ * pin snapshot is the only safe source.
+ */
+ ret = hv_uio_mmap_get_region(&pdata, TXRX_RING_MAP, ®ion);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_PTR_EQ(test, region.buffer, &channel.ringbuffer);
+ KUNIT_EXPECT_NULL(test, region.pages);
+ KUNIT_EXPECT_EQ(test, region.page_count, 0UL);
+ KUNIT_EXPECT_FALSE(test, region.decrypted);
+
+ ret = hv_uio_mmap_get_region(&pdata, INT_PAGE_MAP, ®ion);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_NULL(test, region.buffer);
+ KUNIT_EXPECT_PTR_EQ(test, region.pages, &pdata.int_pages[0]);
+ KUNIT_EXPECT_FALSE(test, region.decrypted);
+
+ ret = hv_uio_mmap_get_region(&pdata, MON_PAGE_MAP, ®ion);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_NULL(test, region.buffer);
+ KUNIT_EXPECT_PTR_EQ(test, region.pages, &pdata.monitor_pages[0]);
+ KUNIT_EXPECT_TRUE(test, region.decrypted);
+
+ ret = hv_uio_mmap_get_region(&pdata, RECV_BUF_MAP, ®ion);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_PTR_EQ(test, region.buffer, &pdata.recv_buffer);
+ KUNIT_EXPECT_NULL(test, region.pages);
+ KUNIT_EXPECT_EQ(test, region.page_count, 0UL);
+ KUNIT_EXPECT_FALSE(test, region.decrypted);
+
+ ret = hv_uio_mmap_get_region(&pdata, SEND_BUF_MAP, ®ion);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_PTR_EQ(test, region.buffer, &pdata.send_buffer);
+ KUNIT_EXPECT_NULL(test, region.pages);
+ KUNIT_EXPECT_EQ(test, region.page_count, 0UL);
+ KUNIT_EXPECT_FALSE(test, region.decrypted);
+
+ KUNIT_EXPECT_EQ(test,
+ hv_uio_mmap_get_region(&pdata, MAX_UIO_MAPS, ®ion),
+ -EINVAL);
+ channel.state = CHANNEL_OPEN_STATE;
+ KUNIT_EXPECT_EQ(test,
+ hv_uio_mmap_get_region(&pdata, TXRX_RING_MAP, ®ion),
+ -ENODEV);
+}
+
+static void hv_uio_mmap_prepare_test(struct kunit *test)
+{
+ struct hv_uio_private_data *pdata;
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ };
+ struct hv_device dev = {
+ .channel = &channel,
+ };
+ struct page *pages[4] = {};
+ struct page *chunks[1] = {};
+ struct vm_area_desc desc = {
+ .start = PAGE_SIZE,
+ .end = 3 * PAGE_SIZE,
+ .pgoff = RECV_BUF_MAP,
+ .page_prot = PAGE_SHARED,
+ };
+ unsigned int i;
+ int ret;
+
+ /* Mapping pins real pages; the array cannot hold NULL entries. */
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+
+ pdata = kunit_kzalloc(test, sizeof(*pdata), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pdata);
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ pdata->device = &dev;
+ pdata->info.priv = pdata;
+ pdata->info.mem[RECV_BUF_MAP].size = ARRAY_SIZE(pages) * PAGE_SIZE;
+ pdata->recv_buffer.pages = pages;
+ pdata->recv_buffer.page_cnt = ARRAY_SIZE(pages);
+ pdata->recv_buffer.chunks = chunks;
+
+ ret = hv_uio_mmap_prepare(&pdata->info, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_EQ(test, desc.action.type, MMAP_MAP_KERNEL_PAGES);
+ KUNIT_EXPECT_PTR_EQ(test, desc.action.map_kernel.pages, &pages[0]);
+ KUNIT_EXPECT_EQ(test, desc.action.map_kernel.nr_pages, 2UL);
+ /*
+ * The protection state comes from the locked pin snapshot, not
+ * from a second read of the buffer descriptor. chunks is set, so
+ * the snapshot says decrypted.
+ */
+ KUNIT_EXPECT_EQ(test, pgprot_val(desc.page_prot),
+ pgprot_val(pgprot_decrypted(PAGE_SHARED)));
+ KUNIT_EXPECT_TRUE(test, vma_desc_test(&desc, VMA_DONTEXPAND_BIT));
+ KUNIT_EXPECT_TRUE(test, vma_desc_test(&desc, VMA_DONTDUMP_BIT));
+ /*
+ * Character-device maps can carry an abandon path: kernfs is not
+ * in this path, so .close is legal here and illegal on "ring".
+ */
+ KUNIT_EXPECT_PTR_EQ(test, desc.vm_ops, &hv_uio_pin_vm_ops);
+ KUNIT_ASSERT_NOT_NULL(test, desc.vm_ops->mapped);
+ KUNIT_ASSERT_NOT_NULL(test, desc.vm_ops->close);
+ /* Buffer-backed maps leave the pin held for mapped()/close(). */
+ KUNIT_ASSERT_NOT_NULL(test, desc.private_data);
+ hv_uio_pin_release(desc.private_data);
+ desc.private_data = NULL;
+
+ desc.vma_flags = EMPTY_VMA_FLAGS;
+ KUNIT_EXPECT_EQ(test, hv_uio_mmap_prepare(&pdata->info, &desc), -EINVAL);
+
+ desc.pgoff = INT_PAGE_MAP;
+ pdata->info.mem[INT_PAGE_MAP].size = HV_HYP_PAGE_SIZE;
+ pdata->int_pages[0] = pages[0];
+ desc.end = 2 * PAGE_SIZE;
+ desc.page_prot = PAGE_SHARED;
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ ret = hv_uio_mmap_prepare(&pdata->info, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_PTR_EQ(test, desc.action.map_kernel.pages,
+ &pdata->int_pages[0]);
+ KUNIT_EXPECT_EQ(test, pgprot_val(desc.page_prot),
+ pgprot_val(PAGE_SHARED));
+ /* int_pages is not a vmbus_buffer, so nothing is pinned. */
+ KUNIT_EXPECT_NULL(test, desc.private_data);
+
+ desc.pgoff = MON_PAGE_MAP;
+ pdata->info.mem[MON_PAGE_MAP].size = HV_HYP_PAGE_SIZE;
+ pdata->monitor_pages[0] = pages[1];
+ desc.page_prot = PAGE_SHARED;
+ ret = hv_uio_mmap_prepare(&pdata->info, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_PTR_EQ(test, desc.action.map_kernel.pages,
+ &pdata->monitor_pages[0]);
+ KUNIT_EXPECT_EQ(test, pgprot_val(desc.page_prot),
+ pgprot_val(pgprot_decrypted(PAGE_SHARED)));
+ KUNIT_EXPECT_NULL(test, desc.private_data);
+
+ desc.pgoff = SEND_BUF_MAP;
+ pdata->info.mem[SEND_BUF_MAP].size = PAGE_SIZE;
+ pdata->send_buffer.pages = pages;
+ pdata->send_buffer.page_cnt = ARRAY_SIZE(pages);
+ desc.page_prot = PAGE_SHARED;
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ ret = hv_uio_mmap_prepare(&pdata->info, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ /* send_buffer has no chunks, so the snapshot says encrypted. */
+ KUNIT_EXPECT_EQ(test, pgprot_val(desc.page_prot),
+ pgprot_val(PAGE_SHARED));
+ KUNIT_ASSERT_NOT_NULL(test, desc.private_data);
+ hv_uio_pin_release(desc.private_data);
+ desc.private_data = NULL;
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+}
+
+/*
+ * The reclaim race this driver closes is a lifetime question: the pin has
+ * to span the gap between mmap_prepare() returning and insert_pages()
+ * taking its own folio references. Dropping it inside the prepare hook
+ * leaves the pages freeable before the mapping exists.
+ *
+ * This is the character-device contract: mapped() and close() share the
+ * release, and release is one-shot.
+ */
+static void hv_uio_pin_lifetime_test(struct kunit *test)
+{
+ struct hv_uio_private_data *pdata;
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ };
+ struct hv_device dev = {
+ .channel = &channel,
+ };
+ struct page *pages[4] = {};
+ struct page *chunks[1] = {};
+ struct vm_area_desc desc = {
+ .start = 0,
+ .end = 4 * PAGE_SIZE,
+ .pgoff = RECV_BUF_MAP,
+ .page_prot = PAGE_SHARED,
+ };
+ struct vm_area_struct *vma;
+ unsigned int i;
+ int ret;
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+
+ pdata = kunit_kzalloc(test, sizeof(*pdata), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pdata);
+ pdata->device = &dev;
+ pdata->info.priv = pdata;
+ pdata->info.mem[RECV_BUF_MAP].size = ARRAY_SIZE(pages) * PAGE_SIZE;
+ pdata->recv_buffer.pages = pages;
+ pdata->recv_buffer.page_cnt = ARRAY_SIZE(pages);
+ pdata->recv_buffer.chunks = chunks;
+
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ ret = hv_uio_mmap_prepare(&pdata->info, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+
+ /*
+ * Prepare must hold a reference on every page: the mapping is not
+ * installed until insert_pages() runs after the hook returns.
+ */
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 2);
+
+ KUNIT_EXPECT_PTR_EQ(test, desc.vm_ops, &hv_uio_pin_vm_ops);
+ KUNIT_ASSERT_NOT_NULL(test, desc.private_data);
+
+ /* mapped(): the mapping holds its own refs, so the pin is released. */
+ ret = desc.vm_ops->mapped(desc.start, desc.end, 0, NULL,
+ &desc.private_data);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_NULL(test, desc.private_data);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ /* close() after mapped() must not unpin again. */
+ vma = kunit_kzalloc(test, sizeof(*vma), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, vma);
+ vma->vm_ops = desc.vm_ops;
+ vma->vm_private_data = desc.private_data;
+ desc.vm_ops->close(vma);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+}
+
+/*
+ * The abandonment path: when the mapping is torn down before mapped() ran
+ * (insert_pages() failure, or a merge that never calls mapped()), close()
+ * must still drop the pin. A foreign private_data must be left alone.
+ */
+static void hv_uio_pin_close_before_mapped_test(struct kunit *test)
+{
+ struct hv_uio_private_data *pdata;
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ };
+ struct hv_device dev = {
+ .channel = &channel,
+ };
+ struct page *pages[2] = {};
+ struct page *chunks[1] = {};
+ struct vm_area_desc desc = {
+ .start = 0,
+ .end = 2 * PAGE_SIZE,
+ .pgoff = RECV_BUF_MAP,
+ .page_prot = PAGE_SHARED,
+ };
+ struct vm_area_struct *vma;
+ unsigned int i;
+ int ret;
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+
+ pdata = kunit_kzalloc(test, sizeof(*pdata), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pdata);
+ pdata->device = &dev;
+ pdata->info.priv = pdata;
+ pdata->info.mem[RECV_BUF_MAP].size = ARRAY_SIZE(pages) * PAGE_SIZE;
+ pdata->recv_buffer.pages = pages;
+ pdata->recv_buffer.page_cnt = ARRAY_SIZE(pages);
+ pdata->recv_buffer.chunks = chunks;
+
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ ret = hv_uio_mmap_prepare(&pdata->info, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_ASSERT_NOT_NULL(test, desc.private_data);
+
+ vma = kunit_kzalloc(test, sizeof(*vma), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, vma);
+ vma->vm_ops = desc.vm_ops;
+ vma->vm_private_data = desc.private_data;
+ desc.vm_ops->close(vma);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ /* A second close is a no-op, not a double unpin. */
+ desc.vm_ops->close(vma);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ /* A pointer this driver did not install must not be freed. */
+ vma->vm_private_data = &pages[0];
+ desc.vm_ops->close(vma);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+}
+
+/*
+ * The kernfs contract, named. fs/kernfs/file.c returns -EINVAL for a VMA
+ * whose vm_ops carry .close, and it performs that check after the .mmap
+ * hook has already succeeded. A successful sysfs mapping therefore cannot
+ * carry an abandon path in vm_ops at all.
+ */
+static void sysfs_ring_mmap_ops_have_no_close(struct kunit *test)
+{
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ .ringbuffer = {
+ .page_cnt = 1,
+ },
+ };
+ struct page *pages[1] = {};
+ struct page *chunks[1] = {};
+ struct vmbus_buffer_pin *pin;
+ struct vm_area_desc desc = {
+ .start = 0,
+ .end = PAGE_SIZE,
+ .pgoff = 0,
+ .page_prot = PAGE_SHARED,
+ };
+ int ret;
+
+ pages[0] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[0]);
+ channel.ringbuffer.pages = pages;
+ channel.ringbuffer.chunks = chunks;
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+
+ ret = hv_uio_ring_mmap_prepare(&channel, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+
+ /* The exact predicate kernfs_fop_mmap() applies must be false. */
+ KUNIT_EXPECT_NULL(test, desc.vm_ops);
+ KUNIT_EXPECT_FALSE(test, desc.vm_ops && desc.vm_ops->close);
+
+ pin = desc.private_data;
+ KUNIT_ASSERT_NOT_NULL(test, pin);
+ sysfs_pin_release(pin);
+
+ __free_page(pages[0]);
+}
+
+/*
+ * The character-device counterpart: a buffer-backed map prepares and
+ * installs the ops that carry the abandon path, then mapped() releases.
+ */
+static void uio_buffer_mmap_success(struct kunit *test)
+{
+ struct hv_uio_private_data *pdata;
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ };
+ struct hv_device dev = {
+ .channel = &channel,
+ };
+ struct page *pages[2] = {};
+ struct page *chunks[1] = {};
+ struct vm_area_desc desc = {
+ .start = 0,
+ .end = 2 * PAGE_SIZE,
+ .pgoff = RECV_BUF_MAP,
+ .page_prot = PAGE_SHARED,
+ };
+ unsigned int i;
+ int ret;
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+
+ pdata = kunit_kzalloc(test, sizeof(*pdata), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pdata);
+ pdata->device = &dev;
+ pdata->info.priv = pdata;
+ pdata->info.mem[RECV_BUF_MAP].size = ARRAY_SIZE(pages) * PAGE_SIZE;
+ pdata->recv_buffer.pages = pages;
+ pdata->recv_buffer.page_cnt = ARRAY_SIZE(pages);
+ pdata->recv_buffer.chunks = chunks;
+
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ ret = hv_uio_mmap_prepare(&pdata->info, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_EQ(test, desc.action.type, MMAP_MAP_KERNEL_PAGES);
+ KUNIT_EXPECT_PTR_EQ(test, desc.vm_ops, &hv_uio_pin_vm_ops);
+ KUNIT_ASSERT_NOT_NULL(test, desc.private_data);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 2);
+
+ ret = desc.vm_ops->mapped(desc.start, desc.end, 0, NULL,
+ &desc.private_data);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_NULL(test, desc.private_data);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+}
+
+/*
+ * A prepare that refuses must not leave references behind. Each failure
+ * mode is checked against the page reference count, which is the only
+ * externally visible record of a leaked pin.
+ */
+static void sysfs_mmap_prepare_failure_releases_pin(struct kunit *test)
+{
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ .ringbuffer = {
+ .page_cnt = 2,
+ },
+ };
+ struct page *pages[2] = {};
+ struct page *chunks[1] = {};
+ struct vm_area_desc desc = {
+ .start = 0,
+ .end = 2 * PAGE_SIZE,
+ .pgoff = 0,
+ .page_prot = PAGE_SHARED,
+ };
+ unsigned int i;
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+ channel.ringbuffer.pages = pages;
+ channel.ringbuffer.chunks = chunks;
+
+ /* Not shared: refused before anything is pinned. */
+ KUNIT_EXPECT_EQ(test, hv_uio_ring_mmap_prepare(&channel, &desc), -EINVAL);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ /* pgoff past the end of the pin: refused before anything is pinned. */
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+ desc.pgoff = 2;
+ KUNIT_EXPECT_EQ(test, hv_uio_ring_mmap_prepare(&channel, &desc), -EINVAL);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ /* Channel not open: refused before anything is pinned. */
+ desc.pgoff = 0;
+ channel.state = CHANNEL_OPEN_STATE;
+ KUNIT_EXPECT_EQ(test, hv_uio_ring_mmap_prepare(&channel, &desc), -ENODEV);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+}
+
+/*
+ * Prepare succeeded and the pin is installed, but the mapping never
+ * becomes established (insert_pages() fails part-way, or the attempt is
+ * otherwise abandoned). The wrapper is the only owner on this surface and
+ * releases with one unpin plus one free, whether or not any VMA ever held
+ * the pages. This proves the object it is handed is fully releasable that
+ * way, and that releasing twice is not possible to observe.
+ */
+static void sysfs_mmap_partial_insert_failure_releases_pin(struct kunit *test)
+{
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ .ringbuffer = {
+ .page_cnt = 3,
+ },
+ };
+ struct page *pages[3] = {};
+ struct page *chunks[1] = {};
+ struct vmbus_buffer_pin *pin;
+ struct vm_area_desc desc = {
+ .start = 0,
+ .end = 3 * PAGE_SIZE,
+ .pgoff = 0,
+ .page_prot = PAGE_SHARED,
+ };
+ unsigned int i;
+ int ret;
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+ channel.ringbuffer.pages = pages;
+ channel.ringbuffer.chunks = chunks;
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+
+ ret = hv_uio_ring_mmap_prepare(&channel, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ pin = desc.private_data;
+ KUNIT_ASSERT_NOT_NULL(test, pin);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 2);
+
+ /* The wrapper's abandon path: one unpin, one free. */
+ sysfs_pin_release(pin);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+}
+
+/*
+ * The window the pin exists to close. Between the snapshot and the point
+ * where the mapping holds its own references, a concurrent
+ * vmbus_release_buffer() can drop the buffer's reference to each page.
+ * The pin must keep the pages alive for the whole window, and must be the
+ * thing that drops the last reference.
+ */
+static void sysfs_mmap_release_during_insert_keeps_pages_alive(struct kunit *test)
+{
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ .ringbuffer = {
+ .page_cnt = 2,
+ },
+ };
+ struct page *pages[2] = {};
+ struct page *chunks[1] = {};
+ struct vmbus_buffer_pin *pin;
+ struct vm_area_desc desc = {
+ .start = 0,
+ .end = 2 * PAGE_SIZE,
+ .pgoff = 0,
+ .page_prot = PAGE_SHARED,
+ };
+ unsigned int i;
+ int ret;
+
+ for (i = 0; i < ARRAY_SIZE(pages); i++) {
+ pages[i] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[i]);
+ }
+ channel.ringbuffer.pages = pages;
+ channel.ringbuffer.chunks = chunks;
+ vma_desc_set_flags(&desc, VMA_SHARED_BIT);
+
+ ret = hv_uio_ring_mmap_prepare(&channel, &desc);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ pin = desc.private_data;
+ KUNIT_ASSERT_NOT_NULL(test, pin);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 2);
+
+ /*
+ * Concurrent release: the buffer drops its own reference to each
+ * page. The pages must still be alive, held only by the pin, so a
+ * mapping that is mid-insert cannot walk freed memory.
+ */
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ __free_page(pages[i]);
+ for (i = 0; i < ARRAY_SIZE(pages); i++)
+ KUNIT_EXPECT_EQ(test, page_ref_count(pages[i]), 1);
+ KUNIT_EXPECT_PTR_EQ(test, pin->pages, &pages[0]);
+ KUNIT_EXPECT_EQ(test, pin->page_count, ARRAY_SIZE(pages));
+
+ /* The pin is what drops the last reference. */
+ sysfs_pin_release(pin);
+}
+
+/*
+ * KCA-23: the protection state has to travel with the pin. The snapshot is
+ * taken in the same critical section as the page array, so a later
+ * vmbus_release_buffer() clearing the descriptor cannot change what this
+ * mapping is told about the pages it holds.
+ */
+static void ring_mmap_release_preserves_protection_snapshot(struct kunit *test)
+{
+ struct vmbus_channel channel = {
+ .state = CHANNEL_OPENED_STATE,
+ .ringbuffer = {
+ .page_cnt = 1,
+ },
+ };
+ struct page *pages[1] = {};
+ struct page *chunks[1] = {};
+ struct vmbus_buffer_pin pin = {};
+ int ret;
+
+ pages[0] = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, pages[0]);
+ channel.ringbuffer.pages = pages;
+ channel.ringbuffer.chunks = chunks;
+
+ ret = vmbus_buffer_pin_pages(&channel.ringbuffer, &pin);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_TRUE(test, pin.decrypted);
+ KUNIT_EXPECT_EQ(test, pin.page_count, 1UL);
+
+ /*
+ * A concurrent release clears the descriptor. The pin must still
+ * describe the pages it holds, not the cleared descriptor.
+ */
+ channel.ringbuffer.chunks = NULL;
+ channel.ringbuffer.pages = NULL;
+ channel.ringbuffer.page_cnt = 0;
+ KUNIT_EXPECT_TRUE(test, pin.decrypted);
+ KUNIT_EXPECT_PTR_EQ(test, pin.pages, &pages[0]);
+ KUNIT_EXPECT_EQ(test, pin.page_count, 1UL);
+
+ vmbus_buffer_unpin_pages(&pin);
+ KUNIT_EXPECT_NULL(test, pin.pages);
+ KUNIT_EXPECT_EQ(test, pin.page_count, 0UL);
+
+ /* A buffer that never took the shared-page path is not decrypted. */
+ channel.ringbuffer.pages = pages;
+ channel.ringbuffer.page_cnt = 1;
+ channel.ringbuffer.chunks = NULL;
+ ret = vmbus_buffer_pin_pages(&channel.ringbuffer, &pin);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+ KUNIT_EXPECT_FALSE(test, pin.decrypted);
+ vmbus_buffer_unpin_pages(&pin);
+
+ __free_page(pages[0]);
+}
+
+static unsigned int hv_uio_disconnect_calls;
+static int hv_uio_disconnect_result;
+
+static int hv_uio_test_disconnect(struct vmbus_channel *channel)
+{
+ hv_uio_disconnect_calls++;
+ channel->state = CHANNEL_OPEN_STATE;
+ return hv_uio_disconnect_result;
+}
+
+static void hv_uio_remove_disconnects_open_channel_test(struct kunit *test)
+{
+ struct vmbus_channel channel = { .state = CHANNEL_OPEN_STATE };
+
+ hv_uio_disconnect_calls = 0;
+ hv_uio_disconnect_result = 0;
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_if_open(&channel,
+ hv_uio_test_disconnect), 0);
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_calls, 0U);
+ channel.state = CHANNEL_OPENED_STATE;
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_if_open(&channel,
+ hv_uio_test_disconnect), 0);
+ KUNIT_EXPECT_EQ(test, channel.state, CHANNEL_OPEN_STATE);
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_calls, 1U);
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_if_open(&channel,
+ hv_uio_test_disconnect), 0);
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_calls, 1U);
+
+ channel.state = CHANNEL_OPENED_STATE;
+ hv_uio_disconnect_result = -EIO;
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_if_open(&channel,
+ hv_uio_test_disconnect), -EIO);
+ KUNIT_EXPECT_EQ(test, hv_uio_disconnect_calls, 2U);
+}
+
+static int hv_uio_test_reader_wake(struct wait_queue_entry *wait,
+ unsigned int mode, int flags, void *key)
+{
+ unsigned int *wakes = wait->private;
+
+ (*wakes)++;
+ return 1;
+}
+
+static void hv_uio_unregister_wakes_reader_test(struct kunit *test)
+{
+ struct uio_info info = {
+ .name = "hv-uio-withdrawal-test",
+ .version = "1",
+ .irq = UIO_IRQ_CUSTOM,
+ };
+ struct uio_device *idev;
+ struct device *parent;
+ wait_queue_entry_t wait;
+ unsigned int wakes = 0;
+ int ret;
+
+ parent = root_device_register("hv-uio-withdrawal-test");
+ KUNIT_ASSERT_FALSE(test, IS_ERR(parent));
+ ret = uio_register_device(parent, &info);
+ if (ret) {
+ root_device_unregister(parent);
+ KUNIT_FAIL(test, "UIO registration failed: %d", ret);
+ return;
+ }
+ idev = info.uio_dev;
+ get_device(&idev->dev);
+ init_waitqueue_func_entry(&wait, hv_uio_test_reader_wake);
+ wait.private = &wakes;
+ add_wait_queue(&idev->wait, &wait);
+
+ uio_unregister_device(&info);
+ KUNIT_EXPECT_PTR_EQ(test, idev->info, NULL);
+ KUNIT_EXPECT_EQ(test, wakes, 1U);
+ remove_wait_queue(&idev->wait, &wait);
+ put_device(&idev->dev);
+ root_device_unregister(parent);
+}
+
+static struct kunit_case hv_uio_ring_mmap_test_cases[] = {
+ KUNIT_CASE(hv_uio_unregister_wakes_reader_test),
+ KUNIT_CASE(hv_uio_remove_disconnects_open_channel_test),
+ KUNIT_CASE(hv_uio_ring_mmap_range_test),
+ KUNIT_CASE(hv_uio_ring_mmap_prepare_test),
+ KUNIT_CASE(hv_uio_mmap_region_select_test),
+ KUNIT_CASE(hv_uio_mmap_prepare_test),
+ KUNIT_CASE(hv_uio_pin_lifetime_test),
+ KUNIT_CASE(hv_uio_pin_close_before_mapped_test),
+ KUNIT_CASE(sysfs_ring_mmap_ops_have_no_close),
+ KUNIT_CASE(uio_buffer_mmap_success),
+ KUNIT_CASE(sysfs_mmap_prepare_failure_releases_pin),
+ KUNIT_CASE(sysfs_mmap_partial_insert_failure_releases_pin),
+ KUNIT_CASE(sysfs_mmap_release_during_insert_keeps_pages_alive),
+ KUNIT_CASE(ring_mmap_release_preserves_protection_snapshot),
+ {}
+};
+
+static struct kunit_suite hv_uio_ring_mmap_test_suite = {
+ .name = "hyperv-uio-mmap",
+ .test_cases = hv_uio_ring_mmap_test_cases,
+};
+
+kunit_test_suite(hv_uio_ring_mmap_test_suite);
diff --git a/include/linux/hyperv.h b/include/linux/hyperv.h
index 0f482ed5c776..f0e8f1af8e75 100644
--- a/include/linux/hyperv.h
+++ b/include/linux/hyperv.h
@@ -1248,6 +1248,30 @@ int vmbus_alloc_buffer_owned(struct vmbus_channel *channel,
void vmbus_release_buffer(struct vmbus_buffer *buffer);
+/**
+ * struct vmbus_buffer_pin - locked snapshot of a buffer's mappable pages
+ * @pages: page array owned by the buffer
+ * @page_count: number of pages in @pages
+ * @decrypted: backing pages are in the decrypted (host-shared) state
+ *
+ * vmbus_buffer_pin_pages() fills every field under
+ * vmbus_buffer_owners_lock. A mapper reads only this snapshot and must
+ * not re-read the buffer descriptor afterwards: vmbus_release_buffer()
+ * clears the descriptor under that same lock while the pin still holds
+ * the pages alive, so a later descriptor read describes the cleared
+ * descriptor rather than the pages this mapper holds.
+ */
+struct vmbus_buffer_pin {
+ struct page **pages;
+ unsigned long page_count;
+ bool decrypted;
+};
+
+int vmbus_buffer_pin_pages(struct vmbus_buffer *buffer,
+ struct vmbus_buffer_pin *pin);
+
+void vmbus_buffer_unpin_pages(struct vmbus_buffer_pin *pin);
+
void vmbus_reset_channel_cb(struct vmbus_channel *channel);
extern int vmbus_recvpacket(struct vmbus_channel *channel,
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 11/14] hv: vmbus: vmalloc requestor metadata
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (9 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 10/14] hv: vmbus: pin buffer pages across UIO mmap to close the reclaim race Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 12/14] hv: netvsc: allocate RNDIS request descriptors with kvzalloc_obj() Emerson Busson
` (2 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
Channel opens allocate request-ID arrays from the ring size. With the
128-page default ring, kvcalloc() can request an order-7 allocation. The
bitmap created by bitmap_zalloc() is another physically contiguous
allocation; at normal NetVSC sizes it can require order-1 pages. Either
allocation can fail under buddy fragmentation while order-0 pages remain
available.
Always allocate both guest-private structures with vzalloc(). Reject a
zero requestor count and check byte-size calculations before allocating,
then pair the allocations with vfree() on rollback and teardown. Neither
structure is exposed to the host, so neither needs to be decrypted for
Confidential Computing.
Extend KUnit coverage to require vmalloc backing for a typical 8 KiB
requestor array, a bitmap larger than one page, and an array above
KMALLOC_MAX_SIZE. Existing requestor cases continue to cover ID lifecycle
and cleanup.
Rollback trigger: revert if requestor metadata is exposed to the host, if
request-ID allocation or teardown regresses on a healthy channel open, or
if hyperv-vmbus-buffer KUnit fails.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/hv/channel.c | 39 ++++++--
drivers/hv/hyperv_vmbus.h | 11 +++
drivers/hv/vmbus_buffer_test.c | 160 +++++++++++++++++++++++++++++++++
3 files changed, 201 insertions(+), 9 deletions(-)
diff --git a/drivers/hv/channel.c b/drivers/hv/channel.c
index 2793ca1b7f32..c7f3f5c6c0d2 100644
--- a/drivers/hv/channel.c
+++ b/drivers/hv/channel.c
@@ -1411,14 +1411,26 @@ EXPORT_SYMBOL_GPL(vmbus_alloc_buffer);
* keeps track of the next available slot in the array. Initially, each
* slot points to the next one (as in a Linked List). The last slot
* does not point to anything, so its value is U64_MAX by default.
+ *
+ * Allocated with vzalloc() rather than kvcalloc(). kvcalloc() can use a
+ * higher-order kmalloc allocation for arrays up to KMALLOC_MAX_SIZE, so
+ * both small requestor arrays and the default 128-page ring array may need
+ * contiguous pages while opening a channel under buddy fragmentation. The
+ * array is guest-private request bookkeeping -- the host never sees the
+ * slot values -- so a vmalloc-backed mapping carries no Confidential
+ * Computing implication and needs no set_memory_decrypted().
* @size: The size of the array
*/
-static u64 *request_arr_init(u32 size)
+u64 *request_arr_init(u32 size)
{
- int i;
+ size_t bytes;
+ u32 i;
u64 *req_arr;
- req_arr = kcalloc(size, sizeof(u64), GFP_KERNEL);
+ if (!size || check_mul_overflow((size_t)size, sizeof(*req_arr), &bytes))
+ return NULL;
+
+ req_arr = vzalloc(bytes);
if (!req_arr)
return NULL;
@@ -1436,8 +1448,9 @@ static u64 *request_arr_init(u32 size)
* Index 0 is the first free slot
* @size: Size of the requestor array
*/
-static int vmbus_alloc_requestor(struct vmbus_requestor *rqstor, u32 size)
+int vmbus_alloc_requestor(struct vmbus_requestor *rqstor, u32 size)
{
+ size_t bitmap_longs, bitmap_bytes;
u64 *rqst_arr;
unsigned long *bitmap;
@@ -1445,9 +1458,17 @@ static int vmbus_alloc_requestor(struct vmbus_requestor *rqstor, u32 size)
if (!rqst_arr)
return -ENOMEM;
- bitmap = bitmap_zalloc(size, GFP_KERNEL);
+ bitmap_longs = size / BITS_PER_LONG;
+ if (size % BITS_PER_LONG)
+ bitmap_longs++;
+ if (check_mul_overflow(bitmap_longs, sizeof(*bitmap), &bitmap_bytes)) {
+ vfree(rqst_arr);
+ return -ENOMEM;
+ }
+
+ bitmap = vzalloc(bitmap_bytes);
if (!bitmap) {
- kfree(rqst_arr);
+ vfree(rqst_arr);
return -ENOMEM;
}
@@ -1464,10 +1485,10 @@ static int vmbus_alloc_requestor(struct vmbus_requestor *rqstor, u32 size)
* vmbus_free_requestor - Frees memory allocated for @rqstor
* @rqstor: Pointer to the requestor struct
*/
-static void vmbus_free_requestor(struct vmbus_requestor *rqstor)
+void vmbus_free_requestor(struct vmbus_requestor *rqstor)
{
- kfree(rqstor->req_arr);
- bitmap_free(rqstor->req_bitmap);
+ vfree(rqstor->req_arr);
+ vfree(rqstor->req_bitmap);
}
static int __vmbus_open(struct vmbus_channel *newchannel,
diff --git a/drivers/hv/hyperv_vmbus.h b/drivers/hv/hyperv_vmbus.h
index 2edeb7988bdc..9819a6b686ce 100644
--- a/drivers/hv/hyperv_vmbus.h
+++ b/drivers/hv/hyperv_vmbus.h
@@ -648,4 +648,15 @@ int vmbus_gpadl_teardown_request(struct vmbus_channel *channel,
void vmbus_complete_gpadl_teardown(struct vmbus_connection *connection,
struct vmbus_channel_gpadl_torndown *response);
+/*
+ * Requestor array lifetime helpers, shared with vmbus_buffer_test.c for
+ * the same reason as the sizing helpers. Defined in channel.c, unexported.
+ * request_arr_init() returns a vmalloc-backed array the caller must vfree();
+ * vmbus_alloc_requestor() owns both vmalloc-backed objects on success, and
+ * vmbus_free_requestor() releases them.
+ */
+u64 *request_arr_init(u32 size);
+int vmbus_alloc_requestor(struct vmbus_requestor *rqstor, u32 size);
+void vmbus_free_requestor(struct vmbus_requestor *rqstor);
+
#endif /* _HYPERV_VMBUS_H */
diff --git a/drivers/hv/vmbus_buffer_test.c b/drivers/hv/vmbus_buffer_test.c
index 5c8e70d861ad..d0102dabef73 100644
--- a/drivers/hv/vmbus_buffer_test.c
+++ b/drivers/hv/vmbus_buffer_test.c
@@ -10,6 +10,7 @@
#include <linux/completion.h>
#include <linux/hyperv.h>
#include <linux/mm.h>
+#include <linux/sizes.h>
#include <linux/slab.h>
#include <linux/vmalloc.h>
@@ -776,6 +777,160 @@ static void vmbus_buffer_order_zero_allocation_test(struct kunit *test)
KUNIT_EXPECT_EQ(test, context.attempts, (unsigned int)MAX_PAGE_ORDER + 1);
}
+/*
+ * Requestor metadata must use vmalloc backing because both the array and its
+ * bitmap can require multiple pages. The requestor is guest-private
+ * bookkeeping: the host never sees the slot values, so this mapping has no
+ * Confidential Computing implication and needs no set_memory_decrypted().
+ * Cover a typical 8 KiB array, a requestor whose bitmap exceeds one page,
+ * and a size beyond KMALLOC_MAX_SIZE without memory pressure.
+ */
+static void vmbus_requestor_alloc_free_test(struct kunit *test)
+{
+ struct vmbus_requestor rqstor = {};
+
+ KUNIT_ASSERT_EQ(test, vmbus_alloc_requestor(&rqstor, 4), 0);
+ KUNIT_EXPECT_EQ(test, rqstor.size, 4U);
+ KUNIT_EXPECT_EQ(test, rqstor.next_request_id, 0U);
+ KUNIT_EXPECT_NOT_NULL(test, rqstor.req_arr);
+ KUNIT_EXPECT_NOT_NULL(test, rqstor.req_bitmap);
+
+ /* The free list links 0->1->2->3 and terminates in U64_MAX. */
+ KUNIT_EXPECT_EQ(test, rqstor.req_arr[0], 1U);
+ KUNIT_EXPECT_EQ(test, rqstor.req_arr[1], 2U);
+ KUNIT_EXPECT_EQ(test, rqstor.req_arr[2], 3U);
+ KUNIT_EXPECT_EQ(test, rqstor.req_arr[3], U64_MAX);
+
+ vmbus_free_requestor(&rqstor);
+}
+
+static void vmbus_requestor_vmalloc_backing_test(struct kunit *test)
+{
+ /*
+ * Cover the small array, a requestor whose bitmap exceeds one page,
+ * and a size beyond KMALLOC_MAX_SIZE. All requestor metadata must use
+ * vmalloc backing regardless of size or allocator pressure.
+ */
+ u32 small_size = 1024;
+ u32 bitmap_size = (PAGE_SIZE / sizeof(unsigned long)) * BITS_PER_LONG + 1;
+ u32 size = (KMALLOC_MAX_SIZE / sizeof(u64)) + 1;
+ struct vmbus_requestor rqstor = {};
+ u64 *req_arr;
+
+ KUNIT_EXPECT_PTR_EQ(test, request_arr_init(0), NULL);
+
+ req_arr = request_arr_init(small_size);
+ KUNIT_ASSERT_NOT_NULL(test, req_arr);
+ KUNIT_EXPECT_TRUE(test, is_vmalloc_addr(req_arr));
+ KUNIT_EXPECT_EQ(test, req_arr[0], 1U);
+ KUNIT_EXPECT_EQ(test, req_arr[small_size - 1], U64_MAX);
+ vfree(req_arr);
+
+ req_arr = request_arr_init(size);
+ KUNIT_ASSERT_NOT_NULL(test, req_arr);
+ KUNIT_EXPECT_TRUE(test, is_vmalloc_addr(req_arr));
+ KUNIT_EXPECT_EQ(test, req_arr[0], 1U);
+ KUNIT_EXPECT_EQ(test, req_arr[size - 1], U64_MAX);
+ vfree(req_arr);
+
+ KUNIT_ASSERT_EQ(test, vmbus_alloc_requestor(&rqstor, bitmap_size), 0);
+ KUNIT_EXPECT_TRUE(test, is_vmalloc_addr(rqstor.req_arr));
+ KUNIT_EXPECT_TRUE(test, is_vmalloc_addr(rqstor.req_bitmap));
+ vmbus_free_requestor(&rqstor);
+}
+
+static void vmbus_requestor_id_lifecycle_test(struct kunit *test)
+{
+ struct vmbus_channel channel = { .rqstor_size = 4 };
+ struct vmbus_requestor *rqstor = &channel.requestor;
+ u64 id0, id1, id2, id3, addr;
+
+ KUNIT_ASSERT_EQ(test, vmbus_alloc_requestor(rqstor, 4), 0);
+
+ /* IDs are 1-based; 0 is reserved for unsolicited host messages. */
+ id0 = vmbus_next_request_id(&channel, 0x1000);
+ KUNIT_EXPECT_EQ(test, id0, 1U);
+ id1 = vmbus_next_request_id(&channel, 0x2000);
+ KUNIT_EXPECT_EQ(test, id1, 2U);
+
+ /* Consumption returns the registered address and frees the slot. */
+ addr = vmbus_request_addr_match(&channel, id0, VMBUS_RQST_ADDR_ANY);
+ KUNIT_EXPECT_EQ(test, addr, 0x1000U);
+ addr = vmbus_request_addr_match(&channel, id1, 0x2000);
+ KUNIT_EXPECT_EQ(test, addr, 0x2000U);
+
+ /*
+ * Consumed slots are reusable. The free list is LIFO -- each
+ * consume pushes its slot onto the head -- so the slot freed last
+ * is the one handed out first. Freeing id0 then id1 leaves slot 1
+ * at the head, and the next ID is therefore id1 again, not id0.
+ */
+ id2 = vmbus_next_request_id(&channel, 0x3000);
+ KUNIT_EXPECT_EQ(test, id2, id1);
+ id3 = vmbus_next_request_id(&channel, 0x4000);
+ KUNIT_EXPECT_EQ(test, id3, id0);
+
+ vmbus_free_requestor(rqstor);
+}
+
+static void vmbus_requestor_invalid_ids_test(struct kunit *test)
+{
+ struct vmbus_channel channel = { .rqstor_size = 4 };
+ struct vmbus_requestor *rqstor = &channel.requestor;
+ u64 id, addr;
+
+ KUNIT_ASSERT_EQ(test, vmbus_alloc_requestor(rqstor, 4), 0);
+
+ /* ID 0 is the unsolicited-message sentinel and is never in the set. */
+ KUNIT_EXPECT_EQ(test, vmbus_request_addr_match(&channel, 0, 0),
+ VMBUS_RQST_ERROR);
+
+ /* Out-of-range IDs are refused, not wrapped. */
+ KUNIT_EXPECT_EQ(test, vmbus_request_addr_match(&channel, 5, 0),
+ VMBUS_RQST_ERROR);
+ KUNIT_EXPECT_EQ(test, vmbus_request_addr_match(&channel, U64_MAX, 0),
+ VMBUS_RQST_ERROR);
+
+ id = vmbus_next_request_id(&channel, 0xABCD);
+ KUNIT_ASSERT_EQ(test, id, 1U);
+
+ /* A wrong expected address does not consume the slot. */
+ addr = vmbus_request_addr_match(&channel, id, 0x9999);
+ KUNIT_EXPECT_EQ(test, addr, 0xABCDU);
+
+ /* The slot is still live and a matching lookup consumes it once. */
+ addr = vmbus_request_addr_match(&channel, id, 0xABCD);
+ KUNIT_EXPECT_EQ(test, addr, 0xABCDU);
+
+ /* Replay: the slot is gone. */
+ addr = vmbus_request_addr_match(&channel, id, VMBUS_RQST_ADDR_ANY);
+ KUNIT_EXPECT_EQ(test, addr, VMBUS_RQST_ERROR);
+
+ vmbus_free_requestor(rqstor);
+}
+
+static void vmbus_requestor_exhaustion_test(struct kunit *test)
+{
+ struct vmbus_channel channel = { .rqstor_size = 2 };
+ struct vmbus_requestor *rqstor = &channel.requestor;
+
+ KUNIT_ASSERT_EQ(test, vmbus_alloc_requestor(rqstor, 2), 0);
+
+ KUNIT_EXPECT_EQ(test, vmbus_next_request_id(&channel, 0x1), 1U);
+ KUNIT_EXPECT_EQ(test, vmbus_next_request_id(&channel, 0x2), 2U);
+
+ /* Both slots taken: the free list is empty and the API says so. */
+ KUNIT_EXPECT_EQ(test, vmbus_next_request_id(&channel, 0x3),
+ VMBUS_RQST_ERROR);
+
+ /* An uninitialized requestor is not a full one. */
+ channel.rqstor_size = 0;
+ KUNIT_EXPECT_EQ(test, vmbus_next_request_id(&channel, 0x4),
+ VMBUS_NO_RQSTOR);
+
+ vmbus_free_requestor(rqstor);
+}
+
static struct kunit_case vmbus_buffer_test_cases[] = {
KUNIT_CASE(vmbus_buffer_size_rounding_test),
KUNIT_CASE(vmbus_buffer_size_overflow_test),
@@ -805,6 +960,11 @@ static struct kunit_case vmbus_buffer_test_cases[] = {
KUNIT_CASE(vmbus_gpadl_post_success_test),
KUNIT_CASE(vmbus_gpadl_response_state_test),
KUNIT_CASE(vmbus_gpadl_teardown_post_failure_test),
+ KUNIT_CASE(vmbus_requestor_alloc_free_test),
+ KUNIT_CASE(vmbus_requestor_vmalloc_backing_test),
+ KUNIT_CASE(vmbus_requestor_id_lifecycle_test),
+ KUNIT_CASE(vmbus_requestor_invalid_ids_test),
+ KUNIT_CASE(vmbus_requestor_exhaustion_test),
{}
};
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 12/14] hv: netvsc: allocate RNDIS request descriptors with kvzalloc_obj()
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (10 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 11/14] hv: vmbus: vmalloc requestor metadata Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 13/14] hv: netvsc: handle a NULL request address on empty completions Emerson Busson
2026-10-07 19:07 ` [PATCH v2 14/14] hv: netvsc: use kvzalloc for device state Emerson Busson
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
struct rndis_request embeds two RNDIS_EXT_LEN protocol tails.
RNDIS_EXT_LEN is HV_HYP_PAGE_SIZE, so the two tails alone are 8 KiB and
the struct is just over KMALLOC_MAX_CACHE_SIZE. kzalloc_obj() therefore
needs an order-2 compound page for every RNDIS control request, which can
fail under buddy fragmentation.
Use kvzalloc_obj() and kvfree() so allocation can fall back to vmalloc.
RNDIS control payloads are sent as GPA-direct page buffers, however, so a
vmalloc-backed descriptor cannot be described by one virt_to_phys() PFN.
Build one hv_page_buffer per Linux/Hyper-V page and resolve vmalloc backing
with vmalloc_to_page(). Linear allocations use virt_to_page(). This keeps
the GPA list valid and preserves the existing per-range DMA mapping for
isolated guests.
Keep both protocol tails unchanged. The response is copied after
request->response_msg into response_ext, within the existing
sizeof(struct rndis_message) + RNDIS_EXT_LEN bound.
Adds a fifth named RNDIS request KUnit case. It builds GPA descriptors from
a vmalloc buffer that crosses a page boundary and checks every PFN, offset,
and length against the backing pages. The existing cases retain the
allocation-size, response-layout, zeroing, request-id, and cleanup checks.
Rollback trigger: revert if the vmalloc page-buffer KUnit case fails or an
exact-source fragmented rebind shows malformed GPA descriptors from this
path.
Fixes: 5b54dac856cb ("hyperv: Add support for virtual Receive Side Scaling (vRSS)")
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/net/hyperv/Kconfig | 12 ++
drivers/net/hyperv/Makefile | 1 +
drivers/net/hyperv/hyperv_net.h | 43 ++++-
drivers/net/hyperv/rndis_filter.c | 117 ++++++++-----
drivers/net/hyperv/rndis_request_test.c | 216 ++++++++++++++++++++++++
5 files changed, 349 insertions(+), 40 deletions(-)
create mode 100644 drivers/net/hyperv/rndis_request_test.c
diff --git a/drivers/net/hyperv/Kconfig b/drivers/net/hyperv/Kconfig
index 982964c1a9fb..226728fd77d7 100644
--- a/drivers/net/hyperv/Kconfig
+++ b/drivers/net/hyperv/Kconfig
@@ -6,3 +6,15 @@ config HYPERV_NET
select NLS
help
Select this option to enable the Hyper-V virtual network driver.
+
+config HYPERV_NET_KUNIT_TEST
+ bool "Build Hyper-V netvsc KUnit tests"
+ depends on HYPERV_NET && KUNIT
+ default KUNIT_ALL_TESTS
+ help
+ Build the netvsc KUnit test suites into the hv_netvsc object.
+ Built into the module rather than as a separate one so the
+ cases can reach the internal helpers declared in hyperv_net.h
+ without exporting them.
+
+ If unsure, say N.
diff --git a/drivers/net/hyperv/Makefile b/drivers/net/hyperv/Makefile
index 0db7ccaec4a4..6f1abc756fde 100644
--- a/drivers/net/hyperv/Makefile
+++ b/drivers/net/hyperv/Makefile
@@ -2,3 +2,4 @@
obj-$(CONFIG_HYPERV_NET) += hv_netvsc.o
hv_netvsc-y := netvsc_drv.o netvsc.o rndis_filter.o netvsc_trace.o netvsc_bpf.o
+hv_netvsc-$(CONFIG_HYPERV_NET_KUNIT_TEST) += rndis_request_test.o
diff --git a/drivers/net/hyperv/hyperv_net.h b/drivers/net/hyperv/hyperv_net.h
index a15cb2460344..6492e7e93ded 100644
--- a/drivers/net/hyperv/hyperv_net.h
+++ b/drivers/net/hyperv/hyperv_net.h
@@ -210,14 +210,25 @@ struct rndis_device {
u8 rss_key[NETVSC_HASH_KEYLEN];
};
+#define RNDIS_EXT_LEN HV_HYP_PAGE_SIZE
-/* Interface */
+/* Full definitions are below, after the RNDIS message format. */
+struct rndis_request;
struct rndis_message;
+
+/* Interface */
struct ndis_offload_params;
struct netvsc_device;
struct netvsc_channel;
struct net_device_context;
+struct rndis_request *get_rndis_request(struct rndis_device *dev,
+ u32 msg_type, u32 msg_len);
+void put_rndis_request(struct rndis_device *dev, struct rndis_request *req);
+int rndis_build_page_buffers(const void *data, u32 len,
+ struct hv_page_buffer *page_bufs,
+ u32 *page_buf_cnt);
+
extern u32 netvsc_ring_bytes;
int netvsc_workqueue_init(void);
@@ -1779,6 +1790,36 @@ struct rndis_message {
union rndis_message_container msg;
};
+/*
+ * RNDIS request descriptor. Declared here so the KUnit cases can reach
+ * the layout the response-copy bound in rndis_filter_receive_resp()
+ * relies on. The two RNDIS_EXT_LEN tails are what make the object
+ * larger than KMALLOC_MAX_CACHE_SIZE.
+ */
+struct rndis_request {
+ struct list_head list_ent;
+ struct completion wait_event;
+
+ struct rndis_message response_msg;
+ /*
+ * The buffer for extended info after the RNDIS response message. It's
+ * referenced based on the data offset in the RNDIS message. Its size
+ * is enough for current needs, and should be sufficient for the near
+ * future.
+ */
+ u8 response_ext[RNDIS_EXT_LEN];
+
+ /* Simplify allocation by having a netvsc packet inline */
+ struct hv_netvsc_packet pkt;
+
+ struct rndis_message request_msg;
+ /*
+ * The buffer for the extended info after the RNDIS request message.
+ * It is referenced and sized in a similar way as response_ext.
+ */
+ u8 request_ext[RNDIS_EXT_LEN];
+};
+
/* Handy macros */
diff --git a/drivers/net/hyperv/rndis_filter.c b/drivers/net/hyperv/rndis_filter.c
index 9b6c44979b4e..1ce40baa2d23 100644
--- a/drivers/net/hyperv/rndis_filter.c
+++ b/drivers/net/hyperv/rndis_filter.c
@@ -16,6 +16,7 @@
#include <linux/if_ether.h>
#include <linux/netdevice.h>
#include <linux/if_vlan.h>
+#include <linux/mm.h>
#include <linux/nls.h>
#include <linux/vmalloc.h>
#include <linux/rtnetlink.h>
@@ -27,31 +28,6 @@
static void rndis_set_multicast(struct work_struct *w);
-#define RNDIS_EXT_LEN HV_HYP_PAGE_SIZE
-struct rndis_request {
- struct list_head list_ent;
- struct completion wait_event;
-
- struct rndis_message response_msg;
- /*
- * The buffer for extended info after the RNDIS response message. It's
- * referenced based on the data offset in the RNDIS message. Its size
- * is enough for current needs, and should be sufficient for the near
- * future.
- */
- u8 response_ext[RNDIS_EXT_LEN];
-
- /* Simplify allocation by having a netvsc packet inline */
- struct hv_netvsc_packet pkt;
-
- struct rndis_message request_msg;
- /*
- * The buffer for the extended info after the RNDIS request message.
- * It is referenced and sized in a similar way as response_ext.
- */
- u8 request_ext[RNDIS_EXT_LEN];
-};
-
static const u8 netvsc_hash_key[NETVSC_HASH_KEYLEN] = {
0x6d, 0x5a, 0x56, 0xda, 0x25, 0x5b, 0x0e, 0xc2,
0x41, 0x67, 0x25, 0x3d, 0x43, 0xa3, 0x8f, 0xb0,
@@ -78,16 +54,27 @@ static struct rndis_device *get_rndis_device(void)
return device;
}
-static struct rndis_request *get_rndis_request(struct rndis_device *dev,
- u32 msg_type,
- u32 msg_len)
+struct rndis_request *get_rndis_request(struct rndis_device *dev,
+ u32 msg_type,
+ u32 msg_len)
{
struct rndis_request *request;
struct rndis_message *rndis_msg;
struct rndis_set_request *set;
unsigned long flags;
- request = kzalloc_obj(struct rndis_request);
+ /*
+ * struct rndis_request embeds two RNDIS_EXT_LEN (4 KiB) protocol
+ * tails and is larger than KMALLOC_MAX_CACHE_SIZE, so
+ * kzalloc_obj() needs an order-2 compound page. That high-order
+ * allocation can fail under buddy fragmentation. kvzalloc_obj()
+ * keeps the zeroing semantics while allowing vmalloc backing when
+ * a suitable kmalloc allocation is unavailable. The request payload
+ * is sent as GPADL-direct page buffers, so rndis_filter_send_request()
+ * must describe each backing page instead of assuming virt_to_phys()
+ * yields one contiguous physical range.
+ */
+ request = kvzalloc_obj(struct rndis_request);
if (!request)
return NULL;
@@ -115,8 +102,8 @@ static struct rndis_request *get_rndis_request(struct rndis_device *dev,
return request;
}
-static void put_rndis_request(struct rndis_device *dev,
- struct rndis_request *req)
+void put_rndis_request(struct rndis_device *dev,
+ struct rndis_request *req)
{
unsigned long flags;
@@ -124,7 +111,8 @@ static void put_rndis_request(struct rndis_device *dev,
list_del(&req->list_ent);
spin_unlock_irqrestore(&dev->request_lock, flags);
- kfree(req);
+ /* Paired with the kvzalloc_obj() in get_rndis_request(). */
+ kvfree(req);
}
static void dump_rndis_message(struct net_device *netdev,
@@ -221,27 +209,78 @@ static void dump_rndis_message(struct net_device *netdev,
}
}
+int rndis_build_page_buffers(const void *data, u32 len,
+ struct hv_page_buffer *page_bufs,
+ u32 *page_buf_cnt)
+{
+ const u8 *addr = data;
+ u32 count = 0;
+
+ if (!data || !len || !page_bufs || !page_buf_cnt)
+ return -EINVAL;
+
+ *page_buf_cnt = 0;
+
+ while (len) {
+ struct page *page;
+ phys_addr_t phys;
+ u32 offset, chunk;
+ unsigned long page_offset = offset_in_page(addr);
+
+ if (count == MAX_PAGE_BUFFER_COUNT)
+ return -E2BIG;
+
+ if (is_vmalloc_addr(addr))
+ page = vmalloc_to_page(addr);
+ else if (virt_addr_valid(addr))
+ page = virt_to_page(addr);
+ else
+ return -EFAULT;
+
+ if (!page)
+ return -EFAULT;
+
+ phys = page_to_phys(page) + page_offset;
+ offset = offset_in_hvpage(phys);
+ chunk = min_t(u32, len, PAGE_SIZE - page_offset);
+ chunk = min_t(u32, chunk, HV_HYP_PAGE_SIZE - offset);
+
+ page_bufs[count].pfn = phys >> HV_HYP_PAGE_SHIFT;
+ page_bufs[count].offset = offset;
+ page_bufs[count].len = chunk;
+
+ addr += chunk;
+ len -= chunk;
+ count++;
+ }
+
+ *page_buf_cnt = count;
+ return 0;
+}
+
static int rndis_filter_send_request(struct rndis_device *dev,
struct rndis_request *req)
{
struct hv_netvsc_packet *packet;
- struct hv_page_buffer pb;
+ struct hv_page_buffer page_bufs[MAX_PAGE_BUFFER_COUNT];
+ u32 page_buf_cnt;
int ret;
/* Setup the packet to send it */
packet = &req->pkt;
packet->total_data_buflen = req->request_msg.msg_len;
- packet->page_buf_cnt = 1;
-
- pb.pfn = virt_to_phys(&req->request_msg) >> HV_HYP_PAGE_SHIFT;
- pb.len = req->request_msg.msg_len;
- pb.offset = offset_in_hvpage(&req->request_msg);
+ ret = rndis_build_page_buffers(&req->request_msg,
+ req->request_msg.msg_len,
+ page_bufs, &page_buf_cnt);
+ if (ret)
+ return ret;
+ packet->page_buf_cnt = page_buf_cnt;
trace_rndis_send(dev->ndev, 0, &req->request_msg);
rcu_read_lock_bh();
- ret = netvsc_send(dev->ndev, packet, NULL, &pb, NULL, false);
+ ret = netvsc_send(dev->ndev, packet, NULL, page_bufs, NULL, false);
rcu_read_unlock_bh();
return ret;
diff --git a/drivers/net/hyperv/rndis_request_test.c b/drivers/net/hyperv/rndis_request_test.c
new file mode 100644
index 000000000000..6affa916e74e
--- /dev/null
+++ b/drivers/net/hyperv/rndis_request_test.c
@@ -0,0 +1,216 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * KUnit tests for RNDIS request descriptor allocation.
+ *
+ * Built into the hv_netvsc object rather than a separate module so the
+ * cases can reach the internal helpers declared in hyperv_net.h
+ * without exporting them.
+ */
+#include <kunit/test.h>
+#include <linux/hyperv.h>
+#include <linux/mm.h>
+#include <linux/slab.h>
+#include <linux/stddef.h>
+#include <linux/vmalloc.h>
+
+#include "hyperv_net.h"
+
+static int rndis_test_device_init(struct rndis_device *dev)
+{
+ memset(dev, 0, sizeof(*dev));
+ atomic_set(&dev->new_req_id, 0);
+ spin_lock_init(&dev->request_lock);
+ INIT_LIST_HEAD(&dev->req_list);
+ return 0;
+}
+
+/*
+ * The premise of the kvzalloc_obj() conversion: the descriptor embeds
+ * two RNDIS_EXT_LEN protocol tails and sits past
+ * KMALLOC_MAX_CACHE_SIZE, so kzalloc_obj() reached kmalloc_large() and
+ * took an order-2 compound page on the subchannel open path. If the
+ * struct ever drops back inside the cache, the conversion stops being
+ * about high-order allocation and this assertion is the reminder.
+ */
+static void rndis_request_size_premise_test(struct kunit *test)
+{
+ KUNIT_EXPECT_GT(test, sizeof(struct rndis_request),
+ (size_t)KMALLOC_MAX_CACHE_SIZE);
+ /*
+ * The copy bound in rndis_filter_receive_resp() is
+ * sizeof(struct rndis_message) + RNDIS_EXT_LEN. That is exactly
+ * the response_msg + response_ext tail of the object. A reorder
+ * or a resize of either member makes the bound lie.
+ */
+ KUNIT_EXPECT_EQ(test,
+ offsetofend(struct rndis_request, response_ext),
+ offsetof(struct rndis_request, response_msg) +
+ sizeof(struct rndis_message) + RNDIS_EXT_LEN);
+ KUNIT_EXPECT_EQ(test,
+ offsetofend(struct rndis_request, request_ext),
+ offsetof(struct rndis_request, request_msg) +
+ sizeof(struct rndis_message) + RNDIS_EXT_LEN);
+}
+
+/*
+ * kvzalloc_obj() keeps the zeroing kzalloc_obj() provided. The two
+ * protocol tails are the fields a response is copied into; a stale
+ * tail is a stale response.
+ */
+static void rndis_request_alloc_zeroed_test(struct kunit *test)
+{
+ struct rndis_device dev;
+ struct rndis_request *req;
+
+ KUNIT_ASSERT_EQ(test, rndis_test_device_init(&dev), 0);
+ req = get_rndis_request(&dev, RNDIS_MSG_INIT, 0x20);
+ KUNIT_ASSERT_NOT_NULL(test, req);
+
+ KUNIT_EXPECT_EQ(test, req->response_ext[0], 0);
+ KUNIT_EXPECT_EQ(test, req->response_ext[RNDIS_EXT_LEN - 1], 0);
+ KUNIT_EXPECT_EQ(test, req->request_ext[0], 0);
+ KUNIT_EXPECT_EQ(test, req->request_ext[RNDIS_EXT_LEN - 1], 0);
+ KUNIT_EXPECT_EQ(test, req->response_msg.msg_len, 0U);
+
+ put_rndis_request(&dev, req);
+}
+
+/* Request ids come from a per-device counter and the list tracks the live set. */
+static void rndis_request_id_lifecycle_test(struct kunit *test)
+{
+ struct rndis_device dev;
+ struct rndis_request *req0, *req1;
+ struct rndis_request *cursor;
+ unsigned int n = 0;
+
+ KUNIT_ASSERT_EQ(test, rndis_test_device_init(&dev), 0);
+
+ req0 = get_rndis_request(&dev, RNDIS_MSG_INIT, 0x20);
+ KUNIT_ASSERT_NOT_NULL(test, req0);
+ req1 = get_rndis_request(&dev, RNDIS_MSG_QUERY, 0x20);
+ KUNIT_ASSERT_NOT_NULL(test, req1);
+
+ /*
+ * get_rndis_request() stamps the id through the set_req template,
+ * which shares the union slot with every other request body. Read
+ * it back through the same member the producer used.
+ */
+ KUNIT_EXPECT_EQ(test,
+ req0->request_msg.msg.set_req.req_id, 1U);
+ KUNIT_EXPECT_EQ(test,
+ req1->request_msg.msg.set_req.req_id, 2U);
+
+ list_for_each_entry(cursor, &dev.req_list, list_ent)
+ n++;
+ KUNIT_EXPECT_EQ(test, n, 2U);
+
+ /*
+ * put_rndis_request() is the cleanup path every error and
+ * timeout branch takes. It unlinks before it frees.
+ */
+ put_rndis_request(&dev, req0);
+ n = 0;
+ list_for_each_entry(cursor, &dev.req_list, list_ent)
+ n++;
+ KUNIT_EXPECT_EQ(test, n, 1U);
+ KUNIT_EXPECT_TRUE(test,
+ list_first_entry(&dev.req_list, struct rndis_request,
+ list_ent) == req1);
+
+ put_rndis_request(&dev, req1);
+ KUNIT_EXPECT_TRUE(test, list_empty(&dev.req_list));
+}
+
+/*
+ * A request left on the list after a completed exchange is a leak of
+ * the descriptor and of its slot in the outstanding set. Repeated
+ * get/put cycles must return to the empty list every time.
+ */
+static void rndis_request_put_cleanup_test(struct kunit *test)
+{
+ struct rndis_device dev;
+ struct rndis_request *req;
+ int i;
+
+ KUNIT_ASSERT_EQ(test, rndis_test_device_init(&dev), 0);
+
+ for (i = 0; i < 4; i++) {
+ req = get_rndis_request(&dev, RNDIS_MSG_INIT, 0x20);
+ KUNIT_ASSERT_NOT_NULL(test, req);
+ KUNIT_EXPECT_FALSE(test, list_empty(&dev.req_list));
+ put_rndis_request(&dev, req);
+ KUNIT_EXPECT_TRUE(test, list_empty(&dev.req_list));
+ }
+}
+
+/*
+ * kvzalloc_obj() may return a vmalloc address under fragmentation. RNDIS
+ * requests are GPA-direct packets, so each backing page must be represented
+ * explicitly instead of deriving a single PFN with virt_to_phys().
+ */
+static void rndis_request_vmalloc_page_buffers_test(struct kunit *test)
+{
+ struct hv_page_buffer page_bufs[MAX_PAGE_BUFFER_COUNT];
+ const u32 data_len = 16;
+ u8 *allocation = vzalloc(2 * PAGE_SIZE);
+ const u8 *cursor;
+ phys_addr_t phys;
+ u32 count = 0, remaining = data_len, i;
+ int ret;
+
+ if (!allocation) {
+ KUNIT_FAIL(test, "could not allocate vmalloc test buffer");
+ return;
+ }
+
+ cursor = allocation + PAGE_SIZE - 8;
+ ret = rndis_build_page_buffers(cursor, data_len, page_bufs, &count);
+ KUNIT_EXPECT_EQ(test, ret, 0);
+ if (ret)
+ goto out;
+
+ KUNIT_EXPECT_EQ(test, count, 2U);
+ for (i = 0; i < count; i++) {
+ struct page *page = vmalloc_to_page(cursor);
+ unsigned long page_offset = offset_in_page(cursor);
+ u32 offset, chunk;
+
+ if (!page) {
+ KUNIT_FAIL(test, "vmalloc test buffer has no backing page");
+ goto out;
+ }
+
+ phys = page_to_phys(page) + page_offset;
+ offset = offset_in_hvpage(phys);
+ chunk = min_t(u32, remaining, PAGE_SIZE - page_offset);
+ chunk = min_t(u32, chunk, HV_HYP_PAGE_SIZE - offset);
+
+ KUNIT_EXPECT_EQ(test, page_bufs[i].pfn,
+ (u64)(phys >> HV_HYP_PAGE_SHIFT));
+ KUNIT_EXPECT_EQ(test, page_bufs[i].offset, offset);
+ KUNIT_EXPECT_EQ(test, page_bufs[i].len, chunk);
+
+ cursor += chunk;
+ remaining -= chunk;
+ }
+ KUNIT_EXPECT_EQ(test, remaining, 0U);
+
+out:
+ vfree(allocation);
+}
+
+static struct kunit_case rndis_request_test_cases[] = {
+ KUNIT_CASE(rndis_request_size_premise_test),
+ KUNIT_CASE(rndis_request_alloc_zeroed_test),
+ KUNIT_CASE(rndis_request_id_lifecycle_test),
+ KUNIT_CASE(rndis_request_put_cleanup_test),
+ KUNIT_CASE(rndis_request_vmalloc_page_buffers_test),
+ {}
+};
+
+static struct kunit_suite rndis_request_test_suite = {
+ .name = "hyperv-rndis-request",
+ .test_cases = rndis_request_test_cases,
+};
+
+kunit_test_suite(rndis_request_test_suite);
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 13/14] hv: netvsc: handle a NULL request address on empty completions
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (11 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 12/14] hv: netvsc: allocate RNDIS request descriptors with kvzalloc_obj() Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
2026-10-07 19:07 ` [PATCH v2 14/14] hv: netvsc: use kvzalloc for device state Emerson Busson
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
netvsc_send_pkt() registers (ulong)skb as the VMBus request address.
Control RNDIS sends carry no skb -- rndis_filter.c calls netvsc_send()
with skb == NULL -- so the registered address is NULL. That is
intentional: there is no guest object to hand back, and the request
slot is still allocated so the completion can reclaim it.
netvsc_send_tx_complete() already tolerates that with if (likely(skb)).
netvsc_send_completion()'s empty-payload branch does not. It casts the
request address to struct nvsp_message * and reads hdr.msg_type
unconditionally, so a completion that resolves to NULL is a fatal NULL
dereference in NAPI/softirq context.
The empty-payload branch exists for NVSP_MSG4_TYPE_SWITCH_DATA_PATH,
which netvsc_switch_datapath() sends with a real nvsp_message address.
A NULL request address is not that message and must not be
dereferenced. Validate VMBUS_NO_RQSTOR alongside VMBUS_RQST_ERROR in
both completion paths -- request_addr_callback() returns it when the
channel has no requestor -- and account a NULL-address completion the
same way the payload path accounts a NULL skb. The request id is
consumed by the lookup, so that completion is the one that owns the
queue_sends decrement and the possible queue wake.
This is reached by the runtime drill in this series once channel open
survives buddy fragmentation: control RNDIS traffic proceeds where it
used to fail with -ENOMEM, and an empty completion for one of those
requests takes the unguarded path.
Revert this patch if an empty completion with a NULL request address
dereferences again, if a netvsc TX stall or an "Invalid transaction
ID" flood appears under this patch, or if a SWITCH_DATA_PATH
completion is shown to be dropped instead of taken through its
nvsp_message path.
Adds eleven named netvsc completion KUnit cases covering the decision
layer of the empty-payload branch: a NULL request address owning the
queue_sends decrement, both invalid sentinels rejected without
accounting, SWITCH_DATA_PATH completing channel_init_wait without
accounting, an unexpected message type taking neither path, a
replayed transaction id accounting exactly once, N completions
decrementing the queue exactly N times, the destroy drain wait waking
only at zero, and netvsc_send_tx_complete() applying the same
sentinel rejection while still accounting a NULL skb. Two skb-backed
payload cases exercise the production completion entry with success
and error status, a published send slot and queue 1. They verify
synchronous skb consumption, exact packet/byte statistics, selected
queue accounting and duplicate transaction rejection. Confidential
DMA unmapping remains platform-specific and is not exercised here.
netvsc_send_acct() is extracted so a test can observe the decrement
and the drain wake directly; it is not a behaviour change. The
helpers lose static and are declared in hyperv_net.h so the cases
reach them without a new EXPORT_SYMBOL_GPL, matching how the VMBus
buffer tests are built into hv_vmbus.
Fixes: 8b31f8c982b7 ("hv_netvsc: Wait for completion on request SWITCH_DATA_PATH")
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/net/hyperv/Makefile | 1 +
drivers/net/hyperv/hyperv_net.h | 15 +
drivers/net/hyperv/netvsc.c | 81 ++--
drivers/net/hyperv/netvsc_completion_test.c | 439 ++++++++++++++++++++
4 files changed, 507 insertions(+), 29 deletions(-)
create mode 100644 drivers/net/hyperv/netvsc_completion_test.c
diff --git a/drivers/net/hyperv/Makefile b/drivers/net/hyperv/Makefile
index 6f1abc756fde..b2d22855fd8a 100644
--- a/drivers/net/hyperv/Makefile
+++ b/drivers/net/hyperv/Makefile
@@ -2,4 +2,5 @@
obj-$(CONFIG_HYPERV_NET) += hv_netvsc.o
hv_netvsc-y := netvsc_drv.o netvsc.o rndis_filter.o netvsc_trace.o netvsc_bpf.o
+hv_netvsc-$(CONFIG_HYPERV_NET_KUNIT_TEST) += netvsc_completion_test.o
hv_netvsc-$(CONFIG_HYPERV_NET_KUNIT_TEST) += rndis_request_test.o
diff --git a/drivers/net/hyperv/hyperv_net.h b/drivers/net/hyperv/hyperv_net.h
index 6492e7e93ded..be3a63a168cf 100644
--- a/drivers/net/hyperv/hyperv_net.h
+++ b/drivers/net/hyperv/hyperv_net.h
@@ -229,6 +229,21 @@ int rndis_build_page_buffers(const void *data, u32 len,
struct hv_page_buffer *page_bufs,
u32 *page_buf_cnt);
+void netvsc_send_acct(struct net_device *ndev,
+ struct netvsc_device *net_device,
+ struct vmbus_channel *channel,
+ u16 q_idx);
+void netvsc_send_tx_complete(struct net_device *ndev,
+ struct netvsc_device *net_device,
+ struct vmbus_channel *channel,
+ const struct vmpacket_descriptor *desc,
+ int budget);
+void netvsc_send_completion(struct net_device *ndev,
+ struct netvsc_device *net_device,
+ struct vmbus_channel *incoming_channel,
+ const struct vmpacket_descriptor *desc,
+ int budget);
+
extern u32 netvsc_ring_bytes;
int netvsc_workqueue_init(void);
diff --git a/drivers/net/hyperv/netvsc.c b/drivers/net/hyperv/netvsc.c
index fd8aa7a3dcb3..0d017f836b9e 100644
--- a/drivers/net/hyperv/netvsc.c
+++ b/drivers/net/hyperv/netvsc.c
@@ -764,20 +764,45 @@ static inline void netvsc_free_send_slot(struct netvsc_device *net_device,
sync_change_bit(index, net_device->send_section_map);
}
-static void netvsc_send_tx_complete(struct net_device *ndev,
- struct netvsc_device *net_device,
- struct vmbus_channel *channel,
- const struct vmpacket_descriptor *desc,
- int budget)
+void netvsc_send_acct(struct net_device *ndev,
+ struct netvsc_device *net_device,
+ struct vmbus_channel *channel,
+ u16 q_idx)
+{
+ struct net_device_context *ndev_ctx = netdev_priv(ndev);
+ int queue_sends;
+
+ queue_sends =
+ atomic_dec_return(&net_device->chan_table[q_idx].queue_sends);
+
+ if (unlikely(net_device->destroy)) {
+ if (queue_sends == 0)
+ wake_up(&net_device->wait_drain);
+ } else {
+ struct netdev_queue *txq = netdev_get_tx_queue(ndev, q_idx);
+
+ if (netif_tx_queue_stopped(txq) && !net_device->tx_disable &&
+ (hv_get_avail_to_write_percent(&channel->outbound) >
+ RING_AVAIL_PERCENT_HIWATER || queue_sends < 1)) {
+ netif_tx_wake_queue(txq);
+ ndev_ctx->eth_stats.wake_queue++;
+ }
+ }
+}
+
+void netvsc_send_tx_complete(struct net_device *ndev,
+ struct netvsc_device *net_device,
+ struct vmbus_channel *channel,
+ const struct vmpacket_descriptor *desc,
+ int budget)
{
struct net_device_context *ndev_ctx = netdev_priv(ndev);
struct sk_buff *skb;
u16 q_idx = 0;
- int queue_sends;
u64 cmd_rqst;
cmd_rqst = channel->request_addr_callback(channel, desc->trans_id);
- if (cmd_rqst == VMBUS_RQST_ERROR) {
+ if (cmd_rqst == VMBUS_RQST_ERROR || cmd_rqst == VMBUS_NO_RQSTOR) {
netdev_err(ndev, "Invalid transaction ID %llx\n", desc->trans_id);
return;
}
@@ -806,29 +831,14 @@ static void netvsc_send_tx_complete(struct net_device *ndev,
napi_consume_skb(skb, budget);
}
- queue_sends =
- atomic_dec_return(&net_device->chan_table[q_idx].queue_sends);
-
- if (unlikely(net_device->destroy)) {
- if (queue_sends == 0)
- wake_up(&net_device->wait_drain);
- } else {
- struct netdev_queue *txq = netdev_get_tx_queue(ndev, q_idx);
-
- if (netif_tx_queue_stopped(txq) && !net_device->tx_disable &&
- (hv_get_avail_to_write_percent(&channel->outbound) >
- RING_AVAIL_PERCENT_HIWATER || queue_sends < 1)) {
- netif_tx_wake_queue(txq);
- ndev_ctx->eth_stats.wake_queue++;
- }
- }
+ netvsc_send_acct(ndev, net_device, channel, q_idx);
}
-static void netvsc_send_completion(struct net_device *ndev,
- struct netvsc_device *net_device,
- struct vmbus_channel *incoming_channel,
- const struct vmpacket_descriptor *desc,
- int budget)
+void netvsc_send_completion(struct net_device *ndev,
+ struct netvsc_device *net_device,
+ struct vmbus_channel *incoming_channel,
+ const struct vmpacket_descriptor *desc,
+ int budget)
{
const struct nvsp_message *nvsp_packet;
u32 msglen = hv_pkt_datalen(desc);
@@ -840,11 +850,24 @@ static void netvsc_send_completion(struct net_device *ndev,
if (!msglen) {
cmd_rqst = incoming_channel->request_addr_callback(incoming_channel,
desc->trans_id);
- if (cmd_rqst == VMBUS_RQST_ERROR) {
+ if (cmd_rqst == VMBUS_RQST_ERROR || cmd_rqst == VMBUS_NO_RQSTOR) {
netdev_err(ndev, "Invalid transaction ID %llx\n", desc->trans_id);
return;
}
+ /*
+ * netvsc_send_pkt() registers (ulong)skb as the request
+ * address. Control RNDIS sends carry no skb, so the
+ * registered address is NULL and there is no nvsp_message
+ * to inspect. The request id is consumed above, so this
+ * completion owns the send accounting -- the same thing
+ * netvsc_send_tx_complete() does when it sees a NULL skb.
+ */
+ if (!cmd_rqst) {
+ netvsc_send_acct(ndev, net_device, incoming_channel, 0);
+ return;
+ }
+
pkt_rqst = (struct nvsp_message *)(uintptr_t)cmd_rqst;
switch (pkt_rqst->hdr.msg_type) {
case NVSP_MSG4_TYPE_SWITCH_DATA_PATH:
diff --git a/drivers/net/hyperv/netvsc_completion_test.c b/drivers/net/hyperv/netvsc_completion_test.c
new file mode 100644
index 000000000000..2614b1ece232
--- /dev/null
+++ b/drivers/net/hyperv/netvsc_completion_test.c
@@ -0,0 +1,439 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * KUnit tests for empty and skb-backed netvsc completion handling.
+ *
+ * Built into the hv_netvsc object rather than a separate module so the
+ * cases can reach the internal helpers declared in hyperv_net.h
+ * without exporting them.
+ */
+#include <kunit/test.h>
+#include <linux/etherdevice.h>
+#include <linux/hyperv.h>
+#include <linux/netdevice.h>
+#include <linux/skbuff.h>
+#include <linux/slab.h>
+#include <linux/wait.h>
+
+#include "hyperv_net.h"
+
+/*
+ * Scripted requestor. request_addr_callback() only receives the channel
+ * and the transaction id, so the fixture is hung off a file-scope
+ * pointer. KUnit runs suite cases serially; one active fixture is
+ * enough and keeps the callback signature untouched.
+ */
+struct netvsc_completion_fixture {
+ struct net_device *ndev;
+ struct netvsc_device *nvdev;
+ struct vmbus_channel *channel;
+ /* Address returned on first use, then VMBUS_NO_RQSTOR: the id is
+ * consumed exactly once, as vmbus_request_addr_match() does.
+ */
+ u64 once_addr;
+ unsigned int calls;
+ u64 seen_ids[8];
+ unsigned int skb_frees;
+};
+
+static struct netvsc_completion_fixture *active_fx;
+
+static u64 test_request_addr(struct vmbus_channel *channel, u64 rqst_id)
+{
+ struct netvsc_completion_fixture *fx = active_fx;
+
+ if (fx->calls < ARRAY_SIZE(fx->seen_ids))
+ fx->seen_ids[fx->calls] = rqst_id;
+ fx->calls++;
+
+ if (fx->calls == 1)
+ return fx->once_addr;
+ return VMBUS_NO_RQSTOR;
+}
+
+/*
+ * A completion carrying no payload: hv_pkt_datalen() == 0. The
+ * empty-completion path looks the request up by trans_id and reads its
+ * message type, so the request itself is a separate object the fixture
+ * hands back through request_addr_callback().
+ */
+struct netvsc_empty_desc {
+ struct vmpacket_descriptor desc;
+} __packed;
+
+static void make_empty_desc(struct netvsc_empty_desc *pkt, u64 trans_id)
+{
+ memset(pkt, 0, sizeof(*pkt));
+ pkt->desc.offset8 = sizeof(pkt->desc) / 8;
+ pkt->desc.len8 = pkt->desc.offset8;
+ pkt->desc.trans_id = trans_id;
+}
+
+/*
+ * The production path never dereferences a net_device queue while
+ * tx_disable is set: netif_tx_queue_stopped() && !tx_disable
+ * short-circuits before hv_get_avail_to_write_percent() reads the
+ * ring. Set it here so a test needs no live outbound ring.
+ */
+static int netvsc_completion_fixture_init(struct netvsc_completion_fixture *fx)
+{
+ fx->ndev = alloc_netdev_mqs(sizeof(struct net_device_context),
+ "hvcompl%d", NET_NAME_UNKNOWN,
+ ether_setup, 2, 2);
+ if (!fx->ndev)
+ return -ENOMEM;
+
+ fx->nvdev = kzalloc_obj(*fx->nvdev, GFP_KERNEL);
+ if (!fx->nvdev) {
+ free_netdev(fx->ndev);
+ return -ENOMEM;
+ }
+
+ fx->channel = kzalloc_obj(*fx->channel, GFP_KERNEL);
+ if (!fx->channel) {
+ kfree(fx->nvdev);
+ free_netdev(fx->ndev);
+ return -ENOMEM;
+ }
+
+ init_waitqueue_head(&fx->nvdev->wait_drain);
+ init_completion(&fx->nvdev->channel_init_wait);
+ fx->nvdev->tx_disable = true;
+ fx->channel->request_addr_callback = test_request_addr;
+ fx->once_addr = 0;
+ fx->calls = 0;
+ memset(fx->seen_ids, 0, sizeof(fx->seen_ids));
+ active_fx = fx;
+ return 0;
+}
+
+static void netvsc_completion_fixture_exit(struct netvsc_completion_fixture *fx)
+{
+ active_fx = NULL;
+ kfree(fx->channel);
+ kfree(fx->nvdev);
+ free_netdev(fx->ndev);
+}
+
+static int netvsc_send_sends(struct netvsc_completion_fixture *fx, u16 q_idx)
+{
+ return atomic_read(&fx->nvdev->chan_table[q_idx].queue_sends);
+}
+
+/* Empty completion, request address NULL: owns the send accounting. */
+static void netvsc_completion_null_address_acct_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 3);
+ make_empty_desc(&pkt, 0x11);
+ /* once_addr = 0 is the control-path NULL skb case. */
+ fx.once_addr = 0;
+
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 2);
+ KUNIT_EXPECT_FALSE(test, completion_done(&fx.nvdev->channel_init_wait));
+ KUNIT_EXPECT_EQ(test, fx.calls, 1U);
+ KUNIT_EXPECT_EQ(test, fx.seen_ids[0], 0x11U);
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+/* VMBUS_RQST_ERROR is rejected: no accounting, no completion. */
+static void netvsc_completion_rqst_error_sentinel_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 3);
+ make_empty_desc(&pkt, 0x22);
+ fx.once_addr = VMBUS_RQST_ERROR;
+
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 3);
+ KUNIT_EXPECT_FALSE(test, completion_done(&fx.nvdev->channel_init_wait));
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+/* VMBUS_NO_RQSTOR is rejected the same way. */
+static void netvsc_completion_no_rqstor_sentinel_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 3);
+ make_empty_desc(&pkt, 0x33);
+ fx.once_addr = VMBUS_NO_RQSTOR;
+
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 3);
+ KUNIT_EXPECT_FALSE(test, completion_done(&fx.nvdev->channel_init_wait));
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+/*
+ * A real request address whose message type is SWITCH_DATA_PATH
+ * completes the channel-init wait. It is not send accounting: the
+ * control request was not a send.
+ */
+static void netvsc_completion_switch_data_path_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+ struct nvsp_message req = {};
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 3);
+ make_empty_desc(&pkt, 0x44);
+ req.hdr.msg_type = NVSP_MSG4_TYPE_SWITCH_DATA_PATH;
+ fx.once_addr = (u64)(unsigned long)&req;
+
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+
+ KUNIT_EXPECT_TRUE(test, completion_done(&fx.nvdev->channel_init_wait));
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 3);
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+/* An unexpected message type takes neither the complete nor the acct path. */
+static void netvsc_completion_unknown_msg_type_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+ struct nvsp_message req = {};
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 3);
+ make_empty_desc(&pkt, 0x55);
+ req.hdr.msg_type = 0xdead;
+ fx.once_addr = (u64)(unsigned long)&req;
+
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+
+ KUNIT_EXPECT_FALSE(test, completion_done(&fx.nvdev->channel_init_wait));
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 3);
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+/*
+ * A repeated completion for an already-consumed transaction id must not
+ * account twice. The requestor returns VMBUS_NO_RQSTOR on the second
+ * use and the empty-completion path rejects it.
+ */
+static void netvsc_completion_duplicate_completion_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 3);
+ make_empty_desc(&pkt, 0x66);
+ fx.once_addr = 0;
+
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 2);
+ KUNIT_EXPECT_EQ(test, fx.calls, 1U);
+
+ /* Second completion for the same id: already consumed. */
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 2);
+ KUNIT_EXPECT_EQ(test, fx.calls, 2U);
+ KUNIT_EXPECT_EQ(test, fx.seen_ids[0], 0x66U);
+ KUNIT_EXPECT_EQ(test, fx.seen_ids[1], 0x66U);
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+/* N completions decrement the queue exactly N times. */
+static void netvsc_completion_single_decrement_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+ int i;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 5);
+
+ for (i = 0; i < 5; i++) {
+ make_empty_desc(&pkt, 0x100 + i);
+ fx.once_addr = 0;
+ fx.calls = 0;
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel,
+ &pkt.desc, 0);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 4 - i);
+ KUNIT_EXPECT_EQ(test, fx.calls, 1U);
+ }
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 0);
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+static int netvsc_test_wake(struct wait_queue_entry *wq_entry,
+ unsigned int mode, int flags, void *key)
+{
+ unsigned int *woken = wq_entry->private;
+
+ (*woken)++;
+ return 0;
+}
+
+/* destroy + a decrement that reaches zero wakes the drain wait. */
+static void netvsc_send_acct_drain_wake_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct wait_queue_entry entry;
+ unsigned int woken = 0;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ fx.nvdev->destroy = true;
+
+ memset(&entry, 0, sizeof(entry));
+ entry.func = netvsc_test_wake;
+ entry.private = &woken;
+ add_wait_queue(&fx.nvdev->wait_drain, &entry);
+
+ /* Non-zero remainder: the drain wait is not woken. */
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 2);
+ netvsc_send_acct(fx.ndev, fx.nvdev, fx.channel, 0);
+ KUNIT_EXPECT_EQ(test, woken, 0U);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 1);
+
+ /* Exactly zero: the drain wait is woken once. */
+ netvsc_send_acct(fx.ndev, fx.nvdev, fx.channel, 0);
+ KUNIT_EXPECT_EQ(test, woken, 1U);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 0);
+
+ remove_wait_queue(&fx.nvdev->wait_drain, &entry);
+ netvsc_completion_fixture_exit(&fx);
+}
+
+/*
+ * netvsc_send_tx_complete() applies the same sentinel rejection and
+ * still accounts a NULL skb.
+ */
+static void netvsc_tx_complete_null_skb_acct_test(struct kunit *test)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct netvsc_empty_desc pkt;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 3);
+ make_empty_desc(&pkt, 0x77);
+ fx.once_addr = 0;
+
+ netvsc_send_tx_complete(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 2);
+
+ /* The sentinel is rejected before any accounting. */
+ make_empty_desc(&pkt, 0x78);
+ fx.once_addr = VMBUS_RQST_ERROR;
+ fx.calls = 0;
+ netvsc_send_tx_complete(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 2);
+
+ netvsc_completion_fixture_exit(&fx);
+}
+
+static void netvsc_test_skb_destructor(struct sk_buff *skb)
+{
+ active_fx->skb_frees++;
+}
+
+static void netvsc_tx_complete_skb(struct kunit *test, u32 status)
+{
+ struct netvsc_completion_fixture fx = {};
+ struct {
+ struct vmpacket_descriptor desc;
+ struct nvsp_message message;
+ } pkt = {};
+ unsigned long send_slots = BIT(3);
+ struct hv_netvsc_packet *packet;
+ struct netvsc_stats_tx *stats;
+ struct sk_buff *skb;
+
+ KUNIT_ASSERT_EQ(test, netvsc_completion_fixture_init(&fx), 0);
+ skb = alloc_skb(64, GFP_KERNEL);
+ if (!skb) {
+ KUNIT_FAIL(test, "skb allocation failed");
+ goto out;
+ }
+ skb->destructor = netvsc_test_skb_destructor;
+ packet = (struct hv_netvsc_packet *)skb->cb;
+ memset(packet, 0, sizeof(*packet));
+ packet->send_buf_index = 3;
+ packet->q_idx = 1;
+ packet->total_packets = 2;
+ packet->total_bytes = 64;
+ fx.nvdev->send_section_map = &send_slots;
+ stats = &fx.nvdev->chan_table[1].tx_stats;
+ u64_stats_init(&stats->syncp);
+ atomic_set(&fx.nvdev->chan_table[0].queue_sends, 7);
+ atomic_set(&fx.nvdev->chan_table[1].queue_sends, 3);
+ fx.once_addr = (u64)(unsigned long)skb;
+ pkt.desc.offset8 = sizeof(pkt.desc) / 8;
+ pkt.desc.len8 = DIV_ROUND_UP(sizeof(pkt), 8);
+ pkt.desc.trans_id = 0x79;
+ pkt.message.hdr.msg_type = NVSP_MSG1_TYPE_SEND_RNDIS_PKT_COMPLETE;
+ pkt.message.msg.v1_msg.send_rndis_pkt_complete.status = status;
+
+ /* Budget zero consumes the real skb synchronously, outside NAPI. */
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+ KUNIT_EXPECT_EQ(test, send_slots, 0UL);
+ KUNIT_EXPECT_EQ(test, fx.skb_frees, 1U);
+ KUNIT_EXPECT_EQ(test, stats->packets, 2ULL);
+ KUNIT_EXPECT_EQ(test, stats->bytes, 64ULL);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 0), 7);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 1), 2);
+
+ /* The consumed transaction cannot free or account the skb twice. */
+ netvsc_send_completion(fx.ndev, fx.nvdev, fx.channel, &pkt.desc, 0);
+ KUNIT_EXPECT_EQ(test, send_slots, 0UL);
+ KUNIT_EXPECT_EQ(test, fx.skb_frees, 1U);
+ KUNIT_EXPECT_EQ(test, stats->packets, 2ULL);
+ KUNIT_EXPECT_EQ(test, stats->bytes, 64ULL);
+ KUNIT_EXPECT_EQ(test, netvsc_send_sends(&fx, 1), 2);
+out:
+ netvsc_completion_fixture_exit(&fx);
+}
+
+static void netvsc_tx_complete_skb_success_test(struct kunit *test)
+{
+ netvsc_tx_complete_skb(test, NVSP_STAT_SUCCESS);
+}
+
+static void netvsc_tx_complete_skb_error_test(struct kunit *test)
+{
+ netvsc_tx_complete_skb(test, NVSP_STAT_FAIL);
+}
+
+static struct kunit_case netvsc_completion_test_cases[] = {
+ KUNIT_CASE(netvsc_completion_null_address_acct_test),
+ KUNIT_CASE(netvsc_completion_rqst_error_sentinel_test),
+ KUNIT_CASE(netvsc_completion_no_rqstor_sentinel_test),
+ KUNIT_CASE(netvsc_completion_switch_data_path_test),
+ KUNIT_CASE(netvsc_completion_unknown_msg_type_test),
+ KUNIT_CASE(netvsc_completion_duplicate_completion_test),
+ KUNIT_CASE(netvsc_completion_single_decrement_test),
+ KUNIT_CASE(netvsc_send_acct_drain_wake_test),
+ KUNIT_CASE(netvsc_tx_complete_null_skb_acct_test),
+ KUNIT_CASE(netvsc_tx_complete_skb_success_test),
+ KUNIT_CASE(netvsc_tx_complete_skb_error_test),
+ {}
+};
+
+static struct kunit_suite netvsc_completion_test_suite = {
+ .name = "hyperv-netvsc-completion",
+ .test_cases = netvsc_completion_test_cases,
+};
+
+kunit_test_suite(netvsc_completion_test_suite);
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH v2 14/14] hv: netvsc: use kvzalloc for device state
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
` (12 preceding siblings ...)
2026-10-07 19:07 ` [PATCH v2 13/14] hv: netvsc: handle a NULL request address on empty completions Emerson Busson
@ 2026-10-07 19:07 ` Emerson Busson
13 siblings, 0 replies; 15+ messages in thread
From: Emerson Busson @ 2026-10-07 19:07 UTC (permalink / raw)
To: mhklinux
Cc: kys, haiyangz, wei.liu, decui, andrew+netdev, davem, edumazet,
kuba, pabeni, gregkh, linux-kernel, linux-hyperv, netdev
alloc_net_device() allocates the fixed-channel netvsc_device with
kzalloc_obj(). The runtime fragmentation drill reaches this allocation
from netvsc_device_add(): windows-2025 reports an order-7 allocation
failure with about 99 MiB free, while windows-latest stalls at channel
open until the guest timeout.
This object is guest-private control-plane metadata, not DMA backing.
Use kvzalloc_obj() so the allocation can fall back to vmalloc when a
high-order kmalloc cannot be satisfied. Zero initialization is preserved.
The RCU workqueue cleanup uses kvfree() to release either backing.
Rollback trigger: revert if channel setup fails with vmalloc capacity
available or teardown leaves netvsc_device allocations outstanding.
Signed-off-by: Emerson Busson <emersonbusson@gmail.com>
---
drivers/net/hyperv/netvsc.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/drivers/net/hyperv/netvsc.c b/drivers/net/hyperv/netvsc.c
index 0d017f836b9e..8ea7c993baac 100644
--- a/drivers/net/hyperv/netvsc.c
+++ b/drivers/net/hyperv/netvsc.c
@@ -144,7 +144,7 @@ static void __free_netvsc_device(struct netvsc_device *nvdev)
vfree(nvdev->chan_table[i].mrc.slots);
}
- kfree(nvdev);
+ kvfree(nvdev);
}
static void free_netvsc_device(struct work_struct *w)
@@ -171,7 +171,7 @@ static struct netvsc_device *alloc_net_device(void)
{
struct netvsc_device *net_device;
- net_device = kzalloc_obj(struct netvsc_device);
+ net_device = kvzalloc_obj(struct netvsc_device);
if (!net_device)
return NULL;
--
2.43.0
^ permalink raw reply [flat|nested] 15+ messages in thread
end of thread, other threads:[~2026-10-07 19:09 UTC | newest]
Thread overview: 15+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-07 19:07 [PATCH v2 0/14] hv: vmbus: make rings and host-visible buffers survive buddy fragmentation Emerson Busson
2026-10-07 19:07 ` [PATCH v2 01/14] hv: vmbus: convert ring backing through the chunk allocator Emerson Busson
2026-10-07 19:07 ` [PATCH v2 02/14] hv: vmbus: validate chunk buffer allocation and cleanup Emerson Busson
2026-10-07 19:07 ` [PATCH v2 03/14] uio: hv_generic: describe buffers for owned allocation Emerson Busson
2026-10-07 19:07 ` [PATCH v2 04/14] hv: vmbus: add KUnit tests for GPADL post failure injection Emerson Busson
2026-10-07 19:07 ` [PATCH v2 05/14] hv: vmbus: add KUnit test for order-zero allocation fallback Emerson Busson
2026-10-07 19:07 ` [PATCH v2 06/14] hv: vmbus: cover all shared-page policy combinations Emerson Busson
2026-10-07 19:07 ` [PATCH v2 07/14] hv: vmbus: distinguish host rescind from local channel unload Emerson Busson
2026-10-07 19:07 ` [PATCH v2 08/14] hv: vmbus: retain backing until ownership and references clear Emerson Busson
2026-10-07 19:07 ` [PATCH v2 09/14] hv: use owned VMBus buffers in NetVSC and UIO Emerson Busson
2026-10-07 19:07 ` [PATCH v2 10/14] hv: vmbus: pin buffer pages across UIO mmap to close the reclaim race Emerson Busson
2026-10-07 19:07 ` [PATCH v2 11/14] hv: vmbus: vmalloc requestor metadata Emerson Busson
2026-10-07 19:07 ` [PATCH v2 12/14] hv: netvsc: allocate RNDIS request descriptors with kvzalloc_obj() Emerson Busson
2026-10-07 19:07 ` [PATCH v2 13/14] hv: netvsc: handle a NULL request address on empty completions Emerson Busson
2026-10-07 19:07 ` [PATCH v2 14/14] hv: netvsc: use kvzalloc for device state Emerson Busson
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®