* [PATCH net-next 0/2] vsock: receive queue hint for io_uring
@ 2026-09-20 19:47 David Carlier
2026-09-20 19:47 ` [PATCH net-next 1/2] vsock: report pending receive data to io_uring David Carlier
2026-09-20 19:47 ` [PATCH net-next 2/2] vsock/test: cover receive queue hints David Carlier
0 siblings, 2 replies; 3+ messages in thread
From: David Carlier @ 2026-09-20 19:47 UTC (permalink / raw)
To: netdev
Cc: sgarzare, davem, edumazet, kuba, pabeni, horms, virtualization,
linux-kernel, David Carlier
AF_VSOCK stream receives never fill msghdr.msg_inq. io_uring initialises the
hint to -1 and treats that as "unknown", so it cannot set
IORING_CQE_F_SOCK_NONEMPTY, and it speculatively retries a multishot receive
after every completion, including when the queue is already empty.
Patch 1 fills the hint at the common exit of __vsock_connectible_recvmsg(),
reusing vsock_stream_has_data(), the transport callback that SIOCINQ already
uses. Following tcp_inq_hint(), it reports a terminal hint of 1 once the
connection is finished, so the caller still performs the receive that
observes EOF. Patch 2 adds selftest coverage.
Numbers come from a ping-pong over vsock loopback using io_uring multishot
receive, 4 KiB messages, the queue drained between messages, with kernel
receives counted by a kprobe. The baseline is the same tree with patch 1
reverted:
entries per delivered message 1.97 -> 1.00 (5000 messages)
receiver CPU time, median 1.640s -> 1.587s (-3.2%)
receiver wall time, median 3.260s -> 3.183s (-2.4%)
Timing is 25 runs of 50000 messages per kernel; CPU gives Mann-Whitney
p=0.006, wall time does not separate from noise (p=0.12), as expected for a
latency-bound workload.
The hint is gated to SOCK_STREAM; seqpacket is untouched. SO_INQ and SCM_INQ
are deliberately left out: they need their own per-socket state and
ancillary-message semantics, and a byte-count contract would have to account
for transports reporting less than the whole queue. That can follow.
Testing: built and run under virtme-ng on loopback, where the new cases pass
and vsock_test still passes 41 of 41 with a clean dmesg. Other transports
were reviewed but not tested. Hyper-V reports the current packet rather than
the whole queue and can return -EIO; an undercount only understates what is
readable, and a negative result leaves the caller's hint untouched.
David Carlier (2):
vsock: report pending receive data to io_uring
vsock/test: cover receive queue hints
net/vmw_vsock/af_vsock.c | 39 +++
tools/testing/vsock/vsock_uring_test.c | 367 +++++++++++++++++++++++++
2 files changed, 406 insertions(+)
base-commit: cea851523034d83dc61d882b1a6a3ad273480f7b
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH net-next 1/2] vsock: report pending receive data to io_uring
2026-09-20 19:47 [PATCH net-next 0/2] vsock: receive queue hint for io_uring David Carlier
@ 2026-09-20 19:47 ` David Carlier
2026-09-20 19:47 ` [PATCH net-next 2/2] vsock/test: cover receive queue hints David Carlier
1 sibling, 0 replies; 3+ messages in thread
From: David Carlier @ 2026-09-20 19:47 UTC (permalink / raw)
To: netdev
Cc: sgarzare, davem, edumazet, kuba, pabeni, horms, virtualization,
linux-kernel, David Carlier
AF_VSOCK stream receives never fill msghdr.msg_inq, so io_uring cannot set
IORING_CQE_F_SOCK_NONEMPTY and retries a multishot receive even after the
queue has been drained.
Fill the hint at the common receive exit using the transport callback that
SIOCINQ already uses, and report 1 once the connection is finished so the
caller performs the receive which observes EOF, as TCP does after a FIN.
A vsock loopback ping-pong with io_uring multishot receive drops entries
into __vsock_connectible_recvmsg from 1.97 to 1.00 per delivered message,
and receiver CPU time by about 3% (25 runs of 50000 messages, p=0.006).
Signed-off-by: David Carlier <devnexen@gmail.com>
---
net/vmw_vsock/af_vsock.c | 39 +++++++++++++++++++++++++++++++++++++++
1 file changed, 39 insertions(+)
diff --git a/net/vmw_vsock/af_vsock.c b/net/vmw_vsock/af_vsock.c
index f840498b58af..63309708c916 100644
--- a/net/vmw_vsock/af_vsock.c
+++ b/net/vmw_vsock/af_vsock.c
@@ -2543,6 +2543,38 @@ static int __vsock_seqpacket_recvmsg(struct sock *sk, struct msghdr *msg,
return err;
}
+/* Bytes a following receive can consume, 1 if it would only see EOF, or a
+ * negative value to leave the caller's hint alone.
+ *
+ * Called under the socket lock after a nonnegative stream receive, so a NULL
+ * transport implies SOCK_DONE.
+ */
+static int vsock_inq_hint(struct sock *sk)
+{
+ struct vsock_sock *vsk = vsock_sk(sk);
+ s64 data;
+
+ if ((sk->sk_shutdown & RCV_SHUTDOWN) ||
+ (sock_flag(sk, SOCK_DONE) && sk->sk_state != TCP_ESTABLISHED))
+ return 1;
+
+ if (!vsk->transport)
+ return 1;
+
+ data = vsock_stream_has_data(vsk);
+ if (data < 0)
+ return -1;
+ if (data > 0)
+ return (int)min_t(s64, data, INT_MAX);
+
+ /* Empty but finished: keep the caller reading so it sees EOF. */
+ if (sock_flag(sk, SOCK_DONE) ||
+ (READ_ONCE(vsk->peer_shutdown) & SEND_SHUTDOWN))
+ return 1;
+
+ return 0;
+}
+
int
__vsock_connectible_recvmsg(struct socket *sock, struct msghdr *msg, size_t len,
int flags)
@@ -2606,6 +2638,13 @@ __vsock_connectible_recvmsg(struct socket *sock, struct msghdr *msg, size_t len,
err = __vsock_seqpacket_recvmsg(sk, msg, len, flags);
out:
+ if (msg->msg_get_inq && err >= 0 && sk->sk_type == SOCK_STREAM) {
+ int inq = vsock_inq_hint(sk);
+
+ if (inq >= 0)
+ msg->msg_inq = inq;
+ }
+
release_sock(sk);
return err;
}
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH net-next 2/2] vsock/test: cover receive queue hints
2026-09-20 19:47 [PATCH net-next 0/2] vsock: receive queue hint for io_uring David Carlier
2026-09-20 19:47 ` [PATCH net-next 1/2] vsock: report pending receive data to io_uring David Carlier
@ 2026-09-20 19:47 ` David Carlier
1 sibling, 0 replies; 3+ messages in thread
From: David Carlier @ 2026-09-20 19:47 UTC (permalink / raw)
To: netdev
Cc: sgarzare, davem, edumazet, kuba, pabeni, horms, virtualization,
linux-kernel, David Carlier
Add io_uring receive tests checking IORING_CQE_F_SOCK_NONEMPTY across a
partial receive, draining while the peer stays connected, EOF, a
nonblocking receive on an empty queue, a zero-length request, and a
multishot receive with provided buffers.
Signed-off-by: David Carlier <devnexen@gmail.com>
---
tools/testing/vsock/vsock_uring_test.c | 367 +++++++++++++++++++++++++
1 file changed, 367 insertions(+)
diff --git a/tools/testing/vsock/vsock_uring_test.c b/tools/testing/vsock/vsock_uring_test.c
index 5c3078969659..318e17bd28bc 100644
--- a/tools/testing/vsock/vsock_uring_test.c
+++ b/tools/testing/vsock/vsock_uring_test.c
@@ -13,7 +13,10 @@
#include <liburing.h>
#include <unistd.h>
#include <sys/mman.h>
+#include <sys/ioctl.h>
#include <linux/kernel.h>
+#include <linux/sockios.h>
+#include <errno.h>
#include <error.h>
#include "util.h"
@@ -28,6 +31,10 @@
#define VSOCK_TEST_DATA_MAX_IOV 3
+#define HINT_CHUNK_SIZE 4096
+#define HINT_BUF_GROUP 1
+#define HINT_BUF_ENTRIES 4
+
struct vsock_io_uring_test {
/* Number of valid elements in 'vecs'. */
int vecs_cnt;
@@ -211,6 +218,341 @@ void test_stream_uring_msg_zc_client(const struct test_opts *opts)
vsock_io_uring_client(opts, &test_data_array[i], true);
}
+struct uring_inq_ctx {
+ struct io_uring ring;
+ int fd;
+};
+
+static void inq_server_init(struct uring_inq_ctx *ctx,
+ const struct test_opts *opts)
+{
+ ctx->fd = vsock_stream_accept(VMADDR_CID_ANY, opts->peer_port, NULL);
+ if (ctx->fd < 0) {
+ perror("accept");
+ exit(EXIT_FAILURE);
+ }
+
+ if (io_uring_queue_init(RING_ENTRIES_NUM, &ctx->ring, 0))
+ error(1, errno, "io_uring_queue_init");
+}
+
+static void inq_server_exit(struct uring_inq_ctx *ctx)
+{
+ io_uring_queue_exit(&ctx->ring);
+ close(ctx->fd);
+}
+
+/* Submit a single receive and report both its result and its CQE flags. */
+static int inq_recv(struct uring_inq_ctx *ctx, void *buf, size_t len,
+ int flags, unsigned int *cflags)
+{
+ struct io_uring_sqe *sqe;
+ struct io_uring_cqe *cqe;
+ int res;
+
+ sqe = io_uring_get_sqe(&ctx->ring);
+ io_uring_prep_recv(sqe, ctx->fd, buf, len, flags);
+
+ if (io_uring_submit(&ctx->ring) != 1)
+ error(1, errno, "io_uring_submit");
+
+ if (io_uring_wait_cqe(&ctx->ring, &cqe))
+ error(1, errno, "io_uring_wait_cqe");
+
+ res = cqe->res;
+ *cflags = cqe->flags;
+ io_uring_cqe_seen(&ctx->ring, cqe);
+
+ return res;
+}
+
+static void expect_res(int res, int expected, const char *what)
+{
+ if (res != expected) {
+ fprintf(stderr, "%s: expected %d, got %d\n", what, expected,
+ res);
+ exit(EXIT_FAILURE);
+ }
+}
+
+static void expect_nonempty(unsigned int cflags, bool expected,
+ const char *what)
+{
+ bool nonempty = !!(cflags & IORING_CQE_F_SOCK_NONEMPTY);
+
+ if (nonempty != expected) {
+ fprintf(stderr, "%s: expected SOCK_NONEMPTY %d, got %d\n",
+ what, expected, nonempty);
+ exit(EXIT_FAILURE);
+ }
+}
+
+/* Wait until the whole payload is queued, so the hint is deterministic. */
+static void inq_wait_queued(int fd, int len)
+{
+ if (!vsock_ioctl_int(fd, SIOCINQ, len)) {
+ fprintf(stderr, "SIOCINQ not supported\n");
+ exit(EXIT_FAILURE);
+ }
+}
+
+static void inq_send_chunks(const struct test_opts *opts, int chunks)
+{
+ char buf[HINT_CHUNK_SIZE];
+ int fd, i;
+
+ fd = vsock_stream_connect(opts->peer_cid, opts->peer_port);
+ if (fd < 0) {
+ perror("connect");
+ exit(EXIT_FAILURE);
+ }
+
+ memset(buf, 0xa5, sizeof(buf));
+ for (i = 0; i < chunks; i++)
+ send_buf(fd, buf, sizeof(buf), 0, sizeof(buf));
+
+ control_writeln("SENT");
+ control_expectln("DONE");
+ close(fd);
+}
+
+static void test_stream_uring_inq_client(const struct test_opts *opts)
+{
+ inq_send_chunks(opts, 2);
+}
+
+static void test_stream_uring_inq_server(const struct test_opts *opts)
+{
+ char buf[HINT_CHUNK_SIZE];
+ struct uring_inq_ctx ctx;
+ unsigned int cflags;
+ int res;
+
+ inq_server_init(&ctx, opts);
+
+ control_expectln("SENT");
+ inq_wait_queued(ctx.fd, 2 * HINT_CHUNK_SIZE);
+
+ /* Data remains after this receive, so the flag must be set. */
+ res = inq_recv(&ctx, buf, sizeof(buf), 0, &cflags);
+ expect_res(res, HINT_CHUNK_SIZE, "partial receive");
+ expect_nonempty(cflags, true, "partial receive");
+
+ /* This receive drains the queue while the peer stays connected. */
+ res = inq_recv(&ctx, buf, sizeof(buf), 0, &cflags);
+ expect_res(res, HINT_CHUNK_SIZE, "draining receive");
+ expect_nonempty(cflags, false, "draining receive");
+
+ control_writeln("DONE");
+ inq_server_exit(&ctx);
+}
+
+static void test_stream_uring_inq_eof_client(const struct test_opts *opts)
+{
+ char buf[HINT_CHUNK_SIZE];
+ int fd;
+
+ fd = vsock_stream_connect(opts->peer_cid, opts->peer_port);
+ if (fd < 0) {
+ perror("connect");
+ exit(EXIT_FAILURE);
+ }
+
+ memset(buf, 0x5a, sizeof(buf));
+ send_buf(fd, buf, sizeof(buf), 0, sizeof(buf));
+ control_writeln("SENT");
+
+ control_expectln("DRAINED");
+ close(fd);
+ control_writeln("CLOSED");
+
+ control_expectln("DONE");
+}
+
+static void test_stream_uring_inq_eof_server(const struct test_opts *opts)
+{
+ char buf[HINT_CHUNK_SIZE];
+ struct uring_inq_ctx ctx;
+ unsigned int cflags;
+ int res;
+
+ inq_server_init(&ctx, opts);
+
+ control_expectln("SENT");
+ inq_wait_queued(ctx.fd, HINT_CHUNK_SIZE);
+
+ res = inq_recv(&ctx, buf, sizeof(buf), 0, &cflags);
+ expect_res(res, HINT_CHUNK_SIZE, "drain before EOF");
+ expect_nonempty(cflags, false, "drain before EOF");
+
+ control_writeln("DRAINED");
+ control_expectln("CLOSED");
+
+ /* The queue is empty and the peer is gone. The hint stays non-zero
+ * so that this receive happens and reports EOF, as TCP does after a
+ * FIN.
+ */
+ res = inq_recv(&ctx, buf, sizeof(buf), 0, &cflags);
+ expect_res(res, 0, "receive at EOF");
+ expect_nonempty(cflags, true, "receive at EOF");
+
+ control_writeln("DONE");
+ inq_server_exit(&ctx);
+}
+
+static void test_stream_uring_inq_empty_client(const struct test_opts *opts)
+{
+ int fd;
+
+ fd = vsock_stream_connect(opts->peer_cid, opts->peer_port);
+ if (fd < 0) {
+ perror("connect");
+ exit(EXIT_FAILURE);
+ }
+
+ control_writeln("READY");
+ control_expectln("DONE");
+ close(fd);
+}
+
+static void test_stream_uring_inq_empty_server(const struct test_opts *opts)
+{
+ char buf[HINT_CHUNK_SIZE];
+ struct uring_inq_ctx ctx;
+ unsigned int cflags;
+ int res;
+
+ inq_server_init(&ctx, opts);
+
+ control_expectln("READY");
+
+ /* A failed receive must not leave a stale positive hint. */
+ res = inq_recv(&ctx, buf, sizeof(buf), MSG_DONTWAIT, &cflags);
+ expect_res(res, -EAGAIN, "empty nonblocking receive");
+ expect_nonempty(cflags, false, "empty nonblocking receive");
+
+ control_writeln("DONE");
+ inq_server_exit(&ctx);
+}
+
+static void test_stream_uring_inq_zerolen_client(const struct test_opts *opts)
+{
+ inq_send_chunks(opts, 1);
+}
+
+static void test_stream_uring_inq_zerolen_server(const struct test_opts *opts)
+{
+ char buf[HINT_CHUNK_SIZE];
+ struct uring_inq_ctx ctx;
+ unsigned int cflags;
+ int res;
+
+ inq_server_init(&ctx, opts);
+
+ control_expectln("SENT");
+ inq_wait_queued(ctx.fd, HINT_CHUNK_SIZE);
+
+ /* A zero-length request is not an error and still describes the
+ * queue behind it.
+ */
+ res = inq_recv(&ctx, buf, 0, 0, &cflags);
+ expect_res(res, 0, "zero-length receive");
+ expect_nonempty(cflags, true, "zero-length receive");
+
+ control_writeln("DONE");
+ inq_server_exit(&ctx);
+}
+
+static void test_stream_uring_inq_mshot_client(const struct test_opts *opts)
+{
+ char buf[HINT_CHUNK_SIZE];
+ int fd;
+
+ fd = vsock_stream_connect(opts->peer_cid, opts->peer_port);
+ if (fd < 0) {
+ perror("connect");
+ exit(EXIT_FAILURE);
+ }
+
+ memset(buf, 0x3c, sizeof(buf));
+ send_buf(fd, buf, sizeof(buf), 0, sizeof(buf));
+ control_writeln("SENT");
+
+ control_expectln("DRAINED");
+ close(fd);
+ control_writeln("CLOSED");
+
+ control_expectln("DONE");
+}
+
+static void test_stream_uring_inq_mshot_server(const struct test_opts *opts)
+{
+ static char bufs[HINT_BUF_ENTRIES][HINT_CHUNK_SIZE];
+ struct io_uring_buf_ring *br;
+ struct uring_inq_ctx ctx;
+ struct io_uring_sqe *sqe;
+ struct io_uring_cqe *cqe;
+ int i, ret;
+
+ inq_server_init(&ctx, opts);
+
+ br = io_uring_setup_buf_ring(&ctx.ring, HINT_BUF_ENTRIES,
+ HINT_BUF_GROUP, 0, &ret);
+ if (!br) {
+ fprintf(stderr, "io_uring_setup_buf_ring: %d\n", ret);
+ exit(EXIT_FAILURE);
+ }
+
+ for (i = 0; i < HINT_BUF_ENTRIES; i++)
+ io_uring_buf_ring_add(br, bufs[i], HINT_CHUNK_SIZE, i,
+ io_uring_buf_ring_mask(HINT_BUF_ENTRIES),
+ i);
+ io_uring_buf_ring_advance(br, HINT_BUF_ENTRIES);
+
+ sqe = io_uring_get_sqe(&ctx.ring);
+ io_uring_prep_recv_multishot(sqe, ctx.fd, NULL, 0, 0);
+ sqe->flags |= IOSQE_BUFFER_SELECT;
+ sqe->buf_group = HINT_BUF_GROUP;
+
+ if (io_uring_submit(&ctx.ring) != 1)
+ error(1, errno, "io_uring_submit");
+
+ control_expectln("SENT");
+
+ /* The payload completion drains the queue and keeps the request
+ * armed, so the hint must report the queue as empty.
+ */
+ if (io_uring_wait_cqe(&ctx.ring, &cqe))
+ error(1, errno, "io_uring_wait_cqe");
+
+ expect_res(cqe->res, HINT_CHUNK_SIZE, "multishot payload");
+ expect_nonempty(cqe->flags, false, "multishot payload");
+ if (!(cqe->flags & IORING_CQE_F_MORE)) {
+ fprintf(stderr, "multishot payload: request not rearmed\n");
+ exit(EXIT_FAILURE);
+ }
+ io_uring_cqe_seen(&ctx.ring, cqe);
+
+ control_writeln("DRAINED");
+ control_expectln("CLOSED");
+
+ /* EOF ends multishot regardless of the hint. */
+ if (io_uring_wait_cqe(&ctx.ring, &cqe))
+ error(1, errno, "io_uring_wait_cqe");
+
+ expect_res(cqe->res, 0, "multishot EOF");
+ if (cqe->flags & IORING_CQE_F_MORE) {
+ fprintf(stderr, "multishot EOF: request still armed\n");
+ exit(EXIT_FAILURE);
+ }
+ io_uring_cqe_seen(&ctx.ring, cqe);
+
+ control_writeln("DONE");
+ io_uring_free_buf_ring(&ctx.ring, br, HINT_BUF_ENTRIES,
+ HINT_BUF_GROUP);
+ inq_server_exit(&ctx);
+}
+
static struct test_case test_cases[] = {
{
.name = "SOCK_STREAM io_uring test",
@@ -222,6 +564,31 @@ static struct test_case test_cases[] = {
.run_server = test_stream_uring_msg_zc_server,
.run_client = test_stream_uring_msg_zc_client,
},
+ {
+ .name = "SOCK_STREAM io_uring receive queue hint",
+ .run_server = test_stream_uring_inq_server,
+ .run_client = test_stream_uring_inq_client,
+ },
+ {
+ .name = "SOCK_STREAM io_uring receive hint at EOF",
+ .run_server = test_stream_uring_inq_eof_server,
+ .run_client = test_stream_uring_inq_eof_client,
+ },
+ {
+ .name = "SOCK_STREAM io_uring receive hint on empty queue",
+ .run_server = test_stream_uring_inq_empty_server,
+ .run_client = test_stream_uring_inq_empty_client,
+ },
+ {
+ .name = "SOCK_STREAM io_uring receive hint zero-length",
+ .run_server = test_stream_uring_inq_zerolen_server,
+ .run_client = test_stream_uring_inq_zerolen_client,
+ },
+ {
+ .name = "SOCK_STREAM io_uring multishot receive hint",
+ .run_server = test_stream_uring_inq_mshot_server,
+ .run_client = test_stream_uring_inq_mshot_client,
+ },
{},
};
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-20 19:47 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-20 19:47 [PATCH net-next 0/2] vsock: receive queue hint for io_uring David Carlier
2026-09-20 19:47 ` [PATCH net-next 1/2] vsock: report pending receive data to io_uring David Carlier
2026-09-20 19:47 ` [PATCH net-next 2/2] vsock/test: cover receive queue hints David Carlier
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®