* [PATCH 0/5] vsock: reduce RX per-packet socket overhead
@ 2026-10-02 7:45 physicalmtea
2026-10-02 10:39 ` Stefano Garzarella
0 siblings, 1 reply; 3+ messages in thread
From: physicalmtea @ 2026-10-02 7:45 UTC (permalink / raw)
To: stefanha, sgarzare, mst
Cc: jasowangio, eperezma, xuanzhuo, davem, edumazet, kuba, pabeni,
horms, virtualization, kvm, netdev, linux-kernel
From: Jia Jia <physicalmtea@gmail.com>
This series reduces per-packet socket overhead in the virtio-vsock RX
worker.
Patch 1: separate socket lookup from locked packet processing.
Patch 2: keep one socket lock across a bounded run of established
STREAM/RW packets for the same native socket.
Patch 3: reuse the locked socket for later packets with the same
address tuple.
Patch 4: coalesce default write-space notifications while that
lock is held.
Patch 5: defer the default readable callback until after the batch
unlocks. Custom callbacks retain per-packet notification behavior.
Performance:
Tested on single-stream AF_VSOCK (1 reader on Guest CPU0, virtio-vsock IRQ
on Guest CPU1, fresh QEMU guest per state, 16 balanced AB/BA pairs per
payload).
End-to-end RX throughput (64K receiver buffer, SO_RCVLOWAT=1):
payload baseline RX (G/s) PATCH RX (G/s) throughput change (95% CI)
256B 0.5722 1.0286 +79.76% [+75.40%, +84.23%]
512B 1.1464 1.9779 +72.53% [+69.50%, +75.63%]
1K 2.2714 3.6368 +60.12% [+56.92%, +63.37%]
4K 3.9095 7.7160 +97.37% [+92.09%, +102.79%]
Fixed-syscall stress test (receiver buffer = sender write size; RX in G/s):
payload baseline RX PATCH RX throughput change (95% CI) recv() base/patch
256B 0.2885 0.6737 +133.50% [+123.58%, +143.86%] 1048577/1048577
512B 0.5631 1.2091 +114.71% [+104.97%, +124.92%] 1048577/1048577
1K 1.5361 2.9678 +93.20% [+80.15%, +107.20%] 1048577/1048577
4K 2.8378 6.1878 +118.05% [+109.15%, +127.33%] 538802/525645
Request-response latency (Ping-Pong RTT, 10,000 requests):
payload baseline mean RTT (us) PATCH mean RTT (us) change (95% CI)
256B 126.968 124.541 -1.91% [-3.35%, -0.45%]
4K 130.362 126.664 -2.84% [-5.12%, -0.50%]
Diagnostic Guest CPU1 system-wide cycles (fixed-64K receiver, with perf):
Buffer size baseline cycles/B PATCH cycles/B cycles change (95% CI)
256B 15.287 13.630 -10.84% [-12.51%, -9.14%]
512B 7.306 6.611 -9.51% [-10.92%, -8.08%]
1K 3.811 3.458 -9.25% [-10.96%, -7.50%]
4K 1.850 1.669 -9.75% [-11.78%, -7.67%]
Jia Jia (5):
vsock: split socket lookup from locked RX processing
vsock: amortize RX socket locking for stream packets
vsock: reuse same-flow socket lookup in RX batches
vsock: coalesce RX write-space notifications in lock batches
vsock: defer RX readable notifications until batch unlock
include/linux/virtio_vsock.h | 14 ++
include/net/af_vsock.h | 3 +
net/vmw_vsock/af_vsock.c | 2 +
net/vmw_vsock/virtio_transport.c | 48 ++++-
net/vmw_vsock/virtio_transport_common.c | 341 +++++++++++++++++++++++++++-----
5 files changed, 357 insertions(+), 51 deletions(-)
base-commit: f5f84daefcd92d7a630066635ecea1433ed5eac7
--
2.53.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH 0/5] vsock: reduce RX per-packet socket overhead
2026-10-02 7:45 [PATCH 0/5] vsock: reduce RX per-packet socket overhead physicalmtea
@ 2026-10-02 10:39 ` Stefano Garzarella
2026-10-02 13:22 ` Jia Jia
0 siblings, 1 reply; 3+ messages in thread
From: Stefano Garzarella @ 2026-10-02 10:39 UTC (permalink / raw)
To: physicalmtea
Cc: stefanha, mst, jasowangio, eperezma, xuanzhuo, davem, edumazet,
kuba, pabeni, horms, virtualization, kvm, netdev, linux-kernel
On Fri, Oct 02, 2026 at 07:45:46AM +0000, physicalmtea@gmail.com wrote:
>From: Jia Jia <physicalmtea@gmail.com>
>
>This series reduces per-packet socket overhead in the virtio-vsock RX
>worker.
I did a quick look but IMO this is hard to review, please try to keep
patches as small as possible, e.g. where you are introducing ctx, store
ctx->net in a net variable, so you don't need to touch the entire
function, etc.
Please check https://docs.kernel.org/process/coding-assistants.html
and also https://docs.kernel.org/process/submitting-patches.html
ans also https://docs.kernel.org/process/maintainer-netdev.html#git-trees-and-patch-flow
(I guess this is net-next material).
I don't think this series was sent properly, I tried
`b4 am 20261002074551.318789-1-physicalmtea@gmail.com` without success
and also looking at
https://lore.kernel.org/virtualization/20261002074551.318789-1-physicalmtea@gmail.com/
I can't see the full series, please check your workflow.
Also, is this for vsock core or just virtio-vsock? if it's just
virtio-vsock transports, please use "vsock/virtio" prefix.
>
>Patch 1: separate socket lookup from locked packet processing.
>
>Patch 2: keep one socket lock across a bounded run of established
>STREAM/RW packets for the same native socket.
>
>Patch 3: reuse the locked socket for later packets with the same
>address tuple.
>
>Patch 4: coalesce default write-space notifications while that
>lock is held.
>
>Patch 5: defer the default readable callback until after the batch
>unlocks. Custom callbacks retain per-packet notification behavior.
>
>Performance:
>
>Tested on single-stream AF_VSOCK (1 reader on Guest CPU0, virtio-vsock IRQ
>on Guest CPU1, fresh QEMU guest per state, 16 balanced AB/BA pairs per
>payload).
What the host is using? vhost-vsock?
>
>End-to-end RX throughput (64K receiver buffer, SO_RCVLOWAT=1):
>
> payload baseline RX (G/s) PATCH RX (G/s) throughput change (95% CI)
What G is? Gbit? Gbytes?
> 256B 0.5722 1.0286 +79.76% [+75.40%, +84.23%]
> 512B 1.1464 1.9779 +72.53% [+69.50%, +75.63%]
> 1K 2.2714 3.6368 +60.12% [+56.92%, +63.37%]
> 4K 3.9095 7.7160 +97.37% [+92.09%, +102.79%]
Why stop to 4k?
>
>Fixed-syscall stress test (receiver buffer = sender write size; RX in G/s):
>
> payload baseline RX PATCH RX throughput change (95% CI) recv() base/patch
> 256B 0.2885 0.6737 +133.50% [+123.58%, +143.86%] 1048577/1048577
> 512B 0.5631 1.2091 +114.71% [+104.97%, +124.92%] 1048577/1048577
> 1K 1.5361 2.9678 +93.20% [+80.15%, +107.20%] 1048577/1048577
> 4K 2.8378 6.1878 +118.05% [+109.15%, +127.33%] 538802/525645
>
>Request-response latency (Ping-Pong RTT, 10,000 requests):
What tool did you used for both throughput and latency?
>
> payload baseline mean RTT (us) PATCH mean RTT (us) change (95% CI)
> 256B 126.968 124.541 -1.91% [-3.35%, -0.45%]
> 4K 130.362 126.664 -2.84% [-5.12%, -0.50%]
Whys only this payload cases? For latency we should use even try small
payloads IMO like 64 bytes.
Thanks,
Stefano
>
>Diagnostic Guest CPU1 system-wide cycles (fixed-64K receiver, with perf):
>
> Buffer size baseline cycles/B PATCH cycles/B cycles change (95% CI)
> 256B 15.287 13.630 -10.84% [-12.51%, -9.14%]
> 512B 7.306 6.611 -9.51% [-10.92%, -8.08%]
> 1K 3.811 3.458 -9.25% [-10.96%, -7.50%]
> 4K 1.850 1.669 -9.75% [-11.78%, -7.67%]
>
>Jia Jia (5):
> vsock: split socket lookup from locked RX processing
> vsock: amortize RX socket locking for stream packets
> vsock: reuse same-flow socket lookup in RX batches
> vsock: coalesce RX write-space notifications in lock batches
> vsock: defer RX readable notifications until batch unlock
>
> include/linux/virtio_vsock.h | 14 ++
> include/net/af_vsock.h | 3 +
> net/vmw_vsock/af_vsock.c | 2 +
> net/vmw_vsock/virtio_transport.c | 48 ++++-
> net/vmw_vsock/virtio_transport_common.c | 341 +++++++++++++++++++++++++++-----
> 5 files changed, 357 insertions(+), 51 deletions(-)
>
>base-commit: f5f84daefcd92d7a630066635ecea1433ed5eac7
>--
>2.53.0
>
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH 0/5] vsock: reduce RX per-packet socket overhead
2026-10-02 10:39 ` Stefano Garzarella
@ 2026-10-02 13:22 ` Jia Jia
0 siblings, 0 replies; 3+ messages in thread
From: Jia Jia @ 2026-10-02 13:22 UTC (permalink / raw)
To: Stefano Garzarella
Cc: stefanha, mst, jasowangio, eperezma, virtualization, netdev,
linux-kernel
>
> On Fri, Oct 02, 2026 at 07:45:46AM +0000, physicalmtea@gmail.com wrote:
> >From: Jia Jia <physicalmtea@gmail.com>
> >
> >This series reduces per-packet socket overhead in the virtio-vsock RX
> >worker.
>
> I did a quick look but IMO this is hard to review, please try to keep
> patches as small as possible, e.g. where you are introducing ctx, store
> ctx->net in a net variable, so you don't need to touch the entire
> function, etc.
Thanks for the review and suggestions.
Will do. I will unpack ctx into local variables to minimize the diff.
>
> Please check https://docs.kernel.org/process/coding-assistants.html
> and also https://docs.kernel.org/process/submitting-patches.html
> ans also https://docs.kernel.org/process/maintainer-netdev.html#git-trees-and-patch-flow
> (I guess this is net-next material).
>
Thanks, I'll reread them. Targeting net-next for v2.
> I don't think this series was sent properly, I tried
> `b4 am 20261002074551.318789-1-physicalmtea@gmail.com` without success
> and also looking at
> https://lore.kernel.org/virtualization/20261002074551.318789-1-physicalmtea@gmail.com/
> I can't see the full series, please check your workflow.
>
Sorry for the broken thread. I will fix the git-send-email setup for v2.
> Also, is this for vsock core or just virtio-vsock? if it's just
> virtio-vsock transports, please use "vsock/virtio" prefix.
>
>
It is for virtio-vsock. Will use the "vsock/virtio:" prefix.
> >
> >Patch 1: separate socket lookup from locked packet processing.
> >
> >Patch 2: keep one socket lock across a bounded run of established
> >STREAM/RW packets for the same native socket.
> >
> >Patch 3: reuse the locked socket for later packets with the same
> >address tuple.
> >
> >Patch 4: coalesce default write-space notifications while that
> >lock is held.
> >
> >Patch 5: defer the default readable callback until after the batch
> >unlocks. Custom callbacks retain per-packet notification behavior.
> >
> >Performance:
> >
> >Tested on single-stream AF_VSOCK (1 reader on Guest CPU0, virtio-vsock IRQ
> >on Guest CPU1, fresh QEMU guest per state, 16 balanced AB/BA pairs per
> >payload).
>
> What the host is using? vhost-vsock?
>
Yes, vhost-vsock.
> >
> >End-to-end RX throughput (64K receiver buffer, SO_RCVLOWAT=1):
> >
> > payload baseline RX (G/s) PATCH RX (G/s) throughput change (95% CI)
>
> What G is? Gbit? Gbytes?
>
Gbit/s. Will fix the table headers.
> > 256B 0.5722 1.0286 +79.76% [+75.40%, +84.23%]
> > 512B 1.1464 1.9779 +72.53% [+69.50%, +75.63%]
> > 1K 2.2714 3.6368 +60.12% [+56.92%, +63.37%]
> > 4K 3.9095 7.7160 +97.37% [+92.09%, +102.79%]
>
> Why stop to 4k?
>
Stopped at 4K because the throughput gains were steadily increasing.
Will test up to 64K and include the results in v2.
> >
> >Fixed-syscall stress test (receiver buffer = sender write size; RX in G/s):
> >
> > payload baseline RX PATCH RX throughput change (95% CI) recv() base/patch
> > 256B 0.2885 0.6737 +133.50% [+123.58%, +143.86%] 1048577/1048577
> > 512B 0.5631 1.2091 +114.71% [+104.97%, +124.92%] 1048577/1048577
> > 1K 1.5361 2.9678 +93.20% [+80.15%, +107.20%] 1048577/1048577
> > 4K 2.8378 6.1878 +118.05% [+109.15%, +127.33%] 538802/525645
> >
> >Request-response latency (Ping-Pong RTT, 10,000 requests):
>
> What tool did you used for both throughput and latency?
>
Throughput: tools/testing/vsock/vsock_perf
Latency: custom userspace ping-pong RTT tool (CLOCK_MONOTONIC).
Will document both in the cover letter.
> >
> > payload baseline mean RTT (us) PATCH mean RTT (us) change (95% CI)
> > 256B 126.968 124.541 -1.91% [-3.35%, -0.45%]
> > 4K 130.362 126.664 -2.84% [-5.12%, -0.50%]
>
> Whys only this payload cases? For latency we should use even try small
> payloads IMO like 64 bytes.
>
Thanks for the suggestion. I'll add a 64B payload test for latency in v2.
> Thanks,
> Stefano
>
> >
> >Diagnostic Guest CPU1 system-wide cycles (fixed-64K receiver, with perf):
> >
> > Buffer size baseline cycles/B PATCH cycles/B cycles change (95% CI)
> > 256B 15.287 13.630 -10.84% [-12.51%, -9.14%]
> > 512B 7.306 6.611 -9.51% [-10.92%, -8.08%]
> > 1K 3.811 3.458 -9.25% [-10.96%, -7.50%]
> > 4K 1.850 1.669 -9.75% [-11.78%, -7.67%]
> >
> >Jia Jia (5):
> > vsock: split socket lookup from locked RX processing
> > vsock: amortize RX socket locking for stream packets
> > vsock: reuse same-flow socket lookup in RX batches
> > vsock: coalesce RX write-space notifications in lock batches
> > vsock: defer RX readable notifications until batch unlock
> >
> > include/linux/virtio_vsock.h | 14 ++
> > include/net/af_vsock.h | 3 +
> > net/vmw_vsock/af_vsock.c | 2 +
> > net/vmw_vsock/virtio_transport.c | 48 ++++-
> > net/vmw_vsock/virtio_transport_common.c | 341 +++++++++++++++++++++++++++-----
> > 5 files changed, 357 insertions(+), 51 deletions(-)
> >
> >base-commit: f5f84daefcd92d7a630066635ecea1433ed5eac7
> >--
> >2.53.0
> >
>
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-02 13:22 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 7:45 [PATCH 0/5] vsock: reduce RX per-packet socket overhead physicalmtea
2026-10-02 10:39 ` Stefano Garzarella
2026-10-02 13:22 ` Jia Jia
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®