Re: [PATCH 0/5] vsock: reduce RX per-packet socket overhead

From: Stefano Garzarella

Date: Fri Oct 02 2026 - 06:50:50 EST


On Fri, Oct 02, 2026 at 07:45:46AM +0000, physicalmtea@xxxxxxxxx wrote:
From: Jia Jia <physicalmtea@xxxxxxxxx>

This series reduces per-packet socket overhead in the virtio-vsock RX
worker.

I did a quick look but IMO this is hard to review, please try to keep patches as small as possible, e.g. where you are introducing ctx, store ctx->net in a net variable, so you don't need to touch the entire function, etc.

Please check https://docs.kernel.org/process/coding-assistants.html
and also https://docs.kernel.org/process/submitting-patches.html
ans also https://docs.kernel.org/process/maintainer-netdev.html#git-trees-and-patch-flow
(I guess this is net-next material).

I don't think this series was sent properly, I tried
`b4 am 20261002074551.318789-1-physicalmtea@xxxxxxxxx` without success and also looking at https://lore.kernel.org/virtualization/20261002074551.318789-1-physicalmtea@xxxxxxxxx/ I can't see the full series, please check your workflow.

Also, is this for vsock core or just virtio-vsock? if it's just virtio-vsock transports, please use "vsock/virtio" prefix.



Patch 1: separate socket lookup from locked packet processing.

Patch 2: keep one socket lock across a bounded run of established
STREAM/RW packets for the same native socket.

Patch 3: reuse the locked socket for later packets with the same
address tuple.

Patch 4: coalesce default write-space notifications while that
lock is held.

Patch 5: defer the default readable callback until after the batch
unlocks. Custom callbacks retain per-packet notification behavior.

Performance:

Tested on single-stream AF_VSOCK (1 reader on Guest CPU0, virtio-vsock IRQ
on Guest CPU1, fresh QEMU guest per state, 16 balanced AB/BA pairs per
payload).

What the host is using? vhost-vsock?


End-to-end RX throughput (64K receiver buffer, SO_RCVLOWAT=1):

payload baseline RX (G/s) PATCH RX (G/s) throughput change (95% CI)

What G is? Gbit? Gbytes?

256B 0.5722 1.0286 +79.76% [+75.40%, +84.23%]
512B 1.1464 1.9779 +72.53% [+69.50%, +75.63%]
1K 2.2714 3.6368 +60.12% [+56.92%, +63.37%]
4K 3.9095 7.7160 +97.37% [+92.09%, +102.79%]

Why stop to 4k?


Fixed-syscall stress test (receiver buffer = sender write size; RX in G/s):

payload baseline RX PATCH RX throughput change (95% CI) recv() base/patch
256B 0.2885 0.6737 +133.50% [+123.58%, +143.86%] 1048577/1048577
512B 0.5631 1.2091 +114.71% [+104.97%, +124.92%] 1048577/1048577
1K 1.5361 2.9678 +93.20% [+80.15%, +107.20%] 1048577/1048577
4K 2.8378 6.1878 +118.05% [+109.15%, +127.33%] 538802/525645

Request-response latency (Ping-Pong RTT, 10,000 requests):

What tool did you used for both throughput and latency?


payload baseline mean RTT (us) PATCH mean RTT (us) change (95% CI)
256B 126.968 124.541 -1.91% [-3.35%, -0.45%]
4K 130.362 126.664 -2.84% [-5.12%, -0.50%]

Whys only this payload cases? For latency we should use even try small payloads IMO like 64 bytes.

Thanks,
Stefano


Diagnostic Guest CPU1 system-wide cycles (fixed-64K receiver, with perf):

Buffer size baseline cycles/B PATCH cycles/B cycles change (95% CI)
256B 15.287 13.630 -10.84% [-12.51%, -9.14%]
512B 7.306 6.611 -9.51% [-10.92%, -8.08%]
1K 3.811 3.458 -9.25% [-10.96%, -7.50%]
4K 1.850 1.669 -9.75% [-11.78%, -7.67%]

Jia Jia (5):
vsock: split socket lookup from locked RX processing
vsock: amortize RX socket locking for stream packets
vsock: reuse same-flow socket lookup in RX batches
vsock: coalesce RX write-space notifications in lock batches
vsock: defer RX readable notifications until batch unlock

include/linux/virtio_vsock.h | 14 ++
include/net/af_vsock.h | 3 +
net/vmw_vsock/af_vsock.c | 2 +
net/vmw_vsock/virtio_transport.c | 48 ++++-
net/vmw_vsock/virtio_transport_common.c | 341 +++++++++++++++++++++++++++-----
5 files changed, 357 insertions(+), 51 deletions(-)

base-commit: f5f84daefcd92d7a630066635ecea1433ed5eac7
--
2.53.0