[PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full

From: lollipopkit

Date: Sun Sep 27 2026 - 03:43:59 EST


SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
once the CQ ring is full. The loop then stops, splices the unprocessed
skbs back onto sk_error_queue and returns -EAGAIN.

-EAGAIN is IOU_RETRY, so the request goes idle until the next poll
event. io_arm_apoll() forces EPOLLET and draining the CQ does not re-run
the command, so the spliced-back timestamps stay queued until an
unrelated later timestamp produces a new edge.

End the multishot when a CQE cannot be posted, as multishot poll and
recv already do. There is no partial result to report, so complete with
-ENOBUFS: no CQ space is left for the timestamp CQEs. This command does
not use provided buffers, so the value cannot be confused with a buffer
ring running empty. The terminal CQE still reaches userspace because
request completions can go to the overflow list. The application sees a
CQE without IORING_CQE_F_MORE, which it already has to handle, and the
first issue of the re-armed request delivers the queued timestamps.

A timestamp that cannot be extracted keeps its current behaviour.

This was found with LLM assistance while reviewing io_uring multishot
handlers for the pattern behind a RECV_ZC stall reported earlier (see
Link), and confirmed with a reproducer on v7.3-rc2 and an A/B run of
this patch.

Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
Link: https://lore.kernel.org/all/010001a0cd914529-23df2f94-bbe0-49cd-ae5d-05922356fa12-000000@xxxxxxxxxxxxxxxxxxx/
Cc: stable@xxxxxxxxxxxxxxx
Assisted-by: Codex:gpt-6-sol
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: lollipopkit <a@xxxxxxxxxx>
---
Reproduction and validation (measured on this revision):
- Loopback, one process, software timestamps, CQE32 ring with 16 CQ
entries, a burst of 200 SO_TIMESTAMPING sends. The filling edge
delivers exactly 16 CQEs. With no further sends nothing more arrives,
although the application drains the ring and keeps entering io_uring;
more timestamped sends release them.
A burst smaller than the ring leaves nothing over and does not stall.
- A/B in a KVM guest on axboe/io_uring-7.3 (a3bdf68feecc) + this patch,
with the new return behind a test-only runtime switch so both arms run
the same binary, 10 runs per arm, identical numbers on every run:
switch off (upstream behaviour): 10/10 stalled, 184 timestamps left
queued, no terminal CQE
switch on (this patch): 10/10 terminal CQE with res -ENOBUFS; after
re-arming, the queued timestamps were delivered

io_uring/cmd_net.c | 19 ++++++++++++++-----
1 file changed, 14 insertions(+), 5 deletions(-)

diff --git a/io_uring/cmd_net.c b/io_uring/cmd_net.c
index 7cd411fc4f33..90d4ec7cc761 100644
--- a/io_uring/cmd_net.c
+++ b/io_uring/cmd_net.c
@@ -69,8 +69,8 @@ static inline int io_uring_cmd_setsockopt(struct socket *sock,
optlen);
}

-static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
- struct sk_buff *skb, unsigned issue_flags)
+static int io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
+ struct sk_buff *skb, unsigned int issue_flags)
{
struct sock_exterr_skb *serr = SKB_EXT_ERR(skb);
struct io_uring_cqe cqe[2];
@@ -83,7 +83,7 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,

ret = skb_get_tx_timestamp(skb, sk, &ts);
if (ret < 0)
- return false;
+ return ret;

tskey = serr->ee.ee_data;
tstype = serr->ee.ee_info;
@@ -98,7 +98,9 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
iots = (struct io_timespec *)&cqe[1];
iots->tv_sec = ts.tv_sec;
iots->tv_nsec = ts.tv_nsec;
- return io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe);
+ if (!io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe))
+ return -ENOBUFS;
+ return 0;
}

static int io_uring_cmd_timestamp(struct socket *sock,
@@ -135,7 +137,8 @@ static int io_uring_cmd_timestamp(struct socket *sock,
skb = skb_peek(&list);
if (!skb)
break;
- if (!io_process_timestamp_skb(cmd, sk, skb, issue_flags))
+ ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
+ if (ret)
break;
__skb_dequeue(&list);
consume_skb(skb);
@@ -145,6 +148,12 @@ static int io_uring_cmd_timestamp(struct socket *sock,
scoped_guard(spinlock_irqsave, &q->lock)
skb_queue_splice(&list, q);
}
+ /*
+ * Aux CQEs cannot overflow and the poll is edge triggered, so nothing
+ * re-runs the command once the CQ drains. End the multishot instead.
+ */
+ if (ret == -ENOBUFS)
+ return -ENOBUFS;
return -EAGAIN;
}


base-commit: a3bdf68feecc57af5c11fb599f860ac9790ffad9
--
2.54.0