* [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
@ 2026-09-27 7:42 lollipopkit
2026-09-27 19:16 ` Pavel Begunkov
2026-09-28 14:31 ` Jens Axboe
0 siblings, 2 replies; 3+ messages in thread
From: lollipopkit @ 2026-09-27 7:42 UTC (permalink / raw)
To: Jens Axboe
Cc: Pavel Begunkov, Willem de Bruijn, io-uring, netdev, linux-kernel,
lollipopkit, stable
SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
once the CQ ring is full. The loop then stops, splices the unprocessed
skbs back onto sk_error_queue and returns -EAGAIN.
-EAGAIN is IOU_RETRY, so the request goes idle until the next poll
event. io_arm_apoll() forces EPOLLET and draining the CQ does not re-run
the command, so the spliced-back timestamps stay queued until an
unrelated later timestamp produces a new edge.
End the multishot when a CQE cannot be posted, as multishot poll and
recv already do. There is no partial result to report, so complete with
-ENOBUFS: no CQ space is left for the timestamp CQEs. This command does
not use provided buffers, so the value cannot be confused with a buffer
ring running empty. The terminal CQE still reaches userspace because
request completions can go to the overflow list. The application sees a
CQE without IORING_CQE_F_MORE, which it already has to handle, and the
first issue of the re-armed request delivers the queued timestamps.
A timestamp that cannot be extracted keeps its current behaviour.
This was found with LLM assistance while reviewing io_uring multishot
handlers for the pattern behind a RECV_ZC stall reported earlier (see
Link), and confirmed with a reproducer on v7.3-rc2 and an A/B run of
this patch.
Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
Link: https://lore.kernel.org/all/010001a0cd914529-23df2f94-bbe0-49cd-ae5d-05922356fa12-000000@email.amazonses.com/
Cc: stable@vger.kernel.org
Assisted-by: Codex:gpt-6-sol
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: lollipopkit <a@lolli.tech>
---
Reproduction and validation (measured on this revision):
- Loopback, one process, software timestamps, CQE32 ring with 16 CQ
entries, a burst of 200 SO_TIMESTAMPING sends. The filling edge
delivers exactly 16 CQEs. With no further sends nothing more arrives,
although the application drains the ring and keeps entering io_uring;
more timestamped sends release them.
A burst smaller than the ring leaves nothing over and does not stall.
- A/B in a KVM guest on axboe/io_uring-7.3 (a3bdf68feecc) + this patch,
with the new return behind a test-only runtime switch so both arms run
the same binary, 10 runs per arm, identical numbers on every run:
switch off (upstream behaviour): 10/10 stalled, 184 timestamps left
queued, no terminal CQE
switch on (this patch): 10/10 terminal CQE with res -ENOBUFS; after
re-arming, the queued timestamps were delivered
io_uring/cmd_net.c | 19 ++++++++++++++-----
1 file changed, 14 insertions(+), 5 deletions(-)
diff --git a/io_uring/cmd_net.c b/io_uring/cmd_net.c
index 7cd411fc4f33..90d4ec7cc761 100644
--- a/io_uring/cmd_net.c
+++ b/io_uring/cmd_net.c
@@ -69,8 +69,8 @@ static inline int io_uring_cmd_setsockopt(struct socket *sock,
optlen);
}
-static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
- struct sk_buff *skb, unsigned issue_flags)
+static int io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
+ struct sk_buff *skb, unsigned int issue_flags)
{
struct sock_exterr_skb *serr = SKB_EXT_ERR(skb);
struct io_uring_cqe cqe[2];
@@ -83,7 +83,7 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
ret = skb_get_tx_timestamp(skb, sk, &ts);
if (ret < 0)
- return false;
+ return ret;
tskey = serr->ee.ee_data;
tstype = serr->ee.ee_info;
@@ -98,7 +98,9 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
iots = (struct io_timespec *)&cqe[1];
iots->tv_sec = ts.tv_sec;
iots->tv_nsec = ts.tv_nsec;
- return io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe);
+ if (!io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe))
+ return -ENOBUFS;
+ return 0;
}
static int io_uring_cmd_timestamp(struct socket *sock,
@@ -135,7 +137,8 @@ static int io_uring_cmd_timestamp(struct socket *sock,
skb = skb_peek(&list);
if (!skb)
break;
- if (!io_process_timestamp_skb(cmd, sk, skb, issue_flags))
+ ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
+ if (ret)
break;
__skb_dequeue(&list);
consume_skb(skb);
@@ -145,6 +148,12 @@ static int io_uring_cmd_timestamp(struct socket *sock,
scoped_guard(spinlock_irqsave, &q->lock)
skb_queue_splice(&list, q);
}
+ /*
+ * Aux CQEs cannot overflow and the poll is edge triggered, so nothing
+ * re-runs the command once the CQ drains. End the multishot instead.
+ */
+ if (ret == -ENOBUFS)
+ return -ENOBUFS;
return -EAGAIN;
}
base-commit: a3bdf68feecc57af5c11fb599f860ac9790ffad9
--
2.54.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
2026-09-27 7:42 [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full lollipopkit
@ 2026-09-27 19:16 ` Pavel Begunkov
2026-09-28 14:31 ` Jens Axboe
1 sibling, 0 replies; 3+ messages in thread
From: Pavel Begunkov @ 2026-09-27 19:16 UTC (permalink / raw)
To: lollipopkit, Jens Axboe
Cc: Willem de Bruijn, io-uring, netdev, linux-kernel, stable
On 9/27/26 08:42, lollipopkit wrote:
> SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
> each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
> CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
> once the CQ ring is full. The loop then stops, splices the unprocessed
> skbs back onto sk_error_queue and returns -EAGAIN.
...> static int io_uring_cmd_timestamp(struct socket *sock,
> @@ -135,7 +137,8 @@ static int io_uring_cmd_timestamp(struct socket *sock,
> skb = skb_peek(&list);
> if (!skb)
> break;
> - if (!io_process_timestamp_skb(cmd, sk, skb, issue_flags))
> + ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
> + if (ret)
> break;
> __skb_dequeue(&list);
> consume_skb(skb);
> @@ -145,6 +148,12 @@ static int io_uring_cmd_timestamp(struct socket *sock,
> scoped_guard(spinlock_irqsave, &q->lock)
> skb_queue_splice(&list, q);
> }
> + /*
> + * Aux CQEs cannot overflow and the poll is edge triggered, so nothing
> + * re-runs the command once the CQ drains. End the multishot instead.
> + */
> + if (ret == -ENOBUFS)
> + return -ENOBUFS;
I'd make it a 3-state return instead of relying on
skb_get_tx_timestamp() not returning specific error codes, but it
should be fine for now.
Reviewed-by: Pavel Begunkov <asml.silence@gmail.com>
> return -EAGAIN;
> }
>
>
> base-commit: a3bdf68feecc57af5c11fb599f860ac9790ffad9
--
Pavel Begunkov
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
2026-09-27 7:42 [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full lollipopkit
2026-09-27 19:16 ` Pavel Begunkov
@ 2026-09-28 14:31 ` Jens Axboe
1 sibling, 0 replies; 3+ messages in thread
From: Jens Axboe @ 2026-09-28 14:31 UTC (permalink / raw)
To: lollipopkit
Cc: Pavel Begunkov, Willem de Bruijn, io-uring, netdev, linux-kernel, stable
On Sun, 27 Sep 2026 15:42:06 +0800, lollipopkit wrote:
> SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
> each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
> CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
> once the CQ ring is full. The loop then stops, splices the unprocessed
> skbs back onto sk_error_queue and returns -EAGAIN.
>
> -EAGAIN is IOU_RETRY, so the request goes idle until the next poll
> event. io_arm_apoll() forces EPOLLET and draining the CQ does not re-run
> the command, so the spliced-back timestamps stay queued until an
> unrelated later timestamp produces a new edge.
>
> [...]
Applied, thanks!
[1/1] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
commit: 3a3d93070c5ec8b8049d57025987ba313420bda9
Best regards,
--
Jens Axboe
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-28 14:31 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 7:42 [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full lollipopkit
2026-09-27 19:16 ` Pavel Begunkov
2026-09-28 14:31 ` Jens Axboe
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®