mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
@ 2026-09-27  7:42 lollipopkit
  2026-09-27 19:16 ` Pavel Begunkov
  2026-09-28 14:31 ` Jens Axboe
  0 siblings, 2 replies; 3+ messages in thread
From: lollipopkit @ 2026-09-27  7:42 UTC (permalink / raw)
  To: Jens Axboe
  Cc: Pavel Begunkov, Willem de Bruijn, io-uring, netdev, linux-kernel,
	lollipopkit, stable

SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
once the CQ ring is full. The loop then stops, splices the unprocessed
skbs back onto sk_error_queue and returns -EAGAIN.

-EAGAIN is IOU_RETRY, so the request goes idle until the next poll
event. io_arm_apoll() forces EPOLLET and draining the CQ does not re-run
the command, so the spliced-back timestamps stay queued until an
unrelated later timestamp produces a new edge.

End the multishot when a CQE cannot be posted, as multishot poll and
recv already do. There is no partial result to report, so complete with
-ENOBUFS: no CQ space is left for the timestamp CQEs. This command does
not use provided buffers, so the value cannot be confused with a buffer
ring running empty. The terminal CQE still reaches userspace because
request completions can go to the overflow list. The application sees a
CQE without IORING_CQE_F_MORE, which it already has to handle, and the
first issue of the re-armed request delivers the queued timestamps.

A timestamp that cannot be extracted keeps its current behaviour.

This was found with LLM assistance while reviewing io_uring multishot
handlers for the pattern behind a RECV_ZC stall reported earlier (see
Link), and confirmed with a reproducer on v7.3-rc2 and an A/B run of
this patch.

Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
Link: https://lore.kernel.org/all/010001a0cd914529-23df2f94-bbe0-49cd-ae5d-05922356fa12-000000@email.amazonses.com/
Cc: stable@vger.kernel.org
Assisted-by: Codex:gpt-6-sol
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: lollipopkit <a@lolli.tech>
---
Reproduction and validation (measured on this revision):
- Loopback, one process, software timestamps, CQE32 ring with 16 CQ
  entries, a burst of 200 SO_TIMESTAMPING sends. The filling edge
  delivers exactly 16 CQEs. With no further sends nothing more arrives,
  although the application drains the ring and keeps entering io_uring;
  more timestamped sends release them.
  A burst smaller than the ring leaves nothing over and does not stall.
- A/B in a KVM guest on axboe/io_uring-7.3 (a3bdf68feecc) + this patch,
  with the new return behind a test-only runtime switch so both arms run
  the same binary, 10 runs per arm, identical numbers on every run:
    switch off (upstream behaviour): 10/10 stalled, 184 timestamps left
      queued, no terminal CQE
    switch on (this patch): 10/10 terminal CQE with res -ENOBUFS; after
      re-arming, the queued timestamps were delivered

 io_uring/cmd_net.c | 19 ++++++++++++++-----
 1 file changed, 14 insertions(+), 5 deletions(-)

diff --git a/io_uring/cmd_net.c b/io_uring/cmd_net.c
index 7cd411fc4f33..90d4ec7cc761 100644
--- a/io_uring/cmd_net.c
+++ b/io_uring/cmd_net.c
@@ -69,8 +69,8 @@ static inline int io_uring_cmd_setsockopt(struct socket *sock,
 				  optlen);
 }
 
-static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
-				     struct sk_buff *skb, unsigned issue_flags)
+static int io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
+				    struct sk_buff *skb, unsigned int issue_flags)
 {
 	struct sock_exterr_skb *serr = SKB_EXT_ERR(skb);
 	struct io_uring_cqe cqe[2];
@@ -83,7 +83,7 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
 
 	ret = skb_get_tx_timestamp(skb, sk, &ts);
 	if (ret < 0)
-		return false;
+		return ret;
 
 	tskey = serr->ee.ee_data;
 	tstype = serr->ee.ee_info;
@@ -98,7 +98,9 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
 	iots = (struct io_timespec *)&cqe[1];
 	iots->tv_sec = ts.tv_sec;
 	iots->tv_nsec = ts.tv_nsec;
-	return io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe);
+	if (!io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe))
+		return -ENOBUFS;
+	return 0;
 }
 
 static int io_uring_cmd_timestamp(struct socket *sock,
@@ -135,7 +137,8 @@ static int io_uring_cmd_timestamp(struct socket *sock,
 		skb = skb_peek(&list);
 		if (!skb)
 			break;
-		if (!io_process_timestamp_skb(cmd, sk, skb, issue_flags))
+		ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
+		if (ret)
 			break;
 		__skb_dequeue(&list);
 		consume_skb(skb);
@@ -145,6 +148,12 @@ static int io_uring_cmd_timestamp(struct socket *sock,
 		scoped_guard(spinlock_irqsave, &q->lock)
 			skb_queue_splice(&list, q);
 	}
+	/*
+	 * Aux CQEs cannot overflow and the poll is edge triggered, so nothing
+	 * re-runs the command once the CQ drains. End the multishot instead.
+	 */
+	if (ret == -ENOBUFS)
+		return -ENOBUFS;
 	return -EAGAIN;
 }
 

base-commit: a3bdf68feecc57af5c11fb599f860ac9790ffad9
-- 
2.54.0


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
  2026-09-27  7:42 [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full lollipopkit
@ 2026-09-27 19:16 ` Pavel Begunkov
  2026-09-28 14:31 ` Jens Axboe
  1 sibling, 0 replies; 3+ messages in thread
From: Pavel Begunkov @ 2026-09-27 19:16 UTC (permalink / raw)
  To: lollipopkit, Jens Axboe
  Cc: Willem de Bruijn, io-uring, netdev, linux-kernel, stable

On 9/27/26 08:42, lollipopkit wrote:
> SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
> each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
> CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
> once the CQ ring is full. The loop then stops, splices the unprocessed
> skbs back onto sk_error_queue and returns -EAGAIN.
...>   static int io_uring_cmd_timestamp(struct socket *sock,
> @@ -135,7 +137,8 @@ static int io_uring_cmd_timestamp(struct socket *sock,
>   		skb = skb_peek(&list);
>   		if (!skb)
>   			break;
> -		if (!io_process_timestamp_skb(cmd, sk, skb, issue_flags))
> +		ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
> +		if (ret)
>   			break;
>   		__skb_dequeue(&list);
>   		consume_skb(skb);
> @@ -145,6 +148,12 @@ static int io_uring_cmd_timestamp(struct socket *sock,
>   		scoped_guard(spinlock_irqsave, &q->lock)
>   			skb_queue_splice(&list, q);
>   	}
> +	/*
> +	 * Aux CQEs cannot overflow and the poll is edge triggered, so nothing
> +	 * re-runs the command once the CQ drains. End the multishot instead.
> +	 */
> +	if (ret == -ENOBUFS)
> +		return -ENOBUFS;

I'd make it a 3-state return instead of relying on
skb_get_tx_timestamp() not returning specific error codes, but it
should be fine for now.

Reviewed-by: Pavel Begunkov <asml.silence@gmail.com>

>   	return -EAGAIN;
>   }
>   
> 
> base-commit: a3bdf68feecc57af5c11fb599f860ac9790ffad9

-- 
Pavel Begunkov


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
  2026-09-27  7:42 [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full lollipopkit
  2026-09-27 19:16 ` Pavel Begunkov
@ 2026-09-28 14:31 ` Jens Axboe
  1 sibling, 0 replies; 3+ messages in thread
From: Jens Axboe @ 2026-09-28 14:31 UTC (permalink / raw)
  To: lollipopkit
  Cc: Pavel Begunkov, Willem de Bruijn, io-uring, netdev, linux-kernel, stable


On Sun, 27 Sep 2026 15:42:06 +0800, lollipopkit wrote:
> SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
> each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
> CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
> once the CQ ring is full. The loop then stops, splices the unprocessed
> skbs back onto sk_error_queue and returns -EAGAIN.
> 
> -EAGAIN is IOU_RETRY, so the request goes idle until the next poll
> event. io_arm_apoll() forces EPOLLET and draining the CQ does not re-run
> the command, so the spliced-back timestamps stay queued until an
> unrelated later timestamp produces a new edge.
> 
> [...]

Applied, thanks!

[1/1] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
      commit: 3a3d93070c5ec8b8049d57025987ba313420bda9

Best regards,
-- 
Jens Axboe




^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-28 14:31 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27  7:42 [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full lollipopkit
2026-09-27 19:16 ` Pavel Begunkov
2026-09-28 14:31 ` Jens Axboe

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®