From: lollipopkit <a@lolli.tech>
To: Jens Axboe <axboe@kernel.dk>
Cc: Pavel Begunkov <asml.silence@gmail.com>,
Willem de Bruijn <willemb@google.com>,
io-uring@vger.kernel.org, netdev@vger.kernel.org,
linux-kernel@vger.kernel.org, lollipopkit <a@lolli.tech>,
stable@vger.kernel.org
Subject: [PATCH] io_uring/cmd_net: end TX_TIMESTAMP multishot when the CQ is full
Date: Sun, 27 Sep 2026 15:42:06 +0800 [thread overview]
Message-ID: <20260927074206.65427-1-a@lolli.tech> (raw)
SOCKET_URING_OP_TX_TIMESTAMP arms an EPOLLERR apoll multishot and, on
each error queue edge, posts one 32 byte aux CQE per timestamp skb. Aux
CQEs have no overflow backing, so io_uring_cmd_post_mshot_cqe32() fails
once the CQ ring is full. The loop then stops, splices the unprocessed
skbs back onto sk_error_queue and returns -EAGAIN.
-EAGAIN is IOU_RETRY, so the request goes idle until the next poll
event. io_arm_apoll() forces EPOLLET and draining the CQ does not re-run
the command, so the spliced-back timestamps stay queued until an
unrelated later timestamp produces a new edge.
End the multishot when a CQE cannot be posted, as multishot poll and
recv already do. There is no partial result to report, so complete with
-ENOBUFS: no CQ space is left for the timestamp CQEs. This command does
not use provided buffers, so the value cannot be confused with a buffer
ring running empty. The terminal CQE still reaches userspace because
request completions can go to the overflow list. The application sees a
CQE without IORING_CQE_F_MORE, which it already has to handle, and the
first issue of the re-armed request delivers the queued timestamps.
A timestamp that cannot be extracted keeps its current behaviour.
This was found with LLM assistance while reviewing io_uring multishot
handlers for the pattern behind a RECV_ZC stall reported earlier (see
Link), and confirmed with a reproducer on v7.3-rc2 and an A/B run of
this patch.
Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
Link: https://lore.kernel.org/all/010001a0cd914529-23df2f94-bbe0-49cd-ae5d-05922356fa12-000000@email.amazonses.com/
Cc: stable@vger.kernel.org
Assisted-by: Codex:gpt-6-sol
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: lollipopkit <a@lolli.tech>
---
Reproduction and validation (measured on this revision):
- Loopback, one process, software timestamps, CQE32 ring with 16 CQ
entries, a burst of 200 SO_TIMESTAMPING sends. The filling edge
delivers exactly 16 CQEs. With no further sends nothing more arrives,
although the application drains the ring and keeps entering io_uring;
more timestamped sends release them.
A burst smaller than the ring leaves nothing over and does not stall.
- A/B in a KVM guest on axboe/io_uring-7.3 (a3bdf68feecc) + this patch,
with the new return behind a test-only runtime switch so both arms run
the same binary, 10 runs per arm, identical numbers on every run:
switch off (upstream behaviour): 10/10 stalled, 184 timestamps left
queued, no terminal CQE
switch on (this patch): 10/10 terminal CQE with res -ENOBUFS; after
re-arming, the queued timestamps were delivered
io_uring/cmd_net.c | 19 ++++++++++++++-----
1 file changed, 14 insertions(+), 5 deletions(-)
diff --git a/io_uring/cmd_net.c b/io_uring/cmd_net.c
index 7cd411fc4f33..90d4ec7cc761 100644
--- a/io_uring/cmd_net.c
+++ b/io_uring/cmd_net.c
@@ -69,8 +69,8 @@ static inline int io_uring_cmd_setsockopt(struct socket *sock,
optlen);
}
-static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
- struct sk_buff *skb, unsigned issue_flags)
+static int io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
+ struct sk_buff *skb, unsigned int issue_flags)
{
struct sock_exterr_skb *serr = SKB_EXT_ERR(skb);
struct io_uring_cqe cqe[2];
@@ -83,7 +83,7 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
ret = skb_get_tx_timestamp(skb, sk, &ts);
if (ret < 0)
- return false;
+ return ret;
tskey = serr->ee.ee_data;
tstype = serr->ee.ee_info;
@@ -98,7 +98,9 @@ static bool io_process_timestamp_skb(struct io_uring_cmd *cmd, struct sock *sk,
iots = (struct io_timespec *)&cqe[1];
iots->tv_sec = ts.tv_sec;
iots->tv_nsec = ts.tv_nsec;
- return io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe);
+ if (!io_uring_cmd_post_mshot_cqe32(cmd, issue_flags, cqe))
+ return -ENOBUFS;
+ return 0;
}
static int io_uring_cmd_timestamp(struct socket *sock,
@@ -135,7 +137,8 @@ static int io_uring_cmd_timestamp(struct socket *sock,
skb = skb_peek(&list);
if (!skb)
break;
- if (!io_process_timestamp_skb(cmd, sk, skb, issue_flags))
+ ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
+ if (ret)
break;
__skb_dequeue(&list);
consume_skb(skb);
@@ -145,6 +148,12 @@ static int io_uring_cmd_timestamp(struct socket *sock,
scoped_guard(spinlock_irqsave, &q->lock)
skb_queue_splice(&list, q);
}
+ /*
+ * Aux CQEs cannot overflow and the poll is edge triggered, so nothing
+ * re-runs the command once the CQ drains. End the multishot instead.
+ */
+ if (ret == -ENOBUFS)
+ return -ENOBUFS;
return -EAGAIN;
}
base-commit: a3bdf68feecc57af5c11fb599f860ac9790ffad9
--
2.54.0
next reply other threads:[~2026-09-27 7:43 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-27 7:42 lollipopkit [this message]
2026-09-27 19:16 ` Pavel Begunkov
2026-09-28 14:31 ` Jens Axboe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260927074206.65427-1-a@lolli.tech \
--to=a@lolli.tech \
--cc=asml.silence@gmail.com \
--cc=axboe@kernel.dk \
--cc=io-uring@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=stable@vger.kernel.org \
--cc=willemb@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®