* [PATCH] io_uring/cmd_net: drop unextractable TX_TIMESTAMP skbs instead of requeueing
@ 2026-09-29 4:54 lollipopkit
2026-09-29 15:06 ` Willem de Bruijn
0 siblings, 1 reply; 4+ messages in thread
From: lollipopkit @ 2026-09-29 4:54 UTC (permalink / raw)
To: Jens Axboe
Cc: Pavel Begunkov, Willem de Bruijn, Richard Cochran, io-uring,
netdev, linux-kernel, lollipopkit, stable
SOCKET_URING_OP_TX_TIMESTAMP selects error-queue skbs with
skb_has_tx_timestamp() but extracts the timestamp with
skb_get_tx_timestamp(). The two use different checks and can disagree:
a skb the walk accepted can be rejected with -ENOENT forever. When
that happens the processing loop stops and skb_queue_splice() puts the
unprocessed skbs back onto the HEAD of sk_error_queue. The rejected
skb is selected first on every following run, fails again, and blocks
every skb behind it: new timestamps keep arriving and keep re-running
the command, but nothing is ever delivered again.
The disagreement is reachable without a race. With
SOF_TIMESTAMPING_BIND_PHC the walk never looks at sk_bind_phc, but
skb_get_tx_timestamp() converts the hardware timestamp through
ptp_convert_timestamp(), which returns 0 when no virtual clock matches
sk_bind_phc (for example when the vclock was destroyed after the
socket bound to it). ktime_to_timespec64_cond() then fails and the
extractor returns -ENOENT permanently. Any skb with a non-zero
hwtstamp is selected while SOF_TIMESTAMPING_RAW_HARDWARE is set, so a
single such skb blocks even the software timestamps queued behind it.
Measured on real hardware, an I219-V (e1000e), using
/sys/class/ptp/ptpN/n_vclocks to create and destroy the vclock the
socket was bound to: 22 error-queue skbs generated (16 software + 6
hardware). Of the 21 extractor calls, 20 failed with -ENOENT and the
one that succeeded delivered 1 software timestamp; deliveries then
froze while 8 further sends kept re-running the command, and the
remaining 21 skbs (15 software + 6 hardware) sat in the error queue.
Both control arms (vclock kept alive; no BIND_PHC) delivered
everything.
With this patch applied on top of the CQ-full fix ("io_uring/cmd_net:
end TX_TIMESTAMP multishot when the CQ is full"), A/B tested in a KVM
guest from a single kernel binary, with the new behaviour behind a
test-only runtime switch (the switch is test scaffolding and is not
part of this patch). The guest needs no special hardware: QEMU's igb
model emulates hardware TX timestamping (TSYNCTXCTL/TXSTMPL/TXSTMPH)
and the igb driver registers a PTP clock, so the guest reproduces the
same failure the same way. The two arms are separate runs with their
own sockets and sends; the device holds one pending hardware
timestamp, so how many of the 16 sends get a hardware skb varies per
run (12 and 8 here). Switch off: 28 skbs generated (16 software + 12
hardware), deliveries froze at 1 and the remaining 27 (15 software +
12 hardware) sat in the error queue. Switch on: 24 skbs generated (16
software + 8 hardware); all 16 software timestamps were delivered, the
8 hardware skbs the extractor rejected with -ENOENT were dropped, and
the error queue drained.
recvmsg(MSG_ERRQUEUE) already treats such a skb as skippable:
sock_recv_errqueue() dequeues it, returns the error message without a
timestamp, and frees it. Do the same here and keep the deliverable
timestamps moving: drop the skb the extractor rejected and continue
with the rest of the list. The CQE-full case keeps its
requeue-and-end-multishot behaviour.
Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
Cc: stable@vger.kernel.org # needs the CQ-full fix above
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: lollipopkit <a@lolli.tech>
---
This applies on top of axboe/for-next (3769bc31a5b7). The CQ-full fix
is 3a3d93070c5e in io_uring-7.3; the two touch adjacent lines of the
same loop, hence the stable note.
Reproducer: a single-file liburing program that runs the three arms
(destroyed vclock, live vclock, no BIND_PHC) against any NIC with
hardware TX timestamps, plus a KVM guest setup using QEMU's igb model.
Happy to post it or turn it into a selftest if that is useful.
One behaviour change worth calling out: skb_has_tx_timestamp() and
skb_get_tx_timestamp() read sk_tsflags independently, so a concurrent
setsockopt(SO_TIMESTAMPING) can also make the extractor reject a skb
the walk selected. Before this patch that skb was delivered on a later
run; now it is dropped, as recvmsg(MSG_ERRQUEUE) would. With a thread
switching sk_tsflags about 800k times against 20000 sends on loopback,
this patch dropped 320 such timestamps. Reading sk_tsflags once per
run and passing it to both helpers removes that case (0 rejections, 0
lost in the same test); it changes both helpers' signatures in
net/socket.c, so I left it out of this fix. I can send it as a
preceding patch if you prefer the two together.
io_uring/cmd_net.c | 15 +++++++++++++++
1 file changed, 15 insertions(+)
diff --git a/io_uring/cmd_net.c b/io_uring/cmd_net.c
index 90d4ec7cc761..406fbe8b7d52 100644
--- a/io_uring/cmd_net.c
+++ b/io_uring/cmd_net.c
@@ -138,6 +138,21 @@ static int io_uring_cmd_timestamp(struct socket *sock,
if (!skb)
break;
ret = io_process_timestamp_skb(cmd, sk, skb, issue_flags);
+ if (ret && ret != -ENOBUFS) {
+ /*
+ * The walk selected this skb, but skb_get_tx_timestamp()
+ * cannot extract a timestamp from it: the two use
+ * different checks. Splicing it back would put it at
+ * the head of sk_error_queue, where every later run
+ * selects it first and fails on it again, so nothing
+ * behind it is ever delivered. recvmsg(MSG_ERRQUEUE)
+ * frees such a skb without reporting a timestamp; drop
+ * it here as well.
+ */
+ __skb_dequeue(&list);
+ kfree_skb(skb);
+ continue;
+ }
if (ret)
break;
__skb_dequeue(&list);
base-commit: 3769bc31a5b79ec406df085592fd9dc14ce7df12
--
2.54.0
^ permalink raw reply [flat|nested] 4+ messages in thread* Re: [PATCH] io_uring/cmd_net: drop unextractable TX_TIMESTAMP skbs instead of requeueing
2026-09-29 4:54 [PATCH] io_uring/cmd_net: drop unextractable TX_TIMESTAMP skbs instead of requeueing lollipopkit
@ 2026-09-29 15:06 ` Willem de Bruijn
2026-09-29 15:35 ` lollipopkit
0 siblings, 1 reply; 4+ messages in thread
From: Willem de Bruijn @ 2026-09-29 15:06 UTC (permalink / raw)
To: lollipopkit, Jens Axboe
Cc: Pavel Begunkov, Willem de Bruijn, Richard Cochran, io-uring,
netdev, linux-kernel, lollipopkit, stable
lollipopkit wrote:
> SOCKET_URING_OP_TX_TIMESTAMP selects error-queue skbs with
> skb_has_tx_timestamp() but extracts the timestamp with
> skb_get_tx_timestamp(). The two use different checks and can disagree:
> a skb the walk accepted can be rejected with -ENOENT forever. When
> that happens the processing loop stops and skb_queue_splice() puts the
> unprocessed skbs back onto the HEAD of sk_error_queue. The rejected
> skb is selected first on every following run, fails again, and blocks
> every skb behind it: new timestamps keep arriving and keep re-running
> the command, but nothing is ever delivered again.
>
> The disagreement is reachable without a race. With
> SOF_TIMESTAMPING_BIND_PHC the walk never looks at sk_bind_phc, but
> skb_get_tx_timestamp() converts the hardware timestamp through
> ptp_convert_timestamp(), which returns 0 when no virtual clock matches
> sk_bind_phc (for example when the vclock was destroyed after the
> socket bound to it). ktime_to_timespec64_cond() then fails and the
> extractor returns -ENOENT permanently. Any skb with a non-zero
> hwtstamp is selected while SOF_TIMESTAMPING_RAW_HARDWARE is set, so a
> single such skb blocks even the software timestamps queued behind it.
>
> Measured on real hardware, an I219-V (e1000e), using
> /sys/class/ptp/ptpN/n_vclocks to create and destroy the vclock the
> socket was bound to: 22 error-queue skbs generated (16 software + 6
> hardware). Of the 21 extractor calls, 20 failed with -ENOENT and the
> one that succeeded delivered 1 software timestamp; deliveries then
> froze while 8 further sends kept re-running the command, and the
> remaining 21 skbs (15 software + 6 hardware) sat in the error queue.
> Both control arms (vclock kept alive; no BIND_PHC) delivered
> everything.
>
> With this patch applied on top of the CQ-full fix ("io_uring/cmd_net:
> end TX_TIMESTAMP multishot when the CQ is full"), A/B tested in a KVM
> guest from a single kernel binary, with the new behaviour behind a
> test-only runtime switch (the switch is test scaffolding and is not
> part of this patch). The guest needs no special hardware: QEMU's igb
> model emulates hardware TX timestamping (TSYNCTXCTL/TXSTMPL/TXSTMPH)
> and the igb driver registers a PTP clock, so the guest reproduces the
> same failure the same way. The two arms are separate runs with their
> own sockets and sends; the device holds one pending hardware
> timestamp, so how many of the 16 sends get a hardware skb varies per
> run (12 and 8 here). Switch off: 28 skbs generated (16 software + 12
> hardware), deliveries froze at 1 and the remaining 27 (15 software +
> 12 hardware) sat in the error queue. Switch on: 24 skbs generated (16
> software + 8 hardware); all 16 software timestamps were delivered, the
> 8 hardware skbs the extractor rejected with -ENOENT were dropped, and
> the error queue drained.
>
> recvmsg(MSG_ERRQUEUE) already treats such a skb as skippable:
> sock_recv_errqueue() dequeues it, returns the error message without a
> timestamp, and frees it. Do the same here and keep the deliverable
> timestamps moving: drop the skb the extractor rejected and continue
> with the rest of the list. The CQE-full case keeps its
> requeue-and-end-multishot behaviour.
>
> Fixes: 9e4ed359b8ef ("io_uring/netcmd: add tx timestamping cmd support")
> Cc: stable@vger.kernel.org # needs the CQ-full fix above
> Assisted-by: Claude:claude-opus-5-5
> Signed-off-by: lollipopkit <a@lolli.tech>
Is this a real name?
https://docs.kernel.org/process/submitting-patches.html#sign-your-work-the-developer-s-certificate-of-origin
"a known identity (sorry, no anonymous contributions.)"
^ permalink raw reply [flat|nested] 4+ messages in thread* Re: [PATCH] io_uring/cmd_net: drop unextractable TX_TIMESTAMP skbs instead of requeueing
2026-09-29 15:06 ` Willem de Bruijn
@ 2026-09-29 15:35 ` lollipopkit
2026-09-30 14:24 ` Willem de Bruijn
0 siblings, 1 reply; 4+ messages in thread
From: lollipopkit @ 2026-09-29 15:35 UTC (permalink / raw)
To: Willem de Bruijn
Cc: Jens Axboe, Pavel Begunkov, Willem de Bruijn, Richard Cochran,
io-uring, netdev, linux-kernel, stable
On Tue, 29 Sep 2026 11:06:26 -0400, Willem de Bruijn wrote:
> lollipopkit wrote:
>> Assisted-by: Claude:claude-opus-5-5
>> Signed-off-by: lollipopkit <a@lolli.tech>
>
> Is this a real name?
>
> https://docs.kernel.org/process/submitting-patches.html#sign-your-work-the-developer-s-certificate-of-origin
>
> "a known identity (sorry, no anonymous contributions.)"
It is not my legal name. It is the handle I have used publicly for
years, including on GitHub (https://github.com/lollipopkit), and it is
the identity I sign all my kernel work with. My understanding is that
this is what d4563201f33a ("Documentation: simplify and clarify DCO
contribution example language") intended by "known identity".
The earlier CQ-full fix (3a3d93070c5e) went in under the same name.
If that is not sufficient for this tree, let me know what you need.
^ permalink raw reply [flat|nested] 4+ messages in thread* Re: [PATCH] io_uring/cmd_net: drop unextractable TX_TIMESTAMP skbs instead of requeueing
2026-09-29 15:35 ` lollipopkit
@ 2026-09-30 14:24 ` Willem de Bruijn
0 siblings, 0 replies; 4+ messages in thread
From: Willem de Bruijn @ 2026-09-30 14:24 UTC (permalink / raw)
To: lollipopkit, Willem de Bruijn
Cc: Jens Axboe, Pavel Begunkov, Willem de Bruijn, Richard Cochran,
io-uring, netdev, linux-kernel, stable
lollipopkit wrote:
> On Tue, 29 Sep 2026 11:06:26 -0400, Willem de Bruijn wrote:
> > lollipopkit wrote:
> >> Assisted-by: Claude:claude-opus-5-5
> >> Signed-off-by: lollipopkit <a@lolli.tech>
> >
> > Is this a real name?
> >
> > https://docs.kernel.org/process/submitting-patches.html#sign-your-work-the-developer-s-certificate-of-origin
> >
> > "a known identity (sorry, no anonymous contributions.)"
>
> It is not my legal name. It is the handle I have used publicly for
> years, including on GitHub (https://github.com/lollipopkit), and it is
> the identity I sign all my kernel work with. My understanding is that
> this is what d4563201f33a ("Documentation: simplify and clarify DCO
> contribution example language") intended by "known identity".
> The earlier CQ-full fix (3a3d93070c5e) went in under the same name.
>
> If that is not sufficient for this tree, let me know what you need.
The commit you mention did remove the explicit reference to pseudonyms.
I don't know whether a github handle counts as "something we could check
back with" or where exactly the line gets drawn.
If anything, this does seem an outlier compared to other sign offs.
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-30 14:24 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 4:54 [PATCH] io_uring/cmd_net: drop unextractable TX_TIMESTAMP skbs instead of requeueing lollipopkit
2026-09-29 15:06 ` Willem de Bruijn
2026-09-29 15:35 ` lollipopkit
2026-09-30 14:24 ` Willem de Bruijn
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®