* [PATCH net v2 0/2] tcp: preserve RACK tracking across partial undo @ 2026-10-06 5:15 nramaswamy 2026-10-06 5:15 ` [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss nramaswamy 2026-10-06 5:15 ` [PATCH net v2 2/2] selftests: net: packetdrill: test RACK after partial undo nramaswamy 0 siblings, 2 replies; 7+ messages in thread From: nramaswamy @ 2026-10-06 5:15 UTC (permalink / raw) To: netdev Cc: Neil Ramaswamy, Eric Dumazet, Neal Cardwell, Kuniyuki Iwashima, Yuchung Cheng, Jiayuan Chen, David S. Miller, Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan, linux-kernel, linux-kselftest From: Neil Ramaswamy <nramaswamy@openai.com> Partial undo can clear the TCPCB_LOST flag on segments already removed from RACK's list, which prevents subsequent RACK loss detection, leading to segments only being retransmitted after the RTO. Restoring them to the RACK list as part of partial undo makes sure we can properly reconsider them for fast retransmission in the future. To do this, we first sort (by transmission time) the segments whose TCPCB_LOST flag is being cleared and then linearly insert those back into the RACK list, which is sorted by transmission time. My investigation started from seeing repeated TCP stalls in prod and the mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to 96 KiB. It also seems like others have seen similar symptoms before [1]. [1] https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@gmail.com/ Changes in v2: - Added production symptoms - Removed the redundant TCPCB_LOST comparison - Fixed lint/long lines Patch 2 is unchanged. v1: https://lore.kernel.org/netdev/20260926002520.42955-4-nramaswamy@openai.com/ Neil Ramaswamy (2): tcp: restore RACK list membership when undoing loss selftests: net: packetdrill: test RACK after partial undo net/ipv4/tcp_input.c | 35 ++++++++++++ ...tcp_partial_undo-restores-to-rack-list.pkt | 54 +++++++++++++++++++ 2 files changed, 89 insertions(+) create mode 100644 tools/testing/selftests/net/packetdrill/tcp_partial_undo-restores-to-rack-list.pkt base-commit: 11536ee3d3e0b1bd35b6f3f8df55a6053eb0c71d -- 2.55.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss 2026-10-06 5:15 [PATCH net v2 0/2] tcp: preserve RACK tracking across partial undo nramaswamy @ 2026-10-06 5:15 ` nramaswamy 2026-10-06 6:42 ` Eric Dumazet 2026-10-06 5:15 ` [PATCH net v2 2/2] selftests: net: packetdrill: test RACK after partial undo nramaswamy 1 sibling, 1 reply; 7+ messages in thread From: nramaswamy @ 2026-10-06 5:15 UTC (permalink / raw) To: netdev Cc: Neil Ramaswamy, Eric Dumazet, Neal Cardwell, Kuniyuki Iwashima, Yuchung Cheng, Jiayuan Chen, David S. Miller, Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan, linux-kernel, linux-kselftest From: Neil Ramaswamy <nramaswamy@openai.com> Partial undo can clear the TCPCB_LOST flag on segments already removed from RACK's list, which prevents subsequent RACK loss detection, leading to segments only being retransmitted after the RTO. Restoring them to the RACK list as part of partial undo makes sure we can properly reconsider them for fast retransmission in the future. To do this, we first sort (by transmission time) the segments whose TCPCB_LOST flag is being cleared and then linearly insert those back into the RACK list, which is sorted by transmission time. My investigation started from seeing repeated TCP stalls in prod and the mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to 96 KiB. It also seems like others have seen similar symptoms before [1]. [1] https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@gmail.com/ Fixes: 043b87d7599e ("tcp: more efficient RACK loss detection") Signed-off-by: Neil Ramaswamy <nramaswamy@openai.com> Assisted-by: LLM sparse --- net/ipv4/tcp_input.c | 35 +++++++++++++++++++++++++++++++++++ 1 file changed, 35 insertions(+) diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c index 92bc60716f33..51f04c7474dc 100644 --- a/net/ipv4/tcp_input.c +++ b/net/ipv4/tcp_input.c @@ -69,6 +69,7 @@ #include <linux/module.h> #include <linux/sysctl.h> #include <linux/kernel.h> +#include <linux/list_sort.h> #include <linux/prefetch.h> #include <linux/bitops.h> #include <net/dst.h> @@ -2840,16 +2841,50 @@ static void DBGUNDO(struct sock *sk, const char *msg) #endif } +static int tcp_rack_skb_cmp(void *priv, const struct list_head *a, + const struct list_head *b) +{ + const struct sk_buff *skb_a = list_entry(a, struct sk_buff, + tcp_tsorted_anchor); + const struct sk_buff *skb_b = list_entry(b, struct sk_buff, + tcp_tsorted_anchor); + + return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a), + tcp_skb_timestamp_us(skb_b), + TCP_SKB_CB(skb_a)->end_seq, + TCP_SKB_CB(skb_b)->end_seq); +} + static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss) { struct tcp_sock *tp = tcp_sk(sk); if (unmark_loss) { + LIST_HEAD(restored); struct sk_buff *skb; skb_rbtree_walk(skb, &sk->tcp_rtx_queue) { + if (TCP_SKB_CB(skb)->sacked & TCPCB_LOST) + list_move_tail(&skb->tcp_tsorted_anchor, + &restored); TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST; } + if (!list_empty(&restored)) { + struct list_head *pos = &tp->tsorted_sent_queue; + + /* Ensure lost skbs are added in transmission order */ + list_sort(NULL, &restored, tcp_rack_skb_cmp); + while (!list_empty(&restored)) { + struct list_head *entry = restored.next; + + while (pos->next != &tp->tsorted_sent_queue && + !tcp_rack_skb_cmp(NULL, pos->next, + entry)) + pos = pos->next; + list_move(entry, pos); + pos = entry; + } + } tp->lost_out = 0; tcp_clear_all_retrans_hints(tp); } -- 2.55.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss 2026-10-06 5:15 ` [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss nramaswamy @ 2026-10-06 6:42 ` Eric Dumazet 2026-10-06 11:23 ` Eric Dumazet 0 siblings, 1 reply; 7+ messages in thread From: Eric Dumazet @ 2026-10-06 6:42 UTC (permalink / raw) To: nramaswamy, netdev Cc: Neal Cardwell, Kuniyuki Iwashima, Yuchung Cheng, Jiayuan Chen, David S. Miller, Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan, linux-kernel, linux-kselftest On 10/6/26 07:15, nramaswamy@openai.com wrote: > From: Neil Ramaswamy <nramaswamy@openai.com> > > Partial undo can clear the TCPCB_LOST flag on segments already removed > from RACK's list, which prevents subsequent RACK loss detection, leading > to segments only being retransmitted after the RTO. Restoring them to > the RACK list as part of partial undo makes sure we can properly > reconsider them for fast retransmission in the future. > > To do this, we first sort (by transmission time) the segments whose > TCPCB_LOST flag is being cleared and then linearly insert those back > into the RACK list, which is sorted by transmission time. > > My investigation started from seeing repeated TCP stalls in prod and the > mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to > 96 KiB. It also seems like others have seen similar symptoms before [1]. > > [1] > https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@gmail.com/ > > Fixes: 043b87d7599e ("tcp: more efficient RACK loss detection") > Signed-off-by: Neil Ramaswamy <nramaswamy@openai.com> > Assisted-by: LLM sparse > --- > net/ipv4/tcp_input.c | 35 +++++++++++++++++++++++++++++++++++ > 1 file changed, 35 insertions(+) > > diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c > index 92bc60716f33..51f04c7474dc 100644 > --- a/net/ipv4/tcp_input.c > +++ b/net/ipv4/tcp_input.c > @@ -69,6 +69,7 @@ > #include <linux/module.h> > #include <linux/sysctl.h> > #include <linux/kernel.h> > +#include <linux/list_sort.h> > #include <linux/prefetch.h> > #include <linux/bitops.h> > #include <net/dst.h> > @@ -2840,16 +2841,50 @@ static void DBGUNDO(struct sock *sk, const char *msg) > #endif > } > > +static int tcp_rack_skb_cmp(void *priv, const struct list_head *a, > + const struct list_head *b) > +{ > + const struct sk_buff *skb_a = list_entry(a, struct sk_buff, > + tcp_tsorted_anchor); > + const struct sk_buff *skb_b = list_entry(b, struct sk_buff, > + tcp_tsorted_anchor); > + > + return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a), > + tcp_skb_timestamp_us(skb_b), > + TCP_SKB_CB(skb_a)->end_seq, > + TCP_SKB_CB(skb_b)->end_seq); Patch LGTM, but perhaps we could avoid the two div_u64() calls here. return tcp_skb_sent_after(skb_a->skb_mstamp_ns, skb_b->skb_mstamp_ns, TCP_SKB_CB(skb_a)->end_seq, TCP_SKB_CB(skb_b)->end_seq); > +} > + > static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss) > { > struct tcp_sock *tp = tcp_sk(sk); > > if (unmark_loss) { > + LIST_HEAD(restored); > struct sk_buff *skb; > > skb_rbtree_walk(skb, &sk->tcp_rtx_queue) { > + if (TCP_SKB_CB(skb)->sacked & TCPCB_LOST) > + list_move_tail(&skb->tcp_tsorted_anchor, > + &restored); > TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST; > } > + if (!list_empty(&restored)) { > + struct list_head *pos = &tp->tsorted_sent_queue; > + > + /* Ensure lost skbs are added in transmission order */ > + list_sort(NULL, &restored, tcp_rack_skb_cmp); > + while (!list_empty(&restored)) { > + struct list_head *entry = restored.next; > + > + while (pos->next != &tp->tsorted_sent_queue && > + !tcp_rack_skb_cmp(NULL, pos->next, > + entry)) > + pos = pos->next; > + list_move(entry, pos); > + pos = entry; > + } > + } > tp->lost_out = 0; > tcp_clear_all_retrans_hints(tp); > } ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss 2026-10-06 6:42 ` Eric Dumazet @ 2026-10-06 11:23 ` Eric Dumazet 2026-10-06 14:13 ` Neal Cardwell 0 siblings, 1 reply; 7+ messages in thread From: Eric Dumazet @ 2026-10-06 11:23 UTC (permalink / raw) To: nramaswamy, netdev Cc: Neal Cardwell, Kuniyuki Iwashima, Yuchung Cheng, Jiayuan Chen, David S. Miller, Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan, linux-kernel, linux-kselftest Le mar. 6 oct. 2026 à 08:42, Eric Dumazet <edumazet@kernel.org> a écrit : > > > > On 10/6/26 07:15, nramaswamy@openai.com wrote: > > From: Neil Ramaswamy <nramaswamy@openai.com> > > > > Partial undo can clear the TCPCB_LOST flag on segments already removed > > from RACK's list, which prevents subsequent RACK loss detection, leading > > to segments only being retransmitted after the RTO. Restoring them to > > the RACK list as part of partial undo makes sure we can properly > > reconsider them for fast retransmission in the future. > > > > To do this, we first sort (by transmission time) the segments whose > > TCPCB_LOST flag is being cleared and then linearly insert those back > > into the RACK list, which is sorted by transmission time. > > > > My investigation started from seeing repeated TCP stalls in prod and the > > mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to > > 96 KiB. It also seems like others have seen similar symptoms before [1]. > > > > [1] > > https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@gmail.com/ > > > > Fixes: 043b87d7599e ("tcp: more efficient RACK loss detection") > > Signed-off-by: Neil Ramaswamy <nramaswamy@openai.com> > > Assisted-by: LLM sparse > > --- > > net/ipv4/tcp_input.c | 35 +++++++++++++++++++++++++++++++++++ > > 1 file changed, 35 insertions(+) > > > > diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c > > index 92bc60716f33..51f04c7474dc 100644 > > --- a/net/ipv4/tcp_input.c > > +++ b/net/ipv4/tcp_input.c > > @@ -69,6 +69,7 @@ > > #include <linux/module.h> > > #include <linux/sysctl.h> > > #include <linux/kernel.h> > > +#include <linux/list_sort.h> > > #include <linux/prefetch.h> > > #include <linux/bitops.h> > > #include <net/dst.h> > > @@ -2840,16 +2841,50 @@ static void DBGUNDO(struct sock *sk, const char *msg) > > #endif > > } > > > > +static int tcp_rack_skb_cmp(void *priv, const struct list_head *a, > > + const struct list_head *b) > > +{ > > + const struct sk_buff *skb_a = list_entry(a, struct sk_buff, > > + tcp_tsorted_anchor); > > + const struct sk_buff *skb_b = list_entry(b, struct sk_buff, > > + tcp_tsorted_anchor); > > + > > + return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a), > > + tcp_skb_timestamp_us(skb_b), > > + TCP_SKB_CB(skb_a)->end_seq, > > + TCP_SKB_CB(skb_b)->end_seq); > > Patch LGTM, but perhaps we could avoid the two div_u64() calls here. > > > return tcp_skb_sent_after(skb_a->skb_mstamp_ns, > skb_b->skb_mstamp_ns, > TCP_SKB_CB(skb_a)->end_seq, > TCP_SKB_CB(skb_b)->end_seq); > > > > > +} > > + > > static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss) > > { > > struct tcp_sock *tp = tcp_sk(sk); > > > > if (unmark_loss) { > > + LIST_HEAD(restored); > > struct sk_buff *skb; > > > > skb_rbtree_walk(skb, &sk->tcp_rtx_queue) { > > + if (TCP_SKB_CB(skb)->sacked & TCPCB_LOST) > > + list_move_tail(&skb->tcp_tsorted_anchor, > > + &restored); > > TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST; > > } > > + if (!list_empty(&restored)) { > > + struct list_head *pos = &tp->tsorted_sent_queue; > > + > > + /* Ensure lost skbs are added in transmission order */ > > + list_sort(NULL, &restored, tcp_rack_skb_cmp); Second thoughts (sorry I am currently attending LPC in Prague, little time this week for reviews) pw-bot: cr I am quite concerned about calling list_sort() inside tcp_undo_cwnd_reduction(). In high-BDP environments (like 100G/200G/400G fabrics), a flow can easily have thousands of packets in flight. If a loss burst affects, say, 4000 skbs: 1. list_sort() will perform ~45,000 comparisons. 2. With indirect function calls (retpoline cost) and cache misses across 1MB+ of sk_buff structures, this can freeze the CPU in SoftIRQ under the socket lock for multiple milliseconds.. We could avoid list_sort() completely by maintaining a second list, e.g. tp->tsorted_lost_queue. Consider the following: 1. When RACK detects loss in tcp_rack_detect_loss(), instead of calling list_del_init(&skb->tcp_tsorted_anchor), move the skb to the lost queue: --- a/net/ipv4/tcp_recovery.c +++ b/net/ipv4/tcp_recovery.c @@ -101,7 +101,7 @@ static void tcp_rack_detect_loss(struct sock *sk, u32 *reo_timeout) remaining = tcp_rack_skb_timeout(tp, skb, reo_wnd); if (remaining <= 0) { tcp_mark_skb_lost(sk, skb); - list_del_init(&skb->tcp_tsorted_anchor); + list_move_tail(&skb->tcp_tsorted_anchor, &tp->tsorted_lost_queue); } else { Because RACK scans tsorted_sent_queue from oldest to newest, packets appended to tsorted_lost_queue are already guaranteed to be in strictly increasing departure timestamp order. 2. When a packet is retransmitted, tcp_update_skb_after_send() already calls: list_move_tail(&skb->tcp_tsorted_anchor, &tp->tsorted_sent_queue); So retransmitted packets automatically leave tsorted_lost_queue and re-enter the tail of tsorted_sent_queue in O(1). 3. When a partial undo occurs in tcp_undo_cwnd_reduction(), any un-retransmitted packets remaining in tsorted_lost_queue are already in sorted order. Since these lost packets were originally sent prior to any subsequent transmissions in tsorted_sent_queue, restoring them can often be done with an O(1) list splice at the head: if (unmark_loss && !list_empty(&tp->tsorted_lost_queue)) list_splice_init(&tp->tsorted_lost_queue, &tp->tsorted_sent_queue); > > + while (!list_empty(&restored)) { > > + struct list_head *entry = restored.next; > > + > > + while (pos->next != &tp->tsorted_sent_queue && > > + !tcp_rack_skb_cmp(NULL, pos->next, > > + entry)) > > + pos = pos->next; > > + list_move(entry, pos); > > + pos = entry; > > + } > > + } > > tp->lost_out = 0; > > tcp_clear_all_retrans_hints(tp); > > } > ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss 2026-10-06 11:23 ` Eric Dumazet @ 2026-10-06 14:13 ` Neal Cardwell 0 siblings, 0 replies; 7+ messages in thread From: Neal Cardwell @ 2026-10-06 14:13 UTC (permalink / raw) To: edumazet Cc: davem, horms, jiayuan.chen, kuba, kuniyu, linux-kernel, linux-kselftest, ncardwell, netdev, nramaswamy, pabeni, shuah, ycheng From: Neal Cardwell <ncardwell@google.com> Hi Neil, Thanks for your TCP patch and packetdrill test! I agree this is worth improving. I share Eric's concerns about the performance costs of this current TCP patch. And Yuchung and I were chatting about this out of band a few days ago; we were both concerned about the performance costs of that approach. Yuchung and I were thinking that perhaps a better trade-off would be to leverage the fact that most of the time when there is an undo there are no lost packets in the scoreboard that were retransmitted, in which case all skbs that are marked as lost have a transmission order that is the same as their sequence order. That means that in most cases in the existing skb_rbtree_walk iterating through the tcp_rtx_queue we can do a kind of fast "merge sort" of the (already time-sorted) lost skbs in tcp_rtx_queue into the (already time-sorted) skbs in tp->tsorted_sent_queue. This should allow us to keep the cost of tcp_undo_cwnd_reduction() as O(packets_out) (without any extra list_sort() cost), while still integrating all the lost packets back into tp->tsorted_sent_queue in the vast majority of cases. And in the rare case when somehow there is an undo while there are lost packets that were EVER_RETRANS, it would be OK to fall back to an RTO for this rare case. That is no worse than the behavior we have been living with for a long time. And with this approach we would not need to add any per-socket or per-skb state. Along those lines, what do you think about something like the following (which compiles and passes your nice packetdrill test): diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c index 92bc60716..cd7c3e408 100644 --- a/net/ipv4/tcp_input.c +++ b/net/ipv4/tcp_input.c @@ -2840,15 +2840,62 @@ static void DBGUNDO(struct sock *sk, const char *msg) #endif } +/* Is skb @a after skb @b in tp->tsorted_sent_queue (send) order? */ +static bool tcp_tsorted_after(const struct list_head *a, + const struct list_head *b) +{ + const struct sk_buff *skb_a = list_entry(a, struct sk_buff, + tcp_tsorted_anchor); + const struct sk_buff *skb_b = list_entry(b, struct sk_buff, + tcp_tsorted_anchor); + + return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a), + tcp_skb_timestamp_us(skb_b), + TCP_SKB_CB(skb_a)->end_seq, + TCP_SKB_CB(skb_b)->end_seq); +} + +/* Link skb back into tp->tsorted_sent_queue in send order, at or after + * pos, and advance pos to it. Leave skb alone if it was sent before + * pos. As pos only moves forward, each call is amortized O(1). + */ +static void tcp_tsorted_relink_skb(struct tcp_sock *tp, struct sk_buff *skb, + struct list_head **pos_ptr) +{ + struct list_head *head = &tp->tsorted_sent_queue; + struct list_head *node = &skb->tcp_tsorted_anchor; + struct list_head *pos = *pos_ptr; + + if (pos != head && !tcp_tsorted_after(node, pos)) + return; + while (pos->next != head && !tcp_tsorted_after(pos->next, node)) + pos = pos->next; + if (pos != node) /* not already linked in place */ + list_move(node, pos); + *pos_ptr = node; +} + static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss) { struct tcp_sock *tp = tcp_sk(sk); if (unmark_loss) { + struct list_head *pos = &tp->tsorted_sent_queue; struct sk_buff *skb; skb_rbtree_walk(skb, &sk->tcp_rtx_queue) { - TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST; + u8 sacked = TCP_SKB_CB(skb)->sacked; + + TCP_SKB_CB(skb)->sacked = sacked & ~TCPCB_LOST; + /* RACK unlinked the skbs it marked lost. Skbs never + * retransmitted keep their original send times, which + * increase with sequence, so in one forward pass we + * relink them all. For the rare case of undo after + * lost retransmissions, we will fall back to RTO. + */ + if ((sacked & (TCPCB_LOST | TCPCB_EVER_RETRANS)) == + TCPCB_LOST) + tcp_tsorted_relink_skb(tp, skb, &pos); } tp->lost_out = 0; tcp_clear_all_retrans_hints(tp); ^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH net v2 2/2] selftests: net: packetdrill: test RACK after partial undo 2026-10-06 5:15 [PATCH net v2 0/2] tcp: preserve RACK tracking across partial undo nramaswamy 2026-10-06 5:15 ` [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss nramaswamy @ 2026-10-06 5:15 ` nramaswamy 2026-10-06 7:04 ` Eric Dumazet 1 sibling, 1 reply; 7+ messages in thread From: nramaswamy @ 2026-10-06 5:15 UTC (permalink / raw) To: netdev Cc: Neil Ramaswamy, Eric Dumazet, Neal Cardwell, Kuniyuki Iwashima, Yuchung Cheng, Jiayuan Chen, David S. Miller, Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan, linux-kernel, linux-kselftest From: Neil Ramaswamy <nramaswamy@openai.com> Reproduces a partial undo bug where a segment is unmarked as lost but is never returned to RACK's timestamp sorted list, which makes it ineligible for future fast retransmission. Signed-off-by: Neil Ramaswamy <nramaswamy@openai.com> Assisted-by: LLM sparse --- ...tcp_partial_undo-restores-to-rack-list.pkt | 54 +++++++++++++++++++ 1 file changed, 54 insertions(+) create mode 100644 tools/testing/selftests/net/packetdrill/tcp_partial_undo-restores-to-rack-list.pkt diff --git a/tools/testing/selftests/net/packetdrill/tcp_partial_undo-restores-to-rack-list.pkt b/tools/testing/selftests/net/packetdrill/tcp_partial_undo-restores-to-rack-list.pkt new file mode 100644 index 000000000000..c07a2f8a5cbd --- /dev/null +++ b/tools/testing/selftests/net/packetdrill/tcp_partial_undo-restores-to-rack-list.pkt @@ -0,0 +1,54 @@ +// SPDX-License-Identifier: GPL-2.0 +// +// Test that a segment unmarked as lost during partial undo is eligible +// for future fast retransmission. + +`./defaults.sh` + +// Establish a connection with a 100 ms RTT and a 1000-byte payload MSS. +// Linux subtracts the 12-byte timestamp options from the advertised MSS. + 0 socket(..., SOCK_STREAM, IPPROTO_TCP) = 3 + +0 setsockopt(3, SOL_SOCKET, SO_REUSEADDR, [1], 4) = 0 + +0 bind(3, ..., ...) = 0 + +0 listen(3, 1) = 0 + + +.1 < S 0:0(0) win 20000 <mss 1012,sackOK,TS val 1000 ecr 0> + +0 > S. 0:0(0) ack 1 <mss 1460,sackOK,TS val 100 ecr 1000> + +.1 < . 1:1(0) ack 1 win 20000 <nop,nop,TS val 1100 ecr 100> + +0 accept(3, ..., ...) = 4 + +// Send A, B, C and D, then E 47 ms after D. + +.01 write(4, ..., 1000) = 1000 + +0 > P. 1:1001(1000) ack 1 <nop,nop,TS val 210 ecr 1100> ++.001 write(4, ..., 1000) = 1000 + +0 > P. 1001:2001(1000) ack 1 <...> ++.001 write(4, ..., 1000) = 1000 + +0 > P. 2001:3001(1000) ack 1 <...> ++.001 write(4, ..., 1000) = 1000 + +0 > P. 3001:4001(1000) ack 1 <...> ++.047 write(4, ..., 1000) = 1000 + +0 > P. 4001:5001(1000) ack 1 <...> + +// SACK C and E together. RACK marks A, B and D lost. D is old enough +// to be retransmitted, but this ACK reports only two newly delivered +// segments, allowing A and B to be retransmitted while D waits. + +.12 < . 1:1(0) ack 1 win 20000 <TS val 1280 ecr 100,sack 4001:5001 2001:3001> + +0 > P. 1:1001(1000) ack 1 <...> + +0 > P. 1001:2001(1000) ack 1 <...> + +0 %{ +assert tcpi_ca_state == TCP_CA_Recovery, tcpi_ca_state +assert tcpi_lost == 3, tcpi_lost +assert tcpi_retrans == 2, tcpi_retrans +}% + +// Deliver retransmitted B and then the original A. D stays missing, +// and we ACK through C with A's original timestamp, which is before +// retransmission started. This triggers partial undo, but D should +// remain eligible for fast retransmission (critically, the timeout +// retransmission counter should be 0). + +.07 < . 1:1(0) ack 3001 win 20000 <TS val 1350 ecr 210,sack 4001:5001> + +0~+.05 > P. 3001:4001(1000) ack 1 <nop,nop,TS val 450 ecr 1350> + +0 %{ assert tcpi_retransmits == 0, tcpi_retransmits }% + +// Acknowledge all five segments. + +.01 < . 1:1(0) ack 5001 win 20000 <nop,nop,TS val 1360 ecr 450> -- 2.55.0 ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH net v2 2/2] selftests: net: packetdrill: test RACK after partial undo 2026-10-06 5:15 ` [PATCH net v2 2/2] selftests: net: packetdrill: test RACK after partial undo nramaswamy @ 2026-10-06 7:04 ` Eric Dumazet 0 siblings, 0 replies; 7+ messages in thread From: Eric Dumazet @ 2026-10-06 7:04 UTC (permalink / raw) To: nramaswamy, netdev Cc: Neal Cardwell, Kuniyuki Iwashima, Yuchung Cheng, Jiayuan Chen, David S. Miller, Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan, linux-kernel, linux-kselftest On 10/6/26 07:15, nramaswamy@openai.com wrote: > From: Neil Ramaswamy <nramaswamy@openai.com> > > Reproduces a partial undo bug where a segment is unmarked as lost but is > never returned to RACK's timestamp sorted list, which makes it ineligible > for future fast retransmission. > > Signed-off-by: Neil Ramaswamy <nramaswamy@openai.com> > Assisted-by: LLM sparse > --- Thanks for this test! Reviewed-by: Eric Dumazet <edumazet@kernel.org> ^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-06 14:13 UTC | newest] Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-10-06 5:15 [PATCH net v2 0/2] tcp: preserve RACK tracking across partial undo nramaswamy 2026-10-06 5:15 ` [PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss nramaswamy 2026-10-06 6:42 ` Eric Dumazet 2026-10-06 11:23 ` Eric Dumazet 2026-10-06 14:13 ` Neal Cardwell 2026-10-06 5:15 ` [PATCH net v2 2/2] selftests: net: packetdrill: test RACK after partial undo nramaswamy 2026-10-06 7:04 ` Eric Dumazet
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®