From: Eric Dumazet <edumazet@kernel.org>
To: nramaswamy@openai.com, netdev@vger.kernel.org
Cc: Neal Cardwell <ncardwell@google.com>,
Neal Cardwell <ncardwell.sw@gmail.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Yuchung Cheng <ycheng@google.com>,
Jiayuan Chen <jiayuan.chen@linux.dev>,
"David S. Miller" <davem@davemloft.net>,
Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
Simon Horman <horms@kernel.org>, Shuah Khan <shuah@kernel.org>,
linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org
Subject: Re: [PATCH net v3 1/3] tcp: restore RACK list membership when undoing loss
Date: Fri, 9 Oct 2026 07:07:30 +0200 [thread overview]
Message-ID: <aceb97a1-2493-437e-926d-ba8fe68bbf48@kernel.org> (raw)
In-Reply-To: <e11ac4c9231e2ff7b288a10b48baa5378f2e6cc1.1791506907.git.nramaswamy@openai.com>
On 10/9/26 07:00, nramaswamy@openai.com wrote:
> From: Neil Ramaswamy <nramaswamy@openai.com>
>
> Partial undo can clear the TCPCB_LOST flag on segments already removed
> from RACK's list, which prevents subsequent RACK loss detection, leading
> to segments only being retransmitted after the RTO.
>
> This patch uses the implementation provided by Neal Cardwell to linearly
> insert lost segments that have not been retransmitted back into the RACK
> list. The core observation is that segments that are lost but not ever
> retransmitted are already sorted by their transmission timestamp, which
> allows us to insert them into RACK's list linearly, without needing to
> pre-sort them.
>
> Segments with TCPCB_EVER_RETRANS are left for RTO recovery. Their
> transmission order can differ from their sequence order, so reinserting
> them would require addiitonal work (e.g. sorting) before reinsertion into
> the RACK list. We assume that this case is rare, and allow those segments
> to be transmitted by the RTO.
>
> My investigation started from seeing repeated TCP stalls in prod and the
> mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to
> 96 KiB. It also seems like others have seen similar symptoms before [1].
>
> [1]
> https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@gmail.com/
>
> Fixes: 043b87d7599e ("tcp: more efficient RACK loss detection")
> Suggested-by: Neal Cardwell <ncardwell@google.com>
> Suggested-by: Yuchung Cheng <ycheng@google.com>
> Link: https://lore.kernel.org/netdev/20261006141300.1722466-1-ncardwell.sw@gmail.com/
> Signed-off-by: Neil Ramaswamy <nramaswamy@openai.com>
> Assisted-by: LLM sparse
> ---
> net/ipv4/tcp_input.c | 50 +++++++++++++++++++++++++++++++++++++++++++-
> 1 file changed, 49 insertions(+), 1 deletion(-)
>
> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> index 92bc60716f33..d0a1e8899c4b 100644
> --- a/net/ipv4/tcp_input.c
> +++ b/net/ipv4/tcp_input.c
> @@ -2840,15 +2840,63 @@ static void DBGUNDO(struct sock *sk, const char *msg)
> #endif
> }
>
> +/* Is skb @a after skb @b in tp->tsorted_sent_queue (send) order? */
> +static bool tcp_tsorted_after(const struct list_head *a,
> + const struct list_head *b)
> +{
> + const struct sk_buff *skb_a = list_entry(a, struct sk_buff,
> + tcp_tsorted_anchor);
> + const struct sk_buff *skb_b = list_entry(b, struct sk_buff,
> + tcp_tsorted_anchor);
> +
> + return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a),
> + tcp_skb_timestamp_us(skb_b),
> + TCP_SKB_CB(skb_a)->end_seq,
> + TCP_SKB_CB(skb_b)->end_seq);
pw-bot: cr
You ignored my initial feedback.
https://lore.kernel.org/netdev/627cd81c-9824-4d1b-ae47-f7f929a96596@kernel.org/
Thanks.
> +}
> +
> +/* Link skb back into tp->tsorted_sent_queue in send order, at or after
> + * pos, and advance pos to it. Leave skb alone if it was sent before
> + * pos. During undo, all relink calls together traverse the RACK list
> + * at most once, as pos only moves forward.
> + */
> +static void tcp_tsorted_relink_skb(struct tcp_sock *tp, struct sk_buff *skb,
> + struct list_head **pos_ptr)
> +{
> + struct list_head *head = &tp->tsorted_sent_queue;
> + struct list_head *node = &skb->tcp_tsorted_anchor;
> + struct list_head *pos = *pos_ptr;
> +
> + if (pos != head && !tcp_tsorted_after(node, pos))
> + return;
> + while (pos->next != head && !tcp_tsorted_after(pos->next, node))
> + pos = pos->next;
> + if (pos != node) /* not already linked in place */
> + list_move(node, pos);
> + *pos_ptr = node;
> +}
> +
> static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss)
> {
> struct tcp_sock *tp = tcp_sk(sk);
>
> if (unmark_loss) {
> + struct list_head *pos = &tp->tsorted_sent_queue;
> struct sk_buff *skb;
>
> skb_rbtree_walk(skb, &sk->tcp_rtx_queue) {
> - TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST;
> + u8 sacked = TCP_SKB_CB(skb)->sacked;
> +
> + TCP_SKB_CB(skb)->sacked = sacked & ~TCPCB_LOST;
> + /* RACK unlinked the skbs it marked lost. Skbs never
> + * retransmitted keep their original send times, which
> + * increase with sequence, so in one forward pass we
> + * relink them all. For the rare case of undo after
> + * lost retransmissions, we will fall back to RTO.
> + */
> + if ((sacked & (TCPCB_LOST | TCPCB_EVER_RETRANS)) ==
> + TCPCB_LOST)
> + tcp_tsorted_relink_skb(tp, skb, &pos);
> }
> tp->lost_out = 0;
> tcp_clear_all_retrans_hints(tp);
next prev parent reply other threads:[~2026-10-09 5:07 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 5:00 [PATCH net v3 0/3] tcp: preserve RACK tracking across partial undo nramaswamy
2026-10-09 5:00 ` [PATCH net v3 1/3] tcp: restore RACK list membership when undoing loss nramaswamy
2026-10-09 5:07 ` Eric Dumazet [this message]
2026-10-09 5:00 ` [PATCH net v3 2/3] selftests: net: packetdrill: test RACK after partial undo nramaswamy
2026-10-09 5:17 ` Eric Dumazet
2026-10-09 5:00 ` [PATCH net v3 3/3] selftests: net: packetdrill: test RTO fallback " nramaswamy
2026-10-09 5:47 ` Eric Dumazet
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aceb97a1-2493-437e-926d-ba8fe68bbf48@kernel.org \
--to=edumazet@kernel.org \
--cc=davem@davemloft.net \
--cc=horms@kernel.org \
--cc=jiayuan.chen@linux.dev \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=ncardwell.sw@gmail.com \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=nramaswamy@openai.com \
--cc=pabeni@redhat.com \
--cc=shuah@kernel.org \
--cc=ycheng@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®