mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Eric Dumazet <edumazet@kernel.org>
To: nramaswamy@openai.com, netdev@vger.kernel.org
Cc: Neal Cardwell <ncardwell@google.com>,
	Neal Cardwell <ncardwell.sw@gmail.com>,
	Kuniyuki Iwashima <kuniyu@google.com>,
	Yuchung Cheng <ycheng@google.com>,
	Jiayuan Chen <jiayuan.chen@linux.dev>,
	"David S. Miller" <davem@davemloft.net>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>, Shuah Khan <shuah@kernel.org>,
	linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org
Subject: Re: [PATCH net v3 1/3] tcp: restore RACK list membership when undoing loss
Date: Fri, 9 Oct 2026 07:07:30 +0200	[thread overview]
Message-ID: <aceb97a1-2493-437e-926d-ba8fe68bbf48@kernel.org> (raw)
In-Reply-To: <e11ac4c9231e2ff7b288a10b48baa5378f2e6cc1.1791506907.git.nramaswamy@openai.com>



On 10/9/26 07:00, nramaswamy@openai.com wrote:
> From: Neil Ramaswamy <nramaswamy@openai.com>
> 
> Partial undo can clear the TCPCB_LOST flag on segments already removed
> from RACK's list, which prevents subsequent RACK loss detection, leading
> to segments only being retransmitted after the RTO.
> 
> This patch uses the implementation provided by Neal Cardwell to linearly
> insert lost segments that have not been retransmitted back into the RACK
> list. The core observation is that segments that are lost but not ever
> retransmitted are already sorted by their transmission timestamp, which
> allows us to insert them into RACK's list linearly, without needing to
> pre-sort them.
> 
> Segments with TCPCB_EVER_RETRANS are left for RTO recovery. Their
> transmission order can differ from their sequence order, so reinserting
> them would require addiitonal work (e.g. sorting) before reinsertion into
> the RACK list. We assume that this case is rare, and allow those segments
> to be transmitted by the RTO.
> 
> My investigation started from seeing repeated TCP stalls in prod and the
> mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to
> 96 KiB. It also seems like others have seen similar symptoms before [1].
> 
> [1]
> https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@gmail.com/
> 
> Fixes: 043b87d7599e ("tcp: more efficient RACK loss detection")
> Suggested-by: Neal Cardwell <ncardwell@google.com>
> Suggested-by: Yuchung Cheng <ycheng@google.com>
> Link: https://lore.kernel.org/netdev/20261006141300.1722466-1-ncardwell.sw@gmail.com/
> Signed-off-by: Neil Ramaswamy <nramaswamy@openai.com>
> Assisted-by: LLM sparse
> ---
>   net/ipv4/tcp_input.c | 50 +++++++++++++++++++++++++++++++++++++++++++-
>   1 file changed, 49 insertions(+), 1 deletion(-)
> 
> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> index 92bc60716f33..d0a1e8899c4b 100644
> --- a/net/ipv4/tcp_input.c
> +++ b/net/ipv4/tcp_input.c
> @@ -2840,15 +2840,63 @@ static void DBGUNDO(struct sock *sk, const char *msg)
>   #endif
>   }
>   
> +/* Is skb @a after skb @b in tp->tsorted_sent_queue (send) order? */
> +static bool tcp_tsorted_after(const struct list_head *a,
> +			      const struct list_head *b)
> +{
> +	const struct sk_buff *skb_a = list_entry(a, struct sk_buff,
> +						 tcp_tsorted_anchor);
> +	const struct sk_buff *skb_b = list_entry(b, struct sk_buff,
> +						 tcp_tsorted_anchor);
> +
> +	return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a),
> +				  tcp_skb_timestamp_us(skb_b),
> +				  TCP_SKB_CB(skb_a)->end_seq,
> +				  TCP_SKB_CB(skb_b)->end_seq);

pw-bot: cr

You ignored my initial feedback.

https://lore.kernel.org/netdev/627cd81c-9824-4d1b-ae47-f7f929a96596@kernel.org/

Thanks.

> +}
> +
> +/* Link skb back into tp->tsorted_sent_queue in send order, at or after
> + * pos, and advance pos to it. Leave skb alone if it was sent before
> + * pos. During undo, all relink calls together traverse the RACK list
> + * at most once, as pos only moves forward.
> + */
> +static void tcp_tsorted_relink_skb(struct tcp_sock *tp, struct sk_buff *skb,
> +				   struct list_head **pos_ptr)
> +{
> +	struct list_head *head = &tp->tsorted_sent_queue;
> +	struct list_head *node = &skb->tcp_tsorted_anchor;
> +	struct list_head *pos = *pos_ptr;
> +
> +	if (pos != head && !tcp_tsorted_after(node, pos))
> +		return;
> +	while (pos->next != head && !tcp_tsorted_after(pos->next, node))
> +		pos = pos->next;
> +	if (pos != node)		/* not already linked in place */
> +		list_move(node, pos);
> +	*pos_ptr = node;
> +}
> +
>   static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss)
>   {
>   	struct tcp_sock *tp = tcp_sk(sk);
>   
>   	if (unmark_loss) {
> +		struct list_head *pos = &tp->tsorted_sent_queue;
>   		struct sk_buff *skb;
>   
>   		skb_rbtree_walk(skb, &sk->tcp_rtx_queue) {
> -			TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST;
> +			u8 sacked = TCP_SKB_CB(skb)->sacked;
> +
> +			TCP_SKB_CB(skb)->sacked = sacked & ~TCPCB_LOST;
> +			/* RACK unlinked the skbs it marked lost. Skbs never
> +			 * retransmitted keep their original send times, which
> +			 * increase with sequence, so in one forward pass we
> +			 * relink them all. For the rare case of undo after
> +			 * lost retransmissions, we will fall back to RTO.
> +			 */
> +			if ((sacked & (TCPCB_LOST | TCPCB_EVER_RETRANS)) ==
> +			    TCPCB_LOST)
> +				tcp_tsorted_relink_skb(tp, skb, &pos);
>   		}
>   		tp->lost_out = 0;
>   		tcp_clear_all_retrans_hints(tp);


  reply	other threads:[~2026-10-09  5:07 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09  5:00 [PATCH net v3 0/3] tcp: preserve RACK tracking across partial undo nramaswamy
2026-10-09  5:00 ` [PATCH net v3 1/3] tcp: restore RACK list membership when undoing loss nramaswamy
2026-10-09  5:07   ` Eric Dumazet [this message]
2026-10-09  5:00 ` [PATCH net v3 2/3] selftests: net: packetdrill: test RACK after partial undo nramaswamy
2026-10-09  5:17   ` Eric Dumazet
2026-10-09  5:00 ` [PATCH net v3 3/3] selftests: net: packetdrill: test RTO fallback " nramaswamy
2026-10-09  5:47   ` Eric Dumazet

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aceb97a1-2493-437e-926d-ba8fe68bbf48@kernel.org \
    --to=edumazet@kernel.org \
    --cc=davem@davemloft.net \
    --cc=horms@kernel.org \
    --cc=jiayuan.chen@linux.dev \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=ncardwell.sw@gmail.com \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=nramaswamy@openai.com \
    --cc=pabeni@redhat.com \
    --cc=shuah@kernel.org \
    --cc=ycheng@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®