From: Petr Tesarik <ptesarik@suse.com>
To: Paolo Abeni <pabeni@redhat.com>,
"David S. Miller" <davem@davemloft.net>,
Eric Dumazet <edumazet@google.com>,
Neal Cardwell <ncardwell@google.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
netdev@vger.kernel.org (open list:NETWORKING [TCP])
Cc: David Ahern <dsahern@kernel.org>,
Jakub Kicinski <kuba@kernel.org>,
linux-kernel@vger.kernel.org (open list),
Petr Tesarik <ptesarik@suse.com>
Subject: [PATCH net v2 1/2] tcp_metrics: set congestion window clamp from the dst entry
Date: Fri, 20 Jun 2025 14:56:43 +0200 [thread overview]
Message-ID: <20250620125644.1045603-2-ptesarik@suse.com> (raw)
In-Reply-To: <20250620125644.1045603-1-ptesarik@suse.com>
If RTAX_CWND is locked, always initialize tp->snd_cwnd_clamp from the
corresponding dst entry.
Note that an unlocked RTAX_CWND does not have any effect on the kernel.
This behavior is even documented in the manual page of ip-route(8):
cwnd NUMBER (Linux 2.3.15+ only)
the clamp for congestion window. It is ignored if the lock
flag is not used.
An unlocked RTAX_CWND was updated by tcp_update_metrics() until v3.6. Since
then, only the newly introduced TCP metric (TCP_METRIC_CWND) has been
updated, rendering unlocked RTAX_CWND useless.
TCP metrics are updated after a TCP connection finishes. If there are no
metrics for a given destination when a new connection is created, default
values are used instead.
This means there are two issues with setting tp->snd_cwnd_clamp from the
TCP metric:
1. If the cwnd option is changed in the routing table, the new value is not
used for new connections as long as there is a cached TCP metric for the
destination. An existing cached metric is not updated from the routing
table unless it has seen no update for longer than TCP_METRICS_TIMEOUT
(1 hour).
2. After evicting the corresponding cached metric, the new value from the
routing table is still not used for new connections until one connection
finishes, and a new cached entry is created.
As a result, the following shenanigan is required to set a new locked cwnd
clamp:
- update the route (``ip route replace ... cwnd lock $value``)
- flush any existing TCP metric entry (``ip tcp_metrics flush $dest``)
- create and finish a dummy connection to the destination to create a TCP
metric entry with the new value
- *next* connection to this destination will use the new value
The above does not seem to be intentional.
NB there is also an initcwnd route parameter (RTAX_INITCWND) to set the
initial size of the congestion window; this patch does not change anything
about the handling of that parameter.
Fixes: 51c5d0c4b169 ("tcp: Maintain dynamic metrics in local cache.")
Signed-off-by: Petr Tesarik <ptesarik@suse.com>
---
net/ipv4/tcp_metrics.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/net/ipv4/tcp_metrics.c b/net/ipv4/tcp_metrics.c
index 4251670e328c..dd8f3457bd72 100644
--- a/net/ipv4/tcp_metrics.c
+++ b/net/ipv4/tcp_metrics.c
@@ -477,6 +477,9 @@ void tcp_init_metrics(struct sock *sk)
if (!dst)
goto reset;
+ if (dst_metric_locked(dst, RTAX_CWND))
+ tp->snd_cwnd_clamp = dst_metric(dst, RTAX_CWND);
+
rcu_read_lock();
tm = tcp_get_metrics(sk, dst, false);
if (!tm) {
@@ -484,9 +487,6 @@ void tcp_init_metrics(struct sock *sk)
goto reset;
}
- if (tcp_metric_locked(tm, TCP_METRIC_CWND))
- tp->snd_cwnd_clamp = tcp_metric_get(tm, TCP_METRIC_CWND);
-
val = READ_ONCE(net->ipv4.sysctl_tcp_no_ssthresh_metrics_save) ?
0 : tcp_metric_get(tm, TCP_METRIC_SSTHRESH);
if (val) {
--
2.49.0
next prev parent reply other threads:[~2025-06-20 12:57 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-06-20 12:56 [PATCH net v2 0/2] tcp_metrics: fix hanlding of route options Petr Tesarik
2025-06-20 12:56 ` Petr Tesarik [this message]
2025-06-20 12:56 ` [PATCH net v2 2/2] tcp_metrics: use ssthresh value from dst if there is no metrics Petr Tesarik
2025-06-20 13:24 ` [PATCH net v2 0/2] tcp_metrics: fix hanlding of route options Eric Dumazet
2025-06-23 7:36 ` Petr Tesarik
2025-06-23 11:44 ` Eric Dumazet
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20250620125644.1045603-2-ptesarik@suse.com \
--to=ptesarik@suse.com \
--cc=davem@davemloft.net \
--cc=dsahern@kernel.org \
--cc=edumazet@google.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®