mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next 0/5] mptcp: misc improvements for v7.4
@ 2026-09-26 15:30 Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 1/5] mptcp: remove thmac from subflow ctx Matthieu Baerts (NGI0)
                   ` (4 more replies)
  0 siblings, 5 replies; 10+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-26 15:30 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0),
	Quanye Yang, Hangbin Liu, Shuah Khan, linux-kselftest

Here are some unrelated improvements for MPTCP and its selftests:

- Patch 1: reduce struct mptcp_subflow_context size: no need to store
  the truncated HMAC for the whole connection.

- Patches 2-3: shrink struct mptcp_options_received used to parse
  incoming MPTCP options, from 136 to 72 bytes on x86_64.

- Patches 4-5: convert IPTables to NFTables in MPTCP selftests.

Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
Hangbin Liu (2):
      selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh
      selftests: mptcp: convert iptables to nftables for mptcp_join.sh

Matthieu Baerts (NGI0) (1):
      mptcp: remove thmac from subflow ctx

Quanye Yang (2):
      mptcp: split FASTCLOSE key from rcvr_key
      mptcp: shrink struct mptcp_options_received

 net/mptcp/options.c                                |   6 +-
 net/mptcp/protocol.h                               |  45 ++++---
 net/mptcp/subflow.c                                |  18 +--
 tools/testing/selftests/net/mptcp/config           |   4 +-
 tools/testing/selftests/net/mptcp/mptcp_join.sh    | 149 ++++++++-------------
 tools/testing/selftests/net/mptcp/mptcp_lib.sh     |   2 +-
 tools/testing/selftests/net/mptcp/mptcp_sockopt.sh |  58 ++++----
 7 files changed, 136 insertions(+), 146 deletions(-)
---
base-commit: 42a9fb3382fc2573e92f41d203b095d9a372cfc9
change-id: 20260925-net-next-mptcp-misc-feat-7-4-3252e498c334

Best regards,
--  
Matthieu Baerts (NGI0) <matttbe@kernel.org>


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH net-next 1/5] mptcp: remove thmac from subflow ctx
  2026-09-26 15:30 [PATCH net-next 0/5] mptcp: misc improvements for v7.4 Matthieu Baerts (NGI0)
@ 2026-09-26 15:30 ` Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 2/5] mptcp: split FASTCLOSE key from rcvr_key Matthieu Baerts (NGI0)
                   ` (3 subsequent siblings)
  4 siblings, 0 replies; 10+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-26 15:30 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0)

This entry is only used in subflow_finish_connect().

Instead, use the original value from mp_opt, and pass it to
subflow_thmac_valid() to do the validation with the given truncated
hmac.

While at it, rename the variables in subflow_thmac_valid() to avoid
confusions about the received one vs the expected one.

Reviewed-by: Geliang Tang <geliang@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
 net/mptcp/protocol.h |  1 -
 net/mptcp/subflow.c  | 18 +++++++++---------
 2 files changed, 9 insertions(+), 10 deletions(-)

diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h
index 0384d6a023f9..dd9957d334b6 100644
--- a/net/mptcp/protocol.h
+++ b/net/mptcp/protocol.h
@@ -593,7 +593,6 @@ struct mptcp_subflow_context {
 	bool	fully_established;  /* path validated */
 	u32	lent_mem_frag;
 	u32	remote_nonce;
-	u64	thmac;
 	u32	local_nonce;
 	u32	remote_token;
 	union {
diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
index f0a6725d2c37..14aa82647c3d 100644
--- a/net/mptcp/subflow.c
+++ b/net/mptcp/subflow.c
@@ -408,20 +408,21 @@ static struct dst_entry *subflow_v6_route_req(const struct sock *sk,
 #endif
 
 /* validate received truncated hmac and create hmac for third ACK */
-static bool subflow_thmac_valid(struct mptcp_subflow_context *subflow)
+static bool subflow_thmac_valid(struct mptcp_subflow_context *subflow,
+				u64 thmac)
 {
 	u8 hmac[SHA256_DIGEST_SIZE];
-	u64 thmac;
+	u64 expected_thmac;
 
 	subflow_generate_hmac(subflow->remote_key, subflow->local_key,
 			      subflow->remote_nonce, subflow->local_nonce,
 			      hmac);
 
-	thmac = get_unaligned_be64(hmac);
-	pr_debug("subflow=%p, token=%u, thmac=%llu, subflow->thmac=%llu\n",
-		 subflow, subflow->token, thmac, subflow->thmac);
+	expected_thmac = get_unaligned_be64(hmac);
+	pr_debug("subflow=%p, token=%u, expected_thmac=%llu, thmac=%llu\n",
+		 subflow, subflow->token, expected_thmac, thmac);
 
-	return thmac == subflow->thmac;
+	return expected_thmac == thmac;
 }
 
 void mptcp_subflow_reset(struct sock *ssk)
@@ -575,14 +576,13 @@ static void subflow_finish_connect(struct sock *sk, const struct sk_buff *skb)
 		}
 
 		subflow->backup = mp_opt.backup;
-		subflow->thmac = mp_opt.thmac;
 		subflow->remote_nonce = mp_opt.nonce;
 		WRITE_ONCE(subflow->remote_id, mp_opt.join_id);
 		pr_debug("subflow=%p, thmac=%llu, remote_nonce=%u backup=%d\n",
-			 subflow, subflow->thmac, subflow->remote_nonce,
+			 subflow, mp_opt.thmac, subflow->remote_nonce,
 			 subflow->backup);
 
-		if (!subflow_thmac_valid(subflow)) {
+		if (!subflow_thmac_valid(subflow, mp_opt.thmac)) {
 			MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_JOINSYNACKMAC);
 			subflow->reset_reason = MPTCP_RST_EMPTCP;
 			goto do_reset;

-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH net-next 2/5] mptcp: split FASTCLOSE key from rcvr_key
  2026-09-26 15:30 [PATCH net-next 0/5] mptcp: misc improvements for v7.4 Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 1/5] mptcp: remove thmac from subflow ctx Matthieu Baerts (NGI0)
@ 2026-09-26 15:30 ` Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 3/5] mptcp: shrink struct mptcp_options_received Matthieu Baerts (NGI0)
                   ` (2 subsequent siblings)
  4 siblings, 0 replies; 10+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-26 15:30 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0), Quanye Yang

From: Quanye Yang <quanyeyang@proton.me>

MP_CAPABLE and MP_FASTCLOSE both stored their key in rcvr_key. Give
FASTCLOSE a dedicated fc_recv_key overlapped with rcvr_key in a union.

This helps shrink struct mptcp_options_received.

Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
 net/mptcp/options.c  | 6 +++---
 net/mptcp/protocol.h | 5 ++++-
 2 files changed, 7 insertions(+), 4 deletions(-)

diff --git a/net/mptcp/options.c b/net/mptcp/options.c
index ce0de02f5a3a..c2f4d7bd080c 100644
--- a/net/mptcp/options.c
+++ b/net/mptcp/options.c
@@ -374,10 +374,10 @@ static void mptcp_parse_option(const struct sk_buff *skb,
 			break;
 
 		ptr += 2;
-		mp_opt->rcvr_key = get_unaligned_be64(ptr);
+		mp_opt->fc_recv_key = get_unaligned_be64(ptr);
 		ptr += 8;
 		mp_opt->suboptions |= OPTION_MPTCP_FASTCLOSE;
-		pr_debug("MP_FASTCLOSE: recv_key=%llu\n", mp_opt->rcvr_key);
+		pr_debug("MP_FASTCLOSE: recv_key=%llu\n", mp_opt->fc_recv_key);
 		break;
 
 	case MPTCPOPT_RST:
@@ -1258,7 +1258,7 @@ bool mptcp_incoming_options(struct sock *sk, struct sk_buff *skb)
 
 	if (unlikely(mp_opt.suboptions != OPTION_MPTCP_DSS)) {
 		if ((mp_opt.suboptions & OPTION_MPTCP_FASTCLOSE) &&
-		    READ_ONCE(msk->local_key) == mp_opt.rcvr_key) {
+		    READ_ONCE(msk->local_key) == mp_opt.fc_recv_key) {
 			WRITE_ONCE(msk->rcv_fastclose, true);
 			mptcp_schedule_work((struct sock *)msk);
 			MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_MPFASTCLOSERX);
diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h
index dd9957d334b6..6b9bffc5992b 100644
--- a/net/mptcp/protocol.h
+++ b/net/mptcp/protocol.h
@@ -146,7 +146,10 @@ static inline bool before64(__u64 seq1, __u64 seq2)
 
 struct mptcp_options_received {
 	u64	sndr_key;
-	u64	rcvr_key;
+	union {
+		u64	rcvr_key;
+		u64	fc_recv_key;
+	};
 	u64	data_ack;
 	u64	data_seq;
 	u32	subflow_seq;

-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH net-next 3/5] mptcp: shrink struct mptcp_options_received
  2026-09-26 15:30 [PATCH net-next 0/5] mptcp: misc improvements for v7.4 Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 1/5] mptcp: remove thmac from subflow ctx Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 2/5] mptcp: split FASTCLOSE key from rcvr_key Matthieu Baerts (NGI0)
@ 2026-09-26 15:30 ` Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 4/5] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh Matthieu Baerts (NGI0)
  2026-09-26 15:30 ` [PATCH net-next 5/5] selftests: mptcp: convert iptables to nftables for mptcp_join.sh Matthieu Baerts (NGI0)
  4 siblings, 0 replies; 10+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-26 15:30 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0), Quanye Yang

From: Quanye Yang <quanyeyang@proton.me>

struct mptcp_options_received is allocated on the stack while parsing
incoming MPTCP options. Several suboptions are mutually exclusive, as
enforced by mptcp_parse_option(), so their payloads can overlap.

Group fields by suboption and place the mutually exclusive payloads in
an anonymous union. Keep DSS and rm_list outside the union: they can
be combined with other suboptions. Move join_id into the MP_JOIN
group, and overlap token, thmac and hmac inside that group.

Further shrinking would require changing the parser so currently
coexisting fields (DSS mapping vs ACK, rm_list, status flags) can
overlap. That adds complexity for little gain, since the outer union is
already dominated by the MP_JOIN / ADD_ADDR members.

This reduces the structure size from 136 to 72 bytes on x86_64.

Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/625
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
 net/mptcp/protocol.h | 45 ++++++++++++++++++++++++++++-----------------
 1 file changed, 28 insertions(+), 17 deletions(-)

diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h
index 6b9bffc5992b..29b03275a59e 100644
--- a/net/mptcp/protocol.h
+++ b/net/mptcp/protocol.h
@@ -145,16 +145,13 @@ static inline bool before64(__u64 seq1, __u64 seq2)
 #define after64(seq2, seq1)	before64(seq1, seq2)
 
 struct mptcp_options_received {
-	u64	sndr_key;
-	union {
-		u64	rcvr_key;
-		u64	fc_recv_key;
+	struct { /* DSS, also used by MP_CAPABLE with data */
+		u64	data_ack;
+		u64	data_seq;
+		u32	subflow_seq;
+		u16	data_len;
+		__sum16	csum;
 	};
-	u64	data_ack;
-	u64	data_seq;
-	u32	subflow_seq;
-	u16	data_len;
-	__sum16	csum;
 	struct_group(status,
 		u16 suboptions;
 		u16 use_map:1,
@@ -170,15 +167,29 @@ struct mptcp_options_received {
 		    deny_join_id0:1,
 		    __unused:2;
 	);
-	u8	join_id;
-	u32	token;
-	u32	nonce;
-	u64	thmac;
-	u8	hmac[MPTCPOPT_HMAC_LEN];
-	struct mptcp_addr_info addr;
 	struct mptcp_rm_list rm_list;
-	u64	ahmac;
-	u64	fail_seq;
+	/* Options below are mutually exclusive, see mptcp_parse_option() */
+	union {
+		struct { /* MP_CAPABLE */
+			u64	sndr_key;
+			u64	rcvr_key;
+		};
+		struct { /* MP_JOIN */
+			u32	nonce;
+			u8	join_id;
+			union {
+				u32	token;			 /* SYN */
+				u64	thmac;			 /* SYN + ACK */
+				u8	hmac[MPTCPOPT_HMAC_LEN]; /* ACK */
+			};
+		};
+		struct { /* ADD_ADDR */
+			struct mptcp_addr_info addr;
+			u64	ahmac;
+		};
+		u64	fail_seq;	/* MP_FAIL */
+		u64	fc_recv_key;	/* MP_FASTCLOSE */
+	};
 };
 
 static inline __be32 mptcp_option(u8 subopt, u8 len, u8 nib, u8 field)

-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH net-next 4/5] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh
  2026-09-26 15:30 [PATCH net-next 0/5] mptcp: misc improvements for v7.4 Matthieu Baerts (NGI0)
                   ` (2 preceding siblings ...)
  2026-09-26 15:30 ` [PATCH net-next 3/5] mptcp: shrink struct mptcp_options_received Matthieu Baerts (NGI0)
@ 2026-09-26 15:30 ` Matthieu Baerts (NGI0)
  2026-09-28  8:00   ` netdev-bot+sashiko
  2026-09-26 15:30 ` [PATCH net-next 5/5] selftests: mptcp: convert iptables to nftables for mptcp_join.sh Matthieu Baerts (NGI0)
  4 siblings, 1 reply; 10+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-26 15:30 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0),
	Hangbin Liu, Shuah Khan, linux-kselftest

From: Hangbin Liu <liuhangbin@kylinos.cn>

During the conversion, we retain the same filter and chain names previously
used by iptables/ip6tables. Counters are not added to accept rules because
the test does not inspect them. After conversion, the generated output
matches the original iptables/ip6tables behavior.

Signed-off-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
To: Shuah Khan <shuah@kernel.org>
Cc: linux-kselftest@vger.kernel.org
---
 tools/testing/selftests/net/mptcp/mptcp_lib.sh     |  2 +-
 tools/testing/selftests/net/mptcp/mptcp_sockopt.sh | 58 ++++++++++++----------
 2 files changed, 34 insertions(+), 26 deletions(-)

diff --git a/tools/testing/selftests/net/mptcp/mptcp_lib.sh b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
index b9d14647f401..e65b4ebee06a 100644
--- a/tools/testing/selftests/net/mptcp/mptcp_lib.sh
+++ b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
@@ -528,7 +528,7 @@ mptcp_lib_check_tools() {
 				exit ${KSFT_SKIP}
 			fi
 			;;
-		"iptables"* | "ip6tables"*)
+		"iptables"* | "ip6tables"* | "nft" | "jq")
 			if ! "${tool}" -V &> /dev/null; then
 				mptcp_lib_pr_skip "Could not run all tests without ${tool}"
 				exit ${KSFT_SKIP}
diff --git a/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh b/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
index e850a87429b6..b4e3eddc9cf2 100755
--- a/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
+++ b/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
@@ -15,8 +15,6 @@ cin=""
 cout=""
 timeout_poll=30
 timeout_test=$((timeout_poll * 2 + 1))
-iptables="iptables"
-ip6tables="ip6tables"
 
 ns1=""
 ns2=""
@@ -49,16 +47,27 @@ add_mark_rules()
 	local ns=$1
 	local m=$2
 
-	local t
-	for t in ${iptables} ${ip6tables}; do
-		# just to debug: check we have multiple subflows connection requests
-		ip netns exec $ns $t -A OUTPUT -p tcp --syn -m mark --mark $m -j ACCEPT
+	local table
+	for table in ip ip6; do
+		ip netns exec "$ns" nft -f - <<-EOF
+			add table $table filter
+			add chain $table filter OUTPUT \
+				{ type filter hook output priority 0; policy accept; }
 
-		# RST packets might be handled by a internal dummy socket
-		ip netns exec $ns $t -A OUTPUT -p tcp --tcp-flags RST RST -m mark --mark 0 -j ACCEPT
+			# just to debug: check we have multiple subflows connection requests
+			add rule $table filter OUTPUT \
+				tcp flags & (fin | syn | rst | ack) == syn \
+				meta mark $m accept
 
-		ip netns exec $ns $t -A OUTPUT -p tcp -m mark --mark $m -j ACCEPT
-		ip netns exec $ns $t -A OUTPUT -p tcp -m mark --mark 0 -j DROP
+			# RST packets might be handled by a internal dummy socket
+			add rule $table filter OUTPUT \
+				tcp flags & rst == rst meta mark 0x0 accept
+
+			add rule $table filter OUTPUT \
+				meta l4proto tcp meta mark $m accept
+			add rule $table filter OUTPUT \
+				meta l4proto tcp meta mark 0 counter drop
+		EOF
 	done
 }
 
@@ -105,32 +114,31 @@ cleanup()
 
 mptcp_lib_check_mptcp
 mptcp_lib_check_kallsyms
-mptcp_lib_check_tools ip "${iptables}" "${ip6tables}"
+mptcp_lib_check_tools ip nft jq
 
 check_mark()
 {
 	local ns=$1
 	local af=$2
 
-	local tables=${iptables}
+	local tables="ip"
 
 	if [ $af -eq 6 ];then
-		tables=${ip6tables}
+		tables="ip6"
 	fi
 
-	local counters values
-	counters=$(ip netns exec $ns $tables -v -L OUTPUT | grep DROP)
-	values=${counters%DROP*}
+	local drops
+	drops=$(ip netns exec "$ns" nft -j list table "$tables" filter | \
+		jq '.nftables[] | select(has("rule")) | .rule |
+			select (.chain=="OUTPUT" and any(.expr[]; has("drop"))) |
+			.expr[] | select(has("counter")) | .counter.packets')
 
-	local v
-	for v in $values; do
-		if [ $v -ne 0 ]; then
-			mptcp_lib_pr_fail "got $tables $values in ns $ns," \
-					  "not 0 - not all expected packets marked"
-			ret=${KSFT_FAIL}
-			return 1
-		fi
-	done
+	if [ -z "$drops" ] || [ "$drops" -ne 0 ]; then
+		mptcp_lib_pr_fail "got $tables $drops in ns $ns," \
+				  "not 0 - not all expected packets marked"
+		ret=${KSFT_FAIL}
+		return 1
+	fi
 
 	return 0
 }

-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH net-next 5/5] selftests: mptcp: convert iptables to nftables for mptcp_join.sh
  2026-09-26 15:30 [PATCH net-next 0/5] mptcp: misc improvements for v7.4 Matthieu Baerts (NGI0)
                   ` (3 preceding siblings ...)
  2026-09-26 15:30 ` [PATCH net-next 4/5] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh Matthieu Baerts (NGI0)
@ 2026-09-26 15:30 ` Matthieu Baerts (NGI0)
  2026-09-28  8:00   ` netdev-bot+sashiko
  4 siblings, 1 reply; 10+ messages in thread
From: Matthieu Baerts (NGI0) @ 2026-09-26 15:30 UTC (permalink / raw)
  To: Mat Martineau, Geliang Tang, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman
  Cc: netdev, mptcp, linux-kernel, Matthieu Baerts (NGI0),
	Hangbin Liu, Shuah Khan, linux-kselftest

From: Hangbin Liu <liuhangbin@kylinos.cn>

During conversion, we retain the same table and chain names used by the
original iptables/ip6tables setup, so rule output is identical to the
former iptables/ip6tables output. Add new function init_nftables so we
only do nft table setup when test need set nft rules.

Unlike iptables, nftables cannot match rules based on their full
specification. Since endpoint_tests is the only test that delete rules,
which inserts and deletes rules one by one. Flushing the table
directly is safe and easier than using nft_handle.

The cBPF bytecode matching MPTCP add‑addr and remove‑addr suboptions
is replaced with native nft matching using "tcp option mptcp subtype".

The config file adds CONFIG_NFT_NUMGEN (replaces iptables statistic nth),
CONFIG_NFT_REJECT and CONFIG_NFT_REJECT_INET for reject‑related rules.
Remove CONFIG_NFT_COMPAT since we don't need it now.

Remove the iptables/ip6tables check in mptcp_lib.sh since no script use
it now.

Signed-off-by: Hangbin Liu <liuhangbin@kylinos.cn>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
---
To: Shuah Khan <shuah@kernel.org>
Cc: linux-kselftest@vger.kernel.org
Cc: bpf@vger.kernel.org  ## Note: no eBPF code has been modified
---
 tools/testing/selftests/net/mptcp/config        |   4 +-
 tools/testing/selftests/net/mptcp/mptcp_join.sh | 149 +++++++++---------------
 tools/testing/selftests/net/mptcp/mptcp_lib.sh  |   2 +-
 3 files changed, 62 insertions(+), 93 deletions(-)

diff --git a/tools/testing/selftests/net/mptcp/config b/tools/testing/selftests/net/mptcp/config
index bb9c4c97c620..61057a066719 100644
--- a/tools/testing/selftests/net/mptcp/config
+++ b/tools/testing/selftests/net/mptcp/config
@@ -29,7 +29,9 @@ CONFIG_NET_SCH_INGRESS=m
 CONFIG_NET_SCH_NETEM=m
 CONFIG_NF_TABLES=m
 CONFIG_NF_TABLES_INET=y
-CONFIG_NFT_COMPAT=m
+CONFIG_NFT_NUMGEN=m
+CONFIG_NFT_REJECT=m
+CONFIG_NFT_REJECT_INET=m
 CONFIG_NFT_SOCKET=m
 CONFIG_NFT_TPROXY=m
 CONFIG_SYN_COOKIES=y
diff --git a/tools/testing/selftests/net/mptcp/mptcp_join.sh b/tools/testing/selftests/net/mptcp/mptcp_join.sh
index 18ce7136a2b0..b16e24418e73 100755
--- a/tools/testing/selftests/net/mptcp/mptcp_join.sh
+++ b/tools/testing/selftests/net/mptcp/mptcp_join.sh
@@ -26,8 +26,6 @@ capout=""
 cappid=""
 ns1=""
 ns2=""
-iptables="iptables"
-ip6tables="ip6tables"
 timeout_poll=30
 timeout_test=$((timeout_poll * 2 + 1))
 capture=false
@@ -99,42 +97,6 @@ unset add_addr_tx_nr
 unset add_addr_echo_tx_nr
 unset add_addr_drop_tx_nr
 
-# generated using "nfbpf_compile '(ip && (ip[54] & 0xf0) == 0x30) ||
-#				  (ip6 && (ip6[74] & 0xf0) == 0x30)'"
-CBPF_MPTCP_SUBOPTION_ADD_ADDR="14,
-			       48 0 0 0,
-			       84 0 0 240,
-			       21 0 3 64,
-			       48 0 0 54,
-			       84 0 0 240,
-			       21 6 7 48,
-			       48 0 0 0,
-			       84 0 0 240,
-			       21 0 4 96,
-			       48 0 0 74,
-			       84 0 0 240,
-			       21 0 1 48,
-			       6 0 0 65535,
-			       6 0 0 0"
-
-# IPv4: TCP hdr of 48B, a first suboption of 12B (DACK8), the RM_ADDR suboption
-# generated using "nfbpf_compile '(ip[32] & 0xf0) == 0xc0 && ip[53] == 0x0c &&
-#				  (ip[66] & 0xf0) == 0x40'"
-CBPF_MPTCP_SUBOPTION_RM_ADDR="13,
-			      48 0 0 0,
-			      84 0 0 240,
-			      21 0 9 64,
-			      48 0 0 32,
-			      84 0 0 240,
-			      21 0 6 192,
-			      48 0 0 53,
-			      21 0 4 12,
-			      48 0 0 66,
-			      84 0 0 240,
-			      21 0 1 64,
-			      6 0 0 65535,
-			      6 0 0 0"
-
 init_partial()
 {
 	capout=$(mktemp)
@@ -184,6 +146,26 @@ init_shapers()
 	done
 }
 
+init_nftables()
+{
+	local netns table
+	for netns in "$ns1" "$ns2"; do
+		for table in ip ip6; do
+			ip netns exec "$netns" nft -f - <<-EOF
+				add table $table filter
+				add chain $table filter INPUT \
+					{ type filter hook input priority filter; policy accept; }
+				add chain $table filter OUTPUT \
+					{ type filter hook output priority filter; policy accept; }
+
+				add table $table mangle
+				add chain $table mangle OUTPUT \
+					{ type route hook output priority mangle; policy accept; }
+			EOF
+		done
+	done
+}
+
 cleanup_partial()
 {
 	rm -f "$capout"
@@ -196,7 +178,7 @@ init() {
 
 	mptcp_lib_check_mptcp
 	mptcp_lib_check_kallsyms
-	mptcp_lib_check_tools ip tc ss "${iptables}" "${ip6tables}"
+	mptcp_lib_check_tools ip tc ss nft
 
 	sin=$(mktemp)
 	sout=$(mktemp)
@@ -380,24 +362,18 @@ reset_with_cookies()
 # $1: test name
 reset_with_add_addr_timeout()
 {
-	local ip="${2:-4}"
-	local tables
+	local ip="${2:-}"
 
 	reset "${1}" || return 1
-
-	tables="${iptables}"
-	if [ $ip -eq 6 ]; then
-		tables="${ip6tables}"
-	fi
+	init_nftables
 
 	# set a maximum, to avoid too long timeout with exponential backoff
 	ip netns exec $ns1 sysctl -q net.mptcp.add_addr_timeout=1
 
-	if ! ip netns exec $ns2 $tables -A OUTPUT -p tcp \
-			-m tcp --tcp-option 30 \
-			-m bpf --bytecode \
-			"$CBPF_MPTCP_SUBOPTION_ADD_ADDR" \
-			-j DROP; then
+	if ! ip netns exec "$ns2" nft add rule \
+			ip"$ip" filter OUTPUT meta l4proto tcp \
+			tcp option mptcp subtype add-addr \
+			drop; then
 		mark_as_skipped "unable to set the 'add addr' rule"
 		return 1
 	fi
@@ -449,22 +425,14 @@ setup_fail_rules()
 	check_invert=1
 	validate_checksum=true
 	local i="$1"
-	local ip="${2:-4}"
-	local tables
+	local ip="${2:-}"
 
-	tables="${iptables}"
-	if [ $ip -eq 6 ]; then
-		tables="${ip6tables}"
-	fi
-
-	ip netns exec $ns2 $tables \
-		-t mangle \
-		-A OUTPUT \
-		-o ns2eth$i \
-		-p tcp \
-		-m length --length 150:9999 \
-		-m statistic --mode nth --packet 1 --every 99999 \
-		-j MARK --set-mark 42 || return ${KSFT_SKIP}
+	init_nftables
+	ip netns exec "$ns2" nft add rule \
+		ip"$ip" mangle OUTPUT oifname ns2eth$i \
+		meta l4proto tcp \
+		meta length 150-9999 numgen inc mod 99999 1 \
+		meta mark set 42 || return ${KSFT_SKIP}
 
 	tc -n $ns2 qdisc add dev ns2eth$i clsact || return ${KSFT_SKIP}
 	tc -n $ns2 filter add dev ns2eth$i egress \
@@ -510,16 +478,15 @@ reset_with_tcp_filter()
 	reset "${1}" || return 1
 	shift
 
+	init_nftables
+
 	local ns="${!1}"
 	local src="${2}"
 	local target="${3}"
 	local chain="${4:-INPUT}"
 
-	if ! ip netns exec "${ns}" ${iptables} \
-			-A "${chain}" \
-			-s "${src}" \
-			-p tcp \
-			-j "${target}"; then
+	if ! ip netns exec "$ns" nft add rule ip filter "${chain}" \
+			ip saddr "${src}" meta l4proto tcp "${target,,}"; then
 		mark_as_skipped "unable to set the filter rules"
 		return 1
 	fi
@@ -4313,12 +4280,15 @@ userspace_tests()
 		chk_mptcp_info subflows 1 subflows 1
 		chk_subflows_total 2 2
 
+		init_nftables
 		# force quick loss
 		ip netns exec $ns2 sysctl -q net.ipv4.tcp_syn_retries=1
-		if ip netns exec "${ns1}" ${iptables} -A INPUT -s "10.0.1.2" \
-		      -p tcp --tcp-option 30 -j REJECT --reject-with tcp-reset &&
-		   ip netns exec "${ns2}" ${iptables} -A INPUT -d "10.0.1.2" \
-		      -p tcp --tcp-option 30 -j REJECT --reject-with tcp-reset; then
+		if ip netns exec "${ns1}" nft add rule ip filter INPUT \
+				ip saddr "10.0.1.2" meta l4proto tcp \
+				tcp option mptcp exists reject with tcp reset &&
+		   ip netns exec "${ns2}" nft add rule ip filter INPUT \
+				ip daddr "10.0.1.2" meta l4proto tcp \
+				tcp option mptcp exists reject with tcp reset; then
 			wait_event ns2 MPTCP_LIB_EVENT_SUB_CLOSED 1
 			wait_event ns1 MPTCP_LIB_EVENT_SUB_CLOSED 1
 			chk_subflows_total 1 1
@@ -4393,7 +4363,7 @@ endpoint_tests()
 		chk_subflow_nr "after new reject" 2
 		chk_mptcp_info subflows 1 subflows 1
 
-		ip netns exec "${ns2}" ${iptables} -D OUTPUT -s "10.0.3.2" -p tcp -j REJECT
+		ip netns exec "${ns2}" nft flush chain ip filter OUTPUT
 		pm_nl_del_endpoint $ns2 3 10.0.3.2
 		pm_nl_add_endpoint $ns2 10.0.3.2 id 3 flags subflow
 		wait_mpj 3
@@ -4402,12 +4372,10 @@ endpoint_tests()
 
 		# To make sure RM_ADDR are sent over a different subflow, but
 		# allow the rest to quickly and cleanly close the subflow
-		local ipt=1
-		ip netns exec "${ns2}" ${iptables} -I OUTPUT -s "10.0.1.2" \
-			-p tcp -m tcp --tcp-option 30 \
-			-m bpf --bytecode \
-			"$CBPF_MPTCP_SUBOPTION_RM_ADDR" \
-			-j DROP || ipt=0
+		local nft=1
+		ip netns exec "${ns2}" nft insert rule ip filter OUTPUT \
+			ip saddr 10.0.1.2 meta l4proto tcp \
+			tcp option mptcp subtype remove-addr drop || nft=0
 		local i
 		for i in $(seq 3); do
 			pm_nl_del_endpoint $ns2 1 10.0.1.2
@@ -4420,7 +4388,7 @@ endpoint_tests()
 			chk_subflow_nr "after re-add id 0 ($i)" 3
 			chk_mptcp_info subflows 3 subflows 3
 		done
-		[ ${ipt} = 1 ] && ip netns exec "${ns2}" ${iptables} -D OUTPUT 1
+		[ "${nft}" = 1 ] && ip netns exec "${ns2}" nft flush chain ip filter OUTPUT
 
 		mptcp_lib_kill_group_wait $tests_pid
 
@@ -4480,20 +4448,19 @@ endpoint_tests()
 		chk_mptcp_info subflows 2 subflows 2
 		chk_mptcp_info add_addr_signal 2 add_addr_accepted 2
 
+		init_nftables
 		# To make sure RM_ADDR are sent over a different subflow, but
 		# allow the rest to quickly and cleanly close the subflow
-		local ipt=1
-		ip netns exec "${ns1}" ${iptables} -I OUTPUT -s "10.0.1.1" \
-			-p tcp -m tcp --tcp-option 30 \
-			-m bpf --bytecode \
-			"$CBPF_MPTCP_SUBOPTION_RM_ADDR" \
-			-j DROP || ipt=0
+		local nft=1
+		ip netns exec "${ns1}" nft insert rule ip filter OUTPUT \
+			ip saddr 10.0.1.1 meta l4proto tcp \
+			tcp option mptcp subtype remove-addr drop || nft=0
 		pm_nl_del_endpoint $ns1 42 10.0.1.1
 		sleep 0.5
 		chk_subflow_nr "after delete ID 0" 2
 		chk_mptcp_info subflows 2 subflows 2
 		chk_mptcp_info add_addr_signal 2 add_addr_accepted 2
-		[ ${ipt} = 1 ] && ip netns exec "${ns1}" ${iptables} -D OUTPUT 1
+		[ "${nft}" = 1 ] && ip netns exec "${ns1}" nft flush chain ip filter OUTPUT
 
 		pm_nl_add_endpoint $ns1 10.0.1.1 id 42 flags signal
 		wait_mpj 4
@@ -4555,7 +4522,7 @@ endpoint_tests()
 		pm_nl_flush_endpoint $ns2
 		pm_nl_flush_endpoint $ns1
 		wait_rm_addr $ns2 0
-		ip netns exec "${ns2}" ${iptables} -D OUTPUT -s "10.0.3.2" -p tcp -j REJECT
+		ip netns exec "${ns2}" nft flush chain ip filter OUTPUT
 		pm_nl_add_endpoint $ns2 10.0.3.2 id 3 flags subflow
 		wait_mpj 1
 		pm_nl_add_endpoint $ns1 10.0.3.1 id 2 flags signal
diff --git a/tools/testing/selftests/net/mptcp/mptcp_lib.sh b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
index e65b4ebee06a..0559bb168203 100644
--- a/tools/testing/selftests/net/mptcp/mptcp_lib.sh
+++ b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
@@ -528,7 +528,7 @@ mptcp_lib_check_tools() {
 				exit ${KSFT_SKIP}
 			fi
 			;;
-		"iptables"* | "ip6tables"* | "nft" | "jq")
+		"nft" | "jq")
 			if ! "${tool}" -V &> /dev/null; then
 				mptcp_lib_pr_skip "Could not run all tests without ${tool}"
 				exit ${KSFT_SKIP}

-- 
2.55.0


^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH net-next 4/5] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh
  2026-09-26 15:30 ` [PATCH net-next 4/5] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh Matthieu Baerts (NGI0)
@ 2026-09-28  8:00   ` netdev-bot+sashiko
  2026-09-28 11:24     ` Matthieu Baerts
  0 siblings, 1 reply; 10+ messages in thread
From: netdev-bot+sashiko @ 2026-09-28  8:00 UTC (permalink / raw)
  To: matttbe
  Cc: martineau, geliang, davem, edumazet, kuba, pabeni, horms, netdev,
	mptcp, linux-kernel, quanyeyang

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The nft/jq tool check added to mptcp_lib_check_tools()
  (tools/testing/selftests/net/mptcp/mptcp_lib.sh:531-532) only runs…
- [Low] In add_mark_rules()
  (tools/testing/selftests/net/mptcp/mptcp_sockopt.sh:57-60), the…

--

selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh

This commit makes mptcp_sockopt.sh use nftables instead of
iptables/ip6tables. It creates ip/ip6 filter tables with an OUTPUT chain,
and check_mark() now reads the drop counter through nft -j and jq. It also
extends mptcp_lib_check_tools() to probe for nft and jq.

> diff --git a/tools/testing/selftests/net/mptcp/mptcp_lib.sh b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
> index b9d14647f4019..e65b4ebee06a6 100644
> --- a/tools/testing/selftests/net/mptcp/mptcp_lib.sh
> +++ b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
> @@ -528,7 +528,7 @@ mptcp_lib_check_tools() {
>  				exit ${KSFT_SKIP}
>  			fi
>  			;;
> -		"iptables"* | "ip6tables"*)
> +		"iptables"* | "ip6tables"* | "nft" | "jq")
>  			if ! "${tool}" -V &> /dev/null; then

[Severity: Medium]
Is nft -V the right probe for nft here?

For nft, -V prints extended build information, and only newer nftables
releases have it. The portable version flag is -v or --version, which is
what the other selftests use. For example, the netfilter/*.sh scripts call
nft --version.

If an older nft rejects -V, would mptcp_lib_check_tools() exit with
KSFT_SKIP from the top-level call in mptcp_sockopt.sh?

    mptcp_lib_check_tools ip nft jq

That would also skip the transfer, sockopt and TCP_INQ subtests, which
don't use nft at all.

The opposite case is also a concern. check_mark() needs an nft that was
built with JSON support:

    drops=$(ip netns exec "$ns" nft -j list table "$tables" filter | \
    ...
    if [ -z "$drops" ] || [ "$drops" -ne 0 ]; then
            mptcp_lib_pr_fail "got $tables $drops in ns $ns," \
                              "not 0 - not all expected packets marked"
            ret=${KSFT_FAIL}

On an nftables build without libjansson, nft -V still exits 0 and just
reports json: no. add_mark_rules() works, because it only uses nft -f.

nft -j then fails with "JSON support not compiled-in", so drops is empty.
The test then reports "not all expected packets marked" and sets
KSFT_FAIL. Should a missing userspace feature give a skip here instead?

The later patch in the series, "selftests: mptcp: convert iptables to
nftables for mptcp_join.sh", keeps the same "nft" | "jq" case with the -V
probe. So it looks like this is still present at the end of the series.

Would a separate nft check match the real dependency better? For example,
nft --version plus a JSON check such as nft -j list ruleset.

>  				mptcp_lib_pr_skip "Could not run all tests without ${tool}"
>  				exit ${KSFT_SKIP}
> diff --git a/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh b/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
> index e850a87429b61..b4e3eddc9cf21 100755
> --- a/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh
> +++ b/tools/testing/selftests/net/mptcp/mptcp_sockopt.sh

[ ... ]

> @@ -49,16 +47,27 @@ add_mark_rules()
>  	local ns=$1
>  	local m=$2
>  
> -	local t
> -	for t in ${iptables} ${ip6tables}; do
> -		# just to debug: check we have multiple subflows connection requests
> -		ip netns exec $ns $t -A OUTPUT -p tcp --syn -m mark --mark $m -j ACCEPT
> -
> -		# RST packets might be handled by a internal dummy socket
> -		ip netns exec $ns $t -A OUTPUT -p tcp --tcp-flags RST RST -m mark --mark 0 -j ACCEPT
> -
> -		ip netns exec $ns $t -A OUTPUT -p tcp -m mark --mark $m -j ACCEPT
> -		ip netns exec $ns $t -A OUTPUT -p tcp -m mark --mark 0 -j DROP
> +	local table
> +	for table in ip ip6; do
> +		ip netns exec "$ns" nft -f - <<-EOF
> +			add table $table filter
> +			add chain $table filter OUTPUT \
> +				{ type filter hook output priority 0; policy accept; }
> +
> +			# just to debug: check we have multiple subflows connection requests
> +			add rule $table filter OUTPUT \
> +				tcp flags & (fin | syn | rst | ack) == syn \
> +				meta mark $m accept

[Severity: Low]
This isn't a bug, but is this comment still accurate after the conversion?

With iptables every rule had implicit counters, so iptables -v -L could
show how many subflow SYNs were sent. In nftables a rule only counts
packets if it has an explicit counter statement, and this rule has none.

Every packet this rule matches is also accepted by the later "meta
l4proto tcp meta mark $m accept" rule, so the rule now does nothing.

The commit message says:

  "Counters are not added to accept rules because the test does not
  inspect them."

However, this rule was meant for manual debugging rather than for the test
itself. Should it get a counter, or should the rule and its comment be
removed?

> +
> +			# RST packets might be handled by a internal dummy socket
> +			add rule $table filter OUTPUT \
> +				tcp flags & rst == rst meta mark 0x0 accept

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260926-net-next-mptcp-misc-feat-7-4-v1-0-67af4ab37406%40kernel.org

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH net-next 5/5] selftests: mptcp: convert iptables to nftables for mptcp_join.sh
  2026-09-26 15:30 ` [PATCH net-next 5/5] selftests: mptcp: convert iptables to nftables for mptcp_join.sh Matthieu Baerts (NGI0)
@ 2026-09-28  8:00   ` netdev-bot+sashiko
  2026-09-28 11:24     ` Matthieu Baerts
  0 siblings, 1 reply; 10+ messages in thread
From: netdev-bot+sashiko @ 2026-09-28  8:00 UTC (permalink / raw)
  To: matttbe
  Cc: martineau, geliang, davem, edumazet, kuba, pabeni, horms, netdev,
	mptcp, linux-kernel, quanyeyang

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] In endpoint_tests()
  (tools/testing/selftests/net/mptcp/mptcp_join.sh), the new rule `tcp…
- [Low] The patch removes CONFIG_NFT_COMPAT from
  tools/testing/selftests/net/mptcp/config, "since we don't need it now".

--

selftests: mptcp: convert iptables to nftables for mptcp_join.sh

This converts the iptables/ip6tables rules in mptcp_join.sh to nftables.
It keeps the same table and chain names and adds init_nftables() to set
up the tables only when a test needs them. The cBPF bytecode that matched
the ADD_ADDR and RM_ADDR suboptions is replaced with "tcp option mptcp
subtype" matching. In the config fragment, CONFIG_NFT_COMPAT is swapped
for the nft modules that are now needed.

> diff --git a/tools/testing/selftests/net/mptcp/config b/tools/testing/selftests/net/mptcp/config
> index bb9c4c97c6206..61057a0667192 100644
> --- a/tools/testing/selftests/net/mptcp/config
> +++ b/tools/testing/selftests/net/mptcp/config
> @@ -29,7 +29,9 @@ CONFIG_NET_SCH_INGRESS=m
>  CONFIG_NET_SCH_NETEM=m
>  CONFIG_NF_TABLES=m
>  CONFIG_NF_TABLES_INET=y
> -CONFIG_NFT_COMPAT=m
> +CONFIG_NFT_NUMGEN=m
> +CONFIG_NFT_REJECT=m
> +CONFIG_NFT_REJECT_INET=m

[Severity: Low]
This isn't a bug, but CONFIG_NFT_COMPAT is removed here "since we don't
need it now". Should the other xtables-only options in this fragment be
removed too?

After this patch the config still has:

CONFIG_IP6_NF_FILTER=m
CONFIG_IP6_NF_TARGET_REJECT=m
CONFIG_IP_NF_FILTER=m
CONFIG_IP_NF_MANGLE=m
CONFIG_IP_NF_TARGET_REJECT=m
CONFIG_NETFILTER_XTABLES=m
CONFIG_NETFILTER_XTABLES_LEGACY=y
CONFIG_NETFILTER_XT_MATCH_BPF=m
CONFIG_NETFILTER_XT_MATCH_LENGTH=m
CONFIG_NETFILTER_XT_MATCH_STATISTIC=m
CONFIG_NETFILTER_XT_TARGET_MARK=m

Only the removed -m bpf, -m length, -m statistic, -j MARK and REJECT
rules needed these. Nothing under tools/testing/selftests/net/mptcp/
uses iptables or ip6tables anymore.

>  CONFIG_NFT_SOCKET=m
>  CONFIG_NFT_TPROXY=m
>  CONFIG_SYN_COOKIES=y
> diff --git a/tools/testing/selftests/net/mptcp/mptcp_join.sh b/tools/testing/selftests/net/mptcp/mptcp_join.sh
> index 18ce7136a2b0e..b16e24418e737 100755
> --- a/tools/testing/selftests/net/mptcp/mptcp_join.sh
> +++ b/tools/testing/selftests/net/mptcp/mptcp_join.sh

[ ... ]

> @@ -4402,12 +4372,10 @@ endpoint_tests()
>  
>  		# To make sure RM_ADDR are sent over a different subflow, but
>  		# allow the rest to quickly and cleanly close the subflow
> -		local ipt=1
> -		ip netns exec "${ns2}" ${iptables} -I OUTPUT -s "10.0.1.2" \
> -			-p tcp -m tcp --tcp-option 30 \
> -			-m bpf --bytecode \
> -			"$CBPF_MPTCP_SUBOPTION_RM_ADDR" \
> -			-j DROP || ipt=0
> +		local nft=1
> +		ip netns exec "${ns2}" nft insert rule ip filter OUTPUT \
> +			ip saddr 10.0.1.2 meta l4proto tcp \
> +			tcp option mptcp subtype remove-addr drop || nft=0

[Severity: Medium]
Can this rule ever match an RM_ADDR packet on an established connection?

nft_exthdr_tcp_eval() stops at the first TCP option whose kind matches
(30 for MPTCP). The subtype check therefore only ever sees the first
MPTCP suboption:

net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
    ...
	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
		optl = optlen(opt, i);

		if (priv->type != opt[i])
			continue;
		...
		return;
	}
    ...
}

Once can_ack is set, mptcp_established_options_dss() always selects a
DSS ack:

net/mptcp/options.c:mptcp_established_options_dss() {
    ...
	opts->ext_copy.use_ack = 1;
	opts->suboptions = OPTION_MPTCP_DSS;
    ...
}

mptcp_established_options_rm_addr() then ORs OPTION_MPTCP_RM_ADDR into
suboptions, so the DSS is kept. mptcp_write_options() writes the DSS
first and writes RM_ADDR later, after the mp_capable_done label:

net/mptcp/options.c:mptcp_write_options() {
    ...
	if (likely(OPTION_MPTCP_DSS & opts->suboptions)) {
    ...
mp_capable_done:
	if (OPTION_MPTCP_RM_ADDR & opts->suboptions) {
    ...
}

On these packets the first kind 30 option is the DSS (subtype 0x2), so
wouldn't "subtype remove-addr" (0x4) always fail to match?

The removed CBPF_MPTCP_SUBOPTION_RM_ADDR bytecode was written for this
layout. Its comment said "a first suboption of 12B (DACK8), the RM_ADDR
suboption". It checked ip[53] == 0x0c for the length of the first
suboption and (ip[66] & 0xf0) == 0x40 for the subtype of the second one.

The same problem affects the ns1 rule on 10.0.1.1 later in
endpoint_tests(), which is installed right before
"pm_nl_del_endpoint $ns1 42 10.0.1.1".

In both places the rule installs without error, so the "|| nft=0"
fallback never triggers. The tests keep passing, but RM_ADDR is no
longer forced onto a different subflow.

The commit message says the bytecode "is replaced with native nft
matching using "tcp option mptcp subtype"". That is true for ADD_ADDR,
because mptcp_write_options() emits ADD_ADDR in an else-if branch with
the DSS. Is it also true for RM_ADDR?

Would a raw payload match on the second MPTCP option (for example with
@th,...) be needed here instead?

>  		local i
>  		for i in $(seq 3); do
>  			pm_nl_del_endpoint $ns2 1 10.0.1.2

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260926-net-next-mptcp-misc-feat-7-4-v1-0-67af4ab37406%40kernel.org

^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH net-next 4/5] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh
  2026-09-28  8:00   ` netdev-bot+sashiko
@ 2026-09-28 11:24     ` Matthieu Baerts
  0 siblings, 0 replies; 10+ messages in thread
From: Matthieu Baerts @ 2026-09-28 11:24 UTC (permalink / raw)
  To: netdev-bot+sashiko
  Cc: martineau, geliang, davem, edumazet, kuba, pabeni, horms, netdev,
	mptcp, linux-kernel, quanyeyang

Hello,

On 28/09/2026 10:00, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential
> issue(s) to consider.
> 
> Critical: 0 · High: 0 · Medium: 1 · Low: 1
> 
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.

TL;DR: False positive: no new version needed here.

> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
> 
> - [Medium] The nft/jq tool check added to mptcp_lib_check_tools()
>   (tools/testing/selftests/net/mptcp/mptcp_lib.sh:531-532) only runs…
> - [Low] In add_mark_rules()
>   (tools/testing/selftests/net/mptcp/mptcp_sockopt.sh:57-60), the…
> 
> --
> 
> selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh
> 
> This commit makes mptcp_sockopt.sh use nftables instead of
> iptables/ip6tables. It creates ip/ip6 filter tables with an OUTPUT chain,
> and check_mark() now reads the drop counter through nft -j and jq. It also
> extends mptcp_lib_check_tools() to probe for nft and jq.
> 
>> diff --git a/tools/testing/selftests/net/mptcp/mptcp_lib.sh b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
>> index b9d14647f4019..e65b4ebee06a6 100644
>> --- a/tools/testing/selftests/net/mptcp/mptcp_lib.sh
>> +++ b/tools/testing/selftests/net/mptcp/mptcp_lib.sh
>> @@ -528,7 +528,7 @@ mptcp_lib_check_tools() {
>>  				exit ${KSFT_SKIP}
>>  			fi
>>  			;;
>> -		"iptables"* | "ip6tables"*)
>> +		"iptables"* | "ip6tables"* | "nft" | "jq")
>>  			if ! "${tool}" -V &> /dev/null; then
> 
> [Severity: Medium]
> Is nft -V the right probe for nft here?
> 
> For nft, -V prints extended build information, and only newer nftables
> releases have it. The portable version flag is -v or --version, which is
> what the other selftests use. For example, the netfilter/*.sh scripts call
> nft --version.

It requires nft >= 0.9.4, released in 2020. I think we are good.

> If an older nft rejects -V, would mptcp_lib_check_tools() exit with
> KSFT_SKIP from the top-level call in mptcp_sockopt.sh?
> 
>     mptcp_lib_check_tools ip nft jq
> 
> That would also skip the transfer, sockopt and TCP_INQ subtests, which
> don't use nft at all.
> 
> The opposite case is also a concern. check_mark() needs an nft that was
> built with JSON support:
> 
>     drops=$(ip netns exec "$ns" nft -j list table "$tables" filter | \
>     ...
>     if [ -z "$drops" ] || [ "$drops" -ne 0 ]; then
>             mptcp_lib_pr_fail "got $tables $drops in ns $ns," \
>                               "not 0 - not all expected packets marked"
>             ret=${KSFT_FAIL}
> 
> On an nftables build without libjansson, nft -V still exits 0 and just
> reports json: no. add_mark_rules() works, because it only uses nft -f.
> 
> nft -j then fails with "JSON support not compiled-in", so drops is empty.
> The test then reports "not all expected packets marked" and sets
> KSFT_FAIL. Should a missing userspace feature give a skip here instead?
> 
> The later patch in the series, "selftests: mptcp: convert iptables to
> nftables for mptcp_join.sh", keeps the same "nft" | "jq" case with the -V
> probe. So it looks like this is still present at the end of the series.
> 
> Would a separate nft check match the real dependency better? For example,
> nft --version plus a JSON check such as nft -j list ruleset.

I don't expect packet distributions not building nft with JSON support.

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 10+ messages in thread

* Re: [PATCH net-next 5/5] selftests: mptcp: convert iptables to nftables for mptcp_join.sh
  2026-09-28  8:00   ` netdev-bot+sashiko
@ 2026-09-28 11:24     ` Matthieu Baerts
  0 siblings, 0 replies; 10+ messages in thread
From: Matthieu Baerts @ 2026-09-28 11:24 UTC (permalink / raw)
  To: netdev-bot+sashiko
  Cc: martineau, geliang, davem, edumazet, kuba, pabeni, horms, netdev,
	mptcp, linux-kernel, quanyeyang

Hello,

On 28/09/2026 10:00, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential
> issue(s) to consider.
> 
> Critical: 0 · High: 0 · Medium: 1 · Low: 1
> 
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.

TL;DR: It looks like there are issues with NFTables, but no new version
needed here for these MPTCP patches.

> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
> 
> - [Medium] In endpoint_tests()
>   (tools/testing/selftests/net/mptcp/mptcp_join.sh), the new rule `tcp…
> - [Low] The patch removes CONFIG_NFT_COMPAT from
>   tools/testing/selftests/net/mptcp/config, "since we don't need it now".
> 
> --
> 
> selftests: mptcp: convert iptables to nftables for mptcp_join.sh
> 
> This converts the iptables/ip6tables rules in mptcp_join.sh to nftables.
> It keeps the same table and chain names and adds init_nftables() to set
> up the tables only when a test needs them. The cBPF bytecode that matched
> the ADD_ADDR and RM_ADDR suboptions is replaced with "tcp option mptcp
> subtype" matching. In the config fragment, CONFIG_NFT_COMPAT is swapped
> for the nft modules that are now needed.

(...)

>> diff --git a/tools/testing/selftests/net/mptcp/mptcp_join.sh b/tools/testing/selftests/net/mptcp/mptcp_join.sh
>> index 18ce7136a2b0e..b16e24418e737 100755
>> --- a/tools/testing/selftests/net/mptcp/mptcp_join.sh
>> +++ b/tools/testing/selftests/net/mptcp/mptcp_join.sh
> 
> [ ... ]
> 
>> @@ -4402,12 +4372,10 @@ endpoint_tests()
>>  
>>  		# To make sure RM_ADDR are sent over a different subflow, but
>>  		# allow the rest to quickly and cleanly close the subflow
>> -		local ipt=1
>> -		ip netns exec "${ns2}" ${iptables} -I OUTPUT -s "10.0.1.2" \
>> -			-p tcp -m tcp --tcp-option 30 \
>> -			-m bpf --bytecode \
>> -			"$CBPF_MPTCP_SUBOPTION_RM_ADDR" \
>> -			-j DROP || ipt=0
>> +		local nft=1
>> +		ip netns exec "${ns2}" nft insert rule ip filter OUTPUT \
>> +			ip saddr 10.0.1.2 meta l4proto tcp \
>> +			tcp option mptcp subtype remove-addr drop || nft=0
> 
> [Severity: Medium]
> Can this rule ever match an RM_ADDR packet on an established connection?
> 
> nft_exthdr_tcp_eval() stops at the first TCP option whose kind matches
> (30 for MPTCP). The subtype check therefore only ever sees the first
> MPTCP suboption:
> 
> net/netfilter/nft_exthdr.c:nft_exthdr_tcp_eval() {
>     ...
> 	for (i = sizeof(*tcph); i < tcphdr_len - 1; i += optl) {
> 		optl = optlen(opt, i);
> 
> 		if (priv->type != opt[i])
> 			continue;
> 		...
> 		return;
> 	}
>     ...
> }

Indeed, the RM_ADDR will be added in a second MPTCP option. It looks
like Netfilter doesn't handle that case. But that seems to be an issue
on the Netfilter side, rather than with the command that should work. I
will follow up with Netfilter devs. If a fix cannot be added on their
side, I will change the nft command to look at a specific offset.

Note that the command here is just to check there is no RM_ADDR sent on
the wrong side: it shouldn't catch any packets here anyway.

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.


^ permalink raw reply	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-09-28 11:24 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-26 15:30 [PATCH net-next 0/5] mptcp: misc improvements for v7.4 Matthieu Baerts (NGI0)
2026-09-26 15:30 ` [PATCH net-next 1/5] mptcp: remove thmac from subflow ctx Matthieu Baerts (NGI0)
2026-09-26 15:30 ` [PATCH net-next 2/5] mptcp: split FASTCLOSE key from rcvr_key Matthieu Baerts (NGI0)
2026-09-26 15:30 ` [PATCH net-next 3/5] mptcp: shrink struct mptcp_options_received Matthieu Baerts (NGI0)
2026-09-26 15:30 ` [PATCH net-next 4/5] selftests: mptcp: convert iptables to nftables for mptcp_sockopt.sh Matthieu Baerts (NGI0)
2026-09-28  8:00   ` netdev-bot+sashiko
2026-09-28 11:24     ` Matthieu Baerts
2026-09-26 15:30 ` [PATCH net-next 5/5] selftests: mptcp: convert iptables to nftables for mptcp_join.sh Matthieu Baerts (NGI0)
2026-09-28  8:00   ` netdev-bot+sashiko
2026-09-28 11:24     ` Matthieu Baerts

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®