mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH net-next] netfilter: bridge: refragment reassembled packets passed up
@ 2026-10-09 15:02 Gabriel Goller
  2026-10-09 15:04 ` netdev-bot+sinfo
  0 siblings, 1 reply; 2+ messages in thread
From: Gabriel Goller @ 2026-10-09 15:02 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter,
	Nikolay Aleksandrov, Ido Schimmel, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan
  Cc: linux-kernel, netfilter-devel, coreteam, bridge, netdev, linux-kselftest

A bridge can have an upper device that is a port of another bridge,
e.g. a VLAN device on br0 that is enslaved to br1:

  eth0 -- br0 -- br0.241 (VLAN device) -- br1 -- eth1

br0 reassembles fragments in prerouting so conntrack sees the whole
packet. If br0 forwards the packet, postrouting splits it into the
original fragments again. If br0 passes it up instead, it stays whole.
br1 then cannot refragment it, because frag_max_size is kept in
skb->cb, which br1 clears on input. Packets larger than the MTU of
eth1 are dropped.

Refragment in local input as well, unless the packet is addressed to
the bridge itself, i.e. pkt_type is PACKET_HOST. That is the case if
the destination is the MAC address of the bridge or of the port the
frame arrived on. Other packets reach local input if they are
broadcast or multicast, sent to the MAC address of another port, or
sent to any address while the bridge is promiscuous. The bridge is
promiscuous if an upper device is a port of another bridge, has its
own MAC address (e.g. a macvlan device), or if a packet capture is
running.

Passing a packet up leaves the bridge just like forwarding, so handle
it the same way: reassemble only to track, then hand on the received
fragments. The layer above sees the same packets as without
conntrack, whatever kind of device it is, and nothing has to be
looked up per packet. Packets for the bridge itself are not bridged
again. They stay whole, so the inet stack can use their conntrack
entry without reassembling again. Fragments are passed up without
conntrack, as the layer above would otherwise skip reassembly and
tracking.

Changes for existing setups with bridge conntrack:

  - Upper devices with their own MAC address and local receivers of
    fragmented broadcast or multicast get fragments and reassemble
    them again. (not good, but also not very common)
  - Packet captures on the bridge see fragments, as without conntrack.
    (not sure if this is a big issue)
  - Packets that cannot be refragmented still stay whole: those whose
    fragments exceed the bridge MTU, IPv6 packets with fragments below
    IPV6_MIN_MTU, and packets that are no longer plain IP because VLAN
    ingress moved a tag of a foreign protocol into the payload.

A selftest covers bridge, VLAN, bridge and bridge, macvlan, bridge
topologies. Fragmented IPv4 and IPv6 pings across them fail in both
directions without the fix.

Fixes: 3c171f496ef5 ("netfilter: bridge: add connection tracking system")
Signed-off-by: Gabriel Goller <g.goller@proxmox.com>
---

RFC: I'm not confident in this patch. Refragmenting on local input
affects every setup with something above a conntrack-enabled bridge,
not just stacked bridges, and may break things I haven't foreseen.

An alternative would be to keep frag_max_size with the packet in a
new skb extension, as br_netfilter does, and let br1 restore it. That
would avoid reassembling twice, but costs an extension id and an
allocation per packet passed up, and the extension has to be removed
on every path that does not use it.

 net/bridge/netfilter/nf_conntrack_bridge.c    |  66 ++++++++
 .../testing/selftests/net/netfilter/Makefile  |   1 +
 .../net/netfilter/bridge_conntrack_frag.sh    | 157 ++++++++++++++++++
 tools/testing/selftests/net/netfilter/config  |   1 +
 4 files changed, 225 insertions(+)
 create mode 100755 tools/testing/selftests/net/netfilter/bridge_conntrack_frag.sh

diff --git a/net/bridge/netfilter/nf_conntrack_bridge.c b/net/bridge/netfilter/nf_conntrack_bridge.c
index b5444335b86f..d738ef5f6dc8 100644
--- a/net/bridge/netfilter/nf_conntrack_bridge.c
+++ b/net/bridge/netfilter/nf_conntrack_bridge.c
@@ -295,6 +295,41 @@ static unsigned int nf_ct_bridge_pre(void *priv, struct sk_buff *skb,
 	return nf_conntrack_in(skb, &bridge_state);
 }
 
+static unsigned int
+nf_ct_bridge_refrag(struct sk_buff *skb, const struct nf_hook_state *state,
+		    int (*output)(struct net *, struct sock *sk,
+				  const struct nf_bridge_frag_data *data,
+				  struct sk_buff *));
+static int nf_ct_bridge_refrag_in(struct net *net, struct sock *sk,
+				  const struct nf_bridge_frag_data *data,
+				  struct sk_buff *skb);
+
+/* Prerouting reassembled the packet for tracking. Unless it is for the bridge
+ * itself, an upper device may bridge it again without frag_max_size and drop
+ * it, so pass up the original fragments. Keep the packet whole if
+ * refragmenting would drop it.
+ */
+static bool nf_ct_bridge_refrag_up(const struct sk_buff *skb)
+{
+	u16 frag_max_size = BR_INPUT_SKB_CB(skb)->frag_max_size;
+
+	if (skb->pkt_type == PACKET_HOST || !frag_max_size ||
+	    frag_max_size > skb->dev->mtu)
+		return false;
+
+	/* VLAN ingress may have moved a tag of a foreign protocol into the
+	 * payload since prerouting, so the packet is no longer plain IP.
+	 */
+	switch (skb->protocol) {
+	case htons(ETH_P_IP):
+		return true;
+	case htons(ETH_P_IPV6):
+		return frag_max_size >= IPV6_MIN_MTU;
+	default:
+		return false;
+	}
+}
+
 static unsigned int nf_ct_bridge_in(void *priv, struct sk_buff *skb,
 				    const struct nf_hook_state *state)
 {
@@ -302,6 +337,14 @@ static unsigned int nf_ct_bridge_in(void *priv, struct sk_buff *skb,
 	struct nf_conntrack *nfct = skb_nfct(skb);
 	struct nf_conn *ct;
 
+	if (nf_ct_bridge_refrag_up(skb)) {
+		/* Fragments carrying conntrack would make the layer above skip
+		 * reassembly and tracking.
+		 */
+		nf_reset_ct(skb);
+		return nf_ct_bridge_refrag(skb, state, nf_ct_bridge_refrag_in);
+	}
+
 	if (promisc) {
 		nf_reset_ct(skb);
 		return NF_ACCEPT;
@@ -387,6 +430,29 @@ static int nf_ct_bridge_frag_restore(struct sk_buff *skb,
 	return 0;
 }
 
+static int nf_ct_bridge_refrag_in(struct net *net, struct sock *sk,
+				  const struct nf_bridge_frag_data *data,
+				  struct sk_buff *skb)
+{
+	int err;
+
+	err = nf_ct_bridge_frag_restore(skb, data);
+	if (err < 0)
+		return err;
+
+	/* The fast path reuses the reassembled head as the first fragment. Its
+	 * ignore_df from defrag would let IPv6 forwarding refragment it.
+	 */
+	skb->ignore_df = 0;
+
+	/* Unlike the transmit path, receive keeps the Ethernet header pulled. */
+	skb_set_mac_header(skb, -ETH_HLEN);
+	br_drop_fake_rtable(skb);
+	netif_receive_skb(skb);
+
+	return 0;
+}
+
 static int nf_ct_bridge_refrag_post(struct net *net, struct sock *sk,
 				    const struct nf_bridge_frag_data *data,
 				    struct sk_buff *skb)
diff --git a/tools/testing/selftests/net/netfilter/Makefile b/tools/testing/selftests/net/netfilter/Makefile
index df3c20c90f5d..d3bc881deaf7 100644
--- a/tools/testing/selftests/net/netfilter/Makefile
+++ b/tools/testing/selftests/net/netfilter/Makefile
@@ -12,6 +12,7 @@ TEST_PROGS := \
 	br_netfilter.sh \
 	br_netfilter_queue.sh \
 	bridge_brouter.sh \
+	bridge_conntrack_frag.sh \
 	conntrack_clash.sh \
 	conntrack_dump_flush.sh \
 	conntrack_icmp_related.sh \
diff --git a/tools/testing/selftests/net/netfilter/bridge_conntrack_frag.sh b/tools/testing/selftests/net/netfilter/bridge_conntrack_frag.sh
new file mode 100755
index 000000000000..04402f0f9fdf
--- /dev/null
+++ b/tools/testing/selftests/net/netfilter/bridge_conntrack_frag.sh
@@ -0,0 +1,157 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (C) 2026 Proxmox Server Solutions GmbH, Gabriel Goller <g.goller@proxmox.com>
+#
+# Bridge conntrack reassembles fragments. Packets passed up a bridge must be
+# refragmented, or a stacked bridge drops them if they exceed its egress MTU:
+#
+# vlan:    sender                   bridge                       receiver
+#          eth0.241 -- eth0 <-> veth0 -- br0 -- br0.241 -- br1 -- veth1 <-> eth0
+#
+# macvlan: sender             bridge                             receiver
+#          eth0 <-> veth0 -- br0 -- mv0 (passthru) -- br1 -- veth1 <-> eth0
+#
+# br0 is promiscuous because its upper device is a bridge port.
+
+source lib.sh
+
+checktool "nft --version" "run test without nft"
+
+ret=0
+
+trap cleanup_all_ns EXIT
+
+add_dev()
+{
+	if ! ip "$@"; then
+		# Do not hide failures of an earlier topology.
+		[ "$ret" -eq 0 ] || exit "$ret"
+		echo "SKIP: Cannot create required network device"
+		exit $ksft_skip
+	fi
+}
+
+setup_topology()
+{
+	local upper=$1 knob
+
+	setup_ns sender nsbr receiver || exit $ksft_skip
+
+	add_dev link add veth0 netns "$nsbr" type veth peer name eth0 netns "$sender"
+	add_dev link add veth1 netns "$nsbr" type veth peer name eth0 netns "$receiver"
+	add_dev -net "$nsbr" link add br0 type bridge
+	add_dev -net "$nsbr" link add br1 type bridge
+
+	case "$upper" in
+	vlan)
+		add_dev -net "$nsbr" link add link br0 name br0.241 type vlan id 241
+		add_dev -net "$sender" link add link eth0 name eth0.241 type vlan id 241
+		upper_dev=br0.241
+		sender_dev=eth0.241
+		;;
+	macvlan)
+		add_dev -net "$nsbr" link add link br0 name mv0 type macvlan mode passthru
+		upper_dev=mv0
+		sender_dev=eth0
+		;;
+	esac
+
+	# br1 could otherwise take over br0's address from its upper device port,
+	# and traffic to br1 would no longer reach br0 via promiscuous delivery.
+	ip -net "$nsbr" link set br1 address 02:00:00:00:00:20
+
+	# br_netfilter would call the inet hooks on bridged traffic.
+	for knob in bridge-nf-call-iptables bridge-nf-call-ip6tables; do
+		ip netns exec "$nsbr" sysctl -qw "net.bridge.$knob=0" 2>/dev/null
+	done
+
+	ip -net "$nsbr" link set veth0 master br0
+	ip -net "$nsbr" link set "$upper_dev" master br1
+	ip -net "$nsbr" link set veth1 master br1
+
+	for dev in veth0 veth1 br0 "$upper_dev" br1; do
+		ip -net "$nsbr" link set "$dev" up
+	done
+	ip -net "$sender" link set eth0 up
+	ip -net "$sender" link set "$sender_dev" up
+	ip -net "$receiver" link set eth0 up
+
+	ip -net "$sender" addr add 192.0.2.1/24 dev "$sender_dev"
+	ip -net "$receiver" addr add 192.0.2.2/24 dev eth0
+	ip -net "$nsbr" addr add 192.0.2.3/24 dev br1
+	ip -net "$sender" addr add 2001:db8:1::1/64 dev "$sender_dev" nodad
+	ip -net "$receiver" addr add 2001:db8:1::2/64 dev eth0 nodad
+	ip -net "$nsbr" addr add 2001:db8:1::3/64 dev br1 nodad
+}
+
+check_ping()
+{
+	local ns=$1 dst=$2 desc=$3
+
+	if ip netns exec "$ns" ping -q -c 2 -i 0.1 -w 3 -M dont -s 4000 "$dst" \
+		>/dev/null 2>&1; then
+		echo "PASS: $phase $desc"
+	else
+		echo "FAIL: $phase $desc"
+		ret=1
+	fi
+}
+
+check_paths()
+{
+	local prefix
+
+	for prefix in 192.0.2. 2001:db8:1::; do
+		check_ping "$sender" "${prefix}2" "sender -> receiver ${prefix}2"
+		check_ping "$receiver" "${prefix}1" "receiver -> sender ${prefix}1"
+		check_ping "$sender" "${prefix}3" "sender -> bridge host ${prefix}3"
+	done
+}
+
+run_test()
+{
+	local upper=$1 family
+
+	setup_topology "$upper"
+
+	phase="$upper without conntrack:"
+	check_paths
+	# A broken setup must not be hidden by a later skip.
+	[ "$ret" -eq 0 ] || exit "$ret"
+
+	# The ct expression enables the bridge conntrack hooks. Count oversized
+	# packets at br0 input to make sure they were reassembled there.
+	if ! ip netns exec "$nsbr" nft -f - <<'NFT'
+table bridge br_frag {
+	counter tracked4 { }
+	counter tracked6 { }
+	chain input {
+		type filter hook input priority 0; policy accept;
+		iifname "veth0" ct state new,established meta length > 1500 \
+			ip saddr 192.0.2.1 counter name tracked4
+		iifname "veth0" ct state new,established meta length > 1500 \
+			ip6 saddr 2001:db8:1::1 counter name tracked6
+	}
+}
+NFT
+	then
+		echo "SKIP: Bridge conntrack is not available"
+		exit $ksft_skip
+	fi
+
+	phase="$upper with conntrack:"
+	check_paths
+
+	for family in 4 6; do
+		if ! ip netns exec "$nsbr" nft list counter bridge br_frag "tracked$family" |
+			grep -q 'packets [1-9]'; then
+			echo "FAIL: $phase no reassembled IPv$family packets at br0 input"
+			ret=1
+		fi
+	done
+}
+
+run_test vlan
+run_test macvlan
+
+exit "$ret"
diff --git a/tools/testing/selftests/net/netfilter/config b/tools/testing/selftests/net/netfilter/config
index 629b2b29ed1a..91985320c292 100644
--- a/tools/testing/selftests/net/netfilter/config
+++ b/tools/testing/selftests/net/netfilter/config
@@ -62,6 +62,7 @@ CONFIG_NET_SCH_HTB=m
 CONFIG_NET_SCH_NETEM=m
 CONFIG_NET_VRF=y
 CONFIG_NF_CONNTRACK=m
+CONFIG_NF_CONNTRACK_BRIDGE=m
 CONFIG_NF_CONNTRACK_EVENTS=y
 CONFIG_NF_CONNTRACK_FTP=m
 CONFIG_NF_CONNTRACK_MARK=y
-- 
2.47.3



^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [RFC PATCH net-next] netfilter: bridge: refragment reassembled packets passed up
  2026-10-09 15:02 [RFC PATCH net-next] netfilter: bridge: refragment reassembled packets passed up Gabriel Goller
@ 2026-10-09 15:04 ` netdev-bot+sinfo
  0 siblings, 0 replies; 2+ messages in thread
From: netdev-bot+sinfo @ 2026-10-09 15:04 UTC (permalink / raw)
  To: Gabriel Goller
  Cc: Pablo Neira Ayuso, Florian Westphal, Phil Sutter,
	Nikolay Aleksandrov, Ido Schimmel, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, Shuah Khan,
	linux-kernel, netfilter-devel, coreteam, bridge, netdev,
	linux-kselftest

Hi!

This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:

 - How the issue was discovered, e.g. hit in production, hit during
   development, syzbot report, manual code inspection, LLM or static
   analysis tool scan.

Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.

The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-09 15:04 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 15:02 [RFC PATCH net-next] netfilter: bridge: refragment reassembled packets passed up Gabriel Goller
2026-10-09 15:04 ` netdev-bot+sinfo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®