mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Omar Ramadan <omar@blockcast.net>
To: Taehee Yoo <ap420073@gmail.com>,
	Andrew Lunn <andrew+netdev@lunn.ch>,
	"David S . Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@kernel.org>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Shuah Khan <shuah@kernel.org>
Cc: Simon Horman <horms@kernel.org>,
	netdev@vger.kernel.org, linux-kselftest@vger.kernel.org,
	linux-kernel@vger.kernel.org
Subject: [PATCH net-next 06/13] amt: forward multicast data over IPv6
Date: Fri,  9 Oct 2026 12:24:19 +0000	[thread overview]
Message-ID: <20261009122426.551178-7-omar@blockcast.net> (raw)
In-Reply-To: <20261009122426.551178-1-omar@blockcast.net>

amt_send_multicast_data() routes and transmits the relay's copy of a
multicast packet only over IPv4, so an IPv6 relay could establish
tunnels but not deliver traffic through them.

Size the copy's headroom for the outer family and send it through
amt_udp_xmit(), whose new @data argument marks Multicast Data, still
routed without a DSCP. A copy that cannot be sent is now counted in
tx_dropped and freed with a specific drop reason; before, it was freed
with SKB_DROP_REASON_NOT_SPECIFIED and not counted, as amt_dev_xmit()
returns NETDEV_TX_OK for the packet.

RFC 7450 forbids a Fragment header on Multicast Data sent over IPv6
(s5.3.3.6.3.2) and wants an IPv6 payload above the tunnel MTU dropped,
with a Packet Too Big to its source (s5.3.3.6.2.2). ip6_fragment() would
drop it, but send the Packet Too Big to the relay itself. Instead,
amt_tmtu_exceeded() compares the payload with the route's MTU less the
outer headers, per tunnel since each gateway has its own route, and
sends the Packet Too Big through icmpv6_ndo_send(), as ip6_tunnel does,
with at least IPV6_MIN_MTU, as a source ignores less. Packet Too Big is
exempt from the ICMPv6 rate limit unless the icmpv6_ratemask sysctl
includes it, so a packet draws one for each gateway with a smaller path
MTU. skb_tunnel_check_pmtu() is not used: with reply false it sends
nothing. A GSO skb, which a UDP_SEGMENT source sends once transmit
checksum offload is on, is judged by its segments, and marking it as a
UDP tunnel packet would leave the inner UDP header out of them, so the
marking moves here from the caller, after the check.

The tunnel MTU follows the route's path MTU, as on every IPv6 UDP
tunnel. RFC 7450 s5.3.3.6.1 asks for a way to turn that off, but IPv6
has none: udpv6_err() lowers the route MTU before it looks up a socket,
and "mtu lock" does not stop it. A forged Packet Too Big can lower the
route MTU to IPV6_MIN_MTU at most. An IPv4 payload that does not fit an
IPv6 tunnel is dropped without an ICMP error, as icmp_send() does not
answer multicast, and it is not fragmented before encapsulation as
s5.3.3.6.2.1 asks for DF=0; the IPv4 tunnel does not do that either.

Assisted-by: LLM
Signed-off-by: Omar Ramadan <omar@blockcast.net>
---
 drivers/net/amt.c | 93 +++++++++++++++++++++++++++++------------------
 1 file changed, 58 insertions(+), 35 deletions(-)

diff --git a/drivers/net/amt.c b/drivers/net/amt.c
index b977b00..412d23c 100644
--- a/drivers/net/amt.c
+++ b/drivers/net/amt.c
@@ -1103,15 +1103,43 @@ static void amt_req_work(struct work_struct *work)
 				 msecs_to_jiffies(100));
 }
 
+/* RFC 7450 s5.3.3.6: Multicast Data must not be fragmented over IPv6, so
+ * a payload above the tunnel MTU is dropped, and an IPv6 source is sent a
+ * Packet Too Big of at least IPV6_MIN_MTU, as it ignores less. icmp_send()
+ * never answers an IPv4 multicast datagram. A GSO skb is judged by its
+ * segments, so @skb must not be marked as a tunnel packet yet:
+ * skb->encapsulation would leave the inner UDP header out of their size.
+ */
+static bool amt_tmtu_exceeded(struct sk_buff *skb,
+			      const struct dst_entry *dst)
+{
+	int off = skb_network_offset(skb);
+	u32 mtu;
+
+	mtu = dst_mtu(dst) - sizeof(struct ipv6hdr) - sizeof(struct udphdr) -
+	      off;
+	if (skb_is_gso(skb) ? skb_gso_validate_network_len(skb, mtu) :
+			      skb->len - off <= mtu)
+		return false;
+
+	if (skb->protocol == htons(ETH_P_IPV6))
+		icmpv6_ndo_send(skb, ICMPV6_PKT_TOOBIG, 0,
+				max_t(u32, mtu, IPV6_MIN_MTU));
+	return true;
+}
+
 /* Route an AMT-encapsulated skb to @daddr and send it over the device's
- * outer family. The caller has already pushed the AMT header.
+ * outer family. The caller has already pushed the AMT header. @data marks
+ * Multicast Data, which may be a GSO skb: it gets the IPv6 tunnel MTU
+ * check, then the UDP tunnel offload marking.
  */
 static int amt_udp_xmit(struct amt_dev *amt, struct sock *sk,
 			struct sk_buff *skb, const union amt_addr *daddr,
-			__be16 sport, __be16 dport)
+			__be16 sport, __be16 dport, bool data)
 {
 	struct rtable *rt;
 	struct flowi4 fl4;
+	int err;
 
 	if (amt_v6(amt)) {
 		struct dst_entry *dst;
@@ -1123,6 +1151,14 @@ static int amt_udp_xmit(struct amt_dev *amt, struct sock *sk,
 				   &daddr->ip6);
 			return PTR_ERR(dst);
 		}
+		if (data) {
+			err = amt_tmtu_exceeded(skb, dst) ? -EMSGSIZE :
+			      udp_tunnel_handle_offloads(skb, true);
+			if (err) {
+				dst_release(dst);
+				return err;
+			}
+		}
 		/* iptunnel_xmit() scrubs the IPv4 tunnel skb, but
 		 * udp_tunnel6_xmit_skb() does not, so drop the inner
 		 * packet's conntrack and extensions here. Links cannot
@@ -1136,11 +1172,17 @@ static int amt_udp_xmit(struct amt_dev *amt, struct sock *sk,
 		return 0;
 	}
 
+	if (data) {
+		err = udp_tunnel_handle_offloads(skb, true);
+		if (err)
+			return err;
+	}
+
 	memset(&fl4, 0, sizeof(struct flowi4));
 	fl4.flowi4_oif         = amt->stream_dev->ifindex;
 	fl4.daddr              = daddr->ip4;
 	fl4.saddr              = amt->local_ip;
-	fl4.flowi4_dscp        = inet_dsfield_to_dscp(AMT_TOS);
+	fl4.flowi4_dscp        = data ? 0 : inet_dsfield_to_dscp(AMT_TOS);
 	fl4.flowi4_proto       = IPPROTO_UDP;
 	rt = ip_route_output_key(amt->net, &fl4);
 	if (IS_ERR(rt)) {
@@ -1229,16 +1271,14 @@ static void amt_send_multicast_data(struct amt_dev *amt,
 {
 	struct amt_header_mcast_data *amtmd;
 	struct sk_buff *skb;
-	struct iphdr *iph;
-	struct flowi4 fl4;
-	struct rtable *rt;
 	struct sock *sk;
+	int err;
 
 	sk = rcu_dereference_bh(amt->sk);
 	if (!sk)
 		return;
 
-	skb = skb_copy_expand(oskb, sizeof(*amtmd) + sizeof(*iph) +
+	skb = skb_copy_expand(oskb, sizeof(*amtmd) + amt_ip_hlen(amt) +
 			      sizeof(struct udphdr), 0, GFP_ATOMIC);
 	if (!skb)
 		return;
@@ -1249,22 +1289,6 @@ static void amt_send_multicast_data(struct amt_dev *amt,
 	 */
 	skb_reset_mac_header(skb);
 	skb_reset_inner_headers(skb);
-	if (udp_tunnel_handle_offloads(skb, true)) {
-		kfree_skb(skb);
-		return;
-	}
-
-	memset(&fl4, 0, sizeof(struct flowi4));
-	fl4.flowi4_oif         = amt->stream_dev->ifindex;
-	fl4.daddr              = tunnel->addr.ip4;
-	fl4.saddr              = amt->local_ip;
-	fl4.flowi4_proto       = IPPROTO_UDP;
-	rt = ip_route_output_key(amt->net, &fl4);
-	if (IS_ERR(rt)) {
-		netdev_dbg(amt->dev, "no route to %pI4\n", &tunnel->addr.ip4);
-		kfree_skb(skb);
-		return;
-	}
 
 	amtmd = skb_push(skb, sizeof(*amtmd));
 	amtmd->version = 0;
@@ -1275,17 +1299,16 @@ static void amt_send_multicast_data(struct amt_dev *amt,
 		skb_set_inner_protocol(skb, htons(ETH_P_IP));
 	else
 		skb_set_inner_protocol(skb, htons(ETH_P_IPV6));
-	udp_tunnel_xmit_skb(rt, sk, skb,
-			    fl4.saddr,
-			    fl4.daddr,
-			    AMT_TOS,
-			    ip4_dst_hoplimit(&rt->dst),
-			    0,
-			    amt->relay_port,
-			    tunnel->source_port,
-			    false,
-			    false,
-			    0);
+	err = amt_udp_xmit(amt, sk, skb, &tunnel->addr, amt->relay_port,
+			   tunnel->source_port, true);
+	if (err) {
+		DEV_STATS_INC(amt->dev, tx_dropped);
+		kfree_skb_reason(skb, err == -EMSGSIZE ?
+				      SKB_DROP_REASON_PKT_TOO_BIG :
+				      err == -ENOMEM ?
+				      SKB_DROP_REASON_NOMEM :
+				      SKB_DROP_REASON_IP_OUTNOROUTES);
+	}
 }
 
 static bool amt_send_membership_query(struct amt_dev *amt,
@@ -1321,7 +1344,7 @@ static bool amt_send_membership_query(struct amt_dev *amt,
 	else
 		skb_set_inner_protocol(skb, htons(ETH_P_IPV6));
 	if (amt_udp_xmit(amt, sk, skb, &tunnel->addr, amt->relay_port,
-			 tunnel->source_port))
+			 tunnel->source_port, false))
 		return true;
 	amt_update_relay_status(tunnel, AMT_STATUS_SENT_QUERY, true);
 	return false;
-- 
2.43.0


  parent reply	other threads:[~2026-10-09 12:24 UTC|newest]

Thread overview: 23+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09 12:24 [PATCH net-next 00/13] amt: add an IPv6 outer transport Omar Ramadan
2026-10-09 12:24 ` [PATCH net-next 01/13] amt: create an AF_INET6 encapsulation socket for an IPv6 outer address Omar Ramadan
2026-10-09 12:24 ` [PATCH net-next 02/13] amt: send the Relay Advertisement over IPv6 Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 03/13] amt: key relay tunnels on a union amt_addr endpoint Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 04/13] amt: send the Membership Query over IPv6 Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 05/13] amt: match the Membership Update tunnel by outer family Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` Omar Ramadan [this message]
2026-10-10 12:41   ` [PATCH net-next 06/13] amt: forward multicast data over IPv6 netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 07/13] amt: size the encapsulation headroom by the outer IP version Omar Ramadan
2026-10-09 12:24 ` [PATCH net-next 08/13] amt: send the AMT gateway control plane over IPv6 Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 09/13] amt: receive " Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 10/13] amt: add netlink attributes for an IPv6 outer transport Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 11/13] MAINTAINERS: amt: cover the amt headers and selftests Omar Ramadan
2026-10-09 12:24 ` [PATCH net-next 12/13] selftests: net: add amt_v6.sh for an IPv6 outer transport Omar Ramadan
2026-10-10 12:41   ` netdev-bot+sashiko
2026-10-09 12:24 ` [PATCH net-next 13/13] selftests: net: add amt_gw_v6.sh for the IPv6 netlink attributes Omar Ramadan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261009122426.551178-7-omar@blockcast.net \
    --to=omar@blockcast.net \
    --cc=andrew+netdev@lunn.ch \
    --cc=ap420073@gmail.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@kernel.org \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=shuah@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®