mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons
@ 2026-09-30 18:38 Anton Danilov
  2026-09-30 18:38 ` [PATCH net-next v5 01/14] vxlan: rename the drop reasons for use by other tunnels Anton Danilov
                   ` (13 more replies)
  0 siblings, 14 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:38 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

Only vxlan reports drop reasons among the tunnel drivers today. The ones
converted here free what they drop with kfree_skb(), which drop_monitor
and the skb:kfree_skb tracepoint do report, but as NOT_SPECIFIED, with
the call site as the only hint at which check failed. That hint does not
go far: all the failures of ip_tunnel_rcv() end at one call site, and so
do nearly all of those of each transmit function. The call site is not a
stable key either: its offset moves with the compiler, inlining and the
configuration, and the debug info that maps it back to a source line
only serves that one build, as the line moves with any change to the
file. So a filter on it has to follow every rebuild. The reason does not
move with the build: NET_DM_ATTR_REASON reports it, and a BPF program
can match it by name. The device counters group the failures coarsely
too: rx_errors and tx_errors each lump together unrelated conditions.

This series covers the generic paths shared by ipip, sit, gre and their
IPv6 counterparts, plus the GRE and ip6tnl specific code, in both
directions; the code specific to ipip and sit keeps kfree_skb(). So does
the tunnel4 and tunnel6 demux, shared with the xfrm tunnels, which frees
the packets that no ipip, sit or ip6tnl device takes. ip6tnl is in
because its transmit handler goes through ip6_tnl_xmit(), which changes
along with ip6_gre, and its receive side is converted to match. The
series also makes two vxlan reasons generic, so that GRE reports the
same ones.

Patch 1 renames VXLAN_INVALID_HDR and VXLAN_VNI_NOT_FOUND to
TUNNEL_INVALID_HDR and TUNNEL_NOT_FOUND. Their values stay the same;
the names vxlan drops are reported with change.

Patch 2 makes __iptunnel_pull_header() and iptunnel_pull_header() return
a drop reason. Until now they returned -ENOMEM for any failure, so a
packet too short to pull looked like an allocation failure: vxlan
reports it as NOMEM, and ip6_gre would do the same for a packet from any
sender whose inner Ethernet header or WCCPv2 word is cut short, since it
pulls the header before the tunnel lookup. Patch 3 has vxlan report the
reason these functions now return.

Patches 4-5 convert the generic receive paths, ip_tunnel_rcv() and
__ip6_tnl_rcv(); patch 5 also converts ipxip6_rcv(), the ip6tnl handler
in front of the latter. Two reasons are added:

  TUNNEL_OPT_MISMATCH  the options a packet carries do not match the
                       tunnel configuration
  TUNNEL_OLD_SEQ       the sequence number is older than the one the
                       tunnel expects, like TCP_OLD_SEQUENCE for TCP

The second one has a failure mode worth naming: when a peer reboots, its
outgoing sequence number restarts at zero, and the receiver can drop
everything until the peer's numbers get past the last one the receiver
accepted. By the counters alone that looks like a misconfiguration: a
packet without the sequence number option bumps the same rx_fifo_errors.

Patches 6-8 convert the GRE specific receive path. gre_parse_header()
returns -EINVAL for every failure, and the only detail its callers could
get was a csum_err flag that none of them read: both ip_gre and ip6_gre
declared it, passed it in and ignored it. gre_parse_header() now returns
a drop reason instead, and its callers take the header length from
tpi->hdr_len. The receive helpers below gre_rcv() return the drop reason
instead of a PACKET_* code, and the PACKET_* codes go away. The GRE
receive paths report TUNNEL_INVALID_HDR and TUNNEL_NOT_FOUND, the length
checks report what pskb_may_pull_reason() returns, and one reason is
added, GRE_CSUM, like TCP_CSUM and UDP_CSUM.

Patches 9-13 convert the transmit side of ip_tunnel, ip_gre, ip6_tunnel
and ip6_gre. One reason is added, TUNNEL_ENCAP, for a failure to build
the encapsulation header. Patch 11 is a small preparation:
prepare_ip6gre_xmit_other() cannot fail, so it is made void rather than
given a drop reason for a branch that never runs. ip6_tnl_xmit() leaves
freeing the packet to its callers, so patch 12 makes it and the helpers
between it and the ndo_start_xmit handlers return the drop reason
instead of an error, and patch 13 converts the ip6_gre handlers. Patch
14 has vxlan report a route that goes out of the vxlan device itself as
RECURSION_LIMIT, as ip_tunnel and ip6_tunnel now do, instead of
IP_OUTNOROUTES.

The transmit side has its own case worth naming: tnl_update_pmtu()
returns -E2BIG after it has already sent the ICMP error back, which is
path MTU discovery working exactly as intended, yet among the device
counters the drop only bumps tx_errors, like a failed encapsulation and
a few other failures do. If the ICMP error never reaches the sender,
the resulting MTU black hole cannot be told from those by the device
counters; the reason tells them apart.

Drop reasons on transmit are not new: vxlan already reports several from
its xmit path, and ip_tunnel_core.c reports RECURSION_LIMIT. They are
most useful for forwarded packets, which is what a tunnel gateway mostly
transmits: the sender is another host, which gets an ICMP error for only
some of these failures, so the drop has to be explained on the gateway.

The reason variable follows one rule in the functions this series
converts that have a drop label: it is initialized to
SKB_DROP_REASON_NOT_SPECIFIED unless the first statement of the function
sets it, and where a helper stores its result in it, the drop label
falls back to SKB_DROP_REASON_NOT_SPECIFIED, as in vxlan_rcv(). Together
they keep a drop without a reason of its own, including one that a later
change adds, from freeing a packet with SKB_NOT_DROPPED_YET. Every drop
path of the series sets a reason, except the looped back multicast check
in ip_gre's gre_rcv().

Changes since v4:
- new patch 1 renames VXLAN_INVALID_HDR and VXLAN_VNI_NOT_FOUND to
  TUNNEL_INVALID_HDR and TUNNEL_NOT_FOUND, and GRE reports those instead
  of the GRE_INVALID_HDR and GRE_TUNNEL_NOT_FOUND added by v4 (Ido)
- the TNL_* reasons are renamed to TUNNEL_*, to match
- the length checks report what pskb_may_pull_reason() returns instead
  of HDR_TRUNC; the WCCP check, which uses skb_header_pointer(), reports
  PKT_TOO_SMALL (Ido)
- gre_rcv() in the demux reports TUNNEL_INVALID_HDR instead of
  UNHANDLED_PROTO for a version it cannot dispatch (Ido)
- __iptunnel_pull_header() and iptunnel_pull_header() return the drop
  reason themselves, instead of the __iptunnel_pull_header_reason() v4
  added; the three callers that tested the result with "< 0" test for a
  non-zero value (Ido)
- new patch 3 has vxlan report the reason __iptunnel_pull_header()
  returns (Ido)
- ip_tunnel_xmit() sets the reason for an NBMA payload it cannot derive
  a destination from in the branch that uses it (Ido); the reason is now
  NO_TX_TARGET, as for the other NBMA packets without a destination
- the gre_parse_header() argument icmp_err is called ignore_csum_err,
  after what it does, and the function has a kernel-doc comment
- the cover letter no longer announces a series for other tunnel
  drivers (Ido)
- ipxip6_rcv() reports drop reasons too (patch 5), so the ip6tnl code
  is converted in both directions
- the ip6_tunnel transmit patch is split in two: ip6_tnl_xmit() and the
  ip6_gre helpers (patch 12), then the ip6_gre handlers (patch 13)
- the reason is initialized to NOT_SPECIFIED only where the first
  statement of a function does not set it, and a drop label reached
  after a helper falls back to NOT_SPECIFIED, as in vxlan_rcv()
- new patch 14 has vxlan report a circular route as RECURSION_LIMIT, as
  ip_tunnel and ip6_tunnel do
- the description of RECURSION_LIMIT gives the tunnel case as an
  example (patch 9)
- a transmit through an NBMA tunnel that finds no endpoint for the
  packet, and through an ip6_gre device without a remote, reports
  NO_TX_TARGET, as on the IPv4 side, instead of DEV_READY; the message
  of patch 12 describes the cases DEV_READY still covers
- ip6gre_tunnel_xmit() drops a packet without IPv6 metadata on a
  collect_md device as TUNNEL_TXINFO right after it looks the metadata
  up, as ip6erspan does, so the DAD probes of an ip6gretap device are no
  longer reported as RECURSION_LIMIT (patch 13)
- the Assisted-by tags take the form that
  Documentation/process/coding-assistants.rst gives

- v4: https://lore.kernel.org/netdev/20260922221507.3268127-1-littlesmilingcloud@gmail.com/
- v3: https://lore.kernel.org/netdev/20260916143717.1875082-1-littlesmilingcloud@gmail.com/
- v2: https://lore.kernel.org/netdev/20260913034937.875068-1-littlesmilingcloud@gmail.com/
- v1: https://lore.kernel.org/netdev/20260831215137.549324-1-littlesmilingcloud@gmail.com/

Anton Danilov (14):
  vxlan: rename the drop reasons for use by other tunnels
  ip_tunnel: make __iptunnel_pull_header() return a drop reason
  vxlan: report the drop reason of __iptunnel_pull_header()
  ip_tunnel: add drop reasons to the generic RX path
  ip6_tunnel: add drop reasons to the receive path
  gre: make gre_parse_header() report a drop reason
  ip_gre: add drop reasons to the RX path
  ip6_gre: add drop reasons to the RX path
  ip_tunnel: add drop reasons to the transmit path
  ip_gre: add drop reasons to the transmit path
  ip6_gre: make prepare_ip6gre_xmit_other() void
  ip6_tunnel: make ip6_tnl_xmit() return a drop reason
  ip6_gre: add drop reasons to the transmit path
  vxlan: report a circular route as SKB_DROP_REASON_RECURSION_LIMIT

 drivers/net/vxlan/vxlan_core.c |  26 ++--
 include/net/dropreason-core.h  |  53 ++++++--
 include/net/gre.h              |   5 +-
 include/net/ip6_tunnel.h       |   5 +-
 include/net/ip_tunnels.h       |  14 +-
 net/ipv4/gre_demux.c           |  65 ++++++---
 net/ipv4/ip_gre.c              | 199 +++++++++++++++++----------
 net/ipv4/ip_tunnel.c           |  64 ++++++---
 net/ipv4/ip_tunnel_core.c      |  22 ++-
 net/ipv6/ip6_gre.c             | 236 ++++++++++++++++++++-------------
 net/ipv6/ip6_tunnel.c          | 153 +++++++++++++--------
 11 files changed, 556 insertions(+), 286 deletions(-)

-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 01/14] vxlan: rename the drop reasons for use by other tunnels
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
@ 2026-09-30 18:38 ` Anton Danilov
  2026-09-30 18:38 ` [PATCH net-next v5 02/14] ip_tunnel: make __iptunnel_pull_header() return a drop reason Anton Danilov
                   ` (12 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:38 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

VXLAN_INVALID_HDR and VXLAN_VNI_NOT_FOUND describe conditions that are
not specific to vxlan: a malformed tunnel header, and a packet for
which no tunnel device is found. The GRE receive paths converted later
in this series drop packets for the same two conditions.

Rename them to TUNNEL_INVALID_HDR and TUNNEL_NOT_FOUND and reword their
descriptions, the way VXLAN_NO_REMOTE became NO_TX_TARGET in commit
46e0ccfb88f0 ("net: vxlan: rename SKB_DROP_REASON_VXLAN_NO_REMOTE").
The numeric values stay the same; the names that the skb:kfree_skb
tracepoint and drop_monitor report change.

Suggested-by: Ido Schimmel <idosch@nvidia.com>
Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 drivers/net/vxlan/vxlan_core.c |  8 ++++----
 include/net/dropreason-core.h  | 19 +++++++++++--------
 2 files changed, 15 insertions(+), 12 deletions(-)

diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index 27b0b6567d52..35a47cc27200 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -1700,7 +1700,7 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 	if (!(vh->vx_flags & VXLAN_HF_VNI)) {
 		netdev_dbg(skb->dev, "invalid vxlan flags=%#x vni=%#x\n",
 			   ntohl(vh->vx_flags), ntohl(vh->vx_vni));
-		reason = SKB_DROP_REASON_VXLAN_INVALID_HDR;
+		reason = SKB_DROP_REASON_TUNNEL_INVALID_HDR;
 		/* Return non vxlan pkt */
 		goto drop;
 	}
@@ -1713,7 +1713,7 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 
 	vxlan = vxlan_vs_find_vni(vs, skb->dev->ifindex, vni, &vninode);
 	if (!vxlan) {
-		reason = SKB_DROP_REASON_VXLAN_VNI_NOT_FOUND;
+		reason = SKB_DROP_REASON_TUNNEL_NOT_FOUND;
 		goto drop;
 	}
 
@@ -1729,7 +1729,7 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		 * previous stack code, and also is more robust and provides a
 		 * little more security in adding extensions to VXLAN.
 		 */
-		reason = SKB_DROP_REASON_VXLAN_INVALID_HDR;
+		reason = SKB_DROP_REASON_TUNNEL_INVALID_HDR;
 		DEV_STATS_INC(vxlan->dev, rx_frame_errors);
 		DEV_STATS_INC(vxlan->dev, rx_errors);
 		vxlan_vnifilter_count(vxlan, cfg, vni, vninode,
@@ -2389,7 +2389,7 @@ static int encap_bypass_if_local(struct sk_buff *skb, struct net_device *dev,
 			DEV_STATS_INC(dev, tx_errors);
 			vxlan_vnifilter_count(vxlan, cfg, vni, NULL,
 					      VXLAN_VNI_STATS_TX_ERRORS, 0);
-			kfree_skb_reason(skb, SKB_DROP_REASON_VXLAN_VNI_NOT_FOUND);
+			kfree_skb_reason(skb, SKB_DROP_REASON_TUNNEL_NOT_FOUND);
 
 			return -ENOENT;
 		}
diff --git a/include/net/dropreason-core.h b/include/net/dropreason-core.h
index 12f909651591..40a27d8887af 100644
--- a/include/net/dropreason-core.h
+++ b/include/net/dropreason-core.h
@@ -111,8 +111,8 @@
 	FN(PACKET_SOCK_ERROR)		\
 	FN(TC_CHAIN_NOTFOUND)		\
 	FN(TC_RECLASSIFY_LOOP)		\
-	FN(VXLAN_INVALID_HDR)		\
-	FN(VXLAN_VNI_NOT_FOUND)		\
+	FN(TUNNEL_INVALID_HDR)		\
+	FN(TUNNEL_NOT_FOUND)		\
 	FN(MAC_INVALID_SOURCE)		\
 	FN(VXLAN_ENTRY_EXISTS)		\
 	FN(NO_TX_TARGET)		\
@@ -539,13 +539,16 @@ enum skb_drop_reason {
 	 */
 	SKB_DROP_REASON_TC_RECLASSIFY_LOOP,
 	/**
-	 * @SKB_DROP_REASON_VXLAN_INVALID_HDR: VXLAN header is invalid. E.g.:
-	 * 1) reserved fields are not zero
-	 * 2) "I" flag is not set
+	 * @SKB_DROP_REASON_TUNNEL_INVALID_HDR: tunnel header is invalid. E.g.:
+	 * 1) VXLAN reserved fields are not zero
+	 * 2) VXLAN "I" flag is not set
 	 */
-	SKB_DROP_REASON_VXLAN_INVALID_HDR,
-	/** @SKB_DROP_REASON_VXLAN_VNI_NOT_FOUND: no VXLAN device found for VNI */
-	SKB_DROP_REASON_VXLAN_VNI_NOT_FOUND,
+	SKB_DROP_REASON_TUNNEL_INVALID_HDR,
+	/**
+	 * @SKB_DROP_REASON_TUNNEL_NOT_FOUND: no tunnel device found for the
+	 * packet, e.g. no VXLAN device for its VNI
+	 */
+	SKB_DROP_REASON_TUNNEL_NOT_FOUND,
 	/** @SKB_DROP_REASON_MAC_INVALID_SOURCE: source mac is invalid */
 	SKB_DROP_REASON_MAC_INVALID_SOURCE,
 	/**
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 02/14] ip_tunnel: make __iptunnel_pull_header() return a drop reason
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
  2026-09-30 18:38 ` [PATCH net-next v5 01/14] vxlan: rename the drop reasons for use by other tunnels Anton Danilov
@ 2026-09-30 18:38 ` Anton Danilov
  2026-09-30 18:38 ` [PATCH net-next v5 03/14] vxlan: report the drop reason of __iptunnel_pull_header() Anton Danilov
                   ` (11 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:38 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

__iptunnel_pull_header() returns -ENOMEM whenever it fails. It can
fail in two pskb_may_pull() calls, one for the tunnel header and one
for the inner Ethernet header of ETH_P_TEB, and in the skb_unclone()
done for GSO packets. pskb_may_pull() fails when the packet is shorter
than the requested length as well as when pulling from the frags cannot
allocate, so a truncated packet and an allocation failure look the same
to the callers. The ones that report a drop reason can only pick
SKB_DROP_REASON_NOMEM, as vxlan_rcv() does, and so would the GRE
receive paths converted by the following patches. In ip6_gre,
gre_rcv() pulls the header before the tunnel lookup, so a packet from
any sender whose ETH_P_TEB inner Ethernet header or WCCPv2 extra word
is cut short would be reported as an out of memory condition.

Make __iptunnel_pull_header() and iptunnel_pull_header() return the
reason pskb_may_pull_reason() already computes, SKB_DROP_REASON_NOMEM
when skb_unclone() fails, and SKB_NOT_DROPPED_YET on success. A failure
is still non-zero, so the callers that only test the result keep
working. Three callers, in ip_gre and ip6_gre, test it with "< 0"
instead; the enum has no negative values, so that test would always be
false, and the compiler does not warn about it. Make them test for a
non-zero value.

Suggested-by: Ido Schimmel <idosch@nvidia.com>
Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 include/net/ip_tunnels.h  | 10 ++++++----
 net/ipv4/ip_gre.c         |  4 ++--
 net/ipv4/ip_tunnel_core.c | 22 +++++++++++++++-------
 net/ipv6/ip6_gre.c        |  2 +-
 4 files changed, 24 insertions(+), 14 deletions(-)

diff --git a/include/net/ip_tunnels.h b/include/net/ip_tunnels.h
index 7102aa11fae2..5cccf4c0e691 100644
--- a/include/net/ip_tunnels.h
+++ b/include/net/ip_tunnels.h
@@ -614,11 +614,13 @@ static inline u8 ip_tunnel_ecn_encap(u8 tos, const struct iphdr *iph,
 	return INET_ECN_encapsulate(tos, inner);
 }
 
-int __iptunnel_pull_header(struct sk_buff *skb, int hdr_len,
-			   __be16 inner_proto, bool raw_proto, bool xnet);
+enum skb_drop_reason
+__iptunnel_pull_header(struct sk_buff *skb, int hdr_len,
+		       __be16 inner_proto, bool raw_proto, bool xnet);
 
-static inline int iptunnel_pull_header(struct sk_buff *skb, int hdr_len,
-				       __be16 inner_proto, bool xnet)
+static inline enum skb_drop_reason
+iptunnel_pull_header(struct sk_buff *skb, int hdr_len,
+		     __be16 inner_proto, bool xnet)
 {
 	return __iptunnel_pull_header(skb, hdr_len, inner_proto, false, xnet);
 }
diff --git a/net/ipv4/ip_gre.c b/net/ipv4/ip_gre.c
index 51fcd603939c..3ee437fa3eae 100644
--- a/net/ipv4/ip_gre.c
+++ b/net/ipv4/ip_gre.c
@@ -312,7 +312,7 @@ static int erspan_rcv(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 		if (__iptunnel_pull_header(skb,
 					   len,
 					   htons(ETH_P_TEB),
-					   false, false) < 0)
+					   false, false))
 			goto drop;
 
 		if (tunnel->collect_md) {
@@ -378,7 +378,7 @@ static int __ipgre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi,
 		const struct iphdr *tnl_params;
 
 		if (__iptunnel_pull_header(skb, hdr_len, tpi->proto,
-					   raw_proto, false) < 0)
+					   raw_proto, false))
 			goto drop;
 
 		/* Special case for ipgre_header_parse(), which expects the
diff --git a/net/ipv4/ip_tunnel_core.c b/net/ipv4/ip_tunnel_core.c
index bab42b9e277f..6c2855adff28 100644
--- a/net/ipv4/ip_tunnel_core.c
+++ b/net/ipv4/ip_tunnel_core.c
@@ -106,19 +106,24 @@ void iptunnel_xmit(struct sock *sk, struct rtable *rt, struct sk_buff *skb,
 }
 EXPORT_SYMBOL_GPL(iptunnel_xmit);
 
-int __iptunnel_pull_header(struct sk_buff *skb, int hdr_len,
-			   __be16 inner_proto, bool raw_proto, bool xnet)
+enum skb_drop_reason
+__iptunnel_pull_header(struct sk_buff *skb, int hdr_len,
+		       __be16 inner_proto, bool raw_proto, bool xnet)
 {
-	if (unlikely(!pskb_may_pull(skb, hdr_len)))
-		return -ENOMEM;
+	enum skb_drop_reason reason;
+
+	reason = pskb_may_pull_reason(skb, hdr_len);
+	if (unlikely(reason))
+		return reason;
 
 	skb_pull_rcsum(skb, hdr_len);
 
 	if (!raw_proto && inner_proto == htons(ETH_P_TEB)) {
 		struct ethhdr *eh;
 
-		if (unlikely(!pskb_may_pull(skb, ETH_HLEN)))
-			return -ENOMEM;
+		reason = pskb_may_pull_reason(skb, ETH_HLEN);
+		if (unlikely(reason))
+			return reason;
 
 		eh = (struct ethhdr *)skb->data;
 		if (likely(eth_proto_is_802_3(eh->h_proto)))
@@ -135,7 +140,10 @@ int __iptunnel_pull_header(struct sk_buff *skb, int hdr_len,
 	skb_set_queue_mapping(skb, 0);
 	skb_scrub_packet(skb, xnet);
 
-	return iptunnel_pull_offloads(skb);
+	if (unlikely(iptunnel_pull_offloads(skb)))
+		return SKB_DROP_REASON_NOMEM;
+
+	return SKB_NOT_DROPPED_YET;
 }
 EXPORT_SYMBOL_GPL(__iptunnel_pull_header);
 
diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c
index 258239a7c53b..0392f6ba862b 100644
--- a/net/ipv6/ip6_gre.c
+++ b/net/ipv6/ip6_gre.c
@@ -512,7 +512,7 @@ static int ip6erspan_rcv(struct sk_buff *skb,
 
 		if (__iptunnel_pull_header(skb, len,
 					   htons(ETH_P_TEB),
-					   false, false) < 0)
+					   false, false))
 			return PACKET_REJECT;
 
 		if (tunnel->parms.collect_md) {
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 03/14] vxlan: report the drop reason of __iptunnel_pull_header()
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
  2026-09-30 18:38 ` [PATCH net-next v5 01/14] vxlan: rename the drop reasons for use by other tunnels Anton Danilov
  2026-09-30 18:38 ` [PATCH net-next v5 02/14] ip_tunnel: make __iptunnel_pull_header() return a drop reason Anton Danilov
@ 2026-09-30 18:38 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 04/14] ip_tunnel: add drop reasons to the generic RX path Anton Danilov
                   ` (10 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:38 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

vxlan_rcv() reports SKB_DROP_REASON_NOMEM whenever
__iptunnel_pull_header() fails, as the helper used to return -ENOMEM
for any failure. The VXLAN header is already in the linear area at this
point, but the inner Ethernet header of a packet without GPE may not
be, so a packet whose inner frame is shorter than an Ethernet header is
reported as an out of memory condition.

Report the reason __iptunnel_pull_header() now returns instead:
PKT_TOO_SMALL for such a packet, NOMEM when an allocation fails.

Suggested-by: Ido Schimmel <idosch@nvidia.com>
Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 drivers/net/vxlan/vxlan_core.c | 8 ++++----
 1 file changed, 4 insertions(+), 4 deletions(-)

diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index 35a47cc27200..0d98e29e2d72 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -1743,11 +1743,11 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		raw_proto = true;
 	}
 
-	if (__iptunnel_pull_header(skb, VXLAN_HLEN, protocol, raw_proto,
-				   !net_eq(vxlan->net, dev_net(vxlan->dev)))) {
-		reason = SKB_DROP_REASON_NOMEM;
+	reason = __iptunnel_pull_header(skb, VXLAN_HLEN, protocol, raw_proto,
+					!net_eq(vxlan->net,
+						dev_net(vxlan->dev)));
+	if (reason)
 		goto drop;
-	}
 
 	if (cfg->flags & VXLAN_F_REMCSUM_RX) {
 		reason = vxlan_remcsum(skb, cfg->flags);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 04/14] ip_tunnel: add drop reasons to the generic RX path
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (2 preceding siblings ...)
  2026-09-30 18:38 ` [PATCH net-next v5 03/14] vxlan: report the drop reason of __iptunnel_pull_header() Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 05/14] ip6_tunnel: add drop reasons to the receive path Anton Danilov
                   ` (9 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

ip_tunnel_rcv() collapses four distinct failures into the single
kfree_skb() under its drop label:

 - the tunnel options carried by the packet do not match the tunnel
   configuration (checksum or sequence number),
 - the sequence number is older than the expected one,
 - the inner network header cannot be pulled,
 - the ECN decapsulation check fails (RFC 6040).

drop_monitor and the skb:kfree_skb tracepoint do see these packets, but
all four as SKB_DROP_REASON_NOT_SPECIFIED from the same call site, so
neither the reason nor the location tells the failures apart. Only the
device error counters (rx_crc_errors, rx_fifo_errors, rx_length_errors,
rx_frame_errors) hint at the cause, and they are not tied to the
dropped packet.

Add two drop reasons for the tunnel specific cases and reuse the
existing ones for the rest:

 - SKB_DROP_REASON_TUNNEL_OPT_MISMATCH is used when the packet lacks the
   checksum or the sequence number option the tunnel is configured for,
   or carries a checksum the tunnel is not configured for, as the
   checksum check compares both ways. This is a configuration mismatch
   between the two endpoints rather than a corrupted checksum: the
   checksum itself is validated earlier, in gre_parse_header().

 - SKB_DROP_REASON_TUNNEL_OLD_SEQ is used when the sequence number is
   older than the expected one. Unlike the previous one, this is a
   property of the received traffic: a remote endpoint that restarts and
   resets its sequence numbering can have all of its packets dropped
   until its sequence numbers catch up with i_seqno.

 - pskb_inet_may_pull_reason() already computes a drop reason,
   SKB_DROP_REASON_PKT_TOO_SMALL or SKB_DROP_REASON_NOMEM, which has so
   far been discarded.

 - SKB_DROP_REASON_IP_TUNNEL_ECN already exists and documents exactly
   this check, but until now it was only used by vxlan.

The sequence number check is split in two so that the two cases can be
told apart. The error counters are left unchanged.

ip_tunnel_rcv() is the RX path of ip_gre, ipip and sit, the latter for
IPv4 and MPLS payloads only: ipip6_rcv() handles IPv6 in IPv4 on its
own. The checksum and the sequence number options only exist for GRE,
so the two new reasons are meant for ip_gre. ipip and sit carry
neither option and only hit SKB_DROP_REASON_TUNNEL_OPT_MISMATCH if their
i_flags are given those bits, for instance through IFLA_IPTUN_FLAGS;
such a device then drops every packet that reaches ip_tunnel_rcv(),
with or without this patch. The ECN reason applies to all three, while
the length ones can only be hit through ip_gre: tunnel4_rcv() has
already pulled the inner IPv4 header for ipip and sit, and
pskb_inet_may_pull_reason() has nothing to pull for an MPLS payload.

ip_tunnel_rcv() initializes the reason to SKB_DROP_REASON_NOT_SPECIFIED,
since its first statement does not set it, and its drop label falls back
to SKB_DROP_REASON_NOT_SPECIFIED, as in vxlan_rcv(), since
pskb_inet_may_pull_reason() stores its result in the reason. Together
they keep a drop without a reason of its own, including one that a
later change adds, from freeing the packet with SKB_NOT_DROPPED_YET. The
rest of the series follows the same rule wherever a function has a drop
label: an initializer unless the first statement sets the reason, and
the fallback wherever a helper stores its result in it.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 include/net/dropreason-core.h | 16 ++++++++++++++++
 net/ipv4/ip_tunnel.c          | 20 ++++++++++++++++----
 2 files changed, 32 insertions(+), 4 deletions(-)

diff --git a/include/net/dropreason-core.h b/include/net/dropreason-core.h
index 40a27d8887af..85aaa58c3ecb 100644
--- a/include/net/dropreason-core.h
+++ b/include/net/dropreason-core.h
@@ -129,6 +129,8 @@
 	FN(PSP_INPUT)			\
 	FN(PSP_OUTPUT)			\
 	FN(RECURSION_LIMIT)		\
+	FN(TUNNEL_OPT_MISMATCH)		\
+	FN(TUNNEL_OLD_SEQ)		\
 	FNe(MAX)
 
 /**
@@ -615,6 +617,20 @@ enum skb_drop_reason {
 	SKB_DROP_REASON_PSP_OUTPUT,
 	/** @SKB_DROP_REASON_RECURSION_LIMIT: Dead loop on virtual device. */
 	SKB_DROP_REASON_RECURSION_LIMIT,
+	/**
+	 * @SKB_DROP_REASON_TUNNEL_OPT_MISMATCH: the tunnel options carried by
+	 * the packet do not match the tunnel configuration, e.g. a GRE tunnel
+	 * configured with 'icsum' or 'iseq' received a packet with no checksum
+	 * or no sequence number, or a GRE tunnel without 'icsum' received a
+	 * packet with a checksum.
+	 */
+	SKB_DROP_REASON_TUNNEL_OPT_MISMATCH,
+	/**
+	 * @SKB_DROP_REASON_TUNNEL_OLD_SEQ: the sequence number carried by the
+	 * packet is older than the one expected by the tunnel, e.g. after the
+	 * remote endpoint restarted and reset its sequence numbering.
+	 */
+	SKB_DROP_REASON_TUNNEL_OLD_SEQ,
 	/**
 	 * @SKB_DROP_REASON_MAX: the maximum of core drop reasons, which
 	 * shouldn't be used as a real 'reason' - only for tracing code gen
diff --git a/net/ipv4/ip_tunnel.c b/net/ipv4/ip_tunnel.c
index 0875474a578a..6d500751f837 100644
--- a/net/ipv4/ip_tunnel.c
+++ b/net/ipv4/ip_tunnel.c
@@ -384,6 +384,7 @@ int ip_tunnel_rcv(struct ip_tunnel *tunnel, struct sk_buff *skb,
 		  const struct tnl_ptk_info *tpi, struct metadata_dst *tun_dst,
 		  bool log_ecn_error)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	const struct iphdr *iph = ip_hdr(skb);
 	int nh, err;
 
@@ -398,14 +399,22 @@ int ip_tunnel_rcv(struct ip_tunnel *tunnel, struct sk_buff *skb,
 	    test_bit(IP_TUNNEL_CSUM_BIT, tpi->flags)) {
 		DEV_STATS_INC(tunnel->dev, rx_crc_errors);
 		DEV_STATS_INC(tunnel->dev, rx_errors);
+		reason = SKB_DROP_REASON_TUNNEL_OPT_MISMATCH;
 		goto drop;
 	}
 
 	if (test_bit(IP_TUNNEL_SEQ_BIT, tunnel->parms.i_flags)) {
-		if (!test_bit(IP_TUNNEL_SEQ_BIT, tpi->flags) ||
-		    (tunnel->i_seqno && (s32)(ntohl(tpi->seq) - tunnel->i_seqno) < 0)) {
+		if (!test_bit(IP_TUNNEL_SEQ_BIT, tpi->flags)) {
 			DEV_STATS_INC(tunnel->dev, rx_fifo_errors);
 			DEV_STATS_INC(tunnel->dev, rx_errors);
+			reason = SKB_DROP_REASON_TUNNEL_OPT_MISMATCH;
+			goto drop;
+		}
+		if (tunnel->i_seqno &&
+		    (s32)(ntohl(tpi->seq) - tunnel->i_seqno) < 0) {
+			DEV_STATS_INC(tunnel->dev, rx_fifo_errors);
+			DEV_STATS_INC(tunnel->dev, rx_errors);
+			reason = SKB_DROP_REASON_TUNNEL_OLD_SEQ;
 			goto drop;
 		}
 		tunnel->i_seqno = ntohl(tpi->seq) + 1;
@@ -419,7 +428,8 @@ int ip_tunnel_rcv(struct ip_tunnel *tunnel, struct sk_buff *skb,
 
 	skb_set_network_header(skb, (tunnel->dev->type == ARPHRD_ETHER) ? ETH_HLEN : 0);
 
-	if (!pskb_inet_may_pull(skb)) {
+	reason = pskb_inet_may_pull_reason(skb);
+	if (reason) {
 		DEV_STATS_INC(tunnel->dev, rx_length_errors);
 		DEV_STATS_INC(tunnel->dev, rx_errors);
 		goto drop;
@@ -434,6 +444,7 @@ int ip_tunnel_rcv(struct ip_tunnel *tunnel, struct sk_buff *skb,
 		if (err > 1) {
 			DEV_STATS_INC(tunnel->dev, rx_frame_errors);
 			DEV_STATS_INC(tunnel->dev, rx_errors);
+			reason = SKB_DROP_REASON_IP_TUNNEL_ECN;
 			goto drop;
 		}
 	}
@@ -455,9 +466,10 @@ int ip_tunnel_rcv(struct ip_tunnel *tunnel, struct sk_buff *skb,
 	return 0;
 
 drop:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	if (tun_dst)
 		dst_release((struct dst_entry *)tun_dst);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return 0;
 }
 EXPORT_SYMBOL_GPL(ip_tunnel_rcv);
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 05/14] ip6_tunnel: add drop reasons to the receive path
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (3 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 04/14] ip_tunnel: add drop reasons to the generic RX path Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 06/14] gre: make gre_parse_header() report a drop reason Anton Danilov
                   ` (8 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

__ip6_tnl_rcv() mirrors its IPv4 counterpart: all of its failures share
a single plain kfree_skb(). Reuse the drop reasons that ip_tunnel_rcv()
now reports and the ones the length helpers already return.

Note that skb_vlan_inet_prepare() returns an enum skb_drop_reason that
has so far been discarded, and that the ETH_HLEN check gets one by
calling pskb_may_pull_reason() instead of pskb_may_pull().

__ip6_tnl_rcv() is reached two ways: ip6_gre (ip6gre, ip6gretap,
ip6erspan) goes through the exported ip6_tnl_rcv(), while the
ip6_tunnel encapsulations (ip4ip6, ip6ip6, mplsip6) reach it from
ipxip6_rcv(). Only ip6_gre sets the checksum and sequence number bits:
tpi_v4, tpi_v6 and tpi_mpls carry nothing but .proto, and unlike ipip
and sit, ip6_tunnel stores IFLA_IPTUN_FLAGS in parms.flags, not in
parms.i_flags. So the option mismatch and the old sequence reasons are
reachable through ip6_gre alone.

ipxip6_rcv() drops the packets it does not hand to __ip6_tnl_rcv() with
a plain kfree_skb() as well. Give them a reason too:
SKB_DROP_REASON_UNHANDLED_PROTO when the payload is not the one the
tunnel mode carries, the reason the transmit side uses later in the
series, SKB_DROP_REASON_XFRM_POLICY when the xfrm policy check fails,
SKB_DROP_REASON_DEV_READY when ip6_tnl_rcv_ctl() refuses the packet,
the reason iptunnel_pull_header() returns and SKB_DROP_REASON_NOMEM
when the metadata dst cannot be allocated.
ip6_tnl_rcv_ctl() refuses when the outer destination is not a usable
local address, such as a tentative one, when the outer source belongs
to this host, and when the tunnel cannot receive with the addresses of
the packet. It returns only 0 or 1, so DEV_READY stands for all three,
as it does for ip6_tnl_xmit_ctl() on the transmit side later in the
series, although it describes only the first. Telling them apart needs
both helpers to return a reason, which is left for a follow-up.

ipxip6_rcv() and __ip6_tnl_rcv() follow the rule of ip_tunnel_rcv() for
the reason: it is initialized to SKB_DROP_REASON_NOT_SPECIFIED, and the
drop label falls back to it, as a helper stores its result in the
reason.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 net/ipv6/ip6_tunnel.c | 47 ++++++++++++++++++++++++++++++++-----------
 1 file changed, 35 insertions(+), 12 deletions(-)

diff --git a/net/ipv6/ip6_tunnel.c b/net/ipv6/ip6_tunnel.c
index d5ff50a2ac01..52f6a38657b9 100644
--- a/net/ipv6/ip6_tunnel.c
+++ b/net/ipv6/ip6_tunnel.c
@@ -813,6 +813,7 @@ static int __ip6_tnl_rcv(struct ip6_tnl *tunnel, struct sk_buff *skb,
 						struct sk_buff *skb),
 			 bool log_ecn_err)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	const struct ipv6hdr *ipv6h;
 	int nh, err;
 
@@ -820,15 +821,22 @@ static int __ip6_tnl_rcv(struct ip6_tnl *tunnel, struct sk_buff *skb,
 	    test_bit(IP_TUNNEL_CSUM_BIT, tpi->flags)) {
 		DEV_STATS_INC(tunnel->dev, rx_crc_errors);
 		DEV_STATS_INC(tunnel->dev, rx_errors);
+		reason = SKB_DROP_REASON_TUNNEL_OPT_MISMATCH;
 		goto drop;
 	}
 
 	if (test_bit(IP_TUNNEL_SEQ_BIT, tunnel->parms.i_flags)) {
-		if (!test_bit(IP_TUNNEL_SEQ_BIT, tpi->flags) ||
-		    (tunnel->i_seqno &&
-		     (s32)(ntohl(tpi->seq) - tunnel->i_seqno) < 0)) {
+		if (!test_bit(IP_TUNNEL_SEQ_BIT, tpi->flags)) {
 			DEV_STATS_INC(tunnel->dev, rx_fifo_errors);
 			DEV_STATS_INC(tunnel->dev, rx_errors);
+			reason = SKB_DROP_REASON_TUNNEL_OPT_MISMATCH;
+			goto drop;
+		}
+		if (tunnel->i_seqno &&
+		    (s32)(ntohl(tpi->seq) - tunnel->i_seqno) < 0) {
+			DEV_STATS_INC(tunnel->dev, rx_fifo_errors);
+			DEV_STATS_INC(tunnel->dev, rx_errors);
+			reason = SKB_DROP_REASON_TUNNEL_OLD_SEQ;
 			goto drop;
 		}
 		tunnel->i_seqno = ntohl(tpi->seq) + 1;
@@ -838,7 +846,8 @@ static int __ip6_tnl_rcv(struct ip6_tnl *tunnel, struct sk_buff *skb,
 
 	/* Warning: All skb pointers will be invalidated! */
 	if (tunnel->dev->type == ARPHRD_ETHER) {
-		if (!pskb_may_pull(skb, ETH_HLEN)) {
+		reason = pskb_may_pull_reason(skb, ETH_HLEN);
+		if (reason) {
 			DEV_STATS_INC(tunnel->dev, rx_length_errors);
 			DEV_STATS_INC(tunnel->dev, rx_errors);
 			goto drop;
@@ -859,7 +868,8 @@ static int __ip6_tnl_rcv(struct ip6_tnl *tunnel, struct sk_buff *skb,
 
 	skb_reset_network_header(skb);
 
-	if (skb_vlan_inet_prepare(skb, true)) {
+	reason = skb_vlan_inet_prepare(skb, true);
+	if (reason) {
 		DEV_STATS_INC(tunnel->dev, rx_length_errors);
 		DEV_STATS_INC(tunnel->dev, rx_errors);
 		goto drop;
@@ -881,6 +891,7 @@ static int __ip6_tnl_rcv(struct ip6_tnl *tunnel, struct sk_buff *skb,
 		if (err > 1) {
 			DEV_STATS_INC(tunnel->dev, rx_frame_errors);
 			DEV_STATS_INC(tunnel->dev, rx_errors);
+			reason = SKB_DROP_REASON_IP_TUNNEL_ECN;
 			goto drop;
 		}
 	}
@@ -896,9 +907,10 @@ static int __ip6_tnl_rcv(struct ip6_tnl *tunnel, struct sk_buff *skb,
 	return 0;
 
 drop:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	if (tun_dst)
 		dst_release((struct dst_entry *)tun_dst);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return 0;
 }
 
@@ -941,6 +953,7 @@ static int ipxip6_rcv(struct sk_buff *skb, u8 ipproto,
 						  const struct ipv6hdr *ipv6h,
 						  struct sk_buff *skb))
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct ip6_tnl *t;
 	const struct ipv6hdr *ipv6h = ipv6_hdr(skb);
 	struct metadata_dst *tun_dst = NULL;
@@ -952,21 +965,30 @@ static int ipxip6_rcv(struct sk_buff *skb, u8 ipproto,
 	if (t) {
 		u8 tproto = READ_ONCE(t->parms.proto);
 
-		if (tproto != ipproto && tproto != 0)
+		if (tproto != ipproto && tproto != 0) {
+			reason = SKB_DROP_REASON_UNHANDLED_PROTO;
 			goto drop;
-		if (!xfrm6_policy_check(NULL, XFRM_POLICY_IN, skb))
+		}
+		if (!xfrm6_policy_check(NULL, XFRM_POLICY_IN, skb)) {
+			reason = SKB_DROP_REASON_XFRM_POLICY;
 			goto drop;
+		}
 		ipv6h = ipv6_hdr(skb);
-		if (!ip6_tnl_rcv_ctl(t, &ipv6h->daddr, &ipv6h->saddr))
+		if (!ip6_tnl_rcv_ctl(t, &ipv6h->daddr, &ipv6h->saddr)) {
+			reason = SKB_DROP_REASON_DEV_READY;
 			goto drop;
-		if (iptunnel_pull_header(skb, 0, tpi->proto, false))
+		}
+		reason = iptunnel_pull_header(skb, 0, tpi->proto, false);
+		if (reason)
 			goto drop;
 		if (t->parms.collect_md) {
 			IP_TUNNEL_DECLARE_FLAGS(flags) = { };
 
 			tun_dst = ipv6_tun_rx_dst(skb, flags, 0, 0);
-			if (!tun_dst)
+			if (!tun_dst) {
+				reason = SKB_DROP_REASON_NOMEM;
 				goto drop;
+			}
 		}
 		ret = __ip6_tnl_rcv(t, skb, tpi, tun_dst, dscp_ecn_decapsulate,
 				    log_ecn_error);
@@ -978,7 +1000,8 @@ static int ipxip6_rcv(struct sk_buff *skb, u8 ipproto,
 
 drop:
 	rcu_read_unlock();
-	kfree_skb(skb);
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
+	kfree_skb_reason(skb, reason);
 	return 0;
 }
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 06/14] gre: make gre_parse_header() report a drop reason
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (4 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 05/14] ip6_tunnel: add drop reasons to the receive path Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 07/14] ip_gre: add drop reasons to the RX path Anton Danilov
                   ` (7 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

gre_parse_header() returns -EINVAL for every failure and its receive
callers turn that into a plain kfree_skb(). The only detail they could
get so far was the csum_err flag, which none of them actually reads:
both ip_gre and ip6_gre declare it, pass it in and then ignore it.

Make gre_parse_header() return the drop reason instead, with
SKB_NOT_DROPPED_YET for a valid header, and let the two receive paths
report it. The header length it used to return is already stored in
tpi->hdr_len, so the callers take it from there. The reasons are:

 - SKB_DROP_REASON_TUNNEL_INVALID_HDR, for a header carrying an
   unsupported version or the routing bit; its description gets a GRE
   example,

 - SKB_DROP_REASON_GRE_CSUM, a new reason for a checksum error, like the
   existing TCP_CSUM, UDP_CSUM, ICMP_CSUM and IP_CSUM,

 - the reason pskb_may_pull_reason() returns when the header cannot be
   pulled: SKB_DROP_REASON_PKT_TOO_SMALL for a short packet and
   SKB_DROP_REASON_NOMEM for an allocation failure. The byte that tells
   WCCPv1 from WCCPv2 is read with skb_header_pointer() instead, which
   only fails on a short packet, so that check returns
   SKB_DROP_REASON_PKT_TOO_SMALL.

gre_rcv() in the demux reports SKB_DROP_REASON_TUNNEL_INVALID_HDR for a
version it cannot dispatch at all, as gre_parse_header() does for a
nonzero version in ip6_gre, which does not go through the demux. A
valid version whose handler is not registered, such as PPTP, is
reported as SKB_DROP_REASON_UNHANDLED_PROTO.

csum_err had a second use: the ICMP error handlers pass NULL for it, so
that a checksum failure does not reject the header and the rest of it is
still parsed, as they only get a part of the original packet. That came
with commit b0350d51f001 ("ip_gre: fix parsing gre header in ipgre_err")
and becomes an explicit ignore_csum_err argument. The checksum is
still computed either way. As the function is exported, the comment
above it becomes a kernel-doc comment, which describes the argument.

Tunnel lookup failures and the other drops of the receive handlers
still report SKB_DROP_REASON_NOT_SPECIFIED at this point; the following
patches convert them.

The reason follows the rule of ip_tunnel_rcv(). gre_rcv() of the demux
and of ip6_gre sets it in its first statement, so it is not initialized
there; ip_gre's gre_rcv() initializes it, which its looped back
multicast check relies on, as it drops before the header is parsed. All
three store the result of a helper in the reason, so their drop labels
fall back to SKB_DROP_REASON_NOT_SPECIFIED.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 include/net/dropreason-core.h |  4 +++
 include/net/gre.h             |  5 +--
 net/ipv4/gre_demux.c          | 65 ++++++++++++++++++++++++-----------
 net/ipv4/ip_gre.c             | 18 +++++-----
 net/ipv6/ip6_gre.c            | 18 +++++-----
 5 files changed, 70 insertions(+), 40 deletions(-)

diff --git a/include/net/dropreason-core.h b/include/net/dropreason-core.h
index 85aaa58c3ecb..0d963f18a21e 100644
--- a/include/net/dropreason-core.h
+++ b/include/net/dropreason-core.h
@@ -131,6 +131,7 @@
 	FN(RECURSION_LIMIT)		\
 	FN(TUNNEL_OPT_MISMATCH)		\
 	FN(TUNNEL_OLD_SEQ)		\
+	FN(GRE_CSUM)			\
 	FNe(MAX)
 
 /**
@@ -544,6 +545,7 @@ enum skb_drop_reason {
 	 * @SKB_DROP_REASON_TUNNEL_INVALID_HDR: tunnel header is invalid. E.g.:
 	 * 1) VXLAN reserved fields are not zero
 	 * 2) VXLAN "I" flag is not set
+	 * 3) GRE version is not supported or the routing bit is set
 	 */
 	SKB_DROP_REASON_TUNNEL_INVALID_HDR,
 	/**
@@ -631,6 +633,8 @@ enum skb_drop_reason {
 	 * remote endpoint restarted and reset its sequence numbering.
 	 */
 	SKB_DROP_REASON_TUNNEL_OLD_SEQ,
+	/** @SKB_DROP_REASON_GRE_CSUM: GRE checksum error */
+	SKB_DROP_REASON_GRE_CSUM,
 	/**
 	 * @SKB_DROP_REASON_MAX: the maximum of core drop reasons, which
 	 * shouldn't be used as a real 'reason' - only for tracing code gen
diff --git a/include/net/gre.h b/include/net/gre.h
index b55f67ecd2fc..b5dbe00ab15e 100644
--- a/include/net/gre.h
+++ b/include/net/gre.h
@@ -32,8 +32,9 @@ struct gre_protocol {
 int gre_add_protocol(const struct gre_protocol *proto, u8 version);
 int gre_del_protocol(const struct gre_protocol *proto, u8 version);
 
-int gre_parse_header(struct sk_buff *skb, struct tnl_ptk_info *tpi,
-		     bool *csum_err, __be16 proto, int nhs);
+enum skb_drop_reason
+gre_parse_header(struct sk_buff *skb, struct tnl_ptk_info *tpi,
+		 bool ignore_csum_err, __be16 proto, int nhs);
 
 static inline bool netif_is_gretap(const struct net_device *dev)
 {
diff --git a/net/ipv4/gre_demux.c b/net/ipv4/gre_demux.c
index 96fd7dc6d82d..fccdaa240f46 100644
--- a/net/ipv4/gre_demux.c
+++ b/net/ipv4/gre_demux.c
@@ -56,28 +56,47 @@ int gre_del_protocol(const struct gre_protocol *proto, u8 version)
 }
 EXPORT_SYMBOL_GPL(gre_del_protocol);
 
-/* Fills in tpi and returns header length to be pulled.
- * Note that caller must use pskb_may_pull() before pulling GRE header.
+/**
+ * gre_parse_header() - parse the GRE header of a packet
+ * @skb: the packet; the GRE header starts @nhs bytes past skb->data
+ * @tpi: filled in with the fields of the header, and with the length of
+ *	the header to be pulled in tpi->hdr_len
+ * @ignore_csum_err: do not reject the header when its checksum does not
+ *	match. The ICMP error handlers set it, as they only get a part of the
+ *	original packet; the checksum is still computed and the rest of the
+ *	header is parsed.
+ * @proto: the protocol to report for a WCCP payload, ETH_P_IP or
+ *	ETH_P_IPV6
+ * @nhs: the offset of the GRE header from skb->data
+ *
+ * The caller must use pskb_may_pull() before pulling the GRE header.
+ *
+ * Return: SKB_NOT_DROPPED_YET, or the reason to drop the packet if the
+ * header is rejected.
  */
-int gre_parse_header(struct sk_buff *skb, struct tnl_ptk_info *tpi,
-		     bool *csum_err, __be16 proto, int nhs)
+enum skb_drop_reason
+gre_parse_header(struct sk_buff *skb, struct tnl_ptk_info *tpi,
+		 bool ignore_csum_err, __be16 proto, int nhs)
 {
 	const struct gre_base_hdr *greh;
+	enum skb_drop_reason reason;
 	__be32 *options;
 	int hdr_len;
 
-	if (unlikely(!pskb_may_pull(skb, nhs + sizeof(struct gre_base_hdr))))
-		return -EINVAL;
+	reason = pskb_may_pull_reason(skb, nhs + sizeof(struct gre_base_hdr));
+	if (unlikely(reason))
+		return reason;
 
 	greh = (struct gre_base_hdr *)(skb->data + nhs);
 	if (unlikely(greh->flags & (GRE_VERSION | GRE_ROUTING)))
-		return -EINVAL;
+		return SKB_DROP_REASON_TUNNEL_INVALID_HDR;
 
 	gre_flags_to_tnl_flags(tpi->flags, greh->flags);
 	hdr_len = gre_calc_hlen(tpi->flags);
 
-	if (!pskb_may_pull(skb, nhs + hdr_len))
-		return -EINVAL;
+	reason = pskb_may_pull_reason(skb, nhs + hdr_len);
+	if (reason)
+		return reason;
 
 	greh = (struct gre_base_hdr *)(skb->data + nhs);
 	tpi->proto = greh->protocol;
@@ -87,9 +106,8 @@ int gre_parse_header(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 		if (!skb_checksum_simple_validate(skb)) {
 			skb_checksum_try_convert(skb, IPPROTO_GRE,
 						 null_compute_pseudo);
-		} else if (csum_err) {
-			*csum_err = true;
-			return -EINVAL;
+		} else if (!ignore_csum_err) {
+			return SKB_DROP_REASON_GRE_CSUM;
 		}
 
 		options++;
@@ -117,7 +135,7 @@ int gre_parse_header(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 		val = skb_header_pointer(skb, nhs + hdr_len,
 					 sizeof(_val), &_val);
 		if (!val)
-			return -EINVAL;
+			return SKB_DROP_REASON_PKT_TOO_SMALL;
 		tpi->proto = proto;
 		if ((*val & 0xF0) != 0x40)
 			hdr_len += 4;
@@ -132,29 +150,35 @@ int gre_parse_header(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 	    greh->protocol == htons(ETH_P_ERSPAN2)) {
 		struct erspan_base_hdr *ershdr;
 
-		if (!pskb_may_pull(skb, nhs + hdr_len + sizeof(*ershdr)))
-			return -EINVAL;
+		reason = pskb_may_pull_reason(skb,
+					      nhs + hdr_len + sizeof(*ershdr));
+		if (reason)
+			return reason;
 
 		ershdr = (struct erspan_base_hdr *)(skb->data + nhs + hdr_len);
 		tpi->key = cpu_to_be32(get_session_id(ershdr));
 	}
 
-	return hdr_len;
+	return SKB_NOT_DROPPED_YET;
 }
 EXPORT_SYMBOL(gre_parse_header);
 
 static int gre_rcv(struct sk_buff *skb)
 {
 	const struct gre_protocol *proto;
+	enum skb_drop_reason reason;
 	u8 ver;
 	int ret;
 
-	if (!pskb_may_pull(skb, 12))
+	reason = pskb_may_pull_reason(skb, 12);
+	if (reason)
 		goto drop;
 
 	ver = skb->data[1]&0x7f;
-	if (ver >= GREPROTO_MAX)
+	if (ver >= GREPROTO_MAX) {
+		reason = SKB_DROP_REASON_TUNNEL_INVALID_HDR;
 		goto drop;
+	}
 
 	rcu_read_lock();
 	proto = rcu_dereference(gre_proto[ver]);
@@ -167,11 +191,12 @@ static int gre_rcv(struct sk_buff *skb)
 drop_nohandler:
 	rcu_read_unlock();
 	dev_core_stats_rx_nohandler_inc(skb->dev);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, SKB_DROP_REASON_UNHANDLED_PROTO);
 	return NET_RX_DROP;
 drop:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	dev_core_stats_rx_dropped_inc(skb->dev);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return NET_RX_DROP;
 }
 
diff --git a/net/ipv4/ip_gre.c b/net/ipv4/ip_gre.c
index 3ee437fa3eae..a1b44007b1d7 100644
--- a/net/ipv4/ip_gre.c
+++ b/net/ipv4/ip_gre.c
@@ -237,8 +237,8 @@ static void gre_err(struct sk_buff *skb, u32 info)
 	const int code = icmp_hdr(skb)->code;
 	struct tnl_ptk_info tpi;
 
-	if (gre_parse_header(skb, &tpi, NULL, htons(ETH_P_IP),
-			     iph->ihl * 4) < 0)
+	if (gre_parse_header(skb, &tpi, true, htons(ETH_P_IP),
+			     iph->ihl * 4))
 		return;
 
 	if (type == ICMP_DEST_UNREACH && code == ICMP_FRAG_NEEDED) {
@@ -439,9 +439,8 @@ static int ipgre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi,
 
 static int gre_rcv(struct sk_buff *skb)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct tnl_ptk_info tpi;
-	bool csum_err = false;
-	int hdr_len;
 
 #ifdef CONFIG_NET_IPGRE_BROADCAST
 	if (ipv4_is_multicast(ip_hdr(skb)->daddr)) {
@@ -451,25 +450,26 @@ static int gre_rcv(struct sk_buff *skb)
 	}
 #endif
 
-	hdr_len = gre_parse_header(skb, &tpi, &csum_err, htons(ETH_P_IP), 0);
-	if (hdr_len < 0)
+	reason = gre_parse_header(skb, &tpi, false, htons(ETH_P_IP), 0);
+	if (reason)
 		goto drop;
 
 	if (unlikely(tpi.proto == htons(ETH_P_ERSPAN) ||
 		     tpi.proto == htons(ETH_P_ERSPAN2))) {
-		if (erspan_rcv(skb, &tpi, hdr_len) == PACKET_RCVD)
+		if (erspan_rcv(skb, &tpi, tpi.hdr_len) == PACKET_RCVD)
 			return 0;
 		goto out;
 	}
 
-	if (ipgre_rcv(skb, &tpi, hdr_len) == PACKET_RCVD)
+	if (ipgre_rcv(skb, &tpi, tpi.hdr_len) == PACKET_RCVD)
 		return 0;
 
 out:
 	icmp_send(skb, ICMP_DEST_UNREACH, ICMP_PORT_UNREACH, 0);
 drop:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	dev_core_stats_rx_dropped_inc(skb->dev);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return 0;
 }
 
diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c
index 0392f6ba862b..25fd290082c4 100644
--- a/net/ipv6/ip6_gre.c
+++ b/net/ipv6/ip6_gre.c
@@ -389,8 +389,8 @@ static int ip6gre_err(struct sk_buff *skb, struct inet6_skb_parm *opt,
 	struct tnl_ptk_info tpi;
 	struct ip6_tnl *t;
 
-	if (gre_parse_header(skb, &tpi, NULL, htons(ETH_P_IPV6),
-			     offset) < 0)
+	if (gre_parse_header(skb, &tpi, true, htons(ETH_P_IPV6),
+			     offset))
 		return -EINVAL;
 
 	ipv6h = (const struct ipv6hdr *)skb->data;
@@ -566,20 +566,19 @@ static int ip6erspan_rcv(struct sk_buff *skb,
 
 static int gre_rcv(struct sk_buff *skb)
 {
+	enum skb_drop_reason reason;
 	struct tnl_ptk_info tpi;
-	bool csum_err = false;
-	int hdr_len;
 
-	hdr_len = gre_parse_header(skb, &tpi, &csum_err, htons(ETH_P_IPV6), 0);
-	if (hdr_len < 0)
+	reason = gre_parse_header(skb, &tpi, false, htons(ETH_P_IPV6), 0);
+	if (reason)
 		goto drop;
 
-	if (iptunnel_pull_header(skb, hdr_len, tpi.proto, false))
+	if (iptunnel_pull_header(skb, tpi.hdr_len, tpi.proto, false))
 		goto drop;
 
 	if (unlikely(tpi.proto == htons(ETH_P_ERSPAN) ||
 		     tpi.proto == htons(ETH_P_ERSPAN2))) {
-		if (ip6erspan_rcv(skb, &tpi, hdr_len) == PACKET_RCVD)
+		if (ip6erspan_rcv(skb, &tpi, tpi.hdr_len) == PACKET_RCVD)
 			return 0;
 		goto out;
 	}
@@ -590,8 +589,9 @@ static int gre_rcv(struct sk_buff *skb)
 out:
 	icmpv6_send(skb, ICMPV6_DEST_UNREACH, ICMPV6_PORT_UNREACH, 0);
 drop:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	dev_core_stats_rx_dropped_inc(skb->dev);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return 0;
 }
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 07/14] ip_gre: add drop reasons to the RX path
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (5 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 06/14] gre: make gre_parse_header() report a drop reason Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 08/14] ip6_gre: " Anton Danilov
                   ` (6 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

A packet that reaches gre_rcv() and does not belong to any tunnel is
dropped, after an ICMP port unreachable is sent back, as
SKB_DROP_REASON_NOT_SPECIFIED from the same kfree_skb_reason() call as
a truncated ERSPAN header or a failed metadata allocation. This is the
GRE counterpart of a UDP packet hitting no socket, yet neither the
reason nor the call site tells it apart. The fallback devices such as
gre0 only take such a packet while they are up, and they are created
down.

Report SKB_DROP_REASON_TUNNEL_NOT_FOUND for it, which vxlan reports when
no device matches the VNI, and add GRE to its description.

erspan_rcv(), __ipgre_rcv() and ipgre_rcv() now return
SKB_NOT_DROPPED_YET where they returned PACKET_RCVD, including the paths
that free the packet themselves, and a drop reason where they returned
PACKET_REJECT or PACKET_NEXT. __ipgre_rcv() only returned PACKET_NEXT
when no tunnel matched, so the retry ipgre_rcv() does for ETH_P_TEB is
now taken on SKB_DROP_REASON_TUNNEL_NOT_FOUND.

The length checks use pskb_may_pull_reason(), and the headers are
pulled with __iptunnel_pull_header(), which returns a drop reason since
an earlier patch of this series. Both report a packet too short to pull
as SKB_DROP_REASON_PKT_TOO_SMALL and an allocation failure as
SKB_DROP_REASON_NOMEM. The metadata allocation failures reuse
SKB_DROP_REASON_NOMEM as well.

erspan_rcv() and __ipgre_rcv() follow the same rule for the reason as
gre_rcv(): it is initialized to SKB_DROP_REASON_NOT_SPECIFIED, and their
drop labels fall back to it, as the pull helpers store their result in
the reason.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 include/net/dropreason-core.h |  3 +-
 net/ipv4/ip_gre.c             | 81 ++++++++++++++++++++---------------
 2 files changed, 49 insertions(+), 35 deletions(-)

diff --git a/include/net/dropreason-core.h b/include/net/dropreason-core.h
index 0d963f18a21e..6f5273e16548 100644
--- a/include/net/dropreason-core.h
+++ b/include/net/dropreason-core.h
@@ -550,7 +550,8 @@ enum skb_drop_reason {
 	SKB_DROP_REASON_TUNNEL_INVALID_HDR,
 	/**
 	 * @SKB_DROP_REASON_TUNNEL_NOT_FOUND: no tunnel device found for the
-	 * packet, e.g. no VXLAN device for its VNI
+	 * packet, e.g. no VXLAN device for its VNI or no GRE tunnel for its
+	 * endpoints and key
 	 */
 	SKB_DROP_REASON_TUNNEL_NOT_FOUND,
 	/** @SKB_DROP_REASON_MAC_INVALID_SOURCE: source mac is invalid */
diff --git a/net/ipv4/ip_gre.c b/net/ipv4/ip_gre.c
index a1b44007b1d7..afd8ece02d7f 100644
--- a/net/ipv4/ip_gre.c
+++ b/net/ipv4/ip_gre.c
@@ -264,9 +264,11 @@ static bool is_erspan_type1(int gre_hdr_len)
 	return gre_hdr_len == 4;
 }
 
-static int erspan_rcv(struct sk_buff *skb, struct tnl_ptk_info *tpi,
-		      int gre_hdr_len)
+static enum skb_drop_reason erspan_rcv(struct sk_buff *skb,
+				       struct tnl_ptk_info *tpi,
+				       int gre_hdr_len)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct net *net = dev_net(skb->dev);
 	struct metadata_dst *tun_dst = NULL;
 	struct erspan_base_hdr *ershdr;
@@ -288,9 +290,10 @@ static int erspan_rcv(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 		tunnel = ip_tunnel_lookup(itn, skb->dev->ifindex, flags,
 					  iph->saddr, iph->daddr, 0);
 	} else {
-		if (unlikely(!pskb_may_pull(skb,
-					    gre_hdr_len + sizeof(*ershdr))))
-			return PACKET_REJECT;
+		reason = pskb_may_pull_reason(skb,
+					      gre_hdr_len + sizeof(*ershdr));
+		if (unlikely(reason))
+			return reason;
 
 		ershdr = (struct erspan_base_hdr *)(skb->data + gre_hdr_len);
 		ver = ershdr->ver;
@@ -306,13 +309,13 @@ static int erspan_rcv(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 		else
 			len = gre_hdr_len + erspan_hdr_len(ver);
 
-		if (unlikely(!pskb_may_pull(skb, len)))
-			return PACKET_REJECT;
+		reason = pskb_may_pull_reason(skb, len);
+		if (unlikely(reason))
+			return reason;
 
-		if (__iptunnel_pull_header(skb,
-					   len,
-					   htons(ETH_P_TEB),
-					   false, false))
+		reason = __iptunnel_pull_header(skb, len, htons(ETH_P_TEB),
+						false, false);
+		if (reason)
 			goto drop;
 
 		if (tunnel->collect_md) {
@@ -328,7 +331,7 @@ static int erspan_rcv(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 			tun_dst = ip_tun_rx_dst(skb, flags,
 						tun_id, sizeof(*md));
 			if (!tun_dst)
-				return PACKET_REJECT;
+				return SKB_DROP_REASON_NOMEM;
 
 			/* MUST set options_len before referencing options */
 			info = &tun_dst->u.tun_info;
@@ -354,18 +357,22 @@ static int erspan_rcv(struct sk_buff *skb, struct tnl_ptk_info *tpi,
 
 		skb_reset_mac_header(skb);
 		ip_tunnel_rcv(tunnel, skb, tpi, tun_dst, log_ecn_error);
-		return PACKET_RCVD;
+		return SKB_NOT_DROPPED_YET;
 	}
-	return PACKET_REJECT;
+	return SKB_DROP_REASON_TUNNEL_NOT_FOUND;
 
 drop:
-	kfree_skb(skb);
-	return PACKET_RCVD;
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
+	kfree_skb_reason(skb, reason);
+	return SKB_NOT_DROPPED_YET;
 }
 
-static int __ipgre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi,
-		       struct ip_tunnel_net *itn, int hdr_len, bool raw_proto)
+static enum skb_drop_reason __ipgre_rcv(struct sk_buff *skb,
+					const struct tnl_ptk_info *tpi,
+					struct ip_tunnel_net *itn,
+					int hdr_len, bool raw_proto)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct metadata_dst *tun_dst = NULL;
 	const struct iphdr *iph;
 	struct ip_tunnel *tunnel;
@@ -377,8 +384,9 @@ static int __ipgre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi,
 	if (tunnel) {
 		const struct iphdr *tnl_params;
 
-		if (__iptunnel_pull_header(skb, hdr_len, tpi->proto,
-					   raw_proto, false))
+		reason = __iptunnel_pull_header(skb, hdr_len, tpi->proto,
+						raw_proto, false);
+		if (reason)
 			goto drop;
 
 		/* Special case for ipgre_header_parse(), which expects the
@@ -401,40 +409,43 @@ static int __ipgre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi,
 			tun_id = key32_to_tunnel_id(tpi->key);
 			tun_dst = ip_tun_rx_dst(skb, flags, tun_id, 0);
 			if (!tun_dst)
-				return PACKET_REJECT;
+				return SKB_DROP_REASON_NOMEM;
 		}
 
 		ip_tunnel_rcv(tunnel, skb, tpi, tun_dst, log_ecn_error);
-		return PACKET_RCVD;
+		return SKB_NOT_DROPPED_YET;
 	}
-	return PACKET_NEXT;
+	return SKB_DROP_REASON_TUNNEL_NOT_FOUND;
 
 drop:
-	kfree_skb(skb);
-	return PACKET_RCVD;
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
+	kfree_skb_reason(skb, reason);
+	return SKB_NOT_DROPPED_YET;
 }
 
-static int ipgre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi,
-		     int hdr_len)
+static enum skb_drop_reason ipgre_rcv(struct sk_buff *skb,
+				      const struct tnl_ptk_info *tpi,
+				      int hdr_len)
 {
 	struct net *net = dev_net(skb->dev);
+	enum skb_drop_reason reason;
 	struct ip_tunnel_net *itn;
-	int res;
 
 	if (tpi->proto == htons(ETH_P_TEB))
 		itn = net_generic(net, gre_tap_net_id);
 	else
 		itn = net_generic(net, ipgre_net_id);
 
-	res = __ipgre_rcv(skb, tpi, itn, hdr_len, false);
-	if (res == PACKET_NEXT && tpi->proto == htons(ETH_P_TEB)) {
+	reason = __ipgre_rcv(skb, tpi, itn, hdr_len, false);
+	if (reason == SKB_DROP_REASON_TUNNEL_NOT_FOUND &&
+	    tpi->proto == htons(ETH_P_TEB)) {
 		/* ipgre tunnels in collect metadata mode should receive
 		 * also ETH_P_TEB traffic.
 		 */
 		itn = net_generic(net, ipgre_net_id);
-		res = __ipgre_rcv(skb, tpi, itn, hdr_len, true);
+		reason = __ipgre_rcv(skb, tpi, itn, hdr_len, true);
 	}
-	return res;
+	return reason;
 }
 
 static int gre_rcv(struct sk_buff *skb)
@@ -456,12 +467,14 @@ static int gre_rcv(struct sk_buff *skb)
 
 	if (unlikely(tpi.proto == htons(ETH_P_ERSPAN) ||
 		     tpi.proto == htons(ETH_P_ERSPAN2))) {
-		if (erspan_rcv(skb, &tpi, tpi.hdr_len) == PACKET_RCVD)
+		reason = erspan_rcv(skb, &tpi, tpi.hdr_len);
+		if (!reason)
 			return 0;
 		goto out;
 	}
 
-	if (ipgre_rcv(skb, &tpi, tpi.hdr_len) == PACKET_RCVD)
+	reason = ipgre_rcv(skb, &tpi, tpi.hdr_len);
+	if (!reason)
 		return 0;
 
 out:
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 08/14] ip6_gre: add drop reasons to the RX path
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (6 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 07/14] ip_gre: add drop reasons to the RX path Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 09/14] ip_tunnel: add drop reasons to the transmit path Anton Danilov
                   ` (5 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

Mirror the previous patch on the IPv6 side. So far gre_rcv() reported
the drops below as SKB_DROP_REASON_NOT_SPECIFIED, all from the same
kfree_skb_reason() call. Now it reports:

 - SKB_DROP_REASON_TUNNEL_NOT_FOUND when no tunnel matches,
 - what pskb_may_pull_reason() returns when an ERSPAN header cannot be
   pulled,
 - what iptunnel_pull_header() and __iptunnel_pull_header() return
   when the pull fails,
 - SKB_DROP_REASON_NOMEM when the metadata dst cannot be allocated.

Unlike ip_gre, gre_rcv() calls the helper before the tunnel lookup, so a
packet whose ETH_P_TEB inner Ethernet header or WCCPv2 extra word is cut
short is reported as SKB_DROP_REASON_PKT_TOO_SMALL whether a tunnel
matches or not.

Like their IPv4 counterparts, ip6gre_rcv() and ip6erspan_rcv() now
return SKB_NOT_DROPPED_YET where they returned PACKET_RCVD and a drop
reason where they returned PACKET_REJECT. That leaves PACKET_RCVD,
PACKET_REJECT and PACKET_NEXT without users, so remove them.
PACKET_REJECT has the value of SKB_CONSUMED, and a stray one in a
function that now returns a drop reason would build silently and be
traced as a consumed packet.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 include/net/ip_tunnels.h |  4 ----
 net/ipv6/ip6_gre.c       | 49 +++++++++++++++++++++++-----------------
 2 files changed, 28 insertions(+), 25 deletions(-)

diff --git a/include/net/ip_tunnels.h b/include/net/ip_tunnels.h
index 5cccf4c0e691..792690ae4f26 100644
--- a/include/net/ip_tunnels.h
+++ b/include/net/ip_tunnels.h
@@ -207,10 +207,6 @@ struct tnl_ptk_info {
 	int hdr_len;
 };
 
-#define PACKET_RCVD	0
-#define PACKET_REJECT	1
-#define PACKET_NEXT	2
-
 #define IP_TNL_HASH_BITS   7
 #define IP_TNL_HASH_SIZE   (1 << IP_TNL_HASH_BITS)
 
diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c
index 25fd290082c4..3333a0ab180c 100644
--- a/net/ipv6/ip6_gre.c
+++ b/net/ipv6/ip6_gre.c
@@ -451,7 +451,8 @@ static int ip6gre_err(struct sk_buff *skb, struct inet6_skb_parm *opt,
 	return 0;
 }
 
-static int ip6gre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi)
+static enum skb_drop_reason ip6gre_rcv(struct sk_buff *skb,
+				       const struct tnl_ptk_info *tpi)
 {
 	const struct ipv6hdr *ipv6h;
 	struct ip6_tnl *tunnel;
@@ -471,31 +472,33 @@ static int ip6gre_rcv(struct sk_buff *skb, const struct tnl_ptk_info *tpi)
 
 			tun_dst = ipv6_tun_rx_dst(skb, flags, tun_id, 0);
 			if (!tun_dst)
-				return PACKET_REJECT;
+				return SKB_DROP_REASON_NOMEM;
 
 			ip6_tnl_rcv(tunnel, skb, tpi, tun_dst, log_ecn_error);
 		} else {
 			ip6_tnl_rcv(tunnel, skb, tpi, NULL, log_ecn_error);
 		}
 
-		return PACKET_RCVD;
+		return SKB_NOT_DROPPED_YET;
 	}
 
-	return PACKET_REJECT;
+	return SKB_DROP_REASON_TUNNEL_NOT_FOUND;
 }
 
-static int ip6erspan_rcv(struct sk_buff *skb,
-			 struct tnl_ptk_info *tpi,
-			 int gre_hdr_len)
+static enum skb_drop_reason ip6erspan_rcv(struct sk_buff *skb,
+					  struct tnl_ptk_info *tpi,
+					  int gre_hdr_len)
 {
 	struct erspan_base_hdr *ershdr;
 	const struct ipv6hdr *ipv6h;
+	enum skb_drop_reason reason;
 	struct erspan_md2 *md2;
 	struct ip6_tnl *tunnel;
 	u8 ver;
 
-	if (unlikely(!pskb_may_pull(skb, sizeof(*ershdr))))
-		return PACKET_REJECT;
+	reason = pskb_may_pull_reason(skb, sizeof(*ershdr));
+	if (unlikely(reason))
+		return reason;
 
 	ipv6h = ipv6_hdr(skb);
 	ershdr = (struct erspan_base_hdr *)skb->data;
@@ -507,13 +510,14 @@ static int ip6erspan_rcv(struct sk_buff *skb,
 	if (tunnel) {
 		int len = erspan_hdr_len(ver);
 
-		if (unlikely(!pskb_may_pull(skb, len)))
-			return PACKET_REJECT;
+		reason = pskb_may_pull_reason(skb, len);
+		if (unlikely(reason))
+			return reason;
 
-		if (__iptunnel_pull_header(skb, len,
-					   htons(ETH_P_TEB),
-					   false, false))
-			return PACKET_REJECT;
+		reason = __iptunnel_pull_header(skb, len, htons(ETH_P_TEB),
+						false, false);
+		if (reason)
+			return reason;
 
 		if (tunnel->parms.collect_md) {
 			struct erspan_metadata *pkt_md, *md;
@@ -530,7 +534,7 @@ static int ip6erspan_rcv(struct sk_buff *skb,
 			tun_dst = ipv6_tun_rx_dst(skb, flags, tun_id,
 						  sizeof(*md));
 			if (!tun_dst)
-				return PACKET_REJECT;
+				return SKB_DROP_REASON_NOMEM;
 
 			/* MUST set options_len before referencing options */
 			info = &tun_dst->u.tun_info;
@@ -558,10 +562,10 @@ static int ip6erspan_rcv(struct sk_buff *skb,
 			ip6_tnl_rcv(tunnel, skb, tpi, NULL, log_ecn_error);
 		}
 
-		return PACKET_RCVD;
+		return SKB_NOT_DROPPED_YET;
 	}
 
-	return PACKET_REJECT;
+	return SKB_DROP_REASON_TUNNEL_NOT_FOUND;
 }
 
 static int gre_rcv(struct sk_buff *skb)
@@ -573,17 +577,20 @@ static int gre_rcv(struct sk_buff *skb)
 	if (reason)
 		goto drop;
 
-	if (iptunnel_pull_header(skb, tpi.hdr_len, tpi.proto, false))
+	reason = iptunnel_pull_header(skb, tpi.hdr_len, tpi.proto, false);
+	if (reason)
 		goto drop;
 
 	if (unlikely(tpi.proto == htons(ETH_P_ERSPAN) ||
 		     tpi.proto == htons(ETH_P_ERSPAN2))) {
-		if (ip6erspan_rcv(skb, &tpi, tpi.hdr_len) == PACKET_RCVD)
+		reason = ip6erspan_rcv(skb, &tpi, tpi.hdr_len);
+		if (!reason)
 			return 0;
 		goto out;
 	}
 
-	if (ip6gre_rcv(skb, &tpi) == PACKET_RCVD)
+	reason = ip6gre_rcv(skb, &tpi);
+	if (!reason)
 		return 0;
 
 out:
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 09/14] ip_tunnel: add drop reasons to the transmit path
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (7 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 08/14] ip6_gre: " Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 10/14] ip_gre: " Anton Danilov
                   ` (4 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

ip_tunnel_xmit() and ip_md_tunnel_xmit() encapsulate the packets sent
through the tunnel, and every failure these two functions detect ends in
a plain kfree_skb(). The device counters separate them a little, but
they are too coarse to act on: tx_errors counts an encapsulation
failure, a routing failure, a routing loop and a packet that is simply
too big alike.

The packet that is too big deserves attention. tnl_update_pmtu()
returns -E2BIG for a non-GSO packet larger than the MTU if it is IPv4
with the DF bit set, or IPv6 and the MTU is at least IPV6_MIN_MTU, after
it has already sent the ICMP error back to the sender. That is path MTU
discovery working as intended, yet among the device counters the drop
only bumps tx_errors, like a failed encapsulation and a few other
failures do. If the ICMP error never reaches the sender, the resulting
MTU black hole cannot be told from those by the device counters. A
failed route lookup and a routing loop do have counters of their own,
tx_carrier_errors and collisions, but in both functions all of these
packets end up in the same kfree_skb() call.

No new reason is needed for most of it:

 - SKB_DROP_REASON_PKT_TOO_BIG for the case above,
 - SKB_DROP_REASON_IP_OUTNOROUTES when the route lookup fails,
 - SKB_DROP_REASON_RECURSION_LIMIT when the route points back at the
   tunnel device itself, which is the "dead loop on virtual device" that
   reason describes; its description now gives this case as an example,
 - SKB_DROP_REASON_NOMEM when the headroom cannot be expanded,
 - SKB_DROP_REASON_NEIGH_CREATEFAIL when the NBMA neighbour lookup
   fails, and SKB_DROP_REASON_NO_TX_TARGET when no destination can be
   derived at all, which includes a payload that is neither IPv4 nor
   IPv6,
 - SKB_DROP_REASON_TUNNEL_TXINFO, which already documents a packet
   reaching an external mode device without metadata, for the
   collect_md path.

Only the encapsulation failure has no fitting reason, so add
SKB_DROP_REASON_TUNNEL_ENCAP for it.

Drop reasons on transmit are not new: vxlan already reports several of
them from its xmit path, and ip_tunnel_core.c reports
SKB_DROP_REASON_RECURSION_LIMIT. They are most useful for forwarded
packets, which is what a tunnel gateway mostly transmits: the sender is
another host, which gets an ICMP error for only some of these failures,
so the drop has to be explained on the gateway.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 include/net/dropreason-core.h | 13 ++++++++++-
 net/ipv4/ip_tunnel.c          | 44 +++++++++++++++++++++++++----------
 2 files changed, 44 insertions(+), 13 deletions(-)

diff --git a/include/net/dropreason-core.h b/include/net/dropreason-core.h
index 6f5273e16548..40f23548f668 100644
--- a/include/net/dropreason-core.h
+++ b/include/net/dropreason-core.h
@@ -132,6 +132,7 @@
 	FN(TUNNEL_OPT_MISMATCH)		\
 	FN(TUNNEL_OLD_SEQ)		\
 	FN(GRE_CSUM)			\
+	FN(TUNNEL_ENCAP)		\
 	FNe(MAX)
 
 /**
@@ -618,7 +619,11 @@ enum skb_drop_reason {
 	SKB_DROP_REASON_PSP_INPUT,
 	/** @SKB_DROP_REASON_PSP_OUTPUT: PSP output checks failed */
 	SKB_DROP_REASON_PSP_OUTPUT,
-	/** @SKB_DROP_REASON_RECURSION_LIMIT: Dead loop on virtual device. */
+	/**
+	 * @SKB_DROP_REASON_RECURSION_LIMIT: Dead loop on virtual device, e.g. a
+	 * tunnel whose route to its remote end goes out of the tunnel device
+	 * itself.
+	 */
 	SKB_DROP_REASON_RECURSION_LIMIT,
 	/**
 	 * @SKB_DROP_REASON_TUNNEL_OPT_MISMATCH: the tunnel options carried by
@@ -636,6 +641,12 @@ enum skb_drop_reason {
 	SKB_DROP_REASON_TUNNEL_OLD_SEQ,
 	/** @SKB_DROP_REASON_GRE_CSUM: GRE checksum error */
 	SKB_DROP_REASON_GRE_CSUM,
+	/**
+	 * @SKB_DROP_REASON_TUNNEL_ENCAP: failed to build the encapsulation
+	 * header of a tunnel, e.g. an unknown or unregistered encapsulation
+	 * type.
+	 */
+	SKB_DROP_REASON_TUNNEL_ENCAP,
 	/**
 	 * @SKB_DROP_REASON_MAX: the maximum of core drop reasons, which
 	 * shouldn't be used as a real 'reason' - only for tracing code gen
diff --git a/net/ipv4/ip_tunnel.c b/net/ipv4/ip_tunnel.c
index 6d500751f837..98ebfadbf64b 100644
--- a/net/ipv4/ip_tunnel.c
+++ b/net/ipv4/ip_tunnel.c
@@ -587,6 +587,7 @@ static int tnl_update_pmtu(struct net_device *dev, struct sk_buff *skb,
 void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 		       u8 proto, int tunnel_hlen)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct ip_tunnel *tunnel = netdev_priv(dev);
 	u32 headroom = sizeof(struct iphdr);
 	struct ip_tunnel_info *tun_info;
@@ -600,8 +601,10 @@ void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 
 	tun_info = skb_tunnel_info(skb);
 	if (unlikely(!tun_info || !(tun_info->mode & IP_TUNNEL_INFO_TX) ||
-		     ip_tunnel_info_af(tun_info) != AF_INET))
+		     ip_tunnel_info_af(tun_info) != AF_INET)) {
+		reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 		goto tx_error;
+	}
 	key = &tun_info->key;
 	memset(&(IPCB(skb)->opt), 0, sizeof(IPCB(skb)->opt));
 	inner_iph = (const struct iphdr *)skb_inner_network_header(skb);
@@ -620,8 +623,10 @@ void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 	if (!tunnel_hlen)
 		tunnel_hlen = ip_encap_hlen(&tun_info->encap);
 
-	if (ip_tunnel_encap(skb, &tun_info->encap, &proto, &fl4) < 0)
+	if (ip_tunnel_encap(skb, &tun_info->encap, &proto, &fl4) < 0) {
+		reason = SKB_DROP_REASON_TUNNEL_ENCAP;
 		goto tx_error;
+	}
 
 	use_cache = ip_tunnel_dst_cache_usable(skb, tun_info);
 	if (use_cache)
@@ -630,6 +635,7 @@ void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 		rt = ip_route_output_key(tunnel->net, &fl4);
 		if (IS_ERR(rt)) {
 			DEV_STATS_INC(dev, tx_carrier_errors);
+			reason = SKB_DROP_REASON_IP_OUTNOROUTES;
 			goto tx_error;
 		}
 		if (use_cache)
@@ -639,6 +645,7 @@ void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 	if (rt->dst.dev == dev) {
 		ip_rt_put(rt);
 		DEV_STATS_INC(dev, collisions);
+		reason = SKB_DROP_REASON_RECURSION_LIMIT;
 		goto tx_error;
 	}
 
@@ -647,6 +654,7 @@ void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 	if (tnl_update_pmtu(dev, skb, rt, df, inner_iph, tunnel_hlen,
 			    key->u.ipv4.dst, true)) {
 		ip_rt_put(rt);
+		reason = SKB_DROP_REASON_PKT_TOO_BIG;
 		goto tx_error;
 	}
 
@@ -664,6 +672,7 @@ void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 	headroom += LL_RESERVED_SPACE(rt->dst.dev) + rt->dst.header_len;
 	if (skb_cow_head(skb, headroom)) {
 		ip_rt_put(rt);
+		reason = SKB_DROP_REASON_NOMEM;
 		goto tx_dropped;
 	}
 
@@ -678,13 +687,14 @@ void ip_md_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 tx_dropped:
 	DEV_STATS_INC(dev, tx_dropped);
 kfree:
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 }
 EXPORT_SYMBOL_GPL(ip_md_tunnel_xmit);
 
 void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 		    const struct iphdr *tnl_params, u8 protocol)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct ip_tunnel *tunnel = netdev_priv(dev);
 	struct ip_tunnel_info *tun_info = NULL;
 	const struct iphdr *inner_iph;
@@ -712,6 +722,7 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 
 		if (!skb_dst(skb)) {
 			DEV_STATS_INC(dev, tx_fifo_errors);
+			reason = SKB_DROP_REASON_NO_TX_TARGET;
 			goto tx_error;
 		}
 
@@ -725,9 +736,8 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 		} else if (payload_protocol == htons(ETH_P_IP)) {
 			rt = skb_rtable(skb);
 			dst = rt_nexthop(rt, inner_iph->daddr);
-		}
 #if IS_ENABLED(CONFIG_IPV6)
-		else if (payload_protocol == htons(ETH_P_IPV6)) {
+		} else if (payload_protocol == htons(ETH_P_IPV6)) {
 			const struct in6_addr *addr6;
 			struct neighbour *neigh;
 			bool do_tx_error_icmp;
@@ -735,8 +745,10 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 
 			neigh = dst_neigh_lookup(skb_dst(skb),
 						 &ipv6_hdr(skb)->daddr);
-			if (!neigh)
+			if (!neigh) {
+				reason = SKB_DROP_REASON_NEIGH_CREATEFAIL;
 				goto tx_error;
+			}
 
 			addr6 = (const struct in6_addr *)&neigh->primary_key;
 			addr_type = ipv6_addr_type(addr6);
@@ -753,12 +765,15 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 				dst = addr6->s6_addr32[3];
 			}
 			neigh_release(neigh);
-			if (do_tx_error_icmp)
+			if (do_tx_error_icmp) {
+				reason = SKB_DROP_REASON_NO_TX_TARGET;
 				goto tx_error_icmp;
-		}
+			}
 #endif
-		else
+		} else {
+			reason = SKB_DROP_REASON_NO_TX_TARGET;
 			goto tx_error;
+		}
 
 		if (!md)
 			connected = false;
@@ -781,8 +796,10 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 			    tunnel->net, READ_ONCE(tunnel->parms.link),
 			    tunnel->fwmark, skb_get_hash(skb), 0);
 
-	if (ip_tunnel_encap(skb, &tunnel->encap, &protocol, &fl4) < 0)
+	if (ip_tunnel_encap(skb, &tunnel->encap, &protocol, &fl4) < 0) {
+		reason = SKB_DROP_REASON_TUNNEL_ENCAP;
 		goto tx_error;
+	}
 
 	if (connected && md) {
 		use_cache = ip_tunnel_dst_cache_usable(skb, tun_info);
@@ -799,6 +816,7 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 
 		if (IS_ERR(rt)) {
 			DEV_STATS_INC(dev, tx_carrier_errors);
+			reason = SKB_DROP_REASON_IP_OUTNOROUTES;
 			goto tx_error;
 		}
 		if (use_cache)
@@ -812,6 +830,7 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 	if (rt->dst.dev == dev) {
 		ip_rt_put(rt);
 		DEV_STATS_INC(dev, collisions);
+		reason = SKB_DROP_REASON_RECURSION_LIMIT;
 		goto tx_error;
 	}
 
@@ -821,6 +840,7 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 
 	if (tnl_update_pmtu(dev, skb, rt, df, inner_iph, 0, 0, false)) {
 		ip_rt_put(rt);
+		reason = SKB_DROP_REASON_PKT_TOO_BIG;
 		goto tx_error;
 	}
 
@@ -855,7 +875,7 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 	if (skb_cow_head(skb, max_headroom)) {
 		ip_rt_put(rt);
 		DEV_STATS_INC(dev, tx_dropped);
-		kfree_skb(skb);
+		kfree_skb_reason(skb, SKB_DROP_REASON_NOMEM);
 		return;
 	}
 
@@ -871,7 +891,7 @@ void ip_tunnel_xmit(struct sk_buff *skb, struct net_device *dev,
 #endif
 tx_error:
 	DEV_STATS_INC(dev, tx_errors);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 }
 EXPORT_SYMBOL_GPL(ip_tunnel_xmit);
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 10/14] ip_gre: add drop reasons to the transmit path
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (8 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 09/14] ip_tunnel: add drop reasons to the transmit path Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 11/14] ip6_gre: make prepare_ip6gre_xmit_other() void Anton Danilov
                   ` (3 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

Each transmit function of ip_gre ends all of its failures in one
kfree_skb() and a tx_dropped increment, so a drop can be traced to the
function and to nothing more precise than "the tunnel did not send it".

No new reason is needed. The length helpers already compute one, so
pskb_inet_may_pull_reason() and pskb_may_pull_reason() are used instead
of their boolean wrappers, and the rest reuses:

 - SKB_DROP_REASON_NOMEM for the headroom expansions, the offload
   handling and the trims,
 - SKB_DROP_REASON_TUNNEL_TXINFO for the collect_md paths, when the
   metadata is missing or incomplete,
 - SKB_DROP_REASON_UNHANDLED_PROTO for an ERSPAN version that is not
   implemented,
 - SKB_DROP_REASON_SKB_CSUM when the checksum starts before the data
   the tunnel is about to send.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 net/ipv4/ip_gre.c | 104 ++++++++++++++++++++++++++++++++++------------
 1 file changed, 77 insertions(+), 27 deletions(-)

diff --git a/net/ipv4/ip_gre.c b/net/ipv4/ip_gre.c
index afd8ece02d7f..58234e857d6c 100644
--- a/net/ipv4/ip_gre.c
+++ b/net/ipv4/ip_gre.c
@@ -509,6 +509,7 @@ static int gre_handle_offloads(struct sk_buff *skb, bool csum)
 static void gre_fb_xmit(struct sk_buff *skb, struct net_device *dev,
 			__be16 proto)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct ip_tunnel *tunnel = netdev_priv(dev);
 	IP_TUNNEL_DECLARE_FLAGS(flags) = { };
 	struct ip_tunnel_info *tun_info;
@@ -517,19 +518,25 @@ static void gre_fb_xmit(struct sk_buff *skb, struct net_device *dev,
 
 	tun_info = skb_tunnel_info(skb);
 	if (unlikely(!tun_info || !(tun_info->mode & IP_TUNNEL_INFO_TX) ||
-		     ip_tunnel_info_af(tun_info) != AF_INET))
+		     ip_tunnel_info_af(tun_info) != AF_INET)) {
+		reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 		goto err_free_skb;
+	}
 
 	key = &tun_info->key;
 	tunnel_hlen = gre_calc_hlen(key->tun_flags);
 
-	if (skb_cow_head(skb, dev->needed_headroom))
+	if (skb_cow_head(skb, dev->needed_headroom)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto err_free_skb;
+	}
 
 	/* Push Tunnel header. */
 	if (gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT,
-					      tunnel->parms.o_flags)))
+					      tunnel->parms.o_flags))) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto err_free_skb;
+	}
 
 	__set_bit(IP_TUNNEL_CSUM_BIT, flags);
 	__set_bit(IP_TUNNEL_KEY_BIT, flags);
@@ -546,12 +553,13 @@ static void gre_fb_xmit(struct sk_buff *skb, struct net_device *dev,
 	return;
 
 err_free_skb:
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	DEV_STATS_INC(dev, tx_dropped);
 }
 
 static void erspan_fb_xmit(struct sk_buff *skb, struct net_device *dev)
 {
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	struct ip_tunnel *tunnel = netdev_priv(dev);
 	IP_TUNNEL_DECLARE_FLAGS(flags) = { };
 	struct ip_tunnel_info *tun_info;
@@ -565,29 +573,41 @@ static void erspan_fb_xmit(struct sk_buff *skb, struct net_device *dev)
 
 	tun_info = skb_tunnel_info(skb);
 	if (unlikely(!tun_info || !(tun_info->mode & IP_TUNNEL_INFO_TX) ||
-		     ip_tunnel_info_af(tun_info) != AF_INET))
+		     ip_tunnel_info_af(tun_info) != AF_INET)) {
+		reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 		goto err_free_skb;
+	}
 
 	key = &tun_info->key;
-	if (!test_bit(IP_TUNNEL_ERSPAN_OPT_BIT, tun_info->key.tun_flags))
+	if (!test_bit(IP_TUNNEL_ERSPAN_OPT_BIT, tun_info->key.tun_flags)) {
+		reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 		goto err_free_skb;
-	if (tun_info->options_len < sizeof(*md))
+	}
+	if (tun_info->options_len < sizeof(*md)) {
+		reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 		goto err_free_skb;
+	}
 	md = ip_tunnel_info_opts(tun_info);
 
 	/* ERSPAN has fixed 8 byte GRE header */
 	version = md->version;
 	tunnel_hlen = 8 + erspan_hdr_len(version);
 
-	if (skb_cow_head(skb, dev->needed_headroom))
+	if (skb_cow_head(skb, dev->needed_headroom)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto err_free_skb;
+	}
 
-	if (gre_handle_offloads(skb, false))
+	if (gre_handle_offloads(skb, false)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto err_free_skb;
+	}
 
 	if (skb->len > dev->mtu + dev->hard_header_len) {
-		if (pskb_trim(skb, dev->mtu + dev->hard_header_len))
+		if (pskb_trim(skb, dev->mtu + dev->hard_header_len)) {
+			reason = SKB_DROP_REASON_NOMEM;
 			goto err_free_skb;
+		}
 		truncate = true;
 	}
 
@@ -619,6 +639,7 @@ static void erspan_fb_xmit(struct sk_buff *skb, struct net_device *dev)
 				       truncate, true);
 		proto = htons(ETH_P_ERSPAN2);
 	} else {
+		reason = SKB_DROP_REASON_UNHANDLED_PROTO;
 		goto err_free_skb;
 	}
 
@@ -631,7 +652,7 @@ static void erspan_fb_xmit(struct sk_buff *skb, struct net_device *dev)
 	return;
 
 err_free_skb:
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	DEV_STATS_INC(dev, tx_dropped);
 }
 
@@ -665,8 +686,10 @@ static netdev_tx_t ipgre_xmit(struct sk_buff *skb,
 	struct ip_tunnel *tunnel = netdev_priv(dev);
 	IP_TUNNEL_DECLARE_FLAGS(flags);
 	const struct iphdr *tnl_params;
+	enum skb_drop_reason reason;
 
-	if (!pskb_inet_may_pull(skb))
+	reason = pskb_inet_may_pull_reason(skb);
+	if (reason)
 		goto free_skb;
 
 	if (tunnel->collect_md) {
@@ -677,10 +700,13 @@ static netdev_tx_t ipgre_xmit(struct sk_buff *skb,
 	if (dev->header_ops) {
 		int pull_len = tunnel->hlen + sizeof(struct iphdr);
 
-		if (skb_cow_head(skb, 0))
+		if (skb_cow_head(skb, 0)) {
+			reason = SKB_DROP_REASON_NOMEM;
 			goto free_skb;
+		}
 
-		if (!pskb_may_pull(skb, pull_len))
+		reason = pskb_may_pull_reason(skb, pull_len);
+		if (reason)
 			goto free_skb;
 
 		tnl_params = (const struct iphdr *)skb->data;
@@ -690,25 +716,32 @@ static netdev_tx_t ipgre_xmit(struct sk_buff *skb,
 		skb_reset_mac_header(skb);
 
 		if (skb->ip_summed == CHECKSUM_PARTIAL &&
-		    skb_checksum_start(skb) < skb->data)
+		    skb_checksum_start(skb) < skb->data) {
+			reason = SKB_DROP_REASON_SKB_CSUM;
 			goto free_skb;
+		}
 	} else {
-		if (skb_cow_head(skb, dev->needed_headroom))
+		if (skb_cow_head(skb, dev->needed_headroom)) {
+			reason = SKB_DROP_REASON_NOMEM;
 			goto free_skb;
+		}
 
 		tnl_params = &tunnel->parms.iph;
 	}
 
 	ip_tunnel_flags_copy(flags, tunnel->parms.o_flags);
 
-	if (gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT, flags)))
+	if (gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT, flags))) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto free_skb;
+	}
 
 	__gre_xmit(skb, dev, tnl_params, skb->protocol, flags);
 	return NETDEV_TX_OK;
 
 free_skb:
-	kfree_skb(skb);
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
+	kfree_skb_reason(skb, reason);
 	DEV_STATS_INC(dev, tx_dropped);
 	return NETDEV_TX_OK;
 }
@@ -718,10 +751,12 @@ static netdev_tx_t erspan_xmit(struct sk_buff *skb,
 {
 	struct ip_tunnel *tunnel = netdev_priv(dev);
 	IP_TUNNEL_DECLARE_FLAGS(flags);
+	enum skb_drop_reason reason;
 	bool truncate = false;
 	__be16 proto;
 
-	if (!pskb_inet_may_pull(skb))
+	reason = pskb_inet_may_pull_reason(skb);
+	if (reason)
 		goto free_skb;
 
 	if (tunnel->collect_md) {
@@ -729,15 +764,21 @@ static netdev_tx_t erspan_xmit(struct sk_buff *skb,
 		return NETDEV_TX_OK;
 	}
 
-	if (gre_handle_offloads(skb, false))
+	if (gre_handle_offloads(skb, false)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto free_skb;
+	}
 
-	if (skb_cow_head(skb, dev->needed_headroom))
+	if (skb_cow_head(skb, dev->needed_headroom)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto free_skb;
+	}
 
 	if (skb->len > dev->mtu + dev->hard_header_len) {
-		if (pskb_trim(skb, dev->mtu + dev->hard_header_len))
+		if (pskb_trim(skb, dev->mtu + dev->hard_header_len)) {
+			reason = SKB_DROP_REASON_NOMEM;
 			goto free_skb;
+		}
 		truncate = true;
 	}
 
@@ -758,6 +799,7 @@ static netdev_tx_t erspan_xmit(struct sk_buff *skb,
 				       truncate, true);
 		proto = htons(ETH_P_ERSPAN2);
 	} else {
+		reason = SKB_DROP_REASON_UNHANDLED_PROTO;
 		goto free_skb;
 	}
 
@@ -766,7 +808,8 @@ static netdev_tx_t erspan_xmit(struct sk_buff *skb,
 	return NETDEV_TX_OK;
 
 free_skb:
-	kfree_skb(skb);
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
+	kfree_skb_reason(skb, reason);
 	DEV_STATS_INC(dev, tx_dropped);
 	return NETDEV_TX_OK;
 }
@@ -776,8 +819,10 @@ static netdev_tx_t gre_tap_xmit(struct sk_buff *skb,
 {
 	struct ip_tunnel *tunnel = netdev_priv(dev);
 	IP_TUNNEL_DECLARE_FLAGS(flags);
+	enum skb_drop_reason reason;
 
-	if (!pskb_inet_may_pull(skb))
+	reason = pskb_inet_may_pull_reason(skb);
+	if (reason)
 		goto free_skb;
 
 	if (tunnel->collect_md) {
@@ -787,17 +832,22 @@ static netdev_tx_t gre_tap_xmit(struct sk_buff *skb,
 
 	ip_tunnel_flags_copy(flags, tunnel->parms.o_flags);
 
-	if (gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT, flags)))
+	if (gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT, flags))) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto free_skb;
+	}
 
-	if (skb_cow_head(skb, dev->needed_headroom))
+	if (skb_cow_head(skb, dev->needed_headroom)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto free_skb;
+	}
 
 	__gre_xmit(skb, dev, &tunnel->parms.iph, htons(ETH_P_TEB), flags);
 	return NETDEV_TX_OK;
 
 free_skb:
-	kfree_skb(skb);
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
+	kfree_skb_reason(skb, reason);
 	DEV_STATS_INC(dev, tx_dropped);
 	return NETDEV_TX_OK;
 }
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 11/14] ip6_gre: make prepare_ip6gre_xmit_other() void
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (9 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 10/14] ip_gre: " Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 12/14] ip6_tunnel: make ip6_tnl_xmit() return a drop reason Anton Danilov
                   ` (2 subsequent siblings)
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

prepare_ip6gre_xmit_other() copies the flow template of the tunnel and
picks up its encapsulation limit, DS field and mark. Unlike its IPv6
sibling, which fails when the packet's tunnel encapsulation limit option
is 0 and so forbids encapsulating it again, it has nothing to fail on:
its only return statement is "return 0", and it has been that way since
commit 41337f52b967 ("ip6_gre: set DSCP for non-IP") added the function.
Its caller still checks the result and bails out on a branch that never
runs.

Make it void and drop the check, the way prepare_ip6gre_xmit_ipv4() is
already called. The next two patches give every failing branch of the
transmit path a drop reason, and this one would otherwise get a reason
it can never report.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 net/ipv6/ip6_gre.c | 16 +++++++---------
 1 file changed, 7 insertions(+), 9 deletions(-)

diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c
index 3333a0ab180c..ba080a58ecb2 100644
--- a/net/ipv6/ip6_gre.c
+++ b/net/ipv6/ip6_gre.c
@@ -681,10 +681,10 @@ static int prepare_ip6gre_xmit_ipv6(struct sk_buff *skb,
 	return 0;
 }
 
-static int prepare_ip6gre_xmit_other(struct sk_buff *skb,
-				     struct net_device *dev,
-				     struct flowi6 *fl6, __u8 *dsfield,
-				     int *encap_limit)
+static void prepare_ip6gre_xmit_other(struct sk_buff *skb,
+				      struct net_device *dev,
+				      struct flowi6 *fl6, __u8 *dsfield,
+				      int *encap_limit)
 {
 	struct ip6_tnl *t = netdev_priv(dev);
 
@@ -704,8 +704,6 @@ static int prepare_ip6gre_xmit_other(struct sk_buff *skb,
 		fl6->flowi6_mark = t->parms.fwmark;
 
 	fl6->flowi6_uid = sock_net_uid(dev_net(dev), NULL);
-
-	return 0;
 }
 
 static struct ip_tunnel_info *skb_tunnel_info_txcheck(struct sk_buff *skb)
@@ -866,9 +864,9 @@ static int ip6gre_xmit_other(struct sk_buff *skb, struct net_device *dev)
 	__u32 mtu;
 	int err;
 
-	if (!t->parms.collect_md &&
-	    prepare_ip6gre_xmit_other(skb, dev, &fl6, &dsfield, &encap_limit))
-		return -1;
+	if (!t->parms.collect_md)
+		prepare_ip6gre_xmit_other(skb, dev, &fl6,
+					  &dsfield, &encap_limit);
 
 	err = gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT,
 						t->parms.o_flags));
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 12/14] ip6_tunnel: make ip6_tnl_xmit() return a drop reason
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (10 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 11/14] ip6_gre: make prepare_ip6gre_xmit_other() void Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 13/14] ip6_gre: add drop reasons to the transmit path Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 14/14] vxlan: report a circular route as SKB_DROP_REASON_RECURSION_LIMIT Anton Danilov
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

Give the IPv6 tunnel transmit path drop reasons, as ip_tunnel_xmit()
and the ip_gre transmit handlers now have. The difference is that
ip6_tnl_xmit() does not free the packet itself: it returns an error and
its callers free it, so the reason has to travel with it.

Make ip6_tnl_xmit() return the drop reason, SKB_NOT_DROPPED_YET on
success, instead of 0, -1 or an errno, and do the same for
ipxip6_tnl_xmit(), __gre6_xmit() and the ip6gre_xmit_*() helpers, so
that the ndo_start_xmit handlers, where the packet is actually freed,
can report it. The callers only told success from failure and looked
for -EMSGSIZE to send an ICMP error back. ip6_tnl_xmit() returns
-EMSGSIZE only for a packet that exceeds the path MTU, which is exactly
when it reports SKB_DROP_REASON_PKT_TOO_BIG, so the ICMP error is now
sent when that reason is returned. A failed xfrm lookup never returns
-EMSGSIZE, so its errno only ever meant failure and goes away.
__gre6_xmit() was declared as returning netdev_tx_t while it returned an
errno, and now returns the reason as well.

The ip6_gre helpers change here because they pass on the return value of
ip6_tnl_xmit(). The ip6_gre handlers take the reason from them but still
free the packet with kfree_skb(); the next patch converts them.
ip6_tnl_start_xmit() frees with the reason and, as in ip_gre, takes the
length reason from pskb_inet_may_pull_reason() instead of
pskb_inet_may_pull(). The other reasons are the ones already used on the
IPv4 side:

 - SKB_DROP_REASON_PKT_TOO_BIG for a packet that exceeds the path MTU,
 - SKB_DROP_REASON_IP_OUTNOROUTES for the route and xfrm lookups,
   including the source address selection that a collect_md tunnel does
   when the flow has no source address,
 - SKB_DROP_REASON_NO_TX_TARGET when an NBMA tunnel gets an skb with no
   destination to derive its endpoint from, or finds no endpoint for it,
   such as an IPv4 packet whose route has no IPv6 gateway or an MPLS
   packet,
 - SKB_DROP_REASON_NEIGH_CREATEFAIL when the NBMA neighbour lookup
   fails,
 - SKB_DROP_REASON_RECURSION_LIMIT for a route pointing back at the
   tunnel, and for the trivial tunnelling loop ip6_tnl_addr_conflict()
   guards against, a packet whose source is the exit point of the
   tunnel,
 - SKB_DROP_REASON_NOMEM for the allocations,
 - SKB_DROP_REASON_TUNNEL_TXINFO for the collect_md metadata checks,
 - SKB_DROP_REASON_TUNNEL_ENCAP when the encapsulation header cannot be
   built, and for a collect_md tunnel that has an encapsulation
   configured, which ip6_tnl_xmit() does not support,
 - SKB_DROP_REASON_UNHANDLED_PROTO for a payload the tunnel does not
   carry, either by its mode or because it is neither IPv4, IPv6 nor
   MPLS.

Two more are IPv6 specific. SKB_DROP_REASON_IPV6_BAD_EXTHDR is used when
the packet's tunnel encapsulation limit option is 0, which forbids
encapsulating it again. SKB_DROP_REASON_DEV_READY is used when
ip6_tnl_xmit_ctl() refuses the transmit. It refuses when the local
address is not configured yet, when the remote address belongs to this
host (the routing loop it warns about), and when the tunnel cannot
transmit with the addresses of the packet. DEV_READY describes the first
case; the other two are reported under it as well, because
ip6_tnl_xmit_ctl() returns only 0 or 1. Telling them apart needs it to
return a reason, which is left for a follow-up, together with
ip6_tnl_rcv_ctl() on the receive side. An NBMA tunnel that found no
endpoint for the packet would be refused there as well; it is caught
before the call, as SKB_DROP_REASON_NO_TX_TARGET.

A collect_md tunnel has no fixed exit point and its raddr is normally
::, as is that of an NBMA ip6tnl device. ip6_tnl_addr_conflict() and,
on a collect_md device, the same check in ip6gre_xmit_ipv6() therefore
also drop the packets such a tunnel sends from ::. They were dropped
before as well; SKB_DROP_REASON_RECURSION_LIMIT names the check that
drops them.
The next patch has the ip6_gre handler drop a packet that carries no
IPv6 tunnel metadata before that check, as
SKB_DROP_REASON_TUNNEL_TXINFO, so that the DAD probes of an ip6gretap
device are reported like its other packets without metadata.
ip6_tnl_start_xmit() still reports such a packet from :: as
SKB_DROP_REASON_RECURSION_LIMIT; an ip6tnl device is IFF_NOARP and does
no DAD, so its own probes do not end up there.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 include/net/ip6_tunnel.h |   5 +-
 net/ipv6/ip6_gre.c       |  84 ++++++++++++++++---------------
 net/ipv6/ip6_tunnel.c    | 106 ++++++++++++++++++++++++---------------
 3 files changed, 113 insertions(+), 82 deletions(-)

diff --git a/include/net/ip6_tunnel.h b/include/net/ip6_tunnel.h
index d1f0a427e9c8..363eed61b296 100644
--- a/include/net/ip6_tunnel.h
+++ b/include/net/ip6_tunnel.h
@@ -143,8 +143,9 @@ int ip6_tnl_rcv(struct ip6_tnl *tunnel, struct sk_buff *skb,
 		bool log_ecn_error);
 int ip6_tnl_xmit_ctl(struct ip6_tnl *t, const struct in6_addr *laddr,
 		     const struct in6_addr *raddr);
-int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
-		 struct flowi6 *fl6, int encap_limit, __u32 *pmtu, __u8 proto);
+enum skb_drop_reason
+ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
+	     struct flowi6 *fl6, int encap_limit, __u32 *pmtu, __u8 proto);
 __u16 ip6_tnl_parse_tlv_enc_lim(struct sk_buff *skb, __u8 *raw);
 __u32 ip6_tnl_get_cap(struct ip6_tnl *t, const struct in6_addr *laddr,
 			     const struct in6_addr *raddr);
diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c
index ba080a58ecb2..9e94ea6b6c20 100644
--- a/net/ipv6/ip6_gre.c
+++ b/net/ipv6/ip6_gre.c
@@ -717,10 +717,10 @@ static struct ip_tunnel_info *skb_tunnel_info_txcheck(struct sk_buff *skb)
 	return tun_info;
 }
 
-static netdev_tx_t __gre6_xmit(struct sk_buff *skb,
-			       struct net_device *dev, __u8 dsfield,
-			       struct flowi6 *fl6, int encap_limit,
-			       __u32 *pmtu, __be16 proto)
+static enum skb_drop_reason __gre6_xmit(struct sk_buff *skb,
+					struct net_device *dev, __u8 dsfield,
+					struct flowi6 *fl6, int encap_limit,
+					__u32 *pmtu, __be16 proto)
 {
 	struct ip6_tnl *tunnel = netdev_priv(dev);
 	IP_TUNNEL_DECLARE_FLAGS(flags);
@@ -745,7 +745,7 @@ static netdev_tx_t __gre6_xmit(struct sk_buff *skb,
 		tun_info = skb_tunnel_info_txcheck(skb);
 		if (IS_ERR(tun_info) ||
 		    unlikely(ip_tunnel_info_af(tun_info) != AF_INET6))
-			return -EINVAL;
+			return SKB_DROP_REASON_TUNNEL_TXINFO;
 
 		key = &tun_info->key;
 		memset(fl6, 0, sizeof(*fl6));
@@ -764,7 +764,7 @@ static netdev_tx_t __gre6_xmit(struct sk_buff *skb,
 		tun_hlen = gre_calc_hlen(flags);
 
 		if (skb_cow_head(skb, dev->needed_headroom ?: tun_hlen + tunnel->encap_hlen))
-			return -ENOMEM;
+			return SKB_DROP_REASON_NOMEM;
 
 		gre_build_header(skb, tun_hlen,
 				 flags, protocol,
@@ -775,7 +775,7 @@ static netdev_tx_t __gre6_xmit(struct sk_buff *skb,
 
 	} else {
 		if (skb_cow_head(skb, dev->needed_headroom ?: tunnel->hlen))
-			return -ENOMEM;
+			return SKB_DROP_REASON_NOMEM;
 
 		ip_tunnel_flags_copy(flags, tunnel->parms.o_flags);
 
@@ -790,9 +790,11 @@ static netdev_tx_t __gre6_xmit(struct sk_buff *skb,
 			    NEXTHDR_GRE);
 }
 
-static inline int ip6gre_xmit_ipv4(struct sk_buff *skb, struct net_device *dev)
+static inline enum skb_drop_reason ip6gre_xmit_ipv4(struct sk_buff *skb,
+						    struct net_device *dev)
 {
 	struct ip6_tnl *t = netdev_priv(dev);
+	enum skb_drop_reason reason;
 	int encap_limit = -1;
 	struct flowi6 fl6;
 	__u8 dsfield = 0;
@@ -808,54 +810,56 @@ static inline int ip6gre_xmit_ipv4(struct sk_buff *skb, struct net_device *dev)
 	err = gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT,
 						t->parms.o_flags));
 	if (err)
-		return -1;
+		return SKB_DROP_REASON_NOMEM;
 
-	err = __gre6_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu,
-			  skb->protocol);
-	if (err != 0) {
+	reason = __gre6_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu,
+			     skb->protocol);
+	if (reason) {
 		/* XXX: send ICMP error even if DF is not set. */
-		if (err == -EMSGSIZE)
+		if (reason == SKB_DROP_REASON_PKT_TOO_BIG)
 			icmp_ndo_send(skb, ICMP_DEST_UNREACH, ICMP_FRAG_NEEDED,
 				      htonl(mtu));
-		return -1;
+		return reason;
 	}
 
-	return 0;
+	return SKB_NOT_DROPPED_YET;
 }
 
-static inline int ip6gre_xmit_ipv6(struct sk_buff *skb, struct net_device *dev)
+static inline enum skb_drop_reason ip6gre_xmit_ipv6(struct sk_buff *skb,
+						    struct net_device *dev)
 {
 	struct ip6_tnl *t = netdev_priv(dev);
 	struct ipv6hdr *ipv6h = ipv6_hdr(skb);
+	enum skb_drop_reason reason;
 	int encap_limit = -1;
 	struct flowi6 fl6;
 	__u8 dsfield = 0;
 	__u32 mtu;
-	int err;
 
 	if (ipv6_addr_equal(&t->parms.raddr, &ipv6h->saddr))
-		return -1;
+		return SKB_DROP_REASON_RECURSION_LIMIT;
 
 	if (!t->parms.collect_md &&
 	    prepare_ip6gre_xmit_ipv6(skb, dev, &fl6, &dsfield, &encap_limit))
-		return -1;
+		return SKB_DROP_REASON_IPV6_BAD_EXTHDR;
 
 	if (gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT,
 					      t->parms.o_flags)))
-		return -1;
+		return SKB_DROP_REASON_NOMEM;
 
-	err = __gre6_xmit(skb, dev, dsfield, &fl6, encap_limit,
-			  &mtu, skb->protocol);
-	if (err != 0) {
-		if (err == -EMSGSIZE)
+	reason = __gre6_xmit(skb, dev, dsfield, &fl6, encap_limit,
+			     &mtu, skb->protocol);
+	if (reason) {
+		if (reason == SKB_DROP_REASON_PKT_TOO_BIG)
 			icmpv6_ndo_send(skb, ICMPV6_PKT_TOOBIG, 0, mtu);
-		return -1;
+		return reason;
 	}
 
-	return 0;
+	return SKB_NOT_DROPPED_YET;
 }
 
-static int ip6gre_xmit_other(struct sk_buff *skb, struct net_device *dev)
+static enum skb_drop_reason ip6gre_xmit_other(struct sk_buff *skb,
+					      struct net_device *dev)
 {
 	struct ip6_tnl *t = netdev_priv(dev);
 	int encap_limit = -1;
@@ -871,10 +875,10 @@ static int ip6gre_xmit_other(struct sk_buff *skb, struct net_device *dev)
 	err = gre_handle_offloads(skb, test_bit(IP_TUNNEL_CSUM_BIT,
 						t->parms.o_flags));
 	if (err)
-		return err;
-	err = __gre6_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu, skb->protocol);
+		return SKB_DROP_REASON_NOMEM;
 
-	return err;
+	return __gre6_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu,
+			   skb->protocol);
 }
 
 static netdev_tx_t ip6gre_tunnel_xmit(struct sk_buff *skb,
@@ -882,8 +886,8 @@ static netdev_tx_t ip6gre_tunnel_xmit(struct sk_buff *skb,
 {
 	struct ip_tunnel_info *tun_info = NULL;
 	struct ip6_tnl *t = netdev_priv(dev);
+	enum skb_drop_reason reason;
 	__be16 payload_protocol;
-	int ret;
 
 	if (!pskb_inet_may_pull(skb))
 		goto tx_err;
@@ -897,17 +901,17 @@ static netdev_tx_t ip6gre_tunnel_xmit(struct sk_buff *skb,
 	payload_protocol = skb_protocol(skb, true);
 	switch (payload_protocol) {
 	case htons(ETH_P_IP):
-		ret = ip6gre_xmit_ipv4(skb, dev);
+		reason = ip6gre_xmit_ipv4(skb, dev);
 		break;
 	case htons(ETH_P_IPV6):
-		ret = ip6gre_xmit_ipv6(skb, dev);
+		reason = ip6gre_xmit_ipv6(skb, dev);
 		break;
 	default:
-		ret = ip6gre_xmit_other(skb, dev);
+		reason = ip6gre_xmit_other(skb, dev);
 		break;
 	}
 
-	if (ret < 0)
+	if (reason)
 		goto tx_err;
 
 	return NETDEV_TX_OK;
@@ -927,11 +931,11 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 	struct ip6_tnl *t = netdev_priv(dev);
 	struct dst_entry *dst = skb_dst(skb);
 	IP_TUNNEL_DECLARE_FLAGS(flags) = { };
+	enum skb_drop_reason reason;
 	bool truncate = false;
 	int encap_limit = -1;
 	__u8 dsfield = false;
 	struct flowi6 fl6;
-	int err = -EINVAL;
 	__be16 proto;
 	__u32 mtu;
 	int nhoff;
@@ -1066,11 +1070,11 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 		if (dst_mtu(dst) > mtu)
 			dst->ops->update_pmtu(dst, NULL, skb, mtu, false);
 	}
-	err = ip6_tnl_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu,
-			   NEXTHDR_GRE);
-	if (err != 0) {
+	reason = ip6_tnl_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu,
+			      NEXTHDR_GRE);
+	if (reason) {
 		/* XXX: send ICMP error even if DF is not set. */
-		if (err == -EMSGSIZE) {
+		if (reason == SKB_DROP_REASON_PKT_TOO_BIG) {
 			if (skb->protocol == htons(ETH_P_IP))
 				icmp_ndo_send(skb, ICMP_DEST_UNREACH,
 					      ICMP_FRAG_NEEDED, htonl(mtu));
diff --git a/net/ipv6/ip6_tunnel.c b/net/ipv6/ip6_tunnel.c
index 52f6a38657b9..7e444115f431 100644
--- a/net/ipv6/ip6_tunnel.c
+++ b/net/ipv6/ip6_tunnel.c
@@ -1115,14 +1115,14 @@ EXPORT_SYMBOL_GPL(ip6_tnl_xmit_ctl);
  *   it.
  *
  * Return:
- *   0 on success
- *   -1 fail
- *   %-EMSGSIZE message too big. return mtu in this case.
+ *   %SKB_NOT_DROPPED_YET on success, otherwise the drop reason.
+ *   %SKB_DROP_REASON_PKT_TOO_BIG means the message is too big, the path MTU
+ *   is stored in @pmtu in this case.
  **/
 
-int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
-		 struct flowi6 *fl6, int encap_limit, __u32 *pmtu,
-		 __u8 proto)
+enum skb_drop_reason
+ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
+	     struct flowi6 *fl6, int encap_limit, __u32 *pmtu, __u8 proto)
 {
 	struct ip6_tnl *t = netdev_priv(dev);
 	struct net *net = t->net;
@@ -1133,11 +1133,11 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 	int err_count, mtu;
 	unsigned int eth_hlen = t->dev->type == ARPHRD_ETHER ? ETH_HLEN : 0;
 	unsigned int psh_hlen = sizeof(struct ipv6hdr) + t->encap_hlen;
+	enum skb_drop_reason reason = SKB_DROP_REASON_NOT_SPECIFIED;
 	unsigned int max_headroom = psh_hlen;
 	__be16 payload_protocol;
 	bool use_cache = false;
 	u8 hop_limit;
-	int err = -1;
 
 	payload_protocol = skb_protocol(skb, true);
 
@@ -1155,13 +1155,17 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 			struct neighbour *neigh;
 			int addr_type;
 
-			if (!skb_dst(skb))
+			if (!skb_dst(skb)) {
+				reason = SKB_DROP_REASON_NO_TX_TARGET;
 				goto tx_err_link_failure;
+			}
 
 			neigh = dst_neigh_lookup(skb_dst(skb),
 						 &ipv6_hdr(skb)->daddr);
-			if (!neigh)
+			if (!neigh) {
+				reason = SKB_DROP_REASON_NEIGH_CREATEFAIL;
 				goto tx_err_link_failure;
+			}
 
 			addr6 = (struct in6_addr *)&neigh->primary_key;
 			addr_type = ipv6_addr_type(addr6);
@@ -1174,12 +1178,19 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 		} else if (payload_protocol == htons(ETH_P_IP)) {
 			const struct rtable *rt = skb_rtable(skb);
 
-			if (!rt)
+			if (!rt) {
+				reason = SKB_DROP_REASON_NO_TX_TARGET;
 				goto tx_err_link_failure;
+			}
 
 			if (rt->rt_gw_family == AF_INET6)
 				memcpy(&fl6->daddr, &rt->rt_gw6, sizeof(fl6->daddr));
 		}
+
+		if (ipv6_addr_any(&fl6->daddr)) {
+			reason = SKB_DROP_REASON_NO_TX_TARGET;
+			goto tx_err_link_failure;
+		}
 	} else if (t->parms.proto != 0 && !(t->parms.flags &
 					    (IP6_TNL_F_USE_ORIG_TCLASS |
 					     IP6_TNL_F_USE_ORIG_FWMARK))) {
@@ -1192,8 +1203,10 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 	if (use_cache)
 		dst = dst_cache_get(&t->dst_cache);
 
-	if (!ip6_tnl_xmit_ctl(t, &fl6->saddr, &fl6->daddr))
+	if (!ip6_tnl_xmit_ctl(t, &fl6->saddr, &fl6->daddr)) {
+		reason = SKB_DROP_REASON_DEV_READY;
 		goto tx_err_link_failure;
+	}
 
 	if (!dst) {
 route_lookup:
@@ -1202,18 +1215,22 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 
 		dst = ip6_route_output(net, NULL, fl6);
 
-		if (dst->error)
+		if (dst->error) {
+			reason = SKB_DROP_REASON_IP_OUTNOROUTES;
 			goto tx_err_link_failure;
+		}
 		dst = xfrm_lookup(net, dst, flowi6_to_flowi(fl6), NULL, 0);
 		if (IS_ERR(dst)) {
-			err = PTR_ERR(dst);
 			dst = NULL;
+			reason = SKB_DROP_REASON_IP_OUTNOROUTES;
 			goto tx_err_link_failure;
 		}
 		if (t->parms.collect_md && ipv6_addr_any(&fl6->saddr) &&
 		    ipv6_dev_get_saddr(net, ip6_dst_idev(dst)->dev,
-				       &fl6->daddr, 0, &fl6->saddr))
+				       &fl6->daddr, 0, &fl6->saddr)) {
+			reason = SKB_DROP_REASON_IP_OUTNOROUTES;
 			goto tx_err_link_failure;
+		}
 		ndst = dst;
 	}
 
@@ -1223,6 +1240,7 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 		DEV_STATS_INC(dev, collisions);
 		net_warn_ratelimited("%s: Local routing loop detected!\n",
 				     t->parms.name);
+		reason = SKB_DROP_REASON_RECURSION_LIMIT;
 		goto tx_err_dst_release;
 	}
 	mtu = dst6_mtu(dst) - eth_hlen - psh_hlen - t->tun_hlen;
@@ -1236,7 +1254,7 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 	skb_dst_update_pmtu_no_confirm(skb, mtu);
 	if (skb->len - t->tun_hlen - eth_hlen > mtu && !skb_is_gso(skb)) {
 		*pmtu = mtu;
-		err = -EMSGSIZE;
+		reason = SKB_DROP_REASON_PKT_TOO_BIG;
 		goto tx_err_dst_release;
 	}
 
@@ -1259,12 +1277,16 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 	 */
 	max_headroom += LL_RESERVED_SPACE(tdev);
 
-	if (skb_cow_head(skb, max_headroom))
+	if (skb_cow_head(skb, max_headroom)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto tx_err_dst_release;
+	}
 
 	if (t->parms.collect_md) {
-		if (t->encap.type != TUNNEL_ENCAP_NONE)
+		if (t->encap.type != TUNNEL_ENCAP_NONE) {
+			reason = SKB_DROP_REASON_TUNNEL_ENCAP;
 			goto tx_err_dst_release;
+		}
 	} else {
 		if (use_cache && ndst)
 			dst_cache_set_ip6(&t->dst_cache, ndst, &fl6->saddr);
@@ -1287,9 +1309,8 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 			+ dst->header_len + t->hlen;
 	ip_tunnel_adj_headroom(dev, max_headroom);
 
-	err = ip6_tnl_encap(skb, t, &proto, fl6);
-	if (err)
-		return err;
+	if (ip6_tnl_encap(skb, t, &proto, fl6))
+		return SKB_DROP_REASON_TUNNEL_ENCAP;
 
 	if (encap_limit >= 0) {
 		init_tel_txopt(&opt, encap_limit);
@@ -1306,21 +1327,22 @@ int ip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev, __u8 dsfield,
 	ipv6h->saddr = fl6->saddr;
 	ipv6h->daddr = fl6->daddr;
 	ip6tunnel_xmit(NULL, skb, dev, 0);
-	return 0;
+	return SKB_NOT_DROPPED_YET;
 tx_err_link_failure:
 	DEV_STATS_INC(dev, tx_carrier_errors);
 	dst_link_failure(skb);
 tx_err_dst_release:
 	dst_release(dst);
-	return err;
+	return reason;
 }
 EXPORT_SYMBOL(ip6_tnl_xmit);
 
-static inline int
+static inline enum skb_drop_reason
 ipxip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev,
 		u8 protocol)
 {
 	struct ip6_tnl *t = netdev_priv(dev);
+	enum skb_drop_reason reason;
 	struct ipv6hdr *ipv6h;
 	const struct iphdr  *iph;
 	int encap_limit = -1;
@@ -1329,11 +1351,10 @@ ipxip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev,
 	__u8 dsfield, orig_dsfield;
 	__u32 mtu;
 	u8 tproto;
-	int err;
 
 	tproto = READ_ONCE(t->parms.proto);
 	if (tproto != protocol && tproto != 0)
-		return -1;
+		return SKB_DROP_REASON_UNHANDLED_PROTO;
 
 	if (t->parms.collect_md) {
 		struct ip_tunnel_info *tun_info;
@@ -1342,7 +1363,7 @@ ipxip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev,
 		tun_info = skb_tunnel_info(skb);
 		if (unlikely(!tun_info || !(tun_info->mode & IP_TUNNEL_INFO_TX) ||
 			     ip_tunnel_info_af(tun_info) != AF_INET6))
-			return -1;
+			return SKB_DROP_REASON_TUNNEL_TXINFO;
 		key = &tun_info->key;
 		memset(&fl6, 0, sizeof(fl6));
 		fl6.flowi6_proto = protocol;
@@ -1379,7 +1400,7 @@ ipxip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev,
 				if (tel->encap_limit == 0) {
 					icmpv6_ndo_send(skb, ICMPV6_PARAMPROB,
 							ICMPV6_HDR_FIELD, offset + 2);
-					return -1;
+					return SKB_DROP_REASON_IPV6_BAD_EXTHDR;
 				}
 				encap_limit = tel->encap_limit - 1;
 			}
@@ -1421,15 +1442,15 @@ ipxip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev,
 	dsfield = INET_ECN_encapsulate(dsfield, orig_dsfield);
 
 	if (iptunnel_handle_offloads(skb, SKB_GSO_IPXIP6))
-		return -1;
+		return SKB_DROP_REASON_NOMEM;
 
 	skb_set_inner_ipproto(skb, protocol);
 
-	err = ip6_tnl_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu,
-			   protocol);
-	if (err != 0) {
+	reason = ip6_tnl_xmit(skb, dev, dsfield, &fl6, encap_limit, &mtu,
+			      protocol);
+	if (reason) {
 		/* XXX: send ICMP error even if DF is not set. */
-		if (err == -EMSGSIZE)
+		if (reason == SKB_DROP_REASON_PKT_TOO_BIG)
 			switch (protocol) {
 			case IPPROTO_IPIP:
 				icmp_ndo_send(skb, ICMP_DEST_UNREACH,
@@ -1441,20 +1462,21 @@ ipxip6_tnl_xmit(struct sk_buff *skb, struct net_device *dev,
 			default:
 				break;
 			}
-		return -1;
+		return reason;
 	}
 
-	return 0;
+	return SKB_NOT_DROPPED_YET;
 }
 
 static netdev_tx_t
 ip6_tnl_start_xmit(struct sk_buff *skb, struct net_device *dev)
 {
 	struct ip6_tnl *t = netdev_priv(dev);
+	enum skb_drop_reason reason;
 	u8 ipproto;
-	int ret;
 
-	if (!pskb_inet_may_pull(skb))
+	reason = pskb_inet_may_pull_reason(skb);
+	if (reason)
 		goto tx_err;
 
 	switch (skb->protocol) {
@@ -1462,27 +1484,31 @@ ip6_tnl_start_xmit(struct sk_buff *skb, struct net_device *dev)
 		ipproto = IPPROTO_IPIP;
 		break;
 	case htons(ETH_P_IPV6):
-		if (ip6_tnl_addr_conflict(t, ipv6_hdr(skb)))
+		if (ip6_tnl_addr_conflict(t, ipv6_hdr(skb))) {
+			reason = SKB_DROP_REASON_RECURSION_LIMIT;
 			goto tx_err;
+		}
 		ipproto = IPPROTO_IPV6;
 		break;
 	case htons(ETH_P_MPLS_UC):
 		ipproto = IPPROTO_MPLS;
 		break;
 	default:
+		reason = SKB_DROP_REASON_UNHANDLED_PROTO;
 		goto tx_err;
 	}
 
-	ret = ipxip6_tnl_xmit(skb, dev, ipproto);
-	if (ret < 0)
+	reason = ipxip6_tnl_xmit(skb, dev, ipproto);
+	if (reason)
 		goto tx_err;
 
 	return NETDEV_TX_OK;
 
 tx_err:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	DEV_STATS_INC(dev, tx_errors);
 	DEV_STATS_INC(dev, tx_dropped);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return NETDEV_TX_OK;
 }
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 13/14] ip6_gre: add drop reasons to the transmit path
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (11 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 12/14] ip6_tunnel: make ip6_tnl_xmit() return a drop reason Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  2026-09-30 18:39 ` [PATCH net-next v5 14/14] vxlan: report a circular route as SKB_DROP_REASON_RECURSION_LIMIT Anton Danilov
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

The previous patch made the ip6_gre helpers return the drop reason of
ip6_tnl_xmit() and their own, but ip6gre_tunnel_xmit() and
ip6erspan_tunnel_xmit() still free the packet with a plain kfree_skb().
Free it with the reason instead, and give the drops of the handlers
themselves one, from the same set:

 - the length reason from pskb_inet_may_pull_reason(), as in ip_gre,
 - SKB_DROP_REASON_NO_TX_TARGET for a device without a remote address
   that is not collect_md, every packet of which ip6_tnl_xmit_ctl()
   refuses, as the handlers pass it the addresses of the device;
   ip6_tnl_xmit() reports an NBMA tunnel that found no endpoint the
   same way,
 - SKB_DROP_REASON_DEV_READY for the other refusals of
   ip6_tnl_xmit_ctl(), which, as the previous patch explains, include a
   routing loop,
 - SKB_DROP_REASON_NOMEM for the offload setup, the trim of a packet
   longer than the MTU and the headroom,
 - SKB_DROP_REASON_TUNNEL_TXINFO for the collect_md metadata checks of
   ip6erspan, and for a packet without IPv6 metadata in ip6gre (see
   below),
 - SKB_DROP_REASON_UNHANDLED_PROTO for an ERSPAN version that is not
   implemented,
 - SKB_DROP_REASON_RECURSION_LIMIT and SKB_DROP_REASON_IPV6_BAD_EXTHDR
   for the checks that ip6erspan_tunnel_xmit() does itself, the same as
   in ip6gre_xmit_ipv6().

ip6gre_tunnel_xmit() looks up the metadata of a collect_md device before
it passes the packet to a helper, but leaves the drop of a packet
without usable metadata, missing or not of the IPv6 family, to the
helpers, and ip6gre_xmit_ipv6() checks the source address first. The
raddr of such a device is ::, so such a packet sent from ::, like the
device's own DAD probe, would be reported as
SKB_DROP_REASON_RECURSION_LIMIT, and other such packets as
SKB_DROP_REASON_TUNNEL_TXINFO. Drop it right after the lookup instead,
as SKB_DROP_REASON_TUNNEL_TXINFO, with the test ip6erspan_tunnel_xmit()
uses; the IPv4 handlers also test presence and family together. The
helpers dropped such a packet before ip6_tnl_xmit() in any case, and
tx_err counts it by the result of the lookup, so only the reason
changes; the test in __gre6_xmit() stays as a safeguard. A packet from
:: that carries IPv6 metadata still meets the check in
ip6gre_xmit_ipv6().

Such an ip6gre device is an NBMA tunnel: ip6gre_header() puts the
endpoint of the packet in its outer header, and __gre6_xmit() would
take it from there, but ip6gre_tunnel_xmit() refuses the packet before
that, so the device sends nothing. This patch does not change that; it
reports these drops like those of an ip6gretap or ip6erspan device
without a remote address, which has no endpoint at all.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 net/ipv6/ip6_gre.c | 73 ++++++++++++++++++++++++++++++++++++----------
 1 file changed, 58 insertions(+), 15 deletions(-)

diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c
index 9e94ea6b6c20..dbca78230815 100644
--- a/net/ipv6/ip6_gre.c
+++ b/net/ipv6/ip6_gre.c
@@ -889,14 +889,28 @@ static netdev_tx_t ip6gre_tunnel_xmit(struct sk_buff *skb,
 	enum skb_drop_reason reason;
 	__be16 payload_protocol;
 
-	if (!pskb_inet_may_pull(skb))
+	reason = pskb_inet_may_pull_reason(skb);
+	if (reason)
 		goto tx_err;
 
-	if (!ip6_tnl_xmit_ctl(t, &t->parms.laddr, &t->parms.raddr))
+	if (!t->parms.collect_md && ipv6_addr_any(&t->parms.raddr)) {
+		reason = SKB_DROP_REASON_NO_TX_TARGET;
 		goto tx_err;
+	}
 
-	if (t->parms.collect_md)
+	if (!ip6_tnl_xmit_ctl(t, &t->parms.laddr, &t->parms.raddr)) {
+		reason = SKB_DROP_REASON_DEV_READY;
+		goto tx_err;
+	}
+
+	if (t->parms.collect_md) {
 		tun_info = skb_tunnel_info_txcheck(skb);
+		if (IS_ERR(tun_info) ||
+		    unlikely(ip_tunnel_info_af(tun_info) != AF_INET6)) {
+			reason = SKB_DROP_REASON_TUNNEL_TXINFO;
+			goto tx_err;
+		}
+	}
 
 	payload_protocol = skb_protocol(skb, true);
 	switch (payload_protocol) {
@@ -917,10 +931,11 @@ static netdev_tx_t ip6gre_tunnel_xmit(struct sk_buff *skb,
 	return NETDEV_TX_OK;
 
 tx_err:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	if (!IS_ERR(tun_info))
 		DEV_STATS_INC(dev, tx_errors);
 	DEV_STATS_INC(dev, tx_dropped);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return NETDEV_TX_OK;
 }
 
@@ -940,18 +955,30 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 	__u32 mtu;
 	int nhoff;
 
-	if (!pskb_inet_may_pull(skb))
+	reason = pskb_inet_may_pull_reason(skb);
+	if (reason)
 		goto tx_err;
 
-	if (!ip6_tnl_xmit_ctl(t, &t->parms.laddr, &t->parms.raddr))
+	if (!t->parms.collect_md && ipv6_addr_any(&t->parms.raddr)) {
+		reason = SKB_DROP_REASON_NO_TX_TARGET;
 		goto tx_err;
+	}
 
-	if (gre_handle_offloads(skb, false))
+	if (!ip6_tnl_xmit_ctl(t, &t->parms.laddr, &t->parms.raddr)) {
+		reason = SKB_DROP_REASON_DEV_READY;
 		goto tx_err;
+	}
+
+	if (gre_handle_offloads(skb, false)) {
+		reason = SKB_DROP_REASON_NOMEM;
+		goto tx_err;
+	}
 
 	if (skb->len > dev->mtu + dev->hard_header_len) {
-		if (pskb_trim(skb, dev->mtu + dev->hard_header_len))
+		if (pskb_trim(skb, dev->mtu + dev->hard_header_len)) {
+			reason = SKB_DROP_REASON_NOMEM;
 			goto tx_err;
+		}
 		truncate = true;
 	}
 
@@ -971,8 +998,10 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 			truncate = true;
 	}
 
-	if (skb_cow_head(skb, dev->needed_headroom ?: t->hlen))
+	if (skb_cow_head(skb, dev->needed_headroom ?: t->hlen)) {
+		reason = SKB_DROP_REASON_NOMEM;
 		goto tx_err;
+	}
 
 	IPCB(skb)->flags = 0;
 
@@ -986,8 +1015,10 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 
 		tun_info = skb_tunnel_info_txcheck(skb);
 		if (IS_ERR(tun_info) ||
-		    unlikely(ip_tunnel_info_af(tun_info) != AF_INET6))
+		    unlikely(ip_tunnel_info_af(tun_info) != AF_INET6)) {
+			reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 			goto tx_err;
+		}
 
 		key = &tun_info->key;
 		memset(&fl6, 0, sizeof(fl6));
@@ -999,10 +1030,14 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 
 		dsfield = key->tos;
 		if (!test_bit(IP_TUNNEL_ERSPAN_OPT_BIT,
-			      tun_info->key.tun_flags))
+			      tun_info->key.tun_flags)) {
+			reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 			goto tx_err;
-		if (tun_info->options_len < sizeof(*md))
+		}
+		if (tun_info->options_len < sizeof(*md)) {
+			reason = SKB_DROP_REASON_TUNNEL_TXINFO;
 			goto tx_err;
+		}
 		md = ip_tunnel_info_opts(tun_info);
 
 		tun_id = tunnel_id_to_key32(key->tun_id);
@@ -1020,6 +1055,7 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 					       truncate, false);
 			proto = htons(ETH_P_ERSPAN2);
 		} else {
+			reason = SKB_DROP_REASON_UNHANDLED_PROTO;
 			goto tx_err;
 		}
 	} else {
@@ -1030,11 +1066,16 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 						 &dsfield, &encap_limit);
 			break;
 		case htons(ETH_P_IPV6):
-			if (ipv6_addr_equal(&t->parms.raddr, &ipv6_hdr(skb)->saddr))
+			if (ipv6_addr_equal(&t->parms.raddr,
+					    &ipv6_hdr(skb)->saddr)) {
+				reason = SKB_DROP_REASON_RECURSION_LIMIT;
 				goto tx_err;
+			}
 			if (prepare_ip6gre_xmit_ipv6(skb, dev, &fl6,
-						     &dsfield, &encap_limit))
+						     &dsfield, &encap_limit)) {
+				reason = SKB_DROP_REASON_IPV6_BAD_EXTHDR;
 				goto tx_err;
+			}
 			break;
 		default:
 			memcpy(&fl6, &t->fl.u.ip6, sizeof(fl6));
@@ -1053,6 +1094,7 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 					       truncate, false);
 			proto = htons(ETH_P_ERSPAN2);
 		} else {
+			reason = SKB_DROP_REASON_UNHANDLED_PROTO;
 			goto tx_err;
 		}
 
@@ -1087,10 +1129,11 @@ static netdev_tx_t ip6erspan_tunnel_xmit(struct sk_buff *skb,
 	return NETDEV_TX_OK;
 
 tx_err:
+	reason = reason ?: SKB_DROP_REASON_NOT_SPECIFIED;
 	if (!IS_ERR(tun_info))
 		DEV_STATS_INC(dev, tx_errors);
 	DEV_STATS_INC(dev, tx_dropped);
-	kfree_skb(skb);
+	kfree_skb_reason(skb, reason);
 	return NETDEV_TX_OK;
 }
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [PATCH net-next v5 14/14] vxlan: report a circular route as SKB_DROP_REASON_RECURSION_LIMIT
  2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
                   ` (12 preceding siblings ...)
  2026-09-30 18:39 ` [PATCH net-next v5 13/14] ip6_gre: add drop reasons to the transmit path Anton Danilov
@ 2026-09-30 18:39 ` Anton Danilov
  13 siblings, 0 replies; 15+ messages in thread
From: Anton Danilov @ 2026-09-30 18:39 UTC (permalink / raw)
  To: netdev
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, David Ahern, Ido Schimmel, Andrew Lunn,
	linux-kernel

udp_tunnel_dst_lookup() and udp_tunnel6_dst_lookup() return -ELOOP
when the route to the remote end goes out of the vxlan device itself,
and vxlan_xmit_one() counts that in collisions, apart from the
-ENETUNREACH of a failed lookup. The drop reason does not tell them
apart: both are reported as SKB_DROP_REASON_IP_OUTNOROUTES, which
describes a failed route lookup.

ip_tunnel and ip6_tunnel make the same check, rt->dst.dev == dev in
ip_tunnel_xmit() and ip_md_tunnel_xmit() and tdev == dev in
ip6_tnl_xmit(), and report it as SKB_DROP_REASON_RECURSION_LIMIT since
earlier in this series. Report the vxlan case the same way, so that
vxlan and these tunnels name this misconfiguration with one reason.

Assisted-by: LLM
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
---
 drivers/net/vxlan/vxlan_core.c | 10 ++++++++--
 1 file changed, 8 insertions(+), 2 deletions(-)

diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index 0d98e29e2d72..7c94a630d184 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -2547,7 +2547,10 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 					   tos, use_cache ? dst_cache : NULL);
 		if (IS_ERR(rt)) {
 			err = PTR_ERR(rt);
-			reason = SKB_DROP_REASON_IP_OUTNOROUTES;
+			if (err == -ELOOP)
+				reason = SKB_DROP_REASON_RECURSION_LIMIT;
+			else
+				reason = SKB_DROP_REASON_IP_OUTNOROUTES;
 			goto tx_error;
 		}
 
@@ -2633,7 +2636,10 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		if (IS_ERR(ndst)) {
 			err = PTR_ERR(ndst);
 			ndst = NULL;
-			reason = SKB_DROP_REASON_IP_OUTNOROUTES;
+			if (err == -ELOOP)
+				reason = SKB_DROP_REASON_RECURSION_LIMIT;
+			else
+				reason = SKB_DROP_REASON_IP_OUTNOROUTES;
 			goto tx_error;
 		}
 
-- 
2.47.3


^ permalink raw reply	[flat|nested] 15+ messages in thread

end of thread, other threads:[~2026-09-30 18:39 UTC | newest]

Thread overview: 15+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 18:38 [PATCH net-next v5 00/14] tunnels: add core and gre drop reasons Anton Danilov
2026-09-30 18:38 ` [PATCH net-next v5 01/14] vxlan: rename the drop reasons for use by other tunnels Anton Danilov
2026-09-30 18:38 ` [PATCH net-next v5 02/14] ip_tunnel: make __iptunnel_pull_header() return a drop reason Anton Danilov
2026-09-30 18:38 ` [PATCH net-next v5 03/14] vxlan: report the drop reason of __iptunnel_pull_header() Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 04/14] ip_tunnel: add drop reasons to the generic RX path Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 05/14] ip6_tunnel: add drop reasons to the receive path Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 06/14] gre: make gre_parse_header() report a drop reason Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 07/14] ip_gre: add drop reasons to the RX path Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 08/14] ip6_gre: " Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 09/14] ip_tunnel: add drop reasons to the transmit path Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 10/14] ip_gre: " Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 11/14] ip6_gre: make prepare_ip6gre_xmit_other() void Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 12/14] ip6_tunnel: make ip6_tnl_xmit() return a drop reason Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 13/14] ip6_gre: add drop reasons to the transmit path Anton Danilov
2026-09-30 18:39 ` [PATCH net-next v5 14/14] vxlan: report a circular route as SKB_DROP_REASON_RECURSION_LIMIT Anton Danilov

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®