mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net v4 0/2] net: prevent partial checksums from modifying IPv4 headers
@ 2026-09-20  0:47 Paulos Yibelo
  2026-09-20  0:47 ` [PATCH net v4 1/2] net: validate virtio checksum start after network header Paulos Yibelo
                   ` (2 more replies)
  0 siblings, 3 replies; 8+ messages in thread
From: Paulos Yibelo @ 2026-09-20  0:47 UTC (permalink / raw)
  To: netdev
  Cc: mst, jasowangio, eperezma, xuanzhuo, virtualization, dsahern,
	idosch, davem, edumazet, kuba, pabeni, horms, willemb, hannes,
	linux-kernel

A TUN or virtio-net user can supply CHECKSUM_PARTIAL metadata whose
checksum start resolves inside the IPv4 header after link-layer headers
are removed. On the IPv4 fragmentation path, skb_checksum_help() may then
modify an IHL which was already parsed and validated. ip_do_fragment()
subsequently trusts the changed IHL and can copy beyond the skb's logical
linear head into emitted IPv4 options.

Patch 1 validates the checksum start relative to skb_network_header().
Patch 2 independently validates and retains the IPv4 header length before
checksum completion.

The issue and this series were reviewed privately. The source reproducer
and complete runtime evidence remain available privately.

Validation included strict checkpatch, focused W=1 builds, a complete
build, two test boots, a legitimate fragmented CHECKSUM_PARTIAL control,
and both forged cases.

Paulos Yibelo (2):
  net: validate virtio checksum start after network header
  ipv4: reject partial checksums covering the IP header

 include/linux/virtio_net.h |  5 +++--
 net/ipv4/ip_output.c       | 23 +++++++++++++++++------
 2 files changed, 20 insertions(+), 8 deletions(-)

base-commit: 9d565b6b72fe3f41fd43636e143072848105189f

^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH net v4 1/2] net: validate virtio checksum start after network header
  2026-09-20  0:47 [PATCH net v4 0/2] net: prevent partial checksums from modifying IPv4 headers Paulos Yibelo
@ 2026-09-20  0:47 ` Paulos Yibelo
  2026-09-20  1:11   ` David Ahern
  2026-09-20  0:47 ` [PATCH net v4 2/2] ipv4: reject partial checksums covering the IP header Paulos Yibelo
  2026-09-21  2:53 ` [PATCH net v5 0/2] net: prevent partial checksums from modifying network headers Paulos Yibelo
  2 siblings, 1 reply; 8+ messages in thread
From: Paulos Yibelo @ 2026-09-20  0:47 UTC (permalink / raw)
  To: netdev
  Cc: mst, jasowangio, eperezma, xuanzhuo, virtualization, dsahern,
	idosch, davem, edumazet, kuba, pabeni, horms, willemb, hannes,
	linux-kernel

__virtio_net_hdr_to_skb() rejects a CHECKSUM_PARTIAL start smaller than an
estimated minimum network-header length. The comparison currently uses the
offset from skb->data rather than the offset from skb_network_header().

For an AF_PACKET frame, skb->data can still point at the Ethernet header
while skb_network_header() points past nested link-layer headers. A
checksum start at the network header can therefore pass, then target byte
zero after those headers are removed.

This does not require a virtual-machine guest. A TUN device with
virtio-net header support can supply the same checksum metadata.

Keep the existing data-relative lower bound and also require checksum
start to follow the estimated minimum relative to skb_network_header().

Fixes: 49d14b54a527 ("net: test for not too small csum_start in virtio_net_hdr_to_skb()")
Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Paulos Yibelo <habte.yibelo@gmail.com>
Acked-by: Michael S. Tsirkin <mst@redhat.com>
---
Changes in v4:
- State explicitly that a TUN device is sufficient and no guest is required,
  as noted by Michael S. Tsirkin. No code changes.

Changes in v3:
- Keep the network-relative comparison on one line for readability, as
  requested by David Ahern.

Changes in v2:
- Make nh_min_len an int and remove the casts, as suggested by Michael S.
  Tsirkin.

 include/linux/virtio_net.h | 5 +++--
 1 file changed, 3 insertions(+), 2 deletions(-)

diff --git a/include/linux/virtio_net.h b/include/linux/virtio_net.h
index c381b91..a95ad46 100644
--- a/include/linux/virtio_net.h
+++ b/include/linux/virtio_net.h
@@ -52,7 +52,7 @@ static inline int __virtio_net_hdr_to_skb(struct sk_buff *skb,
 					  const struct virtio_net_hdr *hdr,
 					  bool little_endian, u8 hdr_gso_type)
 {
-	unsigned int nh_min_len = sizeof(struct iphdr);
+	int nh_min_len = sizeof(struct iphdr);
 	unsigned int gso_type = 0;
 	unsigned int thlen = 0;
 	unsigned int p_off = 0;
@@ -104,7 +104,8 @@ static inline int __virtio_net_hdr_to_skb(struct sk_buff *skb,
 
 		if (!skb_partial_csum_set(skb, start, off))
 			return -EINVAL;
-		if (skb_transport_offset(skb) < nh_min_len)
+		if (skb_transport_offset(skb) < nh_min_len ||
+		    skb_transport_offset(skb) - skb_network_offset(skb) < nh_min_len)
 			return -EINVAL;
 
 		nh_min_len = skb_transport_offset(skb);

^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH net v4 2/2] ipv4: reject partial checksums covering the IP header
  2026-09-20  0:47 [PATCH net v4 0/2] net: prevent partial checksums from modifying IPv4 headers Paulos Yibelo
  2026-09-20  0:47 ` [PATCH net v4 1/2] net: validate virtio checksum start after network header Paulos Yibelo
@ 2026-09-20  0:47 ` Paulos Yibelo
  2026-09-20  1:12   ` David Ahern
  2026-09-21  2:53 ` [PATCH net v5 0/2] net: prevent partial checksums from modifying network headers Paulos Yibelo
  2 siblings, 1 reply; 8+ messages in thread
From: Paulos Yibelo @ 2026-09-20  0:47 UTC (permalink / raw)
  To: netdev
  Cc: mst, jasowangio, eperezma, xuanzhuo, virtualization, dsahern,
	idosch, davem, edumazet, kuba, pabeni, horms, willemb, hannes,
	linux-kernel

ip_do_fragment() completes a CHECKSUM_PARTIAL skb before reading the IPv4
header length. A virtualization interface can supply a checksum start that
still points inside the IPv4 header after link-layer removal.

This does not require a virtual-machine guest. A TUN device with
virtio-net header support is sufficient to reach this path.

skb_checksum_help() can then change iph->ihl after the packet was parsed
and routed. Fragmentation trusts the changed IHL and can copy beyond the
skb's logical linear head into transmitted IPv4 options.

Read and validate IHL before checksum completion, reject a checksum start
inside that header, retain the validated length, and reacquire iph after
skb_checksum_help().

Fixes: dbd3393c56a8 ("ipv4: add defensive check for CHECKSUM_PARTIAL skbs in ip_fragment")
Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Paulos Yibelo <habte.yibelo@gmail.com>
Acked-by: Michael S. Tsirkin <mst@redhat.com>
---
Changes in v4:
- State explicitly that a TUN device is sufficient and no guest is required,
  as noted by Michael S. Tsirkin. No code changes.

Changes in v3:
- No code changes.

Changes in v2:
- No code changes.

 net/ipv4/ip_output.c | 23 +++++++++++++++++------
 1 file changed, 17 insertions(+), 6 deletions(-)

diff --git a/net/ipv4/ip_output.c b/net/ipv4/ip_output.c
index a24cc8e..ff902a2 100644
--- a/net/ipv4/ip_output.c
+++ b/net/ipv4/ip_output.c
@@ -770,17 +770,29 @@ int ip_do_fragment(struct net *net, struct sock *sk, struct sk_buff *skb,
 	struct ip_frag_state state;
 	int err = 0;
 
-	/* for offloaded checksums cleanup checksum before fragmentation */
-	if (skb->ip_summed == CHECKSUM_PARTIAL &&
-	    (err = skb_checksum_help(skb)))
-		goto fail;
-
 	/*
 	 *	Point into the IP datagram header.
 	 */
 
 	iph = ip_hdr(skb);
+	hlen = iph->ihl * 4;
+	if (unlikely(hlen < sizeof(*iph) || hlen > skb_headlen(skb))) {
+		err = -EINVAL;
+		goto fail;
+	}
 
+	/* Complete offloaded checksums only after the validated IP header. */
+	if (skb->ip_summed == CHECKSUM_PARTIAL) {
+		if (unlikely(skb_checksum_start_offset(skb) < hlen)) {
+			err = -EINVAL;
+			goto fail;
+		}
+		err = skb_checksum_help(skb);
+		if (err)
+			goto fail;
+		iph = ip_hdr(skb);
+	}
+
 	mtu = ip_skb_dst_mtu(sk, skb);
 	if (IPCB(skb)->frag_max_size && IPCB(skb)->frag_max_size < mtu)
 		mtu = IPCB(skb)->frag_max_size;
@@ -789,7 +801,6 @@ int ip_do_fragment(struct net *net, struct sock *sk, struct sk_buff *skb,
 	 *	Setup starting values.
 	 */
 
-	hlen = iph->ihl * 4;
 	if (mtu < hlen + 8) {
 		err = -EMSGSIZE;
 		goto fail;

^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH net v4 1/2] net: validate virtio checksum start after network header
  2026-09-20  0:47 ` [PATCH net v4 1/2] net: validate virtio checksum start after network header Paulos Yibelo
@ 2026-09-20  1:11   ` David Ahern
  0 siblings, 0 replies; 8+ messages in thread
From: David Ahern @ 2026-09-20  1:11 UTC (permalink / raw)
  To: Paulos Yibelo, netdev
  Cc: mst, jasowangio, eperezma, xuanzhuo, virtualization, idosch,
	davem, edumazet, kuba, pabeni, horms, willemb, hannes,
	linux-kernel

On 9/19/26 6:47 PM, Paulos Yibelo wrote:
> __virtio_net_hdr_to_skb() rejects a CHECKSUM_PARTIAL start smaller than an
> estimated minimum network-header length. The comparison currently uses the
> offset from skb->data rather than the offset from skb_network_header().
> 
> For an AF_PACKET frame, skb->data can still point at the Ethernet header
> while skb_network_header() points past nested link-layer headers. A
> checksum start at the network header can therefore pass, then target byte
> zero after those headers are removed.
> 
> This does not require a virtual-machine guest. A TUN device with
> virtio-net header support can supply the same checksum metadata.
> 
> Keep the existing data-relative lower bound and also require checksum
> start to follow the estimated minimum relative to skb_network_header().
> 
> Fixes: 49d14b54a527 ("net: test for not too small csum_start in virtio_net_hdr_to_skb()")
> Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Paulos Yibelo <habte.yibelo@gmail.com>
> Acked-by: Michael S. Tsirkin <mst@redhat.com>
> ---
> Changes in v4:
> - State explicitly that a TUN device is sufficient and no guest is required,
>   as noted by Michael S. Tsirkin. No code changes.
> 
> Changes in v3:
> - Keep the network-relative comparison on one line for readability, as
>   requested by David Ahern.
> 
> Changes in v2:
> - Make nh_min_len an int and remove the casts, as suggested by Michael S.
>   Tsirkin.
> 
>  include/linux/virtio_net.h | 5 +++--
>  1 file changed, 3 insertions(+), 2 deletions(-)
> 

Reviewed-by: David Ahern <dsahern@kernel.org>



^ permalink raw reply	[flat|nested] 8+ messages in thread

* Re: [PATCH net v4 2/2] ipv4: reject partial checksums covering the IP header
  2026-09-20  0:47 ` [PATCH net v4 2/2] ipv4: reject partial checksums covering the IP header Paulos Yibelo
@ 2026-09-20  1:12   ` David Ahern
  0 siblings, 0 replies; 8+ messages in thread
From: David Ahern @ 2026-09-20  1:12 UTC (permalink / raw)
  To: Paulos Yibelo, netdev
  Cc: mst, jasowangio, eperezma, xuanzhuo, virtualization, idosch,
	davem, edumazet, kuba, pabeni, horms, willemb, hannes,
	linux-kernel

On 9/19/26 6:47 PM, Paulos Yibelo wrote:
> ip_do_fragment() completes a CHECKSUM_PARTIAL skb before reading the IPv4
> header length. A virtualization interface can supply a checksum start that
> still points inside the IPv4 header after link-layer removal.
> 
> This does not require a virtual-machine guest. A TUN device with
> virtio-net header support is sufficient to reach this path.
> 
> skb_checksum_help() can then change iph->ihl after the packet was parsed
> and routed. Fragmentation trusts the changed IHL and can copy beyond the
> skb's logical linear head into transmitted IPv4 options.
> 
> Read and validate IHL before checksum completion, reject a checksum start
> inside that header, retain the validated length, and reacquire iph after
> skb_checksum_help().
> 
> Fixes: dbd3393c56a8 ("ipv4: add defensive check for CHECKSUM_PARTIAL skbs in ip_fragment")
> Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Paulos Yibelo <habte.yibelo@gmail.com>
> Acked-by: Michael S. Tsirkin <mst@redhat.com>
> ---
> Changes in v4:
> - State explicitly that a TUN device is sufficient and no guest is required,
>   as noted by Michael S. Tsirkin. No code changes.
> 
> Changes in v3:
> - No code changes.
> 
> Changes in v2:
> - No code changes.
> 
>  net/ipv4/ip_output.c | 23 +++++++++++++++++------
>  1 file changed, 17 insertions(+), 6 deletions(-)
> 
>

Reviewed-by: David Ahern <dsahern@kernel.org>


^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH net v5 0/2] net: prevent partial checksums from modifying network headers
  2026-09-20  0:47 [PATCH net v4 0/2] net: prevent partial checksums from modifying IPv4 headers Paulos Yibelo
  2026-09-20  0:47 ` [PATCH net v4 1/2] net: validate virtio checksum start after network header Paulos Yibelo
  2026-09-20  0:47 ` [PATCH net v4 2/2] ipv4: reject partial checksums covering the IP header Paulos Yibelo
@ 2026-09-21  2:53 ` Paulos Yibelo
  2026-09-21  2:53   ` [PATCH net v5 1/2] net: validate virtio checksum start after network header Paulos Yibelo
  2026-09-21  2:53   ` [PATCH net v5 2/2] ip: reject partial checksums covering network headers Paulos Yibelo
  2 siblings, 2 replies; 8+ messages in thread
From: Paulos Yibelo @ 2026-09-21  2:53 UTC (permalink / raw)
  To: netdev
  Cc: richard, anton.ivanov, johannes, willemdebruijn.kernel,
	jasowangio, mst, eperezma, xuanzhuo, andrew+netdev, pablo, fw,
	phil, razor, idosch, dsahern, davem, edumazet, kuba, pabeni,
	horms, linux-um, virtualization, netfilter-devel, coreteam,
	bridge, linux-kernel

A virtio-net header can supply CHECKSUM_PARTIAL metadata whose checksum
start resolves inside the network header after link-layer removal.
Software checksum completion can then modify header bytes which the stack
has already parsed.

Patch 1 validates the checksum start against an explicit data-relative L3
origin. It covers TUN/TAP, virtio-net, AF_PACKET, UML, nested VLAN
headers, and tunnel metadata. It does not rely on skb header state which
may not yet be established.

Patch 2 independently validates the checksum start against the parsed
IPv4 or IPv6 header length in all four IP fragmentation implementations
which complete partial checksums.

The v4 Sashiko findings were correct. Patch 1 used
skb_network_offset() before all receive callers had established it.
Patch 2 compared a signed checksum offset with an unsigned IPv4 header
length. This revision fixes both findings and covers the corresponding
bridge and IPv6 fragmentation paths.

Validation included strict checkpatch, focused x86 and UML W=1 builds,
an offset-boundary model, and application of the exact mail series to the
stated base.

Changes in v5:
- Pass an explicit data-relative L3 origin through the virtio-net
  converter and audit every in-tree caller.
- Parse Ethernet and nested VLAN headers without mutating skb header
  state.
- Propagate virtio-header conversion failures in UML.
- Keep the IPv4 comparison signed and add matching parsed-header checks
  to the IPv4/IPv6 output and bridge-netfilter fragmentation paths.
- Drop Michael S. Tsirkin's Acked-by and David Ahern's Reviewed-by tags
  because both patches changed materially.

Link: https://lore.kernel.org/netdev/20260920004733.6473-1-habte.yibelo@gmail.com/

Paulos Yibelo (2):
  net: validate virtio checksum start after network header
  ip: reject partial checksums covering network headers

 arch/um/drivers/vector_transports.c        | 10 ++-
 drivers/net/tun_vnet.h                     | 28 +++++++-
 drivers/net/virtio_net.c                   |  8 ++-
 include/linux/virtio_net.h                 | 76 ++++++++++++++++++----
 net/bridge/netfilter/nf_conntrack_bridge.c | 21 ++++--
 net/ipv4/ip_output.c                       | 23 +++++--
 net/ipv6/ip6_output.c                      | 12 +++-
 net/ipv6/netfilter.c                       | 12 +++-
 net/packet/af_packet.c                     |  6 +-
 9 files changed, 157 insertions(+), 39 deletions(-)


base-commit: 1e24c4f2ee44be0eee94092b5d13cbdb4bdf0d60
-- 
2.46.0

^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH net v5 1/2] net: validate virtio checksum start after network header
  2026-09-21  2:53 ` [PATCH net v5 0/2] net: prevent partial checksums from modifying network headers Paulos Yibelo
@ 2026-09-21  2:53   ` Paulos Yibelo
  2026-09-21  2:53   ` [PATCH net v5 2/2] ip: reject partial checksums covering network headers Paulos Yibelo
  1 sibling, 0 replies; 8+ messages in thread
From: Paulos Yibelo @ 2026-09-21  2:53 UTC (permalink / raw)
  To: netdev
  Cc: richard, anton.ivanov, johannes, willemdebruijn.kernel,
	jasowangio, mst, eperezma, xuanzhuo, andrew+netdev, pablo, fw,
	phil, razor, idosch, dsahern, davem, edumazet, kuba, pabeni,
	horms, linux-um, virtualization, netfilter-devel, coreteam,
	bridge, linux-kernel

__virtio_net_hdr_to_skb() rejects a CHECKSUM_PARTIAL start smaller than
an estimated minimum network-header length. Its input offsets are relative
to skb->data.

Using skb_network_offset() here is unsafe. TUN/TAP, virtio-net, and UML
parse a received virtio header before skb->network_header is established.
On an skb with headroom, the resulting negative offset enlarges the
apparent distance to the transport header and can admit a checksum start
inside the network header.

Pass the data-relative L3 offset to the converter explicitly. IFF_TUN uses
zero, AF_PACKET supplies its established network offset, and Ethernet
receive paths parse Ethernet and nested VLAN headers with
skb_header_pointer(), without changing skb state. Use the same origin for
tunnel-offset validation, and make UML propagate conversion failures.

This does not require a virtual-machine guest. A TUN or TAP device with
virtio-net header support is sufficient to reach these paths.

Fixes: 49d14b54a527 ("net: test for not too small csum_start in virtio_net_hdr_to_skb()")
Fixes: a2fb4bc4e2a6 ("net: implement virtio helpers to handle UDP GSO tunneling.")
Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
Link: https://lore.kernel.org/netdev/20260920004733.6473-2-habte.yibelo@gmail.com/
Cc: stable@vger.kernel.org
Signed-off-by: Paulos Yibelo <habte.yibelo@gmail.com>
---
Changes in v5:
- Replace the not-yet-established skb network-header offset with an
  explicit data-relative L3 origin.
- Cover all in-tree callers, including Ethernet/VLAN receive paths,
  tunnel metadata, and UML error propagation.
- Drop the prior Acked-by and Reviewed-by tags because the code changed.

Changes in v4:
- State that a TUN device is sufficient and no guest is required, as
  noted by Michael S. Tsirkin.

Changes in v3:
- Keep the network-relative comparison on one line for readability, as
  requested by David Ahern.

Changes in v2:
- Make nh_min_len an int and remove the casts, as suggested by Michael S.
  Tsirkin.

 arch/um/drivers/vector_transports.c | 10 +++-
 drivers/net/tun_vnet.h              | 28 ++++++++++-
 drivers/net/virtio_net.c            |  8 ++-
 include/linux/virtio_net.h          | 76 +++++++++++++++++++++++------
 net/packet/af_packet.c              |  6 ++-
 5 files changed, 106 insertions(+), 22 deletions(-)

diff --git a/arch/um/drivers/vector_transports.c b/arch/um/drivers/vector_transports.c
index ddd127ee9..79bc05fc6 100644
--- a/arch/um/drivers/vector_transports.c
+++ b/arch/um/drivers/vector_transports.c
@@ -197,6 +197,7 @@ static int raw_verify_header(
 	uint8_t *header, struct sk_buff *skb, struct vector_private *vp)
 {
 	struct virtio_net_hdr *vheader = (struct virtio_net_hdr *) header;
+	int network_offset;
 
 	if ((vheader->gso_type != VIRTIO_NET_HDR_GSO_NONE) &&
 		(vp->req_size != 65536)) {
@@ -209,8 +210,13 @@ static int raw_verify_header(
 	if ((vheader->flags & VIRTIO_NET_HDR_F_DATA_VALID) > 0)
 		return 1;
 
-	virtio_net_hdr_to_skb(skb, vheader, virtio_legacy_is_little_endian());
-	return 0;
+	network_offset = virtio_net_hdr_get_l3_offset(skb, vheader);
+	if (network_offset < 0)
+		return network_offset;
+
+	return virtio_net_hdr_to_skb(skb, vheader,
+				     virtio_legacy_is_little_endian(),
+				     network_offset);
 }
 
 static bool get_uint_param(
diff --git a/drivers/net/tun_vnet.h b/drivers/net/tun_vnet.h
index f4c652b1f..1c83c359d 100644
--- a/drivers/net/tun_vnet.h
+++ b/drivers/net/tun_vnet.h
@@ -177,10 +177,27 @@ static inline int tun_vnet_hdr_put(int sz, struct iov_iter *iter,
 	return __tun_vnet_hdr_put(sz, 0, iter, hdr);
 }
 
+static inline int
+tun_vnet_hdr_get_l3_offset(unsigned int flags, const struct sk_buff *skb,
+			   const struct virtio_net_hdr *hdr)
+{
+	if ((flags & TUN_TYPE_MASK) != IFF_TAP)
+		return 0;
+
+	return virtio_net_hdr_get_l3_offset(skb, hdr);
+}
+
 static inline int tun_vnet_hdr_to_skb(unsigned int flags, struct sk_buff *skb,
 				      const struct virtio_net_hdr *hdr)
 {
-	return virtio_net_hdr_to_skb(skb, hdr, tun_vnet_is_little_endian(flags));
+	int network_offset = tun_vnet_hdr_get_l3_offset(flags, skb, hdr);
+
+	if (network_offset < 0)
+		return network_offset;
+
+	return virtio_net_hdr_to_skb(skb, hdr,
+				     tun_vnet_is_little_endian(flags),
+				     network_offset);
 }
 
 /*
@@ -199,10 +216,17 @@ tun_vnet_hdr_tnl_to_skb(unsigned int flags, netdev_features_t features,
 			struct sk_buff *skb,
 			const struct virtio_net_hdr_v1_hash_tunnel *hdr)
 {
+	const struct virtio_net_hdr *vnet_hdr = (const struct virtio_net_hdr *)hdr;
+	int network_offset = tun_vnet_hdr_get_l3_offset(flags, skb, vnet_hdr);
+
+	if (network_offset < 0)
+		return network_offset;
+
 	return virtio_net_hdr_tnl_to_skb(skb, hdr,
 				features & NETIF_F_GSO_UDP_TUNNEL,
 				features & NETIF_F_GSO_UDP_TUNNEL_CSUM,
-				tun_vnet_is_little_endian(flags));
+				tun_vnet_is_little_endian(flags),
+				network_offset);
 }
 
 static inline int tun_vnet_hdr_from_skb(unsigned int flags,
diff --git a/drivers/net/virtio_net.c b/drivers/net/virtio_net.c
index e34c52d05..059eeb18e 100644
--- a/drivers/net/virtio_net.c
+++ b/drivers/net/virtio_net.c
@@ -2502,6 +2502,7 @@ static void virtnet_receive_done(struct virtnet_info *vi, struct receive_queue *
 {
 	struct virtio_net_common_hdr *hdr;
 	struct net_device *dev = vi->dev;
+	int network_offset;
 
 	hdr = skb_vnet_common_hdr(skb);
 	if (dev->features & NETIF_F_RXHASH && vi->has_rss_hash_report)
@@ -2515,9 +2516,12 @@ static void virtnet_receive_done(struct virtnet_info *vi, struct receive_queue *
 		goto frame_err;
 	}
 
-	if (virtio_net_hdr_tnl_to_skb(skb, &hdr->tnl_hdr, vi->rx_tnl,
+	network_offset = virtio_net_hdr_get_l3_offset(skb, &hdr->hdr);
+	if (network_offset < 0 ||
+	    virtio_net_hdr_tnl_to_skb(skb, &hdr->tnl_hdr, vi->rx_tnl,
 				      vi->rx_tnl_csum,
-				      virtio_is_little_endian(vi->vdev))) {
+				      virtio_is_little_endian(vi->vdev),
+				      network_offset)) {
 		net_warn_ratelimited("%s: bad gso: type: %x, size: %u, flags %x tunnel %d tnl csum %d\n",
 				     dev->name, hdr->hdr.gso_type,
 				     hdr->hdr.gso_size, hdr->hdr.flags,
diff --git a/include/linux/virtio_net.h b/include/linux/virtio_net.h
index c381b916c..a4c005796 100644
--- a/include/linux/virtio_net.h
+++ b/include/linux/virtio_net.h
@@ -48,11 +48,49 @@ static inline int virtio_net_hdr_set_proto(struct sk_buff *skb,
 	return 0;
 }
 
+/*
+ * Return the L3 offset of an Ethernet frame starting at skb->data.
+ * The offset is unused without NEEDS_CSUM, so avoid parsing and return zero.
+ */
+static inline int
+virtio_net_hdr_get_l3_offset(const struct sk_buff *skb,
+			     const struct virtio_net_hdr *hdr)
+{
+	unsigned int parse_depth = VLAN_MAX_DEPTH;
+	const struct ethhdr *eth;
+	struct ethhdr ethbuf;
+	__be16 protocol;
+	int depth = ETH_HLEN;
+
+	if (!(hdr->flags & VIRTIO_NET_HDR_F_NEEDS_CSUM))
+		return 0;
+
+	eth = skb_header_pointer(skb, 0, sizeof(ethbuf), &ethbuf);
+	if (!eth)
+		return -EINVAL;
+
+	protocol = eth->h_proto;
+	while (eth_type_vlan(protocol)) {
+		const struct vlan_hdr *vh;
+		struct vlan_hdr vhdr;
+
+		vh = skb_header_pointer(skb, depth, sizeof(vhdr), &vhdr);
+		if (!vh || !--parse_depth)
+			return -EINVAL;
+
+		protocol = vh->h_vlan_encapsulated_proto;
+		depth += VLAN_HLEN;
+	}
+
+	return depth;
+}
+
 static inline int __virtio_net_hdr_to_skb(struct sk_buff *skb,
 					  const struct virtio_net_hdr *hdr,
-					  bool little_endian, u8 hdr_gso_type)
+					  bool little_endian, u8 hdr_gso_type,
+					  int network_offset)
 {
-	unsigned int nh_min_len = sizeof(struct iphdr);
+	int nh_min_len = sizeof(struct iphdr);
 	unsigned int gso_type = 0;
 	unsigned int thlen = 0;
 	unsigned int p_off = 0;
@@ -98,16 +136,20 @@ static inline int __virtio_net_hdr_to_skb(struct sk_buff *skb,
 		u32 start = __virtio16_to_cpu(little_endian, hdr->csum_start);
 		u32 off = __virtio16_to_cpu(little_endian, hdr->csum_offset);
 		u32 needed = start + max_t(u32, thlen, off + sizeof(__sum16));
+		int transport_offset;
 
 		if (!pskb_may_pull(skb, needed))
 			return -EINVAL;
 
 		if (!skb_partial_csum_set(skb, start, off))
 			return -EINVAL;
-		if (skb_transport_offset(skb) < nh_min_len)
+
+		transport_offset = skb_transport_offset(skb);
+		if (transport_offset < nh_min_len || network_offset < 0 ||
+		    network_offset > transport_offset - nh_min_len)
 			return -EINVAL;
 
-		nh_min_len = skb_transport_offset(skb);
+		nh_min_len = transport_offset;
 		p_off = nh_min_len + thlen;
 		if (!pskb_may_pull(skb, p_off))
 			return -EINVAL;
@@ -206,9 +248,11 @@ static inline int __virtio_net_hdr_to_skb(struct sk_buff *skb,
 
 static inline int virtio_net_hdr_to_skb(struct sk_buff *skb,
 					const struct virtio_net_hdr *hdr,
-					bool little_endian)
+					bool little_endian,
+					int network_offset)
 {
-	return __virtio_net_hdr_to_skb(skb, hdr, little_endian, hdr->gso_type);
+	return __virtio_net_hdr_to_skb(skb, hdr, little_endian, hdr->gso_type,
+				       network_offset);
 }
 
 /* This function must be called after virtio_net_hdr_from_skb(). */
@@ -287,7 +331,7 @@ static inline int virtio_net_hdr_from_skb(const struct sk_buff *skb,
 	return 0;
 }
 
-static inline unsigned int virtio_l3min(bool is_ipv6)
+static inline int virtio_l3min(bool is_ipv6)
 {
 	return is_ipv6 ? sizeof(struct ipv6hdr) : sizeof(struct iphdr);
 }
@@ -297,18 +341,19 @@ virtio_net_hdr_tnl_to_skb(struct sk_buff *skb,
 			  const struct virtio_net_hdr_v1_hash_tunnel *vhdr,
 			  bool tnl_hdr_negotiated,
 			  bool tnl_csum_negotiated,
-			  bool little_endian)
+			  bool little_endian, int network_offset)
 {
 	const struct virtio_net_hdr *hdr = (const struct virtio_net_hdr *)vhdr;
-	unsigned int inner_nh, outer_th, inner_th;
-	unsigned int inner_l3min, outer_l3min;
 	u8 gso_inner_type, gso_tunnel_type;
 	bool outer_isv6, inner_isv6;
+	int inner_nh, outer_th, inner_th;
+	int inner_l3min, outer_l3min;
 	int ret;
 
 	gso_tunnel_type = hdr->gso_type & VIRTIO_NET_HDR_GSO_UDP_TUNNEL;
 	if (!gso_tunnel_type)
-		return virtio_net_hdr_to_skb(skb, hdr, little_endian);
+		return virtio_net_hdr_to_skb(skb, hdr, little_endian,
+					     network_offset);
 
 	/* Tunnel not supported/negotiated, but the hdr asks for it. */
 	if (!tnl_hdr_negotiated)
@@ -332,19 +377,22 @@ virtio_net_hdr_tnl_to_skb(struct sk_buff *skb,
 	outer_isv6 = gso_tunnel_type & VIRTIO_NET_HDR_GSO_UDP_TUNNEL_IPV6;
 	inner_isv6 = gso_inner_type == VIRTIO_NET_HDR_GSO_TCPV6;
 	inner_l3min = virtio_l3min(inner_isv6);
-	outer_l3min = ETH_HLEN + virtio_l3min(outer_isv6);
+	outer_l3min = virtio_l3min(outer_isv6);
 
 	inner_th = __virtio16_to_cpu(little_endian, hdr->csum_start);
 	inner_nh = le16_to_cpu(vhdr->inner_nh_offset);
 	outer_th = le16_to_cpu(vhdr->outer_th_offset);
-	if (outer_th < outer_l3min ||
+	if (network_offset < 0 ||
+	    outer_th < outer_l3min ||
+	    network_offset > outer_th - outer_l3min ||
 	    inner_nh < outer_th + sizeof(struct udphdr) ||
 	    inner_th < inner_nh + inner_l3min)
 		return -EINVAL;
 
 	/* Let the basic parsing deal with plain GSO features. */
 	ret = __virtio_net_hdr_to_skb(skb, hdr, true,
-				      hdr->gso_type & ~gso_tunnel_type);
+				      hdr->gso_type & ~gso_tunnel_type,
+				      network_offset);
 	if (ret)
 		return ret;
 
diff --git a/net/packet/af_packet.c b/net/packet/af_packet.c
index 50cae32ae..04c80e23d 100644
--- a/net/packet/af_packet.c
+++ b/net/packet/af_packet.c
@@ -2901,7 +2901,8 @@ static int tpacket_snd(struct packet_sock *po, struct msghdr *msg)
 		}
 
 		if (has_vnet_hdr) {
-			if (virtio_net_hdr_to_skb(skb, &vnet_hdr, vio_le())) {
+			if (virtio_net_hdr_to_skb(skb, &vnet_hdr, vio_le(),
+						  skb_network_offset(skb))) {
 				tp_len = -EINVAL;
 				goto tpacket_error;
 			}
@@ -3103,7 +3104,8 @@ static int packet_snd(struct socket *sock, struct msghdr *msg, size_t len)
 	packet_parse_headers(skb, sock);
 
 	if (vnet_hdr_sz) {
-		err = virtio_net_hdr_to_skb(skb, &vnet_hdr, vio_le());
+		err = virtio_net_hdr_to_skb(skb, &vnet_hdr, vio_le(),
+					    skb_network_offset(skb));
 		if (err)
 			goto out_free;
 		len += vnet_hdr_sz;
-- 
2.46.0

^ permalink raw reply	[flat|nested] 8+ messages in thread

* [PATCH net v5 2/2] ip: reject partial checksums covering network headers
  2026-09-21  2:53 ` [PATCH net v5 0/2] net: prevent partial checksums from modifying network headers Paulos Yibelo
  2026-09-21  2:53   ` [PATCH net v5 1/2] net: validate virtio checksum start after network header Paulos Yibelo
@ 2026-09-21  2:53   ` Paulos Yibelo
  1 sibling, 0 replies; 8+ messages in thread
From: Paulos Yibelo @ 2026-09-21  2:53 UTC (permalink / raw)
  To: netdev
  Cc: richard, anton.ivanov, johannes, willemdebruijn.kernel,
	jasowangio, mst, eperezma, xuanzhuo, andrew+netdev, pablo, fw,
	phil, razor, idosch, dsahern, davem, edumazet, kuba, pabeni,
	horms, linux-um, virtualization, netfilter-devel, coreteam,
	bridge, linux-kernel

ip_do_fragment() and nf_br_ip_fragment() complete a CHECKSUM_PARTIAL skb
before reading the IPv4 header length. ip6_fragment() and br_ip6_fragment()
complete one after parsing the IPv6 header chain. A virtualization
interface can supply a checksum start which, after link-layer removal,
still points inside that parsed network header.

skb_checksum_help() then writes the completed checksum into header bytes
the stack has already consumed. For IPv4, changing iph->ihl after routing
and validation can make fragmentation copy beyond the skb's logical linear
head into transmitted options. A negative checksum-start offset is rejected
by skb_checksum_help(), but only after a WARN_ONCE which can panic a
panic_on_warn system.

Validate the checksum start against the parsed header length before
completing it. For IPv4, read and validate IHL first, retain it, and
reacquire iph after skb_checksum_help() in both implementations. For IPv6,
use the length returned by ip6_find_1stfragopt() in both implementations.

Compare the signed checksum-start offset with the bounded signed header
length so integer promotion cannot bypass either boundary.

Fixes: dbd3393c56a8 ("ipv4: add defensive check for CHECKSUM_PARTIAL skbs in ip_fragment")
Fixes: 405c92f7a541 ("ipv6: add defensive check for CHECKSUM_PARTIAL skbs in ip_fragment")
Fixes: 3c171f496ef5 ("netfilter: bridge: add connection tracking system")
Fixes: 764dd163ac92 ("netfilter: nf_conntrack_bridge: add support for IPv6")
Reported-by: Paulos Yibelo <habte.yibelo@gmail.com>
Link: https://lore.kernel.org/netdev/20260920004733.6473-3-habte.yibelo@gmail.com/
Cc: stable@vger.kernel.org
Signed-off-by: Paulos Yibelo <habte.yibelo@gmail.com>
---
Changes in v5:
- Compare the checksum-start offset and IPv4 header length as signed
  values.
- Add parsed-header checks to the IPv4/IPv6 output and bridge-netfilter
  fragmentation paths.
- Drop the prior Acked-by and Reviewed-by tags because the code changed.

Changes in v4:
- State that a TUN device is sufficient and no guest is required, as
  noted by Michael S. Tsirkin.

Changes in v3:
- No code changes.

Changes in v2:
- No code changes.

 net/bridge/netfilter/nf_conntrack_bridge.c | 21 +++++++++++++++-----
 net/ipv4/ip_output.c                       | 23 ++++++++++++++++------
 net/ipv6/ip6_output.c                      | 12 ++++++++---
 net/ipv6/netfilter.c                       | 12 ++++++++---
 4 files changed, 51 insertions(+), 17 deletions(-)

diff --git a/net/bridge/netfilter/nf_conntrack_bridge.c b/net/bridge/netfilter/nf_conntrack_bridge.c
index 7ecb8a26b..d81ed8692 100644
--- a/net/bridge/netfilter/nf_conntrack_bridge.c
+++ b/net/bridge/netfilter/nf_conntrack_bridge.c
@@ -38,18 +38,29 @@ static int nf_br_ip_fragment(struct net *net, struct sock *sk,
 	struct iphdr *iph;
 	int err = 0;
 
-	/* for offloaded checksums cleanup checksum before fragmentation */
-	if (skb->ip_summed == CHECKSUM_PARTIAL &&
-	    (err = skb_checksum_help(skb)))
+	iph = ip_hdr(skb);
+	hlen = iph->ihl * 4;
+	if (unlikely(hlen < sizeof(*iph) || hlen > skb_headlen(skb))) {
+		err = -EINVAL;
 		goto blackhole;
+	}
 
-	iph = ip_hdr(skb);
+	/* Complete offloaded checksums only after the validated IP header. */
+	if (skb->ip_summed == CHECKSUM_PARTIAL) {
+		if (unlikely(skb_checksum_start_offset(skb) < (int)hlen)) {
+			err = -EINVAL;
+			goto blackhole;
+		}
+		err = skb_checksum_help(skb);
+		if (err)
+			goto blackhole;
+		iph = ip_hdr(skb);
+	}
 
 	/*
 	 *	Setup starting values
 	 */
 
-	hlen = iph->ihl * 4;
 	frag_max_size -= hlen;
 	ll_rs = LL_RESERVED_SPACE(skb->dev);
 	mtu = skb->dev->mtu;
diff --git a/net/ipv4/ip_output.c b/net/ipv4/ip_output.c
index a24cc8ee1..fa6a74d20 100644
--- a/net/ipv4/ip_output.c
+++ b/net/ipv4/ip_output.c
@@ -770,16 +770,28 @@ int ip_do_fragment(struct net *net, struct sock *sk, struct sk_buff *skb,
 	struct ip_frag_state state;
 	int err = 0;
 
-	/* for offloaded checksums cleanup checksum before fragmentation */
-	if (skb->ip_summed == CHECKSUM_PARTIAL &&
-	    (err = skb_checksum_help(skb)))
-		goto fail;
-
 	/*
 	 *	Point into the IP datagram header.
 	 */
 
 	iph = ip_hdr(skb);
+	hlen = iph->ihl * 4;
+	if (unlikely(hlen < sizeof(*iph) || hlen > skb_headlen(skb))) {
+		err = -EINVAL;
+		goto fail;
+	}
+
+	/* Complete offloaded checksums only after the validated IP header. */
+	if (skb->ip_summed == CHECKSUM_PARTIAL) {
+		if (unlikely(skb_checksum_start_offset(skb) < (int)hlen)) {
+			err = -EINVAL;
+			goto fail;
+		}
+		err = skb_checksum_help(skb);
+		if (err)
+			goto fail;
+		iph = ip_hdr(skb);
+	}
 
 	mtu = ip_skb_dst_mtu(sk, skb);
 	if (IPCB(skb)->frag_max_size && IPCB(skb)->frag_max_size < mtu)
@@ -789,7 +801,6 @@ int ip_do_fragment(struct net *net, struct sock *sk, struct sk_buff *skb,
 	 *	Setup starting values.
 	 */
 
-	hlen = iph->ihl * 4;
 	if (mtu < hlen + 8) {
 		err = -EMSGSIZE;
 		goto fail;
diff --git a/net/ipv6/ip6_output.c b/net/ipv6/ip6_output.c
index 550965058..d157b6ade 100644
--- a/net/ipv6/ip6_output.c
+++ b/net/ipv6/ip6_output.c
@@ -942,9 +942,15 @@ int ip6_fragment(struct net *net, struct sock *sk, struct sk_buff *skb,
 	frag_id = ipv6_select_ident(net, &ipv6_hdr(skb)->daddr,
 				    &ipv6_hdr(skb)->saddr);
 
-	if (skb->ip_summed == CHECKSUM_PARTIAL &&
-	    (err = skb_checksum_help(skb)))
-		goto fail;
+	if (skb->ip_summed == CHECKSUM_PARTIAL) {
+		if (unlikely(skb_checksum_start_offset(skb) < (int)hlen)) {
+			err = -EINVAL;
+			goto fail;
+		}
+		err = skb_checksum_help(skb);
+		if (err)
+			goto fail;
+	}
 
 	prevhdr = skb_network_header(skb) + nexthdr_offset;
 	hroom = LL_RESERVED_SPACE(rt->dst.dev);
diff --git a/net/ipv6/netfilter.c b/net/ipv6/netfilter.c
index a7025ec87..da7ada12f 100644
--- a/net/ipv6/netfilter.c
+++ b/net/ipv6/netfilter.c
@@ -144,9 +144,15 @@ int br_ip6_fragment(struct net *net, struct sock *sk, struct sk_buff *skb,
 	frag_id = ipv6_select_ident(net, &ipv6_hdr(skb)->daddr,
 				    &ipv6_hdr(skb)->saddr);
 
-	if (skb->ip_summed == CHECKSUM_PARTIAL &&
-	    (err = skb_checksum_help(skb)))
-		goto blackhole;
+	if (skb->ip_summed == CHECKSUM_PARTIAL) {
+		if (unlikely(skb_checksum_start_offset(skb) < (int)hlen)) {
+			err = -EINVAL;
+			goto blackhole;
+		}
+		err = skb_checksum_help(skb);
+		if (err)
+			goto blackhole;
+	}
 
 	prevhdr = skb_network_header(skb) + nexthdr_offset;
 	hroom = LL_RESERVED_SPACE(skb->dev);
-- 
2.46.0

^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-09-21  2:53 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-20  0:47 [PATCH net v4 0/2] net: prevent partial checksums from modifying IPv4 headers Paulos Yibelo
2026-09-20  0:47 ` [PATCH net v4 1/2] net: validate virtio checksum start after network header Paulos Yibelo
2026-09-20  1:11   ` David Ahern
2026-09-20  0:47 ` [PATCH net v4 2/2] ipv4: reject partial checksums covering the IP header Paulos Yibelo
2026-09-20  1:12   ` David Ahern
2026-09-21  2:53 ` [PATCH net v5 0/2] net: prevent partial checksums from modifying network headers Paulos Yibelo
2026-09-21  2:53   ` [PATCH net v5 1/2] net: validate virtio checksum start after network header Paulos Yibelo
2026-09-21  2:53   ` [PATCH net v5 2/2] ip: reject partial checksums covering network headers Paulos Yibelo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®