* [PATCH bpf v3 1/2] bpf: reject incompatible socket assignments
@ 2026-09-30 16:40 Shihuang Liu
2026-09-30 16:40 ` [PATCH bpf v3 2/2] bpf: reject incompatible protocol changes Shihuang Liu
0 siblings, 1 reply; 3+ messages in thread
From: Shihuang Liu @ 2026-09-30 16:40 UTC (permalink / raw)
To: bpf
Cc: ast, daniel, andrii, eddyz87, memxor, martin.lau, song,
yonghong.song, jolsa, emil, ihor.solodrai, john.fastabend, sdf,
davem, edumazet, kuba, pabeni, horms, linux-kernel, netdev,
Shihuang Liu
bpf_sk_assign() permits TC ingress programs to associate an IPv6 packet
with an AF_INET socket. The receive path can then interpret IPv6 skb
control data as IPv4 metadata. When IP_RETOPTS is enabled, this can cause
__ip_options_echo() to copy beyond its stack buffer.
Reject incompatible packet and socket families before attaching the socket,
while preserving IPv4 assignments to dual-stack AF_INET6 sockets. Validate
request, mapped, and time-wait sockets using their effective family.
Use the packet's effective protocol for VLAN checks and share the same
predicate across bpf_sk_assign() and bpf_sk_assign_tcp_reqsk().
Fixes: cf7fbe660f2d ("bpf: Add socket assign support")
Assisted-by: LLM
Signed-off-by: Shihuang Liu <shlomojune6@gmail.com>
---
Changes since v2:
- Use the effective packet protocol for VLAN-aware family validation.
- Handle IPv4-mapped AF_INET6 sockets, TCP children, and TIME_WAIT sockets.
- Share the protocol-based check with bpf_sk_assign_tcp_reqsk().
v2:
https://lore.kernel.org/netdev/20260911172313.64009-1-shlomojune6@gmail.com/
Changes since v1:
- Move family validation out of the IPv6 receive fast path and into
bpf_sk_assign() and bpf_sk_assign_tcp_reqsk().
- Preserve IPv4 assignments to dual-stack AF_INET6 sockets.
- Check request sockets using rsk_ops->family.
- Split the fix into two patches and target the BPF fixes tree.
v1:
https://lore.kernel.org/netdev/20260823101809.26802-1-shlomojune6@gmail.com/
---
include/uapi/linux/bpf.h | 4 +++
net/core/filter.c | 63 ++++++++++++++++++++++++++++++++++
tools/include/uapi/linux/bpf.h | 4 +++
3 files changed, 71 insertions(+)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 732b35cc08d1c..5d8f5e2c8db38 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -4568,6 +4568,10 @@ union bpf_attr {
* **-EOPNOTSUPP** if the operation is not supported, for example
* a call from outside of TC ingress.
*
+ * **-EAFNOSUPPORT** if the socket family is not compatible with
+ * the network layer of the packet, for example an **AF_INET**
+ * socket and an IPv6 packet.
+ *
* long bpf_sk_assign(struct bpf_sk_lookup *ctx, struct bpf_sock *sk, u64 flags)
* Description
* Helper is overloaded depending on BPF program type. This
diff --git a/net/core/filter.c b/net/core/filter.c
index 61940e7535523..5f64065523584 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -3491,6 +3491,64 @@ static int bpf_skb_proto_xlat(struct sk_buff *skb, __be16 to_proto)
return -ENOTSUPP;
}
+static bool bpf_sk_assign_family_ok_proto(const struct sock *sk, __be16 proto)
+{
+ const struct inet_connection_sock_af_ops *af_ops;
+ unsigned short family;
+
+ switch (proto) {
+ case htons(ETH_P_IP):
+ family = AF_INET;
+ break;
+ case htons(ETH_P_IPV6):
+ family = AF_INET6;
+ break;
+ default:
+ return true;
+ }
+
+ /* Requests inherit the listener family, but have family-specific ops. */
+ if (sk->sk_state == TCP_NEW_SYN_RECV)
+ return inet_reqsk(sk)->rsk_ops->family == family;
+
+ /* A dual-stack listener accepts both packet families. */
+ if (sk->sk_state == TCP_LISTEN && sk->sk_family == AF_INET6)
+ return family == AF_INET6 || !ipv6_only_sock(sk);
+
+#if IS_ENABLED(CONFIG_IPV6)
+ /* IPv4-mapped and pure IPv6 time-wait sockets retain AF_INET6. */
+ if (sk->sk_state == TCP_TIME_WAIT && sk->sk_family == AF_INET6) {
+ const struct inet_timewait_sock *tw = inet_twsk(sk);
+ bool mapped;
+
+ mapped = ipv6_addr_v4mapped(&tw->tw_v6_daddr) &&
+ ipv6_addr_v4mapped(&tw->tw_v6_rcv_saddr);
+ return family == (mapped ? AF_INET : AF_INET6);
+ }
+#endif
+
+ /* IPv4-mapped and pure IPv6 TCP children keep AF_INET6 in sk_family. */
+ if (sk_fullsock(sk) && sk->sk_family == AF_INET6 && sk_is_tcp(sk)) {
+ af_ops = READ_ONCE(inet_csk(sk)->icsk_af_ops);
+ if ((family == AF_INET6 &&
+ af_ops->net_header_len == sizeof(struct iphdr)) ||
+ (family == AF_INET &&
+ af_ops->net_header_len == sizeof(struct ipv6hdr)))
+ return false;
+ }
+
+ return sk->sk_family == family ||
+ (family == AF_INET &&
+ sk->sk_family == AF_INET6 &&
+ !ipv6_only_sock(sk));
+}
+
+static bool bpf_sk_assign_family_ok(const struct sk_buff *skb,
+ const struct sock *sk)
+{
+ return bpf_sk_assign_family_ok_proto(sk, skb_protocol(skb, true));
+}
+
BPF_CALL_3(bpf_skb_change_proto, struct sk_buff *, skb, __be16, proto,
u64, flags)
{
@@ -7989,6 +8047,8 @@ BPF_CALL_3(bpf_sk_assign, struct sk_buff *, skb, struct sock *, sk, u64, flags)
return -ENETUNREACH;
if (sk_unhashed(sk))
return -EOPNOTSUPP;
+ if (!bpf_sk_assign_family_ok(skb, sk))
+ return -EAFNOSUPPORT;
if (sk_is_refcounted(sk) &&
unlikely(!refcount_inc_not_zero(&sk->sk_refcnt)))
return -ENOENT;
@@ -12526,6 +12586,9 @@ __bpf_kfunc int bpf_sk_assign_tcp_reqsk(struct __sk_buff *s, struct sock *sk,
if (net != sock_net(sk))
return -ENETUNREACH;
+ if (!bpf_sk_assign_family_ok(skb, sk))
+ return -EAFNOSUPPORT;
+
switch (skb->protocol) {
case htons(ETH_P_IP):
ops = &tcp_request_sock_ops;
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 732b35cc08d1c..5d8f5e2c8db38 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -4568,6 +4568,10 @@ union bpf_attr {
* **-EOPNOTSUPP** if the operation is not supported, for example
* a call from outside of TC ingress.
*
+ * **-EAFNOSUPPORT** if the socket family is not compatible with
+ * the network layer of the packet, for example an **AF_INET**
+ * socket and an IPv6 packet.
+ *
* long bpf_sk_assign(struct bpf_sk_lookup *ctx, struct bpf_sock *sk, u64 flags)
* Description
* Helper is overloaded depending on BPF program type. This
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH bpf v3 2/2] bpf: reject incompatible protocol changes
2026-09-30 16:40 [PATCH bpf v3 1/2] bpf: reject incompatible socket assignments Shihuang Liu
@ 2026-09-30 16:40 ` Shihuang Liu
2026-09-30 17:26 ` bot+bpf-ci
0 siblings, 1 reply; 3+ messages in thread
From: Shihuang Liu @ 2026-09-30 16:40 UTC (permalink / raw)
To: bpf
Cc: ast, daniel, andrii, eddyz87, memxor, martin.lau, song,
yonghong.song, jolsa, emil, ihor.solodrai, john.fastabend, sdf,
davem, edumazet, kuba, pabeni, horms, linux-kernel, netdev,
Shihuang Liu
An skb assigned to a socket by bpf_sk_assign() can later have its L3
protocol changed by bpf_skb_change_proto() or bpf_skb_adjust_room().
Without revalidation, this bypasses the assignment-time family check.
Reject incompatible protocol changes before modifying the skb, including L3
encapsulation and decapsulation. Reuse the assignment family predicate and
preserve the socket assignment when the change is rejected.
Fixes: cf7fbe660f2d ("bpf: Add socket assign support")
Assisted-by: LLM
Signed-off-by: Shihuang Liu <shlomojune6@gmail.com>
---
Changes since v2:
- Check compatibility before changing the skb, preserving the socket
assignment instead of calling skb_orphan().
- Cover L3 encapsulation and decapsulation in bpf_skb_adjust_room(), as well
as bpf_skb_change_proto().
Changes since v1:
- Revalidate prefetched sockets after bpf_skb_change_proto() changes the
packet protocol, closing a bypass of assignment-time validation.
- Preserve compatible dual-stack assignments and release incompatible
assignments through their existing skb destructor.
- Split the fix into two patches and target the BPF fixes tree.
v2:
https://lore.kernel.org/netdev/20260911172313.64009-2-shlomojune6@gmail.com/
v1:
https://lore.kernel.org/netdev/20260823101809.26802-1-shlomojune6@gmail.com/
---
include/uapi/linux/bpf.h | 10 ++++++++++
net/core/filter.c | 23 +++++++++++++++++++++++
tools/include/uapi/linux/bpf.h | 10 ++++++++++
3 files changed, 43 insertions(+)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 5d8f5e2c8db38..d1897ee13a66c 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -2659,6 +2659,11 @@ union bpf_attr {
* checked and segments are recalculated by the GSO/GRO engine.
* The size for GSO target is adapted as well.
*
+ * If the skb was assigned to a socket by **bpf_sk_assign** and
+ * the requested protocol is incompatible with that socket, the
+ * helper returns **-EAFNOSUPPORT** before translating the packet.
+ * The skb assignment is preserved.
+ *
* All values for *flags* are reserved for future usage, and must
* be left at zero.
*
@@ -3067,6 +3072,11 @@ union bpf_attr {
* removed from the packet. This handles cases where all tunnel
* layers have been decapsulated.
*
+ * If an L3 encapsulation or decapsulation would produce a
+ * protocol incompatible with a socket assigned by
+ * **bpf_sk_assign**, the helper returns **-EAFNOSUPPORT** before
+ * changing the packet. The skb assignment is preserved.
+ *
* A call to this helper is susceptible to change the underlying
* packet buffer. Therefore, at load time, all checks on pointers
* previously done by the verifier are invalidated and must be
diff --git a/net/core/filter.c b/net/core/filter.c
index 5f64065523584..09816da113554 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -3549,6 +3549,12 @@ static bool bpf_sk_assign_family_ok(const struct sk_buff *skb,
return bpf_sk_assign_family_ok_proto(sk, skb_protocol(skb, true));
}
+static bool bpf_skb_proto_change_sk_ok(struct sk_buff *skb, __be16 proto)
+{
+ return !skb_sk_is_prefetched(skb) ||
+ bpf_sk_assign_family_ok_proto(skb->sk, proto);
+}
+
BPF_CALL_3(bpf_skb_change_proto, struct sk_buff *, skb, __be16, proto,
u64, flags)
{
@@ -3556,6 +3562,12 @@ BPF_CALL_3(bpf_skb_change_proto, struct sk_buff *, skb, __be16, proto,
if (unlikely(flags))
return -EINVAL;
+ if (((skb->protocol == htons(ETH_P_IP) &&
+ proto == htons(ETH_P_IPV6)) ||
+ (skb->protocol == htons(ETH_P_IPV6) &&
+ proto == htons(ETH_P_IP))) &&
+ !bpf_skb_proto_change_sk_ok(skb, proto))
+ return -EAFNOSUPPORT;
/* General idea is that this helper does the basic groundwork
* needed for changing the protocol, and eBPF program fills the
@@ -3699,6 +3711,11 @@ static int bpf_skb_net_grow(struct sk_buff *skb, u32 off, u32 len_diff,
if (inner_mac_len > len_diff)
return -EINVAL;
inner_trans = skb->transport_header;
+
+ if (!bpf_skb_proto_change_sk_ok(skb,
+ flags & BPF_F_ADJ_ROOM_ENCAP_L3_IPV6 ?
+ htons(ETH_P_IPV6) : htons(ETH_P_IP)))
+ return -EAFNOSUPPORT;
}
ret = bpf_skb_net_hdr_push(skb, off, len_diff);
@@ -3786,6 +3803,12 @@ static int bpf_skb_net_shrink(struct sk_buff *skb, u32 off, u32 len_diff,
return -ENOTSUPP;
}
+ if (decap &&
+ !bpf_skb_proto_change_sk_ok(skb,
+ flags & BPF_F_ADJ_ROOM_DECAP_L3_IPV6 ?
+ htons(ETH_P_IPV6) : htons(ETH_P_IP)))
+ return -EAFNOSUPPORT;
+
ret = skb_unclone(skb, GFP_ATOMIC);
if (unlikely(ret < 0))
return ret;
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 5d8f5e2c8db38..d1897ee13a66c 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -2659,6 +2659,11 @@ union bpf_attr {
* checked and segments are recalculated by the GSO/GRO engine.
* The size for GSO target is adapted as well.
*
+ * If the skb was assigned to a socket by **bpf_sk_assign** and
+ * the requested protocol is incompatible with that socket, the
+ * helper returns **-EAFNOSUPPORT** before translating the packet.
+ * The skb assignment is preserved.
+ *
* All values for *flags* are reserved for future usage, and must
* be left at zero.
*
@@ -3067,6 +3072,11 @@ union bpf_attr {
* removed from the packet. This handles cases where all tunnel
* layers have been decapsulated.
*
+ * If an L3 encapsulation or decapsulation would produce a
+ * protocol incompatible with a socket assigned by
+ * **bpf_sk_assign**, the helper returns **-EAFNOSUPPORT** before
+ * changing the packet. The skb assignment is preserved.
+ *
* A call to this helper is susceptible to change the underlying
* packet buffer. Therefore, at load time, all checks on pointers
* previously done by the verifier are invalidated and must be
--
2.43.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH bpf v3 2/2] bpf: reject incompatible protocol changes
2026-09-30 16:40 ` [PATCH bpf v3 2/2] bpf: reject incompatible protocol changes Shihuang Liu
@ 2026-09-30 17:26 ` bot+bpf-ci
0 siblings, 0 replies; 3+ messages in thread
From: bot+bpf-ci @ 2026-09-30 17:26 UTC (permalink / raw)
To: shlomojune6, bpf
Cc: ast, daniel, andrii, eddyz87, memxor, martin.lau, song,
yonghong.song, jolsa, emil, ihor.solodrai, john.fastabend, sdf,
davem, edumazet, kuba, pabeni, horms, linux-kernel, netdev,
shlomojune6, ast, andrii, daniel, martin.lau, eddyz87,
yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3477 bytes --]
> bpf: reject incompatible protocol changes
TL;DR: The family revalidation looks incomplete: LWT ingress encap
(bpf_lwt_push_ip_encap, seg6) can still change L3 protocol and leave a
mismatched assigned socket. Extend the check, or narrow the commit message's
claim?
> diff --git a/net/core/filter.c b/net/core/filter.c
> index 173349bd2a3f..22e2ecb0cf2f 100644
> --- a/net/core/filter.c
> +++ b/net/core/filter.c
[ ... ]
> @@ -3549,6 +3549,12 @@ static bool bpf_sk_assign_family_ok(const struct sk_buff *skb,
> return bpf_sk_assign_family_ok_proto(sk, skb_protocol(skb, true));
> }
>
> +static bool bpf_skb_proto_change_sk_ok(struct sk_buff *skb, __be16 proto)
> +{
> + return !skb_sk_is_prefetched(skb) ||
> + bpf_sk_assign_family_ok_proto(skb->sk, proto);
> +}
> +
> BPF_CALL_3(bpf_skb_change_proto, struct sk_buff *, skb, __be16, proto,
> u64, flags)
> {
The commit message says "Without revalidation, this bypasses the
assignment-time family check. Reject incompatible protocol changes before
modifying the skb, including L3 encapsulation and decapsulation."
Is this revalidation complete? It is only wired into bpf_skb_change_proto(),
bpf_skb_net_grow() and bpf_skb_net_shrink(), but another path can still change
the L3 protocol of an skb that keeps its bpf_sk_assign()ed socket.
Here is one concrete sequence:
1. A tc ingress program calls bpf_sk_assign() on an IPv4 packet with an
AF_INET TCP socket. bpf_sk_assign_family_ok() passes and the program
returns TC_ACT_OK.
2. ip_rcv_core() keeps the prefetched socket:
/* Must drop socket now because of tproxy. */
if (!skb_sk_is_prefetched(skb))
skb_orphan(skb);
3. ip_rcv_finish_core() skips early demux because skb->sk is set, then calls
ip_route_input_noref(). A route with "encap bpf in" sets
dst->input = lwtunnel_input.
4. bpf_input() runs the LWT_IN program, which calls
bpf_lwt_in_push_encap(BPF_LWT_ENCAP_IP) with an IPv6 header. That reaches
bpf_lwt_push_ip_encap(..., ingress=true):
} else {
skb->protocol = htons(ETH_P_IPV6);
}
The skb->sk / sock_pfree assignment is left untouched.
BPF_LWT_ENCAP_SEG6 behaves the same way.
5. The program returns BPF_LWT_REROUTE. bpf_lwt_input_reroute() calls
ip6_route_input() and dst_input(), so the packet goes through ip6_input
and on to tcp_v6_rcv().
6. inet6_steal_sock() -> skb_steal_sock() returns the prefetched AF_INET
socket for the IPv6 packet. This is the family mismatch that
bpf_skb_proto_change_sk_ok() is meant to prevent.
In-kernel IP tunnel decapsulation in the same netns keeps the prefetched
socket across a family change in the same way. For example, with sit:
ipip6_rcv()
iptunnel_pull_header() /* sets skb->protocol = ETH_P_IPV6 */
skb_scrub_packet(skb, false) /* does not orphan */
netif_rx()
Should bpf_lwt_push_ip_encap() and bpf_push_seg6_encap() on ingress also call
bpf_skb_proto_change_sk_ok() (or orphan a prefetched socket)? Or should the
commit message narrow its claim to the two tc helpers?
This is not introduced by this patch, and every step above needs
CAP_NET_ADMIN (tc and LWT programs). But it limits the protection the patch
claims to provide.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36746681022
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-30 17:26 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 16:40 [PATCH bpf v3 1/2] bpf: reject incompatible socket assignments Shihuang Liu
2026-09-30 16:40 ` [PATCH bpf v3 2/2] bpf: reject incompatible protocol changes Shihuang Liu
2026-09-30 17:26 ` bot+bpf-ci
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®