* Re: [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
2026-09-27 7:10 [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu() Chengfeng Ye
@ 2026-09-28 3:17 ` Hangbin Liu
2026-09-28 12:34 ` Ido Schimmel
2026-09-29 12:10 ` netdev-bot+sashiko
2 siblings, 0 replies; 5+ messages in thread
From: Hangbin Liu @ 2026-09-28 3:17 UTC (permalink / raw)
To: Chengfeng Ye
Cc: David Ahern, Ido Schimmel, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, John Heffner, netdev,
linux-kernel, stable
On Sun, Sep 27, 2026 at 03:10:51PM +0800, Chengfeng Ye wrote:
> When a socket ignores the path MTU, ip_skb_dst_mtu() reads the device
> MTU without RCU protection. The multicast and broadcast output path
> through ip_mc_output() can reach this helper without an RCU read lock.
> Netfilter releases its internal read lock before calling
> ip_finish_output().
>
> A sender can load dst->dev and then be preempted before reading dev->mtu.
> Concurrent device unregistration replaces dst->dev with blackhole_netdev
> and drops the old device reference. After the grace period and remaining
> references drain, the device can be freed before the sender resumes and
> reads its MTU, causing a use-after-free.
>
> KASAN reported:
>
> BUG: KASAN: slab-use-after-free in ip_skb_dst_mtu+0x634/0x740
> Read of size 4 at addr ffff88810921c038 by task poc/98
> Call Trace:
> ip_skb_dst_mtu+0x634/0x740
> __ip_finish_output.part.0+0x22/0x2c0
> ip_mc_output+0x287/0x930
> ip_send_skb+0x11d/0x150
> udp_send_skb+0x63e/0xdf0
> udp_sendmsg+0x1235/0x1da0
> __sys_sendto+0x32c/0x3a0
>
> Allocated by task 97:
> __kvmalloc_node_noprof+0x1d4/0x620
> alloc_netdev_mqs+0x81/0x12c0
> rtnl_create_link+0xaa3/0xe30
> rtnl_newlink+0xabc/0x1f70
>
> Freed by task 99:
> kfree+0x131/0x3c0
> device_release+0xc8/0x240
> kobject_put+0x14d/0x280
> netdev_run_todo+0x4cb/0xc70
> rtnl_dellink+0x362/0xa90
>
> Protect the device lookup and MTU read with RCU and use dst_dev_rcu() to
> access the device pointer. Keep the MTU limit and headroom calculation
> unchanged.
>
> Fixes: 628a5c561890 ("[INET]: Add IP(V6)_PMTUDISC_RPOBE")
> Cc: stable@vger.kernel.org
> Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
> ---
> include/net/ip.h | 4 +++-
> 1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/include/net/ip.h b/include/net/ip.h
> index 6f602df72ee6..14d6f77f3bb8 100644
> --- a/include/net/ip.h
> +++ b/include/net/ip.h
> @@ -543,7 +543,9 @@ static inline unsigned int ip_skb_dst_mtu(struct sock *sk,
> return ip_dst_mtu_maybe_forward(dst, forwarding);
> }
>
> - mtu = min(READ_ONCE(dst_dev(dst)->mtu), IP_MAX_MTU);
> + rcu_read_lock();
> + mtu = min(READ_ONCE(dst_dev_rcu(dst)->mtu), IP_MAX_MTU);
> + rcu_read_unlock();
> return mtu - lwtunnel_headroom(dst->lwtstate, mtu);
> }
>
> --
> 2.43.0
>
Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
2026-09-27 7:10 [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu() Chengfeng Ye
2026-09-28 3:17 ` Hangbin Liu
@ 2026-09-28 12:34 ` Ido Schimmel
2026-09-28 16:29 ` Chengfeng Ye
2026-09-29 12:10 ` netdev-bot+sashiko
2 siblings, 1 reply; 5+ messages in thread
From: Ido Schimmel @ 2026-09-28 12:34 UTC (permalink / raw)
To: Chengfeng Ye
Cc: David Ahern, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, John Heffner, netdev, linux-kernel,
stable
On Sun, Sep 27, 2026 at 03:10:51PM +0800, Chengfeng Ye wrote:
> When a socket ignores the path MTU, ip_skb_dst_mtu() reads the device
> MTU without RCU protection. The multicast and broadcast output path
> through ip_mc_output() can reach this helper without an RCU read lock.
> Netfilter releases its internal read lock before calling
> ip_finish_output().
>
> A sender can load dst->dev and then be preempted before reading dev->mtu.
> Concurrent device unregistration replaces dst->dev with blackhole_netdev
> and drops the old device reference. After the grace period and remaining
> references drain, the device can be freed before the sender resumes and
> reads its MTU, causing a use-after-free.
>
> KASAN reported:
>
> BUG: KASAN: slab-use-after-free in ip_skb_dst_mtu+0x634/0x740
> Read of size 4 at addr ffff88810921c038 by task poc/98
> Call Trace:
> ip_skb_dst_mtu+0x634/0x740
> __ip_finish_output.part.0+0x22/0x2c0
> ip_mc_output+0x287/0x930
> ip_send_skb+0x11d/0x150
> udp_send_skb+0x63e/0xdf0
> udp_sendmsg+0x1235/0x1da0
> __sys_sendto+0x32c/0x3a0
>
> Allocated by task 97:
> __kvmalloc_node_noprof+0x1d4/0x620
> alloc_netdev_mqs+0x81/0x12c0
> rtnl_create_link+0xaa3/0xe30
> rtnl_newlink+0xabc/0x1f70
>
> Freed by task 99:
> kfree+0x131/0x3c0
> device_release+0xc8/0x240
> kobject_put+0x14d/0x280
> netdev_run_todo+0x4cb/0xc70
> rtnl_dellink+0x362/0xa90
>
> Protect the device lookup and MTU read with RCU and use dst_dev_rcu() to
> access the device pointer. Keep the MTU limit and headroom calculation
> unchanged.
This still leaves other potential UAFs in the ip_mc_output() path.
Better to fix it by adding an RCU read-side critical section in
ip_mc_output() in a similar fashion to its unicast counterpart. See
commit 1dbf1d590d10 ("net: Add locking to protect skb->dev access in
ip_output"). You can blame 4a6ce2b6f2ec ("net: introduce a new function
dst_dev_put()").
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
2026-09-27 7:10 [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu() Chengfeng Ye
2026-09-28 3:17 ` Hangbin Liu
2026-09-28 12:34 ` Ido Schimmel
@ 2026-09-29 12:10 ` netdev-bot+sashiko
2 siblings, 0 replies; 5+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 12:10 UTC (permalink / raw)
To: nicoyip.dev
Cc: dsahern, idosch, davem, edumazet, kuba, pabeni, horms, jheffner,
netdev, linux-kernel, stable
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 1 · Medium: 0 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] This fix is incomplete.
--
ipv4: Fix device use-after-free in ip_skb_dst_mtu()
This patch wraps the dev->mtu read in ip_skb_dst_mtu() in
rcu_read_lock()/rcu_read_unlock() and switches it to dst_dev_rcu(). It
targets a KASAN-reported net_device use-after-free on the multicast and
broadcast output path, where ip_mc_output() reaches ip_finish_output()
without an RCU read lock held.
> diff --git a/include/net/ip.h b/include/net/ip.h
> index 6f602df72ee62..14d6f77f3bb89 100644
> --- a/include/net/ip.h
> +++ b/include/net/ip.h
> @@ -543,7 +543,9 @@ static inline unsigned int ip_skb_dst_mtu(struct sock *sk,
> return ip_dst_mtu_maybe_forward(dst, forwarding);
> }
>
> - mtu = min(READ_ONCE(dst_dev(dst)->mtu), IP_MAX_MTU);
> + rcu_read_lock();
> + mtu = min(READ_ONCE(dst_dev_rcu(dst)->mtu), IP_MAX_MTU);
> + rcu_read_unlock();
[Severity: High]
Is this fix complete? The commit message names the root cause as
ip_mc_output() reaching ip_finish_output() with no RCU read lock held.
That is still the case after this patch, so the other device reads on the
same path are still unprotected.
ip_mc_output() in net/ipv4/ip_output.c does a plain load of the device.
It takes no RCU and no reference, then stores the pointer in skb->dev and
passes it as outdev:
net/ipv4/ip_output.c:ip_mc_output() {
struct rtable *rt = skb_rtable(skb);
struct net_device *dev = rt->dst.dev;
...
skb->dev = dev;
...
return NF_HOOK_COND(NFPROTO_IPV4, NF_INET_POST_ROUTING,
net, sk, skb, NULL, skb->dev,
ip_finish_output,
!(IPCB(skb)->flags & IPSKB_REROUTED));
}
Right after ip_skb_dst_mtu() returns, __ip_finish_output()->
ip_finish_output2() reads the device again, before its own
rcu_read_lock():
net/ipv4/ip_output.c:ip_finish_output2() {
struct net_device *dev = dst_dev(dst);
unsigned int hh_len = LL_RESERVED_SPACE(dev);
...
if (unlikely(skb_headroom(skb) < hh_len && dev->header_ops)) {
...
rcu_read_lock();
neigh = ip_neigh_for_gw(rt, skb, &is_v6gw);
...
}
Can the same race described in the commit message now hit freed memory in
ip_finish_output2() instead of in ip_skb_dst_mtu()?
The same seems to apply to ip_do_fragment(). It calls ip_skb_dst_mtu() and
then reads the device outside RCU when reached from ip_mc_output() via
ip_fragment():
ll_rs = LL_RESERVED_SPACE(rt->dst.dev);
There are more users of the pointer that ip_mc_output() saved in skb->dev.
On the GSO path, ip_finish_output_gso() calls netif_skb_features(skb).
The multicast and broadcast loopback clones go through
ip_mc_finish_output()->dev_loopback_xmit().
No caller on the UDP path from the KASAN trace provides RCU either.
ip_send_skb() calls ip_local_out() directly:
net/ipv4/ip_output.c:ip_send_skb() {
err = ip_local_out(net, skb->sk, skb);
...
}
ip6_send_skb() wraps ip6_local_out() in rcu_read_lock().
The unicast path already has the full-coverage pattern. It was added to
ip_output() by commit 1dbf1d590d10a6 ("net: Add locking to protect
skb->dev access in ip_output"):
rcu_read_lock();
dev = skb_dst_dev_rcu(skb);
skb->dev = dev;
...
ret_val = NF_HOOK_COND(NFPROTO_IPV4, NF_INET_POST_ROUTING,
net, sk, skb, indev, dev,
ip_finish_output,
!(IPCB(skb)->flags & IPSKB_REROUTED));
rcu_read_unlock();
Would it be better to apply the same outer RCU section in ip_mc_output(),
and maybe also in ip_send_skb() to match ip6_send_skb()? That would cover
the whole ip_finish_output() chain instead of only the dev->mtu read here.
> return mtu - lwtunnel_headroom(dst->lwtstate, mtu);
> }
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260927071051.3693368-1-nicoyip.dev%40gmail.com
^ permalink raw reply [flat|nested] 5+ messages in thread