mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
@ 2026-09-27  7:10 Chengfeng Ye
  2026-09-28  3:17 ` Hangbin Liu
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: Chengfeng Ye @ 2026-09-27  7:10 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, John Heffner
  Cc: netdev, linux-kernel, Chengfeng Ye, stable

When a socket ignores the path MTU, ip_skb_dst_mtu() reads the device
MTU without RCU protection. The multicast and broadcast output path
through ip_mc_output() can reach this helper without an RCU read lock.
Netfilter releases its internal read lock before calling
ip_finish_output().

A sender can load dst->dev and then be preempted before reading dev->mtu.
Concurrent device unregistration replaces dst->dev with blackhole_netdev
and drops the old device reference. After the grace period and remaining
references drain, the device can be freed before the sender resumes and
reads its MTU, causing a use-after-free.

KASAN reported:

  BUG: KASAN: slab-use-after-free in ip_skb_dst_mtu+0x634/0x740
  Read of size 4 at addr ffff88810921c038 by task poc/98
  Call Trace:
   ip_skb_dst_mtu+0x634/0x740
   __ip_finish_output.part.0+0x22/0x2c0
   ip_mc_output+0x287/0x930
   ip_send_skb+0x11d/0x150
   udp_send_skb+0x63e/0xdf0
   udp_sendmsg+0x1235/0x1da0
   __sys_sendto+0x32c/0x3a0

  Allocated by task 97:
   __kvmalloc_node_noprof+0x1d4/0x620
   alloc_netdev_mqs+0x81/0x12c0
   rtnl_create_link+0xaa3/0xe30
   rtnl_newlink+0xabc/0x1f70

  Freed by task 99:
   kfree+0x131/0x3c0
   device_release+0xc8/0x240
   kobject_put+0x14d/0x280
   netdev_run_todo+0x4cb/0xc70
   rtnl_dellink+0x362/0xa90

Protect the device lookup and MTU read with RCU and use dst_dev_rcu() to
access the device pointer. Keep the MTU limit and headroom calculation
unchanged.

Fixes: 628a5c561890 ("[INET]: Add IP(V6)_PMTUDISC_RPOBE")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
---
 include/net/ip.h | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/include/net/ip.h b/include/net/ip.h
index 6f602df72ee6..14d6f77f3bb8 100644
--- a/include/net/ip.h
+++ b/include/net/ip.h
@@ -543,7 +543,9 @@ static inline unsigned int ip_skb_dst_mtu(struct sock *sk,
 		return ip_dst_mtu_maybe_forward(dst, forwarding);
 	}
 
-	mtu = min(READ_ONCE(dst_dev(dst)->mtu), IP_MAX_MTU);
+	rcu_read_lock();
+	mtu = min(READ_ONCE(dst_dev_rcu(dst)->mtu), IP_MAX_MTU);
+	rcu_read_unlock();
 	return mtu - lwtunnel_headroom(dst->lwtstate, mtu);
 }
 
-- 
2.43.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
  2026-09-27  7:10 [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu() Chengfeng Ye
@ 2026-09-28  3:17 ` Hangbin Liu
  2026-09-28 12:34 ` Ido Schimmel
  2026-09-29 12:10 ` netdev-bot+sashiko
  2 siblings, 0 replies; 5+ messages in thread
From: Hangbin Liu @ 2026-09-28  3:17 UTC (permalink / raw)
  To: Chengfeng Ye
  Cc: David Ahern, Ido Schimmel, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, John Heffner, netdev,
	linux-kernel, stable

On Sun, Sep 27, 2026 at 03:10:51PM +0800, Chengfeng Ye wrote:
> When a socket ignores the path MTU, ip_skb_dst_mtu() reads the device
> MTU without RCU protection. The multicast and broadcast output path
> through ip_mc_output() can reach this helper without an RCU read lock.
> Netfilter releases its internal read lock before calling
> ip_finish_output().
> 
> A sender can load dst->dev and then be preempted before reading dev->mtu.
> Concurrent device unregistration replaces dst->dev with blackhole_netdev
> and drops the old device reference. After the grace period and remaining
> references drain, the device can be freed before the sender resumes and
> reads its MTU, causing a use-after-free.
> 
> KASAN reported:
> 
>   BUG: KASAN: slab-use-after-free in ip_skb_dst_mtu+0x634/0x740
>   Read of size 4 at addr ffff88810921c038 by task poc/98
>   Call Trace:
>    ip_skb_dst_mtu+0x634/0x740
>    __ip_finish_output.part.0+0x22/0x2c0
>    ip_mc_output+0x287/0x930
>    ip_send_skb+0x11d/0x150
>    udp_send_skb+0x63e/0xdf0
>    udp_sendmsg+0x1235/0x1da0
>    __sys_sendto+0x32c/0x3a0
> 
>   Allocated by task 97:
>    __kvmalloc_node_noprof+0x1d4/0x620
>    alloc_netdev_mqs+0x81/0x12c0
>    rtnl_create_link+0xaa3/0xe30
>    rtnl_newlink+0xabc/0x1f70
> 
>   Freed by task 99:
>    kfree+0x131/0x3c0
>    device_release+0xc8/0x240
>    kobject_put+0x14d/0x280
>    netdev_run_todo+0x4cb/0xc70
>    rtnl_dellink+0x362/0xa90
> 
> Protect the device lookup and MTU read with RCU and use dst_dev_rcu() to
> access the device pointer. Keep the MTU limit and headroom calculation
> unchanged.
> 
> Fixes: 628a5c561890 ("[INET]: Add IP(V6)_PMTUDISC_RPOBE")
> Cc: stable@vger.kernel.org
> Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
> ---
>  include/net/ip.h | 4 +++-
>  1 file changed, 3 insertions(+), 1 deletion(-)
> 
> diff --git a/include/net/ip.h b/include/net/ip.h
> index 6f602df72ee6..14d6f77f3bb8 100644
> --- a/include/net/ip.h
> +++ b/include/net/ip.h
> @@ -543,7 +543,9 @@ static inline unsigned int ip_skb_dst_mtu(struct sock *sk,
>  		return ip_dst_mtu_maybe_forward(dst, forwarding);
>  	}
>  
> -	mtu = min(READ_ONCE(dst_dev(dst)->mtu), IP_MAX_MTU);
> +	rcu_read_lock();
> +	mtu = min(READ_ONCE(dst_dev_rcu(dst)->mtu), IP_MAX_MTU);
> +	rcu_read_unlock();
>  	return mtu - lwtunnel_headroom(dst->lwtstate, mtu);
>  }
>  
> -- 
> 2.43.0
> 

Reviewed-by: Hangbin Liu <liuhangbin@kylinos.cn>

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
  2026-09-27  7:10 [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu() Chengfeng Ye
  2026-09-28  3:17 ` Hangbin Liu
@ 2026-09-28 12:34 ` Ido Schimmel
  2026-09-28 16:29   ` Chengfeng Ye
  2026-09-29 12:10 ` netdev-bot+sashiko
  2 siblings, 1 reply; 5+ messages in thread
From: Ido Schimmel @ 2026-09-28 12:34 UTC (permalink / raw)
  To: Chengfeng Ye
  Cc: David Ahern, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, John Heffner, netdev, linux-kernel,
	stable

On Sun, Sep 27, 2026 at 03:10:51PM +0800, Chengfeng Ye wrote:
> When a socket ignores the path MTU, ip_skb_dst_mtu() reads the device
> MTU without RCU protection. The multicast and broadcast output path
> through ip_mc_output() can reach this helper without an RCU read lock.
> Netfilter releases its internal read lock before calling
> ip_finish_output().
> 
> A sender can load dst->dev and then be preempted before reading dev->mtu.
> Concurrent device unregistration replaces dst->dev with blackhole_netdev
> and drops the old device reference. After the grace period and remaining
> references drain, the device can be freed before the sender resumes and
> reads its MTU, causing a use-after-free.
> 
> KASAN reported:
> 
>   BUG: KASAN: slab-use-after-free in ip_skb_dst_mtu+0x634/0x740
>   Read of size 4 at addr ffff88810921c038 by task poc/98
>   Call Trace:
>    ip_skb_dst_mtu+0x634/0x740
>    __ip_finish_output.part.0+0x22/0x2c0
>    ip_mc_output+0x287/0x930
>    ip_send_skb+0x11d/0x150
>    udp_send_skb+0x63e/0xdf0
>    udp_sendmsg+0x1235/0x1da0
>    __sys_sendto+0x32c/0x3a0
> 
>   Allocated by task 97:
>    __kvmalloc_node_noprof+0x1d4/0x620
>    alloc_netdev_mqs+0x81/0x12c0
>    rtnl_create_link+0xaa3/0xe30
>    rtnl_newlink+0xabc/0x1f70
> 
>   Freed by task 99:
>    kfree+0x131/0x3c0
>    device_release+0xc8/0x240
>    kobject_put+0x14d/0x280
>    netdev_run_todo+0x4cb/0xc70
>    rtnl_dellink+0x362/0xa90
> 
> Protect the device lookup and MTU read with RCU and use dst_dev_rcu() to
> access the device pointer. Keep the MTU limit and headroom calculation
> unchanged.

This still leaves other potential UAFs in the ip_mc_output() path.
Better to fix it by adding an RCU read-side critical section in
ip_mc_output() in a similar fashion to its unicast counterpart. See
commit 1dbf1d590d10 ("net: Add locking to protect skb->dev access in
ip_output"). You can blame 4a6ce2b6f2ec ("net: introduce a new function
dst_dev_put()").

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
  2026-09-28 12:34 ` Ido Schimmel
@ 2026-09-28 16:29   ` Chengfeng Ye
  0 siblings, 0 replies; 5+ messages in thread
From: Chengfeng Ye @ 2026-09-28 16:29 UTC (permalink / raw)
  To: Ido Schimmel
  Cc: David Ahern, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, John Heffner, netdev, linux-kernel,
	stable

On Mon, Sep 28, 2026 at 8:34 PM Ido Schimmel <idosch@nvidia.com> wrote:
>
> On Sun, Sep 27, 2026 at 03:10:51PM +0800, Chengfeng Ye wrote:
> > When a socket ignores the path MTU, ip_skb_dst_mtu() reads the device
> > MTU without RCU protection. The multicast and broadcast output path
> > through ip_mc_output() can reach this helper without an RCU read lock.
> > Netfilter releases its internal read lock before calling
> > ip_finish_output().
> >
> > A sender can load dst->dev and then be preempted before reading dev->mtu.
> > Concurrent device unregistration replaces dst->dev with blackhole_netdev
> > and drops the old device reference. After the grace period and remaining
> > references drain, the device can be freed before the sender resumes and
> > reads its MTU, causing a use-after-free.
> >
> > KASAN reported:
> >
> >   BUG: KASAN: slab-use-after-free in ip_skb_dst_mtu+0x634/0x740
> >   Read of size 4 at addr ffff88810921c038 by task poc/98
> >   Call Trace:
> >    ip_skb_dst_mtu+0x634/0x740
> >    __ip_finish_output.part.0+0x22/0x2c0
> >    ip_mc_output+0x287/0x930
> >    ip_send_skb+0x11d/0x150
> >    udp_send_skb+0x63e/0xdf0
> >    udp_sendmsg+0x1235/0x1da0
> >    __sys_sendto+0x32c/0x3a0
> >
> >   Allocated by task 97:
> >    __kvmalloc_node_noprof+0x1d4/0x620
> >    alloc_netdev_mqs+0x81/0x12c0
> >    rtnl_create_link+0xaa3/0xe30
> >    rtnl_newlink+0xabc/0x1f70
> >
> >   Freed by task 99:
> >    kfree+0x131/0x3c0
> >    device_release+0xc8/0x240
> >    kobject_put+0x14d/0x280
> >    netdev_run_todo+0x4cb/0xc70
> >    rtnl_dellink+0x362/0xa90
> >
> > Protect the device lookup and MTU read with RCU and use dst_dev_rcu() to
> > access the device pointer. Keep the MTU limit and headroom calculation
> > unchanged.
>
> This still leaves other potential UAFs in the ip_mc_output() path.
> Better to fix it by adding an RCU read-side critical section in
> ip_mc_output() in a similar fashion to its unicast counterpart. See
> commit 1dbf1d590d10 ("net: Add locking to protect skb->dev access in
> ip_output"). You can blame 4a6ce2b6f2ec ("net: introduce a new function
> dst_dev_put()").

Thanks for the reviews. The v2 patch has just been sent by applying
the method of 1dbf1d590d10 to ip_mc_output().

Best regards,
Chengfeng

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu()
  2026-09-27  7:10 [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu() Chengfeng Ye
  2026-09-28  3:17 ` Hangbin Liu
  2026-09-28 12:34 ` Ido Schimmel
@ 2026-09-29 12:10 ` netdev-bot+sashiko
  2 siblings, 0 replies; 5+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 12:10 UTC (permalink / raw)
  To: nicoyip.dev
  Cc: dsahern, idosch, davem, edumazet, kuba, pabeni, horms, jheffner,
	netdev, linux-kernel, stable

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 0 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] This fix is incomplete.

--

ipv4: Fix device use-after-free in ip_skb_dst_mtu()

This patch wraps the dev->mtu read in ip_skb_dst_mtu() in
rcu_read_lock()/rcu_read_unlock() and switches it to dst_dev_rcu(). It
targets a KASAN-reported net_device use-after-free on the multicast and
broadcast output path, where ip_mc_output() reaches ip_finish_output()
without an RCU read lock held.

> diff --git a/include/net/ip.h b/include/net/ip.h
> index 6f602df72ee62..14d6f77f3bb89 100644
> --- a/include/net/ip.h
> +++ b/include/net/ip.h
> @@ -543,7 +543,9 @@ static inline unsigned int ip_skb_dst_mtu(struct sock *sk,
>  		return ip_dst_mtu_maybe_forward(dst, forwarding);
>  	}
>  
> -	mtu = min(READ_ONCE(dst_dev(dst)->mtu), IP_MAX_MTU);
> +	rcu_read_lock();
> +	mtu = min(READ_ONCE(dst_dev_rcu(dst)->mtu), IP_MAX_MTU);
> +	rcu_read_unlock();

[Severity: High]
Is this fix complete? The commit message names the root cause as
ip_mc_output() reaching ip_finish_output() with no RCU read lock held.
That is still the case after this patch, so the other device reads on the
same path are still unprotected.

ip_mc_output() in net/ipv4/ip_output.c does a plain load of the device.
It takes no RCU and no reference, then stores the pointer in skb->dev and
passes it as outdev:

net/ipv4/ip_output.c:ip_mc_output() {
	struct rtable *rt = skb_rtable(skb);
	struct net_device *dev = rt->dst.dev;
	...
	skb->dev = dev;
	...
	return NF_HOOK_COND(NFPROTO_IPV4, NF_INET_POST_ROUTING,
			    net, sk, skb, NULL, skb->dev,
			    ip_finish_output,
			    !(IPCB(skb)->flags & IPSKB_REROUTED));
}

Right after ip_skb_dst_mtu() returns, __ip_finish_output()->
ip_finish_output2() reads the device again, before its own
rcu_read_lock():

net/ipv4/ip_output.c:ip_finish_output2() {
	struct net_device *dev = dst_dev(dst);
	unsigned int hh_len = LL_RESERVED_SPACE(dev);
	...
	if (unlikely(skb_headroom(skb) < hh_len && dev->header_ops)) {
	...
	rcu_read_lock();
	neigh = ip_neigh_for_gw(rt, skb, &is_v6gw);
	...
}

Can the same race described in the commit message now hit freed memory in
ip_finish_output2() instead of in ip_skb_dst_mtu()?

The same seems to apply to ip_do_fragment(). It calls ip_skb_dst_mtu() and
then reads the device outside RCU when reached from ip_mc_output() via
ip_fragment():

	ll_rs = LL_RESERVED_SPACE(rt->dst.dev);

There are more users of the pointer that ip_mc_output() saved in skb->dev.
On the GSO path, ip_finish_output_gso() calls netif_skb_features(skb).
The multicast and broadcast loopback clones go through
ip_mc_finish_output()->dev_loopback_xmit().

No caller on the UDP path from the KASAN trace provides RCU either.
ip_send_skb() calls ip_local_out() directly:

net/ipv4/ip_output.c:ip_send_skb() {
	err = ip_local_out(net, skb->sk, skb);
	...
}

ip6_send_skb() wraps ip6_local_out() in rcu_read_lock().

The unicast path already has the full-coverage pattern. It was added to
ip_output() by commit 1dbf1d590d10a6 ("net: Add locking to protect
skb->dev access in ip_output"):

	rcu_read_lock();
	dev = skb_dst_dev_rcu(skb);
	skb->dev = dev;
	...
	ret_val = NF_HOOK_COND(NFPROTO_IPV4, NF_INET_POST_ROUTING,
				net, sk, skb, indev, dev,
				ip_finish_output,
				!(IPCB(skb)->flags & IPSKB_REROUTED));
	rcu_read_unlock();

Would it be better to apply the same outer RCU section in ip_mc_output(),
and maybe also in ip_send_skb() to match ip6_send_skb()? That would cover
the whole ip_finish_output() chain instead of only the dev->mtu read here.

>  	return mtu - lwtunnel_headroom(dst->lwtstate, mtu);
>  }

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260927071051.3693368-1-nicoyip.dev%40gmail.com

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-29 12:11 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27  7:10 [PATCH net] ipv4: Fix device use-after-free in ip_skb_dst_mtu() Chengfeng Ye
2026-09-28  3:17 ` Hangbin Liu
2026-09-28 12:34 ` Ido Schimmel
2026-09-28 16:29   ` Chengfeng Ye
2026-09-29 12:10 ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®