mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net v2 0/3] ipv4: handle nexthop group shrink races
@ 2026-10-02  4:50 Daehyeon Ko
  2026-10-02  4:50 ` [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks Daehyeon Ko
                   ` (2 more replies)
  0 siblings, 3 replies; 4+ messages in thread
From: Daehyeon Ko @ 2026-10-02  4:50 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, Vladimir Vdovin, Donald Hunter, Amit Cohen,
	Roopa Prabhu, netdev, linux-kernel, Daehyeon Ko

fib_info_num_path() and fib_info_nhc() can observe different RCU
generations of a replaceable nexthop group.  An indexed consumer can
therefore accept an index from a larger group and receive NULL after a
concurrent shrink.

Patch 1 is unchanged from v1.  Patch 2 addresses the analogous
fib_dump_info_fnhe() race Ido identified.  Auditing the remaining
accessor pairs found the same issue in the RCU-only hardware-flag
notification path, fixed by patch 3.  Separate patches retain the
correct Fixes tag for each concurrency boundary.

Thanks to Ido for the review and follow-up pointer.

---
v2:
- Carry Ido's Reviewed-by on unchanged patch 1.
- Add separate exception-dump and notification-sizing fixes.

v1: https://lore.kernel.org/netdev/20261001010550.2742297-1-4ncienth@gmail.com/
review: https://lore.kernel.org/netdev/20261001170513.GA1657889@shredder/

Validation:
- Patch 1 retains its deterministic vulnerable/fixed result.
- Patches 2 and 3 are source-audited only.  No new allyesconfig or
  allmodconfig W=1 build or runtime test was run.

Daehyeon Ko (3):
  ipv4: stop PMTU walk when nexthop group shrinks
  ipv4: stop exception dump when nexthop group shrinks
  ipv4: stop route notification sizing when nexthop group shrinks

 net/ipv4/fib_semantics.c | 3 +++
 net/ipv4/route.c         | 5 +++++
 2 files changed, 8 insertions(+)


base-commit: 28bc1ef699610ee09ce3d46a00552f5a0a0144bd
-- 
2.55.0

^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks
  2026-10-02  4:50 [PATCH net v2 0/3] ipv4: handle nexthop group shrink races Daehyeon Ko
@ 2026-10-02  4:50 ` Daehyeon Ko
  2026-10-02  4:50 ` [PATCH net v2 2/3] ipv4: stop exception dump " Daehyeon Ko
  2026-10-02  4:51 ` [PATCH net v2 3/3] ipv4: stop route notification sizing " Daehyeon Ko
  2 siblings, 0 replies; 4+ messages in thread
From: Daehyeon Ko @ 2026-10-02  4:50 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, Vladimir Vdovin, Donald Hunter, Amit Cohen,
	Roopa Prabhu, netdev, linux-kernel, Daehyeon Ko

Commit 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if
multipath enabled") made __ip_rt_update_pmtu() update every path.

For nexthop objects, fib_info_num_path() and fib_info_nhc()
independently load nh->nh_grp.  RCU protects each group's lifetime but
does not make the loads observe the same group.  If replacement shrinks
the group after the loop accepts an index, fib_info_nhc() returns NULL
and update_or_create_fnhe() dereferences it.

A deterministic probe only widened the existing window between the real
operations.  On v7.2, an RTM_NEWNEXTHOP replacement published a
one-member group after the PMTU reader accepted index 1 from a two-member
group, producing:

  KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
  RIP: update_or_create_fnhe+0x45/0x15b0

Stop when the indexed path is absent.  Groups are dense, so no later
index exists in that snapshot.  The same real-writer window completed
114,423 PMTU iterations without an oops with this check.  A reproducer
is available on request.

Replacement requires CAP_NET_ADMIN in the network namespace.  Where
unprivileged user namespaces are permitted, a local user can obtain it
in a new user and network namespace.

Fixes: 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if multipath enabled")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
---
 net/ipv4/route.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index 37674d76f90f0..5f2197874bebc 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1082,6 +1082,8 @@ static void __ip_rt_update_pmtu(struct rtable *rt, struct flowi4 *fl4, u32 mtu)
 
 			for (nhsel = 0; nhsel < fib_info_num_path(res.fi); nhsel++) {
 				nhc = fib_info_nhc(res.fi, nhsel);
+				if (!nhc)
+					break;
 				update_or_create_fnhe(nhc, fl4->daddr, 0, mtu, lock,
 						      jiffies + net->ipv4.ip_rt_mtu_expires);
 			}
-- 
2.55.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH net v2 2/3] ipv4: stop exception dump when nexthop group shrinks
  2026-10-02  4:50 [PATCH net v2 0/3] ipv4: handle nexthop group shrink races Daehyeon Ko
  2026-10-02  4:50 ` [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks Daehyeon Ko
@ 2026-10-02  4:50 ` Daehyeon Ko
  2026-10-02  4:51 ` [PATCH net v2 3/3] ipv4: stop route notification sizing " Daehyeon Ko
  2 siblings, 0 replies; 4+ messages in thread
From: Daehyeon Ko @ 2026-10-02  4:50 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, Vladimir Vdovin, Donald Hunter, Amit Cohen,
	Roopa Prabhu, netdev, linux-kernel, Daehyeon Ko

Commit 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU
protection") allowed IPv4 FIB dumps to run concurrently with nexthop
replacement.

fib_dump_info_fnhe() uses fib_info_num_path() as the loop bound,
then fib_info_nhc() independently reloads nh->nh_grp. If
RTM_NEWNEXTHOP replaces a group with fewer paths between those loads,
fib_info_nhc() returns NULL and the dump dereferences it through
nhc_flags.

Stop the walk when the indexed path is absent. Nexthop groups are
dense, so the current snapshot has no later path.

Fixes: 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU protection")
Reported-by: Ido Schimmel <idosch@nvidia.com>
Link: https://lore.kernel.org/netdev/20261001170513.GA1657889@shredder/
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
---
 net/ipv4/route.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index 5f2197874bebc..268d8821f7e09 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -3170,6 +3170,9 @@ int fib_dump_info_fnhe(struct sk_buff *skb, struct netlink_callback *cb,
 		struct fnhe_hash_bucket *bucket;
 		int err;
 
+		if (!nhc)
+			break;
+
 		if (nhc->nhc_flags & RTNH_F_DEAD)
 			continue;
 
-- 
2.55.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH net v2 3/3] ipv4: stop route notification sizing when nexthop group shrinks
  2026-10-02  4:50 [PATCH net v2 0/3] ipv4: handle nexthop group shrink races Daehyeon Ko
  2026-10-02  4:50 ` [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks Daehyeon Ko
  2026-10-02  4:50 ` [PATCH net v2 2/3] ipv4: stop exception dump " Daehyeon Ko
@ 2026-10-02  4:51 ` Daehyeon Ko
  2 siblings, 0 replies; 4+ messages in thread
From: Daehyeon Ko @ 2026-10-02  4:51 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel
  Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Simon Horman, Vladimir Vdovin, Donald Hunter, Amit Cohen,
	Roopa Prabhu, netdev, linux-kernel, Daehyeon Ko

Commit 680aea08e78c ("net: ipv4: Emit notification when fib
hardware flags are changed") added an RCU-only fib_nlmsg_size() call
for asynchronous hardware flag notifications.

fib_nlmsg_size() checks fib_info_num_path() before each iteration,
then fib_info_nhc() independently reloads nh->nh_grp. If
RTM_NEWNEXTHOP replaces a group with fewer paths between those loads,
fib_info_nhc() returns NULL and fib_nexthop_nlmsg_size() dereferences
it.

Stop sizing when the indexed path is absent. Nexthop groups are
dense, so the current snapshot has no later path.

Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
---
 net/ipv4/fib_semantics.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/net/ipv4/fib_semantics.c b/net/ipv4/fib_semantics.c
index 5c9021ea3a799..47469b7758082 100644
--- a/net/ipv4/fib_semantics.c
+++ b/net/ipv4/fib_semantics.c
@@ -542,6 +542,9 @@ size_t fib_nlmsg_size(struct fib_info *fi)
 			struct fib_nh_common *nhc = fib_info_nhc(fi, i);
 			size_t nhsize;
 
+			if (!nhc)
+				break;
+
 			nhsize = fib_nexthop_nlmsg_size(nhc, nhs != 1);
 
 			if (nhs != 1)
-- 
2.55.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-02  4:51 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02  4:50 [PATCH net v2 0/3] ipv4: handle nexthop group shrink races Daehyeon Ko
2026-10-02  4:50 ` [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks Daehyeon Ko
2026-10-02  4:50 ` [PATCH net v2 2/3] ipv4: stop exception dump " Daehyeon Ko
2026-10-02  4:51 ` [PATCH net v2 3/3] ipv4: stop route notification sizing " Daehyeon Ko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®