* [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks
2026-10-02 4:50 [PATCH net v2 0/3] ipv4: handle nexthop group shrink races Daehyeon Ko
@ 2026-10-02 4:50 ` Daehyeon Ko
2026-10-02 4:50 ` [PATCH net v2 2/3] ipv4: stop exception dump " Daehyeon Ko
2026-10-02 4:51 ` [PATCH net v2 3/3] ipv4: stop route notification sizing " Daehyeon Ko
2 siblings, 0 replies; 4+ messages in thread
From: Daehyeon Ko @ 2026-10-02 4:50 UTC (permalink / raw)
To: David Ahern, Ido Schimmel
Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Vladimir Vdovin, Donald Hunter, Amit Cohen,
Roopa Prabhu, netdev, linux-kernel, Daehyeon Ko
Commit 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if
multipath enabled") made __ip_rt_update_pmtu() update every path.
For nexthop objects, fib_info_num_path() and fib_info_nhc()
independently load nh->nh_grp. RCU protects each group's lifetime but
does not make the loads observe the same group. If replacement shrinks
the group after the loop accepts an index, fib_info_nhc() returns NULL
and update_or_create_fnhe() dereferences it.
A deterministic probe only widened the existing window between the real
operations. On v7.2, an RTM_NEWNEXTHOP replacement published a
one-member group after the PMTU reader accepted index 1 from a two-member
group, producing:
KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
RIP: update_or_create_fnhe+0x45/0x15b0
Stop when the indexed path is absent. Groups are dense, so no later
index exists in that snapshot. The same real-writer window completed
114,423 PMTU iterations without an oops with this check. A reproducer
is available on request.
Replacement requires CAP_NET_ADMIN in the network namespace. Where
unprivileged user namespaces are permitted, a local user can obtain it
in a new user and network namespace.
Fixes: 7d3f3b4367f3 ("net: ipv4: Cache pmtu for all packet paths if multipath enabled")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
---
net/ipv4/route.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index 37674d76f90f0..5f2197874bebc 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1082,6 +1082,8 @@ static void __ip_rt_update_pmtu(struct rtable *rt, struct flowi4 *fl4, u32 mtu)
for (nhsel = 0; nhsel < fib_info_num_path(res.fi); nhsel++) {
nhc = fib_info_nhc(res.fi, nhsel);
+ if (!nhc)
+ break;
update_or_create_fnhe(nhc, fl4->daddr, 0, mtu, lock,
jiffies + net->ipv4.ip_rt_mtu_expires);
}
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH net v2 2/3] ipv4: stop exception dump when nexthop group shrinks
2026-10-02 4:50 [PATCH net v2 0/3] ipv4: handle nexthop group shrink races Daehyeon Ko
2026-10-02 4:50 ` [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks Daehyeon Ko
@ 2026-10-02 4:50 ` Daehyeon Ko
2026-10-02 4:51 ` [PATCH net v2 3/3] ipv4: stop route notification sizing " Daehyeon Ko
2 siblings, 0 replies; 4+ messages in thread
From: Daehyeon Ko @ 2026-10-02 4:50 UTC (permalink / raw)
To: David Ahern, Ido Schimmel
Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Vladimir Vdovin, Donald Hunter, Amit Cohen,
Roopa Prabhu, netdev, linux-kernel, Daehyeon Ko
Commit 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU
protection") allowed IPv4 FIB dumps to run concurrently with nexthop
replacement.
fib_dump_info_fnhe() uses fib_info_num_path() as the loop bound,
then fib_info_nhc() independently reloads nh->nh_grp. If
RTM_NEWNEXTHOP replaces a group with fewer paths between those loads,
fib_info_nhc() returns NULL and the dump dereferences it through
nhc_flags.
Stop the walk when the indexed path is absent. Nexthop groups are
dense, so the current snapshot has no later path.
Fixes: 4ce5dc9316de ("inet: switch inet_dump_fib() to RCU protection")
Reported-by: Ido Schimmel <idosch@nvidia.com>
Link: https://lore.kernel.org/netdev/20261001170513.GA1657889@shredder/
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
---
net/ipv4/route.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index 5f2197874bebc..268d8821f7e09 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -3170,6 +3170,9 @@ int fib_dump_info_fnhe(struct sk_buff *skb, struct netlink_callback *cb,
struct fnhe_hash_bucket *bucket;
int err;
+ if (!nhc)
+ break;
+
if (nhc->nhc_flags & RTNH_F_DEAD)
continue;
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH net v2 3/3] ipv4: stop route notification sizing when nexthop group shrinks
2026-10-02 4:50 [PATCH net v2 0/3] ipv4: handle nexthop group shrink races Daehyeon Ko
2026-10-02 4:50 ` [PATCH net v2 1/3] ipv4: stop PMTU walk when nexthop group shrinks Daehyeon Ko
2026-10-02 4:50 ` [PATCH net v2 2/3] ipv4: stop exception dump " Daehyeon Ko
@ 2026-10-02 4:51 ` Daehyeon Ko
2 siblings, 0 replies; 4+ messages in thread
From: Daehyeon Ko @ 2026-10-02 4:51 UTC (permalink / raw)
To: David Ahern, Ido Schimmel
Cc: David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Vladimir Vdovin, Donald Hunter, Amit Cohen,
Roopa Prabhu, netdev, linux-kernel, Daehyeon Ko
Commit 680aea08e78c ("net: ipv4: Emit notification when fib
hardware flags are changed") added an RCU-only fib_nlmsg_size() call
for asynchronous hardware flag notifications.
fib_nlmsg_size() checks fib_info_num_path() before each iteration,
then fib_info_nhc() independently reloads nh->nh_grp. If
RTM_NEWNEXTHOP replaces a group with fewer paths between those loads,
fib_info_nhc() returns NULL and fib_nexthop_nlmsg_size() dereferences
it.
Stop sizing when the indexed path is absent. Nexthop groups are
dense, so the current snapshot has no later path.
Fixes: 680aea08e78c ("net: ipv4: Emit notification when fib hardware flags are changed")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Daehyeon Ko <4ncienth@gmail.com>
---
net/ipv4/fib_semantics.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/net/ipv4/fib_semantics.c b/net/ipv4/fib_semantics.c
index 5c9021ea3a799..47469b7758082 100644
--- a/net/ipv4/fib_semantics.c
+++ b/net/ipv4/fib_semantics.c
@@ -542,6 +542,9 @@ size_t fib_nlmsg_size(struct fib_info *fi)
struct fib_nh_common *nhc = fib_info_nhc(fi, i);
size_t nhsize;
+ if (!nhc)
+ break;
+
nhsize = fib_nexthop_nlmsg_size(nhc, nhs != 1);
if (nhs != 1)
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread