* [PATCH net] ipv6: fix fib6 walker UAF on NLM_F_REPLACE
@ 2026-10-09 22:51 Cen Zhang (Microsoft)
2026-10-10 22:55 ` netdev-bot+sashiko
0 siblings, 1 reply; 2+ messages in thread
From: Cen Zhang (Microsoft) @ 2026-10-09 22:51 UTC (permalink / raw)
To: David Ahern, Ido Schimmel, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni
Cc: Simon Horman, Matti Vaittinen, Ben Greear, netdev, linux-kernel,
stable, AutonomousCodeSecurity, tgopinath, Cen Zhang
fib6_walker uses w->leaf to remember the next fib6_info to dump, so an
RTM_GETROUTE dump can continue there on the next recvmsg(). This
assumes that whoever removes a route from the tree also updates the
walkers pointing at it. The NLM_F_REPLACE branch of fib6_add_rt2node()
can break this assumption. It removes the old route, fib6_add() frees
it, and w->leaf keeps pointing at the freed object.
The fn_sernum version check in fib6_dump_table() normally protects
against resuming on a stale w->leaf, but it can be bypassed:
ip6_link_failure() resets fn_sernum to -1 when the default route
fails, which an unprivileged user in a user+net namespace can cause:
1) link failure fn_sernum = -1
2) dump pauses on route R
3) replace R fn_sernum changed, R freed
4) link failure fn_sernum = -1
5) dump resumes -1 == -1, reads freed R <- UAF
BUG: KASAN: slab-use-after-free in rt6_fill_node+0x1c22/0x2cd0
Read of size 4 at addr ff11000011fa0064 by task poc/101
rt6_fill_node <- rt6_dump_route
Fix by making the walker adjustment of fib6_del_route() as a common
function fib6_walkers_skip_rt() and call it from both fib6_del_route()
and the NLM_F_REPLACE branch, for the replaced route and each removed
ECMP sibling.
Fixes: 4a287eba2de3 ("IPv6 routing, NLM_F_* flag support: REPLACE and EXCL flags support, warn about missing CREATE flag")
Reported-by: AutonomousCodeSecurity@microsoft.com
Assisted-by: LLM
Signed-off-by: Cen Zhang (Microsoft) <cenzhang@linux.microsoft.com>
---
net/ipv6/ip6_fib.c | 39 ++++++++++++++++++++++++++-------------
1 file changed, 26 insertions(+), 13 deletions(-)
diff --git a/net/ipv6/ip6_fib.c b/net/ipv6/ip6_fib.c
index 9ff761962b45..a5699809fe04 100644
--- a/net/ipv6/ip6_fib.c
+++ b/net/ipv6/ip6_fib.c
@@ -89,6 +89,29 @@ static void fib6_walker_unlink(struct net *net, struct fib6_walker *w)
write_unlock_bh(&net->ipv6.fib6_walker_lock);
}
+/* Move the walkers resting on @rt to the next route, or up the tree if
+ * @rt is the last one. Called with tb6_lock held.
+ */
+static void fib6_walkers_skip_rt(struct net *net, struct fib6_info *rt)
+{
+ struct fib6_info *next;
+ struct fib6_walker *w;
+
+ next = rcu_dereference_protected(rt->fib6_next,
+ lockdep_is_held(&rt->fib6_table->tb6_lock));
+
+ read_lock(&net->ipv6.fib6_walker_lock);
+ FOR_WALKERS(net, w) {
+ if (w->state == FWS_C && w->leaf == rt) {
+ pr_debug("walker %p adjusted by unlink of %p\n", w, rt);
+ w->leaf = next;
+ if (!next)
+ w->state = FWS_U;
+ }
+ }
+ read_unlock(&net->ipv6.fib6_walker_lock);
+}
+
static int fib6_new_sernum(struct net *net)
{
int new, old = atomic_read(&net->ipv6.fib6_sernum);
@@ -1315,6 +1338,7 @@ static int fib6_add_rt2node(struct fib6_node *fn, struct fib6_info *rt,
fib6_info_hold(rt);
rcu_assign_pointer(rt->fib6_node, fn);
rt->fib6_next = iter->fib6_next;
+ fib6_walkers_skip_rt(info->nl_net, iter);
rcu_assign_pointer(*ins, rt);
if (!info->skip_notify)
inet6_rt_notify(RTM_NEWROUTE, rt, info, NLM_F_REPLACE);
@@ -1337,6 +1361,7 @@ static int fib6_add_rt2node(struct fib6_node *fn, struct fib6_info *rt,
if (iter->fib6_metric > rt->fib6_metric)
break;
if (rt6_qualify_for_ecmp(iter)) {
+ fib6_walkers_skip_rt(info->nl_net, iter);
*ins = iter->fib6_next;
iter->fib6_node = NULL;
list_add(&iter->purge_link, purge_list);
@@ -1971,7 +1996,6 @@ static void fib6_del_route(struct fib6_table *table, struct fib6_node *fn,
enum rt_del_reason del_reason)
{
struct fib6_info *leaf, *replace_rt = NULL;
- struct fib6_walker *w;
struct fib6_info *rt = rcu_dereference_protected(*rtp,
lockdep_is_held(&table->tb6_lock));
struct net *net = info->nl_net;
@@ -2022,18 +2046,7 @@ static void fib6_del_route(struct fib6_table *table, struct fib6_node *fn,
rt6_multipath_rebalance(next_sibling);
}
- /* Adjust walkers */
- read_lock(&net->ipv6.fib6_walker_lock);
- FOR_WALKERS(net, w) {
- if (w->state == FWS_C && w->leaf == rt) {
- pr_debug("walker %p adjusted by delroute\n", w);
- w->leaf = rcu_dereference_protected(rt->fib6_next,
- lockdep_is_held(&table->tb6_lock));
- if (!w->leaf)
- w->state = FWS_U;
- }
- }
- read_unlock(&net->ipv6.fib6_walker_lock);
+ fib6_walkers_skip_rt(net, rt);
/* If it was last route, call fib6_repair_tree() to:
* 1. For root node, put back null_entry as how the table was created.
--
2.55.0
^ permalink raw reply [flat|nested] 2+ messages in thread* Re: [PATCH net] ipv6: fix fib6 walker UAF on NLM_F_REPLACE
2026-10-09 22:51 [PATCH net] ipv6: fix fib6 walker UAF on NLM_F_REPLACE Cen Zhang (Microsoft)
@ 2026-10-10 22:55 ` netdev-bot+sashiko
0 siblings, 0 replies; 2+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 22:55 UTC (permalink / raw)
To: cenzhang
Cc: dsahern, idosch, davem, edumazet, kuba, pabeni, horms,
mazziesaccount, greearb, netdev, linux-kernel, stable,
AutonomousCodeSecurity, tgopinath
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 2 (1) · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] fib6_walkers_skip_rt() (net/ipv6/ip6_fib.c:95-113) is
documented as moving walkers that rest on @rt to the next route.
- [Medium] In the NLM_F_REPLACE branch of fib6_add_rt2node()
(net/ipv6/ip6_fib.c:1338-1342) the new route rt takes over iter's slot…
Pre-existing issues:
- [Medium] When a dump resumes and the sernum has changed,
fib6_dump_table() (net/ipv6/ip6_fib.c:632-638) writes w->state =…
--
ipv6: fix fib6 walker UAF on NLM_F_REPLACE
This patch moves the walker adjustment out of fib6_del_route() into a new
helper, fib6_walkers_skip_rt(). The helper is also called from the
NLM_F_REPLACE branch of fib6_add_rt2node() for the replaced route and for
each removed ECMP sibling. The goal is to keep an RTM_GETROUTE dump from
resuming on a freed fib6_info when the fn_sernum check is bypassed.
> diff --git a/net/ipv6/ip6_fib.c b/net/ipv6/ip6_fib.c
> index 9ff761962b45c..a5699809fe049 100644
> --- a/net/ipv6/ip6_fib.c
> +++ b/net/ipv6/ip6_fib.c
> @@ -89,6 +89,29 @@ static void fib6_walker_unlink(struct net *net, struct fib6_walker *w)
[ ... ]
> +/* Move the walkers resting on @rt to the next route, or up the tree if
> + * @rt is the last one. Called with tb6_lock held.
> + */
> +static void fib6_walkers_skip_rt(struct net *net, struct fib6_info *rt)
> +{
[ ... ]
> + read_lock(&net->ipv6.fib6_walker_lock);
> + FOR_WALKERS(net, w) {
> + if (w->state == FWS_C && w->leaf == rt) {
> + pr_debug("walker %p adjusted by unlink of %p\n", w, rt);
> + w->leaf = next;
[Severity: Medium]
Should w->skip_in_node be reset here too?
This moves the walker to a different route, but w->skip_in_node still
holds the number of entries already sent for the old route.
fib6_dump_node() keeps that count per w->leaf:
net/ipv6/ip6_fib.c:fib6_dump_node() {
...
res = rt6_dump_route(rt, w->args, w->skip_in_node);
if (res >= 0) {
/* Frame is full, suspend walking */
w->leaf = rt;
...
w->skip_in_node += res;
...
}
fib6_dump_table() clears it only on a sernum mismatch. When the sernum
matches, it does just this:
} else
w->skip = 0;
Suppose the dump paused partway through the exceptions of route R
(skip_in_node >= 1), and R is then removed through this helper. The
resumed dump calls rt6_dump_route() on R->fib6_next with the stale skip:
net/ipv6/route.c:rt6_dump_route() {
...
if (filter->dump_routes) {
if (skip) {
skip--;
} else {
if (rt6_fill_node(net, arg->skb, rt, NULL, NULL, NULL,
...
}
That drops the next route's own RTM_NEWROUTE message, and the leftover
skip hides its first exceptions. In the FWS_U case, the first route of a
later node is dropped instead. NLM_F_DUMP_INTR is not set in either case.
With this patch, the same thing can happen for a replaced route and for
each removed ECMP sibling in fib6_add_rt2node(). That needs the resume to
go through the fn_sernum == -1 collision described in the commit message.
fib6_del() does not bump the root sernum, so the delete path behaved this
way on a normal resume even before the code was moved out of
fib6_del_route().
Would adding w->skip_in_node = 0 where w->leaf is reassigned make the
helper match its comment?
> + if (!next)
> + w->state = FWS_U;
[Severity: Medium]
This is a pre-existing issue and was not introduced by this patch, but can
this store race with the lockless walker reset in fib6_dump_table()?
When the sernum has changed, fib6_dump_table() resets the walker before it
takes tb6_lock. The walker is still linked on net->ipv6.fib6_walkers at
that point:
net/ipv6/ip6_fib.c:fib6_dump_table() {
...
if (cb->args[5] != sernum) {
/* Begin at the root if the tree changed */
cb->args[5] = sernum;
w->state = FWS_INIT;
w->node = w->root;
w->skip = w->count;
w->skip_in_node = 0;
} else
w->skip = 0;
spin_lock_bh(&table->tb6_lock);
...
}
At the same time, this helper holds only tb6_lock and
read_lock(fib6_walker_lock), and it can still see w->state == FWS_C and
w->leaf == rt.
If its FWS_U store lands after the FWS_INIT store, fib6_walk_continue()
starts at w->root in state FWS_U:
case FWS_U:
if (fn == w->root)
return 0;
The rest of the table is then silently left out of the dump, with no
NLM_F_DUMP_INTR.
The same race already existed with fib6_del_route() as the writer. This
patch adds the replace paths in fib6_add_rt2node() as more writers. I don't
see a use-after-free here, because FWS_INIT and FWS_U do not dereference
w->leaf.
[ ... ]
> @@ -1315,6 +1338,7 @@ static int fib6_add_rt2node(struct fib6_node *fn, struct fib6_info *rt,
> fib6_info_hold(rt);
> rcu_assign_pointer(rt->fib6_node, fn);
> rt->fib6_next = iter->fib6_next;
> + fib6_walkers_skip_rt(info->nl_net, iter);
> rcu_assign_pointer(*ins, rt);
[Severity: Medium]
Can a resumed dump skip the replacement route rt here?
In this branch, rt takes over iter's slot in the leaf list. The helper
treats that as a deletion. It moves walkers on iter to iter->fib6_next,
which comes after rt, or sets FWS_U if iter was the last route in the
node.
fib6_dump_node() leaves w->leaf on the route it hadn't finished. In the
common case, the pause happened at rt6_fill_node() of iter itself with
skip_in_node == 0, so nothing from iter has been sent yet.
Suppose the dump then resumes on the sernum-match path (the
fn_sernum == -1 case from the commit message). fib6_dump_node() starts at
w->leaf and never visits rt. The dump then contains neither the old route
nor its replacement for that prefix and metric, and NLM_F_DUMP_INTR is not
set.
On the sernum-mismatch path, the node is walked again from fn->leaf and rt
is dumped, so the two resume paths give different results.
The ECMP sibling loop further down looks fine, since those routes really
are removed.
For the replaced slot itself, would pointing walkers on iter at rt (and
clearing skip_in_node) keep the replacement in the dump?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009225142.47005-1-cenzhang%40linux.microsoft.com
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-10 22:55 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 22:51 [PATCH net] ipv6: fix fib6 walker UAF on NLM_F_REPLACE Cen Zhang (Microsoft)
2026-10-10 22:55 ` netdev-bot+sashiko
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®