* [PATCH net] netlink: do not copy to user space under netlink_lock_table()
@ 2026-10-08 19:48 Cen Zhang (Microsoft)
2026-10-08 20:42 ` Eric Dumazet
0 siblings, 1 reply; 3+ messages in thread
From: Cen Zhang (Microsoft) @ 2026-10-08 19:48 UTC (permalink / raw)
To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Simon Horman, Kees Cook, Jeff Layton, Nicolas Dichtel,
Ilya Maximets, Kexin Sun, Enrico Pozzobon, Breno Leitao,
David Rheinsberg, netdev, linux-kernel, stable,
AutonomousCodeSecurity, tgopinath, Cen Zhang
netlink_getsockopt() handles NETLINK_LIST_MEMBERSHIPS by copying the
socket's group bitmap to user space with copy_to_iter() under
netlink_lock_table(). However, copy_to_iter() writes to a user address
and can trigger a page fault, and how long that fault takes is up to the
owner of the page. An unprivileged user can make it take as long as they
want, e.g. by mapping a memfd page and keeping it hole punched from
another thread. While the fault is pending, the caller still counts as a
holder of nl_table_users.
nl_table_users is one global counter shared by every netlink protocol
and every network namespace, and netlink_table_grab() waits for it to
drop to zero in TASK_UNINTERRUPTIBLE with no timeout and no signal
check. Therefore, by holding one page fault open, an unprivileged user
blocks every task on the machine that needs netlink_table_grab(): e.g.,
close of a netlink socket, multicast group join, and network namespace
creation. The blocked tasks stay in an unkillable D state until the
attacker lets the fault finish, and hung task detector reports is as
following:
INFO: task V1-nl-close:103 blocked for more than 122 seconds.
task:V1-nl-close state:D stack:14448 pid:103 ...
Call Trace:
schedule+0x36/0xf0
netlink_table_grab.part.0+0x66/0xe0
netlink_release+0x6d4/0x780
__x64_sys_close+0x38/0x80
Fix this by taking a kmemdup() snapshot of the bitmap under
netlink_lock_table() and copying the snapshot to user space after the
lock is released.
Fixes: 47191d65b647 ("netlink: fix locking around NETLINK_LIST_MEMBERSHIPS")
Reported-by: AutonomousCodeSecurity@microsoft.com
Signed-off-by: Cen Zhang (Microsoft) <cenzhang@linux.microsoft.com>
---
net/netlink/af_netlink.c | 20 ++++++++++++++++----
1 file changed, 16 insertions(+), 4 deletions(-)
diff --git a/net/netlink/af_netlink.c b/net/netlink/af_netlink.c
index 9fdf964224ab..f61ff699761f 100644
--- a/net/netlink/af_netlink.c
+++ b/net/netlink/af_netlink.c
@@ -1767,24 +1767,36 @@ static int netlink_getsockopt(struct socket *sock, int level, int optname,
flag = NETLINK_F_RECV_NO_ENOBUFS;
break;
case NETLINK_LIST_MEMBERSHIPS: {
+ unsigned long *groups = NULL;
int pos, idx, shift, err = 0;
+ unsigned int ngroups;
+ /* copy_to_iter() may sleep in a page fault, do it unlocked. */
netlink_lock_table();
- for (pos = 0; pos * 8 < nlk->ngroups; pos += sizeof(u32)) {
+ ngroups = nlk->ngroups;
+ if (ngroups)
+ groups = kmemdup(nlk->groups, NLGRPSZ(ngroups),
+ GFP_KERNEL);
+ netlink_unlock_table();
+
+ if (ngroups && !groups)
+ return -ENOMEM;
+
+ for (pos = 0; pos * 8 < ngroups; pos += sizeof(u32)) {
if (len - pos < sizeof(u32))
break;
idx = pos / sizeof(unsigned long);
shift = (pos % sizeof(unsigned long)) * 8;
- group = (u32)(nlk->groups[idx] >> shift);
+ group = (u32)(groups[idx] >> shift);
if (copy_to_iter(&group, sizeof(u32),
&opt->iter_out) != sizeof(u32)) {
err = -EFAULT;
break;
}
}
- opt->optlen = ALIGN(BITS_TO_BYTES(nlk->ngroups), sizeof(u32));
- netlink_unlock_table();
+ opt->optlen = ALIGN(BITS_TO_BYTES(ngroups), sizeof(u32));
+ kfree(groups);
return err;
}
case NETLINK_LISTEN_ALL_NSID:
--
2.55.0
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH net] netlink: do not copy to user space under netlink_lock_table()
2026-10-08 19:48 [PATCH net] netlink: do not copy to user space under netlink_lock_table() Cen Zhang (Microsoft)
@ 2026-10-08 20:42 ` Eric Dumazet
2026-10-08 21:31 ` Cen Zhang (Microsoft)
0 siblings, 1 reply; 3+ messages in thread
From: Eric Dumazet @ 2026-10-08 20:42 UTC (permalink / raw)
To: Cen Zhang (Microsoft)
Cc: David S. Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
Kees Cook, Jeff Layton, Nicolas Dichtel, Ilya Maximets,
Kexin Sun, Enrico Pozzobon, Breno Leitao, David Rheinsberg,
netdev, linux-kernel, stable, AutonomousCodeSecurity, tgopinath
Le jeu. 8 oct. 2026 à 21:48, Cen Zhang (Microsoft)
<cenzhang@linux.microsoft.com> a écrit :
>
> netlink_getsockopt() handles NETLINK_LIST_MEMBERSHIPS by copying the
> socket's group bitmap to user space with copy_to_iter() under
> netlink_lock_table(). However, copy_to_iter() writes to a user address
> and can trigger a page fault, and how long that fault takes is up to the
> owner of the page. An unprivileged user can make it take as long as they
> want, e.g. by mapping a memfd page and keeping it hole punched from
> another thread. While the fault is pending, the caller still counts as a
> holder of nl_table_users.
>
> nl_table_users is one global counter shared by every netlink protocol
> and every network namespace, and netlink_table_grab() waits for it to
> drop to zero in TASK_UNINTERRUPTIBLE with no timeout and no signal
> check. Therefore, by holding one page fault open, an unprivileged user
> blocks every task on the machine that needs netlink_table_grab(): e.g.,
> close of a netlink socket, multicast group join, and network namespace
> creation. The blocked tasks stay in an unkillable D state until the
> attacker lets the fault finish, and hung task detector reports is as
> following:
>
> INFO: task V1-nl-close:103 blocked for more than 122 seconds.
> task:V1-nl-close state:D stack:14448 pid:103 ...
> Call Trace:
> schedule+0x36/0xf0
> netlink_table_grab.part.0+0x66/0xe0
> netlink_release+0x6d4/0x780
> __x64_sys_close+0x38/0x80
>
> Fix this by taking a kmemdup() snapshot of the bitmap under
> netlink_lock_table() and copying the snapshot to user space after the
> lock is released.
>
> Fixes: 47191d65b647 ("netlink: fix locking around NETLINK_LIST_MEMBERSHIPS")
> Reported-by: AutonomousCodeSecurity@microsoft.com
> Signed-off-by: Cen Zhang (Microsoft) <cenzhang@linux.microsoft.com>
> ---
This is certainly not a net candidate. I call this AI hallucination.
This is not specific to netlink_lock_table().
We have many places where copy_{from,to}_user() runs while a mutex or
a socket lock is held, and some of these locks are global.
One example among many: do_ip_getsockopt() handles IP_MSFILTER with
copy_from_sockptr(), then copy_to_sockptr() in ip_mc_msfget(), all of
it under rtnl_lock(). No privilege needed.
Even in af_netlink.c, your patch does not close the issue you describe:
netlink_bind() calls netlink_insert() under netlink_lock_table(), and
netlink_insert() starts with lock_sock(sk). Another thread can hold
that socket lock across a user copy, e.g. setsockopt(SO_ATTACH_FILTER),
where __get_filter() calls copy_from_user() under sockopt_lock_sock().
So if "user controlled page fault while holding a lock others wait on"
is the bug, a real fix is probably hundreds of patches all over the
tree, not this one.
I do not think we want to do this one call site at a time, and
certainly not as fixes for net and stable.
What is the plan here?
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH net] netlink: do not copy to user space under netlink_lock_table()
2026-10-08 20:42 ` Eric Dumazet
@ 2026-10-08 21:31 ` Cen Zhang (Microsoft)
0 siblings, 0 replies; 3+ messages in thread
From: Cen Zhang (Microsoft) @ 2026-10-08 21:31 UTC (permalink / raw)
To: edumazet
Cc: AutonomousCodeSecurity, cenzhang, davem, david, enrico.pozzobon,
horms, i.maximets, jlayton, kees, kexinsun, kuba, leitao,
linux-kernel, netdev, nicolas.dichtel, pabeni, stable, tgopinath
Hi Eric,
Thanks for your time and the detailed explanation. I withdraw the patch,
and will take the points you raised into account when validating future
findings.
Thanks again, and sorry for the noise.
Cen
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-08 21:31 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-08 19:48 [PATCH net] netlink: do not copy to user space under netlink_lock_table() Cen Zhang (Microsoft)
2026-10-08 20:42 ` Eric Dumazet
2026-10-08 21:31 ` Cen Zhang (Microsoft)
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®