mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t
@ 2026-10-09  8:53 Breno Leitao
  2026-10-09  8:53 ` [PATCH net-next v2 1/7] ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt() Breno Leitao
                   ` (6 more replies)
  0 siblings, 7 replies; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

This series continues the migration to sockopt_t, as described in [1].
do_ip_getsockopt() and do_ipv6_getsockopt() are the last two switches
still taking a sockptr_t pair, and everything they forward to now takes
a sockopt_t, so convert them both.

Patch 1 settles what a negative optlen means on SOL_IPV6 before the
conversion makes it matter. do_ip_getsockopt() rejects it, but
do_ipv6_getsockopt() never has, and the int options answer it as 4.
Treat it as 4, as David Laight suggested, instead of failing callers
that rely on it.

Patch 2 fixes a latent bug in sockopt_expand_out() [2] itself: it caps
the expansion at INT_MAX, but iov_iter_revert() refuses to unroll past
the smaller MAX_RW_COUNT.

MCAST_MSFILTER has to move before the switches above it. It is the last
option with the quirk sockopt_expand_out() was added for [2]: optlen only
covers the fixed part, and the real reply size comes from the gf_numsrc
field inside it. Patches 3 and 4 convert it on both families, and store
the reply length only after the final copy succeeds, so a copy failure
can't report a changed optlen.

Patches 5 and 6 convert the switches themselves. Each wrapper keeps its
__user prototype and builds the sockopt_t, so the proto layer is
untouched. ipv6_getsockopt() cannot use sockopt_init_user(), which
rejects a negative optlen, so patch 6 moves the clamp of patch 1 up
there and builds the sockopt_t after it.

Patch 7 grows the selftest with an ip and an ipv6 fixture, pinning the
returned length, errno and reply bytes.

The temporary sockptr_t-to-sockopt_t glue can go away once the higher
level callers take sockopt_t too. For now, the series keeps the proto
callbacks unchanged and converts only the leaf helpers.

Link: https://lore.kernel.org/all/20260401-getsockopt-v2-0-611df6771aff@debian.org/ [1]
Link: https://lore.kernel.org/all/20260914-getsockopt_phase6-v2-0-e48befc9602e@debian.org/ [2]

To: David Ahern <dsahern@kernel.org>
To: Ido Schimmel <idosch@nvidia.com>
To: David S. Miller <davem@davemloft.net>
To: Eric Dumazet <edumazet@google.com>
To: Jakub Kicinski <kuba@kernel.org>
To: Paolo Abeni <pabeni@redhat.com>
To: Simon Horman <horms@kernel.org>
To: Alexei Starovoitov <ast@kernel.org>
To: Daniel Borkmann <daniel@iogearbox.net>
To: Andrii Nakryiko <andrii@kernel.org>
To: Eduard Zingerman <eddyz87@gmail.com>
To: Kumar Kartikeya Dwivedi <memxor@gmail.com>
To: Martin KaFai Lau <martin.lau@linux.dev>
To: Song Liu <song@kernel.org>
To: Yonghong Song <yonghong.song@linux.dev>
To: Jiri Olsa <jolsa@kernel.org>
To: Emil Tsalapatis <emil@etsalapatis.com>
To: Ihor Solodrai <ihor.solodrai@linux.dev>
To: John Fastabend <john.fastabend@gmail.com>
To: Stanislav Fomichev <sdf@fomichev.me>
To: Shuah Khan <shuah@kernel.org>
Cc: linux-kernel@vger.kernel.org
Cc: netdev@vger.kernel.org
Cc: bpf@vger.kernel.org
Cc: linux-kselftest@vger.kernel.org
To: Stanislav Fomichev <sdf@fomichev.me>
Cc: david.laight.linux@gmail.com

Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v2:
- Treat a negative optlen as 4 instead of rejecting it with
  -EINVAL as v1 did, per David Laight and Stanislav Fomichev. 
- Fix sashiko findings
- Link to v1: https://patch.msgid.link/20260925-sockopt_expand_out_v2-v1-0-c3ef2e3bb5c0@debian.org

---
Breno Leitao (7):
      ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt()
      net: cap sockopt_expand_out() at MAX_RW_COUNT
      ipv6: mcast: convert ip6_mc_msfget() to sockopt_t
      ipv4: igmp: convert ip_mc_gsfget() to sockopt_t
      ipv4: convert do_ip_getsockopt() to sockopt_t
      ipv6: convert do_ipv6_getsockopt() to sockopt_t
      selftests: net: getsockopt_iter: cover ip and ipv6

 include/linux/igmp.h                          |   2 +-
 include/linux/mroute.h                        |   4 +-
 include/linux/mroute6.h                       |   5 +-
 include/linux/net.h                           |  27 +-
 include/net/ip.h                              |   3 +-
 include/net/ipv6.h                            |   4 +-
 net/core/filter.c                             |  36 ++-
 net/ipv4/igmp.c                               |  23 +-
 net/ipv4/ip_sockglue.c                        | 118 ++++-----
 net/ipv4/ipmr.c                               |  11 +-
 net/ipv6/ip6mr.c                              |  13 +-
 net/ipv6/ipv6_sockglue.c                      | 117 +++++----
 net/ipv6/mcast.c                              |  19 +-
 tools/testing/selftests/net/getsockopt_iter.c | 359 ++++++++++++++++++++++++++
 14 files changed, 577 insertions(+), 164 deletions(-)
---
base-commit: aac26bee2287c88af5be5a5ff96d783b19a28790
change-id: 20260924-sockopt_expand_out_v2-ee4bc7ba4cea

Best regards,
--  
Breno Leitao <leitao@debian.org>


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH net-next v2 1/7] ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt()
  2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
@ 2026-10-09  8:53 ` Breno Leitao
  2026-10-10  9:12   ` netdev-bot+sashiko
  2026-10-09  8:53 ` [PATCH net-next v2 2/7] net: cap sockopt_expand_out() at MAX_RW_COUNT Breno Leitao
                   ` (5 subsequent siblings)
  6 siblings, 1 reply; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

do_ipv6_getsockopt() reads optlen into an int and never checks its
sign. Most of what consumes len compares it as unsigned, so a negative
optlen reads as a huge buffer, and the reply depends on the option.

The int options clamp it to their own size:

	len = min_t(unsigned int, sizeof(int), len);

so getsockopt(fd, SOL_IPV6, IPV6_TCLASS, buf, &len) with len set to -1
answers 4 bytes and reports 4, rather than failing.

The other options see a huge buffer. The sticky options reply with
their whole header, IPV6_2292PKTOPTIONS writes every pending control
message without a limit, and IPV6_PATHMTU and IPV6_FLOWLABEL_MGR pass
their length check.

do_ip_getsockopt() rejects a negative optlen, but doing that here would
turn a call that works today into an error for a caller that passes an
uninitialized length. Answer it as 4, as the int options do, for every
option.

That leaves the int options as they are. The sticky headers are cut to
4 bytes, IPV6_2292PKTOPTIONS gets a 4 byte control buffer, and
IPV6_PATHMTU and IPV6_FLOWLABEL_MGR fail with -EINVAL, as they do for
any other undersized buffer.

The conversion of this function to sockopt_t builds its buffers from
optlen, which the standard helpers reject when negative, so the answer
has to be settled before that.

Suggested-by: David Laight <david.laight.linux@gmail.com>
Signed-off-by: Breno Leitao <leitao@debian.org>
---
 net/ipv6/ipv6_sockglue.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/net/ipv6/ipv6_sockglue.c b/net/ipv6/ipv6_sockglue.c
index 4b3536571c9804..4091d1521cc09f 100644
--- a/net/ipv6/ipv6_sockglue.c
+++ b/net/ipv6/ipv6_sockglue.c
@@ -1002,6 +1002,11 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 
 	if (copy_from_sockptr(&len, optlen, sizeof(int)))
 		return -EFAULT;
+	/* Historic bug compatibility: the int options have always taken a
+	 * negative optlen as 4, so take it as 4 everywhere.
+	 */
+	if (len < 0)
+		len = 4;
 	switch (optname) {
 	case MCAST_MSFILTER:
 		if (in_compat_syscall())

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH net-next v2 2/7] net: cap sockopt_expand_out() at MAX_RW_COUNT
  2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
  2026-10-09  8:53 ` [PATCH net-next v2 1/7] ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt() Breno Leitao
@ 2026-10-09  8:53 ` Breno Leitao
  2026-10-10  9:12   ` netdev-bot+sashiko
  2026-10-09  8:53 ` [PATCH net-next v2 3/7] ipv6: mcast: convert ip6_mc_msfget() to sockopt_t Breno Leitao
                   ` (4 subsequent siblings)
  6 siblings, 1 reply; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

sockopt_expand_out() grows the output iterator up to INT_MAX, but its
callers rewind it with iov_iter_revert() once they know the real reply
size. iov_iter_revert() refuses to unroll past MAX_RW_COUNT:

	if (WARN_ON(unroll > MAX_RW_COUNT))
		return;

MAX_RW_COUNT is INT_MAX rounded down to a page boundary, so a caller
that expands into the (MAX_RW_COUNT, INT_MAX] range and then reverts
hits that WARN_ON, and the iterator is left where it was instead of
rewound. So does a caller whose optlen already covers such a size.

Cap the size at MAX_RW_COUNT instead, before the early return for an
optlen that already covers it, so a size that could never be reverted
later is rejected up front.

This was detected by sashiko, and I fits in this patchset/net-next,
given this doesn't seem to be a big deal.

Fixes: 0093f7db9c47 ("net: add sockopt_expand_out()")
Signed-off-by: Breno Leitao <leitao@debian.org>
---
 include/linux/net.h | 10 +++++++---
 1 file changed, 7 insertions(+), 3 deletions(-)

diff --git a/include/linux/net.h b/include/linux/net.h
index e2a866fbcfa5ea..461137bdc8817e 100644
--- a/include/linux/net.h
+++ b/include/linux/net.h
@@ -84,12 +84,16 @@ static inline int sockopt_init_user(sockopt_t *opt, char __user *optval,
  */
 static inline int sockopt_expand_out(sockopt_t *opt, size_t size)
 {
+	/* iov_iter_revert() refuses to unroll past MAX_RW_COUNT, so a size
+	 * beyond that could never be reverted back to the fixed part later,
+	 * even when optlen already covers it.
+	 */
+	if (size > MAX_RW_COUNT)
+		return -EINVAL;
+
 	if (size <= (size_t)opt->optlen)
 		return 0;
 
-	if (size > INT_MAX)
-		return -EINVAL;
-
 	/* Re-anchoring reads iter_out.ubuf, so the iterator has to be a user
 	 * buffer that nothing has written through yet.
 	 */

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH net-next v2 3/7] ipv6: mcast: convert ip6_mc_msfget() to sockopt_t
  2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
  2026-10-09  8:53 ` [PATCH net-next v2 1/7] ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt() Breno Leitao
  2026-10-09  8:53 ` [PATCH net-next v2 2/7] net: cap sockopt_expand_out() at MAX_RW_COUNT Breno Leitao
@ 2026-10-09  8:53 ` Breno Leitao
  2026-10-09  8:53 ` [PATCH net-next v2 4/7] ipv4: igmp: convert ip_mc_gsfget() " Breno Leitao
                   ` (3 subsequent siblings)
  6 siblings, 0 replies; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

MCAST_MSFILTER reads its reply through ip6_mc_msfget(), reached from
do_ipv6_getsockopt() and from nowhere else. Convert it, and build the
sockopt_t at the call site for as long as the caller still carries a
sockptr_t pair.

optlen only has to cover the fixed part, and the real reply size comes
from the gf_numsrc field inside it. Userspace relies on that, so
sockopt_expand_out() grows optval past optlen, for a user address only
and only far enough for the sources the socket has.

ip6_mc_msfget() now advances over the fixed part and writes the source
list through iter_out. Its callers rewind by the reply length they
already compute, which is exactly what the callee consumed, and land
back where they used to write: offset 0 for the native reply, gf_fmode
for the compat one.

ipv6_get_msfilter() and compat_ipv6_get_msfilter() only store *optlen
after their own copy succeeds, so a copy failure can't report a
changed length: -EINVAL, -EADDRNOTAVAIL and -EFAULT all leave the
caller's optlen word unchanged.

The old code stored *optlen before writing the fixed part, so an
unwritable optlen word stopped the fixed part from being written. Now
the fixed part is written first.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 include/net/ipv6.h       |  2 +-
 net/ipv6/ipv6_sockglue.c | 64 ++++++++++++++++++++++++++++++++----------------
 net/ipv6/mcast.c         | 19 +++++++++++---
 3 files changed, 59 insertions(+), 26 deletions(-)

diff --git a/include/net/ipv6.h b/include/net/ipv6.h
index 3de07e738538f7..9bb68d75890364 100644
--- a/include/net/ipv6.h
+++ b/include/net/ipv6.h
@@ -1195,7 +1195,7 @@ int ip6_mc_source(int add, int omode, struct sock *sk,
 int ip6_mc_msfilter(struct sock *sk, struct group_filter *gsf,
 		  struct sockaddr_storage *list);
 int ip6_mc_msfget(struct sock *sk, struct group_filter *gsf,
-		  sockptr_t optval, size_t ss_offset);
+		  sockopt_t *opt, size_t ss_offset);
 
 #ifdef CONFIG_PROC_FS
 int ac6_proc_init(struct net *net);
diff --git a/net/ipv6/ipv6_sockglue.c b/net/ipv6/ipv6_sockglue.c
index 4091d1521cc09f..b488391496af3c 100644
--- a/net/ipv6/ipv6_sockglue.c
+++ b/net/ipv6/ipv6_sockglue.c
@@ -922,48 +922,52 @@ static int ipv6_getsockopt_sticky(struct sock *sk, struct ipv6_txoptions *opt,
 	return len;
 }
 
-static int ipv6_get_msfilter(struct sock *sk, sockptr_t optval,
-			     sockptr_t optlen, int len)
+static int ipv6_get_msfilter(struct sock *sk, sockopt_t *opt)
 {
 	const int size0 = offsetof(struct group_filter, gf_slist_flex);
 	struct group_filter gsf;
-	int num;
+	int num, len;
 	int err;
 
-	if (len < size0)
+	if (opt->optlen < size0)
 		return -EINVAL;
-	if (copy_from_sockptr(&gsf, optval, size0))
+	if (copy_from_iter(&gsf, size0, &opt->iter_in) != size0)
 		return -EFAULT;
 	if (gsf.gf_group.ss_family != AF_INET6)
 		return -EADDRNOTAVAIL;
 	num = gsf.gf_numsrc;
 	sockopt_lock_sock(sk);
-	err = ip6_mc_msfget(sk, &gsf, optval, size0);
+	err = ip6_mc_msfget(sk, &gsf, opt, size0);
 	if (!err) {
 		if (num > gsf.gf_numsrc)
 			num = gsf.gf_numsrc;
 		len = GROUP_FILTER_SIZE(num);
-		if (copy_to_sockptr(optlen, &len, sizeof(int)) ||
-		    copy_to_sockptr(optval, &gsf, size0))
+
+		/* ip6_mc_msfget() consumed the whole reply; rewind to the
+		 * fixed part.
+		 */
+		iov_iter_revert(&opt->iter_out, len);
+		if (copy_to_iter(&gsf, size0, &opt->iter_out) != size0)
 			err = -EFAULT;
+		else
+			opt->optlen = len;
 	}
 	sockopt_release_sock(sk);
 	return err;
 }
 
-static int compat_ipv6_get_msfilter(struct sock *sk, sockptr_t optval,
-				    sockptr_t optlen, int len)
+static int compat_ipv6_get_msfilter(struct sock *sk, sockopt_t *opt)
 {
 	const int size0 = offsetof(struct compat_group_filter, gf_slist_flex);
 	struct compat_group_filter gf32;
 	struct group_filter gf;
 	int err;
-	int num;
+	int num, len;
 
-	if (len < size0)
+	if (opt->optlen < size0)
 		return -EINVAL;
 
-	if (copy_from_sockptr(&gf32, optval, size0))
+	if (copy_from_iter(&gf32, size0, &opt->iter_in) != size0)
 		return -EFAULT;
 	gf.gf_interface = gf32.gf_interface;
 	gf.gf_fmode = gf32.gf_fmode;
@@ -974,19 +978,23 @@ static int compat_ipv6_get_msfilter(struct sock *sk, sockptr_t optval,
 		return -EADDRNOTAVAIL;
 
 	sockopt_lock_sock(sk);
-	err = ip6_mc_msfget(sk, &gf, optval, size0);
+	err = ip6_mc_msfget(sk, &gf, opt, size0);
 	sockopt_release_sock(sk);
 	if (err)
 		return err;
 	if (num > gf.gf_numsrc)
 		num = gf.gf_numsrc;
 	len = GROUP_FILTER_SIZE(num) - (sizeof(gf)-sizeof(gf32));
-	if (copy_to_sockptr(optlen, &len, sizeof(int)) ||
-	    copy_to_sockptr_offset(optval, offsetof(struct compat_group_filter, gf_fmode),
-				   &gf.gf_fmode, sizeof(gf32.gf_fmode)) ||
-	    copy_to_sockptr_offset(optval, offsetof(struct compat_group_filter, gf_numsrc),
-				   &gf.gf_numsrc, sizeof(gf32.gf_numsrc)))
+
+	/* Rewind to gf_fmode, which gf_numsrc follows. */
+	iov_iter_revert(&opt->iter_out,
+			len - offsetof(struct compat_group_filter, gf_fmode));
+	if (copy_to_iter(&gf.gf_fmode, sizeof(gf32.gf_fmode),
+			 &opt->iter_out) != sizeof(gf32.gf_fmode) ||
+	    copy_to_iter(&gf.gf_numsrc, sizeof(gf32.gf_numsrc),
+			 &opt->iter_out) != sizeof(gf32.gf_numsrc))
 		return -EFAULT;
+	opt->optlen = len;
 	return 0;
 }
 
@@ -1009,9 +1017,23 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 		len = 4;
 	switch (optname) {
 	case MCAST_MSFILTER:
+	{
+		struct kvec kvec;
+		sockopt_t opt;
+		int err;
+
+		err = sockptr_to_sockopt(&opt, optval, optlen, &kvec);
+		if (err)
+			return err;
+
 		if (in_compat_syscall())
-			return compat_ipv6_get_msfilter(sk, optval, optlen, len);
-		return ipv6_get_msfilter(sk, optval, optlen, len);
+			err = compat_ipv6_get_msfilter(sk, &opt);
+		else
+			err = ipv6_get_msfilter(sk, &opt);
+		if (!err && copy_to_sockptr(optlen, &opt.optlen, sizeof(int)))
+			err = -EFAULT;
+		return err;
+	}
 	case IPV6_2292PKTOPTIONS:
 	{
 		struct msghdr msg;
diff --git a/net/ipv6/mcast.c b/net/ipv6/mcast.c
index ecef55f261890c..4ca2d77811f4eb 100644
--- a/net/ipv6/mcast.c
+++ b/net/ipv6/mcast.c
@@ -600,14 +600,14 @@ int ip6_mc_msfilter(struct sock *sk, struct group_filter *gsf,
 }
 
 int ip6_mc_msfget(struct sock *sk, struct group_filter *gsf,
-		  sockptr_t optval, size_t ss_offset)
+		  sockopt_t *opt, size_t ss_offset)
 {
 	struct ipv6_pinfo *inet6 = inet6_sk(sk);
 	const struct in6_addr *group;
 	struct ipv6_mc_socklist *pmc;
 	struct ip6_sf_socklist *psl;
+	int i, copycount, err;
 	unsigned int count;
-	int i, copycount;
 
 	group = &((struct sockaddr_in6 *)&gsf->gf_group)->sin6_addr;
 
@@ -629,6 +629,18 @@ int ip6_mc_msfget(struct sock *sk, struct group_filter *gsf,
 
 	copycount = min(count, gsf->gf_numsrc);
 	gsf->gf_numsrc = count;
+
+	/* The source list is sized by the gf_numsrc the caller left in optval,
+	 * not by optlen, which only has to cover the fixed part.
+	 */
+	err = sockopt_expand_out(opt, ss_offset +
+				 copycount * sizeof(struct sockaddr_storage));
+	if (err)
+		return err;
+
+	/* The caller fills the fixed part in once it knows gf_numsrc. */
+	iov_iter_advance(&opt->iter_out, ss_offset);
+
 	for (i = 0; i < copycount; i++) {
 		struct sockaddr_in6 *psin6;
 		struct sockaddr_storage ss;
@@ -637,9 +649,8 @@ int ip6_mc_msfget(struct sock *sk, struct group_filter *gsf,
 		memset(&ss, 0, sizeof(ss));
 		psin6->sin6_family = AF_INET6;
 		psin6->sin6_addr = psl->sl_addr[i];
-		if (copy_to_sockptr_offset(optval, ss_offset, &ss, sizeof(ss)))
+		if (copy_to_iter(&ss, sizeof(ss), &opt->iter_out) != sizeof(ss))
 			return -EFAULT;
-		ss_offset += sizeof(ss);
 	}
 	return 0;
 }

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH net-next v2 4/7] ipv4: igmp: convert ip_mc_gsfget() to sockopt_t
  2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
                   ` (2 preceding siblings ...)
  2026-10-09  8:53 ` [PATCH net-next v2 3/7] ipv6: mcast: convert ip6_mc_msfget() to sockopt_t Breno Leitao
@ 2026-10-09  8:53 ` Breno Leitao
  2026-10-09  8:53 ` [PATCH net-next v2 5/7] ipv4: convert do_ip_getsockopt() " Breno Leitao
                   ` (2 subsequent siblings)
  6 siblings, 0 replies; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

MCAST_MSFILTER reads its reply through ip_mc_gsfget(), reached from
do_ip_getsockopt() and from nowhere else. Convert it the same way as
the ipv6 side, and build the sockopt_t at the call site for as long as
the caller still carries a sockptr_t pair.

optlen here only has to cover the fixed part, and the real reply size
comes from the gf_numsrc field inside it. This is nasty, but userspace
relies on it, so sockopt_expand_out() preserves the same mechanism: it
grows optval only for a user address, and only far enough for the
sources the socket has.

ip_mc_gsfget() writes the source list, and its two callers write the
fixed part afterwards, at the head of optval. iter_out only moves
forward, so the callee advances over the fixed part rather than
addressing each source by offset.

The callers then rewind. The reply length they already compute is
exactly what the callee consumed, ss_offset plus the sources it wrote,
so both land back where they used to write: offset 0 for the native
reply, gf_fmode for the compat one. The bytes are the same.

ip_get_mcast_msfilter() and compat_ip_get_mcast_msfilter() only store
*optlen after their own copy succeeds, so a copy failure can't report
a changed length: -EINVAL, -EADDRNOTAVAIL and -EFAULT all leave the
caller's optlen word unchanged.

The old code stored *optlen before writing the fixed part, so an
unwritable optlen word stopped the fixed part from being written. Now
the fixed part is written first.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 include/linux/igmp.h   |  2 +-
 net/ipv4/igmp.c        | 20 +++++++++++++-----
 net/ipv4/ip_sockglue.c | 57 +++++++++++++++++++++++++++++++-------------------
 3 files changed, 52 insertions(+), 27 deletions(-)

diff --git a/include/linux/igmp.h b/include/linux/igmp.h
index e075611344ef3b..0a1abf3552d2bb 100644
--- a/include/linux/igmp.h
+++ b/include/linux/igmp.h
@@ -276,7 +276,7 @@ extern int ip_mc_msfilter(struct sock *sk, struct ip_msfilter *msf,int ifindex);
 extern int ip_mc_msfget(struct sock *sk, struct ip_msfilter *msf,
 			sockopt_t *opt);
 extern int ip_mc_gsfget(struct sock *sk, struct group_filter *gsf,
-			sockptr_t optval, size_t offset);
+			sockopt_t *opt, size_t offset);
 extern int ip_mc_sf_allow(const struct sock *sk, __be32 local, __be32 rmt,
 			  int dif, int sdif);
 extern void ip_mc_init_dev(struct in_device *);
diff --git a/net/ipv4/igmp.c b/net/ipv4/igmp.c
index 144fca158adcb0..d573c5bf8f038b 100644
--- a/net/ipv4/igmp.c
+++ b/net/ipv4/igmp.c
@@ -2775,9 +2775,9 @@ int ip_mc_msfget(struct sock *sk, struct ip_msfilter *msf, sockopt_t *opt)
 }
 
 int ip_mc_gsfget(struct sock *sk, struct group_filter *gsf,
-		 sockptr_t optval, size_t ss_offset)
+		 sockopt_t *opt, size_t ss_offset)
 {
-	int i, count, copycount;
+	int i, count, copycount, err;
 	struct sockaddr_in *psin;
 	__be32 addr;
 	struct ip_mc_socklist *pmc;
@@ -2805,6 +2805,18 @@ int ip_mc_gsfget(struct sock *sk, struct group_filter *gsf,
 	count = psl ? psl->sl_count : 0;
 	copycount = count < gsf->gf_numsrc ? count : gsf->gf_numsrc;
 	gsf->gf_numsrc = count;
+
+	/* The source list is sized by the gf_numsrc the caller left in optval,
+	 * not by optlen, which only has to cover the fixed part.
+	 */
+	err = sockopt_expand_out(opt, ss_offset +
+				 copycount * sizeof(struct sockaddr_storage));
+	if (err)
+		return err;
+
+	/* The caller fills the fixed part in once it knows gf_numsrc. */
+	iov_iter_advance(&opt->iter_out, ss_offset);
+
 	for (i = 0; i < copycount; i++) {
 		struct sockaddr_storage ss;
 
@@ -2812,10 +2824,8 @@ int ip_mc_gsfget(struct sock *sk, struct group_filter *gsf,
 		memset(&ss, 0, sizeof(ss));
 		psin->sin_family = AF_INET;
 		psin->sin_addr.s_addr = psl->sl_addr[i];
-		if (copy_to_sockptr_offset(optval, ss_offset,
-					   &ss, sizeof(ss)))
+		if (copy_to_iter(&ss, sizeof(ss), &opt->iter_out) != sizeof(ss))
 			return -EFAULT;
-		ss_offset += sizeof(ss);
 	}
 	return 0;
 }
diff --git a/net/ipv4/ip_sockglue.c b/net/ipv4/ip_sockglue.c
index e06c1f48ecad6e..1b7461f1569c32 100644
--- a/net/ipv4/ip_sockglue.c
+++ b/net/ipv4/ip_sockglue.c
@@ -1442,45 +1442,46 @@ static bool getsockopt_needs_rtnl(int optname)
 	return false;
 }
 
-static int ip_get_mcast_msfilter(struct sock *sk, sockptr_t optval,
-				 sockptr_t optlen, int len)
+static int ip_get_mcast_msfilter(struct sock *sk, sockopt_t *opt)
 {
 	const int size0 = offsetof(struct group_filter, gf_slist_flex);
 	struct group_filter gsf;
 	int num, gsf_size;
 	int err;
 
-	if (len < size0)
+	if (opt->optlen < size0)
 		return -EINVAL;
-	if (copy_from_sockptr(&gsf, optval, size0))
+	if (copy_from_iter(&gsf, size0, &opt->iter_in) != size0)
 		return -EFAULT;
 
 	num = gsf.gf_numsrc;
-	err = ip_mc_gsfget(sk, &gsf, optval,
+	err = ip_mc_gsfget(sk, &gsf, opt,
 			   offsetof(struct group_filter, gf_slist_flex));
 	if (err)
 		return err;
 	if (gsf.gf_numsrc < num)
 		num = gsf.gf_numsrc;
 	gsf_size = GROUP_FILTER_SIZE(num);
-	if (copy_to_sockptr(optlen, &gsf_size, sizeof(int)) ||
-	    copy_to_sockptr(optval, &gsf, size0))
+
+	/* ip_mc_gsfget() consumed the whole reply; rewind to the fixed part. */
+	iov_iter_revert(&opt->iter_out, gsf_size);
+	if (copy_to_iter(&gsf, size0, &opt->iter_out) != size0)
 		return -EFAULT;
+	opt->optlen = gsf_size;
 	return 0;
 }
 
-static int compat_ip_get_mcast_msfilter(struct sock *sk, sockptr_t optval,
-					sockptr_t optlen, int len)
+static int compat_ip_get_mcast_msfilter(struct sock *sk, sockopt_t *opt)
 {
 	const int size0 = offsetof(struct compat_group_filter, gf_slist_flex);
 	struct compat_group_filter gf32;
 	struct group_filter gf;
-	int num;
+	int num, len;
 	int err;
 
-	if (len < size0)
+	if (opt->optlen < size0)
 		return -EINVAL;
-	if (copy_from_sockptr(&gf32, optval, size0))
+	if (copy_from_iter(&gf32, size0, &opt->iter_in) != size0)
 		return -EFAULT;
 
 	gf.gf_interface = gf32.gf_interface;
@@ -1488,19 +1489,23 @@ static int compat_ip_get_mcast_msfilter(struct sock *sk, sockptr_t optval,
 	num = gf.gf_numsrc = gf32.gf_numsrc;
 	gf.gf_group = gf32.gf_group;
 
-	err = ip_mc_gsfget(sk, &gf, optval,
+	err = ip_mc_gsfget(sk, &gf, opt,
 			   offsetof(struct compat_group_filter, gf_slist_flex));
 	if (err)
 		return err;
 	if (gf.gf_numsrc < num)
 		num = gf.gf_numsrc;
 	len = GROUP_FILTER_SIZE(num) - (sizeof(gf) - sizeof(gf32));
-	if (copy_to_sockptr(optlen, &len, sizeof(int)) ||
-	    copy_to_sockptr_offset(optval, offsetof(struct compat_group_filter, gf_fmode),
-				   &gf.gf_fmode, sizeof(gf.gf_fmode)) ||
-	    copy_to_sockptr_offset(optval, offsetof(struct compat_group_filter, gf_numsrc),
-				   &gf.gf_numsrc, sizeof(gf.gf_numsrc)))
+
+	/* Rewind to gf_fmode, which gf_numsrc follows. */
+	iov_iter_revert(&opt->iter_out,
+			len - offsetof(struct compat_group_filter, gf_fmode));
+	if (copy_to_iter(&gf.gf_fmode, sizeof(gf32.gf_fmode),
+			 &opt->iter_out) != sizeof(gf32.gf_fmode) ||
+	    copy_to_iter(&gf.gf_numsrc, sizeof(gf32.gf_numsrc),
+			 &opt->iter_out) != sizeof(gf32.gf_numsrc))
 		return -EFAULT;
+	opt->optlen = len;
 	return 0;
 }
 
@@ -1727,12 +1732,22 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 		goto out;
 	}
 	case MCAST_MSFILTER:
+	{
+		struct kvec kvec;
+		sockopt_t opt;
+
+		err = sockptr_to_sockopt(&opt, optval, optlen, &kvec);
+		if (err)
+			goto out;
+
 		if (in_compat_syscall())
-			err = compat_ip_get_mcast_msfilter(sk, optval, optlen,
-							   len);
+			err = compat_ip_get_mcast_msfilter(sk, &opt);
 		else
-			err = ip_get_mcast_msfilter(sk, optval, optlen, len);
+			err = ip_get_mcast_msfilter(sk, &opt);
+		if (!err && copy_to_sockptr(optlen, &opt.optlen, sizeof(int)))
+			err = -EFAULT;
 		goto out;
+	}
 	case IP_PROTOCOL:
 		val = inet_sk(sk)->inet_num;
 		break;

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH net-next v2 5/7] ipv4: convert do_ip_getsockopt() to sockopt_t
  2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
                   ` (3 preceding siblings ...)
  2026-10-09  8:53 ` [PATCH net-next v2 4/7] ipv4: igmp: convert ip_mc_gsfget() " Breno Leitao
@ 2026-10-09  8:53 ` Breno Leitao
  2026-10-10  9:12   ` netdev-bot+sashiko
  2026-10-09  8:53 ` [PATCH net-next v2 6/7] ipv6: convert do_ipv6_getsockopt() " Breno Leitao
  2026-10-09  8:53 ` [PATCH net-next v2 7/7] selftests: net: getsockopt_iter: cover ip and ipv6 Breno Leitao
  6 siblings, 1 reply; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

Convert the SOL_IP switch and ip_mroute_getsockopt() to sockopt_t, its
last leaf on a sockptr_t pair. The IP_MSFILTER and MCAST_MSFILTER
bridges go away.

ip_getsockopt() builds the sockopt_t with sockopt_init_user() and
writes optlen back unconditionally

The write-back also happens before the netfilter fallback, so an
unwritable optlen turns any error into -EFAULT and skips the fallback.

This extra code in sol_ip_sockopt() will go away when we convert
sol_ip_sockopt() to talk sockopt as well, but for now we are converting
the leaves.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 include/linux/mroute.h |  4 +--
 include/net/ip.h       |  3 +-
 net/core/filter.c      | 18 ++++++++---
 net/ipv4/igmp.c        |  3 +-
 net/ipv4/ip_sockglue.c | 87 +++++++++++++++++++-------------------------------
 net/ipv4/ipmr.c        | 11 +++----
 6 files changed, 56 insertions(+), 70 deletions(-)

diff --git a/include/linux/mroute.h b/include/linux/mroute.h
index 4c5003afee6c51..c1c21e61c81c37 100644
--- a/include/linux/mroute.h
+++ b/include/linux/mroute.h
@@ -17,7 +17,7 @@ static inline int ip_mroute_opt(int opt)
 }
 
 int ip_mroute_setsockopt(struct sock *, int, sockptr_t, unsigned int);
-int ip_mroute_getsockopt(struct sock *, int, sockptr_t, sockptr_t);
+int ip_mroute_getsockopt(struct sock *sk, int optname, sockopt_t *opt);
 int ipmr_ioctl(struct sock *sk, int cmd, void *arg);
 int ipmr_compat_ioctl(struct sock *sk, unsigned int cmd, void __user *arg);
 int ip_mr_init(void);
@@ -31,7 +31,7 @@ static inline int ip_mroute_setsockopt(struct sock *sock, int optname,
 }
 
 static inline int ip_mroute_getsockopt(struct sock *sk, int optname,
-				       sockptr_t optval, sockptr_t optlen)
+				       sockopt_t *opt)
 {
 	return -ENOPROTOOPT;
 }
diff --git a/include/net/ip.h b/include/net/ip.h
index 604018cbf0f417..ac6dc64d4ac177 100644
--- a/include/net/ip.h
+++ b/include/net/ip.h
@@ -826,8 +826,7 @@ int do_ip_setsockopt(struct sock *sk, int level, int optname, sockptr_t optval,
 		     unsigned int optlen);
 int ip_setsockopt(struct sock *sk, int level, int optname, sockptr_t optval,
 		  unsigned int optlen);
-int do_ip_getsockopt(struct sock *sk, int level, int optname,
-		     sockptr_t optval, sockptr_t optlen);
+int do_ip_getsockopt(struct sock *sk, int level, int optname, sockopt_t *sopt);
 int ip_getsockopt(struct sock *sk, int level, int optname, char __user *optval,
 		  int __user *optlen);
 int ip_ra_control(struct sock *sk, unsigned char on,
diff --git a/net/core/filter.c b/net/core/filter.c
index ba536be2915fb9..1bc194a32337ac 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -5737,10 +5737,20 @@ static int sol_ip_sockopt(struct sock *sk, int optname,
 		return -EINVAL;
 	}
 
-	if (getopt)
-		return do_ip_getsockopt(sk, SOL_IP, optname,
-					KERNEL_SOCKPTR(optval),
-					KERNEL_SOCKPTR(optlen));
+	if (getopt) {
+		struct kvec kvec;
+		sockopt_t opt;
+		int err;
+
+		err = sockptr_to_sockopt(&opt, KERNEL_SOCKPTR(optval),
+					 KERNEL_SOCKPTR(optlen), &kvec);
+		if (err)
+			return err;
+
+		err = do_ip_getsockopt(sk, SOL_IP, optname, &opt);
+		*optlen = opt.optlen;
+		return err;
+	}
 
 	return do_ip_setsockopt(sk, SOL_IP, optname,
 				KERNEL_SOCKPTR(optval), *optlen);
diff --git a/net/ipv4/igmp.c b/net/ipv4/igmp.c
index d573c5bf8f038b..bf00810cdd7a13 100644
--- a/net/ipv4/igmp.c
+++ b/net/ipv4/igmp.c
@@ -2763,12 +2763,13 @@ int ip_mc_msfget(struct sock *sk, struct ip_msfilter *msf, sockopt_t *opt)
 	if (err)
 		return err;
 
-	opt->optlen = msf_size;
 	if (copy_to_iter(msf, IP_MSFILTER_SIZE(0), &opt->iter_out) !=
 	    IP_MSFILTER_SIZE(0))
 		return -EFAULT;
 	if (len && copy_to_iter(psl->sl_addr, len, &opt->iter_out) != len)
 		return -EFAULT;
+
+	opt->optlen = msf_size;
 	return 0;
 done:
 	return err;
diff --git a/net/ipv4/ip_sockglue.c b/net/ipv4/ip_sockglue.c
index 1b7461f1569c32..90d0b5bee68b59 100644
--- a/net/ipv4/ip_sockglue.c
+++ b/net/ipv4/ip_sockglue.c
@@ -1509,8 +1509,7 @@ static int compat_ip_get_mcast_msfilter(struct sock *sk, sockopt_t *opt)
 	return 0;
 }
 
-int do_ip_getsockopt(struct sock *sk, int level, int optname,
-		     sockptr_t optval, sockptr_t optlen)
+int do_ip_getsockopt(struct sock *sk, int level, int optname, sockopt_t *sopt)
 {
 	struct inet_sock *inet = inet_sk(sk);
 	bool needs_rtnl = getsockopt_needs_rtnl(optname);
@@ -1521,10 +1520,9 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 		return -EOPNOTSUPP;
 
 	if (ip_mroute_opt(optname))
-		return ip_mroute_getsockopt(sk, optname, optval, optlen);
+		return ip_mroute_getsockopt(sk, optname, sopt);
 
-	if (copy_from_sockptr(&len, optlen, sizeof(int)))
-		return -EFAULT;
+	len = sopt->optlen;
 	if (len < 0)
 		return -EINVAL;
 
@@ -1620,16 +1618,15 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 		rcu_read_unlock();
 
 		if (opt->optlen == 0) {
-			len = 0;
-			return copy_to_sockptr(optlen, &len, sizeof(int));
+			sopt->optlen = 0;
+			return 0;
 		}
 
 		ip_options_undo(opt);
 
 		len = min_t(unsigned int, len, opt->optlen);
-		if (copy_to_sockptr(optlen, &len, sizeof(int)))
-			return -EFAULT;
-		if (copy_to_sockptr(optval, opt->__data, len))
+		sopt->optlen = len;
+		if (copy_to_iter(opt->__data, len, &sopt->iter_out) != len)
 			return -EFAULT;
 		return 0;
 	}
@@ -1653,12 +1650,12 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 		if (sk->sk_type != SOCK_STREAM)
 			return -ENOPROTOOPT;
 
-		if (optval.is_kernel) {
+		if (iov_iter_is_kvec(&sopt->iter_out)) {
 			msg.msg_control_is_user = false;
-			msg.msg_control = optval.kernel;
+			msg.msg_control = sopt->iter_out.kvec->iov_base;
 		} else {
 			msg.msg_control_is_user = true;
-			msg.msg_control_user = optval.user;
+			msg.msg_control_user = sopt->iter_out.ubuf;
 		}
 		msg.msg_controllen = len;
 		msg.msg_flags = in_compat_syscall() ? MSG_CMSG_COMPAT : 0;
@@ -1680,8 +1677,8 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 			int tos = READ_ONCE(inet->rcv_tos);
 			put_cmsg(&msg, SOL_IP, IP_TOS, sizeof(tos), &tos);
 		}
-		len -= msg.msg_controllen;
-		return copy_to_sockptr(optlen, &len, sizeof(int));
+		sopt->optlen = len - msg.msg_controllen;
+		return 0;
 	}
 	case IP_UNICAST_IF:
 		val = (__force int)htonl((__u32) READ_ONCE(inet->uc_index));
@@ -1692,9 +1689,8 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 		len = min_t(unsigned int, len, sizeof(struct in_addr));
 		addr.s_addr = READ_ONCE(inet->mc_addr);
 
-		if (copy_to_sockptr(optlen, &len, sizeof(int)))
-			return -EFAULT;
-		if (copy_to_sockptr(optval, &addr, len))
+		sopt->optlen = len;
+		if (copy_to_iter(&addr, len, &sopt->iter_out) != len)
 			return -EFAULT;
 		return 0;
 	}
@@ -1711,43 +1707,25 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 	case IP_MSFILTER:
 	{
 		struct ip_msfilter msf;
-		struct kvec kvec;
-		sockopt_t opt;
 
 		if (len < IP_MSFILTER_SIZE(0)) {
 			err = -EINVAL;
 			goto out;
 		}
-		if (copy_from_sockptr(&msf, optval, IP_MSFILTER_SIZE(0))) {
+		if (copy_from_iter(&msf, IP_MSFILTER_SIZE(0), &sopt->iter_in) !=
+		    IP_MSFILTER_SIZE(0)) {
 			err = -EFAULT;
 			goto out;
 		}
-		err = sockptr_to_sockopt(&opt, optval, optlen, &kvec);
-		if (err)
-			goto out;
-
-		err = ip_mc_msfget(sk, &msf, &opt);
-		if (!err && copy_to_sockptr(optlen, &opt.optlen, sizeof(int)))
-			err = -EFAULT;
+		err = ip_mc_msfget(sk, &msf, sopt);
 		goto out;
 	}
 	case MCAST_MSFILTER:
-	{
-		struct kvec kvec;
-		sockopt_t opt;
-
-		err = sockptr_to_sockopt(&opt, optval, optlen, &kvec);
-		if (err)
-			goto out;
-
 		if (in_compat_syscall())
-			err = compat_ip_get_mcast_msfilter(sk, &opt);
+			err = compat_ip_get_mcast_msfilter(sk, sopt);
 		else
-			err = ip_get_mcast_msfilter(sk, &opt);
-		if (!err && copy_to_sockptr(optlen, &opt.optlen, sizeof(int)))
-			err = -EFAULT;
+			err = ip_get_mcast_msfilter(sk, sopt);
 		goto out;
-	}
 	case IP_PROTOCOL:
 		val = inet_sk(sk)->inet_num;
 		break;
@@ -1759,16 +1737,14 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 copyval:
 	if (len < sizeof(int) && len > 0 && val >= 0 && val <= 255) {
 		unsigned char ucval = (unsigned char)val;
-		len = 1;
-		if (copy_to_sockptr(optlen, &len, sizeof(int)))
-			return -EFAULT;
-		if (copy_to_sockptr(optval, &ucval, 1))
+
+		sopt->optlen = 1;
+		if (copy_to_iter(&ucval, 1, &sopt->iter_out) != 1)
 			return -EFAULT;
 	} else {
 		len = min_t(unsigned int, sizeof(int), len);
-		if (copy_to_sockptr(optlen, &len, sizeof(int)))
-			return -EFAULT;
-		if (copy_to_sockptr(optval, &val, len))
+		sopt->optlen = len;
+		if (copy_to_iter(&val, len, &sopt->iter_out) != len)
 			return -EFAULT;
 	}
 	return 0;
@@ -1783,19 +1759,22 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
 int ip_getsockopt(struct sock *sk, int level,
 		  int optname, char __user *optval, int __user *optlen)
 {
+	sockopt_t opt;
 	int err;
 
-	err = do_ip_getsockopt(sk, level, optname,
-			       USER_SOCKPTR(optval), USER_SOCKPTR(optlen));
+	err = sockopt_init_user(&opt, optval, optlen);
+	if (err)
+		return err;
+
+	err = do_ip_getsockopt(sk, level, optname, &opt);
+	if (put_user(opt.optlen, optlen))
+		return -EFAULT;
 
 #ifdef CONFIG_NETFILTER
 	/* we need to exclude all possible ENOPROTOOPTs except default case */
 	if (err == -ENOPROTOOPT && optname != IP_PKTOPTIONS &&
 			!ip_mroute_opt(optname)) {
-		int len;
-
-		if (get_user(len, optlen))
-			return -EFAULT;
+		int len = opt.optlen;
 
 		err = nf_getsockopt(sk, PF_INET, optname, optval, &len);
 		if (err >= 0)
diff --git a/net/ipv4/ipmr.c b/net/ipv4/ipmr.c
index e4c51ca473019c..f249569b4bf8b9 100644
--- a/net/ipv4/ipmr.c
+++ b/net/ipv4/ipmr.c
@@ -1638,8 +1638,7 @@ int ipmr_sk_ioctl(struct sock *sk, unsigned int cmd, void __user *arg)
 }
 
 /* Getsock opt support for the multicast routing system. */
-int ip_mroute_getsockopt(struct sock *sk, int optname, sockptr_t optval,
-			 sockptr_t optlen)
+int ip_mroute_getsockopt(struct sock *sk, int optname, sockopt_t *opt)
 {
 	int olr;
 	int val;
@@ -1670,16 +1669,14 @@ int ip_mroute_getsockopt(struct sock *sk, int optname, sockptr_t optval,
 		return -ENOPROTOOPT;
 	}
 
-	if (copy_from_sockptr(&olr, optlen, sizeof(int)))
-		return -EFAULT;
+	olr = opt->optlen;
 	if (olr < 0)
 		return -EINVAL;
 
 	olr = min_t(unsigned int, olr, sizeof(int));
 
-	if (copy_to_sockptr(optlen, &olr, sizeof(int)))
-		return -EFAULT;
-	if (copy_to_sockptr(optval, &val, olr))
+	opt->optlen = olr;
+	if (copy_to_iter(&val, olr, &opt->iter_out) != olr)
 		return -EFAULT;
 	return 0;
 }

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH net-next v2 6/7] ipv6: convert do_ipv6_getsockopt() to sockopt_t
  2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
                   ` (4 preceding siblings ...)
  2026-10-09  8:53 ` [PATCH net-next v2 5/7] ipv4: convert do_ip_getsockopt() " Breno Leitao
@ 2026-10-09  8:53 ` Breno Leitao
  2026-10-10  9:12   ` netdev-bot+sashiko
  2026-10-09  8:53 ` [PATCH net-next v2 7/7] selftests: net: getsockopt_iter: cover ip and ipv6 Breno Leitao
  6 siblings, 1 reply; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

Convert the SOL_IPV6 switch and its two remaining sockptr_t leaves,
ip6_mroute_getsockopt() and ipv6_getsockopt_sticky().

ipv6_getsockopt() writes optlen back unconditionally, and
sol_ipv6_sockopt() builds the sockopt_t with sockptr_to_sockopt() and
stores optlen back as well.

ipv6_getsockopt() cannot use sockopt_init_user() as its IPv4 sibling
does, since that rejects a negative optlen and this level answers it as
4. See discussion in the previous patch/commit.

The write-back also happens before the netfilter fallback, so an
unwritable optlen turns any error into -EFAULT and skips the fallback.

The extra code in sol_ipv6_sockopt() will go away when
sol_ipv6_sockopt() receives sockopt_t, but, for now, it only moves the
leaves.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 include/linux/mroute6.h  |  5 ++-
 include/linux/net.h      | 17 +++++++--
 include/net/ipv6.h       |  2 +-
 net/core/filter.c        | 18 +++++++---
 net/ipv6/ip6mr.c         | 13 +++----
 net/ipv6/ipv6_sockglue.c | 90 ++++++++++++++++++++----------------------------
 6 files changed, 73 insertions(+), 72 deletions(-)

diff --git a/include/linux/mroute6.h b/include/linux/mroute6.h
index fddafdc168f733..ab4d206fb32065 100644
--- a/include/linux/mroute6.h
+++ b/include/linux/mroute6.h
@@ -27,7 +27,7 @@ struct sock;
 
 #ifdef CONFIG_IPV6_MROUTE
 extern int ip6_mroute_setsockopt(struct sock *, int, sockptr_t, unsigned int);
-extern int ip6_mroute_getsockopt(struct sock *, int, sockptr_t, sockptr_t);
+int ip6_mroute_getsockopt(struct sock *sk, int optname, sockopt_t *sopt);
 extern int ip6_mr_input(struct sk_buff *skb);
 extern int ip6mr_compat_ioctl(struct sock *sk, unsigned int cmd, void __user *arg);
 extern int ip6_mr_init(void);
@@ -42,8 +42,7 @@ static inline int ip6_mroute_setsockopt(struct sock *sock, int optname,
 }
 
 static inline
-int ip6_mroute_getsockopt(struct sock *sock,
-			  int optname, sockptr_t optval, sockptr_t optlen)
+int ip6_mroute_getsockopt(struct sock *sock, int optname, sockopt_t *sopt)
 {
 	return -ENOPROTOOPT;
 }
diff --git a/include/linux/net.h b/include/linux/net.h
index 461137bdc8817e..3fe84daaa5e2e2 100644
--- a/include/linux/net.h
+++ b/include/linux/net.h
@@ -50,6 +50,19 @@ typedef struct sockopt {
 int sockptr_to_sockopt(sockopt_t *opt, sockptr_t optval, sockptr_t optlen,
 		       struct kvec *kvec);
 
+/*
+ * Point a sockopt_t at a user-backed (optval, len) pair whose length is
+ * already known, e.g. because the caller applied its own optlen quirks
+ * before this point.
+ */
+static inline void sockopt_set_user(sockopt_t *opt, char __user *optval,
+				    int len)
+{
+	iov_iter_ubuf(&opt->iter_out, ITER_DEST, optval, len);
+	iov_iter_ubuf(&opt->iter_in, ITER_SOURCE, optval, len);
+	opt->optlen = len;
+}
+
 /*
  * Initialize a user-backed sockopt_t from the (optval, optlen) __user pair of
  * a getsockopt() callback. Used by transitional __user getsockopt wrappers
@@ -66,9 +79,7 @@ static inline int sockopt_init_user(sockopt_t *opt, char __user *optval,
 	if (len < 0)
 		return -EINVAL;
 
-	iov_iter_ubuf(&opt->iter_out, ITER_DEST, optval, len);
-	iov_iter_ubuf(&opt->iter_in, ITER_SOURCE, optval, len);
-	opt->optlen = len;
+	sockopt_set_user(opt, optval, len);
 
 	return 0;
 }
diff --git a/include/net/ipv6.h b/include/net/ipv6.h
index 9bb68d75890364..a1e1de7da8c70d 100644
--- a/include/net/ipv6.h
+++ b/include/net/ipv6.h
@@ -1142,7 +1142,7 @@ int do_ipv6_setsockopt(struct sock *sk, int level, int optname, sockptr_t optval
 int ipv6_setsockopt(struct sock *sk, int level, int optname, sockptr_t optval,
 		    unsigned int optlen);
 int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
-		       sockptr_t optval, sockptr_t optlen);
+		       sockopt_t *sopt);
 int ipv6_getsockopt(struct sock *sk, int level, int optname,
 		    char __user *optval, int __user *optlen);
 
diff --git a/net/core/filter.c b/net/core/filter.c
index 1bc194a32337ac..8606d5d9a3181b 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -5774,10 +5774,20 @@ static int sol_ipv6_sockopt(struct sock *sk, int optname,
 		return -EINVAL;
 	}
 
-	if (getopt)
-		return do_ipv6_getsockopt(sk, SOL_IPV6, optname,
-					  KERNEL_SOCKPTR(optval),
-					  KERNEL_SOCKPTR(optlen));
+	if (getopt) {
+		struct kvec kvec;
+		sockopt_t opt;
+		int err;
+
+		err = sockptr_to_sockopt(&opt, KERNEL_SOCKPTR(optval),
+					 KERNEL_SOCKPTR(optlen), &kvec);
+		if (err)
+			return err;
+
+		err = do_ipv6_getsockopt(sk, SOL_IPV6, optname, &opt);
+		*optlen = opt.optlen;
+		return err;
+	}
 
 	return do_ipv6_setsockopt(sk, SOL_IPV6, optname,
 				  KERNEL_SOCKPTR(optval), *optlen);
diff --git a/net/ipv6/ip6mr.c b/net/ipv6/ip6mr.c
index 36f117ad367081..9513258b95fb88 100644
--- a/net/ipv6/ip6mr.c
+++ b/net/ipv6/ip6mr.c
@@ -1893,8 +1893,7 @@ int ip6_mroute_setsockopt(struct sock *sk, int optname, sockptr_t optval,
  *	Getsock opt support for the multicast routing system.
  */
 
-int ip6_mroute_getsockopt(struct sock *sk, int optname, sockptr_t optval,
-			  sockptr_t optlen)
+int ip6_mroute_getsockopt(struct sock *sk, int optname, sockopt_t *sopt)
 {
 	int olr;
 	int val;
@@ -1925,16 +1924,12 @@ int ip6_mroute_getsockopt(struct sock *sk, int optname, sockptr_t optval,
 		return -ENOPROTOOPT;
 	}
 
-	if (copy_from_sockptr(&olr, optlen, sizeof(int)))
-		return -EFAULT;
-
-	olr = min_t(int, olr, sizeof(int));
+	olr = min_t(int, sopt->optlen, sizeof(int));
 	if (olr < 0)
 		return -EINVAL;
 
-	if (copy_to_sockptr(optlen, &olr, sizeof(int)))
-		return -EFAULT;
-	if (copy_to_sockptr(optval, &val, olr))
+	sopt->optlen = olr;
+	if (copy_to_iter(&val, olr, &sopt->iter_out) != olr)
 		return -EFAULT;
 	return 0;
 }
diff --git a/net/ipv6/ipv6_sockglue.c b/net/ipv6/ipv6_sockglue.c
index b488391496af3c..c9787579531f64 100644
--- a/net/ipv6/ipv6_sockglue.c
+++ b/net/ipv6/ipv6_sockglue.c
@@ -889,7 +889,7 @@ int ipv6_setsockopt(struct sock *sk, int level, int optname, sockptr_t optval,
 EXPORT_SYMBOL(ipv6_setsockopt);
 
 static int ipv6_getsockopt_sticky(struct sock *sk, struct ipv6_txoptions *opt,
-				  int optname, sockptr_t optval, int len)
+				  int optname, sockopt_t *sopt, int len)
 {
 	struct ipv6_opt_hdr *hdr;
 
@@ -917,7 +917,7 @@ static int ipv6_getsockopt_sticky(struct sock *sk, struct ipv6_txoptions *opt,
 		return 0;
 
 	len = min_t(unsigned int, len, ipv6_optlen(hdr));
-	if (copy_to_sockptr(optval, hdr, len))
+	if (copy_to_iter(hdr, len, &sopt->iter_out) != len)
 		return -EFAULT;
 	return len;
 }
@@ -998,42 +998,21 @@ static int compat_ipv6_get_msfilter(struct sock *sk, sockopt_t *opt)
 	return 0;
 }
 
-int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
-		       sockptr_t optval, sockptr_t optlen)
+int do_ipv6_getsockopt(struct sock *sk, int level, int optname, sockopt_t *sopt)
 {
 	struct ipv6_pinfo *np = inet6_sk(sk);
 	int len;
 	int val;
 
 	if (ip6_mroute_opt(optname))
-		return ip6_mroute_getsockopt(sk, optname, optval, optlen);
+		return ip6_mroute_getsockopt(sk, optname, sopt);
 
-	if (copy_from_sockptr(&len, optlen, sizeof(int)))
-		return -EFAULT;
-	/* Historic bug compatibility: the int options have always taken a
-	 * negative optlen as 4, so take it as 4 everywhere.
-	 */
-	if (len < 0)
-		len = 4;
+	len = sopt->optlen;
 	switch (optname) {
 	case MCAST_MSFILTER:
-	{
-		struct kvec kvec;
-		sockopt_t opt;
-		int err;
-
-		err = sockptr_to_sockopt(&opt, optval, optlen, &kvec);
-		if (err)
-			return err;
-
 		if (in_compat_syscall())
-			err = compat_ipv6_get_msfilter(sk, &opt);
-		else
-			err = ipv6_get_msfilter(sk, &opt);
-		if (!err && copy_to_sockptr(optlen, &opt.optlen, sizeof(int)))
-			err = -EFAULT;
-		return err;
-	}
+			return compat_ipv6_get_msfilter(sk, sopt);
+		return ipv6_get_msfilter(sk, sopt);
 	case IPV6_2292PKTOPTIONS:
 	{
 		struct msghdr msg;
@@ -1042,12 +1021,12 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 		if (sk->sk_type != SOCK_STREAM)
 			return -ENOPROTOOPT;
 
-		if (optval.is_kernel) {
+		if (iov_iter_is_kvec(&sopt->iter_out)) {
 			msg.msg_control_is_user = false;
-			msg.msg_control = optval.kernel;
+			msg.msg_control = sopt->iter_out.kvec->iov_base;
 		} else {
 			msg.msg_control_is_user = true;
-			msg.msg_control_user = optval.user;
+			msg.msg_control_user = sopt->iter_out.ubuf;
 		}
 		msg.msg_controllen = len;
 		msg.msg_flags = 0;
@@ -1098,8 +1077,8 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 				put_cmsg(&msg, SOL_IPV6, IPV6_FLOWINFO, sizeof(flowinfo), &flowinfo);
 			}
 		}
-		len -= msg.msg_controllen;
-		return copy_to_sockptr(optlen, &len, sizeof(int));
+		sopt->optlen = len - msg.msg_controllen;
+		return 0;
 	}
 	case IPV6_MTU:
 	{
@@ -1154,12 +1133,13 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 		sockopt_lock_sock(sk);
 		opt = rcu_dereference_protected(np->opt,
 						lockdep_sock_is_held(sk));
-		len = ipv6_getsockopt_sticky(sk, opt, optname, optval, len);
+		len = ipv6_getsockopt_sticky(sk, opt, optname, sopt, len);
 		sockopt_release_sock(sk);
 		/* check if ipv6_getsockopt_sticky() returns err code */
 		if (len < 0)
 			return len;
-		return copy_to_sockptr(optlen, &len, sizeof(int));
+		sopt->optlen = len;
+		return 0;
 	}
 
 	case IPV6_RECVHOPOPTS:
@@ -1213,9 +1193,8 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 		if (!mtuinfo.ip6m_mtu)
 			return -ENOTCONN;
 
-		if (copy_to_sockptr(optlen, &len, sizeof(int)))
-			return -EFAULT;
-		if (copy_to_sockptr(optval, &mtuinfo, len))
+		sopt->optlen = len;
+		if (copy_to_iter(&mtuinfo, len, &sopt->iter_out) != len)
 			return -EFAULT;
 
 		return 0;
@@ -1292,7 +1271,8 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 		if (len < sizeof(freq))
 			return -EINVAL;
 
-		if (copy_from_sockptr(&freq, optval, sizeof(freq)))
+		if (copy_from_iter(&freq, sizeof(freq), &sopt->iter_in) !=
+		    sizeof(freq))
 			return -EFAULT;
 
 		if (freq.flr_action != IPV6_FL_A_GET)
@@ -1307,9 +1287,8 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 		if (val < 0)
 			return val;
 
-		if (copy_to_sockptr(optlen, &len, sizeof(int)))
-			return -EFAULT;
-		if (copy_to_sockptr(optval, &freq, len))
+		sopt->optlen = len;
+		if (copy_to_iter(&freq, len, &sopt->iter_out) != len)
 			return -EFAULT;
 
 		return 0;
@@ -1367,9 +1346,8 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 		return -ENOPROTOOPT;
 	}
 	len = min_t(unsigned int, sizeof(int), len);
-	if (copy_to_sockptr(optlen, &len, sizeof(int)))
-		return -EFAULT;
-	if (copy_to_sockptr(optval, &val, len))
+	sopt->optlen = len;
+	if (copy_to_iter(&val, len, &sopt->iter_out) != len)
 		return -EFAULT;
 	return 0;
 }
@@ -1377,7 +1355,8 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
 int ipv6_getsockopt(struct sock *sk, int level, int optname,
 		    char __user *optval, int __user *optlen)
 {
-	int err;
+	sockopt_t sopt;
+	int len, err;
 
 	if (level == SOL_IP && sk->sk_type != SOCK_RAW)
 		return ip_getsockopt(sk, level, optname, optval, optlen);
@@ -1385,15 +1364,22 @@ int ipv6_getsockopt(struct sock *sk, int level, int optname,
 	if (level != SOL_IPV6)
 		return -ENOPROTOOPT;
 
-	err = do_ipv6_getsockopt(sk, level, optname,
-				 USER_SOCKPTR(optval), USER_SOCKPTR(optlen));
+	if (get_user(len, optlen))
+		return -EFAULT;
+	/* Historic bug compatibility: the int options have always taken a
+	 * negative optlen as 4, so take it as 4 everywhere.
+	 */
+	if (len < 0)
+		len = 4;
+	sockopt_set_user(&sopt, optval, len);
+
+	err = do_ipv6_getsockopt(sk, level, optname, &sopt);
+	if (put_user(sopt.optlen, optlen))
+		return -EFAULT;
 #ifdef CONFIG_NETFILTER
 	/* we need to exclude all possible ENOPROTOOPTs except default case */
 	if (err == -ENOPROTOOPT && optname != IPV6_2292PKTOPTIONS) {
-		int len;
-
-		if (get_user(len, optlen))
-			return -EFAULT;
+		int len = sopt.optlen;
 
 		err = nf_getsockopt(sk, PF_INET6, optname, optval, &len);
 		if (err >= 0)

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH net-next v2 7/7] selftests: net: getsockopt_iter: cover ip and ipv6
  2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
                   ` (5 preceding siblings ...)
  2026-10-09  8:53 ` [PATCH net-next v2 6/7] ipv6: convert do_ipv6_getsockopt() " Breno Leitao
@ 2026-10-09  8:53 ` Breno Leitao
  2026-10-10  9:12   ` netdev-bot+sashiko
  6 siblings, 1 reply; 13+ messages in thread
From: Breno Leitao @ 2026-10-09  8:53 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S. Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Ihor Solodrai, John Fastabend,
	Stanislav Fomichev, Shuah Khan, Eric Dumazet
  Cc: netdev, linux-kernel, bpf, linux-kselftest, david.laight.linux,
	Breno Leitao, kernel-team

Add an ip and an ipv6 fixture, pinning the returned length and errno
across buffer sizes, the branches that answer empty, an unknown
optname and the options dispatched before the switch. The clamp tests
also compare the bytes copied back, not just the length.

SOL_IP answers a sub-int buffer with one byte where SOL_IPV6 clamps the
int. IP_PKTOPTIONS and IPV6_2292PKTOPTIONS want a stream socket;
MRT_*/MRT6_* want a raw one.

Only the two hop-by-hop tests set IPV6_HOPOPTS, which needs
CAP_NET_RAW, so the other ipv6 tests run unprivileged.

The MRT_*/MRT6_* tests skip on a missing /proc/net/ip_mr_vif or
ip6_mr_vif, rather than on the ENOPROTOOPT the switch also answers for
an unrelated regression. A kernel with CONFIG_IP_MROUTE or
CONFIG_IPV6_MROUTE built in no longer hides that regression behind a
skip.

MRT_VERSION and MRT6_VERSION are spelled out; linux/mroute.h does not
coexist with netinet/in.h here.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 tools/testing/selftests/net/getsockopt_iter.c | 359 ++++++++++++++++++++++++++
 1 file changed, 359 insertions(+)

diff --git a/tools/testing/selftests/net/getsockopt_iter.c b/tools/testing/selftests/net/getsockopt_iter.c
index 6c2408df461232..252a46d44e522b 100644
--- a/tools/testing/selftests/net/getsockopt_iter.c
+++ b/tools/testing/selftests/net/getsockopt_iter.c
@@ -55,6 +55,13 @@
 #ifndef TCP_ULP
 #define TCP_ULP 31
 #endif
+/* linux/mroute.h does not coexist with netinet/in.h here. */
+#ifndef MRT_VERSION
+#define MRT_VERSION 206
+#endif
+#ifndef MRT6_VERSION
+#define MRT6_VERSION 206
+#endif
 
 /* ---------- netlink ---------- */
 
@@ -492,6 +499,358 @@ TEST_F(rawv6, bad_optname)
 	ASSERT_EQ(sizeof(val), optlen);
 }
 
+/* ---------- ip (SOL_IP) ---------- */
+
+FIXTURE(ip)
+{
+	int fd;
+};
+
+FIXTURE_SETUP(ip)
+{
+	/* a router alert option, so IP_OPTIONS has something to answer with */
+	static const unsigned char ipopts[4] = { 0x94, 0x04, 0x00, 0x00 };
+	int ttl = 42;
+
+	self->fd = socket(AF_INET, SOCK_DGRAM, 0);
+	if (self->fd < 0)
+		SKIP(return, "AF_INET dgram socket: %s", strerror(errno));
+
+	if (setsockopt(self->fd, SOL_IP, IP_TTL, &ttl, sizeof(ttl)) < 0)
+		SKIP(return, "set IP_TTL: %s", strerror(errno));
+
+	if (setsockopt(self->fd, SOL_IP, IP_OPTIONS, ipopts,
+		       sizeof(ipopts)) < 0)
+		SKIP(return, "set IP_OPTIONS: %s", strerror(errno));
+}
+
+FIXTURE_TEARDOWN(ip)
+{
+	if (self->fd >= 0)
+		close(self->fd);
+}
+
+TEST_F(ip, ttl_exact)
+{
+	socklen_t optlen = sizeof(int);
+	int val = 0;
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IP, IP_TTL, &val, &optlen));
+	ASSERT_EQ(sizeof(int), optlen);
+	ASSERT_EQ(42, val);
+}
+
+TEST_F(ip, ttl_oversize_clamped)
+{
+	socklen_t optlen = 64;
+	int buf[16] = {};
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IP, IP_TTL, buf, &optlen));
+	ASSERT_EQ(sizeof(int), optlen);
+	ASSERT_EQ(42, buf[0]);
+}
+
+/* SOL_IP answers a sub-int buffer with a single byte when the value fits
+ * in one, rather than clamping the int down.
+ */
+TEST_F(ip, ttl_single_byte)
+{
+	unsigned char buf[3] = {};
+	socklen_t optlen = sizeof(buf);
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IP, IP_TTL, buf, &optlen));
+	ASSERT_EQ(1, optlen);
+	ASSERT_EQ(42, buf[0]);
+}
+
+TEST_F(ip, ttl_zero_len)
+{
+	socklen_t optlen = 0;
+	int val;
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IP, IP_TTL, &val, &optlen));
+	ASSERT_EQ(0, optlen);
+}
+
+TEST_F(ip, negative_optlen)
+{
+	socklen_t optlen = (socklen_t)-1;
+	int val;
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IP, IP_TTL, &val, &optlen));
+	ASSERT_EQ(EINVAL, errno);
+}
+
+TEST_F(ip, options_roundtrip)
+{
+	static const unsigned char ipopts[4] = { 0x94, 0x04, 0x00, 0x00 };
+	unsigned char buf[40] = {};
+	socklen_t optlen = sizeof(buf);
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IP, IP_OPTIONS, buf, &optlen));
+	ASSERT_EQ(sizeof(ipopts), optlen);
+	ASSERT_EQ(0, memcmp(buf, ipopts, sizeof(ipopts)));
+}
+
+TEST_F(ip, options_undersize_clamped)
+{
+	static const unsigned char ipopts[2] = { 0x94, 0x04 };
+	unsigned char buf[2] = {};
+	socklen_t optlen = sizeof(buf);
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IP, IP_OPTIONS, buf, &optlen));
+	ASSERT_EQ(sizeof(buf), optlen);
+	ASSERT_EQ(0, memcmp(buf, ipopts, sizeof(buf)));
+}
+
+/* With no option set the reply is empty and the call still succeeds. */
+TEST_F(ip, options_absent)
+{
+	unsigned char buf[40] = {};
+	socklen_t optlen = sizeof(buf);
+	int fd;
+
+	fd = socket(AF_INET, SOCK_DGRAM, 0);
+	if (fd < 0)
+		SKIP(return, "AF_INET dgram socket: %s", strerror(errno));
+
+	ASSERT_EQ(0, getsockopt(fd, SOL_IP, IP_OPTIONS, buf, &optlen));
+	ASSERT_EQ(0, optlen);
+	close(fd);
+}
+
+TEST_F(ip, multicast_if_oversize_clamped)
+{
+	struct in_addr any = { .s_addr = INADDR_ANY };
+	socklen_t optlen = 64;
+	unsigned char buf[64];
+
+	memset(buf, 0xaa, sizeof(buf));
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IP, IP_MULTICAST_IF, buf,
+				&optlen));
+	ASSERT_EQ(sizeof(struct in_addr), optlen);
+	ASSERT_EQ(0, memcmp(buf, &any, sizeof(any)));
+	ASSERT_EQ(0xaa, buf[sizeof(any)]);
+}
+
+/* IP_PKTOPTIONS only answers on a stream socket. */
+TEST_F(ip, pktoptions_wrong_type)
+{
+	socklen_t optlen = 64;
+	char buf[64];
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IP, IP_PKTOPTIONS, buf,
+				 &optlen));
+	ASSERT_EQ(ENOPROTOOPT, errno);
+}
+
+/* The MRT_* options are dispatched ahead of the rest of the switch and
+ * want a raw IGMP socket. Without CONFIG_IP_MROUTE they are not
+ * dispatched at all and the switch also answers ENOPROTOOPT, so the
+ * config check has to come from somewhere other than this errno.
+ */
+TEST_F(ip, mroute_wrong_type)
+{
+	socklen_t optlen = sizeof(int);
+	int val;
+
+	if (access("/proc/net/ip_mr_vif", F_OK))
+		SKIP(return, "CONFIG_IP_MROUTE disabled");
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IP, MRT_VERSION, &val,
+				 &optlen));
+	ASSERT_EQ(EOPNOTSUPP, errno);
+}
+
+TEST_F(ip, bad_optname)
+{
+	socklen_t optlen = sizeof(int);
+	int val;
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IP, 0x7fff, &val, &optlen));
+	ASSERT_EQ(ENOPROTOOPT, errno);
+	ASSERT_EQ(sizeof(int), optlen);
+}
+
+/* ---------- ipv6 (SOL_IPV6) ---------- */
+
+/* an 8 byte hop-by-hop header, so the sticky options answer */
+static const unsigned char hopopt[8] = { 0, 0, 1, 4, 0, 0, 0, 0 };
+
+FIXTURE(ipv6)
+{
+	int fd;
+};
+
+FIXTURE_SETUP(ipv6)
+{
+	int hops = 42;
+
+	self->fd = socket(AF_INET6, SOCK_DGRAM, 0);
+	if (self->fd < 0)
+		SKIP(return, "AF_INET6 dgram socket: %s", strerror(errno));
+
+	if (setsockopt(self->fd, SOL_IPV6, IPV6_UNICAST_HOPS, &hops,
+		       sizeof(hops)) < 0)
+		SKIP(return, "set IPV6_UNICAST_HOPS: %s", strerror(errno));
+}
+
+FIXTURE_TEARDOWN(ipv6)
+{
+	if (self->fd >= 0)
+		close(self->fd);
+}
+
+TEST_F(ipv6, hops_exact)
+{
+	socklen_t optlen = sizeof(int);
+	int val = 0;
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_UNICAST_HOPS, &val,
+				&optlen));
+	ASSERT_EQ(sizeof(int), optlen);
+	ASSERT_EQ(42, val);
+}
+
+TEST_F(ipv6, hops_oversize_clamped)
+{
+	socklen_t optlen = 64;
+	int buf[16] = {};
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_UNICAST_HOPS, buf,
+				&optlen));
+	ASSERT_EQ(sizeof(int), optlen);
+	ASSERT_EQ(42, buf[0]);
+}
+
+/* Unlike SOL_IP, SOL_IPV6 clamps a sub-int buffer down to what fits
+ * rather than answering a single byte.
+ */
+TEST_F(ipv6, hops_sub_int)
+{
+	unsigned char buf[3];
+	socklen_t optlen = sizeof(buf);
+	int hops = 42;
+
+	memset(buf, 0xaa, sizeof(buf));
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_UNICAST_HOPS, buf,
+				&optlen));
+	ASSERT_EQ(sizeof(buf), optlen);
+	ASSERT_EQ(0, memcmp(buf, &hops, sizeof(buf)));
+}
+
+TEST_F(ipv6, hops_zero_len)
+{
+	socklen_t optlen = 0;
+	int val;
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_UNICAST_HOPS, &val,
+				&optlen));
+	ASSERT_EQ(0, optlen);
+}
+
+/* Unlike SOL_IP, a negative optlen has always behaved as 4 here. */
+TEST_F(ipv6, negative_optlen)
+{
+	socklen_t optlen = (socklen_t)-1;
+	int val = 0;
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_UNICAST_HOPS, &val,
+				&optlen));
+	ASSERT_EQ(sizeof(int), optlen);
+	ASSERT_EQ(42, val);
+}
+
+TEST_F(ipv6, hopopts_roundtrip)
+{
+	unsigned char buf[64] = {};
+	socklen_t optlen = sizeof(buf);
+
+	if (setsockopt(self->fd, SOL_IPV6, IPV6_HOPOPTS, hopopt,
+		       sizeof(hopopt)) < 0)
+		SKIP(return, "set IPV6_HOPOPTS: %s", strerror(errno));
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_HOPOPTS, buf,
+				&optlen));
+	ASSERT_EQ(sizeof(hopopt), optlen);
+	ASSERT_EQ(0, memcmp(buf, hopopt, sizeof(hopopt)));
+}
+
+TEST_F(ipv6, hopopts_undersize_clamped)
+{
+	unsigned char buf[4] = {};
+	socklen_t optlen = sizeof(buf);
+
+	if (setsockopt(self->fd, SOL_IPV6, IPV6_HOPOPTS, hopopt,
+		       sizeof(hopopt)) < 0)
+		SKIP(return, "set IPV6_HOPOPTS: %s", strerror(errno));
+
+	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_HOPOPTS, buf,
+				&optlen));
+	ASSERT_EQ(sizeof(buf), optlen);
+	ASSERT_EQ(0, memcmp(buf, hopopt, sizeof(buf)));
+}
+
+/* With no header set the reply is empty and the call still succeeds. */
+TEST_F(ipv6, hopopts_absent)
+{
+	unsigned char buf[64] = {};
+	socklen_t optlen = sizeof(buf);
+	int fd;
+
+	fd = socket(AF_INET6, SOCK_DGRAM, 0);
+	if (fd < 0)
+		SKIP(return, "AF_INET6 dgram socket: %s", strerror(errno));
+
+	ASSERT_EQ(0, getsockopt(fd, SOL_IPV6, IPV6_HOPOPTS, buf, &optlen));
+	ASSERT_EQ(0, optlen);
+	close(fd);
+}
+
+/* IPV6_PATHMTU wants room for the whole struct ip6_mtuinfo. */
+TEST_F(ipv6, pathmtu_undersize)
+{
+	socklen_t optlen = 8;
+	char buf[8];
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IPV6, IPV6_PATHMTU, buf,
+				 &optlen));
+	ASSERT_EQ(EINVAL, errno);
+}
+
+TEST_F(ipv6, pktoptions_wrong_type)
+{
+	socklen_t optlen = 64;
+	char buf[64];
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IPV6, IPV6_2292PKTOPTIONS, buf,
+				 &optlen));
+	ASSERT_EQ(ENOPROTOOPT, errno);
+}
+
+TEST_F(ipv6, mroute_wrong_type)
+{
+	socklen_t optlen = sizeof(int);
+	int val;
+
+	if (access("/proc/net/ip6_mr_vif", F_OK))
+		SKIP(return, "CONFIG_IPV6_MROUTE disabled");
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IPV6, MRT6_VERSION, &val,
+				 &optlen));
+	ASSERT_EQ(EOPNOTSUPP, errno);
+}
+
+TEST_F(ipv6, bad_optname)
+{
+	socklen_t optlen = sizeof(int);
+	int val;
+
+	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IPV6, 0x7fff, &val, &optlen));
+	ASSERT_EQ(ENOPROTOOPT, errno);
+	ASSERT_EQ(sizeof(int), optlen);
+}
+
 /* ---------- tls ---------- */
 
 FIXTURE(tls)

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH net-next v2 1/7] ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt()
  2026-10-09  8:53 ` [PATCH net-next v2 1/7] ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt() Breno Leitao
@ 2026-10-10  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10  9:12 UTC (permalink / raw)
  To: leitao
  Cc: dsahern, idosch, davem, kuba, pabeni, horms, ast, daniel, andrii,
	eddyz87, memxor, martin.lau, song, yonghong.song, jolsa, emil,
	ihor.solodrai, john.fastabend, sdf, shuah, edumazet, netdev,
	linux-kernel, bpf, linux-kselftest, david.laight.linux,
	kernel-team

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] This commit's message says a negative optlen is answered as 4
  'for every option', and its new comment says 'take it as 4 everywhere'.
- [Low] The commit fixes old behaviour where do_ipv6_getsockopt() treated
  a negative optlen as a huge buffer: IPV6_2292PKTOPTIONS wrote control…

--

ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt()

This patch makes do_ipv6_getsockopt() treat a negative optlen as 4. The
int options already behaved this way through their unsigned min_t()
clamp. Sticky headers and the IPV6_2292PKTOPTIONS control buffer are now
cut to 4 bytes, and IPV6_PATHMTU and IPV6_FLOWLABEL_MGR now fail with
-EINVAL, as preparation for the sockopt_t conversion.

> Suggested-by: David Laight <david.laight.linux@gmail.com>
> Signed-off-by: Breno Leitao <leitao@debian.org>

[Severity: Low]
The commit message describes old behaviour that this patch fixes:

- IPV6_2292PKTOPTIONS writes control messages without a limit.
- IPV6_PATHMTU and IPV6_FLOWLABEL_MGR pass their size checks.
- Sticky headers are copied in full.

There is no Fixes: tag or stable guidance, though. Should one be added,
or a note saying this is not meant for stable?

The code involved predates git (1da177e4c3f4). The old behaviour only
wrote to buffers that userspace supplied. This patch also changes results
that userspace can see. For example, IPV6_PATHMTU with a negative optlen
now returns -EINVAL.

Another patch in this series, "net: cap sockopt_expand_out() at
MAX_RW_COUNT", does carry a Fixes: tag.

> diff --git a/net/ipv6/ipv6_sockglue.c b/net/ipv6/ipv6_sockglue.c
> index 4b3536571c980..4091d1521cc09 100644
> --- a/net/ipv6/ipv6_sockglue.c
> +++ b/net/ipv6/ipv6_sockglue.c
> @@ -1002,6 +1002,11 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
>  
>  	if (copy_from_sockptr(&len, optlen, sizeof(int)))
>  		return -EFAULT;
> +	/* Historic bug compatibility: the int options have always taken a
> +	 * negative optlen as 4, so take it as 4 everywhere.
> +	 */
> +	if (len < 0)
> +		len = 4;

[Severity: Low]
Is "take it as 4 everywhere" accurate at this commit? The commit message
also says a negative optlen is answered as 4 "for every option".

The mroute dispatch at the top of do_ipv6_getsockopt() returns before
the new clamp runs:

net/ipv6/ipv6_sockglue.c:do_ipv6_getsockopt() {
    ...
	if (ip6_mroute_opt(optname))
		return ip6_mroute_getsockopt(sk, optname, optval, optlen);
    ...
}

ip6_mroute_getsockopt() then reads optlen again and rejects a negative
value:

net/ipv6/ip6mr.c:ip6_mroute_getsockopt() {
    ...
	olr = min_t(int, olr, sizeof(int));
	if (olr < 0)
		return -EINVAL;
    ...
}

So on a raw ICMPv6 socket, MRT6_VERSION, MRT6_PIM and MRT6_ASSERT still
fail with len = -1, even though they are int options. IPV6_TCLASS
answers 4 bytes in the same case.

This also seems to contradict the part of the comment that says "the int
options have always taken a negative optlen as 4". The MRT6_* int options
have always returned -EINVAL.

The clamp also only changes the local len. When do_ipv6_getsockopt()
returns -ENOPROTOOPT, ipv6_getsockopt() reads the raw optlen again before
it calls nf_getsockopt():

net/ipv6/ipv6_sockglue.c:ipv6_getsockopt() {
    ...
		if (get_user(len, optlen))
			return -EFAULT;

		err = nf_getsockopt(sk, PF_INET6, optname, optval, &len);
    ...
}

The ip6_tables and conntrack handlers reached from there seem to reject a
negative length themselves. So this looks like a mismatch in wording, not
a functional problem.

A later patch in this series, "ipv6: convert do_ipv6_getsockopt() to
sockopt_t", moves the clamp into ipv6_getsockopt(). There it runs before
both the mroute dispatch and the netfilter fallback, which closes the gap.

Could the comment and commit message here be changed to describe what
this commit actually does?

>  	switch (optname) {
>  	case MCAST_MSFILTER:
>  		if (in_compat_syscall())

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-sockopt_expand_out_v2-v2-0-8ac08c469ecb%40debian.org

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH net-next v2 2/7] net: cap sockopt_expand_out() at MAX_RW_COUNT
  2026-10-09  8:53 ` [PATCH net-next v2 2/7] net: cap sockopt_expand_out() at MAX_RW_COUNT Breno Leitao
@ 2026-10-10  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10  9:12 UTC (permalink / raw)
  To: leitao
  Cc: dsahern, idosch, davem, kuba, pabeni, horms, ast, daniel, andrii,
	eddyz87, memxor, martin.lau, song, yonghong.song, jolsa, emil,
	ihor.solodrai, john.fastabend, sdf, shuah, edumazet, netdev,
	linux-kernel, bpf, linux-kselftest, david.laight.linux,
	kernel-team

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Most of this concern does not hold once the whole series is
  applied.

--

net: cap sockopt_expand_out() at MAX_RW_COUNT

This patch moves the size check in sockopt_expand_out() so it runs before
the optlen early return. It also lowers the limit from INT_MAX to
MAX_RW_COUNT. Any size that iov_iter_revert() could not unroll later is
now rejected up front.

> sockopt_expand_out() grows the output iterator up to INT_MAX, but its
> callers rewind it with iov_iter_revert() once they know the real reply
> size. iov_iter_revert() refuses to unroll past MAX_RW_COUNT:

[ ... ]

> Fixes: 0093f7db9c47 ("net: add sockopt_expand_out()")

[Severity: Low]
Is the present tense accurate at this point in the series?

At this commit, ip_mc_msfget() in net/ipv4/igmp.c is the only caller of
sockopt_expand_out(). It writes forward through opt->iter_out and never
calls iov_iter_revert(). The same holds at 0093f7db9c47, so neither
commit can reach the WARN_ON in iov_iter_revert().

The callers that rewind come from later patches in this series: "ipv6:
mcast: convert ip6_mc_msfget() to sockopt_t" and "ipv4: igmp: convert
ip_mc_gsfget() to sockopt_t". For example:

ip_get_mcast_msfilter()
    iov_iter_revert(&opt->iter_out, gsf_size);

ipv6_get_msfilter()
    iov_iter_revert(&opt->iter_out, len);

The same applies to compat_ip_get_mcast_msfilter() and
compat_ipv6_get_msfilter(). Once those patches land, the new cap does
protect these callers.

Would it be more accurate to say the cap is for callers converted later
in the series? Another option is to move this patch next to those
conversions. The Fixes: tag pointing at 0093f7db9c47 has the same
problem.

Also, the header comment above sockopt_expand_out() in include/linux/net.h
still only describes the kernel-backed -EINVAL case:

 * the head of optval. Only a user buffer can be longer than the optlen the
 * caller declared, so a kernel-backed optval is refused with -EINVAL.

Could it mention the MAX_RW_COUNT rejection as well? The new comment
inside the function already explains it.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-sockopt_expand_out_v2-v2-0-8ac08c469ecb%40debian.org

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH net-next v2 5/7] ipv4: convert do_ip_getsockopt() to sockopt_t
  2026-10-09  8:53 ` [PATCH net-next v2 5/7] ipv4: convert do_ip_getsockopt() " Breno Leitao
@ 2026-10-10  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10  9:12 UTC (permalink / raw)
  To: leitao
  Cc: dsahern, idosch, davem, kuba, pabeni, horms, ast, daniel, andrii,
	eddyz87, memxor, martin.lau, song, yonghong.song, jolsa, emil,
	ihor.solodrai, john.fastabend, sdf, shuah, edumazet, netdev,
	linux-kernel, bpf, linux-kselftest, david.laight.linux,
	kernel-team

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] Undisclosed errno change on calls that already fail.
- [Low] The commit message leaves out two behaviour-relevant parts of the
  patch.

--

ipv4: convert do_ip_getsockopt() to sockopt_t

This converts do_ip_getsockopt() and ip_mroute_getsockopt() from a
sockptr_t pair to sockopt_t and removes the IP_MSFILTER and
MCAST_MSFILTER bridges. ip_getsockopt() now builds the sockopt_t with
sockopt_init_user() and always writes optlen back. sol_ip_sockopt()
converts its kernel buffers with sockptr_to_sockopt().

> diff --git a/net/ipv4/igmp.c b/net/ipv4/igmp.c
> index d573c5bf8f038..bf00810cdd7a1 100644
> --- a/net/ipv4/igmp.c
> +++ b/net/ipv4/igmp.c
> @@ -2763,12 +2763,13 @@ int ip_mc_msfget(struct sock *sk, struct ip_msfilter *msf, sockopt_t *opt)
>  	if (err)
>  		return err;
>  
> -	opt->optlen = msf_size;
>  	if (copy_to_iter(msf, IP_MSFILTER_SIZE(0), &opt->iter_out) !=
>  	    IP_MSFILTER_SIZE(0))
>  		return -EFAULT;
>  	if (len && copy_to_iter(psl->sl_addr, len, &opt->iter_out) != len)
>  		return -EFAULT;
> +
> +	opt->optlen = msf_size;
>  	return 0;
>  done:
>  	return err;

[Severity: Low]

This isn't a bug, but the commit message doesn't explain this hunk. The
move appears to be needed because ip_getsockopt() now always writes
optlen back. The old IP_MSFILTER bridge only wrote it back on success:

    err = ip_mc_msfget(sk, &msf, &opt);
    if (!err && copy_to_sockptr(optlen, &opt.optlen, sizeof(int)))
        err = -EFAULT;

Without the move, a fault while copying the reply would report msf_size
to userspace along with -EFAULT. The previous patch in the series
("ipv4: igmp: convert ip_mc_gsfget() to sockopt_t") explained the same
reasoning for MCAST_MSFILTER. Could this commit message say so as well?

There is also a user-visible fix here that the message doesn't mention.
The old IP_OPTIONS empty path and the old IP_PKTOPTIONS tail in
do_ip_getsockopt() returned the result of copy_to_sockptr() directly:

    return copy_to_sockptr(optlen, &len, sizeof(int));

For a user pointer, that result is the number of bytes not copied. So if
optlen was readable but not writable, getsockopt() could return 4
instead of -EFAULT. Now those paths return 0, and put_user() in
ip_getsockopt() turns the failed write into -EFAULT. Would it be worth
noting this in the changelog?

Also, this sentence in the commit message is missing its final period:

    ip_getsockopt() builds the sockopt_t with sockopt_init_user() and
    writes optlen back unconditionally

> diff --git a/net/ipv4/ip_sockglue.c b/net/ipv4/ip_sockglue.c
> index 1b7461f1569c3..90d0b5bee68b5 100644
> --- a/net/ipv4/ip_sockglue.c
> +++ b/net/ipv4/ip_sockglue.c

[ ... ]

> @@ -1783,19 +1759,22 @@ int do_ip_getsockopt(struct sock *sk, int level, int optname,
>  int ip_getsockopt(struct sock *sk, int level,
>  		  int optname, char __user *optval, int __user *optlen)
>  {
> +	sockopt_t opt;
>  	int err;
>  
> -	err = do_ip_getsockopt(sk, level, optname,
> -			       USER_SOCKPTR(optval), USER_SOCKPTR(optlen));
> +	err = sockopt_init_user(&opt, optval, optlen);
> +	if (err)
> +		return err;
> +
> +	err = do_ip_getsockopt(sk, level, optname, &opt);
> +	if (put_user(opt.optlen, optlen))
> +		return -EFAULT;

[Severity: Low]

Reading optlen in sockopt_init_user() before do_ip_getsockopt() runs
seems to change the errno for calls that were already going to fail.
Is that intended?

Before this patch, do_ip_getsockopt() returned before touching optlen in
these cases:

    if (level != SOL_IP)
        return -EOPNOTSUPP;

    if (ip_mroute_opt(optname))
        return ip_mroute_getsockopt(sk, optname, optval, optlen);

ip_mroute_getsockopt() did not read optlen either before returning
-EOPNOTSUPP (not a raw IGMP socket), -ENOENT (no mr table) or
-ENOPROTOOPT (unknown MRT_* option).

Any level other than the socket's own protocol level ends up here. For
example, tcp_getsockopt() does:

    if (level != SOL_TCP)
        return READ_ONCE(icsk->icsk_af_ops)->getsockopt(sk, level,
                                    optname, optval, optlen);

udp_getsockopt() and raw_getsockopt() do the same for levels other than
SOL_UDP and SOL_RAW.

With this patch, getsockopt(tcp4_fd, SOL_IPV6, X, buf, &len) with
len < 0 returns -EINVAL instead of -EOPNOTSUPP. The same applies to an
MRT_* option on a socket that is not raw IGMP. An unreadable optlen
(for example NULL) now returns -EFAULT.

The commit message only describes the unwritable-optlen case. Could it
mention these errno changes too?

The selftest added later in the series ("selftests: net: getsockopt_iter:
cover ip and ipv6") only checks MRT_VERSION returning EOPNOTSUPP with a
valid optlen, so it would not catch this. The code is unchanged at the
end of the series.

Also, these negative checks look unreachable now. In
do_ip_getsockopt():

    len = sopt->optlen;
    if (len < 0)
        return -EINVAL;

and in ip_mroute_getsockopt():

    olr = opt->optlen;
    if (olr < 0)
        return -EINVAL;

Both callers of do_ip_getsockopt() already reject a negative optlen:
ip_getsockopt() through sockopt_init_user() and sol_ip_sockopt()
through sockptr_to_sockopt(). ip_mroute_getsockopt() is only called
from do_ip_getsockopt(). Are these checks still needed?

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-sockopt_expand_out_v2-v2-0-8ac08c469ecb%40debian.org

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH net-next v2 6/7] ipv6: convert do_ipv6_getsockopt() to sockopt_t
  2026-10-09  8:53 ` [PATCH net-next v2 6/7] ipv6: convert do_ipv6_getsockopt() " Breno Leitao
@ 2026-10-10  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10  9:12 UTC (permalink / raw)
  To: leitao
  Cc: dsahern, idosch, davem, kuba, pabeni, horms, ast, daniel, andrii,
	eddyz87, memxor, martin.lau, song, yonghong.song, jolsa, emil,
	ihor.solodrai, john.fastabend, sdf, shuah, edumazet, netdev,
	linux-kernel, bpf, linux-kselftest, david.laight.linux,
	kernel-team

Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 3

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] ipv6_getsockopt() now reads optlen and replaces a negative
  value with 4 before it calls do_ipv6_getsockopt().
- [Low] When the caller passes a negative optlen, ipv6_getsockopt() sets
  sopt.optlen to 4 and then always runs `put_user(sopt.optlen, optlen)`,…
- [Low] The commit message explains why sockopt_init_user() cannot be
  used by pointing to 'See discussion in the previous patch/commit.' The…
- [Low] ipv6_getsockopt() now declares `int len, err;` at function scope
  (line 1359) and then declares `int len = sopt.optlen;` again inside…

--

ipv6: convert do_ipv6_getsockopt() to sockopt_t

This commit converts do_ipv6_getsockopt(), ip6_mroute_getsockopt() and
ipv6_getsockopt_sticky() from sockptr_t to sockopt_t. ipv6_getsockopt()
now reads optlen itself and treats a negative value as 4. It builds a
user-backed sockopt_t with a new sockopt_set_user() helper and always
writes optlen back, while sol_ipv6_sockopt() builds its sockopt_t with
sockptr_to_sockopt().

> ipv6_getsockopt() cannot use sockopt_init_user() as its IPv4 sibling
> does, since that rejects a negative optlen and this level answers it as
> 4. See discussion in the previous patch/commit.

[Severity: Low]

Which commit is "the previous patch/commit"? The patch right before this
one in the series is "ipv4: convert do_ip_getsockopt() to sockopt_t", and
it does not discuss negative optlen.

The reasoning appears to be in "ipv6: treat a negative optlen as 4 in
do_ipv6_getsockopt()", which is several patches earlier. Could the message
name that commit directly?

The message also doesn't mention the new sockopt_set_user() helper. It is
added to include/linux/net.h by splitting it out of sockopt_init_user().

> diff --git a/net/ipv6/ipv6_sockglue.c b/net/ipv6/ipv6_sockglue.c
> index b488391496af3..c9787579531f6 100644
> --- a/net/ipv6/ipv6_sockglue.c
> +++ b/net/ipv6/ipv6_sockglue.c

[ ... ]

> @@ -998,42 +998,21 @@ static int compat_ipv6_get_msfilter(struct sock *sk, sockopt_t *opt)
>  	return 0;
>  }
>  
> -int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
> -		       sockptr_t optval, sockptr_t optlen)
> +int do_ipv6_getsockopt(struct sock *sk, int level, int optname, sockopt_t *sopt)
>  {
>  	struct ipv6_pinfo *np = inet6_sk(sk);
>  	int len;
>  	int val;
>  
>  	if (ip6_mroute_opt(optname))
> -		return ip6_mroute_getsockopt(sk, optname, optval, optlen);
> +		return ip6_mroute_getsockopt(sk, optname, sopt);
>  
> -	if (copy_from_sockptr(&len, optlen, sizeof(int)))
> -		return -EFAULT;
> -	/* Historic bug compatibility: the int options have always taken a
> -	 * negative optlen as 4, so take it as 4 everywhere.
> -	 */
> -	if (len < 0)
> -		len = 4;
> +	len = sopt->optlen;

[Severity: Medium]

Before this patch, the negative-optlen-as-4 handling ran after the
ip6_mroute_opt() dispatch, so MRT6_VERSION, MRT6_PIM and MRT6_ASSERT never
saw it. Now that the handling lives in ipv6_getsockopt(), does it also
apply to the mroute options?

ipv6_getsockopt()
    if (len < 0)
        len = 4;
    sockopt_set_user(&sopt, optval, len);
  do_ipv6_getsockopt()
    ip6_mroute_getsockopt()
        olr = min_t(int, sopt->optlen, sizeof(int));
        if (olr < 0)
            return -EINVAL;

The old ip6_mroute_getsockopt() read the raw user optlen and returned
-EINVAL for a negative value. When the caller passes -1, olr is now 4.
On a raw ICMPv6 socket, getsockopt(SOL_IPV6, MRT6_VERSION) with
optlen = -1 now succeeds, writes 4 bytes into optval and sets optlen
to 4.

The other caller, sol_ipv6_sockopt(), goes through sockptr_to_sockopt().
That already rejects a negative length, so the olr < 0 check in
ip6_mroute_getsockopt() looks unreachable now.

The IPv4 side still rejects this input. ip_getsockopt() uses
sockopt_init_user(), which returns -EINVAL for len < 0, and
ip_mroute_getsockopt() keeps its own olr < 0 check. Is the difference
between IPv4 and IPv6 intended?

The comment moved into ipv6_getsockopt() says "the int options have
always taken a negative optlen as 4", which was not true for MRT6_*.

The order of operations also changes for these options. On a socket that
isn't raw ICMPv6, the old code returned -EOPNOTSUPP without touching
optlen. Now get_user() and put_user() on optlen run first, so an
unreadable or unwritable optlen gives -EFAULT instead.

Should the mroute options skip the negative-to-4 conversion? If not,
could the commit message describe this behavior change?

[ ... ]

> @@ -1377,7 +1355,8 @@ int do_ipv6_getsockopt(struct sock *sk, int level, int optname,
>  int ipv6_getsockopt(struct sock *sk, int level, int optname,
>  		    char __user *optval, int __user *optlen)
>  {
> -	int err;
> +	sockopt_t sopt;
> +	int len, err;
>  
>  	if (level == SOL_IP && sk->sk_type != SOCK_RAW)
>  		return ip_getsockopt(sk, level, optname, optval, optlen);
> @@ -1385,15 +1364,22 @@ int ipv6_getsockopt(struct sock *sk, int level, int optname,
>  	if (level != SOL_IPV6)
>  		return -ENOPROTOOPT;
>  
> -	err = do_ipv6_getsockopt(sk, level, optname,
> -				 USER_SOCKPTR(optval), USER_SOCKPTR(optlen));
> +	if (get_user(len, optlen))
> +		return -EFAULT;
> +	/* Historic bug compatibility: the int options have always taken a
> +	 * negative optlen as 4, so take it as 4 everywhere.
> +	 */
> +	if (len < 0)
> +		len = 4;
> +	sockopt_set_user(&sopt, optval, len);
> +
> +	err = do_ipv6_getsockopt(sk, level, optname, &sopt);
> +	if (put_user(sopt.optlen, optlen))
> +		return -EFAULT;

[Severity: Low]

When the caller passes a negative optlen, sopt.optlen is 4 at this point,
and put_user() runs whatever err is. Do failing calls now change the
caller's optlen from -1 to 4?

For example, IPV6_PATHMTU with optlen = -1 fails the
"len < sizeof(mtuinfo)" check and returns -EINVAL, and then
put_user(4, optlen) runs. The same happens for:

- -ENOTCONN from IPV6_MTU
- -EINVAL from IPV6_FLOWLABEL_MGR and MCAST_MSFILTER
- -EOPNOTSUPP and -ENOENT from the MRT6_* options
- the default -ENOPROTOOPT

In the parent commit, none of these paths wrote optlen. For MCAST_MSFILTER,
sockptr_to_sockopt() rejected a negative optlen and left optlen alone.
Earlier patches in the series described this as leaving "the caller's
optlen word unchanged".

The netfilter fallback below now gets 4 instead of the caller's original
value. If nf_getsockopt() fails, optlen has already been overwritten.

The commit message mentions the unconditional write-back. It doesn't say
that the stored value can differ from the caller's input on failure. Is
that intended?

>  #ifdef CONFIG_NETFILTER
>  	/* we need to exclude all possible ENOPROTOOPTs except default case */
>  	if (err == -ENOPROTOOPT && optname != IPV6_2292PKTOPTIONS) {
> -		int len;
> -
> -		if (get_user(len, optlen))
> -			return -EFAULT;
> +		int len = sopt.optlen;

[Severity: Low]

This isn't a bug, but this inner int len shadows the int len that this
patch adds at the top of ipv6_getsockopt(). Could it simply be
"len = sopt.optlen;"? W=2 builds with -Wshadow will warn here.

>  
>  		err = nf_getsockopt(sk, PF_INET6, optname, optval, &len);
>  		if (err >= 0)

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-sockopt_expand_out_v2-v2-0-8ac08c469ecb%40debian.org

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH net-next v2 7/7] selftests: net: getsockopt_iter: cover ip and ipv6
  2026-10-09  8:53 ` [PATCH net-next v2 7/7] selftests: net: getsockopt_iter: cover ip and ipv6 Breno Leitao
@ 2026-10-10  9:12   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10  9:12 UTC (permalink / raw)
  To: leitao
  Cc: dsahern, idosch, davem, kuba, pabeni, horms, ast, daniel, andrii,
	eddyz87, memxor, martin.lau, song, yonghong.song, jolsa, emil,
	ihor.solodrai, john.fastabend, sdf, shuah, edumazet, netdev,
	linux-kernel, bpf, linux-kselftest, david.laight.linux,
	kernel-team

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] In tools/testing/selftests/net/getsockopt_iter.c, the ipv6
  fixture's negative_optlen test only exercises IPV6_UNICAST_HOPS.
- [Low] The new TEST_F(ip, mroute_wrong_type) needs CONFIG_IP_MROUTE: it
  skips when /proc/net/ip_mr_vif is missing.

--

selftests: net: getsockopt_iter: cover ip and ipv6

This adds ip and ipv6 fixtures to getsockopt_iter. They check the
returned length and errno for several buffer sizes, for the branches
that answer empty, for an unknown optname, and for the MRT_*/MRT6_*
options dispatched before the switch. The MRT tests skip when
/proc/net/ip_mr_vif or /proc/net/ip6_mr_vif is missing.

> diff --git a/tools/testing/selftests/net/getsockopt_iter.c b/tools/testing/selftests/net/getsockopt_iter.c
> index 6c2408df46123..252a46d44e522 100644
> --- a/tools/testing/selftests/net/getsockopt_iter.c
> +++ b/tools/testing/selftests/net/getsockopt_iter.c

[ ... ]

> +/* The MRT_* options are dispatched ahead of the rest of the switch and
> + * want a raw IGMP socket. Without CONFIG_IP_MROUTE they are not
> + * dispatched at all and the switch also answers ENOPROTOOPT, so the
> + * config check has to come from somewhere other than this errno.
> + */
> +TEST_F(ip, mroute_wrong_type)
> +{
> +	socklen_t optlen = sizeof(int);
> +	int val;
> +
> +	if (access("/proc/net/ip_mr_vif", F_OK))
> +		SKIP(return, "CONFIG_IP_MROUTE disabled");

[Severity: Low]
Should tools/testing/selftests/net/config also get CONFIG_IP_MROUTE=y?

The config fragment already has this line:

    CONFIG_IPV6_MROUTE=y

It has no CONFIG_IP_MROUTE entry, though, and nothing in net/ selects
IP_MROUTE. So whether this test runs depends on the base config.
x86_64 and i386 defconfig enable it, but arm64 defconfig does not. On
an arm64 kselftest-merge build, /proc/net/ip_mr_vif is missing and
ip.mroute_wrong_type always skips, while ipv6.mroute_wrong_type runs.

The commit message says:

    A kernel with CONFIG_IP_MROUTE or CONFIG_IPV6_MROUTE built in no
    longer hides that regression behind a skip.

With the current fragment, the selftest config only builds such a
kernel for IPv6.

> +
> +	ASSERT_EQ(-1, getsockopt(self->fd, SOL_IP, MRT_VERSION, &val,
> +				 &optlen));
> +	ASSERT_EQ(EOPNOTSUPP, errno);
> +}

[ ... ]

> +/* Unlike SOL_IP, a negative optlen has always behaved as 4 here. */
> +TEST_F(ipv6, negative_optlen)
> +{
> +	socklen_t optlen = (socklen_t)-1;
> +	int val = 0;
> +
> +	ASSERT_EQ(0, getsockopt(self->fd, SOL_IPV6, IPV6_UNICAST_HOPS, &val,
> +				&optlen));

[Severity: Low]
Does this test cover the negative optlen change made earlier in this
series?

IPV6_UNICAST_HOPS is an int option. Before the series, it already
turned a negative optlen into 4 through:

    len = min_t(unsigned int, sizeof(int), len);

The earlier patch "ipv6: treat a negative optlen as 4 in
do_ipv6_getsockopt()" added this in ipv6_getsockopt() for every option:

net/ipv6/ipv6_sockglue.c:ipv6_getsockopt() {
    ...
	/* Historic bug compatibility: the int options have always taken a
	 * negative optlen as 4, so take it as 4 everywhere.
	 */
	if (len < 0)
		len = 4;
    ...
}

According to that patch's commit message, a negative optlen now cuts
sticky headers such as IPV6_HOPOPTS to 4 bytes. It also makes
IPV6_PATHMTU and IPV6_FLOWLABEL_MGR fail with -EINVAL.

hopopts_roundtrip, hopopts_undersize_clamped and pathmtu_undersize all
pass a positive optlen. If the clamp were removed or broken later,
ipv6.negative_optlen would still pass.

Could a negative optlen case be added for IPV6_HOPOPTS or IPV6_PATHMTU?

Also, the kernel comment limits "always" to the int options. The
comment above this test does not, so a reader may think SOL_IPV6 never
changed how it handles a negative optlen. Could the comment say this
only holds for the int options?

> +	ASSERT_EQ(sizeof(int), optlen);
> +	ASSERT_EQ(42, val);
> +}

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261009-sockopt_expand_out_v2-v2-0-8ac08c469ecb%40debian.org

^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-10-10  9:13 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09  8:53 [PATCH net-next v2 0/7] ipv4,ipv6: convert the getsockopt switches to sockopt_t Breno Leitao
2026-10-09  8:53 ` [PATCH net-next v2 1/7] ipv6: treat a negative optlen as 4 in do_ipv6_getsockopt() Breno Leitao
2026-10-10  9:12   ` netdev-bot+sashiko
2026-10-09  8:53 ` [PATCH net-next v2 2/7] net: cap sockopt_expand_out() at MAX_RW_COUNT Breno Leitao
2026-10-10  9:12   ` netdev-bot+sashiko
2026-10-09  8:53 ` [PATCH net-next v2 3/7] ipv6: mcast: convert ip6_mc_msfget() to sockopt_t Breno Leitao
2026-10-09  8:53 ` [PATCH net-next v2 4/7] ipv4: igmp: convert ip_mc_gsfget() " Breno Leitao
2026-10-09  8:53 ` [PATCH net-next v2 5/7] ipv4: convert do_ip_getsockopt() " Breno Leitao
2026-10-10  9:12   ` netdev-bot+sashiko
2026-10-09  8:53 ` [PATCH net-next v2 6/7] ipv6: convert do_ipv6_getsockopt() " Breno Leitao
2026-10-10  9:12   ` netdev-bot+sashiko
2026-10-09  8:53 ` [PATCH net-next v2 7/7] selftests: net: getsockopt_iter: cover ip and ipv6 Breno Leitao
2026-10-10  9:12   ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®