From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-lf1-f45.google.com (mail-lf1-f45.google.com [209.85.167.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2C13F35A928 for ; Fri, 30 Jan 2026 22:52:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.45 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769813544; cv=none; b=KSQiTBIJ0PjR84lGwrJrHXaDiznYLeR0+z5b7XCJdF23aR2b0mDoVszKMpmXIseoOLZ9dfO2lSfqseiAAJA2NBELf5Z7ogo0ByGTWvp66iCOvOS5sJwA036BSFDJPPubaIBFGOCDIxQBjo/Ngf6icymqq1mthpFPvhf0cp/v06I= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769813544; c=relaxed/simple; bh=KcUQ1q2ALIp6GQbsuyiFX7GWULQTAiWrMXNKUDp6gDA=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=OBQudJL8zF7yumxS9lCAVmHWXlaShhDP885/eVJY7A35zeA2l//B46PulA2ymjdnJWkfBURoDVda4Y5hlrTRO2TPW3/tS74q/sRDhAdjbfwwZ2fk5UKf+o+faGbDwsN9Hxh0MYodo9BXnPmv7qZtk/JjO10bViTu2gzFX4Of03Q= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BcWQdbTn; arc=none smtp.client-ip=209.85.167.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BcWQdbTn" Received: by mail-lf1-f45.google.com with SMTP id 2adb3069b0e04-59dd7bfeb8aso3028573e87.0 for ; Fri, 30 Jan 2026 14:52:22 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1769813541; x=1770418341; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:subject:cc:to:from:date:from:to:cc:subject:date :message-id:reply-to; bh=ZkVCe2Fm0X1zwm2m79SvkIhPjHfte2ufPAVZoEONEas=; b=BcWQdbTntSuTOZmQGvMUdCz6wKFwehzwoe+PgcxH+q05UV4xfQjVaJCYtSxN4+EgFn KURBdLRZh82aZ4KY1xY81NAU3xNhraKek52nm9enWwZIPOfn/M5XUR+/U2ccZsSSf0m9 nO3F16VE2r0NJCTm7ozYpmnkcI9eiq5xgykzlz9KvIqZQ6z37mK9Yi95m42wnb7h9ArG Av2b0Rz/SW/9bmLRglJti6dMHhIqB9VpLPg6ncHmAXUs0F0rpsRlPpP+2vRELlRh4tpW gexIx80e6olOL++dIhkV9t/DZLWIqwCXqV9EbJZOFzMVQ/fKaFG3nKmP1i2dODPqtx89 o1+Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1769813541; x=1770418341; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=ZkVCe2Fm0X1zwm2m79SvkIhPjHfte2ufPAVZoEONEas=; b=D1yYvRNixY9Cgf+26uc+pYhXIQ+54EOQlVTnsz+gskDj2/jl6/s1eM1zpVgZzyG0Rd /50M4yKIswU0yu5w3KCjy9JL8Y6R4lY4XWtSsby/AFFWfFOwl/mmNAsn/EtrNDVu1RBs qSY2D1tWQyH3fwYJXO2c3go3WL4BpScgjb6TjdyVFIwv4P7K7QzZ4Edn/rgIhVAWpOIl 7TCidMstFtuRj94ru6u42Y52nmjZEzk6OZ8yavrxSvmiA3HSR0aIA9p0M4bkdTBXBMTv 0cVQVG2TTCYtDpS8wrVjB2rIt63hjWT9X+7DqK29BFLVmagw+WASOiAnbgc58H1hgagR h/dw== X-Forwarded-Encrypted: i=1; AJvYcCUrbeqeXNCuvlIq5jGyn5yJyvc4dL0l0i4YnhrVzanJNcJl9ZEakFeTsQv8hV/VZA6eGgHhMMhlxdtn5Ao=@vger.kernel.org X-Gm-Message-State: AOJu0YxlM/H7JkIeFhzaV2ZWESyckkemN9Z21lRQPcryG+ck4Oit+5M9 FEzsKVPBJR3+TsAy5TEPzm/7jOL1StAGOkewv5A/VTUid3y2e7gLeB6ud+BTBQ== X-Gm-Gg: AZuq6aL+59ItstMiBO6eFCDzVq/kaSOnzohzAfJKoxszg9rCNEUXkfz3IKmIvU03+Rd JyIJqkpEbDuXULc+LsJZ6Ojjzb8seHFuWmXJS4ImMvHPHIqdsaG7LwHimW2HP5exFC7JHecM/tJ ykZw82VxunJO+3PG1xcMYNUhbjL3fUhCFeey3N/BrSQDMpnk9o0sAW+IqZ8v4sNyLZ75e6dxFov sHw0GSyJPLANi9IwubawhAy3H+ANRnFgXcl8WGTP4wWker11j+p7ZXHyNiLlBjyPiWuO1Zw3qeG Qp4Dz1L0ijA3h8oBINKSiTwQfEhyUiu1VtrkRX+N6YLdW8QBPdq9H1YXgMtymqTob3tJxgLgr/H 2KQbPZIyIbH/HTxdfJrEDBdyM2vMq+/IoM8QXvtmyLCWZPpvxFTM1Wa6Z3ML2YglKLx6hLEaal4 GcXibYAdTKLIGyaQaw1hmhozT3hBiJbDmbmsD5lEQ80ZVsDUlvPQqZ X-Received: by 2002:a05:6000:2004:b0:42f:b690:6788 with SMTP id ffacd0b85a97d-435f3a6baa6mr6239935f8f.10.1769806350269; Fri, 30 Jan 2026 12:52:30 -0800 (PST) Received: from pumpkin (82-69-66-36.dsl.in-addr.zen.co.uk. [82.69.66.36]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-482e267aad1sm21831845e9.15.2026.01.30.12.52.29 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 30 Jan 2026 12:52:29 -0800 (PST) Date: Fri, 30 Jan 2026 20:52:27 +0000 From: David Laight To: Breno Leitao Cc: "David S. Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Kuniyuki Iwashima , Willem de Bruijn , metze@samba.org, axboe@kernel.dk, Stanislav Fomichev , io-uring@vger.kernel.org, bpf@vger.kernel.org, netdev@vger.kernel.org, Linus Torvalds , linux-kernel@vger.kernel.org, kernel-team@meta.com Subject: Re: [PATCH net-next RFC 0/3] net: move .getsockopt away from __user buffers Message-ID: <20260130205227.6fb1d9ad@pumpkin> In-Reply-To: <20260130-getsockopt-v1-0-9154fcff6f95@debian.org> References: <20260130-getsockopt-v1-0-9154fcff6f95@debian.org> X-Mailer: Claws Mail 4.1.1 (GTK 3.24.38; arm-unknown-linux-gnueabihf) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Fri, 30 Jan 2026 10:46:16 -0800 Breno Leitao wrote: > Currently, .getsockopt callback cannot be called with kernel buffers > because it requires userspace addresses: > > int (*getsockopt)(struct socket *sock, int level, > int optname, char __user *optval, int __user *optlen); > > This prevents kernel callers (io_uring, BPF, etc) from using getsockopt > on levels other than SOL_SOCKET, since they pass kernel pointers rather > than __user pointers. I had thoughts about this as well. I think using iov_iter is over the top and may have measurable performance impact for some paths. I think the first thing to do is sort out 'optlen'. There is absolutely no reason for the user pointer being passed into all the per-protocol functions. (and the code that changes that use sockptr_t are just stupid...) The system call wrapper can do the user copies, it can also suppress the write if the value is unchanged (which matters with clac/slac). The obvious change would be to pass the length itself and make the return value -ERRNO or the size. The annoyance is the few places that want to return an error and change optlen. That might be best addresses by something like: #define GETSOCKOPT_RVAL(errval, size) (1 << 31 | (errval) << 20 | (size)) which would get picked in the rval < 0 path. It would also let 'return 0' mean 'don't change the size' requiring a special return for the one (or two?) places that want to set the size to zero and return success. The length passed should also be 'unsigned int' - with a check for negative values in the system call wrapper. (There are many broken drivers that treat negative lengths as 4.) There is not much point making the 'optval' parameter more than a structure of a user and kernel address - one of which will be NULL. (This is safer than sockptr_t's discriminant union.) You can't police the length because it is sometimes only the length of a header (and in some recent code as well). I have looked at some of this change - it is enormous. David > > Following Linus' suggestion [0], this series introduces a wrapper > around iov_iter (sockopt_t) and a temporary getsockopt_iter callback: > > typedef struct sockopt { > struct iov_iter iter; > int optlen; > } sockopt_t; > > Note: optlen was not suggested by Linus' but I believe it is needed, given > random values could be passed by protocols back to userspace. > > And the callback becomes: > > int (*getsockopt_iter)(struct socket *sock, int level, > int optname, sockopt_t *opt); > > The sockopt_t structure encapsulates: > - An iov_iter for reading/writing option data (works with both user > and kernel buffers) > - An optlen field for buffer size (input) and returned data size > (output) > > The plan is to enable getsockopt to leverage kernel buffers initially, > but then move .setsockopt from sockptr_t into this as well. > > This series: > > 1. Adds the sockopt_t type and getsockopt_iter callback to proto_ops > 2. Adds do_sock_getsockopt_iter() helper that prefers getsockopt_iter > 3. Converts one protocol (netlink) to use getsockopt_iter as a proof of > concept > > This is what I have in mind for this work stream, to make it more > digestible: > > * Keep the temporary getsockopt_iter callback allows protocols to > migrate gradually. > * Once all protocols have been converted, getsockopt can be removed and > getsockopt_iter renamed back to getsockopt with the new API. > * Once the protocols are converted, the SOL_SOCKET limitation in > io_uring_cmd_getsockopt() will be removed. > * Covert setsockopt() to also use a similar strategy, moving it away > from sockptr_t. > * Remove sockptr_t in the front end (do_sock_getsockopt(), > io_uring_cmd_getsockopt()) and start with sockopt_t (instead of > sockptr_t) in __sys_getsockopt() and io_uring_cmd_getsockopt() > > Link: https://lore.kernel.org/all/CAHk-=whmzrO-BMU=uSVXbuoLi-3tJsO=0kHj1BCPBE3F2kVhTA@mail.gmail.com/ [0] > --- > Breno Leitao (3): > net: add getsockopt_iter callback to proto_ops > net: prefer getsockopt_iter in do_sock_getsockopt > netlink: convert to getsockopt_iter > > include/linux/net.h | 19 +++++++++++++++++++ > net/netlink/af_netlink.c | 22 ++++++++++++---------- > net/socket.c | 42 +++++++++++++++++++++++++++++++++++++++--- > 3 files changed, 70 insertions(+), 13 deletions(-) > --- > base-commit: 4d310797262f0ddf129e76c2aad2b950adaf1fda > change-id: 20260130-getsockopt-9f36625eedcb > > Best regards, > -- > Breno Leitao > >