From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C1C8A4A64F4 for ; Mon, 21 Sep 2026 13:39:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789997999; cv=none; b=rx+YHG24zxQHx6xtbf96rjPj4d0LgklT6a5g6Pfkj+t9M/9t5VgVHwI5cpuE6wBsHnBDjppdtwBKE1FxTKWMRCB3ds97tPv7yddDpOQS3zjOJXniP4PShthj8+V3TLOw+1eaxXNnuoT/pzOltuMA/Oh6rHCe9QRMLiWXR2zn3Bk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789997999; c=relaxed/simple; bh=FAuNY7f+2IN9h+0fEp3doBi1UJ2umEK7JqOTw5NHGgE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XgaxsjE/D8crWOYFTgwaXJyhHmK/CMUOgbtdTtTNnVsvB1PhV3mwU7zHkFsszCo9/cZ+1obJnA1h/k26K9mb7i7MR9SW3LUkWzEDIy8IKw4pn2GDDNybcbBbRhMxpydWkByg7OSOcOhyZtZ7Zv8WwJ075xajcAIW6+GzAEJeYqo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Keh0fKHa; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Keh0fKHa" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49e79a408deso16965715e9.2 for ; Mon, 21 Sep 2026 06:39:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789997995; x=1790602795; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=FutSyws4EmX9/QkptnDh8TTYgLLGT7M1LNenj3p2Roc=; b=Keh0fKHaMyar4vCS+5xjzHGUiEWWO9NXVY+9gcea4IWUKxAcY/n9H6E2J9kzCi6stL Icd9ZHvK5Hu/7Ine5phjVDuKIyqJ9xudPb0cUbYNxdR2d4aNYFrTgsL3DjWZWxnninKq 5y7eV2+KupNbXrsaoPbS8vFfySoylJcS/0KLvgjo8YHwf1RPSCTk3YtVpp+laeVfg1cU gMGYztxpKbeMr1u2KsG2RlkyaqS1+5cFQhjrxwA36EbJKE/ATQSrBFek+xc3wLVST3I/ igpYEBgGtEG4JYq81lqIOzjnFcObs0yReSs5++RXP5SAmflGGISI3ENtCqbDIZo3ZtoR tliA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789997995; x=1790602795; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=FutSyws4EmX9/QkptnDh8TTYgLLGT7M1LNenj3p2Roc=; b=e2oGirq+pJu/7E9v+SKcXQvXS1lrIS+G8vAxuv/8lWqRl5qVXZ2gQlvLZLIcaV86lI Q+iAQy4M2ncRziMLTJHKMQJHi3cSjmTeMy6SyLqY3J+4j5+PY/XJiwvaxQtsISPtojOg Z8ZrMxR8nAtzxCd+pKUUH+AQI1oLezGjvk1T0CClkklDDipiNACPZ140ABDiHWEdNYF8 cMMLI6Y60JIEB4SMsKdjyLxsrOdhVTQFWhNyZR/r5vUEIJYYpUupKj9SMflXQ1zz9DXX cuAYA/sgj/GSnX3TiPfG8RDsoY7XPYe6jlrvXAHIKESSFPtYj/Z0hOiC/p/2vanaQd4J NCKQ== X-Forwarded-Encrypted: i=1; AKwUvBwCnGxGtkkrZfJvdRgE7wLPXwig5/9SlqS+rE/jd1Il2UpZLGGBIWR3XA0jTmDwNQQszJiWyU6JmcdnL6E=@vger.kernel.org X-Gm-Message-State: AFuF++mXa6K73r6UHeRoisAojyAfco40M4fYVdsBF+K8/sABWe0V1nCP BLrSpTSYuU+1v3EPPrrNhIFJgOyxNlYvUL2/nthLVSff+UinItisWD0w X-Gm-Gg: AYBFou0P8rzeRfB+0rEZ9kLbXbK2MGAGlsHm4udPjdDoZJ4dtwopR83os8IrdREwI10 79q2XSPYYK5xwGMmJ/kxCTfwqyv62LKrf+3itV02Td2bhUZKgFzNlOB2ashpLgMrQJYMkm3iOTK MJM7+fsDmqQZdcMSHS8x97zV1iA6zab7X+aZ0eYf2lLOYuZgBj0PLXcbP9tbb++UO8dFGsVnMNE GP6xF+6I6yKuikBFKSm7OsEcl8Je77JFogoR/58h4XmZekh/nNbwM5UeIXIYtkZbz2WPDjfqf4n 7Wiw88lpR8TQqYTEE3CbgF7/dcb+yZoB5zBEMBdRnAf0cQUYG2xbC3c7Ng1/R0mz4RFsPP/KpZf zT2CB8S8henc0Xm5lOM4NjmayPwTxOSE9pOBXD+S3n3HQjJMLikWqqkeF5xoMHm+T4NtGbvXauK Wd+i5EZB4ihcXCn9Jrq08xsuM+a4hZIFOt/dY1YrTRUngFzM7OVt6ggL9mZO1KclTk+Sg/rAxqw IKOJERGVtPcpoCA28AIF8M11X6qRCCDXKYKHw1cnF5BrP7I34EfYiDOa1U7U6ZDuL4uU79CXuAb RLMTNr8y5E5BXIXWt9qIrMk+Z3K7/fw= X-Received: by 2002:a05:600c:a013:b0:49e:6891:27a8 with SMTP id 5b1f17b1804b1-49fc574a9bcmr133446705e9.25.1789997994878; Mon, 21 Sep 2026 06:39:54 -0700 (PDT) Received: from 127.0.0.1localhost (82-132-213-26.dab.02.net. [82.132.213.26]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fc585920fsm617171605e9.4.2026.09.21.06.39.51 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 06:39:54 -0700 (PDT) From: Pavel Begunkov To: linux-block@vger.kernel.org Cc: asml.silence@gmail.com, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, linux-nvme@lists.infradead.org, linux-fsdevel@vger.kernel.org, io-uring@vger.kernel.org, Christoph Hellwig , Sumit Semwal , =?UTF-8?q?Christian=20K=C3=B6nig?= , Keith Busch , Sagi Grimberg , Alexander Viro , Christian Brauner , Jan Kara , Andrew Morton , Jens Axboe , Nitesh Shetty , Kanchan Joshi , Anuj Gupta , Tushar Gohad , William Power , Phil Cayton , Matthew Brost , Alasdair Kergon , Mike Snitzer , Mikulas Patocka , Benjamin Marzinski , dm-devel@lists.linux.dev Subject: [PATCH v6 10/13] io_uring/rsrc: extend buffer update Date: Mon, 21 Sep 2026 14:38:54 +0100 Message-ID: <947ff623f7e6844af4435f258ad939fa2192c0e4.1789997898.git.asml.silence@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit We need to pass more information to buffer registration than we can fit into a single struct iovec. This patch allows users to optionally pass struct io_uring_regbuf_desc. Apart from having more space for future use cases, it also introduces registration types. Currently, the type can be either of IO_REGBUF_TYPE_UADDR, which mirrors the iovec path, or IO_REGBUF_TYPE_EMPTY for leaving a buffer table slot empty. The next patch introduces a dmabuf backed type, and can be useful for other extensions like splicing a list of user addresses (i.e. iovec[]), interoperability with zcrx, kernel allocated memory like was brough up by Cristoph. Note, the type only represents a registration option, which is distinct from how io_uring internally stores it. The flags field is not used yet but always useful to have, e.g. we can encode read-only / write-only restrictions using it. Signed-off-by: Pavel Begunkov --- include/uapi/linux/io_uring.h | 27 +++++++++++++- io_uring/rsrc.c | 69 ++++++++++++++++++++++------------- 2 files changed, 69 insertions(+), 27 deletions(-) diff --git a/include/uapi/linux/io_uring.h b/include/uapi/linux/io_uring.h index 909fb7aea638..98b259901185 100644 --- a/include/uapi/linux/io_uring.h +++ b/include/uapi/linux/io_uring.h @@ -790,13 +790,38 @@ struct io_uring_rsrc_update { struct io_uring_rsrc_update2 { __u32 offset; - __u32 resv; + __u32 flags; __aligned_u64 data; __aligned_u64 tags; __u32 nr; __u32 resv2; }; +/* struct io_uring_rsrc_update2::flags */ +enum io_uring_rsrc_reg_flags { + /* + * Use the extended descriptor format for buffer updates, + * see struct io_uring_regbuf_desc + */ + IORING_RSRC_UPDATE_EXTENDED = 1U << 1, +}; + +/* Buffer registration type, passed in struct io_uring_regbuf_desc::type */ +enum io_uring_regbuf_type { + IO_REGBUF_TYPE_EMPTY, + IO_REGBUF_TYPE_UADDR, + + __IO_REGBUF_TYPE_MAX, +}; + +struct io_uring_regbuf_desc { + __u32 type; /* enum io_uring_regbuf_type */ + __u32 flags; + __u64 size; + __u64 uaddr; + __u64 __resv[7]; +}; + /* Skip updating fd indexes set to this value in the fd table */ #define IORING_REGISTER_FILES_SKIP (-2) diff --git a/io_uring/rsrc.c b/io_uring/rsrc.c index da767b34eb55..83a7ad9f5bc7 100644 --- a/io_uring/rsrc.c +++ b/io_uring/rsrc.c @@ -27,11 +27,6 @@ struct io_rsrc_update { u32 offset; }; -struct io_uring_regbuf_desc { - __u64 uaddr; - __u64 size; -}; - static struct io_rsrc_node *io_sqe_buffer_register(struct io_ring_ctx *ctx, struct io_uring_regbuf_desc *desc); @@ -90,9 +85,12 @@ static void io_iov_to_regbuf_desc(const struct iovec *iov, struct io_uring_regbuf_desc *desc) { *desc = (struct io_uring_regbuf_desc) { + .type = IO_REGBUF_TYPE_UADDR, .uaddr = (u64)(uintptr_t)iov->iov_base, .size = iov->iov_len, }; + if (!desc->uaddr) + desc->type = IO_REGBUF_TYPE_EMPTY; } int __io_account_mem(struct user_struct *user, unsigned long nr_pages) @@ -323,6 +321,8 @@ static int __io_sqe_files_update(struct io_ring_ctx *ctx, return -ENXIO; if (up->offset + nr_args > ctx->file_table.data.nr) return -EINVAL; + if (up->flags) + return -EINVAL; for (done = 0; done < nr_args; done++) { u64 tag = 0; @@ -382,9 +382,8 @@ static int __io_sqe_buffers_update(struct io_ring_ctx *ctx, struct io_uring_rsrc_update2 *up, unsigned int nr_args) { + bool extended = up->flags & IORING_RSRC_UPDATE_EXTENDED; u64 __user *tags = u64_to_user_ptr(up->tags); - struct iovec fast_iov, *iov; - struct iovec __user *uvec; u64 user_data = up->data; __u32 done; int i, err; @@ -393,29 +392,49 @@ static int __io_sqe_buffers_update(struct io_ring_ctx *ctx, return -ENXIO; if (up->offset + nr_args > ctx->buf_table.nr) return -EINVAL; + if (up->flags & ~IORING_RSRC_UPDATE_EXTENDED) + return -EINVAL; for (done = 0; done < nr_args; done++) { struct io_uring_regbuf_desc desc; struct io_rsrc_node *node; u64 tag = 0; - uvec = u64_to_user_ptr(user_data); - iov = iovec_from_user(uvec, 1, 1, &fast_iov, io_is_compat(ctx)); - if (IS_ERR(iov)) { - err = PTR_ERR(iov); - break; - } if (tags && copy_from_user(&tag, &tags[done], sizeof(tag))) { err = -EFAULT; break; } - io_iov_to_regbuf_desc(iov, &desc); + if (extended) { + if (copy_from_user(&desc, u64_to_user_ptr(user_data), + sizeof(desc))) { + err = -EFAULT; + break; + } + user_data += sizeof(desc); + } else { + struct iovec __user *uvec = u64_to_user_ptr(user_data); + struct iovec fast_iov, *iov; + + if (io_is_compat(ctx)) + user_data += sizeof(struct compat_iovec); + else + user_data += sizeof(struct iovec); + + iov = iovec_from_user(uvec, 1, 1, &fast_iov, io_is_compat(ctx)); + if (IS_ERR(iov)) { + err = PTR_ERR(iov); + break; + } + io_iov_to_regbuf_desc(iov, &desc); + } + node = io_sqe_buffer_register(ctx, &desc); if (IS_ERR(node)) { err = PTR_ERR(node); break; } + if (tag) { if (!node) { err = -EINVAL; @@ -426,10 +445,6 @@ static int __io_sqe_buffers_update(struct io_ring_ctx *ctx, i = array_index_nospec(up->offset + done, ctx->buf_table.nr); io_reset_rsrc_node(ctx, &ctx->buf_table, i); ctx->buf_table.nodes[i] = node; - if (io_is_compat(ctx)) - user_data += sizeof(struct compat_iovec); - else - user_data += sizeof(struct iovec); } return done ? done : err; } @@ -464,7 +479,7 @@ int io_register_files_update(struct io_ring_ctx *ctx, void __user *arg, memset(&up, 0, sizeof(up)); if (copy_from_user(&up, arg, sizeof(struct io_uring_rsrc_update))) return -EFAULT; - if (up.resv || up.resv2) + if (up.resv2) return -EINVAL; return __io_register_rsrc_update(ctx, IORING_RSRC_FILE, &up, nr_args); } @@ -478,7 +493,7 @@ int io_register_rsrc_update(struct io_ring_ctx *ctx, void __user *arg, return -EINVAL; if (copy_from_user(&up, arg, sizeof(up))) return -EFAULT; - if (!up.nr || up.resv || up.resv2) + if (!up.nr || up.resv2) return -EINVAL; return __io_register_rsrc_update(ctx, type, &up, up.nr); } @@ -578,12 +593,9 @@ int io_files_update(struct io_kiocb *req, unsigned int issue_flags) struct io_uring_rsrc_update2 up2; int ret; + memset(&up2, 0, sizeof(up2)); up2.offset = up->offset; up2.data = up->arg; - up2.nr = 0; - up2.tags = 0; - up2.resv = 0; - up2.resv2 = 0; if (up->offset == IORING_FILE_INDEX_ALLOC) { ret = io_files_update_with_index_alloc(req, issue_flags); @@ -882,8 +894,13 @@ static struct io_rsrc_node *io_sqe_buffer_register(struct io_ring_ctx *ctx, struct io_imu_folio_data data; bool coalesced = false; - if (!uaddr) { - if (size) + if (desc->type >= __IO_REGBUF_TYPE_MAX) + return ERR_PTR(-EINVAL); + if (!mem_is_zero(&desc->__resv, sizeof(desc->__resv)) || desc->flags) + return ERR_PTR(-EINVAL); + + if (desc->type == IO_REGBUF_TYPE_EMPTY) { + if (uaddr || size) return ERR_PTR(-EFAULT); /* remove the buffer without installing a new one */ return NULL; -- 2.54.0