From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f53.google.com (mail-wm1-f53.google.com [209.85.128.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3B84F3431E3 for ; Wed, 7 Oct 2026 01:43:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.53 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791337394; cv=none; b=BSpLKoMiLfLD/JGmutg0hc5ksmX0u3AZkkOZbSsM+yYz1BBN5f6+uoIAym32CZmdwAlJGqyzeN+YVvCbJHH1eQg/umR0wl0y1x4+O6rLMLzDBq1nRt6Ien/vd04IYIPr/kSahxCcdKkLzSepExXdsI8y+BOwjZBrfT14MEsorl4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791337394; c=relaxed/simple; bh=74XT/YcpL6u3pCQJh393Sbt1vl2bLCbf4FmkhI3qmoI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=tITCZlqV2ZUT6iH7Qx81Ij587LqoTUHsAJiCNgl72VtVwpicWJJiHfGYI3WjbDBSKUID5Rk1vrKEhdUl5L1SjhP8en60q9QewmdYG2jgomiIbF0xnkhLMifH8TXYJNRZwHWywcsnSLRnjf7YsuZ+WFremv4D5TxqQqdFfdqt4wY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=W96iVFeW; arc=none smtp.client-ip=209.85.128.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="W96iVFeW" Received: by mail-wm1-f53.google.com with SMTP id 5b1f17b1804b1-4a16bc2278aso9551185e9.1 for ; Tue, 06 Oct 2026 18:43:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1791337390; x=1791942190; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=A+VjBFsuZYeJUSxMs311afilNm9e5ulL53j6D1D3BK4=; b=W96iVFeWOzmqA6XJi5NljeTN4azUolJRUgjDDT3st0QZod2unj0NodvUtlhWuL0FL+ /uylAVmnbshvk5Zs+p9mtt9OK+yKy6VTNxfBsf1dGAjoapr5pBoSz24oZEzHm/Ob6Pac Vn4msF78A3EPoAn5cY5lUodWNGFfPuvf5jMcLXktg//6nVHmHq8iUaX29hqw/fYHjO0P /9BdCO07ySsIq6xpd43QRBAQpnkDtzd1e9nPMcd7R+F48F+1CXaJt2rVQFaomJqqyst3 9gDsd7WTBvt+vUSb2EN5g49uRAeb58VJp2G9LE3a+dGdFgYcymcFwGfgCXn69F9N2SlO eJsg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791337390; x=1791942190; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=A+VjBFsuZYeJUSxMs311afilNm9e5ulL53j6D1D3BK4=; b=cSSe9VG03DcWRYn0o93JpunKU7pYNs875YQpRyNVAeEz4zqFS6hyARo+Z1ajLACKH0 /lvrDkbZD7830nPJj0ynn1hvK8vdY7XvwQoLRyl2fApxO+UUPArbUZtwLoqCmlNZ0joM 6S3j4nu53uecVg/ut4docpB62mBA0JCwx/dT5PekwuevQhzMLWh30rGoXNXW35Dn06v9 ULFFrXID+DVYrBOTFfR8IzOOy1iaBQQzsVDItfM3Zf41URp8AgR1QndsV5kQ8rBWAfOt 9kNmbKR5bRkCiXqR5da84QOWkqDxcZfLvtc3KJwCkxMPK6EUvcJcgkL35NXW2j94cwIt S+3w== X-Forwarded-Encrypted: i=1; AKwUvBxAo41/r6PuK/VqvL6yu3ejzdUqVyod3NXa/vrRqBWRQXgZotG0KJp0bXFtrQtuhtIkj0Qh0LDjWSkv/qE=@vger.kernel.org X-Gm-Message-State: AFuF++k486NtVWl7xrCsH9sFIKi+swTMxYS3WfFqV2sK6mdNY6lCp8n+ Rh1qfKNp1+TjjBocHoALjeEdey2bZ8ZvOPdM61v96Zel/dVq09NtyT/e X-Gm-Gg: AYBFou2RAVZ7NIutEvZtt6x0a30YhA1TB9RmI+iXWvDZN1yXAd3uVZlQq5uh3ZDcWou FK9GZ3sToC3bM6d/zJlVK/nfbFzHla6undYlw15JdFyINviNVjszqekZOxRJaa/+9aW/+LZRcwJ 13wH5ISPUsV9MYV2VR4iL9PmHuEDW1+n6mp8HCIUwUG5Ynlqb5cmH/kR3VxE9p5QYMI4ltuxOgv mIuTiuWUBNHI/5HGSvj1NT1S2M+KQAXn2XKGpx+vUWoD6t1xq1hKPUfjsl6wf1bvSWvZL5VIpms UUB+Mv0hJ8Ad42dx53CEHabjAGpxDlDdXYgqChfrMVAdwNO07qoDBAkeqvP/+vXuhz9EmiubDSB F3hyKCffXFrAEimFslRw70DCC60CrqGkjT30ZoJmJGCb5BRWaFvn6njc5YrWtuOHB+BD3V3oVfv vlJ1NfBnNitB1cgNYMg+WN2rLX0i+4Kr6Ysc73lNLVE7M7TrWc121QNO7Uh1iIHFPKPKl937DOW O/yNlzBQL+MkbA12GJrSKu7YsR9EvtVbJ+lsTex6lxQNBNNF/y0a+9zIyKcUOk7nBR0i+qzkyTh D4p1d5Ce1yw= X-Received: by 2002:a05:600c:a00e:b0:4a0:1a18:b742 with SMTP id 5b1f17b1804b1-4a1800d3de4mr8466785e9.2.1791337390371; Tue, 06 Oct 2026 18:43:10 -0700 (PDT) Received: from 127.mynet ([2a01:4b00:bd21:4f00:7cc6:d3ca:494:116c]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4a17f653d7csm26423895e9.14.2026.10.06.18.43.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 06 Oct 2026 18:43:09 -0700 (PDT) From: Pavel Begunkov To: linux-block@vger.kernel.org Cc: asml.silence@gmail.com, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, linux-nvme@lists.infradead.org, linux-fsdevel@vger.kernel.org, io-uring@vger.kernel.org, Christoph Hellwig , Sumit Semwal , =?UTF-8?q?Christian=20K=C3=B6nig?= , Keith Busch , Sagi Grimberg , Alexander Viro , Christian Brauner , Jan Kara , Andrew Morton , Jens Axboe , Nitesh Shetty , Kanchan Joshi , Anuj Gupta , Tushar Gohad , William Power , Matthew Brost , Alasdair Kergon , Mike Snitzer , Mikulas Patocka , Benjamin Marzinski , dm-devel@lists.linux.dev Subject: [PATCH v9 01/13] dma-buf: introduce initial file I/O infrastructure Date: Wed, 7 Oct 2026 02:42:44 +0100 Message-ID: X-Mailer: git-send-email 2.54.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit The goal is to be able to natively use dma-buf in the read-write / IO path. This patch adds basic building blocks serving as a glue and API between drivers and upper layer subsystems providing the uAPI. Later patches implement it for NVMe raw block devices and expose it to the user space via io_uring. There are two main objects. struct dma_buf_io_ctx and struct dma_buf_io_map. The ctx is used during initial registration and serves as an interface between the upper layer user like io_uring and to the importer subsystem / driver. The map represents the actual dma map established for the target device[s] with dma_buf_map_attachment() and stored in a device specific format. The context is created via a new file operation ->init_dma_buf_io_ctx. The ctx-map separation exists to facilitate map invalidation (see dma_buf_io_invalidate_mappings()). A ctx can create multiple maps during its lifetime, but there can only be no more than one (active) map attached to it. Invalidation drops the active map if present, and any new I/O request will try to create a new one. The primary task of dma_buf_io_map is to count requests using it and wait for their completion when we want to destroy the DMA map. [un]mapping and any work with dma addresses is delegated to the importer driver via an ops table stored in the ctx, see struct dma_buf_io_ops. Only the target driver / subsystem knows about devices it wants to use the dma-buf with, especially in case of multi-device filesystems or stacking in the future. Signed-off-by: Pavel Begunkov --- drivers/dma-buf/Makefile | 2 +- drivers/dma-buf/dma-buf-io.c | 183 +++++++++++++++++++++++++++++++++++ include/linux/dma-buf-io.h | 105 ++++++++++++++++++++ include/linux/fs.h | 2 + 4 files changed, 291 insertions(+), 1 deletion(-) create mode 100644 drivers/dma-buf/dma-buf-io.c create mode 100644 include/linux/dma-buf-io.h diff --git a/drivers/dma-buf/Makefile b/drivers/dma-buf/Makefile index b25d7550bacf..523731b0f83e 100644 --- a/drivers/dma-buf/Makefile +++ b/drivers/dma-buf/Makefile @@ -1,6 +1,6 @@ # SPDX-License-Identifier: GPL-2.0-only obj-y := dma-buf.o dma-fence.o dma-fence-array.o dma-fence-chain.o \ - dma-fence-unwrap.o dma-resv.o dma-buf-mapping.o + dma-fence-unwrap.o dma-resv.o dma-buf-mapping.o dma-buf-io.o obj-$(CONFIG_DMABUF_HEAPS) += dma-heap.o obj-$(CONFIG_DMABUF_HEAPS) += heaps/ obj-$(CONFIG_SYNC_FILE) += sync_file.o diff --git a/drivers/dma-buf/dma-buf-io.c b/drivers/dma-buf/dma-buf-io.c new file mode 100644 index 000000000000..9ba9c17f900a --- /dev/null +++ b/drivers/dma-buf/dma-buf-io.c @@ -0,0 +1,183 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Common infrastructure for supporing dma-buf in the I/O path. + * + * Copyright (C) 2026 Pavel Begunkov + */ +#include +#include + +static void dma_buf_io_map_release(struct dma_buf_io_map *map) +{ + struct dma_buf_io_ctx *ctx = map->ctx; + + dma_resv_assert_held(ctx->dmabuf->resv); + + percpu_ref_exit(&map->refs); + ctx->dev_ops->unmap(ctx, map); +} + +static void dma_buf_io_map_refs_cb(struct percpu_ref *ref) +{ + struct dma_buf_io_map *map = container_of(ref, struct dma_buf_io_map, refs); + + complete(&map->drained); +} + +static void dma_buf_io_kill_map(struct dma_buf_io_ctx *ctx) +{ + struct dma_buf_io_map *map; + + dma_resv_assert_held(ctx->dmabuf->resv); + + map = rcu_dereference_protected(ctx->map, + dma_resv_held(ctx->dmabuf->resv)); + if (!map) + return; + + rcu_assign_pointer(ctx->map, NULL); + percpu_ref_kill(&map->refs); + /* make sure the map is not visible via ctx->map */ + synchronize_rcu_expedited(); + wait_for_completion(&map->drained); + dma_buf_io_map_release(map); +} + +int dma_buf_io_init_map(struct dma_buf_io_ctx *ctx, struct dma_buf_io_map *map, + struct sg_table *sgt) +{ + unsigned seg_shift = ~0U; + struct scatterlist *sg; + unsigned long tmp; + int ret; + + for_each_sgtable_dma_sg(sgt, sg, tmp) + seg_shift = min(seg_shift, __ffs(sg_dma_len(sg))); + + ret = percpu_ref_init(&map->refs, dma_buf_io_map_refs_cb, 0, GFP_KERNEL); + if (ret) + return ret; + init_completion(&map->drained); + map->min_seg_shift = seg_shift; + map->ctx = ctx; + return 0; +} +EXPORT_SYMBOL_NS_GPL(dma_buf_io_init_map, "DMA_BUF"); + +static struct dma_buf_io_map *__dma_buf_io_create_map(struct dma_buf_io_ctx *ctx) +{ + struct dma_buf *dmabuf = ctx->dmabuf; + struct dma_buf_io_map *map; + long ret; + + dma_resv_assert_held(dmabuf->resv); + + if (ctx->is_dead) + return ERR_PTR(-ENOENT); + + map = __dma_buf_io_get_map(ctx); + if (map) + return map; + + map = ctx->dev_ops->map(ctx); + if (IS_ERR(map)) + return map; + + if (WARN_ON_ONCE(!map->min_seg_shift)) + return ERR_PTR(-EFAULT); + + ret = dma_resv_wait_timeout(dmabuf->resv, DMA_RESV_USAGE_KERNEL, + true, MAX_SCHEDULE_TIMEOUT); + if (ret <= 0) { + if (!ret) + ret = -EAGAIN; + dma_buf_io_map_release(map); + return ERR_PTR(ret); + } + + /* get a reference for the caller */ + percpu_ref_get(&map->refs); + rcu_assign_pointer(ctx->map, map); + return map; +} + +struct dma_buf_io_map *dma_buf_io_create_map(struct dma_buf_io_ctx *ctx) +{ + struct dma_buf_io_map *map; + int ret; + + ret = dma_resv_lock_interruptible(ctx->dmabuf->resv, NULL); + if (ret) + return ERR_PTR(ret); + + map = __dma_buf_io_create_map(ctx); + dma_resv_unlock(ctx->dmabuf->resv); + return map; +} + +static void dma_buf_io_ctx_shutdown(struct dma_buf_io_ctx *ctx) +{ + dma_resv_lock(ctx->dmabuf->resv, NULL); + ctx->is_dead = true; + dma_buf_io_kill_map(ctx); + dma_resv_unlock(ctx->dmabuf->resv); +} + +int dma_buf_io_ctx_create(struct file *file, + struct dma_buf *dmabuf, + enum dma_data_direction dir, + struct dma_buf_io_ctx **out_ctx) +{ + struct dma_buf_io_ctx *ctx; + int ret; + + if (!file->f_op->init_dma_buf_io_ctx) + return -EOPNOTSUPP; + + ctx = kzalloc_obj(*ctx); + if (!ctx) + return -ENOMEM; + ctx->dir = dir; + ctx->dmabuf = dmabuf; + get_dma_buf(dmabuf); + + ret = file->f_op->init_dma_buf_io_ctx(file, ctx); + if (ret) { + kfree(ctx); + dma_buf_put(dmabuf); + return ret; + } + + if (WARN_ON_ONCE(!ctx->dev_ops || + !ctx->dev_ops->map || + !ctx->dev_ops->unmap || + !ctx->dev_ops->release)) + return -EINVAL; + + *out_ctx = ctx; + return 0; +} + +void dma_buf_io_ctx_release(struct dma_buf_io_ctx *ctx) +{ + dma_buf_io_ctx_shutdown(ctx); + + if (WARN_ON_ONCE(rcu_dereference_protected(ctx->map, true))) + return; + + ctx->dev_ops->release(ctx); + dma_buf_put(ctx->dmabuf); + kfree(ctx); +} + +void dma_buf_io_invalidate_mappings(struct dma_buf_io_ctx *ctx) +{ + dma_buf_io_kill_map(ctx); +} +EXPORT_SYMBOL_NS_GPL(dma_buf_io_invalidate_mappings, "DMA_BUF"); + +void dma_buf_io_detach(struct dma_buf_io_ctx *ctx) +{ + dma_buf_io_ctx_shutdown(ctx); +} +EXPORT_SYMBOL_NS_GPL(dma_buf_io_detach, "DMA_BUF"); diff --git a/include/linux/dma-buf-io.h b/include/linux/dma-buf-io.h new file mode 100644 index 000000000000..8ac338df51ff --- /dev/null +++ b/include/linux/dma-buf-io.h @@ -0,0 +1,105 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef __DMA_BUF_IO_H__ +#define __DMA_BUF_IO_H__ + +#include + +struct dma_buf_io_ctx; +struct dma_buf_io_map; + +struct dma_buf_io_ops { + /* + * Create a new map for the given ctx. Called with the reservation + * lock held. + */ + struct dma_buf_io_map *(*map)(struct dma_buf_io_ctx *ctx); + + /* + * Clean up device specific parts of the @map. Called with the + * reservation lock held. + */ + void (*unmap)(struct dma_buf_io_ctx *ctx, struct dma_buf_io_map *map); + + /* + * The user tries to destroy the ctx. Release all device specific + * parts of the token. + */ + void (*release)(struct dma_buf_io_ctx *); +}; + +struct dma_buf_io_map { + /* + * Counts attached requests and other users. Device specific unmapping + * is deferred until all refs are dropped. + */ + struct percpu_ref refs; + /* + * Shift for the minimum segment size of the mapping. + */ + unsigned min_seg_shift; + + struct dma_buf_io_ctx *ctx; + struct completion drained; +}; + +struct dma_buf_io_ctx { + struct dma_buf_io_map __rcu *map; + struct dma_buf *dmabuf; + enum dma_data_direction dir; + bool is_dead; + + void *dev_priv; + const struct dma_buf_io_ops *dev_ops; +}; + +int dma_buf_io_ctx_create(struct file *file, + struct dma_buf *dmabuf, + enum dma_data_direction dir, + struct dma_buf_io_ctx **ctx); +void dma_buf_io_ctx_release(struct dma_buf_io_ctx *ctx); + +struct dma_buf_io_map *dma_buf_io_create_map(struct dma_buf_io_ctx *ctx); + +static inline struct dma_buf_io_map * +__dma_buf_io_get_map(struct dma_buf_io_ctx *ctx) +{ + struct dma_buf_io_map *map; + + guard(rcu)(); + + map = rcu_dereference(ctx->map); + if (unlikely(!map || !percpu_ref_tryget_live_rcu(&map->refs))) + return NULL; + + return map; +} + +static inline struct dma_buf_io_map * +dma_buf_io_get_map(struct dma_buf_io_ctx *ctx, bool nowait) +{ + struct dma_buf_io_map *map; + + map = __dma_buf_io_get_map(ctx); + if (likely(map)) + return map; + + if (nowait) + return ERR_PTR(-EAGAIN); + return dma_buf_io_create_map(ctx); +} + +static inline void dma_buf_io_map_drop(struct dma_buf_io_map *map) +{ + percpu_ref_put(&map->refs); +} + +/* + * Device API + */ + +void dma_buf_io_invalidate_mappings(struct dma_buf_io_ctx *ctx); +int dma_buf_io_init_map(struct dma_buf_io_ctx *ctx, struct dma_buf_io_map *map, + struct sg_table *sgt); +void dma_buf_io_detach(struct dma_buf_io_ctx *ctx); + +#endif diff --git a/include/linux/fs.h b/include/linux/fs.h index f9d1e05e8ae6..05c1ff495732 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -1914,6 +1914,7 @@ struct dir_context { #define COPY_FILE_SPLICE (1 << 0) struct io_uring_cmd; +struct dma_buf_io_ctx; struct offset_ctx; struct file_operations { @@ -1960,6 +1961,7 @@ struct file_operations { int (*uring_cmd_iopoll)(struct io_uring_cmd *, struct io_comp_batch *, unsigned int poll_flags); int (*mmap_prepare)(struct vm_area_desc *); + int (*init_dma_buf_io_ctx)(struct file *, struct dma_buf_io_ctx *); } __randomize_layout; /* Supports async buffered reads */ -- 2.54.0