From: Pavel Begunkov <asml.silence@gmail.com>
To: linux-block@vger.kernel.org
Cc: asml.silence@gmail.com, linux-kernel@vger.kernel.org,
linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org,
linaro-mm-sig@lists.linaro.org, linux-nvme@lists.infradead.org,
linux-fsdevel@vger.kernel.org, io-uring@vger.kernel.org,
"Christoph Hellwig" <hch@lst.de>,
"Sumit Semwal" <sumit.semwal@linaro.org>,
"Christian König" <christian.koenig@amd.com>,
"Keith Busch" <kbusch@kernel.org>,
"Sagi Grimberg" <sagi@grimberg.me>,
"Alexander Viro" <viro@zeniv.linux.org.uk>,
"Christian Brauner" <brauner@kernel.org>,
"Jan Kara" <jack@suse.cz>,
"Andrew Morton" <akpm@linux-foundation.org>,
"Jens Axboe" <axboe@kernel.dk>,
"Nitesh Shetty" <nj.shetty@samsung.com>,
"Kanchan Joshi" <joshi.k@samsung.com>,
"Anuj Gupta" <anuj20.g@samsung.com>,
"Tushar Gohad" <tushar.gohad@intel.com>,
"William Power" <william.power@intel.com>,
"Phil Cayton" <phil.cayton@intel.com>,
"Matthew Brost" <matthew.brost@intel.com>,
"Alasdair Kergon" <agk@redhat.com>,
"Mike Snitzer" <snitzer@kernel.org>,
"Mikulas Patocka" <mpatocka@redhat.com>,
"Benjamin Marzinski" <bmarzins@redhat.com>,
dm-devel@lists.linux.dev
Subject: [PATCH v6 00/13] Add dmabuf read/write via io_uring
Date: Mon, 21 Sep 2026 14:38:44 +0100 [thread overview]
Message-ID: <cover.1789997898.git.asml.silence@gmail.com> (raw)
The patch set allows to register a dmabuf to an io_uring instance for
a specified file and use it with io_uring read / write requests. The
infrastructure is not tied to io_uring and there could be more users
in the future. A similar idea was attempted some years ago by Keith [1],
from where I borrowed a number of changes. Later it was brough back up
to life by Tushar and Vishal.
It's an opt-in feature for files, and they need to implement a new
file operation to use it. Only NVMe block devices are supported in this
series. The user API is built on top of io_uring's "registered buffers",
where a dmabuf is registered in a special way, but after it can be used
as any other "registered buffer" with IORING_OP_{READ,WRITE}_FIXED
requests. It's created via a new file operation and the resulted map is
then passed through the I/O stack in a new iterator type. There is some
additional infrastructure to glue it together, count requests, manage
lifetime and implement invalidation.
Tushar, William, Phil did a lot of testing and experimentation on various
devices with previous versions of the patch set, and I've received lots
of help from Anuj, Kanchan and Nitesh with investigations, testing, and
patching. Earlier benchmarks by Anuj for IOMMU optimisations with udmabuf
showed:
STRICT: before = 570 KIOPS, after = 5.01 MIOPS
LAZY: before = 1.93 MIOPS, after = 5.01 MIOPS
PASSTHROUGH: before = 5.01 MIOPS, after = 5.01 MIOPS
# Patch set structure:
- Patches 1-2 introduce internal API and infrastructure mediating
io_uring and target subsystem / devices
- Patches 3-5: block layer support + prep patches
- Patches 6-8 implement NVMe support
- Patches 9-13 add io_uring support and uapi.
There are some liburing tests that can serve as an example:
git: https://github.com/isilence/liburing.git rw-dmabuf-tests-v5
url: https://github.com/isilence/liburing/tree/rw-dmabuf-tests-v5
The patches are based on Jens' for-next. Also available as a branch:
git: https://github.com/isilence/linux.git rw-dmabuf-v6
url: https://github.com/isilence/linux/tree/rw-dmabuf-v6
[1] https://lore.kernel.org/io-uring/20220805162444.3985535-1-kbusch@fb.com/
v6: - remove fences and wait on invalidation synchronously
- report all map requests are detached earlier, outside of wq
- handle nvme removal
- gate nvme_free_descriptors() on iod->nr_descriptors
- fix error handling in nvme_rq_setup_dmabuf_map()
- fix iov_iter_alignment()
- relax the 1G io_uring registered dma-buf limit
- fix type overflow issue in bio split
- allocate struct dma_buf_io_ctx inside the dma-buf-io code
- harden block checks against dma-buf + buffered io
v5: - Reject dma-buf with buffered IO for raw bdev
- Add lim->max_segments bio splitting
- Add bio_iov_iter_set() helper
- Fix io_uring uapi validation
- Other minor changes
NVMe:
- Rename nvme_pci_sgl_set_data() to nvme_pci_dma_iter_set_sgl(), split
into its own prep patch.
- Convert segment walk loops to do-while.
- Drop adjacent segment coalescing logic.
- Drop first_dma/first_len from nvme_pci_dmabuf_sgl_nents()
- Remove entries > NVME_MAX_SEGS bailout; moving to the block layer.
- Factor SGL vs PRP decision making into a helper.
v4: - https://lore.kernel.org/all/cover.1785274111.git.asml.silence@gmail.com/
- Add sgl support from Anuj
- Move it under drivers/dma-buf/ and rename
- Fix mis-sized allocations
- Fix io_uring re-import mishandling
- Drop map before io_uring "task work"
- Move blk-mq callback to block_device_operations
- Convert bio flag to REQ_OP*
- Other small changes
v3: https://lore.kernel.org/io-uring/cover.1777475843.git.asml.silence@gmail.com/
- Rework io_uring registration
- Move token/map infrastructure code out of blk-mq
- Simplify callbacks: remove a separate blk-mq table, which was
mostly just forwarding calls (to nvme).
- Don't skip dma sync depending on request direction
- Fix a couple of hangs
- Rename s/dma/dmabuf/
- Other small changes
v2: - Don't pass raw dma addresses, wrap it into a driver specific object
- Split into two objects: token and map
- Implement move_notify
Anuj Gupta (2):
nvme-pci: rename nvme_pci_sgl_set_data to nvme_pci_dma_iter_set_sgl
nvme-pci: add SGL support for the dmabuf path
Pavel Begunkov (11):
dma-buf: introduce initial file I/O infrastructure
iov_iter: add iterator type for dmabuf maps
block: always adjust bi_offset on bio_advance_iter
block: introduce dma map backed bio type
block: add dma-buf support for raw bdev
nvme-pci: implement dma-buf backed requests
io_uring/rsrc: introduce buf registration structure
io_uring/rsrc: extend buffer update
io_uring/rsrc: add uncloneable regbuf flag
io_uring/rsrc: add regbuf import flags
io_uring/rsrc: add dmabuf backed registered buffers
block/bio.c | 15 +-
block/blk-merge.c | 45 +++
block/fops.c | 24 +-
drivers/dma-buf/Makefile | 2 +-
drivers/dma-buf/dma-buf-io.c | 219 +++++++++++++++
drivers/md/dm-io-rewind.c | 6 +-
drivers/nvme/host/core.c | 12 +
drivers/nvme/host/nvme.h | 2 +
drivers/nvme/host/pci.c | 484 ++++++++++++++++++++++++++++++++-
include/linux/bio.h | 21 +-
include/linux/blk-mq.h | 7 +
include/linux/blk_types.h | 14 +-
include/linux/blkdev.h | 2 +
include/linux/bvec.h | 3 +-
include/linux/dma-buf-io.h | 113 ++++++++
include/linux/fs.h | 2 +
include/linux/io_uring_types.h | 5 +
include/linux/uio.h | 11 +
include/uapi/linux/io_uring.h | 31 ++-
io_uring/io_uring.c | 3 +-
io_uring/rsrc.c | 274 ++++++++++++++++---
io_uring/rsrc.h | 45 ++-
io_uring/rw.c | 6 +-
lib/iov_iter.c | 29 +-
24 files changed, 1295 insertions(+), 80 deletions(-)
create mode 100644 drivers/dma-buf/dma-buf-io.c
create mode 100644 include/linux/dma-buf-io.h
--
2.54.0
next reply other threads:[~2026-09-21 13:39 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 13:38 Pavel Begunkov [this message]
2026-09-21 13:38 ` [PATCH v6 01/13] dma-buf: introduce initial file I/O infrastructure Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 02/13] iov_iter: add iterator type for dmabuf maps Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 03/13] block: always adjust bi_offset on bio_advance_iter Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 04/13] block: introduce dma map backed bio type Pavel Begunkov
2026-09-22 13:16 ` Christoph Hellwig
2026-09-22 13:54 ` Pavel Begunkov
2026-09-23 4:51 ` Christoph Hellwig
2026-09-21 13:38 ` [PATCH v6 05/13] block: add dma-buf support for raw bdev Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 06/13] nvme-pci: implement dma-buf backed requests Pavel Begunkov
2026-09-22 13:20 ` Christoph Hellwig
2026-09-22 13:37 ` Pavel Begunkov
2026-09-23 4:55 ` Christoph Hellwig
2026-09-21 13:38 ` [PATCH v6 07/13] nvme-pci: rename nvme_pci_sgl_set_data to nvme_pci_dma_iter_set_sgl Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 08/13] nvme-pci: add SGL support for the dmabuf path Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 09/13] io_uring/rsrc: introduce buf registration structure Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 10/13] io_uring/rsrc: extend buffer update Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 11/13] io_uring/rsrc: add uncloneable regbuf flag Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 12/13] io_uring/rsrc: add regbuf import flags Pavel Begunkov
2026-09-21 13:38 ` [PATCH v6 13/13] io_uring/rsrc: add dmabuf backed registered buffers Pavel Begunkov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cover.1789997898.git.asml.silence@gmail.com \
--to=asml.silence@gmail.com \
--cc=agk@redhat.com \
--cc=akpm@linux-foundation.org \
--cc=anuj20.g@samsung.com \
--cc=axboe@kernel.dk \
--cc=bmarzins@redhat.com \
--cc=brauner@kernel.org \
--cc=christian.koenig@amd.com \
--cc=dm-devel@lists.linux.dev \
--cc=dri-devel@lists.freedesktop.org \
--cc=hch@lst.de \
--cc=io-uring@vger.kernel.org \
--cc=jack@suse.cz \
--cc=joshi.k@samsung.com \
--cc=kbusch@kernel.org \
--cc=linaro-mm-sig@lists.linaro.org \
--cc=linux-block@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-media@vger.kernel.org \
--cc=linux-nvme@lists.infradead.org \
--cc=matthew.brost@intel.com \
--cc=mpatocka@redhat.com \
--cc=nj.shetty@samsung.com \
--cc=phil.cayton@intel.com \
--cc=sagi@grimberg.me \
--cc=snitzer@kernel.org \
--cc=sumit.semwal@linaro.org \
--cc=tushar.gohad@intel.com \
--cc=viro@zeniv.linux.org.uk \
--cc=william.power@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®