From: "Christian König" <christian.koenig@amd.com>
To: Fred Griffoul <griffoul@gmail.com>,
Paolo Bonzini <pbonzini@redhat.com>,
Sean Christopherson <seanjc@google.com>,
Marc Zyngier <maz@kernel.org>, Oliver Upton <oupton@kernel.org>,
Sumit Semwal <sumit.semwal@linaro.org>,
Jason Gunthorpe <jgg@ziepe.ca>, Kevin Tian <kevin.tian@intel.com>
Cc: David Woodhouse <dwmw2@infradead.org>,
Ackerley Tng <ackerleytng@google.com>,
Joey Gouly <joey.gouly@arm.com>,
Suzuki K Poulose <suzuki.poulose@arm.com>,
Zenghui Yu <yuzenghui@huawei.com>,
Steffen Eiden <seiden@linux.ibm.com>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>, Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
"H . Peter Anvin" <hpa@zytor.com>, Joerg Roedel <joro@8bytes.org>,
Robin Murphy <robin.murphy@arm.com>,
Alex Williamson <alex@shazbot.org>, Shuah Khan <shuah@kernel.org>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
kvmarm@lists.linux.dev, linux-arm-kernel@lists.infradead.org,
iommu@lists.linux.dev, linux-media@vger.kernel.org,
dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org,
linux-kselftest@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, x86@kernel.org
Subject: Re: [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run
Date: Mon, 5 Oct 2026 12:07:00 +0200 [thread overview]
Message-ID: <a04f17cc-f59a-4240-81c8-29776eee3a2e@amd.com> (raw)
In-Reply-To: <20261005095552.52748-3-griffoul@gmail.com>
On 10/5/26 11:55, Fred Griffoul wrote:
> From: Fred Griffoul <fgriffo@amazon.co.uk>
>
> iommufd and KVM write physical addresses into their own page tables.
> To do so, they must ask the exporter which frames back an offset of a
> dma-buf, whether that memory is RAM or MMIO, and whether it may be
> written.
Well filling page tables by the importer is an absolutely clear NO-GO for the DMA-buf design, we have gotten down that path already and it took us years to remove this functionality again.
The problem you are facing here is that dma_buf_mmap() doesn't work because you don't have a VMA.
I can understand the reasoning that you don't want to have a VMA, but I don't think that this is a valid justification to add complexity to DMA-buf and especially bring an approach back which we have already deprecated.
A possible solution might be Jasons patch set to directly negotiate exposing PCI BAR regions as DMA-buf, but that certainly needs more discussion.
But the approach outlined here is an absolutely clear NAK from my side.
Regards,
Christian.
>
> Add a get_phys() operation. The importer passes an offset and a maximum
> length. The exporter reports one run: the frames that start at the
> offset, are backed and physically contiguous, and share one attribute
> word. The run never exceeds the length. The exporter may end it early,
> so importers must not assume that it is the longest possible run.
>
> get_phys() returns -ENOENT when the byte at the offset is not backed,
> and -ENODEV when the buffer is revoked. The caller holds the
> reservation, and either pins the attachment or handles revocation. A
> reported frame stays valid until an invalidation that covers it
> returns.
>
> The attribute word holds the memory type and a READONLY flag. Zero
> means writable RAM. Importers refuse unknown types, reserved bits and
> unknown flags, so attributes added later fail safely. Two flag bits are
> reserved: one for holes and one for confidential memory.
>
> Convert vfio-pci, the iommufd selftest exporter and the KVM sample.
> iommufd behaves as before: it maps a buffer only when one writable run
> covers all of it.
>
> Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
> ---
> drivers/dma-buf/dma-buf.c | 47 +++++++++++
> drivers/iommu/iommufd/iommufd_private.h | 8 --
> drivers/iommu/iommufd/iommufd_test.h | 16 ++++
> drivers/iommu/iommufd/pages.c | 74 ++++-------------
> drivers/iommu/iommufd/selftest.c | 106 ++++++++++++++++++++----
> drivers/vfio/pci/vfio_pci_dmabuf.c | 58 ++++++-------
> include/linux/dma-buf.h | 54 ++++++++++++
> include/linux/vfio_pci_core.h | 3 -
> samples/kvm/gmem_provider.c | 39 ++++-----
> 9 files changed, 267 insertions(+), 138 deletions(-)
>
> diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c
> index d504c636dc29..66b85d53ed22 100644
> --- a/drivers/dma-buf/dma-buf.c
> +++ b/drivers/dma-buf/dma-buf.c
> @@ -1389,6 +1389,53 @@ void dma_buf_invalidate_mappings(struct dma_buf *dmabuf)
> }
> EXPORT_SYMBOL_NS_GPL(dma_buf_invalidate_mappings, "DMA_BUF");
>
> +/**
> + * dma_buf_get_phys - describe the run that starts at an offset
> + * @attach: attachment to query
> + * @offset: first buffer byte to describe
> + * @len: maximum number of bytes to describe
> + * @phys: physical address and length of the run
> + * @attr: DMA_BUF_PHYS_ATTR_* word of the run
> + *
> + * A run is the longest stretch of backed bytes starting at @offset whose
> + * frames are physically contiguous and share one attribute word. On success
> + * *@phys starts at the byte at @offset and covers at most @len bytes; it may
> + * be shorter than the run.
> + *
> + * The dma-buf reservation must be held. The attachment must be pinned or have
> + * revocable importer operations. A frame remains valid until a covering
> + * invalidation callback returns; the exporter must invalidate a changed range
> + * before reusing its old frames.
> + *
> + * Returns:
> + *
> + * 0 on success, -ENOENT if the byte at @offset is not backed, -ENODEV if the
> + * buffer is revoked, -EOPNOTSUPP if the exporter cannot describe itself this
> + * way, or another negative error code.
> + */
> +int dma_buf_get_phys(struct dma_buf_attachment *attach, u64 offset, u64 len,
> + struct phys_vec *phys, u32 *attr)
> +{
> + u64 end;
> + int ret;
> +
> + if (WARN_ON_ONCE(!attach || !attach->dmabuf || !phys || !attr))
> + return -EINVAL;
> + if (!len || check_add_overflow(offset, len, &end) ||
> + end > attach->dmabuf->size)
> + return -EINVAL;
> +
> + dma_resv_assert_held(attach->dmabuf->resv);
> + if (!attach->dmabuf->ops->get_phys)
> + return -EOPNOTSUPP;
> +
> + ret = attach->dmabuf->ops->get_phys(attach, offset, len, phys, attr);
> + if (!ret && WARN_ON_ONCE(!phys->len || phys->len > len))
> + return -EIO;
> + return ret;
> +}
> +EXPORT_SYMBOL_NS_GPL(dma_buf_get_phys, "DMA_BUF");
> +
> /**
> * DOC: cpu access
> *
> diff --git a/drivers/iommu/iommufd/iommufd_private.h b/drivers/iommu/iommufd/iommufd_private.h
> index 43fbc5bed8de..5cded585c227 100644
> --- a/drivers/iommu/iommufd/iommufd_private.h
> +++ b/drivers/iommu/iommufd/iommufd_private.h
> @@ -716,8 +716,6 @@ bool iommufd_should_fail(void);
> int __init iommufd_test_init(void);
> void iommufd_test_exit(void);
> bool iommufd_selftest_is_mock_dev(struct device *dev);
> -int iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
> - struct phys_vec *phys);
> #else
> static inline void iommufd_test_syz_conv_iova_id(struct iommufd_ucmd *ucmd,
> unsigned int ioas_id,
> @@ -739,11 +737,5 @@ static inline bool iommufd_selftest_is_mock_dev(struct device *dev)
> {
> return false;
> }
> -static inline int
> -iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
> - struct phys_vec *phys)
> -{
> - return -EOPNOTSUPP;
> -}
> #endif
> #endif
> diff --git a/drivers/iommu/iommufd/iommufd_test.h b/drivers/iommu/iommufd/iommufd_test.h
> index 52b78cbcc920..28fd9c43edc4 100644
> --- a/drivers/iommu/iommufd/iommufd_test.h
> +++ b/drivers/iommu/iommufd/iommufd_test.h
> @@ -31,6 +31,8 @@ enum {
> IOMMU_TEST_OP_PASID_CHECK_HWPT,
> IOMMU_TEST_OP_DMABUF_GET,
> IOMMU_TEST_OP_DMABUF_REVOKE,
> + IOMMU_TEST_OP_MD_CHECK_MAPPED,
> + IOMMU_TEST_OP_MD_IOVA_TO_PHYS,
> };
>
> enum {
> @@ -193,6 +195,20 @@ struct iommu_test_cmd {
> __s32 dmabuf_fd;
> __u32 revoked;
> } dmabuf_revoke;
> + struct {
> + /*
> + * 1: every page in [iova, iova+length) must be mapped;
> + * 0: none of them may be. Mixed is an error.
> + */
> + __u32 mapped;
> + __u32 __reserved;
> + __aligned_u64 iova;
> + __aligned_u64 length;
> + } check_mapped;
> + struct {
> + __aligned_u64 iova;
> + __aligned_u64 out_phys; /* 0 if unmapped */
> + } iova_to_phys;
> };
> __u32 last;
> };
> diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c
> index f9b2ae6d7e96..196d1bb330c2 100644
> --- a/drivers/iommu/iommufd/pages.c
> +++ b/drivers/iommu/iommufd/pages.c
> @@ -1463,68 +1463,12 @@ static const struct dma_buf_attach_ops iopt_dmabuf_attach_revoke_ops = {
> .invalidate_mappings = iopt_revoke_notify,
> };
>
> -/*
> - * iommufd and vfio have a circular dependency. Future work for a phys
> - * based private interconnect will remove this.
> - */
> -/*
> - * Look up the exporter's phys accessor for iommufd's private-interconnect
> - * path. Also fills *is_cpu_ram: true if the exporter's memory is normal
> - * cache-coherent RAM (needs BATCH_CPU_MEMORY / IOMMU_CACHE), false for MMIO
> - * (needs BATCH_MMIO / IOMMU_MMIO). This will be replaced by a formal
> - * exporter op that returns phys + memory type together.
> - */
> -static int
> -sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
> - struct phys_vec *phys, bool *is_cpu_ram)
> -{
> - typeof(&vfio_pci_dma_buf_iommufd_map) fn;
> - int rc;
> -
> - rc = iommufd_test_dma_buf_iommufd_map(attachment, phys);
> - if (rc != -EOPNOTSUPP) {
> - *is_cpu_ram = false; /* test hook mimics VFIO MMIO */
> - return rc;
> - }
> -
> - /*
> - * Prototype: try the sample gmem provider's dma-buf exporter. This
> - * mirrors the vfio-pci private-interconnect hook, and (like it) is
> - * meant to be replaced by a formal negotiated exporter op returning
> - * phys + memory type. The provider serves RAM, so mark it CPU_RAM.
> - */
> - {
> - extern int gmem_provider_dma_buf_iommufd_map(
> - struct dma_buf_attachment *, struct phys_vec *);
> - typeof(&gmem_provider_dma_buf_iommufd_map) gfn;
> -
> - gfn = symbol_get(gmem_provider_dma_buf_iommufd_map);
> - if (gfn) {
> - rc = gfn(attachment, phys);
> - symbol_put(gmem_provider_dma_buf_iommufd_map);
> - if (rc != -EOPNOTSUPP) {
> - *is_cpu_ram = true;
> - return rc;
> - }
> - }
> - }
> -
> - if (!IS_ENABLED(CONFIG_VFIO_PCI_DMABUF))
> - return -EOPNOTSUPP;
> -
> - fn = symbol_get(vfio_pci_dma_buf_iommufd_map);
> - if (!fn)
> - return -EOPNOTSUPP;
> - rc = fn(attachment, phys);
> - symbol_put(vfio_pci_dma_buf_iommufd_map);
> - *is_cpu_ram = false; /* VFIO PCI dma-buf carries BAR (MMIO) memory */
> - return rc;
> -}
> -
> static int iopt_map_dmabuf(struct iommufd_ctx *ictx, struct iopt_pages *pages,
> struct dma_buf *dmabuf)
> {
> struct dma_buf_attachment *attach;
> + struct phys_vec pv;
> + u32 attr;
> int rc;
>
> attach = dma_buf_dynamic_attach(dmabuf, iommufd_global_device(),
> @@ -1546,10 +1490,20 @@ static int iopt_map_dmabuf(struct iommufd_ctx *ictx, struct iopt_pages *pages,
> if (rc)
> goto err_detach;
>
> - rc = sym_vfio_pci_dma_buf_iommufd_map(attach, &pages->dmabuf.phys,
> - &pages->dmabuf.is_cpu_ram);
> + /* One backed, writable run covering the buffer: refuse the rest. */
> + rc = dma_buf_get_phys(attach, 0, dmabuf->size, &pv, &attr);
> + if (rc == -ENOENT)
> + rc = -EOPNOTSUPP;
> if (rc)
> goto err_unpin;
> + if (pv.len != dmabuf->size || !dma_buf_phys_attr_known(attr) ||
> + (attr & DMA_BUF_PHYS_ATTR_FLAGS_MASK)) {
> + rc = -EOPNOTSUPP;
> + goto err_unpin;
> + }
> + pages->dmabuf.phys = pv;
> + pages->dmabuf.is_cpu_ram =
> + dma_buf_phys_attr_type(attr) == DMA_BUF_PHYS_ATTR_RAM;
>
> dma_resv_unlock(dmabuf->resv);
>
> diff --git a/drivers/iommu/iommufd/selftest.c b/drivers/iommu/iommufd/selftest.c
> index af07c642a526..0899272d1e66 100644
> --- a/drivers/iommu/iommufd/selftest.c
> +++ b/drivers/iommu/iommufd/selftest.c
> @@ -1962,32 +1962,31 @@ static void iommufd_test_dma_buf_release(struct dma_buf *dmabuf)
> kfree(priv);
> }
>
> -static const struct dma_buf_ops iommufd_test_dmabuf_ops = {
> - .attach = iommufd_test_dma_buf_attach,
> - .detach = iommufd_test_dma_buf_detach,
> - .map_dma_buf = iommufd_test_dma_buf_map,
> - .release = iommufd_test_dma_buf_release,
> - .unmap_dma_buf = iommufd_test_dma_buf_unmap,
> -};
> -
> -int iommufd_test_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
> - struct phys_vec *phys)
> +static int iommufd_test_dma_buf_get_phys(struct dma_buf_attachment *attachment,
> + u64 offset, u64 len,
> + struct phys_vec *phys, u32 *attr)
> {
> struct iommufd_test_dma_buf *priv = attachment->dmabuf->priv;
>
> dma_resv_assert_held(attachment->dmabuf->resv);
> -
> - if (attachment->dmabuf->ops != &iommufd_test_dmabuf_ops)
> - return -EOPNOTSUPP;
> -
> if (priv->revoked)
> return -ENODEV;
>
> - phys->paddr = virt_to_phys(priv->memory);
> - phys->len = priv->length;
> + phys->paddr = virt_to_phys(priv->memory) + offset;
> + phys->len = len;
> + *attr = DMA_BUF_PHYS_ATTR_MMIO;
> return 0;
> }
>
> +static const struct dma_buf_ops iommufd_test_dmabuf_ops = {
> + .attach = iommufd_test_dma_buf_attach,
> + .detach = iommufd_test_dma_buf_detach,
> + .map_dma_buf = iommufd_test_dma_buf_map,
> + .release = iommufd_test_dma_buf_release,
> + .unmap_dma_buf = iommufd_test_dma_buf_unmap,
> + .get_phys = iommufd_test_dma_buf_get_phys,
> +};
> +
> static int iommufd_test_dmabuf_get(struct iommufd_ucmd *ucmd,
> unsigned int open_flags,
> size_t len)
> @@ -2031,6 +2030,73 @@ static int iommufd_test_dmabuf_get(struct iommufd_ucmd *ucmd,
> return rc;
> }
>
> +/*
> + * Report the physical address the mock domain resolves @iova to, or 0 if
> + * it is unmapped. Lets a test check that two IOVAs share one frame (a
> + * scratch substitution) without knowing the frame in advance.
> + */
> +static int iommufd_test_md_iova_to_phys(struct iommufd_ucmd *ucmd,
> + unsigned int mockpt_id,
> + unsigned long iova)
> +{
> + struct iommu_test_cmd *cmd = ucmd->cmd;
> + struct iommufd_hw_pagetable *hwpt;
> + struct mock_iommu_domain *mock;
> + unsigned int page_size;
> + int rc;
> +
> + hwpt = get_md_pagetable(ucmd, mockpt_id, &mock);
> + if (IS_ERR(hwpt))
> + return PTR_ERR(hwpt);
> +
> + page_size = 1 << __ffs(mock->domain.pgsize_bitmap);
> + if (iova % page_size) {
> + rc = -EINVAL;
> + goto out_put;
> + }
> + cmd->iova_to_phys.out_phys =
> + mock->domain.ops->iova_to_phys(&mock->domain, iova);
> + rc = iommufd_ucmd_respond(ucmd, sizeof(*cmd));
> +out_put:
> + iommufd_put_object(ucmd->ictx, &hwpt->obj);
> + return rc;
> +}
> +
> +static int iommufd_test_md_check_mapped(struct iommufd_ucmd *ucmd,
> + unsigned int mockpt_id,
> + unsigned long iova, size_t length,
> + bool mapped)
> +{
> + struct iommufd_hw_pagetable *hwpt;
> + struct mock_iommu_domain *mock;
> + unsigned int page_size;
> + int rc = 0;
> +
> + hwpt = get_md_pagetable(ucmd, mockpt_id, &mock);
> + if (IS_ERR(hwpt))
> + return PTR_ERR(hwpt);
> +
> + page_size = 1 << __ffs(mock->domain.pgsize_bitmap);
> + if (iova % page_size || length % page_size || !length) {
> + rc = -EINVAL;
> + goto out_put;
> + }
> +
> + for (; length; length -= page_size, iova += page_size) {
> + bool is_mapped =
> + mock->domain.ops->iova_to_phys(&mock->domain, iova) != 0;
> +
> + if (is_mapped != mapped) {
> + rc = -ENOENT;
> + goto out_put;
> + }
> + }
> +
> +out_put:
> + iommufd_put_object(ucmd->ictx, &hwpt->obj);
> + return rc;
> +}
> +
> static int iommufd_test_dmabuf_revoke(struct iommufd_ucmd *ucmd, int fd,
> bool revoked)
> {
> @@ -2143,6 +2209,14 @@ int iommufd_test(struct iommufd_ucmd *ucmd)
> return iommufd_test_dmabuf_revoke(ucmd,
> cmd->dmabuf_revoke.dmabuf_fd,
> cmd->dmabuf_revoke.revoked);
> + case IOMMU_TEST_OP_MD_CHECK_MAPPED:
> + return iommufd_test_md_check_mapped(ucmd, cmd->id,
> + cmd->check_mapped.iova,
> + cmd->check_mapped.length,
> + cmd->check_mapped.mapped);
> + case IOMMU_TEST_OP_MD_IOVA_TO_PHYS:
> + return iommufd_test_md_iova_to_phys(ucmd, cmd->id,
> + cmd->iova_to_phys.iova);
> default:
> return -EOPNOTSUPP;
> }
> diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
> index c16f460c01d6..381c3d338e8c 100644
> --- a/drivers/vfio/pci/vfio_pci_dmabuf.c
> +++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
> @@ -99,46 +99,46 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf)
> kfree(priv);
> }
>
> -static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
> - .attach = vfio_pci_dma_buf_attach,
> - .map_dma_buf = vfio_pci_dma_buf_map,
> - .unmap_dma_buf = vfio_pci_dma_buf_unmap,
> - .release = vfio_pci_dma_buf_release,
> -};
> -
> /*
> - * This is a temporary "private interconnect" between VFIO DMABUF and iommufd.
> - * It allows the two co-operating drivers to exchange the physical address of
> - * the BAR. This is to be replaced with a formal DMABUF system for negotiated
> - * interconnect types.
> + * Report the BAR's physical range for importers which program it into their own
> + * translation tables, such as iommufd. A BAR is MMIO, never cache-coherent RAM.
> *
> - * If this function succeeds the following are true:
> - * - There is one physical range and it is pointing to MMIO
> - * - When move_notify is called it means revoke, not move, vfio_dma_buf_map
> - * will fail if it is currently revoked
> + * When move_notify is called it means revoke, not move, so this fails while
> + * revoked and vfio_dma_buf_map() does the same.
> */
> -int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
> - struct phys_vec *phys)
> +static int vfio_pci_dma_buf_get_phys(struct dma_buf_attachment *attachment,
> + u64 offset, u64 len,
> + struct phys_vec *phys, u32 *attr)
> {
> - struct vfio_pci_dma_buf *priv;
> + struct vfio_pci_dma_buf *priv = attachment->dmabuf->priv;
> + u32 i;
>
> dma_resv_assert_held(attachment->dmabuf->resv);
> -
> - if (attachment->dmabuf->ops != &vfio_pci_dmabuf_ops)
> - return -EOPNOTSUPP;
> -
> - priv = attachment->dmabuf->priv;
> if (priv->revoked)
> return -ENODEV;
>
> - /* More than one range to iommufd will require proper DMABUF support */
> - if (priv->nr_ranges != 1)
> - return -EOPNOTSUPP;
> -
> - *phys = priv->phys_vec[0];
> + /* Report from @offset to the end of the BAR range containing it. */
> + for (i = 0; i < priv->nr_ranges; i++) {
> + if (offset < priv->phys_vec[i].len)
> + break;
> + offset -= priv->phys_vec[i].len;
> + }
> + if (i == priv->nr_ranges)
> + return -EINVAL;
> + phys->paddr = priv->phys_vec[i].paddr + offset;
> + phys->len = min_t(u64, priv->phys_vec[i].len - offset, len);
> + *attr = DMA_BUF_PHYS_ATTR_MMIO;
> return 0;
> }
> -EXPORT_SYMBOL_FOR_MODULES(vfio_pci_dma_buf_iommufd_map, "iommufd");
> +
> +static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
> + .attach = vfio_pci_dma_buf_attach,
> + .map_dma_buf = vfio_pci_dma_buf_map,
> + .unmap_dma_buf = vfio_pci_dma_buf_unmap,
> + .release = vfio_pci_dma_buf_release,
> + .get_phys = vfio_pci_dma_buf_get_phys,
> +};
> +
>
> int vfio_pci_core_fill_phys_vec(struct phys_vec *phys_vec,
> struct vfio_region_dma_range *dma_ranges,
> diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h
> index d1203da56fc5..b223962e20c2 100644
> --- a/include/linux/dma-buf.h
> +++ b/include/linux/dma-buf.h
> @@ -13,6 +13,7 @@
> #ifndef __DMA_BUF_H__
> #define __DMA_BUF_H__
>
> +#include <linux/bitfield.h>
> #include <linux/iosys-map.h>
> #include <linux/file.h>
> #include <linux/err.h>
> @@ -23,6 +24,7 @@
> #include <linux/dma-fence.h>
> #include <linux/wait.h>
> #include <linux/pci-p2pdma.h>
> +#include <linux/types.h>
>
> struct device;
> struct dma_buf;
> @@ -182,6 +184,29 @@ struct dma_buf_ops {
> struct sg_table *,
> enum dma_data_direction);
>
> + /**
> + * @get_phys:
> + *
> + * Describe the run that starts at @offset, for an importer that
> + * programs its own translation tables. A run is the longest stretch
> + * of backed bytes whose frames are physically contiguous and share
> + * one attribute word. Report it in *@phys, starting at the byte at
> + * @offset and clipped at @offset + @len, with its DMA_BUF_PHYS_ATTR_*
> + * word in *@attr. The exporter may stop before the end of the run;
> + * importers must not assume the reported run is maximal.
> + *
> + * Return 0 on success, -ENOENT if the byte at @offset is not backed,
> + * -ENODEV if the buffer is revoked, or another negative error. Do not
> + * wait for memory to become available.
> + *
> + * The dma-buf reservation is held. The attachment must be pinned or
> + * have revocable importer operations. A reported frame remains valid
> + * until a covering invalidation callback returns; an exporter must
> + * invalidate every change before reusing an old frame.
> + */
> + int (*get_phys)(struct dma_buf_attachment *attach, u64 offset, u64 len,
> + struct phys_vec *phys, u32 *attr);
> +
> /* TODO: Add try_map_dma_buf version, to return immed with -EBUSY
> * if the call would block.
> */
> @@ -576,6 +601,35 @@ void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,
> enum dma_data_direction);
> void dma_buf_invalidate_mappings(struct dma_buf *dma_buf);
> bool dma_buf_attach_revocable(struct dma_buf_attachment *attach);
> +/* bits 0-7: memory type (a value, not flags) */
> +#define DMA_BUF_PHYS_ATTR_TYPE_MASK GENMASK(7, 0)
> +#define DMA_BUF_PHYS_ATTR_RAM 0x00 /* cache-coherent system RAM */
> +#define DMA_BUF_PHYS_ATTR_MMIO 0x01 /* device MMIO, uncached */
> +/* bits 8-15: reserved for a second value field; must be zero */
> +#define DMA_BUF_PHYS_ATTR_RSVD_MASK GENMASK(15, 8)
> +/* bits 16-31: flags; undefined bits must be zero */
> +#define DMA_BUF_PHYS_ATTR_READONLY BIT(16)
> +/* BIT(17): reserved (hole, for importers that walk across gaps) */
> +/* BIT(18): reserved (private, for confidential computing) */
> +#define DMA_BUF_PHYS_ATTR_FLAGS_MASK (DMA_BUF_PHYS_ATTR_READONLY)
> +
> +static inline u32 dma_buf_phys_attr_type(u32 attrs)
> +{
> + return FIELD_GET(DMA_BUF_PHYS_ATTR_TYPE_MASK, attrs);
> +}
> +
> +static inline bool dma_buf_phys_attr_known(u32 attrs)
> +{
> + return dma_buf_phys_attr_type(attrs) <= DMA_BUF_PHYS_ATTR_MMIO &&
> + !(attrs & DMA_BUF_PHYS_ATTR_RSVD_MASK) &&
> + !(attrs & ~(DMA_BUF_PHYS_ATTR_TYPE_MASK |
> + DMA_BUF_PHYS_ATTR_RSVD_MASK |
> + DMA_BUF_PHYS_ATTR_FLAGS_MASK));
> +}
> +
> +int dma_buf_get_phys(struct dma_buf_attachment *attach, u64 offset, u64 len,
> + struct phys_vec *phys, u32 *attr);
> +
> int dma_buf_begin_cpu_access(struct dma_buf *dma_buf,
> enum dma_data_direction dir);
> int dma_buf_end_cpu_access(struct dma_buf *dma_buf,
> diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
> index 9a1674c152aa..2a1d13abdb0c 100644
> --- a/include/linux/vfio_pci_core.h
> +++ b/include/linux/vfio_pci_core.h
> @@ -257,7 +257,4 @@ vfio_pci_core_get_iomap(struct vfio_pci_core_device *vdev, unsigned int bar)
> return vdev->barmap[bar];
> }
>
> -int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
> - struct phys_vec *phys);
> -
> #endif /* VFIO_PCI_CORE_H */
> diff --git a/samples/kvm/gmem_provider.c b/samples/kvm/gmem_provider.c
> index 75197c088762..b6824fe5d228 100644
> --- a/samples/kvm/gmem_provider.c
> +++ b/samples/kvm/gmem_provider.c
> @@ -487,35 +487,30 @@ static void gmem_dma_buf_release(struct dma_buf *dmabuf)
> kfree(priv);
> }
>
> -static const struct dma_buf_ops gmem_dma_buf_ops = {
> - .attach = gmem_dma_buf_attach,
> - .map_dma_buf = gmem_dma_buf_map,
> - .unmap_dma_buf = gmem_dma_buf_unmap,
> - .release = gmem_dma_buf_release,
> -};
> -
> -/*
> - * Private interconnect for iommufd (mirrors vfio_pci_dma_buf_iommufd_map).
> - * Returns the single contiguous phys range for the exported region so iommufd
> - * can program the IOMMU directly, bypassing the DMA API.
> - */
> -int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach,
> - struct phys_vec *phys);
> -int gmem_provider_dma_buf_iommufd_map(struct dma_buf_attachment *attach,
> - struct phys_vec *phys)
> +/* Report this flat sample region through the generic dma-buf operation. */
> +static int gmem_dma_buf_get_phys(struct dma_buf_attachment *attach,
> + u64 offset, u64 len,
> + struct phys_vec *phys, u32 *attr)
> {
> - struct gmem_dmabuf *priv;
> + struct gmem_dmabuf *priv = attach->dmabuf->priv;
>
> dma_resv_assert_held(attach->dmabuf->resv);
> - if (attach->dmabuf->ops != &gmem_dma_buf_ops)
> - return -EOPNOTSUPP;
> - priv = attach->dmabuf->priv;
> if (priv->revoked)
> return -ENODEV;
> - *phys = priv->phys;
> +
> + phys->paddr = priv->phys.paddr + offset;
> + phys->len = len;
> + *attr = DMA_BUF_PHYS_ATTR_RAM;
> return 0;
> }
> -EXPORT_SYMBOL_FOR_MODULES(gmem_provider_dma_buf_iommufd_map, "iommufd");
> +
> +static const struct dma_buf_ops gmem_dma_buf_ops = {
> + .attach = gmem_dma_buf_attach,
> + .map_dma_buf = gmem_dma_buf_map,
> + .unmap_dma_buf = gmem_dma_buf_unmap,
> + .release = gmem_dma_buf_release,
> + .get_phys = gmem_dma_buf_get_phys,
> +};
>
> /* Called with info->dmabufs_lock held on the revoke path. */
> static void gmem_dma_buf_revoke_all(struct gmem_info *info)
> --
> 2.47.3
>
next prev parent reply other threads:[~2026-10-05 10:07 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 03/11] KVM: SEV: Remove struct page dependency from SNP gmem paths David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 04/11] KVM: guest_memfd: Introduce guest memory ops and route native gmem through them David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 05/11] iommufd: Look up private-interconnect phys via exporter symbols David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 06/11] iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 07/11] KVM: guest_memfd: Add ops-driven page revocation David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 08/11] samples/kvm: Add guest_memfd backing sample David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 09/11] selftests/kvm: gmem_provider KVM-only tests David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 10/11] selftests/kvm: gmem_provider iommufd tests David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 11/11] samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64 David Woodhouse
2026-07-20 15:11 ` [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory Paolo Bonzini
2026-07-20 16:39 ` David Woodhouse
2026-07-23 0:24 ` Ackerley Tng
2026-07-23 9:40 ` David Woodhouse
2026-07-23 16:01 ` Ackerley Tng
2026-10-05 9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run Fred Griffoul
2026-10-05 10:07 ` Christian König [this message]
2026-10-05 13:20 ` Fred Griffoul
2026-10-05 14:53 ` Christian König
2026-10-05 9:55 ` [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation Fred Griffoul
2026-10-05 10:08 ` Christian König
2026-10-05 9:55 ` [RFC PATCH 4/6] dma-buf: Allow dynamic attach without a device Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 5/6] KVM: guest_memfd: Add dma-buf backing Fred Griffoul
2026-10-05 9:55 ` [RFC PATCH 6/6] samples/kvm, selftests/kvm: Exercise " Fred Griffoul
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a04f17cc-f59a-4240-81c8-29776eee3a2e@amd.com \
--to=christian.koenig@amd.com \
--cc=ackerleytng@google.com \
--cc=alex@shazbot.org \
--cc=bp@alien8.de \
--cc=catalin.marinas@arm.com \
--cc=dave.hansen@linux.intel.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=dwmw2@infradead.org \
--cc=griffoul@gmail.com \
--cc=hpa@zytor.com \
--cc=iommu@lists.linux.dev \
--cc=jgg@ziepe.ca \
--cc=joey.gouly@arm.com \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=kvmarm@lists.linux.dev \
--cc=linaro-mm-sig@lists.linaro.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-media@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=maz@kernel.org \
--cc=mhiramat@kernel.org \
--cc=mingo@redhat.com \
--cc=oupton@kernel.org \
--cc=pbonzini@redhat.com \
--cc=robin.murphy@arm.com \
--cc=rostedt@goodmis.org \
--cc=seanjc@google.com \
--cc=seiden@linux.ibm.com \
--cc=shuah@kernel.org \
--cc=sumit.semwal@linaro.org \
--cc=suzuki.poulose@arm.com \
--cc=tglx@kernel.org \
--cc=will@kernel.org \
--cc=x86@kernel.org \
--cc=yuzenghui@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®