From: <mhonap@nvidia.com>
To: <alex@shazbot.org>, <jgg@ziepe.ca>, <ankita@nvidia.com>,
<jic23@kernel.org>, <dave.jiang@intel.com>,
<alejandro.lucero-palau@amd.com>, <smadhavan@nvidia.com>,
<corbet@lwn.net>, <skhan@linuxfoundation.org>,
<dave@stgolabs.net>, <alison.schofield@intel.com>,
<vishal.l.verma@intel.com>, <iweiny@kernel.org>,
<ming.li@zohomail.com>, <yishaih@nvidia.com>,
<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
<bhelgaas@google.com>, <dmatlack@google.com>, <kees@kernel.org>,
<gustavoars@kernel.org>
Cc: <cjia@nvidia.com>, <kjaju@nvidia.com>, <vsethi@nvidia.com>,
<zhiw@nvidia.com>, <mhonap@nvidia.com>,
<linux-doc@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<kvm@vger.kernel.org>, <linux-cxl@vger.kernel.org>,
<linux-pci@vger.kernel.org>, <linux-kselftest@vger.kernel.org>,
<linux-hardening@vger.kernel.org>
Subject: [PATCH v5 24/27] vfio/cxl: Export the HDM memory region as a dma-buf
Date: Thu, 17 Sep 2026 00:05:37 +0530 [thread overview]
Message-ID: <20260916183540.3813685-25-mhonap@nvidia.com> (raw)
In-Reply-To: <20260916183540.3813685-1-mhonap@nvidia.com>
From: Manish Honap <mhonap@nvidia.com>
A Type-2 accelerator issues ATS-translated DMA to addresses inside its
own HDM window, so that coherent host range must be present in the
guest's IOAS (the iommufd IOAS backing the nested SMMU stage-2). iommufd
maps a struct-page-less range only by fd, via IOMMU_IOAS_MAP_FILE over a
dma-buf; a userspace-VA IOMMU_IOAS_MAP of the HDM mmap is rejected
because the VMA is VM_IO | VM_PFNMAP. Without a dma-buf the range could
only be mapped through an out-of-tree PFNMAP work-around.
vfio-pci already exports BAR memory as a P2P dma-buf, but the exporter
is BAR-only: vfio_pci_core_feature_dma_buf() rejects any region index at
or above the ROM index, and vfio_pci_core_get_dmabuf_phys() resolves the
physical range from a PCI BAR. The HDM memory region is a dynamic
device-specific region, not a BAR.
Let a device-specific region reach the device's get_dmabuf_phys(): a
region index at or above VFIO_PCI_NUM_REGIONS skips the BAR-resource
check and is validated by the driver instead, bounded to the regions
that exist. Install a CXL-aware get_dmabuf_phys() in the vfio-cxl
provider that returns cxl->hpa_range for the HDM memory region and
delegates real BARs to the core, keeping the BAR path unchanged and the
core free of CXL knowledge.
The HDM window is coherent host memory with no p2pdma provider of its
own, so borrow BAR 0's, matching nvgrace-gpu's handling of its non-BAR
device memory. The iommufd importer does not consume the provider; the
scatterlist map path (real peer DMA) is left to a follow-up once
upstream grows a negotiated interconnect for coherent CXL memory.
Assisted-by: LLM
Signed-off-by: Manish Honap <mhonap@nvidia.com>
---
drivers/vfio/pci/cxl/vfio_cxl_core.c | 60 ++++++++++++++++++++++++++++
drivers/vfio/pci/vfio_pci_dmabuf.c | 27 +++++++++++--
2 files changed, 83 insertions(+), 4 deletions(-)
diff --git a/drivers/vfio/pci/cxl/vfio_cxl_core.c b/drivers/vfio/pci/cxl/vfio_cxl_core.c
index 5fe8e35c63c4..55fa1f86850d 100644
--- a/drivers/vfio/pci/cxl/vfio_cxl_core.c
+++ b/drivers/vfio/pci/cxl/vfio_cxl_core.c
@@ -11,6 +11,7 @@
#include <linux/mm.h>
#include <linux/module.h>
#include <linux/pci.h>
+#include <linux/pci-p2pdma.h>
#include <linux/range.h>
#include <linux/slab.h>
#include <linux/uaccess.h>
@@ -27,6 +28,7 @@
* @hdm_regs: mapped HDM decoder registers, read live by the decoder region
* @hdm_len: length of the HDM decoder register block
* @hdm_valid: true when host CPU access to the HDM range is safe; under memory_lock
+ * @mem_region_index: vfio region index of the mmap-able HDM memory region
*/
struct vfio_cxl_state {
struct cxl_dev_state cxlds;
@@ -36,6 +38,7 @@ struct vfio_cxl_state {
void __iomem *hdm_regs;
u32 hdm_len;
bool hdm_valid;
+ unsigned int mem_region_index;
};
static unsigned long vfio_cxl_mem_pgoff(struct vm_area_struct *vma,
@@ -289,6 +292,50 @@ static void vfio_cxl_release_hpa(void *data)
release_mem_region(cxl->hpa_range.start, range_len(&cxl->hpa_range));
}
+/*
+ * Resolve the physical range that backs a dma-buf export. The core exporter
+ * only knows BARs; teach it the HDM memory region so a guest IOAS can map the
+ * coherent window by fd (IOMMU_IOAS_MAP_FILE) instead of the removed PFNMAP
+ * work-around. Real BARs stay on the byte-identical core path.
+ */
+static int vfio_cxl_get_dmabuf_phys(struct vfio_pci_core_device *vdev,
+ struct p2pdma_provider **provider,
+ unsigned int region_index,
+ struct phys_vec *phys_vec,
+ struct vfio_region_dma_range *dma_ranges,
+ size_t nr_ranges)
+{
+ struct vfio_cxl_state *cxl = vdev->cxl;
+
+ /* Real BARs go through the core P2P exporter unchanged. */
+ if (region_index < VFIO_PCI_NUM_REGIONS)
+ return vfio_pci_core_get_dmabuf_phys(vdev, provider,
+ region_index, phys_vec,
+ dma_ranges, nr_ranges);
+
+ /* Of the device regions, only the HDM memory window is exportable. */
+ if (region_index != cxl->mem_region_index)
+ return -EINVAL;
+
+ /*
+ * The HDM window is coherent host memory, not BAR MMIO, so it has no
+ * p2pdma provider of its own. Borrow BAR 0's: the P2P properties match
+ * and the iommufd importer does not consume the provider. The sgt map
+ * path (real peer DMA) is not supported for the HDM window.
+ */
+ *provider = pcim_p2pdma_provider(vdev->pdev, 0);
+ if (!*provider)
+ return -EINVAL;
+
+ return vfio_pci_core_fill_phys_vec(phys_vec, dma_ranges, nr_ranges,
+ cxl->hpa_range.start,
+ range_len(&cxl->hpa_range));
+}
+
+static const struct vfio_pci_device_ops vfio_cxl_pci_dev_ops = {
+ .get_dmabuf_phys = vfio_cxl_get_dmabuf_phys,
+};
+
static int vfio_cxl_init_device(struct vfio_pci_core_device *vdev)
{
struct pci_dev *pdev = vdev->pdev;
@@ -483,6 +530,19 @@ static int vfio_cxl_open_device(struct vfio_pci_core_device *vdev)
if (ret)
return ret;
+ /* Record where the HDM memory region landed for the dma-buf export. */
+ cxl->mem_region_index = VFIO_PCI_NUM_REGIONS + vdev->num_regions - 1;
+
+ /*
+ * Override the device ops so a dma-buf export of the HDM memory region
+ * resolves to the coherent host range. This is done at open, not init:
+ * vfio_pci_probe() resets pci_ops after vfio_alloc_device() returns, so
+ * an override installed during init would be clobbered. Only a CXL device
+ * reaches this hook (cxl_ops is set on init success), so a fallback to
+ * plain vfio-pci keeps the core ops.
+ */
+ vdev->pci_ops = &vfio_cxl_pci_dev_ops;
+
ret = vfio_cxl_add_region(vdev, VFIO_REGION_SUBTYPE_CXL_COMP_REGS,
&vfio_cxl_comp_regops, cxl->hdm_len,
VFIO_REGION_INFO_FLAG_READ |
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c16f460c01d6..436c616d5b66 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -178,6 +178,15 @@ int vfio_pci_core_get_dmabuf_phys(struct vfio_pci_core_device *vdev,
{
struct pci_dev *pdev = vdev->pdev;
+ /*
+ * This resolver only handles PCI BARs. A device-specific region index
+ * (>= PCI_STD_NUM_BARS) would index pdev->resource[] out of bounds via
+ * pcim_p2pdma_provider(), so reject it; a driver that exports such a
+ * region installs its own get_dmabuf_phys.
+ */
+ if (region_index >= PCI_STD_NUM_BARS)
+ return -EINVAL;
+
*provider = pcim_p2pdma_provider(pdev, region_index);
if (!*provider)
return -EINVAL;
@@ -227,6 +236,7 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
struct vfio_pci_dma_buf *priv;
size_t length;
+ u32 index;
int ret;
if (!vdev->pci_ops || !vdev->pci_ops->get_dmabuf_phys)
@@ -243,13 +253,22 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
if (!get_dma_buf.nr_ranges || get_dma_buf.flags)
return -EINVAL;
+ index = get_dma_buf.region_index;
+
/*
- * For PCI the region_index is the BAR number like everything
- * else. Check that PCI resources have been claimed for it.
+ * A fixed region index is the BAR number; only a BAR can be exported
+ * and its PCI resource must be claimed. A device-specific region (index
+ * >= VFIO_PCI_NUM_REGIONS) has no BAR resource and is validated by the
+ * device's get_dmabuf_phys instead, but the index must name a region
+ * that exists.
*/
- if (get_dma_buf.region_index >= VFIO_PCI_ROM_REGION_INDEX ||
- IS_ERR(vfio_pci_core_get_iomap(vdev, get_dma_buf.region_index)))
+ if (index < VFIO_PCI_NUM_REGIONS) {
+ if (index >= VFIO_PCI_ROM_REGION_INDEX ||
+ IS_ERR(vfio_pci_core_get_iomap(vdev, index)))
+ return -ENODEV;
+ } else if (index - VFIO_PCI_NUM_REGIONS >= vdev->num_regions) {
return -ENODEV;
+ }
dma_ranges = memdup_array_user(&arg->dma_ranges, get_dma_buf.nr_ranges,
sizeof(*dma_ranges));
--
2.25.1
next prev parent reply other threads:[~2026-09-16 18:40 UTC|newest]
Thread overview: 31+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 18:35 [PATCH v5 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-09-16 18:35 ` [PATCH v5 01/27] cxl/regs: Split the BAR block request and ioremap helpers mhonap
2026-09-16 18:35 ` [PATCH v5 02/27] cxl/regs: Let a BAR-owning driver own the component register block mhonap
2026-09-16 18:35 ` [PATCH v5 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-09-16 18:35 ` [PATCH v5 04/27] cxl: Add cxl_reset_dvsec_sequence() for vfio-pci mhonap
2026-09-16 18:35 ` [PATCH v5 05/27] vfio/pci: Add the CXL provider ops registration interface mhonap
2026-09-16 18:35 ` [PATCH v5 06/27] vfio/pci: Detect CXL devices and load the CXL provider on demand mhonap
2026-09-16 18:35 ` [PATCH v5 07/27] vfio/pci: Honor -EPROBE_DEFER from CXL provider probe mhonap
2026-09-16 18:35 ` [PATCH v5 08/27] vfio/pci: Fall back to plain vfio-pci when CXL init fails mhonap
2026-09-16 18:35 ` [PATCH v5 09/27] vfio/pci: Add a generic excluded-range list mhonap
2026-09-16 18:35 ` [PATCH v5 10/27] vfio/pci: Migrate MSI-X exclusion onto the " mhonap
2026-09-16 18:35 ` [PATCH v5 11/27] vfio/pci: Virtualize the CXL DVSEC in vfio_pci_config.c mhonap
2026-09-16 18:35 ` [PATCH v5 12/27] vfio/pci: Call the CXL open and close hooks around device use mhonap
2026-09-16 18:35 ` [PATCH v5 13/27] vfio/pci: Bracket PCI resets with the CXL reset hooks mhonap
2026-09-16 18:35 ` [PATCH v5 14/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-09-16 18:35 ` [PATCH v5 15/27] vfio/cxl: Add the vfio-cxl provider module skeleton mhonap
2026-09-16 18:35 ` [PATCH v5 16/27] vfio/cxl: Create the CXL memdev and set media ready at bind mhonap
2026-09-16 18:35 ` [PATCH v5 17/27] vfio/cxl: Own the whole component register BAR mhonap
2026-09-16 18:35 ` [PATCH v5 18/27] vfio/cxl: Expose the HDM memory region to the guest mhonap
2026-09-16 18:35 ` [PATCH v5 19/27] vfio/cxl: Contain HDM memory errors with memory_failure() mhonap
2026-09-16 18:35 ` [PATCH v5 20/27] vfio/cxl: Expose the HDM decoder registers read-only to the guest mhonap
2026-09-16 18:35 ` [PATCH v5 21/27] vfio/cxl: Exclude the HDM decoder registers from direct BAR access mhonap
2026-09-17 7:28 ` Richard Cheng
2026-09-16 18:35 ` [PATCH v5 22/27] vfio/cxl: Clear the HDM access gate after a hot reset mhonap
2026-09-16 18:35 ` [PATCH v5 23/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-09-16 18:35 ` mhonap [this message]
2026-09-17 7:55 ` [PATCH v5 24/27] vfio/cxl: Export the HDM memory region as a dma-buf Richard Cheng
2026-09-16 18:35 ` [PATCH v5 25/27] vfio/cxl: Run the CXL reset at the vfio reset points mhonap
2026-09-16 18:35 ` [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-09-16 19:33 ` Gregory Price
2026-09-16 18:35 ` [PATCH v5 27/27] selftests/vfio: Add CXL Type-2 passthrough tests mhonap
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260916183540.3813685-25-mhonap@nvidia.com \
--to=mhonap@nvidia.com \
--cc=alejandro.lucero-palau@amd.com \
--cc=alex@shazbot.org \
--cc=alison.schofield@intel.com \
--cc=ankita@nvidia.com \
--cc=bhelgaas@google.com \
--cc=cjia@nvidia.com \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=dmatlack@google.com \
--cc=gustavoars@kernel.org \
--cc=iweiny@kernel.org \
--cc=jgg@ziepe.ca \
--cc=jic23@kernel.org \
--cc=kees@kernel.org \
--cc=kevin.tian@intel.com \
--cc=kjaju@nvidia.com \
--cc=kvm@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-hardening@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=ming.li@zohomail.com \
--cc=skhan@linuxfoundation.org \
--cc=skolothumtho@nvidia.com \
--cc=smadhavan@nvidia.com \
--cc=vishal.l.verma@intel.com \
--cc=vsethi@nvidia.com \
--cc=yishaih@nvidia.com \
--cc=zhiw@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®