* [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources
@ 2026-10-10 7:24 Aneesh Kumar K.V (Arm)
2026-10-10 7:24 ` [RFC PATCH v1 1/6] KVM: guest_memfd: attach and bind device resources Aneesh Kumar K.V (Arm)
` (5 more replies)
0 siblings, 6 replies; 7+ messages in thread
From: Aneesh Kumar K.V (Arm) @ 2026-10-10 7:24 UTC (permalink / raw)
To: linux-coco, linux-kernel
Cc: Aneesh Kumar K.V (Arm),
Ackerley Tng, Alex Williamson, David Woodhouse,
David Hildenbrand, Jason Gunthorpe, Joerg Roedel (AMD),
Kevin Tian, Paolo Bonzini, Robin Murphy, Sean Christopherson,
Will Deacon, Alexey Kardashevskiy, Xu Yilun, Catalin Marinas,
Suzuki K Poulose, Steven Price, Fred Griffoul, iommu, kvm
This series adds a guest_memfd device-resource framework for exposing PCI
BAR memory to guests. The intended use is confidential device assignment,
where KVM needs to coordinate private device mappings with shared access
through VFIO. The Arm CCA backend is provided separately; these patches
contain the guest_memfd, VFIO and IOMMUFD interfaces it uses.
The guest_memfd framework and Arm CCA backend are available together at:
https://git.gitlab.arm.com/linux-arm/linux-cca.git cca/topics/cca-da-gmem-backend
Userspace creates a device-backed guest_memfd using the ordinary VFIO
device fd as resource_fd, with GUEST_MEMFD_FLAG_USE_RESOURCE,
GUEST_MEMFD_FLAG_DEVICE and GUEST_MEMFD_FLAG_INIT_SHARED. Eligible BAR
intervals are registered as guest_memfd memslots. File offsets use the
existing VFIO region-offset namespace, allowing one guest_memfd to cover
multiple BAR intervals. Device-backed guest_memfd rejects mmap and
fallocate; shared guest faults obtain bare BAR PFNs from the provider
instead of resolving the memslot HVA or allocating folios.
IOMMUFD resolves and pins the device's existing vDEVICE for the VM and
serializes MMIO conversion through its MMIO mutex. VFIO supplies BAR PFNs,
checks eligible ranges and excludes independent dma-buf exports while a
guest_memfd provider is attached. VFIO host mmap faults, read/write and
ioeventfd accesses check the requested range's guest_memfd attributes.
Conversion revokes cached host mappings and shared guest mappings before
installing private mappings. Other shared intervals remain accessible.
This series depends on:
* Ackerley Tng's "Allow guest_memfd to be created using a resource (pool)
fd" RFC, which provides the tmpfs resource-provider API:
https://lore.kernel.org/all/20260925-gmem-tmpfs-backend-v1-0-d36159822d18@google.com/
* The IOMMUFD/TSM series, which provides the vIOMMU and vDEVICE interfaces:
https://lore.kernel.org/all/20261008055955.4014342-1-aneesh.kumar@kernel.org/
Notes:
* While developing this framework, I looked at David Woodhouse's
guest_memfd provider work, including the gmem-provider-v3 branch. It
was used for comparison and was not merged. The posted RFC v2 is here:
https://lore.kernel.org/all/20260720111259.122911-1-dwmw2@infradead.org/
* Codex assisted with implementation, review and series organization.
Cc: Ackerley Tng <ackerleytng@google.com>
Cc: Alex Williamson <alex@shazbot.org>
Cc: David Woodhouse <dwmw2@infradead.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: "Joerg Roedel (AMD)" <joro@8bytes.org>
Cc: Kevin Tian <kevin.tian@intel.com>
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: Robin Murphy <robin.murphy@arm.com>
Cc: Sean Christopherson <seanjc@google.com>
Cc: Will Deacon <will@kernel.org>
Cc: Alexey Kardashevskiy <aik@amd.com>
Cc: Xu Yilun <yilun.xu@linux.intel.com>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Suzuki K Poulose <suzuki.poulose@arm.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Fred Griffoul <griffoul@gmail.com>
Cc: iommu@lists.linux.dev
Cc: kvm@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Aneesh Kumar K.V (Arm) (6):
KVM: guest_memfd: attach and bind device resources
KVM: guest_memfd: Support private/shared conversion of device memory
iommufd: Support MMIO provider attachment to vDEVICEs
vfio/pci: Provide guest_memfd backing for PCI BARs
KVM/VFIO: Remove device mappings during guest_memfd teardown
KVM: Invalidate guest_memfd mappings before removing memslot bindings
Documentation/virt/kvm/api.rst | 4 +
drivers/iommu/iommufd/viommu.c | 90 ++++
drivers/vfio/device_cdev.c | 3 +-
drivers/vfio/group.c | 4 +-
drivers/vfio/pci/Kconfig | 11 +
drivers/vfio/pci/Makefile | 2 +
drivers/vfio/pci/vfio_pci.c | 3 +
drivers/vfio/pci/vfio_pci_core.c | 13 +
drivers/vfio/pci/vfio_pci_dmabuf.c | 10 +
drivers/vfio/pci/vfio_pci_gmem.c | 349 +++++++++++++++
drivers/vfio/pci/vfio_pci_gmem.h | 36 ++
drivers/vfio/pci/vfio_pci_gmem_access.c | 138 ++++++
drivers/vfio/pci/vfio_pci_rdwr.c | 13 +-
drivers/vfio/vfio.h | 8 +
drivers/vfio/vfio_main.c | 34 +-
include/linux/guest_memfd.h | 50 +++
include/linux/iommufd.h | 33 ++
include/linux/kvm_host.h | 14 +-
include/linux/vfio.h | 3 +
include/linux/vfio_pci_core.h | 23 +
include/uapi/linux/kvm.h | 1 +
virt/kvm/guest_memfd.c | 545 +++++++++++++++++++++++-
virt/kvm/guest_memfd.h | 6 +
virt/kvm/kvm_main.c | 4 +
24 files changed, 1362 insertions(+), 35 deletions(-)
create mode 100644 drivers/vfio/pci/vfio_pci_gmem.c
create mode 100644 drivers/vfio/pci/vfio_pci_gmem.h
create mode 100644 drivers/vfio/pci/vfio_pci_gmem_access.c
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH v1 1/6] KVM: guest_memfd: attach and bind device resources
2026-10-10 7:24 [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources Aneesh Kumar K.V (Arm)
@ 2026-10-10 7:24 ` Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 2/6] KVM: guest_memfd: Support private/shared conversion of device memory Aneesh Kumar K.V (Arm)
` (4 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Aneesh Kumar K.V (Arm) @ 2026-10-10 7:24 UTC (permalink / raw)
To: linux-coco, linux-kernel
Cc: Aneesh Kumar K.V (Arm),
Ackerley Tng, Alex Williamson, David Woodhouse,
David Hildenbrand, Jason Gunthorpe, Joerg Roedel (AMD),
Kevin Tian, Paolo Bonzini, Robin Murphy, Sean Christopherson,
Will Deacon, Alexey Kardashevskiy, Xu Yilun, Catalin Marinas,
Suzuki K Poulose, Steven Price, Fred Griffoul, iommu, kvm
Introduce the device resource flag and registration of a provider file
operations identity. KVM checks file->f_op before interpreting
private_data as guest_memfd_device_operations.
Add attach() to establish the provider, bind() to associate eligible
file offsets with memslots, get_pfn() for shared device PFNs, release()
for the provider lifetime.
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: Sean Christopherson <seanjc@google.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: kvm@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Assisted-by: Codex
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
---
Documentation/virt/kvm/api.rst | 4 +
include/linux/guest_memfd.h | 34 +++++++
include/linux/kvm_host.h | 9 +-
include/uapi/linux/kvm.h | 1 +
virt/kvm/guest_memfd.c | 167 ++++++++++++++++++++++++++++++---
5 files changed, 202 insertions(+), 13 deletions(-)
diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index 5ece28765d2c..a4abee9f1ce1 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6527,6 +6527,10 @@ specified via KVM_CREATE_GUEST_MEMFD. Currently defined flags:
must be a supported resource (e.g. tmpfs mount
root directory). Currently, tmpfs resource pools
require noswap and huge=never mount options.
+ GUEST_MEMFD_FLAG_DEVICE Use a device resource file as resource_fd. Requires
+ GUEST_MEMFD_FLAG_USE_RESOURCE and
+ GUEST_MEMFD_FLAG_INIT_SHARED; mmap is disallowed.
+ The fd must expose guest_memfd device operations.
============================= ================================================
When the KVM MMU performs a PFN lookup to service a guest fault, the fault will
diff --git a/include/linux/guest_memfd.h b/include/linux/guest_memfd.h
index 60eb4f008c24..4f232612a525 100644
--- a/include/linux/guest_memfd.h
+++ b/include/linux/guest_memfd.h
@@ -5,8 +5,42 @@
#include <linux/types.h>
struct file;
+struct file_operations;
struct folio;
struct mempolicy;
+struct kvm;
+struct guest_memfd_device;
+
+/* Binding data stays valid only while the device guest_memfd is active. */
+struct guest_memfd_device_binding {
+ /* Architecture-specific mapping owner, e.g. an RMM vDEVICE handle. */
+ unsigned long owner;
+ u64 vdev_id;
+};
+
+/* Embedded by a device provider and returned from attach(). */
+struct guest_memfd_device_context {
+ struct guest_memfd_device *device;
+};
+
+struct guest_memfd_device_operations {
+ struct guest_memfd_device_context *(*attach)(struct file *resource,
+ struct guest_memfd_device *gdev);
+ int (*bind)(void *data, u64 offset, u64 size,
+ struct guest_memfd_device_binding *binding);
+ int (*get_pfn)(void *data, u64 offset, unsigned long *pfn);
+ void (*release)(void *data);
+};
+
+/* A registered file type stores device operations in file->private_data. */
+int guest_memfd_register_device_fops(const struct file_operations *fops);
+void guest_memfd_unregister_device_fops(const struct file_operations *fops);
+
+struct guest_memfd_device {
+ struct kvm *kvm;
+ u64 size;
+ void *core;
+};
/**
* struct guest_memfd_provider_operations - Operations for external memory providers
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 1dda54e8dee8..f1675a3e00b6 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -56,6 +56,7 @@
*/
#define KVM_MEMSLOT_INVALID (1UL << 16)
#define KVM_MEMSLOT_GMEM_ONLY (1UL << 17)
+#define KVM_MEMSLOT_GMEM_DEVICE (1UL << 18)
/*
* Bit 63 of the memslot generation number is an "update in-progress flag",
@@ -2605,16 +2606,22 @@ static inline bool kvm_is_private_gfn(struct kvm *kvm, gfn_t gfn)
#endif /* kvm_arch_has_private_mem */
#ifdef CONFIG_KVM_GUEST_MEMFD
+struct guest_memfd_device_binding;
+
bool kvm_gmem_is_private_gfn(struct kvm *kvm, gfn_t gfn);
bool kvm_gmem_range_has_attributes(struct kvm_memory_slot *slot, gfn_t start,
gfn_t end, u64 attributes);
+bool kvm_arch_gmem_device_supported(struct kvm *kvm);
+void kvm_arch_gmem_device_bind(struct kvm_memory_slot *slot,
+ const struct guest_memfd_device_binding *binding);
int kvm_gmem_set_attributes(struct kvm_memory_slot *slot, gfn_t start,
gfn_t end, u64 attributes);
bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm);
static inline u64 kvm_gmem_get_supported_flags(struct kvm *kvm)
{
- u64 flags = GUEST_MEMFD_FLAG_MMAP | GUEST_MEMFD_FLAG_USE_RESOURCE;
+ u64 flags = GUEST_MEMFD_FLAG_MMAP | GUEST_MEMFD_FLAG_USE_RESOURCE |
+ GUEST_MEMFD_FLAG_DEVICE;
if (!kvm || kvm_arch_supports_gmem_init_shared(kvm))
flags |= GUEST_MEMFD_FLAG_INIT_SHARED;
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index 9408ddde164a..b0d853f88ca8 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1681,6 +1681,7 @@ struct kvm_memory_attributes2 {
#define GUEST_MEMFD_FLAG_MMAP (1ULL << 0)
#define GUEST_MEMFD_FLAG_INIT_SHARED (1ULL << 1)
#define GUEST_MEMFD_FLAG_USE_RESOURCE (1ULL << 2)
+#define GUEST_MEMFD_FLAG_DEVICE (1ULL << 3)
struct kvm_create_guest_memfd {
__u64 size;
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index a65a1eda8b90..c1fd232a83dd 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -2,11 +2,13 @@
#include <linux/anon_inodes.h>
#include <linux/backing-dev.h>
#include <linux/falloc.h>
+#include <linux/export.h>
#include <linux/fs.h>
#include <linux/guest_memfd.h>
#include <linux/kvm_host.h>
#include <linux/maple_tree.h>
#include <linux/mempolicy.h>
+#include <linux/mutex.h>
#include <linux/pseudo_fs.h>
#include <linux/pagemap.h>
#include <linux/swap.h>
@@ -16,6 +18,46 @@
static struct vfsmount *kvm_gmem_mnt;
+static const struct file_operations *kvm_gmem_device_fops;
+static DEFINE_MUTEX(kvm_gmem_device_fops_lock);
+
+int guest_memfd_register_device_fops(const struct file_operations *fops)
+{
+ int ret = 0;
+
+ mutex_lock(&kvm_gmem_device_fops_lock);
+ if (kvm_gmem_device_fops)
+ ret = -EBUSY;
+ else
+ kvm_gmem_device_fops = fops;
+ mutex_unlock(&kvm_gmem_device_fops_lock);
+
+ return ret;
+}
+EXPORT_SYMBOL_GPL(guest_memfd_register_device_fops);
+
+void guest_memfd_unregister_device_fops(const struct file_operations *fops)
+{
+ mutex_lock(&kvm_gmem_device_fops_lock);
+ if (kvm_gmem_device_fops == fops)
+ kvm_gmem_device_fops = NULL;
+ mutex_unlock(&kvm_gmem_device_fops_lock);
+}
+EXPORT_SYMBOL_GPL(guest_memfd_unregister_device_fops);
+
+static const struct guest_memfd_device_operations *
+kvm_gmem_device_ops(struct file *file)
+{
+ const struct guest_memfd_device_operations *ops = NULL;
+
+ mutex_lock(&kvm_gmem_device_fops_lock);
+ /* The caller's file reference keeps private_data alive. */
+ if (file->f_op == kvm_gmem_device_fops)
+ ops = file->private_data;
+ mutex_unlock(&kvm_gmem_device_fops_lock);
+ return ops;
+}
+
/*
* A guest_memfd instance can be associated multiple VMs, each with its own
* "view" of the underlying physical memory.
@@ -38,6 +80,8 @@ struct gmem_inode {
struct inode vfs_inode;
struct list_head gmem_file_list;
+ struct guest_memfd_device device;
+ const struct guest_memfd_device_operations *device_ops;
void *provider;
const struct guest_memfd_provider_operations *provider_ops;
@@ -75,6 +119,10 @@ static inline void gmem_provider_invalidate_folio(struct gmem_inode *gi,
static inline void gmem_provider_release(struct gmem_inode *gi)
{
+ if (gi->device_ops) {
+ gi->device_ops->release(gi->provider);
+ kvm_put_kvm(gi->device.kvm);
+ }
if (gi->provider_ops && gi->provider_ops->release)
gi->provider_ops->release(gi->provider);
}
@@ -189,6 +237,9 @@ static struct folio *kvm_gmem_get_folio(struct inode *inode, pgoff_t index)
struct mempolicy *policy;
struct folio *folio;
+ if (gi->device_ops)
+ return ERR_PTR(-EOPNOTSUPP);
+
/*
* Fast-path: See if folio is already present in mapping to avoid
* policy_lookup.
@@ -237,6 +288,8 @@ static struct folio *kvm_gmem_get_folio(struct inode *inode, pgoff_t index)
static enum kvm_gfn_range_filter kvm_gmem_get_all_gfns_filter(struct inode *inode)
{
+ if (GMEM_I(inode)->device_ops)
+ return KVM_FILTER_SHARED;
if (gmem_in_place_conversion)
return KVM_FILTER_SHARED | KVM_FILTER_PRIVATE;
@@ -393,6 +446,9 @@ static long kvm_gmem_fallocate(struct file *file, int mode, loff_t offset,
{
int ret;
+ if (GMEM_I(file_inode(file))->device_ops)
+ return -EOPNOTSUPP;
+
if (!(mode & FALLOC_FL_KEEP_SIZE))
return -EOPNOTSUPP;
@@ -436,8 +492,9 @@ static int kvm_gmem_release(struct inode *inode, struct file *file)
filemap_invalidate_lock(inode->i_mapping);
- xa_for_each(&f->bindings, index, slot)
+ xa_for_each(&f->bindings, index, slot) {
WRITE_ONCE(slot->gmem.file, NULL);
+ }
/*
* All in-flight operations are gone and new bindings can be created.
@@ -477,6 +534,8 @@ DEFINE_CLASS(gmem_get_file, struct file *, if (_T) fput(_T),
static bool kvm_gmem_supports_mmap(struct inode *inode)
{
+ if (GMEM_I(inode)->device_ops)
+ return false;
return GMEM_I(inode)->flags & GUEST_MEMFD_FLAG_MMAP;
}
@@ -733,6 +792,9 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
struct ma_state mas;
int r = 0;
+ if (gi->device_ops)
+ return -EOPNOTSUPP;
+
mt = &gi->attributes;
filemap_invalidate_lock(mapping);
@@ -1007,7 +1069,18 @@ static int kvm_gmem_init_inode(struct inode *inode, loff_t size, u64 flags)
return r;
}
-static int kvm_gmem_attach_resource(struct inode *inode, int resource_fd)
+bool __weak kvm_arch_gmem_device_supported(struct kvm *kvm)
+{
+ return false;
+}
+
+void __weak kvm_arch_gmem_device_bind(struct kvm_memory_slot *slot,
+ const struct guest_memfd_device_binding *binding)
+{
+}
+
+static int kvm_gmem_attach_resource(struct kvm *kvm,
+ struct inode *inode, int resource_fd)
{
struct file *resource_file;
struct inode *res_inode;
@@ -1018,6 +1091,39 @@ static int kvm_gmem_attach_resource(struct inode *inode, int resource_fd)
if (!resource_file)
return -EBADF;
+ if (GMEM_I(inode)->flags & GUEST_MEMFD_FLAG_DEVICE) {
+ struct gmem_inode *gi = GMEM_I(inode);
+ const struct guest_memfd_device_operations *dev_ops;
+
+ dev_ops = kvm_gmem_device_ops(resource_file);
+ if (!dev_ops || !dev_ops->attach) {
+ fput(resource_file);
+ return -EOPNOTSUPP;
+ }
+
+ if (!kvm_arch_gmem_device_supported(kvm)) {
+ fput(resource_file);
+ return -EOPNOTSUPP;
+ }
+
+ if (!(gi->flags & GUEST_MEMFD_FLAG_INIT_SHARED) ||
+ (gi->flags & GUEST_MEMFD_FLAG_MMAP)) {
+ fput(resource_file);
+ return -EOPNOTSUPP;
+ }
+
+ gi->device.kvm = kvm;
+ gi->device.size = i_size_read(inode);
+ provider = dev_ops->attach(resource_file, &gi->device);
+ fput(resource_file);
+ if (IS_ERR(provider))
+ return PTR_ERR(provider);
+ kvm_get_kvm(kvm);
+ gi->provider = provider;
+ gi->device_ops = dev_ops;
+ return 0;
+ }
+
res_inode = file_inode(resource_file);
ops = res_inode->i_sb->s_op->gmem_provider_ops;
@@ -1072,7 +1178,7 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags, int resour
goto err_inode;
if (flags & GUEST_MEMFD_FLAG_USE_RESOURCE) {
- err = kvm_gmem_attach_resource(inode, resource_fd);
+ err = kvm_gmem_attach_resource(kvm, inode, resource_fd);
if (err)
goto err_inode;
}
@@ -1091,6 +1197,7 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags, int resour
xa_init(&f->bindings);
list_add(&f->entry, &GMEM_I(inode)->gmem_file_list);
+ WRITE_ONCE(GMEM_I(inode)->device.core, file);
fd_install(fd, file);
return fd;
@@ -1118,6 +1225,9 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args)
if (!(flags & GUEST_MEMFD_FLAG_USE_RESOURCE) && args->resource_fd)
return -EINVAL;
+ if ((flags & GUEST_MEMFD_FLAG_DEVICE) &&
+ !(flags & GUEST_MEMFD_FLAG_USE_RESOURCE))
+ return -EINVAL;
return __kvm_gmem_create(kvm, size, flags, args->resource_fd);
}
@@ -1125,6 +1235,7 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args)
int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
unsigned int fd, uoff_t offset)
{
+ struct guest_memfd_device_binding binding = {};
uoff_t size = slot->npages << PAGE_SHIFT;
unsigned long start, end;
struct gmem_file *f;
@@ -1140,16 +1251,16 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
return -EBADF;
if (file->f_op != &kvm_gmem_fops)
- goto err;
+ goto out_put;
f = file->private_data;
if (f->kvm != kvm)
- goto err;
+ goto out_put;
inode = file_inode(file);
if (!PAGE_ALIGNED(offset) || offset + size > i_size_read(inode))
- goto err;
+ goto out_put;
filemap_invalidate_lock(inode->i_mapping);
@@ -1159,9 +1270,23 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
if (!xa_empty(&f->bindings) &&
xa_find(&f->bindings, &start, end - 1, XA_PRESENT)) {
r = -EEXIST;
- filemap_invalidate_unlock(inode->i_mapping);
- goto err;
+ goto out_unlock;
+ }
+
+ if (GMEM_I(inode)->device_ops) {
+ if (slot->flags & (KVM_MEM_LOG_DIRTY_PAGES | KVM_MEM_READONLY)) {
+ r = -EINVAL;
+ goto out_unlock;
+ }
+ r = GMEM_I(inode)->device_ops->bind(GMEM_I(inode)->provider, offset,
+ size, &binding);
+ if (r)
+ goto out_unlock;
+ slot->flags |= KVM_MEMSLOT_GMEM_ONLY | KVM_MEMSLOT_GMEM_DEVICE;
}
+ r = xa_err(xa_store_range(&f->bindings, start, end - 1, slot, GFP_KERNEL));
+ if (r)
+ goto out_unlock;
/*
* memslots of flag KVM_MEM_GUEST_MEMFD are immutable to change, so
@@ -1170,19 +1295,20 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
*/
WRITE_ONCE(slot->gmem.file, file);
slot->gmem.pgoff = start;
+ if (slot->flags & KVM_MEMSLOT_GMEM_DEVICE)
+ kvm_arch_gmem_device_bind(slot, &binding);
if (gmem_in_place_conversion || kvm_gmem_supports_mmap(inode))
slot->flags |= KVM_MEMSLOT_GMEM_ONLY;
- xa_store_range(&f->bindings, start, end - 1, slot, GFP_KERNEL);
+ r = 0;
+out_unlock:
filemap_invalidate_unlock(inode->i_mapping);
-
+out_put:
/*
* Drop the reference to the file, even on success. The file pins KVM,
* not the other way 'round. Active bindings are invalidated if the
* file is closed before memslots are destroyed.
*/
- r = 0;
-err:
fput(file);
return r;
}
@@ -1289,6 +1415,21 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot,
filemap_invalidate_lock_shared(file_inode(file)->i_mapping);
+ if (GMEM_I(file_inode(file))->device_ops) {
+ struct gmem_inode *gi = GMEM_I(file_inode(file));
+ unsigned long device_pfn;
+
+ *max_order = 0;
+ if (kvm_gmem_is_private_mem(file_inode(file), index))
+ r = -EAGAIN;
+ else
+ r = gi->device_ops->get_pfn(gi->provider,
+ (u64)index << PAGE_SHIFT, &device_pfn);
+ if (!r)
+ *pfn = device_pfn;
+ goto out;
+ }
+
folio = __kvm_gmem_get_pfn(file, slot, index, pfn, max_order);
if (IS_ERR(folio)) {
r = PTR_ERR(folio);
@@ -1443,6 +1584,8 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb)
gi->flags = 0;
gi->provider = NULL;
gi->provider_ops = NULL;
+ gi->device_ops = NULL;
+ memset(&gi->device, 0, sizeof(gi->device));
INIT_LIST_HEAD(&gi->gmem_file_list);
return &gi->vfs_inode;
}
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH v1 2/6] KVM: guest_memfd: Support private/shared conversion of device memory
2026-10-10 7:24 [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources Aneesh Kumar K.V (Arm)
2026-10-10 7:24 ` [RFC PATCH v1 1/6] KVM: guest_memfd: attach and bind device resources Aneesh Kumar K.V (Arm)
@ 2026-10-10 7:25 ` Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 3/6] iommufd: Support MMIO provider attachment to vDEVICEs Aneesh Kumar K.V (Arm)
` (3 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Aneesh Kumar K.V (Arm) @ 2026-10-10 7:25 UTC (permalink / raw)
To: linux-coco, linux-kernel
Cc: Aneesh Kumar K.V (Arm),
Ackerley Tng, Alex Williamson, David Woodhouse,
David Hildenbrand, Jason Gunthorpe, Joerg Roedel (AMD),
Kevin Tian, Paolo Bonzini, Robin Murphy, Sean Christopherson,
Will Deacon, Alexey Kardashevskiy, Xu Yilun, Catalin Marinas,
Suzuki K Poulose, Steven Price, Fred Griffoul, iommu, kvm
Allow architectures to convert device-backed guest_memfd ranges to
private memory. Coordinate conversion with the device provider and
update guest_memfd attributes only after the architecture operation
succeeds.
Hold the inode invalidate lock across conversion and the attribute
update. Provider callbacks allow a device conversion lock to cover the
same operation, with the inode lock acquired first. This prevents faults
and competing device operations from observing an incomplete conversion.
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: Sean Christopherson <seanjc@google.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: linux-kernel@vger.kernel.org
Cc: kvm@vger.kernel.org
Assisted-by: Codex
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
---
include/linux/guest_memfd.h | 12 +++
include/linux/kvm_host.h | 5 +
virt/kvm/guest_memfd.c | 186 +++++++++++++++++++++++++++++++++++-
3 files changed, 201 insertions(+), 2 deletions(-)
diff --git a/include/linux/guest_memfd.h b/include/linux/guest_memfd.h
index 4f232612a525..b566c581f04d 100644
--- a/include/linux/guest_memfd.h
+++ b/include/linux/guest_memfd.h
@@ -11,6 +11,13 @@ struct mempolicy;
struct kvm;
struct guest_memfd_device;
+/* Claims come from a pending architecture exit, never from a userspace PA. */
+struct guest_memfd_device_request {
+ u64 gpa;
+ u64 pa;
+ u64 vdev_id;
+};
+
/* Binding data stays valid only while the device guest_memfd is active. */
struct guest_memfd_device_binding {
/* Architecture-specific mapping owner, e.g. an RMM vDEVICE handle. */
@@ -29,6 +36,11 @@ struct guest_memfd_device_operations {
int (*bind)(void *data, u64 offset, u64 size,
struct guest_memfd_device_binding *binding);
int (*get_pfn)(void *data, u64 offset, unsigned long *pfn);
+ /* A protected request prepares private mapping; NULL prepares sharing. */
+ int (*prepare_conversion)(struct guest_memfd_device_context *context,
+ const struct guest_memfd_device_request *req);
+ void (*finish_conversion)(struct guest_memfd_device_context *context,
+ const struct guest_memfd_device_request *req);
void (*release)(void *data);
};
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index f1675a3e00b6..32bd56f39f59 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -2614,6 +2614,11 @@ bool kvm_gmem_range_has_attributes(struct kvm_memory_slot *slot, gfn_t start,
bool kvm_arch_gmem_device_supported(struct kvm *kvm);
void kvm_arch_gmem_device_bind(struct kvm_memory_slot *slot,
const struct guest_memfd_device_binding *binding);
+int kvm_arch_gmem_device_map(struct kvm *kvm,
+ const struct kvm_memory_slot *slot,
+ u64 gpa, u64 size, u64 pa);
+int kvm_arch_gmem_device_unmap(struct kvm *kvm, u64 gpa, u64 size);
+int kvm_gmem_device_map(struct kvm *kvm, u64 gpa, u64 size, u64 pa, u64 vdev_id);
int kvm_gmem_set_attributes(struct kvm_memory_slot *slot, gfn_t start,
gfn_t end, u64 attributes);
bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm);
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index c1fd232a83dd..246099caa75f 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -82,6 +82,13 @@ struct gmem_inode {
struct guest_memfd_device device;
const struct guest_memfd_device_operations *device_ops;
+ /*
+ * Whether the completed attribute tree contains any private device range.
+ * Updated under the inode invalidate and provider conversion locks;
+ * provider/TDI readers use READ_ONCE() without taking the inode lock.
+ * This summary avoids reversing the inode -> vDEVICE MMIO lock order.
+ */
+ bool device_has_private;
void *provider;
const struct guest_memfd_provider_operations *provider_ops;
@@ -371,6 +378,10 @@ static void kvm_gmem_invalidate_end(struct inode *inode, pgoff_t start,
__kvm_gmem_invalidate_end(f, start, end);
}
+static int kvm_gmem_device_convert(struct inode *inode, pgoff_t start,
+ pgoff_t nr_pages, const struct guest_memfd_device_request *req,
+ const struct kvm_memory_slot *slot);
+
static long kvm_gmem_punch_hole(struct inode *inode, loff_t offset, loff_t len)
{
enum kvm_gfn_range_filter filter = kvm_gmem_get_all_gfns_filter(inode);
@@ -792,8 +803,12 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
struct ma_state mas;
int r = 0;
- if (gi->device_ops)
- return -EOPNOTSUPP;
+ if (gi->device_ops) {
+ *err_index = start;
+ if (to_private)
+ return -EPERM;
+ return kvm_gmem_device_convert(inode, start, nr_pages, NULL, NULL);
+ }
mt = &gi->attributes;
@@ -1079,6 +1094,172 @@ void __weak kvm_arch_gmem_device_bind(struct kvm_memory_slot *slot,
{
}
+int __weak kvm_arch_gmem_device_map(struct kvm *kvm,
+ const struct kvm_memory_slot *slot,
+ u64 gpa, u64 size, u64 pa)
+{
+ return -EOPNOTSUPP;
+}
+
+int __weak kvm_arch_gmem_device_unmap(struct kvm *kvm, u64 gpa, u64 size)
+{
+ return -EOPNOTSUPP;
+}
+
+static void kvm_gmem_device_update_private(struct gmem_inode *gi)
+{
+ MA_STATE(mas, &gi->attributes, 0, 0);
+ void *entry;
+ bool private = false;
+
+ mas_for_each(&mas, entry, ULONG_MAX) {
+ if (kvm_gmem_get_attributes(&gi->vfs_inode, entry) &
+ KVM_MEMORY_ATTRIBUTE_PRIVATE) {
+ private = true;
+ break;
+ }
+ }
+ WRITE_ONCE(gi->device_has_private, private);
+}
+
+/* The protected request and its slot identify exactly one private range. */
+static int kvm_gmem_device_make_private(struct inode *inode,
+ const struct kvm_memory_slot *slot,
+ const struct guest_memfd_device_request *req,
+ pgoff_t nr_pages)
+{
+ return kvm_arch_gmem_device_map(GMEM_I(inode)->device.kvm, slot,
+ req->gpa, (u64)nr_pages << PAGE_SHIFT,
+ req->pa);
+}
+
+/* The inode lock protects both attributes and the offset-to-GPA bindings. */
+static int kvm_gmem_device_make_shared(struct inode *inode,
+ pgoff_t start, pgoff_t end)
+{
+ struct gmem_inode *gi = GMEM_I(inode);
+ struct gmem_file *f;
+ void *entry;
+
+ MA_STATE(mas, &gi->attributes, start, start);
+
+ mas_for_each(&mas, entry, end - 1) {
+ pgoff_t first = max(start, mas.index);
+ pgoff_t last = min(end, mas.last + 1);
+
+ if (!(kvm_gmem_get_attributes(inode, entry) &
+ KVM_MEMORY_ATTRIBUTE_PRIVATE))
+ continue;
+ while (first < last) {
+ struct kvm_memory_slot *slot = NULL;
+ pgoff_t high;
+ u64 gpa;
+ int ret;
+
+ kvm_gmem_for_each_file(f, inode) {
+ slot = xa_load(&f->bindings, first);
+ if (slot)
+ break;
+ }
+ if (!slot)
+ return -EINVAL;
+ high = min(last, slot->gmem.pgoff + slot->npages);
+ gpa = (slot->base_gfn + first - slot->gmem.pgoff) << PAGE_SHIFT;
+ ret = kvm_arch_gmem_device_unmap(gi->device.kvm, gpa,
+ (high - first) << PAGE_SHIFT);
+ if (ret)
+ return ret;
+ first = high;
+ }
+ }
+ return 0;
+}
+
+static int kvm_gmem_device_convert(struct inode *inode, pgoff_t start,
+ pgoff_t nr_pages, const struct guest_memfd_device_request *req,
+ const struct kvm_memory_slot *slot)
+{
+ struct gmem_inode *gi = GMEM_I(inode);
+ u64 attrs = req ? KVM_MEMORY_ATTRIBUTE_PRIVATE : 0;
+ bool prepared = false;
+ int ret;
+
+ MA_STATE(mas, &gi->attributes, start, start + nr_pages - 1);
+
+ filemap_invalidate_lock(inode->i_mapping);
+ if (__kvm_gmem_range_has_attributes(inode, start, nr_pages, attrs)) {
+ ret = 0;
+ goto out;
+ }
+ if (req && !__kvm_gmem_range_has_attributes(inode, start, nr_pages, 0)) {
+ ret = -EEXIST;
+ goto out;
+ }
+ ret = kvm_gmem_mas_preallocate(&mas, attrs, start, nr_pages);
+ if (ret)
+ goto out;
+ kvm_gmem_invalidate_start(inode, start, start + nr_pages, KVM_FILTER_SHARED);
+ if (req && READ_ONCE(gi->device.kvm->vm_dead))
+ ret = -EIO;
+ else
+ ret = gi->device_ops->prepare_conversion(gi->provider, req);
+ if (!ret) {
+ prepared = true;
+ if (req)
+ ret = kvm_gmem_device_make_private(inode, slot, req, nr_pages);
+ else
+ ret = kvm_gmem_device_make_shared(inode, start, start + nr_pages);
+ }
+ if (!ret) {
+ mas_store_prealloc(&mas, xa_mk_value(attrs));
+ kvm_gmem_device_update_private(gi);
+ } else {
+ mas_destroy(&mas);
+ }
+ kvm_gmem_invalidate_end(inode, start, start + nr_pages);
+ if (prepared)
+ gi->device_ops->finish_conversion(gi->provider, req);
+out:
+ filemap_invalidate_unlock(inode->i_mapping);
+ return ret;
+}
+
+int kvm_gmem_device_map(struct kvm *kvm, u64 gpa, u64 size, u64 pa, u64 vdev_id)
+{
+ struct guest_memfd_device_request req = { gpa, pa, vdev_id };
+ struct kvm_memory_slot *slot;
+ gfn_t gfn = gpa >> PAGE_SHIFT;
+ struct file *file;
+ int ret;
+
+ if (!size)
+ return -EINVAL;
+
+ if (!PAGE_ALIGNED(gpa | size | pa))
+ return -EINVAL;
+
+ /* Caller holds SRCU across the pending architecture request. */
+ slot = gfn_to_memslot(kvm, gfn);
+ if (!slot || (slot->flags & KVM_MEMSLOT_INVALID))
+ return -EINVAL;
+
+ if (!(slot->flags & KVM_MEMSLOT_GMEM_DEVICE))
+ return -EINVAL;
+
+ if ((size >> PAGE_SHIFT) > slot->npages - (gfn - slot->base_gfn))
+ return -EINVAL;
+
+ file = kvm_gmem_get_file(slot);
+ if (!file)
+ return -ENOENT;
+ ret = kvm_gmem_device_convert(file_inode(file),
+ kvm_gmem_get_index(slot, gfn),
+ size >> PAGE_SHIFT, &req, slot);
+ fput(file);
+ return ret;
+}
+EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_device_map);
+
static int kvm_gmem_attach_resource(struct kvm *kvm,
struct inode *inode, int resource_fd)
{
@@ -1585,6 +1766,7 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb)
gi->provider = NULL;
gi->provider_ops = NULL;
gi->device_ops = NULL;
+ gi->device_has_private = false;
memset(&gi->device, 0, sizeof(gi->device));
INIT_LIST_HEAD(&gi->gmem_file_list);
return &gi->vfs_inode;
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH v1 3/6] iommufd: Support MMIO provider attachment to vDEVICEs
2026-10-10 7:24 [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources Aneesh Kumar K.V (Arm)
2026-10-10 7:24 ` [RFC PATCH v1 1/6] KVM: guest_memfd: attach and bind device resources Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 2/6] KVM: guest_memfd: Support private/shared conversion of device memory Aneesh Kumar K.V (Arm)
@ 2026-10-10 7:25 ` Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 4/6] vfio/pci: Provide guest_memfd backing for PCI BARs Aneesh Kumar K.V (Arm)
` (2 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Aneesh Kumar K.V (Arm) @ 2026-10-10 7:25 UTC (permalink / raw)
To: linux-coco, linux-kernel
Cc: Aneesh Kumar K.V (Arm),
Ackerley Tng, Alex Williamson, David Woodhouse,
David Hildenbrand, Jason Gunthorpe, Joerg Roedel (AMD),
Kevin Tian, Paolo Bonzini, Robin Murphy, Sean Christopherson,
Will Deacon, Alexey Kardashevskiy, Xu Yilun, Catalin Marinas,
Suzuki K Poulose, Steven Price, Fred Griffoul, iommu, kvm
Allow a device MMIO provider to attach to the existing vDEVICE after the
backend validates its association with the VM. Keep the vDEVICE and
IOMMUFD context alive until the provider detaches.
Add a per-vDEVICE MMIO mutex to serialize provider attachment and
detachment, and allow providers and backends to use the same lock for
memory conversion and device state changes. Provide callbacks to
invalidate MMIO mappings and query whether private ranges remain.
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Kevin Tian <kevin.tian@intel.com>
Cc: "Joerg Roedel (AMD)" <joro@8bytes.org>
Cc: Will Deacon <will@kernel.org>
Cc: Robin Murphy <robin.murphy@arm.com>
Cc: iommu@lists.linux.dev
Cc: linux-kernel@vger.kernel.org
Assisted-by: Codex
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
---
drivers/iommu/iommufd/viommu.c | 90 ++++++++++++++++++++++++++++++++++
include/linux/iommufd.h | 33 +++++++++++++
2 files changed, 123 insertions(+)
diff --git a/drivers/iommu/iommufd/viommu.c b/drivers/iommu/iommufd/viommu.c
index 9155d0dcb4b4..24a817fb2ae8 100644
--- a/drivers/iommu/iommufd/viommu.c
+++ b/drivers/iommu/iommufd/viommu.c
@@ -250,6 +250,7 @@ int iommufd_vdevice_alloc_ioctl(struct iommufd_ucmd *ucmd)
goto out_unlock_igroup;
}
+ mutex_init(&vdev->mmio_lock);
vdev->virt_id = virt_id;
vdev->viommu = viommu;
refcount_inc(&viommu->obj.users);
@@ -487,3 +488,92 @@ int iommufd_hw_queue_alloc_ioctl(struct iommufd_ucmd *ucmd)
iommufd_put_object(ucmd->ictx, &viommu->obj);
return rc;
}
+
+struct iommufd_vdevice *
+iommufd_device_attach_mmio_provider(struct iommufd_device *idev,
+ struct kvm *kvm, void *data,
+ void (*invalidate_mmio)(void *),
+ bool (*has_private_mmio)(void *),
+ unsigned long *owner)
+{
+ struct iommufd_vdevice *vdev;
+ int ret;
+
+ if (!idev)
+ return ERR_PTR(-ENODEV);
+
+ guard(mutex)(&idev->igroup->lock);
+ if (idev->destroying)
+ return ERR_PTR(-EOPNOTSUPP);
+
+ vdev = idev->vdev;
+ if (!vdev)
+ return ERR_PTR(-EOPNOTSUPP);
+
+ if (!vdev->viommu->ops ||
+ !vdev->viommu->ops->vdevice_get_mmio_owner)
+ return ERR_PTR(-EOPNOTSUPP);
+
+ guard(mutex)(&vdev->mmio_lock);
+ if (vdev->mmio_provider_data)
+ return ERR_PTR(-EBUSY);
+
+ ret = vdev->viommu->ops->vdevice_get_mmio_owner(vdev, kvm, owner);
+ if (ret)
+ return ERR_PTR(ret);
+
+ /* IOMMU_DESTROY must wait for the attached gmem provider to detach. */
+ ret = iommufd_try_inc_users(idev->ictx, &vdev->obj);
+ if (ret)
+ return ERR_PTR(ret);
+
+ iommufd_ctx_get(idev->ictx);
+ vdev->mmio_provider_data = data;
+ vdev->invalidate_mmio = invalidate_mmio;
+ vdev->has_private_mmio = has_private_mmio;
+ invalidate_mmio(data);
+ return vdev;
+}
+EXPORT_SYMBOL_NS_GPL(iommufd_device_attach_mmio_provider, "IOMMUFD");
+
+void iommufd_vdevice_detach_mmio_provider(struct iommufd_vdevice *vdev)
+{
+ struct iommufd_ctx *ictx = vdev->viommu->ictx;
+
+ scoped_guard(mutex, &vdev->mmio_lock) {
+ WARN_ON(iommufd_vdevice_has_private_mmio(vdev));
+ vdev->mmio_provider_data = NULL;
+ vdev->invalidate_mmio = NULL;
+ vdev->has_private_mmio = NULL;
+ }
+ refcount_dec(&vdev->obj.users);
+ iommufd_ctx_put(ictx);
+}
+EXPORT_SYMBOL_NS_GPL(iommufd_vdevice_detach_mmio_provider, "IOMMUFD");
+
+void iommufd_vdevice_mmio_lock(struct iommufd_vdevice *vdev)
+{
+ mutex_lock(&vdev->mmio_lock);
+}
+EXPORT_SYMBOL_NS_GPL(iommufd_vdevice_mmio_lock, "IOMMUFD");
+
+void iommufd_vdevice_mmio_unlock(struct iommufd_vdevice *vdev)
+{
+ mutex_unlock(&vdev->mmio_lock);
+}
+EXPORT_SYMBOL_NS_GPL(iommufd_vdevice_mmio_unlock, "IOMMUFD");
+
+/**
+ * iommufd_vdevice_has_private_mmio() - check for private MMIO ranges
+ * @vdev: vDEVICE whose private MMIO state is being checked
+ *
+ * Return: true if the attached provider has committed private ranges.
+ */
+bool iommufd_vdevice_has_private_mmio(struct iommufd_vdevice *vdev)
+{
+ lockdep_assert_held(&vdev->mmio_lock);
+ if (!vdev->has_private_mmio)
+ return false;
+ return vdev->has_private_mmio(vdev->mmio_provider_data);
+}
+EXPORT_SYMBOL_NS_GPL(iommufd_vdevice_has_private_mmio, "IOMMUFD");
diff --git a/include/linux/iommufd.h b/include/linux/iommufd.h
index 2e67846d6f35..3f3b9d1dd22a 100644
--- a/include/linux/iommufd.h
+++ b/include/linux/iommufd.h
@@ -7,10 +7,12 @@
#define __LINUX_IOMMUFD_H
#include <linux/bits.h>
+#include <linux/cleanup.h>
#include <linux/err.h>
#include <linux/errno.h>
#include <linux/iommu.h>
#include <linux/refcount.h>
+#include <linux/mutex.h>
#include <linux/types.h>
#include <linux/xarray.h>
#include <uapi/linux/iommufd.h>
@@ -26,6 +28,7 @@ struct iommufd_device;
struct iommufd_viommu_ops;
struct tsm_dev;
struct page;
+struct kvm;
struct tsm_guest_req_info;
enum iommufd_object_type {
@@ -73,6 +76,20 @@ int iommufd_device_replace(struct iommufd_device *idev, ioasid_t pasid,
u32 *pt_id);
void iommufd_device_detach(struct iommufd_device *idev, ioasid_t pasid);
+/* The returned opaque vDEVICE pins its IOMMUFD context and physical owner. */
+struct iommufd_vdevice *
+iommufd_device_attach_mmio_provider(struct iommufd_device *idev,
+ struct kvm *kvm, void *data,
+ void (*invalidate_mmio)(void *),
+ bool (*has_private_mmio)(void *),
+ unsigned long *owner);
+void iommufd_vdevice_detach_mmio_provider(struct iommufd_vdevice *vdev);
+void iommufd_vdevice_mmio_lock(struct iommufd_vdevice *vdev);
+void iommufd_vdevice_mmio_unlock(struct iommufd_vdevice *vdev);
+DEFINE_GUARD(iommufd_vdevice_mmio, struct iommufd_vdevice *,
+ iommufd_vdevice_mmio_lock(_T), iommufd_vdevice_mmio_unlock(_T))
+bool iommufd_vdevice_has_private_mmio(struct iommufd_vdevice *vdev);
+
struct iommufd_ctx *iommufd_device_to_ictx(struct iommufd_device *idev);
u32 iommufd_device_to_id(struct iommufd_device *idev);
@@ -130,6 +147,18 @@ struct iommufd_vdevice {
*/
u64 virt_id;
+ /* Serializes provider attachment and conversion. */
+ struct mutex mmio_lock;
+ /* Attached provider and callbacks, protected by mmio_lock until detach. */
+ void *mmio_provider_data;
+ /*
+ * Revoke VFIO userspace BAR mappings and shared guest stage-2 mappings
+ * before delegation. Private guest mappings remain intact; guest_memfd
+ * removes them through its architecture unmap operation.
+ */
+ void (*invalidate_mmio)(void *data);
+ bool (*has_private_mmio)(void *data);
+
/* Guest TSM requests accepted by this vdevice; set by vdevice_init(). */
u64 tsm_req_op_mask;
u32 tsm_tvm_arch;
@@ -187,6 +216,7 @@ struct iommufd_hw_queue {
* include/uapi/linux/iommufd.h)
* If driver has a deinit function to revert what vdevice_init op
* does, it should set it to the @vdev->destroy function pointer
+ * @vdevice_get_mmio_owner: Validate the VM and return its MMIO owner.
* @vdevice_tsm_req: Forward a guest TSM request to a driver-owned vDEVICE
* @get_hw_queue_size: Get the size of a driver-defined HW queue structure for a
* given @viommu corresponding to @queue_type. Driver should
@@ -218,6 +248,9 @@ struct iommufd_viommu_ops {
const struct iommu_user_data *user_data);
int (*cache_invalidate)(struct iommufd_viommu *viommu,
struct iommu_user_data_array *array);
+ /* Owner lookup runs with vdevice->mmio_lock held. */
+ int (*vdevice_get_mmio_owner)(struct iommufd_vdevice *vdev,
+ struct kvm *kvm, unsigned long *owner);
const size_t vdevice_size;
int (*vdevice_init)(struct iommufd_vdevice *vdev);
ssize_t (*vdevice_tsm_req)(struct iommufd_vdevice *vdev,
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH v1 4/6] vfio/pci: Provide guest_memfd backing for PCI BARs
2026-10-10 7:24 [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources Aneesh Kumar K.V (Arm)
` (2 preceding siblings ...)
2026-10-10 7:25 ` [RFC PATCH v1 3/6] iommufd: Support MMIO provider attachment to vDEVICEs Aneesh Kumar K.V (Arm)
@ 2026-10-10 7:25 ` Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 5/6] KVM/VFIO: Remove device mappings during guest_memfd teardown Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 6/6] KVM: Invalidate guest_memfd mappings before removing memslot bindings Aneesh Kumar K.V (Arm)
5 siblings, 0 replies; 7+ messages in thread
From: Aneesh Kumar K.V (Arm) @ 2026-10-10 7:25 UTC (permalink / raw)
To: linux-coco, linux-kernel
Cc: Aneesh Kumar K.V (Arm),
Ackerley Tng, Alex Williamson, David Woodhouse,
David Hildenbrand, Jason Gunthorpe, Joerg Roedel (AMD),
Kevin Tian, Paolo Bonzini, Robin Murphy, Sean Christopherson,
Will Deacon, Alexey Kardashevskiy, Xu Yilun, Catalin Marinas,
Suzuki K Poulose, Steven Price, Fred Griffoul, iommu, kvm
Allow an ordinary VFIO PCI device fd to supply BAR memory to
guest_memfd. Validate eligible BAR ranges and provide their PFNs for
shared guest mappings. Keep the VFIO file and its IOMMUFD vDEVICE alive
while the provider is attached.
Revoke host and shared guest mappings before private conversion, and
check guest_memfd attributes before allowing host BAR access. Exclude
independent dma-buf exports while attached because their mappings bypass
these access checks.
Hold the vDEVICE MMIO mutex across conversion and attribute updates.
Use the VFIO memory lock to serialize BAR access and revocation. Host
access checks try the inode invalidate lock without waiting, avoiding
deadlock with conversion, which acquires that lock first.
Cc: Alex Williamson <alex@shazbot.org>
Cc: Sean Christopherson <seanjc@google.com>
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: linux-kernel@vger.kernel.org
Cc: kvm@vger.kernel.org
Assisted-by: Codex
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
---
drivers/vfio/device_cdev.c | 3 +-
drivers/vfio/group.c | 4 +-
drivers/vfio/pci/Kconfig | 11 +
drivers/vfio/pci/Makefile | 2 +
drivers/vfio/pci/vfio_pci.c | 3 +
drivers/vfio/pci/vfio_pci_core.c | 13 +
drivers/vfio/pci/vfio_pci_dmabuf.c | 10 +
drivers/vfio/pci/vfio_pci_gmem.c | 322 ++++++++++++++++++++++++
drivers/vfio/pci/vfio_pci_gmem.h | 30 +++
drivers/vfio/pci/vfio_pci_gmem_access.c | 134 ++++++++++
drivers/vfio/pci/vfio_pci_rdwr.c | 13 +-
drivers/vfio/vfio.h | 8 +
drivers/vfio/vfio_main.c | 34 ++-
include/linux/guest_memfd.h | 3 +
include/linux/vfio.h | 3 +
include/linux/vfio_pci_core.h | 23 ++
virt/kvm/guest_memfd.c | 105 ++++++++
17 files changed, 702 insertions(+), 19 deletions(-)
create mode 100644 drivers/vfio/pci/vfio_pci_gmem.c
create mode 100644 drivers/vfio/pci/vfio_pci_gmem.h
create mode 100644 drivers/vfio/pci/vfio_pci_gmem_access.c
diff --git a/drivers/vfio/device_cdev.c b/drivers/vfio/device_cdev.c
index 1d9515c967b0..3f251120b1cf 100644
--- a/drivers/vfio/device_cdev.c
+++ b/drivers/vfio/device_cdev.c
@@ -46,7 +46,7 @@ int vfio_device_fops_cdev_open(struct inode *inode, struct file *filep)
goto err_put_registration;
}
- filep->private_data = df;
+ filep->private_data = &df->gmem_ops;
/*
* Use the pseudo fs inode on the device to link all mmaps
@@ -54,7 +54,6 @@ int vfio_device_fops_cdev_open(struct inode *inode, struct file *filep)
* associated to this device using unmap_mapping_range().
*/
filep->f_mapping = device->inode->i_mapping;
-
return 0;
err_put_registration:
diff --git a/drivers/vfio/group.c b/drivers/vfio/group.c
index 5bf8cbdff377..5f45d30ac62a 100644
--- a/drivers/vfio/group.c
+++ b/drivers/vfio/group.c
@@ -271,7 +271,8 @@ static struct file *vfio_device_open_file(struct vfio_device *device)
goto err_free;
filep = anon_inode_getfile_fmode("[vfio-device]", &vfio_device_fops,
- df, O_RDWR, FMODE_PREAD | FMODE_PWRITE);
+ &df->gmem_ops,
+ O_RDWR, FMODE_PREAD | FMODE_PWRITE);
if (IS_ERR(filep)) {
ret = PTR_ERR(filep);
goto err_close_device;
@@ -282,7 +283,6 @@ static struct file *vfio_device_open_file(struct vfio_device *device)
* associated to this device using unmap_mapping_range().
*/
filep->f_mapping = device->inode->i_mapping;
-
if (device->group->type == VFIO_NO_IOMMU)
dev_warn(device->dev, "vfio-noiommu device opened by user "
"(%s:%d)\n", current->comm, task_pid_nr(current));
diff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig
index 296bf01e185e..477c5afad7a2 100644
--- a/drivers/vfio/pci/Kconfig
+++ b/drivers/vfio/pci/Kconfig
@@ -55,6 +55,17 @@ config VFIO_PCI_ZDEV_KVM
To enable s390x KVM vfio-pci extensions, say Y.
+config VFIO_PCI_GMEM
+ bool "VFIO PCI device guest_memfd resource provider"
+ depends on VFIO_PCI && IOMMUFD && ARM64 && KVM
+ depends on KVM_GUEST_MEMFD
+ help
+ Allow the VFIO PCI device fd to provide BAR PFNs to guest_memfd,
+ including confidential-device range conversion through its
+ existing IOMMUFD vDEVICE. This requires a capable TSM backend.
+ Shared access is revoked before confidential range mapping and is
+ restored only after private mappings and device state permit it.
+
config VFIO_PCI_DMABUF
def_bool y if VFIO_PCI_CORE && PCI_P2PDMA && DMA_SHARED_BUFFER
diff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile
index 6138f1bf241d..1abb7448c8b1 100644
--- a/drivers/vfio/pci/Makefile
+++ b/drivers/vfio/pci/Makefile
@@ -3,6 +3,8 @@
vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o
vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o
vfio-pci-core-$(CONFIG_VFIO_PCI_DMABUF) += vfio_pci_dmabuf.o
+vfio-pci-core-$(CONFIG_VFIO_PCI_GMEM) += vfio_pci_gmem_access.o
+vfio-pci-core-$(CONFIG_VFIO_PCI_GMEM) += vfio_pci_gmem.o
obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o
vfio-pci-y := vfio_pci.o
diff --git a/drivers/vfio/pci/vfio_pci.c b/drivers/vfio/pci/vfio_pci.c
index 830369ff878d..675357255a01 100644
--- a/drivers/vfio/pci/vfio_pci.c
+++ b/drivers/vfio/pci/vfio_pci.c
@@ -147,6 +147,9 @@ static int vfio_pci_init_dev(struct vfio_device *core_vdev)
}
static const struct vfio_device_ops vfio_pci_ops = {
+#if IS_ENABLED(CONFIG_VFIO_PCI_GMEM)
+ .gmem_ops = &vfio_pci_gmem_ops,
+#endif
.name = "vfio-pci",
.init = vfio_pci_init_dev,
.release = vfio_pci_core_release_dev,
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 6757054e9d87..1f4bdea4546d 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -1714,6 +1714,7 @@ static void vfio_pci_zap_bars(struct vfio_pci_core_device *vdev)
loff_t len = end - start;
unmap_mapping_range(core_vdev->inode->i_mapping, start, len, true);
+ vfio_pci_gmem_invalidate(vdev);
}
void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)
@@ -1763,6 +1764,18 @@ vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,
if (vdev->pm_runtime_engaged || !__vfio_pci_memory_enabled(vdev))
return VM_FAULT_SIGBUS;
+ /* Use base-page mappings so each BAR page gets its own access check. */
+ if (order && vdev->gmem)
+ return VM_FAULT_FALLBACK;
+ /*
+ * A VFIO mmap can outlive conversion to private memory. Recheck
+ * the guest_memfd and VM attributes for this page so a fault
+ * cannot restore a host mapping of a private BAR range.
+ */
+ if (!vfio_pci_gmem_access_allowed(vdev,
+ (u64)vmf->pgoff << PAGE_SHIFT, PAGE_SIZE))
+ return VM_FAULT_SIGBUS;
+
if (!order)
return vmf_insert_pfn(vmf->vma, vmf->address, pfn);
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c16f460c01d6..b3ae2d97629f 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -306,6 +306,16 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
/* dma_buf_put() now frees priv */
INIT_LIST_HEAD(&priv->dmabufs_elm);
down_write(&vdev->memory_lock);
+ /*
+ * dma-buf consumers bypass guest_memfd range access checks, and their
+ * mappings are not revoked by conversion to private memory. Exclude
+ * exports while attached; memory_lock serializes this with attachment.
+ */
+ if (vdev->gmem) {
+ up_write(&vdev->memory_lock);
+ dma_buf_put(priv->dmabuf);
+ return -EBUSY;
+ }
dma_resv_lock(priv->dmabuf->resv, NULL);
priv->revoked = !__vfio_pci_memory_enabled(vdev);
list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
diff --git a/drivers/vfio/pci/vfio_pci_gmem.c b/drivers/vfio/pci/vfio_pci_gmem.c
new file mode 100644
index 000000000000..d232b5c3e539
--- /dev/null
+++ b/drivers/vfio/pci/vfio_pci_gmem.c
@@ -0,0 +1,322 @@
+// SPDX-License-Identifier: GPL-2.0-only
+#include <linux/guest_memfd.h>
+#include <linux/iommufd.h>
+#include <linux/overflow.h>
+#include <linux/pci.h>
+#include <linux/slab.h>
+#include <linux/vfio_pci_core.h>
+
+#include "../vfio.h"
+#include "vfio_pci_priv.h"
+#include "vfio_pci_gmem.h"
+
+/**
+ * vfio_gmem_resolve_bar_range() - validate and resolve a VFIO BAR interval
+ * @ctx: attached VFIO provider
+ * @offset: byte offset in the VFIO region namespace
+ * @size: interval length in bytes
+ * @bar: output BAR index
+ * @bar_offset: output byte offset within that BAR
+ * @pa: output starting physical address
+ *
+ * Check alignment, mmap eligibility, wrapper and BAR bounds, and the physical
+ * resource snapshot taken at attachment. These checks establish the provider
+ * interval; the VMM approves the protected REC physical-address claim.
+ * Output parameters are valid only on success.
+ *
+ * Return: 0 on success, -ESTALE if the BAR resource changed, or -EINVAL for
+ * an ineligible interval.
+ */
+static int vfio_gmem_resolve_bar_range(struct vfio_pci_gmem *ctx, u64 offset,
+ u64 size, u32 *bar, u64 *bar_offset, phys_addr_t *pa)
+{
+ struct pci_dev *pdev = ctx->vdev->pdev;
+ u32 index = offset >> VFIO_PCI_OFFSET_SHIFT;
+ u64 off = offset & ((1ULL << VFIO_PCI_OFFSET_SHIFT) - 1);
+
+ if (!size || !PAGE_ALIGNED(offset | size))
+ return -EINVAL;
+
+ if (offset >= ctx->context.device->size ||
+ size > ctx->context.device->size - offset)
+ return -EINVAL;
+
+ if (index >= PCI_STD_NUM_BARS)
+ return -EINVAL;
+
+ if (!ctx->vdev->bar_mmap_supported[index] ||
+ !(pci_resource_flags(pdev, index) & IORESOURCE_MEM) ||
+ !PAGE_ALIGNED(ctx->start[index]))
+ return -EINVAL;
+
+ if (off >= ctx->len[index] || size > ctx->len[index] - off)
+ return -EINVAL;
+
+ /* The interval must not cross into the next VFIO region. */
+ if (size > (1ULL << VFIO_PCI_OFFSET_SHIFT) - off)
+ return -EINVAL;
+
+ /* BAR resources must still match the snapshot taken at attachment. */
+ if (ctx->start[index] != pci_resource_start(pdev, index) ||
+ ctx->len[index] != pci_resource_len(pdev, index))
+ return -ESTALE;
+
+ if (check_add_overflow(ctx->start[index], off, pa))
+ return -EINVAL;
+
+ *bar = index;
+ *bar_offset = off;
+ return 0;
+}
+
+/**
+ * vfio_gmem_has_private_mmio() - check for private BAR ranges
+ * @data: attached VFIO provider context
+ *
+ * Return: true if the core has committed private BAR ranges.
+ */
+static bool vfio_gmem_has_private_mmio(void *data)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(data);
+
+ lockdep_assert_held(&ctx->ivdev->mmio_lock);
+ return ctx->context.device->has_private(ctx->context.device);
+}
+
+/**
+ * vfio_gmem_attach() - attach the ordinary VFIO PCI fd to guest_memfd
+ * @file: open VFIO device file supplying the BAR namespace
+ * @gdev: guest_memfd core callbacks and owning VM
+ *
+ * Require an opened IOMMUFD-backed file, exclude another provider and dma-buf
+ * exports, and snapshot the BAR resources. Pin the source file and the existing
+ * vDEVICE after its backend validates the VM. Install reverse invalidation
+ * through the IOMMUFD attachment.
+ *
+ * Return: Provider context on success, or an ERR_PTR() on failure.
+ */
+static struct guest_memfd_device_context *
+vfio_gmem_attach(struct file *file, struct guest_memfd_device *gdev)
+{
+ struct vfio_device_file *df = vfio_device_file_from_file(file);
+ struct vfio_pci_core_device *vdev =
+ container_of(df->device, struct vfio_pci_core_device, vdev);
+ struct vfio_pci_gmem *ctx;
+ struct iommufd_vdevice *ivdev;
+ int bar, ret;
+
+ ctx = kzalloc_obj(*ctx);
+ if (!ctx)
+ return ERR_PTR(-ENOMEM);
+
+ ctx->vdev = vdev;
+ ctx->context.device = gdev;
+
+ mutex_lock(&vdev->vdev.dev_set->lock);
+ /* Paired with smp_store_release() after vfio_df_open(). */
+ if (!smp_load_acquire(&df->access_granted) || !df->iommufd) {
+ ret = -EINVAL;
+ goto fail;
+ }
+
+ down_write(&vdev->memory_lock);
+ /*
+ * Allow only one guest_memfd provider per device. Existing dma-buf
+ * exports bypass guest_memfd access checks and cannot be revoked by
+ * conversion to private memory. Check both under memory_lock to
+ * serialize with provider attachment and dma-buf export.
+ */
+ if (vdev->gmem || !list_empty(&vdev->dmabufs)) {
+ up_write(&vdev->memory_lock);
+ ret = -EBUSY;
+ goto fail;
+ }
+ for (bar = 0; bar < PCI_STD_NUM_BARS; bar++) {
+ ctx->start[bar] = pci_resource_start(vdev->pdev, bar);
+ ctx->len[bar] = pci_resource_len(vdev->pdev, bar);
+ }
+ vdev->gmem = ctx;
+ up_write(&vdev->memory_lock);
+ ivdev = iommufd_device_attach_mmio_provider(vdev->vdev.iommufd_device,
+ gdev->kvm, &ctx->context,
+ vfio_gmem_invalidate_mmio,
+ vfio_gmem_has_private_mmio,
+ &ctx->owner);
+ if (IS_ERR(ivdev)) {
+ ret = PTR_ERR(ivdev);
+ scoped_guard(rwsem_write, &vdev->memory_lock)
+ vdev->gmem = NULL;
+ goto fail;
+ }
+ scoped_guard(rwsem_write, &vdev->memory_lock)
+ ctx->ivdev = ivdev;
+ /* Defer VFIO last-close and IOMMUFD unbind until gmem is released. */
+ ctx->resource = get_file(file);
+ mutex_unlock(&vdev->vdev.dev_set->lock);
+ return &ctx->context;
+fail:
+ mutex_unlock(&vdev->vdev.dev_set->lock);
+ kfree(ctx);
+ return ERR_PTR(ret);
+}
+
+/**
+ * vfio_gmem_bind() - bind an eligible BAR interval to a memslot
+ * @data: provider context returned by vfio_gmem_attach()
+ * @offset: byte offset in the VFIO region namespace
+ * @size: length in bytes
+ * @binding: receives the pinned mapping owner and virtual device ID
+ *
+ * Validate the resource snapshot and reserve the BAR through the ordinary
+ * VFIO iomap machinery.
+ *
+ * Called with the inode invalidate lock held;
+ * Return: 0 on success or a negative error code.
+ */
+static int vfio_gmem_bind(void *data, u64 offset, u64 size,
+ struct guest_memfd_device_binding *binding)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(data);
+ phys_addr_t pa;
+ u64 off;
+ u32 bar;
+ int ret;
+
+ guard(iommufd_vdevice_mmio)(ctx->ivdev);
+ ret = vfio_gmem_resolve_bar_range(ctx, offset, size, &bar, &off, &pa);
+ if (ret)
+ return ret;
+
+ /* Reserve the physical resource with the ordinary VFIO BAR machinery. */
+ ret = PTR_ERR_OR_ZERO(vfio_pci_core_get_iomap(ctx->vdev, bar));
+ if (!ret) {
+ binding->owner = ctx->owner;
+ binding->vdev_id = ctx->ivdev->virt_id;
+ }
+ return ret;
+}
+
+/**
+ * vfio_gmem_get_pfn() - obtain a shared BAR PFN for a guest fault
+ * @data: attached provider context
+ * @offset: page-aligned byte offset in the VFIO region namespace
+ * @pfn: receives the bare BAR PFN
+ *
+ * Check shared-access admission, power state and PCI memory decoding under
+ * the VFIO memory lock before resolving the resource. The PFN has no RAM
+ * page reference to release. The core checks completed attributes and the
+ * fault path checks the invalidation sequence before installing a mapping.
+ *
+ * Guest_memfd calls this only for shared device ranges. Private mappings use
+ * the physical address from an approved protected device request through
+ * kvm_arch_gmem_device_map(). Installing private mappings through ordinary
+ * PFN faults would bypass that approval path.
+ *
+ * Called with the inode invalidate lock held
+ * Return: 0 on success, -EAGAIN while shared access is unavailable, -ENODEV
+ * after detachment, or a resource-validation error.
+ */
+static int vfio_gmem_get_pfn(void *data, u64 offset, unsigned long *pfn)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(data);
+ struct vfio_pci_core_device *vdev = ctx->vdev;
+ phys_addr_t pa;
+ u64 off;
+ u32 bar;
+ int ret;
+
+ guard(rwsem_read)(&vdev->memory_lock);
+ if (vdev->pm_runtime_engaged || !__vfio_pci_memory_enabled(vdev))
+ return -EAGAIN;
+
+ ret = vfio_gmem_resolve_bar_range(ctx, offset, PAGE_SIZE,
+ &bar, &off, &pa);
+ if (!ret)
+ *pfn = PHYS_PFN(pa);
+ return ret;
+}
+
+/**
+ * vfio_gmem_prepare_conversion() - serialize device range conversion
+ * @context: attached provider context
+ * @req: protected private request, or NULL for shared cleanup
+ *
+ * Check the virtual device ID and revoke host BAR access and shared stage-2
+ * mappings before private mapping. Shared cleanup only takes the MMIO mutex.
+ * Keep that mutex held across architecture cleanup and attribute publication.
+ *
+ * Called with the inode invalidate lock held. On success,
+ * vfio_gmem_finish_conversion() must release the retained MMIO mutex.
+ * Return: 0 with the MMIO mutex held, or a negative error with it released.
+ */
+static int vfio_gmem_prepare_conversion(struct guest_memfd_device_context *context,
+ const struct guest_memfd_device_request *req)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(context);
+
+ iommufd_vdevice_mmio_lock(ctx->ivdev);
+ if (req && req->vdev_id != ctx->ivdev->virt_id) {
+ iommufd_vdevice_mmio_unlock(ctx->ivdev);
+ return -EPERM;
+ }
+ if (req)
+ vfio_gmem_invalidate_mmio(context);
+ return 0; /* Keep mmio_lock held until finish_conversion(). */
+}
+
+/**
+ * vfio_gmem_finish_conversion() - release the conversion mutex
+ * @context: attached provider context
+ * @req: protected private request, or NULL for shared cleanup
+ *
+ * Release the MMIO mutex after the core publishes completed attributes.
+ * Subsequent faults and host I/O check their own range; no device-wide private
+ * or shared admission state is published here.
+ *
+ * Called with the inode invalidate lock and the vDEVICE MMIO mutex held
+ * after successful preparation.
+ */
+static void vfio_gmem_finish_conversion(struct guest_memfd_device_context *context,
+ const struct guest_memfd_device_request *req)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(context);
+
+ iommufd_vdevice_mmio_unlock(ctx->ivdev);
+}
+
+/**
+ * vfio_gmem_release() - release a provider after successful core cleanup
+ * @data: provider context to release
+ *
+ * Detach from the pinned vDEVICE, clear the VFIO provider association and
+ * drop the source file reference before freeing the context.
+ *
+ * Takes the VFIO device-set lock and the vDEVICE MMIO and VFIO memory
+ * locks while detaching.
+ */
+static void vfio_gmem_release(void *data)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(data);
+ struct vfio_pci_core_device *vdev = ctx->vdev;
+
+ mutex_lock(&vdev->vdev.dev_set->lock);
+ iommufd_vdevice_detach_mmio_provider(ctx->ivdev);
+ down_write(&vdev->memory_lock);
+ vdev->gmem = NULL;
+ up_write(&vdev->memory_lock);
+ mutex_unlock(&vdev->vdev.dev_set->lock);
+ fput(ctx->resource);
+ kfree(ctx);
+}
+
+const struct guest_memfd_device_operations vfio_pci_gmem_ops = {
+ .attach = vfio_gmem_attach,
+ .bind = vfio_gmem_bind,
+ .get_pfn = vfio_gmem_get_pfn,
+ .prepare_conversion = vfio_gmem_prepare_conversion,
+ .finish_conversion = vfio_gmem_finish_conversion,
+ .release = vfio_gmem_release,
+};
+EXPORT_SYMBOL_GPL(vfio_pci_gmem_ops);
+
+MODULE_IMPORT_NS("IOMMUFD");
diff --git a/drivers/vfio/pci/vfio_pci_gmem.h b/drivers/vfio/pci/vfio_pci_gmem.h
new file mode 100644
index 000000000000..e5ad19ba6748
--- /dev/null
+++ b/drivers/vfio/pci/vfio_pci_gmem.h
@@ -0,0 +1,30 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+#ifndef VFIO_PCI_GMEM_H
+#define VFIO_PCI_GMEM_H
+
+#include <linux/container_of.h>
+#include <linux/guest_memfd.h>
+#include <linux/pci.h>
+
+struct iommufd_vdevice;
+struct vfio_pci_core_device;
+
+struct vfio_pci_gmem {
+ struct guest_memfd_device_context context;
+ struct vfio_pci_core_device *vdev;
+ struct iommufd_vdevice *ivdev;
+ struct file *resource;
+ unsigned long owner;
+ resource_size_t start[PCI_STD_NUM_BARS];
+ resource_size_t len[PCI_STD_NUM_BARS];
+};
+
+static inline struct vfio_pci_gmem *
+vfio_pci_gmem_from_context(struct guest_memfd_device_context *context)
+{
+ return container_of_const(context, struct vfio_pci_gmem, context);
+}
+
+void vfio_gmem_invalidate_mmio(void *data);
+
+#endif /* VFIO_PCI_GMEM_H */
diff --git a/drivers/vfio/pci/vfio_pci_gmem_access.c b/drivers/vfio/pci/vfio_pci_gmem_access.c
new file mode 100644
index 000000000000..a3975b536e22
--- /dev/null
+++ b/drivers/vfio/pci/vfio_pci_gmem_access.c
@@ -0,0 +1,134 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Revoke translations before delegation; shared/private state is per range. */
+#include <linux/guest_memfd.h>
+#include <linux/iommufd.h>
+#include <linux/vfio_pci_core.h>
+
+#include "vfio_pci_gmem.h"
+#include "vfio_pci_priv.h"
+
+/**
+ * vfio_pci_gmem_invalidate() - notify KVM of revoked BAR mappings
+ * @vdev: VFIO PCI device whose BAR translations are being revoked
+ *
+ * Called by vfio_pci_zap_bars() after it removes host BAR mappings. Notify
+ * guest_memfd synchronously so KVM removes the corresponding shared guest
+ * mappings. This helper does not remove host mappings or change attributes.
+ * Skip notification if no provider is attached.
+ *
+ * The guest_memfd callback takes SRCU and the KVM MMU lock without taking
+ * the inode invalidate lock, which may already be held during conversion.
+ *
+ */
+void vfio_pci_gmem_invalidate(struct vfio_pci_core_device *vdev)
+{
+ struct vfio_pci_gmem *ctx = vdev->gmem;
+
+ lockdep_assert_held_write(&vdev->memory_lock);
+ if (ctx)
+ ctx->context.device->invalidate(ctx->context.device);
+}
+
+/**
+ * vfio_gmem_invalidate_mmio() - revoke host and shared guest BAR translations
+ * @data: attached provider context
+ *
+ * Initiate BAR revocation for conversion or access refresh. Call
+ * vfio_pci_zap_and_down_write_memory_lock() to remove host BAR mappings;
+ * its BAR zap path also calls vfio_pci_gmem_invalidate() to remove the
+ * corresponding shared guest mappings. Release the VFIO memory lock before
+ * returning.
+ *
+ * During private conversion, the caller holds the inode invalidate lock to
+ * prevent new guest_memfd faults. Host faults and I/O try that lock when
+ * checking range access, so revocation completes before delegation.
+ *
+ * Caller holds the vDEVICE MMIO mutex; takes the VFIO memory lock for
+ * write and releases it before architecture mapping.
+ */
+void vfio_gmem_invalidate_mmio(void *data)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(data);
+ struct vfio_pci_core_device *vdev = ctx->vdev;
+
+ /* IOMMUFD invokes this during attachment before ctx->ivdev is assigned. */
+ if (ctx->ivdev)
+ lockdep_assert_held(&ctx->ivdev->mmio_lock);
+
+ vfio_pci_zap_and_down_write_memory_lock(vdev);
+ up_write(&vdev->memory_lock);
+}
+
+/**
+ * vfio_pci_gmem_access_allowed() - check a host BAR access against its range
+ * @vdev: VFIO PCI device supplying the BAR
+ * @offset: byte offset in the VFIO region namespace
+ * @size: access length in bytes
+ *
+ * Consult completed guest_memfd attributes for the requested interval.
+ * Private mappings in another interval do not block this access. Try the
+ * inode lock instead of waiting in the reverse order from conversion.
+ *
+ * Return: true when the requested interval is accessible.
+ */
+bool vfio_pci_gmem_access_allowed(struct vfio_pci_core_device *vdev,
+ u64 offset, u64 size)
+{
+ struct vfio_pci_gmem *ctx = vdev->gmem;
+
+ lockdep_assert_held(&vdev->memory_lock);
+ if (!ctx)
+ return true;
+
+ /* Attachment publishes the context before resolving its vDEVICE. */
+ if (!ctx->ivdev)
+ return false;
+
+ return ctx->context.device->host_accessible(ctx->context.device,
+ offset, size);
+}
+
+/**
+ * vfio_pci_gmem_io_access_allowed() - resolve a mapped BAR I/O pointer
+ * @vdev: VFIO PCI device supplying the iomap
+ * @io: address used by ordinary VFIO read, write or ioeventfd
+ * @size: access width in bytes
+ *
+ * Standard BAR iomaps are cached in barmap. Resolve those to file offsets
+ * before applying range admission. Generic VFIO's remaining memory I/O uses
+ * a temporary ROM iomap, which cannot be bound to a device guest_memfd.
+ * Callers supply VFIO's own validated iomap pointers.
+ *
+ * Return: true when the I/O interval is accessible.
+ */
+bool vfio_pci_gmem_io_access_allowed(struct vfio_pci_core_device *vdev,
+ void __iomem *io, size_t size)
+{
+ unsigned long addr = (unsigned long)io;
+ int bar;
+
+ lockdep_assert_held(&vdev->memory_lock);
+ if (!vdev->gmem)
+ return vfio_pci_gmem_access_allowed(vdev, 0, size);
+
+ for (bar = 0; bar < PCI_STD_NUM_BARS; bar++) {
+ unsigned long base = (unsigned long)vdev->barmap[bar];
+ u64 offset;
+
+ if (!base || addr < base)
+ continue;
+
+ offset = addr - base;
+ if (offset >= pci_resource_len(vdev->pdev, bar))
+ continue;
+
+ if (size > pci_resource_len(vdev->pdev, bar) - offset)
+ return false;
+
+ return vfio_pci_gmem_access_allowed(vdev,
+ VFIO_PCI_INDEX_TO_OFFSET(bar) + offset, size);
+ }
+ /* The ROM has no private provider mappings, but lifecycle checks apply. */
+ return vfio_pci_gmem_access_allowed(vdev,
+ VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX), size);
+}
diff --git a/drivers/vfio/pci/vfio_pci_rdwr.c b/drivers/vfio/pci/vfio_pci_rdwr.c
index 7f14dd46de17..54c1b3d2f725 100644
--- a/drivers/vfio/pci/vfio_pci_rdwr.c
+++ b/drivers/vfio/pci/vfio_pci_rdwr.c
@@ -44,7 +44,9 @@ int vfio_pci_core_iowrite##size(struct vfio_pci_core_device *vdev, \
{ \
if (test_mem) { \
down_read(&vdev->memory_lock); \
- if (!__vfio_pci_memory_enabled(vdev)) { \
+ if (!__vfio_pci_memory_enabled(vdev) || \
+ !vfio_pci_gmem_io_access_allowed(vdev, io, \
+ sizeof(u##size))) { \
up_read(&vdev->memory_lock); \
return -EIO; \
} \
@@ -70,7 +72,9 @@ int vfio_pci_core_ioread##size(struct vfio_pci_core_device *vdev, \
{ \
if (test_mem) { \
down_read(&vdev->memory_lock); \
- if (!__vfio_pci_memory_enabled(vdev)) { \
+ if (!__vfio_pci_memory_enabled(vdev) || \
+ !vfio_pci_gmem_io_access_allowed(vdev, io, \
+ sizeof(u##size))) { \
up_read(&vdev->memory_lock); \
return -EIO; \
} \
@@ -380,7 +384,10 @@ static int vfio_pci_ioeventfd_handler(void *opaque, void *unused)
if (ioeventfd->test_mem) {
if (!down_read_trylock(&vdev->memory_lock))
return 1; /* Lock contended, use thread */
- if (!__vfio_pci_memory_enabled(vdev)) {
+ if (!__vfio_pci_memory_enabled(vdev) ||
+ !vfio_pci_gmem_access_allowed(vdev,
+ VFIO_PCI_INDEX_TO_OFFSET(ioeventfd->bar) +
+ ioeventfd->pos, ioeventfd->count)) {
up_read(&vdev->memory_lock);
return 0;
}
diff --git a/drivers/vfio/vfio.h b/drivers/vfio/vfio.h
index 9b619951a5d0..af0940a063a9 100644
--- a/drivers/vfio/vfio.h
+++ b/drivers/vfio/vfio.h
@@ -7,6 +7,7 @@
#define __VFIO_VFIO_H__
#include <linux/file.h>
+#include <linux/guest_memfd.h>
#include <linux/device.h>
#include <linux/cdev.h>
#include <linux/module.h>
@@ -17,6 +18,8 @@ struct iommu_group;
struct vfio_container;
struct vfio_device_file {
+ /* Exposed through file->private_data for VFIO device files. */
+ struct guest_memfd_device_operations gmem_ops;
struct vfio_device *device;
struct vfio_group *group;
@@ -27,6 +30,11 @@ struct vfio_device_file {
struct iommufd_ctx *iommufd; /* protected by struct vfio_device_set::lock */
};
+static inline struct vfio_device_file *vfio_device_file_from_file(struct file *file)
+{
+ return container_of(file->private_data, struct vfio_device_file, gmem_ops);
+}
+
void vfio_device_put_registration(struct vfio_device *device);
bool vfio_device_try_get_registration(struct vfio_device *device);
int vfio_df_open(struct vfio_device_file *df);
diff --git a/drivers/vfio/vfio_main.c b/drivers/vfio/vfio_main.c
index ed96acfa8635..c3df4f9b1ba4 100644
--- a/drivers/vfio/vfio_main.c
+++ b/drivers/vfio/vfio_main.c
@@ -509,6 +509,8 @@ vfio_allocate_device_file(struct vfio_device *device)
if (!df)
return ERR_PTR(-ENOMEM);
+ if (device->ops->gmem_ops)
+ df->gmem_ops = *device->ops->gmem_ops;
df->device = device;
spin_lock_init(&df->kvm_ref_lock);
@@ -642,7 +644,7 @@ static inline void vfio_device_pm_runtime_put(struct vfio_device *device)
*/
static int vfio_device_fops_release(struct inode *inode, struct file *filep)
{
- struct vfio_device_file *df = filep->private_data;
+ struct vfio_device_file *df = vfio_device_file_from_file(filep);
struct vfio_device *device = df->device;
if (df->group)
@@ -1340,7 +1342,7 @@ static long vfio_get_region_info(struct vfio_device *device,
static long vfio_device_fops_unl_ioctl(struct file *filep,
unsigned int cmd, unsigned long arg)
{
- struct vfio_device_file *df = filep->private_data;
+ struct vfio_device_file *df = vfio_device_file_from_file(filep);
struct vfio_device *device = df->device;
void __user *uptr = (void __user *)arg;
int ret;
@@ -1393,7 +1395,7 @@ static long vfio_device_fops_unl_ioctl(struct file *filep,
static ssize_t vfio_device_fops_read(struct file *filep, char __user *buf,
size_t count, loff_t *ppos)
{
- struct vfio_device_file *df = filep->private_data;
+ struct vfio_device_file *df = vfio_device_file_from_file(filep);
struct vfio_device *device = df->device;
/* Paired with smp_store_release() following vfio_df_open() */
@@ -1410,7 +1412,7 @@ static ssize_t vfio_device_fops_write(struct file *filep,
const char __user *buf,
size_t count, loff_t *ppos)
{
- struct vfio_device_file *df = filep->private_data;
+ struct vfio_device_file *df = vfio_device_file_from_file(filep);
struct vfio_device *device = df->device;
/* Paired with smp_store_release() following vfio_df_open() */
@@ -1425,7 +1427,7 @@ static ssize_t vfio_device_fops_write(struct file *filep,
static int vfio_device_fops_mmap(struct file *filep, struct vm_area_struct *vma)
{
- struct vfio_device_file *df = filep->private_data;
+ struct vfio_device_file *df = vfio_device_file_from_file(filep);
struct vfio_device *device = df->device;
/* Paired with smp_store_release() following vfio_df_open() */
@@ -1442,7 +1444,7 @@ static int vfio_device_fops_mmap(struct file *filep, struct vm_area_struct *vma)
static void vfio_device_show_fdinfo(struct seq_file *m, struct file *filep)
{
char *path;
- struct vfio_device_file *df = filep->private_data;
+ struct vfio_device_file *df = vfio_device_file_from_file(filep);
struct vfio_device *device = df->device;
path = kobject_get_path(&device->dev->kobj, GFP_KERNEL);
@@ -1470,11 +1472,9 @@ const struct file_operations vfio_device_fops = {
static struct vfio_device *vfio_device_from_file(struct file *file)
{
- struct vfio_device_file *df = file->private_data;
-
if (file->f_op != &vfio_device_fops)
return NULL;
- return df->device;
+ return vfio_device_file_from_file(file)->device;
}
/**
@@ -1517,7 +1517,7 @@ EXPORT_SYMBOL_GPL(vfio_file_enforced_coherent);
static void vfio_device_file_set_kvm(struct file *file, struct file *kvm)
{
- struct vfio_device_file *df = file->private_data;
+ struct vfio_device_file *df = vfio_device_file_from_file(file);
struct file *old;
if (kvm)
@@ -1814,9 +1814,15 @@ static int __init vfio_init(void)
ida_init(&vfio.device_ida);
+ if (IS_ENABLED(CONFIG_KVM_GUEST_MEMFD)) {
+ ret = guest_memfd_register_device_fops(&vfio_device_fops);
+ if (ret)
+ return ret;
+ }
+
ret = vfio_group_init();
if (ret)
- return ret;
+ goto err_group;
ret = vfio_virqfd_init();
if (ret)
@@ -1834,13 +1840,15 @@ static int __init vfio_init(void)
vfio_debugfs_create_root();
pr_info(DRIVER_DESC " version: " DRIVER_VERSION "\n");
return 0;
-
err_alloc_dev_chrdev:
class_unregister(&vfio_device_class);
err_dev_class:
vfio_virqfd_exit();
err_virqfd:
vfio_group_cleanup();
+err_group:
+ if (IS_ENABLED(CONFIG_KVM_GUEST_MEMFD))
+ guest_memfd_unregister_device_fops(&vfio_device_fops);
return ret;
}
@@ -1853,6 +1861,8 @@ static void __exit vfio_cleanup(void)
vfio_virqfd_exit();
vfio_group_cleanup();
xa_destroy(&vfio_device_set_xa);
+ if (IS_ENABLED(CONFIG_KVM_GUEST_MEMFD))
+ guest_memfd_unregister_device_fops(&vfio_device_fops);
}
module_init(vfio_init);
diff --git a/include/linux/guest_memfd.h b/include/linux/guest_memfd.h
index b566c581f04d..24f4b3c79afa 100644
--- a/include/linux/guest_memfd.h
+++ b/include/linux/guest_memfd.h
@@ -52,6 +52,9 @@ struct guest_memfd_device {
struct kvm *kvm;
u64 size;
void *core;
+ void (*invalidate)(struct guest_memfd_device *gdev);
+ bool (*has_private)(struct guest_memfd_device *gdev);
+ bool (*host_accessible)(struct guest_memfd_device *gdev, u64 offset, u64 size);
};
/**
diff --git a/include/linux/vfio.h b/include/linux/vfio.h
index 0cc91c6f96d2..8f71117a2851 100644
--- a/include/linux/vfio.h
+++ b/include/linux/vfio.h
@@ -113,7 +113,10 @@ struct vfio_device {
* this device is attached to.
* @device_feature: Optional, fill in the VFIO_DEVICE_FEATURE ioctl
*/
+struct guest_memfd_device_operations;
+
struct vfio_device_ops {
+ const struct guest_memfd_device_operations *gmem_ops;
char *name;
int (*init)(struct vfio_device *vdev);
void (*release)(struct vfio_device *vdev);
diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index 9a1674c152aa..73294230d54e 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -149,8 +149,31 @@ struct vfio_pci_core_device {
struct notifier_block nb;
struct rw_semaphore memory_lock;
struct list_head dmabufs;
+ /* Attached provider; publication/removal require memory_lock for write. */
+ struct vfio_pci_gmem *gmem;
};
+#if IS_ENABLED(CONFIG_VFIO_PCI_GMEM)
+extern const struct guest_memfd_device_operations vfio_pci_gmem_ops;
+void vfio_pci_gmem_invalidate(struct vfio_pci_core_device *vdev);
+bool vfio_pci_gmem_access_allowed(struct vfio_pci_core_device *vdev,
+ u64 offset, u64 size);
+bool vfio_pci_gmem_io_access_allowed(struct vfio_pci_core_device *vdev,
+ void __iomem *io, size_t size);
+#else
+static inline void vfio_pci_gmem_invalidate(struct vfio_pci_core_device *vdev) {}
+static inline bool vfio_pci_gmem_access_allowed(struct vfio_pci_core_device *vdev,
+ u64 offset, u64 size)
+{
+ return true;
+}
+static inline bool vfio_pci_gmem_io_access_allowed(struct vfio_pci_core_device *vdev,
+ void __iomem *io, size_t size)
+{
+ return true;
+}
+#endif
+
enum vfio_pci_io_width {
VFIO_PCI_IO_WIDTH_1 = 1,
VFIO_PCI_IO_WIDTH_2 = 2,
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 246099caa75f..fc37066bfaa5 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -1106,6 +1106,108 @@ int __weak kvm_arch_gmem_device_unmap(struct kvm *kvm, u64 gpa, u64 size)
return -EOPNOTSUPP;
}
+static bool gmem_device_host_accessible(struct guest_memfd_device *gdev,
+ u64 offset, u64 size)
+{
+ struct file *file = READ_ONCE(gdev->core);
+ struct inode *inode;
+ pgoff_t first, last;
+ bool shared;
+
+ if (!file)
+ return false;
+ /* Unwrapped BAR offsets have no private guest_memfd mappings. */
+ if (offset >= gdev->size)
+ return true;
+ size = min(size, gdev->size - offset);
+ if (!size)
+ return false;
+ first = offset >> PAGE_SHIFT;
+ last = ((offset + size - 1) >> PAGE_SHIFT) + 1;
+ inode = file_inode(file);
+ /*
+ * Keep guest_memfd attributes stable against conversion while checking
+ * host access. This check only reads attributes, so a shared lock lets
+ * other access checks run concurrently.
+ *
+ * VFIO callers hold memory_lock, but conversion takes the invalidate
+ * lock exclusively before taking memory_lock to revoke BAR mappings.
+ * Waiting here could deadlock with conversion. Try the shared lock and
+ * deny access if it cannot be acquired.
+ */
+ if (!filemap_invalidate_trylock_shared(inode->i_mapping))
+ return false;
+ shared = __kvm_gmem_range_has_attributes(inode, first, last - first, 0);
+ filemap_invalidate_unlock_shared(inode->i_mapping);
+ return shared;
+}
+
+/**
+ * gmem_device_invalidate() - revoke shared guest mappings of device memory
+ * @gdev: device guest_memfd whose provider is revoking access
+ *
+ * VFIO invokes this callback when zapping BAR mappings. Removing host BAR
+ * mappings alone leaves guest stage-2 mappings intact, so unmap the shared
+ * ranges in memslots backed by this guest_memfd and flush guest TLBs as needed.
+ * Advance the MMU invalidation sequence so faults that obtained a provider PFN
+ * before revocation retry instead of installing a stale mapping.
+ *
+ * SRCU protects memslot traversal, and the KVM MMU lock serializes unmapping
+ * with guest faults. Do not acquire the inode invalidate lock: conversion may
+ * already hold it, and VFIO callers hold memory_lock in the reverse order.
+ *
+ * Context: Caller holds the provider memory lock to exclude new provider PFN
+ * lookups during invalidation. Takes SRCU and the KVM MMU lock.
+ */
+static void gmem_device_invalidate(struct guest_memfd_device *gdev)
+{
+ struct file *file = gdev->core;
+ struct kvm *kvm = gdev->kvm;
+ struct kvm_memory_slot *slot;
+ struct kvm_memslots *slots;
+ int srcu_idx, bkt, as;
+ bool flush = false, found_memslot = false;
+
+ if (!file)
+ return;
+ srcu_idx = srcu_read_lock(&kvm->srcu);
+ KVM_MMU_LOCK(kvm);
+ for (as = 0; as < kvm_arch_nr_memslot_as_ids(kvm); as++) {
+ slots = __kvm_memslots(kvm, as);
+ kvm_for_each_memslot(slot, bkt, slots) {
+ struct kvm_gfn_range range = {
+ .slot = slot,
+ .start = slot->base_gfn,
+ .end = slot->base_gfn + slot->npages,
+ .attr_filter = KVM_FILTER_SHARED,
+ .may_block = false,
+ };
+
+ if (READ_ONCE(slot->gmem.file) != file)
+ continue;
+ if (!found_memslot) {
+ found_memslot = true;
+ kvm_mmu_invalidate_start(kvm);
+ }
+ flush |= kvm_mmu_unmap_gfn_range(kvm, &range);
+ }
+ }
+ if (flush)
+ kvm_flush_remote_tlbs(kvm);
+ if (found_memslot)
+ kvm_mmu_invalidate_end(kvm);
+ KVM_MMU_UNLOCK(kvm);
+ srcu_read_unlock(&kvm->srcu, srcu_idx);
+}
+
+/* The attribute tree owns mapping state; this summary is read under provider locks. */
+static bool gmem_device_has_private(struct guest_memfd_device *gdev)
+{
+ struct gmem_inode *gi = container_of(gdev, struct gmem_inode, device);
+
+ return READ_ONCE(gi->device_has_private);
+}
+
static void kvm_gmem_device_update_private(struct gmem_inode *gi)
{
MA_STATE(mas, &gi->attributes, 0, 0);
@@ -1295,6 +1397,9 @@ static int kvm_gmem_attach_resource(struct kvm *kvm,
gi->device.kvm = kvm;
gi->device.size = i_size_read(inode);
+ gi->device.invalidate = gmem_device_invalidate;
+ gi->device.host_accessible = gmem_device_host_accessible;
+ gi->device.has_private = gmem_device_has_private;
provider = dev_ops->attach(resource_file, &gi->device);
fput(resource_file);
if (IS_ERR(provider))
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH v1 5/6] KVM/VFIO: Remove device mappings during guest_memfd teardown
2026-10-10 7:24 [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources Aneesh Kumar K.V (Arm)
` (3 preceding siblings ...)
2026-10-10 7:25 ` [RFC PATCH v1 4/6] vfio/pci: Provide guest_memfd backing for PCI BARs Aneesh Kumar K.V (Arm)
@ 2026-10-10 7:25 ` Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 6/6] KVM: Invalidate guest_memfd mappings before removing memslot bindings Aneesh Kumar K.V (Arm)
5 siblings, 0 replies; 7+ messages in thread
From: Aneesh Kumar K.V (Arm) @ 2026-10-10 7:25 UTC (permalink / raw)
To: linux-coco, linux-kernel
Cc: Aneesh Kumar K.V (Arm),
Ackerley Tng, Alex Williamson, David Woodhouse,
David Hildenbrand, Jason Gunthorpe, Joerg Roedel (AMD),
Kevin Tian, Paolo Bonzini, Robin Murphy, Sean Christopherson,
Will Deacon, Alexey Kardashevskiy, Xu Yilun, Catalin Marinas,
Suzuki K Poulose, Steven Price, Fred Griffoul, iommu, kvm
On final guest_memfd close, stop new provider access and callbacks
before removing private device mappings and releasing the provider. Keep
slot bindings available until the mappings have been removed.
Serialize teardown with conversion using the inode invalidate lock. VFIO
takes the vDEVICE MMIO mutex and its memory lock when disabling access
and callbacks, preventing concurrent faults or device operations from
using provider state being torn down.
Cc: Alex Williamson <alex@shazbot.org>
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: Sean Christopherson <seanjc@google.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: kvm@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Assisted-by: Codex
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
---
drivers/vfio/pci/vfio_pci_gmem.c | 27 +++++++++++++++++++++++++
drivers/vfio/pci/vfio_pci_gmem.h | 6 ++++++
drivers/vfio/pci/vfio_pci_gmem_access.c | 12 +++++++----
include/linux/guest_memfd.h | 1 +
virt/kvm/guest_memfd.c | 21 +++++++++++++++++++
5 files changed, 63 insertions(+), 4 deletions(-)
diff --git a/drivers/vfio/pci/vfio_pci_gmem.c b/drivers/vfio/pci/vfio_pci_gmem.c
index d232b5c3e539..3f8630a0eeca 100644
--- a/drivers/vfio/pci/vfio_pci_gmem.c
+++ b/drivers/vfio/pci/vfio_pci_gmem.c
@@ -183,6 +183,8 @@ static int vfio_gmem_bind(void *data, u64 offset, u64 size,
int ret;
guard(iommufd_vdevice_mmio)(ctx->ivdev);
+ if (ctx->callbacks_detached)
+ return -ENODEV;
ret = vfio_gmem_resolve_bar_range(ctx, offset, size, &bar, &off, &pa);
if (ret)
return ret;
@@ -226,6 +228,8 @@ static int vfio_gmem_get_pfn(void *data, u64 offset, unsigned long *pfn)
int ret;
guard(rwsem_read)(&vdev->memory_lock);
+ if (ctx->callbacks_detached)
+ return -ENODEV;
if (vdev->pm_runtime_engaged || !__vfio_pci_memory_enabled(vdev))
return -EAGAIN;
@@ -255,6 +259,10 @@ static int vfio_gmem_prepare_conversion(struct guest_memfd_device_context *conte
struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(context);
iommufd_vdevice_mmio_lock(ctx->ivdev);
+ if (req && ctx->callbacks_detached) {
+ iommufd_vdevice_mmio_unlock(ctx->ivdev);
+ return -ENODEV;
+ }
if (req && req->vdev_id != ctx->ivdev->virt_id) {
iommufd_vdevice_mmio_unlock(ctx->ivdev);
return -EPERM;
@@ -284,6 +292,24 @@ static void vfio_gmem_finish_conversion(struct guest_memfd_device_context *conte
iommufd_vdevice_mmio_unlock(ctx->ivdev);
}
+/**
+ * vfio_gmem_detach() - stop invalidation callbacks before core teardown
+ * @data: attached provider context
+ *
+ * Mark the provider detached under the VFIO memory lock and revoke BAR
+ * translations.
+ */
+static void vfio_gmem_detach(void *data)
+{
+ struct vfio_pci_gmem *ctx = vfio_pci_gmem_from_context(data);
+
+ guard(iommufd_vdevice_mmio)(ctx->ivdev);
+ scoped_guard(rwsem_write, &ctx->vdev->memory_lock) {
+ ctx->callbacks_detached = true;
+ }
+ vfio_gmem_invalidate_mmio(&ctx->context);
+}
+
/**
* vfio_gmem_release() - release a provider after successful core cleanup
* @data: provider context to release
@@ -315,6 +341,7 @@ const struct guest_memfd_device_operations vfio_pci_gmem_ops = {
.get_pfn = vfio_gmem_get_pfn,
.prepare_conversion = vfio_gmem_prepare_conversion,
.finish_conversion = vfio_gmem_finish_conversion,
+ .detach = vfio_gmem_detach,
.release = vfio_gmem_release,
};
EXPORT_SYMBOL_GPL(vfio_pci_gmem_ops);
diff --git a/drivers/vfio/pci/vfio_pci_gmem.h b/drivers/vfio/pci/vfio_pci_gmem.h
index e5ad19ba6748..a47a8d386051 100644
--- a/drivers/vfio/pci/vfio_pci_gmem.h
+++ b/drivers/vfio/pci/vfio_pci_gmem.h
@@ -17,6 +17,12 @@ struct vfio_pci_gmem {
unsigned long owner;
resource_size_t start[PCI_STD_NUM_BARS];
resource_size_t len[PCI_STD_NUM_BARS];
+ /*
+ * Disables callbacks and new bind/private requests. Set once with
+ * both the vDEVICE MMIO and VFIO memory locks held; readers hold
+ * either lock. Shared cleanup remains allowed after detachment.
+ */
+ bool callbacks_detached;
};
static inline struct vfio_pci_gmem *
diff --git a/drivers/vfio/pci/vfio_pci_gmem_access.c b/drivers/vfio/pci/vfio_pci_gmem_access.c
index a3975b536e22..254a311b5072 100644
--- a/drivers/vfio/pci/vfio_pci_gmem_access.c
+++ b/drivers/vfio/pci/vfio_pci_gmem_access.c
@@ -14,7 +14,7 @@
* Called by vfio_pci_zap_bars() after it removes host BAR mappings. Notify
* guest_memfd synchronously so KVM removes the corresponding shared guest
* mappings. This helper does not remove host mappings or change attributes.
- * Skip notification if no provider is attached.
+ * Skip notification if no provider is attached or its callbacks are detached.
*
* The guest_memfd callback takes SRCU and the KVM MMU lock without taking
* the inode invalidate lock, which may already be held during conversion.
@@ -25,7 +25,7 @@ void vfio_pci_gmem_invalidate(struct vfio_pci_core_device *vdev)
struct vfio_pci_gmem *ctx = vdev->gmem;
lockdep_assert_held_write(&vdev->memory_lock);
- if (ctx)
+ if (ctx && !ctx->callbacks_detached)
ctx->context.device->invalidate(ctx->context.device);
}
@@ -33,11 +33,12 @@ void vfio_pci_gmem_invalidate(struct vfio_pci_core_device *vdev)
* vfio_gmem_invalidate_mmio() - revoke host and shared guest BAR translations
* @data: attached provider context
*
- * Initiate BAR revocation for conversion or access refresh. Call
+ * Initiate BAR revocation for conversion, access refresh or teardown. Call
* vfio_pci_zap_and_down_write_memory_lock() to remove host BAR mappings;
* its BAR zap path also calls vfio_pci_gmem_invalidate() to remove the
* corresponding shared guest mappings. Release the VFIO memory lock before
- * returning.
+ * returning. Host mappings are removed even if guest_memfd callbacks are
+ * detached.
*
* During private conversion, the caller holds the inode invalidate lock to
* prevent new guest_memfd faults. Host faults and I/O try that lock when
@@ -68,6 +69,7 @@ void vfio_gmem_invalidate_mmio(void *data)
* Consult completed guest_memfd attributes for the requested interval.
* Private mappings in another interval do not block this access. Try the
* inode lock instead of waiting in the reverse order from conversion.
+ * Detachment excludes access during teardown.
*
* Return: true when the requested interval is accessible.
*/
@@ -84,6 +86,8 @@ bool vfio_pci_gmem_access_allowed(struct vfio_pci_core_device *vdev,
if (!ctx->ivdev)
return false;
+ if (ctx->callbacks_detached)
+ return false;
return ctx->context.device->host_accessible(ctx->context.device,
offset, size);
}
diff --git a/include/linux/guest_memfd.h b/include/linux/guest_memfd.h
index 24f4b3c79afa..8373ef8150c5 100644
--- a/include/linux/guest_memfd.h
+++ b/include/linux/guest_memfd.h
@@ -41,6 +41,7 @@ struct guest_memfd_device_operations {
const struct guest_memfd_device_request *req);
void (*finish_conversion)(struct guest_memfd_device_context *context,
const struct guest_memfd_device_request *req);
+ void (*detach)(void *data);
void (*release)(void *data);
};
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index fc37066bfaa5..f493b452fae8 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -381,6 +381,7 @@ static void kvm_gmem_invalidate_end(struct inode *inode, pgoff_t start,
static int kvm_gmem_device_convert(struct inode *inode, pgoff_t start,
pgoff_t nr_pages, const struct guest_memfd_device_request *req,
const struct kvm_memory_slot *slot);
+static void kvm_gmem_device_close(struct inode *inode);
static long kvm_gmem_punch_hole(struct inode *inode, loff_t offset, loff_t len)
{
@@ -503,6 +504,9 @@ static int kvm_gmem_release(struct inode *inode, struct file *file)
filemap_invalidate_lock(inode->i_mapping);
+ if (GMEM_I(inode)->device_ops)
+ kvm_gmem_device_close(inode);
+
xa_for_each(&f->bindings, index, slot) {
WRITE_ONCE(slot->gmem.file, NULL);
}
@@ -1362,6 +1366,23 @@ int kvm_gmem_device_map(struct kvm *kvm, u64 gpa, u64 size, u64 pa, u64 vdev_id)
}
EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_device_map);
+static void kvm_gmem_device_close(struct inode *inode)
+{
+ struct gmem_inode *gi = GMEM_I(inode);
+ int ret;
+
+ gi->device_ops->detach(gi->provider);
+ ret = gi->device_ops->prepare_conversion(gi->provider, NULL);
+ if (!ret) {
+ ret = kvm_gmem_device_make_shared(inode, 0,
+ i_size_read(inode) >> PAGE_SHIFT);
+ if (!ret)
+ WRITE_ONCE(gi->device_has_private, false);
+ gi->device_ops->finish_conversion(gi->provider, NULL);
+ }
+ WRITE_ONCE(gi->device.core, NULL);
+}
+
static int kvm_gmem_attach_resource(struct kvm *kvm,
struct inode *inode, int resource_fd)
{
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH v1 6/6] KVM: Invalidate guest_memfd mappings before removing memslot bindings
2026-10-10 7:24 [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources Aneesh Kumar K.V (Arm)
` (4 preceding siblings ...)
2026-10-10 7:25 ` [RFC PATCH v1 5/6] KVM/VFIO: Remove device mappings during guest_memfd teardown Aneesh Kumar K.V (Arm)
@ 2026-10-10 7:25 ` Aneesh Kumar K.V (Arm)
5 siblings, 0 replies; 7+ messages in thread
From: Aneesh Kumar K.V (Arm) @ 2026-10-10 7:25 UTC (permalink / raw)
To: linux-coco, linux-kernel
Cc: Aneesh Kumar K.V (Arm),
Ackerley Tng, Alex Williamson, David Woodhouse,
David Hildenbrand, Jason Gunthorpe, Joerg Roedel (AMD),
Kevin Tian, Paolo Bonzini, Robin Murphy, Sean Christopherson,
Will Deacon, Alexey Kardashevskiy, Xu Yilun, Catalin Marinas,
Suzuki K Poulose, Steven Price, Fred Griffoul, iommu, kvm
Unmap guest_memfd ranges before deleting their memslot bindings. Remove
private device mappings before committing deletion, and invalidate
remaining mappings while holding the inode invalidate lock.
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: Sean Christopherson <seanjc@google.com>
Cc: David Hildenbrand <david@kernel.org>
Cc: kvm@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Assisted-by: Codex
Signed-off-by: Aneesh Kumar K.V (Arm) <aneesh.kumar@kernel.org>
---
virt/kvm/guest_memfd.c | 70 ++++++++++++++++++++++++++++++++++++++++--
virt/kvm/guest_memfd.h | 6 ++++
virt/kvm/kvm_main.c | 4 +++
3 files changed, 77 insertions(+), 3 deletions(-)
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index f493b452fae8..f58270d0510d 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -1330,6 +1330,58 @@ static int kvm_gmem_device_convert(struct inode *inode, pgoff_t start,
return ret;
}
+/**
+ * kvm_gmem_prepare_memslot_delete() - remove guest_memfd mappings before deletion
+ * @kvm: VM whose memslot is being removed
+ * @slot: original guest_memfd memslot being removed
+ *
+ * The caller has replaced the active slot with an entry marked
+ * KVM_MEMSLOT_INVALID, preventing new guest faults from creating mappings.
+ * Remove private device mappings, then invalidate the remaining guest_memfd
+ * mappings, including shared mappings.
+ *
+ * Perform this cleanup while the guest_memfd binding still provides the
+ * file-offset-to-GPA translation needed to locate the guest mappings.
+ * The device provider remains attached until guest_memfd close.
+ *
+ * Caller holds slots_lock.
+ * Return: 0 on success or a negative error code.
+ */
+int kvm_gmem_prepare_memslot_delete(const struct kvm_memory_slot *slot)
+{
+ struct file *file = READ_ONCE(slot->gmem.file);
+ pgoff_t start = slot->gmem.pgoff;
+ pgoff_t end = start + slot->npages;
+ struct inode *inode;
+ int ret;
+
+ /* slots_lock keeps the bindings alive until file release can finish. */
+ file = get_file_active(&file);
+ if (!file) {
+ /* Let .release drain the bindings before deletion can discard them. */
+ if (READ_ONCE(slot->gmem.file))
+ return -EBUSY;
+ return 0;
+ }
+ inode = file_inode(file);
+ if (GMEM_I(inode)->device_ops) {
+ ret = kvm_gmem_device_convert(inode, start, slot->npages, NULL, NULL);
+ if (ret)
+ goto out;
+ }
+
+ filemap_invalidate_lock(inode->i_mapping);
+ __kvm_gmem_invalidate_start(file->private_data, start, end,
+ kvm_gmem_get_all_gfns_filter(inode));
+ __kvm_gmem_invalidate_end(file->private_data, start, end);
+ filemap_invalidate_unlock(inode->i_mapping);
+ ret = 0;
+out:
+ fput(file);
+ return ret;
+}
+
+
int kvm_gmem_device_map(struct kvm *kvm, u64 gpa, u64 size, u64 pa, u64 vdev_id)
{
struct guest_memfd_device_request req = { gpa, pa, vdev_id };
@@ -1620,11 +1672,19 @@ int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
return r;
}
-static void __kvm_gmem_unbind(struct kvm_memory_slot *slot, struct gmem_file *f)
+static void __kvm_gmem_unbind(struct kvm_memory_slot *slot, struct file *file)
{
+ struct gmem_file *f = file->private_data;
unsigned long start = slot->gmem.pgoff;
unsigned long end = start + slot->npages;
+ /*
+ * Unmap this slot's guest mappings before removing it from f->bindings.
+ * The invalidation helper walks f->bindings to find the slots to unmap.
+ */
+ __kvm_gmem_invalidate_start(f, start, end,
+ kvm_gmem_get_all_gfns_filter(file_inode(file)));
+ __kvm_gmem_invalidate_end(f, start, end);
xa_store_range(&f->bindings, start, end - 1, NULL, GFP_KERNEL);
/*
@@ -1656,12 +1716,16 @@ void kvm_gmem_unbind(struct kvm_memory_slot *slot)
* until the caller drops slots_lock.
*/
if (!file) {
- __kvm_gmem_unbind(slot, slot->gmem.file->private_data);
+ struct file *closing_file = slot->gmem.file;
+
+ filemap_invalidate_lock(closing_file->f_mapping);
+ __kvm_gmem_unbind(slot, closing_file);
+ filemap_invalidate_unlock(closing_file->f_mapping);
return;
}
filemap_invalidate_lock(file->f_mapping);
- __kvm_gmem_unbind(slot, file->private_data);
+ __kvm_gmem_unbind(slot, file);
filemap_invalidate_unlock(file->f_mapping);
}
diff --git a/virt/kvm/guest_memfd.h b/virt/kvm/guest_memfd.h
index 0f9c6f840838..6fad6adfdadc 100644
--- a/virt/kvm/guest_memfd.h
+++ b/virt/kvm/guest_memfd.h
@@ -11,6 +11,7 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args);
int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot,
unsigned int fd, uoff_t offset);
void kvm_gmem_unbind(struct kvm_memory_slot *slot);
+int kvm_gmem_prepare_memslot_delete(const struct kvm_memory_slot *slot);
#else
static inline int kvm_gmem_init(struct module *module)
{
@@ -29,6 +30,11 @@ static inline void kvm_gmem_unbind(struct kvm_memory_slot *slot)
{
WARN_ON_ONCE(1);
}
+
+static inline int kvm_gmem_prepare_memslot_delete(const struct kvm_memory_slot *slot)
+{
+ return 0;
+}
#endif /* CONFIG_KVM_GUEST_MEMFD */
#endif /* __KVM_GUEST_MEMFD_H__ */
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index f3f46a8563b3..58da57b8a268 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -1696,6 +1696,10 @@ static int kvm_prepare_memory_region(struct kvm *kvm,
}
r = kvm_arch_prepare_memory_region(kvm, old, new, change);
+ if (!r && change == KVM_MR_DELETE) {
+ if (old->flags & KVM_MEM_GUEST_MEMFD)
+ r = kvm_gmem_prepare_memslot_delete(old);
+ }
/* Free the bitmap on failure if it was allocated above. */
if (r && new && new->dirty_bitmap && (!old || !old->dirty_bitmap))
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-10 7:26 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-10 7:24 [RFC PATCH v1 0/6] KVM/VFIO: guest_memfd support for device MMIO resources Aneesh Kumar K.V (Arm)
2026-10-10 7:24 ` [RFC PATCH v1 1/6] KVM: guest_memfd: attach and bind device resources Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 2/6] KVM: guest_memfd: Support private/shared conversion of device memory Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 3/6] iommufd: Support MMIO provider attachment to vDEVICEs Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 4/6] vfio/pci: Provide guest_memfd backing for PCI BARs Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 5/6] KVM/VFIO: Remove device mappings during guest_memfd teardown Aneesh Kumar K.V (Arm)
2026-10-10 7:25 ` [RFC PATCH v1 6/6] KVM: Invalidate guest_memfd mappings before removing memslot bindings Aneesh Kumar K.V (Arm)
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®