mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Fred Griffoul <griffoul@gmail.com>
To: Paolo Bonzini <pbonzini@redhat.com>,
	Sean Christopherson <seanjc@google.com>,
	Marc Zyngier <maz@kernel.org>, Oliver Upton <oupton@kernel.org>,
	Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>,
	Alexander Viro <viro@zeniv.linux.org.uk>,
	Christian Brauner <brauner@kernel.org>, Jan Kara <jack@suse.cz>,
	Jason Gunthorpe <jgg@ziepe.ca>, Kevin Tian <kevin.tian@intel.com>,
	Joerg Roedel <joro@8bytes.org>, Will Deacon <will@kernel.org>,
	Robin Murphy <robin.murphy@arm.com>,
	Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
	Borislav Petkov <bp@alien8.de>,
	Dave Hansen <dave.hansen@linux.intel.com>,
	x86@kernel.org, "H . Peter Anvin" <hpa@zytor.com>,
	Jonathan Corbet <corbet@lwn.net>, Shuah Khan <shuah@kernel.org>
Cc: David Woodhouse <dwmw2@infradead.org>,
	Ackerley Tng <ackerleytng@google.com>,
	Lorenzo Stoakes <ljs@kernel.org>,
	"Liam R . Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>, Joey Gouly <joey.gouly@arm.com>,
	Suzuki K Poulose <suzuki.poulose@arm.com>,
	Zenghui Yu <yuzenghui@huawei.com>,
	Steffen Eiden <seiden@linux.ibm.com>,
	linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
	kvmarm@lists.linux.dev, iommu@lists.linux.dev,
	linux-fsdevel@vger.kernel.org, linux-mm@kvack.org,
	linux-kselftest@vger.kernel.org
Subject: [PATCH 3/9] KVM: guest_memfd: Add a memory provider backing
Date: Tue,  6 Oct 2026 18:32:29 +0000	[thread overview]
Message-ID: <20261006183235.16576-4-griffoul@gmail.com> (raw)
In-Reply-To: <20261006183235.16576-1-griffoul@gmail.com>

From: Fred Griffoul <fgriffo@amazon.co.uk>

A driver can back a guest_memfd today only by implementing
kvm_gmem_ops itself, and must then keep any device mapping of the same
memory consistent on its own.

Add GUEST_MEMFD_FLAG_USE_PROVIDER. KVM_CREATE_GUEST_MEMFD then takes a
memory provider file in provider_fd, which uses the first reserved word
and must be zero without the flag. Guest faults ask the provider for
each frame, and a revoke removes the range from the guest and from the
VMM's mapping. A page that is not backed, or that is not RAM, exits with
KVM_EXIT_MEMORY_FAULT. Read-only pages are mapped read only, and
fallocate() is not supported.

guest_memfd maps itself into userspace from the same frames, so that
a provider needs no fault handler. Pages marked NO_USER_MAP raise
SIGBUS there.

USE_PROVIDER requires GUEST_MEMFD_FLAG_MMAP, so that the memory slot
is gmem-only. An architecture opts in; x86 does so for VMs without
private or encrypted memory, and other architectures refuse the flag
for now. Native guest_memfd files now take the invalidate lock when
a file is added and when a closing file is unbound.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 Documentation/virt/kvm/api.rst                |  45 ++-
 arch/x86/kvm/x86.c                            |  10 +
 include/linux/kvm_host.h                      |   4 +
 include/uapi/linux/kvm.h                      |  11 +-
 tools/include/uapi/linux/kvm.h                |  11 +-
 .../testing/selftests/kvm/guest_memfd_test.c  |   6 +
 virt/kvm/Kconfig                              |   1 +
 virt/kvm/guest_memfd.c                        | 265 +++++++++++++++++-
 8 files changed, 338 insertions(+), 15 deletions(-)

diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index a5f9ee92f43e..d9aaf8a76b59 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6431,7 +6431,9 @@ and cannot be resized  (guest_memfd files do however support PUNCH_HOLE).
   struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	__s32 provider_fd;
+	__u32 pad;
+	__u64 reserved[5];
   };
 
 Conceptually, the inode backing a guest_memfd file represents physical memory,
@@ -6453,15 +6455,38 @@ a single guest_memfd file, but the bound ranges must not overlap).
 The capability KVM_CAP_GUEST_MEMFD_FLAGS enumerates the `flags` that can be
 specified via KVM_CREATE_GUEST_MEMFD.  Currently defined flags:
 
-  ============================ ================================================
-  GUEST_MEMFD_FLAG_MMAP        Enable using mmap() on the guest_memfd file
-                               descriptor.
-  GUEST_MEMFD_FLAG_INIT_SHARED Make all memory in the file shared during
-                               KVM_CREATE_GUEST_MEMFD (memory files created
-                               without INIT_SHARED will be marked private).
-                               Shared memory can be faulted into host userspace
-                               page tables. Private memory cannot.
-  ============================ ================================================
+  ============================= ================================================
+  GUEST_MEMFD_FLAG_MMAP         Enable using mmap() on the guest_memfd file
+                                descriptor.
+  GUEST_MEMFD_FLAG_INIT_SHARED  Make all memory in the file shared during
+                                KVM_CREATE_GUEST_MEMFD (memory files created
+                                without INIT_SHARED will be marked private).
+                                Shared memory can be faulted into host userspace
+                                page tables. Private memory cannot.
+  GUEST_MEMFD_FLAG_USE_PROVIDER Take the memory of the file from the memory
+                                provider behind `provider_fd`, instead of
+                                allocating it.  Requires
+                                GUEST_MEMFD_FLAG_MMAP.  Not supported for VMs
+                                with private or encrypted memory.
+  ============================= ================================================
+
+`pad` must be zero.  With GUEST_MEMFD_FLAG_USE_PROVIDER, `provider_fd` is a file
+handed out by a memory provider (see include/linux/mem_provider.h); without it,
+`provider_fd` must be zero.
+The provider decides which pages exist and what backs them, and can change
+this at any time.  The guest and any host mapping of the guest_memfd follow
+the change.  A guest access to a page that the provider does not back exits
+to userspace with KVM_EXIT_MEMORY_FAULT, and so does an access to a page that
+is not RAM.  A page that the provider marks read only is mapped read only for
+the guest.  fallocate() is not supported.  Only x86 VMs of type
+KVM_X86_DEFAULT_VM support GUEST_MEMFD_FLAG_USE_PROVIDER.
+
+mmap() of the guest_memfd maps the provider's pages into userspace on fault.
+An access raises SIGBUS if the provider does not back the page, if the page is
+not RAM, if the provider marks it as not to be mapped into userspace, or if it
+is a write to a read-only page.
+KVM's own accesses through the memslot's userspace address fail in the same
+cases.
 
 When the KVM MMU performs a PFN lookup to service a guest fault and the backing
 guest_memfd has the GUEST_MEMFD_FLAG_MMAP set, then the fault will always be
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index afcac1042947..3669316f5712 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -14124,6 +14124,16 @@ bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm)
 	return !kvm_arch_has_private_mem(kvm);
 }
 
+/*
+ * Only VMs without encrypted memory: SEV and SEV-ES guests have no private
+ * memory in KVM's sense, but their memory would need reclaiming when a
+ * provider takes it back.
+ */
+bool kvm_arch_gmem_supports_provider(struct kvm *kvm)
+{
+	return !kvm || kvm->arch.vm_type == KVM_X86_DEFAULT_VM;
+}
+
 #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREPARE
 int kvm_arch_gmem_prepare(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, int max_order)
 {
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 7281d0e94121..3ea591642542 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -832,11 +832,15 @@ static inline bool kvm_arch_has_private_mem(struct kvm *kvm)
 
 #ifdef CONFIG_KVM_GUEST_MEMFD
 bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm);
+bool kvm_arch_gmem_supports_provider(struct kvm *kvm);
 
 static inline u64 kvm_gmem_get_supported_flags(struct kvm *kvm)
 {
 	u64 flags = GUEST_MEMFD_FLAG_MMAP;
 
+	if (kvm_arch_gmem_supports_provider(kvm))
+		flags |= GUEST_MEMFD_FLAG_USE_PROVIDER;
+
 	if (!kvm || kvm_arch_supports_gmem_init_shared(kvm))
 		flags |= GUEST_MEMFD_FLAG_INIT_SHARED;
 
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index 419011097fa8..820448f82e22 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1654,11 +1654,20 @@ struct kvm_memory_attributes {
 #define KVM_CREATE_GUEST_MEMFD	_IOWR(KVMIO,  0xd4, struct kvm_create_guest_memfd)
 #define GUEST_MEMFD_FLAG_MMAP		(1ULL << 0)
 #define GUEST_MEMFD_FLAG_INIT_SHARED	(1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_PROVIDER	(1ULL << 2)
 
 struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	/*
+	 * With GUEST_MEMFD_FLAG_USE_PROVIDER: a memory provider file whose
+	 * pages back this guest_memfd.  The provider may take
+	 * any range back at any time; the guest and every host mapping of
+	 * this guest_memfd follow.  Must be 0 otherwise.
+	 */
+	__s32 provider_fd;
+	__u32 pad;
+	__u64 reserved[5];
 };
 
 #define KVM_PRE_FAULT_MEMORY	_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/tools/include/uapi/linux/kvm.h b/tools/include/uapi/linux/kvm.h
index d0c0c8605976..0f4d1dd0e931 100644
--- a/tools/include/uapi/linux/kvm.h
+++ b/tools/include/uapi/linux/kvm.h
@@ -1644,11 +1644,20 @@ struct kvm_memory_attributes {
 #define KVM_CREATE_GUEST_MEMFD	_IOWR(KVMIO,  0xd4, struct kvm_create_guest_memfd)
 #define GUEST_MEMFD_FLAG_MMAP		(1ULL << 0)
 #define GUEST_MEMFD_FLAG_INIT_SHARED	(1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_PROVIDER	(1ULL << 2)
 
 struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	/*
+	 * With GUEST_MEMFD_FLAG_USE_PROVIDER: a memory provider file whose
+	 * pages back this guest_memfd.  The provider may take
+	 * any range back at any time; the guest and every host mapping of
+	 * this guest_memfd follow.  Must be 0 otherwise.
+	 */
+	__s32 provider_fd;
+	__u32 pad;
+	__u64 reserved[5];
 };
 
 #define KVM_PRE_FAULT_MEMORY	_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 2233d871a38f..91ff10ac6274 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -405,6 +405,12 @@ static void test_guest_memfd_flags(struct kvm_vm *vm)
 
 	for (flag = BIT(0); flag; flag <<= 1) {
 		fd = __vm_create_guest_memfd(vm, page_size, flag);
+		/* USE_PROVIDER also needs MMAP and a provider file. */
+		if (flag == GUEST_MEMFD_FLAG_USE_PROVIDER && (flag & valid_flags)) {
+			TEST_ASSERT(fd < 0 && errno == EINVAL,
+				    "guest_memfd() with USE_PROVIDER alone should fail with EINVAL");
+			continue;
+		}
 		if (flag & valid_flags) {
 			TEST_ASSERT(fd >= 0,
 				    "guest_memfd() with flag '0x%lx' should succeed",
diff --git a/virt/kvm/Kconfig b/virt/kvm/Kconfig
index 794976b88c6f..cfb6c4e51128 100644
--- a/virt/kvm/Kconfig
+++ b/virt/kvm/Kconfig
@@ -105,6 +105,7 @@ config KVM_GENERIC_MEMORY_ATTRIBUTES
 
 config KVM_GUEST_MEMFD
        select XARRAY_MULTI
+       select MEM_PROVIDER
        bool
 
 config HAVE_KVM_ARCH_GMEM_PREPARE
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index a509f1a96c0b..aedd8630e3ea 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -4,6 +4,7 @@
 #include <linux/falloc.h>
 #include <linux/fs.h>
 #include <linux/kvm_host.h>
+#include <linux/mem_provider.h>
 #include <linux/mempolicy.h>
 #include <linux/pseudo_fs.h>
 #include <linux/pagemap.h>
@@ -44,6 +45,9 @@ struct gmem_inode {
 	struct list_head gmem_file_list;
 
 	u64 flags;
+
+	/* The provider of the pages, with GUEST_MEMFD_FLAG_USE_PROVIDER. */
+	struct mem_provider_attachment att;
 };
 
 static __always_inline struct gmem_inode *GMEM_I(struct inode *inode)
@@ -654,7 +658,222 @@ static const struct kvm_gmem_ops kvm_gmem_native_ops = {
 	.fallocate = kvm_gmem_native_fallocate,
 };
 
-static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
+/*
+ * Can @kvm use a memory provider?  @kvm is NULL for the system-wide capability.
+ * An architecture opts in by overriding this.  VMs whose memory is private or
+ * encrypted are not supported: each frame would need preparing before use and
+ * reclaiming when the provider takes it back.
+ */
+bool __weak kvm_arch_gmem_supports_provider(struct kvm *kvm)
+{
+	return false;
+}
+
+/*
+ * A guest_memfd backed by a memory provider (include/linux/mem_provider.h).
+ *
+ * The provider owns the frames, and guest_memfd keeps no state for them.  A
+ * guest fault asks the provider for the frame each time.  A fault that
+ * races with a change in the provider is retried by KVM, because the
+ * provider revokes the range after the change and the revoke opens and
+ * closes KVM's invalidation window.  Host mappings are made by guest_memfd
+ * from the same frames and attributes, and removed by the revoke.
+ *
+ * Bindings and release are the same as for the native guest_memfd.
+ */
+static void kvm_gmem_provider_revoke(struct mem_provider_attachment *att,
+				     loff_t offset, loff_t len)
+{
+	struct gmem_inode *gi = container_of(att, struct gmem_inode, att);
+	struct inode *inode = &gi->vfs_inode;
+	pgoff_t start, end;
+	struct gmem_file *f;
+
+	if (offset >= att->size)
+		return;
+	len = min(len, att->size - offset);
+	start = offset >> PAGE_SHIFT;
+	end = DIV_ROUND_UP(offset + len, PAGE_SIZE);
+
+	/*
+	 * The bindings must be stable so that each start is matched by an
+	 * end.  Only VMs without private memory use a provider, but remove
+	 * both kinds of mapping so that none is missed.
+	 */
+	filemap_invalidate_lock(inode->i_mapping);
+	kvm_gmem_for_each_file(f, inode)
+		__kvm_gmem_invalidate_start(f, start, end,
+					    KVM_FILTER_SHARED | KVM_FILTER_PRIVATE);
+	unmap_mapping_range(inode->i_mapping, (loff_t)start << PAGE_SHIFT,
+			    (loff_t)(end - start) << PAGE_SHIFT, 1);
+	kvm_gmem_for_each_file(f, inode)
+		__kvm_gmem_invalidate_end(f, start, end);
+	filemap_invalidate_unlock(inode->i_mapping);
+}
+
+static int kvm_gmem_provider_get_pfn(struct file *file, struct kvm *kvm,
+				     struct kvm_memory_slot *slot, gfn_t gfn,
+				     kvm_pfn_t *pfn, struct page **page,
+				     int *max_order, bool *writable)
+{
+	struct gmem_inode *gi = GMEM_I(file_inode(file));
+	pgoff_t index = kvm_gmem_get_index(slot, gfn);
+	unsigned long frame;
+	int order, ret;
+	u32 attrs;
+
+	if (file != READ_ONCE(slot->gmem.file))
+		return -EFAULT;
+	if (xa_load(&gmem_file_of(file)->bindings, index) != slot)
+		return -EIO;
+
+	/* A large mapping needs the block aligned in both index and gfn. */
+	order = PUD_ORDER;
+	if (index != gfn)
+		order = min_t(int, order, __ffs(index ^ gfn));
+
+	ret = mem_provider_get_page(&gi->att, index, &frame, &order, &attrs);
+	if (ret)
+		return ret;
+
+	/*
+	 * KVM decides the guest's memory type itself, so accept RAM only.
+	 * -EFAULT, so that the VMM gets a memory fault exit for the access.
+	 */
+	if (mem_provider_type(attrs) != MEM_PROVIDER_TYPE_RAM)
+		return -EFAULT;
+
+	if (attrs & MEM_PROVIDER_ATTR_READONLY) {
+		if (!writable)
+			return -EPERM;
+		*writable = false;
+	}
+
+	*pfn = frame;
+	if (max_order)
+		*max_order = order;
+	return 0;
+}
+
+/* The protection for a host mapping of a frame with @attrs. */
+static pgprot_t kvm_gmem_provider_prot(struct vm_area_struct *vma, u32 attrs)
+{
+	vm_flags_t flags = vma->vm_flags;
+
+	if (attrs & MEM_PROVIDER_ATTR_READONLY)
+		flags &= ~VM_WRITE;
+	return vm_get_page_prot(flags);
+}
+
+/*
+ * Map page @vmf->pgoff of a provider-backed guest_memfd for the host.
+ *
+ * The invalidate lock is held shared across get_page() and the insert, and
+ * a revoke holds it exclusive while it unmaps the range, so a frame is never
+ * mapped after the revoke that removes it.  A writable page is mapped
+ * writable, and a read-only page without write permission.  With @mkwrite,
+ * the PTE is read only and the write must be checked: replace it under the
+ * lock, so that a change to read only cannot slip between the check and the
+ * upgrade.
+ */
+static vm_fault_t kvm_gmem_provider_map_host(struct vm_fault *vmf, bool mkwrite)
+{
+	struct vm_area_struct *vma = vmf->vma;
+	struct inode *inode = file_inode(vma->vm_file);
+	unsigned long uaddr = vmf->address & PAGE_MASK;
+	bool write = vmf->flags & FAULT_FLAG_WRITE;
+	unsigned long pfn;
+	int order = 0;
+	vm_fault_t ret;
+	u32 attrs;
+
+	if (((loff_t)vmf->pgoff << PAGE_SHIFT) >= i_size_read(inode))
+		return VM_FAULT_SIGBUS;
+
+	filemap_invalidate_lock_shared(inode->i_mapping);
+	if (mem_provider_get_page(&GMEM_I(inode)->att, vmf->pgoff, &pfn,
+				  &order, &attrs)) {
+		ret = VM_FAULT_SIGBUS;
+	} else if (mem_provider_type(attrs) != MEM_PROVIDER_TYPE_RAM ||
+		   (attrs & MEM_PROVIDER_ATTR_NO_USER_MAP) ||
+		   (write && (attrs & MEM_PROVIDER_ATTR_READONLY))) {
+		ret = VM_FAULT_SIGBUS;
+	} else {
+		if (mkwrite)
+			zap_special_vma_range(vma, uaddr, PAGE_SIZE);
+		ret = vmf_insert_pfn_prot(vma, uaddr, pfn,
+					  kvm_gmem_provider_prot(vma, attrs));
+	}
+	filemap_invalidate_unlock_shared(inode->i_mapping);
+	return ret;
+}
+
+static vm_fault_t kvm_gmem_provider_fault(struct vm_fault *vmf)
+{
+	return kvm_gmem_provider_map_host(vmf, false);
+}
+
+/* A write to a page that a read fault, or mprotect(), left read only. */
+static vm_fault_t kvm_gmem_provider_pfn_mkwrite(struct vm_fault *vmf)
+{
+	return kvm_gmem_provider_map_host(vmf, true);
+}
+
+static const struct vm_operations_struct kvm_gmem_provider_vm_ops = {
+	.fault		= kvm_gmem_provider_fault,
+	.pfn_mkwrite	= kvm_gmem_provider_pfn_mkwrite,
+};
+
+static int kvm_gmem_provider_mmap(struct file *file, struct vm_area_struct *vma)
+{
+	struct inode *inode = file_inode(file);
+	pgoff_t npages = i_size_read(inode) >> PAGE_SHIFT;
+
+	if (!kvm_gmem_supports_mmap(inode))
+		return -ENODEV;
+
+	if ((vma->vm_flags & (VM_SHARED | VM_MAYSHARE)) !=
+	    (VM_SHARED | VM_MAYSHARE))
+		return -EINVAL;
+
+	if (vma->vm_pgoff >= npages || vma_pages(vma) > npages - vma->vm_pgoff)
+		return -EINVAL;
+
+	/*
+	 * The frames may have no struct page.  They are inserted on fault, so
+	 * that a range that is revoked and comes back is reached again.
+	 */
+	vm_flags_set(vma, VM_PFNMAP | VM_IO | VM_DONTEXPAND | VM_DONTDUMP);
+	vma->vm_ops = &kvm_gmem_provider_vm_ops;
+	return 0;
+}
+
+static const struct kvm_gmem_ops kvm_gmem_provider_ops = {
+	.bind    = kvm_gmem_native_bind,
+	.unbind  = kvm_gmem_native_unbind,
+	.get_pfn = kvm_gmem_provider_get_pfn,
+	.release = kvm_gmem_native_release,
+	.mmap    = kvm_gmem_provider_mmap,
+	/* The provider decides which pages exist: no fallocate(). */
+};
+
+static int kvm_gmem_provider_attach(struct inode *inode, int provider_fd)
+{
+	struct file *file;
+	int ret;
+
+	file = fget(provider_fd);
+	if (!file)
+		return -EBADF;
+
+	ret = mem_provider_attach(&GMEM_I(inode)->att, file,
+				  i_size_read(inode), kvm_gmem_provider_revoke);
+	fput(file);
+	return ret;
+}
+
+static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags,
+			     int provider_fd)
 {
 	static const char *name = "[kvm-gmem]";
 	struct gmem_file *f;
@@ -695,6 +914,12 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
 
 	GMEM_I(inode)->flags = flags;
 
+	if (flags & GUEST_MEMFD_FLAG_USE_PROVIDER) {
+		err = kvm_gmem_provider_attach(inode, provider_fd);
+		if (err)
+			goto err_inode;
+	}
+
 	file = alloc_file_pseudo(inode, kvm_gmem_mnt, name, O_RDWR, &kvm_gmem_fops);
 	if (IS_ERR(file)) {
 		err = PTR_ERR(file);
@@ -703,12 +928,18 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
 
 	file->f_flags |= O_LARGEFILE;
 	file->private_data = &f->backing;
-	f->backing.ops = &kvm_gmem_native_ops;
+	if (flags & GUEST_MEMFD_FLAG_USE_PROVIDER)
+		f->backing.ops = &kvm_gmem_provider_ops;
+	else
+		f->backing.ops = &kvm_gmem_native_ops;
 
 	kvm_get_kvm(kvm);
 	f->kvm = kvm;
 	xa_init(&f->bindings);
+	/* A provider can revoke, and walk the list, as soon as it is attached. */
+	filemap_invalidate_lock(inode->i_mapping);
 	list_add(&f->entry, &GMEM_I(inode)->gmem_file_list);
+	filemap_invalidate_unlock(inode->i_mapping);
 
 	fd_install(fd, file);
 	return fd;
@@ -735,7 +966,20 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args)
 	if (size <= 0 || !PAGE_ALIGNED(size))
 		return -EINVAL;
 
-	return __kvm_gmem_create(kvm, size, flags);
+	if (args->pad ||
+	    (!(flags & GUEST_MEMFD_FLAG_USE_PROVIDER) && args->provider_fd))
+		return -EINVAL;
+
+	/*
+	 * Without MMAP the memslot is not gmem-only, and a VM without private
+	 * memory would fault through the slot's userspace address instead of
+	 * the provider.
+	 */
+	if ((flags & GUEST_MEMFD_FLAG_USE_PROVIDER) &&
+	    !(flags & GUEST_MEMFD_FLAG_MMAP))
+		return -EINVAL;
+
+	return __kvm_gmem_create(kvm, size, flags, args->provider_fd);
 }
 
 /*
@@ -911,9 +1155,14 @@ static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm,
 	 * bindings.  I.e. reaching this point means kvm_gmem_release() hasn't
 	 * yet destroyed the bindings or freed the gmem_file, and can't do so
 	 * until the caller drops slots_lock.
+	 *
+	 * A memory provider can still revoke, and walk the bindings, until
+	 * the inode is evicted, so take the invalidate lock in this case too.
 	 */
 	if (!file) {
+		filemap_invalidate_lock(slot_file->f_mapping);
 		__kvm_gmem_unbind(slot, gmem_file_of(slot_file));
+		filemap_invalidate_unlock(slot_file->f_mapping);
 		return;
 	}
 
@@ -1219,6 +1468,7 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb)
 	mpol_shared_policy_init(&gi->policy, NULL);
 
 	gi->flags = 0;
+	gi->att.ops = NULL;
 	INIT_LIST_HEAD(&gi->gmem_file_list);
 	return &gi->vfs_inode;
 }
@@ -1233,10 +1483,19 @@ static void kvm_gmem_free_inode(struct inode *inode)
 	kmem_cache_free(kvm_gmem_inode_cachep, GMEM_I(inode));
 }
 
+static void kvm_gmem_evict_inode(struct inode *inode)
+{
+	/* Every file is closed, so nothing maps the provider's frames. */
+	mem_provider_detach(&GMEM_I(inode)->att);
+	truncate_inode_pages_final(&inode->i_data);
+	clear_inode(inode);
+}
+
 static const struct super_operations kvm_gmem_super_operations = {
 	.statfs		= simple_statfs,
 	.alloc_inode	= kvm_gmem_alloc_inode,
 	.destroy_inode	= kvm_gmem_destroy_inode,
+	.evict_inode	= kvm_gmem_evict_inode,
 	.free_inode	= kvm_gmem_free_inode,
 };
 

  parent reply	other threads:[~2026-10-06 18:32 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-20 11:03 [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory David Woodhouse
2026-07-20 11:03 ` [RFC PATCH v2 01/11] KVM: selftests: sev_smoke_test: Only run VM types the host offers David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 02/11] KVM: selftests: sev_init2_tests: Derive SEV availability from KVM David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 03/11] KVM: SEV: Remove struct page dependency from SNP gmem paths David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 04/11] KVM: guest_memfd: Introduce guest memory ops and route native gmem through them David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 05/11] iommufd: Look up private-interconnect phys via exporter symbols David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 06/11] iommufd: Plumb dma-buf memory-type (RAM vs MMIO) through the phys map David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 07/11] KVM: guest_memfd: Add ops-driven page revocation David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 08/11] samples/kvm: Add guest_memfd backing sample David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 09/11] selftests/kvm: gmem_provider KVM-only tests David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 10/11] selftests/kvm: gmem_provider iommufd tests David Woodhouse
2026-07-20 11:03   ` [RFC PATCH v2 11/11] samples/kvm, selftests/kvm: Allow the gmem_provider NVMe DMA test on arm64 David Woodhouse
2026-07-20 15:11 ` [RFC PATCH v2 00/11] KVM: Allow alternative providers of guest_memfd backed by PFNMAP memory Paolo Bonzini
2026-07-20 16:39   ` David Woodhouse
2026-07-23  0:24 ` Ackerley Tng
2026-07-23  9:40   ` David Woodhouse
2026-07-23 16:01     ` Ackerley Tng
2026-10-05  9:55 ` [RFC PATCH 0/6] KVM: guest_memfd: back guest_memfd with an imported dma-buf Fred Griffoul
2026-10-05  9:55   ` [RFC PATCH 1/6] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
2026-10-05  9:55   ` [RFC PATCH 2/6] dma-buf: Add get_phys() to describe a physical run Fred Griffoul
2026-10-05 10:07     ` Christian König
2026-10-05 13:20       ` Fred Griffoul
2026-10-05 14:53         ` Christian König
2026-10-05  9:55   ` [RFC PATCH 3/6] dma-buf: Add ranged mapping invalidation Fred Griffoul
2026-10-05 10:08     ` Christian König
2026-10-05  9:55   ` [RFC PATCH 4/6] dma-buf: Allow dynamic attach without a device Fred Griffoul
2026-10-05  9:55   ` [RFC PATCH 5/6] KVM: guest_memfd: Add dma-buf backing Fred Griffoul
2026-10-05  9:55   ` [RFC PATCH 6/6] samples/kvm, selftests/kvm: Exercise " Fred Griffoul
2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
2026-10-06 18:32   ` [PATCH 1/9] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
2026-10-06 18:32   ` [PATCH 2/9] mm: Add memory providers Fred Griffoul
2026-10-06 18:32   ` Fred Griffoul [this message]
2026-10-06 18:32   ` [PATCH 4/9] iommufd: Track the domains of pages that are not pinned Fred Griffoul
2026-10-06 18:32   ` [PATCH 5/9] iommufd: Map memory provider files Fred Griffoul
2026-10-06 18:32   ` [PATCH 6/9] iommufd/selftest: Add mock-domain IOVA queries Fred Griffoul
2026-10-06 18:32   ` [PATCH 7/9] iommufd/selftest: Add a mock memory provider Fred Griffoul
2026-10-06 18:32   ` [PATCH 8/9] samples/kvm: Add a memory provider sample Fred Griffoul
2026-10-06 18:32   ` [PATCH 9/9] KVM: selftests: Test a memory provider shared by KVM and iommufd Fred Griffoul

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261006183235.16576-4-griffoul@gmail.com \
    --to=griffoul@gmail.com \
    --cc=ackerleytng@google.com \
    --cc=akpm@linux-foundation.org \
    --cc=bp@alien8.de \
    --cc=brauner@kernel.org \
    --cc=corbet@lwn.net \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@kernel.org \
    --cc=dwmw2@infradead.org \
    --cc=hpa@zytor.com \
    --cc=iommu@lists.linux.dev \
    --cc=jack@suse.cz \
    --cc=jgg@ziepe.ca \
    --cc=joey.gouly@arm.com \
    --cc=joro@8bytes.org \
    --cc=kevin.tian@intel.com \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=maz@kernel.org \
    --cc=mhocko@suse.com \
    --cc=mingo@redhat.com \
    --cc=oupton@kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=robin.murphy@arm.com \
    --cc=rppt@kernel.org \
    --cc=seanjc@google.com \
    --cc=seiden@linux.ibm.com \
    --cc=shuah@kernel.org \
    --cc=surenb@google.com \
    --cc=suzuki.poulose@arm.com \
    --cc=tglx@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=will@kernel.org \
    --cc=x86@kernel.org \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®