mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
@ 2026-09-26  0:50 Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs Ackerley Tng via B4 Relay
                   ` (18 more replies)
  0 siblings, 19 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

== Motivation

guest_memfd is KVM's guest-first memory provider that was created for
CoCo VMs and will support all VMs.

Today, guest_memfd only supports the equivalent of anonymous
PAGE_SIZE-d memfd memory, and has a lot to catch up to (HugeTLBfs,
tmpfs, DAX, etc) since almost anything can be mmap()-ed and handed to
KVM in a memslot.

I'd like to find a generic interface that would fit different
providers.

== Overview

Think of guest_memfd as a permissions (shared/private) and
KVM-know-how layer wrapping memory allocated from some provider.

This RFC illustrates how a resource pool fd can be handed to
guest_memfd to describe how it should allocate memory.

I believe this is extensible to many different kinds of memory
providers. :)

I'd like to discuss this at LPC 2026's KVM MC [1]. See y'all there!

== How userspace will use this feature

Do one of these,

  resource_fd = fsmount();
  resource_fd = open("/root/of/mount");

Then hand it to KVM_CREATE_GUEST_MEMFD,

  gmem_fd = ioctl(KVM_CREATE_GUEST_MEMFD,
                  GUEST_MEMFD_FLAG_USE_RESOURCE, resource_fd);

And then use gmem_fd as before.

== How to become a provider?

Implement this in the provider's struct super_operations,

  struct guest_memfd_provider_operations {
          void *(*attach)(struct file *resource_file);
          void (*release)(void *provider);
          struct folio *(*alloc_folio)(void *provider, pgoff_t index,
                                       struct mempolicy *mpol);
          void (*invalidate_folio)(void *provider, struct folio *folio);
  };

And the existence of these ops when guest_memfd accesses resource_fd
will determine if the provider is supported.

One requirement of this design is to allow new providers at runtime.

== Questions you might have

=== Why implement in struct super_operations?

It's one of the options to allow guest_memfd to find ops, and for an
RFC I think this requires the least new code and makes a simpler
illustration.

We would progressively add support for kernel-internal providers by
implementing guest_memfd_provider_operations, and loadable filesystem
modules can also implement the ops and be picked up at runtime.

Another option would be to have some kind of registry. Please see
changelog of patch 1 for discussion of options.

=== What if my provider doesn't deal with pages?

Frank and David Woodhouse [2] have use cases for providers that use
PFNs, and this is definitely something a generic interface must
support.

Here are some options I can think of:

1. Don't define .alloc_folio(), instead define .alloc_pfn()
2. Refactor .alloc_folio() to .alloc_pfn()

Either way, I think guest_memfd should still be the primary manager of
the memory.

=== Why have guest_memfd be the primary manager of the memory?

==== For folio-based providers

I propose we allocate the folios based on how the allocator was
configured:

+ e.g. for tmpfs,
    + Use mount page size.
    + Respect mount size limits.
    + Use guest_memfd mpol and fall back to mount mpol.
        + Gregory, now you can set mpol upfront on a mount! [3])
+ Use all the config that your provider has for allocation, reuse all
  those interfaces that your provider already has to do the config so
  guest_memfd doesn't have to reinvent those.

Once allocated, follow guest_memfd's rules on:

+ Tracking number of allocated bytes for this guest_memfd
+ memcg charging
+ Mapping into stage 2 page tables
+ Shared/private conversions
+ memory failure handling
+ etc

For the above split, I had the provider allocate, and then insert the
folio into guest_memfd's filemap.

I think of guest_memfd as the primary manager in the sense that
folio->mapping points to guest_memfd's filemap, which will allow
guest_memfd to handle memory failure by guest_memfd's rules.

Memory failure handling by guest_memfd's rules is something core to
guest_memfd. When guest_memfd supports huge pages, we'd like to be
able to split and only unmap the very page with a memory failure from
stage 2, which will keep guests running longer.

The provider can perform more handling of memory failure; we could
work out the ordering between guest_memfd and the provider.

Other than memory failure, there are other users that use
folio->mapping to determine what action to take. For now AFAICT
guest_memfd doesn't support those, but maybe it will need to in
future.

==== For PFN-based providers

For PFN-based providers, guest_memfd can be taught to store PFNs in
the xarray (like DAX does, perhaps?).

=== But my provider must be able to revoke memory from the guest!

I think that can be achieved with a bidirectional .attach(), where the
provider saves a pointer to the guest_memfd instance.

When memory must be revoked, it can tell the guest_memfd instance to
invalidate the page from everywhere. This has the nice benefit that
guest_memfd already respects the KVM MMU invalidation protocol, so the
provider doesn't have to learn the protocol.

== About this RFC

I chose tmpfs because I think it's the closest provider to the
existing anonymous PAGE_SIZE folios that guest_memfd supports.

The patches in this RFC was written with the help of AI and with many
edits from me.

This was built on top of Sean's kvm-x86 coco branch.

[1] https://lpc.events/event/20/contributions/2551/
[2] https://lore.kernel.org/all/20260720111259.122911-1-dwmw2@infradead.org/
[3] https://lore.kernel.org/all/20260902194657.79075-1-gourry@gourry.net/

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
Ackerley Tng (17):
      mm: shmem: Implement guest_memfd provider operations for tmpfs
      KVM: guest_memfd: Support provider folio allocation
      KVM: guest_memfd: Support provider folio invalidation
      KVM: guest_memfd: Add helper to attach resource provider file
      KVM: selftests: Add helper to create guest_memfd with a resource file
      KVM: selftests: Test negative validation of resource_fd argument
      KVM: selftests: Test rejection of unsupported filesystem for resource_fd
      KVM: selftests: Test rejection of tmpfs file for resource_fd
      KVM: selftests: Test rejection of swap tmpfs mounts
      KVM: selftests: Test rejection of hugepage tmpfs mounts
      KVM: selftests: Test guest_memfd with anonymous fsmount() tmpfs pool
      KVM: selftests: Test guest_memfd with mounted tmpfs root directory
      KVM: selftests: Test guest_memfd resource pool sharing across instances
      KVM: selftests: Test memory allocation against shared tmpfs resource pool
      KVM: selftests: Test that shared tmpfs resource pool size limit is respected
      KVM: selftests: Test guest execution with tmpfs-backed guest_memfd
      KVM: selftests: Document testing TODOs

 Documentation/virt/kvm/api.rst                 |  28 +-
 include/linux/fs/super_types.h                 |   4 +
 include/linux/guest_memfd.h                    |  29 ++
 include/linux/kvm_host.h                       |   2 +-
 include/uapi/linux/kvm.h                       |   5 +-
 mm/shmem.c                                     | 112 +++++++
 tools/include/uapi/linux/kvm.h                 |   5 +-
 tools/testing/selftests/kvm/guest_memfd_test.c | 404 +++++++++++++++++++++++--
 tools/testing/selftests/kvm/include/kvm_util.h |  23 +-
 virt/kvm/guest_memfd.c                         | 118 +++++++-
 10 files changed, 673 insertions(+), 57 deletions(-)
---
base-commit: 37e0791600e5f1e2837a270c885296a172f5825e
change-id: 20260924-gmem-tmpfs-backend-075e990294a1

Best regards,
--  
Ackerley Tng <ackerleytng@google.com>



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-10-01 13:07   ` David Woodhouse
  2026-09-26  0:50 ` [PATCH RFC 02/17] KVM: guest_memfd: Support provider folio allocation Ackerley Tng via B4 Relay
                   ` (17 subsequent siblings)
  18 siblings, 1 reply; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Implement guest_memfd provider operations to declare that tmpfs supports
providing memory to guest_memfd.

I'm implementing it directly in tmpfs as an illustration, one alternative I
can think of is that guest_memfd (KVM the module) could provide a registry,
supporting filesystems in the kernel, and loaded filesystems could register
themselves as providers.

Would it introduce ordering issues? Like if KVM were loaded after the
provider? Should the registry be built-in to the kernel?

Perhaps another way to key the provider functions could be TMPFS_MAGIC?

I put guest_memfd_provider_operations as a pointer in
super_operations. There might be a better place, to support different kinds
of providers. What do other providers need? Perhaps it's also okay to have
guest_memfd look up in a few different places, beginning with the resource
fd it was provided.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 include/linux/fs/super_types.h |   4 ++
 include/linux/guest_memfd.h    |  29 +++++++++++
 mm/shmem.c                     | 112 +++++++++++++++++++++++++++++++++++++++++
 3 files changed, 145 insertions(+)

diff --git a/include/linux/fs/super_types.h b/include/linux/fs/super_types.h
index ecd96aeb1cee7..29e502439648a 100644
--- a/include/linux/fs/super_types.h
+++ b/include/linux/fs/super_types.h
@@ -37,6 +37,7 @@ struct workqueue_struct;
 struct writeback_control;
 struct xattr_handler;
 struct fserror_event;
+struct guest_memfd_provider_operations;
 
 extern struct super_block *blockdev_superblock;
 
@@ -130,6 +131,9 @@ struct super_operations {
 
 	/* Report a filesystem error */
 	void (*report_error)(const struct fserror_event *event);
+#ifdef CONFIG_KVM_GUEST_MEMFD
+	const struct guest_memfd_provider_operations *gmem_provider_ops;
+#endif
 };
 
 struct super_block {
diff --git a/include/linux/guest_memfd.h b/include/linux/guest_memfd.h
new file mode 100644
index 0000000000000..60eb4f008c246
--- /dev/null
+++ b/include/linux/guest_memfd.h
@@ -0,0 +1,29 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+#ifndef _LINUX_GUEST_MEMFD_H
+#define _LINUX_GUEST_MEMFD_H
+
+#include <linux/types.h>
+
+struct file;
+struct folio;
+struct mempolicy;
+
+/**
+ * struct guest_memfd_provider_operations - Operations for external memory providers
+ * @attach: Attach to a resource file, perform filesystem-specific validation,
+ *          and return an opaque provider context.
+ * @release: Release provider context and unpin any resources.
+ * @alloc_folio: Allocate an uninserted folio for a given index and NUMA policy.
+ * @invalidate_folio: Notify provider that a folio has been invalidated.
+ *
+ * Used by filesystems and device drivers that provide memory for guest_memfd.
+ */
+struct guest_memfd_provider_operations {
+	void *(*attach)(struct file *resource_file);
+	void (*release)(void *provider);
+	struct folio *(*alloc_folio)(void *provider, pgoff_t index,
+				     struct mempolicy *mpol);
+	void (*invalidate_folio)(void *provider, struct folio *folio);
+};
+
+#endif /* _LINUX_GUEST_MEMFD_H */
diff --git a/mm/shmem.c b/mm/shmem.c
index 897fa2b61346f..279f29a9861c1 100644
--- a/mm/shmem.c
+++ b/mm/shmem.c
@@ -38,6 +38,7 @@
 #include <linux/uio.h>
 #include <linux/hugetlb.h>
 #include <linux/fs_parser.h>
+#include <linux/guest_memfd.h>
 #include <linux/swapfile.h>
 #include <linux/iversion.h>
 #include <linux/unicode.h>
@@ -5234,6 +5235,114 @@ static const struct inode_operations shmem_special_inode_operations = {
 #endif
 };
 
+#ifdef CONFIG_KVM_GUEST_MEMFD
+static void *shmem_gmem_attach(struct file *resource_file)
+{
+	struct inode *inode = file_inode(resource_file);
+	struct shmem_sb_info *sbinfo;
+
+	if (!S_ISDIR(inode->i_mode))
+		return ERR_PTR(-ENOTDIR);
+
+	/*
+	 * Would like comments: Requiring the root directory of the mount
+	 * provides a less confusing interface, although this is not strictly
+	 * necessary.
+	 */
+	if (resource_file->f_path.dentry != resource_file->f_path.mnt->mnt_root)
+		return ERR_PTR(-EINVAL);
+
+	if (__mnt_is_readonly(resource_file->f_path.mnt))
+		return ERR_PTR(-EROFS);
+
+	if (inode_permission(file_mnt_idmap(resource_file), inode,
+			     MAY_WRITE | MAY_EXEC))
+		return ERR_PTR(-EACCES);
+
+	/* Enforce noswap because guest_memfd does not support swapping. */
+	sbinfo = SHMEM_SB(inode->i_sb);
+	if (!sbinfo->noswap)
+		return ERR_PTR(-EINVAL);
+
+#ifdef CONFIG_TRANSPARENT_HUGEPAGE
+	/*
+	 * guest_memfd does not support huge pages yet; this restriction can be
+	 * relaxed in the future.
+	 *
+	 * TODO: Block rebinding with huge page support.
+	 */
+	if (sbinfo->huge != SHMEM_HUGE_NEVER)
+		return ERR_PTR(-EINVAL);
+#endif
+
+	return (void *)mntget(resource_file->f_path.mnt);
+}
+
+static void shmem_gmem_release(void *provider)
+{
+	struct vfsmount *mnt = provider;
+
+	mntput(mnt);
+}
+
+/*
+ * TODO: Consider refactoring shmem allocation helpers to share code between
+ * internal tmpfs folio allocation and guest_memfd provider allocation.
+ */
+static struct folio *shmem_gmem_alloc_folio(void *provider, pgoff_t index,
+					    struct mempolicy *mpol)
+{
+	struct mempolicy *sb_mpol = NULL;
+	struct vfsmount *mnt = provider;
+	struct shmem_sb_info *sbinfo;
+	struct folio *folio;
+
+	sbinfo = SHMEM_SB(mnt->mnt_sb);
+	if (sbinfo->max_blocks &&
+	    !percpu_counter_limited_add(&sbinfo->used_blocks,
+					sbinfo->max_blocks, 1))
+		return ERR_PTR(-ENOSPC);
+
+	if (!mpol) {
+		sb_mpol = shmem_get_sbmpol(sbinfo);
+		mpol = sb_mpol;
+	}
+
+	if (mpol)
+		folio = folio_alloc_mpol(GFP_HIGHUSER_MOVABLE, 0, mpol, index,
+					 numa_node_id());
+	else
+		folio = folio_alloc(GFP_HIGHUSER_MOVABLE, 0);
+
+	mpol_cond_put(sb_mpol);
+
+	if (!folio) {
+		if (sbinfo->max_blocks)
+			percpu_counter_sub(&sbinfo->used_blocks, 1);
+		return ERR_PTR(-ENOMEM);
+	}
+
+	return folio;
+}
+
+static void shmem_gmem_invalidate_folio(void *provider, struct folio *folio)
+{
+	struct vfsmount *mnt = provider;
+	struct shmem_sb_info *sbinfo;
+
+	sbinfo = SHMEM_SB(mnt->mnt_sb);
+	if (sbinfo->max_blocks)
+		percpu_counter_sub(&sbinfo->used_blocks, folio_nr_pages(folio));
+}
+
+static const struct guest_memfd_provider_operations shmem_gmem_provider_ops = {
+	.attach			= shmem_gmem_attach,
+	.release		= shmem_gmem_release,
+	.alloc_folio		= shmem_gmem_alloc_folio,
+	.invalidate_folio	= shmem_gmem_invalidate_folio,
+};
+#endif /* CONFIG_KVM_GUEST_MEMFD */
+
 static const struct super_operations shmem_ops = {
 	.alloc_inode	= shmem_alloc_inode,
 	.free_inode	= shmem_free_in_core_inode,
@@ -5252,6 +5361,9 @@ static const struct super_operations shmem_ops = {
 	.nr_cached_objects	= shmem_unused_huge_count,
 	.free_cached_objects	= shmem_unused_huge_scan,
 #endif
+#ifdef CONFIG_KVM_GUEST_MEMFD
+	.gmem_provider_ops	= &shmem_gmem_provider_ops,
+#endif
 };
 
 static const struct vm_operations_struct shmem_vm_ops = {

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 02/17] KVM: guest_memfd: Support provider folio allocation
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 03/17] KVM: guest_memfd: Support provider folio invalidation Ackerley Tng via B4 Relay
                   ` (16 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

If this guest_memfd was created with a provider, use the provider to
allocate a folio. For folio-based providers, using guest_memfd's filemap is
better because folio->mapping would be guest_memfd's mapping, and so
guest_memfd can handle anything that makes decisions based off
folio->mapping, like memory_failure().

Going along these lines of using gmem's filemap, passing gmem's filemap for
the provider to insert into seems awkward. At least for tmpfs, the
insertion function does a lot of stuff and assumes stuff about the filemap
being a tmpfs one (completely fair). I think it's better to use gmem's
filemap and do charging (memcg) and inode accounting according to
guest_memfd's rules though, hence insertion is done within gmem, in a gmem
filemap.

Going back to folio->mapping pointing to the gmem filemap, we could have
special memory failure handling for guest_memfd folios too, if there's some
kind of central registry of all PFNs belonging to guest_memfd? There are
other usages of folio->mapping, but gmem doesn't participate in those since
gmem doesn't do swap, etc now.

For non-folio-based providers, here's my suggestion: Don't provide
.alloc_folio(), provide some equivalent callback for pfns, track pfns in
gmem. The provider can force gmem to return the folios anytime, the
.attach() can be bidirectional.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 virt/kvm/guest_memfd.c | 44 ++++++++++++++++++++++++++++++++++++++++----
 1 file changed, 40 insertions(+), 4 deletions(-)

diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index a13445c26d9d6..5ac2d558c8dd8 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -3,6 +3,7 @@
 #include <linux/backing-dev.h>
 #include <linux/falloc.h>
 #include <linux/fs.h>
+#include <linux/guest_memfd.h>
 #include <linux/kvm_host.h>
 #include <linux/maple_tree.h>
 #include <linux/mempolicy.h>
@@ -35,6 +36,9 @@ struct gmem_inode {
 	struct inode vfs_inode;
 	struct list_head gmem_file_list;
 
+	void *provider;
+	const struct guest_memfd_provider_operations *provider_ops;
+
 	u64 flags;
 	/*
 	 * Every index in this inode, whether memory is populated or
@@ -50,6 +54,16 @@ static __always_inline struct gmem_inode *GMEM_I(struct inode *inode)
 	return container_of(inode, struct gmem_inode, vfs_inode);
 }
 
+static inline struct folio *gmem_provider_alloc_folio(struct gmem_inode *gi,
+						      pgoff_t index,
+						      struct mempolicy *mpol)
+{
+	if (!gi->provider_ops || !gi->provider_ops->alloc_folio)
+		return ERR_PTR(-EOPNOTSUPP);
+
+	return gi->provider_ops->alloc_folio(gi->provider, index, mpol);
+}
+
 #define kvm_gmem_for_each_file(f, inode) \
 	list_for_each_entry(f, &GMEM_I(inode)->gmem_file_list, entry)
 
@@ -128,6 +142,7 @@ static bool kvm_gmem_range_has_attributes(struct inode *inode,
  */
 static struct folio *kvm_gmem_get_folio(struct inode *inode, pgoff_t index)
 {
+	struct gmem_inode *gi = GMEM_I(inode);
 	/* TODO: Support huge pages. */
 	struct mempolicy *policy;
 	struct folio *folio;
@@ -140,10 +155,29 @@ static struct folio *kvm_gmem_get_folio(struct inode *inode, pgoff_t index)
 	if (!IS_ERR(folio))
 		return folio;
 
-	policy = mpol_shared_policy_lookup(&GMEM_I(inode)->policy, index);
-	folio = __filemap_get_folio_mpol(inode->i_mapping, index,
-					 FGP_LOCK | FGP_CREAT,
-					 mapping_gfp_mask(inode->i_mapping), policy);
+	policy = mpol_shared_policy_lookup(&gi->policy, index);
+	/*
+	 * TODO: Refactor the internal PAGE_SIZE gmem could be an internal
+	 * provider with internal provider_ops.
+	 */
+	if (gi->provider_ops) {
+		folio = gmem_provider_alloc_folio(gi, index, policy);
+		if (!IS_ERR(folio)) {
+			int r = filemap_add_folio(inode->i_mapping, folio,
+						  index, GFP_KERNEL);
+			if (r) {
+				folio_put(folio);
+				folio = ERR_PTR(r);
+			} else {
+				folio_mark_accessed(folio);
+			}
+		}
+	} else {
+		folio = __filemap_get_folio_mpol(inode->i_mapping, index,
+						 FGP_LOCK | FGP_CREAT,
+						 mapping_gfp_mask(inode->i_mapping),
+						 policy);
+	}
 	mpol_cond_put(policy);
 
 	/*
@@ -1334,6 +1368,8 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb)
 	mt_init_flags(&gi->attributes, MT_FLAGS_LOCK_EXTERN | MT_FLAGS_USE_RCU);
 
 	gi->flags = 0;
+	gi->provider = NULL;
+	gi->provider_ops = NULL;
 	INIT_LIST_HEAD(&gi->gmem_file_list);
 	return &gi->vfs_inode;
 }

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 03/17] KVM: guest_memfd: Support provider folio invalidation
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 02/17] KVM: guest_memfd: Support provider folio allocation Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 04/17] KVM: guest_memfd: Add helper to attach resource provider file Ackerley Tng via B4 Relay
                   ` (15 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

tmpfs increments the total number of used pages on the mount at allocation
time, so that number has to be decremented when the folio is removed from
tmpfs' ownership.

In this model where guest_memfd allocates from a tmpfs mount, I believe we
should respect the limits set in the mount, hence I require .alloc_folio()
to increment.

.invalidate_folio() is then the counterpart to .alloc_folio(), for undoing
any provider-specific per-folio work during allocations.

.invalidate_folio() requires mapping_set_release_always(). Not sure how odd
it is to use that mapping flag.

Using .free_folio() is hard because folio->mapping is NULL by then, so I
can't reach provider information. The provider information could also be
stuffed in folio->private? What do people think of that?

Another alternative would be to use a custom truncation function in
guest_memfd, just like shmem_truncate_range() or shmem_undo_range(). The
con is that guest_memfd might at some point support more generic mm stuff
like swap/reclaim, of some form? Or maybe never? Shall we go custom now and
unify later? Or just not think too far ahead?

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 virt/kvm/guest_memfd.c | 23 +++++++++++++++++++++++
 1 file changed, 23 insertions(+)

diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 5ac2d558c8dd8..00f3bd0cf21e6 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -64,6 +64,13 @@ static inline struct folio *gmem_provider_alloc_folio(struct gmem_inode *gi,
 	return gi->provider_ops->alloc_folio(gi->provider, index, mpol);
 }
 
+static inline void gmem_provider_invalidate_folio(struct gmem_inode *gi,
+						  struct folio *folio)
+{
+	if (gi->provider_ops && gi->provider_ops->invalidate_folio)
+		gi->provider_ops->invalidate_folio(gi->provider, folio);
+}
+
 #define kvm_gmem_for_each_file(f, inode) \
 	list_for_each_entry(f, &GMEM_I(inode)->gmem_file_list, entry)
 
@@ -166,6 +173,7 @@ static struct folio *kvm_gmem_get_folio(struct inode *inode, pgoff_t index)
 			int r = filemap_add_folio(inode->i_mapping, folio,
 						  index, GFP_KERNEL);
 			if (r) {
+				gmem_provider_invalidate_folio(gi, folio);
 				folio_put(folio);
 				folio = ERR_PTR(r);
 			} else {
@@ -893,10 +901,24 @@ static void kvm_gmem_free_folio(struct folio *folio)
 }
 #endif
 
+static void kvm_gmem_invalidate_folio(struct folio *folio, size_t offset,
+				      size_t len)
+{
+	struct inode *inode = folio->mapping->host;
+	struct gmem_inode *gi = GMEM_I(inode);
+
+	/* guest_memfd only allows truncating full folios. */
+	if (WARN_ON_ONCE(offset != 0 || len != folio_size(folio)))
+		return;
+
+	gmem_provider_invalidate_folio(gi, folio);
+}
+
 static const struct address_space_operations kvm_gmem_aops = {
 	.dirty_folio = noop_dirty_folio,
 	.migrate_folio	= kvm_gmem_migrate_folio,
 	.error_remove_folio = kvm_gmem_error_folio,
+	.invalidate_folio = kvm_gmem_invalidate_folio,
 #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_RECLAIM
 	.free_folio = kvm_gmem_free_folio,
 #endif
@@ -935,6 +957,7 @@ static int kvm_gmem_init_inode(struct inode *inode, loff_t size, u64 flags)
 	 */
 	mapping_set_inaccessible(inode->i_mapping);
 	WARN_ON_ONCE(!mapping_unevictable(inode->i_mapping));
+	mapping_set_release_always(inode->i_mapping);
 
 	gi->flags = flags;
 

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 04/17] KVM: guest_memfd: Add helper to attach resource provider file
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (2 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 03/17] KVM: guest_memfd: Support provider folio invalidation Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 05/17] KVM: selftests: Add helper to create guest_memfd with a resource file Ackerley Tng via B4 Relay
                   ` (14 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

A refcount on the resource_fd is taken only to keep the reference stable
for guest_memfd to figure out the existence of a provider. The provider is
responsible for pinning anything required to keep the provider's ops and
the actual provided resource alive.

.release() is called when this guest_memfd inode is destroyed to allow the
provider to clean up.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 Documentation/virt/kvm/api.rst | 28 ++++++++++++++---------
 include/linux/kvm_host.h       |  2 +-
 include/uapi/linux/kvm.h       |  5 ++++-
 virt/kvm/guest_memfd.c         | 51 ++++++++++++++++++++++++++++++++++++++++--
 4 files changed, 72 insertions(+), 14 deletions(-)

diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index 67f0f290797ab..b48df2827a2c3 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6470,7 +6470,9 @@ and cannot be resized  (guest_memfd files do however support PUNCH_HOLE).
   struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	__s32 resource_fd;
+	__u32 pad;
+	__u64 reserved[5];
   };
 
 Conceptually, the inode backing a guest_memfd file represents physical memory,
@@ -6492,15 +6494,21 @@ a single guest_memfd file, but the bound ranges must not overlap).
 The capability KVM_CAP_GUEST_MEMFD_FLAGS enumerates the `flags` that can be
 specified via KVM_CREATE_GUEST_MEMFD.  Currently defined flags:
 
-  ============================ ================================================
-  GUEST_MEMFD_FLAG_MMAP        Enable using mmap() on the guest_memfd file
-                               descriptor.
-  GUEST_MEMFD_FLAG_INIT_SHARED Make all memory in the file shared during
-                               KVM_CREATE_GUEST_MEMFD (memory files created
-                               without INIT_SHARED will be marked private).
-                               Shared memory can be faulted into host userspace
-                               page tables. Private memory cannot.
-  ============================ ================================================
+  ============================= ================================================
+  GUEST_MEMFD_FLAG_MMAP         Enable using mmap() on the guest_memfd file
+                                descriptor.
+  GUEST_MEMFD_FLAG_INIT_SHARED  Make all memory in the file shared during
+                                KVM_CREATE_GUEST_MEMFD (memory files created
+                                without INIT_SHARED will be marked private).
+                                Shared memory can be faulted into host userspace
+                                page tables. Private memory cannot.
+  GUEST_MEMFD_FLAG_USE_RESOURCE Consume resource_fd as an external memory
+                                provider or resource pool. When not set,
+                                resource_fd must be '0'. When set, the provider
+                                must be a supported resource (e.g. tmpfs mount
+                                root directory). Currently, tmpfs resource pools
+                                require noswap and huge=never mount options.
+  ============================= ================================================
 
 When the KVM MMU performs a PFN lookup to service a guest fault, the fault will
 always be consumed from guest_memfd, regardless of whether it is a shared or a
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 485f18454eb45..f4241c05650e9 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -2588,7 +2588,7 @@ bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm);
 
 static inline u64 kvm_gmem_get_supported_flags(struct kvm *kvm)
 {
-	u64 flags = GUEST_MEMFD_FLAG_MMAP;
+	u64 flags = GUEST_MEMFD_FLAG_MMAP | GUEST_MEMFD_FLAG_USE_RESOURCE;
 
 	if (!kvm || kvm_arch_supports_gmem_init_shared(kvm))
 		flags |= GUEST_MEMFD_FLAG_INIT_SHARED;
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index 8dff2fc1972e9..f4108bf7e3192 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1674,11 +1674,14 @@ struct kvm_memory_attributes2 {
 #define KVM_CREATE_GUEST_MEMFD	_IOWR(KVMIO,  0xd4, struct kvm_create_guest_memfd)
 #define GUEST_MEMFD_FLAG_MMAP		(1ULL << 0)
 #define GUEST_MEMFD_FLAG_INIT_SHARED	(1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_RESOURCE	(1ULL << 2)
 
 struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	__s32 resource_fd;
+	__u32 pad;
+	__u64 reserved[5];
 };
 
 #define KVM_PRE_FAULT_MEMORY	_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 00f3bd0cf21e6..0e0d7e93f4148 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -71,6 +71,12 @@ static inline void gmem_provider_invalidate_folio(struct gmem_inode *gi,
 		gi->provider_ops->invalidate_folio(gi->provider, folio);
 }
 
+static inline void gmem_provider_release(struct gmem_inode *gi)
+{
+	if (gi->provider_ops && gi->provider_ops->release)
+		gi->provider_ops->release(gi->provider);
+}
+
 #define kvm_gmem_for_each_file(f, inode) \
 	list_for_each_entry(f, &GMEM_I(inode)->gmem_file_list, entry)
 
@@ -984,7 +990,37 @@ static int kvm_gmem_init_inode(struct inode *inode, loff_t size, u64 flags)
 	return r;
 }
 
-static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
+static int kvm_gmem_attach_resource(struct inode *inode, int resource_fd)
+{
+	struct file *resource_file;
+	struct inode *res_inode;
+	const struct guest_memfd_provider_operations *ops = NULL;
+	void *provider;
+
+	resource_file = fget_raw(resource_fd);
+	if (!resource_file)
+		return -EBADF;
+
+	res_inode = file_inode(resource_file);
+	ops = res_inode->i_sb->s_op->gmem_provider_ops;
+
+	if (!ops || !ops->attach || !ops->alloc_folio) {
+		fput(resource_file);
+		return -EOPNOTSUPP;
+	}
+
+	provider = ops->attach(resource_file);
+	fput(resource_file);
+	if (IS_ERR(provider))
+		return PTR_ERR(provider);
+
+	GMEM_I(inode)->provider = provider;
+	GMEM_I(inode)->provider_ops = ops;
+
+	return 0;
+}
+
+static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags, int resource_fd)
 {
 	static const char *name = "[kvm-gmem]";
 	struct gmem_file *f;
@@ -1018,6 +1054,12 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
 	if (err)
 		goto err_inode;
 
+	if (flags & GUEST_MEMFD_FLAG_USE_RESOURCE) {
+		err = kvm_gmem_attach_resource(inode, resource_fd);
+		if (err)
+			goto err_inode;
+	}
+
 	file = alloc_file_pseudo(inode, kvm_gmem_mnt, name, O_RDWR, &kvm_gmem_fops);
 	if (IS_ERR(file)) {
 		err = PTR_ERR(file);
@@ -1057,7 +1099,10 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args)
 	if (size <= 0 || !PAGE_ALIGNED(size))
 		return -EINVAL;
 
-	return __kvm_gmem_create(kvm, size, flags);
+	if (!(flags & GUEST_MEMFD_FLAG_USE_RESOURCE) && args->resource_fd)
+		return -EINVAL;
+
+	return __kvm_gmem_create(kvm, size, flags, args->resource_fd);
 }
 
 int kvm_gmem_prepare_memory_region(struct kvm *kvm, struct kvm_memory_slot *slot,
@@ -1401,6 +1446,8 @@ static void kvm_gmem_destroy_inode(struct inode *inode)
 {
 	struct gmem_inode *gi = GMEM_I(inode);
 
+	gmem_provider_release(gi);
+
 	mpol_free_shared_policy(&gi->policy);
 
 	/*

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 05/17] KVM: selftests: Add helper to create guest_memfd with a resource file
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (3 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 04/17] KVM: guest_memfd: Add helper to attach resource provider file Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 06/17] KVM: selftests: Test negative validation of resource_fd argument Ackerley Tng via B4 Relay
                   ` (13 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Add a helper to create guest_memfd instances with a resource file
descriptor, and update the existing creation helper to pass a zero
resource file descriptor.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/include/uapi/linux/kvm.h                 |  5 ++++-
 tools/testing/selftests/kvm/include/kvm_util.h | 23 ++++++++++++++++++-----
 2 files changed, 22 insertions(+), 6 deletions(-)

diff --git a/tools/include/uapi/linux/kvm.h b/tools/include/uapi/linux/kvm.h
index 419011097fa8e..66c1f40b50c29 100644
--- a/tools/include/uapi/linux/kvm.h
+++ b/tools/include/uapi/linux/kvm.h
@@ -1654,11 +1654,14 @@ struct kvm_memory_attributes {
 #define KVM_CREATE_GUEST_MEMFD	_IOWR(KVMIO,  0xd4, struct kvm_create_guest_memfd)
 #define GUEST_MEMFD_FLAG_MMAP		(1ULL << 0)
 #define GUEST_MEMFD_FLAG_INIT_SHARED	(1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_RESOURCE	(1ULL << 2)
 
 struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	__s32 resource_fd;
+	__u32 pad;
+	__u64 reserved[5];
 };
 
 #define KVM_PRE_FAULT_MEMORY	_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/tools/testing/selftests/kvm/include/kvm_util.h b/tools/testing/selftests/kvm/include/kvm_util.h
index 777fa3dbf88d6..9eeaefe5a1839 100644
--- a/tools/testing/selftests/kvm/include/kvm_util.h
+++ b/tools/testing/selftests/kvm/include/kvm_util.h
@@ -767,26 +767,39 @@ static inline bool is_smt_on(void)
 
 void vm_create_irqchip(struct kvm_vm *vm);
 
-static inline int __vm_create_guest_memfd(struct kvm_vm *vm, u64 size,
-					  u64 flags)
+static inline int __vm_create_guest_memfd_resource(struct kvm_vm *vm, u64 size,
+						   u64 flags, int resource_fd)
 {
 	struct kvm_create_guest_memfd guest_memfd = {
 		.size = size,
 		.flags = flags,
+		.resource_fd = resource_fd,
 	};
 
 	return __vm_ioctl(vm, KVM_CREATE_GUEST_MEMFD, &guest_memfd);
 }
 
-static inline int vm_create_guest_memfd(struct kvm_vm *vm, u64 size,
-					u64 flags)
+static inline int __vm_create_guest_memfd(struct kvm_vm *vm, u64 size,
+					  u64 flags)
+{
+	return __vm_create_guest_memfd_resource(vm, size, flags, 0);
+}
+
+static inline int vm_create_guest_memfd_resource(struct kvm_vm *vm, u64 size,
+						 u64 flags, int resource_fd)
 {
-	int fd = __vm_create_guest_memfd(vm, size, flags);
+	int fd = __vm_create_guest_memfd_resource(vm, size, flags, resource_fd);
 
 	TEST_ASSERT(fd >= 0, KVM_IOCTL_ERROR(KVM_CREATE_GUEST_MEMFD, fd));
 	return fd;
 }
 
+static inline int vm_create_guest_memfd(struct kvm_vm *vm, u64 size,
+					u64 flags)
+{
+	return vm_create_guest_memfd_resource(vm, size, flags, 0);
+}
+
 void vm_set_user_memory_region(struct kvm_vm *vm, u32 slot, u32 flags,
 			       gpa_t gpa, u64 size, void *hva);
 int __vm_set_user_memory_region(struct kvm_vm *vm, u32 slot, u32 flags,

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 06/17] KVM: selftests: Test negative validation of resource_fd argument
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (4 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 05/17] KVM: selftests: Add helper to create guest_memfd with a resource file Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 07/17] KVM: selftests: Test rejection of unsupported filesystem for resource_fd Ackerley Tng via B4 Relay
                   ` (12 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Add negative test coverage verifying that creating a guest_memfd
instance with a non-zero resource_fd when GUEST_MEMFD_FLAG_USE_RESOURCE
is clear fails with -EINVAL, and creating with an invalid resource_fd
fails with -EBADF.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 2233d871a38f4..84b958b9ef715 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -404,6 +404,8 @@ static void test_guest_memfd_flags(struct kvm_vm *vm)
 	int fd;
 
 	for (flag = BIT(0); flag; flag <<= 1) {
+		if (flag == GUEST_MEMFD_FLAG_USE_RESOURCE)
+			continue;
 		fd = __vm_create_guest_memfd(vm, page_size, flag);
 		if (flag & valid_flags) {
 			TEST_ASSERT(fd >= 0,
@@ -418,6 +420,22 @@ static void test_guest_memfd_flags(struct kvm_vm *vm)
 	}
 }
 
+static void test_resource_fd_invalid(struct kvm_vm *vm)
+{
+	int fd;
+
+	/* Non-zero resource_fd without GUEST_MEMFD_FLAG_USE_RESOURCE fails */
+	fd = __vm_create_guest_memfd_resource(vm, page_size, 0, 1);
+	TEST_ASSERT(fd < 0, "guest_memfd with resource_fd but without flag should fail");
+	TEST_ASSERT_EQ(errno, EINVAL);
+
+	/* Bad file descriptor with GUEST_MEMFD_FLAG_USE_RESOURCE fails */
+	fd = __vm_create_guest_memfd_resource(vm, page_size,
+					      GUEST_MEMFD_FLAG_USE_RESOURCE, -1);
+	TEST_ASSERT(fd < 0, "guest_memfd with -1 resource_fd should fail");
+	TEST_ASSERT_EQ(errno, EBADF);
+}
+
 #define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)	\
 do {									\
 	int fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
@@ -476,6 +494,8 @@ static void test_guest_memfd(unsigned long vm_type)
 
 	test_guest_memfd_flags(vm);
 
+	test_resource_fd_invalid(vm);
+
 	__test_guest_memfd(vm, 0);
 
 	flags = vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS);

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 07/17] KVM: selftests: Test rejection of unsupported filesystem for resource_fd
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (5 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 06/17] KVM: selftests: Test negative validation of resource_fd argument Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 08/17] KVM: selftests: Test rejection of tmpfs file " Ackerley Tng via B4 Relay
                   ` (11 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Add test coverage verifying that passing a directory file descriptor
from an unsupported filesystem without provider operations as a resource
file fails with EOPNOTSUPP.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 22 ++++++++++++++++++++++
 1 file changed, 22 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 84b958b9ef715..31bc74c4c26bd 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -436,6 +436,27 @@ static void test_resource_fd_invalid(struct kvm_vm *vm)
 	TEST_ASSERT_EQ(errno, EBADF);
 }
 
+static void test_resource_fd_unsupported_fs(struct kvm_vm *vm)
+{
+	int fd, file_fd;
+
+	/*
+	 * Passing an FD from an unsupported filesystem without provider ops
+	 * (e.g. the /proc mount root directory) is rejected by KVM with
+	 * EOPNOTSUPP.
+	 */
+	file_fd = open("/proc", O_DIRECTORY | O_RDONLY);
+	TEST_ASSERT(file_fd >= 0, "open /proc failed");
+
+	fd = __vm_create_guest_memfd_resource(vm, page_size,
+					      GUEST_MEMFD_FLAG_USE_RESOURCE,
+					      file_fd);
+	TEST_ASSERT(fd < 0, "guest_memfd with /proc fd should fail");
+	TEST_ASSERT_EQ(errno, EOPNOTSUPP);
+
+	close(file_fd);
+}
+
 #define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)	\
 do {									\
 	int fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
@@ -495,6 +516,7 @@ static void test_guest_memfd(unsigned long vm_type)
 	test_guest_memfd_flags(vm);
 
 	test_resource_fd_invalid(vm);
+	test_resource_fd_unsupported_fs(vm);
 
 	__test_guest_memfd(vm, 0);
 

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 08/17] KVM: selftests: Test rejection of tmpfs file for resource_fd
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (6 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 07/17] KVM: selftests: Test rejection of unsupported filesystem for resource_fd Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 09/17] KVM: selftests: Test rejection of swap tmpfs mounts Ackerley Tng via B4 Relay
                   ` (10 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Add test coverage verifying that passing a regular file from tmpfs
as a resource file fails with ENOTDIR. Only directory file descriptors
representing the mount root are supported as resource pools.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 21 +++++++++++++++++++++
 1 file changed, 21 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 31bc74c4c26bd..a7a1b5470cb2c 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -457,6 +457,26 @@ static void test_resource_fd_unsupported_fs(struct kvm_vm *vm)
 	close(file_fd);
 }
 
+static void test_resource_fd_tmpfs_file(struct kvm_vm *vm)
+{
+	int fd, file_fd;
+
+	/*
+	 * Passing a regular file from tmpfs (e.g. via memfd) is rejected by
+	 * the provider attach callback with ENOTDIR because only directory
+	 * FDs representing the mount root are supported resource pools.
+	 */
+	file_fd = memfd_create("test_tmpfs_file", 0);
+	TEST_ASSERT(file_fd >= 0, "memfd_create failed");
+
+	fd = __vm_create_guest_memfd_resource(vm, page_size,
+					      GUEST_MEMFD_FLAG_USE_RESOURCE,
+					      file_fd);
+	TEST_ASSERT(fd < 0, "guest_memfd with tmpfs file fd should fail");
+	TEST_ASSERT_EQ(errno, ENOTDIR);
+	close(file_fd);
+}
+
 #define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)	\
 do {									\
 	int fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
@@ -517,6 +537,7 @@ static void test_guest_memfd(unsigned long vm_type)
 
 	test_resource_fd_invalid(vm);
 	test_resource_fd_unsupported_fs(vm);
+	test_resource_fd_tmpfs_file(vm);
 
 	__test_guest_memfd(vm, 0);
 

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 09/17] KVM: selftests: Test rejection of swap tmpfs mounts
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (7 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 08/17] KVM: selftests: Test rejection of tmpfs file " Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 10/17] KVM: selftests: Test rejection of hugepage " Ackerley Tng via B4 Relay
                   ` (9 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Ensure that guest_memfd creation fails with -EINVAL when passed a tmpfs
resource pool mount descriptor that has swap enabled, as required by
the tmpfs guest_memfd provider implementation.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 50 ++++++++++++++++++++++++++
 1 file changed, 50 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index a7a1b5470cb2c..5d017164d7b45 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -13,6 +13,7 @@
 
 #include <linux/bitmap.h>
 #include <linux/falloc.h>
+#include <linux/mount.h>
 #include <linux/sizes.h>
 #include <sys/types.h>
 #include <sys/stat.h>
@@ -477,6 +478,54 @@ static void test_resource_fd_tmpfs_file(struct kvm_vm *vm)
 	close(file_fd);
 }
 
+static int create_tmpfs_pool_fd(const char *huge, bool noswap, size_t size)
+{
+	char size_str[32];
+	int fs_fd, mnt_fd;
+
+	TEST_ASSERT(size && IS_ALIGNED(size, page_size),
+		    "Pool size '0x%zx' must be positive and page-aligned", size);
+
+	fs_fd = syscall(__NR_fsopen, "tmpfs", FSOPEN_CLOEXEC);
+	TEST_ASSERT(fs_fd >= 0, "fsopen failed");
+
+	if (noswap)
+		TEST_ASSERT(!syscall(__NR_fsconfig, fs_fd, FSCONFIG_SET_FLAG,
+				     "noswap", NULL, 0), "fsconfig noswap failed");
+
+	if (huge)
+		TEST_ASSERT(!syscall(__NR_fsconfig, fs_fd, FSCONFIG_SET_STRING,
+				     "huge", huge, 0), "fsconfig huge failed");
+
+	snprintf(size_str, sizeof(size_str), "%zu", size);
+	TEST_ASSERT(!syscall(__NR_fsconfig, fs_fd, FSCONFIG_SET_STRING,
+			     "size", size_str, 0), "fsconfig size failed");
+
+	TEST_ASSERT(!syscall(__NR_fsconfig, fs_fd, FSCONFIG_CMD_CREATE,
+			     NULL, NULL, 0), "fsconfig create failed");
+
+	mnt_fd = syscall(__NR_fsmount, fs_fd, FSMOUNT_CLOEXEC, 0);
+	TEST_ASSERT(mnt_fd > 0, "fsmount failed");
+
+	close(fs_fd);
+	return mnt_fd;
+}
+
+static void test_resource_fd_tmpfs_swap(struct kvm_vm *vm)
+{
+	int pool_fd, fd;
+
+	pool_fd = create_tmpfs_pool_fd("never", false, page_size);
+
+	fd = __vm_create_guest_memfd_resource(vm, page_size,
+					      GUEST_MEMFD_FLAG_USE_RESOURCE,
+					      pool_fd);
+	TEST_ASSERT(fd < 0, "guest_memfd with swap-enabled tmpfs should fail");
+	TEST_ASSERT_EQ(errno, EINVAL);
+
+	close(pool_fd);
+}
+
 #define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)	\
 do {									\
 	int fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
@@ -538,6 +587,7 @@ static void test_guest_memfd(unsigned long vm_type)
 	test_resource_fd_invalid(vm);
 	test_resource_fd_unsupported_fs(vm);
 	test_resource_fd_tmpfs_file(vm);
+	test_resource_fd_tmpfs_swap(vm);
 
 	__test_guest_memfd(vm, 0);
 

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 10/17] KVM: selftests: Test rejection of hugepage tmpfs mounts
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (8 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 09/17] KVM: selftests: Test rejection of swap tmpfs mounts Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 11/17] KVM: selftests: Test guest_memfd with anonymous fsmount() tmpfs pool Ackerley Tng via B4 Relay
                   ` (8 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Ensure that guest_memfd creation fails with -EINVAL when passed a tmpfs
resource pool mount descriptor that has huge pages enabled, as required
by the tmpfs guest_memfd provider implementation.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 16 ++++++++++++++++
 1 file changed, 16 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 5d017164d7b45..e8199ed17b0f8 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -526,6 +526,21 @@ static void test_resource_fd_tmpfs_swap(struct kvm_vm *vm)
 	close(pool_fd);
 }
 
+static void test_resource_fd_tmpfs_huge(struct kvm_vm *vm)
+{
+	int pool_fd, fd;
+
+	pool_fd = create_tmpfs_pool_fd("always", true, page_size);
+
+	fd = __vm_create_guest_memfd_resource(vm, page_size,
+					      GUEST_MEMFD_FLAG_USE_RESOURCE,
+					      pool_fd);
+	TEST_ASSERT(fd < 0, "guest_memfd with huge-enabled tmpfs should fail");
+	TEST_ASSERT_EQ(errno, EINVAL);
+
+	close(pool_fd);
+}
+
 #define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)	\
 do {									\
 	int fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
@@ -588,6 +603,7 @@ static void test_guest_memfd(unsigned long vm_type)
 	test_resource_fd_unsupported_fs(vm);
 	test_resource_fd_tmpfs_file(vm);
 	test_resource_fd_tmpfs_swap(vm);
+	test_resource_fd_tmpfs_huge(vm);
 
 	__test_guest_memfd(vm, 0);
 

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 11/17] KVM: selftests: Test guest_memfd with anonymous fsmount() tmpfs pool
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (9 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 10/17] KVM: selftests: Test rejection of hugepage " Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:50 ` [PATCH RFC 12/17] KVM: selftests: Test guest_memfd with mounted tmpfs root directory Ackerley Tng via B4 Relay
                   ` (7 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Add test coverage verifying that creating a guest_memfd backed by an
anonymous tmpfs mount descriptor returned by fsmount() succeeds and
reports the requested file size.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 35 ++++++++++++++++++++------
 1 file changed, 27 insertions(+), 8 deletions(-)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index e8199ed17b0f8..73212d5307049 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -541,12 +541,23 @@ static void test_resource_fd_tmpfs_huge(struct kvm_vm *vm)
 	close(pool_fd);
 }
 
-#define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)	\
-do {									\
-	int fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
-									\
-	test_##__test(args);						\
-	close(fd);							\
+#define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)		\
+do {										\
+	int pool_fd = -1;							\
+	int fd;									\
+										\
+	if ((__flags) & GUEST_MEMFD_FLAG_USE_RESOURCE) {			\
+		pool_fd = create_tmpfs_pool_fd("never", true, __gmem_size);	\
+		fd = vm_create_guest_memfd_resource(__vm, __gmem_size,		\
+						    __flags, pool_fd);		\
+	} else {								\
+		fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
+	}									\
+										\
+	test_##__test(args);							\
+	close(fd);								\
+	if (pool_fd >= 0)							\
+		close(pool_fd);							\
 } while (0)
 
 #define __gmem_test(__test, __vm, __flags, __gmem_size)			\
@@ -606,15 +617,23 @@ static void test_guest_memfd(unsigned long vm_type)
 	test_resource_fd_tmpfs_huge(vm);
 
 	__test_guest_memfd(vm, 0);
+	__test_guest_memfd(vm, GUEST_MEMFD_FLAG_USE_RESOURCE);
 
 	flags = vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS);
-	if (flags & GUEST_MEMFD_FLAG_MMAP)
+	if (flags & GUEST_MEMFD_FLAG_MMAP) {
 		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP);
+		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
+				       GUEST_MEMFD_FLAG_USE_RESOURCE);
+	}
 
 	/* MMAP should always be supported if INIT_SHARED is supported. */
-	if (flags & GUEST_MEMFD_FLAG_INIT_SHARED)
+	if (flags & GUEST_MEMFD_FLAG_INIT_SHARED) {
 		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
 				       GUEST_MEMFD_FLAG_INIT_SHARED);
+		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
+				       GUEST_MEMFD_FLAG_INIT_SHARED |
+				       GUEST_MEMFD_FLAG_USE_RESOURCE);
+	}
 
 	kvm_vm_free(vm);
 }

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 12/17] KVM: selftests: Test guest_memfd with mounted tmpfs root directory
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (10 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 11/17] KVM: selftests: Test guest_memfd with anonymous fsmount() tmpfs pool Ackerley Tng via B4 Relay
@ 2026-09-26  0:50 ` Ackerley Tng via B4 Relay
  2026-09-26  0:51 ` [PATCH RFC 13/17] KVM: selftests: Test guest_memfd resource pool sharing across instances Ackerley Tng via B4 Relay
                   ` (6 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:50 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Add test coverage verifying that opening the root directory of an
existing mounted tmpfs filesystem and passing that file descriptor as
the resource pool succeeds across all guest_memfd subtests and flag
combinations.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 137 +++++++++++++++++++------
 1 file changed, 103 insertions(+), 34 deletions(-)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 73212d5307049..94ef64474f2a0 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -10,11 +10,13 @@
 #include <errno.h>
 #include <stdio.h>
 #include <fcntl.h>
+#include <limits.h>
 
 #include <linux/bitmap.h>
 #include <linux/falloc.h>
 #include <linux/mount.h>
 #include <linux/sizes.h>
+#include <sys/mount.h>
 #include <sys/types.h>
 #include <sys/stat.h>
 
@@ -541,66 +543,122 @@ static void test_resource_fd_tmpfs_huge(struct kvm_vm *vm)
 	close(pool_fd);
 }
 
-#define ____gmem_test(__test, __vm, __flags, __gmem_size, args...)		\
+enum gmem_pool_type {
+	GMEM_POOL_NONE,
+	GMEM_POOL_FSMOUNT,
+	GMEM_POOL_MOUNTED_DIR,
+};
+
+struct gmem_pool {
+	int fd;
+	char path[PATH_MAX];
+	bool is_mounted;
+};
+
+static struct gmem_pool create_gmem_pool(size_t size, enum gmem_pool_type type)
+{
+	struct gmem_pool pool = { .fd = -1, .is_mounted = false };
+	int mnt_fd;
+
+	if (type == GMEM_POOL_NONE)
+		return pool;
+
+	mnt_fd = create_tmpfs_pool_fd("never", true, size);
+	TEST_REQUIRE(mnt_fd >= 0);
+
+	if (type == GMEM_POOL_FSMOUNT) {
+		pool.fd = mnt_fd;
+		return pool;
+	}
+
+	strcpy(pool.path, "/tmp/gmem_test_dir_XXXXXX");
+	TEST_ASSERT(mkdtemp(pool.path), "mkdtemp failed");
+
+	if (syscall(__NR_move_mount, mnt_fd, "", AT_FDCWD, pool.path,
+		    MOVE_MOUNT_F_EMPTY_PATH)) {
+		close(mnt_fd);
+		rmdir(pool.path);
+		TEST_REQUIRE(false);
+	}
+	close(mnt_fd);
+
+	pool.fd = open(pool.path, O_RDONLY | O_DIRECTORY);
+	TEST_ASSERT(pool.fd >= 0, "open mounted tmpfs root failed");
+	pool.is_mounted = true;
+
+	return pool;
+}
+
+static void destroy_gmem_pool(struct gmem_pool *pool)
+{
+	if (pool->fd >= 0)
+		close(pool->fd);
+	if (pool->is_mounted) {
+		TEST_ASSERT(!umount(pool->path), "umount failed");
+		TEST_ASSERT(!rmdir(pool->path), "rmdir failed");
+	}
+}
+
+#define ____gmem_test(__test, __vm, __flags, __pool_type, __gmem_size, args...)	\
 do {										\
-	int pool_fd = -1;							\
+	struct gmem_pool pool = { .fd = -1 };					\
 	int fd;									\
 										\
 	if ((__flags) & GUEST_MEMFD_FLAG_USE_RESOURCE) {			\
-		pool_fd = create_tmpfs_pool_fd("never", true, __gmem_size);	\
+		pool = create_gmem_pool(__gmem_size, __pool_type);		\
 		fd = vm_create_guest_memfd_resource(__vm, __gmem_size,		\
-						    __flags, pool_fd);		\
+						    __flags, pool.fd);		\
 	} else {								\
 		fd = vm_create_guest_memfd(__vm, __gmem_size, __flags);	\
 	}									\
 										\
 	test_##__test(args);							\
 	close(fd);								\
-	if (pool_fd >= 0)							\
-		close(pool_fd);							\
+	destroy_gmem_pool(&pool);						\
 } while (0)
 
-#define __gmem_test(__test, __vm, __flags, __gmem_size)			\
-	____gmem_test(__test, __vm, __flags, __gmem_size, fd, __gmem_size)
+#define __gmem_test(__test, __vm, __flags, __pool_type, __gmem_size)		\
+	____gmem_test(__test, __vm, __flags, __pool_type, __gmem_size, fd, __gmem_size)
 
-#define gmem_test(__test, __vm, __flags)				\
-	__gmem_test(__test, __vm, __flags, page_size * 4)
+#define gmem_test(__test, __vm, __flags, __pool_type)				\
+	__gmem_test(__test, __vm, __flags, __pool_type, page_size * 4)
 
-#define __gmem_test_vm(__test, __vm, __flags, __gmem_size)		\
-	____gmem_test(__test, __vm, __flags, __gmem_size, __vm, fd, __gmem_size)
+#define __gmem_test_vm(__test, __vm, __flags, __pool_type, __gmem_size)	\
+	____gmem_test(__test, __vm, __flags, __pool_type, __gmem_size, __vm, fd, __gmem_size)
 
-#define gmem_test_vm(__test, __vm, __flags)				\
-	__gmem_test_vm(__test, __vm, __flags, page_size * 4)
+#define gmem_test_vm(__test, __vm, __flags, __pool_type)			\
+	__gmem_test_vm(__test, __vm, __flags, __pool_type, page_size * 4)
 
-static void __test_guest_memfd(struct kvm_vm *vm, u64 flags)
+static void __test_guest_memfd(struct kvm_vm *vm, u64 flags,
+			       enum gmem_pool_type pool_type)
 {
 	test_create_guest_memfd_multiple(vm);
 	test_create_guest_memfd_invalid_sizes(vm, flags);
 
-	gmem_test(file_read_write, vm, flags);
+	gmem_test(file_read_write, vm, flags, pool_type);
 
 	if (flags & GUEST_MEMFD_FLAG_MMAP) {
 		if (flags & GUEST_MEMFD_FLAG_INIT_SHARED) {
 			size_t pmd_size = get_trans_hugepagesz();
 
-			gmem_test(mmap_supported, vm, flags);
-			gmem_test(fault_overflow, vm, flags);
-			gmem_test(numa_allocation, vm, flags);
-			__gmem_test(collapse, vm, flags, pmd_size);
+			gmem_test(mmap_supported, vm, flags, pool_type);
+			gmem_test(fault_overflow, vm, flags, pool_type);
+			gmem_test(numa_allocation, vm, flags, pool_type);
+			__gmem_test(collapse, vm, flags, pool_type, pmd_size);
 		} else {
-			gmem_test(fault_private, vm, flags);
+			gmem_test(fault_private, vm, flags, pool_type);
 		}
 
-		gmem_test(mmap_cow, vm, flags);
-		gmem_test(mbind, vm, flags);
+		gmem_test(mmap_cow, vm, flags, pool_type);
+		gmem_test(mbind, vm, flags, pool_type);
 	} else {
-		gmem_test(mmap_not_supported, vm, flags);
+		gmem_test(mmap_not_supported, vm, flags, pool_type);
 	}
 
-	gmem_test(file_size, vm, flags);
-	gmem_test(fallocate, vm, flags);
-	gmem_test(invalid_punch_hole, vm, flags);
-	gmem_test_vm(invalid_binding, vm, flags);
+	gmem_test(file_size, vm, flags, pool_type);
+	gmem_test(fallocate, vm, flags, pool_type);
+	gmem_test(invalid_punch_hole, vm, flags, pool_type);
+	gmem_test_vm(invalid_binding, vm, flags, pool_type);
 }
 
 static void test_guest_memfd(unsigned long vm_type)
@@ -616,23 +674,34 @@ static void test_guest_memfd(unsigned long vm_type)
 	test_resource_fd_tmpfs_swap(vm);
 	test_resource_fd_tmpfs_huge(vm);
 
-	__test_guest_memfd(vm, 0);
-	__test_guest_memfd(vm, GUEST_MEMFD_FLAG_USE_RESOURCE);
+	__test_guest_memfd(vm, 0, GMEM_POOL_NONE);
+	__test_guest_memfd(vm, GUEST_MEMFD_FLAG_USE_RESOURCE, GMEM_POOL_FSMOUNT);
+	__test_guest_memfd(vm, GUEST_MEMFD_FLAG_USE_RESOURCE, GMEM_POOL_MOUNTED_DIR);
 
 	flags = vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS);
 	if (flags & GUEST_MEMFD_FLAG_MMAP) {
-		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP);
+		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP, GMEM_POOL_NONE);
+		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
+				       GUEST_MEMFD_FLAG_USE_RESOURCE,
+				   GMEM_POOL_FSMOUNT);
 		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
-				       GUEST_MEMFD_FLAG_USE_RESOURCE);
+				       GUEST_MEMFD_FLAG_USE_RESOURCE,
+				   GMEM_POOL_MOUNTED_DIR);
 	}
 
 	/* MMAP should always be supported if INIT_SHARED is supported. */
 	if (flags & GUEST_MEMFD_FLAG_INIT_SHARED) {
 		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
-				       GUEST_MEMFD_FLAG_INIT_SHARED);
+				       GUEST_MEMFD_FLAG_INIT_SHARED,
+				   GMEM_POOL_NONE);
+		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
+				       GUEST_MEMFD_FLAG_INIT_SHARED |
+				       GUEST_MEMFD_FLAG_USE_RESOURCE,
+				   GMEM_POOL_FSMOUNT);
 		__test_guest_memfd(vm, GUEST_MEMFD_FLAG_MMAP |
 				       GUEST_MEMFD_FLAG_INIT_SHARED |
-				       GUEST_MEMFD_FLAG_USE_RESOURCE);
+				       GUEST_MEMFD_FLAG_USE_RESOURCE,
+				   GMEM_POOL_MOUNTED_DIR);
 	}
 
 	kvm_vm_free(vm);

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 13/17] KVM: selftests: Test guest_memfd resource pool sharing across instances
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (11 preceding siblings ...)
  2026-09-26  0:50 ` [PATCH RFC 12/17] KVM: selftests: Test guest_memfd with mounted tmpfs root directory Ackerley Tng via B4 Relay
@ 2026-09-26  0:51 ` Ackerley Tng via B4 Relay
  2026-09-26  0:51 ` [PATCH RFC 14/17] KVM: selftests: Test memory allocation against shared tmpfs resource pool Ackerley Tng via B4 Relay
                   ` (5 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:51 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Test that multiple guest_memfd instances can be created backed by the
same tmpfs resource pool descriptor, that each file descriptor reflects
its requested size, and that they have distinct inodes.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 36 ++++++++++++++++++++++++++
 1 file changed, 36 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 94ef64474f2a0..81311be059923 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -543,6 +543,41 @@ static void test_resource_fd_tmpfs_huge(struct kvm_vm *vm)
 	close(pool_fd);
 }
 
+static void test_resource_fd_tmpfs_shared_pool(struct kvm_vm *vm)
+{
+	int pool_fd, fd1, fd2;
+	struct stat st1, st2;
+
+	pool_fd = create_tmpfs_pool_fd("never", true, page_size * 2);
+	if (pool_fd < 0)
+		TEST_REQUIRE(false);
+
+	fd1 = vm_create_guest_memfd_resource(vm, page_size,
+					     GUEST_MEMFD_FLAG_USE_RESOURCE,
+					     pool_fd);
+
+	fd2 = vm_create_guest_memfd_resource(vm, page_size * 2,
+					     GUEST_MEMFD_FLAG_USE_RESOURCE,
+					     pool_fd);
+
+	TEST_ASSERT(!fstat(fd1, &st1), "fstat on fd1 should succeed");
+	TEST_ASSERT(st1.st_size == page_size,
+		    "fd1 st_size (%lu) should match requested size (%lu)",
+		    st1.st_size, page_size);
+
+	TEST_ASSERT(!fstat(fd2, &st2), "fstat on fd2 should succeed");
+	TEST_ASSERT(st2.st_size == page_size * 2,
+		    "fd2 st_size (%lu) should match requested size (%lu)",
+		    st2.st_size, page_size * 2);
+
+	TEST_ASSERT(st1.st_ino != st2.st_ino,
+		    "different guest_memfd instances should have distinct inodes");
+
+	close(fd2);
+	close(fd1);
+	close(pool_fd);
+}
+
 enum gmem_pool_type {
 	GMEM_POOL_NONE,
 	GMEM_POOL_FSMOUNT,
@@ -673,6 +708,7 @@ static void test_guest_memfd(unsigned long vm_type)
 	test_resource_fd_tmpfs_file(vm);
 	test_resource_fd_tmpfs_swap(vm);
 	test_resource_fd_tmpfs_huge(vm);
+	test_resource_fd_tmpfs_shared_pool(vm);
 
 	__test_guest_memfd(vm, 0, GMEM_POOL_NONE);
 	__test_guest_memfd(vm, GUEST_MEMFD_FLAG_USE_RESOURCE, GMEM_POOL_FSMOUNT);

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 14/17] KVM: selftests: Test memory allocation against shared tmpfs resource pool
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (12 preceding siblings ...)
  2026-09-26  0:51 ` [PATCH RFC 13/17] KVM: selftests: Test guest_memfd resource pool sharing across instances Ackerley Tng via B4 Relay
@ 2026-09-26  0:51 ` Ackerley Tng via B4 Relay
  2026-09-26  0:51 ` [PATCH RFC 15/17] KVM: selftests: Test that shared tmpfs resource pool size limit is respected Ackerley Tng via B4 Relay
                   ` (4 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:51 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Verify that allocating folios in multiple guest_memfd instances sharing
the same tmpfs resource pool consumes blocks from the underlying pool as
reported by filesystem statistics.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 16 +++++++++++++++-
 1 file changed, 15 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 81311be059923..4610eeab39be6 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -19,6 +19,7 @@
 #include <sys/mount.h>
 #include <sys/types.h>
 #include <sys/stat.h>
+#include <sys/statvfs.h>
 
 #include "kvm_syscalls.h"
 #include "kvm_util.h"
@@ -545,7 +546,8 @@ static void test_resource_fd_tmpfs_huge(struct kvm_vm *vm)
 
 static void test_resource_fd_tmpfs_shared_pool(struct kvm_vm *vm)
 {
-	int pool_fd, fd1, fd2;
+	int pool_fd, fd1, fd2, ret;
+	struct statvfs svfs;
 	struct stat st1, st2;
 
 	pool_fd = create_tmpfs_pool_fd("never", true, page_size * 2);
@@ -573,6 +575,18 @@ static void test_resource_fd_tmpfs_shared_pool(struct kvm_vm *vm)
 	TEST_ASSERT(st1.st_ino != st2.st_ino,
 		    "different guest_memfd instances should have distinct inodes");
 
+	ret = fallocate(fd1, FALLOC_FL_KEEP_SIZE, 0, page_size);
+	TEST_ASSERT(!ret, "fallocate on fd1 should succeed");
+
+	TEST_ASSERT(!fstatvfs(pool_fd, &svfs), "fstatvfs on pool_fd should succeed");
+	TEST_ASSERT_EQ(svfs.f_blocks - svfs.f_bfree, 1);
+
+	ret = fallocate(fd2, FALLOC_FL_KEEP_SIZE, 0, page_size);
+	TEST_ASSERT(!ret, "fallocate on fd2 should succeed");
+
+	TEST_ASSERT(!fstatvfs(pool_fd, &svfs), "fstatvfs on pool_fd should succeed");
+	TEST_ASSERT_EQ(svfs.f_blocks - svfs.f_bfree, 2);
+
 	close(fd2);
 	close(fd1);
 	close(pool_fd);

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 15/17] KVM: selftests: Test that shared tmpfs resource pool size limit is respected
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (13 preceding siblings ...)
  2026-09-26  0:51 ` [PATCH RFC 14/17] KVM: selftests: Test memory allocation against shared tmpfs resource pool Ackerley Tng via B4 Relay
@ 2026-09-26  0:51 ` Ackerley Tng via B4 Relay
  2026-09-26  0:51 ` [PATCH RFC 16/17] KVM: selftests: Test guest execution with tmpfs-backed guest_memfd Ackerley Tng via B4 Relay
                   ` (3 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:51 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Verify that allocating memory across guest_memfd instances sharing a
tmpfs resource pool fails with ENOSPC when the total allocation exceeds
the pool capacity.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 4 ++++
 1 file changed, 4 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 4610eeab39be6..345b41c912a18 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -587,6 +587,10 @@ static void test_resource_fd_tmpfs_shared_pool(struct kvm_vm *vm)
 	TEST_ASSERT(!fstatvfs(pool_fd, &svfs), "fstatvfs on pool_fd should succeed");
 	TEST_ASSERT_EQ(svfs.f_blocks - svfs.f_bfree, 2);
 
+	ret = fallocate(fd2, FALLOC_FL_KEEP_SIZE, page_size, page_size);
+	TEST_ASSERT(ret < 0, "fallocate exceeding pool capacity should fail");
+	TEST_ASSERT_EQ(errno, ENOSPC);
+
 	close(fd2);
 	close(fd1);
 	close(pool_fd);

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 16/17] KVM: selftests: Test guest execution with tmpfs-backed guest_memfd
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (14 preceding siblings ...)
  2026-09-26  0:51 ` [PATCH RFC 15/17] KVM: selftests: Test that shared tmpfs resource pool size limit is respected Ackerley Tng via B4 Relay
@ 2026-09-26  0:51 ` Ackerley Tng via B4 Relay
  2026-09-26  0:51 ` [PATCH RFC 17/17] KVM: selftests: Document testing TODOs Ackerley Tng via B4 Relay
                   ` (2 subsequent siblings)
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:51 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Test that a guest vCPU can execute instructions and read/write memory
mapped into a guest_memfd instance that is backed by a tmpfs resource
pool descriptor.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 57 ++++++++++++++++++++++++++
 1 file changed, 57 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 345b41c912a18..bf369874f2647 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -829,6 +829,62 @@ static void test_guest_memfd_guest(void)
 	kvm_vm_free(vm);
 }
 
+static void test_guest_memfd_guest_resource(void)
+{
+	const gpa_t gpa = SZ_4G;
+	const int slot = 1;
+	struct kvm_vcpu *vcpu;
+	struct kvm_vm *vm;
+	int pool_fd, fd, i;
+	size_t size;
+	u8 *mem;
+
+	if (!kvm_check_cap(KVM_CAP_GUEST_MEMFD_FLAGS))
+		return;
+
+	pool_fd = create_tmpfs_pool_fd("never", true, page_size);
+	if (pool_fd < 0)
+		TEST_REQUIRE(false);
+
+	vm = __vm_create_shape_with_one_vcpu(VM_SHAPE_DEFAULT, &vcpu, 1, guest_code);
+
+	TEST_ASSERT(vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS) & GUEST_MEMFD_FLAG_MMAP,
+		    "Default VM type should support MMAP, supported flags = 0x%x",
+		    vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS));
+	TEST_ASSERT(vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS) & GUEST_MEMFD_FLAG_INIT_SHARED,
+		    "Default VM type should support INIT_SHARED, supported flags = 0x%x",
+		    vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS));
+
+	size = max_t(size_t, vm->page_size, page_size);
+	fd = __vm_create_guest_memfd_resource(vm, size,
+					      GUEST_MEMFD_FLAG_MMAP |
+					      GUEST_MEMFD_FLAG_INIT_SHARED |
+					      GUEST_MEMFD_FLAG_USE_RESOURCE,
+					      pool_fd);
+	TEST_ASSERT(fd >= 0, "guest_memfd with tmpfs pool should succeed");
+
+	vm_set_user_memory_region2(vm, slot, KVM_MEM_GUEST_MEMFD, gpa, size, NULL, fd, 0);
+
+	mem = kvm_mmap(size, PROT_READ | PROT_WRITE, MAP_SHARED, fd);
+	memset(mem, 0xaa, size);
+	kvm_munmap(mem, size);
+
+	virt_map(vm, gpa, gpa, size / vm->page_size);
+	vcpu_args_set(vcpu, 2, gpa, size);
+	vcpu_run(vcpu);
+
+	TEST_ASSERT_EQ(get_ucall(vcpu, NULL), UCALL_DONE);
+
+	mem = kvm_mmap(size, PROT_READ | PROT_WRITE, MAP_SHARED, fd);
+	for (i = 0; i < size; i++)
+		TEST_ASSERT_EQ(mem[i], 0xff);
+
+	close(fd);
+	close(pool_fd);
+	kvm_vm_free(vm);
+}
+
+
 int main(int argc, char *argv[])
 {
 	unsigned long vm_types, vm_type;
@@ -849,4 +905,5 @@ int main(int argc, char *argv[])
 		test_guest_memfd(vm_type);
 
 	test_guest_memfd_guest();
+	test_guest_memfd_guest_resource();
 }

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* [PATCH RFC 17/17] KVM: selftests: Document testing TODOs
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (15 preceding siblings ...)
  2026-09-26  0:51 ` [PATCH RFC 16/17] KVM: selftests: Test guest execution with tmpfs-backed guest_memfd Ackerley Tng via B4 Relay
@ 2026-09-26  0:51 ` Ackerley Tng via B4 Relay
  2026-09-28 16:32 ` [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd David Woodhouse
  2026-10-01 13:43 ` Gregory Price
  18 siblings, 0 replies; 28+ messages in thread
From: Ackerley Tng via B4 Relay @ 2026-09-26  0:51 UTC (permalink / raw)
  To: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	Gregory Price, David Woodhouse, yan.y.zhao, michael.roth,
	suzuki.poulose, Christian Brauner, Jason Gunthorpe, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest, Ackerley Tng

From: Ackerley Tng <ackerleytng@google.com>

Remind myself about all the tests (A)I did not write.

Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
 tools/testing/selftests/kvm/guest_memfd_test.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index bf369874f2647..4e5d00df899a6 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -721,6 +721,16 @@ static void test_guest_memfd(unsigned long vm_type)
 
 	test_guest_memfd_flags(vm);
 
+	/*
+	 * TODO: Test that guest_memfd creation rejects a read-only tmpfs mount
+	 * (e.g. MOUNT_ATTR_RDONLY) with EROFS.
+	 *
+	 * TODO: Test that guest_memfd creation rejects an ID-mapped tmpfs mount
+	 * (e.g. MOUNT_ATTR_IDMAP) lacking write or execute/search permissions
+	 * with EACCES.
+	 *
+	 * TODO: Test mpol fallback to mount config.
+	 */
 	test_resource_fd_invalid(vm);
 	test_resource_fd_unsupported_fs(vm);
 	test_resource_fd_tmpfs_file(vm);

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (16 preceding siblings ...)
  2026-09-26  0:51 ` [PATCH RFC 17/17] KVM: selftests: Document testing TODOs Ackerley Tng via B4 Relay
@ 2026-09-28 16:32 ` David Woodhouse
  2026-09-28 22:59   ` Ackerley Tng
  2026-10-01 13:43 ` Gregory Price
  18 siblings, 1 reply; 28+ messages in thread
From: David Woodhouse @ 2026-09-28 16:32 UTC (permalink / raw)
  To: ackerleytng, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Jason Gunthorpe,
	Nicolin Chen, Xu Yilun, aik, aneesh.kumar, Vlastimil Babka,
	Connor Williamson, Fred Griffoul
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest

[-- Attachment #1: Type: text/plain, Size: 1336 bytes --]

On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
> 
> === What if my provider doesn't deal with pages?
> 
> Frank and David Woodhouse [2] have use cases for providers that use
> PFNs, and this is definitely something a generic interface must
> support.
> 
> Here are some options I can think of:
> 
> 1. Don't define .alloc_folio(), instead define .alloc_pfn()
> 2. Refactor .alloc_folio() to .alloc_pfn()
> 
> Either way, I think guest_memfd should still be the primary manager of
> the memory.

Thanks for working on this.

I'm not sure what you mean by 'primary manager' here but my use case is
for the provider to be the ultimate arbiter of who owns which PFN, and
revocation of the same. Think of it like a filesystem backed by
external memory. Each guest's memory is a file, and a guest can
*donate* specific pages of its memory to other guests (which in the
context of the *interface* only means that we need revocation).

Should I be trying to rework parts of my series to live on top of this?
Things I have working but which I don't see here include:

 • PFN support (which you mentioned).
 • IOMMUFD support via dma-buf.
 • Revocation (which I've tested correctly breaks down large pages in
   both EPT/NPT and IOMMU).
 • AsyncPF support for pages requested by KVM.

[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
  2026-09-28 16:32 ` [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd David Woodhouse
@ 2026-09-28 22:59   ` Ackerley Tng
  2026-09-29 16:17     ` David Woodhouse
  0 siblings, 1 reply; 28+ messages in thread
From: Ackerley Tng @ 2026-09-28 22:59 UTC (permalink / raw)
  To: David Woodhouse, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Jason Gunthorpe,
	Nicolin Chen, Xu Yilun, aik, aneesh.kumar, Vlastimil Babka,
	Connor Williamson, Fred Griffoul
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest

David Woodhouse <dwmw2@infradead.org> writes:

> On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
>>
>> === What if my provider doesn't deal with pages?
>>
>> Frank and David Woodhouse [2] have use cases for providers that use
>> PFNs, and this is definitely something a generic interface must
>> support.
>>
>> Here are some options I can think of:
>>
>> 1. Don't define .alloc_folio(), instead define .alloc_pfn()
>> 2. Refactor .alloc_folio() to .alloc_pfn()
>>
>> Either way, I think guest_memfd should still be the primary manager of
>> the memory.
>
> Thanks for working on this.
>

Thanks for the quick reply!

> I'm not sure what you mean by 'primary manager' here but my use case is

Naming is hard! See comment on revocation below.

> for the provider to be the ultimate arbiter of who owns which PFN, and
> revocation of the same. Think of it like a filesystem backed by
> external memory. Each guest's memory is a file, and a guest can
> *donate* specific pages of its memory to other guests (which in the

Putting donation aside first - I don't have a good picture of how a
guest would tell the host that it's willing to share a page - can't
comment. Would love to find out more.

> context of the *interface* only means that we need revocation).
>

Going with this revocation part, beginning with a more basic use case:
I'm thinking that the provider can notify guest_memfd that the page or
PFN is going away.

guest_memfd cannot say no, but it does have to tell KVM to zap the page
from stage 2, and for CoCo it may need to do more stuff.

The notification needs some pointer, so I was thinking that on .attach()
the provider would also note down which guest_memfd instance the
provider should notify.

The notifier could tell guest_memfd that some (offset, size) is going
away, and breaking down mappings in the EPT/NPT should be something
guest_memfd tells KVM to do, I think. Are you thinking that the provider
would directly break the mappings down in the EPT/NPT?

Does this notification mechanism work for you? IIUC regular DAX would
need it too, like if the device is unplugged for example.

> Should I be trying to rework parts of my series to live on top of this?

I hope to gather more feedback at LPC. It'd be nice to get early
feedback if you think it can work for you! The overall design still
needs feedback though, don't take this as the final direction.

> Things I have working but which I don't see here include:
>

This sounds like 4 or more different features! I think they do work with
this resource fd proposal.

>  • PFN support (which you mentioned).

Yup, guest_memfd needs some way to track PFNs in addition to folios. I
think for tracking we can reference DAX, the part I haven't looked at is
how to let guest_memfd handle events. Like if there's memory failure on
a PFN, how do we let guest_memfd handle it?

guest_memfd needs to at least be notified, so that it can handle the KVM
stuff: zapping and special stuff for CoCo. Whether the provider or
guest_memfd gets to handle it first can be worked out :)

>  • IOMMUFD support via dma-buf.

Is this about having guest_memfd work with dmabuf fds?

Fuad also brought up that Android uses dmabuf as an interface for CMA
memory. [3] IIUC the dmabuf fd is mmap()ed, then the userspace address
is set up in a memslot for the guest. To make this CMA memory
CoCo-friendly, we want to use it with guest_memfd.

Jason response was that guest_memfd should probably just directly get
memory from CMA [3].

I'm not super familiar with CMA or dmabufs, would either have a suitable
"resource fd" to pass to guest_memfd like how HugeTLBfs has a mount fd?
The "resource fd" should ideally have no way to mmap or read/write the
memory directly.

[3] https://lore.kernel.org/all/CA+EHjTxZ0N3Tfnid404B4tkb_E+Z8mODTHTgBPiF6=bwZp7Hjw@mail.gmail.com/

Or is this about letting IOMMUFD get pages to map in IOMMU page tables
via an fd+offset instead of via mmap()-ed userspace addresses as in
IOMMU_IOAS_MAP_FILE, and that guest_memfd isn't a supported kind of fd
(yet)?

I think guest_memfd should support IOMMU_IOAS_MAP_FILE. One use case is
Confidential IO [4], where SNP would need to have private memory
(guest_memfd) mapped in the IO page tables.

The missing parts here are that guest_memfd and iommufd need to
cooperate for conversions (or some other coordination in userspace). If
private memory gets converted to shared, iommufd needs to know to remap
the pages as shared in the IO page tables.

[4] https://lore.kernel.org/all/20260225075211.3353194-1-aik@amd.com/

>  • Revocation (which I've tested correctly breaks down large pages in
>    both EPT/NPT and IOMMU).

See above.

>  • AsyncPF support for pages requested by KVM.

Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness.
KVM MMU checks that the page is not there, then since async was
configured, KVM tells the guest to schedule in something else.

guest_memfd doesn't have support for async #PFs now. I imagine it would
be something like KVM MMU firing off a kvm_gmem_get_pfn() request
asynchronously, and then guest_memfd gets the page or PFN and inserts it
in guest_memfd filemap, and then notifies KVM MMU?

Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from
the provider.

I think it's orthogonal because guest_memfd could
have gotten a PAGE_SIZE page (existing functionality) or gotten a page
from the provider.

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
  2026-09-28 22:59   ` Ackerley Tng
@ 2026-09-29 16:17     ` David Woodhouse
  2026-09-29 23:41       ` Ackerley Tng
  0 siblings, 1 reply; 28+ messages in thread
From: David Woodhouse @ 2026-09-29 16:17 UTC (permalink / raw)
  To: Ackerley Tng, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Jason Gunthorpe,
	Nicolin Chen, Xu Yilun, aik, aneesh.kumar, Vlastimil Babka,
	Connor Williamson, Fred Griffoul
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest

[-- Attachment #1: Type: text/plain, Size: 7854 bytes --]

On Mon, 2026-09-28 at 15:59 -0700, Ackerley Tng wrote:
> David Woodhouse <dwmw2@infradead.org> writes:
> 
> > On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
> > > 
> > > === What if my provider doesn't deal with pages?
> > > 
> > > Frank and David Woodhouse [2] have use cases for providers that use
> > > PFNs, and this is definitely something a generic interface must
> > > support.
> > > 
> > > Here are some options I can think of:
> > > 
> > > 1. Don't define .alloc_folio(), instead define .alloc_pfn()
> > > 2. Refactor .alloc_folio() to .alloc_pfn()
> > > 
> > > Either way, I think guest_memfd should still be the primary manager of
> > > the memory.
> > 
> > Thanks for working on this.
> > 
> 
> Thanks for the quick reply!
> 
> > I'm not sure what you mean by 'primary manager' here but my use case is
> 
> Naming is hard! See comment on revocation below.
> 
> > for the provider to be the ultimate arbiter of who owns which PFN, and
> > revocation of the same. Think of it like a filesystem backed by
> > external memory. Each guest's memory is a file, and a guest can
> > *donate* specific pages of its memory to other guests (which in the
> 
> Putting donation aside first - I don't have a good picture of how a
> guest would tell the host that it's willing to share a page - can't
> comment. Would love to find out more.
>
> > context of the *interface* only means that we need revocation).

That's implementation-specific. Think of things like Nitro Enclaves
where a guest 'donates' its memory back to the hypervisor to be used by
another microvm. But that interface isn't the guest_memfd concern; as I
said, in the context of the guest_memfd interface, there is only
revocation: that page went away, whatever the reason.

> Going with this revocation part, beginning with a more basic use case:
> I'm thinking that the provider can notify guest_memfd that the page or
> PFN is going away.

Right. That part I have working in what I was posting.
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf06

> guest_memfd cannot say no, but it does have to tell KVM to zap the page
> from stage 2, and for CoCo it may need to do more stuff.
>
> The notification needs some pointer, so I was thinking that on .attach()
> the provider would also note down which guest_memfd instance the
> provider should notify.
>
> The notifier could tell guest_memfd that some (offset, size) is going
> away, and breaking down mappings in the EPT/NPT should be something
> guest_memfd tells KVM to do, I think. Are you thinking that the provider
> would directly break the mappings down in the EPT/NPT?

I think telling KVM is the right thing to do.

> Does this notification mechanism work for you? IIUC regular DAX would
> need it too, like if the device is unplugged for example.

Sure. As long as it works and efficiently covers everyone's use cases,
I'm not going to be overly opinionated a priori about *how* it works.

All such bets are off once we see the code, of course :)

> > Should I be trying to rework parts of my series to live on top of this?
> 
> I hope to gather more feedback at LPC. It'd be nice to get early
> feedback if you think it can work for you! The overall design still
> needs feedback though, don't take this as the final direction.
> 
> > Things I have working but which I don't see here include:
> > 
> 
> This sounds like 4 or more different features! I think they do work with
> this resource fd proposal.
> 
> >  • PFN support (which you mentioned).
> 
> Yup, guest_memfd needs some way to track PFNs in addition to folios. I
> think for tracking we can reference DAX, the part I haven't looked at is
> how to let guest_memfd handle events. Like if there's memory failure on
> a PFN, how do we let guest_memfd handle it?

Telling the guest about correctable and uncorrectable errors, you mean?
I would have thought that's up to the VMM, until the point where (see
revocation)?

> guest_memfd needs to at least be notified, so that it can handle the KVM
> stuff: zapping and special stuff for CoCo. Whether the provider or
> guest_memfd gets to handle it first can be worked out :)
> 
> >  • IOMMUFD support via dma-buf.
> 
> Is this about having guest_memfd work with dmabuf fds?

It's about having the provider of memory (again, think of it as a daxfs
if you will) able to tell *both* KVM and IOMMU which pages are where —
and revoke them at will.

The *implementation* involved dmabuf fds, and we already discussed
which wraps which, on the first posting of my series. But maybe we'll
conclude that the right answer is for IOMMUFD to recognise this
guest_memfd 'provider' as a first-class citizen, and we won't need to
masquerade as dmabuf?

> Fuad also brought up that Android uses dmabuf as an interface for CMA
> memory. [3] IIUC the dmabuf fd is mmap()ed, then the userspace address
> is set up in a memslot for the guest. To make this CMA memory
> CoCo-friendly, we want to use it with guest_memfd.
> 
> Jason response was that guest_memfd should probably just directly get
> memory from CMA [3].
> 
> I'm not super familiar with CMA or dmabufs, would either have a suitable
> "resource fd" to pass to guest_memfd like how HugeTLBfs has a mount fd?
> The "resource fd" should ideally have no way to mmap or read/write the
> memory directly.
> 
> [3] https://lore.kernel.org/all/CA+EHjTxZ0N3Tfnid404B4tkb_E+Z8mODTHTgBPiF6=bwZp7Hjw@mail.gmail.com/

I think where the memory comes *from* is an implementation detail for
the provider. It provides a PFN (or folio, if you must). All else is
not the business of KVM or IOMMUFD.

> Or is this about letting IOMMUFD get pages to map in IOMMU page tables
> via an fd+offset instead of via mmap()-ed userspace addresses as in
> IOMMU_IOAS_MAP_FILE, and that guest_memfd isn't a supported kind of fd
> (yet)?
> 
> I think guest_memfd should support IOMMU_IOAS_MAP_FILE. One use case is
> Confidential IO [4], where SNP would need to have private memory
> (guest_memfd) mapped in the IO page tables.

Right, I think that's what I meant with 'first class citizen' above?

> The missing parts here are that guest_memfd and iommufd need to
> cooperate for conversions (or some other coordination in userspace). If
> private memory gets converted to shared, iommufd needs to know to remap
> the pages as shared in the IO page tables.
> 
> [4] https://lore.kernel.org/all/20260225075211.3353194-1-aik@amd.com/
> 
> >  • Revocation (which I've tested correctly breaks down large pages in
> >    both EPT/NPT and IOMMU).
> 
> See above.
> 
> >  • AsyncPF support for pages requested by KVM.
> 
> Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness.
> KVM MMU checks that the page is not there, then since async was
> configured, KVM tells the guest to schedule in something else.
> 
> guest_memfd doesn't have support for async #PFs now. I imagine it would
> be something like KVM MMU firing off a kvm_gmem_get_pfn() request
> asynchronously, and then guest_memfd gets the page or PFN and inserts it
> in guest_memfd filemap, and then notifies KVM MMU?
> 
> Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from
> the provider.
> 
> I think it's orthogonal because guest_memfd could
> have gotten a PAGE_SIZE page (existing functionality) or gotten a page
> from the provider.

Kind of, but I'm focusing on the guest_memfd provider interface, and
that needs to support the asychronous mode: asked for a PFN for a given
guest address, it returns -EAGAIN and then provides it later, and the
guest gets the right asyncpf behaviour:
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34f

[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
  2026-09-29 16:17     ` David Woodhouse
@ 2026-09-29 23:41       ` Ackerley Tng
  2026-09-30  0:08         ` David Woodhouse
  0 siblings, 1 reply; 28+ messages in thread
From: Ackerley Tng @ 2026-09-29 23:41 UTC (permalink / raw)
  To: David Woodhouse, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Jason Gunthorpe,
	Nicolin Chen, Xu Yilun, aik, aneesh.kumar, Vlastimil Babka,
	Connor Williamson, Fred Griffoul
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest

David Woodhouse <dwmw2@infradead.org> writes:

> On Mon, 2026-09-28 at 15:59 -0700, Ackerley Tng wrote:
>> David Woodhouse <dwmw2@infradead.org> writes:
>>
>> > On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
>> > >
>> > > === What if my provider doesn't deal with pages?
>> > >
>> > > Frank and David Woodhouse [2] have use cases for providers that use
>> > > PFNs, and this is definitely something a generic interface must
>> > > support.
>> > >
>> > > Here are some options I can think of:
>> > >
>> > > 1. Don't define .alloc_folio(), instead define .alloc_pfn()
>> > > 2. Refactor .alloc_folio() to .alloc_pfn()
>> > >
>> > > Either way, I think guest_memfd should still be the primary manager of
>> > > the memory.
>> >
>> > Thanks for working on this.
>> >
>>
>> Thanks for the quick reply!
>>
>> > I'm not sure what you mean by 'primary manager' here but my use case is
>>
>> Naming is hard! See comment on revocation below.
>>
>> > for the provider to be the ultimate arbiter of who owns which PFN, and
>> > revocation of the same. Think of it like a filesystem backed by
>> > external memory. Each guest's memory is a file, and a guest can
>> > *donate* specific pages of its memory to other guests (which in the
>>
>> Putting donation aside first - I don't have a good picture of how a
>> guest would tell the host that it's willing to share a page - can't
>> comment. Would love to find out more.
>>
>> > context of the *interface* only means that we need revocation).
>
> That's implementation-specific. Think of things like Nitro Enclaves
> where a guest 'donates' its memory back to the hypervisor to be used by
> another microvm. But that interface isn't the guest_memfd concern; as I
> said, in the context of the guest_memfd interface, there is only
> revocation: that page went away, whatever the reason.
>

Okay we're on the same page then. :) guest_memfd only needs to know that
the page went away, guest_memfd cannot say no.

>> Going with this revocation part, beginning with a more basic use case:
>> I'm thinking that the provider can notify guest_memfd that the page or
>> PFN is going away.
>
> Right. That part I have working in what I was posting.
> [5] https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf06
>

That is similar to what I had in mind except the parameters, I was
thinking inode, offset, size.

kvm_gmem_invalidate_range() in your patch [5] does have a kvm_gmem
prefix, but it's really a direct call to tell KVM MMU to zap the gfn
range, guest_memfd doesn't get to update its own state.

>> guest_memfd cannot say no, but it does have to tell KVM to zap the page
>> from stage 2, and for CoCo it may need to do more stuff.
>>
>> The notification needs some pointer, so I was thinking that on .attach()
>> the provider would also note down which guest_memfd instance the
>> provider should notify.
>>
>> The notifier could tell guest_memfd that some (offset, size) is going
>> away, and breaking down mappings in the EPT/NPT should be something
>> guest_memfd tells KVM to do, I think. Are you thinking that the provider
>> would directly break the mappings down in the EPT/NPT?
>
> I think telling KVM is the right thing to do.
>

Glad to hear that :)

>> Does this notification mechanism work for you? IIUC regular DAX would
>> need it too, like if the device is unplugged for example.
>
> Sure. As long as it works and efficiently covers everyone's use cases,
> I'm not going to be overly opinionated a priori about *how* it works.
>
> All such bets are off once we see the code, of course :)
>
>> > Should I be trying to rework parts of my series to live on top of this?
>>
>> I hope to gather more feedback at LPC. It'd be nice to get early
>> feedback if you think it can work for you! The overall design still
>> needs feedback though, don't take this as the final direction.
>>
>> > Things I have working but which I don't see here include:
>> >
>>
>> This sounds like 4 or more different features! I think they do work with
>> this resource fd proposal.
>>
>> >  • PFN support (which you mentioned).
>>
>> Yup, guest_memfd needs some way to track PFNs in addition to folios. I
>> think for tracking we can reference DAX, the part I haven't looked at is
>> how to let guest_memfd handle events. Like if there's memory failure on
>> a PFN, how do we let guest_memfd handle it?
>
> Telling the guest about correctable and uncorrectable errors, you mean?
> I would have thought that's up to the VMM, until the point where (see
> revocation)?
>

Informing the guest is definitely the userspace VMM's action to
take.

guest_memfd's role here is updating it's own tracking, and perhaps
splitting folios. When guest_memfd supports huge pages, one thing high
on the wishlist is to split the huge folio and zap only the page where
there was an error and not the huge page.

Advantage of splitting: If the failure happened in the last 4K, the
first 1G - 4K can continue to be used by the guest. If the guest never
touches the last 4K, it doesn't have to know about the failure at all.

>> guest_memfd needs to at least be notified, so that it can handle the KVM
>> stuff: zapping and special stuff for CoCo. Whether the provider or
>> guest_memfd gets to handle it first can be worked out :)
>>
>> >  • IOMMUFD support via dma-buf.
>>
>> Is this about having guest_memfd work with dmabuf fds?
>
> It's about having the provider of memory (again, think of it as a daxfs
> if you will) able to tell *both* KVM and IOMMU which pages are where —
> and revoke them at will.
>
> The *implementation* involved dmabuf fds, and we already discussed
> which wraps which, on the first posting of my series. But maybe we'll
> conclude that the right answer is for IOMMUFD to recognise this
> guest_memfd 'provider' as a first-class citizen, and we won't need to
> masquerade as dmabuf?
>
>> Fuad also brought up that Android uses dmabuf as an interface for CMA
>> memory. [3] IIUC the dmabuf fd is mmap()ed, then the userspace address
>> is set up in a memslot for the guest. To make this CMA memory
>> CoCo-friendly, we want to use it with guest_memfd.
>>
>> Jason response was that guest_memfd should probably just directly get
>> memory from CMA [3].
>>
>> I'm not super familiar with CMA or dmabufs, would either have a suitable
>> "resource fd" to pass to guest_memfd like how HugeTLBfs has a mount fd?
>> The "resource fd" should ideally have no way to mmap or read/write the
>> memory directly.
>>
>> [3] https://lore.kernel.org/all/CA+EHjTxZ0N3Tfnid404B4tkb_E+Z8mODTHTgBPiF6=bwZp7Hjw@mail.gmail.com/
>
> I think where the memory comes *from* is an implementation detail for
> the provider. It provides a PFN (or folio, if you must). All else is
> not the business of KVM or IOMMUFD.
>

So we already have

  dmabuf --wraps--> CMA (and others)

We could have

  [a] guest_memfd --wraps--> dmabuf --wraps--> CMA

or

  [b] guest_memfd --wraps--> CMA

I agree guest_memfd shouldn't care where the PFN comes from, but this
does impact where the provider_ops code is implemented though.

If we go with [b], CMA needs to find a non-dmabuf way for userspace to
get some resource fd. If we go with [a], dmabuf folks need to support
this :)

>> Or is this about letting IOMMUFD get pages to map in IOMMU page tables
>> via an fd+offset instead of via mmap()-ed userspace addresses as in
>> IOMMU_IOAS_MAP_FILE, and that guest_memfd isn't a supported kind of fd
>> (yet)?
>>
>> I think guest_memfd should support IOMMU_IOAS_MAP_FILE. One use case is
>> Confidential IO [4], where SNP would need to have private memory
>> (guest_memfd) mapped in the IO page tables.
>
> Right, I think that's what I meant with 'first class citizen' above?
>

I agree, yes, I would like guest_memfd to be first class alongside other
memory providers.

>> The missing parts here are that guest_memfd and iommufd need to
>> cooperate for conversions (or some other coordination in userspace). If
>> private memory gets converted to shared, iommufd needs to know to remap
>> the pages as shared in the IO page tables.
>>
>> [4] https://lore.kernel.org/all/20260225075211.3353194-1-aik@amd.com/
>>
>> >  • Revocation (which I've tested correctly breaks down large pages in
>> >    both EPT/NPT and IOMMU).
>>
>> See above.
>>

I missed considering breaking down huge pages (the folios). guest_memfd
will definitely need to cooperate with the provider to break down the
pages.

I'll need to prototype support with HugeTLBfs, probably after LPC. There
I'll need to split pages on conversion to shared and merge on conversion
to private.

I think you meant breaking down the mappings for the large pages in
stage 2? guest_memfd tells KVM to do that in one of the earlier
prototypes for guest_memfd HugeTLB pages.

>> >  • AsyncPF support for pages requested by KVM.
>>
>> Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness.
>> KVM MMU checks that the page is not there, then since async was
>> configured, KVM tells the guest to schedule in something else.
>>
>> guest_memfd doesn't have support for async #PFs now. I imagine it would
>> be something like KVM MMU firing off a kvm_gmem_get_pfn() request
>> asynchronously, and then guest_memfd gets the page or PFN and inserts it
>> in guest_memfd filemap, and then notifies KVM MMU?
>>
>> Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from
>> the provider.
>>
>> I think it's orthogonal because guest_memfd could
>> have gotten a PAGE_SIZE page (existing functionality) or gotten a page
>> from the provider.
>
> Kind of, but I'm focusing on the guest_memfd provider interface, and
> that needs to support the asychronous mode: asked for a PFN for a given
> guest address, it returns -EAGAIN and then provides it later, and the
> guest gets the right asyncpf behaviour:
> https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34f

I thought the same .alloc_folio would be used, except for an async #PF
the entire kvm_gmem_get_folio(), which calls .alloc_folio() would be
done in some background thread, so the provider doesn't have to have a
different .alloc_folio_async().

If the provider does need to know, then the corresponding thing in this
series could be:

  struct guest_memfd_provider_operations {
	  void *(*attach)(struct file *resource_file);
	  void (*release)(void *provider);
	  struct folio *(*alloc_folio)(void *provider, pgoff_t index,
-				       struct mempolicy *mpol);
+				       struct mempolicy *mpol, bool nowait);
	  void (*invalidate_folio)(void *provider, struct folio *folio);
  };

something like that?

Either way I think it would have to be a later series after the first
one (for upstreaming).

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
  2026-09-29 23:41       ` Ackerley Tng
@ 2026-09-30  0:08         ` David Woodhouse
  0 siblings, 0 replies; 28+ messages in thread
From: David Woodhouse @ 2026-09-30  0:08 UTC (permalink / raw)
  To: Ackerley Tng, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Jason Gunthorpe,
	Nicolin Chen, Xu Yilun, aik, aneesh.kumar, Vlastimil Babka,
	Connor Williamson, Fred Griffoul
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest

[-- Attachment #1: Type: text/plain, Size: 9005 bytes --]

On Tue, 2026-09-29 at 16:41 -0700, Ackerley Tng wrote:
> David Woodhouse <dwmw2@infradead.org> writes:
> 
> > On Mon, 2026-09-28 at 15:59 -0700, Ackerley Tng wrote:
> > > David Woodhouse <dwmw2@infradead.org> writes:
> > > 
> > > > On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
> > > > > 
> > > > > === What if my provider doesn't deal with pages?
> > > > > 
> > > > > Frank and David Woodhouse [2] have use cases for providers that use
> > > > > PFNs, and this is definitely something a generic interface must
> > > > > support.
> > > > > 
> > > > > Here are some options I can think of:
> > > > > 
> > > > > 1. Don't define .alloc_folio(), instead define .alloc_pfn()
> > > > > 2. Refactor .alloc_folio() to .alloc_pfn()
> > > > > 
> > > > > Either way, I think guest_memfd should still be the primary manager of
> > > > > the memory.
> > > > 
> > > > Thanks for working on this.
> > > > 
> > > 
> > > Thanks for the quick reply!
> > > 
> > > > I'm not sure what you mean by 'primary manager' here but my use case is
> > > 
> > > Naming is hard! See comment on revocation below.
> > > 
> > > > for the provider to be the ultimate arbiter of who owns which PFN, and
> > > > revocation of the same. Think of it like a filesystem backed by
> > > > external memory. Each guest's memory is a file, and a guest can
> > > > *donate* specific pages of its memory to other guests (which in the
> > > 
> > > Putting donation aside first - I don't have a good picture of how a
> > > guest would tell the host that it's willing to share a page - can't
> > > comment. Would love to find out more.
> > > 
> > > > context of the *interface* only means that we need revocation).
> > 
> > That's implementation-specific. Think of things like Nitro Enclaves
> > where a guest 'donates' its memory back to the hypervisor to be used by
> > another microvm. But that interface isn't the guest_memfd concern; as I
> > said, in the context of the guest_memfd interface, there is only
> > revocation: that page went away, whatever the reason.
> > 
> 
> Okay we're on the same page then. :) guest_memfd only needs to know that
> the page went away, guest_memfd cannot say no.
> 
> > > Going with this revocation part, beginning with a more basic use case:
> > > I'm thinking that the provider can notify guest_memfd that the page or
> > > PFN is going away.
> > 
> > Right. That part I have working in what I was posting.
> > [5] https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf06
> > 
> 
> That is similar to what I had in mind except the parameters, I was
> thinking inode, offset, size.
> 
> kvm_gmem_invalidate_range() in your patch [5] does have a kvm_gmem
> prefix, but it's really a direct call to tell KVM MMU to zap the gfn
> range, guest_memfd doesn't get to update its own state.

Late here now, but it occurs to me that I need to check if that makes
sense. Does the provider *know* the GFN, or only the offset within its
own object — which might *not* be in a memslot starting at GFN0?

Tangentially: I'd actually quite like the provider to *know* about any
such offset, because if possible I'd like it to be able to be
opinionated about where in the guest pages are mapped.

> > > 
> > > Yup, guest_memfd needs some way to track PFNs in addition to folios. I
> > > think for tracking we can reference DAX, the part I haven't looked at is
> > > how to let guest_memfd handle events. Like if there's memory failure on
> > > a PFN, how do we let guest_memfd handle it?
> > 
> > Telling the guest about correctable and uncorrectable errors, you mean?
> > I would have thought that's up to the VMM, until the point where (see
> > revocation)?
> > 
> 
> Informing the guest is definitely the userspace VMM's action to
> take.
> 
> guest_memfd's role here is updating it's own tracking, and perhaps
> splitting folios. When guest_memfd supports huge pages, one thing high
> on the wishlist is to split the huge folio and zap only the page where
> there was an error and not the huge page.
> 
> Advantage of splitting: If the failure happened in the last 4K, the
> first 1G - 4K can continue to be used by the guest. If the guest never
> touches the last 4K, it doesn't have to know about the failure at all.

Right, but this is still just revocation of a single 4KiB page and
expecting both KVM and IOMMU to shatter the large page accordingly,
isn't it? That much I already had working.

(The IOMMU shatter does need to be atomic and miss-less, which I don't
think is the case on Arm but could be.)

> > 
> > I think where the memory comes *from* is an implementation detail for
> > the provider. It provides a PFN (or folio, if you must). All else is
> > not the business of KVM or IOMMUFD.
> > 
> 
> So we already have
> 
>   dmabuf --wraps--> CMA (and others)
> 
> We could have
> 
>   [a] guest_memfd --wraps--> dmabuf --wraps--> CMA
> 
> or
> 
>   [b] guest_memfd --wraps--> CMA
> 
> I agree guest_memfd shouldn't care where the PFN comes from, but this
> does impact where the provider_ops code is implemented though.
>
> 
> If we go with [b], CMA needs to find a non-dmabuf way for userspace to
> get some resource fd. If we go with [a], dmabuf folks need to support
> this :)

I'm not sure I see how CMA is relevant to the *interface*.

A given *implementation* (provider) of a guest_memfd resource might use
CMA for its backing store. Or might use hugetlbfs. Or something DAX-
like. Once the implementation has provided its guest_memfd operations
with its 'get_pfn' method, nobody else cares *where* it finds the PFNs
that it returns when we invoke that method.

> I missed considering breaking down huge pages (the folios). guest_memfd
> will definitely need to cooperate with the provider to break down the
> pages.

KVM manages this part on its own. The guest_memfd implementation only
has to tell it to revoke. It's been a while, but I don't even remember
having to do anything special.
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf0


> I'll need to prototype support with HugeTLBfs, probably after LPC. There
> I'll need to split pages on conversion to shared and merge on conversion
> to private.
> 
> I think you meant breaking down the mappings for the large pages in
> stage 2? guest_memfd tells KVM to do that in one of the earlier
> prototypes for guest_memfd HugeTLB pages.
> 
> > > >  • AsyncPF support for pages requested by KVM.
> > > 
> > > Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness.
> > > KVM MMU checks that the page is not there, then since async was
> > > configured, KVM tells the guest to schedule in something else.
> > > 
> > > guest_memfd doesn't have support for async #PFs now. I imagine it would
> > > be something like KVM MMU firing off a kvm_gmem_get_pfn() request
> > > asynchronously, and then guest_memfd gets the page or PFN and inserts it
> > > in guest_memfd filemap, and then notifies KVM MMU?
> > > 
> > > Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from
> > > the provider.
> > > 
> > > I think it's orthogonal because guest_memfd could
> > > have gotten a PAGE_SIZE page (existing functionality) or gotten a page
> > > from the provider.
> > 
> > Kind of, but I'm focusing on the guest_memfd provider interface, and
> > that needs to support the asychronous mode: asked for a PFN for a given
> > guest address, it returns -EAGAIN and then provides it later, and the
> > guest gets the right asyncpf behaviour:
> > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34f
> 
> I thought the same .alloc_folio would be used, except for an async #PF
> the entire kvm_gmem_get_folio(), which calls .alloc_folio() would be
> done in some background thread, so the provider doesn't have to have a
> different .alloc_folio_async().
> 

> If the provider does need to know, then the corresponding thing in this
> series could be:
> 
>   struct guest_memfd_provider_operations {
> 	  void *(*attach)(struct file *resource_file);
> 	  void (*release)(void *provider);
> 	  struct folio *(*alloc_folio)(void *provider, pgoff_t index,
> -				       struct mempolicy *mpol);
> +				       struct mempolicy *mpol, bool nowait);
> 	  void (*invalidate_folio)(void *provider, struct folio *folio);
>   };
> 
> something like that?

Yeah, the 'nowait' argument on get_pfn() is how I have it working at
the moment. And then, as you say, KVM calls it again *without* nowait,
in a context which can sleep.

> Either way I think it would have to be a later series after the first
> one (for upstreaming).

I don't have a particular use case for asyncpf right now personally,
but I think we *should* get the design for the interface right. We
should do it on the IOMMU side too, because ATS+PRI exists.

[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs
  2026-09-26  0:50 ` [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs Ackerley Tng via B4 Relay
@ 2026-10-01 13:07   ` David Woodhouse
  2026-10-01 14:17     ` Jason Gunthorpe
  0 siblings, 1 reply; 28+ messages in thread
From: David Woodhouse @ 2026-10-01 13:07 UTC (permalink / raw)
  To: ackerleytng, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Jason Gunthorpe,
	Nicolin Chen, Xu Yilun, aik, aneesh.kumar, Vlastimil Babka,
	Fred Griffoul, Connor Williamson
  Cc: kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest

[-- Attachment #1: Type: text/plain, Size: 1316 bytes --]

On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
> 
> @@ -130,6 +131,9 @@ struct super_operations {
>  
>  	/* Report a filesystem error */
>  	void (*report_error)(const struct fserror_event *event);
> +#ifdef CONFIG_KVM_GUEST_MEMFD
> +	const struct guest_memfd_provider_operations *gmem_provider_ops;
> +#endif
>  };
>  
>  struct super_block {

I think I'd prefer the memory provider ops *not* to explicitly mention
KVM or guest_memfd. 

It is *just* a provider of memory. It manages its own address space,
and gives you the PFN (or folio, if you must) for a given range of its
address space on request. Perhaps with a 'nowait' flag to allow for
asynchronous operation (expecting its caller to try again without that
flag from a context where it *does* want to wait).

It will also want to manage *permissions* for that memory (read-only,
probably shared/private too?). And *revocation* when it wants a page
back.

The *consumers* of this memory provider API will include KVM/guestmemfd
and IOMMUFD. The latter *might* be through masquerading as dmabuf, or
IOMMUFD might treat memory providers as a first class citizen and be
taught to consume them directly.

But let's try to keep the interface, its implementations, and its
consumers cleanly separated.

[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
  2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
                   ` (17 preceding siblings ...)
  2026-09-28 16:32 ` [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd David Woodhouse
@ 2026-10-01 13:43 ` Gregory Price
  18 siblings, 0 replies; 28+ messages in thread
From: Gregory Price @ 2026-10-01 13:43 UTC (permalink / raw)
  To: Ackerley Tng
  Cc: Hugh Dickins, Baolin Wang, Andrew Morton, Sean Christopherson,
	Paolo Bonzini, David Hildenbrand, Jonathan Corbet, Shuah Khan,
	Randy Dunlap, Shuah Khan, vannapurve, erdemaktas, jxgao,
	rientjes, fvdl, jthoughton, tarunsahu, pratyush, fuad.tabba,
	David Woodhouse, yan.y.zhao, michael.roth, suzuki.poulose,
	Christian Brauner, Jason Gunthorpe, Nicolin Chen, Xu Yilun, aik,
	aneesh.kumar, Vlastimil Babka, kernel-team, kernel-team,
	linux-kernel, linux-mm, kvm, linux-doc, linux-kselftest

On Fri, Sep 25, 2026 at 05:50:47PM -0700, Ackerley Tng wrote:
> + e.g. for tmpfs,
>     + Use mount page size.
>     + Respect mount size limits.
>     + Use guest_memfd mpol and fall back to mount mpol.
>         + Gregory, now you can set mpol upfront on a mount! [3])

I haven't had time to quite give this the read it deserves yet (have
been traveling this past week), but I think this mpol/mount issue can
be solved with a new dax driver (which maybe provides this provider
interface) without needing to involve numa policy at all.

Essentially

fd = open("/dev/cxl_dev", ...); /* backed by anondax driver */

then later when a fault on that fd happens, the driver provides an
anon-memory allocation that explicitly requests the node associated
with that device - as opposed to trying to jam this into mempolicy.

If you're planning to be at LPC, happy to discuss this in person, but I
think this provider pattern is likely much cleaner without mempolicy at
all if you can just make the provider do that allocation path for you.

~Gregory

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs
  2026-10-01 13:07   ` David Woodhouse
@ 2026-10-01 14:17     ` Jason Gunthorpe
  2026-10-01 15:47       ` David Woodhouse
  0 siblings, 1 reply; 28+ messages in thread
From: Jason Gunthorpe @ 2026-10-01 14:17 UTC (permalink / raw)
  To: David Woodhouse
  Cc: ackerleytng, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka, Fred Griffoul,
	Connor Williamson, kernel-team, kernel-team, linux-kernel,
	linux-mm, kvm, linux-doc, linux-kselftest

On Thu, Oct 01, 2026 at 02:07:40PM +0100, David Woodhouse wrote:
> The *consumers* of this memory provider API will include KVM/guestmemfd
> and IOMMUFD. The latter *might* be through masquerading as dmabuf, or
> IOMMUFD might treat memory providers as a first class citizen and be
> taught to consume them directly.

If we are reinventing a dmabuf improved for these use cases then lets
just use it directly.

It would be nice if it could handle MMIO and "cachable MMIO" as well.

Jason

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs
  2026-10-01 14:17     ` Jason Gunthorpe
@ 2026-10-01 15:47       ` David Woodhouse
  2026-10-01 17:15         ` Jason Gunthorpe
  0 siblings, 1 reply; 28+ messages in thread
From: David Woodhouse @ 2026-10-01 15:47 UTC (permalink / raw)
  To: Jason Gunthorpe, Fred Griffoul, Connor Williamson
  Cc: ackerleytng, Hugh Dickins, Baolin Wang, Andrew Morton,
	Sean Christopherson, Paolo Bonzini, David Hildenbrand,
	Jonathan Corbet, Shuah Khan, Randy Dunlap, Shuah Khan,
	vannapurve, erdemaktas, jxgao, rientjes, fvdl, jthoughton,
	tarunsahu, pratyush, fuad.tabba, Gregory Price, yan.y.zhao,
	michael.roth, suzuki.poulose, Christian Brauner, Nicolin Chen,
	Xu Yilun, aik, aneesh.kumar, Vlastimil Babka, Fred Griffoul,
	Connor Williamson, kernel-team, kernel-team, linux-kernel,
	linux-mm, kvm, linux-doc, linux-kselftest

[-- Attachment #1: Type: text/plain, Size: 888 bytes --]

On Thu, 2026-10-01 at 11:17 -0300, Jason Gunthorpe wrote:
> On Thu, Oct 01, 2026 at 02:07:40PM +0100, David Woodhouse wrote:
> > The *consumers* of this memory provider API will include KVM/guestmemfd
> > and IOMMUFD. The latter *might* be through masquerading as dmabuf, or
> > IOMMUFD might treat memory providers as a first class citizen and be
> > taught to consume them directly.
> 
> If we are reinventing a dmabuf improved for these use cases then lets
> just use it directly.


Yes, that's literally what my series is doing at the moment.

> It would be nice if it could handle MMIO and "cachable MMIO" as well.

As opposed to just MMIO vs. RAM as implemented in the prototype in
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=a7ca2d8859036 ?

I think Fred has a cleaner version of that about ready to post as the
next version of my series.

[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]

^ permalink raw reply	[flat|nested] 28+ messages in thread

* Re: [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs
  2026-10-01 15:47       ` David Woodhouse
@ 2026-10-01 17:15         ` Jason Gunthorpe
  0 siblings, 0 replies; 28+ messages in thread
From: Jason Gunthorpe @ 2026-10-01 17:15 UTC (permalink / raw)
  To: David Woodhouse
  Cc: Fred Griffoul, Connor Williamson, ackerleytng, Hugh Dickins,
	Baolin Wang, Andrew Morton, Sean Christopherson, Paolo Bonzini,
	David Hildenbrand, Jonathan Corbet, Shuah Khan, Randy Dunlap,
	Shuah Khan, vannapurve, erdemaktas, jxgao, rientjes, fvdl,
	jthoughton, tarunsahu, pratyush, fuad.tabba, Gregory Price,
	yan.y.zhao, michael.roth, suzuki.poulose, Christian Brauner,
	Nicolin Chen, Xu Yilun, aik, aneesh.kumar, Vlastimil Babka,
	kernel-team, kernel-team, linux-kernel, linux-mm, kvm, linux-doc,
	linux-kselftest

On Thu, Oct 01, 2026 at 04:47:54PM +0100, David Woodhouse wrote:
> On Thu, 2026-10-01 at 11:17 -0300, Jason Gunthorpe wrote:

> > It would be nice if it could handle MMIO and "cachable MMIO" as well.
> 
> As opposed to just MMIO vs. RAM as implemented in the prototype in
> https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=a7ca2d8859036 ?

Sort of, RAM is a little more permanent than "cachable mmio", I prefer
the latter because it conveys it is located on a device, is subject to
things like reset. Eg "CXL ram" would be cachable mmio.

I haven't heard of anything needing a unique PTE for this yet, but who
knows down the road.

> I think Fred has a cleaner version of that about ready to post as the
> next version of my series.

Yeah, cleaned up is needed, I want us to resolve this in the dmabuf
code cleanly not add more of the hack mess I made for vfio/iommufd..

Jason

^ permalink raw reply	[flat|nested] 28+ messages in thread

end of thread, other threads:[~2026-10-01 17:15 UTC | newest]

Thread overview: 28+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-26  0:50 [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs Ackerley Tng via B4 Relay
2026-10-01 13:07   ` David Woodhouse
2026-10-01 14:17     ` Jason Gunthorpe
2026-10-01 15:47       ` David Woodhouse
2026-10-01 17:15         ` Jason Gunthorpe
2026-09-26  0:50 ` [PATCH RFC 02/17] KVM: guest_memfd: Support provider folio allocation Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 03/17] KVM: guest_memfd: Support provider folio invalidation Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 04/17] KVM: guest_memfd: Add helper to attach resource provider file Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 05/17] KVM: selftests: Add helper to create guest_memfd with a resource file Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 06/17] KVM: selftests: Test negative validation of resource_fd argument Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 07/17] KVM: selftests: Test rejection of unsupported filesystem for resource_fd Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 08/17] KVM: selftests: Test rejection of tmpfs file " Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 09/17] KVM: selftests: Test rejection of swap tmpfs mounts Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 10/17] KVM: selftests: Test rejection of hugepage " Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 11/17] KVM: selftests: Test guest_memfd with anonymous fsmount() tmpfs pool Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 12/17] KVM: selftests: Test guest_memfd with mounted tmpfs root directory Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 13/17] KVM: selftests: Test guest_memfd resource pool sharing across instances Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 14/17] KVM: selftests: Test memory allocation against shared tmpfs resource pool Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 15/17] KVM: selftests: Test that shared tmpfs resource pool size limit is respected Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 16/17] KVM: selftests: Test guest execution with tmpfs-backed guest_memfd Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 17/17] KVM: selftests: Document testing TODOs Ackerley Tng via B4 Relay
2026-09-28 16:32 ` [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd David Woodhouse
2026-09-28 22:59   ` Ackerley Tng
2026-09-29 16:17     ` David Woodhouse
2026-09-29 23:41       ` Ackerley Tng
2026-09-30  0:08         ` David Woodhouse
2026-10-01 13:43 ` Gregory Price

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®