mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: David Woodhouse <dwmw2@infradead.org>
To: Ackerley Tng <ackerleytng@google.com>,
	Hugh Dickins <hughd@google.com>,
	 Baolin Wang <baolin.wang@linux.alibaba.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	Sean Christopherson <seanjc@google.com>,
	Paolo Bonzini <pbonzini@redhat.com>,
	David Hildenbrand <david@kernel.org>,
	 Jonathan Corbet	 <corbet@lwn.net>,
	Shuah Khan <skhan@linuxfoundation.org>,
	Randy Dunlap	 <rdunlap@infradead.org>,
	Shuah Khan <shuah@kernel.org>,
	vannapurve@google.com, 	erdemaktas@google.com, jxgao@google.com,
	rientjes@google.com, fvdl@google.com, 	jthoughton@google.com,
	tarunsahu@google.com, pratyush@kernel.org, 	fuad.tabba@linux.dev,
	Gregory Price <gourry@gourry.net>,
	yan.y.zhao@intel.com, 	michael.roth@amd.com,
	suzuki.poulose@arm.com, Christian Brauner	 <brauner@kernel.org>,
	Jason Gunthorpe <jgg@ziepe.ca>,
	Nicolin Chen	 <nicolinc@nvidia.com>,
	Xu Yilun <yilun.xu@linux.intel.com>,
	aik@amd.com, 	aneesh.kumar@kernel.org,
	Vlastimil Babka <vbabka@kernel.org>,
	Connor Williamson	 <connordw@amazon.co.uk>,
	Fred Griffoul <fgriffo@amazon.co.uk>
Cc: kernel-team@android.com, kernel-team@meta.com,
	linux-kernel@vger.kernel.org, 	linux-mm@kvack.org,
	kvm@vger.kernel.org, linux-doc@vger.kernel.org,
		linux-kselftest@vger.kernel.org
Subject: Re: [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd
Date: Tue, 29 Sep 2026 17:17:06 +0100	[thread overview]
Message-ID: <ecab388213245df81815e42246db3b13b2d43e2b.camel@infradead.org> (raw)
In-Reply-To: <CAEvNRgGw_RjiGs_bbVnoRAvq49Qukm4h8XhkBd0DXxn1ZzLrQg@mail.gmail.com>

[-- Attachment #1: Type: text/plain, Size: 7854 bytes --]

On Mon, 2026-09-28 at 15:59 -0700, Ackerley Tng wrote:
> David Woodhouse <dwmw2@infradead.org> writes:
> 
> > On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
> > > 
> > > === What if my provider doesn't deal with pages?
> > > 
> > > Frank and David Woodhouse [2] have use cases for providers that use
> > > PFNs, and this is definitely something a generic interface must
> > > support.
> > > 
> > > Here are some options I can think of:
> > > 
> > > 1. Don't define .alloc_folio(), instead define .alloc_pfn()
> > > 2. Refactor .alloc_folio() to .alloc_pfn()
> > > 
> > > Either way, I think guest_memfd should still be the primary manager of
> > > the memory.
> > 
> > Thanks for working on this.
> > 
> 
> Thanks for the quick reply!
> 
> > I'm not sure what you mean by 'primary manager' here but my use case is
> 
> Naming is hard! See comment on revocation below.
> 
> > for the provider to be the ultimate arbiter of who owns which PFN, and
> > revocation of the same. Think of it like a filesystem backed by
> > external memory. Each guest's memory is a file, and a guest can
> > *donate* specific pages of its memory to other guests (which in the
> 
> Putting donation aside first - I don't have a good picture of how a
> guest would tell the host that it's willing to share a page - can't
> comment. Would love to find out more.
>
> > context of the *interface* only means that we need revocation).

That's implementation-specific. Think of things like Nitro Enclaves
where a guest 'donates' its memory back to the hypervisor to be used by
another microvm. But that interface isn't the guest_memfd concern; as I
said, in the context of the guest_memfd interface, there is only
revocation: that page went away, whatever the reason.

> Going with this revocation part, beginning with a more basic use case:
> I'm thinking that the provider can notify guest_memfd that the page or
> PFN is going away.

Right. That part I have working in what I was posting.
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf06

> guest_memfd cannot say no, but it does have to tell KVM to zap the page
> from stage 2, and for CoCo it may need to do more stuff.
>
> The notification needs some pointer, so I was thinking that on .attach()
> the provider would also note down which guest_memfd instance the
> provider should notify.
>
> The notifier could tell guest_memfd that some (offset, size) is going
> away, and breaking down mappings in the EPT/NPT should be something
> guest_memfd tells KVM to do, I think. Are you thinking that the provider
> would directly break the mappings down in the EPT/NPT?

I think telling KVM is the right thing to do.

> Does this notification mechanism work for you? IIUC regular DAX would
> need it too, like if the device is unplugged for example.

Sure. As long as it works and efficiently covers everyone's use cases,
I'm not going to be overly opinionated a priori about *how* it works.

All such bets are off once we see the code, of course :)

> > Should I be trying to rework parts of my series to live on top of this?
> 
> I hope to gather more feedback at LPC. It'd be nice to get early
> feedback if you think it can work for you! The overall design still
> needs feedback though, don't take this as the final direction.
> 
> > Things I have working but which I don't see here include:
> > 
> 
> This sounds like 4 or more different features! I think they do work with
> this resource fd proposal.
> 
> >  • PFN support (which you mentioned).
> 
> Yup, guest_memfd needs some way to track PFNs in addition to folios. I
> think for tracking we can reference DAX, the part I haven't looked at is
> how to let guest_memfd handle events. Like if there's memory failure on
> a PFN, how do we let guest_memfd handle it?

Telling the guest about correctable and uncorrectable errors, you mean?
I would have thought that's up to the VMM, until the point where (see
revocation)?

> guest_memfd needs to at least be notified, so that it can handle the KVM
> stuff: zapping and special stuff for CoCo. Whether the provider or
> guest_memfd gets to handle it first can be worked out :)
> 
> >  • IOMMUFD support via dma-buf.
> 
> Is this about having guest_memfd work with dmabuf fds?

It's about having the provider of memory (again, think of it as a daxfs
if you will) able to tell *both* KVM and IOMMU which pages are where —
and revoke them at will.

The *implementation* involved dmabuf fds, and we already discussed
which wraps which, on the first posting of my series. But maybe we'll
conclude that the right answer is for IOMMUFD to recognise this
guest_memfd 'provider' as a first-class citizen, and we won't need to
masquerade as dmabuf?

> Fuad also brought up that Android uses dmabuf as an interface for CMA
> memory. [3] IIUC the dmabuf fd is mmap()ed, then the userspace address
> is set up in a memslot for the guest. To make this CMA memory
> CoCo-friendly, we want to use it with guest_memfd.
> 
> Jason response was that guest_memfd should probably just directly get
> memory from CMA [3].
> 
> I'm not super familiar with CMA or dmabufs, would either have a suitable
> "resource fd" to pass to guest_memfd like how HugeTLBfs has a mount fd?
> The "resource fd" should ideally have no way to mmap or read/write the
> memory directly.
> 
> [3] https://lore.kernel.org/all/CA+EHjTxZ0N3Tfnid404B4tkb_E+Z8mODTHTgBPiF6=bwZp7Hjw@mail.gmail.com/

I think where the memory comes *from* is an implementation detail for
the provider. It provides a PFN (or folio, if you must). All else is
not the business of KVM or IOMMUFD.

> Or is this about letting IOMMUFD get pages to map in IOMMU page tables
> via an fd+offset instead of via mmap()-ed userspace addresses as in
> IOMMU_IOAS_MAP_FILE, and that guest_memfd isn't a supported kind of fd
> (yet)?
> 
> I think guest_memfd should support IOMMU_IOAS_MAP_FILE. One use case is
> Confidential IO [4], where SNP would need to have private memory
> (guest_memfd) mapped in the IO page tables.

Right, I think that's what I meant with 'first class citizen' above?

> The missing parts here are that guest_memfd and iommufd need to
> cooperate for conversions (or some other coordination in userspace). If
> private memory gets converted to shared, iommufd needs to know to remap
> the pages as shared in the IO page tables.
> 
> [4] https://lore.kernel.org/all/20260225075211.3353194-1-aik@amd.com/
> 
> >  • Revocation (which I've tested correctly breaks down large pages in
> >    both EPT/NPT and IOMMU).
> 
> See above.
> 
> >  • AsyncPF support for pages requested by KVM.
> 
> Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness.
> KVM MMU checks that the page is not there, then since async was
> configured, KVM tells the guest to schedule in something else.
> 
> guest_memfd doesn't have support for async #PFs now. I imagine it would
> be something like KVM MMU firing off a kvm_gmem_get_pfn() request
> asynchronously, and then guest_memfd gets the page or PFN and inserts it
> in guest_memfd filemap, and then notifies KVM MMU?
> 
> Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from
> the provider.
> 
> I think it's orthogonal because guest_memfd could
> have gotten a PAGE_SIZE page (existing functionality) or gotten a page
> from the provider.

Kind of, but I'm focusing on the guest_memfd provider interface, and
that needs to support the asychronous mode: asked for a PFN for a given
guest address, it returns -EAGAIN and then provides it later, and the
guest gets the right asyncpf behaviour:
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34f

[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]

  reply	other threads:[~2026-09-29 16:17 UTC|newest]

Thread overview: 28+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-26  0:50 Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 01/17] mm: shmem: Implement guest_memfd provider operations for tmpfs Ackerley Tng via B4 Relay
2026-10-01 13:07   ` David Woodhouse
2026-10-01 14:17     ` Jason Gunthorpe
2026-10-01 15:47       ` David Woodhouse
2026-10-01 17:15         ` Jason Gunthorpe
2026-09-26  0:50 ` [PATCH RFC 02/17] KVM: guest_memfd: Support provider folio allocation Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 03/17] KVM: guest_memfd: Support provider folio invalidation Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 04/17] KVM: guest_memfd: Add helper to attach resource provider file Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 05/17] KVM: selftests: Add helper to create guest_memfd with a resource file Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 06/17] KVM: selftests: Test negative validation of resource_fd argument Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 07/17] KVM: selftests: Test rejection of unsupported filesystem for resource_fd Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 08/17] KVM: selftests: Test rejection of tmpfs file " Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 09/17] KVM: selftests: Test rejection of swap tmpfs mounts Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 10/17] KVM: selftests: Test rejection of hugepage " Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 11/17] KVM: selftests: Test guest_memfd with anonymous fsmount() tmpfs pool Ackerley Tng via B4 Relay
2026-09-26  0:50 ` [PATCH RFC 12/17] KVM: selftests: Test guest_memfd with mounted tmpfs root directory Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 13/17] KVM: selftests: Test guest_memfd resource pool sharing across instances Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 14/17] KVM: selftests: Test memory allocation against shared tmpfs resource pool Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 15/17] KVM: selftests: Test that shared tmpfs resource pool size limit is respected Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 16/17] KVM: selftests: Test guest execution with tmpfs-backed guest_memfd Ackerley Tng via B4 Relay
2026-09-26  0:51 ` [PATCH RFC 17/17] KVM: selftests: Document testing TODOs Ackerley Tng via B4 Relay
2026-09-28 16:32 ` [PATCH RFC 00/17] Allow guest_memfd to be created using a resource (pool) fd David Woodhouse
2026-09-28 22:59   ` Ackerley Tng
2026-09-29 16:17     ` David Woodhouse [this message]
2026-09-29 23:41       ` Ackerley Tng
2026-09-30  0:08         ` David Woodhouse
2026-10-01 13:43 ` Gregory Price

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ecab388213245df81815e42246db3b13b2d43e2b.camel@infradead.org \
    --to=dwmw2@infradead.org \
    --cc=ackerleytng@google.com \
    --cc=aik@amd.com \
    --cc=akpm@linux-foundation.org \
    --cc=aneesh.kumar@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=brauner@kernel.org \
    --cc=connordw@amazon.co.uk \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=erdemaktas@google.com \
    --cc=fgriffo@amazon.co.uk \
    --cc=fuad.tabba@linux.dev \
    --cc=fvdl@google.com \
    --cc=gourry@gourry.net \
    --cc=hughd@google.com \
    --cc=jgg@ziepe.ca \
    --cc=jthoughton@google.com \
    --cc=jxgao@google.com \
    --cc=kernel-team@android.com \
    --cc=kernel-team@meta.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=michael.roth@amd.com \
    --cc=nicolinc@nvidia.com \
    --cc=pbonzini@redhat.com \
    --cc=pratyush@kernel.org \
    --cc=rdunlap@infradead.org \
    --cc=rientjes@google.com \
    --cc=seanjc@google.com \
    --cc=shuah@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=suzuki.poulose@arm.com \
    --cc=tarunsahu@google.com \
    --cc=vannapurve@google.com \
    --cc=vbabka@kernel.org \
    --cc=yan.y.zhao@intel.com \
    --cc=yilun.xu@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®