mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Sean Christopherson <seanjc@google.com>
To: Yosry Ahmed <yosry@kernel.org>
Cc: Brendan Jackman <brendan.jackman@linux.dev>,
	Brendan Jackman <jackmanb@google.com>,
	 Borislav Petkov <bp@alien8.de>,
	Dave Hansen <dave.hansen@linux.intel.com>,
	 Peter Zijlstra <peterz@infradead.org>,
	Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,  Wei Xu <weixugc@google.com>,
	Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
	 Lorenzo Stoakes <ljs@kernel.org>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	 x86@kernel.org, Sumit Garg <sumit.garg@oss.qualcomm.com>,
	Will Deacon <will@kernel.org>,
	 rientjes@google.com, patrick.roy@linux.dev,
	 Takahiro Itazuri <itazur@amazon.co.uk>,
	Andy Lutomirski <luto@kernel.org>,
	 David Kaplan <david.kaplan@amd.com>,
	Thomas Gleixner <tglx@kernel.org>,
	 Patrick Bellasi <derkling@google.com>,
	Reiji Watanabe <reijiw@google.com>,
	 Nikita Kalyazin <nikita.kalyazin@linux.dev>,
	Ackerley Tng <ackerleytng@google.com>
Subject: Re: [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP
Date: Fri, 7 Aug 2026 07:26:39 -0700	[thread overview]
Message-ID: <anXrH_6oQTtDxzX9@google.com> (raw)
In-Reply-To: <CAO9r8zPBdjxPTMCeaLdeFeH=9Ybb8wn6KJZ6x9t=o5DyXdEHwA@mail.gmail.com>

On Thu, Aug 06, 2026, Yosry Ahmed wrote:
> On Thu, Aug 6, 2026 at 5:19 PM Sean Christopherson <seanjc@google.com> wrote:
> > > >  2. Always access guest memory through userspace mappings, i.e. through uaccess.
> > > >
> > > > #2 sounds nice, but the problem is that it effectively requires hand-coded assembly
> > > > sequences for anything more complex than basic load/store operations.  Which isn't
> > > > a complete non-starter, but it's a pretty big blocker.  E.g. see the mess that is
> > > > record_steal_time(), and then imagine trying to convert something like
> > > > nested_vmx_prepare_msr_bitmap() to use uaccess.
> > > >
> > > > So, unless someone comes up with a clever idea, KVM will need something GUP-like.
> > > > Strictly speaking, it doesn't necessarily need to be exactly GUP, because KVM could
> > > > poke into guest_memfd directly; KVM would "just" need to manually track its own
> > > > mappings.
> > >
> > > Ideally we can have shared infrastructure for this (i.e. the mermap).
> > >
> > > > But on x86 at least, that's not really a viable option because it only
> > > > works for map-rarely, read/write-many use cases.  For one-off accesses, creating
> > > > and destroying (very) shortlived mappings would be too costly, and so we'd want
> > > > those to go through uaccess.
> > >
> > > Not necessarily (I hope). I think the mermap pre-allocates page tables
> > > (or some of them) and defers some TLB flushes, it's aimed at
> > > short-lived mappings (e.g. for read() syscalls to map a file page,
> > > copy to buffer, then unmap).
> >
> > I highly recommend testing shadow paging if you have aspirations of replacing
> > the get_user() in FNAME(walk_addr_generic) with an on-demand mapping.  I would
> > be (pleasantly) shocked if dynamic mappings can provide acceptable performance.
> 
> Oh I was thinking of existing cases where KVM uses kernel mappings (e.g.
> kvm_vcpu_map()), as these are the ones where KVM uses GUP now, and need to
> work for AS_NO_DIRECT_MAP to be usable. I assume the get_user() calls are
> already a problem for guest_memfd.  

They aren't.  KVM doesn't yet support in-place conversion, so when guest_memfd is
used for private memory, the backing for shared memory must come from something
other than guest_memfd.  If the guest does something to prompt a host/KVM access
to memory that is private or doesn't have a valid backing, then it's either a
guest bug or a host userspace VMM bug.

When guest_memfd is being used for shared memory, i.e. was created with
GUEST_MEMFD_FLAG_INIT_SHARED, then get_user() Just Works, because again it's
userspace's responsibility to establish mappings for memory that KVM may need to
access.

All of that holds true for when in-place conversion comes along: if get_user()
hits a fault, either the guest or userspace screwed up.

> AS_NO_DIRECT_MAP will surely make it a bigger problem, but not a new one :P

Well, if it disallows GUP, that will be a new problem.

> > > > At that point, userspace is basically required to
> > > > maintain mappings for all host-accessible guest memory, and if there are userspace
> > > > mappings, then not using GUP doesn't make much sense.
> > > >
> > > > Note, I called out x86 because x86 has the most extensive emulator and shadow
> > > > paging support, which is where the isolated, one-off accesses happen in spades.
> > > > Other architectures might be able to squeak by without userspace mappings, at
> > > > least for now.
> > > >
> > > > So, in all likelihood, KVM will want GUP.
> > >
> > > Yeah I am thinking that the check here to disallow GUP completely for
> > > unmapped pages is aggressive. Maybe it works for now if KVM does not
> > > currently have any use cases for accessing guest_memfd memory. But if it does
> > > (or will very soon), we need to think more about it, otherwise
> > > AS_NO_DIRECT_MAP is not really usable for guest_memfd. Since you said KVM
> > > "will want" GUP, I assume it currently doesn't?
> >
> > Doesn't what?  Have GUP?  KVM heavily uses GUP, including for guest_memfd that
> > can be mapped into userspace.
> 
> Your wording made me think that KVM doesn't currently use GUP for
> guest_memfd, but I was obviously wrong. So IIUC GUP needs to succeed for
> guest_memfd pages with AS_NO_DIRECT_MAP. 

Yes, though as I said early, it doesn't *have* to be exactly GUP, just something
GUP-like.  E.g. it could be a new API, if that's easier/cleaner.  What I don't
think is a good idea though is handling this entirely in KVM/guest_memfd.

> To actually access the memory, I assume the guest_memfd side will need to
> handle this by either using ephemeral mappings (e.g. mermap), restoring and
> zapping direct mappings, or using a userspace mapping. I suppose for the
> purposes of AS_NO_DIRECT_MAP core support we just need GUP to succeed?

And establish a (ephemeral?) kernel mapping, because general users of GUP will
expect that they can access the physical memory through the direct map.  That's
why I didn't want to handle any of this in KVM[*], the rules and handling need
to be kernel-wide.

[*] https://lore.kernel.org/all/aeemS2wm38Cm4qAf@google.com

  reply	other threads:[~2026-08-07 14:26 UTC|newest]

Thread overview: 72+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-26 22:22 [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 01/26] set_memory: add folio_{zap,restore}_direct_map helpers Brendan Jackman
2026-07-27 10:33   ` Mike Rapoport
2026-07-29 11:42     ` Brendan Jackman
2026-07-30 20:34   ` Yosry Ahmed
2026-07-31  5:21     ` Mike Rapoport
2026-07-31 11:57       ` Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 02/26] mm/secretmem: make use of folio_{zap,restore}_direct_map Brendan Jackman
2026-07-27 10:40   ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP Brendan Jackman
2026-07-30 21:06   ` Yosry Ahmed
2026-07-31 12:15     ` Brendan Jackman
2026-07-31 19:28       ` Yosry Ahmed
2026-08-07  0:02         ` Sean Christopherson
2026-08-07  0:13           ` Yosry Ahmed
2026-08-07  0:19             ` Sean Christopherson
2026-08-07  0:29               ` Yosry Ahmed
2026-08-07 14:26                 ` Sean Christopherson [this message]
2026-08-07 18:12                   ` Yosry Ahmed
2026-08-07 18:49                     ` Sean Christopherson
2026-08-02 16:10   ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 04/26] x86/mm: split out preallocate_sub_pgd() Brendan Jackman
2026-07-31 22:10   ` Yosry Ahmed
2026-08-02 16:13   ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 05/26] x86: move PAE PMD preallocation defines to header Brendan Jackman
2026-07-31 23:59   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 06/26] x86/tlb: Expose some flush function declarations to modules Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 07/26] x86/mm: introduce mm-local region Brendan Jackman
2026-08-02 16:27   ` Mike Rapoport
2026-08-03 22:29   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 08/26] x86/mm: move LDT remap into " Brendan Jackman
2026-08-03 22:33   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 09/26] mm: Create flags arg for __apply_to_page_range() Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 10/26] mm: Add more flags " Brendan Jackman
2026-08-04  0:08   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 11/26] x86/mm: introduce the mermap Brendan Jackman
2026-08-02 16:40   ` Mike Rapoport
2026-08-04 18:38   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 12/26] mm: KUnit tests for " Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 13/26] mm: introduce freetype_t Brendan Jackman
2026-08-04 22:23   ` Yosry Ahmed
2026-08-04 23:02   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 14/26] mm: move migratetype definitions to freetype.h Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 15/26] mm/page_alloc: add support for freetypes with no freelist Brendan Jackman
2026-07-31 14:13   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 16/26] mm: add definitions for allocating unmapped pages Brendan Jackman
2026-08-04 19:53   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 17/26] mm: encode freetype flags in pageblock flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 18/26] mm/page_alloc: separate pcplists by freetype flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 19/26] mm/page_alloc: rename ALLOC_NON_BLOCK back to _HARDER Brendan Jackman
2026-07-31 14:52   ` Vlastimil Babka (SUSE)
2026-08-03  9:20     ` Vlastimil Babka (SUSE)
2026-08-04 21:50     ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 20/26] mm/page_alloc: introduce ALLOC_NOBLOCK Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 21/26] mm/page_alloc: implement FREETYPE_UNMAPPED allocations Brendan Jackman
2026-08-03  9:18   ` Vlastimil Babka (SUSE)
2026-08-04 23:41   ` Yosry Ahmed
2026-08-07  0:05     ` Yosry Ahmed
2026-08-04 23:53   ` Yosry Ahmed
2026-08-05 16:13     ` Yosry Ahmed
2026-08-07  0:16   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 22/26] mm: Minimal KUnit tests for some new page_alloc logic Brendan Jackman
2026-08-03  9:30   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 23/26] mm: Split out NR_FREE_PAGES_BLOCKS_[UN]MAPPED Brendan Jackman
2026-08-03  9:32   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs Brendan Jackman
2026-08-03  9:44   ` Vlastimil Babka (SUSE)
2026-08-06 23:29   ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 25/26] mm: plumb alloc flags into some alloc funcs Brendan Jackman
2026-08-03  9:52   ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 26/26] mm: add fast path for AS_NO_DIRECT_MAP Brendan Jackman
2026-07-29 11:52 ` [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anXrH_6oQTtDxzX9@google.com \
    --to=seanjc@google.com \
    --cc=ackerleytng@google.com \
    --cc=akpm@linux-foundation.org \
    --cc=bp@alien8.de \
    --cc=brendan.jackman@linux.dev \
    --cc=dave.hansen@linux.intel.com \
    --cc=david.kaplan@amd.com \
    --cc=david@kernel.org \
    --cc=derkling@google.com \
    --cc=hannes@cmpxchg.org \
    --cc=itazur@amazon.co.uk \
    --cc=jackmanb@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=luto@kernel.org \
    --cc=nikita.kalyazin@linux.dev \
    --cc=patrick.roy@linux.dev \
    --cc=peterz@infradead.org \
    --cc=reijiw@google.com \
    --cc=rientjes@google.com \
    --cc=rppt@kernel.org \
    --cc=sumit.garg@oss.qualcomm.com \
    --cc=tglx@kernel.org \
    --cc=vbabka@kernel.org \
    --cc=weixugc@google.com \
    --cc=will@kernel.org \
    --cc=x86@kernel.org \
    --cc=yosry@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome