From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B1DE8DF59 for ; Fri, 7 Aug 2026 00:19:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786061952; cv=none; b=qLN75O1FCf31rgGK50E+xsfpfsS+Nc60DPT9GcYCF6Pf+Z7qtYVSFA+hYro5uJiziPab1B674Z8GtyMHR7SopsHc1LAUK0JfcG+s0PdNsrMzZwzOGnS1omAXrGLf/uBIZzhXxelYXb4H6PCNKGqBGk7QH64y9YBoRgoqPtryNs8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786061952; c=relaxed/simple; bh=trU9VK+z73UmrATTdFAJ7DIjblAUIoeYIqOqJxcM0Uw=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=TwMrtjFH79wPvS0S4jLVLEF+HosLHxQwesXib1nU8rLDPj28vVInqzr538WxdZrkAizdQiMkI+0rb3Mjwz6rQsE6OBUMd0zWVJeGAO4kF4x6LoyF68M4uzjs7ZNL0vKtkPur2wVnDs7CfAhyVj7TYCTwPsv8gLaqNfwHbYpKDpQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=EtoNRA9O; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="EtoNRA9O" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2cec4226c70so45294885ad.1 for ; Thu, 06 Aug 2026 17:19:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786061950; x=1786666750; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=luhPZAd9yDKFqzqDYzNPoL8PFv2nmAeAhL8OrthRHBw=; b=EtoNRA9OAggCduAG9Zs/wa9x+wSFZWnIkUSF8vZ876stTcAYTjX0i+4/iEDNmQnXLP 9i89WFVY3ptaydplzU8smSahZoVDB/Px+YByYu/Dycm0eOjiDUqmnJOqo2QsljAk2BfW VwUo4EDuFVDzq4svkMdvOyPbNGS0w3xBzQTk+QWuq88Yw3zqI/kPS13vVZBwGhsTK9Q1 IYrTlL/YSJ59Sn8TV9ol387ilTY3wvi46FoO9jnGyaYLEwbuRns77nPNE5j+noLMg33Y TVStN3jZGlNE82dnaUsVLX6cbqAcm/S46SEmwt5NJT/dx1qHQPTyyDMctXr84r+rJ8tC kUXw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786061950; x=1786666750; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=luhPZAd9yDKFqzqDYzNPoL8PFv2nmAeAhL8OrthRHBw=; b=f7VqrzlhvZzUtTkMQHEZKMKiggoVT1UslpFqICyviMGuIqQ1k11mxMcH3sRptZonJe dxAZDum57CW/eFQq+0F368QOrU9dLHNhJB9/jAAPYmpH/68IJka2kZ2Z9T7CM/DQLbbn rmleh8FqserFMrlDY2XmkT2JH2LunTQfXhwooXxfklnfwPOX3O1R2SZMYDw6fjMAPX9C VELt1Q6t0xSGxd28Co4CCyupKRwbdEvD43plwqXLSlzbQCA4Rr8ahp6TZBmBDigBaJ8o /0xfYxaNbHF8S7BuorP8+0eYkONHP6Z3k9AEL9FNq0jubT3tQBeM1nZxwAkACvhdoZjQ aKzA== X-Forwarded-Encrypted: i=1; AHgh+Rp7NJx7tFq2wFagQub6l0+if8sZiGLrBOxgXtz5iWIkrAOJ9G+RWtXsdbUZGbcnZ2+bKUHMr19TO/pW3P4=@vger.kernel.org X-Gm-Message-State: AOJu0Yz2YONDDb4oZNAs/p6ef+dasIhu3KNw2zYcH7gomzG2PYUysqRx C1xa7aL90dBz5t3fr1DZk7rO/RkRe5gu977wHQCyMfvbtSEu2g30h42VPfjE21eIie744URYMf/ YJnqz9Q== X-Received: from plblg8.prod.google.com ([2002:a17:902:fb88:b0:2cc:61be:8bbb]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:f788:b0:2c9:d55d:2d3 with SMTP id d9443c01a7336-2d0ca7834d4mr237958385ad.15.1786061949754; Thu, 06 Aug 2026 17:19:09 -0700 (PDT) Date: Thu, 6 Aug 2026 17:19:08 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260726-page_alloc-unmapped-v3-0-6f5729aa9832@google.com> <20260726-page_alloc-unmapped-v3-3-6f5729aa9832@google.com> Message-ID: Subject: Re: [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP From: Sean Christopherson To: Yosry Ahmed Cc: Brendan Jackman , Brendan Jackman , Borislav Petkov , Dave Hansen , Peter Zijlstra , Andrew Morton , David Hildenbrand , Vlastimil Babka , Mike Rapoport , Wei Xu , Johannes Weiner , Zi Yan , Lorenzo Stoakes , linux-mm@kvack.org, linux-kernel@vger.kernel.org, x86@kernel.org, Sumit Garg , Will Deacon , rientjes@google.com, patrick.roy@linux.dev, Takahiro Itazuri , Andy Lutomirski , David Kaplan , Thomas Gleixner , Patrick Bellasi , Reiji Watanabe , Nikita Kalyazin , Ackerley Tng Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable On Thu, Aug 06, 2026, Yosry Ahmed wrote: > On Thu, Aug 6, 2026 at 5:02=E2=80=AFPM Sean Christopherson wrote: > > > > On Fri, Jul 31, 2026, Yosry Ahmed wrote: > The mermap provides infrastructure to do this, but I think the assumption= is > that mappings are short-lived. Keeping mappings for the lifetime of the V= M > would probably need more thought. For example, migration is disabled whil= e a > mapping is active (as mappings are per-CPU), and I was even proposing > disabling preemption (lol), so that wouldn't work at all with mappings th= at > are kept around for that long. The address space is also currently limit= ed > to 2M per CPU per process, which should be plenty but could also become a > limitation. They don't _need_ to be kept around that long, but I'm skeptical that on-de= mand mapping will actually work for KVM's use cases. > > 2. Always access guest memory through userspace mappings, i.e. through= uaccess. > > > > #2 sounds nice, but the problem is that it effectively requires hand-co= ded assembly > > sequences for anything more complex than basic load/store operations. = Which isn't > > a complete non-starter, but it's a pretty big blocker. E.g. see the me= ss that is > > record_steal_time(), and then imagine trying to convert something like > > nested_vmx_prepare_msr_bitmap() to use uaccess. > > > > So, unless someone comes up with a clever idea, KVM will need something= GUP-like. > > Strictly speaking, it doesn't necessarily need to be exactly GUP, becau= se KVM could > > poke into guest_memfd directly; KVM would "just" need to manually track= its own > > mappings. >=20 > Ideally we can have shared infrastructure for this (i.e. the mermap). >=20 > > But on x86 at least, that's not really a viable option because it only > > works for map-rarely, read/write-many use cases. For one-off accesses,= creating > > and destroying (very) shortlived mappings would be too costly, and so w= e'd want > > those to go through uaccess. >=20 > Not necessarily (I hope). I think the mermap pre-allocates page tables > (or some of them) and defers some TLB flushes, it's aimed at > short-lived mappings (e.g. for read() syscalls to map a file page, > copy to buffer, then unmap). I highly recommend testing shadow paging if you have aspirations of replaci= ng the get_user() in FNAME(walk_addr_generic) with an on-demand mapping. I wo= uld be (pleasantly) shocked if dynamic mappings can provide acceptable performa= nce. >=20 > > At that point, userspace is basically required to > > maintain mappings for all host-accessible guest memory, and if there ar= e userspace > > mappings, then not using GUP doesn't make much sense. > > > > Note, I called out x86 because x86 has the most extensive emulator and = shadow > > paging support, which is where the isolated, one-off accesses happen in= spades. > > Other architectures might be able to squeak by without userspace mappin= gs, at > > least for now. > > > > So, in all likelihood, KVM will want GUP. >=20 > Yeah I am thinking that the check here to disallow GUP completely for > unmapped pages is aggressive. Maybe it works for now if KVM does not > currently have any use cases for accessing guest_memfd memory. But if it = does > (or will very soon), we need to think more about it, otherwise > AS_NO_DIRECT_MAP is not really usable for guest_memfd. Since you said KVM > "will want" GUP, I assume it currently doesn't? Doesn't what? Have GUP? KVM heavily uses GUP, including for guest_memfd t= hat can be mapped into userspace.