From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B08B778F4A for ; Fri, 7 Aug 2026 00:02:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786060974; cv=none; b=hRhogOUhkLtW1fhyltujwkriIlQqiZ2HYVxN4QEGUhbl6rhVVLEiHQ/wpIYeYKNriZV7ZNkscmF24meu+tI4yoHUzNtpj/XKVNNmBgUf9c83iFTJYkBnP/Y9QRYv94QcDaHwQF4SC9z+eO/WCiYCUw+fLw0z+bW4n4Rtzqok7WE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786060974; c=relaxed/simple; bh=Kn71/F9MIOUCpocqLGGIvMq7Cfft/hBlR/g7zc3dqjQ=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=TE9uWtx/Lpt9UzWFRhcI8ZFl+xihuu798AhO8F6WUFFUIGLrI5G0VHcl5DvTDeKTR2OARclb2BeLkAu2AWl5EtrHoZILPeHXAB89iypdZ6mZHOYSLLwqPyVhkH/u+uWsgnHto/8HQrKaccjH77rs/+nahWa+RaWvLPm0/Utb+wI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=o08P0uOh; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="o08P0uOh" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2cc73f47bdcso45667585ad.3 for ; Thu, 06 Aug 2026 17:02:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786060972; x=1786665772; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=uHDxEiric74Ood8IcxlMwgt128fT9/N03w/mVBGilVY=; b=o08P0uOhKgVCpfRPGl0cuO6w4JmiVbjN/B7PaGEslXdDufcVf4A/XcYgc7QJ3qSIKH 7rg2ZhR/iaoOx9s1rMIUgbbGLS6tVEw44W9900klEWhvEOxzJ7MUMZBQ7PvizimTjRMh vvTdhHiWav6t/MEPA8v02oUBOr0WA1IPPwGXapcYi+o3C3/nNtOyOvL2mlb0ZrWC16tn iRzvHuhGWCWMYhxTAhevNhfy3JlN3uYGvOy5r0eHsxir201AvqoHsnbfGoGH2uwTe674 DY81BZgylTmDFhsc+2m2crYtobMD8veA2L+Pm/P7zyOgpD3+1REogv8fnPgfsLU3HWcW OOKQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786060972; x=1786665772; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=uHDxEiric74Ood8IcxlMwgt128fT9/N03w/mVBGilVY=; b=oCeRje21eeJozZpvib/MPP+44Vh2BBcnD8Z3BVbe58QyCpSfcLNVuBVAzpgZnN6+Gj OOf7+cbK8yEMZb+B7A/tF0q9bYSbFG+iju8yMiSuiRmQWyrmVsM8yhQXY60tqTMae9Ou eihSdIqzJyier8YsCv7Y5p3zL6KCfs72ZJb9VZ4CxJBZqNoxKP2y4qR4L5QSpjyhIVXG z4rGPWPjGSw2R5e0xjdaysroXOVM+Ut3rTRT02Bdx0RqG2LtVkxIVeNckV8E6CKqlYDK IEvC4HcXNETSQqdqM+wEcn+n74dxKYTAK2dGZCM0xNhqrg7UCLNZ9N5Q8UlRsOJG+Itd mi5Q== X-Forwarded-Encrypted: i=1; AHgh+Rp4DWMnhSVc4zEyG+PSuGGTMeuVdxL+OqfQtXy3nOefpkot43/+4dEukLx5baQcDKrpnQtp4xWuMKmYbhg=@vger.kernel.org X-Gm-Message-State: AOJu0YwxjGur86YksRhHLc8JIFnxWKVlxjaAq2pyysTTBHe4MEfG3mzF 6DF6XfS376kt++8/JxUYJL7jhCnBj+w4K2QDNmpE0rLM+rKxW6YqT1NOZ4KvksQhX68H7EBuJA7 gI67fdQ== X-Received: from pljf2.prod.google.com ([2002:a17:902:ff02:b0:2ca:f374:c62e]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:ce84:b0:2cf:9f0b:b562 with SMTP id d9443c01a7336-2d0ca9904b7mr198998005ad.22.1786060971371; Thu, 06 Aug 2026 17:02:51 -0700 (PDT) Date: Thu, 6 Aug 2026 17:02:50 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260726-page_alloc-unmapped-v3-0-6f5729aa9832@google.com> <20260726-page_alloc-unmapped-v3-3-6f5729aa9832@google.com> Message-ID: Subject: Re: [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP From: Sean Christopherson To: Yosry Ahmed Cc: Brendan Jackman , Brendan Jackman , Borislav Petkov , Dave Hansen , Peter Zijlstra , Andrew Morton , David Hildenbrand , Vlastimil Babka , Mike Rapoport , Wei Xu , Johannes Weiner , Zi Yan , Lorenzo Stoakes , linux-mm@kvack.org, linux-kernel@vger.kernel.org, x86@kernel.org, Sumit Garg , Will Deacon , rientjes@google.com, patrick.roy@linux.dev, Takahiro Itazuri , Andy Lutomirski , David Kaplan , Thomas Gleixner , Patrick Bellasi , Reiji Watanabe , Nikita Kalyazin , Ackerley Tng Content-Type: text/plain; charset="us-ascii" On Fri, Jul 31, 2026, Yosry Ahmed wrote: > > > [..] > > >> --- a/mm/gup.c > > >> +++ b/mm/gup.c > > >> @@ -11,7 +11,6 @@ > > >> #include > > >> #include > > >> #include > > >> -#include > > >> > > >> #include > > >> #include > > >> @@ -1216,7 +1215,7 @@ static int check_vma_flags(struct vm_area_struct *vma, unsigned long gup_flags) > > >> if ((gup_flags & FOLL_SPLIT_PMD) && is_vm_hugetlb_page(vma)) > > >> return -EOPNOTSUPP; > > >> > > >> - if (vma_is_secretmem(vma)) > > >> + if (vma_has_no_direct_map(vma)) > > > > > > Same here, and for GUP in general. For example, KVM uses kvm_vcpu_map() > > > to map guest memory and access it (e.g. when running nested > > > virtualization), which uses GUP under the hood AFAICT. So KVM will want > > > GUP to succeed, and probably create an ephemeral mapping as well. This has actually been discussed quite heavily, on multiple occassions. Once the direct map is obliterated, there are basically two options: 1. Establish ephemeral mappings (for a fairly loose definition of "ephemeral"; some of the mappings would likely exist for the lifetime of the VM). 2. Always access guest memory through userspace mappings, i.e. through uaccess. #2 sounds nice, but the problem is that it effectively requires hand-coded assembly sequences for anything more complex than basic load/store operations. Which isn't a complete non-starter, but it's a pretty big blocker. E.g. see the mess that is record_steal_time(), and then imagine trying to convert something like nested_vmx_prepare_msr_bitmap() to use uaccess. So, unless someone comes up with a clever idea, KVM will need something GUP-like. Strictly speaking, it doesn't necessarily need to be exactly GUP, because KVM could poke into guest_memfd directly; KVM would "just" need to manually track its own mappings. But on x86 at least, that's not really a viable option because it only works for map-rarely, read/write-many use cases. For one-off accesses, creating and destroying (very) shortlived mappings would be too costly, and so we'd want those to go through uaccess. At that point, userspace is basically required to maintain mappings for all host-accessible guest memory, and if there are userspace mappings, then not using GUP doesn't make much sense. Note, I called out x86 because x86 has the most extensive emulator and shadow paging support, which is where the isolated, one-off accesses happen in spades. Other architectures might be able to squeak by without userspace mappings, at least for now. So, in all likelihood, KVM will want GUP. P.S. In theory, there are other options. IIRC, arm64 hardware provides the ability to access memory via stage-2 page tables, but I also recall it being broken and/or having severe limitations. Using that also seems like it would defeat the purpose of nuking the direct map since the host would have a full view of guest memory, so long as it had the right "key". The other crazy hair super theoretical idea would be for KVM to hoist part of itself into guest context, but that would be a hilariously costly version of uaccess, and would never fly for CoCo VMs.