From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pg1-f199.google.com (mail-pg1-f199.google.com [209.85.215.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 69D113806AF for ; Fri, 7 Aug 2026 18:49:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.199 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786128565; cv=none; b=ld4GLwuEOPQ80NDtmMIAuLOBcBqlk7/qdyRndje8ySyinhzMzRh9EyxLbA8BWk11sjJRMERqgAkycHZ7G3WoNWDN9eRqePvHXwjyCm0MvRn+nXrW7xQVw73W7TeGCGrjphhkuRsxbl0hTg/qcRaKXhCQUuBduqYHg2tiOgGGheY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786128565; c=relaxed/simple; bh=c7vN9hisBVxAREtlC5n6HFpjKI2IFIEavluo42hMK6Y=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=NqFfj/j4aeU5WHmNMxkzDnic0CSSKH2KI6EJSXiT0HudVmUCZ3dRYKoFXjf2pY19Q4+4KJvopLtK+NFwt03FSHLPGpnnlUUtz6AXYANs6aeWoMFCk8cRWB9slrtEskaA7XlWAMMflgD45kyTYp1yzLm2EMjJDBvSKJJX06NvIPw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=dze+2iPB; arc=none smtp.client-ip=209.85.215.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="dze+2iPB" Received: by mail-pg1-f199.google.com with SMTP id 41be03b00d2f7-cb733fc5024so5134277a12.2 for ; Fri, 07 Aug 2026 11:49:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786128564; x=1786733364; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=s/JtcKWinUL6x2ni4/Rf+XfZdlo3vH7HPHxYmzRfVVQ=; b=dze+2iPBR3GddpoYGF5kNY6V0/MxIKPlviZ9T0ObgZGEzxDlqpvHrA02qasQiNuSmR Qs5+xGRU7skDAJ3C7adMRt+xo3yC6JVGTSiY7dQ3clEf72Ihlmr5L2doTuFsLhRNHfnO I4h4VNa90z7LV7Ph3bHR9YPWJiTmIIhz1i02y54Ktfx/1zwe85lv97iBnjfUIOqIOoZo YRwGOpl1mwc83HM58uXgkYxmhj2IGgbCK3OGVQ3wU2lxnzrDCONZJQkQxuGXIif37x/j JMNdguNQS1tBiv45Xb0RIiAMJCtm1nZwfuzDHhuSynDNKWa8MIk7Ovl896adLrV2uqJi cLKg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786128564; x=1786733364; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=s/JtcKWinUL6x2ni4/Rf+XfZdlo3vH7HPHxYmzRfVVQ=; b=lTGz8GLWVTyMxBPX4ekZ+/zeZp18eyO6HJBhb0EtztnJxrvTZp8sbZ86lZSbCDYWXw QsKM+SbrhhSLl0R6TihvGV3sbjNt5apVk4mfvqW6yX1p7NwKhtN4TaBQbayvfPjGuMe3 +BgA8JJBDQVggYXzzGF7NccNCMbx1T7yx9FunvBBGpmK22NxJZzshY1dS1oO0sGRYZmt wq6jDeu4fTxxAIjDFK8tM2hDikcaDEiP6r36Jf2+ngVXi0Dr7eMyr9fl2UMvOosFi3ll 4Ly15i5qjDj6zqqcJ2jWhIoQKJRDXFe/IDYjkFeZShlO2/hkrxrh71heAdx//M2x/FWh xM6Q== X-Forwarded-Encrypted: i=1; AHgh+RpusOJewN3kfRBydQYupe+ILbONanPgQQOwssCRBBh0pddxm34bHboyce9lTOty8fWOBuCeRF37zX6/dwI=@vger.kernel.org X-Gm-Message-State: AOJu0YwdwSz5ACBF4gAZNKv1LyAbMVV5JQUFY7HcutHoy0Ntv2jYNFv5 Z0kimzVnPLJZ9uSC0f47bgR7Y5gwkgAMkS7CUz2TcuGGdw1InNqXuTWOtNq/p+PjeygAoP4xTMQ DNk85RQ== X-Received: from pgla5.prod.google.com ([2002:a63:b45:0:b0:cbe:948f:d62f]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:7f86:b0:3c3:a3fd:db0a with SMTP id adf61e73a8af0-3cb85e3be02mr26267852637.16.1786128563397; Fri, 07 Aug 2026 11:49:23 -0700 (PDT) Date: Fri, 7 Aug 2026 11:49:22 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260726-page_alloc-unmapped-v3-3-6f5729aa9832@google.com> Message-ID: Subject: Re: [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP From: Sean Christopherson To: Yosry Ahmed Cc: Brendan Jackman , Brendan Jackman , Borislav Petkov , Dave Hansen , Peter Zijlstra , Andrew Morton , David Hildenbrand , Vlastimil Babka , Mike Rapoport , Wei Xu , Johannes Weiner , Zi Yan , Lorenzo Stoakes , linux-mm@kvack.org, linux-kernel@vger.kernel.org, x86@kernel.org, Sumit Garg , Will Deacon , rientjes@google.com, patrick.roy@linux.dev, Takahiro Itazuri , Andy Lutomirski , David Kaplan , Thomas Gleixner , Patrick Bellasi , Reiji Watanabe , Nikita Kalyazin , Ackerley Tng Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable On Fri, Aug 07, 2026, Yosry Ahmed wrote: > On Fri, Aug 7, 2026 at 7:26=E2=80=AFAM Sean Christopherson wrote: > > > AS_NO_DIRECT_MAP will surely make it a bigger problem, but not a new = one :P > > > > Well, if it disallows GUP, that will be a new problem. >=20 > Yeah I think we should check here and allow GUP on unmapped pages > (more below). One thing that confuses me is that the > GUEST_MEMFD_FLAG_NO_DIRECT_MAP series [1] seems to also have this > check that disallows GUP. So I am not sure if KVM needs GUP to work > for guest_memfd now (then how does [1] work?) or it will need it to > work in the future? Well, there's a reason that series hasn't been merged. :-) I didn't get far enough to start looking at the GUP stuff, so I genuinely d= on't know if what it proposed is sane/correct. > [1]https://lore.kernel.org/all/20260410151746.61150-7-kalyazin@amazon.com= / >=20 > > > > > > > > At that point, userspace is basically required to > > > > > > maintain mappings for all host-accessible guest memory, and if = there are userspace > > > > > > mappings, then not using GUP doesn't make much sense. > > > > > > > > > > > > Note, I called out x86 because x86 has the most extensive emula= tor and shadow > > > > > > paging support, which is where the isolated, one-off accesses h= appen in spades. > > > > > > Other architectures might be able to squeak by without userspac= e mappings, at > > > > > > least for now. > > > > > > > > > > > > So, in all likelihood, KVM will want GUP. > > > > > > > > > > Yeah I am thinking that the check here to disallow GUP completely= for > > > > > unmapped pages is aggressive. Maybe it works for now if KVM does = not > > > > > currently have any use cases for accessing guest_memfd memory. Bu= t if it does > > > > > (or will very soon), we need to think more about it, otherwise > > > > > AS_NO_DIRECT_MAP is not really usable for guest_memfd. Since you = said KVM > > > > > "will want" GUP, I assume it currently doesn't? > > > > > > > > Doesn't what? Have GUP? KVM heavily uses GUP, including for guest= _memfd that > > > > can be mapped into userspace. > > > > > > Your wording made me think that KVM doesn't currently use GUP for > > > guest_memfd, but I was obviously wrong. So IIUC GUP needs to succeed = for > > > guest_memfd pages with AS_NO_DIRECT_MAP. > > > > Yes, though as I said early, it doesn't *have* to be exactly GUP, just = something > > GUP-like. E.g. it could be a new API, if that's easier/cleaner. What = I don't > > think is a good idea though is handling this entirely in KVM/guest_memf= d. >=20 > Just to clarify, you mean that GUP (or GUP-like) should work in terms of > pinning the page and handing KVM/guest_memfd the pfn/address, but not > actually making the page accessible or establishing mappings, right? No, I'm saying that whatever API the kernel provides needs to ensure there'= s a valid kernel mapping (or provide one as a return value). =20 > Looking at [2], seems like the consensus was that AS_NO_DIRECT_MAP > means folios are not in the direct map, and callers are responsible > for establishing the mappings (e.g. using the mermap). I'm fine with that direction, but in that case GUP _does_ need to be disall= owed. I.e. _if_ we allow GUP, then GUP itself needs to somehow ensure the direct = map is populated. If GUP is not allowed, then IMO the core kernel needs to pro= vide an API to get at "inaccessible" mappings. Or I suppose GUP could take a fl= ag that says "I pinky-swear not to try and access the memory via the direct ma= p". > [2]https://lore.kernel.org/all/aeemS2wm38Cm4qAf@google.com/ >=20 > > > > > To actually access the memory, I assume the guest_memfd side will nee= d to > > > handle this by either using ephemeral mappings (e.g. mermap), restori= ng and > > > zapping direct mappings, or using a userspace mapping. I suppose for = the > > > purposes of AS_NO_DIRECT_MAP core support we just need GUP to succeed= ? > > > > And establish a (ephemeral?) kernel mapping, because general users of G= UP will > > expect that they can access the physical memory through the direct map.= That's > > why I didn't want to handle any of this in KVM[*], the rules and handli= ng need > > to be kernel-wide. >=20 > See above, I am struggling to understand where you think establishing > mappings should lie. =20 Heh, I'm not surprised you're struggling, because I don't really have an op= inion on exactly who/what is responsible for establishing the mappings. What I c= are about at this point is not having guest_memfd itself provide a GUP-like API= : that needs to be a generic kernel API. > The current approach is that AS_NO_DIRECT_MAP just means folios are not > mapped, and users are responsible for establishing the mappings. I assume= you > agree with this (since you suggested this :P), but you don't want KVM to = do > this ad-hoc, but to have a library for it. >=20 > This library should be the mermap. I imagine (for e.g.) kvm_vcpu_map() > using the mermap under the hood if it knows the mappings do not exist > and using the mermap virtual address instead of the direct map > address. This only works (with the current implementation) if we can > disable migration (or even better, preemption) while a mapping is > active, which I imagine would be tricky or just not possible. >=20 > The alternative could be destroying and recreating the mappings when > the vCPU moves between CPUs, which is.. interesting :) >=20 > I imagine we don't have to sort all of this out now. For the purposes > of AS_NO_DIRECT_MAP (and secretmem AFAICT), we just need to provide a > facility to allocate unmapped pages. None of this is user-facing at > this point. >=20 > > > > [*] https://lore.kernel.org/all/aeemS2wm38Cm4qAf@google.com