From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E24753C343A for ; Sun, 2 Aug 2026 16:27:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785688041; cv=none; b=FdELy7kaLrl3yFWnKljh3KZrZD0rsj7XAWaaugCZdbggFP6+K7Ohy59F5/QXYXDAK11Dn0zJPnuId3Txu01LZYxuP7JjElB1k/+UyHboXdoKktDiIDaGTxU8pWKX951rlMC792Y66IyzTolz+52+DeIe0F17T2FodDitTSzUQ2Y= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785688041; c=relaxed/simple; bh=04Z6dp3IDlz5+sFaZgLCtR3lrcC8EjJtCmmB8M8SArk=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Nqp1aEaKhQQHd7F20mVTetSaVoNCmmk7iL58YoR5X97pL0a/x8TLrLGrAbCHp6HKFdCdbOqTK8546fhHcUini2GK8CBNMhBxULAeW1j3mDM6sfoj19zObePkGMY9Rb5ayyhlopwuUy9Q7m5b5Q+gmfQMLx21x6Bj2hhcB22FOvI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=RbsnRO84; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="RbsnRO84" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 909551F000E9; Sun, 2 Aug 2026 16:27:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785688037; bh=QKjc3TaK187Ow1jc4cqwdroCBKbKjzdu63N3TTnN2as=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=RbsnRO84nLwR4FAQU8isvBanq+9uIEjaWVHqRGSnEcSY5h79Hm8Q2OLK4170/JC/B JHQ3uYGTUGHGwlbeMU1gDQxxBSFnidulo0wBL0l8XxXR502oxxhffoYmf8ArRz0DLK fGFPU8X8f42GYccZWwF1HRZSK3uVAy/gEPvC4MsP6KbaML/pIXhV7wGTwHK89JVUVf pjSqvAzVOIy3vTxEbPPOfS6XSdxyuhNHnuRZ3WNMoB0NTA+LQlh8edg/OI3AqhLnfe yz8YAGrmD/gdPAG/7XR/UvqMLRFJFHZzw8kcGLRXlYiAB8BmEyj6zpGKU6XMyInr4J LwTM9v9h7Me5w== Date: Sun, 2 Aug 2026 19:27:05 +0300 From: Mike Rapoport To: Brendan Jackman Cc: Borislav Petkov , Dave Hansen , Peter Zijlstra , Andrew Morton , David Hildenbrand , Vlastimil Babka , Wei Xu , Johannes Weiner , Zi Yan , Lorenzo Stoakes , linux-mm@kvack.org, linux-kernel@vger.kernel.org, x86@kernel.org, Sumit Garg , Will Deacon , rientjes@google.com, "Kalyazin, Nikita" , patrick.roy@linux.dev, "Itazuri, Takahiro" , Andy Lutomirski , David Kaplan , Thomas Gleixner , Yosry Ahmed , Patrick Bellasi , Reiji Watanabe , Sean Christopherson Subject: Re: [PATCH v3 07/26] x86/mm: introduce mm-local region Message-ID: References: <20260726-page_alloc-unmapped-v3-0-6f5729aa9832@google.com> <20260726-page_alloc-unmapped-v3-7-6f5729aa9832@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260726-page_alloc-unmapped-v3-7-6f5729aa9832@google.com> On Sun, Jul 26, 2026 at 10:22:40PM +0000, Brendan Jackman wrote: > Various security features benefit from having process-local address > mappings within the kernel. Examples include no-direct-map guest_memfd > [2] and significant optimizations for ASI [1]. > > With the currently envisaged usecases, there will be many situations > where almost no processes have any need for the mm-local region. > Therefore, avoid its overhead (memory cost of pagetables, alloc/free > overhead during fork/exit) for processes that don't use it by requiring > its users to explicitly initialize it via the new mm_local_* API. > > As pointed out by Andy in [0], x86 already has a PGD entry that is local > to the mm, which is used for the LDT. In a subsequent patch, the LDT > remap will be unified with the general mm-local region, but to help > keep the patch to a manageable size, first just introduce the mm-local > region. > > On 64-bit, give the mm-local region a whole PGD. On 32-bit, just give it > one PMD. No investigation has been done into whether it's feasible to > expand the region on 32-bit. Most likely there is no strong usecase for > that anyway. > > In order to combine the need for an on-demand mm initialisation, with > the desire to transparently handle propagating mappings to userspace > under KPTI, the user and kernel pagetables are shared at the highest > level possible. For PAE that means the PTE table is shared and for > 64-bit the P4D/PUD. This is implemented by pre-allocating the first > shared table when the mm-local region is first initialised. > > [0] https://lore.kernel.org/linux-mm/CALCETrXHbS9VXfZ80kOjiTrreM2EbapYeGp68mvJPbosUtorYA@mail.gmail.com/ > [1] https://linuxasi.dev/ > [2] https://lore.kernel.org/all/20250924151101.2225820-1-patrick.roy@campus.lmu.de > Signed-off-by: Brendan Jackman > --- > arch/x86/Kconfig | 2 + > arch/x86/include/asm/mmu_context.h | 118 +++++++++++++++++++++++++++++++- > arch/x86/include/asm/pgtable_32_areas.h | 9 ++- > arch/x86/include/asm/pgtable_64_types.h | 12 +++- > arch/x86/mm/pgtable.c | 3 + > include/linux/mm.h | 13 ++++ > include/linux/mm_types.h | 2 + > kernel/fork.c | 1 + > mm/Kconfig | 7 ++ > 9 files changed, 161 insertions(+), 6 deletions(-) > > diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig > index fb298e2191792..3efab3524a6cf 100644 > --- a/arch/x86/Kconfig > +++ b/arch/x86/Kconfig > @@ -132,6 +132,8 @@ config X86 > select ARCH_SUPPORTS_LTO_CLANG > select ARCH_SUPPORTS_LTO_CLANG_THIN > select ARCH_SUPPORTS_RT > + # LDT remap temporarily clashes with mm-local region, can't have both. > + select ARCH_SUPPORTS_MM_LOCAL_REGION if X86_64 || X86_PAE && !MODIFY_LDT_SYSCALL > select ARCH_USE_BUILTIN_BSWAP > select ARCH_USE_CMPXCHG_LOCKREF > select ARCH_USE_MEMTEST > diff --git a/arch/x86/include/asm/mmu_context.h b/arch/x86/include/asm/mmu_context.h > index ef5b507de34e2..3d4f54673014f 100644 > --- a/arch/x86/include/asm/mmu_context.h > +++ b/arch/x86/include/asm/mmu_context.h > @@ -8,8 +8,10 @@ > > #include > > +#include > #include > #include > +#include > #include > #include > #include > @@ -223,10 +225,124 @@ static inline int arch_dup_mmap(struct mm_struct *oldmm, struct mm_struct *mm) > return ldt_dup_context(oldmm, mm); > } > > +#ifdef CONFIG_MM_LOCAL_REGION > +static inline void mm_local_region_free(struct mm_struct *mm) > +{ > + if (!mm_local_region_used(mm)) > + return; > + > + struct mmu_gather tlb; > + unsigned long start = MM_LOCAL_BASE_ADDR; > + unsigned long end = MM_LOCAL_END_ADDR; > + > + /* > + * Although free_pgd_range() is intended for freeing user > + * page-tables, it also works out for kernel mappings on x86. > + * Use tlb_gather_mmu_fullmm() to avoid confusing the > + * range-tracking logic in __tlb_adjust_range(). > + */ > + tlb_gather_mmu_fullmm(&tlb, mm); > + free_pgd_range(&tlb, start, end, start, end); > + tlb_finish_mmu(&tlb); > + > + mm_flags_clear(MMF_LOCAL_REGION_USED, mm); > +} > + > +#if defined(CONFIG_MITIGATION_PAGE_TABLE_ISOLATION) && defined(CONFIG_X86_PAE) > +static inline pmd_t *pgd_to_pmd_walk(pgd_t *pgd, unsigned long va) There's very similar mm_find_pmd() in mm/rmap.c and I bet a bunch of other places walk from PGD to PMD and return PMD in the end. Can we put this function into, say, mm/pgtable-generic.c? Finding all the places that do such walk and sticking it there should not be a part of this set IMNHO, but having it in the generic code is a good start for a future cleanup. > +{ > + p4d_t *p4d; > + pud_t *pud; > + > + if (pgd->pgd == 0) > + return NULL; > + > + p4d = p4d_offset(pgd, va); > + if (p4d_none(*p4d)) > + return NULL; > + > + pud = pud_offset(p4d, va); > + if (pud_none(*pud)) > + return NULL; > + > + return pmd_offset(pud, va); > +} -- Sincerely yours, Mike.