From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 877292566D3 for ; Mon, 26 Jan 2026 18:58:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769453917; cv=none; b=lSM1k7yLmxd689VJvR40tubtY5QOf6sR+CrtxImUXryYBM/Ptzy180svJJD874onvJZK11L02cMAG15ZGBS3Jv+08h28DWzzSBvRvhfVvO8xdfFwGRIuW1uHV7E+sPjznhZGoOFF+wXquJ0+V8BVG/B6r+j2PNIxe7v4at7IuoU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769453917; c=relaxed/simple; bh=r4iiZ79RtTH49O1joz4UwCgZlczdoapI1lsvznIqdhU=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=RxRIbIUKkjHB0yKhXJxd7ViRt2l4Da7upUrMYbR2tBYAwYcwp02XENdGbUODDWy3yM2y0/onqgEenXxv2cFnI3z18HJsn2na6hBrhCRCa8f7vucreRbRJuxjn/rg7+beQzjiGC2GhadeLlxMEFW1TunvlWGclVm/l65vDf3j/Do= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=uWN8oyXw; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="uWN8oyXw" Received: by smtp.kernel.org (Postfix) with ESMTPSA id C6474C19422; Mon, 26 Jan 2026 18:58:35 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1769453917; bh=r4iiZ79RtTH49O1joz4UwCgZlczdoapI1lsvznIqdhU=; h=Date:From:To:Cc:Subject:References:In-Reply-To:From; b=uWN8oyXwiJiDrjwBhvNPkBSk0kfqJORR436xEPqCOQ3npm6jp1ricUxdpa6ZTnpMs o/dAx81J7vVsDTKeVuwVJ8VgH/VKEAYPbKvdvxaydi2Jn0ZyseaqR7ivs9UEw3ipr4 +7a7QzuH8HJ4j/LvYHtqf8CnaZT9fLYtgHD7o/vTb6bp90pfUpf74wZWqYmmBtxxIU my12qrqfywJxGdIm0syUPzeLoWlcW/d7NnmB1e2F3SRxJwURnJeP6aJZsuvPnBPXC5 NyYge/y1vePFB7n550rkG28fSlQ1UCsT9zUNMlBYlvHeSYUWMd1ncHM6dKTzLoP5z1 FRY6PC2tJNRVg== Date: Mon, 26 Jan 2026 18:58:32 +0000 From: Will Deacon To: Yang Shi Cc: Ryan Roberts , catalin.marinas@arm.com, cl@gentwo.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org Subject: Re: [v5 PATCH] arm64: mm: show direct mapping use in /proc/meminfo Message-ID: References: <20260107002944.2940963-1-yang@os.amperecomputing.com> <70b37582-fd26-4d23-a6d6-9a98e3f2eecf@os.amperecomputing.com> <04f816f9-5533-4fe1-99b0-cd405caac485@os.amperecomputing.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=iso-8859-1 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <04f816f9-5533-4fe1-99b0-cd405caac485@os.amperecomputing.com> On Mon, Jan 26, 2026 at 09:55:06AM -0800, Yang Shi wrote: > > > On 1/26/26 6:14 AM, Will Deacon wrote: > > On Thu, Jan 22, 2026 at 01:59:54PM -0800, Yang Shi wrote: > > > On 1/22/26 6:43 AM, Ryan Roberts wrote: > > > > On 21/01/2026 22:44, Yang Shi wrote: > > > > > On 1/21/26 9:23 AM, Ryan Roberts wrote: > > > > But it looks like all the higher level users will only ever unplug in the same > > > > granularity that was plugged in (I might be wrong but that's the sense I get). > > > > > > > > arm64 adds the constraint that it won't unplug any memory that was present at > > > > boot - see prevent_bootmem_remove_notifier(). > > > > > > > > So in practice this is probably safe, though perhaps brittle. > > > > > > > > Some options: > > > > > > > > - leave it as is and worry about it if/when something shifts and hits the > > > > problem. > > > Seems like the most simple way :-) > > > > > > > - Enhance prevent_bootmem_remove_notifier() to reject unplugging memory blocks > > > > whose boundaries are within leaf mappings. > > > I don't quite get why we should enhance prevent_bootmem_remove_notifier(). > > > If I read the code correctly, it just simply reject offline boot memory. > > > Offlining a single memory block is fine. If you check the boundaries there, > > > will it prevent from offlining a single memory block? > > > > > > I think you need enhance try_remove_memory(). But kernel may unmap linear > > > mapping by memory blocks if altmap is used. So you should need an extra page > > > table walk with the start and the size of unplugged dimm before removing the > > > memory to tell whether the boundaries are within leaf mappings or not IIUC. > > > Can it be done in arch_remove_memory()? It seems not because > > > arch_remove_memory() may be called on memory block granularity if altmap is > > > used. > > > > > > > - For non-bbml2_noabort systems, map hotplug memory with a new flag to ensure > > > > that leaf mappings are always <= memory_block_size_bytes(). For > > > > bbml2_noabort, split at the block boundaries before doing the unmapping. > > > The linear mapping will be at most 128M (4K page size), it sounds sub > > > optimal IMHO. > > > > > > > Given I don't think this can happen in practice, probably the middle option is > > > > the best? There is no runtime impact and it will give us a warning if it ever > > > > does happen in future. > > > > > > > > What do you think? > > > I agree it can't happen in practice, so why not just take option #1 given > > > the complexity added by option #2? > > It still looks broken in the case that a region that was mapped with the > > contiguous bit is then unmapped. The sequence seems to iterate over > > each contiguous PTE, zapping the entry and doing the TLBI while the > > other entries in the contiguous range remain intact. I don't think > > that's sufficient to guarantee that you don't have stale TLB entries > > once you've finished processing the whole range. > > > > For example, imagine you have an L1 TLB that only supports 4k entries > > and an L2 TLB that supports 64k entries. Let's say that the contiguous > > range is mapped by pte0 ... pte15 and we've zapped and invalidated > > pte0 ... pte14. At that point, I think the hardware is permitted to use > > the last remaining contiguous pte (pte15) to allocate a 64k entry in the > > L2 TLB covering the whole range. A (speculative) walk via one of the > > virtual addresses translated by pte0 ... pte14 could then hit that entry > > and fill a 4k entry into the L1 TLB. So at the end of the sequence, you > > could presumably still access the first 60k of the range thanks to stale > > entries in the L1 TLB? > > It is a little bit hard for me to understand how come a (speculative) walk > could happen when we reach here. > > Before we reach here, IIUC kernel has: > >  * offlined all the page blocks. It means they are freed and isolated from > buddy allocator, even pfn walk (for example, compaction) should not reach > them at all. >  * vmemmap has been eliminated. So no struct page available. > > From kernel point of view, they are nonreachable now. Did I miss and/or > misunderstand something? I'm talking about hardware speculation. It's mapped as normal memory so the CPU can speculate from it. We can't really reason about the bounds of that, especially in a world with branch predictors and history-based prefetchers. Will