From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ed1-f48.google.com (mail-ed1-f48.google.com [209.85.208.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0A083298CAB for ; Wed, 15 Apr 2026 02:30:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.48 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776220220; cv=none; b=lg5TzYvogLjm2ujlVyqThDTIb4Mvs/4cjAGFWhFgE75TglmJfjTVii4xqF2qQDGVtEK7uoz7m4HecoK3Eh0XgijqrbGREsNVq1gO5WryaAdCYZldhs9Kx1DMYaj9VZcz2182CfR8FXhG1va3rj260BckT8eFXUbCNmtHYhfoSKA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1776220220; c=relaxed/simple; bh=wjMIKO7jZ4s+pzgYNVS8ZiV+DeKblc4caHoL5/pXdjo=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=TNWe3iZzX9zpzbzA3sFv52jQRyqZLswc2pCr9mT/PjE3rmAxc0d4bVe1OHCug1kpt8flJ4NsxE7rtp+/chvfxuTl6qldRPCNdjsSwzJxudeIdBc5cJ7G4jdBPRyl/D8avlrv7Oo1jdwW0l+RPaQw11LdAa1FuwauG3lzPPSFzBo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BSFSF0o5; arc=none smtp.client-ip=209.85.208.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BSFSF0o5" Received: by mail-ed1-f48.google.com with SMTP id 4fb4d7f45d1cf-671ff4b716cso2166538a12.1 for ; Tue, 14 Apr 2026 19:30:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1776220217; x=1776825017; darn=vger.kernel.org; h=user-agent:in-reply-to:content-disposition:mime-version:references :reply-to:message-id:subject:cc:to:from:date:from:to:cc:subject:date :message-id:reply-to; bh=StC25PVDWT8M7W9NSgY17XCq6J9v4q/RTTNrlPmFksY=; b=BSFSF0o5Md85XwUL6TWjTpq98l3MNyHkkHseO/NesCWD2osfaSJZpWLBn2MCZm5WX8 aB6ej+IWh0NyLW2n8+LEOnqc5k1s9PTnpKpX7j9TgIh1yaIs3/kMT8h05XMFC0Q7hXrZ 3aoGdoiFgviHyGzGZIOd0spHY3W4f9DqJOXIQebVEHAMgRs7Gyvygofbp5/7YZJ01atf hjsBfi+KfIKXYC8g4anJTZMRX99yGNoMk43hK8ejYpd2lMVUvURc1slh1V51mY5sybOa JBwFhPWDt6t4CVK91bvfZ+jBxmtGR3gwFXzn8NW42pHJQ5N2BCOAHjjcm01WmRCwmXvT PXdA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1776220217; x=1776825017; h=user-agent:in-reply-to:content-disposition:mime-version:references :reply-to:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=StC25PVDWT8M7W9NSgY17XCq6J9v4q/RTTNrlPmFksY=; b=rzeCaM22PiG863y9vChirs0J/1Y+O5tRuJHEdSR2Pf/xURAJovjZWdH2xU101SmswE fpHHxysoBhNduyLwoq2ygJ3CeC7WYItfqlciMm7/Eod1Ik5H81ZcjF+0O0mnPAHvh8D7 BRkYoKzq0lSgENzWThtPwKF7bQh7Y7SxKHIaW9x3tg3JrZOCEs2tDFQ2u0ifDbPO/3je 346yESlVUnD9Vc/M+Op2Hziqq+Dj18NXOrqyaAeQ7bTquEAt55EgXqM41VzUW3P4jk// 0BhjwOkz54kPtyVBgsvrP6S0FjVdJdA5jIeWtM+usS9tWe3X7YCdo8qdMtDadFN0EbZH eBKg== X-Forwarded-Encrypted: i=1; AFNElJ/YDCFDPApt8Hi0Ko9dhmDA4RcrYCXzxJmL+oU5vjspGD2JEdw2OJ6tR0mS1u0bEYdlMUPZ4rLWpRDp9xM=@vger.kernel.org X-Gm-Message-State: AOJu0Yy2+P5jHCeM+dYtQTURoVAdrF4I+zVMqJrJqmGI2dgyTfmCqxd0 CZ5FEAeI55nFnqlNrzyXUrfcaKXUGVmm73L23iHIEe6X3FJQ4ojrxx5K X-Gm-Gg: AeBDies/NKWfrR3tlph3QUgKpLkrfdaAl1jWYQLlEDV301NXpFwAgGEdwRGi6h7OA7W uABvSrOBDy42hpRirV3miCkgGmJHeNIM1UMCQT5HeXzpzD4ZHthNFEXkUY3u8RQ6iH5bexlTovq 5NjdjBsx9ZvzZVukcJRr2tcUk37uHsHHKhLq4sesnmMDQRPLlLbgmxM4mZNeI8OFGmtTA3ghQa5 Zxn/J+FmsuMSlbAeLo7i2RKVHFDJGISRz+Th4eiwbayy53dxaSdPipG5U4qC++KRNvvCSidTx7f tE7WSyUHJdeu82S3wlhg/PFfV7naD+4F6+ov0lL55vNhEbuobgHja0kjYNGrJjhoHCei6taNhLt +5VNFYjA08uP7LCPvEJXkuiN+ze2yEYytwTz82F4s12PLi3/9mgt5SApKEQCYmivyDWy12OQVJ2 3tkk0W2Oael+psbL0ezSf2Fdrb3mnonoSa X-Received: by 2002:a17:907:25cc:b0:b9c:aae6:907e with SMTP id a640c23a62f3a-b9d7260da0emr1116420366b.13.1776220217053; Tue, 14 Apr 2026 19:30:17 -0700 (PDT) Received: from localhost ([185.92.221.13]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-ba1784a808fsm7816466b.61.2026.04.14.19.30.15 (version=TLS1_2 cipher=ECDHE-ECDSA-CHACHA20-POLY1305 bits=256/256); Tue, 14 Apr 2026 19:30:15 -0700 (PDT) Date: Wed, 15 Apr 2026 02:30:15 +0000 From: Wei Yang To: "David Hildenbrand (Arm)" Cc: Wei Yang , Yuan Liu , Oscar Salvador , Mike Rapoport , linux-mm@kvack.org, Yong Hu , Nanhai Zou , Tim Chen , Qiuxu Zhuo , Yu C Chen , Pan Deng , Tianyou Li , Chen Zhang , linux-kernel@vger.kernel.org Subject: Re: [PATCH v3] mm/memory hotplug/unplug: Optimize zone contiguous check when changing pfn range Message-ID: <20260415023015.5i7hrixrk64uzjqj@master> Reply-To: Wei Yang References: <20260408031615.1831922-1-yuan1.liu@intel.com> <20260413130633.knzkliyqvjhuz2kd@master> <1928b6b0-2ec3-43ca-a41b-e880d974af04@kernel.org> <20260414021219.wayysugpfbzirzh6@master> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: User-Agent: NeoMutt/20170113 (1.7.2) On Tue, Apr 14, 2026 at 11:32:13AM +0200, David Hildenbrand (Arm) wrote: >On 4/14/26 04:12, Wei Yang wrote: >> On Mon, Apr 13, 2026 at 08:24:05PM +0200, David Hildenbrand (Arm) wrote: >>>> With the last memblock region fits in Node 1 Zone Normal. >>>> >>>> Then I punch a hole in this region with 2M(subsection) size with following >>>> change, to mimic there is a hole in memory range: >>>> >>>> @@ -1372,5 +1372,8 @@ __init void e820__memblock_setup(void) >>>> /* Throw away partial pages: */ >>>> memblock_trim_memory(PAGE_SIZE); >>>> >>>> + memblock_remove(0x140000000, 0x200000); >>>> + >>>> memblock_dump_all(); >>>> } >>>> >>>> Then the memblock dump shows: >>>> >>>> MEMBLOCK configuration: >>>> memory size = 0x000000017fd7dc00 reserved size = 0x0000000005a97 9c2 >>>> memory.cnt = 0x4 >>>> memory[0x0] [0x0000000000001000-0x000000000009efff], 0x000000000009e000 bytes on node 0 flags: 0x0 >>>> memory[0x1] [0x0000000000100000-0x00000000bffdefff], 0x00000000bfedf000 bytes on node 0 flags: 0x0 >>>> +- memory[0x2] [0x0000000100000000-0x000000013fffffff], 0x0000000040000000 bytes on node 1 flags: 0x0 >>>> +- memory[0x3] [0x0000000140200000-0x00000001bfffffff], 0x000000007fe00000 bytes on node 1 flags: 0x0 >>>> >>>> We can see the original one memblock region is divided into two, with a hole >>>> of 2M in the middle. >>> >>> Yes, that makes sense. >>> >>>> >>>> Not sure this is a reasonable mimic of memory hole. Also I tried to >>>> punch a larger hole, e.g. 10M, still see the behavioral change. >>>> >>>> The /proc/zoneinfo result: >>>> >>>> w/o patch >>>> >>>> Node 1, zone Normal >>>> pages free 469271 >>>> boost 0 >>>> min 8567 >>>> low 10708 >>>> high 12849 >>>> promo 14990 >>>> spanned 786432 >>>> present 785920 >>>> contigu 0 <--- zone is non-contiguous >>>> managed 766024 >>>> cma 0 >>>> >>>> with patch >>>> >>>> Node 1, zone Normal >>>> pages free 121098 >>>> boost 0 >>>> min 8665 >>>> low 10831 >>>> high 12997 >>>> promo 15163 >>>> spanned 786432 >>>> present 785920 >>>> contigu 1 <--- zone is contiguous >>>> managed 773041 >>>> cma 0 >>>> >>>> This shows we treat Node 1 Zone Normal as non-contiguous before, but treat >>>> it a contiguous zone after this patch. >>>> >>>> Reason: >>>> >>>> set_zone_contiguous() >>>> __pageblock_pfn_to_page() >>>> pfn_to_online_page() >>>> pfn_section_valid() <--- check subsection >>>> >>>> When SPARSEMEM_VMEMMEP is set, pfn_section_valid() checks subsection bit to >>>> decide if it is valid. For a hole, the corresponding bit is not set. So it >>>> is non-contiguous before the patch. >>>> >>>> After this patch, the memory map in this hole also contributes to >>>> pages_with_online_memmap, so it is treated as contiguous. >>> >>> That means that mm init code actually initialized a memmap, so there is >>> a memmap there that is properly initialized? >>> >>> So init_unavailable_range()->for_each_valid_pfn() processed these >>> sub-section holes I guess. >>> >> >> Yes, I think so. >> >> When memmap_init()->for_each_mem_pfn_range() iterate on the last memblock >> region, init_unavailable_range() will init the hole. >> >>> subsection_map_init() takes care of initializing the subsections. That >>> happens before memmap_init() in free_area_init(). >>> >> >> Yes. I guess you mean sparse_init_subsection_map(). >> >>> Is there a problem in for_each_valid_pfn()? >>> >>> And I think there is in first_valid_pfn: >>> >> >> You mean there is a problem in first_valid_pfn? > >That is my theory. > >> >>> if (valid_section(ms) && >>> (early_section(ms) || pfn_section_first_valid(ms, &pfn))) { >>> rcu_read_unlock_sched(); >>> return pfn; >>> } >>> >>> The PFN is valid, but we actually care about whether it will be online. >>> So likely, we should skip over sub-sections here also for early sections >>> (even though the memmap exist, nobody should be looking at it, just like >>> for an offline memory section). >>> >> >> And it should be like below? >> >> if (valid_section(ms) && >> pfn_section_first_valid(ms, &pfn)) { >> rcu_read_unlock_sched(); >> return pfn; >> } > >Probably, yes. We have to understand if other users would be negatively >affected. > >> >> IIUC, this would skip hole and leave allocated memory map uninitialized. And >> then those pages won't contribute to pages_with_online_memmap, which further >> leave the zone non-contiguous. > >Yes. > >> >> But we want zone to be contiguous when we have a hole like this, right? > >Not if a subsection is marked invalid. That's why the existing scenario >is that it will not be contiguous. > Let me try to understand. During previous discussion[1], we want to define "zone->contiguous guarantee pfn_to_page() is valid on the complete zone". So this definition is not true now, since we found the behavioral change. Because __pageblock_pfn_to_page() won't treat range with hole as contiguous, detected by pfn_to_online_page() on invalid subsection. [1]: https://lkml.org/lkml/2026/2/9/550 Now you are thinking the problem is in the iteration function, for_each_valid_pfn(), used in init_unavaiable_range(). It should only take valid subsections into consideration, even for early sections. So your concern is if we change first_valid_pfn(), it will affect other users. If the above understanding is correct, maybe we can use spanned == present to do the trick? Because holes are marked subsection invalid and holes are counted into absent. But I see the mirrored_kernel thing, not fully understand yet. This is the reason to prevent spanned == present approach? >Note that Wei reported that it was not contiguous but would now be >contiguous. > > >If you have a DAX device the plugs into that hole through >memremap_pages()->pagemap_range(), I think this could cause problems. I >doubt that this would happen in practice for such small holes, but if >they would be bigger, or at the start/end of the range, it could be >problematic. > Not fully understand yet, will take a look into it. >-- >Cheers, > >David -- Wei Yang Help you, Help me