From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qv1-f49.google.com (mail-qv1-f49.google.com [209.85.219.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4C68C2E62C0 for ; Thu, 18 Dec 2025 20:43:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.49 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766090583; cv=none; b=bUO8b4Ta/kk8BgC+nMMbzo4MlyUE93WU2TC4WbrEWV5h+0LW5wcT6GexqRASCBMKfGAp5bu5Bc6DEZd70jaSduB1GSHCt27WxkNBOmWT6eHXn68e5E6RuW8R4pyWN4d363Z50yzx5JsY4CnDoLLlFh4MRV2lb1XXIOWcZMbNiRE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766090583; c=relaxed/simple; bh=V+mjMbT6LR2sUNPDMLBfImTWBeQpVVdCsPD1mhB5fI0=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=V2beI534XwshGye4/OszOCW3BP6KbE+TQYGrE0YuddL4q/V9Y0T9zzb1LviK2yTmiZKUm+M/GYwQ92L2+bwIUmRA962nPC2YY8QpeY9RnDzoNV5xNFJl8F+33BcGmPGmIS2n1psogTUd6G+0ZkS+fFmLXXhdaXuRoTM68tpAev0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=pfsPI0fu; arc=none smtp.client-ip=209.85.219.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="pfsPI0fu" Received: by mail-qv1-f49.google.com with SMTP id 6a1803df08f44-88a37cb5afdso25711666d6.0 for ; Thu, 18 Dec 2025 12:43:01 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1766090580; x=1766695380; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=SJtWtYdyXJo9JjOuJovQNKlx0p9HWjikwGNiH4lGFaI=; b=pfsPI0fuiKJ6Xq/sok/fr/djZIUjR8U/fXzRDzZ4yg0O8tKzfRgImvALbWkksDKUTY hHFxSnwL4JNhJsN7fgrzmvW2pfNbKLBgN9vciLu+qM4jjvpFzaDRkAMn/I/xmOqIwhVk k8rQJWSXhpLAlw4ENWvCnBWMbe3HLVtaB/e1tZ49GnkSD4H2EMFZT+kFfDPkgd275pvR 51oQMIfdTPonErj0NCjIMsV22I2+D5wG+Boj1FlUiv/lgQOvb/+KKNYRJxPh3sJVaGIm iAW7f/CkC8kqeFWUkzJarbf9G8ucVOcf07+fdNu7atvsyhmtq27URvKU8ZEt0KOoalCy Om/g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1766090580; x=1766695380; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=SJtWtYdyXJo9JjOuJovQNKlx0p9HWjikwGNiH4lGFaI=; b=Zuuw7oenH+wXx3gtgr5lfq3P5SIc8/tDKrdio4iwyDuwAqONVnar1UEYpVm8Uxsw5N f3MP1/pEm/W7nAuD+IKmH0IkyJmY+AOB/QKxeydDBXJyjy+XKtAHcD6G52uknsbgAK+i 1u84aq+HkiOmfs7tnAvf/b0GsMQUNuP+yZlYvRt18gKR0FEeWBNE/aqLUnKd7bww09fX T2KnqJq1BVkF0/y2vIfPATZ8lsFaBlH7uEjfCk7vnsrfNjHb9EvOP2JHeplG6XP0tpbb XMqDUYj8Yj3ja2sttOwBov+y63hPAb+pQcHWIjDO6trOkGlOOd+1l/Sfs9pptl+SdgLA O5GQ== X-Forwarded-Encrypted: i=1; AJvYcCWwxU9qTb+RnZxQEqrenY1je9q6QBPJxt/ehVSntOF/GBeVSyJk6QvhpNKMdRsE98ZZsYynScomPqdKykw=@vger.kernel.org X-Gm-Message-State: AOJu0Yw2f680nS5svIrUOTEN6xr0OClRJaGtUIBsMNPC216I/1yDwwTw VaWYSYR0cCX/XShB/QbIzR/Y/TDyHubEaDrzEGUrZu188gHtyyJZ/UnbQl6OTIM8GrI= X-Gm-Gg: AY/fxX7Oi3UFOLhG9rHHhbsgRwzcR4pW6vNohD8MOppgoNsjp2A9BAN/IQfc92lIEfd IMTXgt2mjSJEb+scnRjcRo+1qKrR4iiuAYODkvwAVHfo8OyXMXg4iqalJdpuaz5naJMG3XKXt16 ou/GomINWbdOaOR2p+7SJND0IWdpOiDaB5DJluYpkwFqcepjha7L8uJCZnIgjU1PT2zmudtxs5N uLBaSFb4dH0at5xGlos54nX72n78vKOdknpFCiIwikghCQCfFQabFRi/k8X2bTpvyn7N3FGqPB6 n7H8RnGz55CndaO2MwbwtVoFN585TkWIp3H67LBD2W0xux8NTjGR+cQIVv9m6r5F+tZYKMypvXI +YAYI8QX8Sc5ShAOymtDsiDLVTmV7sH53chjQBeKuThG/YU305ufsGBjapkpuQhEg1NYd5Vuy2q VfMFM2PioI9UAF30J27Iu5ykU8bgENh7Qo4GBKnC+uHnrDxLyX8e1ELAKc2GvKXW7/7Akj1A== X-Google-Smtp-Source: AGHT+IHe+/76VfBRWfz9lM8LGjmwX7vteYD1kOK/9DSfcBkIo4Wnhl65RcEQlUT7EV8J2ngMFpIHqA== X-Received: by 2002:a05:6214:413:b0:88a:31e5:80fa with SMTP id 6a1803df08f44-88c50bdb5d3mr73839626d6.16.1766090580209; Thu, 18 Dec 2025 12:43:00 -0800 (PST) Received: from gourry-fedora-PF4VCD3F (pool-96-255-20-138.washdc.ftas.verizon.net. [96.255.20.138]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-88d9623fdfdsm3934746d6.5.2025.12.18.12.42.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 18 Dec 2025 12:42:59 -0800 (PST) Date: Thu, 18 Dec 2025 15:42:22 -0500 From: Gregory Price To: Zi Yan Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, vbabka@suse.cz, surenb@google.com, mhocko@suse.com, jackmanb@google.com, hannes@cmpxchg.org, richard.weiyang@gmail.com, osalvador@suse.de, rientjes@google.com, david@redhat.com, joshua.hahnjy@gmail.com, fvdl@google.com Subject: Re: [PATCH v5] page_alloc: allow migration of smaller hugepages during contig_alloc Message-ID: References: <20251218190832.1319797-1-gourry@gourry.net> <0E77F151-99B0-4F67-814A-4D79439C9A88@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <0E77F151-99B0-4F67-814A-4D79439C9A88@nvidia.com> On Thu, Dec 18, 2025 at 02:45:37PM -0500, Zi Yan wrote: > > That can save another scan? And caller can pass hugetlb_search_result if > they care and check its value if pfn_range_valid_contig() returns false. > Well, first, I've generally seen it discouraged to do output-parameters like this for such trivial things. But that aside... We have to scan again either way if we want to prefer allocating non-hugetlb regions in different memory blocks first. This is what Mel was pointing out (we should touch every OTHER block before we attempt HugeTLB migrations). The best optimization you could hope for is something like the following - but honestly, this is ugly, racy (zone contents may have changed between scans), and if you're already in the slow reliable path then we should just be slow and re-scan the non-hugetlb sections as well. Other than this being ugly, I don't have strong feelings. If people would prefer the second pass to ONLY touch hugetlb sections, I'll ship this. static bool pfn_range_valid_contig(struct zone *z, unsigned long start_pfn, unsigned long nr_pages, bool search_hugetlb, bool *hugetlb_found) { bool hugetlb = false; for (i = start_pfn; i < end_pfn; i++) { ... if (PageHuge(page)) { if (hugetlb_found) *hugetlb_found = true; if (!search_hugetlb) return false; ... hugetlb = true; } } /* * If we're searching for hugetlb regions, only return those * Otherwise only return regions without hugetlb reservations */ return !search_hugetlb || hugetlb; } struct page *alloc_contig_pages_noprof(unsigned long nr_pages, gfp_t gfp_mask, int nid, nodemask_t *nodemask) { bool search_hugetlb = false; bool hugetlb_found = false; retry: zonelist = node_zonelist(nid, gfp_mask); for_each_zone_zonelist_nodemask(zone, z, zonelist, gfp_zone(gfp_mask), nodemask) { spin_lock_irqsave(&zone->lock, flags); pfn = ALIGN(zone->zone_start_pfn, nr_pages); while (zone_spans_last_pfn(zone, pfn, nr_pages)) { if (pfn_range_valid_contig(zone, pfn, nr_pages, search_hugetlb, &hugetlb_found)) { ... } } if (IS_ENABLED(CONFIG_ARCH_ENABLE_HUGEPAGE_MIGRATION) && !search_hugetlb && hugetlb_found) { search_hugetlb = true; goto retry; } return NULL; } ~Gregory