From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qk1-f181.google.com (mail-qk1-f181.google.com [209.85.222.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DD3651A5B8A for ; Sun, 21 Dec 2025 12:42:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.181 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766320972; cv=none; b=G2hPgkAqp0q4iWH+vg76qAMZUCDjA81fmpodzIu2CGQm2DHCRx/zqlvnK/jo9vb663pGO5O59MIpe+vJpz5sJCEnIVMuVrHcJHYWZ73mhws+UH9ApmMnaBpHNmYPLkiwf8Ay8VLw0q+zla1HdOrz8et/plZPJ6GXdQM5W5b+BNU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766320972; c=relaxed/simple; bh=UA5ZeDAnvLvzC1bO0anHtshNuL5wkCnP6SFXvjlGKUA=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=cagN6NUNEQukC6GEoxJu832jHb43sSJf7VgAW3cXBcxqM4DgLDMMroMisL7faZqLz2fVm07/lFOsoXCiew4Zbqfgquris4Jj8na1yVzyA3QsDvNh9IARQeOgEVC6N7175GwfAV/H5mevsSe2u+Kw8sk7CxjmL2ZIiJ2KrHXW5XA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=B3nqHx7D; arc=none smtp.client-ip=209.85.222.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="B3nqHx7D" Received: by mail-qk1-f181.google.com with SMTP id af79cd13be357-8b31a665ba5so381809985a.2 for ; Sun, 21 Dec 2025 04:42:48 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1766320968; x=1766925768; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=GzdhnR2YNZXQ+KYGbrDA5WK3eQTd25rFnUcMaLp5Do8=; b=B3nqHx7Dp1TTjaH8vspDQFuTwc5XKusqHY80YtZtksrUpe9ereNRfxsWfj+KO3Rwx5 aBqxtS902VplrU6EVrUHDoPvT1Hu9ch4K/YsGhf10V+KGHuwdoOiYjwjnMQyH6KeeOh2 NmMFi3y5gTEtr2FndtHKDLng2ZkwfwdKmG/uZ9g2++RFOaoheKDnLCQsCR3T1VnYkNOF NeYxYZiuAynDT1BKPHV7QhqZ3D4zqgTCklwlrANbk7v6EAXNvzty2BT+g4FdHZIGdjvi tqoNJf+tVU+jXCTyjUiSNkKUNuZZdvOhar9YWs2+rIV/ujjiwOQBUMtH2R97ysBkxw63 uGCg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1766320968; x=1766925768; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=GzdhnR2YNZXQ+KYGbrDA5WK3eQTd25rFnUcMaLp5Do8=; b=oY4CVLgjeDmh69OZJ0CRQittEeXhew7sER/MUxVVm0Pf9U18vj5763pfBrzaCSJNbM CU0GRxfwoz/zNG04ZpBIYSAoSbyJF0+j9lSxP3QYbnki5x7f4WSoiZogvdjM0E81nriv iiJ9R31hBAjT7aAgmnRGTeJ8RTZGmNFyccMgqeQPAwXPUJI7gHEa/b8i8wCXD9KK6nhR 5P3do8M0l3jDu4aHvOMEFWK1vDC7WIxUGgNKjVJJl3l4frKaujSQaywN9AT0/AuolV3O 3cg6+9HWdzlEOv0G7qDgZxjKqnaB8cackiXJDmgApY80G181UiU8MS/zdo5OnbDV7ms1 fkcA== X-Forwarded-Encrypted: i=1; AJvYcCXCy1Hp6hSlGNF20tf3r5yw2/3u3iOEQ7h1d+gdUZY8ivGdBaom9ED7uvCwEf0r0tpMQVoNCro/YCbvGbg=@vger.kernel.org X-Gm-Message-State: AOJu0YyrtoCt5rrdg5VSuSPA7EC4C6a2Us4CGgAQq3A+Bwu7hINBrvP5 GfdflBRiOJpV3x4VHLDHBc/Osvbf4s52SPF3xANIjsUDHmdaewwKjOnyRaAE/uhqfzk= X-Gm-Gg: AY/fxX4Zs89NGqWorzMyvb+PTo0p8+t6QTJWrdjDgnxyhGTutLOHtNu13arIMtQr5yE efQG7yB+ac4J18cd8G/RSGUj2EmmIUEq0ZGFqSUWiLvHOBuctXX9jcj4xTivanzjLGDB7ExOt1g fuWCXwnAGtYYIfH+ZJNOKyC8qD8EE6BviMgMeeYFqb33kGzUfSdunNhOKh4Jt6kNCG4HwD8DTWk 4dh+CLffGQsykuwlXu8/pLb6CNTGc4UIeBZXp0FCH0tLk47/r8zZAwrWT2ykju5D5S9NcpWq8nJ S8OFLJeYF92sNV9bpb6YK8XkGxv6F62hsrVo13h1zh52XPxRRUPNlGMN/PUWW0ZCv0/Q9UK3ly+ kv63qYiKYiX9IA7xdf+PpUzckKh3uBmaPSugpChJsFOMqYJe1WV1MIHvZPqb6ib/wEHn/28rMd3 E52SOpZNMmKXlRi2gdTO15+uqq2bFj32E0ZPLqWYc8O4JjsvNNkFPtR/076uY1nJTTQ6vewA== X-Google-Smtp-Source: AGHT+IFxDZsXvwA/Wlu/ERs4meUfrrEIdfd18oC9/D5bfx1lMCs5IZH+WWgu2rqfTgF+yU2gaqHpOA== X-Received: by 2002:ac8:7c4d:0:b0:4f1:ea37:cc6c with SMTP id d75a77b69052e-4f4abccef5cmr118092311cf.1.1766320967720; Sun, 21 Dec 2025 04:42:47 -0800 (PST) Received: from gourry-fedora-PF4VCD3F (pool-96-255-20-138.washdc.ftas.verizon.net. [96.255.20.138]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-4f4ac650703sm59927451cf.27.2025.12.21.04.42.46 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 21 Dec 2025 04:42:47 -0800 (PST) Date: Sun, 21 Dec 2025 07:42:10 -0500 From: Gregory Price To: Zi Yan Cc: Wei Yang , linux-mm@kvack.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, vbabka@suse.cz, surenb@google.com, mhocko@suse.com, jackmanb@google.com, hannes@cmpxchg.org, osalvador@suse.de, rientjes@google.com, david@redhat.com, joshua.hahnjy@gmail.com, fvdl@google.com Subject: Re: [PATCH v6] page_alloc: allow migration of smaller hugepages during contig_alloc Message-ID: References: <20251218233804.1395835-1-gourry@gourry.net> <20251219000800.tnpqzvcdyeqcwryt@master> <7EED2D83-AE17-49CB-BDB6-954793EAFDBF@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <7EED2D83-AE17-49CB-BDB6-954793EAFDBF@nvidia.com> On Fri, Dec 19, 2025 at 03:46:25PM -0500, Zi Yan wrote: > On 19 Dec 2025, at 9:26, Gregory Price wrote: > > > On Fri, Dec 19, 2025 at 12:08:00AM +0000, Wei Yang wrote: > >>> + > >>> + page = compound_head(page); > >>> + order = compound_order(page); > >> > >> The order is get from head page. > >> > >>> + if ((order >= MAX_FOLIO_ORDER) || > >>> + (nr_pages <= (1 << order))) > >>> + return false; > >>> + > >>> + /* No need to check the pfns for this page */ > >>> + i += (1 << order) - 1; > >> > >> So this advance should based on "head page" instead of original page, right? > >> > > > > hm, I think the thought here was that since we're moving forward from > > start of an aligned chunk, we'd never hit a non-head page - but this > > may not be true. > > > > Will think about this for a bit. > > The sole caller of pfn_range_valid_contig(), alloc_contig_pages_noprof(), > scans from the beginning of a zone to the end. pfn_range_valid_contig() > should see head pages all the time, except it scans in the middle of > a 1GB hugetlb when alloc_contig_pages_noprof() is asking for a smaller > nr_pages, like 2MB. But in that case, the if above i += (1 << order) - 1 > would return false without reaching it. Basically, to get to > i += ..., pfn_range_valid_contig() needs to search for nr_pages larger > than PageHuge(page) and nr_pages is always power of two based on > alloc_contig_pages_noprof() requirement, but that means > pfn_range_valid_contig() always sees such PageHuge pages as a whole > within nr_pages range, thus cannot see a tail PageHuge page at the > point of i += .... > Thinking about this a bit more, it might be worthwhile to detect this condiition and just skip that hugepage in the external code. while (zone_spans_last_pfn(zone, pfn, nr_pages)) { if (pfn_range_valid_contig(zone, pfn, nr_pages, skip_hugetlb, &skipped_hugetlb)) { ... snip ... } pfn += nr_pages; /* * TODO: If the last scanned page was a hugepage that caused * the zone to be invalid, skip the rest of that page * (e.g. if we hit a 1GB page trying to allocate a 2MB * page, skip the entire 1GB instead of scanning the * same page 1GB/2MB times). */ ... } But this solves a different problem than this patch, so i will defer. ~Gregory