From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 58DA432826F for ; Tue, 16 Dec 2025 18:22:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765909327; cv=none; b=EMQablohwdtKBSYZMCCukfBCNAR99RkxdFQlzqpgDixaZ7f5gQJqPMin27d4tahywh9CJQq29ujT70z7uh5f7ineON4d1wAVcsRYkflJdyCGV0HdXq2ocvXn424Tp3jisQMZm8XO4BPgRGg1Nz9AR7+ztOHME1knfrOpdguaVzI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1765909327; c=relaxed/simple; bh=dIC+ZznaoOTiv5oAaeLi9cu+OEz/nPaAs+LE4Dgg8c4=; h=Date:From:To:Cc:Subject:Message-Id:In-Reply-To:References: Mime-Version:Content-Type; b=FsJn/ei0/sUholGXYaJPxuXZqm25w9t734K1pWLZfKzfWc13A9lON9QVhlp0qPVzSARjasBxJLTUAR5tbaaP16bcGwjtsNTTAfGLWIi12nery94grJHcFx9jhxlnsZuzrZTzjAskxIuV1DRvgL5wm/3PiDZ3gxHvcY70sXxsNPw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b=CBJ7C0XU; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b="CBJ7C0XU" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 6371AC4CEF1; Tue, 16 Dec 2025 18:22:06 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=linux-foundation.org; s=korg; t=1765909326; bh=dIC+ZznaoOTiv5oAaeLi9cu+OEz/nPaAs+LE4Dgg8c4=; h=Date:From:To:Cc:Subject:In-Reply-To:References:From; b=CBJ7C0XUNfD5bTzVnHKEICoyhB7/dC7JEYsxoRIGLNPwwywo49J0nMgR1iaermaLg 5APqDvzJftCztMCRQhrU12WdddqIfjNZu3Kqz8i4Il1gqyMNVv9FeUv6O2WUcxzeKt KDVcoqGwBV4I69sV5NAM8SJLp1K2DwWyL52Wog5k= Date: Tue, 16 Dec 2025 10:22:05 -0800 From: Andrew Morton To: Joshua Hahn Cc: Daniel Palmer , Matthew Wilcox , Brendan Jackman , Johannes Weiner , Michal Hocko , Suren Baghdasaryan , Vlastimil Babka , Zi Yan , linux-kernel@vger.kernel.org, linux-mm@kvack.org, kernel-team@meta.com Subject: Re: [PATCH v4] mm/page_alloc: prevent reporting pcp->batch = 0 Message-Id: <20251216102205.1d008d1432446956b079bfff@linux-foundation.org> In-Reply-To: <20251216144813.3016985-1-joshua.hahnjy@gmail.com> References: <20251216144813.3016985-1-joshua.hahnjy@gmail.com> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Tue, 16 Dec 2025 06:48:11 -0800 Joshua Hahn wrote: > zone_batchsize returns the appropriate value that should be used for > pcp->batch. If it finds a zone with less than 4096 pages or PAGE_SIZE > > 1M, however, it leads to some incorrect math. > > In the case above, we will get an intermediary value of 1, which is then > rounded down to the nearest power of two, and 1 is subtracted from it. > Since 1 is already a power of two, we will get batch = 1-1 = 0: > > batch = rounddown_pow_of_two(batch + batch/2) - 1; > > A pcp->batch value of 0 is nonsensical, for MMU systems. If this were > actually set, then functions like drain_zone_pages would become no-ops, > since they would free 0 pages at a time. > > Of the two callers of zone_batchsize, the one that is actually used to > set pcp->batch works around this by setting pcp->batch to the maximum > of 1 and zone_batchsize. However, the other caller, zone_pcp_init, > incorrectly prints out the batch size of the zone to be 0. > > This is probably rare in a typical zone, but the DMA zone can often have > less than 4096 pages, which means it will print out "LIFO batch:0". > > Before: [ 0.001216] DMA zone: 3998 pages, LIFO batch:0 > After: [ 0.001210] DMA zone: 3998 pages, LIFO batch:1 > > With all of this said, NOMMU differs in two ways. Semantically, it > should report that pcp->batch is 0. At the same time, it can never > really have a pcp->batch size of 0 since it will reach a deadlock in > pcp freeing functions. For this reason, zone_batchsize should still > report 0 for NOMMU, but zone_set_pageset_high_and_batch should still > interpret it as 1, meaning we cannot get rid of max(1, zone_batchsize()) > in zone_set_pageset_high_and_batch. > > Suggested-by: Daniel Palmer > Signed-off-by: Joshua Hahn > --- > Reviewers' note: > > This patch was originally a part of the 6.19-rc1 pr, but Daniel Palmer > kindly reported that this patch causes an issue on NOMMU systems [1]. > Thank you, Daniel! I wasn't sure how to credit here since it was a > report on an unmerged commit so I went with suggested-by. If this is > problematic please let me know and I will change the tag. > > [1] https://lore.kernel.org/all/CAFr9PX=_HaM3_xPtTiBn5Gw5-0xcRpawpJ02NStfdr0khF2k7g@mail.gmail.com/ > > Reviewer's note (to Andrew): > > This replaces commit 2/2 of the series titled "mm/page_alloc: pcp->batch > cleanups" [2]. That series is in mainline. 2783088ef24e ("mm/page_alloc: prevent reporting pcp->batch = 0"). > --- a/mm/page_alloc.c > +++ b/mm/page_alloc.c > @@ -5888,8 +5888,8 @@ static int zone_batchsize(struct zone *zone) > * and zone lock contention. > */ > batch = min(zone_managed_pages(zone) >> 12, SZ_256K / PAGE_SIZE); > - if (batch < 1) > - batch = 1; > + if (batch <= 1) > + return 1; > > /* > * Clamp the batch to a 2^n - 1 value. Having a power > So this doesn't work. Please send along a fix for current -linus and include Fixes: 2783088ef24e ("mm/page_alloc: prevent reporting pcp->batch = 0") and Reported-by:Daniel and the appropriate Closes: Thanks.