From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f70.google.com (mail-wm1-f70.google.com [209.85.128.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AA6243E63A2 for ; Sun, 26 Jul 2026 22:23:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.70 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785104641; cv=none; b=X7NBJS+pze62zKfq2R0ttEFTg9NOvfzQRDol4qmiLgz6CjXOhPXUPZ+MkKKF+y+O0KRbRqqpMZZ7pxaeEhg3+vx3B5ttkJOfJ5p8Z48Y1vj0zsIWtZJqulfSbuf1FCQ6a4LUWmeY5dilPrjhwHY0FRvPbcmfKiGSC5aqJkVkN28= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785104641; c=relaxed/simple; bh=xilIf6dHekuJD/8ooBDTE1D1gv8zUikfdkLRko7miLc=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ce2cBBCLAaMphAXSnSo7+Ox9YWNyFtca+JOfq0GvPFkJDnw6gkJyFiWPw4aMZKIqiCRtiabejERN64KoZ9YxPIfSBGnY6wH62LZrTU5N0PFIuIrxAHR/L3VwdJpFuKd35d9Ih/UFj2cqDjwtOnz9bMMzq9VQUdmAxymzbEGZNOY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jackmanb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=HqCi5mnl; arc=none smtp.client-ip=209.85.128.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jackmanb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="HqCi5mnl" Received: by mail-wm1-f70.google.com with SMTP id 5b1f17b1804b1-49544a639e9so10570835e9.1 for ; Sun, 26 Jul 2026 15:23:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1785104638; x=1785709438; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=m2xuI/zeM0UttuNeFOz41CO3X9EkXjzbyKL0Tm6Mtd0=; b=HqCi5mnlQj7mIYaiwFGE2NnyUHX8n2bqUc3Gw6qY5cYMi18vjuuT0irwmMzR6s912J 9Gb7tKSTMFX/Gk05g08xK9Wwz6tms2l/F+jCCv2mFDQ0t7l/5iQqjoKGQ1GQFZ8SSKbW V0pQ35F0WGwvvpIdv9kJuB2YUdWtVRJUKGis6ypsZSIA6lbA6NXcMk6EyznvXdmU0HI+ RHWwW5XzA+Z8ABVOUIQGZ7+m7zkEI9crkgL5g+vsdzReb7LBpapgF0ssXzYXBQw+ji/f P5wT4FzVaBzjbsGNImQOpfZy8rhMzps1qFlsLhOMlicXneGmhsjDPNwke+niUaSoNzJN Jrdg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785104638; x=1785709438; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=m2xuI/zeM0UttuNeFOz41CO3X9EkXjzbyKL0Tm6Mtd0=; b=N7AJ54Cu6k2R7hOoqcHt/UtsPNxPoUI4Tpyy5wmJHRdlfuWx425CRpIzWG6Sf0JWXs MT5L5pwUqxqijoM1n2qollVpdVidkDftO/M33jkeZafnkcefhtrj913utNedH2moWUaf gOKH31OwFg6j9X8SVMLZnwxLPfcPUDZvNtK7Z2zOsD31+XSGaov0vPUb/MoGb25i5LjD +PYN/VgwLaL6EW5lSSUszEPL4c2lK2MCu/B8SrgB+HBZgmWntMqGHZcTQATcqJj6cSIl uumx/5WFh7s3d77DlWvwozjKYPPoeWGxC0dvOnYpk95z71osUsfsq8t90Tyt7ps5hmrX 0zoA== X-Forwarded-Encrypted: i=1; AHgh+RpVALoILRdBwQnao228u7k4ah6ybwm/oUojTN1a6qN5SbijDDAYtiyi5JDc+K9yDOOWLl+2jSep1W9jhuc=@vger.kernel.org X-Gm-Message-State: AOJu0Yy1AY3YDnHq/TXvwbFVKTKQ5xHZohjjkT6w4JVCNV1ULhKurc8s KgmlDs0QVlfEb1tFFKMs1xA7ZtfU3GuHU9AC+rjZXVZ6LBKSePwxQvAIqTfD1ExJ1yLYiOgdqli UoaYnUfDUB3CXoQ== X-Received: from wmaj2.prod.google.com ([2002:a05:600c:6c02:b0:494:133:ddd3]) (user=jackmanb job=prod-delivery.src-stubby-dispatcher) by 2002:a05:600c:a102:b0:493:c77c:108a with SMTP id 5b1f17b1804b1-496b575ded2mr61716775e9.36.1785104637810; Sun, 26 Jul 2026 15:23:57 -0700 (PDT) Date: Sun, 26 Jul 2026 22:22:56 +0000 In-Reply-To: <20260726-page_alloc-unmapped-v3-0-6f5729aa9832@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260726-page_alloc-unmapped-v3-0-6f5729aa9832@google.com> X-Mailer: b4 0.16-dev Message-ID: <20260726-page_alloc-unmapped-v3-23-6f5729aa9832@google.com> Subject: [PATCH v3 23/26] mm: Split out NR_FREE_PAGES_BLOCKS_[UN]MAPPED From: Brendan Jackman To: Borislav Petkov , Dave Hansen , Peter Zijlstra , Andrew Morton , David Hildenbrand , Vlastimil Babka , Mike Rapoport , Wei Xu , Johannes Weiner , Zi Yan , Lorenzo Stoakes Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, x86@kernel.org, rppt@kernel.org, Sumit Garg , Will Deacon , rientjes@google.com, "Kalyazin, Nikita" , patrick.roy@linux.dev, "Itazuri, Takahiro" , Andy Lutomirski , David Kaplan , Thomas Gleixner , Yosry Ahmed , Patrick Bellasi , Reiji Watanabe , Sean Christopherson , Brendan Jackman Content-Type: text/plain; charset="utf-8" In order to infer whether compaction is likely to enable an allocation that maps/unmaps a pageblock, we need to know how many fully-free pageblocks of each type there are. Signed-off-by: Brendan Jackman --- include/linux/mmzone.h | 5 ++++- include/linux/vmstat.h | 10 ++++++++++ mm/compaction.c | 5 ++--- mm/page_alloc.c | 13 ++++++++++--- mm/vmscan.c | 49 ++++++++++++++++++++++++++++--------------------- mm/vmstat.c | 3 ++- 6 files changed, 56 insertions(+), 29 deletions(-) diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index 2533a228f71fd..e063a59bbc883 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -181,7 +181,10 @@ enum numa_stat_item { enum zone_stat_item { NR_FREE_PAGES, - NR_FREE_PAGES_BLOCKS, + /* Number of free pages in entirely free pageblocks, with direct map */ + NR_FREE_PAGES_BLOCKS_MAPPED, + /* Ditto, without direct map */ + NR_FREE_PAGES_BLOCKS_UNMAPPED, NR_ZONE_LRU_BASE, /* Used only for compaction and reclaim retry */ NR_ZONE_INACTIVE_ANON = NR_ZONE_LRU_BASE, NR_ZONE_ACTIVE_ANON, diff --git a/include/linux/vmstat.h b/include/linux/vmstat.h index 5b31d8e7ae405..debd593756149 100644 --- a/include/linux/vmstat.h +++ b/include/linux/vmstat.h @@ -224,6 +224,16 @@ static inline unsigned long zone_page_state(struct zone *zone, return x; } +/* + * Approx number of pages in entirely free pageblocks. Due to races this could + * actually return a value more than the number of pages in the zone. + */ +static inline unsigned long zone_free_pages_blocks(struct zone *zone) +{ + return zone_page_state(zone, NR_FREE_PAGES_BLOCKS_MAPPED) + + zone_page_state(zone, NR_FREE_PAGES_BLOCKS_UNMAPPED); +} + /* * More accurate version that also considers the currently pending * deltas. For that we need to loop over all cpus to find the current diff --git a/mm/compaction.c b/mm/compaction.c index c9eb3947ffc79..ed12d2fc6fad3 100644 --- a/mm/compaction.c +++ b/mm/compaction.c @@ -2353,8 +2353,7 @@ static enum compact_result __compact_finished(struct compact_control *cc) if (__zone_watermark_ok(cc->zone, cc->order, high_wmark_pages(cc->zone), cc->highest_zoneidx, cc->alloc_flags, - zone_page_state(cc->zone, - NR_FREE_PAGES_BLOCKS))) + zone_free_pages_blocks(cc->zone))) return COMPACT_SUCCESS; return COMPACT_CONTINUE; @@ -2538,7 +2537,7 @@ compaction_suit_allocation_order(struct zone *zone, unsigned int order, unsigned long watermark; if (kcompactd && defrag_mode) - free_pages = zone_page_state(zone, NR_FREE_PAGES_BLOCKS); + free_pages = zone_free_pages_blocks(zone); else free_pages = zone_page_state(zone, NR_FREE_PAGES); diff --git a/mm/page_alloc.c b/mm/page_alloc.c index ac2f6190117ae..d12ce84662ab7 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -866,6 +866,13 @@ static inline void account_freepages(struct zone *zone, int nr_pages, zone->nr_free_highatomic + nr_pages); } +static inline enum zone_stat_item free_pages_blocks_stat(freetype_t ft) +{ + if (freetype_flags(ft) & FREETYPE_UNMAPPED) + return NR_FREE_PAGES_BLOCKS_UNMAPPED; + return NR_FREE_PAGES_BLOCKS_MAPPED; +} + /* Used for pages not on another list */ static inline void __add_to_free_list(struct page *page, struct zone *zone, unsigned int order, freetype_t freetype, @@ -890,7 +897,7 @@ static inline void __add_to_free_list(struct page *page, struct zone *zone, area->nr_free++; if (order >= pageblock_order && !is_migrate_isolate(free_to_migratetype(freetype))) - __mod_zone_page_state(zone, NR_FREE_PAGES_BLOCKS, nr_pages); + __mod_zone_page_state(zone, free_pages_blocks_stat(freetype), nr_pages); } /* @@ -926,7 +933,7 @@ static inline void move_to_free_list(struct page *page, struct zone *zone, is_migrate_isolate(old_mt) != is_migrate_isolate(new_mt)) { if (!is_migrate_isolate(old_mt)) nr_pages = -nr_pages; - __mod_zone_page_state(zone, NR_FREE_PAGES_BLOCKS, nr_pages); + __mod_zone_page_state(zone, free_pages_blocks_stat(new_ft), nr_pages); } } @@ -954,7 +961,7 @@ static inline void __del_page_from_free_list(struct page *page, struct zone *zon zone->free_area[order].nr_free--; if (order >= pageblock_order && !is_migrate_isolate(free_to_migratetype(freetype))) - __mod_zone_page_state(zone, NR_FREE_PAGES_BLOCKS, -nr_pages); + __mod_zone_page_state(zone, free_pages_blocks_stat(freetype), -nr_pages); } static inline void del_page_from_free_list(struct page *page, struct zone *zone, diff --git a/mm/vmscan.c b/mm/vmscan.c index 566c4e837c7d5..5789c39a0a729 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -6935,6 +6935,28 @@ static bool pgdat_watermark_boosted(pg_data_t *pgdat, int highest_zoneidx) return false; } +/* + * Helper to get *FREE_PAGES* zone stats with accuracy heuristics. + * + * When there is a high number of CPUs in the system, the cumulative error from + * the vmstat per-cpu cache can blur the line between the watermarks. In that + * case, be safe and get an accurate snapshot. + * + * TODO: NR_FREE_PAGES_BLOCKS_* move in steps of pageblock_nr_pages, while the + * vmstat pcp threshold is limited to 125. On many configurations that counter + * won't actually be per-cpu cached. But keep things simple for now; revisit + * when somebody cares. + */ +static inline unsigned long get_free_pages_stat(struct zone *zone, + enum zone_stat_item item) +{ + unsigned long free_pages = zone_page_state(zone, item); + + if (zone->percpu_drift_mark && free_pages < zone->percpu_drift_mark) + return zone_page_state_snapshot(zone, item); + return free_pages; +} + /* * Returns true if there is an eligible zone balanced for the request order * and highest_zoneidx @@ -6950,7 +6972,6 @@ static bool pgdat_balanced(pg_data_t *pgdat, int order, int highest_zoneidx) * meet watermarks. */ for_each_managed_zone_pgdat(zone, pgdat, i, highest_zoneidx) { - enum zone_stat_item item; unsigned long free_pages; if (sysctl_numa_balancing_mode & NUMA_BALANCING_MEMORY_TIERING) @@ -6968,26 +6989,12 @@ static bool pgdat_balanced(pg_data_t *pgdat, int order, int highest_zoneidx) * has dropped order, simply ensure there are enough * base pages for compaction, wake kcompactd & sleep. */ - if (defrag_mode && order) - item = NR_FREE_PAGES_BLOCKS; - else - item = NR_FREE_PAGES; - - /* - * When there is a high number of CPUs in the system, - * the cumulative error from the vmstat per-cpu cache - * can blur the line between the watermarks. In that - * case, be safe and get an accurate snapshot. - * - * TODO: NR_FREE_PAGES_BLOCKS moves in steps of - * pageblock_nr_pages, while the vmstat pcp threshold - * is limited to 125. On many configurations that - * counter won't actually be per-cpu cached. But keep - * things simple for now; revisit when somebody cares. - */ - free_pages = zone_page_state(zone, item); - if (zone->percpu_drift_mark && free_pages < zone->percpu_drift_mark) - free_pages = zone_page_state_snapshot(zone, item); + if (defrag_mode && order) { + free_pages = get_free_pages_stat(zone, NR_FREE_PAGES_BLOCKS_UNMAPPED) + + get_free_pages_stat(zone, NR_FREE_PAGES_BLOCKS_MAPPED); + } else { + free_pages = get_free_pages_stat(zone, NR_FREE_PAGES); + } if (__zone_watermark_ok(zone, order, mark, highest_zoneidx, 0, free_pages)) diff --git a/mm/vmstat.c b/mm/vmstat.c index cb57714539fb5..f5ab6ab641c6d 100644 --- a/mm/vmstat.c +++ b/mm/vmstat.c @@ -1200,7 +1200,8 @@ const char * const vmstat_text[] = { /* enum zone_stat_item counters */ #define I(x) (x) [I(NR_FREE_PAGES)] = "nr_free_pages", - [I(NR_FREE_PAGES_BLOCKS)] = "nr_free_pages_blocks", + [I(NR_FREE_PAGES_BLOCKS_MAPPED)] = "nr_free_pages_blocks_mapped", + [I(NR_FREE_PAGES_BLOCKS_UNMAPPED)] = "nr_free_pages_blocks_unmapped", [I(NR_ZONE_INACTIVE_ANON)] = "nr_zone_inactive_anon", [I(NR_ZONE_ACTIVE_ANON)] = "nr_zone_active_anon", [I(NR_ZONE_INACTIVE_FILE)] = "nr_zone_inactive_file", -- 2.54.0