mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Rik van Riel <riel@surriel.com>
To: linux-kernel@vger.kernel.org
Cc: Andrew Morton <akpm@linux-foundation.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>,
	Brendan Jackman <brendan.jackman@linux.dev>,
	Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
	Kairui Song <kasong@tencent.com>, Qi Zheng <qi.zheng@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	Barry Song <baohua@kernel.org>,
	Axel Rasmussen <axelrasmussen@google.com>,
	Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
	Baoquan He <baoquan.he@linux.dev>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>,
	"Liam R. Howlett" <liam@infradead.org>,
	Mike Rapoport <rppt@kernel.org>,
	linux-mm@kvack.org, Rik van Riel <riel@surriel.com>
Subject: [RFC PATCH 4/5] mm/page_alloc: skip movable blocks during non-movable steals
Date: Tue,  6 Oct 2026 22:12:49 -0400	[thread overview]
Message-ID: <20261007021250.1665929-5-riel@surriel.com> (raw)
In-Reply-To: <20261007021250.1665929-1-riel@surriel.com>

__rmqueue_steal() hands a non-movable allocation pages from a
movable pageblock and leaves the block's type alone. Those pages
can never migrate, so the block reads movable to compaction while
holding content compaction cannot move. On a 3 GB guest running
ls -lR / and a 2.4 GB block device read against a fragmenter, a
temporary counter found 2654, 5341 and 5272 such allocations in
three boots.

__rmqueue_claim() runs first and would convert the block, but
rmqueue_bulk() remembers the mode across its batch. Prevent
those allocations from placing non-movable pages inside
movable pageblocks by refusing __rmqueue_steal() for non-movable
allocations in movable pageblocks.

Movable is last in fallbacks[] for both non-movable types, so
only a larger order is left to try; failing that, the steal
returns NULL and the allocation will loop around to a claim,
another zone, or reclaim.

Keep the steal where mobility grouping is disabled, or where
the block straddles a zone edge. A block that straddles zones
cannot be claimed, because the allocator only holds the lock
for one zone, which leaves stealing as the alloc path there.

Movable allocations can steal, because kcompactd can always
move those pages out of non-movable blocks later.

Over three boots allocstall_normal runs 48 to 129 against 32 to
134 on the base, and compact_stall 324 to 415 against 288 to 387.
Neither direct reclaim nor compaction stalls rise beyond
run-to-run spread.

Assisted-by: LLM
Signed-off-by: Rik van Riel <riel@surriel.com>
---
 mm/page_alloc.c | 19 +++++++++++++++++--
 1 file changed, 17 insertions(+), 2 deletions(-)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 9f1520506ae1d..16d3594d913a0 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -2423,8 +2423,10 @@ __rmqueue_claim(struct zone *zone, int order, int start_migratetype,
 }
 
 /*
- * Try to steal a single page from some fallback migratetype. Leave the rest of
- * the block as its current migratetype, potentially causing fragmentation.
+ * Try to steal one page from a fallback type, leaving the rest of the block
+ * unchanged and possibly fragmented. A non-movable allocation must not steal
+ * from a movable block: without converting its type, the steal would leave
+ * non-movable content under a movable type.
  */
 static __always_inline struct page *
 __rmqueue_steal(struct zone *zone, int order, int start_migratetype)
@@ -2444,6 +2446,19 @@ __rmqueue_steal(struct zone *zone, int order, int start_migratetype)
 			continue;
 
 		page = get_page_from_free_area(area, fallback_mt);
+
+		/*
+		 * Do not allow non-movable allocations in movable
+		 * pageblocks; that could break compaction.
+		 * Non-movable allocations should claim pageblocks, instead.
+		 */
+		if (!is_migrate_movable(start_migratetype) &&
+		    is_migrate_movable(fallback_mt) &&
+		    !page_group_by_mobility_disabled &&
+		    zone_spans_pageblock(zone, page_to_pfn(page))) {
+			continue;
+		}
+
 		page_del_and_expand(zone, page, order, current_order, fallback_mt);
 		trace_mm_page_alloc_extfrag(page, order, current_order,
 					    start_migratetype, fallback_mt);
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-10-07  2:13 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07  2:12 [RFC PATCH 0/5] mm/page_alloc: keep non-movable pages out of movable pageblocks for THP Rik van Riel
2026-10-07  2:12 ` [RFC PATCH 1/5] mm/page_alloc: count guard pages as free again when their buddy merges Rik van Riel
2026-10-07  2:12 ` [RFC PATCH 2/5] mm/page_alloc: count a merged buddy's pages under the merged type Rik van Riel
2026-10-07  2:12 ` [RFC PATCH 3/5] mm/page_alloc: convert any movable pageblock on a non-movable allocation Rik van Riel
2026-10-07  2:12 ` Rik van Riel [this message]
2026-10-07  2:12 ` [RFC PATCH 5/5] mm/page_alloc: claim blocks captured for non-movable use Rik van Riel

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261007021250.1665929-5-riel@surriel.com \
    --to=riel@surriel.com \
    --cc=akpm@linux-foundation.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=baoquan.he@linux.dev \
    --cc=brendan.jackman@linux.dev \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=kasong@tencent.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=qi.zheng@linux.dev \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=weixugc@google.com \
    --cc=yuanchu@google.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®