mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Qiliang Yuan <odys.yuan@gmail.com>
To: Andrew Morton <akpm@linux-foundation.org>,
	 Kairui Song <kasong@tencent.com>, Qi Zheng <qi.zheng@linux.dev>,
	 Shakeel Butt <shakeel.butt@linux.dev>,
	Barry Song <baohua@kernel.org>,
	 Axel Rasmussen <axelrasmussen@google.com>,
	Yuanchu Xie <yuanchu@google.com>,  Wei Xu <weixugc@google.com>,
	Baoquan He <baoquan.he@linux.dev>,
	 Baolin Wang <baolin.wang@linux.alibaba.com>,
	 David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>,
	 "Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 Michal Hocko <mhocko@suse.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	 Masami Hiramatsu <mhiramat@kernel.org>,
	 Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
	 Brendan Jackman <brendan.jackman@linux.dev>,
	 Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	 linux-trace-kernel@vger.kernel.org,
	Qiliang Yuan <odys.yuan@gmail.com>
Subject: [PATCH 2/2] mm/compaction: defer failed async direct compaction
Date: Thu, 01 Oct 2026 23:33:49 +0800	[thread overview]
Message-ID: <20261001-bug-mm-thp-async-compact-defer-v1-2-0174c7923430@gmail.com> (raw)
In-Reply-To: <20261001-bug-mm-thp-async-compact-defer-v1-0-0174c7923430@gmail.com>

A THP fault on a MADV_HUGEPAGE VMA first tries the local node only,
with __GFP_THISNODE | __GFP_NORETRY, and runs a single round of async
direct compaction there before falling back to other nodes. A failed
async run is never deferred.

When the local node has plenty of free memory but no free pageblock,
and what sits between the free pages can't be migrated, such as memory
long-term pinned for RDMA, each of these faults scans the zone again
and fails. Populating a large buffer pays for one failed compaction
per 2M fault. A KV-cache store that registers hundreds of GiB for RDMA
reports registration growing from tens of seconds to tens of minutes,
and works around it with MPOL_INTERLEAVE.

Defer the async state when an async run fails, and check it before
further async runs. Sync compaction keeps its own state and behaves as
before, and a successful compaction or allocation resets both. Keep
resetting the pageblock skip hints only when sync compaction restarts,
as async relies on them.

Populating a 4 GiB MADV_HUGEPAGE buffer from node 0 of a two-node VM,
with node 0 (24 GiB) fragmented by long-term pinning every other page,
median of 3 runs:

                          before     after
  direct compactions      735        15
  time in compaction      122.7 ms   3.6 ms
  population time         380 ms     273 ms
  THPs on node 0          1313       1126

The THPs that no longer land on node 0 come from node 1. With the same
fragmentation but movable memory, compaction still succeeds and the
median run still gets all 2048 THPs on node 0, as before.

Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
 mm/compaction.c | 21 ++++++++++++++-------
 1 file changed, 14 insertions(+), 7 deletions(-)

diff --git a/mm/compaction.c b/mm/compaction.c
index 7f8845d1990aa..25f758bc54f06 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -2600,7 +2600,9 @@ compact_zone(struct compact_control *cc, struct capture_control *capc)
 
 	/*
 	 * Clear pageblock skip if there were failures recently and compaction
-	 * is about to be retried after being deferred.
+	 * is about to be retried after being deferred. Only do it when sync
+	 * compaction restarts: async compaction relies on the skip hints, and
+	 * clearing them on every async retry would rescan the whole zone.
 	 */
 	if (compaction_restarting(cc->zone, cc->order))
 		__reset_isolation_suitable(cc->zone);
@@ -2858,8 +2860,10 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
 			!__cpuset_zone_allowed(zone, gfp_mask))
 				continue;
 
-		if (prio > MIN_COMPACT_PRIORITY
-					&& compaction_deferred(zone, order, true)) {
+		if (prio > MIN_COMPACT_PRIORITY &&
+		    (compaction_deferred(zone, order, true) ||
+		     (prio == COMPACT_PRIO_ASYNC &&
+		      compaction_deferred(zone, order, false)))) {
 			rc = max_t(enum compact_result, COMPACT_DEFERRED, rc);
 			continue;
 		}
@@ -2890,14 +2894,17 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
 			break;
 		}
 
-		if (prio != COMPACT_PRIO_ASYNC && (status == COMPACT_COMPLETE ||
-					status == COMPACT_PARTIAL_SKIPPED))
+		if (status == COMPACT_COMPLETE ||
+		    status == COMPACT_PARTIAL_SKIPPED)
 			/*
 			 * We think that allocation won't succeed in this zone
 			 * so we defer compaction there. If it ends up
-			 * succeeding after all, it will be reset.
+			 * succeeding after all, it will be reset. A failed
+			 * async run only defers further async runs, as sync
+			 * compaction may succeed on pageblocks it skipped.
 			 */
-			defer_compaction(zone, order, true);
+			defer_compaction(zone, order,
+					 prio != COMPACT_PRIO_ASYNC);
 
 		/*
 		 * We might have stopped compacting due to need_resched() in

-- 
2.43.0


  parent reply	other threads:[~2026-10-01 15:34 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 15:33 [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault Qiliang Yuan
2026-10-01 15:33 ` [PATCH 1/2] mm/compaction: keep compaction deferral state per migration mode Qiliang Yuan
2026-10-01 15:33 ` Qiliang Yuan [this message]
2026-10-02 14:21 ` [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault Lorenzo Stoakes (ARM)
2026-10-02 16:17   ` Qiliang Yuan
2026-10-02 17:00     ` Liam R. Howlett

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001-bug-mm-thp-async-compact-defer-v1-2-0174c7923430@gmail.com \
    --to=odys.yuan@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=axelrasmussen@google.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=baoquan.he@linux.dev \
    --cc=brendan.jackman@linux.dev \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=kasong@tencent.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=ljs@kernel.org \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=mhocko@suse.com \
    --cc=qi.zheng@linux.dev \
    --cc=rostedt@goodmis.org \
    --cc=rppt@kernel.org \
    --cc=shakeel.butt@linux.dev \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=weixugc@google.com \
    --cc=yuanchu@google.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®