* [PATCH 1/2] mm/compaction: keep compaction deferral state per migration mode
2026-10-01 15:33 [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault Qiliang Yuan
@ 2026-10-01 15:33 ` Qiliang Yuan
2026-10-01 15:33 ` [PATCH 2/2] mm/compaction: defer failed async direct compaction Qiliang Yuan
2026-10-02 14:21 ` [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault Lorenzo Stoakes (ARM)
2 siblings, 0 replies; 6+ messages in thread
From: Qiliang Yuan @ 2026-10-01 15:33 UTC (permalink / raw)
To: Andrew Morton, Kairui Song, Qi Zheng, Shakeel Butt, Barry Song,
Axel Rasmussen, Yuanchu Xie, Wei Xu, Baoquan He, Baolin Wang,
David Hildenbrand, Lorenzo Stoakes, Liam R. Howlett,
Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
Brendan Jackman, Johannes Weiner, Zi Yan
Cc: linux-mm, linux-kernel, linux-trace-kernel, Qiliang Yuan
try_to_compact_pages() defers a zone only after sync compaction fails,
so a failed async direct compaction is never deferred and repeats the
same futile zone scan on the next attempt. Deferring async compaction
needs state of its own: sync compaction may still succeed on the
pageblocks async skips, so it mustn't be deferred along with it.
Turn compact_considered, compact_defer_shift and compact_order_failed
into arrays indexed by sync, like compact_cached_migrate_pfn, and pass
the mode to the deferral helpers and tracepoints. All callers use the
sync state for now and a reset clears both. The deferral tracepoints
gain a sync field telling which state they report.
Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
include/linux/mmzone.h | 7 +++--
include/trace/events/compaction.h | 27 ++++++++++--------
mm/compaction.c | 60 +++++++++++++++++++++------------------
3 files changed, 51 insertions(+), 43 deletions(-)
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 94f9c3ff54160..91fbaa7aec6f9 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -1140,10 +1140,11 @@ struct zone {
* are skipped before trying again. The number attempted since
* last failure is tracked with compact_considered.
* compact_order_failed is the minimum compaction failed order.
+ * Indexed by sync, like compact_cached_migrate_pfn.
*/
- unsigned int compact_considered;
- unsigned int compact_defer_shift;
- int compact_order_failed;
+ unsigned int compact_considered[ASYNC_AND_SYNC];
+ unsigned int compact_defer_shift[ASYNC_AND_SYNC];
+ int compact_order_failed[ASYNC_AND_SYNC];
#endif
#if defined CONFIG_COMPACTION || defined CONFIG_CMA
diff --git a/include/trace/events/compaction.h b/include/trace/events/compaction.h
index d05759d185389..23ac6e1750623 100644
--- a/include/trace/events/compaction.h
+++ b/include/trace/events/compaction.h
@@ -238,14 +238,15 @@ DEFINE_EVENT(mm_compaction_suitable_template, mm_compaction_suitable,
DECLARE_EVENT_CLASS(mm_compaction_defer_template,
- TP_PROTO(struct zone *zone, int order),
+ TP_PROTO(struct zone *zone, int order, bool sync),
- TP_ARGS(zone, order),
+ TP_ARGS(zone, order, sync),
TP_STRUCT__entry(
__field(int, nid)
__field(enum zone_type, idx)
__field(int, order)
+ __field(bool, sync)
__field(unsigned int, considered)
__field(unsigned int, defer_shift)
__field(int, order_failed)
@@ -255,15 +256,17 @@ DECLARE_EVENT_CLASS(mm_compaction_defer_template,
__entry->nid = zone_to_nid(zone);
__entry->idx = zone_idx(zone);
__entry->order = order;
- __entry->considered = zone->compact_considered;
- __entry->defer_shift = zone->compact_defer_shift;
- __entry->order_failed = zone->compact_order_failed;
+ __entry->sync = sync;
+ __entry->considered = zone->compact_considered[sync];
+ __entry->defer_shift = zone->compact_defer_shift[sync];
+ __entry->order_failed = zone->compact_order_failed[sync];
),
- TP_printk("node=%d zone=%-8s order=%d order_failed=%d consider=%u limit=%lu",
+ TP_printk("node=%d zone=%-8s order=%d sync=%d order_failed=%d consider=%u limit=%lu",
__entry->nid,
__print_symbolic(__entry->idx, ZONE_TYPE),
__entry->order,
+ __entry->sync,
__entry->order_failed,
__entry->considered,
1UL << __entry->defer_shift)
@@ -271,23 +274,23 @@ DECLARE_EVENT_CLASS(mm_compaction_defer_template,
DEFINE_EVENT(mm_compaction_defer_template, mm_compaction_deferred,
- TP_PROTO(struct zone *zone, int order),
+ TP_PROTO(struct zone *zone, int order, bool sync),
- TP_ARGS(zone, order)
+ TP_ARGS(zone, order, sync)
);
DEFINE_EVENT(mm_compaction_defer_template, mm_compaction_defer_compaction,
- TP_PROTO(struct zone *zone, int order),
+ TP_PROTO(struct zone *zone, int order, bool sync),
- TP_ARGS(zone, order)
+ TP_ARGS(zone, order, sync)
);
DEFINE_EVENT(mm_compaction_defer_template, mm_compaction_defer_reset,
- TP_PROTO(struct zone *zone, int order),
+ TP_PROTO(struct zone *zone, int order, bool sync),
- TP_ARGS(zone, order)
+ TP_ARGS(zone, order, sync)
);
TRACE_EVENT(mm_compaction_kcompactd_sleep,
diff --git a/mm/compaction.c b/mm/compaction.c
index a049415512c67..7f8845d1990aa 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -124,35 +124,35 @@ static unsigned long release_free_list(struct list_head *freepages)
* allocation success. 1 << compact_defer_shift, compactions are skipped up
* to a limit of 1 << COMPACT_MAX_DEFER_SHIFT
*/
-static void defer_compaction(struct zone *zone, int order)
+static void defer_compaction(struct zone *zone, int order, bool sync)
{
- zone->compact_considered = 0;
- zone->compact_defer_shift++;
+ zone->compact_considered[sync] = 0;
+ zone->compact_defer_shift[sync]++;
- if (order < zone->compact_order_failed)
- zone->compact_order_failed = order;
+ if (order < zone->compact_order_failed[sync])
+ zone->compact_order_failed[sync] = order;
- if (zone->compact_defer_shift > COMPACT_MAX_DEFER_SHIFT)
- zone->compact_defer_shift = COMPACT_MAX_DEFER_SHIFT;
+ if (zone->compact_defer_shift[sync] > COMPACT_MAX_DEFER_SHIFT)
+ zone->compact_defer_shift[sync] = COMPACT_MAX_DEFER_SHIFT;
- trace_mm_compaction_defer_compaction(zone, order);
+ trace_mm_compaction_defer_compaction(zone, order, sync);
}
/* Returns true if compaction should be skipped this time */
-static bool compaction_deferred(struct zone *zone, int order)
+static bool compaction_deferred(struct zone *zone, int order, bool sync)
{
- unsigned long defer_limit = 1UL << zone->compact_defer_shift;
+ unsigned long defer_limit = 1UL << zone->compact_defer_shift[sync];
- if (order < zone->compact_order_failed)
+ if (order < zone->compact_order_failed[sync])
return false;
/* Avoid possible overflow */
- if (++zone->compact_considered >= defer_limit) {
- zone->compact_considered = defer_limit;
+ if (++zone->compact_considered[sync] >= defer_limit) {
+ zone->compact_considered[sync] = defer_limit;
return false;
}
- trace_mm_compaction_deferred(zone, order);
+ trace_mm_compaction_deferred(zone, order, sync);
return true;
}
@@ -165,24 +165,28 @@ static bool compaction_deferred(struct zone *zone, int order)
void compaction_defer_reset(struct zone *zone, int order,
bool alloc_success)
{
- if (alloc_success) {
- zone->compact_considered = 0;
- zone->compact_defer_shift = 0;
+ int sync;
+
+ for (sync = 0; sync < ASYNC_AND_SYNC; sync++) {
+ if (alloc_success) {
+ zone->compact_considered[sync] = 0;
+ zone->compact_defer_shift[sync] = 0;
+ }
+ if (order >= zone->compact_order_failed[sync])
+ zone->compact_order_failed[sync] = order + 1;
}
- if (order >= zone->compact_order_failed)
- zone->compact_order_failed = order + 1;
- trace_mm_compaction_defer_reset(zone, order);
+ trace_mm_compaction_defer_reset(zone, order, true);
}
-/* Returns true if restarting compaction after many failures */
+/* Returns true if restarting sync compaction after many failures */
static bool compaction_restarting(struct zone *zone, int order)
{
- if (order < zone->compact_order_failed)
+ if (order < zone->compact_order_failed[true])
return false;
- return zone->compact_defer_shift == COMPACT_MAX_DEFER_SHIFT &&
- zone->compact_considered >= 1UL << zone->compact_defer_shift;
+ return zone->compact_defer_shift[true] == COMPACT_MAX_DEFER_SHIFT &&
+ zone->compact_considered[true] >= 1UL << zone->compact_defer_shift[true];
}
/* Returns true if the pageblock should be scanned for pages to isolate. */
@@ -2855,7 +2859,7 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
continue;
if (prio > MIN_COMPACT_PRIORITY
- && compaction_deferred(zone, order)) {
+ && compaction_deferred(zone, order, true)) {
rc = max_t(enum compact_result, COMPACT_DEFERRED, rc);
continue;
}
@@ -2893,7 +2897,7 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
* so we defer compaction there. If it ends up
* succeeding after all, it will be reset.
*/
- defer_compaction(zone, order);
+ defer_compaction(zone, order, true);
/*
* We might have stopped compacting due to need_resched() in
@@ -3111,7 +3115,7 @@ static void kcompactd_do_work(pg_data_t *pgdat)
if (!populated_zone(zone))
continue;
- if (compaction_deferred(zone, cc.order))
+ if (compaction_deferred(zone, cc.order, true))
continue;
ret = compaction_suit_allocation_order(zone,
@@ -3141,7 +3145,7 @@ static void kcompactd_do_work(pg_data_t *pgdat)
* We use sync migration mode here, so we defer like
* sync direct compaction does.
*/
- defer_compaction(zone, cc.order);
+ defer_compaction(zone, cc.order, true);
}
count_compact_events(KCOMPACTD_MIGRATE_SCANNED,
--
2.43.0
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH 2/2] mm/compaction: defer failed async direct compaction
2026-10-01 15:33 [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault Qiliang Yuan
2026-10-01 15:33 ` [PATCH 1/2] mm/compaction: keep compaction deferral state per migration mode Qiliang Yuan
@ 2026-10-01 15:33 ` Qiliang Yuan
2026-10-02 14:21 ` [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault Lorenzo Stoakes (ARM)
2 siblings, 0 replies; 6+ messages in thread
From: Qiliang Yuan @ 2026-10-01 15:33 UTC (permalink / raw)
To: Andrew Morton, Kairui Song, Qi Zheng, Shakeel Butt, Barry Song,
Axel Rasmussen, Yuanchu Xie, Wei Xu, Baoquan He, Baolin Wang,
David Hildenbrand, Lorenzo Stoakes, Liam R. Howlett,
Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
Steven Rostedt, Masami Hiramatsu, Mathieu Desnoyers,
Brendan Jackman, Johannes Weiner, Zi Yan
Cc: linux-mm, linux-kernel, linux-trace-kernel, Qiliang Yuan
A THP fault on a MADV_HUGEPAGE VMA first tries the local node only,
with __GFP_THISNODE | __GFP_NORETRY, and runs a single round of async
direct compaction there before falling back to other nodes. A failed
async run is never deferred.
When the local node has plenty of free memory but no free pageblock,
and what sits between the free pages can't be migrated, such as memory
long-term pinned for RDMA, each of these faults scans the zone again
and fails. Populating a large buffer pays for one failed compaction
per 2M fault. A KV-cache store that registers hundreds of GiB for RDMA
reports registration growing from tens of seconds to tens of minutes,
and works around it with MPOL_INTERLEAVE.
Defer the async state when an async run fails, and check it before
further async runs. Sync compaction keeps its own state and behaves as
before, and a successful compaction or allocation resets both. Keep
resetting the pageblock skip hints only when sync compaction restarts,
as async relies on them.
Populating a 4 GiB MADV_HUGEPAGE buffer from node 0 of a two-node VM,
with node 0 (24 GiB) fragmented by long-term pinning every other page,
median of 3 runs:
before after
direct compactions 735 15
time in compaction 122.7 ms 3.6 ms
population time 380 ms 273 ms
THPs on node 0 1313 1126
The THPs that no longer land on node 0 come from node 1. With the same
fragmentation but movable memory, compaction still succeeds and the
median run still gets all 2048 THPs on node 0, as before.
Signed-off-by: Qiliang Yuan <odys.yuan@gmail.com>
---
mm/compaction.c | 21 ++++++++++++++-------
1 file changed, 14 insertions(+), 7 deletions(-)
diff --git a/mm/compaction.c b/mm/compaction.c
index 7f8845d1990aa..25f758bc54f06 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -2600,7 +2600,9 @@ compact_zone(struct compact_control *cc, struct capture_control *capc)
/*
* Clear pageblock skip if there were failures recently and compaction
- * is about to be retried after being deferred.
+ * is about to be retried after being deferred. Only do it when sync
+ * compaction restarts: async compaction relies on the skip hints, and
+ * clearing them on every async retry would rescan the whole zone.
*/
if (compaction_restarting(cc->zone, cc->order))
__reset_isolation_suitable(cc->zone);
@@ -2858,8 +2860,10 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
!__cpuset_zone_allowed(zone, gfp_mask))
continue;
- if (prio > MIN_COMPACT_PRIORITY
- && compaction_deferred(zone, order, true)) {
+ if (prio > MIN_COMPACT_PRIORITY &&
+ (compaction_deferred(zone, order, true) ||
+ (prio == COMPACT_PRIO_ASYNC &&
+ compaction_deferred(zone, order, false)))) {
rc = max_t(enum compact_result, COMPACT_DEFERRED, rc);
continue;
}
@@ -2890,14 +2894,17 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order,
break;
}
- if (prio != COMPACT_PRIO_ASYNC && (status == COMPACT_COMPLETE ||
- status == COMPACT_PARTIAL_SKIPPED))
+ if (status == COMPACT_COMPLETE ||
+ status == COMPACT_PARTIAL_SKIPPED)
/*
* We think that allocation won't succeed in this zone
* so we defer compaction there. If it ends up
- * succeeding after all, it will be reset.
+ * succeeding after all, it will be reset. A failed
+ * async run only defers further async runs, as sync
+ * compaction may succeed on pageblocks it skipped.
*/
- defer_compaction(zone, order, true);
+ defer_compaction(zone, order,
+ prio != COMPACT_PRIO_ASYNC);
/*
* We might have stopped compacting due to need_resched() in
--
2.43.0
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault
2026-10-01 15:33 [PATCH 0/2] mm/compaction: stop repeating failed async compaction on every THP fault Qiliang Yuan
2026-10-01 15:33 ` [PATCH 1/2] mm/compaction: keep compaction deferral state per migration mode Qiliang Yuan
2026-10-01 15:33 ` [PATCH 2/2] mm/compaction: defer failed async direct compaction Qiliang Yuan
@ 2026-10-02 14:21 ` Lorenzo Stoakes (ARM)
2026-10-02 16:17 ` Qiliang Yuan
2 siblings, 1 reply; 6+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-10-02 14:21 UTC (permalink / raw)
To: Qiliang Yuan
Cc: Andrew Morton, Kairui Song, Qi Zheng, Shakeel Butt, Barry Song,
Axel Rasmussen, Yuanchu Xie, Wei Xu, Baoquan He, Baolin Wang,
David Hildenbrand, Liam R. Howlett, Vlastimil Babka,
Mike Rapoport, Suren Baghdasaryan, Michal Hocko, Steven Rostedt,
Masami Hiramatsu, Mathieu Desnoyers, Brendan Jackman,
Johannes Weiner, Zi Yan, linux-mm, linux-kernel,
linux-trace-kernel
Hi,
NAK.
This series is buggy, and has triggered my AI detection script.
It is kernel policy that you must disclose this with an
Assisted-by tag like:
Assisted-by: LLM
See https://docs.kernel.org/process/coding-assistants.html
It is also kernel policy that you must fully understand and take
responsibility for every patch that you send.
See https://docs.kernel.org/process/generated-content.html most
notably:
If tools permit you to generate a contribution automatically, expect
additional scrutiny in proportion to how much of it was generated.
As with the output of any tooling, the result may be incorrect or
inappropriate. You are expected to understand and to be able to
defend everything you submit. If you are unable to do so, then do
not submit the resulting changes.
If you do so anyway, maintainers are entitled to reject your series
without detailed review.
In general, if you are a newcomer to mm, we expect you to do smaller work
before moving on to larger changes, so you build understanding of both the
technical aspects of mm and how we do things.
Given you sent your first mail on 21st September and have since been
spamming complicated series across multiple different domains, I'm
absolutely not confident that you have any understanding of this.
https://lore.kernel.org/all/?q=f%3Aodys.yuan%40gmail.com
It also noted that you had a previous email, realwujing@gmail.com, which
you have silent switched from, which also bore the hallmarks of AI slop.
The script noted:
- Output rate: since 09-28 he has sent KVM SMM CR3/Hyper-V, ext4+jbd2+quota
"shrinker scan budget accounting" (the same fix applied to three shrinkers;
jbd2 went v1->v5 in 3 days and v1->v3 in 3 hours), blk-mq SRCU tag sets, a
blk-mq sync-run, ublk x2, a 4-patch bpf-next verifier rework, mm/migrate
move_pages (v2 sent 10h after v1) and this series. That is ~9 unrelated deep
areas in 4 days from someone with 4 commits.
Among many other indicators.
It also found several glaring flaws with your series, so it's clearly not
upstreamable.
I am going to have to ask you to focus on smaller changes as a newcomer,
and to acknowledge any LLM tools you have used, please.
Thanks.
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 6+ messages in thread