* [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order
@ 2026-10-06 9:18 Kiryl Shutsemau
2026-10-06 23:41 ` Andrew Morton
2026-10-07 15:53 ` Usama Arif
0 siblings, 2 replies; 5+ messages in thread
From: Kiryl Shutsemau @ 2026-10-06 9:18 UTC (permalink / raw)
To: Andrew Morton, Vlastimil Babka, Johannes Weiner,
David Hildenbrand, Harry Yoo
Cc: Kiryl Shutsemau (Meta),
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, linux-mm, linux-kernel, stable,
kernel-team
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
storm in defrag_mode"), direct reclaim and compaction for non-movable
requests under defrag_mode run at pageblock_order, to produce the whole
blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
still use the request order. An order-0 request can therefore retry
indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
- Reclaim at pageblock_order gives up after one pass as soon as a zone
looks compaction_ready(), and do_try_to_free_pages() then returns 1
even though nothing was reclaimed. It returns before the retry that
would reclaim memory.low-protected cgroups, so when most memory is
protected, the pass that did run finds next to nothing.
- Compaction at pageblock_order fails or is deferred.
- should_reclaim_retry() takes the reported progress as progress for
the order-0 request and resets no_progress_loops. The request
retries.
Order 1-3 requests loop the same way, and should_compact_retry() also
checks their pageblock_order compaction result against the request
order.
On a production host (64G, defrag_mode, memory.low covering most of the
workload), 95% of direct reclaim runs were order-9 runs that returned 1
with nothing reclaimed, at up to 60k runs per second. Across ~200M
should_reclaim_retry() calls in a day, no_progress_loops never left 0.
The spinning allocations were SLUB slab refills for inode and dentry
caches. The time spent registers as memory pressure, and pressure-based
OOM killing takes down both workloads and system services.
For promoted requests:
- Reclaim progress does not reset no_progress_loops, as for costly
orders.
- should_compact_retry() checks the compaction result at the promoted
order and does not retry COMPACT_SKIPPED, since the request can fall
back. The compaction priority floor and the COMPACT_SUCCESS retry
limit stay those of the request order, so a non-costly request still
gets its COMPACT_PRIO_SYNC_FULL pass before it falls back.
When the fallback is taken, reset the retry counters, so that the
fallback attempt gets a full retry budget before the OOM killer is
considered.
__alloc_pages_slowpath() computes the promoted order once per iteration
and passes it to direct reclaim, direct compaction and the two retry
helpers next to the request order, so no callee has to recompute it.
Harry Yoo asked for the two orders to be explicit rather than derived
in each callee.
In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
before after
should_reclaim_retry() calls 63M 293k
peak memory pressure (PSI some avg10) 99% 12%
File creation runs 5.7x faster.
Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
---
v2:
- keep the compaction priority floor and the COMPACT_SUCCESS retry
limit on the request order, so non-costly requests still escalate to
SYNC_FULL before falling back (Harry Yoo, Johannes Weiner)
- compute the promoted order once in __alloc_pages_slowpath() and pass
it to direct reclaim, direct compaction and the retry helpers, instead
of each of them deriving it (Harry Yoo)
mm/page_alloc.c | 123 ++++++++++++++++++++++++++++++------------------
1 file changed, 78 insertions(+), 45 deletions(-)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 12fac9084c48..41830d230f06 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -4127,6 +4127,32 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
return page;
}
+/*
+ * If fallbacks are not permitted (defrag_mode), we either need to
+ * reclaim space in a block of matching type, or clear out an entire
+ * block to allow __rmqueue_claim() to convert.
+ *
+ * Reclaim by itself is primarily freeing space in movable blocks,
+ * since that's where the LRU pages live. So this works for movable
+ * requests, but not for others.
+ *
+ * For those, promote the order of reclaim and compaction to help make
+ * blocks, instead of spinning in reclaim alone unproductively. Retry
+ * decisions based on the outcome of that work - reclaim progress and
+ * compaction results - must account for the promotion as well, so
+ * __alloc_pages_slowpath() computes the promoted order once and passes
+ * it alongside the request order.
+ */
+static inline unsigned int nofrag_promote_order(unsigned int order,
+ unsigned int alloc_flags,
+ const struct alloc_context *ac)
+{
+ if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
+ return max(order, pageblock_order);
+
+ return order;
+}
+
/*
* Maximum number of compaction retries with a progress before OOM
* killer is consider as the only way to move forward.
@@ -4137,8 +4163,9 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
/* Try memory compaction for high-order allocations before reclaim */
static struct page *
__alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
- unsigned int alloc_flags, const struct alloc_context *ac,
- enum compact_priority prio, enum compact_result *compact_result)
+ unsigned int compact_order, unsigned int alloc_flags,
+ const struct alloc_context *ac, enum compact_priority prio,
+ enum compact_result *compact_result)
{
struct page *page = NULL;
unsigned long pflags;
@@ -4149,22 +4176,6 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
.order = order,
.page = NULL,
};
- int compact_order = order;
-
- /*
- * If fallbacks are not permitted (defrag_mode), we either
- * need to reclaim space in a block of matching type, or clear
- * out an entire block to allow __rmqueue_claim() to convert.
- *
- * Reclaim by itself is primarily freeing space in movable
- * blocks, since that's where the LRU pages live. So this
- * works for movable requests, but not for others.
- *
- * For those, promote the order to help make blocks, instead
- * of spinning in reclaim alone unproductively.
- */
- if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
- compact_order = max(order, pageblock_order);
if (!compact_order)
return NULL;
@@ -4246,7 +4257,7 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
static inline bool
should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
- int alloc_flags,
+ int compact_order, int alloc_flags,
enum compact_result compact_result,
enum compact_priority *compact_priority,
int *compaction_retries)
@@ -4257,7 +4268,7 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
int retries = *compaction_retries;
enum compact_priority priority = *compact_priority;
- if (!order)
+ if (!compact_order)
return false;
if (fatal_signal_pending(current))
@@ -4266,10 +4277,14 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
/*
* Compaction was skipped due to a lack of free order-0
* migration targets. Continue if reclaim can help.
+ *
+ * Promoted requests have exhausted their reclaim retries at
+ * this point, and they can fall back instead.
*/
if (compact_result == COMPACT_SKIPPED) {
- ret = compaction_zonelist_suitable(ac, order, alloc_flags,
- gfp_mask);
+ if (compact_order == order)
+ ret = compaction_zonelist_suitable(ac, order, alloc_flags,
+ gfp_mask);
goto out;
}
@@ -4314,8 +4329,9 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
#else
static inline struct page *
__alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
- unsigned int alloc_flags, const struct alloc_context *ac,
- enum compact_priority prio, enum compact_result *compact_result)
+ unsigned int compact_order, unsigned int alloc_flags,
+ const struct alloc_context *ac, enum compact_priority prio,
+ enum compact_result *compact_result)
{
*compact_result = COMPACT_SKIPPED;
return NULL;
@@ -4323,7 +4339,7 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
static inline bool
should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
- int alloc_flags,
+ int compact_order, int alloc_flags,
enum compact_result compact_result,
enum compact_priority *compact_priority,
int *compaction_retries)
@@ -4462,17 +4478,12 @@ __perform_reclaim(gfp_t gfp_mask, unsigned int order,
/* The really slow allocator path where we enter direct reclaim */
static inline struct page *
__alloc_pages_direct_reclaim(gfp_t gfp_mask, unsigned int order,
- unsigned int alloc_flags, const struct alloc_context *ac,
- unsigned long *did_some_progress)
+ unsigned int reclaim_order, unsigned int alloc_flags,
+ const struct alloc_context *ac, unsigned long *did_some_progress)
{
struct page *page = NULL;
unsigned long pflags;
bool drained = false;
- int reclaim_order = order;
-
- /* Match the slowpath compaction promotion in __alloc_pages_direct_compact */
- if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
- reclaim_order = max(order, pageblock_order);
psi_memstall_enter(&pflags);
*did_some_progress = __perform_reclaim(gfp_mask, reclaim_order, ac);
@@ -4638,8 +4649,9 @@ bool gfp_pfmemalloc_allowed(gfp_t gfp_mask)
*/
static inline bool
should_reclaim_retry(gfp_t gfp_mask, unsigned order,
- struct alloc_context *ac, int alloc_flags,
- bool did_some_progress, int *no_progress_loops)
+ unsigned int reclaim_order, struct alloc_context *ac,
+ int alloc_flags, bool did_some_progress,
+ int *no_progress_loops)
{
struct zone *zone;
struct zoneref *z;
@@ -4648,9 +4660,17 @@ should_reclaim_retry(gfp_t gfp_mask, unsigned order,
/*
* Costly allocations might have made a progress but this doesn't mean
* their order will become available due to high fragmentation so
- * always increment the no progress counter for them
+ * always increment the no progress counter for them.
+ *
+ * The same goes for requests whose reclaim is promoted to make whole
+ * blocks. At that order, reclaim also reports progress when it backs
+ * off for compaction without freeing anything.
+ *
+ * The watermark check below stays at the request order: it asks
+ * whether the request itself could succeed after reclaim.
*/
- if (did_some_progress && order <= PAGE_ALLOC_COSTLY_ORDER)
+ if (did_some_progress && order <= PAGE_ALLOC_COSTLY_ORDER &&
+ reclaim_order == order)
*no_progress_loops = 0;
else
(*no_progress_loops)++;
@@ -4790,6 +4810,7 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
const bool costly_order = order > PAGE_ALLOC_COSTLY_ORDER;
struct page *page = NULL;
unsigned int alloc_flags;
+ unsigned int reclaim_order;
unsigned long did_some_progress;
enum compact_priority compact_priority;
enum compact_result compact_result;
@@ -4934,17 +4955,22 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
/* If allocation has taken excessively long, warn about it */
check_alloc_stall_warn(gfp_mask, ac->nodemask, order, alloc_start_time);
+ /* The order reclaim and compaction work at, see nofrag_promote_order() */
+ reclaim_order = nofrag_promote_order(order, alloc_flags, ac);
+
/* Try direct reclaim and then allocating */
if (!compact_first) {
- page = __alloc_pages_direct_reclaim(gfp_mask, order, alloc_flags,
- ac, &did_some_progress);
+ page = __alloc_pages_direct_reclaim(gfp_mask, order, reclaim_order,
+ alloc_flags, ac,
+ &did_some_progress);
if (page)
goto got_pg;
}
/* Try direct compaction and then allocating */
- page = __alloc_pages_direct_compact(gfp_mask, order, alloc_flags, ac,
- compact_priority, &compact_result);
+ page = __alloc_pages_direct_compact(gfp_mask, order, reclaim_order,
+ alloc_flags, ac, compact_priority,
+ &compact_result);
if (page)
goto got_pg;
@@ -4996,8 +5022,9 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
check_retry_zonelist(zonelist_iter_cookie))
goto restart;
- if (should_reclaim_retry(gfp_mask, order, ac, alloc_flags,
- did_some_progress > 0, &no_progress_loops))
+ if (should_reclaim_retry(gfp_mask, order, reclaim_order, ac,
+ alloc_flags, did_some_progress > 0,
+ &no_progress_loops))
goto retry;
/*
@@ -5007,14 +5034,20 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
* of free memory (see __compaction_suitable)
*/
if (did_some_progress > 0 && can_compact &&
- should_compact_retry(gfp_mask, ac, order, alloc_flags,
- compact_result, &compact_priority,
+ should_compact_retry(gfp_mask, ac, order, reclaim_order,
+ alloc_flags, compact_result, &compact_priority,
&compaction_retries))
goto retry;
- /* Reclaim/compaction failed to prevent the fallback */
+ /*
+ * Reclaim/compaction failed to prevent the fallback. The retry
+ * budget was spent on making blocks, not on the request itself;
+ * give the fallback a fresh one before considering OOM.
+ */
if (defrag_mode && (alloc_flags & ALLOC_NOFRAGMENT)) {
alloc_flags &= ~ALLOC_NOFRAGMENT;
+ no_progress_loops = 0;
+ compaction_retries = 0;
goto retry;
}
--
2.54.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-10-06 9:18 [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order Kiryl Shutsemau
@ 2026-10-06 23:41 ` Andrew Morton
2026-10-07 13:01 ` Kiryl Shutsemau
2026-10-07 15:53 ` Usama Arif
1 sibling, 1 reply; 5+ messages in thread
From: Andrew Morton @ 2026-10-06 23:41 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Vlastimil Babka, Johannes Weiner, David Hildenbrand, Harry Yoo,
Kiryl Shutsemau (Meta),
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, linux-mm, linux-kernel, stable,
kernel-team
On Tue, 6 Oct 2026 10:18:13 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> storm in defrag_mode"), direct reclaim and compaction for non-movable
> requests under defrag_mode run at pageblock_order, to produce the whole
> blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
> still use the request order. An order-0 request can therefore retry
> indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
>
> ...
>
> In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
>
> before after
> should_reclaim_retry() calls 63M 293k
> peak memory pressure (PSI some avg10) 99% 12%
>
> File creation runs 5.7x faster.
> Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> ---
> v2:
> - keep the compaction priority floor and the COMPACT_SUCCESS retry
> limit on the request order, so non-costly requests still escalate to
> SYNC_FULL before falling back (Harry Yoo, Johannes Weiner)
> - compute the promoted order once in __alloc_pages_slowpath() and pass
> it to direct reclaim, direct compaction and the retry helpers, instead
> of each of them deriving it (Harry Yoo)
Thanks. That's a fairly large update. v2 diff is below.
Are we sure about cc:stable? Or even 7.3-rcX? It's a pretty big
change.
mm/page_alloc.c | 62 +++++++++++++++++++++++++---------------------
1 file changed, 34 insertions(+), 28 deletions(-)
--- a/mm/page_alloc.c~a
+++ a/mm/page_alloc.c
@@ -4139,8 +4139,9 @@ out:
* For those, promote the order of reclaim and compaction to help make
* blocks, instead of spinning in reclaim alone unproductively. Retry
* decisions based on the outcome of that work - reclaim progress and
- * compaction results - must account for the promotion as well, see
- * should_reclaim_retry() and should_compact_retry().
+ * compaction results - must account for the promotion as well, so
+ * __alloc_pages_slowpath() computes the promoted order once and passes
+ * it alongside the request order.
*/
static inline unsigned int nofrag_promote_order(unsigned int order,
unsigned int alloc_flags,
@@ -4162,8 +4163,9 @@ static inline unsigned int nofrag_promot
/* Try memory compaction for high-order allocations before reclaim */
static struct page *
__alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
- unsigned int alloc_flags, const struct alloc_context *ac,
- enum compact_priority prio, enum compact_result *compact_result)
+ unsigned int compact_order, unsigned int alloc_flags,
+ const struct alloc_context *ac, enum compact_priority prio,
+ enum compact_result *compact_result)
{
struct page *page = NULL;
unsigned long pflags;
@@ -4174,7 +4176,6 @@ __alloc_pages_direct_compact(gfp_t gfp_m
.order = order,
.page = NULL,
};
- unsigned int compact_order = nofrag_promote_order(order, alloc_flags, ac);
if (!compact_order)
return NULL;
@@ -4256,7 +4257,7 @@ __alloc_pages_direct_compact(gfp_t gfp_m
static inline bool
should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
- int alloc_flags,
+ int compact_order, int alloc_flags,
enum compact_result compact_result,
enum compact_priority *compact_priority,
int *compaction_retries)
@@ -4266,10 +4267,7 @@ should_compact_retry(gfp_t gfp_mask, str
bool ret = false;
int retries = *compaction_retries;
enum compact_priority priority = *compact_priority;
- unsigned int compact_order;
- /* Check the compaction result at the order compaction ran at */
- compact_order = nofrag_promote_order(order, alloc_flags, ac);
if (!compact_order)
return false;
@@ -4304,7 +4302,7 @@ should_compact_retry(gfp_t gfp_mask, str
* need much more detailed feedback from compaction to
* make a better decision.
*/
- if (compact_order > PAGE_ALLOC_COSTLY_ORDER)
+ if (order > PAGE_ALLOC_COSTLY_ORDER)
max_retries /= 4;
if (++(*compaction_retries) <= max_retries) {
@@ -4316,7 +4314,7 @@ should_compact_retry(gfp_t gfp_mask, str
/*
* Compaction failed. Retry with increasing priority.
*/
- min_priority = (compact_order > PAGE_ALLOC_COSTLY_ORDER) ?
+ min_priority = (order > PAGE_ALLOC_COSTLY_ORDER) ?
MIN_COMPACT_COSTLY_PRIORITY : MIN_COMPACT_PRIORITY;
if (*compact_priority > min_priority) {
@@ -4331,8 +4329,9 @@ out:
#else
static inline struct page *
__alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
- unsigned int alloc_flags, const struct alloc_context *ac,
- enum compact_priority prio, enum compact_result *compact_result)
+ unsigned int compact_order, unsigned int alloc_flags,
+ const struct alloc_context *ac, enum compact_priority prio,
+ enum compact_result *compact_result)
{
*compact_result = COMPACT_SKIPPED;
return NULL;
@@ -4340,7 +4339,7 @@ __alloc_pages_direct_compact(gfp_t gfp_m
static inline bool
should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
- int alloc_flags,
+ int compact_order, int alloc_flags,
enum compact_result compact_result,
enum compact_priority *compact_priority,
int *compaction_retries)
@@ -4479,13 +4478,12 @@ __perform_reclaim(gfp_t gfp_mask, unsign
/* The really slow allocator path where we enter direct reclaim */
static inline struct page *
__alloc_pages_direct_reclaim(gfp_t gfp_mask, unsigned int order,
- unsigned int alloc_flags, const struct alloc_context *ac,
- unsigned long *did_some_progress)
+ unsigned int reclaim_order, unsigned int alloc_flags,
+ const struct alloc_context *ac, unsigned long *did_some_progress)
{
struct page *page = NULL;
unsigned long pflags;
bool drained = false;
- unsigned int reclaim_order = nofrag_promote_order(order, alloc_flags, ac);
psi_memstall_enter(&pflags);
*did_some_progress = __perform_reclaim(gfp_mask, reclaim_order, ac);
@@ -4651,8 +4649,9 @@ bool gfp_pfmemalloc_allowed(gfp_t gfp_ma
*/
static inline bool
should_reclaim_retry(gfp_t gfp_mask, unsigned order,
- struct alloc_context *ac, int alloc_flags,
- bool did_some_progress, int *no_progress_loops)
+ unsigned int reclaim_order, struct alloc_context *ac,
+ int alloc_flags, bool did_some_progress,
+ int *no_progress_loops)
{
struct zone *zone;
struct zoneref *z;
@@ -4671,7 +4670,7 @@ should_reclaim_retry(gfp_t gfp_mask, uns
* whether the request itself could succeed after reclaim.
*/
if (did_some_progress && order <= PAGE_ALLOC_COSTLY_ORDER &&
- nofrag_promote_order(order, alloc_flags, ac) == order)
+ reclaim_order == order)
*no_progress_loops = 0;
else
(*no_progress_loops)++;
@@ -4811,6 +4810,7 @@ __alloc_pages_slowpath(gfp_t gfp_mask, u
const bool costly_order = order > PAGE_ALLOC_COSTLY_ORDER;
struct page *page = NULL;
unsigned int alloc_flags;
+ unsigned int reclaim_order;
unsigned long did_some_progress;
enum compact_priority compact_priority;
enum compact_result compact_result;
@@ -4955,17 +4955,22 @@ retry:
/* If allocation has taken excessively long, warn about it */
check_alloc_stall_warn(gfp_mask, ac->nodemask, order, alloc_start_time);
+ /* The order reclaim and compaction work at, see nofrag_promote_order() */
+ reclaim_order = nofrag_promote_order(order, alloc_flags, ac);
+
/* Try direct reclaim and then allocating */
if (!compact_first) {
- page = __alloc_pages_direct_reclaim(gfp_mask, order, alloc_flags,
- ac, &did_some_progress);
+ page = __alloc_pages_direct_reclaim(gfp_mask, order, reclaim_order,
+ alloc_flags, ac,
+ &did_some_progress);
if (page)
goto got_pg;
}
/* Try direct compaction and then allocating */
- page = __alloc_pages_direct_compact(gfp_mask, order, alloc_flags, ac,
- compact_priority, &compact_result);
+ page = __alloc_pages_direct_compact(gfp_mask, order, reclaim_order,
+ alloc_flags, ac, compact_priority,
+ &compact_result);
if (page)
goto got_pg;
@@ -5017,8 +5022,9 @@ retry:
check_retry_zonelist(zonelist_iter_cookie))
goto restart;
- if (should_reclaim_retry(gfp_mask, order, ac, alloc_flags,
- did_some_progress > 0, &no_progress_loops))
+ if (should_reclaim_retry(gfp_mask, order, reclaim_order, ac,
+ alloc_flags, did_some_progress > 0,
+ &no_progress_loops))
goto retry;
/*
@@ -5028,8 +5034,8 @@ retry:
* of free memory (see __compaction_suitable)
*/
if (did_some_progress > 0 && can_compact &&
- should_compact_retry(gfp_mask, ac, order, alloc_flags,
- compact_result, &compact_priority,
+ should_compact_retry(gfp_mask, ac, order, reclaim_order,
+ alloc_flags, compact_result, &compact_priority,
&compaction_retries))
goto retry;
_
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-10-06 23:41 ` Andrew Morton
@ 2026-10-07 13:01 ` Kiryl Shutsemau
2026-10-07 18:32 ` Andrew Morton
0 siblings, 1 reply; 5+ messages in thread
From: Kiryl Shutsemau @ 2026-10-07 13:01 UTC (permalink / raw)
To: Andrew Morton
Cc: Vlastimil Babka, Johannes Weiner, David Hildenbrand, Harry Yoo,
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, linux-mm, linux-kernel, stable,
kernel-team
On Tue, Oct 06, 2026 at 04:41:22PM -0700, Andrew Morton wrote:
> On Tue, 6 Oct 2026 10:18:13 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
>
> > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> >
> > Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> > storm in defrag_mode"), direct reclaim and compaction for non-movable
> > requests under defrag_mode run at pageblock_order, to produce the whole
> > blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
> > still use the request order. An order-0 request can therefore retry
> > indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
> >
> > ...
> >
> > In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
> >
> > before after
> > should_reclaim_retry() calls 63M 293k
> > peak memory pressure (PSI some avg10) 99% 12%
> >
> > File creation runs 5.7x faster.
>
> > Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
> > Cc: stable@vger.kernel.org
> > Assisted-by: LLM
> > Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> > ---
> > v2:
> > - keep the compaction priority floor and the COMPACT_SUCCESS retry
> > limit on the request order, so non-costly requests still escalate to
> > SYNC_FULL before falling back (Harry Yoo, Johannes Weiner)
> > - compute the promoted order once in __alloc_pages_slowpath() and pass
> > it to direct reclaim, direct compaction and the retry helpers, instead
> > of each of them deriving it (Harry Yoo)
>
> Thanks. That's a fairly large update. v2 diff is below.
>
> Are we sure about cc:stable? Or even 7.3-rcX? It's a pretty big
> change.
I think defrag_mode is not widely used. Maybe only in Meta.
What about keeping cc:stable, but only submit it to Linus for v7.4-rc1?
By the time we would have the patch battle-tested in the fleet.
--
Kiryl Shutsemau / Kirill A. Shutemov
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-10-06 9:18 [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order Kiryl Shutsemau
2026-10-06 23:41 ` Andrew Morton
@ 2026-10-07 15:53 ` Usama Arif
1 sibling, 0 replies; 5+ messages in thread
From: Usama Arif @ 2026-10-07 15:53 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Usama Arif, Andrew Morton, Vlastimil Babka, Johannes Weiner,
David Hildenbrand, Harry Yoo, Kiryl Shutsemau (Meta),
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, linux-mm, linux-kernel, stable, kernel-team
On Tue, 6 Oct 2026 10:18:13 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> storm in defrag_mode"), direct reclaim and compaction for non-movable
> requests under defrag_mode run at pageblock_order, to produce the whole
> blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
> still use the request order. An order-0 request can therefore retry
> indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
>
> - Reclaim at pageblock_order gives up after one pass as soon as a zone
> looks compaction_ready(), and do_try_to_free_pages() then returns 1
> even though nothing was reclaimed. It returns before the retry that
> would reclaim memory.low-protected cgroups, so when most memory is
> protected, the pass that did run finds next to nothing.
>
> - Compaction at pageblock_order fails or is deferred.
>
> - should_reclaim_retry() takes the reported progress as progress for
> the order-0 request and resets no_progress_loops. The request
> retries.
>
> Order 1-3 requests loop the same way, and should_compact_retry() also
> checks their pageblock_order compaction result against the request
> order.
>
> On a production host (64G, defrag_mode, memory.low covering most of the
> workload), 95% of direct reclaim runs were order-9 runs that returned 1
> with nothing reclaimed, at up to 60k runs per second. Across ~200M
> should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> The spinning allocations were SLUB slab refills for inode and dentry
> caches. The time spent registers as memory pressure, and pressure-based
> OOM killing takes down both workloads and system services.
>
> For promoted requests:
>
> - Reclaim progress does not reset no_progress_loops, as for costly
> orders.
>
> - should_compact_retry() checks the compaction result at the promoted
> order and does not retry COMPACT_SKIPPED, since the request can fall
> back. The compaction priority floor and the COMPACT_SUCCESS retry
> limit stay those of the request order, so a non-costly request still
> gets its COMPACT_PRIO_SYNC_FULL pass before it falls back.
>
> When the fallback is taken, reset the retry counters, so that the
> fallback attempt gets a full retry budget before the OOM killer is
> considered.
>
> __alloc_pages_slowpath() computes the promoted order once per iteration
> and passes it to direct reclaim, direct compaction and the two retry
> helpers next to the request order, so no callee has to recompute it.
> Harry Yoo asked for the two orders to be explicit rather than derived
> in each callee.
>
> In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
>
> before after
> should_reclaim_retry() calls 63M 293k
> peak memory pressure (PSI some avg10) 99% 12%
>
> File creation runs 5.7x faster.
>
> Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
[..]
> @@ -5007,14 +5034,20 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
> * of free memory (see __compaction_suitable)
> */
> if (did_some_progress > 0 && can_compact &&
> - should_compact_retry(gfp_mask, ac, order, alloc_flags,
> - compact_result, &compact_priority,
> + should_compact_retry(gfp_mask, ac, order, reclaim_order,
> + alloc_flags, compact_result, &compact_priority,
> &compaction_retries))
> goto retry;
>
> - /* Reclaim/compaction failed to prevent the fallback */
> + /*
> + * Reclaim/compaction failed to prevent the fallback. The retry
> + * budget was spent on making blocks, not on the request itself;
> + * give the fallback a fresh one before considering OOM.
> + */
> if (defrag_mode && (alloc_flags & ALLOC_NOFRAGMENT)) {
> alloc_flags &= ~ALLOC_NOFRAGMENT;
> + no_progress_loops = 0;
> + compaction_retries = 0;
Should these counters be reset only when the work order was actually
promoted?
For movable requests, and requests already at or above pageblock
order, `reclaim_order == order`. Their retry budget was therefore
spent on the request itself rather than on promoted pageblock
production. Resetting the counters here can give those requests
another 17 no-progress reclaim attempts, plus additional
compaction-success retries, before OOM or allocation failure.
How about:
if (reclaim_order != order) {
no_progress_loops = 0;
compaction_retries = 0;
}
instead?
> goto retry;
> }
>
> --
> 2.54.0
>
>
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-10-07 13:01 ` Kiryl Shutsemau
@ 2026-10-07 18:32 ` Andrew Morton
0 siblings, 0 replies; 5+ messages in thread
From: Andrew Morton @ 2026-10-07 18:32 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Vlastimil Babka, Johannes Weiner, David Hildenbrand, Harry Yoo,
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, linux-mm, linux-kernel, stable,
kernel-team
On Wed, 7 Oct 2026 14:01:33 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
> > > v2:
> > > - keep the compaction priority floor and the COMPACT_SUCCESS retry
> > > limit on the request order, so non-costly requests still escalate to
> > > SYNC_FULL before falling back (Harry Yoo, Johannes Weiner)
> > > - compute the promoted order once in __alloc_pages_slowpath() and pass
> > > it to direct reclaim, direct compaction and the retry helpers, instead
> > > of each of them deriving it (Harry Yoo)
> >
> > Thanks. That's a fairly large update. v2 diff is below.
> >
> > Are we sure about cc:stable? Or even 7.3-rcX? It's a pretty big
> > change.
>
> I think defrag_mode is not widely used. Maybe only in Meta.
>
> What about keeping cc:stable, but only submit it to Linus for v7.4-rc1?
>
> By the time we would have the patch battle-tested in the fleet.
OK, let's do that.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-07 18:32 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 9:18 [PATCH v2] mm: page_alloc: make defrag_mode retries follow the promoted order Kiryl Shutsemau
2026-10-06 23:41 ` Andrew Morton
2026-10-07 13:01 ` Kiryl Shutsemau
2026-10-07 18:32 ` Andrew Morton
2026-10-07 15:53 ` Usama Arif
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®