* [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order
@ 2026-09-29 17:45 Kiryl Shutsemau
2026-09-29 18:39 ` Harry Yoo
2026-09-29 19:58 ` Andrew Morton
0 siblings, 2 replies; 7+ messages in thread
From: Kiryl Shutsemau @ 2026-09-29 17:45 UTC (permalink / raw)
To: Andrew Morton, Vlastimil Babka, Johannes Weiner, David Hildenbrand
Cc: Kiryl Shutsemau (Meta),
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, Harry Yoo, linux-mm, linux-kernel,
stable, kernel-team
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
storm in defrag_mode"), direct reclaim and compaction for non-movable
requests under defrag_mode run at pageblock_order, to produce the whole
blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
still use the request order. An order-0 request can therefore retry
indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
- Reclaim at pageblock_order gives up after one pass as soon as a zone
looks compaction_ready(), and do_try_to_free_pages() then returns 1
even though nothing was reclaimed. It returns before the retry that
would reclaim memory.low-protected cgroups, so when most memory is
protected, the pass that did run finds next to nothing.
- Compaction at pageblock_order fails or is deferred.
- should_reclaim_retry() takes the reported progress as progress for
the order-0 request and resets no_progress_loops. The request
retries.
Order 1-3 requests loop the same way, and should_compact_retry() also
checks their pageblock_order compaction result against the request
order.
On a production host (64G, defrag_mode, memory.low covering most of the
workload), 95% of direct reclaim runs were order-9 runs that returned 1
with nothing reclaimed, at up to 60k runs per second. Across ~200M
should_reclaim_retry() calls in a day, no_progress_loops never left 0.
The spinning allocations were SLUB slab refills for inode and dentry
caches. The time spent registers as memory pressure, and pressure-based
OOM killing takes down both workloads and system services.
Treat promoted requests like costly orders:
- Reclaim progress does not reset no_progress_loops for them.
- should_compact_retry() checks the compaction result at the promoted
order. It does not retry COMPACT_SKIPPED, since the request can fall
back, and it does not escalate compaction to COMPACT_PRIO_SYNC_FULL.
When the fallback is taken, reset the retry counters, so that the
fallback attempt gets a full retry budget before the OOM killer is
considered.
In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
before after
should_reclaim_retry() calls 63M 293k
peak memory pressure (PSI some avg10) 99% 12%
File creation runs 5.7x faster.
Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
Cc: <stable@vger.kernel.org>
Assisted-by: LLM
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
---
mm/page_alloc.c | 85 ++++++++++++++++++++++++++++++++-----------------
1 file changed, 56 insertions(+), 29 deletions(-)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 12fac9084c48..608487672d93 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -4127,6 +4127,31 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
return page;
}
+/*
+ * If fallbacks are not permitted (defrag_mode), we either need to
+ * reclaim space in a block of matching type, or clear out an entire
+ * block to allow __rmqueue_claim() to convert.
+ *
+ * Reclaim by itself is primarily freeing space in movable blocks,
+ * since that's where the LRU pages live. So this works for movable
+ * requests, but not for others.
+ *
+ * For those, promote the order of reclaim and compaction to help make
+ * blocks, instead of spinning in reclaim alone unproductively. Retry
+ * decisions based on the outcome of that work - reclaim progress and
+ * compaction results - must account for the promotion as well, see
+ * should_reclaim_retry() and should_compact_retry().
+ */
+static inline unsigned int nofrag_promote_order(unsigned int order,
+ unsigned int alloc_flags,
+ const struct alloc_context *ac)
+{
+ if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
+ return max(order, pageblock_order);
+
+ return order;
+}
+
/*
* Maximum number of compaction retries with a progress before OOM
* killer is consider as the only way to move forward.
@@ -4149,22 +4174,7 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
.order = order,
.page = NULL,
};
- int compact_order = order;
-
- /*
- * If fallbacks are not permitted (defrag_mode), we either
- * need to reclaim space in a block of matching type, or clear
- * out an entire block to allow __rmqueue_claim() to convert.
- *
- * Reclaim by itself is primarily freeing space in movable
- * blocks, since that's where the LRU pages live. So this
- * works for movable requests, but not for others.
- *
- * For those, promote the order to help make blocks, instead
- * of spinning in reclaim alone unproductively.
- */
- if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
- compact_order = max(order, pageblock_order);
+ unsigned int compact_order = nofrag_promote_order(order, alloc_flags, ac);
if (!compact_order)
return NULL;
@@ -4256,8 +4266,11 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
bool ret = false;
int retries = *compaction_retries;
enum compact_priority priority = *compact_priority;
+ unsigned int compact_order;
- if (!order)
+ /* Check the compaction result at the order compaction ran at */
+ compact_order = nofrag_promote_order(order, alloc_flags, ac);
+ if (!compact_order)
return false;
if (fatal_signal_pending(current))
@@ -4266,10 +4279,14 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
/*
* Compaction was skipped due to a lack of free order-0
* migration targets. Continue if reclaim can help.
+ *
+ * Promoted requests have exhausted their reclaim retries at
+ * this point, and they can fall back instead.
*/
if (compact_result == COMPACT_SKIPPED) {
- ret = compaction_zonelist_suitable(ac, order, alloc_flags,
- gfp_mask);
+ if (compact_order == order)
+ ret = compaction_zonelist_suitable(ac, order, alloc_flags,
+ gfp_mask);
goto out;
}
@@ -4287,7 +4304,7 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
* need much more detailed feedback from compaction to
* make a better decision.
*/
- if (order > PAGE_ALLOC_COSTLY_ORDER)
+ if (compact_order > PAGE_ALLOC_COSTLY_ORDER)
max_retries /= 4;
if (++(*compaction_retries) <= max_retries) {
@@ -4299,7 +4316,7 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
/*
* Compaction failed. Retry with increasing priority.
*/
- min_priority = (order > PAGE_ALLOC_COSTLY_ORDER) ?
+ min_priority = (compact_order > PAGE_ALLOC_COSTLY_ORDER) ?
MIN_COMPACT_COSTLY_PRIORITY : MIN_COMPACT_PRIORITY;
if (*compact_priority > min_priority) {
@@ -4468,11 +4485,7 @@ __alloc_pages_direct_reclaim(gfp_t gfp_mask, unsigned int order,
struct page *page = NULL;
unsigned long pflags;
bool drained = false;
- int reclaim_order = order;
-
- /* Match the slowpath compaction promotion in __alloc_pages_direct_compact */
- if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
- reclaim_order = max(order, pageblock_order);
+ unsigned int reclaim_order = nofrag_promote_order(order, alloc_flags, ac);
psi_memstall_enter(&pflags);
*did_some_progress = __perform_reclaim(gfp_mask, reclaim_order, ac);
@@ -4648,9 +4661,17 @@ should_reclaim_retry(gfp_t gfp_mask, unsigned order,
/*
* Costly allocations might have made a progress but this doesn't mean
* their order will become available due to high fragmentation so
- * always increment the no progress counter for them
+ * always increment the no progress counter for them.
+ *
+ * The same goes for requests whose reclaim is promoted to make whole
+ * blocks. At that order, reclaim also reports progress when it backs
+ * off for compaction without freeing anything.
+ *
+ * The watermark check below stays at the request order: it asks
+ * whether the request itself could succeed after reclaim.
*/
- if (did_some_progress && order <= PAGE_ALLOC_COSTLY_ORDER)
+ if (did_some_progress && order <= PAGE_ALLOC_COSTLY_ORDER &&
+ nofrag_promote_order(order, alloc_flags, ac) == order)
*no_progress_loops = 0;
else
(*no_progress_loops)++;
@@ -5012,9 +5033,15 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
&compaction_retries))
goto retry;
- /* Reclaim/compaction failed to prevent the fallback */
+ /*
+ * Reclaim/compaction failed to prevent the fallback. The retry
+ * budget was spent on making blocks, not on the request itself;
+ * give the fallback a fresh one before considering OOM.
+ */
if (defrag_mode && (alloc_flags & ALLOC_NOFRAGMENT)) {
alloc_flags &= ~ALLOC_NOFRAGMENT;
+ no_progress_loops = 0;
+ compaction_retries = 0;
goto retry;
}
--
2.54.0
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-09-29 17:45 [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order Kiryl Shutsemau
@ 2026-09-29 18:39 ` Harry Yoo
2026-09-30 13:32 ` Kiryl Shutsemau
2026-09-29 19:58 ` Andrew Morton
1 sibling, 1 reply; 7+ messages in thread
From: Harry Yoo @ 2026-09-29 18:39 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Andrew Morton, Vlastimil Babka, Johannes Weiner,
David Hildenbrand, Kiryl Shutsemau (Meta),
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, linux-mm, linux-kernel, stable,
kernel-team
On Tue, Sep 29, 2026 at 06:45:51PM +0100, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> storm in defrag_mode"), direct reclaim and compaction for non-movable
> requests under defrag_mode run at pageblock_order, to produce the whole
> blocks that ALLOC_NOFRAGMENT needs.
> The retry decisions that follow still use the request order.
Indeed, good catch!
> An order-0 request can therefore retry
> indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
>
> - Reclaim at pageblock_order gives up after one pass as soon as a zone
> looks compaction_ready(), and do_try_to_free_pages() then returns 1
> even though nothing was reclaimed. It returns before the retry that
> would reclaim memory.low-protected cgroups, so when most memory is
> protected, the pass that did run finds next to nothing.
>
> - Compaction at pageblock_order fails or is deferred.
>
> - should_reclaim_retry() takes the reported progress as progress for
> the order-0 request and resets no_progress_loops. The request
> retries.
Makes sense to me.
> Order 1-3 requests loop the same way, and should_compact_retry() also
> checks their pageblock_order compaction result against the request
> order.
>
> On a production host (64G, defrag_mode, memory.low covering most of the
> workload), 95% of direct reclaim runs were order-9 runs that returned 1
> with nothing reclaimed, at up to 60k runs per second. Across ~200M
> should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> The spinning allocations were SLUB slab refills for inode and dentry
> caches. The time spent registers as memory pressure, and pressure-based
> OOM killing takes down both workloads and system services.
>
> Treat promoted requests like costly orders:
>
> - Reclaim progress does not reset no_progress_loops for them.
>
> - should_compact_retry() checks the compaction result at the promoted
> order. It does not retry COMPACT_SKIPPED, since the request can fall
> back, and it does not escalate compaction to COMPACT_PRIO_SYNC_FULL.
>
> When the fallback is taken, reset the retry counters, so that the
> fallback attempt gets a full retry budget before the OOM killer is
> considered.
>
> In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
>
> before after
> should_reclaim_retry() calls 63M 293k
> peak memory pressure (PSI some avg10) 99% 12%
>
> File creation runs 5.7x faster.
>
> Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
> Cc: <stable@vger.kernel.org>
> Assisted-by: LLM
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> ---
> mm/page_alloc.c | 85 ++++++++++++++++++++++++++++++++-----------------
> 1 file changed, 56 insertions(+), 29 deletions(-)
>
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index 12fac9084c48..608487672d93 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -4127,6 +4127,31 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> return page;
> }
>
> +/*
> + * If fallbacks are not permitted (defrag_mode), we either need to
> + * reclaim space in a block of matching type, or clear out an entire
> + * block to allow __rmqueue_claim() to convert.
> + *
> + * Reclaim by itself is primarily freeing space in movable blocks,
> + * since that's where the LRU pages live. So this works for movable
> + * requests, but not for others.
> + *
> + * For those, promote the order of reclaim and compaction to help make
> + * blocks, instead of spinning in reclaim alone unproductively. Retry
> + * decisions based on the outcome of that work - reclaim progress and
> + * compaction results - must account for the promotion as well, see
> + * should_reclaim_retry() and should_compact_retry().
> + */
> +static inline unsigned int nofrag_promote_order(unsigned int order,
> + unsigned int alloc_flags,
> + const struct alloc_context *ac)
> +{
> + if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
> + return max(order, pageblock_order);
> +
> + return order;
> +}
I think we should start distinguishing order and compact/reclaim_order
in __alloc_pages_slowpath(). Silently overriding it makes it harder to
follow and easy to make a mistake.
> /*
> * Maximum number of compaction retries with a progress before OOM
> * killer is consider as the only way to move forward.
> @@ -4256,8 +4266,11 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
> bool ret = false;
> int retries = *compaction_retries;
> enum compact_priority priority = *compact_priority;
> + unsigned int compact_order;
>
> - if (!order)
> + /* Check the compaction result at the order compaction ran at */
> + compact_order = nofrag_promote_order(order, alloc_flags, ac);
> + if (!compact_order)
> return false;
>
> if (fatal_signal_pending(current))
> @@ -4299,7 +4316,7 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
> /*
> * Compaction failed. Retry with increasing priority.
> */
> - min_priority = (order > PAGE_ALLOC_COSTLY_ORDER) ?
> + min_priority = (compact_order > PAGE_ALLOC_COSTLY_ORDER) ?
> MIN_COMPACT_COSTLY_PRIORITY : MIN_COMPACT_PRIORITY;
This would change how hard we try to compact with defrag_mode in direct
compaction as it won't try compaction with MIN_COMPACT_PRIORITY anymore.
It doesn't make much sense to change that as part of this fix?
>
> if (*compact_priority > min_priority) {
--
Cheers,
Harry / Hyeonggon
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-09-29 18:39 ` Harry Yoo
@ 2026-09-30 13:32 ` Kiryl Shutsemau
2026-09-30 14:06 ` Johannes Weiner
0 siblings, 1 reply; 7+ messages in thread
From: Kiryl Shutsemau @ 2026-09-30 13:32 UTC (permalink / raw)
To: Harry Yoo, Vlastimil Babka, Johannes Weiner
Cc: Andrew Morton, David Hildenbrand, Suren Baghdasaryan,
Michal Hocko, Brendan Jackman, Zi Yan, Shakeel Butt, Usama Arif,
linux-mm, linux-kernel, stable, kernel-team
On Tue, Sep 29, 2026 at 07:39:41PM +0100, Harry Yoo wrote:
> On Tue, Sep 29, 2026 at 06:45:51PM +0100, Kiryl Shutsemau wrote:
> > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> >
> > Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> > storm in defrag_mode"), direct reclaim and compaction for non-movable
> > requests under defrag_mode run at pageblock_order, to produce the whole
> > blocks that ALLOC_NOFRAGMENT needs.
>
> > The retry decisions that follow still use the request order.
>
> Indeed, good catch!
>
> > An order-0 request can therefore retry
> > indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
> >
> > - Reclaim at pageblock_order gives up after one pass as soon as a zone
> > looks compaction_ready(), and do_try_to_free_pages() then returns 1
> > even though nothing was reclaimed. It returns before the retry that
> > would reclaim memory.low-protected cgroups, so when most memory is
> > protected, the pass that did run finds next to nothing.
> >
> > - Compaction at pageblock_order fails or is deferred.
> >
> > - should_reclaim_retry() takes the reported progress as progress for
> > the order-0 request and resets no_progress_loops. The request
> > retries.
>
> Makes sense to me.
>
> > Order 1-3 requests loop the same way, and should_compact_retry() also
> > checks their pageblock_order compaction result against the request
> > order.
> >
> > On a production host (64G, defrag_mode, memory.low covering most of the
> > workload), 95% of direct reclaim runs were order-9 runs that returned 1
> > with nothing reclaimed, at up to 60k runs per second. Across ~200M
> > should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> > The spinning allocations were SLUB slab refills for inode and dentry
> > caches. The time spent registers as memory pressure, and pressure-based
> > OOM killing takes down both workloads and system services.
> >
> > Treat promoted requests like costly orders:
> >
> > - Reclaim progress does not reset no_progress_loops for them.
> >
> > - should_compact_retry() checks the compaction result at the promoted
> > order. It does not retry COMPACT_SKIPPED, since the request can fall
> > back, and it does not escalate compaction to COMPACT_PRIO_SYNC_FULL.
> >
> > When the fallback is taken, reset the retry counters, so that the
> > fallback attempt gets a full retry budget before the OOM killer is
> > considered.
> >
> > In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
> >
> > before after
> > should_reclaim_retry() calls 63M 293k
> > peak memory pressure (PSI some avg10) 99% 12%
> >
> > File creation runs 5.7x faster.
> >
> > Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
> > Cc: <stable@vger.kernel.org>
> > Assisted-by: LLM
> > Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> > ---
> > mm/page_alloc.c | 85 ++++++++++++++++++++++++++++++++-----------------
> > 1 file changed, 56 insertions(+), 29 deletions(-)
> >
> > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > index 12fac9084c48..608487672d93 100644
> > --- a/mm/page_alloc.c
> > +++ b/mm/page_alloc.c
> > @@ -4127,6 +4127,31 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> > return page;
> > }
> >
> > +/*
> > + * If fallbacks are not permitted (defrag_mode), we either need to
> > + * reclaim space in a block of matching type, or clear out an entire
> > + * block to allow __rmqueue_claim() to convert.
> > + *
> > + * Reclaim by itself is primarily freeing space in movable blocks,
> > + * since that's where the LRU pages live. So this works for movable
> > + * requests, but not for others.
> > + *
> > + * For those, promote the order of reclaim and compaction to help make
> > + * blocks, instead of spinning in reclaim alone unproductively. Retry
> > + * decisions based on the outcome of that work - reclaim progress and
> > + * compaction results - must account for the promotion as well, see
> > + * should_reclaim_retry() and should_compact_retry().
> > + */
> > +static inline unsigned int nofrag_promote_order(unsigned int order,
> > + unsigned int alloc_flags,
> > + const struct alloc_context *ac)
> > +{
> > + if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
> > + return max(order, pageblock_order);
> > +
> > + return order;
> > +}
>
> I think we should start distinguishing order and compact/reclaim_order
> in __alloc_pages_slowpath(). Silently overriding it makes it harder to
> follow and easy to make a mistake.
Agreed, four callers recomputing the same thing is asking for a
mismatch.
I would rather not grow this patch, it has to go to stable.
I will look into a cleanup on top: __alloc_pages_slowpath() computes the
promoted order once per iteration and passes it to direct
reclaim/compaction and the two retry helpers next to the request order,
so the helpers stop knowing about defrag_mode.
> > @@ -4299,7 +4316,7 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
> > /*
> > * Compaction failed. Retry with increasing priority.
> > */
> > - min_priority = (order > PAGE_ALLOC_COSTLY_ORDER) ?
> > + min_priority = (compact_order > PAGE_ALLOC_COSTLY_ORDER) ?
> > MIN_COMPACT_COSTLY_PRIORITY : MIN_COMPACT_PRIORITY;
>
> This would change how hard we try to compact with defrag_mode in direct
> compaction as it won't try compaction with MIN_COMPACT_PRIORITY anymore.
>
> It doesn't make much sense to change that as part of this fix?
It is a choice between compacting harder and falling back, which
fragments a block.
It is a judgement call on what defrag_mode means.
It would also mean that order-0 allocation request promoted to pageblock
can trigger SYNC_FULL compaction. I cannot say I understand the
implications. Will give it a try with the reproducer.
Johannes, Vlastimil, any comments here?
--
Kiryl Shutsemau / Kirill A. Shutemov
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-09-30 13:32 ` Kiryl Shutsemau
@ 2026-09-30 14:06 ` Johannes Weiner
0 siblings, 0 replies; 7+ messages in thread
From: Johannes Weiner @ 2026-09-30 14:06 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Harry Yoo, Vlastimil Babka, Andrew Morton, David Hildenbrand,
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, linux-mm, linux-kernel, stable,
kernel-team
On Wed, Sep 30, 2026 at 02:32:12PM +0100, Kiryl Shutsemau wrote:
> On Tue, Sep 29, 2026 at 07:39:41PM +0100, Harry Yoo wrote:
> > On Tue, Sep 29, 2026 at 06:45:51PM +0100, Kiryl Shutsemau wrote:
> > > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> > >
> > > Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> > > storm in defrag_mode"), direct reclaim and compaction for non-movable
> > > requests under defrag_mode run at pageblock_order, to produce the whole
> > > blocks that ALLOC_NOFRAGMENT needs.
> >
> > > The retry decisions that follow still use the request order.
> >
> > Indeed, good catch!
> >
> > > An order-0 request can therefore retry
> > > indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
> > >
> > > - Reclaim at pageblock_order gives up after one pass as soon as a zone
> > > looks compaction_ready(), and do_try_to_free_pages() then returns 1
> > > even though nothing was reclaimed. It returns before the retry that
> > > would reclaim memory.low-protected cgroups, so when most memory is
> > > protected, the pass that did run finds next to nothing.
> > >
> > > - Compaction at pageblock_order fails or is deferred.
> > >
> > > - should_reclaim_retry() takes the reported progress as progress for
> > > the order-0 request and resets no_progress_loops. The request
> > > retries.
> >
> > Makes sense to me.
> >
> > > Order 1-3 requests loop the same way, and should_compact_retry() also
> > > checks their pageblock_order compaction result against the request
> > > order.
> > >
> > > On a production host (64G, defrag_mode, memory.low covering most of the
> > > workload), 95% of direct reclaim runs were order-9 runs that returned 1
> > > with nothing reclaimed, at up to 60k runs per second. Across ~200M
> > > should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> > > The spinning allocations were SLUB slab refills for inode and dentry
> > > caches. The time spent registers as memory pressure, and pressure-based
> > > OOM killing takes down both workloads and system services.
> > >
> > > Treat promoted requests like costly orders:
> > >
> > > - Reclaim progress does not reset no_progress_loops for them.
> > >
> > > - should_compact_retry() checks the compaction result at the promoted
> > > order. It does not retry COMPACT_SKIPPED, since the request can fall
> > > back, and it does not escalate compaction to COMPACT_PRIO_SYNC_FULL.
> > >
> > > When the fallback is taken, reset the retry counters, so that the
> > > fallback attempt gets a full retry budget before the OOM killer is
> > > considered.
> > >
> > > In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
> > >
> > > before after
> > > should_reclaim_retry() calls 63M 293k
> > > peak memory pressure (PSI some avg10) 99% 12%
> > >
> > > File creation runs 5.7x faster.
> > >
> > > Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
Thanks for fixing this. Ironically, the same workload blowing up again
that I tested the old fix again for 3 weeks :(
Back then, the problem was: allocating thread is spinning on a goal it
doesn't help accomplish.
Now the problem is: allocating thread spinning and helping, but goal
is not achievable.
> > > Cc: <stable@vger.kernel.org>
> > > Assisted-by: LLM
> > > Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> > > ---
> > > mm/page_alloc.c | 85 ++++++++++++++++++++++++++++++++-----------------
> > > 1 file changed, 56 insertions(+), 29 deletions(-)
> > >
> > > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > > index 12fac9084c48..608487672d93 100644
> > > --- a/mm/page_alloc.c
> > > +++ b/mm/page_alloc.c
> > > @@ -4127,6 +4127,31 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> > > return page;
> > > }
> > >
> > > +/*
> > > + * If fallbacks are not permitted (defrag_mode), we either need to
> > > + * reclaim space in a block of matching type, or clear out an entire
> > > + * block to allow __rmqueue_claim() to convert.
> > > + *
> > > + * Reclaim by itself is primarily freeing space in movable blocks,
> > > + * since that's where the LRU pages live. So this works for movable
> > > + * requests, but not for others.
> > > + *
> > > + * For those, promote the order of reclaim and compaction to help make
> > > + * blocks, instead of spinning in reclaim alone unproductively. Retry
> > > + * decisions based on the outcome of that work - reclaim progress and
> > > + * compaction results - must account for the promotion as well, see
> > > + * should_reclaim_retry() and should_compact_retry().
> > > + */
> > > +static inline unsigned int nofrag_promote_order(unsigned int order,
> > > + unsigned int alloc_flags,
> > > + const struct alloc_context *ac)
> > > +{
> > > + if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype != MIGRATE_MOVABLE)
> > > + return max(order, pageblock_order);
> > > +
> > > + return order;
> > > +}
> >
> > I think we should start distinguishing order and compact/reclaim_order
> > in __alloc_pages_slowpath(). Silently overriding it makes it harder to
> > follow and easy to make a mistake.
>
> Agreed, four callers recomputing the same thing is asking for a
> mismatch.
>
> I would rather not grow this patch, it has to go to stable.
>
> I will look into a cleanup on top: __alloc_pages_slowpath() computes the
> promoted order once per iteration and passes it to direct
> reclaim/compaction and the two retry helpers next to the request order,
> so the helpers stop knowing about defrag_mode.
+1
All they really need to know is the split into requested order vs
production order (reclaim_order, compaction_order).
> > > @@ -4299,7 +4316,7 @@ should_compact_retry(gfp_t gfp_mask, struct alloc_context *ac, int order,
> > > /*
> > > * Compaction failed. Retry with increasing priority.
> > > */
> > > - min_priority = (order > PAGE_ALLOC_COSTLY_ORDER) ?
> > > + min_priority = (compact_order > PAGE_ALLOC_COSTLY_ORDER) ?
> > > MIN_COMPACT_COSTLY_PRIORITY : MIN_COMPACT_PRIORITY;
> >
> > This would change how hard we try to compact with defrag_mode in direct
> > compaction as it won't try compaction with MIN_COMPACT_PRIORITY anymore.
> >
> > It doesn't make much sense to change that as part of this fix?
>
> It is a choice between compacting harder and falling back, which
> fragments a block.
>
> It is a judgement call on what defrag_mode means.
>
> It would also mean that order-0 allocation request promoted to pageblock
> can trigger SYNC_FULL compaction. I cannot say I understand the
> implications. Will give it a try with the reproducer.
>
> Johannes, Vlastimil, any comments here?
I would leave this one with requested order unless the reproducer
disagrees.
There is a risk of defrag_mode self defeating over time by raising the
bar for fallbacks but not high enough. Every fallback we let through
will make it harder down the line to compact towards that higher bar.
There is also a non-zero risk of connecting order-0 request contexts
to SYNC compaction which they haven't done before 7e8756d7ad22. But
you traced the problem to retrying, not sync compaction itself.
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-09-29 17:45 [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order Kiryl Shutsemau
2026-09-29 18:39 ` Harry Yoo
@ 2026-09-29 19:58 ` Andrew Morton
2026-09-30 12:35 ` Kiryl Shutsemau
1 sibling, 1 reply; 7+ messages in thread
From: Andrew Morton @ 2026-09-29 19:58 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Vlastimil Babka, Johannes Weiner, David Hildenbrand,
Kiryl Shutsemau (Meta),
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, Harry Yoo, linux-mm, linux-kernel,
stable, kernel-team
On Tue, 29 Sep 2026 18:45:51 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> storm in defrag_mode"), direct reclaim and compaction for non-movable
> requests under defrag_mode run at pageblock_order, to produce the whole
> blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
> still use the request order. An order-0 request can therefore retry
> indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
7e8756d7ad22 is new in 7.3-rcX, so no cc:stable needed.
> - Reclaim at pageblock_order gives up after one pass as soon as a zone
> looks compaction_ready(), and do_try_to_free_pages() then returns 1
> even though nothing was reclaimed. It returns before the retry that
> would reclaim memory.low-protected cgroups, so when most memory is
> protected, the pass that did run finds next to nothing.
>
> - Compaction at pageblock_order fails or is deferred.
>
> - should_reclaim_retry() takes the reported progress as progress for
> the order-0 request and resets no_progress_loops. The request
> retries.
>
> Order 1-3 requests loop the same way, and should_compact_retry() also
> checks their pageblock_order compaction result against the request
> order.
>
> On a production host (64G, defrag_mode, memory.low covering most of the
> workload), 95% of direct reclaim runs were order-9 runs that returned 1
> with nothing reclaimed, at up to 60k runs per second. Across ~200M
> should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> The spinning allocations were SLUB slab refills for inode and dentry
> caches. The time spent registers as memory pressure, and pressure-based
> OOM killing takes down both workloads and system services.
A production host running latest -rc?
> Treat promoted requests like costly orders:
>
> - Reclaim progress does not reset no_progress_loops for them.
>
> - should_compact_retry() checks the compaction result at the promoted
> order. It does not retry COMPACT_SKIPPED, since the request can fall
> back, and it does not escalate compaction to COMPACT_PRIO_SYNC_FULL.
>
> When the fallback is taken, reset the retry counters, so that the
> fallback attempt gets a full retry budget before the OOM killer is
> considered.
>
> In a VM reproducer (32G, defrag_mode, inode churn under memory.low):
>
> before after
> should_reclaim_retry() calls 63M 293k
> peak memory pressure (PSI some avg10) 99% 12%
>
> File creation runs 5.7x faster.
>
> Fixes: 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim storm in defrag_mode")
> Cc: <stable@vger.kernel.org>
Please double-check?
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-09-29 19:58 ` Andrew Morton
@ 2026-09-30 12:35 ` Kiryl Shutsemau
2026-09-30 20:02 ` Andrew Morton
0 siblings, 1 reply; 7+ messages in thread
From: Kiryl Shutsemau @ 2026-09-30 12:35 UTC (permalink / raw)
To: Andrew Morton
Cc: Vlastimil Babka, Johannes Weiner, David Hildenbrand,
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, Harry Yoo, linux-mm, linux-kernel,
stable, kernel-team
On Tue, Sep 29, 2026 at 12:58:16PM -0700, Andrew Morton wrote:
> On Tue, 29 Sep 2026 18:45:51 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
>
> > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> >
> > Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> > storm in defrag_mode"), direct reclaim and compaction for non-movable
> > requests under defrag_mode run at pageblock_order, to produce the whole
> > blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
> > still use the request order. An order-0 request can therefore retry
> > indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
>
> 7e8756d7ad22 is new in 7.3-rcX, so no cc:stable needed.
The commit itself has cc:stable. And it is in 6.18.50 and 7.2.4.
The fix needs to follow it there.
> > - Reclaim at pageblock_order gives up after one pass as soon as a zone
> > looks compaction_ready(), and do_try_to_free_pages() then returns 1
> > even though nothing was reclaimed. It returns before the retry that
> > would reclaim memory.low-protected cgroups, so when most memory is
> > protected, the pass that did run finds next to nothing.
> >
> > - Compaction at pageblock_order fails or is deferred.
> >
> > - should_reclaim_retry() takes the reported progress as progress for
> > the order-0 request and resets no_progress_loops. The request
> > retries.
> >
> > Order 1-3 requests loop the same way, and should_compact_retry() also
> > checks their pageblock_order compaction result against the request
> > order.
> >
> > On a production host (64G, defrag_mode, memory.low covering most of the
> > workload), 95% of direct reclaim runs were order-9 runs that returned 1
> > with nothing reclaimed, at up to 60k runs per second. Across ~200M
> > should_reclaim_retry() calls in a day, no_progress_loops never left 0.
> > The spinning allocations were SLUB slab refills for inode and dentry
> > caches. The time spent registers as memory pressure, and pressure-based
> > OOM killing takes down both workloads and system services.
>
> A production host running latest -rc?
We are running 7.1 with 7e8756d7ad22 backported. We found defrag mode
crucial if we want to use large folios in page cache.
But nobody besides us seems to test it :/
Can we make it default pretty please? :P
--
Kiryl Shutsemau / Kirill A. Shutemov
^ permalink raw reply [flat|nested] 7+ messages in thread* Re: [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order
2026-09-30 12:35 ` Kiryl Shutsemau
@ 2026-09-30 20:02 ` Andrew Morton
0 siblings, 0 replies; 7+ messages in thread
From: Andrew Morton @ 2026-09-30 20:02 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Vlastimil Babka, Johannes Weiner, David Hildenbrand,
Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Zi Yan,
Shakeel Butt, Usama Arif, Harry Yoo, linux-mm, linux-kernel,
stable, kernel-team
On Wed, 30 Sep 2026 13:35:01 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
> On Tue, Sep 29, 2026 at 12:58:16PM -0700, Andrew Morton wrote:
> > On Tue, 29 Sep 2026 18:45:51 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
> >
> > > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> > >
> > > Since commit 7e8756d7ad22 ("mm: page_alloc: fix non-movable reclaim
> > > storm in defrag_mode"), direct reclaim and compaction for non-movable
> > > requests under defrag_mode run at pageblock_order, to produce the whole
> > > blocks that ALLOC_NOFRAGMENT needs. The retry decisions that follow
> > > still use the request order. An order-0 request can therefore retry
> > > indefinitely without ever reaching the ALLOC_NOFRAGMENT fallback:
> >
> > 7e8756d7ad22 is new in 7.3-rcX, so no cc:stable needed.
>
> The commit itself has cc:stable. And it is in 6.18.50 and 7.2.4.
> The fix needs to follow it there.
doh.
> We are running 7.1 with 7e8756d7ad22 backported. We found defrag mode
> crucial if we want to use large folios in page cache.
>
> But nobody besides us seems to test it :/
>
> Can we make it default pretty please? :P
Send a patch ;)
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-09-30 20:03 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 17:45 [PATCH] mm: page_alloc: make defrag_mode retries follow the promoted order Kiryl Shutsemau
2026-09-29 18:39 ` Harry Yoo
2026-09-30 13:32 ` Kiryl Shutsemau
2026-09-30 14:06 ` Johannes Weiner
2026-09-29 19:58 ` Andrew Morton
2026-09-30 12:35 ` Kiryl Shutsemau
2026-09-30 20:02 ` Andrew Morton
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®