From: Yosry Ahmed <yosry@kernel.org>
To: Brendan Jackman <jackmanb@google.com>
Cc: Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
Peter Zijlstra <peterz@infradead.org>,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>, Wei Xu <weixugc@google.com>,
Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
Lorenzo Stoakes <ljs@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
x86@kernel.org, Sumit Garg <sumit.garg@oss.qualcomm.com>,
Will Deacon <will@kernel.org>,
rientjes@google.com, patrick.roy@linux.dev, "Itazuri,
Takahiro" <itazur@amazon.co.uk>,
Andy Lutomirski <luto@kernel.org>,
David Kaplan <david.kaplan@amd.com>,
Thomas Gleixner <tglx@kernel.org>,
Patrick Bellasi <derkling@google.com>,
Reiji Watanabe <reijiw@google.com>,
Sean Christopherson <seanjc@google.com>
Subject: Re: [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs
Date: Thu, 6 Aug 2026 23:29:43 +0000 [thread overview]
Message-ID: <anUUA66kHWAnepLt@google.com> (raw)
In-Reply-To: <20260726-page_alloc-unmapped-v3-24-6f5729aa9832@google.com>
On Sun, Jul 26, 2026 at 10:22:57PM +0000, Brendan Jackman wrote:
> This is the minimal solution for ensuring that compaction can service
> unmapped allocations. Without this, it's possible for compaction to just
> check watermarks and see plenty of free pages, without being aware of
> the direct map state, and thereby cause an ALLOC_UNMAPPED allocation to
> fail unnecessarily.
>
> Instead, with this change, promote compact_order to pageblock order for
> unmapped allocations, much like defrag_mode. Then, check specifically in
> compaction for the presence of wholly mapped blocks that can be unmapped
> once direct compact is complete.
>
> This all takes advantage of a major simplification: since unmapped
> blocks are currently always unmovable, this can be asymmetric. There is
> never a need to promote a !ALLOC_UNMAPPED allocation to compacting at
> pageblock_order, because compaction would be trying to generate a
> currently-unmapped block to map; that will always fail because it would
> require migrating unmapped pages, which is not supported at the moment.
>
> Signed-off-by: Brendan Jackman <jackmanb@google.com>
> ---
> mm/compaction.c | 22 ++++++++++++++++++----
> mm/page_alloc.c | 9 +++++++++
> 2 files changed, 27 insertions(+), 4 deletions(-)
>
> diff --git a/mm/compaction.c b/mm/compaction.c
> index ed12d2fc6fad3..fe1aaf293bbce 100644
> --- a/mm/compaction.c
> +++ b/mm/compaction.c
> @@ -2531,12 +2531,25 @@ bool compaction_zonelist_suitable(struct alloc_context *ac, int order,
> static enum compact_result
> compaction_suit_allocation_order(struct zone *zone, unsigned int order,
> int highest_zoneidx, unsigned int alloc_flags,
> - bool async, bool kcompactd)
> + bool unmapped, bool async, bool kcompactd)
> {
> unsigned long free_pages;
> unsigned long watermark;
>
> - if (kcompactd && defrag_mode)
> + /*
> + * When trying to generate an unmapped block, check the counter for
> + * direct-mapped blocks specifically, since we'll need to unmap the
> + * whole block to service the allocation.
> + *
> + * Why doesn't this apply to the other way around too? (Mightn't we need
> + * to _map_ a whole block, to service a !ALLOC_UNMAPPED allocation?) No,
> + * because of a likely-temporary simplification: currently, unmapped
> + * blocks never contain movable pages, so compaction isn't going to free
> + * up one of those.
> + */
> + if (unmapped)
> + free_pages = zone_page_state(zone, NR_FREE_PAGES_BLOCKS_MAPPED);
> + else if (kcompactd && defrag_mode)
> free_pages = zone_free_pages_blocks(zone);
> else
> free_pages = zone_page_state(zone, NR_FREE_PAGES);
(Sorry in advance for the wall of text, while trying to understand why
we need this I ended up spending a lot of time staring at the compaction
code and coming up with even more questions)
Hmm why do we need to do this here?
free pages is used in this check below:
watermark = wmark_pages(zone, alloc_flags & ALLOC_WMARK_MASK);
if (__zone_watermark_ok(zone, order, watermark, highest_zoneidx,
alloc_flags, free_pages))
return COMPACT_SUCCESS;
IIUC, this basically checks if we can skip compaction because we already
have enough free pages, both in terms of zone watermarks as well as the
availability of pages in the right order.
free_pages is used for the watermark checks, so using
NR_FREE_PAGES_BLOCKS_MAPPED will result in __zone_watermark_ok()
returning false if all free memory is above the watermark, but free
mapped memory specifically is above the watermark.
This kinda makes sense, but:
1. Shouldn't we be checking all free page blocks (i.e.
zone_free_pages_blocks() like defrag_mode), as entirely free unmapped
blocks can also serve the allocation?
2. More importantly, I think for this check the more important part is
checking if we have pages from the correct order (i.e. the loop at the
end of __zone_watermark_ok()). Since we promote compaction requests for
unmapped allocations to pageblock_order, this will essentially check if
we have any free pageblocks, which is ultimately what we want.
Checking the watermark against mapped pages only here seems arbitrary
tbh, since all other callers of __zone_watermark_ok() are oblivious to
mapped vs. unmapped, which seems to be a bigger issue.
For example, compaction_suitable(), which IIUC actually checks if we
have enough free memory as scratch space for compaction will check all
free memory against the watermark, even though it cannot use unmapped
memory. Same probably applies for other callers in
allocation/reclaim/compaction paths.
Since the watermark checks are ignorant of unmapped memory, we can end
up serving allocations when the actual free mapped memory we have is
below watermarks, dipping into reserves and causing OOM kills if we run
out of mapped memory.
Other than the watermark checks, the loop at the end of
__zone_watermark_ok() checking for available free pages of the required
order is also oblivious to unmapped memory:
for (o = order; o < NR_PAGE_ORDERS; o++) {
struct free_area *area = &z->free_area[o];
if (!area->nr_free)
continue;
for (int ft_idx = 0; ft_idx < NR_FREETYPE_IDXS; ft_idx++) {
freetype_t ft = freetype_from_idx(ft_idx);
int mt = free_to_migratetype(ft);
if (list_empty(&area->free_list[ft_idx]))
continue;
if (mt < MIGRATE_PCPTYPES)
return true;
...
}
}
AFAICT free_to_migratetype() will return MIRATE_UNMOVABLE for unmapped
pages, so __zone_watermark_ok() will return true for mapped allocations
if there's a free unmapped page of the same order, even though it cannot
actually be used (unless order >= pageblock_order).
So it seems like __zone_watermark_ok() needs to be reworked to account
for unmapped pages, and potentially other watermark checks.
I wonder if this problem can be side-stepped if we allow sharing
pageblocks between mapped and unmapped pages as a fallback (as I
mentioned in another comment). With this, the checks in
__zone_watermark_ok() can probably stay as-is since all free unmapped
memory can be used for all allocations.
Of course, this is not ideal and would cause performance regressions so
we need this to be the last option, probably by generating free
pageblocks more aggressively (like defrag mode). Perhaps we can somehow
detect the presence of unmapped allocations and treat it like defrag
mode?
> @@ -2599,6 +2612,7 @@ compact_zone(struct compact_control *cc, struct capture_control *capc)
> ret = compaction_suit_allocation_order(cc->zone, cc->order,
> cc->highest_zoneidx,
> cc->alloc_flags,
> + freetype_unmapped(cc->freetype),
> cc->mode == MIGRATE_ASYNC,
> !cc->direct_compaction);
> if (ret != COMPACT_CONTINUE)
> @@ -3084,7 +3098,7 @@ static bool kcompactd_node_suitable(pg_data_t *pgdat)
> ret = compaction_suit_allocation_order(zone,
> pgdat->kcompactd_max_order,
> highest_zoneidx, alloc_flags,
> - false, true);
> + false, false, true);
> if (ret == COMPACT_CONTINUE)
> return true;
> }
> @@ -3127,7 +3141,7 @@ static void kcompactd_do_work(pg_data_t *pgdat)
>
> ret = compaction_suit_allocation_order(zone,
> cc.order, zoneid, cc.alloc_flags,
> - false, true);
> + false, false, true);
> if (ret != COMPACT_CONTINUE)
> continue;
>
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index d12ce84662ab7..5f1dea7eee15b 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -827,6 +827,9 @@ compaction_capture(struct capture_control *capc, struct page *page,
> capc_mt != MIGRATE_MOVABLE)
> return false;
>
> + if (freetype_flags(freetype) != freetype_flags(capc->freetype))
> + return false;
> +
> if (migratetype != capc_mt)
> trace_mm_page_alloc_extfrag(page, capc->order, order,
> capc_mt, migratetype);
> @@ -4523,6 +4526,12 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order,
> if ((alloc_flags & ALLOC_NOFRAGMENT) &&
> free_to_migratetype(ac->freetype) != MIGRATE_MOVABLE)
> compact_order = max(order, pageblock_order);
> + /*
> + * Unmapped allocations benefit from compaction even at order 0, because the
> + * allocator will actually grab a whole block.
> + */
> + if (freetype_flags(ac->freetype) & FREETYPE_UNMAPPED)
> + compact_order = max(order, pageblock_order);
>
> if (!compact_order)
> return NULL;
>
> --
> 2.54.0
>
next prev parent reply other threads:[~2026-08-06 23:29 UTC|newest]
Thread overview: 69+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-26 22:22 [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 01/26] set_memory: add folio_{zap,restore}_direct_map helpers Brendan Jackman
2026-07-27 10:33 ` Mike Rapoport
2026-07-29 11:42 ` Brendan Jackman
2026-07-30 20:34 ` Yosry Ahmed
2026-07-31 5:21 ` Mike Rapoport
2026-07-31 11:57 ` Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 02/26] mm/secretmem: make use of folio_{zap,restore}_direct_map Brendan Jackman
2026-07-27 10:40 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 03/26] mm: introduce AS_NO_DIRECT_MAP Brendan Jackman
2026-07-30 21:06 ` Yosry Ahmed
2026-07-31 12:15 ` Brendan Jackman
2026-07-31 19:28 ` Yosry Ahmed
2026-08-07 0:02 ` Sean Christopherson
2026-08-07 0:13 ` Yosry Ahmed
2026-08-07 0:19 ` Sean Christopherson
2026-08-07 0:29 ` Yosry Ahmed
2026-08-02 16:10 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 04/26] x86/mm: split out preallocate_sub_pgd() Brendan Jackman
2026-07-31 22:10 ` Yosry Ahmed
2026-08-02 16:13 ` Mike Rapoport
2026-07-26 22:22 ` [PATCH v3 05/26] x86: move PAE PMD preallocation defines to header Brendan Jackman
2026-07-31 23:59 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 06/26] x86/tlb: Expose some flush function declarations to modules Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 07/26] x86/mm: introduce mm-local region Brendan Jackman
2026-08-02 16:27 ` Mike Rapoport
2026-08-03 22:29 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 08/26] x86/mm: move LDT remap into " Brendan Jackman
2026-08-03 22:33 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 09/26] mm: Create flags arg for __apply_to_page_range() Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 10/26] mm: Add more flags " Brendan Jackman
2026-08-04 0:08 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 11/26] x86/mm: introduce the mermap Brendan Jackman
2026-08-02 16:40 ` Mike Rapoport
2026-08-04 18:38 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 12/26] mm: KUnit tests for " Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 13/26] mm: introduce freetype_t Brendan Jackman
2026-08-04 22:23 ` Yosry Ahmed
2026-08-04 23:02 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 14/26] mm: move migratetype definitions to freetype.h Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 15/26] mm/page_alloc: add support for freetypes with no freelist Brendan Jackman
2026-07-31 14:13 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 16/26] mm: add definitions for allocating unmapped pages Brendan Jackman
2026-08-04 19:53 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 17/26] mm: encode freetype flags in pageblock flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 18/26] mm/page_alloc: separate pcplists by freetype flags Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 19/26] mm/page_alloc: rename ALLOC_NON_BLOCK back to _HARDER Brendan Jackman
2026-07-31 14:52 ` Vlastimil Babka (SUSE)
2026-08-03 9:20 ` Vlastimil Babka (SUSE)
2026-08-04 21:50 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 20/26] mm/page_alloc: introduce ALLOC_NOBLOCK Brendan Jackman
2026-07-26 22:22 ` [PATCH v3 21/26] mm/page_alloc: implement FREETYPE_UNMAPPED allocations Brendan Jackman
2026-08-03 9:18 ` Vlastimil Babka (SUSE)
2026-08-04 23:41 ` Yosry Ahmed
2026-08-07 0:05 ` Yosry Ahmed
2026-08-04 23:53 ` Yosry Ahmed
2026-08-05 16:13 ` Yosry Ahmed
2026-08-07 0:16 ` Yosry Ahmed
2026-07-26 22:22 ` [PATCH v3 22/26] mm: Minimal KUnit tests for some new page_alloc logic Brendan Jackman
2026-08-03 9:30 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 23/26] mm: Split out NR_FREE_PAGES_BLOCKS_[UN]MAPPED Brendan Jackman
2026-08-03 9:32 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs Brendan Jackman
2026-08-03 9:44 ` Vlastimil Babka (SUSE)
2026-08-06 23:29 ` Yosry Ahmed [this message]
2026-07-26 22:22 ` [PATCH v3 25/26] mm: plumb alloc flags into some alloc funcs Brendan Jackman
2026-08-03 9:52 ` Vlastimil Babka (SUSE)
2026-07-26 22:22 ` [PATCH v3 26/26] mm: add fast path for AS_NO_DIRECT_MAP Brendan Jackman
2026-07-29 11:52 ` [PATCH v3 00/26] mm: Add ALLOC_UNMAPPED and AS_NO_DIRECT_MAP Brendan Jackman
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anUUA66kHWAnepLt@google.com \
--to=yosry@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=bp@alien8.de \
--cc=dave.hansen@linux.intel.com \
--cc=david.kaplan@amd.com \
--cc=david@kernel.org \
--cc=derkling@google.com \
--cc=hannes@cmpxchg.org \
--cc=itazur@amazon.co.uk \
--cc=jackmanb@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=luto@kernel.org \
--cc=patrick.roy@linux.dev \
--cc=peterz@infradead.org \
--cc=reijiw@google.com \
--cc=rientjes@google.com \
--cc=rppt@kernel.org \
--cc=seanjc@google.com \
--cc=sumit.garg@oss.qualcomm.com \
--cc=tglx@kernel.org \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=will@kernel.org \
--cc=x86@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome