From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 853B4383C92 for ; Thu, 6 Aug 2026 23:29:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786058987; cv=none; b=oVBB1v/Xdpk2naso4WgJPETraYTU3pgjVA7SYLd13idZ2do8B02QMHnDzYbdumqQQOqlmWsGRCZdOodAIGd4D5v8IGvtMW76lVnabyXJLjfPiJjMETt1oOslo2nep3eplHbH1qz9RaNMS8lpZxgpeqbbiblTP83eAlvyn/7PgIM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786058987; c=relaxed/simple; bh=z4NnPWYaQGg9vU7aaS+Qw0HS6TWtw+PhvXUQdS8iugM=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=KJTaQ2FPLUi4xTwOkdi8Jp5UXTPI9lYPMiLU3V7e1+Eu9dW+pHyrnuwD8Jlp6FZuoYDGRZTaAAFuy2Yl0dqFI7z4UcIB5gFLdeS1Qi7sNI7i4QyFoKoKQboXlTzUc5HqMnzpzeDH3qrlDqMILCKxasPp3rwtFjLtOXlMNyab3r4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=PPN9I9yD; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="PPN9I9yD" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3CE571F000E9; Thu, 6 Aug 2026 23:29:45 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786058986; bh=sBgeHQZsrsk/VLEduVyme051pQyIT6NonmpdfuoL5k4=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=PPN9I9yDnuc5twCDP0Mpn4tjKqscTci3jF8EhLm3aIE8z24caOeTF3n8zrAIQbnzo B1ErSDuz7D381vVmvFuZ4FookgNVTB1XClSbQZVLCafTVThJWQKbPvD8PR4rZCuwS0 I5WcNm4zngbLyjxZNzVvta2lFJcBx6TxIZkfpdmWF6l7xHqU9Ib+KoR0b2kNZCNO/4 23Vl3xxb6tMChvHI107KVfm59WUWPuv9q8ZQwDInbuU2V0tJcKNgXOT0kvp9Tb6clW D8+BwectjK90Q1DeXmWLjgprTOvW1Cc3DcmXZ6lDCe3YGSJISO6hrbbupkdZkt3bk4 s1iD+2/7FhWNw== Date: Thu, 6 Aug 2026 23:29:43 +0000 From: Yosry Ahmed To: Brendan Jackman Cc: Borislav Petkov , Dave Hansen , Peter Zijlstra , Andrew Morton , David Hildenbrand , Vlastimil Babka , Mike Rapoport , Wei Xu , Johannes Weiner , Zi Yan , Lorenzo Stoakes , linux-mm@kvack.org, linux-kernel@vger.kernel.org, x86@kernel.org, Sumit Garg , Will Deacon , rientjes@google.com, patrick.roy@linux.dev, "Itazuri, Takahiro" , Andy Lutomirski , David Kaplan , Thomas Gleixner , Patrick Bellasi , Reiji Watanabe , Sean Christopherson Subject: Re: [PATCH v3 24/26] mm/page_alloc: always direct compact for unmapped allocs Message-ID: References: <20260726-page_alloc-unmapped-v3-0-6f5729aa9832@google.com> <20260726-page_alloc-unmapped-v3-24-6f5729aa9832@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260726-page_alloc-unmapped-v3-24-6f5729aa9832@google.com> On Sun, Jul 26, 2026 at 10:22:57PM +0000, Brendan Jackman wrote: > This is the minimal solution for ensuring that compaction can service > unmapped allocations. Without this, it's possible for compaction to just > check watermarks and see plenty of free pages, without being aware of > the direct map state, and thereby cause an ALLOC_UNMAPPED allocation to > fail unnecessarily. > > Instead, with this change, promote compact_order to pageblock order for > unmapped allocations, much like defrag_mode. Then, check specifically in > compaction for the presence of wholly mapped blocks that can be unmapped > once direct compact is complete. > > This all takes advantage of a major simplification: since unmapped > blocks are currently always unmovable, this can be asymmetric. There is > never a need to promote a !ALLOC_UNMAPPED allocation to compacting at > pageblock_order, because compaction would be trying to generate a > currently-unmapped block to map; that will always fail because it would > require migrating unmapped pages, which is not supported at the moment. > > Signed-off-by: Brendan Jackman > --- > mm/compaction.c | 22 ++++++++++++++++++---- > mm/page_alloc.c | 9 +++++++++ > 2 files changed, 27 insertions(+), 4 deletions(-) > > diff --git a/mm/compaction.c b/mm/compaction.c > index ed12d2fc6fad3..fe1aaf293bbce 100644 > --- a/mm/compaction.c > +++ b/mm/compaction.c > @@ -2531,12 +2531,25 @@ bool compaction_zonelist_suitable(struct alloc_context *ac, int order, > static enum compact_result > compaction_suit_allocation_order(struct zone *zone, unsigned int order, > int highest_zoneidx, unsigned int alloc_flags, > - bool async, bool kcompactd) > + bool unmapped, bool async, bool kcompactd) > { > unsigned long free_pages; > unsigned long watermark; > > - if (kcompactd && defrag_mode) > + /* > + * When trying to generate an unmapped block, check the counter for > + * direct-mapped blocks specifically, since we'll need to unmap the > + * whole block to service the allocation. > + * > + * Why doesn't this apply to the other way around too? (Mightn't we need > + * to _map_ a whole block, to service a !ALLOC_UNMAPPED allocation?) No, > + * because of a likely-temporary simplification: currently, unmapped > + * blocks never contain movable pages, so compaction isn't going to free > + * up one of those. > + */ > + if (unmapped) > + free_pages = zone_page_state(zone, NR_FREE_PAGES_BLOCKS_MAPPED); > + else if (kcompactd && defrag_mode) > free_pages = zone_free_pages_blocks(zone); > else > free_pages = zone_page_state(zone, NR_FREE_PAGES); (Sorry in advance for the wall of text, while trying to understand why we need this I ended up spending a lot of time staring at the compaction code and coming up with even more questions) Hmm why do we need to do this here? free pages is used in this check below: watermark = wmark_pages(zone, alloc_flags & ALLOC_WMARK_MASK); if (__zone_watermark_ok(zone, order, watermark, highest_zoneidx, alloc_flags, free_pages)) return COMPACT_SUCCESS; IIUC, this basically checks if we can skip compaction because we already have enough free pages, both in terms of zone watermarks as well as the availability of pages in the right order. free_pages is used for the watermark checks, so using NR_FREE_PAGES_BLOCKS_MAPPED will result in __zone_watermark_ok() returning false if all free memory is above the watermark, but free mapped memory specifically is above the watermark. This kinda makes sense, but: 1. Shouldn't we be checking all free page blocks (i.e. zone_free_pages_blocks() like defrag_mode), as entirely free unmapped blocks can also serve the allocation? 2. More importantly, I think for this check the more important part is checking if we have pages from the correct order (i.e. the loop at the end of __zone_watermark_ok()). Since we promote compaction requests for unmapped allocations to pageblock_order, this will essentially check if we have any free pageblocks, which is ultimately what we want. Checking the watermark against mapped pages only here seems arbitrary tbh, since all other callers of __zone_watermark_ok() are oblivious to mapped vs. unmapped, which seems to be a bigger issue. For example, compaction_suitable(), which IIUC actually checks if we have enough free memory as scratch space for compaction will check all free memory against the watermark, even though it cannot use unmapped memory. Same probably applies for other callers in allocation/reclaim/compaction paths. Since the watermark checks are ignorant of unmapped memory, we can end up serving allocations when the actual free mapped memory we have is below watermarks, dipping into reserves and causing OOM kills if we run out of mapped memory. Other than the watermark checks, the loop at the end of __zone_watermark_ok() checking for available free pages of the required order is also oblivious to unmapped memory: for (o = order; o < NR_PAGE_ORDERS; o++) { struct free_area *area = &z->free_area[o]; if (!area->nr_free) continue; for (int ft_idx = 0; ft_idx < NR_FREETYPE_IDXS; ft_idx++) { freetype_t ft = freetype_from_idx(ft_idx); int mt = free_to_migratetype(ft); if (list_empty(&area->free_list[ft_idx])) continue; if (mt < MIGRATE_PCPTYPES) return true; ... } } AFAICT free_to_migratetype() will return MIRATE_UNMOVABLE for unmapped pages, so __zone_watermark_ok() will return true for mapped allocations if there's a free unmapped page of the same order, even though it cannot actually be used (unless order >= pageblock_order). So it seems like __zone_watermark_ok() needs to be reworked to account for unmapped pages, and potentially other watermark checks. I wonder if this problem can be side-stepped if we allow sharing pageblocks between mapped and unmapped pages as a fallback (as I mentioned in another comment). With this, the checks in __zone_watermark_ok() can probably stay as-is since all free unmapped memory can be used for all allocations. Of course, this is not ideal and would cause performance regressions so we need this to be the last option, probably by generating free pageblocks more aggressively (like defrag mode). Perhaps we can somehow detect the presence of unmapped allocations and treat it like defrag mode? > @@ -2599,6 +2612,7 @@ compact_zone(struct compact_control *cc, struct capture_control *capc) > ret = compaction_suit_allocation_order(cc->zone, cc->order, > cc->highest_zoneidx, > cc->alloc_flags, > + freetype_unmapped(cc->freetype), > cc->mode == MIGRATE_ASYNC, > !cc->direct_compaction); > if (ret != COMPACT_CONTINUE) > @@ -3084,7 +3098,7 @@ static bool kcompactd_node_suitable(pg_data_t *pgdat) > ret = compaction_suit_allocation_order(zone, > pgdat->kcompactd_max_order, > highest_zoneidx, alloc_flags, > - false, true); > + false, false, true); > if (ret == COMPACT_CONTINUE) > return true; > } > @@ -3127,7 +3141,7 @@ static void kcompactd_do_work(pg_data_t *pgdat) > > ret = compaction_suit_allocation_order(zone, > cc.order, zoneid, cc.alloc_flags, > - false, true); > + false, false, true); > if (ret != COMPACT_CONTINUE) > continue; > > diff --git a/mm/page_alloc.c b/mm/page_alloc.c > index d12ce84662ab7..5f1dea7eee15b 100644 > --- a/mm/page_alloc.c > +++ b/mm/page_alloc.c > @@ -827,6 +827,9 @@ compaction_capture(struct capture_control *capc, struct page *page, > capc_mt != MIGRATE_MOVABLE) > return false; > > + if (freetype_flags(freetype) != freetype_flags(capc->freetype)) > + return false; > + > if (migratetype != capc_mt) > trace_mm_page_alloc_extfrag(page, capc->order, order, > capc_mt, migratetype); > @@ -4523,6 +4526,12 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned int order, > if ((alloc_flags & ALLOC_NOFRAGMENT) && > free_to_migratetype(ac->freetype) != MIGRATE_MOVABLE) > compact_order = max(order, pageblock_order); > + /* > + * Unmapped allocations benefit from compaction even at order 0, because the > + * allocator will actually grab a whole block. > + */ > + if (freetype_flags(ac->freetype) & FREETYPE_UNMAPPED) > + compact_order = max(order, pageblock_order); > > if (!compact_order) > return NULL; > > -- > 2.54.0 >