* [RFC PROTOTYPE] mm: reliable 1GB page allocation
@ 2026-10-04 5:36 Rik van Riel
0 siblings, 0 replies; only message in thread
From: Rik van Riel @ 2026-10-04 5:36 UTC (permalink / raw)
To: linux-kernel
Cc: linux-mm, Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Kairui Song, Qi Zheng,
Shakeel Butt, Johannes Weiner, Brendan Jackman, Zi Yan
The goal of this series is to keep gigantic (1GB) pages allocatable
after weeks of uptime, by concentrating non-movable pageblocks in a
small number of 1GiB-aligned ranges.
Design document: https://linux-mm.org/GigaBlocks
Code: https://github.com/rikvanriel/linux/tree/riel/gigablock-2026-10-03
A 1GB page needs a 1GB-aligned range in which every page is free or
movable; one unmovable page disqualifies the range. Over time a small
amount of kernel memory spreads across nearly every range. On a 765GB
node after 53 days of build and filesystem work, none of 754 1GB ranges was
free of non-movable pageblocks. On a 253GB test host the same class of
workload mixed 251 of 254 ranges in 9.2 hours, and a request for eight
1GB pages returned none.
This series uses the separate free lists for each migrate type
in zone->free_area[] in the fast path. The idea is to satisfy kernel
allocations from the unmovable and reclaimable free lists, while
movable allocations come from the movable free lists.
For each gigablock, the system tracks whether it contains non-movable
content, or only free and movable page blocks. A gigablock can contain
a mix of non-movable and movable data.
When the kernel free lists run low, the slow path moves pages and page
blocks from inside mixed gigablocks onto the kernel free lists.
Creating kernel free memory inside mixed superblocks is done
by claiming page blocks in a mixed superblock, when the kernel
free lists are exhausted, and a non-movable allocation has to
fall back, or by compaction evacuating movable content from
previously claimed kernel pageblocks that have movable content inside.
Claiming a new pageblock in a mixed superblock is done from the
allocator path, using a bounded scan, only once the kernel free
list is unable to satisfy an allocation. Evacuation is done by
kcompactd, and kicked off when the kernel free list falls below
a watermark. The goal is for the asynchronous path keep the
page allocator on the fast path.
Movable allocations can fall back to the kernel free lists.
That content can always be moved later.
Region size is fixed at 1GB rather than PUD size, geometry is frozen
at boot, and a hotplug span change disables tracking for the zone rather
than resizing metadata.
The diffstat shows that this version adds too much code to page_alloc.c,
which is already too large, and should probably be split up somehow.
One candidate in this series is to split the evacuation code and gigablock
slow path into its own file, moving that out of both page_alloc.c and
compaction.c.
Duplicating a small amount of compaction logic may be preferable to
complicating the existing logic at this point.
This code is not ready to be merged yet. The full design document has
a list of things that still need to be done.
I am posting this in the hopes of getting feedback on the basic design,
so that can be incorporated in the next rounds of tuning and cleanups.
Specifically feedback on design concepts like:
- Using only zone->free_area by migrate type in the fast path
- Moving pages around in the slow path, to keep the kernel free
lists stocked.
- Some amount of scanning in the page allocation code if the
kernel free lists cannot satisfy an allocation, in order
to claim a block in a mixed gigablock.
- Using compaction to evacuate movable content from pageblocks
claimed for kernel use.
- The policy of restricting how much kernel allocations can
claim new pageblocks anywhere.
- The policy of letting movable allocations fall back to the
kernel free lists once the movable free list is exhausted
Documentation/admin-guide/kernel-parameters.txt | 8
include/linux/mmzone.h | 83
include/linux/pageblock-flags.h | 2
include/linux/sched.h | 4
include/linux/vm_event_item.h | 48
mm/compaction.c | 170
mm/folio.c | 5
mm/internal.h | 147
mm/memory_hotplug.c | 3
mm/mm_init.c | 150
mm/page_alloc.c | 5553 +++++++++++++++++-------
mm/page_alloc.h | 6
mm/vmstat.c | 273 +
13 files changed, 5015 insertions(+), 1437 deletions(-)
^ permalink raw reply [flat|nested] only message in thread
only message in thread, other threads:[~2026-10-04 5:37 UTC | newest]
Thread overview: (only message) (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-04 5:36 [RFC PROTOTYPE] mm: reliable 1GB page allocation Rik van Riel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®