* [RFC PATCH 0/6] mm: swap support for hugetlb folios
@ 2026-09-29 7:53 Zongkun Lei
2026-09-29 8:02 ` David Hildenbrand (Arm)
2026-09-29 8:18 ` Lorenzo Stoakes (ARM)
0 siblings, 2 replies; 5+ messages in thread
From: Zongkun Lei @ 2026-09-29 7:53 UTC (permalink / raw)
To: linux-mm
Cc: Andrew Morton, David Hildenbrand, Muchun Song, Oscar Salvador,
Peter Xu, Matthew Wilcox, Kairui Song, Michal Hocko,
Johannes Weiner, Hugh Dickins, Naoya Horiguchi, Zi Yan,
Lorenzo Stoakes, linux-kernel, Zongkun Lei
Hi all,
This series introduces *user-driven* swap support for hugetlb folios,
covering both private anonymous and shared file-backed mappings. It
deliberately keeps the hugetlb reservation semantics intact: the
kernel never reclaims hugetlb pages on its own -- hugetlb folios never
go on an LRU, there is no hot/cold guessing, and nothing hooks into
kswapd or direct reclaim. This is the same boundary the merged
hugetlb demotion keeps: the kernel never reclaims reserved memory by
its own judgement. Demotion is triggered explicitly by the
administrator (writing nr_hugepages) and only touches free pool
pages; this series is triggered explicitly by the workload (madvise)
and acts on its own live mappings. In both cases the decision stays
outside the kernel.
Interface:
- Swap-out: MADV_PAGEOUT via madvise(2)/process_madvise(2) learns to
handle hugetlb mappings. For anonymous mappings the PTE becomes a
swap entry; for shared mappings the PTE is simply cleared and the
location record stays in the hugetlbfs page cache.
- Swap-in: a fault reads the page back from swap.
- Swapoff: all swapped-out hugetlb pages are read back. If the pool
has no free huge pages at swap-in time, the behaviour matches a
first-touch fault on a full pool: after bounded retries the process
receives SIGBUS -- no silent failure, no data loss.
Motivation
==========
Today a hugetlb reservation is a one-way street: once instantiated, a
huge page's memory is completely unreclaimable until the mapping is
torn down. THP and 4K anonymous memory can simply be swapped out;
hugetlb cannot. This gap forces operators to choose between "huge
page performance" and "memory overcommit".
VM overcommit. QEMU guests backed by MAP_SHARED hugetlbfs files
(-mem-path) allocate all guest memory from the host's huge page pool.
Overcommit means selling more guest memory than the pool holds --
fine in practice because guests never all peak at once. But the
extreme case is unavoidable: when every guest's load spikes at the
same time, the pool runs out. Keeping the VMs alive then requires a
swap mechanism that moves the memory the provider judges cold out of
the way. Ballooning cannot save this scenario: it reclaims memory
that is free *inside* the guest, and the extreme case is precisely
that all guests are busy and have nothing to give. The extreme case
is rare, but without an escape hatch providers dare not overcommit
confidently, and pool utilization stays low. This series is that
escape hatch: the VM management stack issues MADV_PAGEOUT on cold
guest memory, VM availability survives the extreme case, and "huge
page performance" vs "memory overcommit" stops being a choice.
Hot-upgrade standby. A hot upgrade first brings up the new instance
and switches traffic before the old one may be touched; to keep
rollback fast, the old instance usually cannot be deleted right away
and stands by for hours or days. During that time it has no traffic
and its memory is cold, yet the process and its mappings still pin
the pool until it is finally removed. With MADV_PAGEOUT the cold
memory can be swapped out during standby and swapped back if a
rollback happens. More generally: whenever the pool is full but
tearing down the mapping is not an option, MADV_PAGEOUT is the only
way to free space.
Design
======
Preparation: tree-wide folio_swap_entry() conversion (patch 1)
--------------------------------------------------------------
Hugetlb flags live in folio->private (a union with folio->swap), so
while a folio sits in the swap cache the low bits of folio->swap.val
can be polluted by HPAGEFLAG (with HVO enabled, bit 4 is set on
virtually every hugetlb folio -- not a theoretical issue). The fix
is folio_swap_entry() -- an identity transform for non-hugetlb folios
(round_down(x, 1) == x for order-0; THP low bits are zero anyway) --
plus a *tree-wide* conversion of all folio->swap readers. What
review must prove shrinks from "N sites x reachability arguments" to
"1 helper + 1 invariant (swap entries are folio-size aligned)", with
zero semantic change to existing paths. It makes *generic* code
hugetlb-safe instead of adding special cases in hugetlb.c, aligned
with the hugetlb genericization direction.
The companion size gate hugetlb_folio_swap_supported() (introduced in
patch 3) clamps both bounds at once: the lower bound (HPAGEFLAG
width, requiring nr_pages >= 128) and the upper bound (non-gigantic,
<= PMD size, so one swap cluster holds the folio).
Private anonymous mappings (patches 2-4)
----------------------------------------
- Swap-in (consumer first, patch 2): hugetlb_anon_swapin() mirrors
do_swap_page(): get_swap_device() pins the device,
swap_cache_get_folio() finds in-flight folios, and on a miss the
folio is allocated from the hugetlb pool and added to the swap
cache via the new generic swap_cache_add_folio() helper (hugetlb
folios do not come from the buddy allocator, so
__swap_cache_alloc() cannot be used). Pool exhaustion retries
briefly, then fails with SIGBUS, matching the
reservation-exhaustion contract of the no-page fault path. The
lifecycle is fully 4K-semantic: fork shares swap entries
(copy_hugetlb_page_range() + swap_dup_entries_direct(), dropping
the exclusive bit); zap/swapin release them; MM_SWAPENTS is
balanced at swapout/fork/swapin/zap per mainline semantics (VmSwap
stays meaningful); swapoff drains via hugetlb_unuse_vma().
The read-back path must merge before the swap-out path: the base
hugetlb_fault() returns ret = 0 for any entry that is neither
migration nor hwpoison, so with swap-out first, touching a
swapped-out page would livelock in an endless fault loop.
- Swap-out (patch 3): hugetlb_reclaim_pages() /
hugetlb_reclaim_folio_list() are trimmed mirrors of
reclaim_pages()/shrink_folio_list() and live in hugetlb.c; the
changelog of patch 3 explains why a mirror and why there. Swap PTE
installation reuses the rmap walker: a new
try_to_unmap_swap_hugetlb_one() (a static callback mirroring
ttu_anon_swapbacked_folio(): single walk, swp_pte_prepare(),
mm_prepare_for_swap_entries(), MM_SWAPENTS accounting) plus the
exported wrapper try_to_unmap_swap_hugetlb(), following the
try_to_unmap_poisoned_hugetlb_one() precedent. Cost to generic
rmap: one added line in rmap.h, zero changes to existing code.
- memory-failure (patch 4): anonymous hugetlb folios can now sit in
the swap cache (the writeback window, the cached copy left after
swap-in) -- a state memory failure has never seen before. Two
adaptations: try_to_unmap() routes hugetlb folios by TTU_HWPOISON
(set -> the poisoned handler; cleared -- the mf case that keeps a
dirty swapcache folio -- -> the swap handler, matching 4K swapcache
semantics), and me_huge_page() intercepts swapcache folios at the
top (mirroring me_swapcache_dirty(): keep the folio in the swap
cache and return MF_DELAYED), fixing the silent eviction of a clean
poisoned folio that would otherwise mean a 2M pool leak plus a lost
poison marker.
Shared file-backed mappings (patch 5)
-------------------------------------
The shmem model: swap-out only clears the PTEs -- no swap entry is
ever installed for a file folio -- and the anchor (a swap value
entry, swp_to_radix_entry()) is left in the hugetlbfs page cache,
holding a dup'ed swap count.
- Swap-out: hugetlbfs_writeout() -- folio_alloc_swap() allocates the
folio-sized slot range, the inode is linked on the per-inode
swaplist (for swapoff traversal), folio_dup_swap() holds a swap
count for the anchor, the page cache slot is replaced by the anchor
in place, and swap_writeout() writes out. The reclaim loop drops
its anon-only gate and madvise drops the VM_MAYSHARE rejection.
- Refault swap-in: hugetlb_no_page() recognizes the anchor (the page
cache lookup reports it as a hole) and swaps the folio back in via
hugetlbfs_do_swapin(), replacing the anchor in place --
indistinguishable from a regular page-cache hit afterwards. The
read-back path must never lag the producer: the base fault path
treats value entries as holes and would silently zero-fill over
swapped-out data.
- read(2) swap-in: hugetlbfs_read_iter() resolves anchors through
hugetlbfs_swapin_read() instead of zero-filling; swapin failures
surface as -EIO/-EFAULT, never stale data.
- Lifecycle: truncate/hole-punch/evict free anchors and their slots
via hugetlbfs_free_swap(), with two UAF guards (the swap cache
entry is removed before the anchor's swap count is dropped; a
refcount race means reclaim is isolating the folio, so the folio
lock is held across swap_put_entries_direct() to wait it out).
Inode eviction drains the swapoff traversal list via
hugetlbfs_evict_drain_swaplist().
- swapoff: try_to_unuse() hooks hugetlbfs_unuse() right after
shmem_unuse(). File folios install no swap PTEs, so the mm walk
cannot reach their slots -- the per-inode swaplist traversal can.
- Symmetric accounting: hugetlb_cgroup, memcg swap/hugetlb charge,
the global reserve and NR_HUGETLB are paired with free_huge_folio()
on every path; the vma-less read/swapoff swap-in charges the
caller's context, the same policy as shmem swapoff.
Base and dependencies
=====================
Based on mm-unstable; no external patch dependencies.
The implementation builds on Kairui Song's swap table work: the
folio-sized contiguous swap slot ranges (folio_alloc_swap() /
folio_dup_swap() / folio_put_swap()) are the primitives that make
this possible at all -- without them a 2M folio would need 512
independent swap entries. The merged phases of that work are already
in mm-unstable [1].
Looking ahead, the swap table roadmap may move swap entries out of
folio->swap; removing folio_swap_entry() then is a mechanical
operation -- the call sites are already concentrated in a single
helper, so it will be easier to delete than what we have today.
This series is complementary to the Reserved THP RFC (Qi Zheng), not
competing with it. Reserved THP explores the endgame -- reserved,
swappable, unsplittable large folios that may absorb most hugetlb use
cases. But the LSFMM 2024 consensus is that the hugetlbfs ABI stays
(migrating away is a 15-20 year scale); as long as the ABI exists,
its unreclaimable-pool problem needs a direct solution. The two also
build on the same swap table foundation, so progress on either side
de-risks the other.
Anticipated questions
=====================
Q1: Why not put hugetlb folios on the LRU and reuse vmscan?
Semantic layer: hugetlb's ABI contract is "reservation = latency
determinism", and the kernel must not tear that up unilaterally --
this is the historical reason hugetlb could not be swapped at all.
Letting the kernel reclaim reserved pages under global memory
pressure based on its own hot/cold guesses (folio_referenced() style
access-pattern inference) injects fault latency exactly where it must
never appear -- the hot-upgrade standby and pre-traffic-peak moments
are precisely why hugetlb is used. So the split in this series is:
the kernel provides the mechanism, the policy belongs to userspace.
MADV_PAGEOUT extends to hugetlb the proactive contract 4K anonymous
memory already has; the workload (orchestrator, JVM, database) knows
its warmup windows and traffic peaks, so hot/cold decisions stay on
the side with the most information. This is the same boundary the
only merged hugetlb reclaim mechanism, demotion, keeps: the kernel
never reclaims reserved memory by its own judgement -- demotion is
administrator-triggered (nr_hugepages) and touches only free pool
pages, this series is workload-triggered (madvise) and touches its
own live mappings; in both cases the decision is outside the kernel.
Technical layer: an LRU is not a neutral queue; it is the input
structure of the kernel's reclaim policy. Once a hugetlb folio is on
an lruvec, kswapd, direct reclaim, MGLRU generation scanning and
cgroup v2 memory.reclaim will *all* see it -- there is no "on the LRU
but exempt from scanning" mode; LRU membership means "in the kernel's
reclaim view". And making the LRU machinery truly understand hugetlb
means adapting per-memcg lruvec accounting, folio referenced/aging,
workingset shadows, MGLRU generations -- which is precisely spraying
if (hugetlb) special paths across generic reclaim code, in direct
conflict with the LSFMM 2024 consensus of cleaning up internals and
removing scattered special cases. A likely suggestion, "on the LRU
but reachable only by memory.reclaim, not kswapd", fails the same
way: the LRU is the shared input of every scanner; there is no
per-consumer visibility.
Implementation layer: under this split the series deliberately takes
the minimal-invasion route -- pool management stays in hugetlb.c,
every reusable generic leaf operation (rmap walk, slot allocation,
writeout, referenced/pin checks) is shared, and only the control flow
is mirrored, with "keep in sync" notes. Today hugetlb folios live on
per-hstate lists carrying pool semantics (reserve, surplus, demotion)
that the LRU machinery knows nothing about; if LRU-ification ever
happens, the mirrored structure folds over mechanically. kswapd
integration is an explicit non-goal, not a "never": the design does
not close the door (the referenced check simply comes back together
with that series' scan_control), but user-space-driven is the right
default for hugetlb, in the same direction as the Reserved THP RFC's
"reserved but reclaimable" roadmap.
Q2: Why extend swap to shared file-backed (hugetlbfs) mappings?
Without it, shared mappings pin the pool absolutely: an instantiated
shared huge page is unreclaimable until truncate/unlink, however cold
the data. Deployments mixing private and shared hugetlb get held
hostage by the shared half -- under pressure the anonymous half can
be swapped out, yet the pool still exhausts and new allocations
(including swap-in itself!) keep failing. Concrete users:
- QEMU guests backed by MAP_SHARED hugetlbfs files (-mem-path): idle
guest memory can finally be reclaimed with MADV_PAGEOUT instead of
pinning the pool for the VM's entire lifetime;
- services that hold shared hugetlbfs files for coordination/staging
but rarely touch them;
- ABI consistency: MADV_PAGEOUT succeeds on private hugetlb and every
4K/THP mapping including shmem -- shared hugetlb was the only
mapping type returning EINVAL, an ABI asymmetry users trip over;
- swapoff reachability: swapoff must be able to drain every kind of
swapped-out page; incomplete coverage leaves swapoff stuck on
anchors it cannot reach.
The anticipated objection -- "shared huge pages are for
performance-critical shared data, why swap them out" -- applies
equally to all shared file pages, yet shmem has always been
swappable. Everything here is opt-in (proactive reclaim only on
request): nothing moves unless userspace asks. Pool semantics hold
end to end: swap-out returns the page to the pool, swap-in competes
for pool memory again, and reservation exhaustion fails with
SIGBUS/-EIO exactly like a first-touch fault.
Testing
=======
New selftest tools/testing/selftests/mm/hugetlb_swap.c (~1120 lines,
37 assertions, 14 test functions), wired into the mm selftest
Makefile and run_vmtests.sh:
- test_pageout_basic: anonymous PAGEOUT -> PTE becomes a swap entry
-> the huge page returns to the pool -> fault swaps back in with
data intact
- test_fork_after_pageout: fork shares swap entries, writes COW
without polluting the parent, parent/child MM_SWAPENTS balance
- test_munmap_releases_swap: zap releases swap entries, SwapFree
recovers
- test_readfault_munmap_drains_swapcache: munmap after a read-fault
swapin drains the leftover swapcache copy (both the pool page and
the slots are returned)
- test_swapoff: swapoff drains swapped-out hugetlb pages
- test_swapin_pool_exhausted: swap-in on an exhausted pool fails with
SIGBUS (same semantics as a first fault; no OOM, no silent errors)
- test_shared_file_swap: shared-mapping pageout (PTE cleared, no swap
entry) -> both pread and fault swap back in with data intact, and
the anchor can be re-established repeatedly
- test_truncate_after_pageout: ftruncate(0) after pageout frees the
anchor and its slots; subsequent access gets SIGBUS
- test_reject_paths: VM_LOCKED / userfaultfd-registered / 1G gigantic
are rejected with EINVAL
- test_mprotect_mremap_smoke: mprotect/mremap over swap entries
- test_vmswap_accounting: VmSwap balances over the whole lifecycle:
+2048 kB at pageout, unchanged for the parent after fork, back to
zero at swapin/munmap (regression test for the unsigned-negation
zap accounting fix)
- test_soft_dirty_swap: soft-dirty bit round-trips between present
PTEs and swap entries, verified with clear_refs isolation
- test_uffd_wp_swap: uffd-wp bit set through the swap entry, carried
back to the present PTE at swapin, write fault delivers the WP
event
- test_hwpoison_swapcache: PAGEOUT -> swap-in leaves the folio in the
swap cache -> MADV_HWPOISON -> the next access must SIGBUS (must
not silently swap stale disk data back in)
Real-machine results (x86-64, 2M huge pages, QEMU + openEuler 24.03,
dedicated swap disk): 37 assertions, 35 pass / 0 fail / 2 expected
skips (mlock is a no-op for hugetlb; no free 1G gigantic page). Both
the VmSwap accounting balance and the hwpoison swapcache interception
are exercised end to end.
Every intermediate patch builds individually with zero warnings; the
!CONFIG_HUGETLB_PAGE configuration builds as well.
Patch split and ordering constraints
====================================
1. folio_swap_entry() helper + tree-wide conversion of all
folio->swap readers (zero semantic change; the changelog documents
the union pollution mechanism, the identity argument, the
alignment invariant, and why the conversion is tree-wide rather
than reachable-sites-only);
2. Anonymous swap-in (read-back first): hugetlb_anon_swapin() + fault
recognition of swap entries + fork/zap/swapoff/MM_SWAPENTS
balancing;
3. Anonymous swap-out core (producer second): the reclaim trio + the
rmap callback (rmap.h +1 line) + the hugetlb_folio_swap_supported()
size gate (with the swap_state.c VM_WARN backstop) + the madvise
MADV_PAGEOUT entry point; the changelog carries the
mirror/placement rationale and the full size-gate argument.
The order must not be flipped: the base hugetlb_fault() returns
ret = 0 for non-migration, non-hwpoison entries, so swap-out first
means an endless fault livelock on swapped-out pages;
4. memory-failure adaptations -- right after patch 3, because the
hole becomes reachable from patch 3 on, and a partial merge of
1-3 must not leave an mf hole behind;
5. File-backed swap as one piece: the shmem-model anchor +
hugetlbfs_writeout() + the refault/read swap-in pair + the
truncate/evict lifecycle + swapoff traversal. Kept together
because none of the four is correct without the others: read-back
alone is all dead code; swap-out without lifecycle management
leaks anchors at truncate; lifecycle without swapoff traversal
leaves swapoff unable to reach the anchors (file folios install no
swap PTEs, so the mm walk cannot find them). Any prefix of it
would be unhealthy, hence a single patch;
6. Selftest + documentation (hugetlbpage.rst).
Comments welcome.
[1] https://lore.kernel.org/r/20250514201729.48420-1-ryncsn@gmail.com
Zongkun Lei (6):
mm/swap: introduce folio_swap_entry() and convert all folio->swap
readers
mm/hugetlb: swap-in support for anonymous hugetlb folios
mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT
mm/memory-failure: handle swapcached hugetlb folios
mm/hugetlb: swap support for file-backed hugetlb folios
selftests/mm: add hugetlb_swap test, document hugetlb swap
Documentation/admin-guide/mm/hugetlbpage.rst | 29 +-
fs/hugetlbfs/inode.c | 128 +-
fs/proc/task_mmu.c | 11 +
include/linux/hugetlb.h | 58 +
include/linux/pagemap.h | 7 +
include/linux/rmap.h | 1 +
include/linux/swap.h | 23 +-
mm/huge_memory.c | 2 +-
mm/hugetlb.c | 1408 +++++++++++++++++-
mm/internal.h | 19 +-
mm/madvise.c | 70 +
mm/memcontrol-v1.c | 4 +-
mm/memcontrol.c | 2 +-
mm/memory-failure.c | 19 +
mm/memory.c | 2 +-
mm/page_io.c | 18 +-
mm/rmap.c | 175 ++-
mm/shmem.c | 6 +-
mm/swap.h | 14 +-
mm/swap_state.c | 55 +-
mm/swapfile.c | 87 +-
mm/userfaultfd.c | 2 +-
mm/util.c | 2 +-
mm/vmscan.c | 4 +-
mm/zswap.c | 4 +-
tools/testing/selftests/mm/Makefile | 1 +
tools/testing/selftests/mm/hugetlb_swap.c | 1118 ++++++++++++++
tools/testing/selftests/mm/run_vmtests.sh | 2 +
28 files changed, 3219 insertions(+), 52 deletions(-)
create mode 100644 tools/testing/selftests/mm/hugetlb_swap.c
--
2.53.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC PATCH 0/6] mm: swap support for hugetlb folios
2026-09-29 7:53 [RFC PATCH 0/6] mm: swap support for hugetlb folios Zongkun Lei
@ 2026-09-29 8:02 ` David Hildenbrand (Arm)
2026-09-29 9:26 ` Zongkun Lei
2026-09-29 8:18 ` Lorenzo Stoakes (ARM)
1 sibling, 1 reply; 5+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-29 8:02 UTC (permalink / raw)
To: Zongkun Lei, linux-mm
Cc: Andrew Morton, Muchun Song, Oscar Salvador, Peter Xu,
Matthew Wilcox, Kairui Song, Michal Hocko, Johannes Weiner,
Hugh Dickins, Naoya Horiguchi, Zi Yan, Lorenzo Stoakes,
linux-kernel
On 9/29/26 09:53, Zongkun Lei wrote:
> Hi all,
>
> This series introduces *user-driven* swap support for hugetlb folios,
> covering both private anonymous and shared file-backed mappings. It
> deliberately keeps the hugetlb reservation semantics intact: the
> kernel never reclaims hugetlb pages on its own -- hugetlb folios never
> go on an LRU, there is no hot/cold guessing, and nothing hooks into
> kswapd or direct reclaim. This is the same boundary the merged
> hugetlb demotion keeps: the kernel never reclaims reserved memory by
> its own judgement. Demotion is triggered explicitly by the
> administrator (writing nr_hugepages) and only touches free pool
> pages; this series is triggered explicitly by the workload (madvise)
> and acts on its own live mappings. In both cases the decision stays
> outside the kernel.
>
> Interface:
>
> - Swap-out: MADV_PAGEOUT via madvise(2)/process_madvise(2) learns to
> handle hugetlb mappings. For anonymous mappings the PTE becomes a
> swap entry; for shared mappings the PTE is simply cleared and the
> location record stays in the hugetlbfs page cache.
> - Swap-in: a fault reads the page back from swap.
> - Swapoff: all swapped-out hugetlb pages are read back. If the pool
> has no free huge pages at swap-in time, the behaviour matches a
> first-touch fault on a full pool: after bounded retries the process
> receives SIGBUS -- no silent failure, no data loss.
>
>
> Motivation
> ==========
>
> Today a hugetlb reservation is a one-way street: once instantiated, a
> huge page's memory is completely unreclaimable until the mapping is
> torn down. THP and 4K anonymous memory can simply be swapped out;
> hugetlb cannot. This gap forces operators to choose between "huge
> page performance" and "memory overcommit".
>
> VM overcommit. QEMU guests backed by MAP_SHARED hugetlbfs files
> (-mem-path) allocate all guest memory from the host's huge page pool.
> Overcommit means selling more guest memory than the pool holds --
> fine in practice because guests never all peak at once. But the
> extreme case is unavoidable: when every guest's load spikes at the
> same time, the pool runs out. Keeping the VMs alive then requires a
> swap mechanism that moves the memory the provider judges cold out of
> the way. Ballooning cannot save this scenario: it reclaims memory
> that is free *inside* the guest, and the extreme case is precisely
> that all guests are busy and have nothing to give. The extreme case
> is rare, but without an escape hatch providers dare not overcommit
> confidently, and pool utilization stays low. This series is that
> escape hatch: the VM management stack issues MADV_PAGEOUT on cold
> guest memory, VM availability survives the extreme case, and "huge
> page performance" vs "memory overcommit" stops being a choice.
>
> Hot-upgrade standby. A hot upgrade first brings up the new instance
> and switches traffic before the old one may be touched; to keep
> rollback fast, the old instance usually cannot be deleted right away
> and stands by for hours or days. During that time it has no traffic
> and its memory is cold, yet the process and its mappings still pin
> the pool until it is finally removed. With MADV_PAGEOUT the cold
> memory can be swapped out during standby and swapped back if a
> rollback happens. More generally: whenever the pool is full but
> tearing down the mapping is not an option, MADV_PAGEOUT is the only
> way to free space.
>
>
> Design
> ======
>
> Preparation: tree-wide folio_swap_entry() conversion (patch 1)
> --------------------------------------------------------------
>
> Hugetlb flags live in folio->private (a union with folio->swap), so
> while a folio sits in the swap cache the low bits of folio->swap.val
> can be polluted by HPAGEFLAG (with HVO enabled, bit 4 is set on
> virtually every hugetlb folio -- not a theoretical issue). The fix
> is folio_swap_entry() -- an identity transform for non-hugetlb folios
> (round_down(x, 1) == x for order-0; THP low bits are zero anyway) --
> plus a *tree-wide* conversion of all folio->swap readers. What
> review must prove shrinks from "N sites x reachability arguments" to
> "1 helper + 1 invariant (swap entries are folio-size aligned)", with
> zero semantic change to existing paths. It makes *generic* code
> hugetlb-safe instead of adding special cases in hugetlb.c, aligned
> with the hugetlb genericization direction.
>
> The companion size gate hugetlb_folio_swap_supported() (introduced in
> patch 3) clamps both bounds at once: the lower bound (HPAGEFLAG
> width, requiring nr_pages >= 128) and the upper bound (non-gigantic,
> <= PMD size, so one swap cluster holds the folio).
>
> Private anonymous mappings (patches 2-4)
> ----------------------------------------
>
> - Swap-in (consumer first, patch 2): hugetlb_anon_swapin() mirrors
> do_swap_page(): get_swap_device() pins the device,
> swap_cache_get_folio() finds in-flight folios, and on a miss the
> folio is allocated from the hugetlb pool and added to the swap
> cache via the new generic swap_cache_add_folio() helper (hugetlb
> folios do not come from the buddy allocator, so
> __swap_cache_alloc() cannot be used). Pool exhaustion retries
> briefly, then fails with SIGBUS, matching the
> reservation-exhaustion contract of the no-page fault path. The
> lifecycle is fully 4K-semantic: fork shares swap entries
> (copy_hugetlb_page_range() + swap_dup_entries_direct(), dropping
> the exclusive bit); zap/swapin release them; MM_SWAPENTS is
> balanced at swapout/fork/swapin/zap per mainline semantics (VmSwap
> stays meaningful); swapoff drains via hugetlb_unuse_vma().
> The read-back path must merge before the swap-out path: the base
> hugetlb_fault() returns ret = 0 for any entry that is neither
> migration nor hwpoison, so with swap-out first, touching a
> swapped-out page would livelock in an endless fault loop.
>
> - Swap-out (patch 3): hugetlb_reclaim_pages() /
> hugetlb_reclaim_folio_list() are trimmed mirrors of
> reclaim_pages()/shrink_folio_list() and live in hugetlb.c; the
> changelog of patch 3 explains why a mirror and why there. Swap PTE
> installation reuses the rmap walker: a new
> try_to_unmap_swap_hugetlb_one() (a static callback mirroring
> ttu_anon_swapbacked_folio(): single walk, swp_pte_prepare(),
> mm_prepare_for_swap_entries(), MM_SWAPENTS accounting) plus the
> exported wrapper try_to_unmap_swap_hugetlb(), following the
> try_to_unmap_poisoned_hugetlb_one() precedent. Cost to generic
> rmap: one added line in rmap.h, zero changes to existing code.
>
> - memory-failure (patch 4): anonymous hugetlb folios can now sit in
> the swap cache (the writeback window, the cached copy left after
> swap-in) -- a state memory failure has never seen before. Two
> adaptations: try_to_unmap() routes hugetlb folios by TTU_HWPOISON
> (set -> the poisoned handler; cleared -- the mf case that keeps a
> dirty swapcache folio -- -> the swap handler, matching 4K swapcache
> semantics), and me_huge_page() intercepts swapcache folios at the
> top (mirroring me_swapcache_dirty(): keep the folio in the swap
> cache and return MF_DELAYED), fixing the silent eviction of a clean
> poisoned folio that would otherwise mean a 2M pool leak plus a lost
> poison marker.
>
> Shared file-backed mappings (patch 5)
> -------------------------------------
>
> The shmem model: swap-out only clears the PTEs -- no swap entry is
> ever installed for a file folio -- and the anchor (a swap value
> entry, swp_to_radix_entry()) is left in the hugetlbfs page cache,
> holding a dup'ed swap count.
>
> - Swap-out: hugetlbfs_writeout() -- folio_alloc_swap() allocates the
> folio-sized slot range, the inode is linked on the per-inode
> swaplist (for swapoff traversal), folio_dup_swap() holds a swap
> count for the anchor, the page cache slot is replaced by the anchor
> in place, and swap_writeout() writes out. The reclaim loop drops
> its anon-only gate and madvise drops the VM_MAYSHARE rejection.
>
> - Refault swap-in: hugetlb_no_page() recognizes the anchor (the page
> cache lookup reports it as a hole) and swaps the folio back in via
> hugetlbfs_do_swapin(), replacing the anchor in place --
> indistinguishable from a regular page-cache hit afterwards. The
> read-back path must never lag the producer: the base fault path
> treats value entries as holes and would silently zero-fill over
> swapped-out data.
>
> - read(2) swap-in: hugetlbfs_read_iter() resolves anchors through
> hugetlbfs_swapin_read() instead of zero-filling; swapin failures
> surface as -EIO/-EFAULT, never stale data.
>
> - Lifecycle: truncate/hole-punch/evict free anchors and their slots
> via hugetlbfs_free_swap(), with two UAF guards (the swap cache
> entry is removed before the anchor's swap count is dropped; a
> refcount race means reclaim is isolating the folio, so the folio
> lock is held across swap_put_entries_direct() to wait it out).
> Inode eviction drains the swapoff traversal list via
> hugetlbfs_evict_drain_swaplist().
>
> - swapoff: try_to_unuse() hooks hugetlbfs_unuse() right after
> shmem_unuse(). File folios install no swap PTEs, so the mm walk
> cannot reach their slots -- the per-inode swaplist traversal can.
>
> - Symmetric accounting: hugetlb_cgroup, memcg swap/hugetlb charge,
> the global reserve and NR_HUGETLB are paired with free_huge_folio()
> on every path; the vma-less read/swapoff swap-in charges the
> caller's context, the same policy as shmem swapoff.
>
>
> Base and dependencies
> =====================
>
> Based on mm-unstable; no external patch dependencies.
>
> The implementation builds on Kairui Song's swap table work: the
> folio-sized contiguous swap slot ranges (folio_alloc_swap() /
> folio_dup_swap() / folio_put_swap()) are the primitives that make
> this possible at all -- without them a 2M folio would need 512
> independent swap entries. The merged phases of that work are already
> in mm-unstable [1].
>
> Looking ahead, the swap table roadmap may move swap entries out of
> folio->swap; removing folio_swap_entry() then is a mechanical
> operation -- the call sites are already concentrated in a single
> helper, so it will be easier to delete than what we have today.
>
> This series is complementary to the Reserved THP RFC (Qi Zheng), not
> competing with it. Reserved THP explores the endgame -- reserved,
> swappable, unsplittable large folios that may absorb most hugetlb use
> cases. But the LSFMM 2024 consensus is that the hugetlbfs ABI stays
> (migrating away is a 15-20 year scale); as long as the ABI exists,
> its unreclaimable-pool problem needs a direct solution. The two also
> build on the same swap table foundation, so progress on either side
> de-risks the other.
>
>
> Anticipated questions
> =====================
>
> Q1: Why not put hugetlb folios on the LRU and reuse vmscan?
>
> Semantic layer: hugetlb's ABI contract is "reservation = latency
> determinism", and the kernel must not tear that up unilaterally --
> this is the historical reason hugetlb could not be swapped at all.
> Letting the kernel reclaim reserved pages under global memory
> pressure based on its own hot/cold guesses (folio_referenced() style
> access-pattern inference) injects fault latency exactly where it must
> never appear -- the hot-upgrade standby and pre-traffic-peak moments
> are precisely why hugetlb is used. So the split in this series is:
> the kernel provides the mechanism, the policy belongs to userspace.
> MADV_PAGEOUT extends to hugetlb the proactive contract 4K anonymous
> memory already has; the workload (orchestrator, JVM, database) knows
> its warmup windows and traffic peaks, so hot/cold decisions stay on
> the side with the most information. This is the same boundary the
> only merged hugetlb reclaim mechanism, demotion, keeps: the kernel
> never reclaims reserved memory by its own judgement -- demotion is
> administrator-triggered (nr_hugepages) and touches only free pool
> pages, this series is workload-triggered (madvise) and touches its
> own live mappings; in both cases the decision is outside the kernel.
>
> Technical layer: an LRU is not a neutral queue; it is the input
> structure of the kernel's reclaim policy. Once a hugetlb folio is on
> an lruvec, kswapd, direct reclaim, MGLRU generation scanning and
> cgroup v2 memory.reclaim will *all* see it -- there is no "on the LRU
> but exempt from scanning" mode; LRU membership means "in the kernel's
> reclaim view". And making the LRU machinery truly understand hugetlb
> means adapting per-memcg lruvec accounting, folio referenced/aging,
> workingset shadows, MGLRU generations -- which is precisely spraying
> if (hugetlb) special paths across generic reclaim code, in direct
> conflict with the LSFMM 2024 consensus of cleaning up internals and
> removing scattered special cases. A likely suggestion, "on the LRU
> but reachable only by memory.reclaim, not kswapd", fails the same
> way: the LRU is the shared input of every scanner; there is no
> per-consumer visibility.
>
> Implementation layer: under this split the series deliberately takes
> the minimal-invasion route -- pool management stays in hugetlb.c,
> every reusable generic leaf operation (rmap walk, slot allocation,
> writeout, referenced/pin checks) is shared, and only the control flow
> is mirrored, with "keep in sync" notes. Today hugetlb folios live on
> per-hstate lists carrying pool semantics (reserve, surplus, demotion)
> that the LRU machinery knows nothing about; if LRU-ification ever
> happens, the mirrored structure folds over mechanically. kswapd
> integration is an explicit non-goal, not a "never": the design does
> not close the door (the referenced check simply comes back together
> with that series' scan_control), but user-space-driven is the right
> default for hugetlb, in the same direction as the Reserved THP RFC's
> "reserved but reclaimable" roadmap.
>
> Q2: Why extend swap to shared file-backed (hugetlbfs) mappings?
>
> Without it, shared mappings pin the pool absolutely: an instantiated
> shared huge page is unreclaimable until truncate/unlink, however cold
> the data. Deployments mixing private and shared hugetlb get held
> hostage by the shared half -- under pressure the anonymous half can
> be swapped out, yet the pool still exhausts and new allocations
> (including swap-in itself!) keep failing. Concrete users:
>
> - QEMU guests backed by MAP_SHARED hugetlbfs files (-mem-path): idle
> guest memory can finally be reclaimed with MADV_PAGEOUT instead of
> pinning the pool for the VM's entire lifetime;
> - services that hold shared hugetlbfs files for coordination/staging
> but rarely touch them;
> - ABI consistency: MADV_PAGEOUT succeeds on private hugetlb and every
> 4K/THP mapping including shmem -- shared hugetlb was the only
> mapping type returning EINVAL, an ABI asymmetry users trip over;
> - swapoff reachability: swapoff must be able to drain every kind of
> swapped-out page; incomplete coverage leaves swapoff stuck on
> anchors it cannot reach.
>
> The anticipated objection -- "shared huge pages are for
> performance-critical shared data, why swap them out" -- applies
> equally to all shared file pages, yet shmem has always been
> swappable. Everything here is opt-in (proactive reclaim only on
> request): nothing moves unless userspace asks. Pool semantics hold
> end to end: swap-out returns the page to the pool, swap-in competes
> for pool memory again, and reservation exhaustion fails with
> SIGBUS/-EIO exactly like a first-touch fault.
>
>
> Testing
> =======
>
> New selftest tools/testing/selftests/mm/hugetlb_swap.c (~1120 lines,
> 37 assertions, 14 test functions), wired into the mm selftest
> Makefile and run_vmtests.sh:
>
> - test_pageout_basic: anonymous PAGEOUT -> PTE becomes a swap entry
> -> the huge page returns to the pool -> fault swaps back in with
> data intact
> - test_fork_after_pageout: fork shares swap entries, writes COW
> without polluting the parent, parent/child MM_SWAPENTS balance
> - test_munmap_releases_swap: zap releases swap entries, SwapFree
> recovers
> - test_readfault_munmap_drains_swapcache: munmap after a read-fault
> swapin drains the leftover swapcache copy (both the pool page and
> the slots are returned)
> - test_swapoff: swapoff drains swapped-out hugetlb pages
> - test_swapin_pool_exhausted: swap-in on an exhausted pool fails with
> SIGBUS (same semantics as a first fault; no OOM, no silent errors)
> - test_shared_file_swap: shared-mapping pageout (PTE cleared, no swap
> entry) -> both pread and fault swap back in with data intact, and
> the anchor can be re-established repeatedly
> - test_truncate_after_pageout: ftruncate(0) after pageout frees the
> anchor and its slots; subsequent access gets SIGBUS
> - test_reject_paths: VM_LOCKED / userfaultfd-registered / 1G gigantic
> are rejected with EINVAL
> - test_mprotect_mremap_smoke: mprotect/mremap over swap entries
> - test_vmswap_accounting: VmSwap balances over the whole lifecycle:
> +2048 kB at pageout, unchanged for the parent after fork, back to
> zero at swapin/munmap (regression test for the unsigned-negation
> zap accounting fix)
> - test_soft_dirty_swap: soft-dirty bit round-trips between present
> PTEs and swap entries, verified with clear_refs isolation
> - test_uffd_wp_swap: uffd-wp bit set through the swap entry, carried
> back to the present PTE at swapin, write fault delivers the WP
> event
> - test_hwpoison_swapcache: PAGEOUT -> swap-in leaves the folio in the
> swap cache -> MADV_HWPOISON -> the next access must SIGBUS (must
> not silently swap stale disk data back in)
>
> Real-machine results (x86-64, 2M huge pages, QEMU + openEuler 24.03,
> dedicated swap disk): 37 assertions, 35 pass / 0 fail / 2 expected
> skips (mlock is a no-op for hugetlb; no free 1G gigantic page). Both
> the VmSwap accounting balance and the hwpoison swapcache interception
> are exercised end to end.
>
> Every intermediate patch builds individually with zero warnings; the
> !CONFIG_HUGETLB_PAGE configuration builds as well.
>
>
> Patch split and ordering constraints
> ====================================
>
> 1. folio_swap_entry() helper + tree-wide conversion of all
> folio->swap readers (zero semantic change; the changelog documents
> the union pollution mechanism, the identity argument, the
> alignment invariant, and why the conversion is tree-wide rather
> than reachable-sites-only);
> 2. Anonymous swap-in (read-back first): hugetlb_anon_swapin() + fault
> recognition of swap entries + fork/zap/swapoff/MM_SWAPENTS
> balancing;
> 3. Anonymous swap-out core (producer second): the reclaim trio + the
> rmap callback (rmap.h +1 line) + the hugetlb_folio_swap_supported()
> size gate (with the swap_state.c VM_WARN backstop) + the madvise
> MADV_PAGEOUT entry point; the changelog carries the
> mirror/placement rationale and the full size-gate argument.
> The order must not be flipped: the base hugetlb_fault() returns
> ret = 0 for non-migration, non-hwpoison entries, so swap-out first
> means an endless fault livelock on swapped-out pages;
> 4. memory-failure adaptations -- right after patch 3, because the
> hole becomes reachable from patch 3 on, and a partial merge of
> 1-3 must not leave an mf hole behind;
> 5. File-backed swap as one piece: the shmem-model anchor +
> hugetlbfs_writeout() + the refault/read swap-in pair + the
> truncate/evict lifecycle + swapoff traversal. Kept together
> because none of the four is correct without the others: read-back
> alone is all dead code; swap-out without lifecycle management
> leaks anchors at truncate; lifecycle without swapoff traversal
> leaves swapoff unable to reach the anchors (file folios install no
> swap PTEs, so the mm walk cannot find them). Any prefix of it
> would be unhealthy, hence a single patch;
> 6. Selftest + documentation (hugetlbpage.rst).
>
> Comments welcome.
>
> [1] https://lore.kernel.org/r/20250514201729.48420-1-ryncsn@gmail.com
>
> Zongkun Lei (6):
> mm/swap: introduce folio_swap_entry() and convert all folio->swap
> readers
> mm/hugetlb: swap-in support for anonymous hugetlb folios
> mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT
> mm/memory-failure: handle swapcached hugetlb folios
> mm/hugetlb: swap support for file-backed hugetlb folios
> selftests/mm: add hugetlb_swap test, document hugetlb swap
>
> Documentation/admin-guide/mm/hugetlbpage.rst | 29 +-
> fs/hugetlbfs/inode.c | 128 +-
> fs/proc/task_mmu.c | 11 +
> include/linux/hugetlb.h | 58 +
> include/linux/pagemap.h | 7 +
> include/linux/rmap.h | 1 +
> include/linux/swap.h | 23 +-
> mm/huge_memory.c | 2 +-
> mm/hugetlb.c | 1408 +++++++++++++++++-
> mm/internal.h | 19 +-
> mm/madvise.c | 70 +
> mm/memcontrol-v1.c | 4 +-
> mm/memcontrol.c | 2 +-
> mm/memory-failure.c | 19 +
> mm/memory.c | 2 +-
> mm/page_io.c | 18 +-
> mm/rmap.c | 175 ++-
> mm/shmem.c | 6 +-
> mm/swap.h | 14 +-
> mm/swap_state.c | 55 +-
> mm/swapfile.c | 87 +-
> mm/userfaultfd.c | 2 +-
> mm/util.c | 2 +-
> mm/vmscan.c | 4 +-
> mm/zswap.c | 4 +-
> tools/testing/selftests/mm/Makefile | 1 +
> tools/testing/selftests/mm/hugetlb_swap.c | 1118 ++++++++++++++
> tools/testing/selftests/mm/run_vmtests.sh | 2 +
> 28 files changed, 3219 insertions(+), 52 deletions(-)
Hugetlb is effectively in a feature freeze. There is no interest in maintaining
even more of the special hugetlb sauce, sorry. :/
--
Cheers,
David
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC PATCH 0/6] mm: swap support for hugetlb folios
2026-09-29 7:53 [RFC PATCH 0/6] mm: swap support for hugetlb folios Zongkun Lei
2026-09-29 8:02 ` David Hildenbrand (Arm)
@ 2026-09-29 8:18 ` Lorenzo Stoakes (ARM)
2026-09-29 9:30 ` Zongkun Lei
1 sibling, 1 reply; 5+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-29 8:18 UTC (permalink / raw)
To: Zongkun Lei
Cc: linux-mm, Andrew Morton, David Hildenbrand, Muchun Song,
Oscar Salvador, Peter Xu, Matthew Wilcox, Kairui Song,
Michal Hocko, Johannes Weiner, Hugh Dickins, Naoya Horiguchi,
Zi Yan, linux-kernel
Hi,
What David said - hugetlb is in freeze so we're unfortunately not interested in
such a big change.
Also you somehow sent the cover letter and the patches separately.
But also:
For one of several possible reasons this mail has triggered an AI
detector script.
Note that, while we are fine with AI assistance, it is kernel
policy that you must disclose this with an Assisted-by tag like:
Assisted-by: LLM
See https://docs.kernel.org/process/coding-assistants.html
It is also kernel policy that you must fully understand and take
responsibility for every patch that you send.
See https://docs.kernel.org/process/generated-content.html most
notably:
If tools permit you to generate a contribution automatically, expect
additional scrutiny in proportion to how much of it was generated.
As with the output of any tooling, the result may be incorrect or
inappropriate. You are expected to understand and to be able to
defend everything you submit. If you are unable to do so, then do
not submit the resulting changes.
If you do so anyway, maintainers are entitled to reject your series
without detailed review.
In general, if you are a newcomer to mm, we expect you to do smaller work
before moving on to larger changes, so you build understanding of both the
technical aspects of mm and how we do things.
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC PATCH 0/6] mm: swap support for hugetlb folios
2026-09-29 8:02 ` David Hildenbrand (Arm)
@ 2026-09-29 9:26 ` Zongkun Lei
0 siblings, 0 replies; 5+ messages in thread
From: Zongkun Lei @ 2026-09-29 9:26 UTC (permalink / raw)
To: David Hildenbrand
Cc: Zongkun Lei, linux-mm, Andrew Morton, Muchun Song,
Oscar Salvador, Peter Xu, Matthew Wilcox, Kairui Song,
Michal Hocko, Johannes Weiner, Hugh Dickins, Naoya Horiguchi,
Zi Yan, Lorenzo Stoakes, linux-kernel
On Tue, 29 Sep 2026 10:02:11 +0200, David Hildenbrand wrote:
> Hugetlb is effectively in a feature freeze. There is no interest in
> maintaining even more of the special hugetlb sauce, sorry. :/
Understood -- thank you for the direct answer.
My mistake was misreading the subsystem's direction. I had read the
earlier list discussions and concluded that the red line was touching
hugetlb's reservation semantics, so I designed the whole series to
keep them intact -- user-driven only, no LRU, no reclaim hooks. I
see now that the actual position is a feature freeze: not "new
features are welcome if reservations stay sacred", but no new
features at all.
Going forward I will stick to work that reduces hugetlb's divergence
from generic code rather than adding to it, and I will float ideas
on the list before putting serious effort into anything big.
Thanks,
Zongkun
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC PATCH 0/6] mm: swap support for hugetlb folios
2026-09-29 8:18 ` Lorenzo Stoakes (ARM)
@ 2026-09-29 9:30 ` Zongkun Lei
0 siblings, 0 replies; 5+ messages in thread
From: Zongkun Lei @ 2026-09-29 9:30 UTC (permalink / raw)
To: Lorenzo Stoakes
Cc: Zongkun Lei, linux-mm, Andrew Morton, David Hildenbrand,
Muchun Song, Oscar Salvador, Peter Xu, Matthew Wilcox,
Kairui Song, Michal Hocko, Johannes Weiner, Hugh Dickins,
Naoya Horiguchi, Zi Yan, linux-kernel
On Tue, 29 Sep 2026 09:18:51 +0100, Lorenzo Stoakes wrote:
> What David said - hugetlb is in freeze so we're unfortunately not
> interested in such a big change.
Understood -- my mistake was misreading the subsystem's direction:
from earlier list discussions I had concluded that keeping hugetlb's
reservation semantics intact was the bar, and designed the series
around that; I see now that the actual position is a broader feature
freeze. Thank you both for the straight answer.
> Also you somehow sent the cover letter and the patches separately.
My apologies for the mess. My SMTP provider (QQ) rewrites the
Message-ID header of outgoing mail but leaves In-Reply-To untouched,
so the patches reference a cover Message-ID that never reached the
list.
> ... triggered an AI detector script ...
Fair, and thank you for raising it openly. Full disclosure: I used
an LLM to help organise and polish the English of the cover letter
and the per-patch commit messages. It also reviewed the code, and
some fixes found during those reviews were applied with its help.
That said, I understand every line of the final code and can defend
all of it in as much detail as anyone wants. Going forward, any
submission from me will carry the disclosure tag that
Documentation/process/coding-assistants.html requires.
> ... we expect you to do smaller work before moving on to larger
> changes ...
Point taken, and it is fair. I will look for smaller mm work to
build that track record -- pointers to a good entry point are
welcome.
Thanks for the detailed reply, and for pointing me in the right
direction.
Zongkun
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-29 9:31 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 7:53 [RFC PATCH 0/6] mm: swap support for hugetlb folios Zongkun Lei
2026-09-29 8:02 ` David Hildenbrand (Arm)
2026-09-29 9:26 ` Zongkun Lei
2026-09-29 8:18 ` Lorenzo Stoakes (ARM)
2026-09-29 9:30 ` Zongkun Lei
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®