mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH 0/6] mm: swap support for hugetlb folios
@ 2026-09-29  7:53 Zongkun Lei
  2026-09-29  8:02 ` David Hildenbrand (Arm)
  2026-09-29  8:18 ` Lorenzo Stoakes (ARM)
  0 siblings, 2 replies; 5+ messages in thread
From: Zongkun Lei @ 2026-09-29  7:53 UTC (permalink / raw)
  To: linux-mm
  Cc: Andrew Morton, David Hildenbrand, Muchun Song, Oscar Salvador,
	Peter Xu, Matthew Wilcox, Kairui Song, Michal Hocko,
	Johannes Weiner, Hugh Dickins, Naoya Horiguchi, Zi Yan,
	Lorenzo Stoakes, linux-kernel, Zongkun Lei

Hi all,

This series introduces *user-driven* swap support for hugetlb folios,
covering both private anonymous and shared file-backed mappings.  It
deliberately keeps the hugetlb reservation semantics intact: the
kernel never reclaims hugetlb pages on its own -- hugetlb folios never
go on an LRU, there is no hot/cold guessing, and nothing hooks into
kswapd or direct reclaim.  This is the same boundary the merged
hugetlb demotion keeps: the kernel never reclaims reserved memory by
its own judgement.  Demotion is triggered explicitly by the
administrator (writing nr_hugepages) and only touches free pool
pages; this series is triggered explicitly by the workload (madvise)
and acts on its own live mappings.  In both cases the decision stays
outside the kernel.

Interface:

- Swap-out: MADV_PAGEOUT via madvise(2)/process_madvise(2) learns to
  handle hugetlb mappings.  For anonymous mappings the PTE becomes a
  swap entry; for shared mappings the PTE is simply cleared and the
  location record stays in the hugetlbfs page cache.
- Swap-in: a fault reads the page back from swap.
- Swapoff: all swapped-out hugetlb pages are read back.  If the pool
  has no free huge pages at swap-in time, the behaviour matches a
  first-touch fault on a full pool: after bounded retries the process
  receives SIGBUS -- no silent failure, no data loss.


Motivation
==========

Today a hugetlb reservation is a one-way street: once instantiated, a
huge page's memory is completely unreclaimable until the mapping is
torn down.  THP and 4K anonymous memory can simply be swapped out;
hugetlb cannot.  This gap forces operators to choose between "huge
page performance" and "memory overcommit".

VM overcommit.  QEMU guests backed by MAP_SHARED hugetlbfs files
(-mem-path) allocate all guest memory from the host's huge page pool.
Overcommit means selling more guest memory than the pool holds --
fine in practice because guests never all peak at once.  But the
extreme case is unavoidable: when every guest's load spikes at the
same time, the pool runs out.  Keeping the VMs alive then requires a
swap mechanism that moves the memory the provider judges cold out of
the way.  Ballooning cannot save this scenario: it reclaims memory
that is free *inside* the guest, and the extreme case is precisely
that all guests are busy and have nothing to give.  The extreme case
is rare, but without an escape hatch providers dare not overcommit
confidently, and pool utilization stays low.  This series is that
escape hatch: the VM management stack issues MADV_PAGEOUT on cold
guest memory, VM availability survives the extreme case, and "huge
page performance" vs "memory overcommit" stops being a choice.

Hot-upgrade standby.  A hot upgrade first brings up the new instance
and switches traffic before the old one may be touched; to keep
rollback fast, the old instance usually cannot be deleted right away
and stands by for hours or days.  During that time it has no traffic
and its memory is cold, yet the process and its mappings still pin
the pool until it is finally removed.  With MADV_PAGEOUT the cold
memory can be swapped out during standby and swapped back if a
rollback happens.  More generally: whenever the pool is full but
tearing down the mapping is not an option, MADV_PAGEOUT is the only
way to free space.


Design
======

Preparation: tree-wide folio_swap_entry() conversion (patch 1)
--------------------------------------------------------------

Hugetlb flags live in folio->private (a union with folio->swap), so
while a folio sits in the swap cache the low bits of folio->swap.val
can be polluted by HPAGEFLAG (with HVO enabled, bit 4 is set on
virtually every hugetlb folio -- not a theoretical issue).  The fix
is folio_swap_entry() -- an identity transform for non-hugetlb folios
(round_down(x, 1) == x for order-0; THP low bits are zero anyway) --
plus a *tree-wide* conversion of all folio->swap readers.  What
review must prove shrinks from "N sites x reachability arguments" to
"1 helper + 1 invariant (swap entries are folio-size aligned)", with
zero semantic change to existing paths.  It makes *generic* code
hugetlb-safe instead of adding special cases in hugetlb.c, aligned
with the hugetlb genericization direction.

The companion size gate hugetlb_folio_swap_supported() (introduced in
patch 3) clamps both bounds at once: the lower bound (HPAGEFLAG
width, requiring nr_pages >= 128) and the upper bound (non-gigantic,
<= PMD size, so one swap cluster holds the folio).

Private anonymous mappings (patches 2-4)
----------------------------------------

- Swap-in (consumer first, patch 2): hugetlb_anon_swapin() mirrors
  do_swap_page(): get_swap_device() pins the device,
  swap_cache_get_folio() finds in-flight folios, and on a miss the
  folio is allocated from the hugetlb pool and added to the swap
  cache via the new generic swap_cache_add_folio() helper (hugetlb
  folios do not come from the buddy allocator, so
  __swap_cache_alloc() cannot be used).  Pool exhaustion retries
  briefly, then fails with SIGBUS, matching the
  reservation-exhaustion contract of the no-page fault path.  The
  lifecycle is fully 4K-semantic: fork shares swap entries
  (copy_hugetlb_page_range() + swap_dup_entries_direct(), dropping
  the exclusive bit); zap/swapin release them; MM_SWAPENTS is
  balanced at swapout/fork/swapin/zap per mainline semantics (VmSwap
  stays meaningful); swapoff drains via hugetlb_unuse_vma().
  The read-back path must merge before the swap-out path: the base
  hugetlb_fault() returns ret = 0 for any entry that is neither
  migration nor hwpoison, so with swap-out first, touching a
  swapped-out page would livelock in an endless fault loop.

- Swap-out (patch 3): hugetlb_reclaim_pages() /
  hugetlb_reclaim_folio_list() are trimmed mirrors of
  reclaim_pages()/shrink_folio_list() and live in hugetlb.c; the
  changelog of patch 3 explains why a mirror and why there.  Swap PTE
  installation reuses the rmap walker: a new
  try_to_unmap_swap_hugetlb_one() (a static callback mirroring
  ttu_anon_swapbacked_folio(): single walk, swp_pte_prepare(),
  mm_prepare_for_swap_entries(), MM_SWAPENTS accounting) plus the
  exported wrapper try_to_unmap_swap_hugetlb(), following the
  try_to_unmap_poisoned_hugetlb_one() precedent.  Cost to generic
  rmap: one added line in rmap.h, zero changes to existing code.

- memory-failure (patch 4): anonymous hugetlb folios can now sit in
  the swap cache (the writeback window, the cached copy left after
  swap-in) -- a state memory failure has never seen before.  Two
  adaptations: try_to_unmap() routes hugetlb folios by TTU_HWPOISON
  (set -> the poisoned handler; cleared -- the mf case that keeps a
  dirty swapcache folio -- -> the swap handler, matching 4K swapcache
  semantics), and me_huge_page() intercepts swapcache folios at the
  top (mirroring me_swapcache_dirty(): keep the folio in the swap
  cache and return MF_DELAYED), fixing the silent eviction of a clean
  poisoned folio that would otherwise mean a 2M pool leak plus a lost
  poison marker.

Shared file-backed mappings (patch 5)
-------------------------------------

The shmem model: swap-out only clears the PTEs -- no swap entry is
ever installed for a file folio -- and the anchor (a swap value
entry, swp_to_radix_entry()) is left in the hugetlbfs page cache,
holding a dup'ed swap count.

- Swap-out: hugetlbfs_writeout() -- folio_alloc_swap() allocates the
  folio-sized slot range, the inode is linked on the per-inode
  swaplist (for swapoff traversal), folio_dup_swap() holds a swap
  count for the anchor, the page cache slot is replaced by the anchor
  in place, and swap_writeout() writes out.  The reclaim loop drops
  its anon-only gate and madvise drops the VM_MAYSHARE rejection.

- Refault swap-in: hugetlb_no_page() recognizes the anchor (the page
  cache lookup reports it as a hole) and swaps the folio back in via
  hugetlbfs_do_swapin(), replacing the anchor in place --
  indistinguishable from a regular page-cache hit afterwards.  The
  read-back path must never lag the producer: the base fault path
  treats value entries as holes and would silently zero-fill over
  swapped-out data.

- read(2) swap-in: hugetlbfs_read_iter() resolves anchors through
  hugetlbfs_swapin_read() instead of zero-filling; swapin failures
  surface as -EIO/-EFAULT, never stale data.

- Lifecycle: truncate/hole-punch/evict free anchors and their slots
  via hugetlbfs_free_swap(), with two UAF guards (the swap cache
  entry is removed before the anchor's swap count is dropped; a
  refcount race means reclaim is isolating the folio, so the folio
  lock is held across swap_put_entries_direct() to wait it out).
  Inode eviction drains the swapoff traversal list via
  hugetlbfs_evict_drain_swaplist().

- swapoff: try_to_unuse() hooks hugetlbfs_unuse() right after
  shmem_unuse().  File folios install no swap PTEs, so the mm walk
  cannot reach their slots -- the per-inode swaplist traversal can.

- Symmetric accounting: hugetlb_cgroup, memcg swap/hugetlb charge,
  the global reserve and NR_HUGETLB are paired with free_huge_folio()
  on every path; the vma-less read/swapoff swap-in charges the
  caller's context, the same policy as shmem swapoff.


Base and dependencies
=====================

Based on mm-unstable; no external patch dependencies.

The implementation builds on Kairui Song's swap table work: the
folio-sized contiguous swap slot ranges (folio_alloc_swap() /
folio_dup_swap() / folio_put_swap()) are the primitives that make
this possible at all -- without them a 2M folio would need 512
independent swap entries.  The merged phases of that work are already
in mm-unstable [1].

Looking ahead, the swap table roadmap may move swap entries out of
folio->swap; removing folio_swap_entry() then is a mechanical
operation -- the call sites are already concentrated in a single
helper, so it will be easier to delete than what we have today.

This series is complementary to the Reserved THP RFC (Qi Zheng), not
competing with it.  Reserved THP explores the endgame -- reserved,
swappable, unsplittable large folios that may absorb most hugetlb use
cases.  But the LSFMM 2024 consensus is that the hugetlbfs ABI stays
(migrating away is a 15-20 year scale); as long as the ABI exists,
its unreclaimable-pool problem needs a direct solution.  The two also
build on the same swap table foundation, so progress on either side
de-risks the other.


Anticipated questions
=====================

Q1: Why not put hugetlb folios on the LRU and reuse vmscan?

Semantic layer: hugetlb's ABI contract is "reservation = latency
determinism", and the kernel must not tear that up unilaterally --
this is the historical reason hugetlb could not be swapped at all.
Letting the kernel reclaim reserved pages under global memory
pressure based on its own hot/cold guesses (folio_referenced() style
access-pattern inference) injects fault latency exactly where it must
never appear -- the hot-upgrade standby and pre-traffic-peak moments
are precisely why hugetlb is used.  So the split in this series is:
the kernel provides the mechanism, the policy belongs to userspace.
MADV_PAGEOUT extends to hugetlb the proactive contract 4K anonymous
memory already has; the workload (orchestrator, JVM, database) knows
its warmup windows and traffic peaks, so hot/cold decisions stay on
the side with the most information.  This is the same boundary the
only merged hugetlb reclaim mechanism, demotion, keeps: the kernel
never reclaims reserved memory by its own judgement -- demotion is
administrator-triggered (nr_hugepages) and touches only free pool
pages, this series is workload-triggered (madvise) and touches its
own live mappings; in both cases the decision is outside the kernel.

Technical layer: an LRU is not a neutral queue; it is the input
structure of the kernel's reclaim policy.  Once a hugetlb folio is on
an lruvec, kswapd, direct reclaim, MGLRU generation scanning and
cgroup v2 memory.reclaim will *all* see it -- there is no "on the LRU
but exempt from scanning" mode; LRU membership means "in the kernel's
reclaim view".  And making the LRU machinery truly understand hugetlb
means adapting per-memcg lruvec accounting, folio referenced/aging,
workingset shadows, MGLRU generations -- which is precisely spraying
if (hugetlb) special paths across generic reclaim code, in direct
conflict with the LSFMM 2024 consensus of cleaning up internals and
removing scattered special cases.  A likely suggestion, "on the LRU
but reachable only by memory.reclaim, not kswapd", fails the same
way: the LRU is the shared input of every scanner; there is no
per-consumer visibility.

Implementation layer: under this split the series deliberately takes
the minimal-invasion route -- pool management stays in hugetlb.c,
every reusable generic leaf operation (rmap walk, slot allocation,
writeout, referenced/pin checks) is shared, and only the control flow
is mirrored, with "keep in sync" notes.  Today hugetlb folios live on
per-hstate lists carrying pool semantics (reserve, surplus, demotion)
that the LRU machinery knows nothing about; if LRU-ification ever
happens, the mirrored structure folds over mechanically.  kswapd
integration is an explicit non-goal, not a "never": the design does
not close the door (the referenced check simply comes back together
with that series' scan_control), but user-space-driven is the right
default for hugetlb, in the same direction as the Reserved THP RFC's
"reserved but reclaimable" roadmap.

Q2: Why extend swap to shared file-backed (hugetlbfs) mappings?

Without it, shared mappings pin the pool absolutely: an instantiated
shared huge page is unreclaimable until truncate/unlink, however cold
the data.  Deployments mixing private and shared hugetlb get held
hostage by the shared half -- under pressure the anonymous half can
be swapped out, yet the pool still exhausts and new allocations
(including swap-in itself!) keep failing.  Concrete users:

- QEMU guests backed by MAP_SHARED hugetlbfs files (-mem-path): idle
  guest memory can finally be reclaimed with MADV_PAGEOUT instead of
  pinning the pool for the VM's entire lifetime;
- services that hold shared hugetlbfs files for coordination/staging
  but rarely touch them;
- ABI consistency: MADV_PAGEOUT succeeds on private hugetlb and every
  4K/THP mapping including shmem -- shared hugetlb was the only
  mapping type returning EINVAL, an ABI asymmetry users trip over;
- swapoff reachability: swapoff must be able to drain every kind of
  swapped-out page; incomplete coverage leaves swapoff stuck on
  anchors it cannot reach.

The anticipated objection -- "shared huge pages are for
performance-critical shared data, why swap them out" -- applies
equally to all shared file pages, yet shmem has always been
swappable.  Everything here is opt-in (proactive reclaim only on
request): nothing moves unless userspace asks.  Pool semantics hold
end to end: swap-out returns the page to the pool, swap-in competes
for pool memory again, and reservation exhaustion fails with
SIGBUS/-EIO exactly like a first-touch fault.


Testing
=======

New selftest tools/testing/selftests/mm/hugetlb_swap.c (~1120 lines,
37 assertions, 14 test functions), wired into the mm selftest
Makefile and run_vmtests.sh:

- test_pageout_basic: anonymous PAGEOUT -> PTE becomes a swap entry
  -> the huge page returns to the pool -> fault swaps back in with
  data intact
- test_fork_after_pageout: fork shares swap entries, writes COW
  without polluting the parent, parent/child MM_SWAPENTS balance
- test_munmap_releases_swap: zap releases swap entries, SwapFree
  recovers
- test_readfault_munmap_drains_swapcache: munmap after a read-fault
  swapin drains the leftover swapcache copy (both the pool page and
  the slots are returned)
- test_swapoff: swapoff drains swapped-out hugetlb pages
- test_swapin_pool_exhausted: swap-in on an exhausted pool fails with
  SIGBUS (same semantics as a first fault; no OOM, no silent errors)
- test_shared_file_swap: shared-mapping pageout (PTE cleared, no swap
  entry) -> both pread and fault swap back in with data intact, and
  the anchor can be re-established repeatedly
- test_truncate_after_pageout: ftruncate(0) after pageout frees the
  anchor and its slots; subsequent access gets SIGBUS
- test_reject_paths: VM_LOCKED / userfaultfd-registered / 1G gigantic
  are rejected with EINVAL
- test_mprotect_mremap_smoke: mprotect/mremap over swap entries
- test_vmswap_accounting: VmSwap balances over the whole lifecycle:
  +2048 kB at pageout, unchanged for the parent after fork, back to
  zero at swapin/munmap (regression test for the unsigned-negation
  zap accounting fix)
- test_soft_dirty_swap: soft-dirty bit round-trips between present
  PTEs and swap entries, verified with clear_refs isolation
- test_uffd_wp_swap: uffd-wp bit set through the swap entry, carried
  back to the present PTE at swapin, write fault delivers the WP
  event
- test_hwpoison_swapcache: PAGEOUT -> swap-in leaves the folio in the
  swap cache -> MADV_HWPOISON -> the next access must SIGBUS (must
  not silently swap stale disk data back in)

Real-machine results (x86-64, 2M huge pages, QEMU + openEuler 24.03,
dedicated swap disk): 37 assertions, 35 pass / 0 fail / 2 expected
skips (mlock is a no-op for hugetlb; no free 1G gigantic page).  Both
the VmSwap accounting balance and the hwpoison swapcache interception
are exercised end to end.

Every intermediate patch builds individually with zero warnings; the
!CONFIG_HUGETLB_PAGE configuration builds as well.


Patch split and ordering constraints
====================================

1. folio_swap_entry() helper + tree-wide conversion of all
   folio->swap readers (zero semantic change; the changelog documents
   the union pollution mechanism, the identity argument, the
   alignment invariant, and why the conversion is tree-wide rather
   than reachable-sites-only);
2. Anonymous swap-in (read-back first): hugetlb_anon_swapin() + fault
   recognition of swap entries + fork/zap/swapoff/MM_SWAPENTS
   balancing;
3. Anonymous swap-out core (producer second): the reclaim trio + the
   rmap callback (rmap.h +1 line) + the hugetlb_folio_swap_supported()
   size gate (with the swap_state.c VM_WARN backstop) + the madvise
   MADV_PAGEOUT entry point; the changelog carries the
   mirror/placement rationale and the full size-gate argument.
   The order must not be flipped: the base hugetlb_fault() returns
   ret = 0 for non-migration, non-hwpoison entries, so swap-out first
   means an endless fault livelock on swapped-out pages;
4. memory-failure adaptations -- right after patch 3, because the
   hole becomes reachable from patch 3 on, and a partial merge of
   1-3 must not leave an mf hole behind;
5. File-backed swap as one piece: the shmem-model anchor +
   hugetlbfs_writeout() + the refault/read swap-in pair + the
   truncate/evict lifecycle + swapoff traversal.  Kept together
   because none of the four is correct without the others: read-back
   alone is all dead code; swap-out without lifecycle management
   leaks anchors at truncate; lifecycle without swapoff traversal
   leaves swapoff unable to reach the anchors (file folios install no
   swap PTEs, so the mm walk cannot find them).  Any prefix of it
   would be unhealthy, hence a single patch;
6. Selftest + documentation (hugetlbpage.rst).

Comments welcome.

[1] https://lore.kernel.org/r/20250514201729.48420-1-ryncsn@gmail.com

Zongkun Lei (6):
  mm/swap: introduce folio_swap_entry() and convert all folio->swap
    readers
  mm/hugetlb: swap-in support for anonymous hugetlb folios
  mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT
  mm/memory-failure: handle swapcached hugetlb folios
  mm/hugetlb: swap support for file-backed hugetlb folios
  selftests/mm: add hugetlb_swap test, document hugetlb swap

 Documentation/admin-guide/mm/hugetlbpage.rst |   29 +-
 fs/hugetlbfs/inode.c                         |  128 +-
 fs/proc/task_mmu.c                           |   11 +
 include/linux/hugetlb.h                      |   58 +
 include/linux/pagemap.h                      |    7 +
 include/linux/rmap.h                         |    1 +
 include/linux/swap.h                         |   23 +-
 mm/huge_memory.c                             |    2 +-
 mm/hugetlb.c                                 | 1408 +++++++++++++++++-
 mm/internal.h                                |   19 +-
 mm/madvise.c                                 |   70 +
 mm/memcontrol-v1.c                           |    4 +-
 mm/memcontrol.c                              |    2 +-
 mm/memory-failure.c                          |   19 +
 mm/memory.c                                  |    2 +-
 mm/page_io.c                                 |   18 +-
 mm/rmap.c                                    |  175 ++-
 mm/shmem.c                                   |    6 +-
 mm/swap.h                                    |   14 +-
 mm/swap_state.c                              |   55 +-
 mm/swapfile.c                                |   87 +-
 mm/userfaultfd.c                             |    2 +-
 mm/util.c                                    |    2 +-
 mm/vmscan.c                                  |    4 +-
 mm/zswap.c                                   |    4 +-
 tools/testing/selftests/mm/Makefile          |    1 +
 tools/testing/selftests/mm/hugetlb_swap.c    | 1118 ++++++++++++++
 tools/testing/selftests/mm/run_vmtests.sh    |    2 +
 28 files changed, 3219 insertions(+), 52 deletions(-)
 create mode 100644 tools/testing/selftests/mm/hugetlb_swap.c

-- 
2.53.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-29  9:31 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29  7:53 [RFC PATCH 0/6] mm: swap support for hugetlb folios Zongkun Lei
2026-09-29  8:02 ` David Hildenbrand (Arm)
2026-09-29  9:26   ` Zongkun Lei
2026-09-29  8:18 ` Lorenzo Stoakes (ARM)
2026-09-29  9:30   ` Zongkun Lei

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®