From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Zongkun Lei <leizongkun@qq.com>, linux-mm@kvack.org
Cc: Andrew Morton <akpm@linux-foundation.org>,
Muchun Song <muchun.song@linux.dev>,
Oscar Salvador <osalvador@suse.de>, Peter Xu <peterx@redhat.com>,
Matthew Wilcox <willy@infradead.org>,
Kairui Song <kasong@tencent.com>, Michal Hocko <mhocko@suse.com>,
Johannes Weiner <hannes@cmpxchg.org>,
Hugh Dickins <hughd@google.com>,
Naoya Horiguchi <nao.horiguchi@gmail.com>,
Zi Yan <ziy@nvidia.com>, Lorenzo Stoakes <ljs@kernel.org>,
linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH 0/6] mm: swap support for hugetlb folios
Date: Tue, 29 Sep 2026 10:02:11 +0200 [thread overview]
Message-ID: <d8646877-4975-4b2d-9255-9285414d65cb@kernel.org> (raw)
In-Reply-To: <tencent_501F9D3CED07B66BAA97D5352FB63EE1A709@qq.com>
On 9/29/26 09:53, Zongkun Lei wrote:
> Hi all,
>
> This series introduces *user-driven* swap support for hugetlb folios,
> covering both private anonymous and shared file-backed mappings. It
> deliberately keeps the hugetlb reservation semantics intact: the
> kernel never reclaims hugetlb pages on its own -- hugetlb folios never
> go on an LRU, there is no hot/cold guessing, and nothing hooks into
> kswapd or direct reclaim. This is the same boundary the merged
> hugetlb demotion keeps: the kernel never reclaims reserved memory by
> its own judgement. Demotion is triggered explicitly by the
> administrator (writing nr_hugepages) and only touches free pool
> pages; this series is triggered explicitly by the workload (madvise)
> and acts on its own live mappings. In both cases the decision stays
> outside the kernel.
>
> Interface:
>
> - Swap-out: MADV_PAGEOUT via madvise(2)/process_madvise(2) learns to
> handle hugetlb mappings. For anonymous mappings the PTE becomes a
> swap entry; for shared mappings the PTE is simply cleared and the
> location record stays in the hugetlbfs page cache.
> - Swap-in: a fault reads the page back from swap.
> - Swapoff: all swapped-out hugetlb pages are read back. If the pool
> has no free huge pages at swap-in time, the behaviour matches a
> first-touch fault on a full pool: after bounded retries the process
> receives SIGBUS -- no silent failure, no data loss.
>
>
> Motivation
> ==========
>
> Today a hugetlb reservation is a one-way street: once instantiated, a
> huge page's memory is completely unreclaimable until the mapping is
> torn down. THP and 4K anonymous memory can simply be swapped out;
> hugetlb cannot. This gap forces operators to choose between "huge
> page performance" and "memory overcommit".
>
> VM overcommit. QEMU guests backed by MAP_SHARED hugetlbfs files
> (-mem-path) allocate all guest memory from the host's huge page pool.
> Overcommit means selling more guest memory than the pool holds --
> fine in practice because guests never all peak at once. But the
> extreme case is unavoidable: when every guest's load spikes at the
> same time, the pool runs out. Keeping the VMs alive then requires a
> swap mechanism that moves the memory the provider judges cold out of
> the way. Ballooning cannot save this scenario: it reclaims memory
> that is free *inside* the guest, and the extreme case is precisely
> that all guests are busy and have nothing to give. The extreme case
> is rare, but without an escape hatch providers dare not overcommit
> confidently, and pool utilization stays low. This series is that
> escape hatch: the VM management stack issues MADV_PAGEOUT on cold
> guest memory, VM availability survives the extreme case, and "huge
> page performance" vs "memory overcommit" stops being a choice.
>
> Hot-upgrade standby. A hot upgrade first brings up the new instance
> and switches traffic before the old one may be touched; to keep
> rollback fast, the old instance usually cannot be deleted right away
> and stands by for hours or days. During that time it has no traffic
> and its memory is cold, yet the process and its mappings still pin
> the pool until it is finally removed. With MADV_PAGEOUT the cold
> memory can be swapped out during standby and swapped back if a
> rollback happens. More generally: whenever the pool is full but
> tearing down the mapping is not an option, MADV_PAGEOUT is the only
> way to free space.
>
>
> Design
> ======
>
> Preparation: tree-wide folio_swap_entry() conversion (patch 1)
> --------------------------------------------------------------
>
> Hugetlb flags live in folio->private (a union with folio->swap), so
> while a folio sits in the swap cache the low bits of folio->swap.val
> can be polluted by HPAGEFLAG (with HVO enabled, bit 4 is set on
> virtually every hugetlb folio -- not a theoretical issue). The fix
> is folio_swap_entry() -- an identity transform for non-hugetlb folios
> (round_down(x, 1) == x for order-0; THP low bits are zero anyway) --
> plus a *tree-wide* conversion of all folio->swap readers. What
> review must prove shrinks from "N sites x reachability arguments" to
> "1 helper + 1 invariant (swap entries are folio-size aligned)", with
> zero semantic change to existing paths. It makes *generic* code
> hugetlb-safe instead of adding special cases in hugetlb.c, aligned
> with the hugetlb genericization direction.
>
> The companion size gate hugetlb_folio_swap_supported() (introduced in
> patch 3) clamps both bounds at once: the lower bound (HPAGEFLAG
> width, requiring nr_pages >= 128) and the upper bound (non-gigantic,
> <= PMD size, so one swap cluster holds the folio).
>
> Private anonymous mappings (patches 2-4)
> ----------------------------------------
>
> - Swap-in (consumer first, patch 2): hugetlb_anon_swapin() mirrors
> do_swap_page(): get_swap_device() pins the device,
> swap_cache_get_folio() finds in-flight folios, and on a miss the
> folio is allocated from the hugetlb pool and added to the swap
> cache via the new generic swap_cache_add_folio() helper (hugetlb
> folios do not come from the buddy allocator, so
> __swap_cache_alloc() cannot be used). Pool exhaustion retries
> briefly, then fails with SIGBUS, matching the
> reservation-exhaustion contract of the no-page fault path. The
> lifecycle is fully 4K-semantic: fork shares swap entries
> (copy_hugetlb_page_range() + swap_dup_entries_direct(), dropping
> the exclusive bit); zap/swapin release them; MM_SWAPENTS is
> balanced at swapout/fork/swapin/zap per mainline semantics (VmSwap
> stays meaningful); swapoff drains via hugetlb_unuse_vma().
> The read-back path must merge before the swap-out path: the base
> hugetlb_fault() returns ret = 0 for any entry that is neither
> migration nor hwpoison, so with swap-out first, touching a
> swapped-out page would livelock in an endless fault loop.
>
> - Swap-out (patch 3): hugetlb_reclaim_pages() /
> hugetlb_reclaim_folio_list() are trimmed mirrors of
> reclaim_pages()/shrink_folio_list() and live in hugetlb.c; the
> changelog of patch 3 explains why a mirror and why there. Swap PTE
> installation reuses the rmap walker: a new
> try_to_unmap_swap_hugetlb_one() (a static callback mirroring
> ttu_anon_swapbacked_folio(): single walk, swp_pte_prepare(),
> mm_prepare_for_swap_entries(), MM_SWAPENTS accounting) plus the
> exported wrapper try_to_unmap_swap_hugetlb(), following the
> try_to_unmap_poisoned_hugetlb_one() precedent. Cost to generic
> rmap: one added line in rmap.h, zero changes to existing code.
>
> - memory-failure (patch 4): anonymous hugetlb folios can now sit in
> the swap cache (the writeback window, the cached copy left after
> swap-in) -- a state memory failure has never seen before. Two
> adaptations: try_to_unmap() routes hugetlb folios by TTU_HWPOISON
> (set -> the poisoned handler; cleared -- the mf case that keeps a
> dirty swapcache folio -- -> the swap handler, matching 4K swapcache
> semantics), and me_huge_page() intercepts swapcache folios at the
> top (mirroring me_swapcache_dirty(): keep the folio in the swap
> cache and return MF_DELAYED), fixing the silent eviction of a clean
> poisoned folio that would otherwise mean a 2M pool leak plus a lost
> poison marker.
>
> Shared file-backed mappings (patch 5)
> -------------------------------------
>
> The shmem model: swap-out only clears the PTEs -- no swap entry is
> ever installed for a file folio -- and the anchor (a swap value
> entry, swp_to_radix_entry()) is left in the hugetlbfs page cache,
> holding a dup'ed swap count.
>
> - Swap-out: hugetlbfs_writeout() -- folio_alloc_swap() allocates the
> folio-sized slot range, the inode is linked on the per-inode
> swaplist (for swapoff traversal), folio_dup_swap() holds a swap
> count for the anchor, the page cache slot is replaced by the anchor
> in place, and swap_writeout() writes out. The reclaim loop drops
> its anon-only gate and madvise drops the VM_MAYSHARE rejection.
>
> - Refault swap-in: hugetlb_no_page() recognizes the anchor (the page
> cache lookup reports it as a hole) and swaps the folio back in via
> hugetlbfs_do_swapin(), replacing the anchor in place --
> indistinguishable from a regular page-cache hit afterwards. The
> read-back path must never lag the producer: the base fault path
> treats value entries as holes and would silently zero-fill over
> swapped-out data.
>
> - read(2) swap-in: hugetlbfs_read_iter() resolves anchors through
> hugetlbfs_swapin_read() instead of zero-filling; swapin failures
> surface as -EIO/-EFAULT, never stale data.
>
> - Lifecycle: truncate/hole-punch/evict free anchors and their slots
> via hugetlbfs_free_swap(), with two UAF guards (the swap cache
> entry is removed before the anchor's swap count is dropped; a
> refcount race means reclaim is isolating the folio, so the folio
> lock is held across swap_put_entries_direct() to wait it out).
> Inode eviction drains the swapoff traversal list via
> hugetlbfs_evict_drain_swaplist().
>
> - swapoff: try_to_unuse() hooks hugetlbfs_unuse() right after
> shmem_unuse(). File folios install no swap PTEs, so the mm walk
> cannot reach their slots -- the per-inode swaplist traversal can.
>
> - Symmetric accounting: hugetlb_cgroup, memcg swap/hugetlb charge,
> the global reserve and NR_HUGETLB are paired with free_huge_folio()
> on every path; the vma-less read/swapoff swap-in charges the
> caller's context, the same policy as shmem swapoff.
>
>
> Base and dependencies
> =====================
>
> Based on mm-unstable; no external patch dependencies.
>
> The implementation builds on Kairui Song's swap table work: the
> folio-sized contiguous swap slot ranges (folio_alloc_swap() /
> folio_dup_swap() / folio_put_swap()) are the primitives that make
> this possible at all -- without them a 2M folio would need 512
> independent swap entries. The merged phases of that work are already
> in mm-unstable [1].
>
> Looking ahead, the swap table roadmap may move swap entries out of
> folio->swap; removing folio_swap_entry() then is a mechanical
> operation -- the call sites are already concentrated in a single
> helper, so it will be easier to delete than what we have today.
>
> This series is complementary to the Reserved THP RFC (Qi Zheng), not
> competing with it. Reserved THP explores the endgame -- reserved,
> swappable, unsplittable large folios that may absorb most hugetlb use
> cases. But the LSFMM 2024 consensus is that the hugetlbfs ABI stays
> (migrating away is a 15-20 year scale); as long as the ABI exists,
> its unreclaimable-pool problem needs a direct solution. The two also
> build on the same swap table foundation, so progress on either side
> de-risks the other.
>
>
> Anticipated questions
> =====================
>
> Q1: Why not put hugetlb folios on the LRU and reuse vmscan?
>
> Semantic layer: hugetlb's ABI contract is "reservation = latency
> determinism", and the kernel must not tear that up unilaterally --
> this is the historical reason hugetlb could not be swapped at all.
> Letting the kernel reclaim reserved pages under global memory
> pressure based on its own hot/cold guesses (folio_referenced() style
> access-pattern inference) injects fault latency exactly where it must
> never appear -- the hot-upgrade standby and pre-traffic-peak moments
> are precisely why hugetlb is used. So the split in this series is:
> the kernel provides the mechanism, the policy belongs to userspace.
> MADV_PAGEOUT extends to hugetlb the proactive contract 4K anonymous
> memory already has; the workload (orchestrator, JVM, database) knows
> its warmup windows and traffic peaks, so hot/cold decisions stay on
> the side with the most information. This is the same boundary the
> only merged hugetlb reclaim mechanism, demotion, keeps: the kernel
> never reclaims reserved memory by its own judgement -- demotion is
> administrator-triggered (nr_hugepages) and touches only free pool
> pages, this series is workload-triggered (madvise) and touches its
> own live mappings; in both cases the decision is outside the kernel.
>
> Technical layer: an LRU is not a neutral queue; it is the input
> structure of the kernel's reclaim policy. Once a hugetlb folio is on
> an lruvec, kswapd, direct reclaim, MGLRU generation scanning and
> cgroup v2 memory.reclaim will *all* see it -- there is no "on the LRU
> but exempt from scanning" mode; LRU membership means "in the kernel's
> reclaim view". And making the LRU machinery truly understand hugetlb
> means adapting per-memcg lruvec accounting, folio referenced/aging,
> workingset shadows, MGLRU generations -- which is precisely spraying
> if (hugetlb) special paths across generic reclaim code, in direct
> conflict with the LSFMM 2024 consensus of cleaning up internals and
> removing scattered special cases. A likely suggestion, "on the LRU
> but reachable only by memory.reclaim, not kswapd", fails the same
> way: the LRU is the shared input of every scanner; there is no
> per-consumer visibility.
>
> Implementation layer: under this split the series deliberately takes
> the minimal-invasion route -- pool management stays in hugetlb.c,
> every reusable generic leaf operation (rmap walk, slot allocation,
> writeout, referenced/pin checks) is shared, and only the control flow
> is mirrored, with "keep in sync" notes. Today hugetlb folios live on
> per-hstate lists carrying pool semantics (reserve, surplus, demotion)
> that the LRU machinery knows nothing about; if LRU-ification ever
> happens, the mirrored structure folds over mechanically. kswapd
> integration is an explicit non-goal, not a "never": the design does
> not close the door (the referenced check simply comes back together
> with that series' scan_control), but user-space-driven is the right
> default for hugetlb, in the same direction as the Reserved THP RFC's
> "reserved but reclaimable" roadmap.
>
> Q2: Why extend swap to shared file-backed (hugetlbfs) mappings?
>
> Without it, shared mappings pin the pool absolutely: an instantiated
> shared huge page is unreclaimable until truncate/unlink, however cold
> the data. Deployments mixing private and shared hugetlb get held
> hostage by the shared half -- under pressure the anonymous half can
> be swapped out, yet the pool still exhausts and new allocations
> (including swap-in itself!) keep failing. Concrete users:
>
> - QEMU guests backed by MAP_SHARED hugetlbfs files (-mem-path): idle
> guest memory can finally be reclaimed with MADV_PAGEOUT instead of
> pinning the pool for the VM's entire lifetime;
> - services that hold shared hugetlbfs files for coordination/staging
> but rarely touch them;
> - ABI consistency: MADV_PAGEOUT succeeds on private hugetlb and every
> 4K/THP mapping including shmem -- shared hugetlb was the only
> mapping type returning EINVAL, an ABI asymmetry users trip over;
> - swapoff reachability: swapoff must be able to drain every kind of
> swapped-out page; incomplete coverage leaves swapoff stuck on
> anchors it cannot reach.
>
> The anticipated objection -- "shared huge pages are for
> performance-critical shared data, why swap them out" -- applies
> equally to all shared file pages, yet shmem has always been
> swappable. Everything here is opt-in (proactive reclaim only on
> request): nothing moves unless userspace asks. Pool semantics hold
> end to end: swap-out returns the page to the pool, swap-in competes
> for pool memory again, and reservation exhaustion fails with
> SIGBUS/-EIO exactly like a first-touch fault.
>
>
> Testing
> =======
>
> New selftest tools/testing/selftests/mm/hugetlb_swap.c (~1120 lines,
> 37 assertions, 14 test functions), wired into the mm selftest
> Makefile and run_vmtests.sh:
>
> - test_pageout_basic: anonymous PAGEOUT -> PTE becomes a swap entry
> -> the huge page returns to the pool -> fault swaps back in with
> data intact
> - test_fork_after_pageout: fork shares swap entries, writes COW
> without polluting the parent, parent/child MM_SWAPENTS balance
> - test_munmap_releases_swap: zap releases swap entries, SwapFree
> recovers
> - test_readfault_munmap_drains_swapcache: munmap after a read-fault
> swapin drains the leftover swapcache copy (both the pool page and
> the slots are returned)
> - test_swapoff: swapoff drains swapped-out hugetlb pages
> - test_swapin_pool_exhausted: swap-in on an exhausted pool fails with
> SIGBUS (same semantics as a first fault; no OOM, no silent errors)
> - test_shared_file_swap: shared-mapping pageout (PTE cleared, no swap
> entry) -> both pread and fault swap back in with data intact, and
> the anchor can be re-established repeatedly
> - test_truncate_after_pageout: ftruncate(0) after pageout frees the
> anchor and its slots; subsequent access gets SIGBUS
> - test_reject_paths: VM_LOCKED / userfaultfd-registered / 1G gigantic
> are rejected with EINVAL
> - test_mprotect_mremap_smoke: mprotect/mremap over swap entries
> - test_vmswap_accounting: VmSwap balances over the whole lifecycle:
> +2048 kB at pageout, unchanged for the parent after fork, back to
> zero at swapin/munmap (regression test for the unsigned-negation
> zap accounting fix)
> - test_soft_dirty_swap: soft-dirty bit round-trips between present
> PTEs and swap entries, verified with clear_refs isolation
> - test_uffd_wp_swap: uffd-wp bit set through the swap entry, carried
> back to the present PTE at swapin, write fault delivers the WP
> event
> - test_hwpoison_swapcache: PAGEOUT -> swap-in leaves the folio in the
> swap cache -> MADV_HWPOISON -> the next access must SIGBUS (must
> not silently swap stale disk data back in)
>
> Real-machine results (x86-64, 2M huge pages, QEMU + openEuler 24.03,
> dedicated swap disk): 37 assertions, 35 pass / 0 fail / 2 expected
> skips (mlock is a no-op for hugetlb; no free 1G gigantic page). Both
> the VmSwap accounting balance and the hwpoison swapcache interception
> are exercised end to end.
>
> Every intermediate patch builds individually with zero warnings; the
> !CONFIG_HUGETLB_PAGE configuration builds as well.
>
>
> Patch split and ordering constraints
> ====================================
>
> 1. folio_swap_entry() helper + tree-wide conversion of all
> folio->swap readers (zero semantic change; the changelog documents
> the union pollution mechanism, the identity argument, the
> alignment invariant, and why the conversion is tree-wide rather
> than reachable-sites-only);
> 2. Anonymous swap-in (read-back first): hugetlb_anon_swapin() + fault
> recognition of swap entries + fork/zap/swapoff/MM_SWAPENTS
> balancing;
> 3. Anonymous swap-out core (producer second): the reclaim trio + the
> rmap callback (rmap.h +1 line) + the hugetlb_folio_swap_supported()
> size gate (with the swap_state.c VM_WARN backstop) + the madvise
> MADV_PAGEOUT entry point; the changelog carries the
> mirror/placement rationale and the full size-gate argument.
> The order must not be flipped: the base hugetlb_fault() returns
> ret = 0 for non-migration, non-hwpoison entries, so swap-out first
> means an endless fault livelock on swapped-out pages;
> 4. memory-failure adaptations -- right after patch 3, because the
> hole becomes reachable from patch 3 on, and a partial merge of
> 1-3 must not leave an mf hole behind;
> 5. File-backed swap as one piece: the shmem-model anchor +
> hugetlbfs_writeout() + the refault/read swap-in pair + the
> truncate/evict lifecycle + swapoff traversal. Kept together
> because none of the four is correct without the others: read-back
> alone is all dead code; swap-out without lifecycle management
> leaks anchors at truncate; lifecycle without swapoff traversal
> leaves swapoff unable to reach the anchors (file folios install no
> swap PTEs, so the mm walk cannot find them). Any prefix of it
> would be unhealthy, hence a single patch;
> 6. Selftest + documentation (hugetlbpage.rst).
>
> Comments welcome.
>
> [1] https://lore.kernel.org/r/20250514201729.48420-1-ryncsn@gmail.com
>
> Zongkun Lei (6):
> mm/swap: introduce folio_swap_entry() and convert all folio->swap
> readers
> mm/hugetlb: swap-in support for anonymous hugetlb folios
> mm/hugetlb: swap out anonymous hugetlb folios via MADV_PAGEOUT
> mm/memory-failure: handle swapcached hugetlb folios
> mm/hugetlb: swap support for file-backed hugetlb folios
> selftests/mm: add hugetlb_swap test, document hugetlb swap
>
> Documentation/admin-guide/mm/hugetlbpage.rst | 29 +-
> fs/hugetlbfs/inode.c | 128 +-
> fs/proc/task_mmu.c | 11 +
> include/linux/hugetlb.h | 58 +
> include/linux/pagemap.h | 7 +
> include/linux/rmap.h | 1 +
> include/linux/swap.h | 23 +-
> mm/huge_memory.c | 2 +-
> mm/hugetlb.c | 1408 +++++++++++++++++-
> mm/internal.h | 19 +-
> mm/madvise.c | 70 +
> mm/memcontrol-v1.c | 4 +-
> mm/memcontrol.c | 2 +-
> mm/memory-failure.c | 19 +
> mm/memory.c | 2 +-
> mm/page_io.c | 18 +-
> mm/rmap.c | 175 ++-
> mm/shmem.c | 6 +-
> mm/swap.h | 14 +-
> mm/swap_state.c | 55 +-
> mm/swapfile.c | 87 +-
> mm/userfaultfd.c | 2 +-
> mm/util.c | 2 +-
> mm/vmscan.c | 4 +-
> mm/zswap.c | 4 +-
> tools/testing/selftests/mm/Makefile | 1 +
> tools/testing/selftests/mm/hugetlb_swap.c | 1118 ++++++++++++++
> tools/testing/selftests/mm/run_vmtests.sh | 2 +
> 28 files changed, 3219 insertions(+), 52 deletions(-)
Hugetlb is effectively in a feature freeze. There is no interest in maintaining
even more of the special hugetlb sauce, sorry. :/
--
Cheers,
David
next prev parent reply other threads:[~2026-09-29 8:02 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 7:53 Zongkun Lei
2026-09-29 8:02 ` David Hildenbrand (Arm) [this message]
2026-09-29 9:26 ` Zongkun Lei
2026-09-29 8:18 ` Lorenzo Stoakes (ARM)
2026-09-29 9:30 ` Zongkun Lei
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=d8646877-4975-4b2d-9255-9285414d65cb@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=hannes@cmpxchg.org \
--cc=hughd@google.com \
--cc=kasong@tencent.com \
--cc=leizongkun@qq.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=nao.horiguchi@gmail.com \
--cc=osalvador@suse.de \
--cc=peterx@redhat.com \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®