* [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs
@ 2026-09-14 11:41 Usama Arif
2026-09-14 11:41 ` [PATCH v7 01/29] mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio() Usama Arif
` (10 more replies)
0 siblings, 11 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif
When reclaim swaps out a PMD-mapped anonymous THP today, the PMD is
split into HPAGE_PMD_NR PTE-level swap entries via TTU_SPLIT_HUGE_PMD
before unmap. This series introduces a PMD-level swap entry so the
huge mapping can survive the swap round-trip and do_huge_pmd_swap_page()
can restore the PMD mapping directly on swap-in, without waiting for
khugepaged to collapse the range later.
The PMD swap entry is a compact page-table encoding for HPAGE_PMD_NR
consecutive swap slots. swap_map accounting remains per-slot and is
unchanged. Importantly, a PMD swap entry does not promise that the swap
cache always contains one PMD-sized folio. While the cache is empty or
contains one PMD-sized folio, PMD-level handling can proceed. Once the
cache has split/per-slot state, users either inspect the individual
slots directly (mincore) or split the PMD swap entry and retry through
the PTE path (fault, swapoff, MADV_WILLNEED, UFFDIO_MOVE). MADV_FREE
does not consult the cache at all: it frees a whole PMD swap entry in
place and only splits when the advised range covers part of the PMD.
Likewise, if any slot is still backed by zswap's per-page
store, PMD-order swap-in consumers split and let the PTE path load the
range page by page; an all-on-disk range can still be read back as one
PMD-sized folio.
The series is ordered so every consumer can handle PMD swap entries
before the swap-out producer starts installing them. The swap-out patch
is the last functional change.
Sashiko reviews on intermediate patches:
Because patch 28 is the first producer of a PMD swap entry, the
consumer code added by patches 9-27 is unreachable at the point it is
introduced. Every previous sashiko review has reported that code as broken
on the basis of a state it cannot yet be in; the series is ordered this
way deliberately. See [1].
Notes on zswap:
Native PMD-order zswap load/store is intentionally left for a follow-up.
Alexandre Ghiti is currently working on this.
This series can still preserve PMD swap entries while zswap is enabled:
zswap stores the THP as order-0 entries, and PMD-order swap-in
consumers split any range that has zswap entries before reading it. If
zswap has written the whole range back to disk, or the swap cache still
contains one PMD-sized folio, PMD-level handling can proceed.
Testing:
The 17 pmd_swap selftests pass on x86_64 with zswap both disabled and
enabled. PMD_SWAP_DEVICE was the sole active swap device, so the
swapoff test ran in both configurations. Note that with zswap enabled
the range may legitimately come back through the PTE fallback, so those
runs skip the PMD-restoration assertions and say so in the log; the
zswap-disabled run is the one that proves PMD restoration.
[1] https://lore.kernel.org/all/282ef982-3e48-4283-9155-73a33fc1c4e8@linux.dev/
v6 -> v7: https://lore.kernel.org/all/20260818131202.494754-1-usama.arif@linux.dev/
- Rebase onto akpm/mm-new at baa8de2f3448 and adapt to the new
get_swap_device() contract and linear_anon_page_index(). The series grows
from 12 to 29 patches.
- Patch 1: code unchanged; reword the commit message and collect review tags.
- Patches 2-9: split the six architecture helpers from generic detection and
add debug_vm_pgtable coverage. Use the s390 RSTE exclusive bit, make the
x86 helper return bool, clear all PMD swap overlay bits before softleaf
decoding, and add the architecture maintainers to Cc.
- Patches 10-11: add a preparatory no-functional-change cleanup of the
migration-splitting API and keep the PMD swap split separate.
- Patch 12: reject a multi-slot duplication range that crosses a swap-cluster
boundary.
- Patch 13: drop the zswap_load() change, now upstream as 1a904e0d3c43, and
check for zswap after swap-cache insertion while all slots are pinned.
- Patch 14: scan every subpage for hardware poison, discard a failed or newly
allocated unmapped PMD-sized folio before PTE fallback, and recheck the PMD
before splitting it.
- Patches 15-23: split the non-present PMD walkers by subsystem. Add the
guard-advice patch so MADV_GUARD_INSTALL/REMOVE leaves PMD swap entries
whole; the other split patches preserve v6 behaviour.
- Patch 24: honour the current THP policy, discard a failed PMD-sized folio,
and recheck the PMD before PTE fallback.
- Patch 25: do not split after cached-folio revalidation loses a race; retry
so a restored present THP remains whole.
- Patches 26-27: separate the independent PTE-batching hardware-poison fix,
scan subpages directly, and discard failed or never-mapped PMD-sized folios
before PTE retry instead of making the whole range fail with SIGBUS.
- Patch 28: keep the normal swap-out path unchanged, but make failed producer
preconditions warn and return -EBUSY rather than BUG or return -EINVAL.
- Patch 29: grow the selftests from 16 to 17, adding
swapin_sync/cache-residency coverage and more robust feature, fallback,
privilege, data-integrity, and swap-device-priority handling.
v5 -> v6: https://lore.kernel.org/all/20260722152043.2273289-1-usama.arif@linux.dev/
- Add patch 1 to rename pmd_to_softleaf_folio() to
pmd_softleaf_to_folio(). No functional change. (Dev Jain)
- Patch 2: warn when pmd_softleaf_to_folio() is given a non-PFN
softleaf rather than silently returning NULL. (Dev Jain)
- Patch 4: bound the fork extend-table fallback to one retry, re-read
the PMD under its lock, normalize unrecoverable copy_huge_pmd() errors
to -ENOMEM so copy_pmd_range() cannot clear and leak the source swap
PMD, and drop a redundant thp_migration_supported() gate.
- Patch 5: check multi-page swap-cache insertions for zswap-backed slots
in __swap_cache_add_check() under the cluster lock, both before
allocation and before insertion, and reject mixed zswap/disk state
with -EBUSY. (Yosry Ahmed, Nhat Pham)
- Patch 6: on a failed non-uptodate PMD-order read, remove the large
folio from swap cache before splitting so order-0
fallback retries individual slots rather than poisoning the whole
2 MiB range; retain hardware-poisoned folios for per-subpage handling.
- Patch 7: make HMM snapshot mode report a PMD swap entry as non-resident,
matching PTE swap entries, rather than HMM_PFN_ERROR. Drop redundant
thp_migration_supported() gates and simplify non-present PMD handling.
- Patch 8: factor PMD MADV_WILLNEED prefetch into
swapin_pmd_swap_entry(), split and retry through PTEs after any
PMD-order swapin failure, and replace the racy folio_test_locked()
plus folio_lock() sequence with folio_trylock().
- Patch 9: guard PMD-swap UFFDIO_MOVE code with CONFIG_THP_SWAP, clarify
RWP marker propagation, and reject a PMD swap entry at the destination
with -EEXIST so UFFDIO_MOVE cannot loop forever on -EAGAIN.
- Patch 10: honor current THP/VMA policy before PMD-order swap-in, recheck
that the PMD is still the original swap entry before splitting for PTE
fallback, and provide the CONFIG_TRANSPARENT_HUGEPAGE wp_huge_pmd()
declaration/stub needed by THP=n builds.
- Patch 11: make an invalid set_pmd_swap_entry() walk context warn and
return -EINVAL instead of falsely reporting success and corrupting the
MM_ANONPAGES/MM_SWAPENTS accounting, and add an exact PMD-size folio
precondition check. (Luiz Capitulino)
- Patch 12: use /proc/swaps for prerequisite detection, check
MADV_HUGEPAGE, and distinguish an environment that cannot allocate a
PMD THP (SKIP) from a swap-out validation failure (FAIL). Add
partial-mprotect and partial-munmap split coverage. Strengthen
munmap/MADV_FREE VmSwap accounting, pagemap slot-offset checks, and
mprotect/mremap swapped-state checks; force mremap to move, check
munmap()'s return, and mark the UFFDIO_MOVE destination MADV_HUGEPAGE
before asserting PMD restoration. Move common setup and cleanup into
one fixture, merge the swapoff fixture, remove the redundant cycles
test, and make the data pattern differ between base pages so the split
tests can detect incorrect slot ordering. Order fork-COW so the parent
writes while the child still holds the untouched shared swap entry.
(Luiz Capitulino)
- Clarify commit messages throughout. Retain TTU_SPLIT_HUGE_PMD after
prototyping its removal: removing it here requires an extra rmap walk
and broadens the series beyond PMD swap entries. (Matthew Wilcox)
- Rebase onto akpm/mm-new from 15 August (4b65683fd25f).
v4 -> v5: https://lore.kernel.org/all/20260713133613.2707815-1-usama.arif@linux.dev/
- Commit message improvements for almost all patches (Yosry for zswap patch)
- Patch 1: make pmd_to_softleaf_folio() reject softleaf entries that do
not encode a PFN, so a PMD swap offset is never interpreted as one.
PMD swap entries remain valid softleaf entries for classification.
(sashiko)
- Patch 2: use the existing pmd_swp_uffd() helper and force freeze=false
for PMD swap entries, which have no struct page for the migration-entry
freeze path. (sashiko)
- Patch 3: document that the caller's page-table or swap-cache reference
pins every source slot while a partial PMD-sized duplication is rolled
back. Keep the pre-existing PTE fork retry behavior outside this
series. (sashiko)
- Patch 5: split to the PTE path rather than mapping a PMD-sized folio
containing a hardware-poisoned subpage, and restore PAGE_NONE when
swapoff restores a UFFD marker in an RWP VMA. (sashiko)
- Patch 6: account SwapPss for a PMD swap entry one slot at a time because
the slots can have different swap reference counts.
- Patch 7: if PMD-order MADV_WILLNEED encounters newly populated per-page
zswap state, revalidate and remove the failed clean PMD-sized cache
folio before retrying through PTEs.
- Patch 8: mark a moved PMD swap entry for UFFD when the UFFDIO_MOVE
destination VMA is RWP-registered. (sashiko)
- Patch 9: restore PAGE_NONE for UFFD RWP swap-in, preserve the original
write-fault state through swap-slot release and COW handling, remove
the unnecessary LRU drain, and prevent PTE batching from mapping a
poisoned subpage. (sashiko)
- Patch 10: add and document thp_swpout_pmd, which counts PMD mappings
replaced by PMD-level swap entries rather than swapped folios.
- Patch 11: register pmd_swap with the default mm selftest runner, preserve
errno across UFFDIO_MOVE cleanup, check swapoff residency before the
first memory access, add a parent-side write and verification to the
fork+COW test, and add RWP regression coverage for swap-in, UFFDIO_MOVE,
and swapoff. (sashiko)
- Keep do_huge_pmd_swap_page() in patch 9. Patches 6 and 7 only add
consumers; patch 10 remains the first producer, so no PMD swap entry
can reach those paths before the fault handler is present. (sashiko)
- Rebase onto latest akpm/mm-new from 22 July (5e0603ba185a)
v3 -> v4: https://lore.kernel.org/all/20260703173903.3789516-1-usama.arif@linux.dev/
- Patch 1: guard the new arch-specific pmd_swp_mkexclusive /
pmd_swp_exclusive / pmd_swp_clear_exclusive helpers on arm64,
loongarch, powerpc, riscv, s390, and x86 with
CONFIG_ARCH_HAS_PMD_SOFTLEAVES, matching the pattern already
used for pmd_swp_soft_dirty. Also fixes the redefinition-vs-
generic-fallback build errors kernel test robot reported on
i386-allnoconfig-bpf and riscv-allnoconfig-bpf, and rewraps
the patch 1 commit message paragraphs to ~75 columns.
(sashiko, kernel test robot, Usama Arif)
- Patch 2: switch the trailing folio_remove_rmap_pmd() gate in
__split_huge_pmd_locked() from *pmd to old_pmd, old_pmd retains
the original present-or-non-present classification for every
branch above. (sashiko)
- Patch 3: teach swap_retry_table_alloc() (and the underlying
swap_extend_table_alloc()) to accept an nr parameter and scan
every slot in [ci_off, ci_off + nr) before committing an
extend-table allocation. (sashiko)
- Patch 4: rename zswap_range_has_entry() to zswap_is_present() so
the same helper serves both single-slot (nr=1) and range queries,
and switch the implementation from XA_STATE + xas_find() to
xa_find(), which handles RCU locking and internal-retry markers
itself. Rename the callers in patches 5, 7, 9. (Yosry)
- Patch 9: refuse to map a swap-cache folio in do_huge_pmd_swap_page()
when folio_contain_hwpoisoned_page() reports a poisoned subpage;
split the PMD swap entry so do_swap_page() can return
VM_FAULT_HWPOISON per subpage instead of the PMD handler mapping
the corrupted memory as one THP. Mirrors the PageHWPoison check
the PTE swap-in path already performs. (sashiko)
- Patch 9: note explicitly in the commit message that PMD-order
swap-in deliberately skips the order-0 readahead paths, order-0
readahead would populate per-page swap-cache state and force the
PMD swap entry to split before the fault could finish. (Kairui)
- Patch 10: move mm_prepare_for_swap_entries() into
set_pmd_swap_entry() between folio_dup_swap() and set_pmd_at()
so this mm is on init_mm.mmlist before any swap PMD referencing
slots with a non-zero swap_map becomes visible. Matches the PTE
swap-out ordering. (sashiko)
- rebase onto latest akpm/mm-new (61cccb8363fcc282d4ae0555b8739dd227f5ad0b)
v2 -> v3: https://lore.kernel.org/all/20260602142537.198755-1-usama.arif@linux.dev/
- Clarified the PMD swap entry rule: it is a compact encoding for
HPAGE_PMD_NR swap slots, not a guarantee that swap cache always has
one PMD-sized folio. (Lance Yang)
- Swapoff, fault, MADV_WILLNEED, and UFFDIO_MOVE now classify the
whole PMD swap-cache range and split/retry through the PTE path for
split/per-slot cache state. (Lance Yang)
- mincore handles PMD swap entries without assuming one lookup covers
a split swap-cache range. (Lance Yang)
- UFFDIO_MOVE rechecks all HPAGE_PMD_NR slots before moving an empty
PMD swap-cache range, avoiding stale rmap metadata for per-slot
cached folios.
- Added a standalone zswap prerequisite patch from Alexandre that
distinguishes all-on-disk large-folio ranges from ranges with
per-page zswap entries.
- Replaced the global zswap-ever-enabled policy with per-range zswap
checks: PMD swap entries can still be installed while zswap is
enabled, and PMD-order swap-in consumers split when the range has
per-page zswap state.
- Added a mincore selftest and updated MADV_WILLNEED coverage so the
test checks that the PMD swap entry remains in place until first
touch. Total pmd_swap coverage is now 14 tests.
v1 -> v2: https://lore.kernel.org/all/20260427100553.2754667-1-usama.arif@linux.dev/
- Patch 1: convert two additional softleaf_to_pmd() callers that
landed in mm-unstable since v1 (mm/debug_vm_pgtable.c,
mm/migrate_device.c) (Dev)
- Patch 2: rename helper ensure_on_mmlist() to
mm_prepare_for_swap_entries() to better describe its purpose
(David)
- Patch 3: drop VM_WARN_ON_ONCE(!pmd_is_migration_entry) as
Dev posted it as a separate patch.
- Patch 5 (new): move softleaf_to_folio() inside the device-private
branch in migrate_vma_collect_pmd(); same class of fix as patch 4
but for the migrate-device PMD walker.
- Patch 6 (new): rename CONFIG_ARCH_ENABLE_THP_MIGRATION to
CONFIG_ARCH_HAS_PMD_SOFTLEAVES so the gate that now drives
swap-entry support too is named for what it actually controls
(PMD softleaf entries), not just migration. (Dev)
- Patch 7: add the missing pmd_swp_exclusive / mkexclusive /
clear_exclusive helpers for powerpc.
- Patches 10 and 14: use upstream swapin_sync() (bundles
swap_cache_alloc_folio + swap_read_folio + the -EEXIST race
retry) instead of the bespoke swapin_alloc_pmd_folio() helper
from v1; do_swap_page and shmem_swapin_folio use the same
helper (Kairui)
- Patch 10: construct a stack vm_fault for the swapoff swap-in so
the allocator can resolve a mempolicy, mirroring how the PTE
swapoff path (unuse_pte_range) already does it.
- Patch 11: extend coverage to check_pmd_state() in khugepaged so a
swapped-out PMD-mapped THP is treated as SCAN_PMD_MAPPED (matches
the existing migration-entry handling). Route PMD swap entries in the
pmd_trans_huge_lock() branch of mincore_pte_range() through
mincore_pmd_swap() so a swapped-out PMD-mapped THP isn't reported as
resident.
- Patch 12 (new): handle PMD swap entries in MADV_WILLNEED via
swapin_sync(BIT(HPAGE_PMD_ORDER)); a naive order-0 read-ahead
would force the subsequent fault to split.
- Patch 13: refuse UFFDIO_MOVE with -EBUSY if the swap-cache folio
was split between swap-out and the move, matching
move_pages_pte()'s rejection of large folios; otherwise only one
of the 512 anon-rmaps would be re-anchored to dst_vma.
- Patch 16: alloc_fill_swap_thp() now uses the existing
mmap_pmd_aligned() helper so tests don't flake/skip based on VA
placement; new MADV_WILLNEED test that watches the PMD-order
mTHP swpin counter; swapoff test restructured to use the
kselftest_harness ASSERT cleanup blocks (no double swapoff, no
verify-after-munmap).
- Collected Acks and Reviews.
Usama Arif (29):
mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio()
arm64: mm: add PMD swap-exclusive helpers
loongarch: mm: add PMD swap-exclusive helpers
powerpc: mm: add PMD swap-exclusive helpers
riscv: mm: add PMD swap-exclusive helpers
s390: mm: add PMD swap-exclusive helpers
x86: mm: add PMD swap-exclusive helpers
mm: recognize PMD swap entries in the softleaf layer
mm/debug_vm_pgtable: test PMD swap-exclusive helpers
mm: make PMD migration-entry splitting explicit
mm: split PMD swap entries into PTE swap entries
mm: handle PMD swap entries in fork path
mm: zswap: reject high-order swap cache allocations backed by zswap
mm: swap in PMD swap entries as whole THPs during swapoff
fs/proc: account PMD swap entries in smaps
mm: handle soft-dirty and uffd-wp on PMD swap entries
mm/hmm: fault PMD swap entries on demand
mm: free PMD swap entries in zap_huge_pmd()
mm/madvise: free PMD swap entries with MADV_FREE
mm/madvise: skip PMD swap entries for MADV_COLD and MADV_PAGEOUT
mm/madvise: keep PMD swap entries whole for MADV_GUARD_INSTALL/REMOVE
mm/mincore: report PMD swap-cache residency
mm/khugepaged: treat PMD swap entries as mapped THPs
mm: handle PMD swap entries in MADV_WILLNEED
mm: handle PMD swap entries in UFFDIO_MOVE
mm: don't PTE-batch a swap-in over a hardware-poisoned subpage
mm: handle PMD swap entry faults on swap-in
mm: install PMD swap entries on swap-out
selftests/mm: add PMD swap entry tests
Documentation/admin-guide/mm/transhuge.rst | 5 +
arch/arm64/include/asm/pgtable.h | 6 +
arch/loongarch/include/asm/pgtable.h | 19 +
arch/powerpc/include/asm/book3s/64/pgtable.h | 17 +
arch/riscv/include/asm/pgtable.h | 15 +
arch/s390/include/asm/pgtable.h | 28 +-
arch/x86/include/asm/pgtable.h | 17 +
fs/proc/task_mmu.c | 46 +-
include/linux/huge_mm.h | 38 +-
include/linux/leafops.h | 44 +-
include/linux/pgtable.h | 17 +
include/linux/swap.h | 4 +-
include/linux/vm_event_item.h | 1 +
include/linux/zswap.h | 6 +
mm/debug_vm_pgtable.c | 40 +
mm/hmm.c | 11 +-
mm/huge_memory.c | 717 +++++++++++++-
mm/internal.h | 58 ++
mm/khugepaged.c | 6 +
mm/madvise.c | 169 +++-
mm/memory.c | 57 +-
mm/migrate_device.c | 7 +-
mm/mincore.c | 47 +-
mm/mprotect.c | 2 +-
mm/rmap.c | 26 +-
mm/swap.h | 22 +-
mm/swap_state.c | 83 +-
mm/swapfile.c | 247 ++++-
mm/userfaultfd.c | 14 +
mm/vmscan.c | 9 +-
mm/vmstat.c | 1 +
mm/zswap.c | 12 +-
tools/testing/selftests/mm/Makefile | 2 +
tools/testing/selftests/mm/ksft_pmd_swap.sh | 4 +
tools/testing/selftests/mm/pmd_swap.c | 989 +++++++++++++++++++
tools/testing/selftests/mm/run_vmtests.sh | 4 +
tools/testing/selftests/mm/vm_util.c | 24 +
tools/testing/selftests/mm/vm_util.h | 2 +
38 files changed, 2636 insertions(+), 180 deletions(-)
create mode 100755 tools/testing/selftests/mm/ksft_pmd_swap.sh
create mode 100644 tools/testing/selftests/mm/pmd_swap.c
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 01/29] mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio()
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
@ 2026-09-14 11:41 ` Usama Arif
2026-09-14 11:41 ` [PATCH v7 02/29] arm64: mm: add PMD swap-exclusive helpers Usama Arif
` (9 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif
pmd_to_softleaf_folio() reads as if it converted a PMD into a folio. What
it does is decode the softleaf entry stored in the PMD and return the
folio that entry references - the direction softleaf_to_folio() already
spells out.
No functional change intended.
Suggested-by: Dev Jain <dev.jain@arm.com>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Reviewed-by: Zi Yan <ziy@nvidia.com>
Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
---
include/linux/leafops.h | 4 ++--
mm/huge_memory.c | 2 +-
2 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/include/linux/leafops.h b/include/linux/leafops.h
index 4c1476ae32343..7c13c58a5e218 100644
--- a/include/linux/leafops.h
+++ b/include/linux/leafops.h
@@ -657,7 +657,7 @@ static inline bool pmd_is_valid_softleaf(pmd_t pmd)
}
/**
- * pmd_to_softleaf_folio() - Convert the PMD entry to a folio.
+ * pmd_softleaf_to_folio() - Convert the PMD softleaf entry to a folio.
* @pmd: PMD entry.
*
* The PMD entry is expected to be a valid PMD softleaf entry.
@@ -665,7 +665,7 @@ static inline bool pmd_is_valid_softleaf(pmd_t pmd)
* Returns: the folio the softleaf entry references if this is a valid softleaf
* entry, otherwise NULL.
*/
-static inline struct folio *pmd_to_softleaf_folio(pmd_t pmd)
+static inline struct folio *pmd_softleaf_to_folio(pmd_t pmd)
{
const softleaf_t entry = softleaf_from_pmd(pmd);
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 7140a1031fb2e..ee8d46827ffdc 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -2518,7 +2518,7 @@ static struct folio *normal_or_softleaf_folio_pmd(struct vm_area_struct *vma,
if (!thp_migration_supported())
WARN_ONCE(1, "Non present huge pmd without pmd migration enabled!");
- return pmd_to_softleaf_folio(pmdval);
+ return pmd_softleaf_to_folio(pmdval);
}
static bool has_deposited_pgtable(struct vm_area_struct *vma, pmd_t pmdval,
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 02/29] arm64: mm: add PMD swap-exclusive helpers
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
2026-09-14 11:41 ` [PATCH v7 01/29] mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio() Usama Arif
@ 2026-09-14 11:41 ` Usama Arif
2026-09-14 11:41 ` [PATCH v7 03/29] loongarch: " Usama Arif
` (8 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif, Catalin Marinas, Will Deacon
A later patch keeps a PMD-mapped anonymous THP mapped by a PMD across the
swap round-trip, so PG_anon_exclusive now has to survive in a swap PMD and
not just in a swap PTE.
arm64 encodes a swap PMD exactly like a swap PTE, so the new helpers wrap
the PTE ones and reuse PTE_SWP_EXCLUSIVE.
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Will Deacon <will@kernel.org>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
arch/arm64/include/asm/pgtable.h | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/arch/arm64/include/asm/pgtable.h b/arch/arm64/include/asm/pgtable.h
index e89ec5f4787b4..d3f53a601aed3 100644
--- a/arch/arm64/include/asm/pgtable.h
+++ b/arch/arm64/include/asm/pgtable.h
@@ -599,6 +599,12 @@ static inline int pmd_protnone(pmd_t pmd)
#define pmd_swp_clear_uffd(pmd) \
pte_pmd(pte_swp_clear_uffd(pmd_pte(pmd)))
#endif /* CONFIG_HAVE_ARCH_USERFAULTFD_WP */
+#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+#define pmd_swp_exclusive(pmd) pte_swp_exclusive(pmd_pte(pmd))
+#define pmd_swp_mkexclusive(pmd) pte_pmd(pte_swp_mkexclusive(pmd_pte(pmd)))
+#define pmd_swp_clear_exclusive(pmd) \
+ pte_pmd(pte_swp_clear_exclusive(pmd_pte(pmd)))
+#endif
#define pmd_write(pmd) pte_write(pmd_pte(pmd))
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 03/29] loongarch: mm: add PMD swap-exclusive helpers
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
2026-09-14 11:41 ` [PATCH v7 01/29] mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio() Usama Arif
2026-09-14 11:41 ` [PATCH v7 02/29] arm64: mm: add PMD swap-exclusive helpers Usama Arif
@ 2026-09-14 11:41 ` Usama Arif
2026-09-14 11:41 ` [PATCH v7 04/29] powerpc: " Usama Arif
` (7 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif, Huacai Chen
A later patch keeps a PMD-mapped anonymous THP mapped by a PMD across the
swap round-trip, so PG_anon_exclusive now has to survive in a swap PMD and
not just in a swap PTE.
A LoongArch swap PMD is the swap PTE value plus _PAGE_HUGE, and
_PAGE_SWP_EXCLUSIVE sits outside both the type and the offset field, so the
PMD helpers can use the same bit.
Cc: Huacai Chen <chenhuacai@kernel.org>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
arch/loongarch/include/asm/pgtable.h | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
diff --git a/arch/loongarch/include/asm/pgtable.h b/arch/loongarch/include/asm/pgtable.h
index cf29a4c8ac593..87fecc3a51001 100644
--- a/arch/loongarch/include/asm/pgtable.h
+++ b/arch/loongarch/include/asm/pgtable.h
@@ -351,6 +351,25 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte)
return pte;
}
+#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd)
+{
+ pmd_val(pmd) |= _PAGE_SWP_EXCLUSIVE;
+ return pmd;
+}
+
+static inline bool pmd_swp_exclusive(pmd_t pmd)
+{
+ return pmd_val(pmd) & _PAGE_SWP_EXCLUSIVE;
+}
+
+static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd)
+{
+ pmd_val(pmd) &= ~_PAGE_SWP_EXCLUSIVE;
+ return pmd;
+}
+#endif
+
#define pte_none(pte) (!(pte_val(pte) & ~_PAGE_GLOBAL))
#define pte_present(pte) (pte_val(pte) & (_PAGE_PRESENT | _PAGE_PROTNONE))
#define pte_no_exec(pte) (pte_val(pte) & _PAGE_NO_EXEC)
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 04/29] powerpc: mm: add PMD swap-exclusive helpers
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (2 preceding siblings ...)
2026-09-14 11:41 ` [PATCH v7 03/29] loongarch: " Usama Arif
@ 2026-09-14 11:41 ` Usama Arif
2026-09-14 11:41 ` [PATCH v7 05/29] riscv: " Usama Arif
` (6 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif, Madhavan Srinivasan
A later patch keeps a PMD-mapped anonymous THP mapped by a PMD across the
swap round-trip, so PG_anon_exclusive now has to survive in a swap PMD and
not just in a swap PTE.
book3s64 builds a swap PMD by running the PTE encoding over pmd_pte(), so
the PMD helpers use the same _PAGE_SWP_EXCLUSIVE bit. It is also the only
powerpc variant that selects ARCH_HAS_PMD_SOFTLEAVES, via PPC_THP.
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
arch/powerpc/include/asm/book3s/64/pgtable.h | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/arch/powerpc/include/asm/book3s/64/pgtable.h b/arch/powerpc/include/asm/book3s/64/pgtable.h
index dff8790a047db..28943ef3c1c80 100644
--- a/arch/powerpc/include/asm/book3s/64/pgtable.h
+++ b/arch/powerpc/include/asm/book3s/64/pgtable.h
@@ -699,6 +699,23 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte)
return __pte_raw(pte_raw(pte) & cpu_to_be64(~_PAGE_SWP_EXCLUSIVE));
}
+#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd)
+{
+ return __pmd_raw(pmd_raw(pmd) | cpu_to_be64(_PAGE_SWP_EXCLUSIVE));
+}
+
+static inline bool pmd_swp_exclusive(pmd_t pmd)
+{
+ return !!(pmd_raw(pmd) & cpu_to_be64(_PAGE_SWP_EXCLUSIVE));
+}
+
+static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd)
+{
+ return __pmd_raw(pmd_raw(pmd) & cpu_to_be64(~_PAGE_SWP_EXCLUSIVE));
+}
+#endif
+
static inline bool check_pte_access(unsigned long access, unsigned long ptev)
{
/*
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 05/29] riscv: mm: add PMD swap-exclusive helpers
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (3 preceding siblings ...)
2026-09-14 11:41 ` [PATCH v7 04/29] powerpc: " Usama Arif
@ 2026-09-14 11:41 ` Usama Arif
2026-09-14 11:41 ` [PATCH v7 06/29] s390: " Usama Arif
` (5 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif, Paul Walmsley, Palmer Dabbelt,
Albert Ou
A later patch keeps a PMD-mapped anonymous THP mapped by a PMD across the
swap round-trip, so PG_anon_exclusive now has to survive in a swap PMD and
not just in a swap PTE.
riscv encodes a swap PMD exactly like a swap PTE, so the new helpers wrap
the PTE ones and reuse _PAGE_SWP_EXCLUSIVE.
Cc: Paul Walmsley <pjw@kernel.org>
Cc: Palmer Dabbelt <palmer@dabbelt.com>
Cc: Albert Ou <aou@eecs.berkeley.edu>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
arch/riscv/include/asm/pgtable.h | 15 +++++++++++++++
1 file changed, 15 insertions(+)
diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
index d48f90140841e..b644db16bda94 100644
--- a/arch/riscv/include/asm/pgtable.h
+++ b/arch/riscv/include/asm/pgtable.h
@@ -1219,6 +1219,21 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte)
}
#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+static inline bool pmd_swp_exclusive(pmd_t pmd)
+{
+ return pte_swp_exclusive(pmd_pte(pmd));
+}
+
+static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd)
+{
+ return pte_pmd(pte_swp_mkexclusive(pmd_pte(pmd)));
+}
+
+static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd)
+{
+ return pte_pmd(pte_swp_clear_exclusive(pmd_pte(pmd)));
+}
+
#define __pmd_to_swp_entry(pmd) ((swp_entry_t) { pmd_val(pmd) })
#define __swp_entry_to_pmd(swp) __pmd((swp).val)
#endif /* CONFIG_ARCH_HAS_PMD_SOFTLEAVES */
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 06/29] s390: mm: add PMD swap-exclusive helpers
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (4 preceding siblings ...)
2026-09-14 11:41 ` [PATCH v7 05/29] riscv: " Usama Arif
@ 2026-09-14 11:41 ` Usama Arif
2026-09-14 11:41 ` [PATCH v7 07/29] x86: " Usama Arif
` (4 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif, Alexander Gordeev,
Gerald Schaefer, Heiko Carstens, Vasily Gorbik
A later patch keeps a PMD-mapped anonymous THP mapped by a PMD across the
swap round-trip, so PG_anon_exclusive now has to survive in a swap PMD and
not just in a swap PTE.
s390 is the one architecture where a swap PMD is not a swap PTE in
disguise: it is an RSTE with its own layout, converted to a fake PTE swap
entry for the common code. Give it its own exclusive bit rather than
borrowing the PTE-format macro. The two happen to have the same value, but
that is a coincidence. Bit 52 was documented as unused; document what it
is now.
Cc: Alexander Gordeev <agordeev@linux.ibm.com>
Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
Cc: Heiko Carstens <hca@linux.ibm.com>
Cc: Vasily Gorbik <gor@linux.ibm.com>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
arch/s390/include/asm/pgtable.h | 28 ++++++++++++++++++++++++++--
1 file changed, 26 insertions(+), 2 deletions(-)
diff --git a/arch/s390/include/asm/pgtable.h b/arch/s390/include/asm/pgtable.h
index 2d5c2ab06de98..0790a0884cfab 100644
--- a/arch/s390/include/asm/pgtable.h
+++ b/arch/s390/include/asm/pgtable.h
@@ -333,6 +333,7 @@ void setup_protection_map(void);
/* Common bits in region and segment table entries, for swap entries */
#define _RST_ENTRY_COMM 0x0010 /* Common-Region/Segment, marks swap entry */
#define _RST_ENTRY_INVALID 0x0020 /* invalid region/segment table entry */
+#define _RST_ENTRY_SWP_EXCLUSIVE 0x0800 /* SW exclusive swap bit, see mk_swap_rste() */
#define _CRST_ENTRIES 2048 /* number of region/segment table entries */
#define _PAGE_ENTRIES 256 /* number of page table entries */
@@ -859,6 +860,28 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte)
return clear_pte_bit(pte, __pgprot(_PAGE_SWP_EXCLUSIVE));
}
+#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+/*
+ * A PMD swap entry is an RSTE, not a PTE, so it needs its own exclusive bit
+ * rather than the PTE-format _PAGE_SWP_EXCLUSIVE. The two happen to have the
+ * same value; see the RSTE swap layout above mk_swap_rste().
+ */
+static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd)
+{
+ return set_pmd_bit(pmd, __pgprot(_RST_ENTRY_SWP_EXCLUSIVE));
+}
+
+static inline bool pmd_swp_exclusive(pmd_t pmd)
+{
+ return pmd_val(pmd) & _RST_ENTRY_SWP_EXCLUSIVE;
+}
+
+static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd)
+{
+ return clear_pmd_bit(pmd, __pgprot(_RST_ENTRY_SWP_EXCLUSIVE));
+}
+#endif
+
static inline int pte_soft_dirty(pte_t pte)
{
return pte_val(pte) & _PAGE_SOFT_DIRTY;
@@ -1900,15 +1923,16 @@ static inline swp_entry_t __swp_entry(unsigned long type, unsigned long offset)
* Bits 59 and 63 are used to indicate the swap entry. Bit 58 marks the rste
* as invalid.
* A swap entry is indicated by bit pattern (rste & 0x011) == 0x010
- * | offset |Xtype |11TT|S0|
+ * | offset |Etype |11TT|S0|
* |0000000000111111111122222222223333333333444444444455|555555|5566|66|
* |0123456789012345678901234567890123456789012345678901|234567|8901|23|
*
* Bits 0-51 store the offset.
+ * Bit 52 (E) is used to remember PG_anon_exclusive
+ * (_RST_ENTRY_SWP_EXCLUSIVE), mirroring bit 52 of a swap pte.
* Bits 53-57 store the type.
* Bit 62 (S) is used for softdirty tracking.
* Bits 60-61 (TT) indicate the table type: 0x01 for REGION3 and 0x00 for SEGMENT.
- * Bit 52 (X) is unused.
*/
#define __SWP_OFFSET_MASK_RSTE ((1UL << 52) - 1)
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 07/29] x86: mm: add PMD swap-exclusive helpers
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (5 preceding siblings ...)
2026-09-14 11:41 ` [PATCH v7 06/29] s390: " Usama Arif
@ 2026-09-14 11:41 ` Usama Arif
2026-09-14 11:42 ` [PATCH v7 08/29] mm: recognize PMD swap entries in the softleaf layer Usama Arif
` (3 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:41 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif, Thomas Gleixner, Ingo Molnar,
Borislav Petkov, Dave Hansen, x86
A later patch keeps a PMD-mapped anonymous THP mapped by a PMD across the
swap round-trip, so PG_anon_exclusive now has to survive in a swap PMD and
not just in a swap PTE.
x86-64 encodes a swap PMD exactly like a swap PTE, so the new helpers reuse
_PAGE_SWP_EXCLUSIVE, bit 3, which the swap-entry layout already reserves
for PG_anon_exclusive. 32-bit x86 aliases that bit to _PAGE_PSE and does
not select ARCH_HAS_PMD_SOFTLEAVES.
Cc: Thomas Gleixner <tglx@kernel.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: x86@kernel.org
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
arch/x86/include/asm/pgtable.h | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/arch/x86/include/asm/pgtable.h b/arch/x86/include/asm/pgtable.h
index d551120a7c889..a2d1cd03cba23 100644
--- a/arch/x86/include/asm/pgtable.h
+++ b/arch/x86/include/asm/pgtable.h
@@ -1525,6 +1525,23 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte)
return pte_clear_flags(pte, _PAGE_SWP_EXCLUSIVE);
}
+#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd)
+{
+ return pmd_set_flags(pmd, _PAGE_SWP_EXCLUSIVE);
+}
+
+static inline bool pmd_swp_exclusive(pmd_t pmd)
+{
+ return pmd_flags(pmd) & _PAGE_SWP_EXCLUSIVE;
+}
+
+static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd)
+{
+ return pmd_clear_flags(pmd, _PAGE_SWP_EXCLUSIVE);
+}
+#endif
+
#ifdef CONFIG_HAVE_ARCH_SOFT_DIRTY
static inline pte_t pte_swp_mksoft_dirty(pte_t pte)
{
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 08/29] mm: recognize PMD swap entries in the softleaf layer
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (6 preceding siblings ...)
2026-09-14 11:41 ` [PATCH v7 07/29] x86: " Usama Arif
@ 2026-09-14 11:42 ` Usama Arif
2026-09-14 11:42 ` [PATCH v7 09/29] mm/debug_vm_pgtable: test PMD swap-exclusive helpers Usama Arif
` (2 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:42 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif
Reclaim splits a PMD-mapped anonymous THP into PTE-level swap entries
before unmapping it, so an ordinary swap entry has never had to appear in a
PMD. Later patches install one there instead, and the softleaf layer is
where every consumer decodes non-present PMDs.
Accept swap entries as valid PMD softleaves and add pmd_is_swap_entry().
A swap entry carries no PFN, so make pmd_softleaf_to_folio() warn and
return NULL rather than interpret a swap offset as a page frame number.
Unlike migration and device-private entries, a PMD swap entry can also
carry the swap-exclusive marker, which softleaf_from_pmd() has to strip
before decoding. Strip all three overlays unconditionally while we are
here: each clear is a plain bit clear, so testing first only buys a branch.
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
include/linux/leafops.h | 40 ++++++++++++++++++++++++++++------------
include/linux/pgtable.h | 17 +++++++++++++++++
2 files changed, 45 insertions(+), 12 deletions(-)
diff --git a/include/linux/leafops.h b/include/linux/leafops.h
index 7c13c58a5e218..ce176c78cefd4 100644
--- a/include/linux/leafops.h
+++ b/include/linux/leafops.h
@@ -98,10 +98,9 @@ static inline softleaf_t softleaf_from_pmd(pmd_t pmd)
if (pmd_present(pmd) || pmd_none(pmd))
return softleaf_mk_none();
- if (pmd_swp_soft_dirty(pmd))
- pmd = pmd_swp_clear_soft_dirty(pmd);
- if (pmd_swp_uffd(pmd))
- pmd = pmd_swp_clear_uffd(pmd);
+ pmd = pmd_swp_clear_soft_dirty(pmd);
+ pmd = pmd_swp_clear_uffd(pmd);
+ pmd = pmd_swp_clear_exclusive(pmd);
arch_entry = __pmd_to_swp_entry(pmd);
/* Temporary until swp_entry_t eliminated. */
@@ -634,18 +633,29 @@ static inline bool pmd_is_migration_entry(pmd_t pmd)
*/
static inline bool softleaf_is_valid_pmd_entry(softleaf_t entry)
{
- /* Only device private, migration entries valid for PMD. */
return softleaf_is_device_private(entry) ||
- softleaf_is_migration(entry);
+ softleaf_is_migration(entry) ||
+ softleaf_is_swap(entry);
+}
+
+/**
+ * pmd_is_swap_entry() - Does this PMD entry encode an actual swap entry?
+ * @pmd: PMD entry.
+ *
+ * Returns: true if the PMD encodes a swap entry, otherwise false.
+ */
+static inline bool pmd_is_swap_entry(pmd_t pmd)
+{
+ return softleaf_is_swap(softleaf_from_pmd(pmd));
}
/**
* pmd_is_valid_softleaf() - Is this PMD entry a valid softleaf entry?
* @pmd: PMD entry.
*
- * PMD leaf entries are valid only if they are device private or migration
- * entries. This function asserts that a PMD leaf entry is valid in this
- * respect.
+ * PMD leaf entries are valid only if they are device private, migration,
+ * or swap entries. This function asserts that a PMD leaf entry is valid
+ * in this respect.
*
* Returns: true if the PMD entry is a valid leaf entry, otherwise false.
*/
@@ -660,10 +670,12 @@ static inline bool pmd_is_valid_softleaf(pmd_t pmd)
* pmd_softleaf_to_folio() - Convert the PMD softleaf entry to a folio.
* @pmd: PMD entry.
*
- * The PMD entry is expected to be a valid PMD softleaf entry.
+ * The PMD entry is expected to be a valid PMD softleaf entry that references a
+ * PFN, that is a migration or device private entry. A PMD swap entry is a valid
+ * softleaf entry but encodes swap slots rather than a PFN, so it has no folio.
*
- * Returns: the folio the softleaf entry references if this is a valid softleaf
- * entry, otherwise NULL.
+ * Returns: the folio the softleaf entry references, or NULL if the entry is not
+ * a valid PMD softleaf entry or does not reference a PFN.
*/
static inline struct folio *pmd_softleaf_to_folio(pmd_t pmd)
{
@@ -673,6 +685,10 @@ static inline struct folio *pmd_softleaf_to_folio(pmd_t pmd)
VM_WARN_ON_ONCE(true);
return NULL;
}
+ if (!softleaf_has_pfn(entry)) {
+ VM_WARN_ON_ONCE(true);
+ return NULL;
+ }
return softleaf_to_folio(entry);
}
diff --git a/include/linux/pgtable.h b/include/linux/pgtable.h
index e3c8ab96941c5..3f955f836abfe 100644
--- a/include/linux/pgtable.h
+++ b/include/linux/pgtable.h
@@ -1917,6 +1917,23 @@ static inline pmd_t pmd_swp_clear_soft_dirty(pmd_t pmd)
}
#endif
+#ifndef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd)
+{
+ return pmd;
+}
+
+static inline bool pmd_swp_exclusive(pmd_t pmd)
+{
+ return false;
+}
+
+static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd)
+{
+ return pmd;
+}
+#endif
+
#ifndef __HAVE_PFNMAP_TRACKING
/*
* Interfaces that can be used by architecture code to keep track of
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 09/29] mm/debug_vm_pgtable: test PMD swap-exclusive helpers
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (7 preceding siblings ...)
2026-09-14 11:42 ` [PATCH v7 08/29] mm: recognize PMD swap entries in the softleaf layer Usama Arif
@ 2026-09-14 11:42 ` Usama Arif
2026-09-14 11:42 ` [PATCH v7 10/29] mm: make PMD migration-entry splitting explicit Usama Arif
2026-09-14 12:46 ` [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:42 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif
An architecture that picked a PMD exclusive bit overlapping the swap type
or offset field would otherwise only be caught by data corruption at
runtime. Mirror pte_swap_exclusive_tests() at PMD level.
Suggested-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
mm/debug_vm_pgtable.c | 40 ++++++++++++++++++++++++++++++++++++++++
1 file changed, 40 insertions(+)
diff --git a/mm/debug_vm_pgtable.c b/mm/debug_vm_pgtable.c
index 2875fd22d7bb0..863111c6d4eb3 100644
--- a/mm/debug_vm_pgtable.c
+++ b/mm/debug_vm_pgtable.c
@@ -802,6 +802,45 @@ static void __init pte_swap_exclusive_tests(struct pgtable_debug_args *args)
WARN_ON(memcmp(&entry, &softleaf, sizeof(entry)));
}
+#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES
+static void __init pmd_swap_exclusive_tests(struct pgtable_debug_args *args)
+{
+ swp_entry_t entry;
+ softleaf_t softleaf;
+ pmd_t pmd;
+
+ if (!has_transparent_hugepage())
+ return;
+
+ pr_debug("Validating PMD swap exclusive\n");
+ entry = args->swp_entry;
+
+ pmd = softleaf_to_pmd(entry);
+ softleaf = softleaf_from_pmd(pmd);
+
+ WARN_ON(pmd_swp_exclusive(pmd));
+ WARN_ON(!softleaf_is_swap(softleaf));
+ WARN_ON(memcmp(&entry, &softleaf, sizeof(entry)));
+
+ pmd = pmd_swp_mkexclusive(pmd);
+ softleaf = softleaf_from_pmd(pmd);
+
+ WARN_ON(!pmd_swp_exclusive(pmd));
+ WARN_ON(!softleaf_is_swap(softleaf));
+ WARN_ON(pmd_swp_soft_dirty(pmd));
+ WARN_ON(memcmp(&entry, &softleaf, sizeof(entry)));
+
+ pmd = pmd_swp_clear_exclusive(pmd);
+ softleaf = softleaf_from_pmd(pmd);
+
+ WARN_ON(pmd_swp_exclusive(pmd));
+ WARN_ON(!softleaf_is_swap(softleaf));
+ WARN_ON(memcmp(&entry, &softleaf, sizeof(entry)));
+}
+#else /* !CONFIG_ARCH_HAS_PMD_SOFTLEAVES */
+static void __init pmd_swap_exclusive_tests(struct pgtable_debug_args *args) { }
+#endif /* CONFIG_ARCH_HAS_PMD_SOFTLEAVES */
+
static void __init pte_swap_tests(struct pgtable_debug_args *args)
{
swp_entry_t arch_entry;
@@ -1322,6 +1361,7 @@ static int __init debug_vm_pgtable(void)
pmd_leaf_soft_dirty_tests(&args);
pte_swap_exclusive_tests(&args);
+ pmd_swap_exclusive_tests(&args);
pte_swap_tests(&args);
pmd_softleaf_tests(&args);
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH v7 10/29] mm: make PMD migration-entry splitting explicit
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (8 preceding siblings ...)
2026-09-14 11:42 ` [PATCH v7 09/29] mm/debug_vm_pgtable: test PMD swap-exclusive helpers Usama Arif
@ 2026-09-14 11:42 ` Usama Arif
2026-09-14 12:46 ` [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 11:42 UTC (permalink / raw)
To: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm
Cc: ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R. Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team, Usama Arif
__split_huge_pmd() and friends take a "freeze" boolean that every caller
has to pass and almost every caller passes as false. The name says nothing
about what it selects, and the one thing it does select - PTE migration
entries instead of PTE mappings - is only ever wanted by the rmap migration
path.
Rename it to use_migration_entries, keep it private to mm/huge_memory.c,
and add split_pmd_to_migration_entries() for try_to_migrate_one(), the only
caller that wants it.
migrate_vma_split_unmapped_folio() also passed freeze=true, but only ever
runs on a PMD that is already a migration entry, which the generic helper
expands into PTE migration entries either way. Its folio_get() only existed
to balance the put_page() that freeze=true performs, so both go.
No functional change intended.
Suggested-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Usama Arif <usama.arif@linux.dev>
---
include/linux/huge_mm.h | 22 ++++++++-------
mm/huge_memory.c | 60 ++++++++++++++++++++++++-----------------
mm/memory.c | 4 +--
mm/migrate_device.c | 7 +----
mm/mprotect.c | 2 +-
mm/rmap.c | 7 +++--
6 files changed, 55 insertions(+), 47 deletions(-)
diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h
index 8ca0fa3be2acb..64b6a2eea899d 100644
--- a/include/linux/huge_mm.h
+++ b/include/linux/huge_mm.h
@@ -430,7 +430,7 @@ int folio_memcg_alloc_deferred(struct folio *folio);
void deferred_split_folio(struct folio *folio, bool partially_mapped);
void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd,
- unsigned long address, bool freeze);
+ unsigned long address);
/**
* pmd_is_huge() - Is this PMD either a huge PMD entry or a software leaf entry?
@@ -462,12 +462,10 @@ static inline bool pmd_is_huge(pmd_t pmd)
do { \
pmd_t *____pmd = (__pmd); \
if (pmd_is_huge(*____pmd)) \
- __split_huge_pmd(__vma, __pmd, __address, \
- false); \
+ __split_huge_pmd(__vma, __pmd, __address); \
} while (0)
-void split_huge_pmd_address(struct vm_area_struct *vma, unsigned long address,
- bool freeze);
+void split_huge_pmd_address(struct vm_area_struct *vma, unsigned long address);
void __split_huge_pud(struct vm_area_struct *vma, pud_t *pud,
unsigned long address);
@@ -590,7 +588,9 @@ static inline bool thp_migration_supported(void)
}
void split_huge_pmd_locked(struct vm_area_struct *vma, unsigned long address,
- pmd_t *pmd, bool freeze);
+ pmd_t *pmd);
+void split_pmd_to_migration_entries(struct vm_area_struct *vma,
+ unsigned long address, pmd_t *pmd);
bool unmap_huge_pmd_locked(struct vm_area_struct *vma, unsigned long addr,
pmd_t *pmdp, struct folio *folio);
void map_anon_folio_pmd_nopf(struct folio *folio, pmd_t *pmd,
@@ -690,12 +690,14 @@ static inline void deferred_split_folio(struct folio *folio, bool partially_mapp
do { } while (0)
static inline void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd,
- unsigned long address, bool freeze) {}
+ unsigned long address) {}
static inline void split_huge_pmd_address(struct vm_area_struct *vma,
- unsigned long address, bool freeze) {}
+ unsigned long address) {}
static inline void split_huge_pmd_locked(struct vm_area_struct *vma,
- unsigned long address, pmd_t *pmd,
- bool freeze) {}
+ unsigned long address, pmd_t *pmd) {}
+static inline void
+split_pmd_to_migration_entries(struct vm_area_struct *vma,
+ unsigned long address, pmd_t *pmd) {}
static inline bool unmap_huge_pmd_locked(struct vm_area_struct *vma,
unsigned long addr, pmd_t *pmdp,
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index ee8d46827ffdc..873887aed0bc2 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -2033,7 +2033,7 @@ int copy_huge_pmd(struct mm_struct *dst_mm, struct mm_struct *src_mm,
pte_free(dst_mm, pgtable);
spin_unlock(src_ptl);
spin_unlock(dst_ptl);
- __split_huge_pmd(src_vma, src_pmd, addr, false);
+ __split_huge_pmd(src_vma, src_pmd, addr);
return -EAGAIN;
}
add_mm_counter(dst_mm, MM_ANONPAGES, HPAGE_PMD_NR);
@@ -2257,7 +2257,7 @@ vm_fault_t do_huge_pmd_wp_page(struct vm_fault *vmf)
folio_unlock(folio);
spin_unlock(vmf->ptl);
fallback:
- __split_huge_pmd(vma, vmf->pmd, vmf->address, false);
+ __split_huge_pmd(vma, vmf->pmd, vmf->address);
return VM_FAULT_FALLBACK;
}
@@ -3190,7 +3190,7 @@ static void __split_huge_zero_page_pmd(struct vm_area_struct *vma,
}
static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
- unsigned long haddr, bool freeze)
+ unsigned long haddr, bool use_migration_entries)
{
struct mm_struct *mm = vma->vm_mm;
struct folio *folio;
@@ -3291,10 +3291,10 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
* folios w.r.t anon exclusive handling. See the comments for
* folio handling and anon_exclusive below.
*/
- if (freeze && anon_exclusive &&
+ if (use_migration_entries && anon_exclusive &&
folio_try_share_anon_rmap_pmd(folio, page))
- freeze = false;
- if (!freeze) {
+ use_migration_entries = false;
+ if (!use_migration_entries) {
rmap_t rmap_flags = RMAP_NONE;
folio_ref_add(folio, HPAGE_PMD_NR - 1);
@@ -3344,11 +3344,11 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
VM_WARN_ON_FOLIO(!folio_test_anon(folio), folio);
/*
- * Without "freeze", we'll simply split the PMD, propagating the
- * PageAnonExclusive() flag for each PTE by setting it for
+ * Without migration entries, we'll simply split the PMD and
+ * propagate the PageAnonExclusive() flag for each PTE by setting it for
* each subpage -- no need to (temporarily) clear.
*
- * With "freeze" we want to replace mapped pages by
+ * With migration entries we want to replace mapped pages by
* migration entries right away. This is only possible if we
* managed to clear PageAnonExclusive() -- see
* set_pmd_migration_entry().
@@ -3359,10 +3359,10 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
* See folio_try_share_anon_rmap_pmd(): invalidate PMD first.
*/
anon_exclusive = PageAnonExclusive(page);
- if (freeze && anon_exclusive &&
+ if (use_migration_entries && anon_exclusive &&
folio_try_share_anon_rmap_pmd(folio, page))
- freeze = false;
- if (!freeze) {
+ use_migration_entries = false;
+ if (!use_migration_entries) {
rmap_t rmap_flags = RMAP_NONE;
folio_ref_add(folio, HPAGE_PMD_NR - 1);
@@ -3387,7 +3387,7 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
* Note that NUMA hinting access restrictions are not transferred to
* avoid any possibility of altering permissions across VMAs.
*/
- if (freeze || pmd_is_migration_entry(old_pmd)) {
+ if (use_migration_entries || pmd_is_migration_entry(old_pmd)) {
pte_t entry;
swp_entry_t swp_entry;
@@ -3420,8 +3420,8 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
for (i = 0, addr = haddr; i < HPAGE_PMD_NR; i++, addr += PAGE_SIZE) {
/*
* anon_exclusive was already propagated to the relevant
- * pages corresponding to the pte entries when freeze
- * is false.
+ * pages corresponding to the pte entries when
+ * use_migration_entries is false.
*/
if (write)
swp_entry = make_writable_device_private_entry(
@@ -3469,7 +3469,7 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
if (!pmd_is_migration_entry(*pmd))
folio_remove_rmap_pmd(folio, page, vma);
- if (freeze)
+ if (use_migration_entries)
put_page(page);
smp_wmb(); /* make pte visible before pmd */
@@ -3477,15 +3477,28 @@ static void __split_huge_pmd_locked(struct vm_area_struct *vma, pmd_t *pmd,
}
void split_huge_pmd_locked(struct vm_area_struct *vma, unsigned long address,
- pmd_t *pmd, bool freeze)
+ pmd_t *pmd)
{
VM_WARN_ON_ONCE(!IS_ALIGNED(address, HPAGE_PMD_SIZE));
if (pmd_trans_huge(*pmd) || pmd_is_valid_softleaf(*pmd))
- __split_huge_pmd_locked(vma, pmd, address, freeze);
+ __split_huge_pmd_locked(vma, pmd, address, false);
+}
+
+/*
+ * Split a present PMD into PTE migration entries, for the rmap migration
+ * walker. Like split_huge_pmd_locked(), the caller must hold the PMD lock and
+ * must already be inside an mmu_notifier invalidate range.
+ */
+void split_pmd_to_migration_entries(struct vm_area_struct *vma,
+ unsigned long address, pmd_t *pmd)
+{
+ VM_WARN_ON_ONCE(!IS_ALIGNED(address, HPAGE_PMD_SIZE));
+ if (pmd_trans_huge(*pmd) || pmd_is_valid_softleaf(*pmd))
+ __split_huge_pmd_locked(vma, pmd, address, true);
}
void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd,
- unsigned long address, bool freeze)
+ unsigned long address)
{
spinlock_t *ptl;
struct mmu_notifier_range range;
@@ -3495,20 +3508,19 @@ void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd,
(address & HPAGE_PMD_MASK) + HPAGE_PMD_SIZE);
mmu_notifier_invalidate_range_start(&range);
ptl = pmd_lock(vma->vm_mm, pmd);
- split_huge_pmd_locked(vma, range.start, pmd, freeze);
+ split_huge_pmd_locked(vma, range.start, pmd);
spin_unlock(ptl);
mmu_notifier_invalidate_range_end(&range);
}
-void split_huge_pmd_address(struct vm_area_struct *vma, unsigned long address,
- bool freeze)
+void split_huge_pmd_address(struct vm_area_struct *vma, unsigned long address)
{
pmd_t *pmd = mm_find_pmd(vma->vm_mm, address);
if (!pmd)
return;
- __split_huge_pmd(vma, pmd, address, freeze);
+ __split_huge_pmd(vma, pmd, address);
}
static inline void split_huge_pmd_if_needed(struct vm_area_struct *vma, unsigned long address)
@@ -3520,7 +3532,7 @@ static inline void split_huge_pmd_if_needed(struct vm_area_struct *vma, unsigned
if (!IS_ALIGNED(address, HPAGE_PMD_SIZE) &&
range_in_vma(vma, ALIGN_DOWN(address, HPAGE_PMD_SIZE),
ALIGN(address, HPAGE_PMD_SIZE)))
- split_huge_pmd_address(vma, address, false);
+ split_huge_pmd_address(vma, address);
}
void vma_adjust_trans_huge(struct vm_area_struct *vma,
diff --git a/mm/memory.c b/mm/memory.c
index 926276d419202..477d7e359b447 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -2096,7 +2096,7 @@ static inline unsigned long zap_pmd_range(struct mmu_gather *tlb,
next = pmd_addr_end(addr, end);
if (pmd_is_huge(*pmd)) {
if (next - addr != HPAGE_PMD_SIZE)
- __split_huge_pmd(vma, pmd, addr, false);
+ __split_huge_pmd(vma, pmd, addr);
else if (zap_huge_pmd(tlb, vma, pmd, addr)) {
addr = next;
continue;
@@ -6382,7 +6382,7 @@ static inline vm_fault_t wp_huge_pmd(struct vm_fault *vmf)
split:
/* COW or write-notify handled on pte level: split pmd. */
- __split_huge_pmd(vma, vmf->pmd, vmf->address, false);
+ __split_huge_pmd(vma, vmf->pmd, vmf->address);
return VM_FAULT_FALLBACK;
}
diff --git a/mm/migrate_device.c b/mm/migrate_device.c
index 0c437004329d9..4a0b61d50d222 100644
--- a/mm/migrate_device.c
+++ b/mm/migrate_device.c
@@ -918,12 +918,7 @@ static int migrate_vma_split_unmapped_folio(struct migrate_vma *migrate,
unsigned long flags;
int ret = 0;
- /*
- * take a reference, since split_huge_pmd_address() with freeze = true
- * drops a reference at the end.
- */
- folio_get(folio);
- split_huge_pmd_address(migrate->vma, addr, true);
+ split_huge_pmd_address(migrate->vma, addr);
ret = folio_split_unmapped(folio, 0);
if (ret)
return ret;
diff --git a/mm/mprotect.c b/mm/mprotect.c
index 2888ee638d872..ee33bbb421008 100644
--- a/mm/mprotect.c
+++ b/mm/mprotect.c
@@ -530,7 +530,7 @@ static inline long change_pmd_range(struct mmu_gather *tlb,
if (pmd_is_huge(_pmd)) {
if ((next - addr != HPAGE_PMD_SIZE) ||
pgtable_split_needed(vma, cp_flags)) {
- __split_huge_pmd(vma, pmd, addr, false);
+ __split_huge_pmd(vma, pmd, addr);
/*
* For file-backed, the pmd could have been
* cleared; make sure pmd populated if
diff --git a/mm/rmap.c b/mm/rmap.c
index 5332c52909be1..feb751e29b992 100644
--- a/mm/rmap.c
+++ b/mm/rmap.c
@@ -2290,7 +2290,7 @@ static bool try_to_unmap_one(struct folio *folio, struct vm_area_struct *vma,
* restart so we can process the PTE-mapped THP.
*/
split_huge_pmd_locked(vma, pvmw.address,
- pvmw.pmd, false);
+ pvmw.pmd);
flags &= ~TTU_SPLIT_HUGE_PMD;
page_vma_mapped_walk_restart(&pvmw);
continue;
@@ -2515,13 +2515,12 @@ static bool try_to_migrate_one(struct folio *folio, struct vm_area_struct *vma,
if (flags & TTU_SPLIT_HUGE_PMD) {
/*
- * split_huge_pmd_locked() might leave the
+ * split_pmd_to_migration_entries() might leave the
* folio mapped through PTEs. Retry the walk
* so we can detect this scenario and properly
* abort the walk.
*/
- split_huge_pmd_locked(vma, pvmw.address,
- pvmw.pmd, true);
+ split_pmd_to_migration_entries(vma, pvmw.address, pvmw.pmd);
flags &= ~TTU_SPLIT_HUGE_PMD;
page_vma_mapped_walk_restart(&pvmw);
continue;
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
` (9 preceding siblings ...)
2026-09-14 11:42 ` [PATCH v7 10/29] mm: make PMD migration-entry splitting explicit Usama Arif
@ 2026-09-14 12:46 ` Usama Arif
10 siblings, 0 replies; 12+ messages in thread
From: Usama Arif @ 2026-09-14 12:46 UTC (permalink / raw)
To: Usama Arif
Cc: Andrew Morton, david, chrisl, kasong, ljs, ziy, linux-mm,
ying.huang, Baoquan He, willy, youngjun.park, hannes, riel,
shakeel.butt, alex, kas, baohua, dev.jain, baolin.wang,
Nico Pache, Liam R . Howlett, ryan.roberts, Vlastimil Babka,
lance.yang, linux-kernel, nphamcs, shikemeng, yosry, qi.zheng,
luizcap, kernel-team
On Mon, 14 Sep 2026 04:41:52 -0700 Usama Arif <usama.arif@linux.dev> wrote:
> When reclaim swaps out a PMD-mapped anonymous THP today, the PMD is
> split into HPAGE_PMD_NR PTE-level swap entries via TTU_SPLIT_HUGE_PMD
> before unmap. This series introduces a PMD-level swap entry so the
> huge mapping can survive the swap round-trip and do_huge_pmd_swap_page()
> can restore the PMD mapping directly on swap-in, without waiting for
> khugepaged to collapse the range later.
>
I somehow messed up sending the v7 series and only 10 patches got sent.
I have resent it in [1]. Apologies for this, please direct all reviews to [1].
[1] https://lore.kernel.org/all/20260914122950.3283997-1-usama.arif@linux.dev/
^ permalink raw reply [flat|nested] 12+ messages in thread
end of thread, other threads:[~2026-09-14 12:46 UTC | newest]
Thread overview: 12+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-14 11:41 [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
2026-09-14 11:41 ` [PATCH v7 01/29] mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio() Usama Arif
2026-09-14 11:41 ` [PATCH v7 02/29] arm64: mm: add PMD swap-exclusive helpers Usama Arif
2026-09-14 11:41 ` [PATCH v7 03/29] loongarch: " Usama Arif
2026-09-14 11:41 ` [PATCH v7 04/29] powerpc: " Usama Arif
2026-09-14 11:41 ` [PATCH v7 05/29] riscv: " Usama Arif
2026-09-14 11:41 ` [PATCH v7 06/29] s390: " Usama Arif
2026-09-14 11:41 ` [PATCH v7 07/29] x86: " Usama Arif
2026-09-14 11:42 ` [PATCH v7 08/29] mm: recognize PMD swap entries in the softleaf layer Usama Arif
2026-09-14 11:42 ` [PATCH v7 09/29] mm/debug_vm_pgtable: test PMD swap-exclusive helpers Usama Arif
2026-09-14 11:42 ` [PATCH v7 10/29] mm: make PMD migration-entry splitting explicit Usama Arif
2026-09-14 12:46 ` [PATCH v7 00/29] mm: PMD-level swap entries for anonymous THPs Usama Arif
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®