* [PATCH v2 01/20] hugetlb: Don't restore vmemmap of non-HVOed folios on bulk restore error
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 02/20] arm64/pgtable: Clear AF with LDCLR on supported systems James Houghton
` (18 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
When hugetlb_vmemmap_restore_folios() fails, it stops processing the
list, so folio_list may still contain folios that were never vmemmap
optimized, e.g. because HVO failed when they were allocated. Their
hugetlb flag has already been cleared by remove_hugetlb_folio().
If no folios could be restored, bulk_vmemmap_restore_error() then calls
hugetlb_vmemmap_restore_folio() on every folio left on folio_list,
including those, which trips the !folio_test_hugetlb() check in
__hugetlb_vmemmap_restore_folio():
page dumped because: VM_WARN_ON_ONCE_FOLIO(!folio_test_hugetlb(folio))
WARNING: mm/hugetlb_vmemmap.c:503 at __hugetlb_vmemmap_restore_folio+0x270/0x300
...
__hugetlb_vmemmap_restore_folio+0x270/0x300 (P)
hugetlb_vmemmap_restore_folio+0x1c/0x30
update_and_free_pages_bulk+0x2c0/0x368
__nr_hugepages_store_common+0x2e8/0x3f0
Otherwise this is benign: the restore is a no-op returning success for
non-optimized folios, which are then freed as intended.
Skip the restore for folios that are not vmemmap optimized; they can be
freed right away.
This is very unlikely to happen without fault injection, as it requires
both HVO and a subsequent vmemmap restore to fail.
Fixes: cfb8c75099db ("hugetlb: perform vmemmap restoration on a list of pages")
Assisted-by: LLM
Signed-off-by: James Houghton <jthoughton@google.com>
---
mm/hugetlb.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index a1b51251103b..4dac7ed1df57 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -1613,9 +1613,14 @@ static void bulk_vmemmap_restore_error(struct hstate *h,
* page is made a surplus page and removed from the list.
* If are able to restore vmemmap and free one hugetlb page, we
* quit processing the list to retry the bulk operation.
+ *
+ * Folios whose vmemmap was never optimized (e.g. because HVO
+ * failed when they were allocated) can be freed right away;
+ * remove_hugetlb_folio() has already cleared their hugetlb flag.
*/
list_for_each_entry_safe(folio, t_folio, folio_list, lru)
- if (hugetlb_vmemmap_restore_folio(h, folio)) {
+ if (folio_test_hugetlb_vmemmap_optimized(folio) &&
+ hugetlb_vmemmap_restore_folio(h, folio)) {
list_del(&folio->lru);
spin_lock_irq(&hugetlb_lock);
add_hugetlb_folio(h, folio, true);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 02/20] arm64/pgtable: Clear AF with LDCLR on supported systems
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
2026-10-03 0:21 ` [PATCH v2 01/20] hugetlb: Don't restore vmemmap of non-HVOed folios on bulk restore error James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 03/20] hugetlb_vmemmap: Always flush TLB if needed upon PTE remapping James Houghton
` (17 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Replace the cmpxchg loop with atomic64_fetch_andnot_relaxed, which, on
systems with LSE atomics, should end up as LDCLR.
This is somewhat overdue. In recent versions of the architecture, if a
system supports FEAT_LSE, forward progress guarantees for exclusive
monitors and cmpxchg loops are relaxed.
Mechanically, I would love to write test_and_clear_bit_relaxed(), but no
such function exists (test_and_clear_bit() has full ordering, which is
also undesirable). Implementing test_and_clear_bit_relaxed() properly is
a pretty large change; use atomic64 instead, as pte_t contains a u64.
Signed-off-by: James Houghton <jthoughton@google.com>
---
arch/arm64/include/asm/pgtable.h | 14 ++++----------
1 file changed, 4 insertions(+), 10 deletions(-)
diff --git a/arch/arm64/include/asm/pgtable.h b/arch/arm64/include/asm/pgtable.h
index 763c5a411d64..e47c3d010715 100644
--- a/arch/arm64/include/asm/pgtable.h
+++ b/arch/arm64/include/asm/pgtable.h
@@ -1252,17 +1252,11 @@ static inline void __pte_clear(struct mm_struct *mm,
static inline bool __ptep_test_and_clear_young(struct vm_area_struct *vma,
unsigned long address, pte_t *ptep)
{
- pte_t old_pte, pte;
+ atomic64_t *pteval = (atomic64_t *)&pte_val(*ptep);
+ s64 af_mask = PTE_AF;
- pte = __ptep_get(ptep);
- do {
- old_pte = pte;
- pte = pte_mkold(pte);
- pte_val(pte) = cmpxchg_relaxed(&pte_val(*ptep),
- pte_val(old_pte), pte_val(pte));
- } while (pte_val(pte) != pte_val(old_pte));
-
- return pte_young(pte);
+ /* Atomically clear PTE_AF, checking that it was set before. */
+ return af_mask & atomic64_fetch_andnot_relaxed(af_mask, pteval);
}
static inline bool __ptep_clear_flush_young(struct vm_area_struct *vma,
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 03/20] hugetlb_vmemmap: Always flush TLB if needed upon PTE remapping
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
2026-10-03 0:21 ` [PATCH v2 01/20] hugetlb: Don't restore vmemmap of non-HVOed folios on bulk restore error James Houghton
2026-10-03 0:21 ` [PATCH v2 02/20] arm64/pgtable: Clear AF with LDCLR on supported systems James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 04/20] hugetlb_vmemmap: Leave pages partially HVOed upon restore failure James Houghton
` (16 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
vmemmap_remap_range() can potentially leave the vmemmap in a partially
optimized state. In this case, a TLB flush should be done if the walk
caller asked for it, even if the optimization was not fully applied.
A follow-up patch will more completely deal with partially-optimized
folios.
Signed-off-by: James Houghton <jthoughton@google.com>
---
mm/hugetlb_vmemmap.c | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 0057fa2a16a8..1a2645c382ba 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -163,13 +163,11 @@ static int vmemmap_remap_range(unsigned long start, unsigned long end,
ret = walk_kernel_page_table_range(start, end, &vmemmap_remap_ops,
NULL, walk);
mmap_read_unlock(&init_mm);
- if (ret)
- return ret;
if (walk->remap_pte && !(walk->flags & VMEMMAP_REMAP_NO_TLB_FLUSH))
flush_tlb_kernel_range(start, end);
- return 0;
+ return ret;
}
/*
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 04/20] hugetlb_vmemmap: Leave pages partially HVOed upon restore failure
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (2 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 03/20] hugetlb_vmemmap: Always flush TLB if needed upon PTE remapping James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 05/20] hugetlb_vmemmap: Use try_update_vmemmap_pte to update in-use PTEs James Houghton
` (15 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Right now it is assumed that the restore routine in vmemmap_remap_free()
will not fail; that is, if we fail to fully optimize a folio, it is
assumed that restoring the folio to its pre-HVO state is guaranteed to
succeed.
Although vmemmap restoration may never fail in practice, properly handle
such failure cases now. This will leave folios in a partially HVOed
state. In a follow-up patch, failure certainly becomes possible.
In the event we end up with partially-HVOed folios, make sure to:
1. Leave the HVO page flag in place.
2. Free any unused vmemmap pages.
Always pass the vmemmap_tail page to the restore routine to avoid
leaking the non-optimized vmemmap pages for a partially HVOed page.
Signed-off-by: James Houghton <jthoughton@google.com>
---
mm/hugetlb_vmemmap.c | 49 ++++++++++++++++++++++++++++++++++----------
1 file changed, 38 insertions(+), 11 deletions(-)
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 1a2645c382ba..90db4d069ff6 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -235,12 +235,15 @@ static void vmemmap_restore_pte(pte_t *pte, unsigned long addr,
{
struct page *src = pte_page(ptep_get(pte)), *dst;
+ if (WARN_ON_ONCE(!walk->vmemmap_tail))
+ return;
+
/*
- * When rolling back vmemmap_remap_free(), keep the copied head page
+ * When restoring a partially-HVOed page, keep the copied head page
* mapping and restore only PTEs currently pointing at the shared tail
* page.
*/
- if (walk->vmemmap_tail && walk->vmemmap_tail != src)
+ if (walk->vmemmap_tail != src)
return;
VM_WARN_ON_ONCE(PageHead((const struct page *)addr));
@@ -276,6 +279,7 @@ static int vmemmap_remap_split(unsigned long start, unsigned long end)
return vmemmap_remap_range(start, end, &walk);
}
+#define VMEMMAP_REMAP_INCOMPLETE 1
/**
* vmemmap_remap_free - remap the vmemmap virtual address range [@start, @end)
* to use @vmemmap_head/tail, then free vmemmap which
@@ -290,7 +294,8 @@ static int vmemmap_remap_split(unsigned long start, unsigned long end)
* responsibility to free pages.
* @flags: modifications to vmemmap_remap_walk flags
*
- * Return: %0 on success, negative error code otherwise.
+ * Return: %0 on success, VMEMMAP_REMAP_INCOMPLETE if the page is incompletely
+ * optimized, negative error code otherwise.
*/
static int vmemmap_remap_free(unsigned long start, unsigned long end,
struct page *vmemmap_head,
@@ -325,7 +330,8 @@ static int vmemmap_remap_free(unsigned long start, unsigned long end,
.flags = 0,
};
- vmemmap_remap_range(start, end, &walk);
+ if (vmemmap_remap_range(start, end, &walk))
+ return VMEMMAP_REMAP_INCOMPLETE;
return ret;
}
@@ -358,6 +364,8 @@ static int alloc_vmemmap_page_list(unsigned long start, unsigned long end,
* vmemmap_remap_alloc - remap the vmemmap virtual address range [@start, end)
* to the page which is from the @vmemmap_pages
* respectively.
+ * @h: the hstate for the folio whose vmemmap is getting remapped
+ * @folio: the folio whose vmemmap is getting remapped
* @start: start address of the vmemmap virtual address range that we want
* to remap.
* @end: end address of the vmemmap virtual address range that we want to
@@ -366,20 +374,35 @@ static int alloc_vmemmap_page_list(unsigned long start, unsigned long end,
*
* Return: %0 on success, negative error code otherwise.
*/
-static int vmemmap_remap_alloc(unsigned long start, unsigned long end,
+static int vmemmap_remap_alloc(const struct hstate *h, struct folio *folio,
+ unsigned long start, unsigned long end,
unsigned long flags)
{
LIST_HEAD(vmemmap_pages);
- struct vmemmap_remap_walk walk = {
+ struct vmemmap_remap_walk walk;
+ struct page *vmemmap_tail;
+ int ret;
+
+ vmemmap_tail = vmemmap_shared_tail_page(h->order, folio_zone(folio));
+ if (WARN_ON_ONCE(!vmemmap_tail))
+ return -ENOMEM;
+
+ if (alloc_vmemmap_page_list(start, end, &vmemmap_pages))
+ return -ENOMEM;
+
+ walk = (struct vmemmap_remap_walk) {
.remap_pte = vmemmap_restore_pte,
+ .vmemmap_tail = vmemmap_tail,
.vmemmap_pages = &vmemmap_pages,
.flags = flags,
};
- if (alloc_vmemmap_page_list(start, end, &vmemmap_pages))
- return -ENOMEM;
+ ret = vmemmap_remap_range(start, end, &walk);
- return vmemmap_remap_range(start, end, &walk);
+ /* Not all pages may have been consumed */
+ free_vmemmap_page_list(&vmemmap_pages);
+
+ return ret;
}
static bool vmemmap_optimize_enabled = IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP_DEFAULT_ON);
@@ -412,7 +435,7 @@ static int __hugetlb_vmemmap_restore_folio(const struct hstate *h,
* When a HugeTLB page is freed to the buddy allocator, previously
* discarded vmemmap pages must be allocated and remapping.
*/
- ret = vmemmap_remap_alloc(vmemmap_start, vmemmap_end, flags);
+ ret = vmemmap_remap_alloc(h, folio, vmemmap_start, vmemmap_end, flags);
if (!ret)
folio_clear_hugetlb_vmemmap_optimized(folio);
@@ -545,7 +568,11 @@ static int __hugetlb_vmemmap_optimize_folio(const struct hstate *h,
vmemmap_head, vmemmap_tail,
vmemmap_pages, flags);
out:
- if (ret)
+ /*
+ * If ret == VMEMMAP_REMAP_INCOMPLETE, the folio might be partially
+ * HVOed. Leave the HVO page folio flag in place.
+ */
+ if (ret < 0)
folio_clear_hugetlb_vmemmap_optimized(folio);
return ret;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 05/20] hugetlb_vmemmap: Use try_update_vmemmap_pte to update in-use PTEs
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (3 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 04/20] hugetlb_vmemmap: Leave pages partially HVOed upon restore failure James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 06/20] hugetlb_vmemmap: Allow architectures to dynamically disallow HVO James Houghton
` (14 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
set_pte_at() cannot be used to replace in-use vmemmap PTEs on arm64, so
replace it with a more specific routine, try_update_vmemmap_pte().
try_update_vmemmap_pte() is used for modifying page table entries such
that there is a guarantee that no fault will be taken.
For the generic implementation, please note one difference: set_pte()
does not invoke page_table_check_ptes_set(), but set_pte_at(), the call
we are about to replace, does. However, this is a functional no-op
because page_table_check_ptes_set() does nothing for init_mm PTEs, which
vmemmap PTEs are.
Signed-off-by: James Houghton <jthoughton@google.com>
---
arch/loongarch/include/asm/pgtable.h | 2 +
arch/riscv/include/asm/pgtable.h | 2 +
arch/x86/include/asm/pgtable.h | 2 +
include/linux/pgtable.h | 22 +++++++++++
mm/hugetlb_vmemmap.c | 56 +++++++++++++++++++---------
5 files changed, 67 insertions(+), 17 deletions(-)
diff --git a/arch/loongarch/include/asm/pgtable.h b/arch/loongarch/include/asm/pgtable.h
index f87603135131..123eda3f6035 100644
--- a/arch/loongarch/include/asm/pgtable.h
+++ b/arch/loongarch/include/asm/pgtable.h
@@ -633,6 +633,8 @@ static inline long pmd_protnone(pmd_t pmd)
#define pmd_leaf(pmd) ((pmd_val(pmd) & _PAGE_HUGE) != 0)
#define pud_leaf(pud) ((pud_val(pud) & _PAGE_HUGE) != 0)
+#define ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PTE
+
/*
* We provide our own get_unmapped area to cope with the virtual aliasing
* constraints placed on us by the cache architecture.
diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
index 4c8fc6845503..6cfae7720882 100644
--- a/arch/riscv/include/asm/pgtable.h
+++ b/arch/riscv/include/asm/pgtable.h
@@ -1165,6 +1165,8 @@ static inline pud_t pud_modify(pud_t pud, pgprot_t newprot)
#endif /* CONFIG_TRANSPARENT_HUGEPAGE */
+#define ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PTE
+
/*
* Encode/decode swap entries and swap PTEs. Swap PTEs are all PTEs that
* are !pte_none() && !pte_present().
diff --git a/arch/x86/include/asm/pgtable.h b/arch/x86/include/asm/pgtable.h
index ef0252a09c28..f60c09cdf1ec 100644
--- a/arch/x86/include/asm/pgtable.h
+++ b/arch/x86/include/asm/pgtable.h
@@ -1368,6 +1368,8 @@ static inline pmd_t pmdp_establish(struct vm_area_struct *vma,
}
#endif
+#define ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PTE
+
#ifdef CONFIG_HAVE_ARCH_TRANSPARENT_HUGEPAGE_PUD
static inline pud_t pudp_establish(struct vm_area_struct *vma,
unsigned long address, pud_t *pudp, pud_t pud)
diff --git a/include/linux/pgtable.h b/include/linux/pgtable.h
index 780fe849ff8b..59ab7ca93548 100644
--- a/include/linux/pgtable.h
+++ b/include/linux/pgtable.h
@@ -457,6 +457,28 @@ static inline void set_ptes(struct mm_struct *mm, unsigned long addr,
#endif
#define set_pte_at(mm, addr, ptep, pte) set_ptes(mm, addr, ptep, pte, 1)
+#ifdef ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PTE
+/*
+ * try_update_vmemmap_pte - Remap PTEs used by the vmemmap.
+ * @addr: Base address of the remapped PTE.
+ * @ptep: Page table pointer to be overwritten.
+ * @pte: Page table entry to write.
+ *
+ * This function is only to be used to update PTEs that map the vmemmap. The
+ * only valid transitions supported by this function are: leaf-level
+ * (PAGE_SIZE), valid-to-valid. The pfn and prot bits may be changed.
+ *
+ * Implementations of this function must ensure that, while the update is taking
+ * place, CPUs will not fault on the remapped virtual address.
+ */
+static inline int try_update_vmemmap_pte(unsigned long addr, pte_t *ptep,
+ pte_t pte)
+{
+ set_pte(ptep, pte);
+ return 0;
+}
+#endif
+
#ifndef __HAVE_ARCH_PTEP_SET_ACCESS_FLAGS
extern int ptep_set_access_flags(struct vm_area_struct *vma,
unsigned long address, pte_t *ptep,
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 90db4d069ff6..3fdb1e4ce1a1 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -33,7 +33,7 @@
* operations.
*/
struct vmemmap_remap_walk {
- void (*remap_pte)(pte_t *pte, unsigned long addr,
+ int (*remap_pte)(pte_t *pte, unsigned long addr,
struct vmemmap_remap_walk *walk);
unsigned long nr_walked;
@@ -140,11 +140,13 @@ static int vmemmap_pte_entry(pte_t *pte, unsigned long addr,
unsigned long next, struct mm_walk *walk)
{
struct vmemmap_remap_walk *vmemmap_walk = walk->private;
+ int ret = 0;
- vmemmap_walk->remap_pte(pte, addr, vmemmap_walk);
- vmemmap_walk->nr_walked++;
+ ret = vmemmap_walk->remap_pte(pte, addr, vmemmap_walk);
+ if (!ret)
+ vmemmap_walk->nr_walked++;
- return 0;
+ return ret;
}
static const struct mm_walk_ops vmemmap_remap_ops = {
@@ -196,18 +198,20 @@ static void free_vmemmap_page_list(struct list_head *list)
free_vmemmap_page(page);
}
-static void vmemmap_remap_pte(pte_t *pte, unsigned long addr,
- struct vmemmap_remap_walk *walk)
+static int vmemmap_remap_pte(pte_t *pte, unsigned long addr,
+ struct vmemmap_remap_walk *walk)
{
struct page *page = pte_page(ptep_get(pte));
pte_t entry;
+ bool head;
+ int ret;
+
+ head = walk->nr_walked == 0 && walk->vmemmap_head;
/* Remapping the head page requires r/w */
- if (unlikely(walk->nr_walked == 0 && walk->vmemmap_head)) {
+ if (unlikely(head)) {
VM_WARN_ON_ONCE(!PageHead((const struct page *)addr));
- list_del(&walk->vmemmap_head->lru);
-
/*
* Makes sure that preceding stores to the page contents from
* vmemmap_remap_free() become visible before the set_pte_at()
@@ -226,17 +230,30 @@ static void vmemmap_remap_pte(pte_t *pte, unsigned long addr,
entry = mk_pte(walk->vmemmap_tail, PAGE_KERNEL_RO);
}
+ ret = try_update_vmemmap_pte(addr, pte, entry);
+ if (ret)
+ return ret;
+
+ /* We successfully overwrote the vmemmap PTE, so we can free
+ * the vmemmap page that was just unmapped, and if we mapped
+ * the new head page, remove it from the list so that it
+ * doesn't get freed later.
+ */
list_add(&page->lru, walk->vmemmap_pages);
- set_pte_at(&init_mm, addr, pte, entry);
+ if (head)
+ list_del(&walk->vmemmap_head->lru);
+
+ return 0;
}
-static void vmemmap_restore_pte(pte_t *pte, unsigned long addr,
- struct vmemmap_remap_walk *walk)
+static int vmemmap_restore_pte(pte_t *pte, unsigned long addr,
+ struct vmemmap_remap_walk *walk)
{
struct page *src = pte_page(ptep_get(pte)), *dst;
+ int ret;
if (WARN_ON_ONCE(!walk->vmemmap_tail))
- return;
+ return -EINVAL;
/*
* When restoring a partially-HVOed page, keep the copied head page
@@ -244,20 +261,25 @@ static void vmemmap_restore_pte(pte_t *pte, unsigned long addr,
* page.
*/
if (walk->vmemmap_tail != src)
- return;
+ return 0;
VM_WARN_ON_ONCE(PageHead((const struct page *)addr));
dst = list_first_entry(walk->vmemmap_pages, struct page, lru);
- list_del(&dst->lru);
copy_page(page_to_virt(dst), page_to_virt(src));
/*
* Makes sure that preceding stores to the page contents become visible
- * before the set_pte_at() write.
+ * before the try_update_vmemmap_pte() write.
*/
smp_wmb();
- set_pte_at(&init_mm, addr, pte, mk_pte(dst, PAGE_KERNEL));
+
+ ret = try_update_vmemmap_pte(addr, pte, mk_pte(dst, PAGE_KERNEL));
+ if (ret)
+ return ret;
+
+ list_del(&dst->lru);
+ return 0;
}
/**
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 06/20] hugetlb_vmemmap: Allow architectures to dynamically disallow HVO
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (4 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 05/20] hugetlb_vmemmap: Use try_update_vmemmap_pte to update in-use PTEs James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 07/20] hugetlb_vmemmap: Disable HVO sysctl if arch doesn't support HVO James Houghton
` (13 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
For arm64, only some platforms support HVO. Add an arch hook to prevent
HVO from taking affect if the architecture lacks support.
It may be the case that arch_supports_hugetlb_vmemmap_optimization()
returns false at first but eventually returns true. This is the case on
arm64. When this happens, bootmem hugepages won't be immediately HVOed,
but they will be optimized in prep_and_add_bootmem_folios().
Signed-off-by: James Houghton <jthoughton@google.com>
---
include/asm-generic/hugetlb.h | 7 +++++++
mm/hugetlb_vmemmap.c | 12 ++++++++++++
2 files changed, 19 insertions(+)
diff --git a/include/asm-generic/hugetlb.h b/include/asm-generic/hugetlb.h
index 635c41cc3479..8d24075d53d4 100644
--- a/include/asm-generic/hugetlb.h
+++ b/include/asm-generic/hugetlb.h
@@ -128,4 +128,11 @@ static inline bool gigantic_page_runtime_supported(void)
}
#endif /* __HAVE_ARCH_GIGANTIC_PAGE_RUNTIME_SUPPORTED */
+#ifndef __HAVE_ARCH_HVO_SUPPORTED
+static inline bool arch_hugetlb_vmemmap_optimization_supported(void)
+{
+ return IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP);
+}
+#endif /* __HAVE_ARCH_HVO_SUPPORTED */
+
#endif /* _ASM_GENERIC_HUGETLB_H */
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 3fdb1e4ce1a1..8596a462fee6 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -16,6 +16,7 @@
#include <linux/pagewalk.h>
#include <linux/pgalloc.h>
#include <linux/vmemmap-optimization.h>
+#include <linux/hugetlb.h>
#include <asm/tlbflush.h>
#include "hugetlb_vmemmap.h"
@@ -529,6 +530,9 @@ static bool vmemmap_should_optimize_folio(const struct hstate *h, struct folio *
if (!READ_ONCE(vmemmap_optimize_enabled))
return false;
+ if (!arch_hugetlb_vmemmap_optimization_supported())
+ return false;
+
if (!hugetlb_vmemmap_optimizable(h))
return false;
@@ -731,6 +735,14 @@ void __init hugetlb_vmemmap_optimize_bootmem_page(unsigned long pfn, unsigned in
if (!READ_ONCE(vmemmap_optimize_enabled))
return;
+ /*
+ * Architectures may return false here but true by the time
+ * hugetlb_init() is called. In this case, although the folios will
+ * not be pre-HVOed, they will be optimized in hugetlb_init().
+ */
+ if (!arch_hugetlb_vmemmap_optimization_supported())
+ return;
+
section_set_compound_order_range(pfn, 1UL << order, order);
}
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 07/20] hugetlb_vmemmap: Disable HVO sysctl if arch doesn't support HVO
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (5 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 06/20] hugetlb_vmemmap: Allow architectures to dynamically disallow HVO James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 08/20] hugetlb: Fully initialize tail struct pages of non-pre-HVOed bootmem folios James Houghton
` (12 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Do not create the vm.hugetlb_optimize_vmemmap sysctl when the
architecture lacks support. This serves only to avoid confusion for
users.
arm64 support for HVO is conditional on certain hardware features being
enabled for all online CPUs. By the time hugetlb_vmemmap_init() runs,
all secondary CPUs will have started, and the system feature set will
have stabilized, including HVO-through-BBM, which
arch_hugetlb_vmemmap_optimization_supported() checks for.
Signed-off-by: James Houghton <jthoughton@google.com>
---
mm/hugetlb_vmemmap.c | 15 +++++++++++++++
1 file changed, 15 insertions(+)
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 8596a462fee6..817ce39d4cf3 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -763,6 +763,21 @@ static int __init hugetlb_vmemmap_init(void)
/* HUGETLB_VMEMMAP_RESERVE_SIZE should cover all used struct pages */
BUILD_BUG_ON(__NR_USED_SUBPAGE > HUGETLB_VMEMMAP_RESERVE_PAGES);
+ /*
+ * Architectures that need to probe CPU support for HVO have done
+ * such probing now. arch_hugetlb_vmemmap_optimization_supported()
+ * will return the correct value.
+ */
+ if (!arch_hugetlb_vmemmap_optimization_supported()) {
+ if (vmemmap_optimize_enabled)
+ pr_info("vmemmap optimization not supported by hardware\n");
+ /*
+ * Return early to avoid setting up the hugetlb vmemmap
+ * sysctls; HVO is not usable.
+ */
+ return 0;
+ }
+
for_each_hstate(h) {
if (hugetlb_vmemmap_optimizable(h)) {
register_sysctl_init("vm", hugetlb_vmemmap_sysctls);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 08/20] hugetlb: Fully initialize tail struct pages of non-pre-HVOed bootmem folios
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (6 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 07/20] hugetlb_vmemmap: Disable HVO sysctl if arch doesn't support HVO James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 09/20] hugetlb_vmemmap: Allow architectures to make HVO enablement boot-time only James Houghton
` (11 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
alloc_bootmem() marks all but the head struct page of a bootmem gigantic
folio as noinit, and gather_bootmem_prealloc_node() then only
initializes the first HUGETLB_VMEMMAP_RESERVE_PAGES struct pages. The
remaining tail struct pages are initialized only if the folio ends up
not being HVOed.
This assumes that a bootmem folio is either pre-HVOed, or will not be
HVOed at all. However, hugetlb_vmemmap_optimize_bootmem_page() may skip
pre-HVO because arch_hugetlb_vmemmap_optimization_supported() does not
yet return true that early in boot (e.g. on arm64, where support depends
on a system-wide CPU capability), while it does return true by the time
hugetlb_vmemmap_optimize_bootmem_folios() runs. Such folios are then
HVOed through the regular remap path with uninitialized tail struct
pages, which trips the PageTail() WARN in vmemmap_remap_pte().
Avoid this by initializing all tail struct pages up front in
gather_bootmem_prealloc_node() for folios that were not pre-HVOed. This
makes the HVO-failure fallback in prep_and_add_bootmem_folios()
unnecessary, as pre-HVOed folios are never passed through the regular
remap path, and all other folios now already have initialized tail
struct pages. A failed optimization either leaves the original vmemmap
in place or restores the tail struct pages from the shared tail page.
Remove the fallback.
For folios that are not HVOed at all, this does not change the amount
of initialization work, only where it is done.
Signed-off-by: James Houghton <jthoughton@google.com>
---
mm/hugetlb.c | 28 +++++++++++++++-------------
1 file changed, 15 insertions(+), 13 deletions(-)
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index 4dac7ed1df57..e971e2362412 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -3303,17 +3303,6 @@ static void __init prep_and_add_bootmem_folios(struct hstate *h,
hugetlb_vmemmap_optimize_bootmem_folios(h, folio_list);
list_for_each_entry_safe(folio, tmp_f, folio_list, lru) {
- if (!folio_test_hugetlb_vmemmap_optimized(folio)) {
- /*
- * If HVO fails, initialize all tail struct pages
- * We do not worry about potential long lock hold
- * time as this is early in boot and there should
- * be no contention.
- */
- hugetlb_folio_init_tail_vmemmap(folio, h,
- HUGETLB_VMEMMAP_RESERVE_PAGES,
- pages_per_huge_page(h));
- }
hugetlb_bootmem_init_migratetype(folio, h);
/* Subdivide locks to achieve better parallel performance */
spin_lock_irqsave(&hugetlb_lock, flags);
@@ -3337,6 +3326,7 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid)
struct page *page = virt_to_page(m);
struct folio *folio = (void *)page;
const unsigned long pfn = folio_pfn(folio);
+ bool pre_hvo;
h = m->hstate;
/*
@@ -3350,11 +3340,23 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid)
VM_BUG_ON(!hstate_is_gigantic(h));
WARN_ON(folio_ref_count(folio) != 1);
+ pre_hvo = vmemmap_optimizable_order(pfn_to_section_compound_order(pfn));
+
+ /*
+ * Pre-HVOed folios have their tail struct pages mirrored from
+ * the shared tail page, so only the first vmemmap page needs
+ * initializing. Otherwise, the tail struct pages (marked noinit
+ * in alloc_bootmem()) must all be initialized now: the folio
+ * may still be HVOed via the regular remap path (e.g. if the
+ * architecture could not determine HVO support at bootmem
+ * allocation time), which expects valid tail pages.
+ */
hugetlb_folio_init_vmemmap(folio, h,
- HUGETLB_VMEMMAP_RESERVE_PAGES);
+ pre_hvo ? HUGETLB_VMEMMAP_RESERVE_PAGES :
+ pages_per_huge_page(h));
init_new_hugetlb_folio(folio);
- if (vmemmap_optimizable_order(pfn_to_section_compound_order(pfn)))
+ if (pre_hvo)
folio_set_hugetlb_vmemmap_optimized(folio);
section_set_compound_order_range(pfn, folio_nr_pages(folio), 0);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 09/20] hugetlb_vmemmap: Allow architectures to make HVO enablement boot-time only
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (7 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 08/20] hugetlb: Fully initialize tail struct pages of non-pre-HVOed bootmem folios James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 10/20] hugetlb_vmemmap: Expose whether HVO is enabled to architecture code James Houghton
` (10 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
arm64 needs all CPUs to support hardware updates of the access flag to
safely update the vmemmap in place, and it will need to refuse to online
CPUs that lack it, but only if HVO may be used. This only works if HVO
cannot be dynamically enabled, otherwise incompatible CPUs may have
already been onlined at enable time.
Add CONFIG_ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT for architectures to
select. When selected, vmemmap_optimize_enabled is __ro_after_init, so
it can only be set with the hugetlb_free_vmemmap= kernel command line
parameter, and the sysctl is made read-only.
Signed-off-by: James Houghton <jthoughton@google.com>
---
Documentation/admin-guide/sysctl/vm.rst | 3 +++
fs/Kconfig | 3 ++-
mm/Kconfig | 9 +++++++++
mm/hugetlb_vmemmap.c | 17 +++++++++++++++--
4 files changed, 29 insertions(+), 3 deletions(-)
diff --git a/Documentation/admin-guide/sysctl/vm.rst b/Documentation/admin-guide/sysctl/vm.rst
index 5b318d17aa4b..d962d563cfd4 100644
--- a/Documentation/admin-guide/sysctl/vm.rst
+++ b/Documentation/admin-guide/sysctl/vm.rst
@@ -692,6 +692,9 @@ pages. So, those surplus pages are still optimized until they are no longer
in use. You would need to wait for those surplus pages to be released before
there are no optimized pages in the system.
+On some architectures, this knob is read-only, and HVO can only be enabled or
+disabled with the hugetlb_free_vmemmap= kernel command line parameter.
+
nr_hugepages_mempolicy
======================
diff --git a/fs/Kconfig b/fs/Kconfig
index 1454b7fe9641..4222db17ba02 100644
--- a/fs/Kconfig
+++ b/fs/Kconfig
@@ -267,7 +267,8 @@ config HUGETLB_PAGE_OPTIMIZE_VMEMMAP_DEFAULT_ON
help
The HugeTLB Vmemmap Optimization (HVO) defaults to off. Say Y here to
enable HVO by default. It can be disabled via hugetlb_free_vmemmap=off
- (boot command line) or hugetlb_optimize_vmemmap (sysctl).
+ (boot command line) or hugetlb_optimize_vmemmap (sysctl, if the
+ architecture allows changing it at runtime).
endif # HUGETLBFS
config HUGETLB_PAGE
diff --git a/mm/Kconfig b/mm/Kconfig
index acefc994d9a8..ad313fddc9da 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -475,6 +475,15 @@ config ARCH_WANT_OPTIMIZE_DAX_VMEMMAP
config ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP
bool
+#
+# Select this config option from the architecture Kconfig if HugeTLB vmemmap
+# optimization may only be enabled or disabled on the kernel command line, e.g.
+# because the architecture configures the system at boot based on whether it
+# is enabled. This makes the hugetlb_optimize_vmemmap sysctl read-only.
+#
+config ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT
+ bool
+
config HAVE_MEMBLOCK_PHYS_MAP
bool
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 817ce39d4cf3..de4148a8aabf 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -428,7 +428,20 @@ static int vmemmap_remap_alloc(const struct hstate *h, struct folio *folio,
return ret;
}
-static bool vmemmap_optimize_enabled = IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP_DEFAULT_ON);
+/*
+ * Architectures selecting CONFIG_ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT rely on
+ * HVO only being enabled or disabled on the kernel command line.
+ */
+#ifdef CONFIG_ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT
+#define __vmemmap_optimize_enabled_attr __ro_after_init
+#define HUGETLB_VMEMMAP_SYSCTL_MODE 0444
+#else
+#define __vmemmap_optimize_enabled_attr
+#define HUGETLB_VMEMMAP_SYSCTL_MODE 0644
+#endif
+
+static bool vmemmap_optimize_enabled __vmemmap_optimize_enabled_attr =
+ IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP_DEFAULT_ON);
static int __init hugetlb_vmemmap_optimize_param(char *buf)
{
return kstrtobool(buf, &vmemmap_optimize_enabled);
@@ -751,7 +764,7 @@ static const struct ctl_table hugetlb_vmemmap_sysctls[] = {
.procname = "hugetlb_optimize_vmemmap",
.data = &vmemmap_optimize_enabled,
.maxlen = sizeof(vmemmap_optimize_enabled),
- .mode = 0644,
+ .mode = HUGETLB_VMEMMAP_SYSCTL_MODE,
.proc_handler = proc_dobool,
},
};
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 10/20] hugetlb_vmemmap: Expose whether HVO is enabled to architecture code
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (8 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 09/20] hugetlb_vmemmap: Allow architectures to make HVO enablement boot-time only James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 11/20] arm64: Add bbm_through_af capability James Houghton
` (9 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Architectures that select CONFIG_ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT
may configure the system at boot depending on whether HVO is enabled,
e.g. arm64 requires all CPUs to support hardware access flag updates
only if HVO is enabled.
Add hugetlb_vmemmap_optimize_enabled() so that they can check. It is
declared before <asm/hugetlb.h> is included, so that it can also be used
from inline helpers there.
It only reflects the hugetlb_free_vmemmap= kernel command line parameter
(and the sysctl, when writable), not whether the architecture supports
HVO.
Signed-off-by: James Houghton <jthoughton@google.com>
---
include/linux/hugetlb.h | 9 +++++++++
mm/hugetlb_vmemmap.c | 18 ++++++++++++++++++
2 files changed, 27 insertions(+)
diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
index 5029c7241863..1ef73b5da664 100644
--- a/include/linux/hugetlb.h
+++ b/include/linux/hugetlb.h
@@ -22,6 +22,15 @@ struct node;
void free_huge_folio(struct folio *folio);
+#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
+bool hugetlb_vmemmap_optimize_enabled(void);
+#else
+static inline bool hugetlb_vmemmap_optimize_enabled(void)
+{
+ return false;
+}
+#endif
+
#ifdef CONFIG_HUGETLB_PAGE
#include <linux/pagemap.h>
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index de4148a8aabf..c1c736dd08b9 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -448,6 +448,24 @@ static int __init hugetlb_vmemmap_optimize_param(char *buf)
}
early_param("hugetlb_free_vmemmap", hugetlb_vmemmap_optimize_param);
+/**
+ * hugetlb_vmemmap_optimize_enabled - whether HVO is enabled
+ *
+ * This only reflects the hugetlb_free_vmemmap= kernel command line parameter
+ * and the vm.hugetlb_optimize_vmemmap sysctl, not whether the architecture
+ * supports HVO (see arch_hugetlb_vmemmap_optimization_supported()).
+ *
+ * The value is final once early parameters have been parsed only if the
+ * architecture selects CONFIG_ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT;
+ * otherwise it may change at any time through the sysctl.
+ *
+ * Return: true if HVO is enabled.
+ */
+bool hugetlb_vmemmap_optimize_enabled(void)
+{
+ return READ_ONCE(vmemmap_optimize_enabled);
+}
+
static int __hugetlb_vmemmap_restore_folio(const struct hstate *h,
struct folio *folio, unsigned long flags)
{
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 11/20] arm64: Add bbm_through_af capability
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (9 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 10/20] hugetlb_vmemmap: Expose whether HVO is enabled to architecture code James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 12/20] arm64: Implement try_update_vmemmap_pte using the AF trick James Houghton
` (8 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
HVO requires BBM Level 3 for the block -> table transitions when
optimizing the vmemmap, and it requires hardware AF management to avoid
taking page faults on the !AF translation that is installed as part of
the RW table -> RO + new OA table transitions.
Call this requirement "BBM through AF"; i.e., all CPUs support BBM-less,
in-place TTD updates. Model this as a system feature, as all CPUs must
support it for the kernel to use it.
It may seem odd to check for HVO support directly in bbm_through_af().
If this weren't done, any system with mismatched HW AF will no longer be
able to start late CPUs that are missing HW AF if all early CPUs have HW
AF (and BBML3).
Signed-off-by: James Houghton <jthoughton@google.com>
---
arch/arm64/include/asm/cpufeature.h | 5 +++++
arch/arm64/kernel/cpufeature.c | 26 ++++++++++++++++++++++++++
arch/arm64/tools/cpucaps | 1 +
3 files changed, 32 insertions(+)
diff --git a/arch/arm64/include/asm/cpufeature.h b/arch/arm64/include/asm/cpufeature.h
index 4f04ad82ea34..fbededd7b748 100644
--- a/arch/arm64/include/asm/cpufeature.h
+++ b/arch/arm64/include/asm/cpufeature.h
@@ -878,6 +878,11 @@ static inline bool system_supports_bbml3(void)
return alternative_has_cap_unlikely(ARM64_HAS_BBML3);
}
+static inline bool system_supports_bbm_through_af(void)
+{
+ return alternative_has_cap_unlikely(ARM64_HAS_BBM_THROUGH_AF);
+}
+
int do_emulate_mrs(struct pt_regs *regs, u32 sys_reg, u32 rt);
bool try_emulate_mrs(struct pt_regs *regs, u32 isn);
diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c
index 32102c3912fa..448d4ef085de 100644
--- a/arch/arm64/kernel/cpufeature.c
+++ b/arch/arm64/kernel/cpufeature.c
@@ -76,6 +76,7 @@
#include <linux/kasan.h>
#include <linux/percpu.h>
#include <linux/sched/isolation.h>
+#include <linux/hugetlb.h>
#include <asm/arm_pmuv3.h>
#include <asm/cpu.h>
@@ -2193,6 +2194,25 @@ static bool has_bbml3(const struct arm64_cpu_capabilities *caps, int scope)
return cpu_supports_bbml3();
}
+static bool has_bbm_through_af(const struct arm64_cpu_capabilities *caps, int scope)
+{
+ /*
+ * BBM-through-AF is only needed for HVO. If HVO is not in use, don't
+ * penalize systems with mismatched support.
+ */
+ if (!hugetlb_vmemmap_optimize_enabled()) {
+ BUILD_BUG_ON(!IS_ENABLED(CONFIG_ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT));
+ return false;
+ }
+
+ /*
+ * We need BBML3 to support Block -> Table transitions without taking
+ * faults, and we need HW AF support to support changing the OA without
+ * taking faults.
+ */
+ return cpu_supports_bbml3() && cpu_has_hw_af();
+}
+
static void cpu_enable_pan(const struct arm64_cpu_capabilities *__unused)
{
/*
@@ -3102,6 +3122,12 @@ static const struct arm64_cpu_capabilities arm64_features[] = {
.type = ARM64_CPUCAP_EARLY_LOCAL_CPU_FEATURE,
.matches = has_bbml3,
},
+ {
+ .desc = "BBM-less TTD updates via Access Flag",
+ .capability = ARM64_HAS_BBM_THROUGH_AF,
+ .type = ARM64_CPUCAP_SYSTEM_FEATURE,
+ .matches = has_bbm_through_af,
+ },
{
.desc = "52-bit Virtual Addressing for KVM (LPA2)",
.capability = ARM64_HAS_LPA2,
diff --git a/arch/arm64/tools/cpucaps b/arch/arm64/tools/cpucaps
index 2775ba3359cf..25a84aa8a363 100644
--- a/arch/arm64/tools/cpucaps
+++ b/arch/arm64/tools/cpucaps
@@ -15,6 +15,7 @@ HAS_ADDRESS_AUTH_IMP_DEF
HAS_AMU_EXTN
HAS_ARMv8_4_TTL
HAS_BBML3
+HAS_BBM_THROUGH_AF
HAS_CACHE_DIC
HAS_CACHE_IDC
HAS_CNP
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 12/20] arm64: Implement try_update_vmemmap_pte using the AF trick
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (10 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 11/20] arm64: Add bbm_through_af capability James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 13/20] arm64: Support hugetlb vmemmap optimization James Houghton
` (7 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
try_update_vmemmap_pte() must modify vmemmap PTEs without introducing a
time window where other CPUs on the system might fault.
Normally a break-before-make sequence is required to avoid conflicts
with cached translations. However, if we can guarantee that the existing
translation cannot be cached, a BBM sequence is not needed.
Translations with the AF unset may not be cached (see Arm ARM Rule
DWZCQ); the implementation of try_update_vmemmap_pte() on arm64 takes
advantage of this fact to replace a PTE without BBM and therefore
without leaving a window open where PE might fault on this translation.
Of course, if some CPUs on the system do not support HW AF management,
clearing the AF will introduce potential faults.
system_supports_bbm_through_af() will return false if any CPUs on the
system do not support HW AF.
Signed-off-by: James Houghton <jthoughton@google.com>
---
arch/arm64/include/asm/pgtable.h | 60 +++++++++++++++++++++++++++++---
1 file changed, 56 insertions(+), 4 deletions(-)
diff --git a/arch/arm64/include/asm/pgtable.h b/arch/arm64/include/asm/pgtable.h
index e47c3d010715..f8e66bb22c0c 100644
--- a/arch/arm64/include/asm/pgtable.h
+++ b/arch/arm64/include/asm/pgtable.h
@@ -1249,14 +1249,24 @@ static inline void __pte_clear(struct mm_struct *mm,
__set_pte(ptep, __pte(0));
}
-static inline bool __ptep_test_and_clear_young(struct vm_area_struct *vma,
- unsigned long address, pte_t *ptep)
+/*
+ * Atomically clear the Accessed flag. Return the old value of the PTE.
+ */
+static inline pte_t __ptep_clear_young(pte_t *ptep)
{
atomic64_t *pteval = (atomic64_t *)&pte_val(*ptep);
s64 af_mask = PTE_AF;
- /* Atomically clear PTE_AF, checking that it was set before. */
- return af_mask & atomic64_fetch_andnot_relaxed(af_mask, pteval);
+ /* Atomically clear PTE_AF. */
+ u64 oldval = atomic64_fetch_andnot_relaxed(af_mask, pteval);
+
+ return __pte(oldval);
+}
+
+static inline bool __ptep_test_and_clear_young(struct vm_area_struct *vma,
+ unsigned long address, pte_t *ptep)
+{
+ return pte_young(__ptep_clear_young(ptep));
}
static inline bool __ptep_clear_flush_young(struct vm_area_struct *vma,
@@ -1734,6 +1744,48 @@ static inline void pte_clear(struct mm_struct *mm,
__pte_clear(mm, addr, ptep);
}
+#define __HAVE_ARCH_TRY_UPDATE_VMEMMAP_PTE
+static inline int try_update_vmemmap_pte(unsigned long addr, pte_t *ptep,
+ const pte_t pte)
+{
+ const int max_attempts = 16;
+ int attempts = 0;
+ pte_t old_pte;
+
+ if (!system_supports_bbm_through_af())
+ return -EOPNOTSUPP;
+
+ /* This routine is only to be used for valid-to-valid transitions. */
+ if (WARN_ON_ONCE(!pte_valid(pte)))
+ return -EINVAL;
+
+ old_pte = __ptep_get(ptep);
+
+ do {
+ if (WARN_ON_ONCE(!pte_valid(old_pte)))
+ return -EINVAL;
+
+ /* We should never get a contiguous PTE here. */
+ if (WARN_ON_ONCE(pte_valid_cont(old_pte)))
+ return -EINVAL;
+
+ if (pte_young(old_pte)) {
+ /* __ptep_clear_young() returns the overwritten PTE */
+ old_pte = pte_mkold(__ptep_clear_young(ptep));
+
+ flush_tlb_kernel_range(addr, addr + PAGE_SIZE);
+ }
+ /*
+ * Translations without AF cannot be cached, so we can replace
+ * them without BBM.
+ */
+ } while (!try_cmpxchg_relaxed(&pte_val(*ptep), &pte_val(old_pte),
+ pte_val(pte)) &&
+ ++attempts < max_attempts);
+
+ return attempts == max_attempts ? -EAGAIN : 0;
+}
+
#define clear_full_ptes clear_full_ptes
static inline void clear_full_ptes(struct mm_struct *mm, unsigned long addr,
pte_t *ptep, unsigned int nr, int full)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 13/20] arm64: Support hugetlb vmemmap optimization
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (11 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 12/20] arm64: Implement try_update_vmemmap_pte using the AF trick James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 14/20] hugetlb_vmemmap: Add fault injection for in-place vmemmap PTE updates James Houghton
` (6 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Support HVO if HW AF and BBML3 are supported. These are the required
features for the vmemmap to be modified in place.
On arm64 systems that do not have the required hardware features to
support HVO, the HVO sysctl will not appear, and any attempts to enable
HVO (either via the command line or via
CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP_DEFAULT_ON=y) will have no effect.
Signed-off-by: James Houghton <jthoughton@google.com>
---
arch/arm64/Kconfig | 2 ++
arch/arm64/include/asm/hugetlb.h | 7 +++++++
mm/hugetlb_vmemmap.c | 4 ++++
3 files changed, 13 insertions(+)
diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
index b51d23a62c92..e4c87487bc7f 100644
--- a/arch/arm64/Kconfig
+++ b/arch/arm64/Kconfig
@@ -92,7 +92,9 @@ config ARM64
select ARCH_WANT_DEFAULT_TOPDOWN_MMAP_LAYOUT
select ARCH_WANT_FRAME_POINTERS
select ARCH_WANT_HUGE_PMD_SHARE if ARM64_4K_PAGES || (ARM64_16K_PAGES && !ARM64_VA_BITS_36)
+ select ARCH_WANT_HUGETLB_VMEMMAP_RO_AFTER_INIT
select ARCH_WANT_LD_ORPHAN_WARN
+ select ARCH_WANT_OPTIMIZE_HUGETLB_VMEMMAP if ARM64_HW_AFDBM
select ARCH_WANTS_EXECMEM_LATE
select ARCH_WANTS_NO_INSTR
select ARCH_WANTS_THP_SWAP if ARM64_4K_PAGES
diff --git a/arch/arm64/include/asm/hugetlb.h b/arch/arm64/include/asm/hugetlb.h
index d038ff14d16c..938d7896a31e 100644
--- a/arch/arm64/include/asm/hugetlb.h
+++ b/arch/arm64/include/asm/hugetlb.h
@@ -11,6 +11,7 @@
#define __ASM_HUGETLB_H
#include <asm/cacheflush.h>
+#include <asm/cpufeature.h>
#include <asm/mte.h>
#include <asm/page.h>
@@ -65,6 +66,12 @@ extern void huge_ptep_modify_prot_commit(struct vm_area_struct *vma,
unsigned long addr, pte_t *ptep,
pte_t old_pte, pte_t new_pte);
+#define __HAVE_ARCH_HVO_SUPPORTED
+static inline bool arch_hugetlb_vmemmap_optimization_supported(void)
+{
+ return system_supports_bbm_through_af();
+}
+
#include <asm-generic/hugetlb.h>
static inline void __flush_hugetlb_tlb_range(struct vm_area_struct *vma,
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index c1c736dd08b9..fabf2b25fe59 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -85,6 +85,10 @@ static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
/* Make pte visible before pmd. See comment in pmd_install(). */
smp_wmb();
+ /*
+ * On arm64, this requires BBML3. Its support has already been
+ * checked.
+ */
pmd_populate_kernel(&init_mm, pmd, pgtable);
if (!(walk->flags & VMEMMAP_SPLIT_NO_TLB_FLUSH))
flush_tlb_kernel_range(start, start + PMD_SIZE);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 14/20] hugetlb_vmemmap: Add fault injection for in-place vmemmap PTE updates
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (12 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 13/20] arm64: Support hugetlb vmemmap optimization James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 15/20] selftests/mm: Add HugeTLB vmemmap optimization stress test James Houghton
` (5 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
try_update_vmemmap_pte() may fail, e.g. on arm64 when the update keeps
racing with hardware access flag updates. HVO handles such failures by
rolling back the optimization, or by leaving folios partially optimized
if the rollback or a later restore fails. These paths are rare to hit
normally.
Add a fault-injection capability, fail_hugetlb_vmemmap_pte, under
CONFIG_FAIL_HUGETLB_VMEMMAP. When a fault is injected, the PTE update
fails with -EAGAIN without touching the page tables, as if the in-place
update had given up.
It can be configured through debugfs or through the
fail_hugetlb_vmemmap_pte= boot option, the latter allowing failures to
be injected when optimizing bootmem folios.
Assisted-by: LLM
Signed-off-by: James Houghton <jthoughton@google.com>
---
.../fault-injection/fault-injection.rst | 6 +++
lib/Kconfig.debug | 9 +++++
mm/hugetlb_vmemmap.c | 38 ++++++++++++++++++-
3 files changed, 51 insertions(+), 2 deletions(-)
diff --git a/Documentation/fault-injection/fault-injection.rst b/Documentation/fault-injection/fault-injection.rst
index c2d3996b5b40..403206645fa4 100644
--- a/Documentation/fault-injection/fault-injection.rst
+++ b/Documentation/fault-injection/fault-injection.rst
@@ -16,6 +16,11 @@ Available fault injection capabilities
injects page allocation failures. (alloc_pages(), get_free_pages(), ...)
+- fail_hugetlb_vmemmap_pte
+
+ injects failures of the in-place vmemmap PTE remaps done by HugeTLB vmemmap
+ optimization. (try_update_vmemmap_pte())
+
- fail_usercopy
injects failures in user memory access functions. (copy_from_user(), get_user(), ...)
@@ -263,6 +268,7 @@ use the boot option::
failslab=
fail_page_alloc=
+ fail_hugetlb_vmemmap_pte=
fail_usercopy=
fail_make_request=
fail_futex=
diff --git a/lib/Kconfig.debug b/lib/Kconfig.debug
index 134b15a44625..96cd1f1da94a 100644
--- a/lib/Kconfig.debug
+++ b/lib/Kconfig.debug
@@ -2066,6 +2066,15 @@ config FAIL_PAGE_ALLOC
help
Provide fault-injection capability for alloc_pages().
+config FAIL_HUGETLB_VMEMMAP
+ bool "Fault-injection capability for HugeTLB vmemmap optimization"
+ depends on FAULT_INJECTION && HUGETLB_PAGE_OPTIMIZE_VMEMMAP
+ help
+ Provide fault-injection capability for the in-place vmemmap page
+ table updates done by HugeTLB vmemmap optimization (HVO), i.e.
+ try_update_vmemmap_pte(). This exercises the rollback and
+ partially-optimized folio paths.
+
config FAULT_INJECTION_USERCOPY
bool "Fault injection capability for usercopy functions"
depends on FAULT_INJECTION
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index fabf2b25fe59..1eca03a3def3 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -17,6 +17,7 @@
#include <linux/pgalloc.h>
#include <linux/vmemmap-optimization.h>
#include <linux/hugetlb.h>
+#include <linux/fault-inject.h>
#include <asm/tlbflush.h>
#include "hugetlb_vmemmap.h"
@@ -50,6 +51,39 @@ struct vmemmap_remap_walk {
unsigned long flags;
};
+#ifdef CONFIG_FAIL_HUGETLB_VMEMMAP
+static DECLARE_FAULT_ATTR(fail_hugetlb_vmemmap_pte);
+
+static int __init setup_fail_hugetlb_vmemmap_pte(char *str)
+{
+ return setup_fault_attr(&fail_hugetlb_vmemmap_pte, str);
+}
+__setup("fail_hugetlb_vmemmap_pte=", setup_fail_hugetlb_vmemmap_pte);
+
+#ifdef CONFIG_FAULT_INJECTION_DEBUG_FS
+static int __init fail_hugetlb_vmemmap_debugfs(void)
+{
+ fault_create_debugfs_attr("fail_hugetlb_vmemmap_pte", NULL,
+ &fail_hugetlb_vmemmap_pte);
+ return 0;
+}
+late_initcall(fail_hugetlb_vmemmap_debugfs);
+#endif /* CONFIG_FAULT_INJECTION_DEBUG_FS */
+
+/*
+ * Inject failures as if the in-place update lost a race too many times
+ * (see the arm64 implementations), without touching the page tables.
+ */
+static int hvo_update_vmemmap_pte(unsigned long addr, pte_t *ptep, pte_t pte)
+{
+ if (should_fail(&fail_hugetlb_vmemmap_pte, PAGE_SIZE))
+ return -EAGAIN;
+ return try_update_vmemmap_pte(addr, ptep, pte);
+}
+#else
+#define hvo_update_vmemmap_pte try_update_vmemmap_pte
+#endif /* CONFIG_FAIL_HUGETLB_VMEMMAP */
+
static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
struct vmemmap_remap_walk *walk)
{
@@ -235,7 +269,7 @@ static int vmemmap_remap_pte(pte_t *pte, unsigned long addr,
entry = mk_pte(walk->vmemmap_tail, PAGE_KERNEL_RO);
}
- ret = try_update_vmemmap_pte(addr, pte, entry);
+ ret = hvo_update_vmemmap_pte(addr, pte, entry);
if (ret)
return ret;
@@ -279,7 +313,7 @@ static int vmemmap_restore_pte(pte_t *pte, unsigned long addr,
*/
smp_wmb();
- ret = try_update_vmemmap_pte(addr, pte, mk_pte(dst, PAGE_KERNEL));
+ ret = hvo_update_vmemmap_pte(addr, pte, mk_pte(dst, PAGE_KERNEL));
if (ret)
return ret;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 15/20] selftests/mm: Add HugeTLB vmemmap optimization stress test
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (13 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 14/20] hugetlb_vmemmap: Add fault injection for in-place vmemmap PTE updates James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 16/20 DO-NOT-MERGE] hugetlb_vmemmap: Use try_populate_vmemmap_pmd for replacing in-use PMDs James Houghton
` (4 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Add a script that stresses HugeTLB vmemmap optimization (HVO), in
particular on architectures that update the vmemmap in place while it
may be concurrently accessed:
- Check that optimizing N folios frees exactly N * (vmemmap pages per
folio - 1) vmemmap pages (per nr_memmap_pages and
nr_memmap_boot_pages), and that restoring them gives them back.
- Repeatedly optimize and restore folios by resizing the hugepage pool,
while concurrently reading struct pages through /proc/kpageflags and
/proc/kpagecount, compacting memory, and optionally reading
page_owner and offlining/onlining memory blocks.
- Do the same with fail_hugetlb_vmemmap_pte fault injection enabled,
if available, to exercise the rollback and partially-optimized folio
paths.
After each phase, the pool must shrink back to its original size and the
memmap accounting must return to its baseline. Finally, the kernel log
must not contain warnings or oopses.
Assisted-by: LLM
Signed-off-by: James Houghton <jthoughton@google.com>
---
tools/testing/selftests/mm/Makefile | 2 +
.../selftests/mm/hugetlb_vmemmap_stress.sh | 347 ++++++++++++++++++
.../selftests/mm/ksft_hugetlb_vmemmap.sh | 4 +
tools/testing/selftests/mm/run_vmtests.sh | 4 +
4 files changed, 357 insertions(+)
create mode 100755 tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh
create mode 100755 tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh
diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/mm/Makefile
index beacc0f87304..51d8fba80c03 100644
--- a/tools/testing/selftests/mm/Makefile
+++ b/tools/testing/selftests/mm/Makefile
@@ -149,6 +149,7 @@ TEST_PROGS += ksft_cow.sh
TEST_PROGS += ksft_gup_test.sh
TEST_PROGS += ksft_hmm.sh
TEST_PROGS += ksft_hugetlb.sh
+TEST_PROGS += ksft_hugetlb_vmemmap.sh
TEST_PROGS += ksft_hugevm.sh
TEST_PROGS += ksft_kmemleak_confirm.sh
TEST_PROGS += ksft_kmemleak_dedup.sh
@@ -180,6 +181,7 @@ TEST_FILES += test_hmm.sh
TEST_FILES += va_high_addr_switch.sh
TEST_FILES += charge_reserved_hugetlb.sh
TEST_FILES += hugetlb_reparenting_test.sh
+TEST_FILES += hugetlb_vmemmap_stress.sh
TEST_FILES += test_page_frag.sh
TEST_FILES += run_vmtests.sh
diff --git a/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh b/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh
new file mode 100755
index 000000000000..94357a4a45c2
--- /dev/null
+++ b/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh
@@ -0,0 +1,347 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+#
+# Stress test for HugeTLB vmemmap optimization (HVO).
+#
+# Phases:
+# 1. accounting: allocate N hugepages, check that nr_memmap_pages +
+# nr_memmap_boot_pages drops by exactly N * (vmemmap pages per folio - 1),
+# then free them and check it returns to the baseline.
+# 2. stress: churn the pool (optimize/restore) while concurrently reading
+# struct pages (/proc/kpageflags, /proc/kpagecount), compacting memory,
+# and optionally reading page_owner and offlining/onlining memory.
+# 3. pte-inject: like 2, with fail_hugetlb_vmemmap_pte enabled; this hits
+# both the optimize rollback and the restore (partial HVO) paths.
+#
+# After each phase the pool is drained back to its original size, memmap
+# accounting must be back at the phase's baseline, and the kernel log must
+# not contain warnings/oopses.
+#
+# The fault-injection phases require CONFIG_FAIL_HUGETLB_VMEMMAP and
+# CONFIG_FAULT_INJECTION_DEBUG_FS, and are skipped otherwise.
+
+set -u
+
+KSFT_PASS=0
+KSFT_FAIL=1
+KSFT_SKIP=4
+
+size_kb=
+nr=16
+duration=60
+prob=20
+readers=4
+struct_page_size=64
+do_page_owner=0
+do_hotplug=0
+
+usage() {
+ cat <<EOF
+Usage: $0 [options]
+ -s KB hugepage size in kB (default: default hugepage size)
+ -n N number of hugepages to churn (default: $nr)
+ -t SEC duration of each stress phase (default: $duration)
+ -p PCT fault-injection probability in percent (default: $prob)
+ -j N number of struct page reader processes (default: $readers)
+ -S BYTES sizeof(struct page) (default: $struct_page_size)
+ -o also read /sys/kernel/debug/page_owner during stress
+ -m also offline/online memory blocks during stress (dissolves
+ free hugepages, exercising the restore path)
+
+For exact accounting checks, run on a hugepage size whose pool is
+initially empty (e.g. no boot-time reservations of that size).
+EOF
+ exit $KSFT_SKIP
+}
+
+while getopts "s:n:t:p:j:S:omh" opt; do
+ case $opt in
+ s) size_kb=$OPTARG ;;
+ n) nr=$OPTARG ;;
+ t) duration=$OPTARG ;;
+ p) prob=$OPTARG ;;
+ j) readers=$OPTARG ;;
+ S) struct_page_size=$OPTARG ;;
+ o) do_page_owner=1 ;;
+ m) do_hotplug=1 ;;
+ *) usage ;;
+ esac
+done
+
+log() { echo "# $*"; }
+skip() { echo "SKIP: $*"; exit $KSFT_SKIP; }
+
+failures=0
+fail() { echo "FAIL: $*"; failures=$((failures + 1)); }
+pass() { echo "PASS: $*"; }
+
+[ "$(id -u)" -eq 0 ] || skip "must be run as root"
+
+[ -r /proc/sys/vm/hugetlb_optimize_vmemmap ] ||
+ skip "HVO not supported (no vm.hugetlb_optimize_vmemmap)"
+[ "$(cat /proc/sys/vm/hugetlb_optimize_vmemmap)" = 1 ] ||
+ skip "HVO disabled (vm.hugetlb_optimize_vmemmap=0)"
+grep -q '^nr_memmap_pages ' /proc/vmstat ||
+ skip "no nr_memmap_pages in /proc/vmstat"
+
+[ -n "$size_kb" ] || size_kb=$(awk '/^Hugepagesize:/ {print $2}' /proc/meminfo)
+hp_dir=/sys/kernel/mm/hugepages/hugepages-${size_kb}kB
+[ -d "$hp_dir" ] || skip "no ${size_kb}kB hugepages"
+
+page_size=$(getconf PAGESIZE)
+vmemmap_pages=$(( size_kb * 1024 / page_size * struct_page_size / page_size ))
+freed_per_folio=$(( vmemmap_pages - 1 ))
+[ "$freed_per_folio" -gt 0 ] ||
+ skip "${size_kb}kB hugepages are not HVO-optimizable"
+
+dbgfs=/sys/kernel/debug
+mountpoint -q $dbgfs || mount -t debugfs none $dbgfs 2>/dev/null
+
+orig_nr=$(cat "$hp_dir/nr_hugepages")
+target_nr=$(( orig_nr + nr ))
+marker="hvo-stress-$$-$(date +%s)"
+tmpdir=$(mktemp -d)
+pids=()
+
+log "hugepage size ${size_kb}kB, page size $page_size"
+log "struct page size $struct_page_size"
+log "vmemmap pages/folio: $vmemmap_pages ($freed_per_folio freed by HVO)"
+log "pool: $orig_nr initially, churning between $orig_nr and $target_nr"
+[ "$orig_nr" -eq 0 ] ||
+ log "WARNING: pool not initially empty, accounting may be inexact"
+
+memmap_total() {
+ awk '/^nr_memmap_(boot_)?pages / {s += $2} END {print s + 0}' \
+ /proc/vmstat
+}
+
+set_nr() {
+ echo "$1" > "$hp_dir/nr_hugepages" 2>/dev/null
+ cat "$hp_dir/nr_hugepages"
+}
+
+# Shrink the pool back to orig_nr. Restore may transiently fail (e.g. with
+# fault injection active), so retry for a while.
+drain() {
+ local i cur surplus
+
+ for i in $(seq 20); do
+ # Pages whose vmemmap could not be restored are kept as free
+ # surplus pages, which shrinking nr_hugepages does not free.
+ # Writing the current size converts them back to persistent
+ # pages first.
+ set_nr "$(cat "$hp_dir/nr_hugepages")" > /dev/null
+ cur=$(set_nr "$orig_nr")
+ [ "$cur" -eq "$orig_nr" ] && return 0
+ sleep 1
+ done
+ surplus=$(cat "$hp_dir/surplus_hugepages")
+ fail "pool stuck at $cur hugepages ($surplus surplus)," \
+ "expected $orig_nr"
+ return 1
+}
+
+fa_dir() { echo "$dbgfs/$1"; }
+
+fa_enable() {
+ local d
+ d=$(fa_dir "$1")
+ echo 0 > "$d/verbose"
+ echo N > "$d/task-filter"
+ echo 1 > "$d/interval"
+ echo 1000000 > "$d/times"
+ echo "$prob" > "$d/probability"
+}
+
+# Disable and print the number of injected failures.
+fa_disable() {
+ local d left
+ d=$(fa_dir "$1")
+ echo 0 > "$d/probability"
+ left=$(cat "$d/times")
+ echo 0 > "$d/times"
+ echo $(( 1000000 - left ))
+}
+
+check_dmesg() {
+ local bad pat
+
+ pat='WARNING:|BUG[: ]|Oops|Unable to handle kernel'
+ pat+='|Internal error|KASAN:|UBSAN:'
+ pat+='|list_(add|del) corruption|page dumped because'
+ bad=$(dmesg | sed -n "/$marker/,\$p" | grep -E "$pat")
+ if [ -n "$bad" ]; then
+ fail "kernel log reports problems:"
+ echo "$bad" | head -20 | sed 's/^/# /'
+ fi
+}
+
+## Workers
+
+churn() {
+ while :; do
+ set_nr "$target_nr" > /dev/null
+ set_nr "$orig_nr" > /dev/null
+ done
+}
+
+kpage_reader() {
+ while :; do
+ dd if=/proc/kpageflags of=/dev/null bs=4M status=none
+ dd if=/proc/kpagecount of=/dev/null bs=4M status=none
+ done
+}
+
+compactor() {
+ while :; do
+ echo 1 > /proc/sys/vm/compact_memory
+ sleep 1
+ done
+}
+
+page_owner_reader() {
+ while :; do
+ cat $dbgfs/page_owner > /dev/null
+ done
+}
+
+hotplugger() {
+ local blk state removable
+
+ while :; do
+ for blk in /sys/devices/system/memory/memory*; do
+ state=$(cat "$blk/state" 2>/dev/null)
+ removable=$(cat "$blk/removable" 2>/dev/null || echo 1)
+ [ "$state" = online ] || continue
+ [ "$removable" = 1 ] || continue
+ echo "$blk" >> "$tmpdir/hotplug"
+ # A signal aborts a pending offline_pages().
+ timeout 10 sh -c "echo offline > $blk/state" 2>/dev/null
+ echo online > "$blk/state" 2>/dev/null
+ sleep 1
+ done
+ done
+}
+
+start_workers() {
+ local i
+
+ churn & pids+=($!)
+ for i in $(seq "$readers"); do
+ kpage_reader & pids+=($!)
+ done
+ compactor & pids+=($!)
+ if [ "$do_page_owner" -eq 1 ]; then
+ if [ -r $dbgfs/page_owner ]; then
+ page_owner_reader & pids+=($!)
+ else
+ log "page_owner not available, not reading it"
+ fi
+ fi
+ [ "$do_hotplug" -eq 1 ] && { hotplugger & pids+=($!); }
+}
+
+stop_workers() {
+ [ "${#pids[@]}" -gt 0 ] || return 0
+ kill "${pids[@]}" 2>/dev/null
+ wait "${pids[@]}" 2>/dev/null
+ pids=()
+}
+
+cleanup() {
+ local blk t
+
+ stop_workers
+ for t in fail_hugetlb_vmemmap_pte; do
+ [ -d "$(fa_dir $t)" ] && fa_disable $t > /dev/null
+ done
+ if [ -f "$tmpdir/hotplug" ]; then
+ sort -u "$tmpdir/hotplug" | while read -r blk; do
+ [ "$(cat "$blk/state")" = online ] ||
+ echo online > "$blk/state" 2>/dev/null
+ done
+ fi
+ set_nr "$orig_nr" > /dev/null
+ rm -rf "$tmpdir"
+}
+trap cleanup EXIT
+trap 'exit $KSFT_FAIL' INT TERM
+
+# Run a stress phase. $1: name, $2: fault attr to enable ("" for none).
+stress_phase() {
+ local name=$1 fa=$2 base after injected
+
+ if [ -n "$fa" ] && [ ! -d "$(fa_dir "$fa")" ]; then
+ echo "SKIP: $name (no $(fa_dir "$fa"))"
+ return
+ fi
+
+ log "phase $name: ${duration}s"
+ base=$(memmap_total)
+ [ -n "$fa" ] && fa_enable "$fa"
+ start_workers
+ sleep "$duration"
+ stop_workers
+ if [ -n "$fa" ]; then
+ injected=$(fa_disable "$fa")
+ log "$name: injected $injected failures"
+ [ "$injected" -gt 0 ] ||
+ log "WARNING: $name: no failures injected"
+ fi
+
+ drain || return
+ after=$(memmap_total)
+ if [ "$after" -ne "$base" ]; then
+ fail "$name: memmap pages $after after drain, expected $base"
+ else
+ pass "$name"
+ fi
+}
+
+accounting_phase() {
+ local base got added after expect
+
+ log "phase accounting"
+ base=$(memmap_total)
+ got=$(set_nr "$target_nr")
+ added=$(( got - orig_nr ))
+ if [ "$added" -le 0 ]; then
+ fail "accounting: could not allocate any ${size_kb}kB hugepages"
+ return
+ fi
+ [ "$added" -eq "$nr" ] || log "accounting: only allocated $added of $nr"
+
+ after=$(memmap_total)
+ expect=$(( base - added * freed_per_folio ))
+ if [ "$after" -ne "$expect" ]; then
+ fail "accounting: memmap pages $after after optimizing" \
+ "$added folios, expected $expect (baseline $base)"
+ else
+ pass "accounting: optimize freed $(( base - after ))" \
+ "vmemmap pages"
+ fi
+
+ drain || return
+ after=$(memmap_total)
+ if [ "$after" -ne "$base" ]; then
+ fail "accounting: memmap pages $after after restore," \
+ "expected $base"
+ else
+ pass "accounting: restore"
+ fi
+}
+
+echo "$marker" > /dev/kmsg
+
+accounting_phase
+stress_phase stress ""
+stress_phase pte-inject fail_hugetlb_vmemmap_pte
+
+check_dmesg
+
+if [ "$failures" -ne 0 ]; then
+ echo "FAILED: $failures check(s)"
+ exit $KSFT_FAIL
+fi
+echo "OK"
+exit $KSFT_PASS
diff --git a/tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh b/tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh
new file mode 100755
index 000000000000..905b75b2cfb4
--- /dev/null
+++ b/tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh
@@ -0,0 +1,4 @@
+#!/bin/sh -e
+# SPDX-License-Identifier: GPL-2.0
+
+./run_vmtests.sh -t hugetlb_vmemmap
diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/selftests/mm/run_vmtests.sh
index a1b45a3dedae..8dcdee7be501 100755
--- a/tools/testing/selftests/mm/run_vmtests.sh
+++ b/tools/testing/selftests/mm/run_vmtests.sh
@@ -77,6 +77,8 @@ separated by spaces:
test transparent huge pages
- hugetlb
test hugetlbfs huge pages
+- hugetlb_vmemmap
+ test the hugetlb vmemmap optimization
- migration
invoke move_pages(2) to exercise the migration entry code
paths in the kernel
@@ -312,6 +314,8 @@ echo "$enable_soft_offline" > /proc/sys/vm/enable_soft_offline
CATEGORY="hugetlb" run_test ./hugetlb-read-hwpoison
fi
+CATEGORY="hugetlb_vmemmap" run_test ./hugetlb_vmemmap_stress.sh
+
if [ $VADDR64 -ne 0 ]; then
# va high address boundary switch test
CATEGORY="hugevm" run_test bash ./va_high_addr_switch.sh
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 16/20 DO-NOT-MERGE] hugetlb_vmemmap: Use try_populate_vmemmap_pmd for replacing in-use PMDs
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (14 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 15/20] selftests/mm: Add HugeTLB vmemmap optimization stress test James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 17/20 DO-NOT-MERGE] arm64: Implement try_populate_vmemmap_pmd using AF trick James Houghton
` (3 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
This routine is to be used for updating in-use, leaf-level PMDs in the
vmemmap without introducing a window where vmemmap accesses might fault.
Because try_populate_vmemmap_pmd() can fail, the split_page() call needs
to be moved to after the PMD update. Because both split_page() and the
PMD update are done under the page table lock, this rearrangement is
safe.
Only leaf-level PMD to non-leaf PMD (mapping the same physical pages) is
the only transition that needs to be supported.
Collapsing a non-leaf PMD to a leaf-level PMD is not used by HVO, so
try_populate_vmemmap_pmd() need not support it.
This patch allows arm64's implementation to replaced with one that does
not require BBML3.
Signed-off-by: James Houghton <jthoughton@google.com>
---
arch/arm64/include/asm/pgalloc.h | 8 ++++++++
arch/loongarch/include/asm/pgalloc.h | 2 ++
arch/riscv/include/asm/pgalloc.h | 2 ++
arch/x86/include/asm/pgalloc.h | 2 ++
include/linux/pgalloc.h | 21 +++++++++++++++++++++
mm/hugetlb_vmemmap.c | 25 ++++++++++++++-----------
6 files changed, 49 insertions(+), 11 deletions(-)
diff --git a/arch/arm64/include/asm/pgalloc.h b/arch/arm64/include/asm/pgalloc.h
index 1b4509d3382c..0263329fb7b3 100644
--- a/arch/arm64/include/asm/pgalloc.h
+++ b/arch/arm64/include/asm/pgalloc.h
@@ -121,4 +121,12 @@ pmd_populate(struct mm_struct *mm, pmd_t *pmdp, pgtable_t ptep)
PMD_TYPE_TABLE | PMD_TABLE_AF | PMD_TABLE_PXN);
}
+static inline int try_populate_vmemmap_pmd(unsigned long addr, pmd_t *pmdp,
+ pte_t *pgtable)
+{
+ /* BBLM3 is required. Its presence has been checked. */
+ pmd_populate_kernel(&init_mm, pmdp, pgtable);
+ return 0;
+}
+
#endif
diff --git a/arch/loongarch/include/asm/pgalloc.h b/arch/loongarch/include/asm/pgalloc.h
index 248f62d0b590..0553addb0afa 100644
--- a/arch/loongarch/include/asm/pgalloc.h
+++ b/arch/loongarch/include/asm/pgalloc.h
@@ -24,6 +24,8 @@ static inline void pmd_populate(struct mm_struct *mm, pmd_t *pmd, pgtable_t pte)
set_pmd(pmd, __pmd((unsigned long)page_address(pte)));
}
+#define ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PMD
+
#ifndef __PAGETABLE_PMD_FOLDED
static inline void pud_populate(struct mm_struct *mm, pud_t *pud, pmd_t *pmd)
diff --git a/arch/riscv/include/asm/pgalloc.h b/arch/riscv/include/asm/pgalloc.h
index 770ce18a7328..b9b4b25fc6f0 100644
--- a/arch/riscv/include/asm/pgalloc.h
+++ b/arch/riscv/include/asm/pgalloc.h
@@ -31,6 +31,8 @@ static inline void pmd_populate(struct mm_struct *mm,
set_pmd(pmd, __pmd((pfn << _PAGE_PFN_SHIFT) | _PAGE_TABLE));
}
+#define ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PMD
+
#ifndef __PAGETABLE_PMD_FOLDED
static inline void pud_populate(struct mm_struct *mm, pud_t *pud, pmd_t *pmd)
{
diff --git a/arch/x86/include/asm/pgalloc.h b/arch/x86/include/asm/pgalloc.h
index c88691b15f3c..f8d24b97cbf5 100644
--- a/arch/x86/include/asm/pgalloc.h
+++ b/arch/x86/include/asm/pgalloc.h
@@ -82,6 +82,8 @@ static inline void pmd_populate(struct mm_struct *mm, pmd_t *pmd,
set_pmd(pmd, __pmd(((pteval_t)pfn << PAGE_SHIFT) | _PAGE_TABLE));
}
+#define ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PMD
+
#if CONFIG_PGTABLE_LEVELS > 2
extern void ___pmd_free_tlb(struct mmu_gather *tlb, pmd_t *pmd);
diff --git a/include/linux/pgalloc.h b/include/linux/pgalloc.h
index 9174fa59bbc5..1dbecb748028 100644
--- a/include/linux/pgalloc.h
+++ b/include/linux/pgalloc.h
@@ -26,4 +26,25 @@
arch_sync_kernel_mappings(addr, addr); \
} while (0)
+#ifdef ARCH_WANTS_GENERIC_POPULATE_VMEMMAP_PMD
+/*
+ * try_populate_vmemmap_pmd - Populate a PMD that is in use by the vmemmap.
+ * @addr: Base address of the remapped PMD.
+ * @pmdp: Page table pointer to be overwritten.
+ * @pgtable: Pointer to the page table that the new PMD will point to.
+ *
+ * This function is only to be used to update PMDs that map the vmemmap to
+ * point to a page of already-populated PTEs that map the same pages.
+ *
+ * Implementations of this function must ensure that, while the update is taking
+ * place, CPUs will not fault on the remapped virtual address range.
+ */
+static inline int try_populate_vmemmap_pmd(unsigned long addr, pmd_t *pmdp,
+ pte_t *pgtable)
+{
+ pmd_populate_kernel(&init_mm, pmdp, pgtable);
+ return 0;
+}
+#endif
+
#endif /* _LINUX_PGALLOC_H */
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 1eca03a3def3..9e0f52474bdb 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -88,6 +88,7 @@ static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
struct vmemmap_remap_walk *walk)
{
pmd_t __pmd;
+ int ret;
int i;
unsigned long addr = start;
pte_t *pgtable;
@@ -107,8 +108,15 @@ static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
set_pte_at(&init_mm, addr, pte, entry);
}
+ ret = 0;
spin_lock(&init_mm.page_table_lock);
if (likely(pmd_leaf(*pmd))) {
+ /* Make pte visible before pmd. See comment in pmd_install(). */
+ smp_wmb();
+ ret = try_populate_vmemmap_pmd(start, pmd, pgtable);
+ if (ret)
+ goto free;
+
/*
* Higher order allocations from buddy allocator must be able to
* be treated as independent small pages (as they can be freed
@@ -117,21 +125,16 @@ static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
if (!PageReserved(head))
split_page(head, get_order(PMD_SIZE));
- /* Make pte visible before pmd. See comment in pmd_install(). */
- smp_wmb();
- /*
- * On arm64, this requires BBML3. Its support has already been
- * checked.
- */
- pmd_populate_kernel(&init_mm, pmd, pgtable);
if (!(walk->flags & VMEMMAP_SPLIT_NO_TLB_FLUSH))
flush_tlb_kernel_range(start, start + PMD_SIZE);
- } else {
- pte_free_kernel(&init_mm, pgtable);
+ goto out;
}
- spin_unlock(&init_mm.page_table_lock);
- return 0;
+free:
+ pte_free_kernel(&init_mm, pgtable);
+out:
+ spin_unlock(&init_mm.page_table_lock);
+ return ret;
}
static int vmemmap_pmd_entry(pmd_t *pmd, unsigned long addr,
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 17/20 DO-NOT-MERGE] arm64: Implement try_populate_vmemmap_pmd using AF trick
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (15 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 16/20 DO-NOT-MERGE] hugetlb_vmemmap: Use try_populate_vmemmap_pmd for replacing in-use PMDs James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 18/20 DO-NOT-MERGE] arm64: Drop BBML3 requirement for HVO James Houghton
` (2 subsequent siblings)
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
This routine is to be used for updating in-use, leaf-level PMDs in the
vmemmap without introducing a window where vmemmap accesses might fault.
Implementing this on arm64 requires some care: use the same access flag
trick that is used for vmemmap PTE updates. HAFT is not needed and the
TLB flushing routine remains identical, as we are not overwriting a
non-leaf PMD.
For systems that support BBML3, there is no need to use the AF
trick.
Signed-off-by: James Houghton <jthoughton@google.com>
---
This patch is the main questionable piece of this "DO-NOT-MERGE" part of
the series. It's unclear if the Arm ARM truly requires hardware
implementations not to cache anything about AF=0 Block translations.
---
arch/arm64/include/asm/pgalloc.h | 51 ++++++++++++++++++++++++++++++--
1 file changed, 48 insertions(+), 3 deletions(-)
diff --git a/arch/arm64/include/asm/pgalloc.h b/arch/arm64/include/asm/pgalloc.h
index 0263329fb7b3..7e304f01ce6d 100644
--- a/arch/arm64/include/asm/pgalloc.h
+++ b/arch/arm64/include/asm/pgalloc.h
@@ -124,9 +124,54 @@ pmd_populate(struct mm_struct *mm, pmd_t *pmdp, pgtable_t ptep)
static inline int try_populate_vmemmap_pmd(unsigned long addr, pmd_t *pmdp,
pte_t *pgtable)
{
- /* BBLM3 is required. Its presence has been checked. */
- pmd_populate_kernel(&init_mm, pmdp, pgtable);
- return 0;
+ const int max_attempts = 16;
+ int attempts = 0;
+ pmd_t old_pmd, new_pmd;
+
+ if (!system_supports_bbm_through_af())
+ return -EOPNOTSUPP;
+
+ if (system_supports_bbml3()) {
+ /*
+ * BBML3 allows block->table transitions if the PTEs underneath
+ * do not conflict with existing, potentially cached
+ * translations.
+ */
+ pmd_populate_kernel(&init_mm, pmdp, pgtable);
+ return 0;
+ }
+
+ new_pmd = __pmd(__phys_to_pmd_val(__pa(pgtable)) |
+ PMD_TYPE_TABLE | PMD_TABLE_AF | PMD_TABLE_UXN);
+
+ old_pmd = pmdp_get(pmdp);
+
+ do {
+ if (WARN_ON_ONCE(!pmd_valid(old_pmd)))
+ return -EINVAL;
+
+ if (WARN_ON_ONCE(!pmd_leaf(old_pmd)))
+ return -EINVAL;
+
+ /* We should never get a contiguous PMD here. */
+ if (WARN_ON_ONCE(pmd_cont(old_pmd)))
+ return -EINVAL;
+
+ if (pmd_young(old_pmd)) {
+ /* __ptep_clear_young() returns the overwritten PTE */
+ old_pmd = pte_pmd(pte_mkold(__ptep_clear_young((pte_t *)pmdp)));
+
+ flush_tlb_kernel_range(addr, addr + PMD_SIZE);
+ }
+ /*
+ * Translations without AF cannot be cached, so we can replace
+ * them without BBM.
+ */
+ } while (!try_cmpxchg_relaxed(&pmd_val(*pmdp), &pmd_val(old_pmd),
+ pmd_val(new_pmd)) &&
+ ++attempts < max_attempts);
+
+ return attempts == max_attempts ? -EAGAIN : 0;
}
#endif
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 18/20 DO-NOT-MERGE] arm64: Drop BBML3 requirement for HVO
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (16 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 17/20 DO-NOT-MERGE] arm64: Implement try_populate_vmemmap_pmd using AF trick James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 19/20 DO-NOT-MERGE] hugetlb_vmemmap: Add fault injection for in-place vmemmap PMD splits James Houghton
2026-10-03 0:21 ` [PATCH v2 20/20 DO-NOT-MERGE] selftests/mm: Add HVO pmd-split fault injection tests James Houghton
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
With PMD updates being done with the Access Flag trick, BBML3 is no
longer needed.
Dropping this requirement vastly widens the pool of systems this
optimization is valid for.
Signed-off-by: James Houghton <jthoughton@google.com>
---
arch/arm64/kernel/cpufeature.c | 7 +++----
1 file changed, 3 insertions(+), 4 deletions(-)
diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c
index 448d4ef085de..f3fd6d0b0e41 100644
--- a/arch/arm64/kernel/cpufeature.c
+++ b/arch/arm64/kernel/cpufeature.c
@@ -2206,11 +2206,10 @@ static bool has_bbm_through_af(const struct arm64_cpu_capabilities *caps, int sc
}
/*
- * We need BBML3 to support Block -> Table transitions without taking
- * faults, and we need HW AF support to support changing the OA without
- * taking faults.
+ * We need HW AF support to support changing the vmemmap mapping level
+ * and OA without taking faults.
*/
- return cpu_supports_bbml3() && cpu_has_hw_af();
+ return cpu_has_hw_af();
}
static void cpu_enable_pan(const struct arm64_cpu_capabilities *__unused)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 19/20 DO-NOT-MERGE] hugetlb_vmemmap: Add fault injection for in-place vmemmap PMD splits
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (17 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 18/20 DO-NOT-MERGE] arm64: Drop BBML3 requirement for HVO James Houghton
@ 2026-10-03 0:21 ` James Houghton
2026-10-03 0:21 ` [PATCH v2 20/20 DO-NOT-MERGE] selftests/mm: Add HVO pmd-split fault injection tests James Houghton
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Like try_update_vmemmap_pte(), try_populate_vmemmap_pmd() may fail,
e.g. on arm64 when replacing the block mapping keeps racing with
hardware access flag updates. HVO handles such failures by not
optimizing the folio.
Add a fault-injection capability, fail_hugetlb_vmemmap_pmd, alongside
fail_hugetlb_vmemmap_pte. When a fault is injected, the PMD split fails
with -EAGAIN without touching the page tables.
It can be configured through debugfs or through the
fail_hugetlb_vmemmap_pmd= boot option. Since vmemmap PMDs are only split
once, the boot option is the most effective way to exercise this path.
Assisted-by: LLM
Signed-off-by: James Houghton <jthoughton@google.com>
---
.../fault-injection/fault-injection.rst | 8 +++++---
lib/Kconfig.debug | 4 ++--
mm/hugetlb_vmemmap.c | 20 ++++++++++++++++++-
3 files changed, 26 insertions(+), 6 deletions(-)
diff --git a/Documentation/fault-injection/fault-injection.rst b/Documentation/fault-injection/fault-injection.rst
index 403206645fa4..5cfabea811da 100644
--- a/Documentation/fault-injection/fault-injection.rst
+++ b/Documentation/fault-injection/fault-injection.rst
@@ -16,10 +16,11 @@ Available fault injection capabilities
injects page allocation failures. (alloc_pages(), get_free_pages(), ...)
-- fail_hugetlb_vmemmap_pte
+- fail_hugetlb_vmemmap_pte, fail_hugetlb_vmemmap_pmd
- injects failures of the in-place vmemmap PTE remaps done by HugeTLB vmemmap
- optimization. (try_update_vmemmap_pte())
+ injects failures of the in-place vmemmap PTE remaps and PMD splits done by
+ HugeTLB vmemmap optimization. (try_update_vmemmap_pte(),
+ try_populate_vmemmap_pmd())
- fail_usercopy
@@ -269,6 +270,7 @@ use the boot option::
failslab=
fail_page_alloc=
fail_hugetlb_vmemmap_pte=
+ fail_hugetlb_vmemmap_pmd=
fail_usercopy=
fail_make_request=
fail_futex=
diff --git a/lib/Kconfig.debug b/lib/Kconfig.debug
index 96cd1f1da94a..7abda377c737 100644
--- a/lib/Kconfig.debug
+++ b/lib/Kconfig.debug
@@ -2072,8 +2072,8 @@ config FAIL_HUGETLB_VMEMMAP
help
Provide fault-injection capability for the in-place vmemmap page
table updates done by HugeTLB vmemmap optimization (HVO), i.e.
- try_update_vmemmap_pte(). This exercises the rollback and
- partially-optimized folio paths.
+ try_update_vmemmap_pte() and try_populate_vmemmap_pmd(). This
+ exercises the rollback and partially-optimized folio paths.
config FAULT_INJECTION_USERCOPY
bool "Fault injection capability for usercopy functions"
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 9e0f52474bdb..f50880dd79a2 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -53,6 +53,7 @@ struct vmemmap_remap_walk {
#ifdef CONFIG_FAIL_HUGETLB_VMEMMAP
static DECLARE_FAULT_ATTR(fail_hugetlb_vmemmap_pte);
+static DECLARE_FAULT_ATTR(fail_hugetlb_vmemmap_pmd);
static int __init setup_fail_hugetlb_vmemmap_pte(char *str)
{
@@ -60,11 +61,19 @@ static int __init setup_fail_hugetlb_vmemmap_pte(char *str)
}
__setup("fail_hugetlb_vmemmap_pte=", setup_fail_hugetlb_vmemmap_pte);
+static int __init setup_fail_hugetlb_vmemmap_pmd(char *str)
+{
+ return setup_fault_attr(&fail_hugetlb_vmemmap_pmd, str);
+}
+__setup("fail_hugetlb_vmemmap_pmd=", setup_fail_hugetlb_vmemmap_pmd);
+
#ifdef CONFIG_FAULT_INJECTION_DEBUG_FS
static int __init fail_hugetlb_vmemmap_debugfs(void)
{
fault_create_debugfs_attr("fail_hugetlb_vmemmap_pte", NULL,
&fail_hugetlb_vmemmap_pte);
+ fault_create_debugfs_attr("fail_hugetlb_vmemmap_pmd", NULL,
+ &fail_hugetlb_vmemmap_pmd);
return 0;
}
late_initcall(fail_hugetlb_vmemmap_debugfs);
@@ -80,8 +89,17 @@ static int hvo_update_vmemmap_pte(unsigned long addr, pte_t *ptep, pte_t pte)
return -EAGAIN;
return try_update_vmemmap_pte(addr, ptep, pte);
}
+
+static int hvo_populate_vmemmap_pmd(unsigned long addr, pmd_t *pmdp,
+ pte_t *pgtable)
+{
+ if (should_fail(&fail_hugetlb_vmemmap_pmd, PMD_SIZE))
+ return -EAGAIN;
+ return try_populate_vmemmap_pmd(addr, pmdp, pgtable);
+}
#else
#define hvo_update_vmemmap_pte try_update_vmemmap_pte
+#define hvo_populate_vmemmap_pmd try_populate_vmemmap_pmd
#endif /* CONFIG_FAIL_HUGETLB_VMEMMAP */
static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
@@ -113,7 +131,7 @@ static int vmemmap_split_pmd(pmd_t *pmd, struct page *head, unsigned long start,
if (likely(pmd_leaf(*pmd))) {
/* Make pte visible before pmd. See comment in pmd_install(). */
smp_wmb();
- ret = try_populate_vmemmap_pmd(start, pmd, pgtable);
+ ret = hvo_populate_vmemmap_pmd(start, pmd, pgtable);
if (ret)
goto free;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread* [PATCH v2 20/20 DO-NOT-MERGE] selftests/mm: Add HVO pmd-split fault injection tests
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
` (18 preceding siblings ...)
2026-10-03 0:21 ` [PATCH v2 19/20 DO-NOT-MERGE] hugetlb_vmemmap: Add fault injection for in-place vmemmap PMD splits James Houghton
@ 2026-10-03 0:21 ` James Houghton
19 siblings, 0 replies; 21+ messages in thread
From: James Houghton @ 2026-10-03 0:21 UTC (permalink / raw)
To: Will Deacon, Catalin Marinas, Muchun Song, Oscar Salvador, Andrew Morton
Cc: Nikos Nikoleris, Linu Cherian, Mark Rutland, David Hildenbrand,
Ryan Roberts, Nanyong Sun, Yu Zhao, Frank van der Linden,
David Rientjes, James Houghton, linux-kernel, linux-arm-kernel,
linux-mm
Run the stress test with PMD splitting failure injection enabled (if
supported). Keep in mind that vmemmap PMDs are only ever split, never
collapsed back to PMDs, so run this step first, before the main stress
test.
Assisted-by: LLM
Signed-off-by: James Houghton <jthoughton@google.com>
---
tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh | 12 ++++++++----
1 file changed, 8 insertions(+), 4 deletions(-)
diff --git a/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh b/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh
index 94357a4a45c2..2c874d15399a 100755
--- a/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh
+++ b/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh
@@ -4,13 +4,16 @@
# Stress test for HugeTLB vmemmap optimization (HVO).
#
# Phases:
-# 1. accounting: allocate N hugepages, check that nr_memmap_pages +
+# 1. pmd-inject: churn the pool with fail_hugetlb_vmemmap_pmd enabled.
+# Runs first because vmemmap PMDs are only split once, so later phases
+# would leave few PMD splits to fail.
+# 2. accounting: allocate N hugepages, check that nr_memmap_pages +
# nr_memmap_boot_pages drops by exactly N * (vmemmap pages per folio - 1),
# then free them and check it returns to the baseline.
-# 2. stress: churn the pool (optimize/restore) while concurrently reading
+# 3. stress: churn the pool (optimize/restore) while concurrently reading
# struct pages (/proc/kpageflags, /proc/kpagecount), compacting memory,
# and optionally reading page_owner and offlining/onlining memory.
-# 3. pte-inject: like 2, with fail_hugetlb_vmemmap_pte enabled; this hits
+# 4. pte-inject: like 2, with fail_hugetlb_vmemmap_pte enabled; this hits
# both the optimize rollback and the restore (partial HVO) paths.
#
# After each phase the pool is drained back to its original size, memmap
@@ -252,7 +255,7 @@ cleanup() {
local blk t
stop_workers
- for t in fail_hugetlb_vmemmap_pte; do
+ for t in fail_hugetlb_vmemmap_pte fail_hugetlb_vmemmap_pmd; do
[ -d "$(fa_dir $t)" ] && fa_disable $t > /dev/null
done
if [ -f "$tmpdir/hotplug" ]; then
@@ -333,6 +336,7 @@ accounting_phase() {
echo "$marker" > /dev/kmsg
+stress_phase pmd-inject fail_hugetlb_vmemmap_pmd
accounting_phase
stress_phase stress ""
stress_phase pte-inject fail_hugetlb_vmemmap_pte
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 21+ messages in thread