From: James Houghton <jthoughton@google.com>
To: Will Deacon <will@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Muchun Song <muchun.song@linux.dev>,
Oscar Salvador <osalvador@suse.de>,
Andrew Morton <akpm@linux-foundation.org>
Cc: Nikos Nikoleris <nikos.nikoleris@arm.com>,
Linu Cherian <linu.cherian@arm.com>,
Mark Rutland <mark.rutland@arm.com>,
David Hildenbrand <david@kernel.org>,
Ryan Roberts <ryan.roberts@arm.com>,
Nanyong Sun <sunnanyong@huawei.com>, Yu Zhao <yuzhao@google.com>,
Frank van der Linden <fvdl@google.com>,
David Rientjes <rientjes@google.com>,
James Houghton <jthoughton@google.com>,
linux-kernel@vger.kernel.org,
linux-arm-kernel@lists.infradead.org, linux-mm@kvack.org
Subject: [PATCH v2 04/20] hugetlb_vmemmap: Leave pages partially HVOed upon restore failure
Date: Sat, 3 Oct 2026 00:21:07 +0000 [thread overview]
Message-ID: <20261003002123.505555-5-jthoughton@google.com> (raw)
In-Reply-To: <20261003002123.505555-1-jthoughton@google.com>
Right now it is assumed that the restore routine in vmemmap_remap_free()
will not fail; that is, if we fail to fully optimize a folio, it is
assumed that restoring the folio to its pre-HVO state is guaranteed to
succeed.
Although vmemmap restoration may never fail in practice, properly handle
such failure cases now. This will leave folios in a partially HVOed
state. In a follow-up patch, failure certainly becomes possible.
In the event we end up with partially-HVOed folios, make sure to:
1. Leave the HVO page flag in place.
2. Free any unused vmemmap pages.
Always pass the vmemmap_tail page to the restore routine to avoid
leaking the non-optimized vmemmap pages for a partially HVOed page.
Signed-off-by: James Houghton <jthoughton@google.com>
---
mm/hugetlb_vmemmap.c | 49 ++++++++++++++++++++++++++++++++++----------
1 file changed, 38 insertions(+), 11 deletions(-)
diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c
index 1a2645c382ba..90db4d069ff6 100644
--- a/mm/hugetlb_vmemmap.c
+++ b/mm/hugetlb_vmemmap.c
@@ -235,12 +235,15 @@ static void vmemmap_restore_pte(pte_t *pte, unsigned long addr,
{
struct page *src = pte_page(ptep_get(pte)), *dst;
+ if (WARN_ON_ONCE(!walk->vmemmap_tail))
+ return;
+
/*
- * When rolling back vmemmap_remap_free(), keep the copied head page
+ * When restoring a partially-HVOed page, keep the copied head page
* mapping and restore only PTEs currently pointing at the shared tail
* page.
*/
- if (walk->vmemmap_tail && walk->vmemmap_tail != src)
+ if (walk->vmemmap_tail != src)
return;
VM_WARN_ON_ONCE(PageHead((const struct page *)addr));
@@ -276,6 +279,7 @@ static int vmemmap_remap_split(unsigned long start, unsigned long end)
return vmemmap_remap_range(start, end, &walk);
}
+#define VMEMMAP_REMAP_INCOMPLETE 1
/**
* vmemmap_remap_free - remap the vmemmap virtual address range [@start, @end)
* to use @vmemmap_head/tail, then free vmemmap which
@@ -290,7 +294,8 @@ static int vmemmap_remap_split(unsigned long start, unsigned long end)
* responsibility to free pages.
* @flags: modifications to vmemmap_remap_walk flags
*
- * Return: %0 on success, negative error code otherwise.
+ * Return: %0 on success, VMEMMAP_REMAP_INCOMPLETE if the page is incompletely
+ * optimized, negative error code otherwise.
*/
static int vmemmap_remap_free(unsigned long start, unsigned long end,
struct page *vmemmap_head,
@@ -325,7 +330,8 @@ static int vmemmap_remap_free(unsigned long start, unsigned long end,
.flags = 0,
};
- vmemmap_remap_range(start, end, &walk);
+ if (vmemmap_remap_range(start, end, &walk))
+ return VMEMMAP_REMAP_INCOMPLETE;
return ret;
}
@@ -358,6 +364,8 @@ static int alloc_vmemmap_page_list(unsigned long start, unsigned long end,
* vmemmap_remap_alloc - remap the vmemmap virtual address range [@start, end)
* to the page which is from the @vmemmap_pages
* respectively.
+ * @h: the hstate for the folio whose vmemmap is getting remapped
+ * @folio: the folio whose vmemmap is getting remapped
* @start: start address of the vmemmap virtual address range that we want
* to remap.
* @end: end address of the vmemmap virtual address range that we want to
@@ -366,20 +374,35 @@ static int alloc_vmemmap_page_list(unsigned long start, unsigned long end,
*
* Return: %0 on success, negative error code otherwise.
*/
-static int vmemmap_remap_alloc(unsigned long start, unsigned long end,
+static int vmemmap_remap_alloc(const struct hstate *h, struct folio *folio,
+ unsigned long start, unsigned long end,
unsigned long flags)
{
LIST_HEAD(vmemmap_pages);
- struct vmemmap_remap_walk walk = {
+ struct vmemmap_remap_walk walk;
+ struct page *vmemmap_tail;
+ int ret;
+
+ vmemmap_tail = vmemmap_shared_tail_page(h->order, folio_zone(folio));
+ if (WARN_ON_ONCE(!vmemmap_tail))
+ return -ENOMEM;
+
+ if (alloc_vmemmap_page_list(start, end, &vmemmap_pages))
+ return -ENOMEM;
+
+ walk = (struct vmemmap_remap_walk) {
.remap_pte = vmemmap_restore_pte,
+ .vmemmap_tail = vmemmap_tail,
.vmemmap_pages = &vmemmap_pages,
.flags = flags,
};
- if (alloc_vmemmap_page_list(start, end, &vmemmap_pages))
- return -ENOMEM;
+ ret = vmemmap_remap_range(start, end, &walk);
- return vmemmap_remap_range(start, end, &walk);
+ /* Not all pages may have been consumed */
+ free_vmemmap_page_list(&vmemmap_pages);
+
+ return ret;
}
static bool vmemmap_optimize_enabled = IS_ENABLED(CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP_DEFAULT_ON);
@@ -412,7 +435,7 @@ static int __hugetlb_vmemmap_restore_folio(const struct hstate *h,
* When a HugeTLB page is freed to the buddy allocator, previously
* discarded vmemmap pages must be allocated and remapping.
*/
- ret = vmemmap_remap_alloc(vmemmap_start, vmemmap_end, flags);
+ ret = vmemmap_remap_alloc(h, folio, vmemmap_start, vmemmap_end, flags);
if (!ret)
folio_clear_hugetlb_vmemmap_optimized(folio);
@@ -545,7 +568,11 @@ static int __hugetlb_vmemmap_optimize_folio(const struct hstate *h,
vmemmap_head, vmemmap_tail,
vmemmap_pages, flags);
out:
- if (ret)
+ /*
+ * If ret == VMEMMAP_REMAP_INCOMPLETE, the folio might be partially
+ * HVOed. Leave the HVO page folio flag in place.
+ */
+ if (ret < 0)
folio_clear_hugetlb_vmemmap_optimized(folio);
return ret;
--
2.56.0.rc1.315.gc6ed9934b7-goog
next prev parent reply other threads:[~2026-10-03 0:21 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-03 0:21 [PATCH v2 00/20] Another attempt at HVO support on arm64 James Houghton
2026-10-03 0:21 ` [PATCH v2 01/20] hugetlb: Don't restore vmemmap of non-HVOed folios on bulk restore error James Houghton
2026-10-03 0:21 ` [PATCH v2 02/20] arm64/pgtable: Clear AF with LDCLR on supported systems James Houghton
2026-10-03 0:21 ` [PATCH v2 03/20] hugetlb_vmemmap: Always flush TLB if needed upon PTE remapping James Houghton
2026-10-03 0:21 ` James Houghton [this message]
2026-10-03 0:21 ` [PATCH v2 05/20] hugetlb_vmemmap: Use try_update_vmemmap_pte to update in-use PTEs James Houghton
2026-10-03 0:21 ` [PATCH v2 06/20] hugetlb_vmemmap: Allow architectures to dynamically disallow HVO James Houghton
2026-10-03 0:21 ` [PATCH v2 07/20] hugetlb_vmemmap: Disable HVO sysctl if arch doesn't support HVO James Houghton
2026-10-03 0:21 ` [PATCH v2 08/20] hugetlb: Fully initialize tail struct pages of non-pre-HVOed bootmem folios James Houghton
2026-10-03 0:21 ` [PATCH v2 09/20] hugetlb_vmemmap: Allow architectures to make HVO enablement boot-time only James Houghton
2026-10-03 0:21 ` [PATCH v2 10/20] hugetlb_vmemmap: Expose whether HVO is enabled to architecture code James Houghton
2026-10-03 0:21 ` [PATCH v2 11/20] arm64: Add bbm_through_af capability James Houghton
2026-10-03 0:21 ` [PATCH v2 12/20] arm64: Implement try_update_vmemmap_pte using the AF trick James Houghton
2026-10-03 0:21 ` [PATCH v2 13/20] arm64: Support hugetlb vmemmap optimization James Houghton
2026-10-03 0:21 ` [PATCH v2 14/20] hugetlb_vmemmap: Add fault injection for in-place vmemmap PTE updates James Houghton
2026-10-03 0:21 ` [PATCH v2 15/20] selftests/mm: Add HugeTLB vmemmap optimization stress test James Houghton
2026-10-03 0:21 ` [PATCH v2 16/20 DO-NOT-MERGE] hugetlb_vmemmap: Use try_populate_vmemmap_pmd for replacing in-use PMDs James Houghton
2026-10-03 0:21 ` [PATCH v2 17/20 DO-NOT-MERGE] arm64: Implement try_populate_vmemmap_pmd using AF trick James Houghton
2026-10-03 0:21 ` [PATCH v2 18/20 DO-NOT-MERGE] arm64: Drop BBML3 requirement for HVO James Houghton
2026-10-03 0:21 ` [PATCH v2 19/20 DO-NOT-MERGE] hugetlb_vmemmap: Add fault injection for in-place vmemmap PMD splits James Houghton
2026-10-03 0:21 ` [PATCH v2 20/20 DO-NOT-MERGE] selftests/mm: Add HVO pmd-split fault injection tests James Houghton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261003002123.505555-5-jthoughton@google.com \
--to=jthoughton@google.com \
--cc=akpm@linux-foundation.org \
--cc=catalin.marinas@arm.com \
--cc=david@kernel.org \
--cc=fvdl@google.com \
--cc=linu.cherian@arm.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mark.rutland@arm.com \
--cc=muchun.song@linux.dev \
--cc=nikos.nikoleris@arm.com \
--cc=osalvador@suse.de \
--cc=rientjes@google.com \
--cc=ryan.roberts@arm.com \
--cc=sunnanyong@huawei.com \
--cc=will@kernel.org \
--cc=yuzhao@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®