* [RFC PATCH 1/4] mm/kmsan: undo the shadow mapping when the origin mapping fails
2026-10-09 6:36 [RFC PATCH 0/4] mm/vmalloc: make the mapping functions undo their partial mappings Hao Ge
@ 2026-10-09 6:36 ` Hao Ge
2026-10-09 6:36 ` [RFC PATCH 2/4] mm/vmalloc: undo partial mappings inside the mapping functions Hao Ge
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Hao Ge @ 2026-10-09 6:36 UTC (permalink / raw)
To: Suren Baghdasaryan, Madhavan Srinivasan, Michael Ellerman,
Nicholas Piggin, Christophe Leroy (CS GROUP),
Ritesh Harjani (IBM),
Shrikanth Hegde, Alexander Potapenko, Marco Elver, Dmitry Vyukov,
Andrew Morton, Dennis Zhou, Tejun Heo, Christoph Lameter,
Uladzislau Rezki
Cc: Hao Ge, linuxppc-dev, linux-kernel, kasan-dev, linux-mm, stable
kmsan_vmap_pages_range_noflush() first maps the shadow pages and then
the origin pages. If the second mapping fails, the first one is left
mapped, and the callers do not roll it back. The next vmap of the
same metadata range hits the stale PTEs again and BUG()s on the huge
mapping path, or gets -EBUSY with a WARN_ON() on the small page one.
Undo the shadow mapping on that error path, the same way
kmsan_ioremap_page_range() cleans up after a partial failure.
__vunmap_range_noflush() only clears the PTEs; nothing has touched
the shadow mapping, so no TLB flush is needed.
Fixes: 47ebd0310e89 ("mm: kmsan: handle alloc failures in kmsan_vmap_pages_range_noflush()")
Cc: stable@vger.kernel.org
Signed-off-by: Hao Ge <hao.ge@linux.dev>
---
mm/kmsan/shadow.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/mm/kmsan/shadow.c b/mm/kmsan/shadow.c
index 0c88d89bf0d6..2166086d3dc3 100644
--- a/mm/kmsan/shadow.c
+++ b/mm/kmsan/shadow.c
@@ -258,6 +258,10 @@ int kmsan_vmap_pages_range_noflush(unsigned long start, unsigned long end,
o_pages, page_shift);
kmsan_leave_runtime();
if (mapped) {
+ /* Undo the shadow mapping set up above. */
+ kmsan_enter_runtime();
+ __vunmap_range_noflush(shadow_start, shadow_end);
+ kmsan_leave_runtime();
err = mapped;
goto ret;
}
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH 2/4] mm/vmalloc: undo partial mappings inside the mapping functions
2026-10-09 6:36 [RFC PATCH 0/4] mm/vmalloc: make the mapping functions undo their partial mappings Hao Ge
2026-10-09 6:36 ` [RFC PATCH 1/4] mm/kmsan: undo the shadow mapping when the origin mapping fails Hao Ge
@ 2026-10-09 6:36 ` Hao Ge
2026-10-09 6:36 ` [RFC PATCH 3/4] powerpc: drop redundant unmaps of failed vmap mappings Hao Ge
2026-10-09 6:36 ` [RFC PATCH 4/4] mm/percpu: stop unmapping the CPU that failed to map Hao Ge
3 siblings, 0 replies; 5+ messages in thread
From: Hao Ge @ 2026-10-09 6:36 UTC (permalink / raw)
To: Suren Baghdasaryan, Madhavan Srinivasan, Michael Ellerman,
Nicholas Piggin, Christophe Leroy (CS GROUP),
Ritesh Harjani (IBM),
Shrikanth Hegde, Alexander Potapenko, Marco Elver, Dmitry Vyukov,
Andrew Morton, Dennis Zhou, Tejun Heo, Christoph Lameter,
Uladzislau Rezki
Cc: Hao Ge, linuxppc-dev, linux-kernel, kasan-dev, linux-mm, Sashiko
__vmap_pages_range_noflush() and friends can install some PTEs before
failing and leave them mapped, and the kernel callers do not agree on
who cleans them up. pcpu_map_pages() and kmsan_ioremap_page_range()
roll back what they mapped before the failure, while
vm_module_tags_populate() and the __GFP_NOFAIL retry loop in
__vmalloc_area_node() rely on the mapping functions cleaning up and
do not call anything like vunmap_range() themselves. When the same
range is mapped again, the attempt hits the leftovers and fails, with
a BUG() in vmap_pte_range() for huge mappings and a warning on the
small-page path.
After discussing with Suren and Ulad, we decided to put the rollback
into the entry points of the vmap API rather than into every low-level
helper. __vmap_pages_range_noflush() undoes the whole range it was
asked to map when it fails, and vmap_page_range() undoes its range
for the ioremap-style mappings.
The rollback is __vunmap_range_noflush(), it only clears the PTEs,
no TLB flush, nothing has touched these mappings.
vmap_pages_range_noflush() drops the KMSAN metadata when the data
mapping fails, and an empty range bails out with -EINVAL now, which
used to BUG_ON() on the small-page path and map nothing at all on
the huge-page one.
Reported-by: Sashiko <sashiko-bot@kernel.org>
Suggested-by: Uladyslau Rezki <urezki@gmail.com>
Signed-off-by: Hao Ge <hao.ge@linux.dev>
---
mm/vmalloc.c | 44 ++++++++++++++++++++++++++++++++------------
1 file changed, 32 insertions(+), 12 deletions(-)
diff --git a/mm/vmalloc.c b/mm/vmalloc.c
index db669103dc66..e1b376f57dd2 100644
--- a/mm/vmalloc.c
+++ b/mm/vmalloc.c
@@ -363,6 +363,9 @@ int vmap_page_range(unsigned long addr, unsigned long end,
if (!err)
err = kmsan_ioremap_page_range(addr, end, phys_addr, prot,
ioremap_max_page_shift);
+ if (err)
+ __vunmap_range_noflush(addr, end);
+
return err;
}
@@ -683,27 +686,35 @@ int __vmap_pages_range_noflush(unsigned long addr, unsigned long end,
pgprot_t prot, struct page **pages, unsigned int page_shift)
{
unsigned int i, nr = (end - addr) >> PAGE_SHIFT;
+ unsigned long start = addr;
+ int err = 0;
+
+ if (WARN_ON_ONCE(addr >= end))
+ return -EINVAL;
if (WARN_ON_ONCE(page_shift < PAGE_SHIFT))
return -EINVAL;
if (!IS_ENABLED(CONFIG_HAVE_ARCH_HUGE_VMALLOC) ||
- page_shift == PAGE_SHIFT)
- return vmap_small_pages_range_noflush(addr, end, prot, pages);
-
- for (i = 0; i < nr; i += 1U << (page_shift - PAGE_SHIFT)) {
- int err;
-
- err = vmap_range_noflush(addr, addr + (1UL << page_shift),
+ page_shift == PAGE_SHIFT) {
+ err = vmap_small_pages_range_noflush(addr, end, prot, pages);
+ } else {
+ for (i = 0; i < nr; i += 1U << (page_shift - PAGE_SHIFT)) {
+ err = vmap_range_noflush(addr, addr + (1UL << page_shift),
page_to_phys(pages[i]), prot,
page_shift);
- if (err)
- return err;
+ if (err)
+ break;
- addr += 1UL << page_shift;
+ addr += 1UL << page_shift;
+ }
}
- return 0;
+ /* Undo the PTEs installed before the failure. */
+ if (err)
+ __vunmap_range_noflush(start, end);
+
+ return err;
}
int vmap_pages_range_noflush(unsigned long addr, unsigned long end,
@@ -715,7 +726,16 @@ int vmap_pages_range_noflush(unsigned long addr, unsigned long end,
if (ret)
return ret;
- return __vmap_pages_range_noflush(addr, end, prot, pages, page_shift);
+
+ ret = __vmap_pages_range_noflush(addr, end, prot, pages, page_shift);
+ /*
+ * The page tables undo themselves on failure. Tear down the
+ * metadata that was fully set up before the mapping failed.
+ */
+ if (ret)
+ kmsan_vunmap_range_noflush(addr, end);
+
+ return ret;
}
static int __vmap_pages_range(unsigned long addr, unsigned long end,
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH 3/4] powerpc: drop redundant unmaps of failed vmap mappings
2026-10-09 6:36 [RFC PATCH 0/4] mm/vmalloc: make the mapping functions undo their partial mappings Hao Ge
2026-10-09 6:36 ` [RFC PATCH 1/4] mm/kmsan: undo the shadow mapping when the origin mapping fails Hao Ge
2026-10-09 6:36 ` [RFC PATCH 2/4] mm/vmalloc: undo partial mappings inside the mapping functions Hao Ge
@ 2026-10-09 6:36 ` Hao Ge
2026-10-09 6:36 ` [RFC PATCH 4/4] mm/percpu: stop unmapping the CPU that failed to map Hao Ge
3 siblings, 0 replies; 5+ messages in thread
From: Hao Ge @ 2026-10-09 6:36 UTC (permalink / raw)
To: Suren Baghdasaryan, Madhavan Srinivasan, Michael Ellerman,
Nicholas Piggin, Christophe Leroy (CS GROUP),
Ritesh Harjani (IBM),
Shrikanth Hegde, Alexander Potapenko, Marco Elver, Dmitry Vyukov,
Andrew Morton, Dennis Zhou, Tejun Heo, Christoph Lameter,
Uladzislau Rezki
Cc: Hao Ge, linuxppc-dev, linux-kernel, kasan-dev, linux-mm
vmap_page_range() and ioremap_page_range() clear the PTEs they
installed when they fail, so the vunmap_range() in remap_isa_base()
and ioremap_phb() has nothing left to remove.
Signed-off-by: Hao Ge <hao.ge@linux.dev>
---
arch/powerpc/kernel/isa-bridge.c | 5 ++---
arch/powerpc/kernel/pci_64.c | 4 +---
2 files changed, 3 insertions(+), 6 deletions(-)
diff --git a/arch/powerpc/kernel/isa-bridge.c b/arch/powerpc/kernel/isa-bridge.c
index 5c064485197a..93029d85e6ea 100644
--- a/arch/powerpc/kernel/isa-bridge.c
+++ b/arch/powerpc/kernel/isa-bridge.c
@@ -46,9 +46,8 @@ static void remap_isa_base(phys_addr_t pa, unsigned long size)
WARN_ON_ONCE(size & ~PAGE_MASK);
if (slab_is_available()) {
- if (vmap_page_range(ISA_IO_BASE, ISA_IO_BASE + size, pa,
- pgprot_noncached(PAGE_KERNEL)))
- vunmap_range(ISA_IO_BASE, ISA_IO_BASE + size);
+ vmap_page_range(ISA_IO_BASE, ISA_IO_BASE + size, pa,
+ pgprot_noncached(PAGE_KERNEL));
} else {
early_ioremap_range(ISA_IO_BASE, pa, size,
pgprot_noncached(PAGE_KERNEL));
diff --git a/arch/powerpc/kernel/pci_64.c b/arch/powerpc/kernel/pci_64.c
index e27342ef128b..f74269636d5e 100644
--- a/arch/powerpc/kernel/pci_64.c
+++ b/arch/powerpc/kernel/pci_64.c
@@ -139,10 +139,8 @@ void __iomem *ioremap_phb(phys_addr_t paddr, unsigned long size)
addr = (unsigned long)area->addr;
if (ioremap_page_range(addr, addr + size, paddr,
- pgprot_noncached(PAGE_KERNEL))) {
- vunmap_range(addr, addr + size);
+ pgprot_noncached(PAGE_KERNEL)))
return NULL;
- }
return (void __iomem *)addr;
}
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread* [RFC PATCH 4/4] mm/percpu: stop unmapping the CPU that failed to map
2026-10-09 6:36 [RFC PATCH 0/4] mm/vmalloc: make the mapping functions undo their partial mappings Hao Ge
` (2 preceding siblings ...)
2026-10-09 6:36 ` [RFC PATCH 3/4] powerpc: drop redundant unmaps of failed vmap mappings Hao Ge
@ 2026-10-09 6:36 ` Hao Ge
3 siblings, 0 replies; 5+ messages in thread
From: Hao Ge @ 2026-10-09 6:36 UTC (permalink / raw)
To: Suren Baghdasaryan, Madhavan Srinivasan, Michael Ellerman,
Nicholas Piggin, Christophe Leroy (CS GROUP),
Ritesh Harjani (IBM),
Shrikanth Hegde, Alexander Potapenko, Marco Elver, Dmitry Vyukov,
Andrew Morton, Dennis Zhou, Tejun Heo, Christoph Lameter,
Uladzislau Rezki
Cc: Hao Ge, linuxppc-dev, linux-kernel, kasan-dev, linux-mm
__pcpu_map_pages() clears the PTEs it installed when it fails, so
pcpu_map_pages() only has to tear down the CPUs it mapped before
the failing one.
Signed-off-by: Hao Ge <hao.ge@linux.dev>
---
mm/percpu-vm.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/mm/percpu-vm.c b/mm/percpu-vm.c
index 509d8901835c..59d3a336a6bd 100644
--- a/mm/percpu-vm.c
+++ b/mm/percpu-vm.c
@@ -252,10 +252,11 @@ static int pcpu_map_pages(struct pcpu_chunk *chunk, struct page **pages,
return 0;
err:
for_each_possible_cpu(tcpu) {
- __pcpu_unmap_pages(pcpu_chunk_addr(chunk, tcpu, page_start),
- page_end - page_start);
+ /* The failing CPU's partial mapping undoes itself */
if (tcpu == cpu)
break;
+ __pcpu_unmap_pages(pcpu_chunk_addr(chunk, tcpu, page_start),
+ page_end - page_start);
}
pcpu_post_unmap_tlb_flush(chunk, page_start, page_end);
return err;
--
2.25.1
^ permalink raw reply [flat|nested] 5+ messages in thread