mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown
@ 2026-10-02 20:05 Dima Koziuk
  2026-10-02 20:05 ` [PATCH v2 2/2] mm: kmsan: fix ioremap error cleanup Dima Koziuk
  2026-10-02 21:12 ` [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown Andrew Morton
  0 siblings, 2 replies; 3+ messages in thread
From: Dima Koziuk @ 2026-10-02 20:05 UTC (permalink / raw)
  To: glider, akpm
  Cc: dmytrokoziuk68, elver, dvyukov, urezki, kasan-dev, linux-mm,
	linux-kernel

kmsan_vmalloc_to_page_or_null() rejects addresses outside the regular
vmalloc and module ranges, including the shadow and origin addresses
passed by its callers. As a result, iounmap does not free the metadata
backing pages. Also, the first iteration of the teardown loop
unmaps the entire metadata range, preventing subsequent iterations from
finding their pages even if the address lookup is fixed.

Factor the page-table walk out of vmalloc_to_page() into
__vmalloc_to_page(), keeping the address check in the public wrapper.
Use the unchecked helper in kmsan_vmalloc_to_page_or_null() and check
for NULL before converting the returned page to a PFN. Mark
__vmalloc_to_page() as __always_inline to avoid an additional function
call in vmalloc_to_page().

Move metadata teardown into kmsan_iounmap_pages(). Free the backing
blocks while their mappings are still available, then unmap each
metadata range once and flush its TLB entries.

Fixes: b073d7f8aee4 ("mm: kmsan: maintain KMSAN metadata for page operations")
Suggested-by: Alexander Potapenko <glider@google.com>
Link: https://lkml.iu.edu/2609.3/12748.html
Signed-off-by: Dima Koziuk <dmytrokoziuk68@gmail.com>
---
Changes in v2:
- Share the vmalloc page-table walk instead of adding a KMSAN PTE walker.
- Keep kmsan_vmalloc_to_page_or_null() and handle absent mappings.

 mm/kmsan/core.c  |  8 +++-----
 mm/kmsan/hooks.c | 51 +++++++++++++++++++++++++++++++-----------------
 mm/vmalloc.c     | 21 ++++++++++++--------
 mm/vmalloc.h     |  1 +
 4 files changed, 50 insertions(+), 31 deletions(-)

diff --git a/mm/kmsan/core.c b/mm/kmsan/core.c
index 90f427b95a21..3ad5f31d56b6 100644
--- a/mm/kmsan/core.c
+++ b/mm/kmsan/core.c
@@ -27,6 +27,7 @@
 #include <linux/vmalloc.h>
 
 #include "../slab.h"
+#include "../vmalloc.h"
 #include "kmsan.h"
 
 bool kmsan_enabled __read_mostly;
@@ -240,11 +241,8 @@ struct page *kmsan_vmalloc_to_page_or_null(void *vaddr)
 {
 	struct page *page;
 
-	if (!kmsan_internal_is_vmalloc_addr(vaddr) &&
-	    !kmsan_internal_is_module_addr(vaddr))
-		return NULL;
-	page = vmalloc_to_page(vaddr);
-	if (pfn_valid(page_to_pfn(page)))
+	page = __vmalloc_to_page(vaddr);
+	if (page && pfn_valid(page_to_pfn(page)))
 		return page;
 	else
 		return NULL;
diff --git a/mm/kmsan/hooks.c b/mm/kmsan/hooks.c
index 5f1b8053f9fa..24f71bc65896 100644
--- a/mm/kmsan/hooks.c
+++ b/mm/kmsan/hooks.c
@@ -20,6 +20,8 @@
 #include <linux/uaccess.h>
 #include <linux/usb.h>
 
+#include <asm/tlbflush.h>
+
 #include "../internal.h"
 #include "../vmalloc.h"
 #include "../slab.h"
@@ -142,6 +144,36 @@ void kmsan_vunmap_range_noflush(unsigned long start, unsigned long end)
 	flush_cache_vmap(vmalloc_origin(start), vmalloc_origin(end));
 }
 
+#define KMSAN_IOREMAP_META_ORDER 1
+
+static void kmsan_iounmap_pages(unsigned long start, unsigned long end)
+{
+	unsigned long shadow_start = vmalloc_shadow(start),
+		      shadow_end = vmalloc_shadow(end);
+	unsigned long origin_start = vmalloc_origin(start),
+		      origin_end = vmalloc_origin(end);
+	unsigned long v_shadow, v_origin;
+	struct page *shadow, *origin;
+	int nr;
+
+	nr = (end - start) / PAGE_SIZE;
+	v_shadow = shadow_start;
+	v_origin = origin_start;
+	for (int i = 0; i < nr;
+	     i++, v_shadow += PAGE_SIZE, v_origin += PAGE_SIZE) {
+		shadow = kmsan_vmalloc_to_page_or_null((void *)v_shadow);
+		origin = kmsan_vmalloc_to_page_or_null((void *)v_origin);
+		if (shadow)
+			__free_pages(shadow, KMSAN_IOREMAP_META_ORDER);
+		if (origin)
+			__free_pages(origin, KMSAN_IOREMAP_META_ORDER);
+	}
+	__vunmap_range_noflush(shadow_start, shadow_end);
+	__vunmap_range_noflush(origin_start, origin_end);
+	flush_tlb_kernel_range(shadow_start, shadow_end);
+	flush_tlb_kernel_range(origin_start, origin_end);
+}
+
 /*
  * This function creates new shadow/origin pages for the physical pages mapped
  * into the virtual memory. If those physical pages already had shadow/origin,
@@ -219,28 +251,11 @@ int kmsan_ioremap_page_range(unsigned long start, unsigned long end,
 
 void kmsan_iounmap_page_range(unsigned long start, unsigned long end)
 {
-	unsigned long v_shadow, v_origin;
-	struct page *shadow, *origin;
-	int nr;
-
 	if (!kmsan_enabled || kmsan_in_runtime())
 		return;
 
-	nr = (end - start) / PAGE_SIZE;
 	kmsan_enter_runtime();
-	v_shadow = (unsigned long)vmalloc_shadow(start);
-	v_origin = (unsigned long)vmalloc_origin(start);
-	for (int i = 0; i < nr;
-	     i++, v_shadow += PAGE_SIZE, v_origin += PAGE_SIZE) {
-		shadow = kmsan_vmalloc_to_page_or_null((void *)v_shadow);
-		origin = kmsan_vmalloc_to_page_or_null((void *)v_origin);
-		__vunmap_range_noflush(v_shadow, vmalloc_shadow(end));
-		__vunmap_range_noflush(v_origin, vmalloc_origin(end));
-		if (shadow)
-			__free_pages(shadow, 1);
-		if (origin)
-			__free_pages(origin, 1);
-	}
+	kmsan_iounmap_pages(start, end);
 	flush_cache_vmap(vmalloc_shadow(start), vmalloc_shadow(end));
 	flush_cache_vmap(vmalloc_origin(start), vmalloc_origin(end));
 	kmsan_leave_runtime();
diff --git a/mm/vmalloc.c b/mm/vmalloc.c
index bea9f76ed7e7..5fe37ad41f38 100644
--- a/mm/vmalloc.c
+++ b/mm/vmalloc.c
@@ -817,9 +817,10 @@ EXPORT_SYMBOL_GPL(is_vmalloc_or_module_addr);
 /*
  * Walk a vmap address to the struct page it maps. Huge vmap mappings will
  * return the tail page that corresponds to the base page address, which
- * matches small vmap mappings.
+ * matches small vmap mappings. Unlike vmalloc_to_page(), this also accepts
+ * addresses outside the vmalloc and module ranges, such as KMSAN metadata.
  */
-struct page *vmalloc_to_page(const void *vmalloc_addr)
+__always_inline struct page *__vmalloc_to_page(const void *vmalloc_addr)
 {
 	unsigned long addr = (unsigned long) vmalloc_addr;
 	struct page *page = NULL;
@@ -829,12 +830,6 @@ struct page *vmalloc_to_page(const void *vmalloc_addr)
 	pmd_t *pmd;
 	pte_t *ptep, pte;
 
-	/*
-	 * XXX we might need to change this if we add VIRTUAL_BUG_ON for
-	 * architectures that do not vmalloc module space
-	 */
-	VIRTUAL_BUG_ON(!is_vmalloc_or_module_addr(vmalloc_addr));
-
 	if (pgd_none(*pgd))
 		return NULL;
 	if (WARN_ON_ONCE(pgd_leaf(*pgd)))
@@ -873,6 +868,16 @@ struct page *vmalloc_to_page(const void *vmalloc_addr)
 
 	return page;
 }
+
+struct page *vmalloc_to_page(const void *vmalloc_addr)
+{
+	/*
+	 * XXX we might need to change this if we add VIRTUAL_BUG_ON for
+	 * architectures that do not vmalloc module space
+	 */
+	VIRTUAL_BUG_ON(!is_vmalloc_or_module_addr(vmalloc_addr));
+	return __vmalloc_to_page(vmalloc_addr);
+}
 EXPORT_SYMBOL(vmalloc_to_page);
 
 /*
diff --git a/mm/vmalloc.h b/mm/vmalloc.h
index 8866ddcff668..bce4c982c88b 100644
--- a/mm/vmalloc.h
+++ b/mm/vmalloc.h
@@ -8,6 +8,7 @@
 #include <linux/vmalloc.h>
 
 #ifdef CONFIG_MMU
+struct page *__vmalloc_to_page(const void *vmalloc_addr);
 void __init vmalloc_init(void);
 int __must_check vmap_pages_range_noflush(unsigned long addr, unsigned long end,
 		pgprot_t prot, struct page **pages,

base-commit: 587858367581b9c55c3690f4e63382ad622719d4

^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH v2 2/2] mm: kmsan: fix ioremap error cleanup
  2026-10-02 20:05 [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown Dima Koziuk
@ 2026-10-02 20:05 ` Dima Koziuk
  2026-10-02 21:12 ` [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown Andrew Morton
  1 sibling, 0 replies; 3+ messages in thread
From: Dima Koziuk @ 2026-10-02 20:05 UTC (permalink / raw)
  To: glider, akpm
  Cc: dmytrokoziuk68, elver, dvyukov, urezki, kasan-dev, linux-mm,
	linux-kernel

Looking further at kmsan_ioremap_page_range(), I found three cases where
error cleanup leaks metadata blocks.

1. If the first iteration fails, clean is zero and cleanup is skipped,
   leaking any allocations that succeeded in that iteration.

2. If shadow mapping succeeds but origin mapping fails, the shadow
   pointer has already been cleared. Removing its mapping loses the
   backing block.

3. On failures after completed iterations, cleanup removes the earlier
   metadata mappings without freeing their backing blocks.

The cleanup needed here is the same as for iounmap, so it makes sense to
reuse kmsan_iounmap_pages(). Track the end of installed mappings with
mapped_end and advance it after each successful shadow mapping. This
includes the current shadow block if origin mapping subsequently fails,
while the helper skips the missing origin mapping.

Run cleanup whenever err is non-zero. Free allocations that have not been
mapped directly, and use the shared helper to unmap and free the installed
metadata, including blocks from completed iterations.

Fixes: fdea03e12aa2 ("mm: kmsan: handle alloc failures in kmsan_ioremap_page_range()")
Reviewed-by: Alexander Potapenko <glider@google.com>
Signed-off-by: Dima Koziuk <dmytrokoziuk68@gmail.com>
---
Changes in v2:
- No code changes. 
- Add Reviewed-by Alexander Potapenko from v1.

I tested this series on Linux 7.3-rc3 under QEMU, 
using ioremap()/iounmap() calls on the QEMU VGA BAR0. 
All tested mappings were torn down without metadata leaks.

 mm/kmsan/hooks.c | 35 ++++++++++++-----------------------
 1 file changed, 12 insertions(+), 23 deletions(-)

diff --git a/mm/kmsan/hooks.c b/mm/kmsan/hooks.c
index 24f71bc65896..a706666db097 100644
--- a/mm/kmsan/hooks.c
+++ b/mm/kmsan/hooks.c
@@ -186,16 +186,17 @@ int kmsan_ioremap_page_range(unsigned long start, unsigned long end,
 	gfp_t gfp_mask = GFP_KERNEL | __GFP_ZERO;
 	struct page *shadow, *origin;
 	unsigned long off = 0;
-	int nr, err = 0, clean = 0, mapped;
+	unsigned long mapped_end = start;
+	int nr, err = 0, mapped;
 
 	if (!kmsan_enabled || kmsan_in_runtime())
 		return 0;
 
 	nr = (end - start) / PAGE_SIZE;
 	kmsan_enter_runtime();
-	for (int i = 0; i < nr; i++, off += PAGE_SIZE, clean = i) {
-		shadow = alloc_pages(gfp_mask, 1);
-		origin = alloc_pages(gfp_mask, 1);
+	for (int i = 0; i < nr; i++, off += PAGE_SIZE) {
+		shadow = alloc_pages(gfp_mask, KMSAN_IOREMAP_META_ORDER);
+		origin = alloc_pages(gfp_mask, KMSAN_IOREMAP_META_ORDER);
 		if (!shadow || !origin) {
 			err = -ENOMEM;
 			goto ret;
@@ -209,39 +210,27 @@ int kmsan_ioremap_page_range(unsigned long start, unsigned long end,
 			goto ret;
 		}
 		shadow = NULL;
+		mapped_end = start + off + PAGE_SIZE;
 		mapped = __vmap_pages_range_noflush(
 			vmalloc_origin(start + off),
 			vmalloc_origin(start + off + PAGE_SIZE), prot, &origin,
 			PAGE_SHIFT);
 		if (mapped) {
-			__vunmap_range_noflush(
-				vmalloc_shadow(start + off),
-				vmalloc_shadow(start + off + PAGE_SIZE));
 			err = mapped;
 			goto ret;
 		}
 		origin = NULL;
 	}
-	/* Page mapping loop finished normally, nothing to clean up. */
-	clean = 0;
 
 ret:
-	if (clean > 0) {
-		/*
-		 * Something went wrong. Clean up shadow/origin pages allocated
-		 * on the last loop iteration, then delete mappings created
-		 * during the previous iterations.
-		 */
+	if (err) {
 		if (shadow)
-			__free_pages(shadow, 1);
+			__free_pages(shadow, KMSAN_IOREMAP_META_ORDER);
 		if (origin)
-			__free_pages(origin, 1);
-		__vunmap_range_noflush(
-			vmalloc_shadow(start),
-			vmalloc_shadow(start + clean * PAGE_SIZE));
-		__vunmap_range_noflush(
-			vmalloc_origin(start),
-			vmalloc_origin(start + clean * PAGE_SIZE));
+			__free_pages(origin, KMSAN_IOREMAP_META_ORDER);
+
+		if (mapped_end > start)
+			kmsan_iounmap_pages(start, mapped_end);
 	}
 	flush_cache_vmap(vmalloc_shadow(start), vmalloc_shadow(end));
 	flush_cache_vmap(vmalloc_origin(start), vmalloc_origin(end));

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown
  2026-10-02 20:05 [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown Dima Koziuk
  2026-10-02 20:05 ` [PATCH v2 2/2] mm: kmsan: fix ioremap error cleanup Dima Koziuk
@ 2026-10-02 21:12 ` Andrew Morton
  1 sibling, 0 replies; 3+ messages in thread
From: Andrew Morton @ 2026-10-02 21:12 UTC (permalink / raw)
  To: Dima Koziuk
  Cc: glider, elver, dvyukov, urezki, kasan-dev, linux-mm, linux-kernel

On Fri,  2 Oct 2026 23:05:07 +0300 Dima Koziuk <dmytrokoziuk68@gmail.com> wrote:

> kmsan_vmalloc_to_page_or_null() rejects addresses outside the regular
> vmalloc and module ranges, including the shadow and origin addresses
> passed by its callers. As a result, iounmap does not free the metadata
> backing pages. Also, the first iteration of the teardown loop
> unmaps the entire metadata range, preventing subsequent iterations from
> finding their pages even if the address lookup is fixed.
> 
> Factor the page-table walk out of vmalloc_to_page() into
> __vmalloc_to_page(), keeping the address check in the public wrapper.
> Use the unchecked helper in kmsan_vmalloc_to_page_or_null() and check
> for NULL before converting the returned page to a PFN. Mark
> __vmalloc_to_page() as __always_inline to avoid an additional function
> call in vmalloc_to_page().
> 
> Move metadata teardown into kmsan_iounmap_pages(). Free the backing
> blocks while their mappings are still available, then unmap each
> metadata range once and flush its TLB entries.

Thanks, I'll queue these for test and further review.

Could people please offer opinions on whether we should backport
one/both into -stable kernels?


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-02 21:12 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-02 20:05 [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown Dima Koziuk
2026-10-02 20:05 ` [PATCH v2 2/2] mm: kmsan: fix ioremap error cleanup Dima Koziuk
2026-10-02 21:12 ` [PATCH v2 1/2] mm: kmsan: fix iounmap metadata teardown Andrew Morton

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®