mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] x86/kmsan: Fix two bugs in the metadata lookup
@ 2026-10-06 18:10 Viorel Cernateanu via B4 Relay
  2026-10-06 18:10 ` [PATCH 1/2] x86/kmsan: Fix CPU entry area " Viorel Cernateanu via B4 Relay
  2026-10-06 18:10 ` [PATCH 2/2] x86/kmsan: Don't call instrumented code from kmsan_virt_addr_valid() Viorel Cernateanu via B4 Relay
  0 siblings, 2 replies; 3+ messages in thread
From: Viorel Cernateanu via B4 Relay @ 2026-10-06 18:10 UTC (permalink / raw)
  To: Alexander Potapenko, Marco Elver, Dmitry Vyukov, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H. Peter Anvin,
	Andrew Morton, Peter Zijlstra, Kees Cook
  Cc: Borislav Petkov, kasan-dev, linux-kernel, Viorel Cernateanu

Two fixes in arch/x86/include/asm/kmsan.h, found while running KMSAN
kernels with KASLR and with CONFIG_DEBUG_PREEMPT.

Patch 1 fixes the CPU lookup for addresses in the CPU entry area, which
is wrong with KASLR, and for the last page of each area without it.

Patch 2 fixes a recursion with CONFIG_DEBUG_PREEMPT: the metadata lookup
ends up in the instrumented preempt_count_add(), and the kernel hangs
right after KMSAN is enabled.

Both were tested on v7.3-rc6 in QEMU/KVM, with x86_64_defconfig plus
x86_debug.config (without LOCK_STAT, LOCKDEP, PROVE_LOCKING, GCOV and
DWARF), KMSAN and UBSAN_BOUNDS, and with a module that poisons and then
checks the KMSAN metadata at the start and at the end of the current
CPU's entry area, so a KMSAN report means the metadata was found.

Patch 1: without it, the end of the area is missed without KASLR, and
the lookup oopses with KASLR. With it, both are found, with and without
KASLR (three boots each). On an Ivy Bridge laptop, v7.3-rc6 with this
patch and KASLR gets past the hard lockup detector setup, and the same
check finds the metadata on all four CPUs. The lookup now scans up to
nr_cpu_ids entries for addresses in the CPU entry area; I have not
measured the cost on machines with many CPUs.

Patch 2, with CONFIG_DEBUG_PREEMPT added: without it, three boots out of
three with KASLR stop after the "ATTENTION: KMSAN is a debugging tool!"
banner and never reach init. With both patches, six boots out of six
reach init, with and without KASLR, and the module above still gets its
KMSAN reports. Not tested with lockdep, PROVE_RCU or
TRACE_PREEMPT_TOGGLE, and not on hardware.

---
Viorel Cernateanu (2):
      x86/kmsan: Fix CPU entry area metadata lookup
      x86/kmsan: Don't call instrumented code from kmsan_virt_addr_valid()

 arch/x86/include/asm/kmsan.h | 76 ++++++++++++++++++++++++++++++--------------
 1 file changed, 52 insertions(+), 24 deletions(-)
---
base-commit: a90ee4305c4a5df72c11b31dacfdc76e00fcf78a
change-id: 20261006-kmsan-serie-e33000b199ce

Best regards,
--  
Viorel Cernateanu <vrilutza@gmail.com>



^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 1/2] x86/kmsan: Fix CPU entry area metadata lookup
  2026-10-06 18:10 [PATCH 0/2] x86/kmsan: Fix two bugs in the metadata lookup Viorel Cernateanu via B4 Relay
@ 2026-10-06 18:10 ` Viorel Cernateanu via B4 Relay
  2026-10-06 18:10 ` [PATCH 2/2] x86/kmsan: Don't call instrumented code from kmsan_virt_addr_valid() Viorel Cernateanu via B4 Relay
  1 sibling, 0 replies; 3+ messages in thread
From: Viorel Cernateanu via B4 Relay @ 2026-10-06 18:10 UTC (permalink / raw)
  To: Alexander Potapenko, Marco Elver, Dmitry Vyukov, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H. Peter Anvin,
	Andrew Morton, Peter Zijlstra, Kees Cook
  Cc: Borislav Petkov, kasan-dev, linux-kernel, Viorel Cernateanu

From: Viorel Cernateanu <vrilutza@gmail.com>

arch_kmsan_get_meta_or_null() works out which CPU owns an address in the
CPU entry area as (addr - CPU_ENTRY_AREA_BASE) / CPU_ENTRY_AREA_SIZE.
That is wrong in two ways:

 - With KASLR, init_cea_offsets() puts each CPU's area in a random slot,
   so the result is a slot number, usually far above NR_CPUS, and
   get_cpu_entry_area() reads past the end of __per_cpu_offset[].

 - The per-CPU areas start one page higher, at CPU_ENTRY_AREA_PER_CPU.
   Even without KASLR, the last page of each area is attributed to the
   next CPU, so KMSAN loses that page, and for the last CPU the lookup
   uses a CPU that does not exist.

On an Ivy Bridge laptop, a KMSAN kernel with KASLR oopses at boot when
the hard lockup detector sets up its perf event. In QEMU, a module that
reads one byte of the current CPU's GDT alias hits the same bug:

  UBSAN: array-index-out-of-bounds in arch/x86/mm/cpu_entry_area.c:25:9
  index 2100006 is out of range for type 'unsigned long[64]'
  Oops: general protection fault, probably for non-canonical address
  RIP: 0010:get_cpu_entry_area+0x14/0x50
  Call Trace:
   kmsan_get_metadata+0xb2/0x150
   kmsan_get_shadow_origin_ptr+0x31/0xa0
   __msan_metadata_ptr_for_load_1+0x22/0x40
   init_module+0x139/0xff0 [cea_repro]

Find the CPU whose entry area contains the address instead.

Fixes: ce732a7520b0 ("x86: kmsan: handle CPU entry area")
Fixes: 97e3d26b5e5f ("x86/mm: Randomize per-cpu entry area")
Signed-off-by: Viorel Cernateanu <vrilutza@gmail.com>
---
 arch/x86/include/asm/kmsan.h | 27 +++++++++++++++++++--------
 1 file changed, 19 insertions(+), 8 deletions(-)

diff --git a/arch/x86/include/asm/kmsan.h b/arch/x86/include/asm/kmsan.h
index d91b37f5b4bb..a71e84d6abab 100644
--- a/arch/x86/include/asm/kmsan.h
+++ b/arch/x86/include/asm/kmsan.h
@@ -13,6 +13,7 @@
 
 #include <asm/cpu_entry_area.h>
 #include <asm/processor.h>
+#include <linux/cpumask.h>
 #include <linux/mmzone.h>
 
 DECLARE_PER_CPU(char[CPU_ENTRY_AREA_SIZE], cpu_entry_area_shadow);
@@ -32,18 +33,28 @@ static inline void *arch_kmsan_get_meta_or_null(void *addr, bool is_origin)
 	unsigned long addr64 = (unsigned long)addr;
 	char *metadata_array;
 	unsigned long off;
-	int cpu;
+	unsigned int cpu;
 
 	if ((addr64 < CPU_ENTRY_AREA_BASE) ||
 	    (addr64 >= (CPU_ENTRY_AREA_BASE + CPU_ENTRY_AREA_MAP_SIZE)))
 		return NULL;
-	cpu = (addr64 - CPU_ENTRY_AREA_BASE) / CPU_ENTRY_AREA_SIZE;
-	off = addr64 - (unsigned long)get_cpu_entry_area(cpu);
-	if ((off < 0) || (off >= CPU_ENTRY_AREA_SIZE))
-		return NULL;
-	metadata_array = is_origin ? cpu_entry_area_origin :
-				     cpu_entry_area_shadow;
-	return &per_cpu(metadata_array[off], cpu);
+
+	/*
+	 * The per-CPU entry areas are not necessarily laid out in CPU order
+	 * (see init_cea_offsets()), so find the area containing @addr instead
+	 * of computing the CPU from the offset.
+	 */
+	for (cpu = 0; cpu < nr_cpu_ids; cpu++) {
+		if (!cpu_possible(cpu))
+			continue;
+		off = addr64 - (unsigned long)get_cpu_entry_area(cpu);
+		if (off >= CPU_ENTRY_AREA_SIZE)
+			continue;
+		metadata_array = is_origin ? cpu_entry_area_origin :
+					     cpu_entry_area_shadow;
+		return &per_cpu(metadata_array[off], cpu);
+	}
+	return NULL;
 }
 
 /*

-- 
2.53.0



^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 2/2] x86/kmsan: Don't call instrumented code from kmsan_virt_addr_valid()
  2026-10-06 18:10 [PATCH 0/2] x86/kmsan: Fix two bugs in the metadata lookup Viorel Cernateanu via B4 Relay
  2026-10-06 18:10 ` [PATCH 1/2] x86/kmsan: Fix CPU entry area " Viorel Cernateanu via B4 Relay
@ 2026-10-06 18:10 ` Viorel Cernateanu via B4 Relay
  1 sibling, 0 replies; 3+ messages in thread
From: Viorel Cernateanu via B4 Relay @ 2026-10-06 18:10 UTC (permalink / raw)
  To: Alexander Potapenko, Marco Elver, Dmitry Vyukov, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H. Peter Anvin,
	Andrew Morton, Peter Zijlstra, Kees Cook
  Cc: Borislav Petkov, kasan-dev, linux-kernel, Viorel Cernateanu

From: Viorel Cernateanu <vrilutza@gmail.com>

With CONFIG_DEBUG_PREEMPT, a KMSAN kernel hangs silently right after
KMSAN is enabled during boot.

kmsan_virt_addr_valid() calls preempt_disable() and pfn_valid(), and
pfn_valid() calls rcu_read_lock_sched(). With CONFIG_DEBUG_PREEMPT both
end up in preempt_count_add(), which is instrumented. Its instrumentation
looks up metadata, which calls kmsan_virt_addr_valid() again, and so on
until the boot stack overflows.

Use a copy of pfn_valid() that disables preemption with the notrace
helpers instead. It is still an RCU-sched read-side critical section,
and it avoids the instrumented preemption and RCU/lockdep helpers on
this path.

Fixes: f6564fce256a ("mm, kmsan: fix infinite recursion due to RCU critical section")
Signed-off-by: Viorel Cernateanu <vrilutza@gmail.com>
Link: https://lore.kernel.org/20240308043448.masllzeqwht45d4j@M910t
---
 arch/x86/include/asm/kmsan.h | 49 +++++++++++++++++++++++++++++---------------
 1 file changed, 33 insertions(+), 16 deletions(-)

diff --git a/arch/x86/include/asm/kmsan.h b/arch/x86/include/asm/kmsan.h
index a71e84d6abab..17effe8be642 100644
--- a/arch/x86/include/asm/kmsan.h
+++ b/arch/x86/include/asm/kmsan.h
@@ -68,6 +68,38 @@ static inline bool kmsan_phys_addr_valid(unsigned long addr)
 		return true;
 }
 
+/*
+ * Same as pfn_valid(), but without rcu_read_lock_sched(). That one calls into
+ * instrumented code: preempt_count_add() with CONFIG_DEBUG_PREEMPT or
+ * CONFIG_TRACE_PREEMPT_TOGGLE, lock_acquire() with CONFIG_DEBUG_LOCK_ALLOC.
+ * Instrumented code looks up metadata itself, so calling it from the metadata
+ * lookup can recurse; with CONFIG_DEBUG_PREEMPT it does, until the stack
+ * overflows.
+ *
+ * Disabling preemption still makes this an RCU-sched read-side critical
+ * section. Preemption is re-enabled without a reschedule, as before, to
+ * avoid entering the scheduler from the metadata lookup.
+ */
+static inline bool kmsan_pfn_valid(unsigned long pfn)
+{
+	struct mem_section *ms;
+	bool ret;
+
+	if (PHYS_PFN(PFN_PHYS(pfn)) != pfn)
+		return false;
+
+	if (pfn_to_section_nr(pfn) >= NR_MEM_SECTIONS)
+		return false;
+	ms = __pfn_to_section(pfn);
+
+	preempt_disable_notrace();
+	ret = valid_section(ms) &&
+	      (early_section(ms) || pfn_section_valid(ms, pfn));
+	preempt_enable_no_resched_notrace();
+
+	return ret;
+}
+
 /*
  * Taken from arch/x86/mm/physaddr.c to avoid using an instrumented version.
  */
@@ -75,7 +107,6 @@ static inline bool kmsan_virt_addr_valid(void *addr)
 {
 	unsigned long x = (unsigned long)addr;
 	unsigned long y = x - __START_KERNEL_map;
-	bool ret;
 
 	/* use the carry flag to determine if x was < __START_KERNEL_map */
 	if (unlikely(x > y)) {
@@ -91,21 +122,7 @@ static inline bool kmsan_virt_addr_valid(void *addr)
 			return false;
 	}
 
-	/*
-	 * pfn_valid() relies on RCU, and may call into the scheduler on exiting
-	 * the critical section. However, this would result in recursion with
-	 * KMSAN. Therefore, disable preemption here, and re-enable preemption
-	 * below while suppressing reschedules to avoid recursion.
-	 *
-	 * Note, this sacrifices occasionally breaking scheduling guarantees.
-	 * Although, a kernel compiled with KMSAN has already given up on any
-	 * performance guarantees due to being heavily instrumented.
-	 */
-	preempt_disable();
-	ret = pfn_valid(x >> PAGE_SHIFT);
-	preempt_enable_no_resched();
-
-	return ret;
+	return kmsan_pfn_valid(x >> PAGE_SHIFT);
 }
 
 #endif /* !MODULE */

-- 
2.53.0



^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-06 18:10 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 18:10 [PATCH 0/2] x86/kmsan: Fix two bugs in the metadata lookup Viorel Cernateanu via B4 Relay
2026-10-06 18:10 ` [PATCH 1/2] x86/kmsan: Fix CPU entry area " Viorel Cernateanu via B4 Relay
2026-10-06 18:10 ` [PATCH 2/2] x86/kmsan: Don't call instrumented code from kmsan_virt_addr_valid() Viorel Cernateanu via B4 Relay

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®