mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 0/2] tracing: Fix early boot trace buffer allocation and tracefs init crash
@ 2026-10-09 16:35 Amit Machhiwal
  2026-10-09 16:35 ` [PATCH 1/2] tracing: Do not initialize tracefs work if tracing is disabled Amit Machhiwal
  2026-10-09 16:35 ` [PATCH 2/2] ring-buffer: Do not check si_mem_available() during SYSTEM_BOOTING Amit Machhiwal
  0 siblings, 2 replies; 3+ messages in thread
From: Amit Machhiwal @ 2026-10-09 16:35 UTC (permalink / raw)
  To: Steven Rostedt, Masami Hiramatsu, linux-trace-kernel, linuxppc-dev
  Cc: Amit Machhiwal, Michal Suchánek, Mathieu Desnoyers,
	Madhavan Srinivasan, Ritesh Harjani (IBM),
	Harsh Prateek Bora, linux-kernel, kvm, kvm-ppc, Mark-PK Tsai,
	stable

This series addresses an early boot crash and trace buffer allocation
failure reported on ppc64le KVM guests [1] (and reproducible under
CONFIG_DEFERRED_STRUCT_PAGE_INIT=y with large static kernel footprints).

Patch 1 fixes the crash (NULL pointer dereference in __find_event_file()):
When trace buffer allocation fails during early boot, tracing is left
disabled.  However, tracer_init_tracefs_work_func() runs asynchronously via
fs_initcall and traverses uninitialized event lists without checking
tracing_disabled.  Checking tracing_disabled early averts the crash.

Patch 2 fixes the root cause of the early allocation failure:
During early_trace_init(), __rb_allocate_pages() performs a heuristic check
using si_mem_available().  Under CONFIG_DEFERRED_STRUCT_PAGE_INIT=y,
NR_FREE_PAGES only reflects the initial non-deferred pool (e.g. 1 section
per node) because the remaining memory has not yet been initialized.  On
kernels with large static binary footprints (debug configs, early
SLUB/vmalloc/ftrace records), this pool is quickly consumed, causing
si_mem_available() to return 0 and trigger a false -ENOMEM before the buddy
allocator has the opportunity to grow the zone on demand via
deferred_grow_zone().  Skipping the check during SYSTEM_BOOTING allows the
allocation to succeed.

Both fixes have been verified by Michal Suchánek.

[1] https://lore.kernel.org/all/arYskzbiaNzBR9MD@kunlun.suse.cz/

Amit Machhiwal (2):
  tracing: Do not initialize tracefs work if tracing is disabled
  ring-buffer: Do not check si_mem_available() during SYSTEM_BOOTING

 kernel/trace/ring_buffer.c | 8 +++++++-
 kernel/trace/trace.c       | 6 ++++++
 2 files changed, 13 insertions(+), 1 deletion(-)


base-commit: af32da41b0327b9c6a37856ba82b6760d6c8d10e
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 1/2] tracing: Do not initialize tracefs work if tracing is disabled
  2026-10-09 16:35 [PATCH 0/2] tracing: Fix early boot trace buffer allocation and tracefs init crash Amit Machhiwal
@ 2026-10-09 16:35 ` Amit Machhiwal
  2026-10-09 16:35 ` [PATCH 2/2] ring-buffer: Do not check si_mem_available() during SYSTEM_BOOTING Amit Machhiwal
  1 sibling, 0 replies; 3+ messages in thread
From: Amit Machhiwal @ 2026-10-09 16:35 UTC (permalink / raw)
  To: Steven Rostedt, Masami Hiramatsu, linux-trace-kernel, linuxppc-dev
  Cc: Amit Machhiwal, Michal Suchánek, Mathieu Desnoyers,
	Madhavan Srinivasan, Ritesh Harjani (IBM),
	Harsh Prateek Bora, linux-kernel, kvm, kvm-ppc, Mark-PK Tsai,
	stable

If trace buffer allocation fails during early boot (e.g.  in
early_trace_init() -> allocate_trace_buffers()), tracing_disabled is
left set to 1 and the global trace array events list remains
uninitialized.

Later during boot, tracer_init_tracefs() (invoked via fs_initcall)
queues tracerfs_init_work asynchronously without checking if tracing was
disabled.  When tracer_init_tracefs_work_func() runs, it calls
event_trace_init() and init_tracer_tracefs(&global_trace, NULL)
unconditionally.  This leads to __find_event_file() dereferencing an
uninitialized tr->events list, causing an early boot NULL pointer
dereference Oops:

  BUG: Kernel NULL pointer dereference on read at 0x00000010
  Faulting instruction address: 0xc000000000485be0
  Oops: Kernel access of bad area, sig: 7 [#1]
  ...
  NIP [c000000000485be0] __find_event_file+0x70/0x3c0
  LR  [c000000000448874] init_tracer_tracefs+0x274/0xc80
  Call Trace:
   init_tracer_tracefs+0x274/0xc80
   tracer_init_tracefs_work_func+0x50/0x320
   process_one_work+0x1e8/0x5c0
   worker_thread+0x1dc/0x3d0
   kthread+0x194/0x1b0
   start_kernel_thread+0x14/0x18

Check tracing_disabled at the beginning of tracer_init_tracefs_work_func()
and return immediately if tracing is disabled or failed to initialize,
preventing the crash.

Fixes: 6621a7004684 ("tracing: make tracer_init_tracefs initcall asynchronous")
Cc: stable@vger.kernel.org
Reported-by: Michal Suchánek <msuchanek@suse.de>
Closes: https://lore.kernel.org/all/arYskzbiaNzBR9MD@kunlun.suse.cz/
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
 kernel/trace/trace.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/kernel/trace/trace.c b/kernel/trace/trace.c
index e4a490d3d08c..7ac9e1c92702 100644
--- a/kernel/trace/trace.c
+++ b/kernel/trace/trace.c
@@ -9287,6 +9287,12 @@ static struct notifier_block trace_module_nb = {
 
 static __init void tracer_init_tracefs_work_func(struct work_struct *work)
 {
+	/*
+	 * Do not attempt to initialize tracefs if tracing was disabled or
+	 * failed to allocate
+	 */
+	if (tracing_disabled)
+		return;
 
 	event_trace_init();
 
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 3+ messages in thread

* [PATCH 2/2] ring-buffer: Do not check si_mem_available() during SYSTEM_BOOTING
  2026-10-09 16:35 [PATCH 0/2] tracing: Fix early boot trace buffer allocation and tracefs init crash Amit Machhiwal
  2026-10-09 16:35 ` [PATCH 1/2] tracing: Do not initialize tracefs work if tracing is disabled Amit Machhiwal
@ 2026-10-09 16:35 ` Amit Machhiwal
  1 sibling, 0 replies; 3+ messages in thread
From: Amit Machhiwal @ 2026-10-09 16:35 UTC (permalink / raw)
  To: Steven Rostedt, Masami Hiramatsu, linux-trace-kernel, linuxppc-dev
  Cc: Amit Machhiwal, Michal Suchánek, Mathieu Desnoyers,
	Madhavan Srinivasan, Ritesh Harjani (IBM),
	Harsh Prateek Bora, linux-kernel, kvm, kvm-ppc, Mark-PK Tsai,
	stable

During early boot (early_trace_init() called from start_kernel()),
tracing allocates initial ring buffers (temp_buffer, array_buffer, and
snapshot_buffer) for CPU 0.

Before attempting page allocation, __rb_allocate_pages() performs a
heuristic check using si_mem_available() to return early with -ENOMEM if
memory appears insufficient.

However, on kernels built with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y,
defer_init() leaves only a single section per node (e.g.  16 MiB with 64
KB pages) initialized up-front.  On kernels with large static binary
footprints (such as debug configurations enabling PAGE_OWNER,
DEBUG_PAGEALLOC, KFENCE, or SLUB_DEBUG), the static kernel image and
early core initialisations (SLUB caches, vmalloc, static ftrace records)
consume virtually all managed pages in this initial pool.

At T=0.000000, watermarks have not yet been established
(totalreserve_pages = 0), so si_mem_available() returns the raw free
page count (often 0-1 pages).  When the snapshot buffer or global trace
buffer attempts to allocate 2 sub-pages, si_mem_available() returns <
nr_pages and prematurely aborts with -ENOMEM.

This failure is false: if the page allocation were actually attempted
via alloc_pages_node(), the page allocator would trigger
deferred_grow_zone() on demand to initialise additional deferred memory
sections.  Checking si_mem_available() before attempting the allocation
short-circuits this on-demand growth.

Skip the si_mem_available() check when system_state == SYSTEM_BOOTING.
Once the system transitions past early boot and page_alloc_init_late()
initialises all deferred memory, si_mem_available() accurately reflects
system-wide free memory and the check operates as intended for runtime
allocations.

Fixes: 2a872fa4e9c8 ("ring-buffer: Check if memory is available before allocation")
Cc: stable@vger.kernel.org
Reported-by: Michal Suchánek <msuchanek@suse.de>
Closes: https://lore.kernel.org/all/arYskzbiaNzBR9MD@kunlun.suse.cz/
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
 kernel/trace/ring_buffer.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c
index 04bb94c29f58..a9f82e2f8fad 100644
--- a/kernel/trace/ring_buffer.c
+++ b/kernel/trace/ring_buffer.c
@@ -2452,9 +2452,15 @@ static int __rb_allocate_pages(struct ring_buffer_per_cpu *cpu_buffer,
 	 * memory. It may not be accurate. But we don't care, we just want
 	 * to prevent doing any allocation when it is obvious that it is
 	 * not going to succeed.
+	 *
+	 * Skip this check during early boot: with CONFIG_DEFERRED_STRUCT_PAGE_INIT,
+	 * NR_FREE_PAGES only reflects the initial non-deferred pool at this
+	 * stage. si_mem_available() returns a false negative while actual
+	 * allocations succeed by growing the zone on demand via
+	 * deferred_grow_zone().
 	 */
 	i = si_mem_available();
-	if (i < nr_pages)
+	if (system_state != SYSTEM_BOOTING && i < nr_pages)
 		return -ENOMEM;
 
 	/*
-- 
2.54.0 (Apple Git-157)


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-09 16:36 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 16:35 [PATCH 0/2] tracing: Fix early boot trace buffer allocation and tracefs init crash Amit Machhiwal
2026-10-09 16:35 ` [PATCH 1/2] tracing: Do not initialize tracefs work if tracing is disabled Amit Machhiwal
2026-10-09 16:35 ` [PATCH 2/2] ring-buffer: Do not check si_mem_available() during SYSTEM_BOOTING Amit Machhiwal

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®