mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Amit Machhiwal <amachhiw@linux.ibm.com>
To: Steven Rostedt <rostedt@goodmis.org>,
	Masami Hiramatsu <mhiramat@kernel.org>,
	linux-trace-kernel@vger.kernel.org,
	linuxppc-dev@lists.ozlabs.org
Cc: "Amit Machhiwal" <amachhiw@linux.ibm.com>,
	"Michal Suchánek" <msuchanek@suse.de>,
	"Mathieu Desnoyers" <mathieu.desnoyers@efficios.com>,
	"Madhavan Srinivasan" <maddy@linux.ibm.com>,
	"Ritesh Harjani (IBM)" <ritesh.list@gmail.com>,
	"Harsh Prateek Bora" <harshpb@linux.ibm.com>,
	linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
	kvm-ppc@vger.kernel.org,
	"Mark-PK Tsai" <mark-pk.tsai@mediatek.com>,
	stable@vger.kernel.org
Subject: [PATCH 2/2] ring-buffer: Do not check si_mem_available() during SYSTEM_BOOTING
Date: Fri,  9 Oct 2026 22:05:57 +0530	[thread overview]
Message-ID: <20261009163557.88467-3-amachhiw@linux.ibm.com> (raw)
In-Reply-To: <20261009163557.88467-1-amachhiw@linux.ibm.com>

During early boot (early_trace_init() called from start_kernel()),
tracing allocates initial ring buffers (temp_buffer, array_buffer, and
snapshot_buffer) for CPU 0.

Before attempting page allocation, __rb_allocate_pages() performs a
heuristic check using si_mem_available() to return early with -ENOMEM if
memory appears insufficient.

However, on kernels built with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y,
defer_init() leaves only a single section per node (e.g.  16 MiB with 64
KB pages) initialized up-front.  On kernels with large static binary
footprints (such as debug configurations enabling PAGE_OWNER,
DEBUG_PAGEALLOC, KFENCE, or SLUB_DEBUG), the static kernel image and
early core initialisations (SLUB caches, vmalloc, static ftrace records)
consume virtually all managed pages in this initial pool.

At T=0.000000, watermarks have not yet been established
(totalreserve_pages = 0), so si_mem_available() returns the raw free
page count (often 0-1 pages).  When the snapshot buffer or global trace
buffer attempts to allocate 2 sub-pages, si_mem_available() returns <
nr_pages and prematurely aborts with -ENOMEM.

This failure is false: if the page allocation were actually attempted
via alloc_pages_node(), the page allocator would trigger
deferred_grow_zone() on demand to initialise additional deferred memory
sections.  Checking si_mem_available() before attempting the allocation
short-circuits this on-demand growth.

Skip the si_mem_available() check when system_state == SYSTEM_BOOTING.
Once the system transitions past early boot and page_alloc_init_late()
initialises all deferred memory, si_mem_available() accurately reflects
system-wide free memory and the check operates as intended for runtime
allocations.

Fixes: 2a872fa4e9c8 ("ring-buffer: Check if memory is available before allocation")
Cc: stable@vger.kernel.org
Reported-by: Michal Suchánek <msuchanek@suse.de>
Closes: https://lore.kernel.org/all/arYskzbiaNzBR9MD@kunlun.suse.cz/
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
 kernel/trace/ring_buffer.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c
index 04bb94c29f58..a9f82e2f8fad 100644
--- a/kernel/trace/ring_buffer.c
+++ b/kernel/trace/ring_buffer.c
@@ -2452,9 +2452,15 @@ static int __rb_allocate_pages(struct ring_buffer_per_cpu *cpu_buffer,
 	 * memory. It may not be accurate. But we don't care, we just want
 	 * to prevent doing any allocation when it is obvious that it is
 	 * not going to succeed.
+	 *
+	 * Skip this check during early boot: with CONFIG_DEFERRED_STRUCT_PAGE_INIT,
+	 * NR_FREE_PAGES only reflects the initial non-deferred pool at this
+	 * stage. si_mem_available() returns a false negative while actual
+	 * allocations succeed by growing the zone on demand via
+	 * deferred_grow_zone().
 	 */
 	i = si_mem_available();
-	if (i < nr_pages)
+	if (system_state != SYSTEM_BOOTING && i < nr_pages)
 		return -ENOMEM;
 
 	/*
-- 
2.54.0 (Apple Git-157)


      parent reply	other threads:[~2026-10-09 16:36 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09 16:35 [PATCH 0/2] tracing: Fix early boot trace buffer allocation and tracefs init crash Amit Machhiwal
2026-10-09 16:35 ` [PATCH 1/2] tracing: Do not initialize tracefs work if tracing is disabled Amit Machhiwal
2026-10-09 16:35 ` Amit Machhiwal [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261009163557.88467-3-amachhiw@linux.ibm.com \
    --to=amachhiw@linux.ibm.com \
    --cc=harshpb@linux.ibm.com \
    --cc=kvm-ppc@vger.kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=linuxppc-dev@lists.ozlabs.org \
    --cc=maddy@linux.ibm.com \
    --cc=mark-pk.tsai@mediatek.com \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mhiramat@kernel.org \
    --cc=msuchanek@suse.de \
    --cc=ritesh.list@gmail.com \
    --cc=rostedt@goodmis.org \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®