From: Amit Machhiwal <amachhiw@linux.ibm.com>
To: Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
linux-trace-kernel@vger.kernel.org,
linuxppc-dev@lists.ozlabs.org
Cc: "Amit Machhiwal" <amachhiw@linux.ibm.com>,
"Michal Suchánek" <msuchanek@suse.de>,
"Mathieu Desnoyers" <mathieu.desnoyers@efficios.com>,
"Madhavan Srinivasan" <maddy@linux.ibm.com>,
"Ritesh Harjani (IBM)" <ritesh.list@gmail.com>,
"Harsh Prateek Bora" <harshpb@linux.ibm.com>,
linux-kernel@vger.kernel.org, kvm@vger.kernel.org,
kvm-ppc@vger.kernel.org,
"Mark-PK Tsai" <mark-pk.tsai@mediatek.com>,
stable@vger.kernel.org
Subject: [PATCH 2/2] ring-buffer: Do not check si_mem_available() during SYSTEM_BOOTING
Date: Fri, 9 Oct 2026 22:05:57 +0530 [thread overview]
Message-ID: <20261009163557.88467-3-amachhiw@linux.ibm.com> (raw)
In-Reply-To: <20261009163557.88467-1-amachhiw@linux.ibm.com>
During early boot (early_trace_init() called from start_kernel()),
tracing allocates initial ring buffers (temp_buffer, array_buffer, and
snapshot_buffer) for CPU 0.
Before attempting page allocation, __rb_allocate_pages() performs a
heuristic check using si_mem_available() to return early with -ENOMEM if
memory appears insufficient.
However, on kernels built with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y,
defer_init() leaves only a single section per node (e.g. 16 MiB with 64
KB pages) initialized up-front. On kernels with large static binary
footprints (such as debug configurations enabling PAGE_OWNER,
DEBUG_PAGEALLOC, KFENCE, or SLUB_DEBUG), the static kernel image and
early core initialisations (SLUB caches, vmalloc, static ftrace records)
consume virtually all managed pages in this initial pool.
At T=0.000000, watermarks have not yet been established
(totalreserve_pages = 0), so si_mem_available() returns the raw free
page count (often 0-1 pages). When the snapshot buffer or global trace
buffer attempts to allocate 2 sub-pages, si_mem_available() returns <
nr_pages and prematurely aborts with -ENOMEM.
This failure is false: if the page allocation were actually attempted
via alloc_pages_node(), the page allocator would trigger
deferred_grow_zone() on demand to initialise additional deferred memory
sections. Checking si_mem_available() before attempting the allocation
short-circuits this on-demand growth.
Skip the si_mem_available() check when system_state == SYSTEM_BOOTING.
Once the system transitions past early boot and page_alloc_init_late()
initialises all deferred memory, si_mem_available() accurately reflects
system-wide free memory and the check operates as intended for runtime
allocations.
Fixes: 2a872fa4e9c8 ("ring-buffer: Check if memory is available before allocation")
Cc: stable@vger.kernel.org
Reported-by: Michal Suchánek <msuchanek@suse.de>
Closes: https://lore.kernel.org/all/arYskzbiaNzBR9MD@kunlun.suse.cz/
Signed-off-by: Amit Machhiwal <amachhiw@linux.ibm.com>
---
kernel/trace/ring_buffer.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c
index 04bb94c29f58..a9f82e2f8fad 100644
--- a/kernel/trace/ring_buffer.c
+++ b/kernel/trace/ring_buffer.c
@@ -2452,9 +2452,15 @@ static int __rb_allocate_pages(struct ring_buffer_per_cpu *cpu_buffer,
* memory. It may not be accurate. But we don't care, we just want
* to prevent doing any allocation when it is obvious that it is
* not going to succeed.
+ *
+ * Skip this check during early boot: with CONFIG_DEFERRED_STRUCT_PAGE_INIT,
+ * NR_FREE_PAGES only reflects the initial non-deferred pool at this
+ * stage. si_mem_available() returns a false negative while actual
+ * allocations succeed by growing the zone on demand via
+ * deferred_grow_zone().
*/
i = si_mem_available();
- if (i < nr_pages)
+ if (system_state != SYSTEM_BOOTING && i < nr_pages)
return -ENOMEM;
/*
--
2.54.0 (Apple Git-157)
prev parent reply other threads:[~2026-10-09 16:36 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 16:35 [PATCH 0/2] tracing: Fix early boot trace buffer allocation and tracefs init crash Amit Machhiwal
2026-10-09 16:35 ` [PATCH 1/2] tracing: Do not initialize tracefs work if tracing is disabled Amit Machhiwal
2026-10-09 16:35 ` Amit Machhiwal [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261009163557.88467-3-amachhiw@linux.ibm.com \
--to=amachhiw@linux.ibm.com \
--cc=harshpb@linux.ibm.com \
--cc=kvm-ppc@vger.kernel.org \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=maddy@linux.ibm.com \
--cc=mark-pk.tsai@mediatek.com \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=msuchanek@suse.de \
--cc=ritesh.list@gmail.com \
--cc=rostedt@goodmis.org \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®