From: Yuanhe Shu <xiangzao@linux.alibaba.com>
To: tglx@kernel.org, mingo@redhat.com, bp@alien8.de,
dave.hansen@linux.intel.com, x86@kernel.org
Cc: hpa@zytor.com, xin@zytor.com, luto@kernel.org,
jpoimboe@kernel.org, peterz@infradead.org, rostedt@goodmis.org,
akpm@linux-foundation.org, rdunlap@infradead.org,
elver@google.com, andreyknvl@gmail.com, glider@google.com,
gor@linux.ibm.com, brads@mainlining.org,
kasan-dev@googlegroups.com, linux-kernel@vger.kernel.org,
Yuanhe Shu <xiangzao@linux.alibaba.com>
Subject: [PATCH v3 0/2] x86/fred: Fix stack depot exhaustion on FRED systems
Date: Tue, 1 Sep 2026 20:40:11 +0800 [thread overview]
Message-ID: <20260901124013.1098560-1-xiangzao@linux.alibaba.com> (raw)
FRED replaces the IDT event entry stubs with asm_fred_entrypoint_user and
asm_fred_entrypoint_kernel, which live in .noinstr.text. The IDT stubs
are the only code covered by __irqentry_text_start..__irqentry_text_end,
which in_irqentry_text() uses to find where an event entered the kernel.
With FRED the detection fails, filter_irq_stacks() no longer truncates
event stacks, and every trace saved from interrupt or exception context
by stack depot users (KASAN alloc/free tracking, SLUB object tracking,
...) becomes a unique combination of "event path x arbitrarily
interrupted context". The depot grows without bound until it is
exhausted.
What we saw: a FRED-capable dual-socket system with KASAN (generic,
inline) and SLUB object tracking enabled hit the limit on both sockets
roughly 80 minutes after boot,
Stack depot reached limit capacity
WARNING: CPU: 286 PID: 128614 at lib/stackdepot.c:271 depot_alloc_stack+0x158/0x170
after which KASAN and SLUB stack tracking silently stop recording (new
traces get handle 0) until reboot. The recorded traces confirmed the
mechanism: interrupt side traces ran through asm_fred_entrypoint_kernel
and continued into the frames of the interrupted task.
Patch 1 adds an optional arch_in_event_entry_text() hook to the generic
entry text check, and renames that check from in_irqentry_text() to
in_event_entry_text() since neither it nor the ranges it tests are
necessarily irq-only; no behavior change by itself. Patch 2
implements it for x86 by bracketing the FRED entry text with
__fred_entry_text_start/end, the same way the IDT stubs are bracketed,
and reporting that range from <asm/sections.h>.
Unlike arm64 and s390, which hit the same symptom and could fix it by
placing their interrupt entry code in .irqentry.text, x86 has no such
section: it emits the markers as labels around the sequentially laid out
IDT stubs and defines __irq_entry to __invalid_section so that nothing
can land in that section. The FRED entry points therefore cannot be
brought inside the existing range, and a second one has to be reported;
see patch 2 for the details.
Testing:
- Build-tested on mainline: full vmlinux builds and link with
CONFIG_X86_FRED=y, CONFIG_X86_KERNEL_IBT=y and KASAN (generic,
inline), with and without CONFIG_KVM_INTEL=y, and with
CONFIG_X86_FRED=n so that the generic fallback is used; objtool link
stage clean.
- Runtime-tested by backporting both patches to the affected kernel (a
6.6 based debug build) and rebooting that system with an unchanged
command line: traces recorded from FRED event context now end at
asm_fred_entrypoint_kernel instead of continuing into the
interrupted task, and the pool count levels off at ~1230 pools
within minutes of boot instead of reaching the 8192 pool limit.
Reproducing on current upstream:
- On FRED hardware, any KASAN or SLUB_DEBUG kernel; on v6.17 and
later, boot with stack_depot_max_pools=1024 to bring the
exhaustion warning up in minutes instead of hours, and watch the
pool count in /sys/kernel/debug/stackdepot/stats grow
monotonically without the fix and converge with it.
- Without FRED hardware, the KVM VMX interrupt forwarding path
(CONFIG_X86_FRED=y + CONFIG_KVM_INTEL) exercises the new filtering
as well, see patch 2.
Both patches carry the same Fixes: tag and are Cc'ed to stable, and 2/2
uses the hook added by 1/2, so they need to be applied together.
Changes in v3:
- 1/2: drop the x86 details from the generic comment and phrase it as an
instruction for future arch maintainers (Dave Hansen)
- 1/2: restore Bradley Morgan's Reviewed-by - he confirmed the tag still
holds for the reworked version
- 2/2: state the caveat that the range cannot tell interrupt entry and
syscall entry apart here rather than in the generic comment
Changes in v2:
- 1/2: drop the new ARCH_HAS_IN_IRQENTRY_TEXT Kconfig symbol and use the
#ifndef macro override pattern instead, with the fallback next to its
only user in kernel/stacktrace.c (Peter Zijlstra)
- 1/2: rename in_irqentry_text() to in_event_entry_text() and name the
hook after it - the ranges are not necessarily irq-only; on x86 they
cover the exception entry stubs too (Peter Zijlstra)
- 1/2: carry the same Fixes: tag as 2/2 (Andrew Morton)
- 1/2: drop Bradley Morgan's Reviewed-by since the patch changed
materially
- 2/2: implement the hook as a static inline in <asm/sections.h> next to
the markers it tests, instead of a Kconfig gated function in
arch/x86/kernel/stacktrace.c; arch/x86/Kconfig is no longer touched
- 2/2: spell out, in the changelog and in a comment on the hook, why the
range covering the ring 3 entry point, and therefore syscalls, does not
change any trace (Peter Zijlstra)
- 2/2: add Reported-by for Xiang Zheng, who found the exhaustion
v1: https://lore.kernel.org/r/20260827150022.1618235-1-xiangzao@linux.alibaba.com
v2: https://lore.kernel.org/r/20260831120221.2874445-1-xiangzao@linux.alibaba.com
Yuanhe Shu (2):
stacktrace: Provide arch_in_event_entry_text() hook
x86/fred: Fix stack depot filtering of FRED event stacks
arch/x86/entry/entry_64_fred.S | 14 ++++++++++++++
arch/x86/include/asm/sections.h | 24 ++++++++++++++++++++++++
kernel/stacktrace.c | 19 ++++++++++++++++---
3 files changed, 54 insertions(+), 3 deletions(-)
base-commit: 45c13f3f9e3bb15fd89ff2864c6f627a3b4b4229
--
2.43.7
next reply other threads:[~2026-09-01 12:40 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-01 12:40 Yuanhe Shu [this message]
2026-09-01 12:40 ` [PATCH v3 1/2] stacktrace: Provide arch_in_event_entry_text() hook Yuanhe Shu
2026-09-02 13:01 ` Alexander Potapenko
2026-09-01 12:40 ` [PATCH v3 2/2] x86/fred: Fix stack depot filtering of FRED event stacks Yuanhe Shu
2026-09-02 13:17 ` Alexander Potapenko
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260901124013.1098560-1-xiangzao@linux.alibaba.com \
--to=xiangzao@linux.alibaba.com \
--cc=akpm@linux-foundation.org \
--cc=andreyknvl@gmail.com \
--cc=bp@alien8.de \
--cc=brads@mainlining.org \
--cc=dave.hansen@linux.intel.com \
--cc=elver@google.com \
--cc=glider@google.com \
--cc=gor@linux.ibm.com \
--cc=hpa@zytor.com \
--cc=jpoimboe@kernel.org \
--cc=kasan-dev@googlegroups.com \
--cc=linux-kernel@vger.kernel.org \
--cc=luto@kernel.org \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rdunlap@infradead.org \
--cc=rostedt@goodmis.org \
--cc=tglx@kernel.org \
--cc=x86@kernel.org \
--cc=xin@zytor.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®