mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Yuanhe Shu <xiangzao@linux.alibaba.com>
To: tglx@kernel.org, mingo@redhat.com, bp@alien8.de,
	dave.hansen@linux.intel.com, x86@kernel.org
Cc: hpa@zytor.com, xin@zytor.com, luto@kernel.org,
	jpoimboe@kernel.org, peterz@infradead.org, rostedt@goodmis.org,
	akpm@linux-foundation.org, rdunlap@infradead.org,
	elver@google.com, andreyknvl@gmail.com, glider@google.com,
	gor@linux.ibm.com, brads@mainlining.org,
	kasan-dev@googlegroups.com, linux-kernel@vger.kernel.org,
	Yuanhe Shu <xiangzao@linux.alibaba.com>
Subject: [PATCH v3 0/2] x86/fred: Fix stack depot exhaustion on FRED systems
Date: Tue,  1 Sep 2026 20:40:11 +0800	[thread overview]
Message-ID: <20260901124013.1098560-1-xiangzao@linux.alibaba.com> (raw)

FRED replaces the IDT event entry stubs with asm_fred_entrypoint_user and
asm_fred_entrypoint_kernel, which live in .noinstr.text.  The IDT stubs
are the only code covered by __irqentry_text_start..__irqentry_text_end,
which in_irqentry_text() uses to find where an event entered the kernel.
With FRED the detection fails, filter_irq_stacks() no longer truncates
event stacks, and every trace saved from interrupt or exception context
by stack depot users (KASAN alloc/free tracking, SLUB object tracking,
...) becomes a unique combination of "event path x arbitrarily
interrupted context".  The depot grows without bound until it is
exhausted.

What we saw: a FRED-capable dual-socket system with KASAN (generic,
inline) and SLUB object tracking enabled hit the limit on both sockets
roughly 80 minutes after boot,

  Stack depot reached limit capacity
  WARNING: CPU: 286 PID: 128614 at lib/stackdepot.c:271 depot_alloc_stack+0x158/0x170

after which KASAN and SLUB stack tracking silently stop recording (new
traces get handle 0) until reboot.  The recorded traces confirmed the
mechanism: interrupt side traces ran through asm_fred_entrypoint_kernel
and continued into the frames of the interrupted task.

Patch 1 adds an optional arch_in_event_entry_text() hook to the generic
entry text check, and renames that check from in_irqentry_text() to
in_event_entry_text() since neither it nor the ranges it tests are
necessarily irq-only; no behavior change by itself.  Patch 2
implements it for x86 by bracketing the FRED entry text with
__fred_entry_text_start/end, the same way the IDT stubs are bracketed,
and reporting that range from <asm/sections.h>.

Unlike arm64 and s390, which hit the same symptom and could fix it by
placing their interrupt entry code in .irqentry.text, x86 has no such
section: it emits the markers as labels around the sequentially laid out
IDT stubs and defines __irq_entry to __invalid_section so that nothing
can land in that section.  The FRED entry points therefore cannot be
brought inside the existing range, and a second one has to be reported;
see patch 2 for the details.

Testing:

  - Build-tested on mainline: full vmlinux builds and link with
    CONFIG_X86_FRED=y, CONFIG_X86_KERNEL_IBT=y and KASAN (generic,
    inline), with and without CONFIG_KVM_INTEL=y, and with
    CONFIG_X86_FRED=n so that the generic fallback is used; objtool link
    stage clean.
  - Runtime-tested by backporting both patches to the affected kernel (a
    6.6 based debug build) and rebooting that system with an unchanged
    command line: traces recorded from FRED event context now end at
    asm_fred_entrypoint_kernel instead of continuing into the
    interrupted task, and the pool count levels off at ~1230 pools
    within minutes of boot instead of reaching the 8192 pool limit.

Reproducing on current upstream:

  - On FRED hardware, any KASAN or SLUB_DEBUG kernel; on v6.17 and
    later, boot with stack_depot_max_pools=1024 to bring the
    exhaustion warning up in minutes instead of hours, and watch the
    pool count in /sys/kernel/debug/stackdepot/stats grow
    monotonically without the fix and converge with it.
  - Without FRED hardware, the KVM VMX interrupt forwarding path
    (CONFIG_X86_FRED=y + CONFIG_KVM_INTEL) exercises the new filtering
    as well, see patch 2.

Both patches carry the same Fixes: tag and are Cc'ed to stable, and 2/2
uses the hook added by 1/2, so they need to be applied together.

Changes in v3:

  - 1/2: drop the x86 details from the generic comment and phrase it as an
    instruction for future arch maintainers (Dave Hansen)
  - 1/2: restore Bradley Morgan's Reviewed-by - he confirmed the tag still
    holds for the reworked version
  - 2/2: state the caveat that the range cannot tell interrupt entry and
    syscall entry apart here rather than in the generic comment

Changes in v2:

  - 1/2: drop the new ARCH_HAS_IN_IRQENTRY_TEXT Kconfig symbol and use the
    #ifndef macro override pattern instead, with the fallback next to its
    only user in kernel/stacktrace.c (Peter Zijlstra)
  - 1/2: rename in_irqentry_text() to in_event_entry_text() and name the
    hook after it - the ranges are not necessarily irq-only; on x86 they
    cover the exception entry stubs too (Peter Zijlstra)
  - 1/2: carry the same Fixes: tag as 2/2 (Andrew Morton)
  - 1/2: drop Bradley Morgan's Reviewed-by since the patch changed
    materially
  - 2/2: implement the hook as a static inline in <asm/sections.h> next to
    the markers it tests, instead of a Kconfig gated function in
    arch/x86/kernel/stacktrace.c; arch/x86/Kconfig is no longer touched
  - 2/2: spell out, in the changelog and in a comment on the hook, why the
    range covering the ring 3 entry point, and therefore syscalls, does not
    change any trace (Peter Zijlstra)
  - 2/2: add Reported-by for Xiang Zheng, who found the exhaustion

v1: https://lore.kernel.org/r/20260827150022.1618235-1-xiangzao@linux.alibaba.com
v2: https://lore.kernel.org/r/20260831120221.2874445-1-xiangzao@linux.alibaba.com

Yuanhe Shu (2):
  stacktrace: Provide arch_in_event_entry_text() hook
  x86/fred: Fix stack depot filtering of FRED event stacks

 arch/x86/entry/entry_64_fred.S  | 14 ++++++++++++++
 arch/x86/include/asm/sections.h | 24 ++++++++++++++++++++++++
 kernel/stacktrace.c             | 19 ++++++++++++++++---
 3 files changed, 54 insertions(+), 3 deletions(-)


base-commit: 45c13f3f9e3bb15fd89ff2864c6f627a3b4b4229
-- 
2.43.7


             reply	other threads:[~2026-09-01 12:40 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-01 12:40 Yuanhe Shu [this message]
2026-09-01 12:40 ` [PATCH v3 1/2] stacktrace: Provide arch_in_event_entry_text() hook Yuanhe Shu
2026-09-02 13:01   ` Alexander Potapenko
2026-09-01 12:40 ` [PATCH v3 2/2] x86/fred: Fix stack depot filtering of FRED event stacks Yuanhe Shu
2026-09-02 13:17   ` Alexander Potapenko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260901124013.1098560-1-xiangzao@linux.alibaba.com \
    --to=xiangzao@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=andreyknvl@gmail.com \
    --cc=bp@alien8.de \
    --cc=brads@mainlining.org \
    --cc=dave.hansen@linux.intel.com \
    --cc=elver@google.com \
    --cc=glider@google.com \
    --cc=gor@linux.ibm.com \
    --cc=hpa@zytor.com \
    --cc=jpoimboe@kernel.org \
    --cc=kasan-dev@googlegroups.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=luto@kernel.org \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rdunlap@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=tglx@kernel.org \
    --cc=x86@kernel.org \
    --cc=xin@zytor.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®