* [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places
@ 2026-09-13 22:28 Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events Arnaldo Carvalho de Melo
` (7 more replies)
0 siblings, 8 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
Hi,
This came out of work on data type profiling, but none of them depends on that
work and nothing in this series needs anything from it, so sending them
separately.
The fixes:
- 'perf report -s type' spins forever, burning all of a CPU with no
output, on the dwz compressed debug info of zlib-ng (libz.so.1):
die_collect_vars() saves the dwarf_dieoffset() of the type DIE,
which is relative to the file that DIE lives in, the dwz common
file for the types shared by more than one CU, and resolving one of
those offsets in the main debug file does not fail, it parses
whatever is at that offset there, in this case a typedef whose
DW_AT_type refers to itself, which is what makes the "follow the
typedefs and qualifiers until a pointer or an array type" loop in
die_get_pointer_type() spin. The offset is now resolved in the file
it was recorded as coming from, which the CU of the collected DIE
tells exactly rather than having to be inferred from what is at the
offset, and the type chases and the struct/union member nesting are
bounded, so a debug info file broken in some other way makes perf
give up on a type with a pr_debug instead of looking like it hung
(patch 7);
- the data type browser prints, in its samples view (-n,
annotate.show_nr_samples), a local variable initialized to zero and
never updated, so every member shows up as having no samples while
the period and percent columns for the same entry are filled in
(patch 4);
- the 'perf data type profiling' shell test turns "this PMU cannot
record these events" into a test failure: the script runs under
'set -e', so the bare 'perf mem record' aborts it through the EXIT
trap, which reports a signal that never happened, instead of
reaching the code right below meant to report the failure (patch 1).
The features:
- 'perf report --progress' prints the phase it is in and how far
along it is to stderr: the TUI shows that with a progress bar, but
in the stdio case the ui_progress updates that are already there are
dropped on the floor. It is what makes a slow run tell itself apart
from a stuck one (patch 5);
- the debuginfo for a DSO, and the kernel symbols, can now be fetched
keyed by the build ID recorded in the perf.data file, using the
debuginfod client, for the cases where they are not available
locally under the name the DSO was opened with: a vmlinux for a
kernel that since got upgraded, or a profile recorded on another
machine. Both are off with --no-debuginfod, with
core.debuginfod=false, and when the build-id cache is turned off
(patches 2 and 3);
- 'perf mem record' asks for PERF_SAMPLE_CPU by default, as without
it the cpu field is the (u32)-1 "no CPU info" sentinel and
per-sample analysis cannot tell accesses from different cores apart
from same-CPU traffic (patch 8).
Which file a saved type DIE offset belongs to is settled with
dwarf_cu_getdwarf(), new in elfutils 0.160, so the libdw feature test
now probes for it and the message in Makefile.config says 0.160 where it
said 0.157: 0.157 to 0.159 is from 2014 and has no such symbol, and
those versions now lose dwarf support with that message instead of
failing to link util/dwarf-aux.c.
Best regards,
- Arnaldo
What changed from v2 (00cffc7df5abfa68):
PATCH 2/8, "perf debuginfo: Fetch debuginfo keyed by build ID using
debuginfod" [sashiko-bot review of PATCH v2 2/8]:
- The conditions to fetch, debuginfod being enabled, the build-id cache
being on and the build ID not having been a miss already, are now
looked at with the fetch lock held, as they are the state the fetch
itself changes: deciding them outside the lock and fetching inside it
let a second thread repeat the fetch the first one had just made, or
repeat one that had just been turned off, and put the same build ID
on the misses list twice.
- A build ID that is already being fetched is waited for, instead of
fetched again: there was nothing that said "this one is in flight", so
a second thread wanting it blocked on the fetch lock and then did its
own complete fetch, which, while the first one is still running, is a
second download of the same file, with a second client, a second
progress line and a second turn at the terminal, for the same answer.
It now sleeps on a condition variable and is woken with whatever the
fetch settled: the file, when there is one, or the not available, when
it found nothing or the user cancelled it, with no retry.
- And a fetch that brought a file back is remembered, so that the ones
that need the same build ID later, e.g. annotating another symbol of
the same DSO, are answered with a strdup() instead of another client:
the file stays in the debuginfod client cache, and its path is checked
before being handed out, in case that cache gets cleaned from under
us.
- The fetches are still one at a time, so a fetch for one build ID waits
for a fetch for another one: the terminal mode, the signal
dispositions, the progress line and the 's'/'d' keys are process
global, and sharing them between concurrent fetches, refcounted, is
left for later, noted in the code. It matters less than it sounds,
now that the build IDs of a workload are answered from memory after
the first pass over them.
PATCH 3/8, "perf symbol: Fall back to fetching the vmlinux by build ID"
[sashiko-bot review of PATCH v2 3/8]:
- The fetch is no longer made with dso->lock held: dso__load() holds it
across dso__load_kernel_sym(), and a fetch can block for a long time
on the network, or on the terminal, waiting for the user, which would
stall every other thread that needs the kernel dso. It is dropped
just around the fetch, nothing of the dso is touched while it is, and
the symbols are loaded with the lock held again, the same way
dso__debuginfo() already does it for the debuginfo of a DSO. If
another thread got the symbols for the dso in the meantime, the file
that came back is not loaded a second time.
PATCH 5/8, "perf report: Add --progress option" [sashiko-bot review of
PATCH v2 5/8]:
- A phase that starts when the progress stack is full is now just not
shown, and its finish() is swallowed, instead of force completing the
phase that encloses it to make room: that one still gets its own
finish() later, which would then complete the phase above it, and so
on, one premature completion per nesting level. Reproduced by
building with STDIO_PROGRESS__MAX_DEPTH set to 1: v2 prints
"Processing events... [100.0%]" while it is at 99.7%, before the
nested phase even runs, v3 warns and keeps the outer phase to
complete when it really does. With the 8 slots in the tree this
cannot be hit, the phases perf has nest at most 2 deep, but the
bookkeeping was wrong.
PATCH 6/8, "perf scripts: Add perf-stuck, to tell where a running perf
is stuck" [sashiko-bot review of PATCH v2 6/8]:
- The perf-dso GDB macro read dloc->ms.map->dso->name and
dloc->ms.sym->name, but struct data_loc_info.ms is a
'struct map_symbol *', so GDB's C evaluator needs -> there: as it
was, 'perf-stuck.sh -g' would have failed to evaluate it.
- The one pre-existing issue in PATCH 7/8, the unbounded member type
chase in die_get_member_type() and the unbounded recursion in
die_find_member(), is not addressed here, it went to the perf TODO
list as item 180, as it is not on the path this series fixes.
Other changes since v2:
- PATCH 2/8: a refetch made because the cached file was cleaned from
under us now updates the entry the same way the first fetch does,
and an entry with no path and no error is fetched again instead of
answered; as it was, the entry kept the first fetch's success with
no path, so the later lookups it existed for returned success with
a NULL path, which debuginfo__new_build_id() cannot open and the
vmlinux fallback in PATCH 3 would take as a file that came back
and dereference.
- PATCH 3/8: the commit message said the ignore_vmlinux_buildid that
'perf record' and 'perf probe' set internally is what keeps them
away from the fallback, but 'perf probe' only sets it for the
commands other than --list, --del and --add when given an offline
vmlinux; the message now says that and that a 'perf probe' that
does not set it, e.g. --funcs or --add without --vmlinux on a
system with a restricted /proc/kallsyms, can have the kernel
debuginfo fetched, like the other tools.
- PATCH 7/8: the kerneldoc of die_get_type_die() said the type would
be looked up in the main file and then in the alt file, using the
one that has a DIE with the saved tag, but the function
deliberately resolves in only the file the offset was recorded as
belonging to, with no fallback, and @from_alt was missing from the
parameter list; it now describes that, and the die_collect_vars()
and die_collect_global_vars() kerneldocs, which still said the
type could be retrieved with dwarf_offdie() and the offset alone,
now point at die_get_type_die().
- PATCH 7/8: the comment on the member->truncated flag said the
browser reads it, but the browser picks between ';' and '{' from
the children list, so a member cut by the nesting limit is drawn
like a complete leaf; the comment now says the flag is for the
JSON exporter added in a later series.
- PATCH 2/8: a search aborted by the user, with 's' or a signal, is
now remembered for the rest of the session: the threads already
waiting for that fetch were told there was nothing to share, but a
request that arrived later started a new download of the same file,
which the user may have skipped for being too big. debuginfod__misses
now records such a cancellation as a cancellation, the 's' message
says the build ID will not be fetched again in this session and the
debug message tells a cancellation and a server miss apart.
- Patches 1, 4 and 8 are unchanged from v2: sashiko-bot found no
issues in them.
What changed from v1 (060dccedab1617dc):
PATCH 2/8, "perf debuginfo: Fetch debuginfo keyed by build ID using
debuginfod" [sashiko-bot review of PATCH 2/8]:
- Included <limits.h> for PATH_MAX, that was coming in through a
glibc transitive include and fails to build with musl libc.
- debuginfod__setup_urls_env() now goes through pthread_once, so that
the setenv() of DEBUGINFOD_URLS happens exactly once: setenv() is
not thread safe and this is on the fetch path, which
dso__debuginfo() deliberately takes outside dso__lock.
- The fetch now runs with debuginfod__fetch_lock held, so that the
process global state it uses — the terminal settings, the SIGINT/
SIGTERM dispositions and the progress and cancellation state — is
only ever touched by one fetch at a time. A second concurrent
fetch would otherwise have taken the first one's raw mode as the
state to restore, leaving the terminal broken when it was done, and
would reset the cancellation state from under the fetch already in
progress. Serializing also keeps two fetches from racing for the
same keypresses and the same progress line, and a second fetch has
nothing to gain from running in parallel with a first one reading
the same kind of file off the same servers.
- An interrupt that lands after the fetch already succeeded is no
longer swallowed: at that point the terminal and the signal
dispositions are the original ones again, so the signal is re-raised,
the same way the failure path does, instead of perf carrying on as
if the user had not asked it to stop.
PATCH 6/8, "perf scripts: Add perf-stuck, to tell where a running perf
is stuck" [sashiko-bot review of PATCH 6/8]:
- /proc/<pid>/stat is now parsed with the parenthesized command name
removed first, so that a process with spaces in its name no longer
shifts every field after it: watching a process named 'my sleep'
used to report "sleep)" as its state and a garbage RSS.
- The CPU time is printed with awk's printf "%d" instead of relying on
its default output format, that switches to scientific notation,
which the bash arithmetic below cannot parse, once the sum of utime
and stime goes past six digits, i.e. some 16 minutes of CPU at
100 Hz.
- The process can go away between the '-d /proc/<pid>' check and the
read of /proc/<pid>/stat, so the read is now checked: a failed or
empty one reports "process gone" instead of 'set -u' aborting the
script on an unbound field, which is what would have hidden the
very exit the script is meant to report.
- The command line printed when the script starts has its control
characters stripped, so that a process started with escape sequences
in its arguments, e.g. one replaying a log line, does not get them
replayed on the terminal of whoever runs this.
- The two pre-existing issues sashiko-bot raised that are outside the
scope of this series, the child entries and their hists arrays
leaked by the data type browser teardown and the missing <stdio.h>
and <stdlib.h> includes in ui/browsers/annotate-data.c, are
recorded as items 178 and 179 of the perf TODO list for follow-up
work.
- Patches 1, 3, 4, 5, 7 and 8 are unchanged from v1: sashiko-bot
found no issues in them.
Arnaldo Carvalho de Melo (8):
perf test: Skip data_type_profiling when the PMU cannot record memory
events
perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod
perf symbol: Fall back to fetching the vmlinux by build ID
perf annotate-data: Show the sample count in the data-type browser
perf report: Add --progress option
perf scripts: Add perf-stuck, to tell where a running perf is stuck
perf annotate-data: Resolve type DIEs in the debug file they came from
perf mem record: Request PERF_SAMPLE_CPU by default
tools/build/feature/test-libdw.c | 15 +-
tools/perf/Documentation/perf-annotate.txt | 10 +
tools/perf/Documentation/perf-config.txt | 19 +
tools/perf/Documentation/perf-mem.txt | 4 +
tools/perf/Documentation/perf-report.txt | 28 ++
tools/perf/Documentation/perf-top.txt | 14 +
tools/perf/Makefile.config | 2 +-
tools/perf/builtin-annotate.c | 2 +
tools/perf/builtin-mem.c | 9 +
tools/perf/builtin-report.c | 17 +
tools/perf/builtin-top.c | 6 +
tools/perf/scripts/perf-stuck.gdb | 104 +++++
tools/perf/scripts/perf-stuck.sh | 194 ++++++++
tools/perf/tests/shell/data_type_profiling.sh | 46 +-
tools/perf/ui/Build | 1 +
tools/perf/ui/browsers/annotate-data.c | 2 +-
tools/perf/ui/progress.h | 2 +
tools/perf/ui/stdio/progress.c | 185 ++++++++
tools/perf/util/annotate-data.c | 59 ++-
tools/perf/util/annotate-data.h | 3 +
tools/perf/util/config.c | 3 +
tools/perf/util/debuginfo.c | 642 ++++++++++++++++++++++++++
tools/perf/util/debuginfo.h | 36 ++
tools/perf/util/dso.c | 19 +
tools/perf/util/dwarf-aux.c | 171 ++++++-
tools/perf/util/dwarf-aux.h | 33 ++
tools/perf/util/ordered-events.c | 17 +-
tools/perf/util/session.c | 12 +-
tools/perf/util/symbol.c | 70 ++-
tools/perf/util/symbol_conf.h | 1 +
30 files changed, 1710 insertions(+), 55 deletions(-)
base-commit: aa18964dd64511305de0711fed912054da6f5d18
v2-head: 00cffc7df5abfa6850beb360deb830d8495163ce
v1-head: 060dccedab1617dcbf6e2be64ed97afe3f3fac63
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
2026-09-14 1:31 ` Namhyung Kim
2026-09-13 22:28 ` [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod Arnaldo Carvalho de Melo
` (6 subsequent siblings)
7 siblings, 1 reply; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
The test records with 'perf mem record' in per-thread mode and its only
guard matches one specific message:
perf mem record -o /dev/null -- true 2>&1 | \
grep -q "failed: no PMU supports the memory events" && exit 2
A PMU that has memory events but refuses them per-thread falls through
it, and AMD IBS does:
$ perf mem record -o /dev/null -- true
Error:
Failure to open event 'ibs_op/ldlat=0/u' on PMU 'ibs_op' which will be removed.
Invalid event (ibs_op/ldlat=0/u) in per-thread mode, enable system wide with '-a'.
Error:
Failure to open any events for recording.
What follows is not a skip either. The script runs under 'set -e', so
the bare 'perf mem record' in test_basic_annotate() aborts it through the
EXIT trap, which reports a signal that never happened and exits 1:
Basic Rust perf annotate test
Unexpected signal in test_basic_annotate
so 'perf test' turns "this PMU cannot record these events" into a test
failure:
$ perf test "data type profiling"
87: perf data type profiling tests : FAILED!
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/tests/shell/data_type_profiling.sh | 46 +++++++++++++++----
1 file changed, 36 insertions(+), 10 deletions(-)
diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
index eca694600a0478d2..a916c410274aa888 100755
--- a/tools/perf/tests/shell/data_type_profiling.sh
+++ b/tools/perf/tests/shell/data_type_profiling.sh
@@ -19,6 +19,15 @@ perfout=$(mktemp /tmp/__perf_test.perf.out.XXXXX)
perf mem record -o /dev/null -- true 2>&1 | \
grep -q "failed: no PMU supports the memory events" && exit 2
+# Skip if per-thread mem record is not supported on this PMU (e.g. AMD IBS
+# needs system-wide '-a'): it is what the test records with below, and a
+# failing record must not be reported as a test failure.
+if ! perf mem record -o /dev/null -- true 2>/dev/null
+then
+ echo "Skip: cannot record memory events on this PMU"
+ exit 2
+fi
+
cleanup() {
rm -rf "${perfdata}" "${perfout}"
rm -rf "${perfdata}".old
@@ -52,25 +61,42 @@ test_basic_annotate() {
index=1 ;;
esac
+ # Under 'set -e' a bare failing command aborts the script through the EXIT
+ # trap, so the commands that report a failure have to be the condition of
+ # an 'if' for that reporting to ever happen.
if [ "x${mode}" == "xBasic" ]
then
- perf mem record -o "${perfdata}" ${testprogs[$index]} 2> /dev/null
+ if ! perf mem record -o "${perfdata}" ${testprogs[$index]} 2> /dev/null
+ then
+ echo "${mode} annotate [Failed: perf record]"
+ err=1
+ return
+ fi
else
- perf mem record -o - ${testprogs[$index]} 2> /dev/null > "${perfdata}"
- fi
- if [ "x$?" != "x0" ]
- then
- echo "${mode} annotate [Failed: perf record]"
- err=1
- return
+ if ! perf mem record -o - ${testprogs[$index]} 2> /dev/null > "${perfdata}"
+ then
+ echo "${mode} annotate [Failed: perf record]"
+ err=1
+ return
+ fi
fi
# Generate the annotated output file
if [ "x${mode}" == "xBasic" ]
then
- perf annotate --code-with-type -i "${perfdata}" --stdio --percent-limit 1 2> /dev/null > "${perfout}"
+ if ! perf annotate --code-with-type -i "${perfdata}" --stdio --percent-limit 1 2> /dev/null > "${perfout}"
+ then
+ echo "${mode} annotate [Failed: perf annotate]"
+ err=1
+ return
+ fi
else
- perf annotate --code-with-type -i - --stdio 2> /dev/null --percent-limit 1 < "${perfdata}" > "${perfout}"
+ if ! perf annotate --code-with-type -i - --stdio 2> /dev/null --percent-limit 1 < "${perfdata}" > "${perfout}"
+ then
+ echo "${mode} annotate [Failed: perf annotate]"
+ err=1
+ return
+ fi
fi
# check if it has the target data type
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
2026-09-14 1:34 ` Namhyung Kim
2026-09-13 22:28 ` [PATCH 3/8] perf symbol: Fall back to fetching the vmlinux by build ID Arnaldo Carvalho de Melo
` (5 subsequent siblings)
7 siblings, 1 reply; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
perf already uses debuginfod to fetch source files when annotating
(via probe-finder.c) and 'perf probe' has open_from_debuginfod(),
which queries debuginfo keyed by build ID when a module's debuginfo
isn't found locally, largely the same thing this adds; eventually
that one could be moved over to the new helper. For the analysis
tools there was no way to obtain the debuginfo for a DSO in a
profile when it isn't available locally under the name the DSO was
opened with, for instance the vmlinux for the kernel a profile was
recorded on when processing it on another machine, or after the
kernel and its debuginfo package got upgraded in between.
Add debuginfo__find_build_id(), that uses the debuginfod client to
locate a debuginfo file keyed by the build ID, checking its local
cache first and then querying the servers in DEBUGINFOD_URLS, and
debuginfo__new_build_id(), that opens the DWARF in the file it finds.
The debuginfod client fails when DEBUGINFOD_URLS isn't set even when
what it wants is in its local cache, and the distro setup scripts that
populate it from /etc/debuginfod don't reach cron jobs, systemd services
and other environments that don't source the profile scripts, so also
set it from the .urls files in /etc/debuginfod when not set.
Querying servers, possibly third party ones, sends off-box the build
IDs of the binaries being analysed and a fetch can take a while, so
this is opt-out: on by default, off with --no-debuginfod, with
core.debuginfod=false, per tool with report.debuginfod and
top.debuginfod, and, since users that set buildid.dir to /dev/null
(e.g. Linus) or otherwise turn the local build-id cache off clearly
don't want fetched files stored on the box, off too in that case.
When a fetch is in progress in a terminal, stdio, the way it prints
progress is how one gets out of it: 's' aborts the current fetch via
the debuginfod client's progress callback protocol and remembers the
build ID, so that the rest of the session doesn't ask for it again,
the user may have skipped it for being too big; 'd' additionally
disables debuginfod for the rest of the session and points at 'perf
config core.debuginfod=false' to make that permanent -- rewriting the
user's ~/.perfconfig from a keypress would silently drop its comments
-- and SIGINT/SIGTERM are intercepted while the terminal is in raw
mode, so that it is restored and the signal is re-raised when the user
interrupts a fetch.
Make dso__debuginfo() use debuginfo__new_build_id() as a fallback,
keyed by the build ID recorded in the perf.data file, so that
consumers such as the data type profiler can resolve the types of
DSOs whose debuginfo can be fetched this way. Do the fetch outside
dso__lock and remember the build IDs that were a miss and the ones
whose search the user cancelled, so that consumers revisiting a set
of DSOs repeatedly, such as the data type profiler on every hist
entry DSO switch, don't pay server round trips per attempt and a
cancelled download, maybe a file the user found too big, isn't
restarted by the next request for the same build ID in the same
session.
debuginfo__new_build_id() needs libdw to open the DWARF, so it lives
in debuginfo.o, built only with CONFIG_LIBDW, while the libdebuginfod
feature check is independent of NO_LIBDW; keep the new build ID
prototypes under HAVE_LIBDW_SUPPORT too, with stubs otherwise, so
that make NO_LIBDW=1 on a system that has the debuginfod client
keeps linking.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/Documentation/perf-annotate.txt | 10 +
tools/perf/Documentation/perf-config.txt | 19 +
tools/perf/Documentation/perf-report.txt | 15 +
tools/perf/Documentation/perf-top.txt | 14 +
tools/perf/builtin-annotate.c | 2 +
tools/perf/builtin-report.c | 6 +
tools/perf/builtin-top.c | 6 +
tools/perf/util/config.c | 3 +
tools/perf/util/debuginfo.c | 681 +++++++++++++++++++++
tools/perf/util/debuginfo.h | 36 ++
tools/perf/util/dso.c | 19 +
tools/perf/util/symbol.c | 2 +
tools/perf/util/symbol_conf.h | 1 +
13 files changed, 814 insertions(+)
diff --git a/tools/perf/Documentation/perf-annotate.txt b/tools/perf/Documentation/perf-annotate.txt
index 1a90b09a12d5abb1..25af1d166dc4c877 100644
--- a/tools/perf/Documentation/perf-annotate.txt
+++ b/tools/perf/Documentation/perf-annotate.txt
@@ -58,6 +58,16 @@ OPTIONS
--ignore-vmlinux::
Ignore vmlinux files.
+--debuginfod::
+--no-debuginfod::
+ Fetch debuginfo keyed by build ID from the debuginfod servers
+ configured in DEBUGINFOD_URLS, checking the local debuginfod
+ client cache first, when it is not available locally, on for
+ these commands by default. See the --debuginfod option of
+ 'perf report' for how to turn it off, including the 's' and
+ 'd' keys that skip a fetch in progress while 'perf' is
+ waiting for it.
+
--itrace::
Options for decoding instruction tracing data. The options are:
diff --git a/tools/perf/Documentation/perf-config.txt b/tools/perf/Documentation/perf-config.txt
index 9b223f8928299945..688306abe847a0df 100644
--- a/tools/perf/Documentation/perf-config.txt
+++ b/tools/perf/Documentation/perf-config.txt
@@ -216,6 +216,14 @@ core.*::
addr2line-timeout::
Sets a timeout (in milliseconds) for parsing 'addr2line'
output. The default timeout is 5s.
+ debuginfod::
+ When set to 'false', disable fetching debuginfo keyed by
+ build ID from the debuginfod servers configured in
+ DEBUGINFOD_URLS. It is on by default, can be overridden per
+ tool with the 'report.debuginfod' and 'top.debuginfod'
+ options and per invocation with --no-debuginfod; it is off
+ too when the local build-id cache is disabled, e.g.
+ 'buildid.dir' set to /dev/null.
tui.*, gtk.*::
Subcommands that can be configured here are 'top', 'report' and 'annotate'.
@@ -562,6 +570,14 @@ report.*::
This option can change default stat behavior with empty results.
If it's set true, 'perf report --stat' will not show 0 stats.
+ report.debuginfod::
+ Fetch debuginfo keyed by build ID from the debuginfod
+ servers configured in DEBUGINFOD_URLS, checking the local
+ debuginfod client cache first, when it is not available
+ locally. On by default, set to 'false' to disable it for
+ 'perf report', globally with 'core.debuginfod=false' or per
+ invocation with --no-debuginfod.
+
top.*::
top.children::
Same as 'report.children'. So if it is enabled, the output of 'top'
@@ -569,6 +585,9 @@ top.*::
column by default.
The default is 'true'.
+ top.debuginfod::
+ Same as 'report.debuginfod', for 'perf top'.
+
top.call-graph::
This is identical to 'call-graph.record-mode', except it is
applicable only for 'top' subcommand. This option ONLY setup
diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index ae68ca402d0b4f0e..fed6af128ff07e4c 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -366,6 +366,21 @@ OPTIONS
--ignore-vmlinux::
Ignore vmlinux files.
+--debuginfod::
+--no-debuginfod::
+ Fetch debuginfo keyed by build ID from the debuginfod servers
+ configured in DEBUGINFOD_URLS, checking the local debuginfod
+ client cache first, when it is not available locally, on for
+ these commands by default. It can be turned off per invocation
+ with --no-debuginfod, per tool with the "report.debuginfod"
+ config option or globally with "core.debuginfod" set to false.
+ While a fetch is in progress in the stdio interface, 's' skips
+ the current fetch and 'd' skips it and disables debuginfod for
+ the rest of the session, pointing at 'perf config' to make that
+ permanent. It is disabled as well when the local build-id cache
+ is turned off, e.g. "buildid.dir" set to /dev/null, as that
+ asks for fetched files not to be kept on the box.
+
--kallsyms=<file>::
kallsyms pathname
diff --git a/tools/perf/Documentation/perf-top.txt b/tools/perf/Documentation/perf-top.txt
index 2da2a16bbf260685..8bcb7b0ac4a2407a 100644
--- a/tools/perf/Documentation/perf-top.txt
+++ b/tools/perf/Documentation/perf-top.txt
@@ -83,6 +83,20 @@ Default is to monitor all CPUS.
--ignore-vmlinux::
Ignore vmlinux files.
+--debuginfod::
+--no-debuginfod::
+ Fetch debuginfo keyed by build ID from the debuginfod servers
+ configured in DEBUGINFOD_URLS, checking the local debuginfod
+ client cache first, when it is not available locally, on for
+ these commands by default. Turn it off per invocation with
+ --no-debuginfod, with the "top.debuginfod" config option or
+ globally with "core.debuginfod" set to false. While a fetch is
+ in progress in the stdio interface, 's' skips the current fetch
+ and 'd' skips it and disables debuginfod for the rest of the
+ session, pointing at 'perf config' to make that permanent.
+ Disabled as well when the build-id cache is off, e.g.
+ "buildid.dir" set to /dev/null.
+
--kallsyms=<file>::
kallsyms pathname
diff --git a/tools/perf/builtin-annotate.c b/tools/perf/builtin-annotate.c
index 4638e6fdc39bb6b7..d14ae7d1c345cb79 100644
--- a/tools/perf/builtin-annotate.c
+++ b/tools/perf/builtin-annotate.c
@@ -733,6 +733,8 @@ int cmd_annotate(int argc, const char **argv)
OPT_BOOLEAN(0, "stdio2", &annotate.use_stdio2, "Use the stdio interface"),
OPT_BOOLEAN(0, "ignore-vmlinux", &symbol_conf.ignore_vmlinux,
"don't load vmlinux even if found"),
+ OPT_BOOLEAN(0, "debuginfod", &symbol_conf.debuginfod,
+ "fetch debuginfo keyed by build ID from the debuginfod servers, on by default, use --no-debuginfod to turn off"),
OPT_STRING('k', "vmlinux", &symbol_conf.vmlinux_name,
"file", "vmlinux pathname"),
OPT_BOOLEAN('m', "modules", &symbol_conf.use_modules,
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index b280ff9ff45e6485..4d3383d1daae2ed9 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -134,6 +134,10 @@ static int report__config(const char *var, const char *value, void *cb)
symbol_conf.event_group = perf_config_bool(var, value);
return 0;
}
+ if (!strcmp(var, "report.debuginfod")) {
+ symbol_conf.debuginfod = perf_config_bool(var, value);
+ return 0;
+ }
if (!strcmp(var, "report.percent-limit")) {
double pcnt = strtof(value, NULL);
@@ -1346,6 +1350,8 @@ int cmd_report(int argc, const char **argv)
"file", "vmlinux pathname"),
OPT_BOOLEAN(0, "ignore-vmlinux", &symbol_conf.ignore_vmlinux,
"don't load vmlinux even if found"),
+ OPT_BOOLEAN(0, "debuginfod", &symbol_conf.debuginfod,
+ "fetch debuginfo keyed by build ID from the debuginfod servers, on by default, use --no-debuginfod to turn off"),
OPT_STRING(0, "kallsyms", &symbol_conf.kallsyms_name,
"file", "kallsyms pathname"),
OPT_BOOLEAN('f', "force", &symbol_conf.force, "don't complain, do it"),
diff --git a/tools/perf/builtin-top.c b/tools/perf/builtin-top.c
index c2562d49be46a9a1..aed45167d95005dd 100644
--- a/tools/perf/builtin-top.c
+++ b/tools/perf/builtin-top.c
@@ -1436,6 +1436,10 @@ static int perf_top_config(const char *var, const char *value, void *cb __maybe_
symbol_conf.cumulate_callchain = perf_config_bool(var, value);
return 0;
}
+ if (!strcmp(var, "top.debuginfod")) {
+ symbol_conf.debuginfod = perf_config_bool(var, value);
+ return 0;
+ }
return 0;
}
@@ -1508,6 +1512,8 @@ int cmd_top(int argc, const char **argv)
"file", "vmlinux pathname"),
OPT_BOOLEAN(0, "ignore-vmlinux", &symbol_conf.ignore_vmlinux,
"don't load vmlinux even if found"),
+ OPT_BOOLEAN(0, "debuginfod", &symbol_conf.debuginfod,
+ "fetch debuginfo keyed by build ID from the debuginfod servers, on by default, use --no-debuginfod to turn off"),
OPT_STRING(0, "kallsyms", &symbol_conf.kallsyms_name,
"file", "kallsyms pathname"),
OPT_BOOLEAN('K', "hide_kernel_symbols", &top.hide_kernel_symbols,
diff --git a/tools/perf/util/config.c b/tools/perf/util/config.c
index b2972c35c1eca68c..31c6618d3b3daf22 100644
--- a/tools/perf/util/config.c
+++ b/tools/perf/util/config.c
@@ -470,6 +470,9 @@ static int perf_default_core_config(const char *var, const char *value)
if (!strcmp(var, "core.addr2line-disable-warn"))
symbol_conf.addr2line_disable_warn = perf_config_bool(var, value);
+ if (!strcmp(var, "core.debuginfod"))
+ symbol_conf.debuginfod = perf_config_bool(var, value);
+
/* Add other config variables here. */
return 0;
}
diff --git a/tools/perf/util/debuginfo.c b/tools/perf/util/debuginfo.c
index 84a78b30ceac1066..21cdd3ec8e139f75 100644
--- a/tools/perf/util/debuginfo.c
+++ b/tools/perf/util/debuginfo.c
@@ -7,17 +7,27 @@
#include <errno.h>
#include <fcntl.h>
+#include <dirent.h>
+#include <limits.h>
+#include <pthread.h>
+#include <signal.h>
+#include <stdbool.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
+#include <termios.h>
#include <unistd.h>
+#include <linux/list.h>
#include <linux/zalloc.h>
+#include <api/fs/fs.h>
#include "build-id.h"
#include "dso.h"
#include "debug.h"
#include "debuginfo.h"
+#include "mutex.h"
#include "symbol.h"
+#include "term.h"
#ifdef HAVE_DEBUGINFOD_SUPPORT
#include <elfutils/debuginfod.h>
@@ -139,6 +149,677 @@ struct debuginfo *debuginfo__new(const char *path)
return __debuginfo__new(buf);
}
+#ifdef HAVE_DEBUGINFOD_SUPPORT
+/*
+ * Set with the use_browser variable in ui/ui.h, not included here to
+ * avoid pulling in the UI headers: when the TUI is in use, printing to
+ * stderr would garble its display.
+ */
+extern int use_browser;
+
+static bool debuginfod_progress_started;
+static bool debuginfod_fetch_cancelled;
+
+/*
+ * A fetch can be interrupted with Ctrl-C/SIGTERM while stdin is in raw
+ * mode: the handler only records the signal, the progress callback
+ * aborts the query, and debuginfo__find_build_id() restores the
+ * terminal and raises the signal again, so that the terminal is never
+ * left in raw mode when perf dies mid-fetch.
+ */
+static volatile sig_atomic_t debuginfod_signal;
+
+static void debuginfod_signal_handler(int sig)
+{
+ debuginfod_signal = sig;
+}
+
+/*
+ * 's': skip this fetch, and remember the build ID so that the rest of
+ * the session doesn't ask for it again, the query is aborted by
+ * returning a non-zero value from the progress callback, as the
+ * debuginfod client docs prescribe. 'd': also disable debuginfod for
+ * the rest of the session, telling how to make that permanent:
+ * rewriting the user's ~/.perfconfig from here would drop its comments,
+ * so point at 'perf config' instead.
+ */
+static void debuginfod__poll_cancel_keys(void)
+{
+ char ch;
+
+ while (read(STDIN_FILENO, &ch, 1) == 1) {
+ if (ch == 's' || ch == 'S') {
+ debuginfod_fetch_cancelled = true;
+ fputs("\nSkipping this debuginfod fetch, this build ID will not be fetched again in this session, press 'd' to also disable it for the other ones\n", stderr);
+ } else if (ch == 'd' || ch == 'D') {
+ debuginfod_fetch_cancelled = true;
+ symbol_conf.debuginfod = false;
+ fputs("\nSkipping this debuginfod fetch and disabling debuginfod for this session, run 'perf config core.debuginfod=false' to also disable it permanently\n", stderr);
+ }
+ }
+}
+
+/*
+ * Print a warning and a progress indicator when the debuginfod client
+ * ends up fetching a file, which can be big, such as the vmlinux for a
+ * kernel profiled on another machine or before it got upgraded, so that
+ * users know perf is not stuck, and let them bail out: 's' skips this
+ * fetch and remembers the build ID, so that the rest of the session
+ * doesn't ask for it again, 'd' also disables debuginfod for the rest of
+ * the session. The client only invokes this once it committed to a
+ * server, so 'a' is the number of bytes fetched so far, 'b' the total
+ * size when the server tells it, -1 otherwise.
+ */
+static int debuginfod_progress_fn(debuginfod_client *c __maybe_unused,
+ long a, long b)
+{
+ if (!isatty(STDERR_FILENO) || use_browser)
+ return 0;
+
+ if (debuginfod_signal)
+ return 1;
+
+ if (isatty(STDIN_FILENO)) {
+ debuginfod__poll_cancel_keys();
+ if (debuginfod_fetch_cancelled)
+ return 1;
+ }
+
+ if (!debuginfod_progress_started) {
+ fprintf(stderr, "Fetching debuginfo by build ID from the debuginfod servers, this may take a while for large files such as the vmlinux, press 's' to skip, 'd' to skip and disable\n");
+ debuginfod_progress_started = true;
+ }
+
+ if (a >= 0) {
+ if (b > 0)
+ fprintf(stderr, " %ld/%ld MiB fetched\r", a >> 20, b >> 20);
+ else
+ fprintf(stderr, " %ld MiB fetched\r", a >> 20);
+ }
+
+ return 0;
+}
+
+/*
+ * The debuginfod client checks its local cache only as part of the
+ * server query flow, so with no servers configured it fails even when
+ * the artifact is in the client cache. Distro setup scripts, e.g.
+ * /etc/profile.d/99-debuginfod.sh, export DEBUGINFOD_URLS from the
+ * .urls files in /etc/debuginfod, but that doesn't reach environments
+ * that don't source the profile scripts, such as cron jobs, systemd
+ * services and CI, so do it here when the variable isn't set. An
+ * explicitly empty DEBUGINFOD_URLS is an opt-out, matching the
+ * perf_debuginfod_setup() handling, and is left alone.
+ *
+ * setenv() is not thread safe and this is on the fetch path, that
+ * dso__debuginfo() takes outside dso__lock, so do it just once, from
+ * whichever fetch gets here first: the value is the same for all of them.
+ */
+static void debuginfod__urls_env_setup(void)
+{
+ char *urls = NULL;
+ DIR *dir;
+ struct dirent *dent;
+
+ if (getenv("DEBUGINFOD_URLS") != NULL)
+ return;
+
+ dir = opendir("/etc/debuginfod");
+ if (dir == NULL)
+ return;
+
+ while ((dent = readdir(dir)) != NULL) {
+ char *content = NULL;
+ char *new_urls;
+ char path[PATH_MAX];
+ size_t len = strlen(dent->d_name), i, size;
+ int n;
+
+ if (len < 5 || strcmp(dent->d_name + len - 5, ".urls"))
+ continue;
+
+ snprintf(path, sizeof(path), "/etc/debuginfod/%s", dent->d_name);
+ if (filename__read_str(path, &content, &size) < 0)
+ continue;
+
+ for (i = 0; i < size; i++)
+ if (content[i] == '\n' || content[i] == '\r')
+ content[i] = ' ';
+
+ if (urls == NULL) {
+ urls = strdup(content);
+ } else {
+ n = asprintf(&new_urls, "%s %s", urls, content);
+ if (n < 0) {
+ free(content);
+ continue;
+ }
+ free(urls);
+ urls = new_urls;
+ }
+ free(content);
+ }
+ closedir(dir);
+
+ if (urls != NULL) {
+ setenv("DEBUGINFOD_URLS", urls, 1);
+ pr_debug("Set DEBUGINFOD_URLS from /etc/debuginfod: %s\n", urls);
+ }
+ free(urls);
+}
+
+static void debuginfod__setup_urls_env(void)
+{
+ static pthread_once_t once = PTHREAD_ONCE_INIT;
+
+ pthread_once(&once, debuginfod__urls_env_setup);
+}
+
+/*
+ * Users can disable the local build-id/.debug cache by setting
+ * buildid.dir to /dev/null, meaning they don't want fetched
+ * binaries/debuginfo stored on the box; the debuginfod client keeps
+ * its own cache in ~/.cache/debuginfod_client, so honour that intent
+ * and don't fetch at all in that case.
+ */
+static bool debuginfod__cache_disabled(void)
+{
+ return !strcmp(buildid_dir, "/dev/null");
+}
+
+/*
+ * Build IDs that shouldn't be searched for again in this session: the
+ * ones already searched for on the debuginfod servers without success,
+ * so that callers that see the same DSO over and over, such as the data
+ * type profiler switching between DSOs on every hist entry, don't pay a
+ * server round trip again for each miss, and the ones whose search the
+ * user cancelled, maybe because what was being downloaded is too big,
+ * so that the next request for the same build ID doesn't restart a
+ * download that was refused. The cache of successes is the debuginfod
+ * client's own, in the local filesystem.
+ */
+struct debuginfod_miss {
+ struct list_head node;
+ struct build_id bid;
+ bool cancelled;
+};
+
+static LIST_HEAD(debuginfod__misses);
+static struct mutex debuginfod__missed_lock;
+
+static void debuginfod__missed_lock_setup(void)
+{
+ mutex_init(&debuginfod__missed_lock);
+}
+
+static void debuginfod__missed_lock_init(void)
+{
+ static pthread_once_t once = PTHREAD_ONCE_INIT;
+
+ pthread_once(&once, debuginfod__missed_lock_setup);
+}
+
+/*
+ * Was the search for this build ID already settled, by the servers
+ * having nothing or by the user cancelling it? When it was, @cancelled
+ * tells the two apart, so that the debug message can say which one it
+ * was.
+ */
+static bool debuginfod__missed(const struct build_id *bid, bool *cancelled)
+{
+ struct debuginfod_miss *miss;
+ bool found = false;
+
+ *cancelled = false;
+
+ debuginfod__missed_lock_init();
+ mutex_lock(&debuginfod__missed_lock);
+ list_for_each_entry(miss, &debuginfod__misses, node) {
+ if (miss->bid.size == bid->size &&
+ !memcmp(miss->bid.data, bid->data, bid->size)) {
+ found = true;
+ *cancelled = miss->cancelled;
+ break;
+ }
+ }
+ mutex_unlock(&debuginfod__missed_lock);
+
+ return found;
+}
+
+static void debuginfod__miss_add(const struct build_id *bid, bool cancelled)
+{
+ struct debuginfod_miss *miss = zalloc(sizeof(*miss));
+
+ if (miss == NULL)
+ return;
+
+ miss->bid = *bid;
+ miss->cancelled = cancelled;
+
+ debuginfod__missed_lock_init();
+ mutex_lock(&debuginfod__missed_lock);
+ list_add(&miss->node, &debuginfod__misses);
+ mutex_unlock(&debuginfod__missed_lock);
+}
+
+/*
+ * One fetch at a time.
+ *
+ * The terminal settings, the signal dispositions and the progress and
+ * cancellation state below are process global, so two concurrent fetches,
+ * which dso__debuginfo() makes possible by taking the fetch out of
+ * dso__lock, would fight over them: the second one would take the first
+ * one's raw mode as the state to restore and leave the terminal broken when
+ * it is done, and resetting the cancellation state would drop the 's'/'d'
+ * keypress that was meant for the fetch already in progress. Serializing
+ * also keeps the two from racing for the same keypresses and for the same
+ * progress line, and a second fetch has nothing to gain from running in
+ * parallel with a first one reading the same kind of file off the same
+ * servers.
+ *
+ * What that costs is that a fetch for one build ID blocks a fetch for
+ * another one, and it is what parallel downloads would fix: the terminal
+ * in raw mode, the signal dispositions, the progress line and the 's'/'d'
+ * keys would have to become per fetch and refcounted, so that N fetches
+ * share one terminal session and one signal handler, with the first one in
+ * setting them up and the last one out putting them back. Worth doing
+ * only if the wait turns out to be long, because it mostly is not: the
+ * lookups below answer the second and later requests for a build ID from
+ * memory, so after the first pass over the build IDs of a workload, which
+ * is the only time anything is fetched at all, the serialization has
+ * nothing left to serialize. Start there if a profile with many DSOs to
+ * fetch shows up in a profile of perf itself.
+ */
+static struct mutex debuginfod__fetch_lock;
+
+static void debuginfod__fetch_lock_setup(void)
+{
+ mutex_init(&debuginfod__fetch_lock);
+}
+
+static void debuginfod__fetch_lock_init(void)
+{
+ static pthread_once_t once = PTHREAD_ONCE_INIT;
+
+ pthread_once(&once, debuginfod__fetch_lock_setup);
+}
+
+/*
+ * The lookups that are in progress, and the ones that brought a file back,
+ * guarded by debuginfod__fetch_lock, which is also the mutex the waiters
+ * below sleep on.
+ *
+ * A second thread that needs a build ID that is already being fetched waits
+ * for that fetch instead of starting another one: while the first one is
+ * still running, the file is not in the debuginfod client cache yet, so the
+ * second one would be a second download of the same file, with a second
+ * client, a second progress line and a second turn at putting the terminal
+ * in raw mode, for the same answer.
+ *
+ * An entry that brought a file back stays, as the answer for whoever needs
+ * the same build ID later, with no client at all: the DSOs in a profile get
+ * asked for over and over, dso__debuginfo() is called per symbol annotated,
+ * and with the path in hand the answer is a strdup(). The file stays in
+ * the debuginfod client cache, so the path stays valid, and like
+ * debuginfod__misses this grows with the number of build IDs in the
+ * workload, one small entry each, and is not trimmed.
+ *
+ * An entry that didn't bring a file back is dropped as soon as whoever was
+ * waiting for it is woken: there is nothing left to share, and the fetch
+ * already recorded it in debuginfod__misses, either as a miss or as a user
+ * cancellation, so the rest of the session doesn't ask for it again.
+ */
+struct debuginfo_lookup {
+ struct list_head node;
+ struct build_id bid;
+ struct cond done;
+ int err; /* 0: 'path' is the file, -1: not available */
+ char *path;
+ int nr_waiters;
+ bool fetching;
+};
+
+static LIST_HEAD(debuginfo_lookups);
+
+static bool build_id__equal(const struct build_id *a, const struct build_id *b)
+{
+ return a->size == b->size && memcmp(a->data, b->data, a->size) == 0;
+}
+
+static struct debuginfo_lookup *debuginfo_lookup__find(const struct build_id *bid)
+{
+ struct debuginfo_lookup *lookup;
+
+ list_for_each_entry(lookup, &debuginfo_lookups, node) {
+ if (build_id__equal(&lookup->bid, bid))
+ return lookup;
+ }
+
+ return NULL;
+}
+
+static void debuginfo_lookup__delete(struct debuginfo_lookup *lookup)
+{
+ list_del(&lookup->node);
+ cond_destroy(&lookup->done);
+ zfree(&lookup->path);
+ free(lookup);
+}
+
+/*
+ * A lookup entry in progress, added to the shared list so that a second
+ * request for the same build ID waits for this one instead of fetching it
+ * again. Called, and the result used, with debuginfod__fetch_lock held.
+ * Out of memory just means not sharing this one, the fetch is the same
+ * without the entry.
+ */
+static struct debuginfo_lookup *debuginfo_lookup__new(const struct build_id *bid)
+{
+ struct debuginfo_lookup *lookup = zalloc(sizeof(*lookup));
+
+ if (lookup != NULL) {
+ lookup->bid = *bid;
+ lookup->err = -1;
+ lookup->fetching = true;
+ cond_init(&lookup->done);
+ list_add(&lookup->node, &debuginfo_lookups);
+ }
+
+ return lookup;
+}
+
+/*
+ * The fetch itself, the terminal in raw mode and the signal dispositions
+ * swapped for the ones that restore it, so that the caller has to hold
+ * debuginfod__fetch_lock for the whole of it, see the comment there.
+ */
+static int debuginfod__fetch(const struct build_id *bid, char **path)
+{
+ char sbuild_id[SBUILD_ID_SIZE];
+ struct termios orig_termios;
+ struct sigaction sa, orig_sigint, orig_sigterm;
+ bool term_set = false, sigint_set = false, sigterm_set = false;
+ debuginfod_client *c;
+ int fd;
+
+ debuginfod__setup_urls_env();
+
+ c = debuginfod_begin();
+ if (c == NULL)
+ return -1;
+
+ debuginfod_set_progressfn(c, debuginfod_progress_fn);
+
+ debuginfod_fetch_cancelled = false;
+ debuginfod_signal = 0;
+
+ /*
+ * Make stdin deliver keypresses without waiting for a newline,
+ * the progress callback above polls it for the 's'/'d' keys,
+ * only in the stdio case with both stdin and stderr being a
+ * terminal, the TUI/pipe cases have no business being poked
+ * here. Intercept SIGINT/SIGTERM so that the terminal is
+ * restored before the process dies, the handler only records
+ * the signal and the callback aborts the query.
+ */
+ if (isatty(STDIN_FILENO) && isatty(STDERR_FILENO) && !use_browser) {
+ set_term_quiet_input(&orig_termios);
+ term_set = true;
+
+ memset(&sa, 0, sizeof(sa));
+ sa.sa_handler = debuginfod_signal_handler;
+ sigemptyset(&sa.sa_mask);
+ if (sigaction(SIGINT, &sa, &orig_sigint) == 0)
+ sigint_set = true;
+ if (sigaction(SIGTERM, &sa, &orig_sigterm) == 0)
+ sigterm_set = true;
+ }
+
+ fd = debuginfod_find_debuginfo(c, bid->data, bid->size, path);
+
+ if (term_set)
+ tcsetattr(STDIN_FILENO, TCSANOW, &orig_termios);
+ if (sigint_set)
+ sigaction(SIGINT, &orig_sigint, NULL);
+ if (sigterm_set)
+ sigaction(SIGTERM, &orig_sigterm, NULL);
+
+ debuginfod_end(c);
+ if (debuginfod_progress_started) {
+ fputc('\n', stderr);
+ debuginfod_progress_started = false;
+ }
+ if (fd < 0) {
+ build_id__snprintf(bid, sbuild_id, sizeof(sbuild_id));
+ if (debuginfod_fetch_cancelled || debuginfod_signal) {
+ pr_debug("debuginfod search for build ID %s cancelled by the user\n",
+ sbuild_id);
+ /*
+ * Remember it so that the rest of the session doesn't
+ * ask for the same file again: the user may have
+ * skipped it for being too big.
+ */
+ debuginfod__miss_add(bid, true);
+ /*
+ * The terminal is restored, die as the user asked;
+ * the original dispositions are back in place.
+ */
+ if (debuginfod_signal)
+ raise(debuginfod_signal);
+ return -1;
+ }
+ pr_debug("No debuginfo found for build ID %s in debuginfod\n",
+ sbuild_id);
+ debuginfod__miss_add(bid, false);
+ return -1;
+ }
+
+ close(fd);
+
+ /*
+ * The interrupt can land after the file is already here, in which
+ * case there is no failure to report, but the user still asked for
+ * perf to stop, and the terminal and the signal dispositions are
+ * back to what they were, so honour it here as well instead of
+ * swallowing it and going on.
+ */
+ if (debuginfod_signal) {
+ build_id__snprintf(bid, sbuild_id, sizeof(sbuild_id));
+ pr_debug("debuginfod found the debuginfo for build ID %s, but the search was interrupted, exiting\n",
+ sbuild_id);
+ raise(debuginfod_signal);
+ }
+
+ return 0;
+}
+
+/*
+ * Look the build ID up, sharing the fetch with whoever else needs it, see
+ * the comment on struct debuginfo_lookup. Called, and left, with
+ * debuginfod__fetch_lock held.
+ */
+static int debuginfo_lookup__find_build_id(const struct build_id *bid, char **path)
+{
+ struct debuginfo_lookup *lookup = debuginfo_lookup__find(bid);
+ bool waited = false;
+ int err;
+
+ if (lookup == NULL) {
+ lookup = debuginfo_lookup__new(bid);
+ if (lookup == NULL)
+ return debuginfod__fetch(bid, path);
+
+ err = debuginfod__fetch(bid, path);
+
+ /*
+ * Publish it: whoever is waiting for this build ID gets the
+ * answer this fetch settled, and, when it brought a file
+ * back, so does whoever needs the same build ID later.
+ */
+ lookup->fetching = false;
+ lookup->err = err;
+ if (err == 0)
+ lookup->path = strdup(*path);
+
+ cond_broadcast(&lookup->done);
+
+ if (lookup->path == NULL && lookup->nr_waiters == 0)
+ debuginfo_lookup__delete(lookup);
+
+ return err;
+ }
+
+ /*
+ * Somebody else got here first: wait for the fetch that is in
+ * progress instead of starting another one, which, while that one
+ * is still running, would download the same file a second time.
+ */
+ if (lookup->fetching) {
+ lookup->nr_waiters++;
+ waited = true;
+
+ while (lookup->fetching)
+ cond_wait(&lookup->done, &debuginfod__fetch_lock);
+ }
+
+ if (lookup->path != NULL) {
+ /*
+ * The file stays in the debuginfod client cache, but that
+ * cache can be cleaned from under us, so check that it is
+ * still there before handing its path out. If it isn't,
+ * forget the path and fetch it again, publishing the new
+ * answer the same way the first fetch does, so that the
+ * next lookup shares it instead of fetching it a third
+ * time.
+ */
+ if (access(lookup->path, R_OK) == 0) {
+ *path = strdup(lookup->path);
+ err = *path != NULL ? 0 : -1;
+ } else {
+ zfree(&lookup->path);
+ err = debuginfod__fetch(bid, path);
+ lookup->err = err;
+ if (err == 0)
+ lookup->path = strdup(*path);
+ }
+ } else if (lookup->err == 0) {
+ /*
+ * No path and no error: the path was fetched but could
+ * not be remembered, look for the file like a caller
+ * with no entry would, and publish the answer as above
+ * so that a success always comes with a path.
+ */
+ err = debuginfod__fetch(bid, path);
+ lookup->err = err;
+ if (err == 0)
+ lookup->path = strdup(*path);
+ } else {
+ /*
+ * Nothing came back and there is nothing to retry: the fetch
+ * was cancelled or interrupted by the user, or found nothing
+ * and said so on the misses list.
+ */
+ err = lookup->err;
+ }
+
+ if (waited)
+ lookup->nr_waiters--;
+
+ /* The last one out drops an entry there is nothing to share. */
+ if (lookup->nr_waiters == 0 && lookup->path == NULL)
+ debuginfo_lookup__delete(lookup);
+
+ return err;
+}
+
+/*
+ * Find a debuginfo file keyed by the build ID, using the debuginfod
+ * client, which checks its local cache first and then queries the
+ * servers in DEBUGINFOD_URLS. Used when the debuginfo is not available
+ * locally under the name the DSO was opened with, for instance the
+ * vmlinux for the kernel the profile was recorded on, when processing
+ * the profile on another machine or after the kernel or its debuginfo
+ * package got upgraded in between.
+ *
+ * Querying servers, possibly third party, sends the build IDs of the
+ * binaries being analysed off the box, so this is opt-out: on by
+ * default, switchable off with --no-debuginfod, with
+ * core.debuginfod=false (what the 'd' key writes), with the per-tool
+ * report.debuginfod/top.debuginfod, and it is off too when the user
+ * disabled the local build-id/.debug cache, e.g. with
+ * buildid.dir = /dev/null, as is the case for users that don't want
+ * any of this stored locally.
+ *
+ * On success the path is stored in *@path and must be freed by the
+ * caller, the file remains available in the debuginfod client cache.
+ */
+int debuginfo__find_build_id(const struct build_id *bid, char **path)
+{
+ int err = -1;
+
+ *path = NULL;
+
+ if (!build_id__is_defined(bid))
+ return -1;
+
+ /*
+ * The checks below have to be made with the lock held, as they look
+ * at the state the fetch changes: debuginfod can be turned off while
+ * a fetch is in progress, by the 'd' key in its progress line, and a
+ * build ID the fetch in progress just settled, as a miss or as a
+ * cancellation, is settled for whoever is waiting for the lock as
+ * well. Deciding here and fetching there would repeat a fetch that
+ * was already made, and put the same build ID on the misses list
+ * twice.
+ */
+ debuginfod__fetch_lock_init();
+ mutex_lock(&debuginfod__fetch_lock);
+
+ if (symbol_conf.debuginfod) {
+ bool cancelled;
+
+ if (debuginfod__cache_disabled()) {
+ pr_debug("Build-id cache disabled (buildid dir is '%s'), not using debuginfod\n",
+ buildid_dir);
+ } else if (debuginfod__missed(bid, &cancelled)) {
+ char sbuild_id[SBUILD_ID_SIZE];
+
+ build_id__snprintf(bid, sbuild_id, sizeof(sbuild_id));
+ pr_debug("Not searching build ID %s in debuginfod again, %s\n",
+ sbuild_id,
+ cancelled ? "the user cancelled the search earlier" :
+ "it was a miss earlier");
+ } else {
+ err = debuginfo_lookup__find_build_id(bid, path);
+ }
+ }
+
+ mutex_unlock(&debuginfod__fetch_lock);
+
+ return err;
+}
+
+struct debuginfo *debuginfo__new_build_id(const struct build_id *bid)
+{
+ char sbuild_id[SBUILD_ID_SIZE];
+ char *path = NULL;
+ struct debuginfo *dbg;
+
+ if (debuginfo__find_build_id(bid, &path))
+ return NULL;
+
+ dbg = __debuginfo__new(path);
+ if (dbg == NULL) {
+ build_id__snprintf(bid, sbuild_id, sizeof(sbuild_id));
+ pr_debug("Failed to open DWARF in debuginfo fetched for build ID %s: %s\n",
+ sbuild_id, path);
+ }
+ free(path);
+ return dbg;
+}
+#endif /* HAVE_DEBUGINFOD_SUPPORT */
+
void debuginfo__delete(struct debuginfo *dbg)
{
if (dbg) {
diff --git a/tools/perf/util/debuginfo.h b/tools/perf/util/debuginfo.h
index a52d69932815cd72..43b211a0dec174f5 100644
--- a/tools/perf/util/debuginfo.h
+++ b/tools/perf/util/debuginfo.h
@@ -5,6 +5,8 @@
#include <errno.h>
#include <linux/compiler.h>
+struct build_id;
+
#ifdef HAVE_LIBDW_SUPPORT
#include "dwarf-aux.h"
@@ -54,6 +56,28 @@ static inline int debuginfo__get_text_offset(struct debuginfo *dbg __maybe_unuse
#ifdef HAVE_DEBUGINFOD_SUPPORT
int get_source_from_debuginfod(const char *raw_path, const char *sbuild_id,
char **new_path);
+
+/*
+ * Finding a debuginfo file keyed by build ID uses the debuginfod client,
+ * but opening the DWARF in it needs libdw, i.e. these live in
+ * debuginfo.o, which is only built with CONFIG_LIBDW.
+ */
+#ifdef HAVE_LIBDW_SUPPORT
+int debuginfo__find_build_id(const struct build_id *bid, char **path);
+struct debuginfo *debuginfo__new_build_id(const struct build_id *bid);
+#else
+static inline int debuginfo__find_build_id(const struct build_id *bid __maybe_unused,
+ char **path __maybe_unused)
+{
+ return -ENOTSUP;
+}
+
+static inline struct debuginfo *
+debuginfo__new_build_id(const struct build_id *bid __maybe_unused)
+{
+ return NULL;
+}
+#endif /* HAVE_LIBDW_SUPPORT */
#else /* HAVE_DEBUGINFOD_SUPPORT */
static inline int get_source_from_debuginfod(const char *raw_path __maybe_unused,
const char *sbuild_id __maybe_unused,
@@ -61,6 +85,18 @@ static inline int get_source_from_debuginfod(const char *raw_path __maybe_unused
{
return -ENOTSUP;
}
+
+static inline int debuginfo__find_build_id(const struct build_id *bid __maybe_unused,
+ char **path __maybe_unused)
+{
+ return -ENOTSUP;
+}
+
+static inline struct debuginfo *
+debuginfo__new_build_id(const struct build_id *bid __maybe_unused)
+{
+ return NULL;
+}
#endif /* HAVE_DEBUGINFOD_SUPPORT */
#endif /* _PERF_DEBUGINFO_H */
diff --git a/tools/perf/util/dso.c b/tools/perf/util/dso.c
index 42bfe30a3b518e80..fe7c3b0265ba633a 100644
--- a/tools/perf/util/dso.c
+++ b/tools/perf/util/dso.c
@@ -32,6 +32,7 @@
#include "string2.h"
#include "vdso.h"
#include "annotate-data.h"
+#include "debuginfo.h"
#include "libdw.h"
static const char * const debuglink_paths[] = {
@@ -2073,5 +2074,23 @@ struct debuginfo *dso__debuginfo(struct dso *dso)
mutex_unlock(dso__lock(dso));
free(name);
+
+ /*
+ * The debuginfo for a DSO in the profile may not be installed
+ * locally, for instance the vmlinux for the kernel the profile was
+ * recorded on when processing it on another machine, or after the
+ * kernel and its debuginfo got upgraded in between. Fall back to
+ * fetching it keyed by the build ID recorded in the perf.data file,
+ * using the debuginfod client, which checks its local cache first.
+ *
+ * Do it outside dso__lock, a fetch from a remote debuginfod server
+ * can take a while and would otherwise block anything else using
+ * this dso, and honour the opt-out, the user may have asked for
+ * no debuginfod via --no-debuginfod, core.debuginfod=false or by
+ * disabling the build-id cache.
+ */
+ if (dinfo == NULL)
+ dinfo = debuginfo__new_build_id(dso__bid(dso));
+
return dinfo;
}
diff --git a/tools/perf/util/symbol.c b/tools/perf/util/symbol.c
index 3587ad243159074f..b1a2684c813c5d8d 100644
--- a/tools/perf/util/symbol.c
+++ b/tools/perf/util/symbol.c
@@ -76,6 +76,8 @@ struct symbol_conf symbol_conf = {
.inline_name = true,
.res_sample = 0,
.addr2line_timeout_ms = 5 * 1000,
+ /* Fetching debuginfo by build ID, off via --no-debuginfod, etc */
+ .debuginfod = true,
};
struct map_list_node {
diff --git a/tools/perf/util/symbol_conf.h b/tools/perf/util/symbol_conf.h
index 71f60081a85bb18d..a56b1d2d843b9ff1 100644
--- a/tools/perf/util/symbol_conf.h
+++ b/tools/perf/util/symbol_conf.h
@@ -45,6 +45,7 @@ struct symbol_conf {
force,
ignore_vmlinux,
ignore_vmlinux_buildid,
+ debuginfod,
show_kernel_path,
use_modules,
allow_aliases,
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 3/8] perf symbol: Fall back to fetching the vmlinux by build ID
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 4/8] perf annotate-data: Show the sample count in the data-type browser Arnaldo Carvalho de Melo
` (4 subsequent siblings)
7 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
A profile recorded on a kernel that is no longer installed, be it
because the machine was rebooted into a new kernel or because the
profile is being processed on another machine, can't have its kernel
symbols resolved: /proc/kallsyms matches the running kernel, not the
one in the profile, and is refused when restricted, e.g. with
kernel.perf_event_paranoid > 1, while the build-id cache may carry
just a kallsyms copy with zeroed addresses. With no kernel symbols
the kernel samples can't be annotated, so they all end up in the
'(unknown)' data type and the kernel is missing from the data type
profile JSON "dsos" entry.
As a last resort, when the kernel symbols can't be found locally and
the user didn't specify a kallsyms file, fetch the vmlinux keyed by
the kernel build ID recorded in the perf.data file using the
debuginfod client and use it for symbols, which also provides the
DWARF debuginfo needed for data type profiling.
Like the other vmlinux sources this honors --ignore-vmlinux and
--ignore-vmlinux_buildid: a fetched vmlinux is still a vmlinux, and
the fetch is keyed by the build ID, so both flags skip it. 'perf
record' sets ignore_vmlinux_buildid internally, and so does 'perf
probe' for the commands other than --list, --del and --add when given
an offline vmlinux, keeping those away from the fetch as well. A
'perf probe' that doesn't set it, e.g. --funcs or --add without
--vmlinux on a system with a restricted /proc/kallsyms, can have the
kernel debuginfo fetched, like the other tools. The fetch itself
follows the opt-out controls added in the previous patch,
--no-debuginfod, core.debuginfod=false and the build-id-cache-is-off
case, and can be skipped with the 's'/'d' keys while it runs. Build
IDs that were a miss are remembered, not queried again on every
dso__load() retry.
With a profile recorded on a system running kernel 7.1.10, later
processed after it was upgraded to 7.1.13, 'perf report -s type
' went from having all 24.91% of the kernel samples as
'(unknown)' data types to resolving struct task_struct, struct rq,
struct tty_struct, struct qspinlock, etc, with the kernel DSO, keyed
by its build ID, appearing in the data type profile JSON.
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/util/symbol.c | 68 +++++++++++++++++++++++++++++++++++++++-
1 file changed, 67 insertions(+), 1 deletion(-)
diff --git a/tools/perf/util/symbol.c b/tools/perf/util/symbol.c
index b1a2684c813c5d8d..79df41732ff25e1c 100644
--- a/tools/perf/util/symbol.c
+++ b/tools/perf/util/symbol.c
@@ -20,6 +20,7 @@
#include "cap.h"
#include "cpumap.h"
#include "debug.h"
+#include "debuginfo.h"
#include "demangle-cxx.h"
#include "demangle-java.h"
#include "demangle-ocaml.h"
@@ -2191,12 +2192,33 @@ static char *dso__find_kallsyms(struct dso *dso, struct map *map)
return strdup(path);
}
+/*
+ * Last resort when the symbols for the kernel the profile was recorded
+ * on can't be found locally: fetch the vmlinux keyed by the build ID
+ * recorded in the perf.data file using the debuginfod client, which
+ * checks its local cache first, e.g. when processing the profile on
+ * another machine or after the kernel and its debuginfo package got
+ * upgraded in between.
+ */
+/*
+ * The fetch itself, that dso__load_kernel_sym() calls with dso->lock
+ * dropped, see the comment there.
+ */
+static int dso__fetch_vmlinux_build_id(struct dso *dso, char **path)
+{
+ if (!dso__has_build_id(dso))
+ return -1;
+
+ return debuginfo__find_build_id(dso__bid(dso), path);
+}
+
static int dso__load_kernel_sym(struct dso *dso, struct map *map)
{
int err;
const char *kallsyms_filename = NULL;
char *kallsyms_allocated_filename = NULL;
char *filename = NULL;
+ bool user_kallsyms = false;
/*
* Step 1: if the user specified a kallsyms or vmlinux filename, use
@@ -2215,6 +2237,7 @@ static int dso__load_kernel_sym(struct dso *dso, struct map *map)
*/
if (symbol_conf.kallsyms_name != NULL) {
kallsyms_filename = symbol_conf.kallsyms_name;
+ user_kallsyms = true;
goto do_kallsyms;
}
@@ -2257,7 +2280,50 @@ static int dso__load_kernel_sym(struct dso *dso, struct map *map)
pr_debug("Using %s for symbols\n", kallsyms_filename);
free(kallsyms_allocated_filename);
- if (err > 0 && !dso__is_kcore(dso)) {
+ /*
+ * The kallsyms may be unavailable or restricted, e.g.
+ * /proc/kallsyms with kernel.perf_event_paranoid > 1, try to fetch
+ * the vmlinux keyed by the build ID using debuginfod as a last
+ * resort, honoring --ignore-vmlinux and --ignore-vmlinux_buildid
+ * like the other vmlinux sources above.
+ */
+ if (err <= 0 && !user_kallsyms &&
+ !symbol_conf.ignore_vmlinux &&
+ !symbol_conf.ignore_vmlinux_buildid) {
+ char *fetched_path = NULL;
+
+ /*
+ * dso__load() holds dso->lock while it calls us, and the
+ * fetch below can take a long time, blocked on the network
+ * or on the terminal, waiting for the user: do it with the
+ * lock dropped, as dso__debuginfo() does for the debuginfo
+ * of a DSO, so that the threads that need this dso don't get
+ * stuck behind a server round trip. Nothing of the dso is
+ * touched by the fetch, the symbols are loaded with the lock
+ * held again, and only if the fetch brought a file back.
+ */
+ mutex_unlock(dso__lock(dso));
+ err = dso__fetch_vmlinux_build_id(dso, &fetched_path);
+ mutex_lock(dso__lock(dso));
+
+ if (err) {
+ zfree(&fetched_path);
+ } else if (dso__loaded(dso)) {
+ /*
+ * Somebody else got the symbols for this dso while
+ * the lock was dropped for the fetch, use those
+ * instead of loading the file that came back a
+ * second time.
+ */
+ pr_debug("%s was loaded while its vmlinux was being fetched, using it\n",
+ dso__name(dso));
+ zfree(&fetched_path);
+ err = 1;
+ } else {
+ /* Takes ownership of 'fetched_path' even when it fails */
+ err = dso__load_vmlinux(dso, map, fetched_path, true);
+ }
+ } else if (err > 0 && !dso__is_kcore(dso)) {
struct maps *kmaps = map__kmaps(map);
dso__set_binary_type(dso, DSO_BINARY_TYPE__KALLSYMS);
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 4/8] perf annotate-data: Show the sample count in the data-type browser
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
` (2 preceding siblings ...)
2026-09-13 22:28 ` [PATCH 3/8] perf symbol: Fall back to fetching the vmlinux by build ID Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 5/8] perf report: Add --progress option Arnaldo Carvalho de Melo
` (3 subsequent siblings)
7 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
The data-type browser has a samples view, selected with -n (or with
annotate.show_nr_samples), in which browser__write_overhead() prints a
local nr_samples variable that is initialized to zero and never
updated, so every member is listed as having no samples while the
period and percent columns for the same entry are filled in.
Print the histogram entry's own count instead.
This predates the load/store counter split, so fix it ahead of that
patch: the split then only has to adapt a line that is already correct,
and this fix can be picked on its own.
Fixes: d001c7a7f4736743 ("perf annotate-data: Add hist_entry__annotate_data_tui()")
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/ui/browsers/annotate-data.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/tools/perf/ui/browsers/annotate-data.c b/tools/perf/ui/browsers/annotate-data.c
index aa8c89fe2e82c1c5..1080ed1a40d2609a 100644
--- a/tools/perf/ui/browsers/annotate-data.c
+++ b/tools/perf/ui/browsers/annotate-data.c
@@ -370,7 +370,7 @@ static void browser__write_overhead(struct ui_browser *uib,
u64 period = hist->period;
double percent = total->period ? (100.0 * period / total->period) : 0;
bool current = ui_browser__is_current_entry(uib, row);
- int nr_samples = 0;
+ int nr_samples = hist->nr_samples;
ui_browser__set_percent_color(uib, percent, current);
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 5/8] perf report: Add --progress option
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
` (3 preceding siblings ...)
2026-09-13 22:28 ` [PATCH 4/8] perf annotate-data: Show the sample count in the data-type browser Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 6/8] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
` (2 subsequent siblings)
7 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
Processing a large data type profiling session, e.g. an AMD IBS one, can
take a long time and, when using the stdio output, 'perf report' gives no
feedback about which phase it is in nor about how far along it is: the
TUI has a progress bar for that, but in the stdio case the ui_progress
updates, that are already there, are dropped on the floor.
Add a --progress option that installs a stdio backend for ui_progress,
ui/stdio/progress.c, printing the phase title, the percentage done and
the current/total counts to stderr:
Processing events... [ 42.3%] 317M / 746M
Merging related events... [ 61.0%] 309026 / 506686
Sorting events for output... [ 98.2%] 14132 / 14387
The first one is the perf.data file size based progress already kept
while reading events, sized with the same unit_number__scnprintf() used
by the TUI progress bar title; the other two are the hist entry based
ones for the hist entry merging (collapse) and output sorting phases,
that print raw counts.
Steps are 1% of the phase total, the default (total / 16) is tuned for
the character cell based TUI bar, and the last update normally stops
short of the total, so finishing a phase prints it as complete.
When stderr is a tty the line is updated in place, with each phase
getting a line of its own, otherwise one line is printed per update, so
that redirecting stderr to a file leaves a readable log of the phases.
Phases can be nested, e.g. the ordered events flushes that take place
while the "Processing events..." phase is still in progress, so the
backend keeps track of the ones started so far to be able to complete
the right one when a phase finishes, as ui_progress__finish() gets no
arguments.
That bookkeeping requires ui_progress__init() and ui_progress__finish()
to be paired, and a few callers had paths returning early without the
finish: do_flush() in ordered-events.c on session_done() and on
deliver() errors, and the ENOMEM paths right after the init in
__perf_session__process_pipe_events() and __perf_session__process_dir_events().
Fix those, with a backend that kept the stale entries it would print
through dangling pointers to returned stack frames, and have the
update() method ignore anything that doesn't match the innermost running
phase, printing nothing is better than printing another phase's numbers
or reading the stack array with index -1.
Committer testing:
⬢ [acme@toolbx perf-tools-next]$ perf record -o perf.data.small -- sleep 0.2
[ perf record: Woken up 2 times to write data ]
[ perf record: Captured and wrote 0.002 MB perf.data.small (7 samples) ]
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -i perf.data.small > /dev/null
Processing events... [ 0.0%] 0B / 2K
Processing events... [ 99.6%] 2040B / 2K
Processing time ordered events... [100.0%] 13 / 13
Processing events... [100.0%] 2K / 2K
Sorting events for output... [100.0%] 5 / 5
⬢ [acme@toolbx perf-tools-next]$
Nothing goes to stderr without --progress:
⬢ [acme@toolbx perf-tools-next]$ perf report --stdio -i perf.data.small > /dev/null 2> stderr.txt
⬢ [acme@toolbx perf-tools-next]$ wc -c stderr.txt
0 stderr.txt
⬢ [acme@toolbx perf-tools-next]$
With a 321 MB perf.data (perf mem record, i7-14700K):
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -i perf.data.i7 > /dev/null
Processing events... [ 0.0%] 0B / 306M
[...]
Sorting events for output... [100.0%] 29996 / 29996
⬢ [acme@toolbx perf-tools-next]$
Which is what prompted this: an AMD IBS data type profile session hangs
in the DWARF type resolution done when merging hist entries, with
--progress it now is visible that it is not the event loading that is
stuck, but the merging, at a specific hist entry:
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -s type -i perf.data.ibs > /dev/null
Processing events... [ 0.0%] 0B / 746M
[...]
Processing events... [100.0%] 746M / 746M
Merging related events... [ 58.0%] 293828 / 506686
Merging related events... [ 59.0%] 298894 / 506686
Merging related events... [ 60.0%] 303960 / 506686
Merging related events... [ 61.0%] 309026 / 506686
⬢ [acme@toolbx perf-tools-next]$ gdb -p $(pidof perf) -batch -ex 'bt 6' -ex detach
#0 0x0000000000773f31 in die_get_pointer_type ()
#1 0x000000000077b817 in find_data_type ()
#2 0x0000000000636f20 in __hist_entry__get_data_type ()
#3 0x0000000000639c65 in hist_entry__get_data_type ()
#4 0x00000000006e57a8 in sort__type_collapse ()
#5 0x00000000006ef0ee in hists__collapse_resort ()
⬢ [acme@toolbx perf-tools-next]$
The report related entries in 'perf test' pass:
⬢ [acme@toolbx perf-tools-next]$ perf test 17 27 30 31 85 88
17: Match and link multiple hists : Ok
27: Filter hist entries : Ok
30: Sort output of hist entries : Ok
31: Cumulate child hist entries : Ok
85: Test that perf report includes file offsets and event type names in diagnostic messages. : Skip
88: Test that perf report handles truncated perf.data gracefully (no crash, no segfault — clean error exit).: Ok
⬢ [acme@toolbx perf-tools-next]$
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/Documentation/perf-report.txt | 13 ++
tools/perf/builtin-report.c | 11 ++
tools/perf/ui/Build | 1 +
tools/perf/ui/progress.h | 2 +
tools/perf/ui/stdio/progress.c | 185 +++++++++++++++++++++++
tools/perf/util/ordered-events.c | 17 ++-
tools/perf/util/session.c | 12 +-
7 files changed, 234 insertions(+), 7 deletions(-)
create mode 100644 tools/perf/ui/stdio/progress.c
diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index fed6af128ff07e4c..2db4b069546130f0 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -29,6 +29,19 @@ OPTIONS
--quiet::
Do not show any warnings or messages. (Suppress -v)
+--progress::
+ Show progress for each of the processing phases, printing the
+ percentage done and the current/total counts: the first phase
+ counts the bytes of the perf.data file processed so far, the
+ merge and sort phases count hist entries, e.g.:
+
+ Processing events... [ 42.3%] 4G / 10G
+ Merging related events... [ 7.1%] 1024 / 14387
+ Sorting events for output... [ 98.2%] 14132 / 14387
+
+ It is a no-op when using the TUI or GTK browsers, that already
+ present progress information.
+
-n::
--show-nr-samples::
Show the number of samples for each symbol
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 4d3383d1daae2ed9..fe59a429b9d6eef9 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -87,6 +87,7 @@ struct report {
bool use_gtk;
#endif
bool use_stdio;
+ bool progress;
bool show_full_info;
bool show_threads;
bool inverted_callchain;
@@ -1373,6 +1374,8 @@ int cmd_report(int argc, const char **argv)
"Use the stdio interface"),
OPT_BOOLEAN(0, "weights", &symbol_conf.annotate_weight,
"Show or hide weight columns in annotation. Default show if non-zero."),
+ OPT_BOOLEAN(0, "progress", &report.progress,
+ "Show progress while processing the perf.data file"),
OPT_BOOLEAN(0, "header", &report.header, "Show data header."),
OPT_BOOLEAN(0, "header-only", &report.header_only,
"Show only data header."),
@@ -1765,6 +1768,14 @@ int cmd_report(int argc, const char **argv)
else
use_browser = 0;
+ /*
+ * The TUI/GTK browsers already show progress, this is for the stdio
+ * case, where we print the percentage of the perf.data file that was
+ * processed so far, for each of the processing phases.
+ */
+ if (report.progress && use_browser == 0)
+ stdio_progress__init();
+
if (report.data_type && use_browser == 1) {
symbol_conf.annotate_data_member = true;
symbol_conf.annotate_data_sample = true;
diff --git a/tools/perf/ui/Build b/tools/perf/ui/Build
index 6005f813c9e3990c..a7b1740d51c80f23 100644
--- a/tools/perf/ui/Build
+++ b/tools/perf/ui/Build
@@ -4,6 +4,7 @@ perf-ui-y += progress.o
perf-ui-y += util.o
perf-ui-y += hist.o
perf-ui-y += stdio/hist.o
+perf-ui-y += stdio/progress.o
CFLAGS_setup.o += -DLIBDIR="BUILD_STR($(LIBDIR))"
diff --git a/tools/perf/ui/progress.h b/tools/perf/ui/progress.h
index 4f52c37b2f099a82..03f1a8bb260ba076 100644
--- a/tools/perf/ui/progress.h
+++ b/tools/perf/ui/progress.h
@@ -23,6 +23,8 @@ void __ui_progress__init(struct ui_progress *p, u64 total,
void ui_progress__update(struct ui_progress *p, u64 adv);
+void stdio_progress__init(void);
+
struct ui_progress_ops {
void (*init)(struct ui_progress *p);
void (*update)(struct ui_progress *p);
diff --git a/tools/perf/ui/stdio/progress.c b/tools/perf/ui/stdio/progress.c
new file mode 100644
index 0000000000000000..b4a6b2732a0546cd
--- /dev/null
+++ b/tools/perf/ui/stdio/progress.c
@@ -0,0 +1,185 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Progress feedback for the stdio (non-TUI/GTK) case, enabled via
+ * 'perf report --progress': the perf.data file size based progress for
+ * the event processing phase, plus the hist entry based ones for the
+ * merging and sorting phases.
+ */
+#include <inttypes.h>
+#include <stdio.h>
+#include <unistd.h>
+#include <linux/kernel.h>
+#include "../../util/debug.h"
+#include "../../util/units.h"
+#include "../progress.h"
+
+/*
+ * Phases can be nested, e.g. the ordered events flushes that take place
+ * while the "Processing events..." phase is still in progress, so keep
+ * track of the ones started so far to be able to complete the right one
+ * when a phase finishes, as ui_progress__finish() gets no arguments. The
+ * bookkeeping of what was last printed is per phase: when the nested one
+ * finishes, the outer one must still know that its own last line was
+ * already the complete one, else it would be printed again at finish time.
+ */
+#define STDIO_PROGRESS__MAX_DEPTH 8
+
+struct stdio_progress_phase {
+ struct ui_progress *p;
+ u64 last_printed;
+ size_t last_len;
+};
+
+static struct stdio_progress_phase stdio_progress__stack[STDIO_PROGRESS__MAX_DEPTH];
+static int stdio_progress__depth;
+static bool stdio_progress__is_tty;
+/*
+ * Phases that started when there was no room left for them on the stack:
+ * they are not shown, and their finish() is still to come, see
+ * stdio_progress__finish().
+ */
+static int stdio_progress__dropped;
+
+static void stdio_progress__print_phase(struct stdio_progress_phase *phase,
+ u64 curr)
+{
+ struct ui_progress *p = phase->p;
+ char buf_cur[20], buf_tot[20], buf[128];
+ double percent = p->total ? 100.0 * (double)curr / (double)p->total : 0.0;
+ size_t len;
+
+ /*
+ * Only the completion line shows 100.0%: a 99.99% progress would
+ * round up to it in the display, making the line printed at finish
+ * time look like a duplicate.
+ */
+ if (curr < p->total && percent > 99.9)
+ percent = 99.9;
+
+ if (p->size) {
+ unit_number__scnprintf(buf_cur, sizeof(buf_cur), curr);
+ unit_number__scnprintf(buf_tot, sizeof(buf_tot), p->total);
+ len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %s / %s",
+ p->title, percent, buf_cur, buf_tot);
+ } else {
+ len = scnprintf(buf, sizeof(buf), "%s [%5.1f%%] %" PRIu64 " / %" PRIu64,
+ p->title, percent, curr, p->total);
+ }
+
+ if (!stdio_progress__is_tty) {
+ fprintf(stderr, "%s\n", buf);
+ goto out;
+ }
+
+ /* Pad to the length of the previous line to erase its leftovers. */
+ fprintf(stderr, "\r%s%*s", buf,
+ (int)(len < phase->last_len ? phase->last_len - len : 0), "");
+ phase->last_len = len;
+out:
+ phase->last_printed = curr;
+ fflush(stderr);
+}
+
+static void __stdio_progress__init(struct ui_progress *p)
+{
+ /*
+ * The default step (total / 16) is meant for the TUI progress
+ * bar, for stdio, where a percentage is printed, use 1% steps.
+ */
+ p->next = p->step = p->total / 100 ?: 1;
+
+ if (stdio_progress__depth == STDIO_PROGRESS__MAX_DEPTH) {
+ /*
+ * Out of room: don't start this phase, stdio_progress__update()
+ * then ignores its updates, as it doesn't match the innermost
+ * running phase, and its finish() is swallowed below, so that
+ * it doesn't complete the phase that encloses it. Completing
+ * that one here, to make room, would leave its own finish()
+ * without a phase to complete, which is what would then get
+ * out of sync, finishing the phases above it one by one.
+ */
+ pr_warning("progress phases nested deeper than %d, not showing progress for %s\n",
+ STDIO_PROGRESS__MAX_DEPTH, p->title);
+ stdio_progress__dropped++;
+ return;
+ }
+
+ /* Start a nested phase in a line of its own. */
+ if (stdio_progress__depth && stdio_progress__is_tty)
+ fputc('\n', stderr);
+
+ stdio_progress__stack[stdio_progress__depth++] =
+ (struct stdio_progress_phase) {
+ .p = p,
+ .last_printed = 0,
+ .last_len = 0,
+ };
+
+ stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+ p->curr);
+}
+
+static void stdio_progress__update(struct ui_progress *p)
+{
+ /*
+ * Phases are started/finished via init/finish, if we get an
+ * update that doesn't match the innermost running phase then
+ * something got out of sync, print nothing rather than some
+ * other phase's numbers, or, with no phases at all, reading
+ * stdio_progress__stack[-1].
+ */
+ if (!stdio_progress__depth ||
+ stdio_progress__stack[stdio_progress__depth - 1].p != p)
+ return;
+
+ stdio_progress__print_phase(&stdio_progress__stack[stdio_progress__depth - 1],
+ p->curr);
+}
+
+static void stdio_progress__finish(void)
+{
+ struct stdio_progress_phase *phase;
+
+ /*
+ * A phase that didn't fit on the stack still gets its finish(), and
+ * being the innermost one, it is the first to come: swallow it, or
+ * it would complete the phase that encloses it.
+ */
+ if (stdio_progress__dropped) {
+ stdio_progress__dropped--;
+ return;
+ }
+
+ if (!stdio_progress__depth)
+ return;
+
+ phase = &stdio_progress__stack[--stdio_progress__depth];
+
+ /*
+ * As we print only at 1% steps, the last line printed may have
+ * stopped short of the total, so close this phase showing it as
+ * complete, unless that was what got printed already.
+ */
+ if (phase->last_printed != phase->p->total)
+ stdio_progress__print_phase(phase, phase->p->total);
+
+ phase->last_printed = 0;
+ phase->last_len = 0;
+
+ if (stdio_progress__is_tty)
+ fputc('\n', stderr);
+
+ fflush(stderr);
+}
+
+static struct ui_progress_ops stdio_progress__ops = {
+ .init = __stdio_progress__init,
+ .update = stdio_progress__update,
+ .finish = stdio_progress__finish,
+};
+
+void stdio_progress__init(void)
+{
+ stdio_progress__is_tty = isatty(STDERR_FILENO) == 1;
+ ui_progress__ops = &stdio_progress__ops;
+}
diff --git a/tools/perf/util/ordered-events.c b/tools/perf/util/ordered-events.c
index a5857f9f5af2d3de..4063403d4978b45e 100644
--- a/tools/perf/util/ordered-events.c
+++ b/tools/perf/util/ordered-events.c
@@ -237,14 +237,16 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
ui_progress__init(&prog, oe->nr_events, "Processing time ordered events...");
list_for_each_entry_safe(iter, tmp, head, list) {
- if (session_done())
- return 0;
+ if (session_done()) {
+ ret = 0;
+ goto out_progress;
+ }
if (iter->timestamp > limit)
break;
ret = oe->deliver(oe, iter);
if (ret < 0)
- return ret;
+ goto out_progress;
ordered_events__delete(oe, iter);
oe->last_flush = iter->timestamp;
@@ -258,10 +260,17 @@ static int do_flush(struct ordered_events *oe, bool show_progress)
else if (last_ts <= limit)
oe->last = list_entry(head->prev, struct ordered_event, list);
+ ret = 0;
+out_progress:
+ /*
+ * Always pair ui_progress__init() with ui_progress__finish(), the
+ * stdio progress backend tracks phases on a stack and an early
+ * return that skipped the finish would leave the dead 'prog' on it.
+ */
if (show_progress)
ui_progress__finish();
- return 0;
+ return ret;
}
static int __ordered_events__flush(struct ordered_events *oe, enum oe_flush how,
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index 3237870a1a34b62c..85166042e899aada 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -3146,8 +3146,12 @@ static int __perf_session__process_pipe_events(struct perf_session *session)
cur_size = sizeof(union perf_event);
buf = malloc(cur_size);
- if (!buf)
- return -errno;
+ if (!buf) {
+ err = -errno;
+ if (update_prog)
+ ui_progress__finish();
+ return err;
+ }
ordered_events__set_copy_on_queue(oe, true);
more:
event = buf;
@@ -3646,8 +3650,10 @@ static int __perf_session__process_dir_events(struct perf_session *session)
}
rd = calloc(nr_readers, sizeof(struct reader));
- if (!rd)
+ if (!rd) {
+ ui_progress__finish();
return -ENOMEM;
+ }
rd[0] = (struct reader) {
.fd = perf_data__fd(session->data),
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 6/8] perf scripts: Add perf-stuck, to tell where a running perf is stuck
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
` (4 preceding siblings ...)
2026-09-13 22:28 ` [PATCH 5/8] perf report: Add --progress option Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 7/8] perf annotate-data: Resolve type DIEs in the debug file they came from Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 8/8] perf mem record: Request PERF_SAMPLE_CPU by default Arnaldo Carvalho de Melo
7 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
A perf that takes forever is hard to tell apart from one that is stuck
in a loop, and when it is stuck there is no way to know where without
attaching gdb to it and looking around, which is what this does, from
the outside, sampling /proc/<pid> at a fixed interval:
⬢ [acme@toolbx perf-tools-next]$ tools/perf/scripts/perf-stuck.sh -i 15 -n 4 -l ibs.log $(pgrep -x perf)
watching 2297313 (perf report --progress -s type -i perf.data.ibs) every 15s
09:39:34 state=R cpu=+0 (0.00s) rss=664295kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB Merging related events... [ 39.0%] 197574 / 506686
09:39:49 state=R cpu=+1496 (14.96s) rss=665383kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=1 Merging related events... [ 39.0%] 197574 / 506686
09:40:04 state=R cpu=+1497 (14.97s) rss=665383kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=2 Merging related events... [ 39.0%] 197574 / 506686
09:40:19 state=R cpu=+1497 (14.97s) rss=502250kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=0 Merging related events... [ 61.0%] 309026 / 506686
The CPU time used grows by a whole interval on every sample while the
[stack] mapping, that would be moving down if this was recursion, stays
put, so that one is spinning, and the last line of the progress log of
'perf report --progress' tells in which phase.
With -g it runs gdb when no progress is made for two consecutive
samples, using the perf-stuck.gdb that sits next to it, which adds the
perf-die-chain command used by the next patch, to print the DIE chain a
DWARF type chasing loop is walking when the perf being watched is stuck
in one of those.
It is a prototype: this wants to become a first class 'perf stuck'
command, sampling a running process from inside perf, instead of this
shell script poking at /proc and shelling out to gdb.
Example:
⬢ [acme@toolbx perf-tools-next]$ tools/perf/scripts/perf-stuck.sh -i 1 -n 3 -g $(pgrep -x sleep)
watching 2342826 (sleep 45 ) every 1s
11:58:36 state=S cpu=+0 (0.00s) rss=492kB stack=7ffe21f89000-7ffe21faa000 size=132kB stuck=2 (no progress log)
... gdb output of 2342826 in /tmp/perf-stuck-gdb.Uz9i5p
⬢ [acme@toolbx perf-tools-next]$ tail -4 /tmp/perf-stuck-gdb.Uz9i5p
#3 0x0000555bc6ddc28f in main ()
not in a DWARF type chaser, try: bt
not in find_data_type()
[Inferior 1 (process 2342826) detached]
⬢ [acme@toolbx perf-tools-next]$
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/scripts/perf-stuck.gdb | 104 ++++++++++++++++
tools/perf/scripts/perf-stuck.sh | 194 ++++++++++++++++++++++++++++++
2 files changed, 298 insertions(+)
create mode 100644 tools/perf/scripts/perf-stuck.gdb
create mode 100755 tools/perf/scripts/perf-stuck.sh
diff --git a/tools/perf/scripts/perf-stuck.gdb b/tools/perf/scripts/perf-stuck.gdb
new file mode 100644
index 0000000000000000..53e9019828b6ef62
--- /dev/null
+++ b/tools/perf/scripts/perf-stuck.gdb
@@ -0,0 +1,104 @@
+# SPDX-License-Identifier: GPL-2.0
+#
+# gdb commands for a perf that is stuck, used by perf-stuck.sh -g and usable
+# directly:
+#
+# gdb -p $(pgrep -x perf) -batch -x perf-stuck.gdb -ex bt
+#
+# PROTOTYPE: part of the perf-stuck.sh stopgap, see the note at the start of
+# that script: this wants to move into a first class 'perf stuck' command,
+# which would print these DIE chains by itself, without gdb.
+#
+# The settings are the ones that keep a batch attach from stopping to ask
+# questions (debuginfod, pagination) and that make the output readable.
+#
+# The commands are for the DWARF type chasers in util/dwarf-aux.c, the
+# functions a data type profiling 'perf report -s type' spins in when a
+# debug info file has a type chain that got into a cycle:
+#
+# perf-die-chain <function> <die variable> [iterations]
+# perf-die-chain-all [iterations]
+# perf-dso
+#
+# For each iteration of the chasing loop they print the DIE address, the
+# CU it came from, its offset in the debug file, its tag and its name: a
+# cycle shows up as the same handful of (addr, cu) pairs repeating, and a
+# CU that changes from one iteration to the next means the chase is
+# hopping between a debug file and its dwz common file.
+
+set pagination off
+set confirm off
+set debuginfod enabled off
+set print pretty on
+set height 0
+set width 0
+
+define perf-die-chain
+ if $argc < 2
+ printf "usage: perf-die-chain <function> <die variable> [iterations]\n"
+ else
+ frame function $arg0
+ if $argc == 3
+ set $perf_die_chain_n = $arg2
+ else
+ set $perf_die_chain_n = 10
+ end
+ set $perf_die_chain_head = $pc
+ set $perf_die_chain_i = 0
+ while $perf_die_chain_i < $perf_die_chain_n
+ printf "chain[%d] die=%p addr=%p cu=%p off=0x%lx tag=%d name=%s\n", $perf_die_chain_i, $arg1, $arg1->addr, $arg1->cu, ((Dwarf_Off) dwarf_dieoffset($arg1)), ((int) dwarf_tag($arg1)), ((char *) dwarf_diename($arg1))
+ until *$perf_die_chain_head
+ set $perf_die_chain_i = $perf_die_chain_i + 1
+ end
+ end
+end
+
+document perf-die-chain
+Print the DIE chain being walked by a DWARF type chasing loop.
+usage: perf-die-chain <function> <die variable> [iterations]
+ perf-die-chain die_get_pointer_type type_die
+ perf-die-chain __die_get_real_type vr_die
+ perf-die-chain die_get_real_type vr_die
+end
+
+define perf-die-chain-all
+ if $argc == 0
+ set $perf_die_chain_n = 10
+ else
+ set $perf_die_chain_n = $arg0
+ end
+ if $_any_caller_is("die_get_pointer_type", 20)
+ printf "stuck in die_get_pointer_type():\n"
+ perf-die-chain die_get_pointer_type type_die $perf_die_chain_n
+ else
+ if $_any_caller_is("__die_get_real_type", 20)
+ printf "stuck in __die_get_real_type():\n"
+ perf-die-chain __die_get_real_type vr_die $perf_die_chain_n
+ else
+ if $_any_caller_is("die_get_real_type", 20)
+ printf "stuck in die_get_real_type():\n"
+ perf-die-chain die_get_real_type vr_die $perf_die_chain_n
+ else
+ printf "not in a DWARF type chaser, try: bt\n"
+ end
+ end
+ end
+end
+
+document perf-die-chain-all
+Find which DWARF type chaser the process is in and print the DIE chain.
+usage: perf-die-chain-all [iterations]
+end
+
+define perf-dso
+ if $_any_caller_is("find_data_type", 20)
+ frame function find_data_type
+ printf "dso=%s ip=0x%lx sym=%s\n", dloc->ms->map->dso->name, dloc->ip, dloc->ms->sym->name
+ else
+ printf "not in find_data_type()\n"
+ end
+end
+
+document perf-dso
+Print the dso, ip and symbol of the data location being resolved.
+end
diff --git a/tools/perf/scripts/perf-stuck.sh b/tools/perf/scripts/perf-stuck.sh
new file mode 100755
index 0000000000000000..70e59192b1d9658e
--- /dev/null
+++ b/tools/perf/scripts/perf-stuck.sh
@@ -0,0 +1,194 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+#
+# perf-stuck - tell a spinning perf apart from a blocked or recursing one
+#
+# Arnaldo Carvalho de Melo <acme@redhat.com>
+#
+# PROTOTYPE: this wants to become a first class 'perf stuck' command, that
+# samples a running perf, or any other process, from inside perf, with the
+# knowledge of the phases perf goes through and of the DWARF type chasing
+# loops built in, instead of this shell script poking at /proc and shelling
+# out to gdb. It is here as a stopgap, to be able to tell where a perf is
+# stuck while looking at hangs such as the one 'perf report -s type' hits
+# on dwz compressed debug info.
+#
+# Samples /proc/<pid> at a fixed interval and prints, for each sample:
+#
+# the CPU time used since the previous sample, so a process burning a
+# full interval's worth of ticks is spinning, while one using none is
+# blocked
+#
+# the [stack] mapping start, which moves down as the stack grows, the
+# giveaway for runaway recursion, together with its size
+#
+# the last line of a progress log, when one is given, e.g. the stderr
+# of 'perf report --progress', to see which phase is stuck
+#
+# A process that burns CPU with a constant stack and no progress is in an
+# unbounded loop, e.g. a die_get_pointer_type() chain that got into a
+# cycle, while one whose [stack] start keeps moving down is recursing.
+#
+# With -g it runs gdb, using the perf-stuck.gdb that sits next to this
+# script, when no progress is made for two consecutive samples, which for
+# a perf in a DWARF type chasing loop prints the DIE chain it is walking.
+#
+# usage: perf-stuck.sh [options] <pid|process-name>
+
+set -u
+
+usage() {
+ cat <<-EOF
+ usage: perf-stuck.sh [options] <pid|process-name>
+
+ -i <secs> sampling interval (default: 10)
+ -n <count> stop after this many samples (default: watch till it exits)
+ -l <file> progress log, its last line is printed with every sample
+ -g run gdb with perf-stuck.gdb when no progress is made for
+ two consecutive samples, writing the output to a temp file
+ -x <file> use this gdb command file instead of perf-stuck.gdb
+ -h this help
+ EOF
+ exit "${1:-0}"
+}
+
+interval=10
+count=0
+progress_log=
+use_gdb=
+gdb_cmds=
+
+while getopts "i:n:l:gx:h" opt; do
+ case "$opt" in
+ i) interval=$OPTARG ;;
+ n) count=$OPTARG ;;
+ l) progress_log=$OPTARG ;;
+ g) use_gdb=1 ;;
+ x) gdb_cmds=$OPTARG ;;
+ h) usage 0 ;;
+ *) usage 1 ;;
+ esac
+done
+shift $((OPTIND - 1))
+
+[ $# -eq 1 ] || usage 1
+
+if [[ "$1" =~ ^[0-9]+$ ]]; then
+ pid=$1
+else
+ pid=$(pgrep -x "$1" | head -1)
+ [ -n "$pid" ] || { echo "no process named '$1'"; exit 1; }
+fi
+
+[ -d /proc/"$pid" ] || { echo "no process $pid"; exit 1; }
+
+if [ -n "$use_gdb" ] && [ -z "$gdb_cmds" ]; then
+ gdb_cmds=$(dirname "$0")/perf-stuck.gdb
+ [ -r "$gdb_cmds" ] || { echo "cannot read $gdb_cmds"; exit 1; }
+fi
+
+hz=$(getconf CLK_TCK)
+psz=$(getconf PAGESIZE)
+prev_cpu=
+prev_stack=
+prev_progress=
+stuck=0
+gdb_done=
+nsample=0
+
+# The command line is whatever the process was started with, so drop the
+# control characters from it: a process started with escape sequences in
+# its arguments, e.g. one replaying a log line, would otherwise get them
+# replayed on the terminal of whoever runs this.
+cmdline=$(tr '\0' ' ' < /proc/"$pid"/cmdline | tr -d '[:cntrl:]')
+
+echo "watching $pid ($cmdline) every ${interval}s"
+
+while :; do
+ if [ ! -d /proc/"$pid" ]; then
+ echo "$(date +%T) process gone"
+ break
+ fi
+
+ # Field 2, the command name, is in parentheses and can contain
+ # spaces, so drop it together with the pid before splitting: the
+ # fields after it then line up, with the state, the utime+stime
+ # pair and the RSS landing where they are read below. Printing the
+ # CPU time with %d instead of relying on awk's default output format
+ # keeps it out of scientific notation, that bash arithmetic cannot
+ # parse, once it goes past six digits, i.e. some 16 minutes of CPU
+ # at 100 Hz.
+ if ! stat_line=$(awk '{ sub(/^[^ ]+ \(.*\) /, "");
+ printf "%s %d %d\n", $1, $12 + $13, $22 }' \
+ /proc/"$pid"/stat 2>/dev/null); then
+ echo "$(date +%T) process gone"
+ break
+ fi
+
+ # The process can be gone between the check above and this read, in
+ # which case there is nothing to report: 'set -u' would otherwise
+ # turn the unbound fields into an aborted script.
+ if [ -z "$stat_line" ]; then
+ echo "$(date +%T) process gone"
+ break
+ fi
+
+ stat=($stat_line)
+ state=${stat[0]}
+ cpu=${stat[1]}
+ # field 24 of /proc/<pid>/stat, the resident set size in pages
+ rss=$(( stat[2] * psz / 1024 ))
+
+ stack=$(awk '/\[stack\]/{print $1; exit}' /proc/"$pid"/maps 2>/dev/null)
+ if [ -n "$stack" ]; then
+ stack_start=0x${stack%-*}
+ stack_size=$(( 0x${stack#*-} - stack_start ))
+ stack_txt="$stack size=$((stack_size / 1024))kB"
+ else
+ stack_start=
+ stack_txt="-"
+ fi
+
+ progress=
+ [ -n "$progress_log" ] && [ -s "$progress_log" ] && progress=$(tail -1 "$progress_log")
+
+ if [ -n "$prev_cpu" ]; then
+ cpu_delta=$(( cpu - prev_cpu ))
+ # With a progress log, count the samples that show no progress,
+ # without one there is no progress to look at, so count them all:
+ # -g then looks at where the process is after two intervals.
+ if [ -z "$progress_log" ] ||
+ { [ -n "$progress" ] && [ "$progress" = "$prev_progress" ]; }; then
+ stuck=$((stuck + 1))
+ else
+ stuck=0
+ fi
+ stuck_txt="stuck=${stuck}"
+ [ "$stack_start" != "$prev_stack" ] && stuck_txt="$stuck_txt STACK"
+ else
+ cpu_delta=0
+ stuck_txt=""
+ fi
+
+ printf '%s state=%s cpu=+%d (%d.%02ds) rss=%dkB stack=%s %s %s\n' \
+ "$(date +%T)" "$state" "$cpu_delta" \
+ $(( cpu_delta / hz )) $(( (cpu_delta % hz) * 100 / hz )) \
+ "$rss" "$stack_txt" "$stuck_txt" "${progress:-(no progress log)}"
+
+ if [ -n "$use_gdb" ] && [ -z "$gdb_done" ] && [ "$stuck" -ge 2 ]; then
+ gdb_log=$(mktemp /tmp/perf-stuck-gdb.XXXXXX)
+ gdb -p "$pid" -batch -x "$gdb_cmds" -ex bt \
+ -ex 'perf-die-chain-all' -ex perf-dso -ex detach > "$gdb_log" 2>&1
+ gdb_done=1
+ echo "... gdb output of $pid in $gdb_log"
+ fi
+
+ prev_cpu=$cpu
+ prev_stack=$stack_start
+ prev_progress=$progress
+
+ nsample=$((nsample + 1))
+ [ "$count" -gt 0 ] && [ "$nsample" -ge "$count" ] && break
+
+ sleep "$interval"
+done
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 7/8] perf annotate-data: Resolve type DIEs in the debug file they came from
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
` (5 preceding siblings ...)
2026-09-13 22:28 ` [PATCH 6/8] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 8/8] perf mem record: Request PERF_SAMPLE_CPU by default Arnaldo Carvalho de Melo
7 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
A 'perf report -s type' on a 783 MB AMD IBS data type profiling session
hangs, burning all of a CPU and producing no output:
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -s type -i perf.data.ibs > /dev/null
Processing events... [100.0%] 746M / 746M
Merging related events... [ 58.0%] 293828 / 506686
Merging related events... [ 59.0%] 298894 / 506686
Merging related events... [ 60.0%] 303960 / 506686
Merging related events... [ 61.0%] 309026 / 506686
[ ... nothing else, ever ... ]
It is a spin and not a slow path: with perf-stuck the CPU time used
grows by a whole interval on every sample, while the [stack] mapping,
which would be moving down if this was recursion, stays put:
⬢ [acme@toolbx perf-tools-next]$ tools/perf/scripts/perf-stuck.sh -i 15 -n 4 -l ibs.log $(pgrep -x perf)
watching 2297313 (perf report --progress -s type -i perf.data.ibs) every 15s
09:39:34 state=R cpu=+0 (0.00s) rss=664295kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB Merging related events... [ 39.0%] 197574 / 506686
09:39:49 state=R cpu=+1496 (14.96s) rss=665383kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=1 Merging related events... [ 39.0%] 197574 / 506686
09:40:04 state=R cpu=+1497 (14.97s) rss=665383kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=2 Merging related events... [ 39.0%] 197574 / 506686
09:40:19 state=R cpu=+1497 (14.97s) rss=502250kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=0 Merging related events... [ 61.0%] 309026 / 506686
Attaching gdb and stepping the loop, printing dwarf_dieoffset() and
dwarf_tag() for the DIE being chased on each trip round it, shows the
chase never moving, on a typedef that refers to itself:
⬢ [acme@toolbx perf-tools-next]$ gdb -p $(pgrep -x perf) -batch -x tools/perf/scripts/perf-stuck.gdb -ex 'perf-die-chain die_get_pointer_type type_die 8' -ex perf-dso -ex detach
stuck in die_get_pointer_type():
#3 0x0000000000772421 in die_get_pointer_type (type_die=0x7ffc3ae4b710, type_die@entry=0x7ffc3ae4b6f0, die_mem=die_mem@entry=0x7ffc3ae4b710) at util/dwarf-aux.c:327
327 type_die = die_get_type(type_die, die_mem);
chain[0] die=0x7ffc3ae4b710 addr=0x7ff6811591ef cu=0x44ef0878 off=0x1f tag=22 name=(null)
chain[1] die=0x7ffc3ae4b710 addr=0x7ff6811591ef cu=0x44ef0878 off=0x1f tag=22 name=(null)
[ ... the very same DIE, forever ... ]
dso=/usr/lib64/libz.so.1.3.1.zlib-ng ip=0xe2e sym=build_tree
The DIE is at offset 0x1f of the debug info of libz.so.1, which is
zlib-ng, and is one of the dwz compressed ones: the type DIEs shared by
more than one CU live in the common file, where 0x1f is a perfectly good
DW_TAG_base_type:
⬢ [acme@toolbx perf-tools-next]$ readelf --debug-dump=info /usr/lib/debug/.dwz/zlib-ng-2.3.3-3.fc44.x86_64 | sed -n '/Compilation Unit @ offset 0:/,/Compilation Unit @ offset 0x5f:/p'
Compilation Unit @ offset 0:
Length: 0x5b (32-bit)
Version: 5
Unit Type: DW_UT_partial (3)
<0><c>: Abbrev Number: 1 (DW_TAG_partial_unit)
[...]
<1><1f>: Abbrev Number: 62 (DW_TAG_base_type)
<20> DW_AT_byte_size : 4
<21> DW_AT_encoding : 5 (signed)
<22> DW_AT_name : int
while in the main debug file that same offset is not a DIE at all, it is
the start of another unit's header:
⬢ [acme@toolbx perf-tools-next]$ readelf --debug-dump=info /usr/lib/debug/usr/lib64/libz.so.1.3.1.zlib-ng-2.3.3-3.fc44.x86_64.debug | grep 'Compilation Unit @' | head -2
Compilation Unit @ offset 0:
Compilation Unit @ offset 0x1f:
die_collect_vars() saves the dwarf_dieoffset() of the type DIE, which is
relative to the file that DIE lives in, and update_var_state() then hands
that offset to dwarf_offdie() together with the main debug file, where
dwarf_offdie() parses whatever is there as a DIE, in this case the one
typedef that refers to itself, which is what makes the "follow the
typedefs and qualifiers until a pointer or an array type" loop in
die_get_pointer_type() spin.
So record, next to the offset, whether the type DIE was in the file the
variable DIE came from or in the dwz common one, and resolve the offset
in the file it came from, which is what the new die_get_type_die() does.
Which of the two that is does not have to be guessed from what is at the
offset: dwz encodes the references into its common file as
DW_FORM_GNU_ref_alt, so elfutils resolves them into the alt Dwarf and
the CU of the resulting type DIE belongs to that other file, so
comparing the Dwarf each of the two CUs belongs to, die_same_file(),
settles it exactly, for any number of hops from the variable DIE.
Then bound the chases themselves: no sane chain of typedefs and
qualifiers is 32 DIEs long, no sane nesting for the struct and union
members that __add_member_cb() follows recursively is 8 deep, and the
same recursion bound covers the type names die_get_typename_from_type()
builds by following pointers and arrays, so that a debug info file
broken in some other way makes perf give up on a type, telling about it
with pr_debug, visible with -v, instead of looking like it hung. The
member nesting bound is reported with pr_debug_dtp (visible with -vvv or
-D type-profile).
Cycles can't occur in member trees from valid DWARF, embedded members
can't be recursive in C, so the nesting bound only ever bites legitimate
depth and the member where the recursion was cut is marked 'truncated',
which the JSON exporter added in a subsequent series reports to its
consumers, so that they can tell a truncated tree from one that really
ends there.
die_get_type_die() then has no fallback to the other file: resolving an
alt file offset in the main file is the misparse above, so it resolves
in the file the offset was recorded as belonging to, and gives up on the
type when the offset does not resolve there. The tag the type DIE had is
kept as a sanity check, a mismatch means the debug info changed under
perf or is broken in yet another way, and the bounds on the type chasers
remain the backstop: a debug info file broken in some other way still
makes perf give up on a type with a pr_debug instead of hanging.
Testing:
Before, on the 783 MB AMD IBS session, killed after some 8 minutes stuck
at 61%:
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -s type -i perf.data.ibs > /dev/null
Processing events... [100.0%] 746M / 746M
Merging related events... [ 61.0%] 309026 / 506686
⬢ [acme@toolbx perf-tools-next]$
After, the whole session is processed, and the zlib-ng types, from the
build_tree() hist entry that used to hang it, show up:
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -s type -i perf.data.ibs > ibs.out
Processing events... [100.0%] 746M / 746M
Merging related events... [100.0%] 506686 / 506686
Sorting events for output... [100.0%] 1428 / 1428
⬢ [acme@toolbx perf-tools-next]$ grep -E "deflate_state|inflate_state|internal_state" ibs.out
0.00% deflate_state
0.00% deflate_state*
0.00% struct inflate_state
0.00% struct inflate_state*
0.00% struct internal_state
⬢ [acme@toolbx perf-tools-next]$ perf test 17 27 30 31 85 88
17: Match and link multiple hists : Ok
27: Filter hist entries : Ok
30: Sort output of hist entries : Ok
31: Cumulate child hist entries : Ok
85: Test that perf report includes file offsets and event type names in diagnostic messages. : Ok
88: Test that perf report handles truncated perf.data gracefully (no crash, no segfault — clean error exit).: Skip
⬢ [acme@toolbx perf-tools-next]$
Requiring elfutils 0.160 for this: dwarf_cu_getdwarf(), the function
that tells which Dwarf a CU belongs to, first appeared in 0.160 ("libdw:
New functions dwarf_cu_getdwarf, dwarf_cu_die", elfutils NEWS), so the
libdw feature test now probes for it, in tools/build/feature/test-libdw.c,
and Makefile.config says 0.160 where it said 0.157. The probe takes the
address instead of calling it, as that is all that is needed to make a
0.157-0.159 elfutils, from 2014 and without the symbol, disable dwarf
support with the existing message rather than fail to link dwarf-aux.c.
Fixes: 06b2ce75386df04b ("perf annotate-data: Maintain variable type info")
Fixes: 55ee3d005d62279d ("perf annotate-data: Add a cache for global variable types")
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/build/feature/test-libdw.c | 15 ++-
tools/perf/Makefile.config | 2 +-
tools/perf/util/annotate-data.c | 59 ++++++++---
tools/perf/util/annotate-data.h | 3 +
tools/perf/util/dwarf-aux.c | 171 +++++++++++++++++++++++++++----
tools/perf/util/dwarf-aux.h | 33 ++++++
6 files changed, 247 insertions(+), 36 deletions(-)
diff --git a/tools/build/feature/test-libdw.c b/tools/build/feature/test-libdw.c
index aabd63ca76b4d7e6..23e1ba6ff3466b9f 100644
--- a/tools/build/feature/test-libdw.c
+++ b/tools/build/feature/test-libdw.c
@@ -49,8 +49,21 @@ int test_elfutils(void)
return 0;
}
+/*
+ * elfutils 0.160 and later: used to tell which debug file a DIE lives in,
+ * the dwz alt file or the main one, see die_same_file() in
+ * tools/perf/util/dwarf-aux.c. Only the symbol is needed, so take its
+ * address instead of calling it.
+ */
+int test_libdw_cu_getdwarf(void)
+{
+ void *sym = (void *)dwarf_cu_getdwarf;
+
+ return sym == NULL;
+}
+
int main(void)
{
return test_libdw() + test_libdw_unwind() + test_libdw_getlocations() +
- test_libdw_getcfi() + test_elfutils();
+ test_libdw_getcfi() + test_libdw_cu_getdwarf() + test_elfutils();
}
diff --git a/tools/perf/Makefile.config b/tools/perf/Makefile.config
index 4d5993da9f94579f..fa78f50db60179f3 100644
--- a/tools/perf/Makefile.config
+++ b/tools/perf/Makefile.config
@@ -470,7 +470,7 @@ else
else
ifneq ($(feature-libdw), 1)
ifndef NO_LIBDW
- $(warning No libdw.h found or old libdw.h found or elfutils is older than 0.157, disables dwarf support. Please install new elfutils-devel/libdw-dev)
+ $(warning No libdw.h found or old libdw.h found or elfutils is older than 0.160, disables dwarf support. Please install new elfutils-devel/libdw-dev)
NO_LIBDW := 1
endif
endif # Dwarf support
diff --git a/tools/perf/util/annotate-data.c b/tools/perf/util/annotate-data.c
index 4e4c587640823c81..06d27868887bdadf 100644
--- a/tools/perf/util/annotate-data.c
+++ b/tools/perf/util/annotate-data.c
@@ -221,6 +221,15 @@ static bool data_type_less(struct rb_node *node_a, const struct rb_node *node_b)
return strcmp(a->self.type_name, b->self.type_name) < 0;
}
+/*
+ * Members of struct/union members are added recursively, and the same DIE
+ * that is not what it looks like, the one that makes the type chasers in
+ * util/dwarf-aux.c spin, can make a member's type point back at one of its
+ * own ancestors, recursing until the stack is gone. Nothing usable comes
+ * out of nesting members this deep anyway.
+ */
+#define MAX_MEMBER_DEPTH 8
+
/* Recursively add new members for struct/union */
static int __add_member_cb(Dwarf_Die *die, void *arg)
{
@@ -235,6 +244,16 @@ static int __add_member_cb(Dwarf_Die *die, void *arg)
if (dwarf_tag(die) != DW_TAG_member)
return DIE_FIND_CB_SIBLING;
+ if (__die_get_real_type(die, &member_type) == NULL)
+ return DIE_FIND_CB_SIBLING;
+
+ if (dwarf_tag(&member_type) == DW_TAG_typedef) {
+ if (die_get_real_type(&member_type, &die_mem) == NULL)
+ return DIE_FIND_CB_SIBLING;
+ } else {
+ die_mem = member_type;
+ }
+
member = zalloc(sizeof(*member));
if (member == NULL)
return DIE_FIND_CB_END;
@@ -242,12 +261,6 @@ static int __add_member_cb(Dwarf_Die *die, void *arg)
strbuf_init(&sb, 32);
die_get_typename(die, &sb);
- __die_get_real_type(die, &member_type);
- if (dwarf_tag(&member_type) == DW_TAG_typedef)
- die_get_real_type(&member_type, &die_mem);
- else
- die_mem = member_type;
-
if (dwarf_aggregate_size(&die_mem, &size) < 0)
size = 0;
@@ -289,10 +302,23 @@ static int __add_member_cb(Dwarf_Die *die, void *arg)
}
member->size = size;
member->offset = loc + parent->offset;
+ member->depth = parent->depth + 1;
INIT_LIST_HEAD(&member->children);
list_add_tail(&member->node, &parent->children);
tag = dwarf_tag(&die_mem);
+ if (member->depth >= MAX_MEMBER_DEPTH) {
+ /*
+ * The JSON exporter added in a later series reports
+ * this to its consumers, so that they can tell a
+ * truncated tree from one that really ends here.
+ */
+ member->truncated = true;
+ pr_debug_dtp("member nesting limit reached at %s\n",
+ member->type_name ?: "(unknown type)");
+ return DIE_FIND_CB_SIBLING;
+ }
+
switch (tag) {
case DW_TAG_structure_type:
case DW_TAG_union_type:
@@ -645,6 +671,8 @@ struct global_var_entry {
u64 start;
u64 end;
u64 die_offset;
+ int die_tag;
+ bool from_alt; /* die_offset is relative to the alt (dwz) file */
};
static int global_var_cmp(const void *_key, const struct rb_node *node)
@@ -682,7 +710,7 @@ static struct global_var_entry *global_var__find(struct data_loc_info *dloc, u64
}
static bool global_var__add(struct data_loc_info *dloc, u64 addr,
- const char *name, Dwarf_Die *type_die)
+ const char *name, Dwarf_Die *type_die, bool from_alt)
{
struct dso *dso = map__dso(dloc->ms->map);
struct global_var_entry *gvar;
@@ -704,6 +732,8 @@ static bool global_var__add(struct data_loc_info *dloc, u64 addr,
gvar->start = addr;
gvar->end = addr + size;
gvar->die_offset = dwarf_dieoffset(type_die);
+ gvar->die_tag = dwarf_tag(type_die);
+ gvar->from_alt = from_alt;
rb_add(&gvar->node, dso__global_vars(dso), global_var_less);
return true;
@@ -778,12 +808,14 @@ static void global_var__collect(struct data_loc_info *dloc)
if (pos->reg != -1)
continue;
- if (!dwarf_offdie(dwarf, pos->die_off, &type_die))
+ if (!die_get_type_die(dwarf, pos->die_off, pos->die_tag,
+ pos->from_alt, &type_die))
continue;
get_global_var_info(dloc, pos->addr, &var_name, &var_offset);
- global_var__add(dloc, pos->addr, var_name, &type_die);
+ global_var__add(dloc, pos->addr, var_name, &type_die,
+ pos->from_alt);
}
delete_var_types(var_types);
@@ -808,7 +840,8 @@ bool get_global_var_type(Dwarf_Die *cu_die, struct data_loc_info *dloc,
gvar = global_var__find(dloc, var_addr);
if (gvar) {
- if (!dwarf_offdie(dloc->di->dbg, gvar->die_offset, type_die))
+ if (!die_get_type_die(dloc->di->dbg, gvar->die_offset,
+ gvar->die_tag, gvar->from_alt, type_die))
return false;
*var_offset = var_addr - gvar->start;
@@ -838,7 +871,8 @@ bool get_global_var_type(Dwarf_Die *cu_die, struct data_loc_info *dloc,
ok:
/* The address should point to the start of the variable */
- global_var__add(dloc, var_addr - *var_offset, var_name, type_die);
+ global_var__add(dloc, var_addr - *var_offset, var_name, type_die,
+ !die_same_file(cu_die, type_die));
return true;
}
@@ -893,7 +927,8 @@ static void update_var_state(struct type_state *state, struct data_loc_info *dlo
continue;
}
/* Get the type DIE using the offset */
- if (!dwarf_offdie(dloc->di->dbg, var->die_off, &mem_die))
+ if (!die_get_type_die(dloc->di->dbg, var->die_off,
+ var->die_tag, var->from_alt, &mem_die))
continue;
if (var->reg == DWARF_REG_FB || var->reg == fbreg || var->reg == state->stack_reg) {
diff --git a/tools/perf/util/annotate-data.h b/tools/perf/util/annotate-data.h
index c26130744260955f..d85866e83fbda69a 100644
--- a/tools/perf/util/annotate-data.h
+++ b/tools/perf/util/annotate-data.h
@@ -57,6 +57,9 @@ struct annotated_member {
char *var_name;
int offset;
int size;
+ unsigned int depth;
+ /* Children not expanded because the nesting limit was reached */
+ bool truncated;
};
/**
diff --git a/tools/perf/util/dwarf-aux.c b/tools/perf/util/dwarf-aux.c
index d7160f87ac7d7ab3..7acb431fd34a8ecb 100644
--- a/tools/perf/util/dwarf-aux.c
+++ b/tools/perf/util/dwarf-aux.c
@@ -266,16 +266,35 @@ Dwarf_Die *die_get_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
return NULL;
}
+/*
+ * The chases below cross typedefs and qualifiers to get to the type that
+ * is actually meant, and a DIE that is not what it looks like, e.g. one
+ * parsed at an offset that is not the start of a DIE in the file it was
+ * resolved in, can have a DW_AT_type that refers back to itself, which
+ * makes them spin forever: 'perf report -s type' did exactly that on the
+ * dwz compressed debug info of zlib-ng (libz.so.1), burning all of a CPU
+ * with no output while resolving a hist entry in build_tree().
+ *
+ * No sane chain is this long, so give up instead of hanging, telling about
+ * it so that the broken debug info can be looked at.
+ */
+#define MAX_TYPE_CHASE 32
+
/* Get a type die, but skip qualifiers */
Dwarf_Die *__die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
{
- int tag;
+ int tag, chase = 0;
do {
vr_die = die_get_type(vr_die, die_mem);
if (!vr_die)
- break;
+ return NULL;
tag = dwarf_tag(vr_die);
+ if (++chase > MAX_TYPE_CHASE) {
+ pr_debug("DWARF: qualifier chase limit reached at DIE 0x%lx\n",
+ (unsigned long)dwarf_dieoffset(vr_die));
+ return NULL;
+ }
} while (tag == DW_TAG_const_type ||
tag == DW_TAG_restrict_type ||
tag == DW_TAG_volatile_type ||
@@ -296,8 +315,15 @@ Dwarf_Die *__die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
*/
Dwarf_Die *die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
{
+ int chase = 0;
+
do {
vr_die = __die_get_real_type(vr_die, die_mem);
+ if (++chase > MAX_TYPE_CHASE) {
+ pr_debug("DWARF: typedef chase limit reached at DIE 0x%lx\n",
+ vr_die ? (unsigned long)dwarf_dieoffset(vr_die) : 0);
+ return NULL;
+ }
} while (vr_die && dwarf_tag(vr_die) == DW_TAG_typedef);
return vr_die;
@@ -314,7 +340,7 @@ Dwarf_Die *die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
*/
Dwarf_Die *die_get_pointer_type(Dwarf_Die *type_die, Dwarf_Die *die_mem)
{
- int tag;
+ int tag, chase = 0;
do {
tag = dwarf_tag(type_die);
@@ -324,6 +350,11 @@ Dwarf_Die *die_get_pointer_type(Dwarf_Die *type_die, Dwarf_Die *die_mem)
tag != DW_TAG_restrict_type && tag != DW_TAG_volatile_type &&
tag != DW_TAG_shared_type)
return NULL;
+ if (++chase > MAX_TYPE_CHASE) {
+ pr_debug("DWARF: pointer type chase limit reached at DIE 0x%lx\n",
+ (unsigned long)dwarf_dieoffset(type_die));
+ return NULL;
+ }
type_die = die_get_type(type_die, die_mem);
} while (type_die);
@@ -1118,17 +1149,27 @@ Dwarf_Die *die_find_member(Dwarf_Die *st_die, const char *name,
die_mem);
}
-/**
- * die_get_typename_from_type - Get the name of given type DIE
- * @type_die: a type DIE
- * @buf: a strbuf for result type name
- *
- * Get the name of @type_die and stores it to @buf. Return 0 if succeeded.
- * and Return -ENOENT if failed to find type name.
- * Note that the result will stores typedef name if possible, and stores
- * "*(function_type)" if the type is a function pointer.
+/*
+ * The name of a pointer or array type is built from the name of the type
+ * it points to or holds, so the recursion below follows DW_AT_type; a
+ * garbage DIE whose DW_AT_type refers back to itself makes it recurse
+ * forever, just like the chases above, so it gets the same bound.
*/
-int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
+static int __die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf,
+ int depth);
+
+static int __die_get_typename(Dwarf_Die *vr_die, struct strbuf *buf, int depth)
+{
+ Dwarf_Die type;
+
+ if (__die_get_real_type(vr_die, &type) == NULL)
+ return -ENOENT;
+
+ return __die_get_typename_from_type(&type, buf, depth);
+}
+
+static int __die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf,
+ int depth)
{
int tag, ret;
const char *tmp = "";
@@ -1155,7 +1196,12 @@ int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
/* Write a base name */
return strbuf_addf(buf, "%s%s", tmp, name ?: "");
}
- ret = die_get_typename(type_die, buf);
+ if (depth >= MAX_TYPE_CHASE) {
+ pr_debug("DWARF: type name recursion limit reached at DIE 0x%lx\n",
+ (unsigned long)dwarf_dieoffset(type_die));
+ return -ENOENT;
+ }
+ ret = __die_get_typename(type_die, buf, depth + 1);
if (ret < 0) {
/* void pointer has no type attribute */
if (tag == DW_TAG_pointer_type && ret == -ENOENT)
@@ -1166,6 +1212,21 @@ int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
return strbuf_addstr(buf, tmp);
}
+/**
+ * die_get_typename_from_type - Get the name of given type DIE
+ * @type_die: a type DIE
+ * @buf: a strbuf for result type name
+ *
+ * Get the name of @type_die and stores it to @buf. Return 0 if succeeded.
+ * and Return -ENOENT if failed to find type name.
+ * Note that the result will stores typedef name if possible, and stores
+ * "*(function_type)" if the type is a function pointer.
+ */
+int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
+{
+ return __die_get_typename_from_type(type_die, buf, 0);
+}
+
/**
* die_get_typename - Get the name of given variable DIE
* @vr_die: a variable DIE
@@ -1178,12 +1239,7 @@ int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
*/
int die_get_typename(Dwarf_Die *vr_die, struct strbuf *buf)
{
- Dwarf_Die type;
-
- if (__die_get_real_type(vr_die, &type) == NULL)
- return -ENOENT;
-
- return die_get_typename_from_type(&type, buf);
+ return __die_get_typename(vr_die, buf, 0);
}
/**
@@ -1632,6 +1688,22 @@ Dwarf_Die *die_find_variable_by_addr(Dwarf_Die *sc_die, Dwarf_Addr addr,
return result;
}
+/*
+ * Whether two DIEs live in the same debug file.
+ *
+ * dwarf_dieoffset() is relative to the file the DIE is in, so this is what
+ * tells an offset that has to be resolved in the dwz alt file, where dwz
+ * moved the type, from one that belongs to the main file: dwz encodes the
+ * references into its common file as DW_FORM_GNU_ref_alt, elfutils resolves
+ * them into the alt Dwarf and the CU of the resulting DIE belongs to that
+ * other file, so comparing the Dwarf each CU belongs to settles it exactly,
+ * rather than inferring it from what happens to be at the offset.
+ */
+bool die_same_file(Dwarf_Die *die_a, Dwarf_Die *die_b)
+{
+ return dwarf_cu_getdwarf(die_a->cu) == dwarf_cu_getdwarf(die_b->cu);
+}
+
static int __die_collect_vars_cb(Dwarf_Die *die_mem, void *arg)
{
struct die_var_type **var_types = arg;
@@ -1676,6 +1748,8 @@ static int __die_collect_vars_cb(Dwarf_Die *die_mem, void *arg)
vt->is_reg_var_addr = true;
vt->die_off = dwarf_dieoffset(&type_die);
+ vt->die_tag = dwarf_tag(&type_die);
+ vt->from_alt = !die_same_file(die_mem, &type_die);
vt->addr = start;
vt->end = end;
vt->has_range = (end != 0 || start != 0);
@@ -1695,7 +1769,8 @@ static int __die_collect_vars_cb(Dwarf_Die *die_mem, void *arg)
*
* Save all variables and parameters in the @sc_die and save them to @var_types.
* The @var_types is a singly-linked list containing type and location info.
- * Actual type can be retrieved using dwarf_offdie() with 'die_off' later.
+ * Actual type can be retrieved using die_get_type_die() with 'die_off',
+ * 'die_tag' and 'from_alt' later.
*
* Callers should free @var_types.
*/
@@ -1741,6 +1816,8 @@ static int __die_collect_global_vars_cb(Dwarf_Die *die_mem, void *arg)
return DIE_FIND_CB_END;
vt->die_off = dwarf_dieoffset(&type_die);
+ vt->die_tag = dwarf_tag(&type_die);
+ vt->from_alt = !die_same_file(die_mem, &type_die);
vt->addr = ops->number;
vt->end = 0;
vt->has_range = false;
@@ -1752,6 +1829,55 @@ static int __die_collect_global_vars_cb(Dwarf_Die *die_mem, void *arg)
return DIE_FIND_CB_SIBLING;
}
+/**
+ * die_get_type_die - Get a type DIE saved by die_collect_vars()
+ * @dbg: the main debug info
+ * @die_off: offset of the type DIE, from dwarf_dieoffset()
+ * @die_tag: tag that DIE had when the offset was saved
+ * @from_alt: whether the type DIE is in the dwz alt file
+ * @die_mem: where to store the resulting DIE
+ *
+ * See the comment in util/dwarf-aux.h: the offset is only meaningful in the
+ * file the DIE was in, which can be the dwz common file, so resolve it in
+ * the file @from_alt says it was in, the main file or its alt file, and use
+ * the DIE only when it has the @die_tag it had when the offset was saved.
+ * There is deliberately no fallback to the other file: resolving an alt
+ * file offset in the main file does not fail, it parses whatever is there
+ * as a DIE, and that is what hung 'perf report -s type'.
+ */
+Dwarf_Die *die_get_type_die(Dwarf *dbg, u64 die_off, int die_tag, bool from_alt,
+ Dwarf_Die *die_mem)
+{
+ Dwarf *target = dbg;
+ Dwarf_Die die;
+
+ if (from_alt) {
+ /*
+ * Deliberately no fallback to the main file when there is
+ * no alt file, or when the offset does not resolve in it:
+ * resolving an alt file offset in the main file does not
+ * fail, it parses whatever is there as a DIE, and that is
+ * what hung 'perf report -s type', see the comment in
+ * util/dwarf-aux.h.
+ */
+ target = dwarf_getalt(dbg);
+ if (target == NULL) {
+ pr_debug("DWARF: no alt (dwz) debug file to resolve the type DIE at offset 0x%lx in\n",
+ (unsigned long)die_off);
+ return NULL;
+ }
+ }
+
+ if (dwarf_offdie(target, die_off, &die) && dwarf_tag(&die) == die_tag) {
+ *die_mem = die;
+ return die_mem;
+ }
+
+ pr_debug("DWARF: no DIE with tag %d at offset 0x%lx in the %s debug file\n",
+ die_tag, (unsigned long)die_off, from_alt ? "alt" : "main");
+ return NULL;
+}
+
/**
* die_collect_global_vars - Save all global variables
* @cu_die: a CU DIE
@@ -1759,7 +1885,8 @@ static int __die_collect_global_vars_cb(Dwarf_Die *die_mem, void *arg)
*
* Save all global variables in the @cu_die and save them to @var_types.
* The @var_types is a singly-linked list containing type and location info.
- * Actual type can be retrieved using dwarf_offdie() with 'die_off' later.
+ * Actual type can be retrieved using die_get_type_die() with 'die_off',
+ * 'die_tag' and 'from_alt' later.
*
* Callers should free @var_types.
*/
diff --git a/tools/perf/util/dwarf-aux.h b/tools/perf/util/dwarf-aux.h
index 161f0bf980b6ee6a..149e0cbf63ccc82a 100644
--- a/tools/perf/util/dwarf-aux.h
+++ b/tools/perf/util/dwarf-aux.h
@@ -152,6 +152,8 @@ int die_get_scopes(Dwarf_Die *cu_die, Dwarf_Addr pc, Dwarf_Die **scopes);
struct die_var_type {
struct die_var_type *next;
u64 die_off;
+ int die_tag;
+ bool from_alt; /* die_off is relative to the alt (dwz) file */
u64 addr;
u64 end; /* end address of location range */
int reg;
@@ -183,6 +185,37 @@ Dwarf_Die *die_find_variable_by_addr(Dwarf_Die *sc_die, Dwarf_Addr addr,
/* Save all variables and parameters in this scope */
void die_collect_vars(Dwarf_Die *sc_die, struct die_var_type **var_types);
+/*
+ * Get the type DIE saved by die_collect_vars()/die_collect_global_vars().
+ *
+ * The offsets those save are the dwarf_dieoffset() of the type DIE, which is
+ * relative to the debug file that DIE lives in: the dwz common file, the alt
+ * file in libdw terms, for the types shared by more than one CU, the main
+ * file for the rest. Resolving an alt file offset in the main file does not
+ * fail: dwarf_offdie() parses whatever is at that offset there, and an offset
+ * that is a CU header in the main file reads back as a typedef whose
+ * DW_AT_type refers to itself, which is what hung 'perf report -s type' on
+ * the dwz compressed debug info of zlib-ng (libz.so.1).
+ *
+ * So @from_alt, recorded when the offset was saved, says which of the two
+ * files to resolve it in. It is not inferred from the DIE contents: dwz
+ * encodes references into the common file as DW_FORM_GNU_ref_alt, so
+ * elfutils resolves them into the alt Dwarf and the CU of the type DIE then
+ * belongs to that other file, which die_same_file() compares exactly.
+ *
+ * There is deliberately no fallback to the other file when the offset does
+ * not resolve: that fallback is the misparse above.
+ *
+ * @die_tag is then only a sanity check: the offset is of a DIE that had this
+ * tag when it was saved, so a mismatch means the debug info changed under us,
+ * or is broken, and giving up on the type is the right answer.
+ */
+Dwarf_Die *die_get_type_die(Dwarf *dbg, u64 die_off, int die_tag, bool from_alt,
+ Dwarf_Die *die_mem);
+
+/* Whether two DIEs live in the same debug file */
+bool die_same_file(Dwarf_Die *die_a, Dwarf_Die *die_b);
+
/* Save all global variables in this CU */
void die_collect_global_vars(Dwarf_Die *cu_die, struct die_var_type **var_types);
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 8/8] perf mem record: Request PERF_SAMPLE_CPU by default
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
` (6 preceding siblings ...)
2026-09-13 22:28 ` [PATCH 7/8] perf annotate-data: Resolve type DIEs in the debug file they came from Arnaldo Carvalho de Melo
@ 2026-09-13 22:28 ` Arnaldo Carvalho de Melo
7 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-13 22:28 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
The data-type profiling per-sample stream keys cross-CPU contention on
sample->cpu: without PERF_SAMPLE_CPU the cpu field is the (u32)-1 "no
CPU info" sentinel, documented as such in perf_session__deliver_event(),
and same-instance reads and writes from different cores are
indistinguishable from same-CPU traffic, so pahole's false-sharing
detector cannot tell them apart.
builtin-record.c already defines --sample-cpu and 'perf mem record'
forwards unknown options to the record parser, so passing it explicitly
works today; make it the default, next to the -d (addr) and -W (weight)
the command already requests, documenting it in perf-mem(1).
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/Documentation/perf-mem.txt | 4 ++++
tools/perf/builtin-mem.c | 9 +++++++++
2 files changed, 13 insertions(+)
diff --git a/tools/perf/Documentation/perf-mem.txt b/tools/perf/Documentation/perf-mem.txt
index 4d164836d0943119..fe51c5e3333dc4a0 100644
--- a/tools/perf/Documentation/perf-mem.txt
+++ b/tools/perf/Documentation/perf-mem.txt
@@ -14,6 +14,10 @@ DESCRIPTION
-----------
"perf mem record" runs a command and gathers memory operation data
from it, into perf.data. Perf record options are accepted and are passed through.
+It also requests the address (-d), the weight (-W, where supported) and the
+CPU id (--sample-cpu) of every sampled access by default; the CPU id is what
+lets per-sample analysis tell reads and writes to the same data from
+different cores apart from same-CPU traffic.
"perf mem report" displays the result. It invokes perf report with the
right set of options to display a memory access profile. By default, loads
diff --git a/tools/perf/builtin-mem.c b/tools/perf/builtin-mem.c
index 6101a26b3a781e69..a708e2549bae4ce7 100644
--- a/tools/perf/builtin-mem.c
+++ b/tools/perf/builtin-mem.c
@@ -135,6 +135,15 @@ static int __cmd_record(int argc, const char **argv, struct perf_mem *mem,
rec_argv[i++] = "-d";
+ /*
+ * The data-type profiling per-sample stream keys cross-CPU
+ * contention on sample->cpu (PERF_SAMPLE_CPU); without it the cpu
+ * field is the (u32)-1 'no CPU info' sentinel and same-instance
+ * reads and writes from different cores are indistinguishable
+ * from same-CPU traffic.
+ */
+ rec_argv[i++] = "--sample-cpu";
+
if (mem->phys_addr)
rec_argv[i++] = "--phys-data";
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events
2026-09-13 22:28 ` [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events Arnaldo Carvalho de Melo
@ 2026-09-14 1:31 ` Namhyung Kim
2026-09-14 22:17 ` Arnaldo Carvalho de Melo
0 siblings, 1 reply; 16+ messages in thread
From: Namhyung Kim @ 2026-09-14 1:31 UTC (permalink / raw)
To: Arnaldo Carvalho de Melo
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
On Sun, Sep 13, 2026 at 07:28:13PM -0300, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
>
> The test records with 'perf mem record' in per-thread mode and its only
> guard matches one specific message:
>
> perf mem record -o /dev/null -- true 2>&1 | \
> grep -q "failed: no PMU supports the memory events" && exit 2
>
> A PMU that has memory events but refuses them per-thread falls through
> it, and AMD IBS does:
>
> $ perf mem record -o /dev/null -- true
> Error:
> Failure to open event 'ibs_op/ldlat=0/u' on PMU 'ibs_op' which will be removed.
> Invalid event (ibs_op/ldlat=0/u) in per-thread mode, enable system wide with '-a'.
> Error:
> Failure to open any events for recording.
Maybe we need to add a fallback logic to add 'swfilt' term for ibs_op.
>
> What follows is not a skip either. The script runs under 'set -e', so
> the bare 'perf mem record' in test_basic_annotate() aborts it through the
> EXIT trap, which reports a signal that never happened and exits 1:
>
> Basic Rust perf annotate test
> Unexpected signal in test_basic_annotate
>
> so 'perf test' turns "this PMU cannot record these events" into a test
> failure:
>
> $ perf test "data type profiling"
> 87: perf data type profiling tests : FAILED!
>
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Acked-by: Namhyung Kim <namhyung@kernel.org>
Thanks,
Namhyung
> ---
> tools/perf/tests/shell/data_type_profiling.sh | 46 +++++++++++++++----
> 1 file changed, 36 insertions(+), 10 deletions(-)
>
> diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
> index eca694600a0478d2..a916c410274aa888 100755
> --- a/tools/perf/tests/shell/data_type_profiling.sh
> +++ b/tools/perf/tests/shell/data_type_profiling.sh
> @@ -19,6 +19,15 @@ perfout=$(mktemp /tmp/__perf_test.perf.out.XXXXX)
> perf mem record -o /dev/null -- true 2>&1 | \
> grep -q "failed: no PMU supports the memory events" && exit 2
>
> +# Skip if per-thread mem record is not supported on this PMU (e.g. AMD IBS
> +# needs system-wide '-a'): it is what the test records with below, and a
> +# failing record must not be reported as a test failure.
> +if ! perf mem record -o /dev/null -- true 2>/dev/null
> +then
> + echo "Skip: cannot record memory events on this PMU"
> + exit 2
> +fi
> +
> cleanup() {
> rm -rf "${perfdata}" "${perfout}"
> rm -rf "${perfdata}".old
> @@ -52,25 +61,42 @@ test_basic_annotate() {
> index=1 ;;
> esac
>
> + # Under 'set -e' a bare failing command aborts the script through the EXIT
> + # trap, so the commands that report a failure have to be the condition of
> + # an 'if' for that reporting to ever happen.
> if [ "x${mode}" == "xBasic" ]
> then
> - perf mem record -o "${perfdata}" ${testprogs[$index]} 2> /dev/null
> + if ! perf mem record -o "${perfdata}" ${testprogs[$index]} 2> /dev/null
> + then
> + echo "${mode} annotate [Failed: perf record]"
> + err=1
> + return
> + fi
> else
> - perf mem record -o - ${testprogs[$index]} 2> /dev/null > "${perfdata}"
> - fi
> - if [ "x$?" != "x0" ]
> - then
> - echo "${mode} annotate [Failed: perf record]"
> - err=1
> - return
> + if ! perf mem record -o - ${testprogs[$index]} 2> /dev/null > "${perfdata}"
> + then
> + echo "${mode} annotate [Failed: perf record]"
> + err=1
> + return
> + fi
> fi
>
> # Generate the annotated output file
> if [ "x${mode}" == "xBasic" ]
> then
> - perf annotate --code-with-type -i "${perfdata}" --stdio --percent-limit 1 2> /dev/null > "${perfout}"
> + if ! perf annotate --code-with-type -i "${perfdata}" --stdio --percent-limit 1 2> /dev/null > "${perfout}"
> + then
> + echo "${mode} annotate [Failed: perf annotate]"
> + err=1
> + return
> + fi
> else
> - perf annotate --code-with-type -i - --stdio 2> /dev/null --percent-limit 1 < "${perfdata}" > "${perfout}"
> + if ! perf annotate --code-with-type -i - --stdio 2> /dev/null --percent-limit 1 < "${perfdata}" > "${perfout}"
> + then
> + echo "${mode} annotate [Failed: perf annotate]"
> + err=1
> + return
> + fi
> fi
>
> # check if it has the target data type
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod
2026-09-13 22:28 ` [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod Arnaldo Carvalho de Melo
@ 2026-09-14 1:34 ` Namhyung Kim
2026-09-14 1:51 ` Arnaldo Carvalho de Melo
0 siblings, 1 reply; 16+ messages in thread
From: Namhyung Kim @ 2026-09-14 1:34 UTC (permalink / raw)
To: Arnaldo Carvalho de Melo
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
On Sun, Sep 13, 2026 at 07:28:14PM -0300, Arnaldo Carvalho de Melo wrote:
> From: Arnaldo Carvalho de Melo <acme@redhat.com>
>
> perf already uses debuginfod to fetch source files when annotating
> (via probe-finder.c) and 'perf probe' has open_from_debuginfod(),
> which queries debuginfo keyed by build ID when a module's debuginfo
> isn't found locally, largely the same thing this adds; eventually
> that one could be moved over to the new helper. For the analysis
> tools there was no way to obtain the debuginfo for a DSO in a
> profile when it isn't available locally under the name the DSO was
> opened with, for instance the vmlinux for the kernel a profile was
> recorded on when processing it on another machine, or after the
> kernel and its debuginfo package got upgraded in between.
>
> Add debuginfo__find_build_id(), that uses the debuginfod client to
> locate a debuginfo file keyed by the build ID, checking its local
> cache first and then querying the servers in DEBUGINFOD_URLS, and
> debuginfo__new_build_id(), that opens the DWARF in the file it finds.
>
> The debuginfod client fails when DEBUGINFOD_URLS isn't set even when
> what it wants is in its local cache, and the distro setup scripts that
> populate it from /etc/debuginfod don't reach cron jobs, systemd services
> and other environments that don't source the profile scripts, so also
> set it from the .urls files in /etc/debuginfod when not set.
>
> Querying servers, possibly third party ones, sends off-box the build
> IDs of the binaries being analysed and a fetch can take a while, so
> this is opt-out: on by default, off with --no-debuginfod, with
> core.debuginfod=false, per tool with report.debuginfod and
> top.debuginfod, and, since users that set buildid.dir to /dev/null
> (e.g. Linus) or otherwise turn the local build-id cache off clearly
> don't want fetched files stored on the box, off too in that case.
>
> When a fetch is in progress in a terminal, stdio, the way it prints
> progress is how one gets out of it: 's' aborts the current fetch via
> the debuginfod client's progress callback protocol and remembers the
> build ID, so that the rest of the session doesn't ask for it again,
> the user may have skipped it for being too big; 'd' additionally
> disables debuginfod for the rest of the session and points at 'perf
> config core.debuginfod=false' to make that permanent -- rewriting the
> user's ~/.perfconfig from a keypress would silently drop its comments
> -- and SIGINT/SIGTERM are intercepted while the terminal is in raw
> mode, so that it is restored and the signal is re-raised when the user
> interrupts a fetch.
>
> Make dso__debuginfo() use debuginfo__new_build_id() as a fallback,
> keyed by the build ID recorded in the perf.data file, so that
> consumers such as the data type profiler can resolve the types of
> DSOs whose debuginfo can be fetched this way. Do the fetch outside
> dso__lock and remember the build IDs that were a miss and the ones
> whose search the user cancelled, so that consumers revisiting a set
> of DSOs repeatedly, such as the data type profiler on every hist
> entry DSO switch, don't pay server round trips per attempt and a
> cancelled download, maybe a file the user found too big, isn't
> restarted by the next request for the same build ID in the same
> session.
>
> debuginfo__new_build_id() needs libdw to open the DWARF, so it lives
> in debuginfo.o, built only with CONFIG_LIBDW, while the libdebuginfod
> feature check is independent of NO_LIBDW; keep the new build ID
> prototypes under HAVE_LIBDW_SUPPORT too, with stubs otherwise, so
> that make NO_LIBDW=1 on a system that has the debuginfod client
> keeps linking.
>
> Assisted-by: LLM
> Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
> ---
[SNIP]
> diff --git a/tools/perf/builtin-annotate.c b/tools/perf/builtin-annotate.c
> index 4638e6fdc39bb6b7..d14ae7d1c345cb79 100644
> --- a/tools/perf/builtin-annotate.c
> +++ b/tools/perf/builtin-annotate.c
> @@ -733,6 +733,8 @@ int cmd_annotate(int argc, const char **argv)
> OPT_BOOLEAN(0, "stdio2", &annotate.use_stdio2, "Use the stdio interface"),
> OPT_BOOLEAN(0, "ignore-vmlinux", &symbol_conf.ignore_vmlinux,
> "don't load vmlinux even if found"),
> + OPT_BOOLEAN(0, "debuginfod", &symbol_conf.debuginfod,
> + "fetch debuginfo keyed by build ID from the debuginfod servers, on by default, use --no-debuginfod to turn off"),
I'm not sure what would be the good default. But with this, it can slow
down the process especially when the binary is not in the debuginfod.
Thanks,
Namhyung
> OPT_STRING('k', "vmlinux", &symbol_conf.vmlinux_name,
> "file", "vmlinux pathname"),
> OPT_BOOLEAN('m', "modules", &symbol_conf.use_modules,
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod
2026-09-14 1:34 ` Namhyung Kim
@ 2026-09-14 1:51 ` Arnaldo Carvalho de Melo
2026-09-14 20:25 ` Namhyung Kim
0 siblings, 1 reply; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-14 1:51 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
On Sun, Sep 13, 2026 at 06:34:44PM -0700, Namhyung Kim wrote:
> On Sun, Sep 13, 2026 at 07:28:14PM -0300, Arnaldo Carvalho de Melo wrote:
> > From: Arnaldo Carvalho de Melo <acme@redhat.com>
<SNIP>
> > Querying servers, possibly third party ones, sends off-box the build
> > IDs of the binaries being analysed and a fetch can take a while, so
> > this is opt-out: on by default, off with --no-debuginfod, with
> > core.debuginfod=false, per tool with report.debuginfod and
> > top.debuginfod, and, since users that set buildid.dir to /dev/null
> > (e.g. Linus) or otherwise turn the local build-id cache off clearly
> > don't want fetched files stored on the box, off too in that case.
> >
> > When a fetch is in progress in a terminal, stdio, the way it prints
> > progress is how one gets out of it: 's' aborts the current fetch via
> > the debuginfod client's progress callback protocol and remembers the
> > build ID, so that the rest of the session doesn't ask for it again,
> > the user may have skipped it for being too big; 'd' additionally
> > disables debuginfod for the rest of the session and points at 'perf
> > config core.debuginfod=false' to make that permanent -- rewriting the
> > user's ~/.perfconfig from a keypress would silently drop its comments
> > -- and SIGINT/SIGTERM are intercepted while the terminal is in raw
> > mode, so that it is restored and the signal is re-raised when the user
> > interrupts a fetch.
<SNIP>
> > +++ b/tools/perf/builtin-annotate.c
> > @@ -733,6 +733,8 @@ int cmd_annotate(int argc, const char **argv)
> > OPT_BOOLEAN(0, "stdio2", &annotate.use_stdio2, "Use the stdio interface"),
> > OPT_BOOLEAN(0, "ignore-vmlinux", &symbol_conf.ignore_vmlinux,
> > "don't load vmlinux even if found"),
> > + OPT_BOOLEAN(0, "debuginfod", &symbol_conf.debuginfod,
> > + "fetch debuginfo keyed by build ID from the debuginfod servers, on by default, use --no-debuginfod to turn off"),
>
> I'm not sure what would be the good default. But with this, it can slow
> down the process especially when the binary is not in the debuginfod.
That is why it allows the user to press 's' to skip it or 'd' to do a
one-time only disablement of this feature.
This is similar to gdb, that at session start asks if the debuginfo
files for the binary and its libraries should be downloaded, well, a bit
better because it allows the user to completely disable this at first
sight by pressing 'd'.
It also already honours configs that disable the ~/.debug cache.
- Arnaldo
> Thanks,
> Namhyung
>
>
> > OPT_STRING('k', "vmlinux", &symbol_conf.vmlinux_name,
> > "file", "vmlinux pathname"),
> > OPT_BOOLEAN('m', "modules", &symbol_conf.use_modules,
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod
2026-09-14 1:51 ` Arnaldo Carvalho de Melo
@ 2026-09-14 20:25 ` Namhyung Kim
2026-09-15 0:19 ` Arnaldo Carvalho de Melo
0 siblings, 1 reply; 16+ messages in thread
From: Namhyung Kim @ 2026-09-14 20:25 UTC (permalink / raw)
To: Arnaldo Carvalho de Melo
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
On Sun, Sep 13, 2026 at 10:51:07PM -0300, Arnaldo Carvalho de Melo wrote:
> On Sun, Sep 13, 2026 at 06:34:44PM -0700, Namhyung Kim wrote:
> > On Sun, Sep 13, 2026 at 07:28:14PM -0300, Arnaldo Carvalho de Melo wrote:
> > > From: Arnaldo Carvalho de Melo <acme@redhat.com>
>
> <SNIP>
>
> > > Querying servers, possibly third party ones, sends off-box the build
> > > IDs of the binaries being analysed and a fetch can take a while, so
> > > this is opt-out: on by default, off with --no-debuginfod, with
> > > core.debuginfod=false, per tool with report.debuginfod and
> > > top.debuginfod, and, since users that set buildid.dir to /dev/null
> > > (e.g. Linus) or otherwise turn the local build-id cache off clearly
> > > don't want fetched files stored on the box, off too in that case.
> > >
> > > When a fetch is in progress in a terminal, stdio, the way it prints
> > > progress is how one gets out of it: 's' aborts the current fetch via
> > > the debuginfod client's progress callback protocol and remembers the
> > > build ID, so that the rest of the session doesn't ask for it again,
> > > the user may have skipped it for being too big; 'd' additionally
> > > disables debuginfod for the rest of the session and points at 'perf
> > > config core.debuginfod=false' to make that permanent -- rewriting the
> > > user's ~/.perfconfig from a keypress would silently drop its comments
> > > -- and SIGINT/SIGTERM are intercepted while the terminal is in raw
> > > mode, so that it is restored and the signal is re-raised when the user
> > > interrupts a fetch.
>
> <SNIP>
>
> > > +++ b/tools/perf/builtin-annotate.c
> > > @@ -733,6 +733,8 @@ int cmd_annotate(int argc, const char **argv)
> > > OPT_BOOLEAN(0, "stdio2", &annotate.use_stdio2, "Use the stdio interface"),
> > > OPT_BOOLEAN(0, "ignore-vmlinux", &symbol_conf.ignore_vmlinux,
> > > "don't load vmlinux even if found"),
> > > + OPT_BOOLEAN(0, "debuginfod", &symbol_conf.debuginfod,
> > > + "fetch debuginfo keyed by build ID from the debuginfod servers, on by default, use --no-debuginfod to turn off"),
> >
> > I'm not sure what would be the good default. But with this, it can slow
> > down the process especially when the binary is not in the debuginfod.
>
> That is why it allows the user to press 's' to skip it or 'd' to do a
> one-time only disablement of this feature.
>
> This is similar to gdb, that at session start asks if the debuginfo
> files for the binary and its libraries should be downloaded, well, a bit
> better because it allows the user to completely disable this at first
> sight by pressing 'd'.
>
> It also already honours configs that disable the ~/.debug cache.
Oh.. I overlooked the details. But then it'd be nice to separate the
logic for the user interaction from the debuginfo fetching.
Thanks,
Namhyung
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events
2026-09-14 1:31 ` Namhyung Kim
@ 2026-09-14 22:17 ` Arnaldo Carvalho de Melo
0 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-14 22:17 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
On Sun, Sep 13, 2026 at 06:31:08PM -0700, Namhyung Kim wrote:
> On Sun, Sep 13, 2026 at 07:28:13PM -0300, Arnaldo Carvalho de Melo wrote:
> > From: Arnaldo Carvalho de Melo <acme@redhat.com>
> >
> > The test records with 'perf mem record' in per-thread mode and its only
> > guard matches one specific message:
> >
> > perf mem record -o /dev/null -- true 2>&1 | \
> > grep -q "failed: no PMU supports the memory events" && exit 2
> >
> > A PMU that has memory events but refuses them per-thread falls through
> > it, and AMD IBS does:
> >
> > $ perf mem record -o /dev/null -- true
> > Error:
> > Failure to open event 'ibs_op/ldlat=0/u' on PMU 'ibs_op' which will be removed.
> > Invalid event (ibs_op/ldlat=0/u) in per-thread mode, enable system wide with '-a'.
> > Error:
> > Failure to open any events for recording.
>
> Maybe we need to add a fallback logic to add 'swfilt' term for ibs_op.
Patch 9 in the latest series uses it, take a look.
- Arnaldo
> > What follows is not a skip either. The script runs under 'set -e', so
> > the bare 'perf mem record' in test_basic_annotate() aborts it through the
> > EXIT trap, which reports a signal that never happened and exits 1:
> >
> > Basic Rust perf annotate test
> > Unexpected signal in test_basic_annotate
> >
> > so 'perf test' turns "this PMU cannot record these events" into a test
> > failure:
> >
> > $ perf test "data type profiling"
> > 87: perf data type profiling tests : FAILED!
> >
> > Assisted-by: LLM
> > Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
>
> Acked-by: Namhyung Kim <namhyung@kernel.org>
>
> Thanks,
> Namhyung
>
> > ---
> > tools/perf/tests/shell/data_type_profiling.sh | 46 +++++++++++++++----
> > 1 file changed, 36 insertions(+), 10 deletions(-)
> >
> > diff --git a/tools/perf/tests/shell/data_type_profiling.sh b/tools/perf/tests/shell/data_type_profiling.sh
> > index eca694600a0478d2..a916c410274aa888 100755
> > --- a/tools/perf/tests/shell/data_type_profiling.sh
> > +++ b/tools/perf/tests/shell/data_type_profiling.sh
> > @@ -19,6 +19,15 @@ perfout=$(mktemp /tmp/__perf_test.perf.out.XXXXX)
> > perf mem record -o /dev/null -- true 2>&1 | \
> > grep -q "failed: no PMU supports the memory events" && exit 2
> >
> > +# Skip if per-thread mem record is not supported on this PMU (e.g. AMD IBS
> > +# needs system-wide '-a'): it is what the test records with below, and a
> > +# failing record must not be reported as a test failure.
> > +if ! perf mem record -o /dev/null -- true 2>/dev/null
> > +then
> > + echo "Skip: cannot record memory events on this PMU"
> > + exit 2
> > +fi
> > +
> > cleanup() {
> > rm -rf "${perfdata}" "${perfout}"
> > rm -rf "${perfdata}".old
> > @@ -52,25 +61,42 @@ test_basic_annotate() {
> > index=1 ;;
> > esac
> >
> > + # Under 'set -e' a bare failing command aborts the script through the EXIT
> > + # trap, so the commands that report a failure have to be the condition of
> > + # an 'if' for that reporting to ever happen.
> > if [ "x${mode}" == "xBasic" ]
> > then
> > - perf mem record -o "${perfdata}" ${testprogs[$index]} 2> /dev/null
> > + if ! perf mem record -o "${perfdata}" ${testprogs[$index]} 2> /dev/null
> > + then
> > + echo "${mode} annotate [Failed: perf record]"
> > + err=1
> > + return
> > + fi
> > else
> > - perf mem record -o - ${testprogs[$index]} 2> /dev/null > "${perfdata}"
> > - fi
> > - if [ "x$?" != "x0" ]
> > - then
> > - echo "${mode} annotate [Failed: perf record]"
> > - err=1
> > - return
> > + if ! perf mem record -o - ${testprogs[$index]} 2> /dev/null > "${perfdata}"
> > + then
> > + echo "${mode} annotate [Failed: perf record]"
> > + err=1
> > + return
> > + fi
> > fi
> >
> > # Generate the annotated output file
> > if [ "x${mode}" == "xBasic" ]
> > then
> > - perf annotate --code-with-type -i "${perfdata}" --stdio --percent-limit 1 2> /dev/null > "${perfout}"
> > + if ! perf annotate --code-with-type -i "${perfdata}" --stdio --percent-limit 1 2> /dev/null > "${perfout}"
> > + then
> > + echo "${mode} annotate [Failed: perf annotate]"
> > + err=1
> > + return
> > + fi
> > else
> > - perf annotate --code-with-type -i - --stdio 2> /dev/null --percent-limit 1 < "${perfdata}" > "${perfout}"
> > + if ! perf annotate --code-with-type -i - --stdio 2> /dev/null --percent-limit 1 < "${perfdata}" > "${perfout}"
> > + then
> > + echo "${mode} annotate [Failed: perf annotate]"
> > + err=1
> > + return
> > + fi
> > fi
> >
> > # check if it has the target data type
> > --
> > 2.55.0
> >
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod
2026-09-14 20:25 ` Namhyung Kim
@ 2026-09-15 0:19 ` Arnaldo Carvalho de Melo
0 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-15 0:19 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
On Mon, Sep 14, 2026 at 01:25:53PM -0700, Namhyung Kim wrote:
> On Sun, Sep 13, 2026 at 10:51:07PM -0300, Arnaldo Carvalho de Melo wrote:
> > On Sun, Sep 13, 2026 at 06:34:44PM -0700, Namhyung Kim wrote:
> > > On Sun, Sep 13, 2026 at 07:28:14PM -0300, Arnaldo Carvalho de Melo wrote:
> > > > +++ b/tools/perf/builtin-annotate.c
> > > > @@ -733,6 +733,8 @@ int cmd_annotate(int argc, const char **argv)
> > > > OPT_BOOLEAN(0, "stdio2", &annotate.use_stdio2, "Use the stdio interface"),
> > > > OPT_BOOLEAN(0, "ignore-vmlinux", &symbol_conf.ignore_vmlinux,
> > > > "don't load vmlinux even if found"),
> > > > + OPT_BOOLEAN(0, "debuginfod", &symbol_conf.debuginfod,
> > > > + "fetch debuginfo keyed by build ID from the debuginfod servers, on by default, use --no-debuginfod to turn off"),
> > > I'm not sure what would be the good default. But with this, it can slow
> > > down the process especially when the binary is not in the debuginfod.
> > That is why it allows the user to press 's' to skip it or 'd' to do a
> > one-time only disablement of this feature.
> > This is similar to gdb, that at session start asks if the debuginfo
> > files for the binary and its libraries should be downloaded, well, a bit
> > better because it allows the user to completely disable this at first
> > sight by pressing 'd'.
> > It also already honours configs that disable the ~/.debug cache.
> Oh.. I overlooked the details. But then it'd be nice to separate the
> logic for the user interaction from the debuginfo fetching.
I will do that tomorrow, as well as have it not just in --stdio, but
also in the TUI.
I need to be more concise and granular, sorry about that.
- Arnaldo
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 7/8] perf annotate-data: Resolve type DIEs in the debug file they came from
2026-09-14 1:35 [PATCH v4 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
@ 2026-09-14 1:36 ` Arnaldo Carvalho de Melo
0 siblings, 0 replies; 16+ messages in thread
From: Arnaldo Carvalho de Melo @ 2026-09-14 1:36 UTC (permalink / raw)
To: Namhyung Kim
Cc: Ingo Molnar, Thomas Gleixner, James Clark, Jiri Olsa, Ian Rogers,
Adrian Hunter, Clark Williams, linux-kernel, linux-perf-users,
Arnaldo Carvalho de Melo
From: Arnaldo Carvalho de Melo <acme@redhat.com>
A 'perf report -s type' on a 783 MB AMD IBS data type profiling session
hangs, burning all of a CPU and producing no output:
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -s type -i perf.data.ibs > /dev/null
Processing events... [100.0%] 746M / 746M
Merging related events... [ 58.0%] 293828 / 506686
Merging related events... [ 59.0%] 298894 / 506686
Merging related events... [ 60.0%] 303960 / 506686
Merging related events... [ 61.0%] 309026 / 506686
[ ... nothing else, ever ... ]
It is a spin and not a slow path: with perf-stuck the CPU time used
grows by a whole interval on every sample, while the [stack] mapping,
which would be moving down if this was recursion, stays put:
⬢ [acme@toolbx perf-tools-next]$ tools/perf/scripts/perf-stuck.sh -i 15 -n 4 -l ibs.log $(pgrep -x perf)
watching 2297313 (perf report --progress -s type -i perf.data.ibs) every 15s
09:39:34 state=R cpu=+0 (0.00s) rss=664295kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB Merging related events... [ 39.0%] 197574 / 506686
09:39:49 state=R cpu=+1496 (14.96s) rss=665383kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=1 Merging related events... [ 39.0%] 197574 / 506686
09:40:04 state=R cpu=+1497 (14.97s) rss=665383kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=2 Merging related events... [ 39.0%] 197574 / 506686
09:40:19 state=R cpu=+1497 (14.97s) rss=502250kB stack=7ffc3ae31000-7ffc3ae52000 size=132kB stuck=0 Merging related events... [ 61.0%] 309026 / 506686
Attaching gdb and stepping the loop, printing dwarf_dieoffset() and
dwarf_tag() for the DIE being chased on each trip round it, shows the
chase never moving, on a typedef that refers to itself:
⬢ [acme@toolbx perf-tools-next]$ gdb -p $(pgrep -x perf) -batch -x tools/perf/scripts/perf-stuck.gdb -ex 'perf-die-chain die_get_pointer_type type_die 8' -ex perf-dso -ex detach
stuck in die_get_pointer_type():
#3 0x0000000000772421 in die_get_pointer_type (type_die=0x7ffc3ae4b710, type_die@entry=0x7ffc3ae4b6f0, die_mem=die_mem@entry=0x7ffc3ae4b710) at util/dwarf-aux.c:327
327 type_die = die_get_type(type_die, die_mem);
chain[0] die=0x7ffc3ae4b710 addr=0x7ff6811591ef cu=0x44ef0878 off=0x1f tag=22 name=(null)
chain[1] die=0x7ffc3ae4b710 addr=0x7ff6811591ef cu=0x44ef0878 off=0x1f tag=22 name=(null)
[ ... the very same DIE, forever ... ]
dso=/usr/lib64/libz.so.1.3.1.zlib-ng ip=0xe2e sym=build_tree
The DIE is at offset 0x1f of the debug info of libz.so.1, which is
zlib-ng, and is one of the dwz compressed ones: the type DIEs shared by
more than one CU live in the common file, where 0x1f is a perfectly good
DW_TAG_base_type:
⬢ [acme@toolbx perf-tools-next]$ readelf --debug-dump=info /usr/lib/debug/.dwz/zlib-ng-2.3.3-3.fc44.x86_64 | sed -n '/Compilation Unit @ offset 0:/,/Compilation Unit @ offset 0x5f:/p'
Compilation Unit @ offset 0:
Length: 0x5b (32-bit)
Version: 5
Unit Type: DW_UT_partial (3)
<0><c>: Abbrev Number: 1 (DW_TAG_partial_unit)
[...]
<1><1f>: Abbrev Number: 62 (DW_TAG_base_type)
<20> DW_AT_byte_size : 4
<21> DW_AT_encoding : 5 (signed)
<22> DW_AT_name : int
while in the main debug file that same offset is not a DIE at all, it is
the start of another unit's header:
⬢ [acme@toolbx perf-tools-next]$ readelf --debug-dump=info /usr/lib/debug/usr/lib64/libz.so.1.3.1.zlib-ng-2.3.3-3.fc44.x86_64.debug | grep 'Compilation Unit @' | head -2
Compilation Unit @ offset 0:
Compilation Unit @ offset 0x1f:
die_collect_vars() saves the dwarf_dieoffset() of the type DIE, which is
relative to the file that DIE lives in, and update_var_state() then hands
that offset to dwarf_offdie() together with the main debug file, where
dwarf_offdie() parses whatever is there as a DIE, in this case the one
typedef that refers to itself, which is what makes the "follow the
typedefs and qualifiers until a pointer or an array type" loop in
die_get_pointer_type() spin.
So record, next to the offset, whether the type DIE was in the file the
variable DIE came from or in the dwz common one, and resolve the offset
in the file it came from, which is what the new die_get_type_die() does.
Which of the two that is does not have to be guessed from what is at the
offset: dwz encodes the references into its common file as
DW_FORM_GNU_ref_alt, so elfutils resolves them into the alt Dwarf and
the CU of the resulting type DIE belongs to that other file, so
comparing the Dwarf each of the two CUs belongs to, die_same_file(),
settles it exactly, for any number of hops from the variable DIE.
Then bound the chases themselves: no sane chain of typedefs and
qualifiers is 32 DIEs long, no sane nesting for the struct and union
members that __add_member_cb() follows recursively is 8 deep, and the
same recursion bound covers the type names die_get_typename_from_type()
builds by following pointers and arrays, so that a debug info file
broken in some other way makes perf give up on a type, telling about it
with pr_debug, visible with -v, instead of looking like it hung. The
member nesting bound is reported with pr_debug_dtp (visible with -vvv or
-D type-profile).
Cycles can't occur in member trees from valid DWARF, embedded members
can't be recursive in C, so the nesting bound only ever bites legitimate
depth and the member where the recursion was cut is marked 'truncated',
which the JSON exporter added in a subsequent series reports to its
consumers, so that they can tell a truncated tree from one that really
ends there.
die_get_type_die() then has no fallback to the other file: resolving an
alt file offset in the main file is the misparse above, so it resolves
in the file the offset was recorded as belonging to, and gives up on the
type when the offset does not resolve there. The tag the type DIE had is
kept as a sanity check, a mismatch means the debug info changed under
perf or is broken in yet another way, and the bounds on the type chasers
remain the backstop: a debug info file broken in some other way still
makes perf give up on a type with a pr_debug instead of hanging.
Testing:
Before, on the 783 MB AMD IBS session, killed after some 8 minutes stuck
at 61%:
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -s type -i perf.data.ibs > /dev/null
Processing events... [100.0%] 746M / 746M
Merging related events... [ 61.0%] 309026 / 506686
⬢ [acme@toolbx perf-tools-next]$
After, the whole session is processed, and the zlib-ng types, from the
build_tree() hist entry that used to hang it, show up:
⬢ [acme@toolbx perf-tools-next]$ perf report --progress -s type -i perf.data.ibs > ibs.out
Processing events... [100.0%] 746M / 746M
Merging related events... [100.0%] 506686 / 506686
Sorting events for output... [100.0%] 1428 / 1428
⬢ [acme@toolbx perf-tools-next]$ grep -E "deflate_state|inflate_state|internal_state" ibs.out
0.00% deflate_state
0.00% deflate_state*
0.00% struct inflate_state
0.00% struct inflate_state*
0.00% struct internal_state
⬢ [acme@toolbx perf-tools-next]$ perf test 17 27 30 31 85 88
17: Match and link multiple hists : Ok
27: Filter hist entries : Ok
30: Sort output of hist entries : Ok
31: Cumulate child hist entries : Ok
85: Test that perf report includes file offsets and event type names in diagnostic messages. : Ok
88: Test that perf report handles truncated perf.data gracefully (no crash, no segfault — clean error exit).: Skip
⬢ [acme@toolbx perf-tools-next]$
Requiring elfutils 0.160 for this: dwarf_cu_getdwarf(), the function
that tells which Dwarf a CU belongs to, first appeared in 0.160 ("libdw:
New functions dwarf_cu_getdwarf, dwarf_cu_die", elfutils NEWS), so the
libdw feature test now probes for it, in tools/build/feature/test-libdw.c,
and Makefile.config says 0.160 where it said 0.157. The probe takes the
address instead of calling it, as that is all that is needed to make a
0.157-0.159 elfutils, from 2014 and without the symbol, disable dwarf
support with the existing message rather than fail to link dwarf-aux.c.
Fixes: 06b2ce75386df04b ("perf annotate-data: Maintain variable type info")
Fixes: 55ee3d005d62279d ("perf annotate-data: Add a cache for global variable types")
Assisted-by: LLM
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/build/feature/test-libdw.c | 15 ++-
tools/perf/Makefile.config | 2 +-
tools/perf/util/annotate-data.c | 59 ++++++++---
tools/perf/util/annotate-data.h | 3 +
tools/perf/util/dwarf-aux.c | 171 +++++++++++++++++++++++++++----
tools/perf/util/dwarf-aux.h | 33 ++++++
6 files changed, 247 insertions(+), 36 deletions(-)
diff --git a/tools/build/feature/test-libdw.c b/tools/build/feature/test-libdw.c
index aabd63ca76b4d7e6..23e1ba6ff3466b9f 100644
--- a/tools/build/feature/test-libdw.c
+++ b/tools/build/feature/test-libdw.c
@@ -49,8 +49,21 @@ int test_elfutils(void)
return 0;
}
+/*
+ * elfutils 0.160 and later: used to tell which debug file a DIE lives in,
+ * the dwz alt file or the main one, see die_same_file() in
+ * tools/perf/util/dwarf-aux.c. Only the symbol is needed, so take its
+ * address instead of calling it.
+ */
+int test_libdw_cu_getdwarf(void)
+{
+ void *sym = (void *)dwarf_cu_getdwarf;
+
+ return sym == NULL;
+}
+
int main(void)
{
return test_libdw() + test_libdw_unwind() + test_libdw_getlocations() +
- test_libdw_getcfi() + test_elfutils();
+ test_libdw_getcfi() + test_libdw_cu_getdwarf() + test_elfutils();
}
diff --git a/tools/perf/Makefile.config b/tools/perf/Makefile.config
index 4d5993da9f94579f..fa78f50db60179f3 100644
--- a/tools/perf/Makefile.config
+++ b/tools/perf/Makefile.config
@@ -470,7 +470,7 @@ else
else
ifneq ($(feature-libdw), 1)
ifndef NO_LIBDW
- $(warning No libdw.h found or old libdw.h found or elfutils is older than 0.157, disables dwarf support. Please install new elfutils-devel/libdw-dev)
+ $(warning No libdw.h found or old libdw.h found or elfutils is older than 0.160, disables dwarf support. Please install new elfutils-devel/libdw-dev)
NO_LIBDW := 1
endif
endif # Dwarf support
diff --git a/tools/perf/util/annotate-data.c b/tools/perf/util/annotate-data.c
index 4e4c587640823c81..06d27868887bdadf 100644
--- a/tools/perf/util/annotate-data.c
+++ b/tools/perf/util/annotate-data.c
@@ -221,6 +221,15 @@ static bool data_type_less(struct rb_node *node_a, const struct rb_node *node_b)
return strcmp(a->self.type_name, b->self.type_name) < 0;
}
+/*
+ * Members of struct/union members are added recursively, and the same DIE
+ * that is not what it looks like, the one that makes the type chasers in
+ * util/dwarf-aux.c spin, can make a member's type point back at one of its
+ * own ancestors, recursing until the stack is gone. Nothing usable comes
+ * out of nesting members this deep anyway.
+ */
+#define MAX_MEMBER_DEPTH 8
+
/* Recursively add new members for struct/union */
static int __add_member_cb(Dwarf_Die *die, void *arg)
{
@@ -235,6 +244,16 @@ static int __add_member_cb(Dwarf_Die *die, void *arg)
if (dwarf_tag(die) != DW_TAG_member)
return DIE_FIND_CB_SIBLING;
+ if (__die_get_real_type(die, &member_type) == NULL)
+ return DIE_FIND_CB_SIBLING;
+
+ if (dwarf_tag(&member_type) == DW_TAG_typedef) {
+ if (die_get_real_type(&member_type, &die_mem) == NULL)
+ return DIE_FIND_CB_SIBLING;
+ } else {
+ die_mem = member_type;
+ }
+
member = zalloc(sizeof(*member));
if (member == NULL)
return DIE_FIND_CB_END;
@@ -242,12 +261,6 @@ static int __add_member_cb(Dwarf_Die *die, void *arg)
strbuf_init(&sb, 32);
die_get_typename(die, &sb);
- __die_get_real_type(die, &member_type);
- if (dwarf_tag(&member_type) == DW_TAG_typedef)
- die_get_real_type(&member_type, &die_mem);
- else
- die_mem = member_type;
-
if (dwarf_aggregate_size(&die_mem, &size) < 0)
size = 0;
@@ -289,10 +302,23 @@ static int __add_member_cb(Dwarf_Die *die, void *arg)
}
member->size = size;
member->offset = loc + parent->offset;
+ member->depth = parent->depth + 1;
INIT_LIST_HEAD(&member->children);
list_add_tail(&member->node, &parent->children);
tag = dwarf_tag(&die_mem);
+ if (member->depth >= MAX_MEMBER_DEPTH) {
+ /*
+ * The JSON exporter added in a later series reports
+ * this to its consumers, so that they can tell a
+ * truncated tree from one that really ends here.
+ */
+ member->truncated = true;
+ pr_debug_dtp("member nesting limit reached at %s\n",
+ member->type_name ?: "(unknown type)");
+ return DIE_FIND_CB_SIBLING;
+ }
+
switch (tag) {
case DW_TAG_structure_type:
case DW_TAG_union_type:
@@ -645,6 +671,8 @@ struct global_var_entry {
u64 start;
u64 end;
u64 die_offset;
+ int die_tag;
+ bool from_alt; /* die_offset is relative to the alt (dwz) file */
};
static int global_var_cmp(const void *_key, const struct rb_node *node)
@@ -682,7 +710,7 @@ static struct global_var_entry *global_var__find(struct data_loc_info *dloc, u64
}
static bool global_var__add(struct data_loc_info *dloc, u64 addr,
- const char *name, Dwarf_Die *type_die)
+ const char *name, Dwarf_Die *type_die, bool from_alt)
{
struct dso *dso = map__dso(dloc->ms->map);
struct global_var_entry *gvar;
@@ -704,6 +732,8 @@ static bool global_var__add(struct data_loc_info *dloc, u64 addr,
gvar->start = addr;
gvar->end = addr + size;
gvar->die_offset = dwarf_dieoffset(type_die);
+ gvar->die_tag = dwarf_tag(type_die);
+ gvar->from_alt = from_alt;
rb_add(&gvar->node, dso__global_vars(dso), global_var_less);
return true;
@@ -778,12 +808,14 @@ static void global_var__collect(struct data_loc_info *dloc)
if (pos->reg != -1)
continue;
- if (!dwarf_offdie(dwarf, pos->die_off, &type_die))
+ if (!die_get_type_die(dwarf, pos->die_off, pos->die_tag,
+ pos->from_alt, &type_die))
continue;
get_global_var_info(dloc, pos->addr, &var_name, &var_offset);
- global_var__add(dloc, pos->addr, var_name, &type_die);
+ global_var__add(dloc, pos->addr, var_name, &type_die,
+ pos->from_alt);
}
delete_var_types(var_types);
@@ -808,7 +840,8 @@ bool get_global_var_type(Dwarf_Die *cu_die, struct data_loc_info *dloc,
gvar = global_var__find(dloc, var_addr);
if (gvar) {
- if (!dwarf_offdie(dloc->di->dbg, gvar->die_offset, type_die))
+ if (!die_get_type_die(dloc->di->dbg, gvar->die_offset,
+ gvar->die_tag, gvar->from_alt, type_die))
return false;
*var_offset = var_addr - gvar->start;
@@ -838,7 +871,8 @@ bool get_global_var_type(Dwarf_Die *cu_die, struct data_loc_info *dloc,
ok:
/* The address should point to the start of the variable */
- global_var__add(dloc, var_addr - *var_offset, var_name, type_die);
+ global_var__add(dloc, var_addr - *var_offset, var_name, type_die,
+ !die_same_file(cu_die, type_die));
return true;
}
@@ -893,7 +927,8 @@ static void update_var_state(struct type_state *state, struct data_loc_info *dlo
continue;
}
/* Get the type DIE using the offset */
- if (!dwarf_offdie(dloc->di->dbg, var->die_off, &mem_die))
+ if (!die_get_type_die(dloc->di->dbg, var->die_off,
+ var->die_tag, var->from_alt, &mem_die))
continue;
if (var->reg == DWARF_REG_FB || var->reg == fbreg || var->reg == state->stack_reg) {
diff --git a/tools/perf/util/annotate-data.h b/tools/perf/util/annotate-data.h
index c26130744260955f..d85866e83fbda69a 100644
--- a/tools/perf/util/annotate-data.h
+++ b/tools/perf/util/annotate-data.h
@@ -57,6 +57,9 @@ struct annotated_member {
char *var_name;
int offset;
int size;
+ unsigned int depth;
+ /* Children not expanded because the nesting limit was reached */
+ bool truncated;
};
/**
diff --git a/tools/perf/util/dwarf-aux.c b/tools/perf/util/dwarf-aux.c
index d7160f87ac7d7ab3..7acb431fd34a8ecb 100644
--- a/tools/perf/util/dwarf-aux.c
+++ b/tools/perf/util/dwarf-aux.c
@@ -266,16 +266,35 @@ Dwarf_Die *die_get_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
return NULL;
}
+/*
+ * The chases below cross typedefs and qualifiers to get to the type that
+ * is actually meant, and a DIE that is not what it looks like, e.g. one
+ * parsed at an offset that is not the start of a DIE in the file it was
+ * resolved in, can have a DW_AT_type that refers back to itself, which
+ * makes them spin forever: 'perf report -s type' did exactly that on the
+ * dwz compressed debug info of zlib-ng (libz.so.1), burning all of a CPU
+ * with no output while resolving a hist entry in build_tree().
+ *
+ * No sane chain is this long, so give up instead of hanging, telling about
+ * it so that the broken debug info can be looked at.
+ */
+#define MAX_TYPE_CHASE 32
+
/* Get a type die, but skip qualifiers */
Dwarf_Die *__die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
{
- int tag;
+ int tag, chase = 0;
do {
vr_die = die_get_type(vr_die, die_mem);
if (!vr_die)
- break;
+ return NULL;
tag = dwarf_tag(vr_die);
+ if (++chase > MAX_TYPE_CHASE) {
+ pr_debug("DWARF: qualifier chase limit reached at DIE 0x%lx\n",
+ (unsigned long)dwarf_dieoffset(vr_die));
+ return NULL;
+ }
} while (tag == DW_TAG_const_type ||
tag == DW_TAG_restrict_type ||
tag == DW_TAG_volatile_type ||
@@ -296,8 +315,15 @@ Dwarf_Die *__die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
*/
Dwarf_Die *die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
{
+ int chase = 0;
+
do {
vr_die = __die_get_real_type(vr_die, die_mem);
+ if (++chase > MAX_TYPE_CHASE) {
+ pr_debug("DWARF: typedef chase limit reached at DIE 0x%lx\n",
+ vr_die ? (unsigned long)dwarf_dieoffset(vr_die) : 0);
+ return NULL;
+ }
} while (vr_die && dwarf_tag(vr_die) == DW_TAG_typedef);
return vr_die;
@@ -314,7 +340,7 @@ Dwarf_Die *die_get_real_type(Dwarf_Die *vr_die, Dwarf_Die *die_mem)
*/
Dwarf_Die *die_get_pointer_type(Dwarf_Die *type_die, Dwarf_Die *die_mem)
{
- int tag;
+ int tag, chase = 0;
do {
tag = dwarf_tag(type_die);
@@ -324,6 +350,11 @@ Dwarf_Die *die_get_pointer_type(Dwarf_Die *type_die, Dwarf_Die *die_mem)
tag != DW_TAG_restrict_type && tag != DW_TAG_volatile_type &&
tag != DW_TAG_shared_type)
return NULL;
+ if (++chase > MAX_TYPE_CHASE) {
+ pr_debug("DWARF: pointer type chase limit reached at DIE 0x%lx\n",
+ (unsigned long)dwarf_dieoffset(type_die));
+ return NULL;
+ }
type_die = die_get_type(type_die, die_mem);
} while (type_die);
@@ -1118,17 +1149,27 @@ Dwarf_Die *die_find_member(Dwarf_Die *st_die, const char *name,
die_mem);
}
-/**
- * die_get_typename_from_type - Get the name of given type DIE
- * @type_die: a type DIE
- * @buf: a strbuf for result type name
- *
- * Get the name of @type_die and stores it to @buf. Return 0 if succeeded.
- * and Return -ENOENT if failed to find type name.
- * Note that the result will stores typedef name if possible, and stores
- * "*(function_type)" if the type is a function pointer.
+/*
+ * The name of a pointer or array type is built from the name of the type
+ * it points to or holds, so the recursion below follows DW_AT_type; a
+ * garbage DIE whose DW_AT_type refers back to itself makes it recurse
+ * forever, just like the chases above, so it gets the same bound.
*/
-int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
+static int __die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf,
+ int depth);
+
+static int __die_get_typename(Dwarf_Die *vr_die, struct strbuf *buf, int depth)
+{
+ Dwarf_Die type;
+
+ if (__die_get_real_type(vr_die, &type) == NULL)
+ return -ENOENT;
+
+ return __die_get_typename_from_type(&type, buf, depth);
+}
+
+static int __die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf,
+ int depth)
{
int tag, ret;
const char *tmp = "";
@@ -1155,7 +1196,12 @@ int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
/* Write a base name */
return strbuf_addf(buf, "%s%s", tmp, name ?: "");
}
- ret = die_get_typename(type_die, buf);
+ if (depth >= MAX_TYPE_CHASE) {
+ pr_debug("DWARF: type name recursion limit reached at DIE 0x%lx\n",
+ (unsigned long)dwarf_dieoffset(type_die));
+ return -ENOENT;
+ }
+ ret = __die_get_typename(type_die, buf, depth + 1);
if (ret < 0) {
/* void pointer has no type attribute */
if (tag == DW_TAG_pointer_type && ret == -ENOENT)
@@ -1166,6 +1212,21 @@ int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
return strbuf_addstr(buf, tmp);
}
+/**
+ * die_get_typename_from_type - Get the name of given type DIE
+ * @type_die: a type DIE
+ * @buf: a strbuf for result type name
+ *
+ * Get the name of @type_die and stores it to @buf. Return 0 if succeeded.
+ * and Return -ENOENT if failed to find type name.
+ * Note that the result will stores typedef name if possible, and stores
+ * "*(function_type)" if the type is a function pointer.
+ */
+int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
+{
+ return __die_get_typename_from_type(type_die, buf, 0);
+}
+
/**
* die_get_typename - Get the name of given variable DIE
* @vr_die: a variable DIE
@@ -1178,12 +1239,7 @@ int die_get_typename_from_type(Dwarf_Die *type_die, struct strbuf *buf)
*/
int die_get_typename(Dwarf_Die *vr_die, struct strbuf *buf)
{
- Dwarf_Die type;
-
- if (__die_get_real_type(vr_die, &type) == NULL)
- return -ENOENT;
-
- return die_get_typename_from_type(&type, buf);
+ return __die_get_typename(vr_die, buf, 0);
}
/**
@@ -1632,6 +1688,22 @@ Dwarf_Die *die_find_variable_by_addr(Dwarf_Die *sc_die, Dwarf_Addr addr,
return result;
}
+/*
+ * Whether two DIEs live in the same debug file.
+ *
+ * dwarf_dieoffset() is relative to the file the DIE is in, so this is what
+ * tells an offset that has to be resolved in the dwz alt file, where dwz
+ * moved the type, from one that belongs to the main file: dwz encodes the
+ * references into its common file as DW_FORM_GNU_ref_alt, elfutils resolves
+ * them into the alt Dwarf and the CU of the resulting DIE belongs to that
+ * other file, so comparing the Dwarf each CU belongs to settles it exactly,
+ * rather than inferring it from what happens to be at the offset.
+ */
+bool die_same_file(Dwarf_Die *die_a, Dwarf_Die *die_b)
+{
+ return dwarf_cu_getdwarf(die_a->cu) == dwarf_cu_getdwarf(die_b->cu);
+}
+
static int __die_collect_vars_cb(Dwarf_Die *die_mem, void *arg)
{
struct die_var_type **var_types = arg;
@@ -1676,6 +1748,8 @@ static int __die_collect_vars_cb(Dwarf_Die *die_mem, void *arg)
vt->is_reg_var_addr = true;
vt->die_off = dwarf_dieoffset(&type_die);
+ vt->die_tag = dwarf_tag(&type_die);
+ vt->from_alt = !die_same_file(die_mem, &type_die);
vt->addr = start;
vt->end = end;
vt->has_range = (end != 0 || start != 0);
@@ -1695,7 +1769,8 @@ static int __die_collect_vars_cb(Dwarf_Die *die_mem, void *arg)
*
* Save all variables and parameters in the @sc_die and save them to @var_types.
* The @var_types is a singly-linked list containing type and location info.
- * Actual type can be retrieved using dwarf_offdie() with 'die_off' later.
+ * Actual type can be retrieved using die_get_type_die() with 'die_off',
+ * 'die_tag' and 'from_alt' later.
*
* Callers should free @var_types.
*/
@@ -1741,6 +1816,8 @@ static int __die_collect_global_vars_cb(Dwarf_Die *die_mem, void *arg)
return DIE_FIND_CB_END;
vt->die_off = dwarf_dieoffset(&type_die);
+ vt->die_tag = dwarf_tag(&type_die);
+ vt->from_alt = !die_same_file(die_mem, &type_die);
vt->addr = ops->number;
vt->end = 0;
vt->has_range = false;
@@ -1752,6 +1829,55 @@ static int __die_collect_global_vars_cb(Dwarf_Die *die_mem, void *arg)
return DIE_FIND_CB_SIBLING;
}
+/**
+ * die_get_type_die - Get a type DIE saved by die_collect_vars()
+ * @dbg: the main debug info
+ * @die_off: offset of the type DIE, from dwarf_dieoffset()
+ * @die_tag: tag that DIE had when the offset was saved
+ * @from_alt: whether the type DIE is in the dwz alt file
+ * @die_mem: where to store the resulting DIE
+ *
+ * See the comment in util/dwarf-aux.h: the offset is only meaningful in the
+ * file the DIE was in, which can be the dwz common file, so resolve it in
+ * the file @from_alt says it was in, the main file or its alt file, and use
+ * the DIE only when it has the @die_tag it had when the offset was saved.
+ * There is deliberately no fallback to the other file: resolving an alt
+ * file offset in the main file does not fail, it parses whatever is there
+ * as a DIE, and that is what hung 'perf report -s type'.
+ */
+Dwarf_Die *die_get_type_die(Dwarf *dbg, u64 die_off, int die_tag, bool from_alt,
+ Dwarf_Die *die_mem)
+{
+ Dwarf *target = dbg;
+ Dwarf_Die die;
+
+ if (from_alt) {
+ /*
+ * Deliberately no fallback to the main file when there is
+ * no alt file, or when the offset does not resolve in it:
+ * resolving an alt file offset in the main file does not
+ * fail, it parses whatever is there as a DIE, and that is
+ * what hung 'perf report -s type', see the comment in
+ * util/dwarf-aux.h.
+ */
+ target = dwarf_getalt(dbg);
+ if (target == NULL) {
+ pr_debug("DWARF: no alt (dwz) debug file to resolve the type DIE at offset 0x%lx in\n",
+ (unsigned long)die_off);
+ return NULL;
+ }
+ }
+
+ if (dwarf_offdie(target, die_off, &die) && dwarf_tag(&die) == die_tag) {
+ *die_mem = die;
+ return die_mem;
+ }
+
+ pr_debug("DWARF: no DIE with tag %d at offset 0x%lx in the %s debug file\n",
+ die_tag, (unsigned long)die_off, from_alt ? "alt" : "main");
+ return NULL;
+}
+
/**
* die_collect_global_vars - Save all global variables
* @cu_die: a CU DIE
@@ -1759,7 +1885,8 @@ static int __die_collect_global_vars_cb(Dwarf_Die *die_mem, void *arg)
*
* Save all global variables in the @cu_die and save them to @var_types.
* The @var_types is a singly-linked list containing type and location info.
- * Actual type can be retrieved using dwarf_offdie() with 'die_off' later.
+ * Actual type can be retrieved using die_get_type_die() with 'die_off',
+ * 'die_tag' and 'from_alt' later.
*
* Callers should free @var_types.
*/
diff --git a/tools/perf/util/dwarf-aux.h b/tools/perf/util/dwarf-aux.h
index 161f0bf980b6ee6a..149e0cbf63ccc82a 100644
--- a/tools/perf/util/dwarf-aux.h
+++ b/tools/perf/util/dwarf-aux.h
@@ -152,6 +152,8 @@ int die_get_scopes(Dwarf_Die *cu_die, Dwarf_Addr pc, Dwarf_Die **scopes);
struct die_var_type {
struct die_var_type *next;
u64 die_off;
+ int die_tag;
+ bool from_alt; /* die_off is relative to the alt (dwz) file */
u64 addr;
u64 end; /* end address of location range */
int reg;
@@ -183,6 +185,37 @@ Dwarf_Die *die_find_variable_by_addr(Dwarf_Die *sc_die, Dwarf_Addr addr,
/* Save all variables and parameters in this scope */
void die_collect_vars(Dwarf_Die *sc_die, struct die_var_type **var_types);
+/*
+ * Get the type DIE saved by die_collect_vars()/die_collect_global_vars().
+ *
+ * The offsets those save are the dwarf_dieoffset() of the type DIE, which is
+ * relative to the debug file that DIE lives in: the dwz common file, the alt
+ * file in libdw terms, for the types shared by more than one CU, the main
+ * file for the rest. Resolving an alt file offset in the main file does not
+ * fail: dwarf_offdie() parses whatever is at that offset there, and an offset
+ * that is a CU header in the main file reads back as a typedef whose
+ * DW_AT_type refers to itself, which is what hung 'perf report -s type' on
+ * the dwz compressed debug info of zlib-ng (libz.so.1).
+ *
+ * So @from_alt, recorded when the offset was saved, says which of the two
+ * files to resolve it in. It is not inferred from the DIE contents: dwz
+ * encodes references into the common file as DW_FORM_GNU_ref_alt, so
+ * elfutils resolves them into the alt Dwarf and the CU of the type DIE then
+ * belongs to that other file, which die_same_file() compares exactly.
+ *
+ * There is deliberately no fallback to the other file when the offset does
+ * not resolve: that fallback is the misparse above.
+ *
+ * @die_tag is then only a sanity check: the offset is of a DIE that had this
+ * tag when it was saved, so a mismatch means the debug info changed under us,
+ * or is broken, and giving up on the type is the right answer.
+ */
+Dwarf_Die *die_get_type_die(Dwarf *dbg, u64 die_off, int die_tag, bool from_alt,
+ Dwarf_Die *die_mem);
+
+/* Whether two DIEs live in the same debug file */
+bool die_same_file(Dwarf_Die *die_a, Dwarf_Die *die_b);
+
/* Save all global variables in this CU */
void die_collect_global_vars(Dwarf_Die *cu_die, struct die_var_type **var_types);
--
2.55.0
^ permalink raw reply [flat|nested] 16+ messages in thread
end of thread, other threads:[~2026-09-15 0:19 UTC | newest]
Thread overview: 16+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-13 22:28 [PATCH v3 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 1/8] perf test: Skip data_type_profiling when the PMU cannot record memory events Arnaldo Carvalho de Melo
2026-09-14 1:31 ` Namhyung Kim
2026-09-14 22:17 ` Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 2/8] perf debuginfo: Fetch debuginfo keyed by build ID using debuginfod Arnaldo Carvalho de Melo
2026-09-14 1:34 ` Namhyung Kim
2026-09-14 1:51 ` Arnaldo Carvalho de Melo
2026-09-14 20:25 ` Namhyung Kim
2026-09-15 0:19 ` Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 3/8] perf symbol: Fall back to fetching the vmlinux by build ID Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 4/8] perf annotate-data: Show the sample count in the data-type browser Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 5/8] perf report: Add --progress option Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 6/8] perf scripts: Add perf-stuck, to tell where a running perf is stuck Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 7/8] perf annotate-data: Resolve type DIEs in the debug file they came from Arnaldo Carvalho de Melo
2026-09-13 22:28 ` [PATCH 8/8] perf mem record: Request PERF_SAMPLE_CPU by default Arnaldo Carvalho de Melo
2026-09-14 1:35 [PATCH v4 0/8] perf tools: Annotate fixes, stdio progress indication, debuginfo-client in more places Arnaldo Carvalho de Melo
2026-09-14 1:36 ` [PATCH 7/8] perf annotate-data: Resolve type DIEs in the debug file they came from Arnaldo Carvalho de Melo
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®