From: Namhyung Kim <namhyung@kernel.org>
To: haghdoost@uber.com
Cc: Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
Arnaldo Carvalho de Melo <acme@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Alexander Shishkin <alexander.shishkin@linux.intel.com>,
Jiri Olsa <jolsa@kernel.org>, Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
James Clark <james.clark@linaro.org>,
Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andriin@fb.com>,
linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v5 4/5] perf script: Add --lazy-load-symbols for lazy symbol loading
Date: Wed, 7 Oct 2026 15:15:50 -0700 [thread overview]
Message-ID: <asbElicX1u_GZu_j@google.com> (raw)
In-Reply-To: <20261005-perf-symbol-memory-send-v5-4-165dceb2b049@uber.com>
On Mon, Oct 05, 2026 at 10:30:34AM -0700, Alireza Haghdoost via B4 Relay wrote:
> From: Alireza Haghdoost <haghdoost@uber.com>
>
> perf script eagerly materializes eligible symbols from every DSO
> encountered in samples and keeps them in an rb-tree until exit.
> Most of those symbols are never sampled. On a large profile this
> turns symbol loading into the main memory cost of perf script, and in a
> memory-constrained cgroup into an OOM kill.
>
> This patch adds --lazy-load-symbols for userspace ELF DSOs. It builds a
> compact address-sorted index and materializes ordinary struct symbols on
> demand. Names are read and demangled from the exact ELF source through
> the existing DSO data cache; resolved symbols are then inserted into the
> DSO's existing rb-tree.
>
> struct symbol is unchanged. Each index entry holds only the address range,
> binding, type and the ELF string-table offset of the name. When a sample
> hits an entry, perf reads and demangles the name and creates a normal
> struct symbol with the name embedded, so code outside the loader never
> sees an unresolved name.
I thought we agreed to just use struct symbol and lazy-load the name
only. With Ian's patch to remove the 'priv' part of symbol and
hopefully RB-tree node as well, the size of symbol would become
manageable, I guess.
Thanks,
Namhyung
>
> On a 120-second cgroup profile of a production database service (54k
> samples across 11 DSOs), lazy mode reduces peak RssAnon from 314 MiB to
> 80 MiB and wall time from 4.49 to 3.75 seconds, with identical output.
> Memory optimizations usually cost time; this one does not because lazy
> loading skips many unnecessary calloc() calls and demangling operations.
>
> Lazy loading is most effective when samples reference only a small
> fraction of the available symbols, such as profiles spanning many large
> DSOs. It still builds an index proportional to the total symbol count.
> Eager loading remains available for dense symbol coverage or cases
> requiring its broader ELF and architecture support.
>
> This does not claim full parity with the eager loader. Lazy loading
> supports the common userspace ELF symtab/dynsym case; .gnu_debugdata and
> PPC64 .opd continue through the eager loader.
>
> To match eager loading, perf first completes zero-sized symbol ranges
> and selects among same-address aliases in .symtab. It then adds .dynsym
> and repeats those steps on the combined index, preserving symbols found
> only in .dynsym. A .dynsym entry that duplicates a .symtab entry is
> omitted because eager duplicate selection would discard it; this
> prevents binaries that export most of .symtab through .dynsym from
> nearly doubling the index during construction.
>
> Lazy loading uses the same duplicate and IFUNC selection as eager
> loading and clips ranges that cross .plt before synthesizing PLT
> symbols. PLT synthesis itself is unchanged. For nested ranges, lookups
> return the innermost symbol containing the address.
>
> Lazy lookups hold the DSO lock while searching the index and inserting
> symbols, but release it before reading names because the DSO data cache
> takes its global lock before DSO locks. Reads from one index are
> serialized because cache-page lookups are otherwise lockless. Before
> building the existing name-sorted symbol array, perf materializes the
> entire lazy index and detaches it from the DSO. A detached index remains
> alive until its last reader finishes, ensuring that no symbols are added
> after the name-sorted array is built.
>
> Signed-off-by: Alireza Haghdoost <haghdoost@uber.com>
> ---
> tools/perf/Documentation/perf-script.txt | 9 +
> tools/perf/builtin-script.c | 2 +
> tools/perf/util/dso.c | 24 +
> tools/perf/util/dso.h | 80 +++
> tools/perf/util/map.c | 5 +-
> tools/perf/util/symbol-elf.c | 920 +++++++++++++++++++++++++++++++
> tools/perf/util/symbol-minimal.c | 9 +
> tools/perf/util/symbol.c | 18 +-
> tools/perf/util/symbol.h | 1 +
> tools/perf/util/symbol_conf.h | 1 +
> 10 files changed, 1067 insertions(+), 2 deletions(-)
>
> diff --git a/tools/perf/Documentation/perf-script.txt b/tools/perf/Documentation/perf-script.txt
> index f0228e784ced..7ef4642803da 100644
> --- a/tools/perf/Documentation/perf-script.txt
> +++ b/tools/perf/Documentation/perf-script.txt
> @@ -343,6 +343,15 @@ include::itrace.txt[]
>
> Default: 127
>
> +--lazy-load-symbols::
> + Resolve symbols lazily instead of eagerly loading the full
> + symbol table of every DSO that appears in a sample. This sharply
> + reduces memory (and usually time) for profiles of large binaries
> + where only a small fraction of the symbol table is referenced. This
> + applies only to userspace ELF DSOs; kernel DSOs, modules, PPC64 DSOs
> + with an .opd section and DSOs whose symbols come from .gnu_debugdata
> + always load eagerly.
> +
> --ns::
> Use 9 decimal places when displaying time (i.e. show the nanoseconds)
>
> diff --git a/tools/perf/builtin-script.c b/tools/perf/builtin-script.c
> index b6691ebb1b4e..404f5a080d45 100644
> --- a/tools/perf/builtin-script.c
> +++ b/tools/perf/builtin-script.c
> @@ -4287,6 +4287,8 @@ int cmd_script(int argc, const char **argv)
> "Set the maximum stack depth when parsing the callchain, "
> "anything beyond the specified depth will be ignored. "
> "Default: kernel.perf_event_max_stack or " __stringify(PERF_MAX_STACK_DEPTH)),
> + OPT_BOOLEAN(0, "lazy-load-symbols", &symbol_conf.lazy_load_symbols,
> + "Resolve symbols lazily instead of loading full symtabs"),
> OPT_BOOLEAN(0, "reltime", &reltime, "Show time stamps relative to start"),
> OPT_BOOLEAN(0, "deltatime", &deltatime, "Show time stamps relative to previous event"),
> OPT_BOOLEAN('I', "show-info", &show_full_info,
> diff --git a/tools/perf/util/dso.c b/tools/perf/util/dso.c
> index 5c4872810ada..d6d2d2ce4157 100644
> --- a/tools/perf/util/dso.c
> +++ b/tools/perf/util/dso.c
> @@ -1721,6 +1721,29 @@ void dso__set_sorted_by_name(struct dso *dso)
> RC_CHK_ACCESS(dso)->sorted_by_name = true;
> }
>
> +struct dso_ondemand *dso_ondemand__new(void)
> +{
> + struct dso_ondemand *od = zalloc(sizeof(*od));
> +
> + if (od)
> + mutex_init(&od->read_lock);
> + return od;
> +}
> +
> +void dso_ondemand__free(struct dso_ondemand *od)
> +{
> + if (!od)
> + return;
> + mutex_destroy(&od->read_lock);
> + free(od->sorted);
> + free(od->outer);
> + if (od->data_dso) {
> + dso__data_close(od->data_dso);
> + dso__put(od->data_dso);
> + }
> + free(od);
> +}
> +
> struct dso *dso__new_id(const char *name, const struct dso_id *id)
> {
> RC_STRUCT(dso) *dso = zalloc(sizeof(*dso) + strlen(name) + 1);
> @@ -1803,6 +1826,7 @@ void dso__delete(struct dso *dso)
>
> dso__data_close(dso);
> auxtrace_cache__free(RC_CHK_ACCESS(dso)->auxtrace_cache);
> + dso_ondemand__free(RC_CHK_ACCESS(dso)->ondemand);
> dso_cache__free(dso);
> zfree(&RC_CHK_ACCESS(dso)->data.path);
> dso__free_a2l(dso);
> diff --git a/tools/perf/util/dso.h b/tools/perf/util/dso.h
> index 7bcd5ec0c312..2d25d43af6e7 100644
> --- a/tools/perf/util/dso.h
> +++ b/tools/perf/util/dso.h
> @@ -283,6 +283,64 @@ struct dso_bpf_prog {
> struct perf_env *env;
> };
>
> +struct sym_idx {
> + u64 start;
> + u64 end;
> + u32 name_off;
> + u8 binding;
> + u8 type;
> + u8 flags;
> +};
> +
> +#define SYM_IDX_FLAG_IFUNC_ALIAS (1 << 0)
> +#define SYM_IDX_FLAG_MATERIALIZED (1 << 1)
> +#define SYM_IDX_FLAG_DYNSTR (1 << 2)
> +
> +#define SYM_IDX_NONE UINT32_MAX
> +
> +struct sym_idx_strtab {
> + u64 offset;
> + u64 size;
> +};
> +
> +/**
> + * struct dso_ondemand - per-DSO state for lazy symbol loading.
> + *
> + * Holds the compact address index and the data-cache handle used to read
> + * symbol names. It stays attached to the parent DSO until name sorting
> + * materializes every symbol and detaches it, or until the DSO is deleted.
> + * The parent DSO lock protects the attachment, the index and @nr_readers.
> + * To read a name, a lookup increments @nr_readers, drops the DSO lock and
> + * reads under @read_lock. It then retakes the DSO lock to insert the symbol
> + * and decrement @nr_readers. The last reader frees a detached index.
> + */
> +struct dso_ondemand {
> + /** @data_dso: Data-cache handle for the exact ELF symbol source. */
> + struct dso *data_dso;
> + /**
> + * @strtab: .strtab and .dynstr file offsets and sizes, indexed by
> + * whether SYM_IDX_FLAG_DYNSTR is set.
> + */
> + struct sym_idx_strtab strtab[2];
> + /** @sorted: Compact address-sorted lazy index. */
> + struct sym_idx *sorted;
> + /**
> + * @outer: Links used to resolve nested ranges, allocated only when
> + * some ranges overlap. For each entry, the nearest earlier entry that
> + * ends after it, or SYM_IDX_NONE.
> + */
> + u32 *outer;
> + /** @nr_sorted: Number of entries in @sorted. */
> + u32 nr_sorted;
> + /** @nr_readers: Lookups reading a name without the DSO lock. */
> + u32 nr_readers;
> + /**
> + * @read_lock: Serializes data-cache page and symbol-name reads. The
> + * data cache looks up pages without a lock.
> + */
> + struct mutex read_lock;
> +};
> +
> struct auxtrace_cache;
>
> DECLARE_RC_STRUCT(dso) {
> @@ -314,6 +372,7 @@ DECLARE_RC_STRUCT(dso) {
> char *symsrc_filename;
> struct nsinfo *nsinfo;
> struct auxtrace_cache *auxtrace_cache;
> + struct dso_ondemand *ondemand;
> union { /* Tool specific area */
> void *priv;
> u64 db_id;
> @@ -473,6 +532,16 @@ static inline void dso__set_auxtrace_cache(struct dso *dso, struct auxtrace_cach
> RC_CHK_ACCESS(dso)->auxtrace_cache = cache;
> }
>
> +static inline struct dso_ondemand *dso__ondemand(struct dso *dso)
> +{
> + return RC_CHK_ACCESS(dso)->ondemand;
> +}
> +
> +static inline void dso__set_ondemand(struct dso *dso, struct dso_ondemand *od)
> +{
> + RC_CHK_ACCESS(dso)->ondemand = od;
> +}
> +
> static inline struct dso_bpf_prog *dso__bpf_prog(struct dso *dso)
> {
> return &RC_CHK_ACCESS(dso)->bpf_prog;
> @@ -849,6 +918,17 @@ char *dso__get_filename(struct dso *dso, const char *root_dir, bool *decomp,
> void dso__put_filename(struct dso *dso, char *filename, bool decomp);
> bool is_kernel_module(const char *pathname, int cpumode);
> bool dso__needs_decompress(struct dso *dso);
> +struct symbol *dso__find_symbol_ondemand(struct dso *dso, u64 addr)
> + LOCKS_EXCLUDED(dso__lock(dso));
> +void dso__materialize_symbols_ondemand(struct dso *dso)
> + LOCKS_EXCLUDED(dso__lock(dso));
> +const char *dso__read_ondemand_symbol_name(struct dso *data_dso,
> + u64 strtab_offset, u64 strtab_size,
> + u64 name_off, char *buf,
> + size_t buflen, char **to_free,
> + unsigned int *nr_reads);
> +struct dso_ondemand *dso_ondemand__new(void);
> +void dso_ondemand__free(struct dso_ondemand *od);
> int dso__decompress_kmodule_fd(struct dso *dso, const char *name);
> int dso__decompress_kmodule_path(struct dso *dso, const char *name,
> char *pathname, size_t len);
> diff --git a/tools/perf/util/map.c b/tools/perf/util/map.c
> index 41cdddc987ee..66626b7f2669 100644
> --- a/tools/perf/util/map.c
> +++ b/tools/perf/util/map.c
> @@ -382,10 +382,13 @@ int map__load(struct map *map)
>
> struct symbol *map__find_symbol(struct map *map, u64 addr)
> {
> + struct dso *dso;
> +
> if (map__load(map) < 0)
> return NULL;
>
> - return dso__find_symbol(map__dso(map), addr);
> + dso = map__dso(map);
> + return dso__find_symbol(dso, addr);
> }
>
> struct symbol *map__find_symbol_by_name_idx(struct map *map, const char *name, size_t *idx)
> diff --git a/tools/perf/util/symbol-elf.c b/tools/perf/util/symbol-elf.c
> index e955c3feddcd..3649b07242f0 100644
> --- a/tools/perf/util/symbol-elf.c
> +++ b/tools/perf/util/symbol-elf.c
> @@ -2,6 +2,7 @@
> #include <fcntl.h>
> #include <stdio.h>
> #include <errno.h>
> +#include <stdint.h>
> #include <stdlib.h>
> #include <string.h>
> #include <unistd.h>
> @@ -12,6 +13,7 @@
> #include "libbfd.h"
> #include "map.h"
> #include "maps.h"
> +#include "namespaces.h"
> #include "symbol.h"
> #include "symsrc.h"
> #include "machine.h"
> @@ -593,6 +595,8 @@ static int dso__synthesize_plt_got_symbols(struct dso *dso, Elf *elf,
> return err;
> }
>
> +static void dso__clip_ondemand_symbols_at(struct dso *dso, u64 addr);
> +
> /*
> * We need to check if we have a .dynsym, so that we can handle the
> * .plt, synthesizing its symbols, that aren't on the symtabs (be it
> @@ -623,6 +627,13 @@ int dso__synthesize_plt_symbols(struct dso *dso, struct symsrc *ss)
> if (!elf_section_by_name(elf, &ehdr, &shdr_plt, ".plt", NULL))
> return 0;
>
> + /*
> + * Zero-sized or oversized ELF symbols can have been extended across
> + * .plt. Clip the index first so lookups cannot attribute PLT addresses
> + * to a preceding symbol before the synthesized PLT symbols are added.
> + */
> + dso__clip_ondemand_symbols_at(dso, shdr_plt.sh_offset);
> +
> /*
> * A symbol from a previous section (e.g. .init) can have been expanded
> * by symbols__fixup_end() to overlap .plt. Truncate it before adding
> @@ -1515,6 +1526,893 @@ static int dso__process_kernel_symbol(struct dso *dso, struct map *map,
> return 0;
> }
>
> +static int cmp_sym_idx(const void *a, const void *b)
> +{
> + const struct sym_idx *sa = a, *sb = b;
> +
> + if (sa->start != sb->start)
> + return sa->start < sb->start ? -1 : 1;
> + /*
> + * qsort is not stable. While sorting, name_off holds the entry's
> + * position, preserving eager's insertion order for equal-start
> + * aliases. sym_idx__sort() restores the name offsets afterwards.
> + */
> + if (sa->name_off != sb->name_off)
> + return sa->name_off < sb->name_off ? -1 : 1;
> + return 0;
> +}
> +
> +static int sym_idx__sort(struct sym_idx *entries, u32 nr)
> +{
> + u32 *name_offs, i;
> +
> + if (!nr)
> + return 0;
> + name_offs = malloc(nr * sizeof(*name_offs));
> + if (!name_offs)
> + return -ENOMEM;
> + for (i = 0; i < nr; i++) {
> + name_offs[i] = entries[i].name_off;
> + entries[i].name_off = i;
> + }
> + qsort(entries, nr, sizeof(*entries), cmp_sym_idx);
> + for (i = 0; i < nr; i++)
> + entries[i].name_off = name_offs[entries[i].name_off];
> + free(name_offs);
> + return 0;
> +}
> +
> +enum sym_idx_bound {
> + SYM_IDX_BOUND_LOWER,
> + SYM_IDX_BOUND_UPPER,
> +};
> +
> +/*
> + * Return the position of the first entry whose start is at or after @addr
> + * (LOWER), or after @addr (UPPER), or @nr if there is none. Unlike bsearch(),
> + * this gives a position when no entry starts at @addr. A range comparator
> + * would not help either: for nested ranges, the inner and the outer entry
> + * both compare equal.
> + */
> +static u32 sym_idx__bound(const struct sym_idx *sorted, u32 nr, u64 addr,
> + enum sym_idx_bound kind)
> +{
> + u32 lo = 0, hi = nr;
> +
> + while (lo < hi) {
> + u32 mid = lo + (hi - lo) / 2;
> +
> + if ((kind == SYM_IDX_BOUND_LOWER) ? sorted[mid].start < addr :
> + sorted[mid].start <= addr)
> + lo = mid + 1;
> + else
> + hi = mid;
> + }
> + return lo;
> +}
> +
> +/*
> + * Return the innermost entry containing @addr, i.e. the last-starting one.
> + * Only the last entry starting at or before @addr can contain it unless
> + * ranges overlap. An earlier entry containing @addr ends after every entry
> + * that does not, so following the outer links, each to the nearest earlier
> + * entry that ends later, reaches the innermost one first.
> + */
> +static u32 sym_idx__find(const struct dso_ondemand *od, u64 addr)
> +{
> + u32 pos = sym_idx__bound(od->sorted, od->nr_sorted, addr,
> + SYM_IDX_BOUND_UPPER);
> +
> + if (!pos)
> + return SYM_IDX_NONE;
> + pos--;
> + while (pos != SYM_IDX_NONE && od->sorted[pos].end <= addr)
> + pos = od->outer ? od->outer[pos] : SYM_IDX_NONE;
> + return pos;
> +}
> +
> +/*
> + * Only allocating @od->outer can fail. A call after clipping never allocates,
> + * as clipping cannot add overlaps.
> + */
> +static int dso_ondemand__link_overlaps(struct dso_ondemand *od)
> +{
> + u32 i, j;
> +
> + if (!od->outer) {
> + for (i = 0; i + 1 < od->nr_sorted; i++) {
> + if (od->sorted[i].end > od->sorted[i + 1].start)
> + break;
> + }
> + if (i + 1 >= od->nr_sorted)
> + return 0;
> + od->outer = malloc(od->nr_sorted * sizeof(*od->outer));
> + if (!od->outer)
> + return -ENOMEM;
> + }
> + for (i = 0; i < od->nr_sorted; i++) {
> + j = i ? i - 1 : SYM_IDX_NONE;
> + while (j != SYM_IDX_NONE && od->sorted[j].end <= od->sorted[i].end)
> + j = od->outer[j];
> + od->outer[i] = j;
> + }
> + return 0;
> +}
> +
> +static struct symbol *symbols__find_start(struct rb_root_cached *symbols, u64 start)
> +{
> + struct rb_node *n = symbols->rb_root.rb_node;
> +
> + while (n) {
> + struct symbol *s = rb_entry(n, struct symbol, rb_node);
> +
> + if (start < s->start)
> + n = n->rb_left;
> + else if (start > s->start)
> + n = n->rb_right;
> + else
> + return s;
> + }
> + return NULL;
> +}
> +
> +static void dso__clip_ondemand_symbols_at(struct dso *dso, u64 addr)
> +{
> + struct dso_ondemand *od = dso__ondemand(dso);
> + u32 lo, i;
> +
> + if (!od)
> + return;
> +
> + lo = sym_idx__bound(od->sorted, od->nr_sorted, addr, SYM_IDX_BOUND_LOWER);
> + for (i = 0; i < lo; i++) {
> + struct sym_idx *idx = &od->sorted[i];
> + struct symbol *sym;
> +
> + if (idx->end <= addr)
> + continue;
> + idx->end = addr;
> + if (!(idx->flags & SYM_IDX_FLAG_MATERIALIZED))
> + continue;
> + sym = symbols__find_start(dso__symbols(dso), idx->start);
> + if (sym && sym->end > addr)
> + sym->end = addr;
> + }
> + if (od->outer)
> + dso_ondemand__link_overlaps(od);
> +}
> +
> +static bool ondemand_sym_ok(Elf *elf, Elf_Data *secstrs,
> + const GElf_Sym *sym, u32 sh_link,
> + uint16_t e_machine)
> +{
> + Elf_Scn *sym_sec;
> + GElf_Shdr sym_shdr;
> + int is_label = elf_sym__is_label(sym);
> + const char *name;
> +
> + if (!is_label && !elf_sym__filter((GElf_Sym *)sym))
> + return false;
> +
> + if (sym->st_shndx == SHN_ABS)
> + return false;
> +
> + sym_sec = elf_getscn(elf, sym->st_shndx);
> + if (!sym_sec)
> + return false;
> + if (!gelf_getshdr(sym_sec, &sym_shdr))
> + return false;
> + if (!(sym_shdr.sh_flags & SHF_ALLOC))
> + return false;
> +
> + if (is_label && (!secstrs || !elf_sec__filter(&sym_shdr, secstrs)))
> + return false;
> +
> + name = elf_strptr(elf, sh_link, sym->st_name);
> + if (!name)
> + return false;
> +
> + /*
> + * Reject ARM/AArch64/RISC-V "mapping symbols" ($a/$d/$t/$x), as
> + * the eager loop does. They are zero-size STT_NOTYPE labels in
> + * allocated sections that would otherwise be indexed and fill
> + * forward over real functions, misattributing everything after
> + * them.
> + */
> + if (e_machine == EM_ARM || e_machine == EM_AARCH64) {
> + if (name[0] == '$' && strchr("adtx", name[1]) &&
> + (name[2] == '\0' || name[2] == '.'))
> + return false;
> + }
> + if (e_machine == EM_RISCV) {
> + if (name[0] == '$' && strchr("dx", name[1]))
> + return false;
> + }
> +
> + return true;
> +}
> +
> +static const char *sym_idx__elf_name(Elf *elf, const size_t *strndx,
> + const struct sym_idx *idx)
> +{
> + return elf_strptr(elf, strndx[!!(idx->flags & SYM_IDX_FLAG_DYNSTR)],
> + idx->name_off);
> +}
> +
> +static void sym_idx__candidate(const struct sym_idx *idx, const char *name,
> + struct symbol_candidate *c)
> +{
> + c->size = idx->end - idx->start;
> + c->name = name;
> + c->type = idx->type;
> + c->binding = idx->binding;
> +}
> +
> +/*
> + * Keep one entry per start address, choosing among aliases and marking IFUNC
> + * aliases with the same policy as symbols__fixup_duplicate().
> + */
> +static u32 sym_idx__dedup_aliases(struct dso *dso, Elf *elf,
> + const size_t *strndx,
> + struct sym_idx *sorted, u32 count)
> +{
> + u32 i, j, k, out = 0;
> +
> + for (i = 0; i < count; i = j) {
> + struct symbol_candidate best, cand;
> + char *best_demangled = NULL, *demangled;
> + const char *name;
> + u32 best_idx = i;
> + bool ifunc_alias;
> +
> + j = i + 1;
> + while (j < count && sorted[j].start == sorted[i].start)
> + j++;
> + if (j == i + 1) {
> + sorted[out++] = sorted[i];
> + continue;
> + }
> +
> + name = sym_idx__elf_name(elf, strndx, &sorted[i]);
> + if (name) {
> + best_demangled = dso__demangle_sym(dso, 0, name);
> + if (best_demangled)
> + name = best_demangled;
> + }
> + sym_idx__candidate(&sorted[i], name, &best);
> + ifunc_alias = sorted[i].flags & SYM_IDX_FLAG_IFUNC_ALIAS;
> +
> + for (k = i + 1; k < j; k++) {
> + name = sym_idx__elf_name(elf, strndx, &sorted[k]);
> + if (!best.name || !name)
> + continue;
> +
> + demangled = dso__demangle_sym(dso, 0, name);
> + if (demangled)
> + name = demangled;
> + sym_idx__candidate(&sorted[k], name, &cand);
> +
> + if (symbol__choose_best(&best, &cand) == SYMBOL_B) {
> + ifunc_alias = (sorted[k].flags & SYM_IDX_FLAG_IFUNC_ALIAS) ||
> + best.type == STT_GNU_IFUNC;
> + free(best_demangled);
> + best_demangled = demangled;
> + best = cand;
> + best_idx = k;
> + } else {
> + ifunc_alias |= cand.type == STT_GNU_IFUNC;
> + free(demangled);
> + }
> + }
> + free(best_demangled);
> +
> + sorted[out] = sorted[best_idx];
> + sorted[out].flags &= ~SYM_IDX_FLAG_IFUNC_ALIAS;
> + if (ifunc_alias)
> + sorted[out].flags |= SYM_IDX_FLAG_IFUNC_ALIAS;
> + out++;
> + }
> + return out;
> +}
> +
> +/*
> + * Whether .dynsym entry @idx is a copy of the .symtab entry at its start, so
> + * that symbols__fixup_duplicate() keeps the .symtab entry. Dropping copies
> + * here keeps a .dynsym that exports most of .symtab from doubling the
> + * index while it is built. The "SyS" check excludes the names on which
> + * arch__choose_best_symbol() does not keep the first of two equal symbols.
> + */
> +static bool sym_idx__dynsym_copy(Elf *elf, const size_t *strndx,
> + const struct sym_idx *symtab, u32 nr,
> + const struct sym_idx *idx)
> +{
> + u32 pos = sym_idx__bound(symtab, nr, idx->start, SYM_IDX_BOUND_LOWER);
> + const struct sym_idx *orig = &symtab[pos];
> + const char *name, *orig_name;
> +
> + if (symbol_conf.allow_aliases || pos == nr ||
> + orig->start != idx->start || orig->end != idx->end ||
> + idx->end == idx->start || orig->type != idx->type ||
> + orig->binding != idx->binding)
> + return false;
> + name = sym_idx__elf_name(elf, strndx, idx);
> + orig_name = sym_idx__elf_name(elf, strndx, orig);
> + return name && orig_name && !strcmp(name, orig_name) &&
> + !strstr(name, "SyS");
> +}
> +
> +/*
> + * Fill @idx from @sym if the eager loader would load it. @symtab holds the
> + * fixed up .symtab entries when loading .dynsym.
> + */
> +static bool sym_idx__from_sym(struct symsrc *syms_ss,
> + struct symsrc *runtime_ss, Elf_Data *secstrs,
> + bool dynsym, const struct sym_idx *symtab,
> + u32 nr, const size_t *strndx,
> + const GElf_Sym *sym, struct sym_idx *idx)
> +{
> + Elf *elf = syms_ss->elf;
> + GElf_Shdr shdr = dynsym ? syms_ss->dynshdr : syms_ss->symshdr;
> + u16 e_machine = syms_ss->ehdr.e_machine;
> + u64 adjusted = sym->st_value;
> + GElf_Phdr phdr;
> +
> + if (!ondemand_sym_ok(elf, secstrs, sym, shdr.sh_link, e_machine))
> + return false;
> +
> + /*
> + * On ARM, symbols for thumb functions have 1 added to the symbol
> + * address as a flag - remove it.
> + */
> + if (e_machine == EM_ARM && GELF_ST_TYPE(sym->st_info) == STT_FUNC &&
> + (adjusted & 1))
> + --adjusted;
> +
> + if (elf_read_program_header(runtime_ss->elf, adjusted, &phdr) == 0) {
> + adjusted -= phdr.p_vaddr - phdr.p_offset;
> + } else {
> + Elf_Scn *sym_sec = elf_getscn(elf, sym->st_shndx);
> + GElf_Shdr sym_shdr;
> +
> + if (sym_sec && gelf_getshdr(sym_sec, &sym_shdr)) {
> + /*
> + * A NOBITS section in a debuginfo file has an invalid
> + * sh_offset; use the runtime section.
> + */
> + if (sym_shdr.sh_type == SHT_NOBITS) {
> + sym_sec = elf_getscn(runtime_ss->elf,
> + sym->st_shndx);
> + if (!sym_sec || !gelf_getshdr(sym_sec, &sym_shdr))
> + return false;
> + }
> + adjusted -= sym_shdr.sh_addr - sym_shdr.sh_offset;
> + }
> + }
> +
> + *idx = (struct sym_idx) {
> + .start = adjusted,
> + .end = adjusted + sym->st_size,
> + .name_off = sym->st_name,
> + .binding = GELF_ST_BIND(sym->st_info),
> + .type = GELF_ST_TYPE(sym->st_info),
> + .flags = dynsym ? SYM_IDX_FLAG_DYNSTR : 0,
> + };
> + return !dynsym || !sym_idx__dynsym_copy(elf, strndx, symtab, nr, idx);
> +}
> +
> +/* Append eligible symbols; record the string table at index @dynsym. */
> +static int sym_idx__add_table(struct symsrc *syms_ss, struct symsrc *runtime_ss,
> + bool dynsym, struct sym_idx **entries, u32 *nr,
> + struct sym_idx_strtab *strtab, size_t *strndx)
> +{
> + Elf *elf = syms_ss->elf;
> + GElf_Ehdr *ehdr = &syms_ss->ehdr;
> + GElf_Shdr shdr = dynsym ? syms_ss->dynshdr : syms_ss->symshdr;
> + GElf_Shdr strshdr;
> + Elf_Scn *strscn, *sec_strndx;
> + Elf_Data *syms, *secstrs = NULL;
> + struct sym_idx *tmp, idx;
> + size_t i, bytes;
> + u64 nr_entries;
> + u32 count = 0, j;
> + GElf_Sym sym;
> +
> + syms = elf_getdata(dynsym ? syms_ss->dynsym : syms_ss->symtab, NULL);
> + if (!syms)
> + return -1;
> +
> + if (!shdr.sh_entsize)
> + return 0;
> +
> + nr_entries = shdr.sh_size / shdr.sh_entsize;
> + if (nr_entries > UINT32_MAX)
> + return -EOVERFLOW;
> +
> + strscn = elf_getscn(elf, shdr.sh_link);
> + if (!strscn || !gelf_getshdr(strscn, &strshdr))
> + return -1;
> + strndx[dynsym] = shdr.sh_link;
> +
> + sec_strndx = elf_getscn(elf, ehdr->e_shstrndx);
> + if (sec_strndx)
> + secstrs = elf_getdata(sec_strndx, NULL);
> +
> + for (i = 0; i < nr_entries; i++) {
> + if (gelf_getsym(syms, i, &sym) &&
> + sym_idx__from_sym(syms_ss, runtime_ss, secstrs, dynsym,
> + *entries, *nr, strndx, &sym, &idx))
> + count++;
> + }
> +
> + if (!count)
> + return 0;
> + if (count > UINT32_MAX - *nr ||
> + check_mul_overflow((size_t)(*nr + count), sizeof(**entries), &bytes))
> + return -EOVERFLOW;
> + tmp = realloc(*entries, bytes);
> + if (!tmp)
> + return -1;
> + *entries = tmp;
> +
> + /*
> + * The file may change between the two passes; never fill more entries
> + * than the first pass counted.
> + */
> + j = *nr;
> + for (i = 0; i < nr_entries && j < *nr + count; i++) {
> + if (!gelf_getsym(syms, i, &sym) ||
> + !sym_idx__from_sym(syms_ss, runtime_ss, secstrs, dynsym,
> + *entries, *nr, strndx, &sym, &idx))
> + continue;
> + tmp[j++] = idx;
> + }
> + *nr = j;
> + strtab[dynsym].offset = strshdr.sh_offset;
> + strtab[dynsym].size = strshdr.sh_size;
> + return 0;
> +}
> +
> +/*
> + * Match symbols__fixup_end() and then symbols__fixup_duplicate(), which the
> + * eager loader runs after inserting a table's symbols.
> + */
> +static int sym_idx__fixup(struct dso *dso, Elf *elf, const size_t *strndx,
> + struct sym_idx *entries, u32 *nr)
> +{
> + int err = sym_idx__sort(entries, *nr);
> + u32 i;
> +
> + if (err)
> + return err;
> + for (i = 0; i < *nr; i++) {
> + if (entries[i].end != entries[i].start)
> + continue;
> + if (i + 1 < *nr)
> + entries[i].end = entries[i + 1].start;
> + else
> + entries[i].end = roundup(entries[i].start, 4096) + 4096;
> + }
> + if (!symbol_conf.allow_aliases)
> + *nr = sym_idx__dedup_aliases(dso, elf, strndx, entries, *nr);
> + return 0;
> +}
> +
> +static struct symbol *sym_idx__new_symbol(struct dso *dso,
> + const struct sym_idx *idx,
> + const char *name)
> +{
> + char *demangled = dso__demangle_sym(dso, 0, name);
> + struct symbol *sym;
> +
> + sym = symbol__new(idx->start, idx->end - idx->start, idx->binding,
> + idx->type, demangled ?: name);
> + free(demangled);
> + if (sym && (idx->flags & SYM_IDX_FLAG_IFUNC_ALIAS))
> + symbol__set_ifunc_alias(sym, true);
> + return sym;
> +}
> +
> +/*
> + * PLT synthesis names IRELATIVE slots after the IFUNC they resolve to, which
> + * it looks up in the rb-tree while dso__load() holds the DSO lock.
> + * Materialize IFUNCs now, with names from the ELF image, so that it finds
> + * them without reading names.
> + */
> +static void sym_idx__materialize_ifuncs(struct dso *dso, Elf *elf,
> + const size_t *strndx)
> +{
> + struct dso_ondemand *od = dso__ondemand(dso);
> + u32 i;
> +
> + for (i = 0; i < od->nr_sorted; i++) {
> + struct sym_idx *idx = &od->sorted[i];
> + const char *name;
> + struct symbol *sym;
> +
> + if (idx->type != STT_GNU_IFUNC &&
> + !(idx->flags & SYM_IDX_FLAG_IFUNC_ALIAS))
> + continue;
> + name = sym_idx__elf_name(elf, strndx, idx);
> + sym = name ? sym_idx__new_symbol(dso, idx, name) : NULL;
> + if (!sym)
> + continue;
> + __symbols__insert(dso__symbols(dso), sym);
> + idx->flags |= SYM_IDX_FLAG_MATERIALIZED;
> + }
> +}
> +
> +/*
> + * Check that @path can be reopened to read names. This runs under the DSO
> + * lock, so it opens the file directly: the DSO data cache takes its global
> + * lock before DSO locks.
> + */
> +static bool dso_ondemand__source_readable(const struct dso_ondemand *od,
> + const char *path)
> +{
> + const struct sym_idx *idx = &od->sorted[0];
> + const struct sym_idx_strtab *strtab;
> + u64 off;
> + u8 probe;
> + bool ok;
> + int fd;
> +
> + strtab = &od->strtab[!!(idx->flags & SYM_IDX_FLAG_DYNSTR)];
> + if (idx->name_off >= strtab->size ||
> + check_add_overflow(strtab->offset, (u64)idx->name_off, &off) ||
> + off > INT64_MAX)
> + return false;
> +
> + fd = open(path, O_RDONLY | O_CLOEXEC);
> + if (fd < 0)
> + return false;
> + ok = pread(fd, &probe, 1, off) == 1;
> + close(fd);
> + return ok;
> +}
> +
> +/*
> + * Index the symbols of .symtab, if @dynsym is zero, and .dynsym. Like
> + * dso__load_sym(), fix up .symtab first and then both tables together.
> + */
> +static int dso__build_ondemand_index(struct dso *dso, struct symsrc *syms_ss,
> + struct symsrc *runtime_ss,
> + int dynsym)
> +{
> + struct sym_idx_strtab strtab[2] = {};
> + struct sym_idx *entries = NULL, *shrunk;
> + struct dso_ondemand *od;
> + size_t strndx[2] = {};
> + u32 nr = 0, prev;
> + bool in_host_ns;
> + int i, err = 0;
> +
> + /*
> + * GNU debugdata is backed by a temporary decompressed fd rather than a
> + * reopenable source path. Keep using the eager loader for that case,
> + * including the runtime .dynsym that dso__load_sym() then loads with
> + * the runtime file as @syms_ss, so that both are fixed up together.
> + */
> + if (dso__symtab_type(dso) == DSO_BINARY_TYPE__GNU_DEBUGDATA)
> + return 0;
> +
> + for (i = dynsym; i < 2; i++) {
> + if (!(i ? syms_ss->dynsym : syms_ss->symtab))
> + continue;
> + prev = nr;
> + err = sym_idx__add_table(syms_ss, runtime_ss, i, &entries, &nr,
> + strtab, strndx);
> + if (!err && nr > prev)
> + err = sym_idx__fixup(dso, syms_ss->elf, strndx, entries, &nr);
> + if (err)
> + goto out_free;
> + }
> + if (!nr)
> + goto out_free;
> +
> + shrunk = realloc(entries, nr * sizeof(*entries));
> + if (shrunk)
> + entries = shrunk;
> +
> + od = dso_ondemand__new();
> + if (!od) {
> + err = -1;
> + goto out_free;
> + }
> + od->sorted = entries;
> + od->nr_sorted = nr;
> + memcpy(od->strtab, strtab, sizeof(strtab));
> + entries = NULL;
> + if (dso_ondemand__link_overlaps(od)) {
> + err = -1;
> + goto out_free_od;
> + }
> +
> + /*
> + * dso__load() has just opened build-id cache files from outside the
> + * mount namespace of the DSO, so read names from them there too.
> + * Other sources were opened in the namespace we are in now.
> + */
> + in_host_ns = syms_ss->type == DSO_BINARY_TYPE__BUILD_ID_CACHE ||
> + syms_ss->type == DSO_BINARY_TYPE__BUILD_ID_CACHE_DEBUGINFO;
> + if (!in_host_ns && !dso_ondemand__source_readable(od, syms_ss->name))
> + goto out_free_od;
> + od->data_dso = dso__new(syms_ss->name);
> + if (!od->data_dso ||
> + dso__data_set_path(od->data_dso, syms_ss->name) < 0)
> + goto out_free_od;
> + dso__set_binary_type(od->data_dso, DSO_BINARY_TYPE__SYSTEM_PATH_DSO);
> + if (!in_host_ns)
> + dso__set_nsinfo(od->data_dso, nsinfo__get(dso__nsinfo(dso)));
> +
> + dso__set_ondemand(dso, od);
> + sym_idx__materialize_ifuncs(dso, syms_ss->elf, strndx);
> +
> + pr_debug("%s: on-demand index: %u symbols (%zu bytes)\n",
> + dso__long_name(dso), nr,
> + nr * (sizeof(*od->sorted) + (od->outer ? sizeof(*od->outer) : 0)));
> + return 1;
> +
> +out_free_od:
> + dso_ondemand__free(od);
> + return err;
> +out_free:
> + free(entries);
> + return err;
> +}
> +
> +const char *dso__read_ondemand_symbol_name(struct dso *data_dso,
> + u64 strtab_offset, u64 strtab_size,
> + u64 name_off, char *buf,
> + size_t buflen, char **to_free,
> + unsigned int *nr_reads)
> +{
> + ssize_t n;
> + u64 remain;
> + u64 file_off;
> + size_t cap, want;
> +
> + *to_free = NULL;
> +
> + if (name_off >= strtab_size)
> + return NULL;
> + if (check_add_overflow(strtab_offset, name_off, &file_off))
> + return NULL;
> + remain = strtab_size - name_off;
> + if (nr_reads)
> + *nr_reads = 0;
> +
> + want = min((u64)(buflen - 1), remain);
> + if (nr_reads)
> + (*nr_reads)++;
> + n = dso__data_read_offset(data_dso, NULL, file_off, (u8 *)buf, want);
> + if (n <= 0)
> + return NULL;
> + buf[n] = '\0';
> + if (memchr(buf, '\0', n))
> + return buf;
> + if ((size_t)n < want)
> + return NULL;
> +
> + cap = 4096;
> + for (;;) {
> + char *tmp;
> +
> + want = cap;
> + if (want > remain)
> + want = remain;
> + if (want == 0)
> + break;
> +
> + tmp = *to_free ? realloc(*to_free, want + 1) : malloc(want + 1);
> + if (!tmp) {
> + free(*to_free);
> + *to_free = NULL;
> + return NULL;
> + }
> + *to_free = tmp;
> +
> + if (nr_reads)
> + (*nr_reads)++;
> + n = dso__data_read_offset(data_dso, NULL, file_off,
> + (u8 *)*to_free, want);
> + if (n <= 0) {
> + free(*to_free);
> + *to_free = NULL;
> + return NULL;
> + }
> + (*to_free)[n] = '\0';
> +
> + if (memchr(*to_free, '\0', n))
> + return *to_free;
> + if ((size_t)n < want)
> + break;
> +
> + if (want >= remain || (u64)n >= remain)
> + break;
> +
> + if (cap > SIZE_MAX / 2)
> + break;
> + cap *= 2;
> + }
> +
> + free(*to_free);
> + *to_free = NULL;
> + return NULL;
> +}
> +
> +/*
> + * A snapshot of an index entry, so that its name can be read without the DSO
> + * lock: name reads go through the DSO data cache, which takes its global lock
> + * and then the lock of the DSO it opens. The index, and its data source, stay
> + * allocated until the last reader is done, even once detached from the DSO.
> + */
> +struct sym_idx_ref {
> + struct dso_ondemand *od;
> + struct sym_idx_strtab strtab;
> + struct sym_idx entry;
> + u32 pos;
> +};
> +
> +static void dso_ondemand__get_ref(struct dso_ondemand *od, u32 pos,
> + struct sym_idx_ref *ref)
> +{
> + ref->od = od;
> + ref->entry = od->sorted[pos];
> + ref->strtab = od->strtab[!!(ref->entry.flags & SYM_IDX_FLAG_DYNSTR)];
> + ref->pos = pos;
> + od->nr_readers++;
> +}
> +
> +static struct dso_ondemand *dso_ondemand__put_ref(struct dso *dso,
> + const struct sym_idx_ref *ref)
> + EXCLUSIVE_LOCKS_REQUIRED(dso__lock(dso))
> +{
> + struct dso_ondemand *od = ref->od;
> +
> + if (--od->nr_readers || dso__ondemand(dso) == od)
> + return NULL;
> + return od;
> +}
> +
> +static struct symbol *sym_idx_ref__read(struct dso *dso,
> + const struct sym_idx_ref *ref)
> + LOCKS_EXCLUDED(dso__lock(dso))
> +{
> + char namebuf[1024];
> + char *name_heap;
> + const char *name;
> + struct symbol *sym;
> +
> + mutex_lock(&ref->od->read_lock);
> + name = dso__read_ondemand_symbol_name(ref->od->data_dso, ref->strtab.offset,
> + ref->strtab.size,
> + ref->entry.name_off, namebuf,
> + sizeof(namebuf), &name_heap, NULL);
> + mutex_unlock(&ref->od->read_lock);
> + if (!name)
> + return NULL;
> + sym = sym_idx__new_symbol(dso, &ref->entry, name);
> + free(name_heap);
> + return sym;
> +}
> +
> +/*
> + * Insert @sym, read for @ref, unless the entry or the whole index was
> + * materialized meanwhile. Return the symbol the tree now has for the entry.
> + */
> +static struct symbol *dso_ondemand__insert(struct dso *dso,
> + const struct sym_idx_ref *ref,
> + struct symbol *sym)
> + EXCLUSIVE_LOCKS_REQUIRED(dso__lock(dso))
> +{
> + struct sym_idx *idx = &ref->od->sorted[ref->pos];
> +
> + if (dso__ondemand(dso) != ref->od ||
> + idx->flags & SYM_IDX_FLAG_MATERIALIZED) {
> + if (sym)
> + symbol__delete(sym);
> + return symbols__find_start(dso__symbols(dso), ref->entry.start);
> + }
> + if (sym) {
> + __symbols__insert(dso__symbols(dso), sym);
> + idx->flags |= SYM_IDX_FLAG_MATERIALIZED;
> + }
> + return sym;
> +}
> +
> +static bool dso_ondemand__next_ref(struct dso *dso, u32 *pos,
> + struct sym_idx_ref *ref)
> + EXCLUSIVE_LOCKS_REQUIRED(dso__lock(dso))
> +{
> + struct dso_ondemand *od = dso__ondemand(dso);
> +
> + if (!od)
> + return false;
> + while (*pos < od->nr_sorted &&
> + (od->sorted[*pos].flags & SYM_IDX_FLAG_MATERIALIZED))
> + (*pos)++;
> + if (*pos >= od->nr_sorted)
> + return false;
> + dso_ondemand__get_ref(od, (*pos)++, ref);
> + return true;
> +}
> +
> +void dso__materialize_symbols_ondemand(struct dso *dso)
> +{
> + struct dso_ondemand *od;
> + struct sym_idx_ref ref;
> + u32 pos = 0, nr_failed = 0;
> + bool more;
> +
> + if (!symbol_conf.lazy_load_symbols)
> + return;
> +
> + for (;;) {
> + struct symbol *sym;
> +
> + mutex_lock(dso__lock(dso));
> + more = dso_ondemand__next_ref(dso, &pos, &ref);
> + mutex_unlock(dso__lock(dso));
> + if (!more)
> + break;
> +
> + sym = sym_idx_ref__read(dso, &ref);
> + if (!sym)
> + nr_failed++;
> + mutex_lock(dso__lock(dso));
> + dso_ondemand__insert(dso, &ref, sym);
> + od = dso_ondemand__put_ref(dso, &ref);
> + mutex_unlock(dso__lock(dso));
> + dso_ondemand__free(od);
> + }
> +
> + /*
> + * Detach the index before the caller builds the name-sorted array, so
> + * that a DSO with that array never changes again.
> + */
> + mutex_lock(dso__lock(dso));
> + od = dso__ondemand(dso);
> + dso__set_ondemand(dso, NULL);
> + if (od && od->nr_readers)
> + od = NULL;
> + mutex_unlock(dso__lock(dso));
> +
> + if (nr_failed)
> + pr_debug("%s: cannot read %u lazily loaded symbol names\n",
> + dso__long_name(dso), nr_failed);
> + dso_ondemand__free(od);
> +}
> +
> +struct symbol *dso__find_symbol_ondemand(struct dso *dso, u64 addr)
> +{
> + struct dso_ondemand *od;
> + struct sym_idx_ref ref;
> + struct symbol *sym;
> + u32 pos;
> +
> + mutex_lock(dso__lock(dso));
> + od = dso__ondemand(dso);
> + pos = od ? sym_idx__find(od, addr) : SYM_IDX_NONE;
> + if (pos == SYM_IDX_NONE) {
> + /* Synthesized PLT symbols, or a fully materialized DSO. */
> + sym = dso__find_symbol_cached(dso, addr);
> + } else if (od->sorted[pos].flags & SYM_IDX_FLAG_MATERIALIZED) {
> + sym = symbols__find_start(dso__symbols(dso), od->sorted[pos].start);
> + } else {
> + dso_ondemand__get_ref(od, pos, &ref);
> + mutex_unlock(dso__lock(dso));
> + sym = sym_idx_ref__read(dso, &ref);
> + mutex_lock(dso__lock(dso));
> + sym = dso_ondemand__insert(dso, &ref, sym);
> + od = dso_ondemand__put_ref(dso, &ref);
> + mutex_unlock(dso__lock(dso));
> + dso_ondemand__free(od);
> + return sym;
> + }
> + mutex_unlock(dso__lock(dso));
> + return sym;
> +}
> +
> static int
> dso__load_sym_internal(struct dso *dso, struct map *map, struct symsrc *syms_ss,
> struct symsrc *runtime_ss, int kmodule, int dynsym)
> @@ -1626,6 +2524,28 @@ dso__load_sym_internal(struct dso *dso, struct map *map, struct symsrc *syms_ss,
> if (kmodule && adjust_kernel_syms)
> max_text_sh_offset = max_text_section(runtime_ss->elf, &runtime_ss->ehdr);
>
> + /*
> + * PPC64 ELFv1 function symbols need the eager loop's .opd descriptor
> + * translation. For symtabs, the selected and runtime sources can differ.
> + */
> + if (symbol_conf.lazy_load_symbols && !dso__kernel(dso) && !kmodule &&
> + !syms_ss->opdsec && (dynsym || !runtime_ss->opdsec)) {
> + int oret = 0;
> +
> + /*
> + * The index covers .symtab and .dynsym. If .symtab was
> + * loaded eagerly, load .dynsym eagerly too.
> + */
> + if (!dynsym || !syms_ss->symtab)
> + oret = dso__build_ondemand_index(dso, syms_ss,
> + runtime_ss, dynsym);
> +
> + if (oret < 0)
> + return oret;
> + if (dso__ondemand(dso))
> + return 1;
> + }
> +
> curr_dso = dso__get(dso);
> elf_symtab__for_each_symbol(syms, nr_syms, idx, sym) {
> struct symbol *f;
> diff --git a/tools/perf/util/symbol-minimal.c b/tools/perf/util/symbol-minimal.c
> index 0a71d1463952..2f51f07b41be 100644
> --- a/tools/perf/util/symbol-minimal.c
> +++ b/tools/perf/util/symbol-minimal.c
> @@ -373,6 +373,15 @@ void symbol__elf_init(void)
> {
> }
>
> +struct symbol *dso__find_symbol_ondemand(struct dso *dso, u64 addr)
> +{
> + return dso__find_symbol_cached(dso, addr);
> +}
> +
> +void dso__materialize_symbols_ondemand(struct dso *dso __maybe_unused)
> +{
> +}
> +
> bool filename__has_section(const char *filename __maybe_unused, const char *sec __maybe_unused)
> {
> return false;
> diff --git a/tools/perf/util/symbol.c b/tools/perf/util/symbol.c
> index f590b69f9f01..39c87835fc1e 100644
> --- a/tools/perf/util/symbol.c
> +++ b/tools/perf/util/symbol.c
> @@ -634,7 +634,7 @@ void dso__delete_symbol(struct dso *dso, struct symbol *sym)
> dso__reset_find_symbol_cache(dso);
> }
>
> -struct symbol *dso__find_symbol(struct dso *dso, u64 addr)
> +struct symbol *dso__find_symbol_cached(struct dso *dso, u64 addr)
> {
> if (dso__last_find_result_addr(dso) != addr || dso__last_find_result_symbol(dso) == NULL) {
> dso__set_last_find_result_addr(dso, addr);
> @@ -644,6 +644,13 @@ struct symbol *dso__find_symbol(struct dso *dso, u64 addr)
> return dso__last_find_result_symbol(dso);
> }
>
> +struct symbol *dso__find_symbol(struct dso *dso, u64 addr)
> +{
> + if (symbol_conf.lazy_load_symbols)
> + return dso__find_symbol_ondemand(dso, addr);
> + return dso__find_symbol_cached(dso, addr);
> +}
> +
> struct symbol *dso__find_symbol_nocache(struct dso *dso, u64 addr)
> {
> return symbols__find(dso__symbols(dso), addr);
> @@ -690,6 +697,7 @@ struct symbol *dso__find_symbol_by_name(struct dso *dso, const char *name, size_
>
> void dso__sort_by_name(struct dso *dso)
> {
> + dso__materialize_symbols_ondemand(dso);
> mutex_lock(dso__lock(dso));
> if (!dso__sorted_by_name(dso)) {
> size_t len = 0;
> @@ -1940,11 +1948,19 @@ int dso__load(struct dso *dso, struct map *map)
> }
>
> #ifdef HAVE_LIBBFD_SUPPORT
> +#ifdef HAVE_LIBELF_SUPPORT
> + if (is_reg && !symbol_conf.lazy_load_symbols)
> +#else
> if (is_reg)
> +#endif
> bfdrc = dso__load_bfd_symbols(dso, name);
> #endif
> if (is_reg && bfdrc < 0)
> sirc = symsrc__init(ss, dso, name, symtab_type);
> +#if defined(HAVE_LIBBFD_SUPPORT) && defined(HAVE_LIBELF_SUPPORT)
> + if (is_reg && symbol_conf.lazy_load_symbols && sirc < 0)
> + bfdrc = dso__load_bfd_symbols(dso, name);
> +#endif
>
> if (nsexit)
> nsinfo__mountns_enter(dso__nsinfo(dso), &nsc);
> diff --git a/tools/perf/util/symbol.h b/tools/perf/util/symbol.h
> index b9fa722a9a14..a2ff8f8327d6 100644
> --- a/tools/perf/util/symbol.h
> +++ b/tools/perf/util/symbol.h
> @@ -200,6 +200,7 @@ void dso__insert_symbol(struct dso *dso,
> void dso__delete_symbol(struct dso *dso,
> struct symbol *sym);
>
> +struct symbol *dso__find_symbol_cached(struct dso *dso, u64 addr);
> struct symbol *dso__find_symbol(struct dso *dso, u64 addr);
> struct symbol *dso__find_symbol_nocache(struct dso *dso, u64 addr);
>
> diff --git a/tools/perf/util/symbol_conf.h b/tools/perf/util/symbol_conf.h
> index 37d35f42dcc1..434b6c51288b 100644
> --- a/tools/perf/util/symbol_conf.h
> +++ b/tools/perf/util/symbol_conf.h
> @@ -77,6 +77,7 @@ struct symbol_conf {
> no_buildid_mmap2,
> guest_code,
> lazy_load_kernel_maps,
> + lazy_load_symbols,
> keep_exited_threads,
> annotate_data_member,
> annotate_data_sample,
>
> --
> Git-157)
>
>
next prev parent reply other threads:[~2026-10-07 22:15 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 17:30 [PATCH v5 0/5] perf script: Lazy " Alireza Haghdoost via B4 Relay
2026-10-05 17:30 ` [PATCH v5 1/5] perf symbols: Fix broken ELF_C_READ_MMAP fallback guard Alireza Haghdoost via B4 Relay
2026-10-05 17:30 ` [PATCH v5 2/5] perf dso: Allow reading DSO data from an explicit file Alireza Haghdoost via B4 Relay
2026-10-07 22:09 ` Namhyung Kim
2026-10-05 17:30 ` [PATCH v5 3/5] perf symbols: Factor out duplicate symbol selection Alireza Haghdoost via B4 Relay
2026-10-05 17:30 ` [PATCH v5 4/5] perf script: Add --lazy-load-symbols for lazy symbol loading Alireza Haghdoost via B4 Relay
2026-10-07 22:15 ` Namhyung Kim [this message]
2026-10-05 17:30 ` [PATCH v5 5/5] perf test: Test " Alireza Haghdoost via B4 Relay
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asbElicX1u_GZu_j@google.com \
--to=namhyung@kernel.org \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=alexander.shishkin@linux.intel.com \
--cc=andriin@fb.com \
--cc=ast@kernel.org \
--cc=haghdoost@uber.com \
--cc=irogers@google.com \
--cc=james.clark@linaro.org \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®