Re: [PATCH v3 0/6] perf script: Bounded and lazy symbol loading

From: Ian Rogers

Date: Fri Sep 25 2026 - 16:25:38 EST


On Fri, Sep 25, 2026 at 12:15 PM Alireza Haghdoost via B4 Relay
<devnull+haghdoost.uber.com@xxxxxxxxxx> wrote:
>
> perf script loads the entire ELF symbol table of every DSO that appears
> in a sample, allocating each symbol into an rb-tree held until process
> exit. Therefore, a large enough profile turns symbol loading into an
> OOM kill. This does not scale to profiling a large cgroup with many large
> binaries on a production system with limited free memory.
>
> This series adds two independent, opt-in mechanisms, a leading
> regression fix, and two preparatory patches:
>
> [1/6] Fix a broken "#ifdef ELF_C_READ_MMAP" guard so perf actually
> mmaps ELF files instead of malloc'ing section data. This is a
> standalone regression fix for 22dd1ac91a77.
>
> [2/6] Let a DSO read its data from one explicit file through the DSO
> data cache. This fixes the split-debuginfo case where offsets from
> the debuginfo file would be applied to the runtime image.
>
> [3/6] Factor duplicate-symbol selection so it works on symbol
> attributes rather than struct symbol. No functional change.
>
> [4/6] --max-symbol-bytes <size>: a byte budget on struct symbol
> allocations (and the lazy index) enforced at the ELF symbol
> loader, degrading to [unknown] with a warning past the cap.
> An unbounded profile doesn't just risk OOM-killing itself. It
> also forces memory pressure on the whole host, pushing the kernel
> to reclaim from co-located latency-sensitive processes. Capping
> it lets the user bound that footprint up front and choose the
> trade-off explicitly.
>
> [5/6] --lazy-load-symbols: build a compact per-DSO sorted index and
> resolve only the sampled addresses, reading names through the DSO
> data cache at lookup time. On the production fixture, peak RssAnon
> drops from 265 MiB to 39 MiB (6.8x) and wall time from 3.1 s to
> 1.85 s (1.7x). Memory optimizations usually cost time; this one
> does not because lazy loading skips a lot of calloc and demangle
> calls.
>
> [6/6] Shell and unit tests for both options.
>
> Lazy loading handles the common userspace ELF symtab/dynsym path. Eager
> loading remains available for dense coverage and for PPC64 .opd and
> .gnu_debugdata.

So currently we have a symbol as:
```
/**
* A symtab entry. When allocated this may be preceded by an annotation (see
* symbol__annotation) and/or a browser_index (see symbol__browser_index).
*/
struct symbol {
struct rb_node rb_node;
/** Range of symbol [start, end). */
u64 start;
u64 end;
/** Length of the string name. */
u16 namelen;
_Atomic uint16_t flags;
/** Architecture specific. Unused except on PPC where it holds
st_other. */
u8 arch_sym;
/** The name of length namelen associated with the symbol. */
char name[];
};
```
and the struct rb_node is a significant size cost:
```
struct rb_node {
unsigned long __rb_parent_color;
struct rb_node *rb_right;
struct rb_node *rb_left;
} __attribute__((aligned(sizeof(long))));
```

A sorted array of symbols would only require 1 pointer per entry,
saving 2 words. The array could reside in the "symbols" variable
passed around something like this:
```
struct symbols {
struct rw_semaphore lock;
struct symbol **symbols;
unsigned int cnt;
unsigned int allocated;
bool sorted;
};
```
This is basically what we do with struct dso and dsos. With a struct
symbols we can add a lock and ensure operations are appropriately
synchronized. On top of this we can build lazy symbol resolution. For
the bounded memory size, I'd prefer something like a reference count
in every symbol (which uses one word we previously saved). We can
periodically scan the symbols array for symbols with a reference count
of 1 and then lower that count, knowing it is safe to release the
symbol's memory because no other thread is referencing it. This would
avoid having symbols appear as "[unknown]". We should really have a
similar scan of the dsos to free up their memory.

If we have a reference count then symbol__annotation and
symbol__browser_index can reference a symbol rather than using
container_of.

I think the patches 4 and 5 do the core work, but they introduce some
warts that I wish didn't have to exist. How do you feel about
reference counting? To avoid reference count leaks we adding a
checking framework that is documented here:
https://perfwiki.github.io/main/reference-count-checking/

Thanks,
Ian








> Changes in v3:
>
> - Rebase onto current perf-tools-next.
> - Pick up Namhyung's Reviewed-by for patch 1.
> - Split the exact-path DSO data support into its own patch (2/6), with a
> DSO data test for reading and reopening through an explicit path.
> - Move the duplicate-selection refactor into a preparatory patch (3/6)
> and factor the whole lazy alias-group handling (traversal, demangling,
> IFUNC propagation, compaction) into one helper.
> - Keep struct symbol::namelen as u16. Charge symbol bytes from the stored
> namelen on both allocation and free, and drop the 64 KiB-name test.
> - Fix lazy-loading races reported by Sashiko: in lazy mode, address
> lookups always take the DSO lock, and building the name-sorted array
> materializes and frees the lazy index even when the budget truncates
> it, so the name array is never invalidated. dso__reset_symbol_names()
> is gone. Add a concurrent budget-truncation test.
> - In lazy mode, when no PT_LOAD covers a symbol and its section is NOBITS
> in the debuginfo file, adjust with the runtime section header as eager
> loading does. Add a lazy/eager symbol parity test and a split-debuginfo
> shell test that exercises this path.
> - Keep each unit test with the code it needs (DSO data in 2/6, budget
> reservation in 4/6); the other tests stay in 6/6.
>
> Link: https://lore.kernel.org/all/20260919-perf-symbol-memory-send-v2-0-495b8f00ad7c@xxxxxxxx/
>
> Changes in v2:
>
> - Replace direct pread() name reads with the exact symbol source's DSO data
> cache, preserving split-debuginfo offsets and descriptor reopen behavior.
> - Drop the byte-identical-output claim and retain eager loading for PPC64
> .opd and .gnu_debugdata.
> - Make the symbol budget atomic and strict, account complete name lengths,
> accept a bare 0 as unlimited, and keep partial zero-sized ranges from
> covering omitted symbols.
> - Align lazy lookup with eager duplicate and IFUNC selection, PLT clipping,
> and name-sorted materialization.
> - Move option documentation into the feature patches. Add unit and shell
> coverage for cache reopen, truncated names, budget truncation, and skip
> handling.
>
> Link: https://lore.kernel.org/all/20260915-perf-symbol-memory-send-v1-0-1d3360e21f07@xxxxxxxx/
> ---
> Alireza Haghdoost (6):
> perf symbols: Fix broken ELF_C_READ_MMAP fallback guard
> perf dso: Allow reading DSO data from an explicit file
> perf symbols: Factor out duplicate symbol selection
> perf script: Add --max-symbol-bytes to bound ELF symbol memory
> perf script: Add --lazy-load-symbols for lazy symbol loading
> perf test: Test lazy symbol loading and symbol memory limits
>
> tools/perf/Documentation/perf-script.txt | 26 +
> tools/perf/arch/powerpc/util/sym-handling.c | 6 +-
> tools/perf/builtin-script.c | 44 ++
> tools/perf/tests/Build | 1 +
> tools/perf/tests/builtin-test.c | 1 +
> tools/perf/tests/dso-data.c | 42 ++
> .../tests/shell/lazy_load_symbols_split_debug.sh | 113 ++++
> tools/perf/tests/shell/script_lazy_load_symbols.sh | 278 ++++++++
> .../tests/shell/script_lazy_load_symbols_skip.sh | 26 +
> tools/perf/tests/symbol-bytes.c | 599 +++++++++++++++++
> tools/perf/tests/tests.h | 1 +
> tools/perf/util/dso.c | 60 +-
> tools/perf/util/dso.h | 47 ++
> tools/perf/util/map.c | 19 +-
> tools/perf/util/symbol-elf.c | 728 ++++++++++++++++++++-
> tools/perf/util/symbol-minimal.c | 16 +
> tools/perf/util/symbol.c | 143 +++-
> tools/perf/util/symbol.h | 30 +-
> tools/perf/util/symbol_conf.h | 2 +
> 19 files changed, 2133 insertions(+), 49 deletions(-)
> ---
> base-commit: edd8a9fe2eca009599e013a29c421c7a6b5ad1b9
> change-id: 20260915-perf-symbol-memory-send-e7cfca1ac3d9
>
> Best regards,
> --
> Alireza Haghdoost <haghdoost@xxxxxxxx>
>
>