Re: [PATCH v3 0/6] perf script: Bounded and lazy symbol loading
From: Alireza Haghdoost
Date: Fri Sep 25 2026 - 17:29:38 EST
> So currently we have a symbol as:
> ```
> /**
> * A symtab entry. When allocated this may be preceded by an annotation (see
> * symbol__annotation) and/or a browser_index (see symbol__browser_index).
> */
> struct symbol {
> struct rb_node rb_node;
> /** Range of symbol [start, end). */
> u64 start;
> u64 end;
> /** Length of the string name. */
> u16 namelen;
> _Atomic uint16_t flags;
> /** Architecture specific. Unused except on PPC where it holds
> st_other. */
> u8 arch_sym;
> /** The name of length namelen associated with the symbol. */
> char name[];
> };
> ```
> and the struct rb_node is a significant size cost:
> ```
> struct rb_node {
> unsigned long __rb_parent_color;
> struct rb_node *rb_right;
> struct rb_node *rb_left;
> } __attribute__((aligned(sizeof(long))));
> ```
>
> A sorted array of symbols would only require 1 pointer per entry,
> saving 2 words. The array could reside in the "symbols" variable
> passed around something like this:
> ```
> struct symbols {
> struct rw_semaphore lock;
> struct symbol **symbols;
> unsigned int cnt;
> unsigned int allocated;
> bool sorted;
> };
> ```
> This is basically what we do with struct dso and dsos. With a struct
> symbols we can add a lock and ensure operations are appropriately
> synchronized. On top of this we can build lazy symbol resolution.
>
I agree that a sorted struct symbols array with its own lock would be
a cleaner container than the rb-tree. However, it would not save much
memory on its own. On a sample fixture, perf script materializes about
765k symbols and peaks at 265 MiB RssAnon. Replacing the rb_node with
one pointer per symbol saves 16 bytes each, about 12 MiB in total.
Most of the memory goes to materializing symbols that are never
sampled. Lazy loading avoids that work and brings the peak down to
39 MiB. Lazily materialized symbols are inserted one at a time between
lookups, which suits a tree better than a sorted array.
If the maintainers prefer the sorted array, I'm open to it, but I'd
like to be clear about the cost. The symbol rb-tree is used at around
100 call sites in 18 files under tools/perf, so the conversion would
be a separate series that has to land first. It should also cover both
eager and lazy loading, so there is one container for materialized
symbols rather than a tree in one mode and an array in the other.
> ...For
> the bounded memory size, I'd prefer something like a reference count
> in every symbol (which uses one word we previously saved). We can
> periodically scan the symbols array for symbols with a reference count
> of 1 and then lower that count, knowing it is safe to release the
> symbol's memory because no other thread is referencing it. This would
> avoid having symbols appear as "[unknown]". We should really have a
> similar scan of the dsos to free up their memory.
>
Eviction would bound the steady state, but it doesn't bound the peak
on its own. Eager loading materializes the whole symtab when the DSO
is loaded, before anything can be evicted, and this series does not
add lazy loading for every loader (PPC64 .opd, .gnu_debugdata, the
kernel and modules still load eagerly). --max-symbol-bytes covers
those paths too. It is opt-in and off by default, and is meant for
hosts where perf shares memory with latency-sensitive services and the
operator needs a guaranteed upper bound. Like perf record --max-size,
what happens past the cap is deterministic: perf warns and reports
[unknown].
>
> I think the patches 4 and 5 do the core work, but they introduce some
> warts that I wish didn't have to exist. How do you feel about
> reference counting? To avoid reference count leaks we adding a
> checking framework that is documented here:
> https://perfwiki.github.io/main/reference-count-checking/
Reference counting would be useful for symbol lifetime and for the
annotation and browser_index cleanup, and freeing unreferenced symbols
would reduce steady-state memory. As above, I see it as a complement
to the cap rather than a replacement, since it can't guarantee an
upper bound.
Refcounting struct symbol touches every holder of a symbol pointer,
so I'd suggest doing it as a separate series using the refcount
checking framework rather than in this series.
If there are specific parts of patches 4 and 5 you see as warts,
please point me at them and I'll address them in v4.
For v4 I currently plan to bound the second pass of
dso__build_ondemand_index() by the allocated count, as reported by
Sashiko.
Thanks,
Alireza