Re: [PATCH v3 0/6] perf script: Bounded and lazy symbol loading

From: Ian Rogers

Date: Fri Sep 25 2026 - 17:56:36 EST


On Fri, Sep 25, 2026 at 2:28 PM Alireza Haghdoost <haghdoost@xxxxxxxx> wrote:
>
> > So currently we have a symbol as:
> > ```
> > /**
> > * A symtab entry. When allocated this may be preceded by an annotation (see
> > * symbol__annotation) and/or a browser_index (see symbol__browser_index).
> > */
> > struct symbol {
> > struct rb_node rb_node;
> > /** Range of symbol [start, end). */
> > u64 start;
> > u64 end;
> > /** Length of the string name. */
> > u16 namelen;
> > _Atomic uint16_t flags;
> > /** Architecture specific. Unused except on PPC where it holds
> > st_other. */
> > u8 arch_sym;
> > /** The name of length namelen associated with the symbol. */
> > char name[];
> > };
> > ```
> > and the struct rb_node is a significant size cost:
> > ```
> > struct rb_node {
> > unsigned long __rb_parent_color;
> > struct rb_node *rb_right;
> > struct rb_node *rb_left;
> > } __attribute__((aligned(sizeof(long))));
> > ```
> >
> > A sorted array of symbols would only require 1 pointer per entry,
> > saving 2 words. The array could reside in the "symbols" variable
> > passed around something like this:
> > ```
> > struct symbols {
> > struct rw_semaphore lock;
> > struct symbol **symbols;
> > unsigned int cnt;
> > unsigned int allocated;
> > bool sorted;
> > };
> > ```
> > This is basically what we do with struct dso and dsos. With a struct
> > symbols we can add a lock and ensure operations are appropriately
> > synchronized. On top of this we can build lazy symbol resolution.
> >
>
> I agree that a sorted struct symbols array with its own lock would be
> a cleaner container than the rb-tree. However, it would not save much
> memory on its own. On a sample fixture, perf script materializes about
> 765k symbols and peaks at 265 MiB RssAnon. Replacing the rb_node with
> one pointer per symbol saves 16 bytes each, about 12 MiB in total.
> Most of the memory goes to materializing symbols that are never
> sampled. Lazy loading avoids that work and brings the peak down to
> 39 MiB. Lazily materialized symbols are inserted one at a time between
> lookups, which suits a tree better than a sorted array.
>
> If the maintainers prefer the sorted array, I'm open to it, but I'd
> like to be clear about the cost. The symbol rb-tree is used at around
> 100 call sites in 18 files under tools/perf, so the conversion would
> be a separate series that has to land first. It should also cover both
> eager and lazy loading, so there is one container for materialized
> symbols rather than a tree in one mode and an array in the other.
>
> > ...For
> > the bounded memory size, I'd prefer something like a reference count
> > in every symbol (which uses one word we previously saved). We can
> > periodically scan the symbols array for symbols with a reference count
> > of 1 and then lower that count, knowing it is safe to release the
> > symbol's memory because no other thread is referencing it. This would
> > avoid having symbols appear as "[unknown]". We should really have a
> > similar scan of the dsos to free up their memory.
> >
>
> Eviction would bound the steady state, but it doesn't bound the peak
> on its own. Eager loading materializes the whole symtab when the DSO
> is loaded, before anything can be evicted, and this series does not
> add lazy loading for every loader (PPC64 .opd, .gnu_debugdata, the
> kernel and modules still load eagerly). --max-symbol-bytes covers
> those paths too. It is opt-in and off by default, and is meant for
> hosts where perf shares memory with latency-sensitive services and the
> operator needs a guaranteed upper bound. Like perf record --max-size,
> what happens past the cap is deterministic: perf warns and reports
> [unknown].
>
> >
> > I think the patches 4 and 5 do the core work, but they introduce some
> > warts that I wish didn't have to exist. How do you feel about
> > reference counting? To avoid reference count leaks we adding a
> > checking framework that is documented here:
> > https://perfwiki.github.io/main/reference-count-checking/
>
> Reference counting would be useful for symbol lifetime and for the
> annotation and browser_index cleanup, and freeing unreferenced symbols
> would reduce steady-state memory. As above, I see it as a complement
> to the cap rather than a replacement, since it can't guarantee an
> upper bound.
>
> Refcounting struct symbol touches every holder of a symbol pointer,
> so I'd suggest doing it as a separate series using the refcount
> checking framework rather than in this series.
>
> If there are specific parts of patches 4 and 5 you see as warts,
> please point me at them and I'll address them in v4.
>
> For v4 I currently plan to bound the second pass of
> dso__build_ondemand_index() by the allocated count, as reported by
> Sashiko.

So, I'm pretty against the complexity of dealing with symbols that may
fail due to a memory pressure issue. I'm also against having memory of
symbols bound, while memory for say debuginfo isn't. Ideally we'd have
a language runtime that'd allow us to express some notion of weakly
holding symbols live. I think we can implement a global function to
reduce the memory footprint by iterating through sessions, machines,
dsos, and symbols, squeezing them when possible. The problem with
implementing that on top of patches 4 and 5 is that it's hard to see
what would be kept, so we'd just revert the changes and start
implementing the reference counting solution again.

In the change there seems to be the addition of helper functions and
structs like symbol_candidate. These appear in header files to be
shared across C files. We have multiple notions of symbols, with and
without libelf. I'd rather the change were more minimal given I think
this is the wrong direction.

Thanks,
Ian

> Thanks,
> Alireza