Re: [PATCH v1 0/7] perf symbol: Reference counting, flat array storage, and LRU shrinking

From: Ian Rogers

Date: Mon Sep 28 2026 - 18:05:22 EST


On Mon, Sep 28, 2026 at 2:51 PM Namhyung Kim <namhyung@xxxxxxxxxx> wrote:
>
> Hello,
>
> On Mon, Sep 28, 2026 at 01:55:26PM -0700, Ian Rogers wrote:
> > On Mon, Sep 28, 2026 at 12:45 PM Alireza Haghdoost <haghdoost@xxxxxxxx> wrote:
> > >
> > > On Mon, Sep 28, 2026 at 12:52 AM Ian Rogers <irogers@xxxxxxxxxx> wrote:
> > > > annotations. Saves 48 bytes per struct symbol (two rb_nodes) and provides
> > > > cache-friendly binary search (bsearch) and callback-based iteration.
> > >
> > > Hi Ian,
> > >
> > > Thanks for posting this and for Cc'ing me. I built it on edd8a9fe2eca0,
> > > the same base as the v3 series I sent last week, and compared it with
> > > v3 in eager mode (today's behavior) and with --lazy-load-symbols.
> > >
> > > All runs use perf script --no-inline --max-stack 127
> > > -F comm,tid,time,ip,sym,dso, with one warm-up and three runs; the
> > > median is shown below. Peak anon is the maximum RssAnon from
> > > /proc/PID/status, sampled every 10 ms.
> > >
> > > 1) Production: 120 s cgroup profile of a storage service on an 80-CPU
> > > host (perf record -a -g -F 99 -G <cgroup>). 54k samples, 11 DSOs,
> > > 5% of lines [unknown].
> > >
> > > time max RSS peak anon
> > > v3 eager 4.5 s 386 MiB 314 MiB
> > > v3 lazy 3.7 s 151 MiB 80 MiB
> > > this series 95.9 s 309 MiB 275 MiB
> > >
> > > 2) A sample fixture from my v3 cover letter, a worst case with 73% of
> > > lines [unknown]:
> > >
> > > time max RSS
> > > v3 eager 3.2 s 334 MiB
> > > v3 lazy 1.9 s 106 MiB
> > > this series 205.3 s 449 MiB
> > >
> > > Most of the time goes to reloading. On the fixture, a profile shows it
> > > under map__find_symbol() -> dso__load() -> dso__load_sym(). After each
> > > shrink, the first miss in a shrunk DSO re-parses its whole symtab, and
> > > addresses that resolve to [unknown] miss every time.
> > >
> > > The peak is still set by dso__load(), since the whole symtab is
> > > materialized before a shrink can run. That is the main blocker for us
> > > running perf script at scale: we can't bound how much memory it will
> > > use, so we have to contain it with memory.max, and then the profiling
> > > job fails with an OOM kill instead of degrading.
> >
> > Hi Alireza, and thanks for the feedback! I'll see what can be done
> > about dso__load in v2.
>
> Thanks for the patches and the analysis.
>
> I think many of this memory overhead come from the symbol string.
> Probably we don't use most of them. Then would it be nice if we can
> lazy-load the strings? Then it'd have an union of a pointer and a file
> offset for symbol strings. The find-by-name API is used rarely for
> user DSOs and the kernel symbols should be fully loaded anyway.
>
> Without the string (and the priv part), now it has a fixed size so it
> should be saved in an array directly. Probably we can add a limit there
> and print it with offset when it doesn't have the name.
>
> I'm a bit skeptical about the refcount approach here. It may be hard to
> determine when it reclaims memory. I guess the lazy-load strings with
> an array would give similar savings like Alireza's work.

Thanks Namhyung. Without reference counting we lack a good way to
discard symbols; we must either assume they always exist as in the
current code or implement reference counting. I don't see a way around
it. Alireza's patches introduced an "unknown" state for when a symbol
limit was hit, but I think that'd be frustrating in practice. We could
use it after trying to shrink memory use, in my opinion.

Fwiw, I never like switching the rbtree to an array. The issue is that
the rbtree has references that really should be managed by a reference
count. Doing that is a challenge in the current code and a sorted
array offers similar performance while simplifying the reference
counting.

We load all symbols instead of lazily because we need the end of a
symbol, which we can only determine by processing the entire symbol
table as the end of one symbol is the start of the next. We generally
translate one address in a DSO into a symbol, so loading all symbols
is overkill. I think we can do better by loading only symbols within a
certain address window rather than loading all symbols in a DSO. We
can then expand this window as more symbols are needed. This at least
bounds the number of symbols but isn't quite lazy loading.

Alireza also made good points about how symbols not being found or
being shrunk leads to thrashing patterns. I think we can fix this by
maintaining extra state in the DSO. I'm working to add this into v2.

Thanks,
Ian

> Thanks,
> Namhyung
>