Re: [PATCH v1 0/7] perf symbol: Reference counting, flat array storage, and LRU shrinking

From: Alireza Haghdoost

Date: Mon Sep 28 2026 - 15:55:59 EST


On Mon, Sep 28, 2026 at 12:52 AM Ian Rogers <irogers@xxxxxxxxxx> wrote:
> annotations. Saves 48 bytes per struct symbol (two rb_nodes) and provides
> cache-friendly binary search (bsearch) and callback-based iteration.

Hi Ian,

Thanks for posting this and for Cc'ing me. I built it on edd8a9fe2eca0,
the same base as the v3 series I sent last week, and compared it with
v3 in eager mode (today's behavior) and with --lazy-load-symbols.

All runs use perf script --no-inline --max-stack 127
-F comm,tid,time,ip,sym,dso, with one warm-up and three runs; the
median is shown below. Peak anon is the maximum RssAnon from
/proc/PID/status, sampled every 10 ms.

1) Production: 120 s cgroup profile of a storage service on an 80-CPU
host (perf record -a -g -F 99 -G <cgroup>). 54k samples, 11 DSOs,
5% of lines [unknown].

time max RSS peak anon
v3 eager 4.5 s 386 MiB 314 MiB
v3 lazy 3.7 s 151 MiB 80 MiB
this series 95.9 s 309 MiB 275 MiB

2) A sample fixture from my v3 cover letter, a worst case with 73% of
lines [unknown]:

time max RSS
v3 eager 3.2 s 334 MiB
v3 lazy 1.9 s 106 MiB
this series 205.3 s 449 MiB

Most of the time goes to reloading. On the fixture, a profile shows it
under map__find_symbol() -> dso__load() -> dso__load_sym(). After each
shrink, the first miss in a shrunk DSO re-parses its whole symtab, and
addresses that resolve to [unknown] miss every time.

The peak is still set by dso__load(), since the whole symtab is
materialized before a shrink can run. That is the main blocker for us
running perf script at scale: we can't bound how much memory it will
use, so we have to contain it with memory.max, and then the profiling
job fails with an OOM kill instead of degrading.

For continuous profiling, a degraded profile is still useful. We build
our view statistically, aggregating many profiles across hosts and
time, so some frames reported as [unknown] in the DSOs that hit the cap
only add a little noise, and the cap reports when it was hit. A lost
profile is worse: the OOM kill removes all its samples and wastes the CPU
and memory already spent. A bounded, reported degradation, like perf record
--max-size, is the trade-off I'd recommend.

The bound also matters before a profile starts. perf runs next to
production workloads, so we have to plan its memory up front: pick a
host with enough headroom and size the job's reservation, without
pushing those sensitive workloads into memory reclaim. Without a bound
we can't do that.

The output of your series matches current eager loading except for a few
frames, and there I think your series is right. For example, a sample inside
nft_do_chain [nf_tables] is printed as dca_sysfs_exit [dca].
symbols__fixup_end() extends the last symbol of a module to
roundup(end + 4096, 4096), which overlaps the next module when it starts
in the next page. The rb-tree lookup can then return the stretched symbol,
while your bsearch picks the right one. This comes from bacefe0c7b77b;
I can send a separate fix that clamps the end to the next symbol's start.

A few questions:

- Which workloads did you have in mind? I'd expect the shrinking to
help most with perf top and long perf report sessions, which I
haven't measured. I'm happy to run those on the same hosts if you
have a setup you'd like me to use.

- Would you consider a finer-grained reload? The index in my patch 5/6
keeps a 24-byte entry per symbol, so a miss after a shrink could
materialize one symbol instead of re-parsing the symtab. The pieces
could fit together: your container and refcounting, the index for
reload, and a byte cap for the peak. If the maintainers prefer your
container first, I'm happy to rebase my v4 onto struct symbols.

- The cover letter says patch 3 saves 48 bytes per symbol (two
rb_nodes). At edd8a9fe2eca0, struct symbol has one rb_node, and the
name index is already a struct symbol ** array on the DSO. Am I
missing a second node?


Thanks,
Alireza