Re: [PATCH v1 0/7] perf symbol: Reference counting, flat array storage, and LRU shrinking

From: Ian Rogers

Date: Mon Sep 28 2026 - 16:58:10 EST


On Mon, Sep 28, 2026 at 12:45 PM Alireza Haghdoost <haghdoost@xxxxxxxx> wrote:
>
> On Mon, Sep 28, 2026 at 12:52 AM Ian Rogers <irogers@xxxxxxxxxx> wrote:
> > annotations. Saves 48 bytes per struct symbol (two rb_nodes) and provides
> > cache-friendly binary search (bsearch) and callback-based iteration.
>
> Hi Ian,
>
> Thanks for posting this and for Cc'ing me. I built it on edd8a9fe2eca0,
> the same base as the v3 series I sent last week, and compared it with
> v3 in eager mode (today's behavior) and with --lazy-load-symbols.
>
> All runs use perf script --no-inline --max-stack 127
> -F comm,tid,time,ip,sym,dso, with one warm-up and three runs; the
> median is shown below. Peak anon is the maximum RssAnon from
> /proc/PID/status, sampled every 10 ms.
>
> 1) Production: 120 s cgroup profile of a storage service on an 80-CPU
> host (perf record -a -g -F 99 -G <cgroup>). 54k samples, 11 DSOs,
> 5% of lines [unknown].
>
> time max RSS peak anon
> v3 eager 4.5 s 386 MiB 314 MiB
> v3 lazy 3.7 s 151 MiB 80 MiB
> this series 95.9 s 309 MiB 275 MiB
>
> 2) A sample fixture from my v3 cover letter, a worst case with 73% of
> lines [unknown]:
>
> time max RSS
> v3 eager 3.2 s 334 MiB
> v3 lazy 1.9 s 106 MiB
> this series 205.3 s 449 MiB
>
> Most of the time goes to reloading. On the fixture, a profile shows it
> under map__find_symbol() -> dso__load() -> dso__load_sym(). After each
> shrink, the first miss in a shrunk DSO re-parses its whole symtab, and
> addresses that resolve to [unknown] miss every time.
>
> The peak is still set by dso__load(), since the whole symtab is
> materialized before a shrink can run. That is the main blocker for us
> running perf script at scale: we can't bound how much memory it will
> use, so we have to contain it with memory.max, and then the profiling
> job fails with an OOM kill instead of degrading.

Hi Alireza, and thanks for the feedback! I'll see what can be done
about dso__load in v2.

> For continuous profiling, a degraded profile is still useful. We build
> our view statistically, aggregating many profiles across hosts and
> time, so some frames reported as [unknown] in the DSOs that hit the cap
> only add a little noise, and the cap reports when it was hit. A lost
> profile is worse: the OOM kill removes all its samples and wastes the CPU
> and memory already spent. A bounded, reported degradation, like perf record
> --max-size, is the trade-off I'd recommend.

I don't disagree. Fwiw, a C++ memory allocation failure inside my
company's code base is a runtime abort. In perf, many code paths
return -ENOMEM or NULL. I wanted to see how far we could get with a
"shrunk" set of data inside perf as even if we fix memory allocation
around symbols, other code paths can still fail.

> The bound also matters before a profile starts. perf runs next to
> production workloads, so we have to plan its memory up front: pick a
> host with enough headroom and size the job's reservation, without
> pushing those sensitive workloads into memory reclaim. Without a bound
> we can't do that.

I'm still not massively sold on having explicit bounds. It reminds me
of old and often seemingly arbitrary numbers like the number of rows
or columns in a spreadsheet. That said, perf has a similar problem
with exhausting file descriptors. As long as the default is
"unlimited," I'm okay with some attempts to limit size, but we should
implement obvious garbage collection as I hope this series introduces.

> The output of your series matches current eager loading except for a few
> frames, and there I think your series is right. For example, a sample inside
> nft_do_chain [nf_tables] is printed as dca_sysfs_exit [dca].
> symbols__fixup_end() extends the last symbol of a module to
> roundup(end + 4096, 4096), which overlaps the next module when it starts
> in the next page. The rb-tree lookup can then return the stretched symbol,
> while your bsearch picks the right one. This comes from bacefe0c7b77b;
> I can send a separate fix that clamps the end to the next symbol's start.

That's surprising as no behavioral change was intended. That then
leads to the question of which symbol is right and whether a latent
bug is being fixed or a new one introduced. Oh what fun :-)

> A few questions:
>
> - Which workloads did you have in mind? I'd expect the shrinking to
> help most with perf top and long perf report sessions, which I
> haven't measured. I'm happy to run those on the same hosts if you
> have a setup you'd like me to use.

So most of my experience with excessive memory usage has come from
perf top, which seems to leak memory and consume all your RAM before
long. That is expected since no DSO is currently unloaded. I'm hoping
the shrinking process can be generic.

> - Would you consider a finer-grained reload? The index in my patch 5/6
> keeps a 24-byte entry per symbol, so a miss after a shrink could
> materialize one symbol instead of re-parsing the symtab. The pieces
> could fit together: your container and refcounting, the index for
> reload, and a byte cap for the peak. If the maintainers prefer your
> container first, I'm happy to rebase my v4 onto struct symbols.

Definitely. I was more focused on the reference counting aspect.
Things like the policy around shrinking can be amended as needed, I
just wanted to establish a policy so the patches made sense. I think
lazier loading, shrinking, .. all make sense for things to be faster
and use less RAM.

> - The cover letter says patch 3 saves 48 bytes per symbol (two
> rb_nodes). At edd8a9fe2eca0, struct symbol has one rb_node, and the
> name index is already a struct symbol ** array on the DSO. Am I
> missing a second node?

No, I'll fix this mistake.

Thanks,
Ian

> Thanks,
> Alireza