Re: [PATCH v1 0/7] perf symbol: Reference counting, flat array storage, and LRU shrinking
From: Alireza Haghdoost
Date: Mon Sep 28 2026 - 20:13:00 EST
On Mon, Sep 28, 2026 at 3:05 PM Ian Rogers <irogers@xxxxxxxxxx> wrote:
> > I think many of this memory overhead come from the symbol string.
> > Probably we don't use most of them. Then would it be nice if we can
> > lazy-load the strings?
Namhyung, your guess is right. I counted live symbol allocations in eager mode
on the sample fixture. At the peak, 986,688 symbols take 258 MiB:
names (demangled, avg 210 bytes) 197 MiB 77%
struct symbol (48 bytes each, 45 MiB 18%
24 of them the rb_node)
malloc overhead 15 MiB 6%
That is most of the 335 MiB max RSS. With --lazy-load-symbols, only
2,928 symbols are created (0.5 MiB), with identical output, and max
RSS is 108 MiB.
> > Without the string (and the priv part), now it has a fixed size so it
> > should be saved in an array directly.
That is close to what patch 5/6 of my v3 does. Its index is a sorted
array of fixed-size entries:
struct sym_idx {
u64 start;
u64 end;
u32 name_off;
u8 binding;
u8 type;
u8 flags;
};
That is 24 bytes per symbol. The name is read from the string table
and demangled only when a sample lands in the symbol. The difference
from your description is that v3 then creates a regular struct symbol
for the hit and inserts it in the rb-tree, because the rest of perf
holds struct symbol pointers. If the entry itself were the symbol,
with the name as a pointer/offset union, that extra step would go
away. Ian's patch 1 would help: once every name access goes through
symbol__name(), that accessor is the only place that has to resolve
an offset.
> > Probably we can add a limit there
> > and print it with offset when it doesn't have the name.
Agreed, that is better than [unknown]. The symbol boundaries are still
known, so the frame can be printed as an offset and resolved offline.
I can do the same for the byte cap in v4.
>
> We load all symbols instead of lazily because we need the end of a
> symbol, which we can only determine by processing the entire symbol
> table as the end of one symbol is the start of the next.
>
Ian, I think most symbols don't need the next one for their end. ELF
symbols have st_size, so the end is start + st_size. Only zero-size
symbols take the next symbol's start. The whole table still has to be
scanned to find the symbol for an address, because .symtab isn't
sorted by address, so an address window would scan it again each time
it grows. The v3 index scans it once, sorts the entries and fixes up
the ends there, but keeps only the 24-byte entries, not the names.
On which kernel symbol is right in the nf_tables example: the sample
is at 0xffffffffc0c4b280, and /proc/kallsyms has nft_do_chain at
0xffffffffc0c4b0b0 and the next nf_tables symbol at
0xffffffffc0c4b4e0, so your series is right. The last dca symbol
starts at 0xffffffffc0c4aa50, and symbols__fixup_end() extends it to
0xffffffffc0c4c000, over the first page of nf_tables. I've sent the
fix separately:
https://lore.kernel.org/all/20260928-haghdoost-perf-symbols-fixup-module-end-v1-1-0a70d1edd401@xxxxxxxx/
Thanks
Alireza