Re: [PATCH v1 0/7] perf symbol: Reference counting, flat array storage, and LRU shrinking
From: Namhyung Kim
Date: Tue Sep 29 2026 - 14:10:32 EST
On Mon, Sep 28, 2026 at 05:12:35PM -0700, Alireza Haghdoost wrote:
> On Mon, Sep 28, 2026 at 3:05 PM Ian Rogers <irogers@xxxxxxxxxx> wrote:
> > > I think many of this memory overhead come from the symbol string.
> > > Probably we don't use most of them. Then would it be nice if we can
> > > lazy-load the strings?
>
> Namhyung, your guess is right. I counted live symbol allocations in eager mode
> on the sample fixture. At the peak, 986,688 symbols take 258 MiB:
>
> names (demangled, avg 210 bytes) 197 MiB 77%
> struct symbol (48 bytes each, 45 MiB 18%
> 24 of them the rb_node)
> malloc overhead 15 MiB 6%
>
> That is most of the 335 MiB max RSS. With --lazy-load-symbols, only
> 2,928 symbols are created (0.5 MiB), with identical output, and max
> RSS is 108 MiB.
Thanks for checking this!
>
> > > Without the string (and the priv part), now it has a fixed size so it
> > > should be saved in an array directly.
>
> That is close to what patch 5/6 of my v3 does. Its index is a sorted
> array of fixed-size entries:
>
> struct sym_idx {
> u64 start;
> u64 end;
> u32 name_off;
> u8 binding;
> u8 type;
> u8 flags;
> };
>
> That is 24 bytes per symbol. The name is read from the string table
> and demangled only when a sample lands in the symbol. The difference
> from your description is that v3 then creates a regular struct symbol
> for the hit and inserts it in the rb-tree, because the rest of perf
> holds struct symbol pointers. If the entry itself were the symbol,
> with the name as a pointer/offset union, that extra step would go
> away. Ian's patch 1 would help: once every name access goes through
> symbol__name(), that accessor is the only place that has to resolve
> an offset.
I'm curious if name_off would work well for PLT symbols which come from
the dynamic symbol table. Probably you need to handle them differently.
>
> > > Probably we can add a limit there
> > > and print it with offset when it doesn't have the name.
>
> Agreed, that is better than [unknown]. The symbol boundaries are still
> known, so the frame can be printed as an offset and resolved offline.
> I can do the same for the byte cap in v4.
>
> >
> > We load all symbols instead of lazily because we need the end of a
> > symbol, which we can only determine by processing the entire symbol
> > table as the end of one symbol is the start of the next.
> >
>
> Ian, I think most symbols don't need the next one for their end. ELF
> symbols have st_size, so the end is start + st_size. Only zero-size
> symbols take the next symbol's start. The whole table still has to be
> scanned to find the symbol for an address, because .symtab isn't
> sorted by address, so an address window would scan it again each time
> it grows. The v3 index scans it once, sorts the entries and fixes up
> the ends there, but keeps only the 24-byte entries, not the names.
>
> On which kernel symbol is right in the nf_tables example: the sample
> is at 0xffffffffc0c4b280, and /proc/kallsyms has nft_do_chain at
> 0xffffffffc0c4b0b0 and the next nf_tables symbol at
> 0xffffffffc0c4b4e0, so your series is right. The last dca symbol
> starts at 0xffffffffc0c4aa50, and symbols__fixup_end() extends it to
> 0xffffffffc0c4c000, over the first page of nf_tables. I've sent the
> fix separately:
> https://lore.kernel.org/all/20260928-haghdoost-perf-symbols-fixup-module-end-v1-1-0a70d1edd401@xxxxxxxx/
Thanks for the fix!
Namhyung