Re: [PATCH v7 0/3] kallsyms: Accelerate symbol name lookups by ~7x

From: Andrew Morton

Date: Tue Sep 29 2026 - 14:46:42 EST


On Tue, 29 Sep 2026 12:07:29 -0600 Jim Cromie via B4 Relay <devnull+jim.cromie.gmail.com@xxxxxxxxxx> wrote:

> In 2022, commit 60443c88f3a8 ("kallsyms: Improve the performance of
> kallsyms_lookup_name()") introduced kallsyms_seqs_of_names[] (+550 KiB
> .rodata), transforming an O(N) linear scan into an O(log N) binary
> search (5.2 ms -> ~7.2 us). While this was a major step forward,
> the binary search inner loop was left decompressing full candidate
> names and scanning across sparse 256:1 markers on every probe.
>
> Modern fleet observability, security daemons (e.g. CrowdStrike Falcon,
> Cilium, Datadog, Falco), and tracing tools resolve thousands of kernel
> functions by name at boot or service start.
>
> CrowdStrike recently hit this in production:
>
> commit 93e8fd1a565e ("ftrace: Use kallsyms binary search for single-symbol lookup")
>
> Attaching just 50 kprobe.session programs caused an 858 ms attach
> stall with 25% CPU burned in kallsyms. That commit routed single-symbol
> libbpf attach directly to kallsyms_lookup_name(). In larger workloads
> (such as the BPF selftest serial_test_kprobe_multi_bench_attach across
> 64,000 symbols), kallsyms_lookup_names() spends ~390 ms in raw CPU spin.
>
> This series accelerates kallsyms_lookup_names() by 7.0x (from 6,102 ns
> down to 866 ns per lookup), cutting 64k-symbol attach from ~390 ms to
> ~55 ms, by fixing two inner-loop bottlenecks:

Thanks, this is awesome.

> 0. Candidate symbols are fully decompressed into a 512-byte stack buffer
> before calling strcmp(), even though ~16 of the 17 search steps
> mismatch at the first differing character (0..N-1, heavily
> front-loaded toward 0-2).
>
> 1. Probes scan sequentially from 256:1 markers in kallsyms_names[],
> decoding an average of 127.5 symbols per probe (~2,170 hops across a
> 17-step search).
>
> The 3-patch progression:
>
> 0. Patch 1 introduces kallsyms_strcmp_symbol() to compare ASCII queries
> against compressed tokens on the fly, bailing out on first mismatch.
> Drops the 512-byte stack buffer and saves ~530 ns.
>
> 1. Patch 2 increases marker density from 256:1 to 16:1, cutting average
> scan distance from 127.5 to 7.5 hops and dropping lookup latency from
> 6,102 ns to 866 ns for +42.2 KiB of .rodata.
>
> 2. Patch 3 inlines and unrolls get_symbol_seq() 24-bit reconstruction.
>
> Results (CONFIG_KALLSYMS_SELFTEST across ~184k symbols):
> - Baseline (256:1): 6,102 ns
> - Patch 1 (strcmp): 5,572 ns (-530 ns)
> - Patch 2 (16:1): 866 ns (7.0x faster)
>
> Trade-offs:
> - .rodata footprint: +42.2 KiB (+10,782 u32 entries for ~184k symbols,
> ~0.1% of loaded kernel image).

That's the bad news. Can anything be done to reduce this aspect?

> ---
> Jim Cromie (3):
> kallsyms: Match compressed tokens on the fly during binary search
> kallsyms: Increase marker density to 16:1 to accelerate lookups
> kallsyms: Unroll 24-bit sequence reconstruction in get_symbol_seq()

I'll queue this in mm.git for testing, would prefer not to go further
without third-party review, please.