Re: [PATCH v5 2/3] kallsyms: Increase marker density to 16:1 to accelerate lookups

From: bot+bpf-ci

Date: Fri Sep 25 2026 - 20:49:05 EST


> kallsyms: Increase marker density to 16:1 to accelerate lookups
>
> kallsyms_lookup_names() resolves symbol names to addresses using a
> binary search over kallsyms_seqs_of_names[]. In baseline, each probe
> invokes get_symbol_offset() to find the symbol in kallsyms_names[].
> Because markers in kallsyms_markers[] are spaced every 256 symbols in
> address order, each probe must sequentially scan and decode ULEB128
> lengths across an average of 127.5 symbols from the nearest marker
> (~2,176 hops across a 17-step binary search).
>
> During bulk symbol resolution (e.g. BPF multi-kprobe and tracing-multi
> attach across tens of thousands of functions), this linear scan penalty
> compounds into substantial latency (~4.7 us to ~6.1 us per lookup).
>
> Increase the static marker density from 256:1 down to 16:1 (1 marker
> every 1 << 4 symbols):
>
> 0. Define KALLSYMS_MARKER_SHIFT as 4 in kernel/kallsyms_internal.h, with
> a matching KALLSYMS_MARKER_MASK of 0x0F. Guard kernel types so host
> scripts/kallsyms.c can include this header directly as the single
> symbolic source of truth.
>
> 1. In scripts/kallsyms.c, emit markers every
> (1 << KALLSYMS_MARKER_SHIFT) symbols into kallsyms_markers[]. For a
> typical kernel with ~184,000 symbols, this increases marker count
> from 719 to 11,501 entries, adding only +42.2 KiB to write-protected
> .rodata (0.002% of vmlinux).
>
> 2. In kernel/kallsyms.c:get_symbol_offset(), compute marker offset via
> pos >> KALLSYMS_MARKER_SHIFT and step through
> pos & KALLSYMS_MARKER_MASK. Because both are compile-time
> constants, GCC emits single bit-shift and AND instructions with zero
> division overhead.
>
> This reduces the maximum sequential scan from 255 down to 15 symbols,
> and cuts the average scan from 127.5 down to 7.5 hops (a 17x reduction
> in sequential loop hops). In-tree selftest measurements across all
> 184,008 symbols show lookup latency dropping from 6,102 ns down to
> 866 ns (a 7.0x speedup) with zero runtime memory allocation, zero RCU
> synchronization, and zero new user-facing APIs or Kconfig options.

This isn't a bug, but could the numbered list be simplified? The items
0-2 restate the diff hunk by hunk, and some of the remarks describe
things the patch doesn't change.

The old code already used 'pos >> 8' and 'pos & 0xFF', so 'zero
division overhead' isn't new to this patch. Similarly, 'zero RCU
synchronization' and 'zero new user-facing APIs' describe things the
patch never touches.

Also, is 0.002% the intended percentage for the extra .rodata? That
would imply a vmlinux of about 2 GiB, but a typical loaded image is
tens of MiB, making the actual percentage closer to 0.1-0.2%.

The useful parts are the lookup-latency motivation and the
memory-versus-speed trade-off. Would dropping the numbered list in
favour of just the rationale make the description clearer?

> diff --git a/kernel/kallsyms_internal.h b/kernel/kallsyms_internal.h
> index 81a867dbe57d..3e6494b7dfc4 100644
> --- a/kernel/kallsyms_internal.h
> +++ b/kernel/kallsyms_internal.h
> @@ -1,7 +1,11 @@
> /* SPDX-License-Identifier: GPL-2.0-only */
> #ifndef LINUX_KALLSYMS_INTERNAL_H_
> #define LINUX_KALLSYMS_INTERNAL_H_
> +#define KALLSYMS_MARKER_SHIFT 4 /* 16:1 sweet spot: +42 KiB .rodata, 17x fewer hops */
> +#define KALLSYMS_MARKER_SIZE (1U << KALLSYMS_MARKER_SHIFT)
> +#define KALLSYMS_MARKER_MASK (KALLSYMS_MARKER_SIZE - 1U)
>
> +#ifdef __KERNEL__
> #include <linux/types.h>

This isn't a bug, but would a short comment explaining the trade-off be
enough here, without the specific measurements? The trailing comment
hard-codes benchmark figures from one configuration (about 184k
symbols). The .rodata cost grows with the symbol count, so '+42 KiB'
will be wrong for other configs and will go stale as the kernel grows.

The same measurements are already in the commit message, where they
usually go. Nothing else in this header has comments like this.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36204745600