Re: [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory
From: Ihor Solodrai
Date: Fri Oct 02 2026 - 16:59:04 EST
Hi Jay, thank you for taking the time to reply.
On 2026-10-02 12:34 a.m., Jay Wang wrote:
A question I had: is it really worth 2k of complicated kernel code to
*maybe sometimes* save <10Mb of memory?
There are many reasons this is worth it. The main ones:
1. This did not start from a number we picked, but from a real use
case: users who run large fleets of small instances, 1 GiB of
memory or less, with their workloads sized to fit. We are not able
to disclose more detail, but for them 5.4 MB on every instance is
real money, and it can be exactly what pushes a workload over its
memory budget and onto the next instance size. That is why we took
this on, even though we knew it would not be a small change.
Ok, good to know there is a real customer for this.
Obviously I don't know any details of the use-case apart from what
you've shared. But bear with me trying the customer's hat on:
- I am running a workload on a legion of 1GiB instances
- I care about memory footprint a lot, because it directly
translates into the amount of compute I pay for
- If I know I don't need BPF on the workload, and these 5.4 MB
bite me, I can just turn off BPF
- If I know I need BPF, then those 5.4 MB must be loaded anyways, so
I have to look for savings elsewhere
What I find strange is for a user of this scale *to not know* whether
BPF will be used in the workload or not, which is what lazy load may
solve.
One thing I can imagine is an opaque workload: untrusted (AI agents
etc), end-user-defined (including BPF usage), or confidential. But in
these cases, what are the chances that BPF is used there? They are
high, BPF is very widespread.
I can also imagine normal workload not needing BPF, but when something
goes wrong an observability or security thing turning on that loads
BPF programs. But this contradicts the "5.4 MB may push workload over
1GiB" premise: you still need to reserve/swap memory for just-in-time
BPF-based tools, otherwise you OOM.
I of course may be missing something, happy to be corrected.
2. Loading code and data only when a system needs them is what kernel
modules exist for. And most of them take well under ~1 MB once
loaded (nf_conntrack, overlay, vfat); even big ones like ext4, kvm
and btrfs stay around 1-2.5 MB. The vmlinux BTF is 5.4 MB.
True. At the same time BPF without vmlinux BTF is barely useful. And
it is trusted by the verifier. Both are addressed by BTF being
built-in in the kernel image.
3. We are not alone in trying to keep BTF out of memory until it is
needed. The inline BTF work [1] plans to deliver its data, which
is even larger, through a module too. The .BTF.link record and the
resolve_btfids option that serve both are already part of this
series (patch 10), so this is not machinery for the vmlinux BTF
alone.
The inline data is different. It's bigger by it's nature, and it has
more specific users: tracing tools. The userspace tools are able to
themselves decide whether to use that data. The kernel only needs to
make it available.
In comparison, as you yourself noted in the cover:
CO-RE, fentry/fexit, kfuncs, struct_ops, sched_ext and bpf-lsm
all depend on vmlinux BTF
Do those tiny VMs that you target run systemd with BPF LSM?
Not necessarily, and that is the image's choice. Even with BPF LSM
built in and bpf in CONFIG_LSM, memory-conscious users can turn it off
at boot with an lsm= list that leaves it out, and then systemd does not
load restrict_fs at all.
And even where user space is in the way, which as above it does not
have to be, the answer is to fix user space so it gets this win too,
not to give up on it.
If the answer is "yes" for the majority of them,
"The majority" is also hard to pin down here: what runs at boot
depends on the user space packages built into each image, and on the
workload.
Yes, it is hard to pin down. Which is why we first need to figure out
whether the alleged memory savings actually help anyone.
Adding code to the kernel is not free. More complexity means bigger
bug surface: more opportunities for concurrency bugs, security bugs
etc. Especially runtime loads with retries and stuff. I'm sure you
understand the future cost of all that.
So the bias shouldn't be "let's implement a big thing that might
hypothetically help someone". It's backwards IMO.
So the question should be what it takes to make this work, not whether
to drop it because it is complicated. If there are ways to cut it
down, or issues found in the implementation, we would be happy to take
them and go through every one.
The question is in the trade-off.
Would you run a million line python program to search for a string in
a text file? No, that's absurd. But you could.
I'm not saying 2k line patch series is unacceptable in principle. It
may be justified. But I am not convinced it is justified in this case.
Let's say we (as in kernel devs) decided that we want to reduce
vmlinux BTF size, and only load it when necessary. Here are a couple
of alternatives, simpler in comparison to CONFIG_DEBUG_INFO_BTF=m,
although still a bit complex:
* Compress it.
$ ls -lah /sys/kernel/btf/vmlinux
-r--r--r-- 1 root root 6.7M Jul 18 01:54 /sys/kernel/btf/vmlinux
$ tar --zstd -cf /tmp/vmlinux.btf.zstd -C /sys/kernel/btf vmlinux
$ ls -lah /tmp/vmlinux.btf.zstd
-rw-r--r-- 1 isolodrai users 2.2M Oct 2 10:11 /tmp/vmlinux.btf.zstd
3x win right there.
Keep zstd-compressed blob in the kernel image and decompress and
parse it synchronously on first use.
Here is a prototype (vibe-coded, obviously):
https://github.com/kernel-patches/bpf/compare/bpf-next_base...theihor:bpf:vmlinux.btf.zstd-20261002
* Use an external blob.
Ship /lib/modules/$r/vmlinux.btf, checked against a SHA-256
recorded in the image. The catch is a synchronous file I/O under
caller locks, so this is less straightforward.
But this should cover embedded case that Alan has mentioned.
Here is a prototype:
https://github.com/kernel-patches/bpf/compare/bpf-next_base...theihor:bpf:vmlinux.btf.external-20261002
Both alternatives avoid kernel-module loading machine. If we think
harder, we may come up with something even more clever.
I don't particularly prefer one approach over the other, including the
module one. Maintainers probably have a better intuition on that.
The point I'm trying to make is that if you could achieve similar
memory savings with say 500-line well encapsulated diff (or even
better: by refactoring, simplifying, deleting code) it would have
already landed.
And if it was a 50-line patch, the discussion on whether anyone cares
wouldn't be relevant, because it's so cheap.
As it stands, it's not clear (to me at least) in what circumstances
we'll get the promised memory savings at all.
[1] https://lore.kernel.org/bpf/20260916074118.1007116-1-alan.maguire@xxxxxxxxxx/