Re: [PATCH bpf-next v3 4/9] bpf: take the vmlinux BTF from the btf_vmlinux module
From: Jay Wang
Date: Thu Oct 01 2026 - 19:52:40 EST
On Sat, Sep 26, 2026 at 08:29:12AM +0000, Alexei Starovoitov wrote:
> On Fri, Sep 25, 2026 at 10:42 PM Jay Wang <wanjay@xxxxxxxxxx> wrote:
> > + if (!data && load) {
> > + /*
> > + * The module notifier installs the BTF before init_module()
> > + * returns, so it is either there after this or the module is
> > + * not available (yet). Not cached: a later call retries,
> > + * e.g. once the module becomes reachable on the root fs.
> > + */
> > + request_module("btf_vmlinux");
>
> This will deadlock.
> event_btf_ids_read() calls btf_get_module_btf(NULL) with event_mutex
> held. With =m and BTF not loaded yet
> cat /sys/kernel/tracing/events/sched/sched_switch/btf_ids
> gets here and waits for modprobe. modprobe gets to
> trace_module_notify() which takes event_mutex.
> lockdep doesn't see it.
>
> print_function_args() gets here for every line of the trace,
> from ftrace_dump() with irqs off too. When the module is not
> installed that is one modprobe per line.
>
> Every caller of bpf_get_btf_vmlinux() and bpf_find_btf_id() was
> written for a function that doesn't wait for user space.
Right. Given that fixing those callers one by one does not work, as
there are too many and every new one would have to know, I went with a
different mechanism in v4:
https://lore.kernel.org/bpf/20261001225214.12351-1-wanjay@xxxxxxxxxx/
In short, the lookups no longer load the BTF; only requests from user
space do, at their start.
The lookups cannot wait for the module, because they run in places that
must not sleep, or that hold locks the module load needs too: under
event_mutex in your btf_ids case, with interrupts off in ftrace_dump(),
inside a running BPF program. Loading the module means waiting for
modprobe, so waiting in any of them can deadlock or sleep where it must
not.
Loading at the start of a request from user space is enough, because
that is where every need for the BTF begins: a program or map that uses
kernel types, a read of the BTF file, a probe event with BTF arguments.
At that point the request holds nothing yet, so it can safely wait for
modprobe. And once it has loaded the BTF, the lookups it makes later
find it there, so they have no reason to wait. Only the kernel's own
early users, the kfunc and struct_ops registrations at boot and the BTF
of modules loaded before it, come before any such request, and they are
queued until the BTF arrives.
With =m, bpf_get_btf_vmlinux() and bpf_find_btf_id() never load the
module and never sleep. Until the BTF is loaded they fail as on a
kernel without BTF (NULL, -EINVAL), so every existing caller keeps the
assumptions it was written with. The only function that loads the BTF
is a new bpf_load_btf_vmlinux(), called at the start of such a request,
with nothing held that loading a module needs:
- on entry to the bpf() syscall: a BPF_PROG_LOAD, BPF_MAP_CREATE or
BPF_BTF_LOAD that failed while the BTF was missing is run once more
after loading it, the way tc and nf_tables retry after loading a
module. The verifier itself never waits, and bpf_sys_bpf() does not
go through that entry;
- BPF_BTF_GET_NEXT_ID with CAP_SYS_ADMIN, and loading a syscall
program, since a light skeleton loader loads its programs while it
runs;
- read() of /sys/kernel/btf/vmlinux. Not mmap(): it runs with the
caller's mmap_lock held, and a uprobe registration holds event_mutex
while it takes the mmap_lock of every mm mapping the probed file,
which would close a loop through the module load. mmap() fails
until the BTF is loaded, and libbpf then falls back to read();
- read() of the sysfs file of a module built against a distilled base
(.BTF.base), through a work item: the reader holds the file's kernfs
active reference, which MODULE_STATE_GOING drains with the module
notifier chain held;
- reading btf_ids (before event_mutex), creating probe events with BTF
arguments (the parser holds only dyn_event_ops_mutex, which no module
load takes), and mounting bpffs with delegate options that name
commands or types.
print_function_args() only uses the BTF if it is already loaded and
never waits for it, so no modprobe per line. Both of your cases now
work as the first BTF user after boot: the cat of btf_ids loads the BTF
before taking event_mutex and prints the ids, and func-args with sysrq-z
prints the functions without arguments, as on a kernel without BTF; the
arguments show up once a request from user space has loaded the BTF.
Jay