Re: [PATCH bpf-next] selftests/bpf: Wait until the .data map of test_mmap is freed

From: bot+bpf-ci

Date: Thu Oct 08 2026 - 22:31:02 EST


> selftests/bpf: Wait until the .data map of test_mmap is freed

TL;DR:
- test_mmap's new unbounded wait may hang test_progs forever if the map leaks;
the watchdog apparently won't kill it
- commit message's watchdog claim and ".data map" wording may be wrong

> diff --git a/tools/testing/selftests/bpf/prog_tests/mmap.c b/tools/testing/selftests/bpf/prog_tests/mmap.c
> index a271d5a0f7ab..0e297ed5da0f 100644
> --- a/tools/testing/selftests/bpf/prog_tests/mmap.c
> +++ b/tools/testing/selftests/bpf/prog_tests/mmap.c
> @@ -280,7 +280,7 @@ void test_mmap(void)
> munmap(tmp1, map_sz);
>
> /* we need to wait for RCU grace period */
> - for (i = 0; i < 10000; i++) {
> + while (true) {
> __u32 id = data_map_id - 1;
> if (bpf_map_get_next_id(id, &id) || id > data_map_id)
> break;

Can this loop spin forever if the map is never freed?

The commit message relies on the watchdog to end the test:

> Poll until the map ID is gone. The watchdog of test_progs kills the
> test if the map is never freed.

Looking at watchdog_timer_func() in test_progs.c, the WD_KILL case only does:

pthread_kill(env.main_thread, SIGSEGV);

It does not re-arm the timer, so it fires once. main() installs the handler
with:

struct sigaction sigact = {
.sa_handler = crash_handler,
.sa_flags = SA_RESETHAND,
};

crash_handler() calls dump_crash_log(), prints a backtrace and returns. It
does not call exit() or abort(). Since this SIGSEGV is sent by
pthread_kill() and is not a bad memory access, no faulting instruction runs
again after the handler returns, and SA_RESETHAND only resets the handler to
the default.

So the thread goes back into the loop, usleep() returns early with EINTR,
which the loop ignores, and polling carries on. The sample output in commit
d9d4d127e813 shows the same behaviour: after "Caught signal #11!" the
send_signal test keeps running and reports "#292 send_signal:FAIL". It only
stopped there because its blocking read() returned EINTR.

If the map is never freed, test_mmap would then loop forever. That is the
mmap/prog/map reference leak this check is meant to catch. Before the patch,
the same case failed after about 10000 iterations with "get_map_by_id_after
unexpectedly succeeded".

test_mmap is a parallel test, so with -j, dispatch_thread() would block in
read_prog_test_msg() waiting for MSG_TEST_DONE from the stuck worker and the
whole test_progs run never finishes. In serial mode every later test is
blocked. With -w 0 the watchdog is disabled and there is no message at all.
In a non-ASAN build the watchdog only prints a stack trace after 120s and the
hang continues.

Would a bounded wait with a much longer limit avoid the flake and still turn
a leak into a test failure? For example,
tools/testing/selftests/bpf/prog_tests/tramp_prog_detach.c:wait_for_map_free()
polls bpf_map_get_fd_by_id() up to 100 times, 100ms apart.

Either way, should the sentence "The watchdog of test_progs kills the test if
the map is never freed." be corrected in the commit message?

This isn't a bug, but the subject and body call this "the .data map". The
map in question is the user-defined BPF_F_MMAPABLE array named data_map in
SEC(".maps"), not a .data section map. Could the wording be adjusted?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37873564350