Re: [PATCH bpf] bpf: Fix smp_processor_id() warning in rhtab_delete_elem()

From: bot+bpf-ci

Date: Mon Sep 28 2026 - 14:21:28 EST


> bpf: Fix smp_processor_id() warning in rhtab_delete_elem()
>
> bpf_mem_cache_free_rcu() uses this_cpu_ptr() which requires preemption
> to be disabled. In rhtab_delete_elem(), this call happens after
> bpf_enable_instrumentation(), so preemption is re-enabled at that
> point and this_cpu_ptr() triggers:
>
> BUG: using smp_processor_id() in preemptible [00000000] code
>
> Fix by moving all post-delete cleanup — rhtab_read_elem_value(),
> check_and_init_map_value(), bpf_obj_cancel_fields(), and
> bpf_mem_cache_free_rcu() — inside the bpf_disable_instrumentation()
> section, before bpf_enable_instrumentation() is called.
>
> This is consistent with __htab_map_lookup_and_delete_batch() which
> wraps the entire batch operation including element freeing under
> bpf_disable_instrumentation().
>
> Fixes: 6905f8601298 ("bpf: Allow special fields in resizable hashtab")
> Reported-by: syzbot+fd7e415d891073b83e1f@xxxxxxxxxxxxxxxxxxxxxxxxx
> Closes: https://syzkaller.appspot.com/bug?extid=fd7e415d891073b83e1f

Is the explanation of how the fix works accurate? The commit message says
bpf_enable_instrumentation() "re-enables preemption," but looking at
include/linux/bpf.h, bpf_disable_instrumentation() is:

static inline void bpf_disable_instrumentation(void)
{
migrate_disable();
this_cpu_inc(bpf_prog_active);
}

and bpf_enable_instrumentation() is:

static inline void bpf_enable_instrumentation(void)
{
this_cpu_dec(bpf_prog_active);
migrate_enable();
}

These functions disable and enable migration, not preemption. After the
patch, bpf_mem_cache_free_rcu() still runs with preemption enabled on the
batch path under CONFIG_PREEMPT_RCU.

The fix works because check_preemption_disabled() in
lib/smp_processor_id.c accepts current->migration_disabled as one of its
early exits, and unit_free_rcu() guards its per-CPU list update with
local_irq_save(). This makes the code both warning-free and correct.

Should the message say that this_cpu_ptr() needs preemption or migration
disabled, and that the fix keeps the free inside the migrate_disable()
section? Otherwise readers might assume the section is non-preemptible,
which it is not.

Also, the message does not identify which caller reaches rhtab_delete_elem()
without migration protection. Looking at the callers:

- rhtab_map_delete_elem() is reached from map_delete_elem()
(kernel/bpf/syscall.c) which wraps the call in
bpf_disable_instrumentation()
- rhtab_map_lookup_and_delete_elem() is reached from
map_lookup_and_delete_elem() (kernel/bpf/syscall.c) which also wraps
the call in bpf_disable_instrumentation()
- BPF program callers run with migration already disabled

The only unprotected caller is __rhtab_map_lookup_and_delete_batch()
(kernel/bpf/hashtab.c), which calls rhtab_delete_elem() in a loop while
holding only rcu_read_lock(). This path reaches rhtab_delete_elem()
through BPF_MAP_LOOKUP_AND_DELETE_BATCH -> bpf_map_do_batch() ->
__rhtab_map_lookup_and_delete_batch().

Could the message name this specific path to clarify that only the batch
operation is affected, not every rhtab_delete_elem() call?

Does the Fixes tag point at the right commit? Looking at 6905f8601298
("bpf: Allow special fields in resizable hashtab"), its diff shows
bpf_mem_cache_free_rcu() and __rhtab_map_lookup_and_delete_batch() only
as unchanged context lines. That commit added rhtab_check_and_free_fields()
but did not change the ordering of operations.

In 16b4d3e2fb24 ("bpf: Implement resizable hashmap basic functions"),
rhtab_delete_elem() already had this structure:

bpf_disable_instrumentation();
err = rhashtable_remove_fast(&rhtab->ht, &elem->node, rhtab_params);
bpf_enable_instrumentation();
...
bpf_mem_cache_free_rcu(&rhtab->ma, elem);

At that commit, every caller was still wrapped in bpf_disable_instrumentation():
map_delete_elem() and map_lookup_and_delete_elem() in syscall.c both wrap
their calls, and BPF programs run with migration disabled.

The first caller that reaches rhtab_delete_elem() holding only
rcu_read_lock() came with 818e00848227 ("bpf: Implement iteration ops for
resizable hashtab"). That commit added __rhtab_map_lookup_and_delete_batch(),
which does:

for (i = 0; i < total; i++)
rhtab_delete_elem(rhtab, del_elems[i], NULL, 0);

under a plain rcu_read_lock(). On CONFIG_PREEMPT_RCU, rcu_read_lock() only
bumps rcu_read_lock_nesting and leaves preempt_count alone, so the task
stays preemptible and this_cpu_ptr() fires the warning.

Should the tag be:

Fixes: 818e00848227 ("bpf: Implement iteration ops for resizable hashtab")

This would tie the fix to the commit that made the bug reachable, so
backport tooling knows where the fix should land.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36460573262