[PATCH bpf 1/2] bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element

From: Donggeun Yoo

Date: Sun Sep 20 2026 - 05:33:45 EST


pcpu_init_value() initializes the per-cpu area of a newly created
[lru_]percpu_hash element. That area is recycled and still holds the
values of whatever element occupied it before, so when the value comes
from a BPF program (onallcpus == false) the function writes the running
CPU's slot and zeroes the rest.

bpf_percpu_hash_update() always passes onallcpus == true, and that arm
calls pcpu_copy_value(), which used to write every CPU. That changed in
commit c6936161fd55 ("bpf: Add BPF_F_CPU and BPF_F_ALL_CPUS flags
support for percpu_hash and lru_percpu_hash maps"): with BPF_F_CPU it
writes the one CPU named in map_flags and returns. On the create
path the remaining slots are left as they were, and a lookup of the new
key hands back the recycled element's values:

update(k1, 0xdeadc0de, BPF_F_ALL_CPUS) every CPU holds 0xdeadc0de
delete(k1) element back on the freelist
update(k2, 0xc0ffee, BPF_F_CPU | 0) creates, writes CPU 0 only
lookup(k2) CPU 0 0xc0ffee, rest 0xdeadc0de

Commit d3bec0138bfb ("bpf: Zero-fill re-used per-cpu map element")
established that a re-used element must not return the previous
tenant's values. BPF_F_CPU is the first way to reach pcpu_init_value()
writing a single CPU with onallcpus set, so extend the zero-filling arm
to cover it, with map_flags >> 32 naming the CPU that receives the
value.

Only creation is affected: pcpu_init_value() is reached from the two
create branches, while an update of an existing element goes straight to
pcpu_copy_value(), where writing one CPU and leaving the others is the
point of the flag.

Fixes: c6936161fd55 ("bpf: Add BPF_F_CPU and BPF_F_ALL_CPUS flags support for percpu_hash and lru_percpu_hash maps")
Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@xxxxxxxxx>
---
Tested on x86_64 under QEMU/KVM against bpf/master a11212910cf0: with
this patch the selftest in 2/2 passes on all three allocation modes,
and without it all three read 0xdeadc0de where they expect 0. Numbers
in the cover letter.

The merged arm no longer calls bpf_obj_cancel_fields() on the named
CPU. That call is inert on this path: it acts only on BPF_TIMER,
BPF_WORKQUEUE and BPF_TASK_WORK, and map_check_btf() rejects all three
for [lru_]percpu_hash.

kernel/bpf/hashtab.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)

diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c
index 4f495dcbf670c..c4683d0e4c149 100644
--- a/kernel/bpf/hashtab.c
+++ b/kernel/bpf/hashtab.c
@@ -1056,12 +1056,12 @@ static void pcpu_init_value(struct bpf_htab *htab, void __percpu *pptr,
* known initial values for cpus other than current one
* (onallcpus=false always when coming from bpf prog).
*/
- if (!onallcpus) {
- int current_cpu = raw_smp_processor_id();
+ if (!onallcpus || (map_flags & BPF_F_CPU)) {
+ int init_cpu = onallcpus ? map_flags >> 32 : raw_smp_processor_id();
int cpu;

for_each_possible_cpu(cpu) {
- if (cpu == current_cpu)
+ if (cpu == init_cpu)
copy_map_value(&htab->map, per_cpu_ptr(pptr, cpu), value);
else /* Since elem is preallocated, we cannot touch special fields */
zero_map_value(&htab->map, per_cpu_ptr(pptr, cpu));
--
2.53.0