[PATCH bpf-next v2] bpf: Avoid soft lockup in __htab_map_lookup_and_delete_batch()
From: Josef Bacik
Date: Wed Sep 09 2026 - 14:36:24 EST
From: "Jose Fernandez (Anthropic)" <jose.fernandez@xxxxxxxxx>
__htab_map_lookup_and_delete_batch() has no rescheduling point. The
batch count bounds how many entries are copied out, not how many
buckets are visited, so one BPF_MAP_LOOKUP_BATCH call can walk the
map end to end. The empty-bucket fast path is worse: it stays inside
a single rcu_read_lock() / bpf_disable_instrumentation() section for
any run of consecutive empty buckets.
That holds up on small maps, but it falls apart at scale. On a
144-CPU arm64 host running a CONFIG_PREEMPT_NONE kernel, periodic
BPF_MAP_LOOKUP_BATCH calls against an LRU hash map with 16,777,216
buckets held a CPU inside the batch op for 77+ seconds and triggered
the soft lockup watchdog.
Commit 75134f16e7dd ("bpf: Add schedule points in batch ops") fixed this
same problem in the generic batch ops, but not in this htab-native path,
which every htab-based hash map variant uses for its lookup[_and_delete]
batch ops.
Complete that fix here. Leave the critical section after 64 consecutive
empty buckets, call cond_resched_tasks_rcu_qs(), and resume at the saved
bucket cursor. No locks are held at that point, and resuming from the
cursor is already the function's behavior for non-empty buckets. Add the
same call to the per-bucket loop after copy_to_user(), where every lock
has been dropped. cond_resched_rcu() is not enough here: sleeping with
bpf_prog_active elevated makes tracing programs on that CPU silently
skip their invocations.
Plain cond_resched() is not enough either. It is a no-op under PREEMPT
and PREEMPT_LAZY, the only models arm64 and x86 have offered since
commit 7dadeaa6e851 ("sched: Further restrict the preemption modes").
It is also never a Tasks RCU quiescent state, in any model: the
reschedule counts as a preemption. The walking task stays a holdout and
stalls every synchronize_rcu_tasks() caller, ftrace and BPF trampoline
teardown included, until the syscall returns [1].
cond_resched_tasks_rcu_qs() is the usual tool for that [2]. It reports
the quiescent state at each yield and still reschedules as
cond_resched() does on PREEMPT_NONE and PREEMPT_VOLUNTARY kernels.
Fixes: 057996380a42 ("bpf: Add batch ops to all htab bpf map")
Cc: "Paul E. McKenney" <paulmck@xxxxxxxxxx>
Cc: Rik van Riel <riel@xxxxxxxxxxx>
Link: https://lore.kernel.org/bpf/20260715215314.44423f47@fangorn/ [1]
Link: https://lore.kernel.org/bpf/9d444098-7c03-4163-af12-bd0a79a51443@paulmck-laptop/ [2]
Assisted-by: LLM
Signed-off-by: Jose Fernandez (Anthropic) <jose.fernandez@xxxxxxxxx>
Signed-off-by: Josef Bacik <josef@xxxxxxxxxxxxxx>
---
Changes in v2:
- Use cond_resched_tasks_rcu_qs() at both yield points so the walk
also reports a Tasks RCU quiescent state, and explain why in the
commit message
- Put the opening /* of both comments on its own line (sashiko review
on v1)
- Cc Paul E. McKenney and Rik van Riel, whose thread the message cites
- Rebase onto bpf-next
- Link to v1: https://lore.kernel.org/bpf/20260709-b4-htab-batch-resched-v1-1-ad7a6b3b4513@xxxxxxxxx
---
kernel/bpf/hashtab.c | 24 +++++++++++++++++++++---
1 file changed, 21 insertions(+), 3 deletions(-)
diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c
index 6f331c80130d..a72dc5b9f184 100644
--- a/kernel/bpf/hashtab.c
+++ b/kernel/bpf/hashtab.c
@@ -1772,6 +1772,12 @@ static int htab_lru_percpu_map_lookup_and_delete_elem(struct bpf_map *map,
flags);
}
+/*
+ * Max consecutive empty buckets to walk in one RCU +
+ * instrumentation-disabled section before rescheduling.
+ */
+#define HTAB_BATCH_EMPTY_RESCHED 64
+
static int
__htab_map_lookup_and_delete_batch(struct bpf_map *map,
const union bpf_attr *attr,
@@ -1793,6 +1799,7 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map,
unsigned long flags = 0;
bool locked = false;
struct htab_elem *l;
+ u32 empty_cnt = 0;
struct bucket *b;
int ret = 0;
@@ -1971,12 +1978,21 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map,
}
next_batch:
- /* If we are not copying data, we can go to next bucket and avoid
- * unlocking the rcu.
+ /*
+ * If we are not copying data, we can go to next bucket and avoid
+ * unlocking the rcu. Bound the walk though: after
+ * HTAB_BATCH_EMPTY_RESCHED consecutive empty buckets, fully exit
+ * the critical section (no locks are held here) and reschedule.
*/
if (!bucket_cnt && (batch + 1 < htab->n_buckets)) {
batch++;
- goto again_nocopy;
+ if (++empty_cnt < HTAB_BATCH_EMPTY_RESCHED)
+ goto again_nocopy;
+ empty_cnt = 0;
+ rcu_read_unlock();
+ bpf_enable_instrumentation();
+ cond_resched_tasks_rcu_qs();
+ goto again;
}
rcu_read_unlock();
@@ -1990,11 +2006,13 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map,
}
total += bucket_cnt;
+ empty_cnt = 0;
batch++;
if (batch >= htab->n_buckets) {
ret = -ENOENT;
goto after_loop;
}
+ cond_resched_tasks_rcu_qs();
goto again;
after_loop:
---
base-commit: af0b84a9215d951d16f26b7ee34353b970cf5d4e
change-id: 20260708-b4-htab-batch-resched-1bce8304766d
Best regards,
--
Josef Bacik <josef@xxxxxxxxxxxxxx>