[PATCH v2 bpf-next] bpf: Claim the per-CPU send_signal irq_work before filling it

From: Yuan Chen

Date: Thu Oct 08 2026 - 22:26:11 EST


irq_work_is_busy() cannot see the per-CPU send_signal_work while
it is being filled: the check only matches after irq_work_queue()
has claimed the work. A concurrent caller therefore passes it, both
callers race for the same irq_work, and the loser's signal is
silently lost along with its task reference while the queued work
runs with a mix of both callers' fields.

kprobe, tracepoint and perf_event programs exclude each other on
the same CPU through bpf_prog_active, so for two callers to meet,
one of the programs has to be a raw tracepoint or an fentry one.
And since commit 87c544108b61 ("bpf: Send signals asynchronously
if !preemptible") this path runs with IRQs enabled too, so a hard
IRQ can interrupt the fill as well, not only an NMI. Found by code
inspection and reproduced with a perf_event NMI program racing an
fentry one.

Claim the work through work->task itself: cmpxchg() it from NULL
to the target task before touching the other fields, and set it
back to NULL once the callback has consumed them. A context
finding the work claimed returns the documented -EBUSY.

Fixes: 1bc7896e9ef4 ("bpf: Fix deadlock with rq_lock in bpf_send_signal()")
Signed-off-by: Yuan Chen <chenyuan@xxxxxxxxxx>
---
Changes in v2:
- Drop the irq_work_queue() return-value handling, which is dead
code: irq_work_single() clears IRQ_WORK_PENDING before the
callback runs and the claim is only released after it, so the
queue cannot fail while the claim is held.
- Use work->task itself as the claim gate, cmpxchg()ing it from
NULL and back, so no extra field is needed.
- Describe which program pairs can actually race and that a hard
IRQ can interrupt the fill since 87c544108b61, not only an NMI.
---
kernel/trace/bpf_trace.c | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)

diff --git a/kernel/trace/bpf_trace.c b/kernel/trace/bpf_trace.c
index 195f78db9bda..5465e7eba238 100644
--- a/kernel/trace/bpf_trace.c
+++ b/kernel/trace/bpf_trace.c
@@ -823,6 +823,7 @@ const struct bpf_func_proto bpf_task_pt_regs_proto = {

struct send_signal_irq_work {
struct irq_work irq_work;
+ /* The claim: NULL marks the slot free for the next caller. */
struct task_struct *task;
u32 sig;
enum pid_type type;
@@ -842,6 +843,8 @@ static void do_bpf_send_signal(struct irq_work *entry)

group_send_sig_info(work->sig, siginfo, work->task, work->type);
put_task_struct(work->task);
+ /* Release once the fields are consumed. */
+ smp_store_release(&work->task, NULL);
}

static int bpf_send_signal_common(u32 sig, enum pid_type type, struct task_struct *task, u64 value)
@@ -885,14 +888,14 @@ static int bpf_send_signal_common(u32 sig, enum pid_type type, struct task_struc
return -EINVAL;

work = this_cpu_ptr(&send_signal_work);
- if (irq_work_is_busy(&work->irq_work))
+ if (cmpxchg(&work->task, NULL, task))
return -EBUSY;

/* Add the current task, which is the target of sending signal,
* to the irq_work. The current task may change when queued
* irq works get executed.
*/
- work->task = get_task_struct(task);
+ get_task_struct(task);
work->has_siginfo = siginfo == &info;
if (work->has_siginfo)
copy_siginfo(&work->info, &info);
--
2.54.0