Re: [PATCH] io_uring/sqpoll: pin task across task-work publication
From: Jens Axboe
Date: Thu Sep 24 2026 - 19:33:05 EST
On 9/24/26 4:59 PM, Jens Axboe wrote:
> On 9/24/26 2:50 PM, J?r?my Jean wrote:
>> SQPOLL can consume a request immediately after mpscq_push(). If this
>> drops the final ring reference, ctx and tctx may be freed before
>> io_req_normal_work_add() finishes using them. KASAN reports:
>>
>> BUG: KASAN: slab-use-after-free in io_req_normal_work_add+0x439/0x510
>> Read of size 4 at addr ff11000000ca0000 by task repro/55
>> (...)
>> BUG: KASAN: slab-use-after-free in queue_work_on+0x25/0x70
>> Write of size 8 at addr ff11000000c3bd00 by task repro/55
>>
>> Handle SQPOLL before publication. Pin tctx->task, publish the request,
>> then use only the pinned task so the publication tail cannot dereference
>> freed contexts.
>>
>> Fixes: af5d68f8892f ("io_uring/sqpoll: manage task_work privately")
>
> Was going to say this isn't correct, as the mpscq is newer than that.
> But I suppose this already existed before that, with the previous switch
> to private task work for SQPOLL? I'll dig into that a bit...
>
>> diff --git a/io_uring/tw.c b/io_uring/tw.c
>> index f573bcc3af6a..2e37a3fd7a33 100644
>> --- a/io_uring/tw.c
>> +++ b/io_uring/tw.c
>> @@ -210,6 +210,19 @@ void io_req_normal_work_add(struct io_kiocb *req)
>> struct io_uring_task *tctx = req->tctx;
>> struct io_ring_ctx *ctx = req->ctx;
>>
>> + /* SQPOLL may consume and retire the request immediately after push. */
>> + if (ctx->flags & IORING_SETUP_SQPOLL) {
>> + struct task_struct *task = tctx->task;
>> + bool first;
>> +
>> + get_task_struct(task);
>> + first = mpscq_push(&tctx->task_list, &req->io_task_work.node);
>> + if (first)
>> + __set_notify_signal(task);
>> + put_task_struct(task);
>> + return;
>> + }
>
> There's no point in having a 'first' variable. LLM's love bools...
>
> if (mpscq_push(&tctx->task_list, &req->io_task_work.node))
> __set_notify_signal(task);
>
> would be better.
Thinking about this a bit more, I think the following fix would be
better:
1) Add a guard(rcu)(); in io_req_normal_work_add() at the top, which is
how we handle this for DEFER_TASKRUN as well.
2) And similarly, expand the synchronize_rcu() run in
io_ring_exit_work() to also include IORING_SETUP_SQPOLL.
I think that's both a cleaner and more efficient fix, rather than fiddle
with task_struct references.
--
Jens Axboe