Re: [PATCH 2/2] futex/requeue: Prevent rcuwait use-after-free during requeue PI
From: Sebastian Andrzej Siewior
Date: Fri Jul 17 2026 - 05:43:10 EST
On 2026-07-17 16:49:22 [+0800], Yao Kai wrote:
> On PREEMPT_RT, FUTEX_CMP_REQUEUE_PI can trigger a KASAN report:
>
> BUG: KASAN: slab-out-of-bounds in _raw_spin_lock_irqsave+0x76/0xe0
> Call Trace:
> _raw_spin_lock_irqsave+0x76/0xe0
> try_to_wake_up+0xab/0x1540
> rcuwait_wake_up+0x39/0x60
> futex_requeue+0x18c3/0x1e10
>
> The futex_q used by futex_wait_requeue_pi() is allocated on the waiter's
> stack. An early wakeup can race with a PI requeue as follows:
>
> waiter requeue task
> ------ ------------
> futex_wait_requeue_pi()
> futex_do_wait()
> schedule()
>
> * timeout/signal wakes waiter *
>
> futex_requeue_pi_wakeup_sync()
> IN_PROGRESS -> WAIT
> rcuwait_wait_event()
> requeue_pi_wake_futex()
> task = READ_ONCE(q->task)
> futex_requeue_pi_complete()
> WAIT -> LOCKED
> return LOCKED
> return
> // q lifetime ends
> rcuwait_wake_up()
>
> futex_requeue_pi_complete() publishes LOCKED before calling
> rcuwait_wake_up(). Once the waiter observes LOCKED, it can return from
> futex_wait_requeue_pi() and let q go out of scope before rcuwait_wake_up()
> reads q->requeue_wait.task and passes the stale pointer to
> try_to_wake_up().
>
> Skip rcuwait_wake_up() for Q_REQUEUE_PI_LOCKED. requeue_pi_wake_futex()
> already saves q->task before publishing LOCKED and wakes the saved task
> afterward.
>
> Fixes: 07d91ef510fb1 ("futex: Prevent requeue_pi() lock nesting issue on RT")
> Cc: stable@xxxxxxxxxxxxxxx
> Signed-off-by: Yao Kai <yaokai34@xxxxxxxxxx>
> ---
> kernel/futex/requeue.c | 9 +++++++--
> 1 file changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/kernel/futex/requeue.c b/kernel/futex/requeue.c
> index abc652b5b2dd..59e587775d9b 100644
> --- a/kernel/futex/requeue.c
> +++ b/kernel/futex/requeue.c
> @@ -155,8 +155,13 @@ static inline void futex_requeue_pi_complete(struct futex_q *q, int locked)
> } while (!atomic_try_cmpxchg(&q->requeue_state, &old, new));
>
> #ifdef CONFIG_PREEMPT_RT
> - /* If the waiter interleaved with the requeue let it know */
> - if (unlikely(old == Q_REQUEUE_PI_WAIT))
> + /*
> + * If the waiter interleaved with the requeue, let it know. For LOCKED,
> + * q may already be invalid, so requeue_pi_wake_futex() wakes the saved
> + * task instead.
> + */
> + if (unlikely(old == Q_REQUEUE_PI_WAIT) &&
> + new != Q_REQUEUE_PI_LOCKED)
> rcuwait_wake_up(&q->requeue_wait);
Your whole assumption is based on the requeue_state in
futex_requeue_pi_wakeup_sync() changes from Q_REQUEUE_PI_IN_PROGRESS to
Q_REQUEUE_PI_WAIT and the rcuwait_wait_event() does not wait because the
condition becomes true before that happens. So the rcuwait_wake_up()
could access q.requeue_wait which is allocated on behalf of the waiter
which is gone. Certainly possible. But if we skip the wait in thise case
we probably miss the 99% cases where the waiter did wait, no?
This looks similar to commit b549113738e8c ("futex: Prevent
use-after-free during requeue-PI").
> #endif
> }
Sebastian