Re: [PATCH] rseq: defer time slice extension yield for sys_futex_wake

From: Alice Ryhl

Date: Mon Aug 31 2026 - 15:24:09 EST


On Mon, Aug 31, 2026 at 09:45:25AM -0400, Mathieu Desnoyers wrote:
> On 2026-08-31 08:57, Alice Ryhl wrote:
> > When a task is granted an rseq scheduler time slice extension, it is
> > expected to finish its critical section and relinquish the CPU via
> > rseq_slice_yield(2). If the task issues any other system call while a
> > grant is active, rseq_syscall_enter_work() forces an immediate
> > reschedule on syscall entry via cond_resched(). This may cause
> > significant latency penalty for userspace lock implementations that use
> > rseq time slice extensions when unlocking the futex.
> >
> > In a userspace mutex unlock sequence:
> > 1. The lock is released in userspace.
> > 2. If there are waiters, the unlocking thread calls sys_futex_wake()
> > to wake a sleeping waiter.
> >
> > Because sys_futex_wake() is currently treated as an arbitrary syscall,
> > rseq_syscall_enter_work() schedules out the unlocking thread upon
> > syscall entry, which is before it has executed the wakeup. Consequently,
> > the lock is free in userspace, but the waiter remains blocked in the
> > kernel while the CPU switches to an unrelated task. The waiter is only
> > woken when the unlocking thread is eventually scheduled back in to
> > finish the syscall, causing lock handoff delays.
> >
> > Thus, update rseq_syscall_enter_work() for sys_futex_wake() so that it
> > does not reschedule during syscall entry. The thread will yield the CPU
> > on the syscall exit path instead.
> >
> > There is no need to apply this optimization to the multiplexed futex()
> > syscall since any userspace code that can invoke rseq_slice_yield() can
> > also invoke futex_wake().
>
> Is the goal there to provide a single blessed way of doing futex wake,
> or to allow the futex multiplexer to keep being used for that wake
> scenario ?
>
> The proposed change exposes two ABIs (multiplexer vs explicit futex
> wake) with very different behaviors. I'm concerned that it would be
> confusing to users.
>
> Thoughts ?

My understanding is that the multiplexed futex syscall is soft
deprecated and the goal is that new code should use the dedicated
syscalls, so I didn't think it was needed to implement the perf
optimizations for the "old" API.

But I do agree it's confusing to do it that way, so I'm happy to also
support the multiplexed one if you think we should.

Note that we can't read the 'op' argument to the multiplexed futex
syscall inside rseq_syscall_enter_work(), so to implement it for that
one too, we would have to skip the cond_resched() for all calls to
sys_futex(), and then re-check inside of sys_futex() itself to call
cond_resched() there if `op != FUTEX_WAKE` and the rseq time slice
extension applies.

Though now that I think about it, maybe we want to skip the
cond_resched() for all futex ops? If you're invoking FUTEX_WAIT, then
there's not really much reason to call cond_resched() if you're calling
schedule() immediately afterwards.

Alice