Re: [PATCH] rseq: defer time slice extension yield for sys_futex_wake

From: Thomas Gleixner

Date: Fri Sep 04 2026 - 01:32:51 EST


On Mon, Aug 31 2026 at 12:57, Alice Ryhl wrote:
> When a task is granted an rseq scheduler time slice extension, it is
> expected to finish its critical section and relinquish the CPU via
> rseq_slice_yield(2). If the task issues any other system call while a
> grant is active, rseq_syscall_enter_work() forces an immediate
> reschedule on syscall entry via cond_resched(). This may cause
> significant latency penalty for userspace lock implementations that use
> rseq time slice extensions when unlocking the futex.
>
> In a userspace mutex unlock sequence:
> 1. The lock is released in userspace.
> 2. If there are waiters, the unlocking thread calls sys_futex_wake()
> to wake a sleeping waiter.
>
> Because sys_futex_wake() is currently treated as an arbitrary syscall,
> rseq_syscall_enter_work() schedules out the unlocking thread upon
> syscall entry, which is before it has executed the wakeup. Consequently,
> the lock is free in userspace, but the waiter remains blocked in the
> kernel while the CPU switches to an unrelated task. The waiter is only
> woken when the unlocking thread is eventually scheduled back in to
> finish the syscall, causing lock handoff delays.
>
> Thus, update rseq_syscall_enter_work() for sys_futex_wake() so that it
> does not reschedule during syscall entry. The thread will yield the CPU
> on the syscall exit path instead.

That's undermining the design and takes control away from the scheduler.

It granted a short extension with well defined semantics and then you
special case futex_wake() which can take arbitrary time to complete.

Thanks,

tglx