Re: [PATCH v2 1/2] futex/requeue: Fix rtmutex schedule preparation for requeue PI

From: Yao Kai

Date: Sun Aug 23 2026 - 23:07:47 EST




On 8/5/2026 4:00 PM, Yao Kai wrote:


On 8/4/2026 8:21 PM, Peter Zijlstra wrote:
On Wed, Jul 22, 2026 at 04:51:39PM +0800, Yao Kai wrote:
A waiter requeued onto a PI futex can reach rt_mutex_wait_proxy_lock()
without rtmutex schedule preparation:

   WARNING: CPU: 0 PID: 293 at kernel/sched/core.c:7606
   RIP: rt_mutex_schedule+0x43/0x50
   Call Trace:
    rt_mutex_slowlock_block.constprop.0+0x5b/0x320
    rt_mutex_wait_proxy_lock+0x3e/0x80
    futex_wait_requeue_pi+0x3ba/0x590
    do_futex+0x171/0x1f0

rt_mutex_schedule() requires current->sched_rt_mutex to be set. Normally,
rt_mutex_pre_schedule() sets it before an rtmutex waiter can schedule. With
requeue PI, another task can enqueue the waiter after its futex_q becomes
visible:

         waiter                          requeue task
         ------                          ------------
futex_wait_requeue_pi()
   futex_wait_setup()
     futex_queue(&q)
                                         futex_requeue()
                                           rt_mutex_start_proxy_lock()
                                             enqueue rt_waiter
                                             install pi_blocked_on
                                           requeue_futex()
                                             plist_del(&q->list)
   futex_do_wait()
     plist_node_empty(&q->list)
     skip schedule()
                                             plist_add(&q->list)
                                           futex_requeue_pi_complete()
                                             IN_PROGRESS -> DONE
   futex_requeue_pi_wakeup_sync() // DONE
   rt_mutex_wait_proxy_lock()
     rt_mutex_schedule()

futex_do_wait() mistakes the temporary removal for a wakeup and skips
schedule(). The proxy waiter can nevertheless remain blocked on the target
rtmutex and subsequently enter rt_mutex_schedule() with
current->sched_rt_mutex clear.

Call rt_mutex_pre_schedule() and rt_mutex_post_schedule() directly around
rt_mutex_wait_proxy_lock() so this second blocking point has the required
scheduler preparation.

Fixes: d14f9e930b90 ("locking/rtmutex: Use rt_mutex specific scheduler helpers")
Suggested-by: Sebastian Andrzej Siewior <bigeasy@xxxxxxxxxxxxx>
Cc: stable@xxxxxxxxxxxxxxx
Signed-off-by: Yao Kai <yaokai34@xxxxxxxxxx>
---
  kernel/futex/requeue.c | 8 ++++++++
  1 file changed, 8 insertions(+)

diff --git a/kernel/futex/requeue.c b/kernel/futex/requeue.c
index 79823ad13683..f7889fb2fce4 100644
--- a/kernel/futex/requeue.c
+++ b/kernel/futex/requeue.c
@@ -1,6 +1,7 @@
  // SPDX-License-Identifier: GPL-2.0-or-later
  #include <linux/plist.h>
+#include <linux/sched/rt.h>
  #include <linux/sched/signal.h>
  #include "futex.h"
@@ -865,7 +866,14 @@ int futex_wait_requeue_pi(u32 __user *uaddr, unsigned int flags,
      case Q_REQUEUE_PI_DONE:
          /* Requeue completed. Current is 'pi_blocked_on' the rtmutex */
          pi_mutex = &q.pi_state->pi_mutex;
+        /*
+         * Requeue temporarily removes q from the hash bucket, so
+         * futex_do_wait() may skip schedule() even though the proxy
+         * waiter still has to block on the rtmutex.
+         */
+        rt_mutex_pre_schedule();
          ret = rt_mutex_wait_proxy_lock(pi_mutex, to, &rt_waiter);
+        rt_mutex_post_schedule();

As per always, I'm totally confused about everything. I mean, futexes
suck, but requeue sucks worse.

So the purpose of rt_mutex_pre_schedule() was to avoid the double waiter
enqueue for rt_mutex on RT, where sched_submit_work() will hit a
spinlock-nee-rtlock.

So rt_mutex_pre_schedule() must happen before the rt_mutex is added as a
waiter. However, AFAICT we're already a waiter at the above spot, no? So
this cannot be right.

The changelogs doesn't at all explain why this is correct. Please help?

You are right that rt_mutex_pre_schedule() is normally required before the
rtmutex waiter is enqueued.

The v2 placement relies on a property of this particular call path that I
failed to explain in the changelog: futex_wait_requeue_pi() is reachable
only via the FUTEX_WAIT_REQUEUE_PI syscall from userspace. Therefore,
'current' cannot be a workqueue/io-wq worker, nor can it carry a live
block plug across the syscall boundary.

Consequently, sched_submit_work() inside rt_mutex_pre_schedule() has no
pending plugged I/O or worker notifications to process here, so it cannot
recurse into an rtlock after the proxy waiter has been installed. In this
specific path, calling rt_mutex_pre_schedule() here effectively only sets
current->sched_rt_mutex before rt_mutex_wait_proxy_lock() calls
rt_mutex_schedule(), without executing any blocking submit_work after enqueue.

This was the rationale for adopting Sebastian's simpler suggestion from the
v1 discussion, but this essential invariant was omitted from the v2
changelog and code comments.

Thanks,
Yao Kai



Gentle ping. If the v2 approach is not acceptable, would the v1 approach be
preferable instead?

https://lore.kernel.org/r/20260717084922.4153317-1-yaokai34@xxxxxxxxxx

Thanks,
Yao Kai