Re: [PATCH] futex: Fix missed wakeup during private hash resize

From: Yao Kai

Date: Tue Aug 04 2026 - 08:47:12 EST




On 8/4/2026 5:16 PM, Peter Zijlstra wrote:
On Mon, Aug 03, 2026 at 06:43:13PM +0800, Yao Kai wrote:
A task performing a custom private hash resize can remain blocked in
uninterruptible sleep indefinitely. The hung-task detector reports:

INFO: task futex-resizer:314 blocked for more than 10 seconds.
task:futex-resizer state:D stack:14824 pid:314 tgid:312 ppid:311

Call Trace:
__schedule+0x521/0xf30
schedule+0x22/0xa0
futex_hash_allocate+0x3db/0x490
__do_sys_prctl+0x6f5/0xbd0
do_syscall_64+0xf9/0x530
entry_SYSCALL_64_after_hwframe+0x77/0x7f

Kernel panic - not syncing: hung_task: blocked tasks

futex_pivot_pending() allows the resize request to continue when
either no replacement hash is pending (hash_new == NULL) or the current
hash reference count has reached zero.

After the final-reference wake, another futex task can complete the
pivot between the two observations:

T1 T2

futex_hash_allocate()
wait_var_event(mm, ...)
futex_pivot_pending(mm)
hash_new != NULL
futex_hash()
futex_ref_get(old) -> false
futex_pivot_hash(mm)
hash_new = NULL
__futex_pivot_hash(mm, new)
rcu_assign_pointer(hash, new)
fph = rcu_dereference(hash) /* new */
futex_ref_is_dead(fph) -> false
schedule()

The pivot changes the state from hash_new != NULL with a dead current
hash to hash_new == NULL with a live current hash. The resize task can
observe hash_new in the pre-pivot state and hash in the post-pivot state,
causing futex_pivot_pending() to return false even though the pivot has
completed. Since a successful pivot does not notify waiters, the task
can go to sleep after the only preceding wakeup has already been
consumed.

Wake waiters after every successful pivot. A full memory barrier before
wake_up_var() pairs with set_current_state() in wait_var_event() and
orders the completed pivot before the lockless waitqueue_active() check
in wake_up_var(). The waiter therefore either observes hash_new == NULL
before sleeping or is made runnable.

Hmm, but isn't the problem a lack of serialization on futex_mm_phash
access?

That is, all of this futex_mm_phash::hash_new and futex_mm_phash::hash
swizzling happens while holding futex_mm_phash::lock, except for
futex_pivot_pending(), that is looking at these values without holding
the lock, resulting in it observing that inconsistent state per the
above.

Taking a mutex in a wait loop is sorta yuck, but it should work. If the
mutex is contended, it sleeps and the wait-loop 'spuriously' doesn't. If
the mutex is uncontended, it doesn't sleep, but the wait-loop will.

Does this work for you?

---

diff --git a/kernel/futex/core.c b/kernel/futex/core.c
index f74ede3df161..72d4698e35fb 100644
--- a/kernel/futex/core.c
+++ b/kernel/futex/core.c
@@ -1778,14 +1778,15 @@ void futex_hash_free(struct mm_struct *mm)
static bool futex_pivot_pending(struct mm_struct *mm)
{
+ struct futex_mm_phash *mmph = &mm->futex.phash;
struct futex_private_hash *fph;
- guard(rcu)();
+ guard(mutex)(&mmph->lock);
- if (!mm->futex.phash.hash_new)
+ if (!mmph->hash_new)
return true;
- fph = rcu_dereference(mm->futex.phash.hash);
+ fph = rcu_dereference_raw(mmph->hash);
return futex_ref_is_dead(fph);
}

Thanks! I verified your patch with my reproducer, and it completely fixes
the issue.
I will send out a v2 patch adopting your approach shortly.

Thanks,
Yao Kai