[PATCH] sched/psi: Remove the RCU wait from trigger destruction

From: Shakeel Butt

Date: Sat Oct 10 2026 - 15:15:27 EST


In the Meta fleet, we found PSI trigger destruction waiting for an RCU
grace period while holding kernfs node_mutex, blocking other users [1].
File close and node drain hold this mutex while calling the release
callback. Cgroup removal and disabling cgroup.pressure also hold
cgroup_mutex during the drain.

The wait was added by commit 0e94682b73bf ("psi: introduce psi monitor")
to protect worker and trigger lookups. Commit 461daba06bdc
("psi: eliminate kthread_worker from psi trigger scheduling mechanism")
replaced the worker queue with a group timer. The scheduler now only
checks whether rtpoll_task is NULL; it uses neither the task nor the
trigger. Commit a06247c6804f ("psi: Fix uaf issue when psi trigger is
destroyed while being polled") removed trigger replacement and polling's
RCU lookup.

Trigger lists are protected by mutexes. Proc poll/select entries are
removed before their file references are dropped. eventpoll_release()
removes epoll entries before the file release callback.
Commit aff037078eca ("sched/psi: use kernfs polling functions for PSI
trigger polling") moved cgroup polling to a kernfs waitqueue whose
lifetime follows the file.

A late timer firing only wakes the group waitqueue. The system group is
static, and commit 5457025fa8ca ("sched/psi: Shut down rtpoll_timer in
psi_cgroup_free()") shuts down the cgroup timer before freeing the group.
Cgroup reclamation already waits for scheduler readers.

Remove synchronize_rcu() from psi_trigger_destroy() and use
rcu_access_pointer() for the worker NULL check. Keep kthread_stop()
outside the trigger mutex.

Tested trigger churn, notifications and cgroup removal in an 8-CPU VM
with KASAN, lockdep, and full and lazy preemption. No warnings were found.

Link: https://github.com/bpftrace/user-tools/tree/master/runnablelockmonitor [1]
Signed-off-by: Shakeel Butt <shakeel.butt@xxxxxxxxx>
---
kernel/sched/psi.c | 25 +++++--------------------
1 file changed, 5 insertions(+), 20 deletions(-)

diff --git a/kernel/sched/psi.c b/kernel/sched/psi.c
index 4e152410653d..7c5381423334 100644
--- a/kernel/sched/psi.c
+++ b/kernel/sched/psi.c
@@ -626,8 +626,6 @@ static void init_rtpoll_triggers(struct psi_group *group, u64 now)
static void psi_schedule_rtpoll_work(struct psi_group *group, unsigned long delay,
bool force)
{
- struct task_struct *task;
-
/*
* atomic_xchg should be called even when !force to provide a
* full memory barrier (see the comment inside psi_rtpoll_work).
@@ -635,19 +633,16 @@ static void psi_schedule_rtpoll_work(struct psi_group *group, unsigned long dela
if (atomic_xchg(&group->rtpoll_scheduled, 1) && !force)
return;

- rcu_read_lock();
-
- task = rcu_dereference(group->rtpoll_task);
/*
- * kworker might be NULL in case psi_trigger_destroy races with
- * psi_task_change (hotpath) which can't use locks
+ * Only test whether a worker is installed; do not dereference the task.
+ * A racing trigger destruction may leave the timer armed.
+ * psi_cgroup_free() shuts it down before freeing the group.
+ * The system PSI group is static.
*/
- if (likely(task))
+ if (likely(rcu_access_pointer(group->rtpoll_task)))
mod_timer(&group->rtpoll_timer, jiffies + delay);
else
atomic_set(&group->rtpoll_scheduled, 0);
-
- rcu_read_unlock();
}

static void psi_rtpoll_work(struct psi_group *group)
@@ -1488,22 +1483,12 @@ void psi_trigger_destroy(struct psi_trigger *t)
mutex_unlock(&group->rtpoll_trigger_lock);
}

- /*
- * Wait for psi_schedule_rtpoll_work RCU to complete its read-side
- * critical section before destroying the trigger and optionally the
- * rtpoll_task.
- */
- synchronize_rcu();
/*
* Stop kthread 'psimon' after releasing rtpoll_trigger_lock to prevent
* a deadlock while waiting for psi_rtpoll_work to acquire
* rtpoll_trigger_lock
*/
if (task_to_destroy) {
- /*
- * After the RCU grace period has expired, the worker
- * can no longer be found through group->rtpoll_task.
- */
kthread_stop(task_to_destroy);
atomic_set(&group->rtpoll_scheduled, 0);
}
--
2.53.0-Meta