Re: [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan
From: Breno Leitao
Date: Thu Jul 30 2026 - 10:07:21 EST
On Mon, Jul 27, 2026 at 06:04:37AM -0700, Breno Leitao wrote:
> I found 3 different cases on Meta fleet, where rcu task stalls show up:
...
> 3) Nvidia driver
> * stuck in nv_procfs_read_lock_params
I've investigated this issue. The nvidia driver contains the following
locking pattern:
static inline int nv_down_read_interruptible(struct rw_semaphore *lock)
{
while (!down_read_trylock(lock))
{
if (signal_pending(current))
return -EINTR;
cond_resched();
}
return 0;
}
Code at: https://github.com/NVIDIA/open-gpu-kernel-modules/blob/452cec62d827034798072827d3866d1881662b77/kernel-open/common/inc/nv-lock.h#L54
The lock is never acquired, and cond_resched() provides insufficient
scheduling relief.
This triggers a more severe issue with ftrace, which holds ftrace_lock
across RCU synchronization:
mutex_lock(&ftrace_lock);
ftrace_shutdown(struct ftrace_ops *ops, int command) {
synchronize_rcu_tasks_rude();
synchronize_rcu_tasks();
}
mutex_unlock(&ftrace_lock);
Since nv_down_read_interruptible() spins indefinitely, ftrace_lock
now becomes completely wedged.
I suppose we want two things:
1) change cond_resched() with cond_resched_tasks_rcu_qs()
2) Avoiding holding the ftrace lock while calling
synchronize_rcu_tasks()? It can take up to 10 minutes on a healthy
system to be releasd, right?
--breno