Re: [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan

From: Paul E. McKenney

Date: Mon Jul 20 2026 - 11:39:31 EST


On Mon, Jul 20, 2026 at 06:23:45AM -0700, Breno Leitao wrote:
> kmemleak_scan() can run for ages on large debug kernels. It was
> causing some soft-lockups which I got fixed with commit
> 3175fcfec8b16baeb ("mm/kmemleak: avoid soft lockup when scanning task
> stacks") with our beloved cond_resched().
>
> I've got the fix above deployed in the Meta fleet, and now I am seeing:
>
> INFO: rcu_tasks detected stalls on tasks:
> task:kmemleak state:R ... nvcsw: 274/274 holdout: 1 idle_cpu: -1/3
> scan_block
> scan_gray_list
> kmemleak_scan
>
> and, worse, blocks the callers waiting on that grace period. Here a BPF
> struct_ops map free, which waits via synchronize_rcu_mult(call_rcu,
> call_rcu_tasks), is stuck long enough to also trip the hung task check:
>
> INFO: task kworker/...:bpf_map_free_deferred blocked for 122 seconds
> __wait_rcu_gp
> bpf_struct_ops_map_free
>
> Then I've learned that cond_resched() is not an RCU-tasks quiescent
> state, so, we need to use stronger primitives.
>
> Use cond_resched_tasks_rcu_qs() at the scan reschedule points so the scan
> reports an RCU-tasks quiescent state as it proceeds.
>
> Inspired by commit b96285e10aad ("tracing: Have osnoise_main() add a
> quiescent state for task rcu").
>
> Signed-off-by: Breno Leitao <leitao@xxxxxxxxxx>

Reviewed-by: Paul E. McKenney <paulmck@xxxxxxxxxx>

> ---
> mm/kmemleak.c | 14 +++++++-------
> 1 file changed, 7 insertions(+), 7 deletions(-)
>
> diff --git a/mm/kmemleak.c b/mm/kmemleak.c
> index 85f18b17e79c4..f63dfacee7ca1 100644
> --- a/mm/kmemleak.c
> +++ b/mm/kmemleak.c
> @@ -1586,7 +1586,7 @@ static int scan_large_block(void *start, void *end)
> if (scan_block(start, next, NULL))
> return 1;
> start = next;
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
> }
>
> return 0;
> @@ -1623,7 +1623,7 @@ static void scan_object(struct kmemleak_object *object)
> scan_block(start, end, object);
>
> raw_spin_unlock_irqrestore(&object->lock, flags);
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
> raw_spin_lock_irqsave(&object->lock, flags);
> if (!(object->flags & OBJECT_ALLOCATED))
> break;
> @@ -1645,7 +1645,7 @@ static void scan_object(struct kmemleak_object *object)
> break;
>
> raw_spin_unlock_irqrestore(&object->lock, flags);
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
> raw_spin_lock_irqsave(&object->lock, flags);
> } while (object->flags & OBJECT_ALLOCATED);
> } else {
> @@ -1673,7 +1673,7 @@ static void scan_gray_list(void)
> */
> object = list_entry(gray_list.next, typeof(*object), gray_list);
> while (&object->gray_list != &gray_list) {
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
>
> /* may add new objects to the list */
> if (!scan_should_stop())
> @@ -1708,7 +1708,7 @@ static void kmemleak_cond_resched(struct kmemleak_object *object)
> raw_spin_unlock_irq(&kmemleak_lock);
>
> rcu_read_unlock();
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
> rcu_read_lock();
>
> raw_spin_lock_irq(&kmemleak_lock);
> @@ -1753,7 +1753,7 @@ static void kmemleak_scan_task_stacks(void)
> }
> put_task_struct(p);
> }
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
> } while (pid && !stop);
> }
>
> @@ -1935,7 +1935,7 @@ static int __kmemleak_scan(bool full)
> struct page *page = pfn_to_online_page(pfn);
>
> if (!(pfn & 63))
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
>
> if (!page)
> continue;
>
> ---
> base-commit: 0718283ab28bc3907e10b61a6b4be6fefa1cbb2f
> change-id: 20260720-kmemleak_rcu_task-6d3ea111b947
>
> Best regards,
> --
> Breno Leitao <leitao@xxxxxxxxxx>
>