Re: [PATCH] mm/kmemleak: report RCU-tasks quiescent states during the scan

From: Breno Leitao

Date: Mon Jul 27 2026 - 09:11:46 EST


> > The document was not a waste of time. it helped people like me to
> > understand what the issue is, and what are the decisions we have ahead
> > of us.
> >
> > Back to what is the best decision/design, I honestly don't have an
> > opinion, but I am happy to get more data about Meta production servers
> > to help with the decision.
>
> Looking forward to seeing what you come up with!

I found 3 different cases on Meta fleet, where rcu task stalls show up:

1) kmemleak -> This patch solves it
2) KVM / kcompactd
* Holdout: kcompactd0 (pid 1216), state:R, nvcsw frozen at
716336/716336 across three reports 10 min apart (stuck ≥20 min in
one compaction pass)
* This is coming from: migrate_pages ->
kvm_mmu_notifier_invalidate_range_start ->
tdp_mmu_next_root -> tdp_mmu_zap_leafs

3) Nvidia driver
* stuck in nv_procfs_read_lock_params

Given we can't do much about 3) and 1) is now closed, I will spend some
time geting more data and possible a fix for 2.