Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
From: Paul E. McKenney
Date: Tue Aug 11 2026 - 13:42:46 EST
On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote:
> On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> >
> > INFO: rcu_tasks detected stalls on tasks:
> > 0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> > task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552
> > Call Trace:
> > shrink_lruvec
> > mem_cgroup_iter
> > shrink_node
> > do_try_to_free_pages
> > try_to_free_pages
> > __alloc_frozen_pages_noprof
> > alloc_pages_noprof
> > pte_alloc_one
> > __pte_alloc
> > handle_mm_fault
> >
> > Nothing promises direct reclaim returns in bounded time, and the scan
> > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> > PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU
> > quiescent state, so the reclaiming task never reports one and becomes a
> > holdout.
>
> I still don't understand why cond_resched() is being treated as involuntary
> preemption but that is orthogonal to this patch.
The history is that cond_resched() was originally intended to be a
preemption point in any otherwise non-preemptible kernel. Therefore,
because it is a preemption point, it counts as an involuntary context
switch.
Thanx, Paul
> > Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> > state even when cond_resched() does nothing.
> >
> > PS: This has been discussed in [1]
> >
> > Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@xxxxxxxxx/ [1]
> > Cc: stable@xxxxxxxxxxxxxxx
> > Signed-off-by: Breno Leitao <leitao@xxxxxxxxxx>
>
> Acked-by: Shakeel Butt <shakeel.butt@xxxxxxxxx>
>