Re: [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()
From: Paul E. McKenney
Date: Mon Jul 27 2026 - 17:16:56 EST
On Mon, Jul 27, 2026 at 01:49:50PM -0700, Andrew Morton wrote:
> On Mon, 27 Jul 2026 06:50:19 -0700 Breno Leitao <leitao@xxxxxxxxxx> wrote:
>
> > migrate_pages_batch() unmaps each folio before moving it, and every
> > unmap runs the mmu_notifier invalidate callbacks. On KVM hosts
> > try_to_migrate() ends up in kvm_mmu_notifier_invalidate_range_start() ->
> > tdp_mmu_zap_leafs(), which is expensive, so unmapping a large batch keeps
> > the CPU busy for a long time.
> >
> > The loop already calls cond_resched(), but on PREEMPTION kernels that is
> > a no-op, and involuntary preemption is not a Tasks-RCU quiescent state.
> >
> > A long batch therefore never reports a quiescent state, and the
> > migrating task (e.g. kcompactd) becomes a Tasks-RCU holdout, stalling the
> > Tasks-RCU grace period for minutes, which is common at Meta fleet:
> >
> > INFO: rcu_tasks detected stalls on tasks:
> > 0000000055349ecc: .. nvcsw: 1157401/1157401 holdout: 1 idle_cpu: -1/56 task:kcompactd0 state:R running task
> > Call Trace:
> > tdp_mmu_zap_leafs
> > tdp_mmu_next_root
> > gfn_to_pfn_cache_invalidate_start
> > kvm_mmu_notifier_invalidate_range_start
> > __mmu_notifier_invalidate_range_start
> > try_to_migrate_one
> > try_to_migrate
> > migrate_pages_batch
> > migrate_pages
> > compact_zone
> > compact_node
> > kcompactd
> > kthread
>
> Well I doubt if users of 7.1 kernels and earlier want to see this. So
> a Fixes: and a cc:stable are needed. The affected code is quite old
> and might even predate the addition of cond_resched_tasks_rcu_qs(). So
> I can't begin to suggest a Fixes: target. Maybe omit it and let the
> -stable team figure it out ;)
>
> > --- a/mm/migrate.c
> > +++ b/mm/migrate.c
> > @@ -1843,7 +1843,7 @@ static int migrate_pages_batch(struct list_head *from,
> > is_thp = folio_test_pmd_mappable(folio);
> > nr_pages = folio_nr_pages(folio);
> >
> > - cond_resched();
> > + cond_resched_tasks_rcu_qs();
> >
>
> Totally off-topic but why the heck was that implemented as a macro.
> Which invokes another macro and another and another and turtles all the
> way down. End result:
>
> do { do { if (!((false)) && ({ do { __attribute__((__noreturn__)) extern void __compiletime_assert_606(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_606(); } while (0); (*(const volatile __typeof_unqual__(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout)); })) do { do { __attribute__((__noreturn__)) extern void __compiletime_assert_607(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof(((
> current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_607(); } while (0); do { *(volatile typeof(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout) = (false); } while (0); } while (0); } while (0); ({ __might_resched("mm/migrate.c", 1846, 0); _cond_resched(); }); } while (0);
>
>
> How much nicer would it be to have
>
> static inline void cond_resched_tasks_rcu_qs(void)
> {
> if (some brief and efficient test)
> some_slow_uninlined_thing();
> }
Sigh. It used to be so simple, then people wanted error checks in
READ_ONCE() and WRITE_ONCE(). ;-)
The point of rcu_tasks_classic_qs() being a macro was to avoid #include
hell. But maybe cond_resched_tasks_rcu_qs() can be a static inline,
as shown below. It builds, so it must be perfect, right? On the other
hand, does that really help, given that cond_resched_tasks_rcu_qs()
gets inlined?
Thanx, Paul
------------------------------------------------------------------------
diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h
index c77b1e02a93cca..9a741ce0588575 100644
--- a/include/linux/rcupdate.h
+++ b/include/linux/rcupdate.h
@@ -217,15 +217,15 @@ static inline void exit_tasks_rcu_finish(void) { }
/**
* cond_resched_tasks_rcu_qs - Report potential quiescent states to RCU
*
- * This macro resembles cond_resched(), except that it is defined to
+ * This function resembles cond_resched(), except that it is defined to
* report potential quiescent states to RCU-tasks even if the cond_resched()
* machinery were to be shut off, as some advocate for PREEMPTION kernels.
*/
-#define cond_resched_tasks_rcu_qs() \
-do { \
- rcu_tasks_qs(current, false); \
- cond_resched(); \
-} while (0)
+static inline void cond_resched_tasks_rcu_qs(void)
+{
+ rcu_tasks_qs(current, false);
+ cond_resched();
+}
/**
* rcu_softirq_qs_periodic - Report RCU and RCU-Tasks quiescent states