Re: [PATCH] mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()
From: Andrew Morton
Date: Mon Jul 27 2026 - 16:49:57 EST
On Mon, 27 Jul 2026 06:50:19 -0700 Breno Leitao <leitao@xxxxxxxxxx> wrote:
> migrate_pages_batch() unmaps each folio before moving it, and every
> unmap runs the mmu_notifier invalidate callbacks. On KVM hosts
> try_to_migrate() ends up in kvm_mmu_notifier_invalidate_range_start() ->
> tdp_mmu_zap_leafs(), which is expensive, so unmapping a large batch keeps
> the CPU busy for a long time.
>
> The loop already calls cond_resched(), but on PREEMPTION kernels that is
> a no-op, and involuntary preemption is not a Tasks-RCU quiescent state.
>
> A long batch therefore never reports a quiescent state, and the
> migrating task (e.g. kcompactd) becomes a Tasks-RCU holdout, stalling the
> Tasks-RCU grace period for minutes, which is common at Meta fleet:
>
> INFO: rcu_tasks detected stalls on tasks:
> 0000000055349ecc: .. nvcsw: 1157401/1157401 holdout: 1 idle_cpu: -1/56 task:kcompactd0 state:R running task
> Call Trace:
> tdp_mmu_zap_leafs
> tdp_mmu_next_root
> gfn_to_pfn_cache_invalidate_start
> kvm_mmu_notifier_invalidate_range_start
> __mmu_notifier_invalidate_range_start
> try_to_migrate_one
> try_to_migrate
> migrate_pages_batch
> migrate_pages
> compact_zone
> compact_node
> kcompactd
> kthread
Well I doubt if users of 7.1 kernels and earlier want to see this. So
a Fixes: and a cc:stable are needed. The affected code is quite old
and might even predate the addition of cond_resched_tasks_rcu_qs(). So
I can't begin to suggest a Fixes: target. Maybe omit it and let the
-stable team figure it out ;)
> --- a/mm/migrate.c
> +++ b/mm/migrate.c
> @@ -1843,7 +1843,7 @@ static int migrate_pages_batch(struct list_head *from,
> is_thp = folio_test_pmd_mappable(folio);
> nr_pages = folio_nr_pages(folio);
>
> - cond_resched();
> + cond_resched_tasks_rcu_qs();
>
Totally off-topic but why the heck was that implemented as a macro.
Which invokes another macro and another and another and turtles all the
way down. End result:
do { do { if (!((false)) && ({ do { __attribute__((__noreturn__)) extern void __compiletime_assert_606(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_606(); } while (0); (*(const volatile __typeof_unqual__(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout)); })) do { do { __attribute__((__noreturn__)) extern void __compiletime_assert_607(void) __attribute__((__error__("Unsupported access size for {READ,WRITE}_ONCE()."))); if (!((sizeof(((current))->rcu_tasks_holdout) == sizeof(char) || sizeof(((current))->rcu_tasks_holdout) == sizeof(short) || sizeof(((current))->rcu_tasks_holdout) == sizeof(int) || sizeof(((
current))->rcu_tasks_holdout) == sizeof(long)) || sizeof(((current))->rcu_tasks_holdout) == sizeof(long long))) __compiletime_assert_607(); } while (0); do { *(volatile typeof(((current))->rcu_tasks_holdout) *)&(((current))->rcu_tasks_holdout) = (false); } while (0); } while (0); } while (0); ({ __might_resched("mm/migrate.c", 1846, 0); _cond_resched(); }); } while (0);
How much nicer would it be to have
static inline void cond_resched_tasks_rcu_qs(void)
{
if (some brief and efficient test)
some_slow_uninlined_thing();
}