Re: [PATCH] sched/cache: honor migrate_llc_task semantics in active load balance

From: Tim Chen

Date: Wed Aug 05 2026 - 12:11:55 EST


On Wed, 2026-08-05 at 10:38 +0800, wanglu15 wrote:
> From: Lu Wang <wanglu.priv@xxxxxxxxx>
>
> Thanks, Tim.
>
> On Tue, 2026-08-04 at 12:42 -0700, Tim Chen wrote:
> > On Tue, 2026-08-04 at 23:07 +0800, Lu Wang wrote:
> > > Thanks, Chenyu.
> > >
> > > On Tue, 2026-08-04 at 16:17 +0800, Chen, Yu C wrote:
> > > > Yes. Besides, if I understand correctly, I suppose Lu Wang was
> > > > referring to the following scenario:
> > > >
> > > > src_rq has 2 runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc),
> > > > while p2 prefers src_rq (src_llc). In this case, migrate_llc_task is
> > > > set because src_rq has at least one task, p1, that wants to migrate
> > > > to dst_rq. In ALB, can_migrate_task() found p2 and returns true for p2
> > > > thus moves p2 out of its preferred LLC.
> > >
> > > That's exactly the scenario I had in mind.
> > >
> > > > Firstly, before ALB is triggered, the generic (passive) load balance is
> > > > triggered. It iterates over p1 and p2 on src_rq to see if it can move any
> > > > one of them to dst_rq, and in most cases it succeeds in moving p1 to
> > > > dst_cpu. As a result, ALB will not be triggered.
> > >
> > > My question is whether p1 is guaranteed to be moved out in passive
> > > LB. can_migrate_task()/migrate_degrades_llc() can reject p1 for
> > > several independent reasons — p1 pinned by cpus_ptr, p1 cache-hot
> > > with nr_balance_failed still below cache_nice_tries, or
> > > can_migrate_llc_task() returning something other than mig_forbid due
> > > to capacity constraints on dst_llc at that instant. If passive LB
> > > rejects p1 for any of these, ALB is still triggered with p1 and p2
> > > both present on src_rq.
> > >
> > > Can we conclude that p1 and p2 never end up on src_rq together when
> > > ALB fires? Or would it help to set up a simple experiment and trace
> > > this path to see whether it actually occurs in practice?
> >
> > With 2 tasks on rq with different preference, active load balance could
> > pick the wrong task as can_migrate_task() checked in active load balance
> > will not consult migrate_degrades_llc(). How about the following patch
> > to fix this issue.
> >
> > [...]
> >
> > + if (env->migration_type == migrate_llc_task &&
> > + env->src_rq->cfs.h_nr_runnable > 1)
> > + return true;
> > +
> > return false;
> > }
>
> Your approach is simpler than mine — it avoids threading
> migration_type across the CPU stopper boundary and doesn't need any
> new rq field.
>
> One thing I'd like to flag, IMO: this approach skips the ALB path
> entirely for migrate_llc_task whenever more than one task is
> runnable, deferring the fix to the next passive LB pass. So it
> trades "delay" for a simpler implementation.
>

If we cannot pull a task from this rq for a migrate_llc_task imbalance 
with more than one runnable task, can_migrate_task() has already 
rejected the candidates — either the task preferring the dst LLC 
is cache-hot or capacity-constrained, or the only movable task 
prefers the source LLC. Forcing ALB here would ignore that.

We keep ALB only for the single-task case, where that
lone task prefers the dst LLC and has no other way to migrate.

I think this is the right thing to do because we shouldn't
force the tasks to move when can_migrate_task() is already
telling us not to.

Tim