Re: [PATCH] sched/cache: honor migrate_llc_task semantics in active load balance

From: wanglu15

Date: Tue Aug 04 2026 - 22:40:04 EST


From: Lu Wang <wanglu.priv@xxxxxxxxx>

Thanks, Tim.

On Tue, 2026-08-04 at 12:42 -0700, Tim Chen wrote:
> On Tue, 2026-08-04 at 23:07 +0800, Lu Wang wrote:
> > Thanks, Chenyu.
> >
> > On Tue, 2026-08-04 at 16:17 +0800, Chen, Yu C wrote:
> > > Yes. Besides, if I understand correctly, I suppose Lu Wang was
> > > referring to the following scenario:
> > >
> > > src_rq has 2 runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc),
> > > while p2 prefers src_rq (src_llc). In this case, migrate_llc_task is
> > > set because src_rq has at least one task, p1, that wants to migrate
> > > to dst_rq. In ALB, can_migrate_task() found p2 and returns true for p2
> > > thus moves p2 out of its preferred LLC.
> >
> > That's exactly the scenario I had in mind.
> >
> > > Firstly, before ALB is triggered, the generic (passive) load balance is
> > > triggered. It iterates over p1 and p2 on src_rq to see if it can move any
> > > one of them to dst_rq, and in most cases it succeeds in moving p1 to
> > > dst_cpu. As a result, ALB will not be triggered.
> >
> > My question is whether p1 is guaranteed to be moved out in passive
> > LB. can_migrate_task()/migrate_degrades_llc() can reject p1 for
> > several independent reasons — p1 pinned by cpus_ptr, p1 cache-hot
> > with nr_balance_failed still below cache_nice_tries, or
> > can_migrate_llc_task() returning something other than mig_forbid due
> > to capacity constraints on dst_llc at that instant. If passive LB
> > rejects p1 for any of these, ALB is still triggered with p1 and p2
> > both present on src_rq.
> >
> > Can we conclude that p1 and p2 never end up on src_rq together when
> > ALB fires? Or would it help to set up a simple experiment and trace
> > this path to see whether it actually occurs in practice?
>
> With 2 tasks on rq with different preference, active load balance could
> pick the wrong task as can_migrate_task() checked in active load balance
> will not consult migrate_degrades_llc(). How about the following patch
> to fix this issue.
>
> [...]
>
> + if (env->migration_type == migrate_llc_task &&
> + env->src_rq->cfs.h_nr_runnable > 1)
> + return true;
> +
> return false;
> }

Your approach is simpler than mine — it avoids threading
migration_type across the CPU stopper boundary and doesn't need any
new rq field.

One thing I'd like to flag, IMO: this approach skips the ALB path
entirely for migrate_llc_task whenever more than one task is
runnable, deferring the fix to the next passive LB pass. So it
trades "delay" for a simpler implementation.

Wang