Re: [PATCH 1/4] sched/cache: Keep nr_pref_llc_running in the runnable domain
From: Tim Chen
Date: Thu Sep 10 2026 - 16:46:21 EST
On Thu, 2026-09-10 at 21:33 +0300, Kayra Cizmeci wrote:
> Hello :>,
>
> > alb_break_llc() decides whether to break LLC preference during active
> > load balance. It does so by testing that every runnable fair task on the
> > source rq prefers its LLC:
> >
> > env->src_rq->nr_pref_llc_running == env->src_rq->cfs.h_nr_runnable
> >
> > But the two counters cover different sets. nr_pref_llc_running is updated
> > in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued,
> > so it follows queued tasks. h_nr_runnable is updated in set_delayed()/
> > clear_delayed() and drops delay-dequeued tasks.
>
> > So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted
> > in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks,
> > alb_break_llc() returns false, and active balance is free to pull a task
> > off its preferred LLC. Active balance only moves runnable tasks, and this
> > is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB
> > skips the per-task test in can_migrate_task(). The runnable set is the one
> > we want.
>
> > Fix it on the counter side. A task should be counted in
> > nr_pref_llc_running exactly while it is both queued on its preferred LLC
> > (pref_llc_queued) and runnable (!sched_delayed). Define that membership
> > once in task_pref_llc_runnable(), and adjust the counter only through
> > pref_llc_running_inc()/pref_llc_running_dec() from the four sites that
> > change either input: account_llc_enqueue(), account_llc_dequeue(),
> > set_delayed() and clear_delayed(). Gating every update on the same
> > predicate keeps the delay, wake and dequeue paths from double-counting
> > or underflowing; see the comments at those sites for the ordering.
>
> > nr_llc_running and sd->llc_counts are not touched and stay on queued
> > semantics.
>
> I have one question tho, can't we combine the checks with h_nr_runnable? On the paper
> if we are updating h_nr_runnable we could check if the nr_pref_llc_running can be
> updated and update it if the condition is right. Because, every nr_pref_llc_running enters
> h_nr_runnable while not every h_nr_runnable enters nr_pref_llc_running.
>
> Why instead we just check the nr_pref_llc_running's conditions on task_pref_llc_runnable()
> and call these dec and inc functions after the h_nr_runnable updates. Wouldn't it be clear that way?
> If possible?
Yes, nr_pref_llc_running is a subset of h_nr_runnable.
We have to keep nr_pref_llc_running accounting apart from h_nr_runnable in set_delayed().
Note that in set_delayed(), pref_llc_running_dec() has to run while the task still looks runnable,
that is before se->sched_delayed = 1, because task_pref_llc_runnable()
gates on !sched_delayed. h_nr_runnable is decremented after the flag is
set:
if (entity_is_task(se))
pref_llc_running_dec(...); /* sched_delayed still 0 */
se->sched_delayed = 1;
...
for_each_sched_entity(se)
cfs_rq->h_nr_runnable--; /* sched_delayed already 1 */
So moving the accounting next to (or after) the h_nr_runnable update
would make task_pref_llc_runnable() return false and skip the
decrement, leaving nr_pref_llc_running too high.
clear_delayed() happens to be safe either way, since it clears
sched_delayed first, but keeping the two symmetric and calling inc/dec
explicitly at each site is what lets the single task_pref_llc_runnable()
predicate stay the one source of truth.
There is also a scope difference: h_nr_runnable is per-cfs_rq and
updated at every level of the hierarchy in the for_each_sched_entity()
loop, while nr_pref_llc_running is a per-rq scalar updated once per
task - which is why the dec sits before the loop, not inside it.
Thanks.
Tim