Re: [PATCH 1/3] sched/fair: Add smt_balance and llc_balance to the decision matrix
From: Jemmy Wong
Date: Fri Oct 02 2026 - 07:06:07 EST
Hi Tim,
Thanks for the review.
> On Oct 2, 2026, at 5:34 AM, Tim Chen <tim.c.chen@xxxxxxxxxxxxxxx> wrote:
>
> On Mon, 2026-09-28 at 12:20 +0800, Jemmy Wong wrote:
>> The group-type matrix was introduced in commit 0b0695f2b34a
>> ("sched/fair: Rework load_balance()"). When commit fee1759e4f04
>> ("sched/fair: Determine active load balance for SMT sched groups")
>> added group_smt_balance and commit f38cc2f0d8a3 ("sched/cache:
>> Prioritize tasks preferring destination LLC during balancing")
>> added group_llc_balance, neither commit updated the matrix table.
>>
>> Both types are only tagged on non-local groups in update_sg_lb_stats(),
>> so their local columns are N/A.
>>
>> As busiest, group_smt_balance is only set when dst_cpu is idle and the
>> SMT group runs more than one task. Against a local has_spare or
>> fully_busy group it goes through the nr_idle checks, where a non-SMT
>> dst group may also force the pull via smt_vs_nonsmt_groups(). Against a
>> local imbalanced or overloaded group the local group is busier and the
>> pair is balanced.
>>
>> As busiest, group_llc_balance is not an unconditional force. A local
>> overloaded group is busier and the pair is balanced; otherwise the
>> nr_idle checks apply, with prefer_sibling still able to force the pull
>> when the local group has spare capacity.
>>
>> No functional change.
>>
>> Signed-off-by: Jemmy Wong <jemmywong512@xxxxxxxxx>
>> ---
>> kernel/sched/fair.c | 16 +++++++++-------
>> 1 file changed, 9 insertions(+), 7 deletions(-)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index 7455a83a6a99..f9ddcecfd19d 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -12947,13 +12947,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s
>> /*
>> * Decision matrix according to the local and busiest group type:
>> *
>> - * busiest \ local has_spare fully_busy misfit asym imbalanced overloaded
>> - * has_spare nr_idle balanced N/A N/A balanced balanced
>> - * fully_busy nr_idle nr_idle N/A N/A balanced balanced
>> - * misfit_task force N/A N/A N/A N/A N/A
>> - * asym_packing force force N/A N/A force force
>> - * imbalanced force force N/A N/A force force
>> - * overloaded force force N/A N/A force avg_load
>> + * busiest \ local has_spare fully_busy misfit smt asym imbalanced llc overloaded
>> + * has_spare nr_idle balanced N/A N/A N/A balanced N/A balanced
>> + * fully_busy nr_idle nr_idle N/A N/A N/A balanced N/A balanced
>> + * misfit_task force N/A N/A N/A N/A N/A N/A N/A
>> + * smt_balance nr_idle nr_idle N/A N/A N/A balanced N/A balanced
>> + * asym_packing force force N/A N/A N/A force N/A force
>> + * imbalanced force force N/A N/A N/A force N/A force
>
> Thanks for picking this up - nice to have the table match the code
> again. I walked the new rows and columns against
> sched_balance_find_src_group(), and they line up, with one exception:
>
>> + * llc_balance nr_idle nr_idle N/A N/A N/A nr_idle N/A balanced
>
> I think local=has_spare, busiest=llc_balance is "force", not "nr_idle":
>
> if (sds.prefer_sibling && local->group_type == group_has_spare &&
> (busiest->group_type == group_llc_balance ||
> sibling_imbalance(env, &sds, busiest, local) > 1))
> goto force_balance;
>
> The group_llc_balance clause short-circuits the sibling_imbalance()
> test, so this pair forces unconditionally once prefer_sibling is set,
> and prefer_sibling is set for LLC-vs-LLC balancing: the groups span
> per-LLC (SD_SHARE_LLC) domains, which keep SD_PREFER_SIBLING (only
> SD_NUMA strips it). That also matches your changelog ("prefer_sibling
> still able to force the pull...") - the table just reads nr_idle where
> the prose says force.
>
> Could you flip that one cell?
>
> * llc_balance force nr_idle N/A N/A N/A nr_idle N/A balanced
You're right. The group_llc_balance clause short-circuits
sibling_imbalance(), so when the local group is has_spare and
prefer_sibling is set, a busiest llc_balance group is always pulled
from. I'll flip that cell to "force" in v2.
I also checked when prefer_sibling is set. It comes from the busiest
group's flags, i.e. the child domain's SD_PREFER_SIBLING (see
update_sd_lb_stats()). sd_init() sets that flag on every level, so SMT,
CLUSTER, MC and PKG all have it, and only SD_NUMA domains clear it. So
the pull is forced unless the child domain is a NUMA one. I'll reword
the changelog and cover letter to match.
> Thanks.
>
> Tim
>
>> + * overloaded force force N/A N/A N/A force N/A avg_load
>> *
>> * N/A : Not Applicable because already filtered while updating
>> * statistics.
Thanks,
Jemmy