[PATCH] sched/cache: Honor asym packing over cache aware schedulin= on hybrid system

From: Tim Chen

Date: Wed Sep 23 2026 - 17:28:40 EST


A regression was reported on an AMD Ryzen AI HX 370 running a cache
intensive Clang full-LTO link. The little cores run at a much lower
frequency (3.3 GHz vs 5.1 GHz) and have only half of the L3 cache
(8 MB vs 16 MB), so pinning such a task to the little-core LLC hurts
twice, and full-LTO builds slow down dramatically compared to
pre-cache-aware-scheduling kernels.

Asym packing and cache aware scheduling express conflicting placement
strategy. Asym packing wants a task to run on the highest priority
CPU, whereas CAS wants to co-locate the tasks of a process on one LLC
regardless of the priority of CPUs in that LLC. When asym packing
is turned on, it is trying to migrate task to an empty core that has
higher priority than source cpu, let asym packing win.

Reported-by: Klaus Kusche <klaus.kusche@xxxxxxxxxxxxxxx>
Closes: https://lore.kernel.org/lkml/2180ea5a-eb28-4152-8d4d-cd00b0c24b2e@c=
omputerix.info/
Signed-off-by: Tim Chen <tim.c.chen@xxxxxxxxxxxxxxx>
---
kernel/sched/fair.c | 21 ++++++++++++++++++---
1 file changed, 18 insertions(+), 3 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index f265de8721db..89eed4fbc4e2 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -10823,6 +10823,8 @@ static inline bool task_misfits_asym_cpu(struct lb_=
env *env, struct task_struct
return false;
}
=20
+static inline bool sched_asym(struct sched_domain *sd, int dst_cpu, int sr=
c_cpu);
+
/*
* Check if task p can migrate from source LLC to
* destination LLC in terms of cache aware load balance.
@@ -10847,6 +10849,10 @@ static enum llc_mig can_migrate_llc_task(struct lb=
_env *env,
if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu))
return mig_unrestricted;
=20
+ /* Prioritize asym packing over cache awareness */
+ if (sched_asym(env->sd, dst_cpu, src_cpu))
+ return mig_unrestricted;
+
/* skip cache aware load balance for too many threads */
if (invalid_llc_nr(grp, p, dst_cpu) ||
exceed_llc_capacity(grp, dst_cpu)) {
@@ -12043,6 +12049,15 @@ static inline bool llc_balance(struct lb_env *env,=
struct sg_lb_stats *sgs,
sgs->group_misfit_task_load)
return false;
=20
+ /*
+ * On asym packing domains, if the destination CPU
+ * has higher priority than all CPUs in the source group,
+ * prioritize asym packing.
+ */=20
+ if ((env->sd->flags & SD_ASYM_PACKING) &&
+ sgs->group_asym_packing)
+ return false;
+
/*
* Skip cache aware tagging if nr_balanced_failed is sufficiently high.
* Threshold of cache_nice_tries is set to 1 higher than nr_balance_faile=
d
@@ -13458,12 +13473,12 @@ static int need_active_balance(struct lb_env *env=
)
{
struct sched_domain *sd =3D env->sd;
=20
- if (alb_break_llc(env))
- return 0;
-
if (asym_active_balance(env))
return 1;
=20
+ if (alb_break_llc(env))
+ return 0;
+
if (imbalanced_active_balance(env))
return 1;
=20
--=20
2.32.0