Re: [PATCH v2] sched/cache: Honor asym packing over cache aware scheduling on hybrid systems

From: Chen Yu

Date: Tue Oct 06 2026 - 11:32:28 EST


On Mon, Oct 05, 2026 at 11:29:53AM -0700, Tim Chen wrote:
> Date: Mon, 5 Oct 2026 11:29:53 -0700
> From: Tim Chen <tim.c.chen@xxxxxxxxxxxxxxx>
> To: Peter Zijlstra <peterz@xxxxxxxxxxxxx>, Ingo Molnar <mingo@xxxxxxxxxx>
> Cc: Tim Chen <tim.c.chen@xxxxxxxxxxxxxxx>, Chen Yu <yu.c.chen@xxxxxxxxx>,
> Mario Limonciello <mario.limonciello@xxxxxxx>, Vishal Badole
> <Vishal.Badole@xxxxxxx>, linux-kernel@xxxxxxxxxxxxxxx, x86@xxxxxxxxxx,
> platform-driver-x86@xxxxxxxxxxxxxxx, K Prateek Nayak
> <KPrateek.Nayak@xxxxxxx>, Ricardo Neri <ricardo.neri@xxxxxxxxx>, Kayra
> Cizmeci <kayracizmeci@xxxxxxxxx>, stable@xxxxxxxxxxxxxxx, Vincent Guittot
> <vincent.guittot@xxxxxxxxxx>, Juri Lelli <juri.lelli@xxxxxxxxxx>, Klaus
> Kusche <klaus.kusche@xxxxxxxxxxxxxxx>
> Subject: [PATCH v2] sched/cache: Honor asym packing over cache aware
> scheduling on hybrid systems
> X-Mailer: git-send-email 2.32.0
>
> A regression was reported on an AMD Ryzen AI HX 370 running a cache
> intensive Clang full-LTO link. The little cores run at a much lower
> frequency (3.3 GHz vs 5.1 GHz) and have only half of the L3 cache
> (8 MB vs 16 MB), so pinning such a task to the little-core LLC
> hurts twice, and full-LTO builds slow down dramatically compared to
> pre-cache-aware-scheduling kernels.
>
> Asym packing and cache aware scheduling express conflicting placement
> strategies. Asym packing wants a task to run on the highest priority CPU,
> whereas cache aware scheduling wants to co-locate the tasks of a process
> on one LLC regardless of the priority of CPUs in that LLC.
>
> When asym packing tries to migrate task to an idle core that has higher
> priority than source cpu, let asym packing win. Moving tasks to a higher
> performing idle core will buy more performance than cache co-location.
>
> Prioritize asym packing over LLC balancing for regular and active load
> balancing.
>
> Fixes: 23b2b5ccc45c ("sched/cache: Introduce helper functions to enforce LLC migration policy")
> Reported-by: Klaus Kusche <klaus.kusche@xxxxxxxxxxxxxxx>
> Closes: https://lore.kernel.org/lkml/2180ea5a-eb28-4152-8d4d-cd00b0c24b2e@xxxxxxxxxxxxxxx/
> Suggested-by: Kayra Cizmeci <kayracizmeci@xxxxxxxxx>
> Tested-by: Klaus Kusche <klaus.kusche@xxxxxxxxxxxxxxx>
> Tested-by: Ricardo Neri <ricardo.neri@xxxxxxxxx>
> Cc: stable@xxxxxxxxxxxxxxx # 7.2.x
> Signed-off-by: Tim Chen <tim.c.chen@xxxxxxxxxxxxxxx>
> ---
>

I leveraged AI to test it on top of 7.3.0-rc2 using an emulated hybrid setup on a symmetric
AMD Ryzen 9 8945HX (Zen 4, 16C/32T, two 32 MB L3):

By hacking the AMD pstate driver and QOS to limit the L3 cache ways:

  - LLC0 "big/fast":   CPUs 0-7,16-23, max freq 5.46 GHz, 16 L3 ways (32 MB),
                       ITMT prefcore ranking 236
  - LLC1 "little/slow": CPUs 8-15,24-31, max freq 3.29 GHz, 8 L3 ways (16 MB),
                       ITMT prefcore ranking 100

Workload: an 8-thread pointer-chase ring (24 MB working set).

This patch works as expected:
Before the patch:
                 cache-aware ON       cache-aware OFF
  throughput     ~230 M/s             ~530 M/s           (-56.6%)
  placement      6 of 8 threads       8 threads LLC0
                 stuck on LLC1


After the patch:
                 cache-aware ON       cache-aware OFF
  throughput     520.7 +- 18.3 M/s    512.9 +- 16.6 M/s   (+1.5%, noise)
  placement      8 threads LLC0       8 threads LLC0


Tested-by: Chen Yu <yu.c.chen@xxxxxxxxx>

thanks,
Chenyu