Re: Cache-aware scheduling does not work well with amd big/little cores

From: Mario Limonciello

Date: Wed Sep 09 2026 - 09:30:37 EST




On 9/9/26 03:59, Klaus Kusche wrote:

Hello,

On 08/09/2026 23:54, Tim Chen wrote:
I did some quick testing (no perfect benchmark environment,
just checking runtime and CPU consumption with "time").

I timed a kernel build (-j 24 and full lto, my own .config)
and an application build (also with a lot of parallelism)
with three different kernels:

a) Cache aware scheduling completely configured off

b) Cache aware scheduling turned on, but without patch

c) Cache aware scheduling turned on, with patch

Big/little scheduling was always on,
Mario's patch was always applied
(without it, results are significantly worse,

Sorry, I am a bit confused. You mentioned later the result for (b)
and (c) are about the same with or without the patch
exposing debugfs (commit c1e7fe5e75ed11fa85368e5a186472afd3858f3a
Mario mentioned in another mail).
But here you say the result is much worse without Mario's patch.
Is Mario's patch the one above or some other patch?

The with/without patch in (b) and (c)
refers to the patch you sent on 31/08/2026
( https://lore.kernel.org/lkml/20260825174112.2580942-1-tim.c.chen@xxxxxxxxxxxxxxx/ ),
not to Mario's patch.


Just to clarify Mario's patch in this context refers to the fix to ITMT/debugfs fixes as Klaus doesn't nominally enable debugfs in Kconfig:

eaece4849991d62fcd6f46637c55dcce00e25d70

because I use kernels without debugfs,
so big/little scheduling is off without the patch).

Results:

There is no significant difference between b) and c)
(<= 1 % wallclock time)

Yes, I don't expect difference between (b) and (c). My
understanding is the patch in question is to only
expose the default cache aware parameters via debugfs
but don't acutally change them.

Sometimes b) is better, sometimes c) is better,
I'd say the differences are below the accuracy of my tests.

But a) was reproducibly better than b) and c)
w.r.t. wallclock time: 2-2.6 %
It was also very slightly better w.r.t. total kernel CPU seconds.
The results w.r.t. total usermode CPU seconds varied too much.
(I always ran the application build twice,
and for all a), b) and c), the second run consumed
significantly more usermode CPU seconds,
but took a little less wallclock time - I don't know why).

Will have to look around to see if we have some similar CPUs
as HX-370. My understanding is that the 4 big cores are in
one L3 and the 8 small cores are in another L3.

Yes, as far as I know, it has 16 MB L3 cache for the 4 big cores
and 8 MB L3 cache for the 8 little cores.

BTW, we have also found two issues with the active load balance
paths for CAS that need fixes. You may want to add those patches
and see if they are helpful to improve things.

Most likely not within the next few days.
I'm still on holiday and quite busy:
Ars Electronica Festival in Linz.

Active load balance fixes:
https://lore.kernel.org/lkml/20260903020656.3793626-1-wanglu.priv@xxxxxxxxx/
https://lore.kernel.org/lkml/2b0a35122ee615c6fa51076e5d79330e633755ac.camel@xxxxxxxxxxxxxxx/