Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems

From: Shrikanth Hegde

Date: Thu Oct 08 2026 - 11:50:40 EST


Hi Mete.

On 10/8/26 8:54 PM, Mete Durlu wrote:
On 08/10/2026 16:58, Shrikanth Hegde wrote:
Hi Meter/Andrea.

On 10/8/26 5:40 PM, Andrea Righi wrote:
Hi Mete,

On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
Summary
===========================================================================
On systems with asymmetric CPU capacities, the scheduler prefers fully idle
cores over idle SMT siblings of busy cores. This works generally well but
virtualized platforms where low capacity cores should be avoided are not
considered. Introduce a new config option and arch hook to prioritize
SMT utilization.

That's true only under physical CPU contention right? or is it always?

Right, because of that s390 treats all CPUs as equal and starts to
assign lower capacities to CPUs with low entitlement once a certain
steal time threshold is crossed (aka physical CPU contention).


Background
===========================================================================
The scheduler has been moving toward better utilization of fully idle cores
and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
broke the behavior s390 is relying on to concentrate workloads on its
high-capacity cores. On s390 CPU capacities represent the CPUs runtime
entitlement assigned by the hypervisor. The idle SMT siblings of busy
high-capacity cores start to perform better than fully idle low-capacity
cores as the whole machine(containing the logical partitions) starts
approaching to a {fully,over}loaded state. Grouping load on high- capacity
cores keeps shared low-capacity cores(which are shared more aggressively)
idle longer, reducing noise to neighboring partitions and improving
overall performance.


[...snip...]

Considerations and Open Questions
===========================================================================
The main goal for this series is adding a simple mechanism for
architectures to switch between idle core and idle SMT priority while
keeping the introduced footprint as small as possible. But there are
some ideas and questions to consider;

1. Should arch_needs_idle_smt_prio() hook be removed?
    The architecture hook is there to allow for any sort of logic to
    dynamically decide when to flip the mechanism, but it can also be
    removed if everyone agrees that this behaviour is not something that
    should be dynamically flipped. It can be simply tied to detection of
    asymmetric capacities and selection of Kconfig option.

2. Is there a simpler way to implement this mechanism?
    The proposed approach is chosen as the other features effecting the
    scheduler's behavior, use the same method. If there is a more
    efficient way to implement the same mechanism I'd be glad to use
    that instead.

3. There are no performance measurements *yet*.
    On s390 the ideal conditions for this mechanism to be beneficial
    usually surface when the whole system is fully loaded and resource
    sharing between the logical partitions starts to get expensive.
    Creating such an environment requires time, therefore no benchmark
    results are available yet.

If it only under contention, then you probably don't want to use low cores for
anything right? If above is yes, then using steal governor would help to avoid low
core if you mark them as non-preferred. (which i guess you guys are exploring already)

IMO it should be dynamically adjustable depending on the contention
level. But yes, on worst case low capacity cores should be avoided by
marking them as non-preferred.

But, find_new_ilb isn't aware of preferred CPU state as of initial patches.
I had thought of making a change to pick a idle preferred CPU instead of a idle non-preferred
core for idle load balancing. It is not done due to below reasons.
- Performance numbers didn't show any major improvements with real life workloads.
- Code becomes quite complex for find_new_ilb.

IIUC, find_new_ilb() finds an idle CPU that can do load balancing
for other idle CPUs. As long as the tasks don't land on non-preferred
CPUs, where ilb occurs should not matter. (At least for s390)


If that is what you want, then steal governor is the right choice.

Even if idle load balancing runs on a non-preferred idle CPU, it is going
to do the idle load balancing for all the idle CPUs. Including preferred and
non-preferred. And actual load balancing bails out early and does not pull any
task towards if it is a non-preferred CPU.


I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
mask's current behavior, soft preference would allow tasks to spill onto other
CPUs once the preferred CPUs have no idle threads available (we have something
like this in sched_ext's scx_cosmos).


cpu_preferred infra won't allow to spill over if preferred CPUs don't have an idle CPU.
It rather enforces packing onto preferred CPUs even if it means rq has more than 1 task.

I think what Andrea has in mind is something similar to what he did with
select_idle_capacity(). Depending on the state of the non-preferred mask
and other factors cores and SMT siblings can be evaluated as ideal
candidate somehow. Wouldn't it be possible to extend the
infrastructure you introduced to achieve this?

That might let us preserve the general preference for fully idle SMT cores while
searching in this order: fully idle preferred cores, idle SMT threads on
preferred cores, then non-preferred cores. This current patch series seems to
drop the first distinction, allowing a partially idle preferred core to win even
when another preferred core is fully idle.

This would need to coexist with the steal governor's stronger use of
cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
rather than a drop-in replacement.

I think this makes a lot of sense. An approach to first fill preferred
CPUs before spilling to the non-preferred cores until contention
forces the evacuation of non-prefered CPUs.

Well, when we see steal time, it usually means there is high contention and
we are using more vCPUs at this moment than possible.
So we mark them as non-preferred. Note that they are idle cpus. But non-preferred.

That means preferred CPUs may have more than 1 task. If we spill over if there
is idle non-preferred CPUs, we will be back to square one. No?


Shrikanth, would it make sense if I try to find a way to extend your
current implementation in this way?

Thanks!
-Mete