Re: [PATCH RFC 0/2] sched: Introduce idle SMT priority for asymmetric capacity systems

From: Tim Chen

Date: Thu Oct 08 2026 - 17:52:27 EST


On Thu, 2026-10-08 at 14:10 +0200, Andrea Righi wrote:
> Hi Mete,
>
> On Thu, Oct 08, 2026 at 10:30:46AM +0200, Mete Durlu wrote:
> > Summary
> > ===========================================================================
> > On systems with asymmetric CPU capacities, the scheduler prefers fully idle
> > cores over idle SMT siblings of busy cores. This works generally well but
> > virtualized platforms where low capacity cores should be avoided are not
> > considered. Introduce a new config option and arch hook to prioritize
> > SMT utilization.
> >
> > Background
> > ===========================================================================
> > The scheduler has been moving toward better utilization of fully idle cores
> > and avoiding SMT penalties. This trend (notably after commit 25a32e400a14
> > ("sched/fair: Prefer fully-idle SMT cores in asym-capacity idle selection")
> > broke the behavior s390 is relying on to concentrate workloads on its
> > high-capacity cores. On s390 CPU capacities represent the CPUs runtime
> > entitlement assigned by the hypervisor. The idle SMT siblings of busy
> > high-capacity cores start to perform better than fully idle low-capacity
> > cores as the whole machine(containing the logical partitions) starts
> > approaching to a {fully,over}loaded state. Grouping load on high-capacity
> > cores keeps shared low-capacity cores(which are shared more aggressively)
> > idle longer, reducing noise to neighboring partitions and improving
> > overall performance.
> >
> > Approach
> > ===========================================================================
> > This series introduces SCHED_IDLE_SMT_PRIO config option and the
> > sched_idle_smt_prio static branch, allowing architectures to treat idle
> > SMT siblings of busy cores as equal candidates during asymmetric capacity
> > load balancing.
> > Static branch checks are placed at paths considering fully idle cores
> > over idle SMT threads in presence of asymmetric CPU capacities within
> > scheduling groups. Inserted checks mostly override hints for idle core
> > selection or cause early exits favoring idle SMT siblings.
> > One more static branch check is added to new task path in order to favor
> > high capacity SMT siblings during initial task placement, therefore the
> > tasks are immediately placed in high capacity SMT siblings instead of
> > considering most idle but low capacity cores.
> >
> > Architectures opt in by selecting ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO and
> > implementing arch_needs_idle_smt_prio(), which is evaluated during each
> > asym_cpu_capacity_scan() to track the state as runtime topology changes.
> >
> > The second patch enables the feature for s390 when running on
> > hardware-backed topology in an LPAR with vertical polarization active
> > and system is approaching to a target state.
> > Otherwise the branch stays disabled and the scheduler falls back to the
> > standard idle-core preference.
> >
> > No functional change on architectures that do not select
> > ARCH_SUPPORTS_SCHED_IDLE_SMT_PRIO.
> >
> > Performance Results
> > ===========================================================================
> > Since this change effects the {fully,over}loaded state it is difficult
> > to test at the moment. Therefore there are no concrete numbers or
> > metrics for now for s390.
> > For architectures which do not opt in to this feature, no performance
> > change is expected.
> >
> > Considerations and Open Questions
> > ===========================================================================
> > The main goal for this series is adding a simple mechanism for
> > architectures to switch between idle core and idle SMT priority while
> > keeping the introduced footprint as small as possible. But there are
> > some ideas and questions to consider;
> >
> > 1. Should arch_needs_idle_smt_prio() hook be removed?
> > The architecture hook is there to allow for any sort of logic to
> > dynamically decide when to flip the mechanism, but it can also be
> > removed if everyone agrees that this behaviour is not something that
> > should be dynamically flipped. It can be simply tied to detection of
> > asymmetric capacities and selection of Kconfig option.
> >
> > 2. Is there a simpler way to implement this mechanism?
> > The proposed approach is chosen as the other features effecting the
> > scheduler's behavior, use the same method. If there is a more
> > efficient way to implement the same mechanism I'd be glad to use
> > that instead.
> >
> > 3. There are no performance measurements *yet*.
> > On s390 the ideal conditions for this mechanism to be beneficial
> > usually surface when the whole system is fully loaded and resource
> > sharing between the logical partitions starts to get expensive.
> > Creating such an environment requires time, therefore no benchmark
> > results are available yet.
>
> I'm wondering if we could build on Shrikanth's cpu_preferred_mask infrastructure
> to express a "soft preference" for the s390 higher-entitlement CPUs. Unlike the
> mask's current behavior, soft preference would allow tasks to spill onto other
> CPUs once the preferred CPUs have no idle threads available (we have something
> like this in sched_ext's scx_cosmos).

Soft preference is okay if contention is low.
But if steal% remains high, we may still need to transition to hard
boundaries/preferences on the cores to run on to bring steal% down.

Tim

>
> That might let us preserve the general preference for fully idle SMT cores while
> searching in this order: fully idle preferred cores, idle SMT threads on
> preferred cores, then non-preferred cores. This current patch series seems to
> drop the first distinction, allowing a partially idle preferred core to win even
> when another preferred core is fully idle.
>
> This would need to coexist with the steal governor's stronger use of
> cpu_preferred_mask, so I'm suggesting it as a potential direction to explore
> rather than a drop-in replacement.
>
> What do you think (Mete / Shrikanth)?
>
> Thanks,
> -Andrea