Re: [PATCH 8/8] sched/eevdf: Add min slice check when selecting CPU
From: Peter Zijlstra
Date: Tue Sep 22 2026 - 06:48:19 EST
On Mon, Sep 21, 2026 at 05:22:38PM +0200, Vincent Guittot wrote:
> Add a new level for selecting CPU when select_task_rq_fair() fails to find
> an idle CPU. This last level will compare the slice to select a CPU where
> the task could run 1st.
> This helps a waking task to select a CPU where a longer slice runs
> instead of one where a task with the same or shorter slice already runs.
>
> Signed-off-by: Vincent Guittot <vincent.guittot@xxxxxxxxxx>
> ---
> kernel/sched/fair.c | 63 ++++++++++++++++++++++++++++++++++++++++++++-
> 1 file changed, 62 insertions(+), 1 deletion(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 452da94289fe..5c9add3e853c 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -9023,6 +9023,59 @@ static inline bool asym_fits_cpu(unsigned long util,
> return true;
> }
>
> +static int select_slice_cpu(struct task_struct *p, struct sched_domain *sd, int target)
> +{
> + unsigned long task_util, util_min, util_max;
> + int cpu, nr = INT_MAX;
> + u64 slice = p->se.slice;
> + struct cpumask *cpus;
> +
> + if (sched_feat(SIS_UTIL) && sd->shared) {
> + /*
> + * Same nr_idle_scan hint as select_idle_cpu(), nr only limits
> + * the scan when not preferring an idle core.
> + */
> + nr = READ_ONCE(sd->shared->nr_idle_scan) + 1;
> + /* overloaded domain is unlikely to have idle cpu/core */
> + if (nr == 1)
> + return -1;
> + }
> +
> + cpus = this_cpu_cpumask_var_ptr(select_rq_mask);
> + cpumask_and(cpus, sched_domain_span(sd), p->cpus_ptr);
> +
> + if (sched_asym_cpucap_active()) {
> + task_util = task_util_est(p);
> + util_min = uclamp_eff_value(p, UCLAMP_MIN);
> + util_max = uclamp_eff_value(p, UCLAMP_MAX);
> + }
This all seems duplicated from select_idle_siblings(), and while I
appreciated the breaking into functions, I do worry about the duplicate
work done too.
> + /* Those CPUs have been tested not being idle and fiting */
> + for_each_cpu_wrap(cpu, cpus, target) {
> + /*
> + * Stop when the nr_idle_scan is exhausted (mirrors
> + * select_idle_cpu() logic).
> + */
> + if (--nr <= 0)
> + return -1;
> +
> + if (slice >= get_rq_min_slice(cpu_rq(cpu)))
> + continue;
> +
> + if (sched_asym_cpucap_active()) {
> + int fits = util_fits_cpu(task_util, util_min, util_max, cpu);
> +
> + /* Perfect fit: capacity satisfies util + uclamp */
> + if (fits > 0)
> + return cpu;
> + } else {
> + return cpu;
> + }
> + }
And it does seem like a waste to re-scan the CPUs we've already visited.
Can't we keep track of the minimal slice CPU that was not idle during
our initial scan?
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 2ad46fb2eafe..d3ab945f50ba 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -9179,6 +9179,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
struct sched_domain *sd;
unsigned long task_util, util_min, util_max;
int i, recent_used_cpu, prev_aff = -1;
+ int best = target;
/*
* On asymmetric system, update task utilization because we will check
@@ -9263,8 +9264,10 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
* capacity path.
*/
if (sd) {
- i = select_idle_capacity(p, sd, target);
- return ((unsigned)i < nr_cpumask_bits) ? i : target;
+ i = select_idle_capacity(p, sd, target, &best)
+ if ((unsigned)i < nr_cpumask_bits)
+ return i;
+ return best;
}
}
@@ -9282,7 +9285,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
}
}
- i = select_idle_cpu(p, sd, has_idle_core, target);
+ i = select_idle_cpu(p, sd, has_idle_core, target, &best);
if ((unsigned)i < nr_cpumask_bits)
return i;
@@ -9297,7 +9300,7 @@ static int select_idle_sibling(struct task_struct *p, int prev, int target)
if ((unsigned int)recent_used_cpu < nr_cpumask_bits)
return recent_used_cpu;
- return target;
+ return best;
}
/**