[tip: sched/core] sched/fair: Randomize equally shallow slow-path candidates

From: tip-bot2 for Christian Loehle

Date: Fri Sep 25 2026 - 06:54:21 EST


The following commit has been merged into the sched/core branch of tip:

Commit-ID: c9ce69fc43bd07dbed9d23a504e56f84604ecf42
Gitweb: https://git.kernel.org/tip/c9ce69fc43bd07dbed9d23a504e56f84604ecf42
Author: Christian Loehle <christian.loehle@xxxxxxx>
AuthorDate: Thu, 17 Sep 2026 16:39:15 +01:00
Committer: Peter Zijlstra <peterz@xxxxxxxxxxxxx>
CommitterDate: Fri, 25 Sep 2026 12:45:58 +02:00

sched/fair: Randomize equally shallow slow-path candidates

Picking the first eligible idle CPU leaves a scan-order bias. Concurrent
slow-path selectors can choose the same CPU before either task is enqueued.

Use reservoir sampling for equal exit latencies, resetting the candidate
count when a shallower candidate appears. Use the per-CPU scheduler PRNG
and reciprocal_scale() to avoid variable division or a second scan.

Use a u64 latency key with U64_MAX for unpublished states. Published
states take precedence; when none are found, sample among the idle CPUs
without a published state.

Signed-off-by: Christian Loehle <christian.loehle@xxxxxxx>
Signed-off-by: Peter Zijlstra (Intel) <peterz@xxxxxxxxxxxxx>
Reviewed-by: Vincent Guittot <vincent.guittot@xxxxxxxxxx>
Link: https://patch.msgid.link/20260917153915.1563875-3-christian.loehle@xxxxxxx
---
kernel/sched/fair.c | 18 ++++++++++++------
1 file changed, 12 insertions(+), 6 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index c4d4b55..2f5c706 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -24,6 +24,7 @@
#include <linux/mmap_lock.h>
#include <linux/hugetlb_inline.h>
#include <linux/jiffies.h>
+#include <linux/math.h>
#include <linux/mm_api.h>
#include <linux/highmem.h>
#include <linux/hrtimer.h>
@@ -8443,7 +8444,8 @@ static int
sched_balance_find_dst_group_cpu(struct sched_group *group, struct task_struct *p, int this_cpu)
{
unsigned long load, min_load = ULONG_MAX;
- unsigned int min_exit_latency = UINT_MAX;
+ u64 min_exit_latency = U64_MAX;
+ unsigned int nr_candidates = 0;
int least_loaded_cpu = this_cpu;
int shallowest_idle_cpu = -1;
int i;
@@ -8464,12 +8466,16 @@ sched_balance_find_dst_group_cpu(struct sched_group *group, struct task_struct *

if (available_idle_cpu(i)) {
struct cpuidle_state *idle = idle_get_state(rq);
- if (idle && idle->exit_latency < min_exit_latency) {
- min_exit_latency = idle->exit_latency;
- shallowest_idle_cpu = i;
- } else if ((!idle || idle->exit_latency == min_exit_latency) &&
- shallowest_idle_cpu == -1) {
+ u64 exit_latency = idle ? idle->exit_latency : U64_MAX;
+
+ if (shallowest_idle_cpu == -1 || exit_latency < min_exit_latency) {
+ min_exit_latency = exit_latency;
shallowest_idle_cpu = i;
+ nr_candidates = 1;
+ } else if (exit_latency == min_exit_latency) {
+ nr_candidates++;
+ if (!reciprocal_scale(sched_rng(), nr_candidates))
+ shallowest_idle_cpu = i;
}
} else if (shallowest_idle_cpu == -1) {
load = cpu_load(cpu_rq(i));