[PATCH v5] sched/fair: Preserve wake-affine CPU for non-SMT reciprocal sync wakeups
From: Shubhang Kaushik (Ampere)
Date: Thu Oct 08 2026 - 18:48:19 EST
For WF_SYNC wakeups, wake_affine() may select the waker CPU, but the CFS
wakeup path still passes that target to select_idle_sibling(). The idle
CPU search can then move the wakee away from the wake-affine target.
Pipe-style ping-pong workloads expose this because the wakee is handed
back and forth between two tasks. In that case, moving the wakee to
another idle CPU can cost more than preserving the wake-affine waker CPU.
Use the existing last_wakee and wake_wide() state to identify narrow
reciprocal WF_SYNC wakeups:
A wakes B
B wakes A
A wakes B
...
Handle only this narrow reciprocal case on non-SMT systems. Once the
wake-affine path has selected or kept the waker CPU, preserve that target
when the waker rq has no other runnable task. Return the waker CPU
before select_idle_sibling() so the idle CPU search does not move this
handoff away from the wake-affine target.
This does not define a generic WF_SYNC placement rule. Generic WF_SYNC
wakeups continue through the existing wake_affine() and
select_idle_sibling() behavior. SMT systems also continue through
select_idle_sibling(), where idle sibling/core placement can be handled
with SMT topology visible.
On asymmetric-capacity systems, still require the wakee to fit on the
waker CPU.
Signed-off-by: Shubhang Kaushik (Ampere) <sh@xxxxxxxxxx>
---
Testing was performed on an 80-core non-SMT Ampere Altra system.
perf bench sched pipe -l 1000000, 30 runs:
default:
3.651 -> 2.204 usec/op mean, about 39.6% improvement
3.733 -> 2.200 usec/op median, about 41.1% improvement
taskset -c 78,79:
3.820 -> 2.691 usec/op mean, about 29.6% improvement
3.887 -> 2.444 usec/op median, about 37.1% improvement
taskset -c 79:
2.292 -> 2.113 usec/op mean, about 7.8% improvement
2.290 -> 2.112 usec/op median, about 7.8% improvement
Baseline: tip/sched/core at 4a3b51aab6e2
Schbench normal mode at 8/40/80/240 workers and pipe mode at
1/2/4/8 workers showed no material regression over three 15s
runs. Additional pipe mode confirmation at one and two workers used
ten 30s runs; worker transfer medians changed by -0.3% and -2.4%,
respectively, with no material wakeup-latency regression.
Hackbench process and thread pipe cases at 1/2/4/8 groups were also
exercised on both kernels. These short runs were not included in the
performance comparison.
---
Changes in v5:
- Rebased on current upstream mainline.
- Refresh perf bench sched pipe results on the rebased kernel.
- WF_SYNC semantics are now documented as advisory; no placement
contract or generic WF_SYNC policy is introduced.
Link to v4: https://lore.kernel.org/r/20260803-b4-sched-sync-wakeup-v4-1-52333b0cfb79@xxxxxxxxxx
Changes in v4:
- Preserve the waker CPU only after the wake-affine path selected or
kept it.
- Clarify that WF_SYNC remains a hint, not a generic placement rule.
- Leave SMT systems on the existing select_idle_sibling() path.
- Refresh testing on tip:sched/core.
Link to v3: https://lore.kernel.org/r/20260727-b4-sched-sync-wakeup-v3-1-90cf481dbd85@xxxxxxxxxx
Changes in v3:
- Limit the direct waker-CPU preference to !sched_smt_active(); SMT
systems continue through the existing wake_affine() and
select_idle_sibling() path.
- Drop the redundant affinity check; want_affine already verifies the
waker CPU is allowed.
- Use a plain p->last_wakee read instead of READ_ONCE().
- Rebase and refresh testing on v7.2-rc5.
Link to v2: https://lore.kernel.org/r/20260722-b4-sched-sync-wakeup-v2-1-f1164560b24b@xxxxxxxxxx
Changes in v2:
- Move the reciprocal handoff preference under the existing
SD_WAKE_AFFINE domain check.
- Drop futex from the changelog motivation.
- Refresh perf bench sched pipe results after rebasing.
Link to v1: https://lore.kernel.org/r/20260721-b4-sched-sync-wakeup-v1-1-dc94f184e27f@xxxxxxxxxx
---
kernel/sched/fair.c | 26 ++++++++++++++++++++++++++
1 file changed, 26 insertions(+)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 57360f5cdde4fa66f4cba2cbb53d4e322134504e..ce624fdc1f4480bc6e7ae710950fcaa93ab32a3e 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -9167,6 +9167,26 @@ static inline bool asym_fits_cpu(unsigned long util,
return true;
}
+/*
+ * For reciprocal WF_SYNC handoffs, prefer the waker CPU when it has no
+ * other runnable task.
+ */
+static bool prefer_sync_pair_cpu(struct task_struct *p, int cpu)
+{
+ struct rq *rq = cpu_rq(cpu);
+
+ if ((rq->nr_running - cfs_h_nr_delayed(rq)) != 1)
+ return false;
+
+ if (sched_asym_cpucap_active()) {
+ sync_entity_load_avg(&p->se);
+ if (!task_fits_cpu(p, cpu))
+ return false;
+ }
+
+ return true;
+}
+
/*
* Try and locate an idle core/thread in the LLC cache domain.
*/
@@ -9955,6 +9975,12 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int wake_flags)
if (cpu != prev_cpu)
new_cpu = wake_affine(tmp, p, cpu, prev_cpu, sync);
+ if (sync && !sched_smt_active() &&
+ new_cpu == cpu &&
+ p->last_wakee == current &&
+ prefer_sync_pair_cpu(p, cpu))
+ return cpu;
+
sd = NULL; /* Prefer wake_affine over balance flags */
break;
}
---
base-commit: 69f80fef3153299d9c72c53d1d71eef6354b6926
change-id: 20260721-b4-sched-sync-wakeup-04d40cbeb1da
Best regards,
--
Shubhang Kaushik (Ampere) <sh@xxxxxxxxxx>