Re: [REGRESSION] [PATCH v3 7/7] sched/eevdf: Move to a single runqueue
From: Aishwarya Rambhadran
Date: Wed Oct 07 2026 - 08:50:27 EST
Hi all,
Thanks for the suggestions. I ran a couple of follow-up tests on the
same AWS Graviton3 (m7g.metal) setup with v7.3-rc4:
1. Disable WA_WEIGHT:
echo NO_WA_WEIGHT > /sys/kernel/debug/sched/features
2. Use the "smp" cgroup mode:
echo smp > /sys/kernel/debug/sched/cgroup_mode
Both tests were run with 20 repeats across 2 boot sessions. Below are
the request latency p99 results for schbench configurations discussed
earlier (usec, smaller is better):
message-thread:16, worker-thread:16
v7.2 = 662869
v7.3-rc4 = 969557
NO_WA_WEIGHT = 958541
smp = 723098
message-thread:64, worker-thread:4
v7.2 = 730453
v7.3-rc4 = 942251
NO_WA_WEIGHT = 740454
smp = 818483
message-thread:32, worker-thread:4
v7.2 = 74251
v7.3-rc4 = 56725
NO_WA_WEIGHT = 59008
smp = 82053
For the m:16 t:16 case that I originally bisected, disabling WA_WEIGHT
does not materially change the regression. Switching to cgroup_mode=smp,
however, recovers most of it.
The m:64 t:4 case behaves differently. Disabling WA_WEIGHT brings the
p99 latency almost back to the v7.2 value, while using smp mode gives
a partial recovery.
There is also an interesting result with m:32 t:4, which was one of the
configurations that improved with v7.3-rc4. That improvement is mostly
preserved with NO_WA_WEIGHT, but disappears with cgroup_mode=smp.
So this does not seem to point to WA_WEIGHT alone as the cause of the
original regression. The results seem consistent with the workload/load-
dependent behavior discussed in this thread, where weight calculation
and resulting placement decisions can help some schbench configurations
while hurting others.
For the original m:16 t:16 regression in particular, the larger recovery
with cgroup_mode=smp also seems consistent with Prateek's observation
that the older weight calculation performs closer to the hierarchical
pick.
Overall, the results suggest that the latency changes are sensitive to
the workload configuration and the weight calculation being used, rather
than being a uniform regression from the single-runqueue change.
It would be interesting to understand whether this variation across
load points is an expected consequence of the new weight calculation,
or whether there is still scope to avoid the larger p99 regressions.
Thanks,
Aishwarya
On 28/09/26 6:30 PM, Mike Galbraith wrote:
On Mon, 2026-09-28 at 11:12 +0800, Chen Yu wrote:
According to my test last week, stacking the waker and wakee on the same CPU bringsDitto CPU wise. The impact of even modest waker/wakee concurrency is
a big improvement on my machine, iff the memory bandwidth is saturated. By comparison,
if the memory bandwidth is low, stacking the wakee on top of the waker causes harm.
considerable. For much of netperf, the post-wakeup tail becomes a CPU
service latency injection for the wakee when unnecessarily stacked.
While cache misses sting, those hurt.
As box approaches saturation, stacking of somewhat synchronous stuff
becomes a better bet. Aiming for having only waker's tail between
wakee and CPU has decent chance of being a winner in a box full of
wider obstacles.
-Mike