Re: [REGRESSION] sched/idle: Sysbench threads regression after f4c31b07b136
From: Joseph Salisbury
Date: Fri Jul 24 2026 - 13:28:18 EST
Hi Rafael, Christian,
On 7/6/26 10:29 AM, Christian Loehle wrote:
On 7/2/26 19:47, Rafael J. Wysocki (Intel) wrote:The guest-visible data does not show the single-idle-state cpuidle case.
Hi,
On Thu, Jul 2, 2026 at 6:30 PM Joseph Salisbury
<joseph.salisbury@xxxxxxxxxx> wrote:
Hi Rafael,I think that this is your case and the tick stops for you sometimes
We are seeing a reproducible MySQL Sysbench threads regression. A
bisect indicated the following commit as the first bad commit:
f4c31b07b136 ("sched: idle: Consolidate the handling of two special cases")
The regression was found in Oracle kernel performance testing on OCI VM
shapes:
VM Details:
* VM.Standard2.1:
x86 OCI VM shape, 1 OCPU / 2 hardware threads, about 14.5 GB RAM
* VM.Standard.A1.Flex.2:
Arm/Ampere A1 flexible VM shape, 2 vCPU threads, about 10.9 GB RAM
The ResultsDB runs show the regression in the Sysbench threads metric:
- VM.Standard2.1: 333 -> 236 (-29.1%)
- VM.Standard.A1.Flex.2: 1286 -> 1152 (-10.4%)
A test kernel was built with f4c31b07b136 reverted and the performance
regression was recovered.
From the code, it is possible the regression is due to the new
previous-wakeup heuristic in the special idle cases. Before the commit:
- no cpuidle driver:
tick_nohz_idle_stop_tick()
default_idle_call()
- one idle state:
tick_nohz_idle_retain_tick()
cpuidle state 0
now while it had never stopped before.
Can you confirm?
Both affected guests report:
/sys/devices/system/cpu/cpuidle/current_driver = none
/sys/devices/system/cpu/cpuidle/current_governor = menu
There are also no /sys/devices/system/cpu/cpu*/cpuidle entries on either guest. So from the guest data, this looks like the no-cpuidle-driver special case rather than the one-idle-state case.
The full revert of f4c31b07b136 recovered the regression. I also tested Rafael's suggested one-line change, applied as:
- idle_call_stop_or_retain_tick(stop_tick);
+ idle_call_stop_or_retain_tick(false);
That test kernel still showed regressed performance on VM.Standard2.1.
For the VMs, there are no guest cpuidle state lists exposed because the cpuidle driver is "none".
Overall, it would be good to know the idle state lists for both the VM
and the host.
I do not currently have the host-side idle-state lists from the OCI hosts. I can try to get that data if it would still be useful.
VM.Standard2.1, x86_64:+1, but also which HZ are you using?
CONFIG_HZ_1000=y
CONFIG_HZ=1000
VM.Standard.A1.Flex.2, aarch64:
CONFIG_HZ_250=y
CONFIG_HZ=250
Both systems reported have 2 logical CPUs then?
Yes:
VM.Standard2.1:
CPU(s): 2
Thread(s) per core: 2
Core(s) per socket: 1
Socket(s): 1
VM.Standard.A1.Flex.2:
CPU(s): 2
Thread(s) per core: 1
Core(s) per socket: 2
Socket(s): 1
Were higher core counts alsoYes. Higher-core-count runs were checked. The regression appears limited to the smaller core-count shapes.
tested and how does it affect them?
The current data shows regressions on:
- VM.Standard2.1
- VM.Standard.A1.Flex.2
- VM.Standard.E4.Flex.1
The larger tested shapes did not show the same regression pattern. The test runs use one sysbench thread per online CPU/core count as encoded in the metric name.