Re: [REGRESSION] sched/idle: Sysbench threads regression after f4c31b07b136

From: Rafael J. Wysocki (Intel)

Date: Thu Jul 02 2026 - 14:48:19 EST


Hi,

On Thu, Jul 2, 2026 at 6:30 PM Joseph Salisbury
<joseph.salisbury@xxxxxxxxxx> wrote:
>
> Hi Rafael,
>
> We are seeing a reproducible MySQL Sysbench threads regression. A
> bisect indicated the following commit as the first bad commit:
> f4c31b07b136 ("sched: idle: Consolidate the handling of two special cases")
>
> The regression was found in Oracle kernel performance testing on OCI VM
> shapes:
>
> VM Details:
> * VM.Standard2.1:
> x86 OCI VM shape, 1 OCPU / 2 hardware threads, about 14.5 GB RAM
>
> * VM.Standard.A1.Flex.2:
> Arm/Ampere A1 flexible VM shape, 2 vCPU threads, about 10.9 GB RAM
>
>
> The ResultsDB runs show the regression in the Sysbench threads metric:
>
> - VM.Standard2.1: 333 -> 236 (-29.1%)
> - VM.Standard.A1.Flex.2: 1286 -> 1152 (-10.4%)
>
> A test kernel was built with f4c31b07b136 reverted and the performance
> regression was recovered.
>
> From the code, it is possible the regression is due to the new
> previous-wakeup heuristic in the special idle cases. Before the commit:
>
> - no cpuidle driver:
> tick_nohz_idle_stop_tick()
> default_idle_call()
>
> - one idle state:
> tick_nohz_idle_retain_tick()
> cpuidle state 0

I think that this is your case and the tick stops for you sometimes
now while it had never stopped before.

Can you confirm?

Overall, it would be good to know the idle state lists for both the VM
and the host.

> After the commit, both paths use:
>
> idle_call_stop_or_retain_tick(got_tick)
>
> Here, got_tick is true if the CPU was woken from the previous idle-loop
> iteration by the scheduler tick, and false otherwise. On this
> workload/platform combination, using that previous wakeup source to
> decide whether to stop or retain the tick appears to change NOHZ
> behavior enough to regress this wakeup-heavy threaded workload.
>
> Do you have any thoughts on whether this is an expected tradeoff of the
> new heuristic,

For the second case, yes. It is expected that stopping the tick more
often may cause performance to drop.

> or whether the special cases need a narrower condition?

Let's first identify the case this happens in.

> For our stable kernels, the immediate candidate fix is to revert the
> backport, but before doing that I wanted to ask whether upstream would
> prefer a targeted adjustment.
>
> I can collect additional data if useful, for example cpuidle
> driver/state information, timer interrupt counts, idle residency, perf
> stat, or scheduler trace data from the affected OCI shapes.

So instead of reverting the commit, please try the appended change
(modulo gmail-induced whitespace breakage).

---
kernel/sched/idle.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)

--- a/kernel/sched/idle.c
+++ b/kernel/sched/idle.c
@@ -250,7 +250,7 @@ static void cpuidle_idle_call(bool stop_
*/
cpuidle_reflect(dev, entered_state);
} else {
- idle_call_stop_or_retain_tick(stop_tick);
+ idle_call_stop_or_retain_tick(false);

/*
* If there is only a single idle state (or none), there is