[REGRESSION] tick/sched: /proc/stat idle time exceeds wall time since v7.2

From: Stian Halseth

Date: Sun Oct 04 2026 - 10:27:53 EST


Hi,

Since v7.2-rc1, /proc/stat counts more idle time than wall time on
NO_HZ_IDLE kernels with TICK_CPU_ACCOUNTING. Monitoring that computes
CPU usage as 1 minus the idle rate shows negative usage on idle
machines.

Measured over 60 s with /proc/stat and the idle_sleeps counters in
/proc/timer_list:

SPARC T7-1, 7.3-rc5, HZ=100, 256 CPUs:
idle per wall second 1.0057 on average, 1.44 on the worst CPU
(92 tick stops/s, 4.8 ms of excess per tick stop)

amd64 (Opteron 1214), 7.2.7, HZ=1000, 2 CPUs:
idle machine: +0.03% and +0.06%
a task waking every 3.7 ms on CPU 1: +13.1% on CPU 1
(285 tick stops/s, 0.46 ms of excess per tick stop)

On both machines the excess is about half a tick per tick stop.

>From reading the code (not bisected), I think the cause is that
cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting") made
dyntick-idle time and tick-sampled time share cpustat[CPUTIME_IDLE].
On idle exit, kcpustat_dyntick_stop() accounts idle time up to now.
The tick is then restarted on its old period, and the first tick
accounts a whole TICK_NSEC through account_process_tick(), although
the part of that period before the idle exit has just been accounted
as idle. Before v7.2, /proc/stat read ts->idle_sleeptime, and the tick
fed a counter it did not use, so nothing was counted twice.

I have a fix that records the overlap in kcpustat_dyntick_stop() and
leaves it out of the first tick. With it, idle per wall second is
1.0000 on both machines, and -0.1% under the wakeup load at 280 tick
stops/s. I am still testing it with IRQ_TIME_ACCOUNTING and with steal
time in a KVM guest, and will send it when that is done.

#regzbot introduced: cf6444c3e1bb7

Stian