Re: [REGRESSION] /proc/stat idle time exceeds wall clock since v7.2

From: Ahmed Shaltout

Date: Wed Sep 30 2026 - 01:01:27 EST


Hi Frederic,


Le Tue, Sep 29, 2026, Frederic Weisbecker a écrit :

> I ran on KVM too, overcommiting the vcpus like you did (16 vCPUs on 8 CPUs)

> and still nothing wrong.

>

> I'm wondering if this is specific to opensuse somehow. Can you try to build

> the latest upstream kernel?


I will, but first one difference that may explain why your guest stays

clean, and a hypothesis that fits our numbers.


Attached again is the 7.2.5 config, plus the 7.1.8 config from the control

box. Both are the stock openSUSE kernel-default files from /boot, unmodified.

Among the scheduler, tick, idle, accounting, preemption and RCU options they

differ only in CONFIG_SCHED_CACHE=y (7.2.5). All seven affected nodes run a

byte-identical config. None of the openSUSE-specific patches on top of

upstream stable touch tick, cputime, nohz, cpuidle or /proc. Going by their

names, they cover lockdown/secure boot, kABI, a few drivers and packaging.


1. There is no cpuidle driver on these guests

---------------------------------------------


On every node, affected or not:


  /sys/devices/system/cpu/cpuidle/current_driver    none

  /sys/devices/system/cpu/cpuidle/current_governor  menu

  /sys/devices/system/cpu/cpu0/cpuidle/             no state* entries


So cpuidle_idle_call() takes the cpuidle_not_available() branch and calls

idle_call_stop_or_retain_tick(got_tick). On the first pass through the idle

loop the tick is retained, and it is stopped only after a tick has fired in

idle. What does your guest report there? If a driver's governor stops the

tick at idle entry, that could be the difference.


2. Hypothesis: the tick after a dyntick-idle window re-counts part of it

------------------------------------------------------------------------


This comes from reading v7.2, not from a test. Take the case with no cpuidle

driver:


  - The retained idle tick is charged to CPUTIME_IDLE by

    account_process_tick(), since kcpustat_idle_dyntick() is still false.

    Its irq exit sets ts->idle_entrytime. The window then opens there via

    kcpustat_dyntick_start(ts->idle_entrytime), so the start is seamless.


  - On wakeup, tick_nohz_idle_exit() closes the window at "now"

    (kcpustat_dyntick_stop). tick_nohz_restart() re-arms the tick at the

    next jiffy boundary. That first tick charges a full TICK_NSEC, although

    the part of that jiffy before "now" is already in the window.


  - The over-count is the wakeup's phase within its jiffy, once per

    window. Up to v7.1, /proc/stat read get_cpu_idle_time_us() for online

    CPUs and ignored tick-charged idle, so the two never met in one counter.


  - When a governor stops the tick at idle entry instead, the window starts

    mid-jiffy. The fragment between the last tick and idle entry is then

    never charged, which would roughly offset the end seam on average. That

    might be why your guest looks right.


I have not verified this on a patched kernel, so please treat it as a lead

only.


3. Measurements, 60 s windows taken today on each node

-------------------------------------------------------


These are the per-CPU cpuN lines of /proc/stat against /proc/uptime's first

field, and .idle_sleeps from /proc/timer_list over the same window.

"sum/wall" is all eight modes divided by uptime x nr_cpus.


  node         cpus kernel  sum/wall  excess ms/s/cpu  sleeps/s/cpu  ms/sleep

  cp-fsn1        8  7.2.5    1.100        100              351        0.285

  cp-nbg1-pps    8  7.2.5    1.080         80              268        0.299

  cp-nbg1-waa    8  7.2.5    1.072         74              245        0.301

  cp-hel1        8  7.2.5    1.081         83              278        0.297

  cp-hel1-b      8  7.2.5    1.076         76              255        0.298

  worker-fsn1    8  7.2.5    1.063         58              200        0.292

  monitoring     4  7.2.5    1.093         95              282        0.338

  staging        8  7.1.8    0.985        -14              263          -


- Across nodes the excess tracks the tick-stop rate at about 0.3 ms per

  .idle_sleeps event. Within one node it does not: a CPU with twice the

  sleeps shows the same excess. .idle_sleeps is incremented on every

  __tick_nohz_idle_stop_tick() call, including re-stops after an IRQ inside

  an already-stopped idle period. So it overstates the number of windows on

  IRQ-heavy CPUs. These guests stop the tick 200-350 times/s per CPU. A

  near-idle test guest would show much less.


- /proc/stat idle and /proc/uptime idle still agree to within 0.02 s on

  every node, over the same window.


- The seven affected nodes were on older kernels on the same VMs until

  19 Sept: five on 7.1.8 and two on 7.0.12. Their daily all-mode sum read

  0.979-0.993 of wall clock. Since rebooting into 7.2.5 it reads 1.04-1.10.


4. Corrections to my previous mail

----------------------------------


- Not every node is AMD EPYC-Rome. cp-hel1 reports "Intel Xeon Processor

  (Skylake, IBRS, no TSX)" and shows the same excess (1.081), so it is not

  vendor-specific. All nodes are the same Hetzner cx43 VM type, except the

  4-vCPU one (cx33).


- cp-hel1 is tainted W, from a boot-time "CPA detected W^X violation"

  warning in __change_page_attr. The other six are untainted and show the

  same excess.


- I wrote that 7.1.8 sums to 0.992-0.997 of the ceiling. Over the last

  three days it is 0.983-0.985, so 7.1.8 slightly under-counts rather than

  matching wall clock exactly.


If the cpuidle difference does not let you reproduce it, I will boot

openSUSE's kernel-vanilla (Kernel:HEAD, currently 7.3-rc5, no distro patches)

on a VM of the same type, with 7.2.5 and 7.1.8 on the same VM as controls,

and send the three results.


Thanks,

Ahmed


On Tue, Sep 29, 2026 at 5:50 PM Frederic Weisbecker <frederic@xxxxxxxxxx> wrote:
Le Tue, Sep 22, 2026 at 07:53:23PM +0400, Ahmed Shaltout a écrit :
> Hi Frederic,
>
> Le Mon, Sep 21, 2026, Frederic Weisbecker a écrit :
> > I can't manage to reproduce that, neither on latest mainline nor on 7.2.7
> >
> > Can you share your whole .config file?
>
> Attached: config-7.2.5-1-default, taken from /boot/config-$(uname -r) on an
> affected node. It is the stock openSUSE Tumbleweed/MicroOS package
> kernel-default-7.2.5-1.1.x86_64, unmodified, no custom build. It differs
> from
> config/x86_64/default in openSUSE:Factory/kernel-source at the 7.2.5 source
> revision only in toolchain-detection and build-metadata lines (CC/AS/LD
> version, LOCALVERSION, MODULE_SIG_KEY, CC_HAS_*), no functional option:
> https://api.opensuse.org/public/source/openSUSE:Factory/kernel-source/config.tar.bz2?rev=4dabae0b715184e35b05b9a168bde6b7

I fear I still can't reproduce with the config in attachment.

>
> The unaffected control box runs kernel-default-7.1.8-1.1.x86_64 from the
> same
> distribution, also stock, same VM type, same command line. Between the two
> configs the only change among the scheduler, tick, accounting, preemption
> and
> RCU options is CONFIG_SCHED_CACHE=y appearing in 7.2.5.
>
> Environment, in case it is what your test box does not share:
>
>   - Every node is a KVM guest (Hetzner Cloud, QEMU, "AMD EPYC-Rome
> Processor",
>     8 vCPUs, one with 4), clocksource kvm-clock.
>   - No steal-time accounting at all: there is no kvm-stealtime line in the
>     boot log and the steal column has read exactly 0 on every node for the
>     75 hours since boot, so the hypervisor does not seem to expose the MSR.
>   - CONFIG_VIRT_CPU_ACCOUNTING_GEN=y, CONFIG_CONTEXT_TRACKING_IDLE=y and
>     CONFIG_NO_HZ_FULL=y are built in (SUSE default) with nohz_full= empty,
>     so the CPUs run idle dynticks, not full.
>   - CONFIG_HZ=1000, CONFIG_PREEMPT_DYNAMIC=y, psi=1 on the command line:
>     BOOT_IMAGE=/boot/vmlinuz-7.2.5-1-default root=UUID=... rd.timeout=60
>     rd.retry=45 quiet systemd.show_status=yes console=ttyS0,115200
>     console=tty0 ignition.platform.id=openstack security=selinux selinux=1
>     psi=1

I ran on KVM too, overcommiting the vcpus like you did (16 vCPUs on 8 CPUs)
and still nothing wrong.

I'm wondering if this is specific to opensuse somehow. Can you try to build
the latest upstream kernel?

Thanks.

--
Frederic Weisbecker
SUSE Labs

Attachment: config-7.2.5-1-default
Description: Binary data

Attachment: config-7.1.8-1-default
Description: Binary data