Re: [REGRESSION] /proc/stat idle time exceeds wall clock since v7.2
From: Ahmed Shaltout
Date: Wed Sep 30 2026 - 01:01:27 EST
Hi Frederic,
Le Tue, Sep 29, 2026, Frederic Weisbecker a écrit :
> I ran on KVM too, overcommiting the vcpus like you did (16 vCPUs on 8 CPUs)
> and still nothing wrong.
>
> I'm wondering if this is specific to opensuse somehow. Can you try to build
> the latest upstream kernel?
I will, but first one difference that may explain why your guest stays
clean, and a hypothesis that fits our numbers.
Attached again is the 7.2.5 config, plus the 7.1.8 config from the control
box. Both are the stock openSUSE kernel-default files from /boot, unmodified.
Among the scheduler, tick, idle, accounting, preemption and RCU options they
differ only in CONFIG_SCHED_CACHE=y (7.2.5). All seven affected nodes run a
byte-identical config. None of the openSUSE-specific patches on top of
upstream stable touch tick, cputime, nohz, cpuidle or /proc. Going by their
names, they cover lockdown/secure boot, kABI, a few drivers and packaging.
1. There is no cpuidle driver on these guests
---------------------------------------------
On every node, affected or not:
/sys/devices/system/cpu/cpuidle/current_driver none
/sys/devices/system/cpu/cpuidle/current_governor menu
/sys/devices/system/cpu/cpu0/cpuidle/ no state* entries
So cpuidle_idle_call() takes the cpuidle_not_available() branch and calls
idle_call_stop_or_retain_tick(got_tick). On the first pass through the idle
loop the tick is retained, and it is stopped only after a tick has fired in
idle. What does your guest report there? If a driver's governor stops the
tick at idle entry, that could be the difference.
2. Hypothesis: the tick after a dyntick-idle window re-counts part of it
------------------------------------------------------------------------
This comes from reading v7.2, not from a test. Take the case with no cpuidle
driver:
- The retained idle tick is charged to CPUTIME_IDLE by
account_process_tick(), since kcpustat_idle_dyntick() is still false.
Its irq exit sets ts->idle_entrytime. The window then opens there via
kcpustat_dyntick_start(ts->idle_entrytime), so the start is seamless.
- On wakeup, tick_nohz_idle_exit() closes the window at "now"
(kcpustat_dyntick_stop). tick_nohz_restart() re-arms the tick at the
next jiffy boundary. That first tick charges a full TICK_NSEC, although
the part of that jiffy before "now" is already in the window.
- The over-count is the wakeup's phase within its jiffy, once per
window. Up to v7.1, /proc/stat read get_cpu_idle_time_us() for online
CPUs and ignored tick-charged idle, so the two never met in one counter.
- When a governor stops the tick at idle entry instead, the window starts
mid-jiffy. The fragment between the last tick and idle entry is then
never charged, which would roughly offset the end seam on average. That
might be why your guest looks right.
I have not verified this on a patched kernel, so please treat it as a lead
only.
3. Measurements, 60 s windows taken today on each node
-------------------------------------------------------
These are the per-CPU cpuN lines of /proc/stat against /proc/uptime's first
field, and .idle_sleeps from /proc/timer_list over the same window.
"sum/wall" is all eight modes divided by uptime x nr_cpus.
node cpus kernel sum/wall excess ms/s/cpu sleeps/s/cpu ms/sleep
cp-fsn1 8 7.2.5 1.100 100 351 0.285
cp-nbg1-pps 8 7.2.5 1.080 80 268 0.299
cp-nbg1-waa 8 7.2.5 1.072 74 245 0.301
cp-hel1 8 7.2.5 1.081 83 278 0.297
cp-hel1-b 8 7.2.5 1.076 76 255 0.298
worker-fsn1 8 7.2.5 1.063 58 200 0.292
monitoring 4 7.2.5 1.093 95 282 0.338
staging 8 7.1.8 0.985 -14 263 -
- Across nodes the excess tracks the tick-stop rate at about 0.3 ms per
.idle_sleeps event. Within one node it does not: a CPU with twice the
sleeps shows the same excess. .idle_sleeps is incremented on every
__tick_nohz_idle_stop_tick() call, including re-stops after an IRQ inside
an already-stopped idle period. So it overstates the number of windows on
IRQ-heavy CPUs. These guests stop the tick 200-350 times/s per CPU. A
near-idle test guest would show much less.
- /proc/stat idle and /proc/uptime idle still agree to within 0.02 s on
every node, over the same window.
- The seven affected nodes were on older kernels on the same VMs until
19 Sept: five on 7.1.8 and two on 7.0.12. Their daily all-mode sum read
0.979-0.993 of wall clock. Since rebooting into 7.2.5 it reads 1.04-1.10.
4. Corrections to my previous mail
----------------------------------
- Not every node is AMD EPYC-Rome. cp-hel1 reports "Intel Xeon Processor
(Skylake, IBRS, no TSX)" and shows the same excess (1.081), so it is not
vendor-specific. All nodes are the same Hetzner cx43 VM type, except the
4-vCPU one (cx33).
- cp-hel1 is tainted W, from a boot-time "CPA detected W^X violation"
warning in __change_page_attr. The other six are untainted and show the
same excess.
- I wrote that 7.1.8 sums to 0.992-0.997 of the ceiling. Over the last
three days it is 0.983-0.985, so 7.1.8 slightly under-counts rather than
matching wall clock exactly.
If the cpuidle difference does not let you reproduce it, I will boot
openSUSE's kernel-vanilla (Kernel:HEAD, currently 7.3-rc5, no distro patches)
on a VM of the same type, with 7.2.5 and 7.1.8 on the same VM as controls,
and send the three results.
Thanks,
Ahmed
Le Tue, Sep 22, 2026 at 07:53:23PM +0400, Ahmed Shaltout a écrit :
> Hi Frederic,
>
> Le Mon, Sep 21, 2026, Frederic Weisbecker a écrit :
> > I can't manage to reproduce that, neither on latest mainline nor on 7.2.7
> >
> > Can you share your whole .config file?
>
> Attached: config-7.2.5-1-default, taken from /boot/config-$(uname -r) on an
> affected node. It is the stock openSUSE Tumbleweed/MicroOS package
> kernel-default-7.2.5-1.1.x86_64, unmodified, no custom build. It differs
> from
> config/x86_64/default in openSUSE:Factory/kernel-source at the 7.2.5 source
> revision only in toolchain-detection and build-metadata lines (CC/AS/LD
> version, LOCALVERSION, MODULE_SIG_KEY, CC_HAS_*), no functional option:
> https://api.opensuse.org/public/source/openSUSE:Factory/kernel-source/config.tar.bz2?rev=4dabae0b715184e35b05b9a168bde6b7
I fear I still can't reproduce with the config in attachment.
>
> The unaffected control box runs kernel-default-7.1.8-1.1.x86_64 from the
> same
> distribution, also stock, same VM type, same command line. Between the two
> configs the only change among the scheduler, tick, accounting, preemption
> and
> RCU options is CONFIG_SCHED_CACHE=y appearing in 7.2.5.
>
> Environment, in case it is what your test box does not share:
>
> - Every node is a KVM guest (Hetzner Cloud, QEMU, "AMD EPYC-Rome
> Processor",
> 8 vCPUs, one with 4), clocksource kvm-clock.
> - No steal-time accounting at all: there is no kvm-stealtime line in the
> boot log and the steal column has read exactly 0 on every node for the
> 75 hours since boot, so the hypervisor does not seem to expose the MSR.
> - CONFIG_VIRT_CPU_ACCOUNTING_GEN=y, CONFIG_CONTEXT_TRACKING_IDLE=y and
> CONFIG_NO_HZ_FULL=y are built in (SUSE default) with nohz_full= empty,
> so the CPUs run idle dynticks, not full.
> - CONFIG_HZ=1000, CONFIG_PREEMPT_DYNAMIC=y, psi=1 on the command line:
> BOOT_IMAGE=/boot/vmlinuz-7.2.5-1-default root=UUID=... rd.timeout=60
> rd.retry=45 quiet systemd.show_status=yes console=ttyS0,115200
> console=tty0 ignition.platform.id=openstack security=selinux selinux=1
> psi=1
I ran on KVM too, overcommiting the vcpus like you did (16 vCPUs on 8 CPUs)
and still nothing wrong.
I'm wondering if this is specific to opensuse somehow. Can you try to build
the latest upstream kernel?
Thanks.
--
Frederic Weisbecker
SUSE Labs
Attachment:
config-7.2.5-1-default
Description: Binary data
Attachment:
config-7.1.8-1-default
Description: Binary data