Re: [PATCH v2] sched/cputime: Don't account idle time twice after dyntick-idle

From: Stian Halseth

Date: Tue Oct 06 2026 - 06:50:59 EST



On Tue, 2026-10-06 at 12:11 +0200, Frederic Weisbecker wrote:
> Le Mon, Oct 05, 2026 at 09:40:10PM +0200, Stian Halseth a écrit :
> > On idle exit the dyntick-idle accounting accounts the time up to
> > now,
> > then the tick is restarted on its old period. The first tick
> > accounts a
> > whole TICK_NSEC to whatever runs, although the part of that period
> > before the idle exit has just been accounted as idle time, or with
> > IRQ_TIME_ACCOUNTING as IRQ time. That is up to a full tick per idle
> > exit, and /proc/stat reports more idle time than wall time.
> >
> > Record how much of the current tick period dyntick-idle has
> > accounted,
> > and leave it out of the tick that ends the period.
> >
> > Fixes: cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting")
> > Link:
> > https://lore.kernel.org/all/20261004142724.3896396-1-stian@xxxxxx/
> > Signed-off-by: Stian Halseth <stian@xxxxxx>
> > ---
> > v2:
> > - Keep the overlap per tick period, so that a second dyntick-idle
> >    period before the tick adds to it instead of replacing it
> > (Frederic)
> > - Read it through a helper in the NO_HZ_COMMON block, with a stub
> > in
> >    its #else, like kcpustat_field_dyntick(). IS_ENABLED() does not
> > build,
> >    as the field only exists with NO_HZ_COMMON. The helper is static
> >    inline, as there is no user with VIRT_CPU_ACCOUNTING_NATIVE.
> >
> > Frederic's case does not happen in my tests on its own. Since
> > f4c31b07b136 ("sched: idle: Consolidate the handling of two special
> > cases"), without a cpuidle driver or with a single state, the tick
> > is
> > only stopped after it woke up the idle loop, and that tick takes
> > the
> > overlap first. With a test-only change that stops the tick on every
> > idle entry, as a governor may, the guest added to a pending overlap
> > about 900 times a second under a pipe ping-pong between two CPUs.
> >
> > So it is not what is left with steal time. That is the same with
> > v2,
> > +0.7% to +1.4%, and only shows when the task also spins for 1 ms
> > after
> > each wakeup.
> >
> > Total CPU time per wall second on the loaded CPU, v2:
> >
> >    SPARC T7-1, HZ=100, busiest CPU*       1.0001
> >    SPARC T7-1, HZ=100, 3.7 ms sleeps      1.0000
> >    SPARC T7-1, HZ=100, 25 ms sleeps       0.9998|
> >    SPARC T7-1, HZ=100, pipe ping-pong     1.0000
> >    x86_64 KVM guest, HZ=250, 3.7 ms       0.997  (1.505 before)
> >    same guest, pipe ping-pong             1.000  (1.152 before)
> >    same guest, 9 ms sleeps                0.985
> >
> >    * under its normal load, about 90 tick stops/s
> >
> > The guest was tested as for v1, with and without
> > IRQ_TIME_ACCOUNTING,
> > with highres=off, with steal time and with the forced restart path.
> > The
> > -1.5% with 9 ms sleeps is the late tick delivery described in v1.
>
> IIUC, there is still an excess of cputime accounted here sometimes,
> right?
> And that never happened before the patchset that rewrote the cputime
> idle
> accounting. Am I understanding correctly?
>
>
Only with steal time, and that is older than the rework. I ran the same
guest test on v7.1 and on ad5a9e14ec8b, the parent of cf6444c3e1bb7.
Total CPU time per wall second on the loaded vCPU:

v7.1 ad5a9e14ec8b v2
3.7 ms sleeps 0.995 0.996 0.998
pipe ping-pong 0.997 0.996 1.000
steal, sleep + spin 1.029 1.028 1.008
steal, pipe ping-pong 1.338 1.341 1.001

Without steal time nothing was counted twice before the rework. With
steal time, v7.1 counted steal time during idle as both idle and steal
time. The rework fixed most of that, and what is left with a busy spin
after each wakeup is smaller than before.


/Stian

Attachment: signature.asc
Description: This is a digitally signed message part