Re: [PATCH] sched/cputime: Don't account idle time twice after dyntick-idle

From: Frederic Weisbecker

Date: Mon Oct 05 2026 - 07:35:20 EST


Le Sun, Oct 04, 2026 at 08:47:01PM +0200, Stian Halseth a écrit :
> On idle exit the dyntick-idle accounting accounts the time up to now,
> then the tick is restarted on its old period. The first tick accounts a
> whole TICK_NSEC to whatever runs, although the part of that period
> before the idle exit has just been accounted as idle time, or with
> IRQ_TIME_ACCOUNTING as IRQ time. That is up to a full tick per idle
> exit, and /proc/stat reports more idle time than wall time.
>
> Record how much of the tick period had passed since the tick was
> stopped, and leave it out of the first tick.
>
> Fixes: cf6444c3e1bb7 ("tick/sched: Unify idle cputime accounting")
> Link: https://lore.kernel.org/all/20261004142724.3896396-1-stian@xxxxxx/
> Signed-off-by: Stian Halseth <stian@xxxxxx>

That makes sense. Some comments below:

> ---
> Tested by comparing /proc/stat with CLOCK_MONOTONIC per CPU over 30 to
> 60 s, with a task on one CPU that sleeps in a loop. Total CPU time per
> wall second on that CPU:
>
> before after
> SPARC T7-1, HZ=100, busiest CPU* 1.44 1.0000
> SPARC T7-1, HZ=100, 3.7 ms sleeps 1.0001
> SPARC T7-1, HZ=100, 25 ms sleeps 0.9996
> Opteron, HZ=1000, 3.7 ms sleeps 1.131 0.999
> x86_64 KVM guest, HZ=250, 3.7 ms 1.53 0.998
> same guest, 9 ms sleeps 1.195 0.982
>
> * under its normal load, about 95 tick stops/s
>
> The guest was tested with and without IRQ_TIME_ACCOUNTING, with
> highres=off, with steal time from a busy loop on the host CPU, and with
> a test-only change that forces tick_nohz_idle_restart_tick() on every
> idle loop iteration with the tick stopped.
>
> The -1.8% left with long sleeps in the guest is time between an idle
> tick's expiry and its delivery to the halted vCPU, which nothing
> accounts. The summed lateness of those ticks matches it within 2 ms/s,
> or within 8 ms/s with steal time, also with highres=off. It is there
> without this patch as well, where the double counting hides it. On the
> T7-1 it is within noise.
>
> With 3.7 ms sleeps and steal time, +0.3% to +1.3% is left, which I
> have not explained.

Perhaps because sometimes idle is reentered shortly after exiting and
kcpustart_dyntick_start() overwrites the previous overlap. Say we have:

TICK_NSEC=100
next_tick=X

tick stop()
entry = X-90
tick_restart()
exit = entry + 10 (which is X-80)
overlap = 10

schedule()
// tick still hasn't fired

tick_stop()
entry = X-10
tick_restart()
exit = entry + 5 (which is X-5)
overlap = 5

tick()
tick accounts TICK_NSEC - 5 but it should also consider the previous
overlap, so it should be TICK_NSEC - 15.

But beware as it only matters for tick X, not for the next-next one that will
fire at X + TICK_NSEC, in which case the previous overlap should be ignored.

But please double check what I'm saying while being sleep deprived :-)


>
> include/linux/kernel_stat.h | 6 ++++--
> kernel/sched/cputime.c | 31 +++++++++++++++++++++++++------
> kernel/time/tick-sched.c | 23 ++++++++++++++++++++---
> 3 files changed, 49 insertions(+), 11 deletions(-)
>
> diff --git a/include/linux/kernel_stat.h b/include/linux/kernel_stat.h
> index 9ca6c2259dfea..950867deea29f 100644
> --- a/include/linux/kernel_stat.h
> +++ b/include/linux/kernel_stat.h
> @@ -40,6 +40,8 @@ struct kernel_cpustat {
> seqcount_t idle_sleeptime_seq;
> u64 idle_entrytime;
> u64 idle_stealtime[2];
> + u64 idle_dyntick_entry;
> + u64 idle_tick_overlap;
> #endif
> u64 cpustat[NR_STATS];
> };
> @@ -111,7 +113,7 @@ static inline unsigned long kstat_cpu_irqs_sum(unsigned int cpu)
> #ifdef CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE
>
> static inline void kcpustat_dyntick_start(u64 now) { }
> -static inline void kcpustat_dyntick_stop(u64 now) { }
> +static inline void kcpustat_dyntick_stop(u64 now, u64 tick_start) { }
> static inline void kcpustat_irq_enter(u64 now) { }
> static inline void kcpustat_irq_exit(u64 now) { }
> static inline bool kcpustat_idle_dyntick(void) { return false; }
> @@ -132,7 +134,7 @@ static inline u64 kcpustat_field_iowait(int cpu)
> #else /* !CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE */
>
> extern void kcpustat_dyntick_start(u64 now);
> -extern void kcpustat_dyntick_stop(u64 now);
> +extern void kcpustat_dyntick_stop(u64 now, u64 tick_start);
> extern void kcpustat_irq_enter(u64 now);
> extern void kcpustat_irq_exit(u64 now);
> extern u64 kcpustat_field_idle(int cpu);
> diff --git a/kernel/sched/cputime.c b/kernel/sched/cputime.c
> index 06bddaa738e52..9205e92680943 100644
> --- a/kernel/sched/cputime.c
> +++ b/kernel/sched/cputime.c
> @@ -358,6 +358,19 @@ void thread_group_cputime(struct task_struct *tsk, struct task_cputime *times)
> }
> }
>
> +/*
> + * The first tick after dyntick-idle covers a period that the dyntick-idle
> + * accounting may already have accounted up to the idle exit.
> + */
> +static u64 tick_cputime(void)
> +{
> +#if defined(CONFIG_NO_HZ_COMMON) && !defined(CONFIG_HAVE_VIRT_CPU_ACCOUNTING_IDLE)
> + return TICK_NSEC - __this_cpu_xchg(kernel_cpustat.idle_tick_overlap, 0);
> +#else

Please use IS_ENABLED()

Thanks.

--
Frederic Weisbecker
SUSE Labs