[REGRESSION] clocksource: delayed watchdog can mark TSC unstable

From: Karl Mehltretter

Date: Sat Oct 10 2026 - 03:00:14 EST


Since the watchdog rewrite in v7.1, the TSC can be marked unstable
after a delayed watchdog run under QEMU TCG. Without injected stalls,
this happened on 5 of 12 boots on a loaded host and 0 of 6 on an idle
host.

The delayed run re-arms the timer with an expiry already in the past.
The watchdog then runs again almost immediately and rejects the TSC
based on a roughly 75 us measurement interval.

#regzbot introduced: 763aacf86f1b

v7.3-rc5, 4 vCPUs, after vCPU 0 was stopped for 0.8 s:

clocksource: Marking clocksource tsc unstable due to frequency skew
clocksource: Watchdog hpet interval: 74940ns
clocksource: Clocksource tsc interval: 74251ns
clocksource: Switched to clocksource hpet

The intervals are 75 us, not the 500 ms watchdog period. A function
trace of clocksource_watchdog() from the same boot:

busybox-79 [000] ..s1. 15.168486: clocksource_watchdog <-call_timer_fn
busybox-79 [000] .Ns.. 16.269726: clocksource_watchdog <-call_timer_fn
busybox-79 [000] .Ns.. 16.269820: clocksource_watchdog <-call_timer_fn

The second call came 1.10 s after the first, the third 94 us after the
second. Re-arming with expires += WATCHDOG_INTERVAL leaves the next
expiry in the past when a run is more than one interval late.

For a 75 us interval, the 500 ppm allowance is only about 37 ns, plus
the readout time. The measured difference was 689 ns.

Before 763aacf86f1b an interval above WATCHDOG_INTERVAL_MAX_NS was
skipped with "Long readout interval, skipping watchdog check" and the
timer re-armed from jiffies, so there was no immediate second run, and
the skew was compared against a fixed margin of at least
2 * WATCHDOG_MAX_SKEW. The rewrite dropped both.

Reproducer: a v7.3-rc5 guest with allnoconfig plus the fragment below,
4 vCPUs, qemu-system-x86_64 -accel tcg, the HPET as watchdog.
PTRACE_ATTACH on the "CPU 0/TCG" thread for 0.8 s, detach, six times
12 s apart. The TSC was marked unstable on the first stall in 2 of 2
runs. Stopping the whole QEMU process did not reproduce it. The guest
clocks stopped with it.

The same stall against other kernels, same config:

TSC marked unstable
v7.3-rc5 2 of 2 runs
v7.3-rc5 + the check below 0 of 2
763aacf86f1b (the rewrite) 3 of 4
763aacf86f1b^ 0 of 2
v7.0 0 of 2

763aacf86f1b^ and v7.0 print the "Long readout interval" line for the
same stalls and keep the TSC. 763aacf86f1b marked it unstable on a
170 us window, hpet 169720 ns against tsc 168974 ns.

The added check in watchdog_check_freq() skips the frequency check and
resets the watchdog when the measured interval is shorter than a quarter
of the period:

if (max_delta < WATCHDOG_INTERVAL * TICK_NSEC / 4)
goto reset;

Re-arming from jiffies instead of from the previous expiry would avoid
the second run altogether.

An LLM agent helped me with the test setup and the trace analysis. I ran
the tests and checked the results.

The relevant part of the config fragment on top of allnoconfig:
CONFIG_64BIT=y
CONFIG_EXPERT=y
CONFIG_SMP=y
CONFIG_NR_CPUS=4
CONFIG_PREEMPT=y
CONFIG_HIGH_RES_TIMERS=y
CONFIG_NO_HZ_FULL=y
CONFIG_HZ_250=y
CONFIG_ACPI=y
CONFIG_PCI=y
CONFIG_VIRTIO_PCI=y
CONFIG_VIRTIO_CONSOLE=y
CONFIG_BLK_DEV_INITRD=y
CONFIG_SERIAL_8250=y
CONFIG_SERIAL_8250_CONSOLE=y
CONFIG_FTRACE=y
CONFIG_FUNCTION_TRACER=y
CONFIG_IRQ_TIME_ACCOUNTING=y
CONFIG_PROVE_LOCKING=y
CONFIG_DEBUG_PREEMPT=y

Thanks,
Karl