Re: [PATCH] x86/tsc: Require a deviating refined calibration to be reproduced

From: Alexander Warth

Date: Sun Oct 04 2026 - 07:55:46 EST


On Sun, 4 Oct 2026 01:37:23 +0330, NvD Ss wrote:
> I appear to be seeing essentially the same TSC issue on another Raven
> Ridge / ASUS A320 system.

Thanks for the report. Your mail is not linked to the patch thread and
dropped part of the Cc list, so I have restored both.

> tsc: Detected 3593.185 MHz processor
> clocksource: Switched to clocksource tsc-early
> tsc: Refined TSC clocksource calibration: 3596.250 MHz
> clocksource: Switched to clocksource tsc
> clocksource: Marking clocksource tsc unstable due to frequency skew
> clocksource: Watchdog hpet interval: 504269467ns
> clocksource: Clocksource tsc interval: 503848443ns

This is the same failure, and it is the case which the patch addresses.

The watchdog sees the TSC 421us short over 504ms, i.e. 835ppm. With
tsc_khz at 3596.250 MHz that puts the real TSC frequency in this
interval at 3593.247 MHz. Your cold boot refines to 3593.248 MHz. So
nothing happens to the TSC in the watchdog interval. The step fell into
the refinement window, the disturbed sample was installed, and the
watchdog rejects the TSC for a frequency error which the kernel
introduced itself.

Your acpi_pm boot agrees: 414us over 496ms is 835ppm again, which
corrects 3596.253 MHz to 3593.250 MHz.

Unpatched 7.2.5 did the same here: 3202.720 MHz via HPET and 3202.698
MHz via acpi_pm instead of 3199.995 MHz, i.e. +852ppm and +845ppm, and
the TSC was lost both times.

If your refinement window is about as long as mine, 1.04 to 1.06s, then
835ppm is a step of about 870 to 885us. Here it is 881 to 888us, on a
TSC which runs at 3200 instead of 3593 MHz. That would be the same step
in time, not in cycles.

The other case exists as well. On a slow debug kernel the step landed
in a watchdog interval here. The refinement was correct and the TSC
interval was 880us longer than the HPET one.

> tsc=reliable clocksource=tsc
[...]
> the normal clocksource validation. systemd-timesyncd initially reached
> the +500 ppm frequency correction limit during that test.

That fits. tsc=reliable disables the watchdog, but the refinement still
runs. That boot kept 3596.25 MHz, so the clock was 835ppm slow, and NTP
cannot correct more than 500ppm.

With the patch, 3596.250 MHz is 853ppm off the early calibration, which
is beyond the 2^-11 (488ppm) limit. The sample is not installed and the
measurement is repeated. The next window should give about 3593.25 MHz,
which matches the early calibration and is installed. The acpi_pm boot
is 732ppm off and takes the same path.

It would help a lot if you could test this:

1. v7.2.8 plus the patch, a few warm reboots. arch/x86/kernel/tsc.c is
identical in v7.2.6, which I tested, and v7.2.8, so the patch
applies as is. It also applies to v6.18.52 with offsets, but I have
not built or booted that.

Expected: the "Refined TSC clocksource calibration" line shows up
about one second later than without the patch, with about 3593.2
MHz, and the clocksource stays tsc. If that is what you get, please
answer with a "Tested-by: Name <address>" line.

CachyOS kernels carry extra patches. A vanilla stable or mainline
build is the more convincing test, but either helps.

If the step lands in a watchdog interval instead, the TSC is still
marked unstable, with a TSC interval which is longer than the
watchdog one. The patch does not address that.

2. Without the patch: one warm reboot with processor.max_cstate=1 on
the kernel command line. That made the step go away here, as did
keeping one CPU busy, one boot each. If it does the same on your
machine, it is the same mechanism and not just the same symptom.

And three questions:

- How many warm reboots have you looked at, and did all of them lose
the TSC?

- The watchdog messages which you quoted are in the 7.2 format. Does
6.18.52 lose the TSC as well? The watchdog code is different there,
so the dmesg lines of such a boot would be useful.

- Did BIOS 5007 behave the same before the update? "remained present"
reads like it. That would mean that the problem has been there since
2019 firmware and is not specific to 6254.

Please leave the timestamps in the dmesg output, and please reply to
all in plain text, so that the list and the maintainers get it.

Thanks,

Alexander