[PATCH 0/5] timekeeping: Reduce magnitude of ntp_error

From: David Woodhouse

Date: Thu Oct 01 2026 - 16:22:13 EST


The timekeeper's ntp_error variable tracks the delta between the ideal
time that the kernel *should* be reporting as CLOCK_REALTIME, and the
time that is *actually* being reported.

A previous patch series in v7.3, up to commit 794ddd6e15cf ("ntp:
Remove tick_length_base, use tick_length directly"), fixed *errors* in
the accounting of ntp_error — cases where the reported time would
diverge from the ideal time *without* the divergence being correctly
tracked in ntp_error for subsequent compensation.

This series is the next step: it reduces the magnitude of the
excursions which are (now) *correctly* tracked in ntp_error.

The ntp_error mechanism works fine when it remains close to zero. The
actual clock is steered back towards the ideal by a single ±1
dithering of the per-cycle 'mult' period, changing between high and
low each time ntp_error crosses zero.

This doesn't work as well when the value of ntp_error gets higher. In
some cases, 83µs of ntp_error was seen. Draining that through the mult
±1 dithering, at typically less than 1ns/s, could take days. And while
it does, the precision of the frequency that can be effected by the
kernel's timekeeping is limited by the integer quantisation, because
it's permanently biased to either mult or mult+1 according to the
direction of the persistent ntp_error, and frequencies between the two
values cannot be achieved by the normal dithering. On top of which,
chrony was attempting to discipline the *reported* clock, even as
the kernel was attempting to steer it as hard as it could (albeit
not very hard) back to the kernel's ideal.

There are two cases where during frequency/phase adjustment, the
change was effectively applied to the 'ideal' vs 'actual' rate at
different moments in time. This accounting for this in ntp_error was
self-consistent, but the better answer is for them to be conceptually
applied at the same moment, without contributing anything to ntp_error
at all.

While the timekeeping code is applying a phase correction, that
applies an additional bias to the per-cycle 'mult' values. If a
tickless kernel goes idle while this skew is being applied, it can
continue to apply that bias *long* after it was intended to finish,
overshooting the intended phase correction (by 500µs/s in the case of
adjtime). Fix this in timekeeping_max_deferment() by requesting a
wakeup at the top of the next second while phase skew is being
applied. (NB: Only for the core timekeeper, not aux clocks which
remain as they were).

Finally, if ntp_error ever does get excessive, we can deliberately
apply a phase skew to return it to zero faster than the normal
dithering would permit. An equivalent mechanism was removed in commit
dc491596f639 ("timekeeping: Rework frequency adjustments to work
better w/ nohz") but especially with the timekeeping_max_deferment()
change mentioned above, this now works effectively even with (and
in fact especially with) tickless idle. But again only for the core
timekeeper... unless we want to wake for aux clocks too?

Tests running at https://david.woodhou.se/ntptest-virt/ on baseline
vs. patched kernels, both tickless and periodic. The effect is seen
most clearly on the tickless kernel, where the baseline experiences
much higher jitter around corrections that should be smooth.

Phase 1 was *correcting* the accounting of ntp_error. Phase 2 is
*reducing* the amount of ntp_error. Tacked on at the end, marked
DO NOT MERGE for now, is phase 3: *correcting* the output of
ktime_get_snapshot_id() for ntp_error at the moment of the snapshot.

This warrants further discussion, which is why it's being held back
for now. The logic is that snapshot users want *accurate* time, and
they don't want to be given a time that we *know* is microseconds
wrong just because we happened to be in a tickless kernel, running
idle for a large period accumulating divergence in ntp_error right
before they asked for their snapshot.

I vacillated about putting the correction as a separate field in the
snapshot for callers to apply only if they want it, but I think it
makes more sense just to return the corrected time in *all* cases,
unconditionally. Nobody *wants* inaccurate time. Except userspace,
where considerations about smoothness and monotonicity override.

David Woodhouse (5):
timekeeping: Allow tick_length changes to apply mid-tick
ntp: Recalculate skew_delta when the phase offset changes
timekeeping: Bound idle sleep while a phase slew is in flight
timekeeping: Reinstate proportional correction of ntp_error
timekeeping: Apply extrapolated ntp_error to clock snapshots

include/linux/timekeeper_internal.h | 10 +-
kernel/time/ntp.c | 144 +++++++++++++++----------
kernel/time/timekeeping.c | 162 ++++++++++++++++++++++++++--
3 files changed, 249 insertions(+), 67 deletions(-)

--
2.43.0