Re: [PATCH v4 3/4] pps: Always use ktime_get_snapshot_id() for pps_get_ts()
From: David Woodhouse
Date: Mon Sep 28 2026 - 09:39:53 EST
On Mon, 2026-09-28 at 09:58 +0200, Rodolfo Giometti wrote:
> On Sat, 2026-09-26 at 21:38 +0100, David Woodhouse wrote:
> > So yes, the specific code path you're looking at *does* get slightly
> > longer (50ns to the counter read instead of 30ns).
> [...]
> > However, they are *entirely* in the noise, as there's about 600 ns of
> > hardware and 2-4 *microseconds* of software latency before we even get
> > there.
>
> Thanks for measuring it, and on real hardware with a real edge. That
> answers my concern: ~20 ns of constant cost against microseconds of
> latency upstream of the handler is not something PPS can see, and
> trading it for the removal of a non-constant error is the right
> trade. It is arm64 rather than the 32-bit board I asked about, but I
> accept the argument holds there too, since the software latency only
> gets larger.
>
> There is still one thing I want to be sure we agree on, about what
> ts_real means for userspace.
>
> Today, without CONFIG_NTP_PPS, ts_real comes from ktime_get_real_ts64(),
> which is the same value userspace reads with
> clock_gettime(CLOCK_REALTIME). A PPS timestamp and a clock_gettime()
> reading are identical by construction: both are the sanitized clock.
> With CONFIG_NTP_PPS the same held so far, since ktime_get_snapshot_id()
> also returned the sanitized value.
>
> After 1/4 and this patch that is no longer true. Your 1/4 says it
> explicitly: callers of ktime_get_snapshot_id() now receive the ideal
> time, "not the sanitized version provided to gettimeofday()". So ts_real
> becomes the ideal line, while clock_gettime(CLOCK_REALTIME) keeps
> returning the sanitized one, and the two differ by ntp_error at the
> instant of the edge.
>
> For hardpps this is clearly what we want, and your 1/4 and 2/4 make
> that case. For userspace I am less sure. chrony and ntpd compare PPS
> timestamps against other sources whose timestamps come from
> clock_gettime(), i.e. from the sanitized clock, so after this series
> the two would be referenced to different lines, ntp_error apart.
> Miroslav, is that something chrony would notice? And David, how large
> can that difference get in practice? Your 1/4 says the divergence can
> span many ticks under NO_HZ.
Theoretically, absent other bugs (qv), ntp_error should rarely be more
than a few tens of nanoseconds and even that is the extreme case.
It swings slightly positive and negative as the reported CLOCK_REALTIME
goes slightly behind, and ahead of, the ideal time.
This happens because of the integer quantisation of the counter period.
The kernel spends a while running with the base 'mult' and getting
behind, and then switches to 'mult+1' to catch up. On a *tickful*
kernel, it flips between them as soon as ntp_error goes
positive/negative each tick.
On the board I'm testing with, that ±1 difference equates to about
0.745ns/s. Right now, its "ideal" mult calibrated against the PPS
signal is about in the middle (1342189164.4992), so that means it has a
choice of running 0.372ns/s slow, or 0.373ns/s fast. Worst case, the
value could come out very close to an integer, and the ±1 choices could
approximate the full 0.745 in one direction, and almost negligible in
the other.
So on a *tickless* kernel when it doesn't course correct every tick, it
could get set to gain 0.745ns/s and then go to sleep for, say, ten
minutes, and accumulate... half a microsecond.
I don't think ten minutes is really that realistic, of course. And in
practice I definitely can't make nohz_full actually contribute a
meaningful amount to ntp_error while *also* waking it up every second
(or even every 5 seconds) to process a pulse.
However, I *have* seen (and fixed) the tick_length changes at chrony
startup introducing 83µs into ntp_error, which would take *days* to
drain through the normal ±1 dithering, and would screw up the actual
frequency settings while it was draining.
I've pushed out my current WIP to my timekeeping branch¹. The fix for
the 83µs ntp_error is the first² in the series — "timekeeping: Allow
tick_length changes to apply mid-tick".
> Since the main users of ts_real are chrony and ntpd, not the kernel,
> I would like an Acked-by from their maintainers before I ack this
> patch. They are the ones who will have to live with the new semantics,
> so they should know what is coming and agree with it.
Absolutely. I've called this out explicitly in the patch which makes
ktime_get_snapshot_id() do that correction. That's the second patch³ in
my branch — "timekeeping: Apply extrapolated ntp_error to clock
snapshots".
I wonder if we should switch PPS to using ktime_get_snapshot_id() in an
*earlier* patch, which wouldn't then include the behavioural change.
Then the note in the 'Apply extrapolated error' patch can then cover
PPS and we consider them all together.
The PPS change does stand alone anyway — regardless of the snapshot
*corrections*, I want PPS using snapshots so that it can report the
actual *counter* values to userspace, like PTP is going to be able to.
Then userspace can choose to discipline the actual *counter* without
worrying about the feedback loop of the kernel's own timekeeping at
all.
¹ https://git.infradead.org/?p=users/dwmw2/linux.git;a=shortlog;h=refs/heads/timekeeping
² https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=3ff34be9c943
³ https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=bac732b0bf23
> So for v5 please put this in the commit message, not only in the
> thread: why ktime_get_real_ts64() is inaccurate (the quantisation of
> the multiplier), that ts_real is now the ideal NTP-disciplined time
> rather than what clock_gettime() returns, the order of magnitude of
> the difference, and a short summary of the latency numbers above.
> Whoever runs git blame on pps_get_ts() in a few years should not have
> to find this thread.
Ack. I discussed the order of magnitude above. Let's see how my
proposed optimisations land, and I'll update the comments on latency
too.
For the "ts_real is suddenly something different" aspect, I'll defer to
Miroslav et al but I do believe that this is either in the noise at
most levels of precision, or a *correction* where it's even noticeable.
The use cases which consume a snapshot are the ones which return a
tuple of that timestamp against something simultaneous — either the raw
counter value, a PPS pulse which is known to be at the top of a second,
a PTP reading from an external clock. (Worst case, PTP sandwiches two
local timestamps around an external clock reading).
Those use cases run within an actual system call (not vDSO), and are
never about the time "now", per se — they're all about the pairing.
They shouldn't be compared directly with a clock_gettime() where
userspace... is preempted and... calls into the vDSO to get the time...
is preempted again and... eventually does something with that timestamp
which it considers to be current, or worse paired with whatever happens
before or after it.
So I'm not too worried about the fact that the two different mechanisms
return time in a *slightly* different form, because the normal
userspace path chooses speed and monotonicity over accuracy.
Attachment:
smime.p7s
Description: S/MIME cryptographic signature