Re: [PATCH v4 3/4] pps: Always use ktime_get_snapshot_id() for pps_get_ts()
From: David Woodhouse
Date: Sat Sep 26 2026 - 16:38:47 EST
On Tue, 2026-09-01 at 17:35 +0200, Rodolfo Giometti wrote:
>
> Separately: are we sure that calling ktime_get_snapshot_id() does not
> introduce much larger delays than ktime_get_real_ts64()? pps_get_ts()
> runs in hard IRQ, and in pps-gpio it is the first statement of the
> handler. Whatever it costs sits between the edge and the timestamp,
> and that is the one thing PPS has to keep short.
>
> I would like to see that measured on something other than an x86 VM
> with a TSC. A small 32-bit ARM board is what I worry about.
Not 32-bit but I had a Banana Pi R64 lying around (Cortex A53, 12.5MHz
arch counter) and I bought it a GPS hat.
So yes, the specific code path you're looking at *does* get slightly
longer (50ns to the counter read instead of 30ns). Some of which is
easily reclaimable with some optimisations...
Firstly, why in $DEITY's name does the Arm kernel not set
CONFIG_ARCH_WANTS_CLOCKSOURCE_READ_INLINE? Setting that and using it
gains us 7-8ns back in both paths. I can get another 2-3ns back by
optimising the CLOCK_REALTIME path through ktime_get_snapshot_id() and
only hitting the case statement for other clock IDs:
┌───────────────┬─────────────────┐
│ ktime_get_ │ ktime_get_ │
│ real_ts64() │ snapshot_id() │
────────────────────────────────────┼───────────────┼─────────────────┤
stock │ 30.2 ns │ 49.9 ns │
+ arm64 inlined clocksource read │ 23.4 ns │ 49.9 ns │
+ inlined in get_snapshot_id() too │ 23.2 ns │ 41.2 ns │
+ CLOCK_REALTIME dispatch bias │ 23.4 ns │ 38.8 ns │
────────────────────────────────────┴───────────────┴─────────────────┘
So I'll probably do those things because I've seen them now.
However, they are *entirely* in the noise, as there's about 600 ns of
hardware and 2-4 *microseconds* of software latency before we even get
there.
I was pondering an IRQF_HWTIMESTAMP which would sample the arch-defined
inline clocksource as early as possible and stash it in the pt_regs, so
I knocked up a proof of concept which just did that unconditionally in
kernel_entry in arch/arm64/kernel/entry.S. That's what gives me the
software latency I cited (~2.1µs from that stamp to the GPIO handler).
To measure the hardware latency I removed the hat and looped the PPS
pin back from an output GPIO, then read the counter the instruction
before the MMIO write to trigger a rising edge. That gave me 0.6µs to
the counter read in entry.S.
pin edge → exception entry stamp: min 0.40µs med 0.64µs p90 0.72µs
entry stamp → pps handler: min 2.08µs med 2.16µs p90 2.56µs
pin edge → pps handler (sum): min 2.72µs med 2.80µs p90 3.12µs
The latency is higher from cold with an actual PPS signal, as opposed
to the warm soak test.
So no, I really don't care about the tiny cost of calling
ktime_get_snapshot_id() to get a more accurate timestamp; at least that
cost is *constant* unlike the sawtoothing ntp_error that it eliminates
from the readings. (Which is actually larger than it should be; I have
more fixes for ntp_error.)
I'm still interested in the IRQF_HWTIMESTAMP thing. As that sample
would be taken outside the tkd->seq count we'd need to use something
like get_system_device_crosststamp() to reliably interpret it, but I'm
going to have to implement that *anyway*. I want to add support for a
hardware module which *latches* the counter value on the pulse, instead
of waiting for the CPU to wake up and do so for itself. One of those
TimeCard devices with PCIe PTM could do that, for example.
Attachment:
smime.p7s
Description: S/MIME cryptographic signature