Re: [PATCH v4 3/4] pps: Always use ktime_get_snapshot_id() for pps_get_ts()

From: David Woodhouse

Date: Fri Oct 02 2026 - 08:30:24 EST


On Fri, 2026-10-02 at 10:07 +0100, David Woodhouse wrote:
>
> I've modified the spin code to take the shape I outlined in email last
> night, just capturing the *counter* bracketing the GPIO read and then
> converting it to 'time' later. I'll have results in a few hours.

That one got to 5 hours so I harvested it. I accidentally ran it
tickless but that doesn't really matter because all of the tickless
accounting issues are basically resolved anyway *and* the poll involves
waking up just before the top of every second to spin anyway, so it
*isn't* idle when the pulses come in.

• https://david.woodhou.se/ntptest-r64/spin1-tickless-1hz/

It doesn't make a lot of difference to the overall statistics. The
outliers are down, from ±1100 to ±800 ns (negative because if we sync
to a late one, then a subsequent pulse appears "early"). I think that
improvement is because of the timekeeper seqcount retry which will bite
*either* ktime_get_real_ns() or ktime_get_snapshot_id(), and which is
gone in my 'spin1' version which takes it all out of the critical path.

┌────────────────────────────┬─────┬──────┬──────┬──────┬─────┐
│ capture method │ p50 │ p95 │ p99 │ max │ σ │
├────────────────────────────┼─────┼──────┼──────┼──────┼─────┤
│ pps-gpio (IRQ) │ 84 │ 1206 │ 4013 │ 4889 │ 688 │
│ polling, stamp after edge │ 206 │ 536 │ 665 │ 1106 │ 286 │
│ polling, bracketed counter │ 206 │ 545 │ 647 │ 811 │ 286 │
│ entry.S counter capture │ 46 │ 109 │ 135 │ 206 │ 58 │
└────────────────────────────┴─────┴──────┴──────┴──────┴─────┘

I think the remaining jitter is largely due to the time it takes, in
the "tight" polling loop, to call gpio_read(). That's twice what I
estimated — it's 600ns. I'm running a quick test of reading the MMIO
directly in the loop... even that is still about 400ns on this
hardware, which isn't going to move the needle very much either.

The interrupt path — and capturing the counter in entry.S — remains the
winner by far. I'm going to do a run of that in a *live* system. My
existing test runs were targeting the tickless behaviour and
deliberately wanted the system to be as idle as possible, but activity
which causes interrupts to be disabled could push the entry.S hack back
down the leaderboard.

But that ends up being all about the *outliers*, and we already said we
need to do a better job of filtering those. With proper filtering, the
occasional pulse being rejected because we had interrupts disabled
shouldn't *matter*.

FWIW my gpio-spin-on-MMIO running now also seems to show a significant
correlation on outliers. For those samples where the half-delta between
the counter before and after the read is two cycles, the p50 phase
offset is 80ns, while when it's three cycles, the p50 is 175ns. We
could just throw *all* of the samples with a wider bracket away (72 of
112 samples, in the abortive 2-minute run before I fixed something and
restarted it). We're better off without them.

Attachment: smime.p7s
Description: S/MIME cryptographic signature