Re: [PATCH 0/2] Optimise ktime_get_snapshot_id() for capture latency
From: David Woodhouse
Date: Fri Oct 02 2026 - 13:15:51 EST
On Fri, 2026-10-02 at 17:41 +0100, Mark Rutland wrote:
> On Fri, Oct 02, 2026 at 04:19:36PM +0100, David Woodhouse wrote:
> > Commit 2e27beeb66e4 ("timekeeping: Allow inlining clocksource::read()")
> > allows an architecture's primary implementation of clocksource::read()
> > to be inlined, apparently because the out-of-line function calls get
> > expensive on x86 with certain speculation mitigations. Use the inline
> > read in tk_clock_read_snapshot(), and wire it up for arm64 too.
> >
> > Also bias the clock_id switch() in ktime_get_snapshot_id() for the
> > likely() case of CLOCK_REALTIME.
> >
> > Measured on a Cortex-A53 at 1.35GHz (12.5MHz arch counter), with
> > inlined clocksource reads enabled: the interval from a raw counter
> > read in the caller to the counter read in ktime_get_snapshot_id()
> > drops from 49.9ns to 41.2ns with the inline read, and to 38.8ns with
> > the dispatch bias, mean of 1M iterations.
> >
> > This is unashamedly a microbenchmark but it actually captures a real
> > world use case: The time it takes for ktime_get_snapshot_id() to read
> > the counter is critical to some users, especially 1PPS signal
> > capture. Every nanosecond helps, and for the inline clocksource read,
> > the actually *complex* part is already in place; this is just using
> > the infrastructure that's already there.
>
> When I mentioned this on IRC, my complaint about complexity wasn't about
> the core code, My conern was with the subtle interactions *within* the
> timer driver, and between the timer driver and arch code.
>
> I don't think it's fair to say that the core code is the singular
> complex part.
This patch doesn't touch the *complex* part of the arch_timer support
with a barge-pole. The inline path is only enabled for the trivial case
where !arch_timer_counter_has_wa() and thus the indirect function
pointer is guaranteed to do *precisely* the same thing.
> Does this show up on any top-level workload, or is the benefit purely
> limited to microbenchmarks and timer synchronization?
For PPS, my latest experiments show that even with a tight loop
*polling* the GPIO directly via MMIO on my test board, the MMIO read
takes ~400ns. So I don't think we can claim that the ~8ns I get from
the inlining is critical to the PPS use case, no.
Rodolfo was concerned about the time taken in ktime_get_snapshot_id();
I responded that I think it's in the noise but there's some low-hanging
fruit for making it faster... I think I've fairly much shown that it's
in the noise *however* hard we try to reduce the rest of the latency.
I think think it's low-hanging fruit though; I guess it comes down to
how we view the complexity.
As also discussed on IRC before you pointed out that it's a long way to
April 1st, the most effective answer for PPS by *far* is to sample the
counter as early as possible in the exception handler, and for the PPS
IRQ handler to pull it out of pt_regs.
https://lore.kernel.org/all/dd7a873caecab6c912e2c0f910d1e46818ea455d.camel@xxxxxxxxxxxxx/
I did say I'd post the inline timer patch for posterity and let you
shoot it down, and here it is. I'm still not posting the entry.S one :)
Attachment:
smime.p7s
Description: S/MIME cryptographic signature