Re: [PATCH 0/2] Optimise ktime_get_snapshot_id() for capture latency
From: Mark Rutland
Date: Fri Oct 02 2026 - 12:41:39 EST
On Fri, Oct 02, 2026 at 04:19:36PM +0100, David Woodhouse wrote:
> Commit 2e27beeb66e4 ("timekeeping: Allow inlining clocksource::read()")
> allows an architecture's primary implementation of clocksource::read()
> to be inlined, apparently because the out-of-line function calls get
> expensive on x86 with certain speculation mitigations. Use the inline
> read in tk_clock_read_snapshot(), and wire it up for arm64 too.
>
> Also bias the clock_id switch() in ktime_get_snapshot_id() for the
> likely() case of CLOCK_REALTIME.
>
> Measured on a Cortex-A53 at 1.35GHz (12.5MHz arch counter), with
> inlined clocksource reads enabled: the interval from a raw counter
> read in the caller to the counter read in ktime_get_snapshot_id()
> drops from 49.9ns to 41.2ns with the inline read, and to 38.8ns with
> the dispatch bias, mean of 1M iterations.
>
> This is unashamedly a microbenchmark but it actually captures a real
> world use case: The time it takes for ktime_get_snapshot_id() to read
> the counter is critical to some users, especially 1PPS signal
> capture. Every nanosecond helps, and for the inline clocksource read,
> the actually *complex* part is already in place; this is just using
> the infrastructure that's already there.
When I mentioned this on IRC, my complaint about complexity wasn't about
the core code, My conern was with the subtle interactions *within* the
timer driver, and between the timer driver and arch code.
I don't think it's fair to say that the core code is the singular
complex part.
Does this show up on any top-level workload, or is the benefit purely
limited to microbenchmarks and timer synchronization?
Mark.