[PATCH 0/2] Optimise ktime_get_snapshot_id() for capture latency
From: David Woodhouse
Date: Fri Oct 02 2026 - 11:20:18 EST
Commit 2e27beeb66e4 ("timekeeping: Allow inlining clocksource::read()")
allows an architecture's primary implementation of clocksource::read()
to be inlined, apparently because the out-of-line function calls get
expensive on x86 with certain speculation mitigations. Use the inline
read in tk_clock_read_snapshot(), and wire it up for arm64 too.
Also bias the clock_id switch() in ktime_get_snapshot_id() for the
likely() case of CLOCK_REALTIME.
Measured on a Cortex-A53 at 1.35GHz (12.5MHz arch counter), with
inlined clocksource reads enabled: the interval from a raw counter
read in the caller to the counter read in ktime_get_snapshot_id()
drops from 49.9ns to 41.2ns with the inline read, and to 38.8ns with
the dispatch bias, mean of 1M iterations.
This is unashamedly a microbenchmark but it actually captures a real
world use case: The time it takes for ktime_get_snapshot_id() to read
the counter is critical to some users, especially 1PPS signal
capture. Every nanosecond helps, and for the inline clocksource read,
the actually *complex* part is already in place; this is just using
the infrastructure that's already there.
A further use case for arch_inlined_clocksource_read() is a polling
mode PPS driver¹, which only needs to snapshot the *counter* around
its GPIO polling, and can then worry about converting it to a
timestamp after the fact.
¹ https://lore.kernel.org/all/24e8134b35cc8be26606d7472295cdb43ea6b38a.camel@xxxxxxxxxxxxx/
David Woodhouse (2):
arm64: Support inlined clocksource reads for the arch counter
timekeeping: Use inlined counter read in ktime_get_snapshot_id()
arch/arm64/Kconfig | 1 +
arch/arm64/include/asm/clock_inlined.h | 26 ++++++
drivers/clocksource/arm_arch_timer.c | 21 +++++
kernel/time/timekeeping.c | 110 ++++++++++++++++---------
4 files changed, 121 insertions(+), 37 deletions(-)
create mode 100644 arch/arm64/include/asm/clock_inlined.h
--
2.43.0