Re: [PATCH] arm64: archrandom: avoid trapping ID register read in __cpu_has_rng()

From: Aman Priyadarshi

Date: Fri Jul 31 2026 - 14:34:09 EST


Hi Will,

Thank you for reviewing the patch!

> On 31 Jul 2026, at 15:43, Will Deacon <will@xxxxxxxxxx> wrote:
>
> Hi Aman,
>
> On Mon, Jul 20, 2026 at 03:06:15PM +0100, Aman Priyadarshi wrote:
>> __cpu_has_rng() has an early-boot fallback, taken before the
>> ARM64_HAS_RNG alternative is patched, that calls
>> this_cpu_has_cap(ARM64_HAS_RNG). With SCOPE_LOCAL_CPU that resolves the
>> capability by reading ID_AA64ISAR0_EL1 directly from hardware via
>> __read_sysreg_by_encoding(), on every invocation.
>>
>> Until the CRNG is seeded, crng_make_state() routes every get_random_*()
>> through extract_entropy(), which drains architectural entropy via
>> arch_get_random_seed_longs()/arch_get_random_longs() and so calls
>> __cpu_has_rng() several times per request. On a direct (non-EFI) boot
>> there is no bootloader seed, so the CRNG stays unseeded for much of
>> boot and essentially every early randomness consumer takes this path.
>>
>> Under virtualization this is costly: the hypervisor traps guest
>> accesses to the ID registers (HCR_EL2.TID3), making each read a vmexit,
>> producing ~200k trapped ID_AA64ISAR0_EL1 reads during boot.
>>
>> The register value is invariant, so read the sanitised feature register
>> instead. read_sanitised_ftr_reg() returns the cached value from
>> arm64_ftr_regs[] with no sysreg access, and hence no trap. That array
>> is populated by cpuinfo_store_boot_cpu() in smp_prepare_boot_cpu(),
>> before the first early RNG use in random_init_early(), so it is always
>> valid here.
>>
>> With this change the trapped reads drop from ~200k to handful number of
>> times, and the boot time drops roughly by 6.3% in the test environment.
>
> Yikes, that's quite a compelling performance improvement.
>
>> diff --git a/arch/arm64/include/asm/archrandom.h b/arch/arm64/include/asm/archrandom.h
>> index 8babfbe31f95..8067e9a35641 100644
>> --- a/arch/arm64/include/asm/archrandom.h
>> +++ b/arch/arm64/include/asm/archrandom.h
>> @@ -61,8 +61,22 @@ static inline bool __arm64_rndrrs(unsigned long *v)
>>
>> static __always_inline bool __cpu_has_rng(void)
>> {
>> - if (unlikely(!system_capabilities_finalized() && !preemptible()))
>> - return this_cpu_has_cap(ARM64_HAS_RNG);
>> + if (unlikely(!system_capabilities_finalized() && !preemptible())) {
>> + /*
>> + * Until the ARM64_HAS_RNG alternative is patched we can't use
>> + * the static-branch form, so consult the feature register
>> + * directly. Don't use this_cpu_has_cap() here: it reads
>> + * ID_AA64ISAR0_EL1 from hardware on every call, under
>> + * virtualization each ID register read traps to the hypervisor
>> + * (HCR_EL2.TID3) -- producing a storm of vmexits during boot.
>> + * The sanitised value is cached in memory.
>> + */
>> + u64 isar0 = read_sanitised_ftr_reg(SYS_ID_AA64ISAR0_EL1);
>> +
>> + return cpuid_feature_extract_unsigned_field(isar0,
>> + ID_AA64ISAR0_EL1_RNDR_SHIFT) >=
>> + ID_AA64ISAR0_EL1_RNDR_IMP;
>> + }
>
> You can probably rewrite this a little more cleanly along the lines of
> the (not even compile-tested) diff below. I was about to do that, but
> then I got a bit confused by the whole thing. The preemptible() check is
> presumably not needed if we're accessing the in-memory feature registers
> rather than the per-CPU id registers, but then how do you handle races
> with concurrent updates to the "safe value" made by CPUs concurrently
> coming online?

I kept preemptible() check for this exact reason: the updates made by CPUs
concurrently coming online will always take the downgrade path (a secondary
CPU can clear RNDR, never set it), and therefore by taking the non-preemptible
branch I can guarantee that the pinned CPU supports the said feature.
I agree, this assumes that a secondary CPU folds its own ID registers into sys_val
before it can ever be a randomness consumer, but looking at the code that seems
the case, please feel free to correct me.
Besides, in my opinion, it's hard to argue correctness of this code without
preemptible() check.

>
> Will
>
> --->8
>
> diff --git a/arch/arm64/include/asm/archrandom.h b/arch/arm64/include/asm/archrandom.h
> index 8babfbe31f95..1c6ccc1776cd 100644
> --- a/arch/arm64/include/asm/archrandom.h
> +++ b/arch/arm64/include/asm/archrandom.h
> @@ -61,8 +61,17 @@ static inline bool __arm64_rndrrs(unsigned long *v)
>
> static __always_inline bool __cpu_has_rng(void)
> {
> - if (unlikely(!system_capabilities_finalized() && !preemptible()))
> - return this_cpu_has_cap(ARM64_HAS_RNG);
> + if (unlikely(!system_capabilities_finalized() && !preemptible())) {
> + /*
> + * Query the in-memory sanitised value to avoid a potential
> + * trap when accessing the ID register under a hypervisor.
> + */
> + u64 isar0 = read_sanitised_ftr_reg(SYS_ID_AA64ISAR0_EL1);
> + u64 rndr = SYS_FIELD_GET(ID_AA64ISAR0_EL1, RNDR, isar0);
> +
> + return rndr >= ID_AA64ISAR0_EL1_RNDR_IMP;
> + }
> +
> return alternative_has_cap_unlikely(ARM64_HAS_RNG);
> }

Agreed, this looks cleaner. Thanks!

- Aman Priyadarshi