Re: [PATCH 0/7] arm64: Batch PSTATE.TCO handling in kernel nofault loops

From: Catalin Marinas

Date: Mon Oct 05 2026 - 10:17:56 EST


On Mon, Aug 24, 2026 at 05:04:45PM +0100, Muhammad Usama Anjum wrote:
> With Hardware Tag-Based KASAN in asynchronous or asymmetric mode, kernel
> nofault loops currently set and clear PSTATE.TCO around every access.
>
> Introduce bare nofault accessors, batching hooks, and an internal scope
> guard. Convert maccess page-fault cleanup to scoped form. Skip
> page-fault setup for zero-sized kernel nofault copies and zero-length
> BPF string operations. Then use bare primitives in the maccess and BPF
> loops. For non-empty operations, this reduces the code-derived dynamic
> MSR TCO execution count from 2N to 2. Two-string BPF comparisons fall
> from 4N to 2.
>
> The BPF changes and an earlier maccess implementation with explicit
> cleanup were tested with QEMU arm64 using Hardware Tag-Based KASAN in
> synchronous, asynchronous, and asymmetric modes. All nine arm64 MTE
> kselftests passed in each mode, as did the 138 focused BPF string_kfuncs
> and varlen subtests. The focused BPF tests also passed with the default
> non-MTE arm64 CPU model. Both final maccess scoped-guard patches were
> arm64 cross-built. Runtime tests have not been rerun after that
> conversion.

TBH, I fail to see the benefit. There's a reduction in the number of
MSR instructions executed in async/asymm mode but does it result in any
improved benchmark numbers? Which workload hits these loops often enough
to matter? They are mostly used by tracing and debug code.

AFAIK, most people wanting to use KASAN in a non-debug environment want
to go for sync mode, where there is no MSR and this series only saves a
few NOPs. We might as well go for a config option to force sync mode (or
off) and remove the unnecessary NOPs, *if* you can show any performance
improvement.

--
Catalin