Re: [PATCH] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
From: Mark Rutland
Date: Tue Sep 15 2026 - 06:21:40 EST
On Tue, Sep 15, 2026 at 02:50:27AM -0700, Breno Leitao wrote:
> Hello Mark,
>
> On Tue, Sep 15, 2026 at 09:39:24AM +0100, Mark Rutland wrote:
> > Hi Breno,
> >
> > I think the change looks reasonable, but the commit message and comments
> > aren't quite right. More on that below.
>
> Thank you very much for your review. I know this is not a trivial one
> (at least from my PoV), I am glad you quickly reviewed it.
>
> I've also dropped few other lines, but kept the benchmark values I've
> collected. Does this look better now?
Yep, that looks good to me, with one minor nit below.
With that fixed up, this all looks good. I assume you'll send a v2.
> Author: Breno Leitao <leitao@xxxxxxxxxx>
> Date: Fri Sep 11 02:57:34 2026 -0700
>
> arm64/sve: Don't zero the SVE state buffer when the SVE state is live
>
> Currently do_sve_acc() always zeroes current->thread.sve_state. This is
> not necessary in the common case, and avoiding the zeroing has a
> measurable impact on some benchmarks.
>
> In the common case where the task is not preempted and its state is not
> altered by a tracer, do_sve_acc() will observe that TIF_FOREIGN_FPSTATE
> is clear. In such cases, only the live register values matter, and the
> in-memory copy is stale regardless of whether it is saved in
> FP_STATE_FPSIMD format or FP_STATE_SVE format.
>
> This is worth doing because the SVE state is discarded on syscall entry,
^^^^^^^^^^^^^^^^^^^^^^^^^^^
That should say something like "It is worth skipping the zeroing
because". We deleted the line saying that skipping the zeroing was safe,
and so it's not clear what "this" is referring to.
Mark.
> so userspace that mixes SVE and syscalls re-traps constantly. A fleet
> profile of arm64 hosts running services whose memset() is SVE shows the
> memset under do_sve_acc() accounting for 29% of the trap handling cost.
>
> Measured on a 72-core Neoverse V2 (SVE VL 128, sve_state_size 546,
> performance governor) with perf bench sched pipe pinned to one CPU, and
> SVE operation on write, so that each loop also takes an SVE access trap.
>
> * -0.99% kernel instructions
> * -1.38% kernel cycles
> * -1.12% wall clock
>
> Signed-off-by: Breno Leitao <leitao@xxxxxxxxxx>
>
> diff --git a/arch/arm64/kernel/fpsimd.c b/arch/arm64/kernel/fpsimd.c
> index e7f1682a3059b..324c9799b0511 100644
> --- a/arch/arm64/kernel/fpsimd.c
> +++ b/arch/arm64/kernel/fpsimd.c
> @@ -1316,7 +1316,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
> return;
> }
>
> - sve_alloc(current, true);
> + sve_alloc(current, false);
> if (!current->thread.sve_state) {
> force_sig(SIGKILL);
> return;
> @@ -1341,6 +1341,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
> sve_flush_live();
> fpsimd_bind_task_to_cpu();
> } else {
> + memset(current->thread.sve_state, 0, sve_state_size(current));
> fpsimd_to_sve(current);
> current->thread.fp_type = FP_STATE_SVE;
> fpsimd_flush_task_state(current);