Re: [PATCH RESEND] xor: add missing vzeroupper to AVX code
From: David Laight
Date: Wed Sep 02 2026 - 12:31:51 EST
On Wed, 2 Sep 2026 15:37:06 +0200
Christoph Hellwig <hch@xxxxxx> wrote:
> On Mon, Aug 31, 2026 at 02:22:48PM -0700, Eric Biggers wrote:
> > Since the AVX optimized XOR code uses YMM registers, execute vzeroupper
> > before returning from it. This is needed to avoid degrading the
> > performance of any later SSE code that may happen to be executed.
> >
> > Fixes: ea4d26ae24e5 ("raid5: add AVX optimized RAID5 checksumming")
> > Cc: stable@xxxxxxxxxxxxxxx
> > Signed-off-by: Eric Biggers <ebiggers@xxxxxxxxxx>
> > ---
> >
> > This didn't get taken through the x86 tree. Andrew, it seems you're
> > taking patches to lib/raid/. Can you apply this one?
>
> Can we do kernel_avx_{begin,end} instead of having to open code
> and document this everywhere, please?
In which case I think you want the vzeroupper in kernel_avx_begin().
Actually, for some cpu at least, you need vzeroupper in kernel_fpu_begin()
even if the code only uses the SSE registers.
See: https://stackoverflow.com/questions/41303780/why-is-this-sse-code-6-times-slower-without-vzeroupper-on-skylake
Basically, on Skylake, all SSE hit a penalty if any ymm high bits might be non-zero.
Don't know what has changed since...
I don't have the Intel optimisation manual downloaded (ok, I might have
it but have NFI where), and the link on that page is broken.
David