Re: [PATCH v2] random: vDSO: Avoid call to memset() when zeroing reserved in __cvdso_getrandom_data()

From: David Laight

Date: Tue Sep 29 2026 - 13:47:37 EST


On Tue, 29 Sep 2026 18:59:39 +0200
Nathan Chancellor <nathan@xxxxxxxxxx> wrote:

> On Sun, Sep 27, 2026 at 08:02:25AM +0100, David Laight wrote:
> > On Sat, 26 Sep 2026 14:32:49 +0200
> > "Jason A. Donenfeld" <Jason@xxxxxxxxx> wrote:
> > > Do we even need to return a value at all? Might as well just make the
> > > function two lines:
> > >
> > > + for (char *d = dst; size--;)
> > > + *d++ = value;
> >
> > Wouldn't it be better to add a barrier() or similar in there to
> > stop the compiler playing unwanted games>
>
> Wouldn't that just result in the same code generation issue that Jason
> pointed out on my v1?
>
> https://lore.kernel.org/aquyVkdYceiWdckX@xxxxxxxxx/

I missed that one going past....
You only get 'rep stosq' because the compiler first converts it to memset().

> Maybe that doesn't matter because modern compilers have better options?

A lot of cpu will run the 'rep stosq' very slowly (it has a big fixed cost).

At a guess the loop is 3 clocks (possibly 2; but that usually needs
you to use negative offsets from the end - and gcc doesn't like that).
That is comparable to a mispredicted branch (and you might get two of them).
It still might actually be faster than the 'rep stosq' version!

I suspect the fastest code is to unroll the loop.
xor %eax, %eax
movl %eax, 12(%rbx)
movq %rax, 16(%rbx)
movq %rax, 24(%rbx)
movq %rax, 32(%rbx)
movq %rax, 40(%rbx)
movq %rax, 48(%rbx)
movq %rax, 56(%rbx)
8 clocks on old cpu, 4 on newer ones.
But more likely to be limited be I-cache reads.

If you are going to use 'rep stos' then you might as well write
13 32bit words.

David