On Tue, 29 Sep 2026 18:59:39 +0200 Nathan Chancellor <[email protected]> wrote:
> On Sun, Sep 27, 2026 at 08:02:25AM +0100, David Laight wrote: > > On Sat, 26 Sep 2026 14:32:49 +0200 > > "Jason A. Donenfeld" <[email protected]> wrote: > > > Do we even need to return a value at all? Might as well just make the > > > function two lines: > > > > > > + for (char *d = dst; size--;) > > > + *d++ = value; > > > > Wouldn't it be better to add a barrier() or similar in there to > > stop the compiler playing unwanted games> > > Wouldn't that just result in the same code generation issue that Jason > pointed out on my v1? > > https://lore.kernel.org/[email protected]/ I missed that one going past.... You only get 'rep stosq' because the compiler first converts it to memset(). > Maybe that doesn't matter because modern compilers have better options? A lot of cpu will run the 'rep stosq' very slowly (it has a big fixed cost). At a guess the loop is 3 clocks (possibly 2; but that usually needs you to use negative offsets from the end - and gcc doesn't like that). That is comparable to a mispredicted branch (and you might get two of them). It still might actually be faster than the 'rep stosq' version! I suspect the fastest code is to unroll the loop. xor %eax, %eax movl %eax, 12(%rbx) movq %rax, 16(%rbx) movq %rax, 24(%rbx) movq %rax, 32(%rbx) movq %rax, 40(%rbx) movq %rax, 48(%rbx) movq %rax, 56(%rbx) 8 clocks on old cpu, 4 on newer ones. But more likely to be limited be I-cache reads. If you are going to use 'rep stos' then you might as well write 13 32bit words. David
