On Fri, Aug 14, 2026 at 10:44 AM H.J. Lu <[email protected]> wrote:
>
> On Fri, Aug 14, 2026 at 4:13 PM Richard Biener
> <[email protected]> wrote:
> >
> > On Fri, Aug 14, 2026 at 9:28 AM H.J. Lu <[email protected]> wrote:
> > >
> > > On Fri, Aug 14, 2026 at 2:49 PM Richard Biener
> > > <[email protected]> wrote:
> > > >
> > > > On Fri, Aug 14, 2026 at 8:24 AM H.J. Lu <[email protected]> wrote:
> > > > >
> > > > > On Fri, Aug 14, 2026 at 2:21 PM Richard Biener
> > > > > <[email protected]> wrote:
> > > > > >
> > > > > > On Thu, Aug 13, 2026 at 10:32 AM H.J. Lu <[email protected]> 
> > > > > > wrote:
> > > > > > >
> > > > > > > On Thu, Aug 13, 2026 at 3:59 PM Richard Biener
> > > > > > > <[email protected]> wrote:
> > > > > > > >
> > > > > > > > On Wed, Aug 12, 2026 at 4:25 PM H.J. Lu <[email protected]> 
> > > > > > > > wrote:
> > > > > > > > >
> > > > > > > > > Double unaligned load and store cost if the vector mode size 
> > > > > > > > > is less
> > > > > > > > > than 8 bytes and the number of vector element is less than 4 
> > > > > > > > > when SSE
> > > > > > > > > is enabled to avoid using vector instructions on unaligned 
> > > > > > > > > short data
> > > > > > > > > with 2 vector elements so that for
> > > > > > > > >
> > > > > > > > > extern char *var1;
> > > > > > > > > extern int var2;
> > > > > > > > >
> > > > > > > > > void
> > > > > > > > > func (void)
> > > > > > > > > {
> > > > > > > > >   var2 = var1[1] + var1[0];
> > > > > > > > > }
> > > > > > > > >
> > > > > > > > > we generate
> > > > > > > > >
> > > > > > > > > movq var1(%rip), %rdx
> > > > > > > > > movsbl 1(%rdx), %eax
> > > > > > > > > movsbl (%rdx), %edx
> > > > > > > > > addl %edx, %eax
> > > > > > > > > movl %eax, var2(%rip)
> > > > > > > > >
> > > > > > > > > instead of
> > > > > > > > >
> > > > > > > > > movq var1(%rip), %rax
> > > > > > > > > pxor %xmm1, %xmm1
> > > > > > > > > pinsrw $0, (%rax), %xmm0
> > > > > > > > > pcmpgtb %xmm0, %xmm1
> > > > > > > > > punpcklbw %xmm1, %xmm0
> > > > > > > > > movdqa %xmm0, %xmm1
> > > > > > > > > psraw $15, %xmm1
> > > > > > > > > punpcklwd %xmm1, %xmm0
> > > > > > > > > movd %xmm0, %edx
> > > > > > > > > pshufd $0xe5, %xmm0, %xmm2
> > > > > > > > > movd %xmm2, %eax
> > > > > > > > > addl %edx, %eax
> > > > > > > > > movl %eax, var2(%rip)
> > > > > > > > >
> > > > > > > > > with -O2 -march=x86-64.
> > > > > > > > >
> > > > > > > > > Compile PR 125100 tests with -mno-sse since unaligned 
> > > > > > > > > V2QImode load is no
> > > > > > > > > longer generated when SSE is enabled.
> > > > > > > >
> > > > > > > > Instead of just doubling the load/store cost can we try to more 
> > > > > > > > accurately
> > > > > > > > model the cost of loading of 1, 2 or 4 byte vectors to SSE 
> > > > > > > > registers?
> > > > > > > > For example with SSE4 we get
> > > > > > > >
> > > > > > > >         movq    var1(%rip), %rax
> > > > > > > >         pinsrw  $0, (%rax), %xmm0
> > > > > > > >         pmovsxbd        %xmm0, %xmm0
> > > > > > > >         movd    %xmm0, %edx
> > > > > > > >         pextrd  $1, %xmm0, %eax
> > > > > > > >         addl    %edx, %eax
> > > > > > > >         movl    %eax, var2(%rip)
> > > > > > > >         ret
> > > > > > >
> > > > > > > Is this really better than
> > > > > > >
> > > > > > >  movq var1(%rip), %rdx
> > > > > > > movswl 2(%rdx), %eax
> > > > > > > movswl (%rdx), %edx
> > > > > > > addl %edx, %eax
> > > > > > > movl %eax, var2(%rip)
> > > > > >
> > > > > > No, but it's better than the SSE2 version ;)  If you just want to 
> > > > > > prevent
> > > > > > these small vectorizations you can deny any scalar_cost < 30 
> > > > > > vectorizations
> > > > > > or so.  I thought the point is to make the costing more precise - 
> > > > > > you
> > > > > > correctly identified at least HImode -> XMM loads/stores to be not 
> > > > > > accurately
> > > > > > costed.
> > > > > >
> > > > >
> > > > > The problems are V2QImode and V2HImode.  Others seem OK.
> > > >
> > > > But I see movd (%rax), %xmm0 for V2HImode, so that seems fine
> > > > for loads and stores.
> > > >
> > > >         movd    (%rdi), %xmm0
> > > >         paddw   %xmm0, %xmm0
> > > >         movd    %xmm0, x(%rip)
> > > >
> > > > For V2QImode:
> > > >
> > > >         pinsrw  $0, (%rdi), %xmm0
> > > >         paddb   %xmm0, %xmm0
> > > >         movd    %xmm0, %eax
> > > >         movw    %ax, x(%rip)
> > > >
> > > > so even SSE2 has pinsrw for the load but nothing for the store (SSE4
> > > > has pextrw there).
> > > >
> > > > So what am I missing?
> > > >
> > >
> > > For
> > >
> > > extern short *var1;
> > > extern int var2;
> > >
> > > void
> > > func (void)
> > > {
> > >   var2 = var1[1] + var1[0];
> > > }
> > >
> > > -O2 generates:
> > >
> > >  movq var1(%rip), %rax
> > > movd (%rax), %xmm0
> > > movdqa %xmm0, %xmm1
> > > psraw $15, %xmm1
> > > punpcklwd %xmm1, %xmm0
> > > movd %xmm0, %edx
> > > pshufd $0xe5, %xmm0, %xmm2
> > > movd %xmm2, %eax
> > > addl %edx, %eax
> > > movl %eax, var2(%rip)
> > >
> > > With SSE4, we get
> > >
> > > movq var1(%rip), %rax
> > > movd (%rax), %xmm0
> > > pmovsxwd %xmm0, %xmm0
> > > movd %xmm0, %edx
> > > pextrd $1, %xmm0, %eax
> > > addl %edx, %eax
> > > movl %eax, var2(%rip)
> > >
> > > It isn't much better than
> > >
> > > movq var1(%rip), %rdx
> > > movswl 2(%rdx), %eax
> > > movswl (%rdx), %edx
> > > addl %edx, %eax
> > > movl %eax, var2(%rip)
> >
> > I'm not arguing about the profitability of the vectorization.  I'm arguing
> > of wheter the costing of the V2HImode load is wrong.  That doesn't seem
> > to be the case?
> >
>
> Load is one part of the computation.  The total cost of V2HImode
> computation is too low.

I think the total cost of the HImode calculation is too high ;)  Or rather,
load and store costs tend to dominate and for scalar we fail to
realize we can issue the two loads in parallel and for stores we fail
to realize they are irrelevant for the computation of latency.

I'll note that x86 cost tables do not tell us the number of loads
(which might depend on load width) that can issue in parallel,
that is, we only have latency information, not throughput.

We can possibly win most of the two-lane BB reduction cases
by doing a DIV_CEIL (.., 2) on the total number of loads.  For 2 scalar
vs 1 vector load that would cancel out.

That said, we make no attempt at identifying dependence chains
so we add individual operation latencies as if we had a single
linear dependence chain.  That's the most fundamental issue
with computing more accurate profitability.

Richard.

>
>
> --
> H.J.

Reply via email to