On Wed, Aug 12, 2026 at 4:25 PM H.J. Lu <[email protected]> wrote:
>
> Double unaligned load and store cost if the vector mode size is less
> than 8 bytes and the number of vector element is less than 4 when SSE
> is enabled to avoid using vector instructions on unaligned short data
> with 2 vector elements so that for
>
> extern char *var1;
> extern int var2;
>
> void
> func (void)
> {
>   var2 = var1[1] + var1[0];
> }
>
> we generate
>
> movq var1(%rip), %rdx
> movsbl 1(%rdx), %eax
> movsbl (%rdx), %edx
> addl %edx, %eax
> movl %eax, var2(%rip)
>
> instead of
>
> movq var1(%rip), %rax
> pxor %xmm1, %xmm1
> pinsrw $0, (%rax), %xmm0
> pcmpgtb %xmm0, %xmm1
> punpcklbw %xmm1, %xmm0
> movdqa %xmm0, %xmm1
> psraw $15, %xmm1
> punpcklwd %xmm1, %xmm0
> movd %xmm0, %edx
> pshufd $0xe5, %xmm0, %xmm2
> movd %xmm2, %eax
> addl %edx, %eax
> movl %eax, var2(%rip)
>
> with -O2 -march=x86-64.
>
> Compile PR 125100 tests with -mno-sse since unaligned V2QImode load is no
> longer generated when SSE is enabled.

Instead of just doubling the load/store cost can we try to more accurately
model the cost of loading of 1, 2 or 4 byte vectors to SSE registers?
For example with SSE4 we get

        movq    var1(%rip), %rax
        pinsrw  $0, (%rax), %xmm0
        pmovsxbd        %xmm0, %xmm0
        movd    %xmm0, %edx
        pextrd  $1, %xmm0, %eax
        addl    %edx, %eax
        movl    %eax, var2(%rip)
        ret

and 4 bytes with SSE can be loaded via

        movq    var1(%rip), %rax
        movd    (%rax), %xmm0

(why the indirection through rax?).  So it's only 2 byte vectors that
are problematic
when SSE4 is not available?

In general we're missing to model the fact that modern x86 uarchs can happily
execute two streams of same scalar integer operations in parallel and using
vector intructions for them a) increases code size due to larger insn encodings,
b) for some uarchs has GPR <-> XMM move penalties.  So I was thinking
that while the vectorizer accepts scalar-cost == vector-cost as being profitable
to vectorize, at least for integer code with only two lanes, there should be a
larger profitability gap.  We can see to adjust the finish_cost hook in case
of BB vectorization to implement such heuristic, like requiring 66% of the
scalar cost?  Benefits/costs of the surrounding code are of course not
and difficult to evaluate (like register pressure on both sides, or EU
occupancy).

So I think you want to be just testing the vector mode size, not element size,
and size 4 looks OK to me.  You also want to handle pinsrw availability and
not double cost (load cost is target dependent) but instead cost a GPR load
plus the relevant vector instruction cost.

Richard.

> gcc/
>
> PR target/126802
> * config/i386/i386.cc (sse_adjust_unaligned_cost): New.
> (ix86_default_vector_cost): Call sse_adjust_unaligned_cost for
> unaligned load and store to adjust unaligned move cost.
>
> testsuite/
>
> PR target/126802
> * gcc.target/i386/pr125100-1.c: Add -mno-sse.
> * gcc.target/i386/pr125100-2.c: Likewise.
> * gcc.target/i386/pr125100-3.c: Likewise.
> * gcc.target/i386/pr126802-1a.c: New test.
> * gcc.target/i386/pr126802-1b.c: Likewise.
> * gcc.target/i386/pr126802-2a.c: Likewise.
> * gcc.target/i386/pr126802-2b.c: Likewise.
> * gcc.target/i386/pr126802-3a.c: Likewise.
> * gcc.target/i386/pr126802-3b.c: Likewise.
> * gcc.target/i386/pr126802-4a.c: Likewise.
> * gcc.target/i386/pr126802-4b.c: Likewise.
>
> Tested on Linux/x86-64 without any regressions.
>
> --
> H.J.

Reply via email to