On Thu, Aug 13, 2026 at 3:59 PM Richard Biener <[email protected]> wrote: > > On Wed, Aug 12, 2026 at 4:25 PM H.J. Lu <[email protected]> wrote: > > > > Double unaligned load and store cost if the vector mode size is less > > than 8 bytes and the number of vector element is less than 4 when SSE > > is enabled to avoid using vector instructions on unaligned short data > > with 2 vector elements so that for > > > > extern char *var1; > > extern int var2; > > > > void > > func (void) > > { > > var2 = var1[1] + var1[0]; > > } > > > > we generate > > > > movq var1(%rip), %rdx > > movsbl 1(%rdx), %eax > > movsbl (%rdx), %edx > > addl %edx, %eax > > movl %eax, var2(%rip) > > > > instead of > > > > movq var1(%rip), %rax > > pxor %xmm1, %xmm1 > > pinsrw $0, (%rax), %xmm0 > > pcmpgtb %xmm0, %xmm1 > > punpcklbw %xmm1, %xmm0 > > movdqa %xmm0, %xmm1 > > psraw $15, %xmm1 > > punpcklwd %xmm1, %xmm0 > > movd %xmm0, %edx > > pshufd $0xe5, %xmm0, %xmm2 > > movd %xmm2, %eax > > addl %edx, %eax > > movl %eax, var2(%rip) > > > > with -O2 -march=x86-64. > > > > Compile PR 125100 tests with -mno-sse since unaligned V2QImode load is no > > longer generated when SSE is enabled. > > Instead of just doubling the load/store cost can we try to more accurately > model the cost of loading of 1, 2 or 4 byte vectors to SSE registers? > For example with SSE4 we get > > movq var1(%rip), %rax > pinsrw $0, (%rax), %xmm0 > pmovsxbd %xmm0, %xmm0 > movd %xmm0, %edx > pextrd $1, %xmm0, %eax > addl %edx, %eax > movl %eax, var2(%rip) > ret
Is this really better than movq var1(%rip), %rdx movswl 2(%rdx), %eax movswl (%rdx), %edx addl %edx, %eax movl %eax, var2(%rip) > and 4 bytes with SSE can be loaded via > > movq var1(%rip), %rax > movd (%rax), %xmm0 > > (why the indirection through rax?). So it's only 2 byte vectors that var1 is a pointer. 2 shorts have the same issue. > are problematic > when SSE4 is not available? > > In general we're missing to model the fact that modern x86 uarchs can happily > execute two streams of same scalar integer operations in parallel and using > vector intructions for them a) increases code size due to larger insn > encodings, > b) for some uarchs has GPR <-> XMM move penalties. So I was thinking > that while the vectorizer accepts scalar-cost == vector-cost as being > profitable > to vectorize, at least for integer code with only two lanes, there should be a > larger profitability gap. We can see to adjust the finish_cost hook in case > of BB vectorization to implement such heuristic, like requiring 66% of the > scalar cost? Benefits/costs of the surrounding code are of course not > and difficult to evaluate (like register pressure on both sides, or EU > occupancy). > > So I think you want to be just testing the vector mode size, not element size, > and size 4 looks OK to me. You also want to handle pinsrw availability and > not double cost (load cost is target dependent) but instead cost a GPR load > plus the relevant vector instruction cost. > > Richard. > > > gcc/ > > > > PR target/126802 > > * config/i386/i386.cc (sse_adjust_unaligned_cost): New. > > (ix86_default_vector_cost): Call sse_adjust_unaligned_cost for > > unaligned load and store to adjust unaligned move cost. > > > > testsuite/ > > > > PR target/126802 > > * gcc.target/i386/pr125100-1.c: Add -mno-sse. > > * gcc.target/i386/pr125100-2.c: Likewise. > > * gcc.target/i386/pr125100-3.c: Likewise. > > * gcc.target/i386/pr126802-1a.c: New test. > > * gcc.target/i386/pr126802-1b.c: Likewise. > > * gcc.target/i386/pr126802-2a.c: Likewise. > > * gcc.target/i386/pr126802-2b.c: Likewise. > > * gcc.target/i386/pr126802-3a.c: Likewise. > > * gcc.target/i386/pr126802-3b.c: Likewise. > > * gcc.target/i386/pr126802-4a.c: Likewise. > > * gcc.target/i386/pr126802-4b.c: Likewise. > > > > Tested on Linux/x86-64 without any regressions. > > > > -- > > H.J. -- H.J.
