Re: [PATCH 1/2] xtensa: Optimize '(x & CST1_POW2) != 0 ? CST2_POW2 : 0'
On Mon, May 22, 2023 at 7:28 PM Max Filippov via Gcc-patches wrote: > > Hi Suwa-san, > > On Mon, May 22, 2023 at 12:06 AM Takayuki 'January June' Suwa > wrote: > > > > This patch decreses one machine instruction from "single bit extraction > > with shifting" operation, and tries to eliminate the conditional > > branch if CST2_POW2 doesn't fit into signed 12 bits with the help > > of ifcvt optimization. > > > > /* example #1 */ > > int test0(int x) { > > return (x & 1048576) != 0 ? 1024 : 0; > > } > > extern int foo(void); > > int test1(void) { > > return (foo() & 1048576) != 0 ? 16777216 : 0; > > } > > > > ;; before > > test0: > > movia9, 0x400 > > sraia2, a2, 10 > > and a2, a2, a9 > > ret.n > > test1: > > addisp, sp, -16 > > s32i.n a0, sp, 12 > > call0 foo > > extui a2, a2, 20, 1 > > sllia2, a2, 20 > > beqz.n a2, .L2 > > movi.n a2, 1 > > sllia2, a2, 24 > > .L2: > > l32i.n a0, sp, 12 > > addisp, sp, 16 > > ret.n > > > > ;; after > > test0: > > extui a2, a2, 20, 1 > > sllia2, a2, 10 > > ret.n > > test1: > > addisp, sp, -16 > > s32i.n a0, sp, 12 > > call0 foo > > l32i.n a0, sp, 12 > > extui a2, a2, 20, 1 > > sllia2, a2, 24 > > addisp, sp, 16 > > ret.n > > > > In addition, if the left shift amount ('exact_log2(CST2_POW2)') is > > between 1 through 3 and a either addition or subtraction with another > > register follows, emit a ADDX[248] or SUBX[248] machine instruction > > instead of separate left shift and add/subtract ones. > > > > /* example #2 */ > > int test2(int x, int y) { > > return ((x & 1048576) != 0 ? 4 : 0) + y; > > } > > int test3(int x, int y) { > > return ((x & 2) != 0 ? 8 : 0) - y; > > } > > > > ;; before > > test2: > > movi.n a9, 4 > > sraia2, a2, 18 > > and a2, a2, a9 > > add.n a2, a2, a3 > > ret.n > > test3: > > movi.n a9, 8 > > sllia2, a2, 2 > > and a2, a2, a9 > > sub a2, a2, a3 > > ret.n > > > > ;; after > > test2: > > extui a2, a2, 20, 1 > > addx4 a2, a2, a3 > > ret.n > > test3: > > extui a2, a2, 1, 1 > > subx8 a2, a2, a3 > > ret.n > > > > gcc/ChangeLog: > > > > * config/xtensa/predicates.md (addsub_operator): New. > > * config/xtensa/xtensa.md (*extzvsi-1bit_ashlsi3, > > *extzvsi-1bit_addsubx): New insn_and_split patterns. > > * config/xtensa/xtensa.cc (xtensa_rtx_costs): > > Add a special case about ifcvt 'noce_try_cmove()' to handle > > constant loads that do not fit into signed 12 bits in the > > patterns added above. > > --- > > gcc/config/xtensa/predicates.md | 3 ++ > > gcc/config/xtensa/xtensa.cc | 3 +- > > gcc/config/xtensa/xtensa.md | 75 + > > 3 files changed, 80 insertions(+), 1 deletion(-) > > This change introduces a bunch of test failures on big endian configuration. > I believe that's because the starting bit position for zero_extract is counted > from different ends depending on the endianness. Yes I ran into something similar just recently when I was improving a similar thing in expand. Thanks, Andrew > > -- > Thanks. > -- Max
Re: [PATCH 1/2] xtensa: Optimize '(x & CST1_POW2) != 0 ? CST2_POW2 : 0'
Hi Suwa-san, On Mon, May 22, 2023 at 12:06 AM Takayuki 'January June' Suwa wrote: > > This patch decreses one machine instruction from "single bit extraction > with shifting" operation, and tries to eliminate the conditional > branch if CST2_POW2 doesn't fit into signed 12 bits with the help > of ifcvt optimization. > > /* example #1 */ > int test0(int x) { > return (x & 1048576) != 0 ? 1024 : 0; > } > extern int foo(void); > int test1(void) { > return (foo() & 1048576) != 0 ? 16777216 : 0; > } > > ;; before > test0: > movia9, 0x400 > sraia2, a2, 10 > and a2, a2, a9 > ret.n > test1: > addisp, sp, -16 > s32i.n a0, sp, 12 > call0 foo > extui a2, a2, 20, 1 > sllia2, a2, 20 > beqz.n a2, .L2 > movi.n a2, 1 > sllia2, a2, 24 > .L2: > l32i.n a0, sp, 12 > addisp, sp, 16 > ret.n > > ;; after > test0: > extui a2, a2, 20, 1 > sllia2, a2, 10 > ret.n > test1: > addisp, sp, -16 > s32i.n a0, sp, 12 > call0 foo > l32i.n a0, sp, 12 > extui a2, a2, 20, 1 > sllia2, a2, 24 > addisp, sp, 16 > ret.n > > In addition, if the left shift amount ('exact_log2(CST2_POW2)') is > between 1 through 3 and a either addition or subtraction with another > register follows, emit a ADDX[248] or SUBX[248] machine instruction > instead of separate left shift and add/subtract ones. > > /* example #2 */ > int test2(int x, int y) { > return ((x & 1048576) != 0 ? 4 : 0) + y; > } > int test3(int x, int y) { > return ((x & 2) != 0 ? 8 : 0) - y; > } > > ;; before > test2: > movi.n a9, 4 > sraia2, a2, 18 > and a2, a2, a9 > add.n a2, a2, a3 > ret.n > test3: > movi.n a9, 8 > sllia2, a2, 2 > and a2, a2, a9 > sub a2, a2, a3 > ret.n > > ;; after > test2: > extui a2, a2, 20, 1 > addx4 a2, a2, a3 > ret.n > test3: > extui a2, a2, 1, 1 > subx8 a2, a2, a3 > ret.n > > gcc/ChangeLog: > > * config/xtensa/predicates.md (addsub_operator): New. > * config/xtensa/xtensa.md (*extzvsi-1bit_ashlsi3, > *extzvsi-1bit_addsubx): New insn_and_split patterns. > * config/xtensa/xtensa.cc (xtensa_rtx_costs): > Add a special case about ifcvt 'noce_try_cmove()' to handle > constant loads that do not fit into signed 12 bits in the > patterns added above. > --- > gcc/config/xtensa/predicates.md | 3 ++ > gcc/config/xtensa/xtensa.cc | 3 +- > gcc/config/xtensa/xtensa.md | 75 + > 3 files changed, 80 insertions(+), 1 deletion(-) This change introduces a bunch of test failures on big endian configuration. I believe that's because the starting bit position for zero_extract is counted from different ends depending on the endianness. -- Thanks. -- Max
[PATCH 1/2] xtensa: Optimize '(x & CST1_POW2) != 0 ? CST2_POW2 : 0'
This patch decreses one machine instruction from "single bit extraction with shifting" operation, and tries to eliminate the conditional branch if CST2_POW2 doesn't fit into signed 12 bits with the help of ifcvt optimization. /* example #1 */ int test0(int x) { return (x & 1048576) != 0 ? 1024 : 0; } extern int foo(void); int test1(void) { return (foo() & 1048576) != 0 ? 16777216 : 0; } ;; before test0: movia9, 0x400 sraia2, a2, 10 and a2, a2, a9 ret.n test1: addisp, sp, -16 s32i.n a0, sp, 12 call0 foo extui a2, a2, 20, 1 sllia2, a2, 20 beqz.n a2, .L2 movi.n a2, 1 sllia2, a2, 24 .L2: l32i.n a0, sp, 12 addisp, sp, 16 ret.n ;; after test0: extui a2, a2, 20, 1 sllia2, a2, 10 ret.n test1: addisp, sp, -16 s32i.n a0, sp, 12 call0 foo l32i.n a0, sp, 12 extui a2, a2, 20, 1 sllia2, a2, 24 addisp, sp, 16 ret.n In addition, if the left shift amount ('exact_log2(CST2_POW2)') is between 1 through 3 and a either addition or subtraction with another register follows, emit a ADDX[248] or SUBX[248] machine instruction instead of separate left shift and add/subtract ones. /* example #2 */ int test2(int x, int y) { return ((x & 1048576) != 0 ? 4 : 0) + y; } int test3(int x, int y) { return ((x & 2) != 0 ? 8 : 0) - y; } ;; before test2: movi.n a9, 4 sraia2, a2, 18 and a2, a2, a9 add.n a2, a2, a3 ret.n test3: movi.n a9, 8 sllia2, a2, 2 and a2, a2, a9 sub a2, a2, a3 ret.n ;; after test2: extui a2, a2, 20, 1 addx4 a2, a2, a3 ret.n test3: extui a2, a2, 1, 1 subx8 a2, a2, a3 ret.n gcc/ChangeLog: * config/xtensa/predicates.md (addsub_operator): New. * config/xtensa/xtensa.md (*extzvsi-1bit_ashlsi3, *extzvsi-1bit_addsubx): New insn_and_split patterns. * config/xtensa/xtensa.cc (xtensa_rtx_costs): Add a special case about ifcvt 'noce_try_cmove()' to handle constant loads that do not fit into signed 12 bits in the patterns added above. --- gcc/config/xtensa/predicates.md | 3 ++ gcc/config/xtensa/xtensa.cc | 3 +- gcc/config/xtensa/xtensa.md | 75 + 3 files changed, 80 insertions(+), 1 deletion(-) diff --git a/gcc/config/xtensa/predicates.md b/gcc/config/xtensa/predicates.md index 2dac193373a..5faf1be8c15 100644 --- a/gcc/config/xtensa/predicates.md +++ b/gcc/config/xtensa/predicates.md @@ -191,6 +191,9 @@ (define_predicate "logical_shift_operator" (match_code "ashift,lshiftrt")) +(define_predicate "addsub_operator" + (match_code "plus,minus")) + (define_predicate "xtensa_cstoresi_operator" (match_code "eq,ne,gt,ge,lt,le")) diff --git a/gcc/config/xtensa/xtensa.cc b/gcc/config/xtensa/xtensa.cc index bb1444c44b6..e3af78cd228 100644 --- a/gcc/config/xtensa/xtensa.cc +++ b/gcc/config/xtensa/xtensa.cc @@ -4355,7 +4355,8 @@ xtensa_rtx_costs (rtx x, machine_mode mode, int outer_code, switch (outer_code) { case SET: - if (xtensa_simm12b (INTVAL (x))) + if (xtensa_simm12b (INTVAL (x)) + || (current_pass && current_pass->tv_id == TV_IFCVT)) { *total = speed ? COSTS_N_INSNS (1) : 0; return true; diff --git a/gcc/config/xtensa/xtensa.md b/gcc/config/xtensa/xtensa.md index 3521fa33b47..bd4614e4be0 100644 --- a/gcc/config/xtensa/xtensa.md +++ b/gcc/config/xtensa/xtensa.md @@ -997,6 +997,81 @@ (set_attr "mode""SI") (set_attr "length" "3")]) +(define_insn_and_split "*extzvsi-1bit_ashlsi3" + [(set (match_operand:SI 0 "register_operand" "=a") + (and:SI (match_operator:SI 4 "logical_shift_operator" + [(match_operand:SI 1 "register_operand" "r") +(match_operand:SI 2 "const_int_operand" "i")]) + (match_operand:SI 3 "const_int_operand" "i")))] + "exact_log2 (INTVAL (operands[3])) > 0" + "#" + "&& 1" + [(set (match_dup 0) + (zero_extract:SI (match_dup 1) +(const_int 1) +(match_dup 2))) + (set (match_dup 0) + (ashift:SI (match_dup 0) + (match_dup 3)))] +{ + int shift = floor_log2 (INTVAL (operands[3])); + switch (GET_CODE (operands[4])) +{ +case ASHIFT: + operands[2] = GEN_INT (shift - INTVAL (operands[2])); + break; +case LSHIFTRT: + operands[2] = GEN_INT (shift + INTVAL (operands[2])); + break; +default: + gcc_unreachable (); +} + operands[3] = GEN_INT (shift); +} + [(set_attr "type""arith") +