Hi Tamar,

> On 4 Aug 2026, at 14:52, Tamar Christina <[email protected]> wrote:
> 
> Hi Kyrill,
> 
>> -----Original Message-----
>> From: [email protected] <[email protected]>
>> Sent: 04 August 2026 13:39
>> To: [email protected]
>> Cc: Tamar Christina <[email protected]>; [email protected]; Kyrylo
>> Tkachov <[email protected]>
>> Subject: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum
>> reductions
>> 
>> From: Kyrylo Tkachov <[email protected]>
>> 
>> The Advanced SIMD widen_[su]sum optabs only cover a single widening step,
>> expanded as a dependent <su>addw + <su>addw2 pair, plus a 4x form that
>> requires dot product.  A reduction into an accumulator that is more than
>> twice as wide as the data therefore has to extend the input explicitly and
>> then issue one widening add per half vector.  Summing bytes into a 64-bit
>> accumulator costs fifteen SIMD operations per 16 bytes of input.
>> 
>> [SU]ADDLP and [SU]ADALP add adjacent lane pairs into the next wider
>> element, so a chain of them expresses any power-of-two widening sum
>> reduction in one operation per step.  The regrouping is exact because the
>> sum of two elements always fits in the doubled element width, and the
>> grouping of lanes inside a reduction accumulator is already unconstrained
>> for WIDEN_SUM_EXPR, which the existing dot product based 4x expander also
>> relies on.
>> 
>> Expand the 2x forms as a single [SU]ADALP, add the missing V4SI <- V16QI
>> and V2SI <- V8QI forms for !TARGET_DOTPROD, and add the V2DI <- V8HI and
>> V2DI <- V16QI forms that no expander covered.  All of them are built by
>> aarch64_expand_widen_sum, which halves the lane count with [SU]ADDLP
>> until
>> one pairwise step remains and then accumulates with [SU]ADALP.
>> 
>> For a sum of unsigned char into long the inner loop changes from
>> 
>> ldr q30, [x1], 16
>> zip1 v28.16b, v30.16b, v29.16b
>> zip2 v30.16b, v30.16b, v29.16b
>> zip1 v26.8h, v28.8h, v29.8h
>> zip2 v28.8h, v28.8h, v29.8h
>> zip1 v27.8h, v30.8h, v29.8h
>> zip2 v30.8h, v30.8h, v29.8h
>> uaddw v31.2d, v31.2d, v26.2s
>> uaddw2 v31.2d, v31.2d, v26.4s
>> ...  (six more uaddw/uaddw2)
>> 
>> to
>> 
>> ldr q31, [x1], 16
>> uaddlp v31.8h, v31.16b
>> uaddlp v31.4s, v31.8h
>> uadalp v30.2d, v31.4s
>> 
> 
> Have you considered considered even for the b to d case to use dotprod
> for the b -> s and then uadalp for the final 4 to d? or is the above codegen
> the fallback for when !dotprod? wasn't quite clear..
> 

It’s not a fallback as the dotprod path exists only for b -> s.
Using dotprod here is interesting, I’ll try it out.
Thanks for the suggestion!
Kyrill

> That should still shave off 1 cycle.
> 
> Thanks,
> Tamar
> 
>> and for a sum of int into long the saddw/saddw2 pair becomes one sadalp.
>> On a Neoverse V2 core with an L1 resident working set this cuts the time of
>> the
>> byte loop by about 88% and of the int loop by about 68%.
>> 
>> Bootstrapped and tested on aarch64-none-linux-gnu.
>> Ok for trunk?
>> Thanks,
>> Kyrill
>> 
>> gcc/ChangeLog:
>> 
>> * config/aarch64/aarch64-protos.h (aarch64_expand_widen_sum):
>> Declare.
>> * config/aarch64/aarch64.cc (aarch64_expand_widen_sum): New
>> function.
>> * config/aarch64/aarch64-simd.md (aarch64_<su>adalp<mode>):
>> Rename
>> to ...
>> (@aarch64_<su>adalp<mode>): ... this.
>> (widen_ssum<Vdblw><mode>3, widen_usum<Vdblw><mode>3):
>> Replace by ...
>> (widen_<su>sum<Vdblw><mode>3): ... this.  Expand to [SU]ADALP.
>> (widen_ssum<mode><vsi2qi>3, widen_usum<mode><vsi2qi>3):
>> Replace
>> by ...
>> (widen_<su>sum<mode><vsi2qi>3): ... this.  Handle
>> !TARGET_DOTPROD.
>> (widen_<su>sumv2di<mode>3): New expander.
>> * config/aarch64/iterators.md (VQ_BH): New mode iterator.
>> 
>> gcc/testsuite/ChangeLog:
>> 
>> * gcc.target/aarch64/pr122069_1.c: Update expected output.
>> * gcc.target/aarch64/pr122069_3.c: Likewise.
>> * gcc.target/aarch64/saddw-1.c: Renamed to...
>> * gcc.target/aarch64/sadalp-1.c: ...this.  Update expected output.
>> * gcc.target/aarch64/saddw-2.c: Renamed to...
>> * gcc.target/aarch64/sadalp-2.c: ...this.  Update expected output.
>> * gcc.target/aarch64/uaddw-1.c: Renamed to...
>> * gcc.target/aarch64/uadalp-1.c: ...this.  Update expected output.
>> * gcc.target/aarch64/uaddw-2.c: Renamed to...
>> * gcc.target/aarch64/uadalp-2.c: ...this.  Update expected output.
>> * gcc.target/aarch64/uaddw-3.c: Renamed to...
>> * gcc.target/aarch64/uadalp-3.c: ...this.  Update expected output.
>> * gcc.target/aarch64/widen_sum_pairwise_1.c: New test.
>> * gcc.target/aarch64/widen_sum_pairwise_2.c: New test.
>> 
>> Signed-off-by: Kyrylo Tkachov <[email protected]>
>> ---
>> gcc/config/aarch64/aarch64-protos.h           |  1 +
>> gcc/config/aarch64/aarch64-simd.md            | 78 ++++++++-----------
>> gcc/config/aarch64/aarch64.cc                 | 27 +++++++
>> gcc/config/aarch64/iterators.md               |  4 +
>> gcc/testsuite/gcc.target/aarch64/pr122069_1.c | 11 +--
>> gcc/testsuite/gcc.target/aarch64/pr122069_3.c |  3 +-
>> .../aarch64/{saddw-1.c => sadalp-1.c}         |  3 +-
>> .../aarch64/{saddw-2.c => sadalp-2.c}         |  3 +-
>> .../aarch64/{uaddw-1.c => uadalp-1.c}         |  3 +-
>> .../aarch64/{uaddw-2.c => uadalp-2.c}         |  3 +-
>> .../aarch64/{uaddw-3.c => uadalp-3.c}         |  3 +-
>> .../gcc.target/aarch64/widen_sum_pairwise_1.c | 39 ++++++++++
>> .../gcc.target/aarch64/widen_sum_pairwise_2.c | 29 +++++++
>> 13 files changed, 141 insertions(+), 66 deletions(-)
>> rename gcc/testsuite/gcc.target/aarch64/{saddw-1.c => sadalp-1.c} (74%)
>> rename gcc/testsuite/gcc.target/aarch64/{saddw-2.c => sadalp-2.c} (74%)
>> rename gcc/testsuite/gcc.target/aarch64/{uaddw-1.c => uadalp-1.c} (75%)
>> rename gcc/testsuite/gcc.target/aarch64/{uaddw-2.c => uadalp-2.c} (75%)
>> rename gcc/testsuite/gcc.target/aarch64/{uaddw-3.c => uadalp-3.c} (74%)
>> create mode 100644
>> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
>> create mode 100644
>> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
>> 
>> diff --git a/gcc/config/aarch64/aarch64-protos.h
>> b/gcc/config/aarch64/aarch64-protos.h
>> index bcc833cfaa1..9303f12f80c 100644
>> --- a/gcc/config/aarch64/aarch64-protos.h
>> +++ b/gcc/config/aarch64/aarch64-protos.h
>> @@ -1066,6 +1066,7 @@ void aarch64_emit_sve_pred_vec_duplicate
>> (machine_mode, rtx, rtx);
>> void aarch64_expand_prologue (void);
>> void aarch64_decompose_vec_struct_index (machine_mode, rtx *, rtx *,
>> bool);
>> void aarch64_expand_vector_init (rtx, rtx);
>> +void aarch64_expand_widen_sum (rtx, rtx, rtx, rtx_code);
>> void aarch64_sve_expand_vector_init_subvector (rtx, rtx);
>> void aarch64_sve_expand_vector_init (rtx, rtx);
>> void aarch64_init_cumulative_args (CUMULATIVE_ARGS *, const_tree, rtx,
>> diff --git a/gcc/config/aarch64/aarch64-simd.md
>> b/gcc/config/aarch64/aarch64-simd.md
>> index 433f16052bf..d119ac17352 100644
>> --- a/gcc/config/aarch64/aarch64-simd.md
>> +++ b/gcc/config/aarch64/aarch64-simd.md
>> @@ -1182,7 +1182,7 @@
>>   }
>> )
>> 
>> -(define_expand "aarch64_<su>adalp<mode>"
>> +(define_expand "@aarch64_<su>adalp<mode>"
>>   [(set (match_operand:<VDBLW> 0 "register_operand")
>> (plus:<VDBLW>
>>   (plus:<VDBLW>
>> @@ -5283,19 +5283,17 @@
>> 
>> ;; <su><addsub>w<q>.
>> 
>> -(define_expand "widen_ssum<Vdblw><mode>3"
>> +;; A widening sum reduction that halves the lane count is a single pairwise
>> +;; widening accumulate.
>> +(define_expand "widen_<su>sum<Vdblw><mode>3"
>>   [(set (match_operand:<VDBLW> 0 "register_operand")
>> - (plus:<VDBLW> (sign_extend:<VDBLW>
>> -         (match_operand:VQW 1 "register_operand"))
>> + (plus:<VDBLW> (ANY_EXTEND:<VDBLW>
>> + (match_operand:VQW 1 "register_operand"))
>>       (match_operand:<VDBLW> 2 "register_operand")))]
>>   "TARGET_SIMD"
>>   {
>> -    rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false);
>> -    rtx temp = gen_reg_rtx (GET_MODE (operands[0]));
>> -
>> -    emit_insn (gen_aarch64_saddw<mode>_internal (temp, operands[2],
>> - operands[1], p));
>> -    emit_insn (gen_aarch64_saddw2<mode> (operands[0], temp,
>> operands[1]));
>> +    emit_insn (gen_aarch64_<su>adalp<mode> (operands[0], operands[2],
>> +     operands[1]));
>>     DONE;
>>   }
>> )
>> @@ -5311,23 +5309,6 @@
>>   DONE;
>> })
>> 
>> -(define_expand "widen_usum<Vdblw><mode>3"
>> -  [(set (match_operand:<VDBLW> 0 "register_operand")
>> - (plus:<VDBLW> (zero_extend:<VDBLW>
>> -         (match_operand:VQW 1 "register_operand"))
>> -       (match_operand:<VDBLW> 2 "register_operand")))]
>> -  "TARGET_SIMD"
>> -  {
>> -    rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false);
>> -    rtx temp = gen_reg_rtx (GET_MODE (operands[0]));
>> -
>> -    emit_insn (gen_aarch64_uaddw<mode>_internal (temp, operands[2],
>> -  operands[1], p));
>> -    emit_insn (gen_aarch64_uaddw2<mode> (operands[0], temp,
>> operands[1]));
>> -    DONE;
>> -  }
>> -)
>> -
>> (define_expand "widen_usum<Vwide><mode>3"
>>   [(set (match_operand:<VWIDE> 0 "register_operand")
>> (plus:<VWIDE> (zero_extend:<VWIDE>
>> @@ -5339,38 +5320,43 @@
>>   DONE;
>> })
>> 
>> -(define_expand "widen_ssum<mode><vsi2qi>3"
>> +;; A widening sum reduction that quarters the lane count.  With dot product
>> +;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * 1).
>> +;; Otherwise it is a pairwise widening add feeding a pairwise widening
>> +;; accumulate.
>> +(define_expand "widen_<su>sum<mode><vsi2qi>3"
>>   [(set (match_operand:VS 0 "register_operand")
>> - (plus:VS (sign_extend:VS
>> + (plus:VS (ANY_EXTEND:VS
>>    (match_operand:<VSI2QI> 1 "register_operand"))
>>  (match_operand:VS 2 "register_operand")))]
>> -  "TARGET_DOTPROD"
>> +  "TARGET_SIMD"
>>   {
>> -    rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
>> -    emit_insn (gen_sdot_prod<mode><vsi2qi> (operands[0], operands[1],
>> ones,
>> -     operands[2]));
>> +    if (TARGET_DOTPROD)
>> +      {
>> + rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
>> + emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0],
>> operands[1],
>> +    ones, operands[2]));
>> +      }
>> +    else
>> +      aarch64_expand_widen_sum (operands[0], operands[2], operands[1],
>> <CODE>);
>>     DONE;
>>   }
>> )
>> 
>> -;; Use dot product to perform double widening sum reductions by
>> -;; changing += a into += (a * 1).  i.e. we seed the multiplication with 1.
>> -(define_expand "widen_usum<mode><vsi2qi>3"
>> -  [(set (match_operand:VS 0 "register_operand")
>> - (plus:VS (zero_extend:VS
>> -         (match_operand:<VSI2QI> 1 "register_operand"))
>> -       (match_operand:VS 2 "register_operand")))]
>> -  "TARGET_DOTPROD"
>> +;; Widening sum reductions into 64-bit elements.  These need two or three
>> +;; pairwise widening steps.
>> +(define_expand "widen_<su>sumv2di<mode>3"
>> +  [(set (match_operand:V2DI 0 "register_operand")
>> + (plus:V2DI (ANY_EXTEND:V2DI
>> +      (match_operand:VQ_BH 1 "register_operand"))
>> +    (match_operand:V2DI 2 "register_operand")))]
>> +  "TARGET_SIMD"
>>   {
>> -    rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
>> -    emit_insn (gen_udot_prod<mode><vsi2qi> (operands[0], operands[1],
>> ones,
>> -     operands[2]));
>> +    aarch64_expand_widen_sum (operands[0], operands[2], operands[1],
>> <CODE>);
>>     DONE;
>>   }
>> )
>> 
>> -;; Use dot product to perform double widening sum reductions by
>> -;; changing += a into += (a * 1).  i.e. we seed the multiplication with 1.
>> (define_insn "aarch64_<ANY_EXTEND:su>subw<mode>"
>>   [(set (match_operand:<VWIDE> 0 "register_operand" "=w")
>> (minus:<VWIDE> (match_operand:<VWIDE> 1 "register_operand"
>> "w")
>> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc
>> index d19ca305d82..628e94e8c40 100644
>> --- a/gcc/config/aarch64/aarch64.cc
>> +++ b/gcc/config/aarch64/aarch64.cc
>> @@ -26327,6 +26327,33 @@ aarch64_expand_vector_init (rtx target, rtx
>> vals)
>>   emit_insn (seq_total_cost < fallback_seq_cost ? seq : fallback_seq);
>> }
>> 
>> +/* Expand the widening sum reduction DEST = ACC + (WIDE) SRC, where the
>> +   Advanced SIMD vector SRC holds an even multiple of the number of lanes
>> +   of the accumulator ACC and of the result DEST.  EXTEND_CODE is
>> +   SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form.
>> +   Halve the lane count with [SU]ADDLP until a single pairwise step is
>> +   left, then accumulate into ACC with [SU]ADALP.  */
>> +
>> +void
>> +aarch64_expand_widen_sum (rtx dest, rtx acc, rtx src, rtx_code extend_code)
>> +{
>> +  unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE
>> (dest)).to_constant ();
>> +  machine_mode mode = GET_MODE (src);
>> +  gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2)
>> == 0);
>> +
>> +  while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2)
>> +    {
>> +      insn_code icode = code_for_aarch64_addlp (extend_code, mode);
>> +      mode = insn_data[icode].operand[0].mode;
>> +      rtx tmp = gen_reg_rtx (mode);
>> +      emit_insn (GEN_FCN (icode) (tmp, src));
>> +      src = tmp;
>> +    }
>> +
>> +  emit_insn (GEN_FCN (code_for_aarch64_adalp (extend_code, mode))
>> (dest, acc,
>> +    src));
>> +}
>> +
>> /* Emit RTL corresponding to:
>>    insr TARGET, ELEM.  */
>> 
>> diff --git a/gcc/config/aarch64/iterators.md
>> b/gcc/config/aarch64/iterators.md
>> index 0d319751430..6a8c93cce37 100644
>> --- a/gcc/config/aarch64/iterators.md
>> +++ b/gcc/config/aarch64/iterators.md
>> @@ -313,6 +313,10 @@
>> ;; All quad integer widen-able modes.
>> (define_mode_iterator VQW [V16QI V8HI V4SI])
>> 
>> +;; Quad integer modes that reach 64-bit elements through more than one
>> +;; pairwise widening step.
>> +(define_mode_iterator VQ_BH [V16QI V8HI])
>> +
>> ;; Double vector modes for combines.
>> (define_mode_iterator VDC [V8QI V4HI V4BF V4HF V2SI V2SF DI DF])
>> 
>> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
>> b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
>> index b2f973261ea..d99b5493ade 100644
>> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
>> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
>> @@ -10,12 +10,8 @@ inline char char_abs(char i) {
>> ** foo_int:
>> **  ...
>> **  sub v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
>> -**  zip1 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
>> -**  zip2 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
>> -**  uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
>> -**  uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
>> -**  uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
>> -**  uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
>> +**  uaddlp v[0-9]+.8h, v[0-9]+.16b
>> +**  uadalp v[0-9]+.4s, v[0-9]+.8h
>> **  ...
>> */
>> int foo_int(unsigned char *x, unsigned char * restrict y) {
>> @@ -29,8 +25,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) {
>> ** foo2_int:
>> **  ...
>> **  add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
>> -**  uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
>> -**  uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
>> +**  uadalp v[0-9]+.4s, v[0-9]+.8h
>> **  ...
>> */
>> int foo2_int(unsigned short *x, unsigned short * restrict y) {
>> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
>> b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
>> index 0e832c43032..f29fc2b2ed4 100644
>> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
>> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
>> @@ -24,8 +24,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) {
>> ** foo2_int:
>> **  ...
>> **  add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
>> -**  uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
>> -**  uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
>> +**  uadalp v[0-9]+.4s, v[0-9]+.8h
>> **  ...
>> */
>> int foo2_int(unsigned short *x, unsigned short * restrict y) {
>> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-1.c
>> b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
>> similarity index 74%
>> rename from gcc/testsuite/gcc.target/aarch64/saddw-1.c
>> rename to gcc/testsuite/gcc.target/aarch64/sadalp-1.c
>> index f8871209b8a..61f9633f1a0 100644
>> --- a/gcc/testsuite/gcc.target/aarch64/saddw-1.c
>> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, short * __restrict x)
>>   return result;
>> }
>> 
>> -/* { dg-final { scan-assembler "saddw" } } */
>> -/* { dg-final { scan-assembler "saddw2" } } */
>> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
>> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-2.c
>> b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
>> similarity index 74%
>> rename from gcc/testsuite/gcc.target/aarch64/saddw-2.c
>> rename to gcc/testsuite/gcc.target/aarch64/sadalp-2.c
>> index b9fc442a2f7..873fda2e1ea 100644
>> --- a/gcc/testsuite/gcc.target/aarch64/saddw-2.c
>> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, int * __restrict x)
>>   return result;
>> }
>> 
>> -/* { dg-final { scan-assembler "saddw" } } */
>> -/* { dg-final { scan-assembler "saddw2" } } */
>> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s} } } */
>> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c
>> b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
>> similarity index 75%
>> rename from gcc/testsuite/gcc.target/aarch64/uaddw-1.c
>> rename to gcc/testsuite/gcc.target/aarch64/uadalp-1.c
>> index 14dff87d7f0..c4034384aae 100644
>> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c
>> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, unsigned short * __restrict x)
>>   return result;
>> }
>> 
>> -/* { dg-final { scan-assembler "uaddw" } } */
>> -/* { dg-final { scan-assembler "uaddw2" } } */
>> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
>> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c
>> b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
>> similarity index 75%
>> rename from gcc/testsuite/gcc.target/aarch64/uaddw-2.c
>> rename to gcc/testsuite/gcc.target/aarch64/uadalp-2.c
>> index 79d0d094fc3..395d36c7c00 100644
>> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c
>> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
>> @@ -14,6 +14,5 @@ t6(int len, void * dummy, unsigned short * __restrict x)
>>   return result;
>> }
>> 
>> -/* { dg-final { scan-assembler "uaddw" } } */
>> -/* { dg-final { scan-assembler "uaddw2" } } */
>> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
>> 
>> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c
>> b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
>> similarity index 74%
>> rename from gcc/testsuite/gcc.target/aarch64/uaddw-3.c
>> rename to gcc/testsuite/gcc.target/aarch64/uadalp-3.c
>> index 39cbd6b6cc2..5fdb1639ab8 100644
>> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c
>> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, char * __restrict x)
>>   return result;
>> }
>> 
>> -/* { dg-final { scan-assembler "uaddw" } } */
>> -/* { dg-final { scan-assembler "uaddw2" } } */
>> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b} } } */
>> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
>> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
>> new file mode 100644
>> index 00000000000..0aec0bf81c8
>> --- /dev/null
>> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
>> @@ -0,0 +1,39 @@
>> +/* { dg-do compile } */
>> +/* { dg-options "-O3 -march=armv8-a -mautovec-preference=asimd-only --
>> param vect-epilogues-nomask=0" } */
>> +
>> +/* Widening sum reductions should use the pairwise widening add and
>> +   accumulate instructions rather than a chain of extensions feeding
>> +   [SU]ADDW pairs.  */
>> +
>> +#define DEF(NAME, ITYPE, OTYPE) \
>> +  OTYPE NAME (const ITYPE *a, long n) \
>> +  { \
>> +    OTYPE s = 0; \
>> +    for (long i = 0; i < n; i++) \
>> +      s += a[i]; \
>> +    return s; \
>> +  }
>> +
>> +DEF (sum_u8_l, unsigned char, long)
>> +DEF (sum_i8_l, signed char, long)
>> +DEF (sum_u16_l, unsigned short, long)
>> +DEF (sum_i16_l, short, long)
>> +DEF (sum_u32_l, unsigned int, long)
>> +DEF (sum_i32_l, int, long)
>> +DEF (sum_u8_i, unsigned char, int)
>> +DEF (sum_i8_i, signed char, int)
>> +DEF (sum_u16_i, unsigned short, int)
>> +DEF (sum_i16_i, short, int)
>> +
>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n}
>> 2 } } */
>> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n}
>> 2 } } */
>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 
>> 2
>> } } */
>> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 
>> 2
>> } } */
>> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 
>> 3
>> } } */
>> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 
>> 3
>> } } */
>> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 
>> 2
>> } } */
>> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 
>> 2
>> } } */
>> +
>> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
>> +/* { dg-final { scan-assembler-not {\tsaddw2?\t} } } */
>> +/* { dg-final { scan-assembler-not {\tzip1\t} } } */
>> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
>> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
>> new file mode 100644
>> index 00000000000..01537deeb9f
>> --- /dev/null
>> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
>> @@ -0,0 +1,29 @@
>> +/* { dg-do compile } */
>> +/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec-
>> preference=asimd-only --param vect-epilogues-nomask=0" } */
>> +
>> +/* With dot product a 4x widening sum stays a single [SU]DOT, while a
>> +   sum into 64-bit elements uses the pairwise widening instructions.  */
>> +
>> +int
>> +sum_u8_i (const unsigned char *a, long n)
>> +{
>> +  int s = 0;
>> +  for (long i = 0; i < n; i++)
>> +    s += a[i];
>> +  return s;
>> +}
>> +
>> +long
>> +sum_u8_l (const unsigned char *a, long n)
>> +{
>> +  long s = 0;
>> +  for (long i = 0; i < n; i++)
>> +    s += a[i];
>> +  return s;
>> +}
>> +
>> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, 
>> v[0-
>> 9]+\.16b\n} 1 } } */
>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n}
>> 1 } } */
>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 
>> 1
>> } } */
>> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 
>> 1
>> } } */
>> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
>> --
>> 2.50.1 (Apple Git-155)


Reply via email to