Hi Kyrill,

> -----Original Message-----
> From: [email protected] <[email protected]>
> Sent: 04 August 2026 13:39
> To: [email protected]
> Cc: Tamar Christina <[email protected]>; [email protected]; Kyrylo
> Tkachov <[email protected]>
> Subject: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum
> reductions
> 
> From: Kyrylo Tkachov <[email protected]>
> 
> The Advanced SIMD widen_[su]sum optabs only cover a single widening step,
> expanded as a dependent <su>addw + <su>addw2 pair, plus a 4x form that
> requires dot product.  A reduction into an accumulator that is more than
> twice as wide as the data therefore has to extend the input explicitly and
> then issue one widening add per half vector.  Summing bytes into a 64-bit
> accumulator costs fifteen SIMD operations per 16 bytes of input.
> 
> [SU]ADDLP and [SU]ADALP add adjacent lane pairs into the next wider
> element, so a chain of them expresses any power-of-two widening sum
> reduction in one operation per step.  The regrouping is exact because the
> sum of two elements always fits in the doubled element width, and the
> grouping of lanes inside a reduction accumulator is already unconstrained
> for WIDEN_SUM_EXPR, which the existing dot product based 4x expander also
> relies on.
> 
> Expand the 2x forms as a single [SU]ADALP, add the missing V4SI <- V16QI
> and V2SI <- V8QI forms for !TARGET_DOTPROD, and add the V2DI <- V8HI and
> V2DI <- V16QI forms that no expander covered.  All of them are built by
> aarch64_expand_widen_sum, which halves the lane count with [SU]ADDLP
> until
> one pairwise step remains and then accumulates with [SU]ADALP.
> 
> For a sum of unsigned char into long the inner loop changes from
> 
>       ldr     q30, [x1], 16
>       zip1    v28.16b, v30.16b, v29.16b
>       zip2    v30.16b, v30.16b, v29.16b
>       zip1    v26.8h, v28.8h, v29.8h
>       zip2    v28.8h, v28.8h, v29.8h
>       zip1    v27.8h, v30.8h, v29.8h
>       zip2    v30.8h, v30.8h, v29.8h
>       uaddw   v31.2d, v31.2d, v26.2s
>       uaddw2  v31.2d, v31.2d, v26.4s
>       ...  (six more uaddw/uaddw2)
> 
> to
> 
>       ldr     q31, [x1], 16
>       uaddlp  v31.8h, v31.16b
>       uaddlp  v31.4s, v31.8h
>       uadalp  v30.2d, v31.4s
> 

Have you considered considered even for the b to d case to use dotprod
for the b -> s and then uadalp for the final 4 to d? or is the above codegen
the fallback for when !dotprod? wasn't quite clear..

That should still shave off 1 cycle.

Thanks,
Tamar

> and for a sum of int into long the saddw/saddw2 pair becomes one sadalp.
> On a Neoverse V2 core with an L1 resident working set this cuts the time of
> the
> byte loop by about 88% and of the int loop by about 68%.
> 
> Bootstrapped and tested on aarch64-none-linux-gnu.
> Ok for trunk?
> Thanks,
> Kyrill
> 
> gcc/ChangeLog:
> 
>       * config/aarch64/aarch64-protos.h (aarch64_expand_widen_sum):
> Declare.
>       * config/aarch64/aarch64.cc (aarch64_expand_widen_sum): New
> function.
>       * config/aarch64/aarch64-simd.md (aarch64_<su>adalp<mode>):
> Rename
>       to ...
>       (@aarch64_<su>adalp<mode>): ... this.
>       (widen_ssum<Vdblw><mode>3, widen_usum<Vdblw><mode>3):
> Replace by ...
>       (widen_<su>sum<Vdblw><mode>3): ... this.  Expand to [SU]ADALP.
>       (widen_ssum<mode><vsi2qi>3, widen_usum<mode><vsi2qi>3):
> Replace
>       by ...
>       (widen_<su>sum<mode><vsi2qi>3): ... this.  Handle
> !TARGET_DOTPROD.
>       (widen_<su>sumv2di<mode>3): New expander.
>       * config/aarch64/iterators.md (VQ_BH): New mode iterator.
> 
> gcc/testsuite/ChangeLog:
> 
>       * gcc.target/aarch64/pr122069_1.c: Update expected output.
>       * gcc.target/aarch64/pr122069_3.c: Likewise.
>       * gcc.target/aarch64/saddw-1.c: Renamed to...
>       * gcc.target/aarch64/sadalp-1.c: ...this.  Update expected output.
>       * gcc.target/aarch64/saddw-2.c: Renamed to...
>       * gcc.target/aarch64/sadalp-2.c: ...this.  Update expected output.
>       * gcc.target/aarch64/uaddw-1.c: Renamed to...
>       * gcc.target/aarch64/uadalp-1.c: ...this.  Update expected output.
>       * gcc.target/aarch64/uaddw-2.c: Renamed to...
>       * gcc.target/aarch64/uadalp-2.c: ...this.  Update expected output.
>       * gcc.target/aarch64/uaddw-3.c: Renamed to...
>       * gcc.target/aarch64/uadalp-3.c: ...this.  Update expected output.
>       * gcc.target/aarch64/widen_sum_pairwise_1.c: New test.
>       * gcc.target/aarch64/widen_sum_pairwise_2.c: New test.
> 
> Signed-off-by: Kyrylo Tkachov <[email protected]>
> ---
>  gcc/config/aarch64/aarch64-protos.h           |  1 +
>  gcc/config/aarch64/aarch64-simd.md            | 78 ++++++++-----------
>  gcc/config/aarch64/aarch64.cc                 | 27 +++++++
>  gcc/config/aarch64/iterators.md               |  4 +
>  gcc/testsuite/gcc.target/aarch64/pr122069_1.c | 11 +--
>  gcc/testsuite/gcc.target/aarch64/pr122069_3.c |  3 +-
>  .../aarch64/{saddw-1.c => sadalp-1.c}         |  3 +-
>  .../aarch64/{saddw-2.c => sadalp-2.c}         |  3 +-
>  .../aarch64/{uaddw-1.c => uadalp-1.c}         |  3 +-
>  .../aarch64/{uaddw-2.c => uadalp-2.c}         |  3 +-
>  .../aarch64/{uaddw-3.c => uadalp-3.c}         |  3 +-
>  .../gcc.target/aarch64/widen_sum_pairwise_1.c | 39 ++++++++++
>  .../gcc.target/aarch64/widen_sum_pairwise_2.c | 29 +++++++
>  13 files changed, 141 insertions(+), 66 deletions(-)
>  rename gcc/testsuite/gcc.target/aarch64/{saddw-1.c => sadalp-1.c} (74%)
>  rename gcc/testsuite/gcc.target/aarch64/{saddw-2.c => sadalp-2.c} (74%)
>  rename gcc/testsuite/gcc.target/aarch64/{uaddw-1.c => uadalp-1.c} (75%)
>  rename gcc/testsuite/gcc.target/aarch64/{uaddw-2.c => uadalp-2.c} (75%)
>  rename gcc/testsuite/gcc.target/aarch64/{uaddw-3.c => uadalp-3.c} (74%)
>  create mode 100644
> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
>  create mode 100644
> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> 
> diff --git a/gcc/config/aarch64/aarch64-protos.h
> b/gcc/config/aarch64/aarch64-protos.h
> index bcc833cfaa1..9303f12f80c 100644
> --- a/gcc/config/aarch64/aarch64-protos.h
> +++ b/gcc/config/aarch64/aarch64-protos.h
> @@ -1066,6 +1066,7 @@ void aarch64_emit_sve_pred_vec_duplicate
> (machine_mode, rtx, rtx);
>  void aarch64_expand_prologue (void);
>  void aarch64_decompose_vec_struct_index (machine_mode, rtx *, rtx *,
> bool);
>  void aarch64_expand_vector_init (rtx, rtx);
> +void aarch64_expand_widen_sum (rtx, rtx, rtx, rtx_code);
>  void aarch64_sve_expand_vector_init_subvector (rtx, rtx);
>  void aarch64_sve_expand_vector_init (rtx, rtx);
>  void aarch64_init_cumulative_args (CUMULATIVE_ARGS *, const_tree, rtx,
> diff --git a/gcc/config/aarch64/aarch64-simd.md
> b/gcc/config/aarch64/aarch64-simd.md
> index 433f16052bf..d119ac17352 100644
> --- a/gcc/config/aarch64/aarch64-simd.md
> +++ b/gcc/config/aarch64/aarch64-simd.md
> @@ -1182,7 +1182,7 @@
>    }
>  )
> 
> -(define_expand "aarch64_<su>adalp<mode>"
> +(define_expand "@aarch64_<su>adalp<mode>"
>    [(set (match_operand:<VDBLW> 0 "register_operand")
>       (plus:<VDBLW>
>         (plus:<VDBLW>
> @@ -5283,19 +5283,17 @@
> 
>  ;; <su><addsub>w<q>.
> 
> -(define_expand "widen_ssum<Vdblw><mode>3"
> +;; A widening sum reduction that halves the lane count is a single pairwise
> +;; widening accumulate.
> +(define_expand "widen_<su>sum<Vdblw><mode>3"
>    [(set (match_operand:<VDBLW> 0 "register_operand")
> -     (plus:<VDBLW> (sign_extend:<VDBLW>
> -                     (match_operand:VQW 1 "register_operand"))
> +     (plus:<VDBLW> (ANY_EXTEND:<VDBLW>
> +                     (match_operand:VQW 1 "register_operand"))
>                     (match_operand:<VDBLW> 2 "register_operand")))]
>    "TARGET_SIMD"
>    {
> -    rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false);
> -    rtx temp = gen_reg_rtx (GET_MODE (operands[0]));
> -
> -    emit_insn (gen_aarch64_saddw<mode>_internal (temp, operands[2],
> -                                             operands[1], p));
> -    emit_insn (gen_aarch64_saddw2<mode> (operands[0], temp,
> operands[1]));
> +    emit_insn (gen_aarch64_<su>adalp<mode> (operands[0], operands[2],
> +                                         operands[1]));
>      DONE;
>    }
>  )
> @@ -5311,23 +5309,6 @@
>    DONE;
>  })
> 
> -(define_expand "widen_usum<Vdblw><mode>3"
> -  [(set (match_operand:<VDBLW> 0 "register_operand")
> -     (plus:<VDBLW> (zero_extend:<VDBLW>
> -                     (match_operand:VQW 1 "register_operand"))
> -                   (match_operand:<VDBLW> 2 "register_operand")))]
> -  "TARGET_SIMD"
> -  {
> -    rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false);
> -    rtx temp = gen_reg_rtx (GET_MODE (operands[0]));
> -
> -    emit_insn (gen_aarch64_uaddw<mode>_internal (temp, operands[2],
> -                                              operands[1], p));
> -    emit_insn (gen_aarch64_uaddw2<mode> (operands[0], temp,
> operands[1]));
> -    DONE;
> -  }
> -)
> -
>  (define_expand "widen_usum<Vwide><mode>3"
>    [(set (match_operand:<VWIDE> 0 "register_operand")
>       (plus:<VWIDE> (zero_extend:<VWIDE>
> @@ -5339,38 +5320,43 @@
>    DONE;
>  })
> 
> -(define_expand "widen_ssum<mode><vsi2qi>3"
> +;; A widening sum reduction that quarters the lane count.  With dot product
> +;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * 1).
> +;; Otherwise it is a pairwise widening add feeding a pairwise widening
> +;; accumulate.
> +(define_expand "widen_<su>sum<mode><vsi2qi>3"
>    [(set (match_operand:VS 0 "register_operand")
> -     (plus:VS (sign_extend:VS
> +     (plus:VS (ANY_EXTEND:VS
>                  (match_operand:<VSI2QI> 1 "register_operand"))
>                (match_operand:VS 2 "register_operand")))]
> -  "TARGET_DOTPROD"
> +  "TARGET_SIMD"
>    {
> -    rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
> -    emit_insn (gen_sdot_prod<mode><vsi2qi> (operands[0], operands[1],
> ones,
> -                                         operands[2]));
> +    if (TARGET_DOTPROD)
> +      {
> +     rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
> +     emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0],
> operands[1],
> +                                                ones, operands[2]));
> +      }
> +    else
> +      aarch64_expand_widen_sum (operands[0], operands[2], operands[1],
> <CODE>);
>      DONE;
>    }
>  )
> 
> -;; Use dot product to perform double widening sum reductions by
> -;; changing += a into += (a * 1).  i.e. we seed the multiplication with 1.
> -(define_expand "widen_usum<mode><vsi2qi>3"
> -  [(set (match_operand:VS 0 "register_operand")
> -     (plus:VS (zero_extend:VS
> -                     (match_operand:<VSI2QI> 1 "register_operand"))
> -                   (match_operand:VS 2 "register_operand")))]
> -  "TARGET_DOTPROD"
> +;; Widening sum reductions into 64-bit elements.  These need two or three
> +;; pairwise widening steps.
> +(define_expand "widen_<su>sumv2di<mode>3"
> +  [(set (match_operand:V2DI 0 "register_operand")
> +     (plus:V2DI (ANY_EXTEND:V2DI
> +                  (match_operand:VQ_BH 1 "register_operand"))
> +                (match_operand:V2DI 2 "register_operand")))]
> +  "TARGET_SIMD"
>    {
> -    rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
> -    emit_insn (gen_udot_prod<mode><vsi2qi> (operands[0], operands[1],
> ones,
> -                                         operands[2]));
> +    aarch64_expand_widen_sum (operands[0], operands[2], operands[1],
> <CODE>);
>      DONE;
>    }
>  )
> 
> -;; Use dot product to perform double widening sum reductions by
> -;; changing += a into += (a * 1).  i.e. we seed the multiplication with 1.
>  (define_insn "aarch64_<ANY_EXTEND:su>subw<mode>"
>    [(set (match_operand:<VWIDE> 0 "register_operand" "=w")
>       (minus:<VWIDE> (match_operand:<VWIDE> 1 "register_operand"
> "w")
> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc
> index d19ca305d82..628e94e8c40 100644
> --- a/gcc/config/aarch64/aarch64.cc
> +++ b/gcc/config/aarch64/aarch64.cc
> @@ -26327,6 +26327,33 @@ aarch64_expand_vector_init (rtx target, rtx
> vals)
>    emit_insn (seq_total_cost < fallback_seq_cost ? seq : fallback_seq);
>  }
> 
> +/* Expand the widening sum reduction DEST = ACC + (WIDE) SRC, where the
> +   Advanced SIMD vector SRC holds an even multiple of the number of lanes
> +   of the accumulator ACC and of the result DEST.  EXTEND_CODE is
> +   SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form.
> +   Halve the lane count with [SU]ADDLP until a single pairwise step is
> +   left, then accumulate into ACC with [SU]ADALP.  */
> +
> +void
> +aarch64_expand_widen_sum (rtx dest, rtx acc, rtx src, rtx_code extend_code)
> +{
> +  unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE
> (dest)).to_constant ();
> +  machine_mode mode = GET_MODE (src);
> +  gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2)
> == 0);
> +
> +  while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2)
> +    {
> +      insn_code icode = code_for_aarch64_addlp (extend_code, mode);
> +      mode = insn_data[icode].operand[0].mode;
> +      rtx tmp = gen_reg_rtx (mode);
> +      emit_insn (GEN_FCN (icode) (tmp, src));
> +      src = tmp;
> +    }
> +
> +  emit_insn (GEN_FCN (code_for_aarch64_adalp (extend_code, mode))
> (dest, acc,
> +                                                                src));
> +}
> +
>  /* Emit RTL corresponding to:
>     insr TARGET, ELEM.  */
> 
> diff --git a/gcc/config/aarch64/iterators.md
> b/gcc/config/aarch64/iterators.md
> index 0d319751430..6a8c93cce37 100644
> --- a/gcc/config/aarch64/iterators.md
> +++ b/gcc/config/aarch64/iterators.md
> @@ -313,6 +313,10 @@
>  ;; All quad integer widen-able modes.
>  (define_mode_iterator VQW [V16QI V8HI V4SI])
> 
> +;; Quad integer modes that reach 64-bit elements through more than one
> +;; pairwise widening step.
> +(define_mode_iterator VQ_BH [V16QI V8HI])
> +
>  ;; Double vector modes for combines.
>  (define_mode_iterator VDC [V8QI V4HI V4BF V4HF V2SI V2SF DI DF])
> 
> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> index b2f973261ea..d99b5493ade 100644
> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c
> @@ -10,12 +10,8 @@ inline char char_abs(char i) {
>  ** foo_int:
>  **   ...
>  **   sub     v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
> -**   zip1    v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
> -**   zip2    v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b
> -**   uaddw   v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -**   uaddw2  v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> -**   uaddw   v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -**   uaddw2  v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> +**   uaddlp  v[0-9]+.8h, v[0-9]+.16b
> +**   uadalp  v[0-9]+.4s, v[0-9]+.8h
>  **   ...
>  */
>  int foo_int(unsigned char *x, unsigned char * restrict y) {
> @@ -29,8 +25,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) {
>  ** foo2_int:
>  **   ...
>  **   add     v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
> -**   uaddw   v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -**   uaddw2  v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> +**   uadalp  v[0-9]+.4s, v[0-9]+.8h
>  **   ...
>  */
>  int foo2_int(unsigned short *x, unsigned short * restrict y) {
> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> index 0e832c43032..f29fc2b2ed4 100644
> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c
> @@ -24,8 +24,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) {
>  ** foo2_int:
>  **   ...
>  **   add     v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h
> -**   uaddw   v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h
> -**   uaddw2  v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h
> +**   uadalp  v[0-9]+.4s, v[0-9]+.8h
>  **   ...
>  */
>  int foo2_int(unsigned short *x, unsigned short * restrict y) {
> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-1.c
> b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
> similarity index 74%
> rename from gcc/testsuite/gcc.target/aarch64/saddw-1.c
> rename to gcc/testsuite/gcc.target/aarch64/sadalp-1.c
> index f8871209b8a..61f9633f1a0 100644
> --- a/gcc/testsuite/gcc.target/aarch64/saddw-1.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, short * __restrict x)
>    return result;
>  }
> 
> -/* { dg-final { scan-assembler "saddw" } } */
> -/* { dg-final { scan-assembler "saddw2" } } */
> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-2.c
> b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
> similarity index 74%
> rename from gcc/testsuite/gcc.target/aarch64/saddw-2.c
> rename to gcc/testsuite/gcc.target/aarch64/sadalp-2.c
> index b9fc442a2f7..873fda2e1ea 100644
> --- a/gcc/testsuite/gcc.target/aarch64/saddw-2.c
> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, int * __restrict x)
>    return result;
>  }
> 
> -/* { dg-final { scan-assembler "saddw" } } */
> -/* { dg-final { scan-assembler "saddw2" } } */
> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c
> b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
> similarity index 75%
> rename from gcc/testsuite/gcc.target/aarch64/uaddw-1.c
> rename to gcc/testsuite/gcc.target/aarch64/uadalp-1.c
> index 14dff87d7f0..c4034384aae 100644
> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c
> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, unsigned short * __restrict x)
>    return result;
>  }
> 
> -/* { dg-final { scan-assembler "uaddw" } } */
> -/* { dg-final { scan-assembler "uaddw2" } } */
> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c
> b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
> similarity index 75%
> rename from gcc/testsuite/gcc.target/aarch64/uaddw-2.c
> rename to gcc/testsuite/gcc.target/aarch64/uadalp-2.c
> index 79d0d094fc3..395d36c7c00 100644
> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c
> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c
> @@ -14,6 +14,5 @@ t6(int len, void * dummy, unsigned short * __restrict x)
>    return result;
>  }
> 
> -/* { dg-final { scan-assembler "uaddw" } } */
> -/* { dg-final { scan-assembler "uaddw2" } } */
> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */
> 
> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c
> b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
> similarity index 74%
> rename from gcc/testsuite/gcc.target/aarch64/uaddw-3.c
> rename to gcc/testsuite/gcc.target/aarch64/uadalp-3.c
> index 39cbd6b6cc2..5fdb1639ab8 100644
> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c
> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c
> @@ -14,5 +14,4 @@ t6(int len, void * dummy, char * __restrict x)
>    return result;
>  }
> 
> -/* { dg-final { scan-assembler "uaddw" } } */
> -/* { dg-final { scan-assembler "uaddw2" } } */
> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
> new file mode 100644
> index 00000000000..0aec0bf81c8
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c
> @@ -0,0 +1,39 @@
> +/* { dg-do compile } */
> +/* { dg-options "-O3 -march=armv8-a -mautovec-preference=asimd-only --
> param vect-epilogues-nomask=0" } */
> +
> +/* Widening sum reductions should use the pairwise widening add and
> +   accumulate instructions rather than a chain of extensions feeding
> +   [SU]ADDW pairs.  */
> +
> +#define DEF(NAME, ITYPE, OTYPE)                              \
> +  OTYPE NAME (const ITYPE *a, long n)                        \
> +  {                                                  \
> +    OTYPE s = 0;                                     \
> +    for (long i = 0; i < n; i++)                     \
> +      s += a[i];                                     \
> +    return s;                                                \
> +  }
> +
> +DEF (sum_u8_l, unsigned char, long)
> +DEF (sum_i8_l, signed char, long)
> +DEF (sum_u16_l, unsigned short, long)
> +DEF (sum_i16_l, short, long)
> +DEF (sum_u32_l, unsigned int, long)
> +DEF (sum_i32_l, int, long)
> +DEF (sum_u8_i, unsigned char, int)
> +DEF (sum_i8_i, signed char, int)
> +DEF (sum_u16_i, unsigned short, int)
> +DEF (sum_i16_i, short, int)
> +
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n}
> 2 } } */
> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n}
> 2 } } */
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2
> } } */
> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2
> } } */
> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 3
> } } */
> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 3
> } } */
> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2
> } } */
> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 2
> } } */
> +
> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
> +/* { dg-final { scan-assembler-not {\tsaddw2?\t} } } */
> +/* { dg-final { scan-assembler-not {\tzip1\t} } } */
> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> new file mode 100644
> index 00000000000..01537deeb9f
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> @@ -0,0 +1,29 @@
> +/* { dg-do compile } */
> +/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec-
> preference=asimd-only --param vect-epilogues-nomask=0" } */
> +
> +/* With dot product a 4x widening sum stays a single [SU]DOT, while a
> +   sum into 64-bit elements uses the pairwise widening instructions.  */
> +
> +int
> +sum_u8_i (const unsigned char *a, long n)
> +{
> +  int s = 0;
> +  for (long i = 0; i < n; i++)
> +    s += a[i];
> +  return s;
> +}
> +
> +long
> +sum_u8_l (const unsigned char *a, long n)
> +{
> +  long s = 0;
> +  for (long i = 0; i < n; i++)
> +    s += a[i];
> +  return s;
> +}
> +
> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, v[0-
> 9]+\.16b\n} 1 } } */
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n}
> 1 } } */
> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 1
> } } */
> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} 1
> } } */
> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
> --
> 2.50.1 (Apple Git-155)

Reply via email to