> -----Original Message-----
> From: Kyrylo Tkachov <[email protected]>
> Sent: 12 August 2026 16:12
> To: Tamar Christina <[email protected]>
> Cc: [email protected]; Wilco Dijkstra <[email protected]>
> Subject: Re: [PATCH 2/2] aarch64: use [SU]DOT for the byte to word step of a
> widening sum
> 
> 
> 
> > On 12 Aug 2026, at 16:45, Tamar Christina <[email protected]>
> wrote:
> >
> >> -----Original Message-----
> >> From: [email protected] <[email protected]>
> >> Sent: 12 August 2026 12:44
> >> To: [email protected]
> >> Cc: Tamar Christina <[email protected]>; Wilco Dijkstra
> >> <[email protected]>; Kyrylo Tkachov <[email protected]>
> >> Subject: [PATCH 2/2] aarch64: use [SU]DOT for the byte to word step of a
> >> widening sum
> >>
> >> From: Kyrylo Tkachov <[email protected]>
> >>
> >> A widening sum from bytes into 64-bit elements spends two [SU]ADDLP
> >> getting from bytes to words.  With dot product that step is a single
> >> [SU]DOT against a vector of ones, which is what the byte to word expander
> >> already does for a 4x reduction.  Each 32-bit element then holds the sum
> >> of four input elements, at most 4 * 255 unsigned and within -512 to 508
> >> signed, so no sum can overflow.
> >>
> >> Move the dot product step into aarch64_expand_reduc_widen_sum, so
> that
> >> any
> >> chain that passes through a byte to word step uses it.  The only shape that
> >> gains is V2DI <- V16QI, because the other shapes either do not start from
> >> bytes or already stop at 32-bit elements:
> >>
> >>  V8HI <- V16QI    [SU]ADALP                 word elements would be too wide
> >>  V4SI <- V8HI     [SU]ADALP                 not a byte source
> >>  V2DI <- V4SI     [SU]ADALP                 not a byte source
> >>  V2SI <- V8QI     [SU]DOT                   unchanged
> >>  V4SI <- V16QI    [SU]DOT                   unchanged
> >>  V2DI <- V8HI     [SU]ADDLP + [SU]ADALP     not a byte source
> >>  V2DI <- V16QI    [SU]DOT + [SU]ADALP       new
> >>
> >> Without dot product every shape keeps the pairwise chain.
> >>
> >> For a sum of unsigned char into long the inner loop changes from
> >>
> >> ldr q31, [x2], 16
> >> uaddlp v31.8h, v31.16b
> >> uaddlp v31.4s, v31.8h
> >> uadalp v30.2d, v31.4s
> >>
> >> to
> >>
> >> ldr q29, [x2], 16
> >> movi v31.4s, 0
> >> udot v31.4s, v29.16b, v27.16b
> >> uadalp v30.2d, v31.4s
> >>
> >> with the vector of ones in v27 hoisted out of the loop.  The instruction
> >> count is unchanged but the vector work is spread better.
> >> A sum of unsigned char into long runs about 24% faster at
> >> -march=armv8.2-a+dotprod, and about 20% faster with an L1 resident
> >> working
> >> set at -mcpu=neoverse-v2, where the vectorizer unrolls the loop by four.
> >>
> >> Bootstrapped and tested on aarch64-none-linux-gnu.
> >> Ok for trunk?
> >
> > This looks good to me, however do we not provide this optab for SVE?
> >
> > https://godbolt.org/z/bxas763ze seems like we don't.  I was expecting
> > it to do a widening load from b -> h and then use udot from h -> d.
> >
> > instead of widening from b -> d, as that's a much lower VF..
> >
> > Would you mind checking?
> >
> 
> I think we’d need either vect_recog_widen_sum_pattern to try an
> intermediate halfword type (it currently looks through the promotion to the
> original byte type and checks only a direct byte-to-doubleword optab)

Hmm I guess it doesn't know about widening types. I would have expected
that given a partial type like VNx8QI that it only checks up to the size of the
inserted conversion.

I can have a look later as well.

> Or we’d need to implement a VNx2DI <- VNx8QI widening sum expander for
> SVE.

That won't work I think until SVE2p3.

> I can look at either as a follow-up, but I’m guessing they are beyond the 
> scope
> of this patch.

Yeah that's fine.

Thanks,
Tamar

> Thanks,
> Kyrill
> 
> 
> > Thanks,
> > Tamar
> >
> >> Thanks,
> >> Kyrill
> >>
> >> gcc/ChangeLog:
> >>
> >> * config/aarch64/aarch64-simd.md
> >> (reduc_widen_<su>sum<mode><vsi2qi>3):
> >> Expand through aarch64_expand_reduc_widen_sum.
> >> * config/aarch64/aarch64.cc (aarch64_expand_reduc_widen_sum):
> >> Use
> >> [SU]DOT for a step from byte to word elements.
> >>
> >> gcc/testsuite/ChangeLog:
> >>
> >> * gcc.target/aarch64/widen_sum_pairwise_2.c: Cover every widening
> >> sum shape and check the dot product sequences.
> >> * gcc.target/aarch64/widen_sum_pairwise_3.c: New test.
> >>
> >> Signed-off-by: Kyrylo Tkachov <[email protected]>
> >> ---
> >> gcc/config/aarch64/aarch64-simd.md            | 16 +---
> >> gcc/config/aarch64/aarch64.cc                 | 41 ++++++++-
> >> .../gcc.target/aarch64/widen_sum_pairwise_2.c | 71 ++++++++++------
> >> .../gcc.target/aarch64/widen_sum_pairwise_3.c | 83
> >> +++++++++++++++++++
> >> 4 files changed, 170 insertions(+), 41 deletions(-)
> >> create mode 100644
> >> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c
> >>
> >> diff --git a/gcc/config/aarch64/aarch64-simd.md
> >> b/gcc/config/aarch64/aarch64-simd.md
> >> index 43be461ef04..6650cbb5f7e 100644
> >> --- a/gcc/config/aarch64/aarch64-simd.md
> >> +++ b/gcc/config/aarch64/aarch64-simd.md
> >> @@ -5320,10 +5320,7 @@
> >>   DONE;
> >> })
> >>
> >> -;; A widening sum reduction that quarters the lane count.  With dot
> product
> >> -;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * 
> >> 1).
> >> -;; Otherwise it is a pairwise widening add feeding a pairwise widening
> >> -;; accumulate.
> >> +;; A widening sum reduction that quarters the lane count.
> >> (define_expand "reduc_widen_<su>sum<mode><vsi2qi>3"
> >>   [(set (match_operand:VS 0 "register_operand")
> >> (plus:VS (ANY_EXTEND:VS
> >> @@ -5331,15 +5328,8 @@
> >>  (match_operand:VS 2 "register_operand")))]
> >>   "TARGET_SIMD"
> >>   {
> >> -    if (TARGET_DOTPROD)
> >> -      {
> >> - rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode));
> >> - emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0],
> >> operands[1],
> >> -    ones, operands[2]));
> >> -      }
> >> -    else
> >> -      aarch64_expand_reduc_widen_sum (operands[0], operands[2],
> >> operands[1],
> >> -       <CODE>);
> >> +    aarch64_expand_reduc_widen_sum (operands[0], operands[2],
> >> operands[1],
> >> +     <CODE>);
> >>     DONE;
> >>   }
> >> )
> >> diff --git a/gcc/config/aarch64/aarch64.cc
> b/gcc/config/aarch64/aarch64.cc
> >> index c1d57ca3964..fdffb13ad22 100644
> >> --- a/gcc/config/aarch64/aarch64.cc
> >> +++ b/gcc/config/aarch64/aarch64.cc
> >> @@ -26331,8 +26331,9 @@ aarch64_expand_vector_init (rtx target, rtx
> vals)
> >>    Advanced SIMD vector SRC holds an even multiple of the number of lanes
> >>    of the accumulator ACC and of the result DEST.  EXTEND_CODE is
> >>    SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form.
> >> -   Halve the lane count with [SU]ADDLP until a single pairwise step is
> >> -   left, then accumulate into ACC with [SU]ADALP.  */
> >> +   Quarter the lane count of a vector of bytes with a [SU]DOT against a
> >> +   vector of ones where that is available, halve it with [SU]ADDLP until a
> >> +   single pairwise step is left, then accumulate into ACC with [SU]ADALP. 
> >>  */
> >>
> >> void
> >> aarch64_expand_reduc_widen_sum (rtx dest, rtx acc, rtx src,
> >> @@ -26340,7 +26341,41 @@ aarch64_expand_reduc_widen_sum (rtx
> dest,
> >> rtx acc, rtx src,
> >> {
> >>   unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE
> >> (dest)).to_constant ();
> >>   machine_mode mode = GET_MODE (src);
> >> -  gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits *
> 2)
> >> == 0);
> >> +  unsigned int nunits = GET_MODE_NUNITS (mode).to_constant ();
> >> +  gcc_assert (nunits % (dest_nunits * 2) == 0);
> >> +
> >> +  /* [SU]DOT against a vector of ones turns += a into += (a * 1), which
> >> +     sums four bytes into each 32-bit element and so covers two halving
> >> +     steps in one operation.  The widest intermediate is 4 * 255, so no
> >> +     product sum can overflow.  Only a step from bytes to words qualifies,
> >> +     and only if the accumulator is at least that wide.  */
> >> +  if (TARGET_DOTPROD
> >> +      && GET_MODE_INNER (mode) == QImode
> >> +      && nunits >= dest_nunits * 4)
> >> +    {
> >> +      machine_mode sum_mode
> >> + = related_vector_mode (mode, SImode, nunits / 4).require ();
> >> +      convert_optab dot = (extend_code == SIGN_EXTEND
> >> +    ? sdot_prod_optab : udot_prod_optab);
> >> +      insn_code icode = convert_optab_handler (dot, sum_mode, mode);
> >> +      rtx ones = force_reg (mode, CONST1_RTX (mode));
> >> +
> >> +      /* A dot product that already reaches the element width of DEST
> >> +  accumulates into ACC itself, otherwise it starts from zero and the
> >> +  remaining steps carry its result into ACC.  */
> >> +      if (sum_mode == GET_MODE (dest))
> >> + {
> >> +   emit_insn (GEN_FCN (icode) (dest, src, ones, acc));
> >> +   return;
> >> + }
> >> +
> >> +      rtx tmp = gen_reg_rtx (sum_mode);
> >> +      emit_insn (GEN_FCN (icode) (tmp, src, ones,
> >> +   force_reg (sum_mode,
> >> +      CONST0_RTX (sum_mode))));
> >> +      src = tmp;
> >> +      mode = sum_mode;
> >> +    }
> >>
> >>   while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2)
> >>     {
> >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> >> index 01537deeb9f..9b3ba07637f 100644
> >> --- a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c
> >> @@ -1,29 +1,50 @@
> >> /* { dg-do compile } */
> >> /* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec-
> >> preference=asimd-only --param vect-epilogues-nomask=0" } */
> >>
> >> -/* With dot product a 4x widening sum stays a single [SU]DOT, while a
> >> -   sum into 64-bit elements uses the pairwise widening instructions.  */
> >> -
> >> -int
> >> -sum_u8_i (const unsigned char *a, long n)
> >> -{
> >> -  int s = 0;
> >> -  for (long i = 0; i < n; i++)
> >> -    s += a[i];
> >> -  return s;
> >> -}
> >> -
> >> -long
> >> -sum_u8_l (const unsigned char *a, long n)
> >> -{
> >> -  long s = 0;
> >> -  for (long i = 0; i < n; i++)
> >> -    s += a[i];
> >> -  return s;
> >> -}
> >> -
> >> -/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b,
> v[0-
> >> 9]+\.16b\n} 1 } } */
> >> -/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-
> 9]+\.16b\n}
> >> 1 } } */
> >> +/* With dot product every widening sum that passes through a byte to
> word
> >> +   step uses one [SU]DOT for that step.  A step that starts or ends
> >> +   somewhere else still uses the pairwise widening instructions.  */
> >> +
> >> +#define DEF(NAME, ITYPE, OTYPE) \
> >> +  OTYPE NAME (const ITYPE *a, long n) \
> >> +  { \
> >> +    OTYPE s = 0; \
> >> +    for (long i = 0; i < n; i++) \
> >> +      s += a[i]; \
> >> +    return s; \
> >> +  }
> >> +
> >> +/* 2x, no dot product: the result elements are too narrow.  */
> >> +DEF (sum_u8_h, unsigned char, unsigned short)
> >> +DEF (sum_i8_h, signed char, short)
> >> +DEF (sum_u16_i, unsigned short, int)
> >> +DEF (sum_i16_i, short, int)
> >> +DEF (sum_u32_l, unsigned int, long)
> >> +DEF (sum_i32_l, int, long)
> >> +
> >> +/* 4x from bytes: one dot product.  */
> >> +DEF (sum_u8_i, unsigned char, int)
> >> +DEF (sum_i8_i, signed char, int)
> >> +
> >> +/* 4x from halfwords: no dot product for that element size.  */
> >> +DEF (sum_u16_l, unsigned short, long)
> >> +DEF (sum_i16_l, short, long)
> >> +
> >> +/* 8x from bytes: a dot product followed by one pairwise accumulate.  */
> >> +DEF (sum_u8_l, unsigned char, long)
> >> +DEF (sum_i8_l, signed char, long)
> >> +
> >> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b,
> v[0-
> >> 9]+\.16b\n} 2 } } */
> >> +/* { dg-final { scan-assembler-times {\tsdot\tv[0-9]+\.4s, v[0-9]+\.16b,
> v[0-
> >> 9]+\.16b\n} 2 } } */
> >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.8h, v[0-
> 9]+\.16b\n}
> >> 1 } } */
> >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.8h, v[0-
> 9]+\.16b\n}
> >> 1 } } */
> >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-
> 9]+\.8h\n} 1
> >> } } */
> >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-
> 9]+\.8h\n} 1
> >> } } */
> >> /* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n}
> 1
> >> } } */
> >> -/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-
> 9]+\.4s\n} 1
> >> } } */
> >> -/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */
> >> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-
> 9]+\.8h\n} 1
> >> } } */
> >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-
> 9]+\.4s\n} 3
> >> } } */
> >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-
> 9]+\.4s\n} 3
> >> } } */
> >> +
> >> +/* The byte to halfword step is what the dot product replaces.  */
> >> +/* { dg-final { scan-assembler-not {\t[su]addlp\tv[0-9]+\.8h, v[0-
> >> 9]+\.16b\n} } } */
> >> +/* { dg-final { scan-assembler-not {\t[su]addw2?\t} } } */
> >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c
> >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c
> >> new file mode 100644
> >> index 00000000000..d2eb8152722
> >> --- /dev/null
> >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c
> >> @@ -0,0 +1,83 @@
> >> +/* { dg-do run } */
> >> +/* { dg-require-effective-target arm_v8_2a_dotprod_neon_hw } */
> >> +/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec-
> >> preference=asimd-only" } */
> >> +
> >> +/* Both expansions of a widening sum reduction, with and without dot
> >> +   product, must agree with a scalar sum for every narrow to wide type
> >> +   pair.  The accumulators are unsigned so that overflow wraps.  */
> >> +
> >> +#define TYPES(X) \
> >> +  X (u8_h,  unsigned char,  unsigned short) \
> >> +  X (i8_h,  signed char,    unsigned short) \
> >> +  X (u8_i,  unsigned char,  unsigned int) \
> >> +  X (i8_i,  signed char,    unsigned int) \
> >> +  X (u8_l,  unsigned char,  unsigned long) \
> >> +  X (i8_l,  signed char,    unsigned long) \
> >> +  X (u16_i, unsigned short, unsigned int) \
> >> +  X (i16_i, short,          unsigned int) \
> >> +  X (u16_l, unsigned short, unsigned long) \
> >> +  X (i16_l, short,          unsigned long) \
> >> +  X (u32_l, unsigned int,   unsigned long) \
> >> +  X (i32_l, int,            unsigned long)
> >> +
> >> +#define SUM(PREFIX, NAME, ITYPE, OTYPE) \
> >> +  __attribute__ ((noipa)) \
> >> +  OTYPE PREFIX##_##NAME (const ITYPE *a, int n) \
> >> +  { \
> >> +    OTYPE s = 0; \
> >> +    for (int i = 0; i < n; i++) \
> >> +      s += a[i]; \
> >> +    return s; \
> >> +  }
> >> +
> >> +#define DOT(NAME, ITYPE, OTYPE) SUM (dot, NAME, ITYPE, OTYPE)
> >> +#define NODOT(NAME, ITYPE, OTYPE) SUM (nodot, NAME, ITYPE, OTYPE)
> >> +
> >> +/* A volatile accumulator keeps this loop scalar.  */
> >> +#define REF(NAME, ITYPE, OTYPE) \
> >> +  __attribute__ ((noipa)) \
> >> +  OTYPE ref_##NAME (const ITYPE *a, int n) \
> >> +  { \
> >> +    volatile OTYPE s = 0; \
> >> +    for (int i = 0; i < n; i++) \
> >> +      s = s + a[i]; \
> >> +    return s; \
> >> +  }
> >> +
> >> +TYPES (DOT)
> >> +TYPES (REF)
> >> +
> >> +#pragma GCC push_options
> >> +#pragma GCC target ("+nodotprod")
> >> +TYPES (NODOT)
> >> +#pragma GCC pop_options
> >> +
> >> +#define BYTES 8192
> >> +static unsigned char buf[BYTES] __attribute__ ((aligned (64)));
> >> +
> >> +#define CHECK(NAME, ITYPE, OTYPE) \
> >> +  { \
> >> +    const ITYPE *p = (const ITYPE *) (buf + off); \
> >> +    OTYPE want = ref_##NAME (p, n); \
> >> +    if (dot_##NAME (p, n) != want || nodot_##NAME (p, n) != want) \
> >> +      __builtin_abort (); \
> >> +  }
> >> +
> >> +int
> >> +main (void)
> >> +{
> >> +  unsigned long x = 1;
> >> +  for (int i = 0; i < BYTES; i++)
> >> +    {
> >> +      x = x * 6364136223846793005UL + 1442695040888963407UL;
> >> +      buf[i] = x >> 40;
> >> +    }
> >> +
> >> +  for (int off = 0; off < 8; off += 4)
> >> +    for (int n = 0; n <= 260; n++)
> >> +      {
> >> + TYPES (CHECK)
> >> +      }
> >> +
> >> +  return 0;
> >> +}
> >> --
> >> 2.50.1 (Apple Git-155)
> 

Reply via email to