> -----Original Message----- > From: Kyrylo Tkachov <[email protected]> > Sent: 12 August 2026 16:12 > To: Tamar Christina <[email protected]> > Cc: [email protected]; Wilco Dijkstra <[email protected]> > Subject: Re: [PATCH 2/2] aarch64: use [SU]DOT for the byte to word step of a > widening sum > > > > > On 12 Aug 2026, at 16:45, Tamar Christina <[email protected]> > wrote: > > > >> -----Original Message----- > >> From: [email protected] <[email protected]> > >> Sent: 12 August 2026 12:44 > >> To: [email protected] > >> Cc: Tamar Christina <[email protected]>; Wilco Dijkstra > >> <[email protected]>; Kyrylo Tkachov <[email protected]> > >> Subject: [PATCH 2/2] aarch64: use [SU]DOT for the byte to word step of a > >> widening sum > >> > >> From: Kyrylo Tkachov <[email protected]> > >> > >> A widening sum from bytes into 64-bit elements spends two [SU]ADDLP > >> getting from bytes to words. With dot product that step is a single > >> [SU]DOT against a vector of ones, which is what the byte to word expander > >> already does for a 4x reduction. Each 32-bit element then holds the sum > >> of four input elements, at most 4 * 255 unsigned and within -512 to 508 > >> signed, so no sum can overflow. > >> > >> Move the dot product step into aarch64_expand_reduc_widen_sum, so > that > >> any > >> chain that passes through a byte to word step uses it. The only shape that > >> gains is V2DI <- V16QI, because the other shapes either do not start from > >> bytes or already stop at 32-bit elements: > >> > >> V8HI <- V16QI [SU]ADALP word elements would be too wide > >> V4SI <- V8HI [SU]ADALP not a byte source > >> V2DI <- V4SI [SU]ADALP not a byte source > >> V2SI <- V8QI [SU]DOT unchanged > >> V4SI <- V16QI [SU]DOT unchanged > >> V2DI <- V8HI [SU]ADDLP + [SU]ADALP not a byte source > >> V2DI <- V16QI [SU]DOT + [SU]ADALP new > >> > >> Without dot product every shape keeps the pairwise chain. > >> > >> For a sum of unsigned char into long the inner loop changes from > >> > >> ldr q31, [x2], 16 > >> uaddlp v31.8h, v31.16b > >> uaddlp v31.4s, v31.8h > >> uadalp v30.2d, v31.4s > >> > >> to > >> > >> ldr q29, [x2], 16 > >> movi v31.4s, 0 > >> udot v31.4s, v29.16b, v27.16b > >> uadalp v30.2d, v31.4s > >> > >> with the vector of ones in v27 hoisted out of the loop. The instruction > >> count is unchanged but the vector work is spread better. > >> A sum of unsigned char into long runs about 24% faster at > >> -march=armv8.2-a+dotprod, and about 20% faster with an L1 resident > >> working > >> set at -mcpu=neoverse-v2, where the vectorizer unrolls the loop by four. > >> > >> Bootstrapped and tested on aarch64-none-linux-gnu. > >> Ok for trunk? > > > > This looks good to me, however do we not provide this optab for SVE? > > > > https://godbolt.org/z/bxas763ze seems like we don't. I was expecting > > it to do a widening load from b -> h and then use udot from h -> d. > > > > instead of widening from b -> d, as that's a much lower VF.. > > > > Would you mind checking? > > > > I think we’d need either vect_recog_widen_sum_pattern to try an > intermediate halfword type (it currently looks through the promotion to the > original byte type and checks only a direct byte-to-doubleword optab)
Hmm I guess it doesn't know about widening types. I would have expected that given a partial type like VNx8QI that it only checks up to the size of the inserted conversion. I can have a look later as well. > Or we’d need to implement a VNx2DI <- VNx8QI widening sum expander for > SVE. That won't work I think until SVE2p3. > I can look at either as a follow-up, but I’m guessing they are beyond the > scope > of this patch. Yeah that's fine. Thanks, Tamar > Thanks, > Kyrill > > > > Thanks, > > Tamar > > > >> Thanks, > >> Kyrill > >> > >> gcc/ChangeLog: > >> > >> * config/aarch64/aarch64-simd.md > >> (reduc_widen_<su>sum<mode><vsi2qi>3): > >> Expand through aarch64_expand_reduc_widen_sum. > >> * config/aarch64/aarch64.cc (aarch64_expand_reduc_widen_sum): > >> Use > >> [SU]DOT for a step from byte to word elements. > >> > >> gcc/testsuite/ChangeLog: > >> > >> * gcc.target/aarch64/widen_sum_pairwise_2.c: Cover every widening > >> sum shape and check the dot product sequences. > >> * gcc.target/aarch64/widen_sum_pairwise_3.c: New test. > >> > >> Signed-off-by: Kyrylo Tkachov <[email protected]> > >> --- > >> gcc/config/aarch64/aarch64-simd.md | 16 +--- > >> gcc/config/aarch64/aarch64.cc | 41 ++++++++- > >> .../gcc.target/aarch64/widen_sum_pairwise_2.c | 71 ++++++++++------ > >> .../gcc.target/aarch64/widen_sum_pairwise_3.c | 83 > >> +++++++++++++++++++ > >> 4 files changed, 170 insertions(+), 41 deletions(-) > >> create mode 100644 > >> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c > >> > >> diff --git a/gcc/config/aarch64/aarch64-simd.md > >> b/gcc/config/aarch64/aarch64-simd.md > >> index 43be461ef04..6650cbb5f7e 100644 > >> --- a/gcc/config/aarch64/aarch64-simd.md > >> +++ b/gcc/config/aarch64/aarch64-simd.md > >> @@ -5320,10 +5320,7 @@ > >> DONE; > >> }) > >> > >> -;; A widening sum reduction that quarters the lane count. With dot > product > >> -;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * > >> 1). > >> -;; Otherwise it is a pairwise widening add feeding a pairwise widening > >> -;; accumulate. > >> +;; A widening sum reduction that quarters the lane count. > >> (define_expand "reduc_widen_<su>sum<mode><vsi2qi>3" > >> [(set (match_operand:VS 0 "register_operand") > >> (plus:VS (ANY_EXTEND:VS > >> @@ -5331,15 +5328,8 @@ > >> (match_operand:VS 2 "register_operand")))] > >> "TARGET_SIMD" > >> { > >> - if (TARGET_DOTPROD) > >> - { > >> - rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); > >> - emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0], > >> operands[1], > >> - ones, operands[2])); > >> - } > >> - else > >> - aarch64_expand_reduc_widen_sum (operands[0], operands[2], > >> operands[1], > >> - <CODE>); > >> + aarch64_expand_reduc_widen_sum (operands[0], operands[2], > >> operands[1], > >> + <CODE>); > >> DONE; > >> } > >> ) > >> diff --git a/gcc/config/aarch64/aarch64.cc > b/gcc/config/aarch64/aarch64.cc > >> index c1d57ca3964..fdffb13ad22 100644 > >> --- a/gcc/config/aarch64/aarch64.cc > >> +++ b/gcc/config/aarch64/aarch64.cc > >> @@ -26331,8 +26331,9 @@ aarch64_expand_vector_init (rtx target, rtx > vals) > >> Advanced SIMD vector SRC holds an even multiple of the number of lanes > >> of the accumulator ACC and of the result DEST. EXTEND_CODE is > >> SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form. > >> - Halve the lane count with [SU]ADDLP until a single pairwise step is > >> - left, then accumulate into ACC with [SU]ADALP. */ > >> + Quarter the lane count of a vector of bytes with a [SU]DOT against a > >> + vector of ones where that is available, halve it with [SU]ADDLP until a > >> + single pairwise step is left, then accumulate into ACC with [SU]ADALP. > >> */ > >> > >> void > >> aarch64_expand_reduc_widen_sum (rtx dest, rtx acc, rtx src, > >> @@ -26340,7 +26341,41 @@ aarch64_expand_reduc_widen_sum (rtx > dest, > >> rtx acc, rtx src, > >> { > >> unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE > >> (dest)).to_constant (); > >> machine_mode mode = GET_MODE (src); > >> - gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * > 2) > >> == 0); > >> + unsigned int nunits = GET_MODE_NUNITS (mode).to_constant (); > >> + gcc_assert (nunits % (dest_nunits * 2) == 0); > >> + > >> + /* [SU]DOT against a vector of ones turns += a into += (a * 1), which > >> + sums four bytes into each 32-bit element and so covers two halving > >> + steps in one operation. The widest intermediate is 4 * 255, so no > >> + product sum can overflow. Only a step from bytes to words qualifies, > >> + and only if the accumulator is at least that wide. */ > >> + if (TARGET_DOTPROD > >> + && GET_MODE_INNER (mode) == QImode > >> + && nunits >= dest_nunits * 4) > >> + { > >> + machine_mode sum_mode > >> + = related_vector_mode (mode, SImode, nunits / 4).require (); > >> + convert_optab dot = (extend_code == SIGN_EXTEND > >> + ? sdot_prod_optab : udot_prod_optab); > >> + insn_code icode = convert_optab_handler (dot, sum_mode, mode); > >> + rtx ones = force_reg (mode, CONST1_RTX (mode)); > >> + > >> + /* A dot product that already reaches the element width of DEST > >> + accumulates into ACC itself, otherwise it starts from zero and the > >> + remaining steps carry its result into ACC. */ > >> + if (sum_mode == GET_MODE (dest)) > >> + { > >> + emit_insn (GEN_FCN (icode) (dest, src, ones, acc)); > >> + return; > >> + } > >> + > >> + rtx tmp = gen_reg_rtx (sum_mode); > >> + emit_insn (GEN_FCN (icode) (tmp, src, ones, > >> + force_reg (sum_mode, > >> + CONST0_RTX (sum_mode)))); > >> + src = tmp; > >> + mode = sum_mode; > >> + } > >> > >> while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2) > >> { > >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c > >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c > >> index 01537deeb9f..9b3ba07637f 100644 > >> --- a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c > >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c > >> @@ -1,29 +1,50 @@ > >> /* { dg-do compile } */ > >> /* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec- > >> preference=asimd-only --param vect-epilogues-nomask=0" } */ > >> > >> -/* With dot product a 4x widening sum stays a single [SU]DOT, while a > >> - sum into 64-bit elements uses the pairwise widening instructions. */ > >> - > >> -int > >> -sum_u8_i (const unsigned char *a, long n) > >> -{ > >> - int s = 0; > >> - for (long i = 0; i < n; i++) > >> - s += a[i]; > >> - return s; > >> -} > >> - > >> -long > >> -sum_u8_l (const unsigned char *a, long n) > >> -{ > >> - long s = 0; > >> - for (long i = 0; i < n; i++) > >> - s += a[i]; > >> - return s; > >> -} > >> - > >> -/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, > v[0- > >> 9]+\.16b\n} 1 } } */ > >> -/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0- > 9]+\.16b\n} > >> 1 } } */ > >> +/* With dot product every widening sum that passes through a byte to > word > >> + step uses one [SU]DOT for that step. A step that starts or ends > >> + somewhere else still uses the pairwise widening instructions. */ > >> + > >> +#define DEF(NAME, ITYPE, OTYPE) \ > >> + OTYPE NAME (const ITYPE *a, long n) \ > >> + { \ > >> + OTYPE s = 0; \ > >> + for (long i = 0; i < n; i++) \ > >> + s += a[i]; \ > >> + return s; \ > >> + } > >> + > >> +/* 2x, no dot product: the result elements are too narrow. */ > >> +DEF (sum_u8_h, unsigned char, unsigned short) > >> +DEF (sum_i8_h, signed char, short) > >> +DEF (sum_u16_i, unsigned short, int) > >> +DEF (sum_i16_i, short, int) > >> +DEF (sum_u32_l, unsigned int, long) > >> +DEF (sum_i32_l, int, long) > >> + > >> +/* 4x from bytes: one dot product. */ > >> +DEF (sum_u8_i, unsigned char, int) > >> +DEF (sum_i8_i, signed char, int) > >> + > >> +/* 4x from halfwords: no dot product for that element size. */ > >> +DEF (sum_u16_l, unsigned short, long) > >> +DEF (sum_i16_l, short, long) > >> + > >> +/* 8x from bytes: a dot product followed by one pairwise accumulate. */ > >> +DEF (sum_u8_l, unsigned char, long) > >> +DEF (sum_i8_l, signed char, long) > >> + > >> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, > v[0- > >> 9]+\.16b\n} 2 } } */ > >> +/* { dg-final { scan-assembler-times {\tsdot\tv[0-9]+\.4s, v[0-9]+\.16b, > v[0- > >> 9]+\.16b\n} 2 } } */ > >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.8h, v[0- > 9]+\.16b\n} > >> 1 } } */ > >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.8h, v[0- > 9]+\.16b\n} > >> 1 } } */ > >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0- > 9]+\.8h\n} 1 > >> } } */ > >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0- > 9]+\.8h\n} 1 > >> } } */ > >> /* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} > 1 > >> } } */ > >> -/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0- > 9]+\.4s\n} 1 > >> } } */ > >> -/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ > >> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0- > 9]+\.8h\n} 1 > >> } } */ > >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0- > 9]+\.4s\n} 3 > >> } } */ > >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0- > 9]+\.4s\n} 3 > >> } } */ > >> + > >> +/* The byte to halfword step is what the dot product replaces. */ > >> +/* { dg-final { scan-assembler-not {\t[su]addlp\tv[0-9]+\.8h, v[0- > >> 9]+\.16b\n} } } */ > >> +/* { dg-final { scan-assembler-not {\t[su]addw2?\t} } } */ > >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c > >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c > >> new file mode 100644 > >> index 00000000000..d2eb8152722 > >> --- /dev/null > >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c > >> @@ -0,0 +1,83 @@ > >> +/* { dg-do run } */ > >> +/* { dg-require-effective-target arm_v8_2a_dotprod_neon_hw } */ > >> +/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec- > >> preference=asimd-only" } */ > >> + > >> +/* Both expansions of a widening sum reduction, with and without dot > >> + product, must agree with a scalar sum for every narrow to wide type > >> + pair. The accumulators are unsigned so that overflow wraps. */ > >> + > >> +#define TYPES(X) \ > >> + X (u8_h, unsigned char, unsigned short) \ > >> + X (i8_h, signed char, unsigned short) \ > >> + X (u8_i, unsigned char, unsigned int) \ > >> + X (i8_i, signed char, unsigned int) \ > >> + X (u8_l, unsigned char, unsigned long) \ > >> + X (i8_l, signed char, unsigned long) \ > >> + X (u16_i, unsigned short, unsigned int) \ > >> + X (i16_i, short, unsigned int) \ > >> + X (u16_l, unsigned short, unsigned long) \ > >> + X (i16_l, short, unsigned long) \ > >> + X (u32_l, unsigned int, unsigned long) \ > >> + X (i32_l, int, unsigned long) > >> + > >> +#define SUM(PREFIX, NAME, ITYPE, OTYPE) \ > >> + __attribute__ ((noipa)) \ > >> + OTYPE PREFIX##_##NAME (const ITYPE *a, int n) \ > >> + { \ > >> + OTYPE s = 0; \ > >> + for (int i = 0; i < n; i++) \ > >> + s += a[i]; \ > >> + return s; \ > >> + } > >> + > >> +#define DOT(NAME, ITYPE, OTYPE) SUM (dot, NAME, ITYPE, OTYPE) > >> +#define NODOT(NAME, ITYPE, OTYPE) SUM (nodot, NAME, ITYPE, OTYPE) > >> + > >> +/* A volatile accumulator keeps this loop scalar. */ > >> +#define REF(NAME, ITYPE, OTYPE) \ > >> + __attribute__ ((noipa)) \ > >> + OTYPE ref_##NAME (const ITYPE *a, int n) \ > >> + { \ > >> + volatile OTYPE s = 0; \ > >> + for (int i = 0; i < n; i++) \ > >> + s = s + a[i]; \ > >> + return s; \ > >> + } > >> + > >> +TYPES (DOT) > >> +TYPES (REF) > >> + > >> +#pragma GCC push_options > >> +#pragma GCC target ("+nodotprod") > >> +TYPES (NODOT) > >> +#pragma GCC pop_options > >> + > >> +#define BYTES 8192 > >> +static unsigned char buf[BYTES] __attribute__ ((aligned (64))); > >> + > >> +#define CHECK(NAME, ITYPE, OTYPE) \ > >> + { \ > >> + const ITYPE *p = (const ITYPE *) (buf + off); \ > >> + OTYPE want = ref_##NAME (p, n); \ > >> + if (dot_##NAME (p, n) != want || nodot_##NAME (p, n) != want) \ > >> + __builtin_abort (); \ > >> + } > >> + > >> +int > >> +main (void) > >> +{ > >> + unsigned long x = 1; > >> + for (int i = 0; i < BYTES; i++) > >> + { > >> + x = x * 6364136223846793005UL + 1442695040888963407UL; > >> + buf[i] = x >> 40; > >> + } > >> + > >> + for (int off = 0; off < 8; off += 4) > >> + for (int n = 0; n <= 260; n++) > >> + { > >> + TYPES (CHECK) > >> + } > >> + > >> + return 0; > >> +} > >> -- > >> 2.50.1 (Apple Git-155) >
