> -----Original Message----- > From: [email protected] <[email protected]> > Sent: 12 August 2026 14:11 > To: [email protected] > Cc: Tamar Christina <[email protected]>; Kyrylo Tkachov > <[email protected]> > Subject: [PATCH] aarch64: complex multiply with FCMLA without -fno-signed- > zeros [PR126589] > > From: Kyrylo Tkachov <[email protected]> > > FCMLA only accumulates, so a plain complex multiply was built by seeding the > accumulator with +0.0. That turns a -0.0 product into +0.0, so the fix for > PR126589 gave the cmul optabs a !HONOR_SIGNED_ZEROS requirement, > which means > FCMLA is only reached with -ffast-math and not with -fcx-limited-range alone. > > The first rotation of the pair computes > > Vd.re = Vd.re + Vn.re * Vm.re > Vd.im = Vd.im + Vn.re * Vm.im > > that is, a multiply of Vm by the real part of Vn broadcast over its own > pair. TRN1 of Vn with itself is that broadcast, so the accumulator is not > needed at all: > > trn1 vt.4s, vn.4s, vn.4s > fmul vd.4s, vt.4s, vm.4s > fcmla vd.4s, vn.4s, vm.4s, #90 > > The sign of a zero now survives, and there is one fewer accumulate. Where > signed zeros are not honoured the previous sequence is kept: MOVI with a > zero > immediate is a zeroing idiom, so seeding the accumulator there is free, while > materialising any other constant is not. > > The same applies to the SVE expander. > > On Neoverse V2, over arrays of 4096 _Complex float at -O3 -mcpu=neoverse- > v2 > -fcx-limited-range, c[i] = a[i] * b[i] drops 5.9% and y[i] += alpha * x[i] > drops 21.8% in run time. -ffast-math is unchanged. > > Bootstrapped and tested on aarch64-none-linux-gnu. > Ok for trunk?
Sure, though this should be reverted once PR121925 lands because it hides the permutations from the vectorizer. For now this is fine though. Thanks, Tamar > Thanks, > Kyrill > > gcc/ChangeLog: > > PR target/126589 > * config/aarch64/aarch64-protos.h (aarch64_emit_cmul_first_half): > Declare. > * config/aarch64/aarch64.cc (aarch64_emit_cmul_first_half): New > function. > * config/aarch64/aarch64-simd.md (cmul<conj_op><mode>3): Use it > when > signed zeros are honoured, and drop the requirement. > * config/aarch64/aarch64-sve.md (cmul<conj_op><mode>3): > Likewise. > > gcc/testsuite/ChangeLog: > > PR target/126589 > * gcc.target/aarch64/complex-mul-signed-zeros.c: New test. > * gcc.target/aarch64/sve/complex-mul-signed-zeros.c: New test. > * gcc.dg/vect/complex/complex-mul-signed-zero-run.c: New test. > > Signed-off-by: Kyrylo Tkachov <[email protected]> > --- > gcc/config/aarch64/aarch64-protos.h | 1 + > gcc/config/aarch64/aarch64-simd.md | 19 +++- > gcc/config/aarch64/aarch64-sve.md | 19 +++- > gcc/config/aarch64/aarch64.cc | 24 ++++ > .../complex/complex-mul-signed-zero-run.c | 105 ++++++++++++++++++ > .../aarch64/complex-mul-signed-zeros.c | 39 +++++++ > .../aarch64/sve/complex-mul-signed-zeros.c | 27 +++++ > 7 files changed, 224 insertions(+), 10 deletions(-) > create mode 100644 gcc/testsuite/gcc.dg/vect/complex/complex-mul- > signed-zero-run.c > create mode 100644 gcc/testsuite/gcc.target/aarch64/complex-mul-signed- > zeros.c > create mode 100644 gcc/testsuite/gcc.target/aarch64/sve/complex-mul- > signed-zeros.c > > diff --git a/gcc/config/aarch64/aarch64-protos.h > b/gcc/config/aarch64/aarch64-protos.h > index bcc833cfaa1..24668aca125 100644 > --- a/gcc/config/aarch64/aarch64-protos.h > +++ b/gcc/config/aarch64/aarch64-protos.h > @@ -1047,6 +1047,7 @@ rtx aarch64_convert_sve_data_to_pred (rtx, rtx); > rtx aarch64_expand_sve_dupq (rtx, machine_mode, rtx); > void aarch64_expand_mov_immediate (rtx, rtx); > rtx aarch64_stack_protect_canary_mem (machine_mode, rtx, > aarch64_salt_type); > +void aarch64_emit_cmul_first_half (rtx, rtx, rtx); > rtx aarch64_ptrue_reg (machine_mode); > rtx aarch64_ptrue_reg (machine_mode, unsigned int); > rtx aarch64_ptrue_reg (machine_mode, machine_mode); > diff --git a/gcc/config/aarch64/aarch64-simd.md > b/gcc/config/aarch64/aarch64-simd.md > index 8ad1ec20f46..48ca6b10d51 100644 > --- a/gcc/config/aarch64/aarch64-simd.md > +++ b/gcc/config/aarch64/aarch64-simd.md > @@ -681,12 +681,23 @@ > (unspec:VHSDF [(match_operand:VHSDF 1 "register_operand") > (match_operand:VHSDF 2 "register_operand")] > FCMUL_OP))] > - "TARGET_COMPLEX && !BYTES_BIG_ENDIAN && !HONOR_SIGNED_ZEROS > (<MODE>mode)" > + "TARGET_COMPLEX && !BYTES_BIG_ENDIAN" > { > - rtx tmp = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode)); > rtx res1 = gen_reg_rtx (<MODE>mode); > - emit_insn (gen_aarch64_fcmla<rotsplit1><mode> (res1, tmp, > - operands[2], operands[1])); > + if (HONOR_SIGNED_ZEROS (<MODE>mode)) > + /* The first rotation multiplies the real part of operand 2 by the whole > + of operand 1, so express it as a multiply rather than as an accumulate > + into zero. Accumulating into +0.0 would turn a -0.0 product into > + +0.0. */ > + aarch64_emit_cmul_first_half (res1, operands[2], operands[1]); > + else > + { > + /* MOVI with a zero immediate is a zeroing idiom, so seeding the > + accumulator here is free. */ > + rtx zero = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode)); > + emit_insn (gen_aarch64_fcmla<rotsplit1><mode> (res1, zero, > + operands[2], > operands[1])); > + } > emit_insn (gen_aarch64_fcmla<rotsplit2><mode> (operands[0], res1, > operands[2], operands[1])); > DONE; > diff --git a/gcc/config/aarch64/aarch64-sve.md > b/gcc/config/aarch64/aarch64-sve.md > index 5f9a19c42e9..539030ab1bc 100644 > --- a/gcc/config/aarch64/aarch64-sve.md > +++ b/gcc/config/aarch64/aarch64-sve.md > @@ -8300,16 +8300,23 @@ > [(match_operand:SVE_FULL_F 1 "register_operand") > (match_operand:SVE_FULL_F 2 "register_operand")] > FCMUL_OP))] > - "TARGET_SVE && !HONOR_SIGNED_ZEROS (<MODE>mode)" > + "TARGET_SVE" > { > rtx pred_reg = aarch64_ptrue_reg (<VPRED>mode); > rtx gp_mode = gen_int_mode (SVE_RELAXED_GP, SImode); > - rtx accum = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode)); > rtx tmp = gen_reg_rtx (<MODE>mode); > - emit_insn > - (gen_aarch64_pred_fcmla<sve_rot1><mode> (tmp, pred_reg, > - operands[2], operands[1], > - accum, gp_mode)); > + if (HONOR_SIGNED_ZEROS (<MODE>mode)) > + /* Accumulating into +0.0 would turn a -0.0 product into +0.0, so do the > + first rotation as the equivalent explicit multiply. */ > + aarch64_emit_cmul_first_half (tmp, operands[2], operands[1]); > + else > + { > + rtx accum = force_reg (<MODE>mode, CONST0_RTX (<MODE>mode)); > + emit_insn > + (gen_aarch64_pred_fcmla<sve_rot1><mode> (tmp, pred_reg, > + operands[2], operands[1], > + accum, gp_mode)); > + } > emit_insn > (gen_aarch64_pred_fcmla<sve_rot2><mode> (operands[0], pred_reg, > operands[2], operands[1], > diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc > index 3041a6ee62a..1a2c578381c 100644 > --- a/gcc/config/aarch64/aarch64.cc > +++ b/gcc/config/aarch64/aarch64.cc > @@ -4229,6 +4229,30 @@ aarch64_ptrue_reg (machine_mode pred_mode, > machine_mode data_mode) > return aarch64_ptrue_reg (pred_mode, size); > } > > +/* Emit into TARGET the first half of the complex multiplication of A by B, > + that is the pair (Are * Bre, Are * Bim) for each complex element. This is > + what FCMLA with a rotation of #0 computes, but as a plain multiply, so no > + additive identity is needed and the sign of a zero product survives. > + TARGET, A and B all have the same floating-point vector mode. */ > + > +void > +aarch64_emit_cmul_first_half (rtx target, rtx a, rtx b) > +{ > + machine_mode mode = GET_MODE (target); > + rtx dup_real = gen_reg_rtx (mode); > + /* TRN1 of a vector with itself broadcasts each even element over the > + odd one, which for complex data is the real part over its own pair. > + The two inputs are the same register, so no big-endian correction is > + needed here. */ > + emit_set_insn (dup_real, > + gen_rtx_UNSPEC (mode, gen_rtvec (2, a, a), UNSPEC_TRN1)); > + rtx res = expand_binop (mode, smul_optab, dup_real, b, target, 0, > + OPTAB_DIRECT); > + gcc_assert (res); > + if (res != target) > + emit_move_insn (target, res); > +} > + > /* Return an all-false predicate register of mode MODE. */ > > rtx > diff --git a/gcc/testsuite/gcc.dg/vect/complex/complex-mul-signed-zero-run.c > b/gcc/testsuite/gcc.dg/vect/complex/complex-mul-signed-zero-run.c > new file mode 100644 > index 00000000000..0e5f98144c6 > --- /dev/null > +++ b/gcc/testsuite/gcc.dg/vect/complex/complex-mul-signed-zero-run.c > @@ -0,0 +1,105 @@ > +/* { dg-do run } */ > +/* { dg-require-effective-target vect_complex_add_double } */ > +/* { dg-additional-options "-fcx-limited-range" } */ > +/* { dg-add-options arm_v8_3a_complex_neon } */ > + > +/* Vectorised complex multiplication must keep the sign of a zero result > + when -fno-signed-zeros is not in effect. */ > + > +#include <math.h> > + > +#define N 64 > + > +_Complex float af[N], bf[N], cf[N], reff[N]; > +_Complex double ad[N], bd[N], cd[N], refd[N]; > + > +__attribute__((noipa)) void > +mulf (_Complex float *__restrict c, _Complex float *__restrict a, > + _Complex float *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +__attribute__((noipa, optimize ("no-tree-vectorize"))) void > +mulf_ref (_Complex float *__restrict c, _Complex float *__restrict a, > + _Complex float *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +__attribute__((noipa)) void > +muld (_Complex double *__restrict c, _Complex double *__restrict a, > + _Complex double *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +__attribute__((noipa, optimize ("no-tree-vectorize"))) void > +muld_ref (_Complex double *__restrict c, _Complex double *__restrict a, > + _Complex double *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +/* Compare bit patterns so that the sign of a zero matters. Any two NaNs > + are equal, their sign and payload are unspecified. */ > + > +static int > +eqf (float x, float y) > +{ > + if (isnan (x) && isnan (y)) > + return 1; > + return __builtin_memcmp (&x, &y, sizeof x) == 0; > +} > + > +static int > +eqd (double x, double y) > +{ > + if (isnan (x) && isnan (y)) > + return 1; > + return __builtin_memcmp (&x, &y, sizeof x) == 0; > +} > + > +static const float vals[] > + = { 0.0f, -0.0f, 1.0f, -1.0f, 3.5f, -2.25f, > + __builtin_inff (), -__builtin_inff (), __builtin_nanf ("") }; > +#define NV ((int) (sizeof (vals) / sizeof (vals[0]))) > + > +int > +main (void) > +{ > + for (int i0 = 0; i0 < NV; i0++) > + for (int i1 = 0; i1 < NV; i1++) > + for (int i2 = 0; i2 < NV; i2++) > + for (int i3 = 0; i3 < NV; i3++) > + { > + for (int k = 0; k < N; k++) > + { > + __real__ af[k] = vals[i0]; > + __imag__ af[k] = vals[i1]; > + __real__ bf[k] = vals[i2]; > + __imag__ bf[k] = vals[i3]; > + __real__ ad[k] = vals[i0]; > + __imag__ ad[k] = vals[i1]; > + __real__ bd[k] = vals[i2]; > + __imag__ bd[k] = vals[i3]; > + } > + > + mulf (cf, af, bf, N); > + mulf_ref (reff, af, bf, N); > + muld (cd, ad, bd, N); > + muld_ref (refd, ad, bd, N); > + > + for (int k = 0; k < N; k++) > + if (!eqf (__real__ cf[k], __real__ reff[k]) > + || !eqf (__imag__ cf[k], __imag__ reff[k]) > + || !eqd (__real__ cd[k], __real__ refd[k]) > + || !eqd (__imag__ cd[k], __imag__ refd[k])) > + __builtin_abort (); > + } > + return 0; > +} > diff --git a/gcc/testsuite/gcc.target/aarch64/complex-mul-signed-zeros.c > b/gcc/testsuite/gcc.target/aarch64/complex-mul-signed-zeros.c > new file mode 100644 > index 00000000000..a15accf8e07 > --- /dev/null > +++ b/gcc/testsuite/gcc.target/aarch64/complex-mul-signed-zeros.c > @@ -0,0 +1,39 @@ > +/* { dg-do compile } */ > +/* { dg-options "-O3 -march=armv8.3-a+nosve -fcx-limited-range" } */ > + > +/* FCMLA only accumulates, so seeding it with +0.0 to get a plain complex > + multiply would turn a -0.0 product into +0.0. Doing the first rotation as > + an explicit multiply keeps the sign of a zero and needs no accumulator, so > + this works without -fno-signed-zeros. */ > + > +void > +mulf (_Complex float *__restrict c, _Complex float *__restrict a, > + _Complex float *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +void > +muld (_Complex double *__restrict c, _Complex double *__restrict a, > + _Complex double *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +void > +mulconjf (_Complex float *__restrict c, _Complex float *__restrict a, > + _Complex float *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * ~b[i]; > +} > + > +/* { dg-final { scan-assembler {trn1\tv[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s} > } } > */ > +/* { dg-final { scan-assembler {trn1\tv[0-9]+\.2d, v[0-9]+\.2d, v[0-9]+\.2d} > } > } */ > +/* { dg-final { scan-assembler {fcmla\tv[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s, > #90} } } */ > +/* { dg-final { scan-assembler {fcmla\tv[0-9]+\.2d, v[0-9]+\.2d, v[0-9]+\.2d, > #90} } } */ > +/* { dg-final { scan-assembler {fcmla\tv[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s, > #270} } } */ > +/* No accumulator is needed, so no zero constant is materialised. */ > +/* { dg-final { scan-assembler-not {fcmla\tv[0-9]+\.[24][sd], v[0- > 9]+\.[24][sd], v[0-9]+\.[24][sd], #0} } } */ > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/complex-mul-signed-zeros.c > b/gcc/testsuite/gcc.target/aarch64/sve/complex-mul-signed-zeros.c > new file mode 100644 > index 00000000000..a79332460f5 > --- /dev/null > +++ b/gcc/testsuite/gcc.target/aarch64/sve/complex-mul-signed-zeros.c > @@ -0,0 +1,27 @@ > +/* { dg-do compile } */ > +/* { dg-options "-O3 -march=armv8.3-a+sve -msve-vector-bits=scalable -fcx- > limited-range" } */ > + > +/* The SVE complex multiply also does the first rotation as an explicit > + multiply, so it does not need -fno-signed-zeros either. */ > + > +void > +mulf (_Complex float *__restrict c, _Complex float *__restrict a, > + _Complex float *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +void > +muld (_Complex double *__restrict c, _Complex double *__restrict a, > + _Complex double *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + c[i] = a[i] * b[i]; > +} > + > +/* { dg-final { scan-assembler {trn1\tz[0-9]+\.s, z[0-9]+\.s, z[0-9]+\.s} } > } */ > +/* { dg-final { scan-assembler {trn1\tz[0-9]+\.d, z[0-9]+\.d, z[0-9]+\.d} } > } */ > +/* { dg-final { scan-assembler {fcmla\tz[0-9]+\.s, p[0-9]+/m, z[0-9]+\.s, > z[0- > 9]+\.s, #90} } } */ > +/* { dg-final { scan-assembler {fcmla\tz[0-9]+\.d, p[0-9]+/m, z[0-9]+\.d, > z[0- > 9]+\.d, #90} } } */ > +/* { dg-final { scan-assembler-not {fcmla\tz[0-9]+\.[sd], p[0-9]+/m, z[0- > 9]+\.[sd], z[0-9]+\.[sd], #0} } } */ > -- > 2.50.1 (Apple Git-155)
