On Tue, Aug 4, 2026 at 2:53 AM <[email protected]> wrote:
>
> From: Kyrylo Tkachov <[email protected]>
>
> A lane of A | B is zero exactly when the corresponding lanes of A and of B
> are both zero, so
>
> (A == 0) & (B == 0) -> (A | B) == 0
>
> and the De Morgan dual for the inequality. One vector comparison goes
> away. Reassociation performs this for scalars, but it never runs on
> vector masks, so the vector form is left alone today.
>
> typedef int v4si __attribute__((vector_size (16)));
> v4si f (v4si a, v4si b) { v4si z = {0,0,0,0}; return (a == z) & (b == z); }
>
> aarch64 -O3 before:
>
> cmeq v0.4s, v0.4s, #0
> cmeq v1.4s, v1.4s, #0
> and v0.16b, v0.16b, v1.16b
>
> after:
>
> orr v0.16b, v0.16b, v1.16b
> cmeq v0.4s, v0.4s, #0
>
> Bootstrapped and tested on aarch64-none-linux-gnu.
> Ok for trunk?
> Thanks,
> Kyrill
>
> gcc/ChangeLog:
>
> * match.pd ((A == 0) & (B == 0), (A != 0) | (B != 0)): New
> simplifications for vector operands.
>
> gcc/testsuite/ChangeLog:
>
> * gcc.dg/tree-ssa/vec-mask-zero-1.c: New test.
>
> Signed-off-by: Kyrylo Tkachov <[email protected]>
> ---
> gcc/match.pd | 14 ++++++++++++++
> gcc/testsuite/gcc.dg/tree-ssa/vec-mask-zero-1.c | 12 ++++++++++++
> 2 files changed, 26 insertions(+)
> create mode 100644 gcc/testsuite/gcc.dg/tree-ssa/vec-mask-zero-1.c
>
> diff --git a/gcc/match.pd b/gcc/match.pd
> index 22202af2cc1..fc81dfc5e66 100644
> --- a/gcc/match.pd
> +++ b/gcc/match.pd
> @@ -3823,6 +3823,20 @@ DEFINE_INT_AND_FLOAT_ROUND_FN (RINT)
> { constant_boolean_node (true, type); })
> ))))))
>
> +/* Combine two vector comparisons against zero into one:
> + (A == 0) & (B == 0) --> (A | B) == 0
> + (A != 0) | (B != 0) --> (A | B) != 0
> + Reassociation does this for scalars only, it never runs on vector
> + masks. */
> +(for eqne (eq ne)
> + bitop (bit_and bit_ior)
> + (simplify
> + (bitop (eqne:s @0 integer_zerop) (eqne:s @1 integer_zerop))
> + (if (VECTOR_INTEGER_TYPE_P (TREE_TYPE (@0))
> + && types_match (TREE_TYPE (@0), TREE_TYPE (@1)))
> + (eqne (bit_ior @0 @1)
> + { build_zero_cst (TREE_TYPE (@0)); }))))
So I think you should be able to combine it with these patterns instead:
```
(for bitop (bit_and bit_ior)
cmp (eq ne)
/* PR35691: Transform
(x == 0 & y == 0) -> (x | typeof(x)(y)) == 0.
(x != 0 | y != 0) -> (x | typeof(x)(y)) != 0. */
(simplify
(bitop (cmp @0 integer_zerop@2) (cmp @1 integer_zerop))
(if (INTEGRAL_TYPE_P (TREE_TYPE (@0))
&& INTEGRAL_TYPE_P (TREE_TYPE (@1))
&& TYPE_PRECISION (TREE_TYPE (@0)) == TYPE_PRECISION (TREE_TYPE (@1)))
(cmp (bit_ior @0 (convert @1)) @2)))
/* Transform:
(x == -1 & y == -1) -> (x & typeof(x)(y)) == -1.
(x != -1 | y != -1) -> (x & typeof(x)(y)) != -1. */
(simplify
(bitop (cmp @0 integer_all_onesp@2) (cmp @1 integer_all_onesp))
(if (INTEGRAL_TYPE_P (TREE_TYPE (@0))
&& INTEGRAL_TYPE_P (TREE_TYPE (@1))
&& TYPE_PRECISION (TREE_TYPE (@0)) == TYPE_PRECISION (TREE_TYPE (@1)))
(cmp (bit_and @0 (convert @1)) @2))))
```
Instead of convert, use view_convert and instead of TYPE_PRECISION
check you need something more like:
```
(if (VECTOR_TYPE_P (type) && VECTOR_TYPE_P (TREE_TYPE (@0))
&& known_eq (TYPE_VECTOR_SUBPARTS (type),
TYPE_VECTOR_SUBPARTS (TREE_TYPE (@0)))
&& tree_nop_conversion_p (TREE_TYPE (type), TREE_TYPE (TREE_TYPE (@0))))))
```
Which might make sense to make in a helper function and use that also
in nop_convert match pattern; maybe vector_nop_conversion_p.
And then also add a testcase for -1.
> +
> /* Optimize (a CMP b) ^ (a CMP b) */
> /* Optimize (a CMP b) != (a CMP b) */
> (for op (bit_xor ne)
> diff --git a/gcc/testsuite/gcc.dg/tree-ssa/vec-mask-zero-1.c
> b/gcc/testsuite/gcc.dg/tree-ssa/vec-mask-zero-1.c
> new file mode 100644
> index 00000000000..f8c669636fd
> --- /dev/null
> +++ b/gcc/testsuite/gcc.dg/tree-ssa/vec-mask-zero-1.c
> @@ -0,0 +1,12 @@
> +/* { dg-do compile } */
> +/* { dg-options "-O2 -fdump-tree-optimized" } */
> +/* A lane of A | B is zero exactly when both lanes are, so a pair of vector
> + comparisons against zero becomes one. Reassociation does this for
> + scalars but never runs on vector masks. */
> +typedef int v4si __attribute__((vector_size (16)));
> +v4si f (v4si a, v4si b) { v4si z = {0,0,0,0}; return (a == z) & (b == z); }
> +v4si g (v4si a, v4si b) { v4si z = {0,0,0,0}; return (a != z) | (b != z); }
> +/* { dg-final { scan-tree-dump-times " == " 1 "optimized" } } */
> +/* { dg-final { scan-tree-dump-times " != " 1 "optimized" } } */
> +/* Each function keeps one OR and one comparison. */
> +/* { dg-final { scan-tree-dump-times " \\| " 2 "optimized" } } */
> --
> 2.50.1 (Apple Git-155)
>