This enables use of a predicate mask or length limit for
vectorization of basic blocks in cases where previously only the
equivalent rolled (i.e. loop) form of some source code would have
been vectorized. Predication is used for groups whose size
is not neatly divisible into vectors of lengths that can be
supported directly by the target.

The initial vector mode for an SLP region is "autodetected" by calling
aarch64_preferred_simd_mode, which prefers SVE modes if supported
unless configured otherwise (e.g. VNx4SI for int). If at least one
profitable subgraph can be scheduled then GCC does not try to vectorize
the region using any other modes, even though their estimated costs
might otherwise have been lower.

For example, if analysis of a 24-byte group succeeds with vector mode
V16QI (using types vector(16) and vector(8) char) then the estimated
cost of the vectorized code is 11+11=22. If analysis of the same group
succeeds with vector mode VNx16QI (using type vector([16,16]) char for
both subtrees) then the estimated cost is 15+15=30. In both cases, the
estimated vectorized cost would beat the estimated scalar cost of
96+48=144, but vector([16,16]) is chosen because VNx16QI is tried
first.

This is mitigated by the fact that a sequence of GIMPLE stmts such as:

vectp.14_86 = x_50(D) + 16;
slp_mask_87 = .WHILE_ULT (0, 8, { 0, ... });
.MASK_STORE (vectp.14_86, 8B, slp_mask_87, vect__34.12_85);

are lowered to a fixed-length vector store (e.g., str d30, [x0, 16]) if
possible, instead of a more literal interpretation such as:

add   x0, x0, 16
ptrue p7.b, vl8
st1b  z30.b, p7, [x0]

The vect_record_max_nunits function used during building of an SLP
tree is updated to prevent it returning failure for BB SLP if the
group size is not an integral multiple of the number of lanes in the
vector type; it now allows such cases if the group size is known to
be less than or equal to the minimum number of lanes in the vector type.

For BB SLP, vect_analyze_slp_instance previously gave up after
building an SLP tree if it could not prove that the group size was
at least the maximum lane count across all of the vector types in
the SLP tree (which is unprovable for scalable vector types), or
attempted to split the group if it could prove that the group size
was greater than this maximum but not exactly divisible by it
(which is also unprovable for scalable vector types).

This function will now provisionally create a new SLP instance if the
group size definitely does not exceed the minimum number of lanes,
even if the group size otherwise satisfies conditions that would
require a loop to be unrolled (e.g., a group of size 3 that uses a
mixture of V4SI and V8HI types). If the group size lies between the
minimum and maximum number of lanes then vectorization is still
abandoned (e.g., a group of size 3 that uses a mixture of
V2DI and V4SI types).

The vect_prologue_cost_for_slp and vect_create_constant_vectors
functions expect to be able to get the number of vectors needed for
an SLP node by dividing the number of scalar elements needed for the
vectorized statements by the number of subparts in the vector type.
This relies on the scalar elements occupying a whole number of vectors
or else one partial vector, which is true of nodes created during
building of the SLP tree but not always true when a vector type is
assigned to an external definition node during analysis of statements.
Update vect_maybe_update_slp_op_vectype to reject late attempts to
assign a vector type that would result in an external definition
that occupies "one and half" vectors or similar.

Update test expectations for gcc.dg/vect/vect-over-widen-*.c,
gcc.target/aarch64/sve/slp_6.c and
gcc.target/aarch64/sve/vec_construct_*.c.

The vec_construct_*.c tests previously expected their output
to use Advanced SIMD instead of SVE despite their use of
vector length agnostic types such as svint16_t and despite
the fact that they are in the aarch64/sve directory. Since
BB SLP can now vectorize these tests using VLA types such
as 'vector([8,8]) char', and because (with one exception) the
resultant code is deemed profitable relative to scalar code,
GCC no longer considers vectorizing using non-VLA types such
as 'vector(8) char' (although the estimated cost with non-VLA
types might have been lower, had it been calculated).
Instruction selection is not the focus of these tests, therefore
I updated them to expect SVE instead (e.g. st1b instead of str)
and added --param=aarch64-autovec-preference=sve-only to reduce
future churn.

Because the cost model takes into account predicate mask
generation for BB SLP with VLA types, the threshold at which
vectorized code wins against scalar code is higher than
before. The number of elements stored by vec_construct_3.c was
increased just enough to allow for that.

gcc/ChangeLog:

        * tree-vect-slp.cc (vect_record_max_nunits): For BB SLP
        vectorization, allow group sizes that are indivisible by the
        vector length if they are known not to exceed it.
        (vect_build_slp_tree_1): If building SLP failed because of an
        unsupported data type or because unrolling was required then use the
        known minimum number of subparts when choosing a splitting point for
        a VLA vector type.
        (vect_build_slp_tree_2): Don't call
        can_duplicate_and_interleave_p when doing basic block SLP
        vectorization.
        (vect_update_slp_min_nunits_for_node): New recursive function.
        Update min_nunits to reflect the minimum number of subparts for
        all of the vector types used by an SLP subgraph.
        (vect_slp_tree_min_nunits): New function. Initialize min_nunits
        then call vect_update_slp_min_nunits_for_node.
        (vect_analyze_slp_instance): For BB SLP vectorization, create
        a new SLP instance if the group size definitely does not exceed
        the minimum number of subparts for all of the vector types used
        in the SLP tree, even if the group size otherwise satisfies
        conditions that would require a loop to be unrolled.
        (vectorizable_slp_permutation_1): Instead of asserting that an
        SLP tree node's number of lanes is compatible with the chosen
        vector width, return a failure indication if incompatible.
        * tree-vect-stmts.cc (get_vectype_for_scalar_type): Execute the
        old algorithm to find a narrower vector type that can be used to
        carve an SLP group into pieces only if partial vector support is
        unavailable for either loads or stores, and only if the number of
        subparts may be greater than the group size rather than also if
        the number of subparts may be equal to the group size.
        Otherwise, assume that using a partial vector is more efficient
        than carving the group into smaller pieces.  If the natural
        vector type is not VLA and it has more subparts than the group
        size then find the smallest vector type that contains the group.
        (vect_maybe_update_slp_op_vectype): Reject external definitions
        when the number of scalar elements represented by the vectorized
        operation is not divisible by the number of subparts in a vector
        type inferred from the scalar type.  This takes into account the
        vectorization factor like vect_get_num_copies although it should
        be 1 for BB SLP vectorization.
        (vect_get_vector_types_for_stmt): Update the description.

gcc/testsuite/ChangeLog:

        * gcc.dg/vect/vect-over-widen-10.c: Update test expectations to
        avoid spurious matching of scan-tree-dump-not pattern.
        * gcc.dg/vect/vect-over-widen-13.c: As above.
        * gcc.dg/vect/vect-over-widen-14.c: As above.
        * gcc.dg/vect/vect-over-widen-17.c: As above.
        * gcc.dg/vect/vect-over-widen-18.c: As above.
        * gcc.dg/vect/vect-over-widen-5.c: As above.
        * gcc.dg/vect/vect-over-widen-6.c: As above.
        * gcc.dg/vect/vect-over-widen-7.c: As above.
        * gcc.dg/vect/vect-over-widen-8.c: As above.
        * gcc.dg/vect/vect-over-widen-9.c: As above.
        * gcc.target/aarch64/sve/slp_6.c: Require ld1b, ld1h and ld1w
        instructions in the output instead of forbidding them because
        discovery of reduction groups now succeeds.  Also require
        the corresponding st1b, st1h and st1w instructions.
        * gcc.target/aarch64/sve/vec_construct_1.c:
        Expect SVE instead of ASIMD instructions and add
        --param=aarch64-autovec-preference=sve-only to stop
        flip-flopping.
        * gcc.target/aarch64/sve/vec_construct_2.c: Forbid SVE
        instructions as well as ASIMD instructions and add
        --param=aarch64-autovec-preference=sve-only to stop
        flip-flopping.
        * gcc.target/aarch64/sve/vec_construct_3.c: Expect SVE instead
        of ASIMD instructions and add
        --param=aarch64-autovec-preference=sve-only to avoid
        flip-flopping. Increase the number of elements stored to
        ensure vectorization using SVE is deemed profitable despite
        predicate mask costs.
        * gcc.target/aarch64/sve/vec_construct_4.c:
        Forbid SVE instructions as well as ASIMD instructions and add
        --param=aarch64-autovec-preference=sve-only to stop
        flip-flopping.  Fix a missing end of comment delimiter.
        * gcc.target/aarch64/sve/vec_construct_5.c: As above.
---
 .../gcc.dg/vect/vect-over-widen-10.c          |   2 +-
 .../gcc.dg/vect/vect-over-widen-13.c          |   2 +-
 .../gcc.dg/vect/vect-over-widen-14.c          |   2 +-
 .../gcc.dg/vect/vect-over-widen-17.c          |   2 +-
 .../gcc.dg/vect/vect-over-widen-18.c          |   2 +-
 gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c |   2 +-
 gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c |   2 +-
 gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c |   2 +-
 gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c |   2 +-
 gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c |   2 +-
 gcc/testsuite/gcc.target/aarch64/sve/slp_6.c  |  18 ++-
 .../gcc.target/aarch64/sve/vec_construct_1.c  |   6 +-
 .../gcc.target/aarch64/sve/vec_construct_2.c  |   4 +-
 .../gcc.target/aarch64/sve/vec_construct_3.c  |  20 +++-
 .../gcc.target/aarch64/sve/vec_construct_4.c  |   5 +-
 .../gcc.target/aarch64/sve/vec_construct_5.c  |   6 +-
 gcc/tree-vect-slp.cc                          | 113 +++++++++++++++---
 gcc/tree-vect-stmts.cc                        |  97 ++++++++++++---
 18 files changed, 228 insertions(+), 61 deletions(-)

diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c
index f0140e4ef6d..6efcf739db9 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-10.c
@@ -16,5 +16,5 @@
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 1} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 2} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* \(unsigned char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c
index 08a65ea5518..720353716cf 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-13.c
@@ -48,5 +48,5 @@ main (void)
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* \+} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* / 2} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* = \(signed char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c
index dfa09f5d2ca..f1d5f95c543 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-14.c
@@ -15,5 +15,5 @@
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* \+} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 1} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* = \(unsigned char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c
index 53fcfd0c06c..ac1a0f86727 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-17.c
@@ -46,5 +46,5 @@ main (void)
    adopts realign_load scheme.  It requires rs6000_builtin_mask_for_load to
    generate mask whose return type is vector char.  */
 /* { dg-final { scan-tree-dump-not {vector[^\n]*char} "vect" { target 
vect_hw_misalign } } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c
index aa58cd1c957..3ebfaa78270 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-18.c
@@ -47,5 +47,5 @@ main (void)
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* |} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* <<} "vect" } } */
 /* { dg-final { scan-tree-dump {vector[^\n]*char} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c
index c2ab11a9d32..1d89789a86d 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-5.c
@@ -49,5 +49,5 @@ main (void)
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* \+ } "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 1} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* \(signed char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c
index bda92c965e0..62d5a52587e 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-6.c
@@ -13,5 +13,5 @@
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* \+ } "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 1} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* \(unsigned char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c
index 1d55e13fb1f..6e09631009a 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-7.c
@@ -51,5 +51,5 @@ main (void)
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* \+ } "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 2} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* \(signed char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c
index 553c0712a79..b6d650beab4 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-8.c
@@ -16,5 +16,5 @@
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* \+ } "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 2} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* \(unsigned char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c 
b/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c
index 36bfc68e053..e82f8a571da 100644
--- a/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c
+++ b/gcc/testsuite/gcc.dg/vect/vect-over-widen-9.c
@@ -56,5 +56,5 @@ main (void)
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 1} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_over_widening_pattern: 
detected:[^\n]* >> 2} "vect" } } */
 /* { dg-final { scan-tree-dump {vect_recog_cast_forwprop_pattern: 
detected:[^\n]* \(signed char\)} "vect" } } */
-/* { dg-final { scan-tree-dump-not {vector[^ ]* int} "vect" } } */
+/* { dg-final { scan-tree-dump-not {vector[^ ]* int vect__} "vect" } } */
 /* { dg-final { scan-tree-dump-times "vectorized 1 loop" 1 "vect" } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c 
b/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c
index 44d128477d2..0ac3f09cbdb 100644
--- a/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c
+++ b/gcc/testsuite/gcc.target/aarch64/sve/slp_6.c
@@ -36,11 +36,21 @@ vec_slp_##TYPE (TYPE *restrict a, TYPE *restrict b, int n)  
\
 
 TEST_ALL (VEC_PERM)
 
-/* These loops can't use SLP.  */
-/* { dg-final { scan-assembler-not {\tld1b\t} } } */
-/* { dg-final { scan-assembler-not {\tld1h\t} } } */
-/* { dg-final { scan-assembler-not {\tld1w\t} } } */
+/* When this test was written, these loops couldn't use SLP.  Since changes to
+   enable tail predication, SLP discovery of size 3 reduction groups instead
+   succeeds, producing predicated loads such as ld1b z25.b, p7/z, [x1] for
+   {b[0], b[1], b[2]}.  */
+
+/* { dg-final { scan-assembler {\tld1b\t} } } */
+/* { dg-final { scan-assembler {\tld1h\t} } } */
+/* { dg-final { scan-assembler {\tld1w\t} } } */
 /* { dg-final { scan-assembler-not {\tld1d\t} } } */
+
+/* { dg-final { scan-assembler {\tst1b\t} } } */
+/* { dg-final { scan-assembler {\tst1h\t} } } */
+/* { dg-final { scan-assembler {\tst1w\t} } } */
+/* { dg-final { scan-assembler-not {\tst1d\t} } } */
+
 /* { dg-final { scan-assembler {\tld3b\t} } } */
 /* { dg-final { scan-assembler {\tld3h\t} } } */
 /* { dg-final { scan-assembler {\tld3w\t} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c 
b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c
index 2f8ce6808a9..eea13c28e49 100644
--- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c
+++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_1.c
@@ -1,5 +1,5 @@
 /* { dg-do compile } */
-/* { dg-options "-O2 -ftree-slp-vectorize" } */
+/* { dg-options "-O2 -ftree-slp-vectorize 
--param=aarch64-autovec-preference=sve-only" } */
 
 /* Test that a group of stores of 8 elements derived from a horizontal
    reduction is vectorized by constructing a vector and storing it.
@@ -30,8 +30,8 @@ foo (svint8_t src0, svint8_t src1, svint8_t src2, svint8_t 
src3, svint8_t src4,
   s.h = svaddv_s8 (all, src7);
 }
 
-/* { dg-final { scan-assembler-times {\tins\tv[0-9]+\.b\[[0-9]+\], 
v[0-9]+\.b\[[0-9]+\]\n} 7 } } */
-/* { dg-final { scan-assembler-times {\tstr\td[0-9]+, } 1 } } */
+/* { dg-final { scan-assembler-times {\tinsr\tz[0-9]+\.h, h[0-9]+\n} 7 } } */
+/* { dg-final { scan-assembler-times {\tst1b\tz[0-9]+\.h, p[0-9]+, 
\[x[0-9]+\]\n} 1 } } */
 
 /* { dg-final { scan-assembler-not {\tstr\tb[0-9]+, } } } */
 /* { dg-final { scan-assembler-not {\tstrb\tw[0-9]+, } } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c 
b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c
index 6715118d7b0..2bf537e13e2 100644
--- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c
+++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_2.c
@@ -1,5 +1,5 @@
 /* { dg-do compile } */
-/* { dg-options "-O2 -ftree-slp-vectorize" } */
+/* { dg-options "-O2 -ftree-slp-vectorize 
--param=aarch64-autovec-preference=sve-only" } */
 
 /* Test that a group of stores of 8 elements derived from the results of calls
    to a function that has only vector parameters and returns a scalar result is
@@ -40,3 +40,5 @@ foo (svint8_t src0, svint8_t src1, svint8_t src2, svint8_t 
src3, svint8_t src4,
 
 /* { dg-final { scan-assembler-not {\tins\tv[0-9]+\.b\[[0-9]+\], w[0-9]+\n} } 
} */
 /* { dg-final { scan-assembler-not {\tstr\td[0-9]+, } } } */
+/* { dg-final { scan-assembler-not {\tfmov\th[0-9]+, h[0-9]+\n} } } */
+/* { dg-final { scan-assembler-not {\tinsr\tz[0-9]+\.b, w[0-9]+\n} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c 
b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c
index 8143d0050ad..ccadaccbcb4 100644
--- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c
+++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_3.c
@@ -1,7 +1,7 @@
 /* { dg-do compile } */
-/* { dg-options "-O2 -ftree-slp-vectorize" } */
+/* { dg-options "-O2 -ftree-slp-vectorize 
--param=aarch64-autovec-preference=sve-only" } */
 
-/* Test that a group of stores of 8 elements derived from a horizontal
+/* Test that a group of stores of 14 elements derived from a horizontal
    reduction is vectorized by constructing a vector and storing it
    even if the results of the reductions are narrowed.
    Since there are no GPR-to-SIMD register transfers, there is no
@@ -13,12 +13,14 @@
 
 struct S
 {
-  char a, b, c, d, e, f, g, h;
+  char a, b, c, d, e, f, g, h, i, j, k, l, m, n;
 } s;
 
 void
 foo (svint16_t src0, svint32_t src1, svint16_t src2, svint32_t src3,
-     svint32_t src4, svint16_t src5, svint32_t src6, svint16_t src7)
+     svint32_t src4, svint16_t src5, svint32_t src6, svint16_t src7,
+     svint16_t src8, svint32_t src9, svint16_t src10, svint32_t src11,
+     svint32_t src12, svint16_t src13)
 {
   svbool_t all16 = svptrue_b16 ();
   svbool_t all32 = svptrue_b32 ();
@@ -30,10 +32,16 @@ foo (svint16_t src0, svint32_t src1, svint16_t src2, 
svint32_t src3,
   s.f = svminv_s16 (all16, src5);
   s.g = svlastb_s32 (svptrue_pat_b32 (SV_VL1), src6);
   s.h = svaddv_s16 (all16, src7);
+  s.i = svmaxv_s16 (all16, src8);
+  s.j = svminv_s32 (all32, src9);
+  s.k = svlastb_s16 (svptrue_pat_b16 (SV_VL1), src10);
+  s.l = svaddv_s32 (all32, src11);
+  s.m = svmaxv_s32 (all32, src12);
+  s.n = svminv_s16 (all16, src13);
 }
 
-/* { dg-final { scan-assembler-times {\tins\tv[0-9]+\.b\[[0-9]+\], 
v[0-9]+\.b\[[0-9]+\]\n} 7 } } */
-/* { dg-final { scan-assembler-times {\tstr\td[0-9]+, } 1 } } */
+/* { dg-final { scan-assembler-times {\tinsr\tz[0-9]+\.b, b[0-9]+\n} 13 } } */
+/* { dg-final { scan-assembler-times {\tst1b\tz[0-9]+\.b, p[0-9]+, 
\[x[0-9]\]\n} 1 } } */
 
 /* { dg-final { scan-assembler-not {\tstr\tb[0-9]+, } } } */
 /* { dg-final { scan-assembler-not {\tstrb\tw[0-9]+, } } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c 
b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c
index 49f8114b64c..3d41af684a3 100644
--- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c
+++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_4.c
@@ -1,5 +1,5 @@
 /* { dg-do compile } */
-/* { dg-options "-O2 -ftree-slp-vectorize" } */
+/* { dg-options "-O2 -ftree-slp-vectorize 
--param=aarch64-autovec-preference=sve-only" } */
 
 /* Test that a group of stores of 8 elements derived from a horizontal
    reduction is not vectorized by constructing a vector and storing it
@@ -33,5 +33,6 @@ foo (svint16_t src0, svint8_t src1, svint16_t src2, svint8_t 
src3,
 /* { dg-final { scan-assembler-times {\tstp\tw[0-9]+, w[0-9]+,} 4 } } */
 
 /* { dg-final { scan-assembler-not {\tins\tv[0-9]+\.s\[[0-9]+\], w[0-9]+\n} } 
} */
-/* { dg-final { scan-assembler-not {\tfmov\ts[0-9]+, w[0-9]+\n} } }
+/* { dg-final { scan-assembler-not {\tfmov\ts[0-9]+, w[0-9]+\n} } } */
 /* { dg-final { scan-assembler-not {\tstp\tq[0-9]+, q[0-9]+,} } } */
+/* { dg-final { scan-assembler-not {\tinsr\tz[0-9]+.s, w[0-9]+\n} } } */
diff --git a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c 
b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c
index 983d6c69ebc..89e57406c0e 100644
--- a/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c
+++ b/gcc/testsuite/gcc.target/aarch64/sve/vec_construct_5.c
@@ -1,5 +1,5 @@
 /* { dg-do compile } */
-/* { dg-options "-O2 -ftree-slp-vectorize" } */
+/* { dg-options "-O2 -ftree-slp-vectorize 
--param=aarch64-autovec-preference=sve-only" } */
 
 /* Test that a group of stores of 8 elements derived from lane extractions is
    vectorized by constructing a vector and storing it.  Since there are no
@@ -30,8 +30,8 @@ foo (svint8_t src0, svint8_t src1, svint8_t src2, svint8_t 
src3, svint8_t src4,
   s.h = svlastb_s8 (p, src7);
 }
 
-/* { dg-final { scan-assembler-times {\tins\tv[0-9]+\.b\[[0-9]+\], 
v[0-9]+\.b\[[0-9]+\]\n} 7 } } */
-/* { dg-final { scan-assembler-times {\tstr\td[0-9]+, } 1 } } */
+/* { dg-final { scan-assembler-times {\tinsr\tz[0-9]+\.h, h[0-9]+\n} 7 } } */
+/* { dg-final { scan-assembler-times {\tst1b\tz[0-9]+\.h, p[0-9]+, 
\[x[0-9]+\]\n} 1 } } */
 
 /* { dg-final { scan-assembler-not {\tstr\tb[0-9]+, } } } */
 /* { dg-final { scan-assembler-not {\tstrb\tw[0-9]+, } } } */
diff --git a/gcc/tree-vect-slp.cc b/gcc/tree-vect-slp.cc
index e919282bd43..10913576bd4 100644
--- a/gcc/tree-vect-slp.cc
+++ b/gcc/tree-vect-slp.cc
@@ -1114,8 +1114,12 @@ vect_record_max_nunits (vec_info *vinfo, stmt_vec_info 
stmt_info,
     }
 
   /* If populating the vector type requires unrolling then fail
-     before adjusting *max_nunits for basic-block vectorization.  */
+     before adjusting *max_nunits for basic-block vectorization.
+     Allow group sizes that are indivisible by the vector length only if they
+     are known not to exceed the vector length.  We may be able to support such
+     cases by generating constant masks.  */
   if (is_a <bb_vec_info> (vinfo)
+      && maybe_gt (group_size, TYPE_VECTOR_SUBPARTS (vectype))
       && !multiple_p (group_size, TYPE_VECTOR_SUBPARTS (vectype)))
     {
       if (dump_enabled_p ())
@@ -1703,16 +1707,22 @@ vect_build_slp_tree_1 (vec_info *vinfo, unsigned char 
*swap,
 
   if (maybe_soft_fail)
     {
-      unsigned HOST_WIDE_INT const_nunits;
-      if (!TYPE_VECTOR_SUBPARTS
-           (soft_fail_nunits_vectype).is_constant (&const_nunits)
-         || const_nunits > group_size)
+      /* Use the known minimum number of subparts for VLA because we still need
+        to choose a splitting point although the choice is more arbitrary.  */
+      unsigned HOST_WIDE_INT const_nunits = constant_lower_bound (
+         TYPE_VECTOR_SUBPARTS (soft_fail_nunits_vectype));
+
+      if (const_nunits > group_size)
        matches[0] = false;
       else
        {
          /* With constant vector elements simulate a mismatch at the
             point we need to split.  */
+         gcc_assert ((const_nunits & (const_nunits - 1)) == 0);
          unsigned tail = group_size & (const_nunits - 1);
+         if (tail == 0)
+           tail = const_nunits;
+         gcc_assert (group_size >= tail);
          memset (&matches[group_size - tail], 0, sizeof (bool) * tail);
        }
       return false;
@@ -2452,13 +2462,21 @@ vect_build_slp_tree_2 (vec_info *vinfo, slp_tree node,
                  /* Check whether we can build the invariant.  If we can't
                     we never will be able to.  */
                  tree type = TREE_TYPE (chains[0][n].op);
-                 if (!GET_MODE_SIZE (vinfo->vector_mode).is_constant ()
-                     && (TREE_CODE (type) == BOOLEAN_TYPE
-                         || !can_duplicate_and_interleave_p (vinfo, group_size,
-                                                             type)))
+                 if (!GET_MODE_SIZE (vinfo->vector_mode).is_constant ())
                    {
-                     matches[0] = false;
-                     goto out;
+                     if (TREE_CODE (type) == BOOLEAN_TYPE)
+                       {
+                         matches[0] = false;
+                         goto out;
+                       }
+
+                     if (!is_a<bb_vec_info> (vinfo)
+                         && !can_duplicate_and_interleave_p (vinfo, group_size,
+                                                             type))
+                       {
+                         matches[0] = false;
+                         goto out;
+                       }
                    }
                }
              else if (dt != vect_internal_def)
@@ -2891,7 +2909,7 @@ out:
                    uniform_val = NULL_TREE;
                    break;
                  }
-             if (!uniform_val
+             if (!uniform_val && !is_a<bb_vec_info> (vinfo)
                  && !can_duplicate_and_interleave_p (vinfo,
                                                      oprnd_info->ops.length (),
                                                      TREE_TYPE (op0)))
@@ -5088,6 +5106,53 @@ vect_analyze_slp_reductions (loop_vec_info loop_vinfo,
   return true;
 }
 
+/* Update MIN_NUNITS to reflect the minimum number of subparts for all of the
+   vector types used by the SLP subgraph rooted at NODE.  VISITED is used to
+   avoid reevaluating any node in the subgraph; it thereby prevents infinite
+   recursion should a cycle be encountered. The value of MIN_NUNITS will only 
be
+   updated if any node in the subgraph has a vector type with a number of
+   subparts that is smaller than the passed-in value of MIN_NUNITS. Before
+   calling this function for the first time, initialize MIN_NUNITS to
+   UINT64_MAX.  */
+
+static void
+vect_update_slp_min_nunits_for_node (slp_tree node, poly_uint64 &min_nunits,
+                                    hash_set<slp_tree> &visited)
+{
+  if (!node || SLP_TREE_DEF_TYPE (node) != vect_internal_def)
+    return;
+
+  if (visited.add (node))
+    return;
+
+  for (slp_tree child : SLP_TREE_CHILDREN (node))
+    vect_update_slp_min_nunits_for_node (child, min_nunits, visited);
+
+  tree vectype = SLP_TREE_VECTYPE (node);
+  if (!vectype)
+    return;
+
+  /* All unit counts have the form vec_info::vector_size * X for some
+     rational X, therefore we know the values are ordered.  */
+  poly_uint64 nunits = TYPE_VECTOR_SUBPARTS (vectype);
+  min_nunits = known_eq (min_nunits, UINT64_MAX)
+                ? nunits
+                : ordered_min (min_nunits, nunits);
+}
+
+/* For NODE, return the minimum number of subparts for all of the vector
+   types used in the given SLP subgraph.  */
+
+static poly_uint64
+vect_slp_tree_min_nunits (slp_tree node)
+{
+  poly_uint64 min_nunits = UINT64_MAX;
+  hash_set<slp_tree> visited;
+  vect_update_slp_min_nunits_for_node (node, min_nunits, visited);
+  gcc_checking_assert (known_ne (min_nunits, UINT64_MAX));
+  return min_nunits;
+}
+
 /* Analyze an SLP instance starting from a group of grouped stores.  Call
    vect_build_slp_tree to build a tree of packed stmts if possible.
    Return FALSE if it's impossible to SLP any stmt in the group.  */
@@ -5157,8 +5222,8 @@ vect_analyze_slp_instance (vec_info *vinfo,
       poly_uint64 unrolling_factor
        = calculate_unrolling_factor (max_nunits, group_size);
 
-      if (maybe_ne (unrolling_factor, 1U)
-         && is_a <bb_vec_info> (vinfo))
+      if (maybe_ne (unrolling_factor, 1U) && is_a<bb_vec_info> (vinfo)
+         && !known_ge (vect_slp_tree_min_nunits (node), group_size))
        {
          unsigned HOST_WIDE_INT const_max_nunits;
          if (!max_nunits.is_constant (&const_max_nunits)
@@ -5244,8 +5309,10 @@ vect_analyze_slp_instance (vec_info *vinfo,
          tree vectype = get_vectype_for_scalar_type (vinfo, scalar_type,
                                                      1 << floor_log2 (i));
          unsigned HOST_WIDE_INT const_nunits;
+         poly_uint64 nunits = TYPE_VECTOR_SUBPARTS (vectype);
          if (vectype
-             && TYPE_VECTOR_SUBPARTS (vectype).is_constant (&const_nunits))
+             && (const_nunits = constant_lower_bound (nunits)) > 1
+             && (i % const_nunits) == 0)
            {
              /* Split into two groups at the first vector boundary.  */
              gcc_assert ((const_nunits & (const_nunits - 1)) == 0);
@@ -11867,7 +11934,21 @@ vectorizable_slp_permutation_1 (vec_info *vinfo, 
gimple_stmt_iterator *gsi,
       unpack_factor = 1;
     }
   unsigned olanes = unpack_factor * ncopies * SLP_TREE_LANES (node);
-  gcc_assert (repeating_p || multiple_p (olanes, nunits));
+
+  /* With fully-predicated BB-SLP, an external node's number of lanes can be
+     incompatible with the chosen vector width (e.g., lane packs of 3 with a
+     natural 2-lane vector type).  */
+  if (!repeating_p && !multiple_p (olanes, nunits))
+    {
+      if (dump_p)
+       dump_printf_loc (MSG_MISSED_OPTIMIZATION, vect_location,
+                        "unsupported permutation %p: vector type %T,"
+                        " nunits=" HOST_WIDE_INT_PRINT_UNSIGNED
+                        " ncopies=%" PRIu64 ", lanes=%u and unpack=%u\n",
+                        (void *) node, vectype, estimated_poly_value (nunits),
+                        ncopies, SLP_TREE_LANES (node), unpack_factor);
+      return -1;
+    }
 
   /* Compute the { { SLP operand, vector index}, lane } permutation sequence
      from the { SLP operand, scalar lane } permutation as recorded in the
diff --git a/gcc/tree-vect-stmts.cc b/gcc/tree-vect-stmts.cc
index ca3b2845713..fb79d683981 100644
--- a/gcc/tree-vect-stmts.cc
+++ b/gcc/tree-vect-stmts.cc
@@ -13821,9 +13821,11 @@ get_related_vectype_for_scalar_type (machine_mode 
prevailing_mode,
 /* Function get_vectype_for_scalar_type.
 
    Returns the vector type corresponding to SCALAR_TYPE as supported
-   by the target.  If GROUP_SIZE is nonzero and we're performing BB
-   vectorization, make sure that the number of elements in the vector
-   is no bigger than GROUP_SIZE.  */
+   by the target, or NULL_TREE if none.  If GROUP_SIZE is nonzero and we're
+   performing BB vectorization, ensure that the number of subparts in the
+   returned type is no bigger than GROUP_SIZE, or no smaller than GROUP_SIZE if
+   a vector mask or length can be used to limit the number of lanes loaded and
+   stored.  */
 
 tree
 get_vectype_for_scalar_type (vec_info *vinfo, tree scalar_type,
@@ -13848,11 +13850,23 @@ get_vectype_for_scalar_type (vec_info *vinfo, tree 
scalar_type,
   if (vectype)
     vinfo->used_vector_modes.add (TYPE_MODE (vectype));
 
-  /* If the natural choice of vector type doesn't satisfy GROUP_SIZE,
-     try again with an explicit number of elements.  */
-  if (vectype
-      && group_size
-      && maybe_ge (TYPE_VECTOR_SUBPARTS (vectype), group_size))
+  if (!vectype || !group_size)
+    return vectype;
+
+  /* The natural choice of vector type satisfies GROUP_SIZE if it can be used 
to
+     store the whole group without a mask or length limit (known_eq) or if it
+     can be used to carve the group into pieces (known_lt).  */
+  const poly_uint64 natural_nelts = TYPE_VECTOR_SUBPARTS (vectype);
+  if (known_le (natural_nelts, group_size))
+    return vectype;
+
+  /* We might be able to use a partial vector to handle a GROUP_SIZE that is 
not
+     satisfied by the natural choice of vector type.  Assume this is more
+     efficient than carving the group into pieces.  */
+  if (vect_get_load_store_partial_vector_style (vectype, true)
+       == vect_load_store_partial_vectors_none
+      || vect_get_load_store_partial_vector_style (vectype, false)
+          == vect_load_store_partial_vectors_none)
     {
       /* Start with the biggest number of units that fits within
         GROUP_SIZE and halve it until we find a valid vector type.
@@ -13874,8 +13888,32 @@ get_vectype_for_scalar_type (vec_info *vinfo, tree 
scalar_type,
          nunits /= 2;
        }
       while (nunits > 1 && !vectype);
+
+      return vectype;
+    }
+
+  /* Interoperability between modes requires one to be a constant multiple of
+     the other, so do not try to substitute a narrower fixed-length vector type
+     for a variable-length vector type.  Do not try to minimize the number of
+     elements in a variable-length type either because it can increase the size
+     of vector constants (e.g., if QImode values are stored in HImode).  */
+  if (!natural_nelts.is_constant ())
+    return vectype;
+
+  /* Try to minimize the number of subparts in the partial vector type.  Start
+     with the smallest number of units that contains GROUP_SIZE and double it
+     until we find a valid vector type.  */
+  unsigned int nunits = 1 << ceil_log2 (group_size);
+  do
+    {
+      vectype = get_related_vectype_for_scalar_type (vinfo->vector_mode,
+                                                    scalar_type, nunits);
+      nunits *= 2;
     }
+  while (known_le (nunits, natural_nelts) && !vectype);
 
+  /* We should have rediscovered the natural vector type if nothing else.  */
+  gcc_assert (vectype != NULL);
   return vectype;
 }
 
@@ -13895,10 +13933,10 @@ get_vectype_for_scalar_type (vec_info *vinfo, tree 
scalar_type, slp_tree node)
 /* Function get_mask_type_for_scalar_type.
 
    Returns the mask type corresponding to a result of comparison
-   of vectors of specified SCALAR_TYPE as supported by target.
-   If GROUP_SIZE is nonzero and we're performing BB vectorization,
-   make sure that the number of elements in the vector is no bigger
-   than GROUP_SIZE.  */
+   of vectors of specified SCALAR_TYPE as supported by target or NULL_TREE if
+   none.  If GROUP_SIZE is nonzero and we're performing BB vectorization, 
ensure
+   that the returned mask type is suitable for use with the vector type that
+   would be chosen by get_vectype_for_scalar_type.  */
 
 tree
 get_mask_type_for_scalar_type (vec_info *vinfo, tree scalar_type,
@@ -14168,7 +14206,33 @@ vect_maybe_update_slp_op_vectype (vec_info *vinfo, 
slp_tree op, tree vectype)
       && SLP_TREE_DEF_TYPE (op) == vect_external_def
       && SLP_TREE_LANES (op) > 1)
     return false;
-  (void) vinfo; /* FORNOW */
+
+  /* When the vectorizer falls back to building vector operands from scalars,
+     it can create SLP trees with external defs that have a number of lanes not
+     divisible by the number of subparts in a vector type naively inferred from
+     the scalar type.  Reject such types to avoid ICE when later computing the
+     prologue cost for invariant operands.  */
+  if (SLP_TREE_DEF_TYPE (op) == vect_external_def)
+    {
+      poly_uint64 vf = vinfo->vectorization_factor;
+
+      vf *= SLP_TREE_LANES (op);
+
+      if (maybe_lt (TYPE_VECTOR_SUBPARTS (vectype), vf)
+         && !multiple_p (vf, TYPE_VECTOR_SUBPARTS (vectype)))
+       {
+         if (dump_enabled_p ())
+           dump_printf_loc (MSG_MISSED_OPTIMIZATION, vect_location,
+                            "lanes=" HOST_WIDE_INT_PRINT_UNSIGNED
+                            " is not divisible by "
+                            "subparts=" HOST_WIDE_INT_PRINT_UNSIGNED ".\n",
+                            estimated_poly_value (vf),
+                            estimated_poly_value (
+                              TYPE_VECTOR_SUBPARTS (vectype)));
+         return false;
+       }
+    }
+
   SLP_TREE_VECTYPE (op) = vectype;
   return true;
 }
@@ -14853,9 +14917,10 @@ vect_gen_while_not (gimple_seq *seq, tree mask_type, 
tree start_index,
 
 /* Try to compute the vector types required to vectorize STMT_INFO,
    returning true on success and false if vectorization isn't possible.
-   If GROUP_SIZE is nonzero and we're performing BB vectorization,
-   take sure that the number of elements in the vectors is no bigger
-   than GROUP_SIZE.
+   If GROUP_SIZE is nonzero and we're performing BB vectorization, ensure that
+   the number of subparts in *STMT_VECTYPE_OUT is no bigger than GROUP_SIZE, or
+   no smaller than GROUP_SIZE if a vector mask or length can be used to limit
+   the number of lanes loaded and stored.
 
    On success:
 
-- 
2.43.0

Reply via email to