https://gcc.gnu.org/bugzilla/show_bug.cgi?id=117379

Avinash Jayakar <avinashd at gcc dot gnu.org> changed:

           What    |Removed                     |Added
----------------------------------------------------------------------------
                 CC|                            |rguenth at gcc dot gnu.org

--- Comment #6 from Avinash Jayakar <avinashd at gcc dot gnu.org> ---
(In reply to Richard Biener from comment #5)

> Yes, we currently are not set up to vectorize multiple pieces of the
> reduction separately (and we still don't handle the MINUS, also not if
> that were uniform).  Both is not difficult, it just requires some extra
> bookkeeping.
> 
> For the case above, rather than two .REDUC_PLUS I'd do
> 
>    vector_a = *a;
>    vector_b = *b;
>    sum = vector_a + vector_b;
>    .REDUC_PLUS (sum);
> 
> so handle it similar as to how we'd vectorize a 8 lane reduction.  The
> 'sum' would be a SLP node with a PLUS operation created by SLP reduction
> discovery when it discovers more than one (same size!) reduction part.
> Having two separate .REDUC_PLUS requires more surgery - like creating
> more than one SLP instance and using separate root stmts which would
> be an 'alternate' scalar input for the other SLP reduction instance.

Trying it with option 1, with one .REDUC_PLUS (sum).
As this improvement was added in
https://gcc.gnu.org/cgit/gcc/commit/?id=3adb3325954177939d59031055fc08a81ae21be3
In similar lines is it possible to recursively call vect_build_slp_instance for
the rest of the statements instead of pushing it to scalar remain. For e.g.,
If the n_matching array is
4 0 0 0 4 0 0 0 
for
_4 = MEM[(u32 *)a_10(D) + 12B];
_3 = MEM[(u32 *)a_10(D) + 8B]; 
_2 = MEM[(u32 *)a_10(D) + 4B];
_1 = *a_10(D);
 _8 = MEM[(u32 *)b_14(D) + 12B];
 _7 = MEM[(u32 *)b_14(D) + 8B];
 _5 = *b_14(D);
 _6 = MEM[(u32 *)b_14(D) + 4B];
try vect_build_slp_instance for both groups instead of largest (since they are
same size)?

Reply via email to