Richard Sandiford <[email protected]> writes:
> Jeffrey Law <[email protected]> writes:
>> On 7/20/2026 6:31 AM, Richard Biener wrote:
>>> The following adds simplification of a scalar mode vec_select wrapping
>>> a vec_concat.  combine sees
>>>
>>> (insn 7 4 8 2 (set (reg:V2DF 106)
>>>          (vec_concat:V2DF (plus:DF (vec_select:DF (reg/v:V2DF 101 [ p ])
>>>                      (parallel [
>>>                              (const_int 0 [0])
>>>                          ]))
>>>                  (vec_select:DF (reg/v:V2DF 101 [ p ])
>>>                      (parallel [
>>>                              (const_int 1 [0x1])
>>>                          ])))
>>>              (plus:DF (vec_select:DF (reg/v:V2DF 101 [ p ])
>>>                      (parallel [
>>>                              (const_int 0 [0])
>>>                          ]))
>>>                  (vec_select:DF (reg/v:V2DF 101 [ p ])
>>>                      (parallel [
>>>                              (const_int 1 [0x1])
>>>                          ]))))) "t.c":5:11 3234 {*sse3_haddv2df3}
>>>       (expr_list:REG_DEAD (reg/v:V2DF 101 [ p ])
>>>          (nil)))
>>> (insn 8 7 9 2 (set (reg:DF 105)
>>>          (vec_select:DF (reg:V2DF 106)
>>>              (parallel [
>>>                      (const_int 0 [0])
>>>                  ]))) "t.c":5:11 7297 {sse2_storelpd}
>>>       (expr_list:REG_DEAD (reg:V2DF 106)
>>>          (nil)))
>>>
>>> and this should simplify and combine to
>>>
>>> (set (reg:DF 105)
>>>      (plus:DF (vec_select:DF (reg:V2DF 109 [ p ])
>>>              (parallel [
>>>                      (const_int 0 [0])
>>>                  ]))
>>>          (vec_select:DF (reg:V2DF 109 [ p ])
>>>              (parallel [
>>>                      (const_int 1 [0x1])
>>>                  ]))))
>>>
>>> which x86 can recognize (but rejects due to consting, see PR126328).
>>> There'll be test coverage in gcc.target/i386/pr54400.c
>>>
>>> Bootstrapped on x86_64-unknown-linux-gnu, testing in progress.
>>>
>>> The code snipped is stripped down from the vector vec_select case
>>> but with the side_effects_p arguments fixed(?)
>>>
>>> OK?
>>>
>>> Thanks,
>>> Richard.
>>>
>>>     * simplify-rtx.cc (simplify_context::simplify_binary_operation_1):
>>>     Simplify scalar vec_select of vec_concat.
>>> ---
>>>   gcc/simplify-rtx.cc | 18 ++++++++++++++++++
>>>   1 file changed, 18 insertions(+)
>>>
>>> diff --git a/gcc/simplify-rtx.cc b/gcc/simplify-rtx.cc
>>> index 882a11c5760..612413bf035 100644
>>> --- a/gcc/simplify-rtx.cc
>>> +++ b/gcc/simplify-rtx.cc
>>> @@ -5148,6 +5148,7 @@ simplify_ashift:
>>>          nested VEC_SELECT expressions.  When input operand is a memory
>>>          operand, this operation can be simplified to a simple scalar
>>>          load from an offsetted memory address.  */
>>> +     int l0, l1;
>>>       int n_elts;
>>>       if (GET_CODE (trueop0) == VEC_SELECT
>>>           && (GET_MODE_NUNITS (GET_MODE (XEXP (trueop0, 0)))
>>> @@ -5210,6 +5211,23 @@ simplify_ashift:
>>>                                 tmp_op, gen_rtx_PARALLEL (VOIDmode, vec));
>>>           return tmp;
>>>         }
>>> +     /* If we select one half of a vec_concat, return that.  */
>>> +     else if (GET_CODE (trueop0) == VEC_CONCAT
>>> +              && (GET_MODE_NUNITS (GET_MODE (XEXP (trueop0, 0)))
>>> +                  .is_constant (&l0))
>>> +              && (GET_MODE_NUNITS (GET_MODE (XEXP (trueop0, 1)))
>>> +                  .is_constant (&l1)))
>> You never use l1, though I guess what you're really testing is that it 
>> is a constant, you don't need to know the precise constant?
>
> Yeah, agreed that this check looks redundant.  Also, AFAIK, VEC_CONCAT
> requires both operands to have the same mode.
>
> But the is_constant seems unnecessary anyway.  It isn't required for:
>
>> +        {
>> +          rtx subop0 = XEXP (trueop0, 0);
>> +          rtx subop1 = XEXP (trueop0, 1);
>> +          machine_mode mode0 = GET_MODE (subop0);
>> +          machine_mode mode1 = GET_MODE (subop1);
>> +          int i0 = INTVAL (XVECEXP (trueop1, 0, 0));
>> +          if (i0 == 0 && mode == mode0 && !side_effects_p (subop1))
>> +            return subop0;
>
> this first simplification, which is valid and potentially useful for VLA
> as well.  And the first check in:

Sorry, scratch the bit about VLA.  This will only trigger for 2-element
vectors.

>> +          if (i0 == l0 && mode == mode1 && !side_effects_p (subop0))
>> +            return subop1;
>
> could be handled with known_eq.
>
> Thanks,
> Richard

Reply via email to