https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946

--- Comment #5 from rguenther at suse dot de <rguenther at suse dot de> ---
> Am 19.08.2026 um 16:35 schrieb ktkachov at gcc dot gnu.org 
> <[email protected]>:
> 
> https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946
> 
> --- Comment #4 from ktkachov at gcc dot gnu.org ---
> (In reply to [email protected] from comment #3)
>>> On Wed, 19 Aug 2026, ktkachov at gcc dot gnu.org wrote:
>>> 
>>> https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946
>>> 
>>> --- Comment #2 from ktkachov at gcc dot gnu.org ---
>>> (In reply to Richard Biener from comment #1)
>>>> So why does RTL if-conversion not produce the fcsel?  (and why do we have
>>>> such strange BB order)
>>>> 
>>>> The COND_EXPR allows the x86 cmov expander to pattern-match its FP MIN/MAX
>>>> operations which match IEEE semantics of m < a ? a : m
>>>> 
>>>> On trunk I do see fcsel being used on aarch64 just fine for your testcase,
>>>> so what "fixed" it there?
>>>> 
>>>> .L3:
>>>>        ldr     s31, [x1, x2, lsl 2]
>>>>        add     x2, x2, 1
>>>>        fabs    s31, s31
>>>>        fcmpe   s31, s0
>>>>        fcsel   s0, s31, s0, gt
>>>>        cmp     x0, x2
>>>>        bne     .L3
>>>>        ret
>>>> 
>>>> We expand from
>>>> 
>>>>  _4 = MEM[(const float *)x_9(D) + _20 * 4];
>>>>  a_10 = ABS_EXPR <_4>;
>>>>  _12 = a_10 > m_16;
>>>>  _11 = _12 ? a_10 : m_16;
>>> 
>>> In this case we do not want fcsel. GCC 13 kept the well-predicted branch
>>> whereas GCC 14 starting using fcsel unconditionally because COND_EXPR 
>>> expansion
>>> goes through the movcc optabs
>> 
>> How do you know the branch is well-predicted?  In general I'd
>> expect its probability to change during the iteration given
>> m grows assuming even distributed x[i].
> 
> In the full application we measured the misprediction rates with HW counters
> and did an A/B comparison with just that fcsel/branch decision changed to
> measure the speedup.
> In this reduced example I think the argument is that the mispredict happens
> whenever the max is updated. Therefore each time it's updated it becomes less
> likely to update again i.e. mispredict since the max is raised so there's 
> fewer
> values left in the domain that are greater than the new max. So for large
> enough n the mispredict rate should be dropping, whereas the fact that the
> running max is a loop recurrence means that a wide core suffers from the
> increased dependency chain of the fcsel.

That would then suggest any such MAX (like with fast-math) are problematic (but
required for vectorization).  So why does the movcc expander not turn this back
to a branch?  Likewise when expanding from max()?

> --
> You are receiving this mail because:
> You are on the CC list for the bug.

Reply via email to