https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946

--- Comment #9 from ktkachov at gcc dot gnu.org ---
(In reply to [email protected] from comment #8)
> On Wed, 19 Aug 2026, ktkachov at gcc dot gnu.org wrote:
> 
> > https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946
> > 
> > --- Comment #6 from ktkachov at gcc dot gnu.org ---
> > (In reply to [email protected] from comment #5)
> > > > Am 19.08.2026 um 16:35 schrieb ktkachov at gcc dot gnu.org 
> > > > <[email protected]>:
> > > > 
> > > > https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946
> > > > 
> > > > --- Comment #4 from ktkachov at gcc dot gnu.org ---
> > > > (In reply to [email protected] from comment #3)
> > > >>> On Wed, 19 Aug 2026, ktkachov at gcc dot gnu.org wrote:
> > > >>> 
> > > >>> https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126946
> > > >>> 
> > > >>> --- Comment #2 from ktkachov at gcc dot gnu.org ---
> > > >>> (In reply to Richard Biener from comment #1)
> > > >>>> So why does RTL if-conversion not produce the fcsel?  (and why do we 
> > > >>>> have
> > > >>>> such strange BB order)
> > > >>>> 
> > > >>>> The COND_EXPR allows the x86 cmov expander to pattern-match its FP 
> > > >>>> MIN/MAX
> > > >>>> operations which match IEEE semantics of m < a ? a : m
> > > >>>> 
> > > >>>> On trunk I do see fcsel being used on aarch64 just fine for your 
> > > >>>> testcase,
> > > >>>> so what "fixed" it there?
> > > >>>> 
> > > >>>> .L3:
> > > >>>>        ldr     s31, [x1, x2, lsl 2]
> > > >>>>        add     x2, x2, 1
> > > >>>>        fabs    s31, s31
> > > >>>>        fcmpe   s31, s0
> > > >>>>        fcsel   s0, s31, s0, gt
> > > >>>>        cmp     x0, x2
> > > >>>>        bne     .L3
> > > >>>>        ret
> > > >>>> 
> > > >>>> We expand from
> > > >>>> 
> > > >>>>  _4 = MEM[(const float *)x_9(D) + _20 * 4];
> > > >>>>  a_10 = ABS_EXPR <_4>;
> > > >>>>  _12 = a_10 > m_16;
> > > >>>>  _11 = _12 ? a_10 : m_16;
> > > >>> 
> > > >>> In this case we do not want fcsel. GCC 13 kept the well-predicted 
> > > >>> branch
> > > >>> whereas GCC 14 starting using fcsel unconditionally because COND_EXPR 
> > > >>> expansion
> > > >>> goes through the movcc optabs
> > > >> 
> > > >> How do you know the branch is well-predicted?  In general I'd
> > > >> expect its probability to change during the iteration given
> > > >> m grows assuming even distributed x[i].
> > > > 
> > > > In the full application we measured the misprediction rates with HW 
> > > > counters
> > > > and did an A/B comparison with just that fcsel/branch decision changed 
> > > > to
> > > > measure the speedup.
> > > > In this reduced example I think the argument is that the mispredict 
> > > > happens
> > > > whenever the max is updated. Therefore each time it's updated it 
> > > > becomes less
> > > > likely to update again i.e. mispredict since the max is raised so 
> > > > there's fewer
> > > > values left in the domain that are greater than the new max. So for 
> > > > large
> > > > enough n the mispredict rate should be dropping, whereas the fact that 
> > > > the
> > > > running max is a loop recurrence means that a wide core suffers from the
> > > > increased dependency chain of the fcsel.
> > > 
> > > That would then suggest any such MAX (like with fast-math) are problematic
> > > (but required for vectorization).  So why does the movcc expander not turn
> > > this back to a branch?  Likewise when expanding from max()?
> > 
> > Yes, with my measurements (on Grace) the scalar fmax is slower than a branch
> > when it's a loop-carried recurrence (about 2x slower vs the 4x slower 
> > version
> > of the fcsel)
> > 
> > I think there is logic during expansion for choosing between branch and 
> > cmov in
> > expand_expr_real_2 e.g. there's expand_cond_expr_using_cmove, but it doesn't
> > take loop recurrence into account
> 
> IMO all this breaking of recurrences for (small) loops should be done
> on RTL rather than on GIMPLE given it's really micro-architectural
> details that influence things.  For example on x86 some uarchs have
> special loop uop machinery that are quite constrained with respect
> to code layout.

Perhaps, I can certainly believe there is a uarch-dependent component. In this
case it seems that for low-iteration loops the CSEL is better but for long
iterations the cbranch is better as the branch predictor trains on the pattern.
so to distinguish between the two we'd need some niter information. Is that
available/deducible at RTL?

Reply via email to