On Mon, 10 Aug 2026, Tamar Christina wrote:

> > -----Original Message-----
> > From: Richard Biener <[email protected]>
> > Sent: 10 August 2026 10:51
> > To: Tamar Christina <[email protected]>
> > Cc: [email protected]; nd <[email protected]>; [email protected]
> > Subject: Re: [patch v3][vect]: fix loop exit profiles with early break 
> > [PR118407]
> > 
> > On Fri, 7 Aug 2026, Tamar Christina wrote:
> > 
> > > It's been a few years since the patches for PR118407 but have been stuck 
> > > on
> > a
> > > review.
> > >
> > > Those patches attempted to make an framework for supporting early break
> > profiles
> > > in not just the vectorizer but in other passes as well.
> > >
> > > This patch takes a different approach. We know that the vectorizer is the 
> > > one
> > > introducing the inaccurate profile information, so have it fix it up.  
> > > There are
> > > other places in the vectorizer where this happens for non-early break too.
> > >
> > > As an example this loop
> > >
> > > #ifndef N
> > > #define N 800
> > > #endif
> > > unsigned vect_a[N];
> > > unsigned vect_b[N];
> > >
> > > unsigned test4(unsigned x)
> > > {
> > >  unsigned ret = 0;
> > >  for (int i = 0; i < N; i++)
> > >  {
> > >    vect_b[i] = x + i;
> > >    if (vect_a[i]*2 != x)
> > >      break;
> > >    vect_a[i] = x;
> > >
> > >  }
> > >  return ret;
> > > }
> > >
> > > generates with -O3 -march=armv9-a the following loop:
> > >
> > > test4:
> > >         cntw    x1
> > >         cmp     w1, 799
> > >         bhi     .L7
> > >         adrp    x4, .LANCHOR0
> > >         add     x3, x4, :lo12:.LANCHOR0
> > >         ptrue   p7.b, all
> > >         add     x5, x3, 3200
> > >         mov     w1, 0
> > >         mov     z30.s, w0
> > >         mov     w6, 800
> > >         index   z29.s, #0, #1
> > >         b       .L4
> > > .L3:
> > >         st1w    z28.s, p7, [x5, x1, lsl 2]
> > >         st1w    z30.s, p7, [x3, x1, lsl 2]
> > >         incw    z29.s
> > >         incw    x1
> > >         whilelo p7.s, w1, w6
> > >         b.none  .L5
> > > .L4:
> > >         ld1w    z31.s, p7/z, [x3, x1, lsl 2]
> > >         add     z28.s, z29.s, z30.s
> > >         add     z31.s, z31.s, z31.s
> > >         cmpne   p15.s, p7/z, z31.s, z30.s
> > >         b.none  .L3
> > >         mov     w5, 800
> > >         mov     w2, w1
> > >         sub     w5, w5, w1
> > > .L2:
> > >
> > > notice the pointless indirect branch to L4, and the commonly taken branch 
> > > of
> > .L3
> > > has to do a branch back to the latch instead of a fall through.  The CFG
> > > becomes increasingly branchy as the number of exits increase.
> > >
> > > With this patch we generate:
> > >
> > > test4:
> > >         cntw    x1
> > >         cmp     w1, 799
> > >         bhi     .L7
> > >         adrp    x3, .LANCHOR0
> > >         add     x3, x3, :lo12:.LANCHOR0
> > >         ptrue   p7.b, all
> > >         add     x4, x3, 3200
> > >         mov     w1, 0
> > >         mov     z30.s, w0
> > >         mov     w5, 800
> > >         index   z29.s, #0, #1
> > >         .p2align 5,,15
> > > .L4:
> > >         ld1w    z31.s, p7/z, [x3, x1, lsl 2]
> > >         add     z28.s, z29.s, z30.s
> > >         add     z31.s, z31.s, z31.s
> > >         cmpne   p15.s, p7/z, z31.s, z30.s
> > >         b.any   .L12
> > >         st1w    z28.s, p7, [x4, x1, lsl 2]
> > >         st1w    z30.s, p7, [x3, x1, lsl 2]
> > >         incw    z29.s
> > >         incw    x1
> > >         whilelo p7.s, w1, w5
> > >         b.any   .L4
> > > .L5:
> > >         mov     w0, 0
> > >         ret
> > >
> > > which is a more natural layout because it now knows the exit is unlikely 
> > > since
> > > they have been scaled correctly.
> > >
> > > The scaling is done by:
> > >
> > > let P be the scalar probability that one iteration takes an early exit 
> > > edge.
> > > 1 - P is then the probabiliy that the exit is not taken.
> > >
> > > Since every vector iteration handles VF scalar iterations, the early exit
> > > branch can be taken when any of the lanes are true.
> > >
> > > We invert this and first calculate the probablity that you reach the 
> > > latch,
> > > i.e. the probability that the vector code does not exit is the probability
> > > Q that none of the lanes are true.
> > >
> > > Q = (1 - P) * (1 - P) * ... * (1 - P)
> > >   = (1 - P) ^ VF
> > >
> > > and so the probability that at least one lane breaks is 1 - Q,
> > > or rather 1 - (i - P) ^ VF.
> > >
> > > For example, if the scalar break probability is 10% and VF is 4:
> > >
> > > P = 0.1
> > > 1 - P = 0.9
> > > (1 - P) ^ VF = 0.9 ^ 4 = 0.6561
> > > 1 - (1 - p) ^ VF = 1 - 0.6561 = 0.3439
> > >
> > > So a scalar early-exit edge that is taken 10% of the time becomes a vector
> > > early-exit edge that is taken about 34.4% of the time, because each vector
> > > iteration gives the break condition four chances to trigger.
> > >
> > > For counted loops we need to exclude the main exit from the calculations
> > > as for those the main exit is counted and so does not have this per-lane
> > > probability.  For uncounted loops, all exits are based on vector compare,
> > > i.e. none of the exits are counted and we have a different loop control.
> > > That means for uncounted loops all exits should be using the per-lane
> > > probability and we should skip the counted exit adjustments in
> > > scale_profile_for_vect_loop.
> > >
> > > NOTE: This does not fully fix PR118407 but only fixes part of the issue.
> > >       there are other adjustments needed but this is the most glaring one.
> > >
> > > Bootstrapped Regtested on aarch64-none-linux-gnu,
> > > arm-none-linux-gnueabihf, x86_64-pc-linux-gnu
> > > -m32, -m64 and no issues.
> > >
> > > Any comments?
> > >
> > > Thanks,
> > > Tamar
> > >
> > > gcc/ChangeLog:
> > >
> > >   PR tree-optimization/118407
> > >   * tree-vect-stmts.cc (vect_update_early_break_profiles): New.
> > >   (scale_profile_for_vect_loop): Use it and pass loop_vinfo.
> > >   (vect_transform_loop): Pass loop_vinfo to
> > scale_profile_for_vect_loop.
> > >
> > > gcc/testsuite/ChangeLog:
> > >
> > >   PR tree-optimization/118407
> > >   * gcc.target/aarch64/sve/noeffect2.c: Update test.
> > >   * gcc.target/aarch64/sve/noeffect3.c: Likewise.
> > >
> > > ---
> > > diff --git a/gcc/testsuite/gcc.target/aarch64/noeffect2.c
> > b/gcc/testsuite/gcc.target/aarch64/noeffect2.c
> > > index
> > 08c531fb18c817b018776a88db4f62eb03b001cb..f0720862f9ced1a8308cb
> > edc1ecfa32d128e011b 100644
> > > --- a/gcc/testsuite/gcc.target/aarch64/noeffect2.c
> > > +++ b/gcc/testsuite/gcc.target/aarch64/noeffect2.c
> > > @@ -15,10 +15,11 @@ int b[N] = {0};
> > >  **       cmgt    v[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s
> > >  **       umaxp   v[0-9]+\.4s, v[0-9]+\.4s, v[0-9]+\.4s
> > >  **       fmov    x[0-9]+, d[0-9]+
> > > -**       cbz     x[0-9]+, \.L[0-9]+
> > > -**       mov     w0, 1
> > > +**       cbn?z   x[0-9]+, \.L[0-9]+
> > > +**       ...
> > > +**       mov     w0, [01]
> > >  **       ret
> > > -**       mov     w0, 0
> > > +**       mov     w0, [01]
> > >  **       ret
> > >  */
> > >  __attribute__ ((noipa, noinline))
> > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/noeffect2.c
> > b/gcc/testsuite/gcc.target/aarch64/sve/noeffect2.c
> > > index
> > 5bb1badde3c4e07d06ba68b2c73ba591e031f611..04b7eb993e36fbae27ff3
> > 849957b2e7e25595c9c 100644
> > > --- a/gcc/testsuite/gcc.target/aarch64/sve/noeffect2.c
> > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/noeffect2.c
> > > @@ -9,16 +9,21 @@ int b[N] = {0};
> > >  /*
> > >  ** foo:
> > >  **       ...
> > > -**       whilelo p[0-9]+\.s, w[0-9]+, w[0-9]+
> > > +**       ptrue   p[0-9]+\.b, all
> > >  **       ...
> > >  **       ld1w    z[0-9]+\.s, p[0-9]+/z, \[x[0-9]+, x[0-9]+, lsl 2\]
> > >  **       ld1w    z[0-9]+\.s, p[0-9]+/z, \[x[0-9]+, x[0-9]+, lsl 2\]
> > >  **       cmpgt   p[0-9]+\.s, p[0-9]+/z, z[0-9]+\.s, z[0-9]+\.s
> > > -**       b\.none \.L[0-9]+
> > > -**       mov     w0, 1
> > > -**       ret
> > > +**       b\.any  \.L[0-9]+
> > > +**       ...
> > > +**       incw    x[0-9]+
> > > +**       whilelo p[0-9]+\.s, w[0-9]+, w[0-9]+
> > > +**       b\.any  \.L[0-9]+
> > >  **       mov     w0, 0
> > >  **       ret
> > > +**       ...
> > > +**       mov     w0, 1
> > > +**       ret
> > >  */
> > >  __attribute__ ((noipa, noinline))
> > >  int foo (void)
> > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/noeffect3.c
> > b/gcc/testsuite/gcc.target/aarch64/sve/noeffect3.c
> > > index
> > c5f81bb0d303a070b482ec40d2921f6e6236ce67..8c20f39d2a467e636cfec
> > 71acacb2da654a49d67 100644
> > > --- a/gcc/testsuite/gcc.target/aarch64/sve/noeffect3.c
> > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/noeffect3.c
> > > @@ -9,11 +9,14 @@ int b[N] = {0};
> > >  /*
> > >  ** foo:
> > >  **       ...
> > > -**       whilelo p[0-9]+\.s, w[0-9]+, w[0-9]+
> > > +**       ptrue   p[0-9]+\.b, all
> > >  **       ...
> > >  **       ld1w    z[0-9]+\.s, p[0-9]+/z, \[x[0-9]+, x[0-9]+, lsl 2\]
> > >  **       ld1w    z[0-9]+\.s, p[0-9]+/z, \[x[0-9]+, x[0-9]+, lsl 2\]
> > >  **       cmpgt   p[0-9]+\.s, p[0-9]+/z, z[0-9]+\.s, z[0-9]+\.s
> > > +**       b\.any  \.L[0-9]+
> > > +**       ...
> > > +**       whilelo p[0-9]+\.s, w[0-9]+, w[0-9]+
> > >  **       ...
> > >  **       ldr     w[0-9]+, \[x[0-9]+, x[0-9]+, lsl 2\]
> > >  **       ldr     w[0-9]+, \[x[0-9]+, x[0-9]+, lsl 2\]
> > > diff --git a/gcc/tree-vect-loop.cc b/gcc/tree-vect-loop.cc
> > > index
> > 31de29c036be5da2a94173e98dccc40a092099be..fbd410544122b2e2fde73
> > 764573d78da3293df0a 100644
> > > --- a/gcc/tree-vect-loop.cc
> > > +++ b/gcc/tree-vect-loop.cc
> > > @@ -10890,14 +10890,88 @@ vect_gen_loop_len_mask (loop_vec_info
> > loop_vinfo, gimple_stmt_iterator *gsi,
> > >    return len_mask;
> > >  }
> > >
> > > +/* Update the exit profile counts scaling them from scalar to vector 
> > > counts.
> > > +
> > > +   The calculations are as follows.
> > > +
> > > +   let P be the scalar probability that one iteration takes an early 
> > > exit edge.
> > > +   1 - P is then the probabiliy that the exit is not taken.
> > > +
> > > +   Since every vector iteration handles VF scalar iterations, the early 
> > > exit
> > > +   branch can be taken when any of the lanes are true.
> > > +
> > > +   We invert this and first calculate the probablity that you reach the 
> > > latch,
> > > +   i.e. the probability that the vector code does not exit is the 
> > > probability
> > > +   Q that none of the lanes are true.
> > > +
> > > +   Q = (1 - P) * (1 - P) * ... * (1 - P)
> > > +     = (1 - P) ^ VF
> > > +
> > > +   and so the probability that at least one lane breaks is 1 - Q,
> > > +   or rather 1 - (i - P) ^ VF.
> > > +
> > > +   For example, if the scalar break probability is 10% and VF is 4:
> > > +
> > > +   P = 0.1
> > > +   1 - P = 0.9
> > > +   (1 - P) ^ VF = 0.9 ^ 4 = 0.6561
> > > +   1 - (1 - p) ^ VF = 1 - 0.6561 = 0.3439
> > > +
> > > +   So a scalar early-exit edge that is taken 10% of the time becomes a 
> > > vector
> > > +   early-exit edge that is taken about 34.4% of the time, because each 
> > > vector
> > > +   iteration gives the break condition four chances to trigger.  */
> > > +
> > > +static void
> > > +vect_update_early_break_profiles (loop_vec_info loop_vinfo)
> > > +{
> > > +  if (!LOOP_VINFO_EARLY_BREAKS (loop_vinfo))
> > > +    return;
> > > +
> > > +  class loop *loop = LOOP_VINFO_LOOP (loop_vinfo);
> > > +  unsigned int vf = vect_vf_for_cost (loop_vinfo);
> > > +
> > > +  auto_vec<std::pair<edge, profile_probability>, 8> updates;
> > > +  for (edge e : get_loop_exit_edges (loop))
> > > +    {
> > > +      if ((e == LOOP_VINFO_MAIN_EXIT (loop_vinfo)
> > > +    && !LOOP_VINFO_NITERS_UNCOUNTED_P (loop_vinfo))
> > > +   || !e->probability.initialized_p ())
> > > + continue;
> > > +
> > > +      /* For any early break exits, the probablity of exiting are based 
> > > on if
> > > +  any lane is true, and so are tied to VF since you have VF chances.
> > > +  As such scale the scalar profile by VF to the the probability for the
> > > +  vector edges using 1 - (1 - p) ^ VF.  */
> > > +      profile_probability probability
> > > + = profile_probability::always () - e->probability.invert ().pow (vf);
> > > +      updates.safe_push ({e, probability});
> > > +    }
> > > +
> > > +  /* First update the edge counts.  */
> > > +  for (auto update : updates)
> > > +    set_edge_probability_and_rescale_others (update.first, 
> > > update.second);
> > 
> > So this still adusts only edge probabilities.  We do expect the counts
> > on exit->dest to remain as they are, right?
> > 
> 
> Yes, since updating that would make it inconsistent with that block's 
> children.
> So it seemed more logical that counts for the blocks after vect don't change.
> 
> This is especially true when we no longer need the scalar loop to perform
> side effects.
> 
> > > +
> > > +  /* And then only update the in loop BBs, i.e. the fall through block 
> > > for
> > > +     the early exits.  */
> > > +  for (auto update : updates)
> > > +    for (edge e : update.first->src->succs)
> > > +      if (!loop_exit_edge_p (loop, e) && single_pred_p (e->dest))
> > > + e->dest->count = e->count ();
> > 
> > Does this then update the in-loop counts correctly, given the order
> > of exits and thus the order of 'updates' isn't necessarily
> > dominator order?
> 
> I don't think the order matters since unlike edge counts which are
> computed, BB counts aren't.  So by now the edge counts are all
> correct and so the order we set the BB counts in don't matter.

Yes, but for

   /\
     \
      A
     /\
       \
        B

when you update the count of B, A->B e->count depends on the count of A,
so you have to update A first?

That said, the loop->header count will have to be updated as well,
from the last exit count?

> > 
> > I also hope Honza will chime in to tell us whether this is the canonical
> > way to update a profile.
> 
> Hopefully.  I also realized I accidentally dropped him from CC of the new
> Version so added now.
> 
> Thanks,
> Tamar
> > 
> > > +}
> > > +
> > >  /* Scale profiling counters by estimation for LOOP which is vectorized
> > >     by factor VF.
> > >     If FLAT is true, the loop we started with had unrealistically flat
> > >     profile.  */
> > >
> > >  static void
> > > -scale_profile_for_vect_loop (class loop *loop, edge exit_e, unsigned vf, 
> > > bool
> > flat)
> > > +scale_profile_for_vect_loop (loop_vec_info loop_vinfo,
> > > +                      class loop *loop, edge exit_e, unsigned vf,
> > > +                      bool flat)
> > >  {
> > > +  /* First scale any early exits.  */
> > > +  vect_update_early_break_profiles (loop_vinfo);
> > > +
> > >    /* For flat profiles do not scale down proportionally by VF and only
> > >       cap by known iteration count bounds.  */
> > >    if (flat)
> > > @@ -10927,14 +11001,19 @@ scale_profile_for_vect_loop (class loop
> > *loop, edge exit_e, unsigned vf, bool fl
> > >        vf /= 2;
> > >      }
> > >
> > > -  if (entry_count.nonzero_p ())
> > > -    set_edge_probability_and_rescale_others
> > > +  if (!LOOP_VINFO_NITERS_UNCOUNTED_P (loop_vinfo))
> > > +    {
> > > +      if (entry_count.nonzero_p ())
> > > + set_edge_probability_and_rescale_others
> > >       (exit_e,
> > >        entry_count.probability_in (loop->header->count / vf));
> > > -  /* Avoid producing very large exit probability when we do not have
> > > -     sensible profile.  */
> > > -  else if (exit_e->probability < profile_probability::always () / (vf * 
> > > 2))
> > > -    set_edge_probability_and_rescale_others (exit_e, exit_e->probability 
> > > *
> > vf);
> > > +      /* Avoid producing very large exit probability when we do not have
> > > +  sensible profile.  */
> > > +      else if (exit_e->probability < profile_probability::always () / 
> > > (vf * 2))
> > > + set_edge_probability_and_rescale_others (exit_e,
> > > +                                          exit_e->probability * vf);
> > > +    }
> > > +
> > >    loop->latch->count = single_pred_edge (loop->latch)->count ();
> > >
> > >    scale_loop_profile (loop, profile_probability::always () / vf,
> > > @@ -11555,8 +11634,9 @@ vect_transform_loop (loop_vec_info
> > loop_vinfo, gimple *loop_vectorized_call)
> > >                     assumed_vf) - 1
> > >    : wi::udiv_floor (loop->nb_iterations_estimate + bias_for_assumed,
> > >                      assumed_vf) - 1);
> > > -  scale_profile_for_vect_loop (loop, LOOP_VINFO_MAIN_EXIT (loop_vinfo),
> > > -                        assumed_vf, flat);
> > > +  scale_profile_for_vect_loop (loop_vinfo, loop,
> > > +                        LOOP_VINFO_MAIN_EXIT (loop_vinfo),
> > assumed_vf,
> > > +                        flat);
> > >
> > >    if (dump_enabled_p ())
> > >      {
> > >
> > >
> > >
> > 
> > --
> > Richard Biener <[email protected]>
> > SUSE Software Solutions Germany GmbH,
> > Frankenstrasse 146, 90461 Nuernberg, Germany;
> > GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG
> > Nuernberg)
> 

-- 
Richard Biener <[email protected]>
SUSE Software Solutions Germany GmbH,
Frankenstrasse 146, 90461 Nuernberg, Germany;
GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG Nuernberg)

Reply via email to