On Thu, 30 Jul 2026, Tamar Christina wrote:

> > -----Original Message-----
> > From: Richard Biener <[email protected]>
> > Sent: 30 July 2026 12:00
> > To: Tamar Christina <[email protected]>
> > Cc: [email protected]; nd <[email protected]>; [email protected]
> > Subject: RE: [patch][vect]: block versioning for VLA when alignment can 
> > never
> > be reached [PR126369]
> > 
> > On Thu, 30 Jul 2026, Tamar Christina wrote:
> > 
> > > > -----Original Message-----
> > > > From: Richard Biener <[email protected]>
> > > > Sent: 30 July 2026 10:06
> > > > To: Tamar Christina <[email protected]>
> > > > Cc: [email protected]; nd <[email protected]>; [email protected]
> > > > Subject: Re: [patch][vect]: block versioning for VLA when alignment can
> > never
> > > > be reached [PR126369]
> > > >
> > > > On Wed, 29 Jul 2026, Tamar Christina wrote:
> > > >
> > > > > The loop
> > > > >
> > > > > char b[100];
> > > > > int c(int a) {
> > > > >   unsigned d = 0;
> > > > >   for (; __builtin_expect (d < a, 1); ++d)
> > > > >     {
> > > > >       if (b[0] + b[d + 1])
> > > > >         return 0;
> > > > >     }
> > > > >  return 1;
> > > > > }
> > > > >
> > > > > compiled with -march=armv8-a+sve -O3 vectorizes with SVE, however to
> > > > vectorize
> > > > > since we don't support peeling for alignment for
> > > > > LOOP_VINFO_EARLY_BREAKS_VECT_PEELED it forces versioning in order
> > to
> > > > vectorize.
> > > > >
> > > > > However versioning will always fail since we know the base access is
> > > > misaligned.
> > > > > i.e. the access to (b+1)[d] is always misaligned.
> > > > >
> > > > > This is happening because VLA will always return misalignment unknown
> > > > from
> > > > > dr_misalignment because this requires to known the size of the vector.
> > > > >
> > > > > To fix this I added a new helper
> > > > vector_versioning_alignment_not_reachable_p
> > > > > that checks if versioning could ever reach alignment.
> > > > >
> > > > > While doing this some unexpected tests started failing. It turns out 
> > > > > that
> > the
> > > > > DR_SCALAR_KNOWN_BOUNDS check which checks to see that even if
> > > > misaligned but
> > > > > all scalar accesses are in bounds of a known fixed size array then 
> > > > > we're OK
> > > > and
> > > > > just need to force masking.
> > > > >
> > > > > The loop it was placed in would exit early after the first misaligned 
> > > > > access
> > and
> > > > > so when you have more than one data access in the loop it wouldn't
> > mark
> > > > the
> > > > > other accesses as safe to speculate.
> > > > >
> > > > > This moves it to its own loop.
> > > > >
> > > > > This removes all the unreachable loops from the testsuite.  I do have
> > patches
> > > > > to enable peeling for alignment for
> > > > LOOP_VINFO_EARLY_BREAKS_VECT_PEELED but I am
> > > > > not yet happy with them (and need to redo prolog peeling LCSSA
> > handling)
> > > > so will
> > > > > submit them later.
> > > > >
> > > > > Bootstrapped Regtested on aarch64-none-linux-gnu,
> > > > > arm-none-linux-gnueabihf, x86_64-pc-linux-gnu
> > > > > -m32, -m64 and no issues.
> > > > >
> > > > > Any comments?
> > > > >
> > > > > Thanks,
> > > > > Tamar
> > > > >
> > > > > gcc/ChangeLog:
> > > > >
> > > > >       * tree-vect-data-refs.cc
> > > > (vector_versioning_alignment_not_reachable_p):
> > > > >       New.
> > > > >       (vect_enhance_data_refs_alignment): Use it.
> > > > >
> > > > > gcc/testsuite/ChangeLog:
> > > > >
> > > > >       * gcc.dg/vect/vect-early-break-no-epilog_7.c: Fix off by 1 
> > > > > error in test
> > > > >       which wasn't noticed before because loop is unreachable.
> > > > >       * gcc.target/aarch64/sve/peeled.c: Mark peeling as xfail till 
> > > > > PFA for
> > > > >       PEELED loops.
> > > > >       * gcc.target/aarch64/sve/peeled1.c: Likewise.
> > > > >
> > > > > ---
> > > > > diff --git a/gcc/testsuite/gcc.dg/vect/vect-early-break-no-epilog_7.c
> > > > b/gcc/testsuite/gcc.dg/vect/vect-early-break-no-epilog_7.c
> > > > > index
> > > >
> > 6d94312f4a60cb5e519c2eb5ee239732a625045c..6986210f67d4ad29c7e4
> > > > 224178ff679d3b8815cd 100644
> > > > > --- a/gcc/testsuite/gcc.dg/vect/vect-early-break-no-epilog_7.c
> > > > > +++ b/gcc/testsuite/gcc.dg/vect/vect-early-break-no-epilog_7.c
> > > > > @@ -10,7 +10,7 @@ int b[N] = {0};
> > > > >
> > > > >  int foo (void)
> > > > >  {
> > > > > -  for (int i = 0; i < (N / 2); i += 2)
> > > > > +  for (int i = 0; i < ((N / 2) - 1); i += 2)
> > > > >      {
> > > > >        if (a[i] > b[i])
> > > > >       return 1;
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled.c
> > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled.c
> > > > > index
> > > >
> > f40ffc8f0bcde40aeb754887ed6d5047e9806bc7..1de20ba79bc47cf7e64464
> > > > a9a1e076884dce0892 100644
> > > > > --- a/gcc/testsuite/gcc.target/aarch64/sve/peeled.c
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled.c
> > > > > @@ -15,6 +15,6 @@ c (int a)
> > > > >    return 1;
> > > > >  }
> > > > >
> > > > > -/* { dg-final { scan-assembler-times {\twhilelo\t} 2 } } */
> > > > > +/* { dg-final { scan-assembler-times {\twhilelo\t} 2 { xfail *-*-* } 
> > > > > } } */
> > > > >  /* { dg-final { scan-assembler-times {\tptest\t} 0 } } */
> > > > > -/* { dg-final { scan-assembler {\tld1b\tz[0-9]+\.s, p[0-9]+/z,} } } 
> > > > > */
> > > > > +/* { dg-final { scan-assembler {\tld1b\tz[0-9]+\.s, p[0-9]+/z,} { 
> > > > > xfail *-*-*
> > } }
> > > > } */
> > > > > diff --git a/gcc/testsuite/gcc.target/aarch64/sve/peeled1.c
> > > > b/gcc/testsuite/gcc.target/aarch64/sve/peeled1.c
> > > > > index
> > > >
> > 92e5bca44c48cefa20e829510a8914e4852145ed..79ca2d8e95ade0bd6510
> > > > a8eff26083855028e48a 100644
> > > > > --- a/gcc/testsuite/gcc.target/aarch64/sve/peeled1.c
> > > > > +++ b/gcc/testsuite/gcc.target/aarch64/sve/peeled1.c
> > > > > @@ -19,6 +19,6 @@ c (int a)
> > > > >    return 1;
> > > > >  }
> > > > >
> > > > > -/* { dg-final { scan-assembler-times {\twhilelo\t} 2 } } */
> > > > > -/* { dg-final { scan-assembler-times {\tptest\t} 2 } } */
> > > > > -/* { dg-final { scan-assembler-times {\tld1b\tz[0-9]+\.h, 
> > > > > p[0-9]+/z,} 2 } }
> > */
> > > > > +/* { dg-final { scan-assembler-times {\twhilelo\t} 2 { xfail *-*-* } 
> > > > > } } */
> > > > > +/* { dg-final { scan-assembler-times {\tptest\t} 2 { xfail *-*-* } } 
> > > > > } */
> > > > > +/* { dg-final { scan-assembler-times {\tld1b\tz[0-9]+\.h, 
> > > > > p[0-9]+/z,} 2 {
> > xfail
> > > > *-*-* } } } */
> > > > > diff --git a/gcc/tree-vect-data-refs.cc b/gcc/tree-vect-data-refs.cc
> > > > > index
> > > >
> > 0e0754769ae441a4d5a505636aa0bb6dbc09ad86..8f7d00ed221be64b639
> > > > 2385dab15a6fa2698061e 100644
> > > > > --- a/gcc/tree-vect-data-refs.cc
> > > > > +++ b/gcc/tree-vect-data-refs.cc
> > > > > @@ -1912,6 +1912,56 @@ vector_alignment_reachable_p (dr_vec_info
> > > > *dr_info, poly_uint64 vf)
> > > > >    return true;
> > > > >  }
> > > > >
> > > > > +/* Return true if DR_INFO is known not to be aligned to its target
> > > > alignment.
> > > > > +
> > > > > +  This handles the case where the target alignment is VLA/poly, so
> > > > > +  dr_misalignment returns misalignment unknown since it doesn't know
> > the
> > > > vector
> > > > > +  length, but the access is known to be misaligned misaligned wrt all
> > possible
> > > > > +  target alignments.  i.e. the base + offset is misaligned vs the 
> > > > > base
> > object.
> > > > > +  */
> > > > > +
> > > > > +static bool
> > > > > +vector_versioning_alignment_not_reachable_p (vec_info *vinfo,
> > > >
> > > > As we have a vector_alignment_reachable_p please avoid a
> > > > similar but negated predicate here, so
> > > > vector_versioning_alignment_reachable_p please.
> > > >
> > > > > +                                          dr_vec_info *dr_info,
> > > > > +                                          poly_int64 offset)
> > > > > +{
> > > > > +  if (DR_TARGET_ALIGNMENT (dr_info).is_constant ())
> > > > > +    return false;
> > > >
> > > > But that's not true, so better check in the caller if the point
> > > > is we've already checked for constant target alignment (did we?).
> > > >
> > >
> > > Yes, for constant vector sizes this would just double reject it,
> > > But this function was intended to handle exact the cases that
> > > dr_misalignment does not handle.
> > >
> > > dr_misalignment does
> > >
> > >   unsigned HOST_WIDE_INT target_alignment_c;
> > >   if (!dr_info->target_alignment.is_constant (&target_alignment_c)
> > >       || !known_misalignment (misalignment, target_alignment_c, 
> > > &misalign))
> > >     return DR_MISALIGNMENT_UNKNOWN;
> > >
> > > which gives you the unknown misalignment because it requires the check
> > > that the misalignment isn't an exact_div of the vector size otherwise 
> > > it's not
> > > actually misaligned.
> > >
> > > The comment on the function says the returned misalignment is WRT to the
> > > VECTYPE.  But if the vectype isn't known It didn't seem correct to modify
> > this.
> > >
> > > When I did, quite a few other alignments calculation went wrong, 
> > > particularly
> > > because without speculative access none of this matters for VLA..
> > >
> > > Then I thought maybe vect_compute_data_ref_alignment should be
> > adjusted
> > > instead, because after all in this case it's the dataref that's 
> > > misaligned.
> > >
> > > And set SET_DR_MISALIGNMENT (dr_info, const_misalignment);
> > >
> > > But that function does
> > >
> > >   unsigned HOST_WIDE_INT vect_align_c;
> > >   if (!vector_alignment.is_constant (&vect_align_c))
> > >     return;
> > >
> > > because it relies on a constant vect_align_c.
> > >
> > > So that's how I got here..
> > >
> > > > > +
> > > > > +  unsigned HOST_WIDE_INT factor
> > > > > +    = known_alignment (DR_TARGET_ALIGNMENT (dr_info));
> > > > > +
> > > > > +  if (factor <= 1)
> > > > > +    return false;
> > > > > +
> > > > > +  HOST_WIDE_INT diff = 0;
> > > > > +  if (STMT_VINFO_GROUPED_ACCESS (dr_info->stmt))
> > > > > +    {
> > > > > +      dr_vec_info *first_dr
> > > > > +     = STMT_VINFO_DR_INFO (DR_GROUP_FIRST_ELEMENT (dr_info-
> > > > >stmt));
> > > > > +
> > > > > +      diff = (TREE_INT_CST_LOW (DR_INIT (dr_info->dr))
> > > > > +           - TREE_INT_CST_LOW (DR_INIT (first_dr->dr)));
> > > > > +      gcc_assert (diff >= 0);
> > > >
> > > > We only get here for the first element because of
> > > > vect_relevant_for_alignment_p.  It doesn't make much sense to
> > > > ask for any sub-element - there'll always be one that cannot
> > > > be aligned.
> > >
> > > I modelled this after dr_misalignment and I thought it was trying to catch
> > > The case where all group loads are misaligned by the same amount, in which
> > > case it's only the first access that matters, but to also handle the part 
> > > where
> > > the first access is aligned and the second entry in the group isn't.
> > >
> > > The rest of the code only looks at the first access so this difference is
> > recorded
> > > this way. At least that's how I understood what dr_misaligment is doing...
> > >
> > > The point is that for VLA we have all the accesses marked with 
> > > misalignment
> > unknown.
> > >
> > > >
> > > > > +
> > > > > +      dr_info = first_dr;
> > > > > +    }
> > > > > +
> > > > > +  innermost_loop_behavior *drb = vect_dr_behavior (vinfo, dr_info);
> > > > > +
> > > > > +  if (drb->offset_alignment < factor)
> > > > > +    return false;
> > > >
> > > > should that have been return true?
> > >
> > > Yeah I inverted the function name at the last minute and forgot invert the
> > > values... doh..
> > >
> > > >
> > > > > +
> > > > > +  poly_offset_int total_misalignment = drb->base_misalignment;
> > > > > +  total_misalignment += wi::to_poly_offset (drb->init);
> > > > > +  total_misalignment += diff;
> > > > > +  total_misalignment += offset;
> > > > > +
> > > >
> > > > So I think the point is why
> > > >
> > > >           int misalignment;
> > > >           if ((misalignment = dr_misalignment (dr_info, vectype, off)) 
> > > > ==
> > > > 0)
> > > >             continue;
> > > >
> > > > does not compute an appropriate misalignment for VLA vectors?  The
> > > > alignment requirement of those is only dependent on the minimal
> > > > number of elements, right?  That is, I expected dr_misalignment (...)
> > > > to compute a misalignment based on DR_INIT.  Why does that not work?
> > > >
> > >
> > > dr_misalignment doesn't support POLY at all. It's punting in all cases
> > > because technically the vect size of a vector doesn't have to be a
> > > multiple of the minimum size (for SVE it is sure) but for general POLY
> > > it just has to be a power of 2.
> > 
> > Huh, it's N + <factor> * M, no?  So you say it dr_misalignment only
> > "fails" at
> > 
> >   unsigned HOST_WIDE_INT target_alignment_c;
> >   if (!dr_info->target_alignment.is_constant (&target_alignment_c)
> >       || !known_misalignment (misalignment, target_alignment_c,
> > &misalign))
> >     return DR_MISALIGNMENT_UNKNOWN;
> > 
> > ?  target_alignment should be constant, right?  otherwise we couldn't
> > be properly aligning vectors?  Oh, for the early-break case where we
> > adjust alignment we have to have actual HW size alignment?  Yes, so
> > that we cannot really improve.
> 
> You don't get there. For VLA dr_info->misalignment is always set to
> UNKNOWN.
> 
> So you exit at 
> 
>   int misalign = dr_info->misalignment;
>   gcc_assert (misalign != DR_MISALIGNMENT_UNINITIALIZED);
>   if (misalign == DR_MISALIGNMENT_UNKNOWN)
>     return misalign;
> 
> But also target_alignment is a POLY and not a constant. i.e. for a
> VNx16QI vector the target_alignment is:
> 
> >>> p dr_info->misalignment
> $2 = -1
> >>> p dr_info->target_alignment
> $3 = {
>   coeffs = {[0] = 16, [1] = 16}
> }
> 
> Or in other words, the vector size.
> 
> My understanding of this, and I could be wrong here (and you and Richard S
> has a similar discussion about this on a PR where you said the alignment was
> wrong and he didn't agree (I'm paraphrasing from memory)) is that for VLA,
> at least SVE the actual alignment is element size, it's expected to always use
> masking.
> 
> For non-early break it means that the alignment always happens at runtime
> If we can't be sure that all accesses are aligned.  Without the speculative
> access we don't really need to align in most cases since the partial final 
> iteration
> means you'd never access what the scalar code didn't.
> 
> So I think that's why it's punting here.  But again I could be and probably am
> wrong but this is likely why we almost never need alignment peeling without
> early break.

So yes, for early break we need to ensure spculative accesses do not trap.
But for non-early break I expect target_alignment to be the constant
element alignment, right?  Or at most AdvSIMD-size constant alignment,
even for VLA SVE?

That said ...

> Thanks,
> Tamar
> 
> > 
> > I'll note that just looking at DR_INIT doesn't make alias versioning
> > never hit - it's the combination of having a DECL and DR_INIT which
> > makes the access always misaligned.  That's true for non-VLA as well,
> > and I'm not sure we catch this.
> > 
> > Consider the testcase accessing the data via an incoming pointer,
> > also offset by 1, there can very well be the case that versioning
> > helps.

This point still holds, so your patch would be too conservative.

Richard.

> > 
> > > The code punts because known_misalignment needs a constant value.
> > >
> > > and because of this I don't' think we can conclusively conclude the 
> > > offset will
> > > cause a misalignment.  Note that for normal VLA this isn't a problem as 
> > > it will
> > then
> > > use peeling for alignment using masks, so it's conservative here.  The
> > intention here
> > > is to only disable versioning
> > >
> > > But for early break we don't yet support that but that is my plan.
> > >
> > > >
> > > > > +  poly_offset_int misalign;
> > > > > +  return (known_misalignment (total_misalignment, factor, &misalign)
> > > > > +       && maybe_ne (misalign, 0));
> > > > > +}
> > > > > +
> > > > >
> > > > >  /* Calculate the cost of the memory access represented by DR_INFO.  
> > > > > */
> > > > >
> > > > > @@ -2455,6 +2505,30 @@ vect_enhance_data_refs_alignment
> > > > (loop_vec_info loop_vinfo)
> > > > >       }
> > > > >      }
> > > > >
> > > > > +  /* See if we can relax the flags on speculative reads for early 
> > > > > break.  Do
> > > > > +     this outside of the other loops below because they can exit 
> > > > > early
> > leading
> > > > > +     to the flag not being cleared for known in bounds cases.  */
> > > > > +  poly_uint64 vf = LOOP_VINFO_VECT_FACTOR (loop_vinfo);
> > > > > +  if (LOOP_VINFO_EARLY_BREAKS (loop_vinfo))
> > > > > +    for (auto dr : datarefs)
> > > > > +      {
> > > > > +     dr_vec_info *dr_info = loop_vinfo->lookup_dr (dr);
> > > > > +     if (!vect_relevant_for_alignment_p (dr_info))
> > > > > +       continue;
> > > > > +
> > > > > +     stmt_vec_info stmt_info = dr_info->stmt;
> > > > > +
> > > > > +     /* With variable VF, unsafe speculative read can be avoided for 
> > > > > known
> > > > > +        inbounds DRs as long as partial vectors are used.  */
> > > > > +     if (!vf.is_constant ()
> > > > > +         && dr_safe_speculative_read_required (stmt_info)
> > > > > +         && DR_SCALAR_KNOWN_BOUNDS (dr_info))
> > > > > +       {
> > > > > +         dr_set_safe_speculative_read_required (stmt_info, false);
> > > > > +         LOOP_VINFO_MUST_USE_PARTIAL_VECTORS_P (loop_vinfo) = true;
> > > > > +       }
> > > > > +      }
> > > > > +
> > > >
> > > > This hunk (and the related below) looks OK, can you split out and
> > > > push separately?
> > >
> > > Will do.
> > >
> > > Thanks,
> > > Tamar
> > >
> > > >
> > > > >    /* While cost model enhancements are expected in the future, the 
> > > > > high
> > > > level
> > > > >       view of the code at this time is as follows:
> > > > >
> > > > > @@ -2495,7 +2569,6 @@ vect_enhance_data_refs_alignment
> > > > (loop_vec_info loop_vinfo)
> > > > >       - The cost of peeling (the extra runtime checks, the increase
> > > > >         in code size).  */
> > > > >
> > > > > -  poly_uint64 vf = LOOP_VINFO_VECT_FACTOR (loop_vinfo);
> > > > >    FOR_EACH_VEC_ELT (datarefs, i, dr)
> > > > >      {
> > > > >        dr_vec_info *dr_info = loop_vinfo->lookup_dr (dr);
> > > > > @@ -2505,16 +2578,6 @@ vect_enhance_data_refs_alignment
> > > > (loop_vec_info loop_vinfo)
> > > > >        stmt_vec_info stmt_info = dr_info->stmt;
> > > > >        tree vectype = STMT_VINFO_VECTYPE (stmt_info);
> > > > >
> > > > > -      /* With variable VF, unsafe speculative read can be avoided 
> > > > > for known
> > > > > -      inbounds DRs as long as partial vectors are used.  */
> > > > > -      if (!vf.is_constant ()
> > > > > -       && dr_safe_speculative_read_required (stmt_info)
> > > > > -       && DR_SCALAR_KNOWN_BOUNDS (dr_info))
> > > > > -     {
> > > > > -       dr_set_safe_speculative_read_required (stmt_info, false);
> > > > > -       LOOP_VINFO_MUST_USE_PARTIAL_VECTORS_P (loop_vinfo) = true;
> > > > > -     }
> > > > > -
> > > > >        do_peeling = vector_alignment_reachable_p (dr_info, vf);
> > > > >        if (do_peeling)
> > > > >          {
> > > > > @@ -3002,6 +3065,26 @@ vect_enhance_data_refs_alignment
> > > > (loop_vec_info loop_vinfo)
> > > > >                    break;
> > > > >                  }
> > > > >
> > > > > +           /* For VLA we generally can't know the exact misalignment 
> > > > > wrt to
> > > > > +              the vector size because we don't know the vector size. 
> > > > >  That
> > > > > +              means misalignment will always be UNKNOWN.  We do know
> > > > the
> > > > > +              vector size must be a power of two so there are some 
> > > > > base
> > > > > +              pointers we know would never be aligned.  The check 
> > > > > below is
> > > > > +              only valid for versioning alone.  For peeling we don't 
> > > > > know yet
> > > > > +              statically whether we can each alignment.  */
> > > > > +           if (!try_peeling_with_versioning
> > > > > +               && misalignment == DR_MISALIGNMENT_UNKNOWN
> > > > > +               && vector_versioning_alignment_not_reachable_p
> > > > (loop_vinfo,
> > > > > +                                                               
> > > > > dr_info, off))
> > > > > +             {
> > > > > +               if (dump_enabled_p ())
> > > > > +                 dump_printf_loc (MSG_MISSED_OPTIMIZATION,
> > > > vect_location,
> > > > > +                                  "alignment versioning will never "
> > > > > +                                  "succeed.\n");
> > > > > +               do_versioning = false;
> > > > > +               break;
> > > > > +             }
> > > > > +
> > > > >             /* Forcing alignment in the first iteration is no good if
> > > > >                we don't keep it across iterations.  For now, just 
> > > > > disable
> > > > >                versioning in this case.
> > > > >
> > > > >
> > > > >
> > > >
> > > > --
> > > > Richard Biener <[email protected]>
> > > > SUSE Software Solutions Germany GmbH,
> > > > Frankenstrasse 146, 90461 Nuernberg, Germany;
> > > > GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG
> > > > Nuernberg)
> > >
> > 
> > --
> > Richard Biener <[email protected]>
> > SUSE Software Solutions Germany GmbH,
> > Frankenstrasse 146, 90461 Nuernberg, Germany;
> > GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG
> > Nuernberg)
> 

-- 
Richard Biener <[email protected]>
SUSE Software Solutions Germany GmbH,
Frankenstrasse 146, 90461 Nuernberg, Germany;
GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG Nuernberg)

Reply via email to