From: Kyrylo Tkachov <[email protected]>
vect_compute_single_scalar_iteration_cost weights inner-loop statements
by LOOP_VINFO_INNER_LOOP_COST_FACTOR through the count it passes to
record_stmt_cost. It used to record them with a vect_prologue location
so that vector_costs::adjust_cost_for_freq would not apply the same
factor again, and the comment there still said so.
r17-129-g7eb6d800c9cc ("Cost scalar into vect_body") switched the
location to vect_body, so for outer-loop vectorisation the scalar
iteration cost is now scaled by the factor squared while the vector body
cost is scaled once.
With the default factor of 50 that makes the scalar cost about 50x too
large, so the profitability denominator stays positive no matter what
the vector cost is and outer loops get vectorised that the cost model
should reject. Only outer-loop vectorisation is affected, since
stmt_in_inner_loop_p is false otherwise.
Skip the weighting in adjust_cost_for_freq when costing the scalar loop.
The vect_body location has to stay: scalar_costs->body_cost () must
remain non-zero for vect_get_known_peeling_cost. The explicit factor
has to stay too, because vect_estimate_min_profitable_iters re-uses
these entries with vect_prologue and vect_epilogue locations, where
adjust_cost_for_freq never applies.
For gcc.target/aarch64/pr121290_1.c the reported scalar iteration cost
goes from 12505 back to 255.
Fixes gcc.target/aarch64/pr121290_1.c, pr121290_2.c and pr121290_3.c.
Bootstrapped and regtested on aarch64-none-linux-gnu.
Ok for trunk?
Thanks,
Kyrill
gcc/ChangeLog:
* tree-vectorizer.cc (vector_costs::adjust_cost_for_freq): Do not
apply the inner loop cost factor when costing the scalar loop.
* tree-vect-loop.cc (vect_compute_single_scalar_iteration_cost):
Update comment.
Signed-off-by: Kyrylo Tkachov <[email protected]>
---
gcc/tree-vect-loop.cc | 8 ++++++--
gcc/tree-vectorizer.cc | 6 +++++-
2 files changed, 11 insertions(+), 3 deletions(-)
diff --git a/gcc/tree-vect-loop.cc b/gcc/tree-vect-loop.cc
index a9335ed68bb..d57508bd66c 100644
--- a/gcc/tree-vect-loop.cc
+++ b/gcc/tree-vect-loop.cc
@@ -1412,8 +1412,12 @@ vect_compute_single_scalar_iteration_cost (loop_vec_info
loop_vinfo)
else
kind = scalar_stmt;
- /* We are using vect_prologue here to avoid scaling twice
- by the inner loop factor. */
+ /* FACTOR already weights inner-loop statements by the inner
+ loop cost factor, so vector_costs::adjust_cost_for_freq must
+ not apply that weighting again for scalar costing. The cost
+ entries recorded here are also re-used by
+ vect_estimate_min_profitable_iters with a vect_prologue or
+ vect_epilogue location, where the weighting never applies. */
record_stmt_cost (&LOOP_VINFO_SCALAR_ITERATION_COST (loop_vinfo),
factor, kind, stmt_info, 0, vect_body);
}
diff --git a/gcc/tree-vectorizer.cc b/gcc/tree-vectorizer.cc
index 8c824e4ee33..43d22b0930e 100644
--- a/gcc/tree-vectorizer.cc
+++ b/gcc/tree-vectorizer.cc
@@ -1893,8 +1893,12 @@ vector_costs::adjust_cost_for_freq (stmt_vec_info
stmt_info,
{
/* Statements in an inner loop relative to the loop being
vectorized are weighted more heavily. The value here is
- arbitrary and could potentially be improved with analysis. */
+ arbitrary and could potentially be improved with analysis.
+ When costing the scalar loop the weighting is already applied by
+ vect_compute_single_scalar_iteration_cost through the count it
+ passes to record_stmt_cost, so do not apply it a second time. */
if (where == vect_body
+ && !m_costing_for_scalar
&& stmt_info
&& stmt_in_inner_loop_p (m_vinfo, stmt_info))
{
--
2.50.1 (Apple Git-155)