github-actions[bot] commented on code in PR #68651:
URL: https://github.com/apache/doris/pull/68651#discussion_r4140092031
##########
fe/fe-core/src/main/java/org/apache/doris/nereids/cost/CostModel.java:
##########
@@ -382,10 +386,17 @@ public Cost visitPhysicalHashAggregate(
inputStatistics.getRowCount() / beNumber, 0);
} else {
int factor = aggregate.getGroupByExpressions().isEmpty() ? 1 :
beNumber;
- // global
+ double rowCost = inputStatistics.getRowCount() / factor;
+ // Bucketed fusion discount: when the one-phase GLOBAL
INPUT_TO_RESULT
+ // aggregate is eligible for translator fusion (correctness +
data-volume
+ // gates are enforced by ChildrenPropertiesRegulator), apply a
discount
+ // to prefer this path over two-phase aggregation.
+ if (aggregate.getAggMode() == AggMode.INPUT_TO_RESULT
Review Comment:
[P2] Apply the bucketed cost discount only when this aggregate can actually
fuse. For `SELECT b, COUNT(DISTINCT a), SUM(v) FROM t GROUP BY b`, the lower
GLOBAL `INPUT_TO_RESULT` dedup aggregate has `SUM(v)` in `INPUT_TO_BUFFER`
mode. This check halves its estimated row cost, but
`containsPartialAggFunction` rejects fusion later, so the plan still exchanges
raw rows. On a large table with few `(a,b)` groups, that bias can select the
raw-row exchange over a LOCAL dedup plan that exchanges far fewer rows. Share
the translator's complete fusion eligibility with costing (or omit the discount
for partial aggregate outputs).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]