Aleksandr Efimov has uploaded this change for review. ( http://gerrit.cloudera.org:8080/24883
Change subject: IMPALA-13052: Estimate reservoir sample memory in aggregations ...................................................................... IMPALA-13052: Estimate reservoir sample memory in aggregations sample(), appx_median() and histogram() keep a ReservoirSampleState per group, allocated by FunctionContext::Allocate() outside the tuple and the reservation, but the planner counted only the 12-byte StringValue slot. The state takes about 4.5KB per call and group even for one row, most of it the mt19937_64 generator, and up to 1.3MB (2.3MB for DECIMAL) at 20000 samples. For the grouping query in the Jira the merge aggregation was estimated at 18MB and used 2.19GB; it is now estimated at 2.27GB. AggregationNode now adds this memory, mirroring the state and its FreePool and MemPool allocations; static asserts in aggregate-functions-ir.cc fail the build if the state or sample sizes change. It takes the rows per group from the first phase so that merges are covered. Running short of this memory never makes an aggregation spill (IMPALA-3304), so it is added after the caps that assume spilling. On lineitem the estimate is within 1% of the measured memory at 4, 600 and 60K rows per group. A streaming preaggregation keeps the planner's usual group count, bounded by PREAGG_BYTES_LIMIT when set, which errs high: 1.5M groups per instance for the Jira query, where each instance holds about 500K. The backend's limit on hash table growth (STREAMING_HT_MIN_REDUCTION) is not mirrored, because it depends on how the input of one instance reduces. The planner expects a reduction of 1.33 both for the Jira query, which reduced fourfold and kept all its groups, and for a COUNT(DISTINCT) over unique rows, whose preaggregation stopped at 98,304 groups. Execution and the row size of the intermediate tuple are unchanged; admission control sees larger estimates for queries with these functions. Testing: - New reservoir-sample-agg.test: the Jira examples, groups at 20000 samples, COUNT(DISTINCT), MEM_ESTIMATE_SCALE_FOR_SPILLING_OPERATOR and PREAGG_BYTES_LIMIT. Reverting each part of the estimate separately fails the cases that part covers. - TpcdsPlannerTest and nine aggregation planner test files produce the same plans. With PLANNER=CALCITE the new estimates are the same. Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d Assisted-by: claude-opus-5 (Claude Code) --- M be/src/exprs/aggregate-functions-ir.cc M fe/src/main/java/org/apache/impala/analysis/FunctionCallExpr.java M fe/src/main/java/org/apache/impala/planner/AggregationNode.java M fe/src/test/java/org/apache/impala/planner/PlannerTest.java A testdata/workloads/functional-planner/queries/PlannerTest/reservoir-sample-agg.test 5 files changed, 618 insertions(+), 5 deletions(-) git pull ssh://gerrit.cloudera.org:29418/Impala-ASF refs/changes/83/24883/1 -- To view, visit http://gerrit.cloudera.org:8080/24883 To unsubscribe, visit http://gerrit.cloudera.org:8080/settings Gerrit-Project: Impala-ASF Gerrit-Branch: master Gerrit-MessageType: newchange Gerrit-Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d Gerrit-Change-Number: 24883 Gerrit-PatchSet: 1 Gerrit-Owner: Aleksandr Efimov <[email protected]>
