Aleksandr Efimov has posted comments on this change. ( 
http://gerrit.cloudera.org:8080/24883 )

Change subject: IMPALA-13052: Estimate reservoir sample memory in aggregations
......................................................................


Patch Set 1:

(1 comment)

http://gerrit.cloudera.org:8080/#/c/24883/1/fe/src/main/java/org/apache/impala/planner/AggregationNode.java
File fe/src/main/java/org/apache/impala/planner/AggregationNode.java:

http://gerrit.cloudera.org:8080/#/c/24883/1/fe/src/main/java/org/apache/impala/planner/AggregationNode.java@88
PS1, Line 88:   // aggregation spill (IMPALA-3304).
> Would it be hard to do it within reservation?
Not in this change: it needs a backend change. Reserving the worst case costs 
1.3MB per call and group (2.3MB for DECIMAL), while a group with one row holds 
about 4.5KB. The Jira query would need about 650GB per instance for one call, 
so the final aggregation would spill most of its input.

Tracking the actual size is the fix Tim described in IMPALA-3304: let 
Allocate() go through, then spill, or pass rows through in a preaggregation, 
once control returns to the aggregator. Partition::Spill() already frees these 
states, so the missing part is that trigger. I'd keep it separate. This change 
at least lets admission control see the memory.



-- 
To view, visit http://gerrit.cloudera.org:8080/24883
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings

Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: comment
Gerrit-Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d
Gerrit-Change-Number: 24883
Gerrit-PatchSet: 1
Gerrit-Owner: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Csaba Ringhofer <[email protected]>
Gerrit-Reviewer: Impala Public Jenkins <[email protected]>
Gerrit-Reviewer: Riza Suminto <[email protected]>
Gerrit-Comment-Date: Mon, 21 Sep 2026 16:58:56 +0000
Gerrit-HasComments: Yes

Reply via email to