Aleksandr Efimov has posted comments on this change. ( http://gerrit.cloudera.org:8080/24883 )
Change subject: IMPALA-13052: Estimate reservoir sample memory in aggregations ...................................................................... Patch Set 1: (1 comment) http://gerrit.cloudera.org:8080/#/c/24883/1/fe/src/main/java/org/apache/impala/planner/AggregationNode.java File fe/src/main/java/org/apache/impala/planner/AggregationNode.java: http://gerrit.cloudera.org:8080/#/c/24883/1/fe/src/main/java/org/apache/impala/planner/AggregationNode.java@88 PS1, Line 88: // aggregation spill (IMPALA-3304). > Would it be hard to do it within reservation? Not in this change: it needs a backend change. Reserving the worst case costs 1.3MB per call and group (2.3MB for DECIMAL), while a group with one row holds about 4.5KB. The Jira query would need about 650GB per instance for one call, so the final aggregation would spill most of its input. Tracking the actual size is the fix Tim described in IMPALA-3304: let Allocate() go through, then spill, or pass rows through in a preaggregation, once control returns to the aggregator. Partition::Spill() already frees these states, so the missing part is that trigger. I'd keep it separate. This change at least lets admission control see the memory. -- To view, visit http://gerrit.cloudera.org:8080/24883 To unsubscribe, visit http://gerrit.cloudera.org:8080/settings Gerrit-Project: Impala-ASF Gerrit-Branch: master Gerrit-MessageType: comment Gerrit-Change-Id: Id6eec3daf7fea9b027ff0be9c95c6ee5a3cf323d Gerrit-Change-Number: 24883 Gerrit-PatchSet: 1 Gerrit-Owner: Aleksandr Efimov <[email protected]> Gerrit-Reviewer: Aleksandr Efimov <[email protected]> Gerrit-Reviewer: Csaba Ringhofer <[email protected]> Gerrit-Reviewer: Impala Public Jenkins <[email protected]> Gerrit-Reviewer: Riza Suminto <[email protected]> Gerrit-Comment-Date: Mon, 21 Sep 2026 16:58:56 +0000 Gerrit-HasComments: Yes
