Rachelint commented on PR #25724:
URL: https://github.com/apache/datafusion/pull/25724#issuecomment-6044777687

   > @Rachelint,
   > 
   > only aggregations whose group keys and aggregate state are all fixed-width 
(primitive or boolean) use buckets or flush their partial table. String, binary 
and nested state keep a single table. BucketedAggregation::supports_state 
replaces the old check, which excluded only nested state.
   > 
   > what do you think on this direction. If looks good, I will start PRs for 
review and landing. If not, I will find other direction to improve aggregation
   
   I agree that we should start by applying the bucket approach only to 
fixed-width types. Despite trying several ideas over the past few days, I 
haven’t yet found a way to reduce the routing overhead enough to make the 
approach broadly beneficial...
   
   I also have a few thoughts we could explore as follow-ups:
   
   - **Perhaps proactive Partial flushing could have its own issue.** It looks 
promising based on my experiments in 
[[#26116](https://github.com/apache/datafusion/pull/26116)](https://github.com/apache/datafusion/pull/26116),
 and I’d be happy to help investigate further.
   
   - **Would it make sense to leave compaction for a follow-up?** It introduces 
some complexity, and I’ve seen cases where it hurts performance while offering 
limited benefit for very high-cardinality workloads like Q32. Starting just 
with the bucket approach only might make the initial implementation easier to 
evaluate.
   
   - **I’m also curious how this would interact with 
[[#24704](https://github.com/apache/datafusion/issues/24704)](https://github.com/apache/datafusion/issues/24704).**
 If proactive flushing keeps Partial output batches small, and Final 
aggregation materializes its output in 64 or more buckets, could that address 
much of the peak-memory concern? I may be missing something here—what remaining 
scenarios would #24704 help with?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to