thinkharderdev commented on PR #25659:
URL: https://github.com/apache/datafusion/pull/25659#issuecomment-5995636366

   > ```
   > Suppose Q1 and Q2 have already been computed, and we cached their partial 
aggregate states:
   > 
   > Q1: select avg(v) from t1
   > 
   > Q2: select k, avg(v) from t2 group by k
   > 
   > Then we want to compute Q3 directly from Q1 and Q2's partial states, 
without recomputing the original inputs:
   > 
   > Q3: select avg(v) from (t1 union t2)
   > ```
   > 
   > If this is the main use case, could Q1 also use `GroupsAccumulator`? It 
seems doable with some refactor, and we can avoid a new contract to all 
aggregate function implementation.
   
   Not exactly this, but something conceptually similar (inspired by 
https://github.com/datafusion-contrib/datafusion-query-cache although 
ultimately quite different in implementation). We can cache and reuse partial 
aggregations to avoid costly scans. This requires adding a "cache key" to 
partial aggregates group by expressions. In the case of approx distinct with no 
group by (in the user query) we then end up with partial aggregation grouped by 
only the cache key and final aggregation with no group by expression which 
blows up. 
   
   > If this is the main use case, could Q1 also use `GroupsAccumulator`?
   
   This was our first thought as well but using the groups accumulator for 
approx distinct for a single group is a ~30% performance regression so its a 
non-starter.
   
   
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to