thinkharderdev commented on PR #25659: URL: https://github.com/apache/datafusion/pull/25659#issuecomment-5995636366
> ``` > Suppose Q1 and Q2 have already been computed, and we cached their partial aggregate states: > > Q1: select avg(v) from t1 > > Q2: select k, avg(v) from t2 group by k > > Then we want to compute Q3 directly from Q1 and Q2's partial states, without recomputing the original inputs: > > Q3: select avg(v) from (t1 union t2) > ``` > > If this is the main use case, could Q1 also use `GroupsAccumulator`? It seems doable with some refactor, and we can avoid a new contract to all aggregate function implementation. Not exactly this, but something conceptually similar (inspired by https://github.com/datafusion-contrib/datafusion-query-cache although ultimately quite different in implementation). We can cache and reuse partial aggregations to avoid costly scans. This requires adding a "cache key" to partial aggregates group by expressions. In the case of approx distinct with no group by (in the user query) we then end up with partial aggregation grouped by only the cache key and final aggregation with no group by expression which blows up. > If this is the main use case, could Q1 also use `GroupsAccumulator`? This was our first thought as well but using the groups accumulator for approx distinct for a single group is a ~30% performance regression so its a non-starter. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
