Jesus Camacho Rodriguez created HIVE-23530:

             Summary: Use SQL functions instead of compute_stats UDAF to 
compute column statistics
                 Key: HIVE-23530
             Project: Hive
          Issue Type: Improvement
          Components: Statistics
            Reporter: Jesus Camacho Rodriguez
            Assignee: Jesus Camacho Rodriguez

Currently we compute column statistics by relying on the {{compute_stats}} 
UDAF. For instance, for a given table {{tbl}}, the query to compute statistics 
for columns is translated internally into:
SELECT compute_stats(c1),
FROM tbl;
{{compute_stats}} produces data for the stats available for each column type, 
e.g., struct<"max":long,"min":long,"countnulls":long,...>.

This issue is to produce a query that relies purely on SQL functions instead:
SELECT max(c1), min(c1), count(case when c1 is null then 1 else null end),
FROM tbl;

This will allow us to deprecate the {{compute_stats}} UDAF since it mostly 
duplicates functionality found in those other functions. Additionally, many of 
those functions already provide a vectorized implementation so the approach 
could potentially improve the performance of column stats collection.

This message was sent by Atlassian Jira

Reply via email to