[ 
https://issues.apache.org/jira/browse/HIVE-27158?focusedWorklogId=852505&page=com.atlassian.jira.plugin.system.issuetabpanels:worklog-tabpanel#worklog-852505
 ]

ASF GitHub Bot logged work on HIVE-27158:
-----------------------------------------

                Author: ASF GitHub Bot
            Created on: 23/Mar/23 12:23
            Start Date: 23/Mar/23 12:23
    Worklog Time Spent: 10m 
      Work Description: simhadri-g commented on code in PR #4131:
URL: https://github.com/apache/hive/pull/4131#discussion_r1146105589


##########
iceberg/iceberg-handler/src/main/java/org/apache/iceberg/mr/hive/HiveIcebergStorageHandler.java:
##########
@@ -349,6 +364,98 @@ public Map<String, String> getBasicStatistics(Partish 
partish) {
     return stats;
   }
 
+
+  @Override
+  public boolean canSetColStatistics() {
+    String statsSource = HiveConf.getVar(conf, 
HiveConf.ConfVars.HIVE_USE_STATS_FROM).toLowerCase();
+    return statsSource.equals(PUFFIN) ? true : false;
+  }
+
+  @Override
+  public boolean 
canProvideColStatistics(org.apache.hadoop.hive.ql.metadata.Table tbl) {
+
+    org.apache.hadoop.hive.ql.metadata.Table hmsTable = tbl;
+    TableDesc tableDesc = Utilities.getTableDesc(hmsTable);
+    Table table = Catalogs.loadTable(conf, tableDesc.getProperties());
+    if (table.currentSnapshot() != null) {
+      String statsSource = HiveConf.getVar(conf, 
HiveConf.ConfVars.HIVE_USE_STATS_FROM).toLowerCase();
+      String statsPath = table.location() + "/stats/" + table.name() + 
table.currentSnapshot().snapshotId();
+      if (statsSource.equals(PUFFIN)) {
+        try (FileSystem fs = new Path(table.location()).getFileSystem(conf)) {
+          if (fs.exists(new Path(statsPath))) {
+            return true;
+          }
+        } catch (IOException e) {
+          LOG.warn(e.getMessage());
+        }
+      }
+    }
+    return false;
+  }
+
+  @Override
+  public List<ColumnStatisticsObj> 
getColStatistics(org.apache.hadoop.hive.ql.metadata.Table tbl) {
+
+    org.apache.hadoop.hive.ql.metadata.Table hmsTable = tbl;
+    TableDesc tableDesc = Utilities.getTableDesc(hmsTable);
+    Table table = Catalogs.loadTable(conf, tableDesc.getProperties());
+    String statsSource = HiveConf.getVar(conf, 
HiveConf.ConfVars.HIVE_USE_STATS_FROM).toLowerCase();
+    switch (statsSource) {
+      case ICEBERG:
+        // Place holder for iceberg stats
+        break;
+      case PUFFIN:
+        String snapshotId = table.name() + 
table.currentSnapshot().snapshotId();

Review Comment:
   Hi Rajesh, 
   Thanks for the review! :) 
   
   Currently, the scope of this patch is for column stats. We can include basic 
stats in this puffin file as well if needed.
   
   But currently, basics stats for iceberg tables are obtained from 
`table.currentSnapshot().summary() `  here : 
https://github.com/apache/hive/blob/master/iceberg/iceberg-handler/src/main/java/org/apache/iceberg/mr/hive/HiveIcebergStorageHandler.java#L316
 .
   I do not think we are querying the metastore for basic stats. Please correct 
me if i am missing something.
   
   
   





Issue Time Tracking
-------------------

    Worklog Id:     (was: 852505)
    Time Spent: 40m  (was: 0.5h)

> Store hive columns stats in puffin files for iceberg tables
> -----------------------------------------------------------
>
>                 Key: HIVE-27158
>                 URL: https://issues.apache.org/jira/browse/HIVE-27158
>             Project: Hive
>          Issue Type: Improvement
>            Reporter: Simhadri Govindappa
>            Assignee: Simhadri Govindappa
>            Priority: Major
>              Labels: pull-request-available
>          Time Spent: 40m
>  Remaining Estimate: 0h
>




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to