[ 
https://issues.apache.org/jira/browse/PARQUET-2261?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17757603#comment-17757603
 ] 

ASF GitHub Bot commented on PARQUET-2261:
-----------------------------------------

mapleFU commented on code in PR #197:
URL: https://github.com/apache/parquet-format/pull/197#discussion_r1302013630


##########
src/main/thrift/parquet.thrift:
##########
@@ -191,6 +191,64 @@ enum FieldRepetitionType {
   REPEATED = 2;
 }
 
+/**
+  * A histogram of repetition and definition levels for either a page or 
column chunk. 
+  *
+  * This is useful for:
+  *   1. Estimating the size of the data when materialized in memory 
+  *   2. For filter push-down on nulls at various levels of nested structures 
and 
+  *      list lengths.
+  */ 
+struct RepetitionDefinitionLevelHistogram {
+   /** 
+     * When present, there is expected to be one element corresponding
+     to each repetition (i.e. size=max repetition_level+1) 

Review Comment:
   ```suggestion
        * to each repetition (i.e. size=max repetition_level+1) 
   ```



##########
src/main/thrift/parquet.thrift:
##########
@@ -191,6 +191,64 @@ enum FieldRepetitionType {
   REPEATED = 2;
 }
 
+/**
+  * A histogram of repetition and definition levels for either a page or 
column chunk. 
+  *
+  * This is useful for:
+  *   1. Estimating the size of the data when materialized in memory 

Review Comment:
   (2) can be easily done with these statistics, but I wonder how can we 
estimate the size of data with rep-def histogram...Is there any formulas?



##########
src/main/thrift/parquet.thrift:
##########
@@ -974,6 +1052,13 @@ struct ColumnIndex {
 
   /** A list containing the number of null values for each page **/
   5: optional list<i64> null_counts
+  /** 
+    * Repetition and definition level histograms for the pages.  
+    *
+    * This contains some redundancy with null_counts, however, to accommodate  
the

Review Comment:
   ```suggestion
       * This contains some redundancy with null_counts, however, to 
accommodate the
   ```





> [Format] Add statistics that reflect decoded size to metadata
> -------------------------------------------------------------
>
>                 Key: PARQUET-2261
>                 URL: https://issues.apache.org/jira/browse/PARQUET-2261
>             Project: Parquet
>          Issue Type: Improvement
>          Components: parquet-format
>            Reporter: Micah Kornfield
>            Assignee: Micah Kornfield
>            Priority: Major
>




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to