[jira] [Commented] (PARQUET-2117) Add rowPosition API in parquet record readers

ASF GitHub Bot (Jira) Mon, 07 Mar 2022 09:23:07 -0800


    [ 
https://issues.apache.org/jira/browse/PARQUET-2117?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17502445#comment-17502445
 ]


ASF GitHub Bot commented on PARQUET-2117:
-----------------------------------------

prakharjain09 commented on a change in pull request #945:
URL: https://github.com/apache/parquet-mr/pull/945#discussion_r820928524



##########
File path: 
parquet-hadoop/src/test/java/org/apache/parquet/hadoop/TestParquetReader.java
##########
@@ -46,10 +47,19 @@
 
   private static final Path FILE_V1 = createTempFile();
   private static final Path FILE_V2 = createTempFile();
-  private static final List<PhoneBookWriter.User> DATA = 
Collections.unmodifiableList(makeUsers(10000));
+  private static final Path STATIC_FILE_WITHOUT_COL_INDEXES = 
createPathFromCP("/test-file-with-no-column-indexes-1.parquet");

Review comment:
       @shangxinli It looks like the 
[column-indexes](https://issues.apache.org/jira/browse/PARQUET-1201) are always 
written in the current version of parquet and are not configurable.
   We are already testing the new row index support with and without the column 
index filtering being triggered (as part of TestColumnIndexFiltering). Also the 
new row index feature doesn't rely on column indexes in any way. So we can skip 
the backward compatibility testing and remove this parquet file from resources. 
What do you think about this?




-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


> Add rowPosition API in parquet record readers
> ---------------------------------------------
>
>                 Key: PARQUET-2117
>                 URL: https://issues.apache.org/jira/browse/PARQUET-2117
>             Project: Parquet
>          Issue Type: New Feature
>          Components: parquet-mr
>            Reporter: Prakhar Jain
>            Priority: Major
>             Fix For: 1.13.0
>
>
> Currently the parquet-mr RecordReader/ParquetFileReader exposes API’s to read 
> parquet file in columnar fashion or record-by-record.
> It will be great to extend them to also support rowPosition API which can 
> tell the position of the current record in the parquet file.
> The rowPosition can be used as a unique row identifier to mark a row. This 
> can be useful to create an index (e.g. B+ tree) over a parquet file/parquet 
> table (e.g.  Spark/Hive).
> There are multiple projects in the parquet eco-system which can benefit from 
> such a functionality: 
>  # Apache Iceberg needs this functionality. It has this implementation 
> already as it relies on low level parquet APIs -  
> [Link1|https://github.com/apache/iceberg/blob/apache-iceberg-0.12.1/parquet/src/main/java/org/apache/iceberg/parquet/ReadConf.java#L171],
>  
> [Link2|https://github.com/apache/iceberg/blob/d4052a73f14b63e1f519aaa722971dc74f8c9796/core/src/main/java/org/apache/iceberg/MetadataColumns.java#L37]
>  # Apache Spark can use this functionality - SPARK-37980



--
This message was sent by Atlassian Jira
(v8.20.1#820001)

[jira] [Commented] (PARQUET-2117) Add rowPosition API in parquet record readers

Reply via email to