Claire McGinty created PARQUET-2443:
---------------------------------------
Summary: Support lazy materialization of row groups in
ParquetFileReader
Key: PARQUET-2443
URL: https://issues.apache.org/jira/browse/PARQUET-2443
Project: Parquet
Issue Type: Improvement
Reporter: Claire McGinty
Motivation: The current behavior of ParquetFilterReader#readNextRowGroup is to
eagerly enumerate all chunks in the row group, then read all pages in the
chunk. For distributed data workloads, this can cause significant memory
pressure, particularly for use cases that require the colocation of multiple
Parquet files on a single worker.
Proposal: A Parquet Configuration option that enables lazy row group reading,
i.e., only a page at a time (plus whatever header is necessary to read that
header). The Configuration option could be either a flag, or an int value for
how many pages/page bytes to buffer at a time.
I think this could be accomplished by modifying
[ParquetFileReader#readAllPages|#L1727]] to re-implement pagesInChunk as an
Iterator<DataPage>, rather than a List<DataPage>. Then, ColumnChunkPageReader
could parse the Configuration option above and decide whether to fully
materialize the iterator or not.
I'm happy to try to create a draft/branch for this to get some early feedback
on the idea!
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]