Dustin Smith created SPARK-60074:
------------------------------------
Summary: Avoid parsing the requested schema twice per split in the
vectorized Parquet reader
Key: SPARK-60074
URL: https://issues.apache.org/jira/browse/SPARK-60074
Project: Spark
Issue Type: Improvement
Components: SQL
Affects Versions: 5.0.0
Reporter: Dustin Smith
For every file split read by the vectorized Parquet reader, the requested Spark
schema JSON (the {{ParquetReadSupport.SPARK_ROW_REQUESTED_SCHEMA}} conf value
set on the driver) is parsed twice with {{StructType.fromString}}: once in
{{ParquetReadSupport.init}}, and again in
{{SpecificParquetRecordReaderBase.initialize}} right after it calls {{init}} on
its own {{ParquetReadSupport}} instance.
The second parse produces an identical schema. Its cost grows with the number
of requested columns: about 0.6 ms per call for 1000 columns.
Measured on 400 Parquet files with 1000 INT columns (local[1], JDK 17, two
interleaved before/after pairs, best time):
* {{SELECT *}}: 5031 / 4987 ms before, 4745 / 4629 ms after (about 4-7% faster)
* {{SELECT c0}} (one requested column, trivial parse): unchanged
Proposed change: when the read support is a {{ParquetReadSupport}}, reuse the
schema its {{init}} already parsed; keep parsing the conf value for any other
read support class. No behavior change.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]