Dustin Smith created SPARK-60074:
------------------------------------

             Summary: Avoid parsing the requested schema twice per split in the 
vectorized Parquet reader
                 Key: SPARK-60074
                 URL: https://issues.apache.org/jira/browse/SPARK-60074
             Project: Spark
          Issue Type: Improvement
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: Dustin Smith


For every file split read by the vectorized Parquet reader, the requested Spark 
schema JSON (the {{ParquetReadSupport.SPARK_ROW_REQUESTED_SCHEMA}} conf value 
set on the driver) is parsed twice with {{StructType.fromString}}: once in 
{{ParquetReadSupport.init}}, and again in 
{{SpecificParquetRecordReaderBase.initialize}} right after it calls {{init}} on 
its own {{ParquetReadSupport}} instance.

The second parse produces an identical schema. Its cost grows with the number 
of requested columns: about 0.6 ms per call for 1000 columns.

Measured on 400 Parquet files with 1000 INT columns (local[1], JDK 17, two 
interleaved before/after pairs, best time):
* {{SELECT *}}: 5031 / 4987 ms before, 4745 / 4629 ms after (about 4-7% faster)
* {{SELECT c0}} (one requested column, trivial parse): unchanged

Proposed change: when the read support is a {{ParquetReadSupport}}, reuse the 
schema its {{init}} already parsed; keep parsing the conf value for any other 
read support class. No behavior change.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to