[ 
https://issues.apache.org/jira/browse/SPARK-60074?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SPARK-60074:
-----------------------------------
    Labels: pull-request-available  (was: )

> Avoid parsing the requested schema twice per split in the vectorized Parquet 
> reader
> -----------------------------------------------------------------------------------
>
>                 Key: SPARK-60074
>                 URL: https://issues.apache.org/jira/browse/SPARK-60074
>             Project: Spark
>          Issue Type: Improvement
>          Components: SQL
>    Affects Versions: 5.0.0
>            Reporter: Dustin Smith
>            Priority: Major
>              Labels: pull-request-available
>
> For every file split read by the vectorized Parquet reader, the requested 
> Spark schema JSON (the {{ParquetReadSupport.SPARK_ROW_REQUESTED_SCHEMA}} conf 
> value set on the driver) is parsed twice with {{StructType.fromString}}: once 
> in {{ParquetReadSupport.init}}, and again in 
> {{SpecificParquetRecordReaderBase.initialize}} right after it calls {{init}} 
> on its own {{ParquetReadSupport}} instance.
> The second parse produces an identical schema. Its cost grows with the number 
> of requested columns: about 0.6 ms per call for 1000 columns.
> Measured on 400 Parquet files with 1000 INT columns (local[1], JDK 17, two 
> interleaved before/after pairs, best time):
> * {{SELECT *}}: 5031 / 4987 ms before, 4745 / 4629 ms after (about 4-7% 
> faster)
> * {{SELECT c0}} (one requested column, trivial parse): unchanged
> Proposed change: when the read support is a {{ParquetReadSupport}}, reuse the 
> schema its {{init}} already parsed; keep parsing the conf value for any other 
> read support class. No behavior change.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to