yihua opened a new issue, #20089: URL: https://github.com/apache/hudi/issues/20089
`HoodieFileGroupReaderBasedFileFormat` writes its vectorized read decision into the session conf `spark.sql.parquet.enableVectorizedReader` on every Hudi scan, and nothing sets it back. After one row-based Hudi scan (for example a wide scan over more than `spark.sql.codegen.maxFields` leaf fields), every later scan in the SparkSession reads row-based, including narrow Hudi scans and plain Parquet tables. A plain Parquet read after a wide Hudi scan was measured at 20 to 25% more executor CPU. The same wide scans are also decoded with parquet-mr, where vanilla Spark's Parquet format decodes them vectorized and returns rows. Proposal: decide vectorized decoding per scan from the scan's own schema and the plan-time batch decision, and never mutate the session conf. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
