yihua opened a new issue, #20089:
URL: https://github.com/apache/hudi/issues/20089

   `HoodieFileGroupReaderBasedFileFormat` writes its vectorized read decision 
into the session conf `spark.sql.parquet.enableVectorizedReader` on every Hudi 
scan, and nothing sets it back. After one row-based Hudi scan (for example a 
wide scan over more than `spark.sql.codegen.maxFields` leaf fields), every 
later scan in the SparkSession reads row-based, including narrow Hudi scans and 
plain Parquet tables. A plain Parquet read after a wide Hudi scan was measured 
at 20 to 25% more executor CPU. The same wide scans are also decoded with 
parquet-mr, where vanilla Spark's Parquet format decodes them vectorized and 
returns rows.
   
   Proposal: decide vectorized decoding per scan from the scan's own schema and 
the plan-time batch decision, and never mutate the session conf.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to