yihua opened a new pull request, #20090: URL: https://github.com/apache/hudi/pull/20090
### Describe the issue this Pull Request addresses closes #20089 part of #20064 `HoodieFileGroupReaderBasedFileFormat` writes `spark.sql.parquet.enableVectorizedReader` into the session conf on every scan. After one row-based Hudi scan (such as a scan wider than `spark.sql.codegen.maxFields`), every later scan in the session reads row-based, including plain Parquet tables: a plain Parquet read after a wide Hudi scan used 20 to 25% more executor CPU. The wide scans themselves also decode with parquet-mr, where vanilla Spark decodes them vectorized. ### Summary and Changelog The format decides vectorized decoding per scan from the scan's output schema, like Spark's `ParquetFileFormat`, and never writes the session conf. Whether to return batches comes from the plan-time `FileFormat.OPTION_RETURNING_BATCH`, which the Parquet and ORC reader builders now follow instead of rechecking the conf. `supportBatch` returns the same answer as before and no longer has side effects. Behavior changes: wide scans (and base-only slices of wide MOR scans) decode vectorized and return rows; on the row path, a file with a type change is read row-based with Cast, so no values change and a nested type change no longer fails; a scan planned for batches gets a vectorized reader even if the conf changes before it runs. MOR file-group merges and every existing exclusion stay row-based. New tests cover the session conf, per-file reader choice, conf changes after planning and the vector/variant exclusions. ### Impact Performance only, no API or config change. Later queries in a session keep their own vectorized decision. Wide scans now hold a vectorized batch (`spark.sql.parquet.columnarReaderBatchSize` rows across the requested columns) per task, as vanilla Spark does for the same schema; lowering that size or disabling the vectorized reader limits it. ### Risk Level Medium. More reads go through Spark's nested vectorized reader and wide scans use more memory per task, matching vanilla Spark. Type-changed files, batch output and MOR merges keep their current path. This textually conflicts with #20079 (Parquet reader and `ParquetSchemaEvolutionUtils` lines) and in one import with #20077; whichever merges second rebases. A pre-existing wrong-value case on narrow batch reads after a schema-on-read type change is tracked separately. ### Documentation Update None. ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
