yihua opened a new pull request, #20102:
URL: https://github.com/apache/hudi/pull/20102

   ### Describe the issue this Pull Request addresses
   
   closes #20100
   part of #20064
   
   For every parquet base file, the Spark readers copy the Hadoop conf three 
times and render the requested schema to JSON three times, although those keys 
depend only on the requested schema. Every task also converts the table schema 
for timestamp repair even when repair is off, the Spark 3.x readers walk every 
footer for shredded variants, and the vectorized reader allocates a full batch 
even for small files.
   
   ### Summary and Changelog
   
   `SparkParquetReaderBase` prepares a `ParquetReadConf` (a `JobConf` with the 
requested schema keys) once per scan and requested schema when the caller 
passes a `SharedScanStorageConfiguration`, as the file group reader format now 
does for base files; other callers get one copy per file instead of three. 
`ParquetSchemaEvolutionUtils.getHadoopAttemptConf` (was `getHadoopConfClone`) 
copies the conf only when the file needs its own requested schema or a pushed 
filter, so the shared conf is never modified. The Spark 3.x readers skip the 
variant walk when no requested struct can be an unshredded variant, vectorized 
batches on Spark 3.5+ are sized to the split's rows, and the timestamp repair 
schema is converted only for tables with a timestamp-millis field. New tests 
check that a COW scan reads all base files with one conf and that shared and 
per-file confs return the same rows, including with pushed filters, type 
changes and 8 concurrent threads.
   
   ### Impact
   
   Lower per-file CPU on Spark parquet base file reads. No API, config or 
output change.
   
   ### Risk Level
   
   low. The files of a scan share a read-only `JobConf`, as vanilla Spark 
already does for footer reads; per-file keys go on a copy, and a concurrent 
test checks the shared conf stays unchanged.
   
   Merging with #20079 (conflicts in `ParquetSchemaEvolutionUtils`, 
`SparkParquetReaderBase` and the six readers): take this PR's 
`getHadoopAttemptConf` and drop #20079's in-place `val hadoopAttemptConf = 
readConf` (`ParquetSchemaEvolutionUtils.scala:89`). With a shared conf, that 
write leaks one file's pushed filter to other files and silently drops their 
rows. Keep 
`TestSparkParquetReaderBase#testAttemptConfIsTheSharedConfUnlessTheFileNeedsItsOwnKeys`
 and `TestSharedScanParquetReadConf` through the merge; they fail on the wrong 
resolution.
   
   This PR also conflicts with #20090 (`ParquetSchemaEvolutionUtils`, the 
readers, `TestBasicSchemaEvolution`) and #20077 
(`HoodieFileGroupReaderBasedFileFormat`, whose two edits move to 
`HoodieFileGroupReaderFunction`); whichever merges second rebases.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to