yihua opened a new pull request, #20103: URL: https://github.com/apache/hudi/pull/20103
### Describe the issue this Pull Request addresses closes #20101 part of #20064 For every parquet base file, `HoodieParquetFileFormatHelper.buildImplicitSchemaChangeInfo` converts the whole footer schema to a Spark schema and compares it with the requested schema to find implicit type changes, even for a query that reads a few columns. The files of a table share a few schemas, so scans with many files repeat the same work. ### Summary and Changelog `buildImplicitSchemaChangeInfo` caches its result per JVM, keyed by the parquet file schema, the requested schema, the Hadoop conf values `ParquetToSparkSchemaConverter` reads and, on Spark 4.1+, `spark.sql.variant.allowReadingShredded`. The cache is bounded by weight (file columns plus requested leaf fields) and at 256 entries, results that relied on a failed Spark adapter check are not cached, and the returned type change map is read-only because files share it. The computation itself is unchanged. `TestHoodieParquetFileFormatHelper` checks that equal schemas reuse one result, that cached results equal the uncached computation across type, nested, nullability and conf changes, and that the key covers every Hadoop conf key the running Spark version's converter reads. ### Impact Lower per-file CPU on Spark parquet base file reads of tables without schema on read, most for narrow queries on wide files. No API, config or output change. ### Risk Level low. The cached value is the previous result for an equal key that covers all of its inputs, and a test fails if a Spark version's converter reads a Hadoop conf key the key misses. Worst-case memory is a few MB. Merges cleanly with #20077, #20079, #20090 and #20102. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
