yihua opened a new pull request, #20085: URL: https://github.com/apache/hudi/pull/20085
### Describe the issue this Pull Request addresses closes #20081 part of #20064 Under schema-on-read, every base file read resolves its file schema from the executor or task manager: `hoodie.properties`, the file's commit file and, when the commit has no internal schema or is not in the valid commits (every new file of a Flink streaming read), a `.hoodie/.schema` listing and read plus a full meta client build. A two-file query issues 78 `.hoodie` accesses from Spark tasks on table version 10. On table version 6 the Spark lookup used the version 2 timeline layout, never found the commit files, and returned an empty file schema for files older than the first schema history version. The Spark incremental relations also wrote schema-on-read settings and the full valid commits list into `sparkContext.hadoopConfiguration` on every streaming batch. Merge order: this PR textually conflicts with #20077 in `HoodieFileGroupReaderBasedFileFormat.scala` and with #20079 in `ParquetSchemaEvolutionUtils.scala`. Whichever merges second gets rebased; the resolution is known. ### Summary and Changelog New `InternalSchemaHistory` (hudi-common) holds the schema history restricted to the query's valid commits plus each commit's own `latest_schema`, read in parallel with the table's timeline layout and cached per commit file, since completed commit files never change. It is loaded once where the query is planned and resolves a version id without I/O, exactly like the timeline lookup. The Spark file group reader format, legacy relations and `SparkReaderContextFactory` ship it in the reader conf; `ParquetSchemaEvolutionUtils` and the Spark 3 and 4 legacy parquet formats resolve from it. Flink's `InternalSchemaManager` holds it instead of a storage conf and table config. Incremental relations pass schema-on-read settings as read options instead of writing the session conf. Tests: a Spark functional test (COW and MOR, table versions 6 and current, rename, type change, add column) checks results and that no Spark task touches `.hoodie` when reading base files; unit tests check equivalence with the timeline lookup, including a commit that completed before an earlier-requested schema change; a Flink test resolves merge schemas with `.hoodie` deleted. ### Impact Schema-on-read base file reads do no per-file metadata I/O on executors and task managers. The driver or job manager reads `.schema` once per query and each commit file once per JVM. On table version 6, files older than the first schema history version now get their commit's schema instead of an empty one. A commit file that exists but cannot be read now fails the query instead of silently falling back. `InternalSchemaManager`'s public constructor changes. MOR log blocks still resolve schemas through the meta client; that moves to the table state in #20078. ### Risk Level medium. It touches the schema-on-read read path of all Spark readers and Flink. Covered by the new tests plus the existing schema evolution suites (Spark DDL, time travel, streaming source, file group reader, Flink `ITTestSchemaEvolution`). The Spark 4 legacy formats get the same edit as Spark 3; CI covers their compile and tests. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
