yihua opened a new pull request, #20086: URL: https://github.com/apache/hudi/pull/20086
### Describe the issue this Pull Request addresses closes #20082 part of #20064 Stacked on #20078 (and #20077), review those first. Until they merge this PR shows their commits plus one commit of this change. In the Hudi Trino connector, every split with log files built a `HoodieTableMetaClient` on the worker. That read `hoodie.properties` and the index definitions (neither goes through the Trino file system cache), resolved the table schema from the timeline and the latest commit metadata unless column name casing resolution was on, and for tables before version 8 loaded the timeline again to check log blocks against the committed instants. A query over N such file groups did this N times. ### Summary and Changelog Workers read file groups from table state that the coordinator ships on the table handle, so they no longer read table metadata from storage. `HudiTableHandle` carries the table config (without the create schema) and, for merge-on-read tables before version 8, the committed instants up to the query's latest commit (new `HudiCommittedInstants`). `HudiMetadata` ships the table schema for every merge-on-read handle. `HudiPageSourceProvider` builds a `FileGroupReaderTableState` from the handle and passes it with the split's storage to the file group reader. Tests: `TestHudiWorkerTableMetadataAccess` asserts through file system tracing spans that workers read log files but nothing under `.hoodie`, for table versions 6 and 8 (it fails on the parent commit), and checks what the coordinator puts on the handle. `TestHudiTableHandle` covers the JSON round trip and the bounded capture of committed instants. ### Impact No worker-side `.hoodie` I/O for splits with log files, down from 3 to 4 metadata reads per split. For merge-on-read handles the coordinator resolves the table schema once per query, even when no split has log files. The handle grows to about 2.5 to 3.2 KB for typical merge-on-read tables (about 9 KB for a 60-column table), plus 20 to 40 bytes per active commit instant for version 6 and 7 merge-on-read tables. ### Risk Level low. The read path gets the same table config, schema and committed-instant semantics as before, captured once on the coordinator. The full hudi-trino suite passes, including the version 6 and 8 merge-on-read smoke and merge-semantics tests. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
