yihua opened a new pull request, #20096: URL: https://github.com/apache/hudi/pull/20096
### Describe the issue this Pull Request addresses closes #20092 part of #20064 Stacked on #20069 (`HoodieEngineContext#broadcast`), review that first. Distributed metadata table lookups in `HoodieBackedTableMetadata` (column stats, expression, record and secondary index) capture the whole table metadata, with both meta clients and the file system view, and a task can list the data table timeline to compute the valid instants. The record index and metadata bloom filter functions on the write path ship the `HoodieTable` and open a new metadata reader in every task. The column stats read ships every candidate file name in every task. ### Summary and Changelog - New serializable `MetadataPartitionReader` holding only driver-resolved state for one metadata table partition (MDT meta client with its timeline, file group reader props, latest MDT instant, valid instants, file slices), exposed through `HoodieTableMetadata#getPartitionReader`. Distributed lookups capture only this reader. - Record index (global and partitioned) and metadata bloom filter probing functions hold an engine context broadcast of the reader instead of the table. `PartitionIdPassthrough` is static, so the lookup partitioner no longer carries the index and write config. Record index keys are encoded through `RecordIndexRawKey` before the lookup. - The on-cluster column stats read broadcasts its candidate file names. - `SparkRDDReadClient` builds a fresh table per lookup, so a reused client sees later commits. - Tests: `TestMetadataTableLookupTasks` with a counting local file system that fails if a lookup task lists a timeline or reads `hoodie.properties`, plus task payload checks; `TestMetadataTableLookupWithPlainKryo`. The I/O and payload tests fail on master. Follow-up: `TaskMetaIOCountingFileSystem.scala` here is a near duplicate of `TaskFileAccessRecordingFileSystem` added for #20093; merge them once both land. ### Impact Smaller task binaries for metadata table lookups (about half on a local test table, more with cluster Hadoop configurations or many candidate files) and no timeline or table config I/O in lookup tasks. Write-path index lookups use the driver's metadata table snapshot for all tasks. `PartitionedRecordIndexFileGroupLookupFunction` and `HoodieMetadataBloomFilterProbingFunction` constructors change. With `hoodie.metadata.metrics.enable=true`, bloom filter probing tasks no longer record `lookup_meta_index_bloom_filters` and its file count from executor JVMs, since they no longer build a table metadata (and a metrics reporter) there. These values never reached the driver's registry and were only visible through a push reporter reachable from executors. Driver-side metadata metrics and the record index lookup metrics are unchanged. ### Risk Level medium. The lookup logic is moved, not rewritten; results are covered by the existing index suites plus new equivalence tests. Write-path lookups now read one driver snapshot of the metadata table. This PR conflicts textually with the PR for #20093 in `HoodieBackedTableMetadata#getRecordsByKeyPrefixes`; whichever merges second rebases. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
