yihua opened a new pull request, #20097: URL: https://github.com/apache/hudi/pull/20097
### Describe the issue this Pull Request addresses closes #20093 part of #20064 Stacked on #20069 (`HoodieEngineContext#broadcast`), review that first. The metadata table writer generates index records with tasks that carry table state and repeat driver-level work. The secondary index update builds a view manager (and without the timeline server a metadata table reader) per written file group and never closes them. The partition stats update ships the table metadata and write config per written partition, builds views per task, reads each written footer twice and reads every column stats file once per written partition. Column stats, bloom filter and record index updates capture the meta client, the tagger ships whole file slices, and the record index bootstrap resolves the schema per file slice. ### Summary and Changelog - Secondary index: previous file slices of the written and replaced file groups are looked up once per commit on the driver (per file group on a remote view, by partition on a local view), the view is closed, and each slice ships with its write stats. - Partition stats: files to read stats for are computed on the driver in batches of 1000 partitions (the consolidated view is closed per batch), column stats are read in one prefix lookup with broadcast prefixes, each written footer is read once, and an unused computation is dropped. - Column stats, bloom filter and record index updates broadcast the meta client; bloom filter tasks get a reader config and the record type instead of the write config. Record index bootstrap resolves its schemas once. The tagger ships record locations instead of file slices. - The broadcasts of one commit's index updates share a new `HoodieBroadcastScope` (the meta client is broadcast once per commit) that is released after the metadata table write. - Tests: `TestMetadataIndexTaskPayload` checks task lineage and the files tasks open and list, on COW and MOR; unit tests for the secondary index view lifecycle, the lookup per commit and the broadcast scope. The task tests fail on the base. ### Impact Performance only: fewer metadata table reader builds, listings and file reads per commit, and smaller task payloads. The metadata table content is unchanged: a dump of all generated records (secondary index, partition stats, column stats, bloom filters, record index update and bootstrap, tagger locations) on COW and MOR fixtures is identical before and after. ### Risk Level low. The lookups that moved to the driver (secondary index previous slices, partition stats file listing, record index bootstrap schema) no longer get Spark task retries, so a transient storage or timeline server failure there fails the metadata table update on the first attempt. Executor tasks share one meta client per executor for reads. Conflicts textually with #20096 in `HoodieBackedTableMetadata#getRecordsByKeyPrefixes` (this PR broadcasts the key prefixes, #20096 rewrites the lookup around a partition reader); whichever merges second rebases. A one-line import conflict with the PR for #20094 in `HoodieBackedTableMetadataWriter` keeps both imports. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
