yihua opened a new pull request, #20095: URL: https://github.com/apache/hudi/pull/20095
### Describe the issue this Pull Request addresses closes #20091 part of #20064 Stacked on #20069 (`HoodieEngineContext#broadcast`), review that first. Index lookups capture the whole `HoodieTable` (and the index with its write config) in their task functions. The simple and global simple indexes ship it in one task per base file to read record keys; `HoodieIndexUtils#getLatestBaseFilesForAllPartitions` runs one task per partition, each building a view; the simple bucket index reloads the timeline in tasks for every partition; global index partition updates fetch every merged slice of a partition per file group; the consistent bucket row writer builds two views per task. ### Summary and Changelog - Simple and global simple lookups run static task functions that capture only the base file, base path, key generator and a broadcast storage configuration (`HoodieKeyLocationFetchHandle#locations` / `#globalLocations`). - `getLatestBaseFilesForAllPartitions`: with the metadata table, one batched metadata lookup on the driver into a short-lived view and no tasks; without it, one task per partition as before. - `HoodieSimpleBucketIndex#tagLocation` uses a broadcast `BucketLocationLoader` built from the writer's latest completed commit and pending instants; each worker builds one view and caches each partition's bucket mapping. No timeline reload in tasks. Subclasses keep the per-partition path and their hooks. - `getExistingRecords` looks up only the requested file group's merged slice and parses the internal schema once. - The consistent bucket row writer reads pending clustering file groups from the timeline and builds at most one view per task. - Tests: task lineage checks that fail when a task carries a table, config, meta client or index, an executor timeline listing counter, a single file group lookup check and a view count check. Each fails on master. Two bloom tests and one bucket test now reload their meta client before looking up files written after the table was built. ### Impact Fewer bytes deserialized per index task and no per-task view, metadata table reader or timeline listing for these lookups; results are unchanged. Bucket index tagging follows the timeline snapshot the write started with instead of a per-task reload. For tables with an HFile base file format, the key-read file opens no longer go through `HoodieWrapperFileSystem`, so they lose the `hoodie.filesystem.operation.retry.*` retries. ### Risk Level medium. Touches tagging for the simple, global simple, bloom and simple bucket indexes. Covered by the new tests and the existing simple, global simple, bloom, global bloom, bucket, consistent bucket and global index partition update suites. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
