yihua opened a new pull request, #20095:
URL: https://github.com/apache/hudi/pull/20095

   ### Describe the issue this Pull Request addresses
   
   closes #20091
   part of #20064
   
   Stacked on #20069 (`HoodieEngineContext#broadcast`), review that first.
   
   Index lookups capture the whole `HoodieTable` (and the index with its write 
config) in their task functions. The simple and global simple indexes ship it 
in one task per base file to read record keys; 
`HoodieIndexUtils#getLatestBaseFilesForAllPartitions` runs one task per 
partition, each building a view; the simple bucket index reloads the timeline 
in tasks for every partition; global index partition updates fetch every merged 
slice of a partition per file group; the consistent bucket row writer builds 
two views per task.
   
   ### Summary and Changelog
   
   - Simple and global simple lookups run static task functions that capture 
only the base file, base path, key generator and a broadcast storage 
configuration (`HoodieKeyLocationFetchHandle#locations` / `#globalLocations`).
   - `getLatestBaseFilesForAllPartitions`: with the metadata table, one batched 
metadata lookup on the driver into a short-lived view and no tasks; without it, 
one task per partition as before.
   - `HoodieSimpleBucketIndex#tagLocation` uses a broadcast 
`BucketLocationLoader` built from the writer's latest completed commit and 
pending instants; each worker builds one view and caches each partition's 
bucket mapping. No timeline reload in tasks. Subclasses keep the per-partition 
path and their hooks.
   - `getExistingRecords` looks up only the requested file group's merged slice 
and parses the internal schema once.
   - The consistent bucket row writer reads pending clustering file groups from 
the timeline and builds at most one view per task.
   - Tests: task lineage checks that fail when a task carries a table, config, 
meta client or index, an executor timeline listing counter, a single file group 
lookup check and a view count check. Each fails on master. Two bloom tests and 
one bucket test now reload their meta client before looking up files written 
after the table was built.
   
   ### Impact
   
   Fewer bytes deserialized per index task and no per-task view, metadata table 
reader or timeline listing for these lookups; results are unchanged. Bucket 
index tagging follows the timeline snapshot the write started with instead of a 
per-task reload. For tables with an HFile base file format, the key-read file 
opens no longer go through `HoodieWrapperFileSystem`, so they lose the 
`hoodie.filesystem.operation.retry.*` retries.
   
   ### Risk Level
   
   medium. Touches tagging for the simple, global simple, bloom and simple 
bucket indexes. Covered by the new tests and the existing simple, global 
simple, bloom, global bloom, bucket, consistent bucket and global index 
partition update suites.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to