yihua opened a new issue, #20091: URL: https://github.com/apache/hudi/issues/20091
Index lookups on the write path run task functions that capture the whole `HoodieTable` and the index with its write config, so each Spark task deserializes them and rebuilds the table's transient state. The simple and global simple indexes ship the table in one task per base file only to read record keys from that file. `HoodieIndexUtils#getLatestBaseFilesForAllPartitions` runs one task per partition, each building its own file system view. The simple bucket index calls `reloadActiveTimeline()` in tasks for every partition, which lists the timeline on executors and lets tasks disagree on pending instants. Global index partition updates fetch every merged file slice of a partition per file group, and the consistent bucket row writer builds two listing-based views per task. Proposal: capture only the small values these tasks read, share the rest through the engine context broadcast, list base files with one metadata table lookup on the driver, and tag the bucket index from the timeline the write started with. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
