yihua opened a new issue, #20093: URL: https://github.com/apache/hudi/issues/20093
The metadata table writer generates index records for every commit with tasks that carry table state and repeat driver-level work. The secondary index update builds a file system view manager (and, without the timeline server, a metadata table reader) in every task and never closes them. The partition stats update ships the table metadata and write config per written partition, builds views per task, reads each written base file footer twice and runs its own column stats lookup, so each column stats file is read once per written partition. Column stats, bloom filter and record index updates capture the data meta client, the record tagger ships whole file slices, and the record index bootstrap resolves the schema per file slice, reading the timeline when the write config has no schema. Proposal: resolve these inputs once per commit on the driver, broadcast the meta client, ship only what the tasks read, and release the broadcasts after the metadata table write. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
