yihua opened a new issue, #20083: URL: https://github.com/apache/hudi/issues/20083
Flink's `MergeOnReadInputFormat`, `CdcInputFormat` and FLIP-27 split reader functions, and Hive's file group record reader, build a meta client for every split. Hive also walks up the path to find the table, loads the timeline and reads commit metadata to get the schema, per split. Two correctness issues sit next to it: Flink lookup joins never see commits made after their first cache load, and Hive copies job settings into the live table config, so a session merge mode can override the persisted one. Proposal: capture the table state where the read is planned and read splits from it, plan lookup join reloads against a reloaded timeline, and keep Hive read options out of the persisted table config. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
