yihua opened a new pull request, #20087:
URL: https://github.com/apache/hudi/pull/20087

   ### Describe the issue this Pull Request addresses
   
   closes #20083
   part of #20064
   
   Stacked on #20078 (and #20077), review those first. Until they merge this PR 
shows their commits plus one commit of this change.
   
   Flink's `MergeOnReadInputFormat`, `CdcInputFormat` and FLIP-27 split reader 
functions, and Hive's file group record reader, built a meta client for every 
split; Hive also walked up the path to find the table, loaded the timeline and 
read commit metadata for the schema. Two correctness issues sit next to it: 
Flink lookup joins never saw commits after their first cache load, and Hive 
copied job settings into the live table config, letting a session merge mode 
override the persisted one.
   
   ### Summary and Changelog
   
   - Flink: `ReaderTableStateProvider` captures the table state where the read 
is planned (one snapshot for bounded reads, a lazily loading state per split 
for streaming reads). Input formats and FLIP-27 reader functions read splits 
from it; the CDC FLIP-27 reader no longer keeps one meta client for the job. 
Lookup join cache reloads plan against a reloaded timeline.
   - Hive: splits carry `HiveReaderTableState` (table state, latest completed 
commit, schema), captured once per table while listing and encoded as plain 
data behind a format byte, capped at 16 KB per split. Read options go to a copy 
of the table properties.
   - `FileGroupReaderTableState.withLazyTimeline`.
   - The test helper `MetaFolderAccessRecordingFileSystem` is byte-identical to 
the copy added by #20072 and #20080, so these PRs merge in any order.
   - Tests cover `.hoodie` access from Flink and Hive read tasks (4 to 21 
accesses per query before, 0 after), lookup join reloads, the Hive split 
encoding and size bound, and version 6 log blocks checked against the shipped 
committed instants.
   
   Merge order: this PR textually conflicts with #20071 in 
`HoodieCdcSplitReaderFunction.java` (one call site takes both the instant 
argument from #20071 and the table state from this PR). Whichever merges second 
gets rebased.
   
   ### Impact
   
   Removes per-split `.hoodie` I/O on Flink and Hive read tasks. Behavior 
changes:
   - Bounded Flink reads of tables before version 8 check log blocks against 
the instants committed at planning, as Spark does after #20078.
   - Flink lookup joins see new commits on reload.
   - `MergeOnReadInputFormat` drops the protected `metaClient` field (source 
break for out-of-tree subclasses).
   - Hive splits carry up to 16 KB of encoded table state (measured 1.9 KB for 
50 columns, 8.8 KB for 1,000 columns). `HoodieRealtimeFileSplit` gains a 
trailing format byte and COW snapshot splits use the new 
`FileSplitWithReaderTableState`, so old readers cannot read new splits; 
mixed-version LLAP daemons must be upgraded together with HS2.
   - Hive session or TBLPROPERTIES `hoodie.*` settings no longer override the 
persisted table config:
     - no longer effective (read through `HoodieTableConfig`): 
`hoodie.record.merge.mode`, `hoodie.record.merge.strategy.id`, 
`hoodie.table.ordering.fields` / `hoodie.table.precombine.field`, 
`hoodie.table.partial.update.mode`, `hoodie.table.recordkey.fields`, 
`hoodie.populate.meta.fields`, `hoodie.table.version`, and the payload class 
when the table persists one;
     - still effective (read from the reader props): 
`hoodie.realtime.merge.skip`, `hoodie.datasource.merge.type`, memory and spill 
settings, record merger implementation classes, and the payload class when the 
table does not persist one.
   
     A table without a persisted ordering field that relied on `set 
hoodie.table.precombine.field=...` now merges by commit time.
   
   ### Risk Level
   
   medium: the Hive split wire format and a new split class, and Flink input 
format fields. When no state was captured (splits not produced by Hudi's 
listing, or an unreadable format byte), readers fall back to the old per-split 
path. Covered by the new tests and the existing hudi-hadoop-mr, 
hudi-java-client Hive reader and hudi-flink read suites.
   
   ### Documentation Update
   
   none. The Hive settings that no longer override the table config go in the 
release notes.
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to