yihua opened a new pull request, #20087:
URL: https://github.com/apache/hudi/pull/20087
### Describe the issue this Pull Request addresses
closes #20083
part of #20064
Stacked on #20078 (and #20077), review those first. Until they merge this PR
shows their commits plus one commit of this change.
Flink's `MergeOnReadInputFormat`, `CdcInputFormat` and FLIP-27 split reader
functions, and Hive's file group record reader, built a meta client for every
split; Hive also walked up the path to find the table, loaded the timeline and
read commit metadata for the schema. Two correctness issues sit next to it:
Flink lookup joins never saw commits after their first cache load, and Hive
copied job settings into the live table config, letting a session merge mode
override the persisted one.
### Summary and Changelog
- Flink: `ReaderTableStateProvider` captures the table state where the read
is planned (one snapshot for bounded reads, a lazily loading state per split
for streaming reads). Input formats and FLIP-27 reader functions read splits
from it; the CDC FLIP-27 reader no longer keeps one meta client for the job.
Lookup join cache reloads plan against a reloaded timeline.
- Hive: splits carry `HiveReaderTableState` (table state, latest completed
commit, schema), captured once per table while listing and encoded as plain
data behind a format byte, capped at 16 KB per split. Read options go to a copy
of the table properties.
- `FileGroupReaderTableState.withLazyTimeline`.
- The test helper `MetaFolderAccessRecordingFileSystem` is byte-identical to
the copy added by #20072 and #20080, so these PRs merge in any order.
- Tests cover `.hoodie` access from Flink and Hive read tasks (4 to 21
accesses per query before, 0 after), lookup join reloads, the Hive split
encoding and size bound, and version 6 log blocks checked against the shipped
committed instants.
Merge order: this PR textually conflicts with #20071 in
`HoodieCdcSplitReaderFunction.java` (one call site takes both the instant
argument from #20071 and the table state from this PR). Whichever merges second
gets rebased.
### Impact
Removes per-split `.hoodie` I/O on Flink and Hive read tasks. Behavior
changes:
- Bounded Flink reads of tables before version 8 check log blocks against
the instants committed at planning, as Spark does after #20078.
- Flink lookup joins see new commits on reload.
- `MergeOnReadInputFormat` drops the protected `metaClient` field (source
break for out-of-tree subclasses).
- Hive splits carry up to 16 KB of encoded table state (measured 1.9 KB for
50 columns, 8.8 KB for 1,000 columns). `HoodieRealtimeFileSplit` gains a
trailing format byte and COW snapshot splits use the new
`FileSplitWithReaderTableState`, so old readers cannot read new splits;
mixed-version LLAP daemons must be upgraded together with HS2.
- Hive session or TBLPROPERTIES `hoodie.*` settings no longer override the
persisted table config:
- no longer effective (read through `HoodieTableConfig`):
`hoodie.record.merge.mode`, `hoodie.record.merge.strategy.id`,
`hoodie.table.ordering.fields` / `hoodie.table.precombine.field`,
`hoodie.table.partial.update.mode`, `hoodie.table.recordkey.fields`,
`hoodie.populate.meta.fields`, `hoodie.table.version`, and the payload class
when the table persists one;
- still effective (read from the reader props):
`hoodie.realtime.merge.skip`, `hoodie.datasource.merge.type`, memory and spill
settings, record merger implementation classes, and the payload class when the
table does not persist one.
A table without a persisted ordering field that relied on `set
hoodie.table.precombine.field=...` now merges by commit time.
### Risk Level
medium: the Hive split wire format and a new split class, and Flink input
format fields. When no state was captured (splits not produced by Hudi's
listing, or an unreadable format byte), readers fall back to the old per-split
path. Covered by the new tests and the existing hudi-hadoop-mr,
hudi-java-client Hive reader and hudi-flink read suites.
### Documentation Update
none. The Hive settings that no longer override the table config go in the
release notes.
### Contributor's checklist
- [ ] Read through [contributor's
guide](https://hudi.apache.org/contribute/how-to-contribute)
- [ ] Enough context is provided in the sections above
- [ ] Adequate tests were added if applicable
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]