yihua opened a new pull request, #20097:
URL: https://github.com/apache/hudi/pull/20097

   ### Describe the issue this Pull Request addresses
   
   closes #20093
   part of #20064
   
   Stacked on #20069 (`HoodieEngineContext#broadcast`), review that first.
   
   The metadata table writer generates index records with tasks that carry 
table state and repeat driver-level work. The secondary index update builds a 
view manager (and without the timeline server a metadata table reader) per 
written file group and never closes them. The partition stats update ships the 
table metadata and write config per written partition, builds views per task, 
reads each written footer twice and reads every column stats file once per 
written partition. Column stats, bloom filter and record index updates capture 
the meta client, the tagger ships whole file slices, and the record index 
bootstrap resolves the schema per file slice.
   
   ### Summary and Changelog
   
   - Secondary index: previous file slices of the written and replaced file 
groups are looked up once per commit on the driver (per file group on a remote 
view, by partition on a local view), the view is closed, and each slice ships 
with its write stats.
   - Partition stats: files to read stats for are computed on the driver in 
batches of 1000 partitions (the consolidated view is closed per batch), column 
stats are read in one prefix lookup with broadcast prefixes, each written 
footer is read once, and an unused computation is dropped.
   - Column stats, bloom filter and record index updates broadcast the meta 
client; bloom filter tasks get a reader config and the record type instead of 
the write config. Record index bootstrap resolves its schemas once. The tagger 
ships record locations instead of file slices.
   - The broadcasts of one commit's index updates share a new 
`HoodieBroadcastScope` (the meta client is broadcast once per commit) that is 
released after the metadata table write.
   - Tests: `TestMetadataIndexTaskPayload` checks task lineage and the files 
tasks open and list, on COW and MOR; unit tests for the secondary index view 
lifecycle, the lookup per commit and the broadcast scope. The task tests fail 
on the base.
   
   ### Impact
   
   Performance only: fewer metadata table reader builds, listings and file 
reads per commit, and smaller task payloads. The metadata table content is 
unchanged: a dump of all generated records (secondary index, partition stats, 
column stats, bloom filters, record index update and bootstrap, tagger 
locations) on COW and MOR fixtures is identical before and after.
   
   ### Risk Level
   
   low. The lookups that moved to the driver (secondary index previous slices, 
partition stats file listing, record index bootstrap schema) no longer get 
Spark task retries, so a transient storage or timeline server failure there 
fails the metadata table update on the first attempt. Executor tasks share one 
meta client per executor for reads.
   
   Conflicts textually with #20096 in 
`HoodieBackedTableMetadata#getRecordsByKeyPrefixes` (this PR broadcasts the key 
prefixes, #20096 rewrites the lookup around a partition reader); whichever 
merges second rebases. A one-line import conflict with the PR for #20094 in 
`HoodieBackedTableMetadataWriter` keeps both imports.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to