yihua opened a new pull request, #20112:
URL: https://github.com/apache/hudi/pull/20112

   ### Describe the issue this Pull Request addresses
   
   closes #20111
   part of #20064
   
   Stacked on #20078 (and #20077), review those first. Until they merge this PR 
shows their two commits plus one commit of this change.
   
   Nothing tests the per-task cost of a Spark read, so regressions like tasks 
reading `.hoodie` on the executors, or the meta client and Hadoop configuration 
riding in every task closure, went unnoticed.
   
   ### Summary and Changelog
   
   `TestSparkReadExecutorFootprint` reads version 6 and current COPY_ON_WRITE 
and MERGE_ON_READ tables with snapshot, read optimized, incremental, time 
travel and CDC queries, with and without the metadata table, with schema on 
read and with data skipping. It fails when a task touches `.hoodie`, when a 
task deserializes the meta client, timeline, storage or Hadoop configuration, 
write config or the file format, or when the task binary grows past 24 KB (1.4x 
the largest size measured on this base, 17544 bytes).
   
   The helpers are reusable by other engines: 
`MetaFolderAccessRecordingFileSystem` (the same file as in #20072 and #20080, 
so the PRs merge in any order), `TaskDeserializationRecorder` (a JVM-wide 
deserialization filter) and `SparkExecutorGuards`.
   
   Rows that need an open PR are `@Disabled` with the PR that enables them:
   - #20080: reads with the metadata table off that list partitions with a 
Spark job (10 rows)
   - #20085: schema-on-read reads (4 rows)
   - #20071: CDC read of a version 6 MERGE_ON_READ table (both guards)
   
   The other 33 rows pass on #20078.
   
   ### Impact
   
   Test only.
   
   ### Risk Level
   
   none
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to