yihua opened a new pull request, #20115: URL: https://github.com/apache/hudi/pull/20115
### Describe the issue this Pull Request addresses closes #20114 part of #20064 Since HUDI-9451 (#13351) `HoodieFileIndex.prepareFileSlices` returns one `PartitionDirectory` per file slice, on the snapshot and incremental paths and in both `shouldEmbedFileSlices` modes. HUDI-9451 did that so a task is shipped only the slice it reads instead of the whole partition's slice mapping, which for partitions with tens of thousands of slices reached 100 MB+ per task. The mapping is only attached to slices with log files or a bootstrap base, though. Base-file-only slices, which is every slice of a copy-on-write table, carry plain partition values, so splitting them per slice gains nothing and changes what every consumer of `FileIndex.listFiles` sees: `FileSourceScanExec` reports the slice count as `numPartitions`, dynamic partition pruning walks one entry per slice, and listeners or stats consumers that enumerate the relation's partitions get one entry per slice. On a table with 2,334 partitions and 15K slices a query listener that logs input partitions wrote 15K entri es per query, and over a long ETL job a 21 MB line that slowed the driver's task scheduling. ### Summary and Changelog `PartitionDirectoryConverter.convertFileSlicesToPartitionDirectories` takes the slices of one partition and returns one directory with plain partition values holding the delegate files of all base-file-only slices, plus one directory per slice that has log files or a bootstrap base, each with a single-entry `HoodiePartitionFileSliceMapping` as before. `prepareFileSlices` takes the slices grouped by partition values, which is how both `listFiles` callers already produce them, and the non-embedded branch puts all files of a partition into one directory again. Non-partitioned tables get the same grouping under their single empty partition value. Split planning is unchanged: Spark flattens the directories to files, sorts them by length and bin-packs them, and the file order is preserved for base-file-only slices. On a mixed partition the base-file-only directory comes before the per-slice directories, so files of equal length may swap places between tasks, with the same task count and task sizes. Tests: `TestPartitionDirectoryConverter.testConvertFileSlicesToPartitionDirectories` checks the directory shapes for a partition mixing base-file-only, base plus log, log-only and bootstrap slices. `TestHoodieFileIndex.testListFilesGroupsFileSlicesByPartition` writes 3 file groups into each of 5 partitions of a COW and a MOR table and checks that `listFiles` returns 5 directories, that after an update the two file groups with log files get their own directories, and that the file partitions Spark plans match those planned from one directory per slice. Without the fix it fails with `expected: <5> but was: <15>`. ### Impact `listFiles` returns one directory per partition for copy-on-write tables and for the base-file-only slices of merge-on-read tables. `numPartitions` in the scan metrics reports partitions again. Fewer driver objects per scan. ### Risk Level low. Slices that need the mapping keep exactly the per-slice shape of HUDI-9451, and the executor read path is untouched. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
