yihua opened a new issue, #20114:
URL: https://github.com/apache/hudi/issues/20114

   Since HUDI-9451 (#13351) the Spark file index returns one 
`PartitionDirectory` per file slice instead of one per partition. 
`HoodieFileIndex.prepareFileSlices` maps each slice to its own directory, on 
the snapshot path (`HoodieFileIndex.listFiles`) and the incremental path 
(`HoodieIncrementalFileIndex.listFiles`), for both `shouldEmbedFileSlices` 
modes.
   
   HUDI-9451 did that to stop shipping the whole partition's slice mapping to 
every task, which for partitions with tens of thousands of slices reached 100 
MB+ per task. But the mapping is only attached to slices that have log files or 
a bootstrap base. Base-file-only slices, which is every slice of a 
copy-on-write table, carry plain partition values, so splitting them per slice 
gains nothing and changes what every consumer of `FileIndex.listFiles` sees:
   
   - `FileSourceScanExec` reports the slice count as `numPartitions` in the SQL 
UI and event logs.
   - Dynamic partition pruning and `selectedPartitions` walk one entry per 
slice.
   - Listeners and catalog stats consumers that enumerate 
`HadoopFsRelation.location` partitions get one entry per slice. On a table with 
2,334 partitions and 15K slices a listener that logs input partitions writes 
15K entries per query, and a 3.5 h ETL job over that table produced a 21 MB log 
line, which slowed the driver's task scheduling for the rest of the job.
   
   Expected: one directory per partition for base-file-only slices, and one 
directory per slice only where the slice mapping is needed.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to