codejoyan commented on issue #2620:
URL: https://github.com/apache/hudi/issues/2620#issuecomment-886077052


   Some additional details for the above runs.
   1. The configs I am using - REGULAR BLOOM.
   2. Max  and Min file size in older partitions - 116 MB and 6 MB respectively
   3. Avg record size - 50 bytes
   4. Avg no of data files in older partitions - Between 157 to 225
   
   I then changed the configs as below to have roughly 100k entries per file. 
But the performance is worse now. It basically gets stuck. Attached Spark Web 
UI screenshot
   1. **Configs** - 
   hoodie.insert.shuffle.parallelism - 1500
   hoodie.upsert.shuffle.parallelism - 1500
   hoodie.parquet.small.file.limit - 4200000
   hoodie.parquet.max.file.size - 5000000
   hoodie.index.type - BLOOM
   hoodie.copyonwrite.record.size.estimate - 50
   hoodie.copyonwrite.insert.split.size - 100000
   hoodie.bloom.index.prune.by.ranges - false
   hoodie.bloom.index.filter.type - DYNAMIC_V0
   hoodie.index.bloom.num_entries - 30000
   
   So basically I am getting relatively better performance (but not within SLA) 
with SIMPLE index but things turn worse with BLOOM index in COW table. 
   
   <img width="1590" alt="Screenshot 2021-07-24 at 10 01 31 PM" 
src="https://user-images.githubusercontent.com/48707638/126875026-8c92b70e-6d83-43e3-966f-26a71c7a156f.png";>
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


Reply via email to