GitHub user cshuo edited a comment on the discussion: Use ZSTD as the default 
compression codec for Parquet files

@danny0405 I think these are two separate concerns. 
* To switch the default to ZSTD, the memory leak in older parquet-java versions 
must be addressed. As mentioned earlier, we can follow Iceberg’s workaround: 
[apache/iceberg#5681](https://github.com/apache/iceberg/pull/5681).
* We should also consider MOR log write throughput. In a preliminary benchmark 
on Alibaba Cloud EMR, for a [MOR 
table](https://github.com/cshuo/streaming-benchmark/blob/main/queries/q2.sql) 
with 20 string fields, with ZSTD enabled, native log write throughput was only 
about 70% of inline log throughput. To improve throughput, we need to tune 
compression parameters specifically for native log files rather than keeping 
them fully aligned with base files. Actually, different compression policies 
across levels or file lifecycles are common in LSM storage. RocksDB, TiKV and 
Paimon support per-level compression, while Cassandra and HBase differentiate 
compression between flush and compaction.

GitHub link: 
https://github.com/apache/hudi/discussions/19615#discussioncomment-18000409

----
This is an automatically sent email for [email protected].
To unsubscribe, please send an email to: [email protected]

Reply via email to