etseidl commented on issue #627: URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5891788106
@zhuqi-lucas those decisions predate my involvement, but having run down this same rabbit hole a few years back I can offer some guesses. What you note about the dictionary page already having high entropy applies even more so to the data pages, which are RLE batches of bit-packed ints. I would suspect the `is_compressed` flag was added to avoid cases where data pages grew in size due to compression that still helped (albeit maybe not so much) the dictionary page. If compression buys you nothing in the dictionary, it's probably just best to skip compression altogether for that column. And the same goes for read performance; compression is always a space/compute tradeoff. If compute is expensive, don't compress. If your main use case is small reads of targeted rows, don't use dictionary. That said, I once saw the need for an `is_compressed` on the dictionary page, and there may still be cases where one is desirable. But as you point out, it would be a forward incompatible change and would need some strong motivation. And it would need to be behind a hard version gate as is being discussed now in the community (too many threads, but see [this one](https://lists.apache.org/thread/rd271soqncrskd11kcr174gchd1vzpkc) for example). If you want more visibility for this discussion, I'd post to the parquet-dev list and link this issue. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
