etseidl commented on issue #627:
URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5891788106

   @zhuqi-lucas those decisions predate my involvement, but having run down 
this same rabbit hole a few years back I can offer some guesses. What you note 
about the dictionary page already having high entropy applies even more so to 
the data pages, which are RLE batches of bit-packed ints. I would suspect the 
`is_compressed` flag was added to avoid cases where data pages grew in size due 
to compression that still helped (albeit maybe not so much) the dictionary 
page. If compression buys you nothing in the dictionary, it's probably just 
best to skip compression altogether for that column. And the same goes for read 
performance; compression is always a space/compute tradeoff. If compute is 
expensive, don't compress. If your main use case is small reads of targeted 
rows, don't use dictionary.
   
   That said, I once saw the need for an `is_compressed` on the dictionary 
page, and there may still be cases where one is desirable. But as you point 
out, it would be a forward incompatible change and would need some strong 
motivation. And it would need to be behind a hard version gate as is being 
discussed now in the community (too many threads, but see [this 
one](https://lists.apache.org/thread/rd271soqncrskd11kcr174gchd1vzpkc) for 
example).
   
   If you want more visibility for this discussion, I'd post to the parquet-dev 
list and link this issue.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to