zhuqi-lucas commented on issue #627:
URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5902744922

   Thanks @alamb — "read-heavy" was the wrong word. The contrast I meant is not 
heavy versus light but **a chunk that decodes nothing versus a chunk that 
decodes at least one value**. Deferral removes the dictionary cost entirely for 
the first and cannot help the second, since by then the dictionary is genuinely 
needed.
   
   Not single-row reads specifically — selective scans. The dictionary is 
decompressed once per column chunk no matter how many rows come out of it, 
while data page cost scales with what is touched, so the fixed part dominates 
as selectivity rises. On ClickBench `hits.parquet` the dictionary is 14.7% of 
compressed bytes, but 63% of *decompressed* bytes when 10% of rows are decoded 
and 94.5% at 1%.
   
   That said, my comment above already walks this back: @etseidl pointed out 
the high-entropy argument applies to the data pages (bit-packed indices), not 
the dictionary (PLAIN raw values), and the encoders confirm it. His two 
existing suggestions — per-column `UNCOMPRESSED`, or turning dictionary 
encoding off where the dictionary dominates — cover our case today with no 
format change. I do not think this clears the two-implementations bar on what I 
can show, so it is not worth more of your time. Left open only so the 
selectivity numbers are on the record.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to