alamb commented on issue #627:
URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5898208798

   Hi @zhuqi-lucas  -- in general I think it would help me to understand this 
proposal if you could provide more details on your usecase
   
   Is the idea that you want to support faster single row reads by allowing 
uncompressed dictionary pages ?
   
   
   > Skip-heavy scans. A reader that skips a column chunk entirely still has to 
decompress its dictionary today, unless the implementation defers it. We 
measured dictionary decompression at 13.6–17.9% of CPU on two production 
query-server classes, all of it reached from the skip path. 
https://github.com/apache/arrow-rs/issues/11154 tracks deferring that decode, 
which removes the cost for chunks that decode nothing — but only for those.
   
   Deferring decode of the dictionary until we know it is needed sounds like 
the right approach to me there
   
   > Read-heavy scans. Any chunk that decodes a value still pays it, and 
deferral cannot help there. An uncompressed dictionary page would, and it is 
the same knob for both cases.
   
   I am not sure what you mean by a read-heavy scan (aren't scans by definition 
read heavy 🤔 ) 
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to