alamb commented on issue #627: URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5898208798
Hi @zhuqi-lucas -- in general I think it would help me to understand this proposal if you could provide more details on your usecase Is the idea that you want to support faster single row reads by allowing uncompressed dictionary pages ? > Skip-heavy scans. A reader that skips a column chunk entirely still has to decompress its dictionary today, unless the implementation defers it. We measured dictionary decompression at 13.6–17.9% of CPU on two production query-server classes, all of it reached from the skip path. https://github.com/apache/arrow-rs/issues/11154 tracks deferring that decode, which removes the cost for chunks that decode nothing — but only for those. Deferring decode of the dictionary until we know it is needed sounds like the right approach to me there > Read-heavy scans. Any chunk that decodes a value still pays it, and deferral cannot help there. An uncompressed dictionary page would, and it is the same knob for both cases. I am not sure what you mean by a read-heavy scan (aren't scans by definition read heavy 🤔 ) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
