zhuqi-lucas opened a new issue, #627:
URL: https://github.com/apache/parquet-format/issues/627

   ### Problem
   
   `DataPageHeaderV2` can declare a page uncompressed independently of the 
column chunk codec:
   
   ```thrift
   struct DataPageHeaderV2 {
     ...
     /** whether the values are compressed.
         Which means the section of the data page between the levels and the 
values is compressed **/
     7: optional bool is_compressed = 1;
   }
   ```
   
   `DictionaryPageHeader` has no equivalent:
   
   ```thrift
   struct DictionaryPageHeader {
     1: required i32 num_values;
     2: required Encoding encoding;
     3: optional bool is_sorted;
   }
   ```
   
   Compression is otherwise a column-chunk-level property 
(`ColumnMetaData.codec`), so a writer that wants zstd for its data pages has no 
way to leave the dictionary page uncompressed. The asymmetry looks incidental 
rather than intentional: V2 gained a per-page escape hatch, dictionary pages 
did not.
   
   ### Why it matters
   
   A dictionary page is a deduplicated set of distinct values. It is exactly 
the part of a column chunk where compression pays least — duplicates are 
already gone, so the codec is working on close-to-unique data — while it sits 
on the critical path of every reader that decodes any value from the chunk.
   
   The cost is also paid at a different granularity than data pages: one 
dictionary page per column per row group. On files with many columns and small 
row groups that is a large count of small, high-entropy decompressions. In one 
file we looked at (14 columns, 12 row groups, every column dictionary-encoded) 
that is 168 dictionary pages, all zstd.
   
   Two concrete readings:
   
   - **Skip-heavy scans.** A reader that skips a column chunk entirely still 
has to decompress its dictionary today, unless the implementation defers it. We 
measured dictionary decompression at 13.6–17.9% of CPU on two production 
query-server classes, all of it reached from the skip path. 
apache/arrow-rs#11154 tracks deferring that decode, which removes the cost for 
chunks that decode nothing — but only for those.
   - **Read-heavy scans.** Any chunk that decodes a value still pays it, and 
deferral cannot help there. An uncompressed dictionary page would, and it is 
the same knob for both cases.
   
   ### Proposal
   
   Add an optional `is_compressed` to `DictionaryPageHeader`, with the same 
default and semantics as the `DataPageHeaderV2` field:
   
   ```thrift
   struct DictionaryPageHeader {
     1: required i32 num_values;
     2: required Encoding encoding;
     3: optional bool is_sorted;
     /** whether the dictionary values are compressed. When absent the page is
         compressed with the column chunk codec, as today. **/
     4: optional bool is_compressed = 1;
   }
   ```
   
   Defaulting to `true` keeps every existing file and writer valid, and readers 
that ignore the field behave exactly as they do now for files that do not set 
it. A reader that honours it needs one branch in the same place 
`DataPageHeaderV2.is_compressed` is already handled.
   
   Writers would then be free to trade a little file size for decode time on 
the page where that trade is most favourable, and to decide it per column — 
leaving it compressed for a large low-cardinality dictionary that compresses 
well, uncompressed for a high-cardinality one that does not.
   
   ### Open questions
   
   - Is the omission from `DictionaryPageHeader` deliberate? I could not find a 
discussion of it, but I may have missed one.
   - Should the same escape hatch be considered for V1 `DataPageHeader`, or is 
V2 the intended migration path for that case?
   - Is per-page opt-out preferable to a separate column-chunk-level 
"dictionary codec" field? The per-page form matches the existing V2 precedent 
and needs no new column metadata, which is why it is proposed here.
   
   Happy to work on the spec change and the arrow-rs side if this direction 
seems reasonable.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to