zhuqi-lucas opened a new issue, #11303:
URL: https://github.com/apache/arrow-rs/issues/11303

   **Describe the bug**
   
   Decoding a byte-array dictionary page copies every value into a fresh buffer 
and UTF-8 validates all of them, even when the read only resolves a handful of 
keys. On a selective read the cost is proportional to the dictionary, not to 
the rows produced: a 100k-entry dictionary costs the same whether the query 
returns 2 rows or 200k.
   
   `set_dict` materializes the whole dictionary:
   
   ```rust
   // parquet/src/arrow/array_reader/byte_array.rs
   fn set_dict(&mut self, buf: Bytes, num_values: u32, encoding: Encoding, 
_is_sorted: bool) -> Result<()> {
       ...
       let mut buffer = OffsetBuffer::with_capacity(0);
       let mut decoder = ByteArrayDecoderPlain::new(buf, num_values as usize, 
Some(num_values as usize), self.validate_utf8);
       decoder.read(&mut buffer, usize::MAX)?;   // every value copied, all of 
them validated
       self.dict = Some(buffer);
       Ok(())
   }
   ```
   
   `OffsetBuffer` owns its bytes, so this is a real allocation plus a full copy 
of the page:
   
   ```rust
   pub struct OffsetBuffer<I: OffsetSizeTrait> {
       pub offsets: Vec<I>,
       pub values: Vec<u8>,
   }
   ```
   
   But the only consumer takes slices and copies just the referenced entries:
   
   ```rust
   // ByteArrayDecoderDictionary::read
   output.extend_from_dictionary(keys, dict.offsets.as_slice(), 
dict.values.as_slice())
   ```
   
   ```rust
   pub fn extend_from_dictionary<K, V>(&mut self, keys: &[K], dict_offsets: 
&[V], dict_values: &[u8]) -> Result<()> {
       for key in keys {
           ...
           // Dictionary values are verified when decoding dictionary page
           
self.values.extend_from_slice(&dict_values[start_offset..end_offset]);
           ...
       }
   }
   ```
   
   So the dictionary is used purely as a lookup table of (offsets, bytes). The 
owned copy built in `set_dict` is never needed as such.
   
   **To Reproduce**
   
   Read a small `RowSelection` (a few rows) from a row group whose column chunk 
has a large dictionary. The dictionary copy and validation dominate, and they 
do not shrink as the selection shrinks.
   
   **Expected behavior**
   
   Building the dictionary should cost a scan of the length prefixes, not a 
copy of the page.
   
   Concretely: keep the decompressed page as the already-refcounted `Bytes` and 
build only the offsets, then let `extend_from_dictionary` copy the referenced 
entries out of that borrowed slice as it already does. That removes the 
allocation and the full copy entirely.
   
   UTF-8 validation would move with it. Today `set_dict` validates the whole 
dictionary and `extend_from_dictionary` skips validation on that basis (see its 
comment). Validating the copied-out entries instead makes the cost proportional 
to rows produced. Note this is a small behavior change: a dictionary entry 
containing invalid UTF-8 that is never referenced would no longer fail the 
read. That seems like an improvement, but it is a change, so worth deciding 
deliberately.
   
   What this does not address is decompression: a compressed dictionary page 
still has to be decompressed as a unit before the offsets can be walked.
   
   **Additional context**
   
   This is the case that the lazy dictionary work does not reach. Skipping the 
dictionary when a column chunk is skipped entirely (#11154, #11168) helps only 
when no rows are needed from the chunk; as soon as one row is needed, the full 
dictionary cost returns. Selective point lookups and small `LIMIT` queries over 
high-cardinality string columns hit this constantly.
   
   Happy to work on this.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to