zhuqi-lucas opened a new issue, #11303:
URL: https://github.com/apache/arrow-rs/issues/11303
**Describe the bug**
Decoding a byte-array dictionary page copies every value into a fresh buffer
and UTF-8 validates all of them, even when the read only resolves a handful of
keys. On a selective read the cost is proportional to the dictionary, not to
the rows produced: a 100k-entry dictionary costs the same whether the query
returns 2 rows or 200k.
`set_dict` materializes the whole dictionary:
```rust
// parquet/src/arrow/array_reader/byte_array.rs
fn set_dict(&mut self, buf: Bytes, num_values: u32, encoding: Encoding,
_is_sorted: bool) -> Result<()> {
...
let mut buffer = OffsetBuffer::with_capacity(0);
let mut decoder = ByteArrayDecoderPlain::new(buf, num_values as usize,
Some(num_values as usize), self.validate_utf8);
decoder.read(&mut buffer, usize::MAX)?; // every value copied, all of
them validated
self.dict = Some(buffer);
Ok(())
}
```
`OffsetBuffer` owns its bytes, so this is a real allocation plus a full copy
of the page:
```rust
pub struct OffsetBuffer<I: OffsetSizeTrait> {
pub offsets: Vec<I>,
pub values: Vec<u8>,
}
```
But the only consumer takes slices and copies just the referenced entries:
```rust
// ByteArrayDecoderDictionary::read
output.extend_from_dictionary(keys, dict.offsets.as_slice(),
dict.values.as_slice())
```
```rust
pub fn extend_from_dictionary<K, V>(&mut self, keys: &[K], dict_offsets:
&[V], dict_values: &[u8]) -> Result<()> {
for key in keys {
...
// Dictionary values are verified when decoding dictionary page
self.values.extend_from_slice(&dict_values[start_offset..end_offset]);
...
}
}
```
So the dictionary is used purely as a lookup table of (offsets, bytes). The
owned copy built in `set_dict` is never needed as such.
**To Reproduce**
Read a small `RowSelection` (a few rows) from a row group whose column chunk
has a large dictionary. The dictionary copy and validation dominate, and they
do not shrink as the selection shrinks.
**Expected behavior**
Building the dictionary should cost a scan of the length prefixes, not a
copy of the page.
Concretely: keep the decompressed page as the already-refcounted `Bytes` and
build only the offsets, then let `extend_from_dictionary` copy the referenced
entries out of that borrowed slice as it already does. That removes the
allocation and the full copy entirely.
UTF-8 validation would move with it. Today `set_dict` validates the whole
dictionary and `extend_from_dictionary` skips validation on that basis (see its
comment). Validating the copied-out entries instead makes the cost proportional
to rows produced. Note this is a small behavior change: a dictionary entry
containing invalid UTF-8 that is never referenced would no longer fail the
read. That seems like an improvement, but it is a change, so worth deciding
deliberately.
What this does not address is decompression: a compressed dictionary page
still has to be decompressed as a unit before the offsets can be walked.
**Additional context**
This is the case that the lazy dictionary work does not reach. Skipping the
dictionary when a column chunk is skipped entirely (#11154, #11168) helps only
when no rows are needed from the chunk; as soon as one row is needed, the full
dictionary cost returns. Selective point lookups and small `LIMIT` queries over
high-cardinality string columns hit this constantly.
Happy to work on this.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]