etseidl commented on code in PR #11262:
URL: https://github.com/apache/arrow-rs/pull/11262#discussion_r4198212454
##########
parquet/src/column/reader.rs:
##########
@@ -484,15 +496,46 @@ where
)?;
offset += bytes_read;
+ if normalize_fixed {
+ // V1 has no non-null count. Both strict
PLAIN validation
+ // and legacy repair need the physical
count: nulls and
+ // nested placeholders have no values in
the data section.
Review Comment:
For the case of FLBA with DLBA encoding, can't we skip the level decoding? I
think we can skip the delta packed lengths fairly quickly, and then get the
number of encoded values based on the size of the remaining bytes.
For PLAIN encoding, I think I'd be fine with scanning ahead and if it's a
repeated `|type_len|type_len bytes|type_len|type_len_bytes...|` I think that's
evidence enough that we've encountered the legacy data and not payload that
just happens to have the same shape. Maybe just check 5 in a row. There too, we
could then infer num_values from the payload size.
Anything to try to avoid the `skip_def_levels` here.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]