HippoBaro opened a new pull request, #11262:
URL: https://github.com/apache/arrow-rs/pull/11262

   # Which issue does this PR close?
   
   <!--
   We generally require a GitHub issue to be filed for all bug fixes and 
enhancements and this helps us generate change logs for our releases. You can 
link an issue to this PR using the GitHub syntax.
   -->
   
   - Closes #11261 
   
   # Rationale for this change
   
   <!--
   Why are you proposing this change? If this is already explained clearly in 
the issue then this section is not needed.
   Explaining clearly why changes are proposed helps reviewers understand your 
changes and offer better suggestions for fixes.
   -->
   
   The writer can encode `Dictionary<K, FixedSizeBinary>` values using the 
`BYTE_ARRAY` wire representation even though the Parquet schema declares 
`FIXED_LEN_BYTE_ARRAY` (FLBA).
   
   This distinction matters for PLAIN encoding, which is also used for 
dictionary entries: `BYTE_ARRAY` prefixes each value with a four-byte length, 
whereas FLBA stores only the value bytes because their width is defined by the 
schema.
   
   # What changes are included in this PR?
   
   <!--
   There is no need to duplicate the description in the issue here but it is 
sometimes worth providing a summary of the individual changes in this PR.
   -->
   
   - Route Arrow `Dictionary<K, FixedSizeBinary>` values through the FLBA 
writer rather than the byte-array writer.
   - Align dictionary decoding, fallback decoding, and Arrow dictionary spill 
conversion with the declared physical type.
   - Stop emitting `BYTE_ARRAY`-specific size statistics for FLBA columns.
   - Require decompressed FLBA dictionary payloads to contain exactly 
`num_values * type_length` bytes, using checked arithmetic, in the typed column 
reader and both Arrow dictionary-decoding paths.
   - Validate PLAIN data-page value sections independently of repetition and 
definition levels, counting physical non-null values.
   - Validate affected pages before exposing their values, including partial 
reads, skips, and PLAIN pages following dictionary-encoded pages.
   
   The reader follows the schema rather than attempting to infer which 
historical writer produced a payload.
   
   # Are these changes tested?
   
   Yes. Regression coverage exercises FLBA dictionary handling and malformed 
dictionary/data-page payloads, including partial reads, skips, and 
dictionary-to-PLAIN transitions.
   
   # Are there any user-facing changes?
   
   Yes. This is an **intentional file-compatibility break for malformed 
historical payloads** produced by the affected dictionary-writing path. Those 
payloads are no longer accepted as valid FLBA data, and reading them result is 
a well-defined error.
   
   New writes use the physical representation declared by the schema. 
Already-conforming files remain supported, including files produced by the 
dense `FixedSizeBinary` writer path. There are no public Rust API signature 
changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to