HippoBaro opened a new pull request, #11278: URL: https://github.com/apache/arrow-rs/pull/11278
# Which issue does this PR close? - Closes #11277. # Rationale for this change The writer produces schemas that do not always match the logical data it stores. The main case involves run-end-encoded arrays. REE arrays have no parent validity bitmap and report a physical null count of zero even when their run values contain nulls. The writer can therefore declare a non-nullable outer REE field as a required Parquet column, then encode nullable run values without the definition levels needed to represent those nulls. This silently replaces nulls with ordinary values. Embedded Arrow metadata has a related consistency problem. It can preserve dictionary wrappers around nested values that the Rust reader cannot reconstruct. The physical Parquet values and levels are valid in these cases, but `ARROW:schema` describes a representation that prevents the default reader from opening the file. # What changes are included in this PR? - Recursively resolve REE value types when deriving the Parquet schema and remove REE wrappers from embedded Arrow metadata, while preserving field names and metadata. - Merge enclosing-field and run-value nullability so nullable run values use optional Parquet fields and the corresponding definition levels. - Preserve supported scalar dictionary hints, including dictionaries of `Utf8View` and `BinaryView`, while removing dictionary wrappers around nested values that the reader cannot reconstruct. - Recurse through dictionary value types when constructing column writers. - Generate nested levels from the writer's target child fields rather than from child declarations on the incoming array. - Compare compatible inputs by logical value type, keeping the writer's declared schema contract independent of alternate dense, dictionary, run-end, offset-width, and view layouts. - Validate resolved map-key nullability during Arrow-to-Parquet schema conversion, rejecting key schemas that are nullable through their encoding wrappers. # Are these changes tested? Yes. Regression coverage includes: - Recursive REE schema conversion and propagation of run-value nullability. - Nested metadata normalization and preservation of supported scalar view-dictionary hints. - Logically nullable map-key schema rejection, with controls for required keys and nullable members inside valid key structures. - Dense batches written under compatible nested dictionary/REE schemas, including struct, list, and map children. # Are there any user-facing changes? Yes. This is a **correctness-driven schema compatibility change**: - Affected REE fields may become optional where they were previously declared required, allowing their logical nulls to be represented correctly. - Embedded Arrow metadata is normalized to reader-supported representations. Unsupported nested dictionary wrappers are no longer preserved, while supported scalar dictionary hints remain intact. - Arrow-to-Parquet schema conversion rejects map-key schemas whose resolved keys are nullable instead of emitting nullable Parquet map keys. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
