HippoBaro opened a new issue, #11277:
URL: https://github.com/apache/arrow-rs/issues/11277

   ### Describe the bug
   
   The Parquet writer does not consistently resolve encoded Arrow layouts when 
deriving the Parquet schema, generating nested levels, and storing the embedded 
`ARROW:schema` metadata. This can produce two different correctness failures:
   
   1. **Logical nulls in run-end-encoded (REE) input can silently become 
ordinary values.** REE arrays do not have a parent validity bitmap, so their 
physical null count is zero even when their run values contain nulls. With an 
outer field marked non-nullable, the writer can declare a required Parquet 
column despite the nullable REE value field. Writing succeeds, but the file has 
no definition levels with which to represent the logical nulls.
   2. **The embedded Arrow schema can prevent reading otherwise valid Parquet 
data.** For example, writing a dense struct under a logically equivalent 
`Dictionary<Int8, Struct<...>>` writer schema succeeds, but the retained 
dictionary hint cannot be reconstructed by the Rust reader. Ignoring Arrow 
metadata allows the values to be read correctly.
   
   ### To Reproduce
   
   ### 1. REE logical nulls are lost
   
   ```rust
   use std::sync::Arc;
   
   use arrow_array::{Array, Int32Array, RecordBatch, RunArray, cast::AsArray, 
types::Int32Type};
   use arrow_schema::{Field, Schema};
   use bytes::Bytes;
   use parquet::arrow::{ArrowWriter, 
arrow_reader::ParquetRecordBatchReaderBuilder};
   
   fn main() -> Result<(), Box<dyn std::error::Error>> {
       let ree = RunArray::<Int32Type>::try_new(
           &Int32Array::from(vec![2, 4, 5]),
           &Int32Array::from(vec![Some(7), None, Some(9)]),
       )?;
       // The outer field is required, but the REE value field is nullable.
       let schema = Arc::new(Schema::new(vec![Field::new(
           "value", ree.data_type().clone(), false,
       )]));
       let batch = RecordBatch::try_new(schema.clone(), vec![Arc::new(ree)])?;
   
       let mut file = Vec::new();
       let mut writer = ArrowWriter::try_new(&mut file, schema, None)?;
       writer.write(&batch)?;
       writer.close()?;
   
       let decoded = 
ParquetRecordBatchReaderBuilder::try_new(Bytes::from(file))?
           .build()?
           .next()
           .unwrap()?;
       assert_eq!(
           decoded.column(0).as_primitive::<Int32Type>(),
           &Int32Array::from(vec![Some(7), Some(7), None, None, Some(9)]),
       );
       Ok(())
   }
   ```
   
   Before the fix, writing and reading both succeed, but the assertion fails: 
the decoded values are `[7, 7, 0, 0, 9]` instead of `[7, 7, null, null, 9]`.
   
   ### 2. A nested dictionary schema hint makes a readable file fail to open
   
   ```rust
   use std::sync::Arc;
   
   use arrow_array::{ArrayRef, Int32Array, RecordBatch, StructArray};
   use arrow_schema::{DataType, Field, Schema};
   use bytes::Bytes;
   use parquet::arrow::{ArrowWriter, 
arrow_reader::ParquetRecordBatchReaderBuilder};
   
   fn main() -> Result<(), Box<dyn std::error::Error>> {
       let column: ArrayRef = Arc::new(StructArray::from(vec![(
           Arc::new(Field::new("item", DataType::Int32, true)),
           Arc::new(Int32Array::from(vec![Some(7), None, Some(9)])) as ArrayRef,
       )]));
       let batch = RecordBatch::try_from_iter([("value", column.clone())])?;
       // Write the dense struct under a logically equivalent dictionary schema.
       let schema = Arc::new(Schema::new(vec![Field::new(
           "value",
           DataType::Dictionary(Box::new(DataType::Int8), 
Box::new(column.data_type().clone())),
           false,
       )]));
   
       let mut file = Vec::new();
       let mut writer = ArrowWriter::try_new(&mut file, schema, None)?;
       writer.write(&batch)?;
       writer.close()?;
   
       let reader = ParquetRecordBatchReaderBuilder::try_new(Bytes::from(file));
       assert!(reader.is_ok(), "reader rejected Arrow schema: {:?}", 
reader.err());
       Ok(())
   }
   ```
   
   Before the fix, the write succeeds, but the assertion fails because reader 
initialization reports:
   
   ```text
   incompatible arrow schema, expected struct got Dictionary(Int8, 
Struct("item": Int32))
   ```
   
   The underlying Parquet values are valid: reading with Arrow metadata ignored 
returns the expected child values `[7, null, 9]`. The failure is in the 
embedded Arrow schema hint.
   
   ### Expected behavior
   
   - Preserve the logical nullability carried by REE value fields. In the first 
example, the resolved Parquet column should be optional and read back `[7, 7, 
null, null, 9]`.
   - Resolve encoded layouts consistently through nested schemas, not only at 
the top level.
   - Store Arrow schema hints that the reader can reconstruct. Dictionary 
wrappers around nested values should not make otherwise valid Parquet data 
unreadable.
   - Preserve supported scalar dictionary hints, including dictionaries of 
`Utf8View` and `BinaryView`.
   - Use the writer's target child-field declarations consistently when 
generating nested levels, even when the input uses a compatible physical layout.
   - Reject map-key schemas that become nullable after resolving dictionary/REE 
wrappers, rather than emitting nullable Parquet map keys.
   
   The reader need not reconstruct the original REE storage layout. It must 
preserve the logical values, nulls, and supported schema information.
   
   
   ### Additional context
   
   _No response_


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to