sdf-jkl opened a new issue, #11191:
URL: https://github.com/apache/arrow-rs/issues/11191

   ## Describe the bug
   
   When requesting Variant output, `variant_get` returns different null 
semantics depending on whether the input object is shredded.
   
   Extracting `a` from `{}` returns an Arrow-null row (SQL NULL) for unshredded 
input, but a valid row containing `Variant::Null` for shredded input. An 
explicitly present `{"a": null}` remains a valid Variant null in both cases.
   
   | Input | Unshredded result | Shredded result | Expected result |
   |---|---|---|---|
   | `{}` | SQL NULL | Variant null | SQL NULL |
   | `{"a": null}` | Variant null | Variant null | Variant null |
   | `{"a": 42}` | Variant integer 42 | Variant integer 42 | Variant integer 42 
|
   
   This changes the result's validity and can affect downstream `IS NULL` and 
`COUNT(expr)` results. The input is generated using `json_to_variant` and 
`shred_variant`; no malformed arrays are constructed manually.
   
   ## To Reproduce
   
   Run this test in `parquet-variant-compute`:
   
   ```rust
   use std::sync::Arc;
   
   use arrow::array::{ArrayRef, StringArray};
   use arrow_schema::{DataType, Field};
   use parquet_variant::VariantPath;
   use parquet_variant_compute::{
       GetOptions, VariantArray, json_to_variant, shred_variant, variant_get,
   };
   
   #[test]
   fn missing_object_field_preserves_null_semantics() {
       let json: ArrayRef =
           Arc::new(StringArray::from(vec!["{}", r#"{"a": null}"#, r#"{"a": 
42}"#]));
       let input = json_to_variant(&json).unwrap();
       let schema = DataType::Struct(vec![Field::new("a", DataType::Int64, 
true)].into());
       let shredded = shred_variant(&input, &schema).unwrap();
       let options = 
GetOptions::new_with_path(VariantPath::try_from("a").unwrap())
           .with_as_type(Some(Arc::new(input.field("result"))));
   
       let before = variant_get(&ArrayRef::from(input), 
options.clone()).unwrap();
       let after = variant_get(&ArrayRef::from(shredded), options).unwrap();
       let before = VariantArray::try_new(&before).unwrap();
       let after = VariantArray::try_new(&after).unwrap();
   
       // Explicit Variant null is a present value.
       assert!(!before.is_null(1));
       assert!(!after.is_null(1));
   
       // A missing object field should produce a missing result.
       assert!(before.is_null(0));
       assert!(after.is_null(0)); // Fails
   }
   ```
   
   The final assertion fails. `with_as_type` requests a Variant result using 
the field's extension metadata.
   
   ## Expected behavior
   
   A missing object field should produce an Arrow-null row regardless of 
shredding, matching the unshredded kernel. An existing field containing 
explicit Variant null should remain a valid row containing `Variant::Null`.
   
   Spark tests this exact distinction for shredded input in its [missing-fields 
test](https://github.com/apache/spark/blob/c470f3dce544ad46ffb748949b6476d3cbc8d443/sql/core/src/test/scala/org/apache/spark/sql/VariantShreddingSuite.scala#L337-L354).
 Its [shredded extraction 
implementation](https://github.com/apache/spark/blob/c470f3dce544ad46ffb748949b6476d3cbc8d443/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/SparkShreddingUtils.scala#L820-L830)
 checks for both children being absent and returns SQL NULL before 
reconstructing the extracted Variant.
   
   ## Additional context
   
   The [Parquet shredding 
spec](https://github.com/apache/parquet-format/blob/master/VariantShredding.md#objects)
 permits both `value` and `typed_value` to be null for a missing object field. 
An explicit Variant null instead has encoded null bytes in `value`. Its 
[required-value recovery 
rule](https://github.com/apache/parquet-format/blob/master/VariantShredding.md#value-shredding)
 is a separate case: a missing required stored value is reconstructed as 
Variant null. The spec does not itself prescribe SQL lookup semantics; this 
report concerns inconsistent lookup results for equivalent shredded and 
unshredded data.
   
   The apparent cause is that object-field traversal finds the field's physical 
columns but does not propagate per-row field absence into the output's outer 
validity. After the path is exhausted, the child becomes a standalone Variant 
row, and the requested Variant output goes through `unshred_variant`. The 
still-valid row with both children null then triggers required-value recovery 
to Variant null. Relevant code: [object 
traversal](https://github.com/apache/arrow-rs/blob/e491ce6c6f343ff9d7a235376d0353d12b365330/parquet-variant-compute/src/variant_get.rs#L147-L158),
 [validity 
propagation](https://github.com/apache/arrow-rs/blob/e491ce6c6f343ff9d7a235376d0353d12b365330/parquet-variant-compute/src/variant_get.rs#L274-L288),
 and 
[reconstruction](https://github.com/apache/arrow-rs/blob/e491ce6c6f343ff9d7a235376d0353d12b365330/parquet-variant-compute/src/unshred_variant.rs#L83-L100).
   
   Related: #11050 reports the analogous problem for out-of-bounds list access. 
Its proposed fix, #11052, carries list-index validity through traversal, but 
currently leaves `path_nulls: None` for object-field steps. This report covers 
the missing-object-field case separately.
   
   **AI usage:** OpenAI Codex assisted with investigating the behavior, 
generating and running the reproducer, checking related implementations and 
issues, and drafting this report.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to