sdf-jkl opened a new issue, #11050:
URL: https://github.com/apache/arrow-rs/issues/11050

   ### Describe the bug
   
   `variant_get` produces different null semantics depending on whether its 
input is shredded.
   
   Looking up `[0]` in `[]` returns an Arrow-null row for unshredded input, but 
a valid row containing 
[Variant::Null](https://docs.rs/parquet-variant/latest/parquet_variant/enum.Variant.html#variant.Null)
 for shredded input.
   
   For a lookup returning Variant:
   
   | Input | Unshredded arrow-rs | Shredded arrow-rs | Spark |
   |---|---|---|---|
   | `[]` | SQL NULL | Variant null | SQL NULL |
   | `[null]` | Variant null | Variant null | Variant null |
   | `[1]` | Variant `1` | Variant `1` | Variant `1` |
   
   This changes the result’s validity and can affect operations such as `IS 
NULL` and `COUNT(expr)`.
   
   ### To Reproduce
   
   Run this test in `parquet-variant-compute`:
   
   ```rust
   use std::sync::Arc;
   
   use arrow::array::{ArrayRef, StringArray};
   use arrow_schema::DataType;
   use parquet_variant::VariantPath;
   use parquet_variant_compute::{
       GetOptions, VariantArray, json_to_variant, shred_variant, variant_get,
   };
   
   #[test]
   fn missing_list_element_preserves_null_semantics() {
       let json: ArrayRef =
           Arc::new(StringArray::from(vec!["[]", "[null]", "[1]"]));
       let input = json_to_variant(&json).unwrap();
   
       let schema = DataType::new_list(DataType::Int64, true);
       let shredded = shred_variant(&input, &schema).unwrap();
   
       let options = GetOptions::new_with_path(VariantPath::from(0));
   
       let before = variant_get(&ArrayRef::from(input), 
options.clone()).unwrap();
       let after = variant_get(&ArrayRef::from(shredded), options).unwrap();
   
       let before = VariantArray::try_new(&before).unwrap();
       let after = VariantArray::try_new(&after).unwrap();
   
       // Explicit Variant null is a present value.
       assert!(!before.is_null(1));
       assert!(!after.is_null(1));
   
       // An out-of-bounds lookup should produce a missing result.
       assert!(before.is_null(0));
       assert!(after.is_null(0)); // Fails
   }
   ```
   
   Found while working on #10635.
   
   ### Expected behavior
   
   An out-of-bounds lookup should return an Arrow-null row regardless of 
shredding. An existing element containing explicit Variant null should remain a 
valid row containing `Variant::Null`.
   
   This matches arrow-rs’s unshredded lookup and Spark:
   
   - Spark documents that 
[`variant_get`](https://spark.apache.org/docs/4.2.0/api/sql/variant-functions/#variant_get)
 returns NULL when the path does not exist.
   - 
[`is_variant_null`](https://spark.apache.org/docs/4.2.0/api/sql/variant-functions/#is_variant_null)
 distinguishes explicit Variant null from SQL NULL.
   - Spark’s [shredded extraction 
implementation](https://github.com/apache/spark/blob/c470f3dce544ad46ffb748949b6476d3cbc8d443/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/SparkShreddingUtils.scala#L831-L838)
 explicitly returns SQL NULL when the requested index is outside the array.
   
   ### Additional context
   
   The [Parquet Variant array 
rules](https://github.com/apache/parquet-format/blob/master/VariantShredding.md#arrays)
 permit empty arrays. The requirement that elements be present applies to 
existing array positions.
   
   Parquet also specifies [Variant-null recovery when a required stored value 
is 
missing](https://github.com/apache/parquet-format/blob/master/VariantShredding.md#value-shredding).
 That concerns an existing required position with missing contents. The spec 
does not explicitly define the result of an out-of-bounds lookup; this report 
concerns inconsistent lookup semantics and Spark parity.
   
   Related discussion: 
[apache/parquet-format#519](https://github.com/apache/parquet-format/issues/519),
 particularly [scovich’s explanation of missing values, SQL NULL, and Variant 
null](https://github.com/apache/parquet-format/issues/519#issuecomment-3397769385).
   
   The likely cause is that `take_list_like_index_as_shredding_state` creates 
null child values for an out-of-bounds index without propagating the missing 
index into the output’s parent validity. The resulting valid row has both 
children null and is interpreted as Variant null. A fix should preserve index 
presence while retaining the required-value recovery behavior.
   
   **AI usage:** OpenAI Codex assisted with investigating this issue, generated 
and ran arrow-rs reproduction tests, and drafted this report. The analysis 
included checking the Parquet shredding specification, arrow-rs’s existing 
lookup behavior, and Spark’s local implementation and documentation.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to