peterxcli opened a new issue, #5426:
URL: https://github.com/apache/datafusion-comet/issues/5426

   ### What is the problem the feature request solves?
   
   #5424 tracks native `variant_get` / `try_variant_get` for literal paths and 
scalar, non-Variant targets, while #5425 tracks the two-argument 
Variant-returning form. Spark also supports a path supplied by a column and 
nested target types:
   
   ```sql
   SELECT variant_get(v, path, 'bigint') FROM t;
   SELECT try_variant_get(v, '$.items', 'array<int>') FROM t;
   SELECT variant_get(v, '$.customer', 'struct<id:bigint,name:string>') FROM t;
   ```
   
   Spark 4.1.3 evaluates a foldable path once but has a separate per-row path 
branch for non-foldable expressions in 
[`VariantGet`](https://github.com/apache/spark/blob/v4.1.3/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/variant/variantExpressions.scala#L263-L325)
 and its [UTF8 path 
helper](https://github.com/apache/spark/blob/v4.1.3/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/variant/variantExpressions.scala#L391-L403).
 It accepts arrays, maps with string keys, and structs recursively as target 
types in 
[`checkDataType`](https://github.com/apache/spark/blob/v4.1.3/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/variant/variantExpressions.scala#L345-L362),
 with the corresponding recursive casts implemented for [arrays, maps, and 
structs](https://github.com/apache/spark/blob/v4.1.3/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/variant/variantExpressions.scala#L495-L535).
   
   Without these paths, otherwise supported Variant extraction queries still 
cross back to Spark.
   
   ### Describe the potential solution
   
   Extend the expression-specific implementation from #5424/#5425 without 
enabling Variant in Comet's general expression gates:
   
   - evaluate and parse a non-foldable string path per non-null row, preserving 
Spark's invalid-path error and SQL NULL behavior;
   - support exactly Spark's recursive target-type set: arrays, maps with 
string keys, and structs whose children are supported;
   - reproduce Spark's missing-field, extra-field, nested null, JSON null, 
decimal, datetime, and time-zone behavior;
   - preserve the strict/try distinction recursively: `variant_get` raises on 
conversion failure, while `try_variant_get` returns SQL NULL at the failing 
value; and
   - keep unsupported target types and shapes on explicit Spark fallback.
   
   Add focused Spark SQL parity and plan tests for column paths, mixed paths 
across rows, invalid and SQL NULL paths, nested arrays/maps/structs, missing 
object fields, conversion failures, and nullable nested values.
   
   ### Additional context
   
   Related work: #4295, #3983, #5407, #5424, and #5425.
   
   Non-goals: predicate pushdown or subfield pruning, Variant constructors, 
general casts, generators, shuffle/spill, writing, Python transport, and 
Iceberg-specific work.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to