exmy opened a new issue, #13064:
URL: https://github.com/apache/gluten/issues/13064

   ### Backend
   
   CH (ClickHouse)
   
   ### Bug description
   
   ### Describe the bug
   
   On the ClickHouse backend, array higher-order functions fail at runtime when 
the array element type in the native plan differs from the type Spark infers 
for the lambda argument, which is the common case for arrays produced by 
`split()`:
   
   org.apache.gluten.exception.GlutenException: Cannot capture column 2 because 
it has incompatible type: got Nullable(String), but String is expected.
       at DB::ColumnFunction::appendArgument
       at DB::FunctionArrayMapped<DB::ArrayFilterImpl, 
DB::NameArrayFilter>::executeImpl
   
   ### To reproduce
   
   CREATE TABLE t (s string) USING parquet;
   INSERT INTO t VALUES ('a_1,b_2'), ('b_1,c_2'), ('a_3'), (null);
   
   SELECT filter(split(s, ','), x -> split(x, '_')[0] = 'a') FROM t;
   
   Same root cause, also affected:
   
   SELECT transform(split(s, ','), (x, i) -> concat(x, cast(i as string))) FROM 
t;
   SELECT aggregate(split(s, ','), '', (acc, x) -> concat(acc, x)) FROM t;
   SELECT zip_with(split(s, ','), split(s, ','), (x, y) -> concat(x, y)) FROM t;
   
   ### Root cause
   
   `CHStringSplitTransformer` declares the result of `split` as `Array(String, 
containsNull = true)` (CH's `splitByXXX` returns an array of nullable strings), 
while Spark's `StringSplit` is `ArrayType(StringType, containsNull = false)`, 
so the lambda variable `x` is inferred as non-nullable `String`.
   
   ClickHouse's function capture requires the appended column type to be 
exactly equal to the lambda argument type (`ColumnFunction::appendArgument`), 
so the array element column (`Nullable(String)`) is rejected.
   
   `transform` (without index) and `array_sort` already align the array element 
type with the lambda argument type; `filter`, `transform` with index, 
`aggregate` (nullable array path) and `zip_with` do not.
   
   ### Expected behavior
   
   The queries above should be executed natively and return the same results as 
vanilla Spark.
   
   ### Gluten version
   
   _No response_
   
   ### Spark version
   
   None
   
   ### Spark configurations
   
   _No response_
   
   ### System information
   
   _No response_
   
   ### Relevant logs
   
   ```bash
   
   ```


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to