exmy opened a new pull request, #13065:
URL: https://github.com/apache/gluten/pull/13065

   CH declares split results as Array(Nullable(String)) while Spark infers 
non-nullable lambda arguments, so native function capture rejects the array 
element column. Align array element types with the lambda argument types for 
filter, transform with index, aggregate and zip_with, and add a regression test.
   
   
   ## What changes are proposed in this pull request?
   
   Fixes #13064.
   
   The CH backend throws `Cannot capture column ... incompatible type` when an 
array higher-order
   function is applied to an array whose native element type differs from the 
type Spark infers for
   the lambda argument. This is the case for arrays produced by `split()`, 
because
   `CHStringSplitTransformer` declares the result as `Array(String, 
containsNull = true)` while
   Spark's `StringSplit` is `ArrayType(StringType, containsNull = false)`, so 
the lambda variable is
   non-nullable `String`.
   
   Align the array element type with the lambda argument type (keeping the 
array's own nullability)
   for `filter`, `transform` with index, `aggregate` (nullable array path) and 
`zip_with`, as
   `transform` without index and `array_sort` already do.
   
   ## How was this patch tested?
   
   - Added `GlutenFunctionValidateSuite#array functions with lambda on nullable 
element array`,
     covering `filter`, `transform` with index, `aggregate` and `zip_with` over 
`split()` results.
   - The failure was reproduced before the fix and confirmed fixed after it on 
the ClickHouse backend
     with Spark 3.3.2 (same code path); `transform` / `aggregate` / `zip_with` 
use the identical
     alignment pattern and still need a CI run on main.
   - Locally: `spotless:check` for `backends-clickhouse`, and clang-format 15 
on the changed C++ lines.
   
   ## Was this patch authored or co-authored using generative AI tooling?
   
   Generated-by: DeepSeek Harness deepseek-v4.1-flash


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to