viirya opened a new issue, #6472:
URL: https://github.com/apache/datafusion-comet/issues/6472

   ### Describe the bug
   
   On Spark 4.x with the default 
`spark.comet.exec.scalaUDF.codegen.enabled=true`, several native string 
expressions return wrong results when their input has a non-`UTF8_BINARY` 
collation. A cast to a collated string type runs through the JVM codegen 
dispatcher, but these expressions above it still run natively and compare or 
search strings by raw bytes, so they ignore the collation:
   
   - `instr`, `substring_index` (`UTF8_LCASE`, `UNICODE_CI`)
   - `trim`, `ltrim`, `rtrim` with an explicit trim string (`UTF8_LCASE`, 
`UNICODE_CI`; `rtrim` also `UTF8_BINARY_RTRIM`)
   - `greatest`, `least` (`UTF8_LCASE`, `UNICODE_CI`; `least` also 
`UTF8_BINARY_RTRIM`)
   
   This is the string counterpart of #6470, which covers the array functions.
   
   ### Steps to reproduce
   
   Parquet table `t` with columns `_1`, `_2` and rows `("a", "A")`, `("x ", 
"x")`, `("b", "c")`, `("Hello World", "hello")`, `("a,B,c", "b")`. With `a = 
CAST(_1 AS STRING COLLATE UTF8_LCASE)` and `b = CAST(_2 AS STRING COLLATE 
UTF8_LCASE)`:
   
   | Expression | Row | Spark | Comet |
   | --- | --- | --- | --- |
   | `instr(a, b)` | `('a', 'A')` | 1 | 0 |
   | `instr(a, b)` | `('a,B,c', 'b')` | 3 | 0 |
   | `substring_index(a, b, 1)` | `('Hello World', 'hello')` | `''` | `'Hello 
World'` |
   | `trim(BOTH b FROM a)` | `('Hello World', 'hello')` | `' World'` | `'Hello 
World'` |
   | `greatest(a, b)` | `('Hello World', 'hello')` | `'Hello World'` | 
`'hello'` |
   | `least(a, b)` | `('a', 'A')` | `'a'` | `'A'` |
   
   The plan shows `cast` dispatched and the string expression native. The 
mismatches reproduce on Spark 4.0 and 4.1.
   
   ### Expected behavior
   
   The results match Spark. Other collation-sensitive string expressions 
already route collated input through the codegen dispatcher, for example 
`contains`, `startswith`, `endswith`, `locate`, `find_in_set`, `replace` and 
`levenshtein`.
   
   ### Additional context
   
   `trim(a)` without a trim string, `upper`, `concat_ws`, `hash` and `xxhash64` 
already match Spark on the same data.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to