viirya opened a new issue, #6472:
URL: https://github.com/apache/datafusion-comet/issues/6472
### Describe the bug
On Spark 4.x with the default
`spark.comet.exec.scalaUDF.codegen.enabled=true`, several native string
expressions return wrong results when their input has a non-`UTF8_BINARY`
collation. A cast to a collated string type runs through the JVM codegen
dispatcher, but these expressions above it still run natively and compare or
search strings by raw bytes, so they ignore the collation:
- `instr`, `substring_index` (`UTF8_LCASE`, `UNICODE_CI`)
- `trim`, `ltrim`, `rtrim` with an explicit trim string (`UTF8_LCASE`,
`UNICODE_CI`; `rtrim` also `UTF8_BINARY_RTRIM`)
- `greatest`, `least` (`UTF8_LCASE`, `UNICODE_CI`; `least` also
`UTF8_BINARY_RTRIM`)
This is the string counterpart of #6470, which covers the array functions.
### Steps to reproduce
Parquet table `t` with columns `_1`, `_2` and rows `("a", "A")`, `("x ",
"x")`, `("b", "c")`, `("Hello World", "hello")`, `("a,B,c", "b")`. With `a =
CAST(_1 AS STRING COLLATE UTF8_LCASE)` and `b = CAST(_2 AS STRING COLLATE
UTF8_LCASE)`:
| Expression | Row | Spark | Comet |
| --- | --- | --- | --- |
| `instr(a, b)` | `('a', 'A')` | 1 | 0 |
| `instr(a, b)` | `('a,B,c', 'b')` | 3 | 0 |
| `substring_index(a, b, 1)` | `('Hello World', 'hello')` | `''` | `'Hello
World'` |
| `trim(BOTH b FROM a)` | `('Hello World', 'hello')` | `' World'` | `'Hello
World'` |
| `greatest(a, b)` | `('Hello World', 'hello')` | `'Hello World'` |
`'hello'` |
| `least(a, b)` | `('a', 'A')` | `'a'` | `'A'` |
The plan shows `cast` dispatched and the string expression native. The
mismatches reproduce on Spark 4.0 and 4.1.
### Expected behavior
The results match Spark. Other collation-sensitive string expressions
already route collated input through the codegen dispatcher, for example
`contains`, `startswith`, `endswith`, `locate`, `find_in_set`, `replace` and
`levenshtein`.
### Additional context
`trim(a)` without a trim string, `upper`, `concat_ws`, `hash` and `xxhash64`
already match Spark on the same data.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]