andygrove opened a new issue, #6460: URL: https://github.com/apache/datafusion-comet/issues/6460
### Describe the bug Spark 4.2 added `BinaryType` to the input types of `Reverse`: `reverse(b)` on a binary column reverses its bytes and returns binary. Spark 4.1 and earlier cast the binary to a string first. `CometReverse` sends every argument that is not an array to the native string `reverse`, so on Spark 4.2 a `reverse` of a binary column in a native plan fails the query: ``` org.apache.comet.CometNativeException: Invalid argument error: Encountered non UTF-8 data: invalid utf-8 sequence of 1 bytes from index 0 ``` ### Steps to reproduce On `main` at 9f68a4144, built with `-Pspark-4.2`: ```sql CREATE TABLE t(b binary) USING parquet; INSERT INTO t VALUES (X'CAFE'), (X''), (NULL), (X'01'); SELECT reverse(b) FROM t; ``` ### Expected behavior Each value's bytes reversed, as Spark returns them: `X'FECA'`, `X''`, `NULL`, `X'01'`. ### Additional context Spark's own `DataFrameFunctionsSuite` test `reverse function - binary`, new in 4.2, hits this only once its DataFrame is cached in Comet's format (#5634). Its input is a local relation, which Comet does not otherwise run natively. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
