andygrove opened a new issue, #6460:
URL: https://github.com/apache/datafusion-comet/issues/6460

   ### Describe the bug
   
   Spark 4.2 added `BinaryType` to the input types of `Reverse`: `reverse(b)` 
on a binary column reverses its bytes and returns binary. Spark 4.1 and earlier 
cast the binary to a string first. `CometReverse` sends every argument that is 
not an array to the native string `reverse`, so on Spark 4.2 a `reverse` of a 
binary column in a native plan fails the query:
   
   ```
   org.apache.comet.CometNativeException: Invalid argument error: Encountered 
non UTF-8 data: invalid utf-8 sequence of 1 bytes from index 0
   ```
   
   ### Steps to reproduce
   
   On `main` at 9f68a4144, built with `-Pspark-4.2`:
   
   ```sql
   CREATE TABLE t(b binary) USING parquet;
   INSERT INTO t VALUES (X'CAFE'), (X''), (NULL), (X'01');
   SELECT reverse(b) FROM t;
   ```
   
   ### Expected behavior
   
   Each value's bytes reversed, as Spark returns them: `X'FECA'`, `X''`, 
`NULL`, `X'01'`.
   
   ### Additional context
   
   Spark's own `DataFrameFunctionsSuite` test `reverse function - binary`, new 
in 4.2, hits this only once its DataFrame is cached in Comet's format (#5634). 
Its input is a local relation, which Comet does not otherwise run natively.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to