andygrove opened a new issue, #6464:
URL: https://github.com/apache/datafusion-comet/issues/6464
### Describe the bug
In 1.1.0-rc1, `explode`, `posexplode` and `explode_outer` of an array of
structs with a boolean field return wrong booleans and misplaced NULLs once an
input batch explodes to more than `spark.comet.batchSize` rows. This happens
with default configs. 1.0.0 returns the right answer for the same queries.
#5362 split the native explode output into chunks of at most
`spark.comet.batchSize` rows, and #5667 slices each chunk out of the child
arrays instead of gathering it with `take`. A sliced struct keeps a bit offset
on its boolean children. Arrow Java ignores that offset when it imports the
batch (#6288), so the JVM reads those booleans from the start of the buffer. An
exploded top-level boolean goes wrong the same way when a native `named_struct`
wraps it or a boolean Scala UDF reads it.
### Steps to reproduce
```scala
withTempPath { dir =>
val path = dir.getCanonicalPath
withSQLConf(CometConf.COMET_ENABLED.key -> "false") {
spark
.range(0, 3000, 1, 1)
.selectExpr(
"id",
"transform(sequence(0, 11), i -> named_struct(" +
"'b', hash(id, i) % 2 = 0, " +
"'bn', IF(hash(id, i, 7) % 5 = 0, NULL, hash(id, i, 3) % 2 = 0), "
+
"'n', id * 12 + i)) AS arr")
.write
.parquet(path)
}
spark.read.parquet(path).createOrReplaceTempView("t")
checkSparkAnswer(sql("SELECT id, s FROM t LATERAL VIEW explode(arr) x AS
s"))
}
```
The native explode runs at both 1.0.0 and rc1. On rc1, 22,824 of the 36,000
rows differ from Spark, starting at output row 8,184, the first row of the
second chunk. `n` is right everywhere, and `b` and `bn` are wrong.
### Expected behavior
The same rows as Spark, as in 1.0.0.
### Additional context
#6339 fixes #6288 on `main` by zeroing boolean offsets at every level before
export, and its `branch-1.1` backport, #6449, fixes all of these queries on
rc1. This was found by the 1.1.0 regression audit in #6399 and is tracked in
#6402.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]