andygrove opened a new issue, #6481:
URL: https://github.com/apache/datafusion-comet/issues/6481
### Describe the bug
`corr`, `covar_pop`, `covar_samp`, `var_pop`, `var_samp`, `stddev_pop` and
`stddev_samp` go wrong under the same two conditions as the `regr_*` aggregates
in #6423:
- a group's variable is constant at a value that binary floating point can't
represent exactly, such as 0.1
- the group's rows are merged from two or more partial aggregates
The constant's variance and covariance come out around 1e-35 and 1e-17
instead of exactly 0, so Comet returns tiny non-zero values where Spark returns
0.0. `corr` returns 0.878 where Spark returns NULL with ANSI off, or fails with
`DIVIDE_BY_ZERO` with ANSI on, the default on Spark 4.x. Comet never raises
that error.
Unlike `regr_*`, these aggregates already ran natively with the same merge
in 1.0.0, so this is not a regression in 1.1.0. It reproduces on 1.0.0,
1.1.0-rc1 and `main` (9c7fcc5aa4).
### Steps to reproduce
```scala
// In a suite extending CometTestBase
spark.range(0, 6, 1, 2)
.selectExpr("CAST(id AS DOUBLE) AS y", "0.1D AS x")
.write
.parquet(path) // two files, so two scan partitions
withParquetTable(path, "t") {
withSQLConf(SQLConf.ANSI_ENABLED.key -> "false") {
checkSparkAnswer(
"SELECT corr(y, x), covar_pop(y, x), covar_samp(y, x), var_pop(x),
var_samp(x), " +
"stddev_pop(x), stddev_samp(x) FROM t")
}
}
```
### Expected behavior
Spark returns `NULL, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0`.
Comet, with `CometHashAggregate` in both the partial and the final stage,
returns `0.8783, 1.04e-17, 1.25e-17, 4.81e-35, 5.78e-35, 6.94e-18, 7.60e-18`
(Spark 4.1 profile). 1.0.0, 1.1.0-rc1 and `main` return bit-identical values.
The sign of `corr` follows the order in which the partial buffers reach the
final aggregate.
### Additional context
The cause is the merge in `native/spark-expr/src/agg_funcs/welford.rs`
(`variance_merge`, `covariance_merge`), as described in #6423. The first merge
into the zero-initialized final buffer turns the partial mean 0.1 into
0.10000000000000002, and the second merge squares that error into `m2`. So
Spark's exact zero checks never fire.
#6451 makes the `regr_*` aggregates fall back for 1.1.0, but leaves these
native, since they already shipped this way. #6076 ports Spark's merge order,
which should fix all of them.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]