LinSimon-901101 commented on issue #6570: URL: https://github.com/apache/datafusion-comet/issues/6570#issuecomment-6000030629
Thanks @kazuyukitanimura. This is a **correctness issue: queries can complete successfully while silently returning incorrect results**. I have now reproduced the original `map(1, spark_partition_id())` example under both `UNION ALL` and `coalesce(1)`. Tested with Spark **4.1.3**, JDK **21.0.12**, and Comet **1.1.0-SNAPSHOT** at commit `714980151d382733c60d00dee58622d327e69676`, using `local[2]` with AQE disabled. Each input branch used `range(0, 8, 1, 2)`. I projected both `spark_partition_id()` and `map(1, spark_partition_id())` before applying union/coalesce, then compared Spark-only execution, Comet with dispatch enabled, and Comet with dispatch disabled. - **UNION ALL:** In the second branch, Spark and Comet’s direct native expression return partition IDs `0` and `1`, but the dispatched map contains `2` and `3`. **8 of 16 rows differ from Spark.** - **coalesce(1):** Rows from input partition `1` retain the correct direct/native partition ID `1`, but the dispatched map contains `0`. **4 of 8 rows differ from Spark.** - **Control:** Without union/coalesce, the dispatch-enabled results match Spark. The physical plans contain `CometUnion` / `CometCoalesce`, with the relevant projects annotated `JVM codegen dispatcher: map`. Runtime dispatcher counters also confirm that these kernels executed. This matches the reported initialization problem: the dispatcher uses `TaskContext.partitionId()` instead of the partition index of the child being computed. The demonstrated silent wrong results support `priority:critical` under the [[triage guide](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md)](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md). I also **verified the workaround**: setting `spark.comet.exec.scalaUDF.codegen.enabled=false` restores results matching Spark in both cases. This disables the shared dispatcher and may affect other dispatched expressions and performance. These results confirm the issue in the tested snapshot. [[Andy’s review](https://github.com/apache/datafusion-comet/pull/5875#pullrequestreview-5401526275)](https://github.com/apache/datafusion-comet/pull/5875#pullrequestreview-5401526275) also identifies it as an existing issue in `branch-1.1`, which is why it is tracked separately from #5875. Reproducer and supporting evidence are attached below, including execution plans, dispatcher counters, and Spark/Comet result comparisons. [issue-6570-evidence.zip](https://github.com/user-attachments/files/33070624/issue-6570-evidence.zip) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
