LinSimon-901101 commented on issue #6570:
URL: 
https://github.com/apache/datafusion-comet/issues/6570#issuecomment-6000030629

   Thanks @kazuyukitanimura. This is a **correctness issue: queries can 
complete successfully while silently returning incorrect results**. I have now 
reproduced the original `map(1, spark_partition_id())` example under both 
`UNION ALL` and `coalesce(1)`.
   
   Tested with Spark **4.1.3**, JDK **21.0.12**, and Comet **1.1.0-SNAPSHOT** 
at commit `714980151d382733c60d00dee58622d327e69676`, using `local[2]` with AQE 
disabled.
   
   Each input branch used `range(0, 8, 1, 2)`. I projected both 
`spark_partition_id()` and `map(1, spark_partition_id())` before applying 
union/coalesce, then compared Spark-only execution, Comet with dispatch 
enabled, and Comet with dispatch disabled.
   
   - **UNION ALL:** In the second branch, Spark and Comet’s direct native 
expression return partition IDs `0` and `1`, but the dispatched map contains 
`2` and `3`. **8 of 16 rows differ from Spark.**
   - **coalesce(1):** Rows from input partition `1` retain the correct 
direct/native partition ID `1`, but the dispatched map contains `0`. **4 of 8 
rows differ from Spark.**
   - **Control:** Without union/coalesce, the dispatch-enabled results match 
Spark.
   
   The physical plans contain `CometUnion` / `CometCoalesce`, with the relevant 
projects annotated `JVM codegen dispatcher: map`. Runtime dispatcher counters 
also confirm that these kernels executed.
   
   This matches the reported initialization problem: the dispatcher uses 
`TaskContext.partitionId()` instead of the partition index of the child being 
computed. The demonstrated silent wrong results support `priority:critical` 
under the [[triage 
guide](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md)](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md).
   
   I also **verified the workaround**: setting 
`spark.comet.exec.scalaUDF.codegen.enabled=false` restores results matching 
Spark in both cases. This disables the shared dispatcher and may affect other 
dispatched expressions and performance.
   
   These results confirm the issue in the tested snapshot. [[Andy’s 
review](https://github.com/apache/datafusion-comet/pull/5875#pullrequestreview-5401526275)](https://github.com/apache/datafusion-comet/pull/5875#pullrequestreview-5401526275)
 also identifies it as an existing issue in `branch-1.1`, which is why it is 
tracked separately from #5875.
   
   Reproducer and supporting evidence are attached below, including execution 
plans, dispatcher counters, and Spark/Comet result comparisons.
   
[issue-6570-evidence.zip](https://github.com/user-attachments/files/33070624/issue-6570-evidence.zip)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to