comphead opened a new pull request, #6700:
URL: https://github.com/apache/datafusion-comet/pull/6700

   ## Which issue does this PR close?
   
   Closes #6570.
   
   ## Rationale for this change
   
   The JVM codegen dispatcher initialized each kernel with 
`TaskContext.partitionId()`. Spark and Comet's native expressions use the index 
of the partition being computed instead. The two differ under a union, a 
coalesce and a cartesian product, so partition-seeded expressions such as 
`rand`, `uuid`, `monotonically_increasing_id` and `spark_partition_id` inside a 
dispatched expression returned different values from Spark. The dispatcher also 
kept one kernel per task, so under a coalesce every parent partition after the 
first reused the first one's kernel and its random state.
   
   Default settings reach this through common expressions. `round` on a double 
and `regexp_replace` dispatch their whole subtree, so `round(rand(42) * 100, 
2)` and `regexp_replace(uuid(), '-', '')` are affected. The issue has a repro.
   
   ## What changes are included in this PR?
   
   - The native planner passes its partition index and plan id 
(`exec_context_id`) to `JvmScalarUdfExpr`, which hands them to 
`CometUdfBridge.evaluate`.
   - `CometUDF` gets an `evaluate` overload that receives both. Its default 
calls the existing method, so other `CometUDF` implementations are unaffected.
   - `CometScalaUDFCodegen` initializes kernels with that partition index and 
caches nondeterministic kernels per plan. Deterministic kernels never read the 
index, so one of them still serves every plan in a task, and a coalesce over 
many partitions does not recompile them.
   
   ## How are these changes tested?
   
   A new test in `CometCodegenSuite` compares a dispatched `map(1, 
spark_partition_id())` and `round(rand(42), 6)` with Spark under a union, a 
coalesce and a cross join. Equivalent queries return different answers from 
Spark on `main`, see the repro in #6570.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to