sunchao commented on code in PR #5867:
URL: https://github.com/apache/datafusion-comet/pull/5867#discussion_r4194624076


##########
spark/src/main/scala/org/apache/comet/serde/arrays.scala:
##########
@@ -925,10 +1013,16 @@ object CometArrayPosition extends 
CometExpressionSerde[ArrayPosition] with Array
   }
 }
 
-object CometArraysZip extends CometExpressionSerde[ArraysZip] {
+object CometArraysZip extends CometExpressionSerde[ArraysZip] with 
CodegenDispatchFallback {

Review Comment:
   [P2] Isolate mutable dispatcher kernels before enabling this route. On a 
single-partition Parquet table with `x = 0..3` and `y = x + 100`, run `SELECT 
arrays_zip(array(map(x, monotonically_increasing_id()))) AS a, 
arrays_zip(array(map(y, monotonically_increasing_id()))) AS b FROM t`. Spark 
produces map values `0,1,2,3` in both columns. The newly dispatched path 
produces `4,5,6,7` in the second column. Binding each subtree independently 
replaces both input columns with `BoundReference(0)`, producing identical 
serialized expressions and vector specifications. `CometUdfBridge` shares one 
dispatcher per task and class, and its cache therefore reuses the first 
expression's mutable counter for the second. At the base, `arrays_zip` rejects 
map elements and this projection falls back to Spark, so this PR exposes a new 
wrong-result case. Please key live kernel state by expression occurrence, while 
sharing compiled code if desired, or retain Spark fallback until that isolation 
exists.
   
   Evidence: Reproduced against freshly compiled head sources with Spark 4.1.3 
in `/tmp/review5867-run-chucxrb0/DistinctInputsProbe.scala`; output is in 
`DistinctInputsProbe.log`. The probe creates the Parquet input, executes the 
Spark reference query, and serializes its optimized expressions through the 
current planner. Assertions confirm identical closure bytes but distinct native 
input references, so the two calls cannot be eliminated as identical native 
expressions. Evaluating both through one current `CometScalaUDFCodegen` 
instance yields second-column map values 4–7. Separate instances yield 0–3 in 
both columns. The shared-instance setup follows `CometUdfBridge.java:198`, and 
the colliding cache key is constructed at `CometScalaUDFCodegen.scala:118`. 
This is a component reproduction; full native execution was not run.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to