sunchao commented on code in PR #5867:
URL: https://github.com/apache/datafusion-comet/pull/5867#discussion_r4194624076
##########
spark/src/main/scala/org/apache/comet/serde/arrays.scala:
##########
@@ -925,10 +1013,16 @@ object CometArrayPosition extends
CometExpressionSerde[ArrayPosition] with Array
}
}
-object CometArraysZip extends CometExpressionSerde[ArraysZip] {
+object CometArraysZip extends CometExpressionSerde[ArraysZip] with
CodegenDispatchFallback {
Review Comment:
[P2] Isolate mutable dispatcher kernels before enabling this route. On a
single-partition Parquet table with `x = 0..3` and `y = x + 100`, run `SELECT
arrays_zip(array(map(x, monotonically_increasing_id()))) AS a,
arrays_zip(array(map(y, monotonically_increasing_id()))) AS b FROM t`. Spark
produces map values `0,1,2,3` in both columns. The newly dispatched path
produces `4,5,6,7` in the second column. Binding each subtree independently
replaces both input columns with `BoundReference(0)`, producing identical
serialized expressions and vector specifications. `CometUdfBridge` shares one
dispatcher per task and class, and its cache therefore reuses the first
expression's mutable counter for the second. At the base, `arrays_zip` rejects
map elements and this projection falls back to Spark, so this PR exposes a new
wrong-result case. Please key live kernel state by expression occurrence, while
sharing compiled code if desired, or retain Spark fallback until that isolation
exists.
Evidence: Reproduced against freshly compiled head sources with Spark 4.1.3
in `/tmp/review5867-run-chucxrb0/DistinctInputsProbe.scala`; output is in
`DistinctInputsProbe.log`. The probe creates the Parquet input, executes the
Spark reference query, and serializes its optimized expressions through the
current planner. Assertions confirm identical closure bytes but distinct native
input references, so the two calls cannot be eliminated as identical native
expressions. Evaluating both through one current `CometScalaUDFCodegen`
instance yields second-column map values 4–7. Separate instances yield 0–3 in
both columns. The shared-instance setup follows `CometUdfBridge.java:198`, and
the colliding cache key is constructed at `CometScalaUDFCodegen.scala:118`.
This is a component reproduction; full native execution was not run.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]