dwsmith1983 commented on code in PR #5867:
URL: https://github.com/apache/datafusion-comet/pull/5867#discussion_r4195242749
##########
spark/src/main/scala/org/apache/comet/serde/arrays.scala:
##########
@@ -925,10 +1013,16 @@ object CometArrayPosition extends
CometExpressionSerde[ArrayPosition] with Array
}
}
-object CometArraysZip extends CometExpressionSerde[ArraysZip] {
+object CometArraysZip extends CometExpressionSerde[ArraysZip] with
CodegenDispatchFallback {
Review Comment:
> Please key live kernel state by expression occurrence, while sharing
compiled code if desired, or retain Spark fallback until that isolation exists.
The shared counter is on `main` for any dispatched route: a Java UDF over
`monotonically_increasing_id()` in two columns returns `a = 0..7, b = 8..15`
per batch where Spark returns equal columns, since a Java UDF has no
per-occurrence encoders and both calls serialize identically. Filed as #6711
and fixed in #6712: a subtree that contains a `Nondeterministic` node is
wrapped in a marker carrying an id from a driver-wide counter before it is
serialized, and the executor strips the marker before compiling, so each
occurrence gets its own cache entry while deterministic subtrees keep sharing
compiled kernels. Your `arrays_zip` query is the same shape and is covered once
#6712 lands; this PR should go in after it, and I will add that query as a
fixture here then.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]