dwsmith1983 commented on code in PR #5867:
URL: https://github.com/apache/datafusion-comet/pull/5867#discussion_r4195242749


##########
spark/src/main/scala/org/apache/comet/serde/arrays.scala:
##########
@@ -925,10 +1013,16 @@ object CometArrayPosition extends 
CometExpressionSerde[ArrayPosition] with Array
   }
 }
 
-object CometArraysZip extends CometExpressionSerde[ArraysZip] {
+object CometArraysZip extends CometExpressionSerde[ArraysZip] with 
CodegenDispatchFallback {

Review Comment:
   > Please key live kernel state by expression occurrence, while sharing 
compiled code if desired, or retain Spark fallback until that isolation exists.
   
   The shared counter is on `main` for any dispatched route: a Java UDF over 
`monotonically_increasing_id()` in two columns returns `a = 0..7, b = 8..15` 
per batch where Spark returns equal columns, since a Java UDF has no 
per-occurrence encoders and both calls serialize identically. Filed as #6711 
and fixed in #6712: a subtree that contains a `Nondeterministic` node is 
wrapped in a marker carrying an id from a driver-wide counter before it is 
serialized, and the executor strips the marker before compiling, so each 
occurrence gets its own cache entry while deterministic subtrees keep sharing 
compiled kernels. Your `arrays_zip` query is the same shape and is covered once 
#6712 lands; this PR should go in after it, and I will add that query as a 
fixture here then.
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to