andygrove opened a new issue, #6705:
URL: https://github.com/apache/datafusion-comet/issues/6705

   ### What is the problem the feature request solves?
   
   On every batch, `CometScalaUDFCodegen.evaluate` copies the serialized bound 
expression out of its first argument (`exprVec.get(0)`), about 7 KB for a 
simple Scala UDF, and hashes the copy into the kernel cache key. The 
`perf-cache-key` TODO on `CometScalaUDFCodegen.CacheKey` already describes this.
   
   @mbutrovich measured it on #6697 
([comment](https://github.com/apache/datafusion-comet/pull/6697#issuecomment-6006552028)).
 He found a fixed cost of about 8 microseconds per batch in the dispatcher, 
about 4 ms over 4M rows at batch size 8192, which grows as batches get smaller. 
About 6 microseconds of it is the difference between calling 
`CometScalaUDFCodegen.evaluate` and calling its generated kernel directly, 
which is where the copy and the hash happen.
   
   ### Describe the potential solution
   
   Compute a hash, or another identifier for the expression, on the driver and 
ship it through the `JvmScalarUdf` proto, so that a batch looks the kernel up 
without copying or hashing the closure. The TODO lists other options: 
per-instance memoization of the last key, or a two-tier cache keyed on the 
generated source.
   
   ### Additional context
   
   The measurements and the benchmark source are in the comment linked above.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to