andygrove opened a new issue, #6705: URL: https://github.com/apache/datafusion-comet/issues/6705
### What is the problem the feature request solves? On every batch, `CometScalaUDFCodegen.evaluate` copies the serialized bound expression out of its first argument (`exprVec.get(0)`), about 7 KB for a simple Scala UDF, and hashes the copy into the kernel cache key. The `perf-cache-key` TODO on `CometScalaUDFCodegen.CacheKey` already describes this. @mbutrovich measured it on #6697 ([comment](https://github.com/apache/datafusion-comet/pull/6697#issuecomment-6006552028)). He found a fixed cost of about 8 microseconds per batch in the dispatcher, about 4 ms over 4M rows at batch size 8192, which grows as batches get smaller. About 6 microseconds of it is the difference between calling `CometScalaUDFCodegen.evaluate` and calling its generated kernel directly, which is where the copy and the hash happen. ### Describe the potential solution Compute a hash, or another identifier for the expression, on the driver and ship it through the `JvmScalarUdf` proto, so that a batch looks the kernel up without copying or hashing the closure. The TODO lists other options: per-instance memoization of the last key, or a two-tier cache keyed on the generated source. ### Additional context The measurements and the benchmark source are in the comment linked above. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
