andygrove commented on PR #4587:
URL: 
https://github.com/apache/datafusion-comet/pull/4587#issuecomment-5876300625

   This is a light fully automated review since there are so many PRs open.
   
   `CometExistenceJoinBenchmark.scala:38` overrides `getSparkSession` with a 
copy of the join benchmarks' session setup from before #6195, and that change 
is now in this branch through the merge from main. #6195 made `isCometLoaded` 
(`CometSparkSessionExtensions.scala:139`) return false when 
`spark.memory.offHeap.enabled` is off and `spark.comet.exec.onHeap.enabled` is 
unset. It also added `.set("spark.comet.exec.onHeap.enabled", "true")` to 
`CometBenchmarkBase.getSparkSession` and to the overrides in 
`CometHashJoinBenchmark`, `CometBroadcastHashJoinBenchmark` and 
`CometSortMergeJoinBenchmark`. This override sets neither, and nothing on the 
`make benchmark-%` path sets them either. So `CometExecRule` and 
`CometScanRule` return the plan unchanged, and both the BHJ and SHJ cases end 
up timing Spark against Spark. The only hint is the "NOT fully Comet native" 
warning from `runExpressionBenchmark`. Could you add the same 
`spark.comet.exec.onHeap.enabled` line to the conf here, so the res
 ults file measures the native existence join?
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to