andygrove opened a new issue, #6463:
URL: https://github.com/apache/datafusion-comet/issues/6463

   ### Describe the bug
   
   Spark builds the plan tree behind the SQL tab's graph, and behind the plans 
that the event log records, with `SparkPlanInfo.fromSparkPlan`. It gives 
Spark's own cache scan the cached plan as a child:
   
   ```scala
   case inMemTab: InMemoryTableScanExec => inMemTab.relation.cachedPlan :: Nil
   ```
   
   `CometInMemoryTableScanExec` is a leaf, so for a relation cached in Comet's 
format the tree ends at the scan. The plan that built the cached relation, and 
its metrics, no longer appear under the scan in the SQL tab, nor in the event 
log that the history server and other tools read.
   
   ### Steps to reproduce
   
   On `main` at 9f68a4144 (Spark 4.2 profile), with `CometPlugin` and Comet's 
cache serializer:
   
   ```scala
   Seq("true", "false").foreach { enabled =>
     spark.conf.set("spark.comet.exec.inMemoryCache.enabled", enabled)
     val cached = spark.range(100).selectExpr("id % 10 AS k", s"'$enabled' AS 
tag")
       .groupBy("k", "tag").count().cache()
     cached.count()
     val q = cached.filter("k > 1")
     q.collect()
     // print the tree of 
SparkPlanInfo.fromSparkPlan(q.queryExecution.executedPlan)
   }
   ```
   
   With Comet's cache scan:
   
   ```
   AdaptiveSparkPlan
     ResultQueryStage
       WholeStageCodegen (1)
         CometColumnarToRow
           InputAdapter
             CometFilter
               TableCacheQueryStage
                 CometInMemoryTableScan
   ```
   
   With Spark's scan reading the same format:
   
   ```
   AdaptiveSparkPlan
     ResultQueryStage
       WholeStageCodegen (1)
         Filter
           InputAdapter
             TableCacheQueryStage
               InMemoryTableScan
                 AdaptiveSparkPlan
                   ResultQueryStage
                     RowToColumnar
                       WholeStageCodegen (2)
                         HashAggregate
                           InputAdapter
                             ShuffleQueryStage
                               Exchange
                                 WholeStageCodegen (1)
                                   HashAggregate
                                     Project
                                       Range
   ```
   
   ### Expected behavior
   
   The cached plan appears below `CometInMemoryTableScan`, as it does below 
Spark's scan.
   
   ### Additional context
   
   Comet probably cannot fix this on its side. `SparkPlanInfo` matches Spark's 
class, and `CometInMemoryTableScanExec` cannot expose the cached plan as a 
child without it becoming part of the query that reads the cache. A fix likely 
needs Spark to recognize other cache scans there.
   
   Spark's `CachedTableSuite` test `SPARK-35332: Make cache plan disable 
configs configurable - check AQE` (4.0 and later) reads the cached plan's final 
stage from this tree, so it fails with Comet's cache format. #5634 skips it 
under Comet.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to