dwsmith1983 commented on code in PR #6268:
URL: https://github.com/apache/datafusion-comet/pull/6268#discussion_r4186652194


##########
spark/src/main/scala/org/apache/spark/sql/comet/operators.scala:
##########
@@ -1318,6 +1325,9 @@ abstract class CometLeafExec extends CometNativeExec with 
LeafExecNode {
  * parent's native execution receives an empty input. 
(`CometIcebergNativeScanExec` does NOT use
  * this trait; it has a dedicated `findAllPlanData` case.)
  *
+ * `perPartitionData.length` is the partition count the native block runs 
with, and the count in

Review Comment:
   > could this doc also say that `outputPartitioning` must not read 
`perPartitionData`, or anything else that runs the scan's DPP subqueries?
   
   Added in 2eb4ff65a. The `CometScanWithPlanData` doc now says 
`perPartitionData.length` is what sizes the native RDD at execution, that 
`outputPartitioning` must not read `perPartitionData` or anything that 
evaluates the scan's DPP subqueries because AQE calls it while optimizing a 
stage before `CometPlanAdaptiveDynamicPruningFilters` has converted the 
placeholders, and that a scan that cannot report a real partitioning without 
them should report `UnknownPartitioning(0)`, as `CometNativeScanExec` does for 
a non-bucketed scan.
   
   I pushed this while the label runs were still going so the doc change would 
not wait on them. It only touches that comment, so the code those runs are 
testing is unchanged.
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to