dwsmith1983 commented on issue #6133:
URL: 
https://github.com/apache/datafusion-comet/issues/6133#issuecomment-5856437388

   The q64 loss looks like a different bug from the q5 failure. I filed it as 
#6264 with a small reproduction.
   
   With AQE on, Comet canonicalizes a scan without its pruning filter while 
that filter is still the adaptive placeholder. A parent exchange over the 2000 
`store_sales` stage can then compare equal to the one over 1999 and be reused. 
The second `cross_sales` reference reads the other year's rows, and they are 
filtered out further up. The scan's canonical form is where the filter goes 
missing, so Spark operators over `CometNativeScan` are exposed too. It only 
happens when no coalesced shuffle read sits between the parent and the scan 
stage, which is normal for shuffles this large. Which stage AQE builds first 
changes between runs, which fits the intermittent results.
   
   Two checks on the SF1000 setup would confirm it:
   
   1. Run q64 with `spark.sql.exchange.reuse=false`. If this is the cause, it 
should never return 0 rows.
   2. In the final plan of a 0-row run, look for a `ReusedExchange` under the 
second `cross_sales` reference that points at the other year's broadcast of 
`store_sales ⋈ store_returns`, with only one of the two `store_sales` scans 
still carrying a pruning filter.
   
   This path needs AQE. If q64 still returns 0 rows with 
`spark.sql.adaptive.enabled=false`, something else is going on as well, and the 
plan from that run would help.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to