dwsmith1983 commented on issue #6133: URL: https://github.com/apache/datafusion-comet/issues/6133#issuecomment-5856437388
The q64 loss looks like a different bug from the q5 failure. I filed it as #6264 with a small reproduction. With AQE on, Comet canonicalizes a scan without its pruning filter while that filter is still the adaptive placeholder. A parent exchange over the 2000 `store_sales` stage can then compare equal to the one over 1999 and be reused. The second `cross_sales` reference reads the other year's rows, and they are filtered out further up. The scan's canonical form is where the filter goes missing, so Spark operators over `CometNativeScan` are exposed too. It only happens when no coalesced shuffle read sits between the parent and the scan stage, which is normal for shuffles this large. Which stage AQE builds first changes between runs, which fits the intermittent results. Two checks on the SF1000 setup would confirm it: 1. Run q64 with `spark.sql.exchange.reuse=false`. If this is the cause, it should never return 0 rows. 2. In the final plan of a 0-row run, look for a `ReusedExchange` under the second `cross_sales` reference that points at the other year's broadcast of `store_sales ⋈ store_returns`, with only one of the two `store_sales` scans still carrying a pruning filter. This path needs AQE. If q64 still returns 0 rows with `spark.sql.adaptive.enabled=false`, something else is going on as well, and the plan from that run would help. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
