LinSimon-901101 opened a new pull request, #6814: URL: https://github.com/apache/datafusion-comet/pull/6814
## Which issue does this PR close? Closes #6667. ## Rationale for this change On Spark 3.4, a single-split Iceberg scan can run a global aggregate without a shuffle. When dynamic partition pruning removes the only split, Comet drops Spark's empty execution partition, preventing the aggregate from running. As a result, `COUNT(*)` and `SUM(amount)` return no rows instead of `(0, NULL)`. ## What changes are included in this PR? - Preserve Spark's empty execution partitions by serializing an empty per-partition Iceberg scan. - Return an empty native stream with the expected schema before initializing storage or requesting catalog credentials when there are no file scan tasks. - Add regression tests for the aggregate result with AQE enabled and disabled, and for an empty scan with an unavailable S3 credential-policy provider. ## How are these changes tested? - Reproduced the original failure on Spark 3.4 with both AQE settings before the fix. - DPP tests passed locally on Spark 3.4 and 3.5, including the new aggregate and credential-provider regressions. - [[Full fork CI passed](https://github.com/LinSimon-901101/datafusion-comet/actions/runs/37885859935)] for commit `c9a808c`, including the Spark SQL and Iceberg test matrices. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
