LinSimon-901101 opened a new pull request, #6814:
URL: https://github.com/apache/datafusion-comet/pull/6814

   ## Which issue does this PR close?
   
   Closes #6667.
   
   ## Rationale for this change
   
   On Spark 3.4, a single-split Iceberg scan can run a global aggregate without 
a shuffle. When dynamic partition pruning removes the only split, Comet drops 
Spark's empty execution partition, preventing the aggregate from running. As a 
result, `COUNT(*)` and `SUM(amount)` return no rows instead of `(0, NULL)`.
   
   ## What changes are included in this PR?
   
   - Preserve Spark's empty execution partitions by serializing an empty 
per-partition Iceberg scan.
   - Return an empty native stream with the expected schema before initializing 
storage or requesting catalog credentials when there are no file scan tasks.
   - Add regression tests for the aggregate result with AQE enabled and 
disabled, and for an empty scan with an unavailable S3 credential-policy 
provider.
   
   ## How are these changes tested?
   
   - Reproduced the original failure on Spark 3.4 with both AQE settings before 
the fix.
   - DPP tests passed locally on Spark 3.4 and 3.5, including the new aggregate 
and credential-provider regressions.
   - [[Full fork CI 
passed](https://github.com/LinSimon-901101/datafusion-comet/actions/runs/37885859935)]
 for commit `c9a808c`, including the Spark SQL and Iceberg test matrices.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to