cloud-fan opened a new pull request, #57859:
URL: https://github.com/apache/spark/pull/57859

   <!--
   Thanks for sending a pull request!  Here are some tips for you:
     1. If this is your first time, please read our contributor guidelines: 
https://spark.apache.org/contributing.html
     2. Ensure you have added or run the appropriate tests for your PR: 
https://spark.apache.org/developer-tools.html
     3. If the PR is unfinished, add '[WIP]' in your PR title, e.g. 
'[WIP][SPARK-XXXX] Your PR title ...'.
     4. Be sure to keep the PR description updated to reflect all changes.
     5. Please write your PR title to summarize what this PR proposes.
     6. If possible, provide a concise example to reproduce the issue for a 
faster review.
     7. If you want to add a new configuration, please read the guideline first 
for naming configurations in
        
'core/src/main/scala/org/apache/spark/internal/config/ConfigEntry.scala'.
     8. If you want to add or modify an error type or message, please read the 
guideline first in
        'common/utils/src/main/resources/error/README.md'.
   -->
   
   ### What changes were proposed in this pull request?
   
   This is a follow-up to #57363. It extends `CombineAdjacentAggregation` for 
adjacent hash
   aggregates in two ways:
   
   - combine `PartialMerge` followed by `Final` into one `Final` aggregate;
   - preserve the lower aggregate's child distribution, streaming, and 
shuffle-partition metadata
     when removing that lower aggregate.
   
   The rule also requires the lower aggregate output to exactly match the final 
aggregate inputs.
   Filtered `Partial` aggregation remains supported, while a `PartialMerge` 
pair with filters is not
   combined.
   
   ### Why are the changes needed?
   
   `CombineAdjacentAggregation` currently handles only `Partial` followed by 
`Final`. A compatible
   adjacent `PartialMerge` and `Final` pair can also be collapsed safely, 
avoiding an unnecessary hash
   aggregation stage. When the lower aggregate is removed, its child-facing 
distribution and streaming
   metadata must move to the combined node because the combined node now reads 
that lower aggregate's
   original child.
   
   These semantics also move the upstream rule closer to covering the safe 
hash-aggregate behavior of
   downstream implementations, reducing the need for overlapping optimizer 
rules.
   
   ### Does this PR introduce _any_ user-facing change?
   
   Yes. Spark may produce a single final hash aggregate instead of adjacent 
partial-merge and final
   hash aggregates when their grouping, lineage, and input/output attributes 
are compatible. Query
   results are unchanged.
   
   ### How was this patch tested?
   
   Added coverage to `CombineAdjacentAggregationSuite` for:
   
   - combining `PartialMerge` and `Final` hash aggregates;
   - preserving distribution, streaming, shuffle, result, and buffer-offset 
metadata;
   - refusing to combine a filtered `PartialMerge` pair.
   
   Ran:
   
   ```
   build/sbt 'sql/testOnly 
org.apache.spark.sql.execution.CombineAdjacentAggregationSuite'
   build/sbt sql/scalastyle
   ```
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   Generated-by: Codex (GPT-5)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to