cloud-fan opened a new pull request, #57859:
URL: https://github.com/apache/spark/pull/57859
<!--
Thanks for sending a pull request! Here are some tips for you:
1. If this is your first time, please read our contributor guidelines:
https://spark.apache.org/contributing.html
2. Ensure you have added or run the appropriate tests for your PR:
https://spark.apache.org/developer-tools.html
3. If the PR is unfinished, add '[WIP]' in your PR title, e.g.
'[WIP][SPARK-XXXX] Your PR title ...'.
4. Be sure to keep the PR description updated to reflect all changes.
5. Please write your PR title to summarize what this PR proposes.
6. If possible, provide a concise example to reproduce the issue for a
faster review.
7. If you want to add a new configuration, please read the guideline first
for naming configurations in
'core/src/main/scala/org/apache/spark/internal/config/ConfigEntry.scala'.
8. If you want to add or modify an error type or message, please read the
guideline first in
'common/utils/src/main/resources/error/README.md'.
-->
### What changes were proposed in this pull request?
This is a follow-up to #57363. It extends `CombineAdjacentAggregation` for
adjacent hash
aggregates in two ways:
- combine `PartialMerge` followed by `Final` into one `Final` aggregate;
- preserve the lower aggregate's child distribution, streaming, and
shuffle-partition metadata
when removing that lower aggregate.
The rule also requires the lower aggregate output to exactly match the final
aggregate inputs.
Filtered `Partial` aggregation remains supported, while a `PartialMerge`
pair with filters is not
combined.
### Why are the changes needed?
`CombineAdjacentAggregation` currently handles only `Partial` followed by
`Final`. A compatible
adjacent `PartialMerge` and `Final` pair can also be collapsed safely,
avoiding an unnecessary hash
aggregation stage. When the lower aggregate is removed, its child-facing
distribution and streaming
metadata must move to the combined node because the combined node now reads
that lower aggregate's
original child.
These semantics also move the upstream rule closer to covering the safe
hash-aggregate behavior of
downstream implementations, reducing the need for overlapping optimizer
rules.
### Does this PR introduce _any_ user-facing change?
Yes. Spark may produce a single final hash aggregate instead of adjacent
partial-merge and final
hash aggregates when their grouping, lineage, and input/output attributes
are compatible. Query
results are unchanged.
### How was this patch tested?
Added coverage to `CombineAdjacentAggregationSuite` for:
- combining `PartialMerge` and `Final` hash aggregates;
- preserving distribution, streaming, shuffle, result, and buffer-offset
metadata;
- refusing to combine a filtered `PartialMerge` pair.
Ran:
```
build/sbt 'sql/testOnly
org.apache.spark.sql.execution.CombineAdjacentAggregationSuite'
build/sbt sql/scalastyle
```
### Was this patch authored or co-authored using generative AI tooling?
Generated-by: Codex (GPT-5)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]