[
https://issues.apache.org/jira/browse/SPARK-59452?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59452:
-----------------------------------
Labels: pull-request-available (was: )
> DeduplicateRelations has quadratic overhead for wide plans
> ----------------------------------------------------------
>
> Key: SPARK-59452
> URL: https://issues.apache.org/jira/browse/SPARK-59452
> Project: Spark
> Issue Type: Improvement
> Components: SQL
> Affects Versions: 4.2.0
> Reporter: Ben Hollis
> Priority: Major
> Labels: pull-request-available
>
> Planning wide UNION and CTE-related queries incurs excessive CPU and
> allocation overhead in DeduplicateRelations.
> For a UNION, the rule compares every branch with every later branch to find
> conflicting expression IDs. A 500-branch UNION therefore performs
> approximately 124,750 branch comparisons and repeatedly rebuilds the
> remaining child sequence. This work grows quadratically with the number of
> branches.
> Related optimizer paths, including CTE inlining and join pushdown through
> UNION, construct synthetic self-joins solely to obtain fresh expression IDs
> for one side. Analyzing these temporary joins performs additional tree
> traversal and operator-resolution work that is discarded immediately
> afterward.
> The overhead is visible during query analysis and optimization, particularly
> for views and other plans containing hundreds of UNION branches or very wide
> CTE outputs. In focused measurements, an existing 500-branch UNION path
> allocated approximately 39 MB, while renewing a 500-column plan through a
> synthetic self-join allocated approximately 4.4 MB.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]