[ 
https://issues.apache.org/jira/browse/SPARK-60067?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SPARK-60067:
-----------------------------------
    Labels: pull-request-available  (was: )

> UpdateFields expression size grows with struct width
> ----------------------------------------------------
>
>                 Key: SPARK-60067
>                 URL: https://issues.apache.org/jira/browse/SPARK-60067
>             Project: Spark
>          Issue Type: Bug
>          Components: SQL
>    Affects Versions: 4.4.0
>            Reporter: Ben Hollis
>            Priority: Major
>              Labels: pull-request-available
>
> Spark represents `Column.withField` and `Column.dropFields` as the Catalyst 
> expression `UpdateFields`. Before execution, Spark replaces it with a new 
> struct containing one Catalyst expression for every output field, including 
> fields that were not changed. A one-field update to an N-field struct 
> therefore creates O(N) expression nodes before any row is evaluated. Wide and 
> nested structs produce large plans that increase optimizer, 
> expression-binding, code-generation, and executor memory costs.
> *Example*
> {code:java}
> val updated = col("wide_struct").withField("field_0", lit(1))
> df.select(updated) {code}
>  
> If `wide_struct` has 1,000 fields, this one-field update becomes a 
> `CreateNamedStruct` with roughly 1,000 field expressions. The query changes 
> one value, but Spark builds and processes a width-sized Catalyst tree before 
> execution. Multiple updates or uses repeat that representation and can 
> produce plans with tens of thousands of nodes.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to