Ben Hollis created SPARK-60067:
----------------------------------

             Summary: UpdateFields expression size grows with struct width
                 Key: SPARK-60067
                 URL: https://issues.apache.org/jira/browse/SPARK-60067
             Project: Spark
          Issue Type: Bug
          Components: SQL
    Affects Versions: 4.4.0
            Reporter: Ben Hollis


Spark represents `Column.withField` and `Column.dropFields` as the Catalyst 
expression `UpdateFields`. Before execution, Spark replaces it with a new 
struct containing one Catalyst expression for every output field, including 
fields that were not changed. A one-field update to an N-field struct therefore 
creates O(N) expression nodes before any row is evaluated. Wide and nested 
structs produce large plans that increase optimizer, expression-binding, 
code-generation, and executor memory costs.
*Example*
{code:java}
val updated = col("wide_struct").withField("field_0", lit(1))
df.select(updated) {code}
 
If `wide_struct` has 1,000 fields, this one-field update becomes a 
`CreateNamedStruct` with roughly 1,000 field expressions. The query changes one 
value, but Spark builds and processes a width-sized Catalyst tree before 
execution. Multiple updates or uses repeat that representation and can produce 
plans with tens of thousands of nodes.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to