Ben Hollis created SPARK-60067:
----------------------------------
Summary: UpdateFields expression size grows with struct width
Key: SPARK-60067
URL: https://issues.apache.org/jira/browse/SPARK-60067
Project: Spark
Issue Type: Bug
Components: SQL
Affects Versions: 4.4.0
Reporter: Ben Hollis
Spark represents `Column.withField` and `Column.dropFields` as the Catalyst
expression `UpdateFields`. Before execution, Spark replaces it with a new
struct containing one Catalyst expression for every output field, including
fields that were not changed. A one-field update to an N-field struct therefore
creates O(N) expression nodes before any row is evaluated. Wide and nested
structs produce large plans that increase optimizer, expression-binding,
code-generation, and executor memory costs.
*Example*
{code:java}
val updated = col("wide_struct").withField("field_0", lit(1))
df.select(updated) {code}
If `wide_struct` has 1,000 fields, this one-field update becomes a
`CreateNamedStruct` with roughly 1,000 field expressions. The query changes one
value, but Spark builds and processes a width-sized Catalyst tree before
execution. Multiple updates or uses repeat that representation and can produce
plans with tens of thousands of nodes.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]