Ignacio MQ created SPARK-60037:
----------------------------------
Summary: Preserve source column order for columns added by MERGE
INTO schema evolution
Key: SPARK-60037
URL: https://issues.apache.org/jira/browse/SPARK-60037
Project: Spark
Issue Type: Improvement
Components: SQL
Affects Versions: 4.1.1
Reporter: Ignacio MQ
When using MERGE INTO ... WITH SCHEMA EVOLUTION (or the equivalent
DataFrame.mergeInto(...).withSchemaEvolution() API), columns and nested struct
fields present in the source but missing in the target are always appended at
the end of the target schema, regardless of their position in the source.
This happens because MergeIntoTable.schemaChanges
(sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/v2Commands.scala)
emits TableChange.addColumn(fieldPath, dataType) without a ColumnPosition, so
connectors that honor positions append them:
{code}
target: (id, a, b)
source: (id, a, new_col, b)
result: (id, a, b, new_col) // appended at end, source order not preserved
{code}
Verified on Spark 4.1.1 and still the case on master (schemaChanges and
ResolveSchemaEvolution emit positionless addColumn).
All the plumbing already exists: TableChange.addColumn accepts a ColumnPosition
(first() / after(name), scoped to the parent struct so nested fields behave the
same), and connectors already honor it, e.g. Apache Iceberg maps After to
moveAfter and First to moveFirst in Spark3Util.applySchemaChanges, applied
within the same UpdateSchema.
h3. Proposal
A flag such as spark.sql.mergeSchemaEvolution.preserveColumnOrder (default
false, keeping current behavior). When enabled, schemaChanges derives each
added field's position from its source position:
* new field at index 0 of its struct -> ColumnPosition.first()
* otherwise -> ColumnPosition.after(<previous source sibling>)
The diff already iterates source fields in order, so runs of consecutive new
fields can anchor to the previous sibling even when that sibling is itself
added by the same change list (changes are applied in order). Anchors should
resolve through the existing toFieldMap / spark.sql.caseSensitiveAnalysis
handling so they use the target's canonical field name. The same logic covers
top-level columns, nested struct fields, array<struct> elements and map values,
since fieldPath is already parent-scoped.
For consistency the same treatment could apply to ResolveSchemaEvolution (the
INSERT ... WITH SCHEMA EVOLUTION / DSv2 write-evolution path).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]