Ignacio MQ created SPARK-60037:
----------------------------------

             Summary: Preserve source column order for columns added by MERGE 
INTO schema evolution
                 Key: SPARK-60037
                 URL: https://issues.apache.org/jira/browse/SPARK-60037
             Project: Spark
          Issue Type: Improvement
          Components: SQL
    Affects Versions: 4.1.1
            Reporter: Ignacio MQ


When using MERGE INTO ... WITH SCHEMA EVOLUTION (or the equivalent 
DataFrame.mergeInto(...).withSchemaEvolution() API), columns and nested struct 
fields present in the source but missing in the target are always appended at 
the end of the target schema, regardless of their position in the source.

This happens because MergeIntoTable.schemaChanges 
(sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/v2Commands.scala)
 emits TableChange.addColumn(fieldPath, dataType) without a ColumnPosition, so 
connectors that honor positions append them:

{code}
target: (id, a, b)
source: (id, a, new_col, b)
result: (id, a, b, new_col)    // appended at end, source order not preserved
{code}

Verified on Spark 4.1.1 and still the case on master (schemaChanges and 
ResolveSchemaEvolution emit positionless addColumn).

All the plumbing already exists: TableChange.addColumn accepts a ColumnPosition 
(first() / after(name), scoped to the parent struct so nested fields behave the 
same), and connectors already honor it, e.g. Apache Iceberg maps After to 
moveAfter and First to moveFirst in Spark3Util.applySchemaChanges, applied 
within the same UpdateSchema.

h3. Proposal

A flag such as spark.sql.mergeSchemaEvolution.preserveColumnOrder (default 
false, keeping current behavior). When enabled, schemaChanges derives each 
added field's position from its source position:

* new field at index 0 of its struct -> ColumnPosition.first()
* otherwise -> ColumnPosition.after(<previous source sibling>)

The diff already iterates source fields in order, so runs of consecutive new 
fields can anchor to the previous sibling even when that sibling is itself 
added by the same change list (changes are applied in order). Anchors should 
resolve through the existing toFieldMap / spark.sql.caseSensitiveAnalysis 
handling so they use the target's canonical field name. The same logic covers 
top-level columns, nested struct fields, array<struct> elements and map values, 
since fieldPath is already parent-scoped.

For consistency the same treatment could apply to ResolveSchemaEvolution (the 
INSERT ... WITH SCHEMA EVOLUTION / DSv2 write-evolution path).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to