szehon-ho commented on PR #17957:
URL: https://github.com/apache/iceberg/pull/17957#issuecomment-5962904077

   Could we add native `INSERT WITH SCHEMA EVOLUTION` integration tests? The 
current end-to-end coverage is for MERGE; the existing DataFrameWriter tests 
exercise legacy `merge-schema` behavior.
   
   With `write.spark.accept-any-schema=false`, cover:
   
   1. Adding a column while casting a SMALLINT source into an existing INT 
column in the same write. Assert that the target stays INT, the new column is 
added, existing rows get null for it, and inserted values are correct.
   2. Widening INT to BIGINT using a value above `Integer.MAX_VALUE`. Assert 
the resulting type and preservation of old and new values.
   3. Writing `map<bigint,string>` into `map<int,string>` with keys that fit 
INT. Assert that the key type stays INT and the write succeeds with the 
expected values.
   
   Please parameterize these for by-name and by-position INSERT. Also, the 
accept-any-schema test currently only calls the hook twice; an actual write 
with native evolution and legacy `merge-schema` enabled would verify their 
interaction. Assert the final schema and rows.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to