szehon-ho commented on PR #17957: URL: https://github.com/apache/iceberg/pull/17957#issuecomment-5962904077
Could we add native `INSERT WITH SCHEMA EVOLUTION` integration tests? The current end-to-end coverage is for MERGE; the existing DataFrameWriter tests exercise legacy `merge-schema` behavior. With `write.spark.accept-any-schema=false`, cover: 1. Adding a column while casting a SMALLINT source into an existing INT column in the same write. Assert that the target stays INT, the new column is added, existing rows get null for it, and inserted values are correct. 2. Widening INT to BIGINT using a value above `Integer.MAX_VALUE`. Assert the resulting type and preservation of old and new values. 3. Writing `map<bigint,string>` into `map<int,string>` with keys that fit INT. Assert that the key type stays INT and the write succeeds with the expected values. Please parameterize these for by-name and by-position INSERT. Also, the accept-any-schema test currently only calls the hook twice; an actual write with native evolution and legacy `merge-schema` enabled would verify their interaction. Assert the final schema and rows. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
