voonhous opened a new pull request, #20126:
URL: https://github.com/apache/hudi/pull/20126

   ### Describe the issue this Pull Request addresses
   
   Validation of Spark 4.1+ `PushVariantIntoScan` on master, schema-evolution 
item 2 of the checklist on #18285 
(https://github.com/apache/hudi/issues/18285#issuecomment-5869224297): implicit 
widening of a top-level column beside a top-level variant.
   
   The #19783 leg uses a nested-only table, where a type change on one member 
of `s` turns the whole struct into an implicit type change and 
`SparkSchemaTransformUtils.addMissingFields` walks every member, including the 
projected variant. A top-level variant beside a widened top-level column takes 
another path: `buildImplicitSchemaChangeInfo` reconciles per column, `v` 
requested as the projection struct is declared equal to the file's 
`VariantType` and is never walked, and only `n` gets a type-change cast. 
Expected to pass, but nothing pinned it. Sibling PR for the MOR arm of the 
nested case: #20125.
   
   ### Summary and Changelog
   
   Test only. `TestVariantShreddingMixedLayouts` gets "Implicit widening of a 
top-level sibling keeps the variant projection":
   
   - table `(id int, v variant, n int, ts long)`: an int base file, a still-int 
SQL update, then one DataFrame upsert with `n` as bigint that widens the table 
and opens a second file group with a bigint base file. On COW the upsert 
carries new keys only, so the int file group stays int; on MOR it also carries 
id 2, so the int base file merges with an int log block and a bigint log block;
   - every slot read back through `variant_get`, one filter per slot, the 
widened column itself and a whole-value cast, with 
`spark.sql.variant.pushVariantIntoScan` on and off and the plan pinned per arm;
   - four legs: COW with both record types, MOR with native parquet log blocks 
(SPARK record type) and with avro data blocks (AVRO record type on table 
version 9, the DataFrame write pinning `hoodie.write.table.version` so 
auto-upgrade does not switch the format);
   - layout pinned: every parquet file shredded, exactly two file groups, and 
on MOR exactly two data blocks of the expected block type.
   
   Result: green on master as-is. No production change.
   
   ### Impact
   
   None.
   
   ### Risk Level
   
   none
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [x] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [x] Enough context is provided in the sections above
   - [x] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to